responsible-request 0.4.0__tar.gz
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- responsible_request-0.4.0/LICENSE +21 -0
- responsible_request-0.4.0/PKG-INFO +319 -0
- responsible_request-0.4.0/README.md +285 -0
- responsible_request-0.4.0/pyproject.toml +89 -0
- responsible_request-0.4.0/pyproject.toml.orig +69 -0
- responsible_request-0.4.0/src/responsible_request/__init__.py +94 -0
- responsible_request-0.4.0/src/responsible_request/_http.py +31 -0
- responsible_request-0.4.0/src/responsible_request/analysis.py +81 -0
- responsible_request-0.4.0/src/responsible_request/cache.py +59 -0
- responsible_request-0.4.0/src/responsible_request/client.py +169 -0
- responsible_request-0.4.0/src/responsible_request/config.py +266 -0
- responsible_request-0.4.0/src/responsible_request/controller.py +79 -0
- responsible_request-0.4.0/src/responsible_request/cost.py +96 -0
- responsible_request-0.4.0/src/responsible_request/estimator.py +113 -0
- responsible_request-0.4.0/src/responsible_request/helpers/__init__.py +21 -0
- responsible_request-0.4.0/src/responsible_request/helpers/batch.py +127 -0
- responsible_request-0.4.0/src/responsible_request/helpers/messages.py +18 -0
- responsible_request-0.4.0/src/responsible_request/helpers/params.py +18 -0
- responsible_request-0.4.0/src/responsible_request/helpers/structured.py +150 -0
- responsible_request-0.4.0/src/responsible_request/limiter.py +165 -0
- responsible_request-0.4.0/src/responsible_request/logfiles.py +120 -0
- responsible_request-0.4.0/src/responsible_request/logging.py +183 -0
- responsible_request-0.4.0/src/responsible_request/providers.py +105 -0
- responsible_request-0.4.0/src/responsible_request/py.typed +0 -0
- responsible_request-0.4.0/src/responsible_request/records.py +265 -0
- responsible_request-0.4.0/src/responsible_request/runs.py +116 -0
- responsible_request-0.4.0/src/responsible_request/sources.py +207 -0
- responsible_request-0.4.0/src/responsible_request/throttle.py +229 -0
- responsible_request-0.4.0/src/responsible_request/transport.py +376 -0
|
@@ -0,0 +1,21 @@
|
|
|
1
|
+
MIT License
|
|
2
|
+
|
|
3
|
+
Copyright (c) 2026 Bern University of Applied Sciences (BFH)
|
|
4
|
+
|
|
5
|
+
Permission is hereby granted, free of charge, to any person obtaining a copy
|
|
6
|
+
of this software and associated documentation files (the "Software"), to deal
|
|
7
|
+
in the Software without restriction, including without limitation the rights
|
|
8
|
+
to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
|
|
9
|
+
copies of the Software, and to permit persons to whom the Software is
|
|
10
|
+
furnished to do so, subject to the following conditions:
|
|
11
|
+
|
|
12
|
+
The above copyright notice and this permission notice shall be included in all
|
|
13
|
+
copies or substantial portions of the Software.
|
|
14
|
+
|
|
15
|
+
THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
|
|
16
|
+
IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
|
|
17
|
+
FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
|
|
18
|
+
AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
|
|
19
|
+
LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
|
|
20
|
+
OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
|
|
21
|
+
SOFTWARE.
|
|
@@ -0,0 +1,319 @@
|
|
|
1
|
+
Metadata-Version: 2.4
|
|
2
|
+
Name: responsible-request
|
|
3
|
+
Version: 0.4.0
|
|
4
|
+
Summary: Polite, reproducible and cost-safe LLM requests for research: load-aware throttling, request logging, caching and budgets for OpenAI-compatible endpoints.
|
|
5
|
+
Keywords: llm,openai,litellm,openrouter,rate-limiting,throttling,logging,caching,reproducibility
|
|
6
|
+
Author: Luca Rolshoven
|
|
7
|
+
Author-email: Luca Rolshoven <luca@rolshoven.io>
|
|
8
|
+
License-Expression: MIT
|
|
9
|
+
License-File: LICENSE
|
|
10
|
+
Classifier: Development Status :: 3 - Alpha
|
|
11
|
+
Classifier: Framework :: AsyncIO
|
|
12
|
+
Classifier: Intended Audience :: Science/Research
|
|
13
|
+
Classifier: Programming Language :: Python :: 3
|
|
14
|
+
Classifier: Programming Language :: Python :: 3.10
|
|
15
|
+
Classifier: Programming Language :: Python :: 3.11
|
|
16
|
+
Classifier: Programming Language :: Python :: 3.12
|
|
17
|
+
Classifier: Programming Language :: Python :: 3.13
|
|
18
|
+
Classifier: Programming Language :: Python :: 3.14
|
|
19
|
+
Classifier: Typing :: Typed
|
|
20
|
+
Classifier: Topic :: Scientific/Engineering :: Artificial Intelligence
|
|
21
|
+
Requires-Dist: loguru>=0.7
|
|
22
|
+
Requires-Dist: openai>=1.40
|
|
23
|
+
Requires-Dist: pydantic>=2
|
|
24
|
+
Requires-Dist: pandas>=2 ; extra == 'pandas'
|
|
25
|
+
Requires-Dist: tqdm>=4 ; extra == 'progress'
|
|
26
|
+
Requires-Dist: zstandard>=0.22 ; python_full_version < '3.14' and extra == 'zstd'
|
|
27
|
+
Requires-Python: >=3.10
|
|
28
|
+
Project-URL: Repository, https://github.com/digital-sustainability/ResponsibleRequest
|
|
29
|
+
Project-URL: Issues, https://github.com/digital-sustainability/ResponsibleRequest/issues
|
|
30
|
+
Provides-Extra: pandas
|
|
31
|
+
Provides-Extra: progress
|
|
32
|
+
Provides-Extra: zstd
|
|
33
|
+
Description-Content-Type: text/markdown
|
|
34
|
+
|
|
35
|
+
# responsible-request
|
|
36
|
+
|
|
37
|
+
Responsible LLM requests for research, against any OpenAI-compatible endpoint (LiteLLM, OpenRouter, vLLM, Ollama, OpenAI, …). It works as a drop-in for `openai.AsyncOpenAI`.
|
|
38
|
+
|
|
39
|
+
- **Polite**: load-aware throttling, so large experiments don't starve the other users of a shared inference server.
|
|
40
|
+
- **Reproducible**: every request is logged (parameters, messages, responses, tokens, timing, provider, system fingerprint) to JSONL and/or SQLite, each client writes a run manifest (versions, git commit, configuration), and `rr.reproducible()` pins the decoding parameters.
|
|
41
|
+
- **Cost-safe**: a response cache means a restarted script only sends what is still missing, and per-request cost tracking with an optional budget stops spending before it gets out of hand, even across restarts.
|
|
42
|
+
|
|
43
|
+
## Load-aware throttling in a nutshell
|
|
44
|
+
|
|
45
|
+
Large experiments can starve the other users of a shared inference server. Hard RPM limits are the usual answer, but they leave the server idle at night and still hurt others at peak times. `responsible-request` instead **watches the latency of your own requests**:
|
|
46
|
+
|
|
47
|
+
- While latency stays close to its baseline, the server is idle, so the client ramps up to `max_rpm`.
|
|
48
|
+
- As soon as latency reaches 3× the baseline (or the server returns 429/5xx/timeouts), other users are active, so the client drops to `min_rpm` at once.
|
|
49
|
+
- Once latency is back to normal, the client ramps up again.
|
|
50
|
+
|
|
51
|
+
For commercial APIs with generous rate limits of their own (OpenAI, OpenRouter, …), load-aware throttling is not needed: use a plain fixed-rate limiter, `rr.ThrottleConfig.fixed(rpm)`. In fixed mode a 429 does not change the rate: it pauses only the affected model (and endpoint) for the time in the `Retry-After` header, or for an exponential backoff (`backoff_initial_s`, doubling up to `backoff_max_s`) if there is none, plus a random jitter of up to `backoff_jitter_s`. The adaptive throttle instead treats a 429 as a sign of other users and drops to `min_rpm` for at least `cooldown_s`.
|
|
52
|
+
|
|
53
|
+
## Installation
|
|
54
|
+
|
|
55
|
+
```bash
|
|
56
|
+
pip install git+https://github.com/digital-sustainability/ResponsibleRequest.git
|
|
57
|
+
# or: uv add git+https://github.com/digital-sustainability/ResponsibleRequest.git
|
|
58
|
+
pip install "responsible-request[pandas,progress] @ git+https://github.com/digital-sustainability/ResponsibleRequest.git"
|
|
59
|
+
```
|
|
60
|
+
|
|
61
|
+
Requires Python ≥ 3.10 and `openai` ≥ 1.40 (it works with both the `httpx`-based 1.x/2.x SDKs and the `httpx2`-based 3.x SDK).
|
|
62
|
+
|
|
63
|
+
## Quickstart
|
|
64
|
+
|
|
65
|
+
```python
|
|
66
|
+
import responsible_request as rr
|
|
67
|
+
|
|
68
|
+
client = rr.AsyncOpenAI( # same arguments as openai.AsyncOpenAI
|
|
69
|
+
base_url="https://inference.example.org/api/v1",
|
|
70
|
+
api_key="sk-...",
|
|
71
|
+
throttle=rr.ThrottleConfig(max_rpm=15, min_rpm=1),
|
|
72
|
+
log=rr.LogConfig(sqlite="requests.db"),
|
|
73
|
+
default_params=rr.reproducible(seed=42), # temperature=0, top_p=1, seed=42 unless set per call
|
|
74
|
+
)
|
|
75
|
+
|
|
76
|
+
response = await client.chat.completions.create(model="gpt-oss:120b", messages=[...])
|
|
77
|
+
print(client.throttle.stats())
|
|
78
|
+
await client.close()
|
|
79
|
+
```
|
|
80
|
+
|
|
81
|
+
As with the OpenAI SDK, `base_url` and `api_key` can come from `OPENAI_BASE_URL` / `OPENAI_API_KEY` instead. For the examples in this repo, copy `.env.example` to `.env` and run them with `uv run --env-file .env python examples/quickstart.py`.
|
|
82
|
+
|
|
83
|
+
For existing code, only the HTTP client has to change:
|
|
84
|
+
|
|
85
|
+
```python
|
|
86
|
+
import openai, responsible_request as rr
|
|
87
|
+
|
|
88
|
+
client = openai.AsyncOpenAI(http_client=rr.http_client(rr.ThrottleConfig(max_rpm=15)))
|
|
89
|
+
```
|
|
90
|
+
|
|
91
|
+
It also works as a plain fixed-rate limiter for everyday use: `rr.ThrottleConfig.fixed(60)`.
|
|
92
|
+
|
|
93
|
+
With OpenRouter (or any other commercial endpoint), a typical setup is:
|
|
94
|
+
|
|
95
|
+
```python
|
|
96
|
+
client = rr.AsyncOpenAI(
|
|
97
|
+
base_url="https://openrouter.ai/api/v1",
|
|
98
|
+
api_key=os.environ["OPENROUTER_API_KEY"],
|
|
99
|
+
throttle=rr.ThrottleConfig.fixed(120),
|
|
100
|
+
log=rr.LogConfig(sqlite="requests.db"),
|
|
101
|
+
cache=True, # a re-run only sends what is missing
|
|
102
|
+
cost=rr.CostConfig(budget_usd=5), # stop at $5, counting earlier runs in requests.db
|
|
103
|
+
run=rr.RunConfig(name="ablation-3"),
|
|
104
|
+
default_params=rr.reproducible(seed=42),
|
|
105
|
+
)
|
|
106
|
+
```
|
|
107
|
+
|
|
108
|
+
See `examples/openrouter.py`.
|
|
109
|
+
|
|
110
|
+
## Helpers
|
|
111
|
+
|
|
112
|
+
```python
|
|
113
|
+
# Many requests, paced by the throttle, results in input order (exceptions returned, not raised)
|
|
114
|
+
results = await rr.run_batch(client, [{"model": m, "messages": msgs} for msgs in dataset], progress=True)
|
|
115
|
+
results = rr.run_batch_sync(client, requests) # from a plain script
|
|
116
|
+
|
|
117
|
+
# Structured output validated with Pydantic; the validation error is fed back to the model on failure
|
|
118
|
+
class Answer(BaseModel):
|
|
119
|
+
label: str
|
|
120
|
+
confidence: float
|
|
121
|
+
|
|
122
|
+
answer = await rr.structured(client, model=m, messages=msgs, schema=Answer, retries=2)
|
|
123
|
+
# (if `content` is empty, the answer is read from `reasoning`, then `reasoning_content`)
|
|
124
|
+
rr.response_format_from_model(Answer) # just the strict json_schema response_format
|
|
125
|
+
|
|
126
|
+
# System prompt marked for OpenRouter's prompt cache (cache_control: ephemeral)
|
|
127
|
+
messages = [rr.cached_system_message(long_instructions), {"role": "user", "content": question}]
|
|
128
|
+
|
|
129
|
+
# Attach metadata to every request record created inside the block
|
|
130
|
+
with rr.tags(experiment="ablation-3", fold=2):
|
|
131
|
+
...
|
|
132
|
+
|
|
133
|
+
# Measure and pin the baseline explicitly, e.g. at night with a representative request
|
|
134
|
+
await rr.calibrate(client, "gpt-oss:120b", n=10, messages=representative_messages)
|
|
135
|
+
```
|
|
136
|
+
|
|
137
|
+
`rr.cached_system_message()` only asks the provider to cache the prompt prefix (cheaper input tokens, the model still runs), whereas `cache=True` replays complete stored responses from your SQLite or JSONL log without sending the request.
|
|
138
|
+
|
|
139
|
+
## How the throttle works
|
|
140
|
+
|
|
141
|
+
```
|
|
142
|
+
ratio >= high_ratio (3.0) or 429/5xx/timeout
|
|
143
|
+
┌────────────────────────────────────────────────────────────┐
|
|
144
|
+
│ ▼
|
|
145
|
+
WARMUP ──baseline known──► NORMAL ◄──at max_rpm── RECOVERING ◄── THROTTLED (min_rpm)
|
|
146
|
+
start_rpm ramp ×1.5 / 30 s ramp ×1.5 / 30 s cooldown ≥ 120 s and
|
|
147
|
+
up to max_rpm ratio < recover_ratio (1.5)
|
|
148
|
+
```
|
|
149
|
+
|
|
150
|
+
- **Pacing**: requests are spaced evenly at the current rate with no bursts, and `max_concurrency` caps the number in flight. One `Throttle` can be shared by several clients, also across threads that each run their own event loop (e.g. `asyncio.run` per thread), and then enforces one rate for all of them:
|
|
151
|
+
|
|
152
|
+
```python
|
|
153
|
+
throttle = rr.Throttle(rr.ThrottleConfig.fixed(60))
|
|
154
|
+
def worker(items): # runs in its own thread
|
|
155
|
+
asyncio.run(work(rr.AsyncOpenAI(throttle=throttle, ...), items))
|
|
156
|
+
```
|
|
157
|
+
- **Per model**: every (endpoint, model) pair has its own lane, because load on one backend says nothing about another. Only `/chat/completions` and `/embeddings` feed the latency estimate (`observe_paths`). Other endpoints (audio, …) are paced and react to errors only. Requests with server-side tools (MCP, web search) are ignored for load estimation, because their latency includes external calls.
|
|
158
|
+
- **Signal**: by default, latency per completion token (`latency / max(completion_tokens, 16)`), so long answers and reasoning traces don't look like load. For streamed requests the signal is the time to first byte. The current value is the median of the last `window` requests.
|
|
159
|
+
- **Baseline**: the 10th percentile of the signal over the last 30 minutes, needing `warmup_requests` samples first. Samples taken while throttled are excluded, so a long busy period does not become the new "normal". You can pin it with `ThrottleConfig(baseline=...)` or `rr.calibrate(...)`, and reset it with `client.throttle.reset_baseline()`.
|
|
160
|
+
- **Dead band**: between `recover_ratio` and `high_ratio` the rate stays where it is, which prevents oscillation.
|
|
161
|
+
|
|
162
|
+
`examples/simulate_load.py` runs the whole loop against a simulated server (time-compressed) where other users appear for a while:
|
|
163
|
+
|
|
164
|
+
```
|
|
165
|
+
t [s] others RPM state load
|
|
166
|
+
8.5 0 1800 normal 1.00
|
|
167
|
+
12.5 48 1800 normal 1.00
|
|
168
|
+
13.5 48 60 throttled 3.56
|
|
169
|
+
...
|
|
170
|
+
30.6 0 60 recovering 1.00
|
|
171
|
+
41.6 0 1800 normal 1.00
|
|
172
|
+
```
|
|
173
|
+
|
|
174
|
+
### Configuration (`ThrottleConfig`)
|
|
175
|
+
|
|
176
|
+
| Parameter | Default | Meaning |
|
|
177
|
+
|---|---|---|
|
|
178
|
+
| `max_rpm` | 15 | rate when the endpoint is idle |
|
|
179
|
+
| `min_rpm` | 1 | rate while others are active (also the probing rate) |
|
|
180
|
+
| `start_rpm` | 2 | rate while the baseline is being established |
|
|
181
|
+
| `max_concurrency` | 32 | max requests in flight per model |
|
|
182
|
+
| `high_ratio` | 3.0 | throttle when latency ≥ this × baseline |
|
|
183
|
+
| `recover_ratio` | 1.5 | ramp up only when latency < this × baseline |
|
|
184
|
+
| `cooldown_s` | 120 | minimum time at `min_rpm` after the last high-load signal |
|
|
185
|
+
| `ramp_factor` / `ramp_interval_s` | 1.5 / 30 | multiplicative ramp-up step and its interval |
|
|
186
|
+
| `metric` | `"latency_per_token"` | `"latency"`, `"ttfb"`, or a callable `RequestRecord -> float` |
|
|
187
|
+
| `window` | 10 | requests in the rolling median |
|
|
188
|
+
| `warmup_requests` | 20 | samples needed before the baseline is trusted |
|
|
189
|
+
| `baseline` | None | pin the baseline instead of estimating it |
|
|
190
|
+
| `baseline_percentile` / `baseline_window_s` | 10 / 1800 | baseline estimator |
|
|
191
|
+
| `observe_paths` | chat, embeddings | endpoints whose latency is used |
|
|
192
|
+
| `estimator_factory` | None | plug in your own `LoadEstimator` (e.g. reading queue depth from Prometheus) |
|
|
193
|
+
| `backoff_initial_s` / `backoff_max_s` | 1 / 60 | fixed mode: pause after a 429 without `Retry-After`, doubling per further 429 |
|
|
194
|
+
| `backoff_jitter_s` | 1 | fixed mode: random extra time added to every 429 pause |
|
|
195
|
+
|
|
196
|
+
## Logging
|
|
197
|
+
|
|
198
|
+
The package logs through [loguru](https://github.com/Delgan/loguru) and, as loguru recommends for libraries, stays silent until a client is created with `log=` (the default `log=True` enables console messages only).
|
|
199
|
+
|
|
200
|
+
- **Console**: rate changes (INFO), switches to throttling (WARNING), and a per-model summary every `summary_interval_s`. Individual requests are never printed. With `LogConfig(console_level="INFO")` the package replaces loguru's default DEBUG handler with a quieter one.
|
|
201
|
+
- **Request records**: one record per HTTP attempt (SDK retries are separate records with `attempt` > 0) goes to `LogConfig(jsonl=...)` and/or `LogConfig(sqlite=...)` (table `requests`). Records are emitted at loguru's TRACE level with `extra["rr_record"]`, so you can also attach your own sinks.
|
|
202
|
+
|
|
203
|
+
| Group | Fields |
|
|
204
|
+
|---|---|
|
|
205
|
+
| core (always) | `request_id`, `timestamp`, `method`, `path`, `model`, `stream`, `attempt`, `status_code`, `error`, `tags`, `cache_key`, `cache_hit`, `run_id` |
|
|
206
|
+
| `timing` | `sent_at`, `first_byte_at`, `finished_at`, `wait_s` (time spent throttled), `ttfb_s`, `latency_s` |
|
|
207
|
+
| `usage` | `prompt_tokens`, `completion_tokens`, `total_tokens`, `cached_tokens`, `reasoning_tokens`, `cost_usd` |
|
|
208
|
+
| `throttle` | `rpm`, `state`, `baseline`, `load_ratio`, `in_flight` |
|
|
209
|
+
| `response_meta` | `response_id`, `response_model`, `finish_reason`, `system_fingerprint`, `provider`, `gateway_request_id`, `upstream_duration_ms`, `provider_meta` (see below) |
|
|
210
|
+
| `params` | all request parameters except messages/input |
|
|
211
|
+
| `request_body` / `response_body` | full request and response (for streams: the concatenated content) |
|
|
212
|
+
|
|
213
|
+
All groups are logged by default. Turn groups off with `LogConfig(fields={"request_body": False})`, or use `LogConfig.minimal()` (no params or bodies). **Request and response bodies can be large and may contain sensitive data.**
|
|
214
|
+
|
|
215
|
+
```python
|
|
216
|
+
df = rr.load_records("requests.db") # pandas DataFrame if pandas is installed, else list of dicts
|
|
217
|
+
```
|
|
218
|
+
|
|
219
|
+
JSONL logs rotate at `rotation="500 MB"` (any loguru rotation). LLM logs are repetitive and compress well, so finished segments can be compressed with `compression="gz"`, `"bz2"`, `"xz"` or `"zst"` (zstd needs Python 3.14+ or `pip install responsible-request[zstd]`); the file being written stays plain. Without rotation, the file is compressed when the client's sinks are removed. `load_records("requests.jsonl")`, the cache and the cost budget read all segments, compressed or not, and `load_records` / `CacheConfig(jsonl=...)` also accept a single archive such as `requests.2026-10-06_12-00-00_000000.jsonl.zst`.
|
|
220
|
+
|
|
221
|
+
```python
|
|
222
|
+
log = rr.LogConfig(jsonl="requests.jsonl", rotation="100 MB", compression="zst")
|
|
223
|
+
```
|
|
224
|
+
|
|
225
|
+
### Provider metadata
|
|
226
|
+
|
|
227
|
+
Gateways and routers report different extra information. *Extractors* map it to the generic `response_meta` fields and put everything else into the `provider_meta` JSON column. All built-in extractors run on every response and only fill in what they find, so nothing has to be configured:
|
|
228
|
+
|
|
229
|
+
| Extractor | Reads |
|
|
230
|
+
|---|---|
|
|
231
|
+
| `openai_compat` | `x-request-id` / `request-id` → `gateway_request_id`, `openai-processing-ms` → `upstream_duration_ms` |
|
|
232
|
+
| `litellm` | `x-litellm-call-id` → `gateway_request_id`, `x-litellm-response-duration-ms` → `upstream_duration_ms`, `x-litellm-response-cost` → `cost_usd`, all other `x-litellm-*` headers → `provider_meta["litellm"]` |
|
|
233
|
+
| `openrouter` | `provider` → `provider`, `usage.cost` → `cost_usd`, `usage.is_byok`/`cost_details` and `native_finish_reason` → `provider_meta["openrouter"]` |
|
|
234
|
+
|
|
235
|
+
Add your own for other endpoints:
|
|
236
|
+
|
|
237
|
+
```python
|
|
238
|
+
def my_gateway(record, headers, body): # body: parsed JSON, stream summary, or None
|
|
239
|
+
if "x-queue-ms" in headers:
|
|
240
|
+
rr.providers.meta(record, "my_gateway")["queue_ms"] = float(headers["x-queue-ms"])
|
|
241
|
+
|
|
242
|
+
client = rr.AsyncOpenAI(..., extractors=[my_gateway])
|
|
243
|
+
```
|
|
244
|
+
|
|
245
|
+
Databases written by version 0.1 keep their `litellm_*` columns; new records leave them empty.
|
|
246
|
+
|
|
247
|
+
### Runs
|
|
248
|
+
|
|
249
|
+
Every client is a *run* with its own `run_id`, which every record carries. With a JSONL or SQLite log, the first request also writes a run manifest: the run name and `metadata`, start time, endpoint (scheme and host only, never the API key), versions of `responsible-request`, `openai` and Python, platform, hostname, command line, working directory, git commit and dirty flag, and the complete throttle, log, cache and cost configuration plus `default_params`.
|
|
250
|
+
|
|
251
|
+
```python
|
|
252
|
+
client = rr.AsyncOpenAI(..., run=rr.RunConfig(name="ablation-3", metadata={"dataset": "v2"}))
|
|
253
|
+
# or simply run="ablation-3"
|
|
254
|
+
|
|
255
|
+
runs = rr.load_runs("requests.db") # the `runs` table, or requests.runs.jsonl for a JSONL log
|
|
256
|
+
df = rr.load_records("requests.db").merge(runs, on="run_id", suffixes=("", "_run"))
|
|
257
|
+
```
|
|
258
|
+
|
|
259
|
+
## Caching
|
|
260
|
+
|
|
261
|
+
With `cache=True`, a request that was already answered successfully is served from the logged records instead of being sent again. Re-running a script (e.g. after a crash) then only sends the requests that are still missing:
|
|
262
|
+
|
|
263
|
+
```python
|
|
264
|
+
client = rr.AsyncOpenAI(
|
|
265
|
+
log=rr.LogConfig(sqlite="requests.db"),
|
|
266
|
+
cache=True, # or rr.CacheConfig(...)
|
|
267
|
+
)
|
|
268
|
+
```
|
|
269
|
+
|
|
270
|
+
- **Key**: a hash of the URL and the complete JSON request body (model, messages and all parameters, after `default_params` are applied). Any change to the prompt or parameters is a miss.
|
|
271
|
+
- **Source**: the records the client writes itself (`LogConfig.sqlite`, else `LogConfig.jsonl`; `response_body` must be logged), or another database or file via `rr.CacheConfig(sqlite=...)` / `rr.CacheConfig(jsonl=...)`.
|
|
272
|
+
- **What is served**: only non-streamed requests whose record has status 200 and no error. Streamed requests are always sent.
|
|
273
|
+
- **Hits** don't wait for or affect the throttle. They are logged with `cache_hit=True` (filter them out when analysing latency or token usage), and the response carries an `x-rr-cache: hit` header. `client.cache.stats()` counts hits and misses.
|
|
274
|
+
- **Repeated sampling**: identical requests share one cached answer. To draw several independent samples, name the tags that belong to the key:
|
|
275
|
+
|
|
276
|
+
```python
|
|
277
|
+
client = rr.AsyncOpenAI(log=..., cache=rr.CacheConfig(key_tags=("sample",)))
|
|
278
|
+
for i in range(5):
|
|
279
|
+
with rr.tags(sample=i): # 5 different keys; a re-run gets the same 5 answers
|
|
280
|
+
await client.chat.completions.create(model=m, messages=msgs, temperature=0.7)
|
|
281
|
+
```
|
|
282
|
+
|
|
283
|
+
Other tags (e.g. `experiment=...`) are not part of the key. Records are written asynchronously, so an identical request sent a few milliseconds after the first one may still go to the server.
|
|
284
|
+
|
|
285
|
+
## Cost and budget
|
|
286
|
+
|
|
287
|
+
With `cost=True` or a `rr.CostConfig`, every record gets a `cost_usd`:
|
|
288
|
+
|
|
289
|
+
1. the cost the provider reports (OpenRouter `usage.cost`, LiteLLM `x-litellm-response-cost`, or a custom extractor), else
|
|
290
|
+
2. the cost computed from token usage and `CostConfig(prices={"model": rr.Price(input=..., output=..., cached_input=...)})` (USD per million tokens; there is no built-in price list, since it would go stale), else
|
|
291
|
+
3. `None`.
|
|
292
|
+
|
|
293
|
+
Cache hits cost `0.0`. `client.cost.stats()` shows what was spent, and the per-model summary line and `client.throttle.stats()` include it.
|
|
294
|
+
|
|
295
|
+
```python
|
|
296
|
+
client = rr.AsyncOpenAI(..., log=rr.LogConfig(sqlite="requests.db"), cost=rr.CostConfig(budget_usd=5))
|
|
297
|
+
```
|
|
298
|
+
|
|
299
|
+
- Once `budget_usd` has been spent, new requests raise `rr.BudgetExceeded` instead of being sent (`run_batch` returns it for each remaining item). Cache hits are still served, so re-running a finished experiment works with an exhausted budget.
|
|
300
|
+
- With `include_logged=True` (the default), the spend already recorded in the log counts too, so restarting a script after a crash does not reset the budget. This sums **all** records in the log file, so use one database per experiment or budget.
|
|
301
|
+
- Requests already in flight when the budget runs out still complete, so the budget can be exceeded by their cost.
|
|
302
|
+
- If a budget is set but a model's responses carry no cost and no price is configured, a warning is logged once: the budget cannot see those requests.
|
|
303
|
+
- With `openai<2`, the SDK retries requests after an exception, so a blocked request is attempted `max_retries` more times (each blocked again, without a network call) before `rr.BudgetExceeded` is raised.
|
|
304
|
+
|
|
305
|
+
## Caveats
|
|
306
|
+
|
|
307
|
+
- **Your own load raises latency too.** If `max_rpm` is high enough to saturate the server by itself, the client will throttle itself. The low-percentile baseline and the dead band absorb moderate self-load. Choose `max_rpm`/`max_concurrency` so that you alone don't saturate the backend.
|
|
308
|
+
- **One process, one event loop per client.** Throttle state is not shared between processes. Several scripts running in parallel each throttle independently.
|
|
309
|
+
- **Detection latency.** At `min_rpm=1` and `window=10`, noticing that others have left takes about ten minutes plus the cooldown. This is deliberate: the client stays polite longer than strictly needed.
|
|
310
|
+
- If the backend becomes permanently slower (e.g. a model is redeployed on smaller hardware), call `client.throttle.reset_baseline()`.
|
|
311
|
+
|
|
312
|
+
## Development
|
|
313
|
+
|
|
314
|
+
```bash
|
|
315
|
+
uv sync
|
|
316
|
+
uv run pytest
|
|
317
|
+
uv run ruff check . && uv run mypy src
|
|
318
|
+
uv run python examples/simulate_load.py
|
|
319
|
+
```
|
|
@@ -0,0 +1,285 @@
|
|
|
1
|
+
# responsible-request
|
|
2
|
+
|
|
3
|
+
Responsible LLM requests for research, against any OpenAI-compatible endpoint (LiteLLM, OpenRouter, vLLM, Ollama, OpenAI, …). It works as a drop-in for `openai.AsyncOpenAI`.
|
|
4
|
+
|
|
5
|
+
- **Polite**: load-aware throttling, so large experiments don't starve the other users of a shared inference server.
|
|
6
|
+
- **Reproducible**: every request is logged (parameters, messages, responses, tokens, timing, provider, system fingerprint) to JSONL and/or SQLite, each client writes a run manifest (versions, git commit, configuration), and `rr.reproducible()` pins the decoding parameters.
|
|
7
|
+
- **Cost-safe**: a response cache means a restarted script only sends what is still missing, and per-request cost tracking with an optional budget stops spending before it gets out of hand, even across restarts.
|
|
8
|
+
|
|
9
|
+
## Load-aware throttling in a nutshell
|
|
10
|
+
|
|
11
|
+
Large experiments can starve the other users of a shared inference server. Hard RPM limits are the usual answer, but they leave the server idle at night and still hurt others at peak times. `responsible-request` instead **watches the latency of your own requests**:
|
|
12
|
+
|
|
13
|
+
- While latency stays close to its baseline, the server is idle, so the client ramps up to `max_rpm`.
|
|
14
|
+
- As soon as latency reaches 3× the baseline (or the server returns 429/5xx/timeouts), other users are active, so the client drops to `min_rpm` at once.
|
|
15
|
+
- Once latency is back to normal, the client ramps up again.
|
|
16
|
+
|
|
17
|
+
For commercial APIs with generous rate limits of their own (OpenAI, OpenRouter, …), load-aware throttling is not needed: use a plain fixed-rate limiter, `rr.ThrottleConfig.fixed(rpm)`. In fixed mode a 429 does not change the rate: it pauses only the affected model (and endpoint) for the time in the `Retry-After` header, or for an exponential backoff (`backoff_initial_s`, doubling up to `backoff_max_s`) if there is none, plus a random jitter of up to `backoff_jitter_s`. The adaptive throttle instead treats a 429 as a sign of other users and drops to `min_rpm` for at least `cooldown_s`.
|
|
18
|
+
|
|
19
|
+
## Installation
|
|
20
|
+
|
|
21
|
+
```bash
|
|
22
|
+
pip install git+https://github.com/digital-sustainability/ResponsibleRequest.git
|
|
23
|
+
# or: uv add git+https://github.com/digital-sustainability/ResponsibleRequest.git
|
|
24
|
+
pip install "responsible-request[pandas,progress] @ git+https://github.com/digital-sustainability/ResponsibleRequest.git"
|
|
25
|
+
```
|
|
26
|
+
|
|
27
|
+
Requires Python ≥ 3.10 and `openai` ≥ 1.40 (it works with both the `httpx`-based 1.x/2.x SDKs and the `httpx2`-based 3.x SDK).
|
|
28
|
+
|
|
29
|
+
## Quickstart
|
|
30
|
+
|
|
31
|
+
```python
|
|
32
|
+
import responsible_request as rr
|
|
33
|
+
|
|
34
|
+
client = rr.AsyncOpenAI( # same arguments as openai.AsyncOpenAI
|
|
35
|
+
base_url="https://inference.example.org/api/v1",
|
|
36
|
+
api_key="sk-...",
|
|
37
|
+
throttle=rr.ThrottleConfig(max_rpm=15, min_rpm=1),
|
|
38
|
+
log=rr.LogConfig(sqlite="requests.db"),
|
|
39
|
+
default_params=rr.reproducible(seed=42), # temperature=0, top_p=1, seed=42 unless set per call
|
|
40
|
+
)
|
|
41
|
+
|
|
42
|
+
response = await client.chat.completions.create(model="gpt-oss:120b", messages=[...])
|
|
43
|
+
print(client.throttle.stats())
|
|
44
|
+
await client.close()
|
|
45
|
+
```
|
|
46
|
+
|
|
47
|
+
As with the OpenAI SDK, `base_url` and `api_key` can come from `OPENAI_BASE_URL` / `OPENAI_API_KEY` instead. For the examples in this repo, copy `.env.example` to `.env` and run them with `uv run --env-file .env python examples/quickstart.py`.
|
|
48
|
+
|
|
49
|
+
For existing code, only the HTTP client has to change:
|
|
50
|
+
|
|
51
|
+
```python
|
|
52
|
+
import openai, responsible_request as rr
|
|
53
|
+
|
|
54
|
+
client = openai.AsyncOpenAI(http_client=rr.http_client(rr.ThrottleConfig(max_rpm=15)))
|
|
55
|
+
```
|
|
56
|
+
|
|
57
|
+
It also works as a plain fixed-rate limiter for everyday use: `rr.ThrottleConfig.fixed(60)`.
|
|
58
|
+
|
|
59
|
+
With OpenRouter (or any other commercial endpoint), a typical setup is:
|
|
60
|
+
|
|
61
|
+
```python
|
|
62
|
+
client = rr.AsyncOpenAI(
|
|
63
|
+
base_url="https://openrouter.ai/api/v1",
|
|
64
|
+
api_key=os.environ["OPENROUTER_API_KEY"],
|
|
65
|
+
throttle=rr.ThrottleConfig.fixed(120),
|
|
66
|
+
log=rr.LogConfig(sqlite="requests.db"),
|
|
67
|
+
cache=True, # a re-run only sends what is missing
|
|
68
|
+
cost=rr.CostConfig(budget_usd=5), # stop at $5, counting earlier runs in requests.db
|
|
69
|
+
run=rr.RunConfig(name="ablation-3"),
|
|
70
|
+
default_params=rr.reproducible(seed=42),
|
|
71
|
+
)
|
|
72
|
+
```
|
|
73
|
+
|
|
74
|
+
See `examples/openrouter.py`.
|
|
75
|
+
|
|
76
|
+
## Helpers
|
|
77
|
+
|
|
78
|
+
```python
|
|
79
|
+
# Many requests, paced by the throttle, results in input order (exceptions returned, not raised)
|
|
80
|
+
results = await rr.run_batch(client, [{"model": m, "messages": msgs} for msgs in dataset], progress=True)
|
|
81
|
+
results = rr.run_batch_sync(client, requests) # from a plain script
|
|
82
|
+
|
|
83
|
+
# Structured output validated with Pydantic; the validation error is fed back to the model on failure
|
|
84
|
+
class Answer(BaseModel):
|
|
85
|
+
label: str
|
|
86
|
+
confidence: float
|
|
87
|
+
|
|
88
|
+
answer = await rr.structured(client, model=m, messages=msgs, schema=Answer, retries=2)
|
|
89
|
+
# (if `content` is empty, the answer is read from `reasoning`, then `reasoning_content`)
|
|
90
|
+
rr.response_format_from_model(Answer) # just the strict json_schema response_format
|
|
91
|
+
|
|
92
|
+
# System prompt marked for OpenRouter's prompt cache (cache_control: ephemeral)
|
|
93
|
+
messages = [rr.cached_system_message(long_instructions), {"role": "user", "content": question}]
|
|
94
|
+
|
|
95
|
+
# Attach metadata to every request record created inside the block
|
|
96
|
+
with rr.tags(experiment="ablation-3", fold=2):
|
|
97
|
+
...
|
|
98
|
+
|
|
99
|
+
# Measure and pin the baseline explicitly, e.g. at night with a representative request
|
|
100
|
+
await rr.calibrate(client, "gpt-oss:120b", n=10, messages=representative_messages)
|
|
101
|
+
```
|
|
102
|
+
|
|
103
|
+
`rr.cached_system_message()` only asks the provider to cache the prompt prefix (cheaper input tokens, the model still runs), whereas `cache=True` replays complete stored responses from your SQLite or JSONL log without sending the request.
|
|
104
|
+
|
|
105
|
+
## How the throttle works
|
|
106
|
+
|
|
107
|
+
```
|
|
108
|
+
ratio >= high_ratio (3.0) or 429/5xx/timeout
|
|
109
|
+
┌────────────────────────────────────────────────────────────┐
|
|
110
|
+
│ ▼
|
|
111
|
+
WARMUP ──baseline known──► NORMAL ◄──at max_rpm── RECOVERING ◄── THROTTLED (min_rpm)
|
|
112
|
+
start_rpm ramp ×1.5 / 30 s ramp ×1.5 / 30 s cooldown ≥ 120 s and
|
|
113
|
+
up to max_rpm ratio < recover_ratio (1.5)
|
|
114
|
+
```
|
|
115
|
+
|
|
116
|
+
- **Pacing**: requests are spaced evenly at the current rate with no bursts, and `max_concurrency` caps the number in flight. One `Throttle` can be shared by several clients, also across threads that each run their own event loop (e.g. `asyncio.run` per thread), and then enforces one rate for all of them:
|
|
117
|
+
|
|
118
|
+
```python
|
|
119
|
+
throttle = rr.Throttle(rr.ThrottleConfig.fixed(60))
|
|
120
|
+
def worker(items): # runs in its own thread
|
|
121
|
+
asyncio.run(work(rr.AsyncOpenAI(throttle=throttle, ...), items))
|
|
122
|
+
```
|
|
123
|
+
- **Per model**: every (endpoint, model) pair has its own lane, because load on one backend says nothing about another. Only `/chat/completions` and `/embeddings` feed the latency estimate (`observe_paths`). Other endpoints (audio, …) are paced and react to errors only. Requests with server-side tools (MCP, web search) are ignored for load estimation, because their latency includes external calls.
|
|
124
|
+
- **Signal**: by default, latency per completion token (`latency / max(completion_tokens, 16)`), so long answers and reasoning traces don't look like load. For streamed requests the signal is the time to first byte. The current value is the median of the last `window` requests.
|
|
125
|
+
- **Baseline**: the 10th percentile of the signal over the last 30 minutes, needing `warmup_requests` samples first. Samples taken while throttled are excluded, so a long busy period does not become the new "normal". You can pin it with `ThrottleConfig(baseline=...)` or `rr.calibrate(...)`, and reset it with `client.throttle.reset_baseline()`.
|
|
126
|
+
- **Dead band**: between `recover_ratio` and `high_ratio` the rate stays where it is, which prevents oscillation.
|
|
127
|
+
|
|
128
|
+
`examples/simulate_load.py` runs the whole loop against a simulated server (time-compressed) where other users appear for a while:
|
|
129
|
+
|
|
130
|
+
```
|
|
131
|
+
t [s] others RPM state load
|
|
132
|
+
8.5 0 1800 normal 1.00
|
|
133
|
+
12.5 48 1800 normal 1.00
|
|
134
|
+
13.5 48 60 throttled 3.56
|
|
135
|
+
...
|
|
136
|
+
30.6 0 60 recovering 1.00
|
|
137
|
+
41.6 0 1800 normal 1.00
|
|
138
|
+
```
|
|
139
|
+
|
|
140
|
+
### Configuration (`ThrottleConfig`)
|
|
141
|
+
|
|
142
|
+
| Parameter | Default | Meaning |
|
|
143
|
+
|---|---|---|
|
|
144
|
+
| `max_rpm` | 15 | rate when the endpoint is idle |
|
|
145
|
+
| `min_rpm` | 1 | rate while others are active (also the probing rate) |
|
|
146
|
+
| `start_rpm` | 2 | rate while the baseline is being established |
|
|
147
|
+
| `max_concurrency` | 32 | max requests in flight per model |
|
|
148
|
+
| `high_ratio` | 3.0 | throttle when latency ≥ this × baseline |
|
|
149
|
+
| `recover_ratio` | 1.5 | ramp up only when latency < this × baseline |
|
|
150
|
+
| `cooldown_s` | 120 | minimum time at `min_rpm` after the last high-load signal |
|
|
151
|
+
| `ramp_factor` / `ramp_interval_s` | 1.5 / 30 | multiplicative ramp-up step and its interval |
|
|
152
|
+
| `metric` | `"latency_per_token"` | `"latency"`, `"ttfb"`, or a callable `RequestRecord -> float` |
|
|
153
|
+
| `window` | 10 | requests in the rolling median |
|
|
154
|
+
| `warmup_requests` | 20 | samples needed before the baseline is trusted |
|
|
155
|
+
| `baseline` | None | pin the baseline instead of estimating it |
|
|
156
|
+
| `baseline_percentile` / `baseline_window_s` | 10 / 1800 | baseline estimator |
|
|
157
|
+
| `observe_paths` | chat, embeddings | endpoints whose latency is used |
|
|
158
|
+
| `estimator_factory` | None | plug in your own `LoadEstimator` (e.g. reading queue depth from Prometheus) |
|
|
159
|
+
| `backoff_initial_s` / `backoff_max_s` | 1 / 60 | fixed mode: pause after a 429 without `Retry-After`, doubling per further 429 |
|
|
160
|
+
| `backoff_jitter_s` | 1 | fixed mode: random extra time added to every 429 pause |
|
|
161
|
+
|
|
162
|
+
## Logging
|
|
163
|
+
|
|
164
|
+
The package logs through [loguru](https://github.com/Delgan/loguru) and, as loguru recommends for libraries, stays silent until a client is created with `log=` (the default `log=True` enables console messages only).
|
|
165
|
+
|
|
166
|
+
- **Console**: rate changes (INFO), switches to throttling (WARNING), and a per-model summary every `summary_interval_s`. Individual requests are never printed. With `LogConfig(console_level="INFO")` the package replaces loguru's default DEBUG handler with a quieter one.
|
|
167
|
+
- **Request records**: one record per HTTP attempt (SDK retries are separate records with `attempt` > 0) goes to `LogConfig(jsonl=...)` and/or `LogConfig(sqlite=...)` (table `requests`). Records are emitted at loguru's TRACE level with `extra["rr_record"]`, so you can also attach your own sinks.
|
|
168
|
+
|
|
169
|
+
| Group | Fields |
|
|
170
|
+
|---|---|
|
|
171
|
+
| core (always) | `request_id`, `timestamp`, `method`, `path`, `model`, `stream`, `attempt`, `status_code`, `error`, `tags`, `cache_key`, `cache_hit`, `run_id` |
|
|
172
|
+
| `timing` | `sent_at`, `first_byte_at`, `finished_at`, `wait_s` (time spent throttled), `ttfb_s`, `latency_s` |
|
|
173
|
+
| `usage` | `prompt_tokens`, `completion_tokens`, `total_tokens`, `cached_tokens`, `reasoning_tokens`, `cost_usd` |
|
|
174
|
+
| `throttle` | `rpm`, `state`, `baseline`, `load_ratio`, `in_flight` |
|
|
175
|
+
| `response_meta` | `response_id`, `response_model`, `finish_reason`, `system_fingerprint`, `provider`, `gateway_request_id`, `upstream_duration_ms`, `provider_meta` (see below) |
|
|
176
|
+
| `params` | all request parameters except messages/input |
|
|
177
|
+
| `request_body` / `response_body` | full request and response (for streams: the concatenated content) |
|
|
178
|
+
|
|
179
|
+
All groups are logged by default. Turn groups off with `LogConfig(fields={"request_body": False})`, or use `LogConfig.minimal()` (no params or bodies). **Request and response bodies can be large and may contain sensitive data.**
|
|
180
|
+
|
|
181
|
+
```python
|
|
182
|
+
df = rr.load_records("requests.db") # pandas DataFrame if pandas is installed, else list of dicts
|
|
183
|
+
```
|
|
184
|
+
|
|
185
|
+
JSONL logs rotate at `rotation="500 MB"` (any loguru rotation). LLM logs are repetitive and compress well, so finished segments can be compressed with `compression="gz"`, `"bz2"`, `"xz"` or `"zst"` (zstd needs Python 3.14+ or `pip install responsible-request[zstd]`); the file being written stays plain. Without rotation, the file is compressed when the client's sinks are removed. `load_records("requests.jsonl")`, the cache and the cost budget read all segments, compressed or not, and `load_records` / `CacheConfig(jsonl=...)` also accept a single archive such as `requests.2026-10-06_12-00-00_000000.jsonl.zst`.
|
|
186
|
+
|
|
187
|
+
```python
|
|
188
|
+
log = rr.LogConfig(jsonl="requests.jsonl", rotation="100 MB", compression="zst")
|
|
189
|
+
```
|
|
190
|
+
|
|
191
|
+
### Provider metadata
|
|
192
|
+
|
|
193
|
+
Gateways and routers report different extra information. *Extractors* map it to the generic `response_meta` fields and put everything else into the `provider_meta` JSON column. All built-in extractors run on every response and only fill in what they find, so nothing has to be configured:
|
|
194
|
+
|
|
195
|
+
| Extractor | Reads |
|
|
196
|
+
|---|---|
|
|
197
|
+
| `openai_compat` | `x-request-id` / `request-id` → `gateway_request_id`, `openai-processing-ms` → `upstream_duration_ms` |
|
|
198
|
+
| `litellm` | `x-litellm-call-id` → `gateway_request_id`, `x-litellm-response-duration-ms` → `upstream_duration_ms`, `x-litellm-response-cost` → `cost_usd`, all other `x-litellm-*` headers → `provider_meta["litellm"]` |
|
|
199
|
+
| `openrouter` | `provider` → `provider`, `usage.cost` → `cost_usd`, `usage.is_byok`/`cost_details` and `native_finish_reason` → `provider_meta["openrouter"]` |
|
|
200
|
+
|
|
201
|
+
Add your own for other endpoints:
|
|
202
|
+
|
|
203
|
+
```python
|
|
204
|
+
def my_gateway(record, headers, body): # body: parsed JSON, stream summary, or None
|
|
205
|
+
if "x-queue-ms" in headers:
|
|
206
|
+
rr.providers.meta(record, "my_gateway")["queue_ms"] = float(headers["x-queue-ms"])
|
|
207
|
+
|
|
208
|
+
client = rr.AsyncOpenAI(..., extractors=[my_gateway])
|
|
209
|
+
```
|
|
210
|
+
|
|
211
|
+
Databases written by version 0.1 keep their `litellm_*` columns; new records leave them empty.
|
|
212
|
+
|
|
213
|
+
### Runs
|
|
214
|
+
|
|
215
|
+
Every client is a *run* with its own `run_id`, which every record carries. With a JSONL or SQLite log, the first request also writes a run manifest: the run name and `metadata`, start time, endpoint (scheme and host only, never the API key), versions of `responsible-request`, `openai` and Python, platform, hostname, command line, working directory, git commit and dirty flag, and the complete throttle, log, cache and cost configuration plus `default_params`.
|
|
216
|
+
|
|
217
|
+
```python
|
|
218
|
+
client = rr.AsyncOpenAI(..., run=rr.RunConfig(name="ablation-3", metadata={"dataset": "v2"}))
|
|
219
|
+
# or simply run="ablation-3"
|
|
220
|
+
|
|
221
|
+
runs = rr.load_runs("requests.db") # the `runs` table, or requests.runs.jsonl for a JSONL log
|
|
222
|
+
df = rr.load_records("requests.db").merge(runs, on="run_id", suffixes=("", "_run"))
|
|
223
|
+
```
|
|
224
|
+
|
|
225
|
+
## Caching
|
|
226
|
+
|
|
227
|
+
With `cache=True`, a request that was already answered successfully is served from the logged records instead of being sent again. Re-running a script (e.g. after a crash) then only sends the requests that are still missing:
|
|
228
|
+
|
|
229
|
+
```python
|
|
230
|
+
client = rr.AsyncOpenAI(
|
|
231
|
+
log=rr.LogConfig(sqlite="requests.db"),
|
|
232
|
+
cache=True, # or rr.CacheConfig(...)
|
|
233
|
+
)
|
|
234
|
+
```
|
|
235
|
+
|
|
236
|
+
- **Key**: a hash of the URL and the complete JSON request body (model, messages and all parameters, after `default_params` are applied). Any change to the prompt or parameters is a miss.
|
|
237
|
+
- **Source**: the records the client writes itself (`LogConfig.sqlite`, else `LogConfig.jsonl`; `response_body` must be logged), or another database or file via `rr.CacheConfig(sqlite=...)` / `rr.CacheConfig(jsonl=...)`.
|
|
238
|
+
- **What is served**: only non-streamed requests whose record has status 200 and no error. Streamed requests are always sent.
|
|
239
|
+
- **Hits** don't wait for or affect the throttle. They are logged with `cache_hit=True` (filter them out when analysing latency or token usage), and the response carries an `x-rr-cache: hit` header. `client.cache.stats()` counts hits and misses.
|
|
240
|
+
- **Repeated sampling**: identical requests share one cached answer. To draw several independent samples, name the tags that belong to the key:
|
|
241
|
+
|
|
242
|
+
```python
|
|
243
|
+
client = rr.AsyncOpenAI(log=..., cache=rr.CacheConfig(key_tags=("sample",)))
|
|
244
|
+
for i in range(5):
|
|
245
|
+
with rr.tags(sample=i): # 5 different keys; a re-run gets the same 5 answers
|
|
246
|
+
await client.chat.completions.create(model=m, messages=msgs, temperature=0.7)
|
|
247
|
+
```
|
|
248
|
+
|
|
249
|
+
Other tags (e.g. `experiment=...`) are not part of the key. Records are written asynchronously, so an identical request sent a few milliseconds after the first one may still go to the server.
|
|
250
|
+
|
|
251
|
+
## Cost and budget
|
|
252
|
+
|
|
253
|
+
With `cost=True` or a `rr.CostConfig`, every record gets a `cost_usd`:
|
|
254
|
+
|
|
255
|
+
1. the cost the provider reports (OpenRouter `usage.cost`, LiteLLM `x-litellm-response-cost`, or a custom extractor), else
|
|
256
|
+
2. the cost computed from token usage and `CostConfig(prices={"model": rr.Price(input=..., output=..., cached_input=...)})` (USD per million tokens; there is no built-in price list, since it would go stale), else
|
|
257
|
+
3. `None`.
|
|
258
|
+
|
|
259
|
+
Cache hits cost `0.0`. `client.cost.stats()` shows what was spent, and the per-model summary line and `client.throttle.stats()` include it.
|
|
260
|
+
|
|
261
|
+
```python
|
|
262
|
+
client = rr.AsyncOpenAI(..., log=rr.LogConfig(sqlite="requests.db"), cost=rr.CostConfig(budget_usd=5))
|
|
263
|
+
```
|
|
264
|
+
|
|
265
|
+
- Once `budget_usd` has been spent, new requests raise `rr.BudgetExceeded` instead of being sent (`run_batch` returns it for each remaining item). Cache hits are still served, so re-running a finished experiment works with an exhausted budget.
|
|
266
|
+
- With `include_logged=True` (the default), the spend already recorded in the log counts too, so restarting a script after a crash does not reset the budget. This sums **all** records in the log file, so use one database per experiment or budget.
|
|
267
|
+
- Requests already in flight when the budget runs out still complete, so the budget can be exceeded by their cost.
|
|
268
|
+
- If a budget is set but a model's responses carry no cost and no price is configured, a warning is logged once: the budget cannot see those requests.
|
|
269
|
+
- With `openai<2`, the SDK retries requests after an exception, so a blocked request is attempted `max_retries` more times (each blocked again, without a network call) before `rr.BudgetExceeded` is raised.
|
|
270
|
+
|
|
271
|
+
## Caveats
|
|
272
|
+
|
|
273
|
+
- **Your own load raises latency too.** If `max_rpm` is high enough to saturate the server by itself, the client will throttle itself. The low-percentile baseline and the dead band absorb moderate self-load. Choose `max_rpm`/`max_concurrency` so that you alone don't saturate the backend.
|
|
274
|
+
- **One process, one event loop per client.** Throttle state is not shared between processes. Several scripts running in parallel each throttle independently.
|
|
275
|
+
- **Detection latency.** At `min_rpm=1` and `window=10`, noticing that others have left takes about ten minutes plus the cooldown. This is deliberate: the client stays polite longer than strictly needed.
|
|
276
|
+
- If the backend becomes permanently slower (e.g. a model is redeployed on smaller hardware), call `client.throttle.reset_baseline()`.
|
|
277
|
+
|
|
278
|
+
## Development
|
|
279
|
+
|
|
280
|
+
```bash
|
|
281
|
+
uv sync
|
|
282
|
+
uv run pytest
|
|
283
|
+
uv run ruff check . && uv run mypy src
|
|
284
|
+
uv run python examples/simulate_load.py
|
|
285
|
+
```
|