responsible-request 0.4.0__tar.gz

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (29) hide show
  1. responsible_request-0.4.0/LICENSE +21 -0
  2. responsible_request-0.4.0/PKG-INFO +319 -0
  3. responsible_request-0.4.0/README.md +285 -0
  4. responsible_request-0.4.0/pyproject.toml +89 -0
  5. responsible_request-0.4.0/pyproject.toml.orig +69 -0
  6. responsible_request-0.4.0/src/responsible_request/__init__.py +94 -0
  7. responsible_request-0.4.0/src/responsible_request/_http.py +31 -0
  8. responsible_request-0.4.0/src/responsible_request/analysis.py +81 -0
  9. responsible_request-0.4.0/src/responsible_request/cache.py +59 -0
  10. responsible_request-0.4.0/src/responsible_request/client.py +169 -0
  11. responsible_request-0.4.0/src/responsible_request/config.py +266 -0
  12. responsible_request-0.4.0/src/responsible_request/controller.py +79 -0
  13. responsible_request-0.4.0/src/responsible_request/cost.py +96 -0
  14. responsible_request-0.4.0/src/responsible_request/estimator.py +113 -0
  15. responsible_request-0.4.0/src/responsible_request/helpers/__init__.py +21 -0
  16. responsible_request-0.4.0/src/responsible_request/helpers/batch.py +127 -0
  17. responsible_request-0.4.0/src/responsible_request/helpers/messages.py +18 -0
  18. responsible_request-0.4.0/src/responsible_request/helpers/params.py +18 -0
  19. responsible_request-0.4.0/src/responsible_request/helpers/structured.py +150 -0
  20. responsible_request-0.4.0/src/responsible_request/limiter.py +165 -0
  21. responsible_request-0.4.0/src/responsible_request/logfiles.py +120 -0
  22. responsible_request-0.4.0/src/responsible_request/logging.py +183 -0
  23. responsible_request-0.4.0/src/responsible_request/providers.py +105 -0
  24. responsible_request-0.4.0/src/responsible_request/py.typed +0 -0
  25. responsible_request-0.4.0/src/responsible_request/records.py +265 -0
  26. responsible_request-0.4.0/src/responsible_request/runs.py +116 -0
  27. responsible_request-0.4.0/src/responsible_request/sources.py +207 -0
  28. responsible_request-0.4.0/src/responsible_request/throttle.py +229 -0
  29. responsible_request-0.4.0/src/responsible_request/transport.py +376 -0
@@ -0,0 +1,21 @@
1
+ MIT License
2
+
3
+ Copyright (c) 2026 Bern University of Applied Sciences (BFH)
4
+
5
+ Permission is hereby granted, free of charge, to any person obtaining a copy
6
+ of this software and associated documentation files (the "Software"), to deal
7
+ in the Software without restriction, including without limitation the rights
8
+ to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
9
+ copies of the Software, and to permit persons to whom the Software is
10
+ furnished to do so, subject to the following conditions:
11
+
12
+ The above copyright notice and this permission notice shall be included in all
13
+ copies or substantial portions of the Software.
14
+
15
+ THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
16
+ IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
17
+ FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
18
+ AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
19
+ LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
20
+ OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
21
+ SOFTWARE.
@@ -0,0 +1,319 @@
1
+ Metadata-Version: 2.4
2
+ Name: responsible-request
3
+ Version: 0.4.0
4
+ Summary: Polite, reproducible and cost-safe LLM requests for research: load-aware throttling, request logging, caching and budgets for OpenAI-compatible endpoints.
5
+ Keywords: llm,openai,litellm,openrouter,rate-limiting,throttling,logging,caching,reproducibility
6
+ Author: Luca Rolshoven
7
+ Author-email: Luca Rolshoven <luca@rolshoven.io>
8
+ License-Expression: MIT
9
+ License-File: LICENSE
10
+ Classifier: Development Status :: 3 - Alpha
11
+ Classifier: Framework :: AsyncIO
12
+ Classifier: Intended Audience :: Science/Research
13
+ Classifier: Programming Language :: Python :: 3
14
+ Classifier: Programming Language :: Python :: 3.10
15
+ Classifier: Programming Language :: Python :: 3.11
16
+ Classifier: Programming Language :: Python :: 3.12
17
+ Classifier: Programming Language :: Python :: 3.13
18
+ Classifier: Programming Language :: Python :: 3.14
19
+ Classifier: Typing :: Typed
20
+ Classifier: Topic :: Scientific/Engineering :: Artificial Intelligence
21
+ Requires-Dist: loguru>=0.7
22
+ Requires-Dist: openai>=1.40
23
+ Requires-Dist: pydantic>=2
24
+ Requires-Dist: pandas>=2 ; extra == 'pandas'
25
+ Requires-Dist: tqdm>=4 ; extra == 'progress'
26
+ Requires-Dist: zstandard>=0.22 ; python_full_version < '3.14' and extra == 'zstd'
27
+ Requires-Python: >=3.10
28
+ Project-URL: Repository, https://github.com/digital-sustainability/ResponsibleRequest
29
+ Project-URL: Issues, https://github.com/digital-sustainability/ResponsibleRequest/issues
30
+ Provides-Extra: pandas
31
+ Provides-Extra: progress
32
+ Provides-Extra: zstd
33
+ Description-Content-Type: text/markdown
34
+
35
+ # responsible-request
36
+
37
+ Responsible LLM requests for research, against any OpenAI-compatible endpoint (LiteLLM, OpenRouter, vLLM, Ollama, OpenAI, …). It works as a drop-in for `openai.AsyncOpenAI`.
38
+
39
+ - **Polite**: load-aware throttling, so large experiments don't starve the other users of a shared inference server.
40
+ - **Reproducible**: every request is logged (parameters, messages, responses, tokens, timing, provider, system fingerprint) to JSONL and/or SQLite, each client writes a run manifest (versions, git commit, configuration), and `rr.reproducible()` pins the decoding parameters.
41
+ - **Cost-safe**: a response cache means a restarted script only sends what is still missing, and per-request cost tracking with an optional budget stops spending before it gets out of hand, even across restarts.
42
+
43
+ ## Load-aware throttling in a nutshell
44
+
45
+ Large experiments can starve the other users of a shared inference server. Hard RPM limits are the usual answer, but they leave the server idle at night and still hurt others at peak times. `responsible-request` instead **watches the latency of your own requests**:
46
+
47
+ - While latency stays close to its baseline, the server is idle, so the client ramps up to `max_rpm`.
48
+ - As soon as latency reaches 3× the baseline (or the server returns 429/5xx/timeouts), other users are active, so the client drops to `min_rpm` at once.
49
+ - Once latency is back to normal, the client ramps up again.
50
+
51
+ For commercial APIs with generous rate limits of their own (OpenAI, OpenRouter, …), load-aware throttling is not needed: use a plain fixed-rate limiter, `rr.ThrottleConfig.fixed(rpm)`. In fixed mode a 429 does not change the rate: it pauses only the affected model (and endpoint) for the time in the `Retry-After` header, or for an exponential backoff (`backoff_initial_s`, doubling up to `backoff_max_s`) if there is none, plus a random jitter of up to `backoff_jitter_s`. The adaptive throttle instead treats a 429 as a sign of other users and drops to `min_rpm` for at least `cooldown_s`.
52
+
53
+ ## Installation
54
+
55
+ ```bash
56
+ pip install git+https://github.com/digital-sustainability/ResponsibleRequest.git
57
+ # or: uv add git+https://github.com/digital-sustainability/ResponsibleRequest.git
58
+ pip install "responsible-request[pandas,progress] @ git+https://github.com/digital-sustainability/ResponsibleRequest.git"
59
+ ```
60
+
61
+ Requires Python ≥ 3.10 and `openai` ≥ 1.40 (it works with both the `httpx`-based 1.x/2.x SDKs and the `httpx2`-based 3.x SDK).
62
+
63
+ ## Quickstart
64
+
65
+ ```python
66
+ import responsible_request as rr
67
+
68
+ client = rr.AsyncOpenAI( # same arguments as openai.AsyncOpenAI
69
+ base_url="https://inference.example.org/api/v1",
70
+ api_key="sk-...",
71
+ throttle=rr.ThrottleConfig(max_rpm=15, min_rpm=1),
72
+ log=rr.LogConfig(sqlite="requests.db"),
73
+ default_params=rr.reproducible(seed=42), # temperature=0, top_p=1, seed=42 unless set per call
74
+ )
75
+
76
+ response = await client.chat.completions.create(model="gpt-oss:120b", messages=[...])
77
+ print(client.throttle.stats())
78
+ await client.close()
79
+ ```
80
+
81
+ As with the OpenAI SDK, `base_url` and `api_key` can come from `OPENAI_BASE_URL` / `OPENAI_API_KEY` instead. For the examples in this repo, copy `.env.example` to `.env` and run them with `uv run --env-file .env python examples/quickstart.py`.
82
+
83
+ For existing code, only the HTTP client has to change:
84
+
85
+ ```python
86
+ import openai, responsible_request as rr
87
+
88
+ client = openai.AsyncOpenAI(http_client=rr.http_client(rr.ThrottleConfig(max_rpm=15)))
89
+ ```
90
+
91
+ It also works as a plain fixed-rate limiter for everyday use: `rr.ThrottleConfig.fixed(60)`.
92
+
93
+ With OpenRouter (or any other commercial endpoint), a typical setup is:
94
+
95
+ ```python
96
+ client = rr.AsyncOpenAI(
97
+ base_url="https://openrouter.ai/api/v1",
98
+ api_key=os.environ["OPENROUTER_API_KEY"],
99
+ throttle=rr.ThrottleConfig.fixed(120),
100
+ log=rr.LogConfig(sqlite="requests.db"),
101
+ cache=True, # a re-run only sends what is missing
102
+ cost=rr.CostConfig(budget_usd=5), # stop at $5, counting earlier runs in requests.db
103
+ run=rr.RunConfig(name="ablation-3"),
104
+ default_params=rr.reproducible(seed=42),
105
+ )
106
+ ```
107
+
108
+ See `examples/openrouter.py`.
109
+
110
+ ## Helpers
111
+
112
+ ```python
113
+ # Many requests, paced by the throttle, results in input order (exceptions returned, not raised)
114
+ results = await rr.run_batch(client, [{"model": m, "messages": msgs} for msgs in dataset], progress=True)
115
+ results = rr.run_batch_sync(client, requests) # from a plain script
116
+
117
+ # Structured output validated with Pydantic; the validation error is fed back to the model on failure
118
+ class Answer(BaseModel):
119
+ label: str
120
+ confidence: float
121
+
122
+ answer = await rr.structured(client, model=m, messages=msgs, schema=Answer, retries=2)
123
+ # (if `content` is empty, the answer is read from `reasoning`, then `reasoning_content`)
124
+ rr.response_format_from_model(Answer) # just the strict json_schema response_format
125
+
126
+ # System prompt marked for OpenRouter's prompt cache (cache_control: ephemeral)
127
+ messages = [rr.cached_system_message(long_instructions), {"role": "user", "content": question}]
128
+
129
+ # Attach metadata to every request record created inside the block
130
+ with rr.tags(experiment="ablation-3", fold=2):
131
+ ...
132
+
133
+ # Measure and pin the baseline explicitly, e.g. at night with a representative request
134
+ await rr.calibrate(client, "gpt-oss:120b", n=10, messages=representative_messages)
135
+ ```
136
+
137
+ `rr.cached_system_message()` only asks the provider to cache the prompt prefix (cheaper input tokens, the model still runs), whereas `cache=True` replays complete stored responses from your SQLite or JSONL log without sending the request.
138
+
139
+ ## How the throttle works
140
+
141
+ ```
142
+ ratio >= high_ratio (3.0) or 429/5xx/timeout
143
+ ┌────────────────────────────────────────────────────────────┐
144
+ │ ▼
145
+ WARMUP ──baseline known──► NORMAL ◄──at max_rpm── RECOVERING ◄── THROTTLED (min_rpm)
146
+ start_rpm ramp ×1.5 / 30 s ramp ×1.5 / 30 s cooldown ≥ 120 s and
147
+ up to max_rpm ratio < recover_ratio (1.5)
148
+ ```
149
+
150
+ - **Pacing**: requests are spaced evenly at the current rate with no bursts, and `max_concurrency` caps the number in flight. One `Throttle` can be shared by several clients, also across threads that each run their own event loop (e.g. `asyncio.run` per thread), and then enforces one rate for all of them:
151
+
152
+ ```python
153
+ throttle = rr.Throttle(rr.ThrottleConfig.fixed(60))
154
+ def worker(items): # runs in its own thread
155
+ asyncio.run(work(rr.AsyncOpenAI(throttle=throttle, ...), items))
156
+ ```
157
+ - **Per model**: every (endpoint, model) pair has its own lane, because load on one backend says nothing about another. Only `/chat/completions` and `/embeddings` feed the latency estimate (`observe_paths`). Other endpoints (audio, …) are paced and react to errors only. Requests with server-side tools (MCP, web search) are ignored for load estimation, because their latency includes external calls.
158
+ - **Signal**: by default, latency per completion token (`latency / max(completion_tokens, 16)`), so long answers and reasoning traces don't look like load. For streamed requests the signal is the time to first byte. The current value is the median of the last `window` requests.
159
+ - **Baseline**: the 10th percentile of the signal over the last 30 minutes, needing `warmup_requests` samples first. Samples taken while throttled are excluded, so a long busy period does not become the new "normal". You can pin it with `ThrottleConfig(baseline=...)` or `rr.calibrate(...)`, and reset it with `client.throttle.reset_baseline()`.
160
+ - **Dead band**: between `recover_ratio` and `high_ratio` the rate stays where it is, which prevents oscillation.
161
+
162
+ `examples/simulate_load.py` runs the whole loop against a simulated server (time-compressed) where other users appear for a while:
163
+
164
+ ```
165
+ t [s] others RPM state load
166
+ 8.5 0 1800 normal 1.00
167
+ 12.5 48 1800 normal 1.00
168
+ 13.5 48 60 throttled 3.56
169
+ ...
170
+ 30.6 0 60 recovering 1.00
171
+ 41.6 0 1800 normal 1.00
172
+ ```
173
+
174
+ ### Configuration (`ThrottleConfig`)
175
+
176
+ | Parameter | Default | Meaning |
177
+ |---|---|---|
178
+ | `max_rpm` | 15 | rate when the endpoint is idle |
179
+ | `min_rpm` | 1 | rate while others are active (also the probing rate) |
180
+ | `start_rpm` | 2 | rate while the baseline is being established |
181
+ | `max_concurrency` | 32 | max requests in flight per model |
182
+ | `high_ratio` | 3.0 | throttle when latency ≥ this × baseline |
183
+ | `recover_ratio` | 1.5 | ramp up only when latency < this × baseline |
184
+ | `cooldown_s` | 120 | minimum time at `min_rpm` after the last high-load signal |
185
+ | `ramp_factor` / `ramp_interval_s` | 1.5 / 30 | multiplicative ramp-up step and its interval |
186
+ | `metric` | `"latency_per_token"` | `"latency"`, `"ttfb"`, or a callable `RequestRecord -> float` |
187
+ | `window` | 10 | requests in the rolling median |
188
+ | `warmup_requests` | 20 | samples needed before the baseline is trusted |
189
+ | `baseline` | None | pin the baseline instead of estimating it |
190
+ | `baseline_percentile` / `baseline_window_s` | 10 / 1800 | baseline estimator |
191
+ | `observe_paths` | chat, embeddings | endpoints whose latency is used |
192
+ | `estimator_factory` | None | plug in your own `LoadEstimator` (e.g. reading queue depth from Prometheus) |
193
+ | `backoff_initial_s` / `backoff_max_s` | 1 / 60 | fixed mode: pause after a 429 without `Retry-After`, doubling per further 429 |
194
+ | `backoff_jitter_s` | 1 | fixed mode: random extra time added to every 429 pause |
195
+
196
+ ## Logging
197
+
198
+ The package logs through [loguru](https://github.com/Delgan/loguru) and, as loguru recommends for libraries, stays silent until a client is created with `log=` (the default `log=True` enables console messages only).
199
+
200
+ - **Console**: rate changes (INFO), switches to throttling (WARNING), and a per-model summary every `summary_interval_s`. Individual requests are never printed. With `LogConfig(console_level="INFO")` the package replaces loguru's default DEBUG handler with a quieter one.
201
+ - **Request records**: one record per HTTP attempt (SDK retries are separate records with `attempt` > 0) goes to `LogConfig(jsonl=...)` and/or `LogConfig(sqlite=...)` (table `requests`). Records are emitted at loguru's TRACE level with `extra["rr_record"]`, so you can also attach your own sinks.
202
+
203
+ | Group | Fields |
204
+ |---|---|
205
+ | core (always) | `request_id`, `timestamp`, `method`, `path`, `model`, `stream`, `attempt`, `status_code`, `error`, `tags`, `cache_key`, `cache_hit`, `run_id` |
206
+ | `timing` | `sent_at`, `first_byte_at`, `finished_at`, `wait_s` (time spent throttled), `ttfb_s`, `latency_s` |
207
+ | `usage` | `prompt_tokens`, `completion_tokens`, `total_tokens`, `cached_tokens`, `reasoning_tokens`, `cost_usd` |
208
+ | `throttle` | `rpm`, `state`, `baseline`, `load_ratio`, `in_flight` |
209
+ | `response_meta` | `response_id`, `response_model`, `finish_reason`, `system_fingerprint`, `provider`, `gateway_request_id`, `upstream_duration_ms`, `provider_meta` (see below) |
210
+ | `params` | all request parameters except messages/input |
211
+ | `request_body` / `response_body` | full request and response (for streams: the concatenated content) |
212
+
213
+ All groups are logged by default. Turn groups off with `LogConfig(fields={"request_body": False})`, or use `LogConfig.minimal()` (no params or bodies). **Request and response bodies can be large and may contain sensitive data.**
214
+
215
+ ```python
216
+ df = rr.load_records("requests.db") # pandas DataFrame if pandas is installed, else list of dicts
217
+ ```
218
+
219
+ JSONL logs rotate at `rotation="500 MB"` (any loguru rotation). LLM logs are repetitive and compress well, so finished segments can be compressed with `compression="gz"`, `"bz2"`, `"xz"` or `"zst"` (zstd needs Python 3.14+ or `pip install responsible-request[zstd]`); the file being written stays plain. Without rotation, the file is compressed when the client's sinks are removed. `load_records("requests.jsonl")`, the cache and the cost budget read all segments, compressed or not, and `load_records` / `CacheConfig(jsonl=...)` also accept a single archive such as `requests.2026-10-06_12-00-00_000000.jsonl.zst`.
220
+
221
+ ```python
222
+ log = rr.LogConfig(jsonl="requests.jsonl", rotation="100 MB", compression="zst")
223
+ ```
224
+
225
+ ### Provider metadata
226
+
227
+ Gateways and routers report different extra information. *Extractors* map it to the generic `response_meta` fields and put everything else into the `provider_meta` JSON column. All built-in extractors run on every response and only fill in what they find, so nothing has to be configured:
228
+
229
+ | Extractor | Reads |
230
+ |---|---|
231
+ | `openai_compat` | `x-request-id` / `request-id` → `gateway_request_id`, `openai-processing-ms` → `upstream_duration_ms` |
232
+ | `litellm` | `x-litellm-call-id` → `gateway_request_id`, `x-litellm-response-duration-ms` → `upstream_duration_ms`, `x-litellm-response-cost` → `cost_usd`, all other `x-litellm-*` headers → `provider_meta["litellm"]` |
233
+ | `openrouter` | `provider` → `provider`, `usage.cost` → `cost_usd`, `usage.is_byok`/`cost_details` and `native_finish_reason` → `provider_meta["openrouter"]` |
234
+
235
+ Add your own for other endpoints:
236
+
237
+ ```python
238
+ def my_gateway(record, headers, body): # body: parsed JSON, stream summary, or None
239
+ if "x-queue-ms" in headers:
240
+ rr.providers.meta(record, "my_gateway")["queue_ms"] = float(headers["x-queue-ms"])
241
+
242
+ client = rr.AsyncOpenAI(..., extractors=[my_gateway])
243
+ ```
244
+
245
+ Databases written by version 0.1 keep their `litellm_*` columns; new records leave them empty.
246
+
247
+ ### Runs
248
+
249
+ Every client is a *run* with its own `run_id`, which every record carries. With a JSONL or SQLite log, the first request also writes a run manifest: the run name and `metadata`, start time, endpoint (scheme and host only, never the API key), versions of `responsible-request`, `openai` and Python, platform, hostname, command line, working directory, git commit and dirty flag, and the complete throttle, log, cache and cost configuration plus `default_params`.
250
+
251
+ ```python
252
+ client = rr.AsyncOpenAI(..., run=rr.RunConfig(name="ablation-3", metadata={"dataset": "v2"}))
253
+ # or simply run="ablation-3"
254
+
255
+ runs = rr.load_runs("requests.db") # the `runs` table, or requests.runs.jsonl for a JSONL log
256
+ df = rr.load_records("requests.db").merge(runs, on="run_id", suffixes=("", "_run"))
257
+ ```
258
+
259
+ ## Caching
260
+
261
+ With `cache=True`, a request that was already answered successfully is served from the logged records instead of being sent again. Re-running a script (e.g. after a crash) then only sends the requests that are still missing:
262
+
263
+ ```python
264
+ client = rr.AsyncOpenAI(
265
+ log=rr.LogConfig(sqlite="requests.db"),
266
+ cache=True, # or rr.CacheConfig(...)
267
+ )
268
+ ```
269
+
270
+ - **Key**: a hash of the URL and the complete JSON request body (model, messages and all parameters, after `default_params` are applied). Any change to the prompt or parameters is a miss.
271
+ - **Source**: the records the client writes itself (`LogConfig.sqlite`, else `LogConfig.jsonl`; `response_body` must be logged), or another database or file via `rr.CacheConfig(sqlite=...)` / `rr.CacheConfig(jsonl=...)`.
272
+ - **What is served**: only non-streamed requests whose record has status 200 and no error. Streamed requests are always sent.
273
+ - **Hits** don't wait for or affect the throttle. They are logged with `cache_hit=True` (filter them out when analysing latency or token usage), and the response carries an `x-rr-cache: hit` header. `client.cache.stats()` counts hits and misses.
274
+ - **Repeated sampling**: identical requests share one cached answer. To draw several independent samples, name the tags that belong to the key:
275
+
276
+ ```python
277
+ client = rr.AsyncOpenAI(log=..., cache=rr.CacheConfig(key_tags=("sample",)))
278
+ for i in range(5):
279
+ with rr.tags(sample=i): # 5 different keys; a re-run gets the same 5 answers
280
+ await client.chat.completions.create(model=m, messages=msgs, temperature=0.7)
281
+ ```
282
+
283
+ Other tags (e.g. `experiment=...`) are not part of the key. Records are written asynchronously, so an identical request sent a few milliseconds after the first one may still go to the server.
284
+
285
+ ## Cost and budget
286
+
287
+ With `cost=True` or a `rr.CostConfig`, every record gets a `cost_usd`:
288
+
289
+ 1. the cost the provider reports (OpenRouter `usage.cost`, LiteLLM `x-litellm-response-cost`, or a custom extractor), else
290
+ 2. the cost computed from token usage and `CostConfig(prices={"model": rr.Price(input=..., output=..., cached_input=...)})` (USD per million tokens; there is no built-in price list, since it would go stale), else
291
+ 3. `None`.
292
+
293
+ Cache hits cost `0.0`. `client.cost.stats()` shows what was spent, and the per-model summary line and `client.throttle.stats()` include it.
294
+
295
+ ```python
296
+ client = rr.AsyncOpenAI(..., log=rr.LogConfig(sqlite="requests.db"), cost=rr.CostConfig(budget_usd=5))
297
+ ```
298
+
299
+ - Once `budget_usd` has been spent, new requests raise `rr.BudgetExceeded` instead of being sent (`run_batch` returns it for each remaining item). Cache hits are still served, so re-running a finished experiment works with an exhausted budget.
300
+ - With `include_logged=True` (the default), the spend already recorded in the log counts too, so restarting a script after a crash does not reset the budget. This sums **all** records in the log file, so use one database per experiment or budget.
301
+ - Requests already in flight when the budget runs out still complete, so the budget can be exceeded by their cost.
302
+ - If a budget is set but a model's responses carry no cost and no price is configured, a warning is logged once: the budget cannot see those requests.
303
+ - With `openai<2`, the SDK retries requests after an exception, so a blocked request is attempted `max_retries` more times (each blocked again, without a network call) before `rr.BudgetExceeded` is raised.
304
+
305
+ ## Caveats
306
+
307
+ - **Your own load raises latency too.** If `max_rpm` is high enough to saturate the server by itself, the client will throttle itself. The low-percentile baseline and the dead band absorb moderate self-load. Choose `max_rpm`/`max_concurrency` so that you alone don't saturate the backend.
308
+ - **One process, one event loop per client.** Throttle state is not shared between processes. Several scripts running in parallel each throttle independently.
309
+ - **Detection latency.** At `min_rpm=1` and `window=10`, noticing that others have left takes about ten minutes plus the cooldown. This is deliberate: the client stays polite longer than strictly needed.
310
+ - If the backend becomes permanently slower (e.g. a model is redeployed on smaller hardware), call `client.throttle.reset_baseline()`.
311
+
312
+ ## Development
313
+
314
+ ```bash
315
+ uv sync
316
+ uv run pytest
317
+ uv run ruff check . && uv run mypy src
318
+ uv run python examples/simulate_load.py
319
+ ```
@@ -0,0 +1,285 @@
1
+ # responsible-request
2
+
3
+ Responsible LLM requests for research, against any OpenAI-compatible endpoint (LiteLLM, OpenRouter, vLLM, Ollama, OpenAI, …). It works as a drop-in for `openai.AsyncOpenAI`.
4
+
5
+ - **Polite**: load-aware throttling, so large experiments don't starve the other users of a shared inference server.
6
+ - **Reproducible**: every request is logged (parameters, messages, responses, tokens, timing, provider, system fingerprint) to JSONL and/or SQLite, each client writes a run manifest (versions, git commit, configuration), and `rr.reproducible()` pins the decoding parameters.
7
+ - **Cost-safe**: a response cache means a restarted script only sends what is still missing, and per-request cost tracking with an optional budget stops spending before it gets out of hand, even across restarts.
8
+
9
+ ## Load-aware throttling in a nutshell
10
+
11
+ Large experiments can starve the other users of a shared inference server. Hard RPM limits are the usual answer, but they leave the server idle at night and still hurt others at peak times. `responsible-request` instead **watches the latency of your own requests**:
12
+
13
+ - While latency stays close to its baseline, the server is idle, so the client ramps up to `max_rpm`.
14
+ - As soon as latency reaches 3× the baseline (or the server returns 429/5xx/timeouts), other users are active, so the client drops to `min_rpm` at once.
15
+ - Once latency is back to normal, the client ramps up again.
16
+
17
+ For commercial APIs with generous rate limits of their own (OpenAI, OpenRouter, …), load-aware throttling is not needed: use a plain fixed-rate limiter, `rr.ThrottleConfig.fixed(rpm)`. In fixed mode a 429 does not change the rate: it pauses only the affected model (and endpoint) for the time in the `Retry-After` header, or for an exponential backoff (`backoff_initial_s`, doubling up to `backoff_max_s`) if there is none, plus a random jitter of up to `backoff_jitter_s`. The adaptive throttle instead treats a 429 as a sign of other users and drops to `min_rpm` for at least `cooldown_s`.
18
+
19
+ ## Installation
20
+
21
+ ```bash
22
+ pip install git+https://github.com/digital-sustainability/ResponsibleRequest.git
23
+ # or: uv add git+https://github.com/digital-sustainability/ResponsibleRequest.git
24
+ pip install "responsible-request[pandas,progress] @ git+https://github.com/digital-sustainability/ResponsibleRequest.git"
25
+ ```
26
+
27
+ Requires Python ≥ 3.10 and `openai` ≥ 1.40 (it works with both the `httpx`-based 1.x/2.x SDKs and the `httpx2`-based 3.x SDK).
28
+
29
+ ## Quickstart
30
+
31
+ ```python
32
+ import responsible_request as rr
33
+
34
+ client = rr.AsyncOpenAI( # same arguments as openai.AsyncOpenAI
35
+ base_url="https://inference.example.org/api/v1",
36
+ api_key="sk-...",
37
+ throttle=rr.ThrottleConfig(max_rpm=15, min_rpm=1),
38
+ log=rr.LogConfig(sqlite="requests.db"),
39
+ default_params=rr.reproducible(seed=42), # temperature=0, top_p=1, seed=42 unless set per call
40
+ )
41
+
42
+ response = await client.chat.completions.create(model="gpt-oss:120b", messages=[...])
43
+ print(client.throttle.stats())
44
+ await client.close()
45
+ ```
46
+
47
+ As with the OpenAI SDK, `base_url` and `api_key` can come from `OPENAI_BASE_URL` / `OPENAI_API_KEY` instead. For the examples in this repo, copy `.env.example` to `.env` and run them with `uv run --env-file .env python examples/quickstart.py`.
48
+
49
+ For existing code, only the HTTP client has to change:
50
+
51
+ ```python
52
+ import openai, responsible_request as rr
53
+
54
+ client = openai.AsyncOpenAI(http_client=rr.http_client(rr.ThrottleConfig(max_rpm=15)))
55
+ ```
56
+
57
+ It also works as a plain fixed-rate limiter for everyday use: `rr.ThrottleConfig.fixed(60)`.
58
+
59
+ With OpenRouter (or any other commercial endpoint), a typical setup is:
60
+
61
+ ```python
62
+ client = rr.AsyncOpenAI(
63
+ base_url="https://openrouter.ai/api/v1",
64
+ api_key=os.environ["OPENROUTER_API_KEY"],
65
+ throttle=rr.ThrottleConfig.fixed(120),
66
+ log=rr.LogConfig(sqlite="requests.db"),
67
+ cache=True, # a re-run only sends what is missing
68
+ cost=rr.CostConfig(budget_usd=5), # stop at $5, counting earlier runs in requests.db
69
+ run=rr.RunConfig(name="ablation-3"),
70
+ default_params=rr.reproducible(seed=42),
71
+ )
72
+ ```
73
+
74
+ See `examples/openrouter.py`.
75
+
76
+ ## Helpers
77
+
78
+ ```python
79
+ # Many requests, paced by the throttle, results in input order (exceptions returned, not raised)
80
+ results = await rr.run_batch(client, [{"model": m, "messages": msgs} for msgs in dataset], progress=True)
81
+ results = rr.run_batch_sync(client, requests) # from a plain script
82
+
83
+ # Structured output validated with Pydantic; the validation error is fed back to the model on failure
84
+ class Answer(BaseModel):
85
+ label: str
86
+ confidence: float
87
+
88
+ answer = await rr.structured(client, model=m, messages=msgs, schema=Answer, retries=2)
89
+ # (if `content` is empty, the answer is read from `reasoning`, then `reasoning_content`)
90
+ rr.response_format_from_model(Answer) # just the strict json_schema response_format
91
+
92
+ # System prompt marked for OpenRouter's prompt cache (cache_control: ephemeral)
93
+ messages = [rr.cached_system_message(long_instructions), {"role": "user", "content": question}]
94
+
95
+ # Attach metadata to every request record created inside the block
96
+ with rr.tags(experiment="ablation-3", fold=2):
97
+ ...
98
+
99
+ # Measure and pin the baseline explicitly, e.g. at night with a representative request
100
+ await rr.calibrate(client, "gpt-oss:120b", n=10, messages=representative_messages)
101
+ ```
102
+
103
+ `rr.cached_system_message()` only asks the provider to cache the prompt prefix (cheaper input tokens, the model still runs), whereas `cache=True` replays complete stored responses from your SQLite or JSONL log without sending the request.
104
+
105
+ ## How the throttle works
106
+
107
+ ```
108
+ ratio >= high_ratio (3.0) or 429/5xx/timeout
109
+ ┌────────────────────────────────────────────────────────────┐
110
+ │ ▼
111
+ WARMUP ──baseline known──► NORMAL ◄──at max_rpm── RECOVERING ◄── THROTTLED (min_rpm)
112
+ start_rpm ramp ×1.5 / 30 s ramp ×1.5 / 30 s cooldown ≥ 120 s and
113
+ up to max_rpm ratio < recover_ratio (1.5)
114
+ ```
115
+
116
+ - **Pacing**: requests are spaced evenly at the current rate with no bursts, and `max_concurrency` caps the number in flight. One `Throttle` can be shared by several clients, also across threads that each run their own event loop (e.g. `asyncio.run` per thread), and then enforces one rate for all of them:
117
+
118
+ ```python
119
+ throttle = rr.Throttle(rr.ThrottleConfig.fixed(60))
120
+ def worker(items): # runs in its own thread
121
+ asyncio.run(work(rr.AsyncOpenAI(throttle=throttle, ...), items))
122
+ ```
123
+ - **Per model**: every (endpoint, model) pair has its own lane, because load on one backend says nothing about another. Only `/chat/completions` and `/embeddings` feed the latency estimate (`observe_paths`). Other endpoints (audio, …) are paced and react to errors only. Requests with server-side tools (MCP, web search) are ignored for load estimation, because their latency includes external calls.
124
+ - **Signal**: by default, latency per completion token (`latency / max(completion_tokens, 16)`), so long answers and reasoning traces don't look like load. For streamed requests the signal is the time to first byte. The current value is the median of the last `window` requests.
125
+ - **Baseline**: the 10th percentile of the signal over the last 30 minutes, needing `warmup_requests` samples first. Samples taken while throttled are excluded, so a long busy period does not become the new "normal". You can pin it with `ThrottleConfig(baseline=...)` or `rr.calibrate(...)`, and reset it with `client.throttle.reset_baseline()`.
126
+ - **Dead band**: between `recover_ratio` and `high_ratio` the rate stays where it is, which prevents oscillation.
127
+
128
+ `examples/simulate_load.py` runs the whole loop against a simulated server (time-compressed) where other users appear for a while:
129
+
130
+ ```
131
+ t [s] others RPM state load
132
+ 8.5 0 1800 normal 1.00
133
+ 12.5 48 1800 normal 1.00
134
+ 13.5 48 60 throttled 3.56
135
+ ...
136
+ 30.6 0 60 recovering 1.00
137
+ 41.6 0 1800 normal 1.00
138
+ ```
139
+
140
+ ### Configuration (`ThrottleConfig`)
141
+
142
+ | Parameter | Default | Meaning |
143
+ |---|---|---|
144
+ | `max_rpm` | 15 | rate when the endpoint is idle |
145
+ | `min_rpm` | 1 | rate while others are active (also the probing rate) |
146
+ | `start_rpm` | 2 | rate while the baseline is being established |
147
+ | `max_concurrency` | 32 | max requests in flight per model |
148
+ | `high_ratio` | 3.0 | throttle when latency ≥ this × baseline |
149
+ | `recover_ratio` | 1.5 | ramp up only when latency < this × baseline |
150
+ | `cooldown_s` | 120 | minimum time at `min_rpm` after the last high-load signal |
151
+ | `ramp_factor` / `ramp_interval_s` | 1.5 / 30 | multiplicative ramp-up step and its interval |
152
+ | `metric` | `"latency_per_token"` | `"latency"`, `"ttfb"`, or a callable `RequestRecord -> float` |
153
+ | `window` | 10 | requests in the rolling median |
154
+ | `warmup_requests` | 20 | samples needed before the baseline is trusted |
155
+ | `baseline` | None | pin the baseline instead of estimating it |
156
+ | `baseline_percentile` / `baseline_window_s` | 10 / 1800 | baseline estimator |
157
+ | `observe_paths` | chat, embeddings | endpoints whose latency is used |
158
+ | `estimator_factory` | None | plug in your own `LoadEstimator` (e.g. reading queue depth from Prometheus) |
159
+ | `backoff_initial_s` / `backoff_max_s` | 1 / 60 | fixed mode: pause after a 429 without `Retry-After`, doubling per further 429 |
160
+ | `backoff_jitter_s` | 1 | fixed mode: random extra time added to every 429 pause |
161
+
162
+ ## Logging
163
+
164
+ The package logs through [loguru](https://github.com/Delgan/loguru) and, as loguru recommends for libraries, stays silent until a client is created with `log=` (the default `log=True` enables console messages only).
165
+
166
+ - **Console**: rate changes (INFO), switches to throttling (WARNING), and a per-model summary every `summary_interval_s`. Individual requests are never printed. With `LogConfig(console_level="INFO")` the package replaces loguru's default DEBUG handler with a quieter one.
167
+ - **Request records**: one record per HTTP attempt (SDK retries are separate records with `attempt` > 0) goes to `LogConfig(jsonl=...)` and/or `LogConfig(sqlite=...)` (table `requests`). Records are emitted at loguru's TRACE level with `extra["rr_record"]`, so you can also attach your own sinks.
168
+
169
+ | Group | Fields |
170
+ |---|---|
171
+ | core (always) | `request_id`, `timestamp`, `method`, `path`, `model`, `stream`, `attempt`, `status_code`, `error`, `tags`, `cache_key`, `cache_hit`, `run_id` |
172
+ | `timing` | `sent_at`, `first_byte_at`, `finished_at`, `wait_s` (time spent throttled), `ttfb_s`, `latency_s` |
173
+ | `usage` | `prompt_tokens`, `completion_tokens`, `total_tokens`, `cached_tokens`, `reasoning_tokens`, `cost_usd` |
174
+ | `throttle` | `rpm`, `state`, `baseline`, `load_ratio`, `in_flight` |
175
+ | `response_meta` | `response_id`, `response_model`, `finish_reason`, `system_fingerprint`, `provider`, `gateway_request_id`, `upstream_duration_ms`, `provider_meta` (see below) |
176
+ | `params` | all request parameters except messages/input |
177
+ | `request_body` / `response_body` | full request and response (for streams: the concatenated content) |
178
+
179
+ All groups are logged by default. Turn groups off with `LogConfig(fields={"request_body": False})`, or use `LogConfig.minimal()` (no params or bodies). **Request and response bodies can be large and may contain sensitive data.**
180
+
181
+ ```python
182
+ df = rr.load_records("requests.db") # pandas DataFrame if pandas is installed, else list of dicts
183
+ ```
184
+
185
+ JSONL logs rotate at `rotation="500 MB"` (any loguru rotation). LLM logs are repetitive and compress well, so finished segments can be compressed with `compression="gz"`, `"bz2"`, `"xz"` or `"zst"` (zstd needs Python 3.14+ or `pip install responsible-request[zstd]`); the file being written stays plain. Without rotation, the file is compressed when the client's sinks are removed. `load_records("requests.jsonl")`, the cache and the cost budget read all segments, compressed or not, and `load_records` / `CacheConfig(jsonl=...)` also accept a single archive such as `requests.2026-10-06_12-00-00_000000.jsonl.zst`.
186
+
187
+ ```python
188
+ log = rr.LogConfig(jsonl="requests.jsonl", rotation="100 MB", compression="zst")
189
+ ```
190
+
191
+ ### Provider metadata
192
+
193
+ Gateways and routers report different extra information. *Extractors* map it to the generic `response_meta` fields and put everything else into the `provider_meta` JSON column. All built-in extractors run on every response and only fill in what they find, so nothing has to be configured:
194
+
195
+ | Extractor | Reads |
196
+ |---|---|
197
+ | `openai_compat` | `x-request-id` / `request-id` → `gateway_request_id`, `openai-processing-ms` → `upstream_duration_ms` |
198
+ | `litellm` | `x-litellm-call-id` → `gateway_request_id`, `x-litellm-response-duration-ms` → `upstream_duration_ms`, `x-litellm-response-cost` → `cost_usd`, all other `x-litellm-*` headers → `provider_meta["litellm"]` |
199
+ | `openrouter` | `provider` → `provider`, `usage.cost` → `cost_usd`, `usage.is_byok`/`cost_details` and `native_finish_reason` → `provider_meta["openrouter"]` |
200
+
201
+ Add your own for other endpoints:
202
+
203
+ ```python
204
+ def my_gateway(record, headers, body): # body: parsed JSON, stream summary, or None
205
+ if "x-queue-ms" in headers:
206
+ rr.providers.meta(record, "my_gateway")["queue_ms"] = float(headers["x-queue-ms"])
207
+
208
+ client = rr.AsyncOpenAI(..., extractors=[my_gateway])
209
+ ```
210
+
211
+ Databases written by version 0.1 keep their `litellm_*` columns; new records leave them empty.
212
+
213
+ ### Runs
214
+
215
+ Every client is a *run* with its own `run_id`, which every record carries. With a JSONL or SQLite log, the first request also writes a run manifest: the run name and `metadata`, start time, endpoint (scheme and host only, never the API key), versions of `responsible-request`, `openai` and Python, platform, hostname, command line, working directory, git commit and dirty flag, and the complete throttle, log, cache and cost configuration plus `default_params`.
216
+
217
+ ```python
218
+ client = rr.AsyncOpenAI(..., run=rr.RunConfig(name="ablation-3", metadata={"dataset": "v2"}))
219
+ # or simply run="ablation-3"
220
+
221
+ runs = rr.load_runs("requests.db") # the `runs` table, or requests.runs.jsonl for a JSONL log
222
+ df = rr.load_records("requests.db").merge(runs, on="run_id", suffixes=("", "_run"))
223
+ ```
224
+
225
+ ## Caching
226
+
227
+ With `cache=True`, a request that was already answered successfully is served from the logged records instead of being sent again. Re-running a script (e.g. after a crash) then only sends the requests that are still missing:
228
+
229
+ ```python
230
+ client = rr.AsyncOpenAI(
231
+ log=rr.LogConfig(sqlite="requests.db"),
232
+ cache=True, # or rr.CacheConfig(...)
233
+ )
234
+ ```
235
+
236
+ - **Key**: a hash of the URL and the complete JSON request body (model, messages and all parameters, after `default_params` are applied). Any change to the prompt or parameters is a miss.
237
+ - **Source**: the records the client writes itself (`LogConfig.sqlite`, else `LogConfig.jsonl`; `response_body` must be logged), or another database or file via `rr.CacheConfig(sqlite=...)` / `rr.CacheConfig(jsonl=...)`.
238
+ - **What is served**: only non-streamed requests whose record has status 200 and no error. Streamed requests are always sent.
239
+ - **Hits** don't wait for or affect the throttle. They are logged with `cache_hit=True` (filter them out when analysing latency or token usage), and the response carries an `x-rr-cache: hit` header. `client.cache.stats()` counts hits and misses.
240
+ - **Repeated sampling**: identical requests share one cached answer. To draw several independent samples, name the tags that belong to the key:
241
+
242
+ ```python
243
+ client = rr.AsyncOpenAI(log=..., cache=rr.CacheConfig(key_tags=("sample",)))
244
+ for i in range(5):
245
+ with rr.tags(sample=i): # 5 different keys; a re-run gets the same 5 answers
246
+ await client.chat.completions.create(model=m, messages=msgs, temperature=0.7)
247
+ ```
248
+
249
+ Other tags (e.g. `experiment=...`) are not part of the key. Records are written asynchronously, so an identical request sent a few milliseconds after the first one may still go to the server.
250
+
251
+ ## Cost and budget
252
+
253
+ With `cost=True` or a `rr.CostConfig`, every record gets a `cost_usd`:
254
+
255
+ 1. the cost the provider reports (OpenRouter `usage.cost`, LiteLLM `x-litellm-response-cost`, or a custom extractor), else
256
+ 2. the cost computed from token usage and `CostConfig(prices={"model": rr.Price(input=..., output=..., cached_input=...)})` (USD per million tokens; there is no built-in price list, since it would go stale), else
257
+ 3. `None`.
258
+
259
+ Cache hits cost `0.0`. `client.cost.stats()` shows what was spent, and the per-model summary line and `client.throttle.stats()` include it.
260
+
261
+ ```python
262
+ client = rr.AsyncOpenAI(..., log=rr.LogConfig(sqlite="requests.db"), cost=rr.CostConfig(budget_usd=5))
263
+ ```
264
+
265
+ - Once `budget_usd` has been spent, new requests raise `rr.BudgetExceeded` instead of being sent (`run_batch` returns it for each remaining item). Cache hits are still served, so re-running a finished experiment works with an exhausted budget.
266
+ - With `include_logged=True` (the default), the spend already recorded in the log counts too, so restarting a script after a crash does not reset the budget. This sums **all** records in the log file, so use one database per experiment or budget.
267
+ - Requests already in flight when the budget runs out still complete, so the budget can be exceeded by their cost.
268
+ - If a budget is set but a model's responses carry no cost and no price is configured, a warning is logged once: the budget cannot see those requests.
269
+ - With `openai<2`, the SDK retries requests after an exception, so a blocked request is attempted `max_retries` more times (each blocked again, without a network call) before `rr.BudgetExceeded` is raised.
270
+
271
+ ## Caveats
272
+
273
+ - **Your own load raises latency too.** If `max_rpm` is high enough to saturate the server by itself, the client will throttle itself. The low-percentile baseline and the dead band absorb moderate self-load. Choose `max_rpm`/`max_concurrency` so that you alone don't saturate the backend.
274
+ - **One process, one event loop per client.** Throttle state is not shared between processes. Several scripts running in parallel each throttle independently.
275
+ - **Detection latency.** At `min_rpm=1` and `window=10`, noticing that others have left takes about ten minutes plus the cooldown. This is deliberate: the client stays polite longer than strictly needed.
276
+ - If the backend becomes permanently slower (e.g. a model is redeployed on smaller hardware), call `client.throttle.reset_baseline()`.
277
+
278
+ ## Development
279
+
280
+ ```bash
281
+ uv sync
282
+ uv run pytest
283
+ uv run ruff check . && uv run mypy src
284
+ uv run python examples/simulate_load.py
285
+ ```