errorbars 0.1.2__py3-none-any.whl

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -0,0 +1,297 @@
1
+ Metadata-Version: 2.5
2
+ Name: errorbars
3
+ Version: 0.1.2
4
+ Summary: Error bars for LLM evals: standard errors, clustering, paired comparisons, leaderboards, and power analysis.
5
+ Project-URL: Homepage, https://antonsoo.github.io/errorbars/
6
+ Project-URL: Repository, https://github.com/antonsoo/errorbars
7
+ Project-URL: Issues, https://github.com/antonsoo/errorbars/issues
8
+ Author-email: Anton Soloviev <anton@praviel.com>
9
+ License-Expression: MIT
10
+ License-File: LICENSE
11
+ Keywords: benchmarking,confidence-intervals,evaluation,language-models,llm-evals,statistics
12
+ Classifier: Development Status :: 4 - Beta
13
+ Classifier: Intended Audience :: Science/Research
14
+ Classifier: License :: OSI Approved :: MIT License
15
+ Classifier: Programming Language :: Python :: 3
16
+ Classifier: Programming Language :: Python :: 3.10
17
+ Classifier: Programming Language :: Python :: 3.11
18
+ Classifier: Programming Language :: Python :: 3.12
19
+ Classifier: Topic :: Scientific/Engineering :: Artificial Intelligence
20
+ Classifier: Typing :: Typed
21
+ Requires-Python: >=3.10
22
+ Requires-Dist: numpy>=1.24
23
+ Provides-Extra: all
24
+ Requires-Dist: matplotlib>=3.7; extra == 'all'
25
+ Requires-Dist: pandas>=1.5; extra == 'all'
26
+ Requires-Dist: rich>=13.0; extra == 'all'
27
+ Provides-Extra: cli
28
+ Requires-Dist: rich>=13.0; extra == 'cli'
29
+ Provides-Extra: inspect
30
+ Requires-Dist: agent-client-protocol<1.0,>=0.12; extra == 'inspect'
31
+ Requires-Dist: httpx<1.0,>=0.27; extra == 'inspect'
32
+ Requires-Dist: inspect-ai>=0.3.268; extra == 'inspect'
33
+ Provides-Extra: pandas
34
+ Requires-Dist: pandas>=1.5; extra == 'pandas'
35
+ Provides-Extra: plot
36
+ Requires-Dist: matplotlib>=3.7; extra == 'plot'
37
+ Description-Content-Type: text/markdown
38
+
39
+ # errorbars
40
+
41
+ **Error bars for LLM evals. Know whether model B is actually better, or you're reading noise.**
42
+
43
+ [![License: MIT](https://img.shields.io/badge/license-MIT-blue.svg)](https://github.com/antonsoo/errorbars/blob/main/LICENSE)
44
+ [![Live demo](https://img.shields.io/badge/live%20demo-calculator-2454a6)](https://antonsoo.github.io/errorbars/)
45
+ [![Hugging Face](https://img.shields.io/badge/Hugging%20Face-workbench-ffd21e)](https://huggingface.co/spaces/antonsoloviev/errorbars)
46
+
47
+ LLM eval results get reported as bare accuracies — "model B scored 71.2%, up from 69.7%" — with
48
+ no error bar, no significance test, and no accounting for the fact that the questions came in
49
+ correlated groups. On a 500-question benchmark, a 1.5-point gain is very often noise: at a 50%
50
+ baseline you'd need roughly a 9-point gap to reliably detect anything at 80% power, α=0.05
51
+ (`errorbars power --n 500 --baseline 0.5` → minimum detectable effect ≈ 8.9 points). Evan Miller's
52
+ "[Adding Error Bars to Evals: A Statistical Approach to Language Model
53
+ Evaluations](https://arxiv.org/abs/2411.00640)" (arXiv:2411.00640, 2024) lays out the right
54
+ practice — CLT and clustered standard errors, variance reduction through pairing, power analysis —
55
+ and this package turns it into a one-liner.
56
+
57
+ ![errorbars leaderboard on a synthetic clustered benchmark](https://raw.githubusercontent.com/antonsoo/errorbars/main/docs/assets/leaderboard-terminal.png)
58
+
59
+ ## Why this exists
60
+
61
+ Run `errorbars leaderboard` on a synthetic benchmark (below) and one apparent 7-point win —
62
+ `tuned-70b` over `baseline-70b`, 0.690 vs. 0.620 — is not statistically significant once it has a
63
+ proper error bar: the paired test already gives p = 0.11. Accounting for the fact that questions
64
+ come 5-to-a-passage, and Holm-correcting across all 6 pairwise comparisons on the board, pushes
65
+ that to p = 0.32. A naive leaderboard that ranks by bare accuracy would still have shipped the gap
66
+ as a win. The full walkthrough, with every number copied from a real command, is in
67
+ [`examples/README.md`](https://github.com/antonsoo/errorbars/blob/main/examples/README.md).
68
+
69
+ ## Quickstart
70
+
71
+ ```bash
72
+ pip install "errorbars[cli]"
73
+ errorbars power --delta 0.03 --baseline 0.5
74
+ ```
75
+
76
+ ```
77
+ Questions needed: 4361
78
+ ```
79
+
80
+ That's how many questions you'd need to reliably detect a 3-point accuracy gap at a 50% baseline,
81
+ with the defaults (80% power, α=0.05, no pairing correlation, no clustering). Clone the repo to
82
+ try it against real per-item scores:
83
+
84
+ ```bash
85
+ git clone https://github.com/antonsoo/errorbars && cd errorbars
86
+ errorbars leaderboard examples/data/reading_comprehension.csv
87
+ ```
88
+
89
+ ## Features
90
+
91
+ - **`summarize`** — mean, SE, and 95% CI (CLT by default, Wilson for small-n binary scores,
92
+ bootstrap on request); clustered SE with design effect and ICC when a `cluster_id` column is
93
+ present; within/between-question variance decomposition when a `sample` column is present.
94
+ - **`compare`** — paired mean difference, SE, CI, p-value; the correlation between the two models'
95
+ per-question scores and how much pairing shrank the SE vs. an unpaired comparison; a
96
+ cluster-robust paired SE/p-value when clusters are present; exact McNemar test for binary scores.
97
+ - **`leaderboard`** — every model with its CI, Holm-corrected pairwise paired tests, and groups of
98
+ statistically indistinguishable models (maximal cliques of the "not significantly different"
99
+ graph); a forest plot (SVG, no dependency; matplotlib if installed).
100
+ - **`power`** — number of questions needed to detect an effect δ at a given α and power, or the
101
+ minimum detectable effect for a given n, accounting for pairing correlation, repeated sampling,
102
+ and cluster design effect.
103
+ - **CLI** — `errorbars summarize|compare|leaderboard|power|import`, `rich` tables by default,
104
+ `--json` for scripting.
105
+ - **Adapters** — `errorbars import lm-eval|inspect` converts lm-evaluation-harness `--log_samples`
106
+ output or an Inspect AI `.eval` log into the canonical format (see below).
107
+ - **Web calculator** — a static "how many eval questions do I need?" power calculator
108
+ ([live demo](https://antonsoo.github.io/errorbars/)), whose TypeScript formulas are checked
109
+ against the Python ones by the test suite.
110
+ - Typed Python ≥3.10, `numpy` the only runtime dependency; `pandas` and `matplotlib` are optional
111
+ extras; `scipy`/`statsmodels` are test-only oracles, never imported at runtime.
112
+
113
+ ## Usage
114
+
115
+ ```bash
116
+ errorbars summarize examples/data/reading_comprehension.csv --model tuned-70b
117
+ ```
118
+
119
+ ```
120
+ summarize: tuned-70b
121
+ ┏━━━━━━━━┳━━━━━━━━━━━━━━━━━━┓
122
+ ┃ metric ┃ value ┃
123
+ ┡━━━━━━━━╇━━━━━━━━━━━━━━━━━━┩
124
+ │ n │ 200 │
125
+ │ mean │ 0.6900 │
126
+ │ SE │ 0.0328 │
127
+ │ 95% CI │ [0.6257, 0.7543] │
128
+ │ method │ clt │
129
+ └────────┴──────────────────┘
130
+ clustering diagnostics
131
+ ┏━━━━━━━━━━━━━━━┳━━━━━━━━━━━━━━━━━━┓
132
+ ┃ metric ┃ value ┃
133
+ ┡━━━━━━━━━━━━━━━╇━━━━━━━━━━━━━━━━━━┩
134
+ │ n clusters │ 40 │
135
+ │ ICC │ 0.0710 │
136
+ │ design effect │ 1.284 │
137
+ │ clustered SE │ 0.0372 │
138
+ │ clustered CI │ [0.6148, 0.7652] │
139
+ └───────────────┴──────────────────┘
140
+ ```
141
+
142
+ ```bash
143
+ errorbars compare examples/data/reading_comprehension.csv \
144
+ --model-a tuned-70b --model-b baseline-70b
145
+ ```
146
+
147
+ ```
148
+ mean diff (A - B) 0.0700
149
+ paired SE 0.0440
150
+ 95% CI [-0.0167, 0.1567]
151
+ p-value 0.1131
152
+ correlation(A, B) 0.1434
153
+ variance reduction from pairing 14.3%
154
+ McNemar exact p-value 0.1405
155
+ ```
156
+
157
+ ```bash
158
+ errorbars power --delta 0.05 --baseline 0.65 --rho 0.14 --cluster-deff 1.28
159
+ ```
160
+
161
+ ```
162
+ Questions needed: 1573
163
+ ```
164
+
165
+ Every number above is copied verbatim from running these commands against
166
+ `examples/data/reading_comprehension.csv` (synthetic — see below). The full leaderboard output and
167
+ the reasoning behind each step are in [`examples/README.md`](https://github.com/antonsoo/errorbars/blob/main/examples/README.md).
168
+
169
+ ### Input format
170
+
171
+ Long-format CSV or JSONL, one row per observation:
172
+
173
+ | question_id | cluster_id (optional) | model | score | sample (optional) |
174
+ |---|---|---|---|---|
175
+ | q1 | passage-003 | tuned-70b | 1 | |
176
+
177
+ Column names are configurable (`--question-col`, `--cluster-col`, etc., or `ColumnMap` in Python).
178
+ Scores can be binary (0/1) or continuous. `cluster_id` groups correlated questions (e.g. several
179
+ questions per reading passage); `sample` marks repeated generations of the same question.
180
+
181
+ ### Importing from lm-evaluation-harness or Inspect AI
182
+
183
+ `errorbars import` converts either tool's own log format into the canonical CSV above:
184
+
185
+ ```bash
186
+ # lm-evaluation-harness: lm_eval run --model <...> --tasks <...> --log_samples --output_path <dir>
187
+ errorbars import lm-eval runs/<dir>/samples_<task>_<timestamp>.jsonl \
188
+ --model my-model-name -o converted.csv
189
+
190
+ # Inspect AI: inspect eval <task> --model <...> (writes a .eval log by default)
191
+ errorbars import inspect logs/<run>.eval -o converted.csv
192
+ ```
193
+
194
+ `--metric` (lm-eval) picks which computed metric to use as the score when a task reports more than
195
+ one (e.g. `acc` vs. `acc_norm`); it defaults to the first one. `--scorer` (Inspect) does the same
196
+ for tasks with multiple scorers. Inspect epochs (`--epochs N`, repeated sampling of the same input)
197
+ land in the `sample` column automatically.
198
+
199
+ Both adapters were built and tested against real, unedited output — not from memory: **lm-eval
200
+ 0.4.13** (`lm_eval run --model dummy --tasks copa --limit 20 --log_samples ...`, plus a
201
+ multi-metric `arc_easy` run) and **inspect-ai 0.3.268** (a 5-sample task through the built-in
202
+ `mockllm/model` provider). The exact log files are committed as test fixtures
203
+ (`tests/fixtures/samples_*.jsonl`, `tests/fixtures/inspect_tiny_qa.eval`) and re-parsed in
204
+ `tests/test_adapters.py` on every run. The Inspect adapter reads logs with Inspect's own
205
+ `inspect_ai.log.read_eval_log` and `inspect_ai.scorer.value_to_float` rather than hand-parsing its
206
+ binary `.eval` format, and needs the `inspect` extra:
207
+ `pip install "errorbars[inspect]"`. The lm-eval
208
+ adapter has no extra dependency — `--log_samples` is already plain JSONL.
209
+
210
+ ## How it works
211
+
212
+ Every statistic is implemented from scratch on `numpy` + the standard library
213
+ (`statistics.NormalDist` for normal quantiles; Student-t tails from a continued-fraction incomplete
214
+ beta function) — no scipy or statsmodels at runtime.
215
+ Full derivations with references are in [`docs/formulas.md`](https://github.com/antonsoo/errorbars/blob/main/docs/formulas.md):
216
+
217
+ 1. CLT and Wilson confidence intervals for a mean
218
+ 2. Cluster-robust standard errors (the CR1 sandwich estimator, matching
219
+ `statsmodels`' `cov_type="cluster"`)
220
+ 3. Intraclass correlation and Kish's design effect
221
+ 4. Within/between-question variance decomposition for repeated sampling
222
+ 5. Paired comparisons, variance reduction from pairing, and exact McNemar
223
+ 6. Holm-Bonferroni correction and maximal-clique grouping for leaderboards
224
+ 7. Power analysis (sample size and minimum detectable effect) under pairing and clustering
225
+
226
+ ## Accuracy and limitations
227
+
228
+ - Cluster-robust SE matches `statsmodels`' `OLS(..., cov_type="cluster")` to 1e-9 on both balanced
229
+ and unbalanced cluster sizes; Wilson intervals match `statsmodels.stats.proportion_confint` to
230
+ 1e-9; the paired t-test matches `scipy.stats.ttest_rel` to 1e-7 at any n, and McNemar's exact
231
+ test is cross-checked against `statsmodels`.
232
+ See `tests/test_stats_vs_oracles.py` and `tests/test_compare_vs_oracles.py`.
233
+ - Monte Carlo coverage tests (`tests/test_coverage_montecarlo.py`) confirm nominal 95% CIs cover
234
+ the true parameter close to 95% of the time — for CLT, Wilson, bootstrap, and cluster-robust
235
+ intervals — with a tolerance sized to the trial count so it won't flake.
236
+ - Paired comparisons use Student's t with n − 1 degrees of freedom, like `scipy.stats.ttest_rel`,
237
+ and clustered SEs use t with G − 1 for G clusters (Cameron & Miller 2015), computed without
238
+ scipy. With few clusters that interval is honest but wide; the per-model `clt` interval stays
239
+ normal-based, which is only right once n is in the dozens.
240
+ - The power formula's samples-per-question adjustment assumes all single-sample variance is
241
+ decoding noise (see `docs/formulas.md` §11 for why, and the caveat on when this is optimistic).
242
+ It's a planning tool for before you run the eval; for a post-hoc measurement with the true
243
+ within/between-question split, use `summarize` on data with a `sample` column instead.
244
+ - Leaderboard groups are maximal cliques of the "not significantly different" graph, which is the
245
+ statistically direct approach — it can produce a model in more than one group, unlike a
246
+ minimal-letters heuristic (e.g. R's `multcompView`).
247
+ - 168 Python tests, 605 TypeScript tests (cross-checking the JS formulas against Python-generated
248
+ vectors), all passing on this box (14 vCPU WSL2 Linux, 48 GB RAM; Python 3.12.3, Node 26.7.0)
249
+ as of 2026-10-01.
250
+
251
+ ## Web calculator
252
+
253
+ [`web/`](https://github.com/antonsoo/errorbars/blob/main/web/) is a static Vite + TypeScript "how many eval questions do I need?" calculator
254
+ ([live demo](https://antonsoo.github.io/errorbars/)). Its power formulas
255
+ (`web/src/power.ts`, `web/src/normal.ts`) are a direct port of `src/errorbars/power.py`; Python
256
+ generates JSON test vectors (`scripts/export_test_vectors.py` → `web/test-vectors.json`) that the
257
+ TypeScript test suite checks against, so the two implementations can't silently drift apart.
258
+
259
+ ![Power calculator](https://raw.githubusercontent.com/antonsoo/errorbars/main/docs/assets/calculator-hero.png)
260
+
261
+ ## Development
262
+
263
+ ```bash
264
+ uv sync --group dev --extra all
265
+ uv run pytest
266
+ uv run ruff check .
267
+ uv run mypy src/errorbars
268
+ ```
269
+
270
+ See [`CONTRIBUTING.md`](https://github.com/antonsoo/errorbars/blob/main/CONTRIBUTING.md) for the web calculator's dev loop and guidelines.
271
+
272
+ ## Contributing
273
+
274
+ Issues and PRs welcome. Any new statistic needs a test against an independent oracle (reference
275
+ library, closed form, or Monte Carlo simulation) — see `tests/` for the pattern.
276
+
277
+ ## Citations
278
+
279
+ ```bibtex
280
+ @misc{miller2024addingerrorbarsevals,
281
+ title = {Adding Error Bars to Evals: A Statistical Approach to Language Model Evaluations},
282
+ author = {Evan Miller},
283
+ year = {2024},
284
+ eprint = {2411.00640},
285
+ archivePrefix = {arXiv},
286
+ primaryClass = {stat.AP},
287
+ url = {https://arxiv.org/abs/2411.00640}
288
+ }
289
+ ```
290
+
291
+ ## License
292
+
293
+ [MIT](https://github.com/antonsoo/errorbars/blob/main/LICENSE) © 2026 Anton Soloviev
294
+
295
+ ---
296
+
297
+ <sub>Part of [Officina](https://antonsoo.github.io/officina/), a set of small open-source tools by [Anton Soloviev](https://github.com/antonsoo).</sub>
@@ -0,0 +1,17 @@
1
+ errorbars/__init__.py,sha256=CRp6GSNgD1xpOmJLHMrAxylNPVnKH6EmhTmfAME_Ol0,547
2
+ errorbars/cli.py,sha256=ubQF9prXpRBAq02cgL1IXP1GJBzDm8BsrvtzSAQz5gY,17448
3
+ errorbars/compare.py,sha256=OZ8clRwaL6ek6KTBRN57_Q3utoIEMKULCKsY97dV-p4,6602
4
+ errorbars/io.py,sha256=a7R4HDQqh5jVkXFLuy8EP5nG-cnCTh1AYxxjE0Xwub0,7615
5
+ errorbars/leaderboard.py,sha256=K19pskSFBqMqr03yC1BIof1EVjg5xDeBhCtCGTsQK3k,6492
6
+ errorbars/plot.py,sha256=NoWT40_OfzdveHLqEw8CPV8HGUcGB98eS3HgA2GA48M,5041
7
+ errorbars/power.py,sha256=1GNpBxR5VpW6xGQLEhIppepjowxp2rQi9aLiqH4m7tM,5429
8
+ errorbars/py.typed,sha256=47DEQpj8HBSa-_TImW-5JCeuQeRkm5NMpJWZG3hSuFU,0
9
+ errorbars/stats.py,sha256=lo0x4TY0Y2UCjLO-3i3Q5WFEZnd-LauDLhpXy763YSM,11807
10
+ errorbars/adapters/__init__.py,sha256=XhWjHhf5OQrODBysnjKDAu_gjX_gY9qCVc6m0B5j1Z8,77
11
+ errorbars/adapters/inspect_ai.py,sha256=5_IACOHXvpF5sJjxj7SzSLlKcBUaFmJskRoGoaU4UvQ,3189
12
+ errorbars/adapters/lm_eval.py,sha256=BOrt30jp5TkiVMRifw9x3I2kkX9XzsXoDTJFhkLQE9o,4023
13
+ errorbars-0.1.2.dist-info/METADATA,sha256=FwiEtauE3MkQuEYwjiTVEOzv4Vsg1wfSdeurfaA38yc,14774
14
+ errorbars-0.1.2.dist-info/WHEEL,sha256=W3fkpkm7-wf9vBI5Z-7s0eWkeM-spu78I8Neb98DeEg,87
15
+ errorbars-0.1.2.dist-info/entry_points.txt,sha256=4tLng6AuvkO_nJeSf4KDkgsGXsEqQcZGkIcjkdmG45E,49
16
+ errorbars-0.1.2.dist-info/licenses/LICENSE,sha256=4mFEvoZz1pwXUSatjc2e14QxRk8LlxILeZ3HXsfttkc,1071
17
+ errorbars-0.1.2.dist-info/RECORD,,
@@ -0,0 +1,4 @@
1
+ Wheel-Version: 1.0
2
+ Generator: hatchling 1.32.4
3
+ Root-Is-Purelib: true
4
+ Tag: py3-none-any
@@ -0,0 +1,2 @@
1
+ [console_scripts]
2
+ errorbars = errorbars.cli:main
@@ -0,0 +1,21 @@
1
+ MIT License
2
+
3
+ Copyright (c) 2026 Anton Soloviev
4
+
5
+ Permission is hereby granted, free of charge, to any person obtaining a copy
6
+ of this software and associated documentation files (the "Software"), to deal
7
+ in the Software without restriction, including without limitation the rights
8
+ to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
9
+ copies of the Software, and to permit persons to whom the Software is
10
+ furnished to do so, subject to the following conditions:
11
+
12
+ The above copyright notice and this permission notice shall be included in all
13
+ copies or substantial portions of the Software.
14
+
15
+ THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
16
+ IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
17
+ FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
18
+ AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
19
+ LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
20
+ OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
21
+ SOFTWARE.