evalstats 0.2.0__tar.gz → 0.2.4__tar.gz

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (70) hide show
  1. evalstats-0.2.4/PKG-INFO +548 -0
  2. evalstats-0.2.4/README.md +494 -0
  3. {evalstats-0.2.0 → evalstats-0.2.4}/evalstats/__init__.py +25 -11
  4. evalstats-0.2.4/evalstats/alignment.py +804 -0
  5. evalstats-0.2.4/evalstats/api.py +1953 -0
  6. {evalstats-0.2.0 → evalstats-0.2.4}/evalstats/cli.py +32 -5
  7. evalstats-0.2.4/evalstats/config.py +314 -0
  8. {evalstats-0.2.0 → evalstats-0.2.4}/evalstats/core/bundles.py +28 -0
  9. {evalstats-0.2.0 → evalstats-0.2.4}/evalstats/core/mixed_effects.py +330 -0
  10. {evalstats-0.2.0 → evalstats-0.2.4}/evalstats/core/paired.py +746 -287
  11. {evalstats-0.2.0 → evalstats-0.2.4}/evalstats/core/ranking.py +106 -3
  12. {evalstats-0.2.0 → evalstats-0.2.4}/evalstats/core/resampling.py +800 -74
  13. {evalstats-0.2.0 → evalstats-0.2.4}/evalstats/core/router.py +175 -39
  14. evalstats-0.2.4/evalstats/core/stats_utils.py +215 -0
  15. {evalstats-0.2.0 → evalstats-0.2.4}/evalstats/core/summary.py +563 -159
  16. {evalstats-0.2.0 → evalstats-0.2.4}/evalstats/core/types.py +1 -1
  17. {evalstats-0.2.0 → evalstats-0.2.4}/evalstats/core/variance.py +107 -15
  18. evalstats-0.2.4/evalstats/loader.py +697 -0
  19. evalstats-0.2.4/evalstats/ppi.py +941 -0
  20. evalstats-0.2.4/evalstats/tests/__init__.py +4412 -0
  21. {evalstats-0.2.0 → evalstats-0.2.4}/evalstats/vis/critical_difference.py +0 -1
  22. {evalstats-0.2.0 → evalstats-0.2.4}/evalstats/vis/forest.py +3 -5
  23. {evalstats-0.2.0 → evalstats-0.2.4}/evalstats/vis/scoreboard.py +8 -9
  24. evalstats-0.2.4/evalstats.egg-info/PKG-INFO +548 -0
  25. {evalstats-0.2.0 → evalstats-0.2.4}/evalstats.egg-info/SOURCES.txt +10 -3
  26. {evalstats-0.2.0 → evalstats-0.2.4}/pyproject.toml +1 -1
  27. evalstats-0.2.4/tests/test_alignment.py +995 -0
  28. {evalstats-0.2.0 → evalstats-0.2.4}/tests/test_analyze.py +24 -1
  29. evalstats-0.2.4/tests/test_auto_ci_routing.py +317 -0
  30. {evalstats-0.2.0 → evalstats-0.2.4}/tests/test_bayes_binary_routing.py +24 -41
  31. evalstats-0.2.4/tests/test_bootstrap_t_pairwise_ranking.py +39 -0
  32. {evalstats-0.2.0 → evalstats-0.2.4}/tests/test_cli.py +117 -10
  33. evalstats-0.2.4/tests/test_compare.py +409 -0
  34. {evalstats-0.2.0 → evalstats-0.2.4}/tests/test_critical_difference_plot.py +10 -9
  35. evalstats-0.2.4/tests/test_ppi_core.py +151 -0
  36. evalstats-0.2.4/tests/test_ppi_corrections.py +2704 -0
  37. {evalstats-0.2.0 → evalstats-0.2.4}/tests/test_resampling.py +2 -2
  38. {evalstats-0.2.0 → evalstats-0.2.4}/tests/test_scoreboard_plot.py +23 -17
  39. {evalstats-0.2.0 → evalstats-0.2.4}/tests/test_set_alpha_ci.py +2 -2
  40. {evalstats-0.2.0 → evalstats-0.2.4}/tests/test_simultaneous_ci.py +291 -20
  41. {evalstats-0.2.0 → evalstats-0.2.4}/tests/test_wilson_newcombe.py +69 -58
  42. evalstats-0.2.0/PKG-INFO +0 -364
  43. evalstats-0.2.0/README.md +0 -310
  44. evalstats-0.2.0/evalstats/compare.py +0 -900
  45. evalstats-0.2.0/evalstats/config.py +0 -29
  46. evalstats-0.2.0/evalstats/core/stats_utils.py +0 -47
  47. evalstats-0.2.0/evalstats.egg-info/PKG-INFO +0 -364
  48. evalstats-0.2.0/tests/test_auto_ci_routing.py +0 -197
  49. evalstats-0.2.0/tests/test_compare_models.py +0 -385
  50. evalstats-0.2.0/tests/test_compare_prompts.py +0 -567
  51. {evalstats-0.2.0 → evalstats-0.2.4}/LICENSE +0 -0
  52. {evalstats-0.2.0 → evalstats-0.2.4}/evalstats/core/__init__.py +0 -0
  53. {evalstats-0.2.0 → evalstats-0.2.4}/evalstats/core/bayes_evals.py +0 -0
  54. {evalstats-0.2.0 → evalstats-0.2.4}/evalstats/io.py +0 -0
  55. {evalstats-0.2.0 → evalstats-0.2.4}/evalstats/vis/__init__.py +0 -0
  56. {evalstats-0.2.0 → evalstats-0.2.4}/evalstats/vis/heatmap.py +0 -0
  57. {evalstats-0.2.0 → evalstats-0.2.4}/evalstats/vis/point_estimates.py +0 -0
  58. {evalstats-0.2.0 → evalstats-0.2.4}/evalstats.egg-info/dependency_links.txt +0 -0
  59. {evalstats-0.2.0 → evalstats-0.2.4}/evalstats.egg-info/entry_points.txt +0 -0
  60. {evalstats-0.2.0 → evalstats-0.2.4}/evalstats.egg-info/requires.txt +0 -0
  61. {evalstats-0.2.0 → evalstats-0.2.4}/evalstats.egg-info/top_level.txt +0 -0
  62. {evalstats-0.2.0 → evalstats-0.2.4}/setup.cfg +0 -0
  63. {evalstats-0.2.0 → evalstats-0.2.4}/tests/test_analyze_factorial.py +0 -0
  64. {evalstats-0.2.0 → evalstats-0.2.4}/tests/test_io.py +0 -0
  65. {evalstats-0.2.0 → evalstats-0.2.4}/tests/test_lmm.py +0 -0
  66. {evalstats-0.2.0 → evalstats-0.2.4}/tests/test_lmm_backend_parity.py +0 -0
  67. {evalstats-0.2.0 → evalstats-0.2.4}/tests/test_lmm_statsmodels.py +0 -0
  68. {evalstats-0.2.0 → evalstats-0.2.4}/tests/test_nig_ci_methods.py +0 -0
  69. {evalstats-0.2.0 → evalstats-0.2.4}/tests/test_p_values.py +0 -0
  70. {evalstats-0.2.0 → evalstats-0.2.4}/tests/test_permutation.py +0 -0
@@ -0,0 +1,548 @@
1
+ Metadata-Version: 2.4
2
+ Name: evalstats
3
+ Version: 0.2.4
4
+ Summary: Statistically sane analysis methods for comparing AI model and prompt performance.
5
+ Author: Ian Arawjo
6
+ License-Expression: MIT
7
+ Project-URL: Homepage, https://github.com/ianarawjo/evalstats
8
+ Project-URL: Repository, https://github.com/ianarawjo/evalstats
9
+ Classifier: Programming Language :: Python :: 3
10
+ Classifier: Programming Language :: Python :: 3 :: Only
11
+ Requires-Python: >=3.9
12
+ Description-Content-Type: text/markdown
13
+ License-File: LICENSE
14
+ Requires-Dist: numpy>=1.22
15
+ Requires-Dist: scipy>=1.9
16
+ Requires-Dist: pandas>=1.5
17
+ Requires-Dist: matplotlib>=3.6
18
+ Requires-Dist: scikit-posthocs
19
+ Requires-Dist: statsmodels>=0.14
20
+ Requires-Dist: plotext
21
+ Provides-Extra: interactive
22
+ Requires-Dist: plotly>=5.0; extra == "interactive"
23
+ Provides-Extra: models
24
+ Requires-Dist: choix>=0.3; extra == "models"
25
+ Provides-Extra: lmm
26
+ Requires-Dist: pymer4>=0.9; extra == "lmm"
27
+ Requires-Dist: pyarrow>=14.0; extra == "lmm"
28
+ Requires-Dist: great_tables; extra == "lmm"
29
+ Requires-Dist: joblib; extra == "lmm"
30
+ Requires-Dist: rpy2; extra == "lmm"
31
+ Requires-Dist: polars; extra == "lmm"
32
+ Requires-Dist: scikit-learn; extra == "lmm"
33
+ Requires-Dist: formulae; extra == "lmm"
34
+ Provides-Extra: xlsx
35
+ Requires-Dist: openpyxl>=3.0; extra == "xlsx"
36
+ Provides-Extra: all
37
+ Requires-Dist: plotly>=5.0; extra == "all"
38
+ Requires-Dist: choix>=0.3; extra == "all"
39
+ Requires-Dist: openpyxl>=3.0; extra == "all"
40
+ Requires-Dist: pymer4>=0.9; extra == "all"
41
+ Requires-Dist: pyarrow>=14.0; extra == "all"
42
+ Requires-Dist: great_tables; extra == "all"
43
+ Requires-Dist: joblib; extra == "all"
44
+ Requires-Dist: rpy2; extra == "all"
45
+ Requires-Dist: polars; extra == "all"
46
+ Requires-Dist: scikit-learn; extra == "all"
47
+ Requires-Dist: formulae; extra == "all"
48
+ Provides-Extra: dev
49
+ Requires-Dist: pytest>=7.0; extra == "dev"
50
+ Requires-Dist: pytest-cov; extra == "dev"
51
+ Requires-Dist: build>=1.0; extra == "dev"
52
+ Requires-Dist: twine>=5.0; extra == "dev"
53
+ Dynamic: license-file
54
+
55
+ # evalstats
56
+
57
+ Rigorous statistical analysis for LLM evaluations, from model and prompt comparisons to statistical tests resilient to LLM judge bias, including in small sample data regimes.
58
+
59
+ `evalstats` helps you answer questions like:
60
+ - Is Prompt A actually better than Prompt B, or just slightly luckier on this dataset?
61
+ - Does Model A beat Model B, or only under a specific prompt phrasing?
62
+ - How sensitive is model performance to prompt wording?
63
+ - Are my performance differences large enough to be meaningful, or just noise?
64
+ - How stable are scores across runs, evaluators, or inputs?
65
+ - Can I trust my LLM-judge scores, or do they need correcting against human labels first?
66
+
67
+ You give `evalstats` your benchmark data, and it runs statistically appropriate analyses that quantify uncertainty and provide confidence bounds on your claims. It does this in two main ways:
68
+
69
+ - **Comparisons**: Comparing models, prompts, or both at once (or any other thing you're comparing, like agent harnesses), and get 95% confidence intervals, pairwise significance tests, and multi-run sensitivity analyses. `evalstats` guides you toward best practices and choose well-calibrated methods and procedures by default, backed by simulations, and was built specifically to fill the gap of statistical knowledge for small-sample size datasets N<100; it will output stats as long as there are at least 15 samples. See [Statistics](#statistics).
70
+ - **PPI-corrected inference**: using a small set of human labels to correct bias in noisy LLM-judge scores, so your means, confidence intervals, and hypothesis tests p-values are accurately calibrated in the face of LLM judge bias. This builds on prediction-powered inference (PPI). See [PPI-Corrected Inference](#ppi-corrected-inference-means-cis-and-tests).
71
+
72
+ In particular, scientists can use our PPI-corrected statistical tests to analyze data for **mixed human-AI subject studies**, where some observations are human-labeled and the rest are graded by an LLM judge. Use `evalstats.tests` directly for LLM-judge-bias-corrected versions of:
73
+
74
+ - t-test (`ttest`, independent or paired; Welch's by default, or Student's equal-variance via `equal_var=True`)
75
+ - Mann–Whitney U (`mannwhitney`)
76
+ - Wilcoxon signed-rank (`wilcoxon`)
77
+ - One-way ANOVA (`anova_oneway`, independent or repeated-measures)
78
+ - Friedman test (`friedman`, repeated-measures rank-based)
79
+ - Kruskal-Wallis (`kruskalwallis`, independent-groups rank-based)
80
+
81
+ As long as the items for human labeling were sampled at random from the full dataset, p-values will stay calibrated even when the LLM judge is biased or miscalibrated. These corrections are validated via extensive Monte Carlo simulations (see `simulations/harness`). To the best of our knowledge, `evalstats` provides the only known implementations of PPI-corrected rank-based nonparametric tests like Wilcoxon.
82
+
83
+ As well,
84
+
85
+ > [!IMPORTANT]
86
+ > We are actively building out this project, both the website/guide and the package.
87
+ > Aside from the package itself, there is a "learning" guide in `website/` which I am building out and will return to after writing up the simulations. This will include simulation- and research-backed examples of statistics for LLM evals, as well as example code (which will, obviously, use `evalstats`, but the lessons hold regardless of implementation).
88
+ > If there's something you'd like to see, or guidance on a specific topic, let us know
89
+ > by raising an Issue.
90
+
91
+ ## Sample output
92
+
93
+ Running `es.compare(evaldata, factors="prompt")` and then `result.summary()` prints a full statistical report to the terminal, including confidence interval line plots, pairwise comparisons between prompt templates, and per-input stability across runs (how stable the model is across multiple runs for the same input). Below is example excerpt from an analysis of a 4-template sentiment-classification benchmark (GPT-4.1-nano, 27 inputs, 3 runs, 3 evaluators):
94
+
95
+ ![Example terminal output](docs/example-output.png)
96
+
97
+ From this output, we can see that Minimal and Instructive are the most promising candidates, but it is statistically unclear which is better. We also see that Chain-of-thought gives the least consistent outputs across multiple runs for the same inputs, compared to the other methods.
98
+
99
+ In the most recent version of `evalstats`, there's also helpful colors to help you see
100
+ this information. For instance, comparing models and prompts at the same time,
101
+ `evalstats` shows a 4-way tie between four combinations of model-prompt:
102
+
103
+ ![Example terminal output with colors](docs/terminal-output-example.jpg)
104
+
105
+ You can also plot within notebook environments (although this feature is being actively built out over time and the least developed at the moment). The `plot_point_estimates` function produces a chart showing each template's absolute mean score with marginal confidence intervals:
106
+
107
+ ![Mean advantage plot](docs/mean_advantage.png)
108
+
109
+ ## Statistics
110
+
111
+ The specific statistical tests that `evalstats.compare()` runs (via the lower-level `analyze()` engine underneath it) are:
112
+
113
+ - **All pairwise prompt comparisons (paired by input)** via `all_pairwise(...)`:
114
+ - Computes mean or median difference (mean by default), bootstrapped 95% confidence interval, and p-value for every prompt template pair.
115
+ - Comparison method defaults to `method="auto"`:
116
+ - **Smoothed bootstrap with a Gaussian KDE** (`method="smooth_bootstrap"`) in situations of non-binary data. It has been verified in our simulations that for eval-type data and small sample sizes especially, smoothed is superior to the other bootstrap methods considered (percentile, BCa, Bayesian).
117
+ - **Bayesian pairwise from [`bayes_evals`](https://github.com/sambowyer/bayes_evals/tree/main) and McNemar's test**: Default methods for binary scores (0 or 1 only). Our simulations showed Bayesian pairwise was superior to bootstrap at small N. Note that Bayesian methods should technically be called credible intervals, but they estimate the confidence interval very closely.
118
+ - Multiple-comparisons correction for p-values (defaults to **Benjamini–Hochberg (fdr_bh)**).
119
+ - Also reports Wilcoxon signed-rank test p-value, in case you need it for people familiar with that test, although p-values from bootstrapped CIs are more robust
120
+
121
+ - **Bootstrap rank distribution** via `bootstrap_ranks(...)`:
122
+ - Estimates each prompt template’s `P(best)` and expected rank among the full list of prompt templates.
123
+
124
+ - **Point estimates** via `robustness_metrics(...)`:
125
+ - Descriptive stats like mean, median, std, CV, IQR, CVaR-10, and key percentiles.
126
+ - Marginal confidence intervals on absolute means/medians.
127
+
128
+ If your benchmark includes repeated runs (`R >= 3`), bootstrap-based analyses above use a **two-level nested bootstrap** (resample inputs, then runs within each input) so run-to-run stochasticity is propagated into CIs and rankings. In that case, `analyze()` also returns a seed/input variance decomposition via `seed_variance_decomposition(...)`.
129
+
130
+ If you set `method="lmm"`, `analyze()` switches to a mixed-effects path (`score ~ template + (1|input)`) with Wald CIs and parametric rank distributions. By default this uses `statsmodels` (pure Python, no additional setup required); pass `backend="pymer4"` to use R's lme4/emmeans instead (requires a separate R installation — see below). **Mixed effects model support is more experimental at the moment.**
131
+
132
+ ## Installation and Quick start CLI
133
+
134
+ ```bash
135
+ pip install evalstats
136
+ ```
137
+
138
+ For Excel (`.xlsx`) input support:
139
+
140
+ ```bash
141
+ pip install "evalstats[xlsx]"
142
+ ```
143
+
144
+ For all optional extras (including mixed-effects/LMM support):
145
+
146
+ ```bash
147
+ pip install "evalstats[all]"
148
+ ```
149
+
150
+ From the command line, `evalstats` can read a CSV or Excel file directly and print a statistical summary:
151
+
152
+ ```bash
153
+ evalstats analyze results.csv
154
+ ```
155
+
156
+ The input file should have a prompt/template column, an item/input column, and a score column (model, run, and evaluator columns are optional) — see the column alias table in the [Python API](#python-api) section below for recognized names. Run `evalstats analyze --help` for the full list of options and supported column aliases.
157
+
158
+ For more complex statistical analysis with mixed effects models, use `method="lmm"`. The default `statsmodels` backend works out of the box; for the optional R-based backend, see below.
159
+
160
+ ## Python API
161
+
162
+ The main entry point is `load_from()` + `compare()`: parse your data once into an `EvalResults` object, then run comparisons against it.
163
+
164
+ ```python
165
+ import pandas as pd
166
+ import evalstats as es
167
+
168
+ df = pd.read_csv("results.csv") # columns: prompt, item, score (model optional)
169
+
170
+ evaldata = es.load_from(df)
171
+ evaldata.summary() # inspect detected structure/column assignments before analyzing
172
+
173
+ result = es.compare(evaldata, factors="prompt")
174
+ result.summary() # full terminal report: CIs, pairwise tests, rank probabilities
175
+ ```
176
+
177
+ `evalstats` expects **long-format** data: one row per (item, score) observation, plus whichever axis you want to compare — `model`, `prompt`, or both — and optionally `run` for repeated runs. Only `item` and `score` are strictly required; you need at least one of `model`/`prompt` too, whichever you pass to `compare(factors=...)`. `load_from()` auto-detects each column's role by matching its name (case-insensitively) against this table:
178
+
179
+ | Role | Canonical name | Recognized aliases | Required? |
180
+ |----------|-----------------|----------------------------------------|------------------------------------------------------|
181
+ | model | `model` | `model_label`, `model_name` | Optional — needed to compare models (`factors="model"`) |
182
+ | prompt | `prompt` | `template`, `prompt_template` | Optional — needed to compare prompts (`factors="prompt"`) |
183
+ | item | `item` | `input`, `example`, `id`, `input_label`| Yes |
184
+ | score | `score` | `value`, `result`, `metric` | Yes |
185
+ | run | `run` | `seed`, `repeat`, `run_id`, `trial` | Optional — add if you have repeated runs per (model/prompt, item) |
186
+
187
+ For example, a minimal CSV comparing prompts:
188
+
189
+ | prompt | item | score |
190
+ |-------------|------|-------|
191
+ | Minimal | q1 | 0.82 |
192
+ | Instructive | q1 | 0.91 |
193
+ | Minimal | q2 | 0.75 |
194
+ | Instructive | q2 | 0.88 |
195
+
196
+ If your columns don't match any of the aliases above, remap them explicitly with `col_map`:
197
+
198
+ ```python
199
+ evaldata = es.load_from(df, col_map={"llm": "model", "variant": "prompt", "q_id": "item"})
200
+ ```
201
+
202
+ `compare()` also handles:
203
+
204
+ - **Comparing models**: `factors="model"`
205
+ - **Factorial designs** (model × prompt): `factors=["model", "prompt"]` (routes to an LMM backend)
206
+ - **Filtering**: any keyword matching a column name acts as a row filter, e.g. `es.compare(evaldata, factors="model", split="test")`
207
+ - **PPI-corrected inference** for noisy LLM-judge scores against a smaller human-labeled subset — see [PPI-Corrected Inference](#ppi-corrected-inference-means-cis-and-tests) below
208
+
209
+ The returned `result` is a `ComparisonResult`. Besides `.summary()`, it has `.to_frame()` / `.to_dict()` for programmatic access, `.plot(method="bar" | "forest" | "cd")` for charts, and `.disagreements()` to surface the items entities disagree on most.
210
+
211
+ ### Advanced: raw score arrays (low-level engine)
212
+
213
+ `compare()` is a wrapper around a lower-level engine, `analyze()`, which operates directly on `BenchmarkResult` / `MultiModelBenchmark` objects (numpy score arrays) rather than a DataFrame. Reach for this path only if you already have scores as arrays and don't want to build a DataFrame first — most use cases should use `compare()` above.
214
+
215
+ ```python
216
+ import numpy as np
217
+ import evalstats as estats
218
+
219
+ # Example raw scores for 4 templates × 3 inputs (single run, single evaluator)
220
+ your_scores = [
221
+ [0.91, 0.88, 0.86],
222
+ [0.90, 0.89, 0.84],
223
+ [0.85, 0.82, 0.80],
224
+ [0.79, 0.76, 0.74],
225
+ ]
226
+ n_templates = 4
227
+ n_inputs = 3
228
+
229
+ # scores shape: (n_templates, n_inputs, n_runs, n_evaluators)
230
+ # For a single evaluator and single run, shape is (N, M, 1, 1)
231
+ scores = np.array(your_scores).reshape(n_templates, n_inputs, 1, 1)
232
+
233
+ result = estats.BenchmarkResult(
234
+ scores=scores,
235
+ template_labels=["Minimal", "Instructive", "Few-shot", "Chain-of-thought"],
236
+ input_labels=[f"input_{i}" for i in range(n_inputs)],
237
+ )
238
+
239
+ analysis = estats.analyze(result, reference="grand_mean", n_bootstrap=5_000)
240
+ analysis.summary() # same terminal report as ComparisonResult.summary()
241
+ ```
242
+
243
+ If you want this lower-level path from a DataFrame (e.g. to inspect the raw `BenchmarkResult` object, or to fine-tune `strict_complete_design`), use `from_dataframe()` instead of `load_from()`. It returns the array-based `BenchmarkResult` / `MultiModelBenchmark` that `analyze()` expects, plus an optional `DataLoadReport` — a data-quality log of coercions/repairs made while parsing (not a statistical report):
244
+
245
+ ```python
246
+ import evalstats as estats
247
+
248
+ benchmark, load_report = estats.from_dataframe(
249
+ df,
250
+ format="auto", # auto / wide / long
251
+ repair=True, # average duplicate cells + fill partial run slots
252
+ strict_complete_design=True, # set False to keep NaNs
253
+ return_report=True,
254
+ )
255
+
256
+ for line in load_report.to_lines():
257
+ print(line)
258
+
259
+ analysis = estats.analyze(benchmark)
260
+ analysis.summary()
261
+ ```
262
+
263
+ To visualize absolute prompt performance directly from a `BenchmarkResult`, bypassing `analyze()` (use `result.plot()` above instead if you're on the `compare()` path):
264
+
265
+ ```python
266
+ fig = estats.plot_point_estimates(result)
267
+ fig.savefig("mean_performance.png", dpi=150, bbox_inches="tight")
268
+ ```
269
+
270
+ ## PPI-Corrected Inference (Means, CIs, and Tests)
271
+
272
+ `evalstats` supports PPI-corrected inference for means, confidence intervals, and common statistical tests.
273
+
274
+ PPI (Prediction-Powered Inference) lets you use lots of cheap LLM
275
+ judgments plus a smaller set of human labels to correct measurement error from the LLM
276
+ judge. This gives you corrected estimates and uncertainty that better reflect what you
277
+ would have gotten from a fully human-labeled study (Angelopoulos et al., 2023).
278
+
279
+ Most PPI correction methods use PPIBoot (bootstrap variant of PPI; Zrnic, 2024).
280
+ Implemented corrections have been battle-tested via simulations (see `simulations/sim_type_i_calibration.py`).
281
+
282
+ > **Important: which items get a human label must be chosen uniformly at
283
+ > random.** PPI correction assumes the labeled subset is representative of
284
+ > the full dataset. If your labeling process instead targets specific items
285
+ > — e.g. "always double-check the borderline or highest-scoring responses,"
286
+ > a common real-world review habit — that's missing-not-at-random (MNAR)
287
+ > selection on the outcome, and PPI correction can stay badly miscalibrated
288
+ > **no matter how many items you label**. This isn't ordinary small-sample
289
+ > noise that more labels fixes; it was confirmed in simulation to persist
290
+ > from 15 up through 300 labeled items out of 400. See
291
+ > `evalstats.ppi.correct`'s docstring for the full analysis. If you can't
292
+ > guarantee random labeling, treat any PPI-corrected result here with
293
+ > caution regardless of the reported CI/p-value.
294
+
295
+ ### Example: Comparing models with corrected LLM judge evals via `compare(..., alignment=...)`
296
+
297
+ ```python
298
+ import evalstats as es
299
+
300
+ # Dataframe columns include:
301
+ # model item llm_score human_score (NaN for unlabeled rows)
302
+ evaldata = es.load_from(df)
303
+
304
+ # Compute alignment between LLM and human judges
305
+ alignment = es.validate_alignment(
306
+ evaldata,
307
+ llm_metric="llm_score",
308
+ human_groundtruth="human_score",
309
+ )
310
+
311
+ # Compare models, using PPI to correct for bias/misalignment with human graders
312
+ result = es.compare(
313
+ evaldata,
314
+ factors="model",
315
+ metric="llm_score",
316
+ alignment={"llm_score": alignment},
317
+ )
318
+
319
+ result.summary()
320
+ ```
321
+
322
+ ### Example: T-test PPI-correction via `evalstats.tests.ttest`
323
+
324
+ Use this for a t-test of mean differences between two groups (or two paired
325
+ conditions when `paired=True`).
326
+
327
+ ```python
328
+ import evalstats as es
329
+
330
+ res = es.tests.ttest(
331
+ a=llm_a,
332
+ b=llm_b,
333
+ a_lab=human_a, # same length as llm_a, NaN where unlabeled
334
+ b_lab=human_b, # same length as llm_b, NaN where unlabeled
335
+ paired=False,
336
+ print_result=False,
337
+ )
338
+
339
+ print(res.p_value, res.corrected_p_value, res.corrected_ci)
340
+ ```
341
+
342
+ ### Example: Mann-Whitney U test PPI-correction via `evalstats.tests.mannwhitney`
343
+
344
+ Use this for a Mann-Whitney U test, a nonparametric two-group comparison based
345
+ on relative ranks rather than assuming normally distributed scores.
346
+
347
+ ```python
348
+ import evalstats as es
349
+
350
+ res = es.tests.mannwhitney(
351
+ x=llm_x,
352
+ y=llm_y,
353
+ x_lab=human_x,
354
+ y_lab=human_y,
355
+ print_result=False,
356
+ )
357
+
358
+ print(res.p_value, res.corrected_p_value, res.corrected_ci)
359
+ ```
360
+
361
+ ### Example: Wilcoxon signed-ranks test PPI-correction via `evalstats.tests.wilcoxon` (paired)
362
+
363
+ Use this for a Wilcoxon signed-rank test, a nonparametric paired test for
364
+ matched observations (before/after, A/B on the same items, etc.).
365
+
366
+ ```python
367
+ import evalstats as es
368
+
369
+ res = es.tests.wilcoxon(
370
+ x=llm_before,
371
+ y=llm_after,
372
+ x_lab=human_before,
373
+ y_lab=human_after,
374
+ print_result=False,
375
+ )
376
+
377
+ print(res.p_value, res.corrected_p_value, res.corrected_ci)
378
+ ```
379
+
380
+ ### Example: One-way ANOVA PPI-correction via `evalstats.tests.anova_oneway`
381
+
382
+ Use this for one-way ANOVA when comparing more than two groups, with
383
+ `repeated=True` for repeated-measures (same subjects across conditions).
384
+
385
+ ```python
386
+ import evalstats as es
387
+
388
+ res = es.tests.anova_oneway(
389
+ llm_g1,
390
+ llm_g2,
391
+ llm_g3,
392
+ groups_lab=[human_g1, human_g2, human_g3],
393
+ repeated=False,
394
+ print_result=False,
395
+ )
396
+
397
+ print(res.p_value, res.corrected_p_value, res.corrected_ci)
398
+ ```
399
+
400
+ ## Motivation
401
+
402
+ Most eval tools in the LLM evaluation space don't help users perform _any_ statistical tests, let alone showcase variances in performance between prompts or models. They instead present bar charts of average performance. Developers then glance at the bar chart and decide that "prompt/model A is better than B." But was it really?
403
+
404
+ Relying purely on bar charts and averages can very, very easily lead to erroneous conclusions—B might actually be more robust than A, or B performs well on an important subset of data, or there's not enough data to conclude one way or the other.
405
+
406
+ Why do people do evals this way? Well, they don't have the time, tools, or knowledge on how to do it better—frequently, they don't even know there's a better way.
407
+
408
+ `evalstats` aims to rectify this with simple, powerful defaults—just throw us your data and we'll run the stats and plot the results for you. Upstream applications, like LLM observability platforms, could take `evalstats` results and plot them in their own front-ends. Prompt optimization tools could also use `evalstats` to decide, e.g., when to cull a candidate prompt and how to present results to users.
409
+
410
+ ## Examples
411
+
412
+ ### Is one prompt "better" than others? Quantify uncertainty
413
+
414
+ When you have scores for multiple prompt templates across a set of inputs, `evalstats` computes bootstrapped 95% confidence intervals and pairwise significance tests so you can see not just which prompt scored highest on average, but how certain you can be about that ranking. It plots these to the terminal so you can check at a glance:
415
+
416
+ ![Comparing across prompts output](docs/compare-prompts-output.png)
417
+
418
+ ### Comparing across models while accounting for prompt sensitivity
419
+
420
+ A common failure mode in LLM benchmarking, both in academic papers and practitioner evaluations, is testing each model with a single prompt template and reporting the resulting scores as if they reflect stable model capabilities. In reality, model rankings can flip under semantically equivalent paraphrases of the same instruction. A benchmark result that says "Model A beats Model B" may be an artifact of prompt phrasing, not a meaningful capability difference.
421
+
422
+ Here, we can see the difference between OpenAI's `gpt-4.1-nano` and MistralAI's `ministral-8b-2512` on a small sentiment classification benchmark, quantified by bootstrapped 95% confidence intervals:
423
+
424
+ ![Comparing across models output](docs/compare-models-output.png)
425
+
426
+ In this run, multiple prompt template variations were considered, making this result more robust than trying a single prompt and calling it a day.
427
+
428
+ ### How stable is the performance across runs?
429
+
430
+ LLMs are stochastic at temperature>0. Will the performance stay similar, even upon multiple runs for the same inputs? `evalstats` offers a helpful "noise plot" which visualizes (in)stability across runs:
431
+
432
+ ![Per-input noise across runs](docs/per-input-noise.png)
433
+
434
+ ## Running Example Scripts
435
+
436
+ We provide multiple standalone example scripts that rig up a simple benchmark, collect LLM responses, and run analyses over them. From the repository root, run any example script directly:
437
+
438
+ ```bash
439
+ python examples/synthetic_mean_advantage.py
440
+ ```
441
+
442
+ Additional examples:
443
+
444
+ ```bash
445
+ # OpenAI sentiment benchmark (single run)
446
+ python examples/sentiment.py
447
+
448
+ # Multi-run variant (captures run-to-run variability)
449
+ python examples/sentiment_multirun.py
450
+
451
+ # Multi-model comparison across prompt templates
452
+ python examples/compare_models_multirun.py
453
+
454
+ # Manual API call walkthrough
455
+ python examples/sentiment_manual_api_calls.py
456
+ ```
457
+
458
+ OpenAI-powered examples require `OPENAI_API_KEY` set in your environment. But, you can easily swap out the model calls to whatever model you prefer.
459
+
460
+ ## Mixed effects models (LMM)
461
+
462
+ > [!IMPORTANT]
463
+ > Mixed effects analysis is experimental, and currently offers only the advantage of gracefully
464
+ > dealing with missing data. In the future, we plan to add factor decomposition across multiple inputs.
465
+ > We recommend only using `method="lmm"` if you need robustness to missing data (`NaN`). Keep
466
+ > in mind that missing data must be reasonably random (i.e., like sampling from a larger distribution).
467
+
468
+ `evalstats` supports mixed-effects models (`score ~ template + (1|input)`) for:
469
+ - Missing data in inputs (some score cells are `NaN`)
470
+ - Factor decomposition when multiple input factors are present
471
+
472
+ ### Default backend: statsmodels (pure Python)
473
+
474
+ No extra setup required — `statsmodels` is included in the standard `pip install evalstats`. Simply pass `method="lmm"`:
475
+
476
+ ```python
477
+ analysis = estats.analyze(result, method="lmm")
478
+ ```
479
+
480
+ `evalstats` fits the model with REML, computes Wald CIs via the delta method, and estimates rank distributions by parametric simulation.
481
+
482
+ ### Optional backend: pymer4 (requires R)
483
+
484
+ For Satterthwaite degrees of freedom and `emmeans`-based pairwise contrasts (R's gold standard for mixed models), pass `backend="pymer4"`:
485
+
486
+ ```python
487
+ analysis = estats.analyze(result, method="lmm", backend="pymer4")
488
+ ```
489
+
490
+ This requires a working R installation with the following packages:
491
+
492
+ ```r
493
+ install.packages(c(
494
+ "lme4",
495
+ "emmeans",
496
+ "tibble",
497
+ "broom",
498
+ "broom.mixed",
499
+ "lmerTest",
500
+ "report",
501
+ "car"
502
+ ))
503
+ ```
504
+
505
+ Then install the Python LMM extra:
506
+
507
+ ```bash
508
+ pip install "evalstats[lmm]"
509
+ ```
510
+
511
+ > [!NOTE]
512
+ > If your environment needs manual dependency pinning, this is the tested equivalent:
513
+ >
514
+ > ```bash
515
+ > pip install "pymer4>=0.9" great_tables joblib rpy2 polars scikit-learn formulae pyarrow
516
+ > ```
517
+
518
+ Installation details may differ on your system.
519
+
520
+ ## Reproducibility: Monte Carlo simulations
521
+
522
+ Claims in this README like "verified in our simulations" are backed by a runnable simulation harness in `simulations/harness/` of this package. We engineered these simulations so that you can run these yourself. For instance:
523
+
524
+ ```bash
525
+ python -m simulations.harness.cli --list-cases
526
+ python -m simulations.harness.cli --official-tests
527
+ python -m simulations.harness.cli ci_single --reps 50 --sizes 10 20
528
+ python -m simulations.harness.cli pvalues --mode ppi --tests ttest wilcoxon anova_rep
529
+ ```
530
+
531
+ `--official-tests will bring up a CLI with options to run specific tests. Each runs each case's canonical, full-scale preset and writes results plus a `manifest.json` (args, output paths, key metrics, pass/fail) to `simulations/out/official_<timestamp>/`. See [`simulations/harness/README.md`](simulations/harness/README.md) for the full case list, scenario library, and verification methodology against the original standalone scripts. Note that *each* simulation can take *very long* to run; even on a MacBook Pro with an M4 Max chip and 64GB RAM, with computation paralellized across 16 CPU cores, it often takes many hours.
532
+ - `ci_single` / `ci_paired` — coverage and width of confidence interval methods (bootstrap, smoothed bootstrap, Bayesian, Wilson, etc.) across synthetic distributions and real benchmark data (OpenEval, Inspect AI).
533
+ - `pvalues --mode pairwise` / `--mode multiarm` — Type-I error and power for pairwise and multi-arm comparisons, including multiple-comparisons correction strategies.
534
+ - `pvalues --mode ppi` — Type-I error calibration and power for every PPI-corrected test in `evalstats.tests`, swept across judge-bias severity, label fraction, and MNAR-labeling scenarios.
535
+
536
+
537
+ ## Development and Contributions
538
+
539
+ For package build, release validation, and maintainer workflows, see [DEVELOPMENT.md](DEVELOPMENT.md).
540
+
541
+ We welcome contributions, especially refinements to our statistical methods. If you're proposing a new correction, CI method, or a fix to an existing one, we encourage battle-testing it against the [simulation harness](#reproducibility-monte-carlo-simulations) first. Please add or extend a scenario and confirm your change holds up on Type-I error and power, not just on the case that motivated it, before opening a PR. The `evalstats` repository already offers a rigorous, expansive synthetic suite that generally has held up against real data.
542
+
543
+ ## License
544
+
545
+ This repository uses two licenses:
546
+
547
+ - **`evalstats` package** (everything outside `website/`) — [MIT](LICENSE).
548
+ - **Stats for Evals Website** (everything in `website/`) — [CC BY-NC-ND 4.0](website/LICENSE). You may share it with attribution non-commercially, but commercial use and derivative works are not permitted.