evalstats 0.2.4__tar.gz → 0.2.6__tar.gz
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- evalstats-0.2.6/PKG-INFO +434 -0
- evalstats-0.2.6/README.md +380 -0
- {evalstats-0.2.4 → evalstats-0.2.6}/evalstats/__init__.py +26 -3
- evalstats-0.2.6/evalstats/alignment.py +2775 -0
- {evalstats-0.2.4 → evalstats-0.2.6}/evalstats/api.py +1291 -139
- {evalstats-0.2.4 → evalstats-0.2.6}/evalstats/cli.py +398 -5
- evalstats-0.2.6/evalstats/config.py +683 -0
- {evalstats-0.2.4 → evalstats-0.2.6}/evalstats/core/bundles.py +36 -0
- evalstats-0.2.6/evalstats/core/design.py +37 -0
- {evalstats-0.2.4 → evalstats-0.2.6}/evalstats/core/mixed_effects.py +37 -15
- {evalstats-0.2.4 → evalstats-0.2.6}/evalstats/core/paired.py +977 -106
- evalstats-0.2.6/evalstats/core/pareto.py +413 -0
- {evalstats-0.2.4 → evalstats-0.2.6}/evalstats/core/ranking.py +65 -0
- {evalstats-0.2.4 → evalstats-0.2.6}/evalstats/core/resampling.py +1246 -102
- {evalstats-0.2.4 → evalstats-0.2.6}/evalstats/core/router.py +256 -69
- {evalstats-0.2.4 → evalstats-0.2.6}/evalstats/core/summary.py +1492 -346
- {evalstats-0.2.4 → evalstats-0.2.6}/evalstats/core/types.py +2 -0
- evalstats-0.2.6/evalstats/core/unpaired.py +882 -0
- {evalstats-0.2.4 → evalstats-0.2.6}/evalstats/core/variance.py +62 -8
- evalstats-0.2.6/evalstats/labeling.py +466 -0
- {evalstats-0.2.4 → evalstats-0.2.6}/evalstats/loader.py +43 -5
- evalstats-0.2.6/evalstats/ppi.py +2102 -0
- evalstats-0.2.6/evalstats/quick.py +1126 -0
- {evalstats-0.2.4 → evalstats-0.2.6}/evalstats/tests/__init__.py +1993 -1069
- {evalstats-0.2.4 → evalstats-0.2.6}/evalstats/vis/__init__.py +4 -0
- evalstats-0.2.6/evalstats/vis/forest.py +926 -0
- evalstats-0.2.6/evalstats/vis/pareto.py +409 -0
- evalstats-0.2.6/evalstats/vis/reliability.py +163 -0
- evalstats-0.2.6/evalstats.egg-info/PKG-INFO +434 -0
- {evalstats-0.2.4 → evalstats-0.2.6}/evalstats.egg-info/SOURCES.txt +18 -0
- {evalstats-0.2.4 → evalstats-0.2.6}/pyproject.toml +1 -1
- {evalstats-0.2.4 → evalstats-0.2.6}/tests/test_alignment.py +301 -97
- {evalstats-0.2.4 → evalstats-0.2.6}/tests/test_analyze.py +159 -3
- {evalstats-0.2.4 → evalstats-0.2.6}/tests/test_auto_ci_routing.py +65 -7
- {evalstats-0.2.4 → evalstats-0.2.6}/tests/test_bayes_binary_routing.py +30 -30
- evalstats-0.2.6/tests/test_bonett_price_multirun.py +350 -0
- evalstats-0.2.6/tests/test_case_cli_preset_agreement.py +61 -0
- evalstats-0.2.6/tests/test_ci_forest_plot.py +448 -0
- {evalstats-0.2.4 → evalstats-0.2.6}/tests/test_cli.py +28 -16
- {evalstats-0.2.4 → evalstats-0.2.6}/tests/test_compare.py +23 -0
- evalstats-0.2.6/tests/test_compare_e2e_dgp.py +172 -0
- evalstats-0.2.6/tests/test_compound_ppi_fwer.py +855 -0
- evalstats-0.2.6/tests/test_design.py +59 -0
- evalstats-0.2.6/tests/test_latex_tables.py +371 -0
- {evalstats-0.2.4 → evalstats-0.2.6}/tests/test_p_values.py +119 -20
- evalstats-0.2.6/tests/test_pareto.py +785 -0
- evalstats-0.2.6/tests/test_ppi_ci_methods.py +311 -0
- {evalstats-0.2.4 → evalstats-0.2.6}/tests/test_ppi_corrections.py +170 -65
- evalstats-0.2.6/tests/test_quick_primitives.py +456 -0
- {evalstats-0.2.4 → evalstats-0.2.6}/tests/test_resampling.py +66 -1
- {evalstats-0.2.4 → evalstats-0.2.6}/tests/test_set_alpha_ci.py +12 -11
- {evalstats-0.2.4 → evalstats-0.2.6}/tests/test_simultaneous_ci.py +240 -51
- evalstats-0.2.6/tests/test_unpaired.py +910 -0
- {evalstats-0.2.4 → evalstats-0.2.6}/tests/test_wilson_newcombe.py +328 -138
- evalstats-0.2.4/PKG-INFO +0 -548
- evalstats-0.2.4/README.md +0 -494
- evalstats-0.2.4/evalstats/alignment.py +0 -804
- evalstats-0.2.4/evalstats/config.py +0 -314
- evalstats-0.2.4/evalstats/ppi.py +0 -941
- evalstats-0.2.4/evalstats/vis/forest.py +0 -270
- evalstats-0.2.4/evalstats.egg-info/PKG-INFO +0 -548
- {evalstats-0.2.4 → evalstats-0.2.6}/LICENSE +0 -0
- {evalstats-0.2.4 → evalstats-0.2.6}/evalstats/core/__init__.py +0 -0
- {evalstats-0.2.4 → evalstats-0.2.6}/evalstats/core/bayes_evals.py +0 -0
- {evalstats-0.2.4 → evalstats-0.2.6}/evalstats/core/stats_utils.py +0 -0
- {evalstats-0.2.4 → evalstats-0.2.6}/evalstats/io.py +0 -0
- {evalstats-0.2.4 → evalstats-0.2.6}/evalstats/vis/critical_difference.py +0 -0
- {evalstats-0.2.4 → evalstats-0.2.6}/evalstats/vis/heatmap.py +0 -0
- {evalstats-0.2.4 → evalstats-0.2.6}/evalstats/vis/point_estimates.py +0 -0
- {evalstats-0.2.4 → evalstats-0.2.6}/evalstats/vis/scoreboard.py +0 -0
- {evalstats-0.2.4 → evalstats-0.2.6}/evalstats.egg-info/dependency_links.txt +0 -0
- {evalstats-0.2.4 → evalstats-0.2.6}/evalstats.egg-info/entry_points.txt +0 -0
- {evalstats-0.2.4 → evalstats-0.2.6}/evalstats.egg-info/requires.txt +0 -0
- {evalstats-0.2.4 → evalstats-0.2.6}/evalstats.egg-info/top_level.txt +0 -0
- {evalstats-0.2.4 → evalstats-0.2.6}/setup.cfg +0 -0
- {evalstats-0.2.4 → evalstats-0.2.6}/tests/test_analyze_factorial.py +0 -0
- {evalstats-0.2.4 → evalstats-0.2.6}/tests/test_bootstrap_t_pairwise_ranking.py +0 -0
- {evalstats-0.2.4 → evalstats-0.2.6}/tests/test_critical_difference_plot.py +0 -0
- {evalstats-0.2.4 → evalstats-0.2.6}/tests/test_io.py +0 -0
- {evalstats-0.2.4 → evalstats-0.2.6}/tests/test_lmm.py +0 -0
- {evalstats-0.2.4 → evalstats-0.2.6}/tests/test_lmm_backend_parity.py +0 -0
- {evalstats-0.2.4 → evalstats-0.2.6}/tests/test_lmm_statsmodels.py +0 -0
- {evalstats-0.2.4 → evalstats-0.2.6}/tests/test_nig_ci_methods.py +0 -0
- {evalstats-0.2.4 → evalstats-0.2.6}/tests/test_permutation.py +0 -0
- {evalstats-0.2.4 → evalstats-0.2.6}/tests/test_ppi_core.py +0 -0
- {evalstats-0.2.4 → evalstats-0.2.6}/tests/test_scoreboard_plot.py +0 -0
evalstats-0.2.6/PKG-INFO
ADDED
|
@@ -0,0 +1,434 @@
|
|
|
1
|
+
Metadata-Version: 2.4
|
|
2
|
+
Name: evalstats
|
|
3
|
+
Version: 0.2.6
|
|
4
|
+
Summary: Statistically sane analysis methods for comparing AI model and prompt performance.
|
|
5
|
+
Author: Ian Arawjo
|
|
6
|
+
License-Expression: MIT
|
|
7
|
+
Project-URL: Homepage, https://github.com/ianarawjo/evalstats
|
|
8
|
+
Project-URL: Repository, https://github.com/ianarawjo/evalstats
|
|
9
|
+
Classifier: Programming Language :: Python :: 3
|
|
10
|
+
Classifier: Programming Language :: Python :: 3 :: Only
|
|
11
|
+
Requires-Python: >=3.9
|
|
12
|
+
Description-Content-Type: text/markdown
|
|
13
|
+
License-File: LICENSE
|
|
14
|
+
Requires-Dist: numpy>=1.22
|
|
15
|
+
Requires-Dist: scipy>=1.9
|
|
16
|
+
Requires-Dist: pandas>=1.5
|
|
17
|
+
Requires-Dist: matplotlib>=3.6
|
|
18
|
+
Requires-Dist: scikit-posthocs
|
|
19
|
+
Requires-Dist: statsmodels>=0.14
|
|
20
|
+
Requires-Dist: plotext
|
|
21
|
+
Provides-Extra: interactive
|
|
22
|
+
Requires-Dist: plotly>=5.0; extra == "interactive"
|
|
23
|
+
Provides-Extra: models
|
|
24
|
+
Requires-Dist: choix>=0.3; extra == "models"
|
|
25
|
+
Provides-Extra: lmm
|
|
26
|
+
Requires-Dist: pymer4>=0.9; extra == "lmm"
|
|
27
|
+
Requires-Dist: pyarrow>=14.0; extra == "lmm"
|
|
28
|
+
Requires-Dist: great_tables; extra == "lmm"
|
|
29
|
+
Requires-Dist: joblib; extra == "lmm"
|
|
30
|
+
Requires-Dist: rpy2; extra == "lmm"
|
|
31
|
+
Requires-Dist: polars; extra == "lmm"
|
|
32
|
+
Requires-Dist: scikit-learn; extra == "lmm"
|
|
33
|
+
Requires-Dist: formulae; extra == "lmm"
|
|
34
|
+
Provides-Extra: xlsx
|
|
35
|
+
Requires-Dist: openpyxl>=3.0; extra == "xlsx"
|
|
36
|
+
Provides-Extra: all
|
|
37
|
+
Requires-Dist: plotly>=5.0; extra == "all"
|
|
38
|
+
Requires-Dist: choix>=0.3; extra == "all"
|
|
39
|
+
Requires-Dist: openpyxl>=3.0; extra == "all"
|
|
40
|
+
Requires-Dist: pymer4>=0.9; extra == "all"
|
|
41
|
+
Requires-Dist: pyarrow>=14.0; extra == "all"
|
|
42
|
+
Requires-Dist: great_tables; extra == "all"
|
|
43
|
+
Requires-Dist: joblib; extra == "all"
|
|
44
|
+
Requires-Dist: rpy2; extra == "all"
|
|
45
|
+
Requires-Dist: polars; extra == "all"
|
|
46
|
+
Requires-Dist: scikit-learn; extra == "all"
|
|
47
|
+
Requires-Dist: formulae; extra == "all"
|
|
48
|
+
Provides-Extra: dev
|
|
49
|
+
Requires-Dist: pytest>=7.0; extra == "dev"
|
|
50
|
+
Requires-Dist: pytest-cov; extra == "dev"
|
|
51
|
+
Requires-Dist: build>=1.0; extra == "dev"
|
|
52
|
+
Requires-Dist: twine>=5.0; extra == "dev"
|
|
53
|
+
Dynamic: license-file
|
|
54
|
+
|
|
55
|
+
# evalstats
|
|
56
|
+
|
|
57
|
+
[](https://pypi.org/project/evalstats/)
|
|
58
|
+
[](LICENSE)
|
|
59
|
+
[](pyproject.toml)
|
|
60
|
+
|
|
61
|
+
Rigorous statistical analysis for LLM evaluations: from model and prompt comparisons to statistical tests resilient to LLM judge bias, including in small-sample data regimes.
|
|
62
|
+
|
|
63
|
+
`evalstats` helps you answer questions like:
|
|
64
|
+
- Is Prompt A actually better than Prompt B, or just slightly luckier on this dataset?
|
|
65
|
+
- Does Model A beat Model B, or only under a specific prompt phrasing?
|
|
66
|
+
- Are my performance differences large enough to be meaningful, or just noise?
|
|
67
|
+
- How stable are scores across runs, evaluators, or inputs?
|
|
68
|
+
- Can I trust my LLM-judge scores, or do they need correcting against human labels first?
|
|
69
|
+
|
|
70
|
+
You give `evalstats` your benchmark data, and it runs statistically appropriate analyses that quantify uncertainty and provide confidence bounds on your claims, in two main ways:
|
|
71
|
+
|
|
72
|
+
- **Comparisons**: compare models, prompts, or both at once, and get 95% confidence intervals, pairwise significance tests, and multi-run sensitivity analyses. `evalstats` picks well-calibrated methods by default, backed by simulations, and was built specifically for small-sample datasets (N<100); it will output stats down to 15 samples. See [Recommended Methods](#recommended-methods).
|
|
73
|
+
- **PPI-corrected inference**: use a small set of human labels to correct bias in noisy LLM-judge scores, so your means, confidence intervals, and hypothesis-test p-values stay calibrated in the face of LLM judge bias. See [PPI-Corrected Inference](#ppi-corrected-inference-means-cis-and-tests).
|
|
74
|
+
|
|
75
|
+
Scientists can use our PPI-corrected statistical tests for **mixed human-AI subject studies**, where some observations are human-labeled and the rest are graded by an LLM judge. `evalstats.tests` gives you LLM-judge-bias-corrected versions of:
|
|
76
|
+
|
|
77
|
+
- t-test (`ttest`, independent or paired; Welch's by default, or Student's equal-variance via `equal_var=True`)
|
|
78
|
+
- Mann–Whitney U (`mannwhitney`)
|
|
79
|
+
- Wilcoxon signed-rank (`wilcoxon`)
|
|
80
|
+
- One-way ANOVA (`anova_oneway`, independent or repeated-measures)
|
|
81
|
+
- Friedman test (`friedman`, repeated-measures rank-based)
|
|
82
|
+
- Kruskal-Wallis (`kruskalwallis`, independent-groups rank-based)
|
|
83
|
+
|
|
84
|
+
As long as items for human labeling were sampled at random from the full dataset, p-values stay calibrated even when the LLM judge is biased or miscalibrated, validated via extensive Monte Carlo simulations (see [`simulations/harness`](simulations/harness)). To the best of our knowledge, `evalstats` provides the only known implementations of PPI-corrected rank-based nonparametric tests like Wilcoxon.
|
|
85
|
+
|
|
86
|
+
> [!IMPORTANT]
|
|
87
|
+
> We are actively building out this project. A paper with the full methodology and simulation-backed validation behind every default is forthcoming: see [Recommended Methods](#recommended-methods) and [Citation](#citation). In the meantime, the [Stats for LLM Evals guide](https://statsforevals.com/) covers the same material in web form. If there's something you'd like to see, let us know by raising an Issue.
|
|
88
|
+
|
|
89
|
+
## Contents
|
|
90
|
+
|
|
91
|
+
- [Installation](#installation)
|
|
92
|
+
- [Quick start](#quick-start)
|
|
93
|
+
- [See it in action](#see-it-in-action)
|
|
94
|
+
- [Recommended Methods](#recommended-methods)
|
|
95
|
+
- [Python API](#python-api)
|
|
96
|
+
- [PPI-Corrected Inference](#ppi-corrected-inference-means-cis-and-tests)
|
|
97
|
+
- [CLI Reference](#cli-reference)
|
|
98
|
+
- [Examples](#examples)
|
|
99
|
+
- [Mixed effects models (LMM)](#mixed-effects-models-lmm)
|
|
100
|
+
- [Reproducibility: Monte Carlo simulations](#reproducibility-monte-carlo-simulations)
|
|
101
|
+
- [Motivation](#motivation)
|
|
102
|
+
- [Development and Contributions](#development-and-contributions)
|
|
103
|
+
- [Citation](#citation)
|
|
104
|
+
- [License](#license)
|
|
105
|
+
|
|
106
|
+
## Installation
|
|
107
|
+
|
|
108
|
+
```bash
|
|
109
|
+
pip install evalstats
|
|
110
|
+
```
|
|
111
|
+
|
|
112
|
+
For Excel (`.xlsx`) input support: `pip install "evalstats[xlsx]"`. For every optional extra (including mixed-effects/LMM support): `pip install "evalstats[all]"`.
|
|
113
|
+
|
|
114
|
+
## Quick start
|
|
115
|
+
|
|
116
|
+
From the command line, `evalstats` can read a CSV or Excel file directly and print a statistical summary:
|
|
117
|
+
|
|
118
|
+
```bash
|
|
119
|
+
evalstats analyze results.csv
|
|
120
|
+
```
|
|
121
|
+
|
|
122
|
+
The input file should have a prompt/template column, an item/input column, and a score column (model, run, and evaluator columns are optional). See the [column alias table](#python-api) below, or run `evalstats analyze --help` for the full option and alias list.
|
|
123
|
+
|
|
124
|
+
From Python, the main entry point is `load_from()` + `compare()`:
|
|
125
|
+
|
|
126
|
+
```python
|
|
127
|
+
import pandas as pd
|
|
128
|
+
import evalstats as es
|
|
129
|
+
|
|
130
|
+
df = pd.read_csv("results.csv") # columns: prompt, item, score (model optional)
|
|
131
|
+
|
|
132
|
+
evaldata = es.load_from(df)
|
|
133
|
+
evaldata.summary() # inspect detected structure/column assignments before analyzing
|
|
134
|
+
|
|
135
|
+
result = es.compare(evaldata, factors="prompt")
|
|
136
|
+
result.summary() # full terminal report: CIs, pairwise tests, rank probabilities
|
|
137
|
+
```
|
|
138
|
+
|
|
139
|
+
See [Python API](#python-api) for the full data-format and `compare()` reference.
|
|
140
|
+
|
|
141
|
+
## See it in action
|
|
142
|
+
|
|
143
|
+
Running `es.compare(evaldata, factors="prompt")` then `result.summary()` prints a full statistical report to the terminal: confidence interval line plots, pairwise comparisons, and per-input stability across runs. Below: a 4-template sentiment-classification benchmark (GPT-4.1-nano, 27 inputs, 3 runs, 3 evaluators).
|
|
144
|
+
|
|
145
|
+

|
|
146
|
+
|
|
147
|
+
From this we can see Minimal and Instructive are the most promising candidates, but it's statistically unclear which is better. Chain-of-thought gives the least consistent outputs across runs.
|
|
148
|
+
|
|
149
|
+
Comparing models and prompts at once, `evalstats` colors in a 4-way tie between four model-prompt combinations:
|
|
150
|
+
|
|
151
|
+

|
|
152
|
+
|
|
153
|
+
You can also plot within notebook environments. `plot_point_estimates` shows each template's absolute mean score with marginal confidence intervals:
|
|
154
|
+
|
|
155
|
+

|
|
156
|
+
|
|
157
|
+
And LLMs are stochastic at temperature > 0: the "noise plot" visualizes (in)stability across runs for the same input:
|
|
158
|
+
|
|
159
|
+

|
|
160
|
+
|
|
161
|
+
## Recommended Methods
|
|
162
|
+
|
|
163
|
+
`evalstats.compare()` defaults to `method="auto"`, which picks a well-calibrated statistical method based on your data's estimand, data type, and sample size. These defaults come from an extensive Monte Carlo simulation study across eval data types, sample sizes, and comparison setups, cross-checked against real LLM eval data, summarized in the two decision trees below. **Boxed methods are the default; gray notes give the multi-run variant and conservative alternatives.**
|
|
164
|
+
|
|
165
|
+

|
|
166
|
+
|
|
167
|
+

|
|
168
|
+
|
|
169
|
+
A paper with the full methodology, simulation results, and justification behind each recommendation is forthcoming (see [Citation](#citation)). Until then, see the [Which Method?](https://statsforevals.com/which-method.html) page on the `evalstats` site for the web version of these trees, and [`simulations/harness`](simulations/harness) to reproduce the underlying simulations yourself.
|
|
170
|
+
|
|
171
|
+
## Python API
|
|
172
|
+
|
|
173
|
+
`evalstats` expects **long-format** data: one row per (item, score) observation, plus whichever axis you want to compare (`model`, `prompt`, or both) and optionally `run` for repeated runs. Only `item` and `score` are strictly required; you need at least one of `model`/`prompt` too, whichever you pass to `compare(factors=...)`. `load_from()` auto-detects each column's role by matching its name (case-insensitively) against this table:
|
|
174
|
+
|
|
175
|
+
| Role | Canonical name | Recognized aliases | Required? |
|
|
176
|
+
|----------|-----------------|----------------------------------------|------------------------------------------------------|
|
|
177
|
+
| model | `model` | `model_label`, `model_name` | Optional: needed to compare models (`factors="model"`) |
|
|
178
|
+
| prompt | `prompt` | `template`, `prompt_template` | Optional: needed to compare prompts (`factors="prompt"`) |
|
|
179
|
+
| item | `item` | `input`, `example`, `id`, `input_label`| Yes |
|
|
180
|
+
| score | `score` | `value`, `result`, `metric` | Yes |
|
|
181
|
+
| run | `run` | `seed`, `repeat`, `run_id`, `trial` | Optional: add if you have repeated runs per (model/prompt, item) |
|
|
182
|
+
|
|
183
|
+
For example, a minimal CSV comparing prompts:
|
|
184
|
+
|
|
185
|
+
| prompt | item | score |
|
|
186
|
+
|-------------|------|-------|
|
|
187
|
+
| Minimal | q1 | 0.82 |
|
|
188
|
+
| Instructive | q1 | 0.91 |
|
|
189
|
+
| Minimal | q2 | 0.75 |
|
|
190
|
+
| Instructive | q2 | 0.88 |
|
|
191
|
+
|
|
192
|
+
If your columns don't match any alias above, remap them explicitly: `es.load_from(df, col_map={"llm": "model", "variant": "prompt", "q_id": "item"})`.
|
|
193
|
+
|
|
194
|
+
`compare()` also handles:
|
|
195
|
+
|
|
196
|
+
- **Comparing models**: `factors="model"`
|
|
197
|
+
- **Factorial designs** (model × prompt): `factors=["model", "prompt"]` (routes to an LMM backend)
|
|
198
|
+
- **Filtering**: any keyword matching a column name acts as a row filter, e.g. `es.compare(evaldata, factors="model", split="test")`
|
|
199
|
+
- **PPI-corrected inference** for noisy LLM-judge scores: see [PPI-Corrected Inference](#ppi-corrected-inference-means-cis-and-tests)
|
|
200
|
+
|
|
201
|
+
The returned `result` is a `ComparisonResult`. Besides `.summary()`, it has `.to_frame()` / `.to_dict()` for programmatic access, `.plot(method="forest" | "bar" | "cd" | "pareto")` for charts, and `.disagreements()` to surface the items entities disagree on most.
|
|
202
|
+
|
|
203
|
+
<details>
|
|
204
|
+
<summary><strong>Advanced: raw score arrays (low-level engine)</strong></summary>
|
|
205
|
+
|
|
206
|
+
`compare()` is a wrapper around a lower-level engine, `analyze()`, which operates directly on `BenchmarkResult` / `MultiModelBenchmark` objects (numpy score arrays) rather than a DataFrame. Reach for this path only if you already have scores as arrays and don't want to build a DataFrame first. Most use cases should use `compare()` above.
|
|
207
|
+
|
|
208
|
+
```python
|
|
209
|
+
import numpy as np
|
|
210
|
+
import evalstats as estats
|
|
211
|
+
|
|
212
|
+
# Example raw scores for 4 templates × 3 inputs (single run, single evaluator)
|
|
213
|
+
your_scores = [
|
|
214
|
+
[0.91, 0.88, 0.86],
|
|
215
|
+
[0.90, 0.89, 0.84],
|
|
216
|
+
[0.85, 0.82, 0.80],
|
|
217
|
+
[0.79, 0.76, 0.74],
|
|
218
|
+
]
|
|
219
|
+
n_templates, n_inputs = 4, 3
|
|
220
|
+
|
|
221
|
+
# scores shape: (n_templates, n_inputs, n_runs, n_evaluators)
|
|
222
|
+
scores = np.array(your_scores).reshape(n_templates, n_inputs, 1, 1)
|
|
223
|
+
|
|
224
|
+
result = estats.BenchmarkResult(
|
|
225
|
+
scores=scores,
|
|
226
|
+
template_labels=["Minimal", "Instructive", "Few-shot", "Chain-of-thought"],
|
|
227
|
+
input_labels=[f"input_{i}" for i in range(n_inputs)],
|
|
228
|
+
)
|
|
229
|
+
|
|
230
|
+
analysis = estats.analyze(result, reference="grand_mean", n_bootstrap=5_000)
|
|
231
|
+
analysis.summary() # same terminal report as ComparisonResult.summary()
|
|
232
|
+
```
|
|
233
|
+
|
|
234
|
+
If you want this lower-level path from a DataFrame (e.g. to inspect the raw `BenchmarkResult` object, or fine-tune `strict_complete_design`), use `from_dataframe()` instead of `load_from()`. It returns the array-based `BenchmarkResult` / `MultiModelBenchmark` plus an optional `DataLoadReport`, a data-quality log of coercions/repairs made while parsing:
|
|
235
|
+
|
|
236
|
+
```python
|
|
237
|
+
import evalstats as estats
|
|
238
|
+
|
|
239
|
+
benchmark, load_report = estats.from_dataframe(
|
|
240
|
+
df, format="auto", repair=True, strict_complete_design=True, return_report=True,
|
|
241
|
+
)
|
|
242
|
+
for line in load_report.to_lines():
|
|
243
|
+
print(line)
|
|
244
|
+
|
|
245
|
+
analysis = estats.analyze(benchmark)
|
|
246
|
+
analysis.summary()
|
|
247
|
+
```
|
|
248
|
+
|
|
249
|
+
To visualize absolute prompt performance directly from a `BenchmarkResult`, bypassing `analyze()`:
|
|
250
|
+
|
|
251
|
+
```python
|
|
252
|
+
fig = estats.plot_point_estimates(result)
|
|
253
|
+
fig.savefig("mean_performance.png", dpi=150, bbox_inches="tight")
|
|
254
|
+
```
|
|
255
|
+
|
|
256
|
+
</details>
|
|
257
|
+
|
|
258
|
+
## PPI-Corrected Inference (Means, CIs, and Tests)
|
|
259
|
+
|
|
260
|
+
PPI (Prediction-Powered Inference) lets you use lots of cheap LLM judgments plus a smaller set of human labels to correct measurement error from the LLM judge, giving you corrected estimates and uncertainty that better reflect what you'd have gotten from a fully human-labeled study (Angelopoulos et al., 2023). Most corrections use PPIBoot (bootstrap variant of PPI; Zrnic, 2024), battle-tested via simulations (see [`simulations/sim_type_i_calibration.py`](simulations/sim_type_i_calibration.py)).
|
|
261
|
+
|
|
262
|
+
> [!IMPORTANT]
|
|
263
|
+
> **Which items get a human label must be chosen uniformly at random.** PPI correction assumes the labeled subset is representative of the full dataset. If your labeling process instead targets specific items — e.g. "always double-check the borderline or highest-scoring responses" — that's missing-not-at-random (MNAR) selection on the outcome, and PPI correction can stay badly miscalibrated **no matter how many items you label**. This isn't ordinary small-sample noise that more labels fixes; confirmed in simulation to persist from 15 up through 300 labeled items out of 400. See `evalstats.ppi.correct`'s docstring for the full analysis, and use `evalstats label` (below) to draw a compliant random sample.
|
|
264
|
+
|
|
265
|
+
### Comparing models with corrected LLM judge evals via `compare(..., alignment=...)`
|
|
266
|
+
|
|
267
|
+
```python
|
|
268
|
+
import evalstats as es
|
|
269
|
+
|
|
270
|
+
# Dataframe columns include: model item llm_score human_score (NaN for unlabeled rows)
|
|
271
|
+
evaldata = es.load_from(df)
|
|
272
|
+
|
|
273
|
+
alignment = es.judge_alignment(evaldata, llm_metric="llm_score", human_groundtruth="human_score")
|
|
274
|
+
|
|
275
|
+
result = es.compare(
|
|
276
|
+
evaldata, factors="model", metric="llm_score", alignment={"llm_score": alignment},
|
|
277
|
+
)
|
|
278
|
+
result.summary()
|
|
279
|
+
```
|
|
280
|
+
|
|
281
|
+
### T-test PPI-correction via `evalstats.tests.ttest`
|
|
282
|
+
|
|
283
|
+
```python
|
|
284
|
+
import evalstats as es
|
|
285
|
+
|
|
286
|
+
res = es.tests.ttest(
|
|
287
|
+
a=llm_a, b=llm_b,
|
|
288
|
+
a_lab=human_a, # same length as llm_a, NaN where unlabeled
|
|
289
|
+
b_lab=human_b, # same length as llm_b, NaN where unlabeled
|
|
290
|
+
paired=False, print_result=False,
|
|
291
|
+
)
|
|
292
|
+
print(res.p_value, res.corrected_p_value, res.corrected_ci)
|
|
293
|
+
```
|
|
294
|
+
|
|
295
|
+
<details>
|
|
296
|
+
<summary><strong>More PPI-corrected tests: Mann-Whitney U, Wilcoxon, one-way ANOVA</strong></summary>
|
|
297
|
+
|
|
298
|
+
**Mann-Whitney U** (`evalstats.tests.mannwhitney`): nonparametric two-group comparison based on relative ranks rather than assuming normally distributed scores:
|
|
299
|
+
|
|
300
|
+
```python
|
|
301
|
+
res = es.tests.mannwhitney(x=llm_x, y=llm_y, x_lab=human_x, y_lab=human_y, print_result=False)
|
|
302
|
+
print(res.p_value, res.corrected_p_value, res.corrected_ci)
|
|
303
|
+
```
|
|
304
|
+
|
|
305
|
+
**Wilcoxon signed-rank** (`evalstats.tests.wilcoxon`): nonparametric paired test for matched observations (before/after, A/B on the same items, etc.):
|
|
306
|
+
|
|
307
|
+
```python
|
|
308
|
+
res = es.tests.wilcoxon(x=llm_before, y=llm_after, x_lab=human_before, y_lab=human_after, print_result=False)
|
|
309
|
+
print(res.p_value, res.corrected_p_value, res.corrected_ci)
|
|
310
|
+
```
|
|
311
|
+
|
|
312
|
+
**One-way ANOVA** (`evalstats.tests.anova_oneway`): more than two groups; pass `repeated=True` for repeated-measures (same subjects across conditions):
|
|
313
|
+
|
|
314
|
+
```python
|
|
315
|
+
res = es.tests.anova_oneway(
|
|
316
|
+
llm_g1, llm_g2, llm_g3,
|
|
317
|
+
groups_lab=[human_g1, human_g2, human_g3], repeated=False, print_result=False,
|
|
318
|
+
)
|
|
319
|
+
print(res.p_value, res.corrected_p_value, res.corrected_ci)
|
|
320
|
+
```
|
|
321
|
+
|
|
322
|
+
</details>
|
|
323
|
+
|
|
324
|
+
## CLI Reference
|
|
325
|
+
|
|
326
|
+
```bash
|
|
327
|
+
evalstats analyze results.csv # full statistical report from a CSV/XLSX file
|
|
328
|
+
evalstats label results.csv # draw a random, MCAR-compliant sample of items for human labeling
|
|
329
|
+
```
|
|
330
|
+
|
|
331
|
+
`evalstats label` picks a uniformly random sample of items per condition (respecting the PPI sample-size floors: 15 minimum, 30 recommended) and writes a CSV/XLSX with a `human_<metric>` column ready for grading, the safe way to build the labeled subset `alignment=`/`*_lab` needs above. Run `evalstats label --help` for the full option list.
|
|
332
|
+
|
|
333
|
+
## Examples
|
|
334
|
+
|
|
335
|
+
`examples/` has 25+ runnable, self-contained scripts covering common workflows: synthetic and OpenAI-backed benchmarks, multi-run comparisons, PPI-corrected judge alignment, factorial designs, and reliability/robustness demos. From the repository root:
|
|
336
|
+
|
|
337
|
+
```bash
|
|
338
|
+
python examples/synthetic_mean_advantage.py # no API key needed
|
|
339
|
+
python examples/sentiment.py # OpenAI sentiment benchmark
|
|
340
|
+
python examples/sentiment_multirun.py # captures run-to-run variability
|
|
341
|
+
python examples/compare_models_multirun.py # multi-model comparison across prompts
|
|
342
|
+
python examples/compare_alignment_ppi.py # PPI-corrected judge comparison
|
|
343
|
+
```
|
|
344
|
+
|
|
345
|
+
OpenAI-powered examples require `OPENAI_API_KEY` set in your environment, but the model calls are easy to swap for whichever provider you prefer.
|
|
346
|
+
|
|
347
|
+
## Mixed effects models (LMM)
|
|
348
|
+
|
|
349
|
+
> [!IMPORTANT]
|
|
350
|
+
> Mixed effects analysis is experimental, currently offering only graceful handling of missing data (assumed reasonably random). Use `method="lmm"` if you need robustness to missing (`NaN`) cells; factor decomposition across multiple input factors is planned.
|
|
351
|
+
|
|
352
|
+
`evalstats` supports mixed-effects models (`score ~ template + (1|input)`) for missing data and multi-factor decomposition. The default backend is pure-Python `statsmodels`, no extra setup required:
|
|
353
|
+
|
|
354
|
+
```python
|
|
355
|
+
analysis = estats.analyze(result, method="lmm")
|
|
356
|
+
```
|
|
357
|
+
|
|
358
|
+
This fits the model with REML, computes Wald CIs via the delta method, and estimates rank distributions by parametric simulation.
|
|
359
|
+
|
|
360
|
+
<details>
|
|
361
|
+
<summary><strong>Optional backend: pymer4 (requires R)</strong></summary>
|
|
362
|
+
|
|
363
|
+
For Satterthwaite degrees of freedom and `emmeans`-based pairwise contrasts (R's gold standard for mixed models), pass `backend="pymer4"`:
|
|
364
|
+
|
|
365
|
+
```python
|
|
366
|
+
analysis = estats.analyze(result, method="lmm", backend="pymer4")
|
|
367
|
+
```
|
|
368
|
+
|
|
369
|
+
This requires a working R installation with:
|
|
370
|
+
|
|
371
|
+
```r
|
|
372
|
+
install.packages(c("lme4", "emmeans", "tibble", "broom", "broom.mixed", "lmerTest", "report", "car"))
|
|
373
|
+
```
|
|
374
|
+
|
|
375
|
+
Then `pip install "evalstats[lmm]"`. If your environment needs manual dependency pinning, this is the tested equivalent:
|
|
376
|
+
|
|
377
|
+
```bash
|
|
378
|
+
pip install "pymer4>=0.9" great_tables joblib rpy2 polars scikit-learn formulae pyarrow
|
|
379
|
+
```
|
|
380
|
+
|
|
381
|
+
Installation details may differ on your system.
|
|
382
|
+
|
|
383
|
+
</details>
|
|
384
|
+
|
|
385
|
+
## Reproducibility: Monte Carlo simulations
|
|
386
|
+
|
|
387
|
+
Claims in this README and on the [`evalstats` site](https://statsforevals.com/) like "verified in our simulations" are backed by a runnable harness in [`simulations/harness/`](simulations/harness):
|
|
388
|
+
|
|
389
|
+
```bash
|
|
390
|
+
python -m simulations.harness.cli --list-cases
|
|
391
|
+
python -m simulations.harness.cli --official-tests
|
|
392
|
+
python -m simulations.harness.cli ci_single --reps 50 --sizes 10 20
|
|
393
|
+
python -m simulations.harness.cli pvalues --mode ppi --tests ttest wilcoxon anova_rep
|
|
394
|
+
```
|
|
395
|
+
|
|
396
|
+
`--official-tests` brings up a CLI to run each case's canonical, full-scale preset, writing results plus a `manifest.json` (args, output paths, key metrics, pass/fail) to `simulations/out/official_<timestamp>/`. See [`simulations/harness/README.md`](simulations/harness/README.md) for the full case list and scenario library. Note that each simulation can take a long time to run. Even parallelized across many cores, official-scale runs can take hours.
|
|
397
|
+
|
|
398
|
+
- `ci_single` / `ci_paired`: coverage and width of confidence interval methods across synthetic distributions and real benchmark data (OpenEval, Inspect AI).
|
|
399
|
+
- `pvalues --mode pairwise` / `--mode multiarm`: Type-I error and power for pairwise and multi-arm comparisons, including FWER correction strategies.
|
|
400
|
+
- `pvalues --mode ppi`: Type-I error calibration and power for every PPI-corrected test in `evalstats.tests`, swept across judge-bias severity, label fraction, and MNAR-labeling scenarios.
|
|
401
|
+
|
|
402
|
+
## Motivation
|
|
403
|
+
|
|
404
|
+
Most eval tools in the LLM evaluation space don't help users perform *any* statistical tests. They present bar charts of average performance, and developers glance at the chart and decide "prompt/model A is better than B." But was it really? Relying on bar charts and averages alone can easily lead to erroneous conclusions: B might be more robust than A, or perform better on an important data subset, or there might not be enough data to conclude either way.
|
|
405
|
+
|
|
406
|
+
People do evals this way because they don't have the time, tools, or statistical knowledge to do better. Often they don't even know there's a better way. `evalstats` aims to rectify this with simple, powerful defaults: throw us your data, and we'll run the stats and plot the results for you.
|
|
407
|
+
|
|
408
|
+
## Development and Contributions
|
|
409
|
+
|
|
410
|
+
For package build, release validation, and maintainer workflows, see [DEVELOPMENT.md](DEVELOPMENT.md).
|
|
411
|
+
|
|
412
|
+
We welcome contributions, especially refinements to our statistical methods. If you're proposing a new correction, CI method, or a fix to an existing one, battle-test it against the [simulation harness](#reproducibility-monte-carlo-simulations) first. Add or extend a scenario and confirm your change holds up on Type-I error and power, not just on the case that motivated it, before opening a PR.
|
|
413
|
+
|
|
414
|
+
## Citation
|
|
415
|
+
|
|
416
|
+
`evalstats` doesn't have a paper yet. One covering the full simulation-backed method validation is forthcoming. Until then, please cite the GitHub repository:
|
|
417
|
+
|
|
418
|
+
```bibtex
|
|
419
|
+
@software{arawjo_evalstats,
|
|
420
|
+
author = {Arawjo, Ian},
|
|
421
|
+
title = {evalstats: Statistically Sound Analysis for LLM Evaluations},
|
|
422
|
+
url = {https://github.com/ianarawjo/evalstats},
|
|
423
|
+
year = {2026}
|
|
424
|
+
}
|
|
425
|
+
```
|
|
426
|
+
|
|
427
|
+
or in prose: Ian Arawjo, *evalstats* (GitHub: [ianarawjo/evalstats](https://github.com/ianarawjo/evalstats)).
|
|
428
|
+
|
|
429
|
+
## License
|
|
430
|
+
|
|
431
|
+
This repository uses two licenses:
|
|
432
|
+
|
|
433
|
+
- **`evalstats` package** (everything outside `website/`): [MIT](LICENSE).
|
|
434
|
+
- **Stats for Evals Website** (everything in `website/`): [CC BY-NC-ND 4.0](website/LICENSE). You may share it with attribution non-commercially, but commercial use and derivative works are not permitted.
|