evalsuite-python 0.1.0__py3-none-any.whl

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -0,0 +1,247 @@
1
+ Metadata-Version: 2.5
2
+ Name: evalsuite-python
3
+ Version: 0.1.0
4
+ Summary: A unified Python framework for machine-learning, clinical, statistical, segmentation, object-detection, uncertainty, and model evaluation.
5
+ Project-URL: Homepage, https://evalsuite-nine.vercel.app
6
+ Project-URL: Documentation, https://evalsuite-nine.vercel.app/docs
7
+ Project-URL: Source, https://github.com/mkcs28/evalsuite-python
8
+ Project-URL: Issues, https://github.com/mkcs28/evalsuite-python/issues
9
+ Project-URL: Changelog, https://github.com/mkcs28/evalsuite-python/blob/main/CHANGELOG.md
10
+ Author: Manoj Kumar C S
11
+ Maintainer: Manoj Kumar C S
12
+ License-Expression: MIT
13
+ License-File: LICENSE
14
+ Keywords: classification,evaluation,machine learning,metrics,regression,reproducibility,statistics
15
+ Classifier: Development Status :: 5 - Production/Stable
16
+ Classifier: Intended Audience :: Developers
17
+ Classifier: Intended Audience :: Science/Research
18
+ Classifier: Operating System :: OS Independent
19
+ Classifier: Programming Language :: Python :: 3
20
+ Classifier: Programming Language :: Python :: 3 :: Only
21
+ Classifier: Programming Language :: Python :: 3.9
22
+ Classifier: Programming Language :: Python :: 3.10
23
+ Classifier: Programming Language :: Python :: 3.11
24
+ Classifier: Programming Language :: Python :: 3.12
25
+ Classifier: Programming Language :: Python :: 3.13
26
+ Classifier: Programming Language :: Python :: 3.14
27
+ Classifier: Topic :: Scientific/Engineering
28
+ Classifier: Topic :: Scientific/Engineering :: Artificial Intelligence
29
+ Classifier: Typing :: Typed
30
+ Requires-Python: >=3.9
31
+ Requires-Dist: numpy>=1.22
32
+ Requires-Dist: pandas>=1.4
33
+ Requires-Dist: scipy>=1.8
34
+ Provides-Extra: all
35
+ Requires-Dist: matplotlib>=3.5; extra == 'all'
36
+ Provides-Extra: dev
37
+ Requires-Dist: build; extra == 'dev'
38
+ Requires-Dist: hypothesis>=6.80; extra == 'dev'
39
+ Requires-Dist: matplotlib>=3.5; extra == 'dev'
40
+ Requires-Dist: mypy>=1.10; extra == 'dev'
41
+ Requires-Dist: pandas-stubs; extra == 'dev'
42
+ Requires-Dist: pip-audit; extra == 'dev'
43
+ Requires-Dist: pytest-cov>=4; extra == 'dev'
44
+ Requires-Dist: pytest>=7; extra == 'dev'
45
+ Requires-Dist: ruff>=0.6; extra == 'dev'
46
+ Requires-Dist: scikit-learn>=1.2; extra == 'dev'
47
+ Requires-Dist: statsmodels>=0.13; extra == 'dev'
48
+ Requires-Dist: twine; extra == 'dev'
49
+ Provides-Extra: plot
50
+ Requires-Dist: matplotlib>=3.5; extra == 'plot'
51
+ Description-Content-Type: text/markdown
52
+
53
+ # EvalSuite
54
+
55
+ [![CI](https://github.com/mkcs28/evalsuite-python/actions/workflows/ci.yml/badge.svg)](https://github.com/mkcs28/evalsuite-python/actions/workflows/ci.yml)
56
+ [![PyPI](https://img.shields.io/pypi/v/evalsuite-python)](https://pypi.org/project/evalsuite-python/)
57
+ [![Python](https://img.shields.io/pypi/pyversions/evalsuite-python)](https://pypi.org/project/evalsuite-python/)
58
+ [![License: MIT](https://img.shields.io/badge/License-MIT-blue.svg)](LICENSE)
59
+
60
+ **Unified, reproducible evaluation for machine learning and research.**
61
+
62
+ EvalSuite brings classification and regression metrics (with clinical, statistical, segmentation and
63
+ object-detection evaluation on the roadmap) into one consistent, validated, documented framework.
64
+
65
+ > **Status: stable (0.1.0).** Every item on the 0.1.0 roadmap is implemented and verified.
66
+
67
+ ## Installation
68
+
69
+ ```bash
70
+ pip install evalsuite-python
71
+ ```
72
+
73
+ The package is installed as `evalsuite-python` and imported as `evalsuite`:
74
+
75
+ ```python
76
+ import evalsuite as es
77
+ ```
78
+
79
+ ## Why EvalSuite
80
+
81
+ - **One consistent API.** Every metric returns a result object that behaves like a number and exports to
82
+ JSON, pandas, Markdown and LaTeX.
83
+ - **Explicit conventions.** Averaging, label order, the positive class and zero-division behaviour are stated
84
+ and recorded in every result, never silently assumed.
85
+ - **Validated.** Each metric is tested against scikit-learn where definitions coincide, plus property-based
86
+ tests and edge cases.
87
+ - **Documented.** Every metric carries its definition, formula, range, input requirements and references,
88
+ available programmatically through `metric_info()`.
89
+ - **Efficient.** `evaluate()` validates inputs once and computes the confusion matrix once for all metrics.
90
+ - **Lightweight.** Requires only NumPy, SciPy and pandas.
91
+
92
+ ## Quick start
93
+
94
+ ```python
95
+ import evalsuite as es
96
+
97
+ y_true = [0, 1, 1, 0, 1, 0]
98
+ y_pred = [0, 1, 0, 0, 1, 1]
99
+ y_prob = [0.1, 0.9, 0.4, 0.2, 0.8, 0.6]
100
+
101
+ result = es.evaluate(y_true, y_pred, y_prob=y_prob)
102
+ print(result.summary())
103
+
104
+ result["f1"] # MetricResult(f1=0.666667)
105
+ f"{result['mcc']:.3f}" # '0.333'
106
+ result.to_latex(caption="Test-set performance")
107
+ result.to_dataframe()
108
+
109
+ es.f1(y_true, y_pred) # individual metrics
110
+ es.roc_auc(y_true, y_prob)
111
+ es.metric_info("classification.mcc").formula # documentation
112
+ es.list_metrics("regression")
113
+ ```
114
+
115
+ ## Comparing models
116
+
117
+ ```python
118
+ result = es.compare(
119
+ y_true,
120
+ {"logistic": pred_lr, "forest": pred_rf, "boosting": pred_gb},
121
+ probabilities={"logistic": prob_lr, "forest": prob_rf, "boosting": prob_gb},
122
+ random_state=0,
123
+ )
124
+ print(result.summary()) # estimates with 95% CIs, paired tests, Holm-adjusted p-values
125
+ result.to_latex(label="tab:models")
126
+
127
+ es.bootstrap_ci("f1", y_true, y_pred, average="macro", random_state=0) # BCa interval for any metric
128
+ es.accuracy_ci(y_true, y_pred) # Wilson interval
129
+ es.delong_test(y_true, prob_a, prob_b) # two correlated AUCs
130
+ es.mcnemar_test(y_true, pred_a, pred_b)
131
+ ```
132
+
133
+ Every model is evaluated on the same bootstrap resamples, so differences are paired. Accuracy is compared
134
+ with McNemar's test, binary ROC AUC with DeLong's test and other metrics with a paired bootstrap test;
135
+ p-values are adjusted for multiple comparisons (Holm by default).
136
+
137
+ ## Classification report
138
+
139
+ ```python
140
+ report = es.classification_report(y_true, y_pred)
141
+ print(report) # per-class precision, recall, F1, specificity, support + averages
142
+ report.save("report.html") # also .csv .md .tex .json .txt
143
+ ```
144
+
145
+ ## Plots
146
+
147
+ ```bash
148
+ pip install "evalsuite-python[plot]" # adds matplotlib; importing evalsuite never loads it
149
+ ```
150
+
151
+ ```python
152
+ es.plot.roc(y_true, {"logistic": prob_lr, "forest": prob_rf}) # AUC in the legend
153
+ es.plot.pr(y_true, prob) # AP and the prevalence line
154
+ es.plot.calibration(y_true, prob) # reliability diagram, ECE, Brier
155
+ es.plot.confusion_matrix(y_true, y_pred, normalize="true")
156
+ es.plot.residuals(y_reg, pred_reg) # or kind="predicted"
157
+ es.plot.comparison(es.compare(...)) # forest plot with CIs
158
+ ```
159
+
160
+ Each function returns a matplotlib `Axes` (pass `ax=` to draw into your own figure). The numbers shown are
161
+ computed with EvalSuite's metrics, so plots and tables always agree. Several models get distinct colours
162
+ *and* line styles, so figures stay readable in greyscale print.
163
+
164
+ ## Exports
165
+
166
+ Every result (`evaluate`, `classification_report`, `compare`, single metrics) exports to `summary()`,
167
+ `to_json()`, `to_csv()`, `to_dataframe()`, `to_markdown()`, `to_latex()` and `to_html()`, and `save(path)` picks
168
+ the format from the extension. HTML pages are standalone (inline CSS, no scripts) and escape all text.
169
+
170
+ ## Command line
171
+
172
+ ```bash
173
+ evalsuite evaluate predictions.csv --y-true label --y-pred pred --y-prob prob
174
+ evalsuite report predictions.csv --y-true label --y-pred pred -o report.html
175
+ evalsuite compare predictions.csv --y-true label --pred lr=pred_lr --pred rf=pred_rf \
176
+ --prob lr=p_lr --prob rf=p_rf --plot comparison.png
177
+ evalsuite plot roc predictions.csv --y-true label --y-prob prob -o roc.png
178
+ evalsuite metrics --category classification
179
+ evalsuite info classification.mcc
180
+ evalsuite benchmark --quick
181
+ ```
182
+
183
+ Input files can be CSV, TSV, Parquet or JSON. Output format follows `--format` or the `-o` extension
184
+ (text, json, csv, markdown, latex, html). Errors are reported in one line with exit code 2.
185
+
186
+ ## Performance
187
+
188
+ Benchmarked against scikit-learn on the same data (fastest of 5 runs; Python 3.12, NumPy 2.5,
189
+ scikit-learn 1.9, Linux x86_64). Every result agrees with scikit-learn to floating-point rounding
190
+ (largest difference 1.1e-16).
191
+
192
+ | Case | n | EvalSuite (ms) | scikit-learn (ms) | Speed-up | Peak memory EvalSuite / sklearn (MiB) |
193
+ | --- | ---: | ---: | ---: | ---: | ---: |
194
+ | 8 binary label metrics via `evaluate()` | 1,000 | 0.38 | 11.70 | **31.2×** | 0.04 / 0.05 |
195
+ | 8 binary label metrics via `evaluate()` | 100,000 | 10.9 | 116.9 | **10.7×** | 3.2 / 3.1 |
196
+ | 8 binary label metrics via `evaluate()` | 1,000,000 | 108.8 | 1043.6 | **9.6×** | 31.5 / 30.5 |
197
+ | macro F1, 10 classes | 1,000,000 | 88.4 | 139.3 | **1.58×** | 30.5 / 21.8 |
198
+ | ROC AUC, binary | 1,000,000 | 247.4 | 352.5 | **1.42×** | 91.6 / 76.3 |
199
+ | MAE, MSE, RMSE, R² via `evaluate()` | 1,000 | 0.12 | 0.90 | **7.2×** | 0.03 / 0.02 |
200
+ | MAE, MSE, RMSE, R² via `evaluate()` | 1,000,000 | 29.3 | 20.1 | 0.69× | 22.9 / 15.3 |
201
+
202
+ `evaluate()` validates inputs once and builds the confusion matrix once for all metrics, which is where the
203
+ speed-up comes from. Large regression arrays are slower because EvalSuite checks every value for NaN,
204
+ infinity, shape and dtype before computing. Reproduce on your machine with `evalsuite benchmark`; full
205
+ table and notes in
206
+ [BENCHMARKS.md](https://github.com/mkcs28/evalsuite-python/blob/main/BENCHMARKS.md).
207
+
208
+ ## Metrics in this release
209
+
210
+ **Classification** (binary, multiclass, multilabel; micro/macro/weighted/samples/per-class averaging;
211
+ sample weights): accuracy, balanced accuracy, precision, recall, specificity, NPV, F1, F-beta, Jaccard,
212
+ MCC, Cohen's kappa (unweighted, linear, quadratic), Hamming loss, confusion matrix, ROC AUC (binary,
213
+ one-vs-rest, one-vs-one), average precision, ROC and PR curves, log loss, Brier score, top-k accuracy,
214
+ calibration curve and expected calibration error.
215
+
216
+ **Regression** (single and multi-output; sample weights): MAE, MSE, RMSE, R², adjusted R², MAPE, sMAPE,
217
+ MSLE, RMSLE, median absolute error, explained variance, max error, mean bias error, quantile (pinball)
218
+ loss, Huber loss, relative absolute error, relative squared error.
219
+
220
+ ## Conventions
221
+
222
+ - `average="auto"` resolves to `"binary"` for binary targets and `"macro"` otherwise; the resolved value
223
+ is stored in `result.params["average"]`.
224
+ - Labels are sorted unless you pass `labels=[...]`; that order defines per-class outputs and the columns
225
+ of 2-D `y_prob`.
226
+ - Undefined ratios (zero denominators) return 0 **with an `UndefinedMetricWarning`**; pass
227
+ `zero_division=np.nan` to propagate NaN, or `0`/`1` to choose silently.
228
+ - Domain violations raise clear errors instead of being patched over (for example MAPE with zero targets).
229
+
230
+ ## Development
231
+
232
+ ```bash
233
+ python -m venv .venv && source .venv/bin/activate
234
+ pip install -e ".[dev]"
235
+ pytest --cov=evalsuite
236
+ ruff check . && ruff format --check . && mypy
237
+ ```
238
+
239
+ ## Links
240
+
241
+ - PyPI: https://pypi.org/project/evalsuite-python/
242
+ - Website and documentation: https://evalsuite-nine.vercel.app
243
+ - Website source: https://github.com/mkcs28/evalsuite
244
+
245
+ ## License
246
+
247
+ MIT. See [LICENSE](LICENSE).
@@ -0,0 +1,35 @@
1
+ evalsuite/__init__.py,sha256=giC_4Uq1bnYCuDuw7lxuxDAp5vFxN7RLkdNTnBt-pPM,3114
2
+ evalsuite/__main__.py,sha256=Nq90_8VlUN1mif-AVgX-kZRL63bd94F5-F7F9UkVZD4,117
3
+ evalsuite/api.py,sha256=ggEc_LM7guktX6pnApYIf-n8AHHMujdiWriqMy6JZdE,9578
4
+ evalsuite/benchmarks.py,sha256=6y0K6-5vwIc7D3ACItQXLSqVCRX6EvqWdVPttVlBkNU,10018
5
+ evalsuite/plot.py,sha256=QqQgcEOjIGkpMSwfgXjO2PzOaLBmj8-dTZrzlqzOHFw,11416
6
+ evalsuite/py.typed,sha256=47DEQpj8HBSa-_TImW-5JCeuQeRkm5NMpJWZG3hSuFU,0
7
+ evalsuite/reporting.py,sha256=-LISv9vvEnltT26GtzLjVlxOCcAwmjv4-HyB8_1UF5o,9615
8
+ evalsuite/version.py,sha256=NhOfuA4GsTmRzoTsxKIfwdcZ0oQ5D93WpPm-V0Nw8ck,98
9
+ evalsuite/classification/__init__.py,sha256=ErBIXbBGnhfW6v-1Cn5SegiERoyxNlmO1tEKcLM15mA,833
10
+ evalsuite/classification/_common.py,sha256=3ANV33ix_mnOzvOAG8glyJO-iYVhhTGO0S3YG0ZddB8,4543
11
+ evalsuite/classification/metrics.py,sha256=8-RPjzuFUKJYlXndrDv6kXlzRdx6aR2lHp7mzKsas2g,36969
12
+ evalsuite/cli/__init__.py,sha256=1x48hORLcg_Is2X08A7CPBT6VDJ1QL94zaJP1HTFxUY,96
13
+ evalsuite/cli/main.py,sha256=ez_P7SKuuOin9E5CAQ0orA5GSHn8sE-oH248RiFsgpI,15021
14
+ evalsuite/core/__init__.py,sha256=oDHl5Dc2GoZf1tCtJueMQ2KcmOq7zhzMZT9XF4uVC6U,93
15
+ evalsuite/core/context.py,sha256=gvRpH3b-S2oAgAGW9W0XWnBVbuydQP3UB9rgWQNM60g,7874
16
+ evalsuite/core/exceptions.py,sha256=9eZI_Lmgwc9fjkSUY9s3fwjAb7KoqgAlcYetrHRhbBg,1503
17
+ evalsuite/core/export.py,sha256=6M_ktQzW3E7usOrwqv5f5zOEdoi48JD6CvgyRpDkGvM,4303
18
+ evalsuite/core/registry.py,sha256=3l2_o2PlckuKY89HxfnxPLf-sDaCe_-McUcAcUmXusk,2884
19
+ evalsuite/core/result.py,sha256=L_fQZRGRgVgzII2jZb_jpRBK2jy-CdO5u9s-6pUk9fE,14927
20
+ evalsuite/core/types.py,sha256=r6bhUk7CYAcrlQzjmf-6xWQsNqO3EOipZD-e-2yAJTI,720
21
+ evalsuite/core/validation.py,sha256=EIAd6m5DIguFM4wLUNVissEYqvjylYEuLGrPq8hOUpw,8350
22
+ evalsuite/regression/__init__.py,sha256=t2XEHUw6e5tyWBHcEeiog2WWd-wwaxf4m5AIX8dJyng,571
23
+ evalsuite/regression/metrics.py,sha256=jVdzmegdpBdyJNCSm1dFR1E-9VHDmQmTVtyrqyjk3KY,22459
24
+ evalsuite/stats/__init__.py,sha256=PcYS9tZh-IptSXx3sX7shmUUdNYsQuD-G7Xkgb3FHiY,745
25
+ evalsuite/stats/_resolve.py,sha256=cK6qdwM7P_mSjzHDEVcV4SekBS58VGGYQ8PuKptgqQs,3768
26
+ evalsuite/stats/compare.py,sha256=DmC53p_IlL5ecP_7sGEYV5fszKjLIyN3dN0JMUp_ri4,17746
27
+ evalsuite/stats/effect.py,sha256=O0-pxY_6o25G3w7ymDar3xb_dtCOO0kG0sLv-A_Yb0Q,4884
28
+ evalsuite/stats/intervals.py,sha256=w2bLuFt3p-TTN_ALH1uKPmDMLcTzBBoV2YvcP-ENAOk,13709
29
+ evalsuite/stats/paired.py,sha256=YK3AFndChKeCEpK512u56VtRlKvhlGdkFLHHYe77zmc,8244
30
+ evalsuite/stats/results.py,sha256=nqXHUiW8l0fQJkbMiunU6Gz-2be1NApZ6Fif2B1K_Yw,4407
31
+ evalsuite_python-0.1.0.dist-info/METADATA,sha256=PMXY0G7f1HRp5SK9mBIMnm8cFBY29t80b1M3QN2ym5s,10743
32
+ evalsuite_python-0.1.0.dist-info/WHEEL,sha256=W3fkpkm7-wf9vBI5Z-7s0eWkeM-spu78I8Neb98DeEg,87
33
+ evalsuite_python-0.1.0.dist-info/entry_points.txt,sha256=YYIwisn4L4sQt3ntzbM0_Z4yZE2KCu3GJ6KfWL1kIwc,54
34
+ evalsuite_python-0.1.0.dist-info/licenses/LICENSE,sha256=GGh6ACzs6Gih4IlfbWxzHvkNKC1E9EQ3eVWW0Njpafk,1072
35
+ evalsuite_python-0.1.0.dist-info/RECORD,,
@@ -0,0 +1,4 @@
1
+ Wheel-Version: 1.0
2
+ Generator: hatchling 1.32.4
3
+ Root-Is-Purelib: true
4
+ Tag: py3-none-any
@@ -0,0 +1,2 @@
1
+ [console_scripts]
2
+ evalsuite = evalsuite.cli.main:main
@@ -0,0 +1,21 @@
1
+ MIT License
2
+
3
+ Copyright (c) 2026 Manoj Kumar C S
4
+
5
+ Permission is hereby granted, free of charge, to any person obtaining a copy
6
+ of this software and associated documentation files (the "Software"), to deal
7
+ in the Software without restriction, including without limitation the rights
8
+ to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
9
+ copies of the Software, and to permit persons to whom the Software is
10
+ furnished to do so, subject to the following conditions:
11
+
12
+ The above copyright notice and this permission notice shall be included in all
13
+ copies or substantial portions of the Software.
14
+
15
+ THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
16
+ IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
17
+ FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
18
+ AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
19
+ LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
20
+ OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
21
+ SOFTWARE.