ssme-lite 0.1.0__tar.gz
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- ssme_lite-0.1.0/CHANGELOG.md +9 -0
- ssme_lite-0.1.0/CONTRIBUTING.md +14 -0
- ssme_lite-0.1.0/LICENSE +21 -0
- ssme_lite-0.1.0/MANIFEST.in +11 -0
- ssme_lite-0.1.0/PKG-INFO +120 -0
- ssme_lite-0.1.0/README.md +80 -0
- ssme_lite-0.1.0/docs/api.md +117 -0
- ssme_lite-0.1.0/docs/guides/cli.md +41 -0
- ssme_lite-0.1.0/docs/guides/inputs.md +67 -0
- ssme_lite-0.1.0/docs/guides/reports.md +49 -0
- ssme_lite-0.1.0/docs/index.md +32 -0
- ssme_lite-0.1.0/docs/installation.md +52 -0
- ssme_lite-0.1.0/docs/method.md +35 -0
- ssme_lite-0.1.0/docs/quickstart.md +55 -0
- ssme_lite-0.1.0/docs/releasing.md +50 -0
- ssme_lite-0.1.0/docs/troubleshooting.md +18 -0
- ssme_lite-0.1.0/docs/tutorials/civilcomments.md +52 -0
- ssme_lite-0.1.0/docs/tutorials/pneumoniamnist.md +54 -0
- ssme_lite-0.1.0/docs/tutorials/transfer.md +51 -0
- ssme_lite-0.1.0/examples/quickstart.py +34 -0
- ssme_lite-0.1.0/examples/sklearn_models.py +36 -0
- ssme_lite-0.1.0/mkdocs.yml +30 -0
- ssme_lite-0.1.0/pyproject.toml +49 -0
- ssme_lite-0.1.0/setup.cfg +4 -0
- ssme_lite-0.1.0/ssme_lite/__init__.py +9 -0
- ssme_lite-0.1.0/ssme_lite/benchmark/__init__.py +13 -0
- ssme_lite-0.1.0/ssme_lite/benchmark/datasets.py +87 -0
- ssme_lite-0.1.0/ssme_lite/benchmark/runner.py +146 -0
- ssme_lite-0.1.0/ssme_lite/cli.py +50 -0
- ssme_lite-0.1.0/ssme_lite/estimation/__init__.py +5 -0
- ssme_lite-0.1.0/ssme_lite/estimation/contribution.py +76 -0
- ssme_lite-0.1.0/ssme_lite/estimation/estimator.py +245 -0
- ssme_lite-0.1.0/ssme_lite/metrics/__init__.py +4 -0
- ssme_lite-0.1.0/ssme_lite/metrics/scores.py +45 -0
- ssme_lite-0.1.0/ssme_lite/prediction/__init__.py +4 -0
- ssme_lite-0.1.0/ssme_lite/prediction/matrix.py +82 -0
- ssme_lite-0.1.0/ssme_lite/py.typed +0 -0
- ssme_lite-0.1.0/ssme_lite/reporting/__init__.py +4 -0
- ssme_lite-0.1.0/ssme_lite/reporting/assets/report.css +12 -0
- ssme_lite-0.1.0/ssme_lite/reporting/assets/report.html +83 -0
- ssme_lite-0.1.0/ssme_lite/reporting/assets/report.js +298 -0
- ssme_lite-0.1.0/ssme_lite/reporting/data.py +205 -0
- ssme_lite-0.1.0/ssme_lite/reporting/html.py +12 -0
- ssme_lite-0.1.0/ssme_lite/reporting/report.py +138 -0
- ssme_lite-0.1.0/ssme_lite/splits/__init__.py +4 -0
- ssme_lite-0.1.0/ssme_lite/splits/split.py +59 -0
- ssme_lite-0.1.0/ssme_lite.egg-info/PKG-INFO +120 -0
- ssme_lite-0.1.0/ssme_lite.egg-info/SOURCES.txt +58 -0
- ssme_lite-0.1.0/ssme_lite.egg-info/dependency_links.txt +1 -0
- ssme_lite-0.1.0/ssme_lite.egg-info/entry_points.txt +2 -0
- ssme_lite-0.1.0/ssme_lite.egg-info/requires.txt +20 -0
- ssme_lite-0.1.0/ssme_lite.egg-info/top_level.txt +1 -0
- ssme_lite-0.1.0/tests/test_core.py +176 -0
- ssme_lite-0.1.0/tests/test_fixed_epochs_contribution.py +84 -0
- ssme_lite-0.1.0/tests/test_official_audit.py +45 -0
- ssme_lite-0.1.0/tests/test_parallel.py +40 -0
- ssme_lite-0.1.0/tests/test_report_data.py +93 -0
- ssme_lite-0.1.0/tests/test_report_density.py +15 -0
- ssme_lite-0.1.0/tests/test_report_structure.py +33 -0
- ssme_lite-0.1.0/tests/test_splits.py +40 -0
|
@@ -0,0 +1,9 @@
|
|
|
1
|
+
# Changelog
|
|
2
|
+
|
|
3
|
+
## 0.1.0
|
|
4
|
+
|
|
5
|
+
- Public estimator, prediction-matrix, split, and report interfaces.
|
|
6
|
+
- Accuracy, ROC AUC, average precision, and calibration-error estimates.
|
|
7
|
+
- Offline interactive reports and CSV/JSON exports.
|
|
8
|
+
- Optional controlled leave-one-model-out influence diagnostics.
|
|
9
|
+
- English documentation, runnable quickstart, real-data tutorials, and release validation.
|
|
@@ -0,0 +1,14 @@
|
|
|
1
|
+
# Contributing
|
|
2
|
+
|
|
3
|
+
Use Python 3.10 or newer and run commands from the repository root.
|
|
4
|
+
|
|
5
|
+
```bash
|
|
6
|
+
python -m pip install -e '.[dev,docs]'
|
|
7
|
+
python -m pytest -q
|
|
8
|
+
python examples/quickstart.py
|
|
9
|
+
mkdocs build --strict
|
|
10
|
+
```
|
|
11
|
+
|
|
12
|
+
Keep the public imports in `ssme_lite` stable. Document new parameters and update runnable examples when behavior changes. Use small synthetic inputs for tests; keep heavyweight datasets and model downloads out of the default test suite. The optional official-reference audit skips when its external dependencies or reference tree are unavailable.
|
|
13
|
+
|
|
14
|
+
Describe the behavior change and validation in pull requests. Do not commit datasets, model weights, credentials, generated distribution files, or local environments. See `docs/releasing.md` for release verification.
|
ssme_lite-0.1.0/LICENSE
ADDED
|
@@ -0,0 +1,21 @@
|
|
|
1
|
+
MIT License
|
|
2
|
+
|
|
3
|
+
Copyright (c) 2026 jason
|
|
4
|
+
|
|
5
|
+
Permission is hereby granted, free of charge, to any person obtaining a copy
|
|
6
|
+
of this software and associated documentation files (the "Software"), to deal
|
|
7
|
+
in the Software without restriction, including without limitation the rights
|
|
8
|
+
to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
|
|
9
|
+
copies of the Software, and to permit persons to whom the Software is
|
|
10
|
+
furnished to do so, subject to the following conditions:
|
|
11
|
+
|
|
12
|
+
The above copyright notice and this permission notice shall be included in all
|
|
13
|
+
copies or substantial portions of the Software.
|
|
14
|
+
|
|
15
|
+
THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
|
|
16
|
+
IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
|
|
17
|
+
FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
|
|
18
|
+
AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
|
|
19
|
+
LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
|
|
20
|
+
OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
|
|
21
|
+
SOFTWARE.
|
|
@@ -0,0 +1,11 @@
|
|
|
1
|
+
include LICENSE README.md CHANGELOG.md CONTRIBUTING.md mkdocs.yml
|
|
2
|
+
recursive-include docs *.md
|
|
3
|
+
recursive-include examples *.py
|
|
4
|
+
recursive-include tests *.py
|
|
5
|
+
recursive-include ssme_lite *.py py.typed *.html *.css *.js
|
|
6
|
+
prune experiments
|
|
7
|
+
prune data
|
|
8
|
+
prune results
|
|
9
|
+
prune site
|
|
10
|
+
prune .github
|
|
11
|
+
global-exclude __pycache__ *.py[cod] .DS_Store
|
ssme_lite-0.1.0/PKG-INFO
ADDED
|
@@ -0,0 +1,120 @@
|
|
|
1
|
+
Metadata-Version: 2.4
|
|
2
|
+
Name: ssme-lite
|
|
3
|
+
Version: 0.1.0
|
|
4
|
+
Summary: Semi-supervised evaluation and selection of existing classifiers
|
|
5
|
+
Author: Junje Shen
|
|
6
|
+
License-Expression: MIT
|
|
7
|
+
Project-URL: Documentation, https://github.com/JasonShen2002/ssme-lite/blob/main/docs/index.md
|
|
8
|
+
Project-URL: Source, https://github.com/JasonShen2002/ssme-lite
|
|
9
|
+
Project-URL: Issues, https://github.com/JasonShen2002/ssme-lite/issues
|
|
10
|
+
Project-URL: Changelog, https://github.com/JasonShen2002/ssme-lite/blob/main/CHANGELOG.md
|
|
11
|
+
Keywords: semi-supervised,model-evaluation,model-selection,machine-learning
|
|
12
|
+
Classifier: Development Status :: 4 - Beta
|
|
13
|
+
Classifier: Intended Audience :: Science/Research
|
|
14
|
+
Classifier: Operating System :: OS Independent
|
|
15
|
+
Classifier: Programming Language :: Python :: 3
|
|
16
|
+
Classifier: Programming Language :: Python :: 3.10
|
|
17
|
+
Classifier: Programming Language :: Python :: 3.11
|
|
18
|
+
Classifier: Programming Language :: Python :: 3.12
|
|
19
|
+
Classifier: Topic :: Scientific/Engineering :: Artificial Intelligence
|
|
20
|
+
Requires-Python: >=3.10
|
|
21
|
+
Description-Content-Type: text/markdown
|
|
22
|
+
License-File: LICENSE
|
|
23
|
+
Requires-Dist: numpy>=1.24
|
|
24
|
+
Requires-Dist: scipy>=1.10
|
|
25
|
+
Requires-Dist: pandas>=2
|
|
26
|
+
Requires-Dist: scikit-learn>=1.3
|
|
27
|
+
Requires-Dist: matplotlib>=3.7
|
|
28
|
+
Requires-Dist: joblib>=1.3
|
|
29
|
+
Provides-Extra: dev
|
|
30
|
+
Requires-Dist: pytest>=7; extra == "dev"
|
|
31
|
+
Provides-Extra: docs
|
|
32
|
+
Requires-Dist: mkdocs<2,>=1.6; extra == "docs"
|
|
33
|
+
Provides-Extra: official
|
|
34
|
+
Requires-Dist: statsmodels>=0.14; extra == "official"
|
|
35
|
+
Provides-Extra: transfer
|
|
36
|
+
Requires-Dist: torch>=2; extra == "transfer"
|
|
37
|
+
Requires-Dist: torchvision>=0.15; extra == "transfer"
|
|
38
|
+
Requires-Dist: pillow>=9; extra == "transfer"
|
|
39
|
+
Dynamic: license-file
|
|
40
|
+
|
|
41
|
+
# SSME-Lite
|
|
42
|
+
|
|
43
|
+
**Evaluate existing classifiers with a few labels and many unlabeled predictions.**
|
|
44
|
+
|
|
45
|
+
SSME-Lite packages semi-supervised model evaluation into a small Python interface: provide model probabilities and partial labels, estimate performance, rank candidates, and export an interactive HTML report that opens offline.
|
|
46
|
+
|
|
47
|
+
- **One estimator:** `fit()` followed by `report()`.
|
|
48
|
+
- **Four metrics:** Accuracy, ROC AUC, average precision (`auprc`), and ECE.
|
|
49
|
+
- **Flexible inputs:** precomputed probabilities or already fitted sklearn-style models.
|
|
50
|
+
- **Portable outputs:** HTML, CSV, JSON, and metric plots.
|
|
51
|
+
- **Reproducible experiments:** explicit seeds, shared splits, and optional model-influence diagnostics.
|
|
52
|
+
- **Lightweight core:** no PyTorch dependency; transfer learning is optional.
|
|
53
|
+
|
|
54
|
+
## Installation
|
|
55
|
+
|
|
56
|
+
Python 3.10 or newer:
|
|
57
|
+
|
|
58
|
+
```bash
|
|
59
|
+
python -m pip install ssme-lite
|
|
60
|
+
```
|
|
61
|
+
|
|
62
|
+
To work from a checkout:
|
|
63
|
+
|
|
64
|
+
```bash
|
|
65
|
+
git clone https://github.com/JasonShen2002/ssme-lite.git
|
|
66
|
+
cd ssme-lite
|
|
67
|
+
python -m pip install -e '.[dev]'
|
|
68
|
+
```
|
|
69
|
+
|
|
70
|
+
## Minimal interface
|
|
71
|
+
|
|
72
|
+
```python
|
|
73
|
+
from ssme_lite import SSMEEstimator
|
|
74
|
+
|
|
75
|
+
# scores: (samples, models), positive-class probabilities for binary tasks.
|
|
76
|
+
# y_partial: one integer label per row; -1 means unknown.
|
|
77
|
+
estimator = SSMEEstimator(random_state=0)
|
|
78
|
+
estimator.fit(scores, y_partial, model_names=model_names)
|
|
79
|
+
report = estimator.report(primary_metric="accuracy")
|
|
80
|
+
print(report.ranking("accuracy"))
|
|
81
|
+
report.save("results/my_report")
|
|
82
|
+
```
|
|
83
|
+
|
|
84
|
+
Every class must have at least one observed label. For multiclass tasks, use `(samples, models, classes)` probabilities and class indices `0..K-1`.
|
|
85
|
+
|
|
86
|
+
For a complete example without external data, run `python examples/quickstart.py` from the checkout. It generates a small synthetic dataset, fits SSME, and saves `results/quickstart/report.html`. This example illustrates the interface, not an empirical performance claim.
|
|
87
|
+
|
|
88
|
+
## Documentation and tutorials
|
|
89
|
+
|
|
90
|
+
- [Documentation home](https://github.com/JasonShen2002/ssme-lite/blob/main/docs/index.md)
|
|
91
|
+
- [Five-minute quickstart](https://github.com/JasonShen2002/ssme-lite/blob/main/docs/quickstart.md)
|
|
92
|
+
- [Use your own models](https://github.com/JasonShen2002/ssme-lite/blob/main/docs/guides/inputs.md)
|
|
93
|
+
- [Understand the report](https://github.com/JasonShen2002/ssme-lite/blob/main/docs/guides/reports.md)
|
|
94
|
+
- [API reference](https://github.com/JasonShen2002/ssme-lite/blob/main/docs/api.md)
|
|
95
|
+
|
|
96
|
+
| Tutorial | What you will learn |
|
|
97
|
+
| --- | --- |
|
|
98
|
+
| [CivilComments](https://github.com/JasonShen2002/ssme-lite/blob/main/docs/tutorials/civilcomments.md) | Quickstart on the original paper's dataset and a qualitative reproduction check |
|
|
99
|
+
| [PneumoniaMNIST](https://github.com/JasonShen2002/ssme-lite/blob/main/docs/tutorials/pneumoniamnist.md) | Main end-to-end demonstration: prediction, evaluation, ranking, and reporting |
|
|
100
|
+
| [Transfer learning](https://github.com/JasonShen2002/ssme-lite/blob/main/docs/tutorials/transfer.md) | Frozen ResNet18 features, five classifier heads, and a shared evaluation interface |
|
|
101
|
+
|
|
102
|
+
[View saved experiment reports](https://jasonshen2002.github.io/ssme-lite/).
|
|
103
|
+
|
|
104
|
+
## Method and interpretation
|
|
105
|
+
|
|
106
|
+
SSME-Lite follows Shanmugam et al., *Evaluating Multiple Models Using Labeled and Unlabeled Data*, NeurIPS 2025. It combines classifier probabilities through additive log-ratio features, weighted class-conditional kernel densities, and iterative posterior updates with observed labels clamped.
|
|
107
|
+
|
|
108
|
+
Report intervals describe latent-label uncertainty conditional on the fitted posterior. Model contribution measures sensitivity to leaving out an input model. See the [method guide](https://github.com/JasonShen2002/ssme-lite/blob/main/docs/method.md) for definitions, implementation choices, and attribution. The current real-data demonstrations are binary classification tasks; the interface also accepts multiclass probabilities.
|
|
109
|
+
|
|
110
|
+
## Development
|
|
111
|
+
|
|
112
|
+
```bash
|
|
113
|
+
python -m pip install -e '.[dev,docs]'
|
|
114
|
+
python -m pytest -q
|
|
115
|
+
mkdocs serve
|
|
116
|
+
```
|
|
117
|
+
|
|
118
|
+
See [Contributing](https://github.com/JasonShen2002/ssme-lite/blob/main/CONTRIBUTING.md) and the [release guide](https://github.com/JasonShen2002/ssme-lite/blob/main/docs/releasing.md).
|
|
119
|
+
|
|
120
|
+
MIT licensed. Maintained by Junje Shen. Please cite the original SSME paper when using the method.
|
|
@@ -0,0 +1,80 @@
|
|
|
1
|
+
# SSME-Lite
|
|
2
|
+
|
|
3
|
+
**Evaluate existing classifiers with a few labels and many unlabeled predictions.**
|
|
4
|
+
|
|
5
|
+
SSME-Lite packages semi-supervised model evaluation into a small Python interface: provide model probabilities and partial labels, estimate performance, rank candidates, and export an interactive HTML report that opens offline.
|
|
6
|
+
|
|
7
|
+
- **One estimator:** `fit()` followed by `report()`.
|
|
8
|
+
- **Four metrics:** Accuracy, ROC AUC, average precision (`auprc`), and ECE.
|
|
9
|
+
- **Flexible inputs:** precomputed probabilities or already fitted sklearn-style models.
|
|
10
|
+
- **Portable outputs:** HTML, CSV, JSON, and metric plots.
|
|
11
|
+
- **Reproducible experiments:** explicit seeds, shared splits, and optional model-influence diagnostics.
|
|
12
|
+
- **Lightweight core:** no PyTorch dependency; transfer learning is optional.
|
|
13
|
+
|
|
14
|
+
## Installation
|
|
15
|
+
|
|
16
|
+
Python 3.10 or newer:
|
|
17
|
+
|
|
18
|
+
```bash
|
|
19
|
+
python -m pip install ssme-lite
|
|
20
|
+
```
|
|
21
|
+
|
|
22
|
+
To work from a checkout:
|
|
23
|
+
|
|
24
|
+
```bash
|
|
25
|
+
git clone https://github.com/JasonShen2002/ssme-lite.git
|
|
26
|
+
cd ssme-lite
|
|
27
|
+
python -m pip install -e '.[dev]'
|
|
28
|
+
```
|
|
29
|
+
|
|
30
|
+
## Minimal interface
|
|
31
|
+
|
|
32
|
+
```python
|
|
33
|
+
from ssme_lite import SSMEEstimator
|
|
34
|
+
|
|
35
|
+
# scores: (samples, models), positive-class probabilities for binary tasks.
|
|
36
|
+
# y_partial: one integer label per row; -1 means unknown.
|
|
37
|
+
estimator = SSMEEstimator(random_state=0)
|
|
38
|
+
estimator.fit(scores, y_partial, model_names=model_names)
|
|
39
|
+
report = estimator.report(primary_metric="accuracy")
|
|
40
|
+
print(report.ranking("accuracy"))
|
|
41
|
+
report.save("results/my_report")
|
|
42
|
+
```
|
|
43
|
+
|
|
44
|
+
Every class must have at least one observed label. For multiclass tasks, use `(samples, models, classes)` probabilities and class indices `0..K-1`.
|
|
45
|
+
|
|
46
|
+
For a complete example without external data, run `python examples/quickstart.py` from the checkout. It generates a small synthetic dataset, fits SSME, and saves `results/quickstart/report.html`. This example illustrates the interface, not an empirical performance claim.
|
|
47
|
+
|
|
48
|
+
## Documentation and tutorials
|
|
49
|
+
|
|
50
|
+
- [Documentation home](https://github.com/JasonShen2002/ssme-lite/blob/main/docs/index.md)
|
|
51
|
+
- [Five-minute quickstart](https://github.com/JasonShen2002/ssme-lite/blob/main/docs/quickstart.md)
|
|
52
|
+
- [Use your own models](https://github.com/JasonShen2002/ssme-lite/blob/main/docs/guides/inputs.md)
|
|
53
|
+
- [Understand the report](https://github.com/JasonShen2002/ssme-lite/blob/main/docs/guides/reports.md)
|
|
54
|
+
- [API reference](https://github.com/JasonShen2002/ssme-lite/blob/main/docs/api.md)
|
|
55
|
+
|
|
56
|
+
| Tutorial | What you will learn |
|
|
57
|
+
| --- | --- |
|
|
58
|
+
| [CivilComments](https://github.com/JasonShen2002/ssme-lite/blob/main/docs/tutorials/civilcomments.md) | Quickstart on the original paper's dataset and a qualitative reproduction check |
|
|
59
|
+
| [PneumoniaMNIST](https://github.com/JasonShen2002/ssme-lite/blob/main/docs/tutorials/pneumoniamnist.md) | Main end-to-end demonstration: prediction, evaluation, ranking, and reporting |
|
|
60
|
+
| [Transfer learning](https://github.com/JasonShen2002/ssme-lite/blob/main/docs/tutorials/transfer.md) | Frozen ResNet18 features, five classifier heads, and a shared evaluation interface |
|
|
61
|
+
|
|
62
|
+
[View saved experiment reports](https://jasonshen2002.github.io/ssme-lite/).
|
|
63
|
+
|
|
64
|
+
## Method and interpretation
|
|
65
|
+
|
|
66
|
+
SSME-Lite follows Shanmugam et al., *Evaluating Multiple Models Using Labeled and Unlabeled Data*, NeurIPS 2025. It combines classifier probabilities through additive log-ratio features, weighted class-conditional kernel densities, and iterative posterior updates with observed labels clamped.
|
|
67
|
+
|
|
68
|
+
Report intervals describe latent-label uncertainty conditional on the fitted posterior. Model contribution measures sensitivity to leaving out an input model. See the [method guide](https://github.com/JasonShen2002/ssme-lite/blob/main/docs/method.md) for definitions, implementation choices, and attribution. The current real-data demonstrations are binary classification tasks; the interface also accepts multiclass probabilities.
|
|
69
|
+
|
|
70
|
+
## Development
|
|
71
|
+
|
|
72
|
+
```bash
|
|
73
|
+
python -m pip install -e '.[dev,docs]'
|
|
74
|
+
python -m pytest -q
|
|
75
|
+
mkdocs serve
|
|
76
|
+
```
|
|
77
|
+
|
|
78
|
+
See [Contributing](https://github.com/JasonShen2002/ssme-lite/blob/main/CONTRIBUTING.md) and the [release guide](https://github.com/JasonShen2002/ssme-lite/blob/main/docs/releasing.md).
|
|
79
|
+
|
|
80
|
+
MIT licensed. Maintained by Junje Shen. Please cite the original SSME paper when using the method.
|
|
@@ -0,0 +1,117 @@
|
|
|
1
|
+
# API reference
|
|
2
|
+
|
|
3
|
+
The stable public imports are:
|
|
4
|
+
|
|
5
|
+
```python
|
|
6
|
+
from ssme_lite import (
|
|
7
|
+
SSMEEstimator, PredictionMatrixGenerator,
|
|
8
|
+
SemiSupervisedSplit, EvaluationReport,
|
|
9
|
+
)
|
|
10
|
+
```
|
|
11
|
+
|
|
12
|
+
## SSMEEstimator
|
|
13
|
+
|
|
14
|
+
```python
|
|
15
|
+
SSMEEstimator(density="kde", bandwidth="scott", max_iter=100,
|
|
16
|
+
tol=1e-3, labeled_weight=10.0, clip=1e-6,
|
|
17
|
+
random_state=0, n_jobs=1, batch_size=2048,
|
|
18
|
+
early_stopping=False)
|
|
19
|
+
```
|
|
20
|
+
|
|
21
|
+
| Parameter | Meaning |
|
|
22
|
+
| --- | --- |
|
|
23
|
+
| `density` | Currently only `"kde"` is supported |
|
|
24
|
+
| `bandwidth` | `"scott"`, `"official"`, or a positive scalar in ALR space |
|
|
25
|
+
| `max_iter` | Positive integer iteration count or early-stopping limit |
|
|
26
|
+
| `tol` | Positive threshold for maximum posterior change |
|
|
27
|
+
| `labeled_weight` | Positive weight on observed samples; unlabeled weight is 1 |
|
|
28
|
+
| `clip` | Probability clipping threshold, strictly between 0 and 0.5 |
|
|
29
|
+
| `random_state` | Seed for initialization and default report sampling |
|
|
30
|
+
| `n_jobs` | joblib parallelism for class-density fitting; 1 is serial |
|
|
31
|
+
| `batch_size` | Positive batch size for posterior evaluation |
|
|
32
|
+
| `early_stopping` | Whether to stop when posterior change falls below `tol` |
|
|
33
|
+
|
|
34
|
+
### fit
|
|
35
|
+
|
|
36
|
+
```python
|
|
37
|
+
fit(scores, y, scores_unlabeled=None, model_names=None,
|
|
38
|
+
initial_responsibilities=None)
|
|
39
|
+
```
|
|
40
|
+
|
|
41
|
+
Returns the fitted estimator. `scores` is binary `(N, M)` or full `(N, M, K)`. `y` has shape `(N,)`, contains class indices or `-1`, and includes every class among observed entries. If supplied, `scores_unlabeled` is appended after the first array. `model_names` contains unique names for the model axis; default names are `model_0`, etc.
|
|
42
|
+
|
|
43
|
+
Advanced: `initial_responsibilities` is a nonnegative normalized `(N_total, K)` posterior used for controlled initialization. Observed labels are still clamped. Ordinary use should leave it unset.
|
|
44
|
+
|
|
45
|
+
Invalid shapes, probabilities, labels, unsupported density choices, or duplicate names raise `ValueError`. The `official` bandwidth raises `ImportError` without statsmodels.
|
|
46
|
+
|
|
47
|
+
### Fitted attributes
|
|
48
|
+
|
|
49
|
+
| Attribute | Meaning |
|
|
50
|
+
| --- | --- |
|
|
51
|
+
| `posterior_` | `(N_total, K)` latent-class probabilities, with observed rows clamped |
|
|
52
|
+
| `scores_`, `y_` | Normalized full probability tensor and aligned partial labels |
|
|
53
|
+
| `model_names_`, `classes_` | Candidate names and class indices |
|
|
54
|
+
| `labeled_indices_`, `unlabeled_indices_` | Indices in the combined fitted pool |
|
|
55
|
+
| `bandwidth_`, `priors_` | Effective scalar bandwidth and fitted class priors |
|
|
56
|
+
| `history_` | Per-iteration posterior-change, objective, and class summaries |
|
|
57
|
+
| `n_iter_`, `converged_`, `stop_reason_` | Actual iterations, tolerance status, and stopping reason |
|
|
58
|
+
|
|
59
|
+
### predict_proba
|
|
60
|
+
|
|
61
|
+
```python
|
|
62
|
+
predict_proba(scores)
|
|
63
|
+
```
|
|
64
|
+
|
|
65
|
+
Returns `(N_new, K)` posterior probabilities from the fitted densities. Input model and class axes must match the fit. This returns inferred labels for samples, not per-model performance. It does not clamp labels for new rows. Use `posterior_` when you need the fitted pool with observed-label clamping.
|
|
66
|
+
|
|
67
|
+
### evaluate / report
|
|
68
|
+
|
|
69
|
+
```python
|
|
70
|
+
evaluate(metrics=("accuracy", "auc", "auprc", "ece"), n_draws=100,
|
|
71
|
+
confidence=0.95, target="all", random_state=None,
|
|
72
|
+
title=None, sample_ids=None, primary_metric=None,
|
|
73
|
+
contribution=False)
|
|
74
|
+
report(**kwargs)
|
|
75
|
+
```
|
|
76
|
+
|
|
77
|
+
Returns `EvaluationReport`; `report` is an alias for `evaluate`. Metrics must be distinct supported keys. `n_draws >= 2`, `0 < confidence < 1`, and target is `"all"` or `"unlabeled"`. The default report seed inherits the estimator seed. `sample_ids`, if provided, identify all fitted rows. `primary_metric` must be among the requested metrics and defaults to the first. Contribution requires at least two models and some unlabeled samples.
|
|
78
|
+
|
|
79
|
+
### contribution
|
|
80
|
+
|
|
81
|
+
`contribution(n_draws=30)` returns a pandas DataFrame of leave-one-model-out diagnostics. `n_draws` is retained for compatibility; this diagnostic uses exact expected Accuracy and does not sample labels. See [interpretation](guides/reports.md).
|
|
82
|
+
|
|
83
|
+
## PredictionMatrixGenerator
|
|
84
|
+
|
|
85
|
+
```python
|
|
86
|
+
PredictionMatrixGenerator(models, clip=1e-6, n_jobs=1)
|
|
87
|
+
```
|
|
88
|
+
|
|
89
|
+
`models` is a mapping of names to already fitted classifiers, or a sequence assigned automatic names. Each classifier must expose `classes_` and `predict_proba(X)`.
|
|
90
|
+
|
|
91
|
+
- `fit(X=None, y=None)` inspects models and aligns class columns; returns self.
|
|
92
|
+
- `transform(X, y_partial=None)` collects probabilities; optional labels record labeled/unlabeled indices.
|
|
93
|
+
- `fit_transform(X, y=None)` combines the two steps without training classifiers.
|
|
94
|
+
|
|
95
|
+
Attributes include `model_names_`, `classes_`, and, after transformation, `sample_ids_`. Binary output is `(N, M)` for `classes_[1]`; multiclass output is `(N, M, K)`.
|
|
96
|
+
|
|
97
|
+
## SemiSupervisedSplit
|
|
98
|
+
|
|
99
|
+
```python
|
|
100
|
+
SemiSupervisedSplit(n_labeled, n_unlabeled, random_state=0)
|
|
101
|
+
```
|
|
102
|
+
|
|
103
|
+
Both counts must be positive integers. `indices(n_samples)` returns `labeled`, `unlabeled`, `estimation`, and `heldout` index arrays, requiring a nonempty held-out set. `partial_labels(y, parts)` returns the partial-label array aligned to `estimation`. `labeled_class_count(y, parts)` counts represented classes. Labels do not affect the permutation.
|
|
104
|
+
|
|
105
|
+
## EvaluationReport
|
|
106
|
+
|
|
107
|
+
Usually obtain this object from `estimator.report()`.
|
|
108
|
+
|
|
109
|
+
- `table`: pandas DataFrame of estimates and uncertainty summaries.
|
|
110
|
+
- `metadata`: settings and interpretation notes.
|
|
111
|
+
- `data`: structured interactive-report payload.
|
|
112
|
+
- `ranking(metric="accuracy")`: sorted DataFrame with a `rank` column; ECE ascending, other metrics descending.
|
|
113
|
+
- `show()`: print estimates and fit status; return self.
|
|
114
|
+
- `plot(metric="accuracy", path=None)`: return a Matplotlib figure and optionally save it.
|
|
115
|
+
- `save(directory)`: create outputs and return the HTML `pathlib.Path`.
|
|
116
|
+
|
|
117
|
+
See the [report guide](guides/reports.md) for table fields and saved-file definitions.
|
|
@@ -0,0 +1,41 @@
|
|
|
1
|
+
# Command-line evaluation
|
|
2
|
+
|
|
3
|
+
The CLI accepts a JSON configuration and can evaluate saved probabilities without model objects.
|
|
4
|
+
|
|
5
|
+
## 1. Save inputs
|
|
6
|
+
|
|
7
|
+
```python
|
|
8
|
+
import numpy as np
|
|
9
|
+
|
|
10
|
+
np.savez("predictions.npz", scores=scores, y=y_partial,
|
|
11
|
+
model_names=np.asarray(model_names, dtype=str))
|
|
12
|
+
```
|
|
13
|
+
|
|
14
|
+
Use the [input contract](inputs.md). Store names as a Unicode array, not a pickled object array. Unknown labels are `-1`.
|
|
15
|
+
|
|
16
|
+
## 2. Write `evaluation.json`
|
|
17
|
+
|
|
18
|
+
```json
|
|
19
|
+
{
|
|
20
|
+
"source": "npz",
|
|
21
|
+
"mode": "evaluate",
|
|
22
|
+
"data": "predictions.npz",
|
|
23
|
+
"output": "results/evaluation",
|
|
24
|
+
"estimator": {"max_iter": 100, "random_state": 0},
|
|
25
|
+
"report": {"n_draws": 100, "primary_metric": "accuracy"}
|
|
26
|
+
}
|
|
27
|
+
```
|
|
28
|
+
|
|
29
|
+
## 3. Run
|
|
30
|
+
|
|
31
|
+
```bash
|
|
32
|
+
ssme-lite evaluation.json
|
|
33
|
+
```
|
|
34
|
+
|
|
35
|
+
Paths resolve relative to the configuration file, not the shell's current directory. The command prints the output directory and saves the report plus `input_config.json`.
|
|
36
|
+
|
|
37
|
+
For a downloadable example, run `python examples/quickstart.py` from the checkout: it also writes `results/quickstart/predictions.npz` and `evaluation.json`. Then run `ssme-lite results/quickstart/evaluation.json`.
|
|
38
|
+
|
|
39
|
+
## Benchmark configurations
|
|
40
|
+
|
|
41
|
+
The repository's `configs/` also includes synthetic and released-score benchmark configurations. Those run repeated experimental comparisons and have different data requirements and costs. The small NPZ example above is the recommended first CLI run. See `ssme-lite --help` for the command syntax.
|
|
@@ -0,0 +1,67 @@
|
|
|
1
|
+
# Use your own predictions
|
|
2
|
+
|
|
3
|
+
## Input contract
|
|
4
|
+
|
|
5
|
+
| Input | Shape / values | Meaning |
|
|
6
|
+
| --- | --- | --- |
|
|
7
|
+
| Binary scores | `(N, M)`, probabilities in `[0, 1]` | Probability of class index 1 |
|
|
8
|
+
| Full probabilities | `(N, M, K)` | Class probabilities summing to one on the last axis |
|
|
9
|
+
| Partial labels | `(N,)`, `-1` or integers `0..K-1` | Unknown labels or observed class indices |
|
|
10
|
+
| Model names | `M` unique strings | Names shown in tables and reports |
|
|
11
|
+
|
|
12
|
+
All models must predict the same samples in the same order and the same class set. Every class needs at least one observed label. Pass probabilities, not logits. Inputs are copied, clipped away from zero and one, and normalized internally.
|
|
13
|
+
|
|
14
|
+
## Runnable sklearn example
|
|
15
|
+
|
|
16
|
+
From the checkout root, run:
|
|
17
|
+
|
|
18
|
+
```bash
|
|
19
|
+
python examples/sklearn_models.py
|
|
20
|
+
```
|
|
21
|
+
|
|
22
|
+
This self-contained script generates data, trains logistic regression and a random forest on 300 training samples, and collects their predictions on 300 different samples. It then fits SSME with 20 observed labels and 180 unlabeled samples, reserving 100 held-out rows. The output is `results/sklearn_models/report.html`, including model-influence diagnostics. No external data or model weights are required.
|
|
23
|
+
|
|
24
|
+
## Already fitted classifiers
|
|
25
|
+
|
|
26
|
+
```python
|
|
27
|
+
from ssme_lite import PredictionMatrixGenerator, SSMEEstimator
|
|
28
|
+
|
|
29
|
+
# Each value is already trained, and exposes classes_ and predict_proba(X).
|
|
30
|
+
generator = PredictionMatrixGenerator({"logistic": fitted_logistic, "forest": fitted_forest})
|
|
31
|
+
scores = generator.fit_transform(X_evaluation)
|
|
32
|
+
estimator = SSMEEstimator(random_state=0).fit(
|
|
33
|
+
scores, y_partial, model_names=generator.model_names_
|
|
34
|
+
)
|
|
35
|
+
report = estimator.report()
|
|
36
|
+
```
|
|
37
|
+
|
|
38
|
+
The generator aligns probability columns to the first model's `classes_`. Its `fit()` inspects models; it does not train them. Binary output uses `generator.classes_[1]` as the positive class.
|
|
39
|
+
|
|
40
|
+
If original labels are strings or nonconsecutive numbers, encode observed labels as positions in `generator.classes_`. Reserve `-1` for unknown entries. For example, classes `['healthy', 'pneumonia']` map to `[0, 1]`.
|
|
41
|
+
|
|
42
|
+
## Separate labeled and unlabeled arrays
|
|
43
|
+
|
|
44
|
+
```python
|
|
45
|
+
estimator.fit(
|
|
46
|
+
labeled_scores, labeled_y,
|
|
47
|
+
scores_unlabeled=unlabeled_scores,
|
|
48
|
+
model_names=model_names,
|
|
49
|
+
)
|
|
50
|
+
```
|
|
51
|
+
|
|
52
|
+
This appends the unlabeled rows internally. Model and class dimensions must match. Reports default to all fitted rows; use `report(target="unlabeled")` to evaluate only the unlabeled portion.
|
|
53
|
+
|
|
54
|
+
## Controlled benchmark splits
|
|
55
|
+
|
|
56
|
+
When full labels are available for a benchmark, use `SemiSupervisedSplit` to hide them reproducibly:
|
|
57
|
+
|
|
58
|
+
```python
|
|
59
|
+
from ssme_lite import SemiSupervisedSplit
|
|
60
|
+
|
|
61
|
+
splitter = SemiSupervisedSplit(20, 400, random_state=0)
|
|
62
|
+
parts = splitter.indices(len(y))
|
|
63
|
+
y_partial = splitter.partial_labels(y, parts)
|
|
64
|
+
estimator.fit(scores[parts["estimation"]], y_partial, model_names=model_names)
|
|
65
|
+
```
|
|
66
|
+
|
|
67
|
+
The remaining samples form `parts['heldout']`. They are not used by the estimator. The splitter uses one uniform permutation and does not inspect labels to repair a missing-class draw. Check `splitter.labeled_class_count(y, parts)` when auditing benchmark splits, and document exclusions or additional annotation. Keep classifier training separate from this evaluation pool.
|
|
@@ -0,0 +1,49 @@
|
|
|
1
|
+
# Understand the report
|
|
2
|
+
|
|
3
|
+
```python
|
|
4
|
+
report = estimator.report(n_draws=100, primary_metric="accuracy")
|
|
5
|
+
report.ranking("accuracy")
|
|
6
|
+
report.table
|
|
7
|
+
report.save("results/evaluation")
|
|
8
|
+
```
|
|
9
|
+
|
|
10
|
+
## Metrics and rankings
|
|
11
|
+
|
|
12
|
+
| Key | Quantity | Better direction |
|
|
13
|
+
| --- | --- | --- |
|
|
14
|
+
| `accuracy` | Fraction of correct class predictions | Higher |
|
|
15
|
+
| `auc` | ROC area; macro one-vs-rest for multiclass | Higher |
|
|
16
|
+
| `auprc` | Average precision; macro average for multiclass | Higher |
|
|
17
|
+
| `ece` | Expected calibration error, 10 equal-width bins | Lower |
|
|
18
|
+
|
|
19
|
+
`auprc` is computed with sklearn's average precision, not trapezoidal integration of a precision–recall curve. Binary ECE uses positive-class probabilities; multiclass ECE uses top-label confidence. Ranking uses the point estimate and assigns the same minimum rank to ties. Choose the metric appropriate to your task before comparing candidates.
|
|
20
|
+
|
|
21
|
+
## Table columns
|
|
22
|
+
|
|
23
|
+
`model`, `metric`, and `estimate` identify each estimate. `interval_low` and `interval_high` are conditional latent-label quantiles at the requested confidence level. `label_std` is the standard deviation across valid label draws; `mc_standard_error` quantifies numerical Monte Carlo error. `valid_draws` records how many finite metric values were available.
|
|
24
|
+
|
|
25
|
+
Accuracy uses its exact conditional expectation, so its point estimate has zero Monte Carlo error. Other metric estimates average valid draws. Accuracy intervals are still obtained from sampled labels. AUC/AP draws missing a class are undefined and omitted; inspect `valid_draws` for small or imbalanced targets.
|
|
26
|
+
|
|
27
|
+
Intervals condition on a fitted posterior. They do not include density-fit, bandwidth, or population sampling uncertainty. Increasing `n_draws` improves numerical precision; it does not account for those additional sources.
|
|
28
|
+
|
|
29
|
+
## Files you can share
|
|
30
|
+
|
|
31
|
+
| File | Purpose |
|
|
32
|
+
| --- | --- |
|
|
33
|
+
| `report.html` | Interactive report, with embedded assets and data; opens offline |
|
|
34
|
+
| `metrics.csv` | Metric table for spreadsheets or further analysis |
|
|
35
|
+
| `metadata.json` | Fit settings, target, seed, and interpretation notes |
|
|
36
|
+
| `report_data.json` | Structured data powering the report |
|
|
37
|
+
| Metric PNGs | Static plots for presentations |
|
|
38
|
+
| `contribution.csv` | Produced when model contribution is enabled |
|
|
39
|
+
|
|
40
|
+
The report summarizes prediction and posterior data. Inspect its contents before sharing outputs from sensitive datasets.
|
|
41
|
+
|
|
42
|
+
## Model influence
|
|
43
|
+
|
|
44
|
+
```python
|
|
45
|
+
report = estimator.report(contribution=True)
|
|
46
|
+
report.save("results/with_contribution")
|
|
47
|
+
```
|
|
48
|
+
|
|
49
|
+
This refits once per removed model, holding visible labels, scalar bandwidth, iteration count, and full-fit initialization fixed. Total variation measures the change in the unlabeled posterior; the report also shows changes to the remaining models' Accuracy estimates. It describes conditional influence, not a causal effect or a guaranteed benefit from dropping a model. Runtime increases by roughly one fit per candidate.
|
|
@@ -0,0 +1,32 @@
|
|
|
1
|
+
# SSME-Lite
|
|
2
|
+
|
|
3
|
+
## From classifier probabilities to an evaluation report
|
|
4
|
+
|
|
5
|
+
SSME-Lite helps you evaluate several already trained classifiers when labels are scarce. It combines a small labeled sample with the models' predictions on unlabeled samples, then produces metric estimates, rankings, and an offline report.
|
|
6
|
+
|
|
7
|
+
**The workflow:** collect probabilities → reveal a few labels → fit `SSMEEstimator` → inspect and export results.
|
|
8
|
+
|
|
9
|
+
## Start here
|
|
10
|
+
|
|
11
|
+
1. [Install the package](installation.md).
|
|
12
|
+
2. [Run the five-minute quickstart](quickstart.md), with no downloads or model weights.
|
|
13
|
+
3. [Connect your own classifiers](guides/inputs.md).
|
|
14
|
+
4. [Read the report](guides/reports.md) and choose your target metric.
|
|
15
|
+
|
|
16
|
+
## Learn through real experiments
|
|
17
|
+
|
|
18
|
+
| Tutorial | Role | Inputs |
|
|
19
|
+
| --- | --- | --- |
|
|
20
|
+
| [CivilComments](tutorials/civilcomments.md) | Original-paper quickstart and reproduction check | Released classifier scores |
|
|
21
|
+
| [PneumoniaMNIST](tutorials/pneumoniamnist.md) | Main end-to-end demonstration | Images and seven released models, or saved probabilities |
|
|
22
|
+
| [Transfer learning](tutorials/transfer.md) | Adapt the workflow to newly trained heads | Frozen ImageNet backbone and five sklearn heads |
|
|
23
|
+
|
|
24
|
+
The core package consumes probabilities. Training models and downloading datasets belong to the tutorials, so you can use the estimator without a deep-learning framework.
|
|
25
|
+
|
|
26
|
+
## Explore further
|
|
27
|
+
|
|
28
|
+
- [Command-line evaluation](guides/cli.md): run from an NPZ file and JSON configuration.
|
|
29
|
+
- [Method and interpretation](method.md): algorithm, metrics, intervals, and model influence.
|
|
30
|
+
- [API reference](api.md): public classes, defaults, inputs, and outputs.
|
|
31
|
+
- [Troubleshooting](troubleshooting.md): common setup and data-alignment problems.
|
|
32
|
+
- [Saved interactive reports](https://jasonshen2002.github.io/ssme-lite/): inspect results before running experiments.
|
|
@@ -0,0 +1,52 @@
|
|
|
1
|
+
# Installation
|
|
2
|
+
|
|
3
|
+
## Core package
|
|
4
|
+
|
|
5
|
+
Use Python 3.10 or newer, preferably in a dedicated virtual environment:
|
|
6
|
+
|
|
7
|
+
```bash
|
|
8
|
+
python -m venv .venv
|
|
9
|
+
# macOS / Linux
|
|
10
|
+
source .venv/bin/activate
|
|
11
|
+
# Windows PowerShell: .venv\Scripts\Activate.ps1
|
|
12
|
+
python -m pip install --upgrade pip
|
|
13
|
+
python -m pip install ssme-lite
|
|
14
|
+
python -c "from ssme_lite import SSMEEstimator; print(SSMEEstimator())"
|
|
15
|
+
```
|
|
16
|
+
|
|
17
|
+
The core uses NumPy, SciPy, pandas, scikit-learn, Matplotlib, and joblib. It does not require PyTorch, model weights, or network access at evaluation time.
|
|
18
|
+
|
|
19
|
+
## Optional features
|
|
20
|
+
|
|
21
|
+
```bash
|
|
22
|
+
# Reference-code bandwidth rule
|
|
23
|
+
python -m pip install 'ssme-lite[official]'
|
|
24
|
+
# Frozen-backbone transfer-learning dependencies
|
|
25
|
+
python -m pip install 'ssme-lite[transfer]'
|
|
26
|
+
```
|
|
27
|
+
|
|
28
|
+
`official` adds statsmodels; it does not download the original research repository. `transfer` adds PyTorch, torchvision, and Pillow; experiment scripts remain in the source repository.
|
|
29
|
+
|
|
30
|
+
## Tutorials and development
|
|
31
|
+
|
|
32
|
+
```bash
|
|
33
|
+
git clone https://github.com/JasonShen2002/ssme-lite.git
|
|
34
|
+
cd ssme-lite
|
|
35
|
+
python -m pip install -e '.[dev,docs]'
|
|
36
|
+
python examples/quickstart.py
|
|
37
|
+
mkdocs serve
|
|
38
|
+
```
|
|
39
|
+
|
|
40
|
+
Run repository commands from the checkout root unless a tutorial says otherwise. Notebooks additionally require Jupyter: `python -m pip install jupyterlab`. The full PneumoniaMNIST inference path also requires `ai-edge-litert`; availability depends on your Python version and platform. Its saved-probability path uses only the core package.
|
|
41
|
+
|
|
42
|
+
## Conda environments
|
|
43
|
+
|
|
44
|
+
You can install the PyPI distribution inside an activated conda environment:
|
|
45
|
+
|
|
46
|
+
```bash
|
|
47
|
+
conda create -n ssme python=3.11 pip
|
|
48
|
+
conda activate ssme
|
|
49
|
+
python -m pip install ssme-lite
|
|
50
|
+
```
|
|
51
|
+
|
|
52
|
+
A draft recipe is included in `conda/`; this does not mean the package has been accepted into conda-forge.
|
|
@@ -0,0 +1,35 @@
|
|
|
1
|
+
# Method and interpretation
|
|
2
|
+
|
|
3
|
+
## Algorithm
|
|
4
|
+
|
|
5
|
+
SSME-Lite implements a semi-supervised evaluation workflow based on Shanmugam et al. (2025). Candidate classifiers are already trained. Their probability vectors form the inputs to the evaluator.
|
|
6
|
+
|
|
7
|
+
1. Clip and normalize probabilities, then apply an additive log-ratio transformation with the final class as reference.
|
|
8
|
+
2. Initialize latent classes by sampling from the mean candidate probabilities; clamp observed labels.
|
|
9
|
+
3. Fit weighted class-conditional Gaussian kernel densities and update class priors.
|
|
10
|
+
4. Update class posteriors and clamp observed labels again.
|
|
11
|
+
5. Repeat for the requested number of iterations, then evaluate candidates under the fitted posterior.
|
|
12
|
+
|
|
13
|
+
Observed samples have default weight 10; unlabeled samples have weight 1. The default is 100 fixed iterations (`early_stopping=False`). With early stopping enabled, the largest posterior change is compared with `tol=1e-3`. `history_`, `n_iter_`, `converged_`, and `stop_reason_` expose fit diagnostics. Fixed-bandwidth KDE updates are an EM-style procedure; the recorded objective is a diagnostic rather than a promised monotonic optimization trace.
|
|
14
|
+
|
|
15
|
+
## Bandwidth choices
|
|
16
|
+
|
|
17
|
+
- `"scott"` (default): mean feature standard deviation multiplied by the multivariate Scott sample-size factor, with a minimum scalar bandwidth of `1e-3`.
|
|
18
|
+
- `"official"`: the smallest per-feature statsmodels normal-reference bandwidth divided by four, following the public reference implementation; install the `official` extra.
|
|
19
|
+
- A positive number: use that fixed scalar bandwidth in ALR space.
|
|
20
|
+
|
|
21
|
+
`official` identifies a reference-code bandwidth rule, not bitwise equivalence of every experiment. For example, the package uses equal-width ECE bins, whereas the paper reports equal-frequency bins.
|
|
22
|
+
|
|
23
|
+
## What is evaluated?
|
|
24
|
+
|
|
25
|
+
By default, reports estimate performance over the fitted estimation pool, including observed and unobserved rows. `target="unlabeled"` restricts the target to unlabeled rows. Held-out benchmark metrics are computed separately and must be labeled as such. The real-data tutorials demonstrate binary tasks; multiclass input support alone is not evidence of benchmark performance on multiclass datasets.
|
|
26
|
+
|
|
27
|
+
The model-influence diagnostic is a controlled leave-one-model-out sensitivity analysis. See [report interpretation](guides/reports.md) for interval and influence definitions.
|
|
28
|
+
|
|
29
|
+
## Attribution
|
|
30
|
+
|
|
31
|
+
SSME-Lite is an independent packaging and workflow implementation, not the authors' official distribution.
|
|
32
|
+
|
|
33
|
+
Divya Shanmugam, Shuvom Sadhuka, Manish Raghavan, John Guttag, Bonnie Berger, and Emma Pierson. **Evaluating Multiple Models Using Labeled and Unlabeled Data.** Advances in Neural Information Processing Systems 38, 2025, pp. 30612–30648. [Paper](https://proceedings.neurips.cc/paper_files/paper/2025/hash/2c0e78cc177dfeca260ef990a8c99209-Abstract-Conference.html) · [Official code](https://github.com/divyashan/SSME).
|
|
34
|
+
|
|
35
|
+
The software is MIT licensed. Dataset and checkpoint terms remain those of their original providers; these assets are not bundled in the Python distribution.
|