uxplain 0.3.0__tar.gz

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -0,0 +1,74 @@
1
+ # Changelog
2
+
3
+ All notable changes to **uxplain** are documented here.
4
+ The format follows [Keep a Changelog](https://keepachangelog.com/),
5
+ and the project adheres to [Semantic Versioning](https://semver.org/).
6
+
7
+ ## [0.3.0] - Unreleased
8
+
9
+ ### Added
10
+ - PNAD Continua and GEIH dataset loaders, without bundling survey microdata.
11
+ - Exact TreeSHAP for supported affine summaries of built-in conformal predictors,
12
+ with generic SHAP as the fallback.
13
+ - Python 3.13 CI coverage and Linux/Windows distribution checks that install the
14
+ wheel and run the tests included in the source archive outside the checkout.
15
+ - Runnable getting-started examples and release-validation instructions.
16
+
17
+ ### Fixed
18
+ - Use label-independent automatic calibration splits for classification, matching
19
+ the usual split-conformal exchangeability argument; no class-based retries.
20
+ - Reject smoothed classification targets in explainers, including with fixed
21
+ seeds, while retaining smoothed prediction as an opt-in.
22
+ - Validate conformal methods, task names, coverage levels and calibration arrays;
23
+ accept single-column targets without broadcasting CQR calibration scores.
24
+ - Return unbounded regression intervals for Mondrian groups absent from calibration.
25
+ - Keep Mondrian assignments deterministic when scores tie and seed permutation
26
+ SHAP without changing the caller's NumPy random state.
27
+ - Restrict the TreeSHAP shortcut to supported algebraic decompositions instead of
28
+ inferring global affinity from agreement on a finite sample.
29
+ - Prevent target leakage for alternative dataset targets and reject ambiguous
30
+ joins of GEIH modules; preserve missing occupation and informality values.
31
+ - Preserve DataFrame feature order and reset pipeline state on refitting.
32
+ - Reject non-finite explanation targets explicitly, while allowing unbounded
33
+ conformal prediction intervals.
34
+ - Plot single-observation SHAP slices and PDP interactions with constant features;
35
+ require an explicit baseline for SHAP waterfall plots.
36
+
37
+ ### Changed
38
+ - Document coverage assumptions, signed CQR spans and pending real-microdata
39
+ revalidation of dataset loaders; label width plots as signed spans.
40
+ - Use SPDX license metadata and include tests in the source distribution.
41
+ - The manuscript's replication results must be regenerated with this version
42
+ before submission; this entry describes a local release candidate.
43
+
44
+ ## [0.2.3]
45
+
46
+ ### Added
47
+ - **PDP-based global feature importance.** `PDPExplanation` now carries an
48
+ `importance` field — the standard deviation of each feature's averaged
49
+ partial-dependence curve (Greenwell et al., 2018). Render it with the new
50
+ `plot_kind="importance"` for the PDP backend.
51
+ - Runnable demo notebooks (`notebooks/01_regression.ipynb`,
52
+ `notebooks/02_classification.ipynb`) covering conformal outputs, all SHAP
53
+ plot kinds via `generate_shap_plots`, PDP / ICE / PDP+ICE / 2D PDP /
54
+ importance, CQR with quantile regressors, and the classification metrics.
55
+
56
+ ### Fixed
57
+ - README: the PDP importance plot kind is documented as `"importance"`
58
+ (was inconsistently shown as `"bar"`, which collided with the SHAP bar plot),
59
+ and the PDP default plot kinds now reflect the actual default (`["pdp"]`).
60
+
61
+ ## [0.2.0]
62
+
63
+ ### Added
64
+ - Classification support: prediction sets, conformal classifiers
65
+ (`CrepesConformalClassifier`), and classification uncertainty metrics
66
+ (`set_size`, `credibility`, `confidence`).
67
+ - Package renamed to `uxplain` and prepared for PyPI.
68
+
69
+ ## [0.1.0]
70
+
71
+ ### Added
72
+ - Initial release: regression conformal predictors (`CrepesConformalPredictor`,
73
+ `CQRConformalPredictor`) with SHAP, PDP, and LIME explainability of
74
+ interval-derived uncertainty metrics.
uxplain-0.3.0/LICENSE ADDED
@@ -0,0 +1,20 @@
1
+ The MIT License (MIT)
2
+ Copyright (c) 2025-2026, Tomas Rodriguez Taborda, Veronica Seguro Varela, Rafael Izbicki, Johnatan Cardona Jimenez
3
+
4
+ Permission is hereby granted, free of charge, to any person obtaining a copy
5
+ of this software and associated documentation files (the "Software"), to deal
6
+ in the Software without restriction, including without limitation the rights
7
+ to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
8
+ copies of the Software, and to permit persons to whom the Software is
9
+ furnished to do so, subject to the following conditions:
10
+
11
+ The above copyright notice and this permission notice shall be included in all
12
+ copies or substantial portions of the Software.
13
+
14
+ THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
15
+ IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
16
+ FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
17
+ AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
18
+ LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
19
+ OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
20
+ SOFTWARE.
uxplain-0.3.0/PKG-INFO ADDED
@@ -0,0 +1,134 @@
1
+ Metadata-Version: 2.4
2
+ Name: uxplain
3
+ Version: 0.3.0
4
+ Summary: Explainability for conformal prediction uncertainty — regression and classification.
5
+ Keywords: conformal prediction,uncertainty quantification,explainability,xai,shap,lime,pdp,machine learning,interpretable ml
6
+ Author: Tomas Rodriguez Taborda, Veronica Seguro Varela, Rafael Izbicki, Johnatan Cardona Jimenez
7
+ Requires-Python: >=3.10
8
+ Description-Content-Type: text/markdown
9
+ License-Expression: MIT
10
+ Classifier: Development Status :: 4 - Beta
11
+ Classifier: Intended Audience :: Science/Research
12
+ Classifier: Intended Audience :: Developers
13
+ Classifier: Programming Language :: Python :: 3
14
+ Classifier: Programming Language :: Python :: 3.10
15
+ Classifier: Programming Language :: Python :: 3.11
16
+ Classifier: Programming Language :: Python :: 3.12
17
+ Classifier: Programming Language :: Python :: 3.13
18
+ Classifier: Operating System :: OS Independent
19
+ Classifier: Topic :: Scientific/Engineering :: Artificial Intelligence
20
+ Classifier: Topic :: Software Development :: Libraries :: Python Modules
21
+ License-File: LICENSE
22
+ Requires-Dist: numpy>=1.26
23
+ Requires-Dist: scikit-learn>=1.4
24
+ Requires-Dist: shap>=0.45
25
+ Requires-Dist: crepes>=0.9.1
26
+ Requires-Dist: lime>=0.2
27
+ Requires-Dist: matplotlib>=3.8
28
+ Requires-Dist: pandas>=1.5
29
+ Requires-Dist: pytest>=8.0 ; extra == "dev"
30
+ Requires-Dist: pytest-cov>=5.0 ; extra == "dev"
31
+ Requires-Dist: ruff>=0.4 ; extra == "dev"
32
+ Requires-Dist: mkdocs>=1.6 ; extra == "docs"
33
+ Requires-Dist: build>=1.2 ; extra == "release"
34
+ Requires-Dist: twine>=6.1 ; extra == "release"
35
+ Project-URL: Bug Tracker, https://github.com/torodriguezt/uxplain/issues
36
+ Project-URL: Documentation, https://github.com/torodriguezt/uxplain/tree/main/docs/docs
37
+ Project-URL: Homepage, https://github.com/torodriguezt/uxplain
38
+ Project-URL: Repository, https://github.com/torodriguezt/uxplain
39
+ Provides-Extra: dev
40
+ Provides-Extra: docs
41
+ Provides-Extra: release
42
+
43
+ # uxplain
44
+
45
+ `uxplain` explains which features make a model more or less uncertain. It wraps
46
+ scikit-learn compatible models with conformal prediction and uses SHAP, PDP/ICE
47
+ or LIME to explain summaries of their prediction intervals or sets.
48
+
49
+ The package supports **regression**, through crepes and conformalized quantile
50
+ regression (CQR), and **classification**, through crepes. A single pipeline handles
51
+ model fitting, calibration and explanation.
52
+
53
+ ## Installation
54
+
55
+ Requires Python 3.10 or later.
56
+
57
+ ```bash
58
+ pip install uxplain
59
+ ```
60
+
61
+ ## Regression
62
+
63
+ Explain prediction interval width with SHAP. `fit()` automatically reserves a
64
+ separate calibration sample.
65
+
66
+ ```python
67
+ from sklearn.datasets import make_regression
68
+ from sklearn.ensemble import RandomForestRegressor
69
+ from sklearn.model_selection import train_test_split
70
+ from uxplain import UncertaintyExplanationPipeline
71
+
72
+ X, y = make_regression(n_samples=300, n_features=4, noise=15, random_state=42)
73
+ X_train, X_test, y_train, y_test = train_test_split(X, y, random_state=42)
74
+
75
+ pipe = UncertaintyExplanationPipeline(
76
+ model=RandomForestRegressor(n_estimators=50, random_state=42),
77
+ confidence=0.9,
78
+ xai_method="shap",
79
+ uncertainty_metric="width",
80
+ random_state=42,
81
+ )
82
+ pipe.fit(X_train, y_train)
83
+ result = pipe.explain(X_test[:5], show_plots=False)
84
+
85
+ print(result.interval_width)
86
+ print(result.explanation_values)
87
+ ```
88
+
89
+ ## Classification
90
+
91
+ Use a classifier and `set_size` to explain how many labels enter the prediction set.
92
+
93
+ ```python
94
+ from sklearn.datasets import load_iris
95
+ from sklearn.ensemble import RandomForestClassifier
96
+
97
+ X, y = load_iris(return_X_y=True)
98
+ X_train, X_test, y_train, y_test = train_test_split(X, y, random_state=42)
99
+
100
+ pipe = UncertaintyExplanationPipeline(
101
+ model=RandomForestClassifier(n_estimators=50, random_state=42),
102
+ uncertainty_metric="set_size",
103
+ random_state=42,
104
+ )
105
+ pipe.fit(X_train, y_train)
106
+ result = pipe.explain(X_test[:5], show_plots=False)
107
+
108
+ print(result.prediction_set)
109
+ print(result.explanation_values)
110
+ ```
111
+
112
+ ## Learn more
113
+
114
+ The [getting-started guide](https://github.com/torodriguezt/uxplain/blob/main/docs/docs/getting-started.md)
115
+ covers calibration, other uncertainty summaries and explainer options. See the
116
+ [notebooks](https://github.com/torodriguezt/uxplain/tree/main/notebooks) for more examples
117
+ and the [release guide](https://github.com/torodriguezt/uxplain/blob/main/docs/docs/getting-started.md#releasing)
118
+ for development and publishing.
119
+
120
+ Coverage guarantees depend on the conformal method's
121
+ [statistical assumptions](https://github.com/torodriguezt/uxplain/blob/main/docs/docs/getting-started.md#statistical-scope-and-limitations);
122
+ they do not extend to the explanations themselves.
123
+
124
+ ## Authors
125
+
126
+ - **Tomas Rodriguez Taborda** — Universidad Nacional de Colombia, Medellín
127
+ - **Veronica Seguro Varela** — Universidad Nacional de Colombia, Medellín
128
+ - **Rafael Izbicki** — Federal University of São Carlos
129
+ - **Johnatan Cardona Jimenez** — Universidad Nacional de Colombia, Medellín
130
+
131
+ ## License
132
+
133
+ [MIT](https://github.com/torodriguezt/uxplain/blob/main/LICENSE).
134
+
@@ -0,0 +1,91 @@
1
+ # uxplain
2
+
3
+ `uxplain` explains which features make a model more or less uncertain. It wraps
4
+ scikit-learn compatible models with conformal prediction and uses SHAP, PDP/ICE
5
+ or LIME to explain summaries of their prediction intervals or sets.
6
+
7
+ The package supports **regression**, through crepes and conformalized quantile
8
+ regression (CQR), and **classification**, through crepes. A single pipeline handles
9
+ model fitting, calibration and explanation.
10
+
11
+ ## Installation
12
+
13
+ Requires Python 3.10 or later.
14
+
15
+ ```bash
16
+ pip install uxplain
17
+ ```
18
+
19
+ ## Regression
20
+
21
+ Explain prediction interval width with SHAP. `fit()` automatically reserves a
22
+ separate calibration sample.
23
+
24
+ ```python
25
+ from sklearn.datasets import make_regression
26
+ from sklearn.ensemble import RandomForestRegressor
27
+ from sklearn.model_selection import train_test_split
28
+ from uxplain import UncertaintyExplanationPipeline
29
+
30
+ X, y = make_regression(n_samples=300, n_features=4, noise=15, random_state=42)
31
+ X_train, X_test, y_train, y_test = train_test_split(X, y, random_state=42)
32
+
33
+ pipe = UncertaintyExplanationPipeline(
34
+ model=RandomForestRegressor(n_estimators=50, random_state=42),
35
+ confidence=0.9,
36
+ xai_method="shap",
37
+ uncertainty_metric="width",
38
+ random_state=42,
39
+ )
40
+ pipe.fit(X_train, y_train)
41
+ result = pipe.explain(X_test[:5], show_plots=False)
42
+
43
+ print(result.interval_width)
44
+ print(result.explanation_values)
45
+ ```
46
+
47
+ ## Classification
48
+
49
+ Use a classifier and `set_size` to explain how many labels enter the prediction set.
50
+
51
+ ```python
52
+ from sklearn.datasets import load_iris
53
+ from sklearn.ensemble import RandomForestClassifier
54
+
55
+ X, y = load_iris(return_X_y=True)
56
+ X_train, X_test, y_train, y_test = train_test_split(X, y, random_state=42)
57
+
58
+ pipe = UncertaintyExplanationPipeline(
59
+ model=RandomForestClassifier(n_estimators=50, random_state=42),
60
+ uncertainty_metric="set_size",
61
+ random_state=42,
62
+ )
63
+ pipe.fit(X_train, y_train)
64
+ result = pipe.explain(X_test[:5], show_plots=False)
65
+
66
+ print(result.prediction_set)
67
+ print(result.explanation_values)
68
+ ```
69
+
70
+ ## Learn more
71
+
72
+ The [getting-started guide](https://github.com/torodriguezt/uxplain/blob/main/docs/docs/getting-started.md)
73
+ covers calibration, other uncertainty summaries and explainer options. See the
74
+ [notebooks](https://github.com/torodriguezt/uxplain/tree/main/notebooks) for more examples
75
+ and the [release guide](https://github.com/torodriguezt/uxplain/blob/main/docs/docs/getting-started.md#releasing)
76
+ for development and publishing.
77
+
78
+ Coverage guarantees depend on the conformal method's
79
+ [statistical assumptions](https://github.com/torodriguezt/uxplain/blob/main/docs/docs/getting-started.md#statistical-scope-and-limitations);
80
+ they do not extend to the explanations themselves.
81
+
82
+ ## Authors
83
+
84
+ - **Tomas Rodriguez Taborda** — Universidad Nacional de Colombia, Medellín
85
+ - **Veronica Seguro Varela** — Universidad Nacional de Colombia, Medellín
86
+ - **Rafael Izbicki** — Federal University of São Carlos
87
+ - **Johnatan Cardona Jimenez** — Universidad Nacional de Colombia, Medellín
88
+
89
+ ## License
90
+
91
+ [MIT](https://github.com/torodriguezt/uxplain/blob/main/LICENSE).
@@ -0,0 +1,109 @@
1
+ [build-system]
2
+ requires = ["flit_core >=3.12,<4"]
3
+ build-backend = "flit_core.buildapi"
4
+
5
+ [project]
6
+ name = "uxplain"
7
+ dynamic = ["version"]
8
+ description = "Explainability for conformal prediction uncertainty — regression and classification."
9
+ authors = [
10
+ { name = "Tomas Rodriguez Taborda" },
11
+ { name = "Veronica Seguro Varela" },
12
+ { name = "Rafael Izbicki" },
13
+ { name = "Johnatan Cardona Jimenez" },
14
+ ]
15
+ license = "MIT"
16
+ license-files = ["LICENSE"]
17
+ readme = "README.md"
18
+ keywords = [
19
+ "conformal prediction",
20
+ "uncertainty quantification",
21
+ "explainability",
22
+ "xai",
23
+ "shap",
24
+ "lime",
25
+ "pdp",
26
+ "machine learning",
27
+ "interpretable ml",
28
+ ]
29
+ classifiers = [
30
+ "Development Status :: 4 - Beta",
31
+ "Intended Audience :: Science/Research",
32
+ "Intended Audience :: Developers",
33
+ "Programming Language :: Python :: 3",
34
+ "Programming Language :: Python :: 3.10",
35
+ "Programming Language :: Python :: 3.11",
36
+ "Programming Language :: Python :: 3.12",
37
+ "Programming Language :: Python :: 3.13",
38
+ "Operating System :: OS Independent",
39
+ "Topic :: Scientific/Engineering :: Artificial Intelligence",
40
+ "Topic :: Software Development :: Libraries :: Python Modules",
41
+ ]
42
+ requires-python = ">=3.10"
43
+ dependencies = [
44
+ "numpy>=1.26",
45
+ "scikit-learn>=1.4",
46
+ "shap>=0.45",
47
+ "crepes>=0.9.1",
48
+ "lime>=0.2",
49
+ "matplotlib>=3.8",
50
+ "pandas>=1.5",
51
+ ]
52
+
53
+ [project.urls]
54
+ Homepage = "https://github.com/torodriguezt/uxplain"
55
+ Repository = "https://github.com/torodriguezt/uxplain"
56
+ Documentation = "https://github.com/torodriguezt/uxplain/tree/main/docs/docs"
57
+ "Bug Tracker" = "https://github.com/torodriguezt/uxplain/issues"
58
+
59
+ [project.optional-dependencies]
60
+ dev = [
61
+ "pytest>=8.0",
62
+ "pytest-cov>=5.0",
63
+ "ruff>=0.4",
64
+ ]
65
+ release = [
66
+ "build>=1.2",
67
+ "twine>=6.1",
68
+ ]
69
+ docs = ["mkdocs>=1.6"]
70
+
71
+ [tool.flit.sdist]
72
+ include = ["LICENSE", "README.md", "CHANGELOG.md", "pyproject.toml", "tests/"]
73
+ exclude = [
74
+ "docs/",
75
+ "notebooks/",
76
+ "models/",
77
+ "reports/",
78
+ "references/",
79
+ "replication/",
80
+ "figures/",
81
+ "uxplain-main/",
82
+ "*.tex",
83
+ "*.bib",
84
+ "*.pdf",
85
+ "*.zip",
86
+ "SUBMISSION_CHECKLIST.md",
87
+ "demo.ipynb",
88
+ "Makefile",
89
+ "requirements.txt",
90
+ ".github/",
91
+ "**/__pycache__/",
92
+ "**/*.pyc",
93
+ ]
94
+
95
+ [tool.pytest.ini_options]
96
+ testpaths = ["tests"]
97
+
98
+ [tool.ruff]
99
+ line-length = 99
100
+ src = ["uxplain"]
101
+ include = ["pyproject.toml", "uxplain/**/*.py"]
102
+ extend-exclude = ["uxplain-main"]
103
+
104
+ [tool.ruff.lint]
105
+ extend-select = ["I"] # Add import sorting
106
+
107
+ [tool.ruff.lint.isort]
108
+ known-first-party = ["uxplain"]
109
+ force-sort-within-sections = true
@@ -0,0 +1,76 @@
1
+ import numpy as np
2
+ import pytest
3
+ from sklearn.datasets import make_classification
4
+ from sklearn.ensemble import GradientBoostingRegressor, RandomForestClassifier
5
+ from sklearn.linear_model import Ridge
6
+ from sklearn.model_selection import train_test_split
7
+
8
+
9
+ @pytest.fixture(scope="module")
10
+ def data():
11
+ rng = np.random.default_rng(0)
12
+ n = 200
13
+ X = rng.standard_normal((n, 4))
14
+ y = X @ np.array([1.5, -2.0, 0.5, 1.0]) + rng.standard_normal(n) * 0.5
15
+ X_tr, X_te, y_tr, y_te = train_test_split(X, y, test_size=0.25, random_state=0)
16
+ X_tr, X_cal, y_tr, y_cal = train_test_split(X_tr, y_tr, test_size=0.25, random_state=0)
17
+ return {
18
+ "X_train": X_tr,
19
+ "X_calib": X_cal,
20
+ "X_test": X_te,
21
+ "y_train": y_tr,
22
+ "y_calib": y_cal,
23
+ "y_test": y_te,
24
+ }
25
+
26
+
27
+ @pytest.fixture
28
+ def ridge():
29
+ return Ridge()
30
+
31
+
32
+ @pytest.fixture
33
+ def quantile_lower():
34
+ return GradientBoostingRegressor(
35
+ loss="quantile", alpha=0.05, n_estimators=30, random_state=0
36
+ )
37
+
38
+
39
+ @pytest.fixture
40
+ def quantile_upper():
41
+ return GradientBoostingRegressor(
42
+ loss="quantile", alpha=0.95, n_estimators=30, random_state=0
43
+ )
44
+
45
+
46
+ @pytest.fixture(scope="module")
47
+ def classification_data():
48
+ X, y = make_classification(
49
+ n_samples=400,
50
+ n_features=4,
51
+ n_informative=3,
52
+ n_redundant=0,
53
+ n_classes=3,
54
+ n_clusters_per_class=1,
55
+ random_state=0,
56
+ )
57
+ X_tr, X_te, y_tr, y_te = train_test_split(
58
+ X, y, test_size=0.25, random_state=0,
59
+ )
60
+ X_tr, X_cal, y_tr, y_cal = train_test_split(
61
+ X_tr, y_tr, test_size=0.25, random_state=0,
62
+ )
63
+ return {
64
+ "X_train": X_tr,
65
+ "X_calib": X_cal,
66
+ "X_test": X_te,
67
+ "y_train": y_tr,
68
+ "y_calib": y_cal,
69
+ "y_test": y_te,
70
+ "n_classes": 3,
71
+ }
72
+
73
+
74
+ @pytest.fixture
75
+ def classifier():
76
+ return RandomForestClassifier(n_estimators=20, random_state=0)