geds-python 0.1.0a3__tar.gz → 0.1.1__tar.gz
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- geds_python-0.1.1/CHANGELOG.md +75 -0
- {geds_python-0.1.0a3 → geds_python-0.1.1}/CITATION.cff +1 -2
- geds_python-0.1.1/PKG-INFO +416 -0
- geds_python-0.1.1/README.md +377 -0
- {geds_python-0.1.0a3 → geds_python-0.1.1}/RELEASING.md +17 -7
- {geds_python-0.1.0a3 → geds_python-0.1.1}/pyproject.toml +2 -2
- geds_python-0.1.1/src/geds/__init__.py +25 -0
- {geds_python-0.1.0a3 → geds_python-0.1.1}/src/geds/_backend.py +176 -8
- geds_python-0.1.1/src/geds/_estimators.py +844 -0
- geds_python-0.1.1/src/geds/_validation.py +148 -0
- geds_python-0.1.1/tests/test_estimators.py +842 -0
- geds_python-0.1.0a3/CHANGELOG.md +0 -31
- geds_python-0.1.0a3/PKG-INFO +0 -208
- geds_python-0.1.0a3/README.md +0 -169
- geds_python-0.1.0a3/src/geds/__init__.py +0 -15
- geds_python-0.1.0a3/src/geds/_estimators.py +0 -385
- geds_python-0.1.0a3/tests/test_estimators.py +0 -218
- {geds_python-0.1.0a3 → geds_python-0.1.1}/.gitignore +0 -0
- {geds_python-0.1.0a3 → geds_python-0.1.1}/LICENSE +0 -0
- {geds_python-0.1.0a3 → geds_python-0.1.1}/src/geds/_plotting.py +0 -0
- {geds_python-0.1.0a3 → geds_python-0.1.1}/src/geds/check.py +0 -0
- {geds_python-0.1.0a3 → geds_python-0.1.1}/src/geds/py.typed +0 -0
|
@@ -0,0 +1,75 @@
|
|
|
1
|
+
# Changelog
|
|
2
|
+
|
|
3
|
+
All notable changes to GeDS for Python are documented in this file.
|
|
4
|
+
|
|
5
|
+
The project follows [Semantic Versioning](https://semver.org/). Versions use
|
|
6
|
+
the Python packaging form of pre-release identifiers, such as `0.1.0a1` for
|
|
7
|
+
the first alpha release.
|
|
8
|
+
|
|
9
|
+
## 0.1.1 - 2026-09-28
|
|
10
|
+
|
|
11
|
+
- Let R select `NGeDS()`'s dimension-specific default stopping rule: `RD`
|
|
12
|
+
for univariate and `SR` for joint bivariate or higher-dimensional fits.
|
|
13
|
+
- Add an executed UK spot-rate vignette with a univariate fit and two
|
|
14
|
+
bivariate sparse-maturity reconstructions, linked from the README.
|
|
15
|
+
- Correct CI installation so integration tests use the newly built wheel,
|
|
16
|
+
rather than a same-version distribution from PyPI.
|
|
17
|
+
|
|
18
|
+
## 0.1.0 - 2026-09-25
|
|
19
|
+
|
|
20
|
+
- Expose R's boosted base-learner importance alongside its existing iteration
|
|
21
|
+
count, and clarify remaining specialized R-only utilities in the feature audit.
|
|
22
|
+
- Add a Gaussian-only Python interface to R's specialized `crossv_GeDS()`
|
|
23
|
+
grid search and its MSE/knot/iteration summaries.
|
|
24
|
+
- Save R's single-learner boosting-iteration visualization as a multipage PDF
|
|
25
|
+
through a Python model method.
|
|
26
|
+
- Expose R's `Derive()`, `Integrate()`, `PPolyRep()`, and `shapeConstrain()`
|
|
27
|
+
through fitted-model methods, with direct R parity tests and explicit model
|
|
28
|
+
restrictions.
|
|
29
|
+
- Check alternate spline orders, generalized-model utilities, mixed-term
|
|
30
|
+
additive constraints, and sequential scikit-learn cross-validation.
|
|
31
|
+
- Add additive GAM and gradient-boosting estimators backed by the R package's
|
|
32
|
+
`NGeDSgam()` and `NGeDSboost()` functions, with independently checked R/Python
|
|
33
|
+
predictions and no duplicate statistical implementation.
|
|
34
|
+
- Support named additive spline terms, joint spline terms, parametric features,
|
|
35
|
+
selected loss families, and individual base-learner predictions.
|
|
36
|
+
|
|
37
|
+
## 0.1.0a4 - release candidate (not published)
|
|
38
|
+
|
|
39
|
+
- Begin an R-to-Python feature audit for the next coordinated release.
|
|
40
|
+
- Add order-specific access to R's deviance, log likelihood, and coefficient
|
|
41
|
+
confidence intervals.
|
|
42
|
+
- Add univariate fit/prediction offsets and named `predict_terms()` output,
|
|
43
|
+
delegated to the locally fixed R package. These require the corresponding
|
|
44
|
+
R GitHub changes before publication.
|
|
45
|
+
- Expand independent R/Python parity tests for weighted normal and Poisson
|
|
46
|
+
fits, bivariate fits and confidence intervals; clarify joint-spline feature
|
|
47
|
+
selection.
|
|
48
|
+
- Copy arrays extracted from R into Python-owned memory so later R calls cannot
|
|
49
|
+
change stored coefficients, knots, or predictions.
|
|
50
|
+
- Require GeDS 0.3.6 with the Python bridge capability marker, and pin the
|
|
51
|
+
tested GitHub commit in installation and CI instructions.
|
|
52
|
+
|
|
53
|
+
## 0.1.0a3 - 2026-09-23
|
|
54
|
+
|
|
55
|
+
- Add `python -m geds.check` for human-readable or JSON environment checks.
|
|
56
|
+
- Add `geds.plot_fit()` for Python-native visualization of univariate fits and
|
|
57
|
+
internal knots.
|
|
58
|
+
- Expand installation and R-library troubleshooting guidance.
|
|
59
|
+
|
|
60
|
+
## 0.1.0a2 - 2026-09-20
|
|
61
|
+
|
|
62
|
+
- Preload R's core numerical DLLs on Windows so embedded R 4.6 can load
|
|
63
|
+
recommended packages without requiring Rtools on the user's `PATH`.
|
|
64
|
+
|
|
65
|
+
## 0.1.0a1 - 2026-09-20
|
|
66
|
+
|
|
67
|
+
Initial alpha release.
|
|
68
|
+
|
|
69
|
+
- Add scikit-learn-style estimators for normal and generalized GeDS models.
|
|
70
|
+
- Delegate all statistical fitting, prediction, coefficient extraction, and
|
|
71
|
+
knot extraction to the GeDS R package.
|
|
72
|
+
- Support pandas and NumPy inputs, numeric and categorical parametric terms,
|
|
73
|
+
sample weights, model serialization, and environment diagnostics.
|
|
74
|
+
- Support Windows and Linux with R 4.6.1 and Python 3.10 or 3.12 in CI.
|
|
75
|
+
- Add a nonlinear regression example illustrating adaptive knot placement.
|
|
@@ -2,8 +2,7 @@ cff-version: 1.2.0
|
|
|
2
2
|
message: "If you use GeDS for Python, please cite this software and the GeDS methodology references."
|
|
3
3
|
title: "GeDS for Python"
|
|
4
4
|
type: software
|
|
5
|
-
version: 0.1.
|
|
6
|
-
date-released: 2026-09-23
|
|
5
|
+
version: 0.1.1
|
|
7
6
|
license: GPL-3.0-only
|
|
8
7
|
repository-code: "https://github.com/emilioluissaenzguillen/GeDS-python"
|
|
9
8
|
authors:
|
|
@@ -0,0 +1,416 @@
|
|
|
1
|
+
Metadata-Version: 2.5
|
|
2
|
+
Name: geds-python
|
|
3
|
+
Version: 0.1.1
|
|
4
|
+
Summary: Python estimators backed by the GeDS R package
|
|
5
|
+
Project-URL: Homepage, https://github.com/emilioluissaenzguillen/GeDS-python
|
|
6
|
+
Project-URL: Repository, https://github.com/emilioluissaenzguillen/GeDS-python
|
|
7
|
+
Project-URL: Issues, https://github.com/emilioluissaenzguillen/GeDS-python/issues
|
|
8
|
+
Project-URL: Changelog, https://github.com/emilioluissaenzguillen/GeDS-python/blob/main/CHANGELOG.md
|
|
9
|
+
Project-URL: R package, https://github.com/emilioluissaenzguillen/GeDS
|
|
10
|
+
Author: Dimitrina S. Dimitrova, Vladimir K. Kaishev, Andrea Lattuada, Emilio L. Sáenz Guillén, Richard J. Verrall
|
|
11
|
+
Maintainer-email: "Emilio L. Sáenz Guillén" <emilioluissaenzguillen@gmail.com>
|
|
12
|
+
License-Expression: GPL-3.0-only
|
|
13
|
+
License-File: LICENSE
|
|
14
|
+
Keywords: R,regression,scikit-learn,splines,statistics
|
|
15
|
+
Classifier: Development Status :: 4 - Beta
|
|
16
|
+
Classifier: Intended Audience :: Science/Research
|
|
17
|
+
Classifier: License :: OSI Approved :: GNU General Public License v3 (GPLv3)
|
|
18
|
+
Classifier: Operating System :: Microsoft :: Windows
|
|
19
|
+
Classifier: Operating System :: POSIX :: Linux
|
|
20
|
+
Classifier: Programming Language :: Python :: 3
|
|
21
|
+
Classifier: Programming Language :: Python :: 3.10
|
|
22
|
+
Classifier: Programming Language :: Python :: 3.11
|
|
23
|
+
Classifier: Programming Language :: Python :: 3.12
|
|
24
|
+
Classifier: Topic :: Scientific/Engineering
|
|
25
|
+
Requires-Python: >=3.10
|
|
26
|
+
Requires-Dist: numpy>=1.24
|
|
27
|
+
Requires-Dist: pandas>=2.0
|
|
28
|
+
Requires-Dist: rpy2<3.7,>=3.6.7
|
|
29
|
+
Requires-Dist: scikit-learn>=1.4
|
|
30
|
+
Provides-Extra: dev
|
|
31
|
+
Requires-Dist: build>=1.2; extra == 'dev'
|
|
32
|
+
Requires-Dist: matplotlib>=3.8; extra == 'dev'
|
|
33
|
+
Requires-Dist: pytest>=8; extra == 'dev'
|
|
34
|
+
Provides-Extra: plot
|
|
35
|
+
Requires-Dist: matplotlib>=3.8; extra == 'plot'
|
|
36
|
+
Provides-Extra: test
|
|
37
|
+
Requires-Dist: pytest>=8; extra == 'test'
|
|
38
|
+
Description-Content-Type: text/markdown
|
|
39
|
+
|
|
40
|
+
# GeDS for Python
|
|
41
|
+
|
|
42
|
+
This package provides a Python interface to the
|
|
43
|
+
[GeDS R package](https://github.com/emilioluissaenzguillen/GeDS). The R package
|
|
44
|
+
is the sole implementation of the statistical methods. Python supplies a
|
|
45
|
+
scikit-learn-style API, pandas/NumPy conversion, environment diagnostics, and
|
|
46
|
+
model serialization.
|
|
47
|
+
|
|
48
|
+
## Requirements
|
|
49
|
+
|
|
50
|
+
- R 4.4 or newer (R 4.6.1 is used for development)
|
|
51
|
+
- GeDS 0.3.6 or newer with the Python bridge fixes
|
|
52
|
+
- Python 3.10 or newer
|
|
53
|
+
|
|
54
|
+
Install the Python package, including the optional plotting dependency used in
|
|
55
|
+
the example:
|
|
56
|
+
|
|
57
|
+
```console
|
|
58
|
+
python -m pip install "geds-python[plot]"
|
|
59
|
+
```
|
|
60
|
+
|
|
61
|
+
Install the R package separately, using R 4.6.1 or another supported R
|
|
62
|
+
installation. Install the tested GeDS 0.3.6 build from GitHub:
|
|
63
|
+
|
|
64
|
+
```r
|
|
65
|
+
install.packages("remotes")
|
|
66
|
+
remotes::install_git(
|
|
67
|
+
"https://github.com/emilioluissaenzguillen/GeDS.git",
|
|
68
|
+
ref = "91b8ddd13aae8f39994c87fc356f05da4799f911",
|
|
69
|
+
dependencies = NA, upgrade = "never"
|
|
70
|
+
)
|
|
71
|
+
```
|
|
72
|
+
|
|
73
|
+
An older GeDS build, even one reporting version `0.3.6`, may lack the fixes
|
|
74
|
+
required by this wrapper. `install.packages("GeDS")` alone is not guaranteed
|
|
75
|
+
to provide them while the CRAN review follows its separate schedule.
|
|
76
|
+
Check the installed R version with `packageVersion("GeDS")`.
|
|
77
|
+
|
|
78
|
+
On Windows, building the GitHub source package requires Rtools compatible
|
|
79
|
+
with the selected R installation. GeDS remains version `0.3.6` on GitHub;
|
|
80
|
+
the wrapper checks an internal compatibility marker for the fit and prediction
|
|
81
|
+
fixes as well as the package version.
|
|
82
|
+
|
|
83
|
+
The wrapper discovers the newest R installation under `Program Files/R` on
|
|
84
|
+
Windows or uses `Rscript` from `PATH` on other platforms. Set `R_HOME` to select
|
|
85
|
+
a particular R installation. If GeDS is installed in a non-default R library,
|
|
86
|
+
set `GEDS_R_LIBRARY` to that library directory before importing `geds`.
|
|
87
|
+
|
|
88
|
+
The Python and R packages have independent release cycles. `geds-python`
|
|
89
|
+
checks the installed GeDS version when its backend first starts and reports the
|
|
90
|
+
selected R installation and package library through `geds.diagnostics()`.
|
|
91
|
+
|
|
92
|
+
Check the backend before fitting:
|
|
93
|
+
|
|
94
|
+
```console
|
|
95
|
+
python -m geds.check
|
|
96
|
+
```
|
|
97
|
+
|
|
98
|
+
For a machine-readable report, use `python -m geds.check --json`. The same
|
|
99
|
+
information is available inside Python:
|
|
100
|
+
|
|
101
|
+
```python
|
|
102
|
+
import geds
|
|
103
|
+
|
|
104
|
+
print(geds.diagnostics())
|
|
105
|
+
```
|
|
106
|
+
|
|
107
|
+
### Selecting R and its package library
|
|
108
|
+
|
|
109
|
+
Usually no configuration is necessary. If several R installations are
|
|
110
|
+
available, select one before starting Python:
|
|
111
|
+
|
|
112
|
+
```powershell
|
|
113
|
+
$env:R_HOME = "C:\Program Files\R\R-4.6.1"
|
|
114
|
+
python -m geds.check
|
|
115
|
+
```
|
|
116
|
+
|
|
117
|
+
```bash
|
|
118
|
+
export R_HOME="/Library/Frameworks/R.framework/Resources" # macOS
|
|
119
|
+
# export R_HOME="/usr/lib/R" # Linux
|
|
120
|
+
python -m geds.check
|
|
121
|
+
```
|
|
122
|
+
|
|
123
|
+
If GeDS is installed in a personal or otherwise non-default R library, set
|
|
124
|
+
`GEDS_R_LIBRARY` to the directory that contains the `GeDS` folder. You can
|
|
125
|
+
find that directory from R with `find.package("GeDS")`; use its parent
|
|
126
|
+
directory as `GEDS_R_LIBRARY`.
|
|
127
|
+
|
|
128
|
+
If the check reports that R is missing, install R or set `R_HOME`. If it finds
|
|
129
|
+
R but not GeDS, start that same R installation and run
|
|
130
|
+
one of the GeDS installation commands above, then rerun the check.
|
|
131
|
+
|
|
132
|
+
## Example
|
|
133
|
+
|
|
134
|
+
For a finance example, see the [UK interest-rate notebook](https://github.com/emilioluissaenzguillen/GeDS-python/blob/main/examples/uk_yield_curves.ipynb).
|
|
135
|
+
It fits a 10-year rate over time and a joint time-by-maturity surface using the
|
|
136
|
+
Bank of England's published nominal spot-rate curves. The notebook downloads
|
|
137
|
+
the source data when run; the repository does not redistribute the archive.
|
|
138
|
+
|
|
139
|
+
Install the optional plotting dependency with
|
|
140
|
+
`python -m pip install "geds-python[plot]"`, then fit and visualize a nonlinear
|
|
141
|
+
regression:
|
|
142
|
+
|
|
143
|
+
```python
|
|
144
|
+
import matplotlib.pyplot as plt
|
|
145
|
+
import numpy as np
|
|
146
|
+
import pandas as pd
|
|
147
|
+
|
|
148
|
+
from geds import GeDSRegressor, plot_fit
|
|
149
|
+
|
|
150
|
+
rng = np.random.RandomState(123)
|
|
151
|
+
n = 500
|
|
152
|
+
|
|
153
|
+
|
|
154
|
+
def f_1(x):
|
|
155
|
+
return (10 * x / (1 + 100 * x**2)) * 4 + 4
|
|
156
|
+
|
|
157
|
+
|
|
158
|
+
x = np.sort(rng.uniform(-2.0, 2.0, size=n))
|
|
159
|
+
means = f_1(x)
|
|
160
|
+
y = rng.normal(means, scale=0.1)
|
|
161
|
+
X = pd.DataFrame({"x": x})
|
|
162
|
+
|
|
163
|
+
model = GeDSRegressor(order=3).fit(X, y)
|
|
164
|
+
knots = np.asarray(model.knots_, dtype=float)
|
|
165
|
+
|
|
166
|
+
print("Internal knots:", knots)
|
|
167
|
+
|
|
168
|
+
fig, ax = plt.subplots()
|
|
169
|
+
plot_fit(model, X, y, ax=ax)
|
|
170
|
+
grid_x = np.linspace(x.min(), x.max(), 500)
|
|
171
|
+
ax.plot(
|
|
172
|
+
grid_x,
|
|
173
|
+
f_1(grid_x),
|
|
174
|
+
color="0.25",
|
|
175
|
+
linestyle=":",
|
|
176
|
+
linewidth=2,
|
|
177
|
+
label="True mean",
|
|
178
|
+
)
|
|
179
|
+
ax.set(ylabel="y")
|
|
180
|
+
ax.legend()
|
|
181
|
+
fig.tight_layout()
|
|
182
|
+
plt.show()
|
|
183
|
+
```
|
|
184
|
+
|
|
185
|
+
With GeDS 0.3.6 and R 4.6.1, this seeded example fits 16 internal knots.
|
|
186
|
+
The dashed vertical lines show how GeDS places more knots around the sharp
|
|
187
|
+
variation near zero while retaining knots across the wider domain.
|
|
188
|
+
|
|
189
|
+
`GeDSRegressor` delegates to `GeDS::NGeDS()`. For exponential-family models,
|
|
190
|
+
use `GeDSGeneralizedRegressor`, which delegates to `GeDS::GGeDS()`.
|
|
191
|
+
By default, `GeDSRegressor` leaves the stopping rule to R: `RD` for one
|
|
192
|
+
spline feature and `SR` for a joint spline with two or more features. Set
|
|
193
|
+
`stop_type="RD"` or `stop_type="SR"` to override it explicitly.
|
|
194
|
+
`GeDSGeneralizedRegressor` uses R's `SR` default. GAM and boosting base
|
|
195
|
+
learners also use their R implementation's dimension-specific rules.
|
|
196
|
+
|
|
197
|
+
For fitted models, `get_deviance(order=...)`, `get_log_likelihood(order=...)`,
|
|
198
|
+
and `get_confidence_intervals(order=..., level=...)` call the corresponding R
|
|
199
|
+
methods. Confidence intervals are returned as a pandas DataFrame with `lower`
|
|
200
|
+
and `upper` columns. As in R, these are coefficient intervals, not confidence
|
|
201
|
+
bands for the fitted curve.
|
|
202
|
+
|
|
203
|
+
The estimators also work with standard scikit-learn tools such as
|
|
204
|
+
`cross_val_score()` and `GridSearchCV`. Use sequential execution (`n_jobs=1`)
|
|
205
|
+
when cross-validating: the wrapper embeds R in the Python process, and
|
|
206
|
+
parallel-worker behavior is not part of the supported interface.
|
|
207
|
+
|
|
208
|
+
R also has a specialized `crossv_GeDS()` routine, which returns a parameter
|
|
209
|
+
grid with cross-validated mean squared error and knot/iteration summaries.
|
|
210
|
+
Use its Python interface when those R-specific results are needed:
|
|
211
|
+
|
|
212
|
+
```python
|
|
213
|
+
from geds import cross_validate_geds
|
|
214
|
+
|
|
215
|
+
cv = cross_validate_geds(
|
|
216
|
+
GeDSRegressor(order=3), X, y,
|
|
217
|
+
{"beta": [0.5, 0.7], "phi": [0.95], "q": [2]},
|
|
218
|
+
n_folds=5, n_cores=1, random_state=123,
|
|
219
|
+
)
|
|
220
|
+
print(cv.best_params)
|
|
221
|
+
print(cv.results)
|
|
222
|
+
```
|
|
223
|
+
|
|
224
|
+
This delegates the entire search to R and does not fit or change the input
|
|
225
|
+
estimator. It currently supports Gaussian models only, accepts the R tuning
|
|
226
|
+
parameters `beta`, `phi`, `q`, and (for boosting) `int_knots_init` and
|
|
227
|
+
`shrinkage`, and defaults to one R worker. R's current non-boost routine does
|
|
228
|
+
not forward other fitting settings; the Python interface rejects custom
|
|
229
|
+
settings it would otherwise silently ignore. Use scikit-learn's grid search
|
|
230
|
+
when you need those settings or a non-Gaussian family.
|
|
231
|
+
|
|
232
|
+
For a fitted univariate spline without extra linear features, R's calculus
|
|
233
|
+
and spline-conversion utilities are available as model methods:
|
|
234
|
+
|
|
235
|
+
```python
|
|
236
|
+
slopes = model.derive([-0.5, 0.0, 0.5], derivative_order=1)
|
|
237
|
+
areas = model.integrate(-1.0, [-0.5, 0.0, 0.5])
|
|
238
|
+
piece_knots, piece_coefficients = model.piecewise_polynomial()
|
|
239
|
+
```
|
|
240
|
+
|
|
241
|
+
`derive()` and `integrate()` operate on the predictor (link) scale, as in R.
|
|
242
|
+
`piecewise_polynomial()` returns the R `PPolyRep()` knot vector and coefficient
|
|
243
|
+
matrix; its last coefficient row is extraneous in R's representation. These
|
|
244
|
+
methods use the estimator's selected spline order unless `order=` is given.
|
|
245
|
+
|
|
246
|
+
For a normal univariate fit, impose a shape constraint without changing the
|
|
247
|
+
original fitted model:
|
|
248
|
+
|
|
249
|
+
```python
|
|
250
|
+
increasing_model = model.shape_constrain("increasing")
|
|
251
|
+
increasing_and_convex = model.shape_constrain(["increasing", "convex"])
|
|
252
|
+
```
|
|
253
|
+
|
|
254
|
+
This calls R's `shapeConstrain()` and returns a new Python estimator. R also
|
|
255
|
+
supports constraints on one selected univariate smoother in Gaussian GAM and
|
|
256
|
+
boosting fits, via `shape_constrain(..., base_learner="f(x)")`. Those additive
|
|
257
|
+
fits must use `normalize_data=False`. Constrained fits do not provide the usual
|
|
258
|
+
unconstrained coefficient confidence intervals.
|
|
259
|
+
|
|
260
|
+
For count data, the generalized estimator uses `GeDS::GGeDS()` and supports
|
|
261
|
+
both response-scale and link-scale prediction:
|
|
262
|
+
|
|
263
|
+
```python
|
|
264
|
+
import numpy as np
|
|
265
|
+
import pandas as pd
|
|
266
|
+
from geds import GeDSGeneralizedRegressor, plot_fit
|
|
267
|
+
|
|
268
|
+
rng = np.random.default_rng(123)
|
|
269
|
+
x = np.sort(rng.uniform(-2, 2, 120))
|
|
270
|
+
X = pd.DataFrame({"x": x})
|
|
271
|
+
counts = rng.poisson(np.exp(1 + np.sin(x)))
|
|
272
|
+
|
|
273
|
+
model = GeDSGeneralizedRegressor(
|
|
274
|
+
family="poisson", beta=0.2, phi=0.95, min_internal_knots=3
|
|
275
|
+
).fit(X, counts)
|
|
276
|
+
mean_counts = model.predict(X)
|
|
277
|
+
log_mean_counts = model.predict_link(X)
|
|
278
|
+
ax = plot_fit(model, X, counts)
|
|
279
|
+
```
|
|
280
|
+
|
|
281
|
+
`min_internal_knots` controls the minimum number of stage-A knots; it is used
|
|
282
|
+
here to make a small sample's fitted spline visible. GeDS determines the
|
|
283
|
+
final knot positions.
|
|
284
|
+
|
|
285
|
+
For a univariate spline with a known offset (for example log exposure in a
|
|
286
|
+
Poisson model), pass one offset value per observation to both fitting and
|
|
287
|
+
prediction. These values are on the link scale:
|
|
288
|
+
|
|
289
|
+
```python
|
|
290
|
+
import numpy as np
|
|
291
|
+
import pandas as pd
|
|
292
|
+
from geds import GeDSGeneralizedRegressor
|
|
293
|
+
|
|
294
|
+
rng = np.random.default_rng(321)
|
|
295
|
+
x = np.linspace(-1.5, 1.5, 90)
|
|
296
|
+
X = pd.DataFrame({"x": x})
|
|
297
|
+
exposure = np.linspace(1.1, 2.0, len(x))
|
|
298
|
+
counts = rng.poisson(exposure * np.exp(1.4 + np.sin(x))) + 1
|
|
299
|
+
log_exposure = np.log(exposure)
|
|
300
|
+
model = GeDSGeneralizedRegressor(
|
|
301
|
+
family="poisson", spline_features=["x"], order=2,
|
|
302
|
+
higher_order=False,
|
|
303
|
+
).fit(X, counts, offset=log_exposure)
|
|
304
|
+
expected_counts = model.predict(X, offset=log_exposure)
|
|
305
|
+
contributions = model.predict_terms(X, offset=log_exposure)
|
|
306
|
+
```
|
|
307
|
+
|
|
308
|
+
`predict_terms()` returns a DataFrame of the R spline and parametric term
|
|
309
|
+
contributions. Its rows sum to the link prediction after adding the offset;
|
|
310
|
+
the offset is not itself a term column. Offset prediction currently supports
|
|
311
|
+
one spline feature only because the R bivariate prediction method does not
|
|
312
|
+
apply new-data offsets consistently. A model fitted with an offset requires
|
|
313
|
+
an offset at prediction time.
|
|
314
|
+
|
|
315
|
+
This offset interface requires the GeDS GitHub fit and prediction fixes. A
|
|
316
|
+
earlier GeDS `0.3.6` installation without those fixes may mishandle
|
|
317
|
+
generalized-model offsets; the backend rejects that version.
|
|
318
|
+
|
|
319
|
+
Choose spline and parametric components explicitly for mixed data:
|
|
320
|
+
|
|
321
|
+
```python
|
|
322
|
+
model = GeDSRegressor(
|
|
323
|
+
spline_features=["x"],
|
|
324
|
+
linear_features=["group"],
|
|
325
|
+
).fit(X, y)
|
|
326
|
+
```
|
|
327
|
+
|
|
328
|
+
Spline features must be numeric. Parametric features may be numeric or
|
|
329
|
+
categorical; their encoding is performed by the R package so fitting and
|
|
330
|
+
prediction use R's native factor semantics.
|
|
331
|
+
If `spline_features` is omitted, all columns are used in a single joint spline
|
|
332
|
+
term. Select `spline_features=["x"]` and `linear_features=["group"]` to keep
|
|
333
|
+
`group` parametric instead. Two spline features create a joint bivariate
|
|
334
|
+
surface, not two separate additive smooths; R's support for more than two
|
|
335
|
+
spline features is experimental. With named pandas columns, prediction may
|
|
336
|
+
receive columns in a different order because the wrapper restores the fitted
|
|
337
|
+
column order before calling R.
|
|
338
|
+
|
|
339
|
+
### Additive GAM and boosting models
|
|
340
|
+
|
|
341
|
+
Use `GeDSGAMRegressor` for R's `NGeDSgam()` and `GeDSBoostRegressor` for
|
|
342
|
+
`NGeDSboost()`. Each entry in `spline_terms` is one additive smooth. Put two
|
|
343
|
+
features in the same entry for a joint surface. If omitted, each non-linear
|
|
344
|
+
feature gets its own smooth; `linear_features` selects parametric terms.
|
|
345
|
+
|
|
346
|
+
```python
|
|
347
|
+
import numpy as np
|
|
348
|
+
import pandas as pd
|
|
349
|
+
from geds import GeDSGAMRegressor, GeDSBoostRegressor
|
|
350
|
+
|
|
351
|
+
x = np.linspace(-2, 2, 100)
|
|
352
|
+
X = pd.DataFrame({"x": x, "z": x**2})
|
|
353
|
+
y = np.sin(x) + 0.3 * x**2
|
|
354
|
+
|
|
355
|
+
gam = GeDSGAMRegressor(
|
|
356
|
+
spline_terms=[("x",), ("z",)], max_iterations=10
|
|
357
|
+
).fit(X, y)
|
|
358
|
+
boost = GeDSBoostRegressor(
|
|
359
|
+
spline_terms=[("x",), ("z",)], max_iterations=20
|
|
360
|
+
).fit(X, y)
|
|
361
|
+
|
|
362
|
+
gam_predictions = gam.predict(X)
|
|
363
|
+
boost_predictions = boost.predict(X)
|
|
364
|
+
x_contribution = gam.predict_component(X, "f(x)")
|
|
365
|
+
importance = boost.get_base_learner_importance()
|
|
366
|
+
```
|
|
367
|
+
|
|
368
|
+
Both estimators expose `predict_link()`, order-specific coefficients, knots,
|
|
369
|
+
deviance, log likelihood, and coefficient confidence intervals through R.
|
|
370
|
+
`predict_component()` delegates a named learner prediction to R. The R
|
|
371
|
+
GAM/boost prediction method does not support `type="terms"`, so these
|
|
372
|
+
estimators do not offer `predict_terms()`. The GAM wrapper supports the
|
|
373
|
+
families accepted by `GeDSGeneralizedRegressor`; for binomial fits it accepts
|
|
374
|
+
0/1 responses and creates the factor required by R. The boosting
|
|
375
|
+
wrapper maps `gaussian`, `poisson`, `binomial`, and `gamma` to mboost families.
|
|
376
|
+
The fitted boosting estimator's `n_iter_` is R's total boosting iteration
|
|
377
|
+
count. `get_base_learner_importance()` returns R's `bl_imp()` in-bag risk
|
|
378
|
+
reductions as a pandas Series with the original Python feature names.
|
|
379
|
+
For a boosted fit with one univariate spline feature, R's iteration plots can
|
|
380
|
+
be saved to a multipage PDF without opening R directly:
|
|
381
|
+
|
|
382
|
+
```python
|
|
383
|
+
single_boost = GeDSBoostRegressor(max_iterations=10).fit(X[["x"]], y)
|
|
384
|
+
single_boost.save_boosting_diagnostics(
|
|
385
|
+
"boosting.pdf", iterations=[0, 1, 2], final_fits=True
|
|
386
|
+
)
|
|
387
|
+
```
|
|
388
|
+
|
|
389
|
+
The method refuses to replace an existing file unless `overwrite=True`.
|
|
390
|
+
For binomial boosting, R expects responses encoded as -1 and 1. Offset
|
|
391
|
+
prediction is not offered for these additive estimators yet.
|
|
392
|
+
|
|
393
|
+
Fitted estimators contain a serialized R model and can be saved with
|
|
394
|
+
`model.save(path)` and restored with `GeDSRegressor.load(path)`. As with any
|
|
395
|
+
pickle-based format, only load files from trusted sources.
|
|
396
|
+
|
|
397
|
+
## Development
|
|
398
|
+
|
|
399
|
+
Clone the repository, then install the development dependencies and run the
|
|
400
|
+
integration tests with:
|
|
401
|
+
|
|
402
|
+
```console
|
|
403
|
+
git clone https://github.com/emilioluissaenzguillen/GeDS-python.git
|
|
404
|
+
cd GeDS-python
|
|
405
|
+
python -m pip install -e ".[dev]"
|
|
406
|
+
python -m pytest
|
|
407
|
+
python -m build
|
|
408
|
+
```
|
|
409
|
+
|
|
410
|
+
The tests start an embedded R session and therefore require a working GeDS
|
|
411
|
+
installation; they do not substitute or reimplement any GeDS calculations.
|
|
412
|
+
|
|
413
|
+
## Contact
|
|
414
|
+
|
|
415
|
+
For questions about the Python interface, contact Emilio L. Sáenz Guillén at
|
|
416
|
+
[emilioluissaenzguillen@gmail.com](mailto:emilioluissaenzguillen@gmail.com).
|