geds-python 0.1.0a3__tar.gz → 0.1.1__tar.gz

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -0,0 +1,75 @@
1
+ # Changelog
2
+
3
+ All notable changes to GeDS for Python are documented in this file.
4
+
5
+ The project follows [Semantic Versioning](https://semver.org/). Versions use
6
+ the Python packaging form of pre-release identifiers, such as `0.1.0a1` for
7
+ the first alpha release.
8
+
9
+ ## 0.1.1 - 2026-09-28
10
+
11
+ - Let R select `NGeDS()`'s dimension-specific default stopping rule: `RD`
12
+ for univariate and `SR` for joint bivariate or higher-dimensional fits.
13
+ - Add an executed UK spot-rate vignette with a univariate fit and two
14
+ bivariate sparse-maturity reconstructions, linked from the README.
15
+ - Correct CI installation so integration tests use the newly built wheel,
16
+ rather than a same-version distribution from PyPI.
17
+
18
+ ## 0.1.0 - 2026-09-25
19
+
20
+ - Expose R's boosted base-learner importance alongside its existing iteration
21
+ count, and clarify remaining specialized R-only utilities in the feature audit.
22
+ - Add a Gaussian-only Python interface to R's specialized `crossv_GeDS()`
23
+ grid search and its MSE/knot/iteration summaries.
24
+ - Save R's single-learner boosting-iteration visualization as a multipage PDF
25
+ through a Python model method.
26
+ - Expose R's `Derive()`, `Integrate()`, `PPolyRep()`, and `shapeConstrain()`
27
+ through fitted-model methods, with direct R parity tests and explicit model
28
+ restrictions.
29
+ - Check alternate spline orders, generalized-model utilities, mixed-term
30
+ additive constraints, and sequential scikit-learn cross-validation.
31
+ - Add additive GAM and gradient-boosting estimators backed by the R package's
32
+ `NGeDSgam()` and `NGeDSboost()` functions, with independently checked R/Python
33
+ predictions and no duplicate statistical implementation.
34
+ - Support named additive spline terms, joint spline terms, parametric features,
35
+ selected loss families, and individual base-learner predictions.
36
+
37
+ ## 0.1.0a4 - release candidate (not published)
38
+
39
+ - Begin an R-to-Python feature audit for the next coordinated release.
40
+ - Add order-specific access to R's deviance, log likelihood, and coefficient
41
+ confidence intervals.
42
+ - Add univariate fit/prediction offsets and named `predict_terms()` output,
43
+ delegated to the locally fixed R package. These require the corresponding
44
+ R GitHub changes before publication.
45
+ - Expand independent R/Python parity tests for weighted normal and Poisson
46
+ fits, bivariate fits and confidence intervals; clarify joint-spline feature
47
+ selection.
48
+ - Copy arrays extracted from R into Python-owned memory so later R calls cannot
49
+ change stored coefficients, knots, or predictions.
50
+ - Require GeDS 0.3.6 with the Python bridge capability marker, and pin the
51
+ tested GitHub commit in installation and CI instructions.
52
+
53
+ ## 0.1.0a3 - 2026-09-23
54
+
55
+ - Add `python -m geds.check` for human-readable or JSON environment checks.
56
+ - Add `geds.plot_fit()` for Python-native visualization of univariate fits and
57
+ internal knots.
58
+ - Expand installation and R-library troubleshooting guidance.
59
+
60
+ ## 0.1.0a2 - 2026-09-20
61
+
62
+ - Preload R's core numerical DLLs on Windows so embedded R 4.6 can load
63
+ recommended packages without requiring Rtools on the user's `PATH`.
64
+
65
+ ## 0.1.0a1 - 2026-09-20
66
+
67
+ Initial alpha release.
68
+
69
+ - Add scikit-learn-style estimators for normal and generalized GeDS models.
70
+ - Delegate all statistical fitting, prediction, coefficient extraction, and
71
+ knot extraction to the GeDS R package.
72
+ - Support pandas and NumPy inputs, numeric and categorical parametric terms,
73
+ sample weights, model serialization, and environment diagnostics.
74
+ - Support Windows and Linux with R 4.6.1 and Python 3.10 or 3.12 in CI.
75
+ - Add a nonlinear regression example illustrating adaptive knot placement.
@@ -2,8 +2,7 @@ cff-version: 1.2.0
2
2
  message: "If you use GeDS for Python, please cite this software and the GeDS methodology references."
3
3
  title: "GeDS for Python"
4
4
  type: software
5
- version: 0.1.0a3
6
- date-released: 2026-09-23
5
+ version: 0.1.1
7
6
  license: GPL-3.0-only
8
7
  repository-code: "https://github.com/emilioluissaenzguillen/GeDS-python"
9
8
  authors:
@@ -0,0 +1,416 @@
1
+ Metadata-Version: 2.5
2
+ Name: geds-python
3
+ Version: 0.1.1
4
+ Summary: Python estimators backed by the GeDS R package
5
+ Project-URL: Homepage, https://github.com/emilioluissaenzguillen/GeDS-python
6
+ Project-URL: Repository, https://github.com/emilioluissaenzguillen/GeDS-python
7
+ Project-URL: Issues, https://github.com/emilioluissaenzguillen/GeDS-python/issues
8
+ Project-URL: Changelog, https://github.com/emilioluissaenzguillen/GeDS-python/blob/main/CHANGELOG.md
9
+ Project-URL: R package, https://github.com/emilioluissaenzguillen/GeDS
10
+ Author: Dimitrina S. Dimitrova, Vladimir K. Kaishev, Andrea Lattuada, Emilio L. Sáenz Guillén, Richard J. Verrall
11
+ Maintainer-email: "Emilio L. Sáenz Guillén" <emilioluissaenzguillen@gmail.com>
12
+ License-Expression: GPL-3.0-only
13
+ License-File: LICENSE
14
+ Keywords: R,regression,scikit-learn,splines,statistics
15
+ Classifier: Development Status :: 4 - Beta
16
+ Classifier: Intended Audience :: Science/Research
17
+ Classifier: License :: OSI Approved :: GNU General Public License v3 (GPLv3)
18
+ Classifier: Operating System :: Microsoft :: Windows
19
+ Classifier: Operating System :: POSIX :: Linux
20
+ Classifier: Programming Language :: Python :: 3
21
+ Classifier: Programming Language :: Python :: 3.10
22
+ Classifier: Programming Language :: Python :: 3.11
23
+ Classifier: Programming Language :: Python :: 3.12
24
+ Classifier: Topic :: Scientific/Engineering
25
+ Requires-Python: >=3.10
26
+ Requires-Dist: numpy>=1.24
27
+ Requires-Dist: pandas>=2.0
28
+ Requires-Dist: rpy2<3.7,>=3.6.7
29
+ Requires-Dist: scikit-learn>=1.4
30
+ Provides-Extra: dev
31
+ Requires-Dist: build>=1.2; extra == 'dev'
32
+ Requires-Dist: matplotlib>=3.8; extra == 'dev'
33
+ Requires-Dist: pytest>=8; extra == 'dev'
34
+ Provides-Extra: plot
35
+ Requires-Dist: matplotlib>=3.8; extra == 'plot'
36
+ Provides-Extra: test
37
+ Requires-Dist: pytest>=8; extra == 'test'
38
+ Description-Content-Type: text/markdown
39
+
40
+ # GeDS for Python
41
+
42
+ This package provides a Python interface to the
43
+ [GeDS R package](https://github.com/emilioluissaenzguillen/GeDS). The R package
44
+ is the sole implementation of the statistical methods. Python supplies a
45
+ scikit-learn-style API, pandas/NumPy conversion, environment diagnostics, and
46
+ model serialization.
47
+
48
+ ## Requirements
49
+
50
+ - R 4.4 or newer (R 4.6.1 is used for development)
51
+ - GeDS 0.3.6 or newer with the Python bridge fixes
52
+ - Python 3.10 or newer
53
+
54
+ Install the Python package, including the optional plotting dependency used in
55
+ the example:
56
+
57
+ ```console
58
+ python -m pip install "geds-python[plot]"
59
+ ```
60
+
61
+ Install the R package separately, using R 4.6.1 or another supported R
62
+ installation. Install the tested GeDS 0.3.6 build from GitHub:
63
+
64
+ ```r
65
+ install.packages("remotes")
66
+ remotes::install_git(
67
+ "https://github.com/emilioluissaenzguillen/GeDS.git",
68
+ ref = "91b8ddd13aae8f39994c87fc356f05da4799f911",
69
+ dependencies = NA, upgrade = "never"
70
+ )
71
+ ```
72
+
73
+ An older GeDS build, even one reporting version `0.3.6`, may lack the fixes
74
+ required by this wrapper. `install.packages("GeDS")` alone is not guaranteed
75
+ to provide them while the CRAN review follows its separate schedule.
76
+ Check the installed R version with `packageVersion("GeDS")`.
77
+
78
+ On Windows, building the GitHub source package requires Rtools compatible
79
+ with the selected R installation. GeDS remains version `0.3.6` on GitHub;
80
+ the wrapper checks an internal compatibility marker for the fit and prediction
81
+ fixes as well as the package version.
82
+
83
+ The wrapper discovers the newest R installation under `Program Files/R` on
84
+ Windows or uses `Rscript` from `PATH` on other platforms. Set `R_HOME` to select
85
+ a particular R installation. If GeDS is installed in a non-default R library,
86
+ set `GEDS_R_LIBRARY` to that library directory before importing `geds`.
87
+
88
+ The Python and R packages have independent release cycles. `geds-python`
89
+ checks the installed GeDS version when its backend first starts and reports the
90
+ selected R installation and package library through `geds.diagnostics()`.
91
+
92
+ Check the backend before fitting:
93
+
94
+ ```console
95
+ python -m geds.check
96
+ ```
97
+
98
+ For a machine-readable report, use `python -m geds.check --json`. The same
99
+ information is available inside Python:
100
+
101
+ ```python
102
+ import geds
103
+
104
+ print(geds.diagnostics())
105
+ ```
106
+
107
+ ### Selecting R and its package library
108
+
109
+ Usually no configuration is necessary. If several R installations are
110
+ available, select one before starting Python:
111
+
112
+ ```powershell
113
+ $env:R_HOME = "C:\Program Files\R\R-4.6.1"
114
+ python -m geds.check
115
+ ```
116
+
117
+ ```bash
118
+ export R_HOME="/Library/Frameworks/R.framework/Resources" # macOS
119
+ # export R_HOME="/usr/lib/R" # Linux
120
+ python -m geds.check
121
+ ```
122
+
123
+ If GeDS is installed in a personal or otherwise non-default R library, set
124
+ `GEDS_R_LIBRARY` to the directory that contains the `GeDS` folder. You can
125
+ find that directory from R with `find.package("GeDS")`; use its parent
126
+ directory as `GEDS_R_LIBRARY`.
127
+
128
+ If the check reports that R is missing, install R or set `R_HOME`. If it finds
129
+ R but not GeDS, start that same R installation and run
130
+ one of the GeDS installation commands above, then rerun the check.
131
+
132
+ ## Example
133
+
134
+ For a finance example, see the [UK interest-rate notebook](https://github.com/emilioluissaenzguillen/GeDS-python/blob/main/examples/uk_yield_curves.ipynb).
135
+ It fits a 10-year rate over time and a joint time-by-maturity surface using the
136
+ Bank of England's published nominal spot-rate curves. The notebook downloads
137
+ the source data when run; the repository does not redistribute the archive.
138
+
139
+ Install the optional plotting dependency with
140
+ `python -m pip install "geds-python[plot]"`, then fit and visualize a nonlinear
141
+ regression:
142
+
143
+ ```python
144
+ import matplotlib.pyplot as plt
145
+ import numpy as np
146
+ import pandas as pd
147
+
148
+ from geds import GeDSRegressor, plot_fit
149
+
150
+ rng = np.random.RandomState(123)
151
+ n = 500
152
+
153
+
154
+ def f_1(x):
155
+ return (10 * x / (1 + 100 * x**2)) * 4 + 4
156
+
157
+
158
+ x = np.sort(rng.uniform(-2.0, 2.0, size=n))
159
+ means = f_1(x)
160
+ y = rng.normal(means, scale=0.1)
161
+ X = pd.DataFrame({"x": x})
162
+
163
+ model = GeDSRegressor(order=3).fit(X, y)
164
+ knots = np.asarray(model.knots_, dtype=float)
165
+
166
+ print("Internal knots:", knots)
167
+
168
+ fig, ax = plt.subplots()
169
+ plot_fit(model, X, y, ax=ax)
170
+ grid_x = np.linspace(x.min(), x.max(), 500)
171
+ ax.plot(
172
+ grid_x,
173
+ f_1(grid_x),
174
+ color="0.25",
175
+ linestyle=":",
176
+ linewidth=2,
177
+ label="True mean",
178
+ )
179
+ ax.set(ylabel="y")
180
+ ax.legend()
181
+ fig.tight_layout()
182
+ plt.show()
183
+ ```
184
+
185
+ With GeDS 0.3.6 and R 4.6.1, this seeded example fits 16 internal knots.
186
+ The dashed vertical lines show how GeDS places more knots around the sharp
187
+ variation near zero while retaining knots across the wider domain.
188
+
189
+ `GeDSRegressor` delegates to `GeDS::NGeDS()`. For exponential-family models,
190
+ use `GeDSGeneralizedRegressor`, which delegates to `GeDS::GGeDS()`.
191
+ By default, `GeDSRegressor` leaves the stopping rule to R: `RD` for one
192
+ spline feature and `SR` for a joint spline with two or more features. Set
193
+ `stop_type="RD"` or `stop_type="SR"` to override it explicitly.
194
+ `GeDSGeneralizedRegressor` uses R's `SR` default. GAM and boosting base
195
+ learners also use their R implementation's dimension-specific rules.
196
+
197
+ For fitted models, `get_deviance(order=...)`, `get_log_likelihood(order=...)`,
198
+ and `get_confidence_intervals(order=..., level=...)` call the corresponding R
199
+ methods. Confidence intervals are returned as a pandas DataFrame with `lower`
200
+ and `upper` columns. As in R, these are coefficient intervals, not confidence
201
+ bands for the fitted curve.
202
+
203
+ The estimators also work with standard scikit-learn tools such as
204
+ `cross_val_score()` and `GridSearchCV`. Use sequential execution (`n_jobs=1`)
205
+ when cross-validating: the wrapper embeds R in the Python process, and
206
+ parallel-worker behavior is not part of the supported interface.
207
+
208
+ R also has a specialized `crossv_GeDS()` routine, which returns a parameter
209
+ grid with cross-validated mean squared error and knot/iteration summaries.
210
+ Use its Python interface when those R-specific results are needed:
211
+
212
+ ```python
213
+ from geds import cross_validate_geds
214
+
215
+ cv = cross_validate_geds(
216
+ GeDSRegressor(order=3), X, y,
217
+ {"beta": [0.5, 0.7], "phi": [0.95], "q": [2]},
218
+ n_folds=5, n_cores=1, random_state=123,
219
+ )
220
+ print(cv.best_params)
221
+ print(cv.results)
222
+ ```
223
+
224
+ This delegates the entire search to R and does not fit or change the input
225
+ estimator. It currently supports Gaussian models only, accepts the R tuning
226
+ parameters `beta`, `phi`, `q`, and (for boosting) `int_knots_init` and
227
+ `shrinkage`, and defaults to one R worker. R's current non-boost routine does
228
+ not forward other fitting settings; the Python interface rejects custom
229
+ settings it would otherwise silently ignore. Use scikit-learn's grid search
230
+ when you need those settings or a non-Gaussian family.
231
+
232
+ For a fitted univariate spline without extra linear features, R's calculus
233
+ and spline-conversion utilities are available as model methods:
234
+
235
+ ```python
236
+ slopes = model.derive([-0.5, 0.0, 0.5], derivative_order=1)
237
+ areas = model.integrate(-1.0, [-0.5, 0.0, 0.5])
238
+ piece_knots, piece_coefficients = model.piecewise_polynomial()
239
+ ```
240
+
241
+ `derive()` and `integrate()` operate on the predictor (link) scale, as in R.
242
+ `piecewise_polynomial()` returns the R `PPolyRep()` knot vector and coefficient
243
+ matrix; its last coefficient row is extraneous in R's representation. These
244
+ methods use the estimator's selected spline order unless `order=` is given.
245
+
246
+ For a normal univariate fit, impose a shape constraint without changing the
247
+ original fitted model:
248
+
249
+ ```python
250
+ increasing_model = model.shape_constrain("increasing")
251
+ increasing_and_convex = model.shape_constrain(["increasing", "convex"])
252
+ ```
253
+
254
+ This calls R's `shapeConstrain()` and returns a new Python estimator. R also
255
+ supports constraints on one selected univariate smoother in Gaussian GAM and
256
+ boosting fits, via `shape_constrain(..., base_learner="f(x)")`. Those additive
257
+ fits must use `normalize_data=False`. Constrained fits do not provide the usual
258
+ unconstrained coefficient confidence intervals.
259
+
260
+ For count data, the generalized estimator uses `GeDS::GGeDS()` and supports
261
+ both response-scale and link-scale prediction:
262
+
263
+ ```python
264
+ import numpy as np
265
+ import pandas as pd
266
+ from geds import GeDSGeneralizedRegressor, plot_fit
267
+
268
+ rng = np.random.default_rng(123)
269
+ x = np.sort(rng.uniform(-2, 2, 120))
270
+ X = pd.DataFrame({"x": x})
271
+ counts = rng.poisson(np.exp(1 + np.sin(x)))
272
+
273
+ model = GeDSGeneralizedRegressor(
274
+ family="poisson", beta=0.2, phi=0.95, min_internal_knots=3
275
+ ).fit(X, counts)
276
+ mean_counts = model.predict(X)
277
+ log_mean_counts = model.predict_link(X)
278
+ ax = plot_fit(model, X, counts)
279
+ ```
280
+
281
+ `min_internal_knots` controls the minimum number of stage-A knots; it is used
282
+ here to make a small sample's fitted spline visible. GeDS determines the
283
+ final knot positions.
284
+
285
+ For a univariate spline with a known offset (for example log exposure in a
286
+ Poisson model), pass one offset value per observation to both fitting and
287
+ prediction. These values are on the link scale:
288
+
289
+ ```python
290
+ import numpy as np
291
+ import pandas as pd
292
+ from geds import GeDSGeneralizedRegressor
293
+
294
+ rng = np.random.default_rng(321)
295
+ x = np.linspace(-1.5, 1.5, 90)
296
+ X = pd.DataFrame({"x": x})
297
+ exposure = np.linspace(1.1, 2.0, len(x))
298
+ counts = rng.poisson(exposure * np.exp(1.4 + np.sin(x))) + 1
299
+ log_exposure = np.log(exposure)
300
+ model = GeDSGeneralizedRegressor(
301
+ family="poisson", spline_features=["x"], order=2,
302
+ higher_order=False,
303
+ ).fit(X, counts, offset=log_exposure)
304
+ expected_counts = model.predict(X, offset=log_exposure)
305
+ contributions = model.predict_terms(X, offset=log_exposure)
306
+ ```
307
+
308
+ `predict_terms()` returns a DataFrame of the R spline and parametric term
309
+ contributions. Its rows sum to the link prediction after adding the offset;
310
+ the offset is not itself a term column. Offset prediction currently supports
311
+ one spline feature only because the R bivariate prediction method does not
312
+ apply new-data offsets consistently. A model fitted with an offset requires
313
+ an offset at prediction time.
314
+
315
+ This offset interface requires the GeDS GitHub fit and prediction fixes. A
316
+ earlier GeDS `0.3.6` installation without those fixes may mishandle
317
+ generalized-model offsets; the backend rejects that version.
318
+
319
+ Choose spline and parametric components explicitly for mixed data:
320
+
321
+ ```python
322
+ model = GeDSRegressor(
323
+ spline_features=["x"],
324
+ linear_features=["group"],
325
+ ).fit(X, y)
326
+ ```
327
+
328
+ Spline features must be numeric. Parametric features may be numeric or
329
+ categorical; their encoding is performed by the R package so fitting and
330
+ prediction use R's native factor semantics.
331
+ If `spline_features` is omitted, all columns are used in a single joint spline
332
+ term. Select `spline_features=["x"]` and `linear_features=["group"]` to keep
333
+ `group` parametric instead. Two spline features create a joint bivariate
334
+ surface, not two separate additive smooths; R's support for more than two
335
+ spline features is experimental. With named pandas columns, prediction may
336
+ receive columns in a different order because the wrapper restores the fitted
337
+ column order before calling R.
338
+
339
+ ### Additive GAM and boosting models
340
+
341
+ Use `GeDSGAMRegressor` for R's `NGeDSgam()` and `GeDSBoostRegressor` for
342
+ `NGeDSboost()`. Each entry in `spline_terms` is one additive smooth. Put two
343
+ features in the same entry for a joint surface. If omitted, each non-linear
344
+ feature gets its own smooth; `linear_features` selects parametric terms.
345
+
346
+ ```python
347
+ import numpy as np
348
+ import pandas as pd
349
+ from geds import GeDSGAMRegressor, GeDSBoostRegressor
350
+
351
+ x = np.linspace(-2, 2, 100)
352
+ X = pd.DataFrame({"x": x, "z": x**2})
353
+ y = np.sin(x) + 0.3 * x**2
354
+
355
+ gam = GeDSGAMRegressor(
356
+ spline_terms=[("x",), ("z",)], max_iterations=10
357
+ ).fit(X, y)
358
+ boost = GeDSBoostRegressor(
359
+ spline_terms=[("x",), ("z",)], max_iterations=20
360
+ ).fit(X, y)
361
+
362
+ gam_predictions = gam.predict(X)
363
+ boost_predictions = boost.predict(X)
364
+ x_contribution = gam.predict_component(X, "f(x)")
365
+ importance = boost.get_base_learner_importance()
366
+ ```
367
+
368
+ Both estimators expose `predict_link()`, order-specific coefficients, knots,
369
+ deviance, log likelihood, and coefficient confidence intervals through R.
370
+ `predict_component()` delegates a named learner prediction to R. The R
371
+ GAM/boost prediction method does not support `type="terms"`, so these
372
+ estimators do not offer `predict_terms()`. The GAM wrapper supports the
373
+ families accepted by `GeDSGeneralizedRegressor`; for binomial fits it accepts
374
+ 0/1 responses and creates the factor required by R. The boosting
375
+ wrapper maps `gaussian`, `poisson`, `binomial`, and `gamma` to mboost families.
376
+ The fitted boosting estimator's `n_iter_` is R's total boosting iteration
377
+ count. `get_base_learner_importance()` returns R's `bl_imp()` in-bag risk
378
+ reductions as a pandas Series with the original Python feature names.
379
+ For a boosted fit with one univariate spline feature, R's iteration plots can
380
+ be saved to a multipage PDF without opening R directly:
381
+
382
+ ```python
383
+ single_boost = GeDSBoostRegressor(max_iterations=10).fit(X[["x"]], y)
384
+ single_boost.save_boosting_diagnostics(
385
+ "boosting.pdf", iterations=[0, 1, 2], final_fits=True
386
+ )
387
+ ```
388
+
389
+ The method refuses to replace an existing file unless `overwrite=True`.
390
+ For binomial boosting, R expects responses encoded as -1 and 1. Offset
391
+ prediction is not offered for these additive estimators yet.
392
+
393
+ Fitted estimators contain a serialized R model and can be saved with
394
+ `model.save(path)` and restored with `GeDSRegressor.load(path)`. As with any
395
+ pickle-based format, only load files from trusted sources.
396
+
397
+ ## Development
398
+
399
+ Clone the repository, then install the development dependencies and run the
400
+ integration tests with:
401
+
402
+ ```console
403
+ git clone https://github.com/emilioluissaenzguillen/GeDS-python.git
404
+ cd GeDS-python
405
+ python -m pip install -e ".[dev]"
406
+ python -m pytest
407
+ python -m build
408
+ ```
409
+
410
+ The tests start an embedded R session and therefore require a working GeDS
411
+ installation; they do not substitute or reimplement any GeDS calculations.
412
+
413
+ ## Contact
414
+
415
+ For questions about the Python interface, contact Emilio L. Sáenz Guillén at
416
+ [emilioluissaenzguillen@gmail.com](mailto:emilioluissaenzguillen@gmail.com).