multivariate-probit 0.1.0__tar.gz → 0.2.0__tar.gz

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (37) hide show
  1. multivariate_probit-0.2.0/.gitignore +27 -0
  2. {multivariate_probit-0.1.0 → multivariate_probit-0.2.0}/PKG-INFO +13 -5
  3. {multivariate_probit-0.1.0 → multivariate_probit-0.2.0}/README.md +8 -4
  4. {multivariate_probit-0.1.0 → multivariate_probit-0.2.0}/docs/api.md +61 -10
  5. {multivariate_probit-0.1.0 → multivariate_probit-0.2.0}/docs/ifm.md +155 -63
  6. {multivariate_probit-0.1.0 → multivariate_probit-0.2.0}/docs/implementation.md +6 -3
  7. multivariate_probit-0.2.0/docs/limitations.md +160 -0
  8. multivariate_probit-0.2.0/docs/studies/README.md +53 -0
  9. multivariate_probit-0.2.0/docs/studies/comparators.md +180 -0
  10. multivariate_probit-0.2.0/docs/studies/index-correlation-shortcut.md +185 -0
  11. multivariate_probit-0.2.0/docs/studies/log-score-evaluator.md +220 -0
  12. multivariate_probit-0.2.0/docs/studies/margin-gaps.md +106 -0
  13. multivariate_probit-0.2.0/docs/studies/mulan-multilabel.md +148 -0
  14. multivariate_probit-0.2.0/docs/studies/multilabel-benchmarks.md +250 -0
  15. multivariate_probit-0.2.0/docs/studies/sparse-label-transfer.md +276 -0
  16. multivariate_probit-0.2.0/docs/studies/uci-credit.md +123 -0
  17. multivariate_probit-0.2.0/hatch_build.py +24 -0
  18. {multivariate_probit-0.1.0 → multivariate_probit-0.2.0}/pyproject.toml +21 -5
  19. {multivariate_probit-0.1.0 → multivariate_probit-0.2.0}/src/multivariate_probit/__init__.py +4 -2
  20. {multivariate_probit-0.1.0 → multivariate_probit-0.2.0}/src/multivariate_probit/_mvn.py +69 -5
  21. {multivariate_probit-0.1.0 → multivariate_probit-0.2.0}/src/multivariate_probit/ifm.py +18 -4
  22. {multivariate_probit-0.1.0 → multivariate_probit-0.2.0}/src/multivariate_probit/inner.py +46 -16
  23. {multivariate_probit-0.1.0 → multivariate_probit-0.2.0}/src/multivariate_probit/linear.py +3 -2
  24. {multivariate_probit-0.1.0 → multivariate_probit-0.2.0}/src/multivariate_probit/model.py +127 -7
  25. multivariate_probit-0.2.0/src/multivariate_probit/orthant/__init__.py +159 -0
  26. {multivariate_probit-0.1.0 → multivariate_probit-0.2.0}/src/multivariate_probit/results.py +20 -5
  27. multivariate_probit-0.2.0/tests/test_calibration.py +153 -0
  28. multivariate_probit-0.2.0/tests/test_contract.py +121 -0
  29. multivariate_probit-0.2.0/tests/test_evaluators.py +158 -0
  30. {multivariate_probit-0.1.0 → multivariate_probit-0.2.0}/tests/test_linear.py +9 -0
  31. {multivariate_probit-0.1.0 → multivariate_probit-0.2.0}/tests/test_xgboost.py +14 -0
  32. multivariate_probit-0.1.0/.gitignore +0 -18
  33. multivariate_probit-0.1.0/docs/limitations.md +0 -80
  34. {multivariate_probit-0.1.0 → multivariate_probit-0.2.0}/LICENSE +0 -0
  35. {multivariate_probit-0.1.0 → multivariate_probit-0.2.0}/src/multivariate_probit/_corr.py +0 -0
  36. {multivariate_probit-0.1.0 → multivariate_probit-0.2.0}/tests/conftest.py +0 -0
  37. {multivariate_probit-0.1.0 → multivariate_probit-0.2.0}/tests/test_correlation_projection.py +0 -0
@@ -0,0 +1,27 @@
1
+ __pycache__/
2
+ *.py[cod]
3
+ *.egg-info/
4
+ build/
5
+ dist/
6
+ .eggs/
7
+ .pytest_cache/
8
+ .coverage
9
+ htmlcov/
10
+ .venv/
11
+ venv/
12
+ .env
13
+ .idea/
14
+ .vscode/
15
+ .DS_Store
16
+
17
+ # Local data, not redistributed
18
+ UCI/
19
+
20
+ # Compiled extensions, except the bundled orthant package
21
+ *.so
22
+ !src/multivariate_probit/orthant/*.so
23
+
24
+ # Local agent notes, never part of the repo
25
+ CLAUDE.md
26
+ CLAUDE.local.md
27
+ PLAN.md
@@ -1,6 +1,6 @@
1
1
  Metadata-Version: 2.5
2
2
  Name: multivariate-probit
3
- Version: 0.1.0
3
+ Version: 0.2.0
4
4
  Summary: Multivariate probit models fitted by Inference Functions for Margins (IFM), with pluggable inner models.
5
5
  Project-URL: Homepage, https://github.com/sign-of-fourier/multivariate-probit
6
6
  Project-URL: Documentation, https://github.com/sign-of-fourier/multivariate-probit/blob/main/docs/ifm.md
@@ -38,12 +38,16 @@ Requires-Python: >=3.9
38
38
  Requires-Dist: numpy>=1.22
39
39
  Requires-Dist: scipy>=1.8
40
40
  Provides-Extra: all
41
+ Requires-Dist: cryptography>=42; extra == 'all'
41
42
  Requires-Dist: scikit-learn>=1.1; extra == 'all'
42
43
  Requires-Dist: xgboost>=1.7; extra == 'all'
43
44
  Provides-Extra: dev
45
+ Requires-Dist: cryptography>=42; extra == 'dev'
44
46
  Requires-Dist: pytest>=7; extra == 'dev'
45
47
  Requires-Dist: scikit-learn>=1.1; extra == 'dev'
46
48
  Requires-Dist: xgboost>=1.7; extra == 'dev'
49
+ Provides-Extra: orthant
50
+ Requires-Dist: cryptography>=42; extra == 'orthant'
47
51
  Provides-Extra: sklearn
48
52
  Requires-Dist: scikit-learn>=1.1; extra == 'sklearn'
49
53
  Provides-Extra: xgboost
@@ -110,10 +114,12 @@ all; everything joint does.
110
114
 
111
115
  ## Everything here is a squashing function over a latent index
112
116
 
113
- The inner model never sees a probability, and never sees another outcome's
114
- labels. It produces an unbounded score η_j(x) on (-∞, ∞); Φ is the only
115
- squashing function applied to it. Any estimator that emits a real-valued score,
116
- or a probability that can be pushed back through Φ⁻¹, is a legal margin.
117
+ The inner model never sees another outcome's labels. What stage two consumes is
118
+ an unbounded index η_j(x) on (-∞, ∞); Φ is the only squashing function applied
119
+ to it. A legal margin either produces that index directly, or emits a
120
+ probability that is pushed back through Φ⁻¹ — a classifier with `predict_proba`.
121
+ An uncalibrated `decision_function` score is not enough, since its scale is
122
+ arbitrary.
117
123
 
118
124
  That is the whole abstraction, and it is why the inner model is swappable
119
125
  without touching the estimation code.
@@ -162,6 +168,8 @@ point estimate. Known gaps are listed in
162
168
  - **[docs/api.md](docs/api.md)** — parameters, attributes, methods, extension
163
169
  points
164
170
  - **[docs/limitations.md](docs/limitations.md)** — known gaps and roadmap
171
+ - **[docs/studies/](docs/studies/)** — the research archive: what was measured,
172
+ what it settled, and what it left open
165
173
 
166
174
  ## Development
167
175
 
@@ -57,10 +57,12 @@ all; everything joint does.
57
57
 
58
58
  ## Everything here is a squashing function over a latent index
59
59
 
60
- The inner model never sees a probability, and never sees another outcome's
61
- labels. It produces an unbounded score η_j(x) on (-∞, ∞); Φ is the only
62
- squashing function applied to it. Any estimator that emits a real-valued score,
63
- or a probability that can be pushed back through Φ⁻¹, is a legal margin.
60
+ The inner model never sees another outcome's labels. What stage two consumes is
61
+ an unbounded index η_j(x) on (-∞, ∞); Φ is the only squashing function applied
62
+ to it. A legal margin either produces that index directly, or emits a
63
+ probability that is pushed back through Φ⁻¹ — a classifier with `predict_proba`.
64
+ An uncalibrated `decision_function` score is not enough, since its scale is
65
+ arbitrary.
64
66
 
65
67
  That is the whole abstraction, and it is why the inner model is swappable
66
68
  without touching the estimation code.
@@ -109,6 +111,8 @@ point estimate. Known gaps are listed in
109
111
  - **[docs/api.md](docs/api.md)** — parameters, attributes, methods, extension
110
112
  points
111
113
  - **[docs/limitations.md](docs/limitations.md)** — known gaps and roadmap
114
+ - **[docs/studies/](docs/studies/)** — the research archive: what was measured,
115
+ what it settled, and what it left open
112
116
 
113
117
  ## Development
114
118
 
@@ -15,6 +15,8 @@ MultivariateProbit(
15
15
  optimizer="Nelder-Mead",
16
16
  project_correlation=True,
17
17
  random_state=None,
18
+ evaluator="quadrature",
19
+ resolution="high",
18
20
  )
19
21
  ```
20
22
 
@@ -26,16 +28,52 @@ MultivariateProbit(
26
28
  | `inner_params` | `None` | Keyword arguments forwarded to the preset factory. Ignored when `inner` is already an instance. |
27
29
  | `dependence` | `"joint"` | `"joint"` maximises the full d-variate likelihood for Σ; `"pairwise"` maximises each pair's bivariate likelihood (composite likelihood, far cheaper). |
28
30
  | `cv` | `5` | Folds used to cross-fit the latent indices that stage two consumes. `None` skips cross-fitting — see the warning below. |
29
- | `n_quad` | `24` | Gauss-Legendre order for the orthant evaluator. Lower it if fitting with many outcomes gets slow. |
31
+ | `n_quad` | `24` | Gauss-Legendre order for `evaluator="quadrature"`. Lower it if fitting with many outcomes gets slow. |
30
32
  | `optimizer` | `"Nelder-Mead"` | Passed to `scipy.optimize.minimize` for `dependence="joint"`. Derivative-free by design. |
31
33
  | `project_correlation` | `True` | Project a pairwise estimate onto the nearest positive-definite correlation matrix. No effect when `dependence="joint"`. |
32
34
  | `random_state` | `None` | Controls the cross-fitting split and `sample`. |
35
+ | `evaluator` | `"quadrature"` | Backend for the orthant probabilities behind `dependence="joint"` and every joint query: `"quadrature"`, `"scipy"` or `"orthant"`. See [Evaluators](#evaluators). |
36
+ | `resolution` | `"high"` | Passed to `evaluator="orthant"`; ignored otherwise. |
37
+
38
+ > **Reading `calibration_`.** This is the Cox calibration slope, a deliberately
39
+ > simple scale check: a probit of each outcome on its own fitted index. A slope
40
+ > near 1 is expected. It is *not* a calibration assessment — a slope of 1 rules
41
+ > out a first-order scale error and nothing more, and a departure from 1 has
42
+ > several possible causes: miscalibrated probabilities from the classifier, a
43
+ > correctly calibrated but noisy index, or an index scored on the rows it was
44
+ > fitted on. For a real assessment of a classifier's probabilities use a
45
+ > reliability curve; calibrating the margins is the caller's job, not the
46
+ > estimator's, and `CalibratedClassifierCV` is the usual way to do it.
47
+ >
48
+ > Slopes outside roughly [0.5, 2.0] emit a `UserWarning` and never raise. The
49
+ > band is a loose convenience, not a test with a calibrated error rate — a
50
+ > slope inside it is not a clean bill of health, and one outside it is a
51
+ > suggestion to look, not a verdict on the fit. A slope that is not identified
52
+ > — a constant or non-finite index, or a single-class outcome — is `nan` and
53
+ > stays silent. A slope near 1 is not evidence that Σ is trustworthy: the check
54
+ > is blind to omitted signal by construction. See
55
+ > [ifm.md](ifm.md#the-calibration-slope-sees-exactly-one-of-them) and
56
+ > [limitations.md](limitations.md).
33
57
 
34
58
  > **`cv=None` is not a neutral speed-up.** In-sample margins drive every fitted
35
59
  > correlation toward +1. It is defensible for the linear default and reckless
36
60
  > for anything that can overfit. See
37
61
  > [ifm.md](ifm.md#why-cross-fitting-is-required).
38
62
 
63
+ ### Evaluators
64
+
65
+ The speed/accuracy trade-off is the caller's. None of the backends falls back to
66
+ another: one that cannot run as asked raises. All results are clipped to
67
+ [0, 1]. `dependence="pairwise"` needs only bivariate probabilities, which are
68
+ always computed in closed form, so it is unaffected during fitting; joint
69
+ queries on a pairwise fit still use the chosen backend.
70
+
71
+ | `evaluator` | What it is | Trade-off |
72
+ | --- | --- | --- |
73
+ | `"quadrature"` | Genz's recursive conditioning with a fixed Gauss-Legendre rule of order `n_quad` (see [implementation.md](implementation.md)). | Deterministic and accurate. Cost grows as `n_quad ** (d - 2)`, so it becomes impractical well before d = 10. |
74
+ | `"scipy"` | `scipy.stats.multivariate_normal.cdf`, one row at a time. | Reaches any d, but is slow per row and randomised: a `dependence="joint"` fit is not exactly reproducible, and its noisy objective can stall the simplex. |
75
+ | `"orthant"` | An optional compiled package bundled as `multivariate_probit.orthant`. Install with `pip install multivariate-probit[orthant]`. | Built for CPython 3.11 and 3.12 on x86-64 Linux only; elsewhere it raises `ImportError`. Without a key it accepts d ≤ 3 at `resolution="low"` only, and raises `ValueError` outside that. A key is read from `$ORTHANT_KEY` or `~/.orthant/key`; see https://quantecarlo.com/orthant_key. |
76
+
39
77
  ### Attributes
40
78
 
41
79
  | Name | Shape | Meaning |
@@ -43,6 +81,7 @@ MultivariateProbit(
43
81
  | `inner_models_` | list, length d | The fitted margins, one per outcome. |
44
82
  | `correlation_` | (d, d) | The fitted Σ. |
45
83
  | `eta_` | (n, d) | The (cross-fitted) latent indices stage two was fitted on. |
84
+ | `calibration_` | (d, 2) | Intercept and slope of a probit of each outcome on its own fitted index — the calibration slope. See the note below. |
46
85
  | `nll_` | float | Negative log-likelihood at the end of the dependence fit. Joint and pairwise fits optimise different objectives, so the values are not comparable across settings. |
47
86
  | `optimize_result_` | OptimizeResult or None | The SciPy result for `dependence="joint"`. |
48
87
  | `n_outcomes_`, `n_features_in_` | int | |
@@ -51,7 +90,7 @@ MultivariateProbit(
51
90
 
52
91
  | Method | Returns | Notes |
53
92
  | --- | --- | --- |
54
- | `fit(X, Y, sample_weight=None)` | self | `Y` is (n, d) and strictly 0/1. Weights are forwarded to margins that accept them. |
93
+ | `fit(X, Y, sample_weight=None)` | self | `Y` is (n, d) and strictly 0/1. Weights are forwarded to every margin, including the wrapped estimator inside `ProbitCalibrated`; a margin whose `fit` takes no `sample_weight` is fitted unweighted with a warning. |
55
94
  | `decision_function(X)` / `transform(X)` | (n, d) | Latent indices η on (-∞, ∞). |
56
95
  | `fit_transform(X, Y)` | (n, d) | |
57
96
  | `predict_proba(X)` | `MultivariateProbitProba` | See below. |
@@ -89,14 +128,26 @@ one-liner.
89
128
  An inner model is anything with `fit(X, y)` and `latent(X) -> (n,)`, where
90
129
  `latent` returns a real-valued index on the probit scale.
91
130
 
92
- Estimators that do not expose `latent` are wrapped automatically by
93
- `ProbitCalibrated`, which supplies it:
94
-
95
- - if the estimator has `predict_proba`, `latent(X)` is `Φ⁻¹(p̂)`, clipped away
96
- from 0 and 1;
97
- - otherwise, if it has `decision_function`, that score is used as the index
98
- as-is — which assumes it is already probit-scaled. Check that assumption
99
- before relying on it.
131
+ Estimators that do not expose `latent` must expose `predict_proba`, and are
132
+ wrapped automatically by `ProbitCalibrated`, whose `latent(X)` is `Φ⁻¹(p̂)`
133
+ clipped away from 0 and 1. Anything else is a `TypeError` at coercion, before
134
+ any margin is fitted: an uncalibrated `decision_function` score is on an
135
+ arbitrary scale, so `Φ` applied to it is not a probability and no diagnostic
136
+ can detect the mismatch. Wrap such an estimator in a calibrator first
137
+ (`CalibratedClassifierCV`, or `SVC(probability=True)`).
138
+
139
+ The probability itself need not be any good. Inverting the link is exact for a
140
+ calibrated `p̂`, and a miscalibrated one is reported by `calibration_` rather
141
+ than repaired.
142
+
143
+ `predict_proba` is an interface requirement, not a quality bar, and the gap
144
+ between the two can be large. A fully grown `DecisionTreeClassifier` satisfies
145
+ the contract and emits only exact 0s and 1s, so every index is pinned at the
146
+ clip and Σ is fitted from thresholds that carry no information — on a synthetic
147
+ bivariate probit with true ρ = 0.5 it returned 0.914, with calibration slopes
148
+ of 0.11. The same tree at `min_samples_leaf=50` returned 0.362. Neither is
149
+ right, and the contract cannot tell them apart; `calibration_` flagged the
150
+ first and not the second.
100
151
 
101
152
  ### Presets
102
153
 
@@ -15,6 +15,8 @@ score (attribute `eta_`). Σ is the latent correlation matrix — a *correlation
15
15
  matrix, with unit diagonal, not a covariance matrix (attribute
16
16
  `correlation_`). Φ and φ are the standard normal CDF and density; Φ_d is the
17
17
  d-variate normal CDF. `d` counts outcomes, `n` counts observations.
18
+ [What Sigma is conditional on](#what-sigma-is-conditional-on) splits η_j into a
19
+ fitted and an ideal version, and says so locally; nowhere else does.
18
20
 
19
21
  ## What IFM is
20
22
 
@@ -175,9 +177,14 @@ Pairwise estimates carry no such constraint — each ρ_jk is fitted in isolatio
175
177
  so the assembled matrix can fail to be positive definite when d ≥ 3. It is
176
178
  projected onto the nearest correlation matrix by Higham's (2002) alternating
177
179
  projections, with a small eigenvalue floor so the result is strictly positive
178
- definite and safe to Cholesky-factor for sampling. In practice the projection
179
- is a no-op unless the pairwise estimates are genuinely contradictory or the
180
- sample is small.
180
+ definite and safe to Cholesky-factor for sampling.
181
+
182
+ How often this engages depends entirely on the data. On the UCI credit study it
183
+ never did — Σ stayed well conditioned throughout. On multi-label benchmark data
184
+ the smallest eigenvalue hit the floor in 11 of 12 fits, so the projection was
185
+ doing real work almost every time
186
+ ([studies/mulan-multilabel.md](studies/mulan-multilabel.md)). Treat it as an
187
+ active part of the estimator at larger d, not a formality.
181
188
 
182
189
  ## Why cross-fitting is required
183
190
 
@@ -220,7 +227,8 @@ them distinct:
220
227
  - **Cross-fitted margins leave a residual downward bias.** An out-of-fold index
221
228
  still carries genuine prediction error, and noise in a regressor attenuates
222
229
  the correlation measured through it — classical errors-in-variables. In the
223
- table, cross-fit estimates land at 0.376 and 0.340 against a true 0.6.
230
+ table, cross-fit estimates land at 0.376 and 0.340 against a true 0.6. The
231
+ algebra is in the next section, as Gap B.
224
232
 
225
233
  So cross-fitting trades a large upward bias for a smaller downward one. Treat a
226
234
  fitted correlation as a **floor** on the true dependence when the margins are
@@ -228,6 +236,127 @@ only moderately predictive; the gap closes as the margins get sharper. Nothing
228
236
  in the estimator corrects for the second effect, and correcting it would
229
237
  require knowing the margins' prediction error on the latent scale.
230
238
 
239
+ ## What Sigma is conditional on
240
+
241
+ Step 4 takes the fitted index as given and evaluates
242
+ `P(Y_j = 1, Y_k = 1 | x) = Φ_2(η_j, η_k; ρ_jk)`. That is valid only if the
243
+ index it is handed satisfies
244
+
245
+ ```
246
+ P(Y_j = 1 | x) = Φ( η_j(x) ) for all x
247
+ ```
248
+
249
+ Two distinct things separate the fitted index from the truth, and **they are
250
+ not the same kind of defect**. Write η*_j for the index a fully specified model
251
+ would use, η_j for the ideal index given only the features this model sees,
252
+
253
+ ```
254
+ η_j(x) = Φ⁻¹( P(Y_j = 1 | x) )
255
+ ```
256
+
257
+ and η̂_j for the one that was actually fitted. Elsewhere in this file η_j is
258
+ the index the estimator works with — `eta_`, the fitted one. This section is
259
+ the one place the three are separated, because the difference between them is
260
+ the subject.
261
+
262
+ ### Gap A — the estimand moves
263
+
264
+ The features and the functional form may omit determinants of the outcome.
265
+ Projecting the full-information index onto the fitted one,
266
+
267
+ ```
268
+ η*_j = c_j η_j + b_j + w_j, w_j independent of η_j, Var(w_j) = v_j
269
+ ```
270
+
271
+ The omitted part w_j joins the latent error, so `w_j + e_j` has variance
272
+ `1 + v_j`. Dividing through to restore the unit diagonal the model requires,
273
+ the correlation stage two can recover is
274
+
275
+ ```
276
+ ρ_recovered = [ ρ_jk + Cov(w_j, w_k) ] / sqrt( (1 + v_j) (1 + v_k) )
277
+ ```
278
+
279
+ **This is not an estimation error.** A margin calibrated with respect to x
280
+ satisfies the display above by construction, and Σ then measures dependence
281
+ *conditional on x*, with the omitted shared variation correctly counted as part
282
+ of it. That is a well-posed quantity, and often the one a caller wants. It is
283
+ simply not the structural correlation of a fully specified model.
284
+
285
+ The two terms pull opposite ways. `Cov(w_j, w_k) > 0` — margins missing the
286
+ *same* thing — raises ρ_recovered, with no bound and no fixed sign in general;
287
+ `v_j > 0` shrinks whatever the numerator is. The denominator has been checked
288
+ against known v on synthetic draws; the numerator has not
289
+ ([studies/margin-gaps.md](studies/margin-gaps.md)).
290
+
291
+ Nothing can correct this, and no diagnostic can detect it. The marginal law of
292
+ (Y_j, η̂_j) depends on c_j and v_j only through `c_j / sqrt(1 + v_j)`: a margin
293
+ that is correctly scaled but incomplete and one that is over-dispersed but
294
+ complete produce identical marginal data while implying different Σ. The
295
+ information is not in the data being fitted, at any level of flexibility.
296
+
297
+ ### Gap B — a real bias
298
+
299
+ The fitted index also differs from the ideal one by finite-sample estimation
300
+ error and by miscalibration of the classifier. This is **different algebra, not
301
+ the same mechanism at a different stage**. Write
302
+
303
+ ```
304
+ η̂_j = η_j + u_j, Var(u_j) = τ_j², Var(η_j) = σ_j²
305
+ ```
306
+
307
+ Here u_j is *part of* η̂_j — the estimator conditions on the noisy index rather
308
+ than on a clean one perturbed afterwards — so this is classical measurement
309
+ error, not the Berkson form of Gap A. The reliability ratio is
310
+
311
+ ```
312
+ λ_j = σ_j² / (σ_j² + τ_j²) < 1
313
+ ```
314
+
315
+ with `E[η_j | η̂_j] = λ_j η̂_j + (1 - λ_j) m_j` and residual variance
316
+ `λ_j τ_j²`. Substituting as before, and for estimation errors independent
317
+ across margins,
318
+
319
+ ```
320
+ ρ_recovered = ρ_jk / sqrt( (1 + λ_j τ_j²) (1 + λ_k τ_k²) )
321
+ ```
322
+
323
+ This is a genuine downward bias — u_j is no part of the conditional estimand —
324
+ and it is the residual attenuation the previous section describes. It is
325
+ derivable rather than empirical.
326
+
327
+ ### The calibration slope sees exactly one of them
328
+
329
+ After step 3 the cross-fitted η̂ and the labels are both in hand, so a
330
+ probit of Y_j on η̂_j costs one one-dimensional fit per outcome and no extra
331
+ inner-model fits. Its slope is the **calibration slope** of prognostic-model
332
+ validation (Cox), reported as `calibration_`, and what it estimates is
333
+ `λ / sqrt(1 + λ τ²)`. That is the reliability of the fitted index, which is
334
+ precisely the quantity separating the two gaps:
335
+
336
+ | | calibration slope | Σ |
337
+ | --- | --- | --- |
338
+ | Gap A (omitted signal) | exactly 1 | moves, legitimately |
339
+ | Gap B (estimation error) | strictly below 1 | biased down |
340
+
341
+ Two consequences follow, and both matter more than the number itself:
342
+
343
+ - A quiet diagnostic is **not** evidence that Σ is trustworthy. It rules out
344
+ one of the two mechanisms, and not the one that moves Σ furthest. On
345
+ synthetic Gap A draws the slope holds at 1.00 while ρ falls by a factor of
346
+ three ([studies/margin-gaps.md](studies/margin-gaps.md)).
347
+ - A slope far *above* 1 is a third thing again: the index has absorbed the
348
+ noise of the rows it is being scored on, so the labels look more predictable
349
+ than the index admits. That is the memorisation signature cross-fitting
350
+ exists to remove.
351
+
352
+ Neither gap is corrected. Gap A cannot be, for the identifiability reason
353
+ above. Gap B could be, since the slope estimates the very reliability that
354
+ attenuates Σ — but the slope is itself estimated, dividing by it amplifies its
355
+ error into Σ, and the result would be a point estimate with no standard error
356
+ to say how far to trust it. Correcting one gap while the other stays silent
357
+ would make `correlation_` harder to interpret, not easier. Report both, correct
358
+ neither. See [limitations.md](limitations.md).
359
+
231
360
  ## Why not full joint MLE (FIML)
232
361
 
233
362
  The other classical estimator for the same model is FIML: optimize the marginal
@@ -276,6 +405,16 @@ and rejected, each variant for a specific, empirically confirmed reason:
276
405
  raw-Y and the correct value, since removing a genuinely positive
277
406
  shared-predictor contribution pulls the number down further.
278
407
 
408
+ - **Rank-transforming the predicted probabilities first** — empirical CDF per
409
+ margin, then Φ⁻¹, then Pearson. This is a monotone reparameterisation of the
410
+ first bullet and inherits its verdict: measured against arm (b) on the same
411
+ cross-fitted indices, the rank step moves the answer by 0.003 on average and
412
+ never more than 0.011. What the number tracks is the overlap between the
413
+ margins' predictors, not ρ. On synthetic draws with a true ρ of 0 and 60%
414
+ shared index variance it returns 0.58; with a true ρ of 0.5 and disjoint
415
+ predictors it returns 0.02
416
+ ([studies/index-correlation-shortcut.md](studies/index-correlation-shortcut.md)).
417
+
279
418
  All three are linear correlation measures applied to a relationship that is
280
419
  nonlinear by construction (a threshold on a latent Gaussian). Only the
281
420
  maximum-likelihood approach above, working through Φ, correctly inverts that
@@ -288,65 +427,18 @@ support, and nothing constrains the result to be consistent with the marginals.
288
427
 
289
428
  ## Validation on real data
290
429
 
291
- Everything above was developed against synthetic draws, where the true Σ is
292
- known. The claim that matters most in practice — that the composite pairwise
293
- objective is a legitimate substitute for the full likelihood — was checked on
294
- the UCI *Default of Credit Card Clients* data (Yeh & Lien, 2009): 30,000
295
- clients, three binary outcomes (any repayment delay in September, July and
296
- April 2005, i.e. `PAY_0`, `PAY_3`, `PAY_6` > 0), five demographic predictors
297
- (`LIMIT_BAL`, `SEX`, `EDUCATION`, `MARRIAGE`, `AGE`), an 80/20 split and
298
- `cv=5`. Outcome prevalence was 22.8 / 14.0 / 10.2 percent. The dataset is not
299
- redistributed with this library.
300
-
301
- Both modes were run on identical margins — same inner model, same seed, same
302
- folds — so the only thing varying is stage two.
303
-
304
- Fitted correlations (ρ_12, ρ_13, ρ_23):
305
-
306
- | inner | joint | pairwise | max entrywise difference |
307
- | --- | --- | --- | --- |
308
- | linear | 0.669, 0.523, 0.682 | 0.673, 0.535, 0.690 | 0.012 |
309
- | xgboost | 0.655, 0.510, 0.667 | 0.661, 0.524, 0.676 | 0.014 |
310
-
311
- Pairwise came out slightly higher in every entry — a small systematic offset,
312
- not noise.
313
-
314
- Cost of stage two alone, on the same cross-fitted η (30,000 rows, d = 3):
315
-
316
- | inner | pairwise | joint | ratio |
317
- | --- | --- | --- | --- |
318
- | linear | 0.53 s | 60.3 s | 114x |
319
- | xgboost | 0.52 s | 45.6 s | 88x |
320
-
321
- Held-out mean joint log-likelihood:
322
-
323
- | inner | joint | pairwise | difference |
324
- | --- | --- | --- | --- |
325
- | linear | -1.09124 | -1.09131 | 0.00007 nats/row |
326
- | xgboost | -1.08451 | -1.08461 | 0.00010 nats/row |
327
-
328
- Per-row `P(all three late)` differed between modes by 0.001 on average, 0.0026
329
- at worst. Marginal AUC, log-loss and Brier are identical across modes by
330
- construction, since the margins are stage one.
331
-
332
- An independent earlier tetrachoric/Gaussian-copula fit on the same data found
333
- correlations in the 0.51-0.66 range, with AUC 0.72, log-loss 0.173 and Brier
334
- 0.043 for the joint event "late in all three months". Both modes reproduce it:
335
- correlations 0.510-0.655, and for that same joint event AUC 0.7224, log-loss
336
- 0.1827, Brier 0.0458.
337
-
338
- **Conclusion.** At d = 3 with moderate positive correlations, `"pairwise"`
339
- gives up about 1e-4 nats/row of held-out likelihood and saves two orders of
340
- magnitude of fitting time. That is a strong result for the composite objective,
341
- but it is evidence from one regime only: Σ was well conditioned throughout
342
- (smallest eigenvalue 0.27 or above), so the projection step never engaged, and
343
- nothing here speaks to large d, near-singular Σ, or strongly mixed-sign
344
- correlations — the settings where a composite likelihood is most likely to
345
- diverge from the full one.
346
-
347
- Note also what this study does *not* establish. Cross-fitting was used
348
- throughout, so these runs assume the case made above rather than retesting it;
349
- the in-sample bias measurement remains synthetic.
430
+ The composite pairwise objective was checked against the full likelihood on the
431
+ UCI *Default of Credit Card Clients* data: 30,000 clients, three binary
432
+ repayment-delay outcomes, five demographic predictors. At d = 3 with moderate
433
+ positive correlations, `"pairwise"` gives up about 1e-4 nats/row of held-out
434
+ likelihood and saves two orders of magnitude of fitting time, and both modes
435
+ reproduce an independent tetrachoric fit on the same data.
436
+
437
+ That is evidence from one regime only — Σ well conditioned throughout, so the
438
+ projection step never engaged — and it validates the *objective*, not Σ itself,
439
+ since the true Σ is unknown there. Numbers, configuration, the sampling noise
440
+ floor and the full list of what it does not establish are in
441
+ [studies/uci-credit.md](studies/uci-credit.md).
350
442
 
351
443
  ## Computational cost
352
444
 
@@ -46,7 +46,8 @@ Drezner-Wesolowsky form of Plackett's identity, and Φ_d for d ≥ 3 by Genz's
46
46
  recursive conditioning down onto it. Both are vectorised over observations and
47
47
  accept a `(d, d, n)` stack so the correlation matrix may vary by row.
48
48
 
49
- `scipy.stats.multivariate_normal.cdf` was deliberately not used:
49
+ `scipy.stats.multivariate_normal.cdf` is not the default, though it can be
50
+ selected with `evaluator="scipy"` (see [api.md](api.md#evaluators)):
50
51
 
51
52
  - It is quasi-Monte-Carlo, so it is **stochastic**. Noise inside an optimizer
52
53
  objective is corrosive — Nelder-Mead cannot distinguish a genuine likelihood
@@ -56,8 +57,9 @@ accept a `(d, d, n)` stack so the correlation matrix may vary by row.
56
57
  - It is about **1e-5** accurate where the closed-form bivariate case reaches
57
58
  about 1e-12.
58
59
 
59
- SciPy's version does appear in the test suite, as an independent cross-check of
60
- the hand-rolled evaluator.
60
+ Selecting it is the caller's speed/accuracy decision; it is the only route to
61
+ joint probabilities once d is too large for quadrature. It also appears in the
62
+ test suite, as an independent cross-check of the hand-rolled evaluator.
61
63
 
62
64
  ### Cross-fitting
63
65
 
@@ -78,6 +80,7 @@ Optimizers and primitives only:
78
80
 
79
81
  - `scipy.optimize.minimize` (Nelder-Mead) — the joint Σ fit
80
82
  - `scipy.optimize.minimize_scalar` (bounded Brent) — each pairwise ρ
83
+ - `scipy.stats.multivariate_normal.cdf` — Φ_d, only with `evaluator="scipy"`
81
84
  - `scipy.stats.norm` — Φ, φ, Φ⁻¹
82
85
  - `numpy.polynomial.legendre.leggauss` — quadrature nodes and weights
83
86
  - `numpy.linalg` — `solve`, `eigh`, `cholesky`