multivariate-probit 0.1.0__tar.gz → 0.2.0__tar.gz
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- multivariate_probit-0.2.0/.gitignore +27 -0
- {multivariate_probit-0.1.0 → multivariate_probit-0.2.0}/PKG-INFO +13 -5
- {multivariate_probit-0.1.0 → multivariate_probit-0.2.0}/README.md +8 -4
- {multivariate_probit-0.1.0 → multivariate_probit-0.2.0}/docs/api.md +61 -10
- {multivariate_probit-0.1.0 → multivariate_probit-0.2.0}/docs/ifm.md +155 -63
- {multivariate_probit-0.1.0 → multivariate_probit-0.2.0}/docs/implementation.md +6 -3
- multivariate_probit-0.2.0/docs/limitations.md +160 -0
- multivariate_probit-0.2.0/docs/studies/README.md +53 -0
- multivariate_probit-0.2.0/docs/studies/comparators.md +180 -0
- multivariate_probit-0.2.0/docs/studies/index-correlation-shortcut.md +185 -0
- multivariate_probit-0.2.0/docs/studies/log-score-evaluator.md +220 -0
- multivariate_probit-0.2.0/docs/studies/margin-gaps.md +106 -0
- multivariate_probit-0.2.0/docs/studies/mulan-multilabel.md +148 -0
- multivariate_probit-0.2.0/docs/studies/multilabel-benchmarks.md +250 -0
- multivariate_probit-0.2.0/docs/studies/sparse-label-transfer.md +276 -0
- multivariate_probit-0.2.0/docs/studies/uci-credit.md +123 -0
- multivariate_probit-0.2.0/hatch_build.py +24 -0
- {multivariate_probit-0.1.0 → multivariate_probit-0.2.0}/pyproject.toml +21 -5
- {multivariate_probit-0.1.0 → multivariate_probit-0.2.0}/src/multivariate_probit/__init__.py +4 -2
- {multivariate_probit-0.1.0 → multivariate_probit-0.2.0}/src/multivariate_probit/_mvn.py +69 -5
- {multivariate_probit-0.1.0 → multivariate_probit-0.2.0}/src/multivariate_probit/ifm.py +18 -4
- {multivariate_probit-0.1.0 → multivariate_probit-0.2.0}/src/multivariate_probit/inner.py +46 -16
- {multivariate_probit-0.1.0 → multivariate_probit-0.2.0}/src/multivariate_probit/linear.py +3 -2
- {multivariate_probit-0.1.0 → multivariate_probit-0.2.0}/src/multivariate_probit/model.py +127 -7
- multivariate_probit-0.2.0/src/multivariate_probit/orthant/__init__.py +159 -0
- {multivariate_probit-0.1.0 → multivariate_probit-0.2.0}/src/multivariate_probit/results.py +20 -5
- multivariate_probit-0.2.0/tests/test_calibration.py +153 -0
- multivariate_probit-0.2.0/tests/test_contract.py +121 -0
- multivariate_probit-0.2.0/tests/test_evaluators.py +158 -0
- {multivariate_probit-0.1.0 → multivariate_probit-0.2.0}/tests/test_linear.py +9 -0
- {multivariate_probit-0.1.0 → multivariate_probit-0.2.0}/tests/test_xgboost.py +14 -0
- multivariate_probit-0.1.0/.gitignore +0 -18
- multivariate_probit-0.1.0/docs/limitations.md +0 -80
- {multivariate_probit-0.1.0 → multivariate_probit-0.2.0}/LICENSE +0 -0
- {multivariate_probit-0.1.0 → multivariate_probit-0.2.0}/src/multivariate_probit/_corr.py +0 -0
- {multivariate_probit-0.1.0 → multivariate_probit-0.2.0}/tests/conftest.py +0 -0
- {multivariate_probit-0.1.0 → multivariate_probit-0.2.0}/tests/test_correlation_projection.py +0 -0
|
@@ -0,0 +1,27 @@
|
|
|
1
|
+
__pycache__/
|
|
2
|
+
*.py[cod]
|
|
3
|
+
*.egg-info/
|
|
4
|
+
build/
|
|
5
|
+
dist/
|
|
6
|
+
.eggs/
|
|
7
|
+
.pytest_cache/
|
|
8
|
+
.coverage
|
|
9
|
+
htmlcov/
|
|
10
|
+
.venv/
|
|
11
|
+
venv/
|
|
12
|
+
.env
|
|
13
|
+
.idea/
|
|
14
|
+
.vscode/
|
|
15
|
+
.DS_Store
|
|
16
|
+
|
|
17
|
+
# Local data, not redistributed
|
|
18
|
+
UCI/
|
|
19
|
+
|
|
20
|
+
# Compiled extensions, except the bundled orthant package
|
|
21
|
+
*.so
|
|
22
|
+
!src/multivariate_probit/orthant/*.so
|
|
23
|
+
|
|
24
|
+
# Local agent notes, never part of the repo
|
|
25
|
+
CLAUDE.md
|
|
26
|
+
CLAUDE.local.md
|
|
27
|
+
PLAN.md
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
Metadata-Version: 2.5
|
|
2
2
|
Name: multivariate-probit
|
|
3
|
-
Version: 0.
|
|
3
|
+
Version: 0.2.0
|
|
4
4
|
Summary: Multivariate probit models fitted by Inference Functions for Margins (IFM), with pluggable inner models.
|
|
5
5
|
Project-URL: Homepage, https://github.com/sign-of-fourier/multivariate-probit
|
|
6
6
|
Project-URL: Documentation, https://github.com/sign-of-fourier/multivariate-probit/blob/main/docs/ifm.md
|
|
@@ -38,12 +38,16 @@ Requires-Python: >=3.9
|
|
|
38
38
|
Requires-Dist: numpy>=1.22
|
|
39
39
|
Requires-Dist: scipy>=1.8
|
|
40
40
|
Provides-Extra: all
|
|
41
|
+
Requires-Dist: cryptography>=42; extra == 'all'
|
|
41
42
|
Requires-Dist: scikit-learn>=1.1; extra == 'all'
|
|
42
43
|
Requires-Dist: xgboost>=1.7; extra == 'all'
|
|
43
44
|
Provides-Extra: dev
|
|
45
|
+
Requires-Dist: cryptography>=42; extra == 'dev'
|
|
44
46
|
Requires-Dist: pytest>=7; extra == 'dev'
|
|
45
47
|
Requires-Dist: scikit-learn>=1.1; extra == 'dev'
|
|
46
48
|
Requires-Dist: xgboost>=1.7; extra == 'dev'
|
|
49
|
+
Provides-Extra: orthant
|
|
50
|
+
Requires-Dist: cryptography>=42; extra == 'orthant'
|
|
47
51
|
Provides-Extra: sklearn
|
|
48
52
|
Requires-Dist: scikit-learn>=1.1; extra == 'sklearn'
|
|
49
53
|
Provides-Extra: xgboost
|
|
@@ -110,10 +114,12 @@ all; everything joint does.
|
|
|
110
114
|
|
|
111
115
|
## Everything here is a squashing function over a latent index
|
|
112
116
|
|
|
113
|
-
The inner model never sees
|
|
114
|
-
|
|
115
|
-
|
|
116
|
-
|
|
117
|
+
The inner model never sees another outcome's labels. What stage two consumes is
|
|
118
|
+
an unbounded index η_j(x) on (-∞, ∞); Φ is the only squashing function applied
|
|
119
|
+
to it. A legal margin either produces that index directly, or emits a
|
|
120
|
+
probability that is pushed back through Φ⁻¹ — a classifier with `predict_proba`.
|
|
121
|
+
An uncalibrated `decision_function` score is not enough, since its scale is
|
|
122
|
+
arbitrary.
|
|
117
123
|
|
|
118
124
|
That is the whole abstraction, and it is why the inner model is swappable
|
|
119
125
|
without touching the estimation code.
|
|
@@ -162,6 +168,8 @@ point estimate. Known gaps are listed in
|
|
|
162
168
|
- **[docs/api.md](docs/api.md)** — parameters, attributes, methods, extension
|
|
163
169
|
points
|
|
164
170
|
- **[docs/limitations.md](docs/limitations.md)** — known gaps and roadmap
|
|
171
|
+
- **[docs/studies/](docs/studies/)** — the research archive: what was measured,
|
|
172
|
+
what it settled, and what it left open
|
|
165
173
|
|
|
166
174
|
## Development
|
|
167
175
|
|
|
@@ -57,10 +57,12 @@ all; everything joint does.
|
|
|
57
57
|
|
|
58
58
|
## Everything here is a squashing function over a latent index
|
|
59
59
|
|
|
60
|
-
The inner model never sees
|
|
61
|
-
|
|
62
|
-
|
|
63
|
-
|
|
60
|
+
The inner model never sees another outcome's labels. What stage two consumes is
|
|
61
|
+
an unbounded index η_j(x) on (-∞, ∞); Φ is the only squashing function applied
|
|
62
|
+
to it. A legal margin either produces that index directly, or emits a
|
|
63
|
+
probability that is pushed back through Φ⁻¹ — a classifier with `predict_proba`.
|
|
64
|
+
An uncalibrated `decision_function` score is not enough, since its scale is
|
|
65
|
+
arbitrary.
|
|
64
66
|
|
|
65
67
|
That is the whole abstraction, and it is why the inner model is swappable
|
|
66
68
|
without touching the estimation code.
|
|
@@ -109,6 +111,8 @@ point estimate. Known gaps are listed in
|
|
|
109
111
|
- **[docs/api.md](docs/api.md)** — parameters, attributes, methods, extension
|
|
110
112
|
points
|
|
111
113
|
- **[docs/limitations.md](docs/limitations.md)** — known gaps and roadmap
|
|
114
|
+
- **[docs/studies/](docs/studies/)** — the research archive: what was measured,
|
|
115
|
+
what it settled, and what it left open
|
|
112
116
|
|
|
113
117
|
## Development
|
|
114
118
|
|
|
@@ -15,6 +15,8 @@ MultivariateProbit(
|
|
|
15
15
|
optimizer="Nelder-Mead",
|
|
16
16
|
project_correlation=True,
|
|
17
17
|
random_state=None,
|
|
18
|
+
evaluator="quadrature",
|
|
19
|
+
resolution="high",
|
|
18
20
|
)
|
|
19
21
|
```
|
|
20
22
|
|
|
@@ -26,16 +28,52 @@ MultivariateProbit(
|
|
|
26
28
|
| `inner_params` | `None` | Keyword arguments forwarded to the preset factory. Ignored when `inner` is already an instance. |
|
|
27
29
|
| `dependence` | `"joint"` | `"joint"` maximises the full d-variate likelihood for Σ; `"pairwise"` maximises each pair's bivariate likelihood (composite likelihood, far cheaper). |
|
|
28
30
|
| `cv` | `5` | Folds used to cross-fit the latent indices that stage two consumes. `None` skips cross-fitting — see the warning below. |
|
|
29
|
-
| `n_quad` | `24` | Gauss-Legendre order for
|
|
31
|
+
| `n_quad` | `24` | Gauss-Legendre order for `evaluator="quadrature"`. Lower it if fitting with many outcomes gets slow. |
|
|
30
32
|
| `optimizer` | `"Nelder-Mead"` | Passed to `scipy.optimize.minimize` for `dependence="joint"`. Derivative-free by design. |
|
|
31
33
|
| `project_correlation` | `True` | Project a pairwise estimate onto the nearest positive-definite correlation matrix. No effect when `dependence="joint"`. |
|
|
32
34
|
| `random_state` | `None` | Controls the cross-fitting split and `sample`. |
|
|
35
|
+
| `evaluator` | `"quadrature"` | Backend for the orthant probabilities behind `dependence="joint"` and every joint query: `"quadrature"`, `"scipy"` or `"orthant"`. See [Evaluators](#evaluators). |
|
|
36
|
+
| `resolution` | `"high"` | Passed to `evaluator="orthant"`; ignored otherwise. |
|
|
37
|
+
|
|
38
|
+
> **Reading `calibration_`.** This is the Cox calibration slope, a deliberately
|
|
39
|
+
> simple scale check: a probit of each outcome on its own fitted index. A slope
|
|
40
|
+
> near 1 is expected. It is *not* a calibration assessment — a slope of 1 rules
|
|
41
|
+
> out a first-order scale error and nothing more, and a departure from 1 has
|
|
42
|
+
> several possible causes: miscalibrated probabilities from the classifier, a
|
|
43
|
+
> correctly calibrated but noisy index, or an index scored on the rows it was
|
|
44
|
+
> fitted on. For a real assessment of a classifier's probabilities use a
|
|
45
|
+
> reliability curve; calibrating the margins is the caller's job, not the
|
|
46
|
+
> estimator's, and `CalibratedClassifierCV` is the usual way to do it.
|
|
47
|
+
>
|
|
48
|
+
> Slopes outside roughly [0.5, 2.0] emit a `UserWarning` and never raise. The
|
|
49
|
+
> band is a loose convenience, not a test with a calibrated error rate — a
|
|
50
|
+
> slope inside it is not a clean bill of health, and one outside it is a
|
|
51
|
+
> suggestion to look, not a verdict on the fit. A slope that is not identified
|
|
52
|
+
> — a constant or non-finite index, or a single-class outcome — is `nan` and
|
|
53
|
+
> stays silent. A slope near 1 is not evidence that Σ is trustworthy: the check
|
|
54
|
+
> is blind to omitted signal by construction. See
|
|
55
|
+
> [ifm.md](ifm.md#the-calibration-slope-sees-exactly-one-of-them) and
|
|
56
|
+
> [limitations.md](limitations.md).
|
|
33
57
|
|
|
34
58
|
> **`cv=None` is not a neutral speed-up.** In-sample margins drive every fitted
|
|
35
59
|
> correlation toward +1. It is defensible for the linear default and reckless
|
|
36
60
|
> for anything that can overfit. See
|
|
37
61
|
> [ifm.md](ifm.md#why-cross-fitting-is-required).
|
|
38
62
|
|
|
63
|
+
### Evaluators
|
|
64
|
+
|
|
65
|
+
The speed/accuracy trade-off is the caller's. None of the backends falls back to
|
|
66
|
+
another: one that cannot run as asked raises. All results are clipped to
|
|
67
|
+
[0, 1]. `dependence="pairwise"` needs only bivariate probabilities, which are
|
|
68
|
+
always computed in closed form, so it is unaffected during fitting; joint
|
|
69
|
+
queries on a pairwise fit still use the chosen backend.
|
|
70
|
+
|
|
71
|
+
| `evaluator` | What it is | Trade-off |
|
|
72
|
+
| --- | --- | --- |
|
|
73
|
+
| `"quadrature"` | Genz's recursive conditioning with a fixed Gauss-Legendre rule of order `n_quad` (see [implementation.md](implementation.md)). | Deterministic and accurate. Cost grows as `n_quad ** (d - 2)`, so it becomes impractical well before d = 10. |
|
|
74
|
+
| `"scipy"` | `scipy.stats.multivariate_normal.cdf`, one row at a time. | Reaches any d, but is slow per row and randomised: a `dependence="joint"` fit is not exactly reproducible, and its noisy objective can stall the simplex. |
|
|
75
|
+
| `"orthant"` | An optional compiled package bundled as `multivariate_probit.orthant`. Install with `pip install multivariate-probit[orthant]`. | Built for CPython 3.11 and 3.12 on x86-64 Linux only; elsewhere it raises `ImportError`. Without a key it accepts d ≤ 3 at `resolution="low"` only, and raises `ValueError` outside that. A key is read from `$ORTHANT_KEY` or `~/.orthant/key`; see https://quantecarlo.com/orthant_key. |
|
|
76
|
+
|
|
39
77
|
### Attributes
|
|
40
78
|
|
|
41
79
|
| Name | Shape | Meaning |
|
|
@@ -43,6 +81,7 @@ MultivariateProbit(
|
|
|
43
81
|
| `inner_models_` | list, length d | The fitted margins, one per outcome. |
|
|
44
82
|
| `correlation_` | (d, d) | The fitted Σ. |
|
|
45
83
|
| `eta_` | (n, d) | The (cross-fitted) latent indices stage two was fitted on. |
|
|
84
|
+
| `calibration_` | (d, 2) | Intercept and slope of a probit of each outcome on its own fitted index — the calibration slope. See the note below. |
|
|
46
85
|
| `nll_` | float | Negative log-likelihood at the end of the dependence fit. Joint and pairwise fits optimise different objectives, so the values are not comparable across settings. |
|
|
47
86
|
| `optimize_result_` | OptimizeResult or None | The SciPy result for `dependence="joint"`. |
|
|
48
87
|
| `n_outcomes_`, `n_features_in_` | int | |
|
|
@@ -51,7 +90,7 @@ MultivariateProbit(
|
|
|
51
90
|
|
|
52
91
|
| Method | Returns | Notes |
|
|
53
92
|
| --- | --- | --- |
|
|
54
|
-
| `fit(X, Y, sample_weight=None)` | self | `Y` is (n, d) and strictly 0/1. Weights are forwarded to
|
|
93
|
+
| `fit(X, Y, sample_weight=None)` | self | `Y` is (n, d) and strictly 0/1. Weights are forwarded to every margin, including the wrapped estimator inside `ProbitCalibrated`; a margin whose `fit` takes no `sample_weight` is fitted unweighted with a warning. |
|
|
55
94
|
| `decision_function(X)` / `transform(X)` | (n, d) | Latent indices η on (-∞, ∞). |
|
|
56
95
|
| `fit_transform(X, Y)` | (n, d) | |
|
|
57
96
|
| `predict_proba(X)` | `MultivariateProbitProba` | See below. |
|
|
@@ -89,14 +128,26 @@ one-liner.
|
|
|
89
128
|
An inner model is anything with `fit(X, y)` and `latent(X) -> (n,)`, where
|
|
90
129
|
`latent` returns a real-valued index on the probit scale.
|
|
91
130
|
|
|
92
|
-
Estimators that do not expose `latent`
|
|
93
|
-
`ProbitCalibrated`,
|
|
94
|
-
|
|
95
|
-
|
|
96
|
-
|
|
97
|
-
|
|
98
|
-
|
|
99
|
-
|
|
131
|
+
Estimators that do not expose `latent` must expose `predict_proba`, and are
|
|
132
|
+
wrapped automatically by `ProbitCalibrated`, whose `latent(X)` is `Φ⁻¹(p̂)`
|
|
133
|
+
clipped away from 0 and 1. Anything else is a `TypeError` at coercion, before
|
|
134
|
+
any margin is fitted: an uncalibrated `decision_function` score is on an
|
|
135
|
+
arbitrary scale, so `Φ` applied to it is not a probability and no diagnostic
|
|
136
|
+
can detect the mismatch. Wrap such an estimator in a calibrator first
|
|
137
|
+
(`CalibratedClassifierCV`, or `SVC(probability=True)`).
|
|
138
|
+
|
|
139
|
+
The probability itself need not be any good. Inverting the link is exact for a
|
|
140
|
+
calibrated `p̂`, and a miscalibrated one is reported by `calibration_` rather
|
|
141
|
+
than repaired.
|
|
142
|
+
|
|
143
|
+
`predict_proba` is an interface requirement, not a quality bar, and the gap
|
|
144
|
+
between the two can be large. A fully grown `DecisionTreeClassifier` satisfies
|
|
145
|
+
the contract and emits only exact 0s and 1s, so every index is pinned at the
|
|
146
|
+
clip and Σ is fitted from thresholds that carry no information — on a synthetic
|
|
147
|
+
bivariate probit with true ρ = 0.5 it returned 0.914, with calibration slopes
|
|
148
|
+
of 0.11. The same tree at `min_samples_leaf=50` returned 0.362. Neither is
|
|
149
|
+
right, and the contract cannot tell them apart; `calibration_` flagged the
|
|
150
|
+
first and not the second.
|
|
100
151
|
|
|
101
152
|
### Presets
|
|
102
153
|
|
|
@@ -15,6 +15,8 @@ score (attribute `eta_`). Σ is the latent correlation matrix — a *correlation
|
|
|
15
15
|
matrix, with unit diagonal, not a covariance matrix (attribute
|
|
16
16
|
`correlation_`). Φ and φ are the standard normal CDF and density; Φ_d is the
|
|
17
17
|
d-variate normal CDF. `d` counts outcomes, `n` counts observations.
|
|
18
|
+
[What Sigma is conditional on](#what-sigma-is-conditional-on) splits η_j into a
|
|
19
|
+
fitted and an ideal version, and says so locally; nowhere else does.
|
|
18
20
|
|
|
19
21
|
## What IFM is
|
|
20
22
|
|
|
@@ -175,9 +177,14 @@ Pairwise estimates carry no such constraint — each ρ_jk is fitted in isolatio
|
|
|
175
177
|
so the assembled matrix can fail to be positive definite when d ≥ 3. It is
|
|
176
178
|
projected onto the nearest correlation matrix by Higham's (2002) alternating
|
|
177
179
|
projections, with a small eigenvalue floor so the result is strictly positive
|
|
178
|
-
definite and safe to Cholesky-factor for sampling.
|
|
179
|
-
|
|
180
|
-
|
|
180
|
+
definite and safe to Cholesky-factor for sampling.
|
|
181
|
+
|
|
182
|
+
How often this engages depends entirely on the data. On the UCI credit study it
|
|
183
|
+
never did — Σ stayed well conditioned throughout. On multi-label benchmark data
|
|
184
|
+
the smallest eigenvalue hit the floor in 11 of 12 fits, so the projection was
|
|
185
|
+
doing real work almost every time
|
|
186
|
+
([studies/mulan-multilabel.md](studies/mulan-multilabel.md)). Treat it as an
|
|
187
|
+
active part of the estimator at larger d, not a formality.
|
|
181
188
|
|
|
182
189
|
## Why cross-fitting is required
|
|
183
190
|
|
|
@@ -220,7 +227,8 @@ them distinct:
|
|
|
220
227
|
- **Cross-fitted margins leave a residual downward bias.** An out-of-fold index
|
|
221
228
|
still carries genuine prediction error, and noise in a regressor attenuates
|
|
222
229
|
the correlation measured through it — classical errors-in-variables. In the
|
|
223
|
-
table, cross-fit estimates land at 0.376 and 0.340 against a true 0.6.
|
|
230
|
+
table, cross-fit estimates land at 0.376 and 0.340 against a true 0.6. The
|
|
231
|
+
algebra is in the next section, as Gap B.
|
|
224
232
|
|
|
225
233
|
So cross-fitting trades a large upward bias for a smaller downward one. Treat a
|
|
226
234
|
fitted correlation as a **floor** on the true dependence when the margins are
|
|
@@ -228,6 +236,127 @@ only moderately predictive; the gap closes as the margins get sharper. Nothing
|
|
|
228
236
|
in the estimator corrects for the second effect, and correcting it would
|
|
229
237
|
require knowing the margins' prediction error on the latent scale.
|
|
230
238
|
|
|
239
|
+
## What Sigma is conditional on
|
|
240
|
+
|
|
241
|
+
Step 4 takes the fitted index as given and evaluates
|
|
242
|
+
`P(Y_j = 1, Y_k = 1 | x) = Φ_2(η_j, η_k; ρ_jk)`. That is valid only if the
|
|
243
|
+
index it is handed satisfies
|
|
244
|
+
|
|
245
|
+
```
|
|
246
|
+
P(Y_j = 1 | x) = Φ( η_j(x) ) for all x
|
|
247
|
+
```
|
|
248
|
+
|
|
249
|
+
Two distinct things separate the fitted index from the truth, and **they are
|
|
250
|
+
not the same kind of defect**. Write η*_j for the index a fully specified model
|
|
251
|
+
would use, η_j for the ideal index given only the features this model sees,
|
|
252
|
+
|
|
253
|
+
```
|
|
254
|
+
η_j(x) = Φ⁻¹( P(Y_j = 1 | x) )
|
|
255
|
+
```
|
|
256
|
+
|
|
257
|
+
and η̂_j for the one that was actually fitted. Elsewhere in this file η_j is
|
|
258
|
+
the index the estimator works with — `eta_`, the fitted one. This section is
|
|
259
|
+
the one place the three are separated, because the difference between them is
|
|
260
|
+
the subject.
|
|
261
|
+
|
|
262
|
+
### Gap A — the estimand moves
|
|
263
|
+
|
|
264
|
+
The features and the functional form may omit determinants of the outcome.
|
|
265
|
+
Projecting the full-information index onto the fitted one,
|
|
266
|
+
|
|
267
|
+
```
|
|
268
|
+
η*_j = c_j η_j + b_j + w_j, w_j independent of η_j, Var(w_j) = v_j
|
|
269
|
+
```
|
|
270
|
+
|
|
271
|
+
The omitted part w_j joins the latent error, so `w_j + e_j` has variance
|
|
272
|
+
`1 + v_j`. Dividing through to restore the unit diagonal the model requires,
|
|
273
|
+
the correlation stage two can recover is
|
|
274
|
+
|
|
275
|
+
```
|
|
276
|
+
ρ_recovered = [ ρ_jk + Cov(w_j, w_k) ] / sqrt( (1 + v_j) (1 + v_k) )
|
|
277
|
+
```
|
|
278
|
+
|
|
279
|
+
**This is not an estimation error.** A margin calibrated with respect to x
|
|
280
|
+
satisfies the display above by construction, and Σ then measures dependence
|
|
281
|
+
*conditional on x*, with the omitted shared variation correctly counted as part
|
|
282
|
+
of it. That is a well-posed quantity, and often the one a caller wants. It is
|
|
283
|
+
simply not the structural correlation of a fully specified model.
|
|
284
|
+
|
|
285
|
+
The two terms pull opposite ways. `Cov(w_j, w_k) > 0` — margins missing the
|
|
286
|
+
*same* thing — raises ρ_recovered, with no bound and no fixed sign in general;
|
|
287
|
+
`v_j > 0` shrinks whatever the numerator is. The denominator has been checked
|
|
288
|
+
against known v on synthetic draws; the numerator has not
|
|
289
|
+
([studies/margin-gaps.md](studies/margin-gaps.md)).
|
|
290
|
+
|
|
291
|
+
Nothing can correct this, and no diagnostic can detect it. The marginal law of
|
|
292
|
+
(Y_j, η̂_j) depends on c_j and v_j only through `c_j / sqrt(1 + v_j)`: a margin
|
|
293
|
+
that is correctly scaled but incomplete and one that is over-dispersed but
|
|
294
|
+
complete produce identical marginal data while implying different Σ. The
|
|
295
|
+
information is not in the data being fitted, at any level of flexibility.
|
|
296
|
+
|
|
297
|
+
### Gap B — a real bias
|
|
298
|
+
|
|
299
|
+
The fitted index also differs from the ideal one by finite-sample estimation
|
|
300
|
+
error and by miscalibration of the classifier. This is **different algebra, not
|
|
301
|
+
the same mechanism at a different stage**. Write
|
|
302
|
+
|
|
303
|
+
```
|
|
304
|
+
η̂_j = η_j + u_j, Var(u_j) = τ_j², Var(η_j) = σ_j²
|
|
305
|
+
```
|
|
306
|
+
|
|
307
|
+
Here u_j is *part of* η̂_j — the estimator conditions on the noisy index rather
|
|
308
|
+
than on a clean one perturbed afterwards — so this is classical measurement
|
|
309
|
+
error, not the Berkson form of Gap A. The reliability ratio is
|
|
310
|
+
|
|
311
|
+
```
|
|
312
|
+
λ_j = σ_j² / (σ_j² + τ_j²) < 1
|
|
313
|
+
```
|
|
314
|
+
|
|
315
|
+
with `E[η_j | η̂_j] = λ_j η̂_j + (1 - λ_j) m_j` and residual variance
|
|
316
|
+
`λ_j τ_j²`. Substituting as before, and for estimation errors independent
|
|
317
|
+
across margins,
|
|
318
|
+
|
|
319
|
+
```
|
|
320
|
+
ρ_recovered = ρ_jk / sqrt( (1 + λ_j τ_j²) (1 + λ_k τ_k²) )
|
|
321
|
+
```
|
|
322
|
+
|
|
323
|
+
This is a genuine downward bias — u_j is no part of the conditional estimand —
|
|
324
|
+
and it is the residual attenuation the previous section describes. It is
|
|
325
|
+
derivable rather than empirical.
|
|
326
|
+
|
|
327
|
+
### The calibration slope sees exactly one of them
|
|
328
|
+
|
|
329
|
+
After step 3 the cross-fitted η̂ and the labels are both in hand, so a
|
|
330
|
+
probit of Y_j on η̂_j costs one one-dimensional fit per outcome and no extra
|
|
331
|
+
inner-model fits. Its slope is the **calibration slope** of prognostic-model
|
|
332
|
+
validation (Cox), reported as `calibration_`, and what it estimates is
|
|
333
|
+
`λ / sqrt(1 + λ τ²)`. That is the reliability of the fitted index, which is
|
|
334
|
+
precisely the quantity separating the two gaps:
|
|
335
|
+
|
|
336
|
+
| | calibration slope | Σ |
|
|
337
|
+
| --- | --- | --- |
|
|
338
|
+
| Gap A (omitted signal) | exactly 1 | moves, legitimately |
|
|
339
|
+
| Gap B (estimation error) | strictly below 1 | biased down |
|
|
340
|
+
|
|
341
|
+
Two consequences follow, and both matter more than the number itself:
|
|
342
|
+
|
|
343
|
+
- A quiet diagnostic is **not** evidence that Σ is trustworthy. It rules out
|
|
344
|
+
one of the two mechanisms, and not the one that moves Σ furthest. On
|
|
345
|
+
synthetic Gap A draws the slope holds at 1.00 while ρ falls by a factor of
|
|
346
|
+
three ([studies/margin-gaps.md](studies/margin-gaps.md)).
|
|
347
|
+
- A slope far *above* 1 is a third thing again: the index has absorbed the
|
|
348
|
+
noise of the rows it is being scored on, so the labels look more predictable
|
|
349
|
+
than the index admits. That is the memorisation signature cross-fitting
|
|
350
|
+
exists to remove.
|
|
351
|
+
|
|
352
|
+
Neither gap is corrected. Gap A cannot be, for the identifiability reason
|
|
353
|
+
above. Gap B could be, since the slope estimates the very reliability that
|
|
354
|
+
attenuates Σ — but the slope is itself estimated, dividing by it amplifies its
|
|
355
|
+
error into Σ, and the result would be a point estimate with no standard error
|
|
356
|
+
to say how far to trust it. Correcting one gap while the other stays silent
|
|
357
|
+
would make `correlation_` harder to interpret, not easier. Report both, correct
|
|
358
|
+
neither. See [limitations.md](limitations.md).
|
|
359
|
+
|
|
231
360
|
## Why not full joint MLE (FIML)
|
|
232
361
|
|
|
233
362
|
The other classical estimator for the same model is FIML: optimize the marginal
|
|
@@ -276,6 +405,16 @@ and rejected, each variant for a specific, empirically confirmed reason:
|
|
|
276
405
|
raw-Y and the correct value, since removing a genuinely positive
|
|
277
406
|
shared-predictor contribution pulls the number down further.
|
|
278
407
|
|
|
408
|
+
- **Rank-transforming the predicted probabilities first** — empirical CDF per
|
|
409
|
+
margin, then Φ⁻¹, then Pearson. This is a monotone reparameterisation of the
|
|
410
|
+
first bullet and inherits its verdict: measured against arm (b) on the same
|
|
411
|
+
cross-fitted indices, the rank step moves the answer by 0.003 on average and
|
|
412
|
+
never more than 0.011. What the number tracks is the overlap between the
|
|
413
|
+
margins' predictors, not ρ. On synthetic draws with a true ρ of 0 and 60%
|
|
414
|
+
shared index variance it returns 0.58; with a true ρ of 0.5 and disjoint
|
|
415
|
+
predictors it returns 0.02
|
|
416
|
+
([studies/index-correlation-shortcut.md](studies/index-correlation-shortcut.md)).
|
|
417
|
+
|
|
279
418
|
All three are linear correlation measures applied to a relationship that is
|
|
280
419
|
nonlinear by construction (a threshold on a latent Gaussian). Only the
|
|
281
420
|
maximum-likelihood approach above, working through Φ, correctly inverts that
|
|
@@ -288,65 +427,18 @@ support, and nothing constrains the result to be consistent with the marginals.
|
|
|
288
427
|
|
|
289
428
|
## Validation on real data
|
|
290
429
|
|
|
291
|
-
|
|
292
|
-
|
|
293
|
-
|
|
294
|
-
|
|
295
|
-
|
|
296
|
-
|
|
297
|
-
|
|
298
|
-
|
|
299
|
-
|
|
300
|
-
|
|
301
|
-
|
|
302
|
-
|
|
303
|
-
|
|
304
|
-
Fitted correlations (ρ_12, ρ_13, ρ_23):
|
|
305
|
-
|
|
306
|
-
| inner | joint | pairwise | max entrywise difference |
|
|
307
|
-
| --- | --- | --- | --- |
|
|
308
|
-
| linear | 0.669, 0.523, 0.682 | 0.673, 0.535, 0.690 | 0.012 |
|
|
309
|
-
| xgboost | 0.655, 0.510, 0.667 | 0.661, 0.524, 0.676 | 0.014 |
|
|
310
|
-
|
|
311
|
-
Pairwise came out slightly higher in every entry — a small systematic offset,
|
|
312
|
-
not noise.
|
|
313
|
-
|
|
314
|
-
Cost of stage two alone, on the same cross-fitted η (30,000 rows, d = 3):
|
|
315
|
-
|
|
316
|
-
| inner | pairwise | joint | ratio |
|
|
317
|
-
| --- | --- | --- | --- |
|
|
318
|
-
| linear | 0.53 s | 60.3 s | 114x |
|
|
319
|
-
| xgboost | 0.52 s | 45.6 s | 88x |
|
|
320
|
-
|
|
321
|
-
Held-out mean joint log-likelihood:
|
|
322
|
-
|
|
323
|
-
| inner | joint | pairwise | difference |
|
|
324
|
-
| --- | --- | --- | --- |
|
|
325
|
-
| linear | -1.09124 | -1.09131 | 0.00007 nats/row |
|
|
326
|
-
| xgboost | -1.08451 | -1.08461 | 0.00010 nats/row |
|
|
327
|
-
|
|
328
|
-
Per-row `P(all three late)` differed between modes by 0.001 on average, 0.0026
|
|
329
|
-
at worst. Marginal AUC, log-loss and Brier are identical across modes by
|
|
330
|
-
construction, since the margins are stage one.
|
|
331
|
-
|
|
332
|
-
An independent earlier tetrachoric/Gaussian-copula fit on the same data found
|
|
333
|
-
correlations in the 0.51-0.66 range, with AUC 0.72, log-loss 0.173 and Brier
|
|
334
|
-
0.043 for the joint event "late in all three months". Both modes reproduce it:
|
|
335
|
-
correlations 0.510-0.655, and for that same joint event AUC 0.7224, log-loss
|
|
336
|
-
0.1827, Brier 0.0458.
|
|
337
|
-
|
|
338
|
-
**Conclusion.** At d = 3 with moderate positive correlations, `"pairwise"`
|
|
339
|
-
gives up about 1e-4 nats/row of held-out likelihood and saves two orders of
|
|
340
|
-
magnitude of fitting time. That is a strong result for the composite objective,
|
|
341
|
-
but it is evidence from one regime only: Σ was well conditioned throughout
|
|
342
|
-
(smallest eigenvalue 0.27 or above), so the projection step never engaged, and
|
|
343
|
-
nothing here speaks to large d, near-singular Σ, or strongly mixed-sign
|
|
344
|
-
correlations — the settings where a composite likelihood is most likely to
|
|
345
|
-
diverge from the full one.
|
|
346
|
-
|
|
347
|
-
Note also what this study does *not* establish. Cross-fitting was used
|
|
348
|
-
throughout, so these runs assume the case made above rather than retesting it;
|
|
349
|
-
the in-sample bias measurement remains synthetic.
|
|
430
|
+
The composite pairwise objective was checked against the full likelihood on the
|
|
431
|
+
UCI *Default of Credit Card Clients* data: 30,000 clients, three binary
|
|
432
|
+
repayment-delay outcomes, five demographic predictors. At d = 3 with moderate
|
|
433
|
+
positive correlations, `"pairwise"` gives up about 1e-4 nats/row of held-out
|
|
434
|
+
likelihood and saves two orders of magnitude of fitting time, and both modes
|
|
435
|
+
reproduce an independent tetrachoric fit on the same data.
|
|
436
|
+
|
|
437
|
+
That is evidence from one regime only — Σ well conditioned throughout, so the
|
|
438
|
+
projection step never engaged — and it validates the *objective*, not Σ itself,
|
|
439
|
+
since the true Σ is unknown there. Numbers, configuration, the sampling noise
|
|
440
|
+
floor and the full list of what it does not establish are in
|
|
441
|
+
[studies/uci-credit.md](studies/uci-credit.md).
|
|
350
442
|
|
|
351
443
|
## Computational cost
|
|
352
444
|
|
|
@@ -46,7 +46,8 @@ Drezner-Wesolowsky form of Plackett's identity, and Φ_d for d ≥ 3 by Genz's
|
|
|
46
46
|
recursive conditioning down onto it. Both are vectorised over observations and
|
|
47
47
|
accept a `(d, d, n)` stack so the correlation matrix may vary by row.
|
|
48
48
|
|
|
49
|
-
`scipy.stats.multivariate_normal.cdf`
|
|
49
|
+
`scipy.stats.multivariate_normal.cdf` is not the default, though it can be
|
|
50
|
+
selected with `evaluator="scipy"` (see [api.md](api.md#evaluators)):
|
|
50
51
|
|
|
51
52
|
- It is quasi-Monte-Carlo, so it is **stochastic**. Noise inside an optimizer
|
|
52
53
|
objective is corrosive — Nelder-Mead cannot distinguish a genuine likelihood
|
|
@@ -56,8 +57,9 @@ accept a `(d, d, n)` stack so the correlation matrix may vary by row.
|
|
|
56
57
|
- It is about **1e-5** accurate where the closed-form bivariate case reaches
|
|
57
58
|
about 1e-12.
|
|
58
59
|
|
|
59
|
-
|
|
60
|
-
|
|
60
|
+
Selecting it is the caller's speed/accuracy decision; it is the only route to
|
|
61
|
+
joint probabilities once d is too large for quadrature. It also appears in the
|
|
62
|
+
test suite, as an independent cross-check of the hand-rolled evaluator.
|
|
61
63
|
|
|
62
64
|
### Cross-fitting
|
|
63
65
|
|
|
@@ -78,6 +80,7 @@ Optimizers and primitives only:
|
|
|
78
80
|
|
|
79
81
|
- `scipy.optimize.minimize` (Nelder-Mead) — the joint Σ fit
|
|
80
82
|
- `scipy.optimize.minimize_scalar` (bounded Brent) — each pairwise ρ
|
|
83
|
+
- `scipy.stats.multivariate_normal.cdf` — Φ_d, only with `evaluator="scipy"`
|
|
81
84
|
- `scipy.stats.norm` — Φ, φ, Φ⁻¹
|
|
82
85
|
- `numpy.polynomial.legendre.leggauss` — quadrature nodes and weights
|
|
83
86
|
- `numpy.linalg` — `solve`, `eigh`, `cholesky`
|