squeeze-kernel 0.6.0__tar.gz → 0.7.1__tar.gz

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -1,6 +1,6 @@
1
1
  Metadata-Version: 2.4
2
2
  Name: squeeze-kernel
3
- Version: 0.6.0
3
+ Version: 0.7.1
4
4
  Summary: Streaming, PSD-by-construction covariance estimator with Fisher-kernel weighting and adaptive shrinkage
5
5
  Keywords: covariance,correlation,ewma,kernel,risk,streaming
6
6
  Author: Robert Kende
@@ -19,6 +19,7 @@ Requires-Dist: numpy>=1.24
19
19
  Requires-Dist: pytest>=7.0 ; extra == 'dev'
20
20
  Requires-Dist: ruff>=0.7 ; extra == 'dev'
21
21
  Requires-Dist: scipy>=1.11 ; extra == 'dev'
22
+ Requires-Dist: mypy>=1.10 ; extra == 'dev'
22
23
  Requires-Dist: scipy>=1.11 ; extra == 'full'
23
24
  Requires-Python: >=3.10
24
25
  Project-URL: Homepage, https://github.com/r0k3/squeeze-kernel
@@ -35,13 +36,13 @@ Description-Content-Type: text/markdown
35
36
  [![Python](https://img.shields.io/pypi/pyversions/squeeze-kernel.svg)](https://pypi.org/project/squeeze-kernel/)
36
37
  [![License: MIT](https://img.shields.io/badge/license-MIT-green.svg)](LICENSE)
37
38
 
38
- A **streaming covariance estimator for panels of daily financial returns**. One `O(n²)` update per day, positive semi-definite **by construction** at every step, missing values handled **natively**, and defaults that require no tuning. Only dependency: NumPy.
39
+ A **streaming covariance estimator for panels of financial returns** that learns fastest on the days that matter. One `O(n²)` update per period, positive semi-definite **by construction** at every step, missing values handled **natively**, and defaults that require no tuning. Only dependency: NumPy.
39
40
 
40
- Reference: *"The Squeeze Kernel Covariance Estimator: Dual-Timescale Tracking with Adaptive Shrinkage"* (Kende, 2026) — [SSRN abstract 6455918](https://ssrn.com/abstract=6455918).
41
+ References: *"The Squeeze Kernel Covariance Estimator: Dual-Timescale Tracking with Adaptive Shrinkage"* (Kende, 2026) — [SSRN abstract 6455918](https://ssrn.com/abstract=6455918) — and its companion *"Cluster-Respecting Shrinkage for Streaming Covariance Estimation"* (Kende, 2026).
41
42
 
42
43
  ## Why
43
44
 
44
- Rolling-window estimators (Ledoit–Wolf, nonlinear shrinkage, RMT denoising) refit over a fixed window each day and cannot adapt within it; multivariate GARCH (DCC) adapts but needs multi-stage estimation and a fragile news coefficient. The Squeeze Kernel is a single streaming recursion that:
45
+ Every standard covariance estimator treats all trading days as equally informative. Markets don't work that way: **correlations reveal themselves when markets move; calm days are mostly noise.** The Squeeze Kernel weighs each day by the information it actually carries — quiet days barely count, dispersion shocks pass through in full — and runs volatility and correlation on separate clocks, so vol spikes never contaminate the correlation estimate. The result is a single streaming recursion that:
45
46
 
46
47
  - **is PSD at every step, structurally** — never needs eigenvalue clipping, nearest-PSD projection, or a solver;
47
48
  - **adapts fastest exactly when it matters** — a Fisher-information kernel up-weights high-dispersion (stress) days, when correlation regimes actually move;
@@ -49,7 +50,37 @@ Rolling-window estimators (Ledoit–Wolf, nonlinear shrinkage, RMT denoising) re
49
50
  - **ingests missing values natively** — listings, delistings, and halts enter as `NaN`; no imputation or complete-case subsetting;
50
51
  - **is fast** — a full 30-year daily pass takes ~0.75 s at n=100 and ~3.4 s at n=300 (single-threaded), 30–40× faster than rolling-window baselines at scale.
51
52
 
52
- On a 30-year S&P 500 panel (n=100, ~7,600 out-of-sample days) it statistically ties DCC on one-step density forecasts and beats EWMA, Ledoit–Wolf, OAS, nonlinear shrinkage, RMT denoising, and the Gerber statistic — and it is the only method in the 90% model confidence set together with DCC. At n=300 it leads every competitor that remains statistically viable.
53
+ **The scoreboard.** On a 30-year S&P 500 panel (~7,600 out-of-sample days) the default single-scale estimator beats EWMA, Ledoit–Wolf, OAS, nonlinear shrinkage, RMT denoising, and the Gerber statistic on one-step density forecasts, and statistically ties DCC — the only two methods in the 90% model confidence set. The headline configuration (multi-scale correlation memory, `corr_half_lives=(43, 173, 693)`) goes further: it **leads every tested method at every universe size and the 90% model confidence set collapses to it alone**, with the margin confirmed out-of-time on an external industry panel. For large equity universes, the cluster shrinkage target (`shrinkage_target="cluster"`) adds a further large gain exactly where shrinkage binds: 25 NLL points at n=300.
54
+
55
+ Against the alternatives:
56
+
57
+ - **RiskMetrics / EWMA** — same one-recursion simplicity, but two clocks and information weighting: better forecasts at zero extra operational cost.
58
+ - **Ledoit–Wolf, nonlinear shrinkage, RMT denoising** — static snapshots refit from scratch on a rolling window each day; the Squeeze Kernel is genuinely dynamic, more accurate on the benchmark, and 30–40× faster at scale.
59
+ - **DCC-GARCH** — matched (single-scale) or beaten (multi-scale) without multi-stage likelihood fitting or the fragile news-impact coefficient; at n=300 DCC needs a multi-year warm-up before its forecasts stabilise and still trails by ~25 NLL.
60
+ - **Gerber statistic** — the Squeeze Kernel is a PSD-by-construction generalization of the same robust-comovement idea: no nearest-PSD repair step, and it wins the head-to-head.
61
+
62
+ ## See the difference
63
+
64
+ Take a passive strategy any allocator would recognize: a long-only minimum-variance portfolio of 300 liquid US stocks, scaled to a 15% volatility target, rebalanced once a month, with 5 bps trading costs. Run it twice on identical data. The only thing that changes between the two runs is the covariance matrix that picks the weights and sets the exposure.
65
+
66
+ <picture>
67
+ <source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/r0k3/squeeze-kernel/main/examples/figures/vol_targeted_portfolio_dark.png">
68
+ <img alt="Vol-targeted long-only minimum-variance portfolio on 300 US equities: Squeeze Kernel vs Ledoit-Wolf equity curves, drawdown, realized volatility, and risk/return profile" src="https://raw.githubusercontent.com/r0k3/squeeze-kernel/main/examples/figures/vol_targeted_portfolio.png">
69
+ </picture>
70
+
71
+ | Method | CAGR | Vol | Sharpe | MaxDD | Calmar | Vol-target RMSE |
72
+ |---|---|---|---|---|---|---|
73
+ | **Squeeze Kernel** | **13.1%** | **13.0%** | **1.00** | **-31.2%** | **0.42** | **6.47%** |
74
+ | Ledoit-Wolf (252d) | 11.7% | 14.6% | 0.80 | -38.7% | 0.30 | 7.51% |
75
+
76
+ You can regenerate this example from the repo alone. The returns panel ships as a parquet (daily returns for 300 US stocks, sourced from Yahoo Finance), and the script prints the table and redraws the figure:
77
+
78
+ ```bash
79
+ pip install squeeze-kernel pandas pyarrow scikit-learn matplotlib
80
+ python examples/vol_targeted_portfolio.py # ~2 minutes
81
+ ```
82
+
83
+ The full protocol and data notes live in [`examples/vol_targeted_portfolio.py`](examples/vol_targeted_portfolio.py) and [`examples/data/build_equity_panel.py`](examples/data/build_equity_panel.py).
53
84
 
54
85
  ## Installation
55
86
 
@@ -116,44 +147,27 @@ kappa = SqueezeKernelEstimator.calibrate_kappa(burn_in_returns, target_weight=0.
116
147
 
117
148
  ## Advanced options
118
149
 
119
- **Adaptive scale-free correlation memory** (`corr_half_lives=(43, 173, 693)`, `corr_theta=0.25`): replaces the single correlation timescale with a positive combination of EWMAs on a geometric half-life ladder — each scale normalized and adaptively shrunk against its own effective sample size, then the covariances blended with weights resting at the prior ∝ half-life^`corr_theta`. By Bernstein's theorem the ladder approximates the power-law memory of financial correlations (the streaming analogue of HAR). For ladders of two or more rungs the blend weights are gated by a sequential surprise detector: a two-sided Page CUSUM on the studentized fast-vs-slow per-rung predictive-score drift (threshold set by Siegmund's average-run-length approximation at ~2 years, no tuned parameters) tilts the weights toward the fast or slow end of the ladder when one side accumulates statistically forced evidence, decaying back at the fastest rung's half-life. Weights equal the prior on all non-alarmed days, the blend stays convex, so PSD holds by construction; `None` (default) reproduces the published single-scale estimator exactly. This is the paper's **headline configuration**: on the S&P 500 benchmark it leads every tested method at every universe size (held-out one-step NLL −4.6 vs single-scale at n=100, −11.1 at n=300 before the cluster target), the 90% model confidence set collapses to it alone, and the detector's margin is confirmed out-of-time on an external industry panel (+0.53 NLL/day, p=1×10⁻⁴). Cost is O(K·n²) per update plus one Cholesky per rung per day for the detector scores. Composes with `shrinkage_target="cluster"`; mutually exclusive with `lambda_corr_fast`.
150
+ All options are off by default; the defaults reproduce the published estimator exactly. Full derivations and benchmark tables are in the papers.
120
151
 
121
- ```python
122
- est = SqueezeKernelEstimator(n_assets=100, corr_half_lives=(43, 173, 693), corr_theta=0.25)
123
- # maximal variant at high dimension:
124
- est = SqueezeKernelEstimator(n_assets=300, corr_half_lives=(43, 173, 693), shrinkage_target="cluster")
125
- ```
152
+ **Multi-scale correlation memory** (`corr_half_lives=(43, 173, 693)`): replaces the single correlation timescale with a ladder of EWMAs whose blend a sequential surprise detector tilts toward fast or slow memory as the evidence demands. This is the papers' headline configuration: it leads every tested method at every universe size, and the blend stays convex so PSD still holds by construction.
126
153
 
127
- **Score-exact weighting** (`weight_statistic="mahalanobis"`, use with `kappa=1.0`): drives the kernel with the Mahalanobis surprise `z'C⁻¹z/N` against the estimator's own correlation instead of the marginal dispersion. Improves accuracy in the moderate-concentration regime — use only when `n / T_eff ≲ 0.5` (e.g. n ≤ 100 at the default `lambda_corr`); at higher concentration the estimated inverse degrades it and the default is strictly better.
154
+ **Cluster shrinkage target** (`shrinkage_target="cluster"`): shrinks toward a target that respects the correlation matrix's own block structure instead of a single equicorrelation, with no clustering algorithm and zero added parameters. The best choice for large equity universes: worth 4 held-out NLL points at n=200 and 25 at n=300 on the benchmark.
128
155
 
129
156
  ```python
130
- est = SqueezeKernelEstimator(n_assets=100, kappa=1.0, weight_statistic="mahalanobis")
157
+ # recommended setup for a large equity universe:
158
+ est = SqueezeKernelEstimator(n_assets=300, corr_half_lives=(43, 173, 693),
159
+ shrinkage_target="cluster")
131
160
  ```
132
161
 
133
- **Score-driven memory** (`lambda_corr_fast=0.99`): lets stress days also *shorten* the correlation memory (decay slides from `lambda_corr` toward `lambda_corr_fast` as the kernel weight rises). Do **not** combine with the Mahalanobis option — they act on the same channel and the combination degrades accuracy.
134
-
135
- **OU volatility anchor** (`vol_anchor_phi=0.995`): mean-reverts each asset's variance prediction toward a slow per-asset anchor (a ~1000-day EWMA of squared returns) before the daily update — a two-timescale, component-style volatility structure. One global parameter with a clean interpretation (deviation half-life ≈ ln 2/(1−φ) days; φ=0.995 ≈ 139 d). On the S&P 500 n=100 benchmark this improved held-out one-step NLL by 3.3 points (4.3 at φ=0.99) and five-step NLL by 3.9 (5.0), with no degradation at n=300. `None` (default) or φ=1 reproduces the published estimator exactly.
162
+ **OU volatility anchor** (`vol_anchor_phi=0.995`): mean-reverts each asset's variance forecast toward a slow per-asset anchor, giving a two-timescale volatility structure with a single parameter (deviation half-life ≈ ln 2/(1−φ) days). Worth 3–5 NLL points on the benchmark.
136
163
 
137
- ```python
138
- est = SqueezeKernelEstimator(n_assets=100, vol_anchor_phi=0.995)
139
- ```
164
+ **Score-exact weighting** (`weight_statistic="mahalanobis"`, use with `kappa=1.0`): drives the kernel with the Mahalanobis surprise against the estimator's own correlation. Use it only in the moderate-concentration regime (`n / T_eff ≲ 0.5`).
140
165
 
141
- **Cluster shrinkage target** (`shrinkage_target="cluster"`): generalizes the equicorrelation shrinkage target to respect the correlation matrix's own block/cluster structure — with **no clustering algorithm**. The target morphs with the shrinkage intensity, T = (1−α)·T_equi + α·[(1−γ)I + γ·(C∘C)], where C∘C is the Hadamard square of the current correlation (positive semi-definite by the Schur product theorem; entries are pairwise shared-variance fractions) and γ is level-matched automatically. Zero added parameters, still O(n²), and as α→0 it reduces exactly to the default estimator. Held-out one-step NLL on the S&P 500 benchmark: ±0.1 at n=100, **−4.2 at n=200, −25.0 at n=300** — recommended whenever the universe size approaches the effective sample size.
166
+ **Score-driven memory** (`lambda_corr_fast=0.99`): lets stress days also shorten the correlation memory. Do not combine with the Mahalanobis option.
142
167
 
143
- ```python
144
- est = SqueezeKernelEstimator(n_assets=300, shrinkage_target="cluster")
145
- ```
168
+ **New-listing usability gate** (`min_obs=60`): exposes a `usable_mask` property marking assets with at least `min_obs` observations, so deployments can exclude cold starts from scoring and optimization; the estimates themselves are unchanged.
146
169
 
147
- **New-listing usability gate** (`min_obs=60`): on an expanding universe, an asset's forecast rows are dominated by its single-observation variance initialization for its first weeks of life and are unusable for scoring or portfolio construction (measured ≈ +2,800 NLL/day on days whose scored set included such assets, on a 42-instrument multi-asset panel). `min_obs` gates nothing inside the estimator — states warm normally, all outputs are unchanged — it exposes a `usable_mask` property marking assets with at least `min_obs` finite observations, so deployments subset with it:
148
-
149
- ```python
150
- est = SqueezeKernelEstimator(n_assets=42, min_obs=60)
151
- # ... update loop ...
152
- m = est.usable_mask
153
- cov_usable = est.get_cov()[np.ix_(m, m)]
154
- ```
155
-
156
- **Alternative kernels**: pass `kernel_fn=kernel_exponential` (with `kernel_kwargs={"gamma": ...}`) or `kernel_chi2_cdf`, or any callable `(d2, *, n_observed, **kw) -> float` mapping to `[0, 1)`. The PSD guarantee holds for any such kernel.
170
+ **Alternative kernels**: pass `kernel_fn=kernel_exponential` or `kernel_chi2_cdf`, or any callable mapping to `[0, 1)`; the PSD guarantee holds for any such kernel.
157
171
 
158
172
  ## How it works
159
173
 
@@ -171,11 +185,10 @@ The complete update is a natural-gradient step on the Gaussian log-likelihood, w
171
185
  uv sync --extra full --extra dev
172
186
  uv run python -m pytest # test suite
173
187
  uv run python -m ruff check . # lint
188
+ uv run mypy # strict type check (src/squeeze_kernel)
174
189
  uv build # build sdist + wheel
175
190
  ```
176
191
 
177
- Releases: publishing a GitHub release from a `v*` tag triggers the [publish workflow](.github/workflows/publish.yml), which builds and uploads to PyPI via trusted publishing.
178
-
179
192
  ## Citation
180
193
 
181
194
  ```bibtex
@@ -0,0 +1,176 @@
1
+ # Squeeze Kernel Covariance Estimator
2
+
3
+ [![CI](https://github.com/r0k3/squeeze-kernel/actions/workflows/ci.yml/badge.svg)](https://github.com/r0k3/squeeze-kernel/actions/workflows/ci.yml)
4
+ [![PyPI](https://img.shields.io/pypi/v/squeeze-kernel.svg)](https://pypi.org/project/squeeze-kernel/)
5
+ [![Python](https://img.shields.io/pypi/pyversions/squeeze-kernel.svg)](https://pypi.org/project/squeeze-kernel/)
6
+ [![License: MIT](https://img.shields.io/badge/license-MIT-green.svg)](LICENSE)
7
+
8
+ A **streaming covariance estimator for panels of financial returns** that learns fastest on the days that matter. One `O(n²)` update per period, positive semi-definite **by construction** at every step, missing values handled **natively**, and defaults that require no tuning. Only dependency: NumPy.
9
+
10
+ References: *"The Squeeze Kernel Covariance Estimator: Dual-Timescale Tracking with Adaptive Shrinkage"* (Kende, 2026) — [SSRN abstract 6455918](https://ssrn.com/abstract=6455918) — and its companion *"Cluster-Respecting Shrinkage for Streaming Covariance Estimation"* (Kende, 2026).
11
+
12
+ ## Why
13
+
14
+ Every standard covariance estimator treats all trading days as equally informative. Markets don't work that way: **correlations reveal themselves when markets move; calm days are mostly noise.** The Squeeze Kernel weighs each day by the information it actually carries — quiet days barely count, dispersion shocks pass through in full — and runs volatility and correlation on separate clocks, so vol spikes never contaminate the correlation estimate. The result is a single streaming recursion that:
15
+
16
+ - **is PSD at every step, structurally** — never needs eigenvalue clipping, nearest-PSD projection, or a solver;
17
+ - **adapts fastest exactly when it matters** — a Fisher-information kernel up-weights high-dispersion (stress) days, when correlation regimes actually move;
18
+ - **regularises itself** — an adaptive equicorrelation shrinkage activates automatically as the asset count approaches the effective sample size, with a provable condition-number bound;
19
+ - **ingests missing values natively** — listings, delistings, and halts enter as `NaN`; no imputation or complete-case subsetting;
20
+ - **is fast** — a full 30-year daily pass takes ~0.75 s at n=100 and ~3.4 s at n=300 (single-threaded), 30–40× faster than rolling-window baselines at scale.
21
+
22
+ **The scoreboard.** On a 30-year S&P 500 panel (~7,600 out-of-sample days) the default single-scale estimator beats EWMA, Ledoit–Wolf, OAS, nonlinear shrinkage, RMT denoising, and the Gerber statistic on one-step density forecasts, and statistically ties DCC — the only two methods in the 90% model confidence set. The headline configuration (multi-scale correlation memory, `corr_half_lives=(43, 173, 693)`) goes further: it **leads every tested method at every universe size and the 90% model confidence set collapses to it alone**, with the margin confirmed out-of-time on an external industry panel. For large equity universes, the cluster shrinkage target (`shrinkage_target="cluster"`) adds a further large gain exactly where shrinkage binds: 25 NLL points at n=300.
23
+
24
+ Against the alternatives:
25
+
26
+ - **RiskMetrics / EWMA** — same one-recursion simplicity, but two clocks and information weighting: better forecasts at zero extra operational cost.
27
+ - **Ledoit–Wolf, nonlinear shrinkage, RMT denoising** — static snapshots refit from scratch on a rolling window each day; the Squeeze Kernel is genuinely dynamic, more accurate on the benchmark, and 30–40× faster at scale.
28
+ - **DCC-GARCH** — matched (single-scale) or beaten (multi-scale) without multi-stage likelihood fitting or the fragile news-impact coefficient; at n=300 DCC needs a multi-year warm-up before its forecasts stabilise and still trails by ~25 NLL.
29
+ - **Gerber statistic** — the Squeeze Kernel is a PSD-by-construction generalization of the same robust-comovement idea: no nearest-PSD repair step, and it wins the head-to-head.
30
+
31
+ ## See the difference
32
+
33
+ Take a passive strategy any allocator would recognize: a long-only minimum-variance portfolio of 300 liquid US stocks, scaled to a 15% volatility target, rebalanced once a month, with 5 bps trading costs. Run it twice on identical data. The only thing that changes between the two runs is the covariance matrix that picks the weights and sets the exposure.
34
+
35
+ <picture>
36
+ <source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/r0k3/squeeze-kernel/main/examples/figures/vol_targeted_portfolio_dark.png">
37
+ <img alt="Vol-targeted long-only minimum-variance portfolio on 300 US equities: Squeeze Kernel vs Ledoit-Wolf equity curves, drawdown, realized volatility, and risk/return profile" src="https://raw.githubusercontent.com/r0k3/squeeze-kernel/main/examples/figures/vol_targeted_portfolio.png">
38
+ </picture>
39
+
40
+ | Method | CAGR | Vol | Sharpe | MaxDD | Calmar | Vol-target RMSE |
41
+ |---|---|---|---|---|---|---|
42
+ | **Squeeze Kernel** | **13.1%** | **13.0%** | **1.00** | **-31.2%** | **0.42** | **6.47%** |
43
+ | Ledoit-Wolf (252d) | 11.7% | 14.6% | 0.80 | -38.7% | 0.30 | 7.51% |
44
+
45
+ You can regenerate this example from the repo alone. The returns panel ships as a parquet (daily returns for 300 US stocks, sourced from Yahoo Finance), and the script prints the table and redraws the figure:
46
+
47
+ ```bash
48
+ pip install squeeze-kernel pandas pyarrow scikit-learn matplotlib
49
+ python examples/vol_targeted_portfolio.py # ~2 minutes
50
+ ```
51
+
52
+ The full protocol and data notes live in [`examples/vol_targeted_portfolio.py`](examples/vol_targeted_portfolio.py) and [`examples/data/build_equity_panel.py`](examples/data/build_equity_panel.py).
53
+
54
+ ## Installation
55
+
56
+ ```bash
57
+ pip install squeeze-kernel # NumPy only
58
+ pip install "squeeze-kernel[full]" # + SciPy (kappa calibration, chi² kernel)
59
+ ```
60
+
61
+ ## Quickstart
62
+
63
+ ```python
64
+ import numpy as np
65
+ from squeeze_kernel import SqueezeKernelEstimator
66
+
67
+ # daily_returns: array of shape (T, n) — may contain NaN for missing assets
68
+ est = SqueezeKernelEstimator(n_assets=daily_returns.shape[1])
69
+
70
+ for r_t in daily_returns: # stream one day at a time
71
+ est.update(r_t)
72
+
73
+ cov = est.get_cov() # (n, n) covariance, PSD by construction
74
+ corr = est.get_corr() # (n, n) correlation
75
+ ```
76
+
77
+ That is the whole API for most uses. The defaults (`lambda_vol=0.98`, `lambda_corr=0.996`, `kappa=0.25`) are the paper-recommended settings for daily returns, selected by time-series cross-validation and robust across a 50× parameter sweep — deploy them as-is.
78
+
79
+ Batch mode, if you prefer the full path in one call:
80
+
81
+ ```python
82
+ from squeeze_kernel import estimate_squeeze_cov
83
+
84
+ cov_path, corr_path, weights = estimate_squeeze_cov(daily_returns, with_weights=True)
85
+ # cov_path: (T, n, n) — the estimate after each day
86
+ ```
87
+
88
+ A complete runnable walkthrough (streaming, missing data, batch) is in [`examples/quickstart.py`](examples/quickstart.py).
89
+
90
+ ## Missing values
91
+
92
+ Pass `NaN` for any asset not observed on a given day — nothing else to do:
93
+
94
+ ```python
95
+ r_t = np.array([0.004, np.nan, -0.011]) # asset 2 not trading today
96
+ est.update(r_t) # PSD preserved, no imputation
97
+ ```
98
+
99
+ ## Parameters
100
+
101
+ | Parameter | Default | Meaning |
102
+ |---|---|---|
103
+ | `lambda_vol` | `0.98` | volatility EWMA decay (half-life ≈ 34 days) |
104
+ | `lambda_corr` | `0.996` | correlation EWMA decay (half-life ≈ 173 days, T_eff ≈ 250) |
105
+ | `kappa` | `0.25` | Fisher kernel saturation; higher = stronger calm-day filtering |
106
+ | `shrinkage` | `"auto"` | adaptive equicorrelation shrinkage (`"none"` or a float to override) |
107
+ | `shrinkage_delta` | `0.10` | concentration threshold at which shrinkage activates |
108
+
109
+ Useful read-only state after each `update()`: `est.weight` (last kernel weight), `est.effective_sample_size` (kernel-weighted T_eff), `est.shrinkage_intensity` (current α).
110
+
111
+ To recalibrate `kappa` for a different asset class (requires the `full` extra):
112
+
113
+ ```python
114
+ kappa = SqueezeKernelEstimator.calibrate_kappa(burn_in_returns, target_weight=0.6)
115
+ ```
116
+
117
+ ## Advanced options
118
+
119
+ All options are off by default; the defaults reproduce the published estimator exactly. Full derivations and benchmark tables are in the papers.
120
+
121
+ **Multi-scale correlation memory** (`corr_half_lives=(43, 173, 693)`): replaces the single correlation timescale with a ladder of EWMAs whose blend a sequential surprise detector tilts toward fast or slow memory as the evidence demands. This is the papers' headline configuration: it leads every tested method at every universe size, and the blend stays convex so PSD still holds by construction.
122
+
123
+ **Cluster shrinkage target** (`shrinkage_target="cluster"`): shrinks toward a target that respects the correlation matrix's own block structure instead of a single equicorrelation, with no clustering algorithm and zero added parameters. The best choice for large equity universes: worth 4 held-out NLL points at n=200 and 25 at n=300 on the benchmark.
124
+
125
+ ```python
126
+ # recommended setup for a large equity universe:
127
+ est = SqueezeKernelEstimator(n_assets=300, corr_half_lives=(43, 173, 693),
128
+ shrinkage_target="cluster")
129
+ ```
130
+
131
+ **OU volatility anchor** (`vol_anchor_phi=0.995`): mean-reverts each asset's variance forecast toward a slow per-asset anchor, giving a two-timescale volatility structure with a single parameter (deviation half-life ≈ ln 2/(1−φ) days). Worth 3–5 NLL points on the benchmark.
132
+
133
+ **Score-exact weighting** (`weight_statistic="mahalanobis"`, use with `kappa=1.0`): drives the kernel with the Mahalanobis surprise against the estimator's own correlation. Use it only in the moderate-concentration regime (`n / T_eff ≲ 0.5`).
134
+
135
+ **Score-driven memory** (`lambda_corr_fast=0.99`): lets stress days also shorten the correlation memory. Do not combine with the Mahalanobis option.
136
+
137
+ **New-listing usability gate** (`min_obs=60`): exposes a `usable_mask` property marking assets with at least `min_obs` observations, so deployments can exclude cold starts from scoring and optimization; the estimates themselves are unchanged.
138
+
139
+ **Alternative kernels**: pass `kernel_fn=kernel_exponential` or `kernel_chi2_cdf`, or any callable mapping to `[0, 1)`; the PSD guarantee holds for any such kernel.
140
+
141
+ ## How it works
142
+
143
+ Three mechanisms in one recursion:
144
+
145
+ 1. **Dual-timescale EWMA** — fast per-asset volatility (`lambda_vol`) is separated from slow correlation dynamics (`lambda_corr`), so variance shocks don't contaminate the correlation estimate.
146
+ 2. **Fisher kernel weighting** — each day's standardized outer product enters with weight `w = d²/(d² + kappa)`, where `d²` is the mean squared standardized return: calm days contribute little, dispersion shocks contribute fully.
147
+ 3. **Adaptive equicorrelation shrinkage** — `alpha = min(1, max(0, n/(2·S) − delta))` blends toward an equicorrelation target using the estimator's own kernel-weighted sample size `S`; it is a no-op at low dimension and provides provably bounded conditioning at high dimension.
148
+
149
+ The complete update is a natural-gradient step on the Gaussian log-likelihood, with the kernel weight acting as an adaptive Riemannian learning rate (paper, Appendix B).
150
+
151
+ ## Development
152
+
153
+ ```bash
154
+ uv sync --extra full --extra dev
155
+ uv run python -m pytest # test suite
156
+ uv run python -m ruff check . # lint
157
+ uv run mypy # strict type check (src/squeeze_kernel)
158
+ uv build # build sdist + wheel
159
+ ```
160
+
161
+ ## Citation
162
+
163
+ ```bibtex
164
+ @article{kende2026squeeze,
165
+ title = {The Squeeze Kernel Covariance Estimator: Dual-Timescale Tracking with Adaptive Shrinkage},
166
+ author = {Kende, Robert},
167
+ year = {2026},
168
+ note = {Available at SSRN: \url{https://ssrn.com/abstract=6455918}}
169
+ }
170
+ ```
171
+
172
+ See also [`CITATION.cff`](CITATION.cff).
173
+
174
+ ## License
175
+
176
+ MIT
@@ -0,0 +1,60 @@
1
+ [build-system]
2
+ requires = ["uv_build>=0.10.6,<0.11.0"]
3
+ build-backend = "uv_build"
4
+
5
+ [project]
6
+ name = "squeeze-kernel"
7
+ version = "0.7.1"
8
+ description = "Streaming, PSD-by-construction covariance estimator with Fisher-kernel weighting and adaptive shrinkage"
9
+ readme = "README.md"
10
+ license = "MIT"
11
+ requires-python = ">=3.10"
12
+ keywords = [
13
+ "covariance",
14
+ "correlation",
15
+ "ewma",
16
+ "kernel",
17
+ "risk",
18
+ "streaming",
19
+ ]
20
+ classifiers = [
21
+ "Development Status :: 4 - Beta",
22
+ "Intended Audience :: Science/Research",
23
+ "Topic :: Scientific/Engineering :: Mathematics",
24
+ "License :: OSI Approved :: MIT License",
25
+ "Programming Language :: Python :: 3",
26
+ "Programming Language :: Python :: 3.10",
27
+ "Programming Language :: Python :: 3.11",
28
+ "Programming Language :: Python :: 3.12",
29
+ "Programming Language :: Python :: 3.13",
30
+ "Operating System :: OS Independent",
31
+ ]
32
+ dependencies = ["numpy>=1.24"]
33
+
34
+ [[project.authors]]
35
+ name = "Robert Kende"
36
+
37
+ [project.urls]
38
+ Homepage = "https://github.com/r0k3/squeeze-kernel"
39
+ Repository = "https://github.com/r0k3/squeeze-kernel"
40
+ Issues = "https://github.com/r0k3/squeeze-kernel/issues"
41
+
42
+ [project.optional-dependencies]
43
+ full = ["scipy>=1.11"]
44
+ dev = [
45
+ "pytest>=7.0",
46
+ "ruff>=0.7",
47
+ "scipy>=1.11",
48
+ "mypy>=1.10",
49
+ ]
50
+
51
+ [tool.pytest.ini_options]
52
+ testpaths = ["tests"]
53
+
54
+ [tool.mypy]
55
+ files = ["src/squeeze_kernel"]
56
+ strict = true
57
+
58
+ [[tool.mypy.overrides]]
59
+ module = "scipy.*"
60
+ ignore_missing_imports = true
@@ -4,7 +4,7 @@ build-backend = "uv_build"
4
4
 
5
5
  [project]
6
6
  name = "squeeze-kernel"
7
- version = "0.6.0"
7
+ version = "0.7.1"
8
8
  description = "Streaming, PSD-by-construction covariance estimator with Fisher-kernel weighting and adaptive shrinkage"
9
9
  readme = "README.md"
10
10
  license = "MIT"
@@ -32,7 +32,15 @@ Issues = "https://github.com/r0k3/squeeze-kernel/issues"
32
32
 
33
33
  [project.optional-dependencies]
34
34
  full = ["scipy>=1.11"]
35
- dev = ["pytest>=7.0", "ruff>=0.7", "scipy>=1.11"]
35
+ dev = ["pytest>=7.0", "ruff>=0.7", "scipy>=1.11", "mypy>=1.10"]
36
36
 
37
37
  [tool.pytest.ini_options]
38
38
  testpaths = ["tests"]
39
+
40
+ [tool.mypy]
41
+ files = ["src/squeeze_kernel"]
42
+ strict = true
43
+
44
+ [[tool.mypy.overrides]]
45
+ module = "scipy.*"
46
+ ignore_missing_imports = true
@@ -34,4 +34,4 @@ __all__ = [
34
34
  "kernel_chi2_cdf",
35
35
  ]
36
36
 
37
- __version__ = "0.6.0"
37
+ __version__ = "0.7.1"
@@ -2,69 +2,59 @@
2
2
 
3
3
  from __future__ import annotations
4
4
 
5
+ from typing import Any
6
+
5
7
  import numpy as np
8
+ from numpy.typing import ArrayLike
6
9
 
7
10
  from squeeze_kernel.estimator import SqueezeKernelEstimator
8
- from squeeze_kernel.kernels import KernelFn
9
11
 
10
12
 
11
13
  def estimate_squeeze_cov(
12
- returns,
14
+ returns: ArrayLike,
13
15
  *,
14
- lambda_vol: float = 0.98,
15
- lambda_corr: float = 0.996,
16
- kappa: float | None = None,
17
- kernel_fn: KernelFn | None = None,
18
- kernel_kwargs: dict[str, object] | None = None,
19
- epsilon: float = 1e-8,
20
- shrinkage: str | float = "auto",
21
- shrinkage_delta: float = 0.10,
22
- impute_missing: bool = False,
23
- impute_threshold: float = 0.6,
24
16
  with_corr: bool = True,
25
17
  with_weights: bool = False,
18
+ **estimator_kwargs: Any,
26
19
  ) -> tuple[np.ndarray, np.ndarray | None, np.ndarray | None]:
27
20
  """Estimate streaming covariance over an entire returns panel.
28
21
 
22
+ Runs one :class:`SqueezeKernelEstimator` over the panel day by day and
23
+ collects the estimate path. ``n_assets`` is taken from the panel shape;
24
+ every other keyword argument is forwarded unchanged to the estimator, so
25
+ batch mode reaches the full estimator surface — ``kappa``, ``shrinkage``,
26
+ ``shrinkage_target``, ``corr_half_lives``, ``weight_statistic``,
27
+ ``vol_anchor_phi``, ``min_obs``, custom kernels, and so on (see the
28
+ estimator's docstring for the complete list and defaults).
29
+
29
30
  Parameters
30
31
  ----------
31
32
  returns : array-like, shape (T, n)
32
33
  2D return matrix. May contain NaN for missing observations.
33
- kappa : float, optional
34
- Saturation parameter for the default Fisher kernel.
35
- kernel_fn : callable, optional
36
- Custom kernel ``(d2, *, n_observed, **kw) -> float``.
37
- kernel_kwargs : dict, optional
38
- Extra keyword arguments forwarded to ``kernel_fn``.
39
34
  with_corr : bool
40
35
  If True, also return the correlation tensor.
41
36
  with_weights : bool
42
37
  If True, also return per-timestamp kernel weights.
38
+ **estimator_kwargs
39
+ Passed through to ``SqueezeKernelEstimator``.
43
40
 
44
41
  Returns
45
42
  -------
46
43
  cov : ndarray, shape (T, n, n)
44
+ The covariance estimate after each day.
47
45
  corr : ndarray or None, shape (T, n, n)
48
46
  weights : ndarray or None, shape (T,)
49
47
  """
50
48
  values = np.asarray(returns, dtype=np.float64)
51
49
  if values.ndim != 2:
52
50
  raise ValueError(f"Expected 2D returns, got shape {values.shape}.")
51
+ if "n_assets" in estimator_kwargs:
52
+ raise ValueError(
53
+ "n_assets is derived from the panel shape and cannot be passed."
54
+ )
53
55
  t_total, n_assets = values.shape
54
56
 
55
- est = SqueezeKernelEstimator(
56
- n_assets,
57
- lambda_vol=lambda_vol,
58
- lambda_corr=lambda_corr,
59
- kappa=kappa,
60
- kernel_fn=kernel_fn,
61
- kernel_kwargs=kernel_kwargs,
62
- epsilon=epsilon,
63
- shrinkage=shrinkage,
64
- shrinkage_delta=shrinkage_delta,
65
- impute_missing=impute_missing,
66
- impute_threshold=impute_threshold,
67
- )
57
+ est = SqueezeKernelEstimator(n_assets, **estimator_kwargs)
68
58
 
69
59
  cov = np.empty((t_total, n_assets, n_assets), dtype=np.float64)
70
60
  corr = np.empty_like(cov) if with_corr else None
@@ -18,11 +18,118 @@ except ImportError: # pragma: no cover
18
18
  from squeeze_kernel.kernels import (
19
19
  KernelFn, kernel_fisher, calibrate_kappa, extract_d2_series,
20
20
  )
21
+ from numpy.typing import ArrayLike
21
22
 
22
23
  if TYPE_CHECKING:
23
24
  from collections.abc import Sequence
24
25
 
25
26
 
27
+ class _SurpriseDetector:
28
+ """Two-sided Page CUSUM on the studentised fast-vs-slow per-rung
29
+ predictive-score drift.
30
+
31
+ Gates the blend weights of the scale-free correlation ladder: on each
32
+ update the day is scored under every rung's previous (t-1) covariance
33
+ — causal, since ``record`` is called with the rung states *after* the
34
+ update — and the drift of the centred rung scores advances the CUSUM.
35
+ An alarm sets a half-magnitude tilt of the theta-prior toward the
36
+ inverse-horizon vector (fast alarm) or the square-root-horizon vector
37
+ (slow alarm); on all other days the tilt decays at the fastest rung's
38
+ half-life. State: five scalars (``gp``, ``gm``, ``tilt``, ``scale``
39
+ and the cached ``prev_sig`` rung covariances).
40
+ """
41
+
42
+ _DRIFT = 0.5
43
+ _THRESHOLD = 4.9721088583 # Siegmund ARL approximation at ~2 years
44
+ _SNAP = 0.5
45
+ _CLIP = 3.0
46
+ _JITTER = 1e-12
47
+
48
+ def __init__(self, half_lives: np.ndarray) -> None:
49
+ hl = half_lives
50
+ self.pi_fast = (1.0 / hl) / (1.0 / hl).sum()
51
+ self.pi_slow = hl ** 0.5 / (hl ** 0.5).sum()
52
+ self._lam_tilt = 2.0 ** (-1.0 / float(hl.min()))
53
+ self._gamma_scale = 2.0 ** (-1.0 / float(np.median(hl)))
54
+ self._k_fast = int(np.argmin(hl))
55
+ self._k_slow = int(np.argmax(hl))
56
+ self.gp = 0.0
57
+ self.gm = 0.0
58
+ self.tilt = 0.0
59
+ self.scale = 1.0
60
+ self.prev_sig: list[np.ndarray] | None = None
61
+
62
+ def score(self, r_t: np.ndarray, finite: np.ndarray) -> np.ndarray | None:
63
+ """Per-rung Gaussian log-likelihood of ``r_t`` under ``prev_sig``,
64
+ restricted to the observed assets. Returns None when there is no
65
+ previous state, nothing is observed, or the day is degenerate
66
+ (e.g. a fresh listing); a degenerate day also decays the tilt."""
67
+ if self.prev_sig is None or not finite.any():
68
+ return None
69
+ oidx = np.flatnonzero(finite)
70
+ r_o = r_t[oidx]
71
+ ell = np.empty(len(self.prev_sig))
72
+ for k, sig in enumerate(self.prev_sig):
73
+ sub = sig[np.ix_(oidx, oidx)]
74
+ sub = (sub + sub.T) * 0.5
75
+ if _cho_factor is not None:
76
+ # One Cholesky per rung: logdet from the factor's
77
+ # diagonal, quadratic form via triangular solves.
78
+ try:
79
+ cf = _cho_factor(sub, lower=True, check_finite=False)
80
+ except np.linalg.LinAlgError:
81
+ self.tilt *= self._lam_tilt
82
+ return None
83
+ logdet = 2.0 * np.log(np.diagonal(cf[0])).sum()
84
+ quad = float(r_o @ _cho_solve(cf, r_o, check_finite=False))
85
+ else:
86
+ sign, logdet = np.linalg.slogdet(sub)
87
+ if sign <= 0:
88
+ self.tilt *= self._lam_tilt
89
+ return None
90
+ try:
91
+ quad = float(r_o @ np.linalg.solve(sub, r_o))
92
+ except np.linalg.LinAlgError:
93
+ self.tilt *= self._lam_tilt
94
+ return None
95
+ ell[k] = -0.5 * (oidx.size * np.log(2 * np.pi) + logdet + quad)
96
+ return ell
97
+
98
+ def advance(self, ell: np.ndarray | None) -> None:
99
+ """Advance the CUSUM with the rung scores for this day. ``None``
100
+ (no clean score) is a no-op: the degenerate-day tilt decay already
101
+ happened inside ``score``."""
102
+ if ell is None:
103
+ return
104
+ dd = ell - ell.mean()
105
+ rms = float(np.sqrt((dd @ dd) / ell.size))
106
+ self.scale = (self._gamma_scale * self.scale
107
+ + (1.0 - self._gamma_scale) * rms)
108
+ zc = np.clip(dd / (self.scale + self._JITTER), -self._CLIP, self._CLIP)
109
+ zfs = float(zc[self._k_fast] - zc[self._k_slow])
110
+ self.gp = max(0.0, self.gp + zfs - self._DRIFT)
111
+ self.gm = max(0.0, self.gm - zfs - self._DRIFT)
112
+ if self.gp > self._THRESHOLD:
113
+ self.tilt, self.gp = self._SNAP, 0.0
114
+ elif self.gm > self._THRESHOLD:
115
+ self.tilt, self.gm = -self._SNAP, 0.0
116
+ else:
117
+ self.tilt *= self._lam_tilt
118
+
119
+ def record(self, rung_covs: list[np.ndarray]) -> None:
120
+ """Cache this step's rung covariances as next step's forecasts."""
121
+ self.prev_sig = rung_covs
122
+
123
+ def blend(self, prior: np.ndarray) -> np.ndarray:
124
+ """Blend weights for the rung covariances: the theta-prior on
125
+ non-alarmed days, tilted half-magnitude toward the horizon vectors
126
+ while an alarm is live. Convex in both regimes."""
127
+ t = self.tilt
128
+ if t >= 0:
129
+ return np.asarray((1.0 - t) * prior + t * self.pi_fast)
130
+ return np.asarray((1.0 + t) * prior + (-t) * self.pi_slow)
131
+
132
+
26
133
  class SqueezeKernelEstimator:
27
134
  """Streaming robust covariance estimator with pluggable kernel weighting.
28
135
 
@@ -251,6 +358,7 @@ class SqueezeKernelEstimator:
251
358
  self.corr_theta = corr_theta
252
359
  self._corr_lam: np.ndarray | None = None
253
360
  self._corr_w: np.ndarray | None = None
361
+ self._detector: _SurpriseDetector | None = None
254
362
  self._Q_list: list[np.ndarray] | None = None
255
363
  self._S_list: list[float] | None = None
256
364
  self._adaptive = False
@@ -271,21 +379,11 @@ class SqueezeKernelEstimator:
271
379
  self._corr_w = w / w.sum()
272
380
  self._Q_list = [np.eye(n_assets, dtype=np.float64) for _ in hl]
273
381
  self._S_list = [float(epsilon) for _ in hl]
274
- # Surprise-gated blend weights (integral for K >= 2): two-sided
275
- # Page CUSUM on the studentised fast-vs-slow per-rung
276
- # predictive-score drift; drift 0.5, threshold from Siegmund's
277
- # ARL approximation at ARL0 = 504 trading days (~2 years).
382
+ # Surprise-gated blend weights (integral for K >= 2): see
383
+ # ``_SurpriseDetector``.
278
384
  self._adaptive = hl.size >= 2
279
385
  if self._adaptive:
280
- self._aw_pi_fast = (1.0 / hl) / (1.0 / hl).sum()
281
- self._aw_pi_slow = hl ** 0.5 / (hl ** 0.5).sum()
282
- self._aw_lam_tilt = 2.0 ** (-1.0 / float(hl.min()))
283
- self._aw_gamma_scale = 2.0 ** (-1.0 / float(np.median(hl)))
284
- self._aw_drift, self._aw_b, self._aw_snap = 0.5, 4.9721088583, 0.5
285
- self._aw_gp = self._aw_gm = 0.0
286
- self._aw_tilt = 0.0
287
- self._aw_scale = 1.0
288
- self._aw_prev_sig: list[np.ndarray] | None = None
386
+ self._detector = _SurpriseDetector(hl)
289
387
 
290
388
  # Resolve shrinkage
291
389
  if isinstance(shrinkage, str):
@@ -328,7 +426,7 @@ class SqueezeKernelEstimator:
328
426
 
329
427
  # ── Public API ────────────────────────────────────────────────────────
330
428
 
331
- def update(self, r_t) -> float:
429
+ def update(self, r_t: ArrayLike) -> float:
332
430
  """Process one return vector and update the covariance estimate.
333
431
 
334
432
  Parameters
@@ -353,77 +451,37 @@ class SqueezeKernelEstimator:
353
451
  # ── Adaptive-weight detector: score r_t under yesterday's per-rung
354
452
  # forecasts, then advance the CUSUM (weights used below therefore
355
453
  # reflect information through r_t only — causal). ──
356
- if (self._adaptive and self._aw_prev_sig is not None
357
- and finite.any()):
358
- oidx = np.flatnonzero(finite)
359
- r_o = r_t[oidx]
360
- ell = np.empty(len(self._aw_prev_sig))
361
- ok = True
362
- for k, sig in enumerate(self._aw_prev_sig):
363
- sub = sig[np.ix_(oidx, oidx)]
364
- sub = (sub + sub.T) * 0.5
365
- if _cho_factor is not None:
366
- # One Cholesky per rung: logdet from the factor's
367
- # diagonal, quadratic form via triangular solves.
368
- try:
369
- cf = _cho_factor(sub, lower=True, check_finite=False)
370
- except np.linalg.LinAlgError:
371
- ok = False # degenerate day (e.g. a fresh
372
- break # listing): no clean score
373
- logdet = 2.0 * np.log(np.diagonal(cf[0])).sum()
374
- quad = float(r_o @ _cho_solve(cf, r_o, check_finite=False))
375
- else:
376
- sign, logdet = np.linalg.slogdet(sub)
377
- if sign <= 0:
378
- ok = False # degenerate day (e.g. a fresh
379
- break # listing): no clean score
380
- try:
381
- quad = float(r_o @ np.linalg.solve(sub, r_o))
382
- except np.linalg.LinAlgError:
383
- ok = False
384
- break
385
- ell[k] = -0.5 * (oidx.size * np.log(2 * np.pi) + logdet + quad)
386
- if not ok:
387
- self._aw_tilt *= self._aw_lam_tilt
388
- ell = None
389
- else:
390
- ell = None
391
- if ell is not None:
392
- dd = ell - ell.mean()
393
- rms = float(np.sqrt((dd @ dd) / ell.size))
394
- self._aw_scale = (self._aw_gamma_scale * self._aw_scale
395
- + (1.0 - self._aw_gamma_scale) * rms)
396
- zc = np.clip(dd / (self._aw_scale + 1e-12), -3.0, 3.0)
397
- k_fast = int(np.argmin(self.corr_half_lives))
398
- k_slow = int(np.argmax(self.corr_half_lives))
399
- zfs = float(zc[k_fast] - zc[k_slow])
400
- self._aw_gp = max(0.0, self._aw_gp + zfs - self._aw_drift)
401
- self._aw_gm = max(0.0, self._aw_gm - zfs - self._aw_drift)
402
- if self._aw_gp > self._aw_b:
403
- self._aw_tilt, self._aw_gp = self._aw_snap, 0.0
404
- elif self._aw_gm > self._aw_b:
405
- self._aw_tilt, self._aw_gm = -self._aw_snap, 0.0
406
- else:
407
- self._aw_tilt *= self._aw_lam_tilt
454
+ ell = None
455
+ detector = self._detector
456
+ if detector is not None:
457
+ ell = detector.score(r_t, finite)
458
+ detector.advance(ell)
408
459
 
409
460
  # ── Volatility update ──
410
- if self._var_t is None:
411
- self._var_t = np.zeros(n, dtype=np.float64)
412
- self._var_init = np.zeros(n, dtype=bool)
461
+ # Locals alias the variance states: the guard initialises the pair
462
+ # together, so both are non-None afterwards, and local aliases keep
463
+ # the narrowing visible to mypy.
464
+ var_t = self._var_t
465
+ var_init = self._var_init
466
+ if var_t is None or var_init is None:
467
+ var_t = np.zeros(n, dtype=np.float64)
468
+ var_init = np.zeros(n, dtype=bool)
469
+ self._var_t = var_t
470
+ self._var_init = var_init
413
471
  if self.vol_anchor_phi is not None:
414
472
  self._var_anchor = np.zeros(n, dtype=np.float64)
415
473
 
416
- first = finite & ~self._var_init
417
- repeat = finite & self._var_init
474
+ first = finite & ~var_init
475
+ repeat = finite & var_init
418
476
  if np.any(first):
419
- self._var_t[first] = r_t[first] ** 2 + eps
420
- self._var_init[first] = True
477
+ var_t[first] = r_t[first] ** 2 + eps
478
+ var_init[first] = True
421
479
  if self._var_anchor is not None:
422
- self._var_anchor[first] = self._var_t[first]
480
+ self._var_anchor[first] = var_t[first]
423
481
  if np.any(repeat):
424
482
  if self.vol_anchor_phi is None:
425
- self._var_t[repeat] = (
426
- self.lambda_vol * self._var_t[repeat]
483
+ var_t[repeat] = (
484
+ self.lambda_vol * var_t[repeat]
427
485
  + (1.0 - self.lambda_vol) * r_t[repeat] ** 2
428
486
  )
429
487
  else:
@@ -431,20 +489,23 @@ class SqueezeKernelEstimator:
431
489
  # per-asset anchor before the measurement update, then update
432
490
  # the anchor itself (order matters and matches the validated
433
491
  # experiment: prediction uses the *old* anchor).
492
+ var_anchor = self._var_anchor
493
+ assert var_anchor is not None
494
+ # allocated together with the anchor on the first update
434
495
  phi = self.vol_anchor_phi
435
496
  lam_bar = self.vol_anchor_decay
436
- anchor = self._var_anchor[repeat]
437
- v_pred = anchor + phi * (self._var_t[repeat] - anchor)
438
- self._var_t[repeat] = (
497
+ anchor = var_anchor[repeat]
498
+ v_pred = anchor + phi * (var_t[repeat] - anchor)
499
+ var_t[repeat] = (
439
500
  self.lambda_vol * v_pred
440
501
  + (1.0 - self.lambda_vol) * r_t[repeat] ** 2
441
502
  )
442
- self._var_anchor[repeat] = (
503
+ var_anchor[repeat] = (
443
504
  lam_bar * anchor + (1.0 - lam_bar) * r_t[repeat] ** 2
444
505
  )
445
506
 
446
507
  vol_t = np.zeros(n, dtype=np.float64)
447
- vol_t[self._var_init] = np.sqrt(self._var_t[self._var_init])
508
+ vol_t[var_init] = np.sqrt(var_t[var_init])
448
509
 
449
510
  # ── Standardized returns ──
450
511
  z_t = np.zeros(n, dtype=np.float64)
@@ -459,8 +520,10 @@ class SqueezeKernelEstimator:
459
520
  # lazy, so bring _corr up to the t-1 state first.
460
521
  if self._dirty:
461
522
  self._materialize()
523
+ corr_prev = self._corr
524
+ assert corr_prev is not None # _materialize always sets it
462
525
  try:
463
- c_sub = self._corr[np.ix_(finite, finite)]
526
+ c_sub = corr_prev[np.ix_(finite, finite)]
464
527
  d2 = float(z_t[finite] @ np.linalg.solve(c_sub, z_t[finite])) / n_obs
465
528
  except np.linalg.LinAlgError:
466
529
  pass
@@ -478,9 +541,12 @@ class SqueezeKernelEstimator:
478
541
  # np.outer otherwise allocates each step. zz' is computed once
479
542
  # and shared by every rung.
480
543
  np.multiply.outer(z_t, z_t, out=self._scratch_outer)
481
- if self._corr_lam is None:
544
+ corr_lam = self._corr_lam
545
+ if corr_lam is None:
482
546
  # ── Single-scale correlation EWMA (published path) ──
483
- lam_c = self.lambda_corr
547
+ Q_t = self._Q_t
548
+ assert Q_t is not None # the single-scale state exists exactly
549
+ lam_c = self.lambda_corr # when the ladder is not configured
484
550
  if self.lambda_corr_fast is not None:
485
551
  # Score-driven memory: stress days (w_t → 1) shorten the memory
486
552
  # toward lambda_corr_fast; calm days keep the slow decay.
@@ -490,23 +556,28 @@ class SqueezeKernelEstimator:
490
556
  # Q <- (1 - eta) Q + eta zz'. With w_t = 0 both S and M decay
491
557
  # by lam_c, so Q is unchanged — no matrix work at all.
492
558
  eta = w_t / self._S_t
493
- self._Q_t *= 1.0 - eta
559
+ Q_t *= 1.0 - eta
494
560
  self._scratch_outer *= eta
495
- self._Q_t += self._scratch_outer
561
+ Q_t += self._scratch_outer
496
562
  else:
497
563
  # ── Scale-free ladder (Mode A) ──
498
564
  # Update K normalised correlation states on the geometric
499
565
  # half-life ladder; extraction (normalise + shrink + blend)
500
566
  # happens lazily in _materialize().
567
+ Q_list = self._Q_list
568
+ S_list = self._S_list
569
+ corr_w = self._corr_w
570
+ assert Q_list is not None and S_list is not None
571
+ assert corr_w is not None # allocated together with corr_lam
501
572
  s_eff = 0.0
502
- for k in range(self._corr_lam.size):
503
- self._S_list[k] = self._corr_lam[k] * self._S_list[k] + w_t
573
+ for k in range(corr_lam.size):
574
+ S_list[k] = corr_lam[k] * S_list[k] + w_t
504
575
  if add:
505
- eta = w_t / self._S_list[k]
506
- self._Q_list[k] *= 1.0 - eta
576
+ eta = w_t / S_list[k]
577
+ Q_list[k] *= 1.0 - eta
507
578
  np.multiply(self._scratch_outer, eta, out=self._scratch_had)
508
- self._Q_list[k] += self._scratch_had
509
- s_eff += self._corr_w[k] * self._S_list[k]
579
+ Q_list[k] += self._scratch_had
580
+ s_eff += corr_w[k] * S_list[k]
510
581
  self._S_t = s_eff # blended effective size (for the property)
511
582
  self._vol_t = vol_t
512
583
  self._dirty = True
@@ -521,7 +592,9 @@ class SqueezeKernelEstimator:
521
592
  raise RuntimeError("Call update() at least once before get_cov().")
522
593
  if self._dirty:
523
594
  self._materialize()
524
- return self._cov.copy()
595
+ cov = self._cov
596
+ assert cov is not None
597
+ return cov.copy()
525
598
 
526
599
  def get_corr(self) -> np.ndarray:
527
600
  """Return the current correlation matrix estimate (n x n)."""
@@ -529,7 +602,9 @@ class SqueezeKernelEstimator:
529
602
  raise RuntimeError("Call update() at least once before get_corr().")
530
603
  if self._dirty:
531
604
  self._materialize()
532
- return self._corr.copy()
605
+ corr = self._corr
606
+ assert corr is not None
607
+ return corr.copy()
533
608
 
534
609
  @property
535
610
  def weight(self) -> float:
@@ -562,7 +637,7 @@ class SqueezeKernelEstimator:
562
637
 
563
638
  @staticmethod
564
639
  def calibrate_kappa(
565
- returns, target_weight: float = 0.5, lambda_vol: float = 0.98,
640
+ returns: ArrayLike, target_weight: float = 0.5, lambda_vol: float = 0.98,
566
641
  ) -> float:
567
642
  """Calibrate κ from burn-in data so E[w_t] ≈ target_weight.
568
643
 
@@ -669,20 +744,24 @@ class SqueezeKernelEstimator:
669
744
  """Adaptive-weight extraction: build per-rung shrunk covariances
670
745
  (kept for the next update's detector scores), blend with the
671
746
  CUSUM-tilted weights, derive _cov/_corr. Runs eagerly."""
672
- eps = self.epsilon
747
+ detector = self._detector
748
+ corr_lam = self._corr_lam
749
+ Q_list = self._Q_list
750
+ S_list = self._S_list
751
+ corr_w = self._corr_w
673
752
  vol_t = self._vol_t
753
+ assert detector is not None # only called when the detector exists
754
+ assert corr_lam is not None and Q_list is not None
755
+ assert S_list is not None and corr_w is not None
756
+ assert vol_t is not None # set on every update before this runs
757
+ eps = self.epsilon
674
758
  vv = np.multiply.outer(vol_t, vol_t)
675
759
  sig_k = []
676
- for k in range(self._corr_lam.size):
677
- corr_k = self._shrunk_corr_from_Q(self._Q_list[k], self._S_list[k])
760
+ for k in range(corr_lam.size):
761
+ corr_k = self._shrunk_corr_from_Q(Q_list[k], S_list[k])
678
762
  sig_k.append(corr_k * vv)
679
- self._aw_prev_sig = sig_k
680
- t = self._aw_tilt
681
- pi = self._corr_w
682
- if t >= 0:
683
- w = (1.0 - t) * pi + t * self._aw_pi_fast
684
- else:
685
- w = (1.0 + t) * pi + (-t) * self._aw_pi_slow
763
+ detector.record(sig_k)
764
+ w = detector.blend(corr_w)
686
765
  cov = w[0] * sig_k[0]
687
766
  for k in range(1, len(sig_k)):
688
767
  cov = cov + w[k] * sig_k[k]
@@ -697,9 +776,12 @@ class SqueezeKernelEstimator:
697
776
  """Extract _cov/_corr from the current state (lazy, on demand)."""
698
777
  eps = self.epsilon
699
778
  vol_t = self._vol_t
779
+ assert vol_t is not None # get_cov/get_corr guard this before calling
700
780
  if self._corr_lam is None:
701
781
  # ── Single scale: shrunk correlation IS the correlation output ──
702
- corr = self._shrunk_corr_from_Q(self._Q_t, self._S_t)
782
+ Q_t = self._Q_t
783
+ assert Q_t is not None # the ladder is not configured on this path
784
+ corr = self._shrunk_corr_from_Q(Q_t, self._S_t)
703
785
  np.multiply.outer(vol_t, vol_t, out=self._scratch_outer)
704
786
  cov = corr * self._scratch_outer
705
787
  cov += cov.T
@@ -712,10 +794,16 @@ class SqueezeKernelEstimator:
712
794
  # ── Ladder: blend per-rung shrunk correlations, then apply the
713
795
  # (shared) volatilities once — algebraically identical to
714
796
  # blending per-rung covariances, K-1 fewer O(n^2) passes.
797
+ corr_lam = self._corr_lam
798
+ Q_list = self._Q_list
799
+ S_list = self._S_list
800
+ corr_w = self._corr_w
801
+ assert Q_list is not None and S_list is not None
802
+ assert corr_w is not None # allocated together with corr_lam
715
803
  mix = np.zeros((self.n_assets, self.n_assets), dtype=np.float64)
716
- for k in range(self._corr_lam.size):
717
- corr_k = self._shrunk_corr_from_Q(self._Q_list[k], self._S_list[k])
718
- corr_k *= self._corr_w[k]
804
+ for k in range(corr_lam.size):
805
+ corr_k = self._shrunk_corr_from_Q(Q_list[k], S_list[k])
806
+ corr_k *= corr_w[k]
719
807
  mix += corr_k
720
808
  cov = mix
721
809
  np.multiply.outer(vol_t, vol_t, out=self._scratch_outer)
@@ -751,7 +839,13 @@ def _resolve_kernel(
751
839
  if "kappa" in kw and kappa is not None:
752
840
  raise ValueError("Pass kappa either as a top-level argument or in kernel_kwargs, not both.")
753
841
 
754
- resolved_kappa = float(kw.get("kappa", 0.25 if kappa is None else kappa))
842
+ if "kappa" in kw:
843
+ kappa_src = kw["kappa"]
844
+ if isinstance(kappa_src, bool) or not isinstance(kappa_src, (int, float)):
845
+ raise ValueError("kappa in kernel_kwargs must be a number.")
846
+ resolved_kappa = float(kappa_src)
847
+ else:
848
+ resolved_kappa = 0.25 if kappa is None else float(kappa)
755
849
  if resolved_kappa <= 0.0:
756
850
  raise ValueError("kappa must be > 0.")
757
851
  kw["kappa"] = resolved_kappa
@@ -5,10 +5,13 @@ from __future__ import annotations
5
5
  import math
6
6
  from typing import Callable
7
7
 
8
+ import numpy as np
9
+ from numpy.typing import ArrayLike
10
+
8
11
  KernelFn = Callable[..., float]
9
12
 
10
13
 
11
- def kernel_fisher(d2: float, /, *, kappa: float, **kw) -> float:
14
+ def kernel_fisher(d2: float, /, *, kappa: float, **kw: object) -> float:
12
15
  """Fisher information saturation kernel: w = d² / (d² + κ).
13
16
 
14
17
  Motivated by the signal-to-noise structure of the Gaussian score.
@@ -19,14 +22,14 @@ def kernel_fisher(d2: float, /, *, kappa: float, **kw) -> float:
19
22
  return d2 / (d2 + kappa)
20
23
 
21
24
 
22
- def kernel_exponential(d2: float, /, *, gamma: float, **kw) -> float:
25
+ def kernel_exponential(d2: float, /, *, gamma: float, **kw: object) -> float:
23
26
  """Rayleigh survival (exponential) kernel: w = 1 − exp(−d²/(2γ²))."""
24
27
  if gamma <= 0.0:
25
28
  raise ValueError("gamma must be > 0.")
26
29
  return 1.0 - math.exp(-d2 / (2.0 * gamma * gamma))
27
30
 
28
31
 
29
- def kernel_chi2_cdf(d2: float, /, *, n_observed: int, **kw) -> float:
32
+ def kernel_chi2_cdf(d2: float, /, *, n_observed: int, **kw: object) -> float:
30
33
  """Chi-squared CDF kernel: w = F_χ²_N(N·d²). Requires scipy."""
31
34
  try:
32
35
  from scipy.stats import chi2
@@ -38,7 +41,7 @@ def kernel_chi2_cdf(d2: float, /, *, n_observed: int, **kw) -> float:
38
41
  return float(chi2.cdf(n_observed * d2, df=n_observed))
39
42
 
40
43
 
41
- def calibrate_kappa(d2_samples, target_mean_weight: float = 0.5) -> float:
44
+ def calibrate_kappa(d2_samples: ArrayLike, target_mean_weight: float = 0.5) -> float:
42
45
  """Find κ such that E[d²/(d²+κ)] ≈ target_mean_weight on burn-in data.
43
46
 
44
47
  This implements the closed-form initializer from Proposition 1 of the paper,
@@ -56,7 +59,6 @@ def calibrate_kappa(d2_samples, target_mean_weight: float = 0.5) -> float:
56
59
  float
57
60
  Calibrated κ value.
58
61
  """
59
- import numpy as np
60
62
  try:
61
63
  from scipy.optimize import brentq
62
64
  except ImportError as exc:
@@ -75,7 +77,7 @@ def calibrate_kappa(d2_samples, target_mean_weight: float = 0.5) -> float:
75
77
  kappa_init = mu * (1.0 / target_mean_weight - 1.0)
76
78
 
77
79
  # Refine via root finding
78
- def residual(kappa):
80
+ def residual(kappa: float) -> float:
79
81
  return float(np.mean(d2 / (d2 + kappa))) - target_mean_weight
80
82
 
81
83
  lo = max(kappa_init * 0.01, 1e-6)
@@ -90,8 +92,8 @@ def calibrate_kappa(d2_samples, target_mean_weight: float = 0.5) -> float:
90
92
 
91
93
 
92
94
  def extract_d2_series(
93
- returns, lambda_vol: float = 0.98, epsilon: float = 1e-8,
94
- ):
95
+ returns: ArrayLike, lambda_vol: float = 0.98, epsilon: float = 1e-8,
96
+ ) -> np.ndarray:
95
97
  """Extract the d̄²_t series from a returns panel for κ calibration.
96
98
 
97
99
  Parameters
@@ -106,7 +108,6 @@ def extract_d2_series(
106
108
  ndarray, shape (T,)
107
109
  Average squared standardized return at each timestep.
108
110
  """
109
- import numpy as np
110
111
 
111
112
  x = np.asarray(returns, dtype=np.float64)
112
113
  t_total, n = x.shape
@@ -1,164 +0,0 @@
1
- # Squeeze Kernel Covariance Estimator
2
-
3
- [![CI](https://github.com/r0k3/squeeze-kernel/actions/workflows/ci.yml/badge.svg)](https://github.com/r0k3/squeeze-kernel/actions/workflows/ci.yml)
4
- [![PyPI](https://img.shields.io/pypi/v/squeeze-kernel.svg)](https://pypi.org/project/squeeze-kernel/)
5
- [![Python](https://img.shields.io/pypi/pyversions/squeeze-kernel.svg)](https://pypi.org/project/squeeze-kernel/)
6
- [![License: MIT](https://img.shields.io/badge/license-MIT-green.svg)](LICENSE)
7
-
8
- A **streaming covariance estimator for panels of daily financial returns**. One `O(n²)` update per day, positive semi-definite **by construction** at every step, missing values handled **natively**, and defaults that require no tuning. Only dependency: NumPy.
9
-
10
- Reference: *"The Squeeze Kernel Covariance Estimator: Dual-Timescale Tracking with Adaptive Shrinkage"* (Kende, 2026) — [SSRN abstract 6455918](https://ssrn.com/abstract=6455918).
11
-
12
- ## Why
13
-
14
- Rolling-window estimators (Ledoit–Wolf, nonlinear shrinkage, RMT denoising) refit over a fixed window each day and cannot adapt within it; multivariate GARCH (DCC) adapts but needs multi-stage estimation and a fragile news coefficient. The Squeeze Kernel is a single streaming recursion that:
15
-
16
- - **is PSD at every step, structurally** — never needs eigenvalue clipping, nearest-PSD projection, or a solver;
17
- - **adapts fastest exactly when it matters** — a Fisher-information kernel up-weights high-dispersion (stress) days, when correlation regimes actually move;
18
- - **regularises itself** — an adaptive equicorrelation shrinkage activates automatically as the asset count approaches the effective sample size, with a provable condition-number bound;
19
- - **ingests missing values natively** — listings, delistings, and halts enter as `NaN`; no imputation or complete-case subsetting;
20
- - **is fast** — a full 30-year daily pass takes ~0.75 s at n=100 and ~3.4 s at n=300 (single-threaded), 30–40× faster than rolling-window baselines at scale.
21
-
22
- On a 30-year S&P 500 panel (n=100, ~7,600 out-of-sample days) it statistically ties DCC on one-step density forecasts and beats EWMA, Ledoit–Wolf, OAS, nonlinear shrinkage, RMT denoising, and the Gerber statistic — and it is the only method in the 90% model confidence set together with DCC. At n=300 it leads every competitor that remains statistically viable.
23
-
24
- ## Installation
25
-
26
- ```bash
27
- pip install squeeze-kernel # NumPy only
28
- pip install "squeeze-kernel[full]" # + SciPy (kappa calibration, chi² kernel)
29
- ```
30
-
31
- ## Quickstart
32
-
33
- ```python
34
- import numpy as np
35
- from squeeze_kernel import SqueezeKernelEstimator
36
-
37
- # daily_returns: array of shape (T, n) — may contain NaN for missing assets
38
- est = SqueezeKernelEstimator(n_assets=daily_returns.shape[1])
39
-
40
- for r_t in daily_returns: # stream one day at a time
41
- est.update(r_t)
42
-
43
- cov = est.get_cov() # (n, n) covariance, PSD by construction
44
- corr = est.get_corr() # (n, n) correlation
45
- ```
46
-
47
- That is the whole API for most uses. The defaults (`lambda_vol=0.98`, `lambda_corr=0.996`, `kappa=0.25`) are the paper-recommended settings for daily returns, selected by time-series cross-validation and robust across a 50× parameter sweep — deploy them as-is.
48
-
49
- Batch mode, if you prefer the full path in one call:
50
-
51
- ```python
52
- from squeeze_kernel import estimate_squeeze_cov
53
-
54
- cov_path, corr_path, weights = estimate_squeeze_cov(daily_returns, with_weights=True)
55
- # cov_path: (T, n, n) — the estimate after each day
56
- ```
57
-
58
- A complete runnable walkthrough (streaming, missing data, batch) is in [`examples/quickstart.py`](examples/quickstart.py).
59
-
60
- ## Missing values
61
-
62
- Pass `NaN` for any asset not observed on a given day — nothing else to do:
63
-
64
- ```python
65
- r_t = np.array([0.004, np.nan, -0.011]) # asset 2 not trading today
66
- est.update(r_t) # PSD preserved, no imputation
67
- ```
68
-
69
- ## Parameters
70
-
71
- | Parameter | Default | Meaning |
72
- |---|---|---|
73
- | `lambda_vol` | `0.98` | volatility EWMA decay (half-life ≈ 34 days) |
74
- | `lambda_corr` | `0.996` | correlation EWMA decay (half-life ≈ 173 days, T_eff ≈ 250) |
75
- | `kappa` | `0.25` | Fisher kernel saturation; higher = stronger calm-day filtering |
76
- | `shrinkage` | `"auto"` | adaptive equicorrelation shrinkage (`"none"` or a float to override) |
77
- | `shrinkage_delta` | `0.10` | concentration threshold at which shrinkage activates |
78
-
79
- Useful read-only state after each `update()`: `est.weight` (last kernel weight), `est.effective_sample_size` (kernel-weighted T_eff), `est.shrinkage_intensity` (current α).
80
-
81
- To recalibrate `kappa` for a different asset class (requires the `full` extra):
82
-
83
- ```python
84
- kappa = SqueezeKernelEstimator.calibrate_kappa(burn_in_returns, target_weight=0.6)
85
- ```
86
-
87
- ## Advanced options
88
-
89
- **Adaptive scale-free correlation memory** (`corr_half_lives=(43, 173, 693)`, `corr_theta=0.25`): replaces the single correlation timescale with a positive combination of EWMAs on a geometric half-life ladder — each scale normalized and adaptively shrunk against its own effective sample size, then the covariances blended with weights resting at the prior ∝ half-life^`corr_theta`. By Bernstein's theorem the ladder approximates the power-law memory of financial correlations (the streaming analogue of HAR). For ladders of two or more rungs the blend weights are gated by a sequential surprise detector: a two-sided Page CUSUM on the studentized fast-vs-slow per-rung predictive-score drift (threshold set by Siegmund's average-run-length approximation at ~2 years, no tuned parameters) tilts the weights toward the fast or slow end of the ladder when one side accumulates statistically forced evidence, decaying back at the fastest rung's half-life. Weights equal the prior on all non-alarmed days, the blend stays convex, so PSD holds by construction; `None` (default) reproduces the published single-scale estimator exactly. This is the paper's **headline configuration**: on the S&P 500 benchmark it leads every tested method at every universe size (held-out one-step NLL −4.6 vs single-scale at n=100, −11.1 at n=300 before the cluster target), the 90% model confidence set collapses to it alone, and the detector's margin is confirmed out-of-time on an external industry panel (+0.53 NLL/day, p=1×10⁻⁴). Cost is O(K·n²) per update plus one Cholesky per rung per day for the detector scores. Composes with `shrinkage_target="cluster"`; mutually exclusive with `lambda_corr_fast`.
90
-
91
- ```python
92
- est = SqueezeKernelEstimator(n_assets=100, corr_half_lives=(43, 173, 693), corr_theta=0.25)
93
- # maximal variant at high dimension:
94
- est = SqueezeKernelEstimator(n_assets=300, corr_half_lives=(43, 173, 693), shrinkage_target="cluster")
95
- ```
96
-
97
- **Score-exact weighting** (`weight_statistic="mahalanobis"`, use with `kappa=1.0`): drives the kernel with the Mahalanobis surprise `z'C⁻¹z/N` against the estimator's own correlation instead of the marginal dispersion. Improves accuracy in the moderate-concentration regime — use only when `n / T_eff ≲ 0.5` (e.g. n ≤ 100 at the default `lambda_corr`); at higher concentration the estimated inverse degrades it and the default is strictly better.
98
-
99
- ```python
100
- est = SqueezeKernelEstimator(n_assets=100, kappa=1.0, weight_statistic="mahalanobis")
101
- ```
102
-
103
- **Score-driven memory** (`lambda_corr_fast=0.99`): lets stress days also *shorten* the correlation memory (decay slides from `lambda_corr` toward `lambda_corr_fast` as the kernel weight rises). Do **not** combine with the Mahalanobis option — they act on the same channel and the combination degrades accuracy.
104
-
105
- **OU volatility anchor** (`vol_anchor_phi=0.995`): mean-reverts each asset's variance prediction toward a slow per-asset anchor (a ~1000-day EWMA of squared returns) before the daily update — a two-timescale, component-style volatility structure. One global parameter with a clean interpretation (deviation half-life ≈ ln 2/(1−φ) days; φ=0.995 ≈ 139 d). On the S&P 500 n=100 benchmark this improved held-out one-step NLL by 3.3 points (4.3 at φ=0.99) and five-step NLL by 3.9 (5.0), with no degradation at n=300. `None` (default) or φ=1 reproduces the published estimator exactly.
106
-
107
- ```python
108
- est = SqueezeKernelEstimator(n_assets=100, vol_anchor_phi=0.995)
109
- ```
110
-
111
- **Cluster shrinkage target** (`shrinkage_target="cluster"`): generalizes the equicorrelation shrinkage target to respect the correlation matrix's own block/cluster structure — with **no clustering algorithm**. The target morphs with the shrinkage intensity, T = (1−α)·T_equi + α·[(1−γ)I + γ·(C∘C)], where C∘C is the Hadamard square of the current correlation (positive semi-definite by the Schur product theorem; entries are pairwise shared-variance fractions) and γ is level-matched automatically. Zero added parameters, still O(n²), and as α→0 it reduces exactly to the default estimator. Held-out one-step NLL on the S&P 500 benchmark: ±0.1 at n=100, **−4.2 at n=200, −25.0 at n=300** — recommended whenever the universe size approaches the effective sample size.
112
-
113
- ```python
114
- est = SqueezeKernelEstimator(n_assets=300, shrinkage_target="cluster")
115
- ```
116
-
117
- **New-listing usability gate** (`min_obs=60`): on an expanding universe, an asset's forecast rows are dominated by its single-observation variance initialization for its first weeks of life and are unusable for scoring or portfolio construction (measured ≈ +2,800 NLL/day on days whose scored set included such assets, on a 42-instrument multi-asset panel). `min_obs` gates nothing inside the estimator — states warm normally, all outputs are unchanged — it exposes a `usable_mask` property marking assets with at least `min_obs` finite observations, so deployments subset with it:
118
-
119
- ```python
120
- est = SqueezeKernelEstimator(n_assets=42, min_obs=60)
121
- # ... update loop ...
122
- m = est.usable_mask
123
- cov_usable = est.get_cov()[np.ix_(m, m)]
124
- ```
125
-
126
- **Alternative kernels**: pass `kernel_fn=kernel_exponential` (with `kernel_kwargs={"gamma": ...}`) or `kernel_chi2_cdf`, or any callable `(d2, *, n_observed, **kw) -> float` mapping to `[0, 1)`. The PSD guarantee holds for any such kernel.
127
-
128
- ## How it works
129
-
130
- Three mechanisms in one recursion:
131
-
132
- 1. **Dual-timescale EWMA** — fast per-asset volatility (`lambda_vol`) is separated from slow correlation dynamics (`lambda_corr`), so variance shocks don't contaminate the correlation estimate.
133
- 2. **Fisher kernel weighting** — each day's standardized outer product enters with weight `w = d²/(d² + kappa)`, where `d²` is the mean squared standardized return: calm days contribute little, dispersion shocks contribute fully.
134
- 3. **Adaptive equicorrelation shrinkage** — `alpha = min(1, max(0, n/(2·S) − delta))` blends toward an equicorrelation target using the estimator's own kernel-weighted sample size `S`; it is a no-op at low dimension and provides provably bounded conditioning at high dimension.
135
-
136
- The complete update is a natural-gradient step on the Gaussian log-likelihood, with the kernel weight acting as an adaptive Riemannian learning rate (paper, Appendix B).
137
-
138
- ## Development
139
-
140
- ```bash
141
- uv sync --extra full --extra dev
142
- uv run python -m pytest # test suite
143
- uv run python -m ruff check . # lint
144
- uv build # build sdist + wheel
145
- ```
146
-
147
- Releases: publishing a GitHub release from a `v*` tag triggers the [publish workflow](.github/workflows/publish.yml), which builds and uploads to PyPI via trusted publishing.
148
-
149
- ## Citation
150
-
151
- ```bibtex
152
- @article{kende2026squeeze,
153
- title = {The Squeeze Kernel Covariance Estimator: Dual-Timescale Tracking with Adaptive Shrinkage},
154
- author = {Kende, Robert},
155
- year = {2026},
156
- note = {Available at SSRN: \url{https://ssrn.com/abstract=6455918}}
157
- }
158
- ```
159
-
160
- See also [`CITATION.cff`](CITATION.cff).
161
-
162
- ## License
163
-
164
- MIT