squeeze-kernel 0.3.0__tar.gz → 0.5.0__tar.gz

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -1,6 +1,6 @@
1
1
  Metadata-Version: 2.4
2
2
  Name: squeeze-kernel
3
- Version: 0.3.0
3
+ Version: 0.5.0
4
4
  Summary: Streaming, PSD-by-construction covariance estimator with Fisher-kernel weighting and adaptive shrinkage
5
5
  Keywords: covariance,correlation,ewma,kernel,risk,streaming
6
6
  Author: Robert Kende
@@ -37,7 +37,7 @@ Description-Content-Type: text/markdown
37
37
 
38
38
  A **streaming covariance estimator for panels of daily financial returns**. One `O(n²)` update per day, positive semi-definite **by construction** at every step, missing values handled **natively**, and defaults that require no tuning. Only dependency: NumPy.
39
39
 
40
- Reference: *"The Squeeze Kernel Covariance Estimator: Dual-Timescale Tracking with Adaptive Shrinkage"* (Kende, 2026).
40
+ Reference: *"The Squeeze Kernel Covariance Estimator: Dual-Timescale Tracking with Adaptive Shrinkage"* (Kende, 2026) — [SSRN abstract 6455918](https://ssrn.com/abstract=6455918).
41
41
 
42
42
  ## Why
43
43
 
@@ -116,6 +116,14 @@ kappa = SqueezeKernelEstimator.calibrate_kappa(burn_in_returns, target_weight=0.
116
116
 
117
117
  ## Advanced options
118
118
 
119
+ **Scale-free correlation memory** (`corr_half_lives=(43, 173, 693)`, `corr_theta=0.25`): replaces the single correlation timescale with a positive combination of EWMAs on a geometric half-life ladder — each scale normalized and adaptively shrunk against its own effective sample size, then the covariances blended with weights ∝ half-life^`corr_theta`. By Bernstein's theorem this approximates the power-law memory of financial correlations (the streaming analogue of HAR); positive weights keep it PSD by construction, and `None` (default) reproduces the published single-scale estimator exactly. This is the paper's **headline configuration**: on the S&P 500 benchmark it leads every tested method at every universe size (held-out one-step NLL −3.9 at n=100, up to −18 at n=300 before the cluster target), and the 90% model confidence set collapses to it alone. Cost is O(K·n²) per update. Composes with `shrinkage_target="cluster"`; mutually exclusive with `lambda_corr_fast`.
120
+
121
+ ```python
122
+ est = SqueezeKernelEstimator(n_assets=100, corr_half_lives=(43, 173, 693), corr_theta=0.25)
123
+ # maximal variant at high dimension:
124
+ est = SqueezeKernelEstimator(n_assets=300, corr_half_lives=(43, 173, 693), shrinkage_target="cluster")
125
+ ```
126
+
119
127
  **Score-exact weighting** (`weight_statistic="mahalanobis"`, use with `kappa=1.0`): drives the kernel with the Mahalanobis surprise `z'C⁻¹z/N` against the estimator's own correlation instead of the marginal dispersion. Improves accuracy in the moderate-concentration regime — use only when `n / T_eff ≲ 0.5` (e.g. n ≤ 100 at the default `lambda_corr`); at higher concentration the estimated inverse degrades it and the default is strictly better.
120
128
 
121
129
  ```python
@@ -130,6 +138,12 @@ est = SqueezeKernelEstimator(n_assets=100, kappa=1.0, weight_statistic="mahalano
130
138
  est = SqueezeKernelEstimator(n_assets=100, vol_anchor_phi=0.995)
131
139
  ```
132
140
 
141
+ **Cluster shrinkage target** (`shrinkage_target="cluster"`): generalizes the equicorrelation shrinkage target to respect the correlation matrix's own block/cluster structure — with **no clustering algorithm**. The target morphs with the shrinkage intensity, T = (1−α)·T_equi + α·[(1−γ)I + γ·(C∘C)], where C∘C is the Hadamard square of the current correlation (positive semi-definite by the Schur product theorem; entries are pairwise shared-variance fractions) and γ is level-matched automatically. Zero added parameters, still O(n²), and as α→0 it reduces exactly to the default estimator. Held-out one-step NLL on the S&P 500 benchmark: ±0.1 at n=100, **−4.2 at n=200, −25.0 at n=300** — recommended whenever the universe size approaches the effective sample size.
142
+
143
+ ```python
144
+ est = SqueezeKernelEstimator(n_assets=300, shrinkage_target="cluster")
145
+ ```
146
+
133
147
  **Alternative kernels**: pass `kernel_fn=kernel_exponential` (with `kernel_kwargs={"gamma": ...}`) or `kernel_chi2_cdf`, or any callable `(d2, *, n_observed, **kw) -> float` mapping to `[0, 1)`. The PSD guarantee holds for any such kernel.
134
148
 
135
149
  ## How it works
@@ -159,7 +173,8 @@ Releases: publishing a GitHub release from a `v*` tag triggers the [publish work
159
173
  @article{kende2026squeeze,
160
174
  title = {The Squeeze Kernel Covariance Estimator: Dual-Timescale Tracking with Adaptive Shrinkage},
161
175
  author = {Kende, Robert},
162
- year = {2026}
176
+ year = {2026},
177
+ note = {Available at SSRN: \url{https://ssrn.com/abstract=6455918}}
163
178
  }
164
179
  ```
165
180
 
@@ -7,7 +7,7 @@
7
7
 
8
8
  A **streaming covariance estimator for panels of daily financial returns**. One `O(n²)` update per day, positive semi-definite **by construction** at every step, missing values handled **natively**, and defaults that require no tuning. Only dependency: NumPy.
9
9
 
10
- Reference: *"The Squeeze Kernel Covariance Estimator: Dual-Timescale Tracking with Adaptive Shrinkage"* (Kende, 2026).
10
+ Reference: *"The Squeeze Kernel Covariance Estimator: Dual-Timescale Tracking with Adaptive Shrinkage"* (Kende, 2026) — [SSRN abstract 6455918](https://ssrn.com/abstract=6455918).
11
11
 
12
12
  ## Why
13
13
 
@@ -86,6 +86,14 @@ kappa = SqueezeKernelEstimator.calibrate_kappa(burn_in_returns, target_weight=0.
86
86
 
87
87
  ## Advanced options
88
88
 
89
+ **Scale-free correlation memory** (`corr_half_lives=(43, 173, 693)`, `corr_theta=0.25`): replaces the single correlation timescale with a positive combination of EWMAs on a geometric half-life ladder — each scale normalized and adaptively shrunk against its own effective sample size, then the covariances blended with weights ∝ half-life^`corr_theta`. By Bernstein's theorem this approximates the power-law memory of financial correlations (the streaming analogue of HAR); positive weights keep it PSD by construction, and `None` (default) reproduces the published single-scale estimator exactly. This is the paper's **headline configuration**: on the S&P 500 benchmark it leads every tested method at every universe size (held-out one-step NLL −3.9 at n=100, up to −18 at n=300 before the cluster target), and the 90% model confidence set collapses to it alone. Cost is O(K·n²) per update. Composes with `shrinkage_target="cluster"`; mutually exclusive with `lambda_corr_fast`.
90
+
91
+ ```python
92
+ est = SqueezeKernelEstimator(n_assets=100, corr_half_lives=(43, 173, 693), corr_theta=0.25)
93
+ # maximal variant at high dimension:
94
+ est = SqueezeKernelEstimator(n_assets=300, corr_half_lives=(43, 173, 693), shrinkage_target="cluster")
95
+ ```
96
+
89
97
  **Score-exact weighting** (`weight_statistic="mahalanobis"`, use with `kappa=1.0`): drives the kernel with the Mahalanobis surprise `z'C⁻¹z/N` against the estimator's own correlation instead of the marginal dispersion. Improves accuracy in the moderate-concentration regime — use only when `n / T_eff ≲ 0.5` (e.g. n ≤ 100 at the default `lambda_corr`); at higher concentration the estimated inverse degrades it and the default is strictly better.
90
98
 
91
99
  ```python
@@ -100,6 +108,12 @@ est = SqueezeKernelEstimator(n_assets=100, kappa=1.0, weight_statistic="mahalano
100
108
  est = SqueezeKernelEstimator(n_assets=100, vol_anchor_phi=0.995)
101
109
  ```
102
110
 
111
+ **Cluster shrinkage target** (`shrinkage_target="cluster"`): generalizes the equicorrelation shrinkage target to respect the correlation matrix's own block/cluster structure — with **no clustering algorithm**. The target morphs with the shrinkage intensity, T = (1−α)·T_equi + α·[(1−γ)I + γ·(C∘C)], where C∘C is the Hadamard square of the current correlation (positive semi-definite by the Schur product theorem; entries are pairwise shared-variance fractions) and γ is level-matched automatically. Zero added parameters, still O(n²), and as α→0 it reduces exactly to the default estimator. Held-out one-step NLL on the S&P 500 benchmark: ±0.1 at n=100, **−4.2 at n=200, −25.0 at n=300** — recommended whenever the universe size approaches the effective sample size.
112
+
113
+ ```python
114
+ est = SqueezeKernelEstimator(n_assets=300, shrinkage_target="cluster")
115
+ ```
116
+
103
117
  **Alternative kernels**: pass `kernel_fn=kernel_exponential` (with `kernel_kwargs={"gamma": ...}`) or `kernel_chi2_cdf`, or any callable `(d2, *, n_observed, **kw) -> float` mapping to `[0, 1)`. The PSD guarantee holds for any such kernel.
104
118
 
105
119
  ## How it works
@@ -129,7 +143,8 @@ Releases: publishing a GitHub release from a `v*` tag triggers the [publish work
129
143
  @article{kende2026squeeze,
130
144
  title = {The Squeeze Kernel Covariance Estimator: Dual-Timescale Tracking with Adaptive Shrinkage},
131
145
  author = {Kende, Robert},
132
- year = {2026}
146
+ year = {2026},
147
+ note = {Available at SSRN: \url{https://ssrn.com/abstract=6455918}}
133
148
  }
134
149
  ```
135
150
 
@@ -4,7 +4,7 @@ build-backend = "uv_build"
4
4
 
5
5
  [project]
6
6
  name = "squeeze-kernel"
7
- version = "0.3.0"
7
+ version = "0.5.0"
8
8
  description = "Streaming, PSD-by-construction covariance estimator with Fisher-kernel weighting and adaptive shrinkage"
9
9
  readme = "README.md"
10
10
  license = "MIT"
@@ -34,4 +34,4 @@ __all__ = [
34
34
  "kernel_chi2_cdf",
35
35
  ]
36
36
 
37
- __version__ = "0.3.0"
37
+ __version__ = "0.5.0"
@@ -2,12 +2,17 @@
2
2
 
3
3
  from __future__ import annotations
4
4
 
5
+ from typing import TYPE_CHECKING
6
+
5
7
  import numpy as np
6
8
 
7
9
  from squeeze_kernel.kernels import (
8
10
  KernelFn, kernel_fisher, calibrate_kappa, extract_d2_series,
9
11
  )
10
12
 
13
+ if TYPE_CHECKING:
14
+ from collections.abc import Sequence
15
+
11
16
 
12
17
  class SqueezeKernelEstimator:
13
18
  """Streaming robust covariance estimator with pluggable kernel weighting.
@@ -93,6 +98,42 @@ class SqueezeKernelEstimator:
93
98
  Decay of the slow per-asset variance anchor (default 0.999,
94
99
  effective memory ≈ 1000 trading days). Only used when
95
100
  ``vol_anchor_phi`` is set.
101
+ shrinkage_target : str
102
+ Geometry of the adaptive shrinkage target. ``'equicorrelation'``
103
+ (default) is the published single-factor target. ``'cluster'``
104
+ uses the concentration-morphing cluster target
105
+ T = (1−α)·T_equi + α·[(1−γ)I + γ·(C∘C)], where C∘C is the Hadamard
106
+ square of the current correlation (PSD by the Schur product
107
+ theorem) and γ = min(1, ρ̄/mean-offdiag(C∘C)) level-matches the
108
+ target to the equicorrelation mass. Respects the correlation
109
+ matrix's own block/cluster structure without any clustering
110
+ algorithm; adds no parameters and stays O(n²). As α → 0 it
111
+ reduces exactly to the published estimator, so behaviour at low
112
+ concentration is unchanged. Held-out one-step NLL on the S&P-500
113
+ benchmark: +0.14 (negligible) at n=100, −4.2 at n=200, −25.0 at
114
+ n=300. Recommended when n approaches the effective sample size.
115
+ corr_half_lives : sequence of float, optional
116
+ Scale-free correlation memory. When set, the single correlation
117
+ timescale is replaced by a positive combination of EWMAs on the
118
+ given geometric half-life ladder (in trading days, e.g.
119
+ ``(43, 173, 693)``): each scale is normalised and adaptively
120
+ shrunk against its own effective sample size, and the resulting
121
+ covariances are blended with weights ∝ half-life\\ :sup:`corr_theta`.
122
+ By Bernstein's theorem this approximates the power-law memory of
123
+ financial correlations (the streaming analogue of HAR). PSD by
124
+ construction (positive combination of PSD matrices). ``None``
125
+ (default) is the published single-scale estimator, bit-for-bit;
126
+ a one-element ladder reduces to a single-scale estimator at that
127
+ half-life. Mutually exclusive with ``lambda_corr_fast``; composes
128
+ with ``shrinkage_target='cluster'``. Cost is O(K·n²) per update.
129
+ Held-out one-step NLL on the S&P-500 benchmark improves at every
130
+ dimension (−3.9 at n=100, up to −18 at n=300 before the cluster
131
+ target); the $90\\%$ model confidence set collapses to this
132
+ configuration alone. Recommended default: the base-centred ladder
133
+ ``(43, 173, 693)`` with ``corr_theta=0.25``.
134
+ corr_theta : float
135
+ Long-memory exponent controlling the ladder weights (default
136
+ 0.25). Only used when ``corr_half_lives`` is set.
96
137
 
97
138
  Examples
98
139
  --------
@@ -123,6 +164,9 @@ class SqueezeKernelEstimator:
123
164
  lambda_corr_fast: float | None = None,
124
165
  vol_anchor_phi: float | None = None,
125
166
  vol_anchor_decay: float = 0.999,
167
+ shrinkage_target: str = "equicorrelation",
168
+ corr_half_lives: "Sequence[float] | None" = None,
169
+ corr_theta: float = 0.25,
126
170
  ):
127
171
  self.n_assets = n_assets
128
172
  self.lambda_vol = lambda_vol
@@ -145,6 +189,37 @@ class SqueezeKernelEstimator:
145
189
  raise ValueError("vol_anchor_decay must be in (0, 1).")
146
190
  self.vol_anchor_phi = vol_anchor_phi
147
191
  self.vol_anchor_decay = vol_anchor_decay
192
+ if shrinkage_target not in ("equicorrelation", "cluster"):
193
+ raise ValueError("shrinkage_target must be 'equicorrelation' or 'cluster'.")
194
+ self.shrinkage_target = shrinkage_target
195
+
196
+ # Scale-free correlation memory (opt-in): replace the single correlation
197
+ # timescale by a positive combination of EWMAs on a geometric half-life
198
+ # ladder, blended per-scale (Mode A). None => single-scale, published
199
+ # behaviour bit-for-bit. See ``corr_half_lives`` in the class docstring.
200
+ self.corr_half_lives = None
201
+ self.corr_theta = corr_theta
202
+ self._corr_lam: np.ndarray | None = None
203
+ self._corr_w: np.ndarray | None = None
204
+ self._M_list: list[np.ndarray] | None = None
205
+ self._S_list: list[float] | None = None
206
+ if corr_half_lives is not None:
207
+ hl = np.asarray(corr_half_lives, dtype=np.float64)
208
+ if hl.ndim != 1 or hl.size < 1 or np.any(hl <= 0.0):
209
+ raise ValueError("corr_half_lives must be a non-empty sequence of positive half-lives.")
210
+ if corr_theta < 0.0:
211
+ raise ValueError("corr_theta must be >= 0.")
212
+ if lambda_corr_fast is not None:
213
+ raise ValueError(
214
+ "corr_half_lives and lambda_corr_fast are mutually exclusive "
215
+ "correlation-memory mechanisms; set at most one."
216
+ )
217
+ self.corr_half_lives = hl
218
+ self._corr_lam = 2.0 ** (-1.0 / hl)
219
+ w = hl ** corr_theta
220
+ self._corr_w = w / w.sum()
221
+ self._M_list = [np.eye(n_assets, dtype=np.float64) * epsilon for _ in hl]
222
+ self._S_list = [float(epsilon) for _ in hl]
148
223
 
149
224
  # Resolve shrinkage
150
225
  if isinstance(shrinkage, str):
@@ -259,22 +334,45 @@ class SqueezeKernelEstimator:
259
334
  if self.impute_missing and 0 < n_obs < n:
260
335
  self._impute(z_t, finite)
261
336
 
262
- # ── Correlation EWMA (in-place to avoid per-step allocations) ──
263
- lam_c = self.lambda_corr
264
- if self.lambda_corr_fast is not None:
265
- # Score-driven memory: stress days (w_t → 1) shorten the memory
266
- # toward lambda_corr_fast; calm days keep the slow decay.
267
- lam_c = self.lambda_corr + (self.lambda_corr_fast - self.lambda_corr) * w_t
268
- self._S_t = lam_c * self._S_t + w_t
269
- self._M_t *= lam_c
270
- if w_t > 0.0 and n_obs > 0:
271
- # np.multiply.outer with out= avoids the temporary that
272
- # np.outer otherwise allocates each step.
273
- np.multiply.outer(z_t, z_t, out=self._scratch_outer)
274
- self._M_t += w_t * self._scratch_outer
275
-
276
- # ── Extract covariance ──
277
- self._cov, self._corr = self._extract(vol_t)
337
+ if self._corr_lam is None:
338
+ # ── Single-scale correlation EWMA (published path, unchanged) ──
339
+ lam_c = self.lambda_corr
340
+ if self.lambda_corr_fast is not None:
341
+ # Score-driven memory: stress days (w_t → 1) shorten the memory
342
+ # toward lambda_corr_fast; calm days keep the slow decay.
343
+ lam_c = self.lambda_corr + (self.lambda_corr_fast - self.lambda_corr) * w_t
344
+ self._S_t = lam_c * self._S_t + w_t
345
+ self._M_t *= lam_c
346
+ if w_t > 0.0 and n_obs > 0:
347
+ # np.multiply.outer with out= avoids the temporary that
348
+ # np.outer otherwise allocates each step.
349
+ np.multiply.outer(z_t, z_t, out=self._scratch_outer)
350
+ self._M_t += w_t * self._scratch_outer
351
+ self._cov, self._corr = self._extract(vol_t)
352
+ else:
353
+ # ── Scale-free ladder (Mode A) ──
354
+ # Update K correlation accumulators on the geometric half-life
355
+ # ladder; normalise and adaptively shrink each against its OWN
356
+ # effective sample size, then blend the per-scale covariances.
357
+ add = w_t > 0.0 and n_obs > 0
358
+ if add:
359
+ np.multiply.outer(z_t, z_t, out=self._scratch_outer)
360
+ cov = np.zeros((n, n), dtype=np.float64)
361
+ s_eff = 0.0
362
+ for k in range(self._corr_lam.size):
363
+ self._S_list[k] = self._corr_lam[k] * self._S_list[k] + w_t
364
+ self._M_list[k] *= self._corr_lam[k]
365
+ if add:
366
+ self._M_list[k] += w_t * self._scratch_outer
367
+ cov_k, _ = self._extract(vol_t, self._M_list[k], self._S_list[k])
368
+ cov += self._corr_w[k] * cov_k
369
+ s_eff += self._corr_w[k] * self._S_list[k]
370
+ cov = 0.5 * (cov + cov.T)
371
+ self._cov = cov
372
+ self._S_t = s_eff # blended effective size (for the property)
373
+ d = np.sqrt(np.maximum(np.diagonal(cov), eps))
374
+ self._corr = cov / np.outer(d, d)
375
+ np.fill_diagonal(self._corr, 1.0)
278
376
  self._last_weight = w_t
279
377
  return w_t
280
378
 
@@ -355,22 +453,23 @@ class SqueezeKernelEstimator:
355
453
  if den > 0.0:
356
454
  z_t[i] = num / den
357
455
 
358
- def _extract(self, vol_t: np.ndarray) -> tuple[np.ndarray, np.ndarray]:
456
+ def _extract(self, vol_t: np.ndarray, M_t=None, S_in=None) -> tuple[np.ndarray, np.ndarray]:
359
457
  eps = self.epsilon
360
- S_t = max(self._S_t, eps)
458
+ M_t = self._M_t if M_t is None else M_t
459
+ S_t = max(self._S_t if S_in is None else S_in, eps)
361
460
  n = self.n_assets
362
461
 
363
462
  # Normalised standardised covariance matrix sigma_z = M_t / S_t.
364
463
  # Compute correlations directly into the cached scratch buffer to
365
464
  # avoid two intermediate allocations (sigma_z and corr).
366
- diag_z = np.diagonal(self._M_t).copy()
465
+ diag_z = np.diagonal(M_t).copy()
367
466
  diag_z /= S_t # in-place
368
467
  inv_diag = 1.0 / np.sqrt(np.maximum(diag_z, eps))
369
468
  # corr_ij = (M_ij / S_t) * inv_diag_i * inv_diag_j; this writes
370
469
  # the rescaled outer-product into _scratch_corr in one pass.
371
470
  np.multiply.outer(inv_diag, inv_diag, out=self._scratch_corr)
372
471
  corr = self._scratch_corr
373
- corr *= self._M_t # in-place; corr = sigma_z * outer(inv_diag,inv_diag)
472
+ corr *= M_t # in-place; corr = sigma_z * outer(inv_diag,inv_diag)
374
473
  corr *= (1.0 / S_t) # absorb the M_t / S_t scale
375
474
  np.fill_diagonal(corr, np.where(diag_z > eps, 1.0, 0.0))
376
475
 
@@ -384,9 +483,25 @@ class SqueezeKernelEstimator:
384
483
  if alpha > 0.0 and n > 1:
385
484
  # Off-diagonal mean: O(n^2) sum, no mask allocation.
386
485
  rho_bar = (corr.sum() - corr.trace()) / self._n_off
387
- corr *= (1.0 - alpha)
388
- corr += alpha * rho_bar
389
- np.fill_diagonal(corr, 1.0)
486
+ if self.shrinkage_target == "equicorrelation" or rho_bar <= 0.0:
487
+ corr *= (1.0 - alpha)
488
+ corr += alpha * rho_bar
489
+ np.fill_diagonal(corr, 1.0)
490
+ else:
491
+ # Cluster (concentration-morphing) target:
492
+ # T = (1-alpha) T_equi + alpha [(1-gamma) I + gamma (C o C)]
493
+ # C o C is the Hadamard square of the raw correlation (PSD by
494
+ # the Schur product theorem, unit diagonal for free); gamma is
495
+ # level-matched so the target carries the same average
496
+ # correlation mass as the equicorrelation target. As
497
+ # alpha -> 0 this reduces exactly to the published estimator.
498
+ had = corr * corr # Hadamard square, O(n^2)
499
+ mean_off = (had.sum() - np.trace(had)) / self._n_off
500
+ gamma = min(1.0, rho_bar / max(mean_off, eps))
501
+ corr *= (1.0 - alpha)
502
+ corr += (alpha * (1.0 - alpha)) * rho_bar
503
+ corr += (alpha * alpha * gamma) * had
504
+ np.fill_diagonal(corr, 1.0)
390
505
 
391
506
  # Build covariance from corr and vol_t. We need an output array
392
507
  # that the caller can keep, so allocate one cov here (cannot reuse