canli-validation-mcp 0.9.1 → 0.10.1
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +30 -9
- package/package.json +1 -1
- package/src/local/js/moments-core.js +16 -2
- package/src/local/js/snooping-core.js +185 -0
- package/src/local/js/validate/haircut-sharpe.js +2 -1
- package/src/local/js/validate/overfitting.js +7 -1
- package/src/local/js/validate/reality-check.js +89 -0
- package/src/local.mjs +2 -0
- package/src/schemas.mjs +25 -0
- package/src/server.mjs +21 -2
package/README.md
CHANGED
|
@@ -7,8 +7,9 @@
|
|
|
7
7
|
|
|
8
8
|
**Is your best backtest real, or just the luckiest of the variants you tried?** This MCP server
|
|
9
9
|
lets Claude, Cursor or any MCP client answer that with the standard corrections: the deflated
|
|
10
|
-
Sharpe ratio, the CSCV probability of backtest overfitting,
|
|
11
|
-
|
|
10
|
+
Sharpe ratio, the CSCV probability of backtest overfitting, data-snooping tests on every variant
|
|
11
|
+
tried (Hansen's SPA, White's Reality Check, Romano-Wolf StepM), the minimum track record length,
|
|
12
|
+
the haircut Sharpe ratio, and luck-equivalent trials. Free and MIT-licensed.
|
|
12
13
|
|
|
13
14
|
## Quick start
|
|
14
15
|
|
|
@@ -34,9 +35,10 @@ other stdio client (Cursor, VS Code) run the same `npx` command; see "Claude Des
|
|
|
34
35
|
## What it does
|
|
35
36
|
|
|
36
37
|
An MCP (Model Context Protocol) server over canlicapital.com's free, keyed validation API. It
|
|
37
|
-
gives a coding agent
|
|
38
|
-
CSCV overfitting,
|
|
39
|
-
|
|
38
|
+
gives a coding agent fifteen tools: issue a free key, run the nine validators (deflated Sharpe,
|
|
39
|
+
CSCV overfitting, data-snooping tests (Hansen's SPA, White's Reality Check and Romano-Wolf StepM),
|
|
40
|
+
paper-evidence conformance, breadth ceiling, minimum track record length, minimum backtest length,
|
|
41
|
+
haircut Sharpe ratio, luck-equivalent trials),
|
|
40
42
|
audit one backtest with three of them in a single call, fetch a stored receipt, verify a
|
|
41
43
|
receipt's signature offline, read service status, and read a company's reported financial history
|
|
42
44
|
from SEC filings. Every validation result carries, beside the number, the sentences that say what
|
|
@@ -80,6 +82,7 @@ repository for the full design.
|
|
|
80
82
|
| `validate_backtest_length` | `POST /api/v1/validate/backtest-length` | yes |
|
|
81
83
|
| `validate_haircut_sharpe` | `POST /api/v1/validate/haircut-sharpe` | yes |
|
|
82
84
|
| `validate_luck_trials` | `POST /api/v1/validate/luck-trials` | yes |
|
|
85
|
+
| `validate_reality_check` | `POST /api/v1/validate/reality-check` (a `matrix_file` is read on your machine and sent as numbers) | yes |
|
|
83
86
|
| `audit_backtest` | the deflated Sharpe, track record and, with `variants`, overfitting routes, one validation each | yes |
|
|
84
87
|
| `verify_receipt` | `GET /api/v1/receipts/{id}` when given an id; the checks run locally | no |
|
|
85
88
|
| `get_receipt` | `GET /api/v1/receipts/{id}` | no |
|
|
@@ -97,6 +100,24 @@ repository for the full design.
|
|
|
97
100
|
Sending fields from both shapes, or from neither, is rejected before any request leaves the
|
|
98
101
|
process; see `src/schemas.mjs`.
|
|
99
102
|
|
|
103
|
+
`validate_reality_check` takes the returns of every variant the search tried, one row per period
|
|
104
|
+
and one column per variant, as `matrix` or as `matrix_file` (a CSV or JSON on your machine), and an
|
|
105
|
+
optional `benchmark` series; with none, variants are tested against zero. It runs three tests on one
|
|
106
|
+
seeded stationary bootstrap:
|
|
107
|
+
|
|
108
|
+
- Hansen's SPA: the chance that the best variant's studentized excess return is this good if no
|
|
109
|
+
variant has an edge, with lower and upper bounds (`spa.p_value`, the consistent p-value, is the
|
|
110
|
+
headline, and `monte_carlo_se` its sampling error);
|
|
111
|
+
- White's Reality Check: the same question without studentizing, so one volatile variant can
|
|
112
|
+
dominate it;
|
|
113
|
+
- Romano and Wolf's StepM: which variants beat the benchmark, with the familywise error held at
|
|
114
|
+
`alpha`.
|
|
115
|
+
|
|
116
|
+
A result reproduces exactly from its `seed`. On three fixed-seed cases the p-values agree with
|
|
117
|
+
Python's `arch` 8.0 and with a numpy transcription of Hansen's formulas within Monte Carlo error, and
|
|
118
|
+
the StepM sets match `arch`'s (`js/snooping-core.test.js`). Only the variants sent are counted: a
|
|
119
|
+
search that tried more than it sends makes luck look smaller than it was.
|
|
120
|
+
|
|
100
121
|
`company_financial_history` is different from the other tools: it reads the public company
|
|
101
122
|
reference at canlicapital.com, not the validation API. Give it a CIK (1 to 10 digits) to list a
|
|
102
123
|
company's available financial histories, or a CIK and a us-gaap concept such as `Revenues` to
|
|
@@ -317,7 +338,7 @@ On a breadth result this is about half the text. Set `CANLI_FULL_ENVELOPE=1` to
|
|
|
317
338
|
## Toolsets (tokens)
|
|
318
339
|
|
|
319
340
|
A client sends the model the whole tool list on every turn, and it is most of each turn's prompt:
|
|
320
|
-
a validation result is a few hundred tokens, the list of all
|
|
341
|
+
a validation result is a few hundred tokens, the list of all fifteen tools several thousand. A
|
|
321
342
|
client that needs one kind of tool can list only that kind, with `CANLI_TOOLSETS` (stdio) or
|
|
322
343
|
`?toolsets=` (hosted endpoint). The default is every tool.
|
|
323
344
|
|
|
@@ -333,9 +354,9 @@ tokenizers give different absolute counts), in the shape an OpenAI-style client
|
|
|
333
354
|
|
|
334
355
|
| CANLI_TOOLSETS | tools | tokens per turn | of all |
|
|
335
356
|
|---|---|---|---|
|
|
336
|
-
| `all` |
|
|
337
|
-
| `validate` |
|
|
338
|
-
| `receipts` | 2 | 312 |
|
|
357
|
+
| `all` | 15 | 4,194 | 100% |
|
|
358
|
+
| `validate` | 11 | 3,526 | 84% |
|
|
359
|
+
| `receipts` | 2 | 312 | 7% |
|
|
339
360
|
| `company` | 1 | 273 | 7% |
|
|
340
361
|
| `status` | 1 | 89 | 2% |
|
|
341
362
|
|
package/package.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "canli-validation-mcp",
|
|
3
|
-
"version": "0.
|
|
3
|
+
"version": "0.10.1",
|
|
4
4
|
"description": "Check whether a backtest is real: deflated Sharpe, CSCV probability of backtest overfitting, minimum track record and backtest length, haircut Sharpe and luck-equivalent trials, as an MCP server. Local mode, hosted endpoint, signed receipts.",
|
|
5
5
|
"keywords": [
|
|
6
6
|
"mcp",
|
|
@@ -7,9 +7,23 @@
|
|
|
7
7
|
// standards/validation-api/vectors.json.
|
|
8
8
|
import { calculateDsr } from "./dsr-core.js";
|
|
9
9
|
|
|
10
|
+
// JSON has no NaN, so a missing value arrives as null, and Number(null) is 0: read that way, a gap
|
|
11
|
+
// becomes a period of zero return, and a malformed string used to be dropped without a word. Each
|
|
12
|
+
// entry must be a finite number or a numeric string; anything else is refused by its position.
|
|
13
|
+
export function finiteNumbers(values, name) {
|
|
14
|
+
if (!Array.isArray(values)) throw new RangeError(`${name} must be an array of numbers`);
|
|
15
|
+
const out = new Array(values.length);
|
|
16
|
+
for (let i = 0; i < values.length; i += 1) {
|
|
17
|
+
const v = values[i];
|
|
18
|
+
const x = typeof v === "number" ? v : typeof v === "string" && v.trim() !== "" ? Number(v) : Number.NaN;
|
|
19
|
+
if (!Number.isFinite(x)) throw new RangeError(`${name}[${i}] is ${v === null ? "null" : v === undefined ? "missing" : typeof v === "number" ? String(v) : JSON.stringify(v)}, not a finite number: remove missing periods rather than sending them, since they would otherwise be read as zero`);
|
|
20
|
+
out[i] = x;
|
|
21
|
+
}
|
|
22
|
+
return out;
|
|
23
|
+
}
|
|
24
|
+
|
|
10
25
|
export function perPeriodMoments(returns) {
|
|
11
|
-
|
|
12
|
-
const x = returns.map(Number).filter((v) => Number.isFinite(v));
|
|
26
|
+
const x = finiteNumbers(returns, "returns");
|
|
13
27
|
const n = x.length;
|
|
14
28
|
if (n < 2) throw new RangeError("need at least 2 finite return observations");
|
|
15
29
|
let min = Infinity;
|
|
@@ -0,0 +1,185 @@
|
|
|
1
|
+
// =============================================================================
|
|
2
|
+
// snooping-core.js
|
|
3
|
+
// -----------------------------------------------------------------------------
|
|
4
|
+
// Data-snooping tests for a search: did the best of K variants beat the benchmark by more than the
|
|
5
|
+
// best of K would by luck? One stationary bootstrap (Politis and Romano 1994) feeds three tests:
|
|
6
|
+
//
|
|
7
|
+
// White's Reality Check (2000) the best mean excess return against the bootstrap distribution
|
|
8
|
+
// of the best recentred mean. Not studentized, so one volatile
|
|
9
|
+
// variant can dominate it.
|
|
10
|
+
// Hansen's SPA (2005) the best t-statistic, each variant studentized by the
|
|
11
|
+
// Politis-Romano variance of its mean. Three recentrings give a
|
|
12
|
+
// lower bound, a consistent p-value and an upper bound: poor
|
|
13
|
+
// variants are not allowed to inflate the p-value.
|
|
14
|
+
// Romano and Wolf's StepM (2005) step-down: which variants beat the benchmark, controlling the
|
|
15
|
+
// familywise error at alpha.
|
|
16
|
+
//
|
|
17
|
+
// d[t][k] is variant k's return minus the benchmark's in period t; higher is better. Every draw
|
|
18
|
+
// comes from makeRandom(seed), so a result reproduces exactly from its inputs. The reference
|
|
19
|
+
// implementation it is checked against (Python's arch 8.0, and a separate numpy transcription of
|
|
20
|
+
// Hansen's formulas for the studentized SPA) draws with numpy, so agreement is within Monte Carlo
|
|
21
|
+
// error, not to the digit; scripts/research/reality-check/ holds the comparison.
|
|
22
|
+
// =============================================================================
|
|
23
|
+
|
|
24
|
+
import { makeRandom } from "./selection-risk-core.js";
|
|
25
|
+
|
|
26
|
+
/** Indices of one stationary-bootstrap resample of n periods with mean block length `block`. */
|
|
27
|
+
export function stationaryIndices(n, block, next, out = new Int32Array(n)) {
|
|
28
|
+
const p = 1 / block;
|
|
29
|
+
let at = Math.floor(next() * n);
|
|
30
|
+
out[0] = at;
|
|
31
|
+
for (let t = 1; t < n; t += 1) {
|
|
32
|
+
at = next() < p ? Math.floor(next() * n) : at + 1 === n ? 0 : at + 1;
|
|
33
|
+
out[t] = at;
|
|
34
|
+
}
|
|
35
|
+
return out;
|
|
36
|
+
}
|
|
37
|
+
|
|
38
|
+
/** round(n^(1/3)), at least 1: the default mean block length. */
|
|
39
|
+
export const defaultBlock = (n) => Math.max(1, Math.round(Math.cbrt(n)));
|
|
40
|
+
|
|
41
|
+
// Weights below this are dropped from the variance sum. Each term they multiply is bounded by the
|
|
42
|
+
// sample variance, so what is dropped is under 2e-14 * n of it (4e-10 at 20,000 periods).
|
|
43
|
+
const KAPPA_FLOOR = 1e-14;
|
|
44
|
+
|
|
45
|
+
/**
|
|
46
|
+
* n times the variance of each column's mean under the stationary bootstrap, in closed form
|
|
47
|
+
* (Politis and Romano 1994, as Hansen 2005 uses it): gamma_0 + 2 * sum_i kappa_i * gamma_i, with
|
|
48
|
+
* kappa_i = (1 - i/n)(1 - p)^i + (i/n)(1 - p)^(n - i). Lags whose weight is below KAPPA_FLOOR are
|
|
49
|
+
* skipped, which keeps the cost near n * k * 30 * block rather than n^2 * k.
|
|
50
|
+
*/
|
|
51
|
+
export function bootstrapVariance(d, n, k, block, means) {
|
|
52
|
+
const p = 1 / block;
|
|
53
|
+
const out = new Float64Array(k);
|
|
54
|
+
const cross = (lag, weight) => {
|
|
55
|
+
for (let t = 0; t + lag < n; t += 1) {
|
|
56
|
+
const a = t * k;
|
|
57
|
+
const b = (t + lag) * k;
|
|
58
|
+
for (let j = 0; j < k; j += 1) out[j] += weight * (d[a + j] - means[j]) * (d[b + j] - means[j]);
|
|
59
|
+
}
|
|
60
|
+
};
|
|
61
|
+
cross(0, 1 / n);
|
|
62
|
+
const q = 1 - p;
|
|
63
|
+
for (let i = 1; i < n; i += 1) {
|
|
64
|
+
const kappa = (1 - i / n) * q ** i + (i / n) * q ** (n - i);
|
|
65
|
+
if (kappa < KAPPA_FLOOR) {
|
|
66
|
+
// The first term has decayed; the wrap-around term (i/n)(1 - p)^(n - i) only reaches the
|
|
67
|
+
// floor again once n - i <= log(floor) / log(1 - p). Skip to there.
|
|
68
|
+
const tail = q > 0 ? Math.log(KAPPA_FLOOR) / Math.log(q) : 0;
|
|
69
|
+
const jump = Math.floor(n - tail);
|
|
70
|
+
if (jump > i + 1) i = Math.min(jump, n) - 1;
|
|
71
|
+
continue;
|
|
72
|
+
}
|
|
73
|
+
cross(i, (2 * kappa) / n);
|
|
74
|
+
}
|
|
75
|
+
return out;
|
|
76
|
+
}
|
|
77
|
+
|
|
78
|
+
/**
|
|
79
|
+
* All three tests on one bootstrap. `d` is a row-major Float64Array of n periods by k variants.
|
|
80
|
+
* Returns means, the variance estimates, the bootstrap maxima under each recentring and the
|
|
81
|
+
* StepM rejections; p-values are shares of `reps` draws strictly above the observed statistic.
|
|
82
|
+
*/
|
|
83
|
+
export function snoopingTests(d, n, k, { block = defaultBlock(n), reps = 1000, seed = 42, alpha = 0.05, studentize = true } = {}) {
|
|
84
|
+
if (!(n >= 2 && k >= 1)) throw new RangeError("need at least 2 periods and 1 variant");
|
|
85
|
+
if (!(block >= 1 && block <= n)) throw new RangeError(`block_length must be from 1 to the number of periods (${n})`);
|
|
86
|
+
if (!Number.isInteger(reps) || reps < 100) throw new RangeError("reps must be an integer of at least 100");
|
|
87
|
+
if (!(alpha > 0 && alpha < 1)) throw new RangeError("alpha must be between 0 and 1");
|
|
88
|
+
const means = new Float64Array(k);
|
|
89
|
+
for (let t = 0; t < n; t += 1) for (let j = 0; j < k; j += 1) means[j] += d[t * k + j];
|
|
90
|
+
for (let j = 0; j < k; j += 1) means[j] /= n;
|
|
91
|
+
const variance = bootstrapVariance(d, n, k, block, means);
|
|
92
|
+
const sd = new Float64Array(k);
|
|
93
|
+
for (let j = 0; j < k; j += 1) sd[j] = Math.sqrt(Math.max(variance[j], 0));
|
|
94
|
+
// A variant whose excess return never varies has no sampling distribution to studentize by.
|
|
95
|
+
const usable = Array.from(sd, (s) => s > 0);
|
|
96
|
+
if (!usable.some(Boolean)) throw new RangeError("every variant's excess return is constant; there is nothing to test");
|
|
97
|
+
const root = Math.sqrt(n);
|
|
98
|
+
const scale = (j) => (studentize ? root / sd[j] : 1);
|
|
99
|
+
// Hansen's recentrings: lower keeps poor variants at zero, consistent drops only those far below
|
|
100
|
+
// zero, upper recentres everything (White's null, the least favourable configuration).
|
|
101
|
+
const threshold = (j) => -Math.sqrt((variance[j] / n) * 2 * Math.log(Math.log(n)));
|
|
102
|
+
const centre = [
|
|
103
|
+
Float64Array.from(means, (m) => Math.max(m, 0)),
|
|
104
|
+
Float64Array.from(means, (m, j) => (m >= threshold(j) ? m : 0)),
|
|
105
|
+
Float64Array.from(means),
|
|
106
|
+
];
|
|
107
|
+
let observed = Number.NEGATIVE_INFINITY;
|
|
108
|
+
let best = -1;
|
|
109
|
+
for (let j = 0; j < k; j += 1) {
|
|
110
|
+
if (!usable[j]) continue;
|
|
111
|
+
const s = means[j] * scale(j);
|
|
112
|
+
if (s > observed) { observed = s; best = j; }
|
|
113
|
+
}
|
|
114
|
+
let rcObserved = Number.NEGATIVE_INFINITY;
|
|
115
|
+
for (let j = 0; j < k; j += 1) if (usable[j]) rcObserved = Math.max(rcObserved, means[j]);
|
|
116
|
+
|
|
117
|
+
const next = makeRandom(seed);
|
|
118
|
+
const idx = new Int32Array(n);
|
|
119
|
+
const star = new Float64Array(k);
|
|
120
|
+
const boot = new Float64Array(reps * k); // bootstrap mean of each variant, per draw
|
|
121
|
+
for (let r = 0; r < reps; r += 1) {
|
|
122
|
+
stationaryIndices(n, block, next, idx);
|
|
123
|
+
star.fill(0);
|
|
124
|
+
for (let t = 0; t < n; t += 1) {
|
|
125
|
+
const row = idx[t] * k;
|
|
126
|
+
for (let j = 0; j < k; j += 1) star[j] += d[row + j];
|
|
127
|
+
}
|
|
128
|
+
for (let j = 0; j < k; j += 1) boot[r * k + j] = star[j] / n;
|
|
129
|
+
}
|
|
130
|
+
|
|
131
|
+
const spaStat = Math.max(observed, 0);
|
|
132
|
+
const exceed = [0, 0, 0];
|
|
133
|
+
let rcExceed = 0;
|
|
134
|
+
for (let r = 0; r < reps; r += 1) {
|
|
135
|
+
const maxima = [Number.NEGATIVE_INFINITY, Number.NEGATIVE_INFINITY, Number.NEGATIVE_INFINITY];
|
|
136
|
+
let rcMax = Number.NEGATIVE_INFINITY;
|
|
137
|
+
for (let j = 0; j < k; j += 1) {
|
|
138
|
+
if (!usable[j]) continue;
|
|
139
|
+
const m = boot[r * k + j];
|
|
140
|
+
for (let c = 0; c < 3; c += 1) maxima[c] = Math.max(maxima[c], (m - centre[c][j]) * scale(j));
|
|
141
|
+
rcMax = Math.max(rcMax, m - means[j]);
|
|
142
|
+
}
|
|
143
|
+
for (let c = 0; c < 3; c += 1) if (Math.max(maxima[c], 0) > spaStat) exceed[c] += 1;
|
|
144
|
+
if (rcMax > rcObserved) rcExceed += 1;
|
|
145
|
+
}
|
|
146
|
+
|
|
147
|
+
// StepM: reject every variant above the (1 - alpha) quantile of the bootstrap maximum over those
|
|
148
|
+
// not yet rejected (consistent recentring), remove them, and repeat until nothing more goes.
|
|
149
|
+
const rejected = new Set();
|
|
150
|
+
const rounds = [];
|
|
151
|
+
for (;;) {
|
|
152
|
+
const live = [];
|
|
153
|
+
for (let j = 0; j < k; j += 1) if (usable[j] && !rejected.has(j)) live.push(j);
|
|
154
|
+
if (!live.length) break;
|
|
155
|
+
const maxima = new Float64Array(reps);
|
|
156
|
+
for (let r = 0; r < reps; r += 1) {
|
|
157
|
+
let m = Number.NEGATIVE_INFINITY;
|
|
158
|
+
for (const j of live) m = Math.max(m, (boot[r * k + j] - centre[1][j]) * scale(j));
|
|
159
|
+
maxima[r] = m;
|
|
160
|
+
}
|
|
161
|
+
maxima.sort();
|
|
162
|
+
const critical = quantile(maxima, 1 - alpha);
|
|
163
|
+
const now = live.filter((j) => means[j] * scale(j) > critical);
|
|
164
|
+
if (!now.length) break;
|
|
165
|
+
for (const j of now) rejected.add(j);
|
|
166
|
+
rounds.push({ critical, rejected: now });
|
|
167
|
+
}
|
|
168
|
+
|
|
169
|
+
return {
|
|
170
|
+
n, k, block, reps, seed, alpha, studentize,
|
|
171
|
+
means, variance, usable,
|
|
172
|
+
best, observed,
|
|
173
|
+
spa: { statistic: spaStat, p_lower: exceed[0] / reps, p_consistent: exceed[1] / reps, p_upper: exceed[2] / reps },
|
|
174
|
+
reality_check: { statistic: rcObserved, p_value: rcExceed / reps },
|
|
175
|
+
stepm: { superior: [...rejected].sort((a, b) => a - b), rounds },
|
|
176
|
+
};
|
|
177
|
+
}
|
|
178
|
+
|
|
179
|
+
// numpy's default percentile (linear interpolation between order statistics), on sorted values.
|
|
180
|
+
function quantile(sorted, q) {
|
|
181
|
+
const h = (sorted.length - 1) * q;
|
|
182
|
+
const lo = Math.floor(h);
|
|
183
|
+
const hi = Math.min(lo + 1, sorted.length - 1);
|
|
184
|
+
return sorted[lo] + (h - lo) * (sorted[hi] - sorted[lo]);
|
|
185
|
+
}
|
|
@@ -1,6 +1,7 @@
|
|
|
1
1
|
// js/validate/haircut-sharpe.js
|
|
2
2
|
// The pure computation behind POST /api/v1/validate/haircut-sharpe, shared by the API route and the
|
|
3
3
|
// MCP package's local mode (mcp/src/local mirrors it byte for byte), so the two cannot disagree.
|
|
4
|
+
import { finiteNumbers } from "../moments-core.js";
|
|
4
5
|
import { haircutSharpe } from "../haircut-core.js";
|
|
5
6
|
|
|
6
7
|
// The haircut Sharpe ratio of Harvey and Liu, "Backtesting" (2015), checked against the authors'
|
|
@@ -13,7 +14,7 @@ const MAX_OTHERS = 10000;
|
|
|
13
14
|
export function compute(body) {
|
|
14
15
|
const missing = ["observed_sharpe_annualized", "periods_per_year", "observations"].filter((k) => body[k] === undefined);
|
|
15
16
|
if (missing.length) throw new RangeError(`Missing required fields: ${missing.join(", ")}`);
|
|
16
|
-
const others = body.other_sharpe_ratios_annualized;
|
|
17
|
+
const others = body.other_sharpe_ratios_annualized === undefined || !Array.isArray(body.other_sharpe_ratios_annualized) ? body.other_sharpe_ratios_annualized : finiteNumbers(body.other_sharpe_ratios_annualized, "other_sharpe_ratios_annualized");
|
|
17
18
|
if (others !== undefined && (!Array.isArray(others) || others.length < 1 || others.length > MAX_OTHERS)) {
|
|
18
19
|
throw new RangeError(`other_sharpe_ratios_annualized must be an array of 1 to ${MAX_OTHERS} Sharpe ratios`);
|
|
19
20
|
}
|
|
@@ -1,6 +1,7 @@
|
|
|
1
1
|
// js/validate/overfitting.js
|
|
2
2
|
// The pure computation behind POST /api/v1/validate/overfitting, shared by the API route and the
|
|
3
3
|
// MCP package's local mode (mcp/src/local mirrors it byte for byte), so the two cannot disagree.
|
|
4
|
+
import { finiteNumbers } from "../moments-core.js";
|
|
4
5
|
import { pboCscv } from "../pbo-core.js";
|
|
5
6
|
import { LIMITS } from "../../api/_lib/limits.js";
|
|
6
7
|
|
|
@@ -12,7 +13,12 @@ export function compute(body) {
|
|
|
12
13
|
if (matrix.length > LIMITS.max_observations) throw new RangeError(`A matrix may hold at most ${LIMITS.max_observations} rows`);
|
|
13
14
|
if (matrix[0].length > LIMITS.max_variants) throw new RangeError(`A matrix may hold at most ${LIMITS.max_variants} variants`);
|
|
14
15
|
if (!Number.isInteger(max_combinations) || max_combinations > LIMITS.max_cscv_combinations) throw new RangeError(`max_combinations may not exceed ${LIMITS.max_cscv_combinations}`);
|
|
15
|
-
const
|
|
16
|
+
const k = matrix[0].length;
|
|
17
|
+
const rows = matrix.map((row, t) => {
|
|
18
|
+
if (!Array.isArray(row) || row.length !== k) throw new RangeError(`matrix row ${t} has ${Array.isArray(row) ? row.length : "no"} values; every row needs ${k}, one per variant`);
|
|
19
|
+
return finiteNumbers(row, `matrix[${t}]`);
|
|
20
|
+
});
|
|
21
|
+
const out = pboCscv(rows, { nSplits: Number(n_splits), maxCombinations: max_combinations, seed: Number(seed) });
|
|
16
22
|
const sorted = [...out.lambdas].sort((a, b) => a - b);
|
|
17
23
|
const q = (p) => sorted[Math.min(sorted.length - 1, Math.floor(p * sorted.length))];
|
|
18
24
|
const degraded = out.is_oos_pairs.filter(([i, o]) => o < i).length / out.is_oos_pairs.length;
|
|
@@ -0,0 +1,89 @@
|
|
|
1
|
+
// js/validate/reality-check.js
|
|
2
|
+
// The pure computation behind POST /api/v1/validate/reality-check, shared by the API route and the
|
|
3
|
+
// MCP package's local mode (mcp/src/local mirrors it byte for byte), so the two cannot disagree.
|
|
4
|
+
//
|
|
5
|
+
// Data-snooping tests on every variant a search tried: Hansen's SPA, White's Reality Check and
|
|
6
|
+
// Romano and Wolf's StepM, all on one seeded stationary bootstrap (js/snooping-core.js).
|
|
7
|
+
import { LIMITS } from "../../api/_lib/limits.js";
|
|
8
|
+
import { defaultBlock, snoopingTests } from "../snooping-core.js";
|
|
9
|
+
|
|
10
|
+
// Bootstrap work (periods x variants x draws) that fits the service's wall-time limit with room to
|
|
11
|
+
// spare: about a second on one core.
|
|
12
|
+
export const MAX_BOOTSTRAP_CELLS = 200_000_000;
|
|
13
|
+
const DEFAULT_REPS = 2000;
|
|
14
|
+
const MIN_REPS = 500;
|
|
15
|
+
const MIN_OBSERVATIONS = 30;
|
|
16
|
+
|
|
17
|
+
const fmt = (x, digits = 3) => (Number.isFinite(x) ? x.toFixed(digits) : String(x));
|
|
18
|
+
|
|
19
|
+
export function compute(body) {
|
|
20
|
+
const { matrix, benchmark } = body;
|
|
21
|
+
if (!Array.isArray(matrix) || !matrix.length || !Array.isArray(matrix[0])) throw new RangeError("matrix must be an array of rows, each a row of variant returns for one period");
|
|
22
|
+
const n = matrix.length;
|
|
23
|
+
const k = matrix[0].length;
|
|
24
|
+
if (n > LIMITS.max_observations) throw new RangeError(`A matrix may hold at most ${LIMITS.max_observations} rows`);
|
|
25
|
+
if (k > LIMITS.max_variants) throw new RangeError(`A matrix may hold at most ${LIMITS.max_variants} variants`);
|
|
26
|
+
if (n < MIN_OBSERVATIONS) throw new RangeError(`A matrix needs at least ${MIN_OBSERVATIONS} rows (periods) for a bootstrap`);
|
|
27
|
+
if (k < 1) throw new RangeError("A matrix needs at least one variant");
|
|
28
|
+
let bench = null;
|
|
29
|
+
if (benchmark !== undefined && benchmark !== null) {
|
|
30
|
+
if (!Array.isArray(benchmark) || benchmark.length !== n) throw new RangeError(`benchmark must be one return per row of matrix (${n})`);
|
|
31
|
+
// JSON null or a string is refused, never read as a number: Number(null) is 0, a silent return.
|
|
32
|
+
const bad = benchmark.findIndex((x) => typeof x !== "number" || !Number.isFinite(x));
|
|
33
|
+
if (bad >= 0) throw new RangeError(`benchmark[${bad}] is not a finite number`);
|
|
34
|
+
bench = benchmark;
|
|
35
|
+
}
|
|
36
|
+
const d = new Float64Array(n * k);
|
|
37
|
+
for (let t = 0; t < n; t += 1) {
|
|
38
|
+
const row = matrix[t];
|
|
39
|
+
if (!Array.isArray(row) || row.length !== k) throw new RangeError(`matrix row ${t} has ${Array.isArray(row) ? row.length : "no"} values; every row needs ${k}, one per variant`);
|
|
40
|
+
for (let j = 0; j < k; j += 1) {
|
|
41
|
+
const x = row[j];
|
|
42
|
+
if (typeof x !== "number" || !Number.isFinite(x)) throw new RangeError(`matrix[${t}][${j}] is not a finite number; a variant with a missing period cannot be compared period by period`);
|
|
43
|
+
d[t * k + j] = x - (bench ? bench[t] : 0);
|
|
44
|
+
}
|
|
45
|
+
}
|
|
46
|
+
const block = body.block_length === undefined ? defaultBlock(n) : Number(body.block_length);
|
|
47
|
+
if (!Number.isInteger(block) || block < 1 || block > Math.floor(n / 2)) throw new RangeError(`block_length must be an integer from 1 to half the number of periods (${Math.floor(n / 2)})`);
|
|
48
|
+
const budget = Math.floor(MAX_BOOTSTRAP_CELLS / (n * k));
|
|
49
|
+
if (budget < MIN_REPS) throw new RangeError(`${n} periods x ${k} variants is too large for ${MIN_REPS} bootstrap draws within the service's limit (${MAX_BOOTSTRAP_CELLS} period-variant draws); send fewer periods (weekly instead of daily) or fewer variants`);
|
|
50
|
+
const reps = body.reps === undefined ? Math.min(DEFAULT_REPS, budget) : Number(body.reps);
|
|
51
|
+
if (!Number.isInteger(reps) || reps < MIN_REPS || reps > budget) throw new RangeError(`reps must be an integer from ${MIN_REPS} to ${budget} for a ${n} x ${k} matrix`);
|
|
52
|
+
const seed = body.seed === undefined ? 42 : Number(body.seed);
|
|
53
|
+
if (!Number.isInteger(seed)) throw new RangeError("seed must be an integer");
|
|
54
|
+
const alpha = body.alpha === undefined ? 0.05 : Number(body.alpha);
|
|
55
|
+
if (!(alpha > 0 && alpha <= 0.5)) throw new RangeError("alpha must be above 0 and at most 0.5");
|
|
56
|
+
|
|
57
|
+
const r = snoopingTests(d, n, k, { block, reps, seed, alpha });
|
|
58
|
+
const best = r.best;
|
|
59
|
+
const tStat = r.observed;
|
|
60
|
+
const p = r.spa.p_consistent;
|
|
61
|
+
const excluded = r.usable.map((ok, j) => (ok ? -1 : j)).filter((j) => j >= 0);
|
|
62
|
+
const superior = r.stepm.superior;
|
|
63
|
+
const against = bench ? "the benchmark" : "zero";
|
|
64
|
+
const verdict = p <= alpha
|
|
65
|
+
? `The best variant beats ${against} by more than the best of ${k} would by luck (p ${fmt(p)} at alpha ${alpha}).`
|
|
66
|
+
: `The best variant does not beat ${against} by more than the best of ${k} could by luck (p ${fmt(p)} at alpha ${alpha}).`;
|
|
67
|
+
return {
|
|
68
|
+
observations: n,
|
|
69
|
+
variants: k,
|
|
70
|
+
benchmark: bench ? "series" : "zero",
|
|
71
|
+
block_length: block,
|
|
72
|
+
block_length_rule: body.block_length === undefined ? "round(n^(1/3)), the default" : "as given",
|
|
73
|
+
reps,
|
|
74
|
+
seed,
|
|
75
|
+
best_variant: { index: best, mean_excess_return: r.means[best], t_stat: tStat },
|
|
76
|
+
spa: {
|
|
77
|
+
statistic: r.spa.statistic,
|
|
78
|
+
p_value: p,
|
|
79
|
+
p_value_lower: r.spa.p_lower,
|
|
80
|
+
p_value_upper: r.spa.p_upper,
|
|
81
|
+
monte_carlo_se: Math.sqrt((p * (1 - p)) / reps),
|
|
82
|
+
},
|
|
83
|
+
reality_check: { statistic: r.reality_check.statistic, p_value: r.reality_check.p_value },
|
|
84
|
+
stepm: { alpha, superior_variants: superior, rounds: r.stepm.rounds.length },
|
|
85
|
+
...(excluded.length ? { constant_variants_excluded: excluded } : {}),
|
|
86
|
+
verdict,
|
|
87
|
+
plain_reading: `Variant ${best} had the best studentized excess return over ${against} (t ${fmt(tStat, 2)}) of the ${k} variants passed. The chance that the best of ${k} variants with no edge looks at least this good is ${fmt(p)} by Hansen's SPA (consistent p-value; bounds ${fmt(r.spa.p_lower)} to ${fmt(r.spa.p_upper)}), and ${fmt(r.reality_check.p_value)} by White's Reality Check, which does not studentize. StepM at ${alpha} finds ${superior.length ? `${superior.length} variant${superior.length === 1 ? "" : "s"} that beat${superior.length === 1 ? "s" : ""} ${against}: ${superior.join(", ")}` : `no variant that beats ${against}`}. Only the variants passed are counted: a search that tried more than it passed makes luck look smaller than it was.`,
|
|
88
|
+
};
|
|
89
|
+
}
|
package/src/local.mjs
CHANGED
|
@@ -10,6 +10,7 @@ import { compute as trackRecord } from "./local/js/validate/track-record.js";
|
|
|
10
10
|
import { compute as backtestLength } from "./local/js/validate/backtest-length.js";
|
|
11
11
|
import { compute as haircutSharpe } from "./local/js/validate/haircut-sharpe.js";
|
|
12
12
|
import { compute as luckTrials } from "./local/js/validate/luck-trials.js";
|
|
13
|
+
import { compute as realityCheck } from "./local/js/validate/reality-check.js";
|
|
13
14
|
|
|
14
15
|
export const LOCAL_VALIDATORS = Object.freeze({
|
|
15
16
|
validate_deflated_sharpe: { endpoint: "validate/deflated-sharpe", compute: deflatedSharpe },
|
|
@@ -20,6 +21,7 @@ export const LOCAL_VALIDATORS = Object.freeze({
|
|
|
20
21
|
validate_backtest_length: { endpoint: "validate/backtest-length", compute: backtestLength },
|
|
21
22
|
validate_haircut_sharpe: { endpoint: "validate/haircut-sharpe", compute: haircutSharpe },
|
|
22
23
|
validate_luck_trials: { endpoint: "validate/luck-trials", compute: luckTrials },
|
|
24
|
+
validate_reality_check: { endpoint: "validate/reality-check", compute: realityCheck },
|
|
23
25
|
});
|
|
24
26
|
|
|
25
27
|
const NOTE = "Computed on this machine in local mode. Nothing was sent to canlicapital.com and no receipt was stored.";
|
package/src/schemas.mjs
CHANGED
|
@@ -24,6 +24,12 @@ export const FIELD_DESCRIPTIONS = Object.freeze({
|
|
|
24
24
|
n_splits: "Even number of blocks, at least 2; default 16.",
|
|
25
25
|
max_combinations: "Most splits evaluated, up to 2000 (default).",
|
|
26
26
|
seed: "Sampling seed; default 42.",
|
|
27
|
+
rc_matrix: "Returns of every variant the search tried, one row per period, one column per variant.",
|
|
28
|
+
matrix_file: "Path to a CSV or JSON with one numeric column per variant, instead of matrix.",
|
|
29
|
+
rc_benchmark: "Benchmark return per period; default zero.",
|
|
30
|
+
block_length: "Mean bootstrap block in periods; default round(n^(1/3)).",
|
|
31
|
+
reps: "Bootstrap draws; default 2000.",
|
|
32
|
+
alpha: "Familywise error for StepM; default 0.05.",
|
|
27
33
|
record: "A canli.paper-evidence.v0 record.",
|
|
28
34
|
sleeve_sharpe: "Annualized Sharpe of one sleeve.",
|
|
29
35
|
average_pairwise_correlation: "Average correlation between sleeves, -1 to 1.",
|
|
@@ -115,6 +121,24 @@ export const overfittingInput = z
|
|
|
115
121
|
})
|
|
116
122
|
.strict();
|
|
117
123
|
|
|
124
|
+
// ---------------------------------------------------------------------------------------------
|
|
125
|
+
// validate_reality_check: every variant's returns, inline or from a file on this machine.
|
|
126
|
+
// ---------------------------------------------------------------------------------------------
|
|
127
|
+
|
|
128
|
+
export const realityCheckToolShape = z
|
|
129
|
+
.object({
|
|
130
|
+
matrix: z.array(z.array(z.number()).max(200)).min(30).max(20000).optional().describe(d.rc_matrix),
|
|
131
|
+
matrix_file: z.string().min(1).max(4096).optional().describe(d.matrix_file),
|
|
132
|
+
benchmark: z.array(z.number()).min(30).max(20000).optional().describe(d.rc_benchmark),
|
|
133
|
+
block_length: z.number().int().min(1).max(10000).optional().describe(d.block_length),
|
|
134
|
+
reps: z.number().int().min(500).max(400000).optional().describe(d.reps),
|
|
135
|
+
seed: z.number().int().optional().describe(d.seed),
|
|
136
|
+
alpha: z.number().gt(0).max(0.5).optional().describe(d.alpha),
|
|
137
|
+
})
|
|
138
|
+
.strict();
|
|
139
|
+
|
|
140
|
+
export const realityCheckInput = realityCheckToolShape.refine((v) => (v.matrix === undefined) !== (v.matrix_file === undefined), "Send exactly one of matrix or matrix_file");
|
|
141
|
+
|
|
118
142
|
// ---------------------------------------------------------------------------------------------
|
|
119
143
|
// validate_paper_evidence: the record shape is the open standard itself, so this schema only
|
|
120
144
|
// checks that a record object was sent; canlicapital.com runs the real conformance check.
|
|
@@ -306,6 +330,7 @@ export const TOOL_DESCRIPTIONS = Object.freeze({
|
|
|
306
330
|
validate_deflated_sharpe: `Deflated Sharpe ratio: the probability (0 to 1) that the selected strategy's Sharpe beats the best that luck gives across the variants tried, with the probabilistic Sharpe and that luck benchmark. Send the seven statistics or a return series. With every variant's returns use validate_overfitting; luck as a trial count, validate_luck_trials; a multiple-testing haircut, validate_haircut_sharpe. ${LIMITS_SENTENCES.notAdmission}`,
|
|
307
331
|
audit_backtest: `One-call audit of a strategy's returns: deflated Sharpe, minimum track record and, with every variant's returns, the probability of backtest overfitting, each the matching validator's result with its own receipt. Point returns_file at the backtest's CSV or JSON instead of pasting long series. Prefer it to calling the validators one by one; one validation per check. ${LIMITS_SENTENCES.notAdmission}`,
|
|
308
332
|
validate_overfitting: `Probability of backtest overfitting (0 to 1) by CSCV: how often the in-sample best variant falls below the out-of-sample median. Needs every variant's returns (periods by variants); with summary statistics only, use validate_deflated_sharpe. ${LIMITS_SENTENCES.notAdmission}`,
|
|
333
|
+
validate_reality_check: `Data-snooping tests on every variant a search tried: Hansen's SPA p-value that the best beat the benchmark only by luck, White's Reality Check, and the variants Romano-Wolf StepM finds better. Send all variants tried, not only the winners. ${LIMITS_SENTENCES.notAdmission}`,
|
|
309
334
|
validate_paper_evidence: `Whether a paper or simulated performance record meets canli.paper-evidence.v0, with a JSON pointer per failure. Checks structure and required disclosures, not whether the returns are good. ${LIMITS_SENTENCES.scope}`,
|
|
310
335
|
validate_backtest_length: `Minimum backtest length (years) before the best of N independent trials is not expected to reach a target Sharpe by luck; with backtest_years, the most trials those years allow. For planning a search; once it has a result, use validate_deflated_sharpe. ${LIMITS_SENTENCES.notAdmission}`,
|
|
311
336
|
validate_luck_trials: `How many skill-less strategies a search would need for its best to reach this Sharpe by luck (Monte Carlo), and with a trial count, the chance it did. States luck as the best of N random tries; for the probability the Sharpe is real, use validate_deflated_sharpe. ${LIMITS_SENTENCES.notAdmission}`,
|
package/src/server.mjs
CHANGED
|
@@ -14,7 +14,7 @@ import { computeLocally } from "./local.mjs";
|
|
|
14
14
|
import { readMatrixFile, readSeriesFile } from "./series-file.mjs";
|
|
15
15
|
import { verifyReceipt } from "./local/js/receipt-statement.js";
|
|
16
16
|
import { StdioServerTransport } from "@modelcontextprotocol/server/stdio";
|
|
17
|
-
import { breadthInput, trackRecordInput, auditBacktestInput, verifyReceiptToolShape, backtestLengthInput, haircutSharpeInput, luckTrialsInput, auditBacktestToolShape, companyHistoryInput, companyHistoryToolShape, deflatedSharpeInput, deflatedSharpeToolShape, emptyInput, getKeyInput, getReceiptInput, overfittingInput, LIMITS_SENTENCES, paperEvidenceInput, TOOL_DESCRIPTIONS, validationOutput, auditOutput, keyOutput, receiptOutput, verifyReceiptOutput, statusOutput, companyHistoryOutput } from "./schemas.mjs";
|
|
17
|
+
import { breadthInput, trackRecordInput, auditBacktestInput, verifyReceiptToolShape, backtestLengthInput, haircutSharpeInput, luckTrialsInput, auditBacktestToolShape, companyHistoryInput, companyHistoryToolShape, deflatedSharpeInput, deflatedSharpeToolShape, emptyInput, getKeyInput, getReceiptInput, overfittingInput, realityCheckInput, realityCheckToolShape, LIMITS_SENTENCES, paperEvidenceInput, TOOL_DESCRIPTIONS, validationOutput, auditOutput, keyOutput, receiptOutput, verifyReceiptOutput, statusOutput, companyHistoryOutput } from "./schemas.mjs";
|
|
18
18
|
|
|
19
19
|
export const DEFAULT_BASE = "https://canlicapital.com";
|
|
20
20
|
export const SERVER_NAME = "canlicapital-validation-mcp";
|
|
@@ -66,7 +66,7 @@ export function configuredLocal(value) {
|
|
|
66
66
|
// is most of each turn's prompt (README, "Toolsets"; bench/tool_list_tokens.py measures it), so a
|
|
67
67
|
// client that needs one kind of tool can load only that kind. Default: all.
|
|
68
68
|
export const TOOLSETS = Object.freeze({
|
|
69
|
-
validate: Object.freeze(["get_key", "validate_deflated_sharpe", "validate_overfitting", "validate_paper_evidence", "validate_breadth", "validate_track_record", "validate_backtest_length", "validate_haircut_sharpe", "validate_luck_trials", "audit_backtest"]),
|
|
69
|
+
validate: Object.freeze(["get_key", "validate_deflated_sharpe", "validate_overfitting", "validate_paper_evidence", "validate_breadth", "validate_track_record", "validate_backtest_length", "validate_haircut_sharpe", "validate_luck_trials", "validate_reality_check", "audit_backtest"]),
|
|
70
70
|
receipts: Object.freeze(["get_receipt", "verify_receipt"]),
|
|
71
71
|
company: Object.freeze(["company_financial_history"]),
|
|
72
72
|
status: Object.freeze(["service_status"]),
|
|
@@ -260,6 +260,20 @@ export async function toolValidateOverfitting(session, args) {
|
|
|
260
260
|
return validationText(session, response);
|
|
261
261
|
}
|
|
262
262
|
|
|
263
|
+
// Every variant's returns, from the call or from a file on this machine (stdio only), tested for
|
|
264
|
+
// data snooping in one call; the file is read here and only its numbers are sent.
|
|
265
|
+
export async function toolValidateRealityCheck(session, args) {
|
|
266
|
+
const input = parseOrThrow(realityCheckInput, args, "validate_reality_check");
|
|
267
|
+
if (input.matrix_file && session.hosted) {
|
|
268
|
+
throw new Error("validate_reality_check: the hosted endpoint cannot read files on your machine; send matrix as numbers, or run the server locally with npx -y canli-validation-mcp.");
|
|
269
|
+
}
|
|
270
|
+
const { matrix_file: file, ...rest } = input;
|
|
271
|
+
const body = file ? { ...rest, matrix: readMatrixFile(file).matrix } : rest;
|
|
272
|
+
if (session.local) { const local = computeLocally("validate_reality_check", body); return validationText(session, local); }
|
|
273
|
+
const response = await validateRemote(session, "validate_reality_check", "/api/v1/validate/reality-check", body);
|
|
274
|
+
return validationText(session, response);
|
|
275
|
+
}
|
|
276
|
+
|
|
263
277
|
export async function toolValidatePaperEvidence(session, args) {
|
|
264
278
|
const body = parseOrThrow(paperEvidenceInput, args, "validate_paper_evidence");
|
|
265
279
|
if (session.local) { const local = computeLocally("validate_paper_evidence", body); return validationText(session, local); }
|
|
@@ -549,6 +563,11 @@ export function registerTools(server, session) {
|
|
|
549
563
|
{ title: "Validate overfitting (CSCV)", annotations: { title: "Validate overfitting (CSCV)", ...WRITES_RECEIPT }, description: TOOL_DESCRIPTIONS.validate_overfitting, inputSchema: overfittingInput, outputSchema: validationOutput },
|
|
550
564
|
(args) => toolValidateOverfitting(session, args),
|
|
551
565
|
);
|
|
566
|
+
register(
|
|
567
|
+
"validate_reality_check",
|
|
568
|
+
{ title: "Data-snooping tests (SPA, Reality Check, StepM)", annotations: { title: "Data-snooping tests (SPA, Reality Check, StepM)", ...WRITES_RECEIPT }, description: TOOL_DESCRIPTIONS.validate_reality_check, inputSchema: realityCheckToolShape, outputSchema: validationOutput },
|
|
569
|
+
(args) => toolValidateRealityCheck(session, args),
|
|
570
|
+
);
|
|
552
571
|
register(
|
|
553
572
|
"validate_paper_evidence",
|
|
554
573
|
{ title: "Validate paper evidence", annotations: { title: "Validate paper evidence", ...WRITES_RECEIPT }, description: TOOL_DESCRIPTIONS.validate_paper_evidence, inputSchema: paperEvidenceInput, outputSchema: validationOutput },
|