canli-validation-mcp 0.9.1 → 0.10.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -7,8 +7,9 @@
7
7
 
8
8
  **Is your best backtest real, or just the luckiest of the variants you tried?** This MCP server
9
9
  lets Claude, Cursor or any MCP client answer that with the standard corrections: the deflated
10
- Sharpe ratio, the CSCV probability of backtest overfitting, the minimum track record length, the
11
- haircut Sharpe ratio, and luck-equivalent trials. Free and MIT-licensed.
10
+ Sharpe ratio, the CSCV probability of backtest overfitting, data-snooping tests on every variant
11
+ tried (Hansen's SPA, White's Reality Check, Romano-Wolf StepM), the minimum track record length,
12
+ the haircut Sharpe ratio, and luck-equivalent trials. Free and MIT-licensed.
12
13
 
13
14
  ## Quick start
14
15
 
@@ -34,9 +35,10 @@ other stdio client (Cursor, VS Code) run the same `npx` command; see "Claude Des
34
35
  ## What it does
35
36
 
36
37
  An MCP (Model Context Protocol) server over canlicapital.com's free, keyed validation API. It
37
- gives a coding agent fourteen tools: issue a free key, run the eight validators (deflated Sharpe,
38
- CSCV overfitting, paper-evidence conformance, breadth ceiling, minimum track record length,
39
- minimum backtest length, haircut Sharpe ratio, luck-equivalent trials),
38
+ gives a coding agent fifteen tools: issue a free key, run the nine validators (deflated Sharpe,
39
+ CSCV overfitting, data-snooping tests (Hansen's SPA, White's Reality Check and Romano-Wolf StepM),
40
+ paper-evidence conformance, breadth ceiling, minimum track record length, minimum backtest length,
41
+ haircut Sharpe ratio, luck-equivalent trials),
40
42
  audit one backtest with three of them in a single call, fetch a stored receipt, verify a
41
43
  receipt's signature offline, read service status, and read a company's reported financial history
42
44
  from SEC filings. Every validation result carries, beside the number, the sentences that say what
@@ -80,6 +82,7 @@ repository for the full design.
80
82
  | `validate_backtest_length` | `POST /api/v1/validate/backtest-length` | yes |
81
83
  | `validate_haircut_sharpe` | `POST /api/v1/validate/haircut-sharpe` | yes |
82
84
  | `validate_luck_trials` | `POST /api/v1/validate/luck-trials` | yes |
85
+ | `validate_reality_check` | `POST /api/v1/validate/reality-check` (a `matrix_file` is read on your machine and sent as numbers) | yes |
83
86
  | `audit_backtest` | the deflated Sharpe, track record and, with `variants`, overfitting routes, one validation each | yes |
84
87
  | `verify_receipt` | `GET /api/v1/receipts/{id}` when given an id; the checks run locally | no |
85
88
  | `get_receipt` | `GET /api/v1/receipts/{id}` | no |
@@ -97,6 +100,24 @@ repository for the full design.
97
100
  Sending fields from both shapes, or from neither, is rejected before any request leaves the
98
101
  process; see `src/schemas.mjs`.
99
102
 
103
+ `validate_reality_check` takes the returns of every variant the search tried, one row per period
104
+ and one column per variant, as `matrix` or as `matrix_file` (a CSV or JSON on your machine), and an
105
+ optional `benchmark` series; with none, variants are tested against zero. It runs three tests on one
106
+ seeded stationary bootstrap:
107
+
108
+ - Hansen's SPA: the chance that the best variant's studentized excess return is this good if no
109
+ variant has an edge, with lower and upper bounds (`spa.p_value`, the consistent p-value, is the
110
+ headline, and `monte_carlo_se` its sampling error);
111
+ - White's Reality Check: the same question without studentizing, so one volatile variant can
112
+ dominate it;
113
+ - Romano and Wolf's StepM: which variants beat the benchmark, with the familywise error held at
114
+ `alpha`.
115
+
116
+ A result reproduces exactly from its `seed`. On three fixed-seed cases the p-values agree with
117
+ Python's `arch` 8.0 and with a numpy transcription of Hansen's formulas within Monte Carlo error, and
118
+ the StepM sets match `arch`'s (`js/snooping-core.test.js`). Only the variants sent are counted: a
119
+ search that tried more than it sends makes luck look smaller than it was.
120
+
100
121
  `company_financial_history` is different from the other tools: it reads the public company
101
122
  reference at canlicapital.com, not the validation API. Give it a CIK (1 to 10 digits) to list a
102
123
  company's available financial histories, or a CIK and a us-gaap concept such as `Revenues` to
@@ -317,7 +338,7 @@ On a breadth result this is about half the text. Set `CANLI_FULL_ENVELOPE=1` to
317
338
  ## Toolsets (tokens)
318
339
 
319
340
  A client sends the model the whole tool list on every turn, and it is most of each turn's prompt:
320
- a validation result is a few hundred tokens, the list of all fourteen tools several thousand. A
341
+ a validation result is a few hundred tokens, the list of all fifteen tools several thousand. A
321
342
  client that needs one kind of tool can list only that kind, with `CANLI_TOOLSETS` (stdio) or
322
343
  `?toolsets=` (hosted endpoint). The default is every tool.
323
344
 
@@ -333,9 +354,9 @@ tokenizers give different absolute counts), in the shape an OpenAI-style client
333
354
 
334
355
  | CANLI_TOOLSETS | tools | tokens per turn | of all |
335
356
  |---|---|---|---|
336
- | `all` | 14 | 3,834 | 100% |
337
- | `validate` | 10 | 3,166 | 83% |
338
- | `receipts` | 2 | 312 | 8% |
357
+ | `all` | 15 | 4,194 | 100% |
358
+ | `validate` | 11 | 3,526 | 84% |
359
+ | `receipts` | 2 | 312 | 7% |
339
360
  | `company` | 1 | 273 | 7% |
340
361
  | `status` | 1 | 89 | 2% |
341
362
 
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "canli-validation-mcp",
3
- "version": "0.9.1",
3
+ "version": "0.10.0",
4
4
  "description": "Check whether a backtest is real: deflated Sharpe, CSCV probability of backtest overfitting, minimum track record and backtest length, haircut Sharpe and luck-equivalent trials, as an MCP server. Local mode, hosted endpoint, signed receipts.",
5
5
  "keywords": [
6
6
  "mcp",
@@ -0,0 +1,185 @@
1
+ // =============================================================================
2
+ // snooping-core.js
3
+ // -----------------------------------------------------------------------------
4
+ // Data-snooping tests for a search: did the best of K variants beat the benchmark by more than the
5
+ // best of K would by luck? One stationary bootstrap (Politis and Romano 1994) feeds three tests:
6
+ //
7
+ // White's Reality Check (2000) the best mean excess return against the bootstrap distribution
8
+ // of the best recentred mean. Not studentized, so one volatile
9
+ // variant can dominate it.
10
+ // Hansen's SPA (2005) the best t-statistic, each variant studentized by the
11
+ // Politis-Romano variance of its mean. Three recentrings give a
12
+ // lower bound, a consistent p-value and an upper bound: poor
13
+ // variants are not allowed to inflate the p-value.
14
+ // Romano and Wolf's StepM (2005) step-down: which variants beat the benchmark, controlling the
15
+ // familywise error at alpha.
16
+ //
17
+ // d[t][k] is variant k's return minus the benchmark's in period t; higher is better. Every draw
18
+ // comes from makeRandom(seed), so a result reproduces exactly from its inputs. The reference
19
+ // implementation it is checked against (Python's arch 8.0, and a separate numpy transcription of
20
+ // Hansen's formulas for the studentized SPA) draws with numpy, so agreement is within Monte Carlo
21
+ // error, not to the digit; scripts/research/reality-check/ holds the comparison.
22
+ // =============================================================================
23
+
24
+ import { makeRandom } from "./selection-risk-core.js";
25
+
26
+ /** Indices of one stationary-bootstrap resample of n periods with mean block length `block`. */
27
+ export function stationaryIndices(n, block, next, out = new Int32Array(n)) {
28
+ const p = 1 / block;
29
+ let at = Math.floor(next() * n);
30
+ out[0] = at;
31
+ for (let t = 1; t < n; t += 1) {
32
+ at = next() < p ? Math.floor(next() * n) : at + 1 === n ? 0 : at + 1;
33
+ out[t] = at;
34
+ }
35
+ return out;
36
+ }
37
+
38
+ /** round(n^(1/3)), at least 1: the default mean block length. */
39
+ export const defaultBlock = (n) => Math.max(1, Math.round(Math.cbrt(n)));
40
+
41
+ // Weights below this are dropped from the variance sum. Each term they multiply is bounded by the
42
+ // sample variance, so what is dropped is under 2e-14 * n of it (4e-10 at 20,000 periods).
43
+ const KAPPA_FLOOR = 1e-14;
44
+
45
+ /**
46
+ * n times the variance of each column's mean under the stationary bootstrap, in closed form
47
+ * (Politis and Romano 1994, as Hansen 2005 uses it): gamma_0 + 2 * sum_i kappa_i * gamma_i, with
48
+ * kappa_i = (1 - i/n)(1 - p)^i + (i/n)(1 - p)^(n - i). Lags whose weight is below KAPPA_FLOOR are
49
+ * skipped, which keeps the cost near n * k * 30 * block rather than n^2 * k.
50
+ */
51
+ export function bootstrapVariance(d, n, k, block, means) {
52
+ const p = 1 / block;
53
+ const out = new Float64Array(k);
54
+ const cross = (lag, weight) => {
55
+ for (let t = 0; t + lag < n; t += 1) {
56
+ const a = t * k;
57
+ const b = (t + lag) * k;
58
+ for (let j = 0; j < k; j += 1) out[j] += weight * (d[a + j] - means[j]) * (d[b + j] - means[j]);
59
+ }
60
+ };
61
+ cross(0, 1 / n);
62
+ const q = 1 - p;
63
+ for (let i = 1; i < n; i += 1) {
64
+ const kappa = (1 - i / n) * q ** i + (i / n) * q ** (n - i);
65
+ if (kappa < KAPPA_FLOOR) {
66
+ // The first term has decayed; the wrap-around term (i/n)(1 - p)^(n - i) only reaches the
67
+ // floor again once n - i <= log(floor) / log(1 - p). Skip to there.
68
+ const tail = q > 0 ? Math.log(KAPPA_FLOOR) / Math.log(q) : 0;
69
+ const jump = Math.floor(n - tail);
70
+ if (jump > i + 1) i = Math.min(jump, n) - 1;
71
+ continue;
72
+ }
73
+ cross(i, (2 * kappa) / n);
74
+ }
75
+ return out;
76
+ }
77
+
78
+ /**
79
+ * All three tests on one bootstrap. `d` is a row-major Float64Array of n periods by k variants.
80
+ * Returns means, the variance estimates, the bootstrap maxima under each recentring and the
81
+ * StepM rejections; p-values are shares of `reps` draws strictly above the observed statistic.
82
+ */
83
+ export function snoopingTests(d, n, k, { block = defaultBlock(n), reps = 1000, seed = 42, alpha = 0.05, studentize = true } = {}) {
84
+ if (!(n >= 2 && k >= 1)) throw new RangeError("need at least 2 periods and 1 variant");
85
+ if (!(block >= 1 && block <= n)) throw new RangeError(`block_length must be from 1 to the number of periods (${n})`);
86
+ if (!Number.isInteger(reps) || reps < 100) throw new RangeError("reps must be an integer of at least 100");
87
+ if (!(alpha > 0 && alpha < 1)) throw new RangeError("alpha must be between 0 and 1");
88
+ const means = new Float64Array(k);
89
+ for (let t = 0; t < n; t += 1) for (let j = 0; j < k; j += 1) means[j] += d[t * k + j];
90
+ for (let j = 0; j < k; j += 1) means[j] /= n;
91
+ const variance = bootstrapVariance(d, n, k, block, means);
92
+ const sd = new Float64Array(k);
93
+ for (let j = 0; j < k; j += 1) sd[j] = Math.sqrt(Math.max(variance[j], 0));
94
+ // A variant whose excess return never varies has no sampling distribution to studentize by.
95
+ const usable = Array.from(sd, (s) => s > 0);
96
+ if (!usable.some(Boolean)) throw new RangeError("every variant's excess return is constant; there is nothing to test");
97
+ const root = Math.sqrt(n);
98
+ const scale = (j) => (studentize ? root / sd[j] : 1);
99
+ // Hansen's recentrings: lower keeps poor variants at zero, consistent drops only those far below
100
+ // zero, upper recentres everything (White's null, the least favourable configuration).
101
+ const threshold = (j) => -Math.sqrt((variance[j] / n) * 2 * Math.log(Math.log(n)));
102
+ const centre = [
103
+ Float64Array.from(means, (m) => Math.max(m, 0)),
104
+ Float64Array.from(means, (m, j) => (m >= threshold(j) ? m : 0)),
105
+ Float64Array.from(means),
106
+ ];
107
+ let observed = Number.NEGATIVE_INFINITY;
108
+ let best = -1;
109
+ for (let j = 0; j < k; j += 1) {
110
+ if (!usable[j]) continue;
111
+ const s = means[j] * scale(j);
112
+ if (s > observed) { observed = s; best = j; }
113
+ }
114
+ let rcObserved = Number.NEGATIVE_INFINITY;
115
+ for (let j = 0; j < k; j += 1) if (usable[j]) rcObserved = Math.max(rcObserved, means[j]);
116
+
117
+ const next = makeRandom(seed);
118
+ const idx = new Int32Array(n);
119
+ const star = new Float64Array(k);
120
+ const boot = new Float64Array(reps * k); // bootstrap mean of each variant, per draw
121
+ for (let r = 0; r < reps; r += 1) {
122
+ stationaryIndices(n, block, next, idx);
123
+ star.fill(0);
124
+ for (let t = 0; t < n; t += 1) {
125
+ const row = idx[t] * k;
126
+ for (let j = 0; j < k; j += 1) star[j] += d[row + j];
127
+ }
128
+ for (let j = 0; j < k; j += 1) boot[r * k + j] = star[j] / n;
129
+ }
130
+
131
+ const spaStat = Math.max(observed, 0);
132
+ const exceed = [0, 0, 0];
133
+ let rcExceed = 0;
134
+ for (let r = 0; r < reps; r += 1) {
135
+ const maxima = [Number.NEGATIVE_INFINITY, Number.NEGATIVE_INFINITY, Number.NEGATIVE_INFINITY];
136
+ let rcMax = Number.NEGATIVE_INFINITY;
137
+ for (let j = 0; j < k; j += 1) {
138
+ if (!usable[j]) continue;
139
+ const m = boot[r * k + j];
140
+ for (let c = 0; c < 3; c += 1) maxima[c] = Math.max(maxima[c], (m - centre[c][j]) * scale(j));
141
+ rcMax = Math.max(rcMax, m - means[j]);
142
+ }
143
+ for (let c = 0; c < 3; c += 1) if (Math.max(maxima[c], 0) > spaStat) exceed[c] += 1;
144
+ if (rcMax > rcObserved) rcExceed += 1;
145
+ }
146
+
147
+ // StepM: reject every variant above the (1 - alpha) quantile of the bootstrap maximum over those
148
+ // not yet rejected (consistent recentring), remove them, and repeat until nothing more goes.
149
+ const rejected = new Set();
150
+ const rounds = [];
151
+ for (;;) {
152
+ const live = [];
153
+ for (let j = 0; j < k; j += 1) if (usable[j] && !rejected.has(j)) live.push(j);
154
+ if (!live.length) break;
155
+ const maxima = new Float64Array(reps);
156
+ for (let r = 0; r < reps; r += 1) {
157
+ let m = Number.NEGATIVE_INFINITY;
158
+ for (const j of live) m = Math.max(m, (boot[r * k + j] - centre[1][j]) * scale(j));
159
+ maxima[r] = m;
160
+ }
161
+ maxima.sort();
162
+ const critical = quantile(maxima, 1 - alpha);
163
+ const now = live.filter((j) => means[j] * scale(j) > critical);
164
+ if (!now.length) break;
165
+ for (const j of now) rejected.add(j);
166
+ rounds.push({ critical, rejected: now });
167
+ }
168
+
169
+ return {
170
+ n, k, block, reps, seed, alpha, studentize,
171
+ means, variance, usable,
172
+ best, observed,
173
+ spa: { statistic: spaStat, p_lower: exceed[0] / reps, p_consistent: exceed[1] / reps, p_upper: exceed[2] / reps },
174
+ reality_check: { statistic: rcObserved, p_value: rcExceed / reps },
175
+ stepm: { superior: [...rejected].sort((a, b) => a - b), rounds },
176
+ };
177
+ }
178
+
179
+ // numpy's default percentile (linear interpolation between order statistics), on sorted values.
180
+ function quantile(sorted, q) {
181
+ const h = (sorted.length - 1) * q;
182
+ const lo = Math.floor(h);
183
+ const hi = Math.min(lo + 1, sorted.length - 1);
184
+ return sorted[lo] + (h - lo) * (sorted[hi] - sorted[lo]);
185
+ }
@@ -0,0 +1,89 @@
1
+ // js/validate/reality-check.js
2
+ // The pure computation behind POST /api/v1/validate/reality-check, shared by the API route and the
3
+ // MCP package's local mode (mcp/src/local mirrors it byte for byte), so the two cannot disagree.
4
+ //
5
+ // Data-snooping tests on every variant a search tried: Hansen's SPA, White's Reality Check and
6
+ // Romano and Wolf's StepM, all on one seeded stationary bootstrap (js/snooping-core.js).
7
+ import { LIMITS } from "../../api/_lib/limits.js";
8
+ import { defaultBlock, snoopingTests } from "../snooping-core.js";
9
+
10
+ // Bootstrap work (periods x variants x draws) that fits the service's wall-time limit with room to
11
+ // spare: about a second on one core.
12
+ export const MAX_BOOTSTRAP_CELLS = 200_000_000;
13
+ const DEFAULT_REPS = 2000;
14
+ const MIN_REPS = 500;
15
+ const MIN_OBSERVATIONS = 30;
16
+
17
+ const fmt = (x, digits = 3) => (Number.isFinite(x) ? x.toFixed(digits) : String(x));
18
+
19
+ export function compute(body) {
20
+ const { matrix, benchmark } = body;
21
+ if (!Array.isArray(matrix) || !matrix.length || !Array.isArray(matrix[0])) throw new RangeError("matrix must be an array of rows, each a row of variant returns for one period");
22
+ const n = matrix.length;
23
+ const k = matrix[0].length;
24
+ if (n > LIMITS.max_observations) throw new RangeError(`A matrix may hold at most ${LIMITS.max_observations} rows`);
25
+ if (k > LIMITS.max_variants) throw new RangeError(`A matrix may hold at most ${LIMITS.max_variants} variants`);
26
+ if (n < MIN_OBSERVATIONS) throw new RangeError(`A matrix needs at least ${MIN_OBSERVATIONS} rows (periods) for a bootstrap`);
27
+ if (k < 1) throw new RangeError("A matrix needs at least one variant");
28
+ let bench = null;
29
+ if (benchmark !== undefined && benchmark !== null) {
30
+ if (!Array.isArray(benchmark) || benchmark.length !== n) throw new RangeError(`benchmark must be one return per row of matrix (${n})`);
31
+ // JSON null or a string is refused, never read as a number: Number(null) is 0, a silent return.
32
+ const bad = benchmark.findIndex((x) => typeof x !== "number" || !Number.isFinite(x));
33
+ if (bad >= 0) throw new RangeError(`benchmark[${bad}] is not a finite number`);
34
+ bench = benchmark;
35
+ }
36
+ const d = new Float64Array(n * k);
37
+ for (let t = 0; t < n; t += 1) {
38
+ const row = matrix[t];
39
+ if (!Array.isArray(row) || row.length !== k) throw new RangeError(`matrix row ${t} has ${Array.isArray(row) ? row.length : "no"} values; every row needs ${k}, one per variant`);
40
+ for (let j = 0; j < k; j += 1) {
41
+ const x = row[j];
42
+ if (typeof x !== "number" || !Number.isFinite(x)) throw new RangeError(`matrix[${t}][${j}] is not a finite number; a variant with a missing period cannot be compared period by period`);
43
+ d[t * k + j] = x - (bench ? bench[t] : 0);
44
+ }
45
+ }
46
+ const block = body.block_length === undefined ? defaultBlock(n) : Number(body.block_length);
47
+ if (!Number.isInteger(block) || block < 1 || block > Math.floor(n / 2)) throw new RangeError(`block_length must be an integer from 1 to half the number of periods (${Math.floor(n / 2)})`);
48
+ const budget = Math.floor(MAX_BOOTSTRAP_CELLS / (n * k));
49
+ if (budget < MIN_REPS) throw new RangeError(`${n} periods x ${k} variants is too large for ${MIN_REPS} bootstrap draws within the service's limit (${MAX_BOOTSTRAP_CELLS} period-variant draws); send fewer periods (weekly instead of daily) or fewer variants`);
50
+ const reps = body.reps === undefined ? Math.min(DEFAULT_REPS, budget) : Number(body.reps);
51
+ if (!Number.isInteger(reps) || reps < MIN_REPS || reps > budget) throw new RangeError(`reps must be an integer from ${MIN_REPS} to ${budget} for a ${n} x ${k} matrix`);
52
+ const seed = body.seed === undefined ? 42 : Number(body.seed);
53
+ if (!Number.isInteger(seed)) throw new RangeError("seed must be an integer");
54
+ const alpha = body.alpha === undefined ? 0.05 : Number(body.alpha);
55
+ if (!(alpha > 0 && alpha <= 0.5)) throw new RangeError("alpha must be above 0 and at most 0.5");
56
+
57
+ const r = snoopingTests(d, n, k, { block, reps, seed, alpha });
58
+ const best = r.best;
59
+ const tStat = r.observed;
60
+ const p = r.spa.p_consistent;
61
+ const excluded = r.usable.map((ok, j) => (ok ? -1 : j)).filter((j) => j >= 0);
62
+ const superior = r.stepm.superior;
63
+ const against = bench ? "the benchmark" : "zero";
64
+ const verdict = p <= alpha
65
+ ? `The best variant beats ${against} by more than the best of ${k} would by luck (p ${fmt(p)} at alpha ${alpha}).`
66
+ : `The best variant does not beat ${against} by more than the best of ${k} could by luck (p ${fmt(p)} at alpha ${alpha}).`;
67
+ return {
68
+ observations: n,
69
+ variants: k,
70
+ benchmark: bench ? "series" : "zero",
71
+ block_length: block,
72
+ block_length_rule: body.block_length === undefined ? "round(n^(1/3)), the default" : "as given",
73
+ reps,
74
+ seed,
75
+ best_variant: { index: best, mean_excess_return: r.means[best], t_stat: tStat },
76
+ spa: {
77
+ statistic: r.spa.statistic,
78
+ p_value: p,
79
+ p_value_lower: r.spa.p_lower,
80
+ p_value_upper: r.spa.p_upper,
81
+ monte_carlo_se: Math.sqrt((p * (1 - p)) / reps),
82
+ },
83
+ reality_check: { statistic: r.reality_check.statistic, p_value: r.reality_check.p_value },
84
+ stepm: { alpha, superior_variants: superior, rounds: r.stepm.rounds.length },
85
+ ...(excluded.length ? { constant_variants_excluded: excluded } : {}),
86
+ verdict,
87
+ plain_reading: `Variant ${best} had the best studentized excess return over ${against} (t ${fmt(tStat, 2)}) of the ${k} variants passed. The chance that the best of ${k} variants with no edge looks at least this good is ${fmt(p)} by Hansen's SPA (consistent p-value; bounds ${fmt(r.spa.p_lower)} to ${fmt(r.spa.p_upper)}), and ${fmt(r.reality_check.p_value)} by White's Reality Check, which does not studentize. StepM at ${alpha} finds ${superior.length ? `${superior.length} variant${superior.length === 1 ? "" : "s"} that beat${superior.length === 1 ? "s" : ""} ${against}: ${superior.join(", ")}` : `no variant that beats ${against}`}. Only the variants passed are counted: a search that tried more than it passed makes luck look smaller than it was.`,
88
+ };
89
+ }
package/src/local.mjs CHANGED
@@ -10,6 +10,7 @@ import { compute as trackRecord } from "./local/js/validate/track-record.js";
10
10
  import { compute as backtestLength } from "./local/js/validate/backtest-length.js";
11
11
  import { compute as haircutSharpe } from "./local/js/validate/haircut-sharpe.js";
12
12
  import { compute as luckTrials } from "./local/js/validate/luck-trials.js";
13
+ import { compute as realityCheck } from "./local/js/validate/reality-check.js";
13
14
 
14
15
  export const LOCAL_VALIDATORS = Object.freeze({
15
16
  validate_deflated_sharpe: { endpoint: "validate/deflated-sharpe", compute: deflatedSharpe },
@@ -20,6 +21,7 @@ export const LOCAL_VALIDATORS = Object.freeze({
20
21
  validate_backtest_length: { endpoint: "validate/backtest-length", compute: backtestLength },
21
22
  validate_haircut_sharpe: { endpoint: "validate/haircut-sharpe", compute: haircutSharpe },
22
23
  validate_luck_trials: { endpoint: "validate/luck-trials", compute: luckTrials },
24
+ validate_reality_check: { endpoint: "validate/reality-check", compute: realityCheck },
23
25
  });
24
26
 
25
27
  const NOTE = "Computed on this machine in local mode. Nothing was sent to canlicapital.com and no receipt was stored.";
package/src/schemas.mjs CHANGED
@@ -24,6 +24,12 @@ export const FIELD_DESCRIPTIONS = Object.freeze({
24
24
  n_splits: "Even number of blocks, at least 2; default 16.",
25
25
  max_combinations: "Most splits evaluated, up to 2000 (default).",
26
26
  seed: "Sampling seed; default 42.",
27
+ rc_matrix: "Returns of every variant the search tried, one row per period, one column per variant.",
28
+ matrix_file: "Path to a CSV or JSON with one numeric column per variant, instead of matrix.",
29
+ rc_benchmark: "Benchmark return per period; default zero.",
30
+ block_length: "Mean bootstrap block in periods; default round(n^(1/3)).",
31
+ reps: "Bootstrap draws; default 2000.",
32
+ alpha: "Familywise error for StepM; default 0.05.",
27
33
  record: "A canli.paper-evidence.v0 record.",
28
34
  sleeve_sharpe: "Annualized Sharpe of one sleeve.",
29
35
  average_pairwise_correlation: "Average correlation between sleeves, -1 to 1.",
@@ -115,6 +121,24 @@ export const overfittingInput = z
115
121
  })
116
122
  .strict();
117
123
 
124
+ // ---------------------------------------------------------------------------------------------
125
+ // validate_reality_check: every variant's returns, inline or from a file on this machine.
126
+ // ---------------------------------------------------------------------------------------------
127
+
128
+ export const realityCheckToolShape = z
129
+ .object({
130
+ matrix: z.array(z.array(z.number()).max(200)).min(30).max(20000).optional().describe(d.rc_matrix),
131
+ matrix_file: z.string().min(1).max(4096).optional().describe(d.matrix_file),
132
+ benchmark: z.array(z.number()).min(30).max(20000).optional().describe(d.rc_benchmark),
133
+ block_length: z.number().int().min(1).max(10000).optional().describe(d.block_length),
134
+ reps: z.number().int().min(500).max(400000).optional().describe(d.reps),
135
+ seed: z.number().int().optional().describe(d.seed),
136
+ alpha: z.number().gt(0).max(0.5).optional().describe(d.alpha),
137
+ })
138
+ .strict();
139
+
140
+ export const realityCheckInput = realityCheckToolShape.refine((v) => (v.matrix === undefined) !== (v.matrix_file === undefined), "Send exactly one of matrix or matrix_file");
141
+
118
142
  // ---------------------------------------------------------------------------------------------
119
143
  // validate_paper_evidence: the record shape is the open standard itself, so this schema only
120
144
  // checks that a record object was sent; canlicapital.com runs the real conformance check.
@@ -306,6 +330,7 @@ export const TOOL_DESCRIPTIONS = Object.freeze({
306
330
  validate_deflated_sharpe: `Deflated Sharpe ratio: the probability (0 to 1) that the selected strategy's Sharpe beats the best that luck gives across the variants tried, with the probabilistic Sharpe and that luck benchmark. Send the seven statistics or a return series. With every variant's returns use validate_overfitting; luck as a trial count, validate_luck_trials; a multiple-testing haircut, validate_haircut_sharpe. ${LIMITS_SENTENCES.notAdmission}`,
307
331
  audit_backtest: `One-call audit of a strategy's returns: deflated Sharpe, minimum track record and, with every variant's returns, the probability of backtest overfitting, each the matching validator's result with its own receipt. Point returns_file at the backtest's CSV or JSON instead of pasting long series. Prefer it to calling the validators one by one; one validation per check. ${LIMITS_SENTENCES.notAdmission}`,
308
332
  validate_overfitting: `Probability of backtest overfitting (0 to 1) by CSCV: how often the in-sample best variant falls below the out-of-sample median. Needs every variant's returns (periods by variants); with summary statistics only, use validate_deflated_sharpe. ${LIMITS_SENTENCES.notAdmission}`,
333
+ validate_reality_check: `Data-snooping tests on every variant a search tried: Hansen's SPA p-value that the best beat the benchmark only by luck, White's Reality Check, and the variants Romano-Wolf StepM finds better. Send all variants tried, not only the winners. ${LIMITS_SENTENCES.notAdmission}`,
309
334
  validate_paper_evidence: `Whether a paper or simulated performance record meets canli.paper-evidence.v0, with a JSON pointer per failure. Checks structure and required disclosures, not whether the returns are good. ${LIMITS_SENTENCES.scope}`,
310
335
  validate_backtest_length: `Minimum backtest length (years) before the best of N independent trials is not expected to reach a target Sharpe by luck; with backtest_years, the most trials those years allow. For planning a search; once it has a result, use validate_deflated_sharpe. ${LIMITS_SENTENCES.notAdmission}`,
311
336
  validate_luck_trials: `How many skill-less strategies a search would need for its best to reach this Sharpe by luck (Monte Carlo), and with a trial count, the chance it did. States luck as the best of N random tries; for the probability the Sharpe is real, use validate_deflated_sharpe. ${LIMITS_SENTENCES.notAdmission}`,
package/src/server.mjs CHANGED
@@ -14,7 +14,7 @@ import { computeLocally } from "./local.mjs";
14
14
  import { readMatrixFile, readSeriesFile } from "./series-file.mjs";
15
15
  import { verifyReceipt } from "./local/js/receipt-statement.js";
16
16
  import { StdioServerTransport } from "@modelcontextprotocol/server/stdio";
17
- import { breadthInput, trackRecordInput, auditBacktestInput, verifyReceiptToolShape, backtestLengthInput, haircutSharpeInput, luckTrialsInput, auditBacktestToolShape, companyHistoryInput, companyHistoryToolShape, deflatedSharpeInput, deflatedSharpeToolShape, emptyInput, getKeyInput, getReceiptInput, overfittingInput, LIMITS_SENTENCES, paperEvidenceInput, TOOL_DESCRIPTIONS, validationOutput, auditOutput, keyOutput, receiptOutput, verifyReceiptOutput, statusOutput, companyHistoryOutput } from "./schemas.mjs";
17
+ import { breadthInput, trackRecordInput, auditBacktestInput, verifyReceiptToolShape, backtestLengthInput, haircutSharpeInput, luckTrialsInput, auditBacktestToolShape, companyHistoryInput, companyHistoryToolShape, deflatedSharpeInput, deflatedSharpeToolShape, emptyInput, getKeyInput, getReceiptInput, overfittingInput, realityCheckInput, realityCheckToolShape, LIMITS_SENTENCES, paperEvidenceInput, TOOL_DESCRIPTIONS, validationOutput, auditOutput, keyOutput, receiptOutput, verifyReceiptOutput, statusOutput, companyHistoryOutput } from "./schemas.mjs";
18
18
 
19
19
  export const DEFAULT_BASE = "https://canlicapital.com";
20
20
  export const SERVER_NAME = "canlicapital-validation-mcp";
@@ -66,7 +66,7 @@ export function configuredLocal(value) {
66
66
  // is most of each turn's prompt (README, "Toolsets"; bench/tool_list_tokens.py measures it), so a
67
67
  // client that needs one kind of tool can load only that kind. Default: all.
68
68
  export const TOOLSETS = Object.freeze({
69
- validate: Object.freeze(["get_key", "validate_deflated_sharpe", "validate_overfitting", "validate_paper_evidence", "validate_breadth", "validate_track_record", "validate_backtest_length", "validate_haircut_sharpe", "validate_luck_trials", "audit_backtest"]),
69
+ validate: Object.freeze(["get_key", "validate_deflated_sharpe", "validate_overfitting", "validate_paper_evidence", "validate_breadth", "validate_track_record", "validate_backtest_length", "validate_haircut_sharpe", "validate_luck_trials", "validate_reality_check", "audit_backtest"]),
70
70
  receipts: Object.freeze(["get_receipt", "verify_receipt"]),
71
71
  company: Object.freeze(["company_financial_history"]),
72
72
  status: Object.freeze(["service_status"]),
@@ -260,6 +260,20 @@ export async function toolValidateOverfitting(session, args) {
260
260
  return validationText(session, response);
261
261
  }
262
262
 
263
+ // Every variant's returns, from the call or from a file on this machine (stdio only), tested for
264
+ // data snooping in one call; the file is read here and only its numbers are sent.
265
+ export async function toolValidateRealityCheck(session, args) {
266
+ const input = parseOrThrow(realityCheckInput, args, "validate_reality_check");
267
+ if (input.matrix_file && session.hosted) {
268
+ throw new Error("validate_reality_check: the hosted endpoint cannot read files on your machine; send matrix as numbers, or run the server locally with npx -y canli-validation-mcp.");
269
+ }
270
+ const { matrix_file: file, ...rest } = input;
271
+ const body = file ? { ...rest, matrix: readMatrixFile(file).matrix } : rest;
272
+ if (session.local) { const local = computeLocally("validate_reality_check", body); return validationText(session, local); }
273
+ const response = await validateRemote(session, "validate_reality_check", "/api/v1/validate/reality-check", body);
274
+ return validationText(session, response);
275
+ }
276
+
263
277
  export async function toolValidatePaperEvidence(session, args) {
264
278
  const body = parseOrThrow(paperEvidenceInput, args, "validate_paper_evidence");
265
279
  if (session.local) { const local = computeLocally("validate_paper_evidence", body); return validationText(session, local); }
@@ -549,6 +563,11 @@ export function registerTools(server, session) {
549
563
  { title: "Validate overfitting (CSCV)", annotations: { title: "Validate overfitting (CSCV)", ...WRITES_RECEIPT }, description: TOOL_DESCRIPTIONS.validate_overfitting, inputSchema: overfittingInput, outputSchema: validationOutput },
550
564
  (args) => toolValidateOverfitting(session, args),
551
565
  );
566
+ register(
567
+ "validate_reality_check",
568
+ { title: "Data-snooping tests (SPA, Reality Check, StepM)", annotations: { title: "Data-snooping tests (SPA, Reality Check, StepM)", ...WRITES_RECEIPT }, description: TOOL_DESCRIPTIONS.validate_reality_check, inputSchema: realityCheckToolShape, outputSchema: validationOutput },
569
+ (args) => toolValidateRealityCheck(session, args),
570
+ );
552
571
  register(
553
572
  "validate_paper_evidence",
554
573
  { title: "Validate paper evidence", annotations: { title: "Validate paper evidence", ...WRITES_RECEIPT }, description: TOOL_DESCRIPTIONS.validate_paper_evidence, inputSchema: paperEvidenceInput, outputSchema: validationOutput },