canli-validation-mcp 0.5.0 → 0.7.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -1,12 +1,19 @@
1
1
  # canli-validation-mcp
2
2
 
3
+ [![npm](https://img.shields.io/npm/v/canli-validation-mcp)](https://www.npmjs.com/package/canli-validation-mcp)
4
+ [![OpenSSF Scorecard](https://api.securityscorecards.dev/projects/github.com/arhancanli/canli-validation-mcp/badge)](https://scorecard.dev/viewer/?uri=github.com/arhancanli/canli-validation-mcp)
5
+ [![OpenSSF Best Practices](https://www.bestpractices.dev/projects/14954/badge)](https://www.bestpractices.dev/projects/14954)
6
+ [![Glama score](https://glama.ai/mcp/servers/arhancanli/canli-validation-mcp/badges/score.svg)](https://glama.ai/mcp/servers/arhancanli/canli-validation-mcp)
7
+
3
8
  An MCP (Model Context Protocol) server over canlicapital.com's free, keyed validation API. It
4
- gives a coding agent nine tools: issue a free key, run the five validators (deflated Sharpe,
5
- CSCV overfitting, paper-evidence conformance, breadth ceiling, minimum track record length),
6
- fetch a stored receipt, read service status, and read a company's reported financial history
7
- from SEC filings. Every tool returns the full API envelope as its result text, success or error,
8
- so the agent cannot see a number without the sentences beside it that say what the number does
9
- not establish.
9
+ gives a coding agent fourteen tools: issue a free key, run the eight validators (deflated Sharpe,
10
+ CSCV overfitting, paper-evidence conformance, breadth ceiling, minimum track record length,
11
+ minimum backtest length, haircut Sharpe ratio, luck-equivalent trials),
12
+ audit one backtest with three of them in a single call, fetch a stored receipt, verify a
13
+ receipt's signature offline, read service status, and read a company's reported financial history
14
+ from SEC filings. Every validation result carries, beside the number, the sentences that say what
15
+ it does not establish and the receipt that records it, success or error, so the agent cannot see a
16
+ number without its limits.
10
17
 
11
18
  This package is published to npm as [`canli-validation-mcp`](https://www.npmjs.com/package/canli-validation-mcp).
12
19
  Also listed on the official MCP Registry (`io.github.arhancanli/canli-validation-mcp`) and
@@ -16,11 +23,18 @@ test this package itself; see "Local checkout" near the bottom.
16
23
  This README describes the version in `package.json`. Unversioned `npx` runs npm's latest
17
24
  release; `npx -y canli-validation-mcp@<version>` pins one.
18
25
 
26
+ ## How well agents use it
27
+
28
+ A fixed benchmark gives models these tools and scores whether they pick the right one and return
29
+ the right answer, against ground truth computed from the same checked code. Results for three
30
+ models, with every run recorded, are in
31
+ [`bench/agent/README.md`](https://github.com/arhancanli/canlicapital/blob/main/mcp/bench/agent/README.md).
32
+
19
33
  ## What the API is (and is not)
20
34
 
21
35
  The engine is the product. The service runs your submitted numbers through the same honesty
22
36
  arithmetic canlicapital.com's own paper record runs on itself and hands back a verdict anyone can
23
- recompute from the receipt. It does not accept market data, does not sign receipts, does not
37
+ recompute from the receipt. It signs every receipt, does not accept market data, does not
24
38
  grade a strategy, and never saw your data source, its costs, or any lookahead in how a series was
25
39
  built. See `docs/superpowers/specs/2026-09-05-developer-key-validation-api-design.md` in the main
26
40
  repository for the full design.
@@ -35,6 +49,11 @@ repository for the full design.
35
49
  | `validate_paper_evidence` | `POST /api/v1/validate/paper-evidence` | yes |
36
50
  | `validate_breadth` | `POST /api/v1/validate/breadth` | yes |
37
51
  | `validate_track_record` | `POST /api/v1/validate/track-record` | yes |
52
+ | `validate_backtest_length` | `POST /api/v1/validate/backtest-length` | yes |
53
+ | `validate_haircut_sharpe` | `POST /api/v1/validate/haircut-sharpe` | yes |
54
+ | `validate_luck_trials` | `POST /api/v1/validate/luck-trials` | yes |
55
+ | `audit_backtest` | the deflated Sharpe, track record and, with `variants`, overfitting routes, one validation each | yes |
56
+ | `verify_receipt` | `GET /api/v1/receipts/{id}` when given an id; the checks run locally | no |
38
57
  | `get_receipt` | `GET /api/v1/receipts/{id}` | no |
39
58
  | `service_status` | `GET /api/v1/validate/status` | no |
40
59
  | `company_financial_history` | `GET /company-data/{cik}.json` (a ticker resolves through `GET /api/v1/company-tickers.json`) | no |
@@ -59,6 +78,24 @@ original SEC response and the record's own boundary sentence: these are accounti
59
78
  reported to the SEC, not market prices, returns or a recommendation. Companies and concepts
60
79
  outside the current release return an error with the available concepts listed.
61
80
 
81
+ ## Auditing a backtest in one call
82
+
83
+ `audit_backtest` takes one strategy's return series, the number of variants tried and their Sharpe
84
+ dispersion, and optionally every variant's returns. It runs `validate_deflated_sharpe` on the
85
+ series, then `validate_track_record` on the Sharpe, skew and kurtosis that check derived, then,
86
+ with `variants`, `validate_overfitting`. Each check is exactly what its own tool returns, with its
87
+ own receipt; the boundary sentences they share are stated once. A check that refuses (a Sharpe
88
+ that cannot beat the benchmark has no minimum track record) is reported as that check's error.
89
+ The audit adds no grade of its own. Through the API it uses one validation per check.
90
+
91
+ When the server runs on your machine, `returns_file` and `variants_file` take the path of the
92
+ backtest's output instead of the numbers: a CSV (comma, semicolon or tab separated, with or without
93
+ a header; date and label columns are ignored, an unnamed or counting index column is skipped and
94
+ reported) or a JSON array. `returns_column` picks the column when there are several. Only the
95
+ numbers are read; the hosted endpoint refuses file paths. An agent copying a long series into a
96
+ call can drop values, and the copy costs tokens; in our agent benchmark, reading the file instead
97
+ took three audit questions from 4 of 9 to 8 of 9 answered correctly on gpt-5.4-mini.
98
+
62
99
  ## Prompts, resources and structured results
63
100
 
64
101
  Clients that show MCP prompts offer two guided workflows: `validate_backtest` (deflated Sharpe, then
@@ -96,7 +133,9 @@ an array of objects.
96
133
  |---|---|---|
97
134
  | `CANLI_API_BASE` | `https://canlicapital.com` | Where the API lives. Point it at a preview deployment for testing. |
98
135
  | `CANLI_KEY` | unset | A key already issued from `POST /api/v1/keys`. When set, `get_key` sends no request and reports the key is already configured; every other tool sends it as `Authorization: Bearer <key>`. |
99
- | `CANLI_LOCAL` | unset | `1` or `true` runs the five validators on this machine (private local mode, below): no key, no network, no receipt. |
136
+ | `CANLI_FULL_ENVELOPE` | unset | `1` or `true` returns each validation's full API envelope instead of the compact result (below). |
137
+ | `CANLI_TOOLSETS` | all | Which tools to list: a comma-separated choice of `validate`, `receipts`, `company` and `status`, or `all`. An unknown name is refused at startup. See "Toolsets" below. |
138
+ | `CANLI_LOCAL` | unset | `1` or `true` runs the eight validators on this machine (private local mode, below): no key, no network, no receipt. |
100
139
 
101
140
  If `CANLI_KEY` is not set and local mode is off, call `get_key` once per session before the validators. The key it
102
141
  returns lives only in this process's memory for the life of the session; it is not written to
@@ -133,6 +172,9 @@ by every hosted caller. For your own quota, issue a free key (see
133
172
  claude mcp add --transport http canli https://canlicapital.com/mcp --header "Authorization: Bearer $CANLI_KEY"
134
173
  ```
135
174
 
175
+ Add `?toolsets=` to the URL to list only some tools (see "Toolsets" below), for example
176
+ `https://canlicapital.com/mcp?toolsets=company`.
177
+
136
178
  The endpoint is stateless. On it, `get_key` issues nothing and says which key is in use, because a
137
179
  key issued there would not reach the next request. A malformed Authorization header is refused
138
180
  rather than replaced with the shared key.
@@ -165,7 +207,7 @@ Run `claude mcp list` to confirm it is registered, and `claude mcp remove canli`
165
207
 
166
208
  ## Private local mode
167
209
 
168
- Set `CANLI_LOCAL=1` and the five validators run on your machine: nothing about the series you
210
+ Set `CANLI_LOCAL=1` and the eight validators run on your machine: nothing about the series you
169
211
  submit is sent to canlicapital.com, no key is needed, and no receipt is stored. The computation is
170
212
  the API's own, shipped byte for byte in `src/local` (a test fails if it drifts), so a local result
171
213
  equals the hosted one; it names no receipt id because none was made.
@@ -211,7 +253,7 @@ const result = await client.callTool({
211
253
  cross_trial_sharpe_sd_annualized: 0.5,
212
254
  },
213
255
  });
214
- console.log(result.content[0].text); // the full envelope, including limits and receipt.url
256
+ console.log(result.content[0].text); // the answer, its limits and its receipt
215
257
 
216
258
  await client.close();
217
259
  ```
@@ -231,8 +273,45 @@ Every envelope this server returns carries these sentences, verbatim, from the A
231
273
  request, 20000 observations per series, 200 variants per matrix.
232
274
 
233
275
  Each tool's description also states one of these sentences, so an agent sees the boundary before
234
- it calls the tool, not only after. No tool in this server strips `limits` or `receipt.url` from
235
- a response; the full envelope is always the result text.
276
+ it calls the tool, not only after.
277
+
278
+ ## Compact results (tokens)
279
+
280
+ A validation result is the answer (`data`), the sentences above except the quota line, and the
281
+ receipt's id and URL; an error keeps its error. The rest of the API envelope (schema, endpoint,
282
+ timestamps, claim and capital class, the human page, the source-file hashes and the quota line)
283
+ describes the service rather than the answer, and an agent pays for every token of it on every
284
+ call. It stays in the stored receipt, which `get_receipt` returns in full, and in `service_status`.
285
+ On a breadth result this is about half the text. Set `CANLI_FULL_ENVELOPE=1` to receive every field.
286
+
287
+ ## Toolsets (tokens)
288
+
289
+ A client sends the model the whole tool list on every turn, and it is most of each turn's prompt:
290
+ a validation result is a few hundred tokens, the list of all fourteen tools several thousand. A
291
+ client that needs one kind of tool can list only that kind, with `CANLI_TOOLSETS` (stdio) or
292
+ `?toolsets=` (hosted endpoint). The default is every tool.
293
+
294
+ | toolset | tools |
295
+ |---|---|
296
+ | `validate` | `get_key`, the eight validators, `audit_backtest` |
297
+ | `receipts` | `get_receipt`, `verify_receipt` |
298
+ | `company` | `company_financial_history` |
299
+ | `status` | `service_status` |
300
+
301
+ Measured with `bench/tool_list_tokens.py` (tokenizer: tiktoken `o200k_base`; other models'
302
+ tokenizers give different absolute counts), in the shape an OpenAI-style client sends the list:
303
+
304
+ | CANLI_TOOLSETS | tools | tokens per turn | of all |
305
+ |---|---|---|---|
306
+ | `all` | 14 | 3,888 | 100% |
307
+ | `validate` | 10 | 3,154 | 81% |
308
+ | `receipts` | 2 | 340 | 9% |
309
+ | `company` | 1 | 302 | 8% |
310
+ | `status` | 1 | 98 | 3% |
311
+
312
+ Providers cache a tool list that is identical from turn to turn and bill the cached part at a
313
+ fraction of the price (`test/tool-list-stable.test.mjs` keeps each list byte-stable); a smaller list
314
+ costs less either way.
236
315
 
237
316
  ## Local checkout
238
317
 
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "canli-validation-mcp",
3
- "version": "0.5.0",
3
+ "version": "0.7.0",
4
4
  "description": "MCP server for canlicapital.com's free validation API: deflated Sharpe, CSCV overfitting, paper-evidence conformance, and breadth ceiling, each returned as the full API envelope so the boundary language cannot be dropped.",
5
5
  "private": false,
6
6
  "type": "module",
@@ -21,7 +21,7 @@
21
21
  },
22
22
  "dependencies": {
23
23
  "@modelcontextprotocol/sdk": "1.30.0",
24
- "zod": "4.5.4"
24
+ "zod": "4.6.5"
25
25
  },
26
26
  "mcpName": "io.github.arhancanli/canli-validation-mcp",
27
27
  "repository": {
@@ -16,7 +16,7 @@ export const LIMITS = Object.freeze({
16
16
  export const LIMITS_TEXT = Object.freeze([
17
17
  "This verdict is about the series exactly as submitted. The service never saw the data source, its costs, survivorship, or any lookahead in how the series was built.",
18
18
  "A deflated Sharpe or overfitting probability above or below any threshold is not admission to anything and is not a forecast.",
19
- "The receipt is content-hashed and reproducible from the open-source core it names. It is not signed.",
19
+ "The receipt is content-hashed, reproducible from the open-source core it names, and signed with Ed25519 by a key published at https://canlicapital.com/.well-known/canli-receipt-keys.json.",
20
20
  `Quotas: ${LIMITS.validations_per_key_per_day} validations per key per UTC day, ${LIMITS.keys_per_client_per_day} keys per client per UTC day, ${LIMITS.max_body_bytes} bytes per validation request, ${LIMITS.max_key_revoke_body_bytes} bytes per key revocation request, ${LIMITS.max_observations} observations per series, ${LIMITS.max_variants} variants per matrix.`,
21
21
  ]);
22
22
 
@@ -125,6 +125,57 @@ function requireFinite(name, value) {
125
125
  if (!Number.isFinite(value)) throw new RangeError(`${name} must be a finite number`);
126
126
  }
127
127
 
128
+ // Expected maximum of N independent standard Normal draws (Bailey, Borwein, López de Prado and Zhu
129
+ // 2014, Proposition 2.1; Bailey and López de Prado 2014, the deflated Sharpe ratio): the Sharpe
130
+ // ratio, in units of its standard deviation, that the best of N skill-less trials is expected to show.
131
+ export function expectedMaxStandardNormal(trials) {
132
+ const n = Number(trials);
133
+ if (!Number.isInteger(n) || n < 2) throw new RangeError("Effective independent trials must be an integer of at least 2");
134
+ return (1 - EULER_MASCHERONI) * normalPpf(1 - 1 / n) + EULER_MASCHERONI * normalPpf(1 - 1 / (n * Math.E));
135
+ }
136
+
137
+ // Minimum Backtest Length (Bailey, Borwein, López de Prado and Zhu 2014, Theorem 3.1): the years of
138
+ // backtest needed so that the best of N skill-less trials is not expected to show an annualized
139
+ // Sharpe of targetSharpe in sample: ((1-g) Z^-1[1-1/N] + g Z^-1[1-1/(Ne)])^2 / targetSharpe^2,
140
+ // bounded above by 2 ln N / targetSharpe^2. Necessary, not sufficient, to avoid overfitting.
141
+ export function minimumBacktestLength({ trials, targetSharpe }) {
142
+ const target = Number(targetSharpe);
143
+ if (!(target > 0 && Number.isFinite(target))) throw new RangeError("The target Sharpe must be a positive number");
144
+ const expectedMax = expectedMaxStandardNormal(trials);
145
+ return {
146
+ years: (expectedMax / target) ** 2,
147
+ upper_bound_years: (2 * Math.log(Number(trials))) / target ** 2,
148
+ expected_max_sharpe_one_year: expectedMax,
149
+ };
150
+ }
151
+
152
+ // The largest number of independent trials whose best is still expected to stay below targetSharpe
153
+ // in sample over `years` of backtest (Eq. 3.1 solved for N). The expected maximum grows with N, so
154
+ // a doubling search then a bisection finds it exactly. Returns 1 when even two trials are too many.
155
+ export function maximumIndependentTrials({ years, targetSharpe }) {
156
+ const y = Number(years);
157
+ const target = Number(targetSharpe);
158
+ if (!(y > 0 && Number.isFinite(y))) throw new RangeError("Backtest years must be a positive number");
159
+ if (!(target > 0 && Number.isFinite(target))) throw new RangeError("The target Sharpe must be a positive number");
160
+ const ceiling = target * Math.sqrt(y);
161
+ const fits = (n) => expectedMaxStandardNormal(n) <= ceiling;
162
+ if (!fits(2)) return 1;
163
+ let lo = 2;
164
+ let hi = 4;
165
+ const LIMIT = 1e15;
166
+ while (fits(hi)) {
167
+ lo = hi;
168
+ if (hi >= LIMIT) return LIMIT;
169
+ hi = Math.min(hi * 2, LIMIT);
170
+ }
171
+ while (hi - lo > 1) {
172
+ const mid = Math.floor((lo + hi) / 2);
173
+ if (fits(mid)) lo = mid;
174
+ else hi = mid;
175
+ }
176
+ return lo;
177
+ }
178
+
128
179
  export function calculateDsr(input) {
129
180
  const values = Object.fromEntries(
130
181
  Object.entries(input).map(([key, value]) => [key, Number(value)]),
@@ -150,10 +201,7 @@ export function calculateDsr(input) {
150
201
  const observedSharpePerPeriod = values.observed_sharpe_annualized / annualizationScale;
151
202
  const trialSdPerPeriod = values.cross_trial_sharpe_sd_annualized / annualizationScale;
152
203
  const trialVariancePerPeriod = trialSdPerPeriod ** 2;
153
- const nTrials = values.effective_independent_trials;
154
- const quantile =
155
- (1 - EULER_MASCHERONI) * normalPpf(1 - 1 / nTrials) +
156
- EULER_MASCHERONI * normalPpf(1 - 1 / (nTrials * Math.E));
204
+ const quantile = expectedMaxStandardNormal(values.effective_independent_trials);
157
205
  const expectedMaxSharpePerPeriod = trialSdPerPeriod * quantile;
158
206
  const nonNormalityVarianceTerm =
159
207
  1 -
@@ -0,0 +1,96 @@
1
+ // The haircut Sharpe ratio: Harvey and Liu, "Backtesting", Journal of Portfolio Management, 2015.
2
+ // A Sharpe ratio found among several tests is converted to a t-statistic, its p-value is adjusted
3
+ // for the number of tests, and the adjusted p-value is converted back to the Sharpe ratio a single
4
+ // test would have needed: the haircut Sharpe ratio. Checked against the authors' own Haircut_SR.m
5
+ // (run in GNU Octave) and against R's p.adjust in js/haircut-paper-vectors.test.js.
6
+ import { studentTQuantileUpper, studentTUpper } from "./student-t.js";
7
+
8
+ // Lo (2002) for returns with first-order autocorrelation rho sampled q times a year: the annualized
9
+ // Sharpe ratio is scaled by [1 + (2 rho / (1 - rho)) (1 - (1 - rho^q) / (q (1 - rho)))]^(-1/2), the
10
+ // form Haircut_SR.m applies to an annualized input.
11
+ export function autocorrelationFactor(rho, periodsPerYear) {
12
+ const r = Number(rho);
13
+ const q = Number(periodsPerYear);
14
+ if (!(r > -1 && r < 1)) throw new RangeError("autocorrelation must be strictly between -1 and 1");
15
+ if (r === 0) return 1;
16
+ const inner = 1 + ((2 * r) / (1 - r)) * (1 - (1 - r ** q) / (q * (1 - r)));
17
+ if (!(inner > 0)) throw new RangeError("This autocorrelation gives no valid annualization factor");
18
+ return inner ** -0.5;
19
+ }
20
+
21
+ function requireCount(name, value, min) {
22
+ const n = Number(value);
23
+ if (!Number.isInteger(n) || n < min) throw new RangeError(`${name} must be an integer of at least ${min}`);
24
+ return n;
25
+ }
26
+
27
+ // Holm (step-down) and Benjamini-Hochberg-Yekutieli adjusted p-values for one member of a family,
28
+ // computed exactly as R's p.adjust computes them (stable order, ties by position).
29
+ export function familyAdjusted(pValues, index) {
30
+ const n = pValues.length;
31
+ const ascending = pValues.map((p, i) => [p, i]).sort((a, b) => a[0] - b[0] || a[1] - b[1]);
32
+ const holm = new Array(n);
33
+ let runningMax = 0;
34
+ ascending.forEach(([p, i], k) => {
35
+ runningMax = Math.max(runningMax, (n - k) * p);
36
+ holm[i] = Math.min(1, runningMax);
37
+ });
38
+ const harmonic = pValues.reduce((sum, _, k) => sum + 1 / (k + 1), 0);
39
+ const descending = pValues.map((p, i) => [p, i]).sort((a, b) => b[0] - a[0] || b[1] - a[1]);
40
+ const bhy = new Array(n);
41
+ let runningMin = Infinity;
42
+ descending.forEach(([p, i], k) => {
43
+ const rank = n - k;
44
+ runningMin = Math.min(runningMin, ((harmonic * n) / rank) * p);
45
+ bhy[i] = Math.min(1, runningMin);
46
+ });
47
+ return { holm: holm[index], bhy: bhy[index] };
48
+ }
49
+
50
+ export function haircutSharpe({ sharpeAnnualized, periodsPerYear, observations, tests, autocorrelation = 0, otherSharpesAnnualized }) {
51
+ const sharpe = Number(sharpeAnnualized);
52
+ const q = Number(periodsPerYear);
53
+ if (!(sharpe > 0 && Number.isFinite(sharpe))) throw new RangeError("The haircut applies to a positive, finite Sharpe ratio");
54
+ if (!(q > 0 && q <= 10000)) throw new RangeError("periods_per_year must be greater than 0 and at most 10000");
55
+ const T = requireCount("observations", observations, 3);
56
+ const others = otherSharpesAnnualized === undefined ? null : otherSharpesAnnualized.map(Number);
57
+ if (others && !others.every(Number.isFinite)) throw new RangeError("Every other Sharpe ratio must be a finite number");
58
+ const m = others ? others.length + 1 : requireCount("tests", tests, 1);
59
+ if (others && tests !== undefined && Number(tests) !== m) throw new RangeError("tests must equal the number of other Sharpe ratios plus one, or be left out");
60
+
61
+ const factor = autocorrelationFactor(autocorrelation, q);
62
+ const df = T - 1;
63
+ const tOf = (annual) => ((annual * factor) / Math.sqrt(q)) * Math.sqrt(T);
64
+ // Two-sided, as in Haircut_SR.m, but from the upper tail directly so that a large t keeps its p-value.
65
+ const pOf = (annual) => {
66
+ const t = tOf(annual);
67
+ return 2 * (t >= 0 ? studentTUpper(t, df) : studentTUpper(-t, df));
68
+ };
69
+ const srCorrected = sharpe * factor;
70
+ const tStat = tOf(sharpe);
71
+ const pSingle = pOf(sharpe);
72
+ const invert = (adjustedP) => {
73
+ const p = Math.min(1, adjustedP);
74
+ const tAdjusted = p >= 1 ? 0 : studentTQuantileUpper(p / 2, df);
75
+ const haircutSharpe = (tAdjusted / Math.sqrt(T)) * Math.sqrt(q);
76
+ return { adjusted_p: p, haircut_sharpe_annualized: haircutSharpe, haircut: (srCorrected - haircutSharpe) / srCorrected };
77
+ };
78
+ const result = {
79
+ sharpe_annualized_corrected: srCorrected,
80
+ autocorrelation_factor: factor,
81
+ t_statistic: tStat,
82
+ degrees_of_freedom: df,
83
+ p_value_single: pSingle,
84
+ tests: m,
85
+ bonferroni: invert(m * pSingle),
86
+ // Harvey and Liu's Eq. 4 for independent tests: 1 - (1 - p)^M.
87
+ independent: invert(-Math.expm1(m * Math.log1p(-pSingle))),
88
+ };
89
+ if (others) {
90
+ const family = [pSingle, ...others.map(pOf)];
91
+ const adjusted = familyAdjusted(family, 0);
92
+ result.holm = invert(adjusted.holm);
93
+ result.bhy = invert(adjusted.bhy);
94
+ }
95
+ return result;
96
+ }
@@ -0,0 +1,97 @@
1
+ // Luck-equivalent trials: how many skill-less strategies a search would have had to try for the best
2
+ // of them to reach the observed Sharpe ratio by luck alone. A reviewer's statistic, stated in the
3
+ // unit a research log records (trials), assembled from known results:
4
+ //
5
+ // - Under the null of no skill and normal returns, the Sharpe ratio's t-statistic, SR * sqrt(T) with
6
+ // the Sharpe per period, is exactly Student t with T - 1 degrees of freedom, so one trial reaches
7
+ // the observed Sharpe with probability p1 = P(t_{T-1} >= SR sqrt(T)).
8
+ // - The best of N independent skill-less trials reaches it with probability 1 - (1 - p1)^N (Sidak),
9
+ // so the N at which that probability equals q is N_q = ln(1 - q) / ln(1 - p1).
10
+ // - The expected maximum of N standard normals (Bailey, Borwein, Lopez de Prado and Zhu 2014), the
11
+ // deflated Sharpe ratio's benchmark, gives the N whose best is expected to reach it.
12
+ //
13
+ // Calibrated by Monte Carlo in js/luck-core.test.js; the full size study, with fixed seeds, is
14
+ // scripts/research/luck-trials-size-study.mjs and its output config/research/luck-trials-size-study.json.
15
+ // The size is correct for normal and for symmetric fat-tailed (Student t4) returns, conservative for
16
+ // positively skewed returns, and too small a count, so too kind to the strategy, for negatively
17
+ // skewed returns: with 252 observations a nominal 5% test rejected 10.8% of skill-less searches at
18
+ // skew -1.3 and 20.8% at skew -3.7.
19
+ import { expectedMaxStandardNormal } from "./dsr-core.js";
20
+ import { autocorrelationFactor } from "./haircut-core.js";
21
+ import { studentTUpper } from "./student-t.js";
22
+
23
+ export const TRIAL_CAP = 1e15;
24
+
25
+ function requirePositiveInteger(name, value, min) {
26
+ const n = Number(value);
27
+ if (!Number.isInteger(n) || n < min) throw new RangeError(`${name} must be an integer of at least ${min}`);
28
+ return n;
29
+ }
30
+
31
+ // P(one skill-less trial shows a Sharpe at least this high): the Student t upper tail.
32
+ export function singleTrialProbability({ sharpe, observations, periodsPerYear }) {
33
+ const sr = Number(sharpe);
34
+ const periods = Number(periodsPerYear);
35
+ const t = requirePositiveInteger("observations", observations, 3);
36
+ if (!Number.isFinite(sr)) throw new RangeError("The Sharpe ratio must be a finite number");
37
+ if (!(periods > 0 && Number.isFinite(periods))) throw new RangeError("periods per year must be a positive number");
38
+ const tStatistic = (sr / Math.sqrt(periods)) * Math.sqrt(t);
39
+ return { t_statistic: tStatistic, probability: studentTUpper(tStatistic, t - 1) };
40
+ }
41
+
42
+ // The N at which the best of N skill-less trials reaches the Sharpe with probability q, as a real
43
+ // number. Below 1 means a single trial already reaches it with probability above q; the count is
44
+ // capped at TRIAL_CAP, beyond which the tail probability is below what double precision resolves.
45
+ export function trialsAtProbability(probability, q) {
46
+ const p = Number(probability);
47
+ const level = Number(q);
48
+ if (!(level > 0 && level < 1)) throw new RangeError("The probability level must be strictly between 0 and 1");
49
+ if (!(p >= 0 && p <= 1)) throw new RangeError("The single-trial probability must be between 0 and 1");
50
+ if (p === 0) return TRIAL_CAP;
51
+ return Math.min(TRIAL_CAP, Math.log1p(-level) / Math.log1p(-p));
52
+ }
53
+
54
+ // P(the best of N skill-less trials shows a Sharpe at least this high): 1 - (1 - p1)^N.
55
+ export function bestOfTrialsProbability(probability, trials) {
56
+ const n = requirePositiveInteger("trials", trials, 1);
57
+ return -Math.expm1(n * Math.log1p(-Number(probability)));
58
+ }
59
+
60
+ // The largest N whose best skill-less trial is expected (the deflated Sharpe ratio's approximation)
61
+ // to stay at or below the t-statistic; 1 when even two are expected to exceed it.
62
+ export function expectedMaximumTrials(tStatistic) {
63
+ const ceiling = Number(tStatistic);
64
+ const fits = (n) => expectedMaxStandardNormal(n) <= ceiling;
65
+ if (!fits(2)) return 1;
66
+ let lo = 2;
67
+ let hi = 4;
68
+ while (fits(hi)) {
69
+ lo = hi;
70
+ if (hi >= TRIAL_CAP) return TRIAL_CAP;
71
+ hi = Math.min(hi * 2, TRIAL_CAP);
72
+ }
73
+ while (hi - lo > 1) {
74
+ const mid = Math.floor((lo + hi) / 2);
75
+ if (fits(mid)) lo = mid;
76
+ else hi = mid;
77
+ }
78
+ return lo;
79
+ }
80
+
81
+ // With the returns' lag-1 autocorrelation, the Sharpe is first corrected as Lo (2002), the form
82
+ // js/haircut-core.js applies; the Null Zoo measured that without it, positively autocorrelated
83
+ // returns (0.2) make every best-of-N test reject about four times as often as its level.
84
+ export function luckEquivalentTrials({ sharpe, observations, periodsPerYear, trials, autocorrelation = 0 }) {
85
+ const factor = autocorrelationFactor(autocorrelation, periodsPerYear);
86
+ const single = singleTrialProbability({ sharpe: Number(sharpe) * factor, observations, periodsPerYear });
87
+ const result = {
88
+ autocorrelation_factor: factor,
89
+ t_statistic: single.t_statistic,
90
+ single_trial_probability: single.probability,
91
+ trials_for_even_odds: trialsAtProbability(single.probability, 0.5),
92
+ trials_for_five_percent: trialsAtProbability(single.probability, 0.05),
93
+ trials_expected_to_match: expectedMaximumTrials(single.t_statistic),
94
+ };
95
+ if (trials !== undefined) result.best_of_trials_probability = bestOfTrialsProbability(single.probability, trials);
96
+ return result;
97
+ }
@@ -0,0 +1,54 @@
1
+ // js/receipt-statement.js
2
+ // What a canlicapital.com receipt signature covers, and how to check one. Shared by the API, which
3
+ // signs (api/_lib/receipt-signature.js), and by the MCP package, which verifies offline
4
+ // (mcp/src/local mirrors this file byte for byte), so the two cannot disagree about the bytes signed.
5
+ //
6
+ // The statement is the canonical JSON of the receipt's content: its id, the validator's endpoint,
7
+ // the sha256 of the input, the sha256 of the output, and the sha256 of every source file that
8
+ // computed it. The id is itself the first 24 hex characters of sha256 over the canonical JSON of
9
+ // {endpoint, input_sha256, output, bindings}, so the signature ties the result to the exact code.
10
+ import { createHash, createPublicKey, verify } from "node:crypto";
11
+
12
+ import { canonicalJson } from "../scripts/canonical-json.mjs";
13
+
14
+ export const SIGNATURE_SCHEMA = "canli.receipt-signature.v1";
15
+ export const KEYS_URL = "https://canlicapital.com/.well-known/canli-receipt-keys.json";
16
+
17
+ const sha256Hex = (text) => createHash("sha256").update(text).digest("hex");
18
+
19
+ export function outputSha256(output) {
20
+ return `sha256:${sha256Hex(canonicalJson(output))}`;
21
+ }
22
+
23
+ export function receiptId({ endpoint, input_sha256, output, bindings }) {
24
+ return sha256Hex(canonicalJson({ endpoint, input_sha256, output, bindings })).slice(0, 24);
25
+ }
26
+
27
+ export function receiptStatement({ id, endpoint, input_sha256, output_sha256, bindings }) {
28
+ return canonicalJson({ schema: SIGNATURE_SCHEMA, id, endpoint, input_sha256, output_sha256, bindings });
29
+ }
30
+
31
+ // The key id is derived from the key: the first 16 hex characters of sha256 over its 32 raw bytes.
32
+ export function keyIdFor(rawPublicKey) {
33
+ return sha256Hex(rawPublicKey).slice(0, 16);
34
+ }
35
+
36
+ // Checks a stored receipt end to end: its output hashes to output_sha256, its content hashes to its
37
+ // id, its signature verifies over the statement, and the signing key is one canlicapital.com
38
+ // publishes. Returns every check so a caller can see which one failed.
39
+ export function verifyReceipt({ id, endpoint, input_sha256, output, bindings, signature }, keys) {
40
+ const checks = { id_matches_content: false, signature_valid: false, key_published: false };
41
+ const computedOutputSha = outputSha256(output);
42
+ checks.id_matches_content = receiptId({ endpoint, input_sha256, output, bindings }) === id;
43
+ const key = signature ? (keys ?? []).find((k) => k.key_id === signature.key_id && k.alg === "Ed25519") : undefined;
44
+ checks.key_published = Boolean(key);
45
+ if (key && signature?.value && signature.schema === SIGNATURE_SCHEMA) {
46
+ const raw = Buffer.from(key.x, "base64url");
47
+ if (keyIdFor(raw) === key.key_id) {
48
+ const publicKey = createPublicKey({ key: { kty: "OKP", crv: "Ed25519", x: key.x }, format: "jwk" });
49
+ const statement = receiptStatement({ id, endpoint, input_sha256, output_sha256: computedOutputSha, bindings });
50
+ checks.signature_valid = verify(null, Buffer.from(statement, "utf8"), publicKey, Buffer.from(signature.value, "base64url"));
51
+ }
52
+ }
53
+ return { valid: checks.id_matches_content && checks.signature_valid && checks.key_published, checks, output_sha256: computedOutputSha, key_id: signature?.key_id ?? null };
54
+ }
@@ -0,0 +1,139 @@
1
+ // Student's t distribution to full double precision, for the haircut Sharpe ratio (Harvey and Liu,
2
+ // "Backtesting", 2015), which tests a Sharpe ratio's t-statistic against t with N - 1 degrees of
3
+ // freedom. Checked against R's pt and qt (js/fixtures/student-t-reference.json).
4
+ //
5
+ // The upper tail is computed directly, never as 1 - cdf: the authors' Haircut_SR.m computes
6
+ // 2 * (1 - tcdf(t, N - 1)), which rounds to 0 once t passes about 8 and makes the haircut Sharpe
7
+ // ratio infinite. P(T > t) = I_x(df / 2, 1 / 2) / 2 with x = df / (df + t^2) keeps full relative
8
+ // precision far into the tail.
9
+
10
+ const HALF_LOG_TWO_PI = 0.5 * Math.log(2 * Math.PI);
11
+
12
+ // Stirling's correction ln Gamma(x) - [(x - 1/2) ln x - x + ln(2 pi) / 2], for x >= 10: the series
13
+ // to the 1/x^13 term, whose error there is below 1e-16.
14
+ function stirlingCorrection(x) {
15
+ const inv = 1 / x;
16
+ const inv2 = inv * inv;
17
+ return inv * (1 / 12 + inv2 * (-1 / 360 + inv2 * (1 / 1260 + inv2 * (-1 / 1680 + inv2 * (1 / 1188 + inv2 * (-691 / 360360 + inv2 * (1 / 156)))))));
18
+ }
19
+
20
+ // ln Gamma(x) for x > 0: shifted to x >= 10, where the series applies; ln Gamma(x) = ln Gamma(x + 1) -
21
+ // ln x undoes the shift.
22
+ export function logGamma(value) {
23
+ let x = Number(value);
24
+ if (!(x > 0) || !Number.isFinite(x)) throw new RangeError("logGamma needs a positive finite number");
25
+ let shift = 0;
26
+ while (x < 10) {
27
+ shift -= Math.log(x);
28
+ x += 1;
29
+ }
30
+ return (x - 0.5) * Math.log(x) - x + HALF_LOG_TWO_PI + stirlingCorrection(x) + shift;
31
+ }
32
+
33
+ // ln Gamma(a + b) - ln Gamma(a) without subtracting two large numbers when a is large (degrees of
34
+ // freedom in the thousands): (a - 1/2) ln(1 + b/a) + b ln(a + b) - b + the two corrections.
35
+ function logGammaRatio(a, b) {
36
+ if (a < 10) return logGamma(a + b) - logGamma(a);
37
+ return (a - 0.5) * Math.log1p(b / a) + b * Math.log(a + b) - b + stirlingCorrection(a + b) - stirlingCorrection(a);
38
+ }
39
+
40
+ // Continued fraction for the incomplete beta function (modified Lentz).
41
+ function betaFraction(a, b, x) {
42
+ const TINY = 1e-300;
43
+ const qab = a + b;
44
+ const qap = a + 1;
45
+ const qam = a - 1;
46
+ let c = 1;
47
+ let d = 1 - (qab * x) / qap;
48
+ if (Math.abs(d) < TINY) d = TINY;
49
+ d = 1 / d;
50
+ let h = d;
51
+ for (let m = 1; m <= 100000; m += 1) {
52
+ const m2 = 2 * m;
53
+ let aa = (m * (b - m) * x) / ((qam + m2) * (a + m2));
54
+ d = 1 + aa * d;
55
+ if (Math.abs(d) < TINY) d = TINY;
56
+ c = 1 + aa / c;
57
+ if (Math.abs(c) < TINY) c = TINY;
58
+ d = 1 / d;
59
+ h *= d * c;
60
+ aa = (-(a + m) * (qab + m) * x) / ((a + m2) * (qap + m2));
61
+ d = 1 + aa * d;
62
+ if (Math.abs(d) < TINY) d = TINY;
63
+ c = 1 + aa / c;
64
+ if (Math.abs(c) < TINY) c = TINY;
65
+ d = 1 / d;
66
+ const delta = d * c;
67
+ h *= delta;
68
+ if (Math.abs(delta - 1) < 1e-16) return h;
69
+ }
70
+ throw new Error("The incomplete beta continued fraction did not converge");
71
+ }
72
+
73
+ // Regularized incomplete beta I_x(a, b) for b = 1/2 (all Student t needs), with y = 1 - x and both
74
+ // logarithms passed separately so that nothing is lost to cancellation when x or y is tiny.
75
+ function incompleteBeta(x, y, logX, logY, a, b) {
76
+ if (x <= 0) return 0;
77
+ if (y <= 0) return 1;
78
+ const front = Math.exp(logGammaRatio(a, b) - logGamma(b) + a * logX + b * logY);
79
+ if (x < (a + 1) / (a + b + 2)) return (front * betaFraction(a, b, x)) / a;
80
+ return 1 - (front * betaFraction(b, a, y)) / b;
81
+ }
82
+
83
+ function requireDf(df) {
84
+ const v = Number(df);
85
+ if (!(v > 0) || !Number.isFinite(v)) throw new RangeError("Degrees of freedom must be a positive finite number");
86
+ return v;
87
+ }
88
+
89
+ // P(T > t) for T ~ t(df).
90
+ export function studentTUpper(t, df) {
91
+ const v = requireDf(df);
92
+ const value = Number(t);
93
+ if (Number.isNaN(value)) return Number.NaN;
94
+ if (value === 0) return 0.5;
95
+ if (value === Infinity) return 0;
96
+ if (value === -Infinity) return 1;
97
+ const t2 = value * value;
98
+ const logX = -Math.log1p(t2 / v);
99
+ const logY = Math.log(t2) - Math.log(v + t2);
100
+ const half = 0.5 * incompleteBeta(Math.exp(logX), t2 / (v + t2), logX, logY, v / 2, 0.5);
101
+ return value > 0 ? half : 1 - half;
102
+ }
103
+
104
+ export function studentTPdf(t, df) {
105
+ const v = requireDf(df);
106
+ const value = Number(t);
107
+ return Math.exp(logGammaRatio(v / 2, 0.5) - 0.5 * Math.log(v * Math.PI) - ((v + 1) / 2) * Math.log1p((value * value) / v));
108
+ }
109
+
110
+ // The t with P(T > t) = upper: bisection on a bracket (the tail can be very heavy for small df),
111
+ // then Newton steps on log P(T > t), which is smooth and keeps relative precision in the tail.
112
+ export function studentTQuantileUpper(upper, df) {
113
+ const v = requireDf(df);
114
+ const p = Number(upper);
115
+ if (!(p > 0 && p < 1)) throw new RangeError("The upper-tail probability must be strictly between 0 and 1");
116
+ if (p === 0.5) return 0;
117
+ if (p > 0.5) return -studentTQuantileUpper(1 - p, v);
118
+ let lo = 0;
119
+ let hi = 1;
120
+ while (studentTUpper(hi, v) > p) {
121
+ lo = hi;
122
+ hi *= 2;
123
+ if (!Number.isFinite(hi)) throw new RangeError("The upper-tail probability is too small for these degrees of freedom");
124
+ }
125
+ for (let i = 0; i < 200 && hi - lo > 1e-12 * hi; i += 1) {
126
+ const mid = 0.5 * (lo + hi);
127
+ if (studentTUpper(mid, v) > p) lo = mid;
128
+ else hi = mid;
129
+ }
130
+ let t = 0.5 * (lo + hi);
131
+ const target = Math.log(p);
132
+ for (let i = 0; i < 4; i += 1) {
133
+ const tail = studentTUpper(t, v);
134
+ const step = (Math.log(tail) - target) / (-studentTPdf(t, v) / tail);
135
+ if (!Number.isFinite(step)) break;
136
+ t -= step;
137
+ }
138
+ return t;
139
+ }