canli-validation-mcp 0.6.0 → 0.7.1

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (3) hide show
  1. package/README.md +65 -1
  2. package/package.json +1 -1
  3. package/src/server.mjs +72 -24
package/README.md CHANGED
@@ -2,8 +2,37 @@
2
2
 
3
3
  [![npm](https://img.shields.io/npm/v/canli-validation-mcp)](https://www.npmjs.com/package/canli-validation-mcp)
4
4
  [![OpenSSF Scorecard](https://api.securityscorecards.dev/projects/github.com/arhancanli/canli-validation-mcp/badge)](https://scorecard.dev/viewer/?uri=github.com/arhancanli/canli-validation-mcp)
5
+ [![OpenSSF Best Practices](https://www.bestpractices.dev/projects/14954/badge)](https://www.bestpractices.dev/projects/14954)
5
6
  [![Glama score](https://glama.ai/mcp/servers/arhancanli/canli-validation-mcp/badges/score.svg)](https://glama.ai/mcp/servers/arhancanli/canli-validation-mcp)
6
7
 
8
+ **Is your best backtest real, or just the luckiest of the variants you tried?** This MCP server
9
+ lets Claude, Cursor or any MCP client answer that with the standard corrections: the deflated
10
+ Sharpe ratio, the CSCV probability of backtest overfitting, the minimum track record length, the
11
+ haircut Sharpe ratio, and luck-equivalent trials. Free and MIT-licensed.
12
+
13
+ ## Quick start
14
+
15
+ ```bash
16
+ # Claude Code, computed on your machine: no key, nothing sent anywhere
17
+ claude mcp add canli-local --env CANLI_LOCAL=1 -- npx -y canli-validation-mcp
18
+
19
+ # the same tools with signed, stored receipts from the free API (a free key is issued on first use)
20
+ claude mcp add canli -- npx -y canli-validation-mcp
21
+
22
+ # nothing to install: the hosted endpoint
23
+ claude mcp add --transport http canli https://canlicapital.com/mcp
24
+ ```
25
+
26
+ Then ask, for example: *"I tried 229 variants and kept the best: an annualised Sharpe of
27
+ 1.5 over 730 daily returns (365 a year), skew -0.5, kurtosis 5, and the variants'
28
+ Sharpe ratios spread by 0.57. Is it real?"* The assistant calls `validate_deflated_sharpe`: the
29
+ best of 229 skill-less variants would reach 1.60 by luck alone, so the probability that the
30
+ Sharpe is above zero falls from 98.1% to 44.4% once the search is counted. Claude Desktop and any
31
+ other stdio client (Cursor, VS Code) run the same `npx` command; see "Claude Desktop" and
32
+ "Generic stdio client" below.
33
+
34
+ ## What it does
35
+
7
36
  An MCP (Model Context Protocol) server over canlicapital.com's free, keyed validation API. It
8
37
  gives a coding agent fourteen tools: issue a free key, run the eight validators (deflated Sharpe,
9
38
  CSCV overfitting, paper-evidence conformance, breadth ceiling, minimum track record length,
@@ -133,6 +162,7 @@ an array of objects.
133
162
  | `CANLI_API_BASE` | `https://canlicapital.com` | Where the API lives. Point it at a preview deployment for testing. |
134
163
  | `CANLI_KEY` | unset | A key already issued from `POST /api/v1/keys`. When set, `get_key` sends no request and reports the key is already configured; every other tool sends it as `Authorization: Bearer <key>`. |
135
164
  | `CANLI_FULL_ENVELOPE` | unset | `1` or `true` returns each validation's full API envelope instead of the compact result (below). |
165
+ | `CANLI_TOOLSETS` | all | Which tools to list: a comma-separated choice of `validate`, `receipts`, `company` and `status`, or `all`. An unknown name is refused at startup. See "Toolsets" below. |
136
166
  | `CANLI_LOCAL` | unset | `1` or `true` runs the eight validators on this machine (private local mode, below): no key, no network, no receipt. |
137
167
 
138
168
  If `CANLI_KEY` is not set and local mode is off, call `get_key` once per session before the validators. The key it
@@ -163,13 +193,18 @@ claude mcp add --transport http canli https://canlicapital.com/mcp
163
193
  ```
164
194
 
165
195
  Without a key, requests run under a shared anonymous key, so the daily validation quota is shared
166
- by every hosted caller. For your own quota, issue a free key (see
196
+ by every hosted caller. When that shared quota is used up for the day, validations are still
197
+ answered, computed by the same code on the hosted endpoint, but without a stored receipt; the
198
+ result says so. For your own quota, issue a free key (see
167
199
  [/developers](https://canlicapital.com/developers#quickstart)) and send it as a header:
168
200
 
169
201
  ```bash
170
202
  claude mcp add --transport http canli https://canlicapital.com/mcp --header "Authorization: Bearer $CANLI_KEY"
171
203
  ```
172
204
 
205
+ Add `?toolsets=` to the URL to list only some tools (see "Toolsets" below), for example
206
+ `https://canlicapital.com/mcp?toolsets=company`.
207
+
173
208
  The endpoint is stateless. On it, `get_key` issues nothing and says which key is in use, because a
174
209
  key issued there would not reach the next request. A malformed Authorization header is refused
175
210
  rather than replaced with the shared key.
@@ -279,6 +314,35 @@ describes the service rather than the answer, and an agent pays for every token
279
314
  call. It stays in the stored receipt, which `get_receipt` returns in full, and in `service_status`.
280
315
  On a breadth result this is about half the text. Set `CANLI_FULL_ENVELOPE=1` to receive every field.
281
316
 
317
+ ## Toolsets (tokens)
318
+
319
+ A client sends the model the whole tool list on every turn, and it is most of each turn's prompt:
320
+ a validation result is a few hundred tokens, the list of all fourteen tools several thousand. A
321
+ client that needs one kind of tool can list only that kind, with `CANLI_TOOLSETS` (stdio) or
322
+ `?toolsets=` (hosted endpoint). The default is every tool.
323
+
324
+ | toolset | tools |
325
+ |---|---|
326
+ | `validate` | `get_key`, the eight validators, `audit_backtest` |
327
+ | `receipts` | `get_receipt`, `verify_receipt` |
328
+ | `company` | `company_financial_history` |
329
+ | `status` | `service_status` |
330
+
331
+ Measured with `bench/tool_list_tokens.py` (tokenizer: tiktoken `o200k_base`; other models'
332
+ tokenizers give different absolute counts), in the shape an OpenAI-style client sends the list:
333
+
334
+ | CANLI_TOOLSETS | tools | tokens per turn | of all |
335
+ |---|---|---|---|
336
+ | `all` | 14 | 3,888 | 100% |
337
+ | `validate` | 10 | 3,154 | 81% |
338
+ | `receipts` | 2 | 340 | 9% |
339
+ | `company` | 1 | 302 | 8% |
340
+ | `status` | 1 | 98 | 3% |
341
+
342
+ Providers cache a tool list that is identical from turn to turn and bill the cached part at a
343
+ fraction of the price (`test/tool-list-stable.test.mjs` keeps each list byte-stable); a smaller list
344
+ costs less either way.
345
+
282
346
  ## Local checkout
283
347
 
284
348
  Only needed to develop or test this package itself, not to run the published one.
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "canli-validation-mcp",
3
- "version": "0.6.0",
3
+ "version": "0.7.1",
4
4
  "description": "MCP server for canlicapital.com's free validation API: deflated Sharpe, CSCV overfitting, paper-evidence conformance, and breadth ceiling, each returned as the full API envelope so the boundary language cannot be dropped.",
5
5
  "private": false,
6
6
  "type": "module",
package/src/server.mjs CHANGED
@@ -66,7 +66,29 @@ export function configuredLocal(value) {
66
66
  return v === "1" || v?.toLowerCase() === "true";
67
67
  }
68
68
 
69
- export function createSession({ base, fetchImpl, envKey, timeoutMs = REQUEST_TIMEOUT_MS, hosted, local, fullEnvelope, receiptKeys } = {}) {
69
+ // Toolsets: which tools the server lists. The tool list is re-sent to the model on every turn and
70
+ // is most of each turn's prompt (README, "Toolsets"; bench/tool_list_tokens.py measures it), so a
71
+ // client that needs one kind of tool can load only that kind. Default: all.
72
+ export const TOOLSETS = Object.freeze({
73
+ validate: Object.freeze(["get_key", "validate_deflated_sharpe", "validate_overfitting", "validate_paper_evidence", "validate_breadth", "validate_track_record", "validate_backtest_length", "validate_haircut_sharpe", "validate_luck_trials", "audit_backtest"]),
74
+ receipts: Object.freeze(["get_receipt", "verify_receipt"]),
75
+ company: Object.freeze(["company_financial_history"]),
76
+ status: Object.freeze(["service_status"]),
77
+ });
78
+
79
+ // CANLI_TOOLSETS=validate,company (or ?toolsets= on the hosted endpoint): a comma-separated list of
80
+ // TOOLSETS names, or "all". Empty or unsubstituted means all; an unknown name is refused, so a typo
81
+ // never silently leaves a client without the tools it asked for.
82
+ export function configuredToolsets(value) {
83
+ const v = configuredKey(value);
84
+ if (!v || v.trim().toLowerCase() === "all") return Object.keys(TOOLSETS);
85
+ const names = [...new Set(v.split(",").map((s) => s.trim().toLowerCase()).filter(Boolean))];
86
+ const unknown = names.filter((n) => !Object.hasOwn(TOOLSETS, n));
87
+ if (unknown.length || !names.length) throw new Error(`Unknown toolset ${unknown.join(", ") || "(none)"}; choose from ${Object.keys(TOOLSETS).join(", ")} or all`);
88
+ return names;
89
+ }
90
+
91
+ export function createSession({ base, fetchImpl, envKey, timeoutMs = REQUEST_TIMEOUT_MS, hosted, local, fullEnvelope, receiptKeys, toolsets } = {}) {
70
92
  if (!Number.isSafeInteger(timeoutMs) || timeoutMs <= 0) throw new Error("Request timeout must be a positive integer");
71
93
  return {
72
94
  base: base ?? process.env.CANLI_API_BASE ?? DEFAULT_BASE,
@@ -81,6 +103,7 @@ export function createSession({ base, fetchImpl, envKey, timeoutMs = REQUEST_TIM
81
103
  fullEnvelope: fullEnvelope ?? configuredFullEnvelope(process.env.CANLI_FULL_ENVELOPE),
82
104
  // Tests pass their own keys; everyone else verifies against the bundled published keys.
83
105
  receiptKeys: receiptKeys ?? undefined,
106
+ toolsets: toolsets ?? configuredToolsets(process.env.CANLI_TOOLSETS),
84
107
  };
85
108
  }
86
109
 
@@ -109,6 +132,29 @@ async function callApi(session, { path, method = "GET", body }) {
109
132
  return { envelope, failed: res.status >= 400 || Boolean(envelope?.error) };
110
133
  }
111
134
 
135
+ // The hosted endpoint's shared anonymous key has one daily quota for every caller who connects
136
+ // without a key of their own. When it is used up, the answer is still computed, by the same code
137
+ // the API runs (src/local, byte for byte), on the hosted endpoint; what the caller loses is the
138
+ // stored, signed receipt, and the note says how to get one. A caller's own key, or a missing
139
+ // shared key, still gets the API's refusal unchanged.
140
+ export const SHARED_QUOTA_NOTE = "The shared anonymous quota of this hosted endpoint is used up for today (it resets at 00:00 UTC), so this result was computed by the same code on the hosted endpoint and no receipt was stored. For a receipt and a quota of your own, get a free key at https://canlicapital.com/developers#quickstart and send it as 'Authorization: Bearer <key>'. To run with no quota at all, on your own machine: npx -y canli-validation-mcp with CANLI_LOCAL=1.";
141
+
142
+ async function validateRemote(session, tool, path, body) {
143
+ let response = await callApi(session, { path, method: "POST", body });
144
+ // stdio without CANLI_KEY: the first validation used to come back 401 and the model had to work
145
+ // out that get_key comes first. Issue the free key once, as get_key would, and retry once.
146
+ if (response.failed && response.envelope?.error?.code === "unauthorized" && !session.hosted && !session.key && !session.envKey) {
147
+ const issued = await callApi(session, { path: "/api/v1/keys", method: "POST", body: { label: "auto" } });
148
+ if (!issued.failed && issued.envelope?.data?.key) {
149
+ session.key = issued.envelope.data.key;
150
+ response = await callApi(session, { path, method: "POST", body });
151
+ }
152
+ }
153
+ if (!response.failed || session.hosted?.keySource !== "shared" || response.envelope?.error?.code !== "quota_exhausted") return response;
154
+ const fallback = computeLocally(tool, body);
155
+ return { ...fallback, envelope: { ...fallback.envelope, computed: "hosted_without_receipt", note: SHARED_QUOTA_NOTE } };
156
+ }
157
+
112
158
  // Compact context (0.3.0): minified JSON. Indentation is whitespace an agent pays for in tokens and
113
159
  // never reads; every field, boundary sentence and provenance value is kept. Measured on the live
114
160
  // Apple StockholdersEquity record (README, "Compact context"): 2,214 -> 1,560 tokens minified,
@@ -207,56 +253,56 @@ export async function toolValidateDeflatedSharpe(session, args) {
207
253
  );
208
254
  }
209
255
  if (session.local) { const local = computeLocally("validate_deflated_sharpe", parsed.data); return validationText(session, local); }
210
- const response = await callApi(session, { path: "/api/v1/validate/deflated-sharpe", method: "POST", body: parsed.data });
256
+ const response = await validateRemote(session, "validate_deflated_sharpe", "/api/v1/validate/deflated-sharpe", parsed.data);
211
257
  return validationText(session, response);
212
258
  }
213
259
 
214
260
  export async function toolValidateOverfitting(session, args) {
215
261
  const body = parseOrThrow(overfittingInput, args, "validate_overfitting");
216
262
  if (session.local) { const local = computeLocally("validate_overfitting", body); return validationText(session, local); }
217
- const response = await callApi(session, { path: "/api/v1/validate/overfitting", method: "POST", body });
263
+ const response = await validateRemote(session, "validate_overfitting", "/api/v1/validate/overfitting", body);
218
264
  return validationText(session, response);
219
265
  }
220
266
 
221
267
  export async function toolValidatePaperEvidence(session, args) {
222
268
  const body = parseOrThrow(paperEvidenceInput, args, "validate_paper_evidence");
223
269
  if (session.local) { const local = computeLocally("validate_paper_evidence", body); return validationText(session, local); }
224
- const response = await callApi(session, { path: "/api/v1/validate/paper-evidence", method: "POST", body });
270
+ const response = await validateRemote(session, "validate_paper_evidence", "/api/v1/validate/paper-evidence", body);
225
271
  return validationText(session, response);
226
272
  }
227
273
 
228
274
  export async function toolValidateBreadth(session, args) {
229
275
  const body = parseOrThrow(breadthInput, args, "validate_breadth");
230
276
  if (session.local) { const local = computeLocally("validate_breadth", body); return validationText(session, local); }
231
- const response = await callApi(session, { path: "/api/v1/validate/breadth", method: "POST", body });
277
+ const response = await validateRemote(session, "validate_breadth", "/api/v1/validate/breadth", body);
232
278
  return validationText(session, response);
233
279
  }
234
280
 
235
281
  export async function toolValidateTrackRecord(session, args) {
236
282
  const body = parseOrThrow(trackRecordInput, args, "validate_track_record");
237
283
  if (session.local) { const local = computeLocally("validate_track_record", body); return validationText(session, local); }
238
- const response = await callApi(session, { path: "/api/v1/validate/track-record", method: "POST", body });
284
+ const response = await validateRemote(session, "validate_track_record", "/api/v1/validate/track-record", body);
239
285
  return validationText(session, response);
240
286
  }
241
287
 
242
288
  export async function toolValidateBacktestLength(session, args) {
243
289
  const body = parseOrThrow(backtestLengthInput, args, "validate_backtest_length");
244
290
  if (session.local) { const local = computeLocally("validate_backtest_length", body); return validationText(session, local); }
245
- const response = await callApi(session, { path: "/api/v1/validate/backtest-length", method: "POST", body });
291
+ const response = await validateRemote(session, "validate_backtest_length", "/api/v1/validate/backtest-length", body);
246
292
  return validationText(session, response);
247
293
  }
248
294
 
249
295
  export async function toolValidateHaircutSharpe(session, args) {
250
296
  const body = parseOrThrow(haircutSharpeInput, args, "validate_haircut_sharpe");
251
297
  if (session.local) { const local = computeLocally("validate_haircut_sharpe", body); return validationText(session, local); }
252
- const response = await callApi(session, { path: "/api/v1/validate/haircut-sharpe", method: "POST", body });
298
+ const response = await validateRemote(session, "validate_haircut_sharpe", "/api/v1/validate/haircut-sharpe", body);
253
299
  return validationText(session, response);
254
300
  }
255
301
 
256
302
  export async function toolValidateLuckTrials(session, args) {
257
303
  const body = parseOrThrow(luckTrialsInput, args, "validate_luck_trials");
258
304
  if (session.local) { const local = computeLocally("validate_luck_trials", body); return validationText(session, local); }
259
- const response = await callApi(session, { path: "/api/v1/validate/luck-trials", method: "POST", body });
305
+ const response = await validateRemote(session, "validate_luck_trials", "/api/v1/validate/luck-trials", body);
260
306
  return validationText(session, response);
261
307
  }
262
308
 
@@ -264,7 +310,7 @@ export async function toolValidateLuckTrials(session, args) {
264
310
  // through the API (one validation of quota, one receipt).
265
311
  async function runValidator(session, tool, path, body) {
266
312
  if (session.local) return computeLocally(tool, body);
267
- return callApi(session, { path, method: "POST", body });
313
+ return validateRemote(session, tool, path, body);
268
314
  }
269
315
 
270
316
  const sameJson = (a, b) => JSON.stringify(a) === JSON.stringify(b);
@@ -476,72 +522,74 @@ const READ_ONLY = { readOnlyHint: true, destructiveHint: false, openWorldHint: t
476
522
  const WRITES_RECEIPT = { readOnlyHint: false, destructiveHint: false, openWorldHint: true };
477
523
 
478
524
  export function registerTools(server, session) {
479
- server.registerTool(
525
+ const enabled = new Set((session.toolsets ?? Object.keys(TOOLSETS)).flatMap((name) => TOOLSETS[name]));
526
+ const register = (name, ...rest) => { if (enabled.has(name)) server.registerTool(name, ...rest); };
527
+ register(
480
528
  "get_key",
481
529
  { title: "Get a free validation key", annotations: { title: "Get a free validation key", readOnlyHint: false, destructiveHint: false, idempotentHint: false, openWorldHint: true }, description: TOOL_DESCRIPTIONS.get_key, inputSchema: getKeyInput },
482
530
  (args) => toolGetKey(session, args),
483
531
  );
484
- server.registerTool(
532
+ register(
485
533
  "validate_deflated_sharpe",
486
534
  { title: "Validate deflated Sharpe", annotations: { title: "Validate deflated Sharpe", ...WRITES_RECEIPT }, description: TOOL_DESCRIPTIONS.validate_deflated_sharpe, inputSchema: deflatedSharpeToolShape },
487
535
  (args) => toolValidateDeflatedSharpe(session, args),
488
536
  );
489
- server.registerTool(
537
+ register(
490
538
  "validate_overfitting",
491
539
  { title: "Validate overfitting (CSCV)", annotations: { title: "Validate overfitting (CSCV)", ...WRITES_RECEIPT }, description: TOOL_DESCRIPTIONS.validate_overfitting, inputSchema: overfittingInput },
492
540
  (args) => toolValidateOverfitting(session, args),
493
541
  );
494
- server.registerTool(
542
+ register(
495
543
  "validate_paper_evidence",
496
544
  { title: "Validate paper evidence", annotations: { title: "Validate paper evidence", ...WRITES_RECEIPT }, description: TOOL_DESCRIPTIONS.validate_paper_evidence, inputSchema: paperEvidenceInput },
497
545
  (args) => toolValidatePaperEvidence(session, args),
498
546
  );
499
- server.registerTool(
547
+ register(
500
548
  "validate_breadth",
501
549
  { title: "Validate breadth ceiling", annotations: { title: "Validate breadth ceiling", ...WRITES_RECEIPT }, description: TOOL_DESCRIPTIONS.validate_breadth, inputSchema: breadthInput },
502
550
  (args) => toolValidateBreadth(session, args),
503
551
  );
504
- server.registerTool(
552
+ register(
505
553
  "validate_track_record",
506
554
  { title: "Minimum track record length", annotations: { title: "Minimum track record length", ...WRITES_RECEIPT }, description: TOOL_DESCRIPTIONS.validate_track_record, inputSchema: trackRecordInput },
507
555
  (args) => toolValidateTrackRecord(session, args),
508
556
  );
509
- server.registerTool(
557
+ register(
510
558
  "validate_backtest_length",
511
559
  { title: "Minimum backtest length", annotations: { title: "Minimum backtest length", ...WRITES_RECEIPT }, description: TOOL_DESCRIPTIONS.validate_backtest_length, inputSchema: backtestLengthInput },
512
560
  (args) => toolValidateBacktestLength(session, args),
513
561
  );
514
- server.registerTool(
562
+ register(
515
563
  "validate_haircut_sharpe",
516
564
  { title: "Haircut Sharpe ratio", annotations: { title: "Haircut Sharpe ratio", ...WRITES_RECEIPT }, description: TOOL_DESCRIPTIONS.validate_haircut_sharpe, inputSchema: haircutSharpeInput },
517
565
  (args) => toolValidateHaircutSharpe(session, args),
518
566
  );
519
- server.registerTool(
567
+ register(
520
568
  "validate_luck_trials",
521
569
  { title: "Luck-equivalent trials", annotations: { title: "Luck-equivalent trials", ...WRITES_RECEIPT }, description: TOOL_DESCRIPTIONS.validate_luck_trials, inputSchema: luckTrialsInput },
522
570
  (args) => toolValidateLuckTrials(session, args),
523
571
  );
524
- server.registerTool(
572
+ register(
525
573
  "audit_backtest",
526
574
  { title: "Audit a backtest", annotations: { title: "Audit a backtest", ...WRITES_RECEIPT }, description: TOOL_DESCRIPTIONS.audit_backtest, inputSchema: auditBacktestToolShape },
527
575
  (args) => toolAuditBacktest(session, args),
528
576
  );
529
- server.registerTool(
577
+ register(
530
578
  "get_receipt",
531
579
  { title: "Get a receipt", annotations: { title: "Get a receipt", ...READ_ONLY }, description: TOOL_DESCRIPTIONS.get_receipt, inputSchema: getReceiptInput },
532
580
  (args) => toolGetReceipt(session, args),
533
581
  );
534
- server.registerTool(
582
+ register(
535
583
  "verify_receipt",
536
584
  { title: "Verify a receipt", annotations: { title: "Verify a receipt", ...READ_ONLY }, description: TOOL_DESCRIPTIONS.verify_receipt, inputSchema: verifyReceiptToolShape },
537
585
  (args) => toolVerifyReceipt(session, args),
538
586
  );
539
- server.registerTool(
587
+ register(
540
588
  "service_status",
541
589
  { title: "Service status", annotations: { title: "Service status", ...READ_ONLY }, description: TOOL_DESCRIPTIONS.service_status, inputSchema: emptyInput },
542
590
  () => toolServiceStatus(session),
543
591
  );
544
- server.registerTool(
592
+ register(
545
593
  "company_financial_history",
546
594
  { title: "Company financial history (SEC)", annotations: { title: "Company financial history (SEC)", ...READ_ONLY }, description: TOOL_DESCRIPTIONS.company_financial_history, inputSchema: companyHistoryToolShape },
547
595
  (args) => toolCompanyFinancialHistory(session, args),