canli-validation-mcp 0.10.1 → 0.11.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -71,6 +71,16 @@ repository for the full design.
71
71
 
72
72
  ## Tools
73
73
 
74
+ Unreleased source adds local journal binding to `validate_paper_evidence`: run
75
+ `CANLI_LOCAL=1 node mcp/src/server.mjs` from the repository, and send exactly one of `record`
76
+ or `record_file`, plus `journal_file` to check source hashes, signatures and recomputed claims.
77
+ `record_file` accepts a standard record, a `{record,signature}` envelope, or the full execution
78
+ export bundle; companion metrics/observations are checked for full bundles. Signed records
79
+ need the detached signature. `bindings.all_match` must be true for source validation to pass.
80
+ Without a journal, the result explicitly reports structure-only conformance. Files and
81
+ signatures are refused in hosted/remote mode and never uploaded. Local results have no Canli
82
+ receipt. The current npm/hosted release retains the existing record-only API.
83
+
74
84
  | Tool | Calls | Key required |
75
85
  |---|---|---|
76
86
  | `get_key` | `POST /api/v1/keys` | no |
@@ -127,6 +137,49 @@ original SEC response and the record's own boundary sentence: these are accounti
127
137
  reported to the SEC, not market prices, returns or a recommendation. Companies and concepts
128
138
  outside the current release return an error with the available concepts listed.
129
139
 
140
+ ## The lab: backtest, summarize, stress, check feasibility
141
+
142
+ Four tools compute in the server process itself, on your machine or on the hosted endpoint, with
143
+ nothing sent to the API and no receipt stored. Each result names the SHA-256 of exactly the numbers
144
+ it was computed from, and states what it does not establish.
145
+
146
+ - **`backtest_strategy`** runs a rule over every parameter set in a grid, on your prices:
147
+ `sma_cross` (fast, slow), `momentum` (lookback, skip), `mean_reversion` (window, entry_z),
148
+ `breakout` (lookback) or `buy_and_hold`, long-only or with `allow_short`, after a cost per unit
149
+ of turnover. The position held over each period is decided from closes up to the one before it,
150
+ so no rule can see the future (a test changes later prices and checks that no earlier position
151
+ moves). It then validates the best variant with the number of variants **this call ran**, not a
152
+ number someone declared: the deflated Sharpe ratio across the grid's measured Sharpe dispersion,
153
+ the CSCV probability of backtest overfitting on every variant's returns, the probabilistic Sharpe
154
+ and the minimum track record, next to buy and hold over the same window. At most 200 variants and
155
+ 20,000 prices per call. No code is ever run: a strategy is a family plus numbers.
156
+ - **`summarize_series`** reads a long price or return series and states it in about a hundred words,
157
+ plus a field for every figure: growth, volatility, Sharpe and drift t-statistic, the worst
158
+ drawdown with its dates and recovery, trend, the volatility regime (terciles of its own 21-period
159
+ volatility), tails, jumps (beyond six robust standard deviations), autocorrelation, stale data and,
160
+ with a benchmark, correlation and beta. A five-year daily series is about 1,260 numbers; the
161
+ summary is a paragraph.
162
+ - **`stress_test`** resamples a strategy's returns with a stationary block bootstrap (seeded, so a run
163
+ reproduces exactly) and reports how often the drawdown limit breaks and the Sharpe turns negative,
164
+ then applies four named scenarios with their rules stated: a crash at the equity high (default the
165
+ larger of 20% and three times the worst period), volatility multiplied, the worst stretch lived
166
+ twice, and positions stuck for several periods after the worst one. The fragility share is the part
167
+ of resampled histories that break the drawdown limit or lose money on a risk-adjusted basis.
168
+ Generative models are not used: trained on one series they memorize it or invent dynamics nobody
169
+ can check.
170
+ - **`check_feasibility`** checks a plan against a real broker before it trades: the broker's published
171
+ order-rate limit against the rebalance's burst of orders (Alpaca 200 a minute, Interactive Brokers
172
+ about 50 a second), the minimum order (Alpaca 1 USD notional), US day-trading rules as they stand
173
+ in 2026 (FINRA retired the pattern day trader rule and its 25,000 USD minimum on 4 June 2026; firms
174
+ may phase in its replacement until 20 October 2027), T+1 settlement in cash accounts, each order's
175
+ share of daily volume, square-root market impact (coefficient Y, default 1) with every copy of the
176
+ same strategy counted, and the capital at which costs and impact eat the expected gross return.
177
+ Every broker fact carries the date it was checked and its source.
178
+
179
+ On your own machine, `prices_file`, `series_file` and `returns_file` take a CSV or JSON path instead
180
+ of the numbers. Only numbers are read, plus an ISO date column when the file has one, so results can
181
+ name the day a drawdown began.
182
+
130
183
  ## Auditing a backtest in one call
131
184
 
132
185
  `audit_backtest` takes one strategy's return series, the number of variants tried and their Sharpe
@@ -147,11 +200,18 @@ took three audit questions from 4 of 9 to 8 of 9 answered correctly on gpt-5.4-m
147
200
 
148
201
  ## Prompts, resources and structured results
149
202
 
150
- Clients that show MCP prompts offer two guided workflows: `validate_backtest` (deflated Sharpe, then
151
- overfitting, then the track record needed, reported with what each number does not establish) and
152
- `track_record_needed`. Two resources can be read: `canli://limits`, the boundary sentences every
153
- result carries, and `canli://sources`, the papers behind each validator and how each is checked
154
- against them. Every tool result carries its envelope both as text and as `structuredContent`.
203
+ Clients that show MCP prompts offer six guided workflows: `validate_backtest` (deflated Sharpe, then
204
+ overfitting, then the track record needed, reported with what each number does not establish),
205
+ `track_record_needed`, `backtest_and_validate`, `stress_my_strategy`, `production_check` and
206
+ `summarize_market_series`. Resources: `canli://limits`, the boundary sentences every result carries;
207
+ `canli://sources`, the papers behind each validator and how each is checked against them;
208
+ `canli://strategy-spec`, the JSON Schema of every strategy family and its parameters; and
209
+ `canli://openapi`, the API's OpenAPI 3.1 document. Two resource templates are for writing code
210
+ against the tools: `canli://schemas/{tool}` returns a tool's exact input and output JSON Schema, and
211
+ `canli://examples/{language}/{tool}` a working call in `python`, `javascript` or `curl` (both
212
+ variables complete). A test compiles every generated Python and JavaScript example and checks that
213
+ every example's arguments are valid input. Every tool result carries its envelope both as text and
214
+ as `structuredContent`.
155
215
 
156
216
  ## Compact context (0.3.0)
157
217
 
@@ -183,7 +243,7 @@ an array of objects.
183
243
  | `CANLI_API_BASE` | `https://canlicapital.com` | Where the API lives. Point it at a preview deployment for testing. |
184
244
  | `CANLI_KEY` | unset | A key already issued from `POST /api/v1/keys`. When set, `get_key` sends no request and reports the key is already configured; every other tool sends it as `Authorization: Bearer <key>`. |
185
245
  | `CANLI_FULL_ENVELOPE` | unset | `1` or `true` returns each validation's full API envelope instead of the compact result (below). |
186
- | `CANLI_TOOLSETS` | all | Which tools to list: a comma-separated choice of `validate`, `receipts`, `company` and `status`, or `all`. An unknown name is refused at startup. See "Toolsets" below. |
246
+ | `CANLI_TOOLSETS` | all | Which tools to list: a comma-separated choice of `validate`, `receipts`, `company`, `status` and `lab`, or `all`. An unknown name is refused at startup. See "Toolsets" below. |
187
247
  | `CANLI_LOCAL` | unset | `1` or `true` runs the eight validators on this machine (private local mode, below): no key, no network, no receipt. |
188
248
 
189
249
  If `CANLI_KEY` is not set and local mode is off, call `get_key` once per session before the validators. The key it
@@ -338,7 +398,7 @@ On a breadth result this is about half the text. Set `CANLI_FULL_ENVELOPE=1` to
338
398
  ## Toolsets (tokens)
339
399
 
340
400
  A client sends the model the whole tool list on every turn, and it is most of each turn's prompt:
341
- a validation result is a few hundred tokens, the list of all fifteen tools several thousand. A
401
+ a validation result is a few hundred tokens, the list of all nineteen tools several thousand. A
342
402
  client that needs one kind of tool can list only that kind, with `CANLI_TOOLSETS` (stdio) or
343
403
  `?toolsets=` (hosted endpoint). The default is every tool.
344
404
 
@@ -348,17 +408,19 @@ client that needs one kind of tool can list only that kind, with `CANLI_TOOLSETS
348
408
  | `receipts` | `get_receipt`, `verify_receipt` |
349
409
  | `company` | `company_financial_history` |
350
410
  | `status` | `service_status` |
411
+ | `lab` | `backtest_strategy`, `summarize_series`, `stress_test`, `check_feasibility` |
351
412
 
352
413
  Measured with `bench/tool_list_tokens.py` (tokenizer: tiktoken `o200k_base`; other models'
353
414
  tokenizers give different absolute counts), in the shape an OpenAI-style client sends the list:
354
415
 
355
416
  | CANLI_TOOLSETS | tools | tokens per turn | of all |
356
417
  |---|---|---|---|
357
- | `all` | 15 | 4,194 | 100% |
358
- | `validate` | 11 | 3,526 | 84% |
359
- | `receipts` | 2 | 312 | 7% |
360
- | `company` | 1 | 273 | 7% |
361
- | `status` | 1 | 89 | 2% |
418
+ | `all` | 19 | 5,946 | 100% |
419
+ | `validate` | 11 | 3,685 | 62% |
420
+ | `receipts` | 2 | 312 | 5% |
421
+ | `company` | 1 | 273 | 5% |
422
+ | `status` | 1 | 89 | 1% |
423
+ | `lab` | 4 | 1,595 | 27% |
362
424
 
363
425
  Providers cache a tool list that is identical from turn to turn and bill the cached part at a
364
426
  fraction of the price (`test/tool-list-stable.test.mjs` keeps each list byte-stable); a smaller list
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "canli-validation-mcp",
3
- "version": "0.10.1",
3
+ "version": "0.11.0",
4
4
  "description": "Check whether a backtest is real: deflated Sharpe, CSCV probability of backtest overfitting, minimum track record and backtest length, haircut Sharpe and luck-equivalent trials, as an MCP server. Local mode, hosted endpoint, signed receipts.",
5
5
  "keywords": [
6
6
  "mcp",
@@ -15,7 +15,10 @@
15
15
  "overfitting",
16
16
  "quantitative-finance",
17
17
  "quant",
18
- "trading"
18
+ "trading",
19
+ "backtest-engine",
20
+ "stress-test",
21
+ "market-impact"
19
22
  ],
20
23
  "private": false,
21
24
  "type": "module",
@@ -0,0 +1,35 @@
1
+ import { resolve } from 'node:path';
2
+ import { computeLocally } from './local.mjs';
3
+ import { MAX_JOURNAL_BYTES, readBoundedFile, readRecordFile } from './local/js/journal-files.js';
4
+ import { journalBindings } from './local/js/trade-journal-export-core.js';
5
+
6
+ export function validateLocalJournalEvidence(input) {
7
+ let record = input.record, signature = input.signature, bundle;
8
+ if (input.record_file) {
9
+ const file = readRecordFile(resolve(input.record_file));
10
+ if (file && typeof file === 'object' && Object.hasOwn(file, 'record')) {
11
+ record = file.record;
12
+ // A {record,signature} envelope is also accepted. Full exports advertise their
13
+ // financial companion fields and those must be recomputed, not ignored.
14
+ if (Object.keys(file).some(k => !['record', 'signature'].includes(k))) bundle = file;
15
+ if (signature !== undefined && file.signature !== undefined) throw new Error('send a detached signature inline or in the record file, not both');
16
+ signature ??= file.signature;
17
+ } else record = file;
18
+ if (!record || typeof record !== 'object' || Array.isArray(record) || record.schema !== 'canli.paper-evidence.v0') throw new Error('record file must contain a canli.paper-evidence.v0 record or export bundle');
19
+ }
20
+ const result = computeLocally('validate_paper_evidence', { record });
21
+ if (result.failed) return result;
22
+ const data = result.envelope.data;
23
+ const unchecked = { checked: false, all_match: null, reason: 'No source journal checked; conformance covers structure and field relationships only.' };
24
+ // Conformance first prevents malformed record members from reaching the financial verifier.
25
+ const bindings = input.journal_file && data.valid
26
+ ? journalBindings(record, readBoundedFile(resolve(input.journal_file), { maxBytes: MAX_JOURNAL_BYTES }), signature, bundle)
27
+ : input.journal_file ? { checked: false, all_match: false, reason: 'Source verification requires a conforming record.' } : unchecked;
28
+ data.bindings = bindings;
29
+ data.conformance_valid = data.valid;
30
+ if (input.journal_file) {
31
+ data.valid = data.valid && bindings.all_match === true;
32
+ data.plain_reading = data.valid ? 'The record conforms and its reconstructable claims match the supplied signed journal and selected range. Key possession is self-attestation; broker authenticity, completeness and trusted time are not established.' : 'Conformance or source-bound recomputation failed; this record has not passed journal verification.';
33
+ } else data.plain_reading += ' No source journal or record signature was checked.';
34
+ return result;
35
+ }
@@ -0,0 +1,122 @@
1
+ // mcp/src/lab-schemas.mjs
2
+ //
3
+ // Input and output schemas for the lab tools (src/lab.mjs): backtest_strategy, summarize_series,
4
+ // stress_test and check_feasibility. Like the validators' descriptions, each tool description
5
+ // carries one sentence of its own result's boundary language verbatim (test/lab.test.mjs checks
6
+ // it against the sentences the computation attaches), so an agent reads what the result cannot be
7
+ // used to claim before it ever calls the tool. Descriptions stay terse: the tool list is re-sent to
8
+ // the model on every turn.
9
+ import { z } from "zod";
10
+ import { BACKTEST_LIMITS_TEXT, FAMILIES } from "./local/js/backtest-core.js";
11
+ import { FEASIBILITY_LIMITS_TEXT } from "./local/js/feasibility-core.js";
12
+ import { SUMMARY_LIMITS_TEXT } from "./local/js/series-summary-core.js";
13
+ import { STRESS_LIMITS_TEXT } from "./local/js/stress-core.js";
14
+
15
+ export const LAB_BOUNDARY = Object.freeze({
16
+ backtest_strategy: BACKTEST_LIMITS_TEXT[3],
17
+ summarize_series: SUMMARY_LIMITS_TEXT[2],
18
+ stress_test: STRESS_LIMITS_TEXT[2],
19
+ check_feasibility: FEASIBILITY_LIMITS_TEXT[0],
20
+ });
21
+
22
+ export const LAB_TOOL_DESCRIPTIONS = Object.freeze({
23
+ backtest_strategy: `Backtest a rule on your prices over a whole parameter grid, then validate the best with the count of variants actually run: deflated Sharpe, overfitting probability, minimum track record, vs buy and hold, after costs. ${LAB_BOUNDARY.backtest_strategy}`,
24
+ summarize_series: `A long price or return series in about 100 words plus fields: growth, risk, dated drawdowns, trend, volatility regime, tails, jumps, stale data. Use instead of reading raw bars. ${LAB_BOUNDARY.summarize_series}`,
25
+ stress_test: `Stress a strategy's returns: bootstrap histories (how often the drawdown limit breaks or Sharpe turns negative) and named crash, volatility, repeat and stuck-position scenarios, with a fragility share. ${LAB_BOUNDARY.stress_test}`,
26
+ check_feasibility: `Check a trading plan before it trades: broker order-rate and minimum-order limits, 2026 US day-trading rules, square-root market impact with crowding, and the capital where costs eat the return. ${LAB_BOUNDARY.check_feasibility}`,
27
+ });
28
+
29
+ const FAMILY_NAMES = Object.keys(FAMILIES);
30
+ // Lengths, counts and magnitudes are checked by the computation itself (src/local/js/*-core.js),
31
+ // with messages that say what to send instead; the advertised schema keeps types, enums and the
32
+ // bounds that tell a model what a value means (fractions, signs), because every bound here is
33
+ // re-sent to the model on every turn.
34
+ const series = () => z.array(z.number());
35
+ const file = (what) => z.string().describe(`CSV or JSON of ${what} on this machine; local server only.`);
36
+ const column = z.union([z.string(), z.number().int()]).describe("Column name or 1-based position.");
37
+ const periods = z.number().describe("252 daily (default), 365 crypto, 52 weekly, 12 monthly.");
38
+ const dates = z.array(z.string()).describe("Optional ISO dates, one per value.");
39
+
40
+ export const backtestInput = z
41
+ .object({
42
+ family: z.enum(FAMILY_NAMES).describe("sma_cross (fast, slow), momentum (lookback, skip), mean_reversion (window, entry_z), breakout (lookback), buy_and_hold."),
43
+ grid: z.record(z.string(), z.union([z.number(), z.array(z.number())])).optional()
44
+ .describe('Values per parameter, e.g. {"fast":[10,20],"slow":[50,200]}; every valid combination runs (max 200).'),
45
+ prices: series().optional().describe("Prices, oldest first."),
46
+ prices_file: file("prices").optional(),
47
+ prices_column: column.optional(),
48
+ dates: dates.optional(),
49
+ periods_per_year: periods.optional(),
50
+ cost_bps: z.number().min(0).optional().describe("Cost per unit of turnover, bps; default 5."),
51
+ allow_short: z.boolean().optional().describe("Short instead of flat; default false."),
52
+ })
53
+ .strict();
54
+
55
+ export const summarizeInput = z
56
+ .object({
57
+ prices: series().optional().describe("Prices, oldest first; or returns, or series_file."),
58
+ returns: z.array(z.number()).optional().describe("Returns as fractions, oldest first."),
59
+ series_file: file("prices or returns").optional(),
60
+ series_column: column.optional(),
61
+ series_kind: z.enum(["prices", "returns"]).optional().describe("series_file holds; default prices."),
62
+ dates: dates.optional(),
63
+ periods_per_year: periods.optional(),
64
+ benchmark_returns: z.array(z.number()).optional().describe("Same-period benchmark returns; adds beta."),
65
+ name: z.string().optional().describe("Label for the text."),
66
+ })
67
+ .strict();
68
+
69
+ export const stressInput = z
70
+ .object({
71
+ returns: z.array(z.number()).optional().describe("Strategy returns as fractions, oldest first."),
72
+ returns_file: file("returns").optional(),
73
+ returns_column: column.optional(),
74
+ periods_per_year: periods.optional(),
75
+ paths: z.number().int().optional().describe("Default 1000."),
76
+ block: z.number().optional().describe("Mean block length; default cube root of n."),
77
+ seed: z.number().int().optional().describe("Default 42; reproduces exactly."),
78
+ drawdown_limit: z.number().gt(0).lt(1).optional().describe("Fraction; default 0.2."),
79
+ crash: z.number().gt(-1).lt(0).optional().describe("Crash size, negative; default min(-0.2, 3x worst period)."),
80
+ volatility_multiplier: z.number().optional().describe("Default 2."),
81
+ outage_periods: z.number().int().optional().describe("Stuck periods after the worst; default 5."),
82
+ })
83
+ .strict();
84
+
85
+ export const feasibilityInput = z
86
+ .object({
87
+ broker: z.enum(["alpaca", "ibkr", "other"]).optional().describe("Default other."),
88
+ asset_class: z.enum(["us_equity", "crypto", "futures", "fx", "options"]).optional(),
89
+ account: z.enum(["cash", "margin"]).optional(),
90
+ capital_usd: z.number(),
91
+ orders_per_rebalance: z.number().int(),
92
+ rebalances_per_year: z.number(),
93
+ turnover_per_year: z.number().min(0).describe("Buys plus sells a year / capital."),
94
+ adv_usd: z.number().describe("Daily dollar volume of a typical holding."),
95
+ daily_volatility: z.number().max(1).describe("Of a typical holding, e.g. 0.02."),
96
+ copies: z.number().int().optional().describe("Agents trading the same signal at once; default 1."),
97
+ expected_gross_return: z.number().min(-1).optional().describe("Annual, before costs; adds capacity."),
98
+ spread_and_fees_bps: z.number().min(0).optional().describe("Per trade; default 0."),
99
+ impact_coefficient: z.number().optional().describe("Square-root law Y; default 1."),
100
+ minutes_per_rebalance: z.number().optional().describe("Default 1."),
101
+ orders_per_minute_limit: z.number().optional().describe("If the broker's is not on file."),
102
+ })
103
+ .strict();
104
+
105
+ const loose = z.looseObject({}).optional();
106
+ const sentences = z.array(z.string()).optional();
107
+
108
+ export const backtestOutput = z
109
+ .looseObject({ variants: loose, selected: loose, buy_and_hold: loose, validation: loose, plain_reading: z.string().optional(), limits: sentences, source: loose })
110
+ .describe("selected holds the best variant's parameters and metrics; validation the deflated Sharpe, overfitting probability and minimum track record counting every variant run; plain_reading states them.");
111
+
112
+ export const summaryOutput = z
113
+ .looseObject({ text: z.string().optional(), growth: loose, risk: loose, drawdown: loose, trend: loose, volatility: z.looseObject({}).nullable().optional(), tails: loose, jumps: loose, limits: sentences, source: loose })
114
+ .describe("text is the summary to read; the other fields hold each figure by the rule named in it.");
115
+
116
+ export const stressOutput = z
117
+ .looseObject({ baseline: loose, resampled: loose, scenarios: z.array(z.unknown()).optional(), fragility: loose, plain_reading: z.string().optional(), limits: sentences, source: loose })
118
+ .describe("resampled holds bootstrap percentiles and breach frequencies; scenarios each named stress and its rule; fragility the share of histories breaking the limit or losing.");
119
+
120
+ export const feasibilityOutput = z
121
+ .looseObject({ verdict: z.string().optional(), checks: z.array(z.unknown()).optional(), to_change: z.array(z.string()).optional(), sources: loose, plain_reading: z.string().optional(), limits: sentences })
122
+ .describe("verdict summarizes the checks; each check has a status (ok, warning, blocking, unknown), its finding and numbers; to_change lists what to fix.");
package/src/lab.mjs ADDED
@@ -0,0 +1,309 @@
1
+ // mcp/src/lab.mjs
2
+ //
3
+ // The lab: four tools that compute in this process, on stdio and on the hosted endpoint alike,
4
+ // from src/local (mirrored byte for byte from the canlicapital repository by
5
+ // scripts/sync-local.mjs). Nothing is sent to an API and no receipt is stored; files are read only
6
+ // by the server on the user's own machine, and only their numbers (and ISO dates) are used.
7
+ //
8
+ // backtest_strategy a rule-based strategy over every parameter set in a grid, validated with
9
+ // the number of variants actually run
10
+ // summarize_series a long series in about a hundred words an agent can reason over
11
+ // stress_test resampled histories and named scenarios, with a fragility share
12
+ // check_feasibility broker limits, day-trading rules, market impact, crowding and capacity
13
+ //
14
+ // Also here: the code-generation resources (every tool's exact JSON Schemas, the strategy spec,
15
+ // runnable client examples, the API's OpenAPI document) and the lab's guided prompts.
16
+ import { createHash } from "node:crypto";
17
+ import { ResourceTemplate } from "@modelcontextprotocol/server";
18
+ import { z } from "zod";
19
+ import { BACKTEST_LIMITS, BACKTEST_LIMITS_TEXT, FAMILIES, runBacktest } from "./local/js/backtest-core.js";
20
+ import { FEASIBILITY_LIMITS_TEXT, checkFeasibility } from "./local/js/feasibility-core.js";
21
+ import { SUMMARY_LIMITS_TEXT, summarizeSeries } from "./local/js/series-summary-core.js";
22
+ import { STRESS_LIMITS_TEXT, stressTest } from "./local/js/stress-core.js";
23
+ import { LAB_TOOL_DESCRIPTIONS, backtestInput, backtestOutput, feasibilityInput, feasibilityOutput, stressInput, stressOutput, summarizeInput, summaryOutput } from "./lab-schemas.mjs";
24
+ import { readSeriesWithDates } from "./series-file.mjs";
25
+
26
+ export const LAB_TOOLS = Object.freeze(["backtest_strategy", "summarize_series", "stress_test", "check_feasibility"]);
27
+
28
+ // Same shape as every other tool's result: the object as text, and as structured content.
29
+ const labText = (value) => ({ content: [{ type: "text", text: JSON.stringify(value) }], structuredContent: value });
30
+
31
+ function parse(schema, args, tool) {
32
+ const result = schema.safeParse(args ?? {});
33
+ if (result.success) return result.data;
34
+ throw new Error(`${tool}: ${result.error.issues.map((i) => `${i.path.join(".") || "(root)"}: ${i.message}`).join("; ")}`);
35
+ }
36
+
37
+ // The series a tool works on: inline, or from a file on this machine. Exactly one of the two.
38
+ function readInput(session, tool, { inline, file, column, inlineName, fileName }) {
39
+ if ((inline === undefined) === (file === undefined)) throw new Error(`${tool}: send exactly one of ${inlineName} or ${fileName}`);
40
+ if (file === undefined) return { values: inline, dates: undefined, source: {} };
41
+ if (session.hosted) throw new Error(`${tool}: the hosted endpoint cannot read files on your machine; send ${inlineName} as numbers, or run the server locally with npx -y canli-validation-mcp.`);
42
+ const read = readSeriesWithDates(file, column);
43
+ return {
44
+ values: read.values,
45
+ dates: read.dates ?? undefined,
46
+ source: { [fileName]: file, column_position: read.column, ...(read.date_column ? { date_column_position: read.date_column } : {}), ...(read.skipped.length ? { skipped_row_counter_columns: read.skipped } : {}) },
47
+ };
48
+ }
49
+
50
+ // The digest of exactly the numbers a result was computed from, so a result can be tied to its input.
51
+ const digest = (values) => createHash("sha256").update(JSON.stringify(values)).digest("hex");
52
+
53
+ // Labels are echoed into results, so each is cut to a date-and-time length; a label is never a
54
+ // channel for long text.
55
+ const labels = (values) => values?.map((value) => String(value).slice(0, 40));
56
+
57
+ export async function toolBacktestStrategy(session, args) {
58
+ const input = parse(backtestInput, args, "backtest_strategy");
59
+ const { values, dates, source } = readInput(session, "backtest_strategy", { inline: input.prices, file: input.prices_file, column: input.prices_column, inlineName: "prices", fileName: "prices_file" });
60
+ const result = runBacktest({ prices: values, family: input.family, grid: input.grid ?? {}, periods_per_year: input.periods_per_year ?? 252, cost_bps: input.cost_bps ?? 5, allow_short: input.allow_short ?? false, dates: labels(input.dates) ?? dates });
61
+ return labText({ ...result, limits: BACKTEST_LIMITS_TEXT, source: { ...source, prices: values.length, input_sha256: digest(values), computed: "locally, in the MCP server process; no receipt" } });
62
+ }
63
+
64
+ export async function toolSummarizeSeries(session, args) {
65
+ const input = parse(summarizeInput, args, "summarize_series");
66
+ const inline = input.prices ?? input.returns;
67
+ if (input.prices !== undefined && input.returns !== undefined) throw new Error("summarize_series: send prices or returns, not both");
68
+ const { values, dates, source } = readInput(session, "summarize_series", { inline, file: input.series_file, column: input.series_column, inlineName: "prices (or returns)", fileName: "series_file" });
69
+ const kind = input.series_file !== undefined ? (input.series_kind ?? "prices") : input.prices !== undefined ? "prices" : "returns";
70
+ const result = summarizeSeries({ [kind]: values, dates: labels(input.dates) ?? dates, periods_per_year: input.periods_per_year ?? 252, benchmark_returns: input.benchmark_returns, name: input.name?.slice(0, 80) });
71
+ return labText({ ...result, limits: SUMMARY_LIMITS_TEXT, source: { ...source, kind, values: values.length, input_sha256: digest(values) } });
72
+ }
73
+
74
+ export async function toolStressTest(session, args) {
75
+ const input = parse(stressInput, args, "stress_test");
76
+ const { values, source } = readInput(session, "stress_test", { inline: input.returns, file: input.returns_file, column: input.returns_column, inlineName: "returns", fileName: "returns_file" });
77
+ const result = stressTest({ ...input, returns: values });
78
+ return labText({ ...result, limits: STRESS_LIMITS_TEXT, source: { ...source, returns: values.length, input_sha256: digest(values) } });
79
+ }
80
+
81
+ export async function toolCheckFeasibility(session, args) {
82
+ const input = parse(feasibilityInput, args, "check_feasibility");
83
+ return labText({ ...checkFeasibility(input), limits: FEASIBILITY_LIMITS_TEXT });
84
+ }
85
+
86
+ // The lab computes locally and changes nothing anywhere: read-only, closed world, idempotent.
87
+ const LAB_ANNOTATIONS = { readOnlyHint: true, destructiveHint: false, idempotentHint: true, openWorldHint: false };
88
+
89
+ export function labToolSpecs(session) {
90
+ const spec = (name, title, input, output, handler) => [name, { title, annotations: { title, ...LAB_ANNOTATIONS }, description: LAB_TOOL_DESCRIPTIONS[name], inputSchema: input, outputSchema: output }, (args) => handler(session, args)];
91
+ return [
92
+ spec("backtest_strategy", "Backtest a rule over a parameter grid, validated", backtestInput, backtestOutput, toolBacktestStrategy),
93
+ spec("summarize_series", "Summarize a price or return series", summarizeInput, summaryOutput, toolSummarizeSeries),
94
+ spec("stress_test", "Stress-test a strategy's returns", stressInput, stressOutput, toolStressTest),
95
+ spec("check_feasibility", "Check a plan against broker limits and impact", feasibilityInput, feasibilityOutput, toolCheckFeasibility),
96
+ ];
97
+ }
98
+
99
+ // ---------------------------------------------------------------------------------------------
100
+ // Code-generation resources. An agent writing code against these tools reads the exact schemas
101
+ // and a working call instead of guessing argument names.
102
+ // ---------------------------------------------------------------------------------------------
103
+
104
+ // One small, valid call per tool: the arguments every generated example sends.
105
+ export const EXAMPLE_ARGS = Object.freeze({
106
+ get_key: {},
107
+ validate_deflated_sharpe: { observed_sharpe_annualized: 1.5, observations: 730, periods_per_year: 365, skew: -0.5, non_excess_kurtosis: 5, effective_independent_trials: 229, cross_trial_sharpe_sd_annualized: 0.57 },
108
+ validate_overfitting: { matrix_note: "every variant's returns, one row per period, one column per variant", matrix: [[0.01, 0.012], [-0.004, -0.002], [0.006, 0.001], [0.002, 0.004], [-0.01, -0.008], [0.007, 0.009], [0.003, -0.001], [0.0, 0.002], [0.005, 0.004], [-0.003, -0.005], [0.008, 0.006], [0.001, 0.003], [-0.002, 0.0], [0.004, 0.002], [0.006, 0.007], [-0.006, -0.004]] },
109
+ validate_reality_check: { matrix_file: "variants.csv" },
110
+ validate_paper_evidence: { record_file: "paper_record.json" },
111
+ validate_breadth: { sleeve_sharpe: 0.8, average_pairwise_correlation: 0.2, sleeves: 10 },
112
+ validate_track_record: { observed_sharpe_annualized: 1.2, periods_per_year: 252, skew: -0.3, non_excess_kurtosis: 4 },
113
+ validate_backtest_length: { target_sharpe_annualized: 1, effective_independent_trials: 50 },
114
+ validate_haircut_sharpe: { observed_sharpe_annualized: 1.5, periods_per_year: 252, observations: 1260, tests: 50 },
115
+ validate_luck_trials: { observed_sharpe_annualized: 1.5, periods_per_year: 252, observations: 1260 },
116
+ audit_backtest: { returns_file: "backtest_returns.csv", periods_per_year: 252, effective_independent_trials: 20, cross_trial_sharpe_sd_annualized: 0.5 },
117
+ get_receipt: { id: "0123456789abcdef01234567" },
118
+ verify_receipt: { id: "0123456789abcdef01234567" },
119
+ service_status: {},
120
+ company_financial_history: { ticker: "AAPL", concept: "Assets", limit: 8 },
121
+ backtest_strategy: { family: "sma_cross", grid: { fast: [10, 20, 50], slow: [100, 200] }, prices_file: "prices.csv", cost_bps: 5 },
122
+ summarize_series: { series_file: "prices.csv", series_kind: "prices", name: "My asset" },
123
+ stress_test: { returns_file: "strategy_returns.csv", drawdown_limit: 0.2, paths: 1000, seed: 42 },
124
+ check_feasibility: { broker: "alpaca", asset_class: "us_equity", account: "margin", capital_usd: 250000, orders_per_rebalance: 40, rebalances_per_year: 52, turnover_per_year: 8, adv_usd: 20000000, daily_volatility: 0.02, copies: 1, expected_gross_return: 0.12, spread_and_fees_bps: 3 },
125
+ });
126
+
127
+ const usesFile = (args) => Object.keys(args).some((k) => k.endsWith("_file"));
128
+ const cleanArgs = (args) => Object.fromEntries(Object.entries(args).filter(([k]) => !k.endsWith("_note")));
129
+
130
+ export function exampleCode(language, tool) {
131
+ const args = cleanArgs(EXAMPLE_ARGS[tool] ?? {});
132
+ const json = JSON.stringify(args, null, 2);
133
+ const fileNote = usesFile(args) ? "Point the *_file argument at your own CSV; the local server reads it, the hosted endpoint cannot." : "";
134
+ if (language === "python") {
135
+ return [
136
+ "# pip install mcp",
137
+ "# Runs the server on this machine (Node.js 20.10+). CANLI_LOCAL=1 computes everything locally, no key.",
138
+ fileNote ? `# ${fileNote}` : null,
139
+ "import asyncio",
140
+ "from mcp import ClientSession, StdioServerParameters",
141
+ "from mcp.client.stdio import stdio_client",
142
+ "",
143
+ 'SERVER = StdioServerParameters(command="npx", args=["-y", "canli-validation-mcp"], env={"CANLI_LOCAL": "1"})',
144
+ "",
145
+ "async def main():",
146
+ " async with stdio_client(SERVER) as (read, write):",
147
+ " async with ClientSession(read, write) as session:",
148
+ " await session.initialize()",
149
+ ` result = await session.call_tool("${tool}", ${json.replace(/\n/g, "\n ").replace(/\btrue\b/g, "True").replace(/\bfalse\b/g, "False").replace(/\bnull\b/g, "None")})`,
150
+ " print(result.structuredContent)",
151
+ "",
152
+ "asyncio.run(main())",
153
+ "",
154
+ ].filter((line) => line !== null).join("\n");
155
+ }
156
+ if (language === "javascript") {
157
+ return [
158
+ "// npm install @modelcontextprotocol/sdk",
159
+ "// Runs the server on this machine (Node.js 20.10+). CANLI_LOCAL=1 computes everything locally, no key.",
160
+ fileNote ? `// ${fileNote}` : null,
161
+ 'import { Client } from "@modelcontextprotocol/sdk/client/index.js";',
162
+ 'import { StdioClientTransport } from "@modelcontextprotocol/sdk/client/stdio.js";',
163
+ "",
164
+ 'const client = new Client({ name: "example", version: "1.0.0" });',
165
+ 'await client.connect(new StdioClientTransport({ command: "npx", args: ["-y", "canli-validation-mcp"], env: { ...process.env, CANLI_LOCAL: "1" } }));',
166
+ `const result = await client.callTool({ name: "${tool}", arguments: ${json} });`,
167
+ "console.log(result.structuredContent);",
168
+ "await client.close();",
169
+ "",
170
+ ].filter((line) => line !== null).join("\n");
171
+ }
172
+ if (language === "curl") {
173
+ if (usesFile(args)) {
174
+ return `# ${tool} reads a file on your machine, which the hosted endpoint cannot do.\n# Use the python or javascript example, or send the numbers inline (see canli://schemas/${tool}).\n`;
175
+ }
176
+ const body = JSON.stringify({ jsonrpc: "2.0", id: 1, method: "tools/call", params: { name: tool, arguments: args } });
177
+ return [
178
+ "# The hosted endpoint: no install, no key needed for a first call. Tools added since the last",
179
+ "# release appear here once that release is pinned; locally they are available at once.",
180
+ "curl -sS https://canlicapital.com/mcp \\",
181
+ " -H 'content-type: application/json' \\",
182
+ " -H 'accept: application/json, text/event-stream' \\",
183
+ ` -d '${body.replace(/'/g, "'\\''")}'`,
184
+ "",
185
+ ].join("\n");
186
+ }
187
+ throw new Error(`No example language ${language}; choose python, javascript or curl`);
188
+ }
189
+
190
+ export const EXAMPLE_LANGUAGES = Object.freeze(["python", "javascript", "curl"]);
191
+
192
+ // Every strategy family as a machine-readable spec: its parameters, their bounds and its rule.
193
+ export function strategySpec() {
194
+ const param = { type: "array", items: { type: "number" }, minItems: 1, maxItems: 60 };
195
+ return {
196
+ $schema: "https://json-schema.org/draft/2020-12/schema",
197
+ title: "backtest_strategy family and grid",
198
+ description: "Pick a family and give each of its parameters a list of values; every valid combination is run (at most 200), positions are decided at each close from prices up to that close, and trade the next period's return.",
199
+ oneOf: Object.entries(FAMILIES).map(([family, spec]) => ({
200
+ title: family,
201
+ description: spec.describe,
202
+ type: "object",
203
+ required: ["family", ...(spec.params.length ? ["grid"] : [])],
204
+ properties: {
205
+ family: { const: family },
206
+ ...(spec.params.length ? { grid: { type: "object", additionalProperties: false, required: spec.params.filter((p) => p !== "skip"), properties: Object.fromEntries(spec.params.map((p) => [p, param])) } } : {}),
207
+ },
208
+ })),
209
+ limits: { max_prices: BACKTEST_LIMITS.max_prices, max_variants: BACKTEST_LIMITS.max_variants, max_parameter_value: BACKTEST_LIMITS.max_parameter_value },
210
+ };
211
+ }
212
+
213
+ // catalog: { toolName: { description, inputSchema, outputSchema } } for every tool the server can list.
214
+ export function registerCodeResources(server, session, catalog) {
215
+ const names = Object.keys(catalog).sort();
216
+ const complete = (list) => (value) => list.filter((n) => n.startsWith(value ?? ""));
217
+ server.registerResource(
218
+ "tool-schema",
219
+ new ResourceTemplate("canli://schemas/{tool}", { list: undefined, complete: { tool: complete(names) } }),
220
+ { title: "A tool's exact JSON Schemas", description: "Input and output JSON Schema and the description of one tool, for writing code against it.", mimeType: "application/json" },
221
+ (uri, { tool }) => {
222
+ const entry = catalog[tool];
223
+ if (!entry) throw new Error(`No tool ${tool}; tools are ${names.join(", ")}`);
224
+ const body = { tool, description: entry.description, input_schema: z.toJSONSchema(entry.inputSchema, { io: "input" }), output_schema: z.toJSONSchema(entry.outputSchema, { io: "output" }) };
225
+ return { contents: [{ uri: uri.href, mimeType: "application/json", text: JSON.stringify(body, null, 2) }] };
226
+ },
227
+ );
228
+ server.registerResource(
229
+ "example",
230
+ new ResourceTemplate("canli://examples/{language}/{tool}", { list: undefined, complete: { language: complete([...EXAMPLE_LANGUAGES]), tool: complete(names) } }),
231
+ { title: "A working call, in Python, JavaScript or curl", description: "Runnable client code calling one tool with valid example arguments.", mimeType: "text/plain" },
232
+ (uri, { language, tool }) => {
233
+ if (!catalog[tool]) throw new Error(`No tool ${tool}; tools are ${names.join(", ")}`);
234
+ return { contents: [{ uri: uri.href, mimeType: "text/plain", text: exampleCode(language, tool) }] };
235
+ },
236
+ );
237
+ server.registerResource(
238
+ "strategy-spec",
239
+ "canli://strategy-spec",
240
+ { title: "Strategy families and their parameters", description: "JSON Schema of backtest_strategy's families, parameters and rules.", mimeType: "application/json" },
241
+ (uri) => ({ contents: [{ uri: uri.href, mimeType: "application/json", text: JSON.stringify(strategySpec(), null, 2) }] }),
242
+ );
243
+ server.registerResource(
244
+ "openapi",
245
+ "canli://openapi",
246
+ { title: "The validation API's OpenAPI 3.1 document", description: "Generated from the endpoints that exist; fetched from canlicapital.com when read.", mimeType: "application/json" },
247
+ async (uri) => {
248
+ const signal = AbortSignal.timeout(session.timeoutMs);
249
+ let res;
250
+ try {
251
+ res = await session.fetchImpl(`${session.base}/api/v1/openapi`, { headers: { Accept: "application/json" }, signal, redirect: "error" });
252
+ } catch {
253
+ throw new Error(`canli://openapi: could not reach ${session.base}/api/v1/openapi; it is also published at https://canlicapital.com/api/v1/openapi`);
254
+ }
255
+ if (!res.ok) throw new Error(`canli://openapi: ${session.base}/api/v1/openapi returned HTTP ${res.status}`);
256
+ return { contents: [{ uri: uri.href, mimeType: "application/json", text: await res.text() }] };
257
+ },
258
+ );
259
+ }
260
+
261
+ // Guided workflows for the lab. The prompt only writes the instructions; the tools do the work.
262
+ export function registerLabPrompts(server) {
263
+ const prompt = (name, title, description, argsSchema, lines) => server.registerPrompt(name, { title, description, argsSchema }, (args) => ({
264
+ messages: [{ role: "user", content: { type: "text", text: lines(args).filter(Boolean).join("\n") } }],
265
+ }));
266
+ prompt(
267
+ "backtest_and_validate",
268
+ "Backtest a rule and check whether the best variant is luck",
269
+ "Run a strategy family over a parameter grid on your prices, then read the deflated Sharpe and overfitting probability that count every variant run.",
270
+ { prices_file: z.string().describe("Path to a CSV of prices on this machine"), family: z.string().optional().describe("sma_cross, momentum, mean_reversion, breakout or buy_and_hold") },
271
+ ({ prices_file, family }) => [
272
+ `Backtest ${family ?? "a sensible rule for this asset"} on the prices in ${prices_file} with backtest_strategy, over a small grid of parameters (read canli://strategy-spec for each family's parameters).`,
273
+ "Report the best variant's Sharpe, drawdown and turnover next to buy and hold, then the deflated Sharpe and the probability of backtest overfitting, which count every variant the call ran.",
274
+ "Say plainly whether the best variant is distinguishable from luck, and quote the limits sentences the result carries.",
275
+ ],
276
+ );
277
+ prompt(
278
+ "stress_my_strategy",
279
+ "How fragile is this strategy?",
280
+ "Resampled histories and named crash, volatility, repeat and outage scenarios for a strategy's returns.",
281
+ { returns_file: z.string().describe("Path to a CSV of the strategy's returns"), drawdown_limit: z.string().optional().describe("The drawdown you could not live with, such as 0.2") },
282
+ ({ returns_file, drawdown_limit }) => [
283
+ `Run stress_test on ${returns_file}${drawdown_limit ? ` with drawdown_limit ${drawdown_limit}` : ""}.`,
284
+ "Explain the fragility share, the 5th-percentile drawdown and which named scenarios break the limit, and what change (lower leverage, a stop, more diversification) would address each.",
285
+ "Say that these are frequencies under stated rules, not forecasts.",
286
+ ],
287
+ );
288
+ prompt(
289
+ "production_check",
290
+ "Will this plan survive a real broker?",
291
+ "Order-rate and minimum-order limits, 2026 day-trading rules, market impact, crowding and capacity for a trading plan.",
292
+ { plan: z.string().describe("Capital, broker, how many orders per rebalance, how often, turnover, and what it trades") },
293
+ ({ plan }) => [
294
+ `Here is the plan: ${plan}.`,
295
+ "Ask me for any of check_feasibility's required inputs you cannot infer (average daily dollar volume and daily volatility of a typical holding), then run it.",
296
+ "List each blocking or warning finding with the change that fixes it, and say when the broker facts were checked.",
297
+ ],
298
+ );
299
+ prompt(
300
+ "summarize_market_series",
301
+ "Read a long series in a hundred words",
302
+ "Summarize a price or return file instead of reading its rows.",
303
+ { series_file: z.string().describe("Path to a CSV of prices or returns"), kind: z.string().optional().describe("prices or returns; default prices") },
304
+ ({ series_file, kind }) => [
305
+ `Call summarize_series on ${series_file} (series_kind ${kind ?? "prices"}) and reason from its text and fields instead of reading the file's rows.`,
306
+ "Point out anything in the data check (stale prices, jumps) that should be fixed before the series is backtested.",
307
+ ],
308
+ );
309
+ }