@booyaka/mcp-vet 0.7.0 → 0.9.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/BENCHMARK.md CHANGED
@@ -1,112 +1,161 @@
1
- # Benchmark: corpus, methodology, and honest limits
2
-
3
- The README claims high precision on real MCP code. This file is the evidence
4
- behind that claim — corpus, pinned commits, counts, labels, and what the
5
- scanner is *known to miss* — so the claim is checkable rather than vibes.
6
-
7
- > Prompted by community feedback on the launch post: *"'0 false positives' is
8
- > encouraging but incomplete without corpus size, commit SHAs, labeled
9
- > negatives, and recall."* Correct. Here they are.
10
-
11
- ## Corpus (pinned)
12
-
13
- Scanned with `mcp-vet` v0.4.0 (`node dist/cli.js <roots> --json`), all rules
14
- enabled, default confidence (`low`), on 2026-07-23:
15
-
16
- | Repo | Commit | Scanned root | Files | LOC |
17
- | --- | --- | --- | --- | --- |
18
- | [modelcontextprotocol/servers](https://github.com/modelcontextprotocol/servers) | `d31124c982401739917fd817c2a59db344529c16` | `src/` | 78 | 14,742 |
19
- | [modelcontextprotocol/typescript-sdk](https://github.com/modelcontextprotocol/typescript-sdk) | `1e1392e3f91583884fe82a0b4b91335875c3fba6` | `examples/` | 144 | 17,224 |
20
- | [modelcontextprotocol/python-sdk](https://github.com/modelcontextprotocol/python-sdk) | `3a6f2996cdd8358957479791e8b26198c07d6a75` | `examples/` | 225 | 12,013 |
21
- | **Total** | | | **447** | **43,979** |
22
-
23
- File counts are candidate files (`.ts/.tsx/.js/.mjs/.cjs/.py`) under the
24
- scanned roots, excluding `node_modules`.
25
-
26
- ## Results
27
-
28
- **105 findings across 41 files** (TypeScript/JavaScript: 93, Python: 12).
29
- By confidence: 66 high, 38 medium, 1 low.
30
-
31
- | Pattern | Findings |
32
- | --- | --- |
33
- | `MCP_SESSION_ID` | 49 |
34
- | `LOGGING_CAP` | 17 |
35
- | `SAMPLING_CAP` | 16 |
36
- | `ROOTS_CAP` | 15 |
37
- | `INITIALIZE_HANDLER` | 4 |
38
- | `TASKS_LEGACY` | 2 |
39
- | `TASKS_RESULT_REMOVED` | 2 |
40
-
41
- ### Labeling
42
-
43
- Every finding was manually reviewed against its source line:
44
-
45
- - **104 / 105 true positives** — real references to a removed or deprecated
46
- protocol surface (session headers/ids, handshake registration, legacy task
47
- methods, deprecated capability declarations and method strings).
48
- - **1 / 105 false positive (0.95%)** —
49
- `stories/json_response/client.py:62` in the typescript-sdk examples:
50
- `assert "mcp-session-id" not in response.headers`. That line is
51
- *already-migrated* test code asserting the header is **absent**; flagging it
52
- as "will break" is wrong. It is exactly what inline suppression
53
- (`# mcp-vet-disable-line MCP_SESSION_ID`) is for, but we count it as a false
54
- positive rather than defining it away. So the honest headline is
55
- **"1 false positive in 44k LOC"**, not zero.
56
-
57
- Notes on reading the numbers:
58
-
59
- - Two occurrences on one line (e.g. `transport.sessionId && sessions.delete(transport.sessionId)`)
60
- are reported as two findings — column-level dedup, not line-level.
61
- - Findings in test files (`__tests__/…`) are counted as true positives: a test
62
- that registers `sampling/createMessage` breaks the same way production code
63
- does.
64
-
65
- ### Labeled negatives
66
-
67
- Files asserted to stay **clean** are part of the repo's test suite and run in CI:
68
-
69
- - `test/fixtures/clean/` — a full server written in the 2026-07-28 style
70
- (per-request `_meta`, `sessionIdGenerator: undefined`, `-32602`).
71
- - `test/fixtures/negatives/` — "false friend" patterns: `sessionId` on plain
72
- app-level objects, `-32002` inside strings/comments, capability-like words
73
- with no capabilities context.
74
- - `test/fixtures/adversarial/caught/` — obfuscations the scanner **must**
75
- catch: aliased imports (TS + Python), namespace-qualified SDK constants,
76
- client transports resuming a `sessionId`.
77
-
78
- Additionally, in the corpus above, comment-only mentions (e.g. `Mcp-Session-Id`
79
- in a comment, `initialize` in prose) produced zero findings — the AST layer
80
- distinguishes executable tokens from comments by construction.
81
-
82
- ## Recall — what the scanner is known to miss
83
-
84
- Static token analysis proves known patterns are **absent**; it cannot prove
85
- your server **speaks the new wire contract**. Recall is bounded by
86
- construction, and the misses are locked into the test suite
87
- (`test/fixtures/adversarial/missed/`, asserted to produce zero findings so any
88
- silent claim-inflation fails CI):
89
-
90
- - split/computed method strings — `'tasks' + '/list'`, `` `tasks/${op}` ``, f-strings
91
- - computed capability keys — `{ ['roo'+'ts']: {} }`
92
- - generated/loop-driven registration from string fragments
93
- - framework-adapter indirection (route tables built at runtime)
94
- - cross-module renames — a wrapper re-exporting an SDK constant under a new
95
- name is flagged in the wrapper file, but a consumer importing only the new
96
- name scans clean on its own
97
-
98
- There is no corpus-wide recall *percentage*: that would require a labeled set
99
- of every legacy usage in the wild, which nobody has. What we can say is: for
100
- the pattern shapes listed in the README, detection is exact; for the shapes
101
- above, it is zero, and the tool says so — pair the scan with runtime checks
102
- (`mcp-vet fixtures`) to cover the difference.
103
-
104
- ## Reproducing
105
-
106
- ```bash
107
- git clone --depth 1 https://github.com/modelcontextprotocol/servers
108
- git clone --depth 1 https://github.com/modelcontextprotocol/typescript-sdk
109
- git clone --depth 1 https://github.com/modelcontextprotocol/python-sdk
110
- # check out the pinned SHAs above, then:
111
- npx @booyaka/mcp-vet servers/src typescript-sdk/examples python-sdk/examples --json --no-files
112
- ```
1
+ # Benchmark: corpus, methodology, and honest limits
2
+
3
+ The README claims high precision on real MCP code. This file is the evidence
4
+ behind that claim — corpus, pinned commits, counts, labels, and what the
5
+ scanner is *known to miss* — so the claim is checkable rather than vibes.
6
+
7
+ > Prompted by community feedback on the launch post: *"'0 false positives' is
8
+ > encouraging but incomplete without corpus size, commit SHAs, labeled
9
+ > negatives, and recall."* Correct. Here they are.
10
+
11
+ ## Corpus (pinned)
12
+
13
+ Originally scanned with `mcp-vet` v0.4.0 (re-scanned with v0.9.0 above) (`node dist/cli.js <roots> --json`), all rules
14
+ enabled, default confidence (`low`), on 2026-07-23:
15
+
16
+ | Repo | Commit | Scanned root | Files | LOC |
17
+ | --- | --- | --- | --- | --- |
18
+ | [modelcontextprotocol/servers](https://github.com/modelcontextprotocol/servers) | `d31124c982401739917fd817c2a59db344529c16` | `src/` | 78 | 14,742 |
19
+ | [modelcontextprotocol/typescript-sdk](https://github.com/modelcontextprotocol/typescript-sdk) | `1e1392e3f91583884fe82a0b4b91335875c3fba6` | `examples/` | 144 | 17,224 |
20
+ | [modelcontextprotocol/python-sdk](https://github.com/modelcontextprotocol/python-sdk) | `3a6f2996cdd8358957479791e8b26198c07d6a75` | `examples/` | 225 | 12,013 |
21
+ | **Total** | | | **447** | **43,979** |
22
+
23
+ File counts are candidate files (`.ts/.tsx/.js/.mjs/.cjs/.py`) under the
24
+ scanned roots, excluding `node_modules`.
25
+
26
+ ## Results — v0.9.0 (18-rule engine, final 2026-07-28 changelog)
27
+
28
+ Re-run 2026-07-28 with `mcp-vet` v0.9.0 on the SAME pinned corpus:
29
+ **243 findings across 66 files** (was 105/41 with the 9-rule engine — the
30
+ final changelog's removals fire heavily on the SDKs' own pre-final examples).
31
+ By confidence: 100 high, 142 medium, 1 low.
32
+
33
+ | Pattern | Findings |
34
+ | --- | --- |
35
+ | `SSE_RESUMABILITY_REMOVED` | 65 |
36
+ | `MCP_SESSION_ID` | 49 |
37
+ | `ELICITATION_COMPLETE_REMOVED` | 44 |
38
+ | `RESOURCE_SUBSCRIBE_REMOVED` | 23 |
39
+ | `SAMPLING_CAP` | 16 |
40
+ | `LOGGING_CAP` | 16 |
41
+ | `ROOTS_CAP` | 10 |
42
+ | `ROOTS_LIST_CHANGED_REMOVED` | 5 |
43
+ | `INITIALIZE_HANDLER` | 4 |
44
+ | `PING_REMOVED` | 3 |
45
+ | `TASKS_LEGACY` | 2 |
46
+ | `TASKS_RESULT_REMOVED` | 2 |
47
+ | `OAUTH_DCR` | 2 |
48
+ | `ERROR_CODE_RENUMBERED` | 1 |
49
+ | `LOGGING_SETLEVEL_REMOVED` | 1 |
50
+
51
+ Labeling of the 138 NEW findings (spot-reviewed per category against source):
52
+
53
+ - **136 true positives.** The corpus repos genuinely implement the removed
54
+ surfaces: the typescript-sdk `everything` example ships an
55
+ `InMemoryEventStore` + `Last-Event-ID` resumability transport (all 65 SSE
56
+ findings sit in those transport/resumability files), the elicitation examples
57
+ register `notifications/elicitation/complete` handlers and read
58
+ `elicitationId`, `everything/resources/subscriptions.ts` and friends register
59
+ `SubscribeRequestSchema`/`UnsubscribeRequestSchema`, and the bearer-auth
60
+ clients send `method: 'ping'`.
61
+ - **2 counted as false positives (honest reading):**
62
+ `guides/serving/sessions-state-scaling.examples.ts:62` returns
63
+ `code: -32001` for "Session not found" — an *implementation-defined* use the
64
+ final policy grandfathers; static analysis cannot distinguish it from the
65
+ renumbered `HeaderMismatch`, so ERROR_CODE_RENUMBERED flags it and we count
66
+ it against ourselves. Plus the pre-existing `mcp-session-id` negative
67
+ assertion (below). During this run a `registerTool('ping', ...)` false
68
+ positive (a tool merely NAMED ping) was found and FIXED before release —
69
+ strict registration context now excludes tool/prompt/resource registration
70
+ calls, locked into `negatives/`.
71
+
72
+ So the v0.9.0 headline on this corpus is **241/243 true positives (2 FP,
73
+ 0.8%)** — same discipline as before: FPs are counted, not defined away.
74
+
75
+ ## Results — v0.4.0 (9-rule engine, release-candidate era)
76
+
77
+ **105 findings across 41 files** (TypeScript/JavaScript: 93, Python: 12).
78
+ By confidence: 66 high, 38 medium, 1 low.
79
+
80
+ | Pattern | Findings |
81
+ | --- | --- |
82
+ | `MCP_SESSION_ID` | 49 |
83
+ | `LOGGING_CAP` | 17 |
84
+ | `SAMPLING_CAP` | 16 |
85
+ | `ROOTS_CAP` | 15 |
86
+ | `INITIALIZE_HANDLER` | 4 |
87
+ | `TASKS_LEGACY` | 2 |
88
+ | `TASKS_RESULT_REMOVED` | 2 |
89
+
90
+ ### Labeling
91
+
92
+ Every finding was manually reviewed against its source line:
93
+
94
+ - **104 / 105 true positives** — real references to a removed or deprecated
95
+ protocol surface (session headers/ids, handshake registration, legacy task
96
+ methods, deprecated capability declarations and method strings).
97
+ - **1 / 105 false positive (0.95%)** —
98
+ `stories/json_response/client.py:62` in the typescript-sdk examples:
99
+ `assert "mcp-session-id" not in response.headers`. That line is
100
+ *already-migrated* test code asserting the header is **absent**; flagging it
101
+ as "will break" is wrong. It is exactly what inline suppression
102
+ (`# mcp-vet-disable-line MCP_SESSION_ID`) is for, but we count it as a false
103
+ positive rather than defining it away. So the honest headline is
104
+ **"1 false positive in 44k LOC"**, not zero.
105
+
106
+ Notes on reading the numbers:
107
+
108
+ - Two occurrences on one line (e.g. `transport.sessionId && sessions.delete(transport.sessionId)`)
109
+ are reported as two findings — column-level dedup, not line-level.
110
+ - Findings in test files (`__tests__/…`) are counted as true positives: a test
111
+ that registers `sampling/createMessage` breaks the same way production code
112
+ does.
113
+
114
+ ### Labeled negatives
115
+
116
+ Files asserted to stay **clean** are part of the repo's test suite and run in CI:
117
+
118
+ - `test/fixtures/clean/` — a full server written in the 2026-07-28 style
119
+ (per-request `_meta`, `sessionIdGenerator: undefined`, `-32602`).
120
+ - `test/fixtures/negatives/` — "false friend" patterns: `sessionId` on plain
121
+ app-level objects, `-32002` inside strings/comments, capability-like words
122
+ with no capabilities context.
123
+ - `test/fixtures/adversarial/caught/` — obfuscations the scanner **must**
124
+ catch: aliased imports (TS + Python), namespace-qualified SDK constants,
125
+ client transports resuming a `sessionId`.
126
+
127
+ Additionally, in the corpus above, comment-only mentions (e.g. `Mcp-Session-Id`
128
+ in a comment, `initialize` in prose) produced zero findings — the AST layer
129
+ distinguishes executable tokens from comments by construction.
130
+
131
+ ## Recall — what the scanner is known to miss
132
+
133
+ Static token analysis proves known patterns are **absent**; it cannot prove
134
+ your server **speaks the new wire contract**. Recall is bounded by
135
+ construction, and the misses are locked into the test suite
136
+ (`test/fixtures/adversarial/missed/`, asserted to produce zero findings so any
137
+ silent claim-inflation fails CI):
138
+
139
+ - split/computed method strings — `'tasks' + '/list'`, `` `tasks/${op}` ``, f-strings
140
+ - computed capability keys — `{ ['roo'+'ts']: {} }`
141
+ - generated/loop-driven registration from string fragments
142
+ - framework-adapter indirection (route tables built at runtime)
143
+ - cross-module renames — a wrapper re-exporting an SDK constant under a new
144
+ name is flagged in the wrapper file, but a consumer importing only the new
145
+ name scans clean on its own
146
+
147
+ There is no corpus-wide recall *percentage*: that would require a labeled set
148
+ of every legacy usage in the wild, which nobody has. What we can say is: for
149
+ the pattern shapes listed in the README, detection is exact; for the shapes
150
+ above, it is zero, and the tool says so — pair the scan with runtime checks
151
+ (`mcp-vet fixtures`) to cover the difference.
152
+
153
+ ## Reproducing
154
+
155
+ ```bash
156
+ git clone --depth 1 https://github.com/modelcontextprotocol/servers
157
+ git clone --depth 1 https://github.com/modelcontextprotocol/typescript-sdk
158
+ git clone --depth 1 https://github.com/modelcontextprotocol/python-sdk
159
+ # check out the pinned SHAs above, then:
160
+ npx @booyaka/mcp-vet servers/src typescript-sdk/examples python-sdk/examples --json --no-files
161
+ ```
package/CHANGELOG.md CHANGED
@@ -4,6 +4,92 @@ All notable changes to `mcp-vet` are documented here. The format is based on
4
4
  [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), and this project adheres
5
5
  to [Semantic Versioning](https://semver.org/spec/v2.0.0.html).
6
6
 
7
+ ## [0.9.0]
8
+
9
+ The final-specification release. mcp-vet was built against the 2026-07-28
10
+ release candidate; the FINAL Key Changes list published 2026-07-28 at
11
+ <https://modelcontextprotocol.io/specification/draft/changelog> is materially
12
+ longer, so a clean 0.8.0 scan was a false all-clear. This release closes that
13
+ gap — every rule now cites a sentence pinned verbatim in
14
+ `docs/SPEC-2026-07-28.md`. (The dated URL
15
+ `/specification/2026-07-28` still returns 404; the final text is served under
16
+ `/specification/draft/`.)
17
+
18
+ ### Added — seven BREAKING static rules (exit 1), in BOTH analyzers
19
+
20
+ - **`PING_REMOVED`** — `ping` in MCP method-registration context,
21
+ `PingRequestSchema`, Python `types.PingRequest`. A `/ping` health route, a
22
+ bare `'ping'` string, or a tool merely *named* ping stays clean.
23
+ - **`RESOURCE_SUBSCRIBE_REMOVED`** — `resources/subscribe`,
24
+ `resources/unsubscribe`, `SubscribeRequestSchema`,
25
+ `UnsubscribeRequestSchema`; the fix points at `subscriptions/listen` and its
26
+ four opt-in types.
27
+ - **`ROOTS_LIST_CHANGED_REMOVED`** and **`LOGGING_SETLEVEL_REMOVED`** — a
28
+ RECLASSIFICATION: 0.8.0 mapped `notifications/roots/list_changed` and
29
+ `logging/setLevel` (+ their SDK schema constants) to the DEPRECATED
30
+ capability rules, reporting two hard removals as exit-0 warnings with a
31
+ grace-period label. They are BREAKING now; the `roots`/`logging` capability
32
+ *keys* stay DEPRECATED, and a test locks the severity split.
33
+ - **`SSE_RESUMABILITY_REMOVED`** — the `Last-Event-ID` header string,
34
+ `lastEventId`, and `eventStore`/`resumptionToken`/`onresumptiontoken` passed
35
+ to a Streamable HTTP transport. Gated on file-level MCP context, so a plain
36
+ SSE client stays clean (locked by `negatives/sse-client.ts`).
37
+ - **`ELICITATION_COMPLETE_REMOVED`** — `notifications/elicitation/complete`
38
+ and the `elicitationId` field.
39
+ - **`ERROR_CODE_RENUMBERED`** — `-32001`/`-32003`/`-32004` → `-32020`/`-32021`/
40
+ `-32022`, flagged ONLY in a JSON-RPC error `code` position (`code:` key,
41
+ `*Error(...)` construction, comparison against `code`) — the changelog
42
+ grandfathers `-32000..-32019` for implementation-defined codes, so a bare
43
+ negative constant is never flagged.
44
+
45
+ ### Added — two DEPRECATED static rules (exit 0)
46
+
47
+ - **`INCLUDE_CONTEXT_VALUES`** — `includeContext` set to `"thisServer"` /
48
+ `"allServers"` (medium; removal "Follows Sampling").
49
+ - **`OAUTH_DCR`** — RFC7591 dynamic-client-registration surfaces
50
+ (`registration_endpoint`, `registration_access_token`,
51
+ `client_id_issued_at`) in favour of Client ID Metadata Documents (medium).
52
+
53
+ ### Added — probe & fixtures
54
+
55
+ - Four checks join the opt-in `mcp-vet probe --spec 2026-07-28` suite (now ten):
56
+ **`missing-result-type`** (ERROR, SEP-2322), **`missing-cacheable-fields`**
57
+ (WARN, SEP-2549 `ttlMs` + `cacheScope`), **`legacy-error-code-renumbered`**
58
+ (ERROR — still answering `-32001`/`-32003`/`-32004`), and
59
+ **`ping-still-answered`** (WARN — `ping` returns a result instead of
60
+ `-32601`). All cross-checked: a dead or non-MCP server is exit 2, never a
61
+ false violation, and every inconclusive outcome is a note.
62
+ - `mcp-vet fixtures` gains **`10-subscriptions-listen`** (opt-in +
63
+ `io.modelcontextprotocol/subscriptionId` tagging + `resources/subscribe` →
64
+ -32601) and **`11-mrtr`** (`resultType: "input_required"` + `inputRequests`,
65
+ retry with `inputResponses`) — eleven fixtures total.
66
+ - `--fix` now rewrites the renumbered codes next to the existing `-32002` →
67
+ `-32602` (same-length, column-anchored); `--dry-run` lists every rewrite.
68
+ - DEPRECATED findings now print the registry's exact removal window ("First
69
+ revision released on or after 2027-07-28", "Follows Sampling") instead of a
70
+ hardcoded 12 months.
71
+ - New `test/fixtures/dirty/` (TS + Python, one instance of every new pattern),
72
+ migrated forms in `clean/`, new true-negatives, computed/split forms in
73
+ `adversarial/missed/`. 78 tests (was 71), none skipped.
74
+ - BENCHMARK.md re-measured with the 18-rule engine on the same pinned corpus.
75
+
76
+ ## [0.8.0]
77
+
78
+ Maintenance release — dependency and CI currency; no rule or probe changes.
79
+
80
+ ### Changed
81
+
82
+ - Dropped `chalk` for a local colouriser; moved to `commander` 15,
83
+ `ts-morph` 28, `@types/node` 26.
84
+ - CI tests on supported Node only (22/24/26); `engines.node` >= 22.
85
+ - `mcp-vet probe` lets the event loop drain instead of calling
86
+ `process.exit()` — fixes an intermittent Windows libuv crash (0xC0000409)
87
+ on process teardown.
88
+ - GitHub Actions moved to latest majors; Dependabot groups minor/patch and
89
+ splits majors; the blocked TypeScript 7 major is ignored until the
90
+ Compiler-API crash upstream clears; our own guards (npm-script-lens audit +
91
+ allowlist drift, ts7-compat-guard, pnpm11-ci-guard) run against this repo.
92
+
7
93
  ## [0.7.0]
8
94
 
9
95
  An opt-in `--spec 2026-07-28` compliance suite for `mcp-vet probe` — six