opentel-mcp 0.8.0 → 0.10.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +161 -0
- package/README.md +508 -19
- package/package.json +14 -2
- package/src/config.js +55 -0
- package/src/fingerprint/classify/auth.js +41 -3
- package/src/fingerprint/classify/channel.js +83 -7
- package/src/fingerprint/classify/validation-paths.js +244 -18
- package/src/index.d.ts +135 -19
- package/src/instrument.js +524 -93
- package/src/registry/instance-registry.js +170 -0
- package/src/registry/types.d.ts +55 -0
- package/src/sdk/detect.js +118 -0
- package/src/thrash/config.js +14 -8
package/CHANGELOG.md
CHANGED
|
@@ -1,5 +1,166 @@
|
|
|
1
1
|
# Changelog
|
|
2
2
|
|
|
3
|
+
## 0.10.0
|
|
4
|
+
|
|
5
|
+
**⚠️ Behavior change, unrelated to the feature below — read this first.**
|
|
6
|
+
`instrumentMcpServer()` now throws for a server object it cannot
|
|
7
|
+
confidently wrap, instead of silently instrumenting nothing.
|
|
8
|
+
`detectServerKind()` (`src/instrument.js`) previously accepted any
|
|
9
|
+
`McpServer`-shaped object whose `.server` merely *had* a
|
|
10
|
+
`setRequestHandler` method — it now additionally requires `.server
|
|
11
|
+
instanceof <Server>` for a real, recognized SDK class. An object that
|
|
12
|
+
satisfies the outer shape but fails that check now throws a new,
|
|
13
|
+
specific error (`UNWRAPPABLE_MCPSERVER_ERROR` — names what was detected
|
|
14
|
+
and the plausible causes: a duplicate/mismatched SDK install, an SDK not
|
|
15
|
+
resolvable from this package's own location, or an unsupported SDK) at
|
|
16
|
+
`instrumentMcpServer()` call time, rather than succeeding and producing
|
|
17
|
+
zero telemetry. This closes a confirmed gap (`docs/known-gaps.md` entry
|
|
18
|
+
7, now marked fixed): an `@modelcontextprotocol/server` (MCP v2) object
|
|
19
|
+
passed to a pre-0.10.0 `instrumentMcpServer()` satisfied the old, looser
|
|
20
|
+
check and appeared to instrument successfully — `getThrashSummary`/
|
|
21
|
+
`getObservationState` attached, no error — while producing zero spans,
|
|
22
|
+
zero metrics, and zero fingerprinting for every tool call. No escape
|
|
23
|
+
hatch was added; see ADR 015 (`docs/adr/015-mcp-v2-support.md`) for the
|
|
24
|
+
full argument against one. **If you're seeing this new error on upgrade**
|
|
25
|
+
and you believe your object genuinely is a real `Server`/`McpServer`
|
|
26
|
+
instance, check for a duplicate/mismatched install of whichever SDK it
|
|
27
|
+
came from (`npm dedupe`, or check for multiple installed copies) — a real
|
|
28
|
+
v1 or v2 `Server`/`McpServer` from a single, consistently-resolved SDK
|
|
29
|
+
install is unaffected by this change.
|
|
30
|
+
|
|
31
|
+
### Added — `@modelcontextprotocol/server` (MCP v2, protocol revision 2026-07-28) support
|
|
32
|
+
|
|
33
|
+
Both the original `@modelcontextprotocol/sdk` ("v1") and the new
|
|
34
|
+
`@modelcontextprotocol/server` ("v2") now work with `instrumentMcpServer()`
|
|
35
|
+
— two separate, OPTIONAL peer dependencies (install whichever one(s) you
|
|
36
|
+
actually use; `package.json`'s `peerDependenciesMeta` marks both
|
|
37
|
+
`optional: true`, verified against real, clean external installs with
|
|
38
|
+
only one, the other, or neither installed — not just `package.json`
|
|
39
|
+
syntax). Same `Server`/`McpServer` API shapes as v1; detection and
|
|
40
|
+
wrapping happen automatically, resolved once per `instrumentMcpServer()`
|
|
41
|
+
call by which SDK the object actually came from. Full design and
|
|
42
|
+
Phase-by-phase implementation notes: ADR 015
|
|
43
|
+
(`docs/adr/015-mcp-v2-support.md`).
|
|
44
|
+
|
|
45
|
+
What works the same as v1: spans, standard attributes (including
|
|
46
|
+
`jsonrpc.request.id`, now read from v2's `ctx.mcpReq.id`), deep failure
|
|
47
|
+
fingerprinting, and `mcp.failure.channel`/`mcp.failure.validation_paths`
|
|
48
|
+
classification (`channel.js`/`validation-paths.js` both gained a
|
|
49
|
+
v2-specific code path — the "MCP error N:" wrapper v1 disguises errors
|
|
50
|
+
with doesn't exist in v2, and v2's rendered validation-issue text uses a
|
|
51
|
+
third, distinct format from either of v1's two).
|
|
52
|
+
|
|
53
|
+
**v2's own `createMcpHandler`/`serveStdio` construct a fresh `Server`/
|
|
54
|
+
`McpServer` per request by default (a factory function you provide), not
|
|
55
|
+
once at process start.** `instrumentMcpServer()` needs to run *inside*
|
|
56
|
+
that factory, on every invocation — see the README's new "MCP v2 support"
|
|
57
|
+
section for a worked example. `instanceKey` (v0.9.0) is the existing
|
|
58
|
+
mechanism for sharing tracker state across those repeated calls; nothing
|
|
59
|
+
new was added for this, since ADR 012's original design already covers
|
|
60
|
+
this exact deployment shape, v2 just makes it the default instead of an
|
|
61
|
+
edge case.
|
|
62
|
+
|
|
63
|
+
**Two gaps not closed this release, both tracked in `docs/known-gaps.md`
|
|
64
|
+
with a "Status update (v0.10.0)" note:** Agent Thrash Detection's fallback
|
|
65
|
+
session id still doesn't survive v2's per-request factory pattern even
|
|
66
|
+
with `instanceKey` set (entry 6 — real session ids work fine either way),
|
|
67
|
+
and `isSingleConnectionTransport()`'s transport-detection heuristic still
|
|
68
|
+
misclassifies the transport `createMcpHandler` builds internally (entry
|
|
69
|
+
8, live as of this release, not merely forward-looking). Both were
|
|
70
|
+
explicitly scoped out of this round, not overlooked; workaround for
|
|
71
|
+
either: `thrashDetection: { enabled: false }`.
|
|
72
|
+
|
|
73
|
+
## 0.9.0
|
|
74
|
+
|
|
75
|
+
**⚠️ Fixed, with a fingerprint behavior change — read this before the
|
|
76
|
+
feature below.** The `auth` failure classifier
|
|
77
|
+
(`src/fingerprint/classify/auth.js`) missed "permission denied" and
|
|
78
|
+
"access denied" — the standard phrasing from Unix, git, AWS IAM, and GCP
|
|
79
|
+
for a permission/authorization failure. It only recognized
|
|
80
|
+
HTTP-status-derived wording (`unauthorized`, `forbidden`,
|
|
81
|
+
`authenticat(e|ion)`) plus 401/403 status codes and a handful of known
|
|
82
|
+
auth-library error names. Messages using the OS/CLI phrasing above fell
|
|
83
|
+
through every classifier and landed in the `internal` catch-all instead.
|
|
84
|
+
Found by running realistic error text through this project's own UI demo
|
|
85
|
+
(`packages/ui/demo/populate.js`) and checking what `DEFAULT_CLASSIFIERS`
|
|
86
|
+
actually returned for it — not assumed. Now also matches "not authorized",
|
|
87
|
+
"permission(s) denied", "access denied", "insufficient permission(s)", and
|
|
88
|
+
Node's own `EACCES`/`EPERM` error codes; still does not match bare
|
|
89
|
+
"authorized" or "permission" alone (see the classifier's own docblock for
|
|
90
|
+
the false-positive cases this deliberately excludes, e.g. "user denied the
|
|
91
|
+
permission request").
|
|
92
|
+
|
|
93
|
+
**This changes fingerprints for affected messages.** `category` is one of
|
|
94
|
+
the hashed inputs `computeFingerprint()` combines into
|
|
95
|
+
`mcp.failure.fingerprint` (see ADR 006). A permission-denied failure that
|
|
96
|
+
previously classified as `internal` now classifies as `auth` — the fields
|
|
97
|
+
feeding the hash change, so the fingerprint itself changes for anyone whose
|
|
98
|
+
tool emits this wording. This does **not** amend the closed 8-category
|
|
99
|
+
taxonomy ADR 006 established (`validation | timeout | network | auth |
|
|
100
|
+
dependency | serialization | internal | unknown`) — `auth` already existed;
|
|
101
|
+
this is a pattern-coverage fix to when the existing category fires, not a
|
|
102
|
+
new category. If you alert or dashboard on a specific `mcp.failure.fingerprint`
|
|
103
|
+
value for a permission error, expect a new value after upgrading.
|
|
104
|
+
|
|
105
|
+
### Added — `instanceKey`: sharing tracker state across `instrumentMcpServer()` calls
|
|
106
|
+
|
|
107
|
+
`instanceKey` (a string option on `instrumentMcpServer()`, or the
|
|
108
|
+
`OTEL_MCP_INSTANCE_KEY` env var — lower precedence than the option) lets
|
|
109
|
+
repeated `instrumentMcpServer()` calls that pass the same key share Agent
|
|
110
|
+
Thrash Detection, budget tracking, schema drift detection, and the
|
|
111
|
+
`ToolOutcome` counter's state, instead of each call constructing all four
|
|
112
|
+
fresh and discarding them. Fixes the gap documented in the README's
|
|
113
|
+
"In-memory tracker state is scoped to one `instrumentMcpServer()` call"
|
|
114
|
+
section and `docs/known-gaps.md` entry 6, under a "stateless" Streamable
|
|
115
|
+
HTTP deployment shape (a fresh `Server`/`McpServer` re-instrumented on
|
|
116
|
+
every incoming request). Full design: ADR 012
|
|
117
|
+
(`docs/adr/012-tracker-lifecycle-and-shared-state.md`).
|
|
118
|
+
|
|
119
|
+
Backed by an internal, bounded, TTL-evicting registry (1000 distinct keys
|
|
120
|
+
per process, 24h TTL renewed on every use — both ADR 012's proposed
|
|
121
|
+
defaults) — fully internal, no new public type. Omit `instanceKey` (the
|
|
122
|
+
default) for behavior byte-identical to every prior version: trackers are
|
|
123
|
+
constructed fresh on every call, and the registry is never touched.
|
|
124
|
+
|
|
125
|
+
**⚠️ Composition requirement: `instanceKey` alone does not fix Agent
|
|
126
|
+
Thrash Detection.** `ThrashDetector` looks episodes up by `(sessionId,
|
|
127
|
+
toolName, fingerprint)` — `instanceKey` shares the tracker object, but
|
|
128
|
+
without a real, transport-provided `extra.sessionId` on every call, each
|
|
129
|
+
`instrumentMcpServer()` call still generates its own random per-connection
|
|
130
|
+
fallback session id, fresh, regardless of `instanceKey`. Sharing the
|
|
131
|
+
tracker doesn't help if the lookup key inside it differs every call —
|
|
132
|
+
each request lands as its own one-off episode instead of contributing to
|
|
133
|
+
one shared loop. Real Streamable HTTP transports provide a real session id
|
|
134
|
+
automatically, so the common case works with `instanceKey` alone — but a
|
|
135
|
+
custom `Transport`, `assumeSingleSession: true`, or anything else on the
|
|
136
|
+
generated-fallback path will set `instanceKey`, see nothing happen, and
|
|
137
|
+
have every reason to think the fix is broken. Same silent-inertness shape
|
|
138
|
+
as the original gap, one layer deeper. See the README's new "instanceKey"
|
|
139
|
+
section for the full explanation and what to do about it — this is not a
|
|
140
|
+
footnote there either.
|
|
141
|
+
|
|
142
|
+
**⚠️ Does not help across process boundaries.** `instanceKey`'s registry
|
|
143
|
+
is one process's in-memory state. On Lambda, Cloud Run, or any
|
|
144
|
+
horizontally-scaled deployment, concurrent/recycled instances each hold
|
|
145
|
+
their own independent registry — passing the identical `instanceKey`
|
|
146
|
+
string everywhere does not change that. Counters remain instance-local and
|
|
147
|
+
best-effort by design; this is a structural limitation, not a
|
|
148
|
+
configuration gap, and this library deliberately does not add an external
|
|
149
|
+
store (Redis/DynamoDB) to close it — see ADR 012's Update section for the
|
|
150
|
+
full reasoning.
|
|
151
|
+
|
|
152
|
+
**Documentation:** the "Metrics" section now covers wiring the Prometheus
|
|
153
|
+
exporter specifically (`@opentelemetry/exporter-prometheus`), not just the
|
|
154
|
+
OTLP example that was already there. Its pull-based text exposition format
|
|
155
|
+
does not attach resource attributes (including `service.name`) to
|
|
156
|
+
individual metric points by default — only to a separate `target_info`
|
|
157
|
+
series — which is invisible with one service but means every series looks
|
|
158
|
+
identical the moment you're scraping more than one instrumented server
|
|
159
|
+
into the same Prometheus. Documents the fix
|
|
160
|
+
(`withResourceConstantLabels: /^service\.name$/`) with a worked example.
|
|
161
|
+
No code change; this behavior was always there, just undocumented. Found
|
|
162
|
+
building `dashboards/grafana-mcp-health.json`'s verification harness.
|
|
163
|
+
|
|
3
164
|
## 0.8.0
|
|
4
165
|
|
|
5
166
|
Three features. **Tool schema drift detection**: a server that silently
|