opentel-mcp 0.9.0 → 0.10.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/CHANGELOG.md CHANGED
@@ -1,5 +1,75 @@
1
1
  # Changelog
2
2
 
3
+ ## 0.10.0
4
+
5
+ **⚠️ Behavior change, unrelated to the feature below — read this first.**
6
+ `instrumentMcpServer()` now throws for a server object it cannot
7
+ confidently wrap, instead of silently instrumenting nothing.
8
+ `detectServerKind()` (`src/instrument.js`) previously accepted any
9
+ `McpServer`-shaped object whose `.server` merely *had* a
10
+ `setRequestHandler` method — it now additionally requires `.server
11
+ instanceof <Server>` for a real, recognized SDK class. An object that
12
+ satisfies the outer shape but fails that check now throws a new,
13
+ specific error (`UNWRAPPABLE_MCPSERVER_ERROR` — names what was detected
14
+ and the plausible causes: a duplicate/mismatched SDK install, an SDK not
15
+ resolvable from this package's own location, or an unsupported SDK) at
16
+ `instrumentMcpServer()` call time, rather than succeeding and producing
17
+ zero telemetry. This closes a confirmed gap (`docs/known-gaps.md` entry
18
+ 7, now marked fixed): an `@modelcontextprotocol/server` (MCP v2) object
19
+ passed to a pre-0.10.0 `instrumentMcpServer()` satisfied the old, looser
20
+ check and appeared to instrument successfully — `getThrashSummary`/
21
+ `getObservationState` attached, no error — while producing zero spans,
22
+ zero metrics, and zero fingerprinting for every tool call. No escape
23
+ hatch was added; see ADR 015 (`docs/adr/015-mcp-v2-support.md`) for the
24
+ full argument against one. **If you're seeing this new error on upgrade**
25
+ and you believe your object genuinely is a real `Server`/`McpServer`
26
+ instance, check for a duplicate/mismatched install of whichever SDK it
27
+ came from (`npm dedupe`, or check for multiple installed copies) — a real
28
+ v1 or v2 `Server`/`McpServer` from a single, consistently-resolved SDK
29
+ install is unaffected by this change.
30
+
31
+ ### Added — `@modelcontextprotocol/server` (MCP v2, protocol revision 2026-07-28) support
32
+
33
+ Both the original `@modelcontextprotocol/sdk` ("v1") and the new
34
+ `@modelcontextprotocol/server` ("v2") now work with `instrumentMcpServer()`
35
+ — two separate, OPTIONAL peer dependencies (install whichever one(s) you
36
+ actually use; `package.json`'s `peerDependenciesMeta` marks both
37
+ `optional: true`, verified against real, clean external installs with
38
+ only one, the other, or neither installed — not just `package.json`
39
+ syntax). Same `Server`/`McpServer` API shapes as v1; detection and
40
+ wrapping happen automatically, resolved once per `instrumentMcpServer()`
41
+ call by which SDK the object actually came from. Full design and
42
+ Phase-by-phase implementation notes: ADR 015
43
+ (`docs/adr/015-mcp-v2-support.md`).
44
+
45
+ What works the same as v1: spans, standard attributes (including
46
+ `jsonrpc.request.id`, now read from v2's `ctx.mcpReq.id`), deep failure
47
+ fingerprinting, and `mcp.failure.channel`/`mcp.failure.validation_paths`
48
+ classification (`channel.js`/`validation-paths.js` both gained a
49
+ v2-specific code path — the "MCP error N:" wrapper v1 disguises errors
50
+ with doesn't exist in v2, and v2's rendered validation-issue text uses a
51
+ third, distinct format from either of v1's two).
52
+
53
+ **v2's own `createMcpHandler`/`serveStdio` construct a fresh `Server`/
54
+ `McpServer` per request by default (a factory function you provide), not
55
+ once at process start.** `instrumentMcpServer()` needs to run *inside*
56
+ that factory, on every invocation — see the README's new "MCP v2 support"
57
+ section for a worked example. `instanceKey` (v0.9.0) is the existing
58
+ mechanism for sharing tracker state across those repeated calls; nothing
59
+ new was added for this, since ADR 012's original design already covers
60
+ this exact deployment shape, v2 just makes it the default instead of an
61
+ edge case.
62
+
63
+ **Two gaps not closed this release, both tracked in `docs/known-gaps.md`
64
+ with a "Status update (v0.10.0)" note:** Agent Thrash Detection's fallback
65
+ session id still doesn't survive v2's per-request factory pattern even
66
+ with `instanceKey` set (entry 6 — real session ids work fine either way),
67
+ and `isSingleConnectionTransport()`'s transport-detection heuristic still
68
+ misclassifies the transport `createMcpHandler` builds internally (entry
69
+ 8, live as of this release, not merely forward-looking). Both were
70
+ explicitly scoped out of this round, not overlooked; workaround for
71
+ either: `thrashDetection: { enabled: false }`.
72
+
3
73
  ## 0.9.0
4
74
 
5
75
  **⚠️ Fixed, with a fingerprint behavior change — read this before the
package/README.md CHANGED
@@ -129,6 +129,111 @@ await server.connect(transport);
129
129
  Runnable versions of both live in `examples/hello-server/` and
130
130
  `examples/hello-mcpserver/`.
131
131
 
132
+ ## MCP v2 (`@modelcontextprotocol/server`) support (v0.10.0+)
133
+
134
+ Both the original SDK and the new one work with `instrumentMcpServer()` —
135
+ they're two separate, OPTIONAL peer dependencies (install whichever one(s)
136
+ you actually use):
137
+
138
+ - `@modelcontextprotocol/sdk` ("v1" throughout this README) — protocol
139
+ revisions through 2025-11-25. The `Server`/`McpServer` APIs above.
140
+ - `@modelcontextprotocol/server` ("v2") — protocol revision 2026-07-28,
141
+ whose headline change is removing the `initialize` handshake and the
142
+ `Mcp-Session-Id` Streamable HTTP header in favor of a stateless,
143
+ self-contained-request model. Same `Server`/`McpServer` shapes, same
144
+ `instrumentMcpServer()` call — detection and wrapping happen
145
+ automatically, resolved once per `instrumentMcpServer()` call by which
146
+ SDK the object you passed in actually came from (ADR 015,
147
+ `docs/adr/015-mcp-v2-support.md`).
148
+
149
+ ```js
150
+ import { McpServer } from '@modelcontextprotocol/server';
151
+ import { instrumentMcpServer } from 'opentel-mcp';
152
+ import { z } from 'zod';
153
+
154
+ const server = new McpServer({ name: 'my-server', version: '1.0.0' });
155
+ instrumentMcpServer(server, { serviceName: 'my-mcp-server' });
156
+
157
+ server.registerTool('echo', { inputSchema: z.object({ text: z.string() }) }, async ({ text }) => ({
158
+ content: [{ type: 'text', text }],
159
+ }));
160
+ ```
161
+
162
+ **The important difference isn't the API — it's the deployment shape.**
163
+ v2's own `createMcpHandler`/`serveStdio` entry points construct a fresh
164
+ `Server`/`McpServer` instance **per request** (via a factory function you
165
+ provide), not once at process start — including for `createMcpHandler`'s
166
+ default stateless HTTP posture, not just an edge case. That means
167
+ `instrumentMcpServer()` has to run **inside the factory**, on every
168
+ invocation, not once at module load the way the v1 examples above do:
169
+
170
+ ```js
171
+ import { createMcpHandler } from '@modelcontextprotocol/server';
172
+ import { McpServer } from '@modelcontextprotocol/server';
173
+ import { instrumentMcpServer } from 'opentel-mcp';
174
+
175
+ const handler = createMcpHandler((ctx) => {
176
+ const server = new McpServer({ name: 'my-server', version: '1.0.0' });
177
+
178
+ // Runs on every request this factory serves. instanceKey is what makes
179
+ // that not mean "trackers reset every time" — see below.
180
+ instrumentMcpServer(server, { serviceName: 'my-mcp-server', instanceKey: 'my-mcp-server' });
181
+
182
+ server.registerTool('echo', { inputSchema: z.object({ text: z.string() }) }, async ({ text }) => ({
183
+ content: [{ type: 'text', text }],
184
+ }));
185
+
186
+ return server;
187
+ });
188
+ ```
189
+
190
+ **`instanceKey` (see the dedicated section below) is the mechanism for
191
+ this** — it's not a new, v2-specific option; it's the same one ADR 012
192
+ built for a v1 "stateless Streamable HTTP" deployment shape that turned
193
+ out to be exactly what v2 makes the *default*, SDK-recommended pattern
194
+ instead of something a host happened to build. Without it, the same four
195
+ in-memory trackers ("In-memory tracker state is scoped to one
196
+ instrumentMcpServer() call" below) reset to empty on every request under
197
+ this pattern, same as they always have for any fresh-instance-per-request
198
+ deployment — v2 doesn't change that mechanism, it just makes hitting it
199
+ the default instead of an edge case.
200
+
201
+ **What works today:** spans, standard attributes (`mcp.method.name`,
202
+ `gen_ai.tool.name`, `jsonrpc.request.id` — read from v2's `ctx.mcpReq.id`),
203
+ deep failure fingerprinting, `mcp.failure.channel`/`mcp.failure.validation_paths`
204
+ classification, and — as of this release — Agent Thrash Detection's
205
+ fallback-session path and transport auto-detection all work the same as
206
+ v1 (ADR 015 Phases 1–3, plus a follow-up round documented in ADR 015's
207
+ Update section):
208
+
209
+ - `isSingleConnectionTransport()` requires positive confirmation
210
+ (`transport.constructor.name === 'StdioServerTransport'`) for v2
211
+ specifically, instead of inferring single-connection from an absent
212
+ `sessionId` property — closing a confirmed false positive against the
213
+ transport `createMcpHandler` builds internally
214
+ (`PerRequestHTTPServerTransport`), which also has no `sessionId`, for
215
+ the opposite reason stdio doesn't. v1 is completely unaffected by this
216
+ change. See `docs/known-gaps.md` entry 8 for the full history.
217
+ - The generated fallback session id is now shared across repeated
218
+ `instrumentMcpServer()` calls when `instanceKey` is set (the same
219
+ registry the four ADR-012 trackers already use), fixing the gap
220
+ "instanceKey alone does not fix thrash detection" below originally
221
+ described for v1's stateless deployment shape, for the v2 factory
222
+ pattern specifically. See `docs/known-gaps.md` entry 6.
223
+
224
+ **One narrower thing still open**, tracked in `docs/known-gaps.md` entry
225
+ 6's own update: `thrashSessionState` (the internal flag tracking "has
226
+ this server ever proven itself session-aware") is not registry-backed the
227
+ way the fallback id now is, so that specific memory still resets on every
228
+ v2 per-request call. This only matters for a deployment that mixes
229
+ real-session-id calls with occasional no-session-id ones under a shared
230
+ `instanceKey` — thrash detection using a **real** session id on every
231
+ call (`ctx.sessionId`) is unaffected either way.
232
+
233
+ Runnable end-to-end coverage lives in `test/instrument.v2.test.js` and
234
+ `test/integration/thrash-v2-transport-detection.test.js`, not a dedicated
235
+ `examples/` directory yet.
236
+
132
237
  ## What gets emitted
133
238
 
134
239
  ### Tool-level failures, specifically
@@ -484,7 +589,13 @@ convenience default, not a maintained price list.
484
589
  one instrumentMcpServer() call" below for what that means under a
485
590
  fresh-`Server`-per-request deployment. Session-scoped limits are also
486
591
  skipped gracefully (not enforced against a fallback key) for transports
487
- with no session id, like stdio.
592
+ with no session id, like stdio. **`perSessionUsd` is unsupported under
593
+ stateless MCP (no real session id available at all) — unchanged from
594
+ ADR 012's conclusion, and there is no fix pending.** Unlike thrash
595
+ detection, this budget tracker has no fingerprint-equivalent attribute
596
+ on the span today, so "Fleet-wide fingerprint frequency" below has
597
+ nothing to substitute with for this scope specifically. `perToolUsd` is
598
+ unaffected — it never depended on session id.
488
599
 
489
600
  ## Agent Thrash Detection (v0.6.0+)
490
601
 
@@ -735,6 +846,22 @@ in order:
735
846
  `assumeSingleSession` left at its default `false` — detection is
736
847
  **skipped silently** for that call rather than guessing.
737
848
 
849
+ **For `@modelcontextprotocol/server` (v2) users:** this structural check —
850
+ "no `sessionId` property on the transport" — is v1-specific. v2's
851
+ transport `createMcpHandler` builds internally
852
+ (`PerRequestHTTPServerTransport`) *also* declares no `sessionId`
853
+ property, but for the opposite reason stdio doesn't: the 2026-07-28
854
+ protocol revision has no session concept at the transport level at all,
855
+ not because each instance is genuinely 1:1 with one client — so for v2
856
+ specifically, this function requires POSITIVE confirmation
857
+ (`transport.constructor.name === 'StdioServerTransport'`) instead of
858
+ inferring single-connection from the absent property. `PerRequestHTTPServerTransport`
859
+ correctly falls through to "undetermined" rather than being misclassified.
860
+ v1's own check is completely unaffected by this. See `docs/known-gaps.md`
861
+ entry 8 for the full history (including the period where this was a
862
+ confirmed, live false positive) and ADR 015's Update section for the
863
+ design argument.
864
+
738
865
  **The risk of getting this wrong:** if you set `assumeSingleSession: true`
739
866
  on a transport that's actually serving multiple concurrent clients (a
740
867
  typical HTTP/SSE deployment behind a load balancer, for instance), their
@@ -798,6 +925,14 @@ Streamable HTTP), consecutive-failure tracking never accumulates past a
798
925
  single call, and `mcp.tool.loop.detected` never fires — silently. Confirmed
799
926
  gap, ADR 012.
800
927
 
928
+ **For deployments with no real session id at all** (see "instanceKey"
929
+ above, specifically the "MCP spec 2026-07-28" subsection, for exactly
930
+ when this applies) — **thrash detection is unsupported, unchanged from
931
+ ADR 012's conclusion.** There is a partial, downstream substitute: "Fleet-
932
+ wide fingerprint frequency" below. Read its caveat before treating it as a
933
+ fix for this gap — it detects how often a bug occurs across every caller,
934
+ not whether one agent is looping.
935
+
801
936
  **A malformed `tools/call` request produces zero telemetry — no span, no
802
937
  fingerprint, nothing.** If a request fails `CallToolRequestSchema`
803
938
  validation itself (e.g. a missing or wrongly-typed `name`/`arguments`
@@ -1457,9 +1592,16 @@ different, also-common shape: **"stateless" Streamable HTTP, where a fresh
1457
1592
  Under that topology, every one of these four trackers is discarded and
1458
1593
  rebuilt from empty before it ever sees a second data point, unless
1459
1594
  `instanceKey` is set — and, for Agent Thrash Detection specifically, a
1460
- real session id is also available on every call. Without both, nothing
1461
- accumulates, nothing crosses a threshold, and nothing warns that this is
1462
- happening — the affected feature is silently inert.
1595
+ real session id is also available on every call. MCP spec 2025-11-25 and
1596
+ earlier transports give you this automatically; **`@modelcontextprotocol/server`
1597
+ (v2, protocol revision 2026-07-28) still has a real, optional `sessionId`
1598
+ field on every call — it isn't removed — but its default,
1599
+ `createMcpHandler`-driven stateless deployment shape usually doesn't
1600
+ populate it, the same "no session id" shape stdio has always had for v1**
1601
+ — see "MCP spec 2026-07-28..." below for exactly what this does and
1602
+ doesn't mean. Without both a shared tracker and a real session id on every
1603
+ call, nothing accumulates, nothing crosses a threshold, and nothing warns
1604
+ that this is happening — the affected feature is silently inert.
1463
1605
 
1464
1606
  **Confirmed, not a hypothetical, and now has a partial fix.** Reproduced
1465
1607
  directly in `test/integration/thrash-stateless-http-lifecycle.test.js`:
@@ -1473,13 +1615,18 @@ trackers, and the `instanceKey` design: ADR 012
1473
1615
  (`docs/adr/012-tracker-lifecycle-and-shared-state.md`). Tracked in
1474
1616
  `docs/known-gaps.md`.
1475
1617
 
1476
- **If you instrument a fresh `Server`/`McpServer` per request: set
1477
- `instanceKey`.** Read the section immediately below in full before relying
1478
- on it — it has one required companion for thrash detection specifically
1479
- (a real session id, not the generated fallback), and it does not help at
1480
- all across multiple processes or containers (Lambda, Cloud Run, or any
1481
- horizontally-scaled deployment). Both are easy to miss and produce the
1482
- exact same silent-inertness symptom as this section describes.
1618
+ **If you instrument a fresh `Server`/`McpServer` per request — including
1619
+ every `@modelcontextprotocol/server` (v2) deployment via `createMcpHandler`/
1620
+ `serveStdio`, whose factory pattern makes this the default, not an edge
1621
+ case — set `instanceKey`.** Read the section immediately below in full
1622
+ before relying on it — it has one required companion for thrash detection
1623
+ specifically (a real session id, not the generated fallback — under v2's
1624
+ default stateless posture this companion usually isn't met, not because
1625
+ it's structurally impossible but because nothing provides one; see "MCP
1626
+ spec 2026-07-28..." below), and it does not help at all across multiple
1627
+ processes or containers (Lambda, Cloud Run, or any horizontally-scaled
1628
+ deployment). All three are easy to miss and produce the exact same
1629
+ silent-inertness symptom as this section describes.
1483
1630
 
1484
1631
  ## instanceKey: sharing tracker state across instrumentMcpServer() calls (v0.9.0+)
1485
1632
 
@@ -1542,19 +1689,25 @@ past 1, and nothing warns you.
1542
1689
  2. a real `extra.sessionId` on every call, so the lookup key inside that
1543
1690
  shared tracker is stable across calls too.
1544
1691
 
1545
- Real Streamable HTTP transports give you (2) automatically — the SDK
1546
- threads a real client session id through `extra.sessionId` on every
1547
- request regardless of whether the `Server` object handling it was just
1548
- constructed, so the common "stateless Streamable HTTP" case works with
1549
- `instanceKey` alone, no extra effort. **You will NOT get (2) for free —
1550
- and thrash detection will stay silently inert despite `instanceKey` being
1551
- set — if:** you're using a custom `Transport` implementation that never
1692
+ Real Streamable HTTP transports built against **MCP spec 2025-11-25 or
1693
+ earlier** give you (2) automatically — the SDK threads a real client
1694
+ session id through `extra.sessionId` on every request regardless of
1695
+ whether the `Server` object handling it was just constructed, so the
1696
+ common "stateless Streamable HTTP" case works with `instanceKey` alone,
1697
+ no extra effort, **for every spec version through 2025-11-25**. That
1698
+ version qualifier is load-bearing, not throat-clearing — see the
1699
+ subsection immediately below. **You will NOT get (2) for free — and
1700
+ thrash detection will stay silently inert despite `instanceKey` being set
1701
+ — if:** you're using a custom `Transport` implementation that never
1552
1702
  exposes a `sessionId`, you've set `assumeSingleSession: true` (which
1553
- exists specifically to opt into the generated fallback), or anything else
1554
- lands on the fallback path described in "Session id resolution" above. If
1555
- you're in any of those cases, you need your own mechanism for threading a
1556
- stable, real session identity into each call — `instanceKey` cannot
1557
- manufacture one for you, and there is no configuration of it that will.
1703
+ exists specifically to opt into the generated fallback), anything else
1704
+ lands on the fallback path described in "Session id resolution" above,
1705
+ **or your transport is built against MCP spec 2026-07-28 or later — see
1706
+ below, this is a different problem `instanceKey` cannot solve at all,
1707
+ not a configuration gap.** If you're in one of the first three cases, you
1708
+ need your own mechanism for threading a stable, real session identity
1709
+ into each call — `instanceKey` cannot manufacture one for you, and there
1710
+ is no configuration of it that will.
1558
1711
 
1559
1712
  This composition requirement is specific to Agent Thrash Detection's
1560
1713
  per-session lookup key. Schema drift detection and the `ToolOutcome`
@@ -1563,6 +1716,70 @@ sufficient for both. Budget tracking's `perToolUsd` scope is also
1563
1716
  session-independent; its `perSessionUsd` scope inherits the identical
1564
1717
  requirement, for the identical reason.
1565
1718
 
1719
+ ### MCP spec 2026-07-28 / `@modelcontextprotocol/server` (v2) — session identity and `instanceKey`
1720
+
1721
+ **Correcting an earlier version of this section:** MCP spec 2026-07-28
1722
+ does NOT remove session identity from this library's reach entirely —
1723
+ `@modelcontextprotocol/server` (v2)'s `ctx.sessionId` is still a real,
1724
+ optional field this library reads (ADR 015 Finding 3).
1725
+
1726
+ [MCP spec 2026-07-28](https://modelcontextprotocol.io/specification/2026-07-28/changelog)
1727
+ removes protocol-level sessions and the `Mcp-Session-Id` header from the
1728
+ Streamable HTTP transport's WIRE format — not deprecated, removed — along
1729
+ with the `initialize`/`notifications/initialized` handshake that used to
1730
+ mint one. That's a real, confirmed spec fact. But it's a statement about
1731
+ the wire protocol, not about the SDK's object model: v2's `ctx.sessionId`
1732
+ (the direct equivalent of v1's `extra.sessionId`) is still there in the
1733
+ type, and this library reads it the same way for both SDKs. The practical
1734
+ effect is that `createMcpHandler`'s default, stateless deployment shape
1735
+ usually just doesn't populate it — the same "no session id" shape stdio
1736
+ has always had for v1, not a new, structurally-impossible-to-meet
1737
+ requirement. A v2 deployment that does provide a real session id (a
1738
+ custom transport, or a future v2 transport with its own session concept)
1739
+ gets ordinary, working thrash detection, no different from v1.
1740
+
1741
+ **For the common case where v2 genuinely provides no real session id**
1742
+ (the `createMcpHandler` stateless default): as of this release, two
1743
+ fixes together make thrash detection work correctly here, the same way
1744
+ it does for v1's equivalent stdio case.
1745
+
1746
+ 1. `isSingleConnectionTransport()` no longer auto-detects
1747
+ `PerRequestHTTPServerTransport` (the transport `createMcpHandler`
1748
+ builds internally) as single-connection at all — see "Known
1749
+ limitations" above and `docs/known-gaps.md` entry 8. So the fallback
1750
+ path below is no longer reached automatically for a typical v2 HTTP
1751
+ deployment; it requires an explicit, informed `assumeSingleSession: true`
1752
+ opt-in (or a genuinely single-connection v2 `serveStdio` deployment,
1753
+ which is positively confirmed the same way stdio always has been).
1754
+ 2. **For whichever deployments do legitimately reach the fallback path**
1755
+ (that explicit opt-in, or v2 stdio): the generated fallback session id
1756
+ is now shared across repeated `instrumentMcpServer()` calls when
1757
+ `instanceKey` is set — the same registry the four ADR-012 trackers
1758
+ already use, one more namespaced entry, no new bound/eviction policy.
1759
+ Previously this id was regenerated fresh on every call regardless of
1760
+ `instanceKey` (per "instanceKey alone does not fix thrash detection"
1761
+ above), so v2's factory pattern (a fresh instance, and a fresh
1762
+ `instrumentMcpServer()` call, per request) meant nothing ever
1763
+ accumulated even for an intentional, correctly-configured
1764
+ single-connection deployment. Fixed: `test/integration/thrash-v2-transport-detection.test.js`
1765
+ drives 5 separate `instrumentMcpServer()` calls sharing one
1766
+ `instanceKey`, each a fresh v2 `Server` with no real session id, and
1767
+ confirms `mcp.tool.loop.detected` now fires by the 5th.
1768
+
1769
+ **One narrower thing still open**, tracked in `docs/known-gaps.md` entry
1770
+ 6's own update: `thrashSessionState` — the flag tracking "has this server
1771
+ ever proven itself session-aware" (used by `resolveThrashSessionId()`'s
1772
+ rule 2, which skips a later no-session call rather than merging it, once
1773
+ a server has shown it hands out real ids) — is still a plain per-call
1774
+ object, not registry-backed. This only matters for a deployment mixing
1775
+ real-session-id calls with occasional no-session-id ones under a shared
1776
+ `instanceKey`; rule 1 (a real session id always wins) is unaffected
1777
+ either way.
1778
+
1779
+ **This package's actual v2 support status:** `@modelcontextprotocol/server`
1780
+ is a supported, optional peer dependency as of v0.10.0 (ADR 015) — see
1781
+ the "MCP v2 support" section above.
1782
+
1566
1783
  ### Registry bounds, and what eviction means
1567
1784
 
1568
1785
  The internal registry `instanceKey` looks trackers up in is bounded, not
@@ -1610,6 +1827,79 @@ via the `OTEL_MCP_INSTANCE_KEY` environment variable (lower precedence
1610
1827
  than the option itself). An empty or whitespace-only value from either
1611
1828
  source is treated the same as omitting it entirely.
1612
1829
 
1830
+ ## Fleet-wide fingerprint frequency (a Tempo recipe, not a library feature)
1831
+
1832
+ Under stateless MCP — no real session id at all, the exact gap the
1833
+ `instanceKey` section above documents in full — Agent Thrash Detection's
1834
+ in-process tracker has nothing to key episodes on, and stays silently
1835
+ inert. That doesn't mean there's nothing to query downstream:
1836
+ `mcp.failure.fingerprint` (see "Failure Fingerprinting" above) is already
1837
+ a plain span attribute on every failed `tools/call` span whenever
1838
+ fingerprinting is enabled, unconditional on session id or any in-process
1839
+ accumulation. A trace backend that can aggregate by an arbitrary span
1840
+ attribute can group by it directly, today, with zero new opentel-mcp
1841
+ emission. Full investigation: ADR 012's third update
1842
+ (`docs/adr/012-tracker-lifecycle-and-shared-state.md`).
1843
+
1844
+ **⚠️ Read this before treating the query below as "thrash detection for
1845
+ stateless deployments" — it is not that.** Agent Thrash Detection's whole
1846
+ premise is *one agent* retrying *one broken call*. Grouping by
1847
+ `mcp.failure.fingerprint` alone has no dimension to separate callers by —
1848
+ it counts occurrences of a bug, full stop. **One agent hitting the same
1849
+ bug 3 times in a row, and three unrelated users each hitting it once, are
1850
+ indistinguishable by this query — both produce a count of 3 for the same
1851
+ fingerprint.** That's a real, useful signal (fleet-wide bug-frequency
1852
+ monitoring) but it is a different question from "is this agent stuck in a
1853
+ loop," and presenting it as the latter would be dishonest. There is
1854
+ currently no supported way to add a caller-identity dimension to this
1855
+ query — see ADR 012's third update for why that would require the library
1856
+ to read an application-chosen tool argument (the MCP spec's own answer
1857
+ for session continuity under 2026-07-28 is a server-minted handle passed
1858
+ back as an ordinary tool argument) and put it on the span, which is a new
1859
+ kind of attribute this library doesn't emit today and a separate design
1860
+ decision, not a small addition.
1861
+
1862
+ ### Why Tempo, and not the other two backends ADR 013 investigated
1863
+
1864
+ ADR 013 investigated attribute/event queryability across SigNoz, Tempo,
1865
+ and Jaeger, but for a different need (rendering panels over single spans
1866
+ or traces). This recipe needs a capability that investigation didn't
1867
+ cover: aggregating — counting, grouping — across many separate traces by
1868
+ an attribute value, not just filtering by one.
1869
+
1870
+ - **Tempo** — TraceQL metrics queries support `by(<attribute>)` grouping
1871
+ over arbitrary span attributes at query time, including high-cardinality
1872
+ ones like `mcp.failure.fingerprint`. The query below is real, pasteable
1873
+ TraceQL, not pseudo-syntax.
1874
+ - **SigNoz** — can express the equivalent, but only as a ClickHouse SQL
1875
+ panel inside a Dashboard, per ADR 013's finding that raw ClickHouse
1876
+ querying is Dashboard-only, not reachable from the ad-hoc query API a
1877
+ live "what's failing right now" view would use.
1878
+ - **Jaeger** — cannot. Its documented `api_v3.QueryService` (`FindTraces`)
1879
+ has no aggregation or grouping parameter of any kind — confirmed
1880
+ against the proto directly, per ADR 013.
1881
+
1882
+ ### A working Tempo query
1883
+
1884
+ ```traceql
1885
+ { span.mcp.failure.fingerprint != "" && status = error }
1886
+ | count_over_time() by (span.mcp.failure.fingerprint)
1887
+ ```
1888
+
1889
+ Run this as a TraceQL metrics query (Grafana Explore, Tempo data source)
1890
+ over your evaluation window. To alert on it, wrap it in a Grafana alert
1891
+ rule with a threshold condition — e.g. fire when any series' value is
1892
+ `>= 3` over a 5-minute window. Tune the threshold and window to your own
1893
+ traffic; `3` matches this library's own default thrash
1894
+ `thrashDetection.threshold` purely for familiarity, not because it's
1895
+ derived for fleet-wide monitoring.
1896
+
1897
+ **Narrow this to one tool or one deployment by adding more attributes to
1898
+ the filter** — e.g. `&& span.gen_ai.tool.name = "lookup_customer"` — the
1899
+ same way you'd narrow any other TraceQL query. This doesn't change the
1900
+ identity-dimension caveat above: narrowing by tool name still can't tell
1901
+ one looping agent apart from several unrelated callers of that same tool.
1902
+
1613
1903
  ## Two modes
1614
1904
 
1615
1905
  ### Quick dev setup
@@ -1662,10 +1952,15 @@ pragmatic choice rather than a spec-pure one.
1662
1952
  - Node.js 20+
1663
1953
  - Windows, macOS, Linux (CI matrix tested)
1664
1954
  - Pure JavaScript, zero native dependencies
1665
- - Supports both low-level `Server` and high-level `McpServer` APIs
1666
- - @modelcontextprotocol/sdk ^1.0.0
1955
+ - Supports both low-level `Server` and high-level `McpServer` APIs, from
1956
+ either of two SDKs — see "Both server APIs" and "MCP v2 support" above
1957
+ - @modelcontextprotocol/sdk ^1.0.0 (optional peer — v1, protocol revisions
1958
+ through 2025-11-25)
1959
+ - @modelcontextprotocol/server ^2.0.0 (optional peer — v2, protocol
1960
+ revision 2026-07-28, v0.10.0+; see "MCP v2 support" above for what's
1961
+ covered and `docs/known-gaps.md` entries 6 and 8 for what isn't yet)
1667
1962
  - @opentelemetry/api ^1.9.0
1668
- - 498 tests (`npm test`) — see `test/`
1963
+ - 748 tests, 744 passing + 4 intentionally skipped (`npm test`) — see `test/`
1669
1964
  - `npm run typecheck` (`tsc --noEmit`) type-checks the public `.d.ts`
1670
1965
  surface (`src/index.d.ts` and friends) — see CONTRIBUTING.md
1671
1966
 
@@ -1720,6 +2015,19 @@ pragmatic choice rather than a spec-pure one.
1720
2015
  a documented, pasteable OpenTelemetry Collector `tailsamplingprocessor`
1721
2016
  config that keeps expensive/budget-exceeded/thrashing traces alongside
1722
2017
  a normal probabilistic sample for everything else.
2018
+ - v0.10.0: `@modelcontextprotocol/server` (MCP v2, protocol revision
2019
+ 2026-07-28) support ✓ — see "MCP v2 support" above and ADR 015
2020
+ (`docs/adr/015-mcp-v2-support.md`). Spans, standard attributes, failure
2021
+ fingerprinting, `mcp.failure.channel`/`validation_paths` classification,
2022
+ and Agent Thrash Detection (including the transport auto-detection
2023
+ heuristic and the fallback session id under `instanceKey`) all work the
2024
+ same as v1. Also hardened `detectServerKind()` to fail loudly instead of
2025
+ silently instrumenting nothing for an unrecognized/unwrappable server
2026
+ object — a behavior change (an object that previously silently no-op'd
2027
+ now throws), separate from v2 support itself; see the CHANGELOG. One
2028
+ narrower thing still open, tracked in `docs/known-gaps.md` entry 6:
2029
+ `thrashSessionState`'s session-awareness memory isn't shared across v2's
2030
+ per-request calls yet, only the fallback id itself.
1723
2031
  - Future: failure clustering + regression detection; recovery hints;
1724
2032
  root-cause chaining across parent spans; alignment with the OTel GenAI
1725
2033
  SIG's MCP semantic conventions when published
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "opentel-mcp",
3
- "version": "0.9.0",
3
+ "version": "0.10.0",
4
4
  "description": "One-line OpenTelemetry instrumentation for Model Context Protocol (MCP) servers",
5
5
  "type": "module",
6
6
  "main": "src/index.js",
@@ -55,14 +55,24 @@
55
55
  "homepage": "https://github.com/Thirumalaiboobathi/opentel-mcp#readme",
56
56
  "peerDependencies": {
57
57
  "@modelcontextprotocol/sdk": ">=1.0.0",
58
+ "@modelcontextprotocol/server": ">=2.0.0",
58
59
  "@opentelemetry/api": "^1.9.0"
59
60
  },
61
+ "peerDependenciesMeta": {
62
+ "@modelcontextprotocol/sdk": {
63
+ "optional": true
64
+ },
65
+ "@modelcontextprotocol/server": {
66
+ "optional": true
67
+ }
68
+ },
60
69
  "dependencies": {
61
70
  "@opentelemetry/exporter-trace-otlp-http": "^0.220.0",
62
71
  "@opentelemetry/resources": "^2.9.0",
63
72
  "@opentelemetry/sdk-trace-node": "^2.9.0"
64
73
  },
65
74
  "devDependencies": {
75
+ "@modelcontextprotocol/server": "^2.0.0",
66
76
  "@opentelemetry/api": "^1.9.0",
67
77
  "@opentelemetry/sdk-metrics": "^2.9.0",
68
78
  "@opentelemetry/sdk-trace-base": "^2.9.0",