opentel-mcp 0.10.0 → 0.12.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/CHANGELOG.md CHANGED
@@ -1,5 +1,272 @@
1
1
  # Changelog
2
2
 
3
+ ## 0.12.0
4
+
5
+ Agent Thrash Detection gains a new, narrower session-identity fallback for
6
+ calls that carry no real session id but do carry a client-propagated W3C
7
+ trace context. This release also fixes two real bugs in the tail-sampling
8
+ recipe v0.8.0 introduced, and adds visibility for a silent budget-tracking
9
+ gap found during that same investigation. Full design for the fallback
10
+ tier: ADR 018 (`docs/adr/018-trace-id-as-thrash-fallback.md`).
11
+
12
+ ### Added — trace id as a thrash-detection session-id fallback (ADR 018)
13
+
14
+ - **New tier in `resolveThrashSessionId()`** (`src/instrument.js`), reached
15
+ only when no real session id has ever been observed on this server, and
16
+ evaluated before the existing generated-UUID/skip fallback: if a
17
+ `tools/call` request's span has a validly-extracted REMOTE parent — i.e.
18
+ `request.params._meta` carried a valid W3C `traceparent` that
19
+ `extractTraceContext()` (ADR 017) turned into a remote `SpanContext` —
20
+ that parent's trace id is used as the session-id candidate for
21
+ thrash-detection grouping. Gated on
22
+ `trace.getSpanContext(parentContext)?.isRemote === true`, never on a
23
+ span's own `traceId` read unconditionally — a root span's trace id is
24
+ freshly, randomly generated on every call, and reading it unconditionally
25
+ would have silently turned today's honest "skip, undetermined" into
26
+ "always produce a session id that never matches the previous call's,"
27
+ which is worse than skipping. See the ADR's "THE TRAP" for the full
28
+ argument.
29
+ - Real `extra.sessionId` still wins unconditionally in every case
30
+ (steps 1–2 of the resolution order untouched, byte for byte) — a trace
31
+ id is only ever consulted for a call that has no real session id at all.
32
+ - Does **not** set `thrashSessionState.hasSeenRealSessionId` — that flag
33
+ means "this transport hands out real session identity," a permanent
34
+ per-server fact; a trace id being present on one call is a per-call fact
35
+ about that one client's behavior, not proof about the transport.
36
+ - No new `diag.warn()` for this tier, deliberately — unlike the
37
+ generated-UUID fallback (which warns because it's this library guessing
38
+ an unproven assumption about transport topology), a trace-id candidate
39
+ is real, client-supplied data with no operator action item to flag, and
40
+ warning on every occurrence — potentially far more often than the
41
+ once-per-server UUID warning — would just train operators to ignore this
42
+ library's warnings generally.
43
+ - No changes to `ThrashDetector`/`thrash/detector.js` — its composite key
44
+ already treats `sessionId` as opaque.
45
+
46
+ **Read the constraint before assuming this closes stateless MCP's
47
+ session gap.** This tier only fires when the calling *client* chooses to
48
+ propagate trace context into `_meta.traceparent` — today, per ADR 018's
49
+ investigation, that means third-party OTel instrumentation
50
+ (`@arizeai/openinference-instrumentation-mcp` and equivalents) wrapping
51
+ **v1-based** SDK clients, not either MCP SDK's own built-in behavior, and
52
+ **no v2-targeting instrumentation was found to exist anywhere**. **This
53
+ does NOT close `docs/known-gaps.md` entry 6's structural finding** — a
54
+ v2/2026-07-28-native deployment whose client doesn't propagate
55
+ `_meta.traceparent` (the default, unconfigured case for essentially
56
+ every v2 client today) gets nothing new from this release: the exact
57
+ same `null`/skip behavior as before. See the README's "Session id
58
+ resolution" section and ADR 018's "Adoption caveat" for the full scope.
59
+
60
+ ### Added — `mcp.tool.schema_drift_detected` span attribute
61
+
62
+ - New boolean span attribute, set alongside (never instead of) the
63
+ existing `mcp.tool.schema_drift.detected` span event
64
+ (`schema-drift/emitter.js`) — the same resolution ADR 011 already
65
+ applied to thrash detection (`mcp.tool.thrash_detected`) for the
66
+ identical event-vs-attribute ambiguity: whether a Collector
67
+ `tailsamplingprocessor`'s `boolean_attribute` policy can match
68
+ span-*event* data (as opposed to top-level span attributes) could not be
69
+ confirmed either way (Go source, not installed in this repository). Only
70
+ ever set to `true`, and only when drift was actually detected — never
71
+ explicitly set `false`.
72
+
73
+ ### Fixed — tail-sampling recipe referenced a non-existent span attribute
74
+
75
+ - The README's `tailsamplingprocessor` recipe recommended keying a
76
+ `boolean_attribute` policy on `mcp.tool.schema_drift.detected` — but
77
+ that string was only ever a span *event* name and a metric counter name
78
+ (`schema-drift/attributes.js`), never passed to `span.setAttribute()`
79
+ anywhere in this package. As documented, that policy could never have
80
+ matched anything. Fixed by the new `mcp.tool.schema_drift_detected`
81
+ attribute above.
82
+ - Also added a `string_attribute` policy on `mcp.tool.pricing_status =
83
+ "unknown"` — v0.11.0's "confidently wrong zero" problem reappearing at
84
+ the sampling layer: an unpriced call has real extracted token usage but
85
+ no `mcp.tool.cost.usd`, indistinguishable from a genuinely free call to
86
+ a numeric-threshold policy.
87
+ - Recipe YAML extracted to `docs/recipes/tail-sampling.yaml` (repo-only,
88
+ not published — the same carve-out `dashboards/` already has), with a
89
+ new cross-check test (`test/recipes/tail-sampling-attributes.test.js`)
90
+ asserting every attribute a policy references is a real, exported
91
+ constant AND actually passed to `span.setAttribute()` — the check that
92
+ would have caught this bug automatically. The README now references the
93
+ file instead of duplicating it, and states plainly that attribute
94
+ *names* are cross-checked but Collector policy *behavior* itself has not
95
+ been run end to end (no Docker/Collector available in this project's dev
96
+ environment).
97
+ - `docs/known-gaps.md` entry 9 (new): an unpriced call never reaches
98
+ `budgetTracker.recordAndCheck()`, so budget guardrails cannot trip on
99
+ unpriced spend regardless of amount. The visibility half is fixed in
100
+ this same release — see "Added" below — but the underlying accounting
101
+ behavior is not; what should happen to an unpriced call's budget
102
+ accounting is a real design question, deliberately left open.
103
+
104
+ ### Added — visibility for unpriced spend against a configured budget (known-gaps entry 9)
105
+
106
+ - **Two new one-time `diag.warn()` diagnostics in `createBudgetTracker()`**
107
+ (`src/cost/budget.js`), no behavior change and no new public surface:
108
+ one fires at construction whenever `perSessionUsd`/`perToolUsd` is
109
+ configured at all, stating plainly that unpriced calls won't count
110
+ toward it; the other — a new `recordUnpriced(model)` method, called from
111
+ `applyCostAttribution()`'s existing `costUsd === null` branch
112
+ (`src/instrument.js`), the same branch that already sets
113
+ `mcp.tool.pricing_status: "unknown"` — fires the first time an unpriced
114
+ call under an active budget is actually observed, naming the model and
115
+ which scope(s) are configured. Both no-op when no budget is configured;
116
+ neither changes `BudgetCheckResult`'s shape or adds a span attribute.
117
+ - **Deliberately diagnostics only — no fallback pricing was added.**
118
+ Making an unpriced call actually count toward a USD budget means
119
+ inventing a number for it, and a wrong invented price is a *different*
120
+ confidently-wrong number, not a fix — the same disease this warning
121
+ exists to flag, one layer up. See `docs/known-gaps.md` entry 9's
122
+ "Status update (v0.12.0)" for the full argument against building that
123
+ now, and why it stays open as a future, ADR-gated decision rather than
124
+ folded into this patch.
125
+ - Both warnings share the budget tracker's own existing
126
+ once-per-tracker-instance granularity (`createBudgetTracker()`'s own
127
+ docblock) — under the default (no `instanceKey`), a fresh-server-per-request
128
+ deployment re-warns on every request for both, the same inherited-caveat
129
+ shape `docs/known-gaps.md` entry 6 documents for the thrash fallback
130
+ warning; a stable `instanceKey` shares one tracker, and one already-armed
131
+ warning, across calls, same as every other registry-backed tracker.
132
+
133
+ ## 0.11.0
134
+
135
+ **⚠️ Type change, not a runtime behavior change — read this first.**
136
+ `ModelPricing` is now a discriminated union
137
+ (`{pricingKind: 'chat', inputPer1M, outputPer1M, currency} |
138
+ {pricingKind: 'embedding', inputPer1M, currency}`) instead of a single
139
+ shape with both token fields always present. A TypeScript consumer with
140
+ an existing custom `pricingTable`/`pricing` object typed against the old
141
+ shape will see a compile error requiring `pricingKind` on each entry.
142
+ **Runtime behavior for those same objects is unchanged**: `calculateCost()`
143
+ treats a missing or unrecognized `pricingKind` as `'chat'`, exactly the
144
+ behavior every pre-v0.11.0 entry already had. See ADR 016
145
+ (`docs/adr/016-pricing-override-and-staleness.md`) point 1.
146
+
147
+ ### Added — Pricing table override and staleness signalling (Phase 1 of 2)
148
+
149
+ `DEFAULT_PRICING` had the same disease this whole library exists to fix
150
+ elsewhere: a hardcoded snapshot with no signal when it's wrong or stale.
151
+ This release closes that. Full design: ADR 016
152
+ (`docs/adr/016-pricing-override-and-staleness.md`).
153
+
154
+ - **Embedding model support.** `DEFAULT_PRICING` gains OpenAI
155
+ `text-embedding-3-small`/`text-embedding-3-large`/`text-embedding-ada-002`,
156
+ Cohere `cohere-embed-v3`, and Bedrock `amazon-titan-embed-v2`.
157
+ Embeddings are input-token-only — rather than modeling that as
158
+ `outputPer1M: 0` (indistinguishable from a data-entry bug), a new
159
+ `pricingKind: 'chat' | 'embedding'` discriminator makes it explicit; an
160
+ `'embedding'` entry has no `outputPer1M` field at all, and
161
+ `calculateCost()` never reads `outputTokens` for one (still validated as
162
+ a non-negative finite number, just never charged for).
163
+ - **`costTracking.pricing`: per-model merge over defaults.** New option,
164
+ a *partial* pricing table merged per-model OVER `pricingTable ??
165
+ DEFAULT_PRICING` — each key you supply replaces that model's entire
166
+ pricing entry, every model you don't name is untouched. This is now the
167
+ recommended way to correct a stale price or add a model
168
+ `DEFAULT_PRICING` doesn't know about, without spreading the whole
169
+ default table by hand (the old workaround the README used to teach).
170
+ `costTracking.pricingTable` keeps its existing full-replace behavior,
171
+ unchanged, for the narrower "I want only my own models" case — see ADR
172
+ 016 point 2 for why both exist.
173
+ - **`mcp.tool.pricing_status` span + metric attribute.** One of `"known"`
174
+ | `"unknown"` | `"user_override"`, set whenever token usage was
175
+ extracted at all — even with no model detected (`"unknown"` in that
176
+ case), unlike the existing model/cost attributes. Added to the
177
+ `mcp.tool.tokens.total` / `mcp.tool.cost.total` metrics too, so a
178
+ dashboard can compute "% of tokens/spend unpriced" as a direct
179
+ aggregation instead of inferring it from missing data. Provenance-based:
180
+ `"user_override"` means the model's key came from your
181
+ `pricing`/`pricingTable`, regardless of whether the numbers you supplied
182
+ happen to match `DEFAULT_PRICING`'s own entry.
183
+ - **Staleness signalling.** `DEFAULT_PRICING_LAST_VERIFIED` (also
184
+ exported) names the table's last-checked date; once it's more than 90
185
+ days old, `instrumentMcpServer()` fires a one-time `diag.warn()`, and,
186
+ when `setupNodeSdk: true`, also attaches an
187
+ `mcp.pricing.default_table_last_verified` resource attribute. Both only
188
+ fire when `DEFAULT_PRICING` is actually contributing to the effective
189
+ table — a caller who fully replaced it via `pricingTable` isn't using
190
+ our defaults, so a warning about them would be misleading.
191
+ `isDefaultPricingStale(now?, thresholdDays?)` (also exported) is the
192
+ pure function behind the warning, for callers who want to check it
193
+ themselves.
194
+ - **Bedrock region caveat, documented not modeled.** Bedrock pricing
195
+ varies by region; `DEFAULT_PRICING`'s Bedrock entries (Nova, and the new
196
+ Titan embedding entry) assume us-east-1 list price and the table is not
197
+ region-keyed — no tool-result usage shape this package recognizes
198
+ carries a region signal to key a lookup on. Documented loudly in the
199
+ README and in `pricing.js`; override via `costTracking.pricing` for a
200
+ different region. See ADR 016 point 5.
201
+ - `calculateCost()` gained defensive validation for malformed pricing
202
+ entries (missing/negative/non-numeric `inputPer1M`/`outputPer1M`),
203
+ since `pricing`/`pricingTable` now make it reachable with
204
+ caller-supplied shapes it previously never had to distrust — degrades
205
+ to `null`, same as an unknown model, never throws.
206
+
207
+ ### Added — W3C Trace Context propagation over MCP `_meta` (Phase 2 of 2, server-side only)
208
+
209
+ The most-complained-about gap in agent observability: an agent's own
210
+ trace (LangGraph or otherwise) and the MCP server's trace for the tool
211
+ call it made were always two disconnected traces, with no edge between
212
+ them. Full design, including the sampling and conflicting-context
213
+ decisions below: ADR 017 (`docs/adr/017-trace-context-propagation.md`).
214
+
215
+ - **`tools/call` requests carrying a valid W3C `traceparent` in
216
+ `params._meta` now become a child of the calling agent's own span**,
217
+ joining what were two disconnected traces into one — under both MCP v1
218
+ and v2, with zero configuration and no new option. `tracestate` is
219
+ propagated too, when present. Works for any client already emitting
220
+ `traceparent` via a standard OTel SDK's `propagation.inject()` in any
221
+ language — this isn't Node/JS-specific on the client side, only on
222
+ which side of the wire this release implements.
223
+ - **The upstream sampling decision is honored automatically, by
224
+ construction, with no sampling logic written for this feature**: the
225
+ extracted `SpanContext` is marked `isRemote: true` with the real parsed
226
+ `traceFlags`, which is exactly what the SDK's own default
227
+ `ParentBasedSampler` already keys its remote-parent decision off of. A
228
+ not-sampled upstream `traceparent` means this tool-call span is not
229
+ recorded or exported, matching the calling agent's own choice — see the
230
+ ADR's "Sampling" section for why forcing sampling regardless was
231
+ considered and rejected.
232
+ - **A `_meta`-extracted context always replaces, never merges with, an
233
+ already-active local context** (e.g. an ambient HTTP-server span from
234
+ auto-instrumentation on a Streamable HTTP transport) — the message-level
235
+ `_meta` context is the semantically correct parent for one tool call,
236
+ full stop, regardless of what transport-level span it happened to
237
+ arrive inside. See the ADR's "Conflicting `_meta.traceparent`" section.
238
+ - **Absent, malformed, or unparseable `_meta`/`traceparent` produces
239
+ behavior that is byte-identical to pre-v0.11.0** — not merely
240
+ equivalent to it: confirmed by reading both `NoopTracer` and the real
241
+ SDK `Tracer`'s own `startActiveSpan()` fallback (`ctx ?? context.active()`),
242
+ which is exactly what this feature's `extractTraceContext()` returns
243
+ for every case that isn't a valid `traceparent`. No `diag.warn()` for
244
+ the common "client doesn't send `_meta.traceparent`" case — see the
245
+ ADR's "no warn spam" constraint.
246
+ - **Zero new dependencies.** `@opentelemetry/core`'s
247
+ `W3CTraceContextPropagator` was the obvious reference implementation
248
+ and was deliberately not taken as a dependency, per
249
+ `CONTRIBUTING.md`'s "no new dependencies without discussion first" —
250
+ everything needed except the traceparent regex itself (~10 lines,
251
+ matching `@opentelemetry/core`'s own validation field-for-field) was
252
+ already available from `@opentelemetry/api`, already a peer dependency
253
+ — including `createTraceState()`, a fully spec-validated `tracestate`
254
+ parser. Full reasoning: ADR 017's "No new dependency" section.
255
+ - **Server-side extraction only.** The client-side shim that would let a
256
+ Node/Python agent framework *set* `_meta.traceparent` on outgoing calls
257
+ is explicitly out of scope for this phase — extraction is independently
258
+ useful today, for free, to any client whose own tooling already sets
259
+ `_meta` in this shape. Tracked as future work, not implied as solved.
260
+
261
+ Also fixed in this release: two places (`index.d.ts`'s
262
+ `instrumentMcpServer()` docblock, and this file's own v0.10.0 entry
263
+ below) still described `docs/known-gaps.md` entries 6/7/8 using language
264
+ that read as still-open, or as scoped out of v0.10.0 — both were stale.
265
+ Entries 7 and 8 have been fully fixed since v0.10.0 with no open caveats;
266
+ entry 6's fallback-session-id half is fixed too, narrowed to a smaller,
267
+ genuinely-still-open remainder (see the corrected v0.10.0 entry below and
268
+ `index.d.ts`'s updated docblock for the accurate, current accounting).
269
+
3
270
  ## 0.10.0
4
271
 
5
272
  **⚠️ Behavior change, unrelated to the feature below — read this first.**
@@ -60,15 +327,32 @@ new was added for this, since ADR 012's original design already covers
60
327
  this exact deployment shape, v2 just makes it the default instead of an
61
328
  edge case.
62
329
 
63
- **Two gaps not closed this release, both tracked in `docs/known-gaps.md`
64
- with a "Status update (v0.10.0)" note:** Agent Thrash Detection's fallback
65
- session id still doesn't survive v2's per-request factory pattern even
66
- with `instanceKey` set (entry 6 — real session ids work fine either way),
67
- and `isSingleConnectionTransport()`'s transport-detection heuristic still
68
- misclassifies the transport `createMcpHandler` builds internally (entry
69
- 8, live as of this release, not merely forward-looking). Both were
70
- explicitly scoped out of this round, not overlooked; workaround for
71
- either: `thrashDetection: { enabled: false }`.
330
+ **Correction (recorded here rather than silently edited): both gaps below
331
+ were actually closed in this same v0.10.0 release, not left open.** The
332
+ paragraph originally here said entries 6 and 8 (`docs/known-gaps.md`)
333
+ were scoped out of this round — true of the round that produced the text
334
+ above, not of what actually shipped. A follow-up investigation, completed
335
+ before v0.10.0 was cut, folded both fixes back in: `isSingleConnectionTransport()`
336
+ no longer misclassifies the transport `createMcpHandler` builds internally
337
+ (entry 8 — fully fixed: v2 now requires positive confirmation,
338
+ `transport.constructor.name === 'StdioServerTransport'`, instead of
339
+ inferring single-connection from an absent `sessionId` property), and
340
+ Agent Thrash Detection's fallback session id is now registry-backed via
341
+ `instanceKey` (entry 6's fallback-id gap — fixed: repeated
342
+ `instrumentMcpServer()` calls sharing an `instanceKey` now reuse the same
343
+ generated id instead of a fresh one per call). Both fixes shipped in the
344
+ same commit, in a specific order — fixing detection (entry 8) before
345
+ sharing the fallback id (entry 6) — since sharing it first would have made
346
+ entry 8's false positive worse, not better.
347
+
348
+ **What remains genuinely open, narrower than either original gap:**
349
+ `thrashSessionState` (whether a server has ever proven itself
350
+ session-aware) still isn't registry-backed, and — structurally, not a
351
+ bug this library can fix — MCP spec 2026-07-28 removes protocol-level
352
+ sessions entirely, so no configuration of this library can produce a
353
+ *real* session id for a spec-2026-07-28-native deployment in the first
354
+ place. See ADR 015's final "Update ... Findings 3 and 8 landed here too"
355
+ section and `docs/known-gaps.md` entries 6 and 8 for the full accounting.
72
356
 
73
357
  ## 0.9.0
74
358