opentel-mcp 0.11.0 → 0.13.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/CHANGELOG.md CHANGED
@@ -1,5 +1,222 @@
1
1
  # Changelog
2
2
 
3
+ ## 0.13.0
4
+
5
+ Closes the two open items from `docs/known-gaps.md` entry 10 (a
6
+ raw-content audit of every attribute/span event this package emits) that
7
+ needed a design decision before a fix — raw exception content on
8
+ `recordException`/`setStatus`, and an unvalidated tool-result model field
9
+ reaching `mcp.tool.model`/`gen_ai.response.model` — plus the two
10
+ lower-risk fixes from the same entry that didn't need one. Full design:
11
+ ADR 019 (`docs/adr/019-raw-content-on-spans.md`). See the README's new
12
+ "What this library records" and "Error recording" sections for the
13
+ operator-facing consequence of each.
14
+
15
+ ### Added — `errorRecording.mode` config (ADR 019 Part 1)
16
+
17
+ - New top-level `errorRecording` option, sibling to `fingerprinting` /
18
+ `costTracking` / `thrashDetection` / `schemaDrift`: `'full'` (default,
19
+ byte-for-byte unchanged from every prior release —
20
+ `recordException(err)` + `setStatus({ message: err.message })` with the
21
+ raw error), `'normalized'` (reuses the existing
22
+ `normalizeMessage()`/`parseAndNormalizeStack()` fingerprinting pipeline
23
+ — no new scrubbing logic — to strip known-sensitive-shaped substrings
24
+ from the message and the local `cwd` prefix from the stack, without
25
+ mutating the original `err`, which both call sites still rethrow), or
26
+ `'none'` (records neither — the same `setStatus({ code: ERROR })`
27
+ no-message pattern already used for tool-level `isError: true`
28
+ failures). Applies to both thrown-exception paths, `tools/call` and
29
+ `tools/list`.
30
+ - New `OTEL_MCP_ERROR_RECORDING_MODE` env var — same
31
+ option-then-env-then-default precedence, and the same silent fallback
32
+ to the default on an unrecognized value, as every other `OTEL_MCP_*`
33
+ config.
34
+ - `error.type`/`exception.type` (both read from `err.name`) are now
35
+ capped at 128 characters unconditionally, in every mode — the same cap
36
+ `mcp.failure.error_class` uses below, reusing its exported constant
37
+ rather than a second, possibly-drifting copy.
38
+ - Default stays `'full'` for all of `0.x`; ADR 019 Part 1 records the
39
+ intent to flip it to `'normalized'` at `1.0`, not before — not decided
40
+ in this release.
41
+
42
+ ### Added — `mcp.tool.model` / `gen_ai.response.model` validation (ADR 019 Part 2)
43
+
44
+ **This is new behavior that can change what a call reports, not just a
45
+ new diagnostic.** Before this release, a tool result's `model` field
46
+ reached the span completely unvalidated, whatever it was. If you have a
47
+ provider/deployment whose model identifiers use a character outside
48
+ `[A-Za-z0-9._:/@-]`, or that (implausibly, but possibly) exceed 256
49
+ characters, upgrading will make `mcp.tool.model`/`gen_ai.response.model`
50
+ disappear from those calls' spans and `mcp.tool.pricing_status` flip from
51
+ whatever it was to `"unknown"` — even for an otherwise-legitimate, real
52
+ model id. The allowlist was deliberately built generous (see below) and
53
+ verified against known provider conventions, but it's still new, and a
54
+ one-time `diag.warn()` names when this happens (shape only, never the
55
+ value) so it's discoverable rather than a silent metric/span change.
56
+
57
+ - A tool result's declared model field is now checked against a 256-
58
+ character length cap and a `[A-Za-z0-9._:/@-]` allowlist — verified
59
+ against every `DEFAULT_PRICING` key and `normalizeModelName()`'s
60
+ documented `provider/model` input contract, deliberately generous —
61
+ before it can reach `mcp.tool.model`, `gen_ai.response.model`,
62
+ `calculateCost()`, or either metric label
63
+ (`mcp.tool.tokens.total`/`mcp.tool.cost.total`).
64
+ - A rejected value is never a silent drop: `mcp.tool.pricing_status` is
65
+ set to `"unknown"` (the same status a legitimately unrecognized model
66
+ already produces), and a one-time `diag.warn()` fires per
67
+ `instrumentMcpServer()` call, reporting shape only — length, and which
68
+ check failed — never the rejected value itself.
69
+ - `budgetTracker.recordUnpriced()` now receives the validated (possibly
70
+ `undefined`) model rather than the raw tool-result value, closing a
71
+ second leak path through its own pre-existing warning that would
72
+ otherwise have echoed a rejected value verbatim.
73
+ - `costTracking.pricing`/`pricingTable` override keys are unaffected —
74
+ operator-authored config, never subject to this gate.
75
+
76
+ ### Fixed — `mcp.failure.error_class` uncapped length, `mcp.failure.validation_paths` dynamic-key leak (known-gaps entry 10)
77
+
78
+ - `mcp.failure.error_class` is now capped at 128 characters
79
+ (`fingerprint/compose.js`'s new `MAX_ERROR_CLASS_LENGTH`) — length-
80
+ bounded only, not pattern-scrubbed; a non-string `.name` is coerced to
81
+ a string before capping rather than thrown. ADR 004's note calling the
82
+ underlying value "low-cardinality" is updated to flag that as an
83
+ assumption, not an enforced property.
84
+ - `mcp.failure.validation_paths` format 1 (the raw Zod issues array, SDK
85
+ ≤1.29.0) now redacts a non-identifier-shaped path segment — a
86
+ `z.record()` schema's runtime key — to the placeholder `<KEY>` instead
87
+ of surfacing it verbatim, matching the identifier-only gate formats 2/3
88
+ already applied. Numeric (array-index) segments are never redacted.
89
+
90
+ ## 0.12.0
91
+
92
+ Agent Thrash Detection gains a new, narrower session-identity fallback for
93
+ calls that carry no real session id but do carry a client-propagated W3C
94
+ trace context. This release also fixes two real bugs in the tail-sampling
95
+ recipe v0.8.0 introduced, and adds visibility for a silent budget-tracking
96
+ gap found during that same investigation. Full design for the fallback
97
+ tier: ADR 018 (`docs/adr/018-trace-id-as-thrash-fallback.md`).
98
+
99
+ ### Added — trace id as a thrash-detection session-id fallback (ADR 018)
100
+
101
+ - **New tier in `resolveThrashSessionId()`** (`src/instrument.js`), reached
102
+ only when no real session id has ever been observed on this server, and
103
+ evaluated before the existing generated-UUID/skip fallback: if a
104
+ `tools/call` request's span has a validly-extracted REMOTE parent — i.e.
105
+ `request.params._meta` carried a valid W3C `traceparent` that
106
+ `extractTraceContext()` (ADR 017) turned into a remote `SpanContext` —
107
+ that parent's trace id is used as the session-id candidate for
108
+ thrash-detection grouping. Gated on
109
+ `trace.getSpanContext(parentContext)?.isRemote === true`, never on a
110
+ span's own `traceId` read unconditionally — a root span's trace id is
111
+ freshly, randomly generated on every call, and reading it unconditionally
112
+ would have silently turned today's honest "skip, undetermined" into
113
+ "always produce a session id that never matches the previous call's,"
114
+ which is worse than skipping. See the ADR's "THE TRAP" for the full
115
+ argument.
116
+ - Real `extra.sessionId` still wins unconditionally in every case
117
+ (steps 1–2 of the resolution order untouched, byte for byte) — a trace
118
+ id is only ever consulted for a call that has no real session id at all.
119
+ - Does **not** set `thrashSessionState.hasSeenRealSessionId` — that flag
120
+ means "this transport hands out real session identity," a permanent
121
+ per-server fact; a trace id being present on one call is a per-call fact
122
+ about that one client's behavior, not proof about the transport.
123
+ - No new `diag.warn()` for this tier, deliberately — unlike the
124
+ generated-UUID fallback (which warns because it's this library guessing
125
+ an unproven assumption about transport topology), a trace-id candidate
126
+ is real, client-supplied data with no operator action item to flag, and
127
+ warning on every occurrence — potentially far more often than the
128
+ once-per-server UUID warning — would just train operators to ignore this
129
+ library's warnings generally.
130
+ - No changes to `ThrashDetector`/`thrash/detector.js` — its composite key
131
+ already treats `sessionId` as opaque.
132
+
133
+ **Read the constraint before assuming this closes stateless MCP's
134
+ session gap.** This tier only fires when the calling *client* chooses to
135
+ propagate trace context into `_meta.traceparent` — today, per ADR 018's
136
+ investigation, that means third-party OTel instrumentation
137
+ (`@arizeai/openinference-instrumentation-mcp` and equivalents) wrapping
138
+ **v1-based** SDK clients, not either MCP SDK's own built-in behavior, and
139
+ **no v2-targeting instrumentation was found to exist anywhere**. **This
140
+ does NOT close `docs/known-gaps.md` entry 6's structural finding** — a
141
+ v2/2026-07-28-native deployment whose client doesn't propagate
142
+ `_meta.traceparent` (the default, unconfigured case for essentially
143
+ every v2 client today) gets nothing new from this release: the exact
144
+ same `null`/skip behavior as before. See the README's "Session id
145
+ resolution" section and ADR 018's "Adoption caveat" for the full scope.
146
+
147
+ ### Added — `mcp.tool.schema_drift_detected` span attribute
148
+
149
+ - New boolean span attribute, set alongside (never instead of) the
150
+ existing `mcp.tool.schema_drift.detected` span event
151
+ (`schema-drift/emitter.js`) — the same resolution ADR 011 already
152
+ applied to thrash detection (`mcp.tool.thrash_detected`) for the
153
+ identical event-vs-attribute ambiguity: whether a Collector
154
+ `tailsamplingprocessor`'s `boolean_attribute` policy can match
155
+ span-*event* data (as opposed to top-level span attributes) could not be
156
+ confirmed either way (Go source, not installed in this repository). Only
157
+ ever set to `true`, and only when drift was actually detected — never
158
+ explicitly set `false`.
159
+
160
+ ### Fixed — tail-sampling recipe referenced a non-existent span attribute
161
+
162
+ - The README's `tailsamplingprocessor` recipe recommended keying a
163
+ `boolean_attribute` policy on `mcp.tool.schema_drift.detected` — but
164
+ that string was only ever a span *event* name and a metric counter name
165
+ (`schema-drift/attributes.js`), never passed to `span.setAttribute()`
166
+ anywhere in this package. As documented, that policy could never have
167
+ matched anything. Fixed by the new `mcp.tool.schema_drift_detected`
168
+ attribute above.
169
+ - Also added a `string_attribute` policy on `mcp.tool.pricing_status =
170
+ "unknown"` — v0.11.0's "confidently wrong zero" problem reappearing at
171
+ the sampling layer: an unpriced call has real extracted token usage but
172
+ no `mcp.tool.cost.usd`, indistinguishable from a genuinely free call to
173
+ a numeric-threshold policy.
174
+ - Recipe YAML extracted to `docs/recipes/tail-sampling.yaml` (repo-only,
175
+ not published — the same carve-out `dashboards/` already has), with a
176
+ new cross-check test (`test/recipes/tail-sampling-attributes.test.js`)
177
+ asserting every attribute a policy references is a real, exported
178
+ constant AND actually passed to `span.setAttribute()` — the check that
179
+ would have caught this bug automatically. The README now references the
180
+ file instead of duplicating it, and states plainly that attribute
181
+ *names* are cross-checked but Collector policy *behavior* itself has not
182
+ been run end to end (no Docker/Collector available in this project's dev
183
+ environment).
184
+ - `docs/known-gaps.md` entry 9 (new): an unpriced call never reaches
185
+ `budgetTracker.recordAndCheck()`, so budget guardrails cannot trip on
186
+ unpriced spend regardless of amount. The visibility half is fixed in
187
+ this same release — see "Added" below — but the underlying accounting
188
+ behavior is not; what should happen to an unpriced call's budget
189
+ accounting is a real design question, deliberately left open.
190
+
191
+ ### Added — visibility for unpriced spend against a configured budget (known-gaps entry 9)
192
+
193
+ - **Two new one-time `diag.warn()` diagnostics in `createBudgetTracker()`**
194
+ (`src/cost/budget.js`), no behavior change and no new public surface:
195
+ one fires at construction whenever `perSessionUsd`/`perToolUsd` is
196
+ configured at all, stating plainly that unpriced calls won't count
197
+ toward it; the other — a new `recordUnpriced(model)` method, called from
198
+ `applyCostAttribution()`'s existing `costUsd === null` branch
199
+ (`src/instrument.js`), the same branch that already sets
200
+ `mcp.tool.pricing_status: "unknown"` — fires the first time an unpriced
201
+ call under an active budget is actually observed, naming the model and
202
+ which scope(s) are configured. Both no-op when no budget is configured;
203
+ neither changes `BudgetCheckResult`'s shape or adds a span attribute.
204
+ - **Deliberately diagnostics only — no fallback pricing was added.**
205
+ Making an unpriced call actually count toward a USD budget means
206
+ inventing a number for it, and a wrong invented price is a *different*
207
+ confidently-wrong number, not a fix — the same disease this warning
208
+ exists to flag, one layer up. See `docs/known-gaps.md` entry 9's
209
+ "Status update (v0.12.0)" for the full argument against building that
210
+ now, and why it stays open as a future, ADR-gated decision rather than
211
+ folded into this patch.
212
+ - Both warnings share the budget tracker's own existing
213
+ once-per-tracker-instance granularity (`createBudgetTracker()`'s own
214
+ docblock) — under the default (no `instanceKey`), a fresh-server-per-request
215
+ deployment re-warns on every request for both, the same inherited-caveat
216
+ shape `docs/known-gaps.md` entry 6 documents for the thrash fallback
217
+ warning; a stable `instanceKey` shares one tracker, and one already-armed
218
+ warning, across calls, same as every other registry-backed tracker.
219
+
3
220
  ## 0.11.0
4
221
 
5
222
  **⚠️ Type change, not a runtime behavior change — read this first.**