opentel-mcp 0.11.0 → 0.13.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +217 -0
- package/README.md +316 -87
- package/package.json +2 -1
- package/src/config.js +15 -0
- package/src/cost/budget.js +90 -3
- package/src/cost/calculator.js +66 -0
- package/src/error-recording/config.js +74 -0
- package/src/error-recording/types.d.ts +55 -0
- package/src/fingerprint/attributes.js +16 -0
- package/src/fingerprint/classify/validation-paths.js +55 -1
- package/src/fingerprint/compose.js +27 -1
- package/src/fingerprint/types.d.ts +1 -1
- package/src/index.d.ts +31 -0
- package/src/instrument.js +322 -29
- package/src/schema-drift/attributes.js +32 -0
- package/src/schema-drift/emitter.js +11 -3
package/CHANGELOG.md
CHANGED
|
@@ -1,5 +1,222 @@
|
|
|
1
1
|
# Changelog
|
|
2
2
|
|
|
3
|
+
## 0.13.0
|
|
4
|
+
|
|
5
|
+
Closes the two open items from `docs/known-gaps.md` entry 10 (a
|
|
6
|
+
raw-content audit of every attribute/span event this package emits) that
|
|
7
|
+
needed a design decision before a fix — raw exception content on
|
|
8
|
+
`recordException`/`setStatus`, and an unvalidated tool-result model field
|
|
9
|
+
reaching `mcp.tool.model`/`gen_ai.response.model` — plus the two
|
|
10
|
+
lower-risk fixes from the same entry that didn't need one. Full design:
|
|
11
|
+
ADR 019 (`docs/adr/019-raw-content-on-spans.md`). See the README's new
|
|
12
|
+
"What this library records" and "Error recording" sections for the
|
|
13
|
+
operator-facing consequence of each.
|
|
14
|
+
|
|
15
|
+
### Added — `errorRecording.mode` config (ADR 019 Part 1)
|
|
16
|
+
|
|
17
|
+
- New top-level `errorRecording` option, sibling to `fingerprinting` /
|
|
18
|
+
`costTracking` / `thrashDetection` / `schemaDrift`: `'full'` (default,
|
|
19
|
+
byte-for-byte unchanged from every prior release —
|
|
20
|
+
`recordException(err)` + `setStatus({ message: err.message })` with the
|
|
21
|
+
raw error), `'normalized'` (reuses the existing
|
|
22
|
+
`normalizeMessage()`/`parseAndNormalizeStack()` fingerprinting pipeline
|
|
23
|
+
— no new scrubbing logic — to strip known-sensitive-shaped substrings
|
|
24
|
+
from the message and the local `cwd` prefix from the stack, without
|
|
25
|
+
mutating the original `err`, which both call sites still rethrow), or
|
|
26
|
+
`'none'` (records neither — the same `setStatus({ code: ERROR })`
|
|
27
|
+
no-message pattern already used for tool-level `isError: true`
|
|
28
|
+
failures). Applies to both thrown-exception paths, `tools/call` and
|
|
29
|
+
`tools/list`.
|
|
30
|
+
- New `OTEL_MCP_ERROR_RECORDING_MODE` env var — same
|
|
31
|
+
option-then-env-then-default precedence, and the same silent fallback
|
|
32
|
+
to the default on an unrecognized value, as every other `OTEL_MCP_*`
|
|
33
|
+
config.
|
|
34
|
+
- `error.type`/`exception.type` (both read from `err.name`) are now
|
|
35
|
+
capped at 128 characters unconditionally, in every mode — the same cap
|
|
36
|
+
`mcp.failure.error_class` uses below, reusing its exported constant
|
|
37
|
+
rather than a second, possibly-drifting copy.
|
|
38
|
+
- Default stays `'full'` for all of `0.x`; ADR 019 Part 1 records the
|
|
39
|
+
intent to flip it to `'normalized'` at `1.0`, not before — not decided
|
|
40
|
+
in this release.
|
|
41
|
+
|
|
42
|
+
### Added — `mcp.tool.model` / `gen_ai.response.model` validation (ADR 019 Part 2)
|
|
43
|
+
|
|
44
|
+
**This is new behavior that can change what a call reports, not just a
|
|
45
|
+
new diagnostic.** Before this release, a tool result's `model` field
|
|
46
|
+
reached the span completely unvalidated, whatever it was. If you have a
|
|
47
|
+
provider/deployment whose model identifiers use a character outside
|
|
48
|
+
`[A-Za-z0-9._:/@-]`, or that (implausibly, but possibly) exceed 256
|
|
49
|
+
characters, upgrading will make `mcp.tool.model`/`gen_ai.response.model`
|
|
50
|
+
disappear from those calls' spans and `mcp.tool.pricing_status` flip from
|
|
51
|
+
whatever it was to `"unknown"` — even for an otherwise-legitimate, real
|
|
52
|
+
model id. The allowlist was deliberately built generous (see below) and
|
|
53
|
+
verified against known provider conventions, but it's still new, and a
|
|
54
|
+
one-time `diag.warn()` names when this happens (shape only, never the
|
|
55
|
+
value) so it's discoverable rather than a silent metric/span change.
|
|
56
|
+
|
|
57
|
+
- A tool result's declared model field is now checked against a 256-
|
|
58
|
+
character length cap and a `[A-Za-z0-9._:/@-]` allowlist — verified
|
|
59
|
+
against every `DEFAULT_PRICING` key and `normalizeModelName()`'s
|
|
60
|
+
documented `provider/model` input contract, deliberately generous —
|
|
61
|
+
before it can reach `mcp.tool.model`, `gen_ai.response.model`,
|
|
62
|
+
`calculateCost()`, or either metric label
|
|
63
|
+
(`mcp.tool.tokens.total`/`mcp.tool.cost.total`).
|
|
64
|
+
- A rejected value is never a silent drop: `mcp.tool.pricing_status` is
|
|
65
|
+
set to `"unknown"` (the same status a legitimately unrecognized model
|
|
66
|
+
already produces), and a one-time `diag.warn()` fires per
|
|
67
|
+
`instrumentMcpServer()` call, reporting shape only — length, and which
|
|
68
|
+
check failed — never the rejected value itself.
|
|
69
|
+
- `budgetTracker.recordUnpriced()` now receives the validated (possibly
|
|
70
|
+
`undefined`) model rather than the raw tool-result value, closing a
|
|
71
|
+
second leak path through its own pre-existing warning that would
|
|
72
|
+
otherwise have echoed a rejected value verbatim.
|
|
73
|
+
- `costTracking.pricing`/`pricingTable` override keys are unaffected —
|
|
74
|
+
operator-authored config, never subject to this gate.
|
|
75
|
+
|
|
76
|
+
### Fixed — `mcp.failure.error_class` uncapped length, `mcp.failure.validation_paths` dynamic-key leak (known-gaps entry 10)
|
|
77
|
+
|
|
78
|
+
- `mcp.failure.error_class` is now capped at 128 characters
|
|
79
|
+
(`fingerprint/compose.js`'s new `MAX_ERROR_CLASS_LENGTH`) — length-
|
|
80
|
+
bounded only, not pattern-scrubbed; a non-string `.name` is coerced to
|
|
81
|
+
a string before capping rather than thrown. ADR 004's note calling the
|
|
82
|
+
underlying value "low-cardinality" is updated to flag that as an
|
|
83
|
+
assumption, not an enforced property.
|
|
84
|
+
- `mcp.failure.validation_paths` format 1 (the raw Zod issues array, SDK
|
|
85
|
+
≤1.29.0) now redacts a non-identifier-shaped path segment — a
|
|
86
|
+
`z.record()` schema's runtime key — to the placeholder `<KEY>` instead
|
|
87
|
+
of surfacing it verbatim, matching the identifier-only gate formats 2/3
|
|
88
|
+
already applied. Numeric (array-index) segments are never redacted.
|
|
89
|
+
|
|
90
|
+
## 0.12.0
|
|
91
|
+
|
|
92
|
+
Agent Thrash Detection gains a new, narrower session-identity fallback for
|
|
93
|
+
calls that carry no real session id but do carry a client-propagated W3C
|
|
94
|
+
trace context. This release also fixes two real bugs in the tail-sampling
|
|
95
|
+
recipe v0.8.0 introduced, and adds visibility for a silent budget-tracking
|
|
96
|
+
gap found during that same investigation. Full design for the fallback
|
|
97
|
+
tier: ADR 018 (`docs/adr/018-trace-id-as-thrash-fallback.md`).
|
|
98
|
+
|
|
99
|
+
### Added — trace id as a thrash-detection session-id fallback (ADR 018)
|
|
100
|
+
|
|
101
|
+
- **New tier in `resolveThrashSessionId()`** (`src/instrument.js`), reached
|
|
102
|
+
only when no real session id has ever been observed on this server, and
|
|
103
|
+
evaluated before the existing generated-UUID/skip fallback: if a
|
|
104
|
+
`tools/call` request's span has a validly-extracted REMOTE parent — i.e.
|
|
105
|
+
`request.params._meta` carried a valid W3C `traceparent` that
|
|
106
|
+
`extractTraceContext()` (ADR 017) turned into a remote `SpanContext` —
|
|
107
|
+
that parent's trace id is used as the session-id candidate for
|
|
108
|
+
thrash-detection grouping. Gated on
|
|
109
|
+
`trace.getSpanContext(parentContext)?.isRemote === true`, never on a
|
|
110
|
+
span's own `traceId` read unconditionally — a root span's trace id is
|
|
111
|
+
freshly, randomly generated on every call, and reading it unconditionally
|
|
112
|
+
would have silently turned today's honest "skip, undetermined" into
|
|
113
|
+
"always produce a session id that never matches the previous call's,"
|
|
114
|
+
which is worse than skipping. See the ADR's "THE TRAP" for the full
|
|
115
|
+
argument.
|
|
116
|
+
- Real `extra.sessionId` still wins unconditionally in every case
|
|
117
|
+
(steps 1–2 of the resolution order untouched, byte for byte) — a trace
|
|
118
|
+
id is only ever consulted for a call that has no real session id at all.
|
|
119
|
+
- Does **not** set `thrashSessionState.hasSeenRealSessionId` — that flag
|
|
120
|
+
means "this transport hands out real session identity," a permanent
|
|
121
|
+
per-server fact; a trace id being present on one call is a per-call fact
|
|
122
|
+
about that one client's behavior, not proof about the transport.
|
|
123
|
+
- No new `diag.warn()` for this tier, deliberately — unlike the
|
|
124
|
+
generated-UUID fallback (which warns because it's this library guessing
|
|
125
|
+
an unproven assumption about transport topology), a trace-id candidate
|
|
126
|
+
is real, client-supplied data with no operator action item to flag, and
|
|
127
|
+
warning on every occurrence — potentially far more often than the
|
|
128
|
+
once-per-server UUID warning — would just train operators to ignore this
|
|
129
|
+
library's warnings generally.
|
|
130
|
+
- No changes to `ThrashDetector`/`thrash/detector.js` — its composite key
|
|
131
|
+
already treats `sessionId` as opaque.
|
|
132
|
+
|
|
133
|
+
**Read the constraint before assuming this closes stateless MCP's
|
|
134
|
+
session gap.** This tier only fires when the calling *client* chooses to
|
|
135
|
+
propagate trace context into `_meta.traceparent` — today, per ADR 018's
|
|
136
|
+
investigation, that means third-party OTel instrumentation
|
|
137
|
+
(`@arizeai/openinference-instrumentation-mcp` and equivalents) wrapping
|
|
138
|
+
**v1-based** SDK clients, not either MCP SDK's own built-in behavior, and
|
|
139
|
+
**no v2-targeting instrumentation was found to exist anywhere**. **This
|
|
140
|
+
does NOT close `docs/known-gaps.md` entry 6's structural finding** — a
|
|
141
|
+
v2/2026-07-28-native deployment whose client doesn't propagate
|
|
142
|
+
`_meta.traceparent` (the default, unconfigured case for essentially
|
|
143
|
+
every v2 client today) gets nothing new from this release: the exact
|
|
144
|
+
same `null`/skip behavior as before. See the README's "Session id
|
|
145
|
+
resolution" section and ADR 018's "Adoption caveat" for the full scope.
|
|
146
|
+
|
|
147
|
+
### Added — `mcp.tool.schema_drift_detected` span attribute
|
|
148
|
+
|
|
149
|
+
- New boolean span attribute, set alongside (never instead of) the
|
|
150
|
+
existing `mcp.tool.schema_drift.detected` span event
|
|
151
|
+
(`schema-drift/emitter.js`) — the same resolution ADR 011 already
|
|
152
|
+
applied to thrash detection (`mcp.tool.thrash_detected`) for the
|
|
153
|
+
identical event-vs-attribute ambiguity: whether a Collector
|
|
154
|
+
`tailsamplingprocessor`'s `boolean_attribute` policy can match
|
|
155
|
+
span-*event* data (as opposed to top-level span attributes) could not be
|
|
156
|
+
confirmed either way (Go source, not installed in this repository). Only
|
|
157
|
+
ever set to `true`, and only when drift was actually detected — never
|
|
158
|
+
explicitly set `false`.
|
|
159
|
+
|
|
160
|
+
### Fixed — tail-sampling recipe referenced a non-existent span attribute
|
|
161
|
+
|
|
162
|
+
- The README's `tailsamplingprocessor` recipe recommended keying a
|
|
163
|
+
`boolean_attribute` policy on `mcp.tool.schema_drift.detected` — but
|
|
164
|
+
that string was only ever a span *event* name and a metric counter name
|
|
165
|
+
(`schema-drift/attributes.js`), never passed to `span.setAttribute()`
|
|
166
|
+
anywhere in this package. As documented, that policy could never have
|
|
167
|
+
matched anything. Fixed by the new `mcp.tool.schema_drift_detected`
|
|
168
|
+
attribute above.
|
|
169
|
+
- Also added a `string_attribute` policy on `mcp.tool.pricing_status =
|
|
170
|
+
"unknown"` — v0.11.0's "confidently wrong zero" problem reappearing at
|
|
171
|
+
the sampling layer: an unpriced call has real extracted token usage but
|
|
172
|
+
no `mcp.tool.cost.usd`, indistinguishable from a genuinely free call to
|
|
173
|
+
a numeric-threshold policy.
|
|
174
|
+
- Recipe YAML extracted to `docs/recipes/tail-sampling.yaml` (repo-only,
|
|
175
|
+
not published — the same carve-out `dashboards/` already has), with a
|
|
176
|
+
new cross-check test (`test/recipes/tail-sampling-attributes.test.js`)
|
|
177
|
+
asserting every attribute a policy references is a real, exported
|
|
178
|
+
constant AND actually passed to `span.setAttribute()` — the check that
|
|
179
|
+
would have caught this bug automatically. The README now references the
|
|
180
|
+
file instead of duplicating it, and states plainly that attribute
|
|
181
|
+
*names* are cross-checked but Collector policy *behavior* itself has not
|
|
182
|
+
been run end to end (no Docker/Collector available in this project's dev
|
|
183
|
+
environment).
|
|
184
|
+
- `docs/known-gaps.md` entry 9 (new): an unpriced call never reaches
|
|
185
|
+
`budgetTracker.recordAndCheck()`, so budget guardrails cannot trip on
|
|
186
|
+
unpriced spend regardless of amount. The visibility half is fixed in
|
|
187
|
+
this same release — see "Added" below — but the underlying accounting
|
|
188
|
+
behavior is not; what should happen to an unpriced call's budget
|
|
189
|
+
accounting is a real design question, deliberately left open.
|
|
190
|
+
|
|
191
|
+
### Added — visibility for unpriced spend against a configured budget (known-gaps entry 9)
|
|
192
|
+
|
|
193
|
+
- **Two new one-time `diag.warn()` diagnostics in `createBudgetTracker()`**
|
|
194
|
+
(`src/cost/budget.js`), no behavior change and no new public surface:
|
|
195
|
+
one fires at construction whenever `perSessionUsd`/`perToolUsd` is
|
|
196
|
+
configured at all, stating plainly that unpriced calls won't count
|
|
197
|
+
toward it; the other — a new `recordUnpriced(model)` method, called from
|
|
198
|
+
`applyCostAttribution()`'s existing `costUsd === null` branch
|
|
199
|
+
(`src/instrument.js`), the same branch that already sets
|
|
200
|
+
`mcp.tool.pricing_status: "unknown"` — fires the first time an unpriced
|
|
201
|
+
call under an active budget is actually observed, naming the model and
|
|
202
|
+
which scope(s) are configured. Both no-op when no budget is configured;
|
|
203
|
+
neither changes `BudgetCheckResult`'s shape or adds a span attribute.
|
|
204
|
+
- **Deliberately diagnostics only — no fallback pricing was added.**
|
|
205
|
+
Making an unpriced call actually count toward a USD budget means
|
|
206
|
+
inventing a number for it, and a wrong invented price is a *different*
|
|
207
|
+
confidently-wrong number, not a fix — the same disease this warning
|
|
208
|
+
exists to flag, one layer up. See `docs/known-gaps.md` entry 9's
|
|
209
|
+
"Status update (v0.12.0)" for the full argument against building that
|
|
210
|
+
now, and why it stays open as a future, ADR-gated decision rather than
|
|
211
|
+
folded into this patch.
|
|
212
|
+
- Both warnings share the budget tracker's own existing
|
|
213
|
+
once-per-tracker-instance granularity (`createBudgetTracker()`'s own
|
|
214
|
+
docblock) — under the default (no `instanceKey`), a fresh-server-per-request
|
|
215
|
+
deployment re-warns on every request for both, the same inherited-caveat
|
|
216
|
+
shape `docs/known-gaps.md` entry 6 documents for the thrash fallback
|
|
217
|
+
warning; a stable `instanceKey` shares one tracker, and one already-armed
|
|
218
|
+
warning, across calls, same as every other registry-backed tracker.
|
|
219
|
+
|
|
3
220
|
## 0.11.0
|
|
4
221
|
|
|
5
222
|
**⚠️ Type change, not a runtime behavior change — read this first.**
|