opentel-mcp 0.12.0 → 0.14.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/CHANGELOG.md CHANGED
@@ -1,5 +1,214 @@
1
1
  # Changelog
2
2
 
3
+ ## 0.14.0
4
+
5
+ Adds `errorRecording.redactor`: a host-supplied hook for content
6
+ `normalizeMessage()`/`parseAndNormalizeStack()` don't recognize — a
7
+ proprietary API key format, an internal account id shape, a customer name
8
+ in prose or in a multi-tenant stack frame path. Full design: ADR 020
9
+ (`docs/adr/020-redactor-hook.md`). Ships in the same release as the
10
+ 0.13.1 fix/docs items below (0.13.1 itself was never tagged/published —
11
+ see that section's own note), but is a logically separate, purely
12
+ additive change from both of them; kept in its own section here rather
13
+ than folded into 0.13.1's for that reason.
14
+
15
+ ### Added — `errorRecording.redactor` hook (ADR 020)
16
+
17
+ - New optional `errorRecording.redactor` field: a synchronous
18
+ `(input: { message, stack }) => { message, stack }` function, consulted
19
+ only under `mode: 'normalized'`, that runs on the raw, uncoerced
20
+ `message`/`stack` — via the same coercion this library's own pipeline
21
+ already uses — **before** `normalizeMessage()`/`parseAndNormalizeStack()`
22
+ ever see them, never the reverse. This library's own scrubbing still
23
+ runs second, over your redactor's output, as a defense-in-depth
24
+ backstop.
25
+ - Configuring a redactor alongside `mode: 'full'` or `'none'` is accepted
26
+ but never invoked — those two modes' entire meaning is "byte-identical,
27
+ regardless of what else is configured" — and produces a one-time
28
+ `diag.warn()` at `instrumentMcpServer()` setup naming the no-op, since
29
+ forgetting to also flip `mode` is an easy, otherwise-silent
30
+ misconfiguration.
31
+ - A redactor that throws, returns a non-string `message`, or returns a
32
+ `stack` that's neither a string nor `undefined` falls back to
33
+ `'none'`-equivalent span output for that one event — status set, no
34
+ `exception` event — **never** to raw/unredacted content, plus a
35
+ one-time, content-free `diag.warn()` naming the failure shape (never
36
+ the message/stack content itself). The returned `message`/`stack` are
37
+ also length-capped defensively before use (reusing
38
+ `normalizeMessage()`'s existing 2048-character limit).
39
+ - **Never reaches `computeFingerprint()`'s hash input, in any
40
+ configuration.** `mcp.failure.fingerprint`/`signature` are computed
41
+ from the real, unmodified thrown error, identically whether or not a
42
+ redactor is configured — a deliberate, non-negotiable design decision
43
+ (ADR 020 Decision 3): the fingerprint is a SHA-256 hash, never a
44
+ plaintext channel, so redacting it would buy no privacy while tying
45
+ fingerprint stability to unversioned host code. Any existing dashboard
46
+ filter or alert keyed on `mcp.failure.fingerprint` keeps working
47
+ unchanged after adding a redactor.
48
+ - **Timing caveat, not just a telemetry one:** the redactor runs
49
+ synchronously, inline, on the error path, before the span ends — there
50
+ is no timeout, and JavaScript's single-threaded execution model means
51
+ one can't be added without a `worker_threads`/`vm.Script` boundary this
52
+ feature deliberately doesn't pay for. A slow redactor (catastrophic
53
+ regex backtracking, an accidental blocking call) adds directly to that
54
+ tool call's own response latency, not just to what shows up in traces.
55
+ See the README's "The redactor hook" section, "Timing," for the full
56
+ writeup, and ADR 020's own "Timing" section for why this is an accepted
57
+ risk rather than a solved one.
58
+ - No environment-variable equivalent — a function can't be expressed as
59
+ an `OTEL_MCP_*` string, same as `costTracking.extractor`/
60
+ `thrashDetection`'s classifier-shaped options already are.
61
+ - New exported types `ErrorRecordingRedactor` and
62
+ `ErrorRecordingRedactorFields` (`src/error-recording/types.d.ts`,
63
+ re-exported from the package root) — a consumer can type a redactor
64
+ function on its own, the same way `UsageExtractor`/`Classifier` already
65
+ let a consumer type an extractor/classifier independently of the option
66
+ object that carries it.
67
+ - Purely additive and default-off: a deployment that never sets
68
+ `errorRecording.redactor` observes zero behavior change.
69
+
70
+ ## 0.13.1
71
+
72
+ Never tagged or published as its own release — ships bundled into
73
+ 0.14.0 above instead, since the redactor work started before this patch
74
+ went out the door. Kept as its own section here (rather than merged into
75
+ 0.14.0's) because both items below are logically independent of the
76
+ redactor: one's a bug fix to existing `'normalized'`-mode behavior, the
77
+ other's a documentation-only correction, and neither has anything to do
78
+ with the new hook.
79
+
80
+ ### Fixed — `'normalized'` mode dropped `exception.message`/status message for non-`Error` throws
81
+
82
+ - Before this fix, `recordThrownException()`'s `'normalized'` branch
83
+ independently re-derived `message`/`stack` via a plain `err?.message`/
84
+ `err?.stack` read — which silently returned `undefined` for a thrown
85
+ **string** or a plain non-`Error` object (neither has a `.message`
86
+ property the way a real `Error` instance does), while
87
+ `computeFingerprint()`'s own, richer `coerceError()` correctly
88
+ recognized those same shapes and still produced a real
89
+ `mcp.failure.fingerprint`. Net effect: a thrown string or non-`Error`
90
+ throw got real fingerprinting but a completely empty `exception.message`
91
+ / status message on the span under `'normalized'` mode — same `err`,
92
+ two independent readers, two different answers, silently.
93
+ - Fixed by extracting the shared `normalizeException()`/`coerceError()`
94
+ pair into their own module (`src/fingerprint/normalize/exception.js`)
95
+ and routing **both** `computeFingerprint()` and `recordThrownException()`
96
+ through the same one computation for the same `err` — the span's
97
+ exception content and the fingerprint's hashed inputs can no longer
98
+ independently drift apart, by construction, not by convention. This is
99
+ the same consolidation ADR 020 (the redactor's own design doc, above)
100
+ repeatedly cites as precedent for keeping the fingerprint path
101
+ structurally isolated from the span-writing path going forward.
102
+ - `'full'` and `'none'` modes were never affected — this bug was specific
103
+ to `'normalized'` mode's own inline coercion.
104
+
105
+ ### Documentation — `errorRecording.mode` scoping correction
106
+
107
+ - Clarified that `errorRecording.mode` controls only the span this
108
+ library creates for the current `tools/call`/`tools/list` — it has no
109
+ effect on exception content any other instrumentation in the same
110
+ process (an APM agent, HTTP or framework auto-instrumentation, anything
111
+ else wrapping the handler) independently records onto its own span for
112
+ the same rethrown error. Since no mode mutates the original `err`, that
113
+ raw content can still land in the same trace, one span up, regardless
114
+ of mode — including `'none'`, which was previously worded in a way that
115
+ could read as a trace-wide guarantee rather than a span-scoped one.
116
+ Confirmed empirically (a real ambient context manager, an outer span
117
+ recording the rethrown error, `InMemorySpanExporter`) — see
118
+ `docs/known-gaps.md` entry 10's 2026-09-01 update for the full
119
+ reproduction and reasoning. README's "Error recording" and "What this
120
+ library records" sections updated accordingly. No code changed — the
121
+ underlying behavior (and the exposure at the default `'full'` mode) was
122
+ already accurate and unchanged; only the description of what
123
+ `errorRecording.mode` scopes to was incomplete.
124
+
125
+ ## 0.13.0
126
+
127
+ Closes the two open items from `docs/known-gaps.md` entry 10 (a
128
+ raw-content audit of every attribute/span event this package emits) that
129
+ needed a design decision before a fix — raw exception content on
130
+ `recordException`/`setStatus`, and an unvalidated tool-result model field
131
+ reaching `mcp.tool.model`/`gen_ai.response.model` — plus the two
132
+ lower-risk fixes from the same entry that didn't need one. Full design:
133
+ ADR 019 (`docs/adr/019-raw-content-on-spans.md`). See the README's new
134
+ "What this library records" and "Error recording" sections for the
135
+ operator-facing consequence of each.
136
+
137
+ ### Added — `errorRecording.mode` config (ADR 019 Part 1)
138
+
139
+ - New top-level `errorRecording` option, sibling to `fingerprinting` /
140
+ `costTracking` / `thrashDetection` / `schemaDrift`: `'full'` (default,
141
+ byte-for-byte unchanged from every prior release —
142
+ `recordException(err)` + `setStatus({ message: err.message })` with the
143
+ raw error), `'normalized'` (reuses the existing
144
+ `normalizeMessage()`/`parseAndNormalizeStack()` fingerprinting pipeline
145
+ — no new scrubbing logic — to strip known-sensitive-shaped substrings
146
+ from the message and the local `cwd` prefix from the stack, without
147
+ mutating the original `err`, which both call sites still rethrow), or
148
+ `'none'` (records neither — the same `setStatus({ code: ERROR })`
149
+ no-message pattern already used for tool-level `isError: true`
150
+ failures). Applies to both thrown-exception paths, `tools/call` and
151
+ `tools/list`.
152
+ - New `OTEL_MCP_ERROR_RECORDING_MODE` env var — same
153
+ option-then-env-then-default precedence, and the same silent fallback
154
+ to the default on an unrecognized value, as every other `OTEL_MCP_*`
155
+ config.
156
+ - `error.type`/`exception.type` (both read from `err.name`) are now
157
+ capped at 128 characters unconditionally, in every mode — the same cap
158
+ `mcp.failure.error_class` uses below, reusing its exported constant
159
+ rather than a second, possibly-drifting copy.
160
+ - Default stays `'full'` for all of `0.x`; ADR 019 Part 1 records the
161
+ intent to flip it to `'normalized'` at `1.0`, not before — not decided
162
+ in this release.
163
+
164
+ ### Added — `mcp.tool.model` / `gen_ai.response.model` validation (ADR 019 Part 2)
165
+
166
+ **This is new behavior that can change what a call reports, not just a
167
+ new diagnostic.** Before this release, a tool result's `model` field
168
+ reached the span completely unvalidated, whatever it was. If you have a
169
+ provider/deployment whose model identifiers use a character outside
170
+ `[A-Za-z0-9._:/@-]`, or that (implausibly, but possibly) exceed 256
171
+ characters, upgrading will make `mcp.tool.model`/`gen_ai.response.model`
172
+ disappear from those calls' spans and `mcp.tool.pricing_status` flip from
173
+ whatever it was to `"unknown"` — even for an otherwise-legitimate, real
174
+ model id. The allowlist was deliberately built generous (see below) and
175
+ verified against known provider conventions, but it's still new, and a
176
+ one-time `diag.warn()` names when this happens (shape only, never the
177
+ value) so it's discoverable rather than a silent metric/span change.
178
+
179
+ - A tool result's declared model field is now checked against a 256-
180
+ character length cap and a `[A-Za-z0-9._:/@-]` allowlist — verified
181
+ against every `DEFAULT_PRICING` key and `normalizeModelName()`'s
182
+ documented `provider/model` input contract, deliberately generous —
183
+ before it can reach `mcp.tool.model`, `gen_ai.response.model`,
184
+ `calculateCost()`, or either metric label
185
+ (`mcp.tool.tokens.total`/`mcp.tool.cost.total`).
186
+ - A rejected value is never a silent drop: `mcp.tool.pricing_status` is
187
+ set to `"unknown"` (the same status a legitimately unrecognized model
188
+ already produces), and a one-time `diag.warn()` fires per
189
+ `instrumentMcpServer()` call, reporting shape only — length, and which
190
+ check failed — never the rejected value itself.
191
+ - `budgetTracker.recordUnpriced()` now receives the validated (possibly
192
+ `undefined`) model rather than the raw tool-result value, closing a
193
+ second leak path through its own pre-existing warning that would
194
+ otherwise have echoed a rejected value verbatim.
195
+ - `costTracking.pricing`/`pricingTable` override keys are unaffected —
196
+ operator-authored config, never subject to this gate.
197
+
198
+ ### Fixed — `mcp.failure.error_class` uncapped length, `mcp.failure.validation_paths` dynamic-key leak (known-gaps entry 10)
199
+
200
+ - `mcp.failure.error_class` is now capped at 128 characters
201
+ (`fingerprint/compose.js`'s new `MAX_ERROR_CLASS_LENGTH`) — length-
202
+ bounded only, not pattern-scrubbed; a non-string `.name` is coerced to
203
+ a string before capping rather than thrown. ADR 004's note calling the
204
+ underlying value "low-cardinality" is updated to flag that as an
205
+ assumption, not an enforced property.
206
+ - `mcp.failure.validation_paths` format 1 (the raw Zod issues array, SDK
207
+ ≤1.29.0) now redacts a non-identifier-shaped path segment — a
208
+ `z.record()` schema's runtime key — to the placeholder `<KEY>` instead
209
+ of surfacing it verbatim, matching the identifier-only gate formats 2/3
210
+ already applied. Numeric (array-index) segments are never redacted.
211
+
3
212
  ## 0.12.0
4
213
 
5
214
  Agent Thrash Detection gains a new, narrower session-identity fallback for