opentel-mcp 0.13.0 → 0.14.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/CHANGELOG.md CHANGED
@@ -1,5 +1,127 @@
1
1
  # Changelog
2
2
 
3
+ ## 0.14.0
4
+
5
+ Adds `errorRecording.redactor`: a host-supplied hook for content
6
+ `normalizeMessage()`/`parseAndNormalizeStack()` don't recognize — a
7
+ proprietary API key format, an internal account id shape, a customer name
8
+ in prose or in a multi-tenant stack frame path. Full design: ADR 020
9
+ (`docs/adr/020-redactor-hook.md`). Ships in the same release as the
10
+ 0.13.1 fix/docs items below (0.13.1 itself was never tagged/published —
11
+ see that section's own note), but is a logically separate, purely
12
+ additive change from both of them; kept in its own section here rather
13
+ than folded into 0.13.1's for that reason.
14
+
15
+ ### Added — `errorRecording.redactor` hook (ADR 020)
16
+
17
+ - New optional `errorRecording.redactor` field: a synchronous
18
+ `(input: { message, stack }) => { message, stack }` function, consulted
19
+ only under `mode: 'normalized'`, that runs on the raw, uncoerced
20
+ `message`/`stack` — via the same coercion this library's own pipeline
21
+ already uses — **before** `normalizeMessage()`/`parseAndNormalizeStack()`
22
+ ever see them, never the reverse. This library's own scrubbing still
23
+ runs second, over your redactor's output, as a defense-in-depth
24
+ backstop.
25
+ - Configuring a redactor alongside `mode: 'full'` or `'none'` is accepted
26
+ but never invoked — those two modes' entire meaning is "byte-identical,
27
+ regardless of what else is configured" — and produces a one-time
28
+ `diag.warn()` at `instrumentMcpServer()` setup naming the no-op, since
29
+ forgetting to also flip `mode` is an easy, otherwise-silent
30
+ misconfiguration.
31
+ - A redactor that throws, returns a non-string `message`, or returns a
32
+ `stack` that's neither a string nor `undefined` falls back to
33
+ `'none'`-equivalent span output for that one event — status set, no
34
+ `exception` event — **never** to raw/unredacted content, plus a
35
+ one-time, content-free `diag.warn()` naming the failure shape (never
36
+ the message/stack content itself). The returned `message`/`stack` are
37
+ also length-capped defensively before use (reusing
38
+ `normalizeMessage()`'s existing 2048-character limit).
39
+ - **Never reaches `computeFingerprint()`'s hash input, in any
40
+ configuration.** `mcp.failure.fingerprint`/`signature` are computed
41
+ from the real, unmodified thrown error, identically whether or not a
42
+ redactor is configured — a deliberate, non-negotiable design decision
43
+ (ADR 020 Decision 3): the fingerprint is a SHA-256 hash, never a
44
+ plaintext channel, so redacting it would buy no privacy while tying
45
+ fingerprint stability to unversioned host code. Any existing dashboard
46
+ filter or alert keyed on `mcp.failure.fingerprint` keeps working
47
+ unchanged after adding a redactor.
48
+ - **Timing caveat, not just a telemetry one:** the redactor runs
49
+ synchronously, inline, on the error path, before the span ends — there
50
+ is no timeout, and JavaScript's single-threaded execution model means
51
+ one can't be added without a `worker_threads`/`vm.Script` boundary this
52
+ feature deliberately doesn't pay for. A slow redactor (catastrophic
53
+ regex backtracking, an accidental blocking call) adds directly to that
54
+ tool call's own response latency, not just to what shows up in traces.
55
+ See the README's "The redactor hook" section, "Timing," for the full
56
+ writeup, and ADR 020's own "Timing" section for why this is an accepted
57
+ risk rather than a solved one.
58
+ - No environment-variable equivalent — a function can't be expressed as
59
+ an `OTEL_MCP_*` string, same as `costTracking.extractor`/
60
+ `thrashDetection`'s classifier-shaped options already are.
61
+ - New exported types `ErrorRecordingRedactor` and
62
+ `ErrorRecordingRedactorFields` (`src/error-recording/types.d.ts`,
63
+ re-exported from the package root) — a consumer can type a redactor
64
+ function on its own, the same way `UsageExtractor`/`Classifier` already
65
+ let a consumer type an extractor/classifier independently of the option
66
+ object that carries it.
67
+ - Purely additive and default-off: a deployment that never sets
68
+ `errorRecording.redactor` observes zero behavior change.
69
+
70
+ ## 0.13.1
71
+
72
+ Never tagged or published as its own release — ships bundled into
73
+ 0.14.0 above instead, since the redactor work started before this patch
74
+ went out the door. Kept as its own section here (rather than merged into
75
+ 0.14.0's) because both items below are logically independent of the
76
+ redactor: one's a bug fix to existing `'normalized'`-mode behavior, the
77
+ other's a documentation-only correction, and neither has anything to do
78
+ with the new hook.
79
+
80
+ ### Fixed — `'normalized'` mode dropped `exception.message`/status message for non-`Error` throws
81
+
82
+ - Before this fix, `recordThrownException()`'s `'normalized'` branch
83
+ independently re-derived `message`/`stack` via a plain `err?.message`/
84
+ `err?.stack` read — which silently returned `undefined` for a thrown
85
+ **string** or a plain non-`Error` object (neither has a `.message`
86
+ property the way a real `Error` instance does), while
87
+ `computeFingerprint()`'s own, richer `coerceError()` correctly
88
+ recognized those same shapes and still produced a real
89
+ `mcp.failure.fingerprint`. Net effect: a thrown string or non-`Error`
90
+ throw got real fingerprinting but a completely empty `exception.message`
91
+ / status message on the span under `'normalized'` mode — same `err`,
92
+ two independent readers, two different answers, silently.
93
+ - Fixed by extracting the shared `normalizeException()`/`coerceError()`
94
+ pair into their own module (`src/fingerprint/normalize/exception.js`)
95
+ and routing **both** `computeFingerprint()` and `recordThrownException()`
96
+ through the same one computation for the same `err` — the span's
97
+ exception content and the fingerprint's hashed inputs can no longer
98
+ independently drift apart, by construction, not by convention. This is
99
+ the same consolidation ADR 020 (the redactor's own design doc, above)
100
+ repeatedly cites as precedent for keeping the fingerprint path
101
+ structurally isolated from the span-writing path going forward.
102
+ - `'full'` and `'none'` modes were never affected — this bug was specific
103
+ to `'normalized'` mode's own inline coercion.
104
+
105
+ ### Documentation — `errorRecording.mode` scoping correction
106
+
107
+ - Clarified that `errorRecording.mode` controls only the span this
108
+ library creates for the current `tools/call`/`tools/list` — it has no
109
+ effect on exception content any other instrumentation in the same
110
+ process (an APM agent, HTTP or framework auto-instrumentation, anything
111
+ else wrapping the handler) independently records onto its own span for
112
+ the same rethrown error. Since no mode mutates the original `err`, that
113
+ raw content can still land in the same trace, one span up, regardless
114
+ of mode — including `'none'`, which was previously worded in a way that
115
+ could read as a trace-wide guarantee rather than a span-scoped one.
116
+ Confirmed empirically (a real ambient context manager, an outer span
117
+ recording the rethrown error, `InMemorySpanExporter`) — see
118
+ `docs/known-gaps.md` entry 10's 2026-09-01 update for the full
119
+ reproduction and reasoning. README's "Error recording" and "What this
120
+ library records" sections updated accordingly. No code changed — the
121
+ underlying behavior (and the exposure at the default `'full'` mode) was
122
+ already accurate and unchanged; only the description of what
123
+ `errorRecording.mode` scopes to was incomplete.
124
+
3
125
  ## 0.13.0
4
126
 
5
127
  Closes the two open items from `docs/known-gaps.md` entry 10 (a
package/README.md CHANGED
@@ -463,12 +463,26 @@ per-call-site wrapping; tracked in the roadmap below.
463
463
 
464
464
  ## Error recording (v0.13.0+)
465
465
 
466
- Every thrown `tools/call`/`tools/list` error goes through
467
- `span.recordException(err)` (an OpenTelemetry SDK method, not one of this
468
- library's own attributes) plus `span.setStatus({ code: ERROR, message:
469
- err.message })` — unconditionally, whether or not `fingerprinting` is
470
- enabled. By default that means `err.message` and `err.stack` land on the
471
- span exactly as thrown. `errorRecording.mode` controls this:
466
+ **What this controls, and who needs it.** Whenever a `tools/call`/
467
+ `tools/list` handler throws, this library records the thrown error's
468
+ message and stack trace onto the span it creates for that call — by
469
+ default, exactly as the tool wrote them, no redaction. If your tools call
470
+ third-party services, wrap unaudited dependencies, or you otherwise don't
471
+ control every string a tool might throw, that's a real channel for
472
+ whatever those errors happen to contain (an email address, an internal
473
+ hostname, a connection string a driver embedded in its own error message)
474
+ to reach wherever your traces get exported. `errorRecording.mode` lets
475
+ you scrub or drop that content instead of recording it verbatim — read
476
+ both caveats below before treating it as a privacy control, though: it
477
+ has two distinct limits, and neither is obvious from the option name
478
+ alone.
479
+
480
+ Mechanically: every thrown error goes through `span.recordException(err)`
481
+ (an OpenTelemetry SDK method, not one of this library's own attributes)
482
+ plus `span.setStatus({ code: ERROR, message: err.message })` —
483
+ unconditionally, whether or not `fingerprinting` is enabled.
484
+ `errorRecording.mode` controls what those two calls actually put on the
485
+ span:
472
486
 
473
487
  | Mode | `exception.message` / status message | `exception.stacktrace` | When to use |
474
488
  |---|---|---|---|
@@ -476,30 +490,197 @@ span exactly as thrown. `errorRecording.mode` controls this:
476
490
  | `'normalized'` | `normalizeMessage(err.message)` — the exact scrubbing pipeline (`src/fingerprint/normalize/message.js`) fingerprinting already runs before hashing: UUIDs, emails, URLs, IPs, timestamps, filesystem paths, hex runs, quoted ids | Reconstructed from `parseAndNormalizeStack()` (`src/fingerprint/normalize/stack.js`) — keeps every function name/file/line, strips only the local `cwd` prefix (and collapses `node_modules` package versions) | Tool results come from third-party or unaudited MCP servers and you want the same scrubbing fingerprinting already trusts, applied to the raw exception content too |
477
491
  | `'none'` | Not set — `span.setStatus({ code: ERROR })` with no message, the same pattern already used for tool-level `isError: true` failures | Not set | You rely entirely on `mcp.failure.*` (category/fingerprint/signature — already hashed/normalized) and don't want any free-text exception content on the span at all |
478
492
 
479
- No mode mutates the original `err` — both call sites rethrow it
480
- afterward, so `'normalized'`/`'none'` build the exception event
481
- independently rather than editing `err.message`/`err.stack` in place.
482
- `error.type`/`exception.type` (`err.name`) is capped at 128 characters
483
- unconditionally in every mode — the same cap `mcp.failure.error_class`
484
- uses — since a length cap on a class-identifier field costs a
485
- well-behaved tool nothing, unlike message/stack content.
486
-
487
- **`'normalized'` is targeted scrubbing, not general-purpose redaction —
488
- read this before treating it as a PII filter.** `normalizeMessage()`
489
- matches specific, structured shapes: UUIDs, email addresses, URLs,
493
+ **Caveat 1 — `'normalized'` is targeted scrubbing, not general-purpose
494
+ redaction. Do not treat it as a PII filter.** `normalizeMessage()`
495
+ matches specific, structured shapes only: UUIDs, email addresses, URLs,
490
496
  IPv4/IPv6 addresses, ISO-8601/Unix timestamps, filesystem paths, long hex
491
497
  runs, and quoted alphanumeric ids (8–64 chars, mixed letters/digits). It
492
- does not recognize sensitive content in general. An API key in a format
498
+ does not recognize sensitive content in general — an API key in a format
493
499
  none of those patterns match (a bare, unquoted token with no digit in it,
494
500
  or a custom prefix scheme), or a customer's name embedded in ordinary
495
501
  prose ("could not process request for Jane Smith"), passes through
496
- `'normalized'` mode completely unchanged — identical to what `'full'`
497
- mode would put on the span. Treat `'normalized'` as "the same scrubbing
502
+ `'normalized'` mode completely unchanged, identical to what `'full'`
503
+ would put on the span. Treat `'normalized'` as "the same scrubbing
498
504
  fingerprinting already trusts for hashing," not as a guarantee that
499
- whatever a tool's error messages contain is safe to record; if a tool's
500
- errors routinely carry sensitive free text these patterns don't happen to
501
- match, `'none'` is the only mode that keeps message/stack content off the
502
- span entirely.
505
+ whatever a tool's error messages contain is safe to record.
506
+
507
+ **Example — the same thrown error, `'full'` vs. `'normalized'`:**
508
+
509
+ ```js
510
+ throw new Error('upstream lookup failed for user jane.doe@example.com');
511
+ ```
512
+
513
+ | Mode | `exception.message` on the span |
514
+ |---|---|
515
+ | `'full'` | `upstream lookup failed for user jane.doe@example.com` |
516
+ | `'normalized'` | `upstream lookup failed for user <EMAIL>` |
517
+ | `'none'` | *(not set — only the ERROR status code is)* |
518
+
519
+ The stack trace scrubs the same way: `'full'` keeps every frame's
520
+ absolute path; `'normalized'` keeps every frame (function name, file,
521
+ line) but strips the local `cwd` prefix.
522
+
523
+ ### The redactor hook (v0.14.0+)
524
+
525
+ Caveat 1 above is the reason this exists: `normalizeMessage()` matches
526
+ eight specific, structured shapes (UUIDs, emails, URLs, IPv4/IPv6,
527
+ timestamps, paths, hex runs, quoted ids), and that is where its
528
+ competence ends. An API key in your own proprietary format
529
+ (`ACME_KEY_[a-f0-9]{8}`), an internal account id shape
530
+ (`ACCT-\d{9}`), or a customer's name embedded in ordinary prose are not
531
+ structurally distinguishable from ordinary text by any pattern this
532
+ library could ship — no ninth built-in pattern fixes that. Only the host
533
+ running the tool knows what their own data looks like.
534
+ `errorRecording.redactor` is a synchronous function you supply that runs
535
+ as part of `mode: 'normalized'`, so you can close that gap without
536
+ forking this library or falling back to `'none'`'s blunt "drop
537
+ everything." See ADR 020 (`docs/adr/020-redactor-hook.md`) for the full
538
+ design rationale; the operator-facing behavior is summarized below.
539
+
540
+ **Worked example:**
541
+
542
+ ```js
543
+ instrumentMcpServer(server, {
544
+ errorRecording: {
545
+ mode: 'normalized',
546
+ redactor: ({ message, stack }) => ({
547
+ message: message
548
+ .replace(/ACME_KEY_[a-f0-9]{8}/g, '[API_KEY]')
549
+ .replace(/ACCT-\d{9}/g, '[ACCOUNT_ID]'),
550
+ stack, // unchanged — this redactor only needs to touch the message
551
+ }),
552
+ },
553
+ });
554
+ ```
555
+
556
+ ```js
557
+ throw new Error(
558
+ 'auth failed for ACME_KEY_7f3a9c2e, ' +
559
+ 'account ACCT-123456789, contact jane.doe@example.com',
560
+ );
561
+ ```
562
+
563
+ | Step | `exception.message` |
564
+ |---|---|
565
+ | Raw, as thrown | `auth failed for ACME_KEY_7f3a9c2e, account ACCT-123456789, contact jane.doe@example.com` |
566
+ | After your redactor (runs first, on the raw text) | `auth failed for [API_KEY], account [ACCOUNT_ID], contact jane.doe@example.com` |
567
+ | After this library's own patterns (run second, on YOUR output) | `auth failed for [API_KEY], account [ACCOUNT_ID], contact <EMAIL>` |
568
+
569
+ **Ordering matters, and it's fixed, not configurable:** your redactor
570
+ always runs first, on the raw, uncoerced `message`/`stack` — before
571
+ `normalizeMessage()`/`parseAndNormalizeStack()` ever see them, never the
572
+ reverse. Two consequences follow directly: (1) this library's own
573
+ scrubbing is still the last thing that touches the content before it
574
+ reaches the span — a defense-in-depth backstop if your redactor misses a
575
+ shape its own patterns don't cover (see the API key example above:
576
+ notice `<EMAIL>` still gets scrubbed even though your redactor never
577
+ mentioned email addresses), and (2) your redactor sees the full,
578
+ untruncated original message — `normalizeMessage()`'s 2048-character
579
+ truncation applies afterward, to your output, not before you ever see the
580
+ content.
581
+
582
+ **Only consulted under `mode: 'normalized'`.** Configure a redactor
583
+ alongside `'full'` or `'none'` and it's accepted but never called — those
584
+ two modes' entire meaning is "byte-identical to their own fixed
585
+ behavior, regardless of what else is configured." Since forgetting to
586
+ also flip `mode` away from its `'full'` default is an easy way to end up
587
+ with a redactor that's silently never invoked, `instrumentMcpServer()`
588
+ logs a one-time `diag.warn()` at setup naming exactly that.
589
+
590
+ **Failure never falls back to raw content.** If your redactor throws,
591
+ returns a non-string `message`, or returns a `stack` that's neither a
592
+ string nor `undefined`, that one event falls back to `'none'`-equivalent
593
+ output — `span.setStatus({ code: ERROR })` only, no `exception` event at
594
+ all — never to `'full'`-equivalent (raw, unredacted) content. A redactor
595
+ you configured specifically to keep something off the span must never
596
+ silently fail open; the safe direction here is *more* redacted, not
597
+ less. This fires a one-time, content-free `diag.warn()` naming the
598
+ failure shape (`threw` / `returned a non-string message` / `returned an
599
+ invalid stack`) — never the message/stack content that triggered it,
600
+ so the warning itself can't become a second, undocumented leak channel.
601
+ Your redactor's returned `message`/`stack` are also length-capped
602
+ defensively before use (reusing `normalizeMessage()`'s own 2048-character
603
+ limit), regardless of what a well-behaved redactor is expected to
604
+ return.
605
+
606
+ **Timing — read this before you write a regex.** The redactor runs
607
+ *synchronously*, inline, on the error path — inside the still-open
608
+ span's own callback, before `span.end()` is called. There is no timeout,
609
+ and there cannot be one: JavaScript is single-threaded, and nothing in
610
+ this library (or in Node itself, short of a `worker_threads` boundary
611
+ this feature deliberately doesn't pay for) can interrupt a synchronous
612
+ function that's already running. **This means a slow redactor doesn't
613
+ just delay telemetry — it delays the tool call's own response to
614
+ whoever's waiting on it.** Catastrophic backtracking in a regex (the
615
+ classic `(a+)+$`-shaped footgun), an accidental synchronous file read, or
616
+ any blocking call inside your redactor adds directly to that tool call's
617
+ latency, on every single call that throws under `'normalized'` mode,
618
+ for as long as the redactor stays slow. Test your redactor's regexes
619
+ against adversarial input the same way you would for any other
620
+ user-facing regex, not just against the examples you had in mind when
621
+ you wrote it.
622
+
623
+ **The fingerprint is completely unaffected — on purpose, always.**
624
+ `mcp.failure.fingerprint`/`mcp.failure.signature` and the values hashed
625
+ into them are computed from the real, unmodified thrown error, identically
626
+ whether or not a redactor is configured. A redactor has no way to reach
627
+ that computation at all — it only ever touches what lands on the span.
628
+ This is deliberate: the fingerprint is a SHA-256 hash, never a plaintext
629
+ channel, so redacting it would buy no privacy while tying fingerprint
630
+ stability to your own, unversioned redactor code. Practically, this means
631
+ any saved dashboard filter or alert keyed on `mcp.failure.fingerprint`
632
+ keeps working exactly as before the moment you add a redactor — nothing
633
+ about adding or changing one ever reshuffles a fingerprint value.
634
+
635
+ The redactor is still scoped to this library's own span, the same as
636
+ every other `mode` value — see Caveat 2 below for what that does and
637
+ doesn't cover; nothing about the redactor changes that boundary.
638
+
639
+ **Caveat 2 — every mode is scoped to the span this library creates, not
640
+ to your trace as a whole.** No mode mutates the original `err` — both
641
+ call sites rethrow it afterward, so `'normalized'`/`'none'` build the
642
+ exception event independently rather than editing `err.message`/
643
+ `err.stack` in place. The consequence: the unmodified, fully raw `err` is
644
+ exactly what propagates to whatever called this handler. If any other
645
+ instrumentation further out in the call stack — an APM agent, HTTP or
646
+ framework auto-instrumentation, anything else wrapping this handler —
647
+ independently calls `recordException(err)`/`setStatus({ message:
648
+ err.message })` on its OWN span when it observes the rethrow, the raw,
649
+ unscrubbed message and stack land in the SAME TRACE as this library's
650
+ (scrubbed) span, one span up, not on it. `'none'` buys no more protection
651
+ here than `'normalized'`: neither mutates `err`, so an outer recorder
652
+ sees the identical raw content either way. Someone choosing `'none'`
653
+ specifically to keep sensitive content out of their tracing backend needs
654
+ to know that protection stops at this library's own span, not at the
655
+ trace boundary.
656
+
657
+ Confirmed directly, not just reasoned about: an instrumented server run
658
+ under a real ambient context manager (`AsyncLocalStorageContextManager` —
659
+ the propagation mechanism production deployments actually use, unlike
660
+ this project's own test suite, which registers most test files with
661
+ `contextManager: null` and therefore can't exercise this path — see
662
+ `docs/known-gaps.md` entry 10's latest update for why that's fine for
663
+ those specific tests), wrapped in an outer span that calls
664
+ `recordException(err)` on catch, with `errorRecording.mode: 'normalized'`
665
+ set and an `InMemorySpanExporter` attached: the tool span's exception
666
+ event carries the scrubbed message, and the ancestor span's exception
667
+ event and status message both carry the original, raw one — same trace
668
+ ID, parent/child span relationship confirmed, no mutation anywhere in
669
+ between.
670
+
671
+ **This is not fixable from inside this library.** The only way to close
672
+ it here would be mutating `err.message`/`err.stack` in place before
673
+ rethrowing — which would corrupt whatever error-handling the host
674
+ application does with that same object afterward, a correctness hazard
675
+ this library isn't willing to trade for a telemetry one. It has authority
676
+ over one span; it has no visibility into, and no control over, what other
677
+ instrumentation attached to the same process does with the error once
678
+ rethrown. See `docs/known-gaps.md` entry 10 for the full writeup.
679
+
680
+ `error.type`/`exception.type` (`err.name`) is capped at 128 characters
681
+ unconditionally in every mode — the same cap `mcp.failure.error_class`
682
+ uses — since a length cap on a class-identifier field costs a
683
+ well-behaved tool nothing, unlike message/stack content.
503
684
 
504
685
  **Default stays `'full'` through all of `0.x`.** Changing it would alter
505
686
  what every trace backend renders for the single highest-traffic failure
@@ -513,6 +694,7 @@ for the full argument, including why the default is expected to flip to
513
694
  | Option | Env var | Type | Default | Description |
514
695
  |---|---|---|---|---|
515
696
  | `mode` | `OTEL_MCP_ERROR_RECORDING_MODE` | `'full'` \| `'normalized'` \| `'none'` | `'full'` | See table above. An unrecognized value falls back to `'full'` silently, same as every other `OTEL_MCP_*` env var |
697
+ | `redactor` | *(none — a function can't be expressed as an env var string)* | `(input: { message: string, stack: string \| undefined }) => { message: string, stack: string \| undefined }` | `undefined` | See "The redactor hook" above. Only consulted when `mode` is `'normalized'`; a non-function value is treated as absent, silently |
516
698
 
517
699
  ## Cost & Token Attribution (v0.5.0)
518
700
 
@@ -1863,23 +2045,145 @@ sensitive information" and still specify recording it by default —
1863
2045
  every other OTel-instrumented library sharing the same trace (an HTTP
1864
2046
  client, a DB driver, a queue consumer) records `err.message`/`err.stack`
1865
2047
  the same way, unscrubbed, through the same `recordException()` call.
1866
- opentel-mcp matches that ecosystem default rather than silently diverging
1867
- from it. What it adds on top: `errorRecording.mode` lets you choose
1868
- `'normalized'` (targeted scrubbing — the same structured patterns
1869
- fingerprinting matches: UUIDs, emails, URLs, IPs, timestamps, paths,
1870
- quoted ids; **not** a general-purpose PII filter, so freeform sensitive
1871
- text — a customer's name in prose, an API key in a format none of those
1872
- patterns match — still reaches the span unchanged, see "Error recording"
1873
- above) or `'none'` (record neither) instead of `'full'` — a level of
1874
- operator control most peer instrumentations don't offer for this path.
1875
- Full reasoning, including why the default doesn't change in this
1876
- release: ADR 019 (`docs/adr/019-raw-content-on-spans.md`).
2048
+ That covers a sharper case than "different libraries recording their own,
2049
+ separate operations," too: an outer wrapper — an APM agent's own request
2050
+ span, or HTTP/framework auto-instrumentation sitting further out in the
2051
+ call stack — independently recording the very same error this library
2052
+ just scrubbed. No `errorRecording.mode` mutates the original `err` (see
2053
+ "Error recording" above), so a rethrown error is exactly what such a
2054
+ wrapper sees; if it calls `recordException(err)` on its own span, the raw
2055
+ content lands one span up in the identical trace, regardless of which
2056
+ mode is set here. opentel-mcp matches that ecosystem default rather than
2057
+ silently diverging from it. What it adds on top: `errorRecording.mode`
2058
+ lets you choose `'normalized'` (targeted scrubbing — the same structured
2059
+ patterns fingerprinting matches: UUIDs, emails, URLs, IPs, timestamps,
2060
+ paths, quoted ids; **not** a general-purpose PII filter, so freeform
2061
+ sensitive text — a customer's name in prose, an API key in a format none
2062
+ of those patterns match — still reaches the span unchanged, see "Error
2063
+ recording" above) or `'none'` (record neither) instead of `'full'` — a
2064
+ level of operator control most peer instrumentations don't offer for this
2065
+ path, scoped to the span this library creates, not a trace-wide
2066
+ guarantee. Full reasoning, including why the default doesn't change in
2067
+ this release: ADR 019 (`docs/adr/019-raw-content-on-spans.md`).
1877
2068
 
1878
2069
  This is a data-handling property of an observability library, not a
1879
2070
  security control opentel-mcp is claiming to provide — it doesn't
1880
2071
  authenticate, encrypt, or restrict who can read your traces; that's your
1881
2072
  tracing backend's job.
1882
2073
 
2074
+ ### Full attribute provenance
2075
+
2076
+ The paragraphs above cover the three channels with real content risk;
2077
+ the table below is the complete picture — every attribute this library
2078
+ sets, anywhere, classified by where its value actually comes from. It was
2079
+ produced by tracing every `span.setAttribute()`/`setAttributes()`/
2080
+ `addEvent()` call site back to its source, generalizing the manual audits
2081
+ `docs/known-gaps.md` entries 10 and 11 already did by hand for the
2082
+ riskiest subset.
2083
+
2084
+ **Why a table instead of a runtime "audit mode":** provenance here is a
2085
+ property of which code path sets an attribute, not of any particular
2086
+ call's runtime state — it's the same answer on every call, for a given
2087
+ release, until the code setting that attribute changes. A live inspection
2088
+ feature can't tell you anything about *category* this table doesn't
2089
+ already say, for free, with no new code and no new place for a value to
2090
+ leak. Where a table genuinely falls short — "what's the *actual* value
2091
+ landing in this attribute for my traffic" — the fix isn't a new library
2092
+ feature; it's the extension hooks documented below, which already exist
2093
+ and were built for exactly this. "Only the host knows what's safe to look
2094
+ at in their own deployment" is the same principle `errorRecording.redactor`
2095
+ itself is built around (ADR 020) — reused here, not re-solved.
2096
+
2097
+ Four categories: **library-computed** (a hash, count, enum, or fixed
2098
+ constant — the safe majority), **host-supplied** (an `instrumentMcpServer()`
2099
+ config value), **SDK-supplied** (a transport/protocol id — session,
2100
+ request, span, trace), and **tool-controlled** (content from a tool's own
2101
+ result or the caller's own arguments — the category every raw-content
2102
+ finding in `docs/known-gaps.md` traces back to). Eight rows below are
2103
+ genuine **blends** — a single word would mislead for these specifically,
2104
+ so they're spelled out instead of forced into one bucket.
2105
+
2106
+ | Attribute | Surface | Provenance | Note |
2107
+ |---|---|---|---|
2108
+ | `mcp.method.name` | span attr | library-computed | fixed string per handler |
2109
+ | `gen_ai.operation.name` | span attr | library-computed | fixed string `"execute_tool"` |
2110
+ | `gen_ai.tool.name` | span attr, metric label | **blend — tool-controlled** | `request.params.name`, unvalidated against the registered tool set (`docs/known-gaps.md` entry 11) |
2111
+ | `jsonrpc.request.id` | span attr | SDK-supplied | transport/protocol correlation id |
2112
+ | `mcp.tool.argument_count` | span attr | tool-controlled | a *count* of `request.params.arguments` keys, deliberately — never the argument values themselves |
2113
+ | `error.type` | span attr | **blend — depends on code path** | the tool's own `err.name` on a thrown failure; the fixed constant `tool_error` on an `isError: true` failure — same key, two different sources depending on which failure path fired. See "Two keys whose provenance changes by code path" below |
2114
+ | `gen_ai.response.model` | span attr | tool-controlled | co-emitted from the same value as `mcp.tool.model`, for dashboard compatibility |
2115
+ | `mcp.tool.tokens.input` / `.output` | span attr | tool-controlled | the tool result's own usage numbers |
2116
+ | `mcp.tool.tokens.total` | span attr | library-computed | sum of the two tool-controlled values above |
2117
+ | `mcp.tool.model` | span attr, metric label | tool-controlled | a tool result's own declared model field, echoed verbatim (`docs/known-gaps.md` entry 10, item 2 — still open) |
2118
+ | `mcp.tool.cost.usd` | span attr | **blend — library-computed over tool-controlled × config** | tool-controlled token counts, multiplied by a pricing table that's either this library's default or a host override |
2119
+ | `mcp.tool.cost.currency` | span attr | library-computed | fixed constant `"USD"` |
2120
+ | `mcp.tool.cost.budget_exceeded` / `.budget_scope` | span attr | library-computed | comparison against a host-supplied budget threshold |
2121
+ | `mcp.tool.pricing_status` | span attr, metric label | **blend — library-computed, depends on tool + host** | which of three enum values depends on a tool-controlled model name resolving (or not) in a library-default-or-host-overridden pricing table |
2122
+ | `mcp.failure.fingerprint` / `.signature` / `.category` / `.origin` / `.channel` | span attr (2 also metric labels) | library-computed | hashes/enums — the governed core this project has kept clean since v0.4.0 |
2123
+ | `mcp.failure.error_class` | span attr | **blend — tool-controlled, capped only** | `err.name`, length-capped at 128 characters but, unlike everything hashed into the fingerprint, *not* pattern-scrubbed — a documented, deliberate tradeoff (`docs/known-gaps.md` entry 10) |
2124
+ | `mcp.failure.validation_paths` | span attr | mostly library-computed (gated) | one legacy Zod-issue-rendering path can still carry a tool-controlled dynamic object key, now redacted to a `<KEY>` placeholder rather than closed structurally (`docs/known-gaps.md` entry 10, item 3) |
2125
+ | `exception.type` (event) | event attr | tool-controlled | length-capped only under `'normalized'` mode; uncapped under `'full'` (the native SDK call) |
2126
+ | `exception.message` / `.stacktrace` (event), status message | event attr + status | **blend — layered** | raw tool content under `'full'`; this library's `normalizeMessage()`/`parseAndNormalizeStack()` transform of that same content under `'normalized'`; optionally a host-supplied `errorRecording.redactor`'s output feeding *that*, first (v0.14.0) — the most layered-provenance attribute in the system |
2127
+ | `mcp.tool.schema_drift_detected` / `mcp.tool.thrash_detected` | boolean span attr | library-computed | only ever set to `true`; omitted on a clean call |
2128
+ | `mcp.tool.schema_drift.type` | event attr, metric label | library-computed | classifier enum |
2129
+ | `mcp.tool.schema_drift.previous_hash` / `.current_hash` | event attr | library-computed | a hash *of* tool-controlled content (the tool's own `inputSchema`), not the schema itself |
2130
+ | `mcp.tool.schema_drift.added_fields` / `.removed_fields` / `.changed_fields` | event attr | **blend — tool-controlled, unbounded** | literal property names from a tool's own schema — unbounded across every tool anyone registers, unlike `argument_count` above, which is deliberately reduced to a number for exactly this reason |
2131
+ | `mcp.loop.length` / `.duration_ms` | event attr | library-computed | count / timestamp arithmetic |
2132
+ | `mcp.loop.wasted_tokens_in` / `_out` / `.wasted_cost_usd` | event attr | library-computed over tool-controlled | same aggregation/pricing chain as `mcp.tool.tokens.*`/`mcp.tool.cost.usd` above |
2133
+ | `mcp.loop.first_span_id` / `.first_trace_id` | event attr | SDK-supplied | OTel SDK-generated span/trace ids |
2134
+ | `mcp.loop.session_id` | event attr | **blend — depends on code path** | a real transport session id when one exists; a library-generated fallback UUID when it doesn't. See "Two keys whose provenance changes by code path" below |
2135
+ | `service.name` | resource attr | host-supplied | the `serviceName` config option (`setupNodeSdk: true` only) |
2136
+ | `mcp.pricing.default_table_last_verified` | resource attr | library-computed | constant, attached only when using the unmodified default pricing table |
2137
+ | `mcp.tool.outcome` | metric label only | library-computed | never a span attribute — `mcp.tool.duration`'s own outcome dimension |
2138
+
2139
+ **Two keys whose provenance changes by code path, not just by
2140
+ configuration:**
2141
+
2142
+ - **`error.type`** carries the tool's own `err.name` when the failure was
2143
+ a thrown exception, but a fixed library constant (`tool_error`) when
2144
+ the failure was a successful JSON-RPC response with `isError: true` —
2145
+ same attribute key, genuinely different source depending on which of
2146
+ the two failure paths produced it.
2147
+ - **`mcp.loop.session_id`** carries a real transport-assigned session id
2148
+ under a session-oriented transport, but a library-generated fallback
2149
+ UUID (`resolveThrashSessionId()`, `src/instrument.js`) under a
2150
+ transport this library can't positively confirm is session-oriented —
2151
+ see "Agent Thrash Detection" → sessionId fallback resolution above. The
2152
+ value is a real id either way; only *who* generated it differs.
2153
+
2154
+ **`mcp.server.name` / `mcp.server.version`** are defined in
2155
+ `src/attributes.js` but are not currently set anywhere in this codebase —
2156
+ dead constants, not a real attribute this library records today. Left
2157
+ out of the table above for that reason.
2158
+
2159
+ **Want to see the actual value landing in an attribute for your own
2160
+ traffic, not just its category?** Don't reach for a new debugging mode —
2161
+ the hooks that let you do this safely already exist, because *you*, not
2162
+ this library, know what's safe to look at in your own deployment:
2163
+
2164
+ - **`errorRecording.redactor`** (ADR 020, `docs/adr/020-redactor-hook.md`)
2165
+ sees the raw `message`/`stack` before this library's own scrubbing runs
2166
+ — a temporary, local redactor that copies what it receives to wherever
2167
+ *you've* decided is safe (your own log line, a breakpoint, a counter)
2168
+ shows you exactly what's landing in `exception.message`/`.stacktrace`
2169
+ for real calls, without this library needing to add a second channel
2170
+ for the same content.
2171
+ - **`costTracking.extractor`** (a `UsageExtractor`) sees the full tool
2172
+ result before token/model/cost attribution runs — the same trick works
2173
+ for `mcp.tool.tokens.*`/`mcp.tool.model`/`mcp.tool.cost.*`.
2174
+ - **A custom `Classifier`** (`opts.classifiers` on `computeFingerprint()`,
2175
+ `src/fingerprint/compose.js` — see "Deep-failure fingerprinting" →
2176
+ "Extending it" above) sees the coerced error before classification.
2177
+ Unlike the two hooks above, this one isn't yet wired through
2178
+ `instrumentMcpServer()`'s own options — using it today means calling
2179
+ `computeFingerprint()` directly rather than relying on the automatic
2180
+ per-call wrapping.
2181
+
2182
+ Each hook is scoped to exactly the content it already has access to; none
2183
+ of them create a new place for that content to reach — they're extension
2184
+ points this library already ships and already documents, not a bespoke
2185
+ audit feature.
2186
+
1883
2187
  ## Configuration
1884
2188
 
1885
2189
  All options passed to `instrumentMcpServer(server, options)`. Source of
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "opentel-mcp",
3
- "version": "0.13.0",
3
+ "version": "0.14.0",
4
4
  "description": "One-line OpenTelemetry instrumentation for Model Context Protocol (MCP) servers",
5
5
  "type": "module",
6
6
  "main": "src/index.js",