opentel-mcp 0.13.0 → 0.14.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +122 -0
- package/README.md +339 -35
- package/package.json +1 -1
- package/src/attributes.js +77 -4
- package/src/error-recording/config.js +84 -11
- package/src/error-recording/redactor.js +140 -0
- package/src/error-recording/types.d.ts +93 -12
- package/src/fingerprint/compose.js +10 -50
- package/src/fingerprint/normalize/exception.js +111 -0
- package/src/fingerprint/normalize/message.js +5 -1
- package/src/index.d.ts +31 -14
- package/src/instrument.js +118 -32
package/CHANGELOG.md
CHANGED
|
@@ -1,5 +1,127 @@
|
|
|
1
1
|
# Changelog
|
|
2
2
|
|
|
3
|
+
## 0.14.0
|
|
4
|
+
|
|
5
|
+
Adds `errorRecording.redactor`: a host-supplied hook for content
|
|
6
|
+
`normalizeMessage()`/`parseAndNormalizeStack()` don't recognize — a
|
|
7
|
+
proprietary API key format, an internal account id shape, a customer name
|
|
8
|
+
in prose or in a multi-tenant stack frame path. Full design: ADR 020
|
|
9
|
+
(`docs/adr/020-redactor-hook.md`). Ships in the same release as the
|
|
10
|
+
0.13.1 fix/docs items below (0.13.1 itself was never tagged/published —
|
|
11
|
+
see that section's own note), but is a logically separate, purely
|
|
12
|
+
additive change from both of them; kept in its own section here rather
|
|
13
|
+
than folded into 0.13.1's for that reason.
|
|
14
|
+
|
|
15
|
+
### Added — `errorRecording.redactor` hook (ADR 020)
|
|
16
|
+
|
|
17
|
+
- New optional `errorRecording.redactor` field: a synchronous
|
|
18
|
+
`(input: { message, stack }) => { message, stack }` function, consulted
|
|
19
|
+
only under `mode: 'normalized'`, that runs on the raw, uncoerced
|
|
20
|
+
`message`/`stack` — via the same coercion this library's own pipeline
|
|
21
|
+
already uses — **before** `normalizeMessage()`/`parseAndNormalizeStack()`
|
|
22
|
+
ever see them, never the reverse. This library's own scrubbing still
|
|
23
|
+
runs second, over your redactor's output, as a defense-in-depth
|
|
24
|
+
backstop.
|
|
25
|
+
- Configuring a redactor alongside `mode: 'full'` or `'none'` is accepted
|
|
26
|
+
but never invoked — those two modes' entire meaning is "byte-identical,
|
|
27
|
+
regardless of what else is configured" — and produces a one-time
|
|
28
|
+
`diag.warn()` at `instrumentMcpServer()` setup naming the no-op, since
|
|
29
|
+
forgetting to also flip `mode` is an easy, otherwise-silent
|
|
30
|
+
misconfiguration.
|
|
31
|
+
- A redactor that throws, returns a non-string `message`, or returns a
|
|
32
|
+
`stack` that's neither a string nor `undefined` falls back to
|
|
33
|
+
`'none'`-equivalent span output for that one event — status set, no
|
|
34
|
+
`exception` event — **never** to raw/unredacted content, plus a
|
|
35
|
+
one-time, content-free `diag.warn()` naming the failure shape (never
|
|
36
|
+
the message/stack content itself). The returned `message`/`stack` are
|
|
37
|
+
also length-capped defensively before use (reusing
|
|
38
|
+
`normalizeMessage()`'s existing 2048-character limit).
|
|
39
|
+
- **Never reaches `computeFingerprint()`'s hash input, in any
|
|
40
|
+
configuration.** `mcp.failure.fingerprint`/`signature` are computed
|
|
41
|
+
from the real, unmodified thrown error, identically whether or not a
|
|
42
|
+
redactor is configured — a deliberate, non-negotiable design decision
|
|
43
|
+
(ADR 020 Decision 3): the fingerprint is a SHA-256 hash, never a
|
|
44
|
+
plaintext channel, so redacting it would buy no privacy while tying
|
|
45
|
+
fingerprint stability to unversioned host code. Any existing dashboard
|
|
46
|
+
filter or alert keyed on `mcp.failure.fingerprint` keeps working
|
|
47
|
+
unchanged after adding a redactor.
|
|
48
|
+
- **Timing caveat, not just a telemetry one:** the redactor runs
|
|
49
|
+
synchronously, inline, on the error path, before the span ends — there
|
|
50
|
+
is no timeout, and JavaScript's single-threaded execution model means
|
|
51
|
+
one can't be added without a `worker_threads`/`vm.Script` boundary this
|
|
52
|
+
feature deliberately doesn't pay for. A slow redactor (catastrophic
|
|
53
|
+
regex backtracking, an accidental blocking call) adds directly to that
|
|
54
|
+
tool call's own response latency, not just to what shows up in traces.
|
|
55
|
+
See the README's "The redactor hook" section, "Timing," for the full
|
|
56
|
+
writeup, and ADR 020's own "Timing" section for why this is an accepted
|
|
57
|
+
risk rather than a solved one.
|
|
58
|
+
- No environment-variable equivalent — a function can't be expressed as
|
|
59
|
+
an `OTEL_MCP_*` string, same as `costTracking.extractor`/
|
|
60
|
+
`thrashDetection`'s classifier-shaped options already are.
|
|
61
|
+
- New exported types `ErrorRecordingRedactor` and
|
|
62
|
+
`ErrorRecordingRedactorFields` (`src/error-recording/types.d.ts`,
|
|
63
|
+
re-exported from the package root) — a consumer can type a redactor
|
|
64
|
+
function on its own, the same way `UsageExtractor`/`Classifier` already
|
|
65
|
+
let a consumer type an extractor/classifier independently of the option
|
|
66
|
+
object that carries it.
|
|
67
|
+
- Purely additive and default-off: a deployment that never sets
|
|
68
|
+
`errorRecording.redactor` observes zero behavior change.
|
|
69
|
+
|
|
70
|
+
## 0.13.1
|
|
71
|
+
|
|
72
|
+
Never tagged or published as its own release — ships bundled into
|
|
73
|
+
0.14.0 above instead, since the redactor work started before this patch
|
|
74
|
+
went out the door. Kept as its own section here (rather than merged into
|
|
75
|
+
0.14.0's) because both items below are logically independent of the
|
|
76
|
+
redactor: one's a bug fix to existing `'normalized'`-mode behavior, the
|
|
77
|
+
other's a documentation-only correction, and neither has anything to do
|
|
78
|
+
with the new hook.
|
|
79
|
+
|
|
80
|
+
### Fixed — `'normalized'` mode dropped `exception.message`/status message for non-`Error` throws
|
|
81
|
+
|
|
82
|
+
- Before this fix, `recordThrownException()`'s `'normalized'` branch
|
|
83
|
+
independently re-derived `message`/`stack` via a plain `err?.message`/
|
|
84
|
+
`err?.stack` read — which silently returned `undefined` for a thrown
|
|
85
|
+
**string** or a plain non-`Error` object (neither has a `.message`
|
|
86
|
+
property the way a real `Error` instance does), while
|
|
87
|
+
`computeFingerprint()`'s own, richer `coerceError()` correctly
|
|
88
|
+
recognized those same shapes and still produced a real
|
|
89
|
+
`mcp.failure.fingerprint`. Net effect: a thrown string or non-`Error`
|
|
90
|
+
throw got real fingerprinting but a completely empty `exception.message`
|
|
91
|
+
/ status message on the span under `'normalized'` mode — same `err`,
|
|
92
|
+
two independent readers, two different answers, silently.
|
|
93
|
+
- Fixed by extracting the shared `normalizeException()`/`coerceError()`
|
|
94
|
+
pair into their own module (`src/fingerprint/normalize/exception.js`)
|
|
95
|
+
and routing **both** `computeFingerprint()` and `recordThrownException()`
|
|
96
|
+
through the same one computation for the same `err` — the span's
|
|
97
|
+
exception content and the fingerprint's hashed inputs can no longer
|
|
98
|
+
independently drift apart, by construction, not by convention. This is
|
|
99
|
+
the same consolidation ADR 020 (the redactor's own design doc, above)
|
|
100
|
+
repeatedly cites as precedent for keeping the fingerprint path
|
|
101
|
+
structurally isolated from the span-writing path going forward.
|
|
102
|
+
- `'full'` and `'none'` modes were never affected — this bug was specific
|
|
103
|
+
to `'normalized'` mode's own inline coercion.
|
|
104
|
+
|
|
105
|
+
### Documentation — `errorRecording.mode` scoping correction
|
|
106
|
+
|
|
107
|
+
- Clarified that `errorRecording.mode` controls only the span this
|
|
108
|
+
library creates for the current `tools/call`/`tools/list` — it has no
|
|
109
|
+
effect on exception content any other instrumentation in the same
|
|
110
|
+
process (an APM agent, HTTP or framework auto-instrumentation, anything
|
|
111
|
+
else wrapping the handler) independently records onto its own span for
|
|
112
|
+
the same rethrown error. Since no mode mutates the original `err`, that
|
|
113
|
+
raw content can still land in the same trace, one span up, regardless
|
|
114
|
+
of mode — including `'none'`, which was previously worded in a way that
|
|
115
|
+
could read as a trace-wide guarantee rather than a span-scoped one.
|
|
116
|
+
Confirmed empirically (a real ambient context manager, an outer span
|
|
117
|
+
recording the rethrown error, `InMemorySpanExporter`) — see
|
|
118
|
+
`docs/known-gaps.md` entry 10's 2026-09-01 update for the full
|
|
119
|
+
reproduction and reasoning. README's "Error recording" and "What this
|
|
120
|
+
library records" sections updated accordingly. No code changed — the
|
|
121
|
+
underlying behavior (and the exposure at the default `'full'` mode) was
|
|
122
|
+
already accurate and unchanged; only the description of what
|
|
123
|
+
`errorRecording.mode` scopes to was incomplete.
|
|
124
|
+
|
|
3
125
|
## 0.13.0
|
|
4
126
|
|
|
5
127
|
Closes the two open items from `docs/known-gaps.md` entry 10 (a
|
package/README.md
CHANGED
|
@@ -463,12 +463,26 @@ per-call-site wrapping; tracked in the roadmap below.
|
|
|
463
463
|
|
|
464
464
|
## Error recording (v0.13.0+)
|
|
465
465
|
|
|
466
|
-
|
|
467
|
-
`
|
|
468
|
-
|
|
469
|
-
|
|
470
|
-
|
|
471
|
-
|
|
466
|
+
**What this controls, and who needs it.** Whenever a `tools/call`/
|
|
467
|
+
`tools/list` handler throws, this library records the thrown error's
|
|
468
|
+
message and stack trace onto the span it creates for that call — by
|
|
469
|
+
default, exactly as the tool wrote them, no redaction. If your tools call
|
|
470
|
+
third-party services, wrap unaudited dependencies, or you otherwise don't
|
|
471
|
+
control every string a tool might throw, that's a real channel for
|
|
472
|
+
whatever those errors happen to contain (an email address, an internal
|
|
473
|
+
hostname, a connection string a driver embedded in its own error message)
|
|
474
|
+
to reach wherever your traces get exported. `errorRecording.mode` lets
|
|
475
|
+
you scrub or drop that content instead of recording it verbatim — read
|
|
476
|
+
both caveats below before treating it as a privacy control, though: it
|
|
477
|
+
has two distinct limits, and neither is obvious from the option name
|
|
478
|
+
alone.
|
|
479
|
+
|
|
480
|
+
Mechanically: every thrown error goes through `span.recordException(err)`
|
|
481
|
+
(an OpenTelemetry SDK method, not one of this library's own attributes)
|
|
482
|
+
plus `span.setStatus({ code: ERROR, message: err.message })` —
|
|
483
|
+
unconditionally, whether or not `fingerprinting` is enabled.
|
|
484
|
+
`errorRecording.mode` controls what those two calls actually put on the
|
|
485
|
+
span:
|
|
472
486
|
|
|
473
487
|
| Mode | `exception.message` / status message | `exception.stacktrace` | When to use |
|
|
474
488
|
|---|---|---|---|
|
|
@@ -476,30 +490,197 @@ span exactly as thrown. `errorRecording.mode` controls this:
|
|
|
476
490
|
| `'normalized'` | `normalizeMessage(err.message)` — the exact scrubbing pipeline (`src/fingerprint/normalize/message.js`) fingerprinting already runs before hashing: UUIDs, emails, URLs, IPs, timestamps, filesystem paths, hex runs, quoted ids | Reconstructed from `parseAndNormalizeStack()` (`src/fingerprint/normalize/stack.js`) — keeps every function name/file/line, strips only the local `cwd` prefix (and collapses `node_modules` package versions) | Tool results come from third-party or unaudited MCP servers and you want the same scrubbing fingerprinting already trusts, applied to the raw exception content too |
|
|
477
491
|
| `'none'` | Not set — `span.setStatus({ code: ERROR })` with no message, the same pattern already used for tool-level `isError: true` failures | Not set | You rely entirely on `mcp.failure.*` (category/fingerprint/signature — already hashed/normalized) and don't want any free-text exception content on the span at all |
|
|
478
492
|
|
|
479
|
-
|
|
480
|
-
|
|
481
|
-
|
|
482
|
-
`error.type`/`exception.type` (`err.name`) is capped at 128 characters
|
|
483
|
-
unconditionally in every mode — the same cap `mcp.failure.error_class`
|
|
484
|
-
uses — since a length cap on a class-identifier field costs a
|
|
485
|
-
well-behaved tool nothing, unlike message/stack content.
|
|
486
|
-
|
|
487
|
-
**`'normalized'` is targeted scrubbing, not general-purpose redaction —
|
|
488
|
-
read this before treating it as a PII filter.** `normalizeMessage()`
|
|
489
|
-
matches specific, structured shapes: UUIDs, email addresses, URLs,
|
|
493
|
+
**Caveat 1 — `'normalized'` is targeted scrubbing, not general-purpose
|
|
494
|
+
redaction. Do not treat it as a PII filter.** `normalizeMessage()`
|
|
495
|
+
matches specific, structured shapes only: UUIDs, email addresses, URLs,
|
|
490
496
|
IPv4/IPv6 addresses, ISO-8601/Unix timestamps, filesystem paths, long hex
|
|
491
497
|
runs, and quoted alphanumeric ids (8–64 chars, mixed letters/digits). It
|
|
492
|
-
does not recognize sensitive content in general
|
|
498
|
+
does not recognize sensitive content in general — an API key in a format
|
|
493
499
|
none of those patterns match (a bare, unquoted token with no digit in it,
|
|
494
500
|
or a custom prefix scheme), or a customer's name embedded in ordinary
|
|
495
501
|
prose ("could not process request for Jane Smith"), passes through
|
|
496
|
-
`'normalized'` mode completely unchanged
|
|
497
|
-
|
|
502
|
+
`'normalized'` mode completely unchanged, identical to what `'full'`
|
|
503
|
+
would put on the span. Treat `'normalized'` as "the same scrubbing
|
|
498
504
|
fingerprinting already trusts for hashing," not as a guarantee that
|
|
499
|
-
whatever a tool's error messages contain is safe to record
|
|
500
|
-
|
|
501
|
-
|
|
502
|
-
|
|
505
|
+
whatever a tool's error messages contain is safe to record.
|
|
506
|
+
|
|
507
|
+
**Example — the same thrown error, `'full'` vs. `'normalized'`:**
|
|
508
|
+
|
|
509
|
+
```js
|
|
510
|
+
throw new Error('upstream lookup failed for user jane.doe@example.com');
|
|
511
|
+
```
|
|
512
|
+
|
|
513
|
+
| Mode | `exception.message` on the span |
|
|
514
|
+
|---|---|
|
|
515
|
+
| `'full'` | `upstream lookup failed for user jane.doe@example.com` |
|
|
516
|
+
| `'normalized'` | `upstream lookup failed for user <EMAIL>` |
|
|
517
|
+
| `'none'` | *(not set — only the ERROR status code is)* |
|
|
518
|
+
|
|
519
|
+
The stack trace scrubs the same way: `'full'` keeps every frame's
|
|
520
|
+
absolute path; `'normalized'` keeps every frame (function name, file,
|
|
521
|
+
line) but strips the local `cwd` prefix.
|
|
522
|
+
|
|
523
|
+
### The redactor hook (v0.14.0+)
|
|
524
|
+
|
|
525
|
+
Caveat 1 above is the reason this exists: `normalizeMessage()` matches
|
|
526
|
+
eight specific, structured shapes (UUIDs, emails, URLs, IPv4/IPv6,
|
|
527
|
+
timestamps, paths, hex runs, quoted ids), and that is where its
|
|
528
|
+
competence ends. An API key in your own proprietary format
|
|
529
|
+
(`ACME_KEY_[a-f0-9]{8}`), an internal account id shape
|
|
530
|
+
(`ACCT-\d{9}`), or a customer's name embedded in ordinary prose are not
|
|
531
|
+
structurally distinguishable from ordinary text by any pattern this
|
|
532
|
+
library could ship — no ninth built-in pattern fixes that. Only the host
|
|
533
|
+
running the tool knows what their own data looks like.
|
|
534
|
+
`errorRecording.redactor` is a synchronous function you supply that runs
|
|
535
|
+
as part of `mode: 'normalized'`, so you can close that gap without
|
|
536
|
+
forking this library or falling back to `'none'`'s blunt "drop
|
|
537
|
+
everything." See ADR 020 (`docs/adr/020-redactor-hook.md`) for the full
|
|
538
|
+
design rationale; the operator-facing behavior is summarized below.
|
|
539
|
+
|
|
540
|
+
**Worked example:**
|
|
541
|
+
|
|
542
|
+
```js
|
|
543
|
+
instrumentMcpServer(server, {
|
|
544
|
+
errorRecording: {
|
|
545
|
+
mode: 'normalized',
|
|
546
|
+
redactor: ({ message, stack }) => ({
|
|
547
|
+
message: message
|
|
548
|
+
.replace(/ACME_KEY_[a-f0-9]{8}/g, '[API_KEY]')
|
|
549
|
+
.replace(/ACCT-\d{9}/g, '[ACCOUNT_ID]'),
|
|
550
|
+
stack, // unchanged — this redactor only needs to touch the message
|
|
551
|
+
}),
|
|
552
|
+
},
|
|
553
|
+
});
|
|
554
|
+
```
|
|
555
|
+
|
|
556
|
+
```js
|
|
557
|
+
throw new Error(
|
|
558
|
+
'auth failed for ACME_KEY_7f3a9c2e, ' +
|
|
559
|
+
'account ACCT-123456789, contact jane.doe@example.com',
|
|
560
|
+
);
|
|
561
|
+
```
|
|
562
|
+
|
|
563
|
+
| Step | `exception.message` |
|
|
564
|
+
|---|---|
|
|
565
|
+
| Raw, as thrown | `auth failed for ACME_KEY_7f3a9c2e, account ACCT-123456789, contact jane.doe@example.com` |
|
|
566
|
+
| After your redactor (runs first, on the raw text) | `auth failed for [API_KEY], account [ACCOUNT_ID], contact jane.doe@example.com` |
|
|
567
|
+
| After this library's own patterns (run second, on YOUR output) | `auth failed for [API_KEY], account [ACCOUNT_ID], contact <EMAIL>` |
|
|
568
|
+
|
|
569
|
+
**Ordering matters, and it's fixed, not configurable:** your redactor
|
|
570
|
+
always runs first, on the raw, uncoerced `message`/`stack` — before
|
|
571
|
+
`normalizeMessage()`/`parseAndNormalizeStack()` ever see them, never the
|
|
572
|
+
reverse. Two consequences follow directly: (1) this library's own
|
|
573
|
+
scrubbing is still the last thing that touches the content before it
|
|
574
|
+
reaches the span — a defense-in-depth backstop if your redactor misses a
|
|
575
|
+
shape its own patterns don't cover (see the API key example above:
|
|
576
|
+
notice `<EMAIL>` still gets scrubbed even though your redactor never
|
|
577
|
+
mentioned email addresses), and (2) your redactor sees the full,
|
|
578
|
+
untruncated original message — `normalizeMessage()`'s 2048-character
|
|
579
|
+
truncation applies afterward, to your output, not before you ever see the
|
|
580
|
+
content.
|
|
581
|
+
|
|
582
|
+
**Only consulted under `mode: 'normalized'`.** Configure a redactor
|
|
583
|
+
alongside `'full'` or `'none'` and it's accepted but never called — those
|
|
584
|
+
two modes' entire meaning is "byte-identical to their own fixed
|
|
585
|
+
behavior, regardless of what else is configured." Since forgetting to
|
|
586
|
+
also flip `mode` away from its `'full'` default is an easy way to end up
|
|
587
|
+
with a redactor that's silently never invoked, `instrumentMcpServer()`
|
|
588
|
+
logs a one-time `diag.warn()` at setup naming exactly that.
|
|
589
|
+
|
|
590
|
+
**Failure never falls back to raw content.** If your redactor throws,
|
|
591
|
+
returns a non-string `message`, or returns a `stack` that's neither a
|
|
592
|
+
string nor `undefined`, that one event falls back to `'none'`-equivalent
|
|
593
|
+
output — `span.setStatus({ code: ERROR })` only, no `exception` event at
|
|
594
|
+
all — never to `'full'`-equivalent (raw, unredacted) content. A redactor
|
|
595
|
+
you configured specifically to keep something off the span must never
|
|
596
|
+
silently fail open; the safe direction here is *more* redacted, not
|
|
597
|
+
less. This fires a one-time, content-free `diag.warn()` naming the
|
|
598
|
+
failure shape (`threw` / `returned a non-string message` / `returned an
|
|
599
|
+
invalid stack`) — never the message/stack content that triggered it,
|
|
600
|
+
so the warning itself can't become a second, undocumented leak channel.
|
|
601
|
+
Your redactor's returned `message`/`stack` are also length-capped
|
|
602
|
+
defensively before use (reusing `normalizeMessage()`'s own 2048-character
|
|
603
|
+
limit), regardless of what a well-behaved redactor is expected to
|
|
604
|
+
return.
|
|
605
|
+
|
|
606
|
+
**Timing — read this before you write a regex.** The redactor runs
|
|
607
|
+
*synchronously*, inline, on the error path — inside the still-open
|
|
608
|
+
span's own callback, before `span.end()` is called. There is no timeout,
|
|
609
|
+
and there cannot be one: JavaScript is single-threaded, and nothing in
|
|
610
|
+
this library (or in Node itself, short of a `worker_threads` boundary
|
|
611
|
+
this feature deliberately doesn't pay for) can interrupt a synchronous
|
|
612
|
+
function that's already running. **This means a slow redactor doesn't
|
|
613
|
+
just delay telemetry — it delays the tool call's own response to
|
|
614
|
+
whoever's waiting on it.** Catastrophic backtracking in a regex (the
|
|
615
|
+
classic `(a+)+$`-shaped footgun), an accidental synchronous file read, or
|
|
616
|
+
any blocking call inside your redactor adds directly to that tool call's
|
|
617
|
+
latency, on every single call that throws under `'normalized'` mode,
|
|
618
|
+
for as long as the redactor stays slow. Test your redactor's regexes
|
|
619
|
+
against adversarial input the same way you would for any other
|
|
620
|
+
user-facing regex, not just against the examples you had in mind when
|
|
621
|
+
you wrote it.
|
|
622
|
+
|
|
623
|
+
**The fingerprint is completely unaffected — on purpose, always.**
|
|
624
|
+
`mcp.failure.fingerprint`/`mcp.failure.signature` and the values hashed
|
|
625
|
+
into them are computed from the real, unmodified thrown error, identically
|
|
626
|
+
whether or not a redactor is configured. A redactor has no way to reach
|
|
627
|
+
that computation at all — it only ever touches what lands on the span.
|
|
628
|
+
This is deliberate: the fingerprint is a SHA-256 hash, never a plaintext
|
|
629
|
+
channel, so redacting it would buy no privacy while tying fingerprint
|
|
630
|
+
stability to your own, unversioned redactor code. Practically, this means
|
|
631
|
+
any saved dashboard filter or alert keyed on `mcp.failure.fingerprint`
|
|
632
|
+
keeps working exactly as before the moment you add a redactor — nothing
|
|
633
|
+
about adding or changing one ever reshuffles a fingerprint value.
|
|
634
|
+
|
|
635
|
+
The redactor is still scoped to this library's own span, the same as
|
|
636
|
+
every other `mode` value — see Caveat 2 below for what that does and
|
|
637
|
+
doesn't cover; nothing about the redactor changes that boundary.
|
|
638
|
+
|
|
639
|
+
**Caveat 2 — every mode is scoped to the span this library creates, not
|
|
640
|
+
to your trace as a whole.** No mode mutates the original `err` — both
|
|
641
|
+
call sites rethrow it afterward, so `'normalized'`/`'none'` build the
|
|
642
|
+
exception event independently rather than editing `err.message`/
|
|
643
|
+
`err.stack` in place. The consequence: the unmodified, fully raw `err` is
|
|
644
|
+
exactly what propagates to whatever called this handler. If any other
|
|
645
|
+
instrumentation further out in the call stack — an APM agent, HTTP or
|
|
646
|
+
framework auto-instrumentation, anything else wrapping this handler —
|
|
647
|
+
independently calls `recordException(err)`/`setStatus({ message:
|
|
648
|
+
err.message })` on its OWN span when it observes the rethrow, the raw,
|
|
649
|
+
unscrubbed message and stack land in the SAME TRACE as this library's
|
|
650
|
+
(scrubbed) span, one span up, not on it. `'none'` buys no more protection
|
|
651
|
+
here than `'normalized'`: neither mutates `err`, so an outer recorder
|
|
652
|
+
sees the identical raw content either way. Someone choosing `'none'`
|
|
653
|
+
specifically to keep sensitive content out of their tracing backend needs
|
|
654
|
+
to know that protection stops at this library's own span, not at the
|
|
655
|
+
trace boundary.
|
|
656
|
+
|
|
657
|
+
Confirmed directly, not just reasoned about: an instrumented server run
|
|
658
|
+
under a real ambient context manager (`AsyncLocalStorageContextManager` —
|
|
659
|
+
the propagation mechanism production deployments actually use, unlike
|
|
660
|
+
this project's own test suite, which registers most test files with
|
|
661
|
+
`contextManager: null` and therefore can't exercise this path — see
|
|
662
|
+
`docs/known-gaps.md` entry 10's latest update for why that's fine for
|
|
663
|
+
those specific tests), wrapped in an outer span that calls
|
|
664
|
+
`recordException(err)` on catch, with `errorRecording.mode: 'normalized'`
|
|
665
|
+
set and an `InMemorySpanExporter` attached: the tool span's exception
|
|
666
|
+
event carries the scrubbed message, and the ancestor span's exception
|
|
667
|
+
event and status message both carry the original, raw one — same trace
|
|
668
|
+
ID, parent/child span relationship confirmed, no mutation anywhere in
|
|
669
|
+
between.
|
|
670
|
+
|
|
671
|
+
**This is not fixable from inside this library.** The only way to close
|
|
672
|
+
it here would be mutating `err.message`/`err.stack` in place before
|
|
673
|
+
rethrowing — which would corrupt whatever error-handling the host
|
|
674
|
+
application does with that same object afterward, a correctness hazard
|
|
675
|
+
this library isn't willing to trade for a telemetry one. It has authority
|
|
676
|
+
over one span; it has no visibility into, and no control over, what other
|
|
677
|
+
instrumentation attached to the same process does with the error once
|
|
678
|
+
rethrown. See `docs/known-gaps.md` entry 10 for the full writeup.
|
|
679
|
+
|
|
680
|
+
`error.type`/`exception.type` (`err.name`) is capped at 128 characters
|
|
681
|
+
unconditionally in every mode — the same cap `mcp.failure.error_class`
|
|
682
|
+
uses — since a length cap on a class-identifier field costs a
|
|
683
|
+
well-behaved tool nothing, unlike message/stack content.
|
|
503
684
|
|
|
504
685
|
**Default stays `'full'` through all of `0.x`.** Changing it would alter
|
|
505
686
|
what every trace backend renders for the single highest-traffic failure
|
|
@@ -513,6 +694,7 @@ for the full argument, including why the default is expected to flip to
|
|
|
513
694
|
| Option | Env var | Type | Default | Description |
|
|
514
695
|
|---|---|---|---|---|
|
|
515
696
|
| `mode` | `OTEL_MCP_ERROR_RECORDING_MODE` | `'full'` \| `'normalized'` \| `'none'` | `'full'` | See table above. An unrecognized value falls back to `'full'` silently, same as every other `OTEL_MCP_*` env var |
|
|
697
|
+
| `redactor` | *(none — a function can't be expressed as an env var string)* | `(input: { message: string, stack: string \| undefined }) => { message: string, stack: string \| undefined }` | `undefined` | See "The redactor hook" above. Only consulted when `mode` is `'normalized'`; a non-function value is treated as absent, silently |
|
|
516
698
|
|
|
517
699
|
## Cost & Token Attribution (v0.5.0)
|
|
518
700
|
|
|
@@ -1863,23 +2045,145 @@ sensitive information" and still specify recording it by default —
|
|
|
1863
2045
|
every other OTel-instrumented library sharing the same trace (an HTTP
|
|
1864
2046
|
client, a DB driver, a queue consumer) records `err.message`/`err.stack`
|
|
1865
2047
|
the same way, unscrubbed, through the same `recordException()` call.
|
|
1866
|
-
|
|
1867
|
-
|
|
1868
|
-
|
|
1869
|
-
|
|
1870
|
-
|
|
1871
|
-
|
|
1872
|
-
|
|
1873
|
-
|
|
1874
|
-
|
|
1875
|
-
|
|
1876
|
-
|
|
2048
|
+
That covers a sharper case than "different libraries recording their own,
|
|
2049
|
+
separate operations," too: an outer wrapper — an APM agent's own request
|
|
2050
|
+
span, or HTTP/framework auto-instrumentation sitting further out in the
|
|
2051
|
+
call stack — independently recording the very same error this library
|
|
2052
|
+
just scrubbed. No `errorRecording.mode` mutates the original `err` (see
|
|
2053
|
+
"Error recording" above), so a rethrown error is exactly what such a
|
|
2054
|
+
wrapper sees; if it calls `recordException(err)` on its own span, the raw
|
|
2055
|
+
content lands one span up in the identical trace, regardless of which
|
|
2056
|
+
mode is set here. opentel-mcp matches that ecosystem default rather than
|
|
2057
|
+
silently diverging from it. What it adds on top: `errorRecording.mode`
|
|
2058
|
+
lets you choose `'normalized'` (targeted scrubbing — the same structured
|
|
2059
|
+
patterns fingerprinting matches: UUIDs, emails, URLs, IPs, timestamps,
|
|
2060
|
+
paths, quoted ids; **not** a general-purpose PII filter, so freeform
|
|
2061
|
+
sensitive text — a customer's name in prose, an API key in a format none
|
|
2062
|
+
of those patterns match — still reaches the span unchanged, see "Error
|
|
2063
|
+
recording" above) or `'none'` (record neither) instead of `'full'` — a
|
|
2064
|
+
level of operator control most peer instrumentations don't offer for this
|
|
2065
|
+
path, scoped to the span this library creates, not a trace-wide
|
|
2066
|
+
guarantee. Full reasoning, including why the default doesn't change in
|
|
2067
|
+
this release: ADR 019 (`docs/adr/019-raw-content-on-spans.md`).
|
|
1877
2068
|
|
|
1878
2069
|
This is a data-handling property of an observability library, not a
|
|
1879
2070
|
security control opentel-mcp is claiming to provide — it doesn't
|
|
1880
2071
|
authenticate, encrypt, or restrict who can read your traces; that's your
|
|
1881
2072
|
tracing backend's job.
|
|
1882
2073
|
|
|
2074
|
+
### Full attribute provenance
|
|
2075
|
+
|
|
2076
|
+
The paragraphs above cover the three channels with real content risk;
|
|
2077
|
+
the table below is the complete picture — every attribute this library
|
|
2078
|
+
sets, anywhere, classified by where its value actually comes from. It was
|
|
2079
|
+
produced by tracing every `span.setAttribute()`/`setAttributes()`/
|
|
2080
|
+
`addEvent()` call site back to its source, generalizing the manual audits
|
|
2081
|
+
`docs/known-gaps.md` entries 10 and 11 already did by hand for the
|
|
2082
|
+
riskiest subset.
|
|
2083
|
+
|
|
2084
|
+
**Why a table instead of a runtime "audit mode":** provenance here is a
|
|
2085
|
+
property of which code path sets an attribute, not of any particular
|
|
2086
|
+
call's runtime state — it's the same answer on every call, for a given
|
|
2087
|
+
release, until the code setting that attribute changes. A live inspection
|
|
2088
|
+
feature can't tell you anything about *category* this table doesn't
|
|
2089
|
+
already say, for free, with no new code and no new place for a value to
|
|
2090
|
+
leak. Where a table genuinely falls short — "what's the *actual* value
|
|
2091
|
+
landing in this attribute for my traffic" — the fix isn't a new library
|
|
2092
|
+
feature; it's the extension hooks documented below, which already exist
|
|
2093
|
+
and were built for exactly this. "Only the host knows what's safe to look
|
|
2094
|
+
at in their own deployment" is the same principle `errorRecording.redactor`
|
|
2095
|
+
itself is built around (ADR 020) — reused here, not re-solved.
|
|
2096
|
+
|
|
2097
|
+
Four categories: **library-computed** (a hash, count, enum, or fixed
|
|
2098
|
+
constant — the safe majority), **host-supplied** (an `instrumentMcpServer()`
|
|
2099
|
+
config value), **SDK-supplied** (a transport/protocol id — session,
|
|
2100
|
+
request, span, trace), and **tool-controlled** (content from a tool's own
|
|
2101
|
+
result or the caller's own arguments — the category every raw-content
|
|
2102
|
+
finding in `docs/known-gaps.md` traces back to). Eight rows below are
|
|
2103
|
+
genuine **blends** — a single word would mislead for these specifically,
|
|
2104
|
+
so they're spelled out instead of forced into one bucket.
|
|
2105
|
+
|
|
2106
|
+
| Attribute | Surface | Provenance | Note |
|
|
2107
|
+
|---|---|---|---|
|
|
2108
|
+
| `mcp.method.name` | span attr | library-computed | fixed string per handler |
|
|
2109
|
+
| `gen_ai.operation.name` | span attr | library-computed | fixed string `"execute_tool"` |
|
|
2110
|
+
| `gen_ai.tool.name` | span attr, metric label | **blend — tool-controlled** | `request.params.name`, unvalidated against the registered tool set (`docs/known-gaps.md` entry 11) |
|
|
2111
|
+
| `jsonrpc.request.id` | span attr | SDK-supplied | transport/protocol correlation id |
|
|
2112
|
+
| `mcp.tool.argument_count` | span attr | tool-controlled | a *count* of `request.params.arguments` keys, deliberately — never the argument values themselves |
|
|
2113
|
+
| `error.type` | span attr | **blend — depends on code path** | the tool's own `err.name` on a thrown failure; the fixed constant `tool_error` on an `isError: true` failure — same key, two different sources depending on which failure path fired. See "Two keys whose provenance changes by code path" below |
|
|
2114
|
+
| `gen_ai.response.model` | span attr | tool-controlled | co-emitted from the same value as `mcp.tool.model`, for dashboard compatibility |
|
|
2115
|
+
| `mcp.tool.tokens.input` / `.output` | span attr | tool-controlled | the tool result's own usage numbers |
|
|
2116
|
+
| `mcp.tool.tokens.total` | span attr | library-computed | sum of the two tool-controlled values above |
|
|
2117
|
+
| `mcp.tool.model` | span attr, metric label | tool-controlled | a tool result's own declared model field, echoed verbatim (`docs/known-gaps.md` entry 10, item 2 — still open) |
|
|
2118
|
+
| `mcp.tool.cost.usd` | span attr | **blend — library-computed over tool-controlled × config** | tool-controlled token counts, multiplied by a pricing table that's either this library's default or a host override |
|
|
2119
|
+
| `mcp.tool.cost.currency` | span attr | library-computed | fixed constant `"USD"` |
|
|
2120
|
+
| `mcp.tool.cost.budget_exceeded` / `.budget_scope` | span attr | library-computed | comparison against a host-supplied budget threshold |
|
|
2121
|
+
| `mcp.tool.pricing_status` | span attr, metric label | **blend — library-computed, depends on tool + host** | which of three enum values depends on a tool-controlled model name resolving (or not) in a library-default-or-host-overridden pricing table |
|
|
2122
|
+
| `mcp.failure.fingerprint` / `.signature` / `.category` / `.origin` / `.channel` | span attr (2 also metric labels) | library-computed | hashes/enums — the governed core this project has kept clean since v0.4.0 |
|
|
2123
|
+
| `mcp.failure.error_class` | span attr | **blend — tool-controlled, capped only** | `err.name`, length-capped at 128 characters but, unlike everything hashed into the fingerprint, *not* pattern-scrubbed — a documented, deliberate tradeoff (`docs/known-gaps.md` entry 10) |
|
|
2124
|
+
| `mcp.failure.validation_paths` | span attr | mostly library-computed (gated) | one legacy Zod-issue-rendering path can still carry a tool-controlled dynamic object key, now redacted to a `<KEY>` placeholder rather than closed structurally (`docs/known-gaps.md` entry 10, item 3) |
|
|
2125
|
+
| `exception.type` (event) | event attr | tool-controlled | length-capped only under `'normalized'` mode; uncapped under `'full'` (the native SDK call) |
|
|
2126
|
+
| `exception.message` / `.stacktrace` (event), status message | event attr + status | **blend — layered** | raw tool content under `'full'`; this library's `normalizeMessage()`/`parseAndNormalizeStack()` transform of that same content under `'normalized'`; optionally a host-supplied `errorRecording.redactor`'s output feeding *that*, first (v0.14.0) — the most layered-provenance attribute in the system |
|
|
2127
|
+
| `mcp.tool.schema_drift_detected` / `mcp.tool.thrash_detected` | boolean span attr | library-computed | only ever set to `true`; omitted on a clean call |
|
|
2128
|
+
| `mcp.tool.schema_drift.type` | event attr, metric label | library-computed | classifier enum |
|
|
2129
|
+
| `mcp.tool.schema_drift.previous_hash` / `.current_hash` | event attr | library-computed | a hash *of* tool-controlled content (the tool's own `inputSchema`), not the schema itself |
|
|
2130
|
+
| `mcp.tool.schema_drift.added_fields` / `.removed_fields` / `.changed_fields` | event attr | **blend — tool-controlled, unbounded** | literal property names from a tool's own schema — unbounded across every tool anyone registers, unlike `argument_count` above, which is deliberately reduced to a number for exactly this reason |
|
|
2131
|
+
| `mcp.loop.length` / `.duration_ms` | event attr | library-computed | count / timestamp arithmetic |
|
|
2132
|
+
| `mcp.loop.wasted_tokens_in` / `_out` / `.wasted_cost_usd` | event attr | library-computed over tool-controlled | same aggregation/pricing chain as `mcp.tool.tokens.*`/`mcp.tool.cost.usd` above |
|
|
2133
|
+
| `mcp.loop.first_span_id` / `.first_trace_id` | event attr | SDK-supplied | OTel SDK-generated span/trace ids |
|
|
2134
|
+
| `mcp.loop.session_id` | event attr | **blend — depends on code path** | a real transport session id when one exists; a library-generated fallback UUID when it doesn't. See "Two keys whose provenance changes by code path" below |
|
|
2135
|
+
| `service.name` | resource attr | host-supplied | the `serviceName` config option (`setupNodeSdk: true` only) |
|
|
2136
|
+
| `mcp.pricing.default_table_last_verified` | resource attr | library-computed | constant, attached only when using the unmodified default pricing table |
|
|
2137
|
+
| `mcp.tool.outcome` | metric label only | library-computed | never a span attribute — `mcp.tool.duration`'s own outcome dimension |
|
|
2138
|
+
|
|
2139
|
+
**Two keys whose provenance changes by code path, not just by
|
|
2140
|
+
configuration:**
|
|
2141
|
+
|
|
2142
|
+
- **`error.type`** carries the tool's own `err.name` when the failure was
|
|
2143
|
+
a thrown exception, but a fixed library constant (`tool_error`) when
|
|
2144
|
+
the failure was a successful JSON-RPC response with `isError: true` —
|
|
2145
|
+
same attribute key, genuinely different source depending on which of
|
|
2146
|
+
the two failure paths produced it.
|
|
2147
|
+
- **`mcp.loop.session_id`** carries a real transport-assigned session id
|
|
2148
|
+
under a session-oriented transport, but a library-generated fallback
|
|
2149
|
+
UUID (`resolveThrashSessionId()`, `src/instrument.js`) under a
|
|
2150
|
+
transport this library can't positively confirm is session-oriented —
|
|
2151
|
+
see "Agent Thrash Detection" → sessionId fallback resolution above. The
|
|
2152
|
+
value is a real id either way; only *who* generated it differs.
|
|
2153
|
+
|
|
2154
|
+
**`mcp.server.name` / `mcp.server.version`** are defined in
|
|
2155
|
+
`src/attributes.js` but are not currently set anywhere in this codebase —
|
|
2156
|
+
dead constants, not a real attribute this library records today. Left
|
|
2157
|
+
out of the table above for that reason.
|
|
2158
|
+
|
|
2159
|
+
**Want to see the actual value landing in an attribute for your own
|
|
2160
|
+
traffic, not just its category?** Don't reach for a new debugging mode —
|
|
2161
|
+
the hooks that let you do this safely already exist, because *you*, not
|
|
2162
|
+
this library, know what's safe to look at in your own deployment:
|
|
2163
|
+
|
|
2164
|
+
- **`errorRecording.redactor`** (ADR 020, `docs/adr/020-redactor-hook.md`)
|
|
2165
|
+
sees the raw `message`/`stack` before this library's own scrubbing runs
|
|
2166
|
+
— a temporary, local redactor that copies what it receives to wherever
|
|
2167
|
+
*you've* decided is safe (your own log line, a breakpoint, a counter)
|
|
2168
|
+
shows you exactly what's landing in `exception.message`/`.stacktrace`
|
|
2169
|
+
for real calls, without this library needing to add a second channel
|
|
2170
|
+
for the same content.
|
|
2171
|
+
- **`costTracking.extractor`** (a `UsageExtractor`) sees the full tool
|
|
2172
|
+
result before token/model/cost attribution runs — the same trick works
|
|
2173
|
+
for `mcp.tool.tokens.*`/`mcp.tool.model`/`mcp.tool.cost.*`.
|
|
2174
|
+
- **A custom `Classifier`** (`opts.classifiers` on `computeFingerprint()`,
|
|
2175
|
+
`src/fingerprint/compose.js` — see "Deep-failure fingerprinting" →
|
|
2176
|
+
"Extending it" above) sees the coerced error before classification.
|
|
2177
|
+
Unlike the two hooks above, this one isn't yet wired through
|
|
2178
|
+
`instrumentMcpServer()`'s own options — using it today means calling
|
|
2179
|
+
`computeFingerprint()` directly rather than relying on the automatic
|
|
2180
|
+
per-call wrapping.
|
|
2181
|
+
|
|
2182
|
+
Each hook is scoped to exactly the content it already has access to; none
|
|
2183
|
+
of them create a new place for that content to reach — they're extension
|
|
2184
|
+
points this library already ships and already documents, not a bespoke
|
|
2185
|
+
audit feature.
|
|
2186
|
+
|
|
1883
2187
|
## Configuration
|
|
1884
2188
|
|
|
1885
2189
|
All options passed to `instrumentMcpServer(server, options)`. Source of
|