opentel-mcp 0.11.0 → 0.12.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/CHANGELOG.md CHANGED
@@ -1,5 +1,135 @@
1
1
  # Changelog
2
2
 
3
+ ## 0.12.0
4
+
5
+ Agent Thrash Detection gains a new, narrower session-identity fallback for
6
+ calls that carry no real session id but do carry a client-propagated W3C
7
+ trace context. This release also fixes two real bugs in the tail-sampling
8
+ recipe v0.8.0 introduced, and adds visibility for a silent budget-tracking
9
+ gap found during that same investigation. Full design for the fallback
10
+ tier: ADR 018 (`docs/adr/018-trace-id-as-thrash-fallback.md`).
11
+
12
+ ### Added — trace id as a thrash-detection session-id fallback (ADR 018)
13
+
14
+ - **New tier in `resolveThrashSessionId()`** (`src/instrument.js`), reached
15
+ only when no real session id has ever been observed on this server, and
16
+ evaluated before the existing generated-UUID/skip fallback: if a
17
+ `tools/call` request's span has a validly-extracted REMOTE parent — i.e.
18
+ `request.params._meta` carried a valid W3C `traceparent` that
19
+ `extractTraceContext()` (ADR 017) turned into a remote `SpanContext` —
20
+ that parent's trace id is used as the session-id candidate for
21
+ thrash-detection grouping. Gated on
22
+ `trace.getSpanContext(parentContext)?.isRemote === true`, never on a
23
+ span's own `traceId` read unconditionally — a root span's trace id is
24
+ freshly, randomly generated on every call, and reading it unconditionally
25
+ would have silently turned today's honest "skip, undetermined" into
26
+ "always produce a session id that never matches the previous call's,"
27
+ which is worse than skipping. See the ADR's "THE TRAP" for the full
28
+ argument.
29
+ - Real `extra.sessionId` still wins unconditionally in every case
30
+ (steps 1–2 of the resolution order untouched, byte for byte) — a trace
31
+ id is only ever consulted for a call that has no real session id at all.
32
+ - Does **not** set `thrashSessionState.hasSeenRealSessionId` — that flag
33
+ means "this transport hands out real session identity," a permanent
34
+ per-server fact; a trace id being present on one call is a per-call fact
35
+ about that one client's behavior, not proof about the transport.
36
+ - No new `diag.warn()` for this tier, deliberately — unlike the
37
+ generated-UUID fallback (which warns because it's this library guessing
38
+ an unproven assumption about transport topology), a trace-id candidate
39
+ is real, client-supplied data with no operator action item to flag, and
40
+ warning on every occurrence — potentially far more often than the
41
+ once-per-server UUID warning — would just train operators to ignore this
42
+ library's warnings generally.
43
+ - No changes to `ThrashDetector`/`thrash/detector.js` — its composite key
44
+ already treats `sessionId` as opaque.
45
+
46
+ **Read the constraint before assuming this closes stateless MCP's
47
+ session gap.** This tier only fires when the calling *client* chooses to
48
+ propagate trace context into `_meta.traceparent` — today, per ADR 018's
49
+ investigation, that means third-party OTel instrumentation
50
+ (`@arizeai/openinference-instrumentation-mcp` and equivalents) wrapping
51
+ **v1-based** SDK clients, not either MCP SDK's own built-in behavior, and
52
+ **no v2-targeting instrumentation was found to exist anywhere**. **This
53
+ does NOT close `docs/known-gaps.md` entry 6's structural finding** — a
54
+ v2/2026-07-28-native deployment whose client doesn't propagate
55
+ `_meta.traceparent` (the default, unconfigured case for essentially
56
+ every v2 client today) gets nothing new from this release: the exact
57
+ same `null`/skip behavior as before. See the README's "Session id
58
+ resolution" section and ADR 018's "Adoption caveat" for the full scope.
59
+
60
+ ### Added — `mcp.tool.schema_drift_detected` span attribute
61
+
62
+ - New boolean span attribute, set alongside (never instead of) the
63
+ existing `mcp.tool.schema_drift.detected` span event
64
+ (`schema-drift/emitter.js`) — the same resolution ADR 011 already
65
+ applied to thrash detection (`mcp.tool.thrash_detected`) for the
66
+ identical event-vs-attribute ambiguity: whether a Collector
67
+ `tailsamplingprocessor`'s `boolean_attribute` policy can match
68
+ span-*event* data (as opposed to top-level span attributes) could not be
69
+ confirmed either way (Go source, not installed in this repository). Only
70
+ ever set to `true`, and only when drift was actually detected — never
71
+ explicitly set `false`.
72
+
73
+ ### Fixed — tail-sampling recipe referenced a non-existent span attribute
74
+
75
+ - The README's `tailsamplingprocessor` recipe recommended keying a
76
+ `boolean_attribute` policy on `mcp.tool.schema_drift.detected` — but
77
+ that string was only ever a span *event* name and a metric counter name
78
+ (`schema-drift/attributes.js`), never passed to `span.setAttribute()`
79
+ anywhere in this package. As documented, that policy could never have
80
+ matched anything. Fixed by the new `mcp.tool.schema_drift_detected`
81
+ attribute above.
82
+ - Also added a `string_attribute` policy on `mcp.tool.pricing_status =
83
+ "unknown"` — v0.11.0's "confidently wrong zero" problem reappearing at
84
+ the sampling layer: an unpriced call has real extracted token usage but
85
+ no `mcp.tool.cost.usd`, indistinguishable from a genuinely free call to
86
+ a numeric-threshold policy.
87
+ - Recipe YAML extracted to `docs/recipes/tail-sampling.yaml` (repo-only,
88
+ not published — the same carve-out `dashboards/` already has), with a
89
+ new cross-check test (`test/recipes/tail-sampling-attributes.test.js`)
90
+ asserting every attribute a policy references is a real, exported
91
+ constant AND actually passed to `span.setAttribute()` — the check that
92
+ would have caught this bug automatically. The README now references the
93
+ file instead of duplicating it, and states plainly that attribute
94
+ *names* are cross-checked but Collector policy *behavior* itself has not
95
+ been run end to end (no Docker/Collector available in this project's dev
96
+ environment).
97
+ - `docs/known-gaps.md` entry 9 (new): an unpriced call never reaches
98
+ `budgetTracker.recordAndCheck()`, so budget guardrails cannot trip on
99
+ unpriced spend regardless of amount. The visibility half is fixed in
100
+ this same release — see "Added" below — but the underlying accounting
101
+ behavior is not; what should happen to an unpriced call's budget
102
+ accounting is a real design question, deliberately left open.
103
+
104
+ ### Added — visibility for unpriced spend against a configured budget (known-gaps entry 9)
105
+
106
+ - **Two new one-time `diag.warn()` diagnostics in `createBudgetTracker()`**
107
+ (`src/cost/budget.js`), no behavior change and no new public surface:
108
+ one fires at construction whenever `perSessionUsd`/`perToolUsd` is
109
+ configured at all, stating plainly that unpriced calls won't count
110
+ toward it; the other — a new `recordUnpriced(model)` method, called from
111
+ `applyCostAttribution()`'s existing `costUsd === null` branch
112
+ (`src/instrument.js`), the same branch that already sets
113
+ `mcp.tool.pricing_status: "unknown"` — fires the first time an unpriced
114
+ call under an active budget is actually observed, naming the model and
115
+ which scope(s) are configured. Both no-op when no budget is configured;
116
+ neither changes `BudgetCheckResult`'s shape or adds a span attribute.
117
+ - **Deliberately diagnostics only — no fallback pricing was added.**
118
+ Making an unpriced call actually count toward a USD budget means
119
+ inventing a number for it, and a wrong invented price is a *different*
120
+ confidently-wrong number, not a fix — the same disease this warning
121
+ exists to flag, one layer up. See `docs/known-gaps.md` entry 9's
122
+ "Status update (v0.12.0)" for the full argument against building that
123
+ now, and why it stays open as a future, ADR-gated decision rather than
124
+ folded into this patch.
125
+ - Both warnings share the budget tracker's own existing
126
+ once-per-tracker-instance granularity (`createBudgetTracker()`'s own
127
+ docblock) — under the default (no `instanceKey`), a fresh-server-per-request
128
+ deployment re-warns on every request for both, the same inherited-caveat
129
+ shape `docs/known-gaps.md` entry 6 documents for the thrash fallback
130
+ warning; a stable `instanceKey` shares one tracker, and one already-armed
131
+ warning, across calls, same as every other registry-backed tracker.
132
+
3
133
  ## 0.11.0
4
134
 
5
135
  **⚠️ Type change, not a runtime behavior change — read this first.**
package/README.md CHANGED
@@ -658,6 +658,28 @@ region.
658
658
  on the span today, so "Fleet-wide fingerprint frequency" below has
659
659
  nothing to substitute with for this scope specifically. `perToolUsd` is
660
660
  unaffected — it never depended on session id.
661
+ - **A budget only ever sees priced spend — an unpriced call (`mcp.tool.pricing_status:
662
+ "unknown"`) never reaches it at all, and never counts toward
663
+ `perSessionUsd`/`perToolUsd`, regardless of how many tokens it burned**
664
+ (`docs/known-gaps.md` entry 9). v0.12.0 makes this loud instead of
665
+ silent, but does not close it: setting a `budget` at all fires a one-time
666
+ `diag.warn()` naming the constraint up front, and the first time an
667
+ actual unpriced call happens while a budget is active, a second one-time
668
+ `diag.warn()` names the model and which scope(s) are configured. Both
669
+ are pure diagnostics — no new span attribute, no change to
670
+ `mcp.tool.cost.budget_exceeded`'s meaning, and an unpriced call still
671
+ contributes nothing to the running total. Deliberate: inventing a
672
+ fallback price for an unpriceable call would trade one confidently-wrong
673
+ number (silent zero) for a different one (a made-up price) — the exact
674
+ failure this attribute/warning pair exists to catch, one layer up. If
675
+ you're seeing either warning, the fix is the same one "Overriding
676
+ pricing" above already documents: add the model to `costTracking.pricing`.
677
+ Both warnings inherit the same per-tracker-instance granularity as the
678
+ budget tracker itself — see the item above and `src/cost/budget.js`'s
679
+ `createBudgetTracker()` docblock for exactly what that means under a
680
+ fresh-`Server`-per-request deployment (a stable `instanceKey` shares one
681
+ tracker, and one already-armed warning, across calls; without one, both
682
+ warnings re-fire on every request).
661
683
 
662
684
  ## Agent Thrash Detection (v0.6.0+)
663
685
 
@@ -895,8 +917,39 @@ in order:
895
917
  shared fallback key, even if `assumeSingleSession` is set. A server
896
918
  that has proven it hands out real session ids doesn't get to fall back
897
919
  just because one particular call lacked one.
898
- 3. **Before any real session id has ever been observed**, a generated
899
- per-connection fallback id is used only when:
920
+ 3. **(v0.12.0, ADR 018) Before any real session id has ever been observed,
921
+ and before the generated-fallback rule below runs**, if this call's
922
+ span has a validly-extracted REMOTE trace parent — i.e.
923
+ `request.params._meta` carried a valid W3C `traceparent` (see "Trace
924
+ Context Propagation" below) that resolved to a remote `SpanContext`,
925
+ not a freshly-generated root span — that parent's **trace id** is used
926
+ as the session-id candidate instead. Never reads a span's own `traceId`
927
+ unconditionally: a root span's trace id is fresh, random, and different
928
+ on every single call, so using it without confirming it was actually
929
+ inherited from a real upstream parent would silently turn "skip,
930
+ undetermined" into "always produce a session id that never matches the
931
+ previous call's" — quieter and worse than skipping. Does **not** mark
932
+ the server session-aware (`hasSeenRealSessionId` stays untouched) — a
933
+ trace id being present on one call is a fact about that one client's
934
+ behavior, not a proof about the transport itself.
935
+
936
+ **⚠️ Read this before assuming it closes the stateless-MCP session gap
937
+ below.** This only fires when the calling *client* chooses to
938
+ propagate trace context into `_meta.traceparent` — today, that means
939
+ third-party OTel instrumentation (e.g.
940
+ `@arizeai/openinference-instrumentation-mcp`) wrapping **v1-based** SDK
941
+ clients, not either MCP SDK's own built-in behavior; no
942
+ v2-targeting instrumentation exists yet. **It does not close
943
+ `docs/known-gaps.md` entry 6** — a v2/2026-07-28-native deployment
944
+ whose client doesn't propagate `_meta.traceparent` (the default,
945
+ unconfigured case for essentially every v2 client today) gets nothing
946
+ new here: the exact same skip behavior as before. Full investigation,
947
+ including why one trace is typically one agent turn (not a protocol
948
+ guarantee) and the residual merge risk this accepts: ADR 018
949
+ (`docs/adr/018-trace-id-as-thrash-fallback.md`).
950
+ 4. **Before any real session id has ever been observed, and no usable
951
+ trace id was found above**, a generated per-connection fallback id is
952
+ used only when:
900
953
  - the transport is **structurally confirmed single-connection** — no
901
954
  `sessionId` property on `server.transport` at all (e.g. stdio's
902
955
  `StdioServerTransport`, which has no session concept whatsoever), or
@@ -1231,6 +1284,25 @@ silently and only the metric still fires.
1231
1284
  | `mcp.tool.schema_drift.removed_fields` | Property names removed — present only when `type` is `field_removed` or `multiple` |
1232
1285
  | `mcp.tool.schema_drift.changed_fields` | Property names whose value changed — present only when `type` is `type_changed` or `multiple` |
1233
1286
 
1287
+ **Also sets a boolean span *attribute*, `mcp.tool.schema_drift_detected: true`**,
1288
+ on that same span, alongside the event above (v0.12.0, ADR 011 —
1289
+ `docs/adr/011-cost-aware-sampling.md`, the same resolution already applied
1290
+ to `mcp.tool.thrash_detected` above, for the identical reason). This
1291
+ exists for one specific consumer: an OpenTelemetry Collector's
1292
+ `tailsamplingprocessor`, whose `boolean_attribute` policy matches
1293
+ top-level span attributes — whether such a policy can also match
1294
+ span-*event* data was investigated and left genuinely unverified (no Go
1295
+ source to check against in this repository), so the attribute exists to
1296
+ remove that uncertainty entirely for anyone wiring up schema-drift-aware
1297
+ tail sampling. See "Cost-aware trace sampling (a Collector recipe, not a
1298
+ library feature)" below. Deliberately a different string from the
1299
+ `mcp.tool.schema_drift.detected` span event/metric name above, for the
1300
+ same reason `mcp.tool.thrash_detected` doesn't reuse `mcp.tool.loop.detected`'s
1301
+ string: the one reader who most needs this name to be unambiguous —
1302
+ someone writing a Collector tail-sampling policy — would otherwise see one
1303
+ bare string with no way to tell which of the two same-named signals
1304
+ they're keying on.
1305
+
1234
1306
  ### Configuration
1235
1307
 
1236
1308
  All fields of `schemaDrift`, each independently overridable by its own
@@ -1493,86 +1565,52 @@ fundamentally bigger, more invasive ask than `instrumentMcpServer(server,
1493
1565
  options)`, and exactly the class of intervention this project has already
1494
1566
  ruled out elsewhere (never override a host's own OpenTelemetry setup).
1495
1567
 
1496
- **What this package does instead: mark, don't decide.** Two of the three
1568
+ **What this package does instead: mark, don't decide.** Four of the five
1497
1569
  signals a tail-sampling policy needs already exist as plain span
1498
- attributes with no changes required — `mcp.tool.cost.usd` (see "Cost &
1499
- Token Attribution" above) and `mcp.tool.cost.budget_exceeded` (a
1500
- cumulative budget guardrail, same section). The third,
1501
- `mcp.tool.thrash_detected` — a boolean span attribute set alongside the
1502
- existing `mcp.loop.detected` span event (see "Agent Thrash Detection" →
1503
- "Span event" above) — is the one new addition this release makes,
1504
- specifically so a tail-sampling policy has an unambiguous, attribute-level
1505
- signal to key on. The actual decision — buffer a trace, evaluate a
1506
- policy, keep or drop the whole thing — belongs to the OpenTelemetry
1507
- Collector's `tailsamplingprocessor`, which already does this correctly,
1508
- already handles the hard parts (per-trace span buffering across a wait
1509
- window, multi-service traces, decision policies), and runs where it can
1510
- see every span in a trace regardless of which process produced it —
1511
- something this library, running inside one MCP server process, never can.
1570
+ attributes with no changes required — `mcp.tool.cost.usd` and
1571
+ `mcp.tool.pricing_status` (see "Cost & Token Attribution" above),
1572
+ `mcp.tool.cost.budget_exceeded` (a cumulative budget guardrail, same
1573
+ section), and `mcp.tool.thrash_detected` — a boolean span attribute set
1574
+ alongside the existing `mcp.loop.detected` span event (see "Agent Thrash
1575
+ Detection" → "Span event" above). The fifth, `mcp.tool.schema_drift_detected`
1576
+ — a boolean span attribute set alongside the existing
1577
+ `mcp.tool.schema_drift.detected` span event (see "Tool schema drift
1578
+ detection" → "Span event" above) — is the one new addition this release
1579
+ makes, for the identical reason `mcp.tool.thrash_detected` was added in
1580
+ the first place: a tail-sampling policy needs an unambiguous,
1581
+ attribute-level signal to key on, and a span *event* isn't confirmed
1582
+ matchable the same way (see below). The actual decision — buffer a trace,
1583
+ evaluate a policy, keep or drop the whole thing — belongs to the
1584
+ OpenTelemetry Collector's `tailsamplingprocessor`, which already does this
1585
+ correctly, already handles the hard parts (per-trace span buffering
1586
+ across a wait window, multi-service traces, decision policies), and runs
1587
+ where it can see every span in a trace regardless of which process
1588
+ produced it — something this library, running inside one MCP server
1589
+ process, never can.
1512
1590
 
1513
1591
  ### A working Collector config
1514
1592
 
1515
- Real, pasteable `tailsamplingprocessor` config — not pseudo-config. Keeps
1516
- any trace containing an expensive call, a budget-exceeded call, or a
1517
- detected thrash loop; everything else gets an ordinary probabilistic
1518
- sample. (Standard OpenTelemetry Collector Contrib syntax — external to
1519
- this repository, so treat field names as this component's own documented
1520
- contract, not something confirmed against code living here.)
1521
-
1522
- ```yaml
1523
- receivers:
1524
- otlp:
1525
- protocols:
1526
- grpc:
1527
- http:
1528
-
1529
- processors:
1530
- tail_sampling:
1531
- decision_wait: 10s
1532
- num_traces: 50000
1533
- expected_new_traces_per_sec: 10
1534
- policies:
1535
- # Keep any trace containing a call that cost more than $0.10.
1536
- - name: expensive-tool-calls
1537
- type: numeric_attribute
1538
- numeric_attribute:
1539
- key: mcp.tool.cost.usd
1540
- min_value: 0.10
1541
-
1542
- # Keep any trace where a configured cost budget was crossed.
1543
- - name: budget-exceeded-calls
1544
- type: boolean_attribute
1545
- boolean_attribute:
1546
- key: mcp.tool.cost.budget_exceeded
1547
- value: true
1548
-
1549
- # Keep any trace containing a detected agent thrash loop.
1550
- - name: thrash-loops
1551
- type: boolean_attribute
1552
- boolean_attribute:
1553
- key: mcp.tool.thrash_detected
1554
- value: true
1555
-
1556
- # Everything else: an ordinary 10% probabilistic sample. Policies
1557
- # are OR'd together by the processor, so this doesn't reduce
1558
- # anything the three policies above already decided to keep — it
1559
- # only adds baseline visibility into the traces none of them matched.
1560
- - name: baseline-sample
1561
- type: probabilistic
1562
- probabilistic:
1563
- sampling_percentage: 10
1564
-
1565
- exporters:
1566
- otlp:
1567
- endpoint: your-backend:4317
1568
-
1569
- service:
1570
- pipelines:
1571
- traces:
1572
- receivers: [otlp]
1573
- processors: [tail_sampling]
1574
- exporters: [otlp]
1575
- ```
1593
+ The full, pasteable `tailsamplingprocessor` config lives in
1594
+ [`docs/recipes/tail-sampling.yaml`](../../docs/recipes/tail-sampling.yaml)
1595
+ — not duplicated here, so there's exactly one copy to keep in sync with
1596
+ this package's actual attribute names. (Standard OpenTelemetry Collector
1597
+ Contrib syntax — external to this repository, so treat field names as
1598
+ that component's own documented contract, not something confirmed against
1599
+ code living here.)
1600
+
1601
+ It keeps any trace containing:
1602
+
1603
+ | Policy | Type | Keys on | Why |
1604
+ |---|---|---|---|
1605
+ | `expensive-tool-calls` | `numeric_attribute` | `mcp.tool.cost.usd` ≥ `0.10` | A single call cost more than your threshold |
1606
+ | `budget-exceeded-calls` | `boolean_attribute` | `mcp.tool.cost.budget_exceeded` = `true` | A configured cumulative budget was crossed |
1607
+ | `thrash-loops` | `boolean_attribute` | `mcp.tool.thrash_detected` = `true` | An agent thrash loop was detected |
1608
+ | `schema-drift-events` | `boolean_attribute` | `mcp.tool.schema_drift_detected` = `true` | A tool's `inputSchema` changed between two `tools/list` calls |
1609
+ | `unpriced-calls` | `string_attribute` | `mcp.tool.pricing_status` = `"unknown"` | See below — a real, unpriced cost, not a cheap one |
1610
+
1611
+ — everything else gets an ordinary 10% probabilistic sample (policies are
1612
+ OR'd together, so this doesn't reduce anything the policies above already
1613
+ kept; it only adds baseline visibility into what none of them matched).
1576
1614
 
1577
1615
  Point `instrumentMcpServer({ exporterUrl: 'http://localhost:4318/v1/traces' })`
1578
1616
  (or your host's own OTLP exporter configuration) at this Collector's
@@ -1580,12 +1618,55 @@ Point `instrumentMcpServer({ exporterUrl: 'http://localhost:4318/v1/traces' })`
1580
1618
  this policy before anything is exported downstream.
1581
1619
 
1582
1620
  **Adjust `min_value`/`sampling_percentage` to your own cost/volume
1583
- profile** — `0.10` and `10%` above are illustrative starting points, not
1584
- recommendations; `decision_wait`/`num_traces` should scale with your
1585
- actual traffic volume (the Collector's own docs cover sizing these). If
1586
- your schema-drift or two-axis observation signals matter for retention
1587
- too, add `mcp.tool.schema_drift.detected` (boolean_attribute) as another
1588
- OR'd policy the same way.
1621
+ profile** — `0.10` and `10%` in the recipe are illustrative starting
1622
+ points, not recommendations; `decision_wait`/`num_traces` should scale
1623
+ with your actual traffic volume (the Collector's own docs cover sizing
1624
+ these).
1625
+
1626
+ **Why `unpriced-calls` is in here — the v0.11.0 "confidently wrong zero"
1627
+ problem, one layer up.** `mcp.tool.pricing_status` (ADR 016 point 4, "Cost
1628
+ & Token Attribution" → "Span attributes" above) exists because an
1629
+ unrecognized model must not be silently reported as costing nothing — it
1630
+ gets `"unknown"`, distinct from a genuinely free/untracked call. But a
1631
+ call with `mcp.tool.pricing_status: "unknown"` has real extracted token
1632
+ usage and *no* `mcp.tool.cost.usd` at all (unpriced calls never reach
1633
+ `mcp.tool.cost.usd`, by construction — see "Cost & Token Attribution"
1634
+ above). To the `expensive-tool-calls` policy above, an attribute that
1635
+ was never set is indistinguishable from a cost of exactly `0` — the same
1636
+ "confidently wrong" failure the pricing-provenance work fixed at the
1637
+ attribute level, reappearing here because a numeric-threshold policy has
1638
+ no way to see the difference between "cheap" and "unknown." The
1639
+ `unpriced-calls` policy closes that gap the same way `mcp.tool.pricing_status`
1640
+ closes it at the attribute level: an explicit signal instead of an
1641
+ inferred absence. (Cost tracking's other kind of absence — a plain-text
1642
+ result with no recognizable token usage at all — has no attribute of any
1643
+ kind, `mcp.tool.pricing_status` included, and is not what this policy is
1644
+ for: that call genuinely has no cost signal, which is expected, not a
1645
+ gap.)
1646
+
1647
+ **What's verified here, and what isn't — read this before trusting this
1648
+ recipe blindly.** Every `mcp.tool.*` key in
1649
+ [`tail-sampling.yaml`](../../docs/recipes/tail-sampling.yaml) is
1650
+ cross-checked by
1651
+ [`test/recipes/tail-sampling-attributes.test.js`](test/recipes/tail-sampling-attributes.test.js)
1652
+ against this package's real, exported attribute constants — and,
1653
+ specifically, against which of them are actually passed to
1654
+ `span.setAttribute()` in `src/`, not merely a span-event or metric name
1655
+ that happens to look like an attribute. That test is what would have
1656
+ caught this recipe's own predecessor bug: an earlier version of this
1657
+ section recommended the schema-drift span *event* name as a
1658
+ `boolean_attribute` policy target, which could never have matched
1659
+ anything. **What is not verified: the actual behavior of this config
1660
+ against a real Collector.** `tailsamplingprocessor` is Contrib-only Go
1661
+ source with no npm package and nothing installed in this repository, and
1662
+ no Docker daemon was reachable in the environment this recipe was last
1663
+ revised in — unlike the Grafana dashboard (`dashboards/README.md`, "Generating
1664
+ sample data / verifying locally"), which was verified end to end against
1665
+ a real Prometheus + Grafana stack before shipping, this recipe has not
1666
+ had the equivalent live-Collector run. Treat the attribute names as
1667
+ trustworthy and the policy semantics as sourced from public OpenTelemetry
1668
+ Collector Contrib documentation, not as something this project has
1669
+ independently confirmed by running it.
1589
1670
 
1590
1671
  ## Trace Context Propagation (v0.11.0+)
1591
1672
 
@@ -2154,7 +2235,19 @@ pragmatic choice rather than a spec-pure one.
2154
2235
  `mcp.tool.cost.usd` / `mcp.tool.cost.budget_exceeded` attributes, plus
2155
2236
  a documented, pasteable OpenTelemetry Collector `tailsamplingprocessor`
2156
2237
  config that keeps expensive/budget-exceeded/thrashing traces alongside
2157
- a normal probabilistic sample for everything else.
2238
+ a normal probabilistic sample for everything else. **Update (v0.12.0):**
2239
+ the recipe YAML moved to `docs/recipes/tail-sampling.yaml` (README
2240
+ references it rather than duplicating it), gained two more policies —
2241
+ `mcp.tool.schema_drift_detected` (a new attribute, the same event/attribute
2242
+ fix applied to schema drift) and `mcp.tool.pricing_status = "unknown"`
2243
+ (closes a v0.11.0-adjacent gap: an unpriced call has real cost but no
2244
+ `mcp.tool.cost.usd`, which a numeric-threshold policy can't tell apart
2245
+ from a genuinely free one) — and is now cross-checked by a test
2246
+ (`test/recipes/tail-sampling-attributes.test.js`) that fixes a real bug
2247
+ this section previously had: it recommended the schema-drift span
2248
+ *event* name as a `boolean_attribute` policy target, which could never
2249
+ have matched anything. See "Cost-aware trace sampling" above for the
2250
+ full detail, including what is and isn't verified.
2158
2251
  - v0.10.0: `@modelcontextprotocol/server` (MCP v2, protocol revision
2159
2252
  2026-07-28) support ✓ — see "MCP v2 support" above and ADR 015
2160
2253
  (`docs/adr/015-mcp-v2-support.md`). Spans, standard attributes, failure
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "opentel-mcp",
3
- "version": "0.11.0",
3
+ "version": "0.12.0",
4
4
  "description": "One-line OpenTelemetry instrumentation for Model Context Protocol (MCP) servers",
5
5
  "type": "module",
6
6
  "main": "src/index.js",
@@ -78,6 +78,7 @@
78
78
  "@opentelemetry/sdk-metrics": "^2.9.0",
79
79
  "@opentelemetry/sdk-trace-base": "^2.9.0",
80
80
  "@vitest/coverage-v8": "^2.1.9",
81
+ "js-yaml": "^4.3.1",
81
82
  "typescript": "^7.0.2",
82
83
  "vitest": "^2.1.8",
83
84
  "zod": "^4.4.3"
@@ -8,8 +8,23 @@
8
8
  * that limit (denying the call, alerting, etc.) is left entirely to
9
9
  * whatever consumes the resulting `mcp.tool.cost.budget_exceeded` span
10
10
  * attribute (see src/attributes.js).
11
+ *
12
+ * v0.12.0 (docs/known-gaps.md entry 9): a call whose model doesn't resolve
13
+ * to a price (`mcp.tool.pricing_status: "unknown"`) never reaches
14
+ * `recordAndCheck()` below at all — `applyCostAttribution()` only calls it
15
+ * when `costUsd !== null`. That's not new behavior here, and this release
16
+ * does not change it (see that entry's "confidently wrong number" reasoning
17
+ * for why inventing a fallback price would trade one silent-failure shape
18
+ * for another, not fix it) — it only adds two diagnostics so the gap is
19
+ * visible instead of silent: a construction-time warning that a configured
20
+ * budget won't see unpriced spend, and a first-occurrence warning naming
21
+ * the model the first time it actually happens. Both are pure
22
+ * `diag.warn()` calls; neither changes `BudgetCheckResult`, adds a span
23
+ * attribute, or alters `recordAndCheck()`'s existing behavior in any way.
11
24
  */
12
25
 
26
+ import { diag } from '@opentelemetry/api';
27
+
13
28
  /**
14
29
  * @typedef {Object} BudgetConfig
15
30
  * @property {number} [perSessionUsd] - Cumulative-cost limit per MCP session id. Calls with no session id
@@ -75,20 +90,57 @@ function accumulate(map, key, amount) {
75
90
  * every time — a confirmed gap, not a hypothetical. See ADR 012,
76
91
  * docs/adr/012-tracker-lifecycle-and-shared-state.md. Restart the process
77
92
  * (or, for a long-lived server, build your own eviction on top) to reset
78
- * intentionally.
93
+ * intentionally. **The two v0.12.0 warnings below inherit this exact same
94
+ * granularity** — both live in this same closure, so both re-arm on every
95
+ * fresh construction. Under the default (no `instanceKey`), a
96
+ * fresh-server-per-request deployment re-warns on *every request* for
97
+ * both — the identical inherited-caveat shape `docs/known-gaps.md` entry 6
98
+ * documents for `thrashSessionState`'s fallback-session warning. Unlike
99
+ * that entry's gap, though, `instrument.js` already constructs this
100
+ * tracker via `getOrCreateTracker(resolved.instanceKey, 'budget', ...)` —
101
+ * so a caller who *does* set a stable `instanceKey` shares one tracker
102
+ * (and therefore one already-armed warning flag) across calls, the same as
103
+ * every other registry-backed tracker; only the no-`instanceKey` default
104
+ * case gets the repeated warning.
79
105
  *
80
106
  * @param {BudgetConfig | undefined} budget - costTracking.budget from config.js. Omitted/undefined limits
81
- * mean that scope is never tracked or checked — recordAndCheck() becomes a pure no-op for it.
82
- * @returns {{ recordAndCheck: (sessionId: unknown, toolName: unknown, costUsd: number) => BudgetCheckResult }}
107
+ * mean that scope is never tracked or checked — recordAndCheck() becomes a pure no-op for it, and neither
108
+ * new v0.12.0 warning below ever fires.
109
+ * @returns {{
110
+ * recordAndCheck: (sessionId: unknown, toolName: unknown, costUsd: number) => BudgetCheckResult,
111
+ * recordUnpriced: (model: unknown) => void,
112
+ * }}
83
113
  */
84
114
  export function createBudgetTracker(budget) {
85
115
  const perSessionUsd = isFiniteNumber(budget?.perSessionUsd) ? budget.perSessionUsd : undefined;
86
116
  const perToolUsd = isFiniteNumber(budget?.perToolUsd) ? budget.perToolUsd : undefined;
117
+ const budgetConfigured = perSessionUsd !== undefined || perToolUsd !== undefined;
118
+
119
+ // v0.12.0, docs/known-gaps.md entry 9: fires at most once per tracker
120
+ // construction (see this function's own docblock for exactly what "once"
121
+ // means here) — a proactive heads-up that this budget, once configured,
122
+ // will silently ignore any call whose model never resolves to a price.
123
+ // Unconditional on whether that ever actually happens; see
124
+ // recordUnpriced() below for the reactive counterpart that fires only
125
+ // when it does.
126
+ if (budgetConfigured) {
127
+ try {
128
+ diag.warn(
129
+ 'opentel-mcp: a costTracking.budget guardrail is configured, but it only accounts for calls whose model ' +
130
+ 'resolves to a known price. A call whose usage is extracted but whose model is unrecognized or unpriced ' +
131
+ '(mcp.tool.pricing_status: "unknown") is never counted toward perSessionUsd/perToolUsd, regardless of its ' +
132
+ 'real token cost — see docs/known-gaps.md entry 9.',
133
+ );
134
+ } catch {
135
+ // Never throw — see module docblock.
136
+ }
137
+ }
87
138
 
88
139
  /** @type {Map<string, number>} */
89
140
  const sessionCostMap = new Map();
90
141
  /** @type {Map<string, number>} */
91
142
  const toolCostMap = new Map();
143
+ let warnedUnpriced = false;
92
144
 
93
145
  return {
94
146
  recordAndCheck(sessionId, toolName, costUsd) {
@@ -118,5 +170,40 @@ export function createBudgetTracker(budget) {
118
170
  return { exceeded: false, scope: null };
119
171
  }
120
172
  },
173
+
174
+ /**
175
+ * v0.12.0, docs/known-gaps.md entry 9: the reactive counterpart to the
176
+ * construction-time warning above. Called from `applyCostAttribution()`
177
+ * (src/instrument.js) in the branch where a call's usage was extracted
178
+ * but its cost never resolved (`costUsd === null`, the same condition
179
+ * that sets `mcp.tool.pricing_status: "unknown"`) — i.e. exactly the
180
+ * calls `recordAndCheck()` above is never invoked for. No-op, silently,
181
+ * when no budget scope is configured (nothing to warn about) or once
182
+ * this tracker has already warned once (see this module's own
183
+ * granularity docblock on `createBudgetTracker()` for what "once"
184
+ * means here). Never throws, matching every other public method here.
185
+ *
186
+ * @param {unknown} model - usage.model from the extractor's result, possibly undefined (no model detected at all).
187
+ */
188
+ recordUnpriced(model) {
189
+ if (!budgetConfigured || warnedUnpriced) return;
190
+ warnedUnpriced = true;
191
+
192
+ try {
193
+ const scopes = [perSessionUsd !== undefined ? 'perSessionUsd' : null, perToolUsd !== undefined ? 'perToolUsd' : null]
194
+ .filter(Boolean)
195
+ .join(', ');
196
+ const modelDescription = isNonEmptyString(model) ? `"${model}"` : '(no model detected)';
197
+ diag.warn(
198
+ `opentel-mcp: a tool call using model ${modelDescription} did not resolve to a price and was not counted ` +
199
+ `toward the configured budget guardrail (${scopes}). Real spend may be higher than ` +
200
+ 'mcp.tool.cost.budget_exceeded/budget_scope report. This warning fires once per tracker instance — see ' +
201
+ "this function's own docblock for what that means under a fresh-server-per-request deployment — and " +
202
+ 'only when a budget is configured. See docs/known-gaps.md entry 9.',
203
+ );
204
+ } catch {
205
+ // Never throw — see module docblock.
206
+ }
207
+ },
121
208
  };
122
209
  }
package/src/instrument.js CHANGED
@@ -764,7 +764,13 @@ function extractSessionAndRequestId(kind, extraOrCtx) {
764
764
  * (src/cost/budget.js) — if that reports the call pushed a configured
765
765
  * budget over its limit, sets mcp.tool.cost.budget_exceeded /
766
766
  * mcp.tool.cost.budget_scope. An unrecognized model silently skips cost
767
- * *and* budget attribution — the token attributes still land.
767
+ * *and* budget attribution — the token attributes still land. v0.12.0
768
+ * (docs/known-gaps.md entry 9): that "silently" is now only true of the
769
+ * span/metrics — the same branch calls `budgetTracker.recordUnpriced()`
770
+ * instead, which fires a one-time diag.warn() naming the model, but only
771
+ * when a budget is actually configured (a pure no-op otherwise). No
772
+ * behavior change: this is diagnostics only, not a new price or a new
773
+ * budget-exceeded condition.
768
774
  *
769
775
  * No-op when `costTracking.enabled` is false. Called from both the
770
776
  * isToolResultError and success branches below (there's a result to read
@@ -840,6 +846,14 @@ function applyCostAttribution(span, metricsRecorder, toolName, sessionId, result
840
846
  span.setAttribute(ATTR_MCP_TOOL_COST_BUDGET_EXCEEDED, true);
841
847
  span.setAttribute(ATTR_MCP_TOOL_COST_BUDGET_SCOPE, budgetResult.scope);
842
848
  }
849
+ } else {
850
+ // v0.12.0, docs/known-gaps.md entry 9: the same condition that set
851
+ // pricingStatus to "unknown" above — usage was extracted, but never
852
+ // reached recordAndCheck() at all, so a configured budget silently
853
+ // never saw it. recordUnpriced() no-ops unless a budget is actually
854
+ // configured, and warns at most once per tracker instance — see its
855
+ // own docblock (src/cost/budget.js).
856
+ budgetTracker.recordUnpriced(usage.model);
843
857
  }
844
858
 
845
859
  return { tokensIn: usage.inputTokens, tokensOut: usage.outputTokens, costUsd: costUsd ?? 0 };
@@ -999,9 +1013,33 @@ function describeFallbackReason(server, kind) {
999
1013
  * is set. A server that has proven it hands out real session ids
1000
1014
  * does not get to fall back just because one particular call lacked
1001
1015
  * one.
1002
- * 3. Before any real sessionId has ever been observed: the generated
1003
- * per-connection fallback (`thrashConnectionFallbackSessionId`) is
1004
- * used when `thrashConfig.assumeSingleSession` is true, or when
1016
+ * 2.5. (ADR 018, docs/adr/018-trace-id-as-thrash-fallback.md) Before any
1017
+ * real sessionId has ever been observed, and before step 3's
1018
+ * UUID/skip fallback runs: if this call's span has a validly
1019
+ * extracted REMOTE parent — i.e. `request.params._meta` carried a
1020
+ * valid W3C `traceparent` that `extractTraceContext()`
1021
+ * (src/tracecontext/extract.js) turned into a remote `SpanContext` —
1022
+ * use that parent's trace id as the session-id candidate. Gated on
1023
+ * `trace.getSpanContext(parentContext)?.isRemote === true`, never on
1024
+ * a span's own `traceId` in isolation: every span, root or child, has
1025
+ * a `traceId`, and a ROOT span's (no valid `traceparent` extracted)
1026
+ * is freshly, randomly generated on every single call — using it
1027
+ * unconditionally would silently turn today's honest "skip,
1028
+ * undetermined" into "always produce a session id that never matches
1029
+ * the previous call's," which is worse than skipping (see the ADR's
1030
+ * "THE TRAP"). Does NOT set `thrashSessionState.hasSeenRealSessionId`
1031
+ * — that flag means "this transport hands out real session
1032
+ * identity," a permanent per-server fact; a trace id being present
1033
+ * on one call is a per-call fact about that one client's behavior,
1034
+ * not proof about the transport (ADR 018's "Constraints accepted").
1035
+ * No diag.warn(): unlike step 3 below, this isn't a guess about
1036
+ * transport topology the operator can correct — it's real,
1037
+ * client-supplied data, strictly finer-grained than the fallback it
1038
+ * sits beside, with no operator action item to flag.
1039
+ * 3. Before any real sessionId has ever been observed, and step 2.5
1040
+ * above didn't apply: the generated per-connection fallback
1041
+ * (`thrashConnectionFallbackSessionId`) is used when
1042
+ * `thrashConfig.assumeSingleSession` is true, or when
1005
1043
  * `isSingleConnectionTransport(server)` reliably determines the
1006
1044
  * transport is single-connection. Otherwise, skip — an undetermined
1007
1045
  * transport is not assumed to be single-connection. The first time
@@ -1016,9 +1054,10 @@ function describeFallbackReason(server, kind) {
1016
1054
  * @param {string} thrashConnectionFallbackSessionId
1017
1055
  * @param {import('./thrash/config.js').ThrashConfig} thrashConfig
1018
1056
  * @param {'v1' | 'v2' | undefined} kind - Threaded through to isSingleConnectionTransport()/describeFallbackReason() — see their docblocks (ADR 015).
1057
+ * @param {import('@opentelemetry/api').Context} parentContext - The span's parent context, as returned by extractTraceContext() (ADR 017) — checked here for a validly-extracted remote SpanContext (ADR 018, step 2.5).
1019
1058
  * @returns {string | null}
1020
1059
  */
1021
- function resolveThrashSessionId(server, sessionId, thrashSessionState, thrashConnectionFallbackSessionId, thrashConfig, kind) {
1060
+ function resolveThrashSessionId(server, sessionId, thrashSessionState, thrashConnectionFallbackSessionId, thrashConfig, kind, parentContext) {
1022
1061
  if (sessionId !== undefined) {
1023
1062
  thrashSessionState.hasSeenRealSessionId = true;
1024
1063
  return sessionId;
@@ -1028,6 +1067,11 @@ function resolveThrashSessionId(server, sessionId, thrashSessionState, thrashCon
1028
1067
  return null;
1029
1068
  }
1030
1069
 
1070
+ const remoteParentSpanContext = trace.getSpanContext(parentContext);
1071
+ if (remoteParentSpanContext?.isRemote === true) {
1072
+ return remoteParentSpanContext.traceId;
1073
+ }
1074
+
1031
1075
  if (thrashConfig.assumeSingleSession || isSingleConnectionTransport(server, kind)) {
1032
1076
  if (!thrashSessionState.hasWarnedFallbackUsed) {
1033
1077
  thrashSessionState.hasWarnedFallbackUsed = true;
@@ -1320,6 +1364,7 @@ function wrapToolCallHandler(
1320
1364
  thrashConnectionFallbackSessionId,
1321
1365
  thrashConfig,
1322
1366
  kind,
1367
+ parentContext,
1323
1368
  );
1324
1369
 
1325
1370
  span.setAttribute(ATTR_MCP_METHOD_NAME, TOOLS_CALL_METHOD);
@@ -32,6 +32,38 @@
32
32
  */
33
33
  export const SPAN_EVENT_NAME_SCHEMA_DRIFT_DETECTED = 'mcp.tool.schema_drift.detected';
34
34
 
35
+ /**
36
+ * Boolean span ATTRIBUTE, distinct from SPAN_EVENT_NAME_SCHEMA_DRIFT_DETECTED
37
+ * above (v0.12.0). Same ambiguity, same resolution ADR 011
38
+ * (docs/adr/011-cost-aware-sampling.md) already applied to thrash detection:
39
+ * whether an OpenTelemetry Collector `tailsamplingprocessor`'s
40
+ * `boolean_attribute` policy can match span-EVENT data (as opposed to
41
+ * top-level span attributes) could not be confirmed either way — the
42
+ * processor is Go source in a separate repository, not installed here. ADR
43
+ * 011 resolved that uncertainty for thrash by adding a parallel attribute,
44
+ * `ATTR_MCP_TOOL_THRASH_DETECTED` (thrash/attributes.js); this is the
45
+ * identical fix applied here, for the identical reason. (This one wasn't
46
+ * added alongside ADR 010/011 originally — a v0.12.0 recipe-verification
47
+ * pass found the README recommending SPAN_EVENT_NAME_SCHEMA_DRIFT_DETECTED
48
+ * itself as a `boolean_attribute` policy target, which cannot work, since
49
+ * that string was never anything but an event/metric name.)
50
+ *
51
+ * Deliberately NOT `mcp.tool.schema_drift.detected` (SPAN_EVENT_NAME_SCHEMA_DRIFT_DETECTED's
52
+ * own string) — reusing it would reintroduce the exact ambiguity this
53
+ * attribute exists to remove: a bare string with no way to tell, from the
54
+ * name alone, whether a given consumer is reading the event or the
55
+ * attribute. Named `mcp.tool.schema_drift_detected` (underscore, no dot
56
+ * before "detected"), mirroring `ATTR_MCP_TOOL_THRASH_DETECTED`'s own
57
+ * naming (thrash/attributes.js: `mcp.tool.thrash_detected`, not
58
+ * `mcp.tool.loop.detected`, for the same reason).
59
+ *
60
+ * Only ever set to `true`, and only when drift was actually detected on
61
+ * this call — never explicitly set `false` for a clean call, matching
62
+ * `ATTR_MCP_TOOL_THRASH_DETECTED`'s same "omit rather than set a
63
+ * negative/empty value" convention.
64
+ */
65
+ export const ATTR_MCP_TOOL_SCHEMA_DRIFT_DETECTED = 'mcp.tool.schema_drift_detected';
66
+
35
67
  /** @type {Readonly<Record<'TYPE' | 'PREVIOUS_HASH' | 'CURRENT_HASH' | 'ADDED_FIELDS' | 'REMOVED_FIELDS' | 'CHANGED_FIELDS', string>>} */
36
68
  export const ATTRIBUTE_KEYS = Object.freeze({
37
69
  /**
@@ -15,7 +15,9 @@
15
15
  * instrument uses (`metrics.getMeter('opentel-mcp', packageVersion)` —
16
16
  * see src/metrics.js's own docblock on why a second getMeter() call
17
17
  * still reports under the same meter to any backend), plus one span
18
- * event on the currently active span.
18
+ * event and one boolean span attribute (v0.12.0, ADR 011,
19
+ * docs/adr/011-cost-aware-sampling.md — see ATTR_MCP_TOOL_SCHEMA_DRIFT_DETECTED's
20
+ * own docblock in ./attributes.js for why) on the currently active span.
19
21
  *
20
22
  * "No active span" handling follows the SAME precedent this module
21
23
  * mirrors, not an invented behavior: ADR 010's "What gets emitted"
@@ -38,7 +40,7 @@
38
40
 
39
41
  import { trace, metrics } from '@opentelemetry/api';
40
42
  import { ATTR_GEN_AI_TOOL_NAME } from '../attributes.js';
41
- import { SPAN_EVENT_NAME_SCHEMA_DRIFT_DETECTED, ATTRIBUTE_KEYS } from './attributes.js';
43
+ import { SPAN_EVENT_NAME_SCHEMA_DRIFT_DETECTED, ATTR_MCP_TOOL_SCHEMA_DRIFT_DETECTED, ATTRIBUTE_KEYS } from './attributes.js';
42
44
 
43
45
  /** @typedef {import('./types.d.ts').SchemaDriftEvent} SchemaDriftEvent */
44
46
 
@@ -66,7 +68,12 @@ export function createSchemaDriftEmitter(packageVersion) {
66
68
  * full detail, including the previous/current hashes and (only when
67
69
  * non-empty, mirroring `mcp.failure.validation_paths`'s
68
70
  * omit-rather-than-set-empty discipline, ADR 009) the changed field
69
- * names. Never throws: metrics and the span event are independently
71
+ * names, plus (v0.12.0, ADR 011) a boolean
72
+ * `mcp.tool.schema_drift_detected` span ATTRIBUTE on that same span —
73
+ * a Collector tail-sampling policy can key on it directly, without
74
+ * depending on whether span-event data is matchable at all (see
75
+ * ATTR_MCP_TOOL_SCHEMA_DRIFT_DETECTED's docblock in ./attributes.js).
76
+ * Never throws: metrics and the span event/attribute are independently
70
77
  * guarded, so a failure in one never suppresses the other, the same
71
78
  * fail-open philosophy as every other emitter in this codebase.
72
79
  *
@@ -98,6 +105,7 @@ export function createSchemaDriftEmitter(packageVersion) {
98
105
  if (event.changedFields?.length > 0) attrs[ATTRIBUTE_KEYS.CHANGED_FIELDS] = event.changedFields;
99
106
 
100
107
  span.addEvent(SPAN_EVENT_NAME_SCHEMA_DRIFT_DETECTED, attrs);
108
+ span.setAttribute(ATTR_MCP_TOOL_SCHEMA_DRIFT_DETECTED, true);
101
109
  } catch {
102
110
  // Never throw — see emit()'s docblock.
103
111
  }