opentel-mcp 0.10.0 → 0.11.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/CHANGELOG.md CHANGED
@@ -1,5 +1,142 @@
1
1
  # Changelog
2
2
 
3
+ ## 0.11.0
4
+
5
+ **⚠️ Type change, not a runtime behavior change — read this first.**
6
+ `ModelPricing` is now a discriminated union
7
+ (`{pricingKind: 'chat', inputPer1M, outputPer1M, currency} |
8
+ {pricingKind: 'embedding', inputPer1M, currency}`) instead of a single
9
+ shape with both token fields always present. A TypeScript consumer with
10
+ an existing custom `pricingTable`/`pricing` object typed against the old
11
+ shape will see a compile error requiring `pricingKind` on each entry.
12
+ **Runtime behavior for those same objects is unchanged**: `calculateCost()`
13
+ treats a missing or unrecognized `pricingKind` as `'chat'`, exactly the
14
+ behavior every pre-v0.11.0 entry already had. See ADR 016
15
+ (`docs/adr/016-pricing-override-and-staleness.md`) point 1.
16
+
17
+ ### Added — Pricing table override and staleness signalling (Phase 1 of 2)
18
+
19
+ `DEFAULT_PRICING` had the same disease this whole library exists to fix
20
+ elsewhere: a hardcoded snapshot with no signal when it's wrong or stale.
21
+ This release closes that. Full design: ADR 016
22
+ (`docs/adr/016-pricing-override-and-staleness.md`).
23
+
24
+ - **Embedding model support.** `DEFAULT_PRICING` gains OpenAI
25
+ `text-embedding-3-small`/`text-embedding-3-large`/`text-embedding-ada-002`,
26
+ Cohere `cohere-embed-v3`, and Bedrock `amazon-titan-embed-v2`.
27
+ Embeddings are input-token-only — rather than modeling that as
28
+ `outputPer1M: 0` (indistinguishable from a data-entry bug), a new
29
+ `pricingKind: 'chat' | 'embedding'` discriminator makes it explicit; an
30
+ `'embedding'` entry has no `outputPer1M` field at all, and
31
+ `calculateCost()` never reads `outputTokens` for one (still validated as
32
+ a non-negative finite number, just never charged for).
33
+ - **`costTracking.pricing`: per-model merge over defaults.** New option,
34
+ a *partial* pricing table merged per-model OVER `pricingTable ??
35
+ DEFAULT_PRICING` — each key you supply replaces that model's entire
36
+ pricing entry, every model you don't name is untouched. This is now the
37
+ recommended way to correct a stale price or add a model
38
+ `DEFAULT_PRICING` doesn't know about, without spreading the whole
39
+ default table by hand (the old workaround the README used to teach).
40
+ `costTracking.pricingTable` keeps its existing full-replace behavior,
41
+ unchanged, for the narrower "I want only my own models" case — see ADR
42
+ 016 point 2 for why both exist.
43
+ - **`mcp.tool.pricing_status` span + metric attribute.** One of `"known"`
44
+ | `"unknown"` | `"user_override"`, set whenever token usage was
45
+ extracted at all — even with no model detected (`"unknown"` in that
46
+ case), unlike the existing model/cost attributes. Added to the
47
+ `mcp.tool.tokens.total` / `mcp.tool.cost.total` metrics too, so a
48
+ dashboard can compute "% of tokens/spend unpriced" as a direct
49
+ aggregation instead of inferring it from missing data. Provenance-based:
50
+ `"user_override"` means the model's key came from your
51
+ `pricing`/`pricingTable`, regardless of whether the numbers you supplied
52
+ happen to match `DEFAULT_PRICING`'s own entry.
53
+ - **Staleness signalling.** `DEFAULT_PRICING_LAST_VERIFIED` (also
54
+ exported) names the table's last-checked date; once it's more than 90
55
+ days old, `instrumentMcpServer()` fires a one-time `diag.warn()`, and,
56
+ when `setupNodeSdk: true`, also attaches an
57
+ `mcp.pricing.default_table_last_verified` resource attribute. Both only
58
+ fire when `DEFAULT_PRICING` is actually contributing to the effective
59
+ table — a caller who fully replaced it via `pricingTable` isn't using
60
+ our defaults, so a warning about them would be misleading.
61
+ `isDefaultPricingStale(now?, thresholdDays?)` (also exported) is the
62
+ pure function behind the warning, for callers who want to check it
63
+ themselves.
64
+ - **Bedrock region caveat, documented not modeled.** Bedrock pricing
65
+ varies by region; `DEFAULT_PRICING`'s Bedrock entries (Nova, and the new
66
+ Titan embedding entry) assume us-east-1 list price and the table is not
67
+ region-keyed — no tool-result usage shape this package recognizes
68
+ carries a region signal to key a lookup on. Documented loudly in the
69
+ README and in `pricing.js`; override via `costTracking.pricing` for a
70
+ different region. See ADR 016 point 5.
71
+ - `calculateCost()` gained defensive validation for malformed pricing
72
+ entries (missing/negative/non-numeric `inputPer1M`/`outputPer1M`),
73
+ since `pricing`/`pricingTable` now make it reachable with
74
+ caller-supplied shapes it previously never had to distrust — degrades
75
+ to `null`, same as an unknown model, never throws.
76
+
77
+ ### Added — W3C Trace Context propagation over MCP `_meta` (Phase 2 of 2, server-side only)
78
+
79
+ The most-complained-about gap in agent observability: an agent's own
80
+ trace (LangGraph or otherwise) and the MCP server's trace for the tool
81
+ call it made were always two disconnected traces, with no edge between
82
+ them. Full design, including the sampling and conflicting-context
83
+ decisions below: ADR 017 (`docs/adr/017-trace-context-propagation.md`).
84
+
85
+ - **`tools/call` requests carrying a valid W3C `traceparent` in
86
+ `params._meta` now become a child of the calling agent's own span**,
87
+ joining what were two disconnected traces into one — under both MCP v1
88
+ and v2, with zero configuration and no new option. `tracestate` is
89
+ propagated too, when present. Works for any client already emitting
90
+ `traceparent` via a standard OTel SDK's `propagation.inject()` in any
91
+ language — this isn't Node/JS-specific on the client side, only on
92
+ which side of the wire this release implements.
93
+ - **The upstream sampling decision is honored automatically, by
94
+ construction, with no sampling logic written for this feature**: the
95
+ extracted `SpanContext` is marked `isRemote: true` with the real parsed
96
+ `traceFlags`, which is exactly what the SDK's own default
97
+ `ParentBasedSampler` already keys its remote-parent decision off of. A
98
+ not-sampled upstream `traceparent` means this tool-call span is not
99
+ recorded or exported, matching the calling agent's own choice — see the
100
+ ADR's "Sampling" section for why forcing sampling regardless was
101
+ considered and rejected.
102
+ - **A `_meta`-extracted context always replaces, never merges with, an
103
+ already-active local context** (e.g. an ambient HTTP-server span from
104
+ auto-instrumentation on a Streamable HTTP transport) — the message-level
105
+ `_meta` context is the semantically correct parent for one tool call,
106
+ full stop, regardless of what transport-level span it happened to
107
+ arrive inside. See the ADR's "Conflicting `_meta.traceparent`" section.
108
+ - **Absent, malformed, or unparseable `_meta`/`traceparent` produces
109
+ behavior that is byte-identical to pre-v0.11.0** — not merely
110
+ equivalent to it: confirmed by reading both `NoopTracer` and the real
111
+ SDK `Tracer`'s own `startActiveSpan()` fallback (`ctx ?? context.active()`),
112
+ which is exactly what this feature's `extractTraceContext()` returns
113
+ for every case that isn't a valid `traceparent`. No `diag.warn()` for
114
+ the common "client doesn't send `_meta.traceparent`" case — see the
115
+ ADR's "no warn spam" constraint.
116
+ - **Zero new dependencies.** `@opentelemetry/core`'s
117
+ `W3CTraceContextPropagator` was the obvious reference implementation
118
+ and was deliberately not taken as a dependency, per
119
+ `CONTRIBUTING.md`'s "no new dependencies without discussion first" —
120
+ everything needed except the traceparent regex itself (~10 lines,
121
+ matching `@opentelemetry/core`'s own validation field-for-field) was
122
+ already available from `@opentelemetry/api`, already a peer dependency
123
+ — including `createTraceState()`, a fully spec-validated `tracestate`
124
+ parser. Full reasoning: ADR 017's "No new dependency" section.
125
+ - **Server-side extraction only.** The client-side shim that would let a
126
+ Node/Python agent framework *set* `_meta.traceparent` on outgoing calls
127
+ is explicitly out of scope for this phase — extraction is independently
128
+ useful today, for free, to any client whose own tooling already sets
129
+ `_meta` in this shape. Tracked as future work, not implied as solved.
130
+
131
+ Also fixed in this release: two places (`index.d.ts`'s
132
+ `instrumentMcpServer()` docblock, and this file's own v0.10.0 entry
133
+ below) still described `docs/known-gaps.md` entries 6/7/8 using language
134
+ that read as still-open, or as scoped out of v0.10.0 — both were stale.
135
+ Entries 7 and 8 have been fully fixed since v0.10.0 with no open caveats;
136
+ entry 6's fallback-session-id half is fixed too, narrowed to a smaller,
137
+ genuinely-still-open remainder (see the corrected v0.10.0 entry below and
138
+ `index.d.ts`'s updated docblock for the accurate, current accounting).
139
+
3
140
  ## 0.10.0
4
141
 
5
142
  **⚠️ Behavior change, unrelated to the feature below — read this first.**
@@ -60,15 +197,32 @@ new was added for this, since ADR 012's original design already covers
60
197
  this exact deployment shape, v2 just makes it the default instead of an
61
198
  edge case.
62
199
 
63
- **Two gaps not closed this release, both tracked in `docs/known-gaps.md`
64
- with a "Status update (v0.10.0)" note:** Agent Thrash Detection's fallback
65
- session id still doesn't survive v2's per-request factory pattern even
66
- with `instanceKey` set (entry 6 — real session ids work fine either way),
67
- and `isSingleConnectionTransport()`'s transport-detection heuristic still
68
- misclassifies the transport `createMcpHandler` builds internally (entry
69
- 8, live as of this release, not merely forward-looking). Both were
70
- explicitly scoped out of this round, not overlooked; workaround for
71
- either: `thrashDetection: { enabled: false }`.
200
+ **Correction (recorded here rather than silently edited): both gaps below
201
+ were actually closed in this same v0.10.0 release, not left open.** The
202
+ paragraph originally here said entries 6 and 8 (`docs/known-gaps.md`)
203
+ were scoped out of this round — true of the round that produced the text
204
+ above, not of what actually shipped. A follow-up investigation, completed
205
+ before v0.10.0 was cut, folded both fixes back in: `isSingleConnectionTransport()`
206
+ no longer misclassifies the transport `createMcpHandler` builds internally
207
+ (entry 8 — fully fixed: v2 now requires positive confirmation,
208
+ `transport.constructor.name === 'StdioServerTransport'`, instead of
209
+ inferring single-connection from an absent `sessionId` property), and
210
+ Agent Thrash Detection's fallback session id is now registry-backed via
211
+ `instanceKey` (entry 6's fallback-id gap — fixed: repeated
212
+ `instrumentMcpServer()` calls sharing an `instanceKey` now reuse the same
213
+ generated id instead of a fresh one per call). Both fixes shipped in the
214
+ same commit, in a specific order — fixing detection (entry 8) before
215
+ sharing the fallback id (entry 6) — since sharing it first would have made
216
+ entry 8's false positive worse, not better.
217
+
218
+ **What remains genuinely open, narrower than either original gap:**
219
+ `thrashSessionState` (whether a server has ever proven itself
220
+ session-aware) still isn't registry-backed, and — structurally, not a
221
+ bug this library can fix — MCP spec 2026-07-28 removes protocol-level
222
+ sessions entirely, so no configuration of this library can produce a
223
+ *real* session id for a spec-2026-07-28-native deployment in the first
224
+ place. See ADR 015's final "Update ... Findings 3 and 8 landed here too"
225
+ section and `docs/known-gaps.md` entries 6 and 8 for the full accounting.
72
226
 
73
227
  ## 0.9.0
74
228
 
package/README.md CHANGED
@@ -495,19 +495,20 @@ point, or JSON-in-text inside `content[0].text`) automatically gets
495
495
  `mcp.tool.tokens.*` / `mcp.tool.model` / `mcp.tool.cost.*` span
496
496
  attributes, priced against `DEFAULT_PRICING`.
497
497
 
498
- ### Advanced: custom pricing, a custom extractor, and a budget guardrail
498
+ ### Advanced: overriding pricing, a custom extractor, and a budget guardrail
499
499
 
500
500
  ```js
501
- import { instrumentMcpServer, DEFAULT_PRICING } from 'opentel-mcp';
501
+ import { instrumentMcpServer } from 'opentel-mcp';
502
502
 
503
503
  instrumentMcpServer(server, {
504
504
  serviceName: 'my-mcp-server',
505
505
  costTracking: {
506
- // Extend or override DEFAULT_PRICING — e.g. price an internal model
507
- // it doesn't know about, or correct stale numbers.
508
- pricingTable: {
509
- ...DEFAULT_PRICING,
510
- 'my-internal-model': { inputPer1M: 1.0, outputPer1M: 2.0, currency: 'USD' },
506
+ // Merged per-model OVER DEFAULT_PRICING — correct a stale price or add
507
+ // a model DEFAULT_PRICING doesn't know about, without having to spread
508
+ // the whole default table yourself. Everything you don't name here is
509
+ // untouched. See "Overriding pricing" below.
510
+ pricing: {
511
+ 'my-internal-model': { pricingKind: 'chat', inputPer1M: 1.0, outputPer1M: 2.0, currency: 'USD' },
511
512
  },
512
513
  // Recognize your own tool result shape. Return null for anything you
513
514
  // don't recognize — never throw (see src/cost/extractor.js).
@@ -526,6 +527,41 @@ instrumentMcpServer(server, {
526
527
  });
527
528
  ```
528
529
 
530
+ ### Overriding pricing (v0.11.0+)
531
+
532
+ Two ways to change what `costTracking` prices against, composing as
533
+ `{ ...(pricingTable ?? DEFAULT_PRICING), ...pricing }` — see ADR 016
534
+ (`docs/adr/016-pricing-override-and-staleness.md`) for the full reasoning:
535
+
536
+ - **`costTracking.pricing`** — a *partial* table, merged **per-model**
537
+ over `DEFAULT_PRICING` (or over `pricingTable`, if you set both). Each
538
+ key you supply replaces that one model's entire pricing entry; every
539
+ model you don't name keeps its default price. This is the recommended
540
+ way to correct a stale number or add a model — enterprises on
541
+ committed-use discounts, AWS EDP, Bedrock provisioned throughput, or
542
+ Azure OpenAI negotiated rates should use this to reflect what they
543
+ actually pay, not list price.
544
+ - **`costTracking.pricingTable`** — *fully replaces* `DEFAULT_PRICING`.
545
+ Use this when you want an effective table containing **only** your own
546
+ models, none of `DEFAULT_PRICING`'s.
547
+
548
+ A model priced via either option reports `pricing_status: 'user_override'`
549
+ (see the span attributes table below) instead of `'known'`.
550
+
551
+ ### Embedding models (v0.11.0+)
552
+
553
+ Embeddings are input-token-only — there's no "output" to price. Rather
554
+ than modeling that as `outputPer1M: 0` (indistinguishable from a
555
+ data-entry bug), `ModelPricing` carries an explicit
556
+ `pricingKind: 'chat' | 'embedding'` discriminator; an `'embedding'` entry
557
+ has no `outputPer1M` field at all, and `calculateCost()` never reads
558
+ `outputTokens` for one. `DEFAULT_PRICING` includes OpenAI
559
+ `text-embedding-3-small`/`text-embedding-3-large`/`text-embedding-ada-002`,
560
+ Cohere `cohere-embed-v3`, and Bedrock `amazon-titan-embed-v2` out of the
561
+ box. A pre-v0.11.0 custom pricing entry with no `pricingKind` field is
562
+ still treated as `'chat'` at runtime — only the TypeScript type is
563
+ stricter, not the runtime.
564
+
529
565
  ### Span attributes
530
566
 
531
567
  | Attribute | Standard OTel? | Description | Example |
@@ -535,19 +571,23 @@ instrumentMcpServer(server, {
535
571
  | `mcp.tool.tokens.total` | Custom | input + output | 1500 |
536
572
  | `mcp.tool.model` | Custom | Detected model name | "claude-sonnet-5" |
537
573
  | `gen_ai.response.model` | Standard (GenAI semconv)[^5] | Same value as `mcp.tool.model`, co-emitted for dashboard compatibility | "claude-sonnet-5" |
574
+ | `mcp.tool.pricing_status` | Custom[^7] | `"known"` \| `"unknown"` \| `"user_override"` — set whenever token usage was extracted, even with no model detected | "known" |
538
575
  | `mcp.tool.cost.usd` | Custom | Estimated cost, from `calculateCost()` | 0.0105 |
539
576
  | `mcp.tool.cost.currency` | Custom | Always `"USD"` today | "USD" |
540
577
  | `mcp.tool.cost.budget_exceeded` | Custom | `true` once a configured `costTracking.budget` limit is crossed | true |
541
578
  | `mcp.tool.cost.budget_scope` | Custom | Which budget scope tripped: `"session"` \| `"tool"` (session wins if both did) | "session" |
542
579
 
543
580
  [^5]: `gen_ai.response.model` is a real OTel GenAI semantic convention attribute ("the name of the model that generated the response") — but this span is an MCP tool-call span (`gen_ai.operation.name: execute_tool`), not a dedicated LLM request/response span, so co-emitting it here is a **pragmatic dashboard-compatibility choice, not a spec-pure emission**. It's set purely so off-the-shelf GenAI dashboards (Grafana, SigNoz, Honeycomb) that filter/group by `gen_ai.response.model` pick these spans up without any opentel-mcp-specific configuration. Full reasoning in `src/attributes.js`'s `ATTR_GEN_AI_RESPONSE_MODEL` docblock.
581
+ [^7]: `mcp.tool.pricing_status` is provenance-based, not value-based: `"user_override"` means the model's key was present in your `costTracking.pricing`/`pricingTable`, whether or not the numbers you supplied happen to match `DEFAULT_PRICING`. `"unknown"` covers both "no model detected" and "model detected but not priceable" (unrecognized, or a malformed override entry) — see ADR 016 point 4.
544
582
 
545
- The four token/model attributes are set together or not at all; the two
546
- cost attributes only appear when a model was detected *and* it resolves
547
- in the configured `pricingTable`; the two budget attributes only appear
548
- when a cost was calculated *and* a configured limit was crossed. Source
549
- of truth: `src/attributes.js` and `src/instrument.js`'s
550
- `applyCostAttribution()`.
583
+ The four token/model attributes are set together or not at all;
584
+ `mcp.tool.pricing_status` is set whenever usage was extracted at all
585
+ (unlike the model/cost attributes, it's present even with no model
586
+ detected); the two cost attributes only appear when a model was detected
587
+ *and* it resolves in the effective pricing table; the two budget
588
+ attributes only appear when a cost was calculated *and* a configured
589
+ limit was crossed. Source of truth: `src/attributes.js` and
590
+ `src/instrument.js`'s `applyCostAttribution()`.
551
591
 
552
592
  ### Metrics
553
593
 
@@ -557,21 +597,38 @@ registered, `enableMetrics: false` opts out of these too.
557
597
 
558
598
  | Metric | Type | Unit | Attributes | Emitted when |
559
599
  |---|---|---|---|---|
560
- | `mcp.tool.tokens.total` | Counter | tokens | `gen_ai.tool.name`, `mcp.tool.model`[^6] | Usage detected in the tool result |
561
- | `mcp.tool.cost.total` | Counter | USD | `gen_ai.tool.name`, `mcp.tool.model`[^6] | Cost calculated (model resolved in `pricingTable`) |
562
-
563
- [^6]: `mcp.tool.model` is only added when a model was detected — the same optional-attribute cardinality pattern `mcp.failure.category` already uses on the other four metrics.
564
-
565
- ### Pricing accuracy
566
-
567
- > **Pricing table last verified 2026-07-29.** Users **MUST** override
568
- > `pricingTable` for production accuracy — provider pricing changes
569
- > frequently and opentel-mcp does not guarantee `DEFAULT_PRICING` stays
570
- > current.
571
-
572
- `DEFAULT_PRICING` (`src/cost/pricing.js`) covers 15+ models across five
573
- providers — Anthropic, OpenAI, Google, AWS Bedrock, and DeepSeek — as a
574
- convenience default, not a maintained price list.
600
+ | `mcp.tool.tokens.total` | Counter | tokens | `gen_ai.tool.name`, `mcp.tool.model`[^6], `mcp.tool.pricing_status` | Usage detected in the tool result |
601
+ | `mcp.tool.cost.total` | Counter | USD | `gen_ai.tool.name`, `mcp.tool.model`[^6], `mcp.tool.pricing_status` | Cost calculated (model resolved in the effective pricing table) |
602
+
603
+ [^6]: `mcp.tool.model` is only added when a model was detected — the same optional-attribute cardinality pattern `mcp.failure.category` already uses on the other four metrics. `mcp.tool.pricing_status` is always added — a fixed, closed 3-value enum, well within this package's metric-label cardinality discipline (see `COST_METRIC_SAFE_ATTRIBUTES` in `src/attributes.js`).
604
+
605
+ Grouping `mcp.tool.tokens.total` by `mcp.tool.pricing_status` answers "what
606
+ fraction of tokens/spend is running through models we can't price" as a
607
+ direct query, instead of inferring it from missing `mcp.tool.cost.*` data.
608
+
609
+ ### Pricing accuracy and staleness
610
+
611
+ > **`DEFAULT_PRICING` is a best-effort snapshot, not a maintained price
612
+ > list.** Provider pricing changes frequently and varies by region/contract
613
+ > — `opentel-mcp` does not guarantee it stays current, and says so at
614
+ > runtime, not just here: once `DEFAULT_PRICING` (checked via
615
+ > `DEFAULT_PRICING_LAST_VERIFIED`, also exported) is more than 90 days
616
+ > past its last-verified date, `instrumentMcpServer()` fires a one-time
617
+ > `diag.warn()` naming that date — and, when `setupNodeSdk: true`, also
618
+ > attaches an `mcp.pricing.default_table_last_verified` resource
619
+ > attribute, so long-running deployments can alert on it directly. Both
620
+ > only fire when `DEFAULT_PRICING` is actually contributing to your
621
+ > effective table (i.e. you haven't fully replaced it via `pricingTable`)
622
+ > — see ADR 016 point 3.
623
+
624
+ `DEFAULT_PRICING` (`src/cost/pricing.js`) covers 20+ models — chat and
625
+ embedding — across six providers: Anthropic, OpenAI, Google, Cohere, AWS
626
+ Bedrock, and DeepSeek. **Bedrock entries (Nova and Titan embeddings)
627
+ assume us-east-1 list pricing** — Bedrock pricing varies by region and
628
+ this table is not region-keyed (no reliable region signal exists in any
629
+ tool-result usage shape this package recognizes to key a lookup on — see
630
+ ADR 016 point 5); override via `costTracking.pricing` for a different
631
+ region.
575
632
 
576
633
  ### Extending it
577
634
 
@@ -581,7 +638,12 @@ convenience default, not a maintained price list.
581
638
  TokenUsage | null`, never throwing) to recognize anything else.
582
639
  - `calculateCost(inputTokens, outputTokens, model, pricingTable)` is also
583
640
  exported directly, for recomputing cost outside the instrumentation
584
- hot path (e.g. over historical spans).
641
+ hot path (e.g. over historical spans). Malformed pricing entries
642
+ (missing/negative/non-numeric `inputPer1M`/`outputPer1M`) degrade to
643
+ `null`, same as an unknown model — never throws.
644
+ - `isDefaultPricingStale(now?, thresholdDays?)` (also exported) is the
645
+ pure function backing the staleness warning above, if you want to check
646
+ it yourself (e.g. in a startup health check).
585
647
  - Disable everything in this section with `costTracking: { enabled:
586
648
  false }`; tracing, metrics, and fingerprinting are all unaffected.
587
649
  - Budget tracking (`costTracking.budget`) is in-memory and scoped to one
@@ -1525,6 +1587,84 @@ your schema-drift or two-axis observation signals matter for retention
1525
1587
  too, add `mcp.tool.schema_drift.detected` (boolean_attribute) as another
1526
1588
  OR'd policy the same way.
1527
1589
 
1590
+ ## Trace Context Propagation (v0.11.0+)
1591
+
1592
+ The most-complained-about gap in agent observability: an agent framework
1593
+ (LangGraph or otherwise) calls a tool on your MCP server, and you get two
1594
+ disconnected traces — the agent's own trace dies at the tool boundary,
1595
+ and the server's trace for what the tool actually did starts fresh, with
1596
+ no edge between them. Debugging a slow or failing agent run means
1597
+ manually correlating timestamps across two separate traces (or two
1598
+ separate services in the same backend) instead of looking at one.
1599
+
1600
+ opentel-mcp closes the server-side half of this: when a `tools/call`
1601
+ request carries a valid [W3C `traceparent`](https://www.w3.org/TR/trace-context/)
1602
+ in `params._meta`, the tool-call span becomes a **child of the calling
1603
+ agent's own span** — the same trace, not two. `tracestate` is propagated
1604
+ too, when present.
1605
+
1606
+ ### Zero-config — nothing to opt into
1607
+
1608
+ There's no `traceContext: { enabled: false }` option, deliberately: this
1609
+ is unconditional, because there's nothing for an operator to want to turn
1610
+ off. Any client that already emits `traceparent` via a standard OTel
1611
+ SDK's `propagation.inject()` — in any language, this isn't Node-specific
1612
+ on the client side — gets linked traces the moment it copies that string
1613
+ into `_meta.traceparent` on its outgoing `tools/call` request:
1614
+
1615
+ ```json
1616
+ {
1617
+ "method": "tools/call",
1618
+ "params": {
1619
+ "name": "search_docs",
1620
+ "arguments": { "query": "..." },
1621
+ "_meta": {
1622
+ "traceparent": "00-4bf92f3577b34da6a3ce929d0e0e4736-00f067aa0ba902b7-01",
1623
+ "tracestate": "vendorname=opaquevalue"
1624
+ }
1625
+ }
1626
+ }
1627
+ ```
1628
+
1629
+ A client that doesn't set `_meta.traceparent` (today's overwhelming
1630
+ majority, until client-side shims exist — see "What's not built yet"
1631
+ below) sees **zero change of any kind** — confirmed byte-identical to
1632
+ pre-v0.11.0 behavior, not merely "should behave the same."
1633
+
1634
+ ### The upstream sampling decision is honored automatically
1635
+
1636
+ If the calling agent's own trace wasn't sampled (the `traceparent`'s
1637
+ flags byte has the sampled bit unset), this tool-call span isn't recorded
1638
+ or exported either — matching the agent's own choice, via the OTel SDK's
1639
+ own default `ParentBasedSampler`, which already inspects a remote
1640
+ parent's `traceFlags` for exactly this. opentel-mcp writes no sampling
1641
+ logic of its own for this — see ADR 017's "Sampling" section
1642
+ (`docs/adr/017-trace-context-propagation.md`) for why forcing sampling
1643
+ regardless was considered and rejected.
1644
+
1645
+ ### Conflicting or pre-existing context
1646
+
1647
+ If your transport's own auto-instrumentation (e.g.
1648
+ `@opentelemetry/instrumentation-http` on a Streamable HTTP server) has
1649
+ already put some other span active by the time this package's handler
1650
+ runs, a valid `_meta.traceparent` **replaces it outright** — it is never
1651
+ merged. The `_meta`-carried context is message-scoped (this one logical
1652
+ agent operation); the transport-level span is connection/request-scoped
1653
+ and isn't the right parent for it. Full reasoning: ADR 017's
1654
+ "Conflicting `_meta.traceparent`" section.
1655
+
1656
+ ### What's not built yet
1657
+
1658
+ **Server-side extraction only.** There is no client-side shim in this
1659
+ package (yet) for a Node or Python agent framework to *set*
1660
+ `_meta.traceparent` on its own outgoing calls — you need a client that
1661
+ already does this itself (or does it via your own glue code:
1662
+ `propagation.inject(context.active(), meta, ...)` from your OTel SDK,
1663
+ writing into `_meta` before the call). Extraction is independently
1664
+ useful today, for free, to any client that already sets `_meta` in this
1665
+ shape; the client-side half is tracked as future work, not implied as
1666
+ solved by this release.
1667
+
1528
1668
  ## Configuration
1529
1669
 
1530
1670
  All options passed to `instrumentMcpServer(server, options)`. Source of
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "opentel-mcp",
3
- "version": "0.10.0",
3
+ "version": "0.11.0",
4
4
  "description": "One-line OpenTelemetry instrumentation for Model Context Protocol (MCP) servers",
5
5
  "type": "module",
6
6
  "main": "src/index.js",
@@ -74,6 +74,7 @@
74
74
  "devDependencies": {
75
75
  "@modelcontextprotocol/server": "^2.0.0",
76
76
  "@opentelemetry/api": "^1.9.0",
77
+ "@opentelemetry/context-async-hooks": "^2.9.0",
77
78
  "@opentelemetry/sdk-metrics": "^2.9.0",
78
79
  "@opentelemetry/sdk-trace-base": "^2.9.0",
79
80
  "@vitest/coverage-v8": "^2.1.9",
package/src/attributes.js CHANGED
@@ -168,3 +168,48 @@ export const MCP_TOOL_COST_BUDGET_SCOPE_SESSION = 'session';
168
168
 
169
169
  /** Well-known mcp.tool.cost.budget_scope value: costTracking.budget.perToolUsd was exceeded. */
170
170
  export const MCP_TOOL_COST_BUDGET_SCOPE_TOOL = 'tool';
171
+
172
+ // --- Pricing provenance (v0.11.0, non-spec, ADR 016) ---
173
+ //
174
+ // docs/adr/016-pricing-override-and-staleness.md. Lets a dashboard answer
175
+ // "what fraction of spend/tokens is unpriced" as a direct query instead of
176
+ // an inference from missing mcp.tool.cost.* attributes.
177
+
178
+ /**
179
+ * Whether the model detected for this tool call resolved to a known price.
180
+ * Set whenever token usage was extracted at all (same gating as
181
+ * ATTR_MCP_TOOL_TOKENS_INPUT/OUTPUT/TOTAL) — present even when no model was
182
+ * detected, unlike ATTR_MCP_TOOL_MODEL/ATTR_MCP_TOOL_COST_USD. One of
183
+ * MCP_TOOL_PRICING_STATUS_KNOWN / _UNKNOWN / _USER_OVERRIDE.
184
+ */
185
+ export const ATTR_MCP_TOOL_PRICING_STATUS = 'mcp.tool.pricing_status';
186
+
187
+ /** Well-known mcp.tool.pricing_status value: no model was detected, or the detected model has no pricing entry. */
188
+ export const MCP_TOOL_PRICING_STATUS_UNKNOWN = 'unknown';
189
+
190
+ /** Well-known mcp.tool.pricing_status value: priced against an unmodified DEFAULT_PRICING entry. */
191
+ export const MCP_TOOL_PRICING_STATUS_KNOWN = 'known';
192
+
193
+ /** Well-known mcp.tool.pricing_status value: priced against a caller-supplied costTracking.pricing/pricingTable entry. */
194
+ export const MCP_TOOL_PRICING_STATUS_USER_OVERRIDE = 'user_override';
195
+
196
+ /**
197
+ * Attribute keys safe to attach to mcp.tool.tokens.total / mcp.tool.cost.total
198
+ * metric labels, for the cost/pricing domain specifically. Mirrors the
199
+ * governance pattern fingerprint/attributes.js's METRIC_SAFE_ATTRIBUTES and
200
+ * schema-drift/attributes.js's own same-named export already established
201
+ * for their domains — this domain gets its own list rather than borrowing
202
+ * either of those (mixing an unrelated domain's cardinality reasoning in
203
+ * here would obscure which ADR covers which attribute). Not re-exported
204
+ * from index.js, same as schema-drift's — internal governance, not public
205
+ * API. mcp.tool.model remains governed the way it already was, an inline
206
+ * cardinality comment in metrics.js, not a list entry here — see ADR 016
207
+ * point 4 for why that attribute's boundedness argument doesn't fit a
208
+ * fixed-enum list cleanly the way this one does.
209
+ *
210
+ * @type {readonly string[]}
211
+ */
212
+ export const COST_METRIC_SAFE_ATTRIBUTES = Object.freeze([ATTR_MCP_TOOL_PRICING_STATUS]);
213
+
214
+ /** Resource attribute (setupNodeSdk: true only) naming DEFAULT_PRICING's lastVerified date. See ADR 016 point 3. */
215
+ export const ATTR_MCP_PRICING_DEFAULT_TABLE_LAST_VERIFIED = 'mcp.pricing.default_table_last_verified';
package/src/config.js CHANGED
@@ -4,8 +4,9 @@
4
4
  */
5
5
 
6
6
  import { diag } from '@opentelemetry/api';
7
- import { DEFAULT_PRICING } from './cost/pricing.js';
7
+ import { DEFAULT_PRICING, DEFAULT_PRICING_LAST_VERIFIED, isDefaultPricingStale } from './cost/pricing.js';
8
8
  import { defaultExtractor } from './cost/extractor.js';
9
+ import { normalizeModelName } from './cost/calculator.js';
9
10
  import { resolveThrashConfig } from './thrash/config.js';
10
11
  import { resolveSchemaDriftConfig } from './schema-drift/config.js';
11
12
 
@@ -13,9 +14,16 @@ import { resolveSchemaDriftConfig } from './schema-drift/config.js';
13
14
  * @typedef {object} CostTrackingOptions
14
15
  * @property {boolean} [enabled=true] - Set to false to disable cost/token span attributes and the
15
16
  * mcp.tool.tokens.total / mcp.tool.cost.total metrics entirely.
16
- * @property {import('./cost/pricing.js').PricingTable} [pricingTable] - Overrides DEFAULT_PRICING
17
- * (src/cost/pricing.js). Supply your own table to price models DEFAULT_PRICING doesn't know about, or to
18
- * correct stale pricing — see that module's docblock.
17
+ * @property {import('./cost/pricing.js').PricingTable} [pricingTable] - Fully replaces DEFAULT_PRICING (or,
18
+ * if `pricing` below is also set, replaces the base table `pricing` is merged over). Supply your own table
19
+ * when you want an effective table containing ONLY your own models. For correcting/adding a few models
20
+ * while keeping the rest of DEFAULT_PRICING, prefer `pricing` instead — see ADR 016
21
+ * (docs/adr/016-pricing-override-and-staleness.md) point 2.
22
+ * @property {Partial<import('./cost/pricing.js').PricingTable>} [pricing] - Partial pricing table, merged
23
+ * per-model OVER `pricingTable ?? DEFAULT_PRICING` — each key you supply replaces that model's entire
24
+ * ModelPricing entry; every model you don't name is untouched. The recommended way to correct stale
25
+ * pricing or add a model DEFAULT_PRICING doesn't know about. A model priced via this option (or via
26
+ * `pricingTable`) reports `pricing_status: 'user_override'` — see ADR 016 point 2.
19
27
  * @property {import('./cost/extractor.js').UsageExtractor} [extractor] - Overrides defaultExtractor
20
28
  * (src/cost/extractor.js). Supply your own to recognize a tool result shape defaultExtractor doesn't.
21
29
  * @property {import('./cost/budget.js').BudgetConfig} [budget] - Per-session and per-tool cumulative-cost
@@ -113,6 +121,24 @@ export function __resetServiceNameWarnedForTests() {
113
121
  warnedServiceNameIgnored = false;
114
122
  }
115
123
 
124
+ // Guards the DEFAULT_PRICING staleness diagnostic below (ADR 016 point 3),
125
+ // same one-per-process pattern as warnedServiceNameIgnored above.
126
+ let warnedPricingStale = false;
127
+
128
+ // Test-only, same purpose as __resetServiceNameWarnedForTests above. Not
129
+ // part of the public API.
130
+ export function __resetPricingStaleWarnedForTests() {
131
+ warnedPricingStale = false;
132
+ }
133
+
134
+ /**
135
+ * @param {unknown} value
136
+ * @returns {value is Record<string, unknown>}
137
+ */
138
+ function isPlainObject(value) {
139
+ return typeof value === 'object' && value !== null && !Array.isArray(value);
140
+ }
141
+
116
142
  // ADR 012, Phase 2: the first env var this codebase reads for a bare
117
143
  // top-level InstrumentOptions field, not one nested inside a feature's own
118
144
  // sub-config (contrast OTEL_MCP_THRASH_*/OTEL_MCP_SCHEMA_DRIFT_*, both
@@ -171,6 +197,50 @@ export function resolveOptions(options) {
171
197
  }
172
198
 
173
199
  const rawCostTracking = opts.costTracking ?? {};
200
+ const costTrackingEnabled = rawCostTracking.enabled ?? true;
201
+
202
+ // ADR 016 point 2: pricingTable (when supplied) is the BASE table,
203
+ // fully replacing DEFAULT_PRICING; pricing (when supplied) is then
204
+ // merged per-model OVER that base — a plain object spread is exactly
205
+ // "per-model, not top-level replacement" here, since each spread key is
206
+ // one model. Non-object pricingTable/pricing values (a caller's typo,
207
+ // e.g. a string or null) are silently ignored rather than spread, same
208
+ // "malformed config degrades, never crashes" discipline calculateCost()
209
+ // itself follows for a malformed individual entry.
210
+ const hasCustomPricingTable = isPlainObject(rawCostTracking.pricingTable);
211
+ const hasPricingOverride = isPlainObject(rawCostTracking.pricing);
212
+ const basePricingTable = hasCustomPricingTable ? rawCostTracking.pricingTable : DEFAULT_PRICING;
213
+ const pricingTable = hasPricingOverride ? { ...basePricingTable, ...rawCostTracking.pricing } : basePricingTable;
214
+
215
+ // ADR 016 point 4: which normalized model keys came from the caller's
216
+ // own config surface (pricingTable and/or pricing), for the
217
+ // mcp.tool.pricing_status 'known' vs 'user_override' distinction —
218
+ // computed once here, not per call. Provenance-based: a key counts as
219
+ // an override because the caller named it, regardless of whether the
220
+ // value they supplied happens to match DEFAULT_PRICING's own entry.
221
+ const pricingOverrideKeys = new Set();
222
+ if (hasCustomPricingTable) {
223
+ for (const key of Object.keys(rawCostTracking.pricingTable)) pricingOverrideKeys.add(normalizeModelName(key));
224
+ }
225
+ if (hasPricingOverride) {
226
+ for (const key of Object.keys(rawCostTracking.pricing)) pricingOverrideKeys.add(normalizeModelName(key));
227
+ }
228
+
229
+ // ADR 016 point 3: only warn when DEFAULT_PRICING is actually
230
+ // contributing to the effective table — a caller who fully replaced it
231
+ // via pricingTable isn't using our defaults at all, so a staleness
232
+ // warning about them would be misleading. Gated on costTrackingEnabled
233
+ // too: no cost tracking happens at all otherwise, so DEFAULT_PRICING's
234
+ // age is moot.
235
+ if (costTrackingEnabled && !hasCustomPricingTable && !warnedPricingStale && isDefaultPricingStale()) {
236
+ warnedPricingStale = true;
237
+ diag.warn(
238
+ `opentel-mcp: DEFAULT_PRICING was last verified ${DEFAULT_PRICING_LAST_VERIFIED}, more than 90 days ago. ` +
239
+ 'Provider list pricing may have changed since. Override costTracking.pricing (merged per-model over ' +
240
+ 'DEFAULT_PRICING) for models whose pricing you need to keep current — see ADR 016, ' +
241
+ 'docs/adr/016-pricing-override-and-staleness.md.',
242
+ );
243
+ }
174
244
 
175
245
  return {
176
246
  serviceName: opts.serviceName,
@@ -180,8 +250,10 @@ export function resolveOptions(options) {
180
250
  setupNodeSdk,
181
251
  fingerprinting: opts.fingerprinting ?? true,
182
252
  costTracking: {
183
- enabled: rawCostTracking.enabled ?? true,
184
- pricingTable: rawCostTracking.pricingTable ?? DEFAULT_PRICING,
253
+ enabled: costTrackingEnabled,
254
+ pricingTable,
255
+ pricingOverrideKeys,
256
+ usingDefaultPricing: !hasCustomPricingTable,
185
257
  extractor: rawCostTracking.extractor ?? defaultExtractor,
186
258
  budget: rawCostTracking.budget,
187
259
  },