opentel-mcp 0.10.0 → 0.11.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +163 -9
- package/README.md +169 -29
- package/package.json +2 -1
- package/src/attributes.js +45 -0
- package/src/config.js +78 -6
- package/src/cost/calculator.d.ts +14 -2
- package/src/cost/calculator.js +48 -17
- package/src/cost/pricing.d.ts +18 -1
- package/src/cost/pricing.js +91 -26
- package/src/cost/types.d.ts +42 -5
- package/src/index.d.ts +28 -11
- package/src/index.js +1 -1
- package/src/instrument.js +70 -16
- package/src/metrics.js +14 -4
- package/src/tracecontext/extract.d.ts +10 -0
- package/src/tracecontext/extract.js +128 -0
package/CHANGELOG.md
CHANGED
|
@@ -1,5 +1,142 @@
|
|
|
1
1
|
# Changelog
|
|
2
2
|
|
|
3
|
+
## 0.11.0
|
|
4
|
+
|
|
5
|
+
**⚠️ Type change, not a runtime behavior change — read this first.**
|
|
6
|
+
`ModelPricing` is now a discriminated union
|
|
7
|
+
(`{pricingKind: 'chat', inputPer1M, outputPer1M, currency} |
|
|
8
|
+
{pricingKind: 'embedding', inputPer1M, currency}`) instead of a single
|
|
9
|
+
shape with both token fields always present. A TypeScript consumer with
|
|
10
|
+
an existing custom `pricingTable`/`pricing` object typed against the old
|
|
11
|
+
shape will see a compile error requiring `pricingKind` on each entry.
|
|
12
|
+
**Runtime behavior for those same objects is unchanged**: `calculateCost()`
|
|
13
|
+
treats a missing or unrecognized `pricingKind` as `'chat'`, exactly the
|
|
14
|
+
behavior every pre-v0.11.0 entry already had. See ADR 016
|
|
15
|
+
(`docs/adr/016-pricing-override-and-staleness.md`) point 1.
|
|
16
|
+
|
|
17
|
+
### Added — Pricing table override and staleness signalling (Phase 1 of 2)
|
|
18
|
+
|
|
19
|
+
`DEFAULT_PRICING` had the same disease this whole library exists to fix
|
|
20
|
+
elsewhere: a hardcoded snapshot with no signal when it's wrong or stale.
|
|
21
|
+
This release closes that. Full design: ADR 016
|
|
22
|
+
(`docs/adr/016-pricing-override-and-staleness.md`).
|
|
23
|
+
|
|
24
|
+
- **Embedding model support.** `DEFAULT_PRICING` gains OpenAI
|
|
25
|
+
`text-embedding-3-small`/`text-embedding-3-large`/`text-embedding-ada-002`,
|
|
26
|
+
Cohere `cohere-embed-v3`, and Bedrock `amazon-titan-embed-v2`.
|
|
27
|
+
Embeddings are input-token-only — rather than modeling that as
|
|
28
|
+
`outputPer1M: 0` (indistinguishable from a data-entry bug), a new
|
|
29
|
+
`pricingKind: 'chat' | 'embedding'` discriminator makes it explicit; an
|
|
30
|
+
`'embedding'` entry has no `outputPer1M` field at all, and
|
|
31
|
+
`calculateCost()` never reads `outputTokens` for one (still validated as
|
|
32
|
+
a non-negative finite number, just never charged for).
|
|
33
|
+
- **`costTracking.pricing`: per-model merge over defaults.** New option,
|
|
34
|
+
a *partial* pricing table merged per-model OVER `pricingTable ??
|
|
35
|
+
DEFAULT_PRICING` — each key you supply replaces that model's entire
|
|
36
|
+
pricing entry, every model you don't name is untouched. This is now the
|
|
37
|
+
recommended way to correct a stale price or add a model
|
|
38
|
+
`DEFAULT_PRICING` doesn't know about, without spreading the whole
|
|
39
|
+
default table by hand (the old workaround the README used to teach).
|
|
40
|
+
`costTracking.pricingTable` keeps its existing full-replace behavior,
|
|
41
|
+
unchanged, for the narrower "I want only my own models" case — see ADR
|
|
42
|
+
016 point 2 for why both exist.
|
|
43
|
+
- **`mcp.tool.pricing_status` span + metric attribute.** One of `"known"`
|
|
44
|
+
| `"unknown"` | `"user_override"`, set whenever token usage was
|
|
45
|
+
extracted at all — even with no model detected (`"unknown"` in that
|
|
46
|
+
case), unlike the existing model/cost attributes. Added to the
|
|
47
|
+
`mcp.tool.tokens.total` / `mcp.tool.cost.total` metrics too, so a
|
|
48
|
+
dashboard can compute "% of tokens/spend unpriced" as a direct
|
|
49
|
+
aggregation instead of inferring it from missing data. Provenance-based:
|
|
50
|
+
`"user_override"` means the model's key came from your
|
|
51
|
+
`pricing`/`pricingTable`, regardless of whether the numbers you supplied
|
|
52
|
+
happen to match `DEFAULT_PRICING`'s own entry.
|
|
53
|
+
- **Staleness signalling.** `DEFAULT_PRICING_LAST_VERIFIED` (also
|
|
54
|
+
exported) names the table's last-checked date; once it's more than 90
|
|
55
|
+
days old, `instrumentMcpServer()` fires a one-time `diag.warn()`, and,
|
|
56
|
+
when `setupNodeSdk: true`, also attaches an
|
|
57
|
+
`mcp.pricing.default_table_last_verified` resource attribute. Both only
|
|
58
|
+
fire when `DEFAULT_PRICING` is actually contributing to the effective
|
|
59
|
+
table — a caller who fully replaced it via `pricingTable` isn't using
|
|
60
|
+
our defaults, so a warning about them would be misleading.
|
|
61
|
+
`isDefaultPricingStale(now?, thresholdDays?)` (also exported) is the
|
|
62
|
+
pure function behind the warning, for callers who want to check it
|
|
63
|
+
themselves.
|
|
64
|
+
- **Bedrock region caveat, documented not modeled.** Bedrock pricing
|
|
65
|
+
varies by region; `DEFAULT_PRICING`'s Bedrock entries (Nova, and the new
|
|
66
|
+
Titan embedding entry) assume us-east-1 list price and the table is not
|
|
67
|
+
region-keyed — no tool-result usage shape this package recognizes
|
|
68
|
+
carries a region signal to key a lookup on. Documented loudly in the
|
|
69
|
+
README and in `pricing.js`; override via `costTracking.pricing` for a
|
|
70
|
+
different region. See ADR 016 point 5.
|
|
71
|
+
- `calculateCost()` gained defensive validation for malformed pricing
|
|
72
|
+
entries (missing/negative/non-numeric `inputPer1M`/`outputPer1M`),
|
|
73
|
+
since `pricing`/`pricingTable` now make it reachable with
|
|
74
|
+
caller-supplied shapes it previously never had to distrust — degrades
|
|
75
|
+
to `null`, same as an unknown model, never throws.
|
|
76
|
+
|
|
77
|
+
### Added — W3C Trace Context propagation over MCP `_meta` (Phase 2 of 2, server-side only)
|
|
78
|
+
|
|
79
|
+
The most-complained-about gap in agent observability: an agent's own
|
|
80
|
+
trace (LangGraph or otherwise) and the MCP server's trace for the tool
|
|
81
|
+
call it made were always two disconnected traces, with no edge between
|
|
82
|
+
them. Full design, including the sampling and conflicting-context
|
|
83
|
+
decisions below: ADR 017 (`docs/adr/017-trace-context-propagation.md`).
|
|
84
|
+
|
|
85
|
+
- **`tools/call` requests carrying a valid W3C `traceparent` in
|
|
86
|
+
`params._meta` now become a child of the calling agent's own span**,
|
|
87
|
+
joining what were two disconnected traces into one — under both MCP v1
|
|
88
|
+
and v2, with zero configuration and no new option. `tracestate` is
|
|
89
|
+
propagated too, when present. Works for any client already emitting
|
|
90
|
+
`traceparent` via a standard OTel SDK's `propagation.inject()` in any
|
|
91
|
+
language — this isn't Node/JS-specific on the client side, only on
|
|
92
|
+
which side of the wire this release implements.
|
|
93
|
+
- **The upstream sampling decision is honored automatically, by
|
|
94
|
+
construction, with no sampling logic written for this feature**: the
|
|
95
|
+
extracted `SpanContext` is marked `isRemote: true` with the real parsed
|
|
96
|
+
`traceFlags`, which is exactly what the SDK's own default
|
|
97
|
+
`ParentBasedSampler` already keys its remote-parent decision off of. A
|
|
98
|
+
not-sampled upstream `traceparent` means this tool-call span is not
|
|
99
|
+
recorded or exported, matching the calling agent's own choice — see the
|
|
100
|
+
ADR's "Sampling" section for why forcing sampling regardless was
|
|
101
|
+
considered and rejected.
|
|
102
|
+
- **A `_meta`-extracted context always replaces, never merges with, an
|
|
103
|
+
already-active local context** (e.g. an ambient HTTP-server span from
|
|
104
|
+
auto-instrumentation on a Streamable HTTP transport) — the message-level
|
|
105
|
+
`_meta` context is the semantically correct parent for one tool call,
|
|
106
|
+
full stop, regardless of what transport-level span it happened to
|
|
107
|
+
arrive inside. See the ADR's "Conflicting `_meta.traceparent`" section.
|
|
108
|
+
- **Absent, malformed, or unparseable `_meta`/`traceparent` produces
|
|
109
|
+
behavior that is byte-identical to pre-v0.11.0** — not merely
|
|
110
|
+
equivalent to it: confirmed by reading both `NoopTracer` and the real
|
|
111
|
+
SDK `Tracer`'s own `startActiveSpan()` fallback (`ctx ?? context.active()`),
|
|
112
|
+
which is exactly what this feature's `extractTraceContext()` returns
|
|
113
|
+
for every case that isn't a valid `traceparent`. No `diag.warn()` for
|
|
114
|
+
the common "client doesn't send `_meta.traceparent`" case — see the
|
|
115
|
+
ADR's "no warn spam" constraint.
|
|
116
|
+
- **Zero new dependencies.** `@opentelemetry/core`'s
|
|
117
|
+
`W3CTraceContextPropagator` was the obvious reference implementation
|
|
118
|
+
and was deliberately not taken as a dependency, per
|
|
119
|
+
`CONTRIBUTING.md`'s "no new dependencies without discussion first" —
|
|
120
|
+
everything needed except the traceparent regex itself (~10 lines,
|
|
121
|
+
matching `@opentelemetry/core`'s own validation field-for-field) was
|
|
122
|
+
already available from `@opentelemetry/api`, already a peer dependency
|
|
123
|
+
— including `createTraceState()`, a fully spec-validated `tracestate`
|
|
124
|
+
parser. Full reasoning: ADR 017's "No new dependency" section.
|
|
125
|
+
- **Server-side extraction only.** The client-side shim that would let a
|
|
126
|
+
Node/Python agent framework *set* `_meta.traceparent` on outgoing calls
|
|
127
|
+
is explicitly out of scope for this phase — extraction is independently
|
|
128
|
+
useful today, for free, to any client whose own tooling already sets
|
|
129
|
+
`_meta` in this shape. Tracked as future work, not implied as solved.
|
|
130
|
+
|
|
131
|
+
Also fixed in this release: two places (`index.d.ts`'s
|
|
132
|
+
`instrumentMcpServer()` docblock, and this file's own v0.10.0 entry
|
|
133
|
+
below) still described `docs/known-gaps.md` entries 6/7/8 using language
|
|
134
|
+
that read as still-open, or as scoped out of v0.10.0 — both were stale.
|
|
135
|
+
Entries 7 and 8 have been fully fixed since v0.10.0 with no open caveats;
|
|
136
|
+
entry 6's fallback-session-id half is fixed too, narrowed to a smaller,
|
|
137
|
+
genuinely-still-open remainder (see the corrected v0.10.0 entry below and
|
|
138
|
+
`index.d.ts`'s updated docblock for the accurate, current accounting).
|
|
139
|
+
|
|
3
140
|
## 0.10.0
|
|
4
141
|
|
|
5
142
|
**⚠️ Behavior change, unrelated to the feature below — read this first.**
|
|
@@ -60,15 +197,32 @@ new was added for this, since ADR 012's original design already covers
|
|
|
60
197
|
this exact deployment shape, v2 just makes it the default instead of an
|
|
61
198
|
edge case.
|
|
62
199
|
|
|
63
|
-
**
|
|
64
|
-
|
|
65
|
-
|
|
66
|
-
|
|
67
|
-
|
|
68
|
-
|
|
69
|
-
|
|
70
|
-
|
|
71
|
-
|
|
200
|
+
**Correction (recorded here rather than silently edited): both gaps below
|
|
201
|
+
were actually closed in this same v0.10.0 release, not left open.** The
|
|
202
|
+
paragraph originally here said entries 6 and 8 (`docs/known-gaps.md`)
|
|
203
|
+
were scoped out of this round — true of the round that produced the text
|
|
204
|
+
above, not of what actually shipped. A follow-up investigation, completed
|
|
205
|
+
before v0.10.0 was cut, folded both fixes back in: `isSingleConnectionTransport()`
|
|
206
|
+
no longer misclassifies the transport `createMcpHandler` builds internally
|
|
207
|
+
(entry 8 — fully fixed: v2 now requires positive confirmation,
|
|
208
|
+
`transport.constructor.name === 'StdioServerTransport'`, instead of
|
|
209
|
+
inferring single-connection from an absent `sessionId` property), and
|
|
210
|
+
Agent Thrash Detection's fallback session id is now registry-backed via
|
|
211
|
+
`instanceKey` (entry 6's fallback-id gap — fixed: repeated
|
|
212
|
+
`instrumentMcpServer()` calls sharing an `instanceKey` now reuse the same
|
|
213
|
+
generated id instead of a fresh one per call). Both fixes shipped in the
|
|
214
|
+
same commit, in a specific order — fixing detection (entry 8) before
|
|
215
|
+
sharing the fallback id (entry 6) — since sharing it first would have made
|
|
216
|
+
entry 8's false positive worse, not better.
|
|
217
|
+
|
|
218
|
+
**What remains genuinely open, narrower than either original gap:**
|
|
219
|
+
`thrashSessionState` (whether a server has ever proven itself
|
|
220
|
+
session-aware) still isn't registry-backed, and — structurally, not a
|
|
221
|
+
bug this library can fix — MCP spec 2026-07-28 removes protocol-level
|
|
222
|
+
sessions entirely, so no configuration of this library can produce a
|
|
223
|
+
*real* session id for a spec-2026-07-28-native deployment in the first
|
|
224
|
+
place. See ADR 015's final "Update ... Findings 3 and 8 landed here too"
|
|
225
|
+
section and `docs/known-gaps.md` entries 6 and 8 for the full accounting.
|
|
72
226
|
|
|
73
227
|
## 0.9.0
|
|
74
228
|
|
package/README.md
CHANGED
|
@@ -495,19 +495,20 @@ point, or JSON-in-text inside `content[0].text`) automatically gets
|
|
|
495
495
|
`mcp.tool.tokens.*` / `mcp.tool.model` / `mcp.tool.cost.*` span
|
|
496
496
|
attributes, priced against `DEFAULT_PRICING`.
|
|
497
497
|
|
|
498
|
-
### Advanced:
|
|
498
|
+
### Advanced: overriding pricing, a custom extractor, and a budget guardrail
|
|
499
499
|
|
|
500
500
|
```js
|
|
501
|
-
import { instrumentMcpServer
|
|
501
|
+
import { instrumentMcpServer } from 'opentel-mcp';
|
|
502
502
|
|
|
503
503
|
instrumentMcpServer(server, {
|
|
504
504
|
serviceName: 'my-mcp-server',
|
|
505
505
|
costTracking: {
|
|
506
|
-
//
|
|
507
|
-
//
|
|
508
|
-
|
|
509
|
-
|
|
510
|
-
|
|
506
|
+
// Merged per-model OVER DEFAULT_PRICING — correct a stale price or add
|
|
507
|
+
// a model DEFAULT_PRICING doesn't know about, without having to spread
|
|
508
|
+
// the whole default table yourself. Everything you don't name here is
|
|
509
|
+
// untouched. See "Overriding pricing" below.
|
|
510
|
+
pricing: {
|
|
511
|
+
'my-internal-model': { pricingKind: 'chat', inputPer1M: 1.0, outputPer1M: 2.0, currency: 'USD' },
|
|
511
512
|
},
|
|
512
513
|
// Recognize your own tool result shape. Return null for anything you
|
|
513
514
|
// don't recognize — never throw (see src/cost/extractor.js).
|
|
@@ -526,6 +527,41 @@ instrumentMcpServer(server, {
|
|
|
526
527
|
});
|
|
527
528
|
```
|
|
528
529
|
|
|
530
|
+
### Overriding pricing (v0.11.0+)
|
|
531
|
+
|
|
532
|
+
Two ways to change what `costTracking` prices against, composing as
|
|
533
|
+
`{ ...(pricingTable ?? DEFAULT_PRICING), ...pricing }` — see ADR 016
|
|
534
|
+
(`docs/adr/016-pricing-override-and-staleness.md`) for the full reasoning:
|
|
535
|
+
|
|
536
|
+
- **`costTracking.pricing`** — a *partial* table, merged **per-model**
|
|
537
|
+
over `DEFAULT_PRICING` (or over `pricingTable`, if you set both). Each
|
|
538
|
+
key you supply replaces that one model's entire pricing entry; every
|
|
539
|
+
model you don't name keeps its default price. This is the recommended
|
|
540
|
+
way to correct a stale number or add a model — enterprises on
|
|
541
|
+
committed-use discounts, AWS EDP, Bedrock provisioned throughput, or
|
|
542
|
+
Azure OpenAI negotiated rates should use this to reflect what they
|
|
543
|
+
actually pay, not list price.
|
|
544
|
+
- **`costTracking.pricingTable`** — *fully replaces* `DEFAULT_PRICING`.
|
|
545
|
+
Use this when you want an effective table containing **only** your own
|
|
546
|
+
models, none of `DEFAULT_PRICING`'s.
|
|
547
|
+
|
|
548
|
+
A model priced via either option reports `pricing_status: 'user_override'`
|
|
549
|
+
(see the span attributes table below) instead of `'known'`.
|
|
550
|
+
|
|
551
|
+
### Embedding models (v0.11.0+)
|
|
552
|
+
|
|
553
|
+
Embeddings are input-token-only — there's no "output" to price. Rather
|
|
554
|
+
than modeling that as `outputPer1M: 0` (indistinguishable from a
|
|
555
|
+
data-entry bug), `ModelPricing` carries an explicit
|
|
556
|
+
`pricingKind: 'chat' | 'embedding'` discriminator; an `'embedding'` entry
|
|
557
|
+
has no `outputPer1M` field at all, and `calculateCost()` never reads
|
|
558
|
+
`outputTokens` for one. `DEFAULT_PRICING` includes OpenAI
|
|
559
|
+
`text-embedding-3-small`/`text-embedding-3-large`/`text-embedding-ada-002`,
|
|
560
|
+
Cohere `cohere-embed-v3`, and Bedrock `amazon-titan-embed-v2` out of the
|
|
561
|
+
box. A pre-v0.11.0 custom pricing entry with no `pricingKind` field is
|
|
562
|
+
still treated as `'chat'` at runtime — only the TypeScript type is
|
|
563
|
+
stricter, not the runtime.
|
|
564
|
+
|
|
529
565
|
### Span attributes
|
|
530
566
|
|
|
531
567
|
| Attribute | Standard OTel? | Description | Example |
|
|
@@ -535,19 +571,23 @@ instrumentMcpServer(server, {
|
|
|
535
571
|
| `mcp.tool.tokens.total` | Custom | input + output | 1500 |
|
|
536
572
|
| `mcp.tool.model` | Custom | Detected model name | "claude-sonnet-5" |
|
|
537
573
|
| `gen_ai.response.model` | Standard (GenAI semconv)[^5] | Same value as `mcp.tool.model`, co-emitted for dashboard compatibility | "claude-sonnet-5" |
|
|
574
|
+
| `mcp.tool.pricing_status` | Custom[^7] | `"known"` \| `"unknown"` \| `"user_override"` — set whenever token usage was extracted, even with no model detected | "known" |
|
|
538
575
|
| `mcp.tool.cost.usd` | Custom | Estimated cost, from `calculateCost()` | 0.0105 |
|
|
539
576
|
| `mcp.tool.cost.currency` | Custom | Always `"USD"` today | "USD" |
|
|
540
577
|
| `mcp.tool.cost.budget_exceeded` | Custom | `true` once a configured `costTracking.budget` limit is crossed | true |
|
|
541
578
|
| `mcp.tool.cost.budget_scope` | Custom | Which budget scope tripped: `"session"` \| `"tool"` (session wins if both did) | "session" |
|
|
542
579
|
|
|
543
580
|
[^5]: `gen_ai.response.model` is a real OTel GenAI semantic convention attribute ("the name of the model that generated the response") — but this span is an MCP tool-call span (`gen_ai.operation.name: execute_tool`), not a dedicated LLM request/response span, so co-emitting it here is a **pragmatic dashboard-compatibility choice, not a spec-pure emission**. It's set purely so off-the-shelf GenAI dashboards (Grafana, SigNoz, Honeycomb) that filter/group by `gen_ai.response.model` pick these spans up without any opentel-mcp-specific configuration. Full reasoning in `src/attributes.js`'s `ATTR_GEN_AI_RESPONSE_MODEL` docblock.
|
|
581
|
+
[^7]: `mcp.tool.pricing_status` is provenance-based, not value-based: `"user_override"` means the model's key was present in your `costTracking.pricing`/`pricingTable`, whether or not the numbers you supplied happen to match `DEFAULT_PRICING`. `"unknown"` covers both "no model detected" and "model detected but not priceable" (unrecognized, or a malformed override entry) — see ADR 016 point 4.
|
|
544
582
|
|
|
545
|
-
The four token/model attributes are set together or not at all;
|
|
546
|
-
|
|
547
|
-
|
|
548
|
-
|
|
549
|
-
|
|
550
|
-
|
|
583
|
+
The four token/model attributes are set together or not at all;
|
|
584
|
+
`mcp.tool.pricing_status` is set whenever usage was extracted at all
|
|
585
|
+
(unlike the model/cost attributes, it's present even with no model
|
|
586
|
+
detected); the two cost attributes only appear when a model was detected
|
|
587
|
+
*and* it resolves in the effective pricing table; the two budget
|
|
588
|
+
attributes only appear when a cost was calculated *and* a configured
|
|
589
|
+
limit was crossed. Source of truth: `src/attributes.js` and
|
|
590
|
+
`src/instrument.js`'s `applyCostAttribution()`.
|
|
551
591
|
|
|
552
592
|
### Metrics
|
|
553
593
|
|
|
@@ -557,21 +597,38 @@ registered, `enableMetrics: false` opts out of these too.
|
|
|
557
597
|
|
|
558
598
|
| Metric | Type | Unit | Attributes | Emitted when |
|
|
559
599
|
|---|---|---|---|---|
|
|
560
|
-
| `mcp.tool.tokens.total` | Counter | tokens | `gen_ai.tool.name`, `mcp.tool.model`[^6] | Usage detected in the tool result |
|
|
561
|
-
| `mcp.tool.cost.total` | Counter | USD | `gen_ai.tool.name`, `mcp.tool.model`[^6] | Cost calculated (model resolved in
|
|
562
|
-
|
|
563
|
-
[^6]: `mcp.tool.model` is only added when a model was detected — the same optional-attribute cardinality pattern `mcp.failure.category` already uses on the other four metrics.
|
|
564
|
-
|
|
565
|
-
|
|
566
|
-
|
|
567
|
-
|
|
568
|
-
|
|
569
|
-
|
|
570
|
-
|
|
571
|
-
|
|
572
|
-
|
|
573
|
-
|
|
574
|
-
|
|
600
|
+
| `mcp.tool.tokens.total` | Counter | tokens | `gen_ai.tool.name`, `mcp.tool.model`[^6], `mcp.tool.pricing_status` | Usage detected in the tool result |
|
|
601
|
+
| `mcp.tool.cost.total` | Counter | USD | `gen_ai.tool.name`, `mcp.tool.model`[^6], `mcp.tool.pricing_status` | Cost calculated (model resolved in the effective pricing table) |
|
|
602
|
+
|
|
603
|
+
[^6]: `mcp.tool.model` is only added when a model was detected — the same optional-attribute cardinality pattern `mcp.failure.category` already uses on the other four metrics. `mcp.tool.pricing_status` is always added — a fixed, closed 3-value enum, well within this package's metric-label cardinality discipline (see `COST_METRIC_SAFE_ATTRIBUTES` in `src/attributes.js`).
|
|
604
|
+
|
|
605
|
+
Grouping `mcp.tool.tokens.total` by `mcp.tool.pricing_status` answers "what
|
|
606
|
+
fraction of tokens/spend is running through models we can't price" as a
|
|
607
|
+
direct query, instead of inferring it from missing `mcp.tool.cost.*` data.
|
|
608
|
+
|
|
609
|
+
### Pricing accuracy and staleness
|
|
610
|
+
|
|
611
|
+
> **`DEFAULT_PRICING` is a best-effort snapshot, not a maintained price
|
|
612
|
+
> list.** Provider pricing changes frequently and varies by region/contract
|
|
613
|
+
> — `opentel-mcp` does not guarantee it stays current, and says so at
|
|
614
|
+
> runtime, not just here: once `DEFAULT_PRICING` (checked via
|
|
615
|
+
> `DEFAULT_PRICING_LAST_VERIFIED`, also exported) is more than 90 days
|
|
616
|
+
> past its last-verified date, `instrumentMcpServer()` fires a one-time
|
|
617
|
+
> `diag.warn()` naming that date — and, when `setupNodeSdk: true`, also
|
|
618
|
+
> attaches an `mcp.pricing.default_table_last_verified` resource
|
|
619
|
+
> attribute, so long-running deployments can alert on it directly. Both
|
|
620
|
+
> only fire when `DEFAULT_PRICING` is actually contributing to your
|
|
621
|
+
> effective table (i.e. you haven't fully replaced it via `pricingTable`)
|
|
622
|
+
> — see ADR 016 point 3.
|
|
623
|
+
|
|
624
|
+
`DEFAULT_PRICING` (`src/cost/pricing.js`) covers 20+ models — chat and
|
|
625
|
+
embedding — across six providers: Anthropic, OpenAI, Google, Cohere, AWS
|
|
626
|
+
Bedrock, and DeepSeek. **Bedrock entries (Nova and Titan embeddings)
|
|
627
|
+
assume us-east-1 list pricing** — Bedrock pricing varies by region and
|
|
628
|
+
this table is not region-keyed (no reliable region signal exists in any
|
|
629
|
+
tool-result usage shape this package recognizes to key a lookup on — see
|
|
630
|
+
ADR 016 point 5); override via `costTracking.pricing` for a different
|
|
631
|
+
region.
|
|
575
632
|
|
|
576
633
|
### Extending it
|
|
577
634
|
|
|
@@ -581,7 +638,12 @@ convenience default, not a maintained price list.
|
|
|
581
638
|
TokenUsage | null`, never throwing) to recognize anything else.
|
|
582
639
|
- `calculateCost(inputTokens, outputTokens, model, pricingTable)` is also
|
|
583
640
|
exported directly, for recomputing cost outside the instrumentation
|
|
584
|
-
hot path (e.g. over historical spans).
|
|
641
|
+
hot path (e.g. over historical spans). Malformed pricing entries
|
|
642
|
+
(missing/negative/non-numeric `inputPer1M`/`outputPer1M`) degrade to
|
|
643
|
+
`null`, same as an unknown model — never throws.
|
|
644
|
+
- `isDefaultPricingStale(now?, thresholdDays?)` (also exported) is the
|
|
645
|
+
pure function backing the staleness warning above, if you want to check
|
|
646
|
+
it yourself (e.g. in a startup health check).
|
|
585
647
|
- Disable everything in this section with `costTracking: { enabled:
|
|
586
648
|
false }`; tracing, metrics, and fingerprinting are all unaffected.
|
|
587
649
|
- Budget tracking (`costTracking.budget`) is in-memory and scoped to one
|
|
@@ -1525,6 +1587,84 @@ your schema-drift or two-axis observation signals matter for retention
|
|
|
1525
1587
|
too, add `mcp.tool.schema_drift.detected` (boolean_attribute) as another
|
|
1526
1588
|
OR'd policy the same way.
|
|
1527
1589
|
|
|
1590
|
+
## Trace Context Propagation (v0.11.0+)
|
|
1591
|
+
|
|
1592
|
+
The most-complained-about gap in agent observability: an agent framework
|
|
1593
|
+
(LangGraph or otherwise) calls a tool on your MCP server, and you get two
|
|
1594
|
+
disconnected traces — the agent's own trace dies at the tool boundary,
|
|
1595
|
+
and the server's trace for what the tool actually did starts fresh, with
|
|
1596
|
+
no edge between them. Debugging a slow or failing agent run means
|
|
1597
|
+
manually correlating timestamps across two separate traces (or two
|
|
1598
|
+
separate services in the same backend) instead of looking at one.
|
|
1599
|
+
|
|
1600
|
+
opentel-mcp closes the server-side half of this: when a `tools/call`
|
|
1601
|
+
request carries a valid [W3C `traceparent`](https://www.w3.org/TR/trace-context/)
|
|
1602
|
+
in `params._meta`, the tool-call span becomes a **child of the calling
|
|
1603
|
+
agent's own span** — the same trace, not two. `tracestate` is propagated
|
|
1604
|
+
too, when present.
|
|
1605
|
+
|
|
1606
|
+
### Zero-config — nothing to opt into
|
|
1607
|
+
|
|
1608
|
+
There's no `traceContext: { enabled: false }` option, deliberately: this
|
|
1609
|
+
is unconditional, because there's nothing for an operator to want to turn
|
|
1610
|
+
off. Any client that already emits `traceparent` via a standard OTel
|
|
1611
|
+
SDK's `propagation.inject()` — in any language, this isn't Node-specific
|
|
1612
|
+
on the client side — gets linked traces the moment it copies that string
|
|
1613
|
+
into `_meta.traceparent` on its outgoing `tools/call` request:
|
|
1614
|
+
|
|
1615
|
+
```json
|
|
1616
|
+
{
|
|
1617
|
+
"method": "tools/call",
|
|
1618
|
+
"params": {
|
|
1619
|
+
"name": "search_docs",
|
|
1620
|
+
"arguments": { "query": "..." },
|
|
1621
|
+
"_meta": {
|
|
1622
|
+
"traceparent": "00-4bf92f3577b34da6a3ce929d0e0e4736-00f067aa0ba902b7-01",
|
|
1623
|
+
"tracestate": "vendorname=opaquevalue"
|
|
1624
|
+
}
|
|
1625
|
+
}
|
|
1626
|
+
}
|
|
1627
|
+
```
|
|
1628
|
+
|
|
1629
|
+
A client that doesn't set `_meta.traceparent` (today's overwhelming
|
|
1630
|
+
majority, until client-side shims exist — see "What's not built yet"
|
|
1631
|
+
below) sees **zero change of any kind** — confirmed byte-identical to
|
|
1632
|
+
pre-v0.11.0 behavior, not merely "should behave the same."
|
|
1633
|
+
|
|
1634
|
+
### The upstream sampling decision is honored automatically
|
|
1635
|
+
|
|
1636
|
+
If the calling agent's own trace wasn't sampled (the `traceparent`'s
|
|
1637
|
+
flags byte has the sampled bit unset), this tool-call span isn't recorded
|
|
1638
|
+
or exported either — matching the agent's own choice, via the OTel SDK's
|
|
1639
|
+
own default `ParentBasedSampler`, which already inspects a remote
|
|
1640
|
+
parent's `traceFlags` for exactly this. opentel-mcp writes no sampling
|
|
1641
|
+
logic of its own for this — see ADR 017's "Sampling" section
|
|
1642
|
+
(`docs/adr/017-trace-context-propagation.md`) for why forcing sampling
|
|
1643
|
+
regardless was considered and rejected.
|
|
1644
|
+
|
|
1645
|
+
### Conflicting or pre-existing context
|
|
1646
|
+
|
|
1647
|
+
If your transport's own auto-instrumentation (e.g.
|
|
1648
|
+
`@opentelemetry/instrumentation-http` on a Streamable HTTP server) has
|
|
1649
|
+
already put some other span active by the time this package's handler
|
|
1650
|
+
runs, a valid `_meta.traceparent` **replaces it outright** — it is never
|
|
1651
|
+
merged. The `_meta`-carried context is message-scoped (this one logical
|
|
1652
|
+
agent operation); the transport-level span is connection/request-scoped
|
|
1653
|
+
and isn't the right parent for it. Full reasoning: ADR 017's
|
|
1654
|
+
"Conflicting `_meta.traceparent`" section.
|
|
1655
|
+
|
|
1656
|
+
### What's not built yet
|
|
1657
|
+
|
|
1658
|
+
**Server-side extraction only.** There is no client-side shim in this
|
|
1659
|
+
package (yet) for a Node or Python agent framework to *set*
|
|
1660
|
+
`_meta.traceparent` on its own outgoing calls — you need a client that
|
|
1661
|
+
already does this itself (or does it via your own glue code:
|
|
1662
|
+
`propagation.inject(context.active(), meta, ...)` from your OTel SDK,
|
|
1663
|
+
writing into `_meta` before the call). Extraction is independently
|
|
1664
|
+
useful today, for free, to any client that already sets `_meta` in this
|
|
1665
|
+
shape; the client-side half is tracked as future work, not implied as
|
|
1666
|
+
solved by this release.
|
|
1667
|
+
|
|
1528
1668
|
## Configuration
|
|
1529
1669
|
|
|
1530
1670
|
All options passed to `instrumentMcpServer(server, options)`. Source of
|
package/package.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "opentel-mcp",
|
|
3
|
-
"version": "0.
|
|
3
|
+
"version": "0.11.0",
|
|
4
4
|
"description": "One-line OpenTelemetry instrumentation for Model Context Protocol (MCP) servers",
|
|
5
5
|
"type": "module",
|
|
6
6
|
"main": "src/index.js",
|
|
@@ -74,6 +74,7 @@
|
|
|
74
74
|
"devDependencies": {
|
|
75
75
|
"@modelcontextprotocol/server": "^2.0.0",
|
|
76
76
|
"@opentelemetry/api": "^1.9.0",
|
|
77
|
+
"@opentelemetry/context-async-hooks": "^2.9.0",
|
|
77
78
|
"@opentelemetry/sdk-metrics": "^2.9.0",
|
|
78
79
|
"@opentelemetry/sdk-trace-base": "^2.9.0",
|
|
79
80
|
"@vitest/coverage-v8": "^2.1.9",
|
package/src/attributes.js
CHANGED
|
@@ -168,3 +168,48 @@ export const MCP_TOOL_COST_BUDGET_SCOPE_SESSION = 'session';
|
|
|
168
168
|
|
|
169
169
|
/** Well-known mcp.tool.cost.budget_scope value: costTracking.budget.perToolUsd was exceeded. */
|
|
170
170
|
export const MCP_TOOL_COST_BUDGET_SCOPE_TOOL = 'tool';
|
|
171
|
+
|
|
172
|
+
// --- Pricing provenance (v0.11.0, non-spec, ADR 016) ---
|
|
173
|
+
//
|
|
174
|
+
// docs/adr/016-pricing-override-and-staleness.md. Lets a dashboard answer
|
|
175
|
+
// "what fraction of spend/tokens is unpriced" as a direct query instead of
|
|
176
|
+
// an inference from missing mcp.tool.cost.* attributes.
|
|
177
|
+
|
|
178
|
+
/**
|
|
179
|
+
* Whether the model detected for this tool call resolved to a known price.
|
|
180
|
+
* Set whenever token usage was extracted at all (same gating as
|
|
181
|
+
* ATTR_MCP_TOOL_TOKENS_INPUT/OUTPUT/TOTAL) — present even when no model was
|
|
182
|
+
* detected, unlike ATTR_MCP_TOOL_MODEL/ATTR_MCP_TOOL_COST_USD. One of
|
|
183
|
+
* MCP_TOOL_PRICING_STATUS_KNOWN / _UNKNOWN / _USER_OVERRIDE.
|
|
184
|
+
*/
|
|
185
|
+
export const ATTR_MCP_TOOL_PRICING_STATUS = 'mcp.tool.pricing_status';
|
|
186
|
+
|
|
187
|
+
/** Well-known mcp.tool.pricing_status value: no model was detected, or the detected model has no pricing entry. */
|
|
188
|
+
export const MCP_TOOL_PRICING_STATUS_UNKNOWN = 'unknown';
|
|
189
|
+
|
|
190
|
+
/** Well-known mcp.tool.pricing_status value: priced against an unmodified DEFAULT_PRICING entry. */
|
|
191
|
+
export const MCP_TOOL_PRICING_STATUS_KNOWN = 'known';
|
|
192
|
+
|
|
193
|
+
/** Well-known mcp.tool.pricing_status value: priced against a caller-supplied costTracking.pricing/pricingTable entry. */
|
|
194
|
+
export const MCP_TOOL_PRICING_STATUS_USER_OVERRIDE = 'user_override';
|
|
195
|
+
|
|
196
|
+
/**
|
|
197
|
+
* Attribute keys safe to attach to mcp.tool.tokens.total / mcp.tool.cost.total
|
|
198
|
+
* metric labels, for the cost/pricing domain specifically. Mirrors the
|
|
199
|
+
* governance pattern fingerprint/attributes.js's METRIC_SAFE_ATTRIBUTES and
|
|
200
|
+
* schema-drift/attributes.js's own same-named export already established
|
|
201
|
+
* for their domains — this domain gets its own list rather than borrowing
|
|
202
|
+
* either of those (mixing an unrelated domain's cardinality reasoning in
|
|
203
|
+
* here would obscure which ADR covers which attribute). Not re-exported
|
|
204
|
+
* from index.js, same as schema-drift's — internal governance, not public
|
|
205
|
+
* API. mcp.tool.model remains governed the way it already was, an inline
|
|
206
|
+
* cardinality comment in metrics.js, not a list entry here — see ADR 016
|
|
207
|
+
* point 4 for why that attribute's boundedness argument doesn't fit a
|
|
208
|
+
* fixed-enum list cleanly the way this one does.
|
|
209
|
+
*
|
|
210
|
+
* @type {readonly string[]}
|
|
211
|
+
*/
|
|
212
|
+
export const COST_METRIC_SAFE_ATTRIBUTES = Object.freeze([ATTR_MCP_TOOL_PRICING_STATUS]);
|
|
213
|
+
|
|
214
|
+
/** Resource attribute (setupNodeSdk: true only) naming DEFAULT_PRICING's lastVerified date. See ADR 016 point 3. */
|
|
215
|
+
export const ATTR_MCP_PRICING_DEFAULT_TABLE_LAST_VERIFIED = 'mcp.pricing.default_table_last_verified';
|
package/src/config.js
CHANGED
|
@@ -4,8 +4,9 @@
|
|
|
4
4
|
*/
|
|
5
5
|
|
|
6
6
|
import { diag } from '@opentelemetry/api';
|
|
7
|
-
import { DEFAULT_PRICING } from './cost/pricing.js';
|
|
7
|
+
import { DEFAULT_PRICING, DEFAULT_PRICING_LAST_VERIFIED, isDefaultPricingStale } from './cost/pricing.js';
|
|
8
8
|
import { defaultExtractor } from './cost/extractor.js';
|
|
9
|
+
import { normalizeModelName } from './cost/calculator.js';
|
|
9
10
|
import { resolveThrashConfig } from './thrash/config.js';
|
|
10
11
|
import { resolveSchemaDriftConfig } from './schema-drift/config.js';
|
|
11
12
|
|
|
@@ -13,9 +14,16 @@ import { resolveSchemaDriftConfig } from './schema-drift/config.js';
|
|
|
13
14
|
* @typedef {object} CostTrackingOptions
|
|
14
15
|
* @property {boolean} [enabled=true] - Set to false to disable cost/token span attributes and the
|
|
15
16
|
* mcp.tool.tokens.total / mcp.tool.cost.total metrics entirely.
|
|
16
|
-
* @property {import('./cost/pricing.js').PricingTable} [pricingTable] -
|
|
17
|
-
*
|
|
18
|
-
*
|
|
17
|
+
* @property {import('./cost/pricing.js').PricingTable} [pricingTable] - Fully replaces DEFAULT_PRICING (or,
|
|
18
|
+
* if `pricing` below is also set, replaces the base table `pricing` is merged over). Supply your own table
|
|
19
|
+
* when you want an effective table containing ONLY your own models. For correcting/adding a few models
|
|
20
|
+
* while keeping the rest of DEFAULT_PRICING, prefer `pricing` instead — see ADR 016
|
|
21
|
+
* (docs/adr/016-pricing-override-and-staleness.md) point 2.
|
|
22
|
+
* @property {Partial<import('./cost/pricing.js').PricingTable>} [pricing] - Partial pricing table, merged
|
|
23
|
+
* per-model OVER `pricingTable ?? DEFAULT_PRICING` — each key you supply replaces that model's entire
|
|
24
|
+
* ModelPricing entry; every model you don't name is untouched. The recommended way to correct stale
|
|
25
|
+
* pricing or add a model DEFAULT_PRICING doesn't know about. A model priced via this option (or via
|
|
26
|
+
* `pricingTable`) reports `pricing_status: 'user_override'` — see ADR 016 point 2.
|
|
19
27
|
* @property {import('./cost/extractor.js').UsageExtractor} [extractor] - Overrides defaultExtractor
|
|
20
28
|
* (src/cost/extractor.js). Supply your own to recognize a tool result shape defaultExtractor doesn't.
|
|
21
29
|
* @property {import('./cost/budget.js').BudgetConfig} [budget] - Per-session and per-tool cumulative-cost
|
|
@@ -113,6 +121,24 @@ export function __resetServiceNameWarnedForTests() {
|
|
|
113
121
|
warnedServiceNameIgnored = false;
|
|
114
122
|
}
|
|
115
123
|
|
|
124
|
+
// Guards the DEFAULT_PRICING staleness diagnostic below (ADR 016 point 3),
|
|
125
|
+
// same one-per-process pattern as warnedServiceNameIgnored above.
|
|
126
|
+
let warnedPricingStale = false;
|
|
127
|
+
|
|
128
|
+
// Test-only, same purpose as __resetServiceNameWarnedForTests above. Not
|
|
129
|
+
// part of the public API.
|
|
130
|
+
export function __resetPricingStaleWarnedForTests() {
|
|
131
|
+
warnedPricingStale = false;
|
|
132
|
+
}
|
|
133
|
+
|
|
134
|
+
/**
|
|
135
|
+
* @param {unknown} value
|
|
136
|
+
* @returns {value is Record<string, unknown>}
|
|
137
|
+
*/
|
|
138
|
+
function isPlainObject(value) {
|
|
139
|
+
return typeof value === 'object' && value !== null && !Array.isArray(value);
|
|
140
|
+
}
|
|
141
|
+
|
|
116
142
|
// ADR 012, Phase 2: the first env var this codebase reads for a bare
|
|
117
143
|
// top-level InstrumentOptions field, not one nested inside a feature's own
|
|
118
144
|
// sub-config (contrast OTEL_MCP_THRASH_*/OTEL_MCP_SCHEMA_DRIFT_*, both
|
|
@@ -171,6 +197,50 @@ export function resolveOptions(options) {
|
|
|
171
197
|
}
|
|
172
198
|
|
|
173
199
|
const rawCostTracking = opts.costTracking ?? {};
|
|
200
|
+
const costTrackingEnabled = rawCostTracking.enabled ?? true;
|
|
201
|
+
|
|
202
|
+
// ADR 016 point 2: pricingTable (when supplied) is the BASE table,
|
|
203
|
+
// fully replacing DEFAULT_PRICING; pricing (when supplied) is then
|
|
204
|
+
// merged per-model OVER that base — a plain object spread is exactly
|
|
205
|
+
// "per-model, not top-level replacement" here, since each spread key is
|
|
206
|
+
// one model. Non-object pricingTable/pricing values (a caller's typo,
|
|
207
|
+
// e.g. a string or null) are silently ignored rather than spread, same
|
|
208
|
+
// "malformed config degrades, never crashes" discipline calculateCost()
|
|
209
|
+
// itself follows for a malformed individual entry.
|
|
210
|
+
const hasCustomPricingTable = isPlainObject(rawCostTracking.pricingTable);
|
|
211
|
+
const hasPricingOverride = isPlainObject(rawCostTracking.pricing);
|
|
212
|
+
const basePricingTable = hasCustomPricingTable ? rawCostTracking.pricingTable : DEFAULT_PRICING;
|
|
213
|
+
const pricingTable = hasPricingOverride ? { ...basePricingTable, ...rawCostTracking.pricing } : basePricingTable;
|
|
214
|
+
|
|
215
|
+
// ADR 016 point 4: which normalized model keys came from the caller's
|
|
216
|
+
// own config surface (pricingTable and/or pricing), for the
|
|
217
|
+
// mcp.tool.pricing_status 'known' vs 'user_override' distinction —
|
|
218
|
+
// computed once here, not per call. Provenance-based: a key counts as
|
|
219
|
+
// an override because the caller named it, regardless of whether the
|
|
220
|
+
// value they supplied happens to match DEFAULT_PRICING's own entry.
|
|
221
|
+
const pricingOverrideKeys = new Set();
|
|
222
|
+
if (hasCustomPricingTable) {
|
|
223
|
+
for (const key of Object.keys(rawCostTracking.pricingTable)) pricingOverrideKeys.add(normalizeModelName(key));
|
|
224
|
+
}
|
|
225
|
+
if (hasPricingOverride) {
|
|
226
|
+
for (const key of Object.keys(rawCostTracking.pricing)) pricingOverrideKeys.add(normalizeModelName(key));
|
|
227
|
+
}
|
|
228
|
+
|
|
229
|
+
// ADR 016 point 3: only warn when DEFAULT_PRICING is actually
|
|
230
|
+
// contributing to the effective table — a caller who fully replaced it
|
|
231
|
+
// via pricingTable isn't using our defaults at all, so a staleness
|
|
232
|
+
// warning about them would be misleading. Gated on costTrackingEnabled
|
|
233
|
+
// too: no cost tracking happens at all otherwise, so DEFAULT_PRICING's
|
|
234
|
+
// age is moot.
|
|
235
|
+
if (costTrackingEnabled && !hasCustomPricingTable && !warnedPricingStale && isDefaultPricingStale()) {
|
|
236
|
+
warnedPricingStale = true;
|
|
237
|
+
diag.warn(
|
|
238
|
+
`opentel-mcp: DEFAULT_PRICING was last verified ${DEFAULT_PRICING_LAST_VERIFIED}, more than 90 days ago. ` +
|
|
239
|
+
'Provider list pricing may have changed since. Override costTracking.pricing (merged per-model over ' +
|
|
240
|
+
'DEFAULT_PRICING) for models whose pricing you need to keep current — see ADR 016, ' +
|
|
241
|
+
'docs/adr/016-pricing-override-and-staleness.md.',
|
|
242
|
+
);
|
|
243
|
+
}
|
|
174
244
|
|
|
175
245
|
return {
|
|
176
246
|
serviceName: opts.serviceName,
|
|
@@ -180,8 +250,10 @@ export function resolveOptions(options) {
|
|
|
180
250
|
setupNodeSdk,
|
|
181
251
|
fingerprinting: opts.fingerprinting ?? true,
|
|
182
252
|
costTracking: {
|
|
183
|
-
enabled:
|
|
184
|
-
pricingTable
|
|
253
|
+
enabled: costTrackingEnabled,
|
|
254
|
+
pricingTable,
|
|
255
|
+
pricingOverrideKeys,
|
|
256
|
+
usingDefaultPricing: !hasCustomPricingTable,
|
|
185
257
|
extractor: rawCostTracking.extractor ?? defaultExtractor,
|
|
186
258
|
budget: rawCostTracking.budget,
|
|
187
259
|
},
|