@glassflow-ai/rius 0.8.0 → 1.0.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +181 -7
- package/dist/index.cjs +2566 -483
- package/dist/index.d.cts +284 -14
- package/dist/index.d.ts +284 -14
- package/dist/index.js +2543 -449
- package/package.json +3 -3
package/README.md
CHANGED
|
@@ -98,6 +98,38 @@ const handleQuery2 = observe(
|
|
|
98
98
|
await handleQuery("hello");
|
|
99
99
|
```
|
|
100
100
|
|
|
101
|
+
That is the CHAIN case, the default. On every other kind the span is named
|
|
102
|
+
the way the GenAI conventions say to, `{operation} {target}`, and the
|
|
103
|
+
wrapped function's name is used as the target only where it IS the target —
|
|
104
|
+
a tool:
|
|
105
|
+
|
|
106
|
+
```ts
|
|
107
|
+
const getWeather = observe(async function getWeather(city: string) { ... }, {
|
|
108
|
+
kind: SpanKind.TOOL,
|
|
109
|
+
});
|
|
110
|
+
await getWeather("Berlin"); // span "execute_tool getWeather", gen_ai.tool.name "getWeather"
|
|
111
|
+
```
|
|
112
|
+
|
|
113
|
+
An AGENT wrapper is named after the agent it invokes (`agentName`, or the
|
|
114
|
+
one `init()` was given), a RETRIEVER after its `dataSourceId`. A function
|
|
115
|
+
name is not an agent or an index, so it is never used as one. Pass `{ name }`
|
|
116
|
+
to override any of this.
|
|
117
|
+
|
|
118
|
+
An AGENT span (from `observe`, `startSpan` or `startAsCurrentSpan`) can also
|
|
119
|
+
carry `agentId` and `agentVersion`, which describe the agent it *invoked*, as
|
|
120
|
+
`gen_ai.agent.id` and `gen_ai.agent.version`. Both are taken verbatim — the
|
|
121
|
+
conventions' own version examples are `1.0.0` and `2025-05-01` — and both are
|
|
122
|
+
ignored on every other kind.
|
|
123
|
+
|
|
124
|
+
An AGENT wrapper also scopes the calls it makes: TOOL spans opened while it
|
|
125
|
+
runs, local or MCP, carry that agent as `gen_ai.agent.name` — on an
|
|
126
|
+
execute-tool span the conventions define that key as the agent *executing*
|
|
127
|
+
the tool, where on an invoke-agent span it is the agent being *invoked*.
|
|
128
|
+
Outside any agent, a tool span falls back to the agent `init()` was given.
|
|
129
|
+
The scope follows async calls, so any nesting depth works; only
|
|
130
|
+
`startAsCurrentSpan` and `observe` open it, since `startSpan` does not make
|
|
131
|
+
its span current.
|
|
132
|
+
|
|
101
133
|
### `startSpan` / `startAsCurrentSpan`: manual spans
|
|
102
134
|
|
|
103
135
|
For finer control than `observe`, create spans directly:
|
|
@@ -118,9 +150,49 @@ automatically and recording any thrown exception. The options argument is
|
|
|
118
150
|
optional in both, so `startAsCurrentSpan("step", async (span) => { ... })`
|
|
119
151
|
works when you have nothing to configure.
|
|
120
152
|
|
|
153
|
+
**The name is optional too.** Omit it and the span is named the way the
|
|
154
|
+
OpenTelemetry GenAI conventions say to — `{operation} {target}` — composed
|
|
155
|
+
from the attributes the span already carries. The callback still comes last,
|
|
156
|
+
so `startAsCurrentSpan(options, fn)` and `startAsCurrentSpan(fn)` are both
|
|
157
|
+
valid:
|
|
158
|
+
|
|
159
|
+
| Kind | Default span name | Without the target |
|
|
160
|
+
| --- | --- | --- |
|
|
161
|
+
| LLM | `chat gpt-4o` (`{operation} {gen_ai.request.model}`) | `chat` |
|
|
162
|
+
| EMBEDDING | `embeddings text-embedding-3-small` | `embeddings` |
|
|
163
|
+
| TOOL | `execute_tool get_weather` | `execute_tool` |
|
|
164
|
+
| AGENT | `invoke_agent planner` | `invoke_agent` |
|
|
165
|
+
| RETRIEVER | `retrieval docs-index` | `retrieval` |
|
|
166
|
+
| CHAIN | the function name for `observe`, else `chain` | — |
|
|
167
|
+
|
|
168
|
+
```ts
|
|
169
|
+
// Named "execute_tool get_weather", with gen_ai.tool.name "get_weather".
|
|
170
|
+
await startAsCurrentSpan(
|
|
171
|
+
{ kind: SpanKind.TOOL, toolName: "get_weather", input: city },
|
|
172
|
+
async (span) => { ... },
|
|
173
|
+
);
|
|
174
|
+
```
|
|
175
|
+
|
|
176
|
+
A name you pass always wins; this changes the default, not your ability to
|
|
177
|
+
choose. The LLM name uses the REQUEST model, never the response model: the
|
|
178
|
+
span is named when it starts, and a partial-span snapshot has to carry the
|
|
179
|
+
same name as the span that replaces it. CHAIN has no operation to compose
|
|
180
|
+
from, so an unnamed CHAIN span is called `chain`.
|
|
181
|
+
|
|
121
182
|
`kind` sets the span's taxonomy (`openinference.span.kind`, our `SpanKind`
|
|
122
183
|
enum: CHAIN by default, or TOOL, RETRIEVER, EMBEDDING, AGENT, LLM). A TOOL
|
|
123
|
-
span also carries `gen_ai.tool.name`, set to the span name
|
|
184
|
+
span also carries `gen_ai.tool.name`, set to the span name unless you pass
|
|
185
|
+
`toolName` (and unset if you pass neither), which both the span helpers and `observe` accept for the case
|
|
186
|
+
where the span name is not the bare tool name. It also takes `toolCallId`
|
|
187
|
+
(`gen_ai.tool.call.id`, the id of the model tool-call this execution answers)
|
|
188
|
+
and `toolType` (`gen_ai.tool.type`: `function`, `extension`, `datastore`).
|
|
189
|
+
Both are yours to pass or omit, never derived, and neither changes the span
|
|
190
|
+
name. The MCP wrapper sets neither on purpose: a tool call id belongs to the
|
|
191
|
+
model's message, not to the protocol exchange, and the wrapper never sees it.
|
|
192
|
+
A RETRIEVER span carries
|
|
193
|
+
`gen_ai.data_source.id` when you pass `dataSourceId`, naming the index or
|
|
194
|
+
collection it searched, and `gen_ai.retrieval.top_k` when you pass `topK`.
|
|
195
|
+
Each kind also
|
|
124
196
|
decides the OpenTelemetry `SpanKind` field the conventions expect: LLM,
|
|
125
197
|
EMBEDDING and RETRIEVER spans are CLIENT, everything else INTERNAL. Pass
|
|
126
198
|
`otelKind` to override it, for example CLIENT for a call to a hosted agent.
|
|
@@ -184,9 +256,27 @@ await startAsCurrentGeneration(
|
|
|
184
256
|
);
|
|
185
257
|
```
|
|
186
258
|
|
|
187
|
-
|
|
188
|
-
|
|
189
|
-
|
|
259
|
+
`modelParameters` are recorded when the span opens, so they also show up on
|
|
260
|
+
the live (pending) view of a call that is still running. A parameter the
|
|
261
|
+
GenAI conventions define is recorded under its canonical `gen_ai.request.*`
|
|
262
|
+
key, and the provider spellings for it are recognised too:
|
|
263
|
+
`max_completion_tokens`, `maxOutputTokens` and `maxTokens` all become
|
|
264
|
+
`gen_ai.request.max_tokens`, `stop` becomes `gen_ai.request.stop_sequences`,
|
|
265
|
+
and `n` becomes `gen_ai.request.choice.count`. Any other parameter is recorded
|
|
266
|
+
under `rius.request.<key>` with its key unchanged. Tool definitions passed as a
|
|
267
|
+
parameter (`tools`, `functions`) are treated as content and are dropped under
|
|
268
|
+
`captureContent: false`. `setFinishReasons` accepts one reason or a list.
|
|
269
|
+
|
|
270
|
+
Two more conventions keys sit on either side of the call. `outputType`
|
|
271
|
+
(`gen_ai.output.type`: `text`, `json`, `image` or `speech`) is an option,
|
|
272
|
+
because the output modality is something the request asks for and so is known
|
|
273
|
+
before the span starts. `setResponseId` (`gen_ai.response.id`) is a setter
|
|
274
|
+
alongside `setModel`, because the provider's id for the completion only exists
|
|
275
|
+
once the completion does.
|
|
276
|
+
|
|
277
|
+
The name is optional here as well: `startAsCurrentGeneration({ model:
|
|
278
|
+
"gpt-4o" }, fn)` names the span `chat gpt-4o`, and an `operation` override
|
|
279
|
+
names it after that operation instead (`embeddings text-embedding-3-small`).
|
|
190
280
|
|
|
191
281
|
## Auto-instrumentation
|
|
192
282
|
|
|
@@ -203,7 +293,7 @@ Install only the ones you use:
|
|
|
203
293
|
- **`openai`** wraps the OpenAI SDK, via `@arizeai/openinference-instrumentation-openai`, turning provider calls into spans.
|
|
204
294
|
- **`anthropic`** wraps the Anthropic SDK, via `@arizeai/openinference-instrumentation-anthropic`, turning Messages calls into LLM spans with model, messages and token counts.
|
|
205
295
|
- **`langchain`** traces LangChain.js chains, models, tools and retrievers, via `@arizeai/openinference-instrumentation-langchain`. It patches the callback manager in `@langchain/core`, which every LangChain.js application already has, and the SDK resolves that module for you so nothing has to be passed in at `init()`.
|
|
206
|
-
- **`vercel-ai`** attaches to spans the Vercel AI SDK's own OpenTelemetry integration produces, via `@arizeai/openinference-vercel`, adding OpenInference attributes to them. This package requires Node 22 or newer.
|
|
296
|
+
- **`vercel-ai`** attaches to spans the Vercel AI SDK's own OpenTelemetry integration produces, via `@arizeai/openinference-vercel`, adding OpenInference attributes to them. It touches only spans in the AI SDK's `ai.*` dialect (v5/v6, recognised by their `ai.operationId`): your own spans, other instrumentations' spans and the GenAI-native spans of AI SDK v7's `@ai-sdk/otel` integration are exported exactly as they would be without it. This package requires Node 22 or newer.
|
|
207
297
|
- **`mcp`** patches `@modelcontextprotocol/sdk`'s `Client.callTool` so every MCP tool call becomes a TOOL span named `execute_tool <tool>`, carrying the tool name, arguments, result, latency, and error status. The span follows the OpenTelemetry MCP conventions: OTel `SpanKind` CLIENT, `mcp.method.name` set to `tools/call`, `mcp.protocol.version` set to the negotiated version when the transport exposes it (Streamable HTTP does; stdio and SSE do not), and `error.type` set to `tool_error` when the server returns an `isError` result.
|
|
208
298
|
|
|
209
299
|
There is no selection option: `init()` always attempts every bundled
|
|
@@ -252,6 +342,7 @@ variables, then the defaults below.
|
|
|
252
342
|
| `endpoint` | `RIUS_ENDPOINT` | `https://ingest.eu.console.rius-glassflow.com` |
|
|
253
343
|
| `apiKey` | `RIUS_API_KEY` | none |
|
|
254
344
|
| `serviceName` | `RIUS_SERVICE_NAME` | `unknown_service` |
|
|
345
|
+
| `serviceVersion` | `RIUS_SERVICE_VERSION` | none (also read from `service.version` in `OTEL_RESOURCE_ATTRIBUTES`; left off the resource when unset, never defaulted) |
|
|
255
346
|
| `disabled` | `RIUS_DISABLED` | `false` |
|
|
256
347
|
| `sampleRate` | `RIUS_SAMPLE_RATE` | `1.0` |
|
|
257
348
|
| `captureContent` | `RIUS_CAPTURE_CONTENT` | `true` |
|
|
@@ -259,12 +350,52 @@ variables, then the defaults below.
|
|
|
259
350
|
| `heartbeat` | `RIUS_HEARTBEAT` | `true` |
|
|
260
351
|
| `heartbeatInterval` | `RIUS_HEARTBEAT_INTERVAL` | `15` (seconds; clamped 5-300) |
|
|
261
352
|
| `agentName` | `RIUS_AGENT_NAME` | `serviceName` |
|
|
353
|
+
| `mainAgentId` | `RIUS_MAIN_AGENT_ID` | none (a STABLE id for the agent this process is; never a transient in-memory id) |
|
|
354
|
+
| `mainAgentDescription` | `RIUS_MAIN_AGENT_DESCRIPTION` | none |
|
|
355
|
+
| `mainAgentVersion` | `RIUS_MAIN_AGENT_VERSION` | none (the version of the AGENT DEFINITION, never derived from `serviceVersion`) |
|
|
262
356
|
| `partialSpans` | `RIUS_PARTIAL_SPANS` | `false` |
|
|
263
357
|
| `partialSpansDelay` | `RIUS_PARTIAL_SPANS_DELAY` | `0` (seconds; clamped 0-60) |
|
|
264
358
|
| `sessionId` | `RIUS_SESSION_ID` | none (process-wide session default; `withSession` overrides it) |
|
|
359
|
+
| `bridgeForeignProvider` | `RIUS_BRIDGE_FOREIGN_PROVIDER` | `false` (also export the spans of another OpenTelemetry SDK that owns the global provider; see below) |
|
|
265
360
|
|
|
266
361
|
Traces are posted to `<endpoint>/v1/traces`.
|
|
267
362
|
|
|
363
|
+
`serviceVersion` is stamped on the resource as `service.version`, so the
|
|
364
|
+
console can tell one build of an agent from another. It has no default on
|
|
365
|
+
purpose: an invented version would put every unversioned process under the
|
|
366
|
+
same wrong number, the way `unknown_service` does for unnamed ones, so the
|
|
367
|
+
key is simply absent until you supply one. This SDK builds its resource
|
|
368
|
+
explicitly, which in OpenTelemetry for JavaScript means the OTel environment
|
|
369
|
+
is not merged into it, so `OTEL_RESOURCE_ATTRIBUTES=service.version=...` is
|
|
370
|
+
read here by hand as the last tier after the option and `RIUS_SERVICE_VERSION`.
|
|
371
|
+
|
|
372
|
+
The agent this process *is* is described on the resource under
|
|
373
|
+
`rius.main_agent.*`: `rius.main_agent.name` (the resolved `agentName`, absent
|
|
374
|
+
when nothing was named and it is still the `unknown_service` placeholder),
|
|
375
|
+
plus the optional `rius.main_agent.id`, `.description` and `.version`.
|
|
376
|
+
`gen_ai.agent.name` is still stamped alongside them, unchanged: the new keys
|
|
377
|
+
are additive so that no deployment loses its agent identity between an SDK
|
|
378
|
+
upgrade and a backend one. The vendor namespace exists because
|
|
379
|
+
`gen_ai.agent.name` has no resource-level meaning in the conventions and on a
|
|
380
|
+
span means the agent being *invoked*, so one key was answering two questions.
|
|
381
|
+
|
|
382
|
+
`mainAgentId` must be a **stable** identifier — a registry id, a hosted
|
|
383
|
+
agent's ARN. A uuid minted at startup or another process-local id identifies
|
|
384
|
+
a *run*, not an agent, and would split one agent into a fresh bucket per
|
|
385
|
+
restart; `service.instance.id` already answers "which process".
|
|
386
|
+
|
|
387
|
+
Three versions travel separately and none is ever derived from another:
|
|
388
|
+
|
|
389
|
+
| Key | Answers |
|
|
390
|
+
| -------------------------- | --------------------------------------------------- |
|
|
391
|
+
| `service.version` | which build of the deployment is running |
|
|
392
|
+
| `rius.main_agent.version` | which version of the agent *definition* it runs — its prompt, tools and policy |
|
|
393
|
+
| `gen_ai.agent.version` | which version of the agent a given AGENT span *invoked* |
|
|
394
|
+
|
|
395
|
+
A service can sit at `2.3.1` while the agent definition it runs is at `7` and
|
|
396
|
+
the agent it calls out to is at `1.0.0`; a prompt change moves the second
|
|
397
|
+
without touching the first.
|
|
398
|
+
|
|
268
399
|
`captureContent` controls whether input/output content (prompts,
|
|
269
400
|
completions, tool arguments) is attached to spans. It defaults to
|
|
270
401
|
`true`. Set it to `false`, or supply a `mask` function, if your spans
|
|
@@ -289,6 +420,27 @@ optional integration is loaded, so no third-party module is patched in your
|
|
|
289
420
|
process. `client.ready` resolves to an empty list. Spans can still be
|
|
290
421
|
created and are simply dropped.
|
|
291
422
|
|
|
423
|
+
### Running next to another OpenTelemetry SDK
|
|
424
|
+
|
|
425
|
+
OpenTelemetry has one global tracer provider per process, and the first
|
|
426
|
+
registration wins. If another SDK (Langfuse v3, for one) claims it before
|
|
427
|
+
`init()`, your own job and tool spans go to that SDK only, while the LLM spans
|
|
428
|
+
inside them still reach Rius — so Rius receives traces whose parents it never
|
|
429
|
+
gets. Rius detects this whatever you configure: those LLM spans carry
|
|
430
|
+
`rius.parent.foreign=true`, the resource records
|
|
431
|
+
`rius.sdk.global_provider=foreign:<class>`, and the first one logs a warning.
|
|
432
|
+
|
|
433
|
+
To send the other SDK's spans to Rius as well, opt in with
|
|
434
|
+
`init({ bridgeForeignProvider: true })` or `RIUS_BRIDGE_FOREIGN_PROVIDER=true`.
|
|
435
|
+
Rius then attaches its own pipeline to that provider and exports its spans too,
|
|
436
|
+
subject to Rius's `sampleRate`, `captureContent` and `mask`, under the Rius
|
|
437
|
+
resource identity. It never modifies what the other SDK exports: that vendor
|
|
438
|
+
receives exactly what it would without Rius.
|
|
439
|
+
|
|
440
|
+
```typescript
|
|
441
|
+
init({ bridgeForeignProvider: true }) // export the other SDK's spans as well
|
|
442
|
+
```
|
|
443
|
+
|
|
292
444
|
## Agent liveness
|
|
293
445
|
|
|
294
446
|
Traces only leave the process when a span ends, so an idle-but-alive agent
|
|
@@ -315,8 +467,9 @@ graceful exit if you want that distinction to show up.
|
|
|
315
467
|
|
|
316
468
|
The heartbeat never keeps the process alive, its timer is unref'd, and never
|
|
317
469
|
throws into your code; delivery failures are logged once per client and
|
|
318
|
-
silent after that. `agentName` identifies the agent in heartbeat payloads
|
|
319
|
-
|
|
470
|
+
silent after that. `agentName` identifies the agent in heartbeat payloads and
|
|
471
|
+
on every span, where it is stamped on the resource as `gen_ai.agent.name`. It
|
|
472
|
+
defaults to `serviceName`, so the two agree unless you set it.
|
|
320
473
|
|
|
321
474
|
### Partial spans
|
|
322
475
|
|
|
@@ -368,6 +521,27 @@ Batching is tunable via the standard OpenTelemetry env vars:
|
|
|
368
521
|
`OTEL_BSP_SCHEDULE_DELAY` (default 5000 ms), `OTEL_BSP_MAX_EXPORT_BATCH_SIZE`
|
|
369
522
|
(default 512), and `OTEL_BSP_EXPORT_TIMEOUT` (default 30000 ms).
|
|
370
523
|
|
|
524
|
+
### Span attribute limit
|
|
525
|
+
|
|
526
|
+
OpenTelemetry caps a span at 128 attributes by default, and a span that
|
|
527
|
+
reaches the cap silently refuses every new key. OpenInference writes one
|
|
528
|
+
attribute per message field and per tool field, so an agent loop with ten
|
|
529
|
+
tools passes 128 after about nine turns, and the keys written last are
|
|
530
|
+
the ones that matter most: the token usage (so cost), the output messages,
|
|
531
|
+
`output.value` and the finish reason.
|
|
532
|
+
|
|
533
|
+
`init()` therefore raises the span attribute count limit to 4096. The
|
|
534
|
+
standard variables still win: if `OTEL_SPAN_ATTRIBUTE_COUNT_LIMIT` or
|
|
535
|
+
`OTEL_ATTRIBUTE_COUNT_LIMIT` is set to a number, `init()` leaves the limit
|
|
536
|
+
to OpenTelemetry, which applies that value. The attribute value-length
|
|
537
|
+
limit is not changed.
|
|
538
|
+
|
|
539
|
+
If you run the OpenInference instrumentations **without** this SDK, on a
|
|
540
|
+
tracer provider of your own, set the limit yourself, for example
|
|
541
|
+
`OTEL_SPAN_ATTRIBUTE_COUNT_LIMIT=4096`, or pass
|
|
542
|
+
`spanLimits: { attributeCountLimit: 4096 }` to your provider. Otherwise
|
|
543
|
+
long agent calls lose their usage and output.
|
|
544
|
+
|
|
371
545
|
## Development
|
|
372
546
|
|
|
373
547
|
```bash
|