@glassflow-ai/rius 0.8.0 → 1.0.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -98,6 +98,38 @@ const handleQuery2 = observe(
98
98
  await handleQuery("hello");
99
99
  ```
100
100
 
101
+ That is the CHAIN case, the default. On every other kind the span is named
102
+ the way the GenAI conventions say to, `{operation} {target}`, and the
103
+ wrapped function's name is used as the target only where it IS the target —
104
+ a tool:
105
+
106
+ ```ts
107
+ const getWeather = observe(async function getWeather(city: string) { ... }, {
108
+ kind: SpanKind.TOOL,
109
+ });
110
+ await getWeather("Berlin"); // span "execute_tool getWeather", gen_ai.tool.name "getWeather"
111
+ ```
112
+
113
+ An AGENT wrapper is named after the agent it invokes (`agentName`, or the
114
+ one `init()` was given), a RETRIEVER after its `dataSourceId`. A function
115
+ name is not an agent or an index, so it is never used as one. Pass `{ name }`
116
+ to override any of this.
117
+
118
+ An AGENT span (from `observe`, `startSpan` or `startAsCurrentSpan`) can also
119
+ carry `agentId` and `agentVersion`, which describe the agent it *invoked*, as
120
+ `gen_ai.agent.id` and `gen_ai.agent.version`. Both are taken verbatim — the
121
+ conventions' own version examples are `1.0.0` and `2025-05-01` — and both are
122
+ ignored on every other kind.
123
+
124
+ An AGENT wrapper also scopes the calls it makes: TOOL spans opened while it
125
+ runs, local or MCP, carry that agent as `gen_ai.agent.name` — on an
126
+ execute-tool span the conventions define that key as the agent *executing*
127
+ the tool, where on an invoke-agent span it is the agent being *invoked*.
128
+ Outside any agent, a tool span falls back to the agent `init()` was given.
129
+ The scope follows async calls, so any nesting depth works; only
130
+ `startAsCurrentSpan` and `observe` open it, since `startSpan` does not make
131
+ its span current.
132
+
101
133
  ### `startSpan` / `startAsCurrentSpan`: manual spans
102
134
 
103
135
  For finer control than `observe`, create spans directly:
@@ -118,9 +150,49 @@ automatically and recording any thrown exception. The options argument is
118
150
  optional in both, so `startAsCurrentSpan("step", async (span) => { ... })`
119
151
  works when you have nothing to configure.
120
152
 
153
+ **The name is optional too.** Omit it and the span is named the way the
154
+ OpenTelemetry GenAI conventions say to — `{operation} {target}` — composed
155
+ from the attributes the span already carries. The callback still comes last,
156
+ so `startAsCurrentSpan(options, fn)` and `startAsCurrentSpan(fn)` are both
157
+ valid:
158
+
159
+ | Kind | Default span name | Without the target |
160
+ | --- | --- | --- |
161
+ | LLM | `chat gpt-4o` (`{operation} {gen_ai.request.model}`) | `chat` |
162
+ | EMBEDDING | `embeddings text-embedding-3-small` | `embeddings` |
163
+ | TOOL | `execute_tool get_weather` | `execute_tool` |
164
+ | AGENT | `invoke_agent planner` | `invoke_agent` |
165
+ | RETRIEVER | `retrieval docs-index` | `retrieval` |
166
+ | CHAIN | the function name for `observe`, else `chain` | — |
167
+
168
+ ```ts
169
+ // Named "execute_tool get_weather", with gen_ai.tool.name "get_weather".
170
+ await startAsCurrentSpan(
171
+ { kind: SpanKind.TOOL, toolName: "get_weather", input: city },
172
+ async (span) => { ... },
173
+ );
174
+ ```
175
+
176
+ A name you pass always wins; this changes the default, not your ability to
177
+ choose. The LLM name uses the REQUEST model, never the response model: the
178
+ span is named when it starts, and a partial-span snapshot has to carry the
179
+ same name as the span that replaces it. CHAIN has no operation to compose
180
+ from, so an unnamed CHAIN span is called `chain`.
181
+
121
182
  `kind` sets the span's taxonomy (`openinference.span.kind`, our `SpanKind`
122
183
  enum: CHAIN by default, or TOOL, RETRIEVER, EMBEDDING, AGENT, LLM). A TOOL
123
- span also carries `gen_ai.tool.name`, set to the span name. Each kind also
184
+ span also carries `gen_ai.tool.name`, set to the span name unless you pass
185
+ `toolName` (and unset if you pass neither), which both the span helpers and `observe` accept for the case
186
+ where the span name is not the bare tool name. It also takes `toolCallId`
187
+ (`gen_ai.tool.call.id`, the id of the model tool-call this execution answers)
188
+ and `toolType` (`gen_ai.tool.type`: `function`, `extension`, `datastore`).
189
+ Both are yours to pass or omit, never derived, and neither changes the span
190
+ name. The MCP wrapper sets neither on purpose: a tool call id belongs to the
191
+ model's message, not to the protocol exchange, and the wrapper never sees it.
192
+ A RETRIEVER span carries
193
+ `gen_ai.data_source.id` when you pass `dataSourceId`, naming the index or
194
+ collection it searched, and `gen_ai.retrieval.top_k` when you pass `topK`.
195
+ Each kind also
124
196
  decides the OpenTelemetry `SpanKind` field the conventions expect: LLM,
125
197
  EMBEDDING and RETRIEVER spans are CLIENT, everything else INTERNAL. Pass
126
198
  `otelKind` to override it, for example CLIENT for a call to a hosted agent.
@@ -184,9 +256,27 @@ await startAsCurrentGeneration(
184
256
  );
185
257
  ```
186
258
 
187
- Each `modelParameters` entry is recorded as `gen_ai.request.<key>`, so use the
188
- provider's own parameter names. `setFinishReasons` accepts one reason or a
189
- list.
259
+ `modelParameters` are recorded when the span opens, so they also show up on
260
+ the live (pending) view of a call that is still running. A parameter the
261
+ GenAI conventions define is recorded under its canonical `gen_ai.request.*`
262
+ key, and the provider spellings for it are recognised too:
263
+ `max_completion_tokens`, `maxOutputTokens` and `maxTokens` all become
264
+ `gen_ai.request.max_tokens`, `stop` becomes `gen_ai.request.stop_sequences`,
265
+ and `n` becomes `gen_ai.request.choice.count`. Any other parameter is recorded
266
+ under `rius.request.<key>` with its key unchanged. Tool definitions passed as a
267
+ parameter (`tools`, `functions`) are treated as content and are dropped under
268
+ `captureContent: false`. `setFinishReasons` accepts one reason or a list.
269
+
270
+ Two more conventions keys sit on either side of the call. `outputType`
271
+ (`gen_ai.output.type`: `text`, `json`, `image` or `speech`) is an option,
272
+ because the output modality is something the request asks for and so is known
273
+ before the span starts. `setResponseId` (`gen_ai.response.id`) is a setter
274
+ alongside `setModel`, because the provider's id for the completion only exists
275
+ once the completion does.
276
+
277
+ The name is optional here as well: `startAsCurrentGeneration({ model:
278
+ "gpt-4o" }, fn)` names the span `chat gpt-4o`, and an `operation` override
279
+ names it after that operation instead (`embeddings text-embedding-3-small`).
190
280
 
191
281
  ## Auto-instrumentation
192
282
 
@@ -203,7 +293,7 @@ Install only the ones you use:
203
293
  - **`openai`** wraps the OpenAI SDK, via `@arizeai/openinference-instrumentation-openai`, turning provider calls into spans.
204
294
  - **`anthropic`** wraps the Anthropic SDK, via `@arizeai/openinference-instrumentation-anthropic`, turning Messages calls into LLM spans with model, messages and token counts.
205
295
  - **`langchain`** traces LangChain.js chains, models, tools and retrievers, via `@arizeai/openinference-instrumentation-langchain`. It patches the callback manager in `@langchain/core`, which every LangChain.js application already has, and the SDK resolves that module for you so nothing has to be passed in at `init()`.
206
- - **`vercel-ai`** attaches to spans the Vercel AI SDK's own OpenTelemetry integration produces, via `@arizeai/openinference-vercel`, adding OpenInference attributes to them. This package requires Node 22 or newer.
296
+ - **`vercel-ai`** attaches to spans the Vercel AI SDK's own OpenTelemetry integration produces, via `@arizeai/openinference-vercel`, adding OpenInference attributes to them. It touches only spans in the AI SDK's `ai.*` dialect (v5/v6, recognised by their `ai.operationId`): your own spans, other instrumentations' spans and the GenAI-native spans of AI SDK v7's `@ai-sdk/otel` integration are exported exactly as they would be without it. This package requires Node 22 or newer.
207
297
  - **`mcp`** patches `@modelcontextprotocol/sdk`'s `Client.callTool` so every MCP tool call becomes a TOOL span named `execute_tool <tool>`, carrying the tool name, arguments, result, latency, and error status. The span follows the OpenTelemetry MCP conventions: OTel `SpanKind` CLIENT, `mcp.method.name` set to `tools/call`, `mcp.protocol.version` set to the negotiated version when the transport exposes it (Streamable HTTP does; stdio and SSE do not), and `error.type` set to `tool_error` when the server returns an `isError` result.
208
298
 
209
299
  There is no selection option: `init()` always attempts every bundled
@@ -252,6 +342,7 @@ variables, then the defaults below.
252
342
  | `endpoint` | `RIUS_ENDPOINT` | `https://ingest.eu.console.rius-glassflow.com` |
253
343
  | `apiKey` | `RIUS_API_KEY` | none |
254
344
  | `serviceName` | `RIUS_SERVICE_NAME` | `unknown_service` |
345
+ | `serviceVersion` | `RIUS_SERVICE_VERSION` | none (also read from `service.version` in `OTEL_RESOURCE_ATTRIBUTES`; left off the resource when unset, never defaulted) |
255
346
  | `disabled` | `RIUS_DISABLED` | `false` |
256
347
  | `sampleRate` | `RIUS_SAMPLE_RATE` | `1.0` |
257
348
  | `captureContent` | `RIUS_CAPTURE_CONTENT` | `true` |
@@ -259,12 +350,52 @@ variables, then the defaults below.
259
350
  | `heartbeat` | `RIUS_HEARTBEAT` | `true` |
260
351
  | `heartbeatInterval` | `RIUS_HEARTBEAT_INTERVAL` | `15` (seconds; clamped 5-300) |
261
352
  | `agentName` | `RIUS_AGENT_NAME` | `serviceName` |
353
+ | `mainAgentId` | `RIUS_MAIN_AGENT_ID` | none (a STABLE id for the agent this process is; never a transient in-memory id) |
354
+ | `mainAgentDescription` | `RIUS_MAIN_AGENT_DESCRIPTION` | none |
355
+ | `mainAgentVersion` | `RIUS_MAIN_AGENT_VERSION` | none (the version of the AGENT DEFINITION, never derived from `serviceVersion`) |
262
356
  | `partialSpans` | `RIUS_PARTIAL_SPANS` | `false` |
263
357
  | `partialSpansDelay` | `RIUS_PARTIAL_SPANS_DELAY` | `0` (seconds; clamped 0-60) |
264
358
  | `sessionId` | `RIUS_SESSION_ID` | none (process-wide session default; `withSession` overrides it) |
359
+ | `bridgeForeignProvider` | `RIUS_BRIDGE_FOREIGN_PROVIDER` | `false` (also export the spans of another OpenTelemetry SDK that owns the global provider; see below) |
265
360
 
266
361
  Traces are posted to `<endpoint>/v1/traces`.
267
362
 
363
+ `serviceVersion` is stamped on the resource as `service.version`, so the
364
+ console can tell one build of an agent from another. It has no default on
365
+ purpose: an invented version would put every unversioned process under the
366
+ same wrong number, the way `unknown_service` does for unnamed ones, so the
367
+ key is simply absent until you supply one. This SDK builds its resource
368
+ explicitly, which in OpenTelemetry for JavaScript means the OTel environment
369
+ is not merged into it, so `OTEL_RESOURCE_ATTRIBUTES=service.version=...` is
370
+ read here by hand as the last tier after the option and `RIUS_SERVICE_VERSION`.
371
+
372
+ The agent this process *is* is described on the resource under
373
+ `rius.main_agent.*`: `rius.main_agent.name` (the resolved `agentName`, absent
374
+ when nothing was named and it is still the `unknown_service` placeholder),
375
+ plus the optional `rius.main_agent.id`, `.description` and `.version`.
376
+ `gen_ai.agent.name` is still stamped alongside them, unchanged: the new keys
377
+ are additive so that no deployment loses its agent identity between an SDK
378
+ upgrade and a backend one. The vendor namespace exists because
379
+ `gen_ai.agent.name` has no resource-level meaning in the conventions and on a
380
+ span means the agent being *invoked*, so one key was answering two questions.
381
+
382
+ `mainAgentId` must be a **stable** identifier — a registry id, a hosted
383
+ agent's ARN. A uuid minted at startup or another process-local id identifies
384
+ a *run*, not an agent, and would split one agent into a fresh bucket per
385
+ restart; `service.instance.id` already answers "which process".
386
+
387
+ Three versions travel separately and none is ever derived from another:
388
+
389
+ | Key | Answers |
390
+ | -------------------------- | --------------------------------------------------- |
391
+ | `service.version` | which build of the deployment is running |
392
+ | `rius.main_agent.version` | which version of the agent *definition* it runs — its prompt, tools and policy |
393
+ | `gen_ai.agent.version` | which version of the agent a given AGENT span *invoked* |
394
+
395
+ A service can sit at `2.3.1` while the agent definition it runs is at `7` and
396
+ the agent it calls out to is at `1.0.0`; a prompt change moves the second
397
+ without touching the first.
398
+
268
399
  `captureContent` controls whether input/output content (prompts,
269
400
  completions, tool arguments) is attached to spans. It defaults to
270
401
  `true`. Set it to `false`, or supply a `mask` function, if your spans
@@ -289,6 +420,27 @@ optional integration is loaded, so no third-party module is patched in your
289
420
  process. `client.ready` resolves to an empty list. Spans can still be
290
421
  created and are simply dropped.
291
422
 
423
+ ### Running next to another OpenTelemetry SDK
424
+
425
+ OpenTelemetry has one global tracer provider per process, and the first
426
+ registration wins. If another SDK (Langfuse v3, for one) claims it before
427
+ `init()`, your own job and tool spans go to that SDK only, while the LLM spans
428
+ inside them still reach Rius — so Rius receives traces whose parents it never
429
+ gets. Rius detects this whatever you configure: those LLM spans carry
430
+ `rius.parent.foreign=true`, the resource records
431
+ `rius.sdk.global_provider=foreign:<class>`, and the first one logs a warning.
432
+
433
+ To send the other SDK's spans to Rius as well, opt in with
434
+ `init({ bridgeForeignProvider: true })` or `RIUS_BRIDGE_FOREIGN_PROVIDER=true`.
435
+ Rius then attaches its own pipeline to that provider and exports its spans too,
436
+ subject to Rius's `sampleRate`, `captureContent` and `mask`, under the Rius
437
+ resource identity. It never modifies what the other SDK exports: that vendor
438
+ receives exactly what it would without Rius.
439
+
440
+ ```typescript
441
+ init({ bridgeForeignProvider: true }) // export the other SDK's spans as well
442
+ ```
443
+
292
444
  ## Agent liveness
293
445
 
294
446
  Traces only leave the process when a span ends, so an idle-but-alive agent
@@ -315,8 +467,9 @@ graceful exit if you want that distinction to show up.
315
467
 
316
468
  The heartbeat never keeps the process alive, its timer is unref'd, and never
317
469
  throws into your code; delivery failures are logged once per client and
318
- silent after that. `agentName` identifies the agent in heartbeat payloads
319
- and defaults to `serviceName`.
470
+ silent after that. `agentName` identifies the agent in heartbeat payloads and
471
+ on every span, where it is stamped on the resource as `gen_ai.agent.name`. It
472
+ defaults to `serviceName`, so the two agree unless you set it.
320
473
 
321
474
  ### Partial spans
322
475
 
@@ -368,6 +521,27 @@ Batching is tunable via the standard OpenTelemetry env vars:
368
521
  `OTEL_BSP_SCHEDULE_DELAY` (default 5000 ms), `OTEL_BSP_MAX_EXPORT_BATCH_SIZE`
369
522
  (default 512), and `OTEL_BSP_EXPORT_TIMEOUT` (default 30000 ms).
370
523
 
524
+ ### Span attribute limit
525
+
526
+ OpenTelemetry caps a span at 128 attributes by default, and a span that
527
+ reaches the cap silently refuses every new key. OpenInference writes one
528
+ attribute per message field and per tool field, so an agent loop with ten
529
+ tools passes 128 after about nine turns, and the keys written last are
530
+ the ones that matter most: the token usage (so cost), the output messages,
531
+ `output.value` and the finish reason.
532
+
533
+ `init()` therefore raises the span attribute count limit to 4096. The
534
+ standard variables still win: if `OTEL_SPAN_ATTRIBUTE_COUNT_LIMIT` or
535
+ `OTEL_ATTRIBUTE_COUNT_LIMIT` is set to a number, `init()` leaves the limit
536
+ to OpenTelemetry, which applies that value. The attribute value-length
537
+ limit is not changed.
538
+
539
+ If you run the OpenInference instrumentations **without** this SDK, on a
540
+ tracer provider of your own, set the limit yourself, for example
541
+ `OTEL_SPAN_ATTRIBUTE_COUNT_LIMIT=4096`, or pass
542
+ `spanLimits: { attributeCountLimit: 4096 }` to your provider. Otherwise
543
+ long agent calls lose their usage and output.
544
+
371
545
  ## Development
372
546
 
373
547
  ```bash