@glassflow-ai/rius 0.7.0 → 1.0.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -98,6 +98,38 @@ const handleQuery2 = observe(
98
98
  await handleQuery("hello");
99
99
  ```
100
100
 
101
+ That is the CHAIN case, the default. On every other kind the span is named
102
+ the way the GenAI conventions say to, `{operation} {target}`, and the
103
+ wrapped function's name is used as the target only where it IS the target —
104
+ a tool:
105
+
106
+ ```ts
107
+ const getWeather = observe(async function getWeather(city: string) { ... }, {
108
+ kind: SpanKind.TOOL,
109
+ });
110
+ await getWeather("Berlin"); // span "execute_tool getWeather", gen_ai.tool.name "getWeather"
111
+ ```
112
+
113
+ An AGENT wrapper is named after the agent it invokes (`agentName`, or the
114
+ one `init()` was given), a RETRIEVER after its `dataSourceId`. A function
115
+ name is not an agent or an index, so it is never used as one. Pass `{ name }`
116
+ to override any of this.
117
+
118
+ An AGENT span (from `observe`, `startSpan` or `startAsCurrentSpan`) can also
119
+ carry `agentId` and `agentVersion`, which describe the agent it *invoked*, as
120
+ `gen_ai.agent.id` and `gen_ai.agent.version`. Both are taken verbatim — the
121
+ conventions' own version examples are `1.0.0` and `2025-05-01` — and both are
122
+ ignored on every other kind.
123
+
124
+ An AGENT wrapper also scopes the calls it makes: TOOL spans opened while it
125
+ runs, local or MCP, carry that agent as `gen_ai.agent.name` — on an
126
+ execute-tool span the conventions define that key as the agent *executing*
127
+ the tool, where on an invoke-agent span it is the agent being *invoked*.
128
+ Outside any agent, a tool span falls back to the agent `init()` was given.
129
+ The scope follows async calls, so any nesting depth works; only
130
+ `startAsCurrentSpan` and `observe` open it, since `startSpan` does not make
131
+ its span current.
132
+
101
133
  ### `startSpan` / `startAsCurrentSpan`: manual spans
102
134
 
103
135
  For finer control than `observe`, create spans directly:
@@ -118,8 +150,70 @@ automatically and recording any thrown exception. The options argument is
118
150
  optional in both, so `startAsCurrentSpan("step", async (span) => { ... })`
119
151
  works when you have nothing to configure.
120
152
 
153
+ **The name is optional too.** Omit it and the span is named the way the
154
+ OpenTelemetry GenAI conventions say to — `{operation} {target}` — composed
155
+ from the attributes the span already carries. The callback still comes last,
156
+ so `startAsCurrentSpan(options, fn)` and `startAsCurrentSpan(fn)` are both
157
+ valid:
158
+
159
+ | Kind | Default span name | Without the target |
160
+ | --- | --- | --- |
161
+ | LLM | `chat gpt-4o` (`{operation} {gen_ai.request.model}`) | `chat` |
162
+ | EMBEDDING | `embeddings text-embedding-3-small` | `embeddings` |
163
+ | TOOL | `execute_tool get_weather` | `execute_tool` |
164
+ | AGENT | `invoke_agent planner` | `invoke_agent` |
165
+ | RETRIEVER | `retrieval docs-index` | `retrieval` |
166
+ | CHAIN | the function name for `observe`, else `chain` | — |
167
+
168
+ ```ts
169
+ // Named "execute_tool get_weather", with gen_ai.tool.name "get_weather".
170
+ await startAsCurrentSpan(
171
+ { kind: SpanKind.TOOL, toolName: "get_weather", input: city },
172
+ async (span) => { ... },
173
+ );
174
+ ```
175
+
176
+ A name you pass always wins; this changes the default, not your ability to
177
+ choose. The LLM name uses the REQUEST model, never the response model: the
178
+ span is named when it starts, and a partial-span snapshot has to carry the
179
+ same name as the span that replaces it. CHAIN has no operation to compose
180
+ from, so an unnamed CHAIN span is called `chain`.
181
+
182
+ `kind` sets the span's taxonomy (`openinference.span.kind`, our `SpanKind`
183
+ enum: CHAIN by default, or TOOL, RETRIEVER, EMBEDDING, AGENT, LLM). A TOOL
184
+ span also carries `gen_ai.tool.name`, set to the span name unless you pass
185
+ `toolName` (and unset if you pass neither), which both the span helpers and `observe` accept for the case
186
+ where the span name is not the bare tool name. It also takes `toolCallId`
187
+ (`gen_ai.tool.call.id`, the id of the model tool-call this execution answers)
188
+ and `toolType` (`gen_ai.tool.type`: `function`, `extension`, `datastore`).
189
+ Both are yours to pass or omit, never derived, and neither changes the span
190
+ name. The MCP wrapper sets neither on purpose: a tool call id belongs to the
191
+ model's message, not to the protocol exchange, and the wrapper never sees it.
192
+ A RETRIEVER span carries
193
+ `gen_ai.data_source.id` when you pass `dataSourceId`, naming the index or
194
+ collection it searched, and `gen_ai.retrieval.top_k` when you pass `topK`.
195
+ Each kind also
196
+ decides the OpenTelemetry `SpanKind` field the conventions expect: LLM,
197
+ EMBEDDING and RETRIEVER spans are CLIENT, everything else INTERNAL. Pass
198
+ `otelKind` to override it, for example CLIENT for a call to a hosted agent.
199
+ Note the name collision: `SpanKind` exported by this package is the taxonomy
200
+ enum; the value for `otelKind` is `SpanKind` from `@opentelemetry/api`, so
201
+ import that one under an alias.
202
+
203
+ ```ts
204
+ import { SpanKind as OtelSpanKind } from "@opentelemetry/api";
205
+ import { SpanKind, startAsCurrentSpan } from "@glassflow-ai/rius";
206
+
207
+ await startAsCurrentSpan(
208
+ "delegate-to-planner",
209
+ { kind: SpanKind.AGENT, otelKind: OtelSpanKind.CLIENT, input: task },
210
+ async (span) => { ... },
211
+ );
212
+ ```
213
+
121
214
  On the manual path, `recordException()` does what the scoped form does for
122
- you: it records the error and sets ERROR status.
215
+ you: it records the error, sets ERROR status and sets `error.type` to the
216
+ error's name.
123
217
 
124
218
  ```ts
125
219
  const span = startSpan("fetch-documents");
@@ -162,9 +256,27 @@ await startAsCurrentGeneration(
162
256
  );
163
257
  ```
164
258
 
165
- Each `modelParameters` entry is recorded as `gen_ai.request.<key>`, so use the
166
- provider's own parameter names. `setFinishReasons` accepts one reason or a
167
- list.
259
+ `modelParameters` are recorded when the span opens, so they also show up on
260
+ the live (pending) view of a call that is still running. A parameter the
261
+ GenAI conventions define is recorded under its canonical `gen_ai.request.*`
262
+ key, and the provider spellings for it are recognised too:
263
+ `max_completion_tokens`, `maxOutputTokens` and `maxTokens` all become
264
+ `gen_ai.request.max_tokens`, `stop` becomes `gen_ai.request.stop_sequences`,
265
+ and `n` becomes `gen_ai.request.choice.count`. Any other parameter is recorded
266
+ under `rius.request.<key>` with its key unchanged. Tool definitions passed as a
267
+ parameter (`tools`, `functions`) are treated as content and are dropped under
268
+ `captureContent: false`. `setFinishReasons` accepts one reason or a list.
269
+
270
+ Two more conventions keys sit on either side of the call. `outputType`
271
+ (`gen_ai.output.type`: `text`, `json`, `image` or `speech`) is an option,
272
+ because the output modality is something the request asks for and so is known
273
+ before the span starts. `setResponseId` (`gen_ai.response.id`) is a setter
274
+ alongside `setModel`, because the provider's id for the completion only exists
275
+ once the completion does.
276
+
277
+ The name is optional here as well: `startAsCurrentGeneration({ model:
278
+ "gpt-4o" }, fn)` names the span `chat gpt-4o`, and an `operation` override
279
+ names it after that operation instead (`embeddings text-embedding-3-small`).
168
280
 
169
281
  ## Auto-instrumentation
170
282
 
@@ -181,8 +293,8 @@ Install only the ones you use:
181
293
  - **`openai`** wraps the OpenAI SDK, via `@arizeai/openinference-instrumentation-openai`, turning provider calls into spans.
182
294
  - **`anthropic`** wraps the Anthropic SDK, via `@arizeai/openinference-instrumentation-anthropic`, turning Messages calls into LLM spans with model, messages and token counts.
183
295
  - **`langchain`** traces LangChain.js chains, models, tools and retrievers, via `@arizeai/openinference-instrumentation-langchain`. It patches the callback manager in `@langchain/core`, which every LangChain.js application already has, and the SDK resolves that module for you so nothing has to be passed in at `init()`.
184
- - **`vercel-ai`** attaches to spans the Vercel AI SDK's own OpenTelemetry integration produces, via `@arizeai/openinference-vercel`, adding OpenInference attributes to them. This package requires Node 22 or newer.
185
- - **`mcp`** patches `@modelcontextprotocol/sdk`'s `Client.callTool` so every MCP tool call becomes a TOOL span, carrying the tool name, arguments, result, latency, and error status.
296
+ - **`vercel-ai`** attaches to spans the Vercel AI SDK's own OpenTelemetry integration produces, via `@arizeai/openinference-vercel`, adding OpenInference attributes to them. It touches only spans in the AI SDK's `ai.*` dialect (v5/v6, recognised by their `ai.operationId`): your own spans, other instrumentations' spans and the GenAI-native spans of AI SDK v7's `@ai-sdk/otel` integration are exported exactly as they would be without it. This package requires Node 22 or newer.
297
+ - **`mcp`** patches `@modelcontextprotocol/sdk`'s `Client.callTool` so every MCP tool call becomes a TOOL span named `execute_tool <tool>`, carrying the tool name, arguments, result, latency, and error status. The span follows the OpenTelemetry MCP conventions: OTel `SpanKind` CLIENT, `mcp.method.name` set to `tools/call`, `mcp.protocol.version` set to the negotiated version when the transport exposes it (Streamable HTTP does; stdio and SSE do not), and `error.type` set to `tool_error` when the server returns an `isError` result.
186
298
 
187
299
  There is no selection option: `init()` always attempts every bundled
188
300
  integration, so which ones actually attach is determined entirely by
@@ -230,6 +342,7 @@ variables, then the defaults below.
230
342
  | `endpoint` | `RIUS_ENDPOINT` | `https://ingest.eu.console.rius-glassflow.com` |
231
343
  | `apiKey` | `RIUS_API_KEY` | none |
232
344
  | `serviceName` | `RIUS_SERVICE_NAME` | `unknown_service` |
345
+ | `serviceVersion` | `RIUS_SERVICE_VERSION` | none (also read from `service.version` in `OTEL_RESOURCE_ATTRIBUTES`; left off the resource when unset, never defaulted) |
233
346
  | `disabled` | `RIUS_DISABLED` | `false` |
234
347
  | `sampleRate` | `RIUS_SAMPLE_RATE` | `1.0` |
235
348
  | `captureContent` | `RIUS_CAPTURE_CONTENT` | `true` |
@@ -237,12 +350,52 @@ variables, then the defaults below.
237
350
  | `heartbeat` | `RIUS_HEARTBEAT` | `true` |
238
351
  | `heartbeatInterval` | `RIUS_HEARTBEAT_INTERVAL` | `15` (seconds; clamped 5-300) |
239
352
  | `agentName` | `RIUS_AGENT_NAME` | `serviceName` |
353
+ | `mainAgentId` | `RIUS_MAIN_AGENT_ID` | none (a STABLE id for the agent this process is; never a transient in-memory id) |
354
+ | `mainAgentDescription` | `RIUS_MAIN_AGENT_DESCRIPTION` | none |
355
+ | `mainAgentVersion` | `RIUS_MAIN_AGENT_VERSION` | none (the version of the AGENT DEFINITION, never derived from `serviceVersion`) |
240
356
  | `partialSpans` | `RIUS_PARTIAL_SPANS` | `false` |
241
357
  | `partialSpansDelay` | `RIUS_PARTIAL_SPANS_DELAY` | `0` (seconds; clamped 0-60) |
242
358
  | `sessionId` | `RIUS_SESSION_ID` | none (process-wide session default; `withSession` overrides it) |
359
+ | `bridgeForeignProvider` | `RIUS_BRIDGE_FOREIGN_PROVIDER` | `false` (also export the spans of another OpenTelemetry SDK that owns the global provider; see below) |
243
360
 
244
361
  Traces are posted to `<endpoint>/v1/traces`.
245
362
 
363
+ `serviceVersion` is stamped on the resource as `service.version`, so the
364
+ console can tell one build of an agent from another. It has no default on
365
+ purpose: an invented version would put every unversioned process under the
366
+ same wrong number, the way `unknown_service` does for unnamed ones, so the
367
+ key is simply absent until you supply one. This SDK builds its resource
368
+ explicitly, which in OpenTelemetry for JavaScript means the OTel environment
369
+ is not merged into it, so `OTEL_RESOURCE_ATTRIBUTES=service.version=...` is
370
+ read here by hand as the last tier after the option and `RIUS_SERVICE_VERSION`.
371
+
372
+ The agent this process *is* is described on the resource under
373
+ `rius.main_agent.*`: `rius.main_agent.name` (the resolved `agentName`, absent
374
+ when nothing was named and it is still the `unknown_service` placeholder),
375
+ plus the optional `rius.main_agent.id`, `.description` and `.version`.
376
+ `gen_ai.agent.name` is still stamped alongside them, unchanged: the new keys
377
+ are additive so that no deployment loses its agent identity between an SDK
378
+ upgrade and a backend one. The vendor namespace exists because
379
+ `gen_ai.agent.name` has no resource-level meaning in the conventions and on a
380
+ span means the agent being *invoked*, so one key was answering two questions.
381
+
382
+ `mainAgentId` must be a **stable** identifier — a registry id, a hosted
383
+ agent's ARN. A uuid minted at startup or another process-local id identifies
384
+ a *run*, not an agent, and would split one agent into a fresh bucket per
385
+ restart; `service.instance.id` already answers "which process".
386
+
387
+ Three versions travel separately and none is ever derived from another:
388
+
389
+ | Key | Answers |
390
+ | -------------------------- | --------------------------------------------------- |
391
+ | `service.version` | which build of the deployment is running |
392
+ | `rius.main_agent.version` | which version of the agent *definition* it runs — its prompt, tools and policy |
393
+ | `gen_ai.agent.version` | which version of the agent a given AGENT span *invoked* |
394
+
395
+ A service can sit at `2.3.1` while the agent definition it runs is at `7` and
396
+ the agent it calls out to is at `1.0.0`; a prompt change moves the second
397
+ without touching the first.
398
+
246
399
  `captureContent` controls whether input/output content (prompts,
247
400
  completions, tool arguments) is attached to spans. It defaults to
248
401
  `true`. Set it to `false`, or supply a `mask` function, if your spans
@@ -267,6 +420,27 @@ optional integration is loaded, so no third-party module is patched in your
267
420
  process. `client.ready` resolves to an empty list. Spans can still be
268
421
  created and are simply dropped.
269
422
 
423
+ ### Running next to another OpenTelemetry SDK
424
+
425
+ OpenTelemetry has one global tracer provider per process, and the first
426
+ registration wins. If another SDK (Langfuse v3, for one) claims it before
427
+ `init()`, your own job and tool spans go to that SDK only, while the LLM spans
428
+ inside them still reach Rius — so Rius receives traces whose parents it never
429
+ gets. Rius detects this whatever you configure: those LLM spans carry
430
+ `rius.parent.foreign=true`, the resource records
431
+ `rius.sdk.global_provider=foreign:<class>`, and the first one logs a warning.
432
+
433
+ To send the other SDK's spans to Rius as well, opt in with
434
+ `init({ bridgeForeignProvider: true })` or `RIUS_BRIDGE_FOREIGN_PROVIDER=true`.
435
+ Rius then attaches its own pipeline to that provider and exports its spans too,
436
+ subject to Rius's `sampleRate`, `captureContent` and `mask`, under the Rius
437
+ resource identity. It never modifies what the other SDK exports: that vendor
438
+ receives exactly what it would without Rius.
439
+
440
+ ```typescript
441
+ init({ bridgeForeignProvider: true }) // export the other SDK's spans as well
442
+ ```
443
+
270
444
  ## Agent liveness
271
445
 
272
446
  Traces only leave the process when a span ends, so an idle-but-alive agent
@@ -293,8 +467,9 @@ graceful exit if you want that distinction to show up.
293
467
 
294
468
  The heartbeat never keeps the process alive, its timer is unref'd, and never
295
469
  throws into your code; delivery failures are logged once per client and
296
- silent after that. `agentName` identifies the agent in heartbeat payloads
297
- and defaults to `serviceName`.
470
+ silent after that. `agentName` identifies the agent in heartbeat payloads and
471
+ on every span, where it is stamped on the resource as `gen_ai.agent.name`. It
472
+ defaults to `serviceName`, so the two agree unless you set it.
298
473
 
299
474
  ### Partial spans
300
475
 
@@ -346,6 +521,27 @@ Batching is tunable via the standard OpenTelemetry env vars:
346
521
  `OTEL_BSP_SCHEDULE_DELAY` (default 5000 ms), `OTEL_BSP_MAX_EXPORT_BATCH_SIZE`
347
522
  (default 512), and `OTEL_BSP_EXPORT_TIMEOUT` (default 30000 ms).
348
523
 
524
+ ### Span attribute limit
525
+
526
+ OpenTelemetry caps a span at 128 attributes by default, and a span that
527
+ reaches the cap silently refuses every new key. OpenInference writes one
528
+ attribute per message field and per tool field, so an agent loop with ten
529
+ tools passes 128 after about nine turns, and the keys written last are
530
+ the ones that matter most: the token usage (so cost), the output messages,
531
+ `output.value` and the finish reason.
532
+
533
+ `init()` therefore raises the span attribute count limit to 4096. The
534
+ standard variables still win: if `OTEL_SPAN_ATTRIBUTE_COUNT_LIMIT` or
535
+ `OTEL_ATTRIBUTE_COUNT_LIMIT` is set to a number, `init()` leaves the limit
536
+ to OpenTelemetry, which applies that value. The attribute value-length
537
+ limit is not changed.
538
+
539
+ If you run the OpenInference instrumentations **without** this SDK, on a
540
+ tracer provider of your own, set the limit yourself, for example
541
+ `OTEL_SPAN_ATTRIBUTE_COUNT_LIMIT=4096`, or pass
542
+ `spanLimits: { attributeCountLimit: 4096 }` to your provider. Otherwise
543
+ long agent calls lose their usage and output.
544
+
349
545
  ## Development
350
546
 
351
547
  ```bash