@runtypelabs/flue-otel 0.4.0 → 0.5.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -10,9 +10,11 @@ token usage, cost, loop iterations, tool calls and stop reason.
10
10
  - Works with Flue `>=1.0.0-beta.9` and 2.x from one entry point.
11
11
  - Depends on `@opentelemetry/api` only. It brings no SDK, provider, exporter or
12
12
  sampler; it writes through whatever your application has registered.
13
- - Emits identifiers, structure and metrics by default. No prompts, completions,
14
- tool arguments, tool results, error messages or stack traces reach the wire
15
- unless you opt in; the one opt-in is [tool content](#tool-content).
13
+ - Emits everything Runtype can show by default: identifiers, structure,
14
+ metrics, the transcript, the system prompt, and tool arguments and results.
15
+ Every content kind has its own off switch, and `content: false` sends shape
16
+ and cost only; see [Content](#content). Error messages and stack traces never
17
+ reach the wire.
16
18
 
17
19
  Full guide: [Instrumenting a Flue agent](https://docs.runtype.com/developer-guides/guides/flue-instrumentation).
18
20
 
@@ -85,11 +87,11 @@ process.on('beforeExit', () => void provider.shutdown())
85
87
 
86
88
  `createRuntypeFlueInstrumentation(options?)` accepts:
87
89
 
88
- | Option | Type | Purpose |
89
- | --------- | ------------------------ | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
90
- | `agents` | `Record<string, string>` | Runtype agent id per Flue agent name, e.g. `{ triage: 'agent_01j...' }`. Stamped as `runtype.agent.id` on that agent's `invoke_agent` span. See [Attribution](#attribution). |
91
- | `tracer` | `Tracer` | Where spans are written. Defaults to `trace.getTracer('@runtypelabs/flue-otel')` on the global provider. Pass one to route Runtype spans through a separate provider. |
92
- | `content` | `FlueContentOptions` | Opt in to tool arguments and results on `execute_tool` spans. Off by default. See [Tool content](#tool-content). |
90
+ | Option | Type | Purpose |
91
+ | --------- | ----------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
92
+ | `agents` | `Record<string, string>` | Runtype agent id per Flue agent name, e.g. `{ triage: 'agent_01j...' }`. Stamped as `runtype.agent.id` on that agent's `invoke_agent` span. See [Attribution](#attribution). |
93
+ | `tracer` | `Tracer` | Where spans are written. Defaults to `trace.getTracer('@runtypelabs/flue-otel')` on the global provider. Pass one to route Runtype spans through a separate provider. |
94
+ | `content` | `FlueContentOptions \| false` | Transcript, system prompt, tool arguments and results. All on by default; switch kinds off individually, or pass `false` for none. See [Content](#content). |
93
95
 
94
96
  The returned object implements Flue's full `FlueInstrumentation` contract:
95
97
  `observe` (builds spans from Flue's observation stream), `interceptor` (makes
@@ -111,7 +113,7 @@ subscriber. Its key is exported as `RUNTYPE_FLUE_INSTRUMENTATION_KEY`.
111
113
  version produced a trace even when the resource is not wired; the resource
112
114
  placement still wins where both are present.
113
115
  - `GEN_AI` and `RUNTYPE`: the attribute-name constants this package emits by
114
- default; `GEN_AI_CONTENT`: the two it emits under the `content` opt-in.
116
+ default; `GEN_AI_CONTENT`: the five content attributes it emits unless switched off.
115
117
  - `DEFAULT_CONTENT_MAX_CHARS` and `INGEST_CONTENT_ATTRIBUTE_CEILING`: the
116
118
  content size defaults, see [Tool content](#tool-content).
117
119
  - `RUNTYPE_STOP_REASONS`, `RUNTYPE_TOOL_TYPES`, `RUNTYPE_SCHEMA_VERSION`,
@@ -188,7 +190,7 @@ backends.
188
190
  | -------------------------------------------------------------- | --------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
189
191
  | `invoke_agent <agent>` | once per agent invocation | `gen_ai.agent.name`, `gen_ai.conversation.id`, the run's request/response model, summed token usage across every turn (`gen_ai.usage.*`), `runtype.stop_reason`, `runtype.tools.reported`, the highest loop iteration (`runtype.iteration`), `runtype.execution.id`, and `runtype.agent.id` when `agents` names the agent |
190
192
  | `chat <model>` | once per model turn | `gen_ai.provider.name`, request/response model, response id, finish reason, per-turn usage, request parameters (`max_tokens`, `temperature`, reasoning level, server address), `runtype.turn.id` / `runtype.turn.index` / `runtype.iteration`, and `runtype.provider.finish_reason` / `runtype.gateway.log_id` when the provider records them (Workers AI attaches both — the gateway log id is a pointer to that exact request in your AI Gateway dashboard) |
191
- | `execute_tool <tool>` | once per model-requested tool call | `gen_ai.tool.name`, `gen_ai.tool.call.id`, the loop position it belongs to, `gen_ai.tool.type` (`function`) for every tool except a sub-agent delegation, and `runtype.tool.type` when the tool's class is known, and `gen_ai.tool.call.arguments` / `gen_ai.tool.call.result` only under the [`content`](#tool-content) opt-in |
193
+ | `execute_tool <tool>` | once per model-requested tool call | `gen_ai.tool.name`, `gen_ai.tool.call.id`, the loop position it belongs to, `gen_ai.tool.type` (`function`) for every tool except a sub-agent delegation, and `runtype.tool.type` when the tool's class is known, and `gen_ai.tool.call.arguments` / `gen_ai.tool.call.result` unless [`content`](#tool-content) switches them off |
192
194
  | `flue.task <agent>`, `flue.compaction`, `flue.operation shell` | delegation, compaction, host shell call | correlation ids only; framework structure, not agent invocations |
193
195
 
194
196
  Every span also carries Flue's own `flue.*` correlation attributes
@@ -229,25 +231,58 @@ Rules the projection follows:
229
231
 
230
232
  An absent attribute costs one column. A wrong one renders as a measurement.
231
233
 
232
- ## Tool content
234
+ ## Content
233
235
 
234
- By default an `execute_tool` span says which tool ran and how long it took, and
235
- Runtype's trace view shows its card as "No tool content in this trace". To see
236
- arguments and results on that card, opt in:
236
+ Everything under `content` is **on by default** and independent: with no
237
+ option at all a run lands in Runtype with its transcript, system prompt, and
238
+ every tool's arguments and results — which is what lets Runtype show the
239
+ conversation and capture it as a replayable eval case. Two families:
240
+ [tool content](#tool-content) on `execute_tool` spans, and
241
+ [messages](#messages) on the `invoke_agent` envelope. Set a switch to `false`
242
+ to drop that one kind, or pass `content: false` to send shape and cost only.
237
243
 
238
244
  ```ts
239
245
  createRuntypeFlueInstrumentation({
240
246
  content: {
241
- toolArguments: true, // gen_ai.tool.call.arguments, set when the span opens
242
- toolResults: true, // gen_ai.tool.call.result, set just before the span ends
247
+ toolArguments: true,
248
+ toolResults: true,
249
+ inputMessages: true,
250
+ outputMessages: true,
251
+ systemInstructions: false, // the one most deployments turn off: large, and rarely per-run
252
+ maxChars: 65_536,
253
+ redact: (value, context) =>
254
+ context.kind === 'arguments' && context.toolName === 'delegate' ? undefined : value,
255
+ },
256
+ })
257
+
258
+ // Shape and cost only, no content of any kind:
259
+ createRuntypeFlueInstrumentation({ content: false })
260
+ ```
261
+
262
+ Content makes spans large. Runtype admits at most 384 KiB of content per span
263
+ and 2 MiB per export request, and rejects a request over 8 MiB whole; past the
264
+ 2 MiB mark the remaining runs in that request land without their transcript.
265
+ With content on, set your `BatchSpanProcessor`'s `maxExportBatchSize` so one
266
+ request stays under those bounds — 16 to 32 spans for chatty agents, 64 for
267
+ short runs — and watch `contentDropCount` on ingested runs.
268
+
269
+ ### Tool content
270
+
271
+ An `execute_tool` span carries the call's arguments and result, so Runtype's
272
+ trace view shows them on the tool card. To drop either:
273
+
274
+ ```ts
275
+ createRuntypeFlueInstrumentation({
276
+ content: {
277
+ toolArguments: false, // gen_ai.tool.call.arguments, set when the span opens
278
+ toolResults: false, // gen_ai.tool.call.result, set just before the span ends
243
279
  maxChars: 65_536, // per value, marker included; this is the default
244
- redact: (value, { kind, toolName }) => value, // optional; return undefined to drop
280
+ redact: (value, { kind, toolName }) => value, // optional; narrow on kind, then toolName
245
281
  },
246
282
  })
247
283
  ```
248
284
 
249
- Each switch is independent and defaults to off, so `{}` changes nothing. The
250
- values come from Flue's stable `tool_start.args` and `tool.result` payloads
285
+ The values come from Flue's stable `tool_start.args` and `tool.result` payloads
251
286
  (2.x's `effectiveResult`, what the model was actually shown, when present).
252
287
 
253
288
  The attribute names are the GenAI semconv `gen_ai.tool.call.arguments` and
@@ -278,23 +313,90 @@ Ceilings worth knowing, because they decide what Runtype keeps:
278
313
  the whole batch (envelope and usage included, not only the content). With
279
314
  both switches on, set `maxExportBatchSize` to something like 64.
280
315
 
281
- What is still never sent, whatever you set:
316
+ What tool content never carries, whatever you set:
282
317
 
283
318
  - **A failed tool's result.** On both Flue lines it is commonly the error
284
319
  payload, and error messages stay off the wire. The span still closes with its
285
320
  error type and exception class name.
286
321
  - **The host's own shell call** (`session.shell()`), which is not a tool row.
287
- - **Prompts and completions.** They live on `AgentMessage`, which Flue marks
288
- unstable. This package will read them once Flue exposes a stable shape. One
289
- edge to know: a delegation to a sub-agent is a tool call, and its
290
- **arguments carry the prompt handed to that sub-agent**. If that matters for
291
- your deployment, drop it with `redact` (return `undefined` when `toolName`
292
- is the delegation tool).
293
-
294
- `redact` runs on the raw value before encoding, so it can address fields by
295
- name; it receives `{ kind: 'arguments' | 'result', toolName }`. Returning
296
- `undefined` drops the attribute for that call. If it throws, the attribute is
297
- dropped and the span is otherwise unaffected.
322
+
323
+ One edge to know: a delegation to a sub-agent is a tool call, and its
324
+ **arguments carry the prompt handed to that sub-agent**. If that matters for
325
+ your deployment, drop it with `redact` (return `undefined` when `toolName` is
326
+ the delegation tool).
327
+
328
+ ### Messages
329
+
330
+ Three switches govern the conversation itself on the `invoke_agent` span, which
331
+ is the span Runtype reads a run's transcript and final output from. All three
332
+ are on unless switched off:
333
+
334
+ ```ts
335
+ createRuntypeFlueInstrumentation({
336
+ content: {
337
+ inputMessages: true, // gen_ai.input.messages
338
+ outputMessages: true, // gen_ai.output.messages
339
+ systemInstructions: false, // gen_ai.system_instructions
340
+ },
341
+ })
342
+ ```
343
+
344
+ - **`inputMessages`** is the conversation as the run's _last_ model turn saw
345
+ it, read from Flue's normalized `turn_request.request.input.messages`: every
346
+ agent turn replaces the capture, so a user message dispatched mid-run (a
347
+ signal delivery, `useDispatchMessage()`) is included, and so is the run's own
348
+ intermediate assistant text. One copy per run, on the envelope. It lands on
349
+ the run as its transcript and is the seed eval capture freezes a case from.
350
+ - **`outputMessages`** is the assistant text of the _last_ agent turn that
351
+ produced any, read from `turn.response.output`. A final step that only called
352
+ tools does not clear an earlier answer, and a failed turn's output is never
353
+ read. Runtype projects it onto the run's final output.
354
+ - **`systemInstructions`** is `turn_request.request.input.systemPrompt`,
355
+ sent as one text part (`[{ type: 'text', content }]`) so a prompt that happens
356
+ to be valid JSON is never re-read as data. It is usually the largest value a
357
+ run carries and rarely differs between runs, so it is the switch most
358
+ deployments turn off; Runtype prepends it to the transcript as a `system` row.
359
+
360
+ Messages are encoded as the GenAI semconv structured-message shape,
361
+ `[{ role, parts: [{ type: 'text', content }] }]`, and they are **text parts
362
+ only**: a `toolResult` message, a tool-call part, a thinking part and an image
363
+ part are dropped, as is any `system` row Flue put in the history (that is what
364
+ `systemInstructions` is for) and any Flue **signal** — Flue delivers its own
365
+ steering (`<signal type="…">…</signal>`, `useAgentFinish` continuations, stream
366
+ recovery) to the model as a user message wrapping one XML element, and those
367
+ are not something the user said. Tool content has its own switches above; the
368
+ rest is never exported. `AgentMessage`, which Flue marks unstable, is still
369
+ never read — the `Llm*` request and response payloads are byte-identical on
370
+ 1.x and 2.x.
371
+
372
+ Over `maxChars`, a message list drops its **oldest** messages whole first, so
373
+ the newest messages survive; a lone message that still does not fit has its
374
+ text cut with the truncation marker, and the JSON stays valid either way. The
375
+ system prompt is cut inside its text part. Because all three share one span,
376
+ each is additionally capped at a third of ingest's per-span ceiling
377
+ (128 KiB), whatever `maxChars` says; at the default ceiling they fit
378
+ comfortably.
379
+
380
+ ### Redaction
381
+
382
+ `redact` runs on every content value before encoding and truncation. It
383
+ receives the value and a context:
384
+
385
+ - `{ kind: 'arguments' | 'result', toolName }` with the raw Flue payload;
386
+ - `{ kind: 'input_messages' | 'output_messages' }` with the projected message
387
+ array (`[{ role, parts: [{ type: 'text', content }] }]`; `{ role, content }`
388
+ with a string or Flue-style `{ type: 'text', text }` parts is accepted back);
389
+ - `{ kind: 'system_instructions' }` with the prompt string; return a string or
390
+ text parts (`[{ type: 'text', content | text }]`) — anything else emits
391
+ nothing, because ingest would discard it silently.
392
+
393
+ `toolName` is present on the tool kinds and `undefined` on the others, so a
394
+ hook written for 0.4 — `({ kind, toolName }) => …` — keeps compiling; narrow on
395
+ `kind` before relying on it.
396
+
397
+ Return the value to keep it, a replacement to substitute it, or `undefined` to
398
+ drop that one attribute. If it throws, the attribute is dropped and the span is
399
+ otherwise unaffected.
298
400
 
299
401
  ## Flue compatibility
300
402
 
package/dist/index.cjs CHANGED
@@ -78,35 +78,156 @@ function mapSettlementOutcome(outcome) {
78
78
  return void 0;
79
79
  }
80
80
 
81
+ // src/messages.ts
82
+ var ROLES = /* @__PURE__ */ new Set(["system", "user", "assistant"]);
83
+ var FLUE_ROLES = /* @__PURE__ */ new Set(["user", "assistant"]);
84
+ var SIGNAL_MESSAGE = /^\s*<([A-Za-z][\w-]*)\b[^>]*>[\s\S]*<\/\1>\s*$/;
85
+ function projectFlueMessages(value) {
86
+ if (!Array.isArray(value)) return [];
87
+ const out = [];
88
+ for (const entry of value) {
89
+ const message = projectFlueMessage(entry);
90
+ if (message) out.push(message);
91
+ }
92
+ return out;
93
+ }
94
+ function projectFlueMessage(value) {
95
+ if (!isRecord(value)) return void 0;
96
+ const role = value.role;
97
+ if (typeof role !== "string" || !FLUE_ROLES.has(role)) return void 0;
98
+ const text = flueMessageText(value.content);
99
+ if (text === "") return void 0;
100
+ if (role === "user" && SIGNAL_MESSAGE.test(text)) return void 0;
101
+ return { role, parts: [{ type: "text", content: text }] };
102
+ }
103
+ function flueMessageText(content) {
104
+ if (typeof content === "string") return content;
105
+ if (!Array.isArray(content)) return "";
106
+ const texts = [];
107
+ for (const part of content) {
108
+ if (!isRecord(part) || part.type !== "text") continue;
109
+ if (typeof part.text === "string" && part.text.length > 0) texts.push(part.text);
110
+ }
111
+ return texts.join("\n");
112
+ }
113
+ function normalizeContentMessages(value) {
114
+ if (!Array.isArray(value)) return [];
115
+ const out = [];
116
+ for (const entry of value) {
117
+ if (!isRecord(entry)) continue;
118
+ const role = entry.role;
119
+ if (typeof role !== "string" || !ROLES.has(role)) continue;
120
+ const text = contentMessageText(entry);
121
+ if (text === "") continue;
122
+ out.push({ role, parts: [{ type: "text", content: text }] });
123
+ }
124
+ return out;
125
+ }
126
+ function contentTextOf(value) {
127
+ if (typeof value === "string") return value;
128
+ return Array.isArray(value) ? contentMessageText({ parts: value }) : "";
129
+ }
130
+ function contentMessageText(message) {
131
+ if (typeof message.content === "string") return message.content;
132
+ const parts = Array.isArray(message.parts) ? message.parts : Array.isArray(message.content) ? message.content : [];
133
+ const texts = [];
134
+ for (const part of parts) {
135
+ if (!isRecord(part) || part.type !== "text") continue;
136
+ const text = typeof part.content === "string" && part.content.length > 0 ? part.content : typeof part.text === "string" && part.text.length > 0 ? part.text : "";
137
+ if (text !== "") texts.push(text);
138
+ }
139
+ return texts.join("\n");
140
+ }
141
+ function isRecord(value) {
142
+ return value !== null && typeof value === "object" && !Array.isArray(value);
143
+ }
144
+
81
145
  // src/content.ts
82
146
  var DEFAULT_CONTENT_MAX_CHARS = 65536;
83
147
  var INGEST_CONTENT_ATTRIBUTE_CEILING = 262144;
84
- function resolveContentPolicy(options) {
148
+ var INGEST_CONTENT_SPAN_CEILING = 393216;
149
+ function resolveContentPolicy(setting) {
150
+ const options = setting === false ? void 0 : setting;
151
+ const on = (value) => setting !== false && value !== false;
85
152
  const requested = options?.maxChars;
86
153
  const maxChars = typeof requested === "number" && Number.isFinite(requested) && requested > 0 ? Math.min(Math.floor(requested), INGEST_CONTENT_ATTRIBUTE_CEILING) : DEFAULT_CONTENT_MAX_CHARS;
87
154
  return {
88
- toolArguments: options?.toolArguments === true,
89
- toolResults: options?.toolResults === true,
155
+ toolArguments: on(options?.toolArguments),
156
+ toolResults: on(options?.toolResults),
157
+ inputMessages: on(options?.inputMessages),
158
+ outputMessages: on(options?.outputMessages),
159
+ systemInstructions: on(options?.systemInstructions),
90
160
  maxChars,
161
+ envelopeMaxChars: Math.min(maxChars, Math.floor(INGEST_CONTENT_SPAN_CEILING / 3)),
91
162
  redact: options?.redact
92
163
  };
93
164
  }
94
165
  function encodeContentValue(value, policy, context3) {
95
166
  if (value === void 0) return void 0;
96
- let redacted = value;
97
- if (policy.redact) {
98
- try {
99
- redacted = policy.redact(value, context3);
100
- } catch {
101
- return void 0;
102
- }
103
- if (redacted === void 0) return void 0;
104
- }
167
+ const redacted = applyRedact(value, policy, context3);
168
+ if (redacted === void 0) return void 0;
105
169
  if (context3.kind === "arguments") return encodeArguments(redacted, policy.maxChars);
106
170
  const encoded = stringify(redacted);
107
171
  if (encoded === void 0) return void 0;
108
172
  return truncate(encoded, policy.maxChars);
109
173
  }
174
+ function encodeMessagesValue(messages, policy, kind) {
175
+ if (messages.length === 0) return void 0;
176
+ const redacted = applyRedact(messages, policy, { kind });
177
+ if (redacted === void 0) return void 0;
178
+ const all = normalizeContentMessages(redacted);
179
+ if (all.length === 0) return void 0;
180
+ const maxChars = policy.envelopeMaxChars;
181
+ let start = all.length;
182
+ let total = 2;
183
+ while (start > 0) {
184
+ const length = JSON.stringify(all[start - 1]).length + (start < all.length ? 1 : 0);
185
+ if (total + length > maxChars) break;
186
+ total += length;
187
+ start -= 1;
188
+ }
189
+ if (start === 0) return JSON.stringify(all);
190
+ if (start < all.length) return JSON.stringify(all.slice(start));
191
+ return encodeTruncatedMessage(all[all.length - 1], maxChars);
192
+ }
193
+ function encodeSystemInstructionsValue(value, policy) {
194
+ if (value.length === 0) return void 0;
195
+ const redacted = applyRedact(value, policy, { kind: "system_instructions" });
196
+ if (redacted === void 0) return void 0;
197
+ const text = contentTextOf(redacted);
198
+ if (text.length === 0) return void 0;
199
+ return encodeTextParts(text, policy.envelopeMaxChars);
200
+ }
201
+ function encodeTruncatedMessage(message, maxChars) {
202
+ const text = message.parts.map((part) => part.content).join("\n");
203
+ return fitJson(
204
+ (cut) => JSON.stringify([{ role: message.role, parts: [{ type: "text", content: cut }] }]),
205
+ text,
206
+ maxChars
207
+ );
208
+ }
209
+ function encodeTextParts(text, maxChars) {
210
+ return fitJson((cut) => JSON.stringify([{ type: "text", content: cut }]), text, maxChars);
211
+ }
212
+ function fitJson(wrap, text, maxChars) {
213
+ let encoded = wrap(text);
214
+ if (encoded.length <= maxChars) return encoded;
215
+ let budget = maxChars - wrap("").length;
216
+ for (let round = 0; round < 8 && budget > 0; round += 1) {
217
+ encoded = wrap(truncate(text, budget));
218
+ if (encoded.length <= maxChars) return encoded;
219
+ budget -= encoded.length - maxChars;
220
+ }
221
+ return void 0;
222
+ }
223
+ function applyRedact(value, policy, context3) {
224
+ if (!policy.redact) return value;
225
+ try {
226
+ return policy.redact(value, context3);
227
+ } catch {
228
+ return void 0;
229
+ }
230
+ }
110
231
  function encodeArguments(value, maxChars) {
111
232
  const shaped = isPlainObject(value) ? value : { value };
112
233
  const encoded = stringify(shaped);
@@ -154,7 +275,7 @@ function isHighSurrogate(code) {
154
275
  }
155
276
 
156
277
  // package.json
157
- var version = "0.4.0";
278
+ var version = "0.5.0";
158
279
 
159
280
  // src/semconv.ts
160
281
  var GEN_AI = {
@@ -183,6 +304,9 @@ var GEN_AI = {
183
304
  serverPort: "server.port"
184
305
  };
185
306
  var GEN_AI_CONTENT = {
307
+ inputMessages: "gen_ai.input.messages",
308
+ outputMessages: "gen_ai.output.messages",
309
+ systemInstructions: "gen_ai.system_instructions",
186
310
  toolCallArguments: "gen_ai.tool.call.arguments",
187
311
  toolCallResult: "gen_ai.tool.call.result"
188
312
  };
@@ -336,9 +460,6 @@ function createFlueProjection(options = {}) {
336
460
  ...isEnvelope ? { [GEN_AI.operationName]: GEN_AI_OPERATION_INVOKE_AGENT } : {},
337
461
  ...isEnvelope && event.agentName ? { [GEN_AI.agentName]: event.agentName } : {},
338
462
  ...isEnvelope && event.conversationId ? { [GEN_AI.conversationId]: event.conversationId } : {},
339
- // The execution id LIFTS the tier and is claim-checked at ingest like
340
- // any inbound id; it does not key the row, which stays derived from
341
- // the trace id.
342
463
  ...isEnvelope && event.submissionId ? { [RUNTYPE.executionId]: event.submissionId } : {},
343
464
  ...isEnvelope && agentId ? { [RUNTYPE.agentId]: agentId } : {},
344
465
  // WHY(README.md): the resource placement is the customer's to wire and is routinely skipped.
@@ -402,9 +523,7 @@ function createFlueProjection(options = {}) {
402
523
  {
403
524
  kind: "open",
404
525
  ref,
405
- // Deliberately NOT `invoke_agent`: see this module's header. A second
406
- // agent-invocation span in the trace is a coin flip over which one
407
- // becomes the run's envelope.
526
+ // WHY(packages/flue-otel/src/projection.ts): Use a task span, not invoke_agent; one trace is persisted as one execution.
408
527
  name: agent ? `flue.task ${agent}` : "flue.task",
409
528
  spanKind: "internal",
410
529
  ...parentRef ? { parentRef } : {},
@@ -444,8 +563,6 @@ function createFlueProjection(options = {}) {
444
563
  ref,
445
564
  name: "flue.compaction",
446
565
  spanKind: "internal",
447
- // `session.compact()` is callable by the host outside any turn, and a
448
- // `compact` operation is not intercepted, so this can be parentless.
449
566
  requiresParent: true,
450
567
  ...parentRef ? { parentRef } : {},
451
568
  ...event.timestamp ? { startTime: event.timestamp } : {},
@@ -510,6 +627,7 @@ function createFlueProjection(options = {}) {
510
627
  if (!envelope.modelPinned && purpose === "agent")
511
628
  envelope.requestModel = request.requestedModel;
512
629
  if (Array.isArray(request.input?.tools)) envelope.toolsReported = true;
630
+ if (position?.atEnvelopeScope) captureInput(envelope, request);
513
631
  }
514
632
  return [
515
633
  {
@@ -555,20 +673,16 @@ function createFlueProjection(options = {}) {
555
673
  if (purpose === "agent" && response.finishReason) {
556
674
  envelope.lastFinishReason = response.finishReason;
557
675
  }
676
+ if (content.outputMessages && purpose === "agent" && !event.taskId && !isError) {
677
+ const output = projectFlueMessage(response.output);
678
+ if (output?.role === "assistant") envelope.outputMessage = output;
679
+ }
558
680
  }
559
681
  const attributes = {
560
682
  ...response.responseModel ? { [GEN_AI.responseModel]: response.responseModel } : {},
561
683
  ...response.responseId ? { [GEN_AI.responseId]: response.responseId } : {},
562
684
  ...response.finishReason ? { [GEN_AI.finishReasons]: [response.finishReason] } : {},
563
685
  ...usageAttributes(response.usage),
564
- // Provider-diagnostic pointers, emitted only when Flue actually provides
565
- // them. Both are optional and provider-dependent (Workers AI attaches
566
- // both today); an absent value costs one column, a synthesized one would
567
- // render as a measurement. `gatewayLogId` is a pointer to the content
568
- // without shipping the content — a customer can click through to that
569
- // exact request in their own AI Gateway dashboard. `providerFinishReason`
570
- // is the provider's exact finish value before normalization, which our
571
- // `GEN_AI.finishReasons` above deliberately hides.
572
686
  ...response.providerFinishReason ? { [RUNTYPE.providerFinishReason]: response.providerFinishReason } : {},
573
687
  ...response.gatewayLogId ? { [RUNTYPE.gatewayLogId]: response.gatewayLogId } : {}
574
688
  };
@@ -599,8 +713,6 @@ function createFlueProjection(options = {}) {
599
713
  ref,
600
714
  name: "flue.operation shell",
601
715
  spanKind: "internal",
602
- // `session.shell()` is callable by the host outside any turn, and a
603
- // `shell` operation is not intercepted, so this can be parentless.
604
716
  requiresParent: true,
605
717
  ...parentRef ? { parentRef } : {},
606
718
  ...event.timestamp ? { startTime: event.timestamp } : {},
@@ -632,10 +744,6 @@ function createFlueProjection(options = {}) {
632
744
  [GEN_AI.toolName]: toolName,
633
745
  [GEN_AI.toolCallId]: toolCallId,
634
746
  ...args !== void 0 ? { [GEN_AI_CONTENT.toolCallArguments]: args } : {},
635
- // The semconv `gen_ai.tool.type` domain is the transport-level one
636
- // (`function` / `extension` / `datastore`); a sub-agent delegation is
637
- // not a plain function call, so it is left unclaimed there while
638
- // `runtype.tool.type` says what it actually is.
639
747
  ...isDelegation ? {} : { [GEN_AI.toolType]: "function" },
640
748
  ...toolType ? { [RUNTYPE.toolType]: toolType } : {},
641
749
  ...event.origin ? { [FLUE.toolOrigin]: event.origin } : {},
@@ -694,10 +802,28 @@ function createFlueProjection(options = {}) {
694
802
  } : {},
695
803
  ...stopReason ? { [RUNTYPE.stopReason]: stopReason } : {},
696
804
  ...envelope.toolsReported ? { [RUNTYPE.toolsReported]: true } : {},
697
- // The highest loop position reached. Ingest records `max + 1` as the run's
698
- // iteration count, so the envelope alone reports it correctly even when
699
- // the batch carrying the model turns flushed earlier.
700
- ...envelope.maxIteration !== null ? { [RUNTYPE.iteration]: envelope.maxIteration } : {}
805
+ ...envelope.maxIteration !== null ? { [RUNTYPE.iteration]: envelope.maxIteration } : {},
806
+ ...envelopeContentAttributes(envelope)
807
+ };
808
+ }
809
+ function captureInput(envelope, request) {
810
+ const input = request.input;
811
+ if (content.inputMessages) {
812
+ const messages = projectFlueMessages(input?.messages);
813
+ if (messages.length > 0) envelope.inputMessages = messages;
814
+ }
815
+ if (content.systemInstructions && typeof input?.systemPrompt === "string") {
816
+ envelope.systemInstructions = input.systemPrompt;
817
+ }
818
+ }
819
+ function envelopeContentAttributes(envelope) {
820
+ const input = envelope.inputMessages ? encodeMessagesValue(envelope.inputMessages, content, "input_messages") : void 0;
821
+ const system = envelope.systemInstructions ? encodeSystemInstructionsValue(envelope.systemInstructions, content) : void 0;
822
+ const output = envelope.outputMessage ? encodeMessagesValue([envelope.outputMessage], content, "output_messages") : void 0;
823
+ return {
824
+ ...input !== void 0 ? { [GEN_AI_CONTENT.inputMessages]: input } : {},
825
+ ...system !== void 0 ? { [GEN_AI_CONTENT.systemInstructions]: system } : {},
826
+ ...output !== void 0 ? { [GEN_AI_CONTENT.outputMessages]: output } : {}
701
827
  };
702
828
  }
703
829
  function resolveEnvelopeRef(event) {
@@ -817,9 +943,7 @@ function positionAttributes(position) {
817
943
  return {
818
944
  ...position.turnId ? { [RUNTYPE.turnId]: position.turnId } : {},
819
945
  [RUNTYPE.turnIndex]: position.index,
820
- // Only the root agent's loop has an ITERATION. A delegated sub-agent's
821
- // turns would otherwise inflate the run's iteration count, which ingest
822
- // computes as the highest position across every span in the trace.
946
+ // WHY(packages/flue-otel/src/projection.ts): Do not stamp delegated turns with root iteration positions; they would inflate the run iteration count.
823
947
  ...position.atEnvelopeScope ? { [RUNTYPE.iteration]: position.index } : {}
824
948
  };
825
949
  }
@@ -845,9 +969,6 @@ function closeIntent(ref, event, isError, errorInfo) {
845
969
  ...isError ? {
846
970
  error: {
847
971
  type: errorInfo?.type ?? "_OTHER",
848
- // The class NAME only. `errorInfo.message` and `.stack` are content
849
- // and this release emits none: a provider message routinely quotes
850
- // the prompt back, and a stack ships filesystem layout.
851
972
  ...errorInfo?.name ? { exceptionType: errorInfo.name } : {}
852
973
  }
853
974
  } : {}