@runtypelabs/flue-otel 0.4.0 → 0.5.1

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -10,9 +10,11 @@ token usage, cost, loop iterations, tool calls and stop reason.
10
10
  - Works with Flue `>=1.0.0-beta.9` and 2.x from one entry point.
11
11
  - Depends on `@opentelemetry/api` only. It brings no SDK, provider, exporter or
12
12
  sampler; it writes through whatever your application has registered.
13
- - Emits identifiers, structure and metrics by default. No prompts, completions,
14
- tool arguments, tool results, error messages or stack traces reach the wire
15
- unless you opt in; the one opt-in is [tool content](#tool-content).
13
+ - Emits everything Runtype can show by default: identifiers, structure,
14
+ metrics, the transcript, the system prompt, and tool arguments and results.
15
+ Every content kind has its own off switch, and `content: false` sends shape
16
+ and cost only; see [Content](#content). Error messages and stack traces never
17
+ reach the wire.
16
18
 
17
19
  Full guide: [Instrumenting a Flue agent](https://docs.runtype.com/developer-guides/guides/flue-instrumentation).
18
20
 
@@ -85,11 +87,11 @@ process.on('beforeExit', () => void provider.shutdown())
85
87
 
86
88
  `createRuntypeFlueInstrumentation(options?)` accepts:
87
89
 
88
- | Option | Type | Purpose |
89
- | --------- | ------------------------ | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
90
- | `agents` | `Record<string, string>` | Runtype agent id per Flue agent name, e.g. `{ triage: 'agent_01j...' }`. Stamped as `runtype.agent.id` on that agent's `invoke_agent` span. See [Attribution](#attribution). |
91
- | `tracer` | `Tracer` | Where spans are written. Defaults to `trace.getTracer('@runtypelabs/flue-otel')` on the global provider. Pass one to route Runtype spans through a separate provider. |
92
- | `content` | `FlueContentOptions` | Opt in to tool arguments and results on `execute_tool` spans. Off by default. See [Tool content](#tool-content). |
90
+ | Option | Type | Purpose |
91
+ | --------- | ----------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
92
+ | `agents` | `Record<string, string>` | Runtype agent id per Flue agent name, e.g. `{ triage: 'agent_01j...' }`. Stamped as `runtype.agent.id` on that agent's `invoke_agent` span. See [Attribution](#attribution). |
93
+ | `tracer` | `Tracer` | Where spans are written. Defaults to `trace.getTracer('@runtypelabs/flue-otel')` on the global provider. Pass one to route Runtype spans through a separate provider. |
94
+ | `content` | `FlueContentOptions \| false` | Transcript, system prompt, tool arguments and results. All on by default; switch kinds off individually, or pass `false` for none. See [Content](#content). |
93
95
 
94
96
  The returned object implements Flue's full `FlueInstrumentation` contract:
95
97
  `observe` (builds spans from Flue's observation stream), `interceptor` (makes
@@ -111,7 +113,7 @@ subscriber. Its key is exported as `RUNTYPE_FLUE_INSTRUMENTATION_KEY`.
111
113
  version produced a trace even when the resource is not wired; the resource
112
114
  placement still wins where both are present.
113
115
  - `GEN_AI` and `RUNTYPE`: the attribute-name constants this package emits by
114
- default; `GEN_AI_CONTENT`: the two it emits under the `content` opt-in.
116
+ default; `GEN_AI_CONTENT`: the five content attributes it emits unless switched off.
115
117
  - `DEFAULT_CONTENT_MAX_CHARS` and `INGEST_CONTENT_ATTRIBUTE_CEILING`: the
116
118
  content size defaults, see [Tool content](#tool-content).
117
119
  - `RUNTYPE_STOP_REASONS`, `RUNTYPE_TOOL_TYPES`, `RUNTYPE_SCHEMA_VERSION`,
@@ -188,7 +190,7 @@ backends.
188
190
  | -------------------------------------------------------------- | --------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
189
191
  | `invoke_agent <agent>` | once per agent invocation | `gen_ai.agent.name`, `gen_ai.conversation.id`, the run's request/response model, summed token usage across every turn (`gen_ai.usage.*`), `runtype.stop_reason`, `runtype.tools.reported`, the highest loop iteration (`runtype.iteration`), `runtype.execution.id`, and `runtype.agent.id` when `agents` names the agent |
190
192
  | `chat <model>` | once per model turn | `gen_ai.provider.name`, request/response model, response id, finish reason, per-turn usage, request parameters (`max_tokens`, `temperature`, reasoning level, server address), `runtype.turn.id` / `runtype.turn.index` / `runtype.iteration`, and `runtype.provider.finish_reason` / `runtype.gateway.log_id` when the provider records them (Workers AI attaches both — the gateway log id is a pointer to that exact request in your AI Gateway dashboard) |
191
- | `execute_tool <tool>` | once per model-requested tool call | `gen_ai.tool.name`, `gen_ai.tool.call.id`, the loop position it belongs to, `gen_ai.tool.type` (`function`) for every tool except a sub-agent delegation, and `runtype.tool.type` when the tool's class is known, and `gen_ai.tool.call.arguments` / `gen_ai.tool.call.result` only under the [`content`](#tool-content) opt-in |
193
+ | `execute_tool <tool>` | once per model-requested tool call | `gen_ai.tool.name`, `gen_ai.tool.call.id`, the loop position it belongs to, `gen_ai.tool.type` (`function`) for every tool except a sub-agent delegation, and `runtype.tool.type` when the tool's class is known, and `gen_ai.tool.call.arguments` / `gen_ai.tool.call.result` unless [`content`](#tool-content) switches them off |
192
194
  | `flue.task <agent>`, `flue.compaction`, `flue.operation shell` | delegation, compaction, host shell call | correlation ids only; framework structure, not agent invocations |
193
195
 
194
196
  Every span also carries Flue's own `flue.*` correlation attributes
@@ -229,25 +231,58 @@ Rules the projection follows:
229
231
 
230
232
  An absent attribute costs one column. A wrong one renders as a measurement.
231
233
 
232
- ## Tool content
234
+ ## Content
233
235
 
234
- By default an `execute_tool` span says which tool ran and how long it took, and
235
- Runtype's trace view shows its card as "No tool content in this trace". To see
236
- arguments and results on that card, opt in:
236
+ Everything under `content` is **on by default** and independent: with no
237
+ option at all a run lands in Runtype with its transcript, system prompt, and
238
+ every tool's arguments and results — which is what lets Runtype show the
239
+ conversation and capture it as a replayable eval case. Two families:
240
+ [tool content](#tool-content) on `execute_tool` spans, and
241
+ [messages](#messages) on the `invoke_agent` envelope. Set a switch to `false`
242
+ to drop that one kind, or pass `content: false` to send shape and cost only.
237
243
 
238
244
  ```ts
239
245
  createRuntypeFlueInstrumentation({
240
246
  content: {
241
- toolArguments: true, // gen_ai.tool.call.arguments, set when the span opens
242
- toolResults: true, // gen_ai.tool.call.result, set just before the span ends
247
+ toolArguments: true,
248
+ toolResults: true,
249
+ inputMessages: true,
250
+ outputMessages: true,
251
+ systemInstructions: false, // the one most deployments turn off: large, and rarely per-run
252
+ maxChars: 65_536,
253
+ redact: (value, context) =>
254
+ context.kind === 'arguments' && context.toolName === 'delegate' ? undefined : value,
255
+ },
256
+ })
257
+
258
+ // Shape and cost only, no content of any kind:
259
+ createRuntypeFlueInstrumentation({ content: false })
260
+ ```
261
+
262
+ Content makes spans large. Runtype admits at most 384 KiB of content per span
263
+ and 2 MiB per export request, and rejects a request over 8 MiB whole; past the
264
+ 2 MiB mark the remaining runs in that request land without their transcript.
265
+ With content on, set your `BatchSpanProcessor`'s `maxExportBatchSize` so one
266
+ request stays under those bounds — 16 to 32 spans for chatty agents, 64 for
267
+ short runs — and watch `contentDropCount` on ingested runs.
268
+
269
+ ### Tool content
270
+
271
+ An `execute_tool` span carries the call's arguments and result, so Runtype's
272
+ trace view shows them on the tool card. To drop either:
273
+
274
+ ```ts
275
+ createRuntypeFlueInstrumentation({
276
+ content: {
277
+ toolArguments: false, // gen_ai.tool.call.arguments, set when the span opens
278
+ toolResults: false, // gen_ai.tool.call.result, set just before the span ends
243
279
  maxChars: 65_536, // per value, marker included; this is the default
244
- redact: (value, { kind, toolName }) => value, // optional; return undefined to drop
280
+ redact: (value, { kind, toolName }) => value, // optional; narrow on kind, then toolName
245
281
  },
246
282
  })
247
283
  ```
248
284
 
249
- Each switch is independent and defaults to off, so `{}` changes nothing. The
250
- values come from Flue's stable `tool_start.args` and `tool.result` payloads
285
+ The values come from Flue's stable `tool_start.args` and `tool.result` payloads
251
286
  (2.x's `effectiveResult`, what the model was actually shown, when present).
252
287
 
253
288
  The attribute names are the GenAI semconv `gen_ai.tool.call.arguments` and
@@ -278,23 +313,103 @@ Ceilings worth knowing, because they decide what Runtype keeps:
278
313
  the whole batch (envelope and usage included, not only the content). With
279
314
  both switches on, set `maxExportBatchSize` to something like 64.
280
315
 
281
- What is still never sent, whatever you set:
316
+ What tool content never carries, whatever you set:
282
317
 
283
318
  - **A failed tool's result.** On both Flue lines it is commonly the error
284
319
  payload, and error messages stay off the wire. The span still closes with its
285
320
  error type and exception class name.
286
321
  - **The host's own shell call** (`session.shell()`), which is not a tool row.
287
- - **Prompts and completions.** They live on `AgentMessage`, which Flue marks
288
- unstable. This package will read them once Flue exposes a stable shape. One
289
- edge to know: a delegation to a sub-agent is a tool call, and its
290
- **arguments carry the prompt handed to that sub-agent**. If that matters for
291
- your deployment, drop it with `redact` (return `undefined` when `toolName`
292
- is the delegation tool).
293
-
294
- `redact` runs on the raw value before encoding, so it can address fields by
295
- name; it receives `{ kind: 'arguments' | 'result', toolName }`. Returning
296
- `undefined` drops the attribute for that call. If it throws, the attribute is
297
- dropped and the span is otherwise unaffected.
322
+
323
+ One edge to know: a delegation to a sub-agent is a tool call, and its
324
+ **arguments carry the prompt handed to that sub-agent**. If that matters for
325
+ your deployment, drop it with `redact` (return `undefined` when `toolName` is
326
+ the delegation tool).
327
+
328
+ ### Messages
329
+
330
+ Three switches govern the conversation itself on the `invoke_agent` span, which
331
+ is the span Runtype reads a run's transcript and final output from. All three
332
+ are on unless switched off:
333
+
334
+ ```ts
335
+ createRuntypeFlueInstrumentation({
336
+ content: {
337
+ inputMessages: true, // gen_ai.input.messages
338
+ outputMessages: true, // gen_ai.output.messages
339
+ systemInstructions: false, // gen_ai.system_instructions
340
+ },
341
+ })
342
+ ```
343
+
344
+ - **`inputMessages`** is the conversation the run was given, read from Flue's
345
+ normalized `turn_request.request.input.messages`: the first agent turn's
346
+ history, plus any `kind: 'user'` message delivered mid-run that a later
347
+ turn's history carries. The run's own assistant rows are not repeated here —
348
+ the tool spans and `outputMessages` carry them — and a turn after a
349
+ compaction, which sees a summary in place of the conversation, never
350
+ replaces an earlier capture (Flue does not flag such a request, so the
351
+ `compaction` observation and the summary row it puts first are the signal).
352
+ One copy per run, on the envelope. It lands on
353
+ the run as its transcript and is the seed eval capture freezes a case from.
354
+ - **`outputMessages`** is the assistant text of the _last_ agent turn that
355
+ produced any, read from `turn.response.output`. A final step that only called
356
+ tools does not clear an earlier answer, and a failed turn's output is never
357
+ read. Runtype projects it onto the run's final output.
358
+ - **`systemInstructions`** is `turn_request.request.input.systemPrompt`,
359
+ sent as one text part (`[{ type: 'text', content }]`) so a prompt that happens
360
+ to be valid JSON is never re-read as data. It is usually the largest value a
361
+ run carries and rarely differs between runs, so it is the switch most
362
+ deployments turn off; Runtype prepends it to the transcript as a `system` row.
363
+
364
+ Messages are encoded as the GenAI semconv structured-message shape,
365
+ `[{ role, parts: [{ type: 'text', content }] }]`, and they are **text parts
366
+ only**: a `toolResult` message, a tool-call part, a thinking part and an image
367
+ part are dropped, as is any `system` row Flue put in the history (that is what
368
+ `systemInstructions` is for). A Flue **signal** — the steering Flue delivers to
369
+ the model as a user message (`<signal type="…">…</signal>` deliveries and
370
+ `dispatch()` inputs, `useAgentFinish` continuations, stream recovery, the
371
+ compaction summary) — is exported as a `system` row, verbatim with its framing,
372
+ so it is never read as something the user said and the reader sees exactly what
373
+ the model saw. Recognition is a text-shape test on Flue's exact framing (one
374
+ element with a leading `type` attribute, body on its own lines, one text part),
375
+ so it can only _reduce_ misclassification: HTML elements that conventionally
376
+ lead with `type` (`<script>`, `<style>`, `<button>`, …) and multi-part messages
377
+ always stay user rows, but a user message that exactly reproduces the framing
378
+ with another tag lands as a `system` row. Tool content has its own switches
379
+ above; the
380
+ rest is never exported. `AgentMessage`, which Flue marks unstable, is still
381
+ never read — the `Llm*` request and response payloads are byte-identical on
382
+ 1.x and 2.x.
383
+
384
+ Over `maxChars`, a message list drops its **oldest** messages whole first, so
385
+ the newest messages survive; a lone message that still does not fit has its
386
+ text cut with the truncation marker, and the JSON stays valid either way. The
387
+ system prompt is cut inside its text part. Because all three share one span,
388
+ each is additionally capped at a third of ingest's per-span ceiling
389
+ (128 KiB), whatever `maxChars` says — as a tool's arguments and result are each
390
+ capped at half of it (192 KiB); at the default ceiling everything fits
391
+ comfortably.
392
+
393
+ ### Redaction
394
+
395
+ `redact` runs on every content value before encoding and truncation. It
396
+ receives the value and a context:
397
+
398
+ - `{ kind: 'arguments' | 'result', toolName }` with the raw Flue payload;
399
+ - `{ kind: 'input_messages' | 'output_messages' }` with the projected message
400
+ array (`[{ role, parts: [{ type: 'text', content }] }]`; `{ role, content }`
401
+ with a string or Flue-style `{ type: 'text', text }` parts is accepted back);
402
+ - `{ kind: 'system_instructions' }` with the prompt string; return a string or
403
+ text parts (`[{ type: 'text', content | text }]`) — anything else emits
404
+ nothing, because ingest would discard it silently.
405
+
406
+ `toolName` is present on the tool kinds and `undefined` on the others, so a
407
+ hook written for 0.4 — `({ kind, toolName }) => …` — keeps compiling; narrow on
408
+ `kind` before relying on it.
409
+
410
+ Return the value to keep it, a replacement to substitute it, or `undefined` to
411
+ drop that one attribute. If it throws, the attribute is dropped and the span is
412
+ otherwise unaffected.
298
413
 
299
414
  ## Flue compatibility
300
415
 
package/dist/index.cjs CHANGED
@@ -78,49 +78,193 @@ function mapSettlementOutcome(outcome) {
78
78
  return void 0;
79
79
  }
80
80
 
81
+ // src/messages.ts
82
+ var ROLES = /* @__PURE__ */ new Set(["system", "user", "assistant"]);
83
+ var FLUE_ROLES = /* @__PURE__ */ new Set(["user", "assistant"]);
84
+ var SIGNAL_MESSAGE = /^<([A-Za-z_][\w.-]*) type="[^"]*"[^>]*>\n[\s\S]*\n<\/\1>$/;
85
+ var COMPACTION_SUMMARY_PREFIX = '<compaction type="context_summary"';
86
+ var HTML_TYPED_TAGS = /* @__PURE__ */ new Set([
87
+ "a",
88
+ "button",
89
+ "embed",
90
+ "input",
91
+ "link",
92
+ "menu",
93
+ "object",
94
+ "ol",
95
+ "script",
96
+ "source",
97
+ "style",
98
+ "ul"
99
+ ]);
100
+ function projectFlueMessages(value) {
101
+ if (!Array.isArray(value)) return [];
102
+ const out = [];
103
+ for (const entry of value) {
104
+ const message = projectFlueMessage(entry);
105
+ if (message) out.push(message);
106
+ }
107
+ return out;
108
+ }
109
+ function projectFlueMessage(value) {
110
+ if (!isRecord(value)) return void 0;
111
+ const role = value.role;
112
+ if (typeof role !== "string" || !FLUE_ROLES.has(role)) return void 0;
113
+ const text = flueMessageText(value.content);
114
+ if (text === "") return void 0;
115
+ if (role === "user" && isFlueSignal(value.content, text)) {
116
+ return { role: "system", parts: [{ type: "text", content: text }] };
117
+ }
118
+ return { role, parts: [{ type: "text", content: text }] };
119
+ }
120
+ function isFlueSignal(content, text) {
121
+ if (!isSingleText(content)) return false;
122
+ const signal = SIGNAL_MESSAGE.exec(text);
123
+ return signal !== null && !HTML_TYPED_TAGS.has(signal[1].toLowerCase());
124
+ }
125
+ function isSingleText(content) {
126
+ if (typeof content === "string") return true;
127
+ return Array.isArray(content) && content.length === 1;
128
+ }
129
+ function flueMessageText(content) {
130
+ return contentMessageText({ content });
131
+ }
132
+ function normalizeContentMessages(value) {
133
+ if (!Array.isArray(value)) return [];
134
+ const out = [];
135
+ for (const entry of value) {
136
+ if (!isRecord(entry)) continue;
137
+ const role = entry.role;
138
+ if (typeof role !== "string" || !ROLES.has(role)) continue;
139
+ const text = contentMessageText(entry);
140
+ if (text === "") continue;
141
+ out.push({ role, parts: [{ type: "text", content: text }] });
142
+ }
143
+ return out;
144
+ }
145
+ function contentTextOf(value) {
146
+ if (typeof value === "string") return value;
147
+ return Array.isArray(value) ? contentMessageText({ parts: value }) : "";
148
+ }
149
+ function contentMessageText(message) {
150
+ if (typeof message.content === "string") return message.content;
151
+ const parts = Array.isArray(message.parts) ? message.parts : Array.isArray(message.content) ? message.content : [];
152
+ const texts = [];
153
+ for (const part of parts) {
154
+ if (!isRecord(part) || part.type !== "text") continue;
155
+ const text = typeof part.content === "string" && part.content.length > 0 ? part.content : typeof part.text === "string" && part.text.length > 0 ? part.text : "";
156
+ if (text !== "") texts.push(text);
157
+ }
158
+ return texts.join("\n");
159
+ }
160
+ function isRecord(value) {
161
+ return value !== null && typeof value === "object" && !Array.isArray(value);
162
+ }
163
+
81
164
  // src/content.ts
82
165
  var DEFAULT_CONTENT_MAX_CHARS = 65536;
83
166
  var INGEST_CONTENT_ATTRIBUTE_CEILING = 262144;
84
- function resolveContentPolicy(options) {
167
+ var INGEST_CONTENT_SPAN_CEILING = 393216;
168
+ function resolveContentPolicy(setting) {
169
+ const options = setting === false ? void 0 : setting;
170
+ const on = (value) => setting !== false && value !== false;
85
171
  const requested = options?.maxChars;
86
172
  const maxChars = typeof requested === "number" && Number.isFinite(requested) && requested > 0 ? Math.min(Math.floor(requested), INGEST_CONTENT_ATTRIBUTE_CEILING) : DEFAULT_CONTENT_MAX_CHARS;
87
173
  return {
88
- toolArguments: options?.toolArguments === true,
89
- toolResults: options?.toolResults === true,
174
+ toolArguments: on(options?.toolArguments),
175
+ toolResults: on(options?.toolResults),
176
+ inputMessages: on(options?.inputMessages),
177
+ outputMessages: on(options?.outputMessages),
178
+ systemInstructions: on(options?.systemInstructions),
90
179
  maxChars,
180
+ envelopeMaxChars: Math.min(maxChars, Math.floor(INGEST_CONTENT_SPAN_CEILING / 3)),
181
+ toolMaxChars: Math.min(maxChars, Math.floor(INGEST_CONTENT_SPAN_CEILING / 2)),
91
182
  redact: options?.redact
92
183
  };
93
184
  }
94
185
  function encodeContentValue(value, policy, context3) {
95
186
  if (value === void 0) return void 0;
96
- let redacted = value;
97
- if (policy.redact) {
98
- try {
99
- redacted = policy.redact(value, context3);
100
- } catch {
101
- return void 0;
102
- }
103
- if (redacted === void 0) return void 0;
104
- }
105
- if (context3.kind === "arguments") return encodeArguments(redacted, policy.maxChars);
187
+ const redacted = applyRedact(value, policy, context3);
188
+ if (redacted === void 0) return void 0;
189
+ if (context3.kind === "arguments") return encodeArguments(redacted, policy.toolMaxChars);
106
190
  const encoded = stringify(redacted);
107
191
  if (encoded === void 0) return void 0;
108
- return truncate(encoded, policy.maxChars);
192
+ return truncate(encoded, policy.toolMaxChars);
193
+ }
194
+ function encodeMessagesValue(messages, policy, kind) {
195
+ if (messages.length === 0) return void 0;
196
+ const redacted = applyRedact(messages, policy, { kind });
197
+ if (redacted === void 0) return void 0;
198
+ const all = normalizeContentMessages(redacted);
199
+ if (all.length === 0) return void 0;
200
+ const maxChars = policy.envelopeMaxChars;
201
+ const perMessage = all.length > 1 ? Math.floor((maxChars - 3) / 2) : maxChars - 2;
202
+ const encoded = [];
203
+ for (const message of all) {
204
+ const one = JSON.stringify(message);
205
+ encoded.push(one.length <= perMessage ? one : encodeTruncatedMessage(message, perMessage) ?? "");
206
+ }
207
+ let start = all.length;
208
+ let total = 2;
209
+ while (start > 0) {
210
+ const length = encoded[start - 1].length + (start < all.length ? 1 : 0);
211
+ if (total + length > maxChars) break;
212
+ total += length;
213
+ start -= 1;
214
+ }
215
+ const kept = encoded.slice(start).filter((one) => one !== "");
216
+ if (kept.length === 0) return void 0;
217
+ return "[" + kept.join(",") + "]";
218
+ }
219
+ function encodeSystemInstructionsValue(value, policy) {
220
+ if (value.length === 0) return void 0;
221
+ const redacted = applyRedact(value, policy, { kind: "system_instructions" });
222
+ if (redacted === void 0) return void 0;
223
+ const text = contentTextOf(redacted);
224
+ if (text.length === 0) return void 0;
225
+ return encodeTextParts(text, policy.envelopeMaxChars);
226
+ }
227
+ function encodeTruncatedMessage(message, maxChars) {
228
+ const text = message.parts.map((part) => part.content).join("\n");
229
+ return fitJson(
230
+ (cut) => JSON.stringify({ role: message.role, parts: [{ type: "text", content: cut }] }),
231
+ text,
232
+ maxChars
233
+ );
234
+ }
235
+ function encodeTextParts(text, maxChars) {
236
+ return fitJson((cut) => JSON.stringify([{ type: "text", content: cut }]), text, maxChars);
237
+ }
238
+ function fitJson(wrap, text, maxChars) {
239
+ let encoded = wrap(text);
240
+ if (encoded.length <= maxChars) return encoded;
241
+ const shell = wrap("").length;
242
+ let budget = maxChars - shell;
243
+ for (let round = 0; round < 8; round += 1) {
244
+ if (budget <= truncationMarker(text.length).length) return void 0;
245
+ const cut = truncate(text, budget);
246
+ encoded = wrap(cut);
247
+ if (encoded.length <= maxChars) return encoded;
248
+ const rescaled = Math.floor(budget * (maxChars - shell) / (encoded.length - shell));
249
+ budget = Math.min(budget - 1, rescaled);
250
+ }
251
+ return void 0;
252
+ }
253
+ function applyRedact(value, policy, context3) {
254
+ if (!policy.redact) return value;
255
+ try {
256
+ return policy.redact(value, context3);
257
+ } catch {
258
+ return void 0;
259
+ }
109
260
  }
110
261
  function encodeArguments(value, maxChars) {
111
262
  const shaped = isPlainObject(value) ? value : { value };
112
263
  const encoded = stringify(shaped);
113
264
  if (encoded === void 0) return void 0;
114
265
  if (encoded.length <= maxChars) return encoded;
115
- let budget = maxChars - TRUNCATED_WRAPPER_OVERHEAD;
116
- for (let round = 0; round < 8 && budget > 0; round += 1) {
117
- const wrapped = JSON.stringify({ truncated: truncate(encoded, budget) });
118
- if (wrapped.length <= maxChars) return wrapped;
119
- budget -= wrapped.length - maxChars;
120
- }
121
- return JSON.stringify({ truncated: truncate(encoded, Math.max(0, budget)) });
266
+ return fitJson((cut) => JSON.stringify({ truncated: cut }), encoded, maxChars);
122
267
  }
123
- var TRUNCATED_WRAPPER_OVERHEAD = 16;
124
268
  function isPlainObject(value) {
125
269
  return value !== null && typeof value === "object" && !Array.isArray(value);
126
270
  }
@@ -154,7 +298,7 @@ function isHighSurrogate(code) {
154
298
  }
155
299
 
156
300
  // package.json
157
- var version = "0.4.0";
301
+ var version = "0.5.1";
158
302
 
159
303
  // src/semconv.ts
160
304
  var GEN_AI = {
@@ -183,6 +327,9 @@ var GEN_AI = {
183
327
  serverPort: "server.port"
184
328
  };
185
329
  var GEN_AI_CONTENT = {
330
+ inputMessages: "gen_ai.input.messages",
331
+ outputMessages: "gen_ai.output.messages",
332
+ systemInstructions: "gen_ai.system_instructions",
186
333
  toolCallArguments: "gen_ai.tool.call.arguments",
187
334
  toolCallResult: "gen_ai.tool.call.result"
188
335
  };
@@ -336,9 +483,6 @@ function createFlueProjection(options = {}) {
336
483
  ...isEnvelope ? { [GEN_AI.operationName]: GEN_AI_OPERATION_INVOKE_AGENT } : {},
337
484
  ...isEnvelope && event.agentName ? { [GEN_AI.agentName]: event.agentName } : {},
338
485
  ...isEnvelope && event.conversationId ? { [GEN_AI.conversationId]: event.conversationId } : {},
339
- // The execution id LIFTS the tier and is claim-checked at ingest like
340
- // any inbound id; it does not key the row, which stays derived from
341
- // the trace id.
342
486
  ...isEnvelope && event.submissionId ? { [RUNTYPE.executionId]: event.submissionId } : {},
343
487
  ...isEnvelope && agentId ? { [RUNTYPE.agentId]: agentId } : {},
344
488
  // WHY(README.md): the resource placement is the customer's to wire and is routinely skipped.
@@ -402,9 +546,7 @@ function createFlueProjection(options = {}) {
402
546
  {
403
547
  kind: "open",
404
548
  ref,
405
- // Deliberately NOT `invoke_agent`: see this module's header. A second
406
- // agent-invocation span in the trace is a coin flip over which one
407
- // becomes the run's envelope.
549
+ // WHY(packages/flue-otel/src/projection.ts): Use a task span, not invoke_agent; one trace is persisted as one execution.
408
550
  name: agent ? `flue.task ${agent}` : "flue.task",
409
551
  spanKind: "internal",
410
552
  ...parentRef ? { parentRef } : {},
@@ -444,8 +586,6 @@ function createFlueProjection(options = {}) {
444
586
  ref,
445
587
  name: "flue.compaction",
446
588
  spanKind: "internal",
447
- // `session.compact()` is callable by the host outside any turn, and a
448
- // `compact` operation is not intercepted, so this can be parentless.
449
589
  requiresParent: true,
450
590
  ...parentRef ? { parentRef } : {},
451
591
  ...event.timestamp ? { startTime: event.timestamp } : {},
@@ -459,6 +599,11 @@ function createFlueProjection(options = {}) {
459
599
  function onCompactionEnd(event) {
460
600
  const ref = compactionKey(event);
461
601
  if (!open.has(ref)) return [];
602
+ if (!event.taskId) {
603
+ const envelope = resolveEnvelopeRef(event);
604
+ const state = envelope ? envelopes.get(envelope) : void 0;
605
+ if (state) state.compacted = true;
606
+ }
462
607
  const intents = sweepDescendants(ref, event.timestamp);
463
608
  intents.push(
464
609
  closeIntent(ref, event, readBoolean(event, "isError") === true, readErrorInfo(event))
@@ -510,6 +655,7 @@ function createFlueProjection(options = {}) {
510
655
  if (!envelope.modelPinned && purpose === "agent")
511
656
  envelope.requestModel = request.requestedModel;
512
657
  if (Array.isArray(request.input?.tools)) envelope.toolsReported = true;
658
+ if (position?.atEnvelopeScope) captureInput(envelope, request);
513
659
  }
514
660
  return [
515
661
  {
@@ -555,20 +701,16 @@ function createFlueProjection(options = {}) {
555
701
  if (purpose === "agent" && response.finishReason) {
556
702
  envelope.lastFinishReason = response.finishReason;
557
703
  }
704
+ if (content.outputMessages && purpose === "agent" && !event.taskId && !isError) {
705
+ const output = projectFlueMessage(response.output);
706
+ if (output?.role === "assistant") envelope.outputMessage = output;
707
+ }
558
708
  }
559
709
  const attributes = {
560
710
  ...response.responseModel ? { [GEN_AI.responseModel]: response.responseModel } : {},
561
711
  ...response.responseId ? { [GEN_AI.responseId]: response.responseId } : {},
562
712
  ...response.finishReason ? { [GEN_AI.finishReasons]: [response.finishReason] } : {},
563
713
  ...usageAttributes(response.usage),
564
- // Provider-diagnostic pointers, emitted only when Flue actually provides
565
- // them. Both are optional and provider-dependent (Workers AI attaches
566
- // both today); an absent value costs one column, a synthesized one would
567
- // render as a measurement. `gatewayLogId` is a pointer to the content
568
- // without shipping the content — a customer can click through to that
569
- // exact request in their own AI Gateway dashboard. `providerFinishReason`
570
- // is the provider's exact finish value before normalization, which our
571
- // `GEN_AI.finishReasons` above deliberately hides.
572
714
  ...response.providerFinishReason ? { [RUNTYPE.providerFinishReason]: response.providerFinishReason } : {},
573
715
  ...response.gatewayLogId ? { [RUNTYPE.gatewayLogId]: response.gatewayLogId } : {}
574
716
  };
@@ -599,8 +741,6 @@ function createFlueProjection(options = {}) {
599
741
  ref,
600
742
  name: "flue.operation shell",
601
743
  spanKind: "internal",
602
- // `session.shell()` is callable by the host outside any turn, and a
603
- // `shell` operation is not intercepted, so this can be parentless.
604
744
  requiresParent: true,
605
745
  ...parentRef ? { parentRef } : {},
606
746
  ...event.timestamp ? { startTime: event.timestamp } : {},
@@ -632,10 +772,6 @@ function createFlueProjection(options = {}) {
632
772
  [GEN_AI.toolName]: toolName,
633
773
  [GEN_AI.toolCallId]: toolCallId,
634
774
  ...args !== void 0 ? { [GEN_AI_CONTENT.toolCallArguments]: args } : {},
635
- // The semconv `gen_ai.tool.type` domain is the transport-level one
636
- // (`function` / `extension` / `datastore`); a sub-agent delegation is
637
- // not a plain function call, so it is left unclaimed there while
638
- // `runtype.tool.type` says what it actually is.
639
775
  ...isDelegation ? {} : { [GEN_AI.toolType]: "function" },
640
776
  ...toolType ? { [RUNTYPE.toolType]: toolType } : {},
641
777
  ...event.origin ? { [FLUE.toolOrigin]: event.origin } : {},
@@ -694,10 +830,43 @@ function createFlueProjection(options = {}) {
694
830
  } : {},
695
831
  ...stopReason ? { [RUNTYPE.stopReason]: stopReason } : {},
696
832
  ...envelope.toolsReported ? { [RUNTYPE.toolsReported]: true } : {},
697
- // The highest loop position reached. Ingest records `max + 1` as the run's
698
- // iteration count, so the envelope alone reports it correctly even when
699
- // the batch carrying the model turns flushed earlier.
700
- ...envelope.maxIteration !== null ? { [RUNTYPE.iteration]: envelope.maxIteration } : {}
833
+ ...envelope.maxIteration !== null ? { [RUNTYPE.iteration]: envelope.maxIteration } : {},
834
+ ...envelopeContentAttributes(envelope)
835
+ };
836
+ }
837
+ function captureInput(envelope, request) {
838
+ const input = request.input;
839
+ if (envelope.inputRaw && (envelope.compacted || isCompactedHistory(input?.messages))) return;
840
+ if (content.inputMessages && Array.isArray(input?.messages) && input.messages.length > 0) {
841
+ envelope.inputRaw = input.messages;
842
+ envelope.inputBaseCount ??= input.messages.length;
843
+ }
844
+ if (content.systemInstructions && typeof input?.systemPrompt === "string") {
845
+ envelope.systemInstructions = input.systemPrompt;
846
+ }
847
+ }
848
+ function isCompactedHistory(messages) {
849
+ if (!Array.isArray(messages) || messages.length === 0) return false;
850
+ const first = messages[0];
851
+ if (!first || first.role !== "user") return false;
852
+ return flueMessageText(first.content).startsWith(COMPACTION_SUMMARY_PREFIX);
853
+ }
854
+ function projectInput(envelope) {
855
+ const raw = envelope.inputRaw ?? [];
856
+ const base = envelope.inputBaseCount ?? raw.length;
857
+ return [
858
+ ...projectFlueMessages(raw.slice(0, base)),
859
+ ...projectFlueMessages(raw.slice(base)).filter((message) => message.role !== "assistant")
860
+ ];
861
+ }
862
+ function envelopeContentAttributes(envelope) {
863
+ const input = envelope.inputRaw ? encodeMessagesValue(projectInput(envelope), content, "input_messages") : void 0;
864
+ const system = envelope.systemInstructions ? encodeSystemInstructionsValue(envelope.systemInstructions, content) : void 0;
865
+ const output = envelope.outputMessage ? encodeMessagesValue([envelope.outputMessage], content, "output_messages") : void 0;
866
+ return {
867
+ ...input !== void 0 ? { [GEN_AI_CONTENT.inputMessages]: input } : {},
868
+ ...system !== void 0 ? { [GEN_AI_CONTENT.systemInstructions]: system } : {},
869
+ ...output !== void 0 ? { [GEN_AI_CONTENT.outputMessages]: output } : {}
701
870
  };
702
871
  }
703
872
  function resolveEnvelopeRef(event) {
@@ -817,9 +986,7 @@ function positionAttributes(position) {
817
986
  return {
818
987
  ...position.turnId ? { [RUNTYPE.turnId]: position.turnId } : {},
819
988
  [RUNTYPE.turnIndex]: position.index,
820
- // Only the root agent's loop has an ITERATION. A delegated sub-agent's
821
- // turns would otherwise inflate the run's iteration count, which ingest
822
- // computes as the highest position across every span in the trace.
989
+ // WHY(packages/flue-otel/src/projection.ts): Do not stamp delegated turns with root iteration positions; they would inflate the run iteration count.
823
990
  ...position.atEnvelopeScope ? { [RUNTYPE.iteration]: position.index } : {}
824
991
  };
825
992
  }
@@ -845,9 +1012,6 @@ function closeIntent(ref, event, isError, errorInfo) {
845
1012
  ...isError ? {
846
1013
  error: {
847
1014
  type: errorInfo?.type ?? "_OTHER",
848
- // The class NAME only. `errorInfo.message` and `.stack` are content
849
- // and this release emits none: a provider message routinely quotes
850
- // the prompt back, and a stack ships filesystem layout.
851
1015
  ...errorInfo?.name ? { exceptionType: errorInfo.name } : {}
852
1016
  }
853
1017
  } : {}