@veryfront/ext-observability-opentelemetry 0.1.1250 → 0.1.1252-rc.13991
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +27 -13
- package/package.json +2 -2
package/README.md
CHANGED
|
@@ -80,22 +80,36 @@ With this extension registered, the workflow executor emits a `workflow.run` spa
|
|
|
80
80
|
execution and a `workflow.node <id>` span per node, and agent spans nest beneath the node
|
|
81
81
|
that produced them.
|
|
82
82
|
|
|
83
|
-
Node spans are named after the node id so a trace reads at a glance.
|
|
84
|
-
|
|
85
|
-
|
|
86
|
-
|
|
87
|
-
|
|
83
|
+
Node spans are named after the node id so a trace reads at a glance. Map and loop children
|
|
84
|
+
name their spans differently, and the difference decides which problem you get:
|
|
85
|
+
|
|
86
|
+
- **Map children carry generated ids.** A `map` over N items builds children `<map>_0`,
|
|
87
|
+
`<map>_1`, and so on, so it emits one span per item _and_ one distinct span name per item.
|
|
88
|
+
The same generated id also lands in `workflow.node.id`.
|
|
89
|
+
- **Loop children keep their authored ids.** A `loop` re-runs the same authored steps once
|
|
90
|
+
per iteration, so every iteration emits a span carrying that step's own id as both its
|
|
91
|
+
name and its `workflow.node.id`. Span names stay bounded no matter how long the loop runs,
|
|
92
|
+
and nothing on the span says which iteration produced it: only the span id and the start
|
|
93
|
+
timestamp separate iteration 0 from iteration 5. The `<loop>_iter_N` ids that appear in run
|
|
94
|
+
state are child-graph run ids, not span names, and no span is ever named after one.
|
|
95
|
+
|
|
96
|
+
Two consequences worth planning for:
|
|
88
97
|
|
|
89
98
|
- **Span volume.** A map over 10,000 items yields at least 10,000 node spans in a single
|
|
90
|
-
trace, before any agent spans nested beneath them. The framework applies no cap
|
|
91
|
-
|
|
92
|
-
`OTEL_BSP_MAX_EXPORT_BATCH_SIZE`
|
|
93
|
-
|
|
94
|
-
|
|
95
|
-
|
|
99
|
+
trace, before any agent spans nested beneath them. The framework applies no cap on `items`,
|
|
100
|
+
so the caller is the only bound on how large a map can get. `OTEL_BSP_MAX_QUEUE_SIZE` and
|
|
101
|
+
`OTEL_BSP_MAX_EXPORT_BATCH_SIZE` govern the SDK's export buffer, not span generation: once
|
|
102
|
+
the queue fills, further spans are dropped rather than exported. Collector-side tail
|
|
103
|
+
sampling decides whether to keep or drop an entire trace, not how many spans it contains.
|
|
104
|
+
Neither control substitutes for keeping map size bounded at the call site.
|
|
105
|
+
- **Name cardinality.** Backends that aggregate by span name, for example Tempo's metrics
|
|
106
|
+
generator, see one series per map item. Loop iterations add no name cardinality. Drop or
|
|
107
|
+
rewrite the `workflow.node <id>` span name at the collector if map fan-out matters for your
|
|
108
|
+
backend; `workflow.node.id` is a separate attribute and rewriting the span name leaves it
|
|
109
|
+
untouched unless you rewrite it too.
|
|
96
110
|
|
|
97
111
|
`workflow.run` is always a trace root. Started from an instrumented HTTP handler, webhook,
|
|
98
|
-
or approval callback it does **not** join that request's trace
|
|
112
|
+
or approval callback it does **not** join that request's trace: a run is durable work that
|
|
99
113
|
outlives whatever started it. Parked on an approval it can resume days later, so nesting it
|
|
100
114
|
under the request would leave an open span inside a finished trace, and OpenTelemetry's
|
|
101
115
|
default parent-based sampler would let a sampled-out request silently drop the entire run.
|
|
@@ -115,7 +129,7 @@ filtering on that attribute reassembles the whole run as it always did.
|
|
|
115
129
|
|
|
116
130
|
The link is built from a W3C `traceparent` persisted on the run record when each execution
|
|
117
131
|
claims it. A run executed with tracing disabled simply stores nothing and the next
|
|
118
|
-
execution links to nothing
|
|
132
|
+
execution links to nothing, so the chain degrades to `workflow.run_id` correlation.
|
|
119
133
|
|
|
120
134
|
Node spans carry `workflow.node.status`, and a failed node or run sets the span status to
|
|
121
135
|
ERROR, so the usual errored-spans filters in Jaeger, Tempo and Datadog work. A cancelled run
|
package/package.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "@veryfront/ext-observability-opentelemetry",
|
|
3
|
-
"version": "0.1.
|
|
3
|
+
"version": "0.1.1252-rc.13991",
|
|
4
4
|
"description": "Veryfront first-party extension package for ext-observability-opentelemetry",
|
|
5
5
|
"keywords": [
|
|
6
6
|
"veryfront",
|
|
@@ -111,7 +111,7 @@
|
|
|
111
111
|
"protobufjs": "7.6.5"
|
|
112
112
|
},
|
|
113
113
|
"peerDependencies": {
|
|
114
|
-
"veryfront": "^0.1.
|
|
114
|
+
"veryfront": "^0.1.1252-rc.13991"
|
|
115
115
|
},
|
|
116
116
|
"type": "module",
|
|
117
117
|
"types": "./esm/index.d.ts",
|