agentcore-dashboard-metrics 0.1.0__tar.gz
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- agentcore_dashboard_metrics-0.1.0/.gitignore +7 -0
- agentcore_dashboard_metrics-0.1.0/LICENSE +21 -0
- agentcore_dashboard_metrics-0.1.0/PKG-INFO +272 -0
- agentcore_dashboard_metrics-0.1.0/README.md +243 -0
- agentcore_dashboard_metrics-0.1.0/docs/agentcore-logging.html +579 -0
- agentcore_dashboard_metrics-0.1.0/pyproject.toml +49 -0
- agentcore_dashboard_metrics-0.1.0/src/agentcore_dashboard_metrics/__init__.py +793 -0
- agentcore_dashboard_metrics-0.1.0/tests/test_core.py +96 -0
- agentcore_dashboard_metrics-0.1.0/tests/test_integrations.py +117 -0
- agentcore_dashboard_metrics-0.1.0/tests/test_tracers.py +276 -0
|
@@ -0,0 +1,21 @@
|
|
|
1
|
+
MIT License
|
|
2
|
+
|
|
3
|
+
Copyright (c) 2026 nakulan
|
|
4
|
+
|
|
5
|
+
Permission is hereby granted, free of charge, to any person obtaining a copy
|
|
6
|
+
of this software and associated documentation files (the "Software"), to deal
|
|
7
|
+
in the Software without restriction, including without limitation the rights
|
|
8
|
+
to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
|
|
9
|
+
copies of the Software, and to permit persons to whom the Software is
|
|
10
|
+
furnished to do so, subject to the following conditions:
|
|
11
|
+
|
|
12
|
+
The above copyright notice and this permission notice shall be included in all
|
|
13
|
+
copies or substantial portions of the Software.
|
|
14
|
+
|
|
15
|
+
THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
|
|
16
|
+
IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
|
|
17
|
+
FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
|
|
18
|
+
AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
|
|
19
|
+
LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
|
|
20
|
+
OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
|
|
21
|
+
SOFTWARE.
|
|
@@ -0,0 +1,272 @@
|
|
|
1
|
+
Metadata-Version: 2.5
|
|
2
|
+
Name: agentcore-dashboard-metrics
|
|
3
|
+
Version: 0.1.0
|
|
4
|
+
Summary: Plug-and-play turn-metrics logging for LangGraph/Strands/CrewAI agents on Bedrock AgentCore, shaped for CloudWatch Logs Insights + Grafana dashboards.
|
|
5
|
+
Author-email: nakulan <nakult721@gmail.com>
|
|
6
|
+
License-Expression: MIT
|
|
7
|
+
License-File: LICENSE
|
|
8
|
+
Keywords: agentcore,bedrock,cloudwatch,crewai,grafana,langchain,langgraph,logging,observability,strands-agents
|
|
9
|
+
Classifier: Development Status :: 3 - Alpha
|
|
10
|
+
Classifier: Intended Audience :: Developers
|
|
11
|
+
Classifier: License :: OSI Approved :: MIT License
|
|
12
|
+
Classifier: Operating System :: OS Independent
|
|
13
|
+
Classifier: Programming Language :: Python :: 3
|
|
14
|
+
Classifier: Programming Language :: Python :: 3 :: Only
|
|
15
|
+
Classifier: Programming Language :: Python :: 3.9
|
|
16
|
+
Classifier: Programming Language :: Python :: 3.10
|
|
17
|
+
Classifier: Programming Language :: Python :: 3.11
|
|
18
|
+
Classifier: Programming Language :: Python :: 3.12
|
|
19
|
+
Classifier: Programming Language :: Python :: 3.13
|
|
20
|
+
Classifier: Topic :: System :: Logging
|
|
21
|
+
Classifier: Topic :: System :: Monitoring
|
|
22
|
+
Classifier: Typing :: Typed
|
|
23
|
+
Requires-Python: >=3.9
|
|
24
|
+
Provides-Extra: dev
|
|
25
|
+
Requires-Dist: build; extra == 'dev'
|
|
26
|
+
Requires-Dist: pytest; extra == 'dev'
|
|
27
|
+
Requires-Dist: twine; extra == 'dev'
|
|
28
|
+
Description-Content-Type: text/markdown
|
|
29
|
+
|
|
30
|
+
# agentcore-dashboard-metrics
|
|
31
|
+
|
|
32
|
+
Plug-and-play turn-metrics logging for agents built on **Amazon Bedrock
|
|
33
|
+
AgentCore** — LangGraph, Strands Agents, or CrewAI — that emits one
|
|
34
|
+
fixed-shape CloudWatch log line your Grafana dashboards can `parse`
|
|
35
|
+
reliably, cheaply, and without ever hand-formatting a log string again.
|
|
36
|
+
|
|
37
|
+
## Why this exists
|
|
38
|
+
|
|
39
|
+
CloudWatch Logs Insights' `parse` command is a literal, positional string
|
|
40
|
+
matcher. It has no notion of "fields" — only fixed text tokens in a fixed
|
|
41
|
+
order. A dashboard panel built around:
|
|
42
|
+
|
|
43
|
+
```
|
|
44
|
+
parse message "UserID: * | SessionID: * | TurnID: * | Status: * | InputTokens: * | OutputTokens: * | TTFTMs: * | LatencyMs: *"
|
|
45
|
+
```
|
|
46
|
+
|
|
47
|
+
silently returns **"No data"** the instant any agent's log line drifts
|
|
48
|
+
from that exact shape — a swapped field, a quoted number, an extra space.
|
|
49
|
+
No error, no warning. Hand-writing that pipe-delimited f-string in every
|
|
50
|
+
agent's entrypoint is exactly the kind of thing that breaks quietly and
|
|
51
|
+
is expensive to debug after the fact.
|
|
52
|
+
|
|
53
|
+
This package is the one place that knows the correct shape. Import a
|
|
54
|
+
function, call it, and every dashboard reading that log group populates —
|
|
55
|
+
whichever framework your agent is written in.
|
|
56
|
+
|
|
57
|
+
## Install
|
|
58
|
+
|
|
59
|
+
```bash
|
|
60
|
+
pip install agentcore-dashboard-metrics
|
|
61
|
+
```
|
|
62
|
+
|
|
63
|
+
Zero required dependencies — the framework wrappers are fully duck-typed
|
|
64
|
+
against whatever `graph`/`agent`/`crew` object you pass in, so installing
|
|
65
|
+
this package never pulls in LangChain, Strands, or CrewAI for you.
|
|
66
|
+
|
|
67
|
+
## Quickstart — tracers (recommended)
|
|
68
|
+
|
|
69
|
+
Instantiate a tracer, call `.attach()` once, then use your agent **exactly as
|
|
70
|
+
you already do**. No wrapper function, no new return value to unpack — the
|
|
71
|
+
log line is emitted as a side effect.
|
|
72
|
+
|
|
73
|
+
### LangGraph (`create_react_agent`)
|
|
74
|
+
|
|
75
|
+
```python
|
|
76
|
+
from agentcore_dashboard_metrics import LangChainTracer
|
|
77
|
+
|
|
78
|
+
async def invoke(payload, context):
|
|
79
|
+
...
|
|
80
|
+
tracer = LangChainTracer(log, user_id=user_id, session_id=session_id)
|
|
81
|
+
tracer.attach(graph)
|
|
82
|
+
|
|
83
|
+
async for event in graph.astream_events({"messages": messages}, version="v2"):
|
|
84
|
+
... # completely unchanged
|
|
85
|
+
```
|
|
86
|
+
|
|
87
|
+
### Strands Agents
|
|
88
|
+
|
|
89
|
+
```python
|
|
90
|
+
from agentcore_dashboard_metrics import StrandsTracer
|
|
91
|
+
|
|
92
|
+
async def invoke(payload, context):
|
|
93
|
+
agent = get_or_create_agent()
|
|
94
|
+
tracer = StrandsTracer(log, user_id=user_id, session_id=session_id)
|
|
95
|
+
tracer.attach(agent)
|
|
96
|
+
|
|
97
|
+
async for event in agent.stream_async(payload.get("prompt")):
|
|
98
|
+
... # completely unchanged
|
|
99
|
+
```
|
|
100
|
+
|
|
101
|
+
### CrewAI
|
|
102
|
+
|
|
103
|
+
```python
|
|
104
|
+
from agentcore_dashboard_metrics import CrewAITracer
|
|
105
|
+
|
|
106
|
+
async def invoke(payload, context):
|
|
107
|
+
crew = get_or_create_crew()
|
|
108
|
+
tracer = CrewAITracer(log, user_id=user_id, session_id=session_id)
|
|
109
|
+
tracer.attach(crew)
|
|
110
|
+
|
|
111
|
+
result = await crew.kickoff_async(inputs={"prompt": payload.get("prompt")})
|
|
112
|
+
# completely unchanged — works for both the non-streaming case and
|
|
113
|
+
# crew.stream=True
|
|
114
|
+
```
|
|
115
|
+
|
|
116
|
+
**How it works:** `.attach()` monkeypatches the exact async streaming method
|
|
117
|
+
each framework already exposes for this (`astream_events` / `stream_async` /
|
|
118
|
+
`kickoff_async`) so every event/result is yielded through completely
|
|
119
|
+
unchanged — verified against a real compiled LangGraph graph, not just a
|
|
120
|
+
mock. Re-attaching a fresh tracer to an already-wrapped object (e.g. a
|
|
121
|
+
cached graph/agent reused across requests, each with a different
|
|
122
|
+
user_id/session_id) swaps which tracer is active instead of stacking
|
|
123
|
+
another layer of wrapping.
|
|
124
|
+
|
|
125
|
+
**The one limitation:** only the wrapped method is traced. `LangChainTracer`
|
|
126
|
+
traces `astream_events`, not a separate `ainvoke` call on the same graph;
|
|
127
|
+
`StrandsTracer` traces `stream_async`, not the sync `agent(prompt)` call.
|
|
128
|
+
Use whichever your entrypoint already calls.
|
|
129
|
+
|
|
130
|
+
## Quickstart — one-shot wrapper functions
|
|
131
|
+
|
|
132
|
+
If you don't want a persistent tracer object — e.g. a graph/agent/crew
|
|
133
|
+
that's rebuilt fresh on every call anyway — call one of these instead. Same
|
|
134
|
+
guarantees, different shape: each does the whole call for you and hands
|
|
135
|
+
back `(output_text, turn_id)`.
|
|
136
|
+
|
|
137
|
+
### LangGraph (`create_react_agent`)
|
|
138
|
+
|
|
139
|
+
```python
|
|
140
|
+
from agentcore_dashboard_metrics import run_agent_turn
|
|
141
|
+
|
|
142
|
+
async def invoke(payload, context):
|
|
143
|
+
...
|
|
144
|
+
output, turn_id = await run_agent_turn(
|
|
145
|
+
graph, messages, log,
|
|
146
|
+
user_id=user_id, session_id=session_id,
|
|
147
|
+
)
|
|
148
|
+
return {"result": output, "turn_id": turn_id}
|
|
149
|
+
```
|
|
150
|
+
|
|
151
|
+
### Strands Agents
|
|
152
|
+
|
|
153
|
+
```python
|
|
154
|
+
from agentcore_dashboard_metrics import run_strands_turn
|
|
155
|
+
|
|
156
|
+
async def invoke(payload, context):
|
|
157
|
+
agent = get_or_create_agent()
|
|
158
|
+
output, turn_id = await run_strands_turn(
|
|
159
|
+
agent, payload.get("prompt"), log,
|
|
160
|
+
user_id=user_id, session_id=session_id,
|
|
161
|
+
)
|
|
162
|
+
return {"result": output, "turn_id": turn_id}
|
|
163
|
+
```
|
|
164
|
+
|
|
165
|
+
### CrewAI
|
|
166
|
+
|
|
167
|
+
```python
|
|
168
|
+
from agentcore_dashboard_metrics import run_crewai_turn
|
|
169
|
+
|
|
170
|
+
async def invoke(payload, context):
|
|
171
|
+
crew = get_or_create_crew()
|
|
172
|
+
output, turn_id = await run_crewai_turn(
|
|
173
|
+
crew, {"prompt": payload.get("prompt")}, log,
|
|
174
|
+
user_id=user_id, session_id=session_id,
|
|
175
|
+
)
|
|
176
|
+
return {"result": output, "turn_id": turn_id}
|
|
177
|
+
```
|
|
178
|
+
|
|
179
|
+
Each wrapper streams the underlying agent, measures time-to-first-token
|
|
180
|
+
and total latency, accumulates token usage across however many model
|
|
181
|
+
calls happen in the turn, sets `status` to `success`/`error`, and emits
|
|
182
|
+
the log line — all of it, so your entrypoint has nothing left to
|
|
183
|
+
hand-format.
|
|
184
|
+
|
|
185
|
+
### Anything else (Strands multi-agent, CrewAI Flows, a custom loop)
|
|
186
|
+
|
|
187
|
+
```python
|
|
188
|
+
from agentcore_dashboard_metrics import TurnMetrics
|
|
189
|
+
|
|
190
|
+
with TurnMetrics(log, user_id=user_id, session_id=session_id) as m:
|
|
191
|
+
result = my_own_agent_call(...)
|
|
192
|
+
m.add_usage(input_tokens=result.input_tokens,
|
|
193
|
+
output_tokens=result.output_tokens)
|
|
194
|
+
m.mark_first_token() # optional, only if you stream
|
|
195
|
+
# the correctly-shaped line is logged automatically on exit,
|
|
196
|
+
# including on an exception (status is set to "error" for you,
|
|
197
|
+
# and the exception is re-raised — never swallowed)
|
|
198
|
+
```
|
|
199
|
+
|
|
200
|
+
### Already computed everything yourself?
|
|
201
|
+
|
|
202
|
+
```python
|
|
203
|
+
from agentcore_dashboard_metrics import log_turn_metrics
|
|
204
|
+
|
|
205
|
+
log_turn_metrics(
|
|
206
|
+
log, user_id=user_id, session_id=session_id, turn_id=turn_id,
|
|
207
|
+
status="success", input_tokens=412, output_tokens=88,
|
|
208
|
+
ttft_ms=640.12, latency_ms=1820.55,
|
|
209
|
+
)
|
|
210
|
+
```
|
|
211
|
+
|
|
212
|
+
## The log line
|
|
213
|
+
|
|
214
|
+
Every path above produces exactly this shape via a single `log.info(...)`
|
|
215
|
+
call:
|
|
216
|
+
|
|
217
|
+
```
|
|
218
|
+
Published metrics — UserID: nakul | SessionID: 826fc1dc-... | TurnID: 3f9c2e1a-... | Status: success | InputTokens: 412 | OutputTokens: 88 | TTFTMs: 640.12 | LatencyMs: 1820.55
|
|
219
|
+
```
|
|
220
|
+
|
|
221
|
+
Build your CloudWatch Logs Insights `parse` pattern against exactly that
|
|
222
|
+
text and every dashboard panel — `stats sum(input_tokens) by user_id`,
|
|
223
|
+
`stats avg(ttft_ms)`, `stats count(*) by status`, etc. — works against any
|
|
224
|
+
agent using this package, regardless of which of the three frameworks it's
|
|
225
|
+
built on.
|
|
226
|
+
|
|
227
|
+
### Field reference
|
|
228
|
+
|
|
229
|
+
| Field | Type | Notes |
|
|
230
|
+
|---|---|---|
|
|
231
|
+
| `UserID` | string | End user identifier |
|
|
232
|
+
| `SessionID` | string | Conversation/session identifier |
|
|
233
|
+
| `TurnID` | string | Unique per invocation; auto-generated (`new_turn_id()`) if not supplied |
|
|
234
|
+
| `Status` | string | Single word, lowercased, no spaces/pipes — `success` or `error` |
|
|
235
|
+
| `InputTokens` / `OutputTokens` | int | Summed across every model call in the turn |
|
|
236
|
+
| `TTFTMs` | float | Time to first streamed token, ms; `-1.00` if the call shape has no token-level stream (e.g. CrewAI's non-streaming `kickoff_async`) |
|
|
237
|
+
| `LatencyMs` | float | Total wall-clock time for the turn |
|
|
238
|
+
|
|
239
|
+
All fields are sanitized defensively — `None`, empty strings, and stray
|
|
240
|
+
`|`/newline characters in ids or status are coerced into something safe
|
|
241
|
+
rather than corrupting the `parse` pattern.
|
|
242
|
+
|
|
243
|
+
### Adding a new field
|
|
244
|
+
|
|
245
|
+
If you need an additional metric (e.g. `ModelID`, `ToolUsed`), append it
|
|
246
|
+
**after `LatencyMs`** in your own logging so it doesn't shift the position
|
|
247
|
+
of any field existing dashboards already parse — `parse` ignores trailing
|
|
248
|
+
fields a given query doesn't ask for.
|
|
249
|
+
|
|
250
|
+
## API reference
|
|
251
|
+
|
|
252
|
+
| Name | Use for |
|
|
253
|
+
|---|---|
|
|
254
|
+
| `LangChainTracer(log, *, user_id, session_id).attach(graph)` | LangGraph `create_react_agent` — instruments `astream_events` |
|
|
255
|
+
| `StrandsTracer(log, *, user_id, session_id).attach(agent)` | Strands `Agent` — instruments `stream_async` |
|
|
256
|
+
| `CrewAITracer(log, *, user_id, session_id).attach(crew)` | CrewAI `Crew` — instruments `kickoff_async` |
|
|
257
|
+
| `run_agent_turn(graph, messages, log, *, user_id, session_id, turn_id=None)` | LangGraph, one-shot |
|
|
258
|
+
| `run_strands_turn(agent, prompt, log, *, user_id, session_id, turn_id=None)` | Strands, one-shot |
|
|
259
|
+
| `run_crewai_turn(crew, inputs, log, *, user_id, session_id, turn_id=None)` | CrewAI, one-shot |
|
|
260
|
+
| `TurnMetrics(log, *, user_id, session_id, turn_id=None)` | Context manager for any other framework/loop |
|
|
261
|
+
| `log_turn_metrics(log, *, user_id, session_id, turn_id, status, input_tokens, output_tokens, ttft_ms, latency_ms)` | Raw formatter, if you've already computed everything |
|
|
262
|
+
| `new_turn_id()` | Generates a fresh turn id |
|
|
263
|
+
|
|
264
|
+
All three `run_*_turn()` functions are `async def` and return
|
|
265
|
+
`(output_text, turn_id)`. On an exception from the underlying
|
|
266
|
+
graph/agent/crew, the metrics line is still logged with
|
|
267
|
+
`status="error"` and the exception is re-raised unchanged — the same
|
|
268
|
+
guarantee applies to the tracers' wrapped methods.
|
|
269
|
+
|
|
270
|
+
## License
|
|
271
|
+
|
|
272
|
+
MIT
|
|
@@ -0,0 +1,243 @@
|
|
|
1
|
+
# agentcore-dashboard-metrics
|
|
2
|
+
|
|
3
|
+
Plug-and-play turn-metrics logging for agents built on **Amazon Bedrock
|
|
4
|
+
AgentCore** — LangGraph, Strands Agents, or CrewAI — that emits one
|
|
5
|
+
fixed-shape CloudWatch log line your Grafana dashboards can `parse`
|
|
6
|
+
reliably, cheaply, and without ever hand-formatting a log string again.
|
|
7
|
+
|
|
8
|
+
## Why this exists
|
|
9
|
+
|
|
10
|
+
CloudWatch Logs Insights' `parse` command is a literal, positional string
|
|
11
|
+
matcher. It has no notion of "fields" — only fixed text tokens in a fixed
|
|
12
|
+
order. A dashboard panel built around:
|
|
13
|
+
|
|
14
|
+
```
|
|
15
|
+
parse message "UserID: * | SessionID: * | TurnID: * | Status: * | InputTokens: * | OutputTokens: * | TTFTMs: * | LatencyMs: *"
|
|
16
|
+
```
|
|
17
|
+
|
|
18
|
+
silently returns **"No data"** the instant any agent's log line drifts
|
|
19
|
+
from that exact shape — a swapped field, a quoted number, an extra space.
|
|
20
|
+
No error, no warning. Hand-writing that pipe-delimited f-string in every
|
|
21
|
+
agent's entrypoint is exactly the kind of thing that breaks quietly and
|
|
22
|
+
is expensive to debug after the fact.
|
|
23
|
+
|
|
24
|
+
This package is the one place that knows the correct shape. Import a
|
|
25
|
+
function, call it, and every dashboard reading that log group populates —
|
|
26
|
+
whichever framework your agent is written in.
|
|
27
|
+
|
|
28
|
+
## Install
|
|
29
|
+
|
|
30
|
+
```bash
|
|
31
|
+
pip install agentcore-dashboard-metrics
|
|
32
|
+
```
|
|
33
|
+
|
|
34
|
+
Zero required dependencies — the framework wrappers are fully duck-typed
|
|
35
|
+
against whatever `graph`/`agent`/`crew` object you pass in, so installing
|
|
36
|
+
this package never pulls in LangChain, Strands, or CrewAI for you.
|
|
37
|
+
|
|
38
|
+
## Quickstart — tracers (recommended)
|
|
39
|
+
|
|
40
|
+
Instantiate a tracer, call `.attach()` once, then use your agent **exactly as
|
|
41
|
+
you already do**. No wrapper function, no new return value to unpack — the
|
|
42
|
+
log line is emitted as a side effect.
|
|
43
|
+
|
|
44
|
+
### LangGraph (`create_react_agent`)
|
|
45
|
+
|
|
46
|
+
```python
|
|
47
|
+
from agentcore_dashboard_metrics import LangChainTracer
|
|
48
|
+
|
|
49
|
+
async def invoke(payload, context):
|
|
50
|
+
...
|
|
51
|
+
tracer = LangChainTracer(log, user_id=user_id, session_id=session_id)
|
|
52
|
+
tracer.attach(graph)
|
|
53
|
+
|
|
54
|
+
async for event in graph.astream_events({"messages": messages}, version="v2"):
|
|
55
|
+
... # completely unchanged
|
|
56
|
+
```
|
|
57
|
+
|
|
58
|
+
### Strands Agents
|
|
59
|
+
|
|
60
|
+
```python
|
|
61
|
+
from agentcore_dashboard_metrics import StrandsTracer
|
|
62
|
+
|
|
63
|
+
async def invoke(payload, context):
|
|
64
|
+
agent = get_or_create_agent()
|
|
65
|
+
tracer = StrandsTracer(log, user_id=user_id, session_id=session_id)
|
|
66
|
+
tracer.attach(agent)
|
|
67
|
+
|
|
68
|
+
async for event in agent.stream_async(payload.get("prompt")):
|
|
69
|
+
... # completely unchanged
|
|
70
|
+
```
|
|
71
|
+
|
|
72
|
+
### CrewAI
|
|
73
|
+
|
|
74
|
+
```python
|
|
75
|
+
from agentcore_dashboard_metrics import CrewAITracer
|
|
76
|
+
|
|
77
|
+
async def invoke(payload, context):
|
|
78
|
+
crew = get_or_create_crew()
|
|
79
|
+
tracer = CrewAITracer(log, user_id=user_id, session_id=session_id)
|
|
80
|
+
tracer.attach(crew)
|
|
81
|
+
|
|
82
|
+
result = await crew.kickoff_async(inputs={"prompt": payload.get("prompt")})
|
|
83
|
+
# completely unchanged — works for both the non-streaming case and
|
|
84
|
+
# crew.stream=True
|
|
85
|
+
```
|
|
86
|
+
|
|
87
|
+
**How it works:** `.attach()` monkeypatches the exact async streaming method
|
|
88
|
+
each framework already exposes for this (`astream_events` / `stream_async` /
|
|
89
|
+
`kickoff_async`) so every event/result is yielded through completely
|
|
90
|
+
unchanged — verified against a real compiled LangGraph graph, not just a
|
|
91
|
+
mock. Re-attaching a fresh tracer to an already-wrapped object (e.g. a
|
|
92
|
+
cached graph/agent reused across requests, each with a different
|
|
93
|
+
user_id/session_id) swaps which tracer is active instead of stacking
|
|
94
|
+
another layer of wrapping.
|
|
95
|
+
|
|
96
|
+
**The one limitation:** only the wrapped method is traced. `LangChainTracer`
|
|
97
|
+
traces `astream_events`, not a separate `ainvoke` call on the same graph;
|
|
98
|
+
`StrandsTracer` traces `stream_async`, not the sync `agent(prompt)` call.
|
|
99
|
+
Use whichever your entrypoint already calls.
|
|
100
|
+
|
|
101
|
+
## Quickstart — one-shot wrapper functions
|
|
102
|
+
|
|
103
|
+
If you don't want a persistent tracer object — e.g. a graph/agent/crew
|
|
104
|
+
that's rebuilt fresh on every call anyway — call one of these instead. Same
|
|
105
|
+
guarantees, different shape: each does the whole call for you and hands
|
|
106
|
+
back `(output_text, turn_id)`.
|
|
107
|
+
|
|
108
|
+
### LangGraph (`create_react_agent`)
|
|
109
|
+
|
|
110
|
+
```python
|
|
111
|
+
from agentcore_dashboard_metrics import run_agent_turn
|
|
112
|
+
|
|
113
|
+
async def invoke(payload, context):
|
|
114
|
+
...
|
|
115
|
+
output, turn_id = await run_agent_turn(
|
|
116
|
+
graph, messages, log,
|
|
117
|
+
user_id=user_id, session_id=session_id,
|
|
118
|
+
)
|
|
119
|
+
return {"result": output, "turn_id": turn_id}
|
|
120
|
+
```
|
|
121
|
+
|
|
122
|
+
### Strands Agents
|
|
123
|
+
|
|
124
|
+
```python
|
|
125
|
+
from agentcore_dashboard_metrics import run_strands_turn
|
|
126
|
+
|
|
127
|
+
async def invoke(payload, context):
|
|
128
|
+
agent = get_or_create_agent()
|
|
129
|
+
output, turn_id = await run_strands_turn(
|
|
130
|
+
agent, payload.get("prompt"), log,
|
|
131
|
+
user_id=user_id, session_id=session_id,
|
|
132
|
+
)
|
|
133
|
+
return {"result": output, "turn_id": turn_id}
|
|
134
|
+
```
|
|
135
|
+
|
|
136
|
+
### CrewAI
|
|
137
|
+
|
|
138
|
+
```python
|
|
139
|
+
from agentcore_dashboard_metrics import run_crewai_turn
|
|
140
|
+
|
|
141
|
+
async def invoke(payload, context):
|
|
142
|
+
crew = get_or_create_crew()
|
|
143
|
+
output, turn_id = await run_crewai_turn(
|
|
144
|
+
crew, {"prompt": payload.get("prompt")}, log,
|
|
145
|
+
user_id=user_id, session_id=session_id,
|
|
146
|
+
)
|
|
147
|
+
return {"result": output, "turn_id": turn_id}
|
|
148
|
+
```
|
|
149
|
+
|
|
150
|
+
Each wrapper streams the underlying agent, measures time-to-first-token
|
|
151
|
+
and total latency, accumulates token usage across however many model
|
|
152
|
+
calls happen in the turn, sets `status` to `success`/`error`, and emits
|
|
153
|
+
the log line — all of it, so your entrypoint has nothing left to
|
|
154
|
+
hand-format.
|
|
155
|
+
|
|
156
|
+
### Anything else (Strands multi-agent, CrewAI Flows, a custom loop)
|
|
157
|
+
|
|
158
|
+
```python
|
|
159
|
+
from agentcore_dashboard_metrics import TurnMetrics
|
|
160
|
+
|
|
161
|
+
with TurnMetrics(log, user_id=user_id, session_id=session_id) as m:
|
|
162
|
+
result = my_own_agent_call(...)
|
|
163
|
+
m.add_usage(input_tokens=result.input_tokens,
|
|
164
|
+
output_tokens=result.output_tokens)
|
|
165
|
+
m.mark_first_token() # optional, only if you stream
|
|
166
|
+
# the correctly-shaped line is logged automatically on exit,
|
|
167
|
+
# including on an exception (status is set to "error" for you,
|
|
168
|
+
# and the exception is re-raised — never swallowed)
|
|
169
|
+
```
|
|
170
|
+
|
|
171
|
+
### Already computed everything yourself?
|
|
172
|
+
|
|
173
|
+
```python
|
|
174
|
+
from agentcore_dashboard_metrics import log_turn_metrics
|
|
175
|
+
|
|
176
|
+
log_turn_metrics(
|
|
177
|
+
log, user_id=user_id, session_id=session_id, turn_id=turn_id,
|
|
178
|
+
status="success", input_tokens=412, output_tokens=88,
|
|
179
|
+
ttft_ms=640.12, latency_ms=1820.55,
|
|
180
|
+
)
|
|
181
|
+
```
|
|
182
|
+
|
|
183
|
+
## The log line
|
|
184
|
+
|
|
185
|
+
Every path above produces exactly this shape via a single `log.info(...)`
|
|
186
|
+
call:
|
|
187
|
+
|
|
188
|
+
```
|
|
189
|
+
Published metrics — UserID: nakul | SessionID: 826fc1dc-... | TurnID: 3f9c2e1a-... | Status: success | InputTokens: 412 | OutputTokens: 88 | TTFTMs: 640.12 | LatencyMs: 1820.55
|
|
190
|
+
```
|
|
191
|
+
|
|
192
|
+
Build your CloudWatch Logs Insights `parse` pattern against exactly that
|
|
193
|
+
text and every dashboard panel — `stats sum(input_tokens) by user_id`,
|
|
194
|
+
`stats avg(ttft_ms)`, `stats count(*) by status`, etc. — works against any
|
|
195
|
+
agent using this package, regardless of which of the three frameworks it's
|
|
196
|
+
built on.
|
|
197
|
+
|
|
198
|
+
### Field reference
|
|
199
|
+
|
|
200
|
+
| Field | Type | Notes |
|
|
201
|
+
|---|---|---|
|
|
202
|
+
| `UserID` | string | End user identifier |
|
|
203
|
+
| `SessionID` | string | Conversation/session identifier |
|
|
204
|
+
| `TurnID` | string | Unique per invocation; auto-generated (`new_turn_id()`) if not supplied |
|
|
205
|
+
| `Status` | string | Single word, lowercased, no spaces/pipes — `success` or `error` |
|
|
206
|
+
| `InputTokens` / `OutputTokens` | int | Summed across every model call in the turn |
|
|
207
|
+
| `TTFTMs` | float | Time to first streamed token, ms; `-1.00` if the call shape has no token-level stream (e.g. CrewAI's non-streaming `kickoff_async`) |
|
|
208
|
+
| `LatencyMs` | float | Total wall-clock time for the turn |
|
|
209
|
+
|
|
210
|
+
All fields are sanitized defensively — `None`, empty strings, and stray
|
|
211
|
+
`|`/newline characters in ids or status are coerced into something safe
|
|
212
|
+
rather than corrupting the `parse` pattern.
|
|
213
|
+
|
|
214
|
+
### Adding a new field
|
|
215
|
+
|
|
216
|
+
If you need an additional metric (e.g. `ModelID`, `ToolUsed`), append it
|
|
217
|
+
**after `LatencyMs`** in your own logging so it doesn't shift the position
|
|
218
|
+
of any field existing dashboards already parse — `parse` ignores trailing
|
|
219
|
+
fields a given query doesn't ask for.
|
|
220
|
+
|
|
221
|
+
## API reference
|
|
222
|
+
|
|
223
|
+
| Name | Use for |
|
|
224
|
+
|---|---|
|
|
225
|
+
| `LangChainTracer(log, *, user_id, session_id).attach(graph)` | LangGraph `create_react_agent` — instruments `astream_events` |
|
|
226
|
+
| `StrandsTracer(log, *, user_id, session_id).attach(agent)` | Strands `Agent` — instruments `stream_async` |
|
|
227
|
+
| `CrewAITracer(log, *, user_id, session_id).attach(crew)` | CrewAI `Crew` — instruments `kickoff_async` |
|
|
228
|
+
| `run_agent_turn(graph, messages, log, *, user_id, session_id, turn_id=None)` | LangGraph, one-shot |
|
|
229
|
+
| `run_strands_turn(agent, prompt, log, *, user_id, session_id, turn_id=None)` | Strands, one-shot |
|
|
230
|
+
| `run_crewai_turn(crew, inputs, log, *, user_id, session_id, turn_id=None)` | CrewAI, one-shot |
|
|
231
|
+
| `TurnMetrics(log, *, user_id, session_id, turn_id=None)` | Context manager for any other framework/loop |
|
|
232
|
+
| `log_turn_metrics(log, *, user_id, session_id, turn_id, status, input_tokens, output_tokens, ttft_ms, latency_ms)` | Raw formatter, if you've already computed everything |
|
|
233
|
+
| `new_turn_id()` | Generates a fresh turn id |
|
|
234
|
+
|
|
235
|
+
All three `run_*_turn()` functions are `async def` and return
|
|
236
|
+
`(output_text, turn_id)`. On an exception from the underlying
|
|
237
|
+
graph/agent/crew, the metrics line is still logged with
|
|
238
|
+
`status="error"` and the exception is re-raised unchanged — the same
|
|
239
|
+
guarantee applies to the tracers' wrapped methods.
|
|
240
|
+
|
|
241
|
+
## License
|
|
242
|
+
|
|
243
|
+
MIT
|