traceburn 0.1.0__tar.gz

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (63) hide show
  1. traceburn-0.1.0/.gitignore +17 -0
  2. traceburn-0.1.0/CITATION.cff +18 -0
  3. traceburn-0.1.0/CONTRIBUTING.md +54 -0
  4. traceburn-0.1.0/LICENSE +21 -0
  5. traceburn-0.1.0/PKG-INFO +259 -0
  6. traceburn-0.1.0/README.md +233 -0
  7. traceburn-0.1.0/assets/flamegraph.png +0 -0
  8. traceburn-0.1.0/assets/trace-tree.png +0 -0
  9. traceburn-0.1.0/assets/waste-report.png +0 -0
  10. traceburn-0.1.0/assets/waterfall.png +0 -0
  11. traceburn-0.1.0/docs/concepts.md +57 -0
  12. traceburn-0.1.0/docs/instrumentation.md +65 -0
  13. traceburn-0.1.0/docs/quickstart.md +82 -0
  14. traceburn-0.1.0/docs/replay.md +49 -0
  15. traceburn-0.1.0/docs/waste-rules.md +123 -0
  16. traceburn-0.1.0/examples/cache_before_after.py +161 -0
  17. traceburn-0.1.0/examples/offline_demo.py +64 -0
  18. traceburn-0.1.0/examples/raw_anthropic.py +41 -0
  19. traceburn-0.1.0/examples/raw_openai.py +64 -0
  20. traceburn-0.1.0/hooks/traceburn_autoinstall.pth +1 -0
  21. traceburn-0.1.0/hooks/traceburn_autoinstall.py +22 -0
  22. traceburn-0.1.0/pyproject.toml +45 -0
  23. traceburn-0.1.0/src/traceburn/__init__.py +62 -0
  24. traceburn-0.1.0/src/traceburn/analyze/__init__.py +3 -0
  25. traceburn-0.1.0/src/traceburn/analyze/diff.py +243 -0
  26. traceburn-0.1.0/src/traceburn/analyze/flamegraph.py +113 -0
  27. traceburn-0.1.0/src/traceburn/analyze/replay.py +351 -0
  28. traceburn-0.1.0/src/traceburn/analyze/waste/__init__.py +88 -0
  29. traceburn-0.1.0/src/traceburn/analyze/waste/_common.py +158 -0
  30. traceburn-0.1.0/src/traceburn/analyze/waste/cache.py +126 -0
  31. traceburn-0.1.0/src/traceburn/analyze/waste/context_bloat.py +111 -0
  32. traceburn-0.1.0/src/traceburn/analyze/waste/duplicates.py +178 -0
  33. traceburn-0.1.0/src/traceburn/analyze/waste/loops.py +121 -0
  34. traceburn-0.1.0/src/traceburn/analyze/waste/model_overkill.py +94 -0
  35. traceburn-0.1.0/src/traceburn/cli.py +336 -0
  36. traceburn-0.1.0/src/traceburn/instrument/__init__.py +47 -0
  37. traceburn-0.1.0/src/traceburn/instrument/_util.py +260 -0
  38. traceburn-0.1.0/src/traceburn/instrument/anthropic.py +428 -0
  39. traceburn-0.1.0/src/traceburn/instrument/openai.py +458 -0
  40. traceburn-0.1.0/src/traceburn/pricing.json +290 -0
  41. traceburn-0.1.0/src/traceburn/pricing.py +180 -0
  42. traceburn-0.1.0/src/traceburn/recorder.py +372 -0
  43. traceburn-0.1.0/src/traceburn/schema.py +158 -0
  44. traceburn-0.1.0/src/traceburn/store.py +393 -0
  45. traceburn-0.1.0/src/traceburn/ui/__init__.py +5 -0
  46. traceburn-0.1.0/src/traceburn/ui/server.py +172 -0
  47. traceburn-0.1.0/src/traceburn/ui/static/app.js +454 -0
  48. traceburn-0.1.0/src/traceburn/ui/static/index.html +48 -0
  49. traceburn-0.1.0/src/traceburn/ui/static/style.css +154 -0
  50. traceburn-0.1.0/tests/conftest.py +42 -0
  51. traceburn-0.1.0/tests/test_autoinstall.py +35 -0
  52. traceburn-0.1.0/tests/test_cli.py +78 -0
  53. traceburn-0.1.0/tests/test_diff.py +122 -0
  54. traceburn-0.1.0/tests/test_flamegraph.py +83 -0
  55. traceburn-0.1.0/tests/test_instrument_anthropic.py +332 -0
  56. traceburn-0.1.0/tests/test_instrument_openai.py +451 -0
  57. traceburn-0.1.0/tests/test_no_network.py +70 -0
  58. traceburn-0.1.0/tests/test_pricing.py +134 -0
  59. traceburn-0.1.0/tests/test_recorder.py +281 -0
  60. traceburn-0.1.0/tests/test_replay.py +370 -0
  61. traceburn-0.1.0/tests/test_store.py +204 -0
  62. traceburn-0.1.0/tests/test_ui.py +191 -0
  63. traceburn-0.1.0/tests/test_waste.py +479 -0
@@ -0,0 +1,17 @@
1
+ __pycache__/
2
+ *.pyc
3
+ *.egg-info/
4
+ .pytest_cache/
5
+ .ruff_cache/
6
+ dist/
7
+ build/
8
+ .venv/
9
+ venv/
10
+ .env
11
+ .env.*
12
+ .traceburn/
13
+ *.db
14
+ *.db-wal
15
+ *.db-shm
16
+ .DS_Store
17
+ node_modules/
@@ -0,0 +1,18 @@
1
+ cff-version: 1.2.0
2
+ message: "If you use traceburn in your work, please cite it as below."
3
+ title: "traceburn: a local-first tracer and efficiency profiler for AI agents"
4
+ type: software
5
+ authors:
6
+ - family-names: "Tran"
7
+ given-names: "Tommy"
8
+ email: "tommy.tranxhec@gmail.com"
9
+ repository-code: "https://github.com/TommyTranX/traceburn"
10
+ license: MIT
11
+ version: 0.1.0
12
+ date-released: "2026-07-05"
13
+ keywords:
14
+ - AI agents
15
+ - LLM observability
16
+ - tracing
17
+ - profiling
18
+ - cost efficiency
@@ -0,0 +1,54 @@
1
+ # Contributing
2
+
3
+ Contributions are welcome, and two kinds are especially wanted: new
4
+ instrumentation adapters and new waste rules. Both are deliberately small
5
+ interfaces; either is an afternoon of work.
6
+
7
+ ## Setup
8
+
9
+ ```
10
+ git clone https://github.com/TommyTranX/traceburn
11
+ cd traceburn
12
+ python -m venv .venv && . .venv/bin/activate
13
+ pip install -e ".[ui]" pytest openai anthropic httpx
14
+ pytest
15
+ ```
16
+
17
+ The test suite makes no network calls. Instrumentation tests run the real
18
+ provider SDKs over `httpx.MockTransport`, so they exercise the exact
19
+ objects the patchers see in production without a key.
20
+
21
+ ## Adding an instrumentation adapter
22
+
23
+ One module in `src/traceburn/instrument/`, three functions:
24
+ `is_available()`, `patch()`, `unpatch()`. The walkthrough with the span
25
+ attribute conventions, streaming wrappers, and the never-break rule is in
26
+ [docs/instrumentation.md](docs/instrumentation.md). Requirements:
27
+
28
+ - lazy imports, so the module is importable when the client library is not
29
+ - capture failures are swallowed and logged, never raised into the host
30
+ - tests via a mock transport: sync, async, streaming, tool calls, an error
31
+ - register it in `instrument/__init__.py`
32
+
33
+ ## Adding a waste rule
34
+
35
+ One module in `src/traceburn/analyze/waste/` with `run(ctx) -> list[Finding]`.
36
+ The contract, helpers, and the precision bar are in
37
+ [docs/waste-rules.md](docs/waste-rules.md). The short version: a positive
38
+ test, a negative test, no dollar figure that does not follow from observed
39
+ tokens and the pricing table, and an honest confidence level.
40
+
41
+ ## Ground rules
42
+
43
+ - `pytest` passes; new behavior comes with tests.
44
+ - The core package imports stdlib only. Anything heavier goes behind an
45
+ extra.
46
+ - No telemetry, no network calls of the tool's own, ever.
47
+ - Plain prose in docs and messages. Cost figures are always labeled
48
+ estimates.
49
+
50
+ ## Updating the pricing table
51
+
52
+ `src/traceburn/pricing.json` carries provider list prices with an `as_of`
53
+ date. Corrections and new models are welcome; cite the public pricing page
54
+ in the PR and update `as_of`.
@@ -0,0 +1,21 @@
1
+ MIT License
2
+
3
+ Copyright (c) 2026 Tommy Tran
4
+
5
+ Permission is hereby granted, free of charge, to any person obtaining a copy
6
+ of this software and associated documentation files (the "Software"), to deal
7
+ in the Software without restriction, including without limitation the rights
8
+ to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
9
+ copies of the Software, and to permit persons to whom the Software is
10
+ furnished to do so, subject to the following conditions:
11
+
12
+ The above copyright notice and this permission notice shall be included in all
13
+ copies or substantial portions of the Software.
14
+
15
+ THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
16
+ IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
17
+ FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
18
+ AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
19
+ LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
20
+ OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
21
+ SOFTWARE.
@@ -0,0 +1,259 @@
1
+ Metadata-Version: 2.4
2
+ Name: traceburn
3
+ Version: 0.1.0
4
+ Summary: Local-first tracer and efficiency profiler for AI agents: cost and latency flamegraphs, deterministic replay, run diffs, and waste detection over one SQLite file.
5
+ Project-URL: Homepage, https://github.com/TommyTranX/traceburn
6
+ Project-URL: Repository, https://github.com/TommyTranX/traceburn
7
+ Project-URL: Issues, https://github.com/TommyTranX/traceburn/issues
8
+ Author-email: Tommy Tran <tommy.tranxhec@gmail.com>
9
+ License: MIT
10
+ License-File: LICENSE
11
+ Keywords: agents,cost,llm,observability,profiling,tracing
12
+ Classifier: Development Status :: 3 - Alpha
13
+ Classifier: Intended Audience :: Developers
14
+ Classifier: License :: OSI Approved :: MIT License
15
+ Classifier: Programming Language :: Python :: 3
16
+ Classifier: Topic :: Software Development :: Debuggers
17
+ Classifier: Topic :: System :: Monitoring
18
+ Requires-Python: >=3.10
19
+ Provides-Extra: all
20
+ Requires-Dist: starlette>=0.37; extra == 'all'
21
+ Requires-Dist: uvicorn>=0.29; extra == 'all'
22
+ Provides-Extra: ui
23
+ Requires-Dist: starlette>=0.37; extra == 'ui'
24
+ Requires-Dist: uvicorn>=0.29; extra == 'ui'
25
+ Description-Content-Type: text/markdown
26
+
27
+ # traceburn
28
+
29
+ ## The finding that made me build this
30
+
31
+ I wrote a small support ticket triage agent: five tickets, one tool call each, a roughly 4,700
32
+ token static policy prompt sent fresh on every call, model claude-haiku-4-5. Running it uncached
33
+ cost $0.0539. traceburn's waste report looked at the trace, noticed that same 4,700 token prefix
34
+ going out uncached on all 10 calls, and estimated that about 82 percent of that spend was
35
+ avoidable. So I added exactly the one cache_control block it suggested and reran the same five
36
+ tickets: $0.0167.
37
+
38
+ That's a 69 percent measured saving. The tool's estimate landed within 18 percent of what actually
39
+ happened, close enough to trust as a first signal, not close enough to treat as gospel. That's
40
+ roughly how I want a cost estimator to behave.
41
+
42
+ Prices change and models get repriced, so do not take my word for it: the reproduction is a few
43
+ cents and a couple of minutes, at
44
+ [examples/cache_before_after.py](https://github.com/TommyTranX/traceburn/blob/main/examples/cache_before_after.py).
45
+ Measured 2026-07-05.
46
+
47
+ ![traceburn's waste report on the uncached run: $0.0539 total, about 82 percent flagged avoidable, with the repeated 5,618-token prefix identified as the cause](https://raw.githubusercontent.com/TommyTranX/traceburn/main/assets/waste-report.png)
48
+
49
+ That's the whole pitch in one run. traceburn is a local-first tracer and efficiency profiler for
50
+ AI agents, built on top of the openai and anthropic Python SDKs: a cost and latency flamegraph, a
51
+ waste report that quantifies avoidable spend instead of just gesturing at it, deterministic replay
52
+ of recorded calls, and run diffs, all backed by one SQLite file on disk. No account, no server to
53
+ stand up for basic use, no telemetry leaving your machine.
54
+
55
+ ## Install
56
+
57
+ Core tracer, stdlib only, zero dependencies:
58
+
59
+ ```bash
60
+ pip install traceburn
61
+ ```
62
+
63
+ Tracer plus the local web viewer (adds starlette and uvicorn):
64
+
65
+ ```bash
66
+ pip install "traceburn[ui]"
67
+ ```
68
+
69
+ Requires Python 3.10 or later.
70
+
71
+ ## Quickstart
72
+
73
+ Patch the SDKs at the top of your agent:
74
+
75
+ ```python
76
+ import traceburn
77
+ traceburn.install()
78
+
79
+ # your existing openai / anthropic code, unchanged
80
+ ```
81
+
82
+ Every sync call, async call, streaming response, tool call, and prompt cache hit on either SDK now
83
+ gets recorded as a span. Traces live in one SQLite file, `./.traceburn/traces.db` by default; set
84
+ the `TRACEBURN_DB` environment variable if you want it somewhere else. If you'd rather not touch
85
+ the source at all, wrap the run instead:
86
+
87
+ ```bash
88
+ TRACEBURN=1 python your_agent.py
89
+ ```
90
+
91
+ Then look at what happened:
92
+
93
+ ```bash
94
+ traceburn ui # opens the web viewer at 127.0.0.1:8765
95
+ traceburn ls # list recorded traces
96
+ traceburn show <id> # inspect one trace
97
+ traceburn waste <id> # run the waste report on one trace
98
+ traceburn diff <a> <b> # compare two traces span by span
99
+ ```
100
+
101
+ No API keys and nothing to configure: `python examples/offline_demo.py` records a simulated agent
102
+ run with realistic token counts, including one deliberate duplicate call, so you can see a real
103
+ trace, a real flamegraph, and a real waste finding inside a minute.
104
+
105
+ ## What it actually does
106
+
107
+ **Trace.** `traceburn.install()` patches both SDKs so every call becomes a span with tokens,
108
+ latency, and cost attached, no code changes required past that one line. Want manual control
109
+ instead, or you're using a framework outside the two supported SDKs? The explicit API, `@trace`,
110
+ `span()`, and `session()`, works by hand with anything.
111
+
112
+ ![traceburn's expandable trace tree, showing an agent's nested spans with per-call tokens and cost](https://raw.githubusercontent.com/TommyTranX/traceburn/main/assets/trace-tree.png)
113
+
114
+ **Flamegraph.** Spans render as a flamegraph you can size two ways: by wall-clock time or by
115
+ dollars spent, with self-time kept separate from time spent in children, so a slow parent span
116
+ doesn't hide which child call actually burned the seconds or the money.
117
+
118
+ **Waste report.** Heuristics look for duplicate calls, unused cache opportunities, bloated
119
+ prompts, model overkill, and retry loops. Each finding ships with a confidence level and, where the
120
+ numbers support it, a dollar figure. More on this below.
121
+
122
+ **Replay.** `traceburn.analyze.replay.replay()` plays recorded provider responses back through the
123
+ real SDK types, so your agent code runs again exactly as before with zero tokens spent. That's
124
+ useful for tests and for debugging without burning a budget. Streaming calls aren't replayable
125
+ yet; replay currently only serves non-streaming recordings.
126
+
127
+ **Diff.** `traceburn diff <a> <b>` lines up two traces span by span and reports the delta in
128
+ tokens, cost, and latency, alongside text diffs of the prompts and responses that changed between
129
+ the runs. Good for answering "did that prompt tweak actually help."
130
+
131
+ ## The web viewer
132
+
133
+ `traceburn ui` starts a small, read-only Starlette API in front of a vendored single-page app: no
134
+ build step, no CDN, everything ships in the package. It gives you the trace tree, the flamegraph,
135
+ a waterfall timeline, the waste report, and the run diff view.
136
+
137
+ It binds to 127.0.0.1 only and checks the request's host header against DNS rebinding, but it has
138
+ no authentication of any kind. That's a deliberate tradeoff: the viewer isn't meant to be reachable
139
+ from anywhere but your own machine.
140
+
141
+ ![traceburn's flamegraph view of an agent run, one row per depth, frame width proportional to latency](https://raw.githubusercontent.com/TommyTranX/traceburn/main/assets/flamegraph.png)
142
+
143
+ ![traceburn's waterfall view of the same run, a timeline of every call with its duration and cost](https://raw.githubusercontent.com/TommyTranX/traceburn/main/assets/waterfall.png)
144
+
145
+ ## Framework support
146
+
147
+ Today, that means the raw `openai` and `anthropic` Python SDKs, patched automatically by
148
+ `traceburn.install()`. If you're on something else, the explicit `span()` / `trace()` / `session()`
149
+ API works with any framework right now, by hand, since it doesn't care what's making the call.
150
+ LangChain, LlamaIndex, and anything already emitting OpenTelemetry GenAI spans aren't instrumented
151
+ automatically yet. That's real, planned work for v0.2, not something already built and just
152
+ undocumented, and it's covered in the roadmap below.
153
+
154
+ ## How the waste rules work
155
+
156
+ The rules live in `traceburn/analyze/waste/` and are documented in full at
157
+ [docs/waste-rules.md](https://github.com/TommyTranX/traceburn/blob/main/docs/waste-rules.md). Five
158
+ kinds of waste get checked for: duplicates (exact and near-duplicate repeated calls), cache (a
159
+ stable prompt prefix resent without ever hitting a provider cache), context_bloat (duplicate
160
+ blocks inside one prompt, or a huge prompt for a tiny output), model_overkill (a frontier-priced
161
+ model spent on trivial short calls, phrased as a suggestion and kept at low confidence on purpose),
162
+ and loops (retry storms, repeated identical tool calls, or a runaway step count). Every finding
163
+ also carries a confidence level: high, medium, low, or info.
164
+
165
+ Two principles govern all of them. A wrong finding does more damage than a missed one, so the
166
+ rules are tuned for precision over recall and would rather stay quiet than guess. And no dollar
167
+ figure is ever printed without observed tokens behind it and a price in the pricing table to
168
+ multiply against; everything the tool prints is labeled an estimate, because it is one.
169
+
170
+ ## Privacy
171
+
172
+ traceburn makes no network calls of its own and sends no telemetry anywhere. The only traffic on
173
+ the wire is your own calls to your own model provider, exactly as they'd happen without traceburn
174
+ installed. A test,
175
+ [tests/test_no_network.py](https://github.com/TommyTranX/traceburn/blob/main/tests/test_no_network.py),
176
+ blocks all socket access at the interpreter level and then runs the recorder, the store, every
177
+ analyzer, and the CLI against that blockade, to prove the point rather than just assert it. Auth
178
+ headers are never recorded in a trace, so an API key cannot end up sitting in a stored span.
179
+
180
+ ## Limitations
181
+
182
+ Instrumentation currently covers only the `openai` and `anthropic` Python SDKs. Within those,
183
+ `parse()` convenience methods and `with_raw_response` calls pass through untraced rather than being
184
+ recorded incorrectly, and multi-choice requests (`n > 1`) only record the first choice.
185
+
186
+ Cost figures come from a dated public pricing table
187
+ ([pricing.json](https://github.com/TommyTranX/traceburn/blob/main/src/traceburn/pricing.json)) and
188
+ don't model long-context pricing tiers or regional surcharges. Token counts prefer whatever the
189
+ provider itself reports as usage; anything estimated is flagged as estimated rather than presented
190
+ as measured.
191
+
192
+ The waste rules are heuristics tuned for precision over recall, which means they'll miss real
193
+ waste sooner than they'll invent fake waste, and every finding states its own confidence so you
194
+ can judge it accordingly. Streaming calls aren't replayable yet; only non-streaming recordings are.
195
+
196
+ The web viewer is read-only, bound to 127.0.0.1 only, and has no authentication.
197
+
198
+ ## Roadmap: v0.2
199
+
200
+ - OpenTelemetry GenAI span ingest plus OTLP export. This is also the path for capturing LangChain
201
+ and LlamaIndex traces, since it rides on their existing OTel instrumentation rather than
202
+ requiring bespoke adapters for each.
203
+ - A pytest plugin built on replay, for deterministic, token-free agent tests.
204
+ - litellm instrumentation.
205
+ - More waste rules.
206
+
207
+ ## Related work
208
+
209
+ [Langfuse](https://github.com/langfuse/langfuse) is a full open-source LLM platform: tracing,
210
+ evals, and prompt management, backed by Postgres and ClickHouse and meant to run as a server.
211
+
212
+ [Arize Phoenix](https://github.com/Arize-ai/phoenix) is the closest neighbor here. It runs locally
213
+ against SQLite with no account needed, and it's strong on tracing and evals, but it runs as a
214
+ server process with a fairly large dependency set, and it doesn't focus on waste detection, a
215
+ cost-weighted flamegraph, deterministic replay, or run diffs.
216
+
217
+ MLflow has been adding GenAI tracing, trace comparison, and efficiency scoring to its tracking
218
+ server; see [mlflow/mlflow](https://github.com/mlflow/mlflow).
219
+
220
+ [LangSmith](https://smith.langchain.com) is LangChain's hosted, proprietary platform.
221
+ [OpenLLMetry](https://github.com/traceloop/openllmetry) takes a different approach: it instruments
222
+ your code and exports OpenTelemetry spans to whatever backend you choose to point it at, rather
223
+ than shipping a backend of its own.
224
+
225
+ [AgentSight](https://github.com/eunomia-bpf/agentsight) renders token flamegraphs of coding agents
226
+ from the system side using eBPF, a genuinely different vantage point, though it's Linux only.
227
+
228
+ Helicone ([Helicone/helicone](https://github.com/Helicone/helicone)), OpenLIT
229
+ ([openlit/openlit](https://github.com/openlit/openlit)), Braintrust
230
+ ([braintrust.dev](https://braintrust.dev)), and Logfire
231
+ ([pydantic.dev/logfire](https://pydantic.dev/logfire)) each pair instrumentation with a server or a
232
+ cloud backend of their own.
233
+
234
+ traceburn's own position is narrower than most of the above: strictly local, one file, no account
235
+ and no server needed for basic use, framework-agnostic at the SDK level, and built around
236
+ efficiency first, meaning the waste report, the dollar-weighted flamegraph, replay, and diff all
237
+ live together in one small package. It's meant to sit next to whatever observability stack you
238
+ already run, not replace it.
239
+
240
+ ## Contributing
241
+
242
+ There are two extension points, each documented and each meant to be roughly an afternoon of work:
243
+ an [instrumentation adapter](https://github.com/TommyTranX/traceburn/blob/main/docs/instrumentation.md)
244
+ for a new SDK or framework, and a
245
+ [waste rule](https://github.com/TommyTranX/traceburn/blob/main/docs/waste-rules.md) for a new
246
+ pattern of avoidable spend. The full guide is at
247
+ [CONTRIBUTING.md](https://github.com/TommyTranX/traceburn/blob/main/CONTRIBUTING.md).
248
+
249
+ ## Citation
250
+
251
+ A citation file is included at
252
+ [CITATION.cff](https://github.com/TommyTranX/traceburn/blob/main/CITATION.cff).
253
+
254
+ ## License
255
+
256
+ MIT. Full text at
257
+ [LICENSE](https://github.com/TommyTranX/traceburn/blob/main/LICENSE).
258
+
259
+ Written by Tommy Tran.
@@ -0,0 +1,233 @@
1
+ # traceburn
2
+
3
+ ## The finding that made me build this
4
+
5
+ I wrote a small support ticket triage agent: five tickets, one tool call each, a roughly 4,700
6
+ token static policy prompt sent fresh on every call, model claude-haiku-4-5. Running it uncached
7
+ cost $0.0539. traceburn's waste report looked at the trace, noticed that same 4,700 token prefix
8
+ going out uncached on all 10 calls, and estimated that about 82 percent of that spend was
9
+ avoidable. So I added exactly the one cache_control block it suggested and reran the same five
10
+ tickets: $0.0167.
11
+
12
+ That's a 69 percent measured saving. The tool's estimate landed within 18 percent of what actually
13
+ happened, close enough to trust as a first signal, not close enough to treat as gospel. That's
14
+ roughly how I want a cost estimator to behave.
15
+
16
+ Prices change and models get repriced, so do not take my word for it: the reproduction is a few
17
+ cents and a couple of minutes, at
18
+ [examples/cache_before_after.py](https://github.com/TommyTranX/traceburn/blob/main/examples/cache_before_after.py).
19
+ Measured 2026-07-05.
20
+
21
+ ![traceburn's waste report on the uncached run: $0.0539 total, about 82 percent flagged avoidable, with the repeated 5,618-token prefix identified as the cause](https://raw.githubusercontent.com/TommyTranX/traceburn/main/assets/waste-report.png)
22
+
23
+ That's the whole pitch in one run. traceburn is a local-first tracer and efficiency profiler for
24
+ AI agents, built on top of the openai and anthropic Python SDKs: a cost and latency flamegraph, a
25
+ waste report that quantifies avoidable spend instead of just gesturing at it, deterministic replay
26
+ of recorded calls, and run diffs, all backed by one SQLite file on disk. No account, no server to
27
+ stand up for basic use, no telemetry leaving your machine.
28
+
29
+ ## Install
30
+
31
+ Core tracer, stdlib only, zero dependencies:
32
+
33
+ ```bash
34
+ pip install traceburn
35
+ ```
36
+
37
+ Tracer plus the local web viewer (adds starlette and uvicorn):
38
+
39
+ ```bash
40
+ pip install "traceburn[ui]"
41
+ ```
42
+
43
+ Requires Python 3.10 or later.
44
+
45
+ ## Quickstart
46
+
47
+ Patch the SDKs at the top of your agent:
48
+
49
+ ```python
50
+ import traceburn
51
+ traceburn.install()
52
+
53
+ # your existing openai / anthropic code, unchanged
54
+ ```
55
+
56
+ Every sync call, async call, streaming response, tool call, and prompt cache hit on either SDK now
57
+ gets recorded as a span. Traces live in one SQLite file, `./.traceburn/traces.db` by default; set
58
+ the `TRACEBURN_DB` environment variable if you want it somewhere else. If you'd rather not touch
59
+ the source at all, wrap the run instead:
60
+
61
+ ```bash
62
+ TRACEBURN=1 python your_agent.py
63
+ ```
64
+
65
+ Then look at what happened:
66
+
67
+ ```bash
68
+ traceburn ui # opens the web viewer at 127.0.0.1:8765
69
+ traceburn ls # list recorded traces
70
+ traceburn show <id> # inspect one trace
71
+ traceburn waste <id> # run the waste report on one trace
72
+ traceburn diff <a> <b> # compare two traces span by span
73
+ ```
74
+
75
+ No API keys and nothing to configure: `python examples/offline_demo.py` records a simulated agent
76
+ run with realistic token counts, including one deliberate duplicate call, so you can see a real
77
+ trace, a real flamegraph, and a real waste finding inside a minute.
78
+
79
+ ## What it actually does
80
+
81
+ **Trace.** `traceburn.install()` patches both SDKs so every call becomes a span with tokens,
82
+ latency, and cost attached, no code changes required past that one line. Want manual control
83
+ instead, or you're using a framework outside the two supported SDKs? The explicit API, `@trace`,
84
+ `span()`, and `session()`, works by hand with anything.
85
+
86
+ ![traceburn's expandable trace tree, showing an agent's nested spans with per-call tokens and cost](https://raw.githubusercontent.com/TommyTranX/traceburn/main/assets/trace-tree.png)
87
+
88
+ **Flamegraph.** Spans render as a flamegraph you can size two ways: by wall-clock time or by
89
+ dollars spent, with self-time kept separate from time spent in children, so a slow parent span
90
+ doesn't hide which child call actually burned the seconds or the money.
91
+
92
+ **Waste report.** Heuristics look for duplicate calls, unused cache opportunities, bloated
93
+ prompts, model overkill, and retry loops. Each finding ships with a confidence level and, where the
94
+ numbers support it, a dollar figure. More on this below.
95
+
96
+ **Replay.** `traceburn.analyze.replay.replay()` plays recorded provider responses back through the
97
+ real SDK types, so your agent code runs again exactly as before with zero tokens spent. That's
98
+ useful for tests and for debugging without burning a budget. Streaming calls aren't replayable
99
+ yet; replay currently only serves non-streaming recordings.
100
+
101
+ **Diff.** `traceburn diff <a> <b>` lines up two traces span by span and reports the delta in
102
+ tokens, cost, and latency, alongside text diffs of the prompts and responses that changed between
103
+ the runs. Good for answering "did that prompt tweak actually help."
104
+
105
+ ## The web viewer
106
+
107
+ `traceburn ui` starts a small, read-only Starlette API in front of a vendored single-page app: no
108
+ build step, no CDN, everything ships in the package. It gives you the trace tree, the flamegraph,
109
+ a waterfall timeline, the waste report, and the run diff view.
110
+
111
+ It binds to 127.0.0.1 only and checks the request's host header against DNS rebinding, but it has
112
+ no authentication of any kind. That's a deliberate tradeoff: the viewer isn't meant to be reachable
113
+ from anywhere but your own machine.
114
+
115
+ ![traceburn's flamegraph view of an agent run, one row per depth, frame width proportional to latency](https://raw.githubusercontent.com/TommyTranX/traceburn/main/assets/flamegraph.png)
116
+
117
+ ![traceburn's waterfall view of the same run, a timeline of every call with its duration and cost](https://raw.githubusercontent.com/TommyTranX/traceburn/main/assets/waterfall.png)
118
+
119
+ ## Framework support
120
+
121
+ Today, that means the raw `openai` and `anthropic` Python SDKs, patched automatically by
122
+ `traceburn.install()`. If you're on something else, the explicit `span()` / `trace()` / `session()`
123
+ API works with any framework right now, by hand, since it doesn't care what's making the call.
124
+ LangChain, LlamaIndex, and anything already emitting OpenTelemetry GenAI spans aren't instrumented
125
+ automatically yet. That's real, planned work for v0.2, not something already built and just
126
+ undocumented, and it's covered in the roadmap below.
127
+
128
+ ## How the waste rules work
129
+
130
+ The rules live in `traceburn/analyze/waste/` and are documented in full at
131
+ [docs/waste-rules.md](https://github.com/TommyTranX/traceburn/blob/main/docs/waste-rules.md). Five
132
+ kinds of waste get checked for: duplicates (exact and near-duplicate repeated calls), cache (a
133
+ stable prompt prefix resent without ever hitting a provider cache), context_bloat (duplicate
134
+ blocks inside one prompt, or a huge prompt for a tiny output), model_overkill (a frontier-priced
135
+ model spent on trivial short calls, phrased as a suggestion and kept at low confidence on purpose),
136
+ and loops (retry storms, repeated identical tool calls, or a runaway step count). Every finding
137
+ also carries a confidence level: high, medium, low, or info.
138
+
139
+ Two principles govern all of them. A wrong finding does more damage than a missed one, so the
140
+ rules are tuned for precision over recall and would rather stay quiet than guess. And no dollar
141
+ figure is ever printed without observed tokens behind it and a price in the pricing table to
142
+ multiply against; everything the tool prints is labeled an estimate, because it is one.
143
+
144
+ ## Privacy
145
+
146
+ traceburn makes no network calls of its own and sends no telemetry anywhere. The only traffic on
147
+ the wire is your own calls to your own model provider, exactly as they'd happen without traceburn
148
+ installed. A test,
149
+ [tests/test_no_network.py](https://github.com/TommyTranX/traceburn/blob/main/tests/test_no_network.py),
150
+ blocks all socket access at the interpreter level and then runs the recorder, the store, every
151
+ analyzer, and the CLI against that blockade, to prove the point rather than just assert it. Auth
152
+ headers are never recorded in a trace, so an API key cannot end up sitting in a stored span.
153
+
154
+ ## Limitations
155
+
156
+ Instrumentation currently covers only the `openai` and `anthropic` Python SDKs. Within those,
157
+ `parse()` convenience methods and `with_raw_response` calls pass through untraced rather than being
158
+ recorded incorrectly, and multi-choice requests (`n > 1`) only record the first choice.
159
+
160
+ Cost figures come from a dated public pricing table
161
+ ([pricing.json](https://github.com/TommyTranX/traceburn/blob/main/src/traceburn/pricing.json)) and
162
+ don't model long-context pricing tiers or regional surcharges. Token counts prefer whatever the
163
+ provider itself reports as usage; anything estimated is flagged as estimated rather than presented
164
+ as measured.
165
+
166
+ The waste rules are heuristics tuned for precision over recall, which means they'll miss real
167
+ waste sooner than they'll invent fake waste, and every finding states its own confidence so you
168
+ can judge it accordingly. Streaming calls aren't replayable yet; only non-streaming recordings are.
169
+
170
+ The web viewer is read-only, bound to 127.0.0.1 only, and has no authentication.
171
+
172
+ ## Roadmap: v0.2
173
+
174
+ - OpenTelemetry GenAI span ingest plus OTLP export. This is also the path for capturing LangChain
175
+ and LlamaIndex traces, since it rides on their existing OTel instrumentation rather than
176
+ requiring bespoke adapters for each.
177
+ - A pytest plugin built on replay, for deterministic, token-free agent tests.
178
+ - litellm instrumentation.
179
+ - More waste rules.
180
+
181
+ ## Related work
182
+
183
+ [Langfuse](https://github.com/langfuse/langfuse) is a full open-source LLM platform: tracing,
184
+ evals, and prompt management, backed by Postgres and ClickHouse and meant to run as a server.
185
+
186
+ [Arize Phoenix](https://github.com/Arize-ai/phoenix) is the closest neighbor here. It runs locally
187
+ against SQLite with no account needed, and it's strong on tracing and evals, but it runs as a
188
+ server process with a fairly large dependency set, and it doesn't focus on waste detection, a
189
+ cost-weighted flamegraph, deterministic replay, or run diffs.
190
+
191
+ MLflow has been adding GenAI tracing, trace comparison, and efficiency scoring to its tracking
192
+ server; see [mlflow/mlflow](https://github.com/mlflow/mlflow).
193
+
194
+ [LangSmith](https://smith.langchain.com) is LangChain's hosted, proprietary platform.
195
+ [OpenLLMetry](https://github.com/traceloop/openllmetry) takes a different approach: it instruments
196
+ your code and exports OpenTelemetry spans to whatever backend you choose to point it at, rather
197
+ than shipping a backend of its own.
198
+
199
+ [AgentSight](https://github.com/eunomia-bpf/agentsight) renders token flamegraphs of coding agents
200
+ from the system side using eBPF, a genuinely different vantage point, though it's Linux only.
201
+
202
+ Helicone ([Helicone/helicone](https://github.com/Helicone/helicone)), OpenLIT
203
+ ([openlit/openlit](https://github.com/openlit/openlit)), Braintrust
204
+ ([braintrust.dev](https://braintrust.dev)), and Logfire
205
+ ([pydantic.dev/logfire](https://pydantic.dev/logfire)) each pair instrumentation with a server or a
206
+ cloud backend of their own.
207
+
208
+ traceburn's own position is narrower than most of the above: strictly local, one file, no account
209
+ and no server needed for basic use, framework-agnostic at the SDK level, and built around
210
+ efficiency first, meaning the waste report, the dollar-weighted flamegraph, replay, and diff all
211
+ live together in one small package. It's meant to sit next to whatever observability stack you
212
+ already run, not replace it.
213
+
214
+ ## Contributing
215
+
216
+ There are two extension points, each documented and each meant to be roughly an afternoon of work:
217
+ an [instrumentation adapter](https://github.com/TommyTranX/traceburn/blob/main/docs/instrumentation.md)
218
+ for a new SDK or framework, and a
219
+ [waste rule](https://github.com/TommyTranX/traceburn/blob/main/docs/waste-rules.md) for a new
220
+ pattern of avoidable spend. The full guide is at
221
+ [CONTRIBUTING.md](https://github.com/TommyTranX/traceburn/blob/main/CONTRIBUTING.md).
222
+
223
+ ## Citation
224
+
225
+ A citation file is included at
226
+ [CITATION.cff](https://github.com/TommyTranX/traceburn/blob/main/CITATION.cff).
227
+
228
+ ## License
229
+
230
+ MIT. Full text at
231
+ [LICENSE](https://github.com/TommyTranX/traceburn/blob/main/LICENSE).
232
+
233
+ Written by Tommy Tran.
Binary file
Binary file
Binary file
Binary file
@@ -0,0 +1,57 @@
1
+ # Concepts
2
+
3
+ ## The data model
4
+
5
+ Three tables, plus content-addressed blobs, in one SQLite file:
6
+
7
+ - A **session** groups related traces (a batch run, a test suite pass).
8
+ - A **trace** is one run: everything under one root span.
9
+ - A **span** is one timed operation with a `kind`: `llm`, `tool`, `agent`,
10
+ `retrieval`, or `custom`. Spans nest via `parent_id`.
11
+
12
+ LLM spans follow the OpenTelemetry GenAI attribute names where they exist
13
+ (`gen_ai.system`, `gen_ai.request.model`, `gen_ai.usage.input_tokens`, and
14
+ so on) plus traceburn extensions: cache token splits, estimated cost, the
15
+ normalized request and response, and a `request_hash` used by replay and
16
+ duplicate detection. The full conventions live in `schema.py` and are a
17
+ public, versioned interface.
18
+
19
+ ## Token accounting
20
+
21
+ `gen_ai.usage.input_tokens` counts only tokens billed at the full input
22
+ rate. Cache reads (`cached_input_tokens`) and cache writes
23
+ (`cache_write_tokens`) are separate, so cost math stays auditable:
24
+
25
+ ```
26
+ cost = input * input_rate + cached * cached_rate
27
+ + cache_write * write_rate + output * output_rate
28
+ ```
29
+
30
+ Rates come from a dated, user-overridable pricing table
31
+ (`pricing.json`, `TRACEBURN_PRICING`). When a provider omits usage (some
32
+ streaming shapes), counts are estimated and the span carries
33
+ `usage_estimated: true`.
34
+
35
+ ## Storage
36
+
37
+ WAL mode, so the viewer reads while your agent writes. Large request and
38
+ response payloads are stored once in a blobs table, keyed by content hash;
39
+ identical payloads across calls cost one row, which is also what makes
40
+ duplicate detection and replay lookups cheap. The schema is versioned via
41
+ a meta row.
42
+
43
+ ## Write path guarantees
44
+
45
+ A failure inside traceburn never breaks your agent: store writes are
46
+ guarded, patcher capture errors degrade to an untraced call, and spans for
47
+ abandoned streams finalize at garbage collection. Exceptions from your own
48
+ code always propagate; the span just records the error first.
49
+
50
+ ## The analyzers
51
+
52
+ Everything in `traceburn/analyze/` reads the store and computes:
53
+ flamegraph folds (latency- or cost-weighted), waterfall timelines, run
54
+ diffs, deterministic replay, and the waste report. None of it makes a
55
+ network call. The CLI and the web viewer are two skins over the same
56
+ functions, and every JSON shape they emit is a public interface you can
57
+ consume without importing the analyzers.