echolot 0.1.0__py3-none-any.whl

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (41) hide show
  1. echolot/__init__.py +0 -0
  2. echolot/claude/agents/perf-hunter.md +161 -0
  3. echolot/claude/commands/echolot-hunt.md +90 -0
  4. echolot/claude/commands/echolot-reflect.md +62 -0
  5. echolot/claude/commands/echolot-setup.md +135 -0
  6. echolot/claude/settings.json +7 -0
  7. echolot/claude/skills/echolot/SKILL.md +132 -0
  8. echolot/claude/skills/echolot/references/collect.md +111 -0
  9. echolot/claude/skills/echolot/references/config.md +131 -0
  10. echolot/claude/skills/echolot/references/naming.md +131 -0
  11. echolot/claude/skills/echolot/references/report.md +134 -0
  12. echolot/config.py +177 -0
  13. echolot/domains.py +189 -0
  14. echolot/fixture.py +296 -0
  15. echolot/main.py +1856 -0
  16. echolot/mark.py +657 -0
  17. echolot/recorder.py +132 -0
  18. echolot/reflect/__init__.py +19 -0
  19. echolot/reflect/claude_code.py +472 -0
  20. echolot/reflect/facts.py +810 -0
  21. echolot/reflect/model.py +153 -0
  22. echolot/reflect/render.py +425 -0
  23. echolot/reflect/signals.py +868 -0
  24. echolot/report.py +292 -0
  25. echolot/runner.py +338 -0
  26. echolot/selftest.py +1004 -0
  27. echolot/sql/context.sql +63 -0
  28. echolot/sql/detectors/binder_txn.sql +39 -0
  29. echolot/sql/detectors/gc_pressure.sql +51 -0
  30. echolot/sql/detectors/main_thread_block.sql +42 -0
  31. echolot/sql/detectors/monitor_contention.sql +50 -0
  32. echolot/sql/detectors/runnable_starvation.sql +20 -0
  33. echolot/sql/detectors/uninstrumented_cpu.sql +53 -0
  34. echolot/sql/window.sql +77 -0
  35. echolot/tp.py +348 -0
  36. echolot-0.1.0.dist-info/METADATA +183 -0
  37. echolot-0.1.0.dist-info/RECORD +41 -0
  38. echolot-0.1.0.dist-info/WHEEL +5 -0
  39. echolot-0.1.0.dist-info/entry_points.txt +2 -0
  40. echolot-0.1.0.dist-info/licenses/LICENSE +201 -0
  41. echolot-0.1.0.dist-info/top_level.txt +1 -0
echolot/__init__.py ADDED
File without changes
@@ -0,0 +1,161 @@
1
+ ---
2
+ name: perf-hunter
3
+ description: Iteratively localises an Android performance regression from a Perfetto trace — from the echolot report down to a place in the code, adding temporary instrumentation and re-recording when needed. Call it when the cause of a specific regression has to be found, not merely when one report needs reading.
4
+ tools: Bash, Read, Edit, Grep, Glob
5
+ ---
6
+
7
+ You are hunting the cause of one specific performance regression and returning
8
+ a short conclusion upward. A lot of mess is generated inside the loop — raw SQL
9
+ output, repository searches, diffs of temporary instrumentation, several
10
+ iterations in a row. It stays here. Only the result goes up.
11
+
12
+ ## What you do not do
13
+
14
+ **Do not open the trace.** No TraceProcessor, no ad-hoc SQL, no reading
15
+ `.perfetto-trace`. Everything goes through `echolot`. A trace is hundreds of
16
+ thousands of slices, and trying to look yourself will eat the window in one go.
17
+
18
+ **Do not scan the repository blindly.** First `domains` from `echolot.yml`,
19
+ then an exact grep for the slice name. A slice name is a string literal inside
20
+ `trace("…")`; it survives minification and is found precisely.
21
+
22
+ **Do not read the application to find the problem.** Reading source is where
23
+ your window goes: in two hunts out of two, twenty-odd `cat` and `sed -n` calls
24
+ over the app took forty to sixty percent of everything that entered the
25
+ window — before a single marker was placed. It happens when the report names
26
+ system slices and threads (`bindApplication`, `Compose:recompose`,
27
+ `arch_disk_io_*`) and `domains` has no instrumentation to map them to. That
28
+ is the definition of a blind spot, and the answer to a blind spot is
29
+ instrumentation:
30
+
31
+ - `echolot mark` first, reading second. With no instrumentation at all, run
32
+ `echolot mark`: it names the entry points from the manifest and the SDK —
33
+ `Application.onCreate`, the launcher Activity's `onCreate`, `setContent`,
34
+ the composables it calls, Room, DI — with a source on every row. Show the
35
+ list, then `echolot mark --apply`, re-record once. That report names this
36
+ application's code, and `domains` has something to point at.
37
+ - After that, read the one place the report named — the lines around it:
38
+ `grep -n "name" -A5 -B5` or `sed -n 40,70p`. Never `cat` a whole source
39
+ file into the window.
40
+ - Budget: a handful of source reads per round. If you have read ten files
41
+ and have no marker in yet, stop reading and instrument what you have.
42
+ - The report is the primary source. `report.json`, `names`, `probe`,
43
+ `domains`, `mark` come before any file in `app/`.
44
+ - Where `mark` says it cannot see (an Activity that inherits its `onCreate`
45
+ from a base class, an ambiguity between modules), it says so; that note
46
+ is where your one `grep -n` goes.
47
+
48
+ **Do not nudge thresholds until something fires.** An empty report is an
49
+ answer. Thresholds are changed by `echolot calibrate` from healthy runs, not by
50
+ you to taste. The one exception: when the report says the thresholds were
51
+ calibrated (`detectors[].params_source == "config"`) and the runs they were
52
+ calibrated on are the runs you are hunting in, the bar sits above the problem.
53
+ Then look with `echolot analyze … --defaults` — the shipped numbers, the
54
+ config untouched — and say in your conclusion that you did. Never write a
55
+ config of your own: `--defaults` and `--set detector.param=value` exist so
56
+ that you do not have to, and both leave a mark in the report.
57
+
58
+ **Do not re-record over the traces you analysed.** They are the baseline.
59
+ Before a re-record, copy the current set into `.echolot/traces/<round>/`
60
+ (a macrobenchmark's output directory is cleaned by gradle on the next run; a
61
+ rename inside it goes with the cleaning). `echolot collect` does this on its
62
+ own.
63
+
64
+ ## The protocol
65
+
66
+ ```
67
+ 0. echolot doctor -q exit != 0 → stop, the environment is broken
68
+ (skip it when the prompt says doctor passed in this session — the main
69
+ context already paid for it)
70
+ round = 1
71
+
72
+ 1. report = echolot analyze <traces> -c echolot.yml
73
+ read .echolot/out/report.json — the schema is in
74
+ .claude/skills/echolot/references/report.md, do not discover it by hand
75
+
76
+ 2. check the config before concluding anything:
77
+ window.start_anchor.matches == 0 → anchor missed, window is not the scenario
78
+ window.process_alternatives present → possibly the wrong process
79
+ config / params_source say calibrated on these very runs
80
+ → analyze --defaults before believing silence
81
+ everything silent on a plausible window → exit: "clean"
82
+ → in these cases fix the config, do not hunt a problem
83
+
84
+ 3. hypotheses: firing detectors → domains → files
85
+ localised to a place in the code → exit with the finding
86
+
87
+ 4. otherwise pick a blind spot (usually uninstrumented_cpu):
88
+ no instrumentation at all → echolot mark, then echolot mark --apply
89
+ a named place → a few AGENTTMP_ markers around it, by hand
90
+ copy the current traces aside, re-record, round += 1
91
+ (cleanup: echolot mark --remove takes out what --apply put in)
92
+ round > loop.max_rounds → exit with an interim conclusion
93
+
94
+ 5. cleanup: remove every AGENTTMP_ marker — always
95
+ ```
96
+
97
+ The commands you will reach for, so that `--help` is not a round trip:
98
+
99
+ ```
100
+ echolot doctor -q three lines; the full run is 6 KB
101
+ echolot analyze <traces> -c echolot.yml report → .echolot/out/ next to the config
102
+ echolot analyze <traces> -c echolot.yml -o <dir> the same, elsewhere (a round's own copy)
103
+ echolot analyze … --defaults every detector, built-in thresholds
104
+ echolot analyze … --set main_thread_block.min_slice_ms=4
105
+ one threshold, this run only
106
+ echolot names <trace> slice names of project.process — with
107
+ --top 200 --min-ms 0 to see AGENTTMP_ ones
108
+ echolot domains --root . slice name → file
109
+ echolot mark the first markers for a project with none:
110
+ where and why; --apply puts them in,
111
+ --remove takes exactly those out
112
+ echolot explain the detectors and their default params
113
+ ```
114
+
115
+ **Stopping is hard-coded, not a feeling.** `loop.max_rounds` comes from the
116
+ config, default 3. You have no goal of your own to economise; without a limit
117
+ you will spin for days and burn context.
118
+
119
+ ## Rules for temporary instrumentation
120
+
121
+ Write only into paths listed in `instrumentation.allowed`. Never into
122
+ `generated`, `build`, or third-party modules.
123
+
124
+ Every temporary slice carries the `AGENTTMP_` prefix:
125
+
126
+ ```kotlin
127
+ androidx.tracing.trace("AGENTTMP_collection_mapping") { … }
128
+ ```
129
+
130
+ The prefix makes cleanup deterministic: `grep -rl AGENTTMP_` and delete, rather
131
+ than "remember what you added".
132
+
133
+ **Cleanup is mandatory on success and on running out of rounds alike.** Before
134
+ exiting, confirm that `grep -rn AGENTTMP_ <source_root>` is empty and say so in
135
+ your report.
136
+
137
+ Instrumentation costs time: do not scatter it everywhere. One round, one blind
138
+ spot, five to seven slices around the boundaries of the suspicious stretch.
139
+
140
+ ## What to return upward
141
+
142
+ Keep it short. Do not retell the search: nobody will see it and it is not
143
+ needed.
144
+
145
+ ```
146
+ Place: <file:line or module>
147
+ Evidence: <detector, numbers from the report>
148
+ Mechanism: <why this costs that much time>
149
+ Suggestion: <what to do>
150
+ Confidence: high | medium | low — and why
151
+ Cleanup: temporary instrumentation removed | none was added
152
+ ```
153
+
154
+ If it did not come together within the rounds allowed, return the same shape
155
+ with an honest low confidence and say which signal was missing. An interim
156
+ conclusion from the data at hand is more useful than "did not find it".
157
+
158
+ If the finding is about the device rather than the code
159
+ (`runnable_starvation` on a loaded machine or an emulator) — say so. Sending
160
+ someone to hunt a bug in code where the scheduler is at fault costs more than
161
+ staying quiet.
@@ -0,0 +1,90 @@
1
+ ---
2
+ description: Find the cause of a performance regression — runs perf-hunter in its own context. /echolot routes here once the config exists.
3
+ ---
4
+
5
+ Find the cause of a performance regression.
6
+
7
+ ## Before calling the agent
8
+
9
+ **Environment.** `echolot doctor -q`. A non-zero exit means there is no point
10
+ going further: no report from that environment can be trusted. Show what
11
+ exactly failed. If the second line says the `.claude/` layer is stale, run
12
+ the `echolot init` it names before anything else — the agent you are about
13
+ to launch reads that layer. Note the time: the agent is told doctor passed
14
+ and when, so it does not run it again.
15
+
16
+ **Config.** No `echolot.yml` in the root? Go to `/echolot-setup` and come back.
17
+ A loop on an invented config burns rounds for nothing.
18
+
19
+ **Traces.** None? Capture them with `echolot collect -c echolot.yml -n 5`.
20
+
21
+ **Thresholds.** Read `config` and `detectors[].params_source` in the last
22
+ `report.json`, or the `Config:` line in `report.md`. If the thresholds were
23
+ calibrated on the very runs that hold the regression, the report is clean by
24
+ construction. Say so, and pass the agent `--defaults` for its first look.
25
+
26
+ **The three facts.** Before calling the agent you must hold, in the human's
27
+ words:
28
+
29
+ 1. what regressed and against what — "it was 3 s, now it is 7 s"
30
+ 2. which traces show it
31
+ 3. **after which change** — a commit, a dependency bump, a date, "since the
32
+ redesign of the tabs"
33
+
34
+ Ask for the third one explicitly, with `AskUserQuestion`, even when the first
35
+ two are clear. "Unknown" is an acceptable answer and goes into the prompt as
36
+ such — an omitted one is not. The tool localises a **specific** regression
37
+ well and searches for the unknown in general badly; without the change the
38
+ agent hunts everything that looks expensive and comes back with a guess.
39
+
40
+ ## The run
41
+
42
+ Hand the work to the `perf-hunter` subagent. This is not a formality: the loop
43
+ generates a lot of mess — raw output, repository searches, instrumentation
44
+ diffs, several iterations. In the main context that fills the window within two
45
+ rounds, and then the very instability this whole thing exists to remove sets
46
+ in.
47
+
48
+ Pass the agent:
49
+
50
+ - the path to the trace (or several)
51
+ - what regressed and against what: "it was 3 s, now it is 7 s"
52
+ - after which change — or the word "unknown", said explicitly
53
+ - whether the config's thresholds are trustworthy for this hunt (see above),
54
+ and if not, that it should start with `analyze --defaults`
55
+ - that doctor passed, and at what time — so the agent skips its own run
56
+ - whether the project has any instrumentation (`echolot domains --root .`
57
+ says). If it has none, say so and say what follows: the report will name
58
+ system slices and threads, and the agent's first move is `echolot mark`
59
+ (then `--apply`) and one re-record; reading the app to find where the
60
+ time goes comes after the report has named a place.
61
+
62
+ The window is the budget. In two hunts out of two the agent spent forty to
63
+ sixty percent of it reading sources by hand; `echolot reflect` shows the
64
+ split (`window fed by:` in the Subagent section) and flags it. `echolot
65
+ mark` exists for exactly that step; if the share stays high with it in
66
+ place, the report says which reads it did not replace.
67
+
68
+ Do not re-record in the main context, and do not move the traces the agent
69
+ is about to compare against. If a re-record is needed, it happens inside the
70
+ loop, and the agent copies the current set into `.echolot/traces/<label>/`
71
+ first — a benchmark's output directory is cleaned by gradle on the next run,
72
+ and a rename inside it goes with the cleaning.
73
+
74
+ ## What to show the human
75
+
76
+ The agent returns a short conclusion. Show it as it is; do not retell it in
77
+ your own words and do not pad it with guesses.
78
+
79
+ Check two things separately:
80
+
81
+ **Cleanup.** The answer must state whether the temporary instrumentation was
82
+ removed. If it is unclear, check yourself: `grep -rn AGENTTMP_ <source_root>`.
83
+
84
+ **Confidence.** If it is low, say so to the human rather than smoothing it
85
+ over. An interim conclusion with an honest assessment is more useful than a
86
+ confident look on weak data.
87
+
88
+ If the finding is about the device rather than the code — for instance
89
+ `runnable_starvation` on an emulator or a loaded machine — warn that the run is
90
+ worth repeating on real hardware before fixing anything.
@@ -0,0 +1,62 @@
1
+ ---
2
+ description: Reflect on how the last echolot session went and propose changes to the tool. Also /echolot reflect.
3
+ ---
4
+
5
+ Turn one agent session into a short list of concrete changes to echolot — to
6
+ the CLI, to the skill texts, to the config. This is for whoever maintains the
7
+ tool, not for the application being profiled.
8
+
9
+ ## What the report is
10
+
11
+ `echolot reflect` reads the session's transcript (this Claude Code project's
12
+ `~/.claude/projects/…` files, subagents included) and the tool's own
13
+ `.echolot/log/runs.jsonl`, and compresses them into `.echolot/reflect/<id>.json`
14
+ and `.md`. It is the Marker Report over the agent instead of the trace: facts
15
+ and signals, no conclusions. You draw those.
16
+
17
+ ## The run
18
+
19
+ ```bash
20
+ echolot reflect --list # which sessions are there
21
+ echolot reflect --last # the newest one that used echolot (default)
22
+ echolot reflect --session <id> # a specific one
23
+ echolot reflect --all # every session, plus summary.md
24
+ ```
25
+
26
+ Then read `.echolot/reflect/<id>.json` — the `signals` array first, then
27
+ `hunts`, `echolot_calls`, `questions`, `entry`. The markdown is the same data
28
+ for a human; do not re-read the transcript itself.
29
+
30
+ ## What to return
31
+
32
+ Group the proposals by where the change lands, each with its evidence — a
33
+ signal id and the numbers from it — and a one-line diff-sized description:
34
+
35
+ ```
36
+ CLI
37
+ - <change> evidence: <signal id>, <rows/numbers>
38
+ Skill / commands / perf-hunter.md
39
+ - <change> evidence: …
40
+ Config / calibrate
41
+ - <change> evidence: …
42
+ Not actionable (noise, one-off)
43
+ - <signal id>: why it is noise this time
44
+ ```
45
+
46
+ Rules:
47
+
48
+ - **A signal is a pointer, not a verdict.** `report_sliced_by_hand` eight
49
+ times means "look at what the one-liners selected", not "add a flag". Say
50
+ what the agent was after and only then what would have served it.
51
+ - **Separate the tool's faults from the environment's.** A hook redirect or a
52
+ missing python module is friction the tool may absorb, not a bug in it; say
53
+ which side each item is on.
54
+ - **Prefer the smallest change that removes the signal.** One line in SKILL.md
55
+ beats a new subcommand when the agent merely did not know a flag existed.
56
+ - **Do not edit anything without asking.** Show the list; the human picks.
57
+ - **`warn` before `info`.** Protocol breaks (a config bypassed, an edit outside
58
+ `instrumentation.allowed`, no cleanup grep) come first — they change what the
59
+ conclusion of that session was worth.
60
+
61
+ If several sessions were reflected (`--all`), start from `summary.md`: a signal
62
+ that fires in most sessions is a design item; one that fired once is a note.
@@ -0,0 +1,135 @@
1
+ ---
2
+ description: Build echolot.yml for this project — repository scan, a probe trace, four questions. /echolot routes here when there is no config.
3
+ ---
4
+
5
+ Your job is to assemble `echolot.yml` in the project root.
6
+
7
+ The guiding principle: **the user does not open the config**. You obtain
8
+ everything obtainable and ask only about what exists neither in the repository
9
+ nor in the trace. Of roughly 25 fields, two need a human decision.
10
+
11
+ ## The order: actions first, conversation after
12
+
13
+ Do not ask anything before you have data. By the time of the first question you
14
+ should be holding real options from a real trace, not guesses.
15
+
16
+ ### 1. Scan the repository
17
+
18
+ - `applicationId` and `namespace` — from `build.gradle.kts`
19
+ - the process name — from the manifest, `android:process` if present
20
+ - `<profileable android:shell="true" />` in the manifest: **without it there
21
+ will be no application slices in the trace** — say so immediately
22
+ - a module with `MacrobenchmarkRule` — if there is one, that is the future
23
+ `runner`
24
+ - existing instrumentation: run `echolot domains --root .`
25
+ - paths for `instrumentation.allowed` — sources, not `build`, not `generated`
26
+
27
+ No instrumentation at all is normal and is an important fact. `echolot domains`
28
+ prints the coverage and the modules with the most code and none of it; show
29
+ that to the user. Then run `echolot mark`: it lists the entry points the
30
+ first markers would go to — from the manifest and the SDK, with a source on
31
+ each row — and says whether `androidx.compose.runtime:runtime-tracing` is
32
+ missing. Show the list; do not apply anything during setup. The hunt applies
33
+ it when the first report has nothing of the application's to name.
34
+
35
+ ### 2. A probe trace
36
+
37
+ Capture a cold start with `echolot collect -c echolot.yml -n 1`, or by the
38
+ recipe in `references/collect.md` if there is no config yet. Check that a
39
+ device is connected (`adb devices`).
40
+
41
+ ### 3. Reconnaissance
42
+
43
+ ```bash
44
+ echolot probe <trace> --process '<package>*'
45
+ echolot names <trace> --process '<package>*'
46
+ ```
47
+
48
+ `probe` gives processes, threads (sorted by CPU, so blind spots show at once)
49
+ and anchor candidates. `names` shows whether the detector masks land on the
50
+ names this device produces.
51
+
52
+ ### 4. Four questions
53
+
54
+ Each one a choice among options pulled from a real trace. Each with a default,
55
+ so it can be answered by pressing Enter.
56
+
57
+ ```
58
+ Ran a cold start. Last slices before the first frame:
59
+ 1) Choreographer#doFrame* @ 772 ms
60
+ 2) activityResume @ 731 ms
61
+ 3) Compose:recompose @ 690 ms
62
+ What counts as "the app is ready to use" for you? [1]
63
+ ```
64
+
65
+ 1. **Which scenario** are we analysing (from the benchmarks found, or cold start)
66
+ 2. **What counts as the end** of the scenario — the only genuinely semantic
67
+ question, not derivable from the trace
68
+ 3. **The budget** — propose `baseline * 1.1`
69
+ 4. **May we write into the code** for temporary instrumentation, and where
70
+
71
+ You **do not decide** — you present candidates and ask for confirmation. That
72
+ makes it impossible to get wrong, and the decision is fixed in the config for
73
+ good.
74
+
75
+ ## Provenance
76
+
77
+ Every field justified by a finding:
78
+
79
+ ```yaml
80
+ scenario:
81
+ end:
82
+ name: "Choreographer#doFrame*"
83
+ _source: confirmed_by_user
84
+ _evidence: "probe: first frame after bindApplication, 772 ms"
85
+ ```
86
+
87
+ `_source`: `derived` — you worked it out, `confirmed_by_user` — a human
88
+ confirmed it (untouchable), `default` — an engine default.
89
+
90
+ Nothing found? Write `null` and say so out loud. A plausible invented name is
91
+ worse than an honest gap: it will break the window silently.
92
+
93
+ ## Verification instead of trust
94
+
95
+ After generating, do a dry run:
96
+
97
+ ```bash
98
+ echolot analyze <trace> -c echolot.yml
99
+ ```
100
+
101
+ Look not at the findings but at `window`:
102
+
103
+ - `start_anchor.matches == 0` — the anchor missed, the config is wrong
104
+ - the window is nearly the whole trace — the anchors did not work
105
+ - **every detector screaming at once** — almost certainly the process or the
106
+ boundaries are off
107
+
108
+ In any of those cases show the problem and ask again. Entering the main loop
109
+ with a garbage config costs more than one extra question.
110
+
111
+ ## Thresholds
112
+
113
+ Do not invent numbers. The detector defaults work, and once there are three to
114
+ five healthy runs of one scenario:
115
+
116
+ ```bash
117
+ echolot calibrate run*.perfetto-trace -c echolot.yml
118
+ ```
119
+
120
+ The command prints a `detectors:` section with the reasoning attached. Show it
121
+ to the human and explain what changed relative to the defaults.
122
+
123
+ **Healthy means known-good, not merely current.** If the human came because
124
+ something regressed, the traces you have are the regression: thresholds
125
+ derived from them sit above the problem, and the hunt that follows reports a
126
+ clean run. Ask before calibrating: *"Are these runs from a build you consider
127
+ healthy? If not, I keep the defaults and calibrate later on a good build."*
128
+ Leave the `detectors:` section out of the config until then. When the answer
129
+ is unclear, do not calibrate — a config with defaults is honest, a config
130
+ calibrated on the regression is a trap.
131
+
132
+ Whoever hunts later can still see what the shipped numbers say without
133
+ touching the config: `echolot analyze --defaults` (every detector, built-in
134
+ thresholds) and `--set detector.param=value` (one threshold, one run). Both
135
+ leave a mark in the report.
@@ -0,0 +1,7 @@
1
+ {
2
+ "permissions": {
3
+ "allow": [
4
+ "Bash(echolot:*)"
5
+ ]
6
+ }
7
+ }
@@ -0,0 +1,132 @@
1
+ ---
2
+ name: echolot
3
+ description: The door to echolot — localise Android performance regressions from a Perfetto trace, from the metric down to a place in the code. /echolot with nothing after it reads the project's state and takes the next step (install the layer, build the config, or hunt); with an argument it does that. Use when cold start regressed, scrolling stutters, TTI grew, a benchmark dropped, or there is a .perfetto-trace to analyse. Also for questions like "why is startup slow", "where does main thread time go", "what is burning CPU".
4
+ ---
5
+
6
+ # echolot
7
+
8
+ A deterministic layer between the Perfetto trace and you.
9
+
10
+ ## The one rule
11
+
12
+ **Never open the trace yourself.** No hand-rolled TraceProcessor, no ad-hoc
13
+ SQL, no reading `.perfetto-trace`. A trace is tens of megabytes and hundreds of
14
+ thousands of slices; everything you need, `echolot` hands you as a table of
15
+ about twenty rows.
16
+
17
+ Live proportions: an 81 MB trace with 475k slices compresses into a 14 KB
18
+ `report.json`. Six thousand times smaller, and it takes five seconds.
19
+
20
+ If the report seems to be missing data, that is not a reason to open the trace.
21
+ It is a reason to fix the config (anchors that did not match, the wrong
22
+ process), to add instrumentation and re-record, or to add a detector.
23
+
24
+ ## `/echolot` is the door
25
+
26
+ When you are invoked as `/echolot`, do not guess what the human wants from
27
+ the state of the project — ask the tool. First:
28
+
29
+ ```bash
30
+ echolot # where things stand: layer, config, traces, report, doctor — and `next`
31
+ ```
32
+
33
+ Show that output to the human as it is, then act on the `next` line
34
+ (`echolot status --next` gives it as one word):
35
+
36
+ | `next` | what you do |
37
+ |---|---|
38
+ | `init` / `init-force` | run `echolot init` (with `--force` when it says so), show the result; then run `echolot` again and continue from its new `next` |
39
+ | `doctor` | run `echolot doctor`, show what failed, stop — no report is trustworthy until it passes |
40
+ | `setup` | invoke the `echolot-setup` skill (the Skill tool) — it builds `echolot.yml` |
41
+ | `fix-config` | show the parse error, ask the human to fix `echolot.yml`, stop |
42
+ | `hunt` | invoke the `echolot-hunt` skill — it asks what regressed and runs perf-hunter |
43
+
44
+ With an argument, the argument wins over the state:
45
+
46
+ - `/echolot init` — run `echolot init`. Not setup: `init` installs or
47
+ updates **this** layer; the config is setup's job.
48
+ - `/echolot setup` — the `echolot-setup` skill.
49
+ - `/echolot hunt <words>` — the `echolot-hunt` skill, with the words as what
50
+ regressed. `/echolot <free text>` about slowness ("why is startup slow",
51
+ "the list stutters since the redesign") means the same.
52
+ - `/echolot doctor`, `/echolot status`, `/echolot analyze …` — run that
53
+ command, show the output.
54
+ - `/echolot reflect` — the `echolot-reflect` skill: how the last session
55
+ went, what to change in the tool.
56
+
57
+ `/echolot-setup`, `/echolot-hunt` and `/echolot-reflect` still exist for
58
+ whoever knows where they are going; `/echolot` is the door for everyone else.
59
+
60
+ ## The order of work
61
+
62
+ ```bash
63
+ echolot doctor -q # does the environment compute correctly?
64
+ echolot analyze <trace> -c echolot.yml
65
+ ```
66
+
67
+ `analyze` writes `.echolot/out/report.json` (for you) and `report.md` (for
68
+ humans), next to the config. Read the **json** — it has a stable schema.
69
+
70
+ No config? That is `next: setup`. No idea what is in the trace?
71
+ `echolot probe <trace> --process '<package>*'`.
72
+
73
+ **Silence is relative to the thresholds.** The report's `Config:` line and
74
+ `detectors[].params_source` say whether the numbers are the shipped defaults
75
+ or calibrated ones. Calibrated on the runs that hold the regression means the
76
+ bar sits above it: look with `echolot analyze … --defaults` before calling a
77
+ run clean.
78
+
79
+ ## Reading the report
80
+
81
+ Details live in `references/report.md`; three things here that you will get
82
+ wrong without them.
83
+
84
+ **Silent detectors matter as much as firing ones.** They stay in the report
85
+ with empty `rows`. Silence means that ground was checked and is clean — do not
86
+ go there.
87
+
88
+ **`self_ms` versus `total_ms`.** Self time, with children subtracted, is where
89
+ the time actually went. `traversal` with `total_ms: 354` and `self_ms: 79` does
90
+ almost nothing itself; dig into its children. Never add `total_ms` across rows:
91
+ they nest inside one another.
92
+
93
+ **Warnings inside `window`.** If `start_anchor.matches == 0`, the window
94
+ expanded to the whole trace and none of the numbers are about your scenario.
95
+ Fix the config rather than hunting a problem. Same for `process_alternatives` —
96
+ you may be analysing the wrong process.
97
+
98
+ ## From a finding to the code
99
+
100
+ 1. A firing detector gives you a `location` — a slice or thread name.
101
+ 2. The `domains` section of `echolot.yml` maps that name to a module and file.
102
+ 3. Not in `domains`? Grep the repository for the slice name: it is a string
103
+ literal inside `trace("...")`, survives minification, and is found exactly.
104
+ 4. Nothing found? The slice is most likely a system one (`bindApplication`,
105
+ `Choreographer#doFrame`, `binder transaction`). See `references/naming.md`.
106
+
107
+ When `uninstrumented_cpu` fires there is no code behind the finding by
108
+ definition: the thread burned CPU with no instrumentation. That is an address
109
+ for adding `trace{}`, not the location of a bug.
110
+
111
+ ## Boundaries
112
+
113
+ The tool **localises one specific regression**: you know "it was 3 s, now it is
114
+ 7 s" and need to find where. It does not search for the unknown across a pile
115
+ of traces.
116
+
117
+ Networking in benchmarks is mocked on purpose — we hunt problems in code, not
118
+ network speed.
119
+
120
+ Thresholds are tied to a device and a scenario. When either changes, run
121
+ `echolot calibrate` on healthy runs instead of nudging numbers by hand.
122
+
123
+ When the question is about the tool rather than the app — "how did that
124
+ session go, what should change in echolot" — that is `/echolot reflect`
125
+ (the `echolot-reflect` skill), not a hunt.
126
+
127
+ ## References
128
+
129
+ - `references/report.md` — the `report.json` schema, all detectors, what each column means
130
+ - `references/config.md` — the `echolot.yml` sections, what the code reads and what it does not
131
+ - `references/naming.md` — how ART names GC, locks and binder; facts from live Android 14
132
+ - `references/collect.md` — capturing a trace: perfetto and adb commands, verified in practice