agent-usage-manager 0.1.2__tar.gz → 0.2.0__tar.gz

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (24) hide show
  1. {agent_usage_manager-0.1.2 → agent_usage_manager-0.2.0}/.github/workflows/ci.yml +3 -0
  2. {agent_usage_manager-0.1.2 → agent_usage_manager-0.2.0}/.gitignore +1 -0
  3. agent_usage_manager-0.2.0/PKG-INFO +316 -0
  4. agent_usage_manager-0.2.0/README.md +297 -0
  5. agent_usage_manager-0.2.0/agent_usage_manager/agents.yaml +76 -0
  6. agent_usage_manager-0.2.0/agent_usage_manager/app.py +1073 -0
  7. agent_usage_manager-0.2.0/agent_usage_manager/cli.py +104 -0
  8. agent_usage_manager-0.2.0/agent_usage_manager/static/index.html +329 -0
  9. agent_usage_manager-0.2.0/docs/dashboard.png +0 -0
  10. {agent_usage_manager-0.1.2 → agent_usage_manager-0.2.0}/pyproject.toml +1 -1
  11. agent_usage_manager-0.2.0/tests/test_smoke.py +344 -0
  12. agent_usage_manager-0.2.0/uv.lock +1260 -0
  13. agent_usage_manager-0.1.2/PKG-INFO +0 -163
  14. agent_usage_manager-0.1.2/README.md +0 -144
  15. agent_usage_manager-0.1.2/agent_usage_manager/agents.yaml +0 -42
  16. agent_usage_manager-0.1.2/agent_usage_manager/app.py +0 -411
  17. agent_usage_manager-0.1.2/agent_usage_manager/cli.py +0 -43
  18. agent_usage_manager-0.1.2/agent_usage_manager/static/index.html +0 -129
  19. agent_usage_manager-0.1.2/tests/test_smoke.py +0 -92
  20. {agent_usage_manager-0.1.2 → agent_usage_manager-0.2.0}/LICENSE +0 -0
  21. {agent_usage_manager-0.1.2 → agent_usage_manager-0.2.0}/agent_usage_manager/__init__.py +0 -0
  22. {agent_usage_manager-0.1.2 → agent_usage_manager-0.2.0}/demo.tape +0 -0
  23. {agent_usage_manager-0.1.2 → agent_usage_manager-0.2.0}/requirements.txt +0 -0
  24. {agent_usage_manager-0.1.2 → agent_usage_manager-0.2.0}/run.sh +0 -0
@@ -5,6 +5,9 @@ on:
5
5
  branches: [main]
6
6
  pull_request:
7
7
 
8
+ permissions:
9
+ contents: read
10
+
8
11
  jobs:
9
12
  test:
10
13
  strategy:
@@ -7,3 +7,4 @@ agents.local.yaml
7
7
  dist/
8
8
  build/
9
9
  *.egg-info/
10
+ .gstack/
@@ -0,0 +1,316 @@
1
+ Metadata-Version: 2.4
2
+ Name: agent-usage-manager
3
+ Version: 0.2.0
4
+ Summary: htop for AI agents — liveness, CPU/mem/GPU usage, and a kill switch for headless agents (openclaw, hermes, ollama, vllm, claude-code).
5
+ Project-URL: Homepage, https://github.com/minglong51/agent-usage-manager
6
+ Project-URL: Repository, https://github.com/minglong51/agent-usage-manager
7
+ License: MIT
8
+ License-File: LICENSE
9
+ Keywords: ai-agents,gpu,llm,monitoring,observability,ollama,vllm
10
+ Requires-Python: >=3.9
11
+ Requires-Dist: fastapi>=0.110
12
+ Requires-Dist: psutil>=5.9
13
+ Requires-Dist: pyyaml>=6.0
14
+ Requires-Dist: uvicorn[standard]>=0.27
15
+ Provides-Extra: dev
16
+ Requires-Dist: httpx>=0.27; extra == 'dev'
17
+ Requires-Dist: pytest>=7; extra == 'dev'
18
+ Description-Content-Type: text/markdown
19
+
20
+ # agent-usage-manager
21
+
22
+ A tiny, single-file web dashboard for **headless AI agents** running on a machine —
23
+ OpenClaw, Hermes, Claude Code, Ollama, vLLM, llama.cpp, or anything you name. It
24
+ shows which agents are alive and what they're costing you (CPU, memory, GPU), and
25
+ gives you a **kill button** per agent. Think `htop`, scoped to just your agents —
26
+ the [screenshot below](docs/dashboard.png) is a real run on a fleet node.
27
+
28
+ No database, no auth layer, no dependencies beyond FastAPI + psutil. Runs on
29
+ macOS and Linux. Meant to be cloned, configured, and run on any node in a fleet.
30
+
31
+ ![agent-usage-manager — live dashboard](docs/dashboard.png)
32
+
33
+ *A real run: ten agents grouped by process tree (`+N` = children rolled up),
34
+ per-agent CPU/memory/uptime, launchd-supervised jobs flagged, and a kill button
35
+ per row.*
36
+
37
+ ```
38
+ AGENT PID STATUS CPU % MEM MB GPU MB UPTIME COMMAND ┆
39
+ openclaw +3 48213 ● running 62.4 1840 7320 2h 11m openclaw serve … [kill] [force]
40
+ claude-code +9 73590 ● running 97.4 7630 — 1h 02m claude --chann … [kill] [force]
41
+ hermes 49001 ● running 18.0 512 — 44m hermes worker … [kill] [force]
42
+ ollama 50122 ● running 3.1 9210 14080 6h 02m ollama runner … [kill] [force]
43
+ ```
44
+ (`+N` = child processes rolled up under the agent; CPU/mem/GPU are tree totals.)
45
+
46
+ > The schematic above shows the GPU column (NVIDIA only); the screenshot is a real
47
+ > run on Apple Silicon, where per-process GPU stats aren't available so that column
48
+ > is hidden. The UI auto-refreshes every 3s.
49
+
50
+ ## What it does
51
+
52
+ - **One row per agent.** Agents are grouped by process tree — the spawned children
53
+ of an agent (inference subprocesses, MCP servers, helpers) are rolled up under it
54
+ with a `+N` badge instead of cluttering the list as separate rows.
55
+ - **Liveness** — green dot = running, red = zombie/dead. Status column shows the OS state.
56
+ - **Usage** — CPU %, resident memory (MB), GPU memory (MB, NVIDIA only), and uptime,
57
+ refreshed every 3s. **CPU/mem/GPU are tree totals** — the agent's true cost including
58
+ everything it spawned.
59
+ - **Kill the tree** — `kill` sends SIGTERM to the agent *and its children* (so spawned
60
+ helpers don't leak resources), `force` sends SIGKILL. SIGTERM auto-escalates to
61
+ SIGKILL after 3s. The confirm dialog tells you how many child processes will stop.
62
+ - **Trends, not just snapshots.** Each row has a CPU sparkline (last ~20 min, sampled
63
+ in the background even with no browser open), plus a **`hot 5m+`** badge when an
64
+ agent has been pegged ≥90% CPU for 5+ minutes, an **`idle 10m+`** badge when a
65
+ long-running agent has done nothing for 10+ minutes, and a **`churn ×N`** badge when
66
+ the same agent has died young 3+ times in 10 minutes — the states worth investigating
67
+ (runaway, possibly wedged, crash-looping under a supervisor). Churn is what hot/idle
68
+ can't see: a crash-looping process is a fresh pid every poll, so no per-process
69
+ window ever fills. A **`leak?`** badge fires when an agent's memory ratchets up
70
+ ≥30% (and ≥128 MB) over 15 minutes without coming back down.
71
+ - **Alerts.** A dashboard only helps while you're looking at it. Add an `alerts:`
72
+ block to `agents.yaml` and any badge appearing runs your command (desktop
73
+ notification, Telegram bot, pager — anything) with the details in `$AUM_*` env
74
+ vars. Fires once per transition with a cooldown, never from the `list` CLI.
75
+ By default only `hot`/`churn`/`leak` alert — `idle` is the normal state of an
76
+ agent fleet that waits for work, so it's opt-in.
77
+
78
+ ```yaml
79
+ alerts:
80
+ command: 'terminal-notifier -title agents -message "$AUM_MSG"'
81
+ cooldown: 600
82
+ flags: [hot, churn, leak] # the default; add idle to opt in
83
+ ```
84
+ - **Prometheus `/metrics`.** Per-agent CPU/mem/instances/restarts and badge states
85
+ in text exposition format, aggregated per label (no pid-churn series bloat) —
86
+ point Grafana or any Prometheus scraper at `http://127.0.0.1:8765/metrics`.
87
+ - **Expand the tree.** Click the `+N` badge to unfold an agent's child processes
88
+ (per-child CPU/mem/command) — see what a kill would actually stop before clicking it.
89
+ - **Config hot-reload.** Edits to `agents.yaml` apply on the next poll, no restart.
90
+ A broken edit keeps the last good config and shows the parse error in the header.
91
+ - **`list` subcommand.** `agent-usage-manager list` (or `list --json`) prints a one-shot
92
+ table to stdout — no server, good for scripts and cron checks.
93
+ - **Kill-safe table.** Rows keep a stable order (sorted by label) and never reorder
94
+ while your pointer is over the table, so the kill button can't shift under your
95
+ cursor mid-click.
96
+
97
+ ## Safety
98
+
99
+ This is the important part — a web page that can kill processes needs guardrails:
100
+
101
+ - **Allowlist only.** Only processes matching a pattern in `agents.yaml` are ever
102
+ listed *or* killable. The kill endpoint re-checks the match server-side before
103
+ sending any signal, so the dashboard can never be used to kill an arbitrary PID.
104
+ - **Protected patterns.** Anything matching `protect:` in `agents.yaml` — plus the
105
+ monitor's own process and PID 1 — shows a disabled, greyed-out kill button and is
106
+ refused server-side.
107
+ - **Secret redaction.** Command lines often carry tokens/keys in env vars or flags
108
+ (`FOO_TOKEN=...`, `--api-key ...`, `sk-...`, `ghp_...`, JWTs). The command column
109
+ redacts these to `***` before they ever reach the browser — safe to screenshot.
110
+ - **Browser guard (CSRF + DNS rebinding).** Binding to localhost doesn't keep
111
+ browsers out — any web page you visit can `fetch()` a localhost port. Requests
112
+ whose `Host` is a non-local DNS name are refused (DNS-rebinding guard), and a
113
+ kill request carrying a foreign `Origin` is refused (CSRF guard) — so a
114
+ malicious page can't kill your agents or read your process list. `curl` and
115
+ the dashboard itself are unaffected.
116
+ - **Bind local by default.** It listens on `127.0.0.1`. Don't expose it to a network
117
+ without putting auth in front of it (reverse proxy + basic auth, SSH tunnel, etc.) —
118
+ it has no built-in authentication.
119
+
120
+ ## Limits & known issues
121
+
122
+ - **GPU column is NVIDIA-only.** Per-process GPU memory comes from
123
+ `nvidia-smi --query-compute-apps` — NVIDIA compute processes (CUDA), in
124
+ practice on Linux. AMD/Intel GPUs aren't read, graphics-only workloads don't
125
+ appear, and Apple Silicon has no per-process GPU accounting API at all, so
126
+ the column is hidden on Macs.
127
+ - **Supervision detection is launchd-only (macOS, user domain).** Root
128
+ `LaunchDaemons` aren't flagged — that needs a privileged
129
+ `launchctl print system/…`. On Linux, systemd-supervised services
130
+ (`Restart=always`) aren't detected either, so killing one looks like it
131
+ failed when systemd respawns it — use `systemctl stop` for those.
132
+ - **Same-user privileges only.** Signals are sent with the server's own
133
+ privileges. Agents running as another user (or root) are listed, but a kill
134
+ won't take (`killed: 0` in the response), and CPU/mem can read as 0 where
135
+ the OS denies access.
136
+ - **History is in-memory.** Sparklines and the `hot`/`idle` flags (~20 min
137
+ window) rebuild from scratch after a server restart.
138
+ - **Windows is untested.** Kill maps to `TerminateProcess` via psutil and may
139
+ work, but CI covers Linux + macOS only.
140
+
141
+ ## Quick start
142
+
143
+ **Recommended — one command, nothing to install first:**
144
+
145
+ ```bash
146
+ uvx agent-usage-manager
147
+ # then open http://127.0.0.1:8765 (it also opens automatically)
148
+ ```
149
+
150
+ `uvx` fetches and runs it in one step — no separate install, no virtualenv, no
151
+ leftovers. Don't have [`uv`](https://github.com/astral-sh/uv) yet? One line:
152
+
153
+ ```bash
154
+ curl -LsSf https://astral.sh/uv/install.sh | sh # macOS / Linux
155
+ # or: pip install uv
156
+ ```
157
+
158
+ <details>
159
+ <summary>Other ways to install</summary>
160
+
161
+ ```bash
162
+ pipx install agent-usage-manager # clean isolated global CLI (needs pipx)
163
+
164
+ pip install agent-usage-manager # universal; use inside a venv —
165
+ # system Python may refuse with
166
+ # "externally-managed-environment"
167
+ ```
168
+
169
+ Then run `agent-usage-manager` (flags below).
170
+ </details>
171
+
172
+ From a clone (for hacking on it):
173
+
174
+ ```bash
175
+ git clone <this-repo> && cd agent-usage-manager
176
+ ./run.sh # venv + editable install, serves on :8765
177
+ ```
178
+
179
+ It opens the dashboard in your browser automatically. Flags: `--host`, `--port`,
180
+ `--config /path/to/agents.yaml`, `--no-browser` (for headless/server use).
181
+
182
+ ## Configure which processes are "agents"
183
+
184
+ Edit `agents.yaml`:
185
+
186
+ ```yaml
187
+ agents:
188
+ - label: openclaw # shown as the badge in the UI
189
+ match: openclaw # case-insensitive substring of the command line
190
+ - label: hermes
191
+ match: hermes
192
+ - label: claude-code
193
+ match: "claude(\\s|$|-code)"
194
+ regex: true # treat `match` as a regex instead of substring
195
+
196
+ protect: # matched + listed, but never killable
197
+ - uvicorn
198
+
199
+ ignore: # never an agent: not listed, not killable
200
+ - crashpad # incidental processes that share a name/bundle
201
+ - shipit # path with a real agent (crash handlers,
202
+ - kiro-cli-term # auto-updaters, integrated-terminal shells, …)
203
+ ```
204
+
205
+ A process matches if the pattern hits its **executable basename + first few
206
+ arguments** — deliberately not the whole command line, so a long embedded arg
207
+ (e.g. a system prompt mentioning "claude") can't misclassify a wrapper. On macOS
208
+ the outermost `.app` **bundle name** is also included, so GUI agents that launch
209
+ a generically-named binary (Kiro.app → `Electron`) are still matched by app name.
210
+
211
+ `protect:` keeps a matched process listed but refuses to kill it; `ignore:`
212
+ drops it from agent classification entirely.
213
+
214
+ **Which `agents.yaml` is used** — resolved once at startup, first hit wins:
215
+
216
+ 1. `AGENTS_CONFIG=/path/to/agents.yaml` env var (the `--config` flag sets this)
217
+ 2. `./agents.yaml` in the directory you launched from
218
+ 3. the default bundled with the package
219
+
220
+ The dashboard header (and `list --json`) shows the resolved path, so you can
221
+ always see which file is live. Hot-reload watches that one file. An
222
+ `AGENTS_CONFIG` path that doesn't exist is an error at startup, not a silent
223
+ fallback.
224
+
225
+ ## launchd-supervised agents (macOS)
226
+
227
+ Some agents run as **launchd services** (a `~/Library/LaunchAgents/*.plist`, or
228
+ anything started by `brew services`). If such a job sets `KeepAlive`, a signal
229
+ can't stop it: the process dies, launchd immediately respawns it under a new PID,
230
+ and the dashboard's "kill" looks like it silently failed.
231
+
232
+ The dashboard detects these (via `launchctl list`) and marks them with a
233
+ **`launchd`** badge. Instead of dead-end kill/force buttons it shows the command
234
+ that actually stops the job — click to copy:
235
+
236
+ ```sh
237
+ launchctl bootout gui/<uid>/<label> # stop now
238
+ launchctl disable gui/<uid>/<label> # …and don't auto-start at login
239
+ ```
240
+
241
+ The kill endpoint refuses signals for these jobs (HTTP 409) and returns the same
242
+ guidance, so the API never lies about a kill that won't stick. The message is
243
+ tailored to the job: `KeepAlive` jobs are told a signal won't stick at all;
244
+ `RunAtLoad`-only jobs are told a signal works now but the job restarts at next
245
+ login. *Limitation:* root `LaunchDaemons` aren't flagged — see
246
+ [Limits & known issues](#limits--known-issues).
247
+
248
+ ## GPU notes
249
+
250
+ Per-process GPU memory comes from `nvidia-smi` when it's on `PATH` (Linux / NVIDIA).
251
+ **Apple Silicon has no per-process GPU accounting API**, so the GPU column stays blank
252
+ on Macs — CPU and memory are the meaningful resource signals there.
253
+
254
+ ## API
255
+
256
+ - `GET /api/agents` → `{ agents: [...], host, cpu_count, mem_total_mb, mem_used_pct, config_path, config_error, ts }`
257
+ — each agent includes `trend` (recent CPU samples) and `flag` (`"hot"` / `"idle"` / `null`)
258
+ - `GET /api/tree/{pid}` → the agent's process subtree (per-child pid/name/cpu/mem/cmdline);
259
+ only works on recognized agents, same authorization as kill
260
+ - `POST /api/kill/{pid}?force=false` → SIGTERM (or SIGKILL with `force=true`)
261
+
262
+ ## Run as a service
263
+
264
+ Linux (systemd), `~/.config/systemd/user/agent-usage-manager.service`:
265
+
266
+ ```ini
267
+ [Unit]
268
+ Description=agent usage manager
269
+ [Service]
270
+ ExecStart=%h/agent-usage-manager/.venv/bin/uvicorn app:app --port 8765
271
+ WorkingDirectory=%h/agent-usage-manager
272
+ Restart=on-failure
273
+ [Install]
274
+ WantedBy=default.target
275
+ ```
276
+
277
+ ```bash
278
+ systemctl --user enable --now agent-usage-manager
279
+ ```
280
+
281
+ ## Development
282
+
283
+ ```bash
284
+ git clone https://github.com/minglong51/agent-usage-manager && cd agent-usage-manager
285
+ pip install -e ".[dev]"
286
+ pytest -q
287
+ ```
288
+
289
+ CI runs the test suite on Linux + macOS (Python 3.9 and 3.12) on every push and PR.
290
+ Cross-platform note: kill uses psutil's `terminate()`/`kill()`, which map to
291
+ SIGTERM/SIGKILL on POSIX and TerminateProcess on Windows.
292
+
293
+ ## Troubleshooting
294
+
295
+ - **`pip install` fails building psutil** — no prebuilt wheel for your
296
+ Python/platform, so pip compiles it: you need a C toolchain and Python
297
+ headers (`xcode-select --install` on macOS; `apt install gcc python3-dev`
298
+ on Debian/Ubuntu). Or skip the problem with `uvx agent-usage-manager`.
299
+ - **Dashboard is empty / "No matching agents running"** — first check which
300
+ config was picked up (resolution order above; the header shows the resolved
301
+ path). Then remember matching is against the executable basename + first few
302
+ arguments, not the full command line — a pattern that only appears deep in
303
+ the args won't match.
304
+ - **Kill "doesn't work" — the agent comes back under a new PID** — it's
305
+ supervised. On macOS the row gets a `launchd` badge with the `launchctl
306
+ bootout` command that actually stops it; on Linux, systemd services aren't
307
+ detected (see limits) — `systemctl stop` them. A `killed: 0` in the kill
308
+ response means nothing was actually signaled (e.g. the agent runs as
309
+ another user).
310
+ - **HTTP 403 on every request** — the DNS-rebinding guard refuses non-local
311
+ hostnames. Use `http://127.0.0.1:8765` (or a bare IP) instead of a custom
312
+ DNS name pointing at the box.
313
+
314
+ ## License
315
+
316
+ MIT
@@ -0,0 +1,297 @@
1
+ # agent-usage-manager
2
+
3
+ A tiny, single-file web dashboard for **headless AI agents** running on a machine —
4
+ OpenClaw, Hermes, Claude Code, Ollama, vLLM, llama.cpp, or anything you name. It
5
+ shows which agents are alive and what they're costing you (CPU, memory, GPU), and
6
+ gives you a **kill button** per agent. Think `htop`, scoped to just your agents —
7
+ the [screenshot below](docs/dashboard.png) is a real run on a fleet node.
8
+
9
+ No database, no auth layer, no dependencies beyond FastAPI + psutil. Runs on
10
+ macOS and Linux. Meant to be cloned, configured, and run on any node in a fleet.
11
+
12
+ ![agent-usage-manager — live dashboard](docs/dashboard.png)
13
+
14
+ *A real run: ten agents grouped by process tree (`+N` = children rolled up),
15
+ per-agent CPU/memory/uptime, launchd-supervised jobs flagged, and a kill button
16
+ per row.*
17
+
18
+ ```
19
+ AGENT PID STATUS CPU % MEM MB GPU MB UPTIME COMMAND ┆
20
+ openclaw +3 48213 ● running 62.4 1840 7320 2h 11m openclaw serve … [kill] [force]
21
+ claude-code +9 73590 ● running 97.4 7630 — 1h 02m claude --chann … [kill] [force]
22
+ hermes 49001 ● running 18.0 512 — 44m hermes worker … [kill] [force]
23
+ ollama 50122 ● running 3.1 9210 14080 6h 02m ollama runner … [kill] [force]
24
+ ```
25
+ (`+N` = child processes rolled up under the agent; CPU/mem/GPU are tree totals.)
26
+
27
+ > The schematic above shows the GPU column (NVIDIA only); the screenshot is a real
28
+ > run on Apple Silicon, where per-process GPU stats aren't available so that column
29
+ > is hidden. The UI auto-refreshes every 3s.
30
+
31
+ ## What it does
32
+
33
+ - **One row per agent.** Agents are grouped by process tree — the spawned children
34
+ of an agent (inference subprocesses, MCP servers, helpers) are rolled up under it
35
+ with a `+N` badge instead of cluttering the list as separate rows.
36
+ - **Liveness** — green dot = running, red = zombie/dead. Status column shows the OS state.
37
+ - **Usage** — CPU %, resident memory (MB), GPU memory (MB, NVIDIA only), and uptime,
38
+ refreshed every 3s. **CPU/mem/GPU are tree totals** — the agent's true cost including
39
+ everything it spawned.
40
+ - **Kill the tree** — `kill` sends SIGTERM to the agent *and its children* (so spawned
41
+ helpers don't leak resources), `force` sends SIGKILL. SIGTERM auto-escalates to
42
+ SIGKILL after 3s. The confirm dialog tells you how many child processes will stop.
43
+ - **Trends, not just snapshots.** Each row has a CPU sparkline (last ~20 min, sampled
44
+ in the background even with no browser open), plus a **`hot 5m+`** badge when an
45
+ agent has been pegged ≥90% CPU for 5+ minutes, an **`idle 10m+`** badge when a
46
+ long-running agent has done nothing for 10+ minutes, and a **`churn ×N`** badge when
47
+ the same agent has died young 3+ times in 10 minutes — the states worth investigating
48
+ (runaway, possibly wedged, crash-looping under a supervisor). Churn is what hot/idle
49
+ can't see: a crash-looping process is a fresh pid every poll, so no per-process
50
+ window ever fills. A **`leak?`** badge fires when an agent's memory ratchets up
51
+ ≥30% (and ≥128 MB) over 15 minutes without coming back down.
52
+ - **Alerts.** A dashboard only helps while you're looking at it. Add an `alerts:`
53
+ block to `agents.yaml` and any badge appearing runs your command (desktop
54
+ notification, Telegram bot, pager — anything) with the details in `$AUM_*` env
55
+ vars. Fires once per transition with a cooldown, never from the `list` CLI.
56
+ By default only `hot`/`churn`/`leak` alert — `idle` is the normal state of an
57
+ agent fleet that waits for work, so it's opt-in.
58
+
59
+ ```yaml
60
+ alerts:
61
+ command: 'terminal-notifier -title agents -message "$AUM_MSG"'
62
+ cooldown: 600
63
+ flags: [hot, churn, leak] # the default; add idle to opt in
64
+ ```
65
+ - **Prometheus `/metrics`.** Per-agent CPU/mem/instances/restarts and badge states
66
+ in text exposition format, aggregated per label (no pid-churn series bloat) —
67
+ point Grafana or any Prometheus scraper at `http://127.0.0.1:8765/metrics`.
68
+ - **Expand the tree.** Click the `+N` badge to unfold an agent's child processes
69
+ (per-child CPU/mem/command) — see what a kill would actually stop before clicking it.
70
+ - **Config hot-reload.** Edits to `agents.yaml` apply on the next poll, no restart.
71
+ A broken edit keeps the last good config and shows the parse error in the header.
72
+ - **`list` subcommand.** `agent-usage-manager list` (or `list --json`) prints a one-shot
73
+ table to stdout — no server, good for scripts and cron checks.
74
+ - **Kill-safe table.** Rows keep a stable order (sorted by label) and never reorder
75
+ while your pointer is over the table, so the kill button can't shift under your
76
+ cursor mid-click.
77
+
78
+ ## Safety
79
+
80
+ This is the important part — a web page that can kill processes needs guardrails:
81
+
82
+ - **Allowlist only.** Only processes matching a pattern in `agents.yaml` are ever
83
+ listed *or* killable. The kill endpoint re-checks the match server-side before
84
+ sending any signal, so the dashboard can never be used to kill an arbitrary PID.
85
+ - **Protected patterns.** Anything matching `protect:` in `agents.yaml` — plus the
86
+ monitor's own process and PID 1 — shows a disabled, greyed-out kill button and is
87
+ refused server-side.
88
+ - **Secret redaction.** Command lines often carry tokens/keys in env vars or flags
89
+ (`FOO_TOKEN=...`, `--api-key ...`, `sk-...`, `ghp_...`, JWTs). The command column
90
+ redacts these to `***` before they ever reach the browser — safe to screenshot.
91
+ - **Browser guard (CSRF + DNS rebinding).** Binding to localhost doesn't keep
92
+ browsers out — any web page you visit can `fetch()` a localhost port. Requests
93
+ whose `Host` is a non-local DNS name are refused (DNS-rebinding guard), and a
94
+ kill request carrying a foreign `Origin` is refused (CSRF guard) — so a
95
+ malicious page can't kill your agents or read your process list. `curl` and
96
+ the dashboard itself are unaffected.
97
+ - **Bind local by default.** It listens on `127.0.0.1`. Don't expose it to a network
98
+ without putting auth in front of it (reverse proxy + basic auth, SSH tunnel, etc.) —
99
+ it has no built-in authentication.
100
+
101
+ ## Limits & known issues
102
+
103
+ - **GPU column is NVIDIA-only.** Per-process GPU memory comes from
104
+ `nvidia-smi --query-compute-apps` — NVIDIA compute processes (CUDA), in
105
+ practice on Linux. AMD/Intel GPUs aren't read, graphics-only workloads don't
106
+ appear, and Apple Silicon has no per-process GPU accounting API at all, so
107
+ the column is hidden on Macs.
108
+ - **Supervision detection is launchd-only (macOS, user domain).** Root
109
+ `LaunchDaemons` aren't flagged — that needs a privileged
110
+ `launchctl print system/…`. On Linux, systemd-supervised services
111
+ (`Restart=always`) aren't detected either, so killing one looks like it
112
+ failed when systemd respawns it — use `systemctl stop` for those.
113
+ - **Same-user privileges only.** Signals are sent with the server's own
114
+ privileges. Agents running as another user (or root) are listed, but a kill
115
+ won't take (`killed: 0` in the response), and CPU/mem can read as 0 where
116
+ the OS denies access.
117
+ - **History is in-memory.** Sparklines and the `hot`/`idle` flags (~20 min
118
+ window) rebuild from scratch after a server restart.
119
+ - **Windows is untested.** Kill maps to `TerminateProcess` via psutil and may
120
+ work, but CI covers Linux + macOS only.
121
+
122
+ ## Quick start
123
+
124
+ **Recommended — one command, nothing to install first:**
125
+
126
+ ```bash
127
+ uvx agent-usage-manager
128
+ # then open http://127.0.0.1:8765 (it also opens automatically)
129
+ ```
130
+
131
+ `uvx` fetches and runs it in one step — no separate install, no virtualenv, no
132
+ leftovers. Don't have [`uv`](https://github.com/astral-sh/uv) yet? One line:
133
+
134
+ ```bash
135
+ curl -LsSf https://astral.sh/uv/install.sh | sh # macOS / Linux
136
+ # or: pip install uv
137
+ ```
138
+
139
+ <details>
140
+ <summary>Other ways to install</summary>
141
+
142
+ ```bash
143
+ pipx install agent-usage-manager # clean isolated global CLI (needs pipx)
144
+
145
+ pip install agent-usage-manager # universal; use inside a venv —
146
+ # system Python may refuse with
147
+ # "externally-managed-environment"
148
+ ```
149
+
150
+ Then run `agent-usage-manager` (flags below).
151
+ </details>
152
+
153
+ From a clone (for hacking on it):
154
+
155
+ ```bash
156
+ git clone <this-repo> && cd agent-usage-manager
157
+ ./run.sh # venv + editable install, serves on :8765
158
+ ```
159
+
160
+ It opens the dashboard in your browser automatically. Flags: `--host`, `--port`,
161
+ `--config /path/to/agents.yaml`, `--no-browser` (for headless/server use).
162
+
163
+ ## Configure which processes are "agents"
164
+
165
+ Edit `agents.yaml`:
166
+
167
+ ```yaml
168
+ agents:
169
+ - label: openclaw # shown as the badge in the UI
170
+ match: openclaw # case-insensitive substring of the command line
171
+ - label: hermes
172
+ match: hermes
173
+ - label: claude-code
174
+ match: "claude(\\s|$|-code)"
175
+ regex: true # treat `match` as a regex instead of substring
176
+
177
+ protect: # matched + listed, but never killable
178
+ - uvicorn
179
+
180
+ ignore: # never an agent: not listed, not killable
181
+ - crashpad # incidental processes that share a name/bundle
182
+ - shipit # path with a real agent (crash handlers,
183
+ - kiro-cli-term # auto-updaters, integrated-terminal shells, …)
184
+ ```
185
+
186
+ A process matches if the pattern hits its **executable basename + first few
187
+ arguments** — deliberately not the whole command line, so a long embedded arg
188
+ (e.g. a system prompt mentioning "claude") can't misclassify a wrapper. On macOS
189
+ the outermost `.app` **bundle name** is also included, so GUI agents that launch
190
+ a generically-named binary (Kiro.app → `Electron`) are still matched by app name.
191
+
192
+ `protect:` keeps a matched process listed but refuses to kill it; `ignore:`
193
+ drops it from agent classification entirely.
194
+
195
+ **Which `agents.yaml` is used** — resolved once at startup, first hit wins:
196
+
197
+ 1. `AGENTS_CONFIG=/path/to/agents.yaml` env var (the `--config` flag sets this)
198
+ 2. `./agents.yaml` in the directory you launched from
199
+ 3. the default bundled with the package
200
+
201
+ The dashboard header (and `list --json`) shows the resolved path, so you can
202
+ always see which file is live. Hot-reload watches that one file. An
203
+ `AGENTS_CONFIG` path that doesn't exist is an error at startup, not a silent
204
+ fallback.
205
+
206
+ ## launchd-supervised agents (macOS)
207
+
208
+ Some agents run as **launchd services** (a `~/Library/LaunchAgents/*.plist`, or
209
+ anything started by `brew services`). If such a job sets `KeepAlive`, a signal
210
+ can't stop it: the process dies, launchd immediately respawns it under a new PID,
211
+ and the dashboard's "kill" looks like it silently failed.
212
+
213
+ The dashboard detects these (via `launchctl list`) and marks them with a
214
+ **`launchd`** badge. Instead of dead-end kill/force buttons it shows the command
215
+ that actually stops the job — click to copy:
216
+
217
+ ```sh
218
+ launchctl bootout gui/<uid>/<label> # stop now
219
+ launchctl disable gui/<uid>/<label> # …and don't auto-start at login
220
+ ```
221
+
222
+ The kill endpoint refuses signals for these jobs (HTTP 409) and returns the same
223
+ guidance, so the API never lies about a kill that won't stick. The message is
224
+ tailored to the job: `KeepAlive` jobs are told a signal won't stick at all;
225
+ `RunAtLoad`-only jobs are told a signal works now but the job restarts at next
226
+ login. *Limitation:* root `LaunchDaemons` aren't flagged — see
227
+ [Limits & known issues](#limits--known-issues).
228
+
229
+ ## GPU notes
230
+
231
+ Per-process GPU memory comes from `nvidia-smi` when it's on `PATH` (Linux / NVIDIA).
232
+ **Apple Silicon has no per-process GPU accounting API**, so the GPU column stays blank
233
+ on Macs — CPU and memory are the meaningful resource signals there.
234
+
235
+ ## API
236
+
237
+ - `GET /api/agents` → `{ agents: [...], host, cpu_count, mem_total_mb, mem_used_pct, config_path, config_error, ts }`
238
+ — each agent includes `trend` (recent CPU samples) and `flag` (`"hot"` / `"idle"` / `null`)
239
+ - `GET /api/tree/{pid}` → the agent's process subtree (per-child pid/name/cpu/mem/cmdline);
240
+ only works on recognized agents, same authorization as kill
241
+ - `POST /api/kill/{pid}?force=false` → SIGTERM (or SIGKILL with `force=true`)
242
+
243
+ ## Run as a service
244
+
245
+ Linux (systemd), `~/.config/systemd/user/agent-usage-manager.service`:
246
+
247
+ ```ini
248
+ [Unit]
249
+ Description=agent usage manager
250
+ [Service]
251
+ ExecStart=%h/agent-usage-manager/.venv/bin/uvicorn app:app --port 8765
252
+ WorkingDirectory=%h/agent-usage-manager
253
+ Restart=on-failure
254
+ [Install]
255
+ WantedBy=default.target
256
+ ```
257
+
258
+ ```bash
259
+ systemctl --user enable --now agent-usage-manager
260
+ ```
261
+
262
+ ## Development
263
+
264
+ ```bash
265
+ git clone https://github.com/minglong51/agent-usage-manager && cd agent-usage-manager
266
+ pip install -e ".[dev]"
267
+ pytest -q
268
+ ```
269
+
270
+ CI runs the test suite on Linux + macOS (Python 3.9 and 3.12) on every push and PR.
271
+ Cross-platform note: kill uses psutil's `terminate()`/`kill()`, which map to
272
+ SIGTERM/SIGKILL on POSIX and TerminateProcess on Windows.
273
+
274
+ ## Troubleshooting
275
+
276
+ - **`pip install` fails building psutil** — no prebuilt wheel for your
277
+ Python/platform, so pip compiles it: you need a C toolchain and Python
278
+ headers (`xcode-select --install` on macOS; `apt install gcc python3-dev`
279
+ on Debian/Ubuntu). Or skip the problem with `uvx agent-usage-manager`.
280
+ - **Dashboard is empty / "No matching agents running"** — first check which
281
+ config was picked up (resolution order above; the header shows the resolved
282
+ path). Then remember matching is against the executable basename + first few
283
+ arguments, not the full command line — a pattern that only appears deep in
284
+ the args won't match.
285
+ - **Kill "doesn't work" — the agent comes back under a new PID** — it's
286
+ supervised. On macOS the row gets a `launchd` badge with the `launchctl
287
+ bootout` command that actually stops it; on Linux, systemd services aren't
288
+ detected (see limits) — `systemctl stop` them. A `killed: 0` in the kill
289
+ response means nothing was actually signaled (e.g. the agent runs as
290
+ another user).
291
+ - **HTTP 403 on every request** — the DNS-rebinding guard refuses non-local
292
+ hostnames. Use `http://127.0.0.1:8765` (or a bare IP) instead of a custom
293
+ DNS name pointing at the box.
294
+
295
+ ## License
296
+
297
+ MIT