agent-usage-manager 0.1.2__tar.gz → 0.2.0__tar.gz
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- {agent_usage_manager-0.1.2 → agent_usage_manager-0.2.0}/.github/workflows/ci.yml +3 -0
- {agent_usage_manager-0.1.2 → agent_usage_manager-0.2.0}/.gitignore +1 -0
- agent_usage_manager-0.2.0/PKG-INFO +316 -0
- agent_usage_manager-0.2.0/README.md +297 -0
- agent_usage_manager-0.2.0/agent_usage_manager/agents.yaml +76 -0
- agent_usage_manager-0.2.0/agent_usage_manager/app.py +1073 -0
- agent_usage_manager-0.2.0/agent_usage_manager/cli.py +104 -0
- agent_usage_manager-0.2.0/agent_usage_manager/static/index.html +329 -0
- agent_usage_manager-0.2.0/docs/dashboard.png +0 -0
- {agent_usage_manager-0.1.2 → agent_usage_manager-0.2.0}/pyproject.toml +1 -1
- agent_usage_manager-0.2.0/tests/test_smoke.py +344 -0
- agent_usage_manager-0.2.0/uv.lock +1260 -0
- agent_usage_manager-0.1.2/PKG-INFO +0 -163
- agent_usage_manager-0.1.2/README.md +0 -144
- agent_usage_manager-0.1.2/agent_usage_manager/agents.yaml +0 -42
- agent_usage_manager-0.1.2/agent_usage_manager/app.py +0 -411
- agent_usage_manager-0.1.2/agent_usage_manager/cli.py +0 -43
- agent_usage_manager-0.1.2/agent_usage_manager/static/index.html +0 -129
- agent_usage_manager-0.1.2/tests/test_smoke.py +0 -92
- {agent_usage_manager-0.1.2 → agent_usage_manager-0.2.0}/LICENSE +0 -0
- {agent_usage_manager-0.1.2 → agent_usage_manager-0.2.0}/agent_usage_manager/__init__.py +0 -0
- {agent_usage_manager-0.1.2 → agent_usage_manager-0.2.0}/demo.tape +0 -0
- {agent_usage_manager-0.1.2 → agent_usage_manager-0.2.0}/requirements.txt +0 -0
- {agent_usage_manager-0.1.2 → agent_usage_manager-0.2.0}/run.sh +0 -0
|
@@ -0,0 +1,316 @@
|
|
|
1
|
+
Metadata-Version: 2.4
|
|
2
|
+
Name: agent-usage-manager
|
|
3
|
+
Version: 0.2.0
|
|
4
|
+
Summary: htop for AI agents — liveness, CPU/mem/GPU usage, and a kill switch for headless agents (openclaw, hermes, ollama, vllm, claude-code).
|
|
5
|
+
Project-URL: Homepage, https://github.com/minglong51/agent-usage-manager
|
|
6
|
+
Project-URL: Repository, https://github.com/minglong51/agent-usage-manager
|
|
7
|
+
License: MIT
|
|
8
|
+
License-File: LICENSE
|
|
9
|
+
Keywords: ai-agents,gpu,llm,monitoring,observability,ollama,vllm
|
|
10
|
+
Requires-Python: >=3.9
|
|
11
|
+
Requires-Dist: fastapi>=0.110
|
|
12
|
+
Requires-Dist: psutil>=5.9
|
|
13
|
+
Requires-Dist: pyyaml>=6.0
|
|
14
|
+
Requires-Dist: uvicorn[standard]>=0.27
|
|
15
|
+
Provides-Extra: dev
|
|
16
|
+
Requires-Dist: httpx>=0.27; extra == 'dev'
|
|
17
|
+
Requires-Dist: pytest>=7; extra == 'dev'
|
|
18
|
+
Description-Content-Type: text/markdown
|
|
19
|
+
|
|
20
|
+
# agent-usage-manager
|
|
21
|
+
|
|
22
|
+
A tiny, single-file web dashboard for **headless AI agents** running on a machine —
|
|
23
|
+
OpenClaw, Hermes, Claude Code, Ollama, vLLM, llama.cpp, or anything you name. It
|
|
24
|
+
shows which agents are alive and what they're costing you (CPU, memory, GPU), and
|
|
25
|
+
gives you a **kill button** per agent. Think `htop`, scoped to just your agents —
|
|
26
|
+
the [screenshot below](docs/dashboard.png) is a real run on a fleet node.
|
|
27
|
+
|
|
28
|
+
No database, no auth layer, no dependencies beyond FastAPI + psutil. Runs on
|
|
29
|
+
macOS and Linux. Meant to be cloned, configured, and run on any node in a fleet.
|
|
30
|
+
|
|
31
|
+

|
|
32
|
+
|
|
33
|
+
*A real run: ten agents grouped by process tree (`+N` = children rolled up),
|
|
34
|
+
per-agent CPU/memory/uptime, launchd-supervised jobs flagged, and a kill button
|
|
35
|
+
per row.*
|
|
36
|
+
|
|
37
|
+
```
|
|
38
|
+
AGENT PID STATUS CPU % MEM MB GPU MB UPTIME COMMAND ┆
|
|
39
|
+
openclaw +3 48213 ● running 62.4 1840 7320 2h 11m openclaw serve … [kill] [force]
|
|
40
|
+
claude-code +9 73590 ● running 97.4 7630 — 1h 02m claude --chann … [kill] [force]
|
|
41
|
+
hermes 49001 ● running 18.0 512 — 44m hermes worker … [kill] [force]
|
|
42
|
+
ollama 50122 ● running 3.1 9210 14080 6h 02m ollama runner … [kill] [force]
|
|
43
|
+
```
|
|
44
|
+
(`+N` = child processes rolled up under the agent; CPU/mem/GPU are tree totals.)
|
|
45
|
+
|
|
46
|
+
> The schematic above shows the GPU column (NVIDIA only); the screenshot is a real
|
|
47
|
+
> run on Apple Silicon, where per-process GPU stats aren't available so that column
|
|
48
|
+
> is hidden. The UI auto-refreshes every 3s.
|
|
49
|
+
|
|
50
|
+
## What it does
|
|
51
|
+
|
|
52
|
+
- **One row per agent.** Agents are grouped by process tree — the spawned children
|
|
53
|
+
of an agent (inference subprocesses, MCP servers, helpers) are rolled up under it
|
|
54
|
+
with a `+N` badge instead of cluttering the list as separate rows.
|
|
55
|
+
- **Liveness** — green dot = running, red = zombie/dead. Status column shows the OS state.
|
|
56
|
+
- **Usage** — CPU %, resident memory (MB), GPU memory (MB, NVIDIA only), and uptime,
|
|
57
|
+
refreshed every 3s. **CPU/mem/GPU are tree totals** — the agent's true cost including
|
|
58
|
+
everything it spawned.
|
|
59
|
+
- **Kill the tree** — `kill` sends SIGTERM to the agent *and its children* (so spawned
|
|
60
|
+
helpers don't leak resources), `force` sends SIGKILL. SIGTERM auto-escalates to
|
|
61
|
+
SIGKILL after 3s. The confirm dialog tells you how many child processes will stop.
|
|
62
|
+
- **Trends, not just snapshots.** Each row has a CPU sparkline (last ~20 min, sampled
|
|
63
|
+
in the background even with no browser open), plus a **`hot 5m+`** badge when an
|
|
64
|
+
agent has been pegged ≥90% CPU for 5+ minutes, an **`idle 10m+`** badge when a
|
|
65
|
+
long-running agent has done nothing for 10+ minutes, and a **`churn ×N`** badge when
|
|
66
|
+
the same agent has died young 3+ times in 10 minutes — the states worth investigating
|
|
67
|
+
(runaway, possibly wedged, crash-looping under a supervisor). Churn is what hot/idle
|
|
68
|
+
can't see: a crash-looping process is a fresh pid every poll, so no per-process
|
|
69
|
+
window ever fills. A **`leak?`** badge fires when an agent's memory ratchets up
|
|
70
|
+
≥30% (and ≥128 MB) over 15 minutes without coming back down.
|
|
71
|
+
- **Alerts.** A dashboard only helps while you're looking at it. Add an `alerts:`
|
|
72
|
+
block to `agents.yaml` and any badge appearing runs your command (desktop
|
|
73
|
+
notification, Telegram bot, pager — anything) with the details in `$AUM_*` env
|
|
74
|
+
vars. Fires once per transition with a cooldown, never from the `list` CLI.
|
|
75
|
+
By default only `hot`/`churn`/`leak` alert — `idle` is the normal state of an
|
|
76
|
+
agent fleet that waits for work, so it's opt-in.
|
|
77
|
+
|
|
78
|
+
```yaml
|
|
79
|
+
alerts:
|
|
80
|
+
command: 'terminal-notifier -title agents -message "$AUM_MSG"'
|
|
81
|
+
cooldown: 600
|
|
82
|
+
flags: [hot, churn, leak] # the default; add idle to opt in
|
|
83
|
+
```
|
|
84
|
+
- **Prometheus `/metrics`.** Per-agent CPU/mem/instances/restarts and badge states
|
|
85
|
+
in text exposition format, aggregated per label (no pid-churn series bloat) —
|
|
86
|
+
point Grafana or any Prometheus scraper at `http://127.0.0.1:8765/metrics`.
|
|
87
|
+
- **Expand the tree.** Click the `+N` badge to unfold an agent's child processes
|
|
88
|
+
(per-child CPU/mem/command) — see what a kill would actually stop before clicking it.
|
|
89
|
+
- **Config hot-reload.** Edits to `agents.yaml` apply on the next poll, no restart.
|
|
90
|
+
A broken edit keeps the last good config and shows the parse error in the header.
|
|
91
|
+
- **`list` subcommand.** `agent-usage-manager list` (or `list --json`) prints a one-shot
|
|
92
|
+
table to stdout — no server, good for scripts and cron checks.
|
|
93
|
+
- **Kill-safe table.** Rows keep a stable order (sorted by label) and never reorder
|
|
94
|
+
while your pointer is over the table, so the kill button can't shift under your
|
|
95
|
+
cursor mid-click.
|
|
96
|
+
|
|
97
|
+
## Safety
|
|
98
|
+
|
|
99
|
+
This is the important part — a web page that can kill processes needs guardrails:
|
|
100
|
+
|
|
101
|
+
- **Allowlist only.** Only processes matching a pattern in `agents.yaml` are ever
|
|
102
|
+
listed *or* killable. The kill endpoint re-checks the match server-side before
|
|
103
|
+
sending any signal, so the dashboard can never be used to kill an arbitrary PID.
|
|
104
|
+
- **Protected patterns.** Anything matching `protect:` in `agents.yaml` — plus the
|
|
105
|
+
monitor's own process and PID 1 — shows a disabled, greyed-out kill button and is
|
|
106
|
+
refused server-side.
|
|
107
|
+
- **Secret redaction.** Command lines often carry tokens/keys in env vars or flags
|
|
108
|
+
(`FOO_TOKEN=...`, `--api-key ...`, `sk-...`, `ghp_...`, JWTs). The command column
|
|
109
|
+
redacts these to `***` before they ever reach the browser — safe to screenshot.
|
|
110
|
+
- **Browser guard (CSRF + DNS rebinding).** Binding to localhost doesn't keep
|
|
111
|
+
browsers out — any web page you visit can `fetch()` a localhost port. Requests
|
|
112
|
+
whose `Host` is a non-local DNS name are refused (DNS-rebinding guard), and a
|
|
113
|
+
kill request carrying a foreign `Origin` is refused (CSRF guard) — so a
|
|
114
|
+
malicious page can't kill your agents or read your process list. `curl` and
|
|
115
|
+
the dashboard itself are unaffected.
|
|
116
|
+
- **Bind local by default.** It listens on `127.0.0.1`. Don't expose it to a network
|
|
117
|
+
without putting auth in front of it (reverse proxy + basic auth, SSH tunnel, etc.) —
|
|
118
|
+
it has no built-in authentication.
|
|
119
|
+
|
|
120
|
+
## Limits & known issues
|
|
121
|
+
|
|
122
|
+
- **GPU column is NVIDIA-only.** Per-process GPU memory comes from
|
|
123
|
+
`nvidia-smi --query-compute-apps` — NVIDIA compute processes (CUDA), in
|
|
124
|
+
practice on Linux. AMD/Intel GPUs aren't read, graphics-only workloads don't
|
|
125
|
+
appear, and Apple Silicon has no per-process GPU accounting API at all, so
|
|
126
|
+
the column is hidden on Macs.
|
|
127
|
+
- **Supervision detection is launchd-only (macOS, user domain).** Root
|
|
128
|
+
`LaunchDaemons` aren't flagged — that needs a privileged
|
|
129
|
+
`launchctl print system/…`. On Linux, systemd-supervised services
|
|
130
|
+
(`Restart=always`) aren't detected either, so killing one looks like it
|
|
131
|
+
failed when systemd respawns it — use `systemctl stop` for those.
|
|
132
|
+
- **Same-user privileges only.** Signals are sent with the server's own
|
|
133
|
+
privileges. Agents running as another user (or root) are listed, but a kill
|
|
134
|
+
won't take (`killed: 0` in the response), and CPU/mem can read as 0 where
|
|
135
|
+
the OS denies access.
|
|
136
|
+
- **History is in-memory.** Sparklines and the `hot`/`idle` flags (~20 min
|
|
137
|
+
window) rebuild from scratch after a server restart.
|
|
138
|
+
- **Windows is untested.** Kill maps to `TerminateProcess` via psutil and may
|
|
139
|
+
work, but CI covers Linux + macOS only.
|
|
140
|
+
|
|
141
|
+
## Quick start
|
|
142
|
+
|
|
143
|
+
**Recommended — one command, nothing to install first:**
|
|
144
|
+
|
|
145
|
+
```bash
|
|
146
|
+
uvx agent-usage-manager
|
|
147
|
+
# then open http://127.0.0.1:8765 (it also opens automatically)
|
|
148
|
+
```
|
|
149
|
+
|
|
150
|
+
`uvx` fetches and runs it in one step — no separate install, no virtualenv, no
|
|
151
|
+
leftovers. Don't have [`uv`](https://github.com/astral-sh/uv) yet? One line:
|
|
152
|
+
|
|
153
|
+
```bash
|
|
154
|
+
curl -LsSf https://astral.sh/uv/install.sh | sh # macOS / Linux
|
|
155
|
+
# or: pip install uv
|
|
156
|
+
```
|
|
157
|
+
|
|
158
|
+
<details>
|
|
159
|
+
<summary>Other ways to install</summary>
|
|
160
|
+
|
|
161
|
+
```bash
|
|
162
|
+
pipx install agent-usage-manager # clean isolated global CLI (needs pipx)
|
|
163
|
+
|
|
164
|
+
pip install agent-usage-manager # universal; use inside a venv —
|
|
165
|
+
# system Python may refuse with
|
|
166
|
+
# "externally-managed-environment"
|
|
167
|
+
```
|
|
168
|
+
|
|
169
|
+
Then run `agent-usage-manager` (flags below).
|
|
170
|
+
</details>
|
|
171
|
+
|
|
172
|
+
From a clone (for hacking on it):
|
|
173
|
+
|
|
174
|
+
```bash
|
|
175
|
+
git clone <this-repo> && cd agent-usage-manager
|
|
176
|
+
./run.sh # venv + editable install, serves on :8765
|
|
177
|
+
```
|
|
178
|
+
|
|
179
|
+
It opens the dashboard in your browser automatically. Flags: `--host`, `--port`,
|
|
180
|
+
`--config /path/to/agents.yaml`, `--no-browser` (for headless/server use).
|
|
181
|
+
|
|
182
|
+
## Configure which processes are "agents"
|
|
183
|
+
|
|
184
|
+
Edit `agents.yaml`:
|
|
185
|
+
|
|
186
|
+
```yaml
|
|
187
|
+
agents:
|
|
188
|
+
- label: openclaw # shown as the badge in the UI
|
|
189
|
+
match: openclaw # case-insensitive substring of the command line
|
|
190
|
+
- label: hermes
|
|
191
|
+
match: hermes
|
|
192
|
+
- label: claude-code
|
|
193
|
+
match: "claude(\\s|$|-code)"
|
|
194
|
+
regex: true # treat `match` as a regex instead of substring
|
|
195
|
+
|
|
196
|
+
protect: # matched + listed, but never killable
|
|
197
|
+
- uvicorn
|
|
198
|
+
|
|
199
|
+
ignore: # never an agent: not listed, not killable
|
|
200
|
+
- crashpad # incidental processes that share a name/bundle
|
|
201
|
+
- shipit # path with a real agent (crash handlers,
|
|
202
|
+
- kiro-cli-term # auto-updaters, integrated-terminal shells, …)
|
|
203
|
+
```
|
|
204
|
+
|
|
205
|
+
A process matches if the pattern hits its **executable basename + first few
|
|
206
|
+
arguments** — deliberately not the whole command line, so a long embedded arg
|
|
207
|
+
(e.g. a system prompt mentioning "claude") can't misclassify a wrapper. On macOS
|
|
208
|
+
the outermost `.app` **bundle name** is also included, so GUI agents that launch
|
|
209
|
+
a generically-named binary (Kiro.app → `Electron`) are still matched by app name.
|
|
210
|
+
|
|
211
|
+
`protect:` keeps a matched process listed but refuses to kill it; `ignore:`
|
|
212
|
+
drops it from agent classification entirely.
|
|
213
|
+
|
|
214
|
+
**Which `agents.yaml` is used** — resolved once at startup, first hit wins:
|
|
215
|
+
|
|
216
|
+
1. `AGENTS_CONFIG=/path/to/agents.yaml` env var (the `--config` flag sets this)
|
|
217
|
+
2. `./agents.yaml` in the directory you launched from
|
|
218
|
+
3. the default bundled with the package
|
|
219
|
+
|
|
220
|
+
The dashboard header (and `list --json`) shows the resolved path, so you can
|
|
221
|
+
always see which file is live. Hot-reload watches that one file. An
|
|
222
|
+
`AGENTS_CONFIG` path that doesn't exist is an error at startup, not a silent
|
|
223
|
+
fallback.
|
|
224
|
+
|
|
225
|
+
## launchd-supervised agents (macOS)
|
|
226
|
+
|
|
227
|
+
Some agents run as **launchd services** (a `~/Library/LaunchAgents/*.plist`, or
|
|
228
|
+
anything started by `brew services`). If such a job sets `KeepAlive`, a signal
|
|
229
|
+
can't stop it: the process dies, launchd immediately respawns it under a new PID,
|
|
230
|
+
and the dashboard's "kill" looks like it silently failed.
|
|
231
|
+
|
|
232
|
+
The dashboard detects these (via `launchctl list`) and marks them with a
|
|
233
|
+
**`launchd`** badge. Instead of dead-end kill/force buttons it shows the command
|
|
234
|
+
that actually stops the job — click to copy:
|
|
235
|
+
|
|
236
|
+
```sh
|
|
237
|
+
launchctl bootout gui/<uid>/<label> # stop now
|
|
238
|
+
launchctl disable gui/<uid>/<label> # …and don't auto-start at login
|
|
239
|
+
```
|
|
240
|
+
|
|
241
|
+
The kill endpoint refuses signals for these jobs (HTTP 409) and returns the same
|
|
242
|
+
guidance, so the API never lies about a kill that won't stick. The message is
|
|
243
|
+
tailored to the job: `KeepAlive` jobs are told a signal won't stick at all;
|
|
244
|
+
`RunAtLoad`-only jobs are told a signal works now but the job restarts at next
|
|
245
|
+
login. *Limitation:* root `LaunchDaemons` aren't flagged — see
|
|
246
|
+
[Limits & known issues](#limits--known-issues).
|
|
247
|
+
|
|
248
|
+
## GPU notes
|
|
249
|
+
|
|
250
|
+
Per-process GPU memory comes from `nvidia-smi` when it's on `PATH` (Linux / NVIDIA).
|
|
251
|
+
**Apple Silicon has no per-process GPU accounting API**, so the GPU column stays blank
|
|
252
|
+
on Macs — CPU and memory are the meaningful resource signals there.
|
|
253
|
+
|
|
254
|
+
## API
|
|
255
|
+
|
|
256
|
+
- `GET /api/agents` → `{ agents: [...], host, cpu_count, mem_total_mb, mem_used_pct, config_path, config_error, ts }`
|
|
257
|
+
— each agent includes `trend` (recent CPU samples) and `flag` (`"hot"` / `"idle"` / `null`)
|
|
258
|
+
- `GET /api/tree/{pid}` → the agent's process subtree (per-child pid/name/cpu/mem/cmdline);
|
|
259
|
+
only works on recognized agents, same authorization as kill
|
|
260
|
+
- `POST /api/kill/{pid}?force=false` → SIGTERM (or SIGKILL with `force=true`)
|
|
261
|
+
|
|
262
|
+
## Run as a service
|
|
263
|
+
|
|
264
|
+
Linux (systemd), `~/.config/systemd/user/agent-usage-manager.service`:
|
|
265
|
+
|
|
266
|
+
```ini
|
|
267
|
+
[Unit]
|
|
268
|
+
Description=agent usage manager
|
|
269
|
+
[Service]
|
|
270
|
+
ExecStart=%h/agent-usage-manager/.venv/bin/uvicorn app:app --port 8765
|
|
271
|
+
WorkingDirectory=%h/agent-usage-manager
|
|
272
|
+
Restart=on-failure
|
|
273
|
+
[Install]
|
|
274
|
+
WantedBy=default.target
|
|
275
|
+
```
|
|
276
|
+
|
|
277
|
+
```bash
|
|
278
|
+
systemctl --user enable --now agent-usage-manager
|
|
279
|
+
```
|
|
280
|
+
|
|
281
|
+
## Development
|
|
282
|
+
|
|
283
|
+
```bash
|
|
284
|
+
git clone https://github.com/minglong51/agent-usage-manager && cd agent-usage-manager
|
|
285
|
+
pip install -e ".[dev]"
|
|
286
|
+
pytest -q
|
|
287
|
+
```
|
|
288
|
+
|
|
289
|
+
CI runs the test suite on Linux + macOS (Python 3.9 and 3.12) on every push and PR.
|
|
290
|
+
Cross-platform note: kill uses psutil's `terminate()`/`kill()`, which map to
|
|
291
|
+
SIGTERM/SIGKILL on POSIX and TerminateProcess on Windows.
|
|
292
|
+
|
|
293
|
+
## Troubleshooting
|
|
294
|
+
|
|
295
|
+
- **`pip install` fails building psutil** — no prebuilt wheel for your
|
|
296
|
+
Python/platform, so pip compiles it: you need a C toolchain and Python
|
|
297
|
+
headers (`xcode-select --install` on macOS; `apt install gcc python3-dev`
|
|
298
|
+
on Debian/Ubuntu). Or skip the problem with `uvx agent-usage-manager`.
|
|
299
|
+
- **Dashboard is empty / "No matching agents running"** — first check which
|
|
300
|
+
config was picked up (resolution order above; the header shows the resolved
|
|
301
|
+
path). Then remember matching is against the executable basename + first few
|
|
302
|
+
arguments, not the full command line — a pattern that only appears deep in
|
|
303
|
+
the args won't match.
|
|
304
|
+
- **Kill "doesn't work" — the agent comes back under a new PID** — it's
|
|
305
|
+
supervised. On macOS the row gets a `launchd` badge with the `launchctl
|
|
306
|
+
bootout` command that actually stops it; on Linux, systemd services aren't
|
|
307
|
+
detected (see limits) — `systemctl stop` them. A `killed: 0` in the kill
|
|
308
|
+
response means nothing was actually signaled (e.g. the agent runs as
|
|
309
|
+
another user).
|
|
310
|
+
- **HTTP 403 on every request** — the DNS-rebinding guard refuses non-local
|
|
311
|
+
hostnames. Use `http://127.0.0.1:8765` (or a bare IP) instead of a custom
|
|
312
|
+
DNS name pointing at the box.
|
|
313
|
+
|
|
314
|
+
## License
|
|
315
|
+
|
|
316
|
+
MIT
|
|
@@ -0,0 +1,297 @@
|
|
|
1
|
+
# agent-usage-manager
|
|
2
|
+
|
|
3
|
+
A tiny, single-file web dashboard for **headless AI agents** running on a machine —
|
|
4
|
+
OpenClaw, Hermes, Claude Code, Ollama, vLLM, llama.cpp, or anything you name. It
|
|
5
|
+
shows which agents are alive and what they're costing you (CPU, memory, GPU), and
|
|
6
|
+
gives you a **kill button** per agent. Think `htop`, scoped to just your agents —
|
|
7
|
+
the [screenshot below](docs/dashboard.png) is a real run on a fleet node.
|
|
8
|
+
|
|
9
|
+
No database, no auth layer, no dependencies beyond FastAPI + psutil. Runs on
|
|
10
|
+
macOS and Linux. Meant to be cloned, configured, and run on any node in a fleet.
|
|
11
|
+
|
|
12
|
+

|
|
13
|
+
|
|
14
|
+
*A real run: ten agents grouped by process tree (`+N` = children rolled up),
|
|
15
|
+
per-agent CPU/memory/uptime, launchd-supervised jobs flagged, and a kill button
|
|
16
|
+
per row.*
|
|
17
|
+
|
|
18
|
+
```
|
|
19
|
+
AGENT PID STATUS CPU % MEM MB GPU MB UPTIME COMMAND ┆
|
|
20
|
+
openclaw +3 48213 ● running 62.4 1840 7320 2h 11m openclaw serve … [kill] [force]
|
|
21
|
+
claude-code +9 73590 ● running 97.4 7630 — 1h 02m claude --chann … [kill] [force]
|
|
22
|
+
hermes 49001 ● running 18.0 512 — 44m hermes worker … [kill] [force]
|
|
23
|
+
ollama 50122 ● running 3.1 9210 14080 6h 02m ollama runner … [kill] [force]
|
|
24
|
+
```
|
|
25
|
+
(`+N` = child processes rolled up under the agent; CPU/mem/GPU are tree totals.)
|
|
26
|
+
|
|
27
|
+
> The schematic above shows the GPU column (NVIDIA only); the screenshot is a real
|
|
28
|
+
> run on Apple Silicon, where per-process GPU stats aren't available so that column
|
|
29
|
+
> is hidden. The UI auto-refreshes every 3s.
|
|
30
|
+
|
|
31
|
+
## What it does
|
|
32
|
+
|
|
33
|
+
- **One row per agent.** Agents are grouped by process tree — the spawned children
|
|
34
|
+
of an agent (inference subprocesses, MCP servers, helpers) are rolled up under it
|
|
35
|
+
with a `+N` badge instead of cluttering the list as separate rows.
|
|
36
|
+
- **Liveness** — green dot = running, red = zombie/dead. Status column shows the OS state.
|
|
37
|
+
- **Usage** — CPU %, resident memory (MB), GPU memory (MB, NVIDIA only), and uptime,
|
|
38
|
+
refreshed every 3s. **CPU/mem/GPU are tree totals** — the agent's true cost including
|
|
39
|
+
everything it spawned.
|
|
40
|
+
- **Kill the tree** — `kill` sends SIGTERM to the agent *and its children* (so spawned
|
|
41
|
+
helpers don't leak resources), `force` sends SIGKILL. SIGTERM auto-escalates to
|
|
42
|
+
SIGKILL after 3s. The confirm dialog tells you how many child processes will stop.
|
|
43
|
+
- **Trends, not just snapshots.** Each row has a CPU sparkline (last ~20 min, sampled
|
|
44
|
+
in the background even with no browser open), plus a **`hot 5m+`** badge when an
|
|
45
|
+
agent has been pegged ≥90% CPU for 5+ minutes, an **`idle 10m+`** badge when a
|
|
46
|
+
long-running agent has done nothing for 10+ minutes, and a **`churn ×N`** badge when
|
|
47
|
+
the same agent has died young 3+ times in 10 minutes — the states worth investigating
|
|
48
|
+
(runaway, possibly wedged, crash-looping under a supervisor). Churn is what hot/idle
|
|
49
|
+
can't see: a crash-looping process is a fresh pid every poll, so no per-process
|
|
50
|
+
window ever fills. A **`leak?`** badge fires when an agent's memory ratchets up
|
|
51
|
+
≥30% (and ≥128 MB) over 15 minutes without coming back down.
|
|
52
|
+
- **Alerts.** A dashboard only helps while you're looking at it. Add an `alerts:`
|
|
53
|
+
block to `agents.yaml` and any badge appearing runs your command (desktop
|
|
54
|
+
notification, Telegram bot, pager — anything) with the details in `$AUM_*` env
|
|
55
|
+
vars. Fires once per transition with a cooldown, never from the `list` CLI.
|
|
56
|
+
By default only `hot`/`churn`/`leak` alert — `idle` is the normal state of an
|
|
57
|
+
agent fleet that waits for work, so it's opt-in.
|
|
58
|
+
|
|
59
|
+
```yaml
|
|
60
|
+
alerts:
|
|
61
|
+
command: 'terminal-notifier -title agents -message "$AUM_MSG"'
|
|
62
|
+
cooldown: 600
|
|
63
|
+
flags: [hot, churn, leak] # the default; add idle to opt in
|
|
64
|
+
```
|
|
65
|
+
- **Prometheus `/metrics`.** Per-agent CPU/mem/instances/restarts and badge states
|
|
66
|
+
in text exposition format, aggregated per label (no pid-churn series bloat) —
|
|
67
|
+
point Grafana or any Prometheus scraper at `http://127.0.0.1:8765/metrics`.
|
|
68
|
+
- **Expand the tree.** Click the `+N` badge to unfold an agent's child processes
|
|
69
|
+
(per-child CPU/mem/command) — see what a kill would actually stop before clicking it.
|
|
70
|
+
- **Config hot-reload.** Edits to `agents.yaml` apply on the next poll, no restart.
|
|
71
|
+
A broken edit keeps the last good config and shows the parse error in the header.
|
|
72
|
+
- **`list` subcommand.** `agent-usage-manager list` (or `list --json`) prints a one-shot
|
|
73
|
+
table to stdout — no server, good for scripts and cron checks.
|
|
74
|
+
- **Kill-safe table.** Rows keep a stable order (sorted by label) and never reorder
|
|
75
|
+
while your pointer is over the table, so the kill button can't shift under your
|
|
76
|
+
cursor mid-click.
|
|
77
|
+
|
|
78
|
+
## Safety
|
|
79
|
+
|
|
80
|
+
This is the important part — a web page that can kill processes needs guardrails:
|
|
81
|
+
|
|
82
|
+
- **Allowlist only.** Only processes matching a pattern in `agents.yaml` are ever
|
|
83
|
+
listed *or* killable. The kill endpoint re-checks the match server-side before
|
|
84
|
+
sending any signal, so the dashboard can never be used to kill an arbitrary PID.
|
|
85
|
+
- **Protected patterns.** Anything matching `protect:` in `agents.yaml` — plus the
|
|
86
|
+
monitor's own process and PID 1 — shows a disabled, greyed-out kill button and is
|
|
87
|
+
refused server-side.
|
|
88
|
+
- **Secret redaction.** Command lines often carry tokens/keys in env vars or flags
|
|
89
|
+
(`FOO_TOKEN=...`, `--api-key ...`, `sk-...`, `ghp_...`, JWTs). The command column
|
|
90
|
+
redacts these to `***` before they ever reach the browser — safe to screenshot.
|
|
91
|
+
- **Browser guard (CSRF + DNS rebinding).** Binding to localhost doesn't keep
|
|
92
|
+
browsers out — any web page you visit can `fetch()` a localhost port. Requests
|
|
93
|
+
whose `Host` is a non-local DNS name are refused (DNS-rebinding guard), and a
|
|
94
|
+
kill request carrying a foreign `Origin` is refused (CSRF guard) — so a
|
|
95
|
+
malicious page can't kill your agents or read your process list. `curl` and
|
|
96
|
+
the dashboard itself are unaffected.
|
|
97
|
+
- **Bind local by default.** It listens on `127.0.0.1`. Don't expose it to a network
|
|
98
|
+
without putting auth in front of it (reverse proxy + basic auth, SSH tunnel, etc.) —
|
|
99
|
+
it has no built-in authentication.
|
|
100
|
+
|
|
101
|
+
## Limits & known issues
|
|
102
|
+
|
|
103
|
+
- **GPU column is NVIDIA-only.** Per-process GPU memory comes from
|
|
104
|
+
`nvidia-smi --query-compute-apps` — NVIDIA compute processes (CUDA), in
|
|
105
|
+
practice on Linux. AMD/Intel GPUs aren't read, graphics-only workloads don't
|
|
106
|
+
appear, and Apple Silicon has no per-process GPU accounting API at all, so
|
|
107
|
+
the column is hidden on Macs.
|
|
108
|
+
- **Supervision detection is launchd-only (macOS, user domain).** Root
|
|
109
|
+
`LaunchDaemons` aren't flagged — that needs a privileged
|
|
110
|
+
`launchctl print system/…`. On Linux, systemd-supervised services
|
|
111
|
+
(`Restart=always`) aren't detected either, so killing one looks like it
|
|
112
|
+
failed when systemd respawns it — use `systemctl stop` for those.
|
|
113
|
+
- **Same-user privileges only.** Signals are sent with the server's own
|
|
114
|
+
privileges. Agents running as another user (or root) are listed, but a kill
|
|
115
|
+
won't take (`killed: 0` in the response), and CPU/mem can read as 0 where
|
|
116
|
+
the OS denies access.
|
|
117
|
+
- **History is in-memory.** Sparklines and the `hot`/`idle` flags (~20 min
|
|
118
|
+
window) rebuild from scratch after a server restart.
|
|
119
|
+
- **Windows is untested.** Kill maps to `TerminateProcess` via psutil and may
|
|
120
|
+
work, but CI covers Linux + macOS only.
|
|
121
|
+
|
|
122
|
+
## Quick start
|
|
123
|
+
|
|
124
|
+
**Recommended — one command, nothing to install first:**
|
|
125
|
+
|
|
126
|
+
```bash
|
|
127
|
+
uvx agent-usage-manager
|
|
128
|
+
# then open http://127.0.0.1:8765 (it also opens automatically)
|
|
129
|
+
```
|
|
130
|
+
|
|
131
|
+
`uvx` fetches and runs it in one step — no separate install, no virtualenv, no
|
|
132
|
+
leftovers. Don't have [`uv`](https://github.com/astral-sh/uv) yet? One line:
|
|
133
|
+
|
|
134
|
+
```bash
|
|
135
|
+
curl -LsSf https://astral.sh/uv/install.sh | sh # macOS / Linux
|
|
136
|
+
# or: pip install uv
|
|
137
|
+
```
|
|
138
|
+
|
|
139
|
+
<details>
|
|
140
|
+
<summary>Other ways to install</summary>
|
|
141
|
+
|
|
142
|
+
```bash
|
|
143
|
+
pipx install agent-usage-manager # clean isolated global CLI (needs pipx)
|
|
144
|
+
|
|
145
|
+
pip install agent-usage-manager # universal; use inside a venv —
|
|
146
|
+
# system Python may refuse with
|
|
147
|
+
# "externally-managed-environment"
|
|
148
|
+
```
|
|
149
|
+
|
|
150
|
+
Then run `agent-usage-manager` (flags below).
|
|
151
|
+
</details>
|
|
152
|
+
|
|
153
|
+
From a clone (for hacking on it):
|
|
154
|
+
|
|
155
|
+
```bash
|
|
156
|
+
git clone <this-repo> && cd agent-usage-manager
|
|
157
|
+
./run.sh # venv + editable install, serves on :8765
|
|
158
|
+
```
|
|
159
|
+
|
|
160
|
+
It opens the dashboard in your browser automatically. Flags: `--host`, `--port`,
|
|
161
|
+
`--config /path/to/agents.yaml`, `--no-browser` (for headless/server use).
|
|
162
|
+
|
|
163
|
+
## Configure which processes are "agents"
|
|
164
|
+
|
|
165
|
+
Edit `agents.yaml`:
|
|
166
|
+
|
|
167
|
+
```yaml
|
|
168
|
+
agents:
|
|
169
|
+
- label: openclaw # shown as the badge in the UI
|
|
170
|
+
match: openclaw # case-insensitive substring of the command line
|
|
171
|
+
- label: hermes
|
|
172
|
+
match: hermes
|
|
173
|
+
- label: claude-code
|
|
174
|
+
match: "claude(\\s|$|-code)"
|
|
175
|
+
regex: true # treat `match` as a regex instead of substring
|
|
176
|
+
|
|
177
|
+
protect: # matched + listed, but never killable
|
|
178
|
+
- uvicorn
|
|
179
|
+
|
|
180
|
+
ignore: # never an agent: not listed, not killable
|
|
181
|
+
- crashpad # incidental processes that share a name/bundle
|
|
182
|
+
- shipit # path with a real agent (crash handlers,
|
|
183
|
+
- kiro-cli-term # auto-updaters, integrated-terminal shells, …)
|
|
184
|
+
```
|
|
185
|
+
|
|
186
|
+
A process matches if the pattern hits its **executable basename + first few
|
|
187
|
+
arguments** — deliberately not the whole command line, so a long embedded arg
|
|
188
|
+
(e.g. a system prompt mentioning "claude") can't misclassify a wrapper. On macOS
|
|
189
|
+
the outermost `.app` **bundle name** is also included, so GUI agents that launch
|
|
190
|
+
a generically-named binary (Kiro.app → `Electron`) are still matched by app name.
|
|
191
|
+
|
|
192
|
+
`protect:` keeps a matched process listed but refuses to kill it; `ignore:`
|
|
193
|
+
drops it from agent classification entirely.
|
|
194
|
+
|
|
195
|
+
**Which `agents.yaml` is used** — resolved once at startup, first hit wins:
|
|
196
|
+
|
|
197
|
+
1. `AGENTS_CONFIG=/path/to/agents.yaml` env var (the `--config` flag sets this)
|
|
198
|
+
2. `./agents.yaml` in the directory you launched from
|
|
199
|
+
3. the default bundled with the package
|
|
200
|
+
|
|
201
|
+
The dashboard header (and `list --json`) shows the resolved path, so you can
|
|
202
|
+
always see which file is live. Hot-reload watches that one file. An
|
|
203
|
+
`AGENTS_CONFIG` path that doesn't exist is an error at startup, not a silent
|
|
204
|
+
fallback.
|
|
205
|
+
|
|
206
|
+
## launchd-supervised agents (macOS)
|
|
207
|
+
|
|
208
|
+
Some agents run as **launchd services** (a `~/Library/LaunchAgents/*.plist`, or
|
|
209
|
+
anything started by `brew services`). If such a job sets `KeepAlive`, a signal
|
|
210
|
+
can't stop it: the process dies, launchd immediately respawns it under a new PID,
|
|
211
|
+
and the dashboard's "kill" looks like it silently failed.
|
|
212
|
+
|
|
213
|
+
The dashboard detects these (via `launchctl list`) and marks them with a
|
|
214
|
+
**`launchd`** badge. Instead of dead-end kill/force buttons it shows the command
|
|
215
|
+
that actually stops the job — click to copy:
|
|
216
|
+
|
|
217
|
+
```sh
|
|
218
|
+
launchctl bootout gui/<uid>/<label> # stop now
|
|
219
|
+
launchctl disable gui/<uid>/<label> # …and don't auto-start at login
|
|
220
|
+
```
|
|
221
|
+
|
|
222
|
+
The kill endpoint refuses signals for these jobs (HTTP 409) and returns the same
|
|
223
|
+
guidance, so the API never lies about a kill that won't stick. The message is
|
|
224
|
+
tailored to the job: `KeepAlive` jobs are told a signal won't stick at all;
|
|
225
|
+
`RunAtLoad`-only jobs are told a signal works now but the job restarts at next
|
|
226
|
+
login. *Limitation:* root `LaunchDaemons` aren't flagged — see
|
|
227
|
+
[Limits & known issues](#limits--known-issues).
|
|
228
|
+
|
|
229
|
+
## GPU notes
|
|
230
|
+
|
|
231
|
+
Per-process GPU memory comes from `nvidia-smi` when it's on `PATH` (Linux / NVIDIA).
|
|
232
|
+
**Apple Silicon has no per-process GPU accounting API**, so the GPU column stays blank
|
|
233
|
+
on Macs — CPU and memory are the meaningful resource signals there.
|
|
234
|
+
|
|
235
|
+
## API
|
|
236
|
+
|
|
237
|
+
- `GET /api/agents` → `{ agents: [...], host, cpu_count, mem_total_mb, mem_used_pct, config_path, config_error, ts }`
|
|
238
|
+
— each agent includes `trend` (recent CPU samples) and `flag` (`"hot"` / `"idle"` / `null`)
|
|
239
|
+
- `GET /api/tree/{pid}` → the agent's process subtree (per-child pid/name/cpu/mem/cmdline);
|
|
240
|
+
only works on recognized agents, same authorization as kill
|
|
241
|
+
- `POST /api/kill/{pid}?force=false` → SIGTERM (or SIGKILL with `force=true`)
|
|
242
|
+
|
|
243
|
+
## Run as a service
|
|
244
|
+
|
|
245
|
+
Linux (systemd), `~/.config/systemd/user/agent-usage-manager.service`:
|
|
246
|
+
|
|
247
|
+
```ini
|
|
248
|
+
[Unit]
|
|
249
|
+
Description=agent usage manager
|
|
250
|
+
[Service]
|
|
251
|
+
ExecStart=%h/agent-usage-manager/.venv/bin/uvicorn app:app --port 8765
|
|
252
|
+
WorkingDirectory=%h/agent-usage-manager
|
|
253
|
+
Restart=on-failure
|
|
254
|
+
[Install]
|
|
255
|
+
WantedBy=default.target
|
|
256
|
+
```
|
|
257
|
+
|
|
258
|
+
```bash
|
|
259
|
+
systemctl --user enable --now agent-usage-manager
|
|
260
|
+
```
|
|
261
|
+
|
|
262
|
+
## Development
|
|
263
|
+
|
|
264
|
+
```bash
|
|
265
|
+
git clone https://github.com/minglong51/agent-usage-manager && cd agent-usage-manager
|
|
266
|
+
pip install -e ".[dev]"
|
|
267
|
+
pytest -q
|
|
268
|
+
```
|
|
269
|
+
|
|
270
|
+
CI runs the test suite on Linux + macOS (Python 3.9 and 3.12) on every push and PR.
|
|
271
|
+
Cross-platform note: kill uses psutil's `terminate()`/`kill()`, which map to
|
|
272
|
+
SIGTERM/SIGKILL on POSIX and TerminateProcess on Windows.
|
|
273
|
+
|
|
274
|
+
## Troubleshooting
|
|
275
|
+
|
|
276
|
+
- **`pip install` fails building psutil** — no prebuilt wheel for your
|
|
277
|
+
Python/platform, so pip compiles it: you need a C toolchain and Python
|
|
278
|
+
headers (`xcode-select --install` on macOS; `apt install gcc python3-dev`
|
|
279
|
+
on Debian/Ubuntu). Or skip the problem with `uvx agent-usage-manager`.
|
|
280
|
+
- **Dashboard is empty / "No matching agents running"** — first check which
|
|
281
|
+
config was picked up (resolution order above; the header shows the resolved
|
|
282
|
+
path). Then remember matching is against the executable basename + first few
|
|
283
|
+
arguments, not the full command line — a pattern that only appears deep in
|
|
284
|
+
the args won't match.
|
|
285
|
+
- **Kill "doesn't work" — the agent comes back under a new PID** — it's
|
|
286
|
+
supervised. On macOS the row gets a `launchd` badge with the `launchctl
|
|
287
|
+
bootout` command that actually stops it; on Linux, systemd services aren't
|
|
288
|
+
detected (see limits) — `systemctl stop` them. A `killed: 0` in the kill
|
|
289
|
+
response means nothing was actually signaled (e.g. the agent runs as
|
|
290
|
+
another user).
|
|
291
|
+
- **HTTP 403 on every request** — the DNS-rebinding guard refuses non-local
|
|
292
|
+
hostnames. Use `http://127.0.0.1:8765` (or a bare IP) instead of a custom
|
|
293
|
+
DNS name pointing at the box.
|
|
294
|
+
|
|
295
|
+
## License
|
|
296
|
+
|
|
297
|
+
MIT
|