agent-usage-manager 0.1.3__tar.gz → 0.2.2__tar.gz

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (26) hide show
  1. {agent_usage_manager-0.1.3 → agent_usage_manager-0.2.2}/.github/workflows/ci.yml +3 -0
  2. {agent_usage_manager-0.1.3 → agent_usage_manager-0.2.2}/.gitignore +1 -0
  3. agent_usage_manager-0.2.2/PKG-INFO +415 -0
  4. agent_usage_manager-0.2.2/README.md +396 -0
  5. {agent_usage_manager-0.1.3 → agent_usage_manager-0.2.2}/agent_usage_manager/agents.yaml +42 -0
  6. agent_usage_manager-0.2.2/agent_usage_manager/app.py +1376 -0
  7. agent_usage_manager-0.2.2/agent_usage_manager/cli.py +137 -0
  8. agent_usage_manager-0.2.2/agent_usage_manager/static/index.html +360 -0
  9. agent_usage_manager-0.2.2/docs/dashboard.png +0 -0
  10. agent_usage_manager-0.2.2/docs/design/HLD.md +115 -0
  11. agent_usage_manager-0.2.2/docs/design/LLD.md +327 -0
  12. {agent_usage_manager-0.1.3 → agent_usage_manager-0.2.2}/pyproject.toml +1 -1
  13. agent_usage_manager-0.2.2/tests/test_smoke.py +413 -0
  14. agent_usage_manager-0.2.2/tests/test_synthetic.py +317 -0
  15. {agent_usage_manager-0.1.3 → agent_usage_manager-0.2.2}/uv.lock +1 -1
  16. agent_usage_manager-0.1.3/PKG-INFO +0 -199
  17. agent_usage_manager-0.1.3/README.md +0 -180
  18. agent_usage_manager-0.1.3/agent_usage_manager/app.py +0 -570
  19. agent_usage_manager-0.1.3/agent_usage_manager/cli.py +0 -43
  20. agent_usage_manager-0.1.3/agent_usage_manager/static/index.html +0 -205
  21. agent_usage_manager-0.1.3/tests/test_smoke.py +0 -92
  22. {agent_usage_manager-0.1.3 → agent_usage_manager-0.2.2}/LICENSE +0 -0
  23. {agent_usage_manager-0.1.3 → agent_usage_manager-0.2.2}/agent_usage_manager/__init__.py +0 -0
  24. {agent_usage_manager-0.1.3 → agent_usage_manager-0.2.2}/demo.tape +0 -0
  25. {agent_usage_manager-0.1.3 → agent_usage_manager-0.2.2}/requirements.txt +0 -0
  26. {agent_usage_manager-0.1.3 → agent_usage_manager-0.2.2}/run.sh +0 -0
@@ -5,6 +5,9 @@ on:
5
5
  branches: [main]
6
6
  pull_request:
7
7
 
8
+ permissions:
9
+ contents: read
10
+
8
11
  jobs:
9
12
  test:
10
13
  strategy:
@@ -7,3 +7,4 @@ agents.local.yaml
7
7
  dist/
8
8
  build/
9
9
  *.egg-info/
10
+ .gstack/
@@ -0,0 +1,415 @@
1
+ Metadata-Version: 2.4
2
+ Name: agent-usage-manager
3
+ Version: 0.2.2
4
+ Summary: htop for AI agents — liveness, CPU/mem/GPU usage, and a kill switch for headless agents (openclaw, hermes, ollama, vllm, claude-code).
5
+ Project-URL: Homepage, https://github.com/minglong51/agent-usage-manager
6
+ Project-URL: Repository, https://github.com/minglong51/agent-usage-manager
7
+ License: MIT
8
+ License-File: LICENSE
9
+ Keywords: ai-agents,gpu,llm,monitoring,observability,ollama,vllm
10
+ Requires-Python: >=3.9
11
+ Requires-Dist: fastapi>=0.110
12
+ Requires-Dist: psutil>=5.9
13
+ Requires-Dist: pyyaml>=6.0
14
+ Requires-Dist: uvicorn[standard]>=0.27
15
+ Provides-Extra: dev
16
+ Requires-Dist: httpx>=0.27; extra == 'dev'
17
+ Requires-Dist: pytest>=7; extra == 'dev'
18
+ Description-Content-Type: text/markdown
19
+
20
+ # agent-usage-manager
21
+
22
+ A tiny, single-file web dashboard for **headless AI agents** running on a machine —
23
+ OpenClaw, Hermes, Claude Code, Ollama, vLLM, llama.cpp, or anything you name. It
24
+ shows which agents are alive and what they're costing you (CPU, memory, GPU), and
25
+ gives you a **kill button** per agent. Think `htop`, scoped to just your agents —
26
+ the [screenshot below](docs/dashboard.png) is a real run on a fleet node.
27
+
28
+ No database, no auth framework (one static token file gates the kill switch),
29
+ no dependencies beyond FastAPI + psutil. Runs on macOS and Linux. It is a
30
+ per-node monitor and guarded local control panel: fleet schedulers may consume
31
+ its read-only telemetry, but should own their own scheduling and actuation.
32
+
33
+ ![agent-usage-manager — live dashboard](docs/dashboard.png)
34
+
35
+ *A real run: ten agents grouped by process tree (`+N` = children rolled up),
36
+ per-agent CPU/memory/uptime, launchd-supervised jobs flagged, and a kill button
37
+ per row.*
38
+
39
+ ```
40
+ AGENT PID STATUS CPU % MEM MB GPU MB UPTIME COMMAND ┆
41
+ openclaw +3 48213 ● running 62.4 1840 7320 2h 11m openclaw serve … [kill] [force]
42
+ claude-code +9 73590 ● running 97.4 7630 — 1h 02m claude --chann … [kill] [force]
43
+ hermes 49001 ● running 18.0 512 — 44m hermes worker … [kill] [force]
44
+ ollama 50122 ● running 3.1 9210 14080 6h 02m ollama runner … [kill] [force]
45
+ ```
46
+ (`+N` = child processes rolled up under the agent; CPU/mem/GPU are tree totals.)
47
+
48
+ > The schematic above shows the GPU column (NVIDIA only); the screenshot is a real
49
+ > run on Apple Silicon, where per-process GPU stats aren't available so that column
50
+ > is hidden. The UI auto-refreshes every 3s.
51
+
52
+ ## What it does
53
+
54
+ - **One row per agent.** Agents are grouped by process tree — the spawned children
55
+ of an agent (inference subprocesses, MCP servers, helpers) are rolled up under it
56
+ with a `+N` badge instead of cluttering the list as separate rows.
57
+ - **Liveness** — green dot = running, red = zombie/dead. Status column shows the OS state.
58
+ - **Usage** — CPU %, resident memory (MB), GPU memory (MB, NVIDIA only), and uptime,
59
+ refreshed every 3s. **CPU/mem/GPU are tree totals** — the agent's true cost including
60
+ everything it spawned.
61
+ - **Kill the tree** — `kill` sends SIGTERM to the agent *and its children* (so spawned
62
+ helpers don't leak resources), `force` sends SIGKILL. SIGTERM auto-escalates to
63
+ SIGKILL after 3s. The confirm dialog tells you how many child processes will stop.
64
+ - **Trends, not just snapshots.** Each row has a CPU sparkline (last ~20 min, sampled
65
+ in the background even with no browser open), plus a **`hot 5m+`** badge when an
66
+ agent has been pegged ≥90% CPU for 5+ minutes, an **`idle 10m+`** badge when a
67
+ long-running agent has done nothing for 10+ minutes, and a **`churn ×N`** badge when
68
+ the same agent has died young 3+ times in 10 minutes — the states worth investigating
69
+ (runaway, possibly wedged, crash-looping under a supervisor). Churn is what hot/idle
70
+ can't see: a crash-looping process is a fresh pid every poll, so no per-process
71
+ window ever fills. A **`leak?`** badge fires when an agent's memory ratchets up
72
+ ≥30% (and ≥128 MB) over 15 minutes without coming back down.
73
+ - **Alerts.** A dashboard only helps while you're looking at it. Add an `alerts:`
74
+ block to `agents.yaml` and any badge appearing runs your command (desktop
75
+ notification, Telegram bot, pager — anything) with the details in `$AUM_*` env
76
+ vars. Fires once per transition with a cooldown, never from the `list` CLI.
77
+ By default only `hot`/`churn`/`leak` alert — `idle` is the normal state of an
78
+ agent fleet that waits for work, so it's opt-in.
79
+
80
+ ```yaml
81
+ alerts:
82
+ command: 'terminal-notifier -title agents -message "$AUM_MSG"'
83
+ cooldown: 600
84
+ flags: [hot, churn, leak] # the default; add idle to opt in
85
+ leak_floor_mb: 1536 # optional: leak alerts only above this RSS
86
+ ```
87
+
88
+ `leak_floor_mb` mutes leak *alerts* (the badge still shows) until the agent's
89
+ absolute footprint crosses the floor — agents that accrue working state, like
90
+ a chat bot growing its session context, ratchet RSS exactly like leakers, so
91
+ relative growth alone can be noisy. Crossing the floor counts as the
92
+ appearance, so a genuine ratchet still alerts.
93
+ - **Prometheus `/metrics`.** Per-agent CPU/mem/instances/restarts and badge states
94
+ in text exposition format, aggregated per label (no pid-churn series bloat) —
95
+ point Grafana or any Prometheus scraper at `http://127.0.0.1:8765/metrics`.
96
+ - **Expand the tree.** Click the `+N` badge to unfold an agent's child processes
97
+ (per-child CPU/mem/command) — see what a kill would actually stop before clicking it.
98
+ - **Config hot-reload.** Edits to `agents.yaml` apply on the next poll, no restart.
99
+ A broken edit keeps the last good config and shows the parse error in the header.
100
+ - **`list` subcommand.** `agent-usage-manager list` (or `list --json`) prints a one-shot
101
+ table to stdout — no server, good for scripts and cron checks.
102
+ - **Kill-safe table.** Rows keep a stable order (sorted by label) and never reorder
103
+ while your pointer is over the table, so the kill button can't shift under your
104
+ cursor mid-click.
105
+
106
+ ## Product boundary
107
+
108
+ `agent-usage-manager` is intentionally **not** a fleet scheduler, dispatcher, or
109
+ multi-host orchestrator. It answers local process questions: what agent process is
110
+ running here, what resources is its process tree using, did it enter a suspicious
111
+ state, and can this local operator safely stop it?
112
+
113
+ If you run a separate fleet control plane, treat AUM as an optional read-only
114
+ input. Scrape `list --json`, `/api/agents`, or `/metrics` for local OS facts, then
115
+ make scheduling, budget, restart, and kill/retire decisions in your own
116
+ deterministic control layer. Do not route irreversible fleet operations through
117
+ AUM's kill endpoint as a central substrate.
118
+
119
+ ## Safety
120
+
121
+ This is the important part — a web page that can kill processes needs guardrails:
122
+
123
+ - **Allowlist only.** Only processes matching a pattern in `agents.yaml` are ever
124
+ listed *or* killable. The kill endpoint re-checks the match server-side before
125
+ sending any signal, so the dashboard can never be used to kill an arbitrary PID.
126
+ - **Protected patterns.** Anything matching `protect:` in `agents.yaml` — plus the
127
+ monitor's own process and PID 1 — shows a disabled, greyed-out kill button and is
128
+ refused server-side.
129
+ - **Secret redaction.** Command lines often carry tokens/keys in env vars or flags
130
+ (`FOO_TOKEN=...`, `--api-key ...`, `sk-...`, `ghp_...`, JWTs). The command column
131
+ redacts these to `***` before they ever reach the browser — safe to screenshot.
132
+ - **Browser guard (CSRF + DNS rebinding).** Binding to localhost doesn't keep
133
+ browsers out — any web page you visit can `fetch()` a localhost port. Requests
134
+ whose `Host` is a non-local DNS name are refused (DNS-rebinding guard), and a
135
+ kill request carrying a foreign `Origin` is refused (CSRF guard) — so a
136
+ malicious page can't kill your agents or read your process list. `curl` and
137
+ the dashboard itself are unaffected.
138
+ - **Kill requires a token (caller authorization).** The allowlist above says what
139
+ *may* be killed; the token says *who* may kill. The monitored agents are
140
+ themselves untrusted HTTP callers — a prompt-injected agent with an HTTP tool
141
+ and localhost reach could otherwise `POST /api/kill` and take down its
142
+ siblings (Origin headers are trivially forged outside a browser). The token is
143
+ auto-generated on first run into a `0600` file —
144
+ `~/Library/Application Support/agent-usage-manager/kill_token` on macOS,
145
+ `$XDG_STATE_HOME/agent-usage-manager/kill_token` (default
146
+ `~/.local/state/…`) elsewhere — and every kill must send it as an
147
+ `X-Kill-Token` header. It is **never served over HTTP** (anything that can
148
+ curl the dashboard could read it): the dashboard asks you to paste it once on
149
+ your first kill and keeps it in the browser's localStorage. Delete the file
150
+ to rotate the token.
151
+ - **Action log.** Every kill attempt — success *and* every refusal — appends a
152
+ JSON line to `actions.log` next to the token file: timestamp, caller address,
153
+ target pid/command, outcome. "What was killed at 3am" and "what's been
154
+ probing the kill endpoint" both have an answer. Append-only, no rotation;
155
+ one line per attempt stays tiny.
156
+ - **Non-loopback binds fail closed.** It listens on `127.0.0.1`; asking it to
157
+ bind anything else (`--host 0.0.0.0`, a LAN IP) refuses to start unless you
158
+ also pass `--unsafe-expose`. Exposing the port means one static token is all
159
+ that stands between the network and your agents — put real auth in front
160
+ (reverse proxy + basic auth, SSH tunnel, etc.) before using that flag.
161
+
162
+ ## Limits & known issues
163
+
164
+ - **GPU column is NVIDIA-only.** Per-process GPU memory comes from
165
+ `nvidia-smi --query-compute-apps` — NVIDIA compute processes (CUDA), in
166
+ practice on Linux. AMD/Intel GPUs aren't read, graphics-only workloads don't
167
+ appear, and Apple Silicon has no per-process GPU accounting API at all, so
168
+ the column is hidden on Macs.
169
+ - **Supervision detection is launchd-only (macOS, user domain).** Root
170
+ `LaunchDaemons` aren't flagged — that needs a privileged
171
+ `launchctl print system/…`. On Linux, systemd-supervised services
172
+ (`Restart=always`) aren't detected either, so killing one looks like it
173
+ failed when systemd respawns it — use `systemctl stop` for those.
174
+ - **Same-user privileges only.** Signals are sent with the server's own
175
+ privileges. Agents running as another user (or root) are listed, but a kill
176
+ won't take (`killed: 0` in the response), and CPU/mem can read as 0 where
177
+ the OS denies access.
178
+ - **History is in-memory.** Sparklines and the `hot`/`idle` flags (~20 min
179
+ window) rebuild from scratch after a server restart.
180
+ - **Windows is untested.** Kill maps to `TerminateProcess` via psutil and may
181
+ work, but CI covers Linux + macOS only.
182
+
183
+ ## Quick start
184
+
185
+ **Recommended — one command, nothing to install first:**
186
+
187
+ ```bash
188
+ uvx agent-usage-manager
189
+ # then open http://127.0.0.1:8765 (it also opens automatically)
190
+ ```
191
+
192
+ `uvx` fetches and runs it in one step — no separate install, no virtualenv, no
193
+ leftovers. Don't have [`uv`](https://github.com/astral-sh/uv) yet? One line:
194
+
195
+ ```bash
196
+ curl -LsSf https://astral.sh/uv/install.sh | sh # macOS / Linux
197
+ # or: pip install uv
198
+ ```
199
+
200
+ <details>
201
+ <summary>Other ways to install</summary>
202
+
203
+ ```bash
204
+ pipx install agent-usage-manager # clean isolated global CLI (needs pipx)
205
+
206
+ pip install agent-usage-manager # universal; use inside a venv —
207
+ # system Python may refuse with
208
+ # "externally-managed-environment"
209
+ ```
210
+
211
+ Then run `agent-usage-manager` (flags below).
212
+ </details>
213
+
214
+ From a clone (for hacking on it):
215
+
216
+ ```bash
217
+ git clone <this-repo> && cd agent-usage-manager
218
+ ./run.sh # venv + editable install, serves on :8765
219
+ ```
220
+
221
+ It opens the dashboard in your browser automatically. Flags: `--host`, `--port`,
222
+ `--config /path/to/agents.yaml`, `--no-browser` (for headless/server use),
223
+ `--unsafe-expose` (required for any non-loopback `--host` — see
224
+ [Safety](#safety)).
225
+
226
+ ## Configure which processes are "agents"
227
+
228
+ Edit `agents.yaml`:
229
+
230
+ ```yaml
231
+ agents:
232
+ - label: openclaw # shown as the badge in the UI
233
+ match: openclaw # case-insensitive substring of the command line
234
+ - label: hermes
235
+ match: hermes
236
+ - label: claude-code
237
+ match: "claude(\\s|$|-code)"
238
+ regex: true # treat `match` as a regex instead of substring
239
+
240
+ protect: # matched + listed, but never killable
241
+ - uvicorn
242
+
243
+ ignore: # never an agent: not listed, not killable
244
+ - crashpad # incidental processes that share a name/bundle
245
+ - shipit # path with a real agent (crash handlers,
246
+ - kiro-cli-term # auto-updaters, integrated-terminal shells, …)
247
+ ```
248
+
249
+ A process matches if the pattern hits its **executable basename + first few
250
+ arguments** — deliberately not the whole command line, so a long embedded arg
251
+ (e.g. a system prompt mentioning "claude") can't misclassify a wrapper. On macOS
252
+ the outermost `.app` **bundle name** is also included, so GUI agents that launch
253
+ a generically-named binary (Kiro.app → `Electron`) are still matched by app name.
254
+
255
+ `protect:` keeps a matched process listed but refuses to kill it; `ignore:`
256
+ drops it from agent classification entirely.
257
+
258
+ **Telling identical agents apart (`tmux_labels:`)** — a fleet of same-binary
259
+ agents (say five `claude` bots, one per tmux session) all hits one `agents:`
260
+ entry and shows N indistinguishable rows; a distinguishing flag deeper in
261
+ their command lines is invisible to matching *by design* (see above). When
262
+ each instance runs in its own tmux session, the session name is its identity:
263
+
264
+ ```yaml
265
+ tmux_labels: "^bot-(.+)$" # session bot-coder_1 → row label coder_1
266
+ ```
267
+
268
+ If a matched agent root (or an ancestor) is a tmux pane process whose session
269
+ name matches the regex, the row's label becomes the first capture group (the
270
+ whole session name if there's no group). Churn tracking, alert transitions,
271
+ and `/metrics` series all use the derived label, so each instance gets its own
272
+ state. Sessions that don't match the regex keep their `agents:` label, and the
273
+ key is ignored where tmux isn't installed or running.
274
+
275
+ **Which `agents.yaml` is used** — resolved once at startup, first hit wins:
276
+
277
+ 1. `AGENTS_CONFIG=/path/to/agents.yaml` env var (the `--config` flag sets this)
278
+ 2. `./agents.yaml` in the directory you launched from
279
+ 3. the default bundled with the package
280
+
281
+ The dashboard header (and `list --json`) shows the resolved path, so you can
282
+ always see which file is live. Hot-reload watches that one file. An
283
+ `AGENTS_CONFIG` path that doesn't exist is an error at startup, not a silent
284
+ fallback.
285
+
286
+ ## launchd-supervised agents (macOS)
287
+
288
+ Some agents run as **launchd services** (a `~/Library/LaunchAgents/*.plist`, or
289
+ anything started by `brew services`). If such a job sets `KeepAlive`, a signal
290
+ can't stop it: the process dies, launchd immediately respawns it under a new PID,
291
+ and the dashboard's "kill" looks like it silently failed.
292
+
293
+ The dashboard detects these (via `launchctl list`) and marks them with a
294
+ **`launchd`** badge. Instead of dead-end kill/force buttons it shows the command
295
+ that actually stops the job — click to copy:
296
+
297
+ ```sh
298
+ launchctl bootout gui/<uid>/<label> # stop now
299
+ launchctl disable gui/<uid>/<label> # …and don't auto-start at login
300
+ ```
301
+
302
+ The kill endpoint refuses signals for these jobs (HTTP 409) and returns the same
303
+ guidance, so the API never lies about a kill that won't stick. The message is
304
+ tailored to the job: `KeepAlive` jobs are told a signal won't stick at all;
305
+ `RunAtLoad`-only jobs are told a signal works now but the job restarts at next
306
+ login. *Limitation:* root `LaunchDaemons` aren't flagged — see
307
+ [Limits & known issues](#limits--known-issues).
308
+
309
+ ## GPU notes
310
+
311
+ Per-process GPU memory comes from `nvidia-smi` when it's on `PATH` (Linux / NVIDIA).
312
+ **Apple Silicon has no per-process GPU accounting API**, so the GPU column stays blank
313
+ on Macs — CPU and memory are the meaningful resource signals there.
314
+
315
+ ## API
316
+
317
+ - `GET /api/agents` → `{ api_version, aum_version, agents: [...], host, cpu_count, mem_total_mb, mem_used_pct, config_path, config_error, ts }`
318
+ — each agent includes read-only telemetry such as `pid`, `create_time`, `label`,
319
+ resource totals, recent CPU `trend`, flags (`hot`, `idle`, `churn`, `leak` when
320
+ present), protection state, and supervised-process guidance. Pair `pid` with
321
+ `create_time` when caching rows so PID reuse cannot alias two different agents.
322
+ This endpoint is suitable as an input to external tools, not as a fleet-control
323
+ contract.
324
+ - `GET /api/tree/{pid}` → the agent's process subtree (per-child pid/name/cpu/mem/cmdline);
325
+ only works on recognized agents, same authorization as kill
326
+ - `POST /api/kill/{pid}?force=false` → SIGTERM (or SIGKILL with `force=true`).
327
+ Requires the `X-Kill-Token` header — the token lives in the `0600` file shown
328
+ in the 403 message (see [Safety](#safety)):
329
+
330
+ ```bash
331
+ curl -X POST -H "X-Kill-Token: $(cat ~/Library/Application\ Support/agent-usage-manager/kill_token)" \
332
+ http://127.0.0.1:8765/api/kill/48213
333
+ ```
334
+
335
+ ## Run as a service
336
+
337
+ Linux (systemd), `~/.config/systemd/user/agent-usage-manager.service`:
338
+
339
+ ```ini
340
+ [Unit]
341
+ Description=agent usage manager
342
+ [Service]
343
+ ExecStart=%h/agent-usage-manager/.venv/bin/uvicorn app:app --port 8765
344
+ WorkingDirectory=%h/agent-usage-manager
345
+ Restart=on-failure
346
+ [Install]
347
+ WantedBy=default.target
348
+ ```
349
+
350
+ ```bash
351
+ systemctl --user enable --now agent-usage-manager
352
+ ```
353
+
354
+ ## Development
355
+
356
+ ```bash
357
+ git clone https://github.com/minglong51/agent-usage-manager && cd agent-usage-manager
358
+ pip install -e ".[dev]"
359
+ pytest -q
360
+ ```
361
+
362
+ CI runs the test suite on Linux + macOS (Python 3.9 and 3.12) on every push and PR.
363
+ Cross-platform note: kill uses psutil's `terminate()`/`kill()`, which map to
364
+ SIGTERM/SIGKILL on POSIX and TerminateProcess on Windows.
365
+
366
+ ## Release notes
367
+
368
+ ### 0.2.2 — unreleased
369
+
370
+ - Added `api_version` and `aum_version` to `/api/agents` and `list --json`.
371
+ - Added per-agent `create_time` so external telemetry consumers can pair it with
372
+ `pid` and avoid PID-reuse aliasing.
373
+ - Added `tmux_labels:` — derive per-instance row labels from tmux session names,
374
+ so a fleet of identical agents stops rendering as N indistinguishable rows.
375
+
376
+ ### 0.2.1 — security and verification hardening
377
+
378
+ - Kill endpoint now requires caller authorization via the static token file.
379
+ - Non-loopback binds fail closed unless `--unsafe-expose` is explicitly passed.
380
+ - Kill attempts and refusals append to the local action log.
381
+ - Added deterministic kill-path regression tests, including pid/create_time pins.
382
+ - Added synthetic hot/idle/churn/leak trace fixtures.
383
+ - Added adversarial matcher cases so lookalike process names stay test-covered.
384
+
385
+ ## Troubleshooting
386
+
387
+ - **`pip install` fails building psutil** — no prebuilt wheel for your
388
+ Python/platform, so pip compiles it: you need a C toolchain and Python
389
+ headers (`xcode-select --install` on macOS; `apt install gcc python3-dev`
390
+ on Debian/Ubuntu). Or skip the problem with `uvx agent-usage-manager`.
391
+ - **Dashboard is empty / "No matching agents running"** — first check which
392
+ config was picked up (resolution order above; the header shows the resolved
393
+ path). Then remember matching is against the executable basename + first few
394
+ arguments, not the full command line — a pattern that only appears deep in
395
+ the args won't match.
396
+ - **Kill "doesn't work" — the agent comes back under a new PID** — it's
397
+ supervised. On macOS the row gets a `launchd` badge with the `launchctl
398
+ bootout` command that actually stops it; on Linux, systemd services aren't
399
+ detected (see limits) — `systemctl stop` them. A `killed: 0` in the kill
400
+ response means nothing was actually signaled (e.g. the agent runs as
401
+ another user).
402
+ - **HTTP 403 on every request** — the DNS-rebinding guard refuses non-local
403
+ hostnames. Use `http://127.0.0.1:8765` (or a bare IP) instead of a custom
404
+ DNS name pointing at the box.
405
+ - **HTTP 403 on kill only ("Kill requires the X-Kill-Token header")** — send
406
+ the token from the file named in the message. In the dashboard, the paste
407
+ prompt reappears on your next kill click (a stored stale token is forgotten
408
+ automatically when the server rejects it).
409
+ - **"refusing to bind …" at startup** — non-loopback `--host` values fail
410
+ closed; add `--unsafe-expose` only with auth in front (see
411
+ [Safety](#safety)).
412
+
413
+ ## License
414
+
415
+ MIT