agent-usage-manager 0.2.4__tar.gz → 0.2.5__tar.gz
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- agent_usage_manager-0.2.5/.github/workflows/publish.yml +41 -0
- agent_usage_manager-0.2.5/AGENTS.md +58 -0
- agent_usage_manager-0.2.5/CLAUDE.md +1 -0
- {agent_usage_manager-0.2.4 → agent_usage_manager-0.2.5}/PKG-INFO +95 -7
- {agent_usage_manager-0.2.4 → agent_usage_manager-0.2.5}/README.md +93 -5
- {agent_usage_manager-0.2.4 → agent_usage_manager-0.2.5}/agent_usage_manager/agents.yaml +35 -2
- {agent_usage_manager-0.2.4 → agent_usage_manager-0.2.5}/agent_usage_manager/app.py +182 -21
- {agent_usage_manager-0.2.4 → agent_usage_manager-0.2.5}/agent_usage_manager/cli.py +87 -2
- {agent_usage_manager-0.2.4 → agent_usage_manager-0.2.5}/agent_usage_manager/static/index.html +98 -38
- agent_usage_manager-0.2.5/demo.tape +23 -0
- agent_usage_manager-0.2.5/docs/dashboard-live.gif +0 -0
- agent_usage_manager-0.2.5/docs/dashboard-live.png +0 -0
- agent_usage_manager-0.2.5/docs/demo.gif +0 -0
- {agent_usage_manager-0.2.4 → agent_usage_manager-0.2.5}/docs/design/HLD.md +40 -35
- {agent_usage_manager-0.2.4 → agent_usage_manager-0.2.5}/docs/design/LLD.md +183 -102
- {agent_usage_manager-0.2.4 → agent_usage_manager-0.2.5}/pyproject.toml +1 -1
- agent_usage_manager-0.2.5/tests/test_design_docs.py +300 -0
- {agent_usage_manager-0.2.4 → agent_usage_manager-0.2.5}/tests/test_smoke.py +185 -2
- {agent_usage_manager-0.2.4 → agent_usage_manager-0.2.5}/tests/test_synthetic.py +117 -0
- agent_usage_manager-0.2.4/demo.tape +0 -21
- {agent_usage_manager-0.2.4 → agent_usage_manager-0.2.5}/.github/workflows/ci.yml +0 -0
- {agent_usage_manager-0.2.4 → agent_usage_manager-0.2.5}/.gitignore +0 -0
- {agent_usage_manager-0.2.4 → agent_usage_manager-0.2.5}/LICENSE +0 -0
- {agent_usage_manager-0.2.4 → agent_usage_manager-0.2.5}/agent_usage_manager/__init__.py +0 -0
- {agent_usage_manager-0.2.4 → agent_usage_manager-0.2.5}/docs/dashboard.png +0 -0
- {agent_usage_manager-0.2.4 → agent_usage_manager-0.2.5}/requirements.txt +0 -0
- {agent_usage_manager-0.2.4 → agent_usage_manager-0.2.5}/run.sh +0 -0
- {agent_usage_manager-0.2.4 → agent_usage_manager-0.2.5}/uv.lock +0 -0
|
@@ -0,0 +1,41 @@
|
|
|
1
|
+
name: Publish to PyPI
|
|
2
|
+
|
|
3
|
+
on:
|
|
4
|
+
release:
|
|
5
|
+
types: [published]
|
|
6
|
+
|
|
7
|
+
permissions:
|
|
8
|
+
contents: read
|
|
9
|
+
|
|
10
|
+
jobs:
|
|
11
|
+
build:
|
|
12
|
+
runs-on: ubuntu-latest
|
|
13
|
+
steps:
|
|
14
|
+
- uses: actions/checkout@v4
|
|
15
|
+
- uses: actions/setup-python@v5
|
|
16
|
+
with:
|
|
17
|
+
python-version: "3.12"
|
|
18
|
+
- run: python -m pip install --upgrade pip build
|
|
19
|
+
- name: Verify tag matches pyproject version
|
|
20
|
+
run: |
|
|
21
|
+
PKG=$(python -c "import tomllib;print(tomllib.load(open('pyproject.toml','rb'))['project']['version'])")
|
|
22
|
+
TAG="${GITHUB_REF_NAME#v}"
|
|
23
|
+
[ "$PKG" = "$TAG" ] || { echo "tag $GITHUB_REF_NAME != pyproject version $PKG"; exit 1; }
|
|
24
|
+
- run: python -m build
|
|
25
|
+
- uses: actions/upload-artifact@v4
|
|
26
|
+
with:
|
|
27
|
+
name: dist
|
|
28
|
+
path: dist/
|
|
29
|
+
|
|
30
|
+
publish:
|
|
31
|
+
needs: build
|
|
32
|
+
runs-on: ubuntu-latest
|
|
33
|
+
environment: pypi
|
|
34
|
+
permissions:
|
|
35
|
+
id-token: write # trusted publishing (OIDC) — no API token stored anywhere
|
|
36
|
+
steps:
|
|
37
|
+
- uses: actions/download-artifact@v4
|
|
38
|
+
with:
|
|
39
|
+
name: dist
|
|
40
|
+
path: dist/
|
|
41
|
+
- uses: pypa/gh-action-pypi-publish@release/v1
|
|
@@ -0,0 +1,58 @@
|
|
|
1
|
+
# agent-usage-manager
|
|
2
|
+
|
|
3
|
+
## Design docs (HLD/LLD)
|
|
4
|
+
|
|
5
|
+
`docs/design/HLD.md` and `docs/design/LLD.md` are the architecture contract for this
|
|
6
|
+
repo. **Where a doc disagrees with the code, the code wins and the doc is stale** — that
|
|
7
|
+
is a bug to fix, not a tie to settle in the doc's favor.
|
|
8
|
+
|
|
9
|
+
A change to mapped code updates the owning doc **in the same PR**. Tiered, so routine
|
|
10
|
+
work does not churn the docs:
|
|
11
|
+
|
|
12
|
+
1. **Update the LLD** when a file, function signature, data model, schema, event shape,
|
|
13
|
+
or config/env surface **the LLD names** changes — or when a module appears or
|
|
14
|
+
disappears under an owned path.
|
|
15
|
+
2. **Update the HLD** only on a boundary change: a component added or retired, a new
|
|
16
|
+
port/service/external dependency, a changed trust boundary or auth gate, a changed
|
|
17
|
+
contract between components, or a subsystem promoted, demoted, or absorbed.
|
|
18
|
+
3. **Update neither** for bugfixes, styling, tests, or refactors that preserve every
|
|
19
|
+
surface the docs name.
|
|
20
|
+
|
|
21
|
+
Do not regenerate a doc wholesale to satisfy this. The `design-doc` skill **overwrites**
|
|
22
|
+
and will destroy hand-written intent (why a gate exists, what the design refuses to do);
|
|
23
|
+
use it to create a missing pair, never to refresh a live one. Edit the affected sections
|
|
24
|
+
and move the `**Refreshed:**` line.
|
|
25
|
+
|
|
26
|
+
`CLAUDE.md` is a **symlink to `AGENTS.md`** — edit `AGENTS.md`, never the symlink.
|
|
27
|
+
Claude Code prefers `CLAUDE.md` and ignores a real `AGENTS.md` when both exist, so the
|
|
28
|
+
symlink keeps one source of truth and stops anything that later writes a `CLAUDE.md`
|
|
29
|
+
from silently shadowing this contract. Codex and Kimi read `AGENTS.md` directly. The
|
|
30
|
+
antigravity lane reads no project file at all — give it self-contained prompts.
|
|
31
|
+
|
|
32
|
+
Ownership map — machine-readable, parsed by `tests/test_design_docs.py`. Add a row when
|
|
33
|
+
you add a package; `none` means "no design contract, deliberately". A dir with no row but
|
|
34
|
+
rows beneath it is a container and is recursed into, so a package dropped inside one
|
|
35
|
+
cannot inherit its parent's coverage.
|
|
36
|
+
|
|
37
|
+
```design-doc-map
|
|
38
|
+
agent_usage_manager/ -> docs/design/HLD.md docs/design/LLD.md
|
|
39
|
+
docs/ -> none
|
|
40
|
+
tests/ -> none
|
|
41
|
+
```
|
|
42
|
+
|
|
43
|
+
`python3 tests/test_design_docs.py` reports drift: per doc, the modules added or removed
|
|
44
|
+
under its owned paths since that doc last changed (test files excluded — a new test needs
|
|
45
|
+
no LLD entry). Added/removed modules are the signal; raw commit counts are noise and print
|
|
46
|
+
as context only. The pytest asserts the map is structurally sound; it deliberately does
|
|
47
|
+
**not** fail on drift, because a doc gate that blocks merges buys rubber-stamp edits,
|
|
48
|
+
not maintained docs. A file *modified* to change its public API will not flag — boundary
|
|
49
|
+
changes still need a human read.
|
|
50
|
+
|
|
51
|
+
|
|
52
|
+
`python3 tests/test_design_docs.py --audit` answers a different question: which tracked modules
|
|
53
|
+
are named in **no** design doc at all. The drift report only diffs forward from a doc's
|
|
54
|
+
own last change, so staleness that predates an incomplete refresh is invisible to it
|
|
55
|
+
permanently — paws described a `scene.js` renderer for three weeks after it became
|
|
56
|
+
`wool.js`, and no number of drift runs could have said so. Run `--audit` when you inherit
|
|
57
|
+
a doc you did not write. It skips tests, `__init__.py`, vendored code, type shims,
|
|
58
|
+
generated migration revisions, and anything under a `-> none` path.
|
|
@@ -0,0 +1 @@
|
|
|
1
|
+
AGENTS.md
|
|
@@ -1,6 +1,6 @@
|
|
|
1
|
-
Metadata-Version: 2.
|
|
1
|
+
Metadata-Version: 2.5
|
|
2
2
|
Name: agent-usage-manager
|
|
3
|
-
Version: 0.2.
|
|
3
|
+
Version: 0.2.5
|
|
4
4
|
Summary: htop for AI agents — liveness, CPU/mem/GPU usage, and a kill switch for headless agents (openclaw, hermes, ollama, vllm, claude-code).
|
|
5
5
|
Project-URL: Homepage, https://github.com/minglong51/agent-usage-manager
|
|
6
6
|
Project-URL: Repository, https://github.com/minglong51/agent-usage-manager
|
|
@@ -49,6 +49,17 @@ ollama 50122 ● running 3.1 9210 14080 6h 02m ollama ru
|
|
|
49
49
|
> run on Apple Silicon, where per-process GPU stats aren't available so that column
|
|
50
50
|
> is hidden. The UI auto-refreshes every 3s.
|
|
51
51
|
|
|
52
|
+
## Quick start
|
|
53
|
+
|
|
54
|
+
```bash
|
|
55
|
+
uvx agent-usage-manager # then open http://127.0.0.1:8765 (opens automatically)
|
|
56
|
+
```
|
|
57
|
+
|
|
58
|
+
One command — no install, no virtualenv, no leftovers. Other install options,
|
|
59
|
+
config, and flags: [Install & run](#install--run).
|
|
60
|
+
|
|
61
|
+

|
|
62
|
+
|
|
52
63
|
## What it does
|
|
53
64
|
|
|
54
65
|
- **One row per agent.** Agents are grouped by process tree — the spawned children
|
|
@@ -64,7 +75,9 @@ ollama 50122 ● running 3.1 9210 14080 6h 02m ollama ru
|
|
|
64
75
|
- **Trends, not just snapshots.** Each row has a CPU sparkline (last ~20 min, sampled
|
|
65
76
|
in the background even with no browser open), plus a **`hot 5m+`** badge when an
|
|
66
77
|
agent has been pegged ≥90% CPU for 5+ minutes, an **`idle 10m+`** badge when a
|
|
67
|
-
long-running agent has done nothing for 10+ minutes
|
|
78
|
+
long-running agent has done nothing for 10+ minutes (suppressed for labels in
|
|
79
|
+
`idle_ok:` — for a fleet that waits for work, idle is the normal state and
|
|
80
|
+
badging it is wallpaper), and a **`churn ×N`** badge when
|
|
68
81
|
the same agent has died young 3+ times in 10 minutes — the states worth investigating
|
|
69
82
|
(runaway, possibly wedged, crash-looping under a supervisor). Churn is what hot/idle
|
|
70
83
|
can't see: a crash-looping process is a fresh pid every poll, so no per-process
|
|
@@ -73,9 +86,15 @@ ollama 50122 ● running 3.1 9210 14080 6h 02m ollama ru
|
|
|
73
86
|
- **Alerts.** A dashboard only helps while you're looking at it. Add an `alerts:`
|
|
74
87
|
block to `agents.yaml` and any badge appearing runs your command (desktop
|
|
75
88
|
notification, Telegram bot, pager — anything) with the details in `$AUM_*` env
|
|
76
|
-
vars.
|
|
89
|
+
vars. `$AUM_MSG` leads with the plain-English verdict ("Codex is crash-looping —
|
|
90
|
+
4 restarts in 10 min"); the machine snapshot trails in brackets. Fires once per
|
|
91
|
+
transition with a cooldown, never from the `list` CLI, and the cooldown is only
|
|
92
|
+
charged when the command exits 0 (a broken notifier doesn't suppress the retry).
|
|
77
93
|
By default only `hot`/`churn`/`leak` alert — `idle` is the normal state of an
|
|
78
94
|
agent fleet that waits for work, so it's opt-in.
|
|
95
|
+
Run `agent-usage-manager test-alert` once after wiring it up: it fires the
|
|
96
|
+
configured command synchronously with a test message and reports the exit
|
|
97
|
+
status, so a broken channel surfaces today instead of during the next incident.
|
|
79
98
|
|
|
80
99
|
```yaml
|
|
81
100
|
alerts:
|
|
@@ -98,11 +117,25 @@ ollama 50122 ● running 3.1 9210 14080 6h 02m ollama ru
|
|
|
98
117
|
- **Config hot-reload.** Edits to `agents.yaml` apply on the next poll, no restart.
|
|
99
118
|
A broken edit keeps the last good config and shows the parse error in the header.
|
|
100
119
|
- **`list` subcommand.** `agent-usage-manager list` (or `list --json`) prints a one-shot
|
|
101
|
-
table to stdout — no server, good for scripts and cron checks.
|
|
120
|
+
table to stdout — no server, good for scripts and cron checks. One-shot mode has no
|
|
121
|
+
history, so the sustained-state flags (`hot`/`idle`/`churn`/`leak`) can never populate
|
|
122
|
+
there (`--json` says `"flags_available": false`); query the running server's
|
|
123
|
+
`/api/agents` when you need flags.
|
|
102
124
|
- **Kill-safe table.** Rows keep a stable order (sorted by label) and never reorder
|
|
103
125
|
while your pointer is over the table, so the kill button can't shift under your
|
|
104
126
|
cursor mid-click.
|
|
105
127
|
|
|
128
|
+
## Why not just htop / Grafana?
|
|
129
|
+
|
|
130
|
+
- **htop sees processes, not agents.** No rollup of a spawned tree under the
|
|
131
|
+
agent that owns it, no agent states (crash-looping, idle, leaking), and an
|
|
132
|
+
unguarded `F9`.
|
|
133
|
+
- **Grafana + Prometheus + an exporter is a stack you operate.** This is one
|
|
134
|
+
command with no database — and it still serves `/metrics` if you want both.
|
|
135
|
+
- **The kill switch is the point.** Allowlist-matched, token-gated, guarded
|
|
136
|
+
against CSRF and DNS rebinding, with an append-only action log — a web page
|
|
137
|
+
that can stop processes needs exactly those.
|
|
138
|
+
|
|
106
139
|
## Product boundary
|
|
107
140
|
|
|
108
141
|
`agent-usage-manager` is intentionally **not** a fleet scheduler, dispatcher, or
|
|
@@ -158,6 +191,13 @@ This is the important part — a web page that can kill processes needs guardrai
|
|
|
158
191
|
also pass `--unsafe-expose`. Exposing the port means one static token is all
|
|
159
192
|
that stands between the network and your agents — put real auth in front
|
|
160
193
|
(reverse proxy + basic auth, SSH tunnel, etc.) before using that flag.
|
|
194
|
+
- **A proxy voids the loopback guarantee.** The production deployment (Ming's
|
|
195
|
+
suite) fronts this loopback bind with `tailscale serve` (`:8448 → 127.0.0.1:8765`),
|
|
196
|
+
so the tailnet reaches the UI despite the loopback bind: reads like
|
|
197
|
+
`/api/agents` are open there, and the static token is the *only* boundary on
|
|
198
|
+
actions. That is deliberate (owner-only tailnet + token-gated actions), but do
|
|
199
|
+
not read "loopback-only" as "not network-reachable" — a proxy in front is not
|
|
200
|
+
a trust boundary.
|
|
161
201
|
|
|
162
202
|
## Limits & known issues
|
|
163
203
|
|
|
@@ -189,7 +229,7 @@ This is the important part — a web page that can kill processes needs guardrai
|
|
|
189
229
|
- **Windows is untested.** Kill maps to `TerminateProcess` via psutil and may
|
|
190
230
|
work, but CI covers Linux + macOS only.
|
|
191
231
|
|
|
192
|
-
##
|
|
232
|
+
## Install & run
|
|
193
233
|
|
|
194
234
|
**Recommended — one command, nothing to install first:**
|
|
195
235
|
|
|
@@ -281,6 +321,19 @@ and `/metrics` series all use the derived label, so each instance gets its own
|
|
|
281
321
|
state. Sessions that don't match the regex keep their `agents:` label, and the
|
|
282
322
|
key is ignored where tmux isn't installed or running.
|
|
283
323
|
|
|
324
|
+
**Supervised fleets (`launchd_labels:`, macOS)** — the same problem for agents
|
|
325
|
+
that run as launchd jobs and never touch tmux (five `hermes` LaunchAgents all
|
|
326
|
+
landing as "hermes"). The launchd job label is their durable identity:
|
|
327
|
+
|
|
328
|
+
```yaml
|
|
329
|
+
launchd_labels: "^ai\\.hermes\\.(?:gateway-)?(.+)$" # ai.hermes.gateway-frontdoor → frontdoor
|
|
330
|
+
```
|
|
331
|
+
|
|
332
|
+
When a matched root's own launchd job label matches, the first capture group
|
|
333
|
+
(the whole label if there's no group) becomes the row label — with the same
|
|
334
|
+
per-instance churn/alert/metrics identity as `tmux_labels:`. `tmux_labels`
|
|
335
|
+
wins when both apply; roots with no matching job keep their `agents:` label.
|
|
336
|
+
|
|
284
337
|
**Which `agents.yaml` is used** — resolved once at startup, first hit wins:
|
|
285
338
|
|
|
286
339
|
1. `AGENTS_CONFIG=/path/to/agents.yaml` env var (the `--config` flag sets this)
|
|
@@ -349,7 +402,7 @@ Linux (systemd), `~/.config/systemd/user/agent-usage-manager.service`:
|
|
|
349
402
|
[Unit]
|
|
350
403
|
Description=agent usage manager
|
|
351
404
|
[Service]
|
|
352
|
-
ExecStart=%h/agent-usage-manager/.venv/bin/uvicorn app:app --port 8765
|
|
405
|
+
ExecStart=%h/agent-usage-manager/.venv/bin/uvicorn agent_usage_manager.app:app --port 8765
|
|
353
406
|
WorkingDirectory=%h/agent-usage-manager
|
|
354
407
|
Restart=on-failure
|
|
355
408
|
[Install]
|
|
@@ -374,6 +427,35 @@ SIGTERM/SIGKILL on POSIX and TerminateProcess on Windows.
|
|
|
374
427
|
|
|
375
428
|
## Release notes
|
|
376
429
|
|
|
430
|
+
### 0.2.5 — per-instance labels for supervised fleets + the honesty batch
|
|
431
|
+
|
|
432
|
+
- `launchd_labels:` config — per-instance row labels from launchd job labels,
|
|
433
|
+
mirroring `tmux_labels:`. A supervised fleet (several LaunchAgents on one
|
|
434
|
+
binary) no longer collapses into one blurred label for churn tracking, alert
|
|
435
|
+
transitions, and `/metrics`. `idle_ok:` suppression matches the base matcher
|
|
436
|
+
label as well, so a renamed instance (`hermes` → `frontdoor`) keeps its
|
|
437
|
+
class-level suppression.
|
|
438
|
+
- New `test-alert` subcommand: fires the configured `alerts.command` once,
|
|
439
|
+
synchronously, and reports the exit status — proves the alert channel before
|
|
440
|
+
an incident depends on it.
|
|
441
|
+
- `ignore:` now covers ChatGPT.app's embedded Codex helpers (renderer/service
|
|
442
|
+
processes, `Resources/codex` app-server) and the `Codex Computer Use`
|
|
443
|
+
desktop-automation app — they matched the `codex` pattern via the bundle
|
|
444
|
+
name and cluttered the dashboard with GUI plumbing.
|
|
445
|
+
- Config deletion is no longer silent: a vanished `agents.yaml` surfaces as a
|
|
446
|
+
`config_error` ("running on the last good config") instead of looking like
|
|
447
|
+
hot-reload still works.
|
|
448
|
+
- The kill confirm now names the agent (label + command line, not just PID)
|
|
449
|
+
and discloses that plain kill escalates SIGTERM → SIGKILL after 3s.
|
|
450
|
+
- The first-kill token prompt names the server's host (and the ssh one-liner)
|
|
451
|
+
for browsers viewing the dashboard over a tunnel.
|
|
452
|
+
- The DNS-rebinding 403 now names the remedy (use a loopback name or bare IP).
|
|
453
|
+
- `list` says so when flag fields can't populate (one-shot mode has no
|
|
454
|
+
history): a stderr note, and `"flags_available": false` in `--json`.
|
|
455
|
+
- `--host ::1` now produces a valid bracketed IPv6 URL for the browser open.
|
|
456
|
+
- Fixed the README's systemd unit (`uvicorn app:app` could never resolve the
|
|
457
|
+
module; it's `agent_usage_manager.app:app`).
|
|
458
|
+
|
|
377
459
|
### 0.2.4 — stop reporting non-agents
|
|
378
460
|
|
|
379
461
|
- `ignore:` now covers Sparkle's `Autoupdate`/`Updater` (Codex.app's equivalent of
|
|
@@ -436,6 +518,12 @@ SIGTERM/SIGKILL on POSIX and TerminateProcess on Windows.
|
|
|
436
518
|
closed; add `--unsafe-expose` only with auth in front (see
|
|
437
519
|
[Safety](#safety)).
|
|
438
520
|
|
|
521
|
+
## Stay in the loop
|
|
522
|
+
|
|
523
|
+
New tools and field notes on running AI agents with discipline go to the
|
|
524
|
+
[Agent Discipline](https://buttondown.com/minglong51) list first — launch
|
|
525
|
+
notes, operational patterns, early access. A few emails a month at most.
|
|
526
|
+
|
|
439
527
|
## License
|
|
440
528
|
|
|
441
529
|
MIT
|
|
@@ -30,6 +30,17 @@ ollama 50122 ● running 3.1 9210 14080 6h 02m ollama ru
|
|
|
30
30
|
> run on Apple Silicon, where per-process GPU stats aren't available so that column
|
|
31
31
|
> is hidden. The UI auto-refreshes every 3s.
|
|
32
32
|
|
|
33
|
+
## Quick start
|
|
34
|
+
|
|
35
|
+
```bash
|
|
36
|
+
uvx agent-usage-manager # then open http://127.0.0.1:8765 (opens automatically)
|
|
37
|
+
```
|
|
38
|
+
|
|
39
|
+
One command — no install, no virtualenv, no leftovers. Other install options,
|
|
40
|
+
config, and flags: [Install & run](#install--run).
|
|
41
|
+
|
|
42
|
+

|
|
43
|
+
|
|
33
44
|
## What it does
|
|
34
45
|
|
|
35
46
|
- **One row per agent.** Agents are grouped by process tree — the spawned children
|
|
@@ -45,7 +56,9 @@ ollama 50122 ● running 3.1 9210 14080 6h 02m ollama ru
|
|
|
45
56
|
- **Trends, not just snapshots.** Each row has a CPU sparkline (last ~20 min, sampled
|
|
46
57
|
in the background even with no browser open), plus a **`hot 5m+`** badge when an
|
|
47
58
|
agent has been pegged ≥90% CPU for 5+ minutes, an **`idle 10m+`** badge when a
|
|
48
|
-
long-running agent has done nothing for 10+ minutes
|
|
59
|
+
long-running agent has done nothing for 10+ minutes (suppressed for labels in
|
|
60
|
+
`idle_ok:` — for a fleet that waits for work, idle is the normal state and
|
|
61
|
+
badging it is wallpaper), and a **`churn ×N`** badge when
|
|
49
62
|
the same agent has died young 3+ times in 10 minutes — the states worth investigating
|
|
50
63
|
(runaway, possibly wedged, crash-looping under a supervisor). Churn is what hot/idle
|
|
51
64
|
can't see: a crash-looping process is a fresh pid every poll, so no per-process
|
|
@@ -54,9 +67,15 @@ ollama 50122 ● running 3.1 9210 14080 6h 02m ollama ru
|
|
|
54
67
|
- **Alerts.** A dashboard only helps while you're looking at it. Add an `alerts:`
|
|
55
68
|
block to `agents.yaml` and any badge appearing runs your command (desktop
|
|
56
69
|
notification, Telegram bot, pager — anything) with the details in `$AUM_*` env
|
|
57
|
-
vars.
|
|
70
|
+
vars. `$AUM_MSG` leads with the plain-English verdict ("Codex is crash-looping —
|
|
71
|
+
4 restarts in 10 min"); the machine snapshot trails in brackets. Fires once per
|
|
72
|
+
transition with a cooldown, never from the `list` CLI, and the cooldown is only
|
|
73
|
+
charged when the command exits 0 (a broken notifier doesn't suppress the retry).
|
|
58
74
|
By default only `hot`/`churn`/`leak` alert — `idle` is the normal state of an
|
|
59
75
|
agent fleet that waits for work, so it's opt-in.
|
|
76
|
+
Run `agent-usage-manager test-alert` once after wiring it up: it fires the
|
|
77
|
+
configured command synchronously with a test message and reports the exit
|
|
78
|
+
status, so a broken channel surfaces today instead of during the next incident.
|
|
60
79
|
|
|
61
80
|
```yaml
|
|
62
81
|
alerts:
|
|
@@ -79,11 +98,25 @@ ollama 50122 ● running 3.1 9210 14080 6h 02m ollama ru
|
|
|
79
98
|
- **Config hot-reload.** Edits to `agents.yaml` apply on the next poll, no restart.
|
|
80
99
|
A broken edit keeps the last good config and shows the parse error in the header.
|
|
81
100
|
- **`list` subcommand.** `agent-usage-manager list` (or `list --json`) prints a one-shot
|
|
82
|
-
table to stdout — no server, good for scripts and cron checks.
|
|
101
|
+
table to stdout — no server, good for scripts and cron checks. One-shot mode has no
|
|
102
|
+
history, so the sustained-state flags (`hot`/`idle`/`churn`/`leak`) can never populate
|
|
103
|
+
there (`--json` says `"flags_available": false`); query the running server's
|
|
104
|
+
`/api/agents` when you need flags.
|
|
83
105
|
- **Kill-safe table.** Rows keep a stable order (sorted by label) and never reorder
|
|
84
106
|
while your pointer is over the table, so the kill button can't shift under your
|
|
85
107
|
cursor mid-click.
|
|
86
108
|
|
|
109
|
+
## Why not just htop / Grafana?
|
|
110
|
+
|
|
111
|
+
- **htop sees processes, not agents.** No rollup of a spawned tree under the
|
|
112
|
+
agent that owns it, no agent states (crash-looping, idle, leaking), and an
|
|
113
|
+
unguarded `F9`.
|
|
114
|
+
- **Grafana + Prometheus + an exporter is a stack you operate.** This is one
|
|
115
|
+
command with no database — and it still serves `/metrics` if you want both.
|
|
116
|
+
- **The kill switch is the point.** Allowlist-matched, token-gated, guarded
|
|
117
|
+
against CSRF and DNS rebinding, with an append-only action log — a web page
|
|
118
|
+
that can stop processes needs exactly those.
|
|
119
|
+
|
|
87
120
|
## Product boundary
|
|
88
121
|
|
|
89
122
|
`agent-usage-manager` is intentionally **not** a fleet scheduler, dispatcher, or
|
|
@@ -139,6 +172,13 @@ This is the important part — a web page that can kill processes needs guardrai
|
|
|
139
172
|
also pass `--unsafe-expose`. Exposing the port means one static token is all
|
|
140
173
|
that stands between the network and your agents — put real auth in front
|
|
141
174
|
(reverse proxy + basic auth, SSH tunnel, etc.) before using that flag.
|
|
175
|
+
- **A proxy voids the loopback guarantee.** The production deployment (Ming's
|
|
176
|
+
suite) fronts this loopback bind with `tailscale serve` (`:8448 → 127.0.0.1:8765`),
|
|
177
|
+
so the tailnet reaches the UI despite the loopback bind: reads like
|
|
178
|
+
`/api/agents` are open there, and the static token is the *only* boundary on
|
|
179
|
+
actions. That is deliberate (owner-only tailnet + token-gated actions), but do
|
|
180
|
+
not read "loopback-only" as "not network-reachable" — a proxy in front is not
|
|
181
|
+
a trust boundary.
|
|
142
182
|
|
|
143
183
|
## Limits & known issues
|
|
144
184
|
|
|
@@ -170,7 +210,7 @@ This is the important part — a web page that can kill processes needs guardrai
|
|
|
170
210
|
- **Windows is untested.** Kill maps to `TerminateProcess` via psutil and may
|
|
171
211
|
work, but CI covers Linux + macOS only.
|
|
172
212
|
|
|
173
|
-
##
|
|
213
|
+
## Install & run
|
|
174
214
|
|
|
175
215
|
**Recommended — one command, nothing to install first:**
|
|
176
216
|
|
|
@@ -262,6 +302,19 @@ and `/metrics` series all use the derived label, so each instance gets its own
|
|
|
262
302
|
state. Sessions that don't match the regex keep their `agents:` label, and the
|
|
263
303
|
key is ignored where tmux isn't installed or running.
|
|
264
304
|
|
|
305
|
+
**Supervised fleets (`launchd_labels:`, macOS)** — the same problem for agents
|
|
306
|
+
that run as launchd jobs and never touch tmux (five `hermes` LaunchAgents all
|
|
307
|
+
landing as "hermes"). The launchd job label is their durable identity:
|
|
308
|
+
|
|
309
|
+
```yaml
|
|
310
|
+
launchd_labels: "^ai\\.hermes\\.(?:gateway-)?(.+)$" # ai.hermes.gateway-frontdoor → frontdoor
|
|
311
|
+
```
|
|
312
|
+
|
|
313
|
+
When a matched root's own launchd job label matches, the first capture group
|
|
314
|
+
(the whole label if there's no group) becomes the row label — with the same
|
|
315
|
+
per-instance churn/alert/metrics identity as `tmux_labels:`. `tmux_labels`
|
|
316
|
+
wins when both apply; roots with no matching job keep their `agents:` label.
|
|
317
|
+
|
|
265
318
|
**Which `agents.yaml` is used** — resolved once at startup, first hit wins:
|
|
266
319
|
|
|
267
320
|
1. `AGENTS_CONFIG=/path/to/agents.yaml` env var (the `--config` flag sets this)
|
|
@@ -330,7 +383,7 @@ Linux (systemd), `~/.config/systemd/user/agent-usage-manager.service`:
|
|
|
330
383
|
[Unit]
|
|
331
384
|
Description=agent usage manager
|
|
332
385
|
[Service]
|
|
333
|
-
ExecStart=%h/agent-usage-manager/.venv/bin/uvicorn app:app --port 8765
|
|
386
|
+
ExecStart=%h/agent-usage-manager/.venv/bin/uvicorn agent_usage_manager.app:app --port 8765
|
|
334
387
|
WorkingDirectory=%h/agent-usage-manager
|
|
335
388
|
Restart=on-failure
|
|
336
389
|
[Install]
|
|
@@ -355,6 +408,35 @@ SIGTERM/SIGKILL on POSIX and TerminateProcess on Windows.
|
|
|
355
408
|
|
|
356
409
|
## Release notes
|
|
357
410
|
|
|
411
|
+
### 0.2.5 — per-instance labels for supervised fleets + the honesty batch
|
|
412
|
+
|
|
413
|
+
- `launchd_labels:` config — per-instance row labels from launchd job labels,
|
|
414
|
+
mirroring `tmux_labels:`. A supervised fleet (several LaunchAgents on one
|
|
415
|
+
binary) no longer collapses into one blurred label for churn tracking, alert
|
|
416
|
+
transitions, and `/metrics`. `idle_ok:` suppression matches the base matcher
|
|
417
|
+
label as well, so a renamed instance (`hermes` → `frontdoor`) keeps its
|
|
418
|
+
class-level suppression.
|
|
419
|
+
- New `test-alert` subcommand: fires the configured `alerts.command` once,
|
|
420
|
+
synchronously, and reports the exit status — proves the alert channel before
|
|
421
|
+
an incident depends on it.
|
|
422
|
+
- `ignore:` now covers ChatGPT.app's embedded Codex helpers (renderer/service
|
|
423
|
+
processes, `Resources/codex` app-server) and the `Codex Computer Use`
|
|
424
|
+
desktop-automation app — they matched the `codex` pattern via the bundle
|
|
425
|
+
name and cluttered the dashboard with GUI plumbing.
|
|
426
|
+
- Config deletion is no longer silent: a vanished `agents.yaml` surfaces as a
|
|
427
|
+
`config_error` ("running on the last good config") instead of looking like
|
|
428
|
+
hot-reload still works.
|
|
429
|
+
- The kill confirm now names the agent (label + command line, not just PID)
|
|
430
|
+
and discloses that plain kill escalates SIGTERM → SIGKILL after 3s.
|
|
431
|
+
- The first-kill token prompt names the server's host (and the ssh one-liner)
|
|
432
|
+
for browsers viewing the dashboard over a tunnel.
|
|
433
|
+
- The DNS-rebinding 403 now names the remedy (use a loopback name or bare IP).
|
|
434
|
+
- `list` says so when flag fields can't populate (one-shot mode has no
|
|
435
|
+
history): a stderr note, and `"flags_available": false` in `--json`.
|
|
436
|
+
- `--host ::1` now produces a valid bracketed IPv6 URL for the browser open.
|
|
437
|
+
- Fixed the README's systemd unit (`uvicorn app:app` could never resolve the
|
|
438
|
+
module; it's `agent_usage_manager.app:app`).
|
|
439
|
+
|
|
358
440
|
### 0.2.4 — stop reporting non-agents
|
|
359
441
|
|
|
360
442
|
- `ignore:` now covers Sparkle's `Autoupdate`/`Updater` (Codex.app's equivalent of
|
|
@@ -417,6 +499,12 @@ SIGTERM/SIGKILL on POSIX and TerminateProcess on Windows.
|
|
|
417
499
|
closed; add `--unsafe-expose` only with auth in front (see
|
|
418
500
|
[Safety](#safety)).
|
|
419
501
|
|
|
502
|
+
## Stay in the loop
|
|
503
|
+
|
|
504
|
+
New tools and field notes on running AI agents with discipline go to the
|
|
505
|
+
[Agent Discipline](https://buttondown.com/minglong51) list first — launch
|
|
506
|
+
notes, operational patterns, early access. A few emails a month at most.
|
|
507
|
+
|
|
420
508
|
## License
|
|
421
509
|
|
|
422
510
|
MIT
|
|
@@ -61,6 +61,20 @@ ignore:
|
|
|
61
61
|
- updater # Sparkle's Updater.app; argv[1] is the app path, so it matches
|
|
62
62
|
- for chrome # "Codex for Chrome" extension host, not the agent
|
|
63
63
|
- tmux attach # tmux clients: `attach -t <session>` carries the agent's name
|
|
64
|
+
- chatgpt codex # ChatGPT.app's embedded Codex helpers ("Codex (Renderer)",
|
|
65
|
+
# "Codex (Service)", Resources/codex app-server) — GUI-app
|
|
66
|
+
# plumbing, matched via the "ChatGPT" bundle prefix; not agents
|
|
67
|
+
- codex computer use # ~/.codex/computer-use desktop-automation helper app
|
|
68
|
+
|
|
69
|
+
# Optional: per-instance labels from launchd job labels — the supervised-fleet
|
|
70
|
+
# counterpart of tmux_labels below. Several LaunchAgents running the same
|
|
71
|
+
# binary (e.g. ai.hermes.gateway, ai.hermes.gateway-frontdoor, …) all hit one
|
|
72
|
+
# agents: entry and land as indistinguishable rows, and tmux never sees them.
|
|
73
|
+
# When a matched root's own launchd job label matches this regex, the first
|
|
74
|
+
# capture group (the whole label if no group) becomes the row label — churn
|
|
75
|
+
# tracking, alert transitions, and /metrics all get per-instance identity.
|
|
76
|
+
# tmux_labels wins when both apply. Empty/absent = off.
|
|
77
|
+
launchd_labels: "^ai\\.hermes\\.(?:gateway-)?(.+)$" # → gateway, frontdoor, engineering_3
|
|
64
78
|
|
|
65
79
|
# Optional: per-instance labels from tmux session names. A fleet of identical
|
|
66
80
|
# agents (e.g. several claude-code bots, one per tmux session bot-*) all hit
|
|
@@ -73,6 +87,20 @@ ignore:
|
|
|
73
87
|
# their agents: label, so incidental tmux use never renames rows.
|
|
74
88
|
tmux_labels: "^bot-(.+)$" # session bot-coder_1 → row label coder_1
|
|
75
89
|
|
|
90
|
+
# Labels whose idle state is NORMAL — agents that wait for work (bots parked
|
|
91
|
+
# on a Telegram poll, gateways waiting for requests). They get no "idle"
|
|
92
|
+
# badge: badging the whole waiting fleet wallpapers the dashboard and trains
|
|
93
|
+
# badge-blindness (12 of 22 rows wore "idle 10m+" before this existed). Same
|
|
94
|
+
# reasoning as `idle` being opt-in for alerts below. Case-insensitive
|
|
95
|
+
# substrings of the row label (incl. tmux-derived ones), like ignore:.
|
|
96
|
+
idle_ok:
|
|
97
|
+
- admin
|
|
98
|
+
- coder # covers coder_1…coder_4
|
|
99
|
+
- hermes # covers the launchd_labels-renamed rows too (gateway,
|
|
100
|
+
# frontdoor, engineering_*) — suppression matches the BASE
|
|
101
|
+
# matcher label as well as the instance label
|
|
102
|
+
- codex
|
|
103
|
+
|
|
76
104
|
# GPU sampling: nvidia-smi is used automatically when present (Linux/NVIDIA).
|
|
77
105
|
# On Apple Silicon there is no per-process GPU API, so the GPU column is blank.
|
|
78
106
|
|
|
@@ -98,9 +126,14 @@ tmux_labels: "^bot-(.+)$" # session bot-coder_1 → row label coder_1
|
|
|
98
126
|
# Badges stay live on the dashboard + machinery panel; the feed line rides the
|
|
99
127
|
# next digest. Genuine breakage has louder, independent signals (mcp-autoheal
|
|
100
128
|
# GAVE UP, cron-health, sync-bot liveness). `hot` dropped from alerting
|
|
101
|
-
# entirely
|
|
129
|
+
# entirely.
|
|
130
|
+
# Ref discipline (2026-07-30): one stable ref per (agent, flag) + --upsert-open,
|
|
131
|
+
# so a repeat firing UPDATES the single open card in place instead of stacking
|
|
132
|
+
# one per cooldown expiry (nine churn cards in five days); the refreshed card
|
|
133
|
+
# carries the latest cpu/mem/pid snapshot. fyi level = feed-only, riding the
|
|
134
|
+
# digest exactly like the retired notify.py --feed-only wiring did.
|
|
102
135
|
alerts:
|
|
103
|
-
command: '
|
|
136
|
+
command: '"$HOME/workspace/founder-agent-os/bin/attention" add --level fyi --source usage-manager --title "$AUM_MSG" --body "$AUM_MSG" --ref "usage-manager:$AUM_LABEL:$AUM_FLAG" --upsert-open'
|
|
104
137
|
cooldown: 7200 # 2 h per (agent, flag)
|
|
105
138
|
flags: [churn, leak]
|
|
106
139
|
leak_floor_mb: 1536 # bots legitimately ratchet RSS as session
|