agent-dispatch 0.13.0__tar.gz → 0.15.0__tar.gz

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (35) hide show
  1. {agent_dispatch-0.13.0 → agent_dispatch-0.15.0}/AGENTS.md +12 -3
  2. {agent_dispatch-0.13.0 → agent_dispatch-0.15.0}/CHANGELOG.md +103 -0
  3. {agent_dispatch-0.13.0 → agent_dispatch-0.15.0}/PKG-INFO +99 -10
  4. {agent_dispatch-0.13.0 → agent_dispatch-0.15.0}/README.md +98 -9
  5. {agent_dispatch-0.13.0 → agent_dispatch-0.15.0}/agents.example.yaml +19 -0
  6. {agent_dispatch-0.13.0 → agent_dispatch-0.15.0}/pyproject.toml +1 -1
  7. {agent_dispatch-0.13.0 → agent_dispatch-0.15.0}/src/agent_dispatch/__init__.py +1 -1
  8. {agent_dispatch-0.13.0 → agent_dispatch-0.15.0}/src/agent_dispatch/cache.py +6 -0
  9. {agent_dispatch-0.13.0 → agent_dispatch-0.15.0}/src/agent_dispatch/cli.py +170 -1
  10. {agent_dispatch-0.13.0 → agent_dispatch-0.15.0}/src/agent_dispatch/config.py +1 -1
  11. {agent_dispatch-0.13.0 → agent_dispatch-0.15.0}/src/agent_dispatch/models.py +24 -0
  12. {agent_dispatch-0.13.0 → agent_dispatch-0.15.0}/src/agent_dispatch/runner.py +325 -52
  13. {agent_dispatch-0.13.0 → agent_dispatch-0.15.0}/src/agent_dispatch/server.py +70 -1
  14. agent_dispatch-0.15.0/src/agent_dispatch/servers.py +152 -0
  15. agent_dispatch-0.15.0/src/agent_dispatch/usage.py +347 -0
  16. agent_dispatch-0.15.0/tests/conftest.py +68 -0
  17. {agent_dispatch-0.13.0 → agent_dispatch-0.15.0}/tests/test_cache.py +15 -0
  18. {agent_dispatch-0.13.0 → agent_dispatch-0.15.0}/tests/test_cli.py +245 -18
  19. {agent_dispatch-0.13.0 → agent_dispatch-0.15.0}/tests/test_config.py +17 -0
  20. {agent_dispatch-0.13.0 → agent_dispatch-0.15.0}/tests/test_models.py +12 -3
  21. {agent_dispatch-0.13.0 → agent_dispatch-0.15.0}/tests/test_runner.py +390 -4
  22. {agent_dispatch-0.13.0 → agent_dispatch-0.15.0}/tests/test_server.py +250 -3
  23. agent_dispatch-0.15.0/tests/test_servers.py +140 -0
  24. agent_dispatch-0.15.0/tests/test_usage.py +294 -0
  25. agent_dispatch-0.13.0/tests/conftest.py +0 -41
  26. {agent_dispatch-0.13.0 → agent_dispatch-0.15.0}/.github/dependabot.yml +0 -0
  27. {agent_dispatch-0.13.0 → agent_dispatch-0.15.0}/.github/workflows/ci.yml +0 -0
  28. {agent_dispatch-0.13.0 → agent_dispatch-0.15.0}/.github/workflows/publish.yml +0 -0
  29. {agent_dispatch-0.13.0 → agent_dispatch-0.15.0}/.gitignore +0 -0
  30. {agent_dispatch-0.13.0 → agent_dispatch-0.15.0}/LICENSE +0 -0
  31. {agent_dispatch-0.13.0 → agent_dispatch-0.15.0}/SECURITY.md +0 -0
  32. {agent_dispatch-0.13.0 → agent_dispatch-0.15.0}/assets/mascot.png +0 -0
  33. {agent_dispatch-0.13.0 → agent_dispatch-0.15.0}/src/agent_dispatch/jobs.py +0 -0
  34. {agent_dispatch-0.13.0 → agent_dispatch-0.15.0}/tests/__init__.py +0 -0
  35. {agent_dispatch-0.13.0 → agent_dispatch-0.15.0}/tests/test_jobs.py +0 -0
@@ -12,6 +12,8 @@ MCP server + CLI that lets Claude Code agents delegate tasks to agents in other
12
12
  |------|------|
13
13
  | `src/agent_dispatch/runner.py` | Sync subprocess wrapper around `claude -p` — the actual work |
14
14
  | `src/agent_dispatch/server.py` | Async FastMCP interface (21 MCP tools), wraps runner in `asyncio.to_thread` + semaphore |
15
+ | `src/agent_dispatch/usage.py` | Append-only dispatch journal (cost/duration/outcome) + aggregation for `stats` and the `typical` block |
16
+ | `src/agent_dispatch/servers.py` | Registry of live `serve` processes and the version each one runs (`doctor` reads it) |
15
17
  | `src/agent_dispatch/cli.py` | Click CLI: `init`, `add`, `update`, `remove`, `list`, `describe`, `test`, `doctor`, `jobs`, `job`, `cancel`, `gc`, `group` (add/list/inspect/update/remove), `serve` |
16
18
  | `src/agent_dispatch/models.py` | Pydantic v2 models (`AgentConfig`, `DispatchGroup`/`GroupMember`, `Settings`, `DispatchResult`) |
17
19
  | `src/agent_dispatch/config.py` | YAML config load/save + project auto-description |
@@ -28,7 +30,7 @@ pip install -e ".[dev]"
28
30
 
29
31
  ```bash
30
32
  ruff check src/ tests/
31
- python3 -m pytest tests/ -v # 578 tests, ~5s
33
+ python3 -m pytest tests/ -v # 707 tests, ~17s
32
34
  ```
33
35
 
34
36
  Tests must **never** invoke the real `claude` CLI. Runner tests mock `shutil.which` + `subprocess.run`/`Popen`; server tests mock `_get_config` + `runner.dispatch`. The one exception is `TestStreamPipeHandling`, which spawns a short-lived *python* subprocess: a pipe deadlock lives in the OS pipe buffer, so a mocked `Popen` structurally cannot reproduce it.
@@ -60,7 +62,14 @@ Tests must **never** invoke the real `claude` CLI. Runner tests mock `shutil.whi
60
62
  - Pydantic does **not** validate on assignment. `Field(ge=...)` guards only the *load* path; every mutation surface (CLI `add`/`update`, MCP `add_agent`/`update_agent`) needs its own boundary check, or the bound escapes as a raw `ValidationError`.
61
63
  - Every state file (`agents.yaml`, job files) is written **temp file + `os.replace`**, never in place, and every load/mutate/save is wrapped in `config.ProcessLock` — the CLI and the MCP server are separate processes writing the same files, so a thread lock alone loses updates.
62
64
  - Anything that changes an agent's config must call `_invalidate_agent_cache` — the cache key holds the agent *name*, not its directory or permissions.
63
- - Only *clean* successes are cached: `cache.put` refuses failures, `denied_tools` results, and `budget_exceeded` results, so the documented "grant access, then re-dispatch" recovery is never short-circuited.
65
+ - Only *clean* successes are cached: `cache.put` refuses failures, `denied_tools` results, `budget_exceeded` results, and results whose `outcome` is `partial`/`blocked`, so the documented "grant access / fix the cause, then re-dispatch" recovery is never short-circuited.
66
+ - **The usage journal is written by a decorator, not at each return.** `runner._journaled` wraps `dispatch`/`dispatch_stream` because each has a dozen early returns; instrumenting them one by one guarantees the next new return path goes unrecorded. It records the outermost call only (`_use_session_flag` is False on the stream's old-CLI retry, which would otherwise double-count one dispatch). It also swallows exceptions itself even though `usage.record` already does: the result in hand is billed, and the guarantee has to hold at the boundary that owns the damage rather than on another module's promise (`test_a_broken_journal_never_breaks_a_paid_dispatch` removes record()'s guard to prove it).
67
+ - **A journal record is one `O_APPEND` write under `PIPE_BUF`**, never a locked read-modify-write: the CLI and every running server share one file, and this user has 14+ servers alive. Every field is capped so the line cannot approach 4096 bytes. Rotation is an atomic rename at ~2 MB, two generations kept.
68
+ - **Profiles are memoized on the journal's (path, size, mtime) and read a NARROWER tail than `stats`.** `list_agents`/`inspect_agent` run on the event-loop thread: a full 2 MB journal cost **24.8 ms** per discovery call at the 512 KB report window — 33x the config parse that 0.13.0 exists to have fixed. Now 3.3 ms cold, ~0 warm. Every profile consumer goes through `usage.profiles()`; never add a caller that re-parses per agent.
69
+ - **`usage_limit` is its own error type** (observed live: `You've hit your session limit · resets 6pm`). It is checked *before* the permission patterns and its hint says the limit is account-wide, so "try another agent" is not a workaround. Text-classified errors attach their advisory through the single `_classification_hint`, not per site.
70
+ - **Server liveness is an advisory lock, not a PID.** `servers.register` holds an exclusive `flock` on `<config dir>/servers/<pid>.json` for the process lifetime; a reader that *takes* that lock owns a dead entry and deletes it. `os.kill(pid, 0)` would believe a recycled PID and would never notice a SIGKILLed server. Never close that fd (re-registering closes the old one first).
71
+ - **The dispatch protocol goes in the system prompt, not the task.** `_build_system_prompt` (runner.py) renders the protocol (non-interactive, time/spend budget, lead with the outcome, trailing `STATUS:` line) plus the agent's `instructions`, and `_build_command` passes it as `--append-system-prompt` — on `--resume` too, since the CLI re-applies an appended prompt on every launch. `_build_prompt` (the `-p` text) is unchanged, so the cache key is unchanged. The text always starts with a `##` header line so it can never be read as a flag. `settings.dispatch_protocol: false` turns the protocol off (for CLIs that predate the flag; `doctor` probes `claude --help` for it); per-agent `instructions` still go through when set.
72
+ - **`outcome` is lifted from the LAST line of the result** (`_split_outcome`), before JSON parsing, by the shared `_build_success_result` (the success-side twin of `_build_error_result` — both `dispatch` and `dispatch_stream` go through it) and by the plain-text fallback tier. A bare marker line is removed from `result`; a marker with a trailing reason stays. `outcome` never flips `success`; `partial`/`blocked` add a `hint` with the `dispatch_session(...)` continuation, after the denial hint. In JSON mode the protocol omits the STATUS instruction (the JSON footer governs), but a STATUS line that arrives anyway is still stripped so `parsed_result` survives.
64
73
  - Remediation text is a contract: a hint that names a flag must name one that exists (`test_printed_budget_hint_is_a_runnable_command` feeds the printed flags back into the CLI). Run the command you print.
65
74
  - The config error sets are declared **once** and in two halves: `config.CONFIG_LOAD_ERRORS` (read) and `config.CONFIG_SAVE_ERRORS` (write — `yaml.dump`'s `RepresenterError` is a `yaml.YAMLError`, therefore neither `OSError` nor `ConfigLoadError`, and used to escape both the MCP guard and the CLI's `_save_or_exit`). Two halves, not one set, because the remediations differ: a failed write is atomic so the old config survives, while a failed read needs the YAML fixed.
66
75
  - MCP tools that load config carry `@_config_guard` under `@mcp.tool()` so a broken `agents.yaml` — or a failed *write* — returns the `{"error": ...}` envelope instead of a raw traceback. The set of load errors lives in one place (`config.CONFIG_LOAD_ERRORS`) because three surfaces handle it: **`UnicodeDecodeError` is a `ValueError`, not an `OSError`**, and listing types per-site is exactly how a cp1251 config slipped past all three.
@@ -85,4 +94,4 @@ Python ≥ 3.10 · `from __future__ import annotations` everywhere · Pydantic v
85
94
 
86
95
  ## More detail
87
96
 
88
- [README.md](README.md) documents every MCP tool with parameter tables, response shapes, and the error-recovery map — it doubles as the behavioral spec. The test suite (`tests/`, 578 tests) encodes the exact expected behavior of every layer: when in doubt, read the tests for the module you're touching (`test_runner.py`, `test_server.py`, `test_cli.py`, ...).
97
+ [README.md](README.md) documents every MCP tool with parameter tables, response shapes, and the error-recovery map — it doubles as the behavioral spec. The test suite (`tests/`, 707 tests) encodes the exact expected behavior of every layer: when in doubt, read the tests for the module you're touching (`test_runner.py`, `test_server.py`, `test_cli.py`, ...).
@@ -7,6 +7,109 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
7
7
 
8
8
  ## [Unreleased]
9
9
 
10
+ ## [0.15.0] - 2026-09-09
11
+
12
+ The measurement round. 0.14.0 told a dispatched agent how to behave; this one
13
+ records what actually happened. Four live dispatches through the new protocol
14
+ came back **4/4 with a STATUS line and 0/4 ending in a clarifying question** —
15
+ including the case it was built for: an ambiguous task ("how many users signed
16
+ up in 30 days?", no project named) where the agent stated its assumption,
17
+ counted across every database, and returned `done` rather than asking. A
18
+ tool-less agent correctly returned `blocked`. No `partial` has been observed in
19
+ the wild yet.
20
+
21
+ ### Added
22
+ - **Usage journal + `agent-dispatch stats`.** Every dispatch appends one line to
23
+ `usage.jsonl` — agent, ok, cost, duration, turns, outcome, error type, caller;
24
+ cache hits too, flagged `cached`. `stats [--days N --agent X --json]` turns it
25
+ into spend, median/p90 durations, outcome and failure counts per agent. Until
26
+ now `cost_usd` and `duration_ms` were returned once and discarded, so "which
27
+ agent burns money" had no answer and the only timing knowledge in the system
28
+ was hand-written prose in agent descriptions. Recording is one `O_APPEND`
29
+ write capped under `PIPE_BUF`, so the CLI and all running servers share one
30
+ file with no lock; it rotates at ~2 MB keeping two generations; a failure to
31
+ write it can never fail a dispatch. Off with `settings.usage_log: false`.
32
+ - **`typical` in `list_agents` and `inspect_agent`** — measured median/p90
33
+ seconds and median cost from that agent's own recent runs, so a caller can
34
+ size `timeout_seconds` from data instead of from a description. Absent below
35
+ three recorded dispatches: no data beats a "typical" derived from two runs.
36
+ Costs ~115 bytes per agent (3.6% of a discovery payload).
37
+ - **`error_type: "usage_limit"`** for the Claude *account's* rate/session limit,
38
+ observed live during this round's evaluation (`You've hit your session limit ·
39
+ resets 6pm`) — it used to land in the generic `cli_error` bucket. The hint
40
+ names the reset time and says the limit is account-wide, so dispatching a
41
+ different agent is not a workaround, and nothing was billed.
42
+ - **`doctor` now names the sessions running stale code.** Each server registers
43
+ `<config dir>/servers/<pid>.json` and holds an advisory lock on it for life;
44
+ `doctor` reports live servers by version and prints the pid and age of every
45
+ one still on an older release. A Python process holds its modules for life, so
46
+ an upgrade reaches an open Claude Code session only when it restarts — on this
47
+ machine 18 servers were running previous code with no way to tell which.
48
+ Liveness is the lock, not the PID: a recycled PID cannot fake a live server
49
+ and a SIGKILLed one cannot linger.
50
+
51
+ ### Changed
52
+ - **A timeout error now suggests a number, not a doubling.** With enough history
53
+ the message names the value that would have covered this agent's real p90
54
+ (padded 50%). Doubling the current timeout was wrong in both directions.
55
+
56
+ ### Fixed
57
+ - Discovery no longer pays for the journal on the event loop: profiles are
58
+ memoized on the journal's (path, size, mtime) and read a narrower tail than
59
+ the report. Measured on a full 2 MB journal, `list_agents` went from
60
+ **24.8 ms to 3.3 ms cold and ~0.01 ms warm** — the 512 KB version would have
61
+ been 33x the config parse that 0.13.0 exists to have eliminated.
62
+
63
+
64
+ ## [0.14.0] - 2026-09-08
65
+
66
+ The delegation round: what a dispatched agent is *told*, and what it tells
67
+ back. Until now `claude -p` was launched with the task and nothing else — it
68
+ did not know it was being driven by another agent, that nobody would answer a
69
+ question, how long it had, or how to report an unfinished job. The failure
70
+ mode was concrete and billed: an ambiguous task ended in *"Could you clarify
71
+ which service you mean?"*, `success: true`, cached for the whole TTL.
72
+
73
+ ### Added
74
+ - **The dispatch protocol.** Every dispatch now appends a short system prompt
75
+ (`--append-system-prompt`, never mixed into the task text) telling the agent
76
+ that it is non-interactive and dispatched by `caller`, that it must state
77
+ assumptions instead of asking, that a denied tool is something to report and
78
+ work around rather than stop on, what its time budget (and spend cap) is, to
79
+ lead with the outcome — which is what makes a `return_ref` summary, the
80
+ *head* of the text, worth reading — and to end with one line:
81
+ `STATUS: done | partial | blocked`. Passed on `--resume` too. Verified live
82
+ against claude 2.1.263 (`agent-dispatch test <agent> --stream`).
83
+ - **`DispatchResult.outcome`** — that trailing line, lifted out of `result`
84
+ into a field: `done`, `partial` or `blocked`, or absent when the agent did
85
+ not report one. It never flips `success`. `partial`/`blocked` come with a
86
+ `hint` carrying the exact `dispatch_session(..., session_id=...)` call to
87
+ continue, ride the `return_ref` payload and `dispatch_jobs` summaries, are
88
+ shown by `agent-dispatch test` / `job <id>`, are labelled for the
89
+ `dispatch_parallel` aggregator so a blocked member is not synthesized as a
90
+ finished one — and are **not cached**: whatever the agent was missing is not
91
+ in the cache key, so a retry after fixing it must run fresh.
92
+ - **Per-agent `instructions`** — standing orders appended after the protocol
93
+ on every dispatch ("read-only SQL only", "never restart a stack unless the
94
+ task says so"). `add_agent`/`update_agent` (`"none"` clears), CLI `add` /
95
+ `update --instructions` (`none` clears), shown by `inspect_agent` and
96
+ `describe`, pruned from YAML when empty. Changing them invalidates the
97
+ agent's cache entries like any other config change.
98
+ - **`settings.dispatch_protocol`** (default `true`) — set to `false` for a raw
99
+ `claude -p` (a CLI that predates `--append-system-prompt`, or an A/B). The
100
+ protocol and the instructions share one flag, so instructions still go
101
+ through when set. `doctor` now probes `claude --help` for the flag and warns
102
+ with the exact remediation when it is missing.
103
+
104
+ ### Changed
105
+ - In `response_format="json"` mode the protocol omits the "lead with the
106
+ outcome" and STATUS bullets — the JSON footer governs the reply shape. A
107
+ STATUS line that arrives anyway is still stripped before parsing, so
108
+ `parsed_result` survives it.
109
+ - `dispatch` and `dispatch_stream` build their success result through one
110
+ `_build_success_result`, the twin of `_build_error_result`.
111
+
112
+
10
113
  ## [0.13.0] - 2026-08-13
11
114
 
12
115
  An efficiency round, measured against a real 38 KB config (4 agents, 6 groups),
@@ -1,6 +1,6 @@
1
1
  Metadata-Version: 2.5
2
2
  Name: agent-dispatch
3
- Version: 0.13.0
3
+ Version: 0.15.0
4
4
  Summary: MCP server that lets Claude Code agents delegate tasks to agents in other project directories
5
5
  Project-URL: Homepage, https://github.com/ginkida/agent-dispatch
6
6
  Project-URL: Repository, https://github.com/ginkida/agent-dispatch
@@ -134,13 +134,16 @@ Lists all configured agents. **Call this first** to see what's available.
134
134
  "capabilities": ["docker_logs", "deploy_debug"],
135
135
  "risky_capabilities": ["restart_services"],
136
136
  "permission_mode": "bypassPermissions",
137
- "allowed_tools": ["Bash", "Read", "Grep"]
137
+ "allowed_tools": ["Bash", "Read", "Grep"],
138
+ "typical": {"median_seconds": 143.3, "p90_seconds": 178.3, "median_cost_usd": 0.2666}
138
139
  }
139
140
  ]
140
141
  ```
141
142
 
142
143
  `mcp_servers`, `stacks`, and `dbs` are detected from the agent's project files (`.mcp.json`, `Dockerfile`, `pyproject.toml`, `Cargo.toml`, `prisma/`, `alembic.ini`, etc.) so callers can pick the right agent without dispatching a probe.
143
144
 
145
+ `typical` is **measured**, not declared: it comes from this agent's own recent dispatches in the [usage journal](#usage-journal--what-your-fleet-actually-costs) and appears only after a few runs. Use it to size `timeout_seconds` and to know what a call will cost before you make it. It is absent when the journal is off or the agent has too little history — no data is better than a typical duration derived from two runs.
146
+
144
147
  ### `inspect_agent`
145
148
 
146
149
  Cheap detailed lookup — reads the agent's files without spawning a `claude` session. Returns the full config (timeout, model, budget, permission mode, allowed/disallowed tools), detected MCP/stacks/DBs, plus short previews of `CLAUDE.md` and `README.md` when present.
@@ -152,6 +155,15 @@ Cheap detailed lookup — reads the agent's files without spawning a `claude` se
152
155
 
153
156
  Use this **before** `dispatch_async`/`dispatch` to confirm an agent has the tools and context for your task — much cheaper than a probe dispatch.
154
157
 
158
+ It also carries the agent's `instructions` (its standing orders) and the full `typical` block — the trimmed one in `list_agents` plus `dispatches`, `failed` and the `outcomes` breakdown:
159
+
160
+ ```json
161
+ "typical": {
162
+ "dispatches": 24, "median_seconds": 143.3, "p90_seconds": 178.3,
163
+ "median_cost_usd": 0.2666, "failed": 1, "outcomes": {"done": 21, "partial": 2}
164
+ }
165
+ ```
166
+
155
167
  ### Groups
156
168
 
157
169
  A **group** bundles related agents into a cross-project working set — typically a few code repos plus capability gateways (an `infra` agent with a Portainer MCP, an `analytics` agent with a browser + Yandex Metrica). It lets one orchestrating session coordinate work that spans code, deploy, and verification.
@@ -220,7 +232,8 @@ dispatch(
220
232
  "session_id": "sess-abc-123",
221
233
  "cost_usd": 0.02,
222
234
  "duration_ms": 5000,
223
- "num_turns": 2
235
+ "num_turns": 2,
236
+ "outcome": "done"
224
237
  }
225
238
 
226
239
  // Response (failure — error_type helps you handle programmatically)
@@ -233,7 +246,7 @@ dispatch(
233
246
  }
234
247
  ```
235
248
 
236
- **`error_type` values:** `permission` (tool/action denied), `timeout`, `recursion` (dispatch depth exceeded), `not_found` (missing directory or CLI), `budget` (the `claude` CLI stopped the session at `max_budget_usd`), `cli_error` (other failures). Permission and budget errors include an actionable hint.
249
+ **`error_type` values:** `permission` (tool/action denied), `timeout`, `recursion` (dispatch depth exceeded), `not_found` (missing directory or CLI), `budget` (the `claude` CLI stopped the session at `max_budget_usd`), `usage_limit` (the Claude *account* hit its rate/session limit), `cli_error` (other failures). Permission, budget, usage-limit and timeout errors include an actionable hint.
237
250
 
238
251
  **Resumable timeouts:** every fresh dispatch pre-assigns a session UUID (`--session-id`), so a timed-out dispatch still returns a `session_id` — the partial transcript survives the kill. The timeout error spells out the recovery: resume with `dispatch_session(agent, "Continue where you left off", session_id=...)`, retry with a bigger `timeout_seconds`, or use `dispatch_async`.
239
252
 
@@ -278,6 +291,30 @@ Error: TypeError at scheduler.py:42
278
291
  Check container logs for recent errors related to the scheduler service
279
292
  ```
280
293
 
294
+ **The dispatch protocol and `outcome`.** `claude -p` on its own does not know it is being driven by another agent: given an ambiguous task it will happily end with *"Could you clarify which service you mean?"* — a billed run that answered nothing, and one that a plain cache would then serve again for the whole TTL. So every dispatch also appends a short **dispatch protocol** to the agent's system prompt (`--append-system-prompt`, never mixed into your task text):
295
+
296
+ - it is running non-interactively, dispatched by `caller`, and nobody will answer a question or approve an action — state the assumption and proceed, take the safer reading when two differ;
297
+ - a denied tool or a missing thing is something to report, not something to stop on — finish everything else that is possible;
298
+ - its time budget is about the agent's timeout (and its spend cap, if one is set) — scope the work to fit; a complete partial answer beats an unfinished perfect one;
299
+ - lead with the outcome, then the evidence; list what could not be done and why (this is what makes a `return_ref` summary — the **head** of the text — worth reading);
300
+ - end with one line, `STATUS: done`, `STATUS: partial` or `STATUS: blocked`.
301
+
302
+ That last line is lifted out of `result` into **`outcome`** — the agent's own verdict, deterministic to check: `"done"` means complete; `"partial"` / `"blocked"` mean the agent says the work is unfinished, the result names what is missing, and a `hint` spells out the continuation (usually `dispatch_session(..., session_id=...)`). Read `outcome` before reading the text. It never flips `success`, `partial` / `blocked` results are **not cached**, and a `dispatch_parallel(..., aggregate=...)` labels such members for the aggregator so a blocked report is not synthesized as a finished one. Absent when the agent did not report one — the protocol is off (`settings.dispatch_protocol: false`), `response_format="json"` was requested (the JSON footer governs the reply shape there), or the agent simply skipped it.
303
+
304
+ **Standing orders.** Per-agent `instructions` (set via `add_agent` / `update_agent` or `agent-dispatch update <name> --instructions "..."`) follow the protocol in the same system prompt on **every** dispatch — "read-only SQL only", "never restart a stack unless the task says so", "answer with exact log lines". Unlike `context`, which is per call, and unlike the project's own `CLAUDE.md`, which is written for an interactive session, these are the rules for *being dispatched*.
305
+
306
+ ```json
307
+ // Response (success, but the agent says it did not finish)
308
+ {
309
+ "agent": "infra",
310
+ "success": true,
311
+ "result": "Restarted horizon. Could not verify the queue drained: the redis container is not reachable from here.",
312
+ "session_id": "sess-abc-123",
313
+ "outcome": "partial",
314
+ "hint": "The agent reports its work is PARTIAL — read the result for what is missing, then continue in the same session via dispatch_session(agent='infra', task='Continue where you left off', session_id='sess-abc-123') or re-dispatch with what it needed."
315
+ }
316
+ ```
317
+
281
318
  ### `dispatch_session`
282
319
 
283
320
  Multi-turn: continue a conversation with an agent. First call starts a session, pass `session_id` back to continue. Never cached.
@@ -395,6 +432,7 @@ Register a new project directory as an agent. Description is auto-generated from
395
432
  | `disallowed_tools` | string | no | Comma-separated disallowed tools |
396
433
  | `capabilities` | string | no | Comma-separated capability labels (e.g. `"docker_logs,deploy_debug"`) |
397
434
  | `risky_capabilities` | string | no | Comma-separated high-risk labels (e.g. `"restart_services"`) |
435
+ | `instructions` | string | no | Standing orders appended to the agent's system prompt on every dispatch (see [the dispatch protocol](#dispatch)) |
398
436
 
399
437
  ### `update_agent`
400
438
 
@@ -412,6 +450,7 @@ Update an existing agent's configuration. Only non-empty fields are changed. Pas
412
450
  | `disallowed_tools` | string | no | Comma-separated. `"none"` to clear |
413
451
  | `capabilities` | string | no | Comma-separated. `"none"` to clear |
414
452
  | `risky_capabilities` | string | no | Comma-separated. `"none"` to clear |
453
+ | `instructions` | string | no | Standing orders for every dispatch (replaces the text). `"none"` to clear |
415
454
 
416
455
  Changing an agent's config drops that agent's cached results — the cache key holds the agent *name*, so a re-pointed or re-permissioned agent would otherwise keep answering from the previous config for the rest of the TTL. The same applies to `add_agent` and `remove_agent`.
417
456
 
@@ -506,15 +545,17 @@ Failures are deterministic: check `success`, then branch on `error_type`.
506
545
  | `error_type` | Meaning | Recovery |
507
546
  |--------------|---------|----------|
508
547
  | `permission` | A tool call was denied | `update_agent(name, allowed_tools="Bash,Read")` (least privilege) or `update_agent(name, permission_mode="bypassPermissions")`, then re-dispatch. The `error` text includes a hint with the exact fix. |
509
- | `timeout` | Process killed at the timeout | Resume the partial work: `dispatch_session(agent, "Continue where you left off", session_id=<from the error text>)`. Or retry with a bigger `timeout_seconds=`, or use `dispatch_async`. A *streaming* dispatch that produced its answer before the deadline returns that answer with a `hint` instead of failing. |
548
+ | `timeout` | Process killed at the timeout | Resume the partial work: `dispatch_session(agent, "Continue where you left off", session_id=<from the error text>)`. Or retry with a bigger `timeout_seconds=` — once the journal has history the error names the value that would have covered this agent's p90 — or use `dispatch_async`. A *streaming* dispatch that produced its answer before the deadline returns that answer with a `hint` instead of failing. |
510
549
  | `not_found` | Agent directory or `claude` CLI missing | `list_agents()` → check `healthy`. Re-add the agent with an existing path, or run `agent-dispatch doctor` to find what's missing. |
511
550
  | `recursion` | Dispatch nesting exceeded `max_dispatch_depth` (default 3) | Don't dispatch from dispatched agents; if the nesting is intentional, raise `max_dispatch_depth` in settings. |
512
551
  | `budget` | The `claude` CLI ended the session at the `max_budget_usd` spend cap — the answer is incomplete | Raise the cap (`update_agent(name, max_budget_usd=2.0)`), switch to a cheaper `model`, or split the task. The partial session is resumable: `dispatch_session(agent, "Continue where you left off", session_id=<from the result>)`. |
552
+ | `usage_limit` | The Claude **account** hit its usage/rate limit — not this agent, and nothing was billed | Wait for the reset named in the `error` text. Dispatching a *different* agent is not a workaround: every agent runs on the same account. If it recurs under load, lower `settings.max_concurrency`. |
513
553
  | `cli_error` | Anything else from the `claude` subprocess | Read the `error` text; run `agent-dispatch doctor` for environment issues; retry once if transient. |
514
554
 
515
555
  Three soft signals that arrive with `success: true`:
516
556
 
517
557
  - **`denied_tools` + `hint`** — the agent finished but some tool calls were blocked; the result may be incomplete. Grant access (see the `permission` row) and re-dispatch.
558
+ - **`outcome: "partial"` / `"blocked"`** — the agent's own verdict that it did not finish; the result names what is missing and the `hint` carries the `dispatch_session(...)` call to continue in the same session. Not cached, so a re-dispatch after fixing the cause runs fresh.
518
559
  - **`parsed_result: null` with `response_format="json"`** — the reply wasn't valid JSON; the raw text is still in `result`. Caveat: an agent that *can't* comply returns `{"error": "<reason>"}` — which parses successfully — so also check `parsed_result` for an `"error"` key.
519
560
  - **`budget_exceeded: true`** — `cost_usd` came in over the agent's `max_budget_usd` (or the settings default) without the CLI stopping the run (the final turn can overshoot the cap). The dispatch is not failed — the money is already spent — but a runaway agent is now visible. Tighten the task, pick a cheaper model, or raise the budget. A run the CLI *did* stop fails with `error_type: "budget"` instead.
520
561
 
@@ -539,6 +580,9 @@ agents:
539
580
  - deploy_debug
540
581
  risky_capabilities: # high-risk labels, surfaced for visibility
541
582
  - restart_services
583
+ # instructions: | # standing orders, appended to the system prompt on every dispatch
584
+ # Read-only: never restart or redeploy unless the task says so.
585
+ # Quote exact log lines with timestamps.
542
586
  # model: sonnet # optional model override
543
587
  # max_budget_usd: 1.0 # cost limit per dispatch
544
588
  # permission_mode: bypassPermissions # one of: default | plan | bypassPermissions
@@ -574,6 +618,11 @@ settings:
574
618
  # - Edit
575
619
  max_dispatch_depth: 3 # recursion protection
576
620
  max_concurrency: 5 # max parallel claude -p processes (per dispatch path)
621
+ # dispatch_protocol: true # send every agent the dispatch protocol (non-interactive,
622
+ # # time budget, STATUS line → `outcome`). false = raw claude -p.
623
+ # usage_log: true # record every dispatch in usage.jsonl (cost, duration, outcome).
624
+ # # Powers `agent-dispatch stats`, the `typical` block and the
625
+ # # measured timeout suggestion. false = record nothing.
577
626
  # job_retention_days: 30 # 0 (default) = never prune. See "Job retention" below.
578
627
  cache:
579
628
  enabled: true
@@ -583,6 +632,42 @@ settings:
583
632
 
584
633
  Config is reloaded on every tool call — add agents without restarting.
585
634
 
635
+ ### Usage journal — what your fleet actually costs
636
+
637
+ Every dispatch appends one line to `~/.config/agent-dispatch/usage.jsonl`
638
+ (override with `AGENT_DISPATCH_USAGE_LOG`): agent, success, cost, duration,
639
+ turns, `outcome`, error type, caller. Cache hits are recorded too, flagged
640
+ `cached`, so the report can show what the cache saves.
641
+
642
+ ```console
643
+ $ agent-dispatch stats --days 7
644
+ Usage (last 7 day(s))
645
+ dispatches: 33, 1 served from cache
646
+ spend: $13.0390
647
+ duration: median 87s, p90 2m40s, max 10m00s
648
+ outcomes: blocked 1, done 27, partial 3
649
+ failures: timeout 1, usage_limit 1
650
+
651
+ Per agent
652
+ analytic 18 runs $ 8.9441 median 2m04s p90 2m51s
653
+ done 15, partial 3
654
+ gitlab 12 runs $ 3.7949 median 60s p90 86s
655
+ done 12 | 1 cached
656
+ ```
657
+
658
+ `--agent NAME` narrows it, `--json` emits the same report as JSON.
659
+
660
+ The journal is what makes three other things work: the `typical` block in
661
+ `list_agents` / `inspect_agent`, the timeout error naming a value derived from
662
+ the agent's real p90 instead of a doubling guess, and any answer at all to
663
+ "which agent is expensive". Turn it off with `usage_log: false` in settings.
664
+
665
+ Properties worth knowing: the file is owner-only (`0o600`); each record is a
666
+ single `O_APPEND` write capped well under `PIPE_BUF`, so the CLI and every
667
+ running server can share it with no lock and no torn lines; it rotates to
668
+ `usage.jsonl.1` at ~2 MB and keeps two generations, so it is bounded at ~4 MB
669
+ forever. A failure to write it can never fail a dispatch.
670
+
586
671
  ### Job retention
587
672
 
588
673
  Every `dispatch_async` **and** every `dispatch(..., return_ref=True)` writes a
@@ -646,8 +731,11 @@ agent-dispatch MCP server
646
731
  ▼
647
732
  New Claude Code session in ~/projects/infra/
648
733
  ├─ Inherits: CLAUDE.md, .mcp.json, project tools
734
+ ├─ System prompt += dispatch protocol + the agent's standing `instructions`
649
735
  ├─ Receives structured prompt with goal/caller/context/task
650
- └─ Returns result → cached for future identical requests
736
+ └─ Returns result (+ its own STATUS → `outcome`) → cached when complete
737
+ │
738
+ └─ one line appended to usage.jsonl (cost, duration, outcome)
651
739
  ```
652
740
 
653
741
  ## Safety
@@ -655,11 +743,11 @@ agent-dispatch MCP server
655
743
  - **Recursion protection** — `AGENT_DISPATCH_DEPTH` env var tracks nesting. Default limit: 3. Best-effort across the subprocess boundary (see [SECURITY.md](SECURITY.md)).
656
744
  - **Argument-injection guard** — structured CLI fields (`session_id`, `model`, `permission_mode`, tool names) that start with `-` are rejected so they can't smuggle extra `claude` flags.
657
745
  - **Path-traversal guard** — caller-supplied `job_id`/`ref` values are validated as 32-char hex before any filesystem access.
658
- - **Owner-only state** — job files (`0o600`) and `agents.yaml` (`0o600`) are written for the owner only; their directories are `0o700`.
746
+ - **Owner-only state** — job files, `agents.yaml`, the usage journal and the server registry are all written `0o600`; their directories are `0o700`.
659
747
  - **Cost control** — `max_budget_usd` per agent or globally is passed to the `claude` CLI as `--max-budget-usd`, so a runaway dispatch is stopped at the cap and comes back as `error_type: "budget"` with a resumable `session_id`. An overshoot that lands over budget without stopping is flagged post-hoc with `budget_exceeded: true` + a hint.
660
748
  - **Concurrency** — `max_concurrency` (default: 5) caps parallel `claude -p` processes. Note: the sync and async dispatch paths use separate semaphores, so the worst-case total is `2 × max_concurrency`.
661
749
  - **Timeout** — per-agent or global (default: 300s). A streaming dispatch runs the agent in its own process group, so the deadline kills the whole tree: a process the agent left running in the background can't hold the dispatch (and its concurrency slot) open past the timeout.
662
- - **Caching** — identical `(agent, task, context, caller, goal, response_format)` requests return cached results, bounded by `cache.max_size` (oldest entry evicted first). Only clean successes are cached: failures, results with `denied_tools`, and results flagged `budget_exceeded` are not, so the documented "grant access / raise the cap, then re-dispatch" recovery is never served a stale crippled answer. Changing an agent's config invalidates its entries. Sessions and dialogues are never cached. A `group=` dispatch folds the group's `shared_context` into `context`, so different groups cache separately and a plain dispatch is unaffected.
750
+ - **Caching** — identical `(agent, task, context, caller, goal, response_format)` requests return cached results, bounded by `cache.max_size` (oldest entry evicted first). Only clean successes are cached: failures, results with `denied_tools`, results flagged `budget_exceeded`, and results the agent itself reported as `partial` / `blocked` are not, so the documented "grant access / raise the cap, then re-dispatch" recovery is never served a stale crippled answer. Changing an agent's config invalidates its entries. Sessions and dialogues are never cached. A `group=` dispatch folds the group's `shared_context` into `context`, so different groups cache separately and a plain dispatch is unaffected.
663
751
  - **Durable config** — `agents.yaml` is written atomically (temp file + rename), so an interrupted write can never truncate it. Every mutation path (CLI and MCP server alike) also takes a cross-process advisory lock, so concurrent edits don't drop one another's agents. The lock is best-effort by design: after waiting 10 seconds it logs a warning and proceeds anyway, because a wedged lock holder must not freeze the MCP server — so on a heavily contended config a lost update is possible, while a truncated one is not.
664
752
 
665
753
  See [SECURITY.md](SECURITY.md) for the full threat model (including the `bypassPermissions` escalation risk and on-disk job files).
@@ -670,13 +758,14 @@ See [SECURITY.md](SECURITY.md) for the full threat model (including the `bypassP
670
758
  |---------|-------------|
671
759
  | `agent-dispatch init` | Create config + register MCP server with Claude Code |
672
760
  | `agent-dispatch add <name> <dir>` | Add an agent (auto-generates description) |
673
- | `agent-dispatch update <name>` | Update agent config (permissions, timeout, model, etc.) |
761
+ | `agent-dispatch update <name>` | Update agent config (permissions, timeout, model, `--instructions`, etc.) |
674
762
  | `agent-dispatch remove <name>` | Remove an agent |
675
763
  | `agent-dispatch list` | List agents with health status and permissions |
676
764
  | `agent-dispatch group <add\|list\|inspect\|update\|remove>` | Manage [groups](#groups) — cross-project working sets of agents |
677
765
  | `agent-dispatch describe <name>` | Show full configuration for one agent (tri-state tools, project files) |
678
766
  | `agent-dispatch test <name> [task] [--stream]` | Test an agent with a dispatch (`--stream` for live progress) |
679
- | `agent-dispatch doctor` | Diagnose installation: Claude CLI, MCP registration, agent health, and group membership |
767
+ | `agent-dispatch stats [--days N --agent X --json]` | What dispatches cost: spend, durations, outcomes and failures per agent |
768
+ | `agent-dispatch doctor` | Diagnose installation: Claude CLI (incl. `--append-system-prompt` support), running servers on stale code, MCP registration, agent health, and group membership |
680
769
  | `agent-dispatch jobs [--status --limit]` | List async dispatch jobs (most recent first) |
681
770
  | `agent-dispatch job <id>` | Show one job: status, progress tail, result preview |
682
771
  | `agent-dispatch cancel <id>` | Cancel a pending job (running jobs: use the `dispatch_cancel` MCP tool) |