agent-dispatch 0.13.0__tar.gz → 0.15.0__tar.gz
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- {agent_dispatch-0.13.0 → agent_dispatch-0.15.0}/AGENTS.md +12 -3
- {agent_dispatch-0.13.0 → agent_dispatch-0.15.0}/CHANGELOG.md +103 -0
- {agent_dispatch-0.13.0 → agent_dispatch-0.15.0}/PKG-INFO +99 -10
- {agent_dispatch-0.13.0 → agent_dispatch-0.15.0}/README.md +98 -9
- {agent_dispatch-0.13.0 → agent_dispatch-0.15.0}/agents.example.yaml +19 -0
- {agent_dispatch-0.13.0 → agent_dispatch-0.15.0}/pyproject.toml +1 -1
- {agent_dispatch-0.13.0 → agent_dispatch-0.15.0}/src/agent_dispatch/__init__.py +1 -1
- {agent_dispatch-0.13.0 → agent_dispatch-0.15.0}/src/agent_dispatch/cache.py +6 -0
- {agent_dispatch-0.13.0 → agent_dispatch-0.15.0}/src/agent_dispatch/cli.py +170 -1
- {agent_dispatch-0.13.0 → agent_dispatch-0.15.0}/src/agent_dispatch/config.py +1 -1
- {agent_dispatch-0.13.0 → agent_dispatch-0.15.0}/src/agent_dispatch/models.py +24 -0
- {agent_dispatch-0.13.0 → agent_dispatch-0.15.0}/src/agent_dispatch/runner.py +325 -52
- {agent_dispatch-0.13.0 → agent_dispatch-0.15.0}/src/agent_dispatch/server.py +70 -1
- agent_dispatch-0.15.0/src/agent_dispatch/servers.py +152 -0
- agent_dispatch-0.15.0/src/agent_dispatch/usage.py +347 -0
- agent_dispatch-0.15.0/tests/conftest.py +68 -0
- {agent_dispatch-0.13.0 → agent_dispatch-0.15.0}/tests/test_cache.py +15 -0
- {agent_dispatch-0.13.0 → agent_dispatch-0.15.0}/tests/test_cli.py +245 -18
- {agent_dispatch-0.13.0 → agent_dispatch-0.15.0}/tests/test_config.py +17 -0
- {agent_dispatch-0.13.0 → agent_dispatch-0.15.0}/tests/test_models.py +12 -3
- {agent_dispatch-0.13.0 → agent_dispatch-0.15.0}/tests/test_runner.py +390 -4
- {agent_dispatch-0.13.0 → agent_dispatch-0.15.0}/tests/test_server.py +250 -3
- agent_dispatch-0.15.0/tests/test_servers.py +140 -0
- agent_dispatch-0.15.0/tests/test_usage.py +294 -0
- agent_dispatch-0.13.0/tests/conftest.py +0 -41
- {agent_dispatch-0.13.0 → agent_dispatch-0.15.0}/.github/dependabot.yml +0 -0
- {agent_dispatch-0.13.0 → agent_dispatch-0.15.0}/.github/workflows/ci.yml +0 -0
- {agent_dispatch-0.13.0 → agent_dispatch-0.15.0}/.github/workflows/publish.yml +0 -0
- {agent_dispatch-0.13.0 → agent_dispatch-0.15.0}/.gitignore +0 -0
- {agent_dispatch-0.13.0 → agent_dispatch-0.15.0}/LICENSE +0 -0
- {agent_dispatch-0.13.0 → agent_dispatch-0.15.0}/SECURITY.md +0 -0
- {agent_dispatch-0.13.0 → agent_dispatch-0.15.0}/assets/mascot.png +0 -0
- {agent_dispatch-0.13.0 → agent_dispatch-0.15.0}/src/agent_dispatch/jobs.py +0 -0
- {agent_dispatch-0.13.0 → agent_dispatch-0.15.0}/tests/__init__.py +0 -0
- {agent_dispatch-0.13.0 → agent_dispatch-0.15.0}/tests/test_jobs.py +0 -0
|
@@ -12,6 +12,8 @@ MCP server + CLI that lets Claude Code agents delegate tasks to agents in other
|
|
|
12
12
|
|------|------|
|
|
13
13
|
| `src/agent_dispatch/runner.py` | Sync subprocess wrapper around `claude -p` — the actual work |
|
|
14
14
|
| `src/agent_dispatch/server.py` | Async FastMCP interface (21 MCP tools), wraps runner in `asyncio.to_thread` + semaphore |
|
|
15
|
+
| `src/agent_dispatch/usage.py` | Append-only dispatch journal (cost/duration/outcome) + aggregation for `stats` and the `typical` block |
|
|
16
|
+
| `src/agent_dispatch/servers.py` | Registry of live `serve` processes and the version each one runs (`doctor` reads it) |
|
|
15
17
|
| `src/agent_dispatch/cli.py` | Click CLI: `init`, `add`, `update`, `remove`, `list`, `describe`, `test`, `doctor`, `jobs`, `job`, `cancel`, `gc`, `group` (add/list/inspect/update/remove), `serve` |
|
|
16
18
|
| `src/agent_dispatch/models.py` | Pydantic v2 models (`AgentConfig`, `DispatchGroup`/`GroupMember`, `Settings`, `DispatchResult`) |
|
|
17
19
|
| `src/agent_dispatch/config.py` | YAML config load/save + project auto-description |
|
|
@@ -28,7 +30,7 @@ pip install -e ".[dev]"
|
|
|
28
30
|
|
|
29
31
|
```bash
|
|
30
32
|
ruff check src/ tests/
|
|
31
|
-
python3 -m pytest tests/ -v #
|
|
33
|
+
python3 -m pytest tests/ -v # 707 tests, ~17s
|
|
32
34
|
```
|
|
33
35
|
|
|
34
36
|
Tests must **never** invoke the real `claude` CLI. Runner tests mock `shutil.which` + `subprocess.run`/`Popen`; server tests mock `_get_config` + `runner.dispatch`. The one exception is `TestStreamPipeHandling`, which spawns a short-lived *python* subprocess: a pipe deadlock lives in the OS pipe buffer, so a mocked `Popen` structurally cannot reproduce it.
|
|
@@ -60,7 +62,14 @@ Tests must **never** invoke the real `claude` CLI. Runner tests mock `shutil.whi
|
|
|
60
62
|
- Pydantic does **not** validate on assignment. `Field(ge=...)` guards only the *load* path; every mutation surface (CLI `add`/`update`, MCP `add_agent`/`update_agent`) needs its own boundary check, or the bound escapes as a raw `ValidationError`.
|
|
61
63
|
- Every state file (`agents.yaml`, job files) is written **temp file + `os.replace`**, never in place, and every load/mutate/save is wrapped in `config.ProcessLock` — the CLI and the MCP server are separate processes writing the same files, so a thread lock alone loses updates.
|
|
62
64
|
- Anything that changes an agent's config must call `_invalidate_agent_cache` — the cache key holds the agent *name*, not its directory or permissions.
|
|
63
|
-
- Only *clean* successes are cached: `cache.put` refuses failures, `denied_tools` results,
|
|
65
|
+
- Only *clean* successes are cached: `cache.put` refuses failures, `denied_tools` results, `budget_exceeded` results, and results whose `outcome` is `partial`/`blocked`, so the documented "grant access / fix the cause, then re-dispatch" recovery is never short-circuited.
|
|
66
|
+
- **The usage journal is written by a decorator, not at each return.** `runner._journaled` wraps `dispatch`/`dispatch_stream` because each has a dozen early returns; instrumenting them one by one guarantees the next new return path goes unrecorded. It records the outermost call only (`_use_session_flag` is False on the stream's old-CLI retry, which would otherwise double-count one dispatch). It also swallows exceptions itself even though `usage.record` already does: the result in hand is billed, and the guarantee has to hold at the boundary that owns the damage rather than on another module's promise (`test_a_broken_journal_never_breaks_a_paid_dispatch` removes record()'s guard to prove it).
|
|
67
|
+
- **A journal record is one `O_APPEND` write under `PIPE_BUF`**, never a locked read-modify-write: the CLI and every running server share one file, and this user has 14+ servers alive. Every field is capped so the line cannot approach 4096 bytes. Rotation is an atomic rename at ~2 MB, two generations kept.
|
|
68
|
+
- **Profiles are memoized on the journal's (path, size, mtime) and read a NARROWER tail than `stats`.** `list_agents`/`inspect_agent` run on the event-loop thread: a full 2 MB journal cost **24.8 ms** per discovery call at the 512 KB report window — 33x the config parse that 0.13.0 exists to have fixed. Now 3.3 ms cold, ~0 warm. Every profile consumer goes through `usage.profiles()`; never add a caller that re-parses per agent.
|
|
69
|
+
- **`usage_limit` is its own error type** (observed live: `You've hit your session limit · resets 6pm`). It is checked *before* the permission patterns and its hint says the limit is account-wide, so "try another agent" is not a workaround. Text-classified errors attach their advisory through the single `_classification_hint`, not per site.
|
|
70
|
+
- **Server liveness is an advisory lock, not a PID.** `servers.register` holds an exclusive `flock` on `<config dir>/servers/<pid>.json` for the process lifetime; a reader that *takes* that lock owns a dead entry and deletes it. `os.kill(pid, 0)` would believe a recycled PID and would never notice a SIGKILLed server. Never close that fd (re-registering closes the old one first).
|
|
71
|
+
- **The dispatch protocol goes in the system prompt, not the task.** `_build_system_prompt` (runner.py) renders the protocol (non-interactive, time/spend budget, lead with the outcome, trailing `STATUS:` line) plus the agent's `instructions`, and `_build_command` passes it as `--append-system-prompt` — on `--resume` too, since the CLI re-applies an appended prompt on every launch. `_build_prompt` (the `-p` text) is unchanged, so the cache key is unchanged. The text always starts with a `##` header line so it can never be read as a flag. `settings.dispatch_protocol: false` turns the protocol off (for CLIs that predate the flag; `doctor` probes `claude --help` for it); per-agent `instructions` still go through when set.
|
|
72
|
+
- **`outcome` is lifted from the LAST line of the result** (`_split_outcome`), before JSON parsing, by the shared `_build_success_result` (the success-side twin of `_build_error_result` — both `dispatch` and `dispatch_stream` go through it) and by the plain-text fallback tier. A bare marker line is removed from `result`; a marker with a trailing reason stays. `outcome` never flips `success`; `partial`/`blocked` add a `hint` with the `dispatch_session(...)` continuation, after the denial hint. In JSON mode the protocol omits the STATUS instruction (the JSON footer governs), but a STATUS line that arrives anyway is still stripped so `parsed_result` survives.
|
|
64
73
|
- Remediation text is a contract: a hint that names a flag must name one that exists (`test_printed_budget_hint_is_a_runnable_command` feeds the printed flags back into the CLI). Run the command you print.
|
|
65
74
|
- The config error sets are declared **once** and in two halves: `config.CONFIG_LOAD_ERRORS` (read) and `config.CONFIG_SAVE_ERRORS` (write — `yaml.dump`'s `RepresenterError` is a `yaml.YAMLError`, therefore neither `OSError` nor `ConfigLoadError`, and used to escape both the MCP guard and the CLI's `_save_or_exit`). Two halves, not one set, because the remediations differ: a failed write is atomic so the old config survives, while a failed read needs the YAML fixed.
|
|
66
75
|
- MCP tools that load config carry `@_config_guard` under `@mcp.tool()` so a broken `agents.yaml` — or a failed *write* — returns the `{"error": ...}` envelope instead of a raw traceback. The set of load errors lives in one place (`config.CONFIG_LOAD_ERRORS`) because three surfaces handle it: **`UnicodeDecodeError` is a `ValueError`, not an `OSError`**, and listing types per-site is exactly how a cp1251 config slipped past all three.
|
|
@@ -85,4 +94,4 @@ Python ≥ 3.10 · `from __future__ import annotations` everywhere · Pydantic v
|
|
|
85
94
|
|
|
86
95
|
## More detail
|
|
87
96
|
|
|
88
|
-
[README.md](README.md) documents every MCP tool with parameter tables, response shapes, and the error-recovery map — it doubles as the behavioral spec. The test suite (`tests/`,
|
|
97
|
+
[README.md](README.md) documents every MCP tool with parameter tables, response shapes, and the error-recovery map — it doubles as the behavioral spec. The test suite (`tests/`, 707 tests) encodes the exact expected behavior of every layer: when in doubt, read the tests for the module you're touching (`test_runner.py`, `test_server.py`, `test_cli.py`, ...).
|
|
@@ -7,6 +7,109 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
|
|
|
7
7
|
|
|
8
8
|
## [Unreleased]
|
|
9
9
|
|
|
10
|
+
## [0.15.0] - 2026-09-09
|
|
11
|
+
|
|
12
|
+
The measurement round. 0.14.0 told a dispatched agent how to behave; this one
|
|
13
|
+
records what actually happened. Four live dispatches through the new protocol
|
|
14
|
+
came back **4/4 with a STATUS line and 0/4 ending in a clarifying question** —
|
|
15
|
+
including the case it was built for: an ambiguous task ("how many users signed
|
|
16
|
+
up in 30 days?", no project named) where the agent stated its assumption,
|
|
17
|
+
counted across every database, and returned `done` rather than asking. A
|
|
18
|
+
tool-less agent correctly returned `blocked`. No `partial` has been observed in
|
|
19
|
+
the wild yet.
|
|
20
|
+
|
|
21
|
+
### Added
|
|
22
|
+
- **Usage journal + `agent-dispatch stats`.** Every dispatch appends one line to
|
|
23
|
+
`usage.jsonl` — agent, ok, cost, duration, turns, outcome, error type, caller;
|
|
24
|
+
cache hits too, flagged `cached`. `stats [--days N --agent X --json]` turns it
|
|
25
|
+
into spend, median/p90 durations, outcome and failure counts per agent. Until
|
|
26
|
+
now `cost_usd` and `duration_ms` were returned once and discarded, so "which
|
|
27
|
+
agent burns money" had no answer and the only timing knowledge in the system
|
|
28
|
+
was hand-written prose in agent descriptions. Recording is one `O_APPEND`
|
|
29
|
+
write capped under `PIPE_BUF`, so the CLI and all running servers share one
|
|
30
|
+
file with no lock; it rotates at ~2 MB keeping two generations; a failure to
|
|
31
|
+
write it can never fail a dispatch. Off with `settings.usage_log: false`.
|
|
32
|
+
- **`typical` in `list_agents` and `inspect_agent`** — measured median/p90
|
|
33
|
+
seconds and median cost from that agent's own recent runs, so a caller can
|
|
34
|
+
size `timeout_seconds` from data instead of from a description. Absent below
|
|
35
|
+
three recorded dispatches: no data beats a "typical" derived from two runs.
|
|
36
|
+
Costs ~115 bytes per agent (3.6% of a discovery payload).
|
|
37
|
+
- **`error_type: "usage_limit"`** for the Claude *account's* rate/session limit,
|
|
38
|
+
observed live during this round's evaluation (`You've hit your session limit ·
|
|
39
|
+
resets 6pm`) — it used to land in the generic `cli_error` bucket. The hint
|
|
40
|
+
names the reset time and says the limit is account-wide, so dispatching a
|
|
41
|
+
different agent is not a workaround, and nothing was billed.
|
|
42
|
+
- **`doctor` now names the sessions running stale code.** Each server registers
|
|
43
|
+
`<config dir>/servers/<pid>.json` and holds an advisory lock on it for life;
|
|
44
|
+
`doctor` reports live servers by version and prints the pid and age of every
|
|
45
|
+
one still on an older release. A Python process holds its modules for life, so
|
|
46
|
+
an upgrade reaches an open Claude Code session only when it restarts — on this
|
|
47
|
+
machine 18 servers were running previous code with no way to tell which.
|
|
48
|
+
Liveness is the lock, not the PID: a recycled PID cannot fake a live server
|
|
49
|
+
and a SIGKILLed one cannot linger.
|
|
50
|
+
|
|
51
|
+
### Changed
|
|
52
|
+
- **A timeout error now suggests a number, not a doubling.** With enough history
|
|
53
|
+
the message names the value that would have covered this agent's real p90
|
|
54
|
+
(padded 50%). Doubling the current timeout was wrong in both directions.
|
|
55
|
+
|
|
56
|
+
### Fixed
|
|
57
|
+
- Discovery no longer pays for the journal on the event loop: profiles are
|
|
58
|
+
memoized on the journal's (path, size, mtime) and read a narrower tail than
|
|
59
|
+
the report. Measured on a full 2 MB journal, `list_agents` went from
|
|
60
|
+
**24.8 ms to 3.3 ms cold and ~0.01 ms warm** — the 512 KB version would have
|
|
61
|
+
been 33x the config parse that 0.13.0 exists to have eliminated.
|
|
62
|
+
|
|
63
|
+
|
|
64
|
+
## [0.14.0] - 2026-09-08
|
|
65
|
+
|
|
66
|
+
The delegation round: what a dispatched agent is *told*, and what it tells
|
|
67
|
+
back. Until now `claude -p` was launched with the task and nothing else — it
|
|
68
|
+
did not know it was being driven by another agent, that nobody would answer a
|
|
69
|
+
question, how long it had, or how to report an unfinished job. The failure
|
|
70
|
+
mode was concrete and billed: an ambiguous task ended in *"Could you clarify
|
|
71
|
+
which service you mean?"*, `success: true`, cached for the whole TTL.
|
|
72
|
+
|
|
73
|
+
### Added
|
|
74
|
+
- **The dispatch protocol.** Every dispatch now appends a short system prompt
|
|
75
|
+
(`--append-system-prompt`, never mixed into the task text) telling the agent
|
|
76
|
+
that it is non-interactive and dispatched by `caller`, that it must state
|
|
77
|
+
assumptions instead of asking, that a denied tool is something to report and
|
|
78
|
+
work around rather than stop on, what its time budget (and spend cap) is, to
|
|
79
|
+
lead with the outcome — which is what makes a `return_ref` summary, the
|
|
80
|
+
*head* of the text, worth reading — and to end with one line:
|
|
81
|
+
`STATUS: done | partial | blocked`. Passed on `--resume` too. Verified live
|
|
82
|
+
against claude 2.1.263 (`agent-dispatch test <agent> --stream`).
|
|
83
|
+
- **`DispatchResult.outcome`** — that trailing line, lifted out of `result`
|
|
84
|
+
into a field: `done`, `partial` or `blocked`, or absent when the agent did
|
|
85
|
+
not report one. It never flips `success`. `partial`/`blocked` come with a
|
|
86
|
+
`hint` carrying the exact `dispatch_session(..., session_id=...)` call to
|
|
87
|
+
continue, ride the `return_ref` payload and `dispatch_jobs` summaries, are
|
|
88
|
+
shown by `agent-dispatch test` / `job <id>`, are labelled for the
|
|
89
|
+
`dispatch_parallel` aggregator so a blocked member is not synthesized as a
|
|
90
|
+
finished one — and are **not cached**: whatever the agent was missing is not
|
|
91
|
+
in the cache key, so a retry after fixing it must run fresh.
|
|
92
|
+
- **Per-agent `instructions`** — standing orders appended after the protocol
|
|
93
|
+
on every dispatch ("read-only SQL only", "never restart a stack unless the
|
|
94
|
+
task says so"). `add_agent`/`update_agent` (`"none"` clears), CLI `add` /
|
|
95
|
+
`update --instructions` (`none` clears), shown by `inspect_agent` and
|
|
96
|
+
`describe`, pruned from YAML when empty. Changing them invalidates the
|
|
97
|
+
agent's cache entries like any other config change.
|
|
98
|
+
- **`settings.dispatch_protocol`** (default `true`) — set to `false` for a raw
|
|
99
|
+
`claude -p` (a CLI that predates `--append-system-prompt`, or an A/B). The
|
|
100
|
+
protocol and the instructions share one flag, so instructions still go
|
|
101
|
+
through when set. `doctor` now probes `claude --help` for the flag and warns
|
|
102
|
+
with the exact remediation when it is missing.
|
|
103
|
+
|
|
104
|
+
### Changed
|
|
105
|
+
- In `response_format="json"` mode the protocol omits the "lead with the
|
|
106
|
+
outcome" and STATUS bullets — the JSON footer governs the reply shape. A
|
|
107
|
+
STATUS line that arrives anyway is still stripped before parsing, so
|
|
108
|
+
`parsed_result` survives it.
|
|
109
|
+
- `dispatch` and `dispatch_stream` build their success result through one
|
|
110
|
+
`_build_success_result`, the twin of `_build_error_result`.
|
|
111
|
+
|
|
112
|
+
|
|
10
113
|
## [0.13.0] - 2026-08-13
|
|
11
114
|
|
|
12
115
|
An efficiency round, measured against a real 38 KB config (4 agents, 6 groups),
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
Metadata-Version: 2.5
|
|
2
2
|
Name: agent-dispatch
|
|
3
|
-
Version: 0.
|
|
3
|
+
Version: 0.15.0
|
|
4
4
|
Summary: MCP server that lets Claude Code agents delegate tasks to agents in other project directories
|
|
5
5
|
Project-URL: Homepage, https://github.com/ginkida/agent-dispatch
|
|
6
6
|
Project-URL: Repository, https://github.com/ginkida/agent-dispatch
|
|
@@ -134,13 +134,16 @@ Lists all configured agents. **Call this first** to see what's available.
|
|
|
134
134
|
"capabilities": ["docker_logs", "deploy_debug"],
|
|
135
135
|
"risky_capabilities": ["restart_services"],
|
|
136
136
|
"permission_mode": "bypassPermissions",
|
|
137
|
-
"allowed_tools": ["Bash", "Read", "Grep"]
|
|
137
|
+
"allowed_tools": ["Bash", "Read", "Grep"],
|
|
138
|
+
"typical": {"median_seconds": 143.3, "p90_seconds": 178.3, "median_cost_usd": 0.2666}
|
|
138
139
|
}
|
|
139
140
|
]
|
|
140
141
|
```
|
|
141
142
|
|
|
142
143
|
`mcp_servers`, `stacks`, and `dbs` are detected from the agent's project files (`.mcp.json`, `Dockerfile`, `pyproject.toml`, `Cargo.toml`, `prisma/`, `alembic.ini`, etc.) so callers can pick the right agent without dispatching a probe.
|
|
143
144
|
|
|
145
|
+
`typical` is **measured**, not declared: it comes from this agent's own recent dispatches in the [usage journal](#usage-journal--what-your-fleet-actually-costs) and appears only after a few runs. Use it to size `timeout_seconds` and to know what a call will cost before you make it. It is absent when the journal is off or the agent has too little history — no data is better than a typical duration derived from two runs.
|
|
146
|
+
|
|
144
147
|
### `inspect_agent`
|
|
145
148
|
|
|
146
149
|
Cheap detailed lookup — reads the agent's files without spawning a `claude` session. Returns the full config (timeout, model, budget, permission mode, allowed/disallowed tools), detected MCP/stacks/DBs, plus short previews of `CLAUDE.md` and `README.md` when present.
|
|
@@ -152,6 +155,15 @@ Cheap detailed lookup — reads the agent's files without spawning a `claude` se
|
|
|
152
155
|
|
|
153
156
|
Use this **before** `dispatch_async`/`dispatch` to confirm an agent has the tools and context for your task — much cheaper than a probe dispatch.
|
|
154
157
|
|
|
158
|
+
It also carries the agent's `instructions` (its standing orders) and the full `typical` block — the trimmed one in `list_agents` plus `dispatches`, `failed` and the `outcomes` breakdown:
|
|
159
|
+
|
|
160
|
+
```json
|
|
161
|
+
"typical": {
|
|
162
|
+
"dispatches": 24, "median_seconds": 143.3, "p90_seconds": 178.3,
|
|
163
|
+
"median_cost_usd": 0.2666, "failed": 1, "outcomes": {"done": 21, "partial": 2}
|
|
164
|
+
}
|
|
165
|
+
```
|
|
166
|
+
|
|
155
167
|
### Groups
|
|
156
168
|
|
|
157
169
|
A **group** bundles related agents into a cross-project working set — typically a few code repos plus capability gateways (an `infra` agent with a Portainer MCP, an `analytics` agent with a browser + Yandex Metrica). It lets one orchestrating session coordinate work that spans code, deploy, and verification.
|
|
@@ -220,7 +232,8 @@ dispatch(
|
|
|
220
232
|
"session_id": "sess-abc-123",
|
|
221
233
|
"cost_usd": 0.02,
|
|
222
234
|
"duration_ms": 5000,
|
|
223
|
-
"num_turns": 2
|
|
235
|
+
"num_turns": 2,
|
|
236
|
+
"outcome": "done"
|
|
224
237
|
}
|
|
225
238
|
|
|
226
239
|
// Response (failure — error_type helps you handle programmatically)
|
|
@@ -233,7 +246,7 @@ dispatch(
|
|
|
233
246
|
}
|
|
234
247
|
```
|
|
235
248
|
|
|
236
|
-
**`error_type` values:** `permission` (tool/action denied), `timeout`, `recursion` (dispatch depth exceeded), `not_found` (missing directory or CLI), `budget` (the `claude` CLI stopped the session at `max_budget_usd`), `cli_error` (other failures). Permission and
|
|
249
|
+
**`error_type` values:** `permission` (tool/action denied), `timeout`, `recursion` (dispatch depth exceeded), `not_found` (missing directory or CLI), `budget` (the `claude` CLI stopped the session at `max_budget_usd`), `usage_limit` (the Claude *account* hit its rate/session limit), `cli_error` (other failures). Permission, budget, usage-limit and timeout errors include an actionable hint.
|
|
237
250
|
|
|
238
251
|
**Resumable timeouts:** every fresh dispatch pre-assigns a session UUID (`--session-id`), so a timed-out dispatch still returns a `session_id` — the partial transcript survives the kill. The timeout error spells out the recovery: resume with `dispatch_session(agent, "Continue where you left off", session_id=...)`, retry with a bigger `timeout_seconds`, or use `dispatch_async`.
|
|
239
252
|
|
|
@@ -278,6 +291,30 @@ Error: TypeError at scheduler.py:42
|
|
|
278
291
|
Check container logs for recent errors related to the scheduler service
|
|
279
292
|
```
|
|
280
293
|
|
|
294
|
+
**The dispatch protocol and `outcome`.** `claude -p` on its own does not know it is being driven by another agent: given an ambiguous task it will happily end with *"Could you clarify which service you mean?"* — a billed run that answered nothing, and one that a plain cache would then serve again for the whole TTL. So every dispatch also appends a short **dispatch protocol** to the agent's system prompt (`--append-system-prompt`, never mixed into your task text):
|
|
295
|
+
|
|
296
|
+
- it is running non-interactively, dispatched by `caller`, and nobody will answer a question or approve an action — state the assumption and proceed, take the safer reading when two differ;
|
|
297
|
+
- a denied tool or a missing thing is something to report, not something to stop on — finish everything else that is possible;
|
|
298
|
+
- its time budget is about the agent's timeout (and its spend cap, if one is set) — scope the work to fit; a complete partial answer beats an unfinished perfect one;
|
|
299
|
+
- lead with the outcome, then the evidence; list what could not be done and why (this is what makes a `return_ref` summary — the **head** of the text — worth reading);
|
|
300
|
+
- end with one line, `STATUS: done`, `STATUS: partial` or `STATUS: blocked`.
|
|
301
|
+
|
|
302
|
+
That last line is lifted out of `result` into **`outcome`** — the agent's own verdict, deterministic to check: `"done"` means complete; `"partial"` / `"blocked"` mean the agent says the work is unfinished, the result names what is missing, and a `hint` spells out the continuation (usually `dispatch_session(..., session_id=...)`). Read `outcome` before reading the text. It never flips `success`, `partial` / `blocked` results are **not cached**, and a `dispatch_parallel(..., aggregate=...)` labels such members for the aggregator so a blocked report is not synthesized as a finished one. Absent when the agent did not report one — the protocol is off (`settings.dispatch_protocol: false`), `response_format="json"` was requested (the JSON footer governs the reply shape there), or the agent simply skipped it.
|
|
303
|
+
|
|
304
|
+
**Standing orders.** Per-agent `instructions` (set via `add_agent` / `update_agent` or `agent-dispatch update <name> --instructions "..."`) follow the protocol in the same system prompt on **every** dispatch — "read-only SQL only", "never restart a stack unless the task says so", "answer with exact log lines". Unlike `context`, which is per call, and unlike the project's own `CLAUDE.md`, which is written for an interactive session, these are the rules for *being dispatched*.
|
|
305
|
+
|
|
306
|
+
```json
|
|
307
|
+
// Response (success, but the agent says it did not finish)
|
|
308
|
+
{
|
|
309
|
+
"agent": "infra",
|
|
310
|
+
"success": true,
|
|
311
|
+
"result": "Restarted horizon. Could not verify the queue drained: the redis container is not reachable from here.",
|
|
312
|
+
"session_id": "sess-abc-123",
|
|
313
|
+
"outcome": "partial",
|
|
314
|
+
"hint": "The agent reports its work is PARTIAL — read the result for what is missing, then continue in the same session via dispatch_session(agent='infra', task='Continue where you left off', session_id='sess-abc-123') or re-dispatch with what it needed."
|
|
315
|
+
}
|
|
316
|
+
```
|
|
317
|
+
|
|
281
318
|
### `dispatch_session`
|
|
282
319
|
|
|
283
320
|
Multi-turn: continue a conversation with an agent. First call starts a session, pass `session_id` back to continue. Never cached.
|
|
@@ -395,6 +432,7 @@ Register a new project directory as an agent. Description is auto-generated from
|
|
|
395
432
|
| `disallowed_tools` | string | no | Comma-separated disallowed tools |
|
|
396
433
|
| `capabilities` | string | no | Comma-separated capability labels (e.g. `"docker_logs,deploy_debug"`) |
|
|
397
434
|
| `risky_capabilities` | string | no | Comma-separated high-risk labels (e.g. `"restart_services"`) |
|
|
435
|
+
| `instructions` | string | no | Standing orders appended to the agent's system prompt on every dispatch (see [the dispatch protocol](#dispatch)) |
|
|
398
436
|
|
|
399
437
|
### `update_agent`
|
|
400
438
|
|
|
@@ -412,6 +450,7 @@ Update an existing agent's configuration. Only non-empty fields are changed. Pas
|
|
|
412
450
|
| `disallowed_tools` | string | no | Comma-separated. `"none"` to clear |
|
|
413
451
|
| `capabilities` | string | no | Comma-separated. `"none"` to clear |
|
|
414
452
|
| `risky_capabilities` | string | no | Comma-separated. `"none"` to clear |
|
|
453
|
+
| `instructions` | string | no | Standing orders for every dispatch (replaces the text). `"none"` to clear |
|
|
415
454
|
|
|
416
455
|
Changing an agent's config drops that agent's cached results — the cache key holds the agent *name*, so a re-pointed or re-permissioned agent would otherwise keep answering from the previous config for the rest of the TTL. The same applies to `add_agent` and `remove_agent`.
|
|
417
456
|
|
|
@@ -506,15 +545,17 @@ Failures are deterministic: check `success`, then branch on `error_type`.
|
|
|
506
545
|
| `error_type` | Meaning | Recovery |
|
|
507
546
|
|--------------|---------|----------|
|
|
508
547
|
| `permission` | A tool call was denied | `update_agent(name, allowed_tools="Bash,Read")` (least privilege) or `update_agent(name, permission_mode="bypassPermissions")`, then re-dispatch. The `error` text includes a hint with the exact fix. |
|
|
509
|
-
| `timeout` | Process killed at the timeout | Resume the partial work: `dispatch_session(agent, "Continue where you left off", session_id=<from the error text>)`. Or retry with a bigger `timeout_seconds
|
|
548
|
+
| `timeout` | Process killed at the timeout | Resume the partial work: `dispatch_session(agent, "Continue where you left off", session_id=<from the error text>)`. Or retry with a bigger `timeout_seconds=` — once the journal has history the error names the value that would have covered this agent's p90 — or use `dispatch_async`. A *streaming* dispatch that produced its answer before the deadline returns that answer with a `hint` instead of failing. |
|
|
510
549
|
| `not_found` | Agent directory or `claude` CLI missing | `list_agents()` → check `healthy`. Re-add the agent with an existing path, or run `agent-dispatch doctor` to find what's missing. |
|
|
511
550
|
| `recursion` | Dispatch nesting exceeded `max_dispatch_depth` (default 3) | Don't dispatch from dispatched agents; if the nesting is intentional, raise `max_dispatch_depth` in settings. |
|
|
512
551
|
| `budget` | The `claude` CLI ended the session at the `max_budget_usd` spend cap — the answer is incomplete | Raise the cap (`update_agent(name, max_budget_usd=2.0)`), switch to a cheaper `model`, or split the task. The partial session is resumable: `dispatch_session(agent, "Continue where you left off", session_id=<from the result>)`. |
|
|
552
|
+
| `usage_limit` | The Claude **account** hit its usage/rate limit — not this agent, and nothing was billed | Wait for the reset named in the `error` text. Dispatching a *different* agent is not a workaround: every agent runs on the same account. If it recurs under load, lower `settings.max_concurrency`. |
|
|
513
553
|
| `cli_error` | Anything else from the `claude` subprocess | Read the `error` text; run `agent-dispatch doctor` for environment issues; retry once if transient. |
|
|
514
554
|
|
|
515
555
|
Three soft signals that arrive with `success: true`:
|
|
516
556
|
|
|
517
557
|
- **`denied_tools` + `hint`** — the agent finished but some tool calls were blocked; the result may be incomplete. Grant access (see the `permission` row) and re-dispatch.
|
|
558
|
+
- **`outcome: "partial"` / `"blocked"`** — the agent's own verdict that it did not finish; the result names what is missing and the `hint` carries the `dispatch_session(...)` call to continue in the same session. Not cached, so a re-dispatch after fixing the cause runs fresh.
|
|
518
559
|
- **`parsed_result: null` with `response_format="json"`** — the reply wasn't valid JSON; the raw text is still in `result`. Caveat: an agent that *can't* comply returns `{"error": "<reason>"}` — which parses successfully — so also check `parsed_result` for an `"error"` key.
|
|
519
560
|
- **`budget_exceeded: true`** — `cost_usd` came in over the agent's `max_budget_usd` (or the settings default) without the CLI stopping the run (the final turn can overshoot the cap). The dispatch is not failed — the money is already spent — but a runaway agent is now visible. Tighten the task, pick a cheaper model, or raise the budget. A run the CLI *did* stop fails with `error_type: "budget"` instead.
|
|
520
561
|
|
|
@@ -539,6 +580,9 @@ agents:
|
|
|
539
580
|
- deploy_debug
|
|
540
581
|
risky_capabilities: # high-risk labels, surfaced for visibility
|
|
541
582
|
- restart_services
|
|
583
|
+
# instructions: | # standing orders, appended to the system prompt on every dispatch
|
|
584
|
+
# Read-only: never restart or redeploy unless the task says so.
|
|
585
|
+
# Quote exact log lines with timestamps.
|
|
542
586
|
# model: sonnet # optional model override
|
|
543
587
|
# max_budget_usd: 1.0 # cost limit per dispatch
|
|
544
588
|
# permission_mode: bypassPermissions # one of: default | plan | bypassPermissions
|
|
@@ -574,6 +618,11 @@ settings:
|
|
|
574
618
|
# - Edit
|
|
575
619
|
max_dispatch_depth: 3 # recursion protection
|
|
576
620
|
max_concurrency: 5 # max parallel claude -p processes (per dispatch path)
|
|
621
|
+
# dispatch_protocol: true # send every agent the dispatch protocol (non-interactive,
|
|
622
|
+
# # time budget, STATUS line → `outcome`). false = raw claude -p.
|
|
623
|
+
# usage_log: true # record every dispatch in usage.jsonl (cost, duration, outcome).
|
|
624
|
+
# # Powers `agent-dispatch stats`, the `typical` block and the
|
|
625
|
+
# # measured timeout suggestion. false = record nothing.
|
|
577
626
|
# job_retention_days: 30 # 0 (default) = never prune. See "Job retention" below.
|
|
578
627
|
cache:
|
|
579
628
|
enabled: true
|
|
@@ -583,6 +632,42 @@ settings:
|
|
|
583
632
|
|
|
584
633
|
Config is reloaded on every tool call — add agents without restarting.
|
|
585
634
|
|
|
635
|
+
### Usage journal — what your fleet actually costs
|
|
636
|
+
|
|
637
|
+
Every dispatch appends one line to `~/.config/agent-dispatch/usage.jsonl`
|
|
638
|
+
(override with `AGENT_DISPATCH_USAGE_LOG`): agent, success, cost, duration,
|
|
639
|
+
turns, `outcome`, error type, caller. Cache hits are recorded too, flagged
|
|
640
|
+
`cached`, so the report can show what the cache saves.
|
|
641
|
+
|
|
642
|
+
```console
|
|
643
|
+
$ agent-dispatch stats --days 7
|
|
644
|
+
Usage (last 7 day(s))
|
|
645
|
+
dispatches: 33, 1 served from cache
|
|
646
|
+
spend: $13.0390
|
|
647
|
+
duration: median 87s, p90 2m40s, max 10m00s
|
|
648
|
+
outcomes: blocked 1, done 27, partial 3
|
|
649
|
+
failures: timeout 1, usage_limit 1
|
|
650
|
+
|
|
651
|
+
Per agent
|
|
652
|
+
analytic 18 runs $ 8.9441 median 2m04s p90 2m51s
|
|
653
|
+
done 15, partial 3
|
|
654
|
+
gitlab 12 runs $ 3.7949 median 60s p90 86s
|
|
655
|
+
done 12 | 1 cached
|
|
656
|
+
```
|
|
657
|
+
|
|
658
|
+
`--agent NAME` narrows it, `--json` emits the same report as JSON.
|
|
659
|
+
|
|
660
|
+
The journal is what makes three other things work: the `typical` block in
|
|
661
|
+
`list_agents` / `inspect_agent`, the timeout error naming a value derived from
|
|
662
|
+
the agent's real p90 instead of a doubling guess, and any answer at all to
|
|
663
|
+
"which agent is expensive". Turn it off with `usage_log: false` in settings.
|
|
664
|
+
|
|
665
|
+
Properties worth knowing: the file is owner-only (`0o600`); each record is a
|
|
666
|
+
single `O_APPEND` write capped well under `PIPE_BUF`, so the CLI and every
|
|
667
|
+
running server can share it with no lock and no torn lines; it rotates to
|
|
668
|
+
`usage.jsonl.1` at ~2 MB and keeps two generations, so it is bounded at ~4 MB
|
|
669
|
+
forever. A failure to write it can never fail a dispatch.
|
|
670
|
+
|
|
586
671
|
### Job retention
|
|
587
672
|
|
|
588
673
|
Every `dispatch_async` **and** every `dispatch(..., return_ref=True)` writes a
|
|
@@ -646,8 +731,11 @@ agent-dispatch MCP server
|
|
|
646
731
|
▼
|
|
647
732
|
New Claude Code session in ~/projects/infra/
|
|
648
733
|
├─ Inherits: CLAUDE.md, .mcp.json, project tools
|
|
734
|
+
├─ System prompt += dispatch protocol + the agent's standing `instructions`
|
|
649
735
|
├─ Receives structured prompt with goal/caller/context/task
|
|
650
|
-
└─ Returns result →
|
|
736
|
+
└─ Returns result (+ its own STATUS → `outcome`) → cached when complete
|
|
737
|
+
│
|
|
738
|
+
└─ one line appended to usage.jsonl (cost, duration, outcome)
|
|
651
739
|
```
|
|
652
740
|
|
|
653
741
|
## Safety
|
|
@@ -655,11 +743,11 @@ agent-dispatch MCP server
|
|
|
655
743
|
- **Recursion protection** — `AGENT_DISPATCH_DEPTH` env var tracks nesting. Default limit: 3. Best-effort across the subprocess boundary (see [SECURITY.md](SECURITY.md)).
|
|
656
744
|
- **Argument-injection guard** — structured CLI fields (`session_id`, `model`, `permission_mode`, tool names) that start with `-` are rejected so they can't smuggle extra `claude` flags.
|
|
657
745
|
- **Path-traversal guard** — caller-supplied `job_id`/`ref` values are validated as 32-char hex before any filesystem access.
|
|
658
|
-
- **Owner-only state** — job files
|
|
746
|
+
- **Owner-only state** — job files, `agents.yaml`, the usage journal and the server registry are all written `0o600`; their directories are `0o700`.
|
|
659
747
|
- **Cost control** — `max_budget_usd` per agent or globally is passed to the `claude` CLI as `--max-budget-usd`, so a runaway dispatch is stopped at the cap and comes back as `error_type: "budget"` with a resumable `session_id`. An overshoot that lands over budget without stopping is flagged post-hoc with `budget_exceeded: true` + a hint.
|
|
660
748
|
- **Concurrency** — `max_concurrency` (default: 5) caps parallel `claude -p` processes. Note: the sync and async dispatch paths use separate semaphores, so the worst-case total is `2 × max_concurrency`.
|
|
661
749
|
- **Timeout** — per-agent or global (default: 300s). A streaming dispatch runs the agent in its own process group, so the deadline kills the whole tree: a process the agent left running in the background can't hold the dispatch (and its concurrency slot) open past the timeout.
|
|
662
|
-
- **Caching** — identical `(agent, task, context, caller, goal, response_format)` requests return cached results, bounded by `cache.max_size` (oldest entry evicted first). Only clean successes are cached: failures, results with `denied_tools`,
|
|
750
|
+
- **Caching** — identical `(agent, task, context, caller, goal, response_format)` requests return cached results, bounded by `cache.max_size` (oldest entry evicted first). Only clean successes are cached: failures, results with `denied_tools`, results flagged `budget_exceeded`, and results the agent itself reported as `partial` / `blocked` are not, so the documented "grant access / raise the cap, then re-dispatch" recovery is never served a stale crippled answer. Changing an agent's config invalidates its entries. Sessions and dialogues are never cached. A `group=` dispatch folds the group's `shared_context` into `context`, so different groups cache separately and a plain dispatch is unaffected.
|
|
663
751
|
- **Durable config** — `agents.yaml` is written atomically (temp file + rename), so an interrupted write can never truncate it. Every mutation path (CLI and MCP server alike) also takes a cross-process advisory lock, so concurrent edits don't drop one another's agents. The lock is best-effort by design: after waiting 10 seconds it logs a warning and proceeds anyway, because a wedged lock holder must not freeze the MCP server — so on a heavily contended config a lost update is possible, while a truncated one is not.
|
|
664
752
|
|
|
665
753
|
See [SECURITY.md](SECURITY.md) for the full threat model (including the `bypassPermissions` escalation risk and on-disk job files).
|
|
@@ -670,13 +758,14 @@ See [SECURITY.md](SECURITY.md) for the full threat model (including the `bypassP
|
|
|
670
758
|
|---------|-------------|
|
|
671
759
|
| `agent-dispatch init` | Create config + register MCP server with Claude Code |
|
|
672
760
|
| `agent-dispatch add <name> <dir>` | Add an agent (auto-generates description) |
|
|
673
|
-
| `agent-dispatch update <name>` | Update agent config (permissions, timeout, model, etc.) |
|
|
761
|
+
| `agent-dispatch update <name>` | Update agent config (permissions, timeout, model, `--instructions`, etc.) |
|
|
674
762
|
| `agent-dispatch remove <name>` | Remove an agent |
|
|
675
763
|
| `agent-dispatch list` | List agents with health status and permissions |
|
|
676
764
|
| `agent-dispatch group <add\|list\|inspect\|update\|remove>` | Manage [groups](#groups) — cross-project working sets of agents |
|
|
677
765
|
| `agent-dispatch describe <name>` | Show full configuration for one agent (tri-state tools, project files) |
|
|
678
766
|
| `agent-dispatch test <name> [task] [--stream]` | Test an agent with a dispatch (`--stream` for live progress) |
|
|
679
|
-
| `agent-dispatch
|
|
767
|
+
| `agent-dispatch stats [--days N --agent X --json]` | What dispatches cost: spend, durations, outcomes and failures per agent |
|
|
768
|
+
| `agent-dispatch doctor` | Diagnose installation: Claude CLI (incl. `--append-system-prompt` support), running servers on stale code, MCP registration, agent health, and group membership |
|
|
680
769
|
| `agent-dispatch jobs [--status --limit]` | List async dispatch jobs (most recent first) |
|
|
681
770
|
| `agent-dispatch job <id>` | Show one job: status, progress tail, result preview |
|
|
682
771
|
| `agent-dispatch cancel <id>` | Cancel a pending job (running jobs: use the `dispatch_cancel` MCP tool) |
|