agent-dispatch 0.9.0__tar.gz → 0.11.0__tar.gz

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (30) hide show
  1. {agent_dispatch-0.9.0 → agent_dispatch-0.11.0}/AGENTS.md +12 -6
  2. {agent_dispatch-0.9.0 → agent_dispatch-0.11.0}/CHANGELOG.md +70 -1
  3. {agent_dispatch-0.9.0 → agent_dispatch-0.11.0}/PKG-INFO +27 -6
  4. {agent_dispatch-0.9.0 → agent_dispatch-0.11.0}/README.md +26 -5
  5. {agent_dispatch-0.9.0 → agent_dispatch-0.11.0}/agents.example.yaml +28 -2
  6. {agent_dispatch-0.9.0 → agent_dispatch-0.11.0}/pyproject.toml +1 -1
  7. {agent_dispatch-0.9.0 → agent_dispatch-0.11.0}/src/agent_dispatch/__init__.py +1 -1
  8. {agent_dispatch-0.9.0 → agent_dispatch-0.11.0}/src/agent_dispatch/cli.py +38 -4
  9. {agent_dispatch-0.9.0 → agent_dispatch-0.11.0}/src/agent_dispatch/config.py +34 -9
  10. {agent_dispatch-0.9.0 → agent_dispatch-0.11.0}/src/agent_dispatch/jobs.py +23 -10
  11. {agent_dispatch-0.9.0 → agent_dispatch-0.11.0}/src/agent_dispatch/models.py +5 -1
  12. {agent_dispatch-0.9.0 → agent_dispatch-0.11.0}/src/agent_dispatch/runner.py +147 -58
  13. {agent_dispatch-0.9.0 → agent_dispatch-0.11.0}/src/agent_dispatch/server.py +4 -2
  14. {agent_dispatch-0.9.0 → agent_dispatch-0.11.0}/tests/test_cli.py +69 -0
  15. {agent_dispatch-0.9.0 → agent_dispatch-0.11.0}/tests/test_config.py +29 -1
  16. {agent_dispatch-0.9.0 → agent_dispatch-0.11.0}/tests/test_jobs.py +72 -9
  17. {agent_dispatch-0.9.0 → agent_dispatch-0.11.0}/tests/test_runner.py +500 -151
  18. {agent_dispatch-0.9.0 → agent_dispatch-0.11.0}/.github/dependabot.yml +0 -0
  19. {agent_dispatch-0.9.0 → agent_dispatch-0.11.0}/.github/workflows/ci.yml +0 -0
  20. {agent_dispatch-0.9.0 → agent_dispatch-0.11.0}/.github/workflows/publish.yml +0 -0
  21. {agent_dispatch-0.9.0 → agent_dispatch-0.11.0}/.gitignore +0 -0
  22. {agent_dispatch-0.9.0 → agent_dispatch-0.11.0}/LICENSE +0 -0
  23. {agent_dispatch-0.9.0 → agent_dispatch-0.11.0}/SECURITY.md +0 -0
  24. {agent_dispatch-0.9.0 → agent_dispatch-0.11.0}/assets/mascot.png +0 -0
  25. {agent_dispatch-0.9.0 → agent_dispatch-0.11.0}/src/agent_dispatch/cache.py +0 -0
  26. {agent_dispatch-0.9.0 → agent_dispatch-0.11.0}/tests/__init__.py +0 -0
  27. {agent_dispatch-0.9.0 → agent_dispatch-0.11.0}/tests/conftest.py +0 -0
  28. {agent_dispatch-0.9.0 → agent_dispatch-0.11.0}/tests/test_cache.py +0 -0
  29. {agent_dispatch-0.9.0 → agent_dispatch-0.11.0}/tests/test_models.py +0 -0
  30. {agent_dispatch-0.9.0 → agent_dispatch-0.11.0}/tests/test_server.py +0 -0
@@ -28,7 +28,7 @@ pip install -e ".[dev]"
28
28
 
29
29
  ```bash
30
30
  ruff check src/ tests/
31
- python3 -m pytest tests/ -v # 466 tests, ~2s — all subprocess calls are mocked
31
+ python3 -m pytest tests/ -v # 495 tests, ~2s — all subprocess calls are mocked
32
32
  ```
33
33
 
34
34
  Tests must **never** invoke the real `claude` CLI. Runner tests mock `shutil.which` + `subprocess.run`/`Popen`; server tests mock `_get_config` + `runner.dispatch`.
@@ -36,14 +36,20 @@ Tests must **never** invoke the real `claude` CLI. Runner tests mock `shutil.whi
36
36
  ## Non-obvious invariants (violating these breaks real behavior)
37
37
 
38
38
  - `allowed_tools` / `disallowed_tools` are **tri-state**: `None` = inherit settings defaults, `[]` = explicitly no tools, `[...]` = exactly these. Check with `is not None`, never `or` — `[]` is falsy but semantically distinct.
39
- - `denied_tools` non-empty + `is_error` ⇒ `error_type="permission"`, regardless of what the error text matches.
39
+ - Error-type precedence on an `is_error` payload (`_build_error_result`): the CLI's own budget stop wins, then `denied_tools` non-empty ⇒ `error_type="permission"` regardless of the error text, then text classification.
40
40
  - **Groups**: a group's `shared_context` is folded into the `context` *string* before the cache/runner calls (`_merge_group_context` in server.py) — runner.py and cache.py are untouched, the cache key disambiguates groups for free, and `group=""` is byte-identical to a plain dispatch. Membership is validated up front (`_validate_group_member`, separate from the pure merge so `dispatch_parallel`'s all-or-nothing pre-check holds). `DispatchConfig` validates only group *keys*, never member existence — a hard cross-ref check would brick config load when a shared gateway agent is removed; dangling refs are flagged (`unknown:true`) at read time instead.
41
41
  - On failure, callers read `DispatchResult.error` + `error_type` — `result` holds the raw agent output even on errors.
42
42
  - `--session-id` and `--resume` conflict — never pass both to `claude`.
43
43
  - Valid permission modes: `default`, `plan`, `bypassPermissions` (`models.py: KNOWN_PERMISSION_MODES`).
44
- - `JobStore.finish`/`fail` refuse already-terminal jobs (returns `None`) — this closes the race with force-cancel; never "fix" it by overwriting.
44
+ - `JobStore.finish`/`fail` refuse already-terminal jobs (returns `None`) — this closes the race with force-cancel; never "fix" it by overwriting. `mark_running` likewise refuses any job that isn't `pending`, so a stale or duplicate worker can't resurrect a finished one.
45
+ - "Is this group member missing?" has exactly one implementation: `DispatchConfig.unknown_group_members()`. Any new surface that lists or validates membership calls it instead of re-deriving the check.
45
46
  - Cancelling a *running* job requires the in-memory `_running_procs` registry (server.py) — the job is marked `cancelled` **before** the subprocess is killed. Don't persist PIDs to disk (PID reuse after restart could kill an unrelated process).
46
- - `max_budget_usd` is **post-hoc**: `_apply_budget` (runner.py) sets `budget_exceeded` + `hint` after the cost is known; it never fails the dispatch.
47
+ - `max_budget_usd` is enforced **by the claude CLI** (`_build_command` passes `--max-budget-usd`): a run stopped at the cap comes back `is_error` with no `result` text, and `_build_error_result` turns it into `error_type="budget"` + `budget_exceeded=True` + a resumable `session_id`. `_apply_budget` is the *secondary*, post-hoc signal for an overshoot that didn't stop the run; it never fails a dispatch.
48
+ - A CLI error payload can have no `result` field at all — the reason lives in `errors` / `subtype`. Read it via `_cli_error_details`, never assume `result` is populated on failure.
49
+
50
+ ## Deliberately not built
51
+
52
+ These were considered — some fully implemented — and cut on purpose: an agent router / auto-dispatch (`recommend_agent` / `dispatch_auto`, removed before 0.8.0 — a keyword scorer adds little over the calling LLM at a handful of agents, and auto-dispatch can spend money or mutate a repo on a guess); groups as an execution engine (they are a descriptive layer — no routing, no per-group settings); an agent-dispatch-side budget ledger across dispatches (the CLI's own `--max-budget-usd` covers a single run; anything cumulative would need state we deliberately don't keep). Please open an issue with the use case before adding any of them.
47
53
 
48
54
  ## Conventions
49
55
 
@@ -51,8 +57,8 @@ Python ≥ 3.10 · `from __future__ import annotations` everywhere · Pydantic v
51
57
 
52
58
  ## When adding a feature, check every layer
53
59
 
54
- `models.py` (data shape) → `runner.py` (dispatch mechanics) → `server.py` (MCP tool) → `cli.py` (CLI flag) → tests for each → `README.md` + `agents.example.yaml` (user docs).
60
+ `models.py` (data shape) → `config.py` (YAML round-trip + empty-collection pruning) → `runner.py` (dispatch mechanics) → `server.py` (MCP tool) → `cli.py` (CLI flag) → tests for each → `README.md` + `agents.example.yaml` (user docs).
55
61
 
56
62
  ## More detail
57
63
 
58
- [README.md](README.md) documents every MCP tool with parameter tables, response shapes, and the error-recovery map — it doubles as the behavioral spec. The test suite (`tests/`, 466 tests) encodes the exact expected behavior of every layer: when in doubt, read the tests for the module you're touching (`test_runner.py`, `test_server.py`, `test_cli.py`, ...).
64
+ [README.md](README.md) documents every MCP tool with parameter tables, response shapes, and the error-recovery map — it doubles as the behavioral spec. The test suite (`tests/`, 495 tests) encodes the exact expected behavior of every layer: when in doubt, read the tests for the module you're touching (`test_runner.py`, `test_server.py`, `test_cli.py`, ...).
@@ -7,6 +7,71 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
7
7
 
8
8
  ## [Unreleased]
9
9
 
10
+ ## [0.11.0] - 2026-07-27
11
+
12
+ The spend cap was already real — now the result says so.
13
+
14
+ ### Added
15
+ - **`error_type: "budget"`.** The `claude` CLI enforces the `--max-budget-usd`
16
+ value agent-dispatch has always passed: when the cap is reached it ends the
17
+ session and reports the failure with no `result` text. That payload used to
18
+ surface as the useless `"reported an error with no details"` fallback with
19
+ `error_type: "cli_error"`. It is now classified as `budget`, carries the
20
+ CLI's own reason (`Reached maximum budget ($X)`), sets
21
+ `budget_exceeded: true`, and spells out the recovery — raise the cap, pick a
22
+ cheaper model, or resume the partial session via the returned `session_id`.
23
+ `agent-dispatch test` prints a matching diagnosis.
24
+
25
+ ### Fixed
26
+ - **Error details are no longer dropped.** CLI-level failures (budget
27
+ exhausted, max turns, execution errors) put their reason in `errors` /
28
+ `subtype` rather than `result`. Both `dispatch` and `dispatch_stream` now
29
+ read those fields when `result` is empty, so the error text names the actual
30
+ problem. Values are capped (5 entries × 300 chars) like every other field
31
+ read back from the subprocess.
32
+ - **Job files can no longer be corrupted by a concurrent writer.** `JobStore`
33
+ wrote through a fixed `<id>.tmp` path; the in-process lock does not cover
34
+ the CLI (`cancel` / `gc`) and the MCP server writing the same job at the
35
+ same time, and the rename could publish interleaved JSON. Each write now
36
+ uses a unique temp name and cleans it up if the write fails.
37
+
38
+ ### Changed
39
+ - The `is_error` handling in `dispatch` and `dispatch_stream` is now one
40
+ shared `_build_error_result()` with an explicit precedence: budget stop >
41
+ denied tools > text classification.
42
+ - Documentation corrected throughout: `max_budget_usd` is enforced per
43
+ dispatch by the CLI, with the `budget_exceeded` flag as the secondary
44
+ post-hoc signal for a run that overshot without being stopped. Previous
45
+ releases described it as post-hoc only.
46
+
47
+ ## [0.10.0] - 2026-07-14
48
+
49
+ `doctor` learns to check groups; three correctness fixes found in review.
50
+
51
+ ### Added
52
+ - **`agent-dispatch doctor` diagnoses group health.** A new "Groups" section
53
+ reports each group's member count, flags dangling members (agents removed
54
+ from config but still referenced) as a failure, and empty groups as a
55
+ warning — each with a concrete remediation command.
56
+
57
+ ### Fixed
58
+ - `auto_describe()`: an empty `"description"` field in `package.json` (a
59
+ common `npm init` placeholder) was no longer being filtered out, producing
60
+ malformed generated descriptions like `" | Stack: Node.js"`. Empty/blank
61
+ descriptions are ignored again.
62
+ - `doctor`'s group remediation hints pointed at `group update --members`,
63
+ a flag that doesn't exist (`group update` only edits `description` /
64
+ `shared_context`; membership is set via `group add --member`). Hints now
65
+ point at the working `group remove` + `group add --member` recreation path.
66
+ - `JobStore.mark_running()` now refuses to start any job that isn't
67
+ `pending` (previously only refused `cancelled`), closing a race where a
68
+ stale/duplicate worker could resurrect an already-finished or failed job.
69
+
70
+ ### Changed
71
+ - Consolidated the "unknown group member" check — previously duplicated
72
+ across five call sites in `cli.py` and `server.py` — into a single
73
+ `DispatchConfig.unknown_group_members()` helper.
74
+
10
75
  ## [0.9.0] - 2026-06-30
11
76
 
12
77
  Coordinate a group of related projects from one session.
@@ -361,7 +426,11 @@ cache bounding, and stale-job recovery.
361
426
  - Dependabot for `pip` + `github-actions`, GitHub Actions pinned to
362
427
  commit SHAs for supply-chain integrity.
363
428
 
364
- [Unreleased]: https://github.com/ginkida/agent-dispatch/compare/v0.6.0...HEAD
429
+ [Unreleased]: https://github.com/ginkida/agent-dispatch/compare/v0.11.0...HEAD
430
+ [0.11.0]: https://github.com/ginkida/agent-dispatch/compare/v0.10.0...v0.11.0
431
+ [0.10.0]: https://github.com/ginkida/agent-dispatch/compare/v0.9.0...v0.10.0
432
+ [0.9.0]: https://github.com/ginkida/agent-dispatch/compare/v0.8.0...v0.9.0
433
+ [0.8.0]: https://github.com/ginkida/agent-dispatch/compare/v0.6.0...v0.8.0
365
434
  [0.6.0]: https://github.com/ginkida/agent-dispatch/compare/v0.5.0...v0.6.0
366
435
  [0.5.0]: https://github.com/ginkida/agent-dispatch/compare/v0.4.0...v0.5.0
367
436
  [0.4.0]: https://github.com/ginkida/agent-dispatch/compare/v0.3.0...v0.4.0
@@ -1,6 +1,6 @@
1
1
  Metadata-Version: 2.4
2
2
  Name: agent-dispatch
3
- Version: 0.9.0
3
+ Version: 0.11.0
4
4
  Summary: MCP server that lets Claude Code agents delegate tasks to agents in other project directories
5
5
  Project-URL: Homepage, https://github.com/ginkida/agent-dispatch
6
6
  Project-URL: Repository, https://github.com/ginkida/agent-dispatch
@@ -43,6 +43,8 @@ Description-Content-Type: text/markdown
43
43
 
44
44
  Each agent runs as a separate `claude -p` session in its own project directory — inheriting that project's MCP servers, CLAUDE.md, and tools. The calling agent just gets the result back.
45
45
 
46
+ Related projects can be bundled into a **[group](#groups)** — a shared brief plus a member list — so one session can coordinate work across them (e.g. code repos + an `infra`/Portainer gateway + an `analytics` gateway).
47
+
46
48
  Works with OAuth, API key, and Claude subscription authentication.
47
49
 
48
50
  > **AI agents:** this README is the canonical doc for *using* the tool — setup: [Quick Start](#quick-start) (every step has a deterministic verify), first call: [`dispatch`](#dispatch), tool selection: [Which Tool to Use](#which-tool-to-use), failure handling: [Error Recovery](#error-recovery). Working *on* this repo instead? See [AGENTS.md](AGENTS.md).
@@ -231,7 +233,7 @@ dispatch(
231
233
  }
232
234
  ```
233
235
 
234
- **`error_type` values:** `permission` (tool/action denied), `timeout`, `recursion` (dispatch depth exceeded), `not_found` (missing directory or CLI), `cli_error` (other failures). Permission errors include an actionable hint.
236
+ **`error_type` values:** `permission` (tool/action denied), `timeout`, `recursion` (dispatch depth exceeded), `not_found` (missing directory or CLI), `budget` (the `claude` CLI stopped the session at `max_budget_usd`), `cli_error` (other failures). Permission and budget errors include an actionable hint.
235
237
 
236
238
  **Resumable timeouts:** every fresh dispatch pre-assigns a session UUID (`--session-id`), so a timed-out dispatch still returns a `session_id` — the partial transcript survives the kill. The timeout error spells out the recovery: resume with `dispatch_session(agent, "Continue where you left off", session_id=...)`, retry with a bigger `timeout_seconds`, or use `dispatch_async`.
237
239
 
@@ -501,13 +503,14 @@ Failures are deterministic: check `success`, then branch on `error_type`.
501
503
  | `timeout` | Process killed at the timeout | Resume the partial work: `dispatch_session(agent, "Continue where you left off", session_id=<from the error text>)`. Or retry with a bigger `timeout_seconds=`, or use `dispatch_async`. |
502
504
  | `not_found` | Agent directory or `claude` CLI missing | `list_agents()` → check `healthy`. Re-add the agent with an existing path, or run `agent-dispatch doctor` to find what's missing. |
503
505
  | `recursion` | Dispatch nesting exceeded `max_dispatch_depth` (default 3) | Don't dispatch from dispatched agents; if the nesting is intentional, raise `max_dispatch_depth` in settings. |
506
+ | `budget` | The `claude` CLI ended the session at the `max_budget_usd` spend cap — the answer is incomplete | Raise the cap (`update_agent(name, max_budget_usd=2.0)`), switch to a cheaper `model`, or split the task. The partial session is resumable: `dispatch_session(agent, "Continue where you left off", session_id=<from the result>)`. |
504
507
  | `cli_error` | Anything else from the `claude` subprocess | Read the `error` text; run `agent-dispatch doctor` for environment issues; retry once if transient. |
505
508
 
506
509
  Three soft signals that arrive with `success: true`:
507
510
 
508
511
  - **`denied_tools` + `hint`** — the agent finished but some tool calls were blocked; the result may be incomplete. Grant access (see the `permission` row) and re-dispatch.
509
512
  - **`parsed_result: null` with `response_format="json"`** — the reply wasn't valid JSON; the raw text is still in `result`. Caveat: an agent that *can't* comply returns `{"error": "<reason>"}` — which parses successfully — so also check `parsed_result` for an `"error"` key.
510
- - **`budget_exceeded: true`** — `cost_usd` exceeded the agent's `max_budget_usd` (or the settings default). The dispatch is not failed — the money is already spent — but a runaway agent is now visible. Tighten the task, pick a cheaper model, or raise the budget.
513
+ - **`budget_exceeded: true`** — `cost_usd` came in over the agent's `max_budget_usd` (or the settings default) without the CLI stopping the run (the final turn can overshoot the cap). The dispatch is not failed — the money is already spent — but a runaway agent is now visible. Tighten the task, pick a cheaper model, or raise the budget. A run the CLI *did* stop fails with `error_type: "budget"` instead.
511
514
 
512
515
  Tool-level errors (unknown agent, malformed input) return a plain envelope instead of a `DispatchResult`:
513
516
 
@@ -539,6 +542,23 @@ agents:
539
542
  # disallowed_tools: # block specific tools
540
543
  # - Write
541
544
 
545
+ # Optional: bundle related agents into a cross-project working set.
546
+ # A descriptive layer — no router; the orchestrating session coordinates
547
+ # with the normal dispatch tools. See the Groups section above.
548
+ groups:
549
+ shop:
550
+ # ORCHESTRATOR-facing: how to coordinate the group. Surfaced by
551
+ # list_groups/inspect_group, NEVER injected into a member's prompt.
552
+ description: "After a code change: deploy via infra, then verify via analytics."
553
+ # MEMBER-facing facts, auto-prepended to dispatch(..., group="shop").
554
+ shared_context: |
555
+ Prod runs in Portainer stack "shop". Metrica counter 12345.
556
+ members: # reference agents above (many-to-many)
557
+ - agent: infra
558
+ use_for: deploy, restart, container logs
559
+ # - agent: backend
560
+ # use_for: orders/payments endpoints
561
+
542
562
  settings:
543
563
  default_timeout: 300
544
564
  # default_permission_mode: bypassPermissions # inherited by all agents
@@ -605,10 +625,10 @@ agent-dispatch MCP server
605
625
  - **Argument-injection guard** — structured CLI fields (`session_id`, `model`, `permission_mode`, tool names) that start with `-` are rejected so they can't smuggle extra `claude` flags.
606
626
  - **Path-traversal guard** — caller-supplied `job_id`/`ref` values are validated as 32-char hex before any filesystem access.
607
627
  - **Owner-only state** — job files (`0o600`) and `agents.yaml` (`0o600`) are written for the owner only; their directories are `0o700`.
608
- - **Cost visibility** — `max_budget_usd` per agent or globally; a dispatch whose cost exceeds it returns `budget_exceeded: true` + a hint (post-hoc — the `claude` CLI has no spend cap, so the overage can be flagged but not prevented).
628
+ - **Cost control** — `max_budget_usd` per agent or globally is passed to the `claude` CLI as `--max-budget-usd`, so a runaway dispatch is stopped at the cap and comes back as `error_type: "budget"` with a resumable `session_id`. An overshoot that lands over budget without stopping is flagged post-hoc with `budget_exceeded: true` + a hint.
609
629
  - **Concurrency** — `max_concurrency` (default: 5) caps parallel `claude -p` processes. Note: the sync and async dispatch paths use separate semaphores, so the worst-case total is `2 × max_concurrency`.
610
630
  - **Timeout** — per-agent or global (default: 300s). Orphaned processes are cleaned up.
611
- - **Caching** — identical `(agent, task, context, caller, goal, response_format)` requests return cached results, bounded by `cache.max_size` (oldest entry evicted first). Only successes are cached. Sessions and dialogues are never cached.
631
+ - **Caching** — identical `(agent, task, context, caller, goal, response_format)` requests return cached results, bounded by `cache.max_size` (oldest entry evicted first). Only successes are cached. Sessions and dialogues are never cached. A `group=` dispatch folds the group's `shared_context` into `context`, so different groups cache separately and a plain dispatch is unaffected.
612
632
 
613
633
  See [SECURITY.md](SECURITY.md) for the full threat model (including the `bypassPermissions` escalation risk and on-disk job files).
614
634
 
@@ -621,9 +641,10 @@ See [SECURITY.md](SECURITY.md) for the full threat model (including the `bypassP
621
641
  | `agent-dispatch update <name>` | Update agent config (permissions, timeout, model, etc.) |
622
642
  | `agent-dispatch remove <name>` | Remove an agent |
623
643
  | `agent-dispatch list` | List agents with health status and permissions |
644
+ | `agent-dispatch group <add\|list\|inspect\|update\|remove>` | Manage [groups](#groups) — cross-project working sets of agents |
624
645
  | `agent-dispatch describe <name>` | Show full configuration for one agent (tri-state tools, project files) |
625
646
  | `agent-dispatch test <name> [task] [--stream]` | Test an agent with a dispatch (`--stream` for live progress) |
626
- | `agent-dispatch doctor` | Diagnose installation: claude CLI, MCP registration, agent health |
647
+ | `agent-dispatch doctor` | Diagnose installation: Claude CLI, MCP registration, agent health, and group membership |
627
648
  | `agent-dispatch jobs [--status --limit]` | List async dispatch jobs (most recent first) |
628
649
  | `agent-dispatch job <id>` | Show one job: status, progress tail, result preview |
629
650
  | `agent-dispatch cancel <id>` | Cancel a pending job (running jobs: use the `dispatch_cancel` MCP tool) |
@@ -13,6 +13,8 @@
13
13
 
14
14
  Each agent runs as a separate `claude -p` session in its own project directory — inheriting that project's MCP servers, CLAUDE.md, and tools. The calling agent just gets the result back.
15
15
 
16
+ Related projects can be bundled into a **[group](#groups)** — a shared brief plus a member list — so one session can coordinate work across them (e.g. code repos + an `infra`/Portainer gateway + an `analytics` gateway).
17
+
16
18
  Works with OAuth, API key, and Claude subscription authentication.
17
19
 
18
20
  > **AI agents:** this README is the canonical doc for *using* the tool — setup: [Quick Start](#quick-start) (every step has a deterministic verify), first call: [`dispatch`](#dispatch), tool selection: [Which Tool to Use](#which-tool-to-use), failure handling: [Error Recovery](#error-recovery). Working *on* this repo instead? See [AGENTS.md](AGENTS.md).
@@ -201,7 +203,7 @@ dispatch(
201
203
  }
202
204
  ```
203
205
 
204
- **`error_type` values:** `permission` (tool/action denied), `timeout`, `recursion` (dispatch depth exceeded), `not_found` (missing directory or CLI), `cli_error` (other failures). Permission errors include an actionable hint.
206
+ **`error_type` values:** `permission` (tool/action denied), `timeout`, `recursion` (dispatch depth exceeded), `not_found` (missing directory or CLI), `budget` (the `claude` CLI stopped the session at `max_budget_usd`), `cli_error` (other failures). Permission and budget errors include an actionable hint.
205
207
 
206
208
  **Resumable timeouts:** every fresh dispatch pre-assigns a session UUID (`--session-id`), so a timed-out dispatch still returns a `session_id` — the partial transcript survives the kill. The timeout error spells out the recovery: resume with `dispatch_session(agent, "Continue where you left off", session_id=...)`, retry with a bigger `timeout_seconds`, or use `dispatch_async`.
207
209
 
@@ -471,13 +473,14 @@ Failures are deterministic: check `success`, then branch on `error_type`.
471
473
  | `timeout` | Process killed at the timeout | Resume the partial work: `dispatch_session(agent, "Continue where you left off", session_id=<from the error text>)`. Or retry with a bigger `timeout_seconds=`, or use `dispatch_async`. |
472
474
  | `not_found` | Agent directory or `claude` CLI missing | `list_agents()` → check `healthy`. Re-add the agent with an existing path, or run `agent-dispatch doctor` to find what's missing. |
473
475
  | `recursion` | Dispatch nesting exceeded `max_dispatch_depth` (default 3) | Don't dispatch from dispatched agents; if the nesting is intentional, raise `max_dispatch_depth` in settings. |
476
+ | `budget` | The `claude` CLI ended the session at the `max_budget_usd` spend cap — the answer is incomplete | Raise the cap (`update_agent(name, max_budget_usd=2.0)`), switch to a cheaper `model`, or split the task. The partial session is resumable: `dispatch_session(agent, "Continue where you left off", session_id=<from the result>)`. |
474
477
  | `cli_error` | Anything else from the `claude` subprocess | Read the `error` text; run `agent-dispatch doctor` for environment issues; retry once if transient. |
475
478
 
476
479
  Three soft signals that arrive with `success: true`:
477
480
 
478
481
  - **`denied_tools` + `hint`** — the agent finished but some tool calls were blocked; the result may be incomplete. Grant access (see the `permission` row) and re-dispatch.
479
482
  - **`parsed_result: null` with `response_format="json"`** — the reply wasn't valid JSON; the raw text is still in `result`. Caveat: an agent that *can't* comply returns `{"error": "<reason>"}` — which parses successfully — so also check `parsed_result` for an `"error"` key.
480
- - **`budget_exceeded: true`** — `cost_usd` exceeded the agent's `max_budget_usd` (or the settings default). The dispatch is not failed — the money is already spent — but a runaway agent is now visible. Tighten the task, pick a cheaper model, or raise the budget.
483
+ - **`budget_exceeded: true`** — `cost_usd` came in over the agent's `max_budget_usd` (or the settings default) without the CLI stopping the run (the final turn can overshoot the cap). The dispatch is not failed — the money is already spent — but a runaway agent is now visible. Tighten the task, pick a cheaper model, or raise the budget. A run the CLI *did* stop fails with `error_type: "budget"` instead.
481
484
 
482
485
  Tool-level errors (unknown agent, malformed input) return a plain envelope instead of a `DispatchResult`:
483
486
 
@@ -509,6 +512,23 @@ agents:
509
512
  # disallowed_tools: # block specific tools
510
513
  # - Write
511
514
 
515
+ # Optional: bundle related agents into a cross-project working set.
516
+ # A descriptive layer — no router; the orchestrating session coordinates
517
+ # with the normal dispatch tools. See the Groups section above.
518
+ groups:
519
+ shop:
520
+ # ORCHESTRATOR-facing: how to coordinate the group. Surfaced by
521
+ # list_groups/inspect_group, NEVER injected into a member's prompt.
522
+ description: "After a code change: deploy via infra, then verify via analytics."
523
+ # MEMBER-facing facts, auto-prepended to dispatch(..., group="shop").
524
+ shared_context: |
525
+ Prod runs in Portainer stack "shop". Metrica counter 12345.
526
+ members: # reference agents above (many-to-many)
527
+ - agent: infra
528
+ use_for: deploy, restart, container logs
529
+ # - agent: backend
530
+ # use_for: orders/payments endpoints
531
+
512
532
  settings:
513
533
  default_timeout: 300
514
534
  # default_permission_mode: bypassPermissions # inherited by all agents
@@ -575,10 +595,10 @@ agent-dispatch MCP server
575
595
  - **Argument-injection guard** — structured CLI fields (`session_id`, `model`, `permission_mode`, tool names) that start with `-` are rejected so they can't smuggle extra `claude` flags.
576
596
  - **Path-traversal guard** — caller-supplied `job_id`/`ref` values are validated as 32-char hex before any filesystem access.
577
597
  - **Owner-only state** — job files (`0o600`) and `agents.yaml` (`0o600`) are written for the owner only; their directories are `0o700`.
578
- - **Cost visibility** — `max_budget_usd` per agent or globally; a dispatch whose cost exceeds it returns `budget_exceeded: true` + a hint (post-hoc — the `claude` CLI has no spend cap, so the overage can be flagged but not prevented).
598
+ - **Cost control** — `max_budget_usd` per agent or globally is passed to the `claude` CLI as `--max-budget-usd`, so a runaway dispatch is stopped at the cap and comes back as `error_type: "budget"` with a resumable `session_id`. An overshoot that lands over budget without stopping is flagged post-hoc with `budget_exceeded: true` + a hint.
579
599
  - **Concurrency** — `max_concurrency` (default: 5) caps parallel `claude -p` processes. Note: the sync and async dispatch paths use separate semaphores, so the worst-case total is `2 × max_concurrency`.
580
600
  - **Timeout** — per-agent or global (default: 300s). Orphaned processes are cleaned up.
581
- - **Caching** — identical `(agent, task, context, caller, goal, response_format)` requests return cached results, bounded by `cache.max_size` (oldest entry evicted first). Only successes are cached. Sessions and dialogues are never cached.
601
+ - **Caching** — identical `(agent, task, context, caller, goal, response_format)` requests return cached results, bounded by `cache.max_size` (oldest entry evicted first). Only successes are cached. Sessions and dialogues are never cached. A `group=` dispatch folds the group's `shared_context` into `context`, so different groups cache separately and a plain dispatch is unaffected.
582
602
 
583
603
  See [SECURITY.md](SECURITY.md) for the full threat model (including the `bypassPermissions` escalation risk and on-disk job files).
584
604
 
@@ -591,9 +611,10 @@ See [SECURITY.md](SECURITY.md) for the full threat model (including the `bypassP
591
611
  | `agent-dispatch update <name>` | Update agent config (permissions, timeout, model, etc.) |
592
612
  | `agent-dispatch remove <name>` | Remove an agent |
593
613
  | `agent-dispatch list` | List agents with health status and permissions |
614
+ | `agent-dispatch group <add\|list\|inspect\|update\|remove>` | Manage [groups](#groups) — cross-project working sets of agents |
594
615
  | `agent-dispatch describe <name>` | Show full configuration for one agent (tri-state tools, project files) |
595
616
  | `agent-dispatch test <name> [task] [--stream]` | Test an agent with a dispatch (`--stream` for live progress) |
596
- | `agent-dispatch doctor` | Diagnose installation: claude CLI, MCP registration, agent health |
617
+ | `agent-dispatch doctor` | Diagnose installation: Claude CLI, MCP registration, agent health, and group membership |
597
618
  | `agent-dispatch jobs [--status --limit]` | List async dispatch jobs (most recent first) |
598
619
  | `agent-dispatch job <id>` | Show one job: status, progress tail, result preview |
599
620
  | `agent-dispatch cancel <id>` | Cancel a pending job (running jobs: use the `dispatch_cancel` MCP tool) |
@@ -10,6 +10,17 @@ agents:
10
10
  - restart_services
11
11
  timeout: 300
12
12
 
13
+ # Analytics gateway — a browser + Yandex Metrica agent, no codebase of its own.
14
+ # Its value is its MCP servers + access, not its source. Shared across groups.
15
+ analytics:
16
+ directory: ~/projects/analytics
17
+ description: "Analytics gateway. MCP servers: browser, yandex-metrica. Pulls funnels, conversion, traffic. Read-only."
18
+ capabilities:
19
+ - funnel_report
20
+ - conversion_metrics
21
+ permission_mode: bypassPermissions # runs non-interactively against read-only sources
22
+ timeout: 300
23
+
13
24
  # Backend agent — source code, tests, database
14
25
  backend:
15
26
  directory: ~/projects/backend
@@ -62,13 +73,28 @@ groups:
62
73
  use_for: orders/payments endpoints, migrations
63
74
  - agent: infra
64
75
  use_for: deploy, restart, container logs
65
- # - agent: analytics # a shared gateway agent could also live here
76
+ - agent: analytics
77
+ use_for: funnel + conversion verification
78
+
79
+ # A second product reusing the SHARED infra/analytics gateways. Membership is
80
+ # many-to-many — gateways are referenced, not owned, so one analytics/infra
81
+ # agent serves every product.
82
+ blog:
83
+ description: "Content site. Same deploy-then-verify loop via the shared gateways."
84
+ shared_context: |
85
+ Production runs in Portainer stack "blog". Metrica counter 67890.
86
+ members:
87
+ - agent: infra
88
+ use_for: deploy, logs
89
+ - agent: analytics
90
+ use_for: pageview + bounce-rate checks
66
91
 
67
92
  settings:
68
93
  default_timeout: 300
69
94
  max_dispatch_depth: 3 # recursion protection: A -> B -> A
70
95
  max_concurrency: 5 # max parallel claude -p processes
71
- # default_max_budget_usd: 1.0 # flags results over this cost (budget_exceeded + hint; post-hoc, can't prevent the spend)
96
+ # default_max_budget_usd: 1.0 # spend cap per dispatch, passed to claude as --max-budget-usd
97
+ # (a run stopped at the cap fails with error_type: budget)
72
98
  # default_permission_mode: bypassPermissions # inherited by agents without override
73
99
  # default_allowed_tools: # inherited by agents without override
74
100
  # - Bash
@@ -1,6 +1,6 @@
1
1
  [project]
2
2
  name = "agent-dispatch"
3
- version = "0.9.0"
3
+ version = "0.11.0"
4
4
  description = "MCP server that lets Claude Code agents delegate tasks to agents in other project directories"
5
5
  readme = "README.md"
6
6
  license = "MIT"
@@ -1,3 +1,3 @@
1
1
  """agent-dispatch: Delegate tasks between Claude Code agents across projects."""
2
2
 
3
- __version__ = "0.9.0"
3
+ __version__ = "0.11.0"
@@ -423,6 +423,12 @@ def test(name: str, task: str, stream: bool, timeout: int | None) -> None:
423
423
  click.echo(click.style("Diagnosis: timeout", fg="yellow"))
424
424
  click.echo(f" agent-dispatch test {name} --timeout 600 # one-off")
425
425
  click.echo(f" agent-dispatch update {name} --timeout 600 # permanent")
426
+ elif result.error_type == "budget":
427
+ click.echo()
428
+ click.echo(click.style("Diagnosis: spend cap reached", fg="yellow"))
429
+ click.echo("The claude CLI stopped the session at --max-budget-usd.")
430
+ click.echo(f" agent-dispatch update {name} --max-budget-usd 2.0 # raise the cap")
431
+ click.echo(f" agent-dispatch update {name} --model haiku # cheaper model")
426
432
  raise SystemExit(1)
427
433
 
428
434
 
@@ -563,12 +569,13 @@ def group_list() -> None:
563
569
  click.echo(f" {click.style(name, bold=True)} ({len(grp.members)} member(s))")
564
570
  if grp.description:
565
571
  click.echo(f" desc: {grp.description}")
572
+ unknown_members = set(config.unknown_group_members(grp))
566
573
  rendered: list[str] = []
567
574
  for m in grp.members:
568
- if m.agent in config.agents:
569
- rendered.append(m.agent)
570
- else:
575
+ if m.agent in unknown_members:
571
576
  rendered.append(click.style(f"{m.agent}(unknown)", fg="red"))
577
+ else:
578
+ rendered.append(m.agent)
572
579
  if rendered:
573
580
  click.echo(f" members: {', '.join(rendered)}")
574
581
  click.echo(f" shared context: {'yes' if grp.shared_context.strip() else 'no'}")
@@ -593,8 +600,9 @@ def group_inspect(name: str) -> None:
593
600
  for line in grp.shared_context.splitlines():
594
601
  click.echo(f" {line}")
595
602
  click.echo(f" members ({len(grp.members)}):")
603
+ unknown_members = set(config.unknown_group_members(grp))
596
604
  for m in grp.members:
597
- marker = "" if m.agent in config.agents else click.style(" (unknown)", fg="red")
605
+ marker = click.style(" (unknown)", fg="red") if m.agent in unknown_members else ""
598
606
  hint = f" — {m.use_for}" if m.use_for else ""
599
607
  click.echo(f" - {m.agent}{marker}{hint}")
600
608
  if not grp.members:
@@ -750,6 +758,32 @@ def doctor() -> None:
750
758
  except OSError as e:
751
759
  fail(f"{name}: directory unreadable - {e}")
752
760
 
761
+ section("Groups")
762
+ if config is None:
763
+ warn("Skipped (config could not be loaded)")
764
+ elif not config.groups:
765
+ ok("No groups configured")
766
+ else:
767
+ for name, group in config.groups.items():
768
+ unknown = config.unknown_group_members(group)
769
+ if unknown:
770
+ missing = ", ".join(unknown)
771
+ fail(f"{name}: unknown member(s): {missing}")
772
+ click.echo(
773
+ f" Fix by recreating: agent-dispatch group remove {name} && "
774
+ f"agent-dispatch group add {name} --member ... (or edit agents.yaml)"
775
+ )
776
+ elif not group.members:
777
+ warn(f"{name}: no members configured")
778
+ click.echo(
779
+ f" Add members by recreating: agent-dispatch group remove {name} && "
780
+ f"agent-dispatch group add {name} --member agent1 --member agent2"
781
+ )
782
+ else:
783
+ count = len(group.members)
784
+ suffix = "member" if count == 1 else "members"
785
+ ok(f"{name}: {count} {suffix}")
786
+
753
787
  section("Summary")
754
788
  issues = counters["issues"]
755
789
  warnings = counters["warnings"]
@@ -80,8 +80,13 @@ def _collect_mcp_servers(directory: Path) -> list[str]:
80
80
  if path.exists():
81
81
  try:
82
82
  data = json.loads(path.read_text(encoding="utf-8"))
83
- servers.extend(data.get("mcpServers", {}).keys())
84
- except (json.JSONDecodeError, KeyError):
83
+ if not isinstance(data, dict):
84
+ raise ValueError("top-level JSON value is not an object")
85
+ configured = data.get("mcpServers", {})
86
+ if not isinstance(configured, dict):
87
+ raise ValueError("mcpServers is not an object")
88
+ servers.extend(str(name) for name in configured)
89
+ except (OSError, UnicodeDecodeError, json.JSONDecodeError, ValueError):
85
90
  logger.debug("Failed to parse MCP config: %s", path)
86
91
  return list(dict.fromkeys(servers)) # deduplicate, preserve order
87
92
 
@@ -137,7 +142,12 @@ def auto_describe(directory: Path) -> str:
137
142
  claude_md = directory / "CLAUDE.md"
138
143
  if claude_md.exists():
139
144
  sentences: list[str] = []
140
- for line in claude_md.read_text(encoding="utf-8").strip().splitlines()[:40]:
145
+ try:
146
+ lines = claude_md.read_text(encoding="utf-8").strip().splitlines()[:40]
147
+ except (OSError, UnicodeDecodeError):
148
+ logger.debug("Failed to read CLAUDE.md: %s", claude_md)
149
+ lines = []
150
+ for line in lines:
141
151
  stripped = line.strip()
142
152
  if stripped and not stripped.startswith("#") and not stripped.startswith("--"):
143
153
  sentences.append(stripped)
@@ -150,7 +160,12 @@ def auto_describe(directory: Path) -> str:
150
160
  if not parts:
151
161
  readme = directory / "README.md"
152
162
  if readme.exists():
153
- for line in readme.read_text(encoding="utf-8").strip().splitlines()[:20]:
163
+ try:
164
+ lines = readme.read_text(encoding="utf-8").strip().splitlines()[:20]
165
+ except (OSError, UnicodeDecodeError):
166
+ logger.debug("Failed to read README.md: %s", readme)
167
+ lines = []
168
+ for line in lines:
154
169
  stripped = line.strip()
155
170
  if (
156
171
  stripped
@@ -165,9 +180,17 @@ def auto_describe(directory: Path) -> str:
165
180
  # pyproject.toml — project description
166
181
  pyproject = directory / "pyproject.toml"
167
182
  if pyproject.exists():
168
- for line in pyproject.read_text(encoding="utf-8").splitlines():
183
+ try:
184
+ lines = pyproject.read_text(encoding="utf-8").splitlines()
185
+ except (OSError, UnicodeDecodeError):
186
+ logger.debug("Failed to read pyproject.toml: %s", pyproject)
187
+ lines = []
188
+ for line in lines:
169
189
  if line.strip().startswith("description"):
170
- desc = line.split("=", 1)[1].strip().strip('"').strip("'")
190
+ _, separator, value = line.partition("=")
191
+ if not separator:
192
+ continue
193
+ desc = value.strip().strip('"').strip("'")
171
194
  if desc:
172
195
  parts.append(desc)
173
196
  break
@@ -177,9 +200,11 @@ def auto_describe(directory: Path) -> str:
177
200
  if pkg_json.exists():
178
201
  try:
179
202
  pkg = json.loads(pkg_json.read_text(encoding="utf-8"))
180
- if pkg.get("description"):
181
- parts.append(pkg["description"])
182
- except (json.JSONDecodeError, KeyError):
203
+ if isinstance(pkg, dict):
204
+ desc = pkg.get("description")
205
+ if isinstance(desc, str) and desc.strip():
206
+ parts.append(desc)
207
+ except (OSError, UnicodeDecodeError, json.JSONDecodeError):
183
208
  logger.debug("Failed to parse package.json: %s", pkg_json)
184
209
 
185
210
  # MCP servers — critical for understanding what tools this agent has
@@ -99,10 +99,22 @@ class JobStore:
99
99
 
100
100
  def _write(self, job: Job) -> None:
101
101
  path = self._path(job.id)
102
- tmp = path.with_suffix(".tmp")
103
- tmp.write_text(job.model_dump_json(indent=2, exclude_none=True), encoding="utf-8")
104
- _chmod_quiet(tmp, 0o600) # owner-only before it becomes visible
105
- os.replace(tmp, path)
102
+ # Unique temp name per write: `self._lock` only serializes writers inside
103
+ # one process, but the CLI (cancel/gc) and the MCP server touch the same
104
+ # files. A shared `<id>.tmp` could be written by both at once and the
105
+ # rename would publish interleaved, unparseable JSON.
106
+ tmp = path.with_name(f"{job.id}.{uuid.uuid4().hex}.tmp")
107
+ try:
108
+ tmp.write_text(job.model_dump_json(indent=2, exclude_none=True), encoding="utf-8")
109
+ _chmod_quiet(tmp, 0o600) # owner-only before it becomes visible
110
+ os.replace(tmp, path)
111
+ except OSError:
112
+ # Don't leave a half-written temp file behind on a failed write.
113
+ try:
114
+ tmp.unlink(missing_ok=True)
115
+ except OSError: # pragma: no cover - best effort
116
+ logger.debug("Failed to clean up temp file %s", tmp)
117
+ raise
106
118
 
107
119
  def create(
108
120
  self,
@@ -192,17 +204,18 @@ class JobStore:
192
204
  def mark_running(self, job_id: str) -> Job | None:
193
205
  """Mark a pending job as running.
194
206
 
195
- Returns the updated job, or None if the job is missing OR has already
196
- been cancelled. Refusing to run a cancelled job closes the race with
197
- ``cancel()``: both take ``self._lock``, so whichever wins, the worker
198
- either sees ``cancelled`` (and skips) or sets ``running`` first (and
199
- cancel then refuses).
207
+ Returns the updated job, or None unless the job exists and is still
208
+ pending. Refusing every non-pending state prevents a duplicate worker
209
+ from resurrecting a completed/failed/cancelled job. It also closes the
210
+ race with ``cancel()``: both take ``self._lock``, so whichever wins,
211
+ the worker either sees ``cancelled`` (and skips) or sets ``running``
212
+ first (and cancel then refuses).
200
213
  """
201
214
  with self._lock:
202
215
  job = self.get(job_id)
203
216
  if job is None:
204
217
  return None
205
- if job.status == "cancelled":
218
+ if job.status != "pending":
206
219
  return None
207
220
  job.status = "running"
208
221
  job.started_at = time.time()
@@ -158,6 +158,10 @@ class DispatchConfig(BaseModel):
158
158
  validate_agent_name(name)
159
159
  return self
160
160
 
161
+ def unknown_group_members(self, group: DispatchGroup) -> list[str]:
162
+ """Member agent names in `group` that aren't in `self.agents` (sorted, deduped)."""
163
+ return sorted({m.agent for m in group.members if m.agent not in self.agents})
164
+
161
165
 
162
166
  class DispatchResult(BaseModel):
163
167
  """Result of a dispatch call."""
@@ -170,7 +174,7 @@ class DispatchResult(BaseModel):
170
174
  duration_ms: int | None = None
171
175
  num_turns: int | None = None
172
176
  error: str | None = None
173
- error_type: str | None = None # permission, timeout, recursion, not_found, cli_error
177
+ error_type: str | None = None # permission, timeout, recursion, not_found, budget, cli_error
174
178
  # Set when response_format="json" was requested AND the agent's result
175
179
  # parsed cleanly. None means: not requested, or requested but unparseable.
176
180
  parsed_result: Any | None = None