agent-dispatch 0.9.0__tar.gz → 0.11.0__tar.gz
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- {agent_dispatch-0.9.0 → agent_dispatch-0.11.0}/AGENTS.md +12 -6
- {agent_dispatch-0.9.0 → agent_dispatch-0.11.0}/CHANGELOG.md +70 -1
- {agent_dispatch-0.9.0 → agent_dispatch-0.11.0}/PKG-INFO +27 -6
- {agent_dispatch-0.9.0 → agent_dispatch-0.11.0}/README.md +26 -5
- {agent_dispatch-0.9.0 → agent_dispatch-0.11.0}/agents.example.yaml +28 -2
- {agent_dispatch-0.9.0 → agent_dispatch-0.11.0}/pyproject.toml +1 -1
- {agent_dispatch-0.9.0 → agent_dispatch-0.11.0}/src/agent_dispatch/__init__.py +1 -1
- {agent_dispatch-0.9.0 → agent_dispatch-0.11.0}/src/agent_dispatch/cli.py +38 -4
- {agent_dispatch-0.9.0 → agent_dispatch-0.11.0}/src/agent_dispatch/config.py +34 -9
- {agent_dispatch-0.9.0 → agent_dispatch-0.11.0}/src/agent_dispatch/jobs.py +23 -10
- {agent_dispatch-0.9.0 → agent_dispatch-0.11.0}/src/agent_dispatch/models.py +5 -1
- {agent_dispatch-0.9.0 → agent_dispatch-0.11.0}/src/agent_dispatch/runner.py +147 -58
- {agent_dispatch-0.9.0 → agent_dispatch-0.11.0}/src/agent_dispatch/server.py +4 -2
- {agent_dispatch-0.9.0 → agent_dispatch-0.11.0}/tests/test_cli.py +69 -0
- {agent_dispatch-0.9.0 → agent_dispatch-0.11.0}/tests/test_config.py +29 -1
- {agent_dispatch-0.9.0 → agent_dispatch-0.11.0}/tests/test_jobs.py +72 -9
- {agent_dispatch-0.9.0 → agent_dispatch-0.11.0}/tests/test_runner.py +500 -151
- {agent_dispatch-0.9.0 → agent_dispatch-0.11.0}/.github/dependabot.yml +0 -0
- {agent_dispatch-0.9.0 → agent_dispatch-0.11.0}/.github/workflows/ci.yml +0 -0
- {agent_dispatch-0.9.0 → agent_dispatch-0.11.0}/.github/workflows/publish.yml +0 -0
- {agent_dispatch-0.9.0 → agent_dispatch-0.11.0}/.gitignore +0 -0
- {agent_dispatch-0.9.0 → agent_dispatch-0.11.0}/LICENSE +0 -0
- {agent_dispatch-0.9.0 → agent_dispatch-0.11.0}/SECURITY.md +0 -0
- {agent_dispatch-0.9.0 → agent_dispatch-0.11.0}/assets/mascot.png +0 -0
- {agent_dispatch-0.9.0 → agent_dispatch-0.11.0}/src/agent_dispatch/cache.py +0 -0
- {agent_dispatch-0.9.0 → agent_dispatch-0.11.0}/tests/__init__.py +0 -0
- {agent_dispatch-0.9.0 → agent_dispatch-0.11.0}/tests/conftest.py +0 -0
- {agent_dispatch-0.9.0 → agent_dispatch-0.11.0}/tests/test_cache.py +0 -0
- {agent_dispatch-0.9.0 → agent_dispatch-0.11.0}/tests/test_models.py +0 -0
- {agent_dispatch-0.9.0 → agent_dispatch-0.11.0}/tests/test_server.py +0 -0
|
@@ -28,7 +28,7 @@ pip install -e ".[dev]"
|
|
|
28
28
|
|
|
29
29
|
```bash
|
|
30
30
|
ruff check src/ tests/
|
|
31
|
-
python3 -m pytest tests/ -v #
|
|
31
|
+
python3 -m pytest tests/ -v # 495 tests, ~2s — all subprocess calls are mocked
|
|
32
32
|
```
|
|
33
33
|
|
|
34
34
|
Tests must **never** invoke the real `claude` CLI. Runner tests mock `shutil.which` + `subprocess.run`/`Popen`; server tests mock `_get_config` + `runner.dispatch`.
|
|
@@ -36,14 +36,20 @@ Tests must **never** invoke the real `claude` CLI. Runner tests mock `shutil.whi
|
|
|
36
36
|
## Non-obvious invariants (violating these breaks real behavior)
|
|
37
37
|
|
|
38
38
|
- `allowed_tools` / `disallowed_tools` are **tri-state**: `None` = inherit settings defaults, `[]` = explicitly no tools, `[...]` = exactly these. Check with `is not None`, never `or` — `[]` is falsy but semantically distinct.
|
|
39
|
-
- `denied_tools` non-empty
|
|
39
|
+
- Error-type precedence on an `is_error` payload (`_build_error_result`): the CLI's own budget stop wins, then `denied_tools` non-empty ⇒ `error_type="permission"` regardless of the error text, then text classification.
|
|
40
40
|
- **Groups**: a group's `shared_context` is folded into the `context` *string* before the cache/runner calls (`_merge_group_context` in server.py) — runner.py and cache.py are untouched, the cache key disambiguates groups for free, and `group=""` is byte-identical to a plain dispatch. Membership is validated up front (`_validate_group_member`, separate from the pure merge so `dispatch_parallel`'s all-or-nothing pre-check holds). `DispatchConfig` validates only group *keys*, never member existence — a hard cross-ref check would brick config load when a shared gateway agent is removed; dangling refs are flagged (`unknown:true`) at read time instead.
|
|
41
41
|
- On failure, callers read `DispatchResult.error` + `error_type` — `result` holds the raw agent output even on errors.
|
|
42
42
|
- `--session-id` and `--resume` conflict — never pass both to `claude`.
|
|
43
43
|
- Valid permission modes: `default`, `plan`, `bypassPermissions` (`models.py: KNOWN_PERMISSION_MODES`).
|
|
44
|
-
- `JobStore.finish`/`fail` refuse already-terminal jobs (returns `None`) — this closes the race with force-cancel; never "fix" it by overwriting.
|
|
44
|
+
- `JobStore.finish`/`fail` refuse already-terminal jobs (returns `None`) — this closes the race with force-cancel; never "fix" it by overwriting. `mark_running` likewise refuses any job that isn't `pending`, so a stale or duplicate worker can't resurrect a finished one.
|
|
45
|
+
- "Is this group member missing?" has exactly one implementation: `DispatchConfig.unknown_group_members()`. Any new surface that lists or validates membership calls it instead of re-deriving the check.
|
|
45
46
|
- Cancelling a *running* job requires the in-memory `_running_procs` registry (server.py) — the job is marked `cancelled` **before** the subprocess is killed. Don't persist PIDs to disk (PID reuse after restart could kill an unrelated process).
|
|
46
|
-
- `max_budget_usd` is **
|
|
47
|
+
- `max_budget_usd` is enforced **by the claude CLI** (`_build_command` passes `--max-budget-usd`): a run stopped at the cap comes back `is_error` with no `result` text, and `_build_error_result` turns it into `error_type="budget"` + `budget_exceeded=True` + a resumable `session_id`. `_apply_budget` is the *secondary*, post-hoc signal for an overshoot that didn't stop the run; it never fails a dispatch.
|
|
48
|
+
- A CLI error payload can have no `result` field at all — the reason lives in `errors` / `subtype`. Read it via `_cli_error_details`, never assume `result` is populated on failure.
|
|
49
|
+
|
|
50
|
+
## Deliberately not built
|
|
51
|
+
|
|
52
|
+
These were considered — some fully implemented — and cut on purpose: an agent router / auto-dispatch (`recommend_agent` / `dispatch_auto`, removed before 0.8.0 — a keyword scorer adds little over the calling LLM at a handful of agents, and auto-dispatch can spend money or mutate a repo on a guess); groups as an execution engine (they are a descriptive layer — no routing, no per-group settings); an agent-dispatch-side budget ledger across dispatches (the CLI's own `--max-budget-usd` covers a single run; anything cumulative would need state we deliberately don't keep). Please open an issue with the use case before adding any of them.
|
|
47
53
|
|
|
48
54
|
## Conventions
|
|
49
55
|
|
|
@@ -51,8 +57,8 @@ Python ≥ 3.10 · `from __future__ import annotations` everywhere · Pydantic v
|
|
|
51
57
|
|
|
52
58
|
## When adding a feature, check every layer
|
|
53
59
|
|
|
54
|
-
`models.py` (data shape) → `runner.py` (dispatch mechanics) → `server.py` (MCP tool) → `cli.py` (CLI flag) → tests for each → `README.md` + `agents.example.yaml` (user docs).
|
|
60
|
+
`models.py` (data shape) → `config.py` (YAML round-trip + empty-collection pruning) → `runner.py` (dispatch mechanics) → `server.py` (MCP tool) → `cli.py` (CLI flag) → tests for each → `README.md` + `agents.example.yaml` (user docs).
|
|
55
61
|
|
|
56
62
|
## More detail
|
|
57
63
|
|
|
58
|
-
[README.md](README.md) documents every MCP tool with parameter tables, response shapes, and the error-recovery map — it doubles as the behavioral spec. The test suite (`tests/`,
|
|
64
|
+
[README.md](README.md) documents every MCP tool with parameter tables, response shapes, and the error-recovery map — it doubles as the behavioral spec. The test suite (`tests/`, 495 tests) encodes the exact expected behavior of every layer: when in doubt, read the tests for the module you're touching (`test_runner.py`, `test_server.py`, `test_cli.py`, ...).
|
|
@@ -7,6 +7,71 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
|
|
|
7
7
|
|
|
8
8
|
## [Unreleased]
|
|
9
9
|
|
|
10
|
+
## [0.11.0] - 2026-07-27
|
|
11
|
+
|
|
12
|
+
The spend cap was already real — now the result says so.
|
|
13
|
+
|
|
14
|
+
### Added
|
|
15
|
+
- **`error_type: "budget"`.** The `claude` CLI enforces the `--max-budget-usd`
|
|
16
|
+
value agent-dispatch has always passed: when the cap is reached it ends the
|
|
17
|
+
session and reports the failure with no `result` text. That payload used to
|
|
18
|
+
surface as the useless `"reported an error with no details"` fallback with
|
|
19
|
+
`error_type: "cli_error"`. It is now classified as `budget`, carries the
|
|
20
|
+
CLI's own reason (`Reached maximum budget ($X)`), sets
|
|
21
|
+
`budget_exceeded: true`, and spells out the recovery — raise the cap, pick a
|
|
22
|
+
cheaper model, or resume the partial session via the returned `session_id`.
|
|
23
|
+
`agent-dispatch test` prints a matching diagnosis.
|
|
24
|
+
|
|
25
|
+
### Fixed
|
|
26
|
+
- **Error details are no longer dropped.** CLI-level failures (budget
|
|
27
|
+
exhausted, max turns, execution errors) put their reason in `errors` /
|
|
28
|
+
`subtype` rather than `result`. Both `dispatch` and `dispatch_stream` now
|
|
29
|
+
read those fields when `result` is empty, so the error text names the actual
|
|
30
|
+
problem. Values are capped (5 entries × 300 chars) like every other field
|
|
31
|
+
read back from the subprocess.
|
|
32
|
+
- **Job files can no longer be corrupted by a concurrent writer.** `JobStore`
|
|
33
|
+
wrote through a fixed `<id>.tmp` path; the in-process lock does not cover
|
|
34
|
+
the CLI (`cancel` / `gc`) and the MCP server writing the same job at the
|
|
35
|
+
same time, and the rename could publish interleaved JSON. Each write now
|
|
36
|
+
uses a unique temp name and cleans it up if the write fails.
|
|
37
|
+
|
|
38
|
+
### Changed
|
|
39
|
+
- The `is_error` handling in `dispatch` and `dispatch_stream` is now one
|
|
40
|
+
shared `_build_error_result()` with an explicit precedence: budget stop >
|
|
41
|
+
denied tools > text classification.
|
|
42
|
+
- Documentation corrected throughout: `max_budget_usd` is enforced per
|
|
43
|
+
dispatch by the CLI, with the `budget_exceeded` flag as the secondary
|
|
44
|
+
post-hoc signal for a run that overshot without being stopped. Previous
|
|
45
|
+
releases described it as post-hoc only.
|
|
46
|
+
|
|
47
|
+
## [0.10.0] - 2026-07-14
|
|
48
|
+
|
|
49
|
+
`doctor` learns to check groups; three correctness fixes found in review.
|
|
50
|
+
|
|
51
|
+
### Added
|
|
52
|
+
- **`agent-dispatch doctor` diagnoses group health.** A new "Groups" section
|
|
53
|
+
reports each group's member count, flags dangling members (agents removed
|
|
54
|
+
from config but still referenced) as a failure, and empty groups as a
|
|
55
|
+
warning — each with a concrete remediation command.
|
|
56
|
+
|
|
57
|
+
### Fixed
|
|
58
|
+
- `auto_describe()`: an empty `"description"` field in `package.json` (a
|
|
59
|
+
common `npm init` placeholder) was no longer being filtered out, producing
|
|
60
|
+
malformed generated descriptions like `" | Stack: Node.js"`. Empty/blank
|
|
61
|
+
descriptions are ignored again.
|
|
62
|
+
- `doctor`'s group remediation hints pointed at `group update --members`,
|
|
63
|
+
a flag that doesn't exist (`group update` only edits `description` /
|
|
64
|
+
`shared_context`; membership is set via `group add --member`). Hints now
|
|
65
|
+
point at the working `group remove` + `group add --member` recreation path.
|
|
66
|
+
- `JobStore.mark_running()` now refuses to start any job that isn't
|
|
67
|
+
`pending` (previously only refused `cancelled`), closing a race where a
|
|
68
|
+
stale/duplicate worker could resurrect an already-finished or failed job.
|
|
69
|
+
|
|
70
|
+
### Changed
|
|
71
|
+
- Consolidated the "unknown group member" check — previously duplicated
|
|
72
|
+
across five call sites in `cli.py` and `server.py` — into a single
|
|
73
|
+
`DispatchConfig.unknown_group_members()` helper.
|
|
74
|
+
|
|
10
75
|
## [0.9.0] - 2026-06-30
|
|
11
76
|
|
|
12
77
|
Coordinate a group of related projects from one session.
|
|
@@ -361,7 +426,11 @@ cache bounding, and stale-job recovery.
|
|
|
361
426
|
- Dependabot for `pip` + `github-actions`, GitHub Actions pinned to
|
|
362
427
|
commit SHAs for supply-chain integrity.
|
|
363
428
|
|
|
364
|
-
[Unreleased]: https://github.com/ginkida/agent-dispatch/compare/v0.
|
|
429
|
+
[Unreleased]: https://github.com/ginkida/agent-dispatch/compare/v0.11.0...HEAD
|
|
430
|
+
[0.11.0]: https://github.com/ginkida/agent-dispatch/compare/v0.10.0...v0.11.0
|
|
431
|
+
[0.10.0]: https://github.com/ginkida/agent-dispatch/compare/v0.9.0...v0.10.0
|
|
432
|
+
[0.9.0]: https://github.com/ginkida/agent-dispatch/compare/v0.8.0...v0.9.0
|
|
433
|
+
[0.8.0]: https://github.com/ginkida/agent-dispatch/compare/v0.6.0...v0.8.0
|
|
365
434
|
[0.6.0]: https://github.com/ginkida/agent-dispatch/compare/v0.5.0...v0.6.0
|
|
366
435
|
[0.5.0]: https://github.com/ginkida/agent-dispatch/compare/v0.4.0...v0.5.0
|
|
367
436
|
[0.4.0]: https://github.com/ginkida/agent-dispatch/compare/v0.3.0...v0.4.0
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
Metadata-Version: 2.4
|
|
2
2
|
Name: agent-dispatch
|
|
3
|
-
Version: 0.
|
|
3
|
+
Version: 0.11.0
|
|
4
4
|
Summary: MCP server that lets Claude Code agents delegate tasks to agents in other project directories
|
|
5
5
|
Project-URL: Homepage, https://github.com/ginkida/agent-dispatch
|
|
6
6
|
Project-URL: Repository, https://github.com/ginkida/agent-dispatch
|
|
@@ -43,6 +43,8 @@ Description-Content-Type: text/markdown
|
|
|
43
43
|
|
|
44
44
|
Each agent runs as a separate `claude -p` session in its own project directory — inheriting that project's MCP servers, CLAUDE.md, and tools. The calling agent just gets the result back.
|
|
45
45
|
|
|
46
|
+
Related projects can be bundled into a **[group](#groups)** — a shared brief plus a member list — so one session can coordinate work across them (e.g. code repos + an `infra`/Portainer gateway + an `analytics` gateway).
|
|
47
|
+
|
|
46
48
|
Works with OAuth, API key, and Claude subscription authentication.
|
|
47
49
|
|
|
48
50
|
> **AI agents:** this README is the canonical doc for *using* the tool — setup: [Quick Start](#quick-start) (every step has a deterministic verify), first call: [`dispatch`](#dispatch), tool selection: [Which Tool to Use](#which-tool-to-use), failure handling: [Error Recovery](#error-recovery). Working *on* this repo instead? See [AGENTS.md](AGENTS.md).
|
|
@@ -231,7 +233,7 @@ dispatch(
|
|
|
231
233
|
}
|
|
232
234
|
```
|
|
233
235
|
|
|
234
|
-
**`error_type` values:** `permission` (tool/action denied), `timeout`, `recursion` (dispatch depth exceeded), `not_found` (missing directory or CLI), `cli_error` (other failures). Permission errors include an actionable hint.
|
|
236
|
+
**`error_type` values:** `permission` (tool/action denied), `timeout`, `recursion` (dispatch depth exceeded), `not_found` (missing directory or CLI), `budget` (the `claude` CLI stopped the session at `max_budget_usd`), `cli_error` (other failures). Permission and budget errors include an actionable hint.
|
|
235
237
|
|
|
236
238
|
**Resumable timeouts:** every fresh dispatch pre-assigns a session UUID (`--session-id`), so a timed-out dispatch still returns a `session_id` — the partial transcript survives the kill. The timeout error spells out the recovery: resume with `dispatch_session(agent, "Continue where you left off", session_id=...)`, retry with a bigger `timeout_seconds`, or use `dispatch_async`.
|
|
237
239
|
|
|
@@ -501,13 +503,14 @@ Failures are deterministic: check `success`, then branch on `error_type`.
|
|
|
501
503
|
| `timeout` | Process killed at the timeout | Resume the partial work: `dispatch_session(agent, "Continue where you left off", session_id=<from the error text>)`. Or retry with a bigger `timeout_seconds=`, or use `dispatch_async`. |
|
|
502
504
|
| `not_found` | Agent directory or `claude` CLI missing | `list_agents()` → check `healthy`. Re-add the agent with an existing path, or run `agent-dispatch doctor` to find what's missing. |
|
|
503
505
|
| `recursion` | Dispatch nesting exceeded `max_dispatch_depth` (default 3) | Don't dispatch from dispatched agents; if the nesting is intentional, raise `max_dispatch_depth` in settings. |
|
|
506
|
+
| `budget` | The `claude` CLI ended the session at the `max_budget_usd` spend cap — the answer is incomplete | Raise the cap (`update_agent(name, max_budget_usd=2.0)`), switch to a cheaper `model`, or split the task. The partial session is resumable: `dispatch_session(agent, "Continue where you left off", session_id=<from the result>)`. |
|
|
504
507
|
| `cli_error` | Anything else from the `claude` subprocess | Read the `error` text; run `agent-dispatch doctor` for environment issues; retry once if transient. |
|
|
505
508
|
|
|
506
509
|
Three soft signals that arrive with `success: true`:
|
|
507
510
|
|
|
508
511
|
- **`denied_tools` + `hint`** — the agent finished but some tool calls were blocked; the result may be incomplete. Grant access (see the `permission` row) and re-dispatch.
|
|
509
512
|
- **`parsed_result: null` with `response_format="json"`** — the reply wasn't valid JSON; the raw text is still in `result`. Caveat: an agent that *can't* comply returns `{"error": "<reason>"}` — which parses successfully — so also check `parsed_result` for an `"error"` key.
|
|
510
|
-
- **`budget_exceeded: true`** — `cost_usd`
|
|
513
|
+
- **`budget_exceeded: true`** — `cost_usd` came in over the agent's `max_budget_usd` (or the settings default) without the CLI stopping the run (the final turn can overshoot the cap). The dispatch is not failed — the money is already spent — but a runaway agent is now visible. Tighten the task, pick a cheaper model, or raise the budget. A run the CLI *did* stop fails with `error_type: "budget"` instead.
|
|
511
514
|
|
|
512
515
|
Tool-level errors (unknown agent, malformed input) return a plain envelope instead of a `DispatchResult`:
|
|
513
516
|
|
|
@@ -539,6 +542,23 @@ agents:
|
|
|
539
542
|
# disallowed_tools: # block specific tools
|
|
540
543
|
# - Write
|
|
541
544
|
|
|
545
|
+
# Optional: bundle related agents into a cross-project working set.
|
|
546
|
+
# A descriptive layer — no router; the orchestrating session coordinates
|
|
547
|
+
# with the normal dispatch tools. See the Groups section above.
|
|
548
|
+
groups:
|
|
549
|
+
shop:
|
|
550
|
+
# ORCHESTRATOR-facing: how to coordinate the group. Surfaced by
|
|
551
|
+
# list_groups/inspect_group, NEVER injected into a member's prompt.
|
|
552
|
+
description: "After a code change: deploy via infra, then verify via analytics."
|
|
553
|
+
# MEMBER-facing facts, auto-prepended to dispatch(..., group="shop").
|
|
554
|
+
shared_context: |
|
|
555
|
+
Prod runs in Portainer stack "shop". Metrica counter 12345.
|
|
556
|
+
members: # reference agents above (many-to-many)
|
|
557
|
+
- agent: infra
|
|
558
|
+
use_for: deploy, restart, container logs
|
|
559
|
+
# - agent: backend
|
|
560
|
+
# use_for: orders/payments endpoints
|
|
561
|
+
|
|
542
562
|
settings:
|
|
543
563
|
default_timeout: 300
|
|
544
564
|
# default_permission_mode: bypassPermissions # inherited by all agents
|
|
@@ -605,10 +625,10 @@ agent-dispatch MCP server
|
|
|
605
625
|
- **Argument-injection guard** — structured CLI fields (`session_id`, `model`, `permission_mode`, tool names) that start with `-` are rejected so they can't smuggle extra `claude` flags.
|
|
606
626
|
- **Path-traversal guard** — caller-supplied `job_id`/`ref` values are validated as 32-char hex before any filesystem access.
|
|
607
627
|
- **Owner-only state** — job files (`0o600`) and `agents.yaml` (`0o600`) are written for the owner only; their directories are `0o700`.
|
|
608
|
-
- **Cost
|
|
628
|
+
- **Cost control** — `max_budget_usd` per agent or globally is passed to the `claude` CLI as `--max-budget-usd`, so a runaway dispatch is stopped at the cap and comes back as `error_type: "budget"` with a resumable `session_id`. An overshoot that lands over budget without stopping is flagged post-hoc with `budget_exceeded: true` + a hint.
|
|
609
629
|
- **Concurrency** — `max_concurrency` (default: 5) caps parallel `claude -p` processes. Note: the sync and async dispatch paths use separate semaphores, so the worst-case total is `2 × max_concurrency`.
|
|
610
630
|
- **Timeout** — per-agent or global (default: 300s). Orphaned processes are cleaned up.
|
|
611
|
-
- **Caching** — identical `(agent, task, context, caller, goal, response_format)` requests return cached results, bounded by `cache.max_size` (oldest entry evicted first). Only successes are cached. Sessions and dialogues are never cached.
|
|
631
|
+
- **Caching** — identical `(agent, task, context, caller, goal, response_format)` requests return cached results, bounded by `cache.max_size` (oldest entry evicted first). Only successes are cached. Sessions and dialogues are never cached. A `group=` dispatch folds the group's `shared_context` into `context`, so different groups cache separately and a plain dispatch is unaffected.
|
|
612
632
|
|
|
613
633
|
See [SECURITY.md](SECURITY.md) for the full threat model (including the `bypassPermissions` escalation risk and on-disk job files).
|
|
614
634
|
|
|
@@ -621,9 +641,10 @@ See [SECURITY.md](SECURITY.md) for the full threat model (including the `bypassP
|
|
|
621
641
|
| `agent-dispatch update <name>` | Update agent config (permissions, timeout, model, etc.) |
|
|
622
642
|
| `agent-dispatch remove <name>` | Remove an agent |
|
|
623
643
|
| `agent-dispatch list` | List agents with health status and permissions |
|
|
644
|
+
| `agent-dispatch group <add\|list\|inspect\|update\|remove>` | Manage [groups](#groups) — cross-project working sets of agents |
|
|
624
645
|
| `agent-dispatch describe <name>` | Show full configuration for one agent (tri-state tools, project files) |
|
|
625
646
|
| `agent-dispatch test <name> [task] [--stream]` | Test an agent with a dispatch (`--stream` for live progress) |
|
|
626
|
-
| `agent-dispatch doctor` | Diagnose installation:
|
|
647
|
+
| `agent-dispatch doctor` | Diagnose installation: Claude CLI, MCP registration, agent health, and group membership |
|
|
627
648
|
| `agent-dispatch jobs [--status --limit]` | List async dispatch jobs (most recent first) |
|
|
628
649
|
| `agent-dispatch job <id>` | Show one job: status, progress tail, result preview |
|
|
629
650
|
| `agent-dispatch cancel <id>` | Cancel a pending job (running jobs: use the `dispatch_cancel` MCP tool) |
|
|
@@ -13,6 +13,8 @@
|
|
|
13
13
|
|
|
14
14
|
Each agent runs as a separate `claude -p` session in its own project directory — inheriting that project's MCP servers, CLAUDE.md, and tools. The calling agent just gets the result back.
|
|
15
15
|
|
|
16
|
+
Related projects can be bundled into a **[group](#groups)** — a shared brief plus a member list — so one session can coordinate work across them (e.g. code repos + an `infra`/Portainer gateway + an `analytics` gateway).
|
|
17
|
+
|
|
16
18
|
Works with OAuth, API key, and Claude subscription authentication.
|
|
17
19
|
|
|
18
20
|
> **AI agents:** this README is the canonical doc for *using* the tool — setup: [Quick Start](#quick-start) (every step has a deterministic verify), first call: [`dispatch`](#dispatch), tool selection: [Which Tool to Use](#which-tool-to-use), failure handling: [Error Recovery](#error-recovery). Working *on* this repo instead? See [AGENTS.md](AGENTS.md).
|
|
@@ -201,7 +203,7 @@ dispatch(
|
|
|
201
203
|
}
|
|
202
204
|
```
|
|
203
205
|
|
|
204
|
-
**`error_type` values:** `permission` (tool/action denied), `timeout`, `recursion` (dispatch depth exceeded), `not_found` (missing directory or CLI), `cli_error` (other failures). Permission errors include an actionable hint.
|
|
206
|
+
**`error_type` values:** `permission` (tool/action denied), `timeout`, `recursion` (dispatch depth exceeded), `not_found` (missing directory or CLI), `budget` (the `claude` CLI stopped the session at `max_budget_usd`), `cli_error` (other failures). Permission and budget errors include an actionable hint.
|
|
205
207
|
|
|
206
208
|
**Resumable timeouts:** every fresh dispatch pre-assigns a session UUID (`--session-id`), so a timed-out dispatch still returns a `session_id` — the partial transcript survives the kill. The timeout error spells out the recovery: resume with `dispatch_session(agent, "Continue where you left off", session_id=...)`, retry with a bigger `timeout_seconds`, or use `dispatch_async`.
|
|
207
209
|
|
|
@@ -471,13 +473,14 @@ Failures are deterministic: check `success`, then branch on `error_type`.
|
|
|
471
473
|
| `timeout` | Process killed at the timeout | Resume the partial work: `dispatch_session(agent, "Continue where you left off", session_id=<from the error text>)`. Or retry with a bigger `timeout_seconds=`, or use `dispatch_async`. |
|
|
472
474
|
| `not_found` | Agent directory or `claude` CLI missing | `list_agents()` → check `healthy`. Re-add the agent with an existing path, or run `agent-dispatch doctor` to find what's missing. |
|
|
473
475
|
| `recursion` | Dispatch nesting exceeded `max_dispatch_depth` (default 3) | Don't dispatch from dispatched agents; if the nesting is intentional, raise `max_dispatch_depth` in settings. |
|
|
476
|
+
| `budget` | The `claude` CLI ended the session at the `max_budget_usd` spend cap — the answer is incomplete | Raise the cap (`update_agent(name, max_budget_usd=2.0)`), switch to a cheaper `model`, or split the task. The partial session is resumable: `dispatch_session(agent, "Continue where you left off", session_id=<from the result>)`. |
|
|
474
477
|
| `cli_error` | Anything else from the `claude` subprocess | Read the `error` text; run `agent-dispatch doctor` for environment issues; retry once if transient. |
|
|
475
478
|
|
|
476
479
|
Three soft signals that arrive with `success: true`:
|
|
477
480
|
|
|
478
481
|
- **`denied_tools` + `hint`** — the agent finished but some tool calls were blocked; the result may be incomplete. Grant access (see the `permission` row) and re-dispatch.
|
|
479
482
|
- **`parsed_result: null` with `response_format="json"`** — the reply wasn't valid JSON; the raw text is still in `result`. Caveat: an agent that *can't* comply returns `{"error": "<reason>"}` — which parses successfully — so also check `parsed_result` for an `"error"` key.
|
|
480
|
-
- **`budget_exceeded: true`** — `cost_usd`
|
|
483
|
+
- **`budget_exceeded: true`** — `cost_usd` came in over the agent's `max_budget_usd` (or the settings default) without the CLI stopping the run (the final turn can overshoot the cap). The dispatch is not failed — the money is already spent — but a runaway agent is now visible. Tighten the task, pick a cheaper model, or raise the budget. A run the CLI *did* stop fails with `error_type: "budget"` instead.
|
|
481
484
|
|
|
482
485
|
Tool-level errors (unknown agent, malformed input) return a plain envelope instead of a `DispatchResult`:
|
|
483
486
|
|
|
@@ -509,6 +512,23 @@ agents:
|
|
|
509
512
|
# disallowed_tools: # block specific tools
|
|
510
513
|
# - Write
|
|
511
514
|
|
|
515
|
+
# Optional: bundle related agents into a cross-project working set.
|
|
516
|
+
# A descriptive layer — no router; the orchestrating session coordinates
|
|
517
|
+
# with the normal dispatch tools. See the Groups section above.
|
|
518
|
+
groups:
|
|
519
|
+
shop:
|
|
520
|
+
# ORCHESTRATOR-facing: how to coordinate the group. Surfaced by
|
|
521
|
+
# list_groups/inspect_group, NEVER injected into a member's prompt.
|
|
522
|
+
description: "After a code change: deploy via infra, then verify via analytics."
|
|
523
|
+
# MEMBER-facing facts, auto-prepended to dispatch(..., group="shop").
|
|
524
|
+
shared_context: |
|
|
525
|
+
Prod runs in Portainer stack "shop". Metrica counter 12345.
|
|
526
|
+
members: # reference agents above (many-to-many)
|
|
527
|
+
- agent: infra
|
|
528
|
+
use_for: deploy, restart, container logs
|
|
529
|
+
# - agent: backend
|
|
530
|
+
# use_for: orders/payments endpoints
|
|
531
|
+
|
|
512
532
|
settings:
|
|
513
533
|
default_timeout: 300
|
|
514
534
|
# default_permission_mode: bypassPermissions # inherited by all agents
|
|
@@ -575,10 +595,10 @@ agent-dispatch MCP server
|
|
|
575
595
|
- **Argument-injection guard** — structured CLI fields (`session_id`, `model`, `permission_mode`, tool names) that start with `-` are rejected so they can't smuggle extra `claude` flags.
|
|
576
596
|
- **Path-traversal guard** — caller-supplied `job_id`/`ref` values are validated as 32-char hex before any filesystem access.
|
|
577
597
|
- **Owner-only state** — job files (`0o600`) and `agents.yaml` (`0o600`) are written for the owner only; their directories are `0o700`.
|
|
578
|
-
- **Cost
|
|
598
|
+
- **Cost control** — `max_budget_usd` per agent or globally is passed to the `claude` CLI as `--max-budget-usd`, so a runaway dispatch is stopped at the cap and comes back as `error_type: "budget"` with a resumable `session_id`. An overshoot that lands over budget without stopping is flagged post-hoc with `budget_exceeded: true` + a hint.
|
|
579
599
|
- **Concurrency** — `max_concurrency` (default: 5) caps parallel `claude -p` processes. Note: the sync and async dispatch paths use separate semaphores, so the worst-case total is `2 × max_concurrency`.
|
|
580
600
|
- **Timeout** — per-agent or global (default: 300s). Orphaned processes are cleaned up.
|
|
581
|
-
- **Caching** — identical `(agent, task, context, caller, goal, response_format)` requests return cached results, bounded by `cache.max_size` (oldest entry evicted first). Only successes are cached. Sessions and dialogues are never cached.
|
|
601
|
+
- **Caching** — identical `(agent, task, context, caller, goal, response_format)` requests return cached results, bounded by `cache.max_size` (oldest entry evicted first). Only successes are cached. Sessions and dialogues are never cached. A `group=` dispatch folds the group's `shared_context` into `context`, so different groups cache separately and a plain dispatch is unaffected.
|
|
582
602
|
|
|
583
603
|
See [SECURITY.md](SECURITY.md) for the full threat model (including the `bypassPermissions` escalation risk and on-disk job files).
|
|
584
604
|
|
|
@@ -591,9 +611,10 @@ See [SECURITY.md](SECURITY.md) for the full threat model (including the `bypassP
|
|
|
591
611
|
| `agent-dispatch update <name>` | Update agent config (permissions, timeout, model, etc.) |
|
|
592
612
|
| `agent-dispatch remove <name>` | Remove an agent |
|
|
593
613
|
| `agent-dispatch list` | List agents with health status and permissions |
|
|
614
|
+
| `agent-dispatch group <add\|list\|inspect\|update\|remove>` | Manage [groups](#groups) — cross-project working sets of agents |
|
|
594
615
|
| `agent-dispatch describe <name>` | Show full configuration for one agent (tri-state tools, project files) |
|
|
595
616
|
| `agent-dispatch test <name> [task] [--stream]` | Test an agent with a dispatch (`--stream` for live progress) |
|
|
596
|
-
| `agent-dispatch doctor` | Diagnose installation:
|
|
617
|
+
| `agent-dispatch doctor` | Diagnose installation: Claude CLI, MCP registration, agent health, and group membership |
|
|
597
618
|
| `agent-dispatch jobs [--status --limit]` | List async dispatch jobs (most recent first) |
|
|
598
619
|
| `agent-dispatch job <id>` | Show one job: status, progress tail, result preview |
|
|
599
620
|
| `agent-dispatch cancel <id>` | Cancel a pending job (running jobs: use the `dispatch_cancel` MCP tool) |
|
|
@@ -10,6 +10,17 @@ agents:
|
|
|
10
10
|
- restart_services
|
|
11
11
|
timeout: 300
|
|
12
12
|
|
|
13
|
+
# Analytics gateway — a browser + Yandex Metrica agent, no codebase of its own.
|
|
14
|
+
# Its value is its MCP servers + access, not its source. Shared across groups.
|
|
15
|
+
analytics:
|
|
16
|
+
directory: ~/projects/analytics
|
|
17
|
+
description: "Analytics gateway. MCP servers: browser, yandex-metrica. Pulls funnels, conversion, traffic. Read-only."
|
|
18
|
+
capabilities:
|
|
19
|
+
- funnel_report
|
|
20
|
+
- conversion_metrics
|
|
21
|
+
permission_mode: bypassPermissions # runs non-interactively against read-only sources
|
|
22
|
+
timeout: 300
|
|
23
|
+
|
|
13
24
|
# Backend agent — source code, tests, database
|
|
14
25
|
backend:
|
|
15
26
|
directory: ~/projects/backend
|
|
@@ -62,13 +73,28 @@ groups:
|
|
|
62
73
|
use_for: orders/payments endpoints, migrations
|
|
63
74
|
- agent: infra
|
|
64
75
|
use_for: deploy, restart, container logs
|
|
65
|
-
|
|
76
|
+
- agent: analytics
|
|
77
|
+
use_for: funnel + conversion verification
|
|
78
|
+
|
|
79
|
+
# A second product reusing the SHARED infra/analytics gateways. Membership is
|
|
80
|
+
# many-to-many — gateways are referenced, not owned, so one analytics/infra
|
|
81
|
+
# agent serves every product.
|
|
82
|
+
blog:
|
|
83
|
+
description: "Content site. Same deploy-then-verify loop via the shared gateways."
|
|
84
|
+
shared_context: |
|
|
85
|
+
Production runs in Portainer stack "blog". Metrica counter 67890.
|
|
86
|
+
members:
|
|
87
|
+
- agent: infra
|
|
88
|
+
use_for: deploy, logs
|
|
89
|
+
- agent: analytics
|
|
90
|
+
use_for: pageview + bounce-rate checks
|
|
66
91
|
|
|
67
92
|
settings:
|
|
68
93
|
default_timeout: 300
|
|
69
94
|
max_dispatch_depth: 3 # recursion protection: A -> B -> A
|
|
70
95
|
max_concurrency: 5 # max parallel claude -p processes
|
|
71
|
-
# default_max_budget_usd: 1.0 #
|
|
96
|
+
# default_max_budget_usd: 1.0 # spend cap per dispatch, passed to claude as --max-budget-usd
|
|
97
|
+
# (a run stopped at the cap fails with error_type: budget)
|
|
72
98
|
# default_permission_mode: bypassPermissions # inherited by agents without override
|
|
73
99
|
# default_allowed_tools: # inherited by agents without override
|
|
74
100
|
# - Bash
|
|
@@ -423,6 +423,12 @@ def test(name: str, task: str, stream: bool, timeout: int | None) -> None:
|
|
|
423
423
|
click.echo(click.style("Diagnosis: timeout", fg="yellow"))
|
|
424
424
|
click.echo(f" agent-dispatch test {name} --timeout 600 # one-off")
|
|
425
425
|
click.echo(f" agent-dispatch update {name} --timeout 600 # permanent")
|
|
426
|
+
elif result.error_type == "budget":
|
|
427
|
+
click.echo()
|
|
428
|
+
click.echo(click.style("Diagnosis: spend cap reached", fg="yellow"))
|
|
429
|
+
click.echo("The claude CLI stopped the session at --max-budget-usd.")
|
|
430
|
+
click.echo(f" agent-dispatch update {name} --max-budget-usd 2.0 # raise the cap")
|
|
431
|
+
click.echo(f" agent-dispatch update {name} --model haiku # cheaper model")
|
|
426
432
|
raise SystemExit(1)
|
|
427
433
|
|
|
428
434
|
|
|
@@ -563,12 +569,13 @@ def group_list() -> None:
|
|
|
563
569
|
click.echo(f" {click.style(name, bold=True)} ({len(grp.members)} member(s))")
|
|
564
570
|
if grp.description:
|
|
565
571
|
click.echo(f" desc: {grp.description}")
|
|
572
|
+
unknown_members = set(config.unknown_group_members(grp))
|
|
566
573
|
rendered: list[str] = []
|
|
567
574
|
for m in grp.members:
|
|
568
|
-
if m.agent in
|
|
569
|
-
rendered.append(m.agent)
|
|
570
|
-
else:
|
|
575
|
+
if m.agent in unknown_members:
|
|
571
576
|
rendered.append(click.style(f"{m.agent}(unknown)", fg="red"))
|
|
577
|
+
else:
|
|
578
|
+
rendered.append(m.agent)
|
|
572
579
|
if rendered:
|
|
573
580
|
click.echo(f" members: {', '.join(rendered)}")
|
|
574
581
|
click.echo(f" shared context: {'yes' if grp.shared_context.strip() else 'no'}")
|
|
@@ -593,8 +600,9 @@ def group_inspect(name: str) -> None:
|
|
|
593
600
|
for line in grp.shared_context.splitlines():
|
|
594
601
|
click.echo(f" {line}")
|
|
595
602
|
click.echo(f" members ({len(grp.members)}):")
|
|
603
|
+
unknown_members = set(config.unknown_group_members(grp))
|
|
596
604
|
for m in grp.members:
|
|
597
|
-
marker =
|
|
605
|
+
marker = click.style(" (unknown)", fg="red") if m.agent in unknown_members else ""
|
|
598
606
|
hint = f" — {m.use_for}" if m.use_for else ""
|
|
599
607
|
click.echo(f" - {m.agent}{marker}{hint}")
|
|
600
608
|
if not grp.members:
|
|
@@ -750,6 +758,32 @@ def doctor() -> None:
|
|
|
750
758
|
except OSError as e:
|
|
751
759
|
fail(f"{name}: directory unreadable - {e}")
|
|
752
760
|
|
|
761
|
+
section("Groups")
|
|
762
|
+
if config is None:
|
|
763
|
+
warn("Skipped (config could not be loaded)")
|
|
764
|
+
elif not config.groups:
|
|
765
|
+
ok("No groups configured")
|
|
766
|
+
else:
|
|
767
|
+
for name, group in config.groups.items():
|
|
768
|
+
unknown = config.unknown_group_members(group)
|
|
769
|
+
if unknown:
|
|
770
|
+
missing = ", ".join(unknown)
|
|
771
|
+
fail(f"{name}: unknown member(s): {missing}")
|
|
772
|
+
click.echo(
|
|
773
|
+
f" Fix by recreating: agent-dispatch group remove {name} && "
|
|
774
|
+
f"agent-dispatch group add {name} --member ... (or edit agents.yaml)"
|
|
775
|
+
)
|
|
776
|
+
elif not group.members:
|
|
777
|
+
warn(f"{name}: no members configured")
|
|
778
|
+
click.echo(
|
|
779
|
+
f" Add members by recreating: agent-dispatch group remove {name} && "
|
|
780
|
+
f"agent-dispatch group add {name} --member agent1 --member agent2"
|
|
781
|
+
)
|
|
782
|
+
else:
|
|
783
|
+
count = len(group.members)
|
|
784
|
+
suffix = "member" if count == 1 else "members"
|
|
785
|
+
ok(f"{name}: {count} {suffix}")
|
|
786
|
+
|
|
753
787
|
section("Summary")
|
|
754
788
|
issues = counters["issues"]
|
|
755
789
|
warnings = counters["warnings"]
|
|
@@ -80,8 +80,13 @@ def _collect_mcp_servers(directory: Path) -> list[str]:
|
|
|
80
80
|
if path.exists():
|
|
81
81
|
try:
|
|
82
82
|
data = json.loads(path.read_text(encoding="utf-8"))
|
|
83
|
-
|
|
84
|
-
|
|
83
|
+
if not isinstance(data, dict):
|
|
84
|
+
raise ValueError("top-level JSON value is not an object")
|
|
85
|
+
configured = data.get("mcpServers", {})
|
|
86
|
+
if not isinstance(configured, dict):
|
|
87
|
+
raise ValueError("mcpServers is not an object")
|
|
88
|
+
servers.extend(str(name) for name in configured)
|
|
89
|
+
except (OSError, UnicodeDecodeError, json.JSONDecodeError, ValueError):
|
|
85
90
|
logger.debug("Failed to parse MCP config: %s", path)
|
|
86
91
|
return list(dict.fromkeys(servers)) # deduplicate, preserve order
|
|
87
92
|
|
|
@@ -137,7 +142,12 @@ def auto_describe(directory: Path) -> str:
|
|
|
137
142
|
claude_md = directory / "CLAUDE.md"
|
|
138
143
|
if claude_md.exists():
|
|
139
144
|
sentences: list[str] = []
|
|
140
|
-
|
|
145
|
+
try:
|
|
146
|
+
lines = claude_md.read_text(encoding="utf-8").strip().splitlines()[:40]
|
|
147
|
+
except (OSError, UnicodeDecodeError):
|
|
148
|
+
logger.debug("Failed to read CLAUDE.md: %s", claude_md)
|
|
149
|
+
lines = []
|
|
150
|
+
for line in lines:
|
|
141
151
|
stripped = line.strip()
|
|
142
152
|
if stripped and not stripped.startswith("#") and not stripped.startswith("--"):
|
|
143
153
|
sentences.append(stripped)
|
|
@@ -150,7 +160,12 @@ def auto_describe(directory: Path) -> str:
|
|
|
150
160
|
if not parts:
|
|
151
161
|
readme = directory / "README.md"
|
|
152
162
|
if readme.exists():
|
|
153
|
-
|
|
163
|
+
try:
|
|
164
|
+
lines = readme.read_text(encoding="utf-8").strip().splitlines()[:20]
|
|
165
|
+
except (OSError, UnicodeDecodeError):
|
|
166
|
+
logger.debug("Failed to read README.md: %s", readme)
|
|
167
|
+
lines = []
|
|
168
|
+
for line in lines:
|
|
154
169
|
stripped = line.strip()
|
|
155
170
|
if (
|
|
156
171
|
stripped
|
|
@@ -165,9 +180,17 @@ def auto_describe(directory: Path) -> str:
|
|
|
165
180
|
# pyproject.toml — project description
|
|
166
181
|
pyproject = directory / "pyproject.toml"
|
|
167
182
|
if pyproject.exists():
|
|
168
|
-
|
|
183
|
+
try:
|
|
184
|
+
lines = pyproject.read_text(encoding="utf-8").splitlines()
|
|
185
|
+
except (OSError, UnicodeDecodeError):
|
|
186
|
+
logger.debug("Failed to read pyproject.toml: %s", pyproject)
|
|
187
|
+
lines = []
|
|
188
|
+
for line in lines:
|
|
169
189
|
if line.strip().startswith("description"):
|
|
170
|
-
|
|
190
|
+
_, separator, value = line.partition("=")
|
|
191
|
+
if not separator:
|
|
192
|
+
continue
|
|
193
|
+
desc = value.strip().strip('"').strip("'")
|
|
171
194
|
if desc:
|
|
172
195
|
parts.append(desc)
|
|
173
196
|
break
|
|
@@ -177,9 +200,11 @@ def auto_describe(directory: Path) -> str:
|
|
|
177
200
|
if pkg_json.exists():
|
|
178
201
|
try:
|
|
179
202
|
pkg = json.loads(pkg_json.read_text(encoding="utf-8"))
|
|
180
|
-
if pkg
|
|
181
|
-
|
|
182
|
-
|
|
203
|
+
if isinstance(pkg, dict):
|
|
204
|
+
desc = pkg.get("description")
|
|
205
|
+
if isinstance(desc, str) and desc.strip():
|
|
206
|
+
parts.append(desc)
|
|
207
|
+
except (OSError, UnicodeDecodeError, json.JSONDecodeError):
|
|
183
208
|
logger.debug("Failed to parse package.json: %s", pkg_json)
|
|
184
209
|
|
|
185
210
|
# MCP servers — critical for understanding what tools this agent has
|
|
@@ -99,10 +99,22 @@ class JobStore:
|
|
|
99
99
|
|
|
100
100
|
def _write(self, job: Job) -> None:
|
|
101
101
|
path = self._path(job.id)
|
|
102
|
-
|
|
103
|
-
|
|
104
|
-
|
|
105
|
-
|
|
102
|
+
# Unique temp name per write: `self._lock` only serializes writers inside
|
|
103
|
+
# one process, but the CLI (cancel/gc) and the MCP server touch the same
|
|
104
|
+
# files. A shared `<id>.tmp` could be written by both at once and the
|
|
105
|
+
# rename would publish interleaved, unparseable JSON.
|
|
106
|
+
tmp = path.with_name(f"{job.id}.{uuid.uuid4().hex}.tmp")
|
|
107
|
+
try:
|
|
108
|
+
tmp.write_text(job.model_dump_json(indent=2, exclude_none=True), encoding="utf-8")
|
|
109
|
+
_chmod_quiet(tmp, 0o600) # owner-only before it becomes visible
|
|
110
|
+
os.replace(tmp, path)
|
|
111
|
+
except OSError:
|
|
112
|
+
# Don't leave a half-written temp file behind on a failed write.
|
|
113
|
+
try:
|
|
114
|
+
tmp.unlink(missing_ok=True)
|
|
115
|
+
except OSError: # pragma: no cover - best effort
|
|
116
|
+
logger.debug("Failed to clean up temp file %s", tmp)
|
|
117
|
+
raise
|
|
106
118
|
|
|
107
119
|
def create(
|
|
108
120
|
self,
|
|
@@ -192,17 +204,18 @@ class JobStore:
|
|
|
192
204
|
def mark_running(self, job_id: str) -> Job | None:
|
|
193
205
|
"""Mark a pending job as running.
|
|
194
206
|
|
|
195
|
-
Returns the updated job, or None
|
|
196
|
-
|
|
197
|
-
|
|
198
|
-
|
|
199
|
-
|
|
207
|
+
Returns the updated job, or None unless the job exists and is still
|
|
208
|
+
pending. Refusing every non-pending state prevents a duplicate worker
|
|
209
|
+
from resurrecting a completed/failed/cancelled job. It also closes the
|
|
210
|
+
race with ``cancel()``: both take ``self._lock``, so whichever wins,
|
|
211
|
+
the worker either sees ``cancelled`` (and skips) or sets ``running``
|
|
212
|
+
first (and cancel then refuses).
|
|
200
213
|
"""
|
|
201
214
|
with self._lock:
|
|
202
215
|
job = self.get(job_id)
|
|
203
216
|
if job is None:
|
|
204
217
|
return None
|
|
205
|
-
if job.status
|
|
218
|
+
if job.status != "pending":
|
|
206
219
|
return None
|
|
207
220
|
job.status = "running"
|
|
208
221
|
job.started_at = time.time()
|
|
@@ -158,6 +158,10 @@ class DispatchConfig(BaseModel):
|
|
|
158
158
|
validate_agent_name(name)
|
|
159
159
|
return self
|
|
160
160
|
|
|
161
|
+
def unknown_group_members(self, group: DispatchGroup) -> list[str]:
|
|
162
|
+
"""Member agent names in `group` that aren't in `self.agents` (sorted, deduped)."""
|
|
163
|
+
return sorted({m.agent for m in group.members if m.agent not in self.agents})
|
|
164
|
+
|
|
161
165
|
|
|
162
166
|
class DispatchResult(BaseModel):
|
|
163
167
|
"""Result of a dispatch call."""
|
|
@@ -170,7 +174,7 @@ class DispatchResult(BaseModel):
|
|
|
170
174
|
duration_ms: int | None = None
|
|
171
175
|
num_turns: int | None = None
|
|
172
176
|
error: str | None = None
|
|
173
|
-
error_type: str | None = None # permission, timeout, recursion, not_found, cli_error
|
|
177
|
+
error_type: str | None = None # permission, timeout, recursion, not_found, budget, cli_error
|
|
174
178
|
# Set when response_format="json" was requested AND the agent's result
|
|
175
179
|
# parsed cleanly. None means: not requested, or requested but unparseable.
|
|
176
180
|
parsed_result: Any | None = None
|