bullswarm 0.35.3 → 0.35.5
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/AGENTS.md +3 -2
- package/CHANGELOG.md +87 -2
- package/README.md +149 -237
- package/data/openrouter-benchmarks.json +9963 -9937
- package/docs/design/providers-0.35.2/CHANGELOG-draft.md +2 -2
- package/docs/design/providers-0.35.2/README.md +6 -6
- package/docs/design/runs-table-0.35.4/README.md +60 -0
- package/docs/design/runs-table-0.35.4/frames/runs-120.txt +70 -0
- package/docs/design/runs-table-0.35.4/frames/runs-200.txt +70 -0
- package/docs/design/runs-table-0.35.4/frames/runs-55.txt +70 -0
- package/docs/design/stats-frames-0.33.2/spending-120.txt +18 -16
- package/docs/design/stats-frames-0.33.2/spending-200.txt +24 -0
- package/docs/design/stats-frames-0.33.2/spending-55.txt +32 -31
- package/docs/design/step-page-0.35.0/frames/failed-120.txt +2 -2
- package/docs/design/step-page-0.35.0/frames/failed-200.txt +2 -2
- package/docs/design/step-page-0.35.0/frames/failed-55.txt +2 -2
- package/docs/design/step-page-0.35.0/frames/real-failed-120.txt +3 -3
- package/docs/design/step-page-0.35.0/frames/real-failed-200.txt +4 -4
- package/docs/design/step-page-0.35.0/frames/real-failed-55.txt +3 -3
- package/docs/design/step-page-0.35.0/frames/rendered-failed-120.txt +3 -3
- package/docs/design/step-page-0.35.0/frames/rendered-failed-200.txt +3 -3
- package/docs/design/step-page-0.35.0/frames/rendered-failed-55.txt +3 -3
- package/docs/design/tidy-0.35.1/frames/colour/real-task-200.txt +3 -3
- package/docs/design/tidy-0.35.1/frames/colour/real-task-55.txt +2 -2
- package/docs/design/tidy-0.35.1/frames/real-stats-model-120.txt +13 -13
- package/docs/design/tidy-0.35.1/frames/real-stats-model-200.txt +13 -13
- package/docs/design/tidy-0.35.1/frames/real-stats-model-55.txt +9 -9
- package/docs/design/tidy-0.35.1/frames/real-stats-spending-120.txt +14 -14
- package/docs/design/tidy-0.35.1/frames/real-stats-spending-200.txt +14 -14
- package/docs/design/tidy-0.35.1/frames/real-stats-spending-55.txt +18 -18
- package/docs/design/tidy-0.35.1/frames/real-task-120.txt +3 -3
- package/docs/design/tidy-0.35.1/frames/real-task-200.txt +3 -3
- package/docs/design/tidy-0.35.1/frames/real-task-55.txt +2 -2
- package/docs/design/tidy-0.35.1/frames/rendered-failed-120.txt +1 -1
- package/docs/design/tidy-0.35.1/frames/rendered-failed-200.txt +1 -1
- package/docs/design/tidy-0.35.1/frames/rendered-failed-55.txt +1 -1
- package/docs/guide/cost.md +1 -1
- package/docs/guide/getting-started.md +62 -16
- package/docs/guide/index.md +44 -6
- package/docs/guide/observing.md +42 -0
- package/docs/guide/playbook.md +120 -0
- package/docs/guide/routing.md +39 -1
- package/docs/index.md +47 -17
- package/docs/integrations/claude-code.md +4 -2
- package/docs/notes/index.md +1 -1
- package/docs/plans/dashboard-0.33.md +1 -1
- package/docs/plans/providers-0.35.2.goal.txt +2 -2
- package/docs/plans/providers-0.35.2.program.json +3 -3
- package/docs/reference/cli.md +40 -25
- package/docs/reference/configuration.md +22 -6
- package/docs/reference/providers.md +103 -4
- package/docs/studies/cost-audit-2026-09-18/claude-actual-vs-recorded.md +1 -1
- package/mods/bullswarm/README.md +14 -8
- package/mods/bullswarm/hooks/names.ts +7 -0
- package/mods/bullswarm/hooks/pane.tsx +33 -8
- package/mods/bullswarm/hooks/pools.ts +2 -0
- package/mods/bullswarm/hooks/register.ts +126 -18
- package/mods/bullswarm/hooks/runs.ts +30 -0
- package/mods/bullswarm/hooks/step.ts +65 -14
- package/mods/bullswarm/types/index.d.ts +4 -0
- package/package.json +4 -2
- package/providers/contrib/command-code/connector-history.json +288 -0
- package/providers/contrib/opencode/connector-history.json +241 -0
- package/providers/contrib/opencode/connector.json +1 -1
- package/skill/references/operations.md +8 -3
- package/src/cli.js +112 -44
- package/src/help.js +85 -44
- package/src/home-cli.js +1 -1
- package/src/lib/cli-flags.js +3 -1
- package/src/lib/config.js +2 -0
- package/src/lib/connector-copies.js +288 -0
- package/src/lib/model-family.js +252 -0
- package/src/lib/pool-labels.js +109 -0
- package/src/lib/providers.js +2 -0
- package/src/lib/reasoning.js +51 -5
- package/src/lib/strategy.js +568 -111
- package/src/lib/tasks.js +37 -6
- package/src/provider-cli.js +63 -6
- package/src/providers/_schema.json +20 -1
- package/src/providers/claude-code/connector-history.json +330 -0
- package/src/providers/claude-code/connector.json +20 -5
- package/src/providers/claude-code/provider.mjs +124 -1
- package/src/providers/codex/connector-history.json +326 -0
- package/src/providers/codex/connector.json +53 -16
- package/src/providers/codex/provider.mjs +115 -0
- package/src/providers/echo/connector-history.json +142 -0
- package/src/providers/grok/connector-history.json +308 -0
- package/src/providers/grok/connector.json +65 -7
- package/src/setup.js +54 -58
- package/src/strategy-cli.js +215 -33
- package/src/strategy-dashboard.js +31 -10
- package/src/workflow/cli.js +60 -3
- package/src/workflow/dash-kit.js +55 -48
- package/src/workflow/dashboard.js +24 -1
- package/src/workflow/history-view.js +293 -141
- package/src/workflow/home-view.js +14 -56
- package/src/workflow/reprice.js +4 -2
- package/src/workflow/runs-cli.js +2 -1
- package/src/workflow/runs-view.js +25 -13
- package/src/workflow/stat-kit.js +118 -20
- package/src/workflow/stats-model.js +1 -0
- package/src/workflow/stats-view.js +27 -32
- package/src/workflow/step-json.js +93 -0
- package/src/workflow/step-model.js +15 -6
- package/src/workflow/watch-cli.js +5 -0
- package/docs/public/favicon.svg +0 -7
- package/src/lib/release.js +0 -56
package/AGENTS.md
CHANGED
|
@@ -69,7 +69,7 @@ is the canonical reference for the CLI surface.
|
|
|
69
69
|
- Every verb must work non-interactively (no TTY). The interactive wizard is
|
|
70
70
|
a human convenience, never a requirement.
|
|
71
71
|
- Version single source: package.json. Release via
|
|
72
|
-
`
|
|
72
|
+
`npm run release -- patch|minor|major [--title "<headline>"]`, then `git push` and
|
|
73
73
|
`git push --tags`
|
|
74
74
|
— CI publishes through npm trusted publishing (OIDC), no tokens.
|
|
75
75
|
|
|
@@ -90,7 +90,8 @@ the authoring method is `skill/references/providers.md`.
|
|
|
90
90
|
## Releasing
|
|
91
91
|
|
|
92
92
|
1. All tests green.
|
|
93
|
-
2. `
|
|
93
|
+
2. `npm run release -- patch --title "<headline>"` (dates `## Unreleased`
|
|
94
|
+
in CHANGELOG.md, bumps package.json, creates commit + tag v*).
|
|
94
95
|
3. `git push && git push --tags`.
|
|
95
96
|
4. GitHub Actions publishes to npm via trusted publishing; verify with
|
|
96
97
|
`npm view bullswarm version`.
|
package/CHANGELOG.md
CHANGED
|
@@ -1,5 +1,90 @@
|
|
|
1
1
|
# bullswarm changelog
|
|
2
2
|
|
|
3
|
+
## Unreleased
|
|
4
|
+
|
|
5
|
+
## 0.35.5 — model choices come from each CLI, newest version wins, setup stops pinning tiers
|
|
6
|
+
|
|
7
|
+
- cli: the maintainer-only `release` verb is gone from the CLI and its help;
|
|
8
|
+
maintainers run `npm run release -- patch|minor|major`, which also dates
|
|
9
|
+
the `## Unreleased` changelog section.
|
|
10
|
+
|
|
11
|
+
- package: the README screenshots and brand images (`docs/public/`) are no
|
|
12
|
+
longer shipped in the npm package; the npm page loads them from GitHub.
|
|
13
|
+
|
|
14
|
+
- grok: every tier now suggests the newest Grok, and the tiers differ only by reasoning. This is an owner decision: Grok has one model line. One family now covers every plain `grok-N.M`, and the newest version wins. A fresh home suggests `grok-4.7` on all three tiers: high at the connector's `xhigh`, then `medium: grok-4.7 · high reasoning` and `low: grok-4.7 · medium reasoning`, both with the reason `one Grok line, lighter reasoning for lighter tiers`. Before, it suggested `grok-4.6` for high, `grok-4.5` for medium, and nothing for low, and it listed `grok-4.7` as unranked. A `grok-4.8` takes over every tier the day `grok models` lists it. `grok-4.7-build-fast` is the same model at twice the price, so it keeps its base model's rank but is never suggested; `strategy set-rung` still selects it. The tier-wide best pick, the Codex and Claude suggestions, and pace-based routing across pools do not change. grok-4.7 and grok-4.5 now have dated price rows from [xAI's models page](https://docs.x.ai/docs/models), and grok-4.6's row was re-checked there (all on 2026-09-23). build-fast stays unpriced: xAI's Pricing page lists its long-context rates as $6.00 / $1.50 / $18.00 per 1M tokens, which is not twice grok-4.7's. The fallback list (`knownModels`) now matches what `grok models` lists.
|
|
15
|
+
|
|
16
|
+
- strategy: `generationFallback` now also covers a tier that no family serves. If a connector opts such a tier in and no exact row ranks a model for it, the tier takes the best-ranked family's newest-generation model at the declared reasoning. That stand-in is the pool's rung, and it stands at rank 0 against other pools, so it is the tier-wide pick only when no other pool serves the tier. A tier may also declare `why`, a plain reason that replaces the generated one; `provider validate` rejects an empty one.
|
|
17
|
+
|
|
18
|
+
- strategy: applying recommendations no longer pins tiers. `strategy apply`, `refresh --apply`, `setup --yes --strategy`, the setup wizard, the TUI's apply key, and the daily auto-refresh now write only each pool's rungs (its model and reasoning per tier). Before, they also pinned every tier to one pool: a fresh `setup --yes --strategy` sent high, medium, and low to Codex even at 100% used, and the auto-refresh pinned them again every 24 hours, so pace-based routing never picked another plan. Now every dispatch picks its pool by spare quota, and `strategy routes` shows each tier moving between pools. `strategy.assignments` stays empty unless you pin a tier yourself. The apply JSON now reports `rungs`, `bestNow` (the tier-wide pick, for display), `unpinned`, and `keptPins`, in place of `applied`.
|
|
19
|
+
|
|
20
|
+
- strategy: pins are explicit only. `strategy assign` makes one, marked `source: "user"`, and `strategy clear-assignment` removes it. Apply and the auto-refresh never remove a pin you set, and apply now keeps the pin's model on that pool's rung, so the pinned tier runs the model you named instead of the pool's recommended one.
|
|
21
|
+
|
|
22
|
+
- strategy: if an earlier setup or apply pinned your tiers, the next `strategy apply` or automatic refresh removes those pins and lists them under `unpinned`. A pin is treated as written by apply when it still has the pool and model that apply recorded in the cached report. Pins you set or changed yourself stay (`keptPins`), as does any pin in a home where apply never ran. Run `bullswarm strategy show` to see which pins remain and what removes each one.
|
|
23
|
+
|
|
24
|
+
- strategy: the tier suggestion reads as a suggestion, not a pin: `high: codex/gpt-6-astra (best now; routing picks by spare quota)` in `strategy show`, `best now (not pins): …` after the wizard and `setup --yes --strategy`, and `Routing now · by spare quota at each dispatch, unless pinned` in the TUI, whose route lines end in `by spare quota` or `pinned to <pool>`. A pinned tier gets its own line in `strategy show`, `pinned to <pool>/<model> by you · high dispatches go there while it is available · strategy clear-assignment high removes it`. `strategy routes` gives each route a `pin` field.
|
|
25
|
+
|
|
26
|
+
- strategy: `strategy show` counts unranked models instead of listing them. It prints one line, such as `unranked: 312 models (command-code 180, opencode 130, grok 2) · never recommended · strategy show --json lists them`, and the setup review screen prints the same count for each provider. `strategy show --json` keeps the full `unranked` list, and the model pickers keep their `(unranked)` tag.
|
|
27
|
+
|
|
28
|
+
- strategy: medium now falls back to the newest generation. When the family that serves a tier has no model in a pool's newest generation, a connector that opts the tier in (`generationFallback`) gives it the next-lower family's newest-generation model, at a declared reasoning level clamped to what that model supports. Codex opts medium in, and this is an owner decision, not a benchmark: 168 medium dispatches as luna at max reasoning ran 99% ok. So a fresh home suggests `medium: gpt-6-luna · max reasoning — no gpt-6 terra yet, newest generation preferred`, and medium returns to terra at the normal medium reasoning once a gpt-6 terra is discovered. High, low, and Claude are unchanged. `strategy apply`, `refresh --apply`, `setup --yes --strategy`, the wizard, and the TUI all write that level into the pool's rung, reported as `max (recommendation)`. It is never written over a level you set, and a later apply removes it when the fallback ends. `strategy show` and the setup screens say why. The cross-pool tier pick also skips a model disabled for its pool now, as each pool's own pick already did, so a tier suggestion can no longer name a model you turned off.
|
|
29
|
+
|
|
30
|
+
- setup, doctor: old copies of packaged connectors in `<home>/connectors/` now follow the package. Those copies were made by installs older than 0.29.0. A copy of a pool the package defines was already ignored, so its old prices (such as Opus 5's on Opus 5.5) never applied; a contrib copy whose provider is not enabled was still used. Every verb now checks each copy against `connector-history.json`, which holds fingerprints of every value each field has shipped with. A copy with no edits moves to `connectors/retired/` so the packaged connector loads, and a contrib provider is enabled first when only the copy defined its pool. A copy you edited is kept, and `bullswarm doctor` (`connector-copies`, a `!` warning) and `strategy show` name its stale fields. The dead copy-to-home step in `setup` is gone.
|
|
31
|
+
|
|
32
|
+
- strategy: connectors now declare model families (`modelFamilies`: a name pattern with a tier and base rank), and the version read from the model id orders each family. So a newly discovered model is ranked on the day it appears: gpt-6-sol and gpt-6-luna are no longer tierless, and a fresh home suggests gpt-6-luna for low instead of gpt-5.6-luna. Inside a pool and family, the newer version always wins, whatever benchmarks, prices, or pace say. Quality is compared on one scale, the rank, and benchmarks only break equal ranks. A model no rule classifies is listed as `unranked` and never recommended. Claude Code prices now follow the rate card line for each model: Opus 5.5 ($4/$20) and Fable 5.1 (cache hits $0.25) no longer inherit Opus 5's or Fable 5's price, and `provider validate` refuses a price on a family.
|
|
33
|
+
|
|
34
|
+
- strategy: Claude Code and Codex model choices now come from bounded, no-prompt installed-CLI handshakes per account; Claude discovery suppresses hooks and preserves explicit 1M selectors, Codex pages `model/list` and records per-model reasoning support, and the static lists remain only as failure fallbacks.
|
|
35
|
+
|
|
36
|
+
- docs: a new README with a banner, the problem, a comparison with the
|
|
37
|
+
alternatives, and dashboard screenshots taken from a made-up demo home;
|
|
38
|
+
a new playbook guide; `scripts/render-readme-shots.sh` rebuilds the
|
|
39
|
+
screenshots and refuses to write them if private terms or source IDs
|
|
40
|
+
survive.
|
|
41
|
+
|
|
42
|
+
- stats: hyphenated project names keep their complete basename when the panel
|
|
43
|
+
has room instead of being reduced to the suffix after the final hyphen.
|
|
44
|
+
|
|
45
|
+
- stats: charts share Home's even whole-unit ticks, zero baseline and daily
|
|
46
|
+
weekday labels, while compact money cells use the dashboard-wide source and
|
|
47
|
+
lower-bound glyphs and return the saved width to pool and project names.
|
|
48
|
+
|
|
49
|
+
- mod: the Step pane follows, selects, and expands the newest turn by default,
|
|
50
|
+
refreshes an open running Step every four seconds, and ticks its projected
|
|
51
|
+
in-flight command name and elapsed time every second without refetching.
|
|
52
|
+
- step: `workflow action show --json` and `workflow task show --json` now emit
|
|
53
|
+
compact Step schema v2: one activity per distinct attempt, one canonical
|
|
54
|
+
event array inside each activity, reference markers for aliases, and the
|
|
55
|
+
unchanged display projection used by the dashboard and Claude mod.
|
|
56
|
+
- pools: per-home display labels (`bullswarm pools label`) now shorten pool ids
|
|
57
|
+
across human CLI output, progress, the dashboard, and the Claude Mod without
|
|
58
|
+
changing credential, routing, meter, history, or workflow-record keys. Pool
|
|
59
|
+
arguments accept either form; JSON retains `pool` and adds `poolLabel`.
|
|
60
|
+
|
|
61
|
+
## 0.35.4 — one Runs table for workflows and tasks
|
|
62
|
+
|
|
63
|
+
- mod: a live single task no longer replaces the selected workflow; each task
|
|
64
|
+
gets its own `task <8-char-tail>` pane button and a real dashboard Step model
|
|
65
|
+
from the read-only `workflow task show` command.
|
|
66
|
+
- runs: workflow and task rows share one set of columns — status, id, project,
|
|
67
|
+
what, kind (`7/7 steps` or `task`), active time, cost and start clock — sized
|
|
68
|
+
once for the whole page, so every day section lines up; long project names
|
|
69
|
+
stop at 18 cells with `…`.
|
|
70
|
+
- runs: a task says what it is about: its first Markdown heading, else a
|
|
71
|
+
labelled subject line (`Outcome:`, `Goal:`, `Task:`, `Defect:` …), else the
|
|
72
|
+
first sentence that is neither setup (`You are …`, `Workspace: …`) nor a
|
|
73
|
+
standing rule (`Edit only …`, `Do not commit.`, `Never …`).
|
|
74
|
+
- runs: tasks recorded before the task ledger show their start clock, their
|
|
75
|
+
project when the cwd was recorded, and a dim `build task on codex`-style
|
|
76
|
+
description instead of a row of dashes.
|
|
77
|
+
- runs: money is one cents cell: `$0.32` measured, `≈$4.01` or `~$25.87`
|
|
78
|
+
estimated, `≥$138.12` when some attempts are unpriced, and `~<$0.01` below
|
|
79
|
+
a cent; rows no longer print `API`, `summed`, `estimated`, `unmeasured` or
|
|
80
|
+
`span`. Day rules read `6 runs · 7 tasks · ≥$374.85 · 3 unpriced`, dropping
|
|
81
|
+
the money and then the unpriced count at narrow widths, never the words.
|
|
82
|
+
- runs: the description shrinks first, then start, kind and project drop as
|
|
83
|
+
the terminal narrows; status, id, time and cost always stay, and no line is
|
|
84
|
+
wider than the page from 40 to 260 columns.
|
|
85
|
+
- runs: the cursor bar is one even inverse band; the grey cost and clock cells
|
|
86
|
+
no longer show as darker blocks inside it.
|
|
87
|
+
|
|
3
88
|
## 0.35.3 — whole right edges in Ghostty and herdr, Enter opens the row you are on, named phases
|
|
4
89
|
|
|
5
90
|
- dashboard: every row is erased before it is painted instead of after, so a
|
|
@@ -145,7 +230,7 @@
|
|
|
145
230
|
sessions read-only; Command Code reads persisted project JSONL only when a
|
|
146
231
|
session transcript exists and carries usage.
|
|
147
232
|
- reprice: transcript lookup follows the loaded provider registry, so
|
|
148
|
-
`opencode2*`, `opencode2:
|
|
233
|
+
`opencode2*`, `opencode2:orbit-*`, and `command-code` pools reach their
|
|
149
234
|
provider-owned readers without a hard-coded provider list. Ambiguous,
|
|
150
235
|
missing, and checkpoint-only records stay unknown rather than becoming
|
|
151
236
|
zero-cost attempts.
|
|
@@ -153,7 +238,7 @@
|
|
|
153
238
|
`--no-session`, so their checkpoints are documented as non-recoverable
|
|
154
239
|
history.
|
|
155
240
|
- pricing: public model cards are retained only with a source and date (the
|
|
156
|
-
cards were checked 2026-09-20). `
|
|
241
|
+
cards were checked 2026-09-20). `orbit/*` and `opencode/union-alpha` relay
|
|
157
242
|
identifiers have no public card, so the underlying OpenAI card is not
|
|
158
243
|
substituted; observed Command Code models use the cited Command Code card.
|
|
159
244
|
- evidence: live OpenCode and Command Code streams captured on 2026-09-20 are
|
package/README.md
CHANGED
|
@@ -1,273 +1,185 @@
|
|
|
1
|
-
|
|
1
|
+
<p align="center">
|
|
2
|
+
<img src="docs/public/brand/bullswarm-banner.png" alt="Bullswarm — use every coding-agent plan you pay for" width="100%">
|
|
3
|
+
</p>
|
|
2
4
|
|
|
3
|
-
|
|
4
|
-
|
|
5
|
-
|
|
5
|
+
<p align="center">
|
|
6
|
+
<img alt="npm version" src="https://img.shields.io/npm/v/bullswarm">
|
|
7
|
+
<img alt="MIT license" src="https://img.shields.io/npm/l/bullswarm">
|
|
8
|
+
<img alt="Node.js 22.12 or later" src="https://img.shields.io/badge/node-%3E%3D22.12-339933?logo=node.js&logoColor=white">
|
|
9
|
+
</p>
|
|
6
10
|
|
|
7
|
-
|
|
8
|
-
across those same CLIs. Every result is judged by what was actually written,
|
|
9
|
-
not by whether the process exited 0.
|
|
11
|
+
<p align="center">Route work across your coding-agent subscriptions, spend quota before it expires, and verify what comes back.</p>
|
|
10
12
|
|
|
11
|
-
|
|
12
|
-
(`/bullswarm`, or `$bullswarm` where a skill uses that syntax) goes straight
|
|
13
|
-
to `bullswarm run` or `bullswarm workflow goal`. A skill here is a short
|
|
14
|
-
instruction file the agent CLI loads. There is no separate preview or
|
|
15
|
-
classifier command to learn first.
|
|
13
|
+
## The problem
|
|
16
14
|
|
|
17
|
-
|
|
15
|
+
You pay for more than one coding agent: a Claude plan or two, Codex, Grok, maybe another.
|
|
18
16
|
|
|
19
|
-
|
|
20
|
-
single coding agent's judgment on whether its own work is done should not be
|
|
21
|
-
the only check in the loop. Bullswarm picks whichever installed agent CLI has
|
|
22
|
-
the most unused quota right now, and treats every delegate's output as
|
|
23
|
-
evidence to be verified — never as an authority to be trusted on its word.
|
|
17
|
+
Every one of them meters you on a clock — five-hour windows, weekly caps, monthly allowances — and unused quota is simply gone when the window resets.
|
|
24
18
|
|
|
25
|
-
|
|
26
|
-
executing everything serially leaves every other installed CLI's quota idle,
|
|
27
|
-
and having that same agent be the sole judge of whether the whole goal is
|
|
28
|
-
done multiplies the risk instead of dividing it. Bullswarm's workflow engine
|
|
29
|
-
runs a graph of dependent actions across whichever pools have quota to spare
|
|
30
|
-
— a pool is one installed agent CLI, or one account of that CLI — and
|
|
31
|
-
computes completion from evidence the graph itself required, not from any one
|
|
32
|
-
delegate's own say-so.
|
|
19
|
+
So one plan runs dry mid-task while the others sit idle. You switch tools by hand just to spend what you already paid for. And the agent that wrote the code is usually the one that decides it is done.
|
|
33
20
|
|
|
34
|
-
##
|
|
21
|
+
## What Bullswarm does
|
|
35
22
|
|
|
36
|
-
|
|
37
|
-
|
|
38
|
-
|
|
23
|
+
Bullswarm turns the agent CLIs already signed in on your machine into **one fleet** that your main agent can command.
|
|
24
|
+
|
|
25
|
+
```mermaid
|
|
26
|
+
flowchart LR
|
|
27
|
+
you([You]) -->|goal| main["Your main agent<br/>Claude Code · Codex · Grok<br/>+ /bullswarm skill"]
|
|
28
|
+
main -->|"bullswarm run · workflow goal"| bs{{"Bullswarm<br/>pace router + workflow kernel"}}
|
|
29
|
+
bs --> c["claude -p"]
|
|
30
|
+
bs --> x["codex exec"]
|
|
31
|
+
bs --> g["grok -p"]
|
|
32
|
+
bs --> o["contributed providers"]
|
|
33
|
+
c & x & g & o -->|output| v["independent verification"]
|
|
34
|
+
v -->|"verdict + evidence"| main
|
|
39
35
|
```
|
|
40
36
|
|
|
41
|
-
|
|
42
|
-
|
|
43
|
-
|
|
44
|
-
|
|
45
|
-
|
|
46
|
-
|
|
47
|
-
|
|
48
|
-
|
|
49
|
-
|
|
50
|
-
|
|
51
|
-
|
|
52
|
-
`mods/bullswarm` is the same routing injected into Claude Code's own engine
|
|
53
|
-
as a Claude Mod (a plugin of TypeScript function hooks, behind
|
|
54
|
-
`CLAUDE_CODE_ENABLE_FUNCTION_HOOKS=1`): the pool meters drawn above the
|
|
55
|
-
prompt, the pools named in the model's context, Claude's general-purpose
|
|
56
|
-
subagents routed to whichever pool has surplus and answered with the
|
|
57
|
-
verified output, and the verdict appended to every `bullswarm run` the model
|
|
58
|
-
runs. The Mod is the dashboard's read-only counterpart: its pane shows Run,
|
|
59
|
-
Step, and the Usage/Pools view in the same meter colours. It does not expose
|
|
60
|
-
the dashboard's Home, Runs, Budget, Stats, Fleet, or Help pages, and
|
|
61
|
-
it has no edit or install action. See [mods/bullswarm/README.md](mods/bullswarm/README.md).
|
|
62
|
-
|
|
63
|
-
Three ways to load it:
|
|
37
|
+
- **Paces quota instead of guessing.** Each task goes to the eligible pool with the most unused quota relative to its reset clock. A pool that is furthest behind pace—or close to resetting with quota left—moves forward.
|
|
38
|
+
- **Routes by the work.** The `analyze`, `build`, and `chore` lanes derive an effort tier; setup maps those tiers to provider and model choices, while `bullswarm setup --wizard` also configures reasoning depth. Or ask your agent to run the non-interactive setup.
|
|
39
|
+
- **Runs one bounded task.** `bullswarm run` routes it to one agent, waits, and returns a verdict.
|
|
40
|
+
- **Executes multi-phase workflows.** Your main agent authors a dependency graph; Bullswarm schedules ready, file-disjoint actions across agents in parallel, then runs integration and independent acceptance when the plan calls for them.
|
|
41
|
+
- **Verifies content, not exit codes.** A delegate's output is evidence. A zero exit code is never enough by itself, and workflow completion stays separate from verified requirements.
|
|
42
|
+
- **Shows the whole system live.** The terminal dashboard has Home, Runs, Run, Step, Budget, Stats, Fleet, and Help views, from portfolio-level quota and history down to individual agent turns.
|
|
43
|
+
- **Reads licence meters.** Built-in readers cover Claude Code, Codex, and Grok; the provider interface extends routing, meters, models, and event streams without putting vendor quirks in the core.
|
|
44
|
+
- **Lets your agent delegate.** The packaged `/bullswarm` skill teaches Claude Code, Codex, and Grok when to use a single run or a workflow.
|
|
45
|
+
- **Lives inside Claude Code too.** The early-access Claude Code Mod adds Bullswarm routing, a run strip, and read-only Run, Step, and Usage panes inside Claude Code.
|
|
46
|
+
- **Changes course while work is live.** Newer steering and plan-revision commands can add, amend, remove, or rerun actions; pause and resume remain explicit.
|
|
64
47
|
|
|
65
|
-
|
|
66
|
-
# with the CLI: links the mod under ~/.claude/skills and sets the flag in ~/.claude/settings.json
|
|
67
|
-
bullswarm integrate install --agents claude --yes
|
|
48
|
+
## Why not just…
|
|
68
49
|
|
|
69
|
-
|
|
70
|
-
|
|
71
|
-
|
|
50
|
+
| Approach | What you give up |
|
|
51
|
+
|---|---|
|
|
52
|
+
| **One agent's built-in subagents or workflows** | Everything draws on that one plan's quota, the same vendor grades its own work, and your process is tied to that vendor's feature. |
|
|
53
|
+
| **An API gateway that re-exposes your subscriptions** | Turning a consumer subscription into a generic API endpoint can conflict with provider terms, and you lose each agent's own tools and harness. |
|
|
54
|
+
| **Switching tools by hand** | You become the scheduler, and the quota you did not get to still expires. |
|
|
55
|
+
| **Bullswarm** | Drives each vendor's own headless CLI—`claude -p`, `codex exec`, `grok -p`—the way those CLIs are meant to be scripted, with the accounts you already signed in. Work lands where quota is spare, and a different agent checks it. |
|
|
72
56
|
|
|
73
|
-
|
|
74
|
-
CLAUDE_CODE_ENABLE_FUNCTION_HOOKS=1 claude --plugin-dir "$(npm root -g)/bullswarm/mods/bullswarm"
|
|
75
|
-
```
|
|
57
|
+
Bullswarm never proxies a subscription as an API and never collects or shares your vendor credentials.
|
|
76
58
|
|
|
77
|
-
The
|
|
78
|
-
the shell or under `env` in `~/.claude/settings.json`, and the `bullswarm`
|
|
79
|
-
CLI on `PATH`. Pick one route: an installed marketplace copy takes precedence
|
|
80
|
-
over the skills-dir link, and Claude says so at startup.
|
|
59
|
+
The controller is portable. Claude Code, Codex, Grok, or any capable caller drives the same CLI and durable workflow kernel, so your workflows are not locked into one vendor's orchestration.
|
|
81
60
|
|
|
82
|
-
##
|
|
61
|
+
## Why now
|
|
62
|
+
|
|
63
|
+
Agent plans now come with hard weekly limits, and serious work spans hours rather than prompts. A workflow can run for one to two hours across several agents—builders in parallel, an integrator, then independent acceptance—while you keep handing new goals to your main agent.
|
|
64
|
+
|
|
65
|
+
With Bullswarm routing by pace, running several goals at once no longer means worrying about wasting one plan's scarce quota: the fleet spends whichever plan is furthest behind.
|
|
66
|
+
|
|
67
|
+
## See it
|
|
83
68
|
|
|
84
|
-
|
|
85
|
-
Use `bullswarm --setup` or `bullswarm setup` to open setup again, and use
|
|
86
|
-
`bullswarm workflow tui` when you want the explicit dashboard command. The
|
|
87
|
-
eight pages answer different questions:
|
|
69
|
+

|
|
88
70
|
|
|
89
|
-
|
|
71
|
+
*Home — today's work, budget position, trends, and recent runs in one view.*
|
|
72
|
+
|
|
73
|
+
| Runs | Run |
|
|
90
74
|
|---|---|
|
|
91
|
-
|
|
|
92
|
-
|
|
|
93
|
-
|
|
94
|
-
| Step |
|
|
95
|
-
| Budget | Pool quota windows, measured worker-minutes versus rest, measured or `≈` money, fit, and biggest workflows. |
|
|
96
|
-
| Stats | Runs, spend, worker-minutes, and verification over 7d/30d/all, by pool, model, and project. |
|
|
97
|
-
| Fleet | Lane/provider model and reasoning rungs, records, meter state, and the setup edit hand-off. |
|
|
98
|
-
| Help | Every key, click, layout rule, and dashboard command. |
|
|
99
|
-
|
|
100
|
-
The visual-fidelity pass keeps the same real numbers while composing and
|
|
101
|
-
colouring these pages like the approved prototype. At 55 columns every Home
|
|
102
|
-
today tile and per-step bar, every Stats model, project and its sparkline, every
|
|
103
|
-
Stats and Fleet pool row, and every run row in the Runs history table keeps one
|
|
104
|
-
row per item on the phone; Budget's per-pool block (a header over `used`,
|
|
105
|
-
`by bullswarm`, `room` and `so far`), Home's `budget · this week` pool (meter
|
|
106
|
-
plus a reset line, as the prototype draws it) and the Runs `active` entry are
|
|
107
|
-
the deliberate exceptions. At 200 columns every page composes to the frame
|
|
108
|
-
rather than capping at the 120-column composition. Estimated figures still
|
|
109
|
-
carry `≈` and their basis; a figure that cannot be measured is a line of words
|
|
110
|
-
saying so — no page draws an empty or dotted track for missing data.
|
|
111
|
-
|
|
112
|
-
Every page has a sticky header, a page tab row, and a sticky bottom nav. The
|
|
113
|
-
tab row is five tabs — Home, Runs, Budget, Stats, Fleet — with Run and Step
|
|
114
|
-
marking Runs, and Help shown only while it is open. The shared key table is:
|
|
115
|
-
|
|
116
|
-
| Key | Does |
|
|
75
|
+
| [](docs/public/screens/runs.png) | [](docs/public/screens/run.png) |
|
|
76
|
+
| Active and historical workflows and single tasks in one table. | Phases, live workers, costs, and the attempt timeline. |
|
|
77
|
+
|
|
78
|
+
| Step | Stats |
|
|
117
79
|
|---|---|
|
|
118
|
-
|
|
|
119
|
-
|
|
|
120
|
-
|
|
121
|
-
|
|
122
|
-
|
|
123
|
-
|
|
124
|
-
|
|
125
|
-
|
|
126
|
-
|
|
127
|
-
|
|
128
|
-
|
|
129
|
-
|
|
130
|
-
|
|
131
|
-
|
|
132
|
-
|
|
133
|
-
|
|
134
|
-
|
|
135
|
-
|
|
136
|
-
|
|
137
|
-
|
|
138
|
-
|
|
139
|
-
|
|
140
|
-
|
|
141
|
-
and
|
|
142
|
-
or controls, or use the wheel to scroll.
|
|
143
|
-
|
|
144
|
-
Home, Runs, Budget, and Stats read a per-run rollup that every finishing
|
|
145
|
-
run appends to `~/.bullswarm/history/runs.jsonl`. After upgrading, backfill
|
|
146
|
-
the runs that finished before 0.33.0 once with `bullswarm workflow reindex`;
|
|
147
|
-
legacy runs get a minimal record and only runs still in flight are skipped.
|
|
148
|
-
|
|
149
|
-
One bounded outcome — a task with a clear finish line:
|
|
80
|
+
| [](docs/public/screens/step.png) | [](docs/public/screens/stats.png) |
|
|
81
|
+
| The live or saved agent transcript, result, task, and cost evidence. | Workflow, spend, worker-time, and verification trends. |
|
|
82
|
+
|
|
83
|
+
[](docs/public/screens/budget.png)
|
|
84
|
+
|
|
85
|
+
*Budget — spend spare quota before each weekly or monthly window resets.*
|
|
86
|
+
|
|
87
|
+
<p align="center">
|
|
88
|
+
<img src="docs/public/screens/home-phone.png" alt="Bullswarm Home dashboard at phone width" width="360">
|
|
89
|
+
<img src="docs/public/screens/run-phone.png" alt="Bullswarm Run page at phone width" width="360">
|
|
90
|
+
</p>
|
|
91
|
+
|
|
92
|
+
*Home and Run retain their core evidence at a 55-column phone width.*
|
|
93
|
+
|
|
94
|
+
## Quick start
|
|
95
|
+
|
|
96
|
+
Requires Node.js 22.12 or later.
|
|
97
|
+
|
|
98
|
+
```bash
|
|
99
|
+
npm i -g bullswarm
|
|
100
|
+
bullswarm setup
|
|
101
|
+
```
|
|
102
|
+
|
|
103
|
+
`setup` discovers installed agent CLIs, shows their quota state, and opens the provider/model control centre. An agent or CI process can use discovered defaults without prompts:
|
|
150
104
|
|
|
151
105
|
```bash
|
|
152
|
-
bullswarm
|
|
106
|
+
bullswarm setup --yes --strategy --integrate
|
|
107
|
+
bullswarm doctor
|
|
153
108
|
```
|
|
154
109
|
|
|
155
|
-
|
|
156
|
-
|
|
157
|
-
|
|
158
|
-
|
|
159
|
-
|
|
160
|
-
|
|
161
|
-
Multi-step work, where you author the plan. The kernel — Bullswarm's own
|
|
162
|
-
runtime, not an agent — validates that program and executes it:
|
|
163
|
-
|
|
164
|
-
```json
|
|
165
|
-
{
|
|
166
|
-
"schemaVersion": "bullswarm.workflow.program.v2",
|
|
167
|
-
"actions": [
|
|
168
|
-
{ "id": "fix", "kind": "implement", "purpose": "Fix the failing tests",
|
|
169
|
-
"dependsOn": [], "ownedFiles": ["src/parser.js"], "affects": ["requirement-1"],
|
|
170
|
-
"evidenceFor": [], "prompt": "In ~/some-repo, fix the failing tests with the smallest correct change." }
|
|
171
|
-
]
|
|
172
|
-
}
|
|
110
|
+
Run one bounded task:
|
|
111
|
+
|
|
112
|
+
```bash
|
|
113
|
+
bullswarm run --lane analyze --add-dir . \
|
|
114
|
+
--prompt "List every TODO in src with file and line number." --json
|
|
173
115
|
```
|
|
174
116
|
|
|
117
|
+
Start a first workflow with an explicitly delegated planner:
|
|
118
|
+
|
|
175
119
|
```bash
|
|
176
|
-
|
|
177
|
-
|
|
178
|
-
--cwd
|
|
179
|
-
bullswarm workflow goal "Fix the failing tests and verify the change" \
|
|
180
|
-
--cwd ~/some-repo --program plan.json # launches, prints a short ID and observation commands, returns
|
|
120
|
+
bullswarm workflow goal \
|
|
121
|
+
"Audit this repository and write a one-page summary" \
|
|
122
|
+
--cwd . --orchestrator auto --watch
|
|
181
123
|
```
|
|
182
124
|
|
|
183
|
-
|
|
184
|
-
default; add `--watch` to follow its low-noise progress in the same terminal
|
|
185
|
-
instead.
|
|
186
|
-
|
|
187
|
-
## How it picks a pool
|
|
188
|
-
|
|
189
|
-
- Work is tagged by lane — read-only analysis, ordinary build work, or
|
|
190
|
-
mechanical chores — not assigned to a fixed pool ahead of time.
|
|
191
|
-
- Among the pools that can do the work, the one furthest behind its own quota
|
|
192
|
-
pace (the most unspent surplus) wins, so quota doesn't expire unused.
|
|
193
|
-
- A pool close to its rolling 5-hour usage ceiling gives way to one with
|
|
194
|
-
headroom only while another pool is actually behind its own pace — it is an
|
|
195
|
-
ordering penalty, not a cutoff. A step may run a pool all the way to 100% of
|
|
196
|
-
that window, because if the provider stops the worker at the wall the retry
|
|
197
|
-
is briefed on what it had already written.
|
|
198
|
-
- A pool whose weekly or monthly subscription window is about to reset gets
|
|
199
|
-
priority for its remaining surplus, so quota doesn't run out the clock
|
|
200
|
-
unspent.
|
|
201
|
-
- A pool already busy with other in-flight work yields to a quieter pool at a
|
|
202
|
-
similar pace, so a burst of parallel work spreads out instead of piling onto
|
|
203
|
-
one pool.
|
|
204
|
-
- A pool that reports a usage-limit error is benched until the provider's own
|
|
205
|
-
reset time and automatically re-tried after that — never left down for good,
|
|
206
|
-
and never retried early.
|
|
207
|
-
|
|
208
|
-
The full mechanics behind each of these are in
|
|
209
|
-
[Routing](https://bulls-work.github.io/bullswarm/guide/routing).
|
|
210
|
-
|
|
211
|
-
## What you get back
|
|
212
|
-
|
|
213
|
-
`bullswarm run` prints a JSON verdict when it finishes:
|
|
214
|
-
|
|
215
|
-
- `keepOnClaude: true` — the router says do this in-session; nothing ran
|
|
216
|
-
- `ok: true` (and `keepOnClaude` is false) — the output passed verification;
|
|
217
|
-
read `outFile`
|
|
218
|
-
- `ok: false` — `why` names the gate that failed
|
|
219
|
-
- `contentUsableDespiteExit: true` — the process exited non-zero but the
|
|
220
|
-
content still verified; read it before re-running
|
|
221
|
-
|
|
222
|
-
A non-zero exit from the delegate is never treated as success on its own. See
|
|
223
|
-
[Result envelope](https://bulls-work.github.io/bullswarm/reference/result) for the full
|
|
224
|
-
verdict shape.
|
|
225
|
-
|
|
226
|
-
A workflow produces a durable, versioned result envelope — a JSON document
|
|
227
|
-
with `runId` / `shortId`, status, per-requirement evidence, per-action
|
|
228
|
-
outcomes, and usage — instead of leaving you to parse a transcript.
|
|
125
|
+
For normal use, your main agent should author `plan.json`, validate it, and launch the exact same goal:
|
|
229
126
|
|
|
230
127
|
```bash
|
|
231
|
-
bullswarm workflow
|
|
232
|
-
|
|
128
|
+
bullswarm workflow plan contract \
|
|
129
|
+
"1. Fix the parser. 2. Add independent verification." --cwd . --json
|
|
130
|
+
|
|
131
|
+
bullswarm workflow plan validate \
|
|
132
|
+
"1. Fix the parser. 2. Add independent verification." \
|
|
133
|
+
--cwd . --program plan.json --json
|
|
134
|
+
|
|
135
|
+
bullswarm workflow goal \
|
|
136
|
+
"1. Fix the parser. 2. Add independent verification." \
|
|
137
|
+
--cwd . --program plan.json --watch
|
|
233
138
|
```
|
|
234
139
|
|
|
235
|
-
|
|
236
|
-
the exact flags to keep polling; `runs result --summary` is what to read once
|
|
237
|
-
a run finishes, and `runs result --json` (no `--summary`) gives the full
|
|
238
|
-
envelope for a failed or partial run. See
|
|
239
|
-
[Observing runs](https://bulls-work.github.io/bullswarm/guide/observing) for the
|
|
240
|
-
full shape of both.
|
|
140
|
+
Open the dashboard at any time. Quitting it does not stop running workflows.
|
|
241
141
|
|
|
242
|
-
|
|
142
|
+
```bash
|
|
143
|
+
bullswarm
|
|
144
|
+
# explicit form:
|
|
145
|
+
bullswarm workflow tui
|
|
146
|
+
```
|
|
243
147
|
|
|
244
|
-
|
|
245
|
-
[bulls-work.github.io/bullswarm](https://bulls-work.github.io/bullswarm/).
|
|
148
|
+
## A practical playbook
|
|
246
149
|
|
|
247
|
-
|
|
248
|
-
|
|
249
|
-
|
|
250
|
-
|
|
150
|
+
1. **Nail down the outcome.** State what must change, what must remain untouched, and what evidence will count as done.
|
|
151
|
+
2. **Hand it to your main agent.** With the `/bullswarm` skill installed, it chooses a bounded run or authors a workflow program with clear territories, dependencies, integration, and acceptance.
|
|
152
|
+
3. **Let the graph fan out.** Ready, file-disjoint actions can run across Codex, Grok, Claude accounts, and contributed providers while routing spends the quota furthest behind pace.
|
|
153
|
+
4. **Stay in control.** Follow the dashboard or `workflow watch`; send guidance, pause, or revise the live plan when the goal changes.
|
|
154
|
+
5. **Sign off on evidence.** Read the durable result envelope and requirement evidence. `completed` and `verified` answer different questions.
|
|
251
155
|
|
|
252
|
-
|
|
253
|
-
|
|
254
|
-
|
|
255
|
-
|
|
256
|
-
|
|
|
257
|
-
|
|
258
|
-
|
|
|
259
|
-
|
|
|
260
|
-
|
|
|
261
|
-
|
|
|
262
|
-
|
|
|
263
|
-
|
|
264
|
-
|
|
265
|
-
|
|
266
|
-
|
|
267
|
-
|
|
268
|
-
|
|
269
|
-
|
|
270
|
-
|
|
271
|
-
|
|
272
|
-
|
|
273
|
-
|
|
156
|
+
You can keep several independent goals running without manually balancing every plan. Bullswarm accounts for in-flight load, while each workflow keeps its own durable state, outputs, events, and result.
|
|
157
|
+
|
|
158
|
+
## Supported agents and meters
|
|
159
|
+
|
|
160
|
+
| Agent CLI | Provider status | Subscription meter declared by the shipped connector | Headless entry point |
|
|
161
|
+
|---|---|---|---|
|
|
162
|
+
| Claude Code | Built in | weekly + 5-hour | `claude -p` |
|
|
163
|
+
| Codex | Built in | weekly | `codex exec` |
|
|
164
|
+
| Grok | Built in | weekly | `grok -p` |
|
|
165
|
+
| OpenCode | Contributed | none in the base connector | `opencode run --auto` |
|
|
166
|
+
| Command Code | Contributed | weekly + monthly + 5-hour | `command-code -p` |
|
|
167
|
+
|
|
168
|
+
Providers declare their own spawn command, meter reader, model discovery, reasoning levels, event decoding, and capabilities. See [Adding a provider](https://bulls-work.github.io/bullswarm/reference/providers).
|
|
169
|
+
|
|
170
|
+
## Status
|
|
171
|
+
|
|
172
|
+
Bullswarm is used daily. The maintainer's local records contained **335 workflow runs** and **507 single tasks** as of September 2026.
|
|
173
|
+
|
|
174
|
+
The routing, content verification, durable workflow kernel, dashboard, and built-in providers are established parts of the project. Mid-run steering and whole-plan revision are newer; use their validation and revision guards, and inspect the resulting evidence.
|
|
175
|
+
|
|
176
|
+
## Learn more
|
|
177
|
+
|
|
178
|
+
- [Documentation](https://bulls-work.github.io/bullswarm/)
|
|
179
|
+
- [Getting started](https://bulls-work.github.io/bullswarm/guide/getting-started)
|
|
180
|
+
- [How routing works](https://bulls-work.github.io/bullswarm/guide/routing)
|
|
181
|
+
- [Authoring workflows](https://bulls-work.github.io/bullswarm/guide/workflows)
|
|
182
|
+
- [Observing runs and the dashboard](https://bulls-work.github.io/bullswarm/guide/observing)
|
|
183
|
+
- [Provider reference](https://bulls-work.github.io/bullswarm/reference/providers)
|
|
184
|
+
- Contributing: [open an issue](https://github.com/Bulls-Work/bullswarm/issues) or [submit a pull request](https://github.com/Bulls-Work/bullswarm/pulls)
|
|
185
|
+
- [MIT licence](LICENSE)
|