bullswarm 0.35.4 → 0.35.6
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/AGENTS.md +3 -2
- package/CHANGELOG.md +84 -2
- package/README.md +104 -257
- package/data/openrouter-benchmarks.json +13349 -12793
- package/docs/design/providers-0.35.2/CHANGELOG-draft.md +2 -2
- package/docs/design/providers-0.35.2/README.md +6 -6
- package/docs/design/redesign-mechanics-principles-options.md +385 -0
- package/docs/design/stats-frames-0.33.2/spending-120.txt +18 -16
- package/docs/design/stats-frames-0.33.2/spending-200.txt +24 -0
- package/docs/design/stats-frames-0.33.2/spending-55.txt +32 -31
- package/docs/design/step-page-0.35.0/frames/failed-120.txt +2 -2
- package/docs/design/step-page-0.35.0/frames/failed-200.txt +2 -2
- package/docs/design/step-page-0.35.0/frames/failed-55.txt +2 -2
- package/docs/design/step-page-0.35.0/frames/real-failed-120.txt +3 -3
- package/docs/design/step-page-0.35.0/frames/real-failed-200.txt +4 -4
- package/docs/design/step-page-0.35.0/frames/real-failed-55.txt +3 -3
- package/docs/design/step-page-0.35.0/frames/rendered-failed-120.txt +3 -3
- package/docs/design/step-page-0.35.0/frames/rendered-failed-200.txt +3 -3
- package/docs/design/step-page-0.35.0/frames/rendered-failed-55.txt +3 -3
- package/docs/design/tidy-0.35.1/frames/real-stats-model-120.txt +13 -13
- package/docs/design/tidy-0.35.1/frames/real-stats-model-200.txt +13 -13
- package/docs/design/tidy-0.35.1/frames/real-stats-model-55.txt +9 -9
- package/docs/design/tidy-0.35.1/frames/real-stats-spending-120.txt +14 -14
- package/docs/design/tidy-0.35.1/frames/real-stats-spending-200.txt +14 -14
- package/docs/design/tidy-0.35.1/frames/real-stats-spending-55.txt +18 -18
- package/docs/design/tidy-0.35.1/frames/rendered-failed-120.txt +1 -1
- package/docs/design/tidy-0.35.1/frames/rendered-failed-200.txt +1 -1
- package/docs/design/tidy-0.35.1/frames/rendered-failed-55.txt +1 -1
- package/docs/guide/cost.md +1 -1
- package/docs/guide/gallery.md +57 -0
- package/docs/guide/getting-started.md +62 -16
- package/docs/guide/index.md +44 -6
- package/docs/guide/observing.md +42 -0
- package/docs/guide/playbook.md +120 -0
- package/docs/guide/routing.md +39 -1
- package/docs/index.md +50 -17
- package/docs/integrations/claude-code.md +4 -2
- package/docs/notes/index.md +1 -1
- package/docs/plans/dashboard-0.33.md +1 -1
- package/docs/plans/providers-0.35.2.goal.txt +2 -2
- package/docs/plans/providers-0.35.2.program.json +3 -3
- package/docs/reference/cli.md +41 -26
- package/docs/reference/configuration.md +24 -7
- package/docs/reference/providers.md +103 -4
- package/docs/studies/cost-audit-2026-09-18/claude-actual-vs-recorded.md +1 -1
- package/mods/bullswarm/README.md +7 -6
- package/mods/bullswarm/hooks/names.ts +7 -0
- package/mods/bullswarm/hooks/pane.tsx +13 -0
- package/mods/bullswarm/hooks/pools.ts +2 -0
- package/mods/bullswarm/hooks/register.ts +72 -14
- package/mods/bullswarm/hooks/runs.ts +1 -0
- package/mods/bullswarm/hooks/step.ts +65 -14
- package/mods/bullswarm/types/index.d.ts +4 -0
- package/package.json +4 -2
- package/providers/contrib/command-code/connector-history.json +288 -0
- package/providers/contrib/opencode/connector-history.json +241 -0
- package/providers/contrib/opencode/connector.json +1 -1
- package/skill/references/operations.md +8 -3
- package/src/cli.js +112 -44
- package/src/help.js +69 -50
- package/src/home-cli.js +1 -1
- package/src/lib/cli-flags.js +3 -3
- package/src/lib/config.js +2 -0
- package/src/lib/connector-copies.js +288 -0
- package/src/lib/model-family.js +252 -0
- package/src/lib/pool-labels.js +109 -0
- package/src/lib/providers.js +2 -0
- package/src/lib/reasoning.js +65 -7
- package/src/lib/strategy.js +628 -120
- package/src/provider-cli.js +63 -6
- package/src/providers/_schema.json +20 -1
- package/src/providers/claude-code/connector-history.json +330 -0
- package/src/providers/claude-code/connector.json +20 -5
- package/src/providers/claude-code/provider.mjs +124 -1
- package/src/providers/codex/connector-history.json +326 -0
- package/src/providers/codex/connector.json +53 -16
- package/src/providers/codex/provider.mjs +115 -0
- package/src/providers/echo/connector-history.json +142 -0
- package/src/providers/grok/connector-history.json +308 -0
- package/src/providers/grok/connector.json +65 -7
- package/src/setup.js +54 -58
- package/src/strategy-cli.js +277 -41
- package/src/strategy-dashboard.js +255 -85
- package/src/workflow/cli.js +11 -4
- package/src/workflow/dash-kit.js +55 -48
- package/src/workflow/dashboard.js +24 -1
- package/src/workflow/home-view.js +14 -56
- package/src/workflow/reprice.js +4 -2
- package/src/workflow/runs-cli.js +2 -1
- package/src/workflow/stat-kit.js +118 -20
- package/src/workflow/stats-model.js +1 -0
- package/src/workflow/stats-view.js +27 -32
- package/src/workflow/step-json.js +93 -0
- package/src/workflow/step-model.js +15 -6
- package/src/workflow/watch-cli.js +5 -0
- package/docs/public/favicon.svg +0 -7
- package/src/lib/release.js +0 -56
package/AGENTS.md
CHANGED
|
@@ -69,7 +69,7 @@ is the canonical reference for the CLI surface.
|
|
|
69
69
|
- Every verb must work non-interactively (no TTY). The interactive wizard is
|
|
70
70
|
a human convenience, never a requirement.
|
|
71
71
|
- Version single source: package.json. Release via
|
|
72
|
-
`
|
|
72
|
+
`npm run release -- patch|minor|major [--title "<headline>"]`, then `git push` and
|
|
73
73
|
`git push --tags`
|
|
74
74
|
— CI publishes through npm trusted publishing (OIDC), no tokens.
|
|
75
75
|
|
|
@@ -90,7 +90,8 @@ the authoring method is `skill/references/providers.md`.
|
|
|
90
90
|
## Releasing
|
|
91
91
|
|
|
92
92
|
1. All tests green.
|
|
93
|
-
2. `
|
|
93
|
+
2. `npm run release -- patch --title "<headline>"` (dates `## Unreleased`
|
|
94
|
+
in CHANGELOG.md, bumps package.json, creates commit + tag v*).
|
|
94
95
|
3. `git push && git push --tags`.
|
|
95
96
|
4. GitHub Actions publishes to npm via trusted publishing; verify with
|
|
96
97
|
`npm view bullswarm version`.
|
package/CHANGELOG.md
CHANGED
|
@@ -1,5 +1,87 @@
|
|
|
1
1
|
# bullswarm changelog
|
|
2
2
|
|
|
3
|
+
## Unreleased
|
|
4
|
+
|
|
5
|
+
## 0.35.6 — set reasoning per tier and per model from the setup screen
|
|
6
|
+
|
|
7
|
+
- setup: the setup screen (`bullswarm setup`, `strategy tui`) can now set
|
|
8
|
+
reasoning. Opening a provider shows each effort tier's model and reasoning
|
|
9
|
+
level with who chose it; ←/→ on a tier steps auto → the levels that CLI
|
|
10
|
+
accepts → "CLI decides", and saves it for that provider as
|
|
11
|
+
`strategy set-rung --reasoning` would. Provider names are no longer cut off,
|
|
12
|
+
and model search now starts with `/`, so typing `f` or `q` in a search no
|
|
13
|
+
longer finishes setup. `strategy inventory --json` lists each provider's
|
|
14
|
+
`reasoningLevels`.
|
|
15
|
+
|
|
16
|
+
- strategy: reasoning can be set per model. When a tier has more than one
|
|
17
|
+
model on a provider, each can think at its own level:
|
|
18
|
+
`strategy set-reasoning --tier high --level max --pool claude-code --model
|
|
19
|
+
claude-fable-5-1 --yes`, `reasoning.models` in `strategy configure`, or the
|
|
20
|
+
indented model rows on the setup screen. It sits above the provider's level
|
|
21
|
+
and below `run --reasoning`; resolved levels report source
|
|
22
|
+
`strategy-model`, and each model in `strategy inventory --json` carries the
|
|
23
|
+
level it would run at.
|
|
24
|
+
|
|
25
|
+
- docs: the README keeps two screenshots and a plainer tone, and its quick
|
|
26
|
+
start is now a prompt you paste into your agent; the other dashboard screens
|
|
27
|
+
moved to a new gallery page on the docs site.
|
|
28
|
+
|
|
29
|
+
## 0.35.5 — model choices come from each CLI, newest version wins, setup stops pinning tiers
|
|
30
|
+
|
|
31
|
+
- cli: the maintainer-only `release` verb is gone from the CLI and its help;
|
|
32
|
+
maintainers run `npm run release -- patch|minor|major`, which also dates
|
|
33
|
+
the `## Unreleased` changelog section.
|
|
34
|
+
|
|
35
|
+
- package: the README screenshots and brand images (`docs/public/`) are no
|
|
36
|
+
longer shipped in the npm package; the npm page loads them from GitHub.
|
|
37
|
+
|
|
38
|
+
- grok: every tier now suggests the newest Grok, and the tiers differ only by reasoning. This is an owner decision: Grok has one model line. One family now covers every plain `grok-N.M`, and the newest version wins. A fresh home suggests `grok-4.7` on all three tiers: high at the connector's `xhigh`, then `medium: grok-4.7 · high reasoning` and `low: grok-4.7 · medium reasoning`, both with the reason `one Grok line, lighter reasoning for lighter tiers`. Before, it suggested `grok-4.6` for high, `grok-4.5` for medium, and nothing for low, and it listed `grok-4.7` as unranked. A `grok-4.8` takes over every tier the day `grok models` lists it. `grok-4.7-build-fast` is the same model at twice the price, so it keeps its base model's rank but is never suggested; `strategy set-rung` still selects it. The tier-wide best pick, the Codex and Claude suggestions, and pace-based routing across pools do not change. grok-4.7 and grok-4.5 now have dated price rows from [xAI's models page](https://docs.x.ai/docs/models), and grok-4.6's row was re-checked there (all on 2026-09-23). build-fast stays unpriced: xAI's Pricing page lists its long-context rates as $6.00 / $1.50 / $18.00 per 1M tokens, which is not twice grok-4.7's. The fallback list (`knownModels`) now matches what `grok models` lists.
|
|
39
|
+
|
|
40
|
+
- strategy: `generationFallback` now also covers a tier that no family serves. If a connector opts such a tier in and no exact row ranks a model for it, the tier takes the best-ranked family's newest-generation model at the declared reasoning. That stand-in is the pool's rung, and it stands at rank 0 against other pools, so it is the tier-wide pick only when no other pool serves the tier. A tier may also declare `why`, a plain reason that replaces the generated one; `provider validate` rejects an empty one.
|
|
41
|
+
|
|
42
|
+
- strategy: applying recommendations no longer pins tiers. `strategy apply`, `refresh --apply`, `setup --yes --strategy`, the setup wizard, the TUI's apply key, and the daily auto-refresh now write only each pool's rungs (its model and reasoning per tier). Before, they also pinned every tier to one pool: a fresh `setup --yes --strategy` sent high, medium, and low to Codex even at 100% used, and the auto-refresh pinned them again every 24 hours, so pace-based routing never picked another plan. Now every dispatch picks its pool by spare quota, and `strategy routes` shows each tier moving between pools. `strategy.assignments` stays empty unless you pin a tier yourself. The apply JSON now reports `rungs`, `bestNow` (the tier-wide pick, for display), `unpinned`, and `keptPins`, in place of `applied`.
|
|
43
|
+
|
|
44
|
+
- strategy: pins are explicit only. `strategy assign` makes one, marked `source: "user"`, and `strategy clear-assignment` removes it. Apply and the auto-refresh never remove a pin you set, and apply now keeps the pin's model on that pool's rung, so the pinned tier runs the model you named instead of the pool's recommended one.
|
|
45
|
+
|
|
46
|
+
- strategy: if an earlier setup or apply pinned your tiers, the next `strategy apply` or automatic refresh removes those pins and lists them under `unpinned`. A pin is treated as written by apply when it still has the pool and model that apply recorded in the cached report. Pins you set or changed yourself stay (`keptPins`), as does any pin in a home where apply never ran. Run `bullswarm strategy show` to see which pins remain and what removes each one.
|
|
47
|
+
|
|
48
|
+
- strategy: the tier suggestion reads as a suggestion, not a pin: `high: codex/gpt-6-astra (best now; routing picks by spare quota)` in `strategy show`, `best now (not pins): …` after the wizard and `setup --yes --strategy`, and `Routing now · by spare quota at each dispatch, unless pinned` in the TUI, whose route lines end in `by spare quota` or `pinned to <pool>`. A pinned tier gets its own line in `strategy show`, `pinned to <pool>/<model> by you · high dispatches go there while it is available · strategy clear-assignment high removes it`. `strategy routes` gives each route a `pin` field.
|
|
49
|
+
|
|
50
|
+
- strategy: `strategy show` counts unranked models instead of listing them. It prints one line, such as `unranked: 312 models (command-code 180, opencode 130, grok 2) · never recommended · strategy show --json lists them`, and the setup review screen prints the same count for each provider. `strategy show --json` keeps the full `unranked` list, and the model pickers keep their `(unranked)` tag.
|
|
51
|
+
|
|
52
|
+
- strategy: medium now falls back to the newest generation. When the family that serves a tier has no model in a pool's newest generation, a connector that opts the tier in (`generationFallback`) gives it the next-lower family's newest-generation model, at a declared reasoning level clamped to what that model supports. Codex opts medium in, and this is an owner decision, not a benchmark: 168 medium dispatches as luna at max reasoning ran 99% ok. So a fresh home suggests `medium: gpt-6-luna · max reasoning — no gpt-6 terra yet, newest generation preferred`, and medium returns to terra at the normal medium reasoning once a gpt-6 terra is discovered. High, low, and Claude are unchanged. `strategy apply`, `refresh --apply`, `setup --yes --strategy`, the wizard, and the TUI all write that level into the pool's rung, reported as `max (recommendation)`. It is never written over a level you set, and a later apply removes it when the fallback ends. `strategy show` and the setup screens say why. The cross-pool tier pick also skips a model disabled for its pool now, as each pool's own pick already did, so a tier suggestion can no longer name a model you turned off.
|
|
53
|
+
|
|
54
|
+
- setup, doctor: old copies of packaged connectors in `<home>/connectors/` now follow the package. Those copies were made by installs older than 0.29.0. A copy of a pool the package defines was already ignored, so its old prices (such as Opus 5's on Opus 5.5) never applied; a contrib copy whose provider is not enabled was still used. Every verb now checks each copy against `connector-history.json`, which holds fingerprints of every value each field has shipped with. A copy with no edits moves to `connectors/retired/` so the packaged connector loads, and a contrib provider is enabled first when only the copy defined its pool. A copy you edited is kept, and `bullswarm doctor` (`connector-copies`, a `!` warning) and `strategy show` name its stale fields. The dead copy-to-home step in `setup` is gone.
|
|
55
|
+
|
|
56
|
+
- strategy: connectors now declare model families (`modelFamilies`: a name pattern with a tier and base rank), and the version read from the model id orders each family. So a newly discovered model is ranked on the day it appears: gpt-6-sol and gpt-6-luna are no longer tierless, and a fresh home suggests gpt-6-luna for low instead of gpt-5.6-luna. Inside a pool and family, the newer version always wins, whatever benchmarks, prices, or pace say. Quality is compared on one scale, the rank, and benchmarks only break equal ranks. A model no rule classifies is listed as `unranked` and never recommended. Claude Code prices now follow the rate card line for each model: Opus 5.5 ($4/$20) and Fable 5.1 (cache hits $0.25) no longer inherit Opus 5's or Fable 5's price, and `provider validate` refuses a price on a family.
|
|
57
|
+
|
|
58
|
+
- strategy: Claude Code and Codex model choices now come from bounded, no-prompt installed-CLI handshakes per account; Claude discovery suppresses hooks and preserves explicit 1M selectors, Codex pages `model/list` and records per-model reasoning support, and the static lists remain only as failure fallbacks.
|
|
59
|
+
|
|
60
|
+
- docs: a new README with a banner, the problem, a comparison with the
|
|
61
|
+
alternatives, and dashboard screenshots taken from a made-up demo home;
|
|
62
|
+
a new playbook guide; `scripts/render-readme-shots.sh` rebuilds the
|
|
63
|
+
screenshots and refuses to write them if private terms or source IDs
|
|
64
|
+
survive.
|
|
65
|
+
|
|
66
|
+
- stats: hyphenated project names keep their complete basename when the panel
|
|
67
|
+
has room instead of being reduced to the suffix after the final hyphen.
|
|
68
|
+
|
|
69
|
+
- stats: charts share Home's even whole-unit ticks, zero baseline and daily
|
|
70
|
+
weekday labels, while compact money cells use the dashboard-wide source and
|
|
71
|
+
lower-bound glyphs and return the saved width to pool and project names.
|
|
72
|
+
|
|
73
|
+
- mod: the Step pane follows, selects, and expands the newest turn by default,
|
|
74
|
+
refreshes an open running Step every four seconds, and ticks its projected
|
|
75
|
+
in-flight command name and elapsed time every second without refetching.
|
|
76
|
+
- step: `workflow action show --json` and `workflow task show --json` now emit
|
|
77
|
+
compact Step schema v2: one activity per distinct attempt, one canonical
|
|
78
|
+
event array inside each activity, reference markers for aliases, and the
|
|
79
|
+
unchanged display projection used by the dashboard and Claude mod.
|
|
80
|
+
- pools: per-home display labels (`bullswarm pools label`) now shorten pool ids
|
|
81
|
+
across human CLI output, progress, the dashboard, and the Claude Mod without
|
|
82
|
+
changing credential, routing, meter, history, or workflow-record keys. Pool
|
|
83
|
+
arguments accept either form; JSON retains `pool` and adds `poolLabel`.
|
|
84
|
+
|
|
3
85
|
## 0.35.4 — one Runs table for workflows and tasks
|
|
4
86
|
|
|
5
87
|
- mod: a live single task no longer replaces the selected workflow; each task
|
|
@@ -172,7 +254,7 @@
|
|
|
172
254
|
sessions read-only; Command Code reads persisted project JSONL only when a
|
|
173
255
|
session transcript exists and carries usage.
|
|
174
256
|
- reprice: transcript lookup follows the loaded provider registry, so
|
|
175
|
-
`opencode2*`, `opencode2:
|
|
257
|
+
`opencode2*`, `opencode2:orbit-*`, and `command-code` pools reach their
|
|
176
258
|
provider-owned readers without a hard-coded provider list. Ambiguous,
|
|
177
259
|
missing, and checkpoint-only records stay unknown rather than becoming
|
|
178
260
|
zero-cost attempts.
|
|
@@ -180,7 +262,7 @@
|
|
|
180
262
|
`--no-session`, so their checkpoints are documented as non-recoverable
|
|
181
263
|
history.
|
|
182
264
|
- pricing: public model cards are retained only with a source and date (the
|
|
183
|
-
cards were checked 2026-09-20). `
|
|
265
|
+
cards were checked 2026-09-20). `orbit/*` and `opencode/union-alpha` relay
|
|
184
266
|
identifiers have no public card, so the underlying OpenAI card is not
|
|
185
267
|
substituted; observed Command Code models use the cited Command Code card.
|
|
186
268
|
- evidence: live OpenCode and Command Code streams captured on 2026-09-20 are
|
package/README.md
CHANGED
|
@@ -1,273 +1,120 @@
|
|
|
1
|
-
|
|
2
|
-
|
|
3
|
-
|
|
4
|
-
|
|
5
|
-
|
|
6
|
-
|
|
7
|
-
|
|
8
|
-
|
|
9
|
-
|
|
10
|
-
|
|
11
|
-
|
|
12
|
-
|
|
13
|
-
|
|
14
|
-
|
|
15
|
-
|
|
16
|
-
|
|
17
|
-
|
|
18
|
-
|
|
19
|
-
|
|
20
|
-
|
|
21
|
-
|
|
22
|
-
the most unused quota right now, and treats every delegate's output as
|
|
23
|
-
evidence to be verified — never as an authority to be trusted on its word.
|
|
24
|
-
|
|
25
|
-
The same problem compounds on multi-step goals: one agent planning and
|
|
26
|
-
executing everything serially leaves every other installed CLI's quota idle,
|
|
27
|
-
and having that same agent be the sole judge of whether the whole goal is
|
|
28
|
-
done multiplies the risk instead of dividing it. Bullswarm's workflow engine
|
|
29
|
-
runs a graph of dependent actions across whichever pools have quota to spare
|
|
30
|
-
— a pool is one installed agent CLI, or one account of that CLI — and
|
|
31
|
-
computes completion from evidence the graph itself required, not from any one
|
|
32
|
-
delegate's own say-so.
|
|
33
|
-
|
|
34
|
-
## Install
|
|
35
|
-
|
|
36
|
-
```bash
|
|
37
|
-
npm i -g bullswarm
|
|
38
|
-
bullswarm setup
|
|
39
|
-
```
|
|
1
|
+
<p align="center">
|
|
2
|
+
<img src="docs/public/brand/bullswarm-banner.jpg" alt="Bullswarm: a fleet of bull agents in tuxedos wearing Claude, OpenAI and Grok pins" width="100%">
|
|
3
|
+
</p>
|
|
4
|
+
|
|
5
|
+
<p align="center">
|
|
6
|
+
<a href="https://www.npmjs.com/package/bullswarm"><img alt="npm version" src="https://img.shields.io/npm/v/bullswarm"></a>
|
|
7
|
+
<a href="LICENSE"><img alt="MIT license" src="https://img.shields.io/npm/l/bullswarm"></a>
|
|
8
|
+
<a href="https://nodejs.org/"><img alt="Node.js 22.12 or later" src="https://img.shields.io/badge/node-%3E%3D22.12-339933?logo=node.js&logoColor=white"></a>
|
|
9
|
+
</p>
|
|
10
|
+
|
|
11
|
+
<p align="center">Route work across your coding-agent subscriptions, spend quota before it expires, and verify what comes back.</p>
|
|
12
|
+
|
|
13
|
+
## The problem
|
|
14
|
+
|
|
15
|
+
You pay for more than one coding agent: a Claude plan or two, Codex, Grok, maybe another.
|
|
16
|
+
|
|
17
|
+
Every one of them meters you on a clock — five-hour windows, weekly caps, monthly allowances — and unused quota is simply gone when the window resets.
|
|
18
|
+
|
|
19
|
+
So one plan runs dry mid-task while the others sit idle. You switch tools by hand just to spend what you already paid for. And the agent that wrote the code is usually the one that decides it is done.
|
|
20
|
+
|
|
21
|
+
## What Bullswarm does
|
|
40
22
|
|
|
41
|
-
|
|
42
|
-
|
|
43
|
-
|
|
44
|
-
|
|
45
|
-
|
|
46
|
-
|
|
47
|
-
|
|
48
|
-
|
|
49
|
-
|
|
50
|
-
|
|
51
|
-
|
|
52
|
-
|
|
53
|
-
as a Claude Mod (a plugin of TypeScript function hooks, behind
|
|
54
|
-
`CLAUDE_CODE_ENABLE_FUNCTION_HOOKS=1`): the pool meters drawn above the
|
|
55
|
-
prompt, the pools named in the model's context, Claude's general-purpose
|
|
56
|
-
subagents routed to whichever pool has surplus and answered with the
|
|
57
|
-
verified output, and the verdict appended to every `bullswarm run` the model
|
|
58
|
-
runs. The Mod is the dashboard's read-only counterpart: its pane shows Run,
|
|
59
|
-
Step, and the Usage/Pools view in the same meter colours. It does not expose
|
|
60
|
-
the dashboard's Home, Runs, Budget, Stats, Fleet, or Help pages, and
|
|
61
|
-
it has no edit or install action. See [mods/bullswarm/README.md](mods/bullswarm/README.md).
|
|
62
|
-
|
|
63
|
-
Three ways to load it:
|
|
64
|
-
|
|
65
|
-
```bash
|
|
66
|
-
# with the CLI: links the mod under ~/.claude/skills and sets the flag in ~/.claude/settings.json
|
|
67
|
-
bullswarm integrate install --agents claude --yes
|
|
68
|
-
|
|
69
|
-
# from Claude Code's plugin marketplace (a copy that `claude plugin update` refreshes)
|
|
70
|
-
claude plugin marketplace add Bulls-Work/bullswarm
|
|
71
|
-
claude plugin install bullswarm@bullswarm
|
|
72
|
-
|
|
73
|
-
# one session only, from the installed package
|
|
74
|
-
CLAUDE_CODE_ENABLE_FUNCTION_HOOKS=1 claude --plugin-dir "$(npm root -g)/bullswarm/mods/bullswarm"
|
|
23
|
+
Bullswarm turns the agent CLIs already signed in on your machine into **one fleet** that your main agent can command.
|
|
24
|
+
|
|
25
|
+
```mermaid
|
|
26
|
+
flowchart LR
|
|
27
|
+
you([You]) -->|goal| main["Your main agent<br/>Claude Code · Codex · Grok<br/>+ /bullswarm skill"]
|
|
28
|
+
main -->|"bullswarm run · workflow goal"| bs{{"Bullswarm<br/>pace router + workflow kernel"}}
|
|
29
|
+
bs --> c["claude -p"]
|
|
30
|
+
bs --> x["codex exec"]
|
|
31
|
+
bs --> g["grok -p"]
|
|
32
|
+
bs --> o["contributed providers"]
|
|
33
|
+
c & x & g & o -->|output| v["independent verification"]
|
|
34
|
+
v -->|"verdict + evidence"| main
|
|
75
35
|
```
|
|
76
36
|
|
|
77
|
-
|
|
78
|
-
|
|
79
|
-
|
|
80
|
-
|
|
37
|
+
- **Spends quota by pace.** Each task goes to the plan with the most spare quota for how far its window has run, so a plan that is behind, or about to reset with quota left, gets used first.
|
|
38
|
+
- **Picks models for you.** Tasks come in three lanes (`analyze`, `build`, `chore`). Setup asks each CLI which models it offers and suggests the newest one for each effort level, so you don't have to update settings every time a vendor ships a model.
|
|
39
|
+
- **Runs one task or a whole workflow.** `bullswarm run` sends one task to one agent and returns a verdict. For bigger goals your main agent writes a plan; Bullswarm runs the independent steps in parallel across agents, then integration and a final check by a different agent.
|
|
40
|
+
- **Checks the work, not the exit code.** A delegate saying "done" isn't enough. Bullswarm reads what it actually produced, and a workflow finishing is kept separate from its requirements being verified.
|
|
41
|
+
- **Shows everything live.** A terminal dashboard covers quota, history, running workflows and each agent's individual turns.
|
|
42
|
+
- **Works with the agent you already use.** The `/bullswarm` skill teaches Claude Code, Codex and Grok when to delegate. There's also an early-access Claude Code Mod that shows runs and usage inside Claude Code.
|
|
43
|
+
- **Lets you steer mid-run.** Add, change, remove or rerun steps while a workflow is running, or pause and resume it.
|
|
44
|
+
- **Extends with providers.** Claude Code, Codex and Grok are built in. Other CLIs can be added as providers without touching the core.
|
|
81
45
|
|
|
82
|
-
##
|
|
46
|
+
## See it
|
|
83
47
|
|
|
84
|
-
|
|
85
|
-
Use `bullswarm --setup` or `bullswarm setup` to open setup again, and use
|
|
86
|
-
`bullswarm workflow tui` when you want the explicit dashboard command. The
|
|
87
|
-
eight pages answer different questions:
|
|
48
|
+

|
|
88
49
|
|
|
89
|
-
|
|
90
|
-
|
|
91
|
-
|
|
92
|
-
|
|
93
|
-
|
|
94
|
-
|
|
95
|
-
|
|
96
|
-
|
|
97
|
-
|
|
98
|
-
|
|
99
|
-
|
|
100
|
-
The visual-fidelity pass keeps the same real numbers while composing and
|
|
101
|
-
colouring these pages like the approved prototype. At 55 columns every Home
|
|
102
|
-
today tile and per-step bar, every Stats model, project and its sparkline, every
|
|
103
|
-
Stats and Fleet pool row, and every run row in the Runs history table keeps one
|
|
104
|
-
row per item on the phone; Budget's per-pool block (a header over `used`,
|
|
105
|
-
`by bullswarm`, `room` and `so far`), Home's `budget · this week` pool (meter
|
|
106
|
-
plus a reset line, as the prototype draws it) and the Runs `active` entry are
|
|
107
|
-
the deliberate exceptions. At 200 columns every page composes to the frame
|
|
108
|
-
rather than capping at the 120-column composition. Estimated figures still
|
|
109
|
-
carry `≈` and their basis; a figure that cannot be measured is a line of words
|
|
110
|
-
saying so — no page draws an empty or dotted track for missing data.
|
|
111
|
-
|
|
112
|
-
Every page has a sticky header, a page tab row, and a sticky bottom nav. The
|
|
113
|
-
tab row is five tabs — Home, Runs, Budget, Stats, Fleet — with Run and Step
|
|
114
|
-
marking Runs, and Help shown only while it is open. The shared key table is:
|
|
115
|
-
|
|
116
|
-
| Key | Does |
|
|
50
|
+
*Home: what's running today, how much quota each plan has left, and what you've spent.*
|
|
51
|
+
|
|
52
|
+

|
|
53
|
+
|
|
54
|
+
*Run: one workflow's plan, which agent took each step, and what it cost.*
|
|
55
|
+
|
|
56
|
+
More screens, including phone-sized ones, are in the [gallery](https://bulls-work.github.io/bullswarm/guide/gallery).
|
|
57
|
+
|
|
58
|
+
## Why not just…
|
|
59
|
+
|
|
60
|
+
| Approach | What you give up |
|
|
117
61
|
|---|---|
|
|
118
|
-
|
|
|
119
|
-
|
|
|
120
|
-
|
|
|
121
|
-
|
|
|
122
|
-
| `f` | open Fleet |
|
|
123
|
-
| `h` | open Home |
|
|
124
|
-
| `?` | open Help |
|
|
125
|
-
| `1`–`9` | open that run from the nav |
|
|
126
|
-
| `Tab` | cycle the current page's sub-tabs |
|
|
127
|
-
| `Shift+Tab` | cycle workflows |
|
|
128
|
-
| `p` | cycle the period on Home and Stats |
|
|
129
|
-
| `Esc` / `←` | move out one page, then Home |
|
|
130
|
-
| `↑`/`k`, `↓`/`j` | move one line |
|
|
131
|
-
| `Enter` / `→` / `l` | open the selected run, step, tab, or action |
|
|
132
|
-
| `PgUp` / `PgDn` | scroll one screen |
|
|
133
|
-
| `Home` / `End` | jump to the top or bottom |
|
|
134
|
-
| `ctrl+s` | copy the screen through OSC 52, falling back to `pbcopy`, `wl-copy`, or `xclip` |
|
|
135
|
-
| `q` | quit the dashboard; workflows keep running |
|
|
136
|
-
|
|
137
|
-
The 0.33.0 rebinding is deliberate: `r` no longer refreshes, `b` no longer
|
|
138
|
-
moves out, and `Tab` no longer cycles workflows. Esc/left moves out, the view
|
|
139
|
-
refreshes itself, and Shift+Tab still cycles workflows. Runs also provides `/`
|
|
140
|
-
(filter), `a` (active/all), and `i` (install); Run provides `o`, `v`, and `t`,
|
|
141
|
-
and Fleet provides `e` for setup. Click tabs, tiles, bars, runs, steps, dates,
|
|
142
|
-
or controls, or use the wheel to scroll.
|
|
143
|
-
|
|
144
|
-
Home, Runs, Budget, and Stats read a per-run rollup that every finishing
|
|
145
|
-
run appends to `~/.bullswarm/history/runs.jsonl`. After upgrading, backfill
|
|
146
|
-
the runs that finished before 0.33.0 once with `bullswarm workflow reindex`;
|
|
147
|
-
legacy runs get a minimal record and only runs still in flight are skipped.
|
|
148
|
-
|
|
149
|
-
One bounded outcome — a task with a clear finish line:
|
|
150
|
-
|
|
151
|
-
```bash
|
|
152
|
-
bullswarm run --lane analyze --add-dir ~/some-repo --prompt "Explain the parser" --json
|
|
153
|
-
```
|
|
62
|
+
| **One agent's built-in subagents or workflows** | Everything draws on that one plan's quota, the same vendor grades its own work, and your process is tied to that vendor's feature. |
|
|
63
|
+
| **An API gateway that re-exposes your subscriptions** | Turning a consumer subscription into a generic API endpoint can conflict with provider terms, and you lose each agent's own tools and harness. |
|
|
64
|
+
| **Switching tools by hand** | You become the scheduler, and the quota you did not get to still expires. |
|
|
65
|
+
| **Bullswarm** | Drives each vendor's own headless CLI—`claude -p`, `codex exec`, `grok -p`—the way those CLIs are meant to be scripted, with the accounts you already signed in. Work lands where quota is spare, and a different agent checks it. |
|
|
154
66
|
|
|
155
|
-
|
|
156
|
-
`build` (edits) and `chore` (mechanical edits). This routes the prompt to
|
|
157
|
-
whichever pool is eligible, dispatches it, watches it to completion, verifies
|
|
158
|
-
the output, and prints one JSON verdict — nothing else runs and nothing is
|
|
159
|
-
left in the background.
|
|
160
|
-
|
|
161
|
-
Multi-step work, where you author the plan. The kernel — Bullswarm's own
|
|
162
|
-
runtime, not an agent — validates that program and executes it:
|
|
163
|
-
|
|
164
|
-
```json
|
|
165
|
-
{
|
|
166
|
-
"schemaVersion": "bullswarm.workflow.program.v2",
|
|
167
|
-
"actions": [
|
|
168
|
-
{ "id": "fix", "kind": "implement", "purpose": "Fix the failing tests",
|
|
169
|
-
"dependsOn": [], "ownedFiles": ["src/parser.js"], "affects": ["requirement-1"],
|
|
170
|
-
"evidenceFor": [], "prompt": "In ~/some-repo, fix the failing tests with the smallest correct change." }
|
|
171
|
-
]
|
|
172
|
-
}
|
|
173
|
-
```
|
|
67
|
+
Bullswarm never proxies a subscription as an API and never collects or shares your vendor credentials.
|
|
174
68
|
|
|
175
|
-
|
|
176
|
-
|
|
177
|
-
|
|
178
|
-
|
|
179
|
-
|
|
180
|
-
|
|
181
|
-
|
|
69
|
+
The controller is portable. Claude Code, Codex, Grok, or any capable caller drives the same CLI and durable workflow kernel, so your workflows are not locked into one vendor's orchestration.
|
|
70
|
+
|
|
71
|
+
## Why now
|
|
72
|
+
|
|
73
|
+
Agent plans now come with hard weekly limits, and serious work spans hours rather than prompts. A workflow can run for one to two hours across several agents—builders in parallel, an integrator, then independent acceptance—while you keep handing new goals to your main agent.
|
|
74
|
+
|
|
75
|
+
With Bullswarm routing by pace, running several goals at once no longer means worrying about wasting one plan's scarce quota: the fleet spends whichever plan is furthest behind.
|
|
182
76
|
|
|
183
|
-
|
|
184
|
-
|
|
185
|
-
|
|
186
|
-
|
|
187
|
-
|
|
188
|
-
|
|
189
|
-
|
|
190
|
-
|
|
191
|
-
|
|
192
|
-
|
|
193
|
-
|
|
194
|
-
|
|
195
|
-
|
|
196
|
-
|
|
197
|
-
|
|
198
|
-
- A pool whose weekly or monthly subscription window is about to reset gets
|
|
199
|
-
priority for its remaining surplus, so quota doesn't run out the clock
|
|
200
|
-
unspent.
|
|
201
|
-
- A pool already busy with other in-flight work yields to a quieter pool at a
|
|
202
|
-
similar pace, so a burst of parallel work spreads out instead of piling onto
|
|
203
|
-
one pool.
|
|
204
|
-
- A pool that reports a usage-limit error is benched until the provider's own
|
|
205
|
-
reset time and automatically re-tried after that — never left down for good,
|
|
206
|
-
and never retried early.
|
|
207
|
-
|
|
208
|
-
The full mechanics behind each of these are in
|
|
209
|
-
[Routing](https://bulls-work.github.io/bullswarm/guide/routing).
|
|
210
|
-
|
|
211
|
-
## What you get back
|
|
212
|
-
|
|
213
|
-
`bullswarm run` prints a JSON verdict when it finishes:
|
|
214
|
-
|
|
215
|
-
- `keepOnClaude: true` — the router says do this in-session; nothing ran
|
|
216
|
-
- `ok: true` (and `keepOnClaude` is false) — the output passed verification;
|
|
217
|
-
read `outFile`
|
|
218
|
-
- `ok: false` — `why` names the gate that failed
|
|
219
|
-
- `contentUsableDespiteExit: true` — the process exited non-zero but the
|
|
220
|
-
content still verified; read it before re-running
|
|
221
|
-
|
|
222
|
-
A non-zero exit from the delegate is never treated as success on its own. See
|
|
223
|
-
[Result envelope](https://bulls-work.github.io/bullswarm/reference/result) for the full
|
|
224
|
-
verdict shape.
|
|
225
|
-
|
|
226
|
-
A workflow produces a durable, versioned result envelope — a JSON document
|
|
227
|
-
with `runId` / `shortId`, status, per-requirement evidence, per-action
|
|
228
|
-
outcomes, and usage — instead of leaving you to parse a transcript.
|
|
229
|
-
|
|
230
|
-
```bash
|
|
231
|
-
bullswarm workflow watch <shortId> --next # wait for the next notable event, then exit
|
|
232
|
-
bullswarm workflow runs result <shortId> --json --summary # compact status once the run is terminal
|
|
77
|
+
## Quick start
|
|
78
|
+
|
|
79
|
+
You need Node.js 22.12 or later and at least one agent CLI you're signed in to (Claude Code, Codex or Grok).
|
|
80
|
+
|
|
81
|
+
The easiest way to set up is to let your agent do it. Paste this into Claude Code, Codex or Grok:
|
|
82
|
+
|
|
83
|
+
```text
|
|
84
|
+
Install and set up Bullswarm for me:
|
|
85
|
+
1. Run `npm i -g bullswarm`.
|
|
86
|
+
2. Run `bullswarm setup --yes --strategy --integrate` to find my agent CLIs,
|
|
87
|
+
pick models, and install the /bullswarm skill for each agent.
|
|
88
|
+
3. Run `bullswarm doctor` and tell me which agents are ready and how much
|
|
89
|
+
quota each one has left.
|
|
90
|
+
4. Read the installed /bullswarm skill so you know when to use
|
|
91
|
+
`bullswarm run` and when to write a workflow.
|
|
233
92
|
```
|
|
234
93
|
|
|
235
|
-
|
|
236
|
-
the exact flags to keep polling; `runs result --summary` is what to read once
|
|
237
|
-
a run finishes, and `runs result --json` (no `--summary`) gives the full
|
|
238
|
-
envelope for a failed or partial run. See
|
|
239
|
-
[Observing runs](https://bulls-work.github.io/bullswarm/guide/observing) for the
|
|
240
|
-
full shape of both.
|
|
94
|
+
After that, just give your agent goals as usual. It will hand work to Bullswarm when that helps. Run `bullswarm` in a terminal to open the dashboard; closing it doesn't stop anything that's running.
|
|
241
95
|
|
|
242
|
-
|
|
96
|
+
To set things up by hand, or to choose models and reasoning yourself, see [Getting started](https://bulls-work.github.io/bullswarm/guide/getting-started).
|
|
243
97
|
|
|
244
|
-
|
|
245
|
-
[bulls-work.github.io/bullswarm](https://bulls-work.github.io/bullswarm/).
|
|
98
|
+
## Supported agents and meters
|
|
246
99
|
|
|
247
|
-
|
|
248
|
-
|
|
249
|
-
|
|
250
|
-
|
|
100
|
+
| Agent CLI | Support | Quota windows tracked | Command it runs |
|
|
101
|
+
|---|---|---|---|
|
|
102
|
+
| Claude Code | Built in | weekly + 5-hour | `claude -p` |
|
|
103
|
+
| Codex | Built in | weekly | `codex exec` |
|
|
104
|
+
| Grok | Built in | weekly | `grok -p` |
|
|
105
|
+
| OpenCode | Contributed | none | `opencode run --auto` |
|
|
106
|
+
| Command Code | Contributed | weekly + monthly + 5-hour | `command-code -p` |
|
|
251
107
|
|
|
252
|
-
|
|
253
|
-
|
|
254
|
-
|
|
255
|
-
|
|
256
|
-
|
|
257
|
-
|
|
258
|
-
|
|
259
|
-
|
|
260
|
-
|
|
261
|
-
|
|
262
|
-
|
|
263
|
-
|
|
264
|
-
|
|
265
|
-
| [Result envelope](https://bulls-work.github.io/bullswarm/reference/result) | Every field of `run --json` and of the workflow result document |
|
|
266
|
-
| [Claude Code](https://bulls-work.github.io/bullswarm/integrations/claude-code) | The packaged skill, the MCP server, and the read-only Claude Mod counterpart under `mods/bullswarm` |
|
|
267
|
-
| [Codex and Grok](https://bulls-work.github.io/bullswarm/integrations/agent-clis) | What `bullswarm integrate` writes for each agent CLI, and how to check it |
|
|
268
|
-
| [Issue watcher](https://bulls-work.github.io/bullswarm/integrations/issue-watcher) | The launchd agent that triages and fixes new GitHub issues |
|
|
269
|
-
| [Historical notes](https://bulls-work.github.io/bullswarm/notes/) | Working notes, audits, and experiment writeups, kept as records |
|
|
270
|
-
|
|
271
|
-
## License
|
|
272
|
-
|
|
273
|
-
MIT
|
|
108
|
+
Each provider describes how to launch its CLI, read its quota, list its models and read its output. To add one, see [Adding a provider](https://bulls-work.github.io/bullswarm/reference/providers).
|
|
109
|
+
|
|
110
|
+
## Learn more
|
|
111
|
+
|
|
112
|
+
- [Documentation](https://bulls-work.github.io/bullswarm/)
|
|
113
|
+
- [Getting started](https://bulls-work.github.io/bullswarm/guide/getting-started)
|
|
114
|
+
- [Day-to-day playbook](https://bulls-work.github.io/bullswarm/guide/playbook)
|
|
115
|
+
- [How routing works](https://bulls-work.github.io/bullswarm/guide/routing)
|
|
116
|
+
- [Authoring workflows](https://bulls-work.github.io/bullswarm/guide/workflows)
|
|
117
|
+
- [The dashboard](https://bulls-work.github.io/bullswarm/guide/gallery) and [observing runs](https://bulls-work.github.io/bullswarm/guide/observing)
|
|
118
|
+
- [Provider reference](https://bulls-work.github.io/bullswarm/reference/providers)
|
|
119
|
+
- Contributing: [open an issue](https://github.com/Bulls-Work/bullswarm/issues) or [submit a pull request](https://github.com/Bulls-Work/bullswarm/pulls)
|
|
120
|
+
- [MIT licence](LICENSE)
|