bullswarm 0.35.4 → 0.35.5

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (95) hide show
  1. package/AGENTS.md +3 -2
  2. package/CHANGELOG.md +60 -2
  3. package/README.md +149 -237
  4. package/data/openrouter-benchmarks.json +9963 -9937
  5. package/docs/design/providers-0.35.2/CHANGELOG-draft.md +2 -2
  6. package/docs/design/providers-0.35.2/README.md +6 -6
  7. package/docs/design/stats-frames-0.33.2/spending-120.txt +18 -16
  8. package/docs/design/stats-frames-0.33.2/spending-200.txt +24 -0
  9. package/docs/design/stats-frames-0.33.2/spending-55.txt +32 -31
  10. package/docs/design/step-page-0.35.0/frames/failed-120.txt +2 -2
  11. package/docs/design/step-page-0.35.0/frames/failed-200.txt +2 -2
  12. package/docs/design/step-page-0.35.0/frames/failed-55.txt +2 -2
  13. package/docs/design/step-page-0.35.0/frames/real-failed-120.txt +3 -3
  14. package/docs/design/step-page-0.35.0/frames/real-failed-200.txt +4 -4
  15. package/docs/design/step-page-0.35.0/frames/real-failed-55.txt +3 -3
  16. package/docs/design/step-page-0.35.0/frames/rendered-failed-120.txt +3 -3
  17. package/docs/design/step-page-0.35.0/frames/rendered-failed-200.txt +3 -3
  18. package/docs/design/step-page-0.35.0/frames/rendered-failed-55.txt +3 -3
  19. package/docs/design/tidy-0.35.1/frames/real-stats-model-120.txt +13 -13
  20. package/docs/design/tidy-0.35.1/frames/real-stats-model-200.txt +13 -13
  21. package/docs/design/tidy-0.35.1/frames/real-stats-model-55.txt +9 -9
  22. package/docs/design/tidy-0.35.1/frames/real-stats-spending-120.txt +14 -14
  23. package/docs/design/tidy-0.35.1/frames/real-stats-spending-200.txt +14 -14
  24. package/docs/design/tidy-0.35.1/frames/real-stats-spending-55.txt +18 -18
  25. package/docs/design/tidy-0.35.1/frames/rendered-failed-120.txt +1 -1
  26. package/docs/design/tidy-0.35.1/frames/rendered-failed-200.txt +1 -1
  27. package/docs/design/tidy-0.35.1/frames/rendered-failed-55.txt +1 -1
  28. package/docs/guide/cost.md +1 -1
  29. package/docs/guide/getting-started.md +62 -16
  30. package/docs/guide/index.md +44 -6
  31. package/docs/guide/observing.md +42 -0
  32. package/docs/guide/playbook.md +120 -0
  33. package/docs/guide/routing.md +39 -1
  34. package/docs/index.md +47 -17
  35. package/docs/integrations/claude-code.md +4 -2
  36. package/docs/notes/index.md +1 -1
  37. package/docs/plans/dashboard-0.33.md +1 -1
  38. package/docs/plans/providers-0.35.2.goal.txt +2 -2
  39. package/docs/plans/providers-0.35.2.program.json +3 -3
  40. package/docs/reference/cli.md +40 -25
  41. package/docs/reference/configuration.md +22 -6
  42. package/docs/reference/providers.md +103 -4
  43. package/docs/studies/cost-audit-2026-09-18/claude-actual-vs-recorded.md +1 -1
  44. package/mods/bullswarm/README.md +7 -6
  45. package/mods/bullswarm/hooks/names.ts +7 -0
  46. package/mods/bullswarm/hooks/pane.tsx +13 -0
  47. package/mods/bullswarm/hooks/pools.ts +2 -0
  48. package/mods/bullswarm/hooks/register.ts +72 -14
  49. package/mods/bullswarm/hooks/runs.ts +1 -0
  50. package/mods/bullswarm/hooks/step.ts +65 -14
  51. package/mods/bullswarm/types/index.d.ts +4 -0
  52. package/package.json +4 -2
  53. package/providers/contrib/command-code/connector-history.json +288 -0
  54. package/providers/contrib/opencode/connector-history.json +241 -0
  55. package/providers/contrib/opencode/connector.json +1 -1
  56. package/skill/references/operations.md +8 -3
  57. package/src/cli.js +112 -44
  58. package/src/help.js +59 -45
  59. package/src/home-cli.js +1 -1
  60. package/src/lib/cli-flags.js +1 -1
  61. package/src/lib/config.js +2 -0
  62. package/src/lib/connector-copies.js +288 -0
  63. package/src/lib/model-family.js +252 -0
  64. package/src/lib/pool-labels.js +109 -0
  65. package/src/lib/providers.js +2 -0
  66. package/src/lib/reasoning.js +51 -5
  67. package/src/lib/strategy.js +568 -111
  68. package/src/provider-cli.js +63 -6
  69. package/src/providers/_schema.json +20 -1
  70. package/src/providers/claude-code/connector-history.json +330 -0
  71. package/src/providers/claude-code/connector.json +20 -5
  72. package/src/providers/claude-code/provider.mjs +124 -1
  73. package/src/providers/codex/connector-history.json +326 -0
  74. package/src/providers/codex/connector.json +53 -16
  75. package/src/providers/codex/provider.mjs +115 -0
  76. package/src/providers/echo/connector-history.json +142 -0
  77. package/src/providers/grok/connector-history.json +308 -0
  78. package/src/providers/grok/connector.json +65 -7
  79. package/src/setup.js +54 -58
  80. package/src/strategy-cli.js +215 -33
  81. package/src/strategy-dashboard.js +31 -10
  82. package/src/workflow/cli.js +11 -4
  83. package/src/workflow/dash-kit.js +55 -48
  84. package/src/workflow/dashboard.js +24 -1
  85. package/src/workflow/home-view.js +14 -56
  86. package/src/workflow/reprice.js +4 -2
  87. package/src/workflow/runs-cli.js +2 -1
  88. package/src/workflow/stat-kit.js +118 -20
  89. package/src/workflow/stats-model.js +1 -0
  90. package/src/workflow/stats-view.js +27 -32
  91. package/src/workflow/step-json.js +93 -0
  92. package/src/workflow/step-model.js +15 -6
  93. package/src/workflow/watch-cli.js +5 -0
  94. package/docs/public/favicon.svg +0 -7
  95. package/src/lib/release.js +0 -56
package/AGENTS.md CHANGED
@@ -69,7 +69,7 @@ is the canonical reference for the CLI surface.
69
69
  - Every verb must work non-interactively (no TTY). The interactive wizard is
70
70
  a human convenience, never a requirement.
71
71
  - Version single source: package.json. Release via
72
- `node bin/bullswarm.js release patch|minor|major`, then `git push` and
72
+ `npm run release -- patch|minor|major [--title "<headline>"]`, then `git push` and
73
73
  `git push --tags`
74
74
  — CI publishes through npm trusted publishing (OIDC), no tokens.
75
75
 
@@ -90,7 +90,8 @@ the authoring method is `skill/references/providers.md`.
90
90
  ## Releasing
91
91
 
92
92
  1. All tests green.
93
- 2. `node bin/bullswarm.js release patch` (creates commit + tag v*).
93
+ 2. `npm run release -- patch --title "<headline>"` (dates `## Unreleased`
94
+ in CHANGELOG.md, bumps package.json, creates commit + tag v*).
94
95
  3. `git push && git push --tags`.
95
96
  4. GitHub Actions publishes to npm via trusted publishing; verify with
96
97
  `npm view bullswarm version`.
package/CHANGELOG.md CHANGED
@@ -1,5 +1,63 @@
1
1
  # bullswarm changelog
2
2
 
3
+ ## Unreleased
4
+
5
+ ## 0.35.5 — model choices come from each CLI, newest version wins, setup stops pinning tiers
6
+
7
+ - cli: the maintainer-only `release` verb is gone from the CLI and its help;
8
+ maintainers run `npm run release -- patch|minor|major`, which also dates
9
+ the `## Unreleased` changelog section.
10
+
11
+ - package: the README screenshots and brand images (`docs/public/`) are no
12
+ longer shipped in the npm package; the npm page loads them from GitHub.
13
+
14
+ - grok: every tier now suggests the newest Grok, and the tiers differ only by reasoning. This is an owner decision: Grok has one model line. One family now covers every plain `grok-N.M`, and the newest version wins. A fresh home suggests `grok-4.7` on all three tiers: high at the connector's `xhigh`, then `medium: grok-4.7 · high reasoning` and `low: grok-4.7 · medium reasoning`, both with the reason `one Grok line, lighter reasoning for lighter tiers`. Before, it suggested `grok-4.6` for high, `grok-4.5` for medium, and nothing for low, and it listed `grok-4.7` as unranked. A `grok-4.8` takes over every tier the day `grok models` lists it. `grok-4.7-build-fast` is the same model at twice the price, so it keeps its base model's rank but is never suggested; `strategy set-rung` still selects it. The tier-wide best pick, the Codex and Claude suggestions, and pace-based routing across pools do not change. grok-4.7 and grok-4.5 now have dated price rows from [xAI's models page](https://docs.x.ai/docs/models), and grok-4.6's row was re-checked there (all on 2026-09-23). build-fast stays unpriced: xAI's Pricing page lists its long-context rates as $6.00 / $1.50 / $18.00 per 1M tokens, which is not twice grok-4.7's. The fallback list (`knownModels`) now matches what `grok models` lists.
15
+
16
+ - strategy: `generationFallback` now also covers a tier that no family serves. If a connector opts such a tier in and no exact row ranks a model for it, the tier takes the best-ranked family's newest-generation model at the declared reasoning. That stand-in is the pool's rung, and it stands at rank 0 against other pools, so it is the tier-wide pick only when no other pool serves the tier. A tier may also declare `why`, a plain reason that replaces the generated one; `provider validate` rejects an empty one.
17
+
18
+ - strategy: applying recommendations no longer pins tiers. `strategy apply`, `refresh --apply`, `setup --yes --strategy`, the setup wizard, the TUI's apply key, and the daily auto-refresh now write only each pool's rungs (its model and reasoning per tier). Before, they also pinned every tier to one pool: a fresh `setup --yes --strategy` sent high, medium, and low to Codex even at 100% used, and the auto-refresh pinned them again every 24 hours, so pace-based routing never picked another plan. Now every dispatch picks its pool by spare quota, and `strategy routes` shows each tier moving between pools. `strategy.assignments` stays empty unless you pin a tier yourself. The apply JSON now reports `rungs`, `bestNow` (the tier-wide pick, for display), `unpinned`, and `keptPins`, in place of `applied`.
19
+
20
+ - strategy: pins are explicit only. `strategy assign` makes one, marked `source: "user"`, and `strategy clear-assignment` removes it. Apply and the auto-refresh never remove a pin you set, and apply now keeps the pin's model on that pool's rung, so the pinned tier runs the model you named instead of the pool's recommended one.
21
+
22
+ - strategy: if an earlier setup or apply pinned your tiers, the next `strategy apply` or automatic refresh removes those pins and lists them under `unpinned`. A pin is treated as written by apply when it still has the pool and model that apply recorded in the cached report. Pins you set or changed yourself stay (`keptPins`), as does any pin in a home where apply never ran. Run `bullswarm strategy show` to see which pins remain and what removes each one.
23
+
24
+ - strategy: the tier suggestion reads as a suggestion, not a pin: `high: codex/gpt-6-astra (best now; routing picks by spare quota)` in `strategy show`, `best now (not pins): …` after the wizard and `setup --yes --strategy`, and `Routing now · by spare quota at each dispatch, unless pinned` in the TUI, whose route lines end in `by spare quota` or `pinned to <pool>`. A pinned tier gets its own line in `strategy show`, `pinned to <pool>/<model> by you · high dispatches go there while it is available · strategy clear-assignment high removes it`. `strategy routes` gives each route a `pin` field.
25
+
26
+ - strategy: `strategy show` counts unranked models instead of listing them. It prints one line, such as `unranked: 312 models (command-code 180, opencode 130, grok 2) · never recommended · strategy show --json lists them`, and the setup review screen prints the same count for each provider. `strategy show --json` keeps the full `unranked` list, and the model pickers keep their `(unranked)` tag.
27
+
28
+ - strategy: medium now falls back to the newest generation. When the family that serves a tier has no model in a pool's newest generation, a connector that opts the tier in (`generationFallback`) gives it the next-lower family's newest-generation model, at a declared reasoning level clamped to what that model supports. Codex opts medium in, and this is an owner decision, not a benchmark: 168 medium dispatches as luna at max reasoning ran 99% ok. So a fresh home suggests `medium: gpt-6-luna · max reasoning — no gpt-6 terra yet, newest generation preferred`, and medium returns to terra at the normal medium reasoning once a gpt-6 terra is discovered. High, low, and Claude are unchanged. `strategy apply`, `refresh --apply`, `setup --yes --strategy`, the wizard, and the TUI all write that level into the pool's rung, reported as `max (recommendation)`. It is never written over a level you set, and a later apply removes it when the fallback ends. `strategy show` and the setup screens say why. The cross-pool tier pick also skips a model disabled for its pool now, as each pool's own pick already did, so a tier suggestion can no longer name a model you turned off.
29
+
30
+ - setup, doctor: old copies of packaged connectors in `<home>/connectors/` now follow the package. Those copies were made by installs older than 0.29.0. A copy of a pool the package defines was already ignored, so its old prices (such as Opus 5's on Opus 5.5) never applied; a contrib copy whose provider is not enabled was still used. Every verb now checks each copy against `connector-history.json`, which holds fingerprints of every value each field has shipped with. A copy with no edits moves to `connectors/retired/` so the packaged connector loads, and a contrib provider is enabled first when only the copy defined its pool. A copy you edited is kept, and `bullswarm doctor` (`connector-copies`, a `!` warning) and `strategy show` name its stale fields. The dead copy-to-home step in `setup` is gone.
31
+
32
+ - strategy: connectors now declare model families (`modelFamilies`: a name pattern with a tier and base rank), and the version read from the model id orders each family. So a newly discovered model is ranked on the day it appears: gpt-6-sol and gpt-6-luna are no longer tierless, and a fresh home suggests gpt-6-luna for low instead of gpt-5.6-luna. Inside a pool and family, the newer version always wins, whatever benchmarks, prices, or pace say. Quality is compared on one scale, the rank, and benchmarks only break equal ranks. A model no rule classifies is listed as `unranked` and never recommended. Claude Code prices now follow the rate card line for each model: Opus 5.5 ($4/$20) and Fable 5.1 (cache hits $0.25) no longer inherit Opus 5's or Fable 5's price, and `provider validate` refuses a price on a family.
33
+
34
+ - strategy: Claude Code and Codex model choices now come from bounded, no-prompt installed-CLI handshakes per account; Claude discovery suppresses hooks and preserves explicit 1M selectors, Codex pages `model/list` and records per-model reasoning support, and the static lists remain only as failure fallbacks.
35
+
36
+ - docs: a new README with a banner, the problem, a comparison with the
37
+ alternatives, and dashboard screenshots taken from a made-up demo home;
38
+ a new playbook guide; `scripts/render-readme-shots.sh` rebuilds the
39
+ screenshots and refuses to write them if private terms or source IDs
40
+ survive.
41
+
42
+ - stats: hyphenated project names keep their complete basename when the panel
43
+ has room instead of being reduced to the suffix after the final hyphen.
44
+
45
+ - stats: charts share Home's even whole-unit ticks, zero baseline and daily
46
+ weekday labels, while compact money cells use the dashboard-wide source and
47
+ lower-bound glyphs and return the saved width to pool and project names.
48
+
49
+ - mod: the Step pane follows, selects, and expands the newest turn by default,
50
+ refreshes an open running Step every four seconds, and ticks its projected
51
+ in-flight command name and elapsed time every second without refetching.
52
+ - step: `workflow action show --json` and `workflow task show --json` now emit
53
+ compact Step schema v2: one activity per distinct attempt, one canonical
54
+ event array inside each activity, reference markers for aliases, and the
55
+ unchanged display projection used by the dashboard and Claude mod.
56
+ - pools: per-home display labels (`bullswarm pools label`) now shorten pool ids
57
+ across human CLI output, progress, the dashboard, and the Claude Mod without
58
+ changing credential, routing, meter, history, or workflow-record keys. Pool
59
+ arguments accept either form; JSON retains `pool` and adds `poolLabel`.
60
+
3
61
  ## 0.35.4 — one Runs table for workflows and tasks
4
62
 
5
63
  - mod: a live single task no longer replaces the selected workflow; each task
@@ -172,7 +230,7 @@
172
230
  sessions read-only; Command Code reads persisted project JSONL only when a
173
231
  session transcript exists and carries usage.
174
232
  - reprice: transcript lookup follows the loaded provider registry, so
175
- `opencode2*`, `opencode2:kaihk-*`, and `command-code` pools reach their
233
+ `opencode2*`, `opencode2:orbit-*`, and `command-code` pools reach their
176
234
  provider-owned readers without a hard-coded provider list. Ambiguous,
177
235
  missing, and checkpoint-only records stay unknown rather than becoming
178
236
  zero-cost attempts.
@@ -180,7 +238,7 @@
180
238
  `--no-session`, so their checkpoints are documented as non-recoverable
181
239
  history.
182
240
  - pricing: public model cards are retained only with a source and date (the
183
- cards were checked 2026-09-20). `kaihk/*` and `opencode/union-alpha` relay
241
+ cards were checked 2026-09-20). `orbit/*` and `opencode/union-alpha` relay
184
242
  identifiers have no public card, so the underlying OpenAI card is not
185
243
  substituted; observed Command Code models use the cited Command Code card.
186
244
  - evidence: live OpenCode and Command Code streams captured on 2026-09-20 are
package/README.md CHANGED
@@ -1,273 +1,185 @@
1
- # bullswarm
1
+ <p align="center">
2
+ <img src="docs/public/brand/bullswarm-banner.png" alt="Bullswarm — use every coding-agent plan you pay for" width="100%">
3
+ </p>
2
4
 
3
- Bullswarm is a CLI that sends a coding task to whichever of your installed
4
- agent CLIs — Claude Code, Codex, Grok, OpenCode, or Command Code — currently
5
- has unused subscription quota, then checks the result by its content.
5
+ <p align="center">
6
+ <img alt="npm version" src="https://img.shields.io/npm/v/bullswarm">
7
+ <img alt="MIT license" src="https://img.shields.io/npm/l/bullswarm">
8
+ <img alt="Node.js 22.12 or later" src="https://img.shields.io/badge/node-%3E%3D22.12-339933?logo=node.js&logoColor=white">
9
+ </p>
6
10
 
7
- One command runs one task. A second command runs a graph of dependent tasks
8
- across those same CLIs. Every result is judged by what was actually written,
9
- not by whether the process exited 0.
11
+ <p align="center">Route work across your coding-agent subscriptions, spend quota before it expires, and verify what comes back.</p>
10
12
 
11
- For an agent already working in a repo, the packaged `bullswarm` skill
12
- (`/bullswarm`, or `$bullswarm` where a skill uses that syntax) goes straight
13
- to `bullswarm run` or `bullswarm workflow goal`. A skill here is a short
14
- instruction file the agent CLI loads. There is no separate preview or
15
- classifier command to learn first.
13
+ ## The problem
16
14
 
17
- ## Why it exists
15
+ You pay for more than one coding agent: a Claude plan or two, Codex, Grok, maybe another.
18
16
 
19
- Subscription quota expires on a clock, whether you spend it or not, and a
20
- single coding agent's judgment on whether its own work is done should not be
21
- the only check in the loop. Bullswarm picks whichever installed agent CLI has
22
- the most unused quota right now, and treats every delegate's output as
23
- evidence to be verified — never as an authority to be trusted on its word.
17
+ Every one of them meters you on a clock — five-hour windows, weekly caps, monthly allowances — and unused quota is simply gone when the window resets.
24
18
 
25
- The same problem compounds on multi-step goals: one agent planning and
26
- executing everything serially leaves every other installed CLI's quota idle,
27
- and having that same agent be the sole judge of whether the whole goal is
28
- done multiplies the risk instead of dividing it. Bullswarm's workflow engine
29
- runs a graph of dependent actions across whichever pools have quota to spare
30
- — a pool is one installed agent CLI, or one account of that CLI — and
31
- computes completion from evidence the graph itself required, not from any one
32
- delegate's own say-so.
19
+ So one plan runs dry mid-task while the others sit idle. You switch tools by hand just to spend what you already paid for. And the agent that wrote the code is usually the one that decides it is done.
33
20
 
34
- ## Install
21
+ ## What Bullswarm does
35
22
 
36
- ```bash
37
- npm i -g bullswarm
38
- bullswarm setup
23
+ Bullswarm turns the agent CLIs already signed in on your machine into **one fleet** that your main agent can command.
24
+
25
+ ```mermaid
26
+ flowchart LR
27
+ you([You]) -->|goal| main["Your main agent<br/>Claude Code · Codex · Grok<br/>+ /bullswarm skill"]
28
+ main -->|"bullswarm run · workflow goal"| bs{{"Bullswarm<br/>pace router + workflow kernel"}}
29
+ bs --> c["claude -p"]
30
+ bs --> x["codex exec"]
31
+ bs --> g["grok -p"]
32
+ bs --> o["contributed providers"]
33
+ c & x & g & o -->|output| v["independent verification"]
34
+ v -->|"verdict + evidence"| main
39
35
  ```
40
36
 
41
- Later, `bullswarm update` upgrades that install to the latest published
42
- version in place (`bullswarm update --check` only reports). Requires Node.js
43
- 22.12 or later. `bullswarm setup` walks through detecting your
44
- installed agent CLIs, showing their quota state, and writing a routing
45
- configuration. See
46
- [Getting started](https://bulls-work.github.io/bullswarm/guide/getting-started)
47
- for integrating Bullswarm's skill into Codex, Claude, and Grok, and for the
48
- full quick-start command list.
49
-
50
- ## Claude Mod (early access)
51
-
52
- `mods/bullswarm` is the same routing injected into Claude Code's own engine
53
- as a Claude Mod (a plugin of TypeScript function hooks, behind
54
- `CLAUDE_CODE_ENABLE_FUNCTION_HOOKS=1`): the pool meters drawn above the
55
- prompt, the pools named in the model's context, Claude's general-purpose
56
- subagents routed to whichever pool has surplus and answered with the
57
- verified output, and the verdict appended to every `bullswarm run` the model
58
- runs. The Mod is the dashboard's read-only counterpart: its pane shows Run,
59
- Step, and the Usage/Pools view in the same meter colours. It does not expose
60
- the dashboard's Home, Runs, Budget, Stats, Fleet, or Help pages, and
61
- it has no edit or install action. See [mods/bullswarm/README.md](mods/bullswarm/README.md).
62
-
63
- Three ways to load it:
37
+ - **Paces quota instead of guessing.** Each task goes to the eligible pool with the most unused quota relative to its reset clock. A pool that is furthest behind pace—or close to resetting with quota left—moves forward.
38
+ - **Routes by the work.** The `analyze`, `build`, and `chore` lanes derive an effort tier; setup maps those tiers to provider and model choices, while `bullswarm setup --wizard` also configures reasoning depth. Or ask your agent to run the non-interactive setup.
39
+ - **Runs one bounded task.** `bullswarm run` routes it to one agent, waits, and returns a verdict.
40
+ - **Executes multi-phase workflows.** Your main agent authors a dependency graph; Bullswarm schedules ready, file-disjoint actions across agents in parallel, then runs integration and independent acceptance when the plan calls for them.
41
+ - **Verifies content, not exit codes.** A delegate's output is evidence. A zero exit code is never enough by itself, and workflow completion stays separate from verified requirements.
42
+ - **Shows the whole system live.** The terminal dashboard has Home, Runs, Run, Step, Budget, Stats, Fleet, and Help views, from portfolio-level quota and history down to individual agent turns.
43
+ - **Reads licence meters.** Built-in readers cover Claude Code, Codex, and Grok; the provider interface extends routing, meters, models, and event streams without putting vendor quirks in the core.
44
+ - **Lets your agent delegate.** The packaged `/bullswarm` skill teaches Claude Code, Codex, and Grok when to use a single run or a workflow.
45
+ - **Lives inside Claude Code too.** The early-access Claude Code Mod adds Bullswarm routing, a run strip, and read-only Run, Step, and Usage panes inside Claude Code.
46
+ - **Changes course while work is live.** Newer steering and plan-revision commands can add, amend, remove, or rerun actions; pause and resume remain explicit.
64
47
 
65
- ```bash
66
- # with the CLI: links the mod under ~/.claude/skills and sets the flag in ~/.claude/settings.json
67
- bullswarm integrate install --agents claude --yes
48
+ ## Why not just…
68
49
 
69
- # from Claude Code's plugin marketplace (a copy that `claude plugin update` refreshes)
70
- claude plugin marketplace add Bulls-Work/bullswarm
71
- claude plugin install bullswarm@bullswarm
50
+ | Approach | What you give up |
51
+ |---|---|
52
+ | **One agent's built-in subagents or workflows** | Everything draws on that one plan's quota, the same vendor grades its own work, and your process is tied to that vendor's feature. |
53
+ | **An API gateway that re-exposes your subscriptions** | Turning a consumer subscription into a generic API endpoint can conflict with provider terms, and you lose each agent's own tools and harness. |
54
+ | **Switching tools by hand** | You become the scheduler, and the quota you did not get to still expires. |
55
+ | **Bullswarm** | Drives each vendor's own headless CLI—`claude -p`, `codex exec`, `grok -p`—the way those CLIs are meant to be scripted, with the accounts you already signed in. Work lands where quota is spare, and a different agent checks it. |
72
56
 
73
- # one session only, from the installed package
74
- CLAUDE_CODE_ENABLE_FUNCTION_HOOKS=1 claude --plugin-dir "$(npm root -g)/bullswarm/mods/bullswarm"
75
- ```
57
+ Bullswarm never proxies a subscription as an API and never collects or shares your vendor credentials.
76
58
 
77
- The marketplace route still needs `CLAUDE_CODE_ENABLE_FUNCTION_HOOKS=1`, in
78
- the shell or under `env` in `~/.claude/settings.json`, and the `bullswarm`
79
- CLI on `PATH`. Pick one route: an installed marketplace copy takes precedence
80
- over the skills-dir link, and Claude says so at startup.
59
+ The controller is portable. Claude Code, Codex, Grok, or any capable caller drives the same CLI and durable workflow kernel, so your workflows are not locked into one vendor's orchestration.
81
60
 
82
- ## Quick start
61
+ ## Why now
62
+
63
+ Agent plans now come with hard weekly limits, and serious work spans hours rather than prompts. A workflow can run for one to two hours across several agents—builders in parallel, an integrator, then independent acceptance—while you keep handing new goals to your main agent.
64
+
65
+ With Bullswarm routing by pace, running several goals at once no longer means worrying about wasting one plan's scarce quota: the fleet spends whichever plan is furthest behind.
66
+
67
+ ## See it
83
68
 
84
- Once setup is complete, bare `bullswarm` opens the dashboard on its Home page.
85
- Use `bullswarm --setup` or `bullswarm setup` to open setup again, and use
86
- `bullswarm workflow tui` when you want the explicit dashboard command. The
87
- eight pages answer different questions:
69
+ ![Bullswarm Home dashboard showing today's work, quota, spend, and recent runs](docs/public/screens/home.png)
88
70
 
89
- | Page | What it answers |
71
+ *Home — today's work, budget position, trends, and recent runs in one view.*
72
+
73
+ | Runs | Run |
90
74
  |---|---|
91
- | Home | What happened today, what is verified, what measured or labelled money/licence data exists, and what is active or recent; includes pool/model/project breakdowns. |
92
- | Runs | Which workflows are active or historical — one `active` block and the History day table below it — plus the product commands; the integration line and `/` filter live here, and `?` carries the agents and `run it` blocks. |
93
- | Run | Where one workflow is in its plan, which workers are live or next, and its ETA and budget shares. |
94
- | Step | What one action is doing: route, attempt, activity, verdict/failure, prompt, usage, events, output, and artifacts. |
95
- | Budget | Pool quota windows, measured worker-minutes versus rest, measured or `≈` money, fit, and biggest workflows. |
96
- | Stats | Runs, spend, worker-minutes, and verification over 7d/30d/all, by pool, model, and project. |
97
- | Fleet | Lane/provider model and reasoning rungs, records, meter state, and the setup edit hand-off. |
98
- | Help | Every key, click, layout rule, and dashboard command. |
99
-
100
- The visual-fidelity pass keeps the same real numbers while composing and
101
- colouring these pages like the approved prototype. At 55 columns every Home
102
- today tile and per-step bar, every Stats model, project and its sparkline, every
103
- Stats and Fleet pool row, and every run row in the Runs history table keeps one
104
- row per item on the phone; Budget's per-pool block (a header over `used`,
105
- `by bullswarm`, `room` and `so far`), Home's `budget · this week` pool (meter
106
- plus a reset line, as the prototype draws it) and the Runs `active` entry are
107
- the deliberate exceptions. At 200 columns every page composes to the frame
108
- rather than capping at the 120-column composition. Estimated figures still
109
- carry `≈` and their basis; a figure that cannot be measured is a line of words
110
- saying so — no page draws an empty or dotted track for missing data.
111
-
112
- Every page has a sticky header, a page tab row, and a sticky bottom nav. The
113
- tab row is five tabs — Home, Runs, Budget, Stats, Fleet — with Run and Step
114
- marking Runs, and Help shown only while it is open. The shared key table is:
115
-
116
- | Key | Does |
75
+ | [![Bullswarm Runs dashboard showing active and historical workflows and tasks](docs/public/screens/runs.png)](docs/public/screens/runs.png) | [![Bullswarm Run page showing phases, live work, spend, and timeline](docs/public/screens/run.png)](docs/public/screens/run.png) |
76
+ | Active and historical workflows and single tasks in one table. | Phases, live workers, costs, and the attempt timeline. |
77
+
78
+ | Step | Stats |
117
79
  |---|---|
118
- | `r` | open Runs; the view refreshes itself, so `r` is no longer refresh |
119
- | `b` | open Budget; it is no longer move-out |
120
- | `s` | open Stats |
121
- | `y` | open Runs at its History table (the first day header), except that it confirms a pending stop |
122
- | `f` | open Fleet |
123
- | `h` | open Home |
124
- | `?` | open Help |
125
- | `1`–`9` | open that run from the nav |
126
- | `Tab` | cycle the current page's sub-tabs |
127
- | `Shift+Tab` | cycle workflows |
128
- | `p` | cycle the period on Home and Stats |
129
- | `Esc` / `←` | move out one page, then Home |
130
- | `↑`/`k`, `↓`/`j` | move one line |
131
- | `Enter` / `→` / `l` | open the selected run, step, tab, or action |
132
- | `PgUp` / `PgDn` | scroll one screen |
133
- | `Home` / `End` | jump to the top or bottom |
134
- | `ctrl+s` | copy the screen through OSC 52, falling back to `pbcopy`, `wl-copy`, or `xclip` |
135
- | `q` | quit the dashboard; workflows keep running |
136
-
137
- The 0.33.0 rebinding is deliberate: `r` no longer refreshes, `b` no longer
138
- moves out, and `Tab` no longer cycles workflows. Esc/left moves out, the view
139
- refreshes itself, and Shift+Tab still cycles workflows. Runs also provides `/`
140
- (filter), `a` (active/all), and `i` (install); Run provides `o`, `v`, and `t`,
141
- and Fleet provides `e` for setup. Click tabs, tiles, bars, runs, steps, dates,
142
- or controls, or use the wheel to scroll.
143
-
144
- Home, Runs, Budget, and Stats read a per-run rollup that every finishing
145
- run appends to `~/.bullswarm/history/runs.jsonl`. After upgrading, backfill
146
- the runs that finished before 0.33.0 once with `bullswarm workflow reindex`;
147
- legacy runs get a minimal record and only runs still in flight are skipped.
148
-
149
- One bounded outcome — a task with a clear finish line:
80
+ | [![Bullswarm Step page showing agent turns, result, task, and cost](docs/public/screens/step.png)](docs/public/screens/step.png) | [![Bullswarm Stats dashboard showing usage and outcome trends](docs/public/screens/stats.png)](docs/public/screens/stats.png) |
81
+ | The live or saved agent transcript, result, task, and cost evidence. | Workflow, spend, worker-time, and verification trends. |
82
+
83
+ [![Bullswarm Budget dashboard showing licence meters, quota pace, and reset windows](docs/public/screens/budget.png)](docs/public/screens/budget.png)
84
+
85
+ *Budget — spend spare quota before each weekly or monthly window resets.*
86
+
87
+ <p align="center">
88
+ <img src="docs/public/screens/home-phone.png" alt="Bullswarm Home dashboard at phone width" width="360">
89
+ <img src="docs/public/screens/run-phone.png" alt="Bullswarm Run page at phone width" width="360">
90
+ </p>
91
+
92
+ *Home and Run retain their core evidence at a 55-column phone width.*
93
+
94
+ ## Quick start
95
+
96
+ Requires Node.js 22.12 or later.
97
+
98
+ ```bash
99
+ npm i -g bullswarm
100
+ bullswarm setup
101
+ ```
102
+
103
+ `setup` discovers installed agent CLIs, shows their quota state, and opens the provider/model control centre. An agent or CI process can use discovered defaults without prompts:
150
104
 
151
105
  ```bash
152
- bullswarm run --lane analyze --add-dir ~/some-repo --prompt "Explain the parser" --json
106
+ bullswarm setup --yes --strategy --integrate
107
+ bullswarm doctor
153
108
  ```
154
109
 
155
- `--lane analyze` tags the work as read-only analysis. The other lanes are
156
- `build` (edits) and `chore` (mechanical edits). This routes the prompt to
157
- whichever pool is eligible, dispatches it, watches it to completion, verifies
158
- the output, and prints one JSON verdict — nothing else runs and nothing is
159
- left in the background.
160
-
161
- Multi-step work, where you author the plan. The kernel — Bullswarm's own
162
- runtime, not an agent — validates that program and executes it:
163
-
164
- ```json
165
- {
166
- "schemaVersion": "bullswarm.workflow.program.v2",
167
- "actions": [
168
- { "id": "fix", "kind": "implement", "purpose": "Fix the failing tests",
169
- "dependsOn": [], "ownedFiles": ["src/parser.js"], "affects": ["requirement-1"],
170
- "evidenceFor": [], "prompt": "In ~/some-repo, fix the failing tests with the smallest correct change." }
171
- ]
172
- }
110
+ Run one bounded task:
111
+
112
+ ```bash
113
+ bullswarm run --lane analyze --add-dir . \
114
+ --prompt "List every TODO in src with file and line number." --json
173
115
  ```
174
116
 
117
+ Start a first workflow with an explicitly delegated planner:
118
+
175
119
  ```bash
176
- # 1. Write plan.json — the bounded action program the kernel will enforce.
177
- bullswarm workflow plan validate "Fix the failing tests and verify the change" \
178
- --cwd ~/some-repo --program plan.json --json # exit 0 valid, exit 2 with the issues; nothing launches
179
- bullswarm workflow goal "Fix the failing tests and verify the change" \
180
- --cwd ~/some-repo --program plan.json # launches, prints a short ID and observation commands, returns
120
+ bullswarm workflow goal \
121
+ "Audit this repository and write a one-page summary" \
122
+ --cwd . --orchestrator auto --watch
181
123
  ```
182
124
 
183
- `workflow goal` starts a durable background run and returns immediately by
184
- default; add `--watch` to follow its low-noise progress in the same terminal
185
- instead.
186
-
187
- ## How it picks a pool
188
-
189
- - Work is tagged by lane — read-only analysis, ordinary build work, or
190
- mechanical chores — not assigned to a fixed pool ahead of time.
191
- - Among the pools that can do the work, the one furthest behind its own quota
192
- pace (the most unspent surplus) wins, so quota doesn't expire unused.
193
- - A pool close to its rolling 5-hour usage ceiling gives way to one with
194
- headroom only while another pool is actually behind its own pace — it is an
195
- ordering penalty, not a cutoff. A step may run a pool all the way to 100% of
196
- that window, because if the provider stops the worker at the wall the retry
197
- is briefed on what it had already written.
198
- - A pool whose weekly or monthly subscription window is about to reset gets
199
- priority for its remaining surplus, so quota doesn't run out the clock
200
- unspent.
201
- - A pool already busy with other in-flight work yields to a quieter pool at a
202
- similar pace, so a burst of parallel work spreads out instead of piling onto
203
- one pool.
204
- - A pool that reports a usage-limit error is benched until the provider's own
205
- reset time and automatically re-tried after that — never left down for good,
206
- and never retried early.
207
-
208
- The full mechanics behind each of these are in
209
- [Routing](https://bulls-work.github.io/bullswarm/guide/routing).
210
-
211
- ## What you get back
212
-
213
- `bullswarm run` prints a JSON verdict when it finishes:
214
-
215
- - `keepOnClaude: true` — the router says do this in-session; nothing ran
216
- - `ok: true` (and `keepOnClaude` is false) — the output passed verification;
217
- read `outFile`
218
- - `ok: false` — `why` names the gate that failed
219
- - `contentUsableDespiteExit: true` — the process exited non-zero but the
220
- content still verified; read it before re-running
221
-
222
- A non-zero exit from the delegate is never treated as success on its own. See
223
- [Result envelope](https://bulls-work.github.io/bullswarm/reference/result) for the full
224
- verdict shape.
225
-
226
- A workflow produces a durable, versioned result envelope — a JSON document
227
- with `runId` / `shortId`, status, per-requirement evidence, per-action
228
- outcomes, and usage — instead of leaving you to parse a transcript.
125
+ For normal use, your main agent should author `plan.json`, validate it, and launch the exact same goal:
229
126
 
230
127
  ```bash
231
- bullswarm workflow watch <shortId> --next # wait for the next notable event, then exit
232
- bullswarm workflow runs result <shortId> --json --summary # compact status once the run is terminal
128
+ bullswarm workflow plan contract \
129
+ "1. Fix the parser. 2. Add independent verification." --cwd . --json
130
+
131
+ bullswarm workflow plan validate \
132
+ "1. Fix the parser. 2. Add independent verification." \
133
+ --cwd . --program plan.json --json
134
+
135
+ bullswarm workflow goal \
136
+ "1. Fix the parser. 2. Add independent verification." \
137
+ --cwd . --program plan.json --watch
233
138
  ```
234
139
 
235
- `workflow watch --next` prints one line per event and relaunches itself with
236
- the exact flags to keep polling; `runs result --summary` is what to read once
237
- a run finishes, and `runs result --json` (no `--summary`) gives the full
238
- envelope for a failed or partial run. See
239
- [Observing runs](https://bulls-work.github.io/bullswarm/guide/observing) for the
240
- full shape of both.
140
+ Open the dashboard at any time. Quitting it does not stop running workflows.
241
141
 
242
- ## Documentation
142
+ ```bash
143
+ bullswarm
144
+ # explicit form:
145
+ bullswarm workflow tui
146
+ ```
243
147
 
244
- The full documentation is published at
245
- [bulls-work.github.io/bullswarm](https://bulls-work.github.io/bullswarm/).
148
+ ## A practical playbook
246
149
 
247
- The dashboard is the main screen after setup: bare `bullswarm` opens Home,
248
- `bullswarm --setup` or `bullswarm setup` opens setup, and
249
- `bullswarm workflow tui` is the explicit form. The [Observing runs](https://bulls-work.github.io/bullswarm/guide/observing)
250
- page maps its pages, keys, and mouse controls.
150
+ 1. **Nail down the outcome.** State what must change, what must remain untouched, and what evidence will count as done.
151
+ 2. **Hand it to your main agent.** With the `/bullswarm` skill installed, it chooses a bounded run or authors a workflow program with clear territories, dependencies, integration, and acceptance.
152
+ 3. **Let the graph fan out.** Ready, file-disjoint actions can run across Codex, Grok, Claude accounts, and contributed providers while routing spends the quota furthest behind pace.
153
+ 4. **Stay in control.** Follow the dashboard or `workflow watch`; send guidance, pause, or revise the live plan when the goal changes.
154
+ 5. **Sign off on evidence.** Read the durable result envelope and requirement evidence. `completed` and `verified` answer different questions.
251
155
 
252
- | Page | What it covers |
253
- |---|---|
254
- | [Introduction](https://bulls-work.github.io/bullswarm/guide/) | What Bullswarm is, the two entry points, and the four rules it never breaks |
255
- | [Getting started](https://bulls-work.github.io/bullswarm/guide/getting-started) | Install, `setup`, `doctor`, agent integration, and your first verified run |
256
- | [Concepts](https://bulls-work.github.io/bullswarm/guide/concepts) | Pools, lanes, surplus, the two windows, verdicts, quarantine, and the run directory |
257
- | [Run one task](https://bulls-work.github.io/bullswarm/guide/run) | Every `bullswarm run` option, and what each verdict asks you to do |
258
- | [Workflows](https://bulls-work.github.io/bullswarm/guide/workflows) | Authoring the program `workflow goal` executes: territories, dependencies, integration, acceptance |
259
- | [Observing runs](https://bulls-work.github.io/bullswarm/guide/observing) | `workflow watch`, the Home/Runs/Run/Step/Budget/Stats/Fleet/Help dashboard pages, keys, mouse, and terminal glyphs |
260
- | [Routing](https://bulls-work.github.io/bullswarm/guide/routing) | How a pool is picked: pace, 5-hour headroom, urgency, load, quarantine |
261
- | [CLI reference](https://bulls-work.github.io/bullswarm/reference/cli) | Every verb and nested subcommand, with its flags and defaults |
262
- | [Workflow program](https://bulls-work.github.io/bullswarm/reference/program) | The `bullswarm.workflow.program.v2` document: action fields, kinds, validation rules |
263
- | [Configuration](https://bulls-work.github.io/bullswarm/reference/configuration) | The Bullswarm home, `state.json`, strategy models and rungs, environment variables |
264
- | [Providers](https://bulls-work.github.io/bullswarm/reference/providers) | Adding your own agent CLI or reseller account as a provider plugin |
265
- | [Result envelope](https://bulls-work.github.io/bullswarm/reference/result) | Every field of `run --json` and of the workflow result document |
266
- | [Claude Code](https://bulls-work.github.io/bullswarm/integrations/claude-code) | The packaged skill, the MCP server, and the read-only Claude Mod counterpart under `mods/bullswarm` |
267
- | [Codex and Grok](https://bulls-work.github.io/bullswarm/integrations/agent-clis) | What `bullswarm integrate` writes for each agent CLI, and how to check it |
268
- | [Issue watcher](https://bulls-work.github.io/bullswarm/integrations/issue-watcher) | The launchd agent that triages and fixes new GitHub issues |
269
- | [Historical notes](https://bulls-work.github.io/bullswarm/notes/) | Working notes, audits, and experiment writeups, kept as records |
270
-
271
- ## License
272
-
273
- MIT
156
+ You can keep several independent goals running without manually balancing every plan. Bullswarm accounts for in-flight load, while each workflow keeps its own durable state, outputs, events, and result.
157
+
158
+ ## Supported agents and meters
159
+
160
+ | Agent CLI | Provider status | Subscription meter declared by the shipped connector | Headless entry point |
161
+ |---|---|---|---|
162
+ | Claude Code | Built in | weekly + 5-hour | `claude -p` |
163
+ | Codex | Built in | weekly | `codex exec` |
164
+ | Grok | Built in | weekly | `grok -p` |
165
+ | OpenCode | Contributed | none in the base connector | `opencode run --auto` |
166
+ | Command Code | Contributed | weekly + monthly + 5-hour | `command-code -p` |
167
+
168
+ Providers declare their own spawn command, meter reader, model discovery, reasoning levels, event decoding, and capabilities. See [Adding a provider](https://bulls-work.github.io/bullswarm/reference/providers).
169
+
170
+ ## Status
171
+
172
+ Bullswarm is used daily. The maintainer's local records contained **335 workflow runs** and **507 single tasks** as of September 2026.
173
+
174
+ The routing, content verification, durable workflow kernel, dashboard, and built-in providers are established parts of the project. Mid-run steering and whole-plan revision are newer; use their validation and revision guards, and inspect the resulting evidence.
175
+
176
+ ## Learn more
177
+
178
+ - [Documentation](https://bulls-work.github.io/bullswarm/)
179
+ - [Getting started](https://bulls-work.github.io/bullswarm/guide/getting-started)
180
+ - [How routing works](https://bulls-work.github.io/bullswarm/guide/routing)
181
+ - [Authoring workflows](https://bulls-work.github.io/bullswarm/guide/workflows)
182
+ - [Observing runs and the dashboard](https://bulls-work.github.io/bullswarm/guide/observing)
183
+ - [Provider reference](https://bulls-work.github.io/bullswarm/reference/providers)
184
+ - Contributing: [open an issue](https://github.com/Bulls-Work/bullswarm/issues) or [submit a pull request](https://github.com/Bulls-Work/bullswarm/pulls)
185
+ - [MIT licence](LICENSE)