bullswarm 0.35.4 → 0.35.6

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (97) hide show
  1. package/AGENTS.md +3 -2
  2. package/CHANGELOG.md +84 -2
  3. package/README.md +104 -257
  4. package/data/openrouter-benchmarks.json +13349 -12793
  5. package/docs/design/providers-0.35.2/CHANGELOG-draft.md +2 -2
  6. package/docs/design/providers-0.35.2/README.md +6 -6
  7. package/docs/design/redesign-mechanics-principles-options.md +385 -0
  8. package/docs/design/stats-frames-0.33.2/spending-120.txt +18 -16
  9. package/docs/design/stats-frames-0.33.2/spending-200.txt +24 -0
  10. package/docs/design/stats-frames-0.33.2/spending-55.txt +32 -31
  11. package/docs/design/step-page-0.35.0/frames/failed-120.txt +2 -2
  12. package/docs/design/step-page-0.35.0/frames/failed-200.txt +2 -2
  13. package/docs/design/step-page-0.35.0/frames/failed-55.txt +2 -2
  14. package/docs/design/step-page-0.35.0/frames/real-failed-120.txt +3 -3
  15. package/docs/design/step-page-0.35.0/frames/real-failed-200.txt +4 -4
  16. package/docs/design/step-page-0.35.0/frames/real-failed-55.txt +3 -3
  17. package/docs/design/step-page-0.35.0/frames/rendered-failed-120.txt +3 -3
  18. package/docs/design/step-page-0.35.0/frames/rendered-failed-200.txt +3 -3
  19. package/docs/design/step-page-0.35.0/frames/rendered-failed-55.txt +3 -3
  20. package/docs/design/tidy-0.35.1/frames/real-stats-model-120.txt +13 -13
  21. package/docs/design/tidy-0.35.1/frames/real-stats-model-200.txt +13 -13
  22. package/docs/design/tidy-0.35.1/frames/real-stats-model-55.txt +9 -9
  23. package/docs/design/tidy-0.35.1/frames/real-stats-spending-120.txt +14 -14
  24. package/docs/design/tidy-0.35.1/frames/real-stats-spending-200.txt +14 -14
  25. package/docs/design/tidy-0.35.1/frames/real-stats-spending-55.txt +18 -18
  26. package/docs/design/tidy-0.35.1/frames/rendered-failed-120.txt +1 -1
  27. package/docs/design/tidy-0.35.1/frames/rendered-failed-200.txt +1 -1
  28. package/docs/design/tidy-0.35.1/frames/rendered-failed-55.txt +1 -1
  29. package/docs/guide/cost.md +1 -1
  30. package/docs/guide/gallery.md +57 -0
  31. package/docs/guide/getting-started.md +62 -16
  32. package/docs/guide/index.md +44 -6
  33. package/docs/guide/observing.md +42 -0
  34. package/docs/guide/playbook.md +120 -0
  35. package/docs/guide/routing.md +39 -1
  36. package/docs/index.md +50 -17
  37. package/docs/integrations/claude-code.md +4 -2
  38. package/docs/notes/index.md +1 -1
  39. package/docs/plans/dashboard-0.33.md +1 -1
  40. package/docs/plans/providers-0.35.2.goal.txt +2 -2
  41. package/docs/plans/providers-0.35.2.program.json +3 -3
  42. package/docs/reference/cli.md +41 -26
  43. package/docs/reference/configuration.md +24 -7
  44. package/docs/reference/providers.md +103 -4
  45. package/docs/studies/cost-audit-2026-09-18/claude-actual-vs-recorded.md +1 -1
  46. package/mods/bullswarm/README.md +7 -6
  47. package/mods/bullswarm/hooks/names.ts +7 -0
  48. package/mods/bullswarm/hooks/pane.tsx +13 -0
  49. package/mods/bullswarm/hooks/pools.ts +2 -0
  50. package/mods/bullswarm/hooks/register.ts +72 -14
  51. package/mods/bullswarm/hooks/runs.ts +1 -0
  52. package/mods/bullswarm/hooks/step.ts +65 -14
  53. package/mods/bullswarm/types/index.d.ts +4 -0
  54. package/package.json +4 -2
  55. package/providers/contrib/command-code/connector-history.json +288 -0
  56. package/providers/contrib/opencode/connector-history.json +241 -0
  57. package/providers/contrib/opencode/connector.json +1 -1
  58. package/skill/references/operations.md +8 -3
  59. package/src/cli.js +112 -44
  60. package/src/help.js +69 -50
  61. package/src/home-cli.js +1 -1
  62. package/src/lib/cli-flags.js +3 -3
  63. package/src/lib/config.js +2 -0
  64. package/src/lib/connector-copies.js +288 -0
  65. package/src/lib/model-family.js +252 -0
  66. package/src/lib/pool-labels.js +109 -0
  67. package/src/lib/providers.js +2 -0
  68. package/src/lib/reasoning.js +65 -7
  69. package/src/lib/strategy.js +628 -120
  70. package/src/provider-cli.js +63 -6
  71. package/src/providers/_schema.json +20 -1
  72. package/src/providers/claude-code/connector-history.json +330 -0
  73. package/src/providers/claude-code/connector.json +20 -5
  74. package/src/providers/claude-code/provider.mjs +124 -1
  75. package/src/providers/codex/connector-history.json +326 -0
  76. package/src/providers/codex/connector.json +53 -16
  77. package/src/providers/codex/provider.mjs +115 -0
  78. package/src/providers/echo/connector-history.json +142 -0
  79. package/src/providers/grok/connector-history.json +308 -0
  80. package/src/providers/grok/connector.json +65 -7
  81. package/src/setup.js +54 -58
  82. package/src/strategy-cli.js +277 -41
  83. package/src/strategy-dashboard.js +255 -85
  84. package/src/workflow/cli.js +11 -4
  85. package/src/workflow/dash-kit.js +55 -48
  86. package/src/workflow/dashboard.js +24 -1
  87. package/src/workflow/home-view.js +14 -56
  88. package/src/workflow/reprice.js +4 -2
  89. package/src/workflow/runs-cli.js +2 -1
  90. package/src/workflow/stat-kit.js +118 -20
  91. package/src/workflow/stats-model.js +1 -0
  92. package/src/workflow/stats-view.js +27 -32
  93. package/src/workflow/step-json.js +93 -0
  94. package/src/workflow/step-model.js +15 -6
  95. package/src/workflow/watch-cli.js +5 -0
  96. package/docs/public/favicon.svg +0 -7
  97. package/src/lib/release.js +0 -56
package/AGENTS.md CHANGED
@@ -69,7 +69,7 @@ is the canonical reference for the CLI surface.
69
69
  - Every verb must work non-interactively (no TTY). The interactive wizard is
70
70
  a human convenience, never a requirement.
71
71
  - Version single source: package.json. Release via
72
- `node bin/bullswarm.js release patch|minor|major`, then `git push` and
72
+ `npm run release -- patch|minor|major [--title "<headline>"]`, then `git push` and
73
73
  `git push --tags`
74
74
  — CI publishes through npm trusted publishing (OIDC), no tokens.
75
75
 
@@ -90,7 +90,8 @@ the authoring method is `skill/references/providers.md`.
90
90
  ## Releasing
91
91
 
92
92
  1. All tests green.
93
- 2. `node bin/bullswarm.js release patch` (creates commit + tag v*).
93
+ 2. `npm run release -- patch --title "<headline>"` (dates `## Unreleased`
94
+ in CHANGELOG.md, bumps package.json, creates commit + tag v*).
94
95
  3. `git push && git push --tags`.
95
96
  4. GitHub Actions publishes to npm via trusted publishing; verify with
96
97
  `npm view bullswarm version`.
package/CHANGELOG.md CHANGED
@@ -1,5 +1,87 @@
1
1
  # bullswarm changelog
2
2
 
3
+ ## Unreleased
4
+
5
+ ## 0.35.6 — set reasoning per tier and per model from the setup screen
6
+
7
+ - setup: the setup screen (`bullswarm setup`, `strategy tui`) can now set
8
+ reasoning. Opening a provider shows each effort tier's model and reasoning
9
+ level with who chose it; ←/→ on a tier steps auto → the levels that CLI
10
+ accepts → "CLI decides", and saves it for that provider as
11
+ `strategy set-rung --reasoning` would. Provider names are no longer cut off,
12
+ and model search now starts with `/`, so typing `f` or `q` in a search no
13
+ longer finishes setup. `strategy inventory --json` lists each provider's
14
+ `reasoningLevels`.
15
+
16
+ - strategy: reasoning can be set per model. When a tier has more than one
17
+ model on a provider, each can think at its own level:
18
+ `strategy set-reasoning --tier high --level max --pool claude-code --model
19
+ claude-fable-5-1 --yes`, `reasoning.models` in `strategy configure`, or the
20
+ indented model rows on the setup screen. It sits above the provider's level
21
+ and below `run --reasoning`; resolved levels report source
22
+ `strategy-model`, and each model in `strategy inventory --json` carries the
23
+ level it would run at.
24
+
25
+ - docs: the README keeps two screenshots and a plainer tone, and its quick
26
+ start is now a prompt you paste into your agent; the other dashboard screens
27
+ moved to a new gallery page on the docs site.
28
+
29
+ ## 0.35.5 — model choices come from each CLI, newest version wins, setup stops pinning tiers
30
+
31
+ - cli: the maintainer-only `release` verb is gone from the CLI and its help;
32
+ maintainers run `npm run release -- patch|minor|major`, which also dates
33
+ the `## Unreleased` changelog section.
34
+
35
+ - package: the README screenshots and brand images (`docs/public/`) are no
36
+ longer shipped in the npm package; the npm page loads them from GitHub.
37
+
38
+ - grok: every tier now suggests the newest Grok, and the tiers differ only by reasoning. This is an owner decision: Grok has one model line. One family now covers every plain `grok-N.M`, and the newest version wins. A fresh home suggests `grok-4.7` on all three tiers: high at the connector's `xhigh`, then `medium: grok-4.7 · high reasoning` and `low: grok-4.7 · medium reasoning`, both with the reason `one Grok line, lighter reasoning for lighter tiers`. Before, it suggested `grok-4.6` for high, `grok-4.5` for medium, and nothing for low, and it listed `grok-4.7` as unranked. A `grok-4.8` takes over every tier the day `grok models` lists it. `grok-4.7-build-fast` is the same model at twice the price, so it keeps its base model's rank but is never suggested; `strategy set-rung` still selects it. The tier-wide best pick, the Codex and Claude suggestions, and pace-based routing across pools do not change. grok-4.7 and grok-4.5 now have dated price rows from [xAI's models page](https://docs.x.ai/docs/models), and grok-4.6's row was re-checked there (all on 2026-09-23). build-fast stays unpriced: xAI's Pricing page lists its long-context rates as $6.00 / $1.50 / $18.00 per 1M tokens, which is not twice grok-4.7's. The fallback list (`knownModels`) now matches what `grok models` lists.
39
+
40
+ - strategy: `generationFallback` now also covers a tier that no family serves. If a connector opts such a tier in and no exact row ranks a model for it, the tier takes the best-ranked family's newest-generation model at the declared reasoning. That stand-in is the pool's rung, and it stands at rank 0 against other pools, so it is the tier-wide pick only when no other pool serves the tier. A tier may also declare `why`, a plain reason that replaces the generated one; `provider validate` rejects an empty one.
41
+
42
+ - strategy: applying recommendations no longer pins tiers. `strategy apply`, `refresh --apply`, `setup --yes --strategy`, the setup wizard, the TUI's apply key, and the daily auto-refresh now write only each pool's rungs (its model and reasoning per tier). Before, they also pinned every tier to one pool: a fresh `setup --yes --strategy` sent high, medium, and low to Codex even at 100% used, and the auto-refresh pinned them again every 24 hours, so pace-based routing never picked another plan. Now every dispatch picks its pool by spare quota, and `strategy routes` shows each tier moving between pools. `strategy.assignments` stays empty unless you pin a tier yourself. The apply JSON now reports `rungs`, `bestNow` (the tier-wide pick, for display), `unpinned`, and `keptPins`, in place of `applied`.
43
+
44
+ - strategy: pins are explicit only. `strategy assign` makes one, marked `source: "user"`, and `strategy clear-assignment` removes it. Apply and the auto-refresh never remove a pin you set, and apply now keeps the pin's model on that pool's rung, so the pinned tier runs the model you named instead of the pool's recommended one.
45
+
46
+ - strategy: if an earlier setup or apply pinned your tiers, the next `strategy apply` or automatic refresh removes those pins and lists them under `unpinned`. A pin is treated as written by apply when it still has the pool and model that apply recorded in the cached report. Pins you set or changed yourself stay (`keptPins`), as does any pin in a home where apply never ran. Run `bullswarm strategy show` to see which pins remain and what removes each one.
47
+
48
+ - strategy: the tier suggestion reads as a suggestion, not a pin: `high: codex/gpt-6-astra (best now; routing picks by spare quota)` in `strategy show`, `best now (not pins): …` after the wizard and `setup --yes --strategy`, and `Routing now · by spare quota at each dispatch, unless pinned` in the TUI, whose route lines end in `by spare quota` or `pinned to <pool>`. A pinned tier gets its own line in `strategy show`, `pinned to <pool>/<model> by you · high dispatches go there while it is available · strategy clear-assignment high removes it`. `strategy routes` gives each route a `pin` field.
49
+
50
+ - strategy: `strategy show` counts unranked models instead of listing them. It prints one line, such as `unranked: 312 models (command-code 180, opencode 130, grok 2) · never recommended · strategy show --json lists them`, and the setup review screen prints the same count for each provider. `strategy show --json` keeps the full `unranked` list, and the model pickers keep their `(unranked)` tag.
51
+
52
+ - strategy: medium now falls back to the newest generation. When the family that serves a tier has no model in a pool's newest generation, a connector that opts the tier in (`generationFallback`) gives it the next-lower family's newest-generation model, at a declared reasoning level clamped to what that model supports. Codex opts medium in, and this is an owner decision, not a benchmark: 168 medium dispatches as luna at max reasoning ran 99% ok. So a fresh home suggests `medium: gpt-6-luna · max reasoning — no gpt-6 terra yet, newest generation preferred`, and medium returns to terra at the normal medium reasoning once a gpt-6 terra is discovered. High, low, and Claude are unchanged. `strategy apply`, `refresh --apply`, `setup --yes --strategy`, the wizard, and the TUI all write that level into the pool's rung, reported as `max (recommendation)`. It is never written over a level you set, and a later apply removes it when the fallback ends. `strategy show` and the setup screens say why. The cross-pool tier pick also skips a model disabled for its pool now, as each pool's own pick already did, so a tier suggestion can no longer name a model you turned off.
53
+
54
+ - setup, doctor: old copies of packaged connectors in `<home>/connectors/` now follow the package. Those copies were made by installs older than 0.29.0. A copy of a pool the package defines was already ignored, so its old prices (such as Opus 5's on Opus 5.5) never applied; a contrib copy whose provider is not enabled was still used. Every verb now checks each copy against `connector-history.json`, which holds fingerprints of every value each field has shipped with. A copy with no edits moves to `connectors/retired/` so the packaged connector loads, and a contrib provider is enabled first when only the copy defined its pool. A copy you edited is kept, and `bullswarm doctor` (`connector-copies`, a `!` warning) and `strategy show` name its stale fields. The dead copy-to-home step in `setup` is gone.
55
+
56
+ - strategy: connectors now declare model families (`modelFamilies`: a name pattern with a tier and base rank), and the version read from the model id orders each family. So a newly discovered model is ranked on the day it appears: gpt-6-sol and gpt-6-luna are no longer tierless, and a fresh home suggests gpt-6-luna for low instead of gpt-5.6-luna. Inside a pool and family, the newer version always wins, whatever benchmarks, prices, or pace say. Quality is compared on one scale, the rank, and benchmarks only break equal ranks. A model no rule classifies is listed as `unranked` and never recommended. Claude Code prices now follow the rate card line for each model: Opus 5.5 ($4/$20) and Fable 5.1 (cache hits $0.25) no longer inherit Opus 5's or Fable 5's price, and `provider validate` refuses a price on a family.
57
+
58
+ - strategy: Claude Code and Codex model choices now come from bounded, no-prompt installed-CLI handshakes per account; Claude discovery suppresses hooks and preserves explicit 1M selectors, Codex pages `model/list` and records per-model reasoning support, and the static lists remain only as failure fallbacks.
59
+
60
+ - docs: a new README with a banner, the problem, a comparison with the
61
+ alternatives, and dashboard screenshots taken from a made-up demo home;
62
+ a new playbook guide; `scripts/render-readme-shots.sh` rebuilds the
63
+ screenshots and refuses to write them if private terms or source IDs
64
+ survive.
65
+
66
+ - stats: hyphenated project names keep their complete basename when the panel
67
+ has room instead of being reduced to the suffix after the final hyphen.
68
+
69
+ - stats: charts share Home's even whole-unit ticks, zero baseline and daily
70
+ weekday labels, while compact money cells use the dashboard-wide source and
71
+ lower-bound glyphs and return the saved width to pool and project names.
72
+
73
+ - mod: the Step pane follows, selects, and expands the newest turn by default,
74
+ refreshes an open running Step every four seconds, and ticks its projected
75
+ in-flight command name and elapsed time every second without refetching.
76
+ - step: `workflow action show --json` and `workflow task show --json` now emit
77
+ compact Step schema v2: one activity per distinct attempt, one canonical
78
+ event array inside each activity, reference markers for aliases, and the
79
+ unchanged display projection used by the dashboard and Claude mod.
80
+ - pools: per-home display labels (`bullswarm pools label`) now shorten pool ids
81
+ across human CLI output, progress, the dashboard, and the Claude Mod without
82
+ changing credential, routing, meter, history, or workflow-record keys. Pool
83
+ arguments accept either form; JSON retains `pool` and adds `poolLabel`.
84
+
3
85
  ## 0.35.4 — one Runs table for workflows and tasks
4
86
 
5
87
  - mod: a live single task no longer replaces the selected workflow; each task
@@ -172,7 +254,7 @@
172
254
  sessions read-only; Command Code reads persisted project JSONL only when a
173
255
  session transcript exists and carries usage.
174
256
  - reprice: transcript lookup follows the loaded provider registry, so
175
- `opencode2*`, `opencode2:kaihk-*`, and `command-code` pools reach their
257
+ `opencode2*`, `opencode2:orbit-*`, and `command-code` pools reach their
176
258
  provider-owned readers without a hard-coded provider list. Ambiguous,
177
259
  missing, and checkpoint-only records stay unknown rather than becoming
178
260
  zero-cost attempts.
@@ -180,7 +262,7 @@
180
262
  `--no-session`, so their checkpoints are documented as non-recoverable
181
263
  history.
182
264
  - pricing: public model cards are retained only with a source and date (the
183
- cards were checked 2026-09-20). `kaihk/*` and `opencode/union-alpha` relay
265
+ cards were checked 2026-09-20). `orbit/*` and `opencode/union-alpha` relay
184
266
  identifiers have no public card, so the underlying OpenAI card is not
185
267
  substituted; observed Command Code models use the cited Command Code card.
186
268
  - evidence: live OpenCode and Command Code streams captured on 2026-09-20 are
package/README.md CHANGED
@@ -1,273 +1,120 @@
1
- # bullswarm
2
-
3
- Bullswarm is a CLI that sends a coding task to whichever of your installed
4
- agent CLIs — Claude Code, Codex, Grok, OpenCode, or Command Code — currently
5
- has unused subscription quota, then checks the result by its content.
6
-
7
- One command runs one task. A second command runs a graph of dependent tasks
8
- across those same CLIs. Every result is judged by what was actually written,
9
- not by whether the process exited 0.
10
-
11
- For an agent already working in a repo, the packaged `bullswarm` skill
12
- (`/bullswarm`, or `$bullswarm` where a skill uses that syntax) goes straight
13
- to `bullswarm run` or `bullswarm workflow goal`. A skill here is a short
14
- instruction file the agent CLI loads. There is no separate preview or
15
- classifier command to learn first.
16
-
17
- ## Why it exists
18
-
19
- Subscription quota expires on a clock, whether you spend it or not, and a
20
- single coding agent's judgment on whether its own work is done should not be
21
- the only check in the loop. Bullswarm picks whichever installed agent CLI has
22
- the most unused quota right now, and treats every delegate's output as
23
- evidence to be verified — never as an authority to be trusted on its word.
24
-
25
- The same problem compounds on multi-step goals: one agent planning and
26
- executing everything serially leaves every other installed CLI's quota idle,
27
- and having that same agent be the sole judge of whether the whole goal is
28
- done multiplies the risk instead of dividing it. Bullswarm's workflow engine
29
- runs a graph of dependent actions across whichever pools have quota to spare
30
- — a pool is one installed agent CLI, or one account of that CLI — and
31
- computes completion from evidence the graph itself required, not from any one
32
- delegate's own say-so.
33
-
34
- ## Install
35
-
36
- ```bash
37
- npm i -g bullswarm
38
- bullswarm setup
39
- ```
1
+ <p align="center">
2
+ <img src="docs/public/brand/bullswarm-banner.jpg" alt="Bullswarm: a fleet of bull agents in tuxedos wearing Claude, OpenAI and Grok pins" width="100%">
3
+ </p>
4
+
5
+ <p align="center">
6
+ <a href="https://www.npmjs.com/package/bullswarm"><img alt="npm version" src="https://img.shields.io/npm/v/bullswarm"></a>
7
+ <a href="LICENSE"><img alt="MIT license" src="https://img.shields.io/npm/l/bullswarm"></a>
8
+ <a href="https://nodejs.org/"><img alt="Node.js 22.12 or later" src="https://img.shields.io/badge/node-%3E%3D22.12-339933?logo=node.js&logoColor=white"></a>
9
+ </p>
10
+
11
+ <p align="center">Route work across your coding-agent subscriptions, spend quota before it expires, and verify what comes back.</p>
12
+
13
+ ## The problem
14
+
15
+ You pay for more than one coding agent: a Claude plan or two, Codex, Grok, maybe another.
16
+
17
+ Every one of them meters you on a clock — five-hour windows, weekly caps, monthly allowances — and unused quota is simply gone when the window resets.
18
+
19
+ So one plan runs dry mid-task while the others sit idle. You switch tools by hand just to spend what you already paid for. And the agent that wrote the code is usually the one that decides it is done.
20
+
21
+ ## What Bullswarm does
40
22
 
41
- Later, `bullswarm update` upgrades that install to the latest published
42
- version in place (`bullswarm update --check` only reports). Requires Node.js
43
- 22.12 or later. `bullswarm setup` walks through detecting your
44
- installed agent CLIs, showing their quota state, and writing a routing
45
- configuration. See
46
- [Getting started](https://bulls-work.github.io/bullswarm/guide/getting-started)
47
- for integrating Bullswarm's skill into Codex, Claude, and Grok, and for the
48
- full quick-start command list.
49
-
50
- ## Claude Mod (early access)
51
-
52
- `mods/bullswarm` is the same routing injected into Claude Code's own engine
53
- as a Claude Mod (a plugin of TypeScript function hooks, behind
54
- `CLAUDE_CODE_ENABLE_FUNCTION_HOOKS=1`): the pool meters drawn above the
55
- prompt, the pools named in the model's context, Claude's general-purpose
56
- subagents routed to whichever pool has surplus and answered with the
57
- verified output, and the verdict appended to every `bullswarm run` the model
58
- runs. The Mod is the dashboard's read-only counterpart: its pane shows Run,
59
- Step, and the Usage/Pools view in the same meter colours. It does not expose
60
- the dashboard's Home, Runs, Budget, Stats, Fleet, or Help pages, and
61
- it has no edit or install action. See [mods/bullswarm/README.md](mods/bullswarm/README.md).
62
-
63
- Three ways to load it:
64
-
65
- ```bash
66
- # with the CLI: links the mod under ~/.claude/skills and sets the flag in ~/.claude/settings.json
67
- bullswarm integrate install --agents claude --yes
68
-
69
- # from Claude Code's plugin marketplace (a copy that `claude plugin update` refreshes)
70
- claude plugin marketplace add Bulls-Work/bullswarm
71
- claude plugin install bullswarm@bullswarm
72
-
73
- # one session only, from the installed package
74
- CLAUDE_CODE_ENABLE_FUNCTION_HOOKS=1 claude --plugin-dir "$(npm root -g)/bullswarm/mods/bullswarm"
23
+ Bullswarm turns the agent CLIs already signed in on your machine into **one fleet** that your main agent can command.
24
+
25
+ ```mermaid
26
+ flowchart LR
27
+ you([You]) -->|goal| main["Your main agent<br/>Claude Code · Codex · Grok<br/>+ /bullswarm skill"]
28
+ main -->|"bullswarm run · workflow goal"| bs{{"Bullswarm<br/>pace router + workflow kernel"}}
29
+ bs --> c["claude -p"]
30
+ bs --> x["codex exec"]
31
+ bs --> g["grok -p"]
32
+ bs --> o["contributed providers"]
33
+ c & x & g & o -->|output| v["independent verification"]
34
+ v -->|"verdict + evidence"| main
75
35
  ```
76
36
 
77
- The marketplace route still needs `CLAUDE_CODE_ENABLE_FUNCTION_HOOKS=1`, in
78
- the shell or under `env` in `~/.claude/settings.json`, and the `bullswarm`
79
- CLI on `PATH`. Pick one route: an installed marketplace copy takes precedence
80
- over the skills-dir link, and Claude says so at startup.
37
+ - **Spends quota by pace.** Each task goes to the plan with the most spare quota for how far its window has run, so a plan that is behind, or about to reset with quota left, gets used first.
38
+ - **Picks models for you.** Tasks come in three lanes (`analyze`, `build`, `chore`). Setup asks each CLI which models it offers and suggests the newest one for each effort level, so you don't have to update settings every time a vendor ships a model.
39
+ - **Runs one task or a whole workflow.** `bullswarm run` sends one task to one agent and returns a verdict. For bigger goals your main agent writes a plan; Bullswarm runs the independent steps in parallel across agents, then integration and a final check by a different agent.
40
+ - **Checks the work, not the exit code.** A delegate saying "done" isn't enough. Bullswarm reads what it actually produced, and a workflow finishing is kept separate from its requirements being verified.
41
+ - **Shows everything live.** A terminal dashboard covers quota, history, running workflows and each agent's individual turns.
42
+ - **Works with the agent you already use.** The `/bullswarm` skill teaches Claude Code, Codex and Grok when to delegate. There's also an early-access Claude Code Mod that shows runs and usage inside Claude Code.
43
+ - **Lets you steer mid-run.** Add, change, remove or rerun steps while a workflow is running, or pause and resume it.
44
+ - **Extends with providers.** Claude Code, Codex and Grok are built in. Other CLIs can be added as providers without touching the core.
81
45
 
82
- ## Quick start
46
+ ## See it
83
47
 
84
- Once setup is complete, bare `bullswarm` opens the dashboard on its Home page.
85
- Use `bullswarm --setup` or `bullswarm setup` to open setup again, and use
86
- `bullswarm workflow tui` when you want the explicit dashboard command. The
87
- eight pages answer different questions:
48
+ ![The Bullswarm Home dashboard: today's work, quota, spend and recent runs](docs/public/screens/home.png)
88
49
 
89
- | Page | What it answers |
90
- |---|---|
91
- | Home | What happened today, what is verified, what measured or labelled money/licence data exists, and what is active or recent; includes pool/model/project breakdowns. |
92
- | Runs | Which workflows are active or historical — one `active` block and the History day table below it — plus the product commands; the integration line and `/` filter live here, and `?` carries the agents and `run it` blocks. |
93
- | Run | Where one workflow is in its plan, which workers are live or next, and its ETA and budget shares. |
94
- | Step | What one action is doing: route, attempt, activity, verdict/failure, prompt, usage, events, output, and artifacts. |
95
- | Budget | Pool quota windows, measured worker-minutes versus rest, measured or `≈` money, fit, and biggest workflows. |
96
- | Stats | Runs, spend, worker-minutes, and verification over 7d/30d/all, by pool, model, and project. |
97
- | Fleet | Lane/provider model and reasoning rungs, records, meter state, and the setup edit hand-off. |
98
- | Help | Every key, click, layout rule, and dashboard command. |
99
-
100
- The visual-fidelity pass keeps the same real numbers while composing and
101
- colouring these pages like the approved prototype. At 55 columns every Home
102
- today tile and per-step bar, every Stats model, project and its sparkline, every
103
- Stats and Fleet pool row, and every run row in the Runs history table keeps one
104
- row per item on the phone; Budget's per-pool block (a header over `used`,
105
- `by bullswarm`, `room` and `so far`), Home's `budget · this week` pool (meter
106
- plus a reset line, as the prototype draws it) and the Runs `active` entry are
107
- the deliberate exceptions. At 200 columns every page composes to the frame
108
- rather than capping at the 120-column composition. Estimated figures still
109
- carry `≈` and their basis; a figure that cannot be measured is a line of words
110
- saying so — no page draws an empty or dotted track for missing data.
111
-
112
- Every page has a sticky header, a page tab row, and a sticky bottom nav. The
113
- tab row is five tabs — Home, Runs, Budget, Stats, Fleet — with Run and Step
114
- marking Runs, and Help shown only while it is open. The shared key table is:
115
-
116
- | Key | Does |
50
+ *Home: what's running today, how much quota each plan has left, and what you've spent.*
51
+
52
+ ![A workflow's Run page: phases, steps, agents and cost](docs/public/screens/run.png)
53
+
54
+ *Run: one workflow's plan, which agent took each step, and what it cost.*
55
+
56
+ More screens, including phone-sized ones, are in the [gallery](https://bulls-work.github.io/bullswarm/guide/gallery).
57
+
58
+ ## Why not just…
59
+
60
+ | Approach | What you give up |
117
61
  |---|---|
118
- | `r` | open Runs; the view refreshes itself, so `r` is no longer refresh |
119
- | `b` | open Budget; it is no longer move-out |
120
- | `s` | open Stats |
121
- | `y` | open Runs at its History table (the first day header), except that it confirms a pending stop |
122
- | `f` | open Fleet |
123
- | `h` | open Home |
124
- | `?` | open Help |
125
- | `1`–`9` | open that run from the nav |
126
- | `Tab` | cycle the current page's sub-tabs |
127
- | `Shift+Tab` | cycle workflows |
128
- | `p` | cycle the period on Home and Stats |
129
- | `Esc` / `←` | move out one page, then Home |
130
- | `↑`/`k`, `↓`/`j` | move one line |
131
- | `Enter` / `→` / `l` | open the selected run, step, tab, or action |
132
- | `PgUp` / `PgDn` | scroll one screen |
133
- | `Home` / `End` | jump to the top or bottom |
134
- | `ctrl+s` | copy the screen through OSC 52, falling back to `pbcopy`, `wl-copy`, or `xclip` |
135
- | `q` | quit the dashboard; workflows keep running |
136
-
137
- The 0.33.0 rebinding is deliberate: `r` no longer refreshes, `b` no longer
138
- moves out, and `Tab` no longer cycles workflows. Esc/left moves out, the view
139
- refreshes itself, and Shift+Tab still cycles workflows. Runs also provides `/`
140
- (filter), `a` (active/all), and `i` (install); Run provides `o`, `v`, and `t`,
141
- and Fleet provides `e` for setup. Click tabs, tiles, bars, runs, steps, dates,
142
- or controls, or use the wheel to scroll.
143
-
144
- Home, Runs, Budget, and Stats read a per-run rollup that every finishing
145
- run appends to `~/.bullswarm/history/runs.jsonl`. After upgrading, backfill
146
- the runs that finished before 0.33.0 once with `bullswarm workflow reindex`;
147
- legacy runs get a minimal record and only runs still in flight are skipped.
148
-
149
- One bounded outcome — a task with a clear finish line:
150
-
151
- ```bash
152
- bullswarm run --lane analyze --add-dir ~/some-repo --prompt "Explain the parser" --json
153
- ```
62
+ | **One agent's built-in subagents or workflows** | Everything draws on that one plan's quota, the same vendor grades its own work, and your process is tied to that vendor's feature. |
63
+ | **An API gateway that re-exposes your subscriptions** | Turning a consumer subscription into a generic API endpoint can conflict with provider terms, and you lose each agent's own tools and harness. |
64
+ | **Switching tools by hand** | You become the scheduler, and the quota you did not get to still expires. |
65
+ | **Bullswarm** | Drives each vendor's own headless CLI—`claude -p`, `codex exec`, `grok -p`—the way those CLIs are meant to be scripted, with the accounts you already signed in. Work lands where quota is spare, and a different agent checks it. |
154
66
 
155
- `--lane analyze` tags the work as read-only analysis. The other lanes are
156
- `build` (edits) and `chore` (mechanical edits). This routes the prompt to
157
- whichever pool is eligible, dispatches it, watches it to completion, verifies
158
- the output, and prints one JSON verdict — nothing else runs and nothing is
159
- left in the background.
160
-
161
- Multi-step work, where you author the plan. The kernel — Bullswarm's own
162
- runtime, not an agent — validates that program and executes it:
163
-
164
- ```json
165
- {
166
- "schemaVersion": "bullswarm.workflow.program.v2",
167
- "actions": [
168
- { "id": "fix", "kind": "implement", "purpose": "Fix the failing tests",
169
- "dependsOn": [], "ownedFiles": ["src/parser.js"], "affects": ["requirement-1"],
170
- "evidenceFor": [], "prompt": "In ~/some-repo, fix the failing tests with the smallest correct change." }
171
- ]
172
- }
173
- ```
67
+ Bullswarm never proxies a subscription as an API and never collects or shares your vendor credentials.
174
68
 
175
- ```bash
176
- # 1. Write plan.json — the bounded action program the kernel will enforce.
177
- bullswarm workflow plan validate "Fix the failing tests and verify the change" \
178
- --cwd ~/some-repo --program plan.json --json # exit 0 valid, exit 2 with the issues; nothing launches
179
- bullswarm workflow goal "Fix the failing tests and verify the change" \
180
- --cwd ~/some-repo --program plan.json # launches, prints a short ID and observation commands, returns
181
- ```
69
+ The controller is portable. Claude Code, Codex, Grok, or any capable caller drives the same CLI and durable workflow kernel, so your workflows are not locked into one vendor's orchestration.
70
+
71
+ ## Why now
72
+
73
+ Agent plans now come with hard weekly limits, and serious work spans hours rather than prompts. A workflow can run for one to two hours across several agents—builders in parallel, an integrator, then independent acceptance—while you keep handing new goals to your main agent.
74
+
75
+ With Bullswarm routing by pace, running several goals at once no longer means worrying about wasting one plan's scarce quota: the fleet spends whichever plan is furthest behind.
182
76
 
183
- `workflow goal` starts a durable background run and returns immediately by
184
- default; add `--watch` to follow its low-noise progress in the same terminal
185
- instead.
186
-
187
- ## How it picks a pool
188
-
189
- - Work is tagged by lane — read-only analysis, ordinary build work, or
190
- mechanical chores — not assigned to a fixed pool ahead of time.
191
- - Among the pools that can do the work, the one furthest behind its own quota
192
- pace (the most unspent surplus) wins, so quota doesn't expire unused.
193
- - A pool close to its rolling 5-hour usage ceiling gives way to one with
194
- headroom only while another pool is actually behind its own pace — it is an
195
- ordering penalty, not a cutoff. A step may run a pool all the way to 100% of
196
- that window, because if the provider stops the worker at the wall the retry
197
- is briefed on what it had already written.
198
- - A pool whose weekly or monthly subscription window is about to reset gets
199
- priority for its remaining surplus, so quota doesn't run out the clock
200
- unspent.
201
- - A pool already busy with other in-flight work yields to a quieter pool at a
202
- similar pace, so a burst of parallel work spreads out instead of piling onto
203
- one pool.
204
- - A pool that reports a usage-limit error is benched until the provider's own
205
- reset time and automatically re-tried after that — never left down for good,
206
- and never retried early.
207
-
208
- The full mechanics behind each of these are in
209
- [Routing](https://bulls-work.github.io/bullswarm/guide/routing).
210
-
211
- ## What you get back
212
-
213
- `bullswarm run` prints a JSON verdict when it finishes:
214
-
215
- - `keepOnClaude: true` — the router says do this in-session; nothing ran
216
- - `ok: true` (and `keepOnClaude` is false) — the output passed verification;
217
- read `outFile`
218
- - `ok: false` — `why` names the gate that failed
219
- - `contentUsableDespiteExit: true` — the process exited non-zero but the
220
- content still verified; read it before re-running
221
-
222
- A non-zero exit from the delegate is never treated as success on its own. See
223
- [Result envelope](https://bulls-work.github.io/bullswarm/reference/result) for the full
224
- verdict shape.
225
-
226
- A workflow produces a durable, versioned result envelope — a JSON document
227
- with `runId` / `shortId`, status, per-requirement evidence, per-action
228
- outcomes, and usage — instead of leaving you to parse a transcript.
229
-
230
- ```bash
231
- bullswarm workflow watch <shortId> --next # wait for the next notable event, then exit
232
- bullswarm workflow runs result <shortId> --json --summary # compact status once the run is terminal
77
+ ## Quick start
78
+
79
+ You need Node.js 22.12 or later and at least one agent CLI you're signed in to (Claude Code, Codex or Grok).
80
+
81
+ The easiest way to set up is to let your agent do it. Paste this into Claude Code, Codex or Grok:
82
+
83
+ ```text
84
+ Install and set up Bullswarm for me:
85
+ 1. Run `npm i -g bullswarm`.
86
+ 2. Run `bullswarm setup --yes --strategy --integrate` to find my agent CLIs,
87
+ pick models, and install the /bullswarm skill for each agent.
88
+ 3. Run `bullswarm doctor` and tell me which agents are ready and how much
89
+ quota each one has left.
90
+ 4. Read the installed /bullswarm skill so you know when to use
91
+ `bullswarm run` and when to write a workflow.
233
92
  ```
234
93
 
235
- `workflow watch --next` prints one line per event and relaunches itself with
236
- the exact flags to keep polling; `runs result --summary` is what to read once
237
- a run finishes, and `runs result --json` (no `--summary`) gives the full
238
- envelope for a failed or partial run. See
239
- [Observing runs](https://bulls-work.github.io/bullswarm/guide/observing) for the
240
- full shape of both.
94
+ After that, just give your agent goals as usual. It will hand work to Bullswarm when that helps. Run `bullswarm` in a terminal to open the dashboard; closing it doesn't stop anything that's running.
241
95
 
242
- ## Documentation
96
+ To set things up by hand, or to choose models and reasoning yourself, see [Getting started](https://bulls-work.github.io/bullswarm/guide/getting-started).
243
97
 
244
- The full documentation is published at
245
- [bulls-work.github.io/bullswarm](https://bulls-work.github.io/bullswarm/).
98
+ ## Supported agents and meters
246
99
 
247
- The dashboard is the main screen after setup: bare `bullswarm` opens Home,
248
- `bullswarm --setup` or `bullswarm setup` opens setup, and
249
- `bullswarm workflow tui` is the explicit form. The [Observing runs](https://bulls-work.github.io/bullswarm/guide/observing)
250
- page maps its pages, keys, and mouse controls.
100
+ | Agent CLI | Support | Quota windows tracked | Command it runs |
101
+ |---|---|---|---|
102
+ | Claude Code | Built in | weekly + 5-hour | `claude -p` |
103
+ | Codex | Built in | weekly | `codex exec` |
104
+ | Grok | Built in | weekly | `grok -p` |
105
+ | OpenCode | Contributed | none | `opencode run --auto` |
106
+ | Command Code | Contributed | weekly + monthly + 5-hour | `command-code -p` |
251
107
 
252
- | Page | What it covers |
253
- |---|---|
254
- | [Introduction](https://bulls-work.github.io/bullswarm/guide/) | What Bullswarm is, the two entry points, and the four rules it never breaks |
255
- | [Getting started](https://bulls-work.github.io/bullswarm/guide/getting-started) | Install, `setup`, `doctor`, agent integration, and your first verified run |
256
- | [Concepts](https://bulls-work.github.io/bullswarm/guide/concepts) | Pools, lanes, surplus, the two windows, verdicts, quarantine, and the run directory |
257
- | [Run one task](https://bulls-work.github.io/bullswarm/guide/run) | Every `bullswarm run` option, and what each verdict asks you to do |
258
- | [Workflows](https://bulls-work.github.io/bullswarm/guide/workflows) | Authoring the program `workflow goal` executes: territories, dependencies, integration, acceptance |
259
- | [Observing runs](https://bulls-work.github.io/bullswarm/guide/observing) | `workflow watch`, the Home/Runs/Run/Step/Budget/Stats/Fleet/Help dashboard pages, keys, mouse, and terminal glyphs |
260
- | [Routing](https://bulls-work.github.io/bullswarm/guide/routing) | How a pool is picked: pace, 5-hour headroom, urgency, load, quarantine |
261
- | [CLI reference](https://bulls-work.github.io/bullswarm/reference/cli) | Every verb and nested subcommand, with its flags and defaults |
262
- | [Workflow program](https://bulls-work.github.io/bullswarm/reference/program) | The `bullswarm.workflow.program.v2` document: action fields, kinds, validation rules |
263
- | [Configuration](https://bulls-work.github.io/bullswarm/reference/configuration) | The Bullswarm home, `state.json`, strategy models and rungs, environment variables |
264
- | [Providers](https://bulls-work.github.io/bullswarm/reference/providers) | Adding your own agent CLI or reseller account as a provider plugin |
265
- | [Result envelope](https://bulls-work.github.io/bullswarm/reference/result) | Every field of `run --json` and of the workflow result document |
266
- | [Claude Code](https://bulls-work.github.io/bullswarm/integrations/claude-code) | The packaged skill, the MCP server, and the read-only Claude Mod counterpart under `mods/bullswarm` |
267
- | [Codex and Grok](https://bulls-work.github.io/bullswarm/integrations/agent-clis) | What `bullswarm integrate` writes for each agent CLI, and how to check it |
268
- | [Issue watcher](https://bulls-work.github.io/bullswarm/integrations/issue-watcher) | The launchd agent that triages and fixes new GitHub issues |
269
- | [Historical notes](https://bulls-work.github.io/bullswarm/notes/) | Working notes, audits, and experiment writeups, kept as records |
270
-
271
- ## License
272
-
273
- MIT
108
+ Each provider describes how to launch its CLI, read its quota, list its models and read its output. To add one, see [Adding a provider](https://bulls-work.github.io/bullswarm/reference/providers).
109
+
110
+ ## Learn more
111
+
112
+ - [Documentation](https://bulls-work.github.io/bullswarm/)
113
+ - [Getting started](https://bulls-work.github.io/bullswarm/guide/getting-started)
114
+ - [Day-to-day playbook](https://bulls-work.github.io/bullswarm/guide/playbook)
115
+ - [How routing works](https://bulls-work.github.io/bullswarm/guide/routing)
116
+ - [Authoring workflows](https://bulls-work.github.io/bullswarm/guide/workflows)
117
+ - [The dashboard](https://bulls-work.github.io/bullswarm/guide/gallery) and [observing runs](https://bulls-work.github.io/bullswarm/guide/observing)
118
+ - [Provider reference](https://bulls-work.github.io/bullswarm/reference/providers)
119
+ - Contributing: [open an issue](https://github.com/Bulls-Work/bullswarm/issues) or [submit a pull request](https://github.com/Bulls-Work/bullswarm/pulls)
120
+ - [MIT licence](LICENSE)