bullswarm 0.35.5 → 0.35.6
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +24 -0
- package/README.md +42 -107
- package/data/openrouter-benchmarks.json +7812 -7282
- package/docs/design/redesign-mechanics-principles-options.md +385 -0
- package/docs/guide/gallery.md +57 -0
- package/docs/index.md +4 -1
- package/docs/reference/cli.md +1 -1
- package/docs/reference/configuration.md +2 -1
- package/package.json +1 -1
- package/src/help.js +10 -5
- package/src/lib/cli-flags.js +2 -2
- package/src/lib/reasoning.js +15 -3
- package/src/lib/strategy.js +60 -9
- package/src/strategy-cli.js +62 -8
- package/src/strategy-dashboard.js +224 -75
package/CHANGELOG.md
CHANGED
|
@@ -2,6 +2,30 @@
|
|
|
2
2
|
|
|
3
3
|
## Unreleased
|
|
4
4
|
|
|
5
|
+
## 0.35.6 — set reasoning per tier and per model from the setup screen
|
|
6
|
+
|
|
7
|
+
- setup: the setup screen (`bullswarm setup`, `strategy tui`) can now set
|
|
8
|
+
reasoning. Opening a provider shows each effort tier's model and reasoning
|
|
9
|
+
level with who chose it; ←/→ on a tier steps auto → the levels that CLI
|
|
10
|
+
accepts → "CLI decides", and saves it for that provider as
|
|
11
|
+
`strategy set-rung --reasoning` would. Provider names are no longer cut off,
|
|
12
|
+
and model search now starts with `/`, so typing `f` or `q` in a search no
|
|
13
|
+
longer finishes setup. `strategy inventory --json` lists each provider's
|
|
14
|
+
`reasoningLevels`.
|
|
15
|
+
|
|
16
|
+
- strategy: reasoning can be set per model. When a tier has more than one
|
|
17
|
+
model on a provider, each can think at its own level:
|
|
18
|
+
`strategy set-reasoning --tier high --level max --pool claude-code --model
|
|
19
|
+
claude-fable-5-1 --yes`, `reasoning.models` in `strategy configure`, or the
|
|
20
|
+
indented model rows on the setup screen. It sits above the provider's level
|
|
21
|
+
and below `run --reasoning`; resolved levels report source
|
|
22
|
+
`strategy-model`, and each model in `strategy inventory --json` carries the
|
|
23
|
+
level it would run at.
|
|
24
|
+
|
|
25
|
+
- docs: the README keeps two screenshots and a plainer tone, and its quick
|
|
26
|
+
start is now a prompt you paste into your agent; the other dashboard screens
|
|
27
|
+
moved to a new gallery page on the docs site.
|
|
28
|
+
|
|
5
29
|
## 0.35.5 — model choices come from each CLI, newest version wins, setup stops pinning tiers
|
|
6
30
|
|
|
7
31
|
- cli: the maintainer-only `release` verb is gone from the CLI and its help;
|
package/README.md
CHANGED
|
@@ -1,11 +1,11 @@
|
|
|
1
1
|
<p align="center">
|
|
2
|
-
<img src="docs/public/brand/bullswarm-banner.
|
|
2
|
+
<img src="docs/public/brand/bullswarm-banner.jpg" alt="Bullswarm: a fleet of bull agents in tuxedos wearing Claude, OpenAI and Grok pins" width="100%">
|
|
3
3
|
</p>
|
|
4
4
|
|
|
5
5
|
<p align="center">
|
|
6
|
-
<img alt="npm version" src="https://img.shields.io/npm/v/bullswarm">
|
|
7
|
-
<img alt="MIT license" src="https://img.shields.io/npm/l/bullswarm">
|
|
8
|
-
<img alt="Node.js 22.12 or later" src="https://img.shields.io/badge/node-%3E%3D22.12-339933?logo=node.js&logoColor=white">
|
|
6
|
+
<a href="https://www.npmjs.com/package/bullswarm"><img alt="npm version" src="https://img.shields.io/npm/v/bullswarm"></a>
|
|
7
|
+
<a href="LICENSE"><img alt="MIT license" src="https://img.shields.io/npm/l/bullswarm"></a>
|
|
8
|
+
<a href="https://nodejs.org/"><img alt="Node.js 22.12 or later" src="https://img.shields.io/badge/node-%3E%3D22.12-339933?logo=node.js&logoColor=white"></a>
|
|
9
9
|
</p>
|
|
10
10
|
|
|
11
11
|
<p align="center">Route work across your coding-agent subscriptions, spend quota before it expires, and verify what comes back.</p>
|
|
@@ -34,16 +34,26 @@ flowchart LR
|
|
|
34
34
|
v -->|"verdict + evidence"| main
|
|
35
35
|
```
|
|
36
36
|
|
|
37
|
-
- **
|
|
38
|
-
- **
|
|
39
|
-
- **Runs one
|
|
40
|
-
- **
|
|
41
|
-
- **
|
|
42
|
-
- **
|
|
43
|
-
- **
|
|
44
|
-
- **
|
|
45
|
-
|
|
46
|
-
|
|
37
|
+
- **Spends quota by pace.** Each task goes to the plan with the most spare quota for how far its window has run, so a plan that is behind, or about to reset with quota left, gets used first.
|
|
38
|
+
- **Picks models for you.** Tasks come in three lanes (`analyze`, `build`, `chore`). Setup asks each CLI which models it offers and suggests the newest one for each effort level, so you don't have to update settings every time a vendor ships a model.
|
|
39
|
+
- **Runs one task or a whole workflow.** `bullswarm run` sends one task to one agent and returns a verdict. For bigger goals your main agent writes a plan; Bullswarm runs the independent steps in parallel across agents, then integration and a final check by a different agent.
|
|
40
|
+
- **Checks the work, not the exit code.** A delegate saying "done" isn't enough. Bullswarm reads what it actually produced, and a workflow finishing is kept separate from its requirements being verified.
|
|
41
|
+
- **Shows everything live.** A terminal dashboard covers quota, history, running workflows and each agent's individual turns.
|
|
42
|
+
- **Works with the agent you already use.** The `/bullswarm` skill teaches Claude Code, Codex and Grok when to delegate. There's also an early-access Claude Code Mod that shows runs and usage inside Claude Code.
|
|
43
|
+
- **Lets you steer mid-run.** Add, change, remove or rerun steps while a workflow is running, or pause and resume it.
|
|
44
|
+
- **Extends with providers.** Claude Code, Codex and Grok are built in. Other CLIs can be added as providers without touching the core.
|
|
45
|
+
|
|
46
|
+
## See it
|
|
47
|
+
|
|
48
|
+

|
|
49
|
+
|
|
50
|
+
*Home: what's running today, how much quota each plan has left, and what you've spent.*
|
|
51
|
+
|
|
52
|
+

|
|
53
|
+
|
|
54
|
+
*Run: one workflow's plan, which agent took each step, and what it cost.*
|
|
55
|
+
|
|
56
|
+
More screens, including phone-sized ones, are in the [gallery](https://bulls-work.github.io/bullswarm/guide/gallery).
|
|
47
57
|
|
|
48
58
|
## Why not just…
|
|
49
59
|
|
|
@@ -64,122 +74,47 @@ Agent plans now come with hard weekly limits, and serious work spans hours rathe
|
|
|
64
74
|
|
|
65
75
|
With Bullswarm routing by pace, running several goals at once no longer means worrying about wasting one plan's scarce quota: the fleet spends whichever plan is furthest behind.
|
|
66
76
|
|
|
67
|
-
## See it
|
|
68
|
-
|
|
69
|
-

|
|
70
|
-
|
|
71
|
-
*Home — today's work, budget position, trends, and recent runs in one view.*
|
|
72
|
-
|
|
73
|
-
| Runs | Run |
|
|
74
|
-
|---|---|
|
|
75
|
-
| [](docs/public/screens/runs.png) | [](docs/public/screens/run.png) |
|
|
76
|
-
| Active and historical workflows and single tasks in one table. | Phases, live workers, costs, and the attempt timeline. |
|
|
77
|
-
|
|
78
|
-
| Step | Stats |
|
|
79
|
-
|---|---|
|
|
80
|
-
| [](docs/public/screens/step.png) | [](docs/public/screens/stats.png) |
|
|
81
|
-
| The live or saved agent transcript, result, task, and cost evidence. | Workflow, spend, worker-time, and verification trends. |
|
|
82
|
-
|
|
83
|
-
[](docs/public/screens/budget.png)
|
|
84
|
-
|
|
85
|
-
*Budget — spend spare quota before each weekly or monthly window resets.*
|
|
86
|
-
|
|
87
|
-
<p align="center">
|
|
88
|
-
<img src="docs/public/screens/home-phone.png" alt="Bullswarm Home dashboard at phone width" width="360">
|
|
89
|
-
<img src="docs/public/screens/run-phone.png" alt="Bullswarm Run page at phone width" width="360">
|
|
90
|
-
</p>
|
|
91
|
-
|
|
92
|
-
*Home and Run retain their core evidence at a 55-column phone width.*
|
|
93
|
-
|
|
94
77
|
## Quick start
|
|
95
78
|
|
|
96
|
-
|
|
79
|
+
You need Node.js 22.12 or later and at least one agent CLI you're signed in to (Claude Code, Codex or Grok).
|
|
97
80
|
|
|
98
|
-
|
|
99
|
-
npm i -g bullswarm
|
|
100
|
-
bullswarm setup
|
|
101
|
-
```
|
|
102
|
-
|
|
103
|
-
`setup` discovers installed agent CLIs, shows their quota state, and opens the provider/model control centre. An agent or CI process can use discovered defaults without prompts:
|
|
104
|
-
|
|
105
|
-
```bash
|
|
106
|
-
bullswarm setup --yes --strategy --integrate
|
|
107
|
-
bullswarm doctor
|
|
108
|
-
```
|
|
109
|
-
|
|
110
|
-
Run one bounded task:
|
|
111
|
-
|
|
112
|
-
```bash
|
|
113
|
-
bullswarm run --lane analyze --add-dir . \
|
|
114
|
-
--prompt "List every TODO in src with file and line number." --json
|
|
115
|
-
```
|
|
116
|
-
|
|
117
|
-
Start a first workflow with an explicitly delegated planner:
|
|
118
|
-
|
|
119
|
-
```bash
|
|
120
|
-
bullswarm workflow goal \
|
|
121
|
-
"Audit this repository and write a one-page summary" \
|
|
122
|
-
--cwd . --orchestrator auto --watch
|
|
123
|
-
```
|
|
81
|
+
The easiest way to set up is to let your agent do it. Paste this into Claude Code, Codex or Grok:
|
|
124
82
|
|
|
125
|
-
|
|
126
|
-
|
|
127
|
-
|
|
128
|
-
bullswarm
|
|
129
|
-
|
|
130
|
-
|
|
131
|
-
|
|
132
|
-
|
|
133
|
-
|
|
134
|
-
|
|
135
|
-
bullswarm workflow goal \
|
|
136
|
-
"1. Fix the parser. 2. Add independent verification." \
|
|
137
|
-
--cwd . --program plan.json --watch
|
|
138
|
-
```
|
|
139
|
-
|
|
140
|
-
Open the dashboard at any time. Quitting it does not stop running workflows.
|
|
141
|
-
|
|
142
|
-
```bash
|
|
143
|
-
bullswarm
|
|
144
|
-
# explicit form:
|
|
145
|
-
bullswarm workflow tui
|
|
83
|
+
```text
|
|
84
|
+
Install and set up Bullswarm for me:
|
|
85
|
+
1. Run `npm i -g bullswarm`.
|
|
86
|
+
2. Run `bullswarm setup --yes --strategy --integrate` to find my agent CLIs,
|
|
87
|
+
pick models, and install the /bullswarm skill for each agent.
|
|
88
|
+
3. Run `bullswarm doctor` and tell me which agents are ready and how much
|
|
89
|
+
quota each one has left.
|
|
90
|
+
4. Read the installed /bullswarm skill so you know when to use
|
|
91
|
+
`bullswarm run` and when to write a workflow.
|
|
146
92
|
```
|
|
147
93
|
|
|
148
|
-
|
|
149
|
-
|
|
150
|
-
1. **Nail down the outcome.** State what must change, what must remain untouched, and what evidence will count as done.
|
|
151
|
-
2. **Hand it to your main agent.** With the `/bullswarm` skill installed, it chooses a bounded run or authors a workflow program with clear territories, dependencies, integration, and acceptance.
|
|
152
|
-
3. **Let the graph fan out.** Ready, file-disjoint actions can run across Codex, Grok, Claude accounts, and contributed providers while routing spends the quota furthest behind pace.
|
|
153
|
-
4. **Stay in control.** Follow the dashboard or `workflow watch`; send guidance, pause, or revise the live plan when the goal changes.
|
|
154
|
-
5. **Sign off on evidence.** Read the durable result envelope and requirement evidence. `completed` and `verified` answer different questions.
|
|
94
|
+
After that, just give your agent goals as usual. It will hand work to Bullswarm when that helps. Run `bullswarm` in a terminal to open the dashboard; closing it doesn't stop anything that's running.
|
|
155
95
|
|
|
156
|
-
|
|
96
|
+
To set things up by hand, or to choose models and reasoning yourself, see [Getting started](https://bulls-work.github.io/bullswarm/guide/getting-started).
|
|
157
97
|
|
|
158
98
|
## Supported agents and meters
|
|
159
99
|
|
|
160
|
-
| Agent CLI |
|
|
100
|
+
| Agent CLI | Support | Quota windows tracked | Command it runs |
|
|
161
101
|
|---|---|---|---|
|
|
162
102
|
| Claude Code | Built in | weekly + 5-hour | `claude -p` |
|
|
163
103
|
| Codex | Built in | weekly | `codex exec` |
|
|
164
104
|
| Grok | Built in | weekly | `grok -p` |
|
|
165
|
-
| OpenCode | Contributed | none
|
|
105
|
+
| OpenCode | Contributed | none | `opencode run --auto` |
|
|
166
106
|
| Command Code | Contributed | weekly + monthly + 5-hour | `command-code -p` |
|
|
167
107
|
|
|
168
|
-
|
|
169
|
-
|
|
170
|
-
## Status
|
|
171
|
-
|
|
172
|
-
Bullswarm is used daily. The maintainer's local records contained **335 workflow runs** and **507 single tasks** as of September 2026.
|
|
173
|
-
|
|
174
|
-
The routing, content verification, durable workflow kernel, dashboard, and built-in providers are established parts of the project. Mid-run steering and whole-plan revision are newer; use their validation and revision guards, and inspect the resulting evidence.
|
|
108
|
+
Each provider describes how to launch its CLI, read its quota, list its models and read its output. To add one, see [Adding a provider](https://bulls-work.github.io/bullswarm/reference/providers).
|
|
175
109
|
|
|
176
110
|
## Learn more
|
|
177
111
|
|
|
178
112
|
- [Documentation](https://bulls-work.github.io/bullswarm/)
|
|
179
113
|
- [Getting started](https://bulls-work.github.io/bullswarm/guide/getting-started)
|
|
114
|
+
- [Day-to-day playbook](https://bulls-work.github.io/bullswarm/guide/playbook)
|
|
180
115
|
- [How routing works](https://bulls-work.github.io/bullswarm/guide/routing)
|
|
181
116
|
- [Authoring workflows](https://bulls-work.github.io/bullswarm/guide/workflows)
|
|
182
|
-
- [
|
|
117
|
+
- [The dashboard](https://bulls-work.github.io/bullswarm/guide/gallery) and [observing runs](https://bulls-work.github.io/bullswarm/guide/observing)
|
|
183
118
|
- [Provider reference](https://bulls-work.github.io/bullswarm/reference/providers)
|
|
184
119
|
- Contributing: [open an issue](https://github.com/Bulls-Work/bullswarm/issues) or [submit a pull request](https://github.com/Bulls-Work/bullswarm/pulls)
|
|
185
120
|
- [MIT licence](LICENSE)
|