bullswarm 0.25.5 → 0.27.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/AGENTS.md +19 -25
- package/CHANGELOG.md +192 -0
- package/README.md +118 -151
- package/connectors/codex.json +3 -2
- package/connectors/command-code.json +229 -27
- package/data/README.md +139 -9
- package/data/epoch-benchmarks.json +6028 -0
- package/docs/studies/portal-token-diet.md +127 -0
- package/docs/workflow-design.md +5 -1
- package/package.json +4 -3
- package/skill/SKILL.md +28 -35
- package/skill/references/operations.md +65 -13
- package/src/cli.js +2 -4
- package/src/help.js +93 -460
- package/src/integrate.js +4 -11
- package/src/lib/epoch-benchmarks.js +394 -0
- package/src/lib/strategy.js +209 -2
- package/src/lib/watch.js +93 -20
- package/src/setup.js +76 -24
- package/src/strategy-cli.js +199 -3
- package/src/workflow/action-validator.js +117 -4
- package/src/workflow/cli.js +114 -284
- package/src/workflow/dashboard.js +153 -967
- package/src/workflow/goal.js +0 -36
- package/src/workflow/runs-cli.js +117 -131
- package/src/workflow/short-id.js +86 -127
- package/src/workflow/steering.js +6 -8
- package/src/workflow/v2-cancellation.js +6 -1
- package/src/workflow/v2-outcome.js +8 -1
- package/src/workflow/v2-planner.js +31 -5
- package/src/workflow/v2-runtime.js +1 -2
- package/src/workflow/v2-state.js +21 -2
- package/src/workflow/watch-cli.js +79 -133
- package/bin/check-output-schema.js +0 -30
- package/src/delegate.js +0 -438
- package/src/workflow/decision.js +0 -357
- package/src/workflow/draft-cli.js +0 -434
- package/src/workflow/draft.js +0 -268
- package/src/workflow/result.js +0 -239
- package/src/workflow/runner.js +0 -1304
- package/src/workflow/runtime.js +0 -1902
- package/src/workflow/schema.js +0 -81
- package/src/workflow/semaphore.js +0 -56
- package/src/workflow/template.js +0 -146
- package/src/workflow/tui.js +0 -297
- package/src/workflow/validate.js +0 -366
- package/workflows/adaptive-code-review.json +0 -45
- package/workflows/agent-model-comparison.json +0 -49
- package/workflows/connector-audit.json +0 -60
- package/workflows/smoke-two-step.json +0 -41
- package/workflows/verify-and-cap.json +0 -62
package/AGENTS.md
CHANGED
|
@@ -20,37 +20,31 @@ content. Published as `bullswarm` on npm.
|
|
|
20
20
|
5. Workflow dispatches must honor the same guarantees as single runs:
|
|
21
21
|
`BULLSWARM_DEPTH` is propagated, burst-gated pools are excluded, and
|
|
22
22
|
auth verdicts quarantine the pool + append to the shared decision log
|
|
23
|
-
(R6/R7/R8 in `src/workflow/
|
|
24
|
-
6. Adversarial verification is a first-class primitive:
|
|
25
|
-
|
|
26
|
-
|
|
27
|
-
|
|
28
|
-
|
|
29
|
-
stored under `~/.bullswarm/drafts/<name>/` and are runnable by name
|
|
30
|
-
without an upfront JSON. JSON is still the durable artifact — drafts
|
|
31
|
-
are JSON documents, just built one mutation at a time.
|
|
32
|
-
8. New goal workflows are caller-planned programs in a shared workspace.
|
|
23
|
+
(R6/R7/R8 in `src/workflow/v2-dispatch.js`).
|
|
24
|
+
6. Adversarial verification is a first-class primitive: an action naming
|
|
25
|
+
requirements in `evidenceFor` is dispatched under an evidence contract and
|
|
26
|
+
judges them from the durable artifact, so a requirement is only verified by
|
|
27
|
+
work someone else inspected (R-skeptic).
|
|
28
|
+
7. New goal workflows are caller-planned programs in a shared workspace.
|
|
33
29
|
`bullswarm workflow goal --program` executes the graph; `--orchestrator`
|
|
34
30
|
explicitly delegates planning. File territories are advisory scheduling
|
|
35
31
|
hints, and the graph finishes without automatic gap rounds. `verified`
|
|
36
32
|
separately records requirement evidence. `--isolation` opts into strict
|
|
37
33
|
per-worker worktrees. Saved V2 runs preserve their original semantics.
|
|
34
|
+
8. Historical authored-graph runs remain visible as read-only `legacy` rows.
|
|
35
|
+
Their executor was removed in 0.27.0; driving commands fail closed before
|
|
36
|
+
dispatch and historical run directories remain untouched.
|
|
38
37
|
|
|
39
38
|
## Development
|
|
40
39
|
|
|
41
40
|
```bash
|
|
42
41
|
npm test # full suite, no network needed (meters read from cache)
|
|
43
42
|
node bin/bullswarm.js doctor --json # readiness report
|
|
44
|
-
node bin/bullswarm.js workflow
|
|
43
|
+
node bin/bullswarm.js workflow goal "Fix the failing tests" --program plan.json
|
|
45
44
|
node bin/bullswarm.js workflow runs # ongoing workflow instances
|
|
46
45
|
node bin/bullswarm.js workflow runs --all # including historical
|
|
47
|
-
|
|
48
|
-
|
|
49
|
-
# Build a workflow from the shell:
|
|
50
|
-
bullswarm workflow draft create my-audit
|
|
51
|
-
bullswarm workflow draft phase add my-audit discover
|
|
52
|
-
bullswarm workflow draft step add my-audit discover list-files --type run --prompt 'List files'
|
|
53
|
-
bullswarm workflow draft run my-audit
|
|
46
|
+
# Validate a caller-authored program before launch:
|
|
47
|
+
bullswarm workflow plan validate "Fix the failing tests" --program plan.json
|
|
54
48
|
# Operate on a run by shortId (6 chars) or full runId (`wf-...`):
|
|
55
49
|
bullswarm workflow runs show <shortId>
|
|
56
50
|
bullswarm workflow runs delete <shortId> --yes
|
|
@@ -59,13 +53,13 @@ bullswarm workflow runs delete <shortId> --yes
|
|
|
59
53
|
## Using bullswarm from another agent
|
|
60
54
|
|
|
61
55
|
If you are an agent that wants to offload bounded work via bullswarm,
|
|
62
|
-
read `skill/SKILL.md` — that's the agent-facing user guide.
|
|
63
|
-
|
|
64
|
-
|
|
65
|
-
|
|
66
|
-
`workflow
|
|
67
|
-
|
|
68
|
-
|
|
56
|
+
read `skill/SKILL.md` — that's the agent-facing user guide. There are
|
|
57
|
+
exactly two ways to start work, and the caller chooses the shape itself: one
|
|
58
|
+
bounded outcome goes to `bullswarm run`; parallel territories, integration,
|
|
59
|
+
or independent acceptance go to `bullswarm workflow goal` with a program you
|
|
60
|
+
author (`bullswarm workflow plan contract` returns the schema). There is no
|
|
61
|
+
classifier or preview step. The skill is published alongside the package and
|
|
62
|
+
is the canonical reference for the CLI surface.
|
|
69
63
|
|
|
70
64
|
- Zero runtime dependencies. Node >= 18. Tests must never require network:
|
|
71
65
|
prime `~/.bullswarm/meters/*.json` caches with fresh timestamps if needed.
|
package/CHANGELOG.md
CHANGED
|
@@ -1,5 +1,197 @@
|
|
|
1
1
|
# bullswarm changelog
|
|
2
2
|
|
|
3
|
+
## 0.27.0 — one workflow engine
|
|
4
|
+
|
|
5
|
+
- Fixed: a worker that floods its stdout could kill the kernel. Every chunk of a
|
|
6
|
+
worker's stdout and stderr was appended to one string; a command-code worker
|
|
7
|
+
whose transcript outgrew Node's maximum string length threw
|
|
8
|
+
`RangeError: Invalid string length` inside the stream handler, the kernel died
|
|
9
|
+
as an uncaught exception, and every worker it supervised died with it (seen
|
|
10
|
+
twice on the same action in one day; the trace is in
|
|
11
|
+
`~/.bullswarm/goals/<runId>/stderr.log`). Streams are now captured through a
|
|
12
|
+
bounded buffer that keeps the first and last 16 MiB of each stream, counts what
|
|
13
|
+
it dropped (`captureTruncated` on the attempt observation), and any exception
|
|
14
|
+
raised while reading a worker now fails that attempt instead of the kernel.
|
|
15
|
+
Fatal-signature matching already looked only at the last 4,000 characters, so
|
|
16
|
+
quota and auth detection are unchanged.
|
|
17
|
+
|
|
18
|
+
- There is now one workflow engine. The authored-graph verbs `workflow run`,
|
|
19
|
+
`validate`, `list`, `draft`, `inspect` and `approval` are gone — each falls to
|
|
20
|
+
the workflow-level unknown-verb message, exits 2 and spawns nothing — and with
|
|
21
|
+
them the eleven V1 modules they drove: `runtime.js` (1,902 lines),
|
|
22
|
+
`runner.js` (1,304), `draft-cli.js` (434), `validate.js` (366),
|
|
23
|
+
`decision.js` (357), `tui.js` (297), `draft.js` (268), `result.js` (239),
|
|
24
|
+
`template.js` (146), `schema.js` (81) and `semaphore.js` (56). The V1 panel
|
|
25
|
+
model, timeline and orchestrator-detail twins came out of `dashboard.js`
|
|
26
|
+
(−965), the V1 branches out of `cli.js` (−291), `runs-cli.js` (−141),
|
|
27
|
+
`short-id.js` (−127) and `watch-cli.js` (−132). `src/workflow/*.js` goes from
|
|
28
|
+
16,941 lines to 10,257. Also deleted: the five saved definitions
|
|
29
|
+
`workflows/adaptive-code-review.json`, `agent-model-comparison.json`,
|
|
30
|
+
`connector-audit.json`, `smoke-two-step.json` and `verify-and-cap.json`
|
|
31
|
+
(257 lines) together with the `workflows/` entry in package.json `files`;
|
|
32
|
+
`scripts/sanity-multi-claude.mjs` (131), which only exercised `runWorkflow`;
|
|
33
|
+
and two more tools that could only reach deleted modules —
|
|
34
|
+
`bin/check-output-schema.js` (30, the V1 `outputSchema` worker preflight) and
|
|
35
|
+
`scripts/planner-contract-probe.mjs` (175, a probe for the V1 `decide`
|
|
36
|
+
planner). `workflow capabilities` now reports `engines.authoredGraphs` as
|
|
37
|
+
`{ retired: '0.27.0', command: null }` instead of advertising a verb that
|
|
38
|
+
exits 2. `bullswarm workflow --help` describes one engine.
|
|
39
|
+
|
|
40
|
+
- `newRunId` moved from `runner.js` into `src/workflow/short-id.js`, same
|
|
41
|
+
behaviour, exported; `v2-runtime.js` imports it from there.
|
|
42
|
+
|
|
43
|
+
- Removed the last V1 remnants that no gate caught because they named no
|
|
44
|
+
deleted symbol: the dead authored-graph planner prompt in
|
|
45
|
+
`src/workflow/goal.js` (`AUTONOMOUS_ORCHESTRATOR_PROMPT`,
|
|
46
|
+
`PLANNER_RULES_SECTION`, `PLANNER_EXAMPLES_SECTION` — 36 lines whose only
|
|
47
|
+
consumer was the deleted `runtime.js`, and which still taught `type`,
|
|
48
|
+
`stepTemplate`, `itemsFrom`, `outputSchema`, `covers` and `completion.when`
|
|
49
|
+
to a planner the V2 validator would reject), and the permanently-zero
|
|
50
|
+
`fanout: { total, ok, failed }` counter that `dashboard.js` still put on every
|
|
51
|
+
row and rendered behind an unreachable branch. `outputSchema`, `itemsFrom` and
|
|
52
|
+
`stepTemplate` now appear nowhere in `src/`. The `skill/references/operations.md`
|
|
53
|
+
"Adversarial verification" section described the removed `{ok, concerns,
|
|
54
|
+
summary}` verify verdict; it now documents the `bullswarm.workflow.evidence.v2`
|
|
55
|
+
envelope the kernel actually enforces. `AGENTS.md` doctrine item 5 pointed at
|
|
56
|
+
the deleted `runtime.js` and now points at `v2-dispatch.js`.
|
|
57
|
+
|
|
58
|
+
- Historical authored-graph runs stay readable, read-only, and nothing tries to
|
|
59
|
+
drive them. A run directory whose `state.json` lacks
|
|
60
|
+
`schemaVersion: 'bullswarm.workflow.state.v2'` is a legacy run: `workflow runs`
|
|
61
|
+
(with `--all` and `--json`) lists it as one row — short id, run id, name or
|
|
62
|
+
goal, status, age — marked `legacy`, reading only those five fields and never
|
|
63
|
+
throwing on a missing one, and the workflow home lists the same row. Every
|
|
64
|
+
driving command — `runs show`, `runs result`, `watch`, `cancel`, `resume`,
|
|
65
|
+
`steer`, `action show`, `tui <runId>` — prints exactly one line, `legacy
|
|
66
|
+
authored-graph run <shortId>: its executor was removed in 0.27.0; files remain
|
|
67
|
+
under <dir>`, and exits 2 before touching anything — including on the older
|
|
68
|
+
directories that hold only a `workflow.json` and never had a `state.json` at
|
|
69
|
+
all; the workflow home shows
|
|
70
|
+
that same line in its detail pane. `events <runId>` still replays the durable
|
|
71
|
+
JSONL and `runs delete <id> --yes` still removes the directory. Historical
|
|
72
|
+
directories are never modified. The stale-owner reconciliation that used to
|
|
73
|
+
run before every dispatch is gone with the V1 liveness model it served.
|
|
74
|
+
|
|
75
|
+
- Tests: 818 -> 661. Six V1-only files were deleted
|
|
76
|
+
(`workflow-adaptive`, `workflow-gaps`, `workflow-draft`, `workflow-schema`,
|
|
77
|
+
`workflow-validate`, `workflow-run` — 151 tests); `workflow-runs`,
|
|
78
|
+
`workflow-watch`, `assignments`, `workflow-interruption`, `workflow-steering`
|
|
79
|
+
and `workflow-goal` were rewritten onto the V2 kernel keeping every assertion
|
|
80
|
+
about shared behaviour; `workflow-dashboard` went from 48 cases to 35 — 21
|
|
81
|
+
V1-only cases removed and 2 added with the rewrite, then 6 re-added against V2
|
|
82
|
+
fixtures for the rendering behaviours the removal had dropped (blocked-action
|
|
83
|
+
naming, one segment header per phase or dependency level, the mid-segment
|
|
84
|
+
continuation header, parallel levels grouped in declared order, the narrow
|
|
85
|
+
layout, and auto-follow); and a new `workflow-legacy-runs` (11 tests) proves
|
|
86
|
+
the legacy contract against a synthetic legacy `state.json`.
|
|
87
|
+
`tests/manual-dynamic-real.mjs` (330 lines), a
|
|
88
|
+
manual real-provider matrix for authored `run`/`decide` graphs, went with the
|
|
89
|
+
executor; it was never part of the suite count. No test dispatches a real
|
|
90
|
+
provider.
|
|
91
|
+
|
|
92
|
+
## 0.26.0 — two entry points, kinds and rungs
|
|
93
|
+
|
|
94
|
+
- There are now exactly two ways to start work, and `bullswarm delegate` is
|
|
95
|
+
gone. `delegate` existed to decide, on the caller's behalf, whether a request
|
|
96
|
+
needed one agent or a workflow — with `--mode auto` it spent a real
|
|
97
|
+
analyze-lane dispatch on a deterministic-then-LLM classifier before doing any
|
|
98
|
+
of the actual work. The calling agent already knows the shape of its own
|
|
99
|
+
request, so that round trip bought latency and quota, not accuracy. `bullswarm
|
|
100
|
+
run` is the single-agent entry point and `bullswarm workflow goal` is the
|
|
101
|
+
workflow entry point; `bullswarm --help` names those two, in that order, and
|
|
102
|
+
`bullswarm delegate` now exits 2 with the standard unknown-verb message and
|
|
103
|
+
dispatches nothing. The packaged skill and the awareness block registered into
|
|
104
|
+
Codex, Claude, and Grok say the same thing in six lines: decide the shape
|
|
105
|
+
yourself, there is no preview step.
|
|
106
|
+
|
|
107
|
+
- `bullswarm run --lane analyze` now defaults to `medium` effort instead of
|
|
108
|
+
`high`. The V2 validator has always defaulted an analyze action to medium, so
|
|
109
|
+
the same lane meant a different tier depending on which entry point you used.
|
|
110
|
+
`run` and the validator now read one exported `DEFAULT_EFFORT_BY_LANE` table,
|
|
111
|
+
and `run --help` states the corrected default.
|
|
112
|
+
|
|
113
|
+
- Program actions can now say what they *are* instead of restating how to route
|
|
114
|
+
them. The optional `kind` field takes one of seven values — `mechanical`,
|
|
115
|
+
`io-read`, `check`, `implement`, `integration`, `architecture`,
|
|
116
|
+
`adversarial-acceptance` — and derives `lane` and `effort` from a table
|
|
117
|
+
exported alongside the existing per-lane default. Resolution is per field: an
|
|
118
|
+
explicit action field wins, then the kind table, then a new optional
|
|
119
|
+
program-level `defaults` object (`effort` and `reasoning` only), then the lane
|
|
120
|
+
default. A kind outside the closed list is a validation error with the allowed
|
|
121
|
+
values named, because it is a typo in the program rather than a runtime
|
|
122
|
+
condition. Programs that use neither `kind` nor `defaults` and state
|
|
123
|
+
`lane` and `effort` on every action normalise byte-identically to before;
|
|
124
|
+
`effort` itself is now optional and falls back to the per-lane default
|
|
125
|
+
where it used to be rejected as missing.
|
|
126
|
+
|
|
127
|
+
- Two non-blocking advisories, never rejections. `all-writers-high` fires when
|
|
128
|
+
three or more build/chore actions run and none is below high effort;
|
|
129
|
+
`docs-at-high` fires when a build/chore action owns only `*.md` files at high
|
|
130
|
+
effort. `workflow plan validate --json` carries them as `advisories` and the
|
|
131
|
+
human output prints `advisory:` lines; `workflow goal --program` prints the
|
|
132
|
+
same lines at launch. Neither changes acceptance or an exit code. The kernel
|
|
133
|
+
records them on the run, and `workflow runs show` lists them.
|
|
134
|
+
|
|
135
|
+
- `workflow runs result`, `runs show`, and `action show` print `kind` next to
|
|
136
|
+
lane and effort when an action has one. `workflow action show` now understands
|
|
137
|
+
autonomous V2 runs at all — it previously read only the V1 action ledger and
|
|
138
|
+
failed on every V2 run.
|
|
139
|
+
|
|
140
|
+
- **Rungs: one pool's model plus its reasoning level, for one effort tier, read
|
|
141
|
+
and written as one thing.** `bullswarm strategy rungs [--json] [--pool <name>]`
|
|
142
|
+
prints one row per enabled pool and configured tier with the effective model
|
|
143
|
+
and its source, the effective reasoning level and the layer that chose it, the
|
|
144
|
+
dated Epoch benchmark evidence for that model *at that level* (`blended`, cost
|
|
145
|
+
per task, tokens per task), and the local record from the decision log for that
|
|
146
|
+
pool and tier (dispatches, median wall minutes, ok share). Absent evidence
|
|
147
|
+
prints `no evidence` and an unmeasured tier prints `no dispatches`; neither is
|
|
148
|
+
ever estimated. `strategy inventory --json` gained the same rows under `rungs`.
|
|
149
|
+
|
|
150
|
+
- `bullswarm strategy set-rung <pool> <tier> --model <model> [--reasoning <level>]
|
|
151
|
+
[--force]` writes both halves of a rung in one atomic state save, so the model
|
|
152
|
+
and the thinking depth can never land separately. A rung is singular per pool
|
|
153
|
+
and tier: the tier moves off whichever model held it, and that model keeps its
|
|
154
|
+
other tiers. A level the connector cannot express is clamped to the strongest
|
|
155
|
+
it accepts and the clamp is printed. An unknown pool or tier exits 2; a model
|
|
156
|
+
absent from the pool's cached discovery exits 2 and lists the known models
|
|
157
|
+
unless `--force` is given. Neither `rungs` nor `set-rung` ever spawns model
|
|
158
|
+
discovery. No state migration: rungs are a projection of `strategy.modelTiers`
|
|
159
|
+
and `strategy.reasoning`, and `~/.bullswarm/state.json` gained no keys.
|
|
160
|
+
|
|
161
|
+
- The setup wizard's tier step now shows each suggested rung with its benchmark
|
|
162
|
+
evidence line and asks one reasoning question per configured tier. **Behavior
|
|
163
|
+
change:** Enter keeps that connector's own per-tier default and writes nothing,
|
|
164
|
+
where the previous question stored a suggested level (`xhigh`/`high`/`medium`)
|
|
165
|
+
on a blank answer. Connector defaults remain the final fallback, so a pool with
|
|
166
|
+
no configured rung behaves exactly as before. Non-TTY and `--yes` paths are
|
|
167
|
+
unchanged.
|
|
168
|
+
|
|
169
|
+
- A dated evidence datapack per model *and reasoning level*, from Epoch AI.
|
|
170
|
+
`data/epoch-benchmarks.json` (schema `bullswarm.epoch.benchmarks.v1`) is built
|
|
171
|
+
by the new `scripts/refresh-epoch-benchmarks.mjs` from Epoch's cursorbench,
|
|
172
|
+
deepswe, arc-agi-2, and critpt exports, and `src/lib/epoch-benchmarks.js`
|
|
173
|
+
reads it with the same cache → bundled → URL fallback as the OpenRouter pack.
|
|
174
|
+
`rungEvidence()` returns the mean of whichever of those four scores exist for a
|
|
175
|
+
(model, level) pair as `blended`, with cost and tokens per task from
|
|
176
|
+
cursorbench; `normalizeModelId()` is how connector model ids match the export.
|
|
177
|
+
The data is used under CC BY 4.0 — Epoch AI, 'AI Benchmarking Hub'. Published
|
|
178
|
+
online at epoch.ai. Retrieved from https://epoch.ai/benchmarks.
|
|
179
|
+
|
|
180
|
+
- The daily refresh job is renamed `.github/workflows/refresh-benchmarks.yml` and
|
|
181
|
+
now refreshes both assets on the rolling `benchmark-data-latest` release: it
|
|
182
|
+
downloads and unzips Epoch's public export, runs the script, runs the new
|
|
183
|
+
tests, and uploads `epoch-benchmarks.json` next to `openrouter-benchmarks.json`.
|
|
184
|
+
The ambiguous `npm run refresh:benchmarks` script is split into
|
|
185
|
+
`refresh:openrouter` and `refresh:epoch`, since only one of the two now
|
|
186
|
+
refreshes "the benchmarks".
|
|
187
|
+
|
|
188
|
+
- The OpenRouter builder is unchanged, and that is a finding rather than an
|
|
189
|
+
omission: the Artificial Analysis records in the 2026-09-08 capture carry only
|
|
190
|
+
`agentic_index`, `coding_index`, and `intelligence_index` per model, with no
|
|
191
|
+
reasoning-effort marker on any of the `reasoning_effort`, `effort`, `variant`,
|
|
192
|
+
`reasoning`, or `reasoning_level` fields checked, so there is no per-effort row
|
|
193
|
+
to keep. `data/README.md` records the field names inspected.
|
|
194
|
+
|
|
3
195
|
## 0.25.5 — forecast-aware routing
|
|
4
196
|
|
|
5
197
|
- Bullswarm now knows what it is already running. Every dispatch registers the
|