bullswarm 0.25.5 → 0.27.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (51) hide show
  1. package/AGENTS.md +19 -25
  2. package/CHANGELOG.md +192 -0
  3. package/README.md +118 -151
  4. package/connectors/codex.json +3 -2
  5. package/connectors/command-code.json +229 -27
  6. package/data/README.md +139 -9
  7. package/data/epoch-benchmarks.json +6028 -0
  8. package/docs/studies/portal-token-diet.md +127 -0
  9. package/docs/workflow-design.md +5 -1
  10. package/package.json +4 -3
  11. package/skill/SKILL.md +28 -35
  12. package/skill/references/operations.md +65 -13
  13. package/src/cli.js +2 -4
  14. package/src/help.js +93 -460
  15. package/src/integrate.js +4 -11
  16. package/src/lib/epoch-benchmarks.js +394 -0
  17. package/src/lib/strategy.js +209 -2
  18. package/src/lib/watch.js +93 -20
  19. package/src/setup.js +76 -24
  20. package/src/strategy-cli.js +199 -3
  21. package/src/workflow/action-validator.js +117 -4
  22. package/src/workflow/cli.js +114 -284
  23. package/src/workflow/dashboard.js +153 -967
  24. package/src/workflow/goal.js +0 -36
  25. package/src/workflow/runs-cli.js +117 -131
  26. package/src/workflow/short-id.js +86 -127
  27. package/src/workflow/steering.js +6 -8
  28. package/src/workflow/v2-cancellation.js +6 -1
  29. package/src/workflow/v2-outcome.js +8 -1
  30. package/src/workflow/v2-planner.js +31 -5
  31. package/src/workflow/v2-runtime.js +1 -2
  32. package/src/workflow/v2-state.js +21 -2
  33. package/src/workflow/watch-cli.js +79 -133
  34. package/bin/check-output-schema.js +0 -30
  35. package/src/delegate.js +0 -438
  36. package/src/workflow/decision.js +0 -357
  37. package/src/workflow/draft-cli.js +0 -434
  38. package/src/workflow/draft.js +0 -268
  39. package/src/workflow/result.js +0 -239
  40. package/src/workflow/runner.js +0 -1304
  41. package/src/workflow/runtime.js +0 -1902
  42. package/src/workflow/schema.js +0 -81
  43. package/src/workflow/semaphore.js +0 -56
  44. package/src/workflow/template.js +0 -146
  45. package/src/workflow/tui.js +0 -297
  46. package/src/workflow/validate.js +0 -366
  47. package/workflows/adaptive-code-review.json +0 -45
  48. package/workflows/agent-model-comparison.json +0 -49
  49. package/workflows/connector-audit.json +0 -60
  50. package/workflows/smoke-two-step.json +0 -41
  51. package/workflows/verify-and-cap.json +0 -62
package/AGENTS.md CHANGED
@@ -20,37 +20,31 @@ content. Published as `bullswarm` on npm.
20
20
  5. Workflow dispatches must honor the same guarantees as single runs:
21
21
  `BULLSWARM_DEPTH` is propagated, burst-gated pools are excluded, and
22
22
  auth verdicts quarantine the pool + append to the shared decision log
23
- (R6/R7/R8 in `src/workflow/runtime.js`).
24
- 6. Adversarial verification is a first-class primitive: a `verify` step
25
- reads a prior outFile and demands a JSON `{ok, concerns, summary}`
26
- verdict before downstream steps can trust the work (R-skeptic).
27
- 7. Workflows can be built incrementally from the shell
28
- (`bullswarm workflow draft create/phase/step/set/...`). Drafts are
29
- stored under `~/.bullswarm/drafts/<name>/` and are runnable by name
30
- without an upfront JSON. JSON is still the durable artifact — drafts
31
- are JSON documents, just built one mutation at a time.
32
- 8. New goal workflows are caller-planned programs in a shared workspace.
23
+ (R6/R7/R8 in `src/workflow/v2-dispatch.js`).
24
+ 6. Adversarial verification is a first-class primitive: an action naming
25
+ requirements in `evidenceFor` is dispatched under an evidence contract and
26
+ judges them from the durable artifact, so a requirement is only verified by
27
+ work someone else inspected (R-skeptic).
28
+ 7. New goal workflows are caller-planned programs in a shared workspace.
33
29
  `bullswarm workflow goal --program` executes the graph; `--orchestrator`
34
30
  explicitly delegates planning. File territories are advisory scheduling
35
31
  hints, and the graph finishes without automatic gap rounds. `verified`
36
32
  separately records requirement evidence. `--isolation` opts into strict
37
33
  per-worker worktrees. Saved V2 runs preserve their original semantics.
34
+ 8. Historical authored-graph runs remain visible as read-only `legacy` rows.
35
+ Their executor was removed in 0.27.0; driving commands fail closed before
36
+ dispatch and historical run directories remain untouched.
38
37
 
39
38
  ## Development
40
39
 
41
40
  ```bash
42
41
  npm test # full suite, no network needed (meters read from cache)
43
42
  node bin/bullswarm.js doctor --json # readiness report
44
- node bin/bullswarm.js workflow list # discover workflows
43
+ node bin/bullswarm.js workflow goal "Fix the failing tests" --program plan.json
45
44
  node bin/bullswarm.js workflow runs # ongoing workflow instances
46
45
  node bin/bullswarm.js workflow runs --all # including historical
47
- node bin/bullswarm.js workflow validate <file> # dry-run
48
- BULLSWARM_HOME=/tmp/bs node bin/bullswarm.js workflow run <file> # sandboxed run
49
- # Build a workflow from the shell:
50
- bullswarm workflow draft create my-audit
51
- bullswarm workflow draft phase add my-audit discover
52
- bullswarm workflow draft step add my-audit discover list-files --type run --prompt 'List files'
53
- bullswarm workflow draft run my-audit
46
+ # Validate a caller-authored program before launch:
47
+ bullswarm workflow plan validate "Fix the failing tests" --program plan.json
54
48
  # Operate on a run by shortId (6 chars) or full runId (`wf-...`):
55
49
  bullswarm workflow runs show <shortId>
56
50
  bullswarm workflow runs delete <shortId> --yes
@@ -59,13 +53,13 @@ bullswarm workflow runs delete <shortId> --yes
59
53
  ## Using bullswarm from another agent
60
54
 
61
55
  If you are an agent that wants to offload bounded work via bullswarm,
62
- read `skill/SKILL.md` — that's the agent-facing user guide. Use
63
- `bullswarm delegate` (or the installed `/bullswarm` skill) by default: it
64
- previews whether one bounded agent or an autonomous workflow is appropriate,
65
- shows the conceptual plan, and executes the chosen engine. Reach for `run`,
66
- `workflow goal`, or a fixed workflow graph directly only when the caller has
67
- already chosen that execution shape. The skill is published alongside the
68
- package and is the canonical reference for the CLI surface.
56
+ read `skill/SKILL.md` — that's the agent-facing user guide. There are
57
+ exactly two ways to start work, and the caller chooses the shape itself: one
58
+ bounded outcome goes to `bullswarm run`; parallel territories, integration,
59
+ or independent acceptance go to `bullswarm workflow goal` with a program you
60
+ author (`bullswarm workflow plan contract` returns the schema). There is no
61
+ classifier or preview step. The skill is published alongside the package and
62
+ is the canonical reference for the CLI surface.
69
63
 
70
64
  - Zero runtime dependencies. Node >= 18. Tests must never require network:
71
65
  prime `~/.bullswarm/meters/*.json` caches with fresh timestamps if needed.
package/CHANGELOG.md CHANGED
@@ -1,5 +1,197 @@
1
1
  # bullswarm changelog
2
2
 
3
+ ## 0.27.0 — one workflow engine
4
+
5
+ - Fixed: a worker that floods its stdout could kill the kernel. Every chunk of a
6
+ worker's stdout and stderr was appended to one string; a command-code worker
7
+ whose transcript outgrew Node's maximum string length threw
8
+ `RangeError: Invalid string length` inside the stream handler, the kernel died
9
+ as an uncaught exception, and every worker it supervised died with it (seen
10
+ twice on the same action in one day; the trace is in
11
+ `~/.bullswarm/goals/<runId>/stderr.log`). Streams are now captured through a
12
+ bounded buffer that keeps the first and last 16 MiB of each stream, counts what
13
+ it dropped (`captureTruncated` on the attempt observation), and any exception
14
+ raised while reading a worker now fails that attempt instead of the kernel.
15
+ Fatal-signature matching already looked only at the last 4,000 characters, so
16
+ quota and auth detection are unchanged.
17
+
18
+ - There is now one workflow engine. The authored-graph verbs `workflow run`,
19
+ `validate`, `list`, `draft`, `inspect` and `approval` are gone — each falls to
20
+ the workflow-level unknown-verb message, exits 2 and spawns nothing — and with
21
+ them the eleven V1 modules they drove: `runtime.js` (1,902 lines),
22
+ `runner.js` (1,304), `draft-cli.js` (434), `validate.js` (366),
23
+ `decision.js` (357), `tui.js` (297), `draft.js` (268), `result.js` (239),
24
+ `template.js` (146), `schema.js` (81) and `semaphore.js` (56). The V1 panel
25
+ model, timeline and orchestrator-detail twins came out of `dashboard.js`
26
+ (−965), the V1 branches out of `cli.js` (−291), `runs-cli.js` (−141),
27
+ `short-id.js` (−127) and `watch-cli.js` (−132). `src/workflow/*.js` goes from
28
+ 16,941 lines to 10,257. Also deleted: the five saved definitions
29
+ `workflows/adaptive-code-review.json`, `agent-model-comparison.json`,
30
+ `connector-audit.json`, `smoke-two-step.json` and `verify-and-cap.json`
31
+ (257 lines) together with the `workflows/` entry in package.json `files`;
32
+ `scripts/sanity-multi-claude.mjs` (131), which only exercised `runWorkflow`;
33
+ and two more tools that could only reach deleted modules —
34
+ `bin/check-output-schema.js` (30, the V1 `outputSchema` worker preflight) and
35
+ `scripts/planner-contract-probe.mjs` (175, a probe for the V1 `decide`
36
+ planner). `workflow capabilities` now reports `engines.authoredGraphs` as
37
+ `{ retired: '0.27.0', command: null }` instead of advertising a verb that
38
+ exits 2. `bullswarm workflow --help` describes one engine.
39
+
40
+ - `newRunId` moved from `runner.js` into `src/workflow/short-id.js`, same
41
+ behaviour, exported; `v2-runtime.js` imports it from there.
42
+
43
+ - Removed the last V1 remnants that no gate caught because they named no
44
+ deleted symbol: the dead authored-graph planner prompt in
45
+ `src/workflow/goal.js` (`AUTONOMOUS_ORCHESTRATOR_PROMPT`,
46
+ `PLANNER_RULES_SECTION`, `PLANNER_EXAMPLES_SECTION` — 36 lines whose only
47
+ consumer was the deleted `runtime.js`, and which still taught `type`,
48
+ `stepTemplate`, `itemsFrom`, `outputSchema`, `covers` and `completion.when`
49
+ to a planner the V2 validator would reject), and the permanently-zero
50
+ `fanout: { total, ok, failed }` counter that `dashboard.js` still put on every
51
+ row and rendered behind an unreachable branch. `outputSchema`, `itemsFrom` and
52
+ `stepTemplate` now appear nowhere in `src/`. The `skill/references/operations.md`
53
+ "Adversarial verification" section described the removed `{ok, concerns,
54
+ summary}` verify verdict; it now documents the `bullswarm.workflow.evidence.v2`
55
+ envelope the kernel actually enforces. `AGENTS.md` doctrine item 5 pointed at
56
+ the deleted `runtime.js` and now points at `v2-dispatch.js`.
57
+
58
+ - Historical authored-graph runs stay readable, read-only, and nothing tries to
59
+ drive them. A run directory whose `state.json` lacks
60
+ `schemaVersion: 'bullswarm.workflow.state.v2'` is a legacy run: `workflow runs`
61
+ (with `--all` and `--json`) lists it as one row — short id, run id, name or
62
+ goal, status, age — marked `legacy`, reading only those five fields and never
63
+ throwing on a missing one, and the workflow home lists the same row. Every
64
+ driving command — `runs show`, `runs result`, `watch`, `cancel`, `resume`,
65
+ `steer`, `action show`, `tui <runId>` — prints exactly one line, `legacy
66
+ authored-graph run <shortId>: its executor was removed in 0.27.0; files remain
67
+ under <dir>`, and exits 2 before touching anything — including on the older
68
+ directories that hold only a `workflow.json` and never had a `state.json` at
69
+ all; the workflow home shows
70
+ that same line in its detail pane. `events <runId>` still replays the durable
71
+ JSONL and `runs delete <id> --yes` still removes the directory. Historical
72
+ directories are never modified. The stale-owner reconciliation that used to
73
+ run before every dispatch is gone with the V1 liveness model it served.
74
+
75
+ - Tests: 818 -> 661. Six V1-only files were deleted
76
+ (`workflow-adaptive`, `workflow-gaps`, `workflow-draft`, `workflow-schema`,
77
+ `workflow-validate`, `workflow-run` — 151 tests); `workflow-runs`,
78
+ `workflow-watch`, `assignments`, `workflow-interruption`, `workflow-steering`
79
+ and `workflow-goal` were rewritten onto the V2 kernel keeping every assertion
80
+ about shared behaviour; `workflow-dashboard` went from 48 cases to 35 — 21
81
+ V1-only cases removed and 2 added with the rewrite, then 6 re-added against V2
82
+ fixtures for the rendering behaviours the removal had dropped (blocked-action
83
+ naming, one segment header per phase or dependency level, the mid-segment
84
+ continuation header, parallel levels grouped in declared order, the narrow
85
+ layout, and auto-follow); and a new `workflow-legacy-runs` (11 tests) proves
86
+ the legacy contract against a synthetic legacy `state.json`.
87
+ `tests/manual-dynamic-real.mjs` (330 lines), a
88
+ manual real-provider matrix for authored `run`/`decide` graphs, went with the
89
+ executor; it was never part of the suite count. No test dispatches a real
90
+ provider.
91
+
92
+ ## 0.26.0 — two entry points, kinds and rungs
93
+
94
+ - There are now exactly two ways to start work, and `bullswarm delegate` is
95
+ gone. `delegate` existed to decide, on the caller's behalf, whether a request
96
+ needed one agent or a workflow — with `--mode auto` it spent a real
97
+ analyze-lane dispatch on a deterministic-then-LLM classifier before doing any
98
+ of the actual work. The calling agent already knows the shape of its own
99
+ request, so that round trip bought latency and quota, not accuracy. `bullswarm
100
+ run` is the single-agent entry point and `bullswarm workflow goal` is the
101
+ workflow entry point; `bullswarm --help` names those two, in that order, and
102
+ `bullswarm delegate` now exits 2 with the standard unknown-verb message and
103
+ dispatches nothing. The packaged skill and the awareness block registered into
104
+ Codex, Claude, and Grok say the same thing in six lines: decide the shape
105
+ yourself, there is no preview step.
106
+
107
+ - `bullswarm run --lane analyze` now defaults to `medium` effort instead of
108
+ `high`. The V2 validator has always defaulted an analyze action to medium, so
109
+ the same lane meant a different tier depending on which entry point you used.
110
+ `run` and the validator now read one exported `DEFAULT_EFFORT_BY_LANE` table,
111
+ and `run --help` states the corrected default.
112
+
113
+ - Program actions can now say what they *are* instead of restating how to route
114
+ them. The optional `kind` field takes one of seven values — `mechanical`,
115
+ `io-read`, `check`, `implement`, `integration`, `architecture`,
116
+ `adversarial-acceptance` — and derives `lane` and `effort` from a table
117
+ exported alongside the existing per-lane default. Resolution is per field: an
118
+ explicit action field wins, then the kind table, then a new optional
119
+ program-level `defaults` object (`effort` and `reasoning` only), then the lane
120
+ default. A kind outside the closed list is a validation error with the allowed
121
+ values named, because it is a typo in the program rather than a runtime
122
+ condition. Programs that use neither `kind` nor `defaults` and state
123
+ `lane` and `effort` on every action normalise byte-identically to before;
124
+ `effort` itself is now optional and falls back to the per-lane default
125
+ where it used to be rejected as missing.
126
+
127
+ - Two non-blocking advisories, never rejections. `all-writers-high` fires when
128
+ three or more build/chore actions run and none is below high effort;
129
+ `docs-at-high` fires when a build/chore action owns only `*.md` files at high
130
+ effort. `workflow plan validate --json` carries them as `advisories` and the
131
+ human output prints `advisory:` lines; `workflow goal --program` prints the
132
+ same lines at launch. Neither changes acceptance or an exit code. The kernel
133
+ records them on the run, and `workflow runs show` lists them.
134
+
135
+ - `workflow runs result`, `runs show`, and `action show` print `kind` next to
136
+ lane and effort when an action has one. `workflow action show` now understands
137
+ autonomous V2 runs at all — it previously read only the V1 action ledger and
138
+ failed on every V2 run.
139
+
140
+ - **Rungs: one pool's model plus its reasoning level, for one effort tier, read
141
+ and written as one thing.** `bullswarm strategy rungs [--json] [--pool <name>]`
142
+ prints one row per enabled pool and configured tier with the effective model
143
+ and its source, the effective reasoning level and the layer that chose it, the
144
+ dated Epoch benchmark evidence for that model *at that level* (`blended`, cost
145
+ per task, tokens per task), and the local record from the decision log for that
146
+ pool and tier (dispatches, median wall minutes, ok share). Absent evidence
147
+ prints `no evidence` and an unmeasured tier prints `no dispatches`; neither is
148
+ ever estimated. `strategy inventory --json` gained the same rows under `rungs`.
149
+
150
+ - `bullswarm strategy set-rung <pool> <tier> --model <model> [--reasoning <level>]
151
+ [--force]` writes both halves of a rung in one atomic state save, so the model
152
+ and the thinking depth can never land separately. A rung is singular per pool
153
+ and tier: the tier moves off whichever model held it, and that model keeps its
154
+ other tiers. A level the connector cannot express is clamped to the strongest
155
+ it accepts and the clamp is printed. An unknown pool or tier exits 2; a model
156
+ absent from the pool's cached discovery exits 2 and lists the known models
157
+ unless `--force` is given. Neither `rungs` nor `set-rung` ever spawns model
158
+ discovery. No state migration: rungs are a projection of `strategy.modelTiers`
159
+ and `strategy.reasoning`, and `~/.bullswarm/state.json` gained no keys.
160
+
161
+ - The setup wizard's tier step now shows each suggested rung with its benchmark
162
+ evidence line and asks one reasoning question per configured tier. **Behavior
163
+ change:** Enter keeps that connector's own per-tier default and writes nothing,
164
+ where the previous question stored a suggested level (`xhigh`/`high`/`medium`)
165
+ on a blank answer. Connector defaults remain the final fallback, so a pool with
166
+ no configured rung behaves exactly as before. Non-TTY and `--yes` paths are
167
+ unchanged.
168
+
169
+ - A dated evidence datapack per model *and reasoning level*, from Epoch AI.
170
+ `data/epoch-benchmarks.json` (schema `bullswarm.epoch.benchmarks.v1`) is built
171
+ by the new `scripts/refresh-epoch-benchmarks.mjs` from Epoch's cursorbench,
172
+ deepswe, arc-agi-2, and critpt exports, and `src/lib/epoch-benchmarks.js`
173
+ reads it with the same cache → bundled → URL fallback as the OpenRouter pack.
174
+ `rungEvidence()` returns the mean of whichever of those four scores exist for a
175
+ (model, level) pair as `blended`, with cost and tokens per task from
176
+ cursorbench; `normalizeModelId()` is how connector model ids match the export.
177
+ The data is used under CC BY 4.0 — Epoch AI, 'AI Benchmarking Hub'. Published
178
+ online at epoch.ai. Retrieved from https://epoch.ai/benchmarks.
179
+
180
+ - The daily refresh job is renamed `.github/workflows/refresh-benchmarks.yml` and
181
+ now refreshes both assets on the rolling `benchmark-data-latest` release: it
182
+ downloads and unzips Epoch's public export, runs the script, runs the new
183
+ tests, and uploads `epoch-benchmarks.json` next to `openrouter-benchmarks.json`.
184
+ The ambiguous `npm run refresh:benchmarks` script is split into
185
+ `refresh:openrouter` and `refresh:epoch`, since only one of the two now
186
+ refreshes "the benchmarks".
187
+
188
+ - The OpenRouter builder is unchanged, and that is a finding rather than an
189
+ omission: the Artificial Analysis records in the 2026-09-08 capture carry only
190
+ `agentic_index`, `coding_index`, and `intelligence_index` per model, with no
191
+ reasoning-effort marker on any of the `reasoning_effort`, `effort`, `variant`,
192
+ `reasoning`, or `reasoning_level` fields checked, so there is no per-effort row
193
+ to keep. `data/README.md` records the field names inspected.
194
+
3
195
  ## 0.25.5 — forecast-aware routing
4
196
 
5
197
  - Bullswarm now knows what it is already running. Every dispatch registers the