bullswarm 0.37.3 → 0.38.1

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (72) hide show
  1. package/AGENTS.md +51 -52
  2. package/CHANGELOG.md +121 -0
  3. package/data/openrouter-benchmarks.json +8783 -8591
  4. package/docs/design/0.38.0-removals.md +123 -0
  5. package/docs/guide/observing.md +2 -2
  6. package/docs/guide/playbook.md +13 -11
  7. package/docs/guide/routing.md +17 -21
  8. package/docs/guide/workflows.md +33 -74
  9. package/docs/reference/cli.md +71 -110
  10. package/docs/reference/program.md +9 -8
  11. package/docs/reference/providers.md +1 -0
  12. package/docs/reference/result.md +15 -12
  13. package/package.json +1 -1
  14. package/providers/contrib/command-code/connector-history.json +36 -0
  15. package/providers/contrib/command-code/connector.json +78 -27
  16. package/skill/SKILL.md +26 -29
  17. package/skill/references/operations.md +126 -236
  18. package/skill/references/patterns.md +5 -5
  19. package/skill/references/program.md +28 -3
  20. package/src/help.js +127 -213
  21. package/src/lib/cli-flags.js +29 -22
  22. package/src/lib/config.js +11 -2
  23. package/src/lib/model-pin.js +9 -1
  24. package/src/lib/provider-errors.js +26 -1
  25. package/src/lib/run-delegate.js +15 -0
  26. package/src/lib/strategy.js +161 -9
  27. package/src/lib/watch.js +43 -1
  28. package/src/provider-cli.js +5 -0
  29. package/src/providers/_schema.json +4 -0
  30. package/src/strategy-cli.js +83 -7
  31. package/src/workflow/action-validator.js +37 -4
  32. package/src/workflow/attempt-bytes.js +5 -7
  33. package/src/workflow/attempt-record.js +3 -0
  34. package/src/workflow/caller-planner.js +9 -135
  35. package/src/workflow/cli-capabilities.js +4 -9
  36. package/src/workflow/cli-goal-document.js +9 -24
  37. package/src/workflow/cli-goal.js +28 -64
  38. package/src/workflow/cli-launch.js +4 -3
  39. package/src/workflow/cli-plan.js +51 -421
  40. package/src/workflow/cli-program-checks.js +17 -13
  41. package/src/workflow/cli-run-lookup.js +61 -5
  42. package/src/workflow/cli-run-verbs.js +36 -27
  43. package/src/workflow/cli-step-verbs.js +4 -4
  44. package/src/workflow/cli-steps.js +37 -6
  45. package/src/workflow/cli.js +9 -4
  46. package/src/workflow/contract-v3.js +2 -2
  47. package/src/workflow/goal-column.js +24 -2
  48. package/src/workflow/goal.js +0 -64
  49. package/src/workflow/home-model.js +1 -1
  50. package/src/workflow/kernel-resume.js +1 -16
  51. package/src/workflow/legacy-verification.js +514 -0
  52. package/src/workflow/no-pool-why.js +8 -4
  53. package/src/workflow/pick-preview.js +7 -5
  54. package/src/workflow/revision-v3.js +128 -13
  55. package/src/workflow/run-control.js +5 -32
  56. package/src/workflow/run-model.js +1 -1
  57. package/src/workflow/run-verdict.js +11 -2
  58. package/src/workflow/run-view.js +1 -1
  59. package/src/workflow/runs-cli.js +13 -13
  60. package/src/workflow/step-prompts.js +3 -89
  61. package/src/workflow/step-route.js +7 -3
  62. package/src/workflow/step-vocabulary.js +5 -2
  63. package/src/workflow/time-box.js +15 -0
  64. package/src/workflow/v2-dispatch.js +301 -539
  65. package/src/workflow/v2-outcome.js +3 -1
  66. package/src/workflow/v2-planner.js +3 -113
  67. package/src/workflow/v2-revision.js +22 -2
  68. package/src/workflow/v2-runtime.js +80 -725
  69. package/src/workflow/v2-state.js +8 -3
  70. package/src/workflow/watch-cli.js +16 -12
  71. package/src/workflow/workflow-flags.js +29 -2
  72. package/src/workflow/verify-rounds.js +0 -1228
package/AGENTS.md CHANGED
@@ -14,31 +14,27 @@ The core is being redesigned around facts-only mechanics, four mandatory
14
14
  principles, caller-chosen options and a pattern library, for any kind of work
15
15
  rather than code only. The design, the decisions taken and the staged build
16
16
  plan are in `docs/design/redesign-mechanics-principles-options.md`, and draft
17
- pattern cards are in `docs/design/patterns/`. Build in the plan's stage order.
18
- The doctrine below stays in force until the stage that changes an item lands.
19
- The redesign rewords items 1, 5, 6 and 7, and each stage updates this file.
20
- Stage 1 (step vocabulary) has landed: program-mode steps may state a role and a
21
- deliverable, each kind belongs to one role and keeps its exact routing, and
22
- the no-op gate is now 'declared deliverable not produced' (failure kind
23
- `not-produced`), measured over the whole step. Stage 2 (evidence v1) has landed:
24
- program steps may declare command and schema `evidence` that the kernel runs
25
- after the worker, a failure is `failed-evidence` with one same-pool retry, and
26
- finished steps in new runs are labelled `proven by …` or `finished · unproven`.
27
- Stage 3 (failure rule and routing constraints) has landed: one automatic retry
28
- per step, then the caller; a usage limit goes straight to the caller, from a
29
- step, the dispatched planner or the preflight scout alike; the
30
- needs-you block with `step rerun --avoid` and `step accept`; the per-step
31
- `route`; `verifyRounds` counts fixes (default 1); reviews are placed only by
32
- route. Runs started earlier keep their rules (`features.json`).
33
- 0.37.0 (program v3, the generic model) has landed: a step is a run, and a
17
+ pattern cards are in `docs/design/patterns/`. Each release updates this file.
18
+
19
+ 0.37.0 (program v3, the generic model) landed: a step is a run, and a
34
20
  workflow composes steps, phases, gates and loops. `bullswarm run` is a
35
- one-step workflow. New programs are `bullswarm.workflow.program.v3`: a step
36
- passes by facts (clean exit, deliverable produced, evidence passed, answer
37
- matching its schema), a gate waits for `workflow continue`, a loop reruns its
38
- steps until one condition holds (at most 5 rounds), and new work is added
39
- with `workflow add`, never by editing the run's steps. v3 reports facts per
40
- step and has no requirement IDs, `evidenceFor` or `verifyRounds`. v2 programs
41
- and saved runs keep running and replaying as before.
21
+ one-step workflow. A step passes by facts (clean exit, deliverable produced,
22
+ evidence passed, answer matching its schema), a gate waits for
23
+ `workflow continue`, a loop reruns its steps until one condition holds (at
24
+ most 5 rounds), and new work is added with `workflow add`, never by editing
25
+ the run's steps. v3 reports facts per step and has no requirement IDs,
26
+ `evidenceFor` or `verifyRounds`.
27
+
28
+ 0.38.0 (the removals release, `docs/design/0.38.0-removals.md`) deleted the
29
+ mechanisms v3 replaced: the preflight scout, the dispatched planner
30
+ (`--orchestrator`), the caller-planner gap turn (`plan show`/`plan submit`),
31
+ whole-plan revise (`plan export`/`plan revise`), the v2 repair loop, the
32
+ kernel-written digest and review tasks, and the old dispatch rules. New
33
+ programs are `bullswarm.workflow.program.v3` only; a v2 program is refused
34
+ before a run folder exists. Only a run marked `programFormat: 3` is driven.
35
+ Every other saved run (v2, stage 1-3, legacy) is view-only: it stays
36
+ listable, showable and countable, and every driving command refuses it
37
+ before doing anything. Their readers and stored formats are unchanged.
42
38
 
43
39
  ## Non-negotiable doctrine
44
40
 
@@ -55,6 +51,12 @@ and saved runs keep running and replaying as before.
55
51
  outlives a step is the 100% refusal marker a usage limit writes when the
56
52
  meter cannot be read, and it counts only when its reset was named or
57
53
  measured, never guessed (`refusalResetKnown` in `src/meters/framework.js`).
54
+ Beside it sits one plan fact: `state.strategy.planExcludedModels[pool]`,
55
+ the models a provider said that pool's subscription does not include
56
+ (failure kind `model-not-in-plan`). It is a fact about the plan, like its
57
+ price, not a spent or dead pool: the pool stays pickable with its other
58
+ models, and the record ends when the subscription changes or the operator
59
+ turns the model back on (`src/lib/strategy.js`).
58
60
  Recursion depth is core-owned via env (`BULLSWARM_DEPTH`).
59
61
  5. Workflow dispatches honor the same guarantees as single runs, in
60
62
  `src/workflow/v2-dispatch.js`: `BULLSWARM_DEPTH` is checked and propagated
@@ -76,34 +78,31 @@ and saved runs keep running and replaying as before.
76
78
  pool inside a run; only a transient rate limit backs off on the same pool,
77
79
  at most twice (20 s, then 60 s, or a named wait of at most 2 minutes), then
78
80
  goes to the caller.
79
- Only a failed step's dependents wait. The dispatched planner and the
80
- preflight scout follow the same usage-limit rule (`usageLimitsToCaller` in
81
- `dispatchV2Action`): they stop and the run tells the caller, with no
82
- automatic move to another pool.
83
- 6. Review is a caller option, recorded as a fact. A step naming requirements
84
- in `evidenceFor` is dispatched under the evidence contract and judges them
85
- from the durable artifact. Where it runs is the caller's choice through
86
- `route` (`independentOf`, `providers`, `pools`); Bullswarm never moves a
87
- review on its own, and records who reviewed (pool, model, provider) and
88
- whether that provider also wrote the work. A caller's `step accept` is
89
- recorded as evidence `choice` and never makes a requirement verified.
90
- Runs started before this rule keep automatic writer avoidance (R12/R13).
91
- 7. New goal workflows are caller-planned programs in a shared workspace.
92
- `bullswarm workflow goal --program` executes the graph; `--orchestrator`
93
- explicitly delegates planning. File territories are advisory scheduling
94
- hints. In a v3 program (the format for new work) a check is an ordinary
95
- step, and fixing until it passes is a loop the caller declares (`loops`,
96
- `until` one condition, `maxRounds` 1-5); a loop out of rounds or a gate
97
- waits for the caller (`workflow continue`), and a failed step goes to the
98
- caller through the watcher's needs-you block: rerun elsewhere, add steps
99
- (`workflow add`), take over, or accept anyway. v3 validate refuses
100
- `defaults.verifyRounds`. In a v2 program a failed check gets one fix step
101
- and one re-review (`defaults.verifyRounds`, default 1, 0-3), then the
102
- caller, and `verified` separately records requirement evidence.
103
- `--isolation` opts into strict per-worker worktrees. Saved V2 runs
104
- preserve their original semantics (`features.json`).
105
- 8. Historical authored-graph runs remain visible as read-only `legacy` rows.
106
- Their executor was removed in 0.27.0; driving commands fail closed before
81
+ Only a failed step's dependents wait.
82
+ 6. Review is a caller option, recorded as a fact. A check is an ordinary
83
+ step with an `answer` and/or `evidence`, and it passes by facts only.
84
+ Where it runs is the caller's choice through `route` (`independentOf`,
85
+ `providers`, `pools`); Bullswarm never moves a review on its own, and
86
+ records who reviewed (pool, model, provider) and whether that provider
87
+ also wrote the work. A caller's `step accept` is recorded as evidence
88
+ `choice`. Saved v2 runs keep showing their requirement verdicts
89
+ (`evidenceFor`, `verified`) as they were recorded.
90
+ 7. New goal workflows are caller-planned v3 programs in a shared workspace.
91
+ `bullswarm workflow goal --program` executes the graph; the caller writes
92
+ the plan (or a step whose answer is a list of steps, appended with
93
+ `workflow add --from-answer`). File territories are advisory scheduling
94
+ hints. A check is an ordinary step, and fixing until it passes is a loop
95
+ the caller declares (`loops`, `until` one condition, `maxRounds` 1-5); a
96
+ loop out of rounds or a gate waits for the caller (`workflow continue`),
97
+ and a failed step goes to the caller through the watcher's needs-you
98
+ block: rerun elsewhere, add steps (`workflow add`), take over, or accept
99
+ anyway. v3 validate refuses `defaults.verifyRounds`. `--isolation` opts
100
+ into strict per-worker worktrees. A saved run that is not v3 is view-only
101
+ (0.38.0): its original semantics are kept for display (`features.json`),
102
+ and nothing drives it again.
103
+ 8. Historical authored-graph runs remain visible as read-only `legacy` rows,
104
+ and `runs show`/`result`/`watch` print a bounded summary of them. Their
105
+ executor was removed in 0.27.0; driving commands fail closed before
107
106
  dispatch and historical run directories remain untouched.
108
107
 
109
108
  ## Development
package/CHANGELOG.md CHANGED
@@ -2,6 +2,127 @@
2
2
 
3
3
  ## Unreleased
4
4
 
5
+ ## 0.38.1 — Fewer wasted retries, honest no-pool reasons
6
+
7
+ - no eligible pool: when a pool lacks a tier only because free models are
8
+ off for it, or because its plan was seen not to include the tier's model,
9
+ the reason now says so (`free models are off for <pool>`, `<pool> has no
10
+ <tier>-tier model its plan includes (plan excludes <models>)`, or `<model>
11
+ is not in <pool>'s plan` for a model the step names), also under a step's
12
+ route and in `run --dry-run`. Before, a routed step read only `no
13
+ enabled pool left has a model on the <tier> tier`.
14
+ - dispatch: a worker can mark a blocker it may not change by starting a `##
15
+ Not done` line with `outside:` (`- outside: tests/router.test.js fails on
16
+ main before this change`); work tasks now ask for this. A step that fails
17
+ `failed-evidence`, `not-produced` or `semantic` with such a line skips its
18
+ one automatic retry and comes back to you at once, its `why` ending `retry
19
+ skipped: the worker reported a blocker outside this step: <item>`; the saved
20
+ attempt lists every such item as `outsideBlockers`. `schema` and process
21
+ failures keep their retry.
22
+ - workflow add: a fragment may carry `blocks: {"<new step id>": ["<existing
23
+ step id>", ...]}` to make existing steps that have not started also wait for
24
+ a step it adds, for example a fix in front of a report (then `step accept`
25
+ the failed step). It is the one change a fragment may make to an existing
26
+ step; a started step, a loop's step, a gate, a loop or a cycle is refused in
27
+ fragment words. A step already blocked behind a failed step stays blocked
28
+ (it is not reset and blocked again). The output prints `waits step <id>
29
+ now also waits for <ids>`, and `--json` carries `waits`.
30
+ - advisories: a new `suite-wider-than-files` advisory names a step that owns
31
+ files but runs a command check naming none of them (a bare `npm test`),
32
+ because a failure elsewhere in the suite would fail a step that cannot fix
33
+ it. Scope the command, or run the suite in a check step. Like every advisory
34
+ it never refuses a program or changes an exit code.
35
+ - advisories: with `--json`, `workflow goal` and `workflow plan validate` no
36
+ longer print advisory text; the advisories are in the JSON as `advisories`
37
+ (the detached launch document always, a foreground result when there are
38
+ any). In human mode `plan validate` now prints them on stderr, as `workflow
39
+ goal` does.
40
+ - run --json: the verdict has top-level `pool` and `model`, the last attempt's
41
+ (null when nothing was dispatched). `pick` keeps them too. Small flat
42
+ objects such as the reasoning record now print on one line, so the verdict
43
+ stays within 60 lines.
44
+ - workflow runs: the goal column skips constraint sentences that follow a
45
+ dropped leading folder (`Work in /abs/path (git branch x). Do not commit.
46
+ Fix the parser crash.` shows `Fix the parser crash.`); a goal that is only
47
+ constraints shows the folder name.
48
+ - strategy: new `bullswarm strategy set-free <allow|never> [--pool <name>]
49
+ --yes` (and `set-free reset --pool <name>`) stops free models being
50
+ suggested, applied or dispatched, for every pool or one; a pool's own
51
+ setting wins. A pool left with only free models for a tier is out of that
52
+ tier with the reason `free models are off for <pool>`: a step no pool can
53
+ take says it in its `why` and `routeCandidates`, `run --dry-run` says it
54
+ too, and `strategy show` prints it under the tier. A model you name
55
+ yourself is still allowed. `strategy show` prints a `free models:` line, and
56
+ the low tier's basis stops putting free models first when every pool has
57
+ them off.
58
+ - providers: a connector may list `modelPlanSignatures`, the text its CLI
59
+ prints when the subscription's plan does not include the requested model
60
+ (Command Code: `MODEL_NOT_IN_PLAN`). Only the provider's error output is
61
+ read: stderr, and on an event stream the records the provider flags as
62
+ errors (read even when the worker exits 0), so a worker's reply that quotes
63
+ the code, even on a line starting `Error:` or in a clean final record, is
64
+ still a reply. A hit is the new failure kind
65
+ `model-not-in-plan` (a process failure, labelled `model not in plan`): the
66
+ step's retry goes to another pool, and `state.strategy.planExcludedModels`
67
+ records the model for that pool. The pool keeps running its other models;
68
+ strategy, rungs, pins and dispatch pass that model over on that pool only.
69
+ `strategy show` prints `not in plan: <pool>: <model>`. `strategy
70
+ include-model <model>`, a changed `strategy set-subscription`, or turning
71
+ the model back on for the pool clears it, and `workflow resume` reruns such
72
+ a step.
73
+ - command-code: models are ranked by family (astra, sol, gpt, terra, luna,
74
+ mini, fable, opus, sonnet, haiku, grok), so a newly listed model is ranked
75
+ the day it appears and the newer version of a family ranks first
76
+ (`gpt-6.1-sol` above `gpt-5.6-sol`, `claude-opus-5-5` above
77
+ `claude-opus-4-7`). Dated prices and the deliberate overrides (`gpt-5.4` and
78
+ `grok-4.5` medium, `minimax-m3` high) stay as rows; kimi and glm stay
79
+ unranked.
80
+
81
+ ## 0.38.0 — The removals release: v3 only, saved runs view-only, less code
82
+
83
+ - removed: the mechanisms program v3 replaced (0.37.0) no longer run. The
84
+ preflight scout (`--scout`, `--no-scout`), the dispatched Workflow Planner
85
+ (`--orchestrator`, `--orchestrator-model`, `--orchestrator-strict`,
86
+ `--strict-orchestrator`, `--suggested-plan`, `--planner-reasoning`), the
87
+ caller-planner gap turn (`workflow plan show`, `workflow plan submit`),
88
+ whole-plan revise (`workflow plan export`, `workflow plan revise`),
89
+ `workflow plan contract --v2`, the v2 repair loop (`defaults.verifyRounds`),
90
+ the kernel-written digest and review tasks, and goal requirement
91
+ extraction are gone. For one release each removed flag and verb exits 2
92
+ with one sentence naming its replacement: a first step your later steps
93
+ depend on instead of the scout; a program you write, or a step whose answer
94
+ is a list of steps appended with `workflow add --from-answer`, instead of
95
+ the planner; `workflow add` and `workflow step rerun` instead of revise.
96
+ - v2 programs: a `bullswarm.workflow.program.v2` is refused by
97
+ `workflow goal --program` and `workflow plan validate` with exit 2, before
98
+ any run folder exists: `bullswarm.workflow.program.v2 is no longer accepted
99
+ for a new run; write a program.v3 (bullswarm workflow plan contract) and
100
+ check it (bullswarm workflow plan validate --program <file.json>)`.
101
+ - saved runs: only a run marked `programFormat: 3` is driven. A run an
102
+ earlier Bullswarm started (v2 and every earlier format, legacy included)
103
+ is view-only: `workflow resume`, `goal --resume`, `pause`, `steer`,
104
+ `step rerun`/`accept`/`restart` and the removed plan verbs exit 2 with
105
+ `run <id> was started by an earlier Bullswarm and is view-only; start a new
106
+ run: …` and change nothing (`workflow add` gives the same sentence with
107
+ exit 1); the kernel refuses to resume one too. `workflow cancel` still finalizes a v2 run left live. Every saved run
108
+ stays listed, shown and counted as before (`runs list --all`, `show`,
109
+ `result`, `--summary`, `watch`, the dashboard, `stats`, History), and a
110
+ saved repair-loop run keeps its recorded `next` text.
111
+ - legacy runs: `workflow runs show`, `runs result`, `watch` and `tui` on a
112
+ run from before 0.27 print a bounded read-only summary built from its saved
113
+ files instead of refusing; minutes or cost the files do not hold stay
114
+ unknown.
115
+ - dispatch: every step follows the one failure rule: one automatic retry,
116
+ then the caller; a usage limit goes straight to the caller; a transient
117
+ rate limit backs off on its own pool only. The rules older runs used are
118
+ gone with those runs' execution.
119
+ - `workflow capabilities` no longer lists `actionRoles` or the dispatched
120
+ planner mode; roles and kinds stay only as the reader of saved v2 runs.
121
+ - The text judge (`judgeContent`) stays: `pools probe`, `health` and
122
+ `contentUsableDespiteExit` in `bullswarm run --json` still use it.
123
+ - dispatch: a worker a signal stopped (the kernel's SIGTERM, a timeout) is
124
+ recorded as `interrupted`, not as a `process` failure, on v3 steps too.
125
+
5
126
  ## 0.37.3 — Price-band model picks, tier-level reasoning, per-step model
6
127
 
7
128
  - strategy: a connector can opt into a price band (`priceBand: true`; Codex