bullswarm 0.37.2 → 0.38.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (70) hide show
  1. package/AGENTS.md +45 -52
  2. package/CHANGELOG.md +75 -0
  3. package/data/openrouter-benchmarks.json +11649 -11485
  4. package/docs/design/0.38.0-removals.md +123 -0
  5. package/docs/guide/observing.md +2 -2
  6. package/docs/guide/playbook.md +13 -11
  7. package/docs/guide/routing.md +17 -21
  8. package/docs/guide/workflows.md +33 -74
  9. package/docs/reference/cli.md +37 -114
  10. package/docs/reference/configuration.md +1 -1
  11. package/docs/reference/program.md +5 -6
  12. package/docs/reference/providers.md +3 -3
  13. package/docs/reference/result.md +15 -12
  14. package/package.json +1 -1
  15. package/skill/SKILL.md +4 -5
  16. package/skill/references/operations.md +113 -236
  17. package/skill/references/patterns.md +5 -5
  18. package/skill/references/program.md +12 -3
  19. package/src/help.js +93 -207
  20. package/src/lib/cli-flags.js +29 -23
  21. package/src/lib/config.js +3 -0
  22. package/src/lib/model-pin.js +68 -0
  23. package/src/lib/price-band.js +90 -0
  24. package/src/lib/run-step.js +3 -0
  25. package/src/lib/strategy.js +57 -2
  26. package/src/provider-cli.js +4 -0
  27. package/src/providers/_schema.json +2 -0
  28. package/src/providers/claude-code/connector-history.json +2 -0
  29. package/src/providers/claude-code/connector.json +1 -1
  30. package/src/providers/claude-code/provider.mjs +9 -3
  31. package/src/providers/codex/connector-history.json +12 -1
  32. package/src/providers/codex/connector.json +27 -1
  33. package/src/providers/grok/connector-history.json +6 -2
  34. package/src/providers/grok/connector.json +7 -7
  35. package/src/workflow/action-validator.js +9 -2
  36. package/src/workflow/attempt-bytes.js +5 -7
  37. package/src/workflow/caller-planner.js +9 -135
  38. package/src/workflow/cli-capabilities.js +4 -9
  39. package/src/workflow/cli-goal-document.js +9 -24
  40. package/src/workflow/cli-goal.js +24 -64
  41. package/src/workflow/cli-plan.js +54 -421
  42. package/src/workflow/cli-pool-checks.js +27 -0
  43. package/src/workflow/cli-program-checks.js +17 -13
  44. package/src/workflow/cli-run-lookup.js +61 -5
  45. package/src/workflow/cli-run-verbs.js +36 -27
  46. package/src/workflow/cli-step-verbs.js +5 -4
  47. package/src/workflow/cli-steps.js +23 -4
  48. package/src/workflow/cli.js +14 -6
  49. package/src/workflow/contract-v3.js +1 -0
  50. package/src/workflow/goal.js +0 -64
  51. package/src/workflow/home-model.js +1 -1
  52. package/src/workflow/kernel-resume.js +1 -16
  53. package/src/workflow/legacy-verification.js +514 -0
  54. package/src/workflow/pick-preview.js +10 -4
  55. package/src/workflow/program-v3.js +4 -1
  56. package/src/workflow/revision-v3.js +19 -3
  57. package/src/workflow/run-control.js +5 -32
  58. package/src/workflow/run-model.js +1 -1
  59. package/src/workflow/run-view.js +1 -1
  60. package/src/workflow/runs-cli.js +13 -13
  61. package/src/workflow/step-prompts.js +3 -89
  62. package/src/workflow/v2-dispatch.js +194 -536
  63. package/src/workflow/v2-outcome.js +1 -1
  64. package/src/workflow/v2-planner.js +2 -113
  65. package/src/workflow/v2-revision.js +6 -2
  66. package/src/workflow/v2-runtime.js +82 -726
  67. package/src/workflow/v2-state.js +3 -3
  68. package/src/workflow/watch-cli.js +16 -12
  69. package/src/workflow/workflow-flags.js +29 -2
  70. package/src/workflow/verify-rounds.js +0 -1228
package/AGENTS.md CHANGED
@@ -14,31 +14,27 @@ The core is being redesigned around facts-only mechanics, four mandatory
14
14
  principles, caller-chosen options and a pattern library, for any kind of work
15
15
  rather than code only. The design, the decisions taken and the staged build
16
16
  plan are in `docs/design/redesign-mechanics-principles-options.md`, and draft
17
- pattern cards are in `docs/design/patterns/`. Build in the plan's stage order.
18
- The doctrine below stays in force until the stage that changes an item lands.
19
- The redesign rewords items 1, 5, 6 and 7, and each stage updates this file.
20
- Stage 1 (step vocabulary) has landed: program-mode steps may state a role and a
21
- deliverable, each kind belongs to one role and keeps its exact routing, and
22
- the no-op gate is now 'declared deliverable not produced' (failure kind
23
- `not-produced`), measured over the whole step. Stage 2 (evidence v1) has landed:
24
- program steps may declare command and schema `evidence` that the kernel runs
25
- after the worker, a failure is `failed-evidence` with one same-pool retry, and
26
- finished steps in new runs are labelled `proven by …` or `finished · unproven`.
27
- Stage 3 (failure rule and routing constraints) has landed: one automatic retry
28
- per step, then the caller; a usage limit goes straight to the caller, from a
29
- step, the dispatched planner or the preflight scout alike; the
30
- needs-you block with `step rerun --avoid` and `step accept`; the per-step
31
- `route`; `verifyRounds` counts fixes (default 1); reviews are placed only by
32
- route. Runs started earlier keep their rules (`features.json`).
33
- 0.37.0 (program v3, the generic model) has landed: a step is a run, and a
17
+ pattern cards are in `docs/design/patterns/`. Each release updates this file.
18
+
19
+ 0.37.0 (program v3, the generic model) landed: a step is a run, and a
34
20
  workflow composes steps, phases, gates and loops. `bullswarm run` is a
35
- one-step workflow. New programs are `bullswarm.workflow.program.v3`: a step
36
- passes by facts (clean exit, deliverable produced, evidence passed, answer
37
- matching its schema), a gate waits for `workflow continue`, a loop reruns its
38
- steps until one condition holds (at most 5 rounds), and new work is added
39
- with `workflow add`, never by editing the run's steps. v3 reports facts per
40
- step and has no requirement IDs, `evidenceFor` or `verifyRounds`. v2 programs
41
- and saved runs keep running and replaying as before.
21
+ one-step workflow. A step passes by facts (clean exit, deliverable produced,
22
+ evidence passed, answer matching its schema), a gate waits for
23
+ `workflow continue`, a loop reruns its steps until one condition holds (at
24
+ most 5 rounds), and new work is added with `workflow add`, never by editing
25
+ the run's steps. v3 reports facts per step and has no requirement IDs,
26
+ `evidenceFor` or `verifyRounds`.
27
+
28
+ 0.38.0 (the removals release, `docs/design/0.38.0-removals.md`) deleted the
29
+ mechanisms v3 replaced: the preflight scout, the dispatched planner
30
+ (`--orchestrator`), the caller-planner gap turn (`plan show`/`plan submit`),
31
+ whole-plan revise (`plan export`/`plan revise`), the v2 repair loop, the
32
+ kernel-written digest and review tasks, and the old dispatch rules. New
33
+ programs are `bullswarm.workflow.program.v3` only; a v2 program is refused
34
+ before a run folder exists. Only a run marked `programFormat: 3` is driven.
35
+ Every other saved run (v2, stage 1-3, legacy) is view-only: it stays
36
+ listable, showable and countable, and every driving command refuses it
37
+ before doing anything. Their readers and stored formats are unchanged.
42
38
 
43
39
  ## Non-negotiable doctrine
44
40
 
@@ -76,34 +72,31 @@ and saved runs keep running and replaying as before.
76
72
  pool inside a run; only a transient rate limit backs off on the same pool,
77
73
  at most twice (20 s, then 60 s, or a named wait of at most 2 minutes), then
78
74
  goes to the caller.
79
- Only a failed step's dependents wait. The dispatched planner and the
80
- preflight scout follow the same usage-limit rule (`usageLimitsToCaller` in
81
- `dispatchV2Action`): they stop and the run tells the caller, with no
82
- automatic move to another pool.
83
- 6. Review is a caller option, recorded as a fact. A step naming requirements
84
- in `evidenceFor` is dispatched under the evidence contract and judges them
85
- from the durable artifact. Where it runs is the caller's choice through
86
- `route` (`independentOf`, `providers`, `pools`); Bullswarm never moves a
87
- review on its own, and records who reviewed (pool, model, provider) and
88
- whether that provider also wrote the work. A caller's `step accept` is
89
- recorded as evidence `choice` and never makes a requirement verified.
90
- Runs started before this rule keep automatic writer avoidance (R12/R13).
91
- 7. New goal workflows are caller-planned programs in a shared workspace.
92
- `bullswarm workflow goal --program` executes the graph; `--orchestrator`
93
- explicitly delegates planning. File territories are advisory scheduling
94
- hints. In a v3 program (the format for new work) a check is an ordinary
95
- step, and fixing until it passes is a loop the caller declares (`loops`,
96
- `until` one condition, `maxRounds` 1-5); a loop out of rounds or a gate
97
- waits for the caller (`workflow continue`), and a failed step goes to the
98
- caller through the watcher's needs-you block: rerun elsewhere, add steps
99
- (`workflow add`), take over, or accept anyway. v3 validate refuses
100
- `defaults.verifyRounds`. In a v2 program a failed check gets one fix step
101
- and one re-review (`defaults.verifyRounds`, default 1, 0-3), then the
102
- caller, and `verified` separately records requirement evidence.
103
- `--isolation` opts into strict per-worker worktrees. Saved V2 runs
104
- preserve their original semantics (`features.json`).
105
- 8. Historical authored-graph runs remain visible as read-only `legacy` rows.
106
- Their executor was removed in 0.27.0; driving commands fail closed before
75
+ Only a failed step's dependents wait.
76
+ 6. Review is a caller option, recorded as a fact. A check is an ordinary
77
+ step with an `answer` and/or `evidence`, and it passes by facts only.
78
+ Where it runs is the caller's choice through `route` (`independentOf`,
79
+ `providers`, `pools`); Bullswarm never moves a review on its own, and
80
+ records who reviewed (pool, model, provider) and whether that provider
81
+ also wrote the work. A caller's `step accept` is recorded as evidence
82
+ `choice`. Saved v2 runs keep showing their requirement verdicts
83
+ (`evidenceFor`, `verified`) as they were recorded.
84
+ 7. New goal workflows are caller-planned v3 programs in a shared workspace.
85
+ `bullswarm workflow goal --program` executes the graph; the caller writes
86
+ the plan (or a step whose answer is a list of steps, appended with
87
+ `workflow add --from-answer`). File territories are advisory scheduling
88
+ hints. A check is an ordinary step, and fixing until it passes is a loop
89
+ the caller declares (`loops`, `until` one condition, `maxRounds` 1-5); a
90
+ loop out of rounds or a gate waits for the caller (`workflow continue`),
91
+ and a failed step goes to the caller through the watcher's needs-you
92
+ block: rerun elsewhere, add steps (`workflow add`), take over, or accept
93
+ anyway. v3 validate refuses `defaults.verifyRounds`. `--isolation` opts
94
+ into strict per-worker worktrees. A saved run that is not v3 is view-only
95
+ (0.38.0): its original semantics are kept for display (`features.json`),
96
+ and nothing drives it again.
97
+ 8. Historical authored-graph runs remain visible as read-only `legacy` rows,
98
+ and `runs show`/`result`/`watch` print a bounded summary of them. Their
99
+ executor was removed in 0.27.0; driving commands fail closed before
107
100
  dispatch and historical run directories remain untouched.
108
101
 
109
102
  ## Development
package/CHANGELOG.md CHANGED
@@ -2,6 +2,81 @@
2
2
 
3
3
  ## Unreleased
4
4
 
5
+ ## 0.38.0 — The removals release: v3 only, saved runs view-only, less code
6
+
7
+ - removed: the mechanisms program v3 replaced (0.37.0) no longer run. The
8
+ preflight scout (`--scout`, `--no-scout`), the dispatched Workflow Planner
9
+ (`--orchestrator`, `--orchestrator-model`, `--orchestrator-strict`,
10
+ `--strict-orchestrator`, `--suggested-plan`, `--planner-reasoning`), the
11
+ caller-planner gap turn (`workflow plan show`, `workflow plan submit`),
12
+ whole-plan revise (`workflow plan export`, `workflow plan revise`),
13
+ `workflow plan contract --v2`, the v2 repair loop (`defaults.verifyRounds`),
14
+ the kernel-written digest and review tasks, and goal requirement
15
+ extraction are gone. For one release each removed flag and verb exits 2
16
+ with one sentence naming its replacement: a first step your later steps
17
+ depend on instead of the scout; a program you write, or a step whose answer
18
+ is a list of steps appended with `workflow add --from-answer`, instead of
19
+ the planner; `workflow add` and `workflow step rerun` instead of revise.
20
+ - v2 programs: a `bullswarm.workflow.program.v2` is refused by
21
+ `workflow goal --program` and `workflow plan validate` with exit 2, before
22
+ any run folder exists: `bullswarm.workflow.program.v2 is no longer accepted
23
+ for a new run; write a program.v3 (bullswarm workflow plan contract) and
24
+ check it (bullswarm workflow plan validate --program <file.json>)`.
25
+ - saved runs: only a run marked `programFormat: 3` is driven. A run an
26
+ earlier Bullswarm started (v2 and every earlier format, legacy included)
27
+ is view-only: `workflow resume`, `goal --resume`, `pause`, `steer`,
28
+ `step rerun`/`accept`/`restart` and the removed plan verbs exit 2 with
29
+ `run <id> was started by an earlier Bullswarm and is view-only; start a new
30
+ run: …` and change nothing (`workflow add` gives the same sentence with
31
+ exit 1); the kernel refuses to resume one too. `workflow cancel` still finalizes a v2 run left live. Every saved run
32
+ stays listed, shown and counted as before (`runs list --all`, `show`,
33
+ `result`, `--summary`, `watch`, the dashboard, `stats`, History), and a
34
+ saved repair-loop run keeps its recorded `next` text.
35
+ - legacy runs: `workflow runs show`, `runs result`, `watch` and `tui` on a
36
+ run from before 0.27 print a bounded read-only summary built from its saved
37
+ files instead of refusing; minutes or cost the files do not hold stay
38
+ unknown.
39
+ - dispatch: every step follows the one failure rule: one automatic retry,
40
+ then the caller; a usage limit goes straight to the caller; a transient
41
+ rate limit backs off on its own pool only. The rules older runs used are
42
+ gone with those runs' execution.
43
+ - `workflow capabilities` no longer lists `actionRoles` or the dispatched
44
+ planner mode; roles and kinds stay only as the reader of saved v2 runs.
45
+ - The text judge (`judgeContent`) stays: `pools probe`, `health` and
46
+ `contentUsableDespiteExit` in `bullswarm run --json` still use it.
47
+ - dispatch: a worker a signal stopped (the kernel's SIGTERM, a timeout) is
48
+ recorded as `interrupted`, not as a `process` failure, on v3 steps too.
49
+
50
+ ## 0.37.3 — Price-band model picks, tier-level reasoning, per-step model
51
+
52
+ - strategy: a connector can opt into a price band (`priceBand: true`; Codex
53
+ does). Each tier then takes the best newest-generation model whose dated
54
+ API price is no more than what served that tier one generation back, at the
55
+ tier's normal reasoning. Codex now suggests gpt-6.1-sol for high (at high
56
+ reasoning) and medium (at medium), since it costs less than gpt-5.6-sol and
57
+ gpt-5.6-terra did; gpt-6-astra costs more than either. Without prices the
58
+ family order decides, as before.
59
+ - codex: dated API prices for gpt-6-astra, gpt-6.1-sol, gpt-6-sol and
60
+ gpt-6-luna, with cache writes at 1.25x input as OpenAI's caching guide
61
+ states.
62
+ - claude: model discovery reads an alias row's model from its display name
63
+ ("Opus 5.5"), as current Claude Code CLIs send it. Before, Opus 5.5, Sonnet
64
+ 5.5 and Haiku 4.5 could go missing, so a tier fell back to an older model or
65
+ had none.
66
+ - claude, grok: each tier runs at its own reasoning level by default (high on
67
+ high, medium on medium, low on low), as Codex already did; it was one level
68
+ higher (xhigh, high, medium). A step's `reasoning` or `--reasoning` still
69
+ sets any level.
70
+ - model: a v3 step may name an exact `model`, and `bullswarm run` takes
71
+ `--model`. Only pools whose model discovery lists it stay eligible, spare
72
+ quota picks among them, and no other model is substituted; validate, goal
73
+ and `workflow add` refuse a model no enabled pool can run, naming each
74
+ pool's reason. `--worker-model` now keeps the same promise: a pool that
75
+ does not list the pinned model is no longer picked for it.
76
+ - workflow step rerun: `--avoid <pool>` works on a program-v3 run. It was
77
+ refused as a step change; now the rerun step's pools route is the one
78
+ change allowed.
79
+
5
80
  ## 0.37.2 — Real-use fixes: goal column, v3 result, routed note
6
81
 
7
82
  - workflow runs: a goal that starts with a labelled folder ("Repo: <path>