bullswarm 0.38.2 → 0.38.5

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/AGENTS.md CHANGED
@@ -123,7 +123,9 @@ bullswarm workflow runs delete <shortId> --yes
123
123
  ## Using bullswarm from another agent
124
124
 
125
125
  If you are an agent that wants to offload bounded work via bullswarm,
126
- read `skill/SKILL.md` — that's the agent-facing user guide. There are
126
+ read `skill/SKILL.md` — that's the agent-facing user guide: a short entry
127
+ point that links its references (`recovery.md`, `program.md`, `patterns.md`,
128
+ `operations.md`, `providers.md`) for the rest. There are
127
129
  exactly two ways to start work, and the caller chooses the shape itself: one
128
130
  bounded outcome goes to `bullswarm run`; parallel territories, integration,
129
131
  or independent acceptance go to `bullswarm workflow goal` with a program you
package/CHANGELOG.md CHANGED
@@ -2,6 +2,62 @@
2
2
 
3
3
  ## Unreleased
4
4
 
5
+ ## 0.38.5 — plan validate --try-checks (0.38.4 republished)
6
+
7
+ - release: 0.38.4 was tagged but never published to npm (a test read only the `## Unreleased` changelog section, which the release empties); 0.38.5 ships 0.38.4's changes. The docs site builds again: three wrapped lines that began with a `<placeholder>` were read as HTML.
8
+
9
+ ## 0.38.4 — plan validate --try-checks
10
+
11
+ - plan validate: `--try-checks` runs each step's command check once, now,
12
+ against the current tree, the way a step runs it (same shell, folder,
13
+ environment and timeout), and prints one `try` line per check with its exit
14
+ and last output line (or `timed out after <n>s`) and the names of any files
15
+ it changed, ignored files and folders outside git included. A check that
16
+ exits non-zero or times out also prints the last 20 lines of its output,
17
+ indented under the try line, so the error is visible. A check that reads the step's output is not tried, and the results never
18
+ change the exit code. Without the flag validate still runs nothing; when the
19
+ program has command checks it now says so (`checks not run · add
20
+ --try-checks …`, and `checks: {tried: false, commands}` in `--json`). Use it
21
+ for checks that are safe to run now: a missing module or command on a `try`
22
+ line means the check itself is wrong.
23
+ - docs: `workflow step accept` on a failed step inside a loop does not end the
24
+ loop; to stop a loop early let it run out of rounds, then
25
+ `workflow continue <run> <loop>` (recorded as condition not met).
26
+
27
+ ## 0.38.3 — Blind reviews and a reorganized skill
28
+
29
+ - steps: a v3 step may declare `"blindTo": ["<step id>", ...]`, naming steps
30
+ that run before it (directly or through others). Its task still lists each
31
+ named dependency, but without that step's output file or checked answer,
32
+ and a loop's `Previous round` block leaves out that step's answer and output
33
+ (it still says the step ran, its status and its evidence); the step waits
34
+ for it as before. Use it on a review of a build step: in a real run a
35
+ builder's answer explained a deviation away with a credible reason and the
36
+ reviewer passed with no findings, while the same reviewer without the
37
+ answer caught it. Validate and `workflow add` refuse a non-array, an unknown
38
+ id, a step that does not run before it, or the step itself, like
39
+ `route.independentOf`, which is unchanged: it still only picks the provider.
40
+ `plan validate` prints `blind to <ids>` on the step's line.
41
+ - docs: the review and critique examples in `patterns.md` and `program.md`
42
+ are strict: the contract as numbered checks, any difference a finding even
43
+ when it looks harmless, intended or justified (the caller decides), and an
44
+ answer with `checks` ({id, holds, evidence}) and `passed`, true only when
45
+ every check holds. Find-then-check stays without `blindTo`, since a check
46
+ works from the list the finder answered.
47
+ - skill: reorganized for progressive disclosure. `skill/SKILL.md` is now the
48
+ short entry point a caller reads every time (25,122 bytes before, 12,198
49
+ after): choosing `bullswarm run` or a workflow, one command block for each,
50
+ a "Driving it well" playbook of ten rules from real runs, and a "Read this
51
+ when" index. The rare paths moved to a new `skill/references/recovery.md`
52
+ (15,988 bytes): the failure rule, needs-you blocks and their options, usage
53
+ limits and no free pool, rate-limit backoff, stale steps and `step restart`,
54
+ pause, resume and cancel, and a partial end, with the same real output.
55
+ `operations.md` drops what now lives there (54,517 bytes before, 44,792
56
+ after) and gains the watch modes and the real `add`, `wait` and `continue`
57
+ output; `program.md` gains a "Writing a program" section (18,536 bytes
58
+ before, 21,577 after). Every command, option and output line the old
59
+ entry point named is still in the skill.
60
+
5
61
  ## 0.38.2 — Three large files split into one-concept modules
6
62
 
7
63
  - internal: the three largest workflow files are split into one-concept
@@ -264,8 +264,8 @@ A real run on grok, parked at its gate, on the Run page at 120 columns:
264
264
  continued without more rounds reads `→ continued by the caller after N of
265
265
  max rounds (condition not met)`: it never passed. Each attempt of a
266
266
  loop step says its round (`write · round 2`).
267
- - **A gate row** heads the phase of the steps behind it and reads `waits after
268
- <steps>`, `waiting for you · <note>` with its `continue` command, `passed ·
267
+ - **A gate row** heads the phase of the steps behind it and reads
268
+ `waits after <steps>`, `waiting for you · <note>` with its `continue` command, `passed ·
269
269
  continued by the caller`, or `skipped · <condition> does not hold` when its
270
270
  `when` condition did not hold.
271
271
  - **Answers.** `answer {…}` is the checked answer (it matched the step's schema);
@@ -198,8 +198,8 @@ well: as `quota` when every reason is a usage limit, else as `unavailable`. Its
198
198
  spare: pool-a at its 5-hour limit until <time>; pool-b at its weekly limit
199
199
  until <time>`, and `back at` is the earliest known return among them. A retry
200
200
  the step was promised (after a crash, a sign-in failure or a failed gate) that
201
- finds no free pool keeps its own failure and ends its `why` with `· no retry:
202
- <pool> <reason>; …`. When another pool that can run the step is free, routing
201
+ finds no free pool keeps its own failure and ends its `why` with
202
+ `· no retry: <pool> <reason>; …`. When another pool that can run the step is free, routing
203
203
  picks it as usual.
204
204
 
205
205
  A run saved by 0.37.x may end with `the workflow planner stopped on a usage
@@ -108,6 +108,7 @@ bullswarm workflow plan validate "Make the acme tests pass" --cwd=/private/tmp/v
108
108
  ✓ program v3 valid: 2 steps, 0 gates, 1 loop (nothing launched)
109
109
  fix build/medium deliverable=files
110
110
  check analyze/medium evidence=command answer after fix
111
+ checks not run · add --try-checks to run each command check once now against the current tree (it may take time and must not change files)
111
112
  loop until-green steps fix, check · until check's evidence passed · at most 3 rounds
112
113
  launch bullswarm workflow goal 'Make the acme tests pass' --cwd /private/tmp/v37fix/acme --program /private/tmp/v37fix/loop.json --json
113
114
  ```
@@ -255,7 +255,7 @@ Trailing `<task text...>` is mutually exclusive with `--prompt` and `--task-file
255
255
  | `--avoid-provider <provider,...>` | never route to pools of these providers | unset |
256
256
  | `--dry-run` | print the kernel's routing pick, the forecast, and the exact command that would be spawned, without spawning, recording a run, registering an assignment, or writing the decision log | off (dispatches for real) |
257
257
  | `--no-caller` | accepted and ignored for one release: the calling agent is never a pool of its own run | removed in 0.37.0 |
258
- | `--json` | print the compact machine-readable verdict; top-level `pool` and `model` name who ran the last attempt (null when nothing was dispatched; `pick` keeps the same two beside the command); its last field, details, is the command for the full record and per-attempt cost (bullswarm workflow runs result <shortId> --json) | human-readable summary ending with the run id and that command |
258
+ | `--json` | print the compact machine-readable verdict; top-level `pool` and `model` name who ran the last attempt (null when nothing was dispatched; `pick` keeps the same two beside the command); its last field, details, is the command for the full record and per-attempt cost (`bullswarm workflow runs result <shortId> --json`) | human-readable summary ending with the run id and that command |
259
259
 
260
260
  A run is a one-step workflow, recorded under `workflows/<id>/` (`bullswarm workflow runs --all`). It gets the step's one automatic retry unless `--no-retry`; a usage limit exits 1 with no retry, and the pool's meter is read again at once, so a window it shows at 100% keeps the pool out of later picks until that window resets. Nothing else about a failed pool is remembered. A build or chore run must change a file (else failure kind `not-produced`). The JSON shape is in [Result envelope](/reference/result).
261
261
 
@@ -44,6 +44,7 @@ Steps, gates and loops share one id space: an id is kebab-case and used once.
44
44
  | `effort` | no | `high`, `medium`, `low`; default by lane: analyze medium, build medium, chore low; a chore step must be low |
45
45
  | `reasoning` | no | `low`, `medium`, `high`, `xhigh`, `max`, or `default` (pass nothing): how hard the picked model thinks |
46
46
  | `route` | no | `{pools: {use, avoid}, providers: {use, avoid}, independentOf: [step ids]}`: a hard filter applied before quota pacing; `independentOf` names steps this step depends on (directly or through others) whose providers it must not use |
47
+ | `blindTo` | no | step ids this step depends on (directly or through others) whose output file and checked answer it is not handed, in its task or a loop's `Previous round` block; the dependency still orders it. Use it on a review of a build step; leave it off a check that must read the list it checks |
47
48
  | `answer` | no | a JSON schema: the worker writes its final answer as JSON to a file Bullswarm names (at most 256 KiB), and that file is checked; a mismatch is failure kind `schema`; the checked answer goes to dependent steps, conditions, `workflow wait`, `watch` and `runs result` |
48
49
  | `evidence` | no | up to 5 checks Bullswarm runs after the worker, as in [Evidence](#evidence-command-and-schema) |
49
50
  | `deliverable` | no | `files`, `report`, `data`, `media`, `outward`, or `{type, paths}` (default `files` for build and chore, `report` for analyze without an answer, none for analyze with an answer); not produced is failure kind `not-produced` |
@@ -366,7 +367,7 @@ Checks receive `BULLSWARM_EVIDENCE=1`, `BULLSWARM_STEP_ID`, `BULLSWARM_STEP_OUTP
366
367
 
367
368
  Checks must be read-only. Before the first item and after each one, Bullswarm hashes the deliverable: the step's `ownedFiles`, declared deliverable paths and saved response, plus every tracked file and HEAD in an isolated copy or for a build or chore step with no `ownedFiles` (it runs alone). A change fails the item with `changed the deliverable: …`, and the later items do not run. Untracked by-products there are recorded as `touched` rather than failed; declare a new file as a deliverable path when it must be protected. In an isolated copy, after the last item, Bullswarm removes the files the checks created and puts back untracked files they rewrote or deleted (up to 16 MB in total), so none of them is merged back. One it cannot put back is named in an attempt note, `check by-product not restored: <paths>`, and is left out of the ownership check and the merge-back. Elsewhere HEAD is not compared: when it moves while an item runs (another step may have committed), the item records `headMoved: true` as a fact and does not fail.
368
369
 
369
- Scope a command to its step. Put a whole-suite command on a step that runs alone or last, and pass `--run` or `CI=1` yourself when a test runner watches files. Run every check once by hand before launch and give it a generous timeout: changing a wrong check amends the step and reruns its worker. A suite that runs longer than 600 seconds cannot be one item: split it into several items or keep it in the step's prompt. To prove finished work without rerunning it, add a `check` step with its own `evidence`; it does not rerun the work. An old kernel refuses `evidence`; pause, revise, then resume.
370
+ Scope a command to its step. Put a whole-suite command on a step that runs alone or last, and pass `--run` or `CI=1` yourself when a test runner watches files. Run every check once before launch (`workflow plan validate --try-checks` runs each command check once in the workspace, for checks that are safe to run now) and give it a generous timeout: changing a wrong check amends the step and reruns its worker. A suite that runs longer than 600 seconds cannot be one item: split it into several items or keep it in the step's prompt. To prove finished work without rerunning it, add a `check` step with its own `evidence`; it does not rerun the work. An old kernel refuses `evidence`; pause, revise, then resume.
370
371
 
371
372
  Each check's result is in the full `bullswarm workflow runs result <id> --json` envelope (not `--summary`) under `actions[].evidenceResults`: `status`, `exit`, `tail`, `why` and the log path. `bullswarm workflow action show <id> <step>` shows the same for each attempt under `attempts[].evidenceResults`. Each item's full output is in `evidence-<step>-attempt-<n>-<k>.log` in the run directory.
372
373
 
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "bullswarm",
3
- "version": "0.38.2",
3
+ "version": "0.38.5",
4
4
  "description": "Route work across coding-agent CLI subscriptions — paced by live quota meters, verified by content, never trusting exit codes.",
5
5
  "type": "module",
6
6
  "bin": {