bullswarm 0.38.2 → 0.38.5
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/AGENTS.md +3 -1
- package/CHANGELOG.md +56 -0
- package/docs/guide/observing.md +2 -2
- package/docs/guide/routing.md +2 -2
- package/docs/guide/workflows.md +1 -0
- package/docs/reference/cli.md +1 -1
- package/docs/reference/program.md +2 -1
- package/package.json +1 -1
- package/skill/SKILL.md +119 -320
- package/skill/agents/openai.yaml +2 -2
- package/skill/references/operations.md +55 -183
- package/skill/references/patterns.md +20 -4
- package/skill/references/program.md +99 -2
- package/skill/references/recovery.md +257 -0
- package/src/help.js +14 -7
- package/src/lib/cli-flags.js +1 -1
- package/src/workflow/action-validator.js +10 -2
- package/src/workflow/cli-plan.js +12 -1
- package/src/workflow/contract-v3.js +12 -2
- package/src/workflow/gates-loops.js +5 -2
- package/src/workflow/plan-try-checks.js +160 -0
- package/src/workflow/program-v3.js +5 -1
- package/src/workflow/revision-v3.js +2 -2
- package/src/workflow/step-prompts.js +10 -0
- package/src/workflow/step-route.js +29 -0
package/AGENTS.md
CHANGED
|
@@ -123,7 +123,9 @@ bullswarm workflow runs delete <shortId> --yes
|
|
|
123
123
|
## Using bullswarm from another agent
|
|
124
124
|
|
|
125
125
|
If you are an agent that wants to offload bounded work via bullswarm,
|
|
126
|
-
read `skill/SKILL.md` — that's the agent-facing user guide
|
|
126
|
+
read `skill/SKILL.md` — that's the agent-facing user guide: a short entry
|
|
127
|
+
point that links its references (`recovery.md`, `program.md`, `patterns.md`,
|
|
128
|
+
`operations.md`, `providers.md`) for the rest. There are
|
|
127
129
|
exactly two ways to start work, and the caller chooses the shape itself: one
|
|
128
130
|
bounded outcome goes to `bullswarm run`; parallel territories, integration,
|
|
129
131
|
or independent acceptance go to `bullswarm workflow goal` with a program you
|
package/CHANGELOG.md
CHANGED
|
@@ -2,6 +2,62 @@
|
|
|
2
2
|
|
|
3
3
|
## Unreleased
|
|
4
4
|
|
|
5
|
+
## 0.38.5 — plan validate --try-checks (0.38.4 republished)
|
|
6
|
+
|
|
7
|
+
- release: 0.38.4 was tagged but never published to npm (a test read only the `## Unreleased` changelog section, which the release empties); 0.38.5 ships 0.38.4's changes. The docs site builds again: three wrapped lines that began with a `<placeholder>` were read as HTML.
|
|
8
|
+
|
|
9
|
+
## 0.38.4 — plan validate --try-checks
|
|
10
|
+
|
|
11
|
+
- plan validate: `--try-checks` runs each step's command check once, now,
|
|
12
|
+
against the current tree, the way a step runs it (same shell, folder,
|
|
13
|
+
environment and timeout), and prints one `try` line per check with its exit
|
|
14
|
+
and last output line (or `timed out after <n>s`) and the names of any files
|
|
15
|
+
it changed, ignored files and folders outside git included. A check that
|
|
16
|
+
exits non-zero or times out also prints the last 20 lines of its output,
|
|
17
|
+
indented under the try line, so the error is visible. A check that reads the step's output is not tried, and the results never
|
|
18
|
+
change the exit code. Without the flag validate still runs nothing; when the
|
|
19
|
+
program has command checks it now says so (`checks not run · add
|
|
20
|
+
--try-checks …`, and `checks: {tried: false, commands}` in `--json`). Use it
|
|
21
|
+
for checks that are safe to run now: a missing module or command on a `try`
|
|
22
|
+
line means the check itself is wrong.
|
|
23
|
+
- docs: `workflow step accept` on a failed step inside a loop does not end the
|
|
24
|
+
loop; to stop a loop early let it run out of rounds, then
|
|
25
|
+
`workflow continue <run> <loop>` (recorded as condition not met).
|
|
26
|
+
|
|
27
|
+
## 0.38.3 — Blind reviews and a reorganized skill
|
|
28
|
+
|
|
29
|
+
- steps: a v3 step may declare `"blindTo": ["<step id>", ...]`, naming steps
|
|
30
|
+
that run before it (directly or through others). Its task still lists each
|
|
31
|
+
named dependency, but without that step's output file or checked answer,
|
|
32
|
+
and a loop's `Previous round` block leaves out that step's answer and output
|
|
33
|
+
(it still says the step ran, its status and its evidence); the step waits
|
|
34
|
+
for it as before. Use it on a review of a build step: in a real run a
|
|
35
|
+
builder's answer explained a deviation away with a credible reason and the
|
|
36
|
+
reviewer passed with no findings, while the same reviewer without the
|
|
37
|
+
answer caught it. Validate and `workflow add` refuse a non-array, an unknown
|
|
38
|
+
id, a step that does not run before it, or the step itself, like
|
|
39
|
+
`route.independentOf`, which is unchanged: it still only picks the provider.
|
|
40
|
+
`plan validate` prints `blind to <ids>` on the step's line.
|
|
41
|
+
- docs: the review and critique examples in `patterns.md` and `program.md`
|
|
42
|
+
are strict: the contract as numbered checks, any difference a finding even
|
|
43
|
+
when it looks harmless, intended or justified (the caller decides), and an
|
|
44
|
+
answer with `checks` ({id, holds, evidence}) and `passed`, true only when
|
|
45
|
+
every check holds. Find-then-check stays without `blindTo`, since a check
|
|
46
|
+
works from the list the finder answered.
|
|
47
|
+
- skill: reorganized for progressive disclosure. `skill/SKILL.md` is now the
|
|
48
|
+
short entry point a caller reads every time (25,122 bytes before, 12,198
|
|
49
|
+
after): choosing `bullswarm run` or a workflow, one command block for each,
|
|
50
|
+
a "Driving it well" playbook of ten rules from real runs, and a "Read this
|
|
51
|
+
when" index. The rare paths moved to a new `skill/references/recovery.md`
|
|
52
|
+
(15,988 bytes): the failure rule, needs-you blocks and their options, usage
|
|
53
|
+
limits and no free pool, rate-limit backoff, stale steps and `step restart`,
|
|
54
|
+
pause, resume and cancel, and a partial end, with the same real output.
|
|
55
|
+
`operations.md` drops what now lives there (54,517 bytes before, 44,792
|
|
56
|
+
after) and gains the watch modes and the real `add`, `wait` and `continue`
|
|
57
|
+
output; `program.md` gains a "Writing a program" section (18,536 bytes
|
|
58
|
+
before, 21,577 after). Every command, option and output line the old
|
|
59
|
+
entry point named is still in the skill.
|
|
60
|
+
|
|
5
61
|
## 0.38.2 — Three large files split into one-concept modules
|
|
6
62
|
|
|
7
63
|
- internal: the three largest workflow files are split into one-concept
|
package/docs/guide/observing.md
CHANGED
|
@@ -264,8 +264,8 @@ A real run on grok, parked at its gate, on the Run page at 120 columns:
|
|
|
264
264
|
continued without more rounds reads `→ continued by the caller after N of
|
|
265
265
|
max rounds (condition not met)`: it never passed. Each attempt of a
|
|
266
266
|
loop step says its round (`write · round 2`).
|
|
267
|
-
- **A gate row** heads the phase of the steps behind it and reads
|
|
268
|
-
<steps>`, `waiting for you · <note>` with its `continue` command, `passed ·
|
|
267
|
+
- **A gate row** heads the phase of the steps behind it and reads
|
|
268
|
+
`waits after <steps>`, `waiting for you · <note>` with its `continue` command, `passed ·
|
|
269
269
|
continued by the caller`, or `skipped · <condition> does not hold` when its
|
|
270
270
|
`when` condition did not hold.
|
|
271
271
|
- **Answers.** `answer {…}` is the checked answer (it matched the step's schema);
|
package/docs/guide/routing.md
CHANGED
|
@@ -198,8 +198,8 @@ well: as `quota` when every reason is a usage limit, else as `unavailable`. Its
|
|
|
198
198
|
spare: pool-a at its 5-hour limit until <time>; pool-b at its weekly limit
|
|
199
199
|
until <time>`, and `back at` is the earliest known return among them. A retry
|
|
200
200
|
the step was promised (after a crash, a sign-in failure or a failed gate) that
|
|
201
|
-
finds no free pool keeps its own failure and ends its `why` with
|
|
202
|
-
<pool> <reason>; …`. When another pool that can run the step is free, routing
|
|
201
|
+
finds no free pool keeps its own failure and ends its `why` with
|
|
202
|
+
`· no retry: <pool> <reason>; …`. When another pool that can run the step is free, routing
|
|
203
203
|
picks it as usual.
|
|
204
204
|
|
|
205
205
|
A run saved by 0.37.x may end with `the workflow planner stopped on a usage
|
package/docs/guide/workflows.md
CHANGED
|
@@ -108,6 +108,7 @@ bullswarm workflow plan validate "Make the acme tests pass" --cwd=/private/tmp/v
|
|
|
108
108
|
✓ program v3 valid: 2 steps, 0 gates, 1 loop (nothing launched)
|
|
109
109
|
fix build/medium deliverable=files
|
|
110
110
|
check analyze/medium evidence=command answer after fix
|
|
111
|
+
checks not run · add --try-checks to run each command check once now against the current tree (it may take time and must not change files)
|
|
111
112
|
loop until-green steps fix, check · until check's evidence passed · at most 3 rounds
|
|
112
113
|
launch bullswarm workflow goal 'Make the acme tests pass' --cwd /private/tmp/v37fix/acme --program /private/tmp/v37fix/loop.json --json
|
|
113
114
|
```
|
package/docs/reference/cli.md
CHANGED
|
@@ -255,7 +255,7 @@ Trailing `<task text...>` is mutually exclusive with `--prompt` and `--task-file
|
|
|
255
255
|
| `--avoid-provider <provider,...>` | never route to pools of these providers | unset |
|
|
256
256
|
| `--dry-run` | print the kernel's routing pick, the forecast, and the exact command that would be spawned, without spawning, recording a run, registering an assignment, or writing the decision log | off (dispatches for real) |
|
|
257
257
|
| `--no-caller` | accepted and ignored for one release: the calling agent is never a pool of its own run | removed in 0.37.0 |
|
|
258
|
-
| `--json` | print the compact machine-readable verdict; top-level `pool` and `model` name who ran the last attempt (null when nothing was dispatched; `pick` keeps the same two beside the command); its last field, details, is the command for the full record and per-attempt cost (bullswarm workflow runs result <shortId> --json) | human-readable summary ending with the run id and that command |
|
|
258
|
+
| `--json` | print the compact machine-readable verdict; top-level `pool` and `model` name who ran the last attempt (null when nothing was dispatched; `pick` keeps the same two beside the command); its last field, details, is the command for the full record and per-attempt cost (`bullswarm workflow runs result <shortId> --json`) | human-readable summary ending with the run id and that command |
|
|
259
259
|
|
|
260
260
|
A run is a one-step workflow, recorded under `workflows/<id>/` (`bullswarm workflow runs --all`). It gets the step's one automatic retry unless `--no-retry`; a usage limit exits 1 with no retry, and the pool's meter is read again at once, so a window it shows at 100% keeps the pool out of later picks until that window resets. Nothing else about a failed pool is remembered. A build or chore run must change a file (else failure kind `not-produced`). The JSON shape is in [Result envelope](/reference/result).
|
|
261
261
|
|
|
@@ -44,6 +44,7 @@ Steps, gates and loops share one id space: an id is kebab-case and used once.
|
|
|
44
44
|
| `effort` | no | `high`, `medium`, `low`; default by lane: analyze medium, build medium, chore low; a chore step must be low |
|
|
45
45
|
| `reasoning` | no | `low`, `medium`, `high`, `xhigh`, `max`, or `default` (pass nothing): how hard the picked model thinks |
|
|
46
46
|
| `route` | no | `{pools: {use, avoid}, providers: {use, avoid}, independentOf: [step ids]}`: a hard filter applied before quota pacing; `independentOf` names steps this step depends on (directly or through others) whose providers it must not use |
|
|
47
|
+
| `blindTo` | no | step ids this step depends on (directly or through others) whose output file and checked answer it is not handed, in its task or a loop's `Previous round` block; the dependency still orders it. Use it on a review of a build step; leave it off a check that must read the list it checks |
|
|
47
48
|
| `answer` | no | a JSON schema: the worker writes its final answer as JSON to a file Bullswarm names (at most 256 KiB), and that file is checked; a mismatch is failure kind `schema`; the checked answer goes to dependent steps, conditions, `workflow wait`, `watch` and `runs result` |
|
|
48
49
|
| `evidence` | no | up to 5 checks Bullswarm runs after the worker, as in [Evidence](#evidence-command-and-schema) |
|
|
49
50
|
| `deliverable` | no | `files`, `report`, `data`, `media`, `outward`, or `{type, paths}` (default `files` for build and chore, `report` for analyze without an answer, none for analyze with an answer); not produced is failure kind `not-produced` |
|
|
@@ -366,7 +367,7 @@ Checks receive `BULLSWARM_EVIDENCE=1`, `BULLSWARM_STEP_ID`, `BULLSWARM_STEP_OUTP
|
|
|
366
367
|
|
|
367
368
|
Checks must be read-only. Before the first item and after each one, Bullswarm hashes the deliverable: the step's `ownedFiles`, declared deliverable paths and saved response, plus every tracked file and HEAD in an isolated copy or for a build or chore step with no `ownedFiles` (it runs alone). A change fails the item with `changed the deliverable: …`, and the later items do not run. Untracked by-products there are recorded as `touched` rather than failed; declare a new file as a deliverable path when it must be protected. In an isolated copy, after the last item, Bullswarm removes the files the checks created and puts back untracked files they rewrote or deleted (up to 16 MB in total), so none of them is merged back. One it cannot put back is named in an attempt note, `check by-product not restored: <paths>`, and is left out of the ownership check and the merge-back. Elsewhere HEAD is not compared: when it moves while an item runs (another step may have committed), the item records `headMoved: true` as a fact and does not fail.
|
|
368
369
|
|
|
369
|
-
Scope a command to its step. Put a whole-suite command on a step that runs alone or last, and pass `--run` or `CI=1` yourself when a test runner watches files. Run every check once
|
|
370
|
+
Scope a command to its step. Put a whole-suite command on a step that runs alone or last, and pass `--run` or `CI=1` yourself when a test runner watches files. Run every check once before launch (`workflow plan validate --try-checks` runs each command check once in the workspace, for checks that are safe to run now) and give it a generous timeout: changing a wrong check amends the step and reruns its worker. A suite that runs longer than 600 seconds cannot be one item: split it into several items or keep it in the step's prompt. To prove finished work without rerunning it, add a `check` step with its own `evidence`; it does not rerun the work. An old kernel refuses `evidence`; pause, revise, then resume.
|
|
370
371
|
|
|
371
372
|
Each check's result is in the full `bullswarm workflow runs result <id> --json` envelope (not `--summary`) under `actions[].evidenceResults`: `status`, `exit`, `tail`, `why` and the log path. `bullswarm workflow action show <id> <step>` shows the same for each attempt under `attempts[].evidenceResults`. Each item's full output is in `evidence-<step>-attempt-<n>-<k>.log` in the run directory.
|
|
372
373
|
|
package/package.json
CHANGED