bullswarm 0.38.2 → 0.38.3

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/AGENTS.md CHANGED
@@ -123,7 +123,9 @@ bullswarm workflow runs delete <shortId> --yes
123
123
  ## Using bullswarm from another agent
124
124
 
125
125
  If you are an agent that wants to offload bounded work via bullswarm,
126
- read `skill/SKILL.md` — that's the agent-facing user guide. There are
126
+ read `skill/SKILL.md` — that's the agent-facing user guide: a short entry
127
+ point that links its references (`recovery.md`, `program.md`, `patterns.md`,
128
+ `operations.md`, `providers.md`) for the rest. There are
127
129
  exactly two ways to start work, and the caller chooses the shape itself: one
128
130
  bounded outcome goes to `bullswarm run`; parallel territories, integration,
129
131
  or independent acceptance go to `bullswarm workflow goal` with a program you
package/CHANGELOG.md CHANGED
@@ -2,6 +2,40 @@
2
2
 
3
3
  ## Unreleased
4
4
 
5
+ ## 0.38.3 — Blind reviews and a reorganized skill
6
+
7
+ - steps: a v3 step may declare `"blindTo": ["<step id>", ...]`, naming steps
8
+ that run before it (directly or through others). Its task still lists each
9
+ named dependency, but without that step's output file or checked answer,
10
+ and a loop's `Previous round` block leaves out that step's answer and output
11
+ (it still says the step ran, its status and its evidence); the step waits
12
+ for it as before. Use it on a review of a build step: in a real run a
13
+ builder's answer explained a deviation away with a credible reason and the
14
+ reviewer passed with no findings, while the same reviewer without the
15
+ answer caught it. Validate and `workflow add` refuse a non-array, an unknown
16
+ id, a step that does not run before it, or the step itself, like
17
+ `route.independentOf`, which is unchanged: it still only picks the provider.
18
+ `plan validate` prints `blind to <ids>` on the step's line.
19
+ - docs: the review and critique examples in `patterns.md` and `program.md`
20
+ are strict: the contract as numbered checks, any difference a finding even
21
+ when it looks harmless, intended or justified (the caller decides), and an
22
+ answer with `checks` ({id, holds, evidence}) and `passed`, true only when
23
+ every check holds. Find-then-check stays without `blindTo`, since a check
24
+ works from the list the finder answered.
25
+ - skill: reorganized for progressive disclosure. `skill/SKILL.md` is now the
26
+ short entry point a caller reads every time (25,122 bytes before, 12,198
27
+ after): choosing `bullswarm run` or a workflow, one command block for each,
28
+ a "Driving it well" playbook of ten rules from real runs, and a "Read this
29
+ when" index. The rare paths moved to a new `skill/references/recovery.md`
30
+ (15,988 bytes): the failure rule, needs-you blocks and their options, usage
31
+ limits and no free pool, rate-limit backoff, stale steps and `step restart`,
32
+ pause, resume and cancel, and a partial end, with the same real output.
33
+ `operations.md` drops what now lives there (54,517 bytes before, 44,792
34
+ after) and gains the watch modes and the real `add`, `wait` and `continue`
35
+ output; `program.md` gains a "Writing a program" section (18,536 bytes
36
+ before, 21,577 after). Every command, option and output line the old
37
+ entry point named is still in the skill.
38
+
5
39
  ## 0.38.2 — Three large files split into one-concept modules
6
40
 
7
41
  - internal: the three largest workflow files are split into one-concept
@@ -44,6 +44,7 @@ Steps, gates and loops share one id space: an id is kebab-case and used once.
44
44
  | `effort` | no | `high`, `medium`, `low`; default by lane: analyze medium, build medium, chore low; a chore step must be low |
45
45
  | `reasoning` | no | `low`, `medium`, `high`, `xhigh`, `max`, or `default` (pass nothing): how hard the picked model thinks |
46
46
  | `route` | no | `{pools: {use, avoid}, providers: {use, avoid}, independentOf: [step ids]}`: a hard filter applied before quota pacing; `independentOf` names steps this step depends on (directly or through others) whose providers it must not use |
47
+ | `blindTo` | no | step ids this step depends on (directly or through others) whose output file and checked answer it is not handed, in its task or a loop's `Previous round` block; the dependency still orders it. Use it on a review of a build step; leave it off a check that must read the list it checks |
47
48
  | `answer` | no | a JSON schema: the worker writes its final answer as JSON to a file Bullswarm names (at most 256 KiB), and that file is checked; a mismatch is failure kind `schema`; the checked answer goes to dependent steps, conditions, `workflow wait`, `watch` and `runs result` |
48
49
  | `evidence` | no | up to 5 checks Bullswarm runs after the worker, as in [Evidence](#evidence-command-and-schema) |
49
50
  | `deliverable` | no | `files`, `report`, `data`, `media`, `outward`, or `{type, paths}` (default `files` for build and chore, `report` for analyze without an answer, none for analyze with an answer); not produced is failure kind `not-produced` |
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "bullswarm",
3
- "version": "0.38.2",
3
+ "version": "0.38.3",
4
4
  "description": "Route work across coding-agent CLI subscriptions — paced by live quota meters, verified by content, never trusting exit codes.",
5
5
  "type": "module",
6
6
  "bin": {
package/skill/SKILL.md CHANGED
@@ -1,6 +1,6 @@
1
1
  ---
2
2
  name: bullswarm
3
- description: Delegate bounded work through Bullswarm, one quota-routed task (bullswarm run) or a workflow you write from steps, phases, gates and loops (bullswarm workflow goal --program); watch it as closely as you choose, and extend it with workflow add. Use for /bullswarm, offloading, independent checks, or requested multi-agent execution.
3
+ description: Load before you hand work to another coding agent through the Bullswarm CLI, as one quota-routed task (bullswarm run) or a workflow you write from steps, phases, gates and loops (bullswarm workflow goal --program), and whenever you watch, answer, extend or recover such a run. Use for /bullswarm, offloading, independent checks or reviews, and requested multi-agent execution; not for work you are asked to do yourself.
4
4
  ---
5
5
 
6
6
  # Bullswarm
@@ -52,13 +52,12 @@ changes. Use `--task-file` for long text. Options you will use:
52
52
  Read the verdict. `ok: true` means read `outFile` (and `answer`) and check the
53
53
  content before you use it. `ok: false` means inspect and report the failure;
54
54
  `failureKind` names it and `why` says it in one line. `pool` and `model` name
55
- who ran it; `shortId` names the run;
56
- the last field, `details`, is the command for its record and cost. A
57
- build or chore run that changes no file (files git ignores count) fails
58
- `not-produced`, also outside git. Do not run `doctor` unless dispatch reports
59
- a readiness problem.
55
+ who ran it; `shortId` names the run; the last field, `details`, is the command
56
+ for its record and cost. A build or chore run that changes no file (files git
57
+ ignores count) fails `not-produced`, also outside git. Do not run `doctor`
58
+ unless dispatch reports a readiness problem.
60
59
 
61
- ## 3. A workflow: four blocks
60
+ ## 3. A workflow: write, validate, launch, watch
62
61
 
63
62
  | Block | What it is | Declared as |
64
63
  |---|---|---|
@@ -67,99 +66,31 @@ a readiness problem.
67
66
  | Gate | the run stops there and waits for you, or only when an answer says so | an entry in `gates`: `{id, dependsOn, when?, note?}` |
68
67
  | Loop | steps that repeat until one step's answer (or evidence) says stop, at most `maxRounds` (1-5) | an entry in `loops`: `{id, steps, until, maxRounds}` |
69
68
 
70
- A program that researches in parallel, rewrites a brief until an independent
71
- critique passes, waits for you, then publishes (`plan validate` accepts it: 5
72
- steps, 1 gate, 1 loop):
69
+ A build, then a strict review by another provider that does not read the
70
+ builder's own account (playbook rule 6):
73
71
 
74
72
  ```json
75
73
  {
76
74
  "schemaVersion": "bullswarm.workflow.program.v3",
77
75
  "steps": [
78
- { "id": "search-a", "phase": "research", "prompt": "In /work/acme, collect the claims sources/a.md makes about widgets.", "answer": { "type": "object", "required": ["claims"], "properties": { "claims": { "type": "array", "items": { "type": "string" } } } } },
79
- { "id": "search-b", "phase": "research", "prompt": "In /work/acme, collect the claims sources/b.md makes about widgets.", "answer": { "type": "object", "required": ["claims"], "properties": { "claims": { "type": "array", "items": { "type": "string" } } } } },
80
- { "id": "draft", "phase": "writing", "dependsOn": ["search-a", "search-b"], "lane": "build", "files": ["brief.md"], "prompt": "In /work/acme, write brief.md from the claims your dependencies answered. From round 2 on, fix the problems the previous critique listed." },
81
- { "id": "critique", "phase": "writing", "dependsOn": ["draft"], "route": { "independentOf": ["draft"] }, "prompt": "In /work/acme, check every claim in brief.md against sources/. List only problems a line of sources/ shows. Answer passed true when you list none.", "answer": { "type": "object", "required": ["passed", "problems"], "properties": { "passed": { "type": "boolean" }, "problems": { "type": "array", "items": { "type": "string" } } } } },
82
- { "id": "post", "phase": "publish", "dependsOn": ["approve"], "deliverable": "outward", "retry": 0, "prompt": "Publish /work/acme/brief.md to the acme wiki, and list the page you created." }
83
- ],
84
- "loops": [{ "id": "polish", "steps": ["draft", "critique"], "until": { "step": "critique", "field": "passed" }, "maxRounds": 2 }],
85
- "gates": [{ "id": "approve", "dependsOn": ["polish"], "note": "Read brief.md and decide whether to publish it" }]
76
+ { "id": "build", "lane": "build", "files": ["src/csv-writer.js", "tests/csv-writer.test.js"],
77
+ "prompt": "In /work/acme, add src/csv-writer.js: writeCsv(rows) returns RFC 4180 CSV text. Add its tests in tests/csv-writer.test.js.",
78
+ "evidence": [{ "type": "command", "cmd": "node --test tests/csv-writer.test.js", "timeoutSec": 300 }] },
79
+ { "id": "review", "dependsOn": ["build"], "route": { "independentOf": ["build"] }, "blindTo": ["build"],
80
+ "prompt": "In /work/acme, review src/csv-writer.js against: 1. a field holding a comma, a quote or a line break is quoted, and a quote inside it is doubled; 2. every line ends with CRLF. Any difference is a finding, even if justified by the author. Record each check as holds true or false with its evidence. Answer passed true only when every check holds. Change no file.",
81
+ "answer": { "type": "object", "required": ["checks", "passed"], "properties": {
82
+ "checks": { "type": "array", "items": { "type": "object", "required": ["id", "holds", "evidence"], "properties": { "id": { "type": "string" }, "holds": { "type": "boolean" }, "evidence": { "type": "string" } } } },
83
+ "passed": { "type": "boolean" } } } }
84
+ ]
86
85
  }
87
86
  ```
88
87
 
89
- [program.md](references/program.md) lists every field.
90
- [patterns.md](references/patterns.md) has five workflows to copy: find then
91
- check each finding, fix until a check passes, draft to publish, parallel
92
- slices, and a triage (usually one run).
93
-
94
- ### Writing the program
95
-
96
- - **Plan as far ahead as you know.** Declare the steps, gates and loops you can
97
- see now. Where the next part depends on an answer (one check per finding, one
98
- build step per slice), stop there and add it later with `workflow add`.
99
- - **A step passes by facts.** Its worker ended cleanly, its deliverable was
100
- produced, its evidence passed, and its answer (when declared) matched the
101
- schema. Declare an `answer` whenever you, a later step or a condition needs
102
- data from the step; a dependent step is handed its dependencies' checked
103
- answer files.
104
- - **The one condition form.** A gate's `when` and a loop's `until` read one
105
- value: `{"step": "critique", "field": "passed"}` (a boolean the step's answer
106
- schema requires; add `"equals": false` to invert it) or `{"step": "check",
107
- "evidence": "passed"}`. No expressions and no else: anything more is your
108
- call, with `workflow wait` and `workflow add`.
109
- - **Loops run every step in every round**, and read their condition when the
110
- round is over. Put the deciding step last, and give every writer in the loop
111
- work each round: a build or chore step that changes no file fails
112
- `not-produced` (rewrite the draft from the critique; do not put a revise
113
- step after a critique that may pass the first time). From round 2 on, each
114
- step's task carries a `Previous round` block with the last round's answers
115
- and evidence. When a loop's `until` is the evidence form (`{"step": "check",
116
- "evidence": "passed"}`), a failed check on that step reads as "not passed"
117
- and the loop goes on. With the field form, a failed check fails the step as
118
- it would outside a loop.
119
- - **A critique asks only for what the sources can show.** A claim that
120
- something is missing cannot cite a line, so a critique that demands one
121
- never passes. Cap `maxRounds` at 2 unless a round is cheap, and decide up
122
- front what you do when it runs out (continuing it is recorded as unmet).
123
- - **Gates stop only what is behind them.** Other branches keep running. When
124
- only waiting gates or loops are left, the run parks with status `waiting`.
125
- - **Independent checks.** `route.independentOf` names steps this step depends
126
- on (directly or through others) whose provider it must not use, so a check
127
- independent of `find` also depends on `find`. It needs a second provider:
128
- with only one enabled, validate and `workflow add` refuse it (`every enabled
129
- pool that could run it … uses that provider; enable a pool of another
130
- provider or drop independentOf`). The other provider needs a model on the
131
- step's tier: a no-pool refusal names each pool's reason.
132
- - **Shared folder.** All workers share one tree. Name each writer's exact
133
- files in `files` (steps whose files overlap run one after the other) and tell
134
- it to keep other workers' edits.
135
- - **Evidence: checks Bullswarm runs.** Add a command or schema check for
136
- anything a machine can check (`"evidence": [{"type": "command", "cmd": "npm
137
- test", "timeoutSec": 300}]`, at most 5 items). Each check has a timeout
138
- (default 120 seconds, at most 600); a suite that runs longer cannot be one
139
- item: split it, or have the step run it and answer with the result. A step
140
- whose command names none of its `files` (a bare `npm test`) gets advisory
141
- `suite-wider-than-files`: scope it, or use a check step. Run a
142
- check by hand before launch, because fixing a wrong check reruns the worker.
143
- Checks are read-only: a change to the deliverable fails the item. If the
144
- worker fails first, no check runs: the step's handback line and watch's
145
- failed line read `evidence not run`, and the JSON has `evidenceResults:
146
- null`. Each check's result is in `runs result <id> --json` under
147
- `actions[].evidenceResults` (`status`, `exit`, `tail`, `why`) and in
148
- `workflow action show <id> <step>`.
149
- - **Steps that must not repeat.** Sending or publishing is `"deliverable":
150
- "outward"` with `"retry": 0`; an outward step is never retried once its worker
151
- started.
152
- - **Prompts are self-contained.** Nothing is substituted: name the absolute
153
- workspace path, the outcome, the files, what to read from dependencies, and
154
- the checks to run. Every task carries a soft time box (`timeBox` minutes, a
155
- guide, never a timeout); a step that lists items under `## Not done` still
156
- succeeds and reads `returned early · N not done`. A failed step whose report
157
- lists `- outside: <blocker>` (something it may not change) skips its retry.
158
-
159
- Old v2 programs (`bullswarm.workflow.program.v2`) still run? No: since 0.38.0
160
- a new run refuses them, and runs they started are view-only.
161
-
162
- ### Validate, then launch
88
+ ```text
89
+ validate:
90
+ ✓ program v3 valid: 2 steps, 0 gates, 0 loops (nothing launched)
91
+ build build/medium deliverable=files evidence=command
92
+ review analyze/medium answer after build route: independent of build blind to build
93
+ ```
163
94
 
164
95
  Keep the goal in a file and pass it as `"$(cat goal.txt)"` to both commands,
165
96
  so validate and launch get identical text. `bullswarm workflow plan contract`
@@ -167,244 +98,104 @@ so validate and launch get identical text. `bullswarm workflow plan contract`
167
98
 
168
99
  ```bash
169
100
  bullswarm workflow plan validate "$(cat goal.txt)" --cwd=<abs-dir> --program=<abs-dir>/plan.json
101
+ bullswarm workflow goal "$(cat goal.txt)" --cwd=<abs-dir> --program=<abs-dir>/plan.json --json
102
+ bullswarm workflow watch <shortId> --until trouble
170
103
  ```
171
104
 
172
- ```text
173
- ✓ program v3 valid: 2 steps, 0 gates, 1 loop (nothing launched)
174
- fix build/medium deliverable=files
175
- check analyze/medium evidence=command answer after fix
176
- loop until-green steps fix, check · until check's evidence passed · at most 3 rounds
177
- launch bullswarm workflow goal 'Make the acme tests pass' --cwd /private/tmp/v37e/acme --program /private/tmp/v37e/loop.json --json
178
- ```
179
-
180
- Exit 2 lists the `issues`: fix them and validate again. Advisories never
181
- block; with `--json` they are in the JSON (`advisories`). Exit 0 prints the
182
- launch line; run it. A goal over 120 characters, or with a line break, shows
183
- as `"<goal>"` in that line (`launch bullswarm workflow goal "<goal>" --cwd
184
- …`): put `"$(cat goal.txt)"` in its place before you run it. The launch
185
- detaches and returns `shortId`; report it.
186
-
187
- ## 4. Watch: choose how close
188
-
189
- | Mode | Wakes you on | Command |
190
- |---|---|---|
191
- | Wake-ups only (the default choice) | a gate waiting, a loop out of rounds, a step that needs you (after its retry, or at once for a usage limit), a pause, a stale step, steering, the end; each loop that finished since the last wake is printed too, without waking | `bullswarm workflow watch <shortId> --until trouble` |
192
- | Every step | each finished step with its answer, loop rounds, plus every wake-up | `bullswarm workflow watch <shortId>` (or `--next` for one step at a time) |
193
- | Named steps | only the steps, gates or loops you name | `bullswarm workflow wait <shortId> <id...>` |
194
-
195
- Start one watch right after launch; it prints nothing more until a wake. Each exit is one wake:
196
- read the output, act, and start the printed `next:` line again. If your
197
- harness cannot wake you when a background process ends, run the watch in the foreground: it blocks until the
198
- wake; give it a `--timeout` under your tool's time limit (`--until trouble
199
- --timeout 100` for 2 minutes). A restart without `--after` attaches at the
200
- newest event and skips wakes in between. Never end your turn while a run you
201
- own is still running. Between wakes do not poll, read the run directory, or
202
- send per-step status replies.
203
-
204
- A gate `ship` after the loop above wakes you like this (real output); the
205
- loop's line comes with the wake:
105
+ Validate exit 2 lists the `issues`: fix them and validate again. Advisories
106
+ never block; with `--json` they are in the JSON (`advisories`). Exit 0 prints
107
+ the launch line; run it. A goal over 120 characters, or with a line break,
108
+ shows as `"<goal>"` in that line (`launch bullswarm workflow goal "<goal>"
109
+ --cwd …`): put `"$(cat goal.txt)"` in its place. The launch detaches and
110
+ returns `shortId`; report it.
206
111
 
207
- ```text
208
- ✓ loop until-green passed in round 1 of 3 · check's evidence passed
209
- ⧖ gate ship waiting · Read the fix and decide whether to write CHANGES.md · continue: bullswarm workflow continue m39i62 ship
210
- outcome: waiting
211
- waiting: gate ship · Read the fix and decide whether to write CHANGES.md
212
- next: bullswarm workflow continue m39i62 ship
213
- ```
214
-
215
- `workflow wait` returns each named step's facts and checked answer (exit 0
216
- when none failed, 1 when one failed or the run stopped short, 2 on a timeout),
217
- after a line for each loop the named ids wait behind:
218
-
219
- ```text
220
- ✓ loop until-green passed · round 1 of 3
221
- ⧖ gate ship waiting · Read the fix and decide whether to write CHANGES.md
222
- continue bullswarm workflow continue m39i62 ship
223
- ```
224
-
225
- ### Gates and loops that wait for you
226
-
227
- `bullswarm workflow continue <shortId> <gate>` passes a waiting gate, and the
228
- steps behind it start. A loop out of rounds waits the same way:
229
- `bullswarm workflow continue <shortId> <loop> --rounds <1-5>` gives it more
230
- rounds; without `--rounds` the steps behind it run, and it reads `→ loop <loop>
231
- continued by the caller after N of N rounds (condition not met)`, never
232
- passed. Before you continue you may add steps. The command relaunches the kernel when none is running:
233
-
234
- ```text
235
- ✓ gate ship passed in m39i62; kernel relaunched
236
- watch bullswarm workflow watch m39i62 --until trouble
237
- ```
238
-
239
- ### When a step needs you
240
-
241
- The one failure rule: a failed step gets one automatic retry (a process
242
- failure on another eligible pool, a failed check on the same pool with the
243
- failure attached), then it comes back to you in a needs-you block. A step with
244
- `retry: 0`, a started outward step, a usage limit, and a step no pool can run
245
- come back at once (`not retried`). Only a failed step's dependents wait; other branches finish. A real block:
246
-
247
- ```text
248
- ✗ lint needs you · command evidence failed after 1 retry
249
- evidence test -f LINT-OK.md → exit 1
250
- try 1 grok · grok-4.7 · 2m33s · 0 files
251
- try 2 same pool, failure attached · 1m33s · 0 files
252
- waiting on this: summary
253
- your call:
254
- retry here bullswarm workflow step rerun hkbbbi lint
255
- add steps bullswarm workflow add hkbbbi --steps part.json
256
- then wait bullswarm workflow wait hkbbbi <added ids>
257
- take over output: /private/tmp/v37e/home/workflows/wf-muli1jve-d48ab9/out-lint-attempt-2.md · diff: /private/tmp/v37e/home/workflows/wf-muli1jve-d48ab9/diff-lint-attempt-2.txt
258
- accept anyway bullswarm workflow step accept hkbbbi lint --reason "…"
259
- next: bullswarm workflow watch hkbbbi --until trouble --after <sequence> --since <iso>
260
- ```
261
-
262
- Choose one option, run it, then start the `next:` watch again:
263
-
264
- | Printed option | When to choose it | What it runs |
265
- |---|---|---|
266
- | `rerun elsewhere` | another eligible pool may succeed | `bullswarm workflow step rerun <shortId> <step> --avoid <pool>`; the pool stays in the step's route |
267
- | `retry here` | printed instead when no other pool could run the step | `bullswarm workflow step rerun <shortId> <step>` |
268
- | `add steps` / `then wait` | the step must be done differently: a v3 run's steps are never edited, so add a new one | `bullswarm workflow add <shortId> --steps part.json`, then `bullswarm workflow wait <shortId> <added ids>` |
269
- | `take over` | the rest is small or needs something only you have | read `output:` (and `diff:`) and do the work yourself |
270
- | `accept anyway` | you keep the failed result as it is | `bullswarm workflow step accept <shortId> <step> --reason "…"`: recorded as your choice, never proof; rerunning the step undoes it |
271
-
272
- ### A usage limit or no free pool
273
-
274
- A usage limit ends the step: a spent 5-hour, weekly or monthly window, or no
275
- credit left. The step comes straight back to you in a needs-you block (`✗
276
- <step> needs you · out of quota …`), even when the notice names no reset.
277
- Nothing waits, moves to another pool or retries by itself, and the rest of the
278
- run keeps going. Bullswarm never remembers a spent or dead pool from one step
279
- to the next: the pool's meter is read again at once, and a window it shows at
280
- 100% keeps the pool out of later steps until that window resets. The same
281
- happens when no pool that can run the step is free when it is picked. Its
282
- `why` names every pool and its reason (`no pool with quota to spare: <pool> at
283
- its 5-hour limit until <time>; …`, `<pool> at its weekly limit until <time>`,
284
- or `no pool free: …` when a reason is not a usage limit, and the header then
285
- reads `no eligible pool`). A retry the step was promised that finds no free
286
- pool keeps its own failure, and its `why` ends `· no retry: <pool> <reason>;
287
- …`.
288
-
289
- A short "too many requests" rate limit is not a usage limit. It backs off on
290
- the same pool at most twice (20 s, then 60 s, or the wait it names when that
291
- is at most 2 minutes), then comes back to you (`✗ <step> needs you · rate
292
- limited · backed off twice`). A try after a backoff reads `· after a
293
- rate-limit backoff`. One that names a longer wait comes back to you at once,
294
- with `back at` at the end of that wait. One whose pool is no longer free for
295
- the backoff (at its 5-hour, weekly or monthly limit, or nearly spent in the
296
- meantime) comes back to you at once too: as `out of quota` when that pool is
297
- out on a usage limit, with `back at` its return when that is known. A sign-in
298
- failure, a provider error or a worker that died at start still gets the step's
299
- one automatic retry by itself, on another free pool when there is one. After a
300
- sign-in failure that retry skips every pool that shares the credential;
301
- nothing is stored, so a later step can pick that pool again. A model the
302
- pool's plan lacks (`model not in plan`) retries on another pool, and that pool
303
- never gets that model again until `strategy include-model <model>`.
304
-
305
- When a return time is known, for this or any other failure, the block prints
306
- `back at <time>` and adds one option:
307
-
308
- | Printed option | When to choose it | What it runs |
309
- |---|---|---|
310
- | `rerun elsewhere` | another pool can run the step now | `bullswarm workflow step rerun <shortId> <step> --avoid <pool>` |
311
- | `wait for it` | the step should run on that pool, or nothing else can run it | `after <time>: bullswarm workflow step rerun <shortId> <step>`: run that rerun yourself after the `back at` time; before then the pool is still out |
312
- | `accept anyway` | you keep the step's result as it is | `bullswarm workflow step accept <shortId> <step> --reason "…"` |
313
-
314
- `add steps` and `take over` are printed too. To stop the whole run instead,
315
- run `bullswarm workflow cancel <shortId>`.
316
-
317
- ### A step that looks stale
318
-
319
- The watcher prints `⚠ <step> looks stale: <reasons>` once per attempt when a
320
- running step's score crosses its threshold (operations.md lists the reasons).
321
- While Bullswarm runs a step's declared
322
- checks, only quiet counts, read from the checks' heartbeat: `no check heartbeat
323
- for <N>m`. Nothing is stopped for you: let it run (start the `next:` watch
324
- again), or restart it with `bullswarm workflow step restart <shortId> <step>
325
- [--pool <pool>]`, which stops that step only and runs it again with a handoff
326
- of what the stopped attempt did.
112
+ Old v2 programs (`bullswarm.workflow.program.v2`) still run? No: since 0.38.0
113
+ a new run refuses them, and runs they started are view-only.
327
114
 
328
- ### When it finishes
115
+ Start one watch right after launch; it prints nothing until a wake (a gate
116
+ waiting, a loop out of rounds, a step that needs you, a pause, a stale step,
117
+ steering, the end). Each exit is one wake: read the output, act, and start the
118
+ printed `next:` line again. If your harness cannot wake you when a background
119
+ process ends, run the watch in the foreground: it blocks until the wake; give
120
+ it a `--timeout` under your tool's time limit (`--until trouble --timeout 100`
121
+ for 2 minutes). A restart without `--after` attaches at the newest event and
122
+ skips wakes in between. Never end your turn while a run you own is still
123
+ running. At a gate: `bullswarm workflow continue <shortId> <gate>`; a loop out
124
+ of rounds takes `--rounds <1-5>`. `watch --until trouble` also wakes on
125
+ `steering received` (a person left guidance: decide what it means and add
126
+ steps).
127
+
128
+ ## 4. Driving it well
129
+
130
+ Lessons from real runs; each holds for this version.
131
+
132
+ 1. **One worker beats chunks.** Split only when one worker cannot hold the
133
+ input (the triage above): priority and consistency are judgements across
134
+ items, and chunks cannot make them.
135
+ 2. **Write the judgement down.** A step whose decisions are spelled out (a
136
+ numbered spec, the rulings) runs well at `effort` `medium` or `low`; open
137
+ design needs `high`. Capability is model and reasoning together: the
138
+ effort tier picks both per pool (`bullswarm strategy rungs --json` shows
139
+ each rung), and a step's `reasoning` overrides only the level.
140
+ 3. **A contract first.** When parallel steps must fit together, a first step
141
+ that writes the shared shape (types, an event format), which the others
142
+ depend on, keeps the join short; without it the integrator becomes the
143
+ author.
144
+ 4. **Files decide parallelism.** Steps whose `files` overlap run one after
145
+ the other. Give each writer its own files and tell it to keep other
146
+ workers' edits; keep a breaking rename in one step, not parallel with its
147
+ consumers, which would build against the old name.
148
+ 5. **Checks are facts.** Put anything a machine can say in `evidence`, and
149
+ run each check by hand before launch: fixing a wrong check reruns the
150
+ worker. Workers share one tree, so a later step can undo what an earlier
151
+ step's check proved; put the final check where nothing runs after it
152
+ (after integration, or on the step a gate waits behind).
153
+ 6. **Reviews that hold.** State the contract as numbered checks, make any
154
+ difference a finding even when the author justifies it, have it record
155
+ `holds` and evidence per check with `passed` true only when every check
156
+ holds, route it `independentOf` the author, and add `blindTo` the author:
157
+ a builder's credible reason talks a reviewer who reads it out of a real
158
+ finding. Keep the brief consistent with the repository's own rules: a
159
+ brief that contradicts a repo test cannot pass honestly.
160
+ 7. **Plan ahead, add later.** Declare the steps, gates and loops you can see.
161
+ Where the next part depends on an answer (one check per finding, one build
162
+ step per slice), stop there and add it with `workflow add` when the answer
163
+ arrives. Put a gate before expensive or outward work, and after a design
164
+ step the owner must approve.
165
+ 8. **Watch by wake-ups.** One `watch --until trouble` per run. Between wakes
166
+ do not poll, read the run directory, or send per-step status replies. At a
167
+ needs-you block, pick one printed option and run it as printed.
168
+ 9. **The worker's report is not proof.** Read the output and the diff.
169
+ `ok: true` and `proven by command` mean the checks passed, not that the
170
+ work is right; `answer checked` is a well-formed claim.
171
+ 10. **Retries are for flakes.** The one automatic retry covers a crash or a
172
+ failed check. A worker whose check or deliverable failed and whose report
173
+ lists `- outside: <blocker>` under `## Not done` comes back at once. Fix
174
+ the cause with `workflow add` rather than rerunning the same step.
175
+
176
+ ## 5. Read the outcome
329
177
 
330
178
  `outcome:` is `completed`, `partial` or `cancelled`, and `reason:` says why in
331
179
  one line. A completed v3 run hands nothing back: every step succeeded. The
332
180
  proof line says what backs each step: `proven by command` or `proven by
333
181
  schema` (a check Bullswarm ran passed), `answer checked` (its answer passed
334
182
  its schema: a well-formed claim, not proof, so it is not counted as proven),
335
- `finished · unproven` (neither), and `accepted by choice`. When you report the outcome, quote the run's proof line
183
+ `finished · unproven` (neither), and `accepted by choice` (your decision,
184
+ never verification). When you report the outcome, quote the run's proof line
336
185
  as printed (`proof: …` at the end of watch, `# proof` in `runs result`)
337
186
  instead of paraphrasing it. Then read the real outputs and answers
338
187
  (`bullswarm workflow runs result <shortId> --json` names every step's output)
339
- and probe the important edge cases yourself.
340
-
341
- A partial run lists each unfinished step (`step <id>: <status> (<kind>) —
342
- <why>`, with `its pool is back at <time>` when that is known) and your options
343
- under `your call:` (real output):
344
-
345
- ```text
346
- step lint: failed (failed-evidence) after 1 retry — test -f LINT-OK.md → exit 1
347
- step summary: blocked (dependency) — failed dependency
348
- your call:
349
- add bullswarm workflow add hkbbbi --steps part.json, then bullswarm workflow wait hkbbbi <added ids> (appends steps, gates or loops and reopens the run)
350
- rerun bullswarm workflow step rerun hkbbbi lint [--avoid <pool>] (runs it again with its last attempt's handoff)
351
- accept bullswarm workflow step accept hkbbbi lint --reason "…" (recorded as your choice, never proof)
352
- take over do the unfinished work yourself; bullswarm workflow runs result hkbbbi --json names every step's output
353
- restart start a new run: bullswarm workflow goal "<goal>" --cwd /private/tmp/v37e/acme --program <file.json>
354
- ```
355
-
356
- | Printed option | When | What to do |
357
- |---|---|---|
358
- | `add` | new work is needed, or a step must be done differently | `bullswarm workflow add <shortId> --steps part.json`, then `bullswarm workflow wait <shortId> <added ids>`; the run reopens |
359
- | `retry` | a step stopped for a reason a retry fixes: a crashed or silent worker, a usage limit, or no pool free | `bullswarm workflow resume <shortId>` (run it after its `back at` time) |
360
- | `rerun` | a step failed | `bullswarm workflow step rerun <shortId> <step> [--avoid <pool>]` |
361
- | `accept` | you keep a failed step as it is | `bullswarm workflow step accept <shortId> <step> --reason "…"` |
362
- | `take over` | the rest is small, or needs something only you have | do it yourself |
363
- | `restart` | the goal or the approach was wrong | `bullswarm workflow goal "<goal>" --cwd <dir> --program <file.json>` |
364
-
365
- `retry` appears only when a step is retryable. When nothing is retryable,
366
- `resume` prints `nothing to retry`, starts nothing, and exits 1.
367
-
368
- **An accept is a choice, never proof.** `step accept` lets the step's
369
- dependents run, and the step reads `accepted by choice` and is counted apart
370
- in the proof line (`N accepted by choice: <steps>`). Report it as your
371
- decision, never as verification.
372
-
373
- ## 5. Extend or steer a workflow
374
-
375
- A v3 run's steps, gates and loops are never edited. You change what happens
376
- next by adding to it:
377
-
378
- - `bullswarm workflow add <shortId> --steps part.json` appends a fragment
379
- `{steps, gates?, loops?}`. New steps may depend on existing steps, gates and
380
- loops, finished or not (on a loop's steps only through the loop's id).
381
- Nothing the run has changes, and a finished run reopens, except that
382
- `blocks: {"fix": ["report"]}` makes existing not-started steps also wait for
383
- a step it adds (then `step accept` the failed step). Added steps do not
384
- take the program's `defaults`: set `lane` and `effort` on each.
385
- `--from-answer <step>` adds the fragment a step answered (read it with
386
- `wait` first). Real output:
387
-
388
- ```text
389
- ✓ added to jcefns · revision 2 (applied directly; kernel relaunched)
390
- reopened the completed run; its earlier result is archived
391
- added step check-first
392
- wait bullswarm workflow wait jcefns check-first
393
- ```
394
-
395
- - `bullswarm workflow step rerun <shortId> <step>` runs a step again with its
396
- last attempt's handoff; `step accept` keeps a failed one; `step restart`
397
- stops a running one and runs it again.
398
- - `bullswarm workflow pause <shortId>` starts nothing new (`--now` also stops
399
- running steps); `bullswarm workflow resume <shortId>` lifts it.
400
- `bullswarm workflow cancel <shortId>` stops the run.
401
- - `watch --until trouble` also wakes on `steering received` (a person left
402
- guidance for you: decide what it means and add steps) and on a pause.
403
-
404
- A stopped step's file edits stay in the shared tree; when they must not
405
- remain, add a step that reverts them. Exit 0 can mean launched, waiting,
406
- paused or completed, so always read the returned status.
407
-
408
- [operations.md](references/operations.md) covers the details: add and wait,
409
- continue, reruns and accepts, pause and resume, watch flags, the handback
410
- fields, routing diagnosis, and reading view-only saved runs.
188
+ and probe the important edge cases yourself. A `partial` run lists what is
189
+ unfinished and your options: see recovery.md.
190
+
191
+ ## 6. Read this when
192
+
193
+ | Situation | Read |
194
+ |---|---|
195
+ | a step needs you (a needs-you block), a rate limit, a usage limit or no free pool | [recovery.md](references/recovery.md) |
196
+ | a step looks stale; pause, restart or cancel; a run ended `partial` | [recovery.md](references/recovery.md) |
197
+ | writing a program: every field, conditions, answers, evidence, deliverables | [program.md](references/program.md) |
198
+ | a review: the strict form and `blindTo` | [program.md](references/program.md) "Reviews", [patterns.md](references/patterns.md) 3 |
199
+ | copying a workflow: find then check, fix until green, draft to publish, parallel slices, triage | [patterns.md](references/patterns.md) |
200
+ | `workflow add` and `blocks`, `wait`, `continue`, watch modes and flags, handback fields, routing and reasoning diagnosis, view-only saved runs | [operations.md](references/operations.md) |
201
+ | adding a provider | [providers.md](references/providers.md) |
@@ -1,7 +1,7 @@
1
1
  interface:
2
2
  display_name: "Bullswarm"
3
- short_description: "Choose one agent or a verified workflow"
4
- default_prompt: "Use $bullswarm to classify, preview, and execute this task with the smallest reliable agent setup."
3
+ short_description: "Delegate one task or a workflow you write"
4
+ default_prompt: "Use $bullswarm to choose one run or a workflow for this task, execute it, and check the result by facts."
5
5
 
6
6
  policy:
7
7
  allow_implicit_invocation: true