bullswarm 0.38.2 → 0.38.3
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/AGENTS.md +3 -1
- package/CHANGELOG.md +34 -0
- package/docs/reference/program.md +1 -0
- package/package.json +1 -1
- package/skill/SKILL.md +111 -320
- package/skill/agents/openai.yaml +2 -2
- package/skill/references/operations.md +55 -183
- package/skill/references/patterns.md +18 -4
- package/skill/references/program.md +96 -2
- package/skill/references/recovery.md +253 -0
- package/src/help.js +2 -2
- package/src/workflow/action-validator.js +10 -2
- package/src/workflow/cli-plan.js +1 -1
- package/src/workflow/contract-v3.js +11 -1
- package/src/workflow/gates-loops.js +5 -2
- package/src/workflow/program-v3.js +5 -1
- package/src/workflow/revision-v3.js +2 -2
- package/src/workflow/step-prompts.js +10 -0
- package/src/workflow/step-route.js +29 -0
package/AGENTS.md
CHANGED
|
@@ -123,7 +123,9 @@ bullswarm workflow runs delete <shortId> --yes
|
|
|
123
123
|
## Using bullswarm from another agent
|
|
124
124
|
|
|
125
125
|
If you are an agent that wants to offload bounded work via bullswarm,
|
|
126
|
-
read `skill/SKILL.md` — that's the agent-facing user guide
|
|
126
|
+
read `skill/SKILL.md` — that's the agent-facing user guide: a short entry
|
|
127
|
+
point that links its references (`recovery.md`, `program.md`, `patterns.md`,
|
|
128
|
+
`operations.md`, `providers.md`) for the rest. There are
|
|
127
129
|
exactly two ways to start work, and the caller chooses the shape itself: one
|
|
128
130
|
bounded outcome goes to `bullswarm run`; parallel territories, integration,
|
|
129
131
|
or independent acceptance go to `bullswarm workflow goal` with a program you
|
package/CHANGELOG.md
CHANGED
|
@@ -2,6 +2,40 @@
|
|
|
2
2
|
|
|
3
3
|
## Unreleased
|
|
4
4
|
|
|
5
|
+
## 0.38.3 — Blind reviews and a reorganized skill
|
|
6
|
+
|
|
7
|
+
- steps: a v3 step may declare `"blindTo": ["<step id>", ...]`, naming steps
|
|
8
|
+
that run before it (directly or through others). Its task still lists each
|
|
9
|
+
named dependency, but without that step's output file or checked answer,
|
|
10
|
+
and a loop's `Previous round` block leaves out that step's answer and output
|
|
11
|
+
(it still says the step ran, its status and its evidence); the step waits
|
|
12
|
+
for it as before. Use it on a review of a build step: in a real run a
|
|
13
|
+
builder's answer explained a deviation away with a credible reason and the
|
|
14
|
+
reviewer passed with no findings, while the same reviewer without the
|
|
15
|
+
answer caught it. Validate and `workflow add` refuse a non-array, an unknown
|
|
16
|
+
id, a step that does not run before it, or the step itself, like
|
|
17
|
+
`route.independentOf`, which is unchanged: it still only picks the provider.
|
|
18
|
+
`plan validate` prints `blind to <ids>` on the step's line.
|
|
19
|
+
- docs: the review and critique examples in `patterns.md` and `program.md`
|
|
20
|
+
are strict: the contract as numbered checks, any difference a finding even
|
|
21
|
+
when it looks harmless, intended or justified (the caller decides), and an
|
|
22
|
+
answer with `checks` ({id, holds, evidence}) and `passed`, true only when
|
|
23
|
+
every check holds. Find-then-check stays without `blindTo`, since a check
|
|
24
|
+
works from the list the finder answered.
|
|
25
|
+
- skill: reorganized for progressive disclosure. `skill/SKILL.md` is now the
|
|
26
|
+
short entry point a caller reads every time (25,122 bytes before, 12,198
|
|
27
|
+
after): choosing `bullswarm run` or a workflow, one command block for each,
|
|
28
|
+
a "Driving it well" playbook of ten rules from real runs, and a "Read this
|
|
29
|
+
when" index. The rare paths moved to a new `skill/references/recovery.md`
|
|
30
|
+
(15,988 bytes): the failure rule, needs-you blocks and their options, usage
|
|
31
|
+
limits and no free pool, rate-limit backoff, stale steps and `step restart`,
|
|
32
|
+
pause, resume and cancel, and a partial end, with the same real output.
|
|
33
|
+
`operations.md` drops what now lives there (54,517 bytes before, 44,792
|
|
34
|
+
after) and gains the watch modes and the real `add`, `wait` and `continue`
|
|
35
|
+
output; `program.md` gains a "Writing a program" section (18,536 bytes
|
|
36
|
+
before, 21,577 after). Every command, option and output line the old
|
|
37
|
+
entry point named is still in the skill.
|
|
38
|
+
|
|
5
39
|
## 0.38.2 — Three large files split into one-concept modules
|
|
6
40
|
|
|
7
41
|
- internal: the three largest workflow files are split into one-concept
|
|
@@ -44,6 +44,7 @@ Steps, gates and loops share one id space: an id is kebab-case and used once.
|
|
|
44
44
|
| `effort` | no | `high`, `medium`, `low`; default by lane: analyze medium, build medium, chore low; a chore step must be low |
|
|
45
45
|
| `reasoning` | no | `low`, `medium`, `high`, `xhigh`, `max`, or `default` (pass nothing): how hard the picked model thinks |
|
|
46
46
|
| `route` | no | `{pools: {use, avoid}, providers: {use, avoid}, independentOf: [step ids]}`: a hard filter applied before quota pacing; `independentOf` names steps this step depends on (directly or through others) whose providers it must not use |
|
|
47
|
+
| `blindTo` | no | step ids this step depends on (directly or through others) whose output file and checked answer it is not handed, in its task or a loop's `Previous round` block; the dependency still orders it. Use it on a review of a build step; leave it off a check that must read the list it checks |
|
|
47
48
|
| `answer` | no | a JSON schema: the worker writes its final answer as JSON to a file Bullswarm names (at most 256 KiB), and that file is checked; a mismatch is failure kind `schema`; the checked answer goes to dependent steps, conditions, `workflow wait`, `watch` and `runs result` |
|
|
48
49
|
| `evidence` | no | up to 5 checks Bullswarm runs after the worker, as in [Evidence](#evidence-command-and-schema) |
|
|
49
50
|
| `deliverable` | no | `files`, `report`, `data`, `media`, `outward`, or `{type, paths}` (default `files` for build and chore, `report` for analyze without an answer, none for analyze with an answer); not produced is failure kind `not-produced` |
|
package/package.json
CHANGED
package/skill/SKILL.md
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
---
|
|
2
2
|
name: bullswarm
|
|
3
|
-
description:
|
|
3
|
+
description: Load before you hand work to another coding agent through the Bullswarm CLI, as one quota-routed task (bullswarm run) or a workflow you write from steps, phases, gates and loops (bullswarm workflow goal --program), and whenever you watch, answer, extend or recover such a run. Use for /bullswarm, offloading, independent checks or reviews, and requested multi-agent execution; not for work you are asked to do yourself.
|
|
4
4
|
---
|
|
5
5
|
|
|
6
6
|
# Bullswarm
|
|
@@ -52,13 +52,12 @@ changes. Use `--task-file` for long text. Options you will use:
|
|
|
52
52
|
Read the verdict. `ok: true` means read `outFile` (and `answer`) and check the
|
|
53
53
|
content before you use it. `ok: false` means inspect and report the failure;
|
|
54
54
|
`failureKind` names it and `why` says it in one line. `pool` and `model` name
|
|
55
|
-
who ran it; `shortId` names the run;
|
|
56
|
-
|
|
57
|
-
|
|
58
|
-
|
|
59
|
-
a readiness problem.
|
|
55
|
+
who ran it; `shortId` names the run; the last field, `details`, is the command
|
|
56
|
+
for its record and cost. A build or chore run that changes no file (files git
|
|
57
|
+
ignores count) fails `not-produced`, also outside git. Do not run `doctor`
|
|
58
|
+
unless dispatch reports a readiness problem.
|
|
60
59
|
|
|
61
|
-
## 3. A workflow:
|
|
60
|
+
## 3. A workflow: write, validate, launch, watch
|
|
62
61
|
|
|
63
62
|
| Block | What it is | Declared as |
|
|
64
63
|
|---|---|---|
|
|
@@ -67,99 +66,31 @@ a readiness problem.
|
|
|
67
66
|
| Gate | the run stops there and waits for you, or only when an answer says so | an entry in `gates`: `{id, dependsOn, when?, note?}` |
|
|
68
67
|
| Loop | steps that repeat until one step's answer (or evidence) says stop, at most `maxRounds` (1-5) | an entry in `loops`: `{id, steps, until, maxRounds}` |
|
|
69
68
|
|
|
70
|
-
A
|
|
71
|
-
|
|
72
|
-
steps, 1 gate, 1 loop):
|
|
69
|
+
A build, then a strict review by another provider that does not read the
|
|
70
|
+
builder's own account (playbook rule 6):
|
|
73
71
|
|
|
74
72
|
```json
|
|
75
73
|
{
|
|
76
74
|
"schemaVersion": "bullswarm.workflow.program.v3",
|
|
77
75
|
"steps": [
|
|
78
|
-
{ "id": "
|
|
79
|
-
|
|
80
|
-
|
|
81
|
-
{ "id": "
|
|
82
|
-
|
|
83
|
-
|
|
84
|
-
|
|
85
|
-
|
|
76
|
+
{ "id": "build", "lane": "build", "files": ["src/csv-writer.js", "tests/csv-writer.test.js"],
|
|
77
|
+
"prompt": "In /work/acme, add src/csv-writer.js: writeCsv(rows) returns RFC 4180 CSV text. Add its tests in tests/csv-writer.test.js.",
|
|
78
|
+
"evidence": [{ "type": "command", "cmd": "node --test tests/csv-writer.test.js", "timeoutSec": 300 }] },
|
|
79
|
+
{ "id": "review", "dependsOn": ["build"], "route": { "independentOf": ["build"] }, "blindTo": ["build"],
|
|
80
|
+
"prompt": "In /work/acme, review src/csv-writer.js against: 1. a field holding a comma, a quote or a line break is quoted, and a quote inside it is doubled; 2. every line ends with CRLF. Any difference is a finding, even if justified by the author. Record each check as holds true or false with its evidence. Answer passed true only when every check holds. Change no file.",
|
|
81
|
+
"answer": { "type": "object", "required": ["checks", "passed"], "properties": {
|
|
82
|
+
"checks": { "type": "array", "items": { "type": "object", "required": ["id", "holds", "evidence"], "properties": { "id": { "type": "string" }, "holds": { "type": "boolean" }, "evidence": { "type": "string" } } } },
|
|
83
|
+
"passed": { "type": "boolean" } } } }
|
|
84
|
+
]
|
|
86
85
|
}
|
|
87
86
|
```
|
|
88
87
|
|
|
89
|
-
|
|
90
|
-
|
|
91
|
-
|
|
92
|
-
|
|
93
|
-
|
|
94
|
-
|
|
95
|
-
|
|
96
|
-
- **Plan as far ahead as you know.** Declare the steps, gates and loops you can
|
|
97
|
-
see now. Where the next part depends on an answer (one check per finding, one
|
|
98
|
-
build step per slice), stop there and add it later with `workflow add`.
|
|
99
|
-
- **A step passes by facts.** Its worker ended cleanly, its deliverable was
|
|
100
|
-
produced, its evidence passed, and its answer (when declared) matched the
|
|
101
|
-
schema. Declare an `answer` whenever you, a later step or a condition needs
|
|
102
|
-
data from the step; a dependent step is handed its dependencies' checked
|
|
103
|
-
answer files.
|
|
104
|
-
- **The one condition form.** A gate's `when` and a loop's `until` read one
|
|
105
|
-
value: `{"step": "critique", "field": "passed"}` (a boolean the step's answer
|
|
106
|
-
schema requires; add `"equals": false` to invert it) or `{"step": "check",
|
|
107
|
-
"evidence": "passed"}`. No expressions and no else: anything more is your
|
|
108
|
-
call, with `workflow wait` and `workflow add`.
|
|
109
|
-
- **Loops run every step in every round**, and read their condition when the
|
|
110
|
-
round is over. Put the deciding step last, and give every writer in the loop
|
|
111
|
-
work each round: a build or chore step that changes no file fails
|
|
112
|
-
`not-produced` (rewrite the draft from the critique; do not put a revise
|
|
113
|
-
step after a critique that may pass the first time). From round 2 on, each
|
|
114
|
-
step's task carries a `Previous round` block with the last round's answers
|
|
115
|
-
and evidence. When a loop's `until` is the evidence form (`{"step": "check",
|
|
116
|
-
"evidence": "passed"}`), a failed check on that step reads as "not passed"
|
|
117
|
-
and the loop goes on. With the field form, a failed check fails the step as
|
|
118
|
-
it would outside a loop.
|
|
119
|
-
- **A critique asks only for what the sources can show.** A claim that
|
|
120
|
-
something is missing cannot cite a line, so a critique that demands one
|
|
121
|
-
never passes. Cap `maxRounds` at 2 unless a round is cheap, and decide up
|
|
122
|
-
front what you do when it runs out (continuing it is recorded as unmet).
|
|
123
|
-
- **Gates stop only what is behind them.** Other branches keep running. When
|
|
124
|
-
only waiting gates or loops are left, the run parks with status `waiting`.
|
|
125
|
-
- **Independent checks.** `route.independentOf` names steps this step depends
|
|
126
|
-
on (directly or through others) whose provider it must not use, so a check
|
|
127
|
-
independent of `find` also depends on `find`. It needs a second provider:
|
|
128
|
-
with only one enabled, validate and `workflow add` refuse it (`every enabled
|
|
129
|
-
pool that could run it … uses that provider; enable a pool of another
|
|
130
|
-
provider or drop independentOf`). The other provider needs a model on the
|
|
131
|
-
step's tier: a no-pool refusal names each pool's reason.
|
|
132
|
-
- **Shared folder.** All workers share one tree. Name each writer's exact
|
|
133
|
-
files in `files` (steps whose files overlap run one after the other) and tell
|
|
134
|
-
it to keep other workers' edits.
|
|
135
|
-
- **Evidence: checks Bullswarm runs.** Add a command or schema check for
|
|
136
|
-
anything a machine can check (`"evidence": [{"type": "command", "cmd": "npm
|
|
137
|
-
test", "timeoutSec": 300}]`, at most 5 items). Each check has a timeout
|
|
138
|
-
(default 120 seconds, at most 600); a suite that runs longer cannot be one
|
|
139
|
-
item: split it, or have the step run it and answer with the result. A step
|
|
140
|
-
whose command names none of its `files` (a bare `npm test`) gets advisory
|
|
141
|
-
`suite-wider-than-files`: scope it, or use a check step. Run a
|
|
142
|
-
check by hand before launch, because fixing a wrong check reruns the worker.
|
|
143
|
-
Checks are read-only: a change to the deliverable fails the item. If the
|
|
144
|
-
worker fails first, no check runs: the step's handback line and watch's
|
|
145
|
-
failed line read `evidence not run`, and the JSON has `evidenceResults:
|
|
146
|
-
null`. Each check's result is in `runs result <id> --json` under
|
|
147
|
-
`actions[].evidenceResults` (`status`, `exit`, `tail`, `why`) and in
|
|
148
|
-
`workflow action show <id> <step>`.
|
|
149
|
-
- **Steps that must not repeat.** Sending or publishing is `"deliverable":
|
|
150
|
-
"outward"` with `"retry": 0`; an outward step is never retried once its worker
|
|
151
|
-
started.
|
|
152
|
-
- **Prompts are self-contained.** Nothing is substituted: name the absolute
|
|
153
|
-
workspace path, the outcome, the files, what to read from dependencies, and
|
|
154
|
-
the checks to run. Every task carries a soft time box (`timeBox` minutes, a
|
|
155
|
-
guide, never a timeout); a step that lists items under `## Not done` still
|
|
156
|
-
succeeds and reads `returned early · N not done`. A failed step whose report
|
|
157
|
-
lists `- outside: <blocker>` (something it may not change) skips its retry.
|
|
158
|
-
|
|
159
|
-
Old v2 programs (`bullswarm.workflow.program.v2`) still run? No: since 0.38.0
|
|
160
|
-
a new run refuses them, and runs they started are view-only.
|
|
161
|
-
|
|
162
|
-
### Validate, then launch
|
|
88
|
+
```text
|
|
89
|
+
validate:
|
|
90
|
+
✓ program v3 valid: 2 steps, 0 gates, 0 loops (nothing launched)
|
|
91
|
+
build build/medium deliverable=files evidence=command
|
|
92
|
+
review analyze/medium answer after build route: independent of build blind to build
|
|
93
|
+
```
|
|
163
94
|
|
|
164
95
|
Keep the goal in a file and pass it as `"$(cat goal.txt)"` to both commands,
|
|
165
96
|
so validate and launch get identical text. `bullswarm workflow plan contract`
|
|
@@ -167,244 +98,104 @@ so validate and launch get identical text. `bullswarm workflow plan contract`
|
|
|
167
98
|
|
|
168
99
|
```bash
|
|
169
100
|
bullswarm workflow plan validate "$(cat goal.txt)" --cwd=<abs-dir> --program=<abs-dir>/plan.json
|
|
101
|
+
bullswarm workflow goal "$(cat goal.txt)" --cwd=<abs-dir> --program=<abs-dir>/plan.json --json
|
|
102
|
+
bullswarm workflow watch <shortId> --until trouble
|
|
170
103
|
```
|
|
171
104
|
|
|
172
|
-
|
|
173
|
-
|
|
174
|
-
|
|
175
|
-
|
|
176
|
-
|
|
177
|
-
|
|
178
|
-
```
|
|
179
|
-
|
|
180
|
-
Exit 2 lists the `issues`: fix them and validate again. Advisories never
|
|
181
|
-
block; with `--json` they are in the JSON (`advisories`). Exit 0 prints the
|
|
182
|
-
launch line; run it. A goal over 120 characters, or with a line break, shows
|
|
183
|
-
as `"<goal>"` in that line (`launch bullswarm workflow goal "<goal>" --cwd
|
|
184
|
-
…`): put `"$(cat goal.txt)"` in its place before you run it. The launch
|
|
185
|
-
detaches and returns `shortId`; report it.
|
|
186
|
-
|
|
187
|
-
## 4. Watch: choose how close
|
|
188
|
-
|
|
189
|
-
| Mode | Wakes you on | Command |
|
|
190
|
-
|---|---|---|
|
|
191
|
-
| Wake-ups only (the default choice) | a gate waiting, a loop out of rounds, a step that needs you (after its retry, or at once for a usage limit), a pause, a stale step, steering, the end; each loop that finished since the last wake is printed too, without waking | `bullswarm workflow watch <shortId> --until trouble` |
|
|
192
|
-
| Every step | each finished step with its answer, loop rounds, plus every wake-up | `bullswarm workflow watch <shortId>` (or `--next` for one step at a time) |
|
|
193
|
-
| Named steps | only the steps, gates or loops you name | `bullswarm workflow wait <shortId> <id...>` |
|
|
194
|
-
|
|
195
|
-
Start one watch right after launch; it prints nothing more until a wake. Each exit is one wake:
|
|
196
|
-
read the output, act, and start the printed `next:` line again. If your
|
|
197
|
-
harness cannot wake you when a background process ends, run the watch in the foreground: it blocks until the
|
|
198
|
-
wake; give it a `--timeout` under your tool's time limit (`--until trouble
|
|
199
|
-
--timeout 100` for 2 minutes). A restart without `--after` attaches at the
|
|
200
|
-
newest event and skips wakes in between. Never end your turn while a run you
|
|
201
|
-
own is still running. Between wakes do not poll, read the run directory, or
|
|
202
|
-
send per-step status replies.
|
|
203
|
-
|
|
204
|
-
A gate `ship` after the loop above wakes you like this (real output); the
|
|
205
|
-
loop's line comes with the wake:
|
|
105
|
+
Validate exit 2 lists the `issues`: fix them and validate again. Advisories
|
|
106
|
+
never block; with `--json` they are in the JSON (`advisories`). Exit 0 prints
|
|
107
|
+
the launch line; run it. A goal over 120 characters, or with a line break,
|
|
108
|
+
shows as `"<goal>"` in that line (`launch bullswarm workflow goal "<goal>"
|
|
109
|
+
--cwd …`): put `"$(cat goal.txt)"` in its place. The launch detaches and
|
|
110
|
+
returns `shortId`; report it.
|
|
206
111
|
|
|
207
|
-
|
|
208
|
-
|
|
209
|
-
⧖ gate ship waiting · Read the fix and decide whether to write CHANGES.md · continue: bullswarm workflow continue m39i62 ship
|
|
210
|
-
outcome: waiting
|
|
211
|
-
waiting: gate ship · Read the fix and decide whether to write CHANGES.md
|
|
212
|
-
next: bullswarm workflow continue m39i62 ship
|
|
213
|
-
```
|
|
214
|
-
|
|
215
|
-
`workflow wait` returns each named step's facts and checked answer (exit 0
|
|
216
|
-
when none failed, 1 when one failed or the run stopped short, 2 on a timeout),
|
|
217
|
-
after a line for each loop the named ids wait behind:
|
|
218
|
-
|
|
219
|
-
```text
|
|
220
|
-
✓ loop until-green passed · round 1 of 3
|
|
221
|
-
⧖ gate ship waiting · Read the fix and decide whether to write CHANGES.md
|
|
222
|
-
continue bullswarm workflow continue m39i62 ship
|
|
223
|
-
```
|
|
224
|
-
|
|
225
|
-
### Gates and loops that wait for you
|
|
226
|
-
|
|
227
|
-
`bullswarm workflow continue <shortId> <gate>` passes a waiting gate, and the
|
|
228
|
-
steps behind it start. A loop out of rounds waits the same way:
|
|
229
|
-
`bullswarm workflow continue <shortId> <loop> --rounds <1-5>` gives it more
|
|
230
|
-
rounds; without `--rounds` the steps behind it run, and it reads `→ loop <loop>
|
|
231
|
-
continued by the caller after N of N rounds (condition not met)`, never
|
|
232
|
-
passed. Before you continue you may add steps. The command relaunches the kernel when none is running:
|
|
233
|
-
|
|
234
|
-
```text
|
|
235
|
-
✓ gate ship passed in m39i62; kernel relaunched
|
|
236
|
-
watch bullswarm workflow watch m39i62 --until trouble
|
|
237
|
-
```
|
|
238
|
-
|
|
239
|
-
### When a step needs you
|
|
240
|
-
|
|
241
|
-
The one failure rule: a failed step gets one automatic retry (a process
|
|
242
|
-
failure on another eligible pool, a failed check on the same pool with the
|
|
243
|
-
failure attached), then it comes back to you in a needs-you block. A step with
|
|
244
|
-
`retry: 0`, a started outward step, a usage limit, and a step no pool can run
|
|
245
|
-
come back at once (`not retried`). Only a failed step's dependents wait; other branches finish. A real block:
|
|
246
|
-
|
|
247
|
-
```text
|
|
248
|
-
✗ lint needs you · command evidence failed after 1 retry
|
|
249
|
-
evidence test -f LINT-OK.md → exit 1
|
|
250
|
-
try 1 grok · grok-4.7 · 2m33s · 0 files
|
|
251
|
-
try 2 same pool, failure attached · 1m33s · 0 files
|
|
252
|
-
waiting on this: summary
|
|
253
|
-
your call:
|
|
254
|
-
retry here bullswarm workflow step rerun hkbbbi lint
|
|
255
|
-
add steps bullswarm workflow add hkbbbi --steps part.json
|
|
256
|
-
then wait bullswarm workflow wait hkbbbi <added ids>
|
|
257
|
-
take over output: /private/tmp/v37e/home/workflows/wf-muli1jve-d48ab9/out-lint-attempt-2.md · diff: /private/tmp/v37e/home/workflows/wf-muli1jve-d48ab9/diff-lint-attempt-2.txt
|
|
258
|
-
accept anyway bullswarm workflow step accept hkbbbi lint --reason "…"
|
|
259
|
-
next: bullswarm workflow watch hkbbbi --until trouble --after <sequence> --since <iso>
|
|
260
|
-
```
|
|
261
|
-
|
|
262
|
-
Choose one option, run it, then start the `next:` watch again:
|
|
263
|
-
|
|
264
|
-
| Printed option | When to choose it | What it runs |
|
|
265
|
-
|---|---|---|
|
|
266
|
-
| `rerun elsewhere` | another eligible pool may succeed | `bullswarm workflow step rerun <shortId> <step> --avoid <pool>`; the pool stays in the step's route |
|
|
267
|
-
| `retry here` | printed instead when no other pool could run the step | `bullswarm workflow step rerun <shortId> <step>` |
|
|
268
|
-
| `add steps` / `then wait` | the step must be done differently: a v3 run's steps are never edited, so add a new one | `bullswarm workflow add <shortId> --steps part.json`, then `bullswarm workflow wait <shortId> <added ids>` |
|
|
269
|
-
| `take over` | the rest is small or needs something only you have | read `output:` (and `diff:`) and do the work yourself |
|
|
270
|
-
| `accept anyway` | you keep the failed result as it is | `bullswarm workflow step accept <shortId> <step> --reason "…"`: recorded as your choice, never proof; rerunning the step undoes it |
|
|
271
|
-
|
|
272
|
-
### A usage limit or no free pool
|
|
273
|
-
|
|
274
|
-
A usage limit ends the step: a spent 5-hour, weekly or monthly window, or no
|
|
275
|
-
credit left. The step comes straight back to you in a needs-you block (`✗
|
|
276
|
-
<step> needs you · out of quota …`), even when the notice names no reset.
|
|
277
|
-
Nothing waits, moves to another pool or retries by itself, and the rest of the
|
|
278
|
-
run keeps going. Bullswarm never remembers a spent or dead pool from one step
|
|
279
|
-
to the next: the pool's meter is read again at once, and a window it shows at
|
|
280
|
-
100% keeps the pool out of later steps until that window resets. The same
|
|
281
|
-
happens when no pool that can run the step is free when it is picked. Its
|
|
282
|
-
`why` names every pool and its reason (`no pool with quota to spare: <pool> at
|
|
283
|
-
its 5-hour limit until <time>; …`, `<pool> at its weekly limit until <time>`,
|
|
284
|
-
or `no pool free: …` when a reason is not a usage limit, and the header then
|
|
285
|
-
reads `no eligible pool`). A retry the step was promised that finds no free
|
|
286
|
-
pool keeps its own failure, and its `why` ends `· no retry: <pool> <reason>;
|
|
287
|
-
…`.
|
|
288
|
-
|
|
289
|
-
A short "too many requests" rate limit is not a usage limit. It backs off on
|
|
290
|
-
the same pool at most twice (20 s, then 60 s, or the wait it names when that
|
|
291
|
-
is at most 2 minutes), then comes back to you (`✗ <step> needs you · rate
|
|
292
|
-
limited · backed off twice`). A try after a backoff reads `· after a
|
|
293
|
-
rate-limit backoff`. One that names a longer wait comes back to you at once,
|
|
294
|
-
with `back at` at the end of that wait. One whose pool is no longer free for
|
|
295
|
-
the backoff (at its 5-hour, weekly or monthly limit, or nearly spent in the
|
|
296
|
-
meantime) comes back to you at once too: as `out of quota` when that pool is
|
|
297
|
-
out on a usage limit, with `back at` its return when that is known. A sign-in
|
|
298
|
-
failure, a provider error or a worker that died at start still gets the step's
|
|
299
|
-
one automatic retry by itself, on another free pool when there is one. After a
|
|
300
|
-
sign-in failure that retry skips every pool that shares the credential;
|
|
301
|
-
nothing is stored, so a later step can pick that pool again. A model the
|
|
302
|
-
pool's plan lacks (`model not in plan`) retries on another pool, and that pool
|
|
303
|
-
never gets that model again until `strategy include-model <model>`.
|
|
304
|
-
|
|
305
|
-
When a return time is known, for this or any other failure, the block prints
|
|
306
|
-
`back at <time>` and adds one option:
|
|
307
|
-
|
|
308
|
-
| Printed option | When to choose it | What it runs |
|
|
309
|
-
|---|---|---|
|
|
310
|
-
| `rerun elsewhere` | another pool can run the step now | `bullswarm workflow step rerun <shortId> <step> --avoid <pool>` |
|
|
311
|
-
| `wait for it` | the step should run on that pool, or nothing else can run it | `after <time>: bullswarm workflow step rerun <shortId> <step>`: run that rerun yourself after the `back at` time; before then the pool is still out |
|
|
312
|
-
| `accept anyway` | you keep the step's result as it is | `bullswarm workflow step accept <shortId> <step> --reason "…"` |
|
|
313
|
-
|
|
314
|
-
`add steps` and `take over` are printed too. To stop the whole run instead,
|
|
315
|
-
run `bullswarm workflow cancel <shortId>`.
|
|
316
|
-
|
|
317
|
-
### A step that looks stale
|
|
318
|
-
|
|
319
|
-
The watcher prints `⚠ <step> looks stale: <reasons>` once per attempt when a
|
|
320
|
-
running step's score crosses its threshold (operations.md lists the reasons).
|
|
321
|
-
While Bullswarm runs a step's declared
|
|
322
|
-
checks, only quiet counts, read from the checks' heartbeat: `no check heartbeat
|
|
323
|
-
for <N>m`. Nothing is stopped for you: let it run (start the `next:` watch
|
|
324
|
-
again), or restart it with `bullswarm workflow step restart <shortId> <step>
|
|
325
|
-
[--pool <pool>]`, which stops that step only and runs it again with a handoff
|
|
326
|
-
of what the stopped attempt did.
|
|
112
|
+
Old v2 programs (`bullswarm.workflow.program.v2`) still run? No: since 0.38.0
|
|
113
|
+
a new run refuses them, and runs they started are view-only.
|
|
327
114
|
|
|
328
|
-
|
|
115
|
+
Start one watch right after launch; it prints nothing until a wake (a gate
|
|
116
|
+
waiting, a loop out of rounds, a step that needs you, a pause, a stale step,
|
|
117
|
+
steering, the end). Each exit is one wake: read the output, act, and start the
|
|
118
|
+
printed `next:` line again. If your harness cannot wake you when a background
|
|
119
|
+
process ends, run the watch in the foreground: it blocks until the wake; give
|
|
120
|
+
it a `--timeout` under your tool's time limit (`--until trouble --timeout 100`
|
|
121
|
+
for 2 minutes). A restart without `--after` attaches at the newest event and
|
|
122
|
+
skips wakes in between. Never end your turn while a run you own is still
|
|
123
|
+
running. At a gate: `bullswarm workflow continue <shortId> <gate>`; a loop out
|
|
124
|
+
of rounds takes `--rounds <1-5>`. `watch --until trouble` also wakes on
|
|
125
|
+
`steering received` (a person left guidance: decide what it means and add
|
|
126
|
+
steps).
|
|
127
|
+
|
|
128
|
+
## 4. Driving it well
|
|
129
|
+
|
|
130
|
+
Lessons from real runs; each holds for this version.
|
|
131
|
+
|
|
132
|
+
1. **One worker beats chunks.** Split only when one worker cannot hold the
|
|
133
|
+
input (the triage above): priority and consistency are judgements across
|
|
134
|
+
items, and chunks cannot make them.
|
|
135
|
+
2. **Write the judgement down.** A step whose decisions are spelled out (a
|
|
136
|
+
numbered spec, the rulings) runs well at `effort` `medium` or `low`; open
|
|
137
|
+
design needs `high`. Capability is model and reasoning together: the
|
|
138
|
+
effort tier picks both per pool (`bullswarm strategy rungs --json` shows
|
|
139
|
+
each rung), and a step's `reasoning` overrides only the level.
|
|
140
|
+
3. **A contract first.** When parallel steps must fit together, a first step
|
|
141
|
+
that writes the shared shape (types, an event format), which the others
|
|
142
|
+
depend on, keeps the join short; without it the integrator becomes the
|
|
143
|
+
author.
|
|
144
|
+
4. **Files decide parallelism.** Steps whose `files` overlap run one after
|
|
145
|
+
the other. Give each writer its own files and tell it to keep other
|
|
146
|
+
workers' edits; keep a breaking rename in one step, not parallel with its
|
|
147
|
+
consumers, which would build against the old name.
|
|
148
|
+
5. **Checks are facts.** Put anything a machine can say in `evidence`, and
|
|
149
|
+
run each check by hand before launch: fixing a wrong check reruns the
|
|
150
|
+
worker. Workers share one tree, so a later step can undo what an earlier
|
|
151
|
+
step's check proved; put the final check where nothing runs after it
|
|
152
|
+
(after integration, or on the step a gate waits behind).
|
|
153
|
+
6. **Reviews that hold.** State the contract as numbered checks, make any
|
|
154
|
+
difference a finding even when the author justifies it, have it record
|
|
155
|
+
`holds` and evidence per check with `passed` true only when every check
|
|
156
|
+
holds, route it `independentOf` the author, and add `blindTo` the author:
|
|
157
|
+
a builder's credible reason talks a reviewer who reads it out of a real
|
|
158
|
+
finding. Keep the brief consistent with the repository's own rules: a
|
|
159
|
+
brief that contradicts a repo test cannot pass honestly.
|
|
160
|
+
7. **Plan ahead, add later.** Declare the steps, gates and loops you can see.
|
|
161
|
+
Where the next part depends on an answer (one check per finding, one build
|
|
162
|
+
step per slice), stop there and add it with `workflow add` when the answer
|
|
163
|
+
arrives. Put a gate before expensive or outward work, and after a design
|
|
164
|
+
step the owner must approve.
|
|
165
|
+
8. **Watch by wake-ups.** One `watch --until trouble` per run. Between wakes
|
|
166
|
+
do not poll, read the run directory, or send per-step status replies. At a
|
|
167
|
+
needs-you block, pick one printed option and run it as printed.
|
|
168
|
+
9. **The worker's report is not proof.** Read the output and the diff.
|
|
169
|
+
`ok: true` and `proven by command` mean the checks passed, not that the
|
|
170
|
+
work is right; `answer checked` is a well-formed claim.
|
|
171
|
+
10. **Retries are for flakes.** The one automatic retry covers a crash or a
|
|
172
|
+
failed check. A worker whose check or deliverable failed and whose report
|
|
173
|
+
lists `- outside: <blocker>` under `## Not done` comes back at once. Fix
|
|
174
|
+
the cause with `workflow add` rather than rerunning the same step.
|
|
175
|
+
|
|
176
|
+
## 5. Read the outcome
|
|
329
177
|
|
|
330
178
|
`outcome:` is `completed`, `partial` or `cancelled`, and `reason:` says why in
|
|
331
179
|
one line. A completed v3 run hands nothing back: every step succeeded. The
|
|
332
180
|
proof line says what backs each step: `proven by command` or `proven by
|
|
333
181
|
schema` (a check Bullswarm ran passed), `answer checked` (its answer passed
|
|
334
182
|
its schema: a well-formed claim, not proof, so it is not counted as proven),
|
|
335
|
-
`finished · unproven` (neither), and `accepted by choice
|
|
183
|
+
`finished · unproven` (neither), and `accepted by choice` (your decision,
|
|
184
|
+
never verification). When you report the outcome, quote the run's proof line
|
|
336
185
|
as printed (`proof: …` at the end of watch, `# proof` in `runs result`)
|
|
337
186
|
instead of paraphrasing it. Then read the real outputs and answers
|
|
338
187
|
(`bullswarm workflow runs result <shortId> --json` names every step's output)
|
|
339
|
-
and probe the important edge cases yourself.
|
|
340
|
-
|
|
341
|
-
|
|
342
|
-
|
|
343
|
-
|
|
344
|
-
|
|
345
|
-
|
|
346
|
-
|
|
347
|
-
|
|
348
|
-
|
|
349
|
-
|
|
350
|
-
|
|
351
|
-
|
|
352
|
-
|
|
353
|
-
restart start a new run: bullswarm workflow goal "<goal>" --cwd /private/tmp/v37e/acme --program <file.json>
|
|
354
|
-
```
|
|
355
|
-
|
|
356
|
-
| Printed option | When | What to do |
|
|
357
|
-
|---|---|---|
|
|
358
|
-
| `add` | new work is needed, or a step must be done differently | `bullswarm workflow add <shortId> --steps part.json`, then `bullswarm workflow wait <shortId> <added ids>`; the run reopens |
|
|
359
|
-
| `retry` | a step stopped for a reason a retry fixes: a crashed or silent worker, a usage limit, or no pool free | `bullswarm workflow resume <shortId>` (run it after its `back at` time) |
|
|
360
|
-
| `rerun` | a step failed | `bullswarm workflow step rerun <shortId> <step> [--avoid <pool>]` |
|
|
361
|
-
| `accept` | you keep a failed step as it is | `bullswarm workflow step accept <shortId> <step> --reason "…"` |
|
|
362
|
-
| `take over` | the rest is small, or needs something only you have | do it yourself |
|
|
363
|
-
| `restart` | the goal or the approach was wrong | `bullswarm workflow goal "<goal>" --cwd <dir> --program <file.json>` |
|
|
364
|
-
|
|
365
|
-
`retry` appears only when a step is retryable. When nothing is retryable,
|
|
366
|
-
`resume` prints `nothing to retry`, starts nothing, and exits 1.
|
|
367
|
-
|
|
368
|
-
**An accept is a choice, never proof.** `step accept` lets the step's
|
|
369
|
-
dependents run, and the step reads `accepted by choice` and is counted apart
|
|
370
|
-
in the proof line (`N accepted by choice: <steps>`). Report it as your
|
|
371
|
-
decision, never as verification.
|
|
372
|
-
|
|
373
|
-
## 5. Extend or steer a workflow
|
|
374
|
-
|
|
375
|
-
A v3 run's steps, gates and loops are never edited. You change what happens
|
|
376
|
-
next by adding to it:
|
|
377
|
-
|
|
378
|
-
- `bullswarm workflow add <shortId> --steps part.json` appends a fragment
|
|
379
|
-
`{steps, gates?, loops?}`. New steps may depend on existing steps, gates and
|
|
380
|
-
loops, finished or not (on a loop's steps only through the loop's id).
|
|
381
|
-
Nothing the run has changes, and a finished run reopens, except that
|
|
382
|
-
`blocks: {"fix": ["report"]}` makes existing not-started steps also wait for
|
|
383
|
-
a step it adds (then `step accept` the failed step). Added steps do not
|
|
384
|
-
take the program's `defaults`: set `lane` and `effort` on each.
|
|
385
|
-
`--from-answer <step>` adds the fragment a step answered (read it with
|
|
386
|
-
`wait` first). Real output:
|
|
387
|
-
|
|
388
|
-
```text
|
|
389
|
-
✓ added to jcefns · revision 2 (applied directly; kernel relaunched)
|
|
390
|
-
reopened the completed run; its earlier result is archived
|
|
391
|
-
added step check-first
|
|
392
|
-
wait bullswarm workflow wait jcefns check-first
|
|
393
|
-
```
|
|
394
|
-
|
|
395
|
-
- `bullswarm workflow step rerun <shortId> <step>` runs a step again with its
|
|
396
|
-
last attempt's handoff; `step accept` keeps a failed one; `step restart`
|
|
397
|
-
stops a running one and runs it again.
|
|
398
|
-
- `bullswarm workflow pause <shortId>` starts nothing new (`--now` also stops
|
|
399
|
-
running steps); `bullswarm workflow resume <shortId>` lifts it.
|
|
400
|
-
`bullswarm workflow cancel <shortId>` stops the run.
|
|
401
|
-
- `watch --until trouble` also wakes on `steering received` (a person left
|
|
402
|
-
guidance for you: decide what it means and add steps) and on a pause.
|
|
403
|
-
|
|
404
|
-
A stopped step's file edits stay in the shared tree; when they must not
|
|
405
|
-
remain, add a step that reverts them. Exit 0 can mean launched, waiting,
|
|
406
|
-
paused or completed, so always read the returned status.
|
|
407
|
-
|
|
408
|
-
[operations.md](references/operations.md) covers the details: add and wait,
|
|
409
|
-
continue, reruns and accepts, pause and resume, watch flags, the handback
|
|
410
|
-
fields, routing diagnosis, and reading view-only saved runs.
|
|
188
|
+
and probe the important edge cases yourself. A `partial` run lists what is
|
|
189
|
+
unfinished and your options: see recovery.md.
|
|
190
|
+
|
|
191
|
+
## 6. Read this when
|
|
192
|
+
|
|
193
|
+
| Situation | Read |
|
|
194
|
+
|---|---|
|
|
195
|
+
| a step needs you (a needs-you block), a rate limit, a usage limit or no free pool | [recovery.md](references/recovery.md) |
|
|
196
|
+
| a step looks stale; pause, restart or cancel; a run ended `partial` | [recovery.md](references/recovery.md) |
|
|
197
|
+
| writing a program: every field, conditions, answers, evidence, deliverables | [program.md](references/program.md) |
|
|
198
|
+
| a review: the strict form and `blindTo` | [program.md](references/program.md) "Reviews", [patterns.md](references/patterns.md) 3 |
|
|
199
|
+
| copying a workflow: find then check, fix until green, draft to publish, parallel slices, triage | [patterns.md](references/patterns.md) |
|
|
200
|
+
| `workflow add` and `blocks`, `wait`, `continue`, watch modes and flags, handback fields, routing and reasoning diagnosis, view-only saved runs | [operations.md](references/operations.md) |
|
|
201
|
+
| adding a provider | [providers.md](references/providers.md) |
|
package/skill/agents/openai.yaml
CHANGED
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
interface:
|
|
2
2
|
display_name: "Bullswarm"
|
|
3
|
-
short_description: "
|
|
4
|
-
default_prompt: "Use $bullswarm to
|
|
3
|
+
short_description: "Delegate one task or a workflow you write"
|
|
4
|
+
default_prompt: "Use $bullswarm to choose one run or a workflow for this task, execute it, and check the result by facts."
|
|
5
5
|
|
|
6
6
|
policy:
|
|
7
7
|
allow_implicit_invocation: true
|