bullswarm 0.10.9 → 0.12.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +117 -0
- package/docs/claude-dynamic-workflow-mechanics.md +343 -0
- package/docs/experiments/2026-08-29-ultracode-vs-bullswarm.md +341 -0
- package/package.json +1 -1
- package/skill/SKILL.md +31 -0
- package/src/help.js +3 -2
- package/src/lib/verify.js +23 -1
- package/src/workflow/cli.js +5 -1
- package/src/workflow/dashboard.js +2 -1
- package/src/workflow/decision.js +98 -7
- package/src/workflow/goal.js +38 -8
- package/src/workflow/runner.js +258 -42
- package/src/workflow/runtime.js +114 -29
- package/src/workflow/template.js +41 -9
- package/src/workflow/validate.js +10 -3
- package/src/workflow/watch-cli.js +18 -3
package/CHANGELOG.md
CHANGED
|
@@ -1,5 +1,122 @@
|
|
|
1
1
|
# bullswarm changelog
|
|
2
2
|
|
|
3
|
+
## 0.12.0 — one decision is a whole program
|
|
4
|
+
|
|
5
|
+
Completes the convergence on Claude Code's dynamic-workflow mechanics
|
|
6
|
+
(`docs/claude-dynamic-workflow-mechanics.md` §1.11 and §4.2): the orchestrator
|
|
7
|
+
is positioned as the compiler of the goal into a program the runtime runs to
|
|
8
|
+
the end, and it is consulted again only at the program boundary (every action
|
|
9
|
+
finished, or the graph blocked).
|
|
10
|
+
|
|
11
|
+
- Data-driven fan-out in proposals: a planner `fanout` may carry
|
|
12
|
+
`itemsFrom: "outputs.<actionId>.outFile"` instead of inline `items`. The
|
|
13
|
+
producer becomes an implicit `dependsOn` and the runtime resolves the item
|
|
14
|
+
list when the producer finishes, so the planner no longer spends a round trip
|
|
15
|
+
waiting to see how many items discovery found. If the producer's output has
|
|
16
|
+
no parseable JSON array, the runtime runs ONE bounded, read-only extraction
|
|
17
|
+
action (`<fanoutId>-items`, `source: "runtime-extraction"`) over that output
|
|
18
|
+
before failing the fan-out truthfully. Resolved lists above
|
|
19
|
+
`maxItemsPerExpansion` fail the fan-out with the count. Events:
|
|
20
|
+
`action.items_resolved`, `action.items_extraction_requested`,
|
|
21
|
+
`action.items_extracted`.
|
|
22
|
+
- Repair policy on verify: `repair: { prompt, maxRounds (1–3), effort? }` on a
|
|
23
|
+
planner `verify`. When the verifier returns `ok:false`, the executor runs a
|
|
24
|
+
fix action (`<verifyId>-repair-<n>`, `source: "repair-policy"`) carrying the
|
|
25
|
+
verifier's concerns verbatim and re-runs the same verify, without a planner
|
|
26
|
+
turn. Dispatch or JSON-parse failures are not repaired. Events:
|
|
27
|
+
`action.repair_started`, `action.reverify_started`, `action.repaired`,
|
|
28
|
+
`action.reverify_rejected`, `action.repair_failed`.
|
|
29
|
+
- Fan-out artifact: a fan-out now writes `out-<id>-summary-*.md` (every item's
|
|
30
|
+
verdict plus an output excerpt) and records it as `outputs.<id>.outFile`, so a
|
|
31
|
+
verify can depend on a fan-out directly and `review` is inferred as usual.
|
|
32
|
+
- Fix: fan-out outputs stored the success COUNT in `ok`, so a dynamic action
|
|
33
|
+
depending on a fan-out could never become ready ("blocked by failed or
|
|
34
|
+
unresolved dependencies") and a fan-out never counted as a successful worker
|
|
35
|
+
for completion evidence. `ok` is now a boolean and the count moved to
|
|
36
|
+
`succeeded`; `fanoutSucceededCount()` reads pre-0.12 state files.
|
|
37
|
+
- Fix: the content gate rejected a worker whose whole answer is a JSON array
|
|
38
|
+
or object as an "announcement without substance", which is exactly what a
|
|
39
|
+
discovery step is told to return. Structured answers now pass
|
|
40
|
+
(`hasStructuredAnswer`).
|
|
41
|
+
- `parseJsonArray` prefers the trailing array, so prose containing brackets
|
|
42
|
+
before the list no longer poisons `itemsFrom`.
|
|
43
|
+
- Planner prompt: PLANNING DOCTRINE rewritten around "you are compiling the
|
|
44
|
+
goal into a program"; new data-driven fanout and repair skeletons; a
|
|
45
|
+
four-action program skeleton (discover → fanout(itemsFrom) → verify(repair)
|
|
46
|
+
→ verify-suite); `executionConstraints.programFeatures` and
|
|
47
|
+
`plannerConsultedOnlyAtProgramBoundary`. The goal orchestrator prompt is
|
|
48
|
+
reframed the same way.
|
|
49
|
+
- Literal double braces in prompts no longer kill actions. Only a known root
|
|
50
|
+
followed by dotted identifiers (`{{item}}`, `{{outputs.<id>.outFile}}`,
|
|
51
|
+
`{{inputs.x}}`, `{{runId}}`, `{{wfDir}}`) is a template ref; any other
|
|
52
|
+
`{{…}}` text (a JSDoc type such as `{{maxLength?: number}}`, Mustache, a JS
|
|
53
|
+
object in a template literal) is left exactly as written by the renderer
|
|
54
|
+
and ignored by the validator. Observed in the 0.11.1 comparison run: a
|
|
55
|
+
planner-authored verify prompt containing `{{maxLength?: number}}` failed
|
|
56
|
+
at render time with zero attempts. Relatedly, `verify` no longer
|
|
57
|
+
template-renders the artifact it reviews: only the reviewer instructions
|
|
58
|
+
are a template; the worker's report is appended verbatim.
|
|
59
|
+
- Doctrine: an `ok:true` verify is accepted; its concerns are informational.
|
|
60
|
+
The 0.11.1 comparison run spent a whole extra program round (7 actions,
|
|
61
|
+
~10 min) polishing "non-blocking" nits reported by verifiers that had
|
|
62
|
+
passed, which Claude's fix stage never does.
|
|
63
|
+
- Scout before compiling: `workflow goal` now starts with a read-only `scout`
|
|
64
|
+
run action (tree, manifest, test status, units of work with the files each
|
|
65
|
+
owns, shared files, risks; ends with a JSON array of unit names) so the
|
|
66
|
+
orchestrator's first program names real files and commands — the counterpart
|
|
67
|
+
of the inline scouting a Claude Code session does before authoring a
|
|
68
|
+
Workflow script. `--no-scout` skips it. A failed scout is non-fatal: the
|
|
69
|
+
planner still runs and sees `outputs.scout.ok=false` with the reason, and
|
|
70
|
+
the scout never counts as a delivery worker.
|
|
71
|
+
- The planner finally sees what workers said: every `outputs.<id>` in the
|
|
72
|
+
durable planner context carries `outputExcerpt` (up to 3 000 chars each,
|
|
73
|
+
36 000 total, newest first) instead of only `ok`/`why`/`outFile`.
|
|
74
|
+
- `workflow goal` default `maxItemsPerExpansion` raised 8 → 24 so a
|
|
75
|
+
data-driven fan-out over a medium repository does not fail on the bound.
|
|
76
|
+
- Known limitation: `itemsFrom` removes the planner turn, not the stage
|
|
77
|
+
barrier. A verify that depends on a data-driven fan-out waits for all items;
|
|
78
|
+
per-item verify overlap on discovered items would need chained
|
|
79
|
+
`stepTemplate`s (not in this release). For known items keep proposing N fix
|
|
80
|
+
+ N verify chains inline, which already overlap under the ready-set
|
|
81
|
+
scheduler.
|
|
82
|
+
|
|
83
|
+
## 0.11.1 — reliable `--watch` handoff
|
|
84
|
+
|
|
85
|
+
- `workflow goal --watch` no longer races the detached child: the watcher now
|
|
86
|
+
waits up to 30 s for the run's `state.json` to appear before attaching, and
|
|
87
|
+
`runWorkflowWatch` accepts `waitForRunMs`. The 0.11.0 tag failed to publish
|
|
88
|
+
because this race made the release-gate test fail on the CI runner; 0.11.1
|
|
89
|
+
carries the full 0.11.0 change set below.
|
|
90
|
+
|
|
91
|
+
## 0.11.0 — plan the whole graph, run it wide
|
|
92
|
+
|
|
93
|
+
Adopts the driving mechanics of Claude Code's dynamic workflow (documented in
|
|
94
|
+
`docs/claude-dynamic-workflow-mechanics.md`) into the autonomous loop.
|
|
95
|
+
|
|
96
|
+
- Ready-set scheduler: every planner action whose dependencies have succeeded
|
|
97
|
+
starts immediately, and a dependent action starts the moment its own inputs
|
|
98
|
+
finish rather than when the whole round finishes. The global
|
|
99
|
+
`settings.concurrency` limiter caps real parallelism. Verify-B now overlaps
|
|
100
|
+
fix-C exactly like a Claude `pipeline()` stage.
|
|
101
|
+
- Planning doctrine: the planner is told to propose the complete dependency
|
|
102
|
+
graph in one decision (per-item fix→verify chains plus one whole-system
|
|
103
|
+
verify), to declare file ownership per action and order same-file edits with
|
|
104
|
+
`dependsOn`, to write self-contained worker prompts, and what a planning round
|
|
105
|
+
trip costs. The goal orchestrator prompt no longer asks for "the smallest
|
|
106
|
+
useful set" of actions. `executionConstraints.concurrency` is exposed.
|
|
107
|
+
- `workflow goal` default `--concurrency` is 8 (was 3; max 16).
|
|
108
|
+
- The planner prompt's shared-working-tree caution now says what is actually
|
|
109
|
+
unsafe (whole-tree mutation, running the full suite while others edit) and
|
|
110
|
+
states that concurrent workers editing disjoint files is the expected mode;
|
|
111
|
+
the 0.10.9 orchestrator had cited the old wording as its reason not to fan
|
|
112
|
+
out ("Implementation is deliberately NOT fanned out … shared-target mutation
|
|
113
|
+
policy").
|
|
114
|
+
- `verify.review` is recovered when a planner puts instructions or a filesystem
|
|
115
|
+
path there: instructions move to `prompt`, the single dependency's artifact is
|
|
116
|
+
inferred, and any `review` that is not `outputs.<actionId>.outFile` is
|
|
117
|
+
rejected at validation (feeding the corrective turn) instead of failing a
|
|
118
|
+
dispatch after a full planning round trip.
|
|
119
|
+
|
|
3
120
|
## 0.10.9 — planner self-correction and honest silence
|
|
4
121
|
|
|
5
122
|
- An invalid or non-JSON orchestrator decision no longer fails the run. The
|
|
@@ -0,0 +1,343 @@
|
|
|
1
|
+
# How Claude Code drives a dynamic workflow — mechanics, and what bullswarm adopts
|
|
2
|
+
|
|
3
|
+
Written 2026-08-29 from inside a Claude Code session that has the `Workflow`
|
|
4
|
+
tool ("ultracode") loaded, by the model that authors those workflows.
|
|
5
|
+
|
|
6
|
+
Every statement is tagged:
|
|
7
|
+
|
|
8
|
+
- **[SPEC]** — quoted or closely paraphrased from the `Workflow` tool contract
|
|
9
|
+
and the `workflow-authoring` reference as loaded into the session on
|
|
10
|
+
2026-08-29. This is the documented behaviour the orchestrating model is
|
|
11
|
+
told to rely on.
|
|
12
|
+
- **[OBSERVED]** — measured in this repository's experiments
|
|
13
|
+
(`docs/experiments/2026-08-29-ultracode-vs-bullswarm.md`).
|
|
14
|
+
- **[INFERRED]** — my reading of how the harness must behave to satisfy the
|
|
15
|
+
spec. Not confirmed by source; treat as a hypothesis.
|
|
16
|
+
|
|
17
|
+
## 0. The one-paragraph shape
|
|
18
|
+
|
|
19
|
+
Claude's orchestrator (the main-loop model) does **not** decide step-by-step at
|
|
20
|
+
runtime. It writes a *program* — a small JavaScript script — that declares the
|
|
21
|
+
phases and calls `agent(prompt, opts)` once per worker, wired together with
|
|
22
|
+
`pipeline()` / `parallel()` / plain loops. The harness executes that program
|
|
23
|
+
deterministically, spawning subagents concurrently up to a cap, and the model
|
|
24
|
+
reads the aggregate return value when the program finishes. Planning is
|
|
25
|
+
front-loaded into one authoring act; parallelism is explicit in the code; every
|
|
26
|
+
worker receives an individually written, self-contained prompt and (usually) a
|
|
27
|
+
JSON schema its answer must satisfy. Re-planning happens either as ordinary
|
|
28
|
+
code (loops, conditionals) inside the script, or between scripts when the model
|
|
29
|
+
reads a result and authors the next one. **[SPEC]**
|
|
30
|
+
|
|
31
|
+
bullswarm's `workflow goal`, by contrast, runs an LLM *at every checkpoint*: a
|
|
32
|
+
`decide` step proposes JSON actions, the runtime validates and executes them,
|
|
33
|
+
then asks the LLM again. **[SPEC — bullswarm source]** The rest of this document
|
|
34
|
+
is about which of Claude's mechanics that loop can adopt without giving up its
|
|
35
|
+
one advantage — the user supplies a goal, never a graph.
|
|
36
|
+
|
|
37
|
+
## 1. Mechanics, one at a time
|
|
38
|
+
|
|
39
|
+
### 1.1 The control plane is code, authored once **[SPEC]**
|
|
40
|
+
|
|
41
|
+
- A script must begin with a pure-literal `export const meta = { name,
|
|
42
|
+
description, phases: [{ title, detail }] }`; the body uses `phase()`,
|
|
43
|
+
`agent()`, `pipeline()`, `parallel()`, `log()`, `args`, `budget`, and
|
|
44
|
+
`workflow()` (one level of nesting).
|
|
45
|
+
- The tool is explicitly for "multi-step orchestration where control flow
|
|
46
|
+
should be deterministic (loops, conditionals, fan-out) rather than
|
|
47
|
+
model-driven."
|
|
48
|
+
- The recommended pattern is **hybrid**: "scout inline first (list the files,
|
|
49
|
+
find the channels, scope the diff) to discover the work-list, then call
|
|
50
|
+
Workflow to pipeline over it. You don't need to know the shape before the
|
|
51
|
+
*task* — only before the *orchestration step*."
|
|
52
|
+
- Multi-phase work is several workflows in sequence: "run several in sequence —
|
|
53
|
+
read each result before deciding the next phase. You stay in the loop; each
|
|
54
|
+
workflow is one well-scoped fan-out."
|
|
55
|
+
|
|
56
|
+
Consequence: between two agents inside one workflow there is **no model
|
|
57
|
+
round-trip**. The next agent starts the instant its inputs exist. **[INFERRED
|
|
58
|
+
from spec; consistent with OBSERVED timings]**
|
|
59
|
+
|
|
60
|
+
### 1.2 Phases are progress groups, not barriers **[SPEC]**
|
|
61
|
+
|
|
62
|
+
- `phase(title)` starts a new phase; "subsequent agent() calls are grouped
|
|
63
|
+
under this title in the progress display". `opts.phase` on `agent()` assigns
|
|
64
|
+
the group explicitly and exists precisely "to avoid races on the global
|
|
65
|
+
phase() state" inside `pipeline()`/`parallel()` stages.
|
|
66
|
+
- Titles are matched exactly against `meta.phases`; an unmatched title "just
|
|
67
|
+
gets its own progress group".
|
|
68
|
+
- Nothing about a phase synchronises work. The **only** barrier is
|
|
69
|
+
`parallel()`.
|
|
70
|
+
|
|
71
|
+
### 1.3 Parallelism primitives **[SPEC]**
|
|
72
|
+
|
|
73
|
+
- `pipeline(items, stage1, stage2, …)`: "run each item through all stages
|
|
74
|
+
independently, NO barrier between stages. Item A can be in stage 3 while item
|
|
75
|
+
B is still in stage 1. This is the DEFAULT for multi-stage work. Wall-clock =
|
|
76
|
+
slowest single-item chain, not sum-of-slowest-per-stage." A stage that throws
|
|
77
|
+
drops that item to `null` and skips its remaining stages.
|
|
78
|
+
- `parallel(thunks)`: "run tasks concurrently. This is a BARRIER … Use ONLY
|
|
79
|
+
when you genuinely need all results together." A barrier is justified only
|
|
80
|
+
when stage N needs cross-item context from all of stage N−1 (dedup/merge,
|
|
81
|
+
early exit on zero count, "compare with the other findings").
|
|
82
|
+
- Anti-barrier guidance is explicit: "'I need to flatten/map/filter first' —
|
|
83
|
+
do it inside a pipeline stage"; "'The stages are conceptually separate' —
|
|
84
|
+
that's what pipeline() models. Separate stages ≠ synchronized stages."
|
|
85
|
+
- Caps: "Concurrent agent() calls are capped at min(16, available CPUs − 2) per
|
|
86
|
+
workflow — excess calls queue and run as slots free up." Lifetime cap 1000
|
|
87
|
+
agents; ≤ 4096 items per `parallel()`/`pipeline()` call.
|
|
88
|
+
|
|
89
|
+
### 1.4 One prompt per agent, structured return **[SPEC]**
|
|
90
|
+
|
|
91
|
+
- `agent(prompt, opts)`; "Subagents are told their final text IS the return
|
|
92
|
+
value (not a human-facing message), so they return raw data."
|
|
93
|
+
- `opts.schema` (JSON Schema): "the subagent is forced to call a
|
|
94
|
+
StructuredOutput tool and agent() returns the validated object — no parsing
|
|
95
|
+
needed … validation happens at the tool-call layer so the model retries on
|
|
96
|
+
mismatch."
|
|
97
|
+
- Other per-agent knobs: `label` (display), `phase` (group), `model`
|
|
98
|
+
(override; "default to omitting it"), `effort` (`low`…`max`), `isolation:
|
|
99
|
+
'worktree'` ("EXPENSIVE … use ONLY when agents mutate files in parallel and
|
|
100
|
+
would otherwise conflict"), `agentType` (custom subagent definition).
|
|
101
|
+
- `agent()` "returns null if the user skips the agent mid-run or the subagent
|
|
102
|
+
dies on a terminal API error after retries (filter with .filter(Boolean))."
|
|
103
|
+
|
|
104
|
+
### 1.5 Failure semantics live in code **[SPEC]**
|
|
105
|
+
|
|
106
|
+
- `parallel()` "never rejects" — a failing thunk becomes `null`.
|
|
107
|
+
- Retry, repair and convergence are ordinary loops: *loop-until-count*,
|
|
108
|
+
*loop-until-budget*, *loop-until-dry* ("keep spawning finders until K
|
|
109
|
+
consecutive rounds return nothing new"). No planner turn is spent deciding to
|
|
110
|
+
retry; the script already says so.
|
|
111
|
+
|
|
112
|
+
### 1.6 Determinism and resume **[SPEC]**
|
|
113
|
+
|
|
114
|
+
- `Date.now()`, `Math.random()`, argless `new Date()` throw in scripts ("they
|
|
115
|
+
would break resume"); timestamps come in via `args`.
|
|
116
|
+
- Every invocation persists its script; the run's `journal.jsonl` "records each
|
|
117
|
+
agent's actual return value". Resume = same script + `resumeFromRunId`: "the
|
|
118
|
+
longest unchanged prefix of agent() calls returns cached results instantly;
|
|
119
|
+
the first edited/new call and everything after it runs live."
|
|
120
|
+
|
|
121
|
+
### 1.7 Budget is a hard ceiling **[SPEC]**
|
|
122
|
+
|
|
123
|
+
- `budget.total / spent() / remaining()`; "The target is a HARD ceiling, not
|
|
124
|
+
advisory: once spent() reaches total, further agent() calls throw."
|
|
125
|
+
- bullswarm deliberately does the opposite (advisory targets, by user
|
|
126
|
+
decision: "a failed run is more costly than a slightly over budget run").
|
|
127
|
+
This difference is kept on purpose.
|
|
128
|
+
|
|
129
|
+
### 1.8 Observability **[SPEC]**
|
|
130
|
+
|
|
131
|
+
- Progress is rendered as a tree grouped by phase, with per-agent labels;
|
|
132
|
+
`log()` "emit[s] a progress message to the user (shown as a narrator line
|
|
133
|
+
above the progress tree)"; `/workflows` shows live progress; the tool result
|
|
134
|
+
carries the `runId` and transcript directory.
|
|
135
|
+
|
|
136
|
+
### 1.9 Quality patterns the orchestrator is told to compose **[SPEC]**
|
|
137
|
+
|
|
138
|
+
Adversarial verify (N skeptics prompted to refute; kill on majority),
|
|
139
|
+
perspective-diverse verify (distinct lenses instead of N identical refuters),
|
|
140
|
+
judge panel (N independent attempts → parallel judges → synthesis), loop-until-
|
|
141
|
+
dry, multi-modal sweep, completeness critic, and "no silent caps: if a workflow
|
|
142
|
+
bounds coverage (top-N, no-retry, sampling), log() what was dropped."
|
|
143
|
+
|
|
144
|
+
### 1.10 Ultracode **[SPEC]**
|
|
145
|
+
|
|
146
|
+
"When a system-reminder confirms ultracode is on, that opt-in is standing:
|
|
147
|
+
author and run a workflow for every substantive task by default … For
|
|
148
|
+
multi-phase work (understand → design → implement → review), that often means
|
|
149
|
+
several workflows in sequence — one per phase — so you stay in the loop between
|
|
150
|
+
them."
|
|
151
|
+
|
|
152
|
+
## 1.11 Upfront plan or mid-flight steering? — the answer
|
|
153
|
+
|
|
154
|
+
The question that decides how bullswarm should converge: does Claude prepare
|
|
155
|
+
all phases and parallelism **before** the workflow starts, or does it keep
|
|
156
|
+
planning **during** execution?
|
|
157
|
+
|
|
158
|
+
**Before. Entirely.** The orchestrator model writes one complete program, the
|
|
159
|
+
harness validates it, and then executes it without consulting the model again.
|
|
160
|
+
Evidence:
|
|
161
|
+
|
|
162
|
+
- **[SPEC]** the script is passed whole and parsed before any agent runs;
|
|
163
|
+
"Use this tool for multi-step orchestration where control flow should be
|
|
164
|
+
deterministic (loops, conditionals, fan-out) rather than model-driven";
|
|
165
|
+
"Workflows run in the background — this tool returns immediately … a
|
|
166
|
+
task-notification arrives when the workflow completes."
|
|
167
|
+
- **[OBSERVED, goal 2]** ~4 minutes of inline scouting and advisor review →
|
|
168
|
+
one 23 k-character script (probe → author → adversarial verify → fix-loop,
|
|
169
|
+
per module) → rejected on a parse error before any agent spawned → corrected
|
|
170
|
+
in 80 s → ~20 minutes of execution with six agents concurrent and **zero
|
|
171
|
+
orchestrator decisions**. The session's own words: "the harness will
|
|
172
|
+
re-invoke me when it completes, so polling would just burn tokens. Waiting."
|
|
173
|
+
|
|
174
|
+
What *looks* like mid-flight steering is pre-authored into the program:
|
|
175
|
+
|
|
176
|
+
1. **Data-driven shape.** A stage's structured result (`schema`) becomes the
|
|
177
|
+
next stage's items: `pipeline(discovery.failures, fix, verify)`,
|
|
178
|
+
loop-until-dry. The author fixes the *policy*; the runtime fixes the *size*.
|
|
179
|
+
2. **Repair as code.** `if (!verify.ok) { fix; verify }` bounded by a counter.
|
|
180
|
+
Every fix → re-verify handoff observed in the goal-2 journal was this `if`.
|
|
181
|
+
3. **Budget as data.** `budget.remaining()` scales loops; the ceiling is hard.
|
|
182
|
+
|
|
183
|
+
Model-level re-planning exists only **at workflow boundaries**: "run several
|
|
184
|
+
in sequence — read each result before deciding the next phase"; the hybrid
|
|
185
|
+
"scout inline first … then call Workflow"; and edit-and-resume ("the longest
|
|
186
|
+
unchanged prefix of agent() calls returns cached results instantly; the first
|
|
187
|
+
edited/new call and everything after it runs live"). Steering means stop, edit
|
|
188
|
+
the program, resume — never a per-step decision.
|
|
189
|
+
|
|
190
|
+
**Consequence for bullswarm.** Converge on *Direction A — a program, not a
|
|
191
|
+
step list*: one planning turn produces the complete graph **plus** the
|
|
192
|
+
adaptation policy, the runtime executes it to completion, and the planner is
|
|
193
|
+
consulted again only at the boundary (complete, or next program). Do not build
|
|
194
|
+
a continuous mid-flight steering loop (Direction B); Claude has none inside a
|
|
195
|
+
workflow, and bullswarm already has `steer`, cancel and resume for the
|
|
196
|
+
boundary-level levers. What the runtime still lacks to express a program is
|
|
197
|
+
listed in §4.2 (0.12.0 scope).
|
|
198
|
+
|
|
199
|
+
## 2. What that looks like from the outside **[OBSERVED]**
|
|
200
|
+
|
|
201
|
+
Filled from the experiment report as runs complete. Numbers here are copied
|
|
202
|
+
from `docs/experiments/2026-08-29-ultracode-vs-bullswarm.md`, never projected.
|
|
203
|
+
|
|
204
|
+
- bullswarm 0.10.9 smoke goal ("create hello.txt and verify it"), single pool
|
|
205
|
+
`claude-code` pinned to `claude-opus-5`: wall 553 s; **4 orchestrator turns =
|
|
206
|
+
438 s (79 % of wall)**; 2 worker attempts = 114 s; max concurrency 1; 6
|
|
207
|
+
dispatches; ~29 k tokens (bullswarm's byte/4 estimate). Two of the four turns
|
|
208
|
+
were spent recovering from a `verify` proposal whose `review` field carried
|
|
209
|
+
instructions instead of an artifact path — a shape the 0.10.9 planner
|
|
210
|
+
skeleton itself had suggested.
|
|
211
|
+
- bullswarm 0.10.9 on the 7-module fixture (baseline, `--concurrency 8`): the
|
|
212
|
+
orchestrator's first decision (152 s) proposed one serial chain
|
|
213
|
+
discover → implement → verify and wrote: "Implementation is deliberately NOT
|
|
214
|
+
fanned out: all fixes land in one shared working tree and converge on
|
|
215
|
+
src/index.js, so concurrent workers would violate the shared-target mutation
|
|
216
|
+
policy and race on the barrel file." The policy it cites was a caution line
|
|
217
|
+
in the planner prompt; a concurrency cap of 8 was available and unused.
|
|
218
|
+
Discovery alone then ran 458 s. (Final numbers: experiment report.)
|
|
219
|
+
- (comparison fixture final results: pending)
|
|
220
|
+
|
|
221
|
+
## 3. bullswarm today, mechanic by mechanic
|
|
222
|
+
|
|
223
|
+
| Mechanic | Claude `Workflow` | bullswarm ≤ 0.10.9 | Gap |
|
|
224
|
+
| --- | --- | --- | --- |
|
|
225
|
+
| Control plane | Code, authored once; no model call between agents | LLM `decide` turn at every checkpoint; each turn is a fresh `claude -p --resume` process reading the full durable context | Structural. Reachable target: **one planning turn per replan-worthy event** (initial DAG; then only on failure/completion), not per action |
|
|
226
|
+
| Phases | Labels for grouping; never synchronise | Forward-only kebab-case names per action; also just labels | None |
|
|
227
|
+
| Parallelism | `pipeline` default, `parallel` barrier; cap min(16, CPUs−2) | `executeActions` ran dependency-ready siblings **serially** (`runner.js:558`); only `fanout` items ran concurrently; goal default concurrency 3 | **Fixed in 0.11.0** — ready-set scheduler + default 8 |
|
|
228
|
+
| Planner bias | Script author is told to fan out and default to pipeline | Goal prompt said "return needs_more_work with the **smallest useful set** of bounded … actions" (`goal.js:18`) and planner prompt said "keep actions cohesive" | **Fixed in 0.11.0** — "propose the COMPLETE dependency graph", per-item fix→verify chains, file ownership, self-contained prompts |
|
|
229
|
+
| Per-agent prompt | Self-contained, plus JSON schema enforced at tool layer | Planner-authored prompt; free-text answer, content-verified by heuristics; `verify` returns JSON verdict | Partial. Schema-enforced worker output is a candidate, not adopted yet |
|
|
230
|
+
| Failure handling | Loops in code; `null` on agent death | Planner replans (costly); 0.10.9 added corrective turns for invalid decisions and 0.11.0 recovers mis-shaped `verify.review` before dispatch | Improved; retry-in-code per action still absent |
|
|
231
|
+
| Determinism / resume | Journal of return values; prefix cache | Durable `state.json` + `events.jsonl` + action ledger; resume skips durable outputs | Equivalent |
|
|
232
|
+
| Data-driven fan-out | `pipeline(discovered.items, …)` — count unknown when the script is written | Decision schema forced inline `items`; the planner spent a turn waiting for discovery | **Fixed in 0.12.0** — `itemsFrom` on proposed fan-outs + one bounded extraction retry |
|
|
233
|
+
| Repair loops | `while`/retry in code | Planner replanned after every failed verify | **Fixed in 0.12.0** — `verify.repair` policy runs fix → re-verify inside the executor |
|
|
234
|
+
| Budget | Hard ceiling | Advisory targets (user decision) | Intentional difference |
|
|
235
|
+
| Observability | Progress tree, narrator, `/workflows` | `watch` heartbeat (semantic quiet + agent-output quiet since 0.10.9), `tui`, events | Comparable |
|
|
236
|
+
| Isolation | `isolation: 'worktree'` per agent | Shared `addDir`; planner-declared file ownership | Candidate |
|
|
237
|
+
|
|
238
|
+
## 4. Adopted into bullswarm 0.11.0
|
|
239
|
+
|
|
240
|
+
1. **Ready-set scheduler** (`src/workflow/runner.js`, `executeActions`).
|
|
241
|
+
Every action whose `dependsOn` have all succeeded is launched immediately;
|
|
242
|
+
a dependent action starts the moment *its own* dependencies finish, not when
|
|
243
|
+
the whole round finishes. The global dispatch limiter
|
|
244
|
+
(`settings.concurrency`) caps real concurrency. This gives `pipeline`
|
|
245
|
+
semantics to any DAG the planner proposes: verify-B overlaps fix-C.
|
|
246
|
+
Test: `dependency-ready sibling actions run concurrently and dependents
|
|
247
|
+
start as soon as their own inputs finish` (`tests/workflow-adaptive.test.js`).
|
|
248
|
+
2. **Planning doctrine in the planner prompt** (`src/workflow/runtime.js`) and
|
|
249
|
+
goal orchestrator prompt (`src/workflow/goal.js`): propose the complete graph
|
|
250
|
+
in one decision; independent actions run concurrently; per-item fix→verify
|
|
251
|
+
chains plus one final whole-system verify; explicit file ownership per
|
|
252
|
+
action and `dependsOn` for any same-file edits; self-contained worker
|
|
253
|
+
prompts with absolute paths and the exact acceptance command. The prompt
|
|
254
|
+
also states the cost of a planning turn so the model can weigh it. The
|
|
255
|
+
context exposes `executionConstraints.concurrency` and
|
|
256
|
+
`readySiblingsRunConcurrently: true`.
|
|
257
|
+
3. **Default `--concurrency` 8** for `workflow goal` (was 3; max 16).
|
|
258
|
+
4. **`verify.review` contract made survivable**: instructions placed in
|
|
259
|
+
`review` are moved to `prompt` and the single dependency's artifact is
|
|
260
|
+
inferred; a `review` that is not `outputs.<actionId>.outFile` is rejected at
|
|
261
|
+
validation (so the 0.10.9 corrective turn fixes it) instead of failing a
|
|
262
|
+
dispatch after a planning round trip.
|
|
263
|
+
|
|
264
|
+
### 4.2 Adopted in 0.12.0 — a *program* expressible in one decision
|
|
265
|
+
|
|
266
|
+
The user's framing for this release: position the orchestrator as the
|
|
267
|
+
**compiler** of the goal into a workflow program; the program drives every
|
|
268
|
+
phase and turn; the model is consulted again only at a boundary that needs
|
|
269
|
+
judgement — the same division of labour Claude Code uses between the script
|
|
270
|
+
author and the `Workflow` runtime.
|
|
271
|
+
|
|
272
|
+
1. **Data-driven fan-out in proposals** — shipped. A proposed `fanout` takes
|
|
273
|
+
`itemsFrom: "outputs.<actionId>.outFile"` (producer may be co-proposed; it
|
|
274
|
+
becomes an implicit `dependsOn`). The runtime resolves the list when the
|
|
275
|
+
producer finishes. This is Claude's `pipeline(discovery.failures, …)`: the
|
|
276
|
+
planner no longer spends a turn waiting to see how many items there are.
|
|
277
|
+
2. **Structured worker output** — shipped in its cheap form. Discovery workers
|
|
278
|
+
are told to end with a JSON array; `parseJsonArray` prefers the trailing
|
|
279
|
+
array; the content gate accepts a bare JSON array/object as substance; and
|
|
280
|
+
if the output still has no array the runtime runs ONE bounded, read-only
|
|
281
|
+
extraction action over it (never re-running the producer, which may have
|
|
282
|
+
mutated files). That is the "schema retry" of Claude's `StructuredOutput`,
|
|
283
|
+
done as a second cheap agent instead of a tool-layer retry. A general
|
|
284
|
+
`outputSchema` on run actions is still open (§5).
|
|
285
|
+
3. **Pre-authored repair** — shipped. `repair: { prompt, maxRounds }` on a
|
|
286
|
+
verify: verify-fail → `<verifyId>-repair-<n>` (concerns verbatim) →
|
|
287
|
+
re-verify, inside the executor. Claude's fix-loop as code.
|
|
288
|
+
4. **Boundary-only consultation** — the loop is now: decision 1 = the program;
|
|
289
|
+
the ready-set executor runs it to completion (fan-outs resolve, repairs run);
|
|
290
|
+
decision 2 = `complete` or the next program. The planner prompt says so
|
|
291
|
+
explicitly (`plannerConsultedOnlyAtProgramBoundary`), and the goal
|
|
292
|
+
orchestrator prompt is reframed as "compile the goal into a complete
|
|
293
|
+
workflow program".
|
|
294
|
+
5. **Scout, then compile** — shipped. Claude's author reads the repo inline
|
|
295
|
+
for ~4 min before writing the script; bullswarm's orchestrator was compiling
|
|
296
|
+
blind (goal text + cwd, forbidden to run commands, and the planner context
|
|
297
|
+
exposed no worker output text at all — only `ok`/`why`/`outFile`). `workflow
|
|
298
|
+
goal` now runs a read-only `scout` action first, and every output in the
|
|
299
|
+
planner context carries an `outputExcerpt`, so the first program is written
|
|
300
|
+
against a real survey and boundary decisions read what workers reported.
|
|
301
|
+
6. Two bugs found on the way that had silently blocked this shape in ≤ 0.11.1:
|
|
302
|
+
fan-out outputs recorded the success *count* in `ok`, so nothing could ever
|
|
303
|
+
depend on a fan-out (the ready-set test is `ok === true`); and the content
|
|
304
|
+
gate rejected a worker whose whole answer was a JSON array as an
|
|
305
|
+
"announcement without substance".
|
|
306
|
+
|
|
307
|
+
7. Two robustness gaps the 0.11.1 comparison run itself exposed, fixed
|
|
308
|
+
before 0.12.0 shipped: (a) a planner-authored verify prompt that quoted a
|
|
309
|
+
JSDoc type literally — `{{maxLength?: number}}` — was parsed as a template
|
|
310
|
+
ref and killed the action at render time with zero attempts, forcing an
|
|
311
|
+
extra planner turn to re-issue it. Only a known root plus dotted
|
|
312
|
+
identifiers is a ref now; other double-brace text is prompt content.
|
|
313
|
+
(Claude never has this class of bug: prompts are JS strings, the runtime
|
|
314
|
+
does no substitution.) (b) The new scout is a failable step ahead of the
|
|
315
|
+
planner; it is non-fatal by construction (`onError: continue`), the planner
|
|
316
|
+
sees `outputs.scout.ok=false` with the reason, and a run where only the
|
|
317
|
+
scout succeeded is `blocked`, never "delivered".
|
|
318
|
+
|
|
319
|
+
**Honest limitation.** `itemsFrom` removes the planner *turn*, not the stage
|
|
320
|
+
*barrier*: a verify depending on a data-driven fan-out waits for all items,
|
|
321
|
+
whereas Claude's `pipeline()` overlaps verify-B with fix-C for discovered items
|
|
322
|
+
too. Per-item overlap on unknown items would need a fan-out whose
|
|
323
|
+
`stepTemplate` is itself a chain — not in 0.12.0. For known items the planner
|
|
324
|
+
proposes N fix + N verify inline and the ready-set scheduler already overlaps
|
|
325
|
+
them.
|
|
326
|
+
|
|
327
|
+
## 5. Not adopted (yet), and why
|
|
328
|
+
|
|
329
|
+
- **Schema-enforced worker output.** bullswarm's content verification and the
|
|
330
|
+
JSON `verify` verdict cover the failure mode today; adding per-action
|
|
331
|
+
`outputSchema` is the next step if planners keep re-asking workers for
|
|
332
|
+
structure.
|
|
333
|
+
- **Per-action worktree isolation.** File ownership declared by the planner is
|
|
334
|
+
cheaper and matches Claude's own guidance ("EXPENSIVE … use ONLY when agents
|
|
335
|
+
mutate files in parallel and would otherwise conflict").
|
|
336
|
+
- **Hard budgets.** Deliberately advisory; see 1.7.
|
|
337
|
+
- **Retry policy in the graph.** Claude writes repair loops as code. A per-
|
|
338
|
+
action `retryPolicy` proposed by the planner would remove one planner turn per
|
|
339
|
+
transient failure; not yet built.
|
|
340
|
+
- **Identical control plane.** bullswarm's orchestrator remains an LLM per
|
|
341
|
+
checkpoint because the user supplies only a goal. The convergence target is
|
|
342
|
+
therefore "few planning turns, full width between them", measured as
|
|
343
|
+
`plannerSec / wallSec` and `maxConcurrentAttempts` in the experiment report.
|