bullswarm 0.19.0 → 0.21.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/AGENTS.md CHANGED
@@ -57,18 +57,21 @@ bullswarm workflow runs delete <shortId> --yes
57
57
  ## Using bullswarm from another agent
58
58
 
59
59
  If you are an agent that wants to offload bounded work via bullswarm,
60
- read `skill/SKILL.md` — that's the agent-facing user guide. It covers
61
- when to reach for `run` vs `workflow draft`, the verify-step pattern,
62
- how to write prompts that pass the content gate, and the failure modes
63
- you'll hit. The skill is published alongside the package and is the
64
- canonical reference for the CLI surface.
60
+ read `skill/SKILL.md` — that's the agent-facing user guide. Use
61
+ `bullswarm delegate` (or the installed `/bullswarm` skill) by default: it
62
+ previews whether one bounded agent or an autonomous workflow is appropriate,
63
+ shows the conceptual plan, and executes the chosen engine. Reach for `run`,
64
+ `workflow goal`, or a fixed workflow graph directly only when the caller has
65
+ already chosen that execution shape. The skill is published alongside the
66
+ package and is the canonical reference for the CLI surface.
65
67
 
66
68
  - Zero runtime dependencies. Node >= 18. Tests must never require network:
67
69
  prime `~/.bullswarm/meters/*.json` caches with fresh timestamps if needed.
68
70
  - Every verb must work non-interactively (no TTY). The interactive wizard is
69
71
  a human convenience, never a requirement.
70
72
  - Version single source: package.json. Release via
71
- `node bin/bullswarm.js release patch|minor|major` then `git push --tags`
73
+ `node bin/bullswarm.js release patch|minor|major`, then `git push` and
74
+ `git push --tags`
72
75
  — CI publishes through npm trusted publishing (OIDC), no tokens.
73
76
 
74
77
  ## Adding a connector
package/CHANGELOG.md CHANGED
@@ -1,5 +1,99 @@
1
1
  # bullswarm changelog
2
2
 
3
+ ## 0.21.0 — unified TUI shell and LLM-first delegation
4
+
5
+ - The interactive workflow viewer is now one application shell instead of
6
+ several screens with their own rules. The workflows list, a run, a phase and
7
+ an agent are four depths of one hierarchy: every screen carries the same
8
+ persistent breadcrumb at the top (`Workflows › hdtdxs · timeline-segments ›
9
+ Verify › verify-renderer`), which drops its deepest segments first when the
10
+ terminal is too narrow to hold the whole path. All four depths share key
11
+ bindings generated from a single key-map definition: Up/Down (or k/j) move
12
+ within the current level, Enter and Right (or l) go one level in, Esc and
13
+ Left (or h) go one level out, and Tab/Shift+Tab jump to the next or previous
14
+ workflow, re-entering the sibling at the same depth when the equivalent phase
15
+ exists. The drill-down layout is uniform too — a left sidebar listing the
16
+ current level beside a right pane previewing the highlighted item — at every
17
+ depth and on narrow terminals as well, including the run list, which was a
18
+ full-width table with no preview pane before.
19
+
20
+ - Workflow timelines now render in phase-segmented sections with continued
21
+ headers for interleaved phases, grouped Preflight scout/planner milestones,
22
+ elapsed or running phase state, and consistent desktop, narrow, and scroll
23
+ continuation behavior without per-line phase prefixes. Phase-completion
24
+ summary rows are retained, so rows such as `└─✓ completed 4/4` remain visible
25
+ beneath their phase headers.
26
+
27
+ - Non-interactive `workflow tui <run-id>` output now includes the same static
28
+ segmented timeline as the interactive viewer, alongside the historical detail
29
+ tree, so real command output can be used to inspect and verify the layout.
30
+
31
+ - Documentation now describes concerns as an attribute of a completed
32
+ outcome — `outcome.concerns` on a delivered result — rather than
33
+ presenting `completed_with_concerns` as its own terminal status to handle
34
+ separately from `completed`. Run records that carry the
35
+ `completed_with_concerns` status value, including legacy runs recorded
36
+ before this framing, remain fully readable: `runs result`, the TUI, and
37
+ `workflow watch` still read it exactly like `completed` — a delivered
38
+ result with concerns to review, never a failure.
39
+
40
+ - Delegation classification now starts with deterministic signals and, in
41
+ automatic execution, lets an LLM refine the choice between a single delegate
42
+ and a workflow. `--dry-run` performs that same bounded low-effort
43
+ classification request (one analyze-lane, low-effort dispatch) before
44
+ printing the plan, so the preview matches what a live run would decide — it
45
+ never dispatches the work itself. `--classify deterministic` bypasses that
46
+ refinement and remains the instant, no-dispatch preview; `--classify llm`
47
+ requires the refinement and fails if no usable LLM decision is available.
48
+ An explicit `--mode single|workflow` remains the caller's choice and bypasses
49
+ automatic LLM classification.
50
+
51
+ - OpenCode connector portability: `connectors/opencode2.json` no longer
52
+ hardcodes `--model kaihk/gpt-5.6-luna` in `spawn.cmd`. A plain OpenCode
53
+ install with no KaiHK provider configured now dispatches with OpenCode's
54
+ own default model instead of failing to resolve a KaiHK-only model.
55
+ `src/lib/opencode-kaihk.js` still injects the explicit
56
+ `--model <providerId>/gpt-5.6-luna` for each discovered KaiHK provider, so
57
+ the primary `opencode2` pool and any extra `opencode2:<id>` pools keep
58
+ dispatching with their pinned per-provider model exactly as before.
59
+
60
+ ## 0.20.0 — common agent delegation entry point
61
+
62
+ - `/bullswarm` and `bullswarm delegate` now give agents one transparent entry
63
+ point for arbitrary self-contained tasks: classify the request as one bounded
64
+ delegate or an autonomous workflow, show the reason and conceptual plan, then
65
+ execute the selected engine. Explicit mode and lane overrides remain
66
+ available, and `--dry-run --json` exposes the decision after its bounded
67
+ classification request without dispatching the work itself.
68
+ - Workflow decisions persist the suggested conceptual plan alongside the
69
+ original intent, while the packaged skill keeps the common path concise and
70
+ moves operational detail into a focused reference.
71
+ - Planner context now labels the preflight scout as completion-ineligible and
72
+ requires the first program to contain a real delivery worker plus its
73
+ verifier, preventing an apparently complete scout report from causing a
74
+ rejected completion and redundant recovery round.
75
+ - Ready siblings now honor a connector-owned soft concurrency preference. The
76
+ OpenCode route prefers one in-flight worker, so additional parallel work is
77
+ spread across healthy subscriptions instead of risking correlated headless
78
+ session exits; a lone eligible pool still runs rather than failing capacity.
79
+ - Agent integration upgrades its managed awareness marker to advertise the
80
+ common interface consistently across Codex, Claude, and Grok.
81
+ - Classification now understands negated and instructional mutation language,
82
+ so read-only requests that discuss how to add or write something do not
83
+ accidentally enter the build lane, while a later affirmative implementation
84
+ request still does.
85
+ - A trailing `help` token remains contextual and side-effect-free even after
86
+ options, matching `-h` and `--help`; several README and setup/help examples
87
+ were also brought back into sync with the real CLI.
88
+ - Historical workflow design documents now identify themselves as dated
89
+ implementation records and list the current `verify`, `decide`, and
90
+ `outputSchema` surfaces instead of presenting resolved gaps as current.
91
+ - Extra KaiHK providers in `~/.config/opencode/opencode.json` (`kaihk-2`, …)
92
+ become `opencode2:<id>` pools, spawned with `--model <id>/gpt-5.6-luna`.
93
+ Spend is read from `GET /api/usage/token` plus
94
+ `/v1/dashboard/billing/usage` (USD = `total_usage / 100`). The HTML wallet
95
+ page still needs a browser session and is not the key API.
96
+
3
97
  ## 0.19.0 — unified workflow dashboard
4
98
 
5
99
  - Running `bullswarm workflow` on an interactive terminal now opens one
package/README.md CHANGED
@@ -5,6 +5,11 @@ orchestrator, build and expand the plan, route bounded worker actions by quota,
5
5
  verify the result, and finish without an initiating agent authoring a graph.
6
6
  Every delegate output is judged by content before it counts.
7
7
 
8
+ For agents, `/bullswarm` (or `$bullswarm` where skills use that syntax) is the
9
+ common entry point. Its durable CLI equivalent is `bullswarm delegate`: it
10
+ first explains whether the request needs one bounded agent or an autonomous
11
+ workflow, shows the conceptual plan, and then executes the selected engine.
12
+
8
13
  Every command and nested subcommand supports contextual `-h` / `--help`
9
14
  without initializing state or executing the command:
10
15
 
@@ -66,6 +71,9 @@ bullswarm # first run: interactive setup wizard
66
71
  bullswarm setup # re-run or repair
67
72
  bullswarm pools # meter state, pace position, quarantine status
68
73
  bullswarm strategy refresh --apply --yes # approve capability-aware tier autopilot
74
+ bullswarm delegate --cwd ~/some-repo --prompt "Explain the parser" # one agent
75
+ bullswarm delegate --cwd ~/some-repo --prompt "Audit all commands, fix help, and independently verify" # workflow
76
+ bullswarm delegate --dry-run --json --cwd ~/some-repo --prompt "Your task" # bounded classification + decision/plan; no work dispatch
69
77
  bullswarm run --lane analyze --add-dir ~/some-repo --task-file /tmp/t.md --json
70
78
  bullswarm run --lane analyze --add-dir ~/some-repo --prompt "Inspect the parser" --json
71
79
  bullswarm workflow goal "Fix the failing tests and verify the change" --cwd ~/some-repo
@@ -78,6 +86,7 @@ bullswarm health # re-judge saved outputs; catch gate failures
78
86
  |---|---|
79
87
  | `setup` | Discover installed agent CLIs, show quota state, toggle pools, suggest a routing table, write config. Approval-gated, idempotent. |
80
88
  | `integrate` | Register or remove the canonical Bullswarm skill and global awareness rules for Codex, Claude, and Grok. |
89
+ | `delegate` | Explain and execute the smallest reliable shape: one content-verified agent or an autonomous verified workflow. |
81
90
  | `run` | route → dispatch → watch → verify → one JSON verdict |
82
91
  | `health` | Re-judge saved outputs against their verdicts; surface verify-gate failures and quarantine clusters |
83
92
  | `pools` | Show each pool's meter state, pace position, quarantine status |
@@ -85,6 +94,22 @@ bullswarm health # re-judge saved outputs; catch gate failures
85
94
  | `doctor` | Machine-readable readiness report; self-heals on first call |
86
95
  | `workflow` | Start an autonomous goal, or run / validate / draft / inspect explicit workflows and their live instances. |
87
96
 
97
+ ### Delegate classification
98
+
99
+ With the default `--mode auto`, `delegate` first uses deterministic task
100
+ signals, then uses an LLM to refine the choice between a single delegate and a
101
+ workflow during execution. If that optional refinement is unavailable or
102
+ unusable, automatic mode uses the deterministic decision.
103
+
104
+ Use `--classify deterministic` to bypass the LLM refinement — this is the
105
+ instant, no-dispatch preview. Use `--classify llm` when an LLM decision is
106
+ required: the command fails if it cannot obtain a usable one. In automatic
107
+ mode, `--dry-run` still performs that same bounded low-effort classification
108
+ request (one analyze-lane, low-effort dispatch) and prints the resulting
109
+ decision — it never dispatches the work itself. An explicit `--mode single` or
110
+ `--mode workflow` is the caller's decision and bypasses automatic LLM
111
+ classification.
112
+
88
113
  Discover and validate workflow definitions without executing them:
89
114
 
90
115
  ```bash
@@ -204,7 +229,18 @@ bullswarm workflow goal "Implement and verify the change" --cwd . \
204
229
  The worker lock covers the scout, ordinary runs, fan-out items, repairs,
205
230
  re-verification, and runtime extraction helpers. A pool that cannot guarantee
206
231
  the requested model is ineligible rather than silently substituting another
207
- model. `--max-agents` and `--max-workflow-seconds` are
232
+ model.
233
+
234
+ The `opencode2` connector itself does not require a KaiHK provider: its base
235
+ spawn command carries no hardcoded model, so a plain OpenCode installation
236
+ dispatches with OpenCode's own configured default. When
237
+ `~/.config/opencode/opencode.json` has one or more KaiHK providers configured,
238
+ Bullswarm discovers them and pins an explicit `--model <providerId>/gpt-5.6-luna`
239
+ per provider — the first as the primary `opencode2` pool, each additional one
240
+ as its own `opencode2:<id>` pool — which is what the `--worker-model
241
+ kaihk/gpt-5.6-luna` example above locks onto.
242
+
243
+ `--max-agents` and `--max-workflow-seconds` are
208
244
  advisory planning targets; `--max-expansion-rounds` is also an advisory
209
245
  convergence target. Hard structural safeguards are adjusted with
210
246
  `--max-actions` and `--max-items-per-expansion`.
@@ -227,7 +263,7 @@ bullswarm workflow draft phase add audit-code discover
227
263
  bullswarm workflow draft phase add audit-code review
228
264
  bullswarm workflow draft step add audit-code discover list-files \
229
265
  --type run --lane chore --prompt "List every .js file in src/" \
230
- --addDir '{{inputs.targetDir}}'
266
+ --add-dir '{{inputs.targetDir}}'
231
267
  bullswarm workflow draft step add audit-code review per-file \
232
268
  --type fanout --items-from 'outputs.list-files.outFile' \
233
269
  --lane analyze --concurrency 2 \
@@ -440,11 +476,16 @@ rather than discarding the run as a blanket failure. Delegates have no
440
476
  implicit wall-clock timeout; set a step's `timeoutSec` (or direct-run
441
477
  `--timeout`) only when an operator explicitly wants a hard termination timer.
442
478
 
443
- An autonomous `complete` remains strictly verified. A planner `stop` produces
444
- `completed_with_concerns` when a useful delivery exists, including unresolved
445
- verification concerns and the stopping reason; it produces `blocked` only
446
- when no useful delivery exists. `workflow runs result` treats the qualified
447
- delivery as ready while reporting `verified:false`.
479
+ An autonomous `complete` remains strictly verified. A planner `stop` still
480
+ delivers a completed outcome when a useful delivery exists: unresolved
481
+ verification concerns and the stopping reason ride along as `outcome.concerns`
482
+ and `outcome.reason`, attributes of that completed outcome rather than a
483
+ separate terminal status. `stop` produces `blocked` only when no useful
484
+ delivery exists. `workflow runs result` treats the completed outcome as ready
485
+ while reporting `verified:false`. The status value `completed_with_concerns`
486
+ still appears on some runs — including legacy ones recorded before this
487
+ framing — and every consumer reads it exactly like `completed`: a delivered
488
+ result with concerns to review, never a failure.
448
489
 
449
490
  The planner returns versioned JSON. It may propose `needs_more_work` with
450
491
  bounded `run`, inline-`fanout`, or `verify` actions. The deterministic runtime
@@ -0,0 +1,30 @@
1
+ #!/usr/bin/env node
2
+
3
+ import { readFileSync } from 'node:fs';
4
+ import { isValidOutputSchema, validateAgainstSchema } from '../src/workflow/schema.js';
5
+
6
+ function argValue(name) {
7
+ const index = process.argv.indexOf(name);
8
+ return index >= 0 ? process.argv[index + 1] : null;
9
+ }
10
+
11
+ const schemaPath = argValue('--schema');
12
+ const valuePath = argValue('--value');
13
+ if (!schemaPath || !valuePath) {
14
+ console.error('Usage: check-output-schema --schema <schema.json> --value <candidate.json>');
15
+ process.exit(2);
16
+ }
17
+
18
+ try {
19
+ const schema = JSON.parse(readFileSync(schemaPath, 'utf8'));
20
+ const value = JSON.parse(readFileSync(valuePath, 'utf8'));
21
+ const schemaValidity = isValidOutputSchema(schema);
22
+ const result = schemaValidity.ok
23
+ ? validateAgainstSchema(value, schema)
24
+ : { ok: false, errors: schemaValidity.issues };
25
+ process.stdout.write(`${JSON.stringify(result)}\n`);
26
+ process.exit(result.ok ? 0 : 1);
27
+ } catch (error) {
28
+ process.stdout.write(`${JSON.stringify({ ok: false, errors: [error.message] })}\n`);
29
+ process.exit(1);
30
+ }
@@ -3,7 +3,7 @@
3
3
  "bin": "opencode",
4
4
  "configDirs": ["~/.config/opencode"],
5
5
  "spawn": {
6
- "cmd": ["opencode", "run", "--auto", "--model", "kaihk/gpt-5.6-luna", "{taskFile}"],
6
+ "cmd": ["opencode", "run", "--auto", "{taskFile}"],
7
7
  "cwdMode": "pwd",
8
8
  "$comment-cwdMode": "QUIRK: resolves its project from $PWD, not the spawn cwd. The watcher MUST set env.PWD and spawn with cwd inside the target repo, or it will silently analyse the wrong repository and answer confidently about it."
9
9
  },
@@ -21,10 +21,12 @@
21
21
  { "match": { "path": "type", "equals": "text" }, "path": "part.text", "mode": "concat", "separator": "\n" }
22
22
  ]
23
23
  },
24
- "$comment-auto": "--auto is required for headless workflow dispatch: task files live under ~/.bullswarm, outside the target repo, and OpenCode otherwise pauses for an interactive permission approval. --model pins the QA/runtime pool to Luna instead of the CLI default.",
24
+ "$comment-auto": "--auto is required for headless workflow dispatch: task files live under ~/.bullswarm, outside the target repo, and OpenCode otherwise pauses for an interactive permission approval. The base connector intentionally leaves --model unset so a plain OpenCode installation uses its own default; opencode-kaihk.js adds --model <providerId>/gpt-5.6-luna only for discovered KaiHK providers, after which modelSelection can replace it for assignments and step locks.",
25
25
  "$comment-exit1": "known failure mode: writes a complete correct answer, then dies with a Console-sync auth error and exit 1. The verdict sets contentUsableDespiteExit instead of discarding the work.",
26
26
  "meter": { "type": "none" },
27
27
  "costRank": 1,
28
+ "preferredConcurrency": 1,
29
+ "$comment-preferredConcurrency": "OpenCode's shared local/provider session can terminate sibling headless runs together under parallel load. Prefer one in-flight OpenCode worker and route other ready siblings to healthy pools; if no alternative is eligible, availability wins and OpenCode may still be used.",
28
30
  "lanes": ["analyze", "build", "chore"],
29
31
  "capabilities": ["strong-analysis", "code-reading", "file-editing", "workflow-planning"],
30
32
  "modelDiscovery": { "cmd": ["opencode", "models"], "parse": "lines", "includePattern": "^[^\\s]+/[^\\s]+$", "timeoutMs": 20000, "maxModels": 250 },
@@ -13,6 +13,8 @@ Every statement is tagged:
13
13
  (`docs/experiments/2026-08-29-ultracode-vs-bullswarm.md`).
14
14
  - **[INFERRED]** — my reading of how the harness must behave to satisfy the
15
15
  spec. Not confirmed by source; treat as a hypothesis.
16
+ - **[IMPLEMENTED]** — behavior shipped in Bullswarm and backed by its source
17
+ and regression suite, rather than a claim about Claude's workflow contract.
16
18
 
17
19
  ## 0. The one-paragraph shape
18
20
 
@@ -349,7 +351,7 @@ author and the `Workflow` runtime.
349
351
  exist — the script's `while (!ok)` loop *is* the evidence — which is the
350
352
  general lesson: every piece of control flow bullswarm moves from planner
351
353
  into runtime needs its evidence rule moved with it.
352
- 11. **[SPEC] Schema-enforced worker output** — a planner `run` action or fan-out
354
+ 11. **[IMPLEMENTED] Schema-enforced worker output** — a planner `run` action or fan-out
353
355
  `stepTemplate` may declare an object-typed `outputSchema` subset. The
354
356
  runtime appends instructions for one trailing matching JSON object, with no
355
357
  prose or markdown fences after it, then parses and validates the object.
@@ -363,7 +365,10 @@ author and the `Workflow` runtime.
363
365
  fields, and `fanout.itemsFrom` can consume `outputs.<id>.data.items` without
364
366
  extraction when it is already an array. Planner decision validation rejects
365
367
  `outputSchema` on a proposed `verify` because verify has a fixed verdict
366
- shape.
368
+ shape. Schema-backed dispatches suppress ordinary same-pool retries so the
369
+ schema contract gets exactly its one bounded correction attempt. On resume,
370
+ a fan-out item is skipped only when both its verdict and declared schema
371
+ are satisfied; the schema must be declared on `stepTemplate.outputSchema`.
367
372
 
368
373
  **Honest limitation.** `itemsFrom` removes the planner *turn*, not the stage
369
374
  *barrier*: a verify depending on a data-driven fan-out waits for all items,
@@ -1,4 +1,4 @@
1
- # Bullswarm Dynamic Workflow Handoff
1
+ # Bullswarm Dynamic Workflow Handoff (Historical)
2
2
 
3
3
  **Purpose:** iteration brief for making bullswarm's workflow system behave like
4
4
  Claude Code's dynamic workflows while preserving bullswarm's provider routing,
@@ -6,8 +6,10 @@ quota pacing, content verification, and agent-friendly CLI contracts.
6
6
 
7
7
  **Audience:** the next implementation agent.
8
8
 
9
- **Status:** build target, based on repository inspection, real OpenCode/Luna QA,
10
- and a Fleetlens inspection of Claude workflow telemetry.
9
+ **Status:** historical implementation brief from 2026-08-21. The gaps and task
10
+ list below describe the state at that date; they are not a current capability
11
+ matrix. For current behavior use `README.md`, `skill/SKILL.md`,
12
+ `docs/claude-dynamic-workflow-mechanics.md`, and the contextual CLI help.
11
13
 
12
14
  ## Executive Summary
13
15
 
@@ -23,10 +25,12 @@ understand request
23
25
  -> repeat until complete
24
26
  ```
25
27
 
26
- Bullswarm currently has a validated JSON plan, sequential phases, dynamic
27
- fan-out, retries, escalation, verification, resume, and a basic dashboard. It
28
- does not yet have the central `observe -> decide -> schedule` loop. Its graph
29
- is fixed after validation; only a fan-out's item count can expand at runtime.
28
+ Bullswarm now has the control loop this brief proposed: a durable orchestrator
29
+ observes completed work, proposes a bounded program, deterministic validation
30
+ accepts or rejects it, and the runtime schedules ready actions before the next
31
+ checkpoint. Static JSON workflows remain supported alongside zero-graph
32
+ `workflow goal` execution. The rest of this document preserves the historical
33
+ evidence and build rationale that led to that implementation.
30
34
 
31
35
  The target is not an uncontrolled mutable DAG and not an LLM that owns the
32
36
  runtime. The target is a hybrid:
@@ -235,6 +239,8 @@ Supported step types:
235
239
  - `run`: one delegate invocation.
236
240
  - `fanout`: one delegate invocation per item.
237
241
  - `verify`: a skeptical review of a prior output artifact.
242
+ - `decide`: a durable adaptive planning gate whose proposal is validated before
243
+ any new action is appended or executed.
238
244
 
239
245
  The implementation is mainly in:
240
246
 
@@ -276,8 +282,19 @@ A later step can reference them with templates:
276
282
  {{outputs.previous.pool}}
277
283
  {{outputs.previous.outputText}}
278
284
  {{outputs.previous.outFile}}
285
+ {{outputs.previous.data.field}}
279
286
  ```
280
287
 
288
+ For a structured `run`, declare an object `outputSchema` when a later step
289
+ needs typed data. The successful state record contains `data` and
290
+ `schemaOk: true`; after the single schema correction retry fails it retains the
291
+ output text and contains `schemaOk: false` and `schemaErrors`. A schema retry is
292
+ observable as `action.output_schema_retry`, followed by
293
+ `action.output_validated` only when the corrected object passes validation.
294
+ The task also supplies an exact local schema-preflight command. The worker uses
295
+ it on a temporary candidate before replying, while the runtime independently
296
+ revalidates the captured response before exposing `data` downstream.
297
+
281
298
  A fan-out can use a prior output file as its item source:
282
299
 
283
300
  ```json
@@ -290,8 +307,13 @@ A fan-out can use a prior output file as its item source:
290
307
  }
291
308
  ```
292
309
 
293
- The runtime reads the referenced file and parses a JSON array. This is dynamic
294
- item expansion, not dynamic workflow graph expansion.
310
+ The runtime first consumes an already-recorded array such as
311
+ `outputs.discover.data.items`. For the legacy `outputs.<id>.outFile` form it
312
+ reads the referenced file and parses a JSON array. This is dynamic item
313
+ expansion, not dynamic workflow graph expansion. Fan-out schemas belong on
314
+ `stepTemplate.outputSchema`; each item stores its own `data`, `schemaOk`, and
315
+ possible `schemaErrors`, and resume re-runs only items whose verdict or schema
316
+ is incomplete.
295
317
 
296
318
  ### Routing and model selection
297
319
 
@@ -335,6 +357,7 @@ Bullswarm currently has:
335
357
  - Recursion-depth propagation and guard
336
358
  - Resume of successful steps
337
359
  - Fan-out resume by item fingerprint
360
+ - Structured worker output with one schema retry and durable schema state
338
361
  - Cooperative cancellation through `state.json`
339
362
  - Heartbeats during long dispatches
340
363
  - Basic interactive dashboard
@@ -346,6 +369,7 @@ Bullswarm currently has:
346
369
  bullswarm doctor --json
347
370
  bullswarm workflow capabilities --json
348
371
  bullswarm workflow inspect <file-or-name>
372
+ bullswarm workflow runs result <shortId> --json
349
373
  bullswarm workflow tui --json
350
374
  bullswarm workflow tui --json <shortId>
351
375
  bullswarm workflow tui --json --cancel <shortId>
@@ -360,7 +384,14 @@ bullswarm workflow inspect <file-or-name>
360
384
  bullswarm workflow tui --json
361
385
  ```
362
386
 
363
- ## 3. Important Gaps, Ordered by Priority
387
+ ## 3. Historical Gaps, Ordered by Priority
388
+
389
+ This section is an as-built checklist from the original handoff. The adaptive
390
+ decision loop, bounded graph expansion, structured decisions, attempt ledger,
391
+ ordered event log, cancellation, and timeline/dashboard surfaces described
392
+ below are implemented now. The past-tense gap text is retained so reviewers can
393
+ trace requirements to the resulting runtime and tests; it must not be read as
394
+ current product status.
364
395
 
365
396
  ### P0: no observe-plan-execute loop
366
397
 
@@ -94,6 +94,19 @@ the internal `review` artifact path. Bullswarm now deterministically infers that
94
94
  path for a verifier with one dependency and always appends the required JSON
95
95
  verdict contract. This is deliberately runtime knowledge, not caller steering.
96
96
 
97
+ ## Structured Output Evidence
98
+
99
+ For an optional object `outputSchema` on a `run`, the runtime appends the
100
+ contract to the worker prompt, parses the trailing object, and allows exactly
101
+ one schema correction retry. The durable `state.json` record exposes
102
+ `outputs.<id>.data` and `schemaOk`; a failed correction retains `schemaErrors`
103
+ and the output text. Fan-out item records use the same fields under
104
+ `outputs.<fanoutId>.items[]`. Offline regressions in
105
+ `tests/workflow-adaptive.test.js` and `tests/workflow-gaps.test.js` verify retry
106
+ events, persisted state, resume behavior, downstream rendering, and data-backed
107
+ fan-out. The stable result envelope returns the durable artifact content; typed
108
+ worker data remains inspectable in `state.outputs.<id>.data`.
109
+
97
110
  ## Acceptance interpretation
98
111
 
99
112
  - Provider output is accepted by content verification, never by exit status alone.
@@ -389,3 +389,165 @@ defect), every repair prompt says to edit only the reviewed work's files and nev
389
389
  requires one owner per file including an existing test the change breaks, and the validator line states run-wide id
390
390
  uniqueness. Direction from the user: "unless it is completely nonsense or unable to finish I don't see a reason to
391
391
  reject so easily". Claim to test on the next rerun: none of the three `8ebi8a` rejection reasons can produce ok:false.
392
+
393
+ ## Run `5cvj72` — bullswarm 0.20.0 builds bullswarm's next release (2026-08-31, real repo, installed runtime)
394
+
395
+ Goal: (R1) remove the hardcoded `kaihk/gpt-5.6-luna` from `connectors/opencode2.json` while keeping KaiHK per-provider
396
+ model injection, (R2) LLM-refined delegation classification on top of the deterministic guess, (R3) docs, (R4) full
397
+ suite green. Launched through `bullswarm delegate` itself (dry-run preview → executed pinned `--mode=workflow`).
398
+
399
+ Result: **13 min 58 s** (838 s), 1 planner turn (73 s, 8.7 %), 12 dispatches, parallelism 1.69, max 3 concurrent,
400
+ 9-action program (3 implementers ∥ → 3 unit verifies covering R1–R3 → suite verify covering R1–R4), auto-completed
401
+ `completed_with_concerns` (1 informational concern), 388/388 tests — verified independently by the operator.
402
+
403
+ Reliability events, all self-healed: `update-documentation` on claude-code/sonnet-5 failed the output substance gate
404
+ ("announcement without substance") → escalated to codex/gpt-5.6-terra, succeeded. `verify-classifier` rejected once —
405
+ legitimately under the lenient bar (R2 demanded BOTH reasons in the decision; `refineDecision()` dropped the
406
+ deterministic one; the focused tests themselves passed 18/18) — one repair round fixed it, re-verify accepted. The
407
+ coverage-evidence flip fired only alongside that real concern; no false rejection this run.
408
+
409
+ Routing: soft `preferredConcurrency:1` spillover worked as designed — 9/12 dispatches on `opencode2 kaihk-2/gpt-5.6-luna`
410
+ (the all-tier assignment), while parallel siblings spilled to claude-code (fable-5 for `implement-classifier`,
411
+ sonnet-5 for the failed docs attempt) and codex (terra). Note: the per-pool assignment pins the model only on its own
412
+ pool; spillover pools use their connector default, so a luna-only run needs either per-pool assignments or no
413
+ concurrent siblings.
414
+
415
+ Post-run verification by the operator: repo connector via a temp home — KaiHK off ⇒ `["opencode","run","--auto",
416
+ "{taskFile}"]` (no `--model`); KaiHK on ⇒ `--model kaihk/gpt-5.6-luna`, `kaihk-2/…`, `kaihk-3/…` injected per provider.
417
+ Live `--classify=llm` end-to-end: decision came back `source: llm-classifier` with both reasons and a sensible verdict.
418
+ Two observations recorded, not fixed: (a) `--dry-run` never consults the LLM, so the canonical skill flow
419
+ (dry-run preview → pinned re-execute) exercises only the deterministic classifier; (b) the substance gate rejected a
420
+ correct 16-char answer ("bullswarm 0.20.0") from the tiny smoke delegation — legitimately terse outputs still fail the
421
+ 40/80-char floors. Installed homes keep the old connector until a setup upgrade copies the new file.
422
+
423
+ ## Run p3jbha — always-LLM classification + cross-agent integration audit (2026-08-31)
424
+
425
+ Goal (4 numbered requirements): make the LLM the deciding classifier for every auto-mode
426
+ `delegate` call including `--dry-run` (deterministic stays as the pre-pass hint fed into the
427
+ LLM prompt); make LLM fallback visible in the envelope instead of silent; update every doc
428
+ that claimed dry-run was deterministic-only; write a read-only cross-agent integration audit
429
+ for Claude/Codex/Grok. Launched through the **repo binary at 8908ef8** (installed 0.20.0
430
+ predates the LLM classifier) via the canonical flow: `delegate --dry-run` preview → pinned
431
+ `--mode=workflow`. Before launch: `claude-fable-5` added to `strategy exclude-model`
432
+ (recorded no-Fable pref; last run's spillover violation), and three dangling
433
+ `bullswarm.broken-20260830` symlinks removed from all three CLIs' skill dirs.
434
+
435
+ **Preview accuracy specimen:** the deterministic classifier scored the goal correctly
436
+ (workflow, score 7) but the phrase "read-only cross-agent integration audit" tripped
437
+ `READ_ONLY_LABEL_RE`, so `hasMutationIntent` returned false and the suggested plan dropped
438
+ its Execute phase entirely (Inspect→Verify→Deliver for a mostly-code-edit task). A live
439
+ one-regex misfire — exactly the case for LLM-decided classification.
440
+
441
+ **Outcome: completed_with_concerns (verified: true), auto-completed.** Wall 1,292 s
442
+ (21m32s), attempt-busy 2,050 s, parallelism 1.59, max 3 concurrent, zero quota wait.
443
+ 14 dispatches: 2 planner turns (99 s total = 7.7%) + 12 worker attempts. Diff: 6 files
444
+ +72/−25 plus the new 411-line audit doc. Independently re-verified by the operator:
445
+ `npm test` 389/389; live `--dry-run` auto → `source: llm-classifier` in 23.6 s with the
446
+ deterministic sub-object preserved; `--dry-run --classify=deterministic` → 0.12 s, no
447
+ dispatch.
448
+
449
+ **Routing:** implement/verify work on opencode2 (kaihk-2/gpt-5.6-luna); docs on
450
+ claude-code/claude-sonnet-5; the audit on claude-code/claude-opus-5 (822 s, the critical
451
+ path). The only "fable" string in the run state is the exclusion entry itself — the
452
+ mitigation held.
453
+
454
+ **Reliability tally — 3 ok:false verdicts, 1 genuinely earned:**
455
+ 1. `verify-integration-audit` round 1: legitimate — the audit really ended with the
456
+ command appendix and lacked the required recommendations section; repair added §7
457
+ (five evidence-linked recommendations).
458
+ 2. `verify-documentation` round 1: cross-ownership overreach — it rejected R3's docs
459
+ because R4's `docs/integration-audit-2026-08-31.md` (owned by a *different, still
460
+ running* action, 822 s) did not exist yet. Under the runtime's lenient acceptance
461
+ standard, later-scheduled work and other actions' files are concerns, not rejections.
462
+ 3. `verify-integration-audit` re-verify after repair: **factually false** — it claimed
463
+ "no concrete recommendations appear at the end" and cited lines 384–411 as the final
464
+ section, while the repaired file (mtime 05:01:14Z, before the verifier started at
465
+ 05:01:50Z) held §7 Recommendations at lines 345–383. The verifier anchored on its
466
+ previous verdict instead of re-reading. This exhausted maxRounds=1, failed the action,
467
+ blocked `verify-suite`, and forced planner turn 2 — which proposed a fresh
468
+ `verify-completion` that passed all four requirements with line-level evidence, then
469
+ auto-completed. Recovery cost ≈ 2 min of wall time.
470
+
471
+ **Observations recorded, not yet fixed:** (a) re-verify verdict anchoring — the reverify
472
+ prompt could require the verifier to re-read the changed files and address the repair's
473
+ report before repeating a rejection; (b) requirement-coverage entanglement keeps making
474
+ verifiers judge files other actions own (second run in a row); (c) the run's own audit
475
+ deliverable found a real product defect: `integrate status` computes skill-link identity
476
+ against the *invoking checkout's* path, so any other valid Bullswarm install reports
477
+ `conflict`/exit 1 and `integrate install` refuses to repair it — stale skill symlinks are
478
+ sticky until removed by hand (see docs/integration-audit-2026-08-31.md §2, §7).
479
+
480
+ ## Run wxfwda — retire completed_with_concerns + two dashboard truthfulness fixes (2026-08-31)
481
+
482
+ User directive: "having concern is not a problem of bullswarm itself but part of the agent
483
+ lifecycle, I do not think we need to formalize it as a feature" — plus two screenshot
484
+ defects from viewing run p3jbha: permanent ✗ phase marks on a delivered run, and
485
+ "Final Verification · 1/1 complete" over a pane saying "Not started yet" for the
486
+ never-dispatched verify-suite. Launched through the repo binary at f11e544; the preview was
487
+ the **first live canonical-flow use of the always-LLM classifier** — `source:
488
+ llm-classifier` in 20.6 s, agreeing with the deterministic hint (workflow, score 7).
489
+
490
+ **Outcome: auto-completed, verified: true.** Wall 1,541 s (25m41s), busy 3,110 s,
491
+ **parallelism 2.02** (best of the series), max 4 concurrent, planner 80 s (5.2%),
492
+ 15 attempts, zero quota wait. Diff: 11 files +444/−67. Operator-verified: `npm test`
493
+ **394/394** (5 new tests), and both defects proven fixed by rendering the real
494
+ wf-mtgr56l1-167281 run dir with the new code — phase 6 renders `✓ Audit Verification 2/2`,
495
+ phase 7 renders `⊘ Final Verification 0/1` with pane "⊘ verify-suite · never dispatched ·
496
+ blocked by verify-integration-audit", planner panel "! Completed with 5 concerns" from the
497
+ outcome envelope.
498
+
499
+ **Design as landed:** new runs always terminate `completed` (or blocked/failed/…); the
500
+ qualification lives in `outcome` (`verified`, `bestEffort`, `concerns`, new
501
+ `qualification: 'verified'|'qualified'`). The stage `delivered_with_concerns` and event
502
+ `run.completed_with_concerns` are gone for new runs. `completed_with_concerns` remains
503
+ parse-only for legacy run dirs (the `budget_exhausted` precedent), rendered through the same
504
+ outcome-driven sentences. Dashboard invariants: a delivered run's phase list carries no
505
+ failure marks (attempt rows keep true history; failed/blocked/interrupted runs keep ✗);
506
+ a never-dispatched dependency-blocked action gets `⊘`, is excluded from "N/N complete",
507
+ and its pane names the failed dependency (`outputs[id].dependencyBlocked`, which the real
508
+ runner already writes).
509
+
510
+ **Reliability tally:** 1 rejection (verify-watcher-compatibility round 1) — cross-ownership
511
+ again: its concerns *praised* the reviewed work (10/10 focused tests, correctly refused to
512
+ touch unowned files) and complained about a missing `// legacy runs` comment in status.js,
513
+ a file the action did not own. Repaired in 36 s, re-verified ok. Third run in a row where
514
+ the only rejections judge files outside the reviewed work's ownership — the
515
+ requirement-coverage entanglement follow-up is now clearly the top reliability fix.
516
+ This run itself was labeled `completed_with_concerns` by the pre-change runtime it ran on —
517
+ expected artifact, not a failed fix; the label class it removes dies with the next release.
518
+
519
+ ## Run hdtdxs — phase-segmented timeline (2026-08-31)
520
+
521
+ User picked the segmented layout from a three-way mockup (headers on phase boundaries,
522
+ `· continued` on interleave, per-line `[Phase: …]` prefixes dropped, chronology preserved).
523
+ Launched through the repo binary at 5064c0f; preview `source: llm-classifier` in 16.6 s.
524
+ **First run executing on the post-removal runtime: it terminated plain `completed`
525
+ (qualification: qualified, 1 concern) — live validation of the wxfwda status collapse.**
526
+
527
+ **Outcome: auto-completed, verified: true.** Wall 2,290 s (38m10s), busy 3,038 s,
528
+ parallelism 1.33, max 2 concurrent, planner 214 s across 4 checkpoints + 1 correction
529
+ (9.3%), 23 attempts. Diff: 3 files +440/−49. Operator-verified: `npm test` **400/400**
530
+ (6 new tests) and `node --test tests/workflow-dashboard.test.js` 32/32 — which makes the
531
+ run's single recorded concern (a focused-test failure on the narrow-width continuation
532
+ header) stale by the time of handoff; the repair that closed verify-final-acceptance fixed
533
+ it. Rendering both real run dirs shows 0 `[Phase:` prefixes and correct headers:
534
+ `── Suite ─── 1m08s ──`, `── Verify · continued ─── 16m13s ──`,
535
+ `── Final Acceptance Verify ─── 6m58s ──`; scrolled viewports re-emit a continuation
536
+ header. Non-interactive `workflow tui <runId>` now prints the same segmented timeline
537
+ (added mid-run after a verifier rejected the render evidence as unverifiable).
538
+
539
+ **Reliability tally — the best verifier showing of the series: 5 rejections, ALL
540
+ legitimate.** verify-renderer r1 (3 stale tests genuinely failing under the acceptance
541
+ command), r2 (narrow variant genuinely still emitted `[Phase:` prefixes), verify-tests r1
542
+ (CHANGELOG entry genuinely missing — this one cascaded: 8 downstream actions
543
+ dependency-blocked across two recovery attempts before planner turn 4 landed the fix),
544
+ verify-final-changelog r1 (the goal's required documentation choice genuinely absent),
545
+ verify-final-acceptance r1 (render evidence genuinely invalid — non-interactive tui had no
546
+ timeline; the repair added it). Zero cross-ownership rejections for the first time in four
547
+ runs. Cost of the cascade: 2 extra planner turns + 1 planner correction (it tried to reuse
548
+ a finished phase name), ~8 min of wall. Routing clean: opencode2/luna everywhere except
549
+ add-tests on claude-code (856 s); the only "fable" string in state is the exclusion entry.
550
+
551
+ **Cosmetic observations for a later pass:** an empty `── Planner · continued ──` header can
552
+ render with no rows beneath it when its events fall outside the viewport; the top status
553
+ line now reads `· done` for a plain completed run (new wording from the status collapse).