bullswarm 0.27.1 → 0.28.1

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/CHANGELOG.md CHANGED
@@ -1,5 +1,246 @@
1
1
  # bullswarm changelog
2
2
 
3
+ ## 0.28.1 — summary bytes on the wire, monthly pacing
4
+
5
+ - `workflow runs result <id> --summary` was budgeted by `fitResultSummary`
6
+ against compact `JSON.stringify(summary)` (`RESULT_SUMMARY_BYTE_BUDGET =
7
+ 4096` in `src/workflow/v2-outcome.js`) but `jsonOut` printed
8
+ `JSON.stringify(obj, null, 2)`, so the bytes on the wire exceeded the
9
+ budget. `--summary` now prints compact single-line JSON
10
+ (`JSON.stringify(obj)` in `src/workflow/runs-cli.js`); `--json` without
11
+ `--summary` still pretty-prints the full envelope, and `--summary --json`
12
+ stays identical to `--summary`. Measured on
13
+ `tests/fixtures/real-result-ze5xz2.json` through the summariser / CLI:
14
+ compact `--summary` is 3,786 bytes
15
+ (`tests/workflow-result-summary.test.js` prints `result-summary size:
16
+ full=57141 summary=3786`; same figure as
17
+ `Buffer.byteLength(JSON.stringify(summarizeV2Result(fixture)))`); the
18
+ full envelope as `--json` prints it is 60,709 bytes
19
+ (`JSON.stringify(envelope, null, 2)` plus the trailing newline
20
+ `console.log` adds — `tests/workflow-result-summary.test.js` prints
21
+ `result-summary cli: prettyFull=60709`).
22
+
23
+ - Routing paced every pool by `windows.seven_day ?? windows.monthly`
24
+ (`paceSnapshot` in `src/meters/framework.js`), so command-code — whose real
25
+ budget is a monthly credit allocation — was paced by its weekly rate-limit
26
+ window. Live meter, captured 2026-09-09T10:45:28Z and evaluated at
27
+ 2026-09-09T11:09:44.982Z: weekly `used 73.1% elapsed 91.8% surplus +18.7`
28
+ against monthly `used 79.4% elapsed 75.3% surplus -4.1`, with 14.43 of 70
29
+ credits left for the 7.66 days to the 2026-09-17T03:06:55Z reset. Routing
30
+ therefore called it "the most-behind capable pool" and kept sending it work
31
+ while its monthly budget was already 4.1 points overspent. Pacing is now per
32
+ pool: `pacingWindowFor({connector, subscription})` resolves
33
+ `state.strategy.subscriptions[pool].quotaWindow`, then
34
+ `connector.subscription.quotaWindow`, normalised to `weekly` | `monthly` |
35
+ `null` (any other label — including a pre-0.28.1 free-text one — is ignored
36
+ for pacing and keeps the old weekly-first order). `paceSnapshot(snapshot,
37
+ nowMs, {pacingWindow})` takes `monthly` → `windows.monthly ??
38
+ windows.seven_day`, `weekly`/`null` → `windows.seven_day ??
39
+ windows.monthly`, and returns `pacingWindow` naming the window actually
40
+ used. `src/lib/config.js` `buildPools` resolves the choice once per pool
41
+ (where the connector and the state both are) and re-paces `usedPct`,
42
+ `elapsedPct`, `pace` and `paceResetsAt` off `reading.windows`, so cache,
43
+ stale and live readings are paced identically; the pool view carries
44
+ `pacingWindow`. The 5h gate is untouched: `command-code` still gates on 5h
45
+ `25.2%`. Connectors already declared this — `command-code` and the kaihk
46
+ pools `monthly`, `claude-code`/`codex`/`grok` `weekly`; nothing read it for
47
+ pacing before.
48
+ - The spend model follows the pacing window. `WINDOW_KEYS` (framework.js)
49
+ gains `monthly: {snapshot: 'monthly', history: 'monthly', windowMs: null}`,
50
+ and `rateForWindow` (`src/lib/spend.js`) derives the bootstrap window start
51
+ with `meta.windowMs ?? monthlyWindowMs(resetsAtMs)` — the calendar month
52
+ ending at the provider's `resets_at` (M2), never an assumed 30 days.
53
+ `attachSpend` now writes `pool.spend.monthly` and `pool.projectedMonthlyPct`
54
+ next to the fiveHour/weekly fields, plus `pool.spend.pacing = {window,
55
+ ratePerMinute, source, samples}` and `pool.projectedPacingPct` for the
56
+ window that paces the pool (default `weekly`, so a pool that declares
57
+ nothing keeps its old numbers). `inflightLoad` (`src/lib/route.js`) charges
58
+ the in-flight penalty from `spend.pacing?.ratePerMinute ??
59
+ spend.weekly?.ratePerMinute` with that rate's own source label, so the
60
+ surplus and the penalty are measured in the same window; candidate rows gain
61
+ `pacingWindow` and `projectedPacingPct` beside the unchanged
62
+ `projectedWeeklyPct`.
63
+ - Operator control and display. `bullswarm strategy set-subscription <pool>
64
+ --quota-window <weekly|monthly>` now selects the window that paces routing
65
+ (help text in `src/help.js`) and validates it: anything else exits 2 with
66
+ `--quota-window must be weekly or monthly (or unknown to clear)`, and
67
+ `unknown` clears the override back to the connector's declaration. Labels
68
+ already stored are ignored for pacing, never rejected on read. `bullswarm
69
+ pools` names the window in the meter column — `cmd-fixture cost=5
70
+ lanes=chore monthly used 79.4% elapsed 75.3% [cache] surplus=-4.1
71
+ inflight=0 5h=25.2% ready` — and `pools --json` entries carry
72
+ `pacingWindow`. `strategy refresh`/`show` print the window on each
73
+ subscription line (`command-code: GOAT · ... · monthly 79.4% used · surplus
74
+ -4.1`) and carry `pacingWindow` next to the free-text `quotaWindow` label in
75
+ `--json`; `strategy inventory --json` carries it per provider.
76
+ - Tests: 738 -> 750, 0 failures. The new behaviour is covered in
77
+ `tests/meters.test.js` (the live command-code snapshot as a fixture, the
78
+ helper's precedence, `buildPools` pacing), `tests/spend.test.js` (monthly
79
+ bootstrap window start, `spend.pacing`/`projectedPacingPct`),
80
+ `tests/route.test.js` (the penalty on the pacing window; weekly-only pools
81
+ unchanged), `tests/strategy-cli.test.js` (`--quota-window` validation) and
82
+ `tests/assignments.test.js` (the `pools` meter column and `--json`
83
+ `pacingWindow`).
84
+
85
+ ## 0.28.0 — context diet
86
+
87
+ - `workflow runs result <id> --summary` prints a compact status-loop
88
+ envelope, `schemaVersion: "bullswarm.workflow.result-summary.v1"`. It
89
+ carries `runId`, `shortId`, `status`, `verified`, `executionMode`,
90
+ `reason`, `finishedAt`, the goal's first line trimmed to 120 characters
91
+ plus `goalBytes`, each requirement as `{ id, status, mandatory,
92
+ evidenceCount, why }` (`why` is the latest evidence's first line trimmed
93
+ to 200 characters), each action as `{ id, kind, lane, effort, status,
94
+ pool, model, reasoning, wallSec, outFile, bytes }`, `concerns: { count,
95
+ first }` (up to three one-liners), the `usage` block, and `next: { full:
96
+ "bullswarm workflow runs result <id> --json", outputs: [<outFile paths>]
97
+ }`. `--summary` implies JSON with or without `--json`; it is not
98
+ TTY-dependent. The full `bullswarm.workflow.result.v2` envelope is
99
+ unchanged and stays the default. Read the full envelope on a failed or
100
+ partial run, or before judging evidence. Flag, help, and example:
101
+ `src/lib/cli-flags.js`, `src/help.js` (`Usage: bullswarm workflow runs
102
+ result <shortId|runId> [--json] [--summary]`; `--summary` "print the
103
+ compact JSON status-loop envelope; implies --json"). A terminal `workflow
104
+ watch` prints `next: bullswarm workflow runs result <shortId> --json
105
+ --summary`.
106
+
107
+ - Byte accounting on every action attempt. `state.attempts[].bytes` is
108
+ `{ taskFile, authorPrompt, kernel, dependencyInputs, output }` — the task
109
+ file the kernel wrote, the action's own prompt text as authored, the
110
+ remainder after subtracting that prompt and any embedded requirement
111
+ text, the sum of the dependency output files the task points at (0 when
112
+ there are none; `digestOf` drill-down paths are pointers, not inputs),
113
+ and the durable out file on completion. The result envelope copies the
114
+ last attempt's `bytes` onto `actions[]` and totals `usage.bytes: {
115
+ taskFiles, dependencyInputs, outputs }`. `workflow runs show` appends
116
+ `in <taskFile>/<dependencyInputs> out <output>` per attempt (the unit
117
+ fixture prints `in 3.1K/60.8K out 14.8K`); missing values stay blank /
118
+ null, never guessed. These are UTF-8 byte counts, never tokens.
119
+
120
+ - `kind: "digest"` (analyze/low) is an extractive condensation of its
121
+ dependencies' outputs so an expensive consumer reads one artifact
122
+ instead of many raw out-files. The kernel writes the whole task: quote
123
+ verbatim (never paraphrase or judge) each source's delivered items,
124
+ validation numbers, commands and their output, unfinished work, and
125
+ every shared-file or integrator request; one section per source headed
126
+ by its absolute output path; no verdicts, no recommendations, no new
127
+ claims; target at most a quarter of the input bytes or 8 KB, whichever
128
+ is larger. The author's prompt is focus guidance only. Validation
129
+ (exit 2 with the reason at `workflow plan validate` and `workflow goal
130
+ --program`): a digest must depend on at least one action, must have
131
+ empty `evidenceFor`, owns no files, needs no `affects`. No evidence
132
+ action may list a digest in `dependsOn` — evidence reads the real
133
+ artifacts. Consumers that depend on a digest receive, in their
134
+ dependency artifacts, the digest entry plus `digestOf: [{ actionId,
135
+ outputFile }]` for each digested source. Use one when three or more
136
+ writers feed a single integrator, or when a consumer's dependency
137
+ outputs would exceed roughly 20 KB; never for evidence.
138
+
139
+ - The planning contract states the kind. `workflow plan contract --json`
140
+ (`Usage: bullswarm workflow plan contract "<goal>" [--cwd <dir>]
141
+ [--json]`) returns 16 rules; `rules[5]` lists `digest=analyze/low` in
142
+ the kind table and `rules[6]` is the extractive digest rule (when to
143
+ insert one, `digestOf`, empty `evidenceFor` / `ownedFiles`, evidence
144
+ must not depend on a digest).
145
+
146
+ - The packaged skill's status loop (`skill/SKILL.md`,
147
+ `skill/references/operations.md`) now recommends `bullswarm workflow
148
+ runs result <shortId> --json --summary`. Read the full envelope with
149
+ `--json` alone when the run is failed or partial, or before judging
150
+ evidence.
151
+
152
+ - Measured numbers, each with its command. On the real 0.27.1 build run
153
+ `ze5xz2` (files under `.diet-inputs/`): `workflow runs result ze5xz2
154
+ --json` is 60,709 bytes on disk (`wc -c .diet-inputs/real-result-ze5xz2.json`;
155
+ the same file is `tests/fixtures/real-result-ze5xz2.json`). The 0.28.0
156
+ goal recorded that envelope's `requirements` as 39,231 bytes and `goal`
157
+ as 11,009 bytes. Re-measuring the same file: `Buffer.byteLength(goal)` =
158
+ 10,996 (the summary's `goalBytes`) and `JSON.stringify(requirements)` =
159
+ 39,288; compact `JSON.stringify` of the parsed envelope is 57,141.
160
+ `summarizeV2Result` of that fixture is 3,786 bytes —
161
+ `tests/workflow-result-summary.test.js` prints `result-summary size:
162
+ full=57141 summary=3786` (numbers re-measured in 0.28.1). The 0.28.0
163
+ goal recorded the integrator's inputs as 60,790 bytes; `wc -c` of the
164
+ seven dependency out-files under `.diet-inputs/` sums to 46,022
165
+ (out-surface 18,659, out-routing-cleanup
166
+ 10,202, out-state-bugs 7,052, out-docs 6,831, out-dead-kernel 1,377,
167
+ out-verify-gate 974, out-dead-code 927) and the integrator task file is
168
+ 14,768 (`wc -c .diet-inputs/task-integrate-attempt-1.md`), which
169
+ together are 60,790.
170
+
171
+ - Fixture measurement, the same goal run twice under a temporary home
172
+ (`tests/workflow-context-diet-measurement.test.js`, which prints both
173
+ lines below). Three writers each padded to a few KB feed one integrator
174
+ directly, then the same three feed a `kind: "digest"` that feeds the
175
+ integrator: the integrator's `bytes.dependencyInputs` falls from
176
+ **12,477 bytes to 63 bytes** (`context diet: integrator dependencyInputs
177
+ without digest=12477 with digest=63`). The 63 is a floor, not a
178
+ condensation ratio — the deterministic fixture worker answers with a
179
+ fixed stub instead of really condensing. The ceiling is the byte target
180
+ the kernel writes into that digest's own task, **8,192 bytes** for this
181
+ input, and the test asserts the saving holds at that ceiling too
182
+ (8,192 < 12,477). The same run's envelopes print as `context diet: run B
183
+ envelope full=<n> summary=<n>` (the full envelope embeds the temporary
184
+ home's absolute paths, so its size varies a few bytes between runs; the
185
+ summary carries one `next.runDir` string plus basenames, so it does not);
186
+ the deterministic envelope comparison to quote is the real-run fixture
187
+ above (57,141 full vs the measured summary).
188
+
189
+ - `workflow capabilities` now reports the closed kind list at
190
+ `engines.autonomousV2.actionKinds`, cloned from the validator's
191
+ `KIND_DEFAULTS` rather than hand-listed, so `digest` and every future
192
+ kind are discoverable by a probing agent
193
+ (`src/workflow/cli.js`). The planning contract's
194
+ `program.actionFields.kind` description is derived from the same table
195
+ (`src/workflow/v2-planner.js`).
196
+
197
+ - `TIER_LANES` (`src/lib/strategy.js`) excludes kernel-owned kinds from the
198
+ effort-tier count, so the map stays `{ high: analyze, medium: build,
199
+ low: chore }`. A digest is a mechanism the kernel writes, not a nature of
200
+ work that should define which lane a tier routes to; counting it would
201
+ have flipped low from `chore` to `analyze` on the strength of an action
202
+ no planner has to reason about. No shipped connector's routing changes
203
+ either way — all six declare all three lanes.
204
+
205
+ - The three KaiHK-backed OpenCode pools can now run gpt-5.6-luna at a
206
+ chosen reasoning level. `connectors/opencode2.json` declares
207
+ `reasoning: { flag: "--variant", levels: [low, medium, high, xhigh,
208
+ max], defaults: { high: high, medium: medium, low: low } }` — the same
209
+ five levels as `connectors/command-code.json`, which fronts the same
210
+ backend — so rungs for these pools stop printing `— (unsupported)`.
211
+ opencode only forwards a `--variant` its config declares for that
212
+ model, so the flag alone would be silently dropped;
213
+ `expandOpenCodeKaihkConnectors` (`src/lib/opencode-kaihk.js`) therefore
214
+ sets `env.OPENCODE_CONFIG_CONTENT` on the base pool and on every clone
215
+ to the variants for that pool's OWN provider id, via the new pure
216
+ helper `kaihkVariantsConfig(providerId, model = KAIHK_OPENCODE_MODEL)`:
217
+ `{"provider":{"kaihk-2":{"models":{"gpt-5.6-luna":{"variants":{"low":{"reasoningEffort":"low"},"medium":{"reasoningEffort":"medium"},"high":{"reasoningEffort":"high"},"xhigh":{"reasoningEffort":"xhigh"},"max":{"reasoningEffort":"max"}}}}}}}`.
218
+ opencode merges that JSON string over the config file, so the API key
219
+ and everything else in `~/.config/opencode/opencode.json` stays in
220
+ force. An `OPENCODE_CONFIG_CONTENT` the operator set by hand in the
221
+ installed connector is never overwritten, on the base pool or on the
222
+ clones. A medium/max dispatch on `opencode2:kaihk-2` composes
223
+ `opencode run --auto --model kaihk-2/gpt-5.6-luna <taskFile> --variant
224
+ max --format json`; a level already pinned in the template is replaced,
225
+ not duplicated. Per pool:
226
+ `bullswarm strategy set-rung opencode2 medium --model kaihk/gpt-5.6-luna --reasoning max`,
227
+ `bullswarm strategy set-rung opencode2:kaihk-2 medium --model kaihk-2/gpt-5.6-luna --reasoning max`,
228
+ `bullswarm strategy set-rung opencode2:kaihk-3 medium --model kaihk-3/gpt-5.6-luna --reasoning max`.
229
+ Existing installations pick the block up through
230
+ `upgradeConnectorMetadata` (`src/setup.js`), which backfills a missing
231
+ `reasoning` block and leaves a customised one alone. Non-KaiHK opencode
232
+ installations get no injected variants, so `--variant` is a no-op there
233
+ rather than an error — recorded in the connector's
234
+ `$comment-reasoning`.
235
+
236
+ - Tests: 713 -> 738, 0 failures
237
+ (`env -u CLAUDE_CONFIG_DIR -u FORCE_COLOR -u NO_COLOR npm test`). Four
238
+ new files carry the new behaviour: `tests/workflow-bytes.test.js`,
239
+ `tests/workflow-digest.test.js`,
240
+ `tests/workflow-result-summary.test.js` (with the
241
+ `tests/fixtures/real-result-ze5xz2.json` envelope it measures) and
242
+ `tests/workflow-context-diet-measurement.test.js`.
243
+
3
244
  ## 0.27.1 — audit cleanup
4
245
 
5
246
  - Deleted the remaining dead symbols the 2026-09-09 audit listed as Tier A:
package/README.md CHANGED
@@ -51,7 +51,12 @@ detaches safely.
51
51
  passing verification.
52
52
  2. **Pace by meter.** The scheduling resource is the subscription window:
53
53
  elapsed% minus used%, most-behind pool wins. Pace may only promote a
54
- *cheaper* pool. Lanes are work-nature, never hard-coded to pools. The
54
+ *cheaper* pool. Lanes are work-nature, never hard-coded to pools. Which
55
+ window paces one pool is the subscription window that pool's connector
56
+ declares (`quotaWindow`: weekly for claude-code, codex and grok; monthly
57
+ for command-code and the kaihk pools), overridable per pool with
58
+ `bullswarm strategy set-subscription <pool> --quota-window <weekly|monthly>`
59
+ — `bullswarm pools` names it in the meter column. The
55
60
  5-hour window never paces — it gates: a pool at or above 75% of it is
56
61
  chosen only when no eligible pool below that line exists, and one at or
57
62
  above 90% is not dispatched at all.
@@ -180,6 +185,25 @@ per pool and tier, so the tier moves off whichever model held it while that
180
185
  model keeps its other tiers. Nothing about `state.json` changed shape: rungs are
181
186
  a view over `strategy.modelTiers` and `strategy.reasoning`.
182
187
 
188
+ The KaiHK-backed OpenCode pools (`opencode2`, `opencode2:kaihk-2`,
189
+ `opencode2:kaihk-3`) express reasoning as opencode's `--variant <level>`, at the
190
+ same five levels as `command-code`. opencode only forwards a variant its own
191
+ config declares for that model, so bullswarm injects them: each pool is spawned
192
+ with `OPENCODE_CONFIG_CONTENT` declaring `low`/`medium`/`high`/`xhigh`/`max` as
193
+ `reasoningEffort` variants of `<providerId>/gpt-5.6-luna`, merged over your
194
+ `~/.config/opencode/opencode.json` (your API keys stay in force). Without that
195
+ injection opencode accepts `--variant` and silently drops it. Set a rung per
196
+ pool, using that pool's own provider prefix:
197
+
198
+ ```bash
199
+ bullswarm strategy set-rung opencode2:kaihk-2 medium \
200
+ --model kaihk-2/gpt-5.6-luna --reasoning max
201
+ ```
202
+
203
+ If you set `OPENCODE_CONFIG_CONTENT` yourself in
204
+ `~/.bullswarm/connectors/opencode2.json`, bullswarm leaves it alone and injects
205
+ nothing — you own the variants from then on.
206
+
183
207
  The benchmark evidence comes from Epoch AI's benchmarking hub, used under
184
208
  CC BY 4.0: Epoch AI, 'AI Benchmarking Hub'. Published online at epoch.ai.
185
209
  Retrieved from <https://epoch.ai/benchmarks>. `blended` is the mean of the
@@ -390,12 +414,25 @@ names that nature once and derives them:
390
414
  | --- | --- | --- |
391
415
  | `mechanical` | chore | low |
392
416
  | `io-read` | analyze | low |
417
+ | `digest` | analyze | low |
393
418
  | `check` | analyze | medium |
394
419
  | `implement` | build | medium |
395
420
  | `integration` | build | high |
396
421
  | `architecture` | analyze | high |
397
422
  | `adversarial-acceptance` | analyze | high |
398
423
 
424
+ `digest` is the one kind whose instructions the kernel supplies in full — your
425
+ prompt for it is focus guidance only. It condenses the
426
+ outputs of the actions it depends on — quoting each source's delivered items,
427
+ validation numbers, commands, unfinished work, and requests verbatim, one
428
+ section per source, with no verdicts of its own — so an expensive consumer
429
+ reads one artifact instead of many raw output files, and the digest entry in
430
+ that consumer's dependency artifacts still names every digested source for
431
+ drill-down. Use one when three or more writers feed a single integrator, or
432
+ when a consumer's dependency outputs would exceed roughly 20 KB. A digest must
433
+ depend on at least one action, owns no files, needs no `affects`, and no
434
+ evidence action may depend on one: evidence reads the real artifacts.
435
+
399
436
  Resolution is per field: an explicit `lane` or `effort` on the action wins,
400
437
  then the kind table, then an optional program-level `defaults` object — which
401
438
  may set only `effort` and `reasoning`, because lane follows the individual
@@ -439,6 +476,7 @@ bullswarm workflow runs show <shortId>
439
476
  bullswarm workflow watch <shortId> # V2: attach, then one line per notable event
440
477
  bullswarm workflow watch <shortId> --next # print the next notable event and exit
441
478
  # relaunch with the --after/--since it prints
479
+ bullswarm workflow runs result <shortId> --json --summary # compact status-loop envelope once terminal
442
480
  bullswarm workflow # unified human workflow home
443
481
  bullswarm workflow tui <shortId> # jump directly to one run timeline
444
482
  bullswarm workflow tui --json <shortId>
@@ -574,7 +612,8 @@ bullswarm workflow runs --historical --since yesterday --until today
574
612
  bullswarm workflow runs --all --from 2026-08-20 --to 2026-08-27
575
613
  bullswarm workflow runs --limit 20 # cap the result count
576
614
  bullswarm workflow runs show <shortId> # state + report + summary
577
- bullswarm workflow runs result <shortId> --json # stable result for the calling agent
615
+ bullswarm workflow runs result <shortId> --json --summary # compact status-loop envelope
616
+ bullswarm workflow runs result <shortId> --json # full envelope (failed/partial, or before judging evidence)
578
617
  bullswarm runs show <shortId> # top-level shorthand
579
618
  bullswarm workflow runs delete <shortId> --yes # remove the run dir
580
619
  ```
@@ -591,8 +630,10 @@ Values accept ISO timestamps, local `YYYY-MM-DD` dates, `today`, `yesterday`,
591
630
  `tomorrow`, `now`, or relative durations such as `30m`, `24h`, `7d`, and `2w`.
592
631
 
593
632
  After a workflow reaches a terminal state, agents should consume
594
- `workflow runs result <id> --json` instead of probing `state.json`, task files,
595
- or provider-specific output. Autonomous V2 returns the versioned
633
+ `workflow runs result <id> --json --summary` for the status loop instead of
634
+ probing `state.json`, task files, or provider-specific output. Read the full
635
+ envelope with `--json` alone when the run is failed or partial, or before
636
+ judging evidence. Autonomous V2's full document is the versioned
596
637
  `bullswarm.workflow.result.v2` envelope with kernel-computed status, fresh
597
638
  requirement evidence, per-action status/failure/output files, explicit gaps,
598
639
  usage, and verification qualification. New programs include `executionMode:
@@ -608,6 +649,75 @@ otherwise the command returns after printing this handoff.
608
649
  Time filters preserve the existing scope, so use `--all` or `--historical` when
609
650
  auditing completed runs.
610
651
 
652
+ ### Context diet
653
+
654
+ The kernel now measures — and can shrink — what it puts in front of a model.
655
+ These are UTF-8 byte counts, never tokens.
656
+
657
+ **Status loop.** Poll with `--summary`; it implies JSON (with or without
658
+ `--json`) and prints `schemaVersion: "bullswarm.workflow.result-summary.v1"`:
659
+ `runId`, `shortId`, `status`, `verified`, `executionMode`, `reason`,
660
+ `finishedAt`, the goal's first line (120 characters) plus `goalBytes`, each
661
+ requirement as `{ id, status, mandatory, evidenceCount, why }`, each action as
662
+ `{ id, kind, lane, effort, status, pool, model, reasoning, wallSec, outFile,
663
+ bytes }`, `concerns: { count, first }`, `usage`, and `next: { full, runDir, outputs }` — every output name is a basename inside `next.runDir`.
664
+ `--summary` is single-line JSON (`JSON.stringify`), so the bytes on the wire
665
+ match the 4,096-byte fitter budget. As printed by the CLI on
666
+ `tests/fixtures/real-result-ze5xz2.json`, the compact summary is 3,786 bytes
667
+ and the pretty full envelope (`--json` alone) is 60,709 bytes.
668
+ The full `bullswarm.workflow.result.v2` envelope is unchanged and remains the
669
+ default. Read it (`--json` alone) on a failed or partial run, or before judging
670
+ evidence. A terminal `workflow watch` prints the same compact command as
671
+ `next:`.
672
+
673
+ ```bash
674
+ bullswarm workflow runs result <shortId> --json --summary
675
+ bullswarm workflow runs result <shortId> --json
676
+ ```
677
+
678
+ `workflow runs result --help` states `Usage: bullswarm workflow runs result
679
+ <shortId|runId> [--json] [--summary]`; `--summary` is "print the compact JSON
680
+ status-loop envelope; implies --json".
681
+
682
+ **Bytes.** Every attempt records `bytes: { taskFile, authorPrompt, kernel,
683
+ dependencyInputs, output }` — the task file the kernel wrote, the action's own
684
+ prompt as authored, the remainder after subtracting that prompt and any
685
+ embedded requirement text, the sum of the dependency output files the task
686
+ points at (0 when there are none), and the durable out file on completion.
687
+ The result envelope copies the last attempt's `bytes` onto `actions[]` and
688
+ totals `usage.bytes: { taskFiles, dependencyInputs, outputs }`.
689
+ `workflow runs show` appends `in <taskFile>/<dependencyInputs> out <output>`
690
+ per attempt (blank when unrecorded). Missing values are null, never guessed.
691
+
692
+ **Digest.** `kind: "digest"` is analyze/low. It is an extractive condensation
693
+ of its dependencies' outputs — quoted delivered items, validation numbers,
694
+ commands, unfinished work, and integrator requests; no verdicts of its own.
695
+ The kernel writes the whole task; the author's prompt is focus guidance only.
696
+ Use one when three or more writers feed a single integrator, or when a
697
+ consumer's dependency outputs would exceed roughly 20 KB. A digest must
698
+ depend on at least one action, owns no files, has empty `evidenceFor`, and
699
+ needs no `affects`. Evidence must not depend on a digest: evidence reads the
700
+ real artifacts. Consumers that depend on a digest receive that digest plus a
701
+ `digestOf` array of `{ actionId, outputFile }` so they can drill down; those
702
+ paths are pointers, not extra `dependencyInputs`. (The kind table above
703
+ derives lane and effort.)
704
+
705
+ ```json
706
+ {
707
+ "schemaVersion": "bullswarm.workflow.program.v2",
708
+ "actions": [
709
+ { "id": "write-a", "kind": "implement", "dependsOn": [], "ownedFiles": ["a.ts"], "affects": ["requirement-1"], "evidenceFor": [], "purpose": "Write slice A", "prompt": "Implement A and report the checks you ran." },
710
+ { "id": "write-b", "kind": "implement", "dependsOn": [], "ownedFiles": ["b.ts"], "affects": ["requirement-2"], "evidenceFor": [], "purpose": "Write slice B", "prompt": "Implement B and report the checks you ran." },
711
+ { "id": "write-c", "kind": "implement", "dependsOn": [], "ownedFiles": ["c.ts"], "affects": ["requirement-3"], "evidenceFor": [], "purpose": "Write slice C", "prompt": "Implement C and report the checks you ran." },
712
+ { "id": "condense", "kind": "digest", "dependsOn": ["write-a", "write-b", "write-c"], "ownedFiles": [], "affects": [], "evidenceFor": [], "purpose": "Condense the writer outputs", "prompt": "Keep every acceptance number and every shared-file request." },
713
+ { "id": "integrate", "kind": "integration", "dependsOn": ["condense"], "ownedFiles": [], "affects": ["requirement-1", "requirement-2", "requirement-3"], "evidenceFor": [], "purpose": "Integrate and run the gates", "prompt": "Apply every request the digest carries and run the repository gates." }
714
+ ]
715
+ }
716
+ ```
717
+
718
+ An evidence action for those requirements depends on `write-a`, `write-b`, and
719
+ `write-c` — never on `condense`.
720
+
611
721
  ### Live workflow dashboard
612
722
 
613
723
  For ordinary observation, use the non-interactive watcher. For V2 runs it
@@ -615,8 +725,9 @@ prints one attach line, then one line per notable event as it happens
615
725
  (action finished/failed/blocked/cancelled, evidence, stage completion,
616
726
  planner turn, stall/recovery, cancellation, and the existing pause and
617
727
  terminal `outcome:` / `next:` lines) and stays silent while work is merely
618
- in progress. Agent starts, mechanical retries, and steering delivery print
619
- only with `--verbose`. A usage-limit failure (`failureKind: 'quota'`) always
728
+ in progress. A terminal watch's `next:` line is
729
+ `bullswarm workflow runs result <shortId> --json --summary`. Agent starts,
730
+ mechanical retries, and steering delivery print only with `--verbose`. A usage-limit failure (`failureKind: 'quota'`) always
620
731
  prints, verbose or not: `⚠ <actionId> usage limit on <pool> · paused until
621
732
  <deadline> · retrying on another pool`, followed once the mechanical retry
622
733
  lands on another pool by `↺ <actionId> now on <pool> · <model>`. The
@@ -34,6 +34,12 @@
34
34
  "capabilities": ["strong-analysis", "code-reading", "file-editing", "workflow-planning"],
35
35
  "modelDiscovery": { "cmd": ["opencode", "models"], "parse": "lines", "includePattern": "^[^\\s]+/[^\\s]+$", "timeoutMs": 20000, "maxModels": 250 },
36
36
  "modelSelection": { "flag": "--model", "mode": "replace-or-append" },
37
+ "$comment-reasoning": "verified 2026-09-09 against opencode 1.18.25: `opencode run --help` documents `--variant model variant (provider-specific reasoning effort, e.g., high, max, minimal)`. opencode only forwards a variant its CONFIG declares for that model, and the owner's opencode.json declares none, so on a bare install `--variant` is accepted and silently dropped. The level reaches the KaiHK API as `reasoning_effort` ONLY because expandOpenCodeKaihkConnectors (src/lib/opencode-kaihk.js) injects the five variants for <providerId>/gpt-5.6-luna through env.OPENCODE_CONFIG_CONTENT on every discovered KaiHK pool. Probed with that variable set: `--variant bogus` (reasoningEffort bogus-level) failed at the API with `level \"bogus-level\" not supported, valid levels: low, medium, high, xhigh, max`, and `--variant max` succeeded; probed again with the flag AFTER the positional task text, exactly where bullswarm appends it (`opencode run --auto --model <id>/gpt-5.6-luna <task> --variant <level> --format json`), with the same API rejection for the bogus level, so the position is proven too. Same five levels as connectors/command-code.json, which fronts the same backend. UNVERIFIED for non-KaiHK opencode providers: they get no injected variants, so the flag is a no-op there rather than an error.",
38
+ "reasoning": {
39
+ "flag": "--variant",
40
+ "levels": ["low", "medium", "high", "xhigh", "max"],
41
+ "defaults": { "high": "high", "medium": "medium", "low": "low" }
42
+ },
37
43
  "modelProfiles": [
38
44
  { "match": "(?:^|/)claude-fable-", "tier": "high", "qualityRank": 6, "autoRecommend": false },
39
45
  { "match": "gpt-5\\.6-sol$", "tier": "high", "qualityRank": 6, "autoRecommend": true },