bullswarm 0.27.1 → 0.28.1
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +241 -0
- package/README.md +117 -6
- package/connectors/opencode2.json +6 -0
- package/data/openrouter-benchmarks.json +9933 -9854
- package/docs/studies/portal-token-diet.md +10 -0
- package/package.json +1 -1
- package/skill/SKILL.md +15 -3
- package/skill/references/operations.md +32 -3
- package/src/cli.js +6 -1
- package/src/help.js +15 -8
- package/src/lib/cli-flags.js +1 -1
- package/src/lib/config.js +44 -6
- package/src/lib/forecast.js +3 -2
- package/src/lib/opencode-kaihk.js +46 -0
- package/src/lib/route.js +36 -19
- package/src/lib/spend.js +40 -20
- package/src/lib/strategy.js +16 -1
- package/src/meters/framework.js +82 -6
- package/src/strategy-cli.js +33 -3
- package/src/workflow/action-validator.js +32 -3
- package/src/workflow/cli.js +4 -1
- package/src/workflow/dashboard.js +2 -2
- package/src/workflow/runs-cli.js +35 -5
- package/src/workflow/v2-dispatch.js +1 -1
- package/src/workflow/v2-outcome.js +224 -3
- package/src/workflow/v2-planner.js +11 -1
- package/src/workflow/v2-runtime.js +130 -5
- package/src/workflow/v2-state.js +27 -2
- package/src/workflow/watch-cli.js +1 -1
package/CHANGELOG.md
CHANGED
|
@@ -1,5 +1,246 @@
|
|
|
1
1
|
# bullswarm changelog
|
|
2
2
|
|
|
3
|
+
## 0.28.1 — summary bytes on the wire, monthly pacing
|
|
4
|
+
|
|
5
|
+
- `workflow runs result <id> --summary` was budgeted by `fitResultSummary`
|
|
6
|
+
against compact `JSON.stringify(summary)` (`RESULT_SUMMARY_BYTE_BUDGET =
|
|
7
|
+
4096` in `src/workflow/v2-outcome.js`) but `jsonOut` printed
|
|
8
|
+
`JSON.stringify(obj, null, 2)`, so the bytes on the wire exceeded the
|
|
9
|
+
budget. `--summary` now prints compact single-line JSON
|
|
10
|
+
(`JSON.stringify(obj)` in `src/workflow/runs-cli.js`); `--json` without
|
|
11
|
+
`--summary` still pretty-prints the full envelope, and `--summary --json`
|
|
12
|
+
stays identical to `--summary`. Measured on
|
|
13
|
+
`tests/fixtures/real-result-ze5xz2.json` through the summariser / CLI:
|
|
14
|
+
compact `--summary` is 3,786 bytes
|
|
15
|
+
(`tests/workflow-result-summary.test.js` prints `result-summary size:
|
|
16
|
+
full=57141 summary=3786`; same figure as
|
|
17
|
+
`Buffer.byteLength(JSON.stringify(summarizeV2Result(fixture)))`); the
|
|
18
|
+
full envelope as `--json` prints it is 60,709 bytes
|
|
19
|
+
(`JSON.stringify(envelope, null, 2)` plus the trailing newline
|
|
20
|
+
`console.log` adds — `tests/workflow-result-summary.test.js` prints
|
|
21
|
+
`result-summary cli: prettyFull=60709`).
|
|
22
|
+
|
|
23
|
+
- Routing paced every pool by `windows.seven_day ?? windows.monthly`
|
|
24
|
+
(`paceSnapshot` in `src/meters/framework.js`), so command-code — whose real
|
|
25
|
+
budget is a monthly credit allocation — was paced by its weekly rate-limit
|
|
26
|
+
window. Live meter, captured 2026-09-09T10:45:28Z and evaluated at
|
|
27
|
+
2026-09-09T11:09:44.982Z: weekly `used 73.1% elapsed 91.8% surplus +18.7`
|
|
28
|
+
against monthly `used 79.4% elapsed 75.3% surplus -4.1`, with 14.43 of 70
|
|
29
|
+
credits left for the 7.66 days to the 2026-09-17T03:06:55Z reset. Routing
|
|
30
|
+
therefore called it "the most-behind capable pool" and kept sending it work
|
|
31
|
+
while its monthly budget was already 4.1 points overspent. Pacing is now per
|
|
32
|
+
pool: `pacingWindowFor({connector, subscription})` resolves
|
|
33
|
+
`state.strategy.subscriptions[pool].quotaWindow`, then
|
|
34
|
+
`connector.subscription.quotaWindow`, normalised to `weekly` | `monthly` |
|
|
35
|
+
`null` (any other label — including a pre-0.28.1 free-text one — is ignored
|
|
36
|
+
for pacing and keeps the old weekly-first order). `paceSnapshot(snapshot,
|
|
37
|
+
nowMs, {pacingWindow})` takes `monthly` → `windows.monthly ??
|
|
38
|
+
windows.seven_day`, `weekly`/`null` → `windows.seven_day ??
|
|
39
|
+
windows.monthly`, and returns `pacingWindow` naming the window actually
|
|
40
|
+
used. `src/lib/config.js` `buildPools` resolves the choice once per pool
|
|
41
|
+
(where the connector and the state both are) and re-paces `usedPct`,
|
|
42
|
+
`elapsedPct`, `pace` and `paceResetsAt` off `reading.windows`, so cache,
|
|
43
|
+
stale and live readings are paced identically; the pool view carries
|
|
44
|
+
`pacingWindow`. The 5h gate is untouched: `command-code` still gates on 5h
|
|
45
|
+
`25.2%`. Connectors already declared this — `command-code` and the kaihk
|
|
46
|
+
pools `monthly`, `claude-code`/`codex`/`grok` `weekly`; nothing read it for
|
|
47
|
+
pacing before.
|
|
48
|
+
- The spend model follows the pacing window. `WINDOW_KEYS` (framework.js)
|
|
49
|
+
gains `monthly: {snapshot: 'monthly', history: 'monthly', windowMs: null}`,
|
|
50
|
+
and `rateForWindow` (`src/lib/spend.js`) derives the bootstrap window start
|
|
51
|
+
with `meta.windowMs ?? monthlyWindowMs(resetsAtMs)` — the calendar month
|
|
52
|
+
ending at the provider's `resets_at` (M2), never an assumed 30 days.
|
|
53
|
+
`attachSpend` now writes `pool.spend.monthly` and `pool.projectedMonthlyPct`
|
|
54
|
+
next to the fiveHour/weekly fields, plus `pool.spend.pacing = {window,
|
|
55
|
+
ratePerMinute, source, samples}` and `pool.projectedPacingPct` for the
|
|
56
|
+
window that paces the pool (default `weekly`, so a pool that declares
|
|
57
|
+
nothing keeps its old numbers). `inflightLoad` (`src/lib/route.js`) charges
|
|
58
|
+
the in-flight penalty from `spend.pacing?.ratePerMinute ??
|
|
59
|
+
spend.weekly?.ratePerMinute` with that rate's own source label, so the
|
|
60
|
+
surplus and the penalty are measured in the same window; candidate rows gain
|
|
61
|
+
`pacingWindow` and `projectedPacingPct` beside the unchanged
|
|
62
|
+
`projectedWeeklyPct`.
|
|
63
|
+
- Operator control and display. `bullswarm strategy set-subscription <pool>
|
|
64
|
+
--quota-window <weekly|monthly>` now selects the window that paces routing
|
|
65
|
+
(help text in `src/help.js`) and validates it: anything else exits 2 with
|
|
66
|
+
`--quota-window must be weekly or monthly (or unknown to clear)`, and
|
|
67
|
+
`unknown` clears the override back to the connector's declaration. Labels
|
|
68
|
+
already stored are ignored for pacing, never rejected on read. `bullswarm
|
|
69
|
+
pools` names the window in the meter column — `cmd-fixture cost=5
|
|
70
|
+
lanes=chore monthly used 79.4% elapsed 75.3% [cache] surplus=-4.1
|
|
71
|
+
inflight=0 5h=25.2% ready` — and `pools --json` entries carry
|
|
72
|
+
`pacingWindow`. `strategy refresh`/`show` print the window on each
|
|
73
|
+
subscription line (`command-code: GOAT · ... · monthly 79.4% used · surplus
|
|
74
|
+
-4.1`) and carry `pacingWindow` next to the free-text `quotaWindow` label in
|
|
75
|
+
`--json`; `strategy inventory --json` carries it per provider.
|
|
76
|
+
- Tests: 738 -> 750, 0 failures. The new behaviour is covered in
|
|
77
|
+
`tests/meters.test.js` (the live command-code snapshot as a fixture, the
|
|
78
|
+
helper's precedence, `buildPools` pacing), `tests/spend.test.js` (monthly
|
|
79
|
+
bootstrap window start, `spend.pacing`/`projectedPacingPct`),
|
|
80
|
+
`tests/route.test.js` (the penalty on the pacing window; weekly-only pools
|
|
81
|
+
unchanged), `tests/strategy-cli.test.js` (`--quota-window` validation) and
|
|
82
|
+
`tests/assignments.test.js` (the `pools` meter column and `--json`
|
|
83
|
+
`pacingWindow`).
|
|
84
|
+
|
|
85
|
+
## 0.28.0 — context diet
|
|
86
|
+
|
|
87
|
+
- `workflow runs result <id> --summary` prints a compact status-loop
|
|
88
|
+
envelope, `schemaVersion: "bullswarm.workflow.result-summary.v1"`. It
|
|
89
|
+
carries `runId`, `shortId`, `status`, `verified`, `executionMode`,
|
|
90
|
+
`reason`, `finishedAt`, the goal's first line trimmed to 120 characters
|
|
91
|
+
plus `goalBytes`, each requirement as `{ id, status, mandatory,
|
|
92
|
+
evidenceCount, why }` (`why` is the latest evidence's first line trimmed
|
|
93
|
+
to 200 characters), each action as `{ id, kind, lane, effort, status,
|
|
94
|
+
pool, model, reasoning, wallSec, outFile, bytes }`, `concerns: { count,
|
|
95
|
+
first }` (up to three one-liners), the `usage` block, and `next: { full:
|
|
96
|
+
"bullswarm workflow runs result <id> --json", outputs: [<outFile paths>]
|
|
97
|
+
}`. `--summary` implies JSON with or without `--json`; it is not
|
|
98
|
+
TTY-dependent. The full `bullswarm.workflow.result.v2` envelope is
|
|
99
|
+
unchanged and stays the default. Read the full envelope on a failed or
|
|
100
|
+
partial run, or before judging evidence. Flag, help, and example:
|
|
101
|
+
`src/lib/cli-flags.js`, `src/help.js` (`Usage: bullswarm workflow runs
|
|
102
|
+
result <shortId|runId> [--json] [--summary]`; `--summary` "print the
|
|
103
|
+
compact JSON status-loop envelope; implies --json"). A terminal `workflow
|
|
104
|
+
watch` prints `next: bullswarm workflow runs result <shortId> --json
|
|
105
|
+
--summary`.
|
|
106
|
+
|
|
107
|
+
- Byte accounting on every action attempt. `state.attempts[].bytes` is
|
|
108
|
+
`{ taskFile, authorPrompt, kernel, dependencyInputs, output }` — the task
|
|
109
|
+
file the kernel wrote, the action's own prompt text as authored, the
|
|
110
|
+
remainder after subtracting that prompt and any embedded requirement
|
|
111
|
+
text, the sum of the dependency output files the task points at (0 when
|
|
112
|
+
there are none; `digestOf` drill-down paths are pointers, not inputs),
|
|
113
|
+
and the durable out file on completion. The result envelope copies the
|
|
114
|
+
last attempt's `bytes` onto `actions[]` and totals `usage.bytes: {
|
|
115
|
+
taskFiles, dependencyInputs, outputs }`. `workflow runs show` appends
|
|
116
|
+
`in <taskFile>/<dependencyInputs> out <output>` per attempt (the unit
|
|
117
|
+
fixture prints `in 3.1K/60.8K out 14.8K`); missing values stay blank /
|
|
118
|
+
null, never guessed. These are UTF-8 byte counts, never tokens.
|
|
119
|
+
|
|
120
|
+
- `kind: "digest"` (analyze/low) is an extractive condensation of its
|
|
121
|
+
dependencies' outputs so an expensive consumer reads one artifact
|
|
122
|
+
instead of many raw out-files. The kernel writes the whole task: quote
|
|
123
|
+
verbatim (never paraphrase or judge) each source's delivered items,
|
|
124
|
+
validation numbers, commands and their output, unfinished work, and
|
|
125
|
+
every shared-file or integrator request; one section per source headed
|
|
126
|
+
by its absolute output path; no verdicts, no recommendations, no new
|
|
127
|
+
claims; target at most a quarter of the input bytes or 8 KB, whichever
|
|
128
|
+
is larger. The author's prompt is focus guidance only. Validation
|
|
129
|
+
(exit 2 with the reason at `workflow plan validate` and `workflow goal
|
|
130
|
+
--program`): a digest must depend on at least one action, must have
|
|
131
|
+
empty `evidenceFor`, owns no files, needs no `affects`. No evidence
|
|
132
|
+
action may list a digest in `dependsOn` — evidence reads the real
|
|
133
|
+
artifacts. Consumers that depend on a digest receive, in their
|
|
134
|
+
dependency artifacts, the digest entry plus `digestOf: [{ actionId,
|
|
135
|
+
outputFile }]` for each digested source. Use one when three or more
|
|
136
|
+
writers feed a single integrator, or when a consumer's dependency
|
|
137
|
+
outputs would exceed roughly 20 KB; never for evidence.
|
|
138
|
+
|
|
139
|
+
- The planning contract states the kind. `workflow plan contract --json`
|
|
140
|
+
(`Usage: bullswarm workflow plan contract "<goal>" [--cwd <dir>]
|
|
141
|
+
[--json]`) returns 16 rules; `rules[5]` lists `digest=analyze/low` in
|
|
142
|
+
the kind table and `rules[6]` is the extractive digest rule (when to
|
|
143
|
+
insert one, `digestOf`, empty `evidenceFor` / `ownedFiles`, evidence
|
|
144
|
+
must not depend on a digest).
|
|
145
|
+
|
|
146
|
+
- The packaged skill's status loop (`skill/SKILL.md`,
|
|
147
|
+
`skill/references/operations.md`) now recommends `bullswarm workflow
|
|
148
|
+
runs result <shortId> --json --summary`. Read the full envelope with
|
|
149
|
+
`--json` alone when the run is failed or partial, or before judging
|
|
150
|
+
evidence.
|
|
151
|
+
|
|
152
|
+
- Measured numbers, each with its command. On the real 0.27.1 build run
|
|
153
|
+
`ze5xz2` (files under `.diet-inputs/`): `workflow runs result ze5xz2
|
|
154
|
+
--json` is 60,709 bytes on disk (`wc -c .diet-inputs/real-result-ze5xz2.json`;
|
|
155
|
+
the same file is `tests/fixtures/real-result-ze5xz2.json`). The 0.28.0
|
|
156
|
+
goal recorded that envelope's `requirements` as 39,231 bytes and `goal`
|
|
157
|
+
as 11,009 bytes. Re-measuring the same file: `Buffer.byteLength(goal)` =
|
|
158
|
+
10,996 (the summary's `goalBytes`) and `JSON.stringify(requirements)` =
|
|
159
|
+
39,288; compact `JSON.stringify` of the parsed envelope is 57,141.
|
|
160
|
+
`summarizeV2Result` of that fixture is 3,786 bytes —
|
|
161
|
+
`tests/workflow-result-summary.test.js` prints `result-summary size:
|
|
162
|
+
full=57141 summary=3786` (numbers re-measured in 0.28.1). The 0.28.0
|
|
163
|
+
goal recorded the integrator's inputs as 60,790 bytes; `wc -c` of the
|
|
164
|
+
seven dependency out-files under `.diet-inputs/` sums to 46,022
|
|
165
|
+
(out-surface 18,659, out-routing-cleanup
|
|
166
|
+
10,202, out-state-bugs 7,052, out-docs 6,831, out-dead-kernel 1,377,
|
|
167
|
+
out-verify-gate 974, out-dead-code 927) and the integrator task file is
|
|
168
|
+
14,768 (`wc -c .diet-inputs/task-integrate-attempt-1.md`), which
|
|
169
|
+
together are 60,790.
|
|
170
|
+
|
|
171
|
+
- Fixture measurement, the same goal run twice under a temporary home
|
|
172
|
+
(`tests/workflow-context-diet-measurement.test.js`, which prints both
|
|
173
|
+
lines below). Three writers each padded to a few KB feed one integrator
|
|
174
|
+
directly, then the same three feed a `kind: "digest"` that feeds the
|
|
175
|
+
integrator: the integrator's `bytes.dependencyInputs` falls from
|
|
176
|
+
**12,477 bytes to 63 bytes** (`context diet: integrator dependencyInputs
|
|
177
|
+
without digest=12477 with digest=63`). The 63 is a floor, not a
|
|
178
|
+
condensation ratio — the deterministic fixture worker answers with a
|
|
179
|
+
fixed stub instead of really condensing. The ceiling is the byte target
|
|
180
|
+
the kernel writes into that digest's own task, **8,192 bytes** for this
|
|
181
|
+
input, and the test asserts the saving holds at that ceiling too
|
|
182
|
+
(8,192 < 12,477). The same run's envelopes print as `context diet: run B
|
|
183
|
+
envelope full=<n> summary=<n>` (the full envelope embeds the temporary
|
|
184
|
+
home's absolute paths, so its size varies a few bytes between runs; the
|
|
185
|
+
summary carries one `next.runDir` string plus basenames, so it does not);
|
|
186
|
+
the deterministic envelope comparison to quote is the real-run fixture
|
|
187
|
+
above (57,141 full vs the measured summary).
|
|
188
|
+
|
|
189
|
+
- `workflow capabilities` now reports the closed kind list at
|
|
190
|
+
`engines.autonomousV2.actionKinds`, cloned from the validator's
|
|
191
|
+
`KIND_DEFAULTS` rather than hand-listed, so `digest` and every future
|
|
192
|
+
kind are discoverable by a probing agent
|
|
193
|
+
(`src/workflow/cli.js`). The planning contract's
|
|
194
|
+
`program.actionFields.kind` description is derived from the same table
|
|
195
|
+
(`src/workflow/v2-planner.js`).
|
|
196
|
+
|
|
197
|
+
- `TIER_LANES` (`src/lib/strategy.js`) excludes kernel-owned kinds from the
|
|
198
|
+
effort-tier count, so the map stays `{ high: analyze, medium: build,
|
|
199
|
+
low: chore }`. A digest is a mechanism the kernel writes, not a nature of
|
|
200
|
+
work that should define which lane a tier routes to; counting it would
|
|
201
|
+
have flipped low from `chore` to `analyze` on the strength of an action
|
|
202
|
+
no planner has to reason about. No shipped connector's routing changes
|
|
203
|
+
either way — all six declare all three lanes.
|
|
204
|
+
|
|
205
|
+
- The three KaiHK-backed OpenCode pools can now run gpt-5.6-luna at a
|
|
206
|
+
chosen reasoning level. `connectors/opencode2.json` declares
|
|
207
|
+
`reasoning: { flag: "--variant", levels: [low, medium, high, xhigh,
|
|
208
|
+
max], defaults: { high: high, medium: medium, low: low } }` — the same
|
|
209
|
+
five levels as `connectors/command-code.json`, which fronts the same
|
|
210
|
+
backend — so rungs for these pools stop printing `— (unsupported)`.
|
|
211
|
+
opencode only forwards a `--variant` its config declares for that
|
|
212
|
+
model, so the flag alone would be silently dropped;
|
|
213
|
+
`expandOpenCodeKaihkConnectors` (`src/lib/opencode-kaihk.js`) therefore
|
|
214
|
+
sets `env.OPENCODE_CONFIG_CONTENT` on the base pool and on every clone
|
|
215
|
+
to the variants for that pool's OWN provider id, via the new pure
|
|
216
|
+
helper `kaihkVariantsConfig(providerId, model = KAIHK_OPENCODE_MODEL)`:
|
|
217
|
+
`{"provider":{"kaihk-2":{"models":{"gpt-5.6-luna":{"variants":{"low":{"reasoningEffort":"low"},"medium":{"reasoningEffort":"medium"},"high":{"reasoningEffort":"high"},"xhigh":{"reasoningEffort":"xhigh"},"max":{"reasoningEffort":"max"}}}}}}}`.
|
|
218
|
+
opencode merges that JSON string over the config file, so the API key
|
|
219
|
+
and everything else in `~/.config/opencode/opencode.json` stays in
|
|
220
|
+
force. An `OPENCODE_CONFIG_CONTENT` the operator set by hand in the
|
|
221
|
+
installed connector is never overwritten, on the base pool or on the
|
|
222
|
+
clones. A medium/max dispatch on `opencode2:kaihk-2` composes
|
|
223
|
+
`opencode run --auto --model kaihk-2/gpt-5.6-luna <taskFile> --variant
|
|
224
|
+
max --format json`; a level already pinned in the template is replaced,
|
|
225
|
+
not duplicated. Per pool:
|
|
226
|
+
`bullswarm strategy set-rung opencode2 medium --model kaihk/gpt-5.6-luna --reasoning max`,
|
|
227
|
+
`bullswarm strategy set-rung opencode2:kaihk-2 medium --model kaihk-2/gpt-5.6-luna --reasoning max`,
|
|
228
|
+
`bullswarm strategy set-rung opencode2:kaihk-3 medium --model kaihk-3/gpt-5.6-luna --reasoning max`.
|
|
229
|
+
Existing installations pick the block up through
|
|
230
|
+
`upgradeConnectorMetadata` (`src/setup.js`), which backfills a missing
|
|
231
|
+
`reasoning` block and leaves a customised one alone. Non-KaiHK opencode
|
|
232
|
+
installations get no injected variants, so `--variant` is a no-op there
|
|
233
|
+
rather than an error — recorded in the connector's
|
|
234
|
+
`$comment-reasoning`.
|
|
235
|
+
|
|
236
|
+
- Tests: 713 -> 738, 0 failures
|
|
237
|
+
(`env -u CLAUDE_CONFIG_DIR -u FORCE_COLOR -u NO_COLOR npm test`). Four
|
|
238
|
+
new files carry the new behaviour: `tests/workflow-bytes.test.js`,
|
|
239
|
+
`tests/workflow-digest.test.js`,
|
|
240
|
+
`tests/workflow-result-summary.test.js` (with the
|
|
241
|
+
`tests/fixtures/real-result-ze5xz2.json` envelope it measures) and
|
|
242
|
+
`tests/workflow-context-diet-measurement.test.js`.
|
|
243
|
+
|
|
3
244
|
## 0.27.1 — audit cleanup
|
|
4
245
|
|
|
5
246
|
- Deleted the remaining dead symbols the 2026-09-09 audit listed as Tier A:
|
package/README.md
CHANGED
|
@@ -51,7 +51,12 @@ detaches safely.
|
|
|
51
51
|
passing verification.
|
|
52
52
|
2. **Pace by meter.** The scheduling resource is the subscription window:
|
|
53
53
|
elapsed% minus used%, most-behind pool wins. Pace may only promote a
|
|
54
|
-
*cheaper* pool. Lanes are work-nature, never hard-coded to pools.
|
|
54
|
+
*cheaper* pool. Lanes are work-nature, never hard-coded to pools. Which
|
|
55
|
+
window paces one pool is the subscription window that pool's connector
|
|
56
|
+
declares (`quotaWindow`: weekly for claude-code, codex and grok; monthly
|
|
57
|
+
for command-code and the kaihk pools), overridable per pool with
|
|
58
|
+
`bullswarm strategy set-subscription <pool> --quota-window <weekly|monthly>`
|
|
59
|
+
— `bullswarm pools` names it in the meter column. The
|
|
55
60
|
5-hour window never paces — it gates: a pool at or above 75% of it is
|
|
56
61
|
chosen only when no eligible pool below that line exists, and one at or
|
|
57
62
|
above 90% is not dispatched at all.
|
|
@@ -180,6 +185,25 @@ per pool and tier, so the tier moves off whichever model held it while that
|
|
|
180
185
|
model keeps its other tiers. Nothing about `state.json` changed shape: rungs are
|
|
181
186
|
a view over `strategy.modelTiers` and `strategy.reasoning`.
|
|
182
187
|
|
|
188
|
+
The KaiHK-backed OpenCode pools (`opencode2`, `opencode2:kaihk-2`,
|
|
189
|
+
`opencode2:kaihk-3`) express reasoning as opencode's `--variant <level>`, at the
|
|
190
|
+
same five levels as `command-code`. opencode only forwards a variant its own
|
|
191
|
+
config declares for that model, so bullswarm injects them: each pool is spawned
|
|
192
|
+
with `OPENCODE_CONFIG_CONTENT` declaring `low`/`medium`/`high`/`xhigh`/`max` as
|
|
193
|
+
`reasoningEffort` variants of `<providerId>/gpt-5.6-luna`, merged over your
|
|
194
|
+
`~/.config/opencode/opencode.json` (your API keys stay in force). Without that
|
|
195
|
+
injection opencode accepts `--variant` and silently drops it. Set a rung per
|
|
196
|
+
pool, using that pool's own provider prefix:
|
|
197
|
+
|
|
198
|
+
```bash
|
|
199
|
+
bullswarm strategy set-rung opencode2:kaihk-2 medium \
|
|
200
|
+
--model kaihk-2/gpt-5.6-luna --reasoning max
|
|
201
|
+
```
|
|
202
|
+
|
|
203
|
+
If you set `OPENCODE_CONFIG_CONTENT` yourself in
|
|
204
|
+
`~/.bullswarm/connectors/opencode2.json`, bullswarm leaves it alone and injects
|
|
205
|
+
nothing — you own the variants from then on.
|
|
206
|
+
|
|
183
207
|
The benchmark evidence comes from Epoch AI's benchmarking hub, used under
|
|
184
208
|
CC BY 4.0: Epoch AI, 'AI Benchmarking Hub'. Published online at epoch.ai.
|
|
185
209
|
Retrieved from <https://epoch.ai/benchmarks>. `blended` is the mean of the
|
|
@@ -390,12 +414,25 @@ names that nature once and derives them:
|
|
|
390
414
|
| --- | --- | --- |
|
|
391
415
|
| `mechanical` | chore | low |
|
|
392
416
|
| `io-read` | analyze | low |
|
|
417
|
+
| `digest` | analyze | low |
|
|
393
418
|
| `check` | analyze | medium |
|
|
394
419
|
| `implement` | build | medium |
|
|
395
420
|
| `integration` | build | high |
|
|
396
421
|
| `architecture` | analyze | high |
|
|
397
422
|
| `adversarial-acceptance` | analyze | high |
|
|
398
423
|
|
|
424
|
+
`digest` is the one kind whose instructions the kernel supplies in full — your
|
|
425
|
+
prompt for it is focus guidance only. It condenses the
|
|
426
|
+
outputs of the actions it depends on — quoting each source's delivered items,
|
|
427
|
+
validation numbers, commands, unfinished work, and requests verbatim, one
|
|
428
|
+
section per source, with no verdicts of its own — so an expensive consumer
|
|
429
|
+
reads one artifact instead of many raw output files, and the digest entry in
|
|
430
|
+
that consumer's dependency artifacts still names every digested source for
|
|
431
|
+
drill-down. Use one when three or more writers feed a single integrator, or
|
|
432
|
+
when a consumer's dependency outputs would exceed roughly 20 KB. A digest must
|
|
433
|
+
depend on at least one action, owns no files, needs no `affects`, and no
|
|
434
|
+
evidence action may depend on one: evidence reads the real artifacts.
|
|
435
|
+
|
|
399
436
|
Resolution is per field: an explicit `lane` or `effort` on the action wins,
|
|
400
437
|
then the kind table, then an optional program-level `defaults` object — which
|
|
401
438
|
may set only `effort` and `reasoning`, because lane follows the individual
|
|
@@ -439,6 +476,7 @@ bullswarm workflow runs show <shortId>
|
|
|
439
476
|
bullswarm workflow watch <shortId> # V2: attach, then one line per notable event
|
|
440
477
|
bullswarm workflow watch <shortId> --next # print the next notable event and exit
|
|
441
478
|
# relaunch with the --after/--since it prints
|
|
479
|
+
bullswarm workflow runs result <shortId> --json --summary # compact status-loop envelope once terminal
|
|
442
480
|
bullswarm workflow # unified human workflow home
|
|
443
481
|
bullswarm workflow tui <shortId> # jump directly to one run timeline
|
|
444
482
|
bullswarm workflow tui --json <shortId>
|
|
@@ -574,7 +612,8 @@ bullswarm workflow runs --historical --since yesterday --until today
|
|
|
574
612
|
bullswarm workflow runs --all --from 2026-08-20 --to 2026-08-27
|
|
575
613
|
bullswarm workflow runs --limit 20 # cap the result count
|
|
576
614
|
bullswarm workflow runs show <shortId> # state + report + summary
|
|
577
|
-
bullswarm workflow runs result <shortId> --json #
|
|
615
|
+
bullswarm workflow runs result <shortId> --json --summary # compact status-loop envelope
|
|
616
|
+
bullswarm workflow runs result <shortId> --json # full envelope (failed/partial, or before judging evidence)
|
|
578
617
|
bullswarm runs show <shortId> # top-level shorthand
|
|
579
618
|
bullswarm workflow runs delete <shortId> --yes # remove the run dir
|
|
580
619
|
```
|
|
@@ -591,8 +630,10 @@ Values accept ISO timestamps, local `YYYY-MM-DD` dates, `today`, `yesterday`,
|
|
|
591
630
|
`tomorrow`, `now`, or relative durations such as `30m`, `24h`, `7d`, and `2w`.
|
|
592
631
|
|
|
593
632
|
After a workflow reaches a terminal state, agents should consume
|
|
594
|
-
`workflow runs result <id> --json`
|
|
595
|
-
or provider-specific output.
|
|
633
|
+
`workflow runs result <id> --json --summary` for the status loop instead of
|
|
634
|
+
probing `state.json`, task files, or provider-specific output. Read the full
|
|
635
|
+
envelope with `--json` alone when the run is failed or partial, or before
|
|
636
|
+
judging evidence. Autonomous V2's full document is the versioned
|
|
596
637
|
`bullswarm.workflow.result.v2` envelope with kernel-computed status, fresh
|
|
597
638
|
requirement evidence, per-action status/failure/output files, explicit gaps,
|
|
598
639
|
usage, and verification qualification. New programs include `executionMode:
|
|
@@ -608,6 +649,75 @@ otherwise the command returns after printing this handoff.
|
|
|
608
649
|
Time filters preserve the existing scope, so use `--all` or `--historical` when
|
|
609
650
|
auditing completed runs.
|
|
610
651
|
|
|
652
|
+
### Context diet
|
|
653
|
+
|
|
654
|
+
The kernel now measures — and can shrink — what it puts in front of a model.
|
|
655
|
+
These are UTF-8 byte counts, never tokens.
|
|
656
|
+
|
|
657
|
+
**Status loop.** Poll with `--summary`; it implies JSON (with or without
|
|
658
|
+
`--json`) and prints `schemaVersion: "bullswarm.workflow.result-summary.v1"`:
|
|
659
|
+
`runId`, `shortId`, `status`, `verified`, `executionMode`, `reason`,
|
|
660
|
+
`finishedAt`, the goal's first line (120 characters) plus `goalBytes`, each
|
|
661
|
+
requirement as `{ id, status, mandatory, evidenceCount, why }`, each action as
|
|
662
|
+
`{ id, kind, lane, effort, status, pool, model, reasoning, wallSec, outFile,
|
|
663
|
+
bytes }`, `concerns: { count, first }`, `usage`, and `next: { full, runDir, outputs }` — every output name is a basename inside `next.runDir`.
|
|
664
|
+
`--summary` is single-line JSON (`JSON.stringify`), so the bytes on the wire
|
|
665
|
+
match the 4,096-byte fitter budget. As printed by the CLI on
|
|
666
|
+
`tests/fixtures/real-result-ze5xz2.json`, the compact summary is 3,786 bytes
|
|
667
|
+
and the pretty full envelope (`--json` alone) is 60,709 bytes.
|
|
668
|
+
The full `bullswarm.workflow.result.v2` envelope is unchanged and remains the
|
|
669
|
+
default. Read it (`--json` alone) on a failed or partial run, or before judging
|
|
670
|
+
evidence. A terminal `workflow watch` prints the same compact command as
|
|
671
|
+
`next:`.
|
|
672
|
+
|
|
673
|
+
```bash
|
|
674
|
+
bullswarm workflow runs result <shortId> --json --summary
|
|
675
|
+
bullswarm workflow runs result <shortId> --json
|
|
676
|
+
```
|
|
677
|
+
|
|
678
|
+
`workflow runs result --help` states `Usage: bullswarm workflow runs result
|
|
679
|
+
<shortId|runId> [--json] [--summary]`; `--summary` is "print the compact JSON
|
|
680
|
+
status-loop envelope; implies --json".
|
|
681
|
+
|
|
682
|
+
**Bytes.** Every attempt records `bytes: { taskFile, authorPrompt, kernel,
|
|
683
|
+
dependencyInputs, output }` — the task file the kernel wrote, the action's own
|
|
684
|
+
prompt as authored, the remainder after subtracting that prompt and any
|
|
685
|
+
embedded requirement text, the sum of the dependency output files the task
|
|
686
|
+
points at (0 when there are none), and the durable out file on completion.
|
|
687
|
+
The result envelope copies the last attempt's `bytes` onto `actions[]` and
|
|
688
|
+
totals `usage.bytes: { taskFiles, dependencyInputs, outputs }`.
|
|
689
|
+
`workflow runs show` appends `in <taskFile>/<dependencyInputs> out <output>`
|
|
690
|
+
per attempt (blank when unrecorded). Missing values are null, never guessed.
|
|
691
|
+
|
|
692
|
+
**Digest.** `kind: "digest"` is analyze/low. It is an extractive condensation
|
|
693
|
+
of its dependencies' outputs — quoted delivered items, validation numbers,
|
|
694
|
+
commands, unfinished work, and integrator requests; no verdicts of its own.
|
|
695
|
+
The kernel writes the whole task; the author's prompt is focus guidance only.
|
|
696
|
+
Use one when three or more writers feed a single integrator, or when a
|
|
697
|
+
consumer's dependency outputs would exceed roughly 20 KB. A digest must
|
|
698
|
+
depend on at least one action, owns no files, has empty `evidenceFor`, and
|
|
699
|
+
needs no `affects`. Evidence must not depend on a digest: evidence reads the
|
|
700
|
+
real artifacts. Consumers that depend on a digest receive that digest plus a
|
|
701
|
+
`digestOf` array of `{ actionId, outputFile }` so they can drill down; those
|
|
702
|
+
paths are pointers, not extra `dependencyInputs`. (The kind table above
|
|
703
|
+
derives lane and effort.)
|
|
704
|
+
|
|
705
|
+
```json
|
|
706
|
+
{
|
|
707
|
+
"schemaVersion": "bullswarm.workflow.program.v2",
|
|
708
|
+
"actions": [
|
|
709
|
+
{ "id": "write-a", "kind": "implement", "dependsOn": [], "ownedFiles": ["a.ts"], "affects": ["requirement-1"], "evidenceFor": [], "purpose": "Write slice A", "prompt": "Implement A and report the checks you ran." },
|
|
710
|
+
{ "id": "write-b", "kind": "implement", "dependsOn": [], "ownedFiles": ["b.ts"], "affects": ["requirement-2"], "evidenceFor": [], "purpose": "Write slice B", "prompt": "Implement B and report the checks you ran." },
|
|
711
|
+
{ "id": "write-c", "kind": "implement", "dependsOn": [], "ownedFiles": ["c.ts"], "affects": ["requirement-3"], "evidenceFor": [], "purpose": "Write slice C", "prompt": "Implement C and report the checks you ran." },
|
|
712
|
+
{ "id": "condense", "kind": "digest", "dependsOn": ["write-a", "write-b", "write-c"], "ownedFiles": [], "affects": [], "evidenceFor": [], "purpose": "Condense the writer outputs", "prompt": "Keep every acceptance number and every shared-file request." },
|
|
713
|
+
{ "id": "integrate", "kind": "integration", "dependsOn": ["condense"], "ownedFiles": [], "affects": ["requirement-1", "requirement-2", "requirement-3"], "evidenceFor": [], "purpose": "Integrate and run the gates", "prompt": "Apply every request the digest carries and run the repository gates." }
|
|
714
|
+
]
|
|
715
|
+
}
|
|
716
|
+
```
|
|
717
|
+
|
|
718
|
+
An evidence action for those requirements depends on `write-a`, `write-b`, and
|
|
719
|
+
`write-c` — never on `condense`.
|
|
720
|
+
|
|
611
721
|
### Live workflow dashboard
|
|
612
722
|
|
|
613
723
|
For ordinary observation, use the non-interactive watcher. For V2 runs it
|
|
@@ -615,8 +725,9 @@ prints one attach line, then one line per notable event as it happens
|
|
|
615
725
|
(action finished/failed/blocked/cancelled, evidence, stage completion,
|
|
616
726
|
planner turn, stall/recovery, cancellation, and the existing pause and
|
|
617
727
|
terminal `outcome:` / `next:` lines) and stays silent while work is merely
|
|
618
|
-
in progress.
|
|
619
|
-
|
|
728
|
+
in progress. A terminal watch's `next:` line is
|
|
729
|
+
`bullswarm workflow runs result <shortId> --json --summary`. Agent starts,
|
|
730
|
+
mechanical retries, and steering delivery print only with `--verbose`. A usage-limit failure (`failureKind: 'quota'`) always
|
|
620
731
|
prints, verbose or not: `⚠ <actionId> usage limit on <pool> · paused until
|
|
621
732
|
<deadline> · retrying on another pool`, followed once the mechanical retry
|
|
622
733
|
lands on another pool by `↺ <actionId> now on <pool> · <model>`. The
|
|
@@ -34,6 +34,12 @@
|
|
|
34
34
|
"capabilities": ["strong-analysis", "code-reading", "file-editing", "workflow-planning"],
|
|
35
35
|
"modelDiscovery": { "cmd": ["opencode", "models"], "parse": "lines", "includePattern": "^[^\\s]+/[^\\s]+$", "timeoutMs": 20000, "maxModels": 250 },
|
|
36
36
|
"modelSelection": { "flag": "--model", "mode": "replace-or-append" },
|
|
37
|
+
"$comment-reasoning": "verified 2026-09-09 against opencode 1.18.25: `opencode run --help` documents `--variant model variant (provider-specific reasoning effort, e.g., high, max, minimal)`. opencode only forwards a variant its CONFIG declares for that model, and the owner's opencode.json declares none, so on a bare install `--variant` is accepted and silently dropped. The level reaches the KaiHK API as `reasoning_effort` ONLY because expandOpenCodeKaihkConnectors (src/lib/opencode-kaihk.js) injects the five variants for <providerId>/gpt-5.6-luna through env.OPENCODE_CONFIG_CONTENT on every discovered KaiHK pool. Probed with that variable set: `--variant bogus` (reasoningEffort bogus-level) failed at the API with `level \"bogus-level\" not supported, valid levels: low, medium, high, xhigh, max`, and `--variant max` succeeded; probed again with the flag AFTER the positional task text, exactly where bullswarm appends it (`opencode run --auto --model <id>/gpt-5.6-luna <task> --variant <level> --format json`), with the same API rejection for the bogus level, so the position is proven too. Same five levels as connectors/command-code.json, which fronts the same backend. UNVERIFIED for non-KaiHK opencode providers: they get no injected variants, so the flag is a no-op there rather than an error.",
|
|
38
|
+
"reasoning": {
|
|
39
|
+
"flag": "--variant",
|
|
40
|
+
"levels": ["low", "medium", "high", "xhigh", "max"],
|
|
41
|
+
"defaults": { "high": "high", "medium": "medium", "low": "low" }
|
|
42
|
+
},
|
|
37
43
|
"modelProfiles": [
|
|
38
44
|
{ "match": "(?:^|/)claude-fable-", "tier": "high", "qualityRank": 6, "autoRecommend": false },
|
|
39
45
|
{ "match": "gpt-5\\.6-sol$", "tier": "high", "qualityRank": 6, "autoRecommend": true },
|