bullswarm 0.13.0 → 0.13.1

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/CHANGELOG.md CHANGED
@@ -1,5 +1,18 @@
1
1
  # bullswarm changelog
2
2
 
3
+ ## 0.13.1 — a repaired verify counts as verification of its repair
4
+
5
+ - `completionEvidenceGaps` accepted a verify as evidence for the latest worker
6
+ only when the verify depended on that worker. The executor's repair loop
7
+ produces the reverse edge — `<verify>-repair-N` depends on `<verify>`, then
8
+ the same verify re-runs — so after a clean repair round every `complete`
9
+ was rejected with "missing a successful verification of latest worker
10
+ <verify>-repair-1", and 0.13.0's `all-actions-ok` auto-completion would
11
+ have been blocked the same way. Observed on goal-2 run `wf-mtdcghw0`
12
+ (2026-08-28): three extra planner turns and one redundant verify (~11 min)
13
+ to prove what the re-verify had already shown. A verify that ended ok:true
14
+ after its own repair action now verifies that repair.
15
+
3
16
  ## 0.13.0 — programs can complete themselves
4
17
 
5
18
  - A planner may attach `completion: { when: "all-actions-ok", reason }` to a
@@ -339,7 +339,56 @@ the runtime parks the dispatch in `waiting_for_quota`, re-reads the meter
339
339
  every 60 s, continues when the window resets, and only fails — naming pool,
340
340
  usage and reset time — after reset + 10 min grace.
341
341
 
342
- (re-run on 0.12.1 pending — it starts as soon as the 5h window resets at 22:30 Z)
342
+ **Second launch, 19:28:03 Z, installed 0.12.1, run `wf-mtdcghw0-bfefc7` on a
343
+ fresh pristine `g2-bs-v3` — observed the wait working.** The meter read 95 %
344
+ at launch, so the runtime parked the first dispatch (the scout) in stage
345
+ `waiting_for_quota` with `until 22:40:00 Z` (reset 22:30 + 10 min grace),
346
+ event `dispatch.waiting_for_quota` carrying pool, usage and reset time. It
347
+ re-read the meter every 60 s for **3 h 2 min 14 s** (`waitedMs 10933739`) and
348
+ emitted `dispatch.quota_available` at **22:30:16.979 Z** — 17 s after the
349
+ provider reset — then dispatched the scout with no operator action. Wall-clock
350
+ numbers below therefore exclude this wait (`metrics-bullswarm.mjs` reports
351
+ `quotaWaitSec` and `wallExclWaitSec` separately); the wait is a provider
352
+ constraint, not execution time.
353
+
354
+ **Outcome: `completed`, `verified: true`, 23:18:38 Z.** Timeline (all Z):
355
+
356
+ | when | what |
357
+ |---|---|
358
+ | 22:30:17 | scout dispatched (read-only survey), 197 s |
359
+ | 22:33:34 → 22:39:14 | planner turn 1, **340 s** → one `needs_more_work` program of 14 actions |
360
+ | 22:39:14 | **7 workers started in the same second**: `build-{csv,duration,intervals,lru,semver,slugify}` + `write-docs-index` |
361
+ | 22:44:45 → 22:46:36 | each `verify-<module>` started the moment its own builder finished — per-chain pipelining, no stage barrier (`verify-intervals` was running while `build-slugify` still built) |
362
+ | 22:52:29 | all six module verifies `ok:true` (`verify-slugify` 354 s — see landmine note) → `verify-full-delivery` |
363
+ | 22:59:02 | `verify-full-delivery` **ok:false**: `docs/README.md` missing → runtime spawned `verify-full-delivery-repair-1` (`source: repair-policy`), no planner turn |
364
+ | 23:01:19 → 23:07:58 | repair wrote `docs/README.md` (137 s); re-verify passed (399 s) |
365
+ | 23:07:58 → 23:09:42 | planner turn 2 (105 s): `complete` — **rejected by the runtime**: "missing a successful verification of latest worker verify-full-delivery-repair-1" |
366
+ | 23:09:42 → 23:12:25 | planner turn 3 (163 s): diagnosed the rejection as mechanical, added one read-only `verify-final-acceptance` depending on the repair |
367
+ | 23:12:25 → 23:17:54 | `verify-final-acceptance` ok:true (328 s) |
368
+ | 23:17:54 → 23:18:38 | planner turn 4 (44 s): `complete`, accepted |
369
+
370
+ Numbers (`metrics-bullswarm.mjs`, wait excluded):
371
+
372
+ | metric | 0.11.1 (`g2-bs-v2`) | **0.12.1 (`g2-bs-v3`)** | Claude #2 |
373
+ |---|---|---|---|
374
+ | execution wall | 41 min 14 s | **48 min 21 s** (2 901 s; +10 934 s quota wait) | 58 min |
375
+ | planner turns / seconds | 3 / 667 s (27 %) | **4 / 652 s (22 %)** — turns 2–4 (312 s) plus `verify-final-acceptance` (328 s) exist only because of the rejection bug below | 0 during execution |
376
+ | dispatches / max concurrent | 22 / 6 | **22 / 7** | 24 / ~10 |
377
+ | parallelism (busy ÷ wall) | 2.74 | **2.11** | 3.1 |
378
+ | actions by source | planner 14+7+0 | **planner 14 + 1, repair-policy 1** | script |
379
+ | tests after | 130/130 | **120/120** (52 + 68 new, 6 files ≥ 9 tests each) | 168/168 |
380
+ | existing tests / src | byte-identical / comment-only | **byte-identical / comment-only** (audit-fixture.sh: 0 non-comment line diffs in all 7 src files) | same |
381
+ | deliverables | all | **all** (6 docs pages, `docs/README.md` 6-row index) | all |
382
+ | estimated tokens | — | 201 568 (utf8/4 estimate) | — |
383
+
384
+ What the run showed:
385
+
386
+ 1. **The brace landmine is closed (controlled A/B).** `g2-bs-v3` is a pristine copy, so `src/slugify.js` still carries the `@param {{maxLength?: number}}` JSDoc that killed 0.11.1's `verify-slugify` with zero attempts. On 0.12.1 the same verify dispatched, reviewed the artifact with the braces intact, and returned `ok:true` (with three informational concerns, none of which spawned a polish action — the doctrine held).
387
+ 2. **The planner compiled a Claude-shaped program on the first turn.** Six independent `build → verify` chains + a parallel docs-index builder + one final gate, each verify carrying `repair {maxRounds: 1}`. The runtime then ran it as a pipeline: verifies started per chain, not after a barrier.
388
+ 3. **The repair loop worked live, and paid for a planner mistake.** `write-docs-index` was compiled with `dependsOn: []`, so it launched with the builders and found no `docs/` to index; the worker refused to invent summaries and returned a status note. The final verify caught the missing file and the runtime's repair round fixed it — no planner turn, ~9 min. In Claude's model the same mistake is an authoring error in the script; here it is a compile error by the planner. Neither runtime can catch it deterministically; both recover through verification.
389
+ 4. **Runtime bug found: a repair action is never "verified".** `completionEvidenceGaps` accepts a verify as evidence for the latest worker only if `verify.dependsOn` includes that worker. A repair action depends on its verify (the reverse edge), and the verify's post-repair re-run *is* its verification, but the check does not know that — so a clean `complete` was rejected and the run spent 3 more turns and ~11 min proving what it already had. The same check gates 0.13.0's `all-actions-ok` auto-completion, which would have been blocked the same way. Fix: 0.13.1 (below).
390
+
391
+ Take the bug and the dependency slip out and this run is ~29 min of execution with two planner turns — the shape the 0.12.0 design targeted.
343
392
 
344
393
  The originally planned 0.10.9 goal-2 run was dropped at the user's request
345
394
  (2026-08-29): the installed latest is the only baseline that matters.
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "bullswarm",
3
- "version": "0.13.0",
3
+ "version": "0.13.1",
4
4
  "description": "Route work across coding-agent CLI subscriptions — paced by live quota meters, verified by content, never trusting exit codes.",
5
5
  "type": "module",
6
6
  "bin": {
@@ -475,6 +475,19 @@ function actionOutputOk(action, outputs) {
475
475
  : action.status === 'succeeded';
476
476
  }
477
477
 
478
+ // A verify is evidence for a worker when the worker feeds it (dependsOn), or
479
+ // when the worker is that verify's own repair action: the executor re-runs the
480
+ // verify after every repair round, so a verify that ended ok:true after its
481
+ // `<verify>-repair-N` has verified the repair even though the dependency edge
482
+ // points the other way. Without this a clean run's `complete` was rejected as
483
+ // "missing a successful verification of latest worker <verify>-repair-1"
484
+ // (observed on goal-2 run wf-mtdcghw0, 2026-08-28) and cost three planner turns.
485
+ export function verifiesWorker(verify, worker) {
486
+ if ((verify.dependsOn ?? []).includes(worker.id)) return true;
487
+ return (worker.dependsOn ?? []).includes(verify.id)
488
+ && new RegExp(`^${verify.id.replace(/[.*+?^${}()|[\]\\]/g, '\\$&')}-repair-\\d+$`).test(worker.id);
489
+ }
490
+
478
491
  export function completionEvidenceGaps(dynamicActions, policy, outputs = {}) {
479
492
  const missing = [];
480
493
  const successfulWorkers = dynamicActions.filter(
@@ -490,7 +503,7 @@ export function completionEvidenceGaps(dynamicActions, policy, outputs = {}) {
490
503
  && action.status === 'succeeded'
491
504
  && actionOutputOk(action, outputs)
492
505
  && latestSuccessfulWorker
493
- && (action.dependsOn ?? []).includes(latestSuccessfulWorker.id),
506
+ && verifiesWorker(action, latestSuccessfulWorker),
494
507
  )) {
495
508
  missing.push(latestSuccessfulWorker
496
509
  ? `a successful verification of latest worker ${latestSuccessfulWorker.id}`