bullswarm 0.12.1 → 0.13.1
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +26 -0
- package/docs/claude-dynamic-workflow-mechanics.md +13 -0
- package/docs/experiments/2026-08-29-ultracode-vs-bullswarm.md +131 -3
- package/package.json +1 -1
- package/skill/SKILL.md +6 -0
- package/src/workflow/decision.js +27 -0
- package/src/workflow/goal.js +1 -1
- package/src/workflow/runner.js +67 -1
- package/src/workflow/runtime.js +2 -1
package/CHANGELOG.md
CHANGED
|
@@ -1,5 +1,31 @@
|
|
|
1
1
|
# bullswarm changelog
|
|
2
2
|
|
|
3
|
+
## 0.13.1 — a repaired verify counts as verification of its repair
|
|
4
|
+
|
|
5
|
+
- `completionEvidenceGaps` accepted a verify as evidence for the latest worker
|
|
6
|
+
only when the verify depended on that worker. The executor's repair loop
|
|
7
|
+
produces the reverse edge — `<verify>-repair-N` depends on `<verify>`, then
|
|
8
|
+
the same verify re-runs — so after a clean repair round every `complete`
|
|
9
|
+
was rejected with "missing a successful verification of latest worker
|
|
10
|
+
<verify>-repair-1", and 0.13.0's `all-actions-ok` auto-completion would
|
|
11
|
+
have been blocked the same way. Observed on goal-2 run `wf-mtdcghw0`
|
|
12
|
+
(2026-08-28): three extra planner turns and one redundant verify (~11 min)
|
|
13
|
+
to prove what the re-verify had already shown. A verify that ended ok:true
|
|
14
|
+
after its own repair action now verifies that repair.
|
|
15
|
+
|
|
16
|
+
## 0.13.0 — programs can complete themselves
|
|
17
|
+
|
|
18
|
+
- A planner may attach `completion: { when: "all-actions-ok", reason }` to a
|
|
19
|
+
program (a `needs_more_work` decision that includes at least one verify).
|
|
20
|
+
When every action of that program — repairs included — finishes ok and the
|
|
21
|
+
completion policy is satisfied, the runtime records the `complete` decision
|
|
22
|
+
itself (`source: "program-completion"`, event `decision.auto_completed`) and
|
|
23
|
+
the run ends without another planner turn. Anything failing emits
|
|
24
|
+
`decision.completion_predicate_unmet` (with the failing action ids) and the
|
|
25
|
+
boundary returns to the planner as before. Measured motivation: in the goal-2
|
|
26
|
+
comparison every clean bullswarm run still paid a final 110–250 s planner
|
|
27
|
+
turn just to say "complete"; Claude's script ends when its code says so.
|
|
28
|
+
|
|
3
29
|
## 0.12.1 — a burst-gated provider is waited for, never failed on the spot
|
|
4
30
|
|
|
5
31
|
- `workflow` runs no longer die with `no eligible pool` when every candidate
|
|
@@ -316,6 +316,19 @@ author and the `Workflow` runtime.
|
|
|
316
316
|
sees `outputs.scout.ok=false` with the reason, and a run where only the
|
|
317
317
|
scout succeeded is `blocked`, never "delivered".
|
|
318
318
|
|
|
319
|
+
8. **Program-level completion** (0.13.0, unreleased at the time of writing) —
|
|
320
|
+
`completion: { when: "all-actions-ok", reason }` on a program. Claude's
|
|
321
|
+
script simply returns when its code is done; bullswarm still spent a final
|
|
322
|
+
planner turn (110–250 s measured) to say `complete` after a clean run. Now
|
|
323
|
+
the runtime records that decision itself (`source: "program-completion"`,
|
|
324
|
+
never below the completion policy) and consults the planner only when
|
|
325
|
+
something failed. With 0.12.0's repair-in-program this makes a clean run
|
|
326
|
+
**one planner turn**: compile, execute, done — Claude's "0 orchestrator turns
|
|
327
|
+
during execution" for the passing case.
|
|
328
|
+
9. **Rate limits are waited for** (0.12.1): a burst-gated provider parks the
|
|
329
|
+
dispatch in `waiting_for_quota` until the window resets instead of failing
|
|
330
|
+
the run in 4 s, which is what the first 0.12.0 comparison launch did.
|
|
331
|
+
|
|
319
332
|
**Honest limitation.** `itemsFrom` removes the planner *turn*, not the stage
|
|
320
333
|
*barrier*: a verify depending on a data-driven fan-out waits for all items,
|
|
321
334
|
whereas Claude's `pipeline()` overlaps verify-B with fix-C for discovered items
|
|
@@ -327,15 +327,143 @@ three fixes on "non-blocking" nits from verifiers that had returned
|
|
|
327
327
|
|
|
328
328
|
### bullswarm 0.12.0 (installed binary) — same goal, fresh copy `g2-bs-v3`
|
|
329
329
|
|
|
330
|
-
|
|
330
|
+
**First launch, 19:09:48 Z, installed 0.12.0 — failed in 4 s.** Both the
|
|
331
|
+
scout and the orchestrator were `failed_terminal` with `no eligible pool`
|
|
332
|
+
before any process was spawned. Cause: the account's Claude 5-hour window read
|
|
333
|
+
**91 %** (live meter at 19:09:49 Z, reset 22:30 Z) after the two long Opus
|
|
334
|
+
runs, so the pace gate excluded the only pool. Not a 0.12.0 regression (the
|
|
335
|
+
gate pre-dates it) but exactly the class of difference this comparison is
|
|
336
|
+
for: Claude Code, rate-limited mid-workflow, waits and retries; bullswarm
|
|
337
|
+
failed the run with no reset time in the reason. Fixed as **0.12.1** (`beeed94`):
|
|
338
|
+
the runtime parks the dispatch in `waiting_for_quota`, re-reads the meter
|
|
339
|
+
every 60 s, continues when the window resets, and only fails — naming pool,
|
|
340
|
+
usage and reset time — after reset + 10 min grace.
|
|
341
|
+
|
|
342
|
+
**Second launch, 19:28:03 Z, installed 0.12.1, run `wf-mtdcghw0-bfefc7` on a
|
|
343
|
+
fresh pristine `g2-bs-v3` — observed the wait working.** The meter read 95 %
|
|
344
|
+
at launch, so the runtime parked the first dispatch (the scout) in stage
|
|
345
|
+
`waiting_for_quota` with `until 22:40:00 Z` (reset 22:30 + 10 min grace),
|
|
346
|
+
event `dispatch.waiting_for_quota` carrying pool, usage and reset time. It
|
|
347
|
+
re-read the meter every 60 s for **3 h 2 min 14 s** (`waitedMs 10933739`) and
|
|
348
|
+
emitted `dispatch.quota_available` at **22:30:16.979 Z** — 17 s after the
|
|
349
|
+
provider reset — then dispatched the scout with no operator action. Wall-clock
|
|
350
|
+
numbers below therefore exclude this wait (`metrics-bullswarm.mjs` reports
|
|
351
|
+
`quotaWaitSec` and `wallExclWaitSec` separately); the wait is a provider
|
|
352
|
+
constraint, not execution time.
|
|
353
|
+
|
|
354
|
+
**Outcome: `completed`, `verified: true`, 23:18:38 Z.** Timeline (all Z):
|
|
355
|
+
|
|
356
|
+
| when | what |
|
|
357
|
+
|---|---|
|
|
358
|
+
| 22:30:17 | scout dispatched (read-only survey), 197 s |
|
|
359
|
+
| 22:33:34 → 22:39:14 | planner turn 1, **340 s** → one `needs_more_work` program of 14 actions |
|
|
360
|
+
| 22:39:14 | **7 workers started in the same second**: `build-{csv,duration,intervals,lru,semver,slugify}` + `write-docs-index` |
|
|
361
|
+
| 22:44:45 → 22:46:36 | each `verify-<module>` started the moment its own builder finished — per-chain pipelining, no stage barrier (`verify-intervals` was running while `build-slugify` still built) |
|
|
362
|
+
| 22:52:29 | all six module verifies `ok:true` (`verify-slugify` 354 s — see landmine note) → `verify-full-delivery` |
|
|
363
|
+
| 22:59:02 | `verify-full-delivery` **ok:false**: `docs/README.md` missing → runtime spawned `verify-full-delivery-repair-1` (`source: repair-policy`), no planner turn |
|
|
364
|
+
| 23:01:19 → 23:07:58 | repair wrote `docs/README.md` (137 s); re-verify passed (399 s) |
|
|
365
|
+
| 23:07:58 → 23:09:42 | planner turn 2 (105 s): `complete` — **rejected by the runtime**: "missing a successful verification of latest worker verify-full-delivery-repair-1" |
|
|
366
|
+
| 23:09:42 → 23:12:25 | planner turn 3 (163 s): diagnosed the rejection as mechanical, added one read-only `verify-final-acceptance` depending on the repair |
|
|
367
|
+
| 23:12:25 → 23:17:54 | `verify-final-acceptance` ok:true (328 s) |
|
|
368
|
+
| 23:17:54 → 23:18:38 | planner turn 4 (44 s): `complete`, accepted |
|
|
369
|
+
|
|
370
|
+
Numbers (`metrics-bullswarm.mjs`, wait excluded):
|
|
371
|
+
|
|
372
|
+
| metric | 0.11.1 (`g2-bs-v2`) | **0.12.1 (`g2-bs-v3`)** | Claude #2 |
|
|
373
|
+
|---|---|---|---|
|
|
374
|
+
| execution wall | 41 min 14 s | **48 min 21 s** (2 901 s; +10 934 s quota wait) | 58 min |
|
|
375
|
+
| planner turns / seconds | 3 / 667 s (27 %) | **4 / 652 s (22 %)** — turns 2–4 (312 s) plus `verify-final-acceptance` (328 s) exist only because of the rejection bug below | 0 during execution |
|
|
376
|
+
| dispatches / max concurrent | 22 / 6 | **22 / 7** | 24 / ~10 |
|
|
377
|
+
| parallelism (busy ÷ wall) | 2.74 | **2.11** | 3.1 |
|
|
378
|
+
| actions by source | planner 14+7+0 | **planner 14 + 1, repair-policy 1** | script |
|
|
379
|
+
| tests after | 130/130 | **120/120** (52 + 68 new, 6 files ≥ 9 tests each) | 168/168 |
|
|
380
|
+
| existing tests / src | byte-identical / comment-only | **byte-identical / comment-only** (audit-fixture.sh: 0 non-comment line diffs in all 7 src files) | same |
|
|
381
|
+
| deliverables | all | **all** (6 docs pages, `docs/README.md` 6-row index) | all |
|
|
382
|
+
| estimated tokens | — | 201 568 (utf8/4 estimate) | — |
|
|
383
|
+
|
|
384
|
+
What the run showed:
|
|
385
|
+
|
|
386
|
+
1. **The brace landmine is closed (controlled A/B).** `g2-bs-v3` is a pristine copy, so `src/slugify.js` still carries the `@param {{maxLength?: number}}` JSDoc that killed 0.11.1's `verify-slugify` with zero attempts. On 0.12.1 the same verify dispatched, reviewed the artifact with the braces intact, and returned `ok:true` (with three informational concerns, none of which spawned a polish action — the doctrine held).
|
|
387
|
+
2. **The planner compiled a Claude-shaped program on the first turn.** Six independent `build → verify` chains + a parallel docs-index builder + one final gate, each verify carrying `repair {maxRounds: 1}`. The runtime then ran it as a pipeline: verifies started per chain, not after a barrier.
|
|
388
|
+
3. **The repair loop worked live, and paid for a planner mistake.** `write-docs-index` was compiled with `dependsOn: []`, so it launched with the builders and found no `docs/` to index; the worker refused to invent summaries and returned a status note. The final verify caught the missing file and the runtime's repair round fixed it — no planner turn, ~9 min. In Claude's model the same mistake is an authoring error in the script; here it is a compile error by the planner. Neither runtime can catch it deterministically; both recover through verification.
|
|
389
|
+
4. **Runtime bug found: a repair action is never "verified".** `completionEvidenceGaps` accepts a verify as evidence for the latest worker only if `verify.dependsOn` includes that worker. A repair action depends on its verify (the reverse edge), and the verify's post-repair re-run *is* its verification, but the check does not know that — so a clean `complete` was rejected and the run spent 3 more turns and ~11 min proving what it already had. The same check gates 0.13.0's `all-actions-ok` auto-completion, which would have been blocked the same way. Fix: 0.13.1 (below).
|
|
390
|
+
|
|
391
|
+
Take the bug and the dependency slip out and this run is ~29 min of execution with two planner turns — the shape the 0.12.0 design targeted.
|
|
331
392
|
|
|
332
393
|
The originally planned 0.10.9 goal-2 run was dropped at the user's request
|
|
333
394
|
(2026-08-29): the installed latest is the only baseline that matters.
|
|
334
395
|
|
|
335
396
|
## Behaviour differences observed
|
|
336
397
|
|
|
337
|
-
(
|
|
398
|
+
Same goal, same fixture, same model (Opus for every worker and for bullswarm's
|
|
399
|
+
planner; the Claude session's author was Opus too). Read left to right: what
|
|
400
|
+
Claude did, what bullswarm 0.11.1 did on the identical run, and what 0.12.x
|
|
401
|
+
now does about it (built and unit-tested; the live re-run below is the
|
|
402
|
+
confirmation).
|
|
403
|
+
|
|
404
|
+
| Dimension | Claude Code `Workflow` (ultracode) — observed | bullswarm 0.11.1 — observed | bullswarm 0.12.0 / 0.12.1 — built |
|
|
405
|
+
| --- | --- | --- | --- |
|
|
406
|
+
| Who plans, and when | The session author read every file and ran the tests inline (4 min), then wrote **one script** (23 k chars, 5 `agent()` sites). **0 orchestrator turns during the 48 min 51 s of execution.** | The planner compiled the **whole 14-action graph in one decision** (253 s) — but blind: goal text + cwd only, no repo survey, no worker output text in its context. Consulted **3 times** (253 s, 304 s, 110 s) = **27 % of wall**. | Read-only `scout` action before the planner; `outputExcerpt` of every finished action in the planner context; prompt reframed as "compile the goal into a PROGRAM"; planner told it is consulted only at the program boundary. |
|
|
407
|
+
| Item discovery | `pipeline(MODULES, probe, author, verify, fix-loop)` over a known list; when a list is unknown Claude discovers it inline *before* writing the script. | Goal named the six modules → inlined them. Nothing to discover here. | `fanout.itemsFrom: "outputs.<discovery>.outFile"` resolved at run time (+ one bounded read-only extraction retry), so an unknown item count never costs a planner turn. |
|
|
408
|
+
| Parallel width and overlap | **6 concurrent** (= six items, cap 8), mean parallelism 3.1. Per-item pipeline: author-B starts the second probe-B ends; no barriers. | **6 concurrent**, mean parallelism 2.74. Ready-set scheduler: each `verify-<m>` started the second its own `module-<m>` finished; `docs-index` waited for all six by design. | Unchanged for known items. Limitation stays: a verify on a *discovered* fan-out waits for all items (no per-item chain inside a fan-out yet). |
|
|
409
|
+
| Verify → fix | Fix loops **pre-authored in code** (`while (!verdict.ok && rounds < N)`): slugify ×2, intervals ×1, all inside the script; 9 verifies, 3 fixes, 0 planner involvement. | A failed/blocked verify came back to the **planner** (turn 2, 304 s), which authored `slugify-recheck` + `verify-slugify-2`. Round trip ≈ 5 min before the fix even started. | `verify.repair { prompt, maxRounds 1–3 }` — the executor runs `<verify>-repair-<n>` with the concerns verbatim and re-runs the same verify; only still-failing verifies return to the planner. |
|
|
410
|
+
| Passing verifies with nits | Schema-forced `{ok, issues}`; the script fixes only when `!ok`. Nits on passing modules were ignored. | Planner spent **2 of 3 remediation fixes** (`polish-semver`, `polish-lru`) on "non-blocking" notes from verifiers that had returned `ok:true` — an extra ~10 min program round. | Doctrine line: an `ok:true` verify is accepted; its concerns are informational. |
|
|
411
|
+
| Robustness to content | Prompts are JS strings; the runtime substitutes nothing. A parse error in the *script* was caught by the harness and corrected inline in 94 s. | The template renderer parsed **any** `{{…}}` — in a planner prompt *and* in the review artifact it appended. `verify-slugify` died at render time with **0 attempts**, blocked `verify-suite`, cost a planner round, and the fix **rewrote fixture source** (JSDoc) to dodge the bug. | Only a known root + dotted identifiers is a template ref; other double braces are text. `verify` appends the reviewed artifact verbatim, never rendered. |
|
|
412
|
+
| Provider rate limits | Agents retry on API errors; a terminal error resolves the agent to `null`, the script keeps going. The session waits. | Pool burst-gated (5h window 91 %) → the whole run **failed in 4 s** with `no eligible pool`, no reset time named (first 0.12.0 launch, 19:09 Z). | 0.12.1: `waiting_for_quota` stage, meter re-read every 60 s, continue when the window resets; fail only after reset + 10 min grace, naming pool / usage / reset time. |
|
|
413
|
+
| Structured worker output | `schema:` forces a `StructuredOutput` tool call; mismatches retry at the tool layer, so the script never parses prose. | Prose "content gate" (`looksLikeWork`); a bare JSON array answer was rejected as an "announcement"; the verify verdict is the only structured channel. | Content gate accepts JSON; `parseJsonArray` prefers the trailing array; one extraction action when discovery output has no array. A general `outputSchema` on run actions is still open. |
|
|
414
|
+
| Failure semantics | In code: `parallel()` never rejects, a throwing stage drops its item to `null`, `.filter(Boolean)`. | Runtime `onError: continue` per step; failed dependencies block dependents; blocked graph → planner. | Same, plus: a failed scout is non-fatal; fan-out `ok` is a boolean so dependents can wait on a whole fan-out. |
|
|
415
|
+
| Outcome quality (audits, read-only) | 168/168 tests (52 + 116 new); existing tests byte-identical; `src` comment-only; every deliverable present. | 130/130 tests (52 + 78 new); existing tests byte-identical; `src` comment-only; every deliverable present. | — (re-run below) |
|
|
416
|
+
| Time and agents | 58 min end to end; 24 agents. | 41 min end to end; 22 dispatches. Faster because it wrote fewer tests per module (11–15 vs 16–24) and skipped Claude's probe stage — not because it orchestrated better. | — |
|
|
417
|
+
|
|
418
|
+
The short version: after 0.11.x the *shape* already matched (one decision =
|
|
419
|
+
whole graph, six in parallel, per-item verify overlap). What still separated
|
|
420
|
+
the two was everything Claude keeps **inside the program** — reading the repo
|
|
421
|
+
before planning, data-driven fan-out, repair loops, tolerating passing-with-
|
|
422
|
+
nits, tolerating rate limits — which bullswarm was still paying a 2–5 minute
|
|
423
|
+
planner round trip for, or failing on. 0.12.0/0.12.1 move each of those into
|
|
424
|
+
the runtime.
|
|
338
425
|
|
|
339
426
|
## What to change in bullswarm
|
|
340
427
|
|
|
341
|
-
(
|
|
428
|
+
**Shipped in this cycle** (0.12.0 `c1b71a8`, 0.12.1 `beeed94`; all covered by
|
|
429
|
+
unit tests, 295/295):
|
|
430
|
+
|
|
431
|
+
1. Orchestrator as **compiler**: prompt reframed; planner consulted only at the
|
|
432
|
+
program boundary; `programFeatures: ['itemsFrom', 'repair']` advertised.
|
|
433
|
+
2. `fanout.itemsFrom` — data-driven fan-out resolved at execution time, one
|
|
434
|
+
bounded read-only extraction retry, `maxItemsPerExpansion` default 24.
|
|
435
|
+
3. `verify.repair { prompt, maxRounds }` — pre-authored fix loops run by the
|
|
436
|
+
executor.
|
|
437
|
+
4. Fan-out summary artifact; fan-out `ok` is a boolean (count in `succeeded`).
|
|
438
|
+
5. `scout` step before the first program (`--no-scout`), non-fatal.
|
|
439
|
+
6. `outputExcerpt` for every output in the planner context.
|
|
440
|
+
7. Content gate accepts structured JSON; `parseJsonArray` prefers the
|
|
441
|
+
trailing array.
|
|
442
|
+
8. Template refs are grammar-checked; review artifacts are never rendered.
|
|
443
|
+
9. Doctrine: `ok:true` verifies are accepted; concerns are informational.
|
|
444
|
+
10. Burst-gated providers are waited for (`waiting_for_quota`), not failed.
|
|
445
|
+
11. **Program-level completion predicate** (`0bcfd21`, 0.13.0 unreleased while
|
|
446
|
+
the 0.12.1 observation run is in flight): `completion: { when:
|
|
447
|
+
"all-actions-ok", reason }` lets a clean program record its own `complete`
|
|
448
|
+
decision — no final planner turn (110 s here) just to say so. Anything
|
|
449
|
+
failing still returns to the planner.
|
|
450
|
+
|
|
451
|
+
**Still open, in priority order** (each is a measured gap, not a guess):
|
|
452
|
+
|
|
453
|
+
1. **Per-item chains for discovered items.** `itemsFrom` removes the planner
|
|
454
|
+
turn but not the stage barrier; a fan-out whose `stepTemplate` is itself a
|
|
455
|
+
chain (`run → verify(repair)` per item) would give Claude's `pipeline()`
|
|
456
|
+
overlap for unknown item lists too.
|
|
457
|
+
2. **General `outputSchema` on run actions** (Claude's `StructuredOutput`):
|
|
458
|
+
validate a worker's JSON at the dispatch layer and retry once, instead of
|
|
459
|
+
the prose gate plus an extraction action.
|
|
460
|
+
3. **Planner latency.** Every planner turn is a fresh `claude -p --resume`
|
|
461
|
+
process reading a large durable context (253–304 s here vs Claude's ~4 min
|
|
462
|
+
*once*). With repair-in-program and self-completion a clean run is one turn;
|
|
463
|
+
the remaining lever is a smaller planner context (excerpts already budgeted
|
|
464
|
+
at 36 k chars).
|
|
465
|
+
4. **Adversarial verification by default.** Claude's script verified every
|
|
466
|
+
module with a reviewer told to *refute*; bullswarm's verify prompt is
|
|
467
|
+
whatever the planner wrote. A default skeptic framing in the verify wrapper
|
|
468
|
+
is cheap and would have caught nothing extra here — listed for parity, not
|
|
469
|
+
urgency.
|
package/package.json
CHANGED
package/skill/SKILL.md
CHANGED
|
@@ -355,6 +355,12 @@ that expressible without extra turns:
|
|
|
355
355
|
`ok:false`, the executor runs `<verifyId>-repair-<n>` (`source:
|
|
356
356
|
"repair-policy"`) with the verifier's concerns verbatim and re-runs the same
|
|
357
357
|
verify, up to `maxRounds` (1–3), without a planner turn.
|
|
358
|
+
- `completion: { when: "all-actions-ok", reason }` (top level of a program) —
|
|
359
|
+
when every action of the program, repairs included, finishes ok and the
|
|
360
|
+
completion policy is met, the runtime records the `complete` decision itself
|
|
361
|
+
(`source: "program-completion"`, event `decision.auto_completed`) and the run
|
|
362
|
+
ends without another planner turn; a failing action emits
|
|
363
|
+
`decision.completion_predicate_unmet` and the boundary returns to the planner.
|
|
358
364
|
|
|
359
365
|
Every fan-out records a summary artifact as `outputs.<id>.outFile` and a
|
|
360
366
|
boolean `ok` (item count in `succeeded`), so a verify may depend on a fan-out
|
package/src/workflow/decision.js
CHANGED
|
@@ -42,6 +42,9 @@ export const REVIEW_PATH_RE = /^outputs\.([A-Za-z0-9_-]+(?:\[\d+\])?)\.outFile$/
|
|
|
42
42
|
// action whose output ends with a JSON array of items.
|
|
43
43
|
export const ITEMS_FROM_RE = /^outputs\.([A-Za-z0-9_-]+)(?:\.outFile)?$/;
|
|
44
44
|
export const REPAIR_MAX_ROUNDS = 3;
|
|
45
|
+
// Program-level completion predicates a planner may attach to a program so the
|
|
46
|
+
// runtime can record completion itself when every action finishes ok.
|
|
47
|
+
export const COMPLETION_PREDICATES = new Set(['all-actions-ok']);
|
|
45
48
|
|
|
46
49
|
export function looksLikeItemsFromPath(value) {
|
|
47
50
|
return typeof value === 'string' && ITEMS_FROM_RE.test(value.trim());
|
|
@@ -117,6 +120,27 @@ export function validateDecisionProposal(proposal, {
|
|
|
117
120
|
if (currentActionCount + safeActions.length > maxActions) {
|
|
118
121
|
issues.push(`proposal would exceed maxActions=${maxActions}`);
|
|
119
122
|
}
|
|
123
|
+
if (proposal.completion !== undefined) {
|
|
124
|
+
const completion = proposal.completion;
|
|
125
|
+
if (!completion || typeof completion !== 'object' || Array.isArray(completion)) {
|
|
126
|
+
issues.push('completion must be an object {"when":"all-actions-ok","reason":"…"}');
|
|
127
|
+
} else {
|
|
128
|
+
if (!COMPLETION_PREDICATES.has(completion.when)) {
|
|
129
|
+
issues.push(`completion.when must be one of: ${[...COMPLETION_PREDICATES].join(', ')}`);
|
|
130
|
+
}
|
|
131
|
+
if (completion.reason !== undefined && (typeof completion.reason !== 'string' || !completion.reason.trim())) {
|
|
132
|
+
issues.push('completion.reason must be a non-empty string when present');
|
|
133
|
+
}
|
|
134
|
+
for (const key of Object.keys(completion)) {
|
|
135
|
+
if (!['when', 'reason'].includes(key)) issues.push(`completion.${key} is not a planner field`);
|
|
136
|
+
}
|
|
137
|
+
if (proposal.decision !== 'needs_more_work') {
|
|
138
|
+
issues.push('completion is only meaningful on a needs_more_work decision (a program)');
|
|
139
|
+
} else if (!safeActions.some((action) => action?.type === 'verify')) {
|
|
140
|
+
issues.push('a self-completing program must include at least one verify action');
|
|
141
|
+
}
|
|
142
|
+
}
|
|
143
|
+
}
|
|
120
144
|
|
|
121
145
|
const known = new Set(knownActionIds);
|
|
122
146
|
const proposedIds = new Set(safeActions.map((action) => action?.id).filter((id) => typeof id === 'string'));
|
|
@@ -246,5 +270,8 @@ export function validateDecisionProposal(proposal, {
|
|
|
246
270
|
decision: proposal.decision,
|
|
247
271
|
reason: proposal.reason.trim(),
|
|
248
272
|
actions: safeActions.map((action) => ({ ...action, dependsOn: [...(action.dependsOn ?? [])] })),
|
|
273
|
+
...(proposal.completion
|
|
274
|
+
? { completion: { when: proposal.completion.when, ...(proposal.completion.reason ? { reason: proposal.completion.reason.trim() } : {}) } }
|
|
275
|
+
: {}),
|
|
249
276
|
};
|
|
250
277
|
}
|
package/src/workflow/goal.js
CHANGED
|
@@ -20,7 +20,7 @@ export const AUTONOMOUS_ORCHESTRATOR_PROMPT = [
|
|
|
20
20
|
'3. Give workers self-contained prompts with the exact goal, absolute working directory, the files they may edit (and that they must not touch others), the expected artifact, and the exact acceptance command. A worker sees only its own prompt.',
|
|
21
21
|
'4. Assign every action a short kebab-case phase name such as discover, fix, verify-items, or verify-suite. Phases are forward-only: never append new work to a phase that already finished.',
|
|
22
22
|
'5. Use dependsOn only for real data or same-file ordering dependencies. For N known items propose N fix actions and N verify actions (each verify depending only on its own fix) plus one final verify depending on all of them; use fanout with inline items when every item needs the identical prompt. When the item count is unknown, propose a discovery run whose prompt ends with "RETURN ONLY a JSON array of <items>" and a fanout with itemsFrom "outputs.<discovery-id>.outFile", so the runtime fans out the moment discovery finishes.',
|
|
23
|
-
' A verify action with exactly one dependency automatically reviews that dependency artifact (a fan-out artifact summarises every item); you do not need to supply a review path. Give every verify a repair policy {"prompt": "<how to fix what the verifier rejects>", "maxRounds": 1-3} so a rejected verdict is fixed and re-checked inside the program instead of costing another checkpoint.',
|
|
23
|
+
' A verify action with exactly one dependency automatically reviews that dependency artifact (a fan-out artifact summarises every item); you do not need to supply a review path. Give every verify a repair policy {"prompt": "<how to fix what the verifier rejects>", "maxRounds": 1-3} so a rejected verdict is fixed and re-checked inside the program instead of costing another checkpoint. When the program ends with verification that would satisfy the goal, add a top-level "completion": {"when": "all-actions-ok", "reason": "<what a clean run proves>"} so a clean run is recorded as complete without another checkpoint.',
|
|
24
24
|
'6. Recover from a failed action with a new bounded action in a new phase when useful; do not repeat an identical failed plan.',
|
|
25
25
|
'7. Require concrete verification of changed behavior. For code changes, obtain relevant test or inspection evidence before completion.',
|
|
26
26
|
'8. Return complete only when durable outputs prove the original goal and its acceptance checks are satisfied.',
|
package/src/workflow/runner.js
CHANGED
|
@@ -475,6 +475,19 @@ function actionOutputOk(action, outputs) {
|
|
|
475
475
|
: action.status === 'succeeded';
|
|
476
476
|
}
|
|
477
477
|
|
|
478
|
+
// A verify is evidence for a worker when the worker feeds it (dependsOn), or
|
|
479
|
+
// when the worker is that verify's own repair action: the executor re-runs the
|
|
480
|
+
// verify after every repair round, so a verify that ended ok:true after its
|
|
481
|
+
// `<verify>-repair-N` has verified the repair even though the dependency edge
|
|
482
|
+
// points the other way. Without this a clean run's `complete` was rejected as
|
|
483
|
+
// "missing a successful verification of latest worker <verify>-repair-1"
|
|
484
|
+
// (observed on goal-2 run wf-mtdcghw0, 2026-08-28) and cost three planner turns.
|
|
485
|
+
export function verifiesWorker(verify, worker) {
|
|
486
|
+
if ((verify.dependsOn ?? []).includes(worker.id)) return true;
|
|
487
|
+
return (worker.dependsOn ?? []).includes(verify.id)
|
|
488
|
+
&& new RegExp(`^${verify.id.replace(/[.*+?^${}()|[\]\\]/g, '\\$&')}-repair-\\d+$`).test(worker.id);
|
|
489
|
+
}
|
|
490
|
+
|
|
478
491
|
export function completionEvidenceGaps(dynamicActions, policy, outputs = {}) {
|
|
479
492
|
const missing = [];
|
|
480
493
|
const successfulWorkers = dynamicActions.filter(
|
|
@@ -490,7 +503,7 @@ export function completionEvidenceGaps(dynamicActions, policy, outputs = {}) {
|
|
|
490
503
|
&& action.status === 'succeeded'
|
|
491
504
|
&& actionOutputOk(action, outputs)
|
|
492
505
|
&& latestSuccessfulWorker
|
|
493
|
-
&& (action
|
|
506
|
+
&& verifiesWorker(action, latestSuccessfulWorker),
|
|
494
507
|
)) {
|
|
495
508
|
missing.push(latestSuccessfulWorker
|
|
496
509
|
? `a successful verification of latest worker ${latestSuccessfulWorker.id}`
|
|
@@ -1042,6 +1055,59 @@ async function runDecisionLoop({ runtime, gate, phase, state, retryAttempts }) {
|
|
|
1042
1055
|
runtime.emit('kernel.checkpointed', { stage: 'executing', gateId: gate.id });
|
|
1043
1056
|
const executed = await executeActions(proposal.actions);
|
|
1044
1057
|
if (!executed.ok) return executed;
|
|
1058
|
+
|
|
1059
|
+
// Program-level completion: the planner said "if every action of this
|
|
1060
|
+
// program (repairs included) finishes ok, that IS completion". Judge it
|
|
1061
|
+
// by data — no planner turn — but never below the completion policy.
|
|
1062
|
+
if (proposal.completion?.when === 'all-actions-ok') {
|
|
1063
|
+
const programActions = (state.plan.actions ?? []).filter((entry) => entry.decisionSequence === decision.sequence);
|
|
1064
|
+
const ledgerById = new Map((state.actionLedger ?? []).map((entry) => [entry.id, entry]));
|
|
1065
|
+
const failing = programActions
|
|
1066
|
+
.filter((entry) => !ledgerById.has(entry.id) || !actionOutputOk(ledgerById.get(entry.id), state.outputs))
|
|
1067
|
+
.map((entry) => entry.id);
|
|
1068
|
+
const dynamicActions = (state.actionLedger ?? []).filter((action) => action.parentId === gate.id);
|
|
1069
|
+
const gaps = failing.length ? [] : completionEvidenceGaps(dynamicActions, state.orchestration?.completionPolicy, state.outputs);
|
|
1070
|
+
if (!failing.length && !gaps.length) {
|
|
1071
|
+
const verifyIds = programActions.filter((entry) => entry.kind === 'verify').map((entry) => entry.id);
|
|
1072
|
+
const reason = proposal.completion.reason?.trim()
|
|
1073
|
+
|| `Program completed: all ${programActions.length} actions finished ok, verified by ${verifyIds.join(', ')}.`;
|
|
1074
|
+
const auto = {
|
|
1075
|
+
sequence: (state.decisions?.length ?? 0) + 1,
|
|
1076
|
+
gateId: gate.id,
|
|
1077
|
+
decision: 'complete',
|
|
1078
|
+
reason,
|
|
1079
|
+
actions: [],
|
|
1080
|
+
artifact: null,
|
|
1081
|
+
createdAt: new Date().toISOString(),
|
|
1082
|
+
accepted: true,
|
|
1083
|
+
source: 'program-completion',
|
|
1084
|
+
predicate: 'all-actions-ok',
|
|
1085
|
+
programSequence: decision.sequence,
|
|
1086
|
+
};
|
|
1087
|
+
state.decisions.push(auto);
|
|
1088
|
+
runtime.emit('decision.created', auto);
|
|
1089
|
+
runtime.emit('decision.auto_completed', {
|
|
1090
|
+
gateId: gate.id, sequence: auto.sequence, programSequence: decision.sequence,
|
|
1091
|
+
actions: programActions.map((entry) => entry.id), reason,
|
|
1092
|
+
});
|
|
1093
|
+
state.outcome = {
|
|
1094
|
+
status: 'completed',
|
|
1095
|
+
verified: true,
|
|
1096
|
+
bestEffort: false,
|
|
1097
|
+
reason,
|
|
1098
|
+
concerns: [],
|
|
1099
|
+
deliveryActionId: dynamicActions.filter((action) =>
|
|
1100
|
+
action.kind !== 'verify' && actionOutputOk(action, state.outputs)).at(-1)?.id ?? null,
|
|
1101
|
+
source: 'program-completion',
|
|
1102
|
+
};
|
|
1103
|
+
state.outputs[gate.id] = { ...state.outputs[gate.id], ok: true, why: reason, autoCompleted: true };
|
|
1104
|
+
runtime.persist();
|
|
1105
|
+
return { ok: true, why: reason, complete: true, decision: { decision: 'complete', reason, actions: [], completion: proposal.completion } };
|
|
1106
|
+
}
|
|
1107
|
+
runtime.emit('decision.completion_predicate_unmet', {
|
|
1108
|
+
gateId: gate.id, programSequence: decision.sequence, failing, gaps,
|
|
1109
|
+
});
|
|
1110
|
+
}
|
|
1045
1111
|
// Loop intentionally returns to observation and invokes the planner again.
|
|
1046
1112
|
}
|
|
1047
1113
|
}
|
package/src/workflow/runtime.js
CHANGED
|
@@ -1122,7 +1122,7 @@ export class WorkflowRuntime {
|
|
|
1122
1122
|
executionConstraints: {
|
|
1123
1123
|
concurrency: Number(this.state.settings?.concurrency ?? 1) || 1,
|
|
1124
1124
|
readySiblingsRunConcurrently: true,
|
|
1125
|
-
programFeatures: ['itemsFrom', 'repair'],
|
|
1125
|
+
programFeatures: ['itemsFrom', 'repair', 'completion'],
|
|
1126
1126
|
plannerConsultedOnlyAtProgramBoundary: true,
|
|
1127
1127
|
actionTimeoutSec: Number(step.actionDefaults?.timeoutSec ?? step.timeoutSec) || null,
|
|
1128
1128
|
actionTimeoutIsExplicitOptIn: step.actionDefaults?.timeoutSec != null || step.timeoutSec != null,
|
|
@@ -1179,6 +1179,7 @@ export class WorkflowRuntime {
|
|
|
1179
1179
|
'- Unknown item count: never spend a decision to learn how many items there are. Propose a discovery run action whose prompt ends with "RETURN ONLY a JSON array of <items>", plus a fanout with "itemsFrom":"outputs.<discovery-id>.outFile" whose stepTemplate.prompt uses {{item}}. The runtime resolves the list when discovery finishes (with one bounded read-only extraction retry if the output is not a clean array) and fans out immediately.',
|
|
1180
1180
|
'- Verification failures: give each verify a "repair" policy {"prompt":"<how to fix what the verifier rejects>","maxRounds":1-3}. When the verifier returns ok:false, the runtime runs a fix action carrying the verifier concerns verbatim and re-runs the same verify, inside the program. Only verifies still failing after their rounds come back to you.',
|
|
1181
1181
|
'- A verify that returned ok:true is accepted. Its concerns are informational (overlaps, wording nits, "non-blocking" notes): do not spend a program round polishing them unless the goal text itself demands it. Only ok:false verifies are work.',
|
|
1182
|
+
'- Self-completing programs: when the program you propose ends with verification that would satisfy the goal, add a top-level "completion": {"when":"all-actions-ok","reason":"<what a clean run proves>"}. If every action of the program (repairs included) finishes ok and the completion policy is met, the runtime records the completion itself and does not consult you again; anything failing brings the boundary back to you. Use it on every program whose clean run would be the finished goal.',
|
|
1182
1183
|
'- Per-item chains: for N known items propose N focused run actions plus N verify actions, each verify depending only on its own run, so verifying one item overlaps with fixing another; add one final verify depending on all of them. For items discovered at run time use the discovery → fanout → verify shape above.',
|
|
1183
1184
|
'- File ownership: every action prompt must name exactly which files it may edit and state that it must not touch any other file. Two actions that must edit the same file MUST be ordered with dependsOn; never let concurrent actions write the same file.',
|
|
1184
1185
|
'- Self-contained prompts: a worker sees only its own prompt, never this context. Each prompt must state the absolute working directory, what to read, what to change, the exact command that proves success, and what to report back. Prefer many small parallel actions over one large serial one.',
|