bullswarm 0.12.0 → 0.13.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +30 -0
- package/docs/claude-dynamic-workflow-mechanics.md +13 -0
- package/docs/experiments/2026-08-29-ultracode-vs-bullswarm.md +82 -3
- package/package.json +1 -1
- package/skill/SKILL.md +12 -0
- package/src/workflow/decision.js +27 -0
- package/src/workflow/goal.js +1 -1
- package/src/workflow/runner.js +57 -0
- package/src/workflow/runtime.js +112 -5
package/CHANGELOG.md
CHANGED
|
@@ -1,5 +1,35 @@
|
|
|
1
1
|
# bullswarm changelog
|
|
2
2
|
|
|
3
|
+
## 0.13.0 — programs can complete themselves
|
|
4
|
+
|
|
5
|
+
- A planner may attach `completion: { when: "all-actions-ok", reason }` to a
|
|
6
|
+
program (a `needs_more_work` decision that includes at least one verify).
|
|
7
|
+
When every action of that program — repairs included — finishes ok and the
|
|
8
|
+
completion policy is satisfied, the runtime records the `complete` decision
|
|
9
|
+
itself (`source: "program-completion"`, event `decision.auto_completed`) and
|
|
10
|
+
the run ends without another planner turn. Anything failing emits
|
|
11
|
+
`decision.completion_predicate_unmet` (with the failing action ids) and the
|
|
12
|
+
boundary returns to the planner as before. Measured motivation: in the goal-2
|
|
13
|
+
comparison every clean bullswarm run still paid a final 110–250 s planner
|
|
14
|
+
turn just to say "complete"; Claude's script ends when its code says so.
|
|
15
|
+
|
|
16
|
+
## 0.12.1 — a burst-gated provider is waited for, never failed on the spot
|
|
17
|
+
|
|
18
|
+
- `workflow` runs no longer die with `no eligible pool` when every candidate
|
|
19
|
+
pool is burst-gated (provider 5-hour window ≥ 90 % used). The runtime parks
|
|
20
|
+
the dispatch in a new `waiting_for_quota` stage (`state.quotaWait` names the
|
|
21
|
+
pool, its 5h usage and reset time; events `dispatch.waiting_for_quota`,
|
|
22
|
+
`dispatch.quota_available`, `dispatch.quota_wait_expired`), re-reads the
|
|
23
|
+
provider meter every 60 s, and continues the moment the gate lifts. It gives
|
|
24
|
+
up — with the pool, usage and reset time in the failure reason — only after
|
|
25
|
+
the known reset time plus 10 min of grace (5 h when no reset time is known).
|
|
26
|
+
The planner's context is composed after the wait, so it never sees an empty
|
|
27
|
+
pool list. Observed 2026-08-28 19:09 Z: the first 0.12.0 comparison run
|
|
28
|
+
failed in 4 s because the account's Claude 5h window read 91 % (reset
|
|
29
|
+
22:30 Z); Claude Code in the same situation waits on the rate limit.
|
|
30
|
+
Options for embedding callers/tests: `quotaPollMs`, `quotaWaitGraceMs`,
|
|
31
|
+
`quotaWaitUnknownResetMs` on `runWorkflow`; `readMeter` injection.
|
|
32
|
+
|
|
3
33
|
## 0.12.0 — one decision is a whole program
|
|
4
34
|
|
|
5
35
|
Completes the convergence on Claude Code's dynamic-workflow mechanics
|
|
@@ -316,6 +316,19 @@ author and the `Workflow` runtime.
|
|
|
316
316
|
sees `outputs.scout.ok=false` with the reason, and a run where only the
|
|
317
317
|
scout succeeded is `blocked`, never "delivered".
|
|
318
318
|
|
|
319
|
+
8. **Program-level completion** (0.13.0, unreleased at the time of writing) —
|
|
320
|
+
`completion: { when: "all-actions-ok", reason }` on a program. Claude's
|
|
321
|
+
script simply returns when its code is done; bullswarm still spent a final
|
|
322
|
+
planner turn (110–250 s measured) to say `complete` after a clean run. Now
|
|
323
|
+
the runtime records that decision itself (`source: "program-completion"`,
|
|
324
|
+
never below the completion policy) and consults the planner only when
|
|
325
|
+
something failed. With 0.12.0's repair-in-program this makes a clean run
|
|
326
|
+
**one planner turn**: compile, execute, done — Claude's "0 orchestrator turns
|
|
327
|
+
during execution" for the passing case.
|
|
328
|
+
9. **Rate limits are waited for** (0.12.1): a burst-gated provider parks the
|
|
329
|
+
dispatch in `waiting_for_quota` until the window resets instead of failing
|
|
330
|
+
the run in 4 s, which is what the first 0.12.0 comparison launch did.
|
|
331
|
+
|
|
319
332
|
**Honest limitation.** `itemsFrom` removes the planner *turn*, not the stage
|
|
320
333
|
*barrier*: a verify depending on a data-driven fan-out waits for all items,
|
|
321
334
|
whereas Claude's `pipeline()` overlaps verify-B with fix-C for discovered items
|
|
@@ -327,15 +327,94 @@ three fixes on "non-blocking" nits from verifiers that had returned
|
|
|
327
327
|
|
|
328
328
|
### bullswarm 0.12.0 (installed binary) — same goal, fresh copy `g2-bs-v3`
|
|
329
329
|
|
|
330
|
-
|
|
330
|
+
**First launch, 19:09:48 Z, installed 0.12.0 — failed in 4 s.** Both the
|
|
331
|
+
scout and the orchestrator were `failed_terminal` with `no eligible pool`
|
|
332
|
+
before any process was spawned. Cause: the account's Claude 5-hour window read
|
|
333
|
+
**91 %** (live meter at 19:09:49 Z, reset 22:30 Z) after the two long Opus
|
|
334
|
+
runs, so the pace gate excluded the only pool. Not a 0.12.0 regression (the
|
|
335
|
+
gate pre-dates it) but exactly the class of difference this comparison is
|
|
336
|
+
for: Claude Code, rate-limited mid-workflow, waits and retries; bullswarm
|
|
337
|
+
failed the run with no reset time in the reason. Fixed as **0.12.1** (`beeed94`):
|
|
338
|
+
the runtime parks the dispatch in `waiting_for_quota`, re-reads the meter
|
|
339
|
+
every 60 s, continues when the window resets, and only fails — naming pool,
|
|
340
|
+
usage and reset time — after reset + 10 min grace.
|
|
341
|
+
|
|
342
|
+
(re-run on 0.12.1 pending — it starts as soon as the 5h window resets at 22:30 Z)
|
|
331
343
|
|
|
332
344
|
The originally planned 0.10.9 goal-2 run was dropped at the user's request
|
|
333
345
|
(2026-08-29): the installed latest is the only baseline that matters.
|
|
334
346
|
|
|
335
347
|
## Behaviour differences observed
|
|
336
348
|
|
|
337
|
-
(
|
|
349
|
+
Same goal, same fixture, same model (Opus for every worker and for bullswarm's
|
|
350
|
+
planner; the Claude session's author was Opus too). Read left to right: what
|
|
351
|
+
Claude did, what bullswarm 0.11.1 did on the identical run, and what 0.12.x
|
|
352
|
+
now does about it (built and unit-tested; the live re-run below is the
|
|
353
|
+
confirmation).
|
|
354
|
+
|
|
355
|
+
| Dimension | Claude Code `Workflow` (ultracode) — observed | bullswarm 0.11.1 — observed | bullswarm 0.12.0 / 0.12.1 — built |
|
|
356
|
+
| --- | --- | --- | --- |
|
|
357
|
+
| Who plans, and when | The session author read every file and ran the tests inline (4 min), then wrote **one script** (23 k chars, 5 `agent()` sites). **0 orchestrator turns during the 48 min 51 s of execution.** | The planner compiled the **whole 14-action graph in one decision** (253 s) — but blind: goal text + cwd only, no repo survey, no worker output text in its context. Consulted **3 times** (253 s, 304 s, 110 s) = **27 % of wall**. | Read-only `scout` action before the planner; `outputExcerpt` of every finished action in the planner context; prompt reframed as "compile the goal into a PROGRAM"; planner told it is consulted only at the program boundary. |
|
|
358
|
+
| Item discovery | `pipeline(MODULES, probe, author, verify, fix-loop)` over a known list; when a list is unknown Claude discovers it inline *before* writing the script. | Goal named the six modules → inlined them. Nothing to discover here. | `fanout.itemsFrom: "outputs.<discovery>.outFile"` resolved at run time (+ one bounded read-only extraction retry), so an unknown item count never costs a planner turn. |
|
|
359
|
+
| Parallel width and overlap | **6 concurrent** (= six items, cap 8), mean parallelism 3.1. Per-item pipeline: author-B starts the second probe-B ends; no barriers. | **6 concurrent**, mean parallelism 2.74. Ready-set scheduler: each `verify-<m>` started the second its own `module-<m>` finished; `docs-index` waited for all six by design. | Unchanged for known items. Limitation stays: a verify on a *discovered* fan-out waits for all items (no per-item chain inside a fan-out yet). |
|
|
360
|
+
| Verify → fix | Fix loops **pre-authored in code** (`while (!verdict.ok && rounds < N)`): slugify ×2, intervals ×1, all inside the script; 9 verifies, 3 fixes, 0 planner involvement. | A failed/blocked verify came back to the **planner** (turn 2, 304 s), which authored `slugify-recheck` + `verify-slugify-2`. Round trip ≈ 5 min before the fix even started. | `verify.repair { prompt, maxRounds 1–3 }` — the executor runs `<verify>-repair-<n>` with the concerns verbatim and re-runs the same verify; only still-failing verifies return to the planner. |
|
|
361
|
+
| Passing verifies with nits | Schema-forced `{ok, issues}`; the script fixes only when `!ok`. Nits on passing modules were ignored. | Planner spent **2 of 3 remediation fixes** (`polish-semver`, `polish-lru`) on "non-blocking" notes from verifiers that had returned `ok:true` — an extra ~10 min program round. | Doctrine line: an `ok:true` verify is accepted; its concerns are informational. |
|
|
362
|
+
| Robustness to content | Prompts are JS strings; the runtime substitutes nothing. A parse error in the *script* was caught by the harness and corrected inline in 94 s. | The template renderer parsed **any** `{{…}}` — in a planner prompt *and* in the review artifact it appended. `verify-slugify` died at render time with **0 attempts**, blocked `verify-suite`, cost a planner round, and the fix **rewrote fixture source** (JSDoc) to dodge the bug. | Only a known root + dotted identifiers is a template ref; other double braces are text. `verify` appends the reviewed artifact verbatim, never rendered. |
|
|
363
|
+
| Provider rate limits | Agents retry on API errors; a terminal error resolves the agent to `null`, the script keeps going. The session waits. | Pool burst-gated (5h window 91 %) → the whole run **failed in 4 s** with `no eligible pool`, no reset time named (first 0.12.0 launch, 19:09 Z). | 0.12.1: `waiting_for_quota` stage, meter re-read every 60 s, continue when the window resets; fail only after reset + 10 min grace, naming pool / usage / reset time. |
|
|
364
|
+
| Structured worker output | `schema:` forces a `StructuredOutput` tool call; mismatches retry at the tool layer, so the script never parses prose. | Prose "content gate" (`looksLikeWork`); a bare JSON array answer was rejected as an "announcement"; the verify verdict is the only structured channel. | Content gate accepts JSON; `parseJsonArray` prefers the trailing array; one extraction action when discovery output has no array. A general `outputSchema` on run actions is still open. |
|
|
365
|
+
| Failure semantics | In code: `parallel()` never rejects, a throwing stage drops its item to `null`, `.filter(Boolean)`. | Runtime `onError: continue` per step; failed dependencies block dependents; blocked graph → planner. | Same, plus: a failed scout is non-fatal; fan-out `ok` is a boolean so dependents can wait on a whole fan-out. |
|
|
366
|
+
| Outcome quality (audits, read-only) | 168/168 tests (52 + 116 new); existing tests byte-identical; `src` comment-only; every deliverable present. | 130/130 tests (52 + 78 new); existing tests byte-identical; `src` comment-only; every deliverable present. | — (re-run below) |
|
|
367
|
+
| Time and agents | 58 min end to end; 24 agents. | 41 min end to end; 22 dispatches. Faster because it wrote fewer tests per module (11–15 vs 16–24) and skipped Claude's probe stage — not because it orchestrated better. | — |
|
|
368
|
+
|
|
369
|
+
The short version: after 0.11.x the *shape* already matched (one decision =
|
|
370
|
+
whole graph, six in parallel, per-item verify overlap). What still separated
|
|
371
|
+
the two was everything Claude keeps **inside the program** — reading the repo
|
|
372
|
+
before planning, data-driven fan-out, repair loops, tolerating passing-with-
|
|
373
|
+
nits, tolerating rate limits — which bullswarm was still paying a 2–5 minute
|
|
374
|
+
planner round trip for, or failing on. 0.12.0/0.12.1 move each of those into
|
|
375
|
+
the runtime.
|
|
338
376
|
|
|
339
377
|
## What to change in bullswarm
|
|
340
378
|
|
|
341
|
-
(
|
|
379
|
+
**Shipped in this cycle** (0.12.0 `c1b71a8`, 0.12.1 `beeed94`; all covered by
|
|
380
|
+
unit tests, 295/295):
|
|
381
|
+
|
|
382
|
+
1. Orchestrator as **compiler**: prompt reframed; planner consulted only at the
|
|
383
|
+
program boundary; `programFeatures: ['itemsFrom', 'repair']` advertised.
|
|
384
|
+
2. `fanout.itemsFrom` — data-driven fan-out resolved at execution time, one
|
|
385
|
+
bounded read-only extraction retry, `maxItemsPerExpansion` default 24.
|
|
386
|
+
3. `verify.repair { prompt, maxRounds }` — pre-authored fix loops run by the
|
|
387
|
+
executor.
|
|
388
|
+
4. Fan-out summary artifact; fan-out `ok` is a boolean (count in `succeeded`).
|
|
389
|
+
5. `scout` step before the first program (`--no-scout`), non-fatal.
|
|
390
|
+
6. `outputExcerpt` for every output in the planner context.
|
|
391
|
+
7. Content gate accepts structured JSON; `parseJsonArray` prefers the
|
|
392
|
+
trailing array.
|
|
393
|
+
8. Template refs are grammar-checked; review artifacts are never rendered.
|
|
394
|
+
9. Doctrine: `ok:true` verifies are accepted; concerns are informational.
|
|
395
|
+
10. Burst-gated providers are waited for (`waiting_for_quota`), not failed.
|
|
396
|
+
11. **Program-level completion predicate** (`0bcfd21`, 0.13.0 unreleased while
|
|
397
|
+
the 0.12.1 observation run is in flight): `completion: { when:
|
|
398
|
+
"all-actions-ok", reason }` lets a clean program record its own `complete`
|
|
399
|
+
decision — no final planner turn (110 s here) just to say so. Anything
|
|
400
|
+
failing still returns to the planner.
|
|
401
|
+
|
|
402
|
+
**Still open, in priority order** (each is a measured gap, not a guess):
|
|
403
|
+
|
|
404
|
+
1. **Per-item chains for discovered items.** `itemsFrom` removes the planner
|
|
405
|
+
turn but not the stage barrier; a fan-out whose `stepTemplate` is itself a
|
|
406
|
+
chain (`run → verify(repair)` per item) would give Claude's `pipeline()`
|
|
407
|
+
overlap for unknown item lists too.
|
|
408
|
+
2. **General `outputSchema` on run actions** (Claude's `StructuredOutput`):
|
|
409
|
+
validate a worker's JSON at the dispatch layer and retry once, instead of
|
|
410
|
+
the prose gate plus an extraction action.
|
|
411
|
+
3. **Planner latency.** Every planner turn is a fresh `claude -p --resume`
|
|
412
|
+
process reading a large durable context (253–304 s here vs Claude's ~4 min
|
|
413
|
+
*once*). With repair-in-program and self-completion a clean run is one turn;
|
|
414
|
+
the remaining lever is a smaller planner context (excerpts already budgeted
|
|
415
|
+
at 36 k chars).
|
|
416
|
+
4. **Adversarial verification by default.** Claude's script verified every
|
|
417
|
+
module with a reviewer told to *refute*; bullswarm's verify prompt is
|
|
418
|
+
whatever the planner wrote. A default skeptic framing in the verify wrapper
|
|
419
|
+
is cheap and would have caught nothing extra here — listed for parity, not
|
|
420
|
+
urgency.
|
package/package.json
CHANGED
package/skill/SKILL.md
CHANGED
|
@@ -355,6 +355,12 @@ that expressible without extra turns:
|
|
|
355
355
|
`ok:false`, the executor runs `<verifyId>-repair-<n>` (`source:
|
|
356
356
|
"repair-policy"`) with the verifier's concerns verbatim and re-runs the same
|
|
357
357
|
verify, up to `maxRounds` (1–3), without a planner turn.
|
|
358
|
+
- `completion: { when: "all-actions-ok", reason }` (top level of a program) —
|
|
359
|
+
when every action of the program, repairs included, finishes ok and the
|
|
360
|
+
completion policy is met, the runtime records the `complete` decision itself
|
|
361
|
+
(`source: "program-completion"`, event `decision.auto_completed`) and the run
|
|
362
|
+
ends without another planner turn; a failing action emits
|
|
363
|
+
`decision.completion_predicate_unmet` and the boundary returns to the planner.
|
|
358
364
|
|
|
359
365
|
Every fan-out records a summary artifact as `outputs.<id>.outFile` and a
|
|
360
366
|
boolean `ok` (item count in `succeeded`), so a verify may depend on a fan-out
|
|
@@ -363,6 +369,12 @@ finished action reported), and `workflow goal` runs a read-only `scout` action
|
|
|
363
369
|
first (`--no-scout` to skip) so the first program is compiled from a real
|
|
364
370
|
survey of the repository rather than from the goal text alone.
|
|
365
371
|
|
|
372
|
+
A burst-gated provider (5-hour window ≥ 90 % used) is a *wait*, not a
|
|
373
|
+
failure: the run parks in stage `waiting_for_quota` (`state.quotaWait` shows
|
|
374
|
+
the pool, usage and reset time), re-reads the meter every 60 s, and continues
|
|
375
|
+
when the window resets. Only after the reset time plus 10 min of grace does the
|
|
376
|
+
action fail, with the gate named in `why`.
|
|
377
|
+
|
|
366
378
|
Allowed planner decisions are `proceed`, `complete`, `needs_more_work`,
|
|
367
379
|
`retry`, `escalate`, `wait_for_approval`, and `stop`. Expansion decisions must
|
|
368
380
|
contain bounded actions; malformed or over-budget output executes nothing.
|
package/src/workflow/decision.js
CHANGED
|
@@ -42,6 +42,9 @@ export const REVIEW_PATH_RE = /^outputs\.([A-Za-z0-9_-]+(?:\[\d+\])?)\.outFile$/
|
|
|
42
42
|
// action whose output ends with a JSON array of items.
|
|
43
43
|
export const ITEMS_FROM_RE = /^outputs\.([A-Za-z0-9_-]+)(?:\.outFile)?$/;
|
|
44
44
|
export const REPAIR_MAX_ROUNDS = 3;
|
|
45
|
+
// Program-level completion predicates a planner may attach to a program so the
|
|
46
|
+
// runtime can record completion itself when every action finishes ok.
|
|
47
|
+
export const COMPLETION_PREDICATES = new Set(['all-actions-ok']);
|
|
45
48
|
|
|
46
49
|
export function looksLikeItemsFromPath(value) {
|
|
47
50
|
return typeof value === 'string' && ITEMS_FROM_RE.test(value.trim());
|
|
@@ -117,6 +120,27 @@ export function validateDecisionProposal(proposal, {
|
|
|
117
120
|
if (currentActionCount + safeActions.length > maxActions) {
|
|
118
121
|
issues.push(`proposal would exceed maxActions=${maxActions}`);
|
|
119
122
|
}
|
|
123
|
+
if (proposal.completion !== undefined) {
|
|
124
|
+
const completion = proposal.completion;
|
|
125
|
+
if (!completion || typeof completion !== 'object' || Array.isArray(completion)) {
|
|
126
|
+
issues.push('completion must be an object {"when":"all-actions-ok","reason":"…"}');
|
|
127
|
+
} else {
|
|
128
|
+
if (!COMPLETION_PREDICATES.has(completion.when)) {
|
|
129
|
+
issues.push(`completion.when must be one of: ${[...COMPLETION_PREDICATES].join(', ')}`);
|
|
130
|
+
}
|
|
131
|
+
if (completion.reason !== undefined && (typeof completion.reason !== 'string' || !completion.reason.trim())) {
|
|
132
|
+
issues.push('completion.reason must be a non-empty string when present');
|
|
133
|
+
}
|
|
134
|
+
for (const key of Object.keys(completion)) {
|
|
135
|
+
if (!['when', 'reason'].includes(key)) issues.push(`completion.${key} is not a planner field`);
|
|
136
|
+
}
|
|
137
|
+
if (proposal.decision !== 'needs_more_work') {
|
|
138
|
+
issues.push('completion is only meaningful on a needs_more_work decision (a program)');
|
|
139
|
+
} else if (!safeActions.some((action) => action?.type === 'verify')) {
|
|
140
|
+
issues.push('a self-completing program must include at least one verify action');
|
|
141
|
+
}
|
|
142
|
+
}
|
|
143
|
+
}
|
|
120
144
|
|
|
121
145
|
const known = new Set(knownActionIds);
|
|
122
146
|
const proposedIds = new Set(safeActions.map((action) => action?.id).filter((id) => typeof id === 'string'));
|
|
@@ -246,5 +270,8 @@ export function validateDecisionProposal(proposal, {
|
|
|
246
270
|
decision: proposal.decision,
|
|
247
271
|
reason: proposal.reason.trim(),
|
|
248
272
|
actions: safeActions.map((action) => ({ ...action, dependsOn: [...(action.dependsOn ?? [])] })),
|
|
273
|
+
...(proposal.completion
|
|
274
|
+
? { completion: { when: proposal.completion.when, ...(proposal.completion.reason ? { reason: proposal.completion.reason.trim() } : {}) } }
|
|
275
|
+
: {}),
|
|
249
276
|
};
|
|
250
277
|
}
|
package/src/workflow/goal.js
CHANGED
|
@@ -20,7 +20,7 @@ export const AUTONOMOUS_ORCHESTRATOR_PROMPT = [
|
|
|
20
20
|
'3. Give workers self-contained prompts with the exact goal, absolute working directory, the files they may edit (and that they must not touch others), the expected artifact, and the exact acceptance command. A worker sees only its own prompt.',
|
|
21
21
|
'4. Assign every action a short kebab-case phase name such as discover, fix, verify-items, or verify-suite. Phases are forward-only: never append new work to a phase that already finished.',
|
|
22
22
|
'5. Use dependsOn only for real data or same-file ordering dependencies. For N known items propose N fix actions and N verify actions (each verify depending only on its own fix) plus one final verify depending on all of them; use fanout with inline items when every item needs the identical prompt. When the item count is unknown, propose a discovery run whose prompt ends with "RETURN ONLY a JSON array of <items>" and a fanout with itemsFrom "outputs.<discovery-id>.outFile", so the runtime fans out the moment discovery finishes.',
|
|
23
|
-
' A verify action with exactly one dependency automatically reviews that dependency artifact (a fan-out artifact summarises every item); you do not need to supply a review path. Give every verify a repair policy {"prompt": "<how to fix what the verifier rejects>", "maxRounds": 1-3} so a rejected verdict is fixed and re-checked inside the program instead of costing another checkpoint.',
|
|
23
|
+
' A verify action with exactly one dependency automatically reviews that dependency artifact (a fan-out artifact summarises every item); you do not need to supply a review path. Give every verify a repair policy {"prompt": "<how to fix what the verifier rejects>", "maxRounds": 1-3} so a rejected verdict is fixed and re-checked inside the program instead of costing another checkpoint. When the program ends with verification that would satisfy the goal, add a top-level "completion": {"when": "all-actions-ok", "reason": "<what a clean run proves>"} so a clean run is recorded as complete without another checkpoint.',
|
|
24
24
|
'6. Recover from a failed action with a new bounded action in a new phase when useful; do not repeat an identical failed plan.',
|
|
25
25
|
'7. Require concrete verification of changed behavior. For code changes, obtain relevant test or inspection evidence before completion.',
|
|
26
26
|
'8. Return complete only when durable outputs prove the original goal and its acceptance checks are satisfied.',
|
package/src/workflow/runner.js
CHANGED
|
@@ -258,6 +258,10 @@ export async function runWorkflow(opts) {
|
|
|
258
258
|
runDir,
|
|
259
259
|
onEvent: opts.onEvent,
|
|
260
260
|
env: opts.env,
|
|
261
|
+
readMeter: opts.readMeter,
|
|
262
|
+
quotaPollMs: opts.quotaPollMs,
|
|
263
|
+
quotaWaitGraceMs: opts.quotaWaitGraceMs,
|
|
264
|
+
quotaWaitUnknownResetMs: opts.quotaWaitUnknownResetMs,
|
|
261
265
|
});
|
|
262
266
|
state.runner = {
|
|
263
267
|
pid: process.pid,
|
|
@@ -1038,6 +1042,59 @@ async function runDecisionLoop({ runtime, gate, phase, state, retryAttempts }) {
|
|
|
1038
1042
|
runtime.emit('kernel.checkpointed', { stage: 'executing', gateId: gate.id });
|
|
1039
1043
|
const executed = await executeActions(proposal.actions);
|
|
1040
1044
|
if (!executed.ok) return executed;
|
|
1045
|
+
|
|
1046
|
+
// Program-level completion: the planner said "if every action of this
|
|
1047
|
+
// program (repairs included) finishes ok, that IS completion". Judge it
|
|
1048
|
+
// by data — no planner turn — but never below the completion policy.
|
|
1049
|
+
if (proposal.completion?.when === 'all-actions-ok') {
|
|
1050
|
+
const programActions = (state.plan.actions ?? []).filter((entry) => entry.decisionSequence === decision.sequence);
|
|
1051
|
+
const ledgerById = new Map((state.actionLedger ?? []).map((entry) => [entry.id, entry]));
|
|
1052
|
+
const failing = programActions
|
|
1053
|
+
.filter((entry) => !ledgerById.has(entry.id) || !actionOutputOk(ledgerById.get(entry.id), state.outputs))
|
|
1054
|
+
.map((entry) => entry.id);
|
|
1055
|
+
const dynamicActions = (state.actionLedger ?? []).filter((action) => action.parentId === gate.id);
|
|
1056
|
+
const gaps = failing.length ? [] : completionEvidenceGaps(dynamicActions, state.orchestration?.completionPolicy, state.outputs);
|
|
1057
|
+
if (!failing.length && !gaps.length) {
|
|
1058
|
+
const verifyIds = programActions.filter((entry) => entry.kind === 'verify').map((entry) => entry.id);
|
|
1059
|
+
const reason = proposal.completion.reason?.trim()
|
|
1060
|
+
|| `Program completed: all ${programActions.length} actions finished ok, verified by ${verifyIds.join(', ')}.`;
|
|
1061
|
+
const auto = {
|
|
1062
|
+
sequence: (state.decisions?.length ?? 0) + 1,
|
|
1063
|
+
gateId: gate.id,
|
|
1064
|
+
decision: 'complete',
|
|
1065
|
+
reason,
|
|
1066
|
+
actions: [],
|
|
1067
|
+
artifact: null,
|
|
1068
|
+
createdAt: new Date().toISOString(),
|
|
1069
|
+
accepted: true,
|
|
1070
|
+
source: 'program-completion',
|
|
1071
|
+
predicate: 'all-actions-ok',
|
|
1072
|
+
programSequence: decision.sequence,
|
|
1073
|
+
};
|
|
1074
|
+
state.decisions.push(auto);
|
|
1075
|
+
runtime.emit('decision.created', auto);
|
|
1076
|
+
runtime.emit('decision.auto_completed', {
|
|
1077
|
+
gateId: gate.id, sequence: auto.sequence, programSequence: decision.sequence,
|
|
1078
|
+
actions: programActions.map((entry) => entry.id), reason,
|
|
1079
|
+
});
|
|
1080
|
+
state.outcome = {
|
|
1081
|
+
status: 'completed',
|
|
1082
|
+
verified: true,
|
|
1083
|
+
bestEffort: false,
|
|
1084
|
+
reason,
|
|
1085
|
+
concerns: [],
|
|
1086
|
+
deliveryActionId: dynamicActions.filter((action) =>
|
|
1087
|
+
action.kind !== 'verify' && actionOutputOk(action, state.outputs)).at(-1)?.id ?? null,
|
|
1088
|
+
source: 'program-completion',
|
|
1089
|
+
};
|
|
1090
|
+
state.outputs[gate.id] = { ...state.outputs[gate.id], ok: true, why: reason, autoCompleted: true };
|
|
1091
|
+
runtime.persist();
|
|
1092
|
+
return { ok: true, why: reason, complete: true, decision: { decision: 'complete', reason, actions: [], completion: proposal.completion } };
|
|
1093
|
+
}
|
|
1094
|
+
runtime.emit('decision.completion_predicate_unmet', {
|
|
1095
|
+
gateId: gate.id, programSequence: decision.sequence, failing, gaps,
|
|
1096
|
+
});
|
|
1097
|
+
}
|
|
1041
1098
|
// Loop intentionally returns to observation and invokes the planner again.
|
|
1042
1099
|
}
|
|
1043
1100
|
}
|
package/src/workflow/runtime.js
CHANGED
|
@@ -35,6 +35,7 @@ import { aggregateUsage } from '../lib/usage.js';
|
|
|
35
35
|
import { classifyAgentProgress, recordAgentAction } from '../lib/agent-events.js';
|
|
36
36
|
import { deliverSteering } from './steering.js';
|
|
37
37
|
import { resolveDispatchModel } from '../lib/strategy.js';
|
|
38
|
+
import { getMeterReading } from '../meters/registry.js';
|
|
38
39
|
|
|
39
40
|
// Cap how much of each step's output we keep inline in state.json.
|
|
40
41
|
// Persisting full transcripts bloat state.json on long workflows. The
|
|
@@ -45,6 +46,12 @@ const FANOUT_ITEM_EXCERPT_BYTES = 6_000;
|
|
|
45
46
|
// What the planner sees of each action's output (per output / all outputs).
|
|
46
47
|
const PLANNER_EXCERPT_CHARS = 3_000;
|
|
47
48
|
const PLANNER_EXCERPT_TOTAL_CHARS = 36_000;
|
|
49
|
+
// A burst gate (provider 5h window >= 90 % used) is a WAIT, never a hard stop:
|
|
50
|
+
// the runtime parks the dispatch until the window resets (+ grace), re-reading
|
|
51
|
+
// the meter every QUOTA_POLL_MS, and only then fails with the reset time named.
|
|
52
|
+
export const BURST_WAIT_GRACE_MS = 10 * 60_000;
|
|
53
|
+
export const BURST_WAIT_UNKNOWN_RESET_MS = 5 * 3600_000;
|
|
54
|
+
export const QUOTA_POLL_MS = 60_000;
|
|
48
55
|
|
|
49
56
|
export function plannerBudgetContext(budget = {}) {
|
|
50
57
|
const dispatchesUsedBeforePlanner = Number(budget.dispatchesUsed ?? 0);
|
|
@@ -110,6 +117,11 @@ export class WorkflowRuntime {
|
|
|
110
117
|
);
|
|
111
118
|
this.limiter = concap;
|
|
112
119
|
this.parentEnv = opts.env ?? process.env;
|
|
120
|
+
// Meter refresh used while waiting on a burst gate; tests inject a fake.
|
|
121
|
+
this.readMeter = opts.readMeter ?? ((name) => getMeterReading(name, { force: true }));
|
|
122
|
+
this.quotaPollMs = Math.max(10, Number(opts.quotaPollMs ?? this.state?.settings?.quotaPollMs ?? QUOTA_POLL_MS));
|
|
123
|
+
this.quotaWaitGraceMs = Math.max(0, Number(opts.quotaWaitGraceMs ?? BURST_WAIT_GRACE_MS));
|
|
124
|
+
this.quotaWaitUnknownResetMs = Math.max(0, Number(opts.quotaWaitUnknownResetMs ?? BURST_WAIT_UNKNOWN_RESET_MS));
|
|
113
125
|
// Counters used for planner-visible advisory budgeting.
|
|
114
126
|
this.state.attempts ??= [];
|
|
115
127
|
this.state.actionLedger ??= [];
|
|
@@ -244,9 +256,13 @@ export class WorkflowRuntime {
|
|
|
244
256
|
* before selection (R8).
|
|
245
257
|
*/
|
|
246
258
|
async dispatch(step, taskText, targetDir, paths, opts = {}) {
|
|
259
|
+
if (this.state.cancelRequested) return { ok: false, keepOnClaude: false, why: 'workflow cancellation requested', pick: { pool: null }, meta: {} };
|
|
260
|
+
const effortTier = step.effort ?? ({ analyze: 'high', build: 'medium', chore: 'low' }[step.lane ?? 'chore']);
|
|
261
|
+
// A burst-gated provider is waited for (outside the concurrency permit),
|
|
262
|
+
// never failed on the spot.
|
|
263
|
+
await this.awaitBurstRoom(step, effortTier, this.actionId(step, opts));
|
|
247
264
|
if (this.state.cancelRequested) return { ok: false, keepOnClaude: false, why: 'workflow cancellation requested', pick: { pool: null }, meta: {} };
|
|
248
265
|
return this.limiter.runWith(async () => {
|
|
249
|
-
const effortTier = step.effort ?? ({ analyze: 'high', build: 'medium', chore: 'low' }[step.lane ?? 'chore']);
|
|
250
266
|
const attemptPools = this.preparePools(step, effortTier);
|
|
251
267
|
let lastVerdict = null;
|
|
252
268
|
const retryAllowance = Math.max(0, Math.min(Number(opts.retryAttempts ?? 1), 3));
|
|
@@ -269,10 +285,13 @@ export class WorkflowRuntime {
|
|
|
269
285
|
});
|
|
270
286
|
if (!route.pick) {
|
|
271
287
|
if (lastVerdict) return lastVerdict;
|
|
288
|
+
const stillGated = attemptPools.length ? [] : this.burstGatedPoolsFor(step, effortTier);
|
|
272
289
|
const refused = {
|
|
273
290
|
ok: false,
|
|
274
291
|
keepOnClaude: false,
|
|
275
|
-
why:
|
|
292
|
+
why: stillGated.length
|
|
293
|
+
? `no eligible pool: every candidate is burst-gated (${WorkflowRuntime.describeBurstGate(stillGated)}) and the wait for the window expired`
|
|
294
|
+
: `no eligible pool (${route.why})`,
|
|
276
295
|
pick: { pool: null },
|
|
277
296
|
meta: { exitCode: null },
|
|
278
297
|
};
|
|
@@ -656,7 +675,7 @@ export class WorkflowRuntime {
|
|
|
656
675
|
});
|
|
657
676
|
}
|
|
658
677
|
|
|
659
|
-
preparePools(step, effortTier = step.effort ?? ({ analyze: 'high', build: 'medium', chore: 'low' }[step.lane ?? 'chore'])) {
|
|
678
|
+
preparePools(step, effortTier = step.effort ?? ({ analyze: 'high', build: 'medium', chore: 'low' }[step.lane ?? 'chore']), { ignoreBurstGate = false } = {}) {
|
|
660
679
|
// Fresh eligible list per dispatch: enabled, not quarantined, and
|
|
661
680
|
// not currently burst-gated (R8). Quarantine has been applied to
|
|
662
681
|
// pool.quarantine by the live buildPools pass; if a previous
|
|
@@ -667,7 +686,7 @@ export class WorkflowRuntime {
|
|
|
667
686
|
if (pool.quarantine && !isQuarantined(pool, now)) pool.quarantine = null;
|
|
668
687
|
}
|
|
669
688
|
return this.pools.filter(
|
|
670
|
-
(p) => p.enabled !== false && !isQuarantined(p, now) && p.burstGate !== true &&
|
|
689
|
+
(p) => p.enabled !== false && !isQuarantined(p, now) && (ignoreBurstGate || p.burstGate !== true) &&
|
|
671
690
|
(step.pool == null || p.name === step.pool) &&
|
|
672
691
|
!(step.avoidPools ?? []).includes(p.name) &&
|
|
673
692
|
(step.requiresCapabilities ?? []).every((capability) =>
|
|
@@ -682,6 +701,92 @@ export class WorkflowRuntime {
|
|
|
682
701
|
}).filter((pool) => pool.modelPolicy.eligible);
|
|
683
702
|
}
|
|
684
703
|
|
|
704
|
+
/** Pools that would serve this step if they were not burst-gated. */
|
|
705
|
+
burstGatedPoolsFor(step, effortTier) {
|
|
706
|
+
return this.preparePools(step, effortTier, { ignoreBurstGate: true }).filter((p) => p.burstGate === true);
|
|
707
|
+
}
|
|
708
|
+
|
|
709
|
+
static describeBurstGate(pools) {
|
|
710
|
+
return pools.map((p) => {
|
|
711
|
+
const w = p.meterSnapshot?.five_hour ?? {};
|
|
712
|
+
const used = Number.isFinite(w.utilization) ? `${Math.round(w.utilization)}% used` : 'usage unknown';
|
|
713
|
+
const resets = w.resets_at ? `resets ${new Date(w.resets_at).toISOString().replace(/\.\d{3}Z$/, 'Z')}` : 'reset time unknown';
|
|
714
|
+
return `${p.name} 5h window ${used}, ${resets}`;
|
|
715
|
+
}).join('; ');
|
|
716
|
+
}
|
|
717
|
+
|
|
718
|
+
/**
|
|
719
|
+
* If every pool that could serve `step` is burst-gated, wait for the gate
|
|
720
|
+
* to lift instead of failing the action: the provider window resets at a
|
|
721
|
+
* known time, the run is durable, and a failed run costs more than a late
|
|
722
|
+
* one. Re-reads the meters every quotaPollMs; gives up (so the caller fails
|
|
723
|
+
* with a clear reason) only after the latest known reset + grace, or after
|
|
724
|
+
* BURST_WAIT_UNKNOWN_RESET_MS when no reset time is known.
|
|
725
|
+
*/
|
|
726
|
+
async awaitBurstRoom(step, effortTier, actionId = step.id) {
|
|
727
|
+
const gatedOnly = () => {
|
|
728
|
+
if (this.preparePools(step, effortTier).length) return null;
|
|
729
|
+
const gated = this.burstGatedPoolsFor(step, effortTier);
|
|
730
|
+
return gated.length ? gated : null;
|
|
731
|
+
};
|
|
732
|
+
let gated = gatedOnly();
|
|
733
|
+
if (!gated) return { waited: false };
|
|
734
|
+
const resetTimes = gated.map((p) => Date.parse(p.meterSnapshot?.five_hour?.resets_at ?? '')).filter(Number.isFinite);
|
|
735
|
+
const startedAt = Date.now();
|
|
736
|
+
const deadline = resetTimes.length
|
|
737
|
+
? Math.max(...resetTimes) + this.quotaWaitGraceMs
|
|
738
|
+
: startedAt + this.quotaWaitUnknownResetMs;
|
|
739
|
+
const previousStage = this.state.stage;
|
|
740
|
+
this.state.stage = 'waiting_for_quota';
|
|
741
|
+
this.state.quotaWait = {
|
|
742
|
+
actionId,
|
|
743
|
+
since: new Date(startedAt).toISOString(),
|
|
744
|
+
until: new Date(deadline).toISOString(),
|
|
745
|
+
pools: gated.map((p) => ({
|
|
746
|
+
name: p.name,
|
|
747
|
+
fiveHourUsedPct: p.meterSnapshot?.five_hour?.utilization ?? null,
|
|
748
|
+
resetsAt: p.meterSnapshot?.five_hour?.resets_at ?? null,
|
|
749
|
+
})),
|
|
750
|
+
};
|
|
751
|
+
this.emit('dispatch.waiting_for_quota', { actionId, ...this.state.quotaWait, detail: WorkflowRuntime.describeBurstGate(gated) });
|
|
752
|
+
const cancelled = () => {
|
|
753
|
+
if (this.state.cancelRequested) return true;
|
|
754
|
+
try { return JSON.parse(readFileSync(join(this.runDir, 'state.json'), 'utf8')).cancelRequested === true; } catch { return false; }
|
|
755
|
+
};
|
|
756
|
+
let lifted = false;
|
|
757
|
+
while (Date.now() < deadline && !cancelled()) {
|
|
758
|
+
await new Promise((resolve) => setTimeout(resolve, Math.min(this.quotaPollMs, Math.max(1, deadline - Date.now()))));
|
|
759
|
+
for (const gatedView of gated) {
|
|
760
|
+
// preparePools hands out copies; the gate lives on the pool itself.
|
|
761
|
+
const pool = this.pools.find((p) => p.name === gatedView.name) ?? gatedView;
|
|
762
|
+
let reading = null;
|
|
763
|
+
try { reading = await this.readMeter(pool.name); } catch { reading = null; }
|
|
764
|
+
// Only a real provider snapshot may open or keep the gate; a pool
|
|
765
|
+
// without a meter reader ('none') leaves the gate as it was.
|
|
766
|
+
if (!reading?.snapshot) continue;
|
|
767
|
+
pool.burstGate = reading.burstGate === true;
|
|
768
|
+
pool.meterSnapshot = reading.snapshot;
|
|
769
|
+
if (reading.pacing) {
|
|
770
|
+
pool.usedPct = reading.pacing.usedPct ?? pool.usedPct;
|
|
771
|
+
pool.elapsedPct = reading.pacing.elapsedPct ?? pool.elapsedPct;
|
|
772
|
+
pool.pace = reading.pacing.surplus ?? pool.pace;
|
|
773
|
+
}
|
|
774
|
+
}
|
|
775
|
+
gated = gatedOnly();
|
|
776
|
+
if (!gated) { lifted = true; break; }
|
|
777
|
+
}
|
|
778
|
+
const waitedMs = Date.now() - startedAt;
|
|
779
|
+
delete this.state.quotaWait;
|
|
780
|
+
this.state.stage = previousStage;
|
|
781
|
+
if (lifted) {
|
|
782
|
+
this.emit('dispatch.quota_available', { actionId, waitedMs });
|
|
783
|
+
} else {
|
|
784
|
+
this.emit('dispatch.quota_wait_expired', { actionId, waitedMs, detail: gated ? WorkflowRuntime.describeBurstGate(gated) : null, cancelled: cancelled() });
|
|
785
|
+
}
|
|
786
|
+
this.persist();
|
|
787
|
+
return { waited: true, lifted, waitedMs, gated };
|
|
788
|
+
}
|
|
789
|
+
|
|
685
790
|
appendDecision(step, poolName, verdict, paths, routing = null) {
|
|
686
791
|
try {
|
|
687
792
|
const coreState = loadState(this.bullswarmDir);
|
|
@@ -934,6 +1039,7 @@ export class WorkflowRuntime {
|
|
|
934
1039
|
|
|
935
1040
|
async runDecision(step, scope, opts = {}) {
|
|
936
1041
|
this.enforceRequiredInputs(step.id);
|
|
1042
|
+
await this.awaitBurstRoom(step, step.effort ?? 'high', step.id);
|
|
937
1043
|
const deliveredSteering = deliverSteering(this.state, this.runDir);
|
|
938
1044
|
for (const steering of deliveredSteering) {
|
|
939
1045
|
this.emit('steering.delivered', {
|
|
@@ -1016,7 +1122,7 @@ export class WorkflowRuntime {
|
|
|
1016
1122
|
executionConstraints: {
|
|
1017
1123
|
concurrency: Number(this.state.settings?.concurrency ?? 1) || 1,
|
|
1018
1124
|
readySiblingsRunConcurrently: true,
|
|
1019
|
-
programFeatures: ['itemsFrom', 'repair'],
|
|
1125
|
+
programFeatures: ['itemsFrom', 'repair', 'completion'],
|
|
1020
1126
|
plannerConsultedOnlyAtProgramBoundary: true,
|
|
1021
1127
|
actionTimeoutSec: Number(step.actionDefaults?.timeoutSec ?? step.timeoutSec) || null,
|
|
1022
1128
|
actionTimeoutIsExplicitOptIn: step.actionDefaults?.timeoutSec != null || step.timeoutSec != null,
|
|
@@ -1073,6 +1179,7 @@ export class WorkflowRuntime {
|
|
|
1073
1179
|
'- Unknown item count: never spend a decision to learn how many items there are. Propose a discovery run action whose prompt ends with "RETURN ONLY a JSON array of <items>", plus a fanout with "itemsFrom":"outputs.<discovery-id>.outFile" whose stepTemplate.prompt uses {{item}}. The runtime resolves the list when discovery finishes (with one bounded read-only extraction retry if the output is not a clean array) and fans out immediately.',
|
|
1074
1180
|
'- Verification failures: give each verify a "repair" policy {"prompt":"<how to fix what the verifier rejects>","maxRounds":1-3}. When the verifier returns ok:false, the runtime runs a fix action carrying the verifier concerns verbatim and re-runs the same verify, inside the program. Only verifies still failing after their rounds come back to you.',
|
|
1075
1181
|
'- A verify that returned ok:true is accepted. Its concerns are informational (overlaps, wording nits, "non-blocking" notes): do not spend a program round polishing them unless the goal text itself demands it. Only ok:false verifies are work.',
|
|
1182
|
+
'- Self-completing programs: when the program you propose ends with verification that would satisfy the goal, add a top-level "completion": {"when":"all-actions-ok","reason":"<what a clean run proves>"}. If every action of the program (repairs included) finishes ok and the completion policy is met, the runtime records the completion itself and does not consult you again; anything failing brings the boundary back to you. Use it on every program whose clean run would be the finished goal.',
|
|
1076
1183
|
'- Per-item chains: for N known items propose N focused run actions plus N verify actions, each verify depending only on its own run, so verifying one item overlaps with fixing another; add one final verify depending on all of them. For items discovered at run time use the discovery → fanout → verify shape above.',
|
|
1077
1184
|
'- File ownership: every action prompt must name exactly which files it may edit and state that it must not touch any other file. Two actions that must edit the same file MUST be ordered with dependsOn; never let concurrent actions write the same file.',
|
|
1078
1185
|
'- Self-contained prompts: a worker sees only its own prompt, never this context. Each prompt must state the absolute working directory, what to read, what to change, the exact command that proves success, and what to report back. Prefer many small parallel actions over one large serial one.',
|