bullswarm 0.14.0 → 0.15.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +54 -0
- package/docs/experiments/2026-08-29-dogfood-bullswarm-builds-bullswarm.md +171 -0
- package/package.json +1 -1
- package/skill/SKILL.md +17 -1
- package/src/lib/claude-accounts.js +202 -0
- package/src/lib/config.js +3 -1
- package/src/lib/watch.js +6 -1
- package/src/meters/claude.js +18 -9
- package/src/meters/registry.js +21 -1
- package/src/workflow/cli.js +3 -1
- package/src/workflow/dashboard.js +49 -17
- package/src/workflow/decision.js +17 -5
- package/src/workflow/draft.js +2 -6
- package/src/workflow/fsjson.js +41 -0
- package/src/workflow/goal.js +3 -3
- package/src/workflow/runner.js +75 -9
- package/src/workflow/runs-cli.js +14 -6
- package/src/workflow/runtime.js +33 -4
- package/src/workflow/short-id.js +3 -2
package/CHANGELOG.md
CHANGED
|
@@ -1,5 +1,59 @@
|
|
|
1
1
|
# bullswarm changelog
|
|
2
2
|
|
|
3
|
+
## 0.15.0 — extra Claude Code logins as separate pools
|
|
4
|
+
|
|
5
|
+
- Claude Code extra logins (`~/.claude-<slug>` / `$CLAUDE_CONFIG_DIR`) become
|
|
6
|
+
their own pools (`claude-code:<slug>`), metered and spawned with
|
|
7
|
+
`CLAUDE_CONFIG_DIR` set. Discovery is dynamic from the filesystem; there
|
|
8
|
+
is no hardcoded extra-profile list. The spawn command for each profile is
|
|
9
|
+
`CLAUDE_CONFIG_DIR=<dir> claude`.
|
|
10
|
+
|
|
11
|
+
## 0.14.1 — the TUI survives its writer; steering lands or expires truthfully
|
|
12
|
+
|
|
13
|
+
Proven on a goal-3 re-run (`d7xyg2`): 1 planner turn / 269 s (baseline 0.13.1: 1 / 294 s), planner context 6.2 k chars (from 32.7 k), auto-completed, deliverable verified, zero observation crashes — `docs/experiments/2026-08-29-dogfood-bullswarm-builds-bullswarm.md`.
|
|
14
|
+
|
|
15
|
+
- Workflow `state.json`/`report.json`/`workflow.json` writes are atomic
|
|
16
|
+
(temp + rename, new `src/workflow/fsjson.js`): a concurrent reader can never
|
|
17
|
+
observe a half-written file. Earned: `workflow tui` crashed with
|
|
18
|
+
"Unterminated string in JSON at position 138968" parsing `state.json`
|
|
19
|
+
mid-write (observed twice, 2026-08-29).
|
|
20
|
+
- Observation readers tolerate torn or missing JSON: the TUI keeps painting
|
|
21
|
+
the last good frame of the same run, and a render or key-handler error is
|
|
22
|
+
shown in the message line instead of killing the process and stranding the
|
|
23
|
+
terminal in alt-screen raw mode. Mutating commands (stop, approval) retry
|
|
24
|
+
the read once and then refuse loudly instead of silently dropping the
|
|
25
|
+
operator's command. `runs delete` treats an unreadable `state.json` as
|
|
26
|
+
ongoing (refuses without `--force`) rather than deleting a possibly-live run.
|
|
27
|
+
- An action being re-run (repair round, re-verify, schema retry) reads as
|
|
28
|
+
`running` and its phase as `active` even when its previous round recorded
|
|
29
|
+
`ok:false`; a failed mark now means failed-and-not-being-retried. (User
|
|
30
|
+
report: the TUI showed ✗ "2/2 complete" beside a live spinner.)
|
|
31
|
+
- Pending operator steering defers program self-completion: a clean program
|
|
32
|
+
with `completion: all-actions-ok` returns to the planner gate (event
|
|
33
|
+
`decision.completion_deferred`), which delivers the steer — instead of
|
|
34
|
+
auto-completing and silently discarding it (defect observed live:
|
|
35
|
+
0 `steering.delivered` events for a queued steer). Steering that can no
|
|
36
|
+
longer reach any gate is marked `expired_undelivered` with event
|
|
37
|
+
`steering.expired` at the terminal transition; interrupted runs keep their
|
|
38
|
+
queue for the resumed run's next gate.
|
|
39
|
+
- Resume re-runs an action the interruption cancelled mid-flight instead of
|
|
40
|
+
re-planning around a phantom failure: cancelled actions and the dependents
|
|
41
|
+
blocked only by them are reopened (event `action.reopened`) and the accepted
|
|
42
|
+
program continues from where it stopped. Observed on a SIGTERM-interrupted
|
|
43
|
+
run: 1 cancelled action → 4 "blocked" → a spurious planner turn.
|
|
44
|
+
- Planner contract: a verify with several `dependsOn` must set `review`
|
|
45
|
+
(rule 4); the program's last worker must be covered by a successful verify
|
|
46
|
+
(rule 8) — both were the causes of extra planner gates on the goal-3 proof
|
|
47
|
+
run. A corrective turn's `validationFeedback.rejectedResponseExcerpt` is
|
|
48
|
+
capped at 2 000 chars, and the rejected proposal is resent as a skeleton
|
|
49
|
+
(ids, shapes, dependsOn; prompts elided) — the planner's thread already
|
|
50
|
+
holds it verbatim.
|
|
51
|
+
- A verify without `review` is no longer grounds to reject a whole program:
|
|
52
|
+
it reviews its single (or last) dependency's artifact, or audits the
|
|
53
|
+
repository directly when it has no `dependsOn` (`reviewScope: repository`).
|
|
54
|
+
Observed on two proof runs: a 9-action program bounced for one field,
|
|
55
|
+
costing a 5-minute correction turn each time.
|
|
56
|
+
|
|
3
57
|
## 0.14.0 — structured worker output, compact planner contract
|
|
4
58
|
|
|
5
59
|
- A verify whose reply cannot be parsed as the verdict JSON gets ONE bounded
|
|
@@ -0,0 +1,171 @@
|
|
|
1
|
+
# Dogfood 2026-08-29 — bullswarm builds bullswarm (outputSchema + planner refactor)
|
|
2
|
+
|
|
3
|
+
Observation log of the two dogfood runs that produced 0.14.0, kept verbatim
|
|
4
|
+
from the driving session's notes; the post-run defects below produced 0.14.1.
|
|
5
|
+
|
|
6
|
+
|
|
7
|
+
Goal file: `goal4.txt` (7 deliverables, single-implementer constraint). Repo branch: `feat/output-schema`.
|
|
8
|
+
Runtime: local worktree `bullswarm-rt` pinned at main (user: redispatch with local runtime, keep committing; no npm wait).
|
|
9
|
+
Home: default `~/.bullswarm` (heterogeneous pools) — user can `bullswarm workflow tui <shortId>`.
|
|
10
|
+
|
|
11
|
+
## Attempt 1 — 01:43:37 Z, installed 0.13.1 — failed at validation, nothing ran
|
|
12
|
+
`autonomous workflow invalid (nothing ran): phases[0].steps[0](scout): template ref "{{outputs.x.data.field}}" cannot resolve`.
|
|
13
|
+
Cause: goal text spliced into the scout prompt; validator parsed a quoted ref in user text.
|
|
14
|
+
Fix (bullswarm defect #1): goal → declared `inputs.goal`, inserted at render time; unresolved grammar-valid refs are
|
|
15
|
+
left literal + `template.unresolved_ref` event instead of fatal. Commit `7badea3`, released 0.13.2 (`a0f0965`).
|
|
16
|
+
Verified: scout task file of attempt 3 contains `{{outputs.x.data.field}}` verbatim; only `{{inputs.goal}}` resolved.
|
|
17
|
+
|
|
18
|
+
## Attempt 2 — 01:47:37 Z — launcher bug (mine, not bullswarm): zsh does not word-split `$BS="node path"`. Fixed script.
|
|
19
|
+
|
|
20
|
+
## Attempt 3 — 01:48:19 Z, runtime 0.13.2 @ a0f0965 — run `wf-mtdq1l9v-ed22fe` / `75t4n2`
|
|
21
|
+
- 01:48:22 scout → pool `opencode2` model `kaihk/gpt-5.6-luna`. Routing: "most-behind capable pool (surplus 0)";
|
|
22
|
+
candidates opencode2 pace 0 (unmetered) > grok −7.5 > claude-code −9.4 > codex −49.1.
|
|
23
|
+
Observation: an unmetered pool reads as exactly on pace (0) and therefore outranks every metered pool that is
|
|
24
|
+
ahead of pace. Not a crash, but quota-unknown pools capture all work whenever metered pools are burning ahead.
|
|
25
|
+
- 01:48:22 → 01:49:54 scout ok (92 s, opencode2).
|
|
26
|
+
- 01:50:06 → 02:00:40 planner turn 1 (**~634 s** — roughly 2× the goal-2/3 turns; goal text is 6.5 k chars and the
|
|
27
|
+
program is 12 actions). Decision: `needs_more_work`, 12 actions, `completion: all-actions-ok` attached.
|
|
28
|
+
Program honours the single-implementer constraint: `impl-src` (all src) ∥ `docs`; `tests-schema`/`tests-adaptive`/
|
|
29
|
+
`tests-gaps` + `verify-impl(repair 2)` fan out after impl-src; per-writer verifies with repair; `final-report` →
|
|
30
|
+
`verify-suite(repair 2)`. Note: `verify-impl`'s repair may edit src while the test writers read it — accepted risk,
|
|
31
|
+
`verify-suite` runs the whole suite at the end.
|
|
32
|
+
- 02:00:31 `impl-src` and `docs` → opencode2 `kaihk/gpt-5.6-luna` (medium effort, "most-behind capable pool (surplus 0)").
|
|
33
|
+
All build work lands on the unmetered pool while claude-code/grok/codex are ahead of pace. Orchestrator stayed on
|
|
34
|
+
claude-code (pinned).
|
|
35
|
+
- 02:00:31 → 02:07:09 `impl-src` ok (398 s, opencode2); `docs` ok (140 s). 02:07:11 five actions started at once
|
|
36
|
+
(3 test writers + verify-impl + verify-docs).
|
|
37
|
+
- 02:08:20 `verify-impl` ok:false → `verify-impl-repair-1` (concerns were concrete and correct: schema JSON omitted from
|
|
38
|
+
the retry task text; fan-out item resume skipped schemaOk; no combined run+fanout skeleton). 02:08:36 `verify-docs`
|
|
39
|
+
ok:false → `verify-docs-repair-1` (doc claimed fan-out persists top-level data/schemaOk; doc claimed static validator
|
|
40
|
+
rejects outputSchema on verify but only decision.js did). Repairs 121 s / 124 s. Both are repair-loop live uses; if the
|
|
41
|
+
program then self-completes this is the first live exercise of the 0.13.1 fix path.
|
|
42
|
+
- 02:12 → 02:18 second round of repairs: `verify-impl-repair-2` (escalation could allow a 2nd schema retry — legit design
|
|
43
|
+
nit; "schema.js untracked so absent from git diff --stat" — verifier misreading), `verify-adaptive-tests-repair-1/2`
|
|
44
|
+
(rejected twice for the same reason: "diff is not purely additive: modifies an existing assertion" — the
|
|
45
|
+
programFeatures assertion HAD to change; an unrepairable process criterion). Both verifies ended ok:false after
|
|
46
|
+
maxRounds → `final-report`/`verify-suite` blocked (`failed_terminal: dynamic actions blocked by failed or unresolved
|
|
47
|
+
dependencies`) → program boundary → planner turn 2 at 02:18:52.
|
|
48
|
+
Observation (behaviour): a verifier that judges process criteria ("purely additive", "untracked file") instead of the
|
|
49
|
+
goal's acceptance checks produces rejections no repair can satisfy; two repair rounds (~10 min) were spent before the
|
|
50
|
+
boundary. Candidate for the prompt audit: verify doctrine "ok:false only for failed acceptance checks; process
|
|
51
|
+
observations are concerns" and/or runtime: identical concerns after a repair → boundary immediately.
|
|
52
|
+
- 02:18:52 → ~02:25:20 planner turn 2 (~390 s): `needs_more_work`, 3 sequential actions — `recheck-src` (verify, repair 2,
|
|
53
|
+
depends on verify-impl-repair-2) → `full-suite-report` (run) → `verify-final` (verify, repair 2). Reason correctly
|
|
54
|
+
notes no whole-suite evidence existed after the repairs (impl-src's npm test predated them).
|
|
55
|
+
- 02:25:52 → 02:32:36 follow-up program: `recheck-src` ok:false → `recheck-src-repair-1` → re-verify ok; `full-suite-report`
|
|
56
|
+
ok; `verify-final` ok → **`decision.auto_completed` (program-completion)** at 02:32:39.
|
|
57
|
+
## Result (attempt 3)
|
|
58
|
+
- Wall **44 min 17 s** (2 654 s), 28 dispatches (26 opencode2 + 2 planner on claude-code), max concurrent 5,
|
|
59
|
+
parallelism 1.5, planner 2 turns / 1 045 s (39 % — turn 1 alone 637 s), repairs 6 (2 repaired ok, 4 re-verify
|
|
60
|
+
rejected), tokens ≈ 108 k (estimate).
|
|
61
|
+
- Deliverable: 10 files changed + 2 new (489+/38−); `npm test` **318/318** (299 + 19) on my own run; validator module,
|
|
62
|
+
decision/validate/runtime/runner/template changes all present; docs + changelog written.
|
|
63
|
+
- My review: sound; one robustness flaw fixed by hand — `readTrailingObject` used a reverse brace/quote scanner whose
|
|
64
|
+
escape handling is wrong scanning backwards (a `\"` inside a string could derail it and waste the single retry);
|
|
65
|
+
replaced with the parse-candidates approach `hasStructuredAnswer` already uses. Double failure now reports the
|
|
66
|
+
retry's errors. Escalation concern from verify-impl judged mistaken (escalation follows failed dispatches only).
|
|
67
|
+
- 0.13.1 fix path: NOT exercised here either — the latest worker at completion was `full-suite-report`, verified by a
|
|
68
|
+
direct edge (`verify-final`).
|
|
69
|
+
## Adversarial review via `bullswarm run --lane analyze` (02:36 → 02:39 Z, 206 s, opencode2 kaihk/gpt-5.6-luna)
|
|
70
|
+
- Verdict "do not release" with **two confirmed, reproduced defects** — both in code I had reviewed and passed:
|
|
71
|
+
(1) my rewritten `readTrailingObject` returned the first schema-valid `{…}` from the right, so a nested object
|
|
72
|
+
could be recorded (`{"wrapper":{"ok":"inner"}}` → data `{"ok":"inner"}`); (2) `schema.js` used `in`, so
|
|
73
|
+
`toString`/`constructor`/`__proto__` counted as present/declared. Cleared: escaped quotes, fenced JSON, resume rules,
|
|
74
|
+
exactly-one retry, validator paths. Fixed + regression tests (320/320); my fix also had an infinite loop when output
|
|
75
|
+
starts with `{` (lastIndexOf clamps negative fromIndex) — caught by the suite hanging, fixed.
|
|
76
|
+
- Evidence for open item "adversarial verification by default": a 3-minute refute-framed review found what the
|
|
77
|
+
run's own verifies (6 rounds) and my manual review both missed.
|
|
78
|
+
## Goal 5 — planner context/contract refactor — run `wf-mtds7tzx-95ab05` / `ejk9w2`, 02:49:09 Z, runtime 0.13.2 local
|
|
79
|
+
- scout 72 s (opencode2); planner turn 1 02:50:55 → 02:58:56 (~480 s): 6 actions, completion attached; noticed the goal's
|
|
80
|
+
stale "318 passing" and used the real 320. Program: impl-src → verify-src(repair 2) → update-tests ∥ update-docs →
|
|
81
|
+
verify-tests(repair 2) → verify-suite(repair 1).
|
|
82
|
+
- 02:58:56 → 03:09:04 impl-src (~610 s). verify-src rejected twice, both times on SUBSTANTIVE spec points (obsolete
|
|
83
|
+
skeleton text left in a comment; 6 JSON examples instead of 2; `<item>` instead of `{{item}}` in examples; excerpt
|
|
84
|
+
policy). One misread to check in the final diff: it called the existing 3 000/36 000-char excerpt caps a violation of
|
|
85
|
+
"full excerpt" although the goal said "(existing budget logic)" — the repair may have removed the caps.
|
|
86
|
+
Ordering tension: verify-src runs before update-tests, so it necessarily sees 5 failing old assertions; the planner
|
|
87
|
+
should either fold assertion updates into impl-src or make verify-src judge src only.
|
|
88
|
+
- Goal 5 finished 04:00:40 Z (71.5 min, 3 planner turns + program-completion, 15 dispatches, plannerSec 1 105+175).
|
|
89
|
+
Where the time went: 18–20 min planner turns; 12.4 min verify-src repair loop enforcing my over-exacting spec and
|
|
90
|
+
judging intermediate state; ~8 min decision-3 nit round (`align-prefix-number` + `confirm-docs`) triggered by an
|
|
91
|
+
UNPARSEABLE final-check verdict; the queued steer (03:57:32, "converge now") was NEVER delivered — 0 steering.delivered events; the run auto-completed (program-completion, 04:00:40) and deliverSteering only runs at planner gates, so auto-completion silently discards pending operator steering. Known issue; convergence came from the runtime, not the steer.
|
|
92
|
+
Deliverable reviewed + committed `15f1534`: 10-rule contract (2.2 k chars) + 2 examples (1.9 k) replace 16.3 k prefix;
|
|
93
|
+
compact ledger rows; id-only failures; 200-char stale excerpts; `decision.context_built` size event; 323/323.
|
|
94
|
+
- My follow-up (commit after 15f1534): verify verdict parse failure → ONE bounded re-ask (`verify.verdict_retry`,
|
|
95
|
+
test with a garbled-once verifier, 324/324); contract amendments (verify scoping, converge-not-polish, restored
|
|
96
|
+
shared-tree/redundant-verification/operatorSteering lines the merge dropped); re-budgeted "full" excerpts (the
|
|
97
|
+
uncapped version could have rebuilt the 163 k contexts).
|
|
98
|
+
- Speed answer to the user: ~35 of 66 min (at question time) was real work; fixes target the rest — re-ask (−8 min),
|
|
99
|
+
verify scoping (−12 min), convergence rule (−nit rounds). Remaining lever: planner turn latency itself (Opus
|
|
100
|
+
high-effort per boundary; context compaction cuts cost ~7×, latency is model thinking time).
|
|
101
|
+
|
|
102
|
+
## Post-run defects → 0.14.1 (fixed directly, dogfooding paused by user direction)
|
|
103
|
+
- `workflow tui` crashed twice (`detailRow` dashboard.js:809 `JSON.parse` of
|
|
104
|
+
state.json mid-write; the throw escaped the repaint timer and killed the TUI,
|
|
105
|
+
stranding the terminal in alt-screen raw mode). Fix: atomic temp+rename
|
|
106
|
+
writes for state/report/workflow.json + torn-read-tolerant observation
|
|
107
|
+
readers + guarded paint/key handlers with last-good-frame fallback.
|
|
108
|
+
- TUI showed phase ✗ "2/2 complete" while a re-verify attempt was live. Fix:
|
|
109
|
+
an action with an active agent reads as running; phase precedence
|
|
110
|
+
active > failed > completed.
|
|
111
|
+
- The goal-5 steer was never delivered: auto-completion bypassed the planner
|
|
112
|
+
gate and silently discarded pending steering (0 steering.delivered events).
|
|
113
|
+
Fix: pending steering defers self-completion to the planner
|
|
114
|
+
(`decision.completion_deferred`); undeliverable steering is marked
|
|
115
|
+
`expired_undelivered` (`steering.expired`) at the terminal transition.
|
|
116
|
+
- Routing concentration on the unmetered pool (26/28 dispatches) confirmed as
|
|
117
|
+
design intent (quota protection outranks diversity) and documented in the
|
|
118
|
+
skill rather than changed.
|
|
119
|
+
|
|
120
|
+
## Proof run 1 — goal-3 re-run `d8pr8s` (wf-mtdvuk9m), 0.14.1-pre @ 0bbe78c
|
|
121
|
+
Baseline (0.13.1, same fixture/goal/pool/flags): 28m42s wall, 1 planner turn, 294 s planner.
|
|
122
|
+
- **completed + verified, deliverable exactly right** (csv+slugify guarded, existing tests byte-identical, 63/63),
|
|
123
|
+
zero crashes, 8 workers, no repairs.
|
|
124
|
+
- Wall **32m19s**, planner **5 dispatches / 463 s** (195+114+58+66+30). Per-turn latency DOWN (max 195 s vs 294 s);
|
|
125
|
+
context per turn 6.2k–48k chars (`decision.context_built` measuring itself) vs the old 16.3k prefix + up-to-178k contexts.
|
|
126
|
+
- The 3 extra gates, each diagnosed and fixed in `c0ff947`:
|
|
127
|
+
1. first proposal rejected — verify with several dependsOn lacked `review` → contract rule 4 now states it;
|
|
128
|
+
2. evidence-policy boundary + rejected `complete` — final-report left as last unverified worker → rule 8 now
|
|
129
|
+
states the LAST worker must be covered by a verify;
|
|
130
|
+
3. one deliberate operator steer — which **live-proved the 0.14.1 steering fix**: `decision.completion_deferred`
|
|
131
|
+
→ `steering.delivered` (first ever observed; the goal-5 defect showed 0) → planner turn honoured it.
|
|
132
|
+
- Also observed and fixed: a corrective turn re-inflated validationFeedback to 24k chars (raw response duplicated
|
|
133
|
+
the parsed proposal) → excerpt capped at 2k.
|
|
134
|
+
- 0.13.1 completion-evidence policy exercised live for the first time: it refused auto-completion twice, correctly.
|
|
135
|
+
|
|
136
|
+
## Proof run 2 — goal-3 re-run `djnjka` (wf-mtdx5htt), 0.14.1-pre @ c0ff947 — interrupted, then cancelled
|
|
137
|
+
- Turn-1 proposal **accepted first try** (no validation rejection, no correction turn): contract fix #1 confirmed.
|
|
138
|
+
Turn-1 context 6,201 chars.
|
|
139
|
+
- At 05:21 the driving session's background task was killed by the harness (not the user); SIGTERM reached the
|
|
140
|
+
runner, which persisted `interrupted` + resumable (1/3 steps, in-flight `triage` cancelled) — the 0.13 interruption
|
|
141
|
+
path working as designed.
|
|
142
|
+
- Resume at 05:54 exposed a **resume defect**: the cancelled `triage` was persisted `ok:false` ("workflow
|
|
143
|
+
cancellation requested"), so its 4 dependents were marked "blocked by failed or unresolved dependencies" and the
|
|
144
|
+
planner was asked to re-plan around a failure that never happened. Cancelled the run; fixed in `459c58c`
|
|
145
|
+
(cancelled actions + dependents blocked only by them are reopened on resume, event `action.reopened`; regression
|
|
146
|
+
test drives a cancel marker into a slow in-flight action and asserts exactly one further planner turn).
|
|
147
|
+
|
|
148
|
+
## Proof run 3 — goal-3 re-run `4t6m5a` (wf-mtdyyqkw), 0.14.1-pre @ 459c58c — cancelled after diagnosis
|
|
149
|
+
- Turn-1 proposal (9 actions, sound shape: 3 disjoint impl workers ∥ audit of untouched modules → per-module
|
|
150
|
+
verifies → final-report → verify-final, completion attached) was **rejected** for one field: a zero-dependsOn
|
|
151
|
+
audit verify carried no `review`. Correction turn cost ~5 min and re-inflated context to 47.8k chars
|
|
152
|
+
(validationFeedback 38k — the 2k cap on rejectedResponseExcerpt was insufficient because rejectedProposal
|
|
153
|
+
itself is 37k). Root cause is the validator's posture, not the planner: rejecting a whole program for a field the
|
|
154
|
+
runtime can default. Fixed in `548eabe`: review defaults to the single/last dependency's artifact, or
|
|
155
|
+
`reviewScope: repository` for a no-dependency audit; contract rule 4 reworded. Run cancelled to re-prove cleanly.
|
|
156
|
+
|
|
157
|
+
## Proof run 4 — goal-3 re-run `d7xyg2` (wf-mtdzhw88), 0.14.1-pre @ 548eabe — **PASS**
|
|
158
|
+
| metric | 0.13.1 baseline (x3x2a2-era, same fixture) | 0.14.1-pre run 4 |
|
|
159
|
+
| --- | ---: | ---: |
|
|
160
|
+
| outcome | completed, verified | completed, verified, **auto-completed** (program-completion) |
|
|
161
|
+
| wall | 28 min 42 s | 30 min 12 s |
|
|
162
|
+
| planner turns / plannerSec | 1 / 294 s | **1 / 269 s** |
|
|
163
|
+
| planner context (turn 1) | 32.7 k chars (measured on a sibling run) | **6.2 k chars** |
|
|
164
|
+
| dispatches (workers) | — | 8 (7) · max concurrent 9 |
|
|
165
|
+
| corrections / rejections / repairs / verdict re-asks | 0 / 0 / 0 / 0 | 0 / 0 / 0 / 0 |
|
|
166
|
+
| deliverable | csv + slugify guarded, 63/63 | csv + slugify guarded, **63/63**, existing test files byte-identical |
|
|
167
|
+
Program: scout → probe (all exports, wrong-type matrix) → fanout guards over the discovered modules → verify-guards →
|
|
168
|
+
report → verify-final, `completion: all-actions-ok`. Wall is within noise of baseline (+90 s, dominated by worker
|
|
169
|
+
model time: probe 4.7 min, guards 5.1 min, verify-guards 4.7 min); the planner side is faster and 5× smaller.
|
|
170
|
+
Zero observation crashes across four runs of TUI/watch/runs/result/static-tui polling and a 20 s stress loop
|
|
171
|
+
(1,681 paints against the live writer, 0 torn, 0 throws).
|
package/package.json
CHANGED
package/skill/SKILL.md
CHANGED
|
@@ -246,7 +246,12 @@ bullswarm workflow steer <shortId> --message "guidance for the next planner chec
|
|
|
246
246
|
`capabilities` reports available pools, supported lanes, configured models,
|
|
247
247
|
meter readings, burst gates, quarantine state, retry limits, and the important
|
|
248
248
|
routing rule. Automatic routing chooses the highest time-adjusted quota surplus
|
|
249
|
-
among capable pools.
|
|
249
|
+
among capable pools. An unmetered pool reads as exactly on pace (surplus 0), so
|
|
250
|
+
whenever every metered pool is burning ahead of its window (negative
|
|
251
|
+
surplus) the unmetered pool wins ALL work — by design: quota protection
|
|
252
|
+
outranks provider diversity (observed 2026-08-29: 26 of 28 dispatches on
|
|
253
|
+
one unmetered pool). If that concentration is unwanted, meter the pool or
|
|
254
|
+
exclude its models via `strategy exclude-model`. For strategic model selection, first run:
|
|
250
255
|
|
|
251
256
|
```bash
|
|
252
257
|
bullswarm strategy refresh
|
|
@@ -406,6 +411,17 @@ Resume by shortId:
|
|
|
406
411
|
bullswarm workflow draft run my-audit --resume <shortId> --json --quiet
|
|
407
412
|
```
|
|
408
413
|
|
|
414
|
+
## Writing goals that converge
|
|
415
|
+
|
|
416
|
+
State outcomes, not measurements. A goal that fixes character counts, exact
|
|
417
|
+
event names, or cosmetic layout turns every verifier into a nit machine:
|
|
418
|
+
observed 2026-08-29 (run `ejk9w2`), a spec with hard numeric limits cost a
|
|
419
|
+
12-minute verify/repair loop enforcing them against an intermediate state.
|
|
420
|
+
Say what must be true at the end (`npm test` passes, the planner receives one
|
|
421
|
+
contract stated once, context stays bounded) and let workers pick the numbers;
|
|
422
|
+
put any hard limit in ONE final acceptance check, not on every intermediate
|
|
423
|
+
verify.
|
|
424
|
+
|
|
409
425
|
## Writing prompts that the verify gate will accept
|
|
410
426
|
|
|
411
427
|
The verify gate (`src/lib/verify.js`) flags outputs as `intent_only`
|
|
@@ -0,0 +1,202 @@
|
|
|
1
|
+
// Discover Claude Code logins on this machine.
|
|
2
|
+
//
|
|
3
|
+
// One account = one config dir (`~/.claude` by default, or CLAUDE_CONFIG_DIR).
|
|
4
|
+
// Extra homes live next to it as `~/.claude-<slug>`. Credentials:
|
|
5
|
+
// macOS Keychain `Claude Code-credentials` for ~/.claude
|
|
6
|
+
// `Claude Code-credentials-<sha256(absPath)[:8]>` for any other home
|
|
7
|
+
// plus `$dir/.credentials.json` on every platform.
|
|
8
|
+
|
|
9
|
+
import { createHash } from 'node:crypto';
|
|
10
|
+
import { existsSync, readdirSync, readFileSync, statSync } from 'node:fs';
|
|
11
|
+
import { homedir, platform } from 'node:os';
|
|
12
|
+
import { basename, join, resolve } from 'node:path';
|
|
13
|
+
import {
|
|
14
|
+
DEFAULT_KEYCHAIN_SERVICE,
|
|
15
|
+
extractCredentials,
|
|
16
|
+
isUsable,
|
|
17
|
+
readFromMacKeychain,
|
|
18
|
+
} from '../meters/claude.js';
|
|
19
|
+
|
|
20
|
+
const HOME_MARKERS = ['.credentials.json', '.claude.json', 'settings.json', 'projects'];
|
|
21
|
+
|
|
22
|
+
export function defaultClaudeHome(homeDir = homedir()) {
|
|
23
|
+
return resolve(join(homeDir, '.claude'));
|
|
24
|
+
}
|
|
25
|
+
|
|
26
|
+
export function keychainServiceForConfigDir(configDir, homeDir = homedir()) {
|
|
27
|
+
const resolved = resolve(configDir);
|
|
28
|
+
if (resolved === defaultClaudeHome(homeDir)) return DEFAULT_KEYCHAIN_SERVICE;
|
|
29
|
+
const hash = createHash('sha256').update(resolved).digest('hex').slice(0, 8);
|
|
30
|
+
return `${DEFAULT_KEYCHAIN_SERVICE}-${hash}`;
|
|
31
|
+
}
|
|
32
|
+
|
|
33
|
+
export function accountSlugForConfigDir(configDir, homeDir = homedir()) {
|
|
34
|
+
const resolved = resolve(configDir);
|
|
35
|
+
if (resolved === defaultClaudeHome(homeDir)) return null;
|
|
36
|
+
const base = basename(resolved);
|
|
37
|
+
if (base.startsWith('.claude-')) {
|
|
38
|
+
const slug = base.slice('.claude-'.length);
|
|
39
|
+
return slug.length > 0 ? slug : 'alt';
|
|
40
|
+
}
|
|
41
|
+
if (base.startsWith('.claude')) {
|
|
42
|
+
const rest = base.slice('.claude'.length).replace(/^-+/, '');
|
|
43
|
+
return rest.length > 0 ? rest : 'alt';
|
|
44
|
+
}
|
|
45
|
+
return base || 'alt';
|
|
46
|
+
}
|
|
47
|
+
|
|
48
|
+
export function poolNameForSlug(slug) {
|
|
49
|
+
return slug ? `claude-code:${slug}` : 'claude-code';
|
|
50
|
+
}
|
|
51
|
+
|
|
52
|
+
export function looksLikeClaudeHome(dir) {
|
|
53
|
+
try {
|
|
54
|
+
if (!statSync(dir).isDirectory()) return false;
|
|
55
|
+
} catch {
|
|
56
|
+
return false;
|
|
57
|
+
}
|
|
58
|
+
return HOME_MARKERS.some((name) => existsSync(join(dir, name)));
|
|
59
|
+
}
|
|
60
|
+
|
|
61
|
+
export function discoverClaudeConfigDirs(opts = {}) {
|
|
62
|
+
const homeDir = opts.homeDir ?? homedir();
|
|
63
|
+
const found = [];
|
|
64
|
+
const seen = new Set();
|
|
65
|
+
const add = (dir) => {
|
|
66
|
+
const resolved = resolve(dir);
|
|
67
|
+
if (seen.has(resolved)) return;
|
|
68
|
+
if (!looksLikeClaudeHome(resolved)) return;
|
|
69
|
+
seen.add(resolved);
|
|
70
|
+
found.push(resolved);
|
|
71
|
+
};
|
|
72
|
+
add(join(homeDir, '.claude'));
|
|
73
|
+
try {
|
|
74
|
+
for (const name of readdirSync(homeDir)) {
|
|
75
|
+
if (!name.startsWith('.claude-')) continue;
|
|
76
|
+
add(join(homeDir, name));
|
|
77
|
+
}
|
|
78
|
+
} catch { /* unreadable home */ }
|
|
79
|
+
const envDir = opts.envConfigDir ?? process.env.CLAUDE_CONFIG_DIR;
|
|
80
|
+
if (envDir && String(envDir).trim()) add(String(envDir).trim());
|
|
81
|
+
return found;
|
|
82
|
+
}
|
|
83
|
+
|
|
84
|
+
function readCredentialsFile(path) {
|
|
85
|
+
try {
|
|
86
|
+
return extractCredentials(readFileSync(path, 'utf8'));
|
|
87
|
+
} catch {
|
|
88
|
+
return null;
|
|
89
|
+
}
|
|
90
|
+
}
|
|
91
|
+
|
|
92
|
+
export function readAccountCredentials(configDir, opts = {}) {
|
|
93
|
+
const homeDir = opts.homeDir ?? homedir();
|
|
94
|
+
const nowMs = opts.nowMs ?? Date.now();
|
|
95
|
+
const os = opts.platform ?? platform();
|
|
96
|
+
const keychainRead = opts.readKeychain ?? readFromMacKeychain;
|
|
97
|
+
const usableOnly = opts.usableOnly !== false;
|
|
98
|
+
|
|
99
|
+
const fileCreds = readCredentialsFile(join(configDir, '.credentials.json'));
|
|
100
|
+
let keychainCreds = null;
|
|
101
|
+
if (os === 'darwin') {
|
|
102
|
+
keychainCreds = keychainRead(keychainServiceForConfigDir(configDir, homeDir));
|
|
103
|
+
}
|
|
104
|
+
const extraFile = resolve(configDir) === defaultClaudeHome(homeDir)
|
|
105
|
+
? readCredentialsFile(join(homeDir, '.config', 'claude', 'credentials.json'))
|
|
106
|
+
: null;
|
|
107
|
+
|
|
108
|
+
const candidates = [];
|
|
109
|
+
if (keychainCreds) candidates.push({ creds: keychainCreds, source: 'keychain' });
|
|
110
|
+
if (fileCreds) candidates.push({ creds: fileCreds, source: 'file' });
|
|
111
|
+
if (extraFile) candidates.push({ creds: extraFile, source: 'file' });
|
|
112
|
+
for (const c of candidates) {
|
|
113
|
+
if (!usableOnly || isUsable(c.creds, nowMs)) return c;
|
|
114
|
+
}
|
|
115
|
+
return null;
|
|
116
|
+
}
|
|
117
|
+
|
|
118
|
+
export function profileCommand(configDir, bin = 'claude') {
|
|
119
|
+
return `CLAUDE_CONFIG_DIR=${configDir} ${bin}`;
|
|
120
|
+
}
|
|
121
|
+
|
|
122
|
+
export function discoverClaudeAccounts(opts = {}) {
|
|
123
|
+
const homeDir = opts.homeDir ?? homedir();
|
|
124
|
+
const dirs = discoverClaudeConfigDirs({
|
|
125
|
+
homeDir,
|
|
126
|
+
envConfigDir: opts.envConfigDir,
|
|
127
|
+
});
|
|
128
|
+
const accounts = [];
|
|
129
|
+
const seenTokens = new Set();
|
|
130
|
+
for (const configDir of dirs) {
|
|
131
|
+
const got = readAccountCredentials(configDir, opts);
|
|
132
|
+
if (!got) continue;
|
|
133
|
+
if (seenTokens.has(got.creds.accessToken)) continue;
|
|
134
|
+
seenTokens.add(got.creds.accessToken);
|
|
135
|
+
const slug = accountSlugForConfigDir(configDir, homeDir);
|
|
136
|
+
accounts.push({
|
|
137
|
+
configDir,
|
|
138
|
+
slug,
|
|
139
|
+
pool: poolNameForSlug(slug),
|
|
140
|
+
command: profileCommand(configDir, opts.bin ?? 'claude'),
|
|
141
|
+
creds: got.creds,
|
|
142
|
+
source: got.source,
|
|
143
|
+
});
|
|
144
|
+
}
|
|
145
|
+
accounts.sort((a, b) => {
|
|
146
|
+
if (a.slug === null) return -1;
|
|
147
|
+
if (b.slug === null) return 1;
|
|
148
|
+
return a.slug.localeCompare(b.slug);
|
|
149
|
+
});
|
|
150
|
+
return accounts;
|
|
151
|
+
}
|
|
152
|
+
|
|
153
|
+
/**
|
|
154
|
+
* Clone the packaged `claude-code` connector once per extra login.
|
|
155
|
+
* The default home keeps the historical pool name `claude-code`.
|
|
156
|
+
* Extra homes become `claude-code:<slug>` with CLAUDE_CONFIG_DIR in env
|
|
157
|
+
* so spawn bills that seat. Discovery is filesystem-driven — never a
|
|
158
|
+
* hardcoded profile list.
|
|
159
|
+
*/
|
|
160
|
+
export function expandClaudeAccountConnectors(connectors, opts = {}) {
|
|
161
|
+
if (opts.disabled === true) return connectors;
|
|
162
|
+
if (process.env.BULLSWARM_DISABLE_CLAUDE_PROFILES === '1') return connectors;
|
|
163
|
+
// node:test (and MCP children it spawns) inherit NODE_TEST_CONTEXT. Do not
|
|
164
|
+
// scan the operator's real extra Claude homes unless the test injects
|
|
165
|
+
// `accounts` or `homeDir`.
|
|
166
|
+
if (process.env.NODE_TEST_CONTEXT && opts.accounts == null && opts.homeDir == null) {
|
|
167
|
+
return connectors;
|
|
168
|
+
}
|
|
169
|
+
const base = connectors['claude-code'];
|
|
170
|
+
if (!base) return connectors;
|
|
171
|
+
const accounts = opts.accounts ?? discoverClaudeAccounts({
|
|
172
|
+
homeDir: opts.homeDir,
|
|
173
|
+
envConfigDir: opts.envConfigDir,
|
|
174
|
+
bin: base.bin ?? 'claude',
|
|
175
|
+
});
|
|
176
|
+
const defaultDir = accounts.find((a) => a.slug == null)?.configDir
|
|
177
|
+
?? defaultClaudeHome(opts.homeDir);
|
|
178
|
+
const bin = base.bin ?? 'claude';
|
|
179
|
+
base.env = { ...(base.env ?? {}), CLAUDE_CONFIG_DIR: defaultDir };
|
|
180
|
+
base.profile = {
|
|
181
|
+
slug: null,
|
|
182
|
+
configDir: defaultDir,
|
|
183
|
+
command: profileCommand(defaultDir, bin),
|
|
184
|
+
};
|
|
185
|
+
for (const account of accounts) {
|
|
186
|
+
if (!account.slug) continue;
|
|
187
|
+
const name = account.pool;
|
|
188
|
+
if (connectors[name]) continue;
|
|
189
|
+
const clone = structuredClone(base);
|
|
190
|
+
clone.name = name;
|
|
191
|
+
clone.env = { ...(base.env ?? {}), CLAUDE_CONFIG_DIR: account.configDir };
|
|
192
|
+
clone.configDirs = [account.configDir];
|
|
193
|
+
clone.flags = { ...(base.flags ?? {}), isCaller: false };
|
|
194
|
+
clone.profile = {
|
|
195
|
+
slug: account.slug,
|
|
196
|
+
configDir: account.configDir,
|
|
197
|
+
command: account.command,
|
|
198
|
+
};
|
|
199
|
+
connectors[name] = clone;
|
|
200
|
+
}
|
|
201
|
+
return connectors;
|
|
202
|
+
}
|
package/src/lib/config.js
CHANGED
|
@@ -13,8 +13,9 @@ import { readFileSync, readdirSync, existsSync } from 'node:fs';
|
|
|
13
13
|
import { join } from 'node:path';
|
|
14
14
|
import { loadState } from './state.js';
|
|
15
15
|
import { paceScore, isQuarantined } from './route.js';
|
|
16
|
+
import { expandClaudeAccountConnectors } from './claude-accounts.js';
|
|
16
17
|
|
|
17
|
-
export function loadConnectors(bullswarmDir) {
|
|
18
|
+
export function loadConnectors(bullswarmDir, opts = {}) {
|
|
18
19
|
const dir = join(bullswarmDir, 'connectors');
|
|
19
20
|
if (!existsSync(dir)) return {};
|
|
20
21
|
const out = {};
|
|
@@ -28,6 +29,7 @@ export function loadConnectors(bullswarmDir) {
|
|
|
28
29
|
// never crash a run
|
|
29
30
|
}
|
|
30
31
|
}
|
|
32
|
+
expandClaudeAccountConnectors(out, opts);
|
|
31
33
|
return out;
|
|
32
34
|
}
|
|
33
35
|
|
package/src/lib/watch.js
CHANGED
|
@@ -78,7 +78,12 @@ export function runDelegate(connector, taskFile, targetDir, opts = {}) {
|
|
|
78
78
|
// hazard). Caller-supplied opts.env takes precedence over
|
|
79
79
|
// process.env so the runtime can inject BULLSWARM_DEPTH (recursion
|
|
80
80
|
// guard) and other core-owned env contracts.
|
|
81
|
-
env: {
|
|
81
|
+
env: {
|
|
82
|
+
...process.env,
|
|
83
|
+
...(connector.env ?? {}),
|
|
84
|
+
...(opts.env ?? {}),
|
|
85
|
+
PWD: resolvedDir,
|
|
86
|
+
},
|
|
82
87
|
stdio: ['ignore', 'pipe', 'pipe'],
|
|
83
88
|
});
|
|
84
89
|
opts.onSpawn?.(child.pid);
|
package/src/meters/claude.js
CHANGED
|
@@ -19,18 +19,24 @@ export class ClaudeMeterError extends Error {
|
|
|
19
19
|
}
|
|
20
20
|
}
|
|
21
21
|
|
|
22
|
+
export const DEFAULT_KEYCHAIN_SERVICE = 'Claude Code-credentials';
|
|
23
|
+
|
|
24
|
+
export function isUsable(creds, now = Date.now(), skewMs = EXPIRY_SKEW_MS) {
|
|
25
|
+
return Boolean(creds) && creds.expiresAt - skewMs > now;
|
|
26
|
+
}
|
|
27
|
+
|
|
22
28
|
export function readOAuthCredentials() {
|
|
23
29
|
if (platform() === 'darwin') {
|
|
24
|
-
return readFromMacKeychain() ?? readFromCredentialsFile();
|
|
30
|
+
return readFromMacKeychain(DEFAULT_KEYCHAIN_SERVICE) ?? readFromCredentialsFile();
|
|
25
31
|
}
|
|
26
32
|
return readFromCredentialsFile();
|
|
27
33
|
}
|
|
28
34
|
|
|
29
|
-
function readFromMacKeychain() {
|
|
35
|
+
export function readFromMacKeychain(service = DEFAULT_KEYCHAIN_SERVICE) {
|
|
30
36
|
try {
|
|
31
37
|
const blob = execFileSync(
|
|
32
38
|
'security',
|
|
33
|
-
['find-generic-password', '-s',
|
|
39
|
+
['find-generic-password', '-s', service, '-w'],
|
|
34
40
|
{ stdio: ['ignore', 'pipe', 'ignore'], encoding: 'utf8' },
|
|
35
41
|
);
|
|
36
42
|
return extractCredentials(blob);
|
|
@@ -73,13 +79,13 @@ function normalizeWindow(raw) {
|
|
|
73
79
|
};
|
|
74
80
|
}
|
|
75
81
|
|
|
76
|
-
export function parseClaudeUsage(body) {
|
|
82
|
+
export function parseClaudeUsage(body, pool = 'claude-code') {
|
|
77
83
|
if (!body || typeof body !== 'object') {
|
|
78
84
|
throw new ClaudeMeterError('Claude usage response missing body', 'parse');
|
|
79
85
|
}
|
|
80
86
|
return {
|
|
81
87
|
captured_at: new Date().toISOString(),
|
|
82
|
-
pool
|
|
88
|
+
pool,
|
|
83
89
|
five_hour: normalizeWindow(body.five_hour) ?? { utilization: null, resets_at: null },
|
|
84
90
|
seven_day: normalizeWindow(body.seven_day) ?? { utilization: null, resets_at: null },
|
|
85
91
|
monthly: null,
|
|
@@ -89,12 +95,11 @@ export function parseClaudeUsage(body) {
|
|
|
89
95
|
};
|
|
90
96
|
}
|
|
91
97
|
|
|
92
|
-
export async function
|
|
93
|
-
const creds = readOAuthCredentials();
|
|
98
|
+
export async function fetchClaudeUsageWithCredentials(creds, pool = 'claude-code') {
|
|
94
99
|
if (!creds) {
|
|
95
100
|
throw new ClaudeMeterError('No Claude Code OAuth token. Run `claude` to log in.', 'no_token');
|
|
96
101
|
}
|
|
97
|
-
if (creds
|
|
102
|
+
if (!isUsable(creds)) {
|
|
98
103
|
throw new ClaudeMeterError(
|
|
99
104
|
'Claude OAuth token expired; open Claude Code to refresh it.',
|
|
100
105
|
'expired',
|
|
@@ -122,5 +127,9 @@ export async function fetchClaudeUsage() {
|
|
|
122
127
|
} catch (err) {
|
|
123
128
|
throw new ClaudeMeterError(`Failed to parse usage response: ${err.message}`, 'parse');
|
|
124
129
|
}
|
|
125
|
-
return parseClaudeUsage(body);
|
|
130
|
+
return parseClaudeUsage(body, pool);
|
|
131
|
+
}
|
|
132
|
+
|
|
133
|
+
export async function fetchClaudeUsage() {
|
|
134
|
+
return fetchClaudeUsageWithCredentials(readOAuthCredentials(), 'claude-code');
|
|
126
135
|
}
|
package/src/meters/registry.js
CHANGED
|
@@ -6,7 +6,8 @@ import { MeterCache, paceSnapshot, FRESH_MS, STALE_MS } from './framework.js';
|
|
|
6
6
|
import { fetchCodexUsage, CodexMeterError } from './codex.js';
|
|
7
7
|
import { fetchGrokUsage, GrokMeterError } from './grok.js';
|
|
8
8
|
import { fetchCommandCodeUsage, CommandCodeMeterError } from './command-code.js';
|
|
9
|
-
import { fetchClaudeUsage, ClaudeMeterError } from './claude.js';
|
|
9
|
+
import { fetchClaudeUsage, fetchClaudeUsageWithCredentials, ClaudeMeterError } from './claude.js';
|
|
10
|
+
import { discoverClaudeAccounts, poolNameForSlug } from '../lib/claude-accounts.js';
|
|
10
11
|
|
|
11
12
|
export const METERS_DIR = () =>
|
|
12
13
|
process.env.BULLSWARM_HOME?.trim() || join(homedir(), '.bullswarm');
|
|
@@ -19,7 +20,26 @@ const READERS = {
|
|
|
19
20
|
claude: fetchClaudeUsage,
|
|
20
21
|
};
|
|
21
22
|
|
|
23
|
+
function claudeReaderFor(pool) {
|
|
24
|
+
return async () => {
|
|
25
|
+
const accounts = discoverClaudeAccounts();
|
|
26
|
+
const slug = pool.startsWith('claude-code:') ? pool.slice('claude-code:'.length) : null;
|
|
27
|
+
const account = accounts.find((a) => poolNameForSlug(a.slug) === pool)
|
|
28
|
+
?? accounts.find((a) => a.slug === slug);
|
|
29
|
+
if (!account) {
|
|
30
|
+
throw new ClaudeMeterError(
|
|
31
|
+
`No Claude Code OAuth token for pool ${pool}. Log in with CLAUDE_CONFIG_DIR pointing at that home.`,
|
|
32
|
+
'no_token',
|
|
33
|
+
);
|
|
34
|
+
}
|
|
35
|
+
return fetchClaudeUsageWithCredentials(account.creds, pool);
|
|
36
|
+
};
|
|
37
|
+
}
|
|
38
|
+
|
|
22
39
|
export function readerFor(pool) {
|
|
40
|
+
if (pool === 'claude-code' || pool === 'claude' || pool.startsWith('claude-code:')) {
|
|
41
|
+
return claudeReaderFor(pool);
|
|
42
|
+
}
|
|
23
43
|
return READERS[pool] ?? null;
|
|
24
44
|
}
|
|
25
45
|
|
package/src/workflow/cli.js
CHANGED
|
@@ -340,7 +340,7 @@ async function wfGoal(opts) {
|
|
|
340
340
|
try {
|
|
341
341
|
doc = existsSync(workflowPath)
|
|
342
342
|
? JSON.parse(readFileSync(workflowPath, 'utf8'))
|
|
343
|
-
:
|
|
343
|
+
: readJsonForUpdate(statePath, 'workflow state')._doc;
|
|
344
344
|
} catch (err) {
|
|
345
345
|
console.error(`✗ cannot load durable workflow for ${resumeRunId}: ${err.message}`);
|
|
346
346
|
return 1;
|
|
@@ -457,6 +457,8 @@ async function wfCapabilities(opts) {
|
|
|
457
457
|
enabled: p.enabled !== false,
|
|
458
458
|
lanes: p.lanes ?? p.connector?.lanes ?? [],
|
|
459
459
|
capabilities: p.capabilities ?? p.connector?.capabilities ?? [],
|
|
460
|
+
command: p.connector?.profile?.command ?? p.connector?.spawn?.cmd?.[0] ?? null,
|
|
461
|
+
configDir: p.connector?.profile?.configDir ?? p.connector?.env?.CLAUDE_CONFIG_DIR ?? null,
|
|
460
462
|
model: (() => {
|
|
461
463
|
const cmd = p.connector?.spawn?.cmd ?? [];
|
|
462
464
|
const i = cmd.indexOf('--model');
|
|
@@ -1,7 +1,8 @@
|
|
|
1
1
|
// Interactive workflow dashboard, inspired by Claude Code's /workflows view.
|
|
2
2
|
// It deliberately uses only ANSI sequences and Node's standard streams.
|
|
3
3
|
|
|
4
|
-
import { readFileSync,
|
|
4
|
+
import { readFileSync, existsSync } from 'node:fs';
|
|
5
|
+
import { readJsonSafe, readJsonForUpdate, writeJsonAtomic } from './fsjson.js';
|
|
5
6
|
import { fanoutSucceededCount } from './runner.js';
|
|
6
7
|
import { join } from 'node:path';
|
|
7
8
|
import { listRuns, resolveRunId } from './short-id.js';
|
|
@@ -44,7 +45,7 @@ export function requestCancel(bullswarmDir, token) {
|
|
|
44
45
|
if (!resolved) throw new Error(`no run found for "${token}"`);
|
|
45
46
|
const statePath = join(resolved.runDir, 'state.json');
|
|
46
47
|
if (!existsSync(statePath)) throw new Error(`run "${token}" has no state.json`);
|
|
47
|
-
const state =
|
|
48
|
+
const state = readJsonForUpdate(statePath, 'workflow state');
|
|
48
49
|
if (state.finishedAt || isTerminalWorkflowStatus(state.status)) {
|
|
49
50
|
return { ...resolved, state, alreadyFinished: true };
|
|
50
51
|
}
|
|
@@ -53,7 +54,7 @@ export function requestCancel(bullswarmDir, token) {
|
|
|
53
54
|
state.status = 'cancelling';
|
|
54
55
|
state.cancellingAt = state.cancelRequestedAt;
|
|
55
56
|
appendEvent(resolved.runDir, state, 'run.cancellation_requested', { requestedAt: state.cancelRequestedAt });
|
|
56
|
-
|
|
57
|
+
writeJsonAtomic(statePath, state);
|
|
57
58
|
return { ...resolved, state, alreadyFinished: false };
|
|
58
59
|
}
|
|
59
60
|
|
|
@@ -228,6 +229,12 @@ function autonomousControlPlane(state) {
|
|
|
228
229
|
}
|
|
229
230
|
|
|
230
231
|
function effectiveActionStatus(action, state) {
|
|
232
|
+
// A re-running action (repair round, re-verify, schema retry) must read as
|
|
233
|
+
// running even when a previous round recorded ok:false — a failed mark on
|
|
234
|
+
// work that is still being retried misreports the run (user report 2026-08-29).
|
|
235
|
+
const active = Object.values(state.activeAgents ?? {}).some((agent) =>
|
|
236
|
+
agent.stepId === action.id || String(agent.stepId ?? '').startsWith(`${action.id}[`));
|
|
237
|
+
if (active || action.status === 'running') return 'running';
|
|
231
238
|
const output = state.outputs?.[action.id];
|
|
232
239
|
if (output?.ok === false) return 'failed_terminal';
|
|
233
240
|
if (output?.ok === true && action.status === 'succeeded') return 'succeeded';
|
|
@@ -282,9 +289,10 @@ export function workflowPanelModel(row, {
|
|
|
282
289
|
const failed = effectiveStatuses.filter((status) => String(status).startsWith('failed')).length;
|
|
283
290
|
const current = name === activePhaseName && !state.finishedAt;
|
|
284
291
|
const active = effectiveStatuses.some((status) => status === 'running');
|
|
285
|
-
const status =
|
|
286
|
-
:
|
|
287
|
-
:
|
|
292
|
+
const status = active ? 'active'
|
|
293
|
+
: failed ? 'failed'
|
|
294
|
+
: actions.length && completed === actions.length ? 'completed'
|
|
295
|
+
: current ? 'waiting' : 'pending';
|
|
288
296
|
return { name, label: phaseLabel(name, orchestrator), status, actions, completed, total: actions.length };
|
|
289
297
|
});
|
|
290
298
|
const selectedPhase = phases[selectedPhaseIndex];
|
|
@@ -806,8 +814,8 @@ function detailRow(bullswarmDir, token) {
|
|
|
806
814
|
if (!resolved) throw new Error(`no run found for "${token}"`);
|
|
807
815
|
const statePath = join(resolved.runDir, 'state.json');
|
|
808
816
|
const reportPath = join(resolved.runDir, 'report.json');
|
|
809
|
-
const state =
|
|
810
|
-
const report =
|
|
817
|
+
const state = readJsonSafe(statePath);
|
|
818
|
+
const report = readJsonSafe(reportPath);
|
|
811
819
|
return { ...resolved, state, report, events: readEvents(resolved.runDir), status: state?.status };
|
|
812
820
|
}
|
|
813
821
|
|
|
@@ -823,6 +831,7 @@ export async function runDashboard(bullswarmDir, {
|
|
|
823
831
|
let selected = 0;
|
|
824
832
|
let detail = Boolean(token);
|
|
825
833
|
let message = null;
|
|
834
|
+
let lastGoodRow = null;
|
|
826
835
|
let rows = dashboardRows(bullswarmDir);
|
|
827
836
|
let selectedRunId = token ? detailRow(bullswarmDir, token).runId : (rows[selected]?.runId ?? null);
|
|
828
837
|
const ui = {
|
|
@@ -838,10 +847,14 @@ export async function runDashboard(bullswarmDir, {
|
|
|
838
847
|
orchestratorVerbose: false,
|
|
839
848
|
spinnerFrame: 0,
|
|
840
849
|
};
|
|
841
|
-
const
|
|
850
|
+
const paintUnsafe = () => {
|
|
842
851
|
if (selected >= rows.length) selected = Math.max(0, rows.length - 1);
|
|
843
852
|
if (detail && selectedRunId) {
|
|
844
|
-
|
|
853
|
+
// A torn read while the runner writes state.json yields state:null for
|
|
854
|
+
// one frame — keep painting the last good snapshot of the same run.
|
|
855
|
+
const fresh = detailRow(bullswarmDir, selectedRunId);
|
|
856
|
+
const row = (fresh.state || lastGoodRow?.runId !== fresh.runId) ? fresh : lastGoodRow;
|
|
857
|
+
if (row === fresh) lastGoodRow = fresh;
|
|
845
858
|
const model = workflowPanelModel(row, {
|
|
846
859
|
phaseIndex: ui.followActivePhase ? null : ui.phaseIndex,
|
|
847
860
|
agentIndex: ui.followActiveAgent ? null : ui.agentIndex,
|
|
@@ -858,8 +871,19 @@ export async function runDashboard(bullswarmDir, {
|
|
|
858
871
|
}
|
|
859
872
|
output.write(renderDashboard({ rows, selected, message }));
|
|
860
873
|
};
|
|
874
|
+
// A render error must never kill the TUI or strand the terminal in
|
|
875
|
+
// alt-screen raw mode (crash observed 2026-08-29 at detailRow via the
|
|
876
|
+
// repaint timer). Show the error in the message line and keep running.
|
|
877
|
+
const paint = () => {
|
|
878
|
+
try { paintUnsafe(); } catch (err) {
|
|
879
|
+
message = `display error: ${err.message}`;
|
|
880
|
+
try { output.write(renderDashboard({ rows, selected, message })); } catch { /* keep the loop alive */ }
|
|
881
|
+
}
|
|
882
|
+
};
|
|
861
883
|
const refresh = () => {
|
|
862
|
-
|
|
884
|
+
try {
|
|
885
|
+
rows = dashboardRows(bullswarmDir);
|
|
886
|
+
} catch (err) { message = `display error: ${err.message}`; }
|
|
863
887
|
if (!selectedRunId) selectedRunId = rows[selected]?.runId ?? null;
|
|
864
888
|
paint();
|
|
865
889
|
};
|
|
@@ -935,7 +959,7 @@ export async function runDashboard(bullswarmDir, {
|
|
|
935
959
|
refresh();
|
|
936
960
|
} catch (err) { message = err.message; ui.confirmCancel = false; paint(); }
|
|
937
961
|
};
|
|
938
|
-
const
|
|
962
|
+
const onDataUnsafe = (buf) => {
|
|
939
963
|
const key = String(buf);
|
|
940
964
|
if (ui.confirmCancel) {
|
|
941
965
|
if (key === 'y' || key === 'Y') return requestSelectedCancel();
|
|
@@ -1019,6 +1043,14 @@ export async function runDashboard(bullswarmDir, {
|
|
|
1019
1043
|
return paint();
|
|
1020
1044
|
}
|
|
1021
1045
|
};
|
|
1046
|
+
const onData = (buf) => {
|
|
1047
|
+
// A key-handler error (e.g. a drill-in racing the writer) must never
|
|
1048
|
+
// kill the TUI; finish() still restores the terminal on q/Ctrl-C.
|
|
1049
|
+
try { onDataUnsafe(buf); } catch (err) {
|
|
1050
|
+
message = `display error: ${err.message}`;
|
|
1051
|
+
paint();
|
|
1052
|
+
}
|
|
1053
|
+
};
|
|
1022
1054
|
const onResize = () => paint();
|
|
1023
1055
|
input.on('data', onData);
|
|
1024
1056
|
output.on?.('resize', onResize);
|
|
@@ -1035,8 +1067,8 @@ export function dashboardJson(bullswarmDir, { all = false, token = null, cancel
|
|
|
1035
1067
|
if (!resolved) throw new Error(`no run found for "${token}"`);
|
|
1036
1068
|
const statePath = join(resolved.runDir, 'state.json');
|
|
1037
1069
|
const reportPath = join(resolved.runDir, 'report.json');
|
|
1038
|
-
const state =
|
|
1039
|
-
const report =
|
|
1070
|
+
const state = readJsonSafe(statePath);
|
|
1071
|
+
const report = readJsonSafe(reportPath);
|
|
1040
1072
|
const events = readEvents(resolved.runDir);
|
|
1041
1073
|
return { action: 'show', ...resolved, state, report, events };
|
|
1042
1074
|
}
|
|
@@ -1048,7 +1080,7 @@ export function actionJson(bullswarmDir, token, actionId) {
|
|
|
1048
1080
|
const resolved = resolveRunId(bullswarmDir, token);
|
|
1049
1081
|
if (!resolved) throw new Error(`no run found for "${token}"`);
|
|
1050
1082
|
const statePath = join(resolved.runDir, 'state.json');
|
|
1051
|
-
const state =
|
|
1083
|
+
const state = readJsonSafe(statePath);
|
|
1052
1084
|
const action = state?.actionLedger?.find((entry) => entry.id === actionId);
|
|
1053
1085
|
if (!action) throw new Error(`run "${token}" has no action "${actionId}"`);
|
|
1054
1086
|
const attempts = (action.attempts ?? []).map((index) => state.attempts?.[index]).filter(Boolean);
|
|
@@ -1062,7 +1094,7 @@ export function decideApproval(bullswarmDir, token, decision) {
|
|
|
1062
1094
|
const resolved = resolveRunId(bullswarmDir, token);
|
|
1063
1095
|
if (!resolved) throw new Error(`no run found for "${token}"`);
|
|
1064
1096
|
const statePath = join(resolved.runDir, 'state.json');
|
|
1065
|
-
const state =
|
|
1097
|
+
const state = readJsonForUpdate(statePath, 'workflow state');
|
|
1066
1098
|
if (state.status !== 'waiting_for_approval' || !state.approval) {
|
|
1067
1099
|
throw new Error(`run "${token}" is not waiting for approval`);
|
|
1068
1100
|
}
|
|
@@ -1085,6 +1117,6 @@ export function decideApproval(bullswarmDir, token, decision) {
|
|
|
1085
1117
|
gateId: state.approval.gateId,
|
|
1086
1118
|
decidedAt: at,
|
|
1087
1119
|
});
|
|
1088
|
-
|
|
1120
|
+
writeJsonAtomic(statePath, state);
|
|
1089
1121
|
return { ...resolved, decision, state };
|
|
1090
1122
|
}
|
package/src/workflow/decision.js
CHANGED
|
@@ -72,10 +72,20 @@ export function normalizeDecisionProposal(proposal) {
|
|
|
72
72
|
};
|
|
73
73
|
}
|
|
74
74
|
if (action?.type !== 'verify') return action;
|
|
75
|
-
const
|
|
76
|
-
|
|
75
|
+
const dependsOn = Array.isArray(action.dependsOn) ? action.dependsOn : [];
|
|
76
|
+
const singleDependency = dependsOn.length === 1 ? dependsOn[0] : null;
|
|
77
77
|
if (action.review == null) {
|
|
78
|
-
|
|
78
|
+
// A verify reviews an artifact; when the planner names none, take the
|
|
79
|
+
// obvious one instead of rejecting a whole program for one field
|
|
80
|
+
// (observed 2026-08-29: a 9-action proposal bounced for a
|
|
81
|
+
// zero-dependency audit verify, costing a 5-minute correction turn).
|
|
82
|
+
// One dependency -> its artifact. Several -> the most downstream
|
|
83
|
+
// (last listed). None -> a repository audit with no artifact.
|
|
84
|
+
if (singleDependency) return { ...action, review: `outputs.${singleDependency}.outFile` };
|
|
85
|
+
if (dependsOn.length > 1) {
|
|
86
|
+
return { ...action, review: `outputs.${dependsOn.at(-1)}.outFile`, reviewDefaultedFrom: 'last-dependency' };
|
|
87
|
+
}
|
|
88
|
+
return { ...action, reviewScope: 'repository' };
|
|
79
89
|
}
|
|
80
90
|
if (looksLikeReviewPath(action.review)) return { ...action, review: action.review.trim() };
|
|
81
91
|
if (typeof action.review === 'string' && singleDependency) {
|
|
@@ -241,8 +251,10 @@ export function validateDecisionProposal(proposal, {
|
|
|
241
251
|
}
|
|
242
252
|
}
|
|
243
253
|
if (action.type === 'verify') {
|
|
244
|
-
if (
|
|
245
|
-
|
|
254
|
+
if (action.review == null && action.reviewScope === 'repository') {
|
|
255
|
+
// Normalized zero-dependency audit: no artifact to review.
|
|
256
|
+
} else if (typeof action.review !== 'string') {
|
|
257
|
+
issues.push(`${at}.review must be a dotted artifact path like "outputs.<actionId>.outFile" when present; omit it to review the single/last dependency's artifact, or to audit the repository when the verify has no dependsOn`);
|
|
246
258
|
} else if (!looksLikeReviewPath(action.review)) {
|
|
247
259
|
issues.push(`${at}.review must be a dotted artifact path like "outputs.<actionId>.outFile" (the artifact to review), not instructions or a filesystem path; put reviewer instructions in ${at}.prompt`);
|
|
248
260
|
} else {
|
package/src/workflow/draft.js
CHANGED
|
@@ -58,12 +58,8 @@ export function draftPaths(bullswarmDir, name) {
|
|
|
58
58
|
return { dir: d, doc: join(d, 'workflow.json'), meta: join(d, 'meta.json') };
|
|
59
59
|
}
|
|
60
60
|
|
|
61
|
-
|
|
62
|
-
|
|
63
|
-
const tmp = `${p}.tmp-${randomBytes(3).toString('hex')}`;
|
|
64
|
-
writeFileSync(tmp, content);
|
|
65
|
-
renameSync(tmp, p);
|
|
66
|
-
}
|
|
61
|
+
// Hoisted to fsjson.js (shared with runtime/runner state persistence).
|
|
62
|
+
import { atomicWriteFileSync as atomicWrite } from './fsjson.js';
|
|
67
63
|
|
|
68
64
|
export function draftExists(bullswarmDir, name) {
|
|
69
65
|
return existsSync(draftPaths(bullswarmDir, name).doc);
|
|
@@ -0,0 +1,41 @@
|
|
|
1
|
+
// Atomic JSON persistence + torn-read-tolerant JSON reads.
|
|
2
|
+
//
|
|
3
|
+
// Doctrine:
|
|
4
|
+
// F1. Every state/report/workflow artifact is written via temp+rename so a
|
|
5
|
+
// concurrent reader can NEVER observe a half-written file (earned:
|
|
6
|
+
// `bullswarm workflow tui` crashed with "Unterminated string in JSON at
|
|
7
|
+
// position 138968" parsing state.json mid-write, 2026-08-29).
|
|
8
|
+
// F2. Display readers tolerate a missing or torn file (returns fallback):
|
|
9
|
+
// observation must never crash on the writer's timing.
|
|
10
|
+
// F3. Mutating read-modify-write readers must NOT silently no-op: they
|
|
11
|
+
// retry once (rename-atomic writes make the second read succeed) and
|
|
12
|
+
// then throw a clear error instead of dropping the user's command.
|
|
13
|
+
|
|
14
|
+
import { readFileSync, writeFileSync, renameSync, mkdirSync, existsSync } from 'node:fs';
|
|
15
|
+
import { dirname } from 'node:path';
|
|
16
|
+
import { randomBytes } from 'node:crypto';
|
|
17
|
+
|
|
18
|
+
export function atomicWriteFileSync(path, content) {
|
|
19
|
+
mkdirSync(dirname(path), { recursive: true });
|
|
20
|
+
const tmp = `${path}.tmp-${randomBytes(3).toString('hex')}`;
|
|
21
|
+
writeFileSync(tmp, content);
|
|
22
|
+
renameSync(tmp, path);
|
|
23
|
+
}
|
|
24
|
+
|
|
25
|
+
export function writeJsonAtomic(path, value) {
|
|
26
|
+
atomicWriteFileSync(path, `${JSON.stringify(value, null, 2)}\n`);
|
|
27
|
+
}
|
|
28
|
+
|
|
29
|
+
/** Display-path read: missing or torn file -> fallback, never a throw. */
|
|
30
|
+
export function readJsonSafe(path, fallback = null) {
|
|
31
|
+
if (!existsSync(path)) return fallback;
|
|
32
|
+
try { return JSON.parse(readFileSync(path, 'utf8')); } catch { return fallback; }
|
|
33
|
+
}
|
|
34
|
+
|
|
35
|
+
/** Mutating-path read: retry once, then throw a clear actionable error. */
|
|
36
|
+
export function readJsonForUpdate(path, what = 'file') {
|
|
37
|
+
try { return JSON.parse(readFileSync(path, 'utf8')); } catch { /* torn or corrupt; retry once */ }
|
|
38
|
+
try { return JSON.parse(readFileSync(path, 'utf8')); } catch (err) {
|
|
39
|
+
throw new Error(`${what} at ${path} is unreadable (${err.message}); it may be mid-write — retry the command`);
|
|
40
|
+
}
|
|
41
|
+
}
|
package/src/workflow/goal.js
CHANGED
|
@@ -11,11 +11,11 @@ export const PLANNER_RULES_SECTION = [
|
|
|
11
11
|
'1. Compile the whole program in one decision: the runtime executes all proposed actions and consults you only at a finished-or-blocked boundary, so deferring decidable work costs another round trip.',
|
|
12
12
|
'2. Make every worker prompt self-contained: include the exact goal, absolute cwd, owned files and a no-other-files boundary, expected artifact, acceptance command, and report format, because workers see only their own prompt.',
|
|
13
13
|
'3. Use short kebab-case, forward-only phases and dependsOn only for real data or same-file ordering; recovery uses a new phase and never repeats an identical failed plan.',
|
|
14
|
-
'4. For known N items, create N run plus N verify actions, each verify depending only on its own run, then one suite verify depending on all;
|
|
14
|
+
'4. For known N items, create N run plus N verify actions, each verify depending only on its own run, then one suite verify depending on all; a verify judges the artifact in review (default: its single or last dependency; a verify with no dependsOn audits the repository directly).',
|
|
15
15
|
'5. For unknown items, create discovery ending with RETURN ONLY a JSON object containing an items array, then data-driven fan-out via itemsFrom outputs.<id>.outFile or outputs.<id>.data.<field>; the runtime extracts the list and retries once read-only if needed.',
|
|
16
16
|
'6. Put outputSchema on workers whose reports are consumed or whose claims the runtime must check, so structured data is durable and can drive later fan-out.',
|
|
17
17
|
'7. Put verify.repair on every verify, and scope each verify to what can be true at its point in the graph: work scheduled later is not a defect, and cosmetic mismatches with the goal text are concerns, never ok:false. An ok:false verdict is repaired and re-checked inside the program; ok:true is accepted and concerns are informational, not extra work.',
|
|
18
|
-
'8. Add completion with all-actions-ok whenever a clean program finishes the goal; when the goal\'s acceptance checks pass, return complete rather than adding polish or alignment actions. Return complete only on durable verified evidence, never proceed, never ask the user, and stop only for a concrete unresolved blocker with a qualified outcome.',
|
|
18
|
+
'8. Add completion with all-actions-ok whenever a clean program finishes the goal; when the goal\'s acceptance checks pass, return complete rather than adding polish or alignment actions. Completion evidence requires the program\'s LAST worker to be covered by a successful verify — never leave a report or other run as the final unverified action. Return complete only on durable verified evidence, never proceed, never ask the user, and stop only for a concrete unresolved blocker with a qualified outcome.',
|
|
19
19
|
'9. Treat agent-count, workflow-duration, and expansion-round budgets as advisory planning targets, never hard stop conditions; the dispatch budget counts this planner call plus workers, verifiers, retries, and escalations. Converge as targets approach, avoid optional work, and exceed a target only for one essential bounded action or required verification.',
|
|
20
20
|
'10. This is a control-plane thread: do not invoke Bullswarm, use tools, modify files, or propose pool, addDir, taskFile, shell authority, or unbounded work; route and process authority belong to the runtime.',
|
|
21
21
|
'Shared working tree: concurrent workers editing DISJOINT files is the normal parallel mode; order shared files (indexes, barrels) after their feeders with dependsOn, and run the full suite once in a final verify — never while other workers still edit. Avoid redundant expensive verification: later verifiers reuse durable clean full-suite evidence unless it is stale or the code changed again. operatorSteering in the context is explicit operator guidance for this checkpoint: apply it within the original intent; it cannot weaken verification or expand authority.',
|
|
@@ -26,7 +26,7 @@ export const PLANNER_EXAMPLES_SECTION = [
|
|
|
26
26
|
'[{"type":"run","phase":"implement","prompt":"..."},{"type":"run","phase":"report","prompt":"...","outputSchema":{"type":"object"}},{"type":"fanout","phase":"fix","items":["alpha"],"stepTemplate":{"prompt":"Handle {{item}}."}},{"type":"fanout","phase":"fix","itemsFrom":"outputs.discover.outFile","stepTemplate":{"prompt":"Handle {{item}}."}},{"type":"verify","phase":"verify","prompt":"Check the artifact.","repair":{"prompt":"Fix rejected concerns.","maxRounds":1}}]',
|
|
27
27
|
'Complete program:',
|
|
28
28
|
'[{"id":"discover","type":"run","phase":"discover","prompt":"In /abs/repo discover items and end with RETURN ONLY a JSON object containing an items array of item names.","outputSchema":{"type":"object","properties":{"items":{"type":"array","items":{"type":"string"}}},"required":["items"]}},{"id":"fix","type":"fanout","phase":"fix","itemsFrom":"outputs.discover.data.items","stepTemplate":{"prompt":"In /abs/repo edit only the files for {{item}} and run its focused acceptance command."},"dependsOn":["discover"]},{"id":"verify-items","type":"verify","phase":"verify-items","prompt":"Independently verify every item artifact.","dependsOn":["fix"],"repair":{"prompt":"Fix each rejected item in /abs/repo and re-run its focused command.","maxRounds":2}},{"id":"verify-suite","type":"verify","phase":"verify-suite","prompt":"Run the full acceptance command in /abs/repo.","dependsOn":["verify-items"],"repair":{"prompt":"Fix the suite failure in /abs/repo and rerun the suite.","maxRounds":1}}],"completion":{"when":"all-actions-ok","reason":"The item checks and final suite verification prove the goal."}]',
|
|
29
|
-
'Rules the validator enforces: action type is run, fanout, or verify; fanout has stepTemplate and either items or itemsFrom; verify.review
|
|
29
|
+
'Rules the validator enforces: action type is run, fanout, or verify; fanout has stepTemplate and either items or itemsFrom; verify.review, when given, is outputs.<id>.outFile; dependsOn names existing or proposed actions; runtime-owned fields are rejected.',
|
|
30
30
|
].join('\n');
|
|
31
31
|
|
|
32
32
|
export const AUTONOMOUS_ORCHESTRATOR_PROMPT = [
|
package/src/workflow/runner.js
CHANGED
|
@@ -5,7 +5,9 @@
|
|
|
5
5
|
// settings.stopOnPhaseFailure: abort after a phase with any failure
|
|
6
6
|
// Resume: steps whose recorded verdict is ok:true are skipped (R2).
|
|
7
7
|
|
|
8
|
-
import { readFileSync,
|
|
8
|
+
import { readFileSync, existsSync, mkdirSync } from 'node:fs';
|
|
9
|
+
import { writeJsonAtomic } from './fsjson.js';
|
|
10
|
+
import { readSteering } from './steering.js';
|
|
9
11
|
import { join } from 'node:path';
|
|
10
12
|
import { randomBytes } from 'node:crypto';
|
|
11
13
|
import { validateWorkflow } from './validate.js';
|
|
@@ -91,7 +93,7 @@ export async function runWorkflow(opts) {
|
|
|
91
93
|
// A generated goal workflow must be restartable without the initiating
|
|
92
94
|
// process or an external draft file. The exact executable definition is
|
|
93
95
|
// therefore a first-class run artifact.
|
|
94
|
-
|
|
96
|
+
writeJsonAtomic(join(runDir, 'workflow.json'), doc);
|
|
95
97
|
}
|
|
96
98
|
|
|
97
99
|
let state;
|
|
@@ -142,7 +144,7 @@ export async function runWorkflow(opts) {
|
|
|
142
144
|
// Commit cleared terminal/control markers before WorkflowRuntime begins
|
|
143
145
|
// merging dashboard-side state. Otherwise the first resume event can
|
|
144
146
|
// re-import the stale cancelRequested marker from the interrupted run.
|
|
145
|
-
|
|
147
|
+
writeJsonAtomic(join(runDir, 'state.json'), state);
|
|
146
148
|
} else {
|
|
147
149
|
if (resuming) {
|
|
148
150
|
throw new Error(`cannot resume: no state.json for run ${runId}`);
|
|
@@ -221,7 +223,7 @@ export async function runWorkflow(opts) {
|
|
|
221
223
|
state.orchestration = doc.orchestration
|
|
222
224
|
? { ...(state.orchestration ?? {}), ...doc.orchestration }
|
|
223
225
|
: state.orchestration;
|
|
224
|
-
|
|
226
|
+
writeJsonAtomic(join(runDir, 'workflow.json'), doc);
|
|
225
227
|
}
|
|
226
228
|
if (opts.inputs && Object.keys(opts.inputs).length) {
|
|
227
229
|
state.inputs = { ...state.inputs, ...opts.inputs };
|
|
@@ -447,6 +449,20 @@ export async function runWorkflow(opts) {
|
|
|
447
449
|
state.recovery = { resumable: true, signal: interruptionSignal, interruptedAt: finishedAt };
|
|
448
450
|
}
|
|
449
451
|
if (abortReason && !interrupted) state.abortReason = abortReason;
|
|
452
|
+
// Steering queued after the last planner gate must end truthfully: on a
|
|
453
|
+
// real terminal transition (not interrupted — a resume still has gates
|
|
454
|
+
// ahead — and not waiting for approval), mark undelivered entries expired.
|
|
455
|
+
if (!waitingForApproval && !interrupted) {
|
|
456
|
+
const undelivered = readSteering(runDir).filter((entry) =>
|
|
457
|
+
!(state.steering ?? []).some((known) => known.id === entry.id));
|
|
458
|
+
if (undelivered.length) {
|
|
459
|
+
state.steering = [
|
|
460
|
+
...(state.steering ?? []),
|
|
461
|
+
...undelivered.map((entry) => ({ ...entry, status: 'expired_undelivered', expiredAt: finishedAt })),
|
|
462
|
+
];
|
|
463
|
+
runtime.emit('steering.expired', { steeringIds: undelivered.map((entry) => entry.id) });
|
|
464
|
+
}
|
|
465
|
+
}
|
|
450
466
|
runtime.persist();
|
|
451
467
|
|
|
452
468
|
const preliminaryReport = buildReport(state, doc, runDir);
|
|
@@ -456,7 +472,7 @@ export async function runWorkflow(opts) {
|
|
|
456
472
|
: state.status === 'blocked' ? 'run.blocked' : 'run.completed';
|
|
457
473
|
runtime.emit(terminalEvent, { runId, status: state.status, report: preliminaryReport.summary, outcome: state.outcome ?? null });
|
|
458
474
|
const report = buildReport(state, doc, runDir);
|
|
459
|
-
|
|
475
|
+
writeJsonAtomic(join(runDir, 'report.json'), report);
|
|
460
476
|
opts.onEvent?.({ type: 'workflow.completed', runId, status: state.status, report: report.summary });
|
|
461
477
|
|
|
462
478
|
return { runId, runDir, state, report };
|
|
@@ -845,6 +861,41 @@ async function runDecisionLoop({ runtime, gate, phase, state, retryAttempts }) {
|
|
|
845
861
|
|
|
846
862
|
// Resume accepted expansion work before asking the planner for a new
|
|
847
863
|
// semantic decision. Successful actions are skipped by durable output.
|
|
864
|
+
// An action the interruption cancelled mid-flight (status `cancelled`,
|
|
865
|
+
// output `ok:false` "workflow cancellation requested") is unfinished work,
|
|
866
|
+
// not a failure: clear its output so it re-runs, and clear the outputs of
|
|
867
|
+
// the dependents that were blocked only because of it — otherwise resume
|
|
868
|
+
// asks the planner to re-plan around a phantom failure (observed on run
|
|
869
|
+
// djnjka, 2026-08-29: 1 cancelled action → 4 "blocked" → spurious turn).
|
|
870
|
+
const ledgerById = new Map((state.actionLedger ?? []).map((entry) => [entry.id, entry]));
|
|
871
|
+
const reopened = new Set();
|
|
872
|
+
for (const entry of state.plan?.actions ?? []) {
|
|
873
|
+
if (entry.source !== 'planner' || !entry.definition) continue;
|
|
874
|
+
if (ledgerById.get(entry.id)?.status === 'cancelled') reopened.add(entry.id);
|
|
875
|
+
}
|
|
876
|
+
let grew = reopened.size > 0;
|
|
877
|
+
while (grew) {
|
|
878
|
+
grew = false;
|
|
879
|
+
for (const entry of state.plan?.actions ?? []) {
|
|
880
|
+
if (entry.source !== 'planner' || !entry.definition || reopened.has(entry.id)) continue;
|
|
881
|
+
const blocked = state.outputs[entry.id]?.dependencyBlocked === true;
|
|
882
|
+
if (blocked && (entry.definition.dependsOn ?? []).some((id) => reopened.has(id))) {
|
|
883
|
+
reopened.add(entry.id);
|
|
884
|
+
grew = true;
|
|
885
|
+
}
|
|
886
|
+
}
|
|
887
|
+
}
|
|
888
|
+
for (const id of reopened) {
|
|
889
|
+
delete state.outputs[id];
|
|
890
|
+
const ledger = ledgerById.get(id);
|
|
891
|
+
if (ledger) {
|
|
892
|
+
ledger.status = 'pending';
|
|
893
|
+
delete ledger.why;
|
|
894
|
+
delete ledger.finishedAt;
|
|
895
|
+
}
|
|
896
|
+
runtime.emit('action.reopened', { actionId: id, reason: 'interrupted before completion' });
|
|
897
|
+
}
|
|
898
|
+
if (reopened.size) runtime.persist();
|
|
848
899
|
const unfinishedAccepted = (state.plan?.actions ?? [])
|
|
849
900
|
.filter((entry) => entry.source === 'planner' && entry.definition &&
|
|
850
901
|
state.outputs[entry.id] == null)
|
|
@@ -1077,7 +1128,13 @@ async function runDecisionLoop({ runtime, gate, phase, state, retryAttempts }) {
|
|
|
1077
1128
|
.map((entry) => entry.id);
|
|
1078
1129
|
const dynamicActions = (state.actionLedger ?? []).filter((action) => action.parentId === gate.id);
|
|
1079
1130
|
const gaps = failing.length ? [] : completionEvidenceGaps(dynamicActions, state.orchestration?.completionPolicy, state.outputs);
|
|
1080
|
-
|
|
1131
|
+
// Pending operator steering blocks self-completion: the documented
|
|
1132
|
+
// contract is delivery at the next planner gate, so a clean program
|
|
1133
|
+
// returns to the planner (which delivers the steer) instead of
|
|
1134
|
+
// silently discarding it (defect observed on wf-mtds7tzx, 2026-08-29).
|
|
1135
|
+
const pendingSteering = (failing.length || gaps.length) ? [] : readSteering(runtime.runDir).filter((entry) =>
|
|
1136
|
+
!(state.steering ?? []).some((known) => known.id === entry.id));
|
|
1137
|
+
if (!failing.length && !gaps.length && !pendingSteering.length) {
|
|
1081
1138
|
const verifyIds = programActions.filter((entry) => entry.kind === 'verify').map((entry) => entry.id);
|
|
1082
1139
|
const reason = proposal.completion.reason?.trim()
|
|
1083
1140
|
|| `Program completed: all ${programActions.length} actions finished ok, verified by ${verifyIds.join(', ')}.`;
|
|
@@ -1114,9 +1171,18 @@ async function runDecisionLoop({ runtime, gate, phase, state, retryAttempts }) {
|
|
|
1114
1171
|
runtime.persist();
|
|
1115
1172
|
return { ok: true, why: reason, complete: true, decision: { decision: 'complete', reason, actions: [], completion: proposal.completion } };
|
|
1116
1173
|
}
|
|
1117
|
-
|
|
1118
|
-
|
|
1119
|
-
|
|
1174
|
+
if (pendingSteering.length) {
|
|
1175
|
+
runtime.emit('decision.completion_deferred', {
|
|
1176
|
+
gateId: gate.id,
|
|
1177
|
+
programSequence: decision.sequence,
|
|
1178
|
+
reason: 'operator steering pending',
|
|
1179
|
+
steeringIds: pendingSteering.map((entry) => entry.id),
|
|
1180
|
+
});
|
|
1181
|
+
} else {
|
|
1182
|
+
runtime.emit('decision.completion_predicate_unmet', {
|
|
1183
|
+
gateId: gate.id, programSequence: decision.sequence, failing, gaps,
|
|
1184
|
+
});
|
|
1185
|
+
}
|
|
1120
1186
|
}
|
|
1121
1187
|
// Loop intentionally returns to observation and invokes the planner again.
|
|
1122
1188
|
}
|
package/src/workflow/runs-cli.js
CHANGED
|
@@ -14,6 +14,7 @@
|
|
|
14
14
|
// resolver in short-id.js maps both to the run directory.
|
|
15
15
|
|
|
16
16
|
import { existsSync, rmSync, readFileSync } from 'node:fs';
|
|
17
|
+
import { readJsonSafe } from './fsjson.js';
|
|
17
18
|
import { join } from 'node:path';
|
|
18
19
|
import { listRuns, resolveRunId, isOngoing } from './short-id.js';
|
|
19
20
|
import { BULLSWARM_DIR } from './cli.js';
|
|
@@ -161,8 +162,8 @@ function runsShow(idToken, opts) {
|
|
|
161
162
|
const { runId, runDir } = resolved;
|
|
162
163
|
const statePath = join(runDir, 'state.json');
|
|
163
164
|
const reportPath = join(runDir, 'report.json');
|
|
164
|
-
const state =
|
|
165
|
-
const report =
|
|
165
|
+
const state = readJsonSafe(statePath);
|
|
166
|
+
const report = readJsonSafe(reportPath);
|
|
166
167
|
const ongoing = isOngoing(runDir, state);
|
|
167
168
|
|
|
168
169
|
if (opts.json) {
|
|
@@ -191,8 +192,8 @@ function runsResult(idToken, opts) {
|
|
|
191
192
|
const { runId, runDir } = resolved;
|
|
192
193
|
const statePath = join(runDir, 'state.json');
|
|
193
194
|
const reportPath = join(runDir, 'report.json');
|
|
194
|
-
const state =
|
|
195
|
-
const report =
|
|
195
|
+
const state = readJsonSafe(statePath);
|
|
196
|
+
const report = readJsonSafe(reportPath);
|
|
196
197
|
const ongoing = isOngoing(runDir, state);
|
|
197
198
|
const result = buildWorkflowResult({
|
|
198
199
|
state, report, runId, shortId: resolved.shortId, ongoing,
|
|
@@ -246,8 +247,15 @@ function runsDelete(idToken, opts, rest) {
|
|
|
246
247
|
// Refuse to delete an ongoing run without --force. Half-finished
|
|
247
248
|
// runs are usually a debugging target, not garbage.
|
|
248
249
|
const statePath = join(runDir, 'state.json');
|
|
249
|
-
|
|
250
|
-
|
|
250
|
+
// Delete guard: an unreadable (mid-write) state means the run may be live —
|
|
251
|
+
// treat it as ongoing rather than deleting a live run; --force still wins.
|
|
252
|
+
let state = null;
|
|
253
|
+
let stateUnreadable = false;
|
|
254
|
+
if (existsSync(statePath)) {
|
|
255
|
+
state = readJsonSafe(statePath, undefined);
|
|
256
|
+
if (state === undefined) { state = null; stateUnreadable = true; }
|
|
257
|
+
}
|
|
258
|
+
const ongoing = stateUnreadable ? true : isOngoing(runDir, state);
|
|
251
259
|
if (ongoing && !opts.force) {
|
|
252
260
|
return err(
|
|
253
261
|
`refusing to delete ongoing run "${runId}" (shortId ${shortId ?? '?'}); ` +
|
package/src/workflow/runtime.js
CHANGED
|
@@ -22,6 +22,7 @@
|
|
|
22
22
|
// text always lives in the per-step outFile.
|
|
23
23
|
|
|
24
24
|
import { mkdirSync, writeFileSync, readFileSync, existsSync } from 'node:fs';
|
|
25
|
+
import { writeJsonAtomic } from './fsjson.js';
|
|
25
26
|
import { join } from 'node:path';
|
|
26
27
|
import { createHash, randomUUID } from 'node:crypto';
|
|
27
28
|
import { pickPool, isQuarantined } from '../lib/route.js';
|
|
@@ -92,6 +93,24 @@ export function plannerBudgetContext(budget = {}) {
|
|
|
92
93
|
};
|
|
93
94
|
}
|
|
94
95
|
|
|
96
|
+
/** Skeleton of a rejected planner proposal: shapes and ids, prompts elided. */
|
|
97
|
+
export function compactRejectedProposal(proposal) {
|
|
98
|
+
if (!proposal || typeof proposal !== 'object') return proposal ?? null;
|
|
99
|
+
const elide = (text) => (typeof text === 'string' && text.length > 160 ? `${text.slice(0, 160)}…[${text.length} chars]` : text);
|
|
100
|
+
return {
|
|
101
|
+
...proposal,
|
|
102
|
+
reason: elide(proposal.reason),
|
|
103
|
+
actions: Array.isArray(proposal.actions) ? proposal.actions.map((action) => {
|
|
104
|
+
if (!action || typeof action !== 'object') return action;
|
|
105
|
+
const out = { ...action };
|
|
106
|
+
if ('prompt' in out) out.prompt = elide(out.prompt);
|
|
107
|
+
if (out.repair && typeof out.repair === 'object') out.repair = { ...out.repair, prompt: elide(out.repair.prompt) };
|
|
108
|
+
if (out.stepTemplate && typeof out.stepTemplate === 'object') out.stepTemplate = { ...out.stepTemplate, prompt: elide(out.stepTemplate.prompt) };
|
|
109
|
+
return out;
|
|
110
|
+
}) : proposal.actions,
|
|
111
|
+
};
|
|
112
|
+
}
|
|
113
|
+
|
|
95
114
|
export class WorkflowRuntime {
|
|
96
115
|
/**
|
|
97
116
|
* @param {object} opts
|
|
@@ -184,7 +203,8 @@ export class WorkflowRuntime {
|
|
|
184
203
|
if (disk.cancellingAt) this.state.cancellingAt = disk.cancellingAt;
|
|
185
204
|
}
|
|
186
205
|
} catch { /* a partial state file should not break the workflow */ }
|
|
187
|
-
|
|
206
|
+
// Atomic: a 1s-interval TUI reads this file while we write it (F1).
|
|
207
|
+
writeJsonAtomic(join(this.runDir, 'state.json'), this.state);
|
|
188
208
|
}
|
|
189
209
|
|
|
190
210
|
refreshCancellation() {
|
|
@@ -1036,13 +1056,16 @@ export class WorkflowRuntime {
|
|
|
1036
1056
|
*/
|
|
1037
1057
|
async runVerify(step, scope, opts = {}) {
|
|
1038
1058
|
this.enforceRequiredInputs(step.id);
|
|
1039
|
-
if (!step.review) {
|
|
1059
|
+
if (!step.review && step.reviewScope !== 'repository') {
|
|
1040
1060
|
throw new Error(
|
|
1041
1061
|
`verify step "${step.id}" needs a "review" path ` +
|
|
1042
1062
|
`(e.g. review: "outputs.<priorStep>.outFile")`,
|
|
1043
1063
|
);
|
|
1044
1064
|
}
|
|
1045
1065
|
const reviewedText = (() => {
|
|
1066
|
+
if (!step.review) {
|
|
1067
|
+
return '(no artifact: this verify has no upstream action — audit the repository state directly with fresh commands)';
|
|
1068
|
+
}
|
|
1046
1069
|
try {
|
|
1047
1070
|
// `review` is a dotted path into the scope, NOT a template
|
|
1048
1071
|
// (matches the design of `fanout.itemsFrom`). Resolve it the
|
|
@@ -1297,8 +1320,14 @@ export class WorkflowRuntime {
|
|
|
1297
1320
|
maxAttempts: opts.correction.maxAttempts,
|
|
1298
1321
|
why: opts.correction.why,
|
|
1299
1322
|
issues: opts.correction.issues ?? [],
|
|
1300
|
-
|
|
1301
|
-
|
|
1323
|
+
// The planner's own thread already holds the full rejected proposal
|
|
1324
|
+
// (worker prompts included); resend only its skeleton so a corrective
|
|
1325
|
+
// turn cannot re-inflate the compacted context (observed: 38k-char
|
|
1326
|
+
// validationFeedback on run 4t6m5a, 37k of it the verbatim proposal).
|
|
1327
|
+
rejectedProposal: compactRejectedProposal(opts.correction.rejectedProposal),
|
|
1328
|
+
rejectedResponseExcerpt: typeof opts.correction.rejectedResponse === 'string'
|
|
1329
|
+
? opts.correction.rejectedResponse.slice(0, 2000)
|
|
1330
|
+
: null,
|
|
1302
1331
|
} : null,
|
|
1303
1332
|
operatorSteering: (this.state.steering ?? []).map((entry) => ({
|
|
1304
1333
|
id: entry.id,
|
package/src/workflow/short-id.js
CHANGED
|
@@ -16,7 +16,8 @@
|
|
|
16
16
|
// not the generator.
|
|
17
17
|
|
|
18
18
|
import { randomBytes } from 'node:crypto';
|
|
19
|
-
import { readdirSync, readFileSync,
|
|
19
|
+
import { readdirSync, readFileSync, existsSync, statSync } from 'node:fs';
|
|
20
|
+
import { writeJsonAtomic } from './fsjson.js';
|
|
20
21
|
import { join } from 'node:path';
|
|
21
22
|
import { appendEvent } from './events.js';
|
|
22
23
|
import { aggregateUsage } from '../lib/usage.js';
|
|
@@ -235,7 +236,7 @@ export function reconcileInterruptedRun(runDir, state, {
|
|
|
235
236
|
delete state.currentPhase;
|
|
236
237
|
delete state.currentStep;
|
|
237
238
|
appendEvent(runDir, state, 'run.interrupted_reconciled', { reason, resumable: true });
|
|
238
|
-
|
|
239
|
+
writeJsonAtomic(statePath, state);
|
|
239
240
|
return state;
|
|
240
241
|
}
|
|
241
242
|
|