bullswarm 0.14.0 → 0.15.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/CHANGELOG.md CHANGED
@@ -1,5 +1,59 @@
1
1
  # bullswarm changelog
2
2
 
3
+ ## 0.15.0 — extra Claude Code logins as separate pools
4
+
5
+ - Claude Code extra logins (`~/.claude-<slug>` / `$CLAUDE_CONFIG_DIR`) become
6
+ their own pools (`claude-code:<slug>`), metered and spawned with
7
+ `CLAUDE_CONFIG_DIR` set. Discovery is dynamic from the filesystem; there
8
+ is no hardcoded extra-profile list. The spawn command for each profile is
9
+ `CLAUDE_CONFIG_DIR=<dir> claude`.
10
+
11
+ ## 0.14.1 — the TUI survives its writer; steering lands or expires truthfully
12
+
13
+ Proven on a goal-3 re-run (`d7xyg2`): 1 planner turn / 269 s (baseline 0.13.1: 1 / 294 s), planner context 6.2 k chars (from 32.7 k), auto-completed, deliverable verified, zero observation crashes — `docs/experiments/2026-08-29-dogfood-bullswarm-builds-bullswarm.md`.
14
+
15
+ - Workflow `state.json`/`report.json`/`workflow.json` writes are atomic
16
+ (temp + rename, new `src/workflow/fsjson.js`): a concurrent reader can never
17
+ observe a half-written file. Earned: `workflow tui` crashed with
18
+ "Unterminated string in JSON at position 138968" parsing `state.json`
19
+ mid-write (observed twice, 2026-08-29).
20
+ - Observation readers tolerate torn or missing JSON: the TUI keeps painting
21
+ the last good frame of the same run, and a render or key-handler error is
22
+ shown in the message line instead of killing the process and stranding the
23
+ terminal in alt-screen raw mode. Mutating commands (stop, approval) retry
24
+ the read once and then refuse loudly instead of silently dropping the
25
+ operator's command. `runs delete` treats an unreadable `state.json` as
26
+ ongoing (refuses without `--force`) rather than deleting a possibly-live run.
27
+ - An action being re-run (repair round, re-verify, schema retry) reads as
28
+ `running` and its phase as `active` even when its previous round recorded
29
+ `ok:false`; a failed mark now means failed-and-not-being-retried. (User
30
+ report: the TUI showed ✗ "2/2 complete" beside a live spinner.)
31
+ - Pending operator steering defers program self-completion: a clean program
32
+ with `completion: all-actions-ok` returns to the planner gate (event
33
+ `decision.completion_deferred`), which delivers the steer — instead of
34
+ auto-completing and silently discarding it (defect observed live:
35
+ 0 `steering.delivered` events for a queued steer). Steering that can no
36
+ longer reach any gate is marked `expired_undelivered` with event
37
+ `steering.expired` at the terminal transition; interrupted runs keep their
38
+ queue for the resumed run's next gate.
39
+ - Resume re-runs an action the interruption cancelled mid-flight instead of
40
+ re-planning around a phantom failure: cancelled actions and the dependents
41
+ blocked only by them are reopened (event `action.reopened`) and the accepted
42
+ program continues from where it stopped. Observed on a SIGTERM-interrupted
43
+ run: 1 cancelled action → 4 "blocked" → a spurious planner turn.
44
+ - Planner contract: a verify with several `dependsOn` must set `review`
45
+ (rule 4); the program's last worker must be covered by a successful verify
46
+ (rule 8) — both were the causes of extra planner gates on the goal-3 proof
47
+ run. A corrective turn's `validationFeedback.rejectedResponseExcerpt` is
48
+ capped at 2 000 chars, and the rejected proposal is resent as a skeleton
49
+ (ids, shapes, dependsOn; prompts elided) — the planner's thread already
50
+ holds it verbatim.
51
+ - A verify without `review` is no longer grounds to reject a whole program:
52
+ it reviews its single (or last) dependency's artifact, or audits the
53
+ repository directly when it has no `dependsOn` (`reviewScope: repository`).
54
+ Observed on two proof runs: a 9-action program bounced for one field,
55
+ costing a 5-minute correction turn each time.
56
+
3
57
  ## 0.14.0 — structured worker output, compact planner contract
4
58
 
5
59
  - A verify whose reply cannot be parsed as the verdict JSON gets ONE bounded
@@ -0,0 +1,171 @@
1
+ # Dogfood 2026-08-29 — bullswarm builds bullswarm (outputSchema + planner refactor)
2
+
3
+ Observation log of the two dogfood runs that produced 0.14.0, kept verbatim
4
+ from the driving session's notes; the post-run defects below produced 0.14.1.
5
+
6
+
7
+ Goal file: `goal4.txt` (7 deliverables, single-implementer constraint). Repo branch: `feat/output-schema`.
8
+ Runtime: local worktree `bullswarm-rt` pinned at main (user: redispatch with local runtime, keep committing; no npm wait).
9
+ Home: default `~/.bullswarm` (heterogeneous pools) — user can `bullswarm workflow tui <shortId>`.
10
+
11
+ ## Attempt 1 — 01:43:37 Z, installed 0.13.1 — failed at validation, nothing ran
12
+ `autonomous workflow invalid (nothing ran): phases[0].steps[0](scout): template ref "{{outputs.x.data.field}}" cannot resolve`.
13
+ Cause: goal text spliced into the scout prompt; validator parsed a quoted ref in user text.
14
+ Fix (bullswarm defect #1): goal → declared `inputs.goal`, inserted at render time; unresolved grammar-valid refs are
15
+ left literal + `template.unresolved_ref` event instead of fatal. Commit `7badea3`, released 0.13.2 (`a0f0965`).
16
+ Verified: scout task file of attempt 3 contains `{{outputs.x.data.field}}` verbatim; only `{{inputs.goal}}` resolved.
17
+
18
+ ## Attempt 2 — 01:47:37 Z — launcher bug (mine, not bullswarm): zsh does not word-split `$BS="node path"`. Fixed script.
19
+
20
+ ## Attempt 3 — 01:48:19 Z, runtime 0.13.2 @ a0f0965 — run `wf-mtdq1l9v-ed22fe` / `75t4n2`
21
+ - 01:48:22 scout → pool `opencode2` model `kaihk/gpt-5.6-luna`. Routing: "most-behind capable pool (surplus 0)";
22
+ candidates opencode2 pace 0 (unmetered) > grok −7.5 > claude-code −9.4 > codex −49.1.
23
+ Observation: an unmetered pool reads as exactly on pace (0) and therefore outranks every metered pool that is
24
+ ahead of pace. Not a crash, but quota-unknown pools capture all work whenever metered pools are burning ahead.
25
+ - 01:48:22 → 01:49:54 scout ok (92 s, opencode2).
26
+ - 01:50:06 → 02:00:40 planner turn 1 (**~634 s** — roughly 2× the goal-2/3 turns; goal text is 6.5 k chars and the
27
+ program is 12 actions). Decision: `needs_more_work`, 12 actions, `completion: all-actions-ok` attached.
28
+ Program honours the single-implementer constraint: `impl-src` (all src) ∥ `docs`; `tests-schema`/`tests-adaptive`/
29
+ `tests-gaps` + `verify-impl(repair 2)` fan out after impl-src; per-writer verifies with repair; `final-report` →
30
+ `verify-suite(repair 2)`. Note: `verify-impl`'s repair may edit src while the test writers read it — accepted risk,
31
+ `verify-suite` runs the whole suite at the end.
32
+ - 02:00:31 `impl-src` and `docs` → opencode2 `kaihk/gpt-5.6-luna` (medium effort, "most-behind capable pool (surplus 0)").
33
+ All build work lands on the unmetered pool while claude-code/grok/codex are ahead of pace. Orchestrator stayed on
34
+ claude-code (pinned).
35
+ - 02:00:31 → 02:07:09 `impl-src` ok (398 s, opencode2); `docs` ok (140 s). 02:07:11 five actions started at once
36
+ (3 test writers + verify-impl + verify-docs).
37
+ - 02:08:20 `verify-impl` ok:false → `verify-impl-repair-1` (concerns were concrete and correct: schema JSON omitted from
38
+ the retry task text; fan-out item resume skipped schemaOk; no combined run+fanout skeleton). 02:08:36 `verify-docs`
39
+ ok:false → `verify-docs-repair-1` (doc claimed fan-out persists top-level data/schemaOk; doc claimed static validator
40
+ rejects outputSchema on verify but only decision.js did). Repairs 121 s / 124 s. Both are repair-loop live uses; if the
41
+ program then self-completes this is the first live exercise of the 0.13.1 fix path.
42
+ - 02:12 → 02:18 second round of repairs: `verify-impl-repair-2` (escalation could allow a 2nd schema retry — legit design
43
+ nit; "schema.js untracked so absent from git diff --stat" — verifier misreading), `verify-adaptive-tests-repair-1/2`
44
+ (rejected twice for the same reason: "diff is not purely additive: modifies an existing assertion" — the
45
+ programFeatures assertion HAD to change; an unrepairable process criterion). Both verifies ended ok:false after
46
+ maxRounds → `final-report`/`verify-suite` blocked (`failed_terminal: dynamic actions blocked by failed or unresolved
47
+ dependencies`) → program boundary → planner turn 2 at 02:18:52.
48
+ Observation (behaviour): a verifier that judges process criteria ("purely additive", "untracked file") instead of the
49
+ goal's acceptance checks produces rejections no repair can satisfy; two repair rounds (~10 min) were spent before the
50
+ boundary. Candidate for the prompt audit: verify doctrine "ok:false only for failed acceptance checks; process
51
+ observations are concerns" and/or runtime: identical concerns after a repair → boundary immediately.
52
+ - 02:18:52 → ~02:25:20 planner turn 2 (~390 s): `needs_more_work`, 3 sequential actions — `recheck-src` (verify, repair 2,
53
+ depends on verify-impl-repair-2) → `full-suite-report` (run) → `verify-final` (verify, repair 2). Reason correctly
54
+ notes no whole-suite evidence existed after the repairs (impl-src's npm test predated them).
55
+ - 02:25:52 → 02:32:36 follow-up program: `recheck-src` ok:false → `recheck-src-repair-1` → re-verify ok; `full-suite-report`
56
+ ok; `verify-final` ok → **`decision.auto_completed` (program-completion)** at 02:32:39.
57
+ ## Result (attempt 3)
58
+ - Wall **44 min 17 s** (2 654 s), 28 dispatches (26 opencode2 + 2 planner on claude-code), max concurrent 5,
59
+ parallelism 1.5, planner 2 turns / 1 045 s (39 % — turn 1 alone 637 s), repairs 6 (2 repaired ok, 4 re-verify
60
+ rejected), tokens ≈ 108 k (estimate).
61
+ - Deliverable: 10 files changed + 2 new (489+/38−); `npm test` **318/318** (299 + 19) on my own run; validator module,
62
+ decision/validate/runtime/runner/template changes all present; docs + changelog written.
63
+ - My review: sound; one robustness flaw fixed by hand — `readTrailingObject` used a reverse brace/quote scanner whose
64
+ escape handling is wrong scanning backwards (a `\"` inside a string could derail it and waste the single retry);
65
+ replaced with the parse-candidates approach `hasStructuredAnswer` already uses. Double failure now reports the
66
+ retry's errors. Escalation concern from verify-impl judged mistaken (escalation follows failed dispatches only).
67
+ - 0.13.1 fix path: NOT exercised here either — the latest worker at completion was `full-suite-report`, verified by a
68
+ direct edge (`verify-final`).
69
+ ## Adversarial review via `bullswarm run --lane analyze` (02:36 → 02:39 Z, 206 s, opencode2 kaihk/gpt-5.6-luna)
70
+ - Verdict "do not release" with **two confirmed, reproduced defects** — both in code I had reviewed and passed:
71
+ (1) my rewritten `readTrailingObject` returned the first schema-valid `{…}` from the right, so a nested object
72
+ could be recorded (`{"wrapper":{"ok":"inner"}}` → data `{"ok":"inner"}`); (2) `schema.js` used `in`, so
73
+ `toString`/`constructor`/`__proto__` counted as present/declared. Cleared: escaped quotes, fenced JSON, resume rules,
74
+ exactly-one retry, validator paths. Fixed + regression tests (320/320); my fix also had an infinite loop when output
75
+ starts with `{` (lastIndexOf clamps negative fromIndex) — caught by the suite hanging, fixed.
76
+ - Evidence for open item "adversarial verification by default": a 3-minute refute-framed review found what the
77
+ run's own verifies (6 rounds) and my manual review both missed.
78
+ ## Goal 5 — planner context/contract refactor — run `wf-mtds7tzx-95ab05` / `ejk9w2`, 02:49:09 Z, runtime 0.13.2 local
79
+ - scout 72 s (opencode2); planner turn 1 02:50:55 → 02:58:56 (~480 s): 6 actions, completion attached; noticed the goal's
80
+ stale "318 passing" and used the real 320. Program: impl-src → verify-src(repair 2) → update-tests ∥ update-docs →
81
+ verify-tests(repair 2) → verify-suite(repair 1).
82
+ - 02:58:56 → 03:09:04 impl-src (~610 s). verify-src rejected twice, both times on SUBSTANTIVE spec points (obsolete
83
+ skeleton text left in a comment; 6 JSON examples instead of 2; `<item>` instead of `{{item}}` in examples; excerpt
84
+ policy). One misread to check in the final diff: it called the existing 3 000/36 000-char excerpt caps a violation of
85
+ "full excerpt" although the goal said "(existing budget logic)" — the repair may have removed the caps.
86
+ Ordering tension: verify-src runs before update-tests, so it necessarily sees 5 failing old assertions; the planner
87
+ should either fold assertion updates into impl-src or make verify-src judge src only.
88
+ - Goal 5 finished 04:00:40 Z (71.5 min, 3 planner turns + program-completion, 15 dispatches, plannerSec 1 105+175).
89
+ Where the time went: 18–20 min planner turns; 12.4 min verify-src repair loop enforcing my over-exacting spec and
90
+ judging intermediate state; ~8 min decision-3 nit round (`align-prefix-number` + `confirm-docs`) triggered by an
91
+ UNPARSEABLE final-check verdict; the queued steer (03:57:32, "converge now") was NEVER delivered — 0 steering.delivered events; the run auto-completed (program-completion, 04:00:40) and deliverSteering only runs at planner gates, so auto-completion silently discards pending operator steering. Known issue; convergence came from the runtime, not the steer.
92
+ Deliverable reviewed + committed `15f1534`: 10-rule contract (2.2 k chars) + 2 examples (1.9 k) replace 16.3 k prefix;
93
+ compact ledger rows; id-only failures; 200-char stale excerpts; `decision.context_built` size event; 323/323.
94
+ - My follow-up (commit after 15f1534): verify verdict parse failure → ONE bounded re-ask (`verify.verdict_retry`,
95
+ test with a garbled-once verifier, 324/324); contract amendments (verify scoping, converge-not-polish, restored
96
+ shared-tree/redundant-verification/operatorSteering lines the merge dropped); re-budgeted "full" excerpts (the
97
+ uncapped version could have rebuilt the 163 k contexts).
98
+ - Speed answer to the user: ~35 of 66 min (at question time) was real work; fixes target the rest — re-ask (−8 min),
99
+ verify scoping (−12 min), convergence rule (−nit rounds). Remaining lever: planner turn latency itself (Opus
100
+ high-effort per boundary; context compaction cuts cost ~7×, latency is model thinking time).
101
+
102
+ ## Post-run defects → 0.14.1 (fixed directly, dogfooding paused by user direction)
103
+ - `workflow tui` crashed twice (`detailRow` dashboard.js:809 `JSON.parse` of
104
+ state.json mid-write; the throw escaped the repaint timer and killed the TUI,
105
+ stranding the terminal in alt-screen raw mode). Fix: atomic temp+rename
106
+ writes for state/report/workflow.json + torn-read-tolerant observation
107
+ readers + guarded paint/key handlers with last-good-frame fallback.
108
+ - TUI showed phase ✗ "2/2 complete" while a re-verify attempt was live. Fix:
109
+ an action with an active agent reads as running; phase precedence
110
+ active > failed > completed.
111
+ - The goal-5 steer was never delivered: auto-completion bypassed the planner
112
+ gate and silently discarded pending steering (0 steering.delivered events).
113
+ Fix: pending steering defers self-completion to the planner
114
+ (`decision.completion_deferred`); undeliverable steering is marked
115
+ `expired_undelivered` (`steering.expired`) at the terminal transition.
116
+ - Routing concentration on the unmetered pool (26/28 dispatches) confirmed as
117
+ design intent (quota protection outranks diversity) and documented in the
118
+ skill rather than changed.
119
+
120
+ ## Proof run 1 — goal-3 re-run `d8pr8s` (wf-mtdvuk9m), 0.14.1-pre @ 0bbe78c
121
+ Baseline (0.13.1, same fixture/goal/pool/flags): 28m42s wall, 1 planner turn, 294 s planner.
122
+ - **completed + verified, deliverable exactly right** (csv+slugify guarded, existing tests byte-identical, 63/63),
123
+ zero crashes, 8 workers, no repairs.
124
+ - Wall **32m19s**, planner **5 dispatches / 463 s** (195+114+58+66+30). Per-turn latency DOWN (max 195 s vs 294 s);
125
+ context per turn 6.2k–48k chars (`decision.context_built` measuring itself) vs the old 16.3k prefix + up-to-178k contexts.
126
+ - The 3 extra gates, each diagnosed and fixed in `c0ff947`:
127
+ 1. first proposal rejected — verify with several dependsOn lacked `review` → contract rule 4 now states it;
128
+ 2. evidence-policy boundary + rejected `complete` — final-report left as last unverified worker → rule 8 now
129
+ states the LAST worker must be covered by a verify;
130
+ 3. one deliberate operator steer — which **live-proved the 0.14.1 steering fix**: `decision.completion_deferred`
131
+ → `steering.delivered` (first ever observed; the goal-5 defect showed 0) → planner turn honoured it.
132
+ - Also observed and fixed: a corrective turn re-inflated validationFeedback to 24k chars (raw response duplicated
133
+ the parsed proposal) → excerpt capped at 2k.
134
+ - 0.13.1 completion-evidence policy exercised live for the first time: it refused auto-completion twice, correctly.
135
+
136
+ ## Proof run 2 — goal-3 re-run `djnjka` (wf-mtdx5htt), 0.14.1-pre @ c0ff947 — interrupted, then cancelled
137
+ - Turn-1 proposal **accepted first try** (no validation rejection, no correction turn): contract fix #1 confirmed.
138
+ Turn-1 context 6,201 chars.
139
+ - At 05:21 the driving session's background task was killed by the harness (not the user); SIGTERM reached the
140
+ runner, which persisted `interrupted` + resumable (1/3 steps, in-flight `triage` cancelled) — the 0.13 interruption
141
+ path working as designed.
142
+ - Resume at 05:54 exposed a **resume defect**: the cancelled `triage` was persisted `ok:false` ("workflow
143
+ cancellation requested"), so its 4 dependents were marked "blocked by failed or unresolved dependencies" and the
144
+ planner was asked to re-plan around a failure that never happened. Cancelled the run; fixed in `459c58c`
145
+ (cancelled actions + dependents blocked only by them are reopened on resume, event `action.reopened`; regression
146
+ test drives a cancel marker into a slow in-flight action and asserts exactly one further planner turn).
147
+
148
+ ## Proof run 3 — goal-3 re-run `4t6m5a` (wf-mtdyyqkw), 0.14.1-pre @ 459c58c — cancelled after diagnosis
149
+ - Turn-1 proposal (9 actions, sound shape: 3 disjoint impl workers ∥ audit of untouched modules → per-module
150
+ verifies → final-report → verify-final, completion attached) was **rejected** for one field: a zero-dependsOn
151
+ audit verify carried no `review`. Correction turn cost ~5 min and re-inflated context to 47.8k chars
152
+ (validationFeedback 38k — the 2k cap on rejectedResponseExcerpt was insufficient because rejectedProposal
153
+ itself is 37k). Root cause is the validator's posture, not the planner: rejecting a whole program for a field the
154
+ runtime can default. Fixed in `548eabe`: review defaults to the single/last dependency's artifact, or
155
+ `reviewScope: repository` for a no-dependency audit; contract rule 4 reworded. Run cancelled to re-prove cleanly.
156
+
157
+ ## Proof run 4 — goal-3 re-run `d7xyg2` (wf-mtdzhw88), 0.14.1-pre @ 548eabe — **PASS**
158
+ | metric | 0.13.1 baseline (x3x2a2-era, same fixture) | 0.14.1-pre run 4 |
159
+ | --- | ---: | ---: |
160
+ | outcome | completed, verified | completed, verified, **auto-completed** (program-completion) |
161
+ | wall | 28 min 42 s | 30 min 12 s |
162
+ | planner turns / plannerSec | 1 / 294 s | **1 / 269 s** |
163
+ | planner context (turn 1) | 32.7 k chars (measured on a sibling run) | **6.2 k chars** |
164
+ | dispatches (workers) | — | 8 (7) · max concurrent 9 |
165
+ | corrections / rejections / repairs / verdict re-asks | 0 / 0 / 0 / 0 | 0 / 0 / 0 / 0 |
166
+ | deliverable | csv + slugify guarded, 63/63 | csv + slugify guarded, **63/63**, existing test files byte-identical |
167
+ Program: scout → probe (all exports, wrong-type matrix) → fanout guards over the discovered modules → verify-guards →
168
+ report → verify-final, `completion: all-actions-ok`. Wall is within noise of baseline (+90 s, dominated by worker
169
+ model time: probe 4.7 min, guards 5.1 min, verify-guards 4.7 min); the planner side is faster and 5× smaller.
170
+ Zero observation crashes across four runs of TUI/watch/runs/result/static-tui polling and a 20 s stress loop
171
+ (1,681 paints against the live writer, 0 torn, 0 throws).
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "bullswarm",
3
- "version": "0.14.0",
3
+ "version": "0.15.0",
4
4
  "description": "Route work across coding-agent CLI subscriptions — paced by live quota meters, verified by content, never trusting exit codes.",
5
5
  "type": "module",
6
6
  "bin": {
package/skill/SKILL.md CHANGED
@@ -246,7 +246,12 @@ bullswarm workflow steer <shortId> --message "guidance for the next planner chec
246
246
  `capabilities` reports available pools, supported lanes, configured models,
247
247
  meter readings, burst gates, quarantine state, retry limits, and the important
248
248
  routing rule. Automatic routing chooses the highest time-adjusted quota surplus
249
- among capable pools. For strategic model selection, first run:
249
+ among capable pools. An unmetered pool reads as exactly on pace (surplus 0), so
250
+ whenever every metered pool is burning ahead of its window (negative
251
+ surplus) the unmetered pool wins ALL work — by design: quota protection
252
+ outranks provider diversity (observed 2026-08-29: 26 of 28 dispatches on
253
+ one unmetered pool). If that concentration is unwanted, meter the pool or
254
+ exclude its models via `strategy exclude-model`. For strategic model selection, first run:
250
255
 
251
256
  ```bash
252
257
  bullswarm strategy refresh
@@ -406,6 +411,17 @@ Resume by shortId:
406
411
  bullswarm workflow draft run my-audit --resume <shortId> --json --quiet
407
412
  ```
408
413
 
414
+ ## Writing goals that converge
415
+
416
+ State outcomes, not measurements. A goal that fixes character counts, exact
417
+ event names, or cosmetic layout turns every verifier into a nit machine:
418
+ observed 2026-08-29 (run `ejk9w2`), a spec with hard numeric limits cost a
419
+ 12-minute verify/repair loop enforcing them against an intermediate state.
420
+ Say what must be true at the end (`npm test` passes, the planner receives one
421
+ contract stated once, context stays bounded) and let workers pick the numbers;
422
+ put any hard limit in ONE final acceptance check, not on every intermediate
423
+ verify.
424
+
409
425
  ## Writing prompts that the verify gate will accept
410
426
 
411
427
  The verify gate (`src/lib/verify.js`) flags outputs as `intent_only`
@@ -0,0 +1,202 @@
1
+ // Discover Claude Code logins on this machine.
2
+ //
3
+ // One account = one config dir (`~/.claude` by default, or CLAUDE_CONFIG_DIR).
4
+ // Extra homes live next to it as `~/.claude-<slug>`. Credentials:
5
+ // macOS Keychain `Claude Code-credentials` for ~/.claude
6
+ // `Claude Code-credentials-<sha256(absPath)[:8]>` for any other home
7
+ // plus `$dir/.credentials.json` on every platform.
8
+
9
+ import { createHash } from 'node:crypto';
10
+ import { existsSync, readdirSync, readFileSync, statSync } from 'node:fs';
11
+ import { homedir, platform } from 'node:os';
12
+ import { basename, join, resolve } from 'node:path';
13
+ import {
14
+ DEFAULT_KEYCHAIN_SERVICE,
15
+ extractCredentials,
16
+ isUsable,
17
+ readFromMacKeychain,
18
+ } from '../meters/claude.js';
19
+
20
+ const HOME_MARKERS = ['.credentials.json', '.claude.json', 'settings.json', 'projects'];
21
+
22
+ export function defaultClaudeHome(homeDir = homedir()) {
23
+ return resolve(join(homeDir, '.claude'));
24
+ }
25
+
26
+ export function keychainServiceForConfigDir(configDir, homeDir = homedir()) {
27
+ const resolved = resolve(configDir);
28
+ if (resolved === defaultClaudeHome(homeDir)) return DEFAULT_KEYCHAIN_SERVICE;
29
+ const hash = createHash('sha256').update(resolved).digest('hex').slice(0, 8);
30
+ return `${DEFAULT_KEYCHAIN_SERVICE}-${hash}`;
31
+ }
32
+
33
+ export function accountSlugForConfigDir(configDir, homeDir = homedir()) {
34
+ const resolved = resolve(configDir);
35
+ if (resolved === defaultClaudeHome(homeDir)) return null;
36
+ const base = basename(resolved);
37
+ if (base.startsWith('.claude-')) {
38
+ const slug = base.slice('.claude-'.length);
39
+ return slug.length > 0 ? slug : 'alt';
40
+ }
41
+ if (base.startsWith('.claude')) {
42
+ const rest = base.slice('.claude'.length).replace(/^-+/, '');
43
+ return rest.length > 0 ? rest : 'alt';
44
+ }
45
+ return base || 'alt';
46
+ }
47
+
48
+ export function poolNameForSlug(slug) {
49
+ return slug ? `claude-code:${slug}` : 'claude-code';
50
+ }
51
+
52
+ export function looksLikeClaudeHome(dir) {
53
+ try {
54
+ if (!statSync(dir).isDirectory()) return false;
55
+ } catch {
56
+ return false;
57
+ }
58
+ return HOME_MARKERS.some((name) => existsSync(join(dir, name)));
59
+ }
60
+
61
+ export function discoverClaudeConfigDirs(opts = {}) {
62
+ const homeDir = opts.homeDir ?? homedir();
63
+ const found = [];
64
+ const seen = new Set();
65
+ const add = (dir) => {
66
+ const resolved = resolve(dir);
67
+ if (seen.has(resolved)) return;
68
+ if (!looksLikeClaudeHome(resolved)) return;
69
+ seen.add(resolved);
70
+ found.push(resolved);
71
+ };
72
+ add(join(homeDir, '.claude'));
73
+ try {
74
+ for (const name of readdirSync(homeDir)) {
75
+ if (!name.startsWith('.claude-')) continue;
76
+ add(join(homeDir, name));
77
+ }
78
+ } catch { /* unreadable home */ }
79
+ const envDir = opts.envConfigDir ?? process.env.CLAUDE_CONFIG_DIR;
80
+ if (envDir && String(envDir).trim()) add(String(envDir).trim());
81
+ return found;
82
+ }
83
+
84
+ function readCredentialsFile(path) {
85
+ try {
86
+ return extractCredentials(readFileSync(path, 'utf8'));
87
+ } catch {
88
+ return null;
89
+ }
90
+ }
91
+
92
+ export function readAccountCredentials(configDir, opts = {}) {
93
+ const homeDir = opts.homeDir ?? homedir();
94
+ const nowMs = opts.nowMs ?? Date.now();
95
+ const os = opts.platform ?? platform();
96
+ const keychainRead = opts.readKeychain ?? readFromMacKeychain;
97
+ const usableOnly = opts.usableOnly !== false;
98
+
99
+ const fileCreds = readCredentialsFile(join(configDir, '.credentials.json'));
100
+ let keychainCreds = null;
101
+ if (os === 'darwin') {
102
+ keychainCreds = keychainRead(keychainServiceForConfigDir(configDir, homeDir));
103
+ }
104
+ const extraFile = resolve(configDir) === defaultClaudeHome(homeDir)
105
+ ? readCredentialsFile(join(homeDir, '.config', 'claude', 'credentials.json'))
106
+ : null;
107
+
108
+ const candidates = [];
109
+ if (keychainCreds) candidates.push({ creds: keychainCreds, source: 'keychain' });
110
+ if (fileCreds) candidates.push({ creds: fileCreds, source: 'file' });
111
+ if (extraFile) candidates.push({ creds: extraFile, source: 'file' });
112
+ for (const c of candidates) {
113
+ if (!usableOnly || isUsable(c.creds, nowMs)) return c;
114
+ }
115
+ return null;
116
+ }
117
+
118
+ export function profileCommand(configDir, bin = 'claude') {
119
+ return `CLAUDE_CONFIG_DIR=${configDir} ${bin}`;
120
+ }
121
+
122
+ export function discoverClaudeAccounts(opts = {}) {
123
+ const homeDir = opts.homeDir ?? homedir();
124
+ const dirs = discoverClaudeConfigDirs({
125
+ homeDir,
126
+ envConfigDir: opts.envConfigDir,
127
+ });
128
+ const accounts = [];
129
+ const seenTokens = new Set();
130
+ for (const configDir of dirs) {
131
+ const got = readAccountCredentials(configDir, opts);
132
+ if (!got) continue;
133
+ if (seenTokens.has(got.creds.accessToken)) continue;
134
+ seenTokens.add(got.creds.accessToken);
135
+ const slug = accountSlugForConfigDir(configDir, homeDir);
136
+ accounts.push({
137
+ configDir,
138
+ slug,
139
+ pool: poolNameForSlug(slug),
140
+ command: profileCommand(configDir, opts.bin ?? 'claude'),
141
+ creds: got.creds,
142
+ source: got.source,
143
+ });
144
+ }
145
+ accounts.sort((a, b) => {
146
+ if (a.slug === null) return -1;
147
+ if (b.slug === null) return 1;
148
+ return a.slug.localeCompare(b.slug);
149
+ });
150
+ return accounts;
151
+ }
152
+
153
+ /**
154
+ * Clone the packaged `claude-code` connector once per extra login.
155
+ * The default home keeps the historical pool name `claude-code`.
156
+ * Extra homes become `claude-code:<slug>` with CLAUDE_CONFIG_DIR in env
157
+ * so spawn bills that seat. Discovery is filesystem-driven — never a
158
+ * hardcoded profile list.
159
+ */
160
+ export function expandClaudeAccountConnectors(connectors, opts = {}) {
161
+ if (opts.disabled === true) return connectors;
162
+ if (process.env.BULLSWARM_DISABLE_CLAUDE_PROFILES === '1') return connectors;
163
+ // node:test (and MCP children it spawns) inherit NODE_TEST_CONTEXT. Do not
164
+ // scan the operator's real extra Claude homes unless the test injects
165
+ // `accounts` or `homeDir`.
166
+ if (process.env.NODE_TEST_CONTEXT && opts.accounts == null && opts.homeDir == null) {
167
+ return connectors;
168
+ }
169
+ const base = connectors['claude-code'];
170
+ if (!base) return connectors;
171
+ const accounts = opts.accounts ?? discoverClaudeAccounts({
172
+ homeDir: opts.homeDir,
173
+ envConfigDir: opts.envConfigDir,
174
+ bin: base.bin ?? 'claude',
175
+ });
176
+ const defaultDir = accounts.find((a) => a.slug == null)?.configDir
177
+ ?? defaultClaudeHome(opts.homeDir);
178
+ const bin = base.bin ?? 'claude';
179
+ base.env = { ...(base.env ?? {}), CLAUDE_CONFIG_DIR: defaultDir };
180
+ base.profile = {
181
+ slug: null,
182
+ configDir: defaultDir,
183
+ command: profileCommand(defaultDir, bin),
184
+ };
185
+ for (const account of accounts) {
186
+ if (!account.slug) continue;
187
+ const name = account.pool;
188
+ if (connectors[name]) continue;
189
+ const clone = structuredClone(base);
190
+ clone.name = name;
191
+ clone.env = { ...(base.env ?? {}), CLAUDE_CONFIG_DIR: account.configDir };
192
+ clone.configDirs = [account.configDir];
193
+ clone.flags = { ...(base.flags ?? {}), isCaller: false };
194
+ clone.profile = {
195
+ slug: account.slug,
196
+ configDir: account.configDir,
197
+ command: account.command,
198
+ };
199
+ connectors[name] = clone;
200
+ }
201
+ return connectors;
202
+ }
package/src/lib/config.js CHANGED
@@ -13,8 +13,9 @@ import { readFileSync, readdirSync, existsSync } from 'node:fs';
13
13
  import { join } from 'node:path';
14
14
  import { loadState } from './state.js';
15
15
  import { paceScore, isQuarantined } from './route.js';
16
+ import { expandClaudeAccountConnectors } from './claude-accounts.js';
16
17
 
17
- export function loadConnectors(bullswarmDir) {
18
+ export function loadConnectors(bullswarmDir, opts = {}) {
18
19
  const dir = join(bullswarmDir, 'connectors');
19
20
  if (!existsSync(dir)) return {};
20
21
  const out = {};
@@ -28,6 +29,7 @@ export function loadConnectors(bullswarmDir) {
28
29
  // never crash a run
29
30
  }
30
31
  }
32
+ expandClaudeAccountConnectors(out, opts);
31
33
  return out;
32
34
  }
33
35
 
package/src/lib/watch.js CHANGED
@@ -78,7 +78,12 @@ export function runDelegate(connector, taskFile, targetDir, opts = {}) {
78
78
  // hazard). Caller-supplied opts.env takes precedence over
79
79
  // process.env so the runtime can inject BULLSWARM_DEPTH (recursion
80
80
  // guard) and other core-owned env contracts.
81
- env: { ...process.env, ...(opts.env ?? {}), PWD: resolvedDir },
81
+ env: {
82
+ ...process.env,
83
+ ...(connector.env ?? {}),
84
+ ...(opts.env ?? {}),
85
+ PWD: resolvedDir,
86
+ },
82
87
  stdio: ['ignore', 'pipe', 'pipe'],
83
88
  });
84
89
  opts.onSpawn?.(child.pid);
@@ -19,18 +19,24 @@ export class ClaudeMeterError extends Error {
19
19
  }
20
20
  }
21
21
 
22
+ export const DEFAULT_KEYCHAIN_SERVICE = 'Claude Code-credentials';
23
+
24
+ export function isUsable(creds, now = Date.now(), skewMs = EXPIRY_SKEW_MS) {
25
+ return Boolean(creds) && creds.expiresAt - skewMs > now;
26
+ }
27
+
22
28
  export function readOAuthCredentials() {
23
29
  if (platform() === 'darwin') {
24
- return readFromMacKeychain() ?? readFromCredentialsFile();
30
+ return readFromMacKeychain(DEFAULT_KEYCHAIN_SERVICE) ?? readFromCredentialsFile();
25
31
  }
26
32
  return readFromCredentialsFile();
27
33
  }
28
34
 
29
- function readFromMacKeychain() {
35
+ export function readFromMacKeychain(service = DEFAULT_KEYCHAIN_SERVICE) {
30
36
  try {
31
37
  const blob = execFileSync(
32
38
  'security',
33
- ['find-generic-password', '-s', 'Claude Code-credentials', '-w'],
39
+ ['find-generic-password', '-s', service, '-w'],
34
40
  { stdio: ['ignore', 'pipe', 'ignore'], encoding: 'utf8' },
35
41
  );
36
42
  return extractCredentials(blob);
@@ -73,13 +79,13 @@ function normalizeWindow(raw) {
73
79
  };
74
80
  }
75
81
 
76
- export function parseClaudeUsage(body) {
82
+ export function parseClaudeUsage(body, pool = 'claude-code') {
77
83
  if (!body || typeof body !== 'object') {
78
84
  throw new ClaudeMeterError('Claude usage response missing body', 'parse');
79
85
  }
80
86
  return {
81
87
  captured_at: new Date().toISOString(),
82
- pool: 'claude-code',
88
+ pool,
83
89
  five_hour: normalizeWindow(body.five_hour) ?? { utilization: null, resets_at: null },
84
90
  seven_day: normalizeWindow(body.seven_day) ?? { utilization: null, resets_at: null },
85
91
  monthly: null,
@@ -89,12 +95,11 @@ export function parseClaudeUsage(body) {
89
95
  };
90
96
  }
91
97
 
92
- export async function fetchClaudeUsage() {
93
- const creds = readOAuthCredentials();
98
+ export async function fetchClaudeUsageWithCredentials(creds, pool = 'claude-code') {
94
99
  if (!creds) {
95
100
  throw new ClaudeMeterError('No Claude Code OAuth token. Run `claude` to log in.', 'no_token');
96
101
  }
97
- if (creds.expiresAt - EXPIRY_SKEW_MS <= Date.now()) {
102
+ if (!isUsable(creds)) {
98
103
  throw new ClaudeMeterError(
99
104
  'Claude OAuth token expired; open Claude Code to refresh it.',
100
105
  'expired',
@@ -122,5 +127,9 @@ export async function fetchClaudeUsage() {
122
127
  } catch (err) {
123
128
  throw new ClaudeMeterError(`Failed to parse usage response: ${err.message}`, 'parse');
124
129
  }
125
- return parseClaudeUsage(body);
130
+ return parseClaudeUsage(body, pool);
131
+ }
132
+
133
+ export async function fetchClaudeUsage() {
134
+ return fetchClaudeUsageWithCredentials(readOAuthCredentials(), 'claude-code');
126
135
  }
@@ -6,7 +6,8 @@ import { MeterCache, paceSnapshot, FRESH_MS, STALE_MS } from './framework.js';
6
6
  import { fetchCodexUsage, CodexMeterError } from './codex.js';
7
7
  import { fetchGrokUsage, GrokMeterError } from './grok.js';
8
8
  import { fetchCommandCodeUsage, CommandCodeMeterError } from './command-code.js';
9
- import { fetchClaudeUsage, ClaudeMeterError } from './claude.js';
9
+ import { fetchClaudeUsage, fetchClaudeUsageWithCredentials, ClaudeMeterError } from './claude.js';
10
+ import { discoverClaudeAccounts, poolNameForSlug } from '../lib/claude-accounts.js';
10
11
 
11
12
  export const METERS_DIR = () =>
12
13
  process.env.BULLSWARM_HOME?.trim() || join(homedir(), '.bullswarm');
@@ -19,7 +20,26 @@ const READERS = {
19
20
  claude: fetchClaudeUsage,
20
21
  };
21
22
 
23
+ function claudeReaderFor(pool) {
24
+ return async () => {
25
+ const accounts = discoverClaudeAccounts();
26
+ const slug = pool.startsWith('claude-code:') ? pool.slice('claude-code:'.length) : null;
27
+ const account = accounts.find((a) => poolNameForSlug(a.slug) === pool)
28
+ ?? accounts.find((a) => a.slug === slug);
29
+ if (!account) {
30
+ throw new ClaudeMeterError(
31
+ `No Claude Code OAuth token for pool ${pool}. Log in with CLAUDE_CONFIG_DIR pointing at that home.`,
32
+ 'no_token',
33
+ );
34
+ }
35
+ return fetchClaudeUsageWithCredentials(account.creds, pool);
36
+ };
37
+ }
38
+
22
39
  export function readerFor(pool) {
40
+ if (pool === 'claude-code' || pool === 'claude' || pool.startsWith('claude-code:')) {
41
+ return claudeReaderFor(pool);
42
+ }
23
43
  return READERS[pool] ?? null;
24
44
  }
25
45
 
@@ -340,7 +340,7 @@ async function wfGoal(opts) {
340
340
  try {
341
341
  doc = existsSync(workflowPath)
342
342
  ? JSON.parse(readFileSync(workflowPath, 'utf8'))
343
- : JSON.parse(readFileSync(statePath, 'utf8'))._doc;
343
+ : readJsonForUpdate(statePath, 'workflow state')._doc;
344
344
  } catch (err) {
345
345
  console.error(`✗ cannot load durable workflow for ${resumeRunId}: ${err.message}`);
346
346
  return 1;
@@ -457,6 +457,8 @@ async function wfCapabilities(opts) {
457
457
  enabled: p.enabled !== false,
458
458
  lanes: p.lanes ?? p.connector?.lanes ?? [],
459
459
  capabilities: p.capabilities ?? p.connector?.capabilities ?? [],
460
+ command: p.connector?.profile?.command ?? p.connector?.spawn?.cmd?.[0] ?? null,
461
+ configDir: p.connector?.profile?.configDir ?? p.connector?.env?.CLAUDE_CONFIG_DIR ?? null,
460
462
  model: (() => {
461
463
  const cmd = p.connector?.spawn?.cmd ?? [];
462
464
  const i = cmd.indexOf('--model');
@@ -1,7 +1,8 @@
1
1
  // Interactive workflow dashboard, inspired by Claude Code's /workflows view.
2
2
  // It deliberately uses only ANSI sequences and Node's standard streams.
3
3
 
4
- import { readFileSync, writeFileSync, existsSync } from 'node:fs';
4
+ import { readFileSync, existsSync } from 'node:fs';
5
+ import { readJsonSafe, readJsonForUpdate, writeJsonAtomic } from './fsjson.js';
5
6
  import { fanoutSucceededCount } from './runner.js';
6
7
  import { join } from 'node:path';
7
8
  import { listRuns, resolveRunId } from './short-id.js';
@@ -44,7 +45,7 @@ export function requestCancel(bullswarmDir, token) {
44
45
  if (!resolved) throw new Error(`no run found for "${token}"`);
45
46
  const statePath = join(resolved.runDir, 'state.json');
46
47
  if (!existsSync(statePath)) throw new Error(`run "${token}" has no state.json`);
47
- const state = JSON.parse(readFileSync(statePath, 'utf8'));
48
+ const state = readJsonForUpdate(statePath, 'workflow state');
48
49
  if (state.finishedAt || isTerminalWorkflowStatus(state.status)) {
49
50
  return { ...resolved, state, alreadyFinished: true };
50
51
  }
@@ -53,7 +54,7 @@ export function requestCancel(bullswarmDir, token) {
53
54
  state.status = 'cancelling';
54
55
  state.cancellingAt = state.cancelRequestedAt;
55
56
  appendEvent(resolved.runDir, state, 'run.cancellation_requested', { requestedAt: state.cancelRequestedAt });
56
- writeFileSync(statePath, `${JSON.stringify(state, null, 2)}\n`);
57
+ writeJsonAtomic(statePath, state);
57
58
  return { ...resolved, state, alreadyFinished: false };
58
59
  }
59
60
 
@@ -228,6 +229,12 @@ function autonomousControlPlane(state) {
228
229
  }
229
230
 
230
231
  function effectiveActionStatus(action, state) {
232
+ // A re-running action (repair round, re-verify, schema retry) must read as
233
+ // running even when a previous round recorded ok:false — a failed mark on
234
+ // work that is still being retried misreports the run (user report 2026-08-29).
235
+ const active = Object.values(state.activeAgents ?? {}).some((agent) =>
236
+ agent.stepId === action.id || String(agent.stepId ?? '').startsWith(`${action.id}[`));
237
+ if (active || action.status === 'running') return 'running';
231
238
  const output = state.outputs?.[action.id];
232
239
  if (output?.ok === false) return 'failed_terminal';
233
240
  if (output?.ok === true && action.status === 'succeeded') return 'succeeded';
@@ -282,9 +289,10 @@ export function workflowPanelModel(row, {
282
289
  const failed = effectiveStatuses.filter((status) => String(status).startsWith('failed')).length;
283
290
  const current = name === activePhaseName && !state.finishedAt;
284
291
  const active = effectiveStatuses.some((status) => status === 'running');
285
- const status = failed ? 'failed'
286
- : actions.length && completed === actions.length ? 'completed'
287
- : active ? 'active' : current ? 'waiting' : 'pending';
292
+ const status = active ? 'active'
293
+ : failed ? 'failed'
294
+ : actions.length && completed === actions.length ? 'completed'
295
+ : current ? 'waiting' : 'pending';
288
296
  return { name, label: phaseLabel(name, orchestrator), status, actions, completed, total: actions.length };
289
297
  });
290
298
  const selectedPhase = phases[selectedPhaseIndex];
@@ -806,8 +814,8 @@ function detailRow(bullswarmDir, token) {
806
814
  if (!resolved) throw new Error(`no run found for "${token}"`);
807
815
  const statePath = join(resolved.runDir, 'state.json');
808
816
  const reportPath = join(resolved.runDir, 'report.json');
809
- const state = existsSync(statePath) ? JSON.parse(readFileSync(statePath, 'utf8')) : null;
810
- const report = existsSync(reportPath) ? JSON.parse(readFileSync(reportPath, 'utf8')) : null;
817
+ const state = readJsonSafe(statePath);
818
+ const report = readJsonSafe(reportPath);
811
819
  return { ...resolved, state, report, events: readEvents(resolved.runDir), status: state?.status };
812
820
  }
813
821
 
@@ -823,6 +831,7 @@ export async function runDashboard(bullswarmDir, {
823
831
  let selected = 0;
824
832
  let detail = Boolean(token);
825
833
  let message = null;
834
+ let lastGoodRow = null;
826
835
  let rows = dashboardRows(bullswarmDir);
827
836
  let selectedRunId = token ? detailRow(bullswarmDir, token).runId : (rows[selected]?.runId ?? null);
828
837
  const ui = {
@@ -838,10 +847,14 @@ export async function runDashboard(bullswarmDir, {
838
847
  orchestratorVerbose: false,
839
848
  spinnerFrame: 0,
840
849
  };
841
- const paint = () => {
850
+ const paintUnsafe = () => {
842
851
  if (selected >= rows.length) selected = Math.max(0, rows.length - 1);
843
852
  if (detail && selectedRunId) {
844
- const row = detailRow(bullswarmDir, selectedRunId);
853
+ // A torn read while the runner writes state.json yields state:null for
854
+ // one frame — keep painting the last good snapshot of the same run.
855
+ const fresh = detailRow(bullswarmDir, selectedRunId);
856
+ const row = (fresh.state || lastGoodRow?.runId !== fresh.runId) ? fresh : lastGoodRow;
857
+ if (row === fresh) lastGoodRow = fresh;
845
858
  const model = workflowPanelModel(row, {
846
859
  phaseIndex: ui.followActivePhase ? null : ui.phaseIndex,
847
860
  agentIndex: ui.followActiveAgent ? null : ui.agentIndex,
@@ -858,8 +871,19 @@ export async function runDashboard(bullswarmDir, {
858
871
  }
859
872
  output.write(renderDashboard({ rows, selected, message }));
860
873
  };
874
+ // A render error must never kill the TUI or strand the terminal in
875
+ // alt-screen raw mode (crash observed 2026-08-29 at detailRow via the
876
+ // repaint timer). Show the error in the message line and keep running.
877
+ const paint = () => {
878
+ try { paintUnsafe(); } catch (err) {
879
+ message = `display error: ${err.message}`;
880
+ try { output.write(renderDashboard({ rows, selected, message })); } catch { /* keep the loop alive */ }
881
+ }
882
+ };
861
883
  const refresh = () => {
862
- rows = dashboardRows(bullswarmDir);
884
+ try {
885
+ rows = dashboardRows(bullswarmDir);
886
+ } catch (err) { message = `display error: ${err.message}`; }
863
887
  if (!selectedRunId) selectedRunId = rows[selected]?.runId ?? null;
864
888
  paint();
865
889
  };
@@ -935,7 +959,7 @@ export async function runDashboard(bullswarmDir, {
935
959
  refresh();
936
960
  } catch (err) { message = err.message; ui.confirmCancel = false; paint(); }
937
961
  };
938
- const onData = (buf) => {
962
+ const onDataUnsafe = (buf) => {
939
963
  const key = String(buf);
940
964
  if (ui.confirmCancel) {
941
965
  if (key === 'y' || key === 'Y') return requestSelectedCancel();
@@ -1019,6 +1043,14 @@ export async function runDashboard(bullswarmDir, {
1019
1043
  return paint();
1020
1044
  }
1021
1045
  };
1046
+ const onData = (buf) => {
1047
+ // A key-handler error (e.g. a drill-in racing the writer) must never
1048
+ // kill the TUI; finish() still restores the terminal on q/Ctrl-C.
1049
+ try { onDataUnsafe(buf); } catch (err) {
1050
+ message = `display error: ${err.message}`;
1051
+ paint();
1052
+ }
1053
+ };
1022
1054
  const onResize = () => paint();
1023
1055
  input.on('data', onData);
1024
1056
  output.on?.('resize', onResize);
@@ -1035,8 +1067,8 @@ export function dashboardJson(bullswarmDir, { all = false, token = null, cancel
1035
1067
  if (!resolved) throw new Error(`no run found for "${token}"`);
1036
1068
  const statePath = join(resolved.runDir, 'state.json');
1037
1069
  const reportPath = join(resolved.runDir, 'report.json');
1038
- const state = existsSync(statePath) ? JSON.parse(readFileSync(statePath, 'utf8')) : null;
1039
- const report = existsSync(reportPath) ? JSON.parse(readFileSync(reportPath, 'utf8')) : null;
1070
+ const state = readJsonSafe(statePath);
1071
+ const report = readJsonSafe(reportPath);
1040
1072
  const events = readEvents(resolved.runDir);
1041
1073
  return { action: 'show', ...resolved, state, report, events };
1042
1074
  }
@@ -1048,7 +1080,7 @@ export function actionJson(bullswarmDir, token, actionId) {
1048
1080
  const resolved = resolveRunId(bullswarmDir, token);
1049
1081
  if (!resolved) throw new Error(`no run found for "${token}"`);
1050
1082
  const statePath = join(resolved.runDir, 'state.json');
1051
- const state = existsSync(statePath) ? JSON.parse(readFileSync(statePath, 'utf8')) : null;
1083
+ const state = readJsonSafe(statePath);
1052
1084
  const action = state?.actionLedger?.find((entry) => entry.id === actionId);
1053
1085
  if (!action) throw new Error(`run "${token}" has no action "${actionId}"`);
1054
1086
  const attempts = (action.attempts ?? []).map((index) => state.attempts?.[index]).filter(Boolean);
@@ -1062,7 +1094,7 @@ export function decideApproval(bullswarmDir, token, decision) {
1062
1094
  const resolved = resolveRunId(bullswarmDir, token);
1063
1095
  if (!resolved) throw new Error(`no run found for "${token}"`);
1064
1096
  const statePath = join(resolved.runDir, 'state.json');
1065
- const state = JSON.parse(readFileSync(statePath, 'utf8'));
1097
+ const state = readJsonForUpdate(statePath, 'workflow state');
1066
1098
  if (state.status !== 'waiting_for_approval' || !state.approval) {
1067
1099
  throw new Error(`run "${token}" is not waiting for approval`);
1068
1100
  }
@@ -1085,6 +1117,6 @@ export function decideApproval(bullswarmDir, token, decision) {
1085
1117
  gateId: state.approval.gateId,
1086
1118
  decidedAt: at,
1087
1119
  });
1088
- writeFileSync(statePath, `${JSON.stringify(state, null, 2)}\n`);
1120
+ writeJsonAtomic(statePath, state);
1089
1121
  return { ...resolved, decision, state };
1090
1122
  }
@@ -72,10 +72,20 @@ export function normalizeDecisionProposal(proposal) {
72
72
  };
73
73
  }
74
74
  if (action?.type !== 'verify') return action;
75
- const singleDependency = Array.isArray(action.dependsOn) && action.dependsOn.length === 1
76
- ? action.dependsOn[0] : null;
75
+ const dependsOn = Array.isArray(action.dependsOn) ? action.dependsOn : [];
76
+ const singleDependency = dependsOn.length === 1 ? dependsOn[0] : null;
77
77
  if (action.review == null) {
78
- return singleDependency ? { ...action, review: `outputs.${singleDependency}.outFile` } : action;
78
+ // A verify reviews an artifact; when the planner names none, take the
79
+ // obvious one instead of rejecting a whole program for one field
80
+ // (observed 2026-08-29: a 9-action proposal bounced for a
81
+ // zero-dependency audit verify, costing a 5-minute correction turn).
82
+ // One dependency -> its artifact. Several -> the most downstream
83
+ // (last listed). None -> a repository audit with no artifact.
84
+ if (singleDependency) return { ...action, review: `outputs.${singleDependency}.outFile` };
85
+ if (dependsOn.length > 1) {
86
+ return { ...action, review: `outputs.${dependsOn.at(-1)}.outFile`, reviewDefaultedFrom: 'last-dependency' };
87
+ }
88
+ return { ...action, reviewScope: 'repository' };
79
89
  }
80
90
  if (looksLikeReviewPath(action.review)) return { ...action, review: action.review.trim() };
81
91
  if (typeof action.review === 'string' && singleDependency) {
@@ -241,8 +251,10 @@ export function validateDecisionProposal(proposal, {
241
251
  }
242
252
  }
243
253
  if (action.type === 'verify') {
244
- if (typeof action.review !== 'string') {
245
- issues.push(`${at}.review is required: a verify with one dependsOn reviews that artifact automatically; with several, set review to "outputs.<actionId>.outFile" and put reviewer instructions in prompt`);
254
+ if (action.review == null && action.reviewScope === 'repository') {
255
+ // Normalized zero-dependency audit: no artifact to review.
256
+ } else if (typeof action.review !== 'string') {
257
+ issues.push(`${at}.review must be a dotted artifact path like "outputs.<actionId>.outFile" when present; omit it to review the single/last dependency's artifact, or to audit the repository when the verify has no dependsOn`);
246
258
  } else if (!looksLikeReviewPath(action.review)) {
247
259
  issues.push(`${at}.review must be a dotted artifact path like "outputs.<actionId>.outFile" (the artifact to review), not instructions or a filesystem path; put reviewer instructions in ${at}.prompt`);
248
260
  } else {
@@ -58,12 +58,8 @@ export function draftPaths(bullswarmDir, name) {
58
58
  return { dir: d, doc: join(d, 'workflow.json'), meta: join(d, 'meta.json') };
59
59
  }
60
60
 
61
- function atomicWrite(p, content) {
62
- mkdirSync(dirname(p), { recursive: true });
63
- const tmp = `${p}.tmp-${randomBytes(3).toString('hex')}`;
64
- writeFileSync(tmp, content);
65
- renameSync(tmp, p);
66
- }
61
+ // Hoisted to fsjson.js (shared with runtime/runner state persistence).
62
+ import { atomicWriteFileSync as atomicWrite } from './fsjson.js';
67
63
 
68
64
  export function draftExists(bullswarmDir, name) {
69
65
  return existsSync(draftPaths(bullswarmDir, name).doc);
@@ -0,0 +1,41 @@
1
+ // Atomic JSON persistence + torn-read-tolerant JSON reads.
2
+ //
3
+ // Doctrine:
4
+ // F1. Every state/report/workflow artifact is written via temp+rename so a
5
+ // concurrent reader can NEVER observe a half-written file (earned:
6
+ // `bullswarm workflow tui` crashed with "Unterminated string in JSON at
7
+ // position 138968" parsing state.json mid-write, 2026-08-29).
8
+ // F2. Display readers tolerate a missing or torn file (returns fallback):
9
+ // observation must never crash on the writer's timing.
10
+ // F3. Mutating read-modify-write readers must NOT silently no-op: they
11
+ // retry once (rename-atomic writes make the second read succeed) and
12
+ // then throw a clear error instead of dropping the user's command.
13
+
14
+ import { readFileSync, writeFileSync, renameSync, mkdirSync, existsSync } from 'node:fs';
15
+ import { dirname } from 'node:path';
16
+ import { randomBytes } from 'node:crypto';
17
+
18
+ export function atomicWriteFileSync(path, content) {
19
+ mkdirSync(dirname(path), { recursive: true });
20
+ const tmp = `${path}.tmp-${randomBytes(3).toString('hex')}`;
21
+ writeFileSync(tmp, content);
22
+ renameSync(tmp, path);
23
+ }
24
+
25
+ export function writeJsonAtomic(path, value) {
26
+ atomicWriteFileSync(path, `${JSON.stringify(value, null, 2)}\n`);
27
+ }
28
+
29
+ /** Display-path read: missing or torn file -> fallback, never a throw. */
30
+ export function readJsonSafe(path, fallback = null) {
31
+ if (!existsSync(path)) return fallback;
32
+ try { return JSON.parse(readFileSync(path, 'utf8')); } catch { return fallback; }
33
+ }
34
+
35
+ /** Mutating-path read: retry once, then throw a clear actionable error. */
36
+ export function readJsonForUpdate(path, what = 'file') {
37
+ try { return JSON.parse(readFileSync(path, 'utf8')); } catch { /* torn or corrupt; retry once */ }
38
+ try { return JSON.parse(readFileSync(path, 'utf8')); } catch (err) {
39
+ throw new Error(`${what} at ${path} is unreadable (${err.message}); it may be mid-write — retry the command`);
40
+ }
41
+ }
@@ -11,11 +11,11 @@ export const PLANNER_RULES_SECTION = [
11
11
  '1. Compile the whole program in one decision: the runtime executes all proposed actions and consults you only at a finished-or-blocked boundary, so deferring decidable work costs another round trip.',
12
12
  '2. Make every worker prompt self-contained: include the exact goal, absolute cwd, owned files and a no-other-files boundary, expected artifact, acceptance command, and report format, because workers see only their own prompt.',
13
13
  '3. Use short kebab-case, forward-only phases and dependsOn only for real data or same-file ordering; recovery uses a new phase and never repeats an identical failed plan.',
14
- '4. For known N items, create N run plus N verify actions, each verify depending only on its own run, then one suite verify depending on all; this exposes safe parallelism while preserving per-item evidence.',
14
+ '4. For known N items, create N run plus N verify actions, each verify depending only on its own run, then one suite verify depending on all; a verify judges the artifact in review (default: its single or last dependency; a verify with no dependsOn audits the repository directly).',
15
15
  '5. For unknown items, create discovery ending with RETURN ONLY a JSON object containing an items array, then data-driven fan-out via itemsFrom outputs.<id>.outFile or outputs.<id>.data.<field>; the runtime extracts the list and retries once read-only if needed.',
16
16
  '6. Put outputSchema on workers whose reports are consumed or whose claims the runtime must check, so structured data is durable and can drive later fan-out.',
17
17
  '7. Put verify.repair on every verify, and scope each verify to what can be true at its point in the graph: work scheduled later is not a defect, and cosmetic mismatches with the goal text are concerns, never ok:false. An ok:false verdict is repaired and re-checked inside the program; ok:true is accepted and concerns are informational, not extra work.',
18
- '8. Add completion with all-actions-ok whenever a clean program finishes the goal; when the goal\'s acceptance checks pass, return complete rather than adding polish or alignment actions. Return complete only on durable verified evidence, never proceed, never ask the user, and stop only for a concrete unresolved blocker with a qualified outcome.',
18
+ '8. Add completion with all-actions-ok whenever a clean program finishes the goal; when the goal\'s acceptance checks pass, return complete rather than adding polish or alignment actions. Completion evidence requires the program\'s LAST worker to be covered by a successful verify — never leave a report or other run as the final unverified action. Return complete only on durable verified evidence, never proceed, never ask the user, and stop only for a concrete unresolved blocker with a qualified outcome.',
19
19
  '9. Treat agent-count, workflow-duration, and expansion-round budgets as advisory planning targets, never hard stop conditions; the dispatch budget counts this planner call plus workers, verifiers, retries, and escalations. Converge as targets approach, avoid optional work, and exceed a target only for one essential bounded action or required verification.',
20
20
  '10. This is a control-plane thread: do not invoke Bullswarm, use tools, modify files, or propose pool, addDir, taskFile, shell authority, or unbounded work; route and process authority belong to the runtime.',
21
21
  'Shared working tree: concurrent workers editing DISJOINT files is the normal parallel mode; order shared files (indexes, barrels) after their feeders with dependsOn, and run the full suite once in a final verify — never while other workers still edit. Avoid redundant expensive verification: later verifiers reuse durable clean full-suite evidence unless it is stale or the code changed again. operatorSteering in the context is explicit operator guidance for this checkpoint: apply it within the original intent; it cannot weaken verification or expand authority.',
@@ -26,7 +26,7 @@ export const PLANNER_EXAMPLES_SECTION = [
26
26
  '[{"type":"run","phase":"implement","prompt":"..."},{"type":"run","phase":"report","prompt":"...","outputSchema":{"type":"object"}},{"type":"fanout","phase":"fix","items":["alpha"],"stepTemplate":{"prompt":"Handle {{item}}."}},{"type":"fanout","phase":"fix","itemsFrom":"outputs.discover.outFile","stepTemplate":{"prompt":"Handle {{item}}."}},{"type":"verify","phase":"verify","prompt":"Check the artifact.","repair":{"prompt":"Fix rejected concerns.","maxRounds":1}}]',
27
27
  'Complete program:',
28
28
  '[{"id":"discover","type":"run","phase":"discover","prompt":"In /abs/repo discover items and end with RETURN ONLY a JSON object containing an items array of item names.","outputSchema":{"type":"object","properties":{"items":{"type":"array","items":{"type":"string"}}},"required":["items"]}},{"id":"fix","type":"fanout","phase":"fix","itemsFrom":"outputs.discover.data.items","stepTemplate":{"prompt":"In /abs/repo edit only the files for {{item}} and run its focused acceptance command."},"dependsOn":["discover"]},{"id":"verify-items","type":"verify","phase":"verify-items","prompt":"Independently verify every item artifact.","dependsOn":["fix"],"repair":{"prompt":"Fix each rejected item in /abs/repo and re-run its focused command.","maxRounds":2}},{"id":"verify-suite","type":"verify","phase":"verify-suite","prompt":"Run the full acceptance command in /abs/repo.","dependsOn":["verify-items"],"repair":{"prompt":"Fix the suite failure in /abs/repo and rerun the suite.","maxRounds":1}}],"completion":{"when":"all-actions-ok","reason":"The item checks and final suite verification prove the goal."}]',
29
- 'Rules the validator enforces: action type is run, fanout, or verify; fanout has stepTemplate and either items or itemsFrom; verify.review is a string when explicit review is needed; dependsOn names existing or proposed actions; runtime-owned fields are rejected.',
29
+ 'Rules the validator enforces: action type is run, fanout, or verify; fanout has stepTemplate and either items or itemsFrom; verify.review, when given, is outputs.<id>.outFile; dependsOn names existing or proposed actions; runtime-owned fields are rejected.',
30
30
  ].join('\n');
31
31
 
32
32
  export const AUTONOMOUS_ORCHESTRATOR_PROMPT = [
@@ -5,7 +5,9 @@
5
5
  // settings.stopOnPhaseFailure: abort after a phase with any failure
6
6
  // Resume: steps whose recorded verdict is ok:true are skipped (R2).
7
7
 
8
- import { readFileSync, writeFileSync, existsSync, mkdirSync } from 'node:fs';
8
+ import { readFileSync, existsSync, mkdirSync } from 'node:fs';
9
+ import { writeJsonAtomic } from './fsjson.js';
10
+ import { readSteering } from './steering.js';
9
11
  import { join } from 'node:path';
10
12
  import { randomBytes } from 'node:crypto';
11
13
  import { validateWorkflow } from './validate.js';
@@ -91,7 +93,7 @@ export async function runWorkflow(opts) {
91
93
  // A generated goal workflow must be restartable without the initiating
92
94
  // process or an external draft file. The exact executable definition is
93
95
  // therefore a first-class run artifact.
94
- writeFileSync(join(runDir, 'workflow.json'), `${JSON.stringify(doc, null, 2)}\n`);
96
+ writeJsonAtomic(join(runDir, 'workflow.json'), doc);
95
97
  }
96
98
 
97
99
  let state;
@@ -142,7 +144,7 @@ export async function runWorkflow(opts) {
142
144
  // Commit cleared terminal/control markers before WorkflowRuntime begins
143
145
  // merging dashboard-side state. Otherwise the first resume event can
144
146
  // re-import the stale cancelRequested marker from the interrupted run.
145
- writeFileSync(join(runDir, 'state.json'), `${JSON.stringify(state, null, 2)}\n`);
147
+ writeJsonAtomic(join(runDir, 'state.json'), state);
146
148
  } else {
147
149
  if (resuming) {
148
150
  throw new Error(`cannot resume: no state.json for run ${runId}`);
@@ -221,7 +223,7 @@ export async function runWorkflow(opts) {
221
223
  state.orchestration = doc.orchestration
222
224
  ? { ...(state.orchestration ?? {}), ...doc.orchestration }
223
225
  : state.orchestration;
224
- writeFileSync(join(runDir, 'workflow.json'), `${JSON.stringify(doc, null, 2)}\n`);
226
+ writeJsonAtomic(join(runDir, 'workflow.json'), doc);
225
227
  }
226
228
  if (opts.inputs && Object.keys(opts.inputs).length) {
227
229
  state.inputs = { ...state.inputs, ...opts.inputs };
@@ -447,6 +449,20 @@ export async function runWorkflow(opts) {
447
449
  state.recovery = { resumable: true, signal: interruptionSignal, interruptedAt: finishedAt };
448
450
  }
449
451
  if (abortReason && !interrupted) state.abortReason = abortReason;
452
+ // Steering queued after the last planner gate must end truthfully: on a
453
+ // real terminal transition (not interrupted — a resume still has gates
454
+ // ahead — and not waiting for approval), mark undelivered entries expired.
455
+ if (!waitingForApproval && !interrupted) {
456
+ const undelivered = readSteering(runDir).filter((entry) =>
457
+ !(state.steering ?? []).some((known) => known.id === entry.id));
458
+ if (undelivered.length) {
459
+ state.steering = [
460
+ ...(state.steering ?? []),
461
+ ...undelivered.map((entry) => ({ ...entry, status: 'expired_undelivered', expiredAt: finishedAt })),
462
+ ];
463
+ runtime.emit('steering.expired', { steeringIds: undelivered.map((entry) => entry.id) });
464
+ }
465
+ }
450
466
  runtime.persist();
451
467
 
452
468
  const preliminaryReport = buildReport(state, doc, runDir);
@@ -456,7 +472,7 @@ export async function runWorkflow(opts) {
456
472
  : state.status === 'blocked' ? 'run.blocked' : 'run.completed';
457
473
  runtime.emit(terminalEvent, { runId, status: state.status, report: preliminaryReport.summary, outcome: state.outcome ?? null });
458
474
  const report = buildReport(state, doc, runDir);
459
- writeFileSync(join(runDir, 'report.json'), `${JSON.stringify(report, null, 2)}\n`);
475
+ writeJsonAtomic(join(runDir, 'report.json'), report);
460
476
  opts.onEvent?.({ type: 'workflow.completed', runId, status: state.status, report: report.summary });
461
477
 
462
478
  return { runId, runDir, state, report };
@@ -845,6 +861,41 @@ async function runDecisionLoop({ runtime, gate, phase, state, retryAttempts }) {
845
861
 
846
862
  // Resume accepted expansion work before asking the planner for a new
847
863
  // semantic decision. Successful actions are skipped by durable output.
864
+ // An action the interruption cancelled mid-flight (status `cancelled`,
865
+ // output `ok:false` "workflow cancellation requested") is unfinished work,
866
+ // not a failure: clear its output so it re-runs, and clear the outputs of
867
+ // the dependents that were blocked only because of it — otherwise resume
868
+ // asks the planner to re-plan around a phantom failure (observed on run
869
+ // djnjka, 2026-08-29: 1 cancelled action → 4 "blocked" → spurious turn).
870
+ const ledgerById = new Map((state.actionLedger ?? []).map((entry) => [entry.id, entry]));
871
+ const reopened = new Set();
872
+ for (const entry of state.plan?.actions ?? []) {
873
+ if (entry.source !== 'planner' || !entry.definition) continue;
874
+ if (ledgerById.get(entry.id)?.status === 'cancelled') reopened.add(entry.id);
875
+ }
876
+ let grew = reopened.size > 0;
877
+ while (grew) {
878
+ grew = false;
879
+ for (const entry of state.plan?.actions ?? []) {
880
+ if (entry.source !== 'planner' || !entry.definition || reopened.has(entry.id)) continue;
881
+ const blocked = state.outputs[entry.id]?.dependencyBlocked === true;
882
+ if (blocked && (entry.definition.dependsOn ?? []).some((id) => reopened.has(id))) {
883
+ reopened.add(entry.id);
884
+ grew = true;
885
+ }
886
+ }
887
+ }
888
+ for (const id of reopened) {
889
+ delete state.outputs[id];
890
+ const ledger = ledgerById.get(id);
891
+ if (ledger) {
892
+ ledger.status = 'pending';
893
+ delete ledger.why;
894
+ delete ledger.finishedAt;
895
+ }
896
+ runtime.emit('action.reopened', { actionId: id, reason: 'interrupted before completion' });
897
+ }
898
+ if (reopened.size) runtime.persist();
848
899
  const unfinishedAccepted = (state.plan?.actions ?? [])
849
900
  .filter((entry) => entry.source === 'planner' && entry.definition &&
850
901
  state.outputs[entry.id] == null)
@@ -1077,7 +1128,13 @@ async function runDecisionLoop({ runtime, gate, phase, state, retryAttempts }) {
1077
1128
  .map((entry) => entry.id);
1078
1129
  const dynamicActions = (state.actionLedger ?? []).filter((action) => action.parentId === gate.id);
1079
1130
  const gaps = failing.length ? [] : completionEvidenceGaps(dynamicActions, state.orchestration?.completionPolicy, state.outputs);
1080
- if (!failing.length && !gaps.length) {
1131
+ // Pending operator steering blocks self-completion: the documented
1132
+ // contract is delivery at the next planner gate, so a clean program
1133
+ // returns to the planner (which delivers the steer) instead of
1134
+ // silently discarding it (defect observed on wf-mtds7tzx, 2026-08-29).
1135
+ const pendingSteering = (failing.length || gaps.length) ? [] : readSteering(runtime.runDir).filter((entry) =>
1136
+ !(state.steering ?? []).some((known) => known.id === entry.id));
1137
+ if (!failing.length && !gaps.length && !pendingSteering.length) {
1081
1138
  const verifyIds = programActions.filter((entry) => entry.kind === 'verify').map((entry) => entry.id);
1082
1139
  const reason = proposal.completion.reason?.trim()
1083
1140
  || `Program completed: all ${programActions.length} actions finished ok, verified by ${verifyIds.join(', ')}.`;
@@ -1114,9 +1171,18 @@ async function runDecisionLoop({ runtime, gate, phase, state, retryAttempts }) {
1114
1171
  runtime.persist();
1115
1172
  return { ok: true, why: reason, complete: true, decision: { decision: 'complete', reason, actions: [], completion: proposal.completion } };
1116
1173
  }
1117
- runtime.emit('decision.completion_predicate_unmet', {
1118
- gateId: gate.id, programSequence: decision.sequence, failing, gaps,
1119
- });
1174
+ if (pendingSteering.length) {
1175
+ runtime.emit('decision.completion_deferred', {
1176
+ gateId: gate.id,
1177
+ programSequence: decision.sequence,
1178
+ reason: 'operator steering pending',
1179
+ steeringIds: pendingSteering.map((entry) => entry.id),
1180
+ });
1181
+ } else {
1182
+ runtime.emit('decision.completion_predicate_unmet', {
1183
+ gateId: gate.id, programSequence: decision.sequence, failing, gaps,
1184
+ });
1185
+ }
1120
1186
  }
1121
1187
  // Loop intentionally returns to observation and invokes the planner again.
1122
1188
  }
@@ -14,6 +14,7 @@
14
14
  // resolver in short-id.js maps both to the run directory.
15
15
 
16
16
  import { existsSync, rmSync, readFileSync } from 'node:fs';
17
+ import { readJsonSafe } from './fsjson.js';
17
18
  import { join } from 'node:path';
18
19
  import { listRuns, resolveRunId, isOngoing } from './short-id.js';
19
20
  import { BULLSWARM_DIR } from './cli.js';
@@ -161,8 +162,8 @@ function runsShow(idToken, opts) {
161
162
  const { runId, runDir } = resolved;
162
163
  const statePath = join(runDir, 'state.json');
163
164
  const reportPath = join(runDir, 'report.json');
164
- const state = existsSync(statePath) ? JSON.parse(readFileSync(statePath, 'utf8')) : null;
165
- const report = existsSync(reportPath) ? JSON.parse(readFileSync(reportPath, 'utf8')) : null;
165
+ const state = readJsonSafe(statePath);
166
+ const report = readJsonSafe(reportPath);
166
167
  const ongoing = isOngoing(runDir, state);
167
168
 
168
169
  if (opts.json) {
@@ -191,8 +192,8 @@ function runsResult(idToken, opts) {
191
192
  const { runId, runDir } = resolved;
192
193
  const statePath = join(runDir, 'state.json');
193
194
  const reportPath = join(runDir, 'report.json');
194
- const state = existsSync(statePath) ? JSON.parse(readFileSync(statePath, 'utf8')) : null;
195
- const report = existsSync(reportPath) ? JSON.parse(readFileSync(reportPath, 'utf8')) : null;
195
+ const state = readJsonSafe(statePath);
196
+ const report = readJsonSafe(reportPath);
196
197
  const ongoing = isOngoing(runDir, state);
197
198
  const result = buildWorkflowResult({
198
199
  state, report, runId, shortId: resolved.shortId, ongoing,
@@ -246,8 +247,15 @@ function runsDelete(idToken, opts, rest) {
246
247
  // Refuse to delete an ongoing run without --force. Half-finished
247
248
  // runs are usually a debugging target, not garbage.
248
249
  const statePath = join(runDir, 'state.json');
249
- const state = existsSync(statePath) ? JSON.parse(readFileSync(statePath, 'utf8')) : null;
250
- const ongoing = isOngoing(runDir, state);
250
+ // Delete guard: an unreadable (mid-write) state means the run may be live —
251
+ // treat it as ongoing rather than deleting a live run; --force still wins.
252
+ let state = null;
253
+ let stateUnreadable = false;
254
+ if (existsSync(statePath)) {
255
+ state = readJsonSafe(statePath, undefined);
256
+ if (state === undefined) { state = null; stateUnreadable = true; }
257
+ }
258
+ const ongoing = stateUnreadable ? true : isOngoing(runDir, state);
251
259
  if (ongoing && !opts.force) {
252
260
  return err(
253
261
  `refusing to delete ongoing run "${runId}" (shortId ${shortId ?? '?'}); ` +
@@ -22,6 +22,7 @@
22
22
  // text always lives in the per-step outFile.
23
23
 
24
24
  import { mkdirSync, writeFileSync, readFileSync, existsSync } from 'node:fs';
25
+ import { writeJsonAtomic } from './fsjson.js';
25
26
  import { join } from 'node:path';
26
27
  import { createHash, randomUUID } from 'node:crypto';
27
28
  import { pickPool, isQuarantined } from '../lib/route.js';
@@ -92,6 +93,24 @@ export function plannerBudgetContext(budget = {}) {
92
93
  };
93
94
  }
94
95
 
96
+ /** Skeleton of a rejected planner proposal: shapes and ids, prompts elided. */
97
+ export function compactRejectedProposal(proposal) {
98
+ if (!proposal || typeof proposal !== 'object') return proposal ?? null;
99
+ const elide = (text) => (typeof text === 'string' && text.length > 160 ? `${text.slice(0, 160)}…[${text.length} chars]` : text);
100
+ return {
101
+ ...proposal,
102
+ reason: elide(proposal.reason),
103
+ actions: Array.isArray(proposal.actions) ? proposal.actions.map((action) => {
104
+ if (!action || typeof action !== 'object') return action;
105
+ const out = { ...action };
106
+ if ('prompt' in out) out.prompt = elide(out.prompt);
107
+ if (out.repair && typeof out.repair === 'object') out.repair = { ...out.repair, prompt: elide(out.repair.prompt) };
108
+ if (out.stepTemplate && typeof out.stepTemplate === 'object') out.stepTemplate = { ...out.stepTemplate, prompt: elide(out.stepTemplate.prompt) };
109
+ return out;
110
+ }) : proposal.actions,
111
+ };
112
+ }
113
+
95
114
  export class WorkflowRuntime {
96
115
  /**
97
116
  * @param {object} opts
@@ -184,7 +203,8 @@ export class WorkflowRuntime {
184
203
  if (disk.cancellingAt) this.state.cancellingAt = disk.cancellingAt;
185
204
  }
186
205
  } catch { /* a partial state file should not break the workflow */ }
187
- writeFileSync(join(this.runDir, 'state.json'), `${JSON.stringify(this.state, null, 2)}\n`);
206
+ // Atomic: a 1s-interval TUI reads this file while we write it (F1).
207
+ writeJsonAtomic(join(this.runDir, 'state.json'), this.state);
188
208
  }
189
209
 
190
210
  refreshCancellation() {
@@ -1036,13 +1056,16 @@ export class WorkflowRuntime {
1036
1056
  */
1037
1057
  async runVerify(step, scope, opts = {}) {
1038
1058
  this.enforceRequiredInputs(step.id);
1039
- if (!step.review) {
1059
+ if (!step.review && step.reviewScope !== 'repository') {
1040
1060
  throw new Error(
1041
1061
  `verify step "${step.id}" needs a "review" path ` +
1042
1062
  `(e.g. review: "outputs.<priorStep>.outFile")`,
1043
1063
  );
1044
1064
  }
1045
1065
  const reviewedText = (() => {
1066
+ if (!step.review) {
1067
+ return '(no artifact: this verify has no upstream action — audit the repository state directly with fresh commands)';
1068
+ }
1046
1069
  try {
1047
1070
  // `review` is a dotted path into the scope, NOT a template
1048
1071
  // (matches the design of `fanout.itemsFrom`). Resolve it the
@@ -1297,8 +1320,14 @@ export class WorkflowRuntime {
1297
1320
  maxAttempts: opts.correction.maxAttempts,
1298
1321
  why: opts.correction.why,
1299
1322
  issues: opts.correction.issues ?? [],
1300
- rejectedProposal: opts.correction.rejectedProposal ?? null,
1301
- rejectedResponseExcerpt: opts.correction.rejectedResponse ?? null,
1323
+ // The planner's own thread already holds the full rejected proposal
1324
+ // (worker prompts included); resend only its skeleton so a corrective
1325
+ // turn cannot re-inflate the compacted context (observed: 38k-char
1326
+ // validationFeedback on run 4t6m5a, 37k of it the verbatim proposal).
1327
+ rejectedProposal: compactRejectedProposal(opts.correction.rejectedProposal),
1328
+ rejectedResponseExcerpt: typeof opts.correction.rejectedResponse === 'string'
1329
+ ? opts.correction.rejectedResponse.slice(0, 2000)
1330
+ : null,
1302
1331
  } : null,
1303
1332
  operatorSteering: (this.state.steering ?? []).map((entry) => ({
1304
1333
  id: entry.id,
@@ -16,7 +16,8 @@
16
16
  // not the generator.
17
17
 
18
18
  import { randomBytes } from 'node:crypto';
19
- import { readdirSync, readFileSync, writeFileSync, existsSync, statSync } from 'node:fs';
19
+ import { readdirSync, readFileSync, existsSync, statSync } from 'node:fs';
20
+ import { writeJsonAtomic } from './fsjson.js';
20
21
  import { join } from 'node:path';
21
22
  import { appendEvent } from './events.js';
22
23
  import { aggregateUsage } from '../lib/usage.js';
@@ -235,7 +236,7 @@ export function reconcileInterruptedRun(runDir, state, {
235
236
  delete state.currentPhase;
236
237
  delete state.currentStep;
237
238
  appendEvent(runDir, state, 'run.interrupted_reconciled', { reason, resumable: true });
238
- writeFileSync(statePath, `${JSON.stringify(state, null, 2)}\n`);
239
+ writeJsonAtomic(statePath, state);
239
240
  return state;
240
241
  }
241
242