bullswarm 0.15.0 → 0.17.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +85 -0
- package/README.md +12 -7
- package/docs/experiments/2026-08-29-dogfood-bullswarm-builds-bullswarm.md +182 -0
- package/docs/planner-prompt-audit-2026-08-29.md +33 -0
- package/package.json +1 -1
- package/skill/SKILL.md +6 -2
- package/src/help.js +4 -2
- package/src/workflow/dashboard.js +414 -24
- package/src/workflow/decision.js +3 -0
- package/src/workflow/goal.js +15 -15
- package/src/workflow/runner.js +14 -2
- package/src/workflow/runtime.js +16 -3
package/CHANGELOG.md
CHANGED
|
@@ -1,5 +1,90 @@
|
|
|
1
1
|
# bullswarm changelog
|
|
2
2
|
|
|
3
|
+
## Unreleased
|
|
4
|
+
|
|
5
|
+
- Workflow timeline (PR #5) hardened after a 16-agent adversarial review against
|
|
6
|
+
real run state (23 findings, 21 confirmed): worker rows now name their phase
|
|
7
|
+
(`├─✓ [Verify] verify-impl`) because concurrent phases interleave in time
|
|
8
|
+
order and the tree glyph alone hung a row under the wrong phase; a phase whose
|
|
9
|
+
actions never started (a blocked tail) is shown as `[Phase: X] blocked`
|
|
10
|
+
instead of vanishing; the header line is truncated so widths down to 20
|
|
11
|
+
columns really hold; PgUp now scrolls the timeline to earlier rows (it was a
|
|
12
|
+
dead key at the newest view) and scroll state resets when the pane changes;
|
|
13
|
+
below 100 columns the footer and status line no longer advertise a timeline
|
|
14
|
+
the narrow layout does not render. Confirmed minors left open are listed on
|
|
15
|
+
PR #5.
|
|
16
|
+
- Reworked the autonomous workflow TUI around a human-readable execution story:
|
|
17
|
+
the existing Workflow Planner and phase sidebar now sits beside a timestamped
|
|
18
|
+
timeline of completed preflight, planner, phase, and worker milestones; active
|
|
19
|
+
workers and the waiting/running planner are isolated in a Live section with
|
|
20
|
+
their latest normalized action and stream heartbeat, and future work stays in
|
|
21
|
+
a distinct Next section. `v` keeps raw action-ledger and event evidence one key
|
|
22
|
+
away without mixing it into the default view.
|
|
23
|
+
- Rule 7: a verify checks the goal's own acceptance criteria and never adds a
|
|
24
|
+
process rule the goal does not state (append-only, existing tests
|
|
25
|
+
untouched); when the implementation changes what an existing assertion
|
|
26
|
+
pins, a worker must own updating it. Earned three times on goal 4 (attempt
|
|
27
|
+
3, `r2vu9i`, `euh622`): the planner wrote "EXTEND BY APPENDING only" into
|
|
28
|
+
the test worker and its verify while the goal said "do NOT modify existing
|
|
29
|
+
tests except to extend them" and item 5 forced `programFeatures` to grow,
|
|
30
|
+
so the assertion at `workflow-adaptive.test.js:206` had no owner, the
|
|
31
|
+
re-verify rejected the mandated extension, and a planner turn recovered.
|
|
32
|
+
- Goal-4 rerun on v0.16.0 (`euh622`): 36 min 00 s, four stage phases in the TUI
|
|
33
|
+
(implement, tests, verify, report) instead of sixteen one-action rows, 22/24
|
|
34
|
+
dispatches on `kaihk/gpt-5.6-luna`, auto-completed, 319/319; two planner
|
|
35
|
+
turns because of the false rejection above (planner turn 2: "an append-only
|
|
36
|
+
rule that the goal itself makes unsatisfiable").
|
|
37
|
+
|
|
38
|
+
## 0.16.0 — the planner sets the width; a re-verify judges the repair
|
|
39
|
+
|
|
40
|
+
- A re-verify after a repair round now receives the concerns it raised and the
|
|
41
|
+
repair's report, and may return ok:false only for an unresolved listed
|
|
42
|
+
concern or a regression; anything newly noticed is informational. Earned on
|
|
43
|
+
`r2vu9i`: `verify-src` round 2 rejected on two concerns round 1 never raised;
|
|
44
|
+
`verify-tests-runtime` round 2 rejected the very edit its round 1 demanded.
|
|
45
|
+
Live-proven on `bizp4s`: the one re-verify rejection was an `ENOENT`
|
|
46
|
+
regression in the acceptance checks, and its verdict opens "the two
|
|
47
|
+
original concerns are repaired".
|
|
48
|
+
- A phase is a pipeline stage: rule 3 now says one kebab-case name shared by
|
|
49
|
+
its actions, never one phase per action, and the complete-program example
|
|
50
|
+
uses five phases for eight actions (`verify` holds verify-fix, verify-tests
|
|
51
|
+
and verify-suite). Earned on `bizp4s`: the planner mirrored the example and
|
|
52
|
+
wrote sixteen one-action phases — no scheduling cost (phases never gate;
|
|
53
|
+
`dependsOn` does), but a TUI phase list carrying no information.
|
|
54
|
+
- Goal-4 rerun on this release (`bizp4s`, runtime `9af8fdf`, workers on
|
|
55
|
+
`kaihk/gpt-5.6-luna`): **25 min 13 s** (attempt 3: 44 min; 0.15.0: 72 min;
|
|
56
|
+
audited contract alone: 37 min), one planner turn (247 s, 16 % of wall),
|
|
57
|
+
parallelism 1.77, 3 repair rounds each fixing a real defect, 0 schema
|
|
58
|
+
retries, 0 corrections, auto-completed, 319/319, existing tests +174/−0.
|
|
59
|
+
- Goal-4 rerun on the audited contract (`r2vu9i`): 36 min 58 s (attempt 3:
|
|
60
|
+
44 min; 0.15.0: 72 min), 5 parallel writers, parallelism 1.55, tests depend
|
|
61
|
+
on the implementation run rather than its verify, 0 schema retries, 0
|
|
62
|
+
corrections, auto-completed, 314/314.
|
|
63
|
+
- Planner contract audited against Claude Code's workflow-authoring reference
|
|
64
|
+
(three-lens review + adversarial verification, run on the real goal-4 task
|
|
65
|
+
text) and rewritten within the same caps (rules 3,999 / examples 2,938
|
|
66
|
+
chars). New in substance: a verdict is never data (depend on the run that
|
|
67
|
+
wrote your files, not on its verify); split to the width the tree allows
|
|
68
|
+
(one worker for N independent files is N chains in series); outputSchema
|
|
69
|
+
only where a later action reads the object, never on prose; a repair edits
|
|
70
|
+
files and cannot rewrite the answer under review; workers run their unit's
|
|
71
|
+
focused command, never the full suite; the planner sets `lane` and `effort`
|
|
72
|
+
per action. The complete-program example is now valid JSON and shows tests
|
|
73
|
+
running beside the src verify. `docs/planner-prompt-audit-2026-08-29.md` §6.
|
|
74
|
+
- Planner-proposed `lane`/`effort`/`requiresCapabilities` now survive the
|
|
75
|
+
gate defaults (`runner.js` spread order let `lane: build` overwrite every
|
|
76
|
+
proposal); `lane` is validated like `effort`.
|
|
77
|
+
- `outputSchema` output reading tolerates a closing markdown fence after the
|
|
78
|
+
trailing JSON object, and the schema instruction says the object is an
|
|
79
|
+
INSTANCE whose keys are the `properties` names (never the schema itself).
|
|
80
|
+
Earned on the goal-4 rerun `ydpjts` (0.15.0): a stray `"type"` key and then
|
|
81
|
+
a `}\n```` tail spent the single schema retry and a 279 s planner turn on
|
|
82
|
+
an otherwise complete report (≈ 11 min).
|
|
83
|
+
- Goal-4 rerun recorded in `docs/experiments/2026-08-29-dogfood-bullswarm-builds-bullswarm.md`:
|
|
84
|
+
0 repairs (attempt 3: 6), planner 17 % of wall (39 %), 326/326 — but 72 min
|
|
85
|
+
vs 44 min because every worker landed on the slowest most-behind pool and
|
|
86
|
+
the program ran serially (parallelism 1.05).
|
|
87
|
+
|
|
3
88
|
## 0.15.0 — extra Claude Code logins as separate pools
|
|
4
89
|
|
|
5
90
|
- Claude Code extra logins (`~/.claude-<slug>` / `$CLAUDE_CONFIG_DIR`) become
|
package/README.md
CHANGED
|
@@ -289,14 +289,19 @@ bullswarm workflow watch <shortId> --verbose # detailed agent/action view
|
|
|
289
289
|
```
|
|
290
290
|
|
|
291
291
|
`workflow tui` is the interactive, Claude-style `/workflows` view. For an
|
|
292
|
-
autonomous goal its left navigation stacks a compact
|
|
293
|
-
the Phases panel; internal planner turns never appear as workers or phases.
|
|
294
|
-
|
|
295
|
-
|
|
292
|
+
autonomous goal its left navigation stacks a compact Workflow Planner panel
|
|
293
|
+
above the Phases panel; internal planner turns never appear as workers or phases.
|
|
294
|
+
The default desktop main panel is a timestamped workflow timeline: completed
|
|
295
|
+
preflight, planner-checkpoint, phase-transition, and worker-result events stay
|
|
296
|
+
above a live section containing the waiting/running Workflow Planner and workers,
|
|
297
|
+
each with its latest semantic action and stream heartbeat. Planned work is kept
|
|
298
|
+
in a separate Next section so it cannot be mistaken for execution evidence.
|
|
299
|
+
Select Workflow Planner and press Enter, or press `o`
|
|
300
|
+
anywhere, to open a summary-first planner overview: what it is doing now,
|
|
296
301
|
its latest decision in plain language, why it chose that path, what happens
|
|
297
|
-
next, progress, and the last three semantic actions. Press `v`
|
|
298
|
-
technical
|
|
299
|
-
and artifact paths. Status marks are consistent throughout the tree: `○` not started,
|
|
302
|
+
next, progress, and the last three semantic actions. Press `v` from the timeline
|
|
303
|
+
for workflow technical state, or from Workflow Planner for provider session,
|
|
304
|
+
every checkpoint turn, usage, prompt, and artifact paths. Status marks are consistent throughout the tree: `○` not started,
|
|
300
305
|
an animated Braille spinner for active work, `⧖` waiting, `✓` finished, and
|
|
301
306
|
`✗` failed or interrupted. The non-emoji `⧖` avoids the inconsistent cell
|
|
302
307
|
width of `⌛` across terminal fonts. It watches ongoing runs from disk and supports `j`/`k` or arrow-key selection, Enter for
|
|
@@ -169,3 +169,185 @@ report → verify-final, `completion: all-actions-ok`. Wall is within noise of b
|
|
|
169
169
|
model time: probe 4.7 min, guards 5.1 min, verify-guards 4.7 min); the planner side is faster and 5× smaller.
|
|
170
170
|
Zero observation crashes across four runs of TUI/watch/runs/result/static-tui polling and a 20 s stress loop
|
|
171
171
|
(1,681 paints against the live writer, 0 torn, 0 throws).
|
|
172
|
+
|
|
173
|
+
## Goal-4 rerun on 0.15.0 — `ydpjts` (wf-mte8azjz-bcb079), 10:19:31 → 11:31:48 Z — reliability PASS, speed FAIL
|
|
174
|
+
|
|
175
|
+
Same goal text (`goal4.txt`), same flags (`--orchestrator claude-code --concurrency 8 --max-agents 40
|
|
176
|
+
--max-expansion-rounds 8`), fresh fixture `g4-bs-v2` = `git archive a0f0965` (v0.13.2, 299/299, no `schema.js`) on
|
|
177
|
+
branch `feat/output-schema`, default `~/.bullswarm` home, runtime `bullswarm-rt` @ 728231d (v0.15.0).
|
|
178
|
+
|
|
179
|
+
| metric | attempt 3 (0.13.2) | **rerun (0.15.0)** |
|
|
180
|
+
| --- | ---: | ---: |
|
|
181
|
+
| outcome | auto-completed after planner turn 2 | **auto-completed** after a 2-action recovery program |
|
|
182
|
+
| wall | 44 min 17 s | **72 min 14 s** (4 334 s) |
|
|
183
|
+
| planner turns / plannerSec | 2 / 1 045 s (39 %; turn 1 = 634 s) | **2 / 755 s (17 %)** — 476 s + 279 s; completion recorded by the runtime |
|
|
184
|
+
| planner context (turn 1 / turn 2) | — | 12.9 k / 41.5 k chars (`decision.context_built`) |
|
|
185
|
+
| dispatches | 28 (26 opencode2 + 2 planner) | **13** — 2 planner (opus-5) + 11 workers (sonnet-5), all on `claude-code` |
|
|
186
|
+
| max concurrent / parallelism | 5 / 1.5 | **2 / 1.05** |
|
|
187
|
+
| repairs | 6 (4 unrepairable: process-criteria rejections) | **0** — every verify ok:true first round |
|
|
188
|
+
| corrections / rejections / verdict re-asks | 0 / 0 / — | 0 / 0 / 0 |
|
|
189
|
+
| schema retries | — | 1 (`report`; both attempts failed, see below) |
|
|
190
|
+
| tests after | 318/318 (299 + 19) | **326/326** (299 + 27); existing test files extended only (+281 / −0); no commit; version untouched |
|
|
191
|
+
|
|
192
|
+
Program (turn 1, 476 s): `impl-src` ∥ `docs` → `verify-src`(repair 2) / `verify-docs`(repair 1) → `tests` →
|
|
193
|
+
`verify-tests`(repair 2) → `report`(outputSchema) → `verify-report`(repair 1, `review` defaulted), `completion`
|
|
194
|
+
attached. Three verifies omitted `review` and were accepted (0.14.1 defaulting) — the same omission cost a 5-minute
|
|
195
|
+
correction turn on proof runs 1 and 3. Rule 8 honoured (last worker covered by a verify).
|
|
196
|
+
|
|
197
|
+
Verifier behaviour is what 0.14.x was meant to produce: `verify-src` passed with a disclosed deviation
|
|
198
|
+
(`programFeatures` literal not extended because `tests/workflow-adaptive.test.js:206` regex-locks it and the goal
|
|
199
|
+
forbids modifying existing assertions) instead of rejecting on a process criterion; `verify-tests` flagged the filler
|
|
200
|
+
prose workaround as a concern, not a failure.
|
|
201
|
+
|
|
202
|
+
**The one failure and its cost (≈ 11.4 min):** `report` carried an `outputSchema`. Attempt 1 ended with an object
|
|
203
|
+
carrying a stray `"type"` key copied from the schema (`additionalProperties:false` → `type is not allowed`, correct).
|
|
204
|
+
The single retry ended `}\n\`\`\`` — a closing markdown fence — and `readTrailingObject` refused it as "did not end
|
|
205
|
+
with a JSON object" → `report` ok:false → `verify-report` blocked → `decision.completion_predicate_unmet` → planner
|
|
206
|
+
turn 2 (279 s), which diagnosed it correctly and re-delivered the report without a schema. Fixed in `393a914`:
|
|
207
|
+
trailing fences are stripped before the object is read, and the instruction says the object is an INSTANCE whose keys
|
|
208
|
+
are the `properties` names. Unit-tested; not yet exercised live.
|
|
209
|
+
|
|
210
|
+
**Where the 72 minutes went — the run was serial (parallelism 1.05):** scout 405 s → planner 476 s → `impl-src`
|
|
211
|
+
**1 231 s** → `verify-src` 330 s → `tests` **951 s** → `verify-tests` 116 s → `report` 141 + 123 s → planner 279 s →
|
|
212
|
+
`final-summary` 174 s → `verify-final-summary` 109 s. Two causes, neither a crash:
|
|
213
|
+
1. Routing put every worker on `claude-code`/`claude-sonnet-5` because it was the most-behind capable pool
|
|
214
|
+
(pace +6.9 vs opencode2 0, `claude-code:wati` −14.8). The same `impl-src` took 398 s on `opencode2`
|
|
215
|
+
(`kaihk/gpt-5.6-luna`) in attempt 3; the read-only scout took 405 s vs 92 s. Worker model speed is not part of the
|
|
216
|
+
surplus formula.
|
|
217
|
+
2. The planner proposed ONE `tests` worker after `verify-src` instead of three file-disjoint test writers in parallel
|
|
218
|
+
with the src verify (attempt 3's shape). Reliability-first, width-second: correct under the goal's single-implementer
|
|
219
|
+
constraint, but it lengthened the critical path by ~16 min.
|
|
220
|
+
|
|
221
|
+
Pass conditions set before the run: ≤ 1 repair round ✓ (0); 0 process-criteria rejections ✓; auto-completed ✓;
|
|
222
|
+
≤ 35 min ✗ (72 min). Reliability at 12-action complexity is now proven on ≥ 0.14.1; speed is worker-bound and
|
|
223
|
+
routing-bound.
|
|
224
|
+
|
|
225
|
+
## Goal-4 rerun on the audited contract — `r2vu9i` (wf-mtefmdie-b39e8f), 13:44:19 → 14:21:21 Z — **structural PASS, 36 min 58 s**
|
|
226
|
+
|
|
227
|
+
Runtime `d14c1fa` (contract from the second audit, §6 of `docs/planner-prompt-audit-2026-08-29.md`, plus the runner
|
|
228
|
+
lane fix); workers pinned to `opencode2`/`kaihk/gpt-5.6-luna` via `strategy assign high|medium|low` (cleared after);
|
|
229
|
+
orchestrator `claude-code`/`claude-opus-5`; fresh fixture `g4-bs-v3` = a0f0965 (299/299).
|
|
230
|
+
|
|
231
|
+
| metric | attempt 3 (0.13.2) | `ydpjts` (0.15.0) | **`r2vu9i` (audited contract)** |
|
|
232
|
+
| --- | ---: | ---: | ---: |
|
|
233
|
+
| wall | 44 min 17 s | 72 min 14 s | **36 min 58 s** (2 219 s) |
|
|
234
|
+
| planner turns / plannerSec | 2 / 1 045 s (39 %) | 2 / 755 s (17 %) | 2 / 648 s (29 %) — 409 s + 238 s |
|
|
235
|
+
| dispatches | 28 | 13 | 26 (24 gpt-5.6-luna workers + 2 opus planner) |
|
|
236
|
+
| max concurrent / parallelism | 5 / 1.5 | 2 / 1.05 | **4 / 1.55** |
|
|
237
|
+
| writers in parallel after planning | 2 (+3 test writers later) | 2 | **5** (impl-src ∥ 3 docs; 2 test writers as soon as impl-src landed) |
|
|
238
|
+
| repairs (rounds / repaired ok / re-verify rejected) | 6 / 2 / 4 | 0 | 4 / 2 / 2 |
|
|
239
|
+
| schema retries / corrections / verdict re-asks | — / 0 / — | 1 / 0 / 0 | **0 / 0 / 0** |
|
|
240
|
+
| tests after | 318 | 326 | **314/314** (299 + 15); existing tests +167 / −1 (one assertion extended, as goal item 5 requires); no commit; version untouched |
|
|
241
|
+
|
|
242
|
+
Structural pass conditions (set before launch): `tests-*` depend on `impl-src`, not `verify-src` ✓ (the planner's own
|
|
243
|
+
reason: "two file-disjoint test workers depend on that run (not its verify)"); more than one file-disjoint writer ✓ (5);
|
|
244
|
+
parallelism ≥ 1.5 ✓ (1.55); 0 corrections ✓; 0 schema retries ✓ (the report carried no schema); auto-completed ✓;
|
|
245
|
+
≥ 299 tests, existing tests extended only ✓. Failed: 0 repairs ✗ (4 rounds). Lane/effort proposed by the planner reached
|
|
246
|
+
dispatch (`docs-changelog` routed `chore`/`low`) — the runner merge fix is live. `accept-suite` was a verify with no
|
|
247
|
+
dependsOn — the first live exercise of the repository-scope branch.
|
|
248
|
+
|
|
249
|
+
Where the four repair rounds went, and what each says:
|
|
250
|
+
- `verify-tests-schema` round 1: missing invalid-`minimum` case → repaired, re-verify ok. Legitimate; cost 49 + 30 s.
|
|
251
|
+
- `verify-src` round 1: a real defect (`recordOutput` dropped `data`/`schemaOk`, breaking resume-safety) → repaired.
|
|
252
|
+
Legitimate. Round 2 then rejected on two concerns round 1 never raised (enum structural equality, root-path naming)
|
|
253
|
+
→ repaired, re-verify ok. **Moving goalposts**: 92 + 268 s spent on nits that should have been round-1 concerns.
|
|
254
|
+
- `verify-tests-runtime` round 1: acceptance command failed on the pre-existing `workflow-adaptive.test.js:206`
|
|
255
|
+
assertion (goal item 5 requires extending it) → the repair extended it → round 2 rejected BECAUSE an existing test
|
|
256
|
+
was modified — a process criterion the planner had written into the verify prompt ("append tests only") that
|
|
257
|
+
contradicts the goal. Unrepairable by construction → `verify-suite`/report blocked → planner turn 2 (238 s), which
|
|
258
|
+
diagnosed "a false rejection" and recovered in 207 s. Cost ≈ 10 min.
|
|
259
|
+
Fix shipped after the run (`9af8fdf`, unit-tested, not yet live): a re-verify receives the concerns it raised and the
|
|
260
|
+
repair's report and may reject only for an unresolved listed concern or a regression — both round-2 rejections above
|
|
261
|
+
become informational under it.
|
|
262
|
+
|
|
263
|
+
Speed accounting vs `ydpjts`: worker pool (gpt-5.6-luna vs sonnet-5) and width together took the critical path from
|
|
264
|
+
72 to ~27 min of productive work; the remaining ~10 min is the verifier-behaviour waste above.
|
|
265
|
+
|
|
266
|
+
## Goal-4 rerun with the re-verify fix — `bizp4s` (wf-mtehhbwd-7db3de), 14:36:26 → 15:01:39 Z — **PASS, 25 min 13 s**
|
|
267
|
+
|
|
268
|
+
Runtime `9af8fdf` (audited contract + runner lane fix + re-verify scoping); workers pinned to
|
|
269
|
+
`opencode2`/`kaihk/gpt-5.6-luna` via `strategy assign high|medium|low` (cleared after, assignments `{}`); orchestrator
|
|
270
|
+
`claude-code`/`claude-opus-5`; fresh fixture `g4-bs-v4` = a0f0965 (299/299, cloned clean from the v3 fixture commit).
|
|
271
|
+
|
|
272
|
+
| metric | attempt 3 (0.13.2) | `ydpjts` (0.15.0) | `r2vu9i` (audited contract) | **`bizp4s` (+ re-verify fix)** |
|
|
273
|
+
| --- | ---: | ---: | ---: | ---: |
|
|
274
|
+
| wall | 44 min 17 s | 72 min 14 s | 36 min 58 s | **25 min 13 s** (1 513 s) |
|
|
275
|
+
| planner turns / plannerSec | 2 / 1 045 s (39 %) | 2 / 755 s (17 %) | 2 / 648 s (29 %) | **1 / 247 s (16 %)** |
|
|
276
|
+
| dispatches | 28 | 13 | 26 | 23 (22 gpt-5.6-luna + 1 opus planner) |
|
|
277
|
+
| max concurrent / parallelism | 5 / 1.5 | 2 / 1.05 | 4 / 1.55 | **4 / 1.77** |
|
|
278
|
+
| writers in parallel after planning | 2 (+3) | 2 | 5 | 4 (impl-src ∥ 3 docs), then 2 test writers ∥ verify-src the moment impl-src landed |
|
|
279
|
+
| repairs (rounds / repaired ok / re-verify rejected) | 6 / 2 / 4 | 0 | 4 / 2 / 2 | 3 / 2 / 1 — every rejection a real defect |
|
|
280
|
+
| schema retries / corrections / verdict re-asks | — / 0 / — | 1 / 0 / 0 | 0 / 0 / 0 | **0 / 0 / 0** |
|
|
281
|
+
| tests after | 318 | 326 | 314/314 | **319/319** (299 + 20); existing tests +174 / −0; no commit; version untouched |
|
|
282
|
+
|
|
283
|
+
Program (one decision, 15 actions + completion): `impl-src`(build/high) ∥ `doc-changelog`(chore/low) ∥
|
|
284
|
+
`doc-skill`(chore/low) ∥ `doc-mechanics`(chore/medium), each doc with its own analyze/low verify; `test-schema`
|
|
285
|
+
(build/medium), `test-runtime`(build/high) and `verify-src`(analyze/high) all `dependsOn: ["impl-src"]`; `verify-suite`
|
|
286
|
+
(analyze/low) on the six unit verifies; `report`(chore/low, no schema) → `verify-report`(analyze/medium);
|
|
287
|
+
`completion: all-actions-ok`. Every planner-set effort tier reached dispatch (`configured <tier> assignment`).
|
|
288
|
+
`decision.auto_completed` — the planner was consulted exactly once.
|
|
289
|
+
|
|
290
|
+
The three repair rounds, and why none is verifier waste:
|
|
291
|
+
- `verify-src` round 1 (ok:false): `validateWorkflow` did not check `stepTemplate.outputSchema` and the runtime ignored it
|
|
292
|
+
during fan-out — the fan-out half of the goal was unimplemented. Repair 197 s.
|
|
293
|
+
- `verify-src` re-verify round 1 (ok:false, `action.reverify_rejected`): the repair's schema-retry path handed dispatch a
|
|
294
|
+
file name that dispatch re-suffixed `-attempt-2`, so the runtime read a nonexistent file (`ENOENT`; 3 focused tests
|
|
295
|
+
failing). A regression in the acceptance checks — exactly the rejection the new scoping still allows. The task text
|
|
296
|
+
carried `RE-VERIFY round 1 of 2 … Concerns you raised (verbatim):` with round 1's concerns, and the verdict's summary
|
|
297
|
+
opens "The two original fan-out concerns are repaired" — the verifier judged the repair, as instructed. Repair 177 s;
|
|
298
|
+
round 2 re-verify ok (52/52 focused).
|
|
299
|
+
- `verify-test-schema` round 1 (ok:false): no invalid-value case per supported keyword. Repair 63 s, re-verify ok.
|
|
300
|
+
- `verify-test-runtime` accepted first time: the goal's item-5 assertion was satisfied without touching an existing test
|
|
301
|
+
(+174 / −0), so the "append only" tension of `r2vu9i` never arose.
|
|
302
|
+
|
|
303
|
+
Pass conditions (set before `r2vu9i`): tests depend on `impl-src` ✓; > 1 file-disjoint writer ✓ (4, then 2 more);
|
|
304
|
+
parallelism ≥ 1.5 ✓ (1.77); 0 corrections ✓; 0 schema retries ✓; auto-completed ✓; ≥ 299 tests, existing tests only
|
|
305
|
+
extended ✓. "0 repairs" ✗ as a literal count (3), but the condition's intent — no repair round that a verifier caused —
|
|
306
|
+
is met: each round fixed a defect the deliverable needed fixed. Prediction before launch was "the two moving-goalpost
|
|
307
|
+
rounds and the 10-min recovery turn vanish, wall ≈ 30 min"; observed 25 min 13 s with one planner turn.
|
|
308
|
+
|
|
309
|
+
Cost: 22 of 23 dispatches on the unmetered opencode2 seat; Claude quota spent on one 247 s planner turn.
|
|
310
|
+
|
|
311
|
+
## Goal-4 rerun on v0.16.0 (phase = stage) — `euh622` (wf-mtej85ws-18a3c0), 15:25:14 → 16:01:14 Z — **36 min 00 s, auto-completed; one planner-authored false rejection cost the recovery turn**
|
|
312
|
+
|
|
313
|
+
Runtime `4bfd7f4` = released v0.16.0 (rule 3 "a phase is a pipeline stage … never one per action"); workers pinned to
|
|
314
|
+
`opencode2`/`kaihk/gpt-5.6-luna` (cleared after); orchestrator `claude-code`/`claude-opus-5`; fresh fixture `g4-bs-v5`.
|
|
315
|
+
|
|
316
|
+
| metric | `r2vu9i` | `bizp4s` | **`euh622`** |
|
|
317
|
+
| --- | ---: | ---: | ---: |
|
|
318
|
+
| wall | 36 min 58 s | 25 min 13 s | 36 min 00 s (2 157 s) |
|
|
319
|
+
| planner turns / plannerSec | 2 / 648 s | 1 / 247 s | 2 / 728 s (34 %) — 462 s + 266 s |
|
|
320
|
+
| dispatches | 26 | 23 | 24 (22 luna + 2 opus) |
|
|
321
|
+
| max concurrent / parallelism | 4 / 1.55 | 4 / 1.77 | 4 / 1.5 |
|
|
322
|
+
| phases in the TUI | 19 one-action rows | 16 one-action rows | **4 stages** (implement 2, tests 3, verify 6 + repairs, report 2) + 2 recovery phases |
|
|
323
|
+
| repairs (rounds / repaired ok / re-verify rejected) | 4 / 2 / 2 | 3 / 2 / 1 | 4 / 2 / 2 |
|
|
324
|
+
| tests after | 314 | 319 | **319/319**; existing tests +116/−1 (the mandated `:206` extension) and +153/−0 |
|
|
325
|
+
|
|
326
|
+
What the phase change did: the planner wrote `implement` (impl ∥ docs), `tests` (three test writers), `verify` (six
|
|
327
|
+
verifies), `report` — the layout asked for, with no scheduling change (impl ∥ docs started together; three test writers
|
|
328
|
+
and verify-impl started the second impl landed; verify-docs ran while impl was still running). Width was one docs
|
|
329
|
+
worker (three files merged; off the critical path) but three test writers — comparable to `bizp4s`.
|
|
330
|
+
|
|
331
|
+
The four repair rounds:
|
|
332
|
+
- `verify-impl` r1: fan-out items dispatched through plain `dispatch()`, so `stepTemplate.outputSchema` was never applied
|
|
333
|
+
— real. Re-verify rejected: the repair tested `step.outputSchema` instead of `itemStep.outputSchema` — the listed
|
|
334
|
+
concern still unresolved, exactly the rejection the re-verify rule permits. r2 repaired; re-verify ok.
|
|
335
|
+
- `verify-test-runtime` r1: the retry case did not assert the `errors` payload of `action.output_schema_retry` — real.
|
|
336
|
+
- `verify-test-refs` r1: the acceptance command failed on the pre-existing assertion at `workflow-adaptive.test.js:206`
|
|
337
|
+
(`programFeatures` pinned to three entries) because goal item 5 mandates adding `outputSchema` to it. The repair
|
|
338
|
+
extended the assertion (+116/−1). Re-verify rejected BECAUSE an existing assertion changed — the planner had written
|
|
339
|
+
"EXTEND BY APPENDING new test cases only" into `test-refs` and "shows APPENDED cases only" into `verify-test-refs`,
|
|
340
|
+
and "Do NOT modify existing tests" into `impl`, while the goal says "do NOT modify existing tests except to extend
|
|
341
|
+
them". No worker owned the assertion; the verifier treated its prompt's rule as an unresolved concern. `verify-suite`,
|
|
342
|
+
`report`, `verify-report` blocked → planner turn 2 (266 s), whose reason is exact: "verify-test-refs ended ok:false
|
|
343
|
+
on an append-only rule that the goal itself makes unsatisfiable — goal item 5 mandates adding 'outputSchema' to
|
|
344
|
+
programFeatures". Recovery program `verify-suite-full` → `final-report` → `verify-final-report`, auto-completed.
|
|
345
|
+
Third occurrence of this shape (attempt 3, `r2vu9i`, here); `bizp4s` avoided it only because its implementation
|
|
346
|
+
happened to keep the old assertion true.
|
|
347
|
+
|
|
348
|
+
Cost of the false rejection: the 266 s planner turn plus the serialised tail ≈ 5–6 min; the rest of the gap to
|
|
349
|
+
`bizp4s` is variance (planner turn 1 462 s vs 247 s on the same contract; `impl` 445 s vs 377 s).
|
|
350
|
+
|
|
351
|
+
Fix committed after the run, unreleased (`71960ae`): rule 7 — "A verify checks the goal's own acceptance criteria …
|
|
352
|
+
never add a process rule the goal does not state (append-only, tests untouched); when the implementation changes what
|
|
353
|
+
an existing assertion pins, a worker must own updating it." Proof pending a rerun on that commit.
|
|
@@ -155,3 +155,36 @@ Deliverables:
|
|
|
155
155
|
context and the separate run-state/TUI attempt view, but does not document a
|
|
156
156
|
renamed/dropped planner-context field shape such as the old attempt records
|
|
157
157
|
or an old `failures` representation.
|
|
158
|
+
|
|
159
|
+
## 6. Second audit — after the goal-4 rerun on 0.15.0 (`ydpjts`)
|
|
160
|
+
|
|
161
|
+
Method: a Claude Code dynamic workflow (12 agents, 16 min) reviewed the COMPLETE turn-1 task text of run `ydpjts`
|
|
162
|
+
(19,577 chars; 6,199-char prefix) and the program the planner wrote, from three lenses — width/critical path,
|
|
163
|
+
worker-prompt authoring, clarity vs Claude's own workflow-authoring reference — and adversarially verified every new
|
|
164
|
+
finding (refute-framed, source-checked). 3 known + 6 new findings survived; 2 were refuted (one because `lane` was
|
|
165
|
+
dead text — see the runtime defect below — one for deleting a must-keep rule).
|
|
166
|
+
|
|
167
|
+
What the planner received said nothing wrong; it said too little in three places, and the example showed a linear
|
|
168
|
+
program:
|
|
169
|
+
|
|
170
|
+
| finding | evidence on `ydpjts` | contract change |
|
|
171
|
+
| --- | --- | --- |
|
|
172
|
+
| a verify's verdict was treated as a dependency | `tests` dependsOn `verify-src`; 330 s idle | rule 3: "a worker depends on the run that wrote its input files, never on that run's verify (a verdict is not data)" |
|
|
173
|
+
| width framed only as "known N items" | one `tests` worker, 951 s, where three file-disjoint writers were allowed | rule 4: "each file-disjoint unit ... one worker for N independent files is N chains in series" |
|
|
174
|
+
| outputSchema invited on a prose report | `report` schema → stray key, then fenced tail → retry + planner turn (≈ 685 s) | rule 6: schema only where a later action reads the object; prose gets none |
|
|
175
|
+
| repair cannot rewrite an answer under review (new) | `verify-report` repair prompt asked the worker to "restate the report" — unreadable by design (`repairAndReverify` re-reads the original outFile) | rule 7: "the repair edits files and cannot rewrite the answer under review, so reject only what a file edit can fix" |
|
|
176
|
+
| workers demanded `npm test` as acceptance (new) | `tests` prompt: "No other worker is editing the tree while you run" | shared-tree line: unit's focused command, never the full suite |
|
|
177
|
+
| no guidance on `effort`/`lane` (new) | every action ran at build/medium | rule 10: planner sets lane and effort; they pick the model tier |
|
|
178
|
+
| example program was linear and not valid JSON (new) | — | example rewritten: `{"actions":[…],"completion":…}`, tests depend on fix and run beside verify-fix, lane/effort shown |
|
|
179
|
+
| runtime: planner `lane` silently overwritten (new) | `runner.js` merged `{...action, ...actionDefaults}` so `lane: build` won | fixed: action overrides gate defaults; addDir stays runtime-owned; `lane` validated |
|
|
180
|
+
|
|
181
|
+
Sizes after: rules **3,999** (cap 4,000), examples **2,938** (cap 3,000). Every rule still stated once with its reason.
|
|
182
|
+
Acceptance: goal-4 rerun on the new contract with workers pinned to `opencode2`/`kaihk/gpt-5.6-luna` (strategy
|
|
183
|
+
assignments high/medium/low, cleared after) — structural pass conditions: tests depend on `impl-src` not
|
|
184
|
+
`verify-src`; more than one file-disjoint writer; parallelism ≥ 1.5; 0 repairs / corrections / schema retries;
|
|
185
|
+
auto-completed; ≥ 299 tests, existing tests extended only; wall reported with the pool mix.
|
|
186
|
+
|
|
187
|
+
Outcome: `r2vu9i` (runtime `d14c1fa`) met every structural condition except "0 repairs" (4 rounds, 2 of them verifier
|
|
188
|
+
moving-goalposts → re-verify scoping `9af8fdf`); `bizp4s` (runtime `9af8fdf`) then ran in 25 min 13 s with one planner
|
|
189
|
+
turn (247 s), parallelism 1.77, 3 repair rounds each fixing a real defect, 0 schema retries / corrections, 319/319.
|
|
190
|
+
Details in `docs/experiments/2026-08-29-dogfood-bullswarm-builds-bullswarm.md`.
|
package/package.json
CHANGED
package/skill/SKILL.md
CHANGED
|
@@ -277,8 +277,12 @@ supports explicit model selection, Bullswarm pins an allowed model in the same
|
|
|
277
277
|
effort tier; otherwise that pool is excluded because its implicit default
|
|
278
278
|
cannot be guaranteed. Restore eligibility with `strategy include-model`.
|
|
279
279
|
|
|
280
|
-
In the human TUI, the autonomous orchestrator
|
|
281
|
-
stacked above the phase tree.
|
|
280
|
+
In the human TUI, the autonomous orchestrator is presented as Workflow Planner
|
|
281
|
+
in a compact selectable panel stacked above the phase tree. The default desktop
|
|
282
|
+
view pairs that unchanged navigation with a timestamped workflow timeline:
|
|
283
|
+
finished events stay above a live Planner/worker section with each participant's
|
|
284
|
+
latest semantic action, while unexecuted work stays in a separate Next section.
|
|
285
|
+
Select Workflow Planner and press Enter, or press
|
|
282
286
|
`o`, to see a summary-first overview of its current role, latest decision,
|
|
283
287
|
reason, next action, progress, and recent semantic activity. Press `v` for the
|
|
284
288
|
durable provider session, checkpoint prompts and turns, usage, and artifact
|
package/src/help.js
CHANGED
|
@@ -676,8 +676,9 @@ const workflowInspectText = rich({
|
|
|
676
676
|
|
|
677
677
|
const workflowTuiText = rich({
|
|
678
678
|
usage: 'bullswarm workflow tui [<runId>] [--json] [--all] [--show <runId>] [--cancel <runId>]',
|
|
679
|
-
purpose: 'Open the interactive full-screen
|
|
680
|
-
+ 'and historical runs, or print a
|
|
679
|
+
purpose: 'Open the interactive full-screen workflow timeline with Workflow Planner, phase, '
|
|
680
|
+
+ 'live-agent, and technical drill-down views for ongoing and historical runs, or print a '
|
|
681
|
+
+ 'static/JSON snapshot for a non-interactive caller.',
|
|
681
682
|
args: [{ name: '[<runId>]', desc: 'shortId or runId to open directly in detail view; omit to see the run picker' }],
|
|
682
683
|
options: [
|
|
683
684
|
{ flag: '--json', desc: "print a JSON snapshot instead of opening the interactive browser (list of ongoing runs, or one run's state/report/events when a runId is given)", default: "opens the interactive browser on a TTY; without a TTY, a given runId instead prints one static text detail tree" },
|
|
@@ -689,6 +690,7 @@ const workflowTuiText = rich({
|
|
|
689
690
|
'interactive mode and the --json/--show/--all views are read-only',
|
|
690
691
|
'--cancel writes state.json (cancelRequested=true, status=cancelling) — cooperative, not a force-kill: the workflow stops at its next safe checkpoint',
|
|
691
692
|
'inside the interactive browser, q detaches without stopping the underlying workflow; c requests the same cancellation with a confirmation prompt',
|
|
693
|
+
'the default timeline is derived from durable state and events; press v for raw action-ledger and event evidence',
|
|
692
694
|
],
|
|
693
695
|
examples: [
|
|
694
696
|
{ cmd: 'bullswarm workflow tui', note: 'interactive run picker' },
|
|
@@ -259,6 +259,8 @@ export function workflowPanelModel(row, {
|
|
|
259
259
|
const isControlAction = (action) => orchestrator.autonomous
|
|
260
260
|
&& action.id === orchestrator.actionId
|
|
261
261
|
&& action.kind === 'decide';
|
|
262
|
+
const isPreflightAction = (action) => orchestrator.autonomous && action.id === 'scout';
|
|
263
|
+
const isNonPhaseAction = (action) => isControlAction(action) || isPreflightAction(action);
|
|
262
264
|
const phaseNames = [];
|
|
263
265
|
const addPhase = (name) => {
|
|
264
266
|
if (name && !phaseNames.includes(name)) phaseNames.push(name);
|
|
@@ -266,12 +268,13 @@ export function workflowPanelModel(row, {
|
|
|
266
268
|
for (const phase of state._doc?.phases ?? []) {
|
|
267
269
|
const controlOnly = orchestrator.autonomous
|
|
268
270
|
&& (phase.steps ?? []).length
|
|
269
|
-
&& (phase.steps ?? []).every((step) =>
|
|
271
|
+
&& (phase.steps ?? []).every((step) =>
|
|
272
|
+
(step.id === orchestrator.actionId && step.type === 'decide') || step.id === 'scout');
|
|
270
273
|
if (!controlOnly) addPhase(phase.name);
|
|
271
274
|
}
|
|
272
|
-
for (const action of ledger) if (!
|
|
273
|
-
for (const step of state.steps ?? []) if (step.stepId !== orchestrator.actionId) addPhase(step.phase);
|
|
274
|
-
if (state.currentStep?.id !== orchestrator.actionId) addPhase(state.currentStep?.phase);
|
|
275
|
+
for (const action of ledger) if (!isNonPhaseAction(action)) addPhase(action.phase);
|
|
276
|
+
for (const step of state.steps ?? []) if (step.stepId !== orchestrator.actionId && step.stepId !== 'scout') addPhase(step.phase);
|
|
277
|
+
if (state.currentStep?.id !== orchestrator.actionId && state.currentStep?.id !== 'scout') addPhase(state.currentStep?.phase);
|
|
275
278
|
if (!orchestrator.autonomous) addPhase(state.currentPhase?.name);
|
|
276
279
|
if (!phaseNames.length) phaseNames.push(orchestrator.autonomous ? 'execution' : 'starting');
|
|
277
280
|
|
|
@@ -283,7 +286,7 @@ export function workflowPanelModel(row, {
|
|
|
283
286
|
phaseNames.length - 1,
|
|
284
287
|
);
|
|
285
288
|
const phases = phaseNames.map((name) => {
|
|
286
|
-
const actions = ledger.filter((action) => action.phase === name && !
|
|
289
|
+
const actions = ledger.filter((action) => action.phase === name && !isNonPhaseAction(action));
|
|
287
290
|
const effectiveStatuses = actions.map((action) => effectiveActionStatus(action, state));
|
|
288
291
|
const completed = effectiveStatuses.filter((status) => TERMINAL_ACTIONS.has(status)).length;
|
|
289
292
|
const failed = effectiveStatuses.filter((status) => String(status).startsWith('failed')).length;
|
|
@@ -322,6 +325,11 @@ export function workflowPanelModel(row, {
|
|
|
322
325
|
}
|
|
323
326
|
for (const [key, active] of Object.entries(state.activeAgents ?? {})) {
|
|
324
327
|
if (representedActiveKeys.has(key)) continue;
|
|
328
|
+
// The autonomous orchestrator is a control-plane thread, not a worker in
|
|
329
|
+
// whichever execution phase happens to be selected. It has its own panel.
|
|
330
|
+
// Without this guard, an active checkpoint is re-added below a completed
|
|
331
|
+
// phase when both share the durable `autonomous-delivery` phase name.
|
|
332
|
+
if (orchestrator.autonomous && active.stepId === orchestrator.actionId) continue;
|
|
325
333
|
const action = ledger.find((entry) => entry.id === active.stepId || active.stepId?.startsWith(`${entry.id}[`));
|
|
326
334
|
if ((action?.phase ?? state.currentPhase?.name) !== selectedPhase.name) continue;
|
|
327
335
|
agents.push({
|
|
@@ -355,10 +363,12 @@ export function renderWorkflowTui(row, {
|
|
|
355
363
|
width = 120, height = 36, focus = 0, phaseIndex = null, agentIndex = null,
|
|
356
364
|
detailScroll = 0, message = null, confirmCancel = false,
|
|
357
365
|
controlSelected = false, orchestratorDetail = false, orchestratorVerbose = false,
|
|
366
|
+
workflowVerbose = false,
|
|
358
367
|
spinnerFrame = 0,
|
|
359
368
|
} = {}) {
|
|
360
|
-
width = Math.max(
|
|
369
|
+
width = Math.max(20, Number(width) || 120);
|
|
361
370
|
height = Math.max(18, Number(height) || 36);
|
|
371
|
+
const narrow = width < 100;
|
|
362
372
|
const model = workflowPanelModel(row, { phaseIndex, agentIndex });
|
|
363
373
|
const state = model.state;
|
|
364
374
|
const status = row?.status ?? state.status ?? 'starting';
|
|
@@ -372,17 +382,21 @@ export function renderWorkflowTui(row, {
|
|
|
372
382
|
const terminalLabel = state.finishedAt ? ` · ${status === 'completed' ? 'done' : status}` : '';
|
|
373
383
|
const runName = state.workflow ?? row?.shortId ?? state.shortId ?? row?.runId ?? 'workflow';
|
|
374
384
|
const header = [
|
|
375
|
-
` ${truncate(runName, Math.max(1, width - agentProgress.length - elapsed.length - terminalLabel.length - 5))} · ${agentProgress}${elapsed}${terminalLabel}`,
|
|
385
|
+
truncate(` ${truncate(runName, Math.max(1, width - agentProgress.length - elapsed.length - terminalLabel.length - 5))} · ${agentProgress}${elapsed}${terminalLabel}`, width),
|
|
376
386
|
` ${truncate(state.intent?.goal ?? state.workflow ?? 'workflow', width - 2)}`,
|
|
377
387
|
];
|
|
378
388
|
const footer = confirmCancel
|
|
379
389
|
? ' Stop this workflow? y confirm · n/Esc keep running'
|
|
380
390
|
: orchestratorDetail
|
|
381
391
|
? ` ↑/↓ scroll · v ${orchestratorVerbose ? 'overview' : 'technical details'} · Esc back · c stop · q detach`
|
|
382
|
-
:
|
|
392
|
+
: workflowVerbose
|
|
393
|
+
? ' ↑/↓ scroll · v overview · Esc back · c stop · q detach'
|
|
394
|
+
: narrow
|
|
395
|
+
? ' ↑/↓ select · Enter inspect · Esc back · o planner · v technical · c stop · q detach'
|
|
396
|
+
: ' ↑/↓ select · PgUp/PgDn timeline · Enter inspect · ←/→ switch · v technical · q detach';
|
|
383
397
|
const rawMessageLine = message
|
|
384
398
|
? ` ${truncate(message, width - 2)}`
|
|
385
|
-
: ` ${orchestratorDetail ? `
|
|
399
|
+
: ` ${orchestratorDetail ? `Workflow Planner ${orchestratorVerbose ? 'technical details' : 'overview'}` : workflowVerbose ? 'Workflow technical details' : focus === 0 ? (narrow ? 'Phases' : 'Timeline · auto-following newest event') : focus === 1 ? 'Agents' : 'Agent activity'} · r refresh · workflow continues after detach`;
|
|
386
400
|
const messageLine = truncate(rawMessageLine, width);
|
|
387
401
|
const bodyHeight = Math.max(10, height - header.length - 3);
|
|
388
402
|
|
|
@@ -414,7 +428,6 @@ export function renderWorkflowTui(row, {
|
|
|
414
428
|
|
|
415
429
|
// Two information-rich panes become counterproductive on typical 80-column
|
|
416
430
|
// SSH/mobile terminals. Keep the drill-down full-width below 100 columns.
|
|
417
|
-
const narrow = width < 100;
|
|
418
431
|
const leftWidth = narrow ? width : Math.max(24, Math.min(34, Math.floor(width * 0.27)));
|
|
419
432
|
const rightWidth = narrow ? width : width - leftWidth;
|
|
420
433
|
const orchestrationLines = orchestratorDetailLines(
|
|
@@ -426,8 +439,10 @@ export function renderWorkflowTui(row, {
|
|
|
426
439
|
const detail = orchestratorDetail
|
|
427
440
|
? orchestrationLines
|
|
428
441
|
: agentDetailLines(model, Math.max(20, rightWidth - 4), spinnerFrame);
|
|
442
|
+
const technical = workflowTechnicalLines(model, Math.max(20, rightWidth - 4));
|
|
429
443
|
const contentHeight = bodyHeight - 2;
|
|
430
|
-
const
|
|
444
|
+
const scrollSource = workflowVerbose ? technical : detail;
|
|
445
|
+
const maxScroll = Math.max(0, scrollSource.length - contentHeight);
|
|
431
446
|
const scroll = clamp(detailScroll, 0, maxScroll);
|
|
432
447
|
const visiblePhases = panelWindow(['', ...phaseLines], model.phaseIndex, 1, contentHeight).slice(1);
|
|
433
448
|
const visibleAgents = panelWindow(['', ...agentLines], model.agentIndex, 1, contentHeight).slice(1);
|
|
@@ -436,7 +451,7 @@ export function renderWorkflowTui(row, {
|
|
|
436
451
|
const orchestrationNavLines = model.orchestrator.autonomous
|
|
437
452
|
? [
|
|
438
453
|
selectLine(
|
|
439
|
-
`${statusIcon(model.orchestrator.active ? 'running' : model.orchestrator.status, spinnerFrame)}
|
|
454
|
+
`${statusIcon(model.orchestrator.active ? 'running' : model.orchestrator.status, spinnerFrame)} ${plannerDisplayStatus(model)}`,
|
|
440
455
|
controlSelected,
|
|
441
456
|
focus === 0 && !orchestratorDetail,
|
|
442
457
|
leftWidth - 2,
|
|
@@ -445,6 +460,7 @@ export function renderWorkflowTui(row, {
|
|
|
445
460
|
[model.orchestrator.pool, model.orchestrator.model].filter(Boolean).join(' · ') || 'select to inspect',
|
|
446
461
|
leftWidth - 2,
|
|
447
462
|
),
|
|
463
|
+
dimLine(plannerUsageSummary(model), leftWidth - 2),
|
|
448
464
|
]
|
|
449
465
|
: [];
|
|
450
466
|
const narrowWorkflowLines = model.orchestrator.autonomous
|
|
@@ -457,7 +473,19 @@ export function renderWorkflowTui(row, {
|
|
|
457
473
|
|
|
458
474
|
let body;
|
|
459
475
|
if (orchestratorDetail) {
|
|
460
|
-
body = renderPanel(`
|
|
476
|
+
body = renderPanel(`Workflow Planner · ${orchestratorVerbose ? 'technical details' : 'overview'}`, visibleDetail, width, bodyHeight);
|
|
477
|
+
} else if (workflowVerbose) {
|
|
478
|
+
const visibleTechnical = technical.slice(scroll, scroll + contentHeight);
|
|
479
|
+
if (narrow) body = renderPanel('Workflow technical details', visibleTechnical, width, bodyHeight);
|
|
480
|
+
else {
|
|
481
|
+
const left = model.orchestrator.autonomous
|
|
482
|
+
? [
|
|
483
|
+
...renderPanel('Workflow Planner', orchestrationNavLines, leftWidth, 5),
|
|
484
|
+
...renderPanel(phaseTitle, visiblePhases, leftWidth, bodyHeight - 5),
|
|
485
|
+
]
|
|
486
|
+
: renderPanel(phaseTitle, visiblePhases, leftWidth, bodyHeight);
|
|
487
|
+
body = joinPanels(left, renderPanel('Workflow technical details', visibleTechnical, rightWidth, bodyHeight));
|
|
488
|
+
}
|
|
461
489
|
} else if (narrow) {
|
|
462
490
|
const mobile = focus === 0
|
|
463
491
|
? {
|
|
@@ -473,15 +501,17 @@ export function renderWorkflowTui(row, {
|
|
|
473
501
|
} else if (focus < 2) {
|
|
474
502
|
const left = model.orchestrator.autonomous
|
|
475
503
|
? [
|
|
476
|
-
...renderPanel('
|
|
504
|
+
...renderPanel('Workflow Planner', orchestrationNavLines, leftWidth, 5),
|
|
477
505
|
...renderPanel(phaseTitle, visiblePhases, leftWidth, bodyHeight - 5),
|
|
478
506
|
]
|
|
479
507
|
: renderPanel(phaseTitle, visiblePhases, leftWidth, bodyHeight);
|
|
480
508
|
body = joinPanels(
|
|
481
509
|
left,
|
|
482
510
|
controlSelected
|
|
483
|
-
? renderPanel(`
|
|
484
|
-
:
|
|
511
|
+
? renderPanel(`Workflow Planner · ${model.orchestrator.status}`, orchestrationLines.slice(0, contentHeight), rightWidth, bodyHeight)
|
|
512
|
+
: focus === 0
|
|
513
|
+
? renderWorkflowOverviewPanel(model, rightWidth, bodyHeight, spinnerFrame, detailScroll)
|
|
514
|
+
: renderPanel(agentTitle, visibleAgents, rightWidth, bodyHeight),
|
|
485
515
|
);
|
|
486
516
|
} else {
|
|
487
517
|
body = joinPanels(
|
|
@@ -508,6 +538,349 @@ function joinPanels(left, right) {
|
|
|
508
538
|
return left.map((line, index) => `${line}${right[index] ?? ''}`);
|
|
509
539
|
}
|
|
510
540
|
|
|
541
|
+
function renderWorkflowOverviewPanel(model, width, height, spinnerFrame, timelineScroll = 0) {
|
|
542
|
+
const inner = Math.max(1, width - 2);
|
|
543
|
+
const timeline = workflowTimelineLines(model, inner);
|
|
544
|
+
const live = workflowLiveLines(model, inner, spinnerFrame);
|
|
545
|
+
const next = workflowNextLines(model, inner);
|
|
546
|
+
const contentRows = Math.max(3, height - 4); // outer border + two section dividers
|
|
547
|
+
const nextRows = Math.min(next.length, 2);
|
|
548
|
+
const liveRows = Math.min(live.lines.length, Math.max(2, Math.floor(contentRows * 0.42)));
|
|
549
|
+
const timelineRows = Math.max(1, contentRows - liveRows - nextRows);
|
|
550
|
+
const maxTimelineScroll = Math.max(0, timeline.lines.length - timelineRows);
|
|
551
|
+
const scroll = clamp(timelineScroll, 0, maxTimelineScroll);
|
|
552
|
+
const start = Math.max(0, timeline.lines.length - timelineRows - scroll);
|
|
553
|
+
let visibleTimeline = timeline.lines.slice(start, start + timelineRows);
|
|
554
|
+
if (start > 0 && visibleTimeline.length) {
|
|
555
|
+
visibleTimeline[0] = dimText(`↑ ${start + 1} earlier timeline rows`, inner);
|
|
556
|
+
}
|
|
557
|
+
if (start + timelineRows < timeline.lines.length && visibleTimeline.length) {
|
|
558
|
+
visibleTimeline[visibleTimeline.length - 1] = dimText(`↓ ${timeline.lines.length - start - timelineRows + 1} newer timeline rows`, inner);
|
|
559
|
+
}
|
|
560
|
+
const visibleLive = live.lines.slice(0, liveRows);
|
|
561
|
+
const visibleNext = next.slice(0, nextRows);
|
|
562
|
+
const title = ` Workflow timeline · ${timeline.milestoneCount} milestone${timeline.milestoneCount === 1 ? '' : 's'} `;
|
|
563
|
+
const rows = [`┌${truncate(title, inner)}${'─'.repeat(Math.max(0, inner - truncate(title, inner).length))}┐`];
|
|
564
|
+
for (const line of visibleTimeline) rows.push(`│${panelCell(line, inner)}│`);
|
|
565
|
+
while (rows.length < 1 + timelineRows) rows.push(`│${panelCell('', inner)}│`);
|
|
566
|
+
rows.push(sectionDivider(`Live · ${live.running} running · ${live.waiting} waiting`, inner));
|
|
567
|
+
for (const line of visibleLive) rows.push(`│${panelCell(line, inner)}│`);
|
|
568
|
+
while (rows.length < 2 + timelineRows + liveRows) rows.push(`│${panelCell('', inner)}│`);
|
|
569
|
+
rows.push(sectionDivider('Next', inner));
|
|
570
|
+
for (const line of visibleNext) rows.push(`│${panelCell(line, inner)}│`);
|
|
571
|
+
while (rows.length < height - 1) rows.push(`│${panelCell('', inner)}│`);
|
|
572
|
+
rows.push(`└${'─'.repeat(inner)}┘`);
|
|
573
|
+
return rows.slice(0, height);
|
|
574
|
+
}
|
|
575
|
+
|
|
576
|
+
function sectionDivider(label, inner) {
|
|
577
|
+
const text = truncate(` ${label} `, inner);
|
|
578
|
+
return `├${text}${'─'.repeat(Math.max(0, inner - text.length))}┤`;
|
|
579
|
+
}
|
|
580
|
+
|
|
581
|
+
function workflowTimelineLines(model, width) {
|
|
582
|
+
const { state, orchestrator } = model;
|
|
583
|
+
const ledger = state.actionLedger ?? [];
|
|
584
|
+
const events = [];
|
|
585
|
+
const add = (at, lines, sequence = Number.MAX_SAFE_INTEGER) => {
|
|
586
|
+
if (!at) return;
|
|
587
|
+
events.push({ at, sequence, lines: Array.isArray(lines) ? lines : [lines] });
|
|
588
|
+
};
|
|
589
|
+
const scout = ledger.find((action) => action.id === 'scout');
|
|
590
|
+
add(state.startedAt, [
|
|
591
|
+
timelineRow(state.startedAt, '● Workflow initiated', '', width),
|
|
592
|
+
timelineDetail(scout ? 'Goal accepted; preparing repository reconnaissance' : 'Execution started', width),
|
|
593
|
+
]);
|
|
594
|
+
|
|
595
|
+
const scoutStartedAt = actionStartedAt(state, scout);
|
|
596
|
+
const scoutFinishedAt = actionFinishedAt(state, scout);
|
|
597
|
+
if (scoutStartedAt) {
|
|
598
|
+
add(scoutStartedAt, [
|
|
599
|
+
timelineRow(scoutStartedAt, '● [Preflight: Scout] started', '', width),
|
|
600
|
+
timelineDetail('Read-only repository and capability inspection', width),
|
|
601
|
+
]);
|
|
602
|
+
}
|
|
603
|
+
if (scoutFinishedAt && TERMINAL_ACTIONS.has(effectiveActionStatus(scout, state))) {
|
|
604
|
+
const attempt = latestAttemptForAction(state, scout);
|
|
605
|
+
const metadata = [attempt?.pool, attempt?.model, tokenText(attempt?.usage)].filter(Boolean).join(' · ');
|
|
606
|
+
add(scoutFinishedAt, [
|
|
607
|
+
timelineRow(scoutFinishedAt, `${statusIcon(effectiveActionStatus(scout, state))} [Preflight: Scout] completed`, durationText(scoutStartedAt, scoutFinishedAt), width),
|
|
608
|
+
...(metadata ? [timelineDetail(metadata, width)] : []),
|
|
609
|
+
]);
|
|
610
|
+
}
|
|
611
|
+
|
|
612
|
+
orchestrator.attempts.forEach((attempt, index) => {
|
|
613
|
+
if (!attempt.finishedAt || !TERMINAL_ACTIONS.has(attempt.status)) return;
|
|
614
|
+
const decision = decisionForPlannerAttempt(state, attempt, index, orchestrator.attempts);
|
|
615
|
+
const summary = decision?.reason ? sentencePreview(decision.reason, Math.max(30, width - 10))
|
|
616
|
+
: decision ? decisionLabel(decision.decision) : 'No accepted decision; correction or retry turn';
|
|
617
|
+
add(attempt.finishedAt, [
|
|
618
|
+
timelineRow(attempt.finishedAt, `◆ [Workflow Planner] checkpoint #${index + 1}`, durationText(attempt.startedAt, attempt.finishedAt), width),
|
|
619
|
+
timelineDetail(summary, width),
|
|
620
|
+
]);
|
|
621
|
+
});
|
|
622
|
+
|
|
623
|
+
const phases = new Map();
|
|
624
|
+
for (const action of ledger) {
|
|
625
|
+
if (action.id === 'scout' || (orchestrator.autonomous && action.id === orchestrator.actionId && action.kind === 'decide')) continue;
|
|
626
|
+
if (!action.phase) continue;
|
|
627
|
+
if (!phases.has(action.phase)) phases.set(action.phase, []);
|
|
628
|
+
phases.get(action.phase).push(action);
|
|
629
|
+
}
|
|
630
|
+
for (const [name, actions] of phases) {
|
|
631
|
+
const realStart = earliestTimestamp(actions.map((action) => actionStartedAt(state, action)));
|
|
632
|
+
const startedAt = realStart ?? earliestTimestamp(actions.map((action) => actionFinishedAt(state, action)));
|
|
633
|
+
if (!startedAt) continue;
|
|
634
|
+
const label = phaseLabel(name, orchestrator);
|
|
635
|
+
add(startedAt, timelineRow(startedAt, `├─ [Phase: ${label}] ${realStart ? 'started' : 'blocked'}`, '', width));
|
|
636
|
+
const finished = actions
|
|
637
|
+
.filter((action) => actionFinishedAt(state, action) && TERMINAL_ACTIONS.has(effectiveActionStatus(action, state)))
|
|
638
|
+
.sort((a, b) => Date.parse(actionFinishedAt(state, a)) - Date.parse(actionFinishedAt(state, b)));
|
|
639
|
+
finished.forEach((action, index) => {
|
|
640
|
+
const terminalPhase = finished.length === actions.length && index === finished.length - 1;
|
|
641
|
+
const branch = terminalPhase ? '│ └─' : '│ ├─';
|
|
642
|
+
const actionFinished = actionFinishedAt(state, action);
|
|
643
|
+
const actionStarted = actionStartedAt(state, action);
|
|
644
|
+
add(actionFinished, timelineRow(
|
|
645
|
+
actionFinished,
|
|
646
|
+
`${branch}${statusIcon(effectiveActionStatus(action, state))} [${label}] ${action.id}`,
|
|
647
|
+
actionStarted ? durationText(actionStarted, actionFinished) : '',
|
|
648
|
+
width,
|
|
649
|
+
));
|
|
650
|
+
});
|
|
651
|
+
if (finished.length === actions.length && actions.length) {
|
|
652
|
+
const finishedAt = latestTimestamp(actions.map((action) => actionFinishedAt(state, action)));
|
|
653
|
+
const failed = actions.some((action) => String(effectiveActionStatus(action, state)).startsWith('failed'));
|
|
654
|
+
add(finishedAt, timelineRow(finishedAt, `└─${failed ? '✗' : '✓'} [Phase: ${label}] completed`, `${finished.length}/${actions.length}`, width));
|
|
655
|
+
}
|
|
656
|
+
}
|
|
657
|
+
|
|
658
|
+
for (const event of model.events) {
|
|
659
|
+
const detail = timelineControlEvent(event, width);
|
|
660
|
+
if (detail) add(event.committedAt, detail, Number(event.sequence));
|
|
661
|
+
}
|
|
662
|
+
|
|
663
|
+
events.sort((a, b) => Date.parse(a.at) - Date.parse(b.at) || a.sequence - b.sequence);
|
|
664
|
+
const lines = [];
|
|
665
|
+
events.forEach((event, index) => {
|
|
666
|
+
if (index) lines.push('');
|
|
667
|
+
lines.push(...event.lines);
|
|
668
|
+
});
|
|
669
|
+
return { lines: lines.length ? lines : ['Waiting for the first durable workflow milestone'], milestoneCount: events.length };
|
|
670
|
+
}
|
|
671
|
+
|
|
672
|
+
function decisionForPlannerAttempt(state, attempt, index, attempts) {
|
|
673
|
+
const decisions = state.decisions ?? [];
|
|
674
|
+
if (attempt?.outFile) {
|
|
675
|
+
const artifactMatch = decisions.find((decision) => decision.artifact === attempt.outFile);
|
|
676
|
+
if (artifactMatch) return artifactMatch;
|
|
677
|
+
}
|
|
678
|
+
const started = Date.parse(attempt?.startedAt ?? '');
|
|
679
|
+
const finished = Date.parse(attempt?.finishedAt ?? '');
|
|
680
|
+
if (Number.isFinite(started) && Number.isFinite(finished)) {
|
|
681
|
+
const timeMatch = decisions.find((decision) => {
|
|
682
|
+
const created = Date.parse(decision.createdAt ?? '');
|
|
683
|
+
return Number.isFinite(created) && created >= started && created <= finished + 2_000;
|
|
684
|
+
});
|
|
685
|
+
if (timeMatch) return timeMatch;
|
|
686
|
+
}
|
|
687
|
+
const hasDurableCorrelation = decisions.some((decision) => decision.artifact || decision.createdAt)
|
|
688
|
+
|| attempts.some((entry) => entry.outFile);
|
|
689
|
+
return hasDurableCorrelation ? null : decisions[index];
|
|
690
|
+
}
|
|
691
|
+
|
|
692
|
+
function timelineControlEvent(event, width) {
|
|
693
|
+
const labels = {
|
|
694
|
+
'decision.rejected': '✗ [Workflow Planner] decision rejected',
|
|
695
|
+
'decision.correction_requested': '⧖ [Workflow Planner] correction requested',
|
|
696
|
+
'decision.orchestrator_escalated': '◆ [Workflow Planner] provider escalated',
|
|
697
|
+
'run.cancellation_requested': '⧖ Workflow cancellation requested',
|
|
698
|
+
'run.cancelling': '⧖ Workflow cancellation requested',
|
|
699
|
+
'run.interruption_requested': '⧖ Workflow interruption requested',
|
|
700
|
+
'workflow.expansion_target_exceeded': '! Advisory expansion target exceeded',
|
|
701
|
+
'workflow.agent_target_exceeded': '! Advisory agent target exceeded',
|
|
702
|
+
};
|
|
703
|
+
const label = labels[event.type];
|
|
704
|
+
if (!label) return null;
|
|
705
|
+
const reason = event.payload?.why ?? event.payload?.reason;
|
|
706
|
+
return [
|
|
707
|
+
timelineRow(event.committedAt, label, '', width),
|
|
708
|
+
...(reason ? [timelineDetail(sentencePreview(reason, Math.max(20, width - 8)), width)] : []),
|
|
709
|
+
];
|
|
710
|
+
}
|
|
711
|
+
|
|
712
|
+
function workflowLiveLines(model, width, spinnerFrame) {
|
|
713
|
+
const { state, orchestrator } = model;
|
|
714
|
+
const activeWorkers = Object.values(state.activeAgents ?? {})
|
|
715
|
+
.filter((agent) => agent.stepId !== orchestrator.actionId)
|
|
716
|
+
.sort((a, b) => String(b.lastEventAt ?? b.lastActivityAt ?? '').localeCompare(String(a.lastEventAt ?? a.lastActivityAt ?? '')));
|
|
717
|
+
const lines = [];
|
|
718
|
+
let running = activeWorkers.length + (orchestrator.active ? 1 : 0);
|
|
719
|
+
let waiting = 0;
|
|
720
|
+
if (orchestrator.autonomous && !state.finishedAt) {
|
|
721
|
+
const plannerWaiting = !orchestrator.active && activeWorkers.length > 0;
|
|
722
|
+
if (plannerWaiting) waiting += 1;
|
|
723
|
+
const plannerStatus = orchestrator.active ? 'planning' : plannerWaiting ? 'waiting' : orchestrator.status;
|
|
724
|
+
lines.push(alignRight(
|
|
725
|
+
`${statusIcon(plannerStatus, spinnerFrame)} [Workflow Planner] · ${orchestrator.pool} · ${orchestrator.model}`,
|
|
726
|
+
plannerStatus,
|
|
727
|
+
width,
|
|
728
|
+
));
|
|
729
|
+
if (plannerWaiting) lines.push(` Waiting for ${activeWorkers.length === 1 ? activeWorkers[0].stepId : `${activeWorkers.length} workers`}`);
|
|
730
|
+
else if (orchestrator.active) lines.push(' Choosing the next smallest useful action');
|
|
731
|
+
else lines.push(` ${humanStatus(orchestrator.status)}`);
|
|
732
|
+
const plannerAction = orchestrator.active?.lastActions?.at(-1) ?? orchestrator.latestAttempt?.lastActions?.at(-1);
|
|
733
|
+
if (plannerAction) lines.push(` ↳ ${friendlyActionKind(plannerAction.kind)}${plannerAction.summary ? ` · ${friendlyActionSummary(plannerAction)}` : ''}`);
|
|
734
|
+
else if (orchestrator.latestDecision) lines.push(` ↳ Decision · ${decisionLabel(orchestrator.latestDecision.decision)}`);
|
|
735
|
+
const plannerStream = streamActivityLine(orchestrator.active);
|
|
736
|
+
if (plannerStream) lines.push(` ${plannerStream}`);
|
|
737
|
+
lines.push('');
|
|
738
|
+
}
|
|
739
|
+
for (const agent of activeWorkers) {
|
|
740
|
+
const action = (state.actionLedger ?? []).find((entry) => entry.id === agent.stepId || agent.stepId?.startsWith(`${entry.id}[`));
|
|
741
|
+
lines.push(alignRight(
|
|
742
|
+
`${statusIcon(agent.status ?? 'running', spinnerFrame)} ${agent.stepId} · ${agent.pool ?? 'unassigned'} · ${agent.model ?? 'connector model'}`,
|
|
743
|
+
durationText(action?.startedAt ?? agent.startedAt),
|
|
744
|
+
width,
|
|
745
|
+
));
|
|
746
|
+
const latest = agent.lastActions?.at(-1);
|
|
747
|
+
lines.push(latest
|
|
748
|
+
? ` ↳ ${friendlyActionKind(latest.kind)}${latest.summary ? ` · ${friendlyActionSummary(latest)}` : ''}`
|
|
749
|
+
: ' ↳ waiting for the first semantic action event');
|
|
750
|
+
const stream = streamActivityLine(agent);
|
|
751
|
+
if (stream) lines.push(` ${stream}`);
|
|
752
|
+
lines.push('');
|
|
753
|
+
}
|
|
754
|
+
if (!lines.length) lines.push(state.finishedAt ? '✓ No live agents · workflow is terminal' : '⧖ Waiting for the next dispatch');
|
|
755
|
+
return { lines, running, waiting };
|
|
756
|
+
}
|
|
757
|
+
|
|
758
|
+
function workflowNextLines(model, width) {
|
|
759
|
+
const { state, orchestrator } = model;
|
|
760
|
+
if (state.finishedAt) return [truncate(`${statusIcon(state.status)} Workflow terminal · obtain the stable result envelope`, width)];
|
|
761
|
+
const ledger = state.actionLedger ?? [];
|
|
762
|
+
const pending = ledger.find((action) => action.id !== 'scout'
|
|
763
|
+
&& action.id !== orchestrator.actionId
|
|
764
|
+
&& !TERMINAL_ACTIONS.has(effectiveActionStatus(action, state))
|
|
765
|
+
&& effectiveActionStatus(action, state) !== 'running');
|
|
766
|
+
if (pending) {
|
|
767
|
+
const blockers = (pending.dependsOn ?? []).filter((id) => state.outputs?.[id]?.ok !== true);
|
|
768
|
+
const wait = blockers.length ? ` · waiting for ${blockers.join(', ')}` : '';
|
|
769
|
+
return [truncate(`○ [Phase: ${phaseLabel(pending.phase, orchestrator)}] · ${pending.id}${wait}`, width)];
|
|
770
|
+
}
|
|
771
|
+
if (Object.keys(state.activeAgents ?? {}).some((key) => state.activeAgents[key]?.stepId !== orchestrator.actionId)) {
|
|
772
|
+
return ['○ [Workflow Planner] will reassess when current work finishes'];
|
|
773
|
+
}
|
|
774
|
+
if (orchestrator.active) return ['○ Awaiting the next [Workflow Planner] decision'];
|
|
775
|
+
return ['○ Awaiting the next [Workflow Planner] decision'];
|
|
776
|
+
}
|
|
777
|
+
|
|
778
|
+
function workflowTechnicalLines(model, width) {
|
|
779
|
+
const { state, orchestrator } = model;
|
|
780
|
+
const lines = [
|
|
781
|
+
`Status · ${state.status ?? 'starting'}`,
|
|
782
|
+
`Current · ${state.currentStep?.id ?? '—'} · ${state.currentStep?.phase ?? state.currentPhase?.name ?? '—'}`,
|
|
783
|
+
`Started · ${state.startedAt ?? '—'}`,
|
|
784
|
+
`Usage · ${compactUsage(state.usage)}`,
|
|
785
|
+
'',
|
|
786
|
+
'Action ledger',
|
|
787
|
+
];
|
|
788
|
+
for (const action of state.actionLedger ?? []) {
|
|
789
|
+
const control = orchestrator.autonomous && action.id === orchestrator.actionId && action.kind === 'decide';
|
|
790
|
+
lines.push(`${statusIcon(effectiveActionStatus(action, state))} ${control ? '[Workflow Planner]' : action.id} · ${action.kind} · ${action.status ?? 'pending'} · ${action.phase ?? '—'}`);
|
|
791
|
+
}
|
|
792
|
+
if (!(state.actionLedger ?? []).length) lines.push('· no actions recorded');
|
|
793
|
+
lines.push('', 'Recent durable events');
|
|
794
|
+
for (const event of model.events.slice(-12)) lines.push(`#${event.sequence} ${event.type}`);
|
|
795
|
+
if (!model.events.length) lines.push('· no events recorded');
|
|
796
|
+
return wrapLines(lines, width);
|
|
797
|
+
}
|
|
798
|
+
|
|
799
|
+
function timelineRow(at, label, right, width) {
|
|
800
|
+
return alignRight(`${clockText(at)} ${label}`, right, width);
|
|
801
|
+
}
|
|
802
|
+
|
|
803
|
+
function timelineDetail(text, width) {
|
|
804
|
+
return truncate(` ${text}`, width);
|
|
805
|
+
}
|
|
806
|
+
|
|
807
|
+
function alignRight(left, right, width) {
|
|
808
|
+
const suffix = right ? String(right) : '';
|
|
809
|
+
if (!suffix) return truncate(left, width);
|
|
810
|
+
const room = Math.max(1, width - suffix.length - 1);
|
|
811
|
+
const lhs = truncate(left, room);
|
|
812
|
+
return `${lhs}${' '.repeat(Math.max(1, width - lhs.length - suffix.length))}${suffix}`;
|
|
813
|
+
}
|
|
814
|
+
|
|
815
|
+
function clockText(value) {
|
|
816
|
+
const date = new Date(value ?? '');
|
|
817
|
+
if (!Number.isFinite(date.getTime())) return '--:--';
|
|
818
|
+
return `${String(date.getHours()).padStart(2, '0')}:${String(date.getMinutes()).padStart(2, '0')}`;
|
|
819
|
+
}
|
|
820
|
+
|
|
821
|
+
function earliestTimestamp(values) {
|
|
822
|
+
return values.filter(Boolean).sort((a, b) => Date.parse(a) - Date.parse(b))[0] ?? null;
|
|
823
|
+
}
|
|
824
|
+
|
|
825
|
+
function latestTimestamp(values) {
|
|
826
|
+
return values.filter(Boolean).sort((a, b) => Date.parse(b) - Date.parse(a))[0] ?? null;
|
|
827
|
+
}
|
|
828
|
+
|
|
829
|
+
function latestAttemptForAction(state, action) {
|
|
830
|
+
if (!action) return null;
|
|
831
|
+
return (action.attempts ?? []).map((index) => state.attempts?.[index]).filter(Boolean).at(-1) ?? null;
|
|
832
|
+
}
|
|
833
|
+
|
|
834
|
+
function actionStartedAt(state, action) {
|
|
835
|
+
if (!action) return null;
|
|
836
|
+
const attempts = (action.attempts ?? []).map((index) => state.attempts?.[index]).filter(Boolean);
|
|
837
|
+
return action.startedAt ?? earliestTimestamp(attempts.map((attempt) => attempt.startedAt));
|
|
838
|
+
}
|
|
839
|
+
|
|
840
|
+
function actionFinishedAt(state, action) {
|
|
841
|
+
if (!action) return null;
|
|
842
|
+
const attempts = (action.attempts ?? []).map((index) => state.attempts?.[index]).filter(Boolean);
|
|
843
|
+
return action.finishedAt ?? latestTimestamp(attempts.map((attempt) => attempt.finishedAt));
|
|
844
|
+
}
|
|
845
|
+
|
|
846
|
+
function streamActivityLine(agent) {
|
|
847
|
+
if (!agent) return '';
|
|
848
|
+
const at = agent.lastEventAt ?? agent.lastActivityAt;
|
|
849
|
+
if (!at) return 'stream waiting for the first provider event';
|
|
850
|
+
return `stream active ${durationText(at)} ago · ${formatBytes(agent.outputBytesObserved ?? 0)} observed`;
|
|
851
|
+
}
|
|
852
|
+
|
|
853
|
+
function formatBytes(value) {
|
|
854
|
+
const bytes = Math.max(0, Number(value) || 0);
|
|
855
|
+
if (bytes < 1024) return `${bytes} B`;
|
|
856
|
+
if (bytes < 1024 * 1024) return `${Math.round(bytes / 1024)} KB`;
|
|
857
|
+
return `${(bytes / (1024 * 1024)).toFixed(1)} MB`;
|
|
858
|
+
}
|
|
859
|
+
|
|
860
|
+
function humanStatus(value) {
|
|
861
|
+
return String(value ?? 'waiting').replaceAll('_', ' ').replace(/^./, (char) => char.toUpperCase());
|
|
862
|
+
}
|
|
863
|
+
|
|
864
|
+
function plannerDisplayStatus(model) {
|
|
865
|
+
const { orchestrator, state } = model;
|
|
866
|
+
if (state.finishedAt) return state.status === 'completed' ? 'Completed' : humanStatus(state.status);
|
|
867
|
+
if (orchestrator.active) return 'Planning next actions';
|
|
868
|
+
const workers = Object.values(state.activeAgents ?? {}).filter((agent) => agent.stepId !== orchestrator.actionId);
|
|
869
|
+
if (workers.length) return 'Waiting for workers';
|
|
870
|
+
if (orchestrator.status === 'reviewing evidence') return 'Reviewing evidence';
|
|
871
|
+
return humanStatus(orchestrator.status);
|
|
872
|
+
}
|
|
873
|
+
|
|
874
|
+
function plannerUsageSummary(model) {
|
|
875
|
+
const checkpoints = model.orchestrator.attempts.length;
|
|
876
|
+
const cost = model.state.usage?.cost?.estimatedUsd ?? model.state.usage?.cost?.knownSubtotalUsd;
|
|
877
|
+
return `Checkpoints ${checkpoints}${Number.isFinite(cost) ? ` · $${cost.toFixed(2)}` : ''}`;
|
|
878
|
+
}
|
|
879
|
+
|
|
880
|
+
function dimText(value, width) {
|
|
881
|
+
return `\x1b[2m${truncate(value, width)}\x1b[0m`;
|
|
882
|
+
}
|
|
883
|
+
|
|
511
884
|
function orchestratorDetailLines(model, width, spinnerFrame, { verbose = false } = {}) {
|
|
512
885
|
const { orchestrator, state } = model;
|
|
513
886
|
if (!orchestrator.autonomous) {
|
|
@@ -520,6 +893,7 @@ function orchestratorDetailLines(model, width, spinnerFrame, { verbose = false }
|
|
|
520
893
|
const conversations = Object.entries(state.orchestration?.conversations ?? {});
|
|
521
894
|
const decisions = state.decisions ?? [];
|
|
522
895
|
const latestDecision = decisions.at(-1);
|
|
896
|
+
const latestLiveAction = liveActions.at(-1) ?? null;
|
|
523
897
|
const workerAttempts = (state.attempts ?? []).filter((attempt) => attempt.actionId !== orchestrator.actionId);
|
|
524
898
|
const completedWorkers = workerAttempts.filter((attempt) => TERMINAL_ACTIONS.has(attempt.status)).length;
|
|
525
899
|
const activeWorkers = Object.values(state.activeAgents ?? {})
|
|
@@ -538,6 +912,14 @@ function orchestratorDetailLines(model, width, spinnerFrame, { verbose = false }
|
|
|
538
912
|
`Now · ${stateLabel}`,
|
|
539
913
|
`Progress · ${completedWorkers}/${workerAttempts.length} worker attempts finished · ${orchestrator.attempts.length} planning checkpoint${orchestrator.attempts.length === 1 ? '' : 's'}`,
|
|
540
914
|
];
|
|
915
|
+
if (active) {
|
|
916
|
+
lines.push(latestLiveAction
|
|
917
|
+
? `Latest action · ${friendlyActionKind(latestLiveAction.kind)}${latestLiveAction.summary ? ` · ${friendlyActionSummary(latestLiveAction)}` : ''}`
|
|
918
|
+
: 'Latest action · waiting for the first semantic event');
|
|
919
|
+
lines.push(active.lastEventAt
|
|
920
|
+
? `Live stream · event ${durationText(active.lastEventAt)} ago · ${active.outputBytesObserved ?? 0} bytes observed`
|
|
921
|
+
: 'Live stream · waiting for the first provider event');
|
|
922
|
+
}
|
|
541
923
|
if (latestDecision) {
|
|
542
924
|
lines.push(`Latest decision · ${decisionLabel(latestDecision.decision)}`);
|
|
543
925
|
if (latestDecision.reason) lines.push(`Why · ${sentencePreview(latestDecision.reason)}`);
|
|
@@ -581,7 +963,7 @@ function orchestratorDetailLines(model, width, spinnerFrame, { verbose = false }
|
|
|
581
963
|
lines.push('', 'Checkpoint turns');
|
|
582
964
|
if (!orchestrator.attempts.length) lines.push('· waiting for the first planning turn');
|
|
583
965
|
orchestrator.attempts.forEach((attempt, index) => {
|
|
584
|
-
const decision = (state
|
|
966
|
+
const decision = decisionForPlannerAttempt(state, attempt, index, orchestrator.attempts);
|
|
585
967
|
lines.push(`#${index + 1} ${statusIcon(attempt.status, spinnerFrame)} ${attempt.status} · ${attempt.pool ?? '—'} · ${attempt.model ?? 'connector model'} · ${durationText(attempt.startedAt, attempt.finishedAt)}`);
|
|
586
968
|
lines.push(` ${compactUsage(attempt.usage)}`);
|
|
587
969
|
if (decision) lines.push(` decision: ${decision.decision} · ${decision.reason}`);
|
|
@@ -845,6 +1227,7 @@ export async function runDashboard(bullswarmDir, {
|
|
|
845
1227
|
controlSelected: false,
|
|
846
1228
|
orchestratorDetail: false,
|
|
847
1229
|
orchestratorVerbose: false,
|
|
1230
|
+
workflowVerbose: false,
|
|
848
1231
|
spinnerFrame: 0,
|
|
849
1232
|
};
|
|
850
1233
|
const paintUnsafe = () => {
|
|
@@ -915,7 +1298,7 @@ export async function runDashboard(bullswarmDir, {
|
|
|
915
1298
|
}
|
|
916
1299
|
const row = detailRow(bullswarmDir, selectedRunId);
|
|
917
1300
|
const model = workflowPanelModel(row, { phaseIndex: ui.phaseIndex, agentIndex: ui.agentIndex });
|
|
918
|
-
if (ui.orchestratorDetail) {
|
|
1301
|
+
if (ui.orchestratorDetail || ui.workflowVerbose) {
|
|
919
1302
|
ui.detailScroll = Math.max(0, ui.detailScroll + delta);
|
|
920
1303
|
return paint();
|
|
921
1304
|
}
|
|
@@ -979,27 +1362,32 @@ export async function runDashboard(bullswarmDir, {
|
|
|
979
1362
|
ui.focus = 0;
|
|
980
1363
|
ui.controlSelected = true;
|
|
981
1364
|
ui.detailScroll = 0;
|
|
1365
|
+
} else if (ui.workflowVerbose) {
|
|
1366
|
+
ui.workflowVerbose = false;
|
|
1367
|
+
ui.detailScroll = 0;
|
|
982
1368
|
} else if (detail && ui.focus > 0) ui.focus -= 1;
|
|
983
1369
|
else if (detail && !token) detail = false;
|
|
984
1370
|
else message = 'At phase level · press q to detach while the workflow keeps running.';
|
|
985
1371
|
return paint();
|
|
986
1372
|
}
|
|
987
1373
|
if (key === '\t' || key === '\u001b[C' || key === 'l') {
|
|
988
|
-
if (detail) ui.focus = (ui.focus + 1) % 3;
|
|
1374
|
+
if (detail) { ui.focus = (ui.focus + 1) % 3; ui.detailScroll = 0; }
|
|
989
1375
|
return paint();
|
|
990
1376
|
}
|
|
991
1377
|
if (key === '\u001b[D' || key === 'h') {
|
|
992
|
-
if (detail) ui.focus = (ui.focus + 2) % 3;
|
|
1378
|
+
if (detail) { ui.focus = (ui.focus + 2) % 3; ui.detailScroll = 0; }
|
|
993
1379
|
return paint();
|
|
994
1380
|
}
|
|
995
1381
|
if (key === '\u001b[A' || key === 'k') return moveVertical(-1);
|
|
996
1382
|
if (key === '\u001b[B' || key === 'j') return moveVertical(1);
|
|
997
|
-
|
|
998
|
-
if (key === '\u001b[
|
|
1383
|
+
const timelineScroll = detail && ui.focus === 0 && !ui.orchestratorDetail && !ui.workflowVerbose;
|
|
1384
|
+
if (key === '\u001b[5~') { ui.detailScroll = timelineScroll ? ui.detailScroll + 8 : Math.max(0, ui.detailScroll - 8); return paint(); }
|
|
1385
|
+
if (key === '\u001b[6~') { ui.detailScroll = timelineScroll ? Math.max(0, ui.detailScroll - 8) : ui.detailScroll + 8; return paint(); }
|
|
999
1386
|
if (key === 'o' && detail) {
|
|
1000
1387
|
const row = detailRow(bullswarmDir, selectedRunId);
|
|
1001
1388
|
const model = workflowPanelModel(row, { phaseIndex: ui.phaseIndex, agentIndex: ui.agentIndex });
|
|
1002
1389
|
if (model.orchestrator.autonomous) {
|
|
1390
|
+
ui.workflowVerbose = false;
|
|
1003
1391
|
ui.orchestratorDetail = true;
|
|
1004
1392
|
ui.orchestratorVerbose = false;
|
|
1005
1393
|
ui.controlSelected = true;
|
|
@@ -1011,8 +1399,9 @@ export async function runDashboard(bullswarmDir, {
|
|
|
1011
1399
|
if (key === '1' && detail) { ui.focus = 0; return paint(); }
|
|
1012
1400
|
if (key === '2' && detail) { ui.focus = 1; return paint(); }
|
|
1013
1401
|
if (key === '3' && detail) { ui.focus = 2; return paint(); }
|
|
1014
|
-
if (key === 'v' &&
|
|
1015
|
-
ui.orchestratorVerbose = !ui.orchestratorVerbose;
|
|
1402
|
+
if (key === 'v' && detail) {
|
|
1403
|
+
if (ui.orchestratorDetail) ui.orchestratorVerbose = !ui.orchestratorVerbose;
|
|
1404
|
+
else ui.workflowVerbose = !ui.workflowVerbose;
|
|
1016
1405
|
ui.detailScroll = 0;
|
|
1017
1406
|
message = null;
|
|
1018
1407
|
return paint();
|
|
@@ -1025,6 +1414,7 @@ export async function runDashboard(bullswarmDir, {
|
|
|
1025
1414
|
ui.followActiveAgent = true;
|
|
1026
1415
|
} else if (ui.focus === 0) {
|
|
1027
1416
|
if (ui.controlSelected) {
|
|
1417
|
+
ui.workflowVerbose = false;
|
|
1028
1418
|
ui.orchestratorDetail = true;
|
|
1029
1419
|
ui.orchestratorVerbose = false;
|
|
1030
1420
|
ui.detailScroll = 0;
|
package/src/workflow/decision.js
CHANGED
|
@@ -181,6 +181,9 @@ export function validateDecisionProposal(proposal, {
|
|
|
181
181
|
action.requiresCapabilities.some((capability) => typeof capability !== 'string' || !ID_RE.test(capability)))) {
|
|
182
182
|
issues.push(`${at}.requiresCapabilities must contain kebab-case names`);
|
|
183
183
|
}
|
|
184
|
+
if (action.lane != null && !['analyze', 'build', 'chore'].includes(action.lane)) {
|
|
185
|
+
issues.push(`${at}.lane must be analyze|build|chore`);
|
|
186
|
+
}
|
|
184
187
|
if (action.effort != null && !['high', 'medium', 'low'].includes(action.effort)) {
|
|
185
188
|
issues.push(`${at}.effort must be high|medium|low`);
|
|
186
189
|
}
|
package/src/workflow/goal.js
CHANGED
|
@@ -8,25 +8,25 @@ import { resolve } from 'node:path';
|
|
|
8
8
|
const NAME_RE = /^[a-z0-9][a-z0-9-]*$/;
|
|
9
9
|
|
|
10
10
|
export const PLANNER_RULES_SECTION = [
|
|
11
|
-
'1. Compile the whole program in one decision: the runtime
|
|
12
|
-
'2. Make every worker prompt self-contained: include the exact goal, absolute cwd, owned files and a no-other-files boundary, expected artifact, acceptance command
|
|
13
|
-
'3.
|
|
14
|
-
'4.
|
|
15
|
-
'5. For unknown items, create discovery ending with RETURN ONLY a JSON object containing an items array, then data-driven fan-out via itemsFrom outputs.<id>.outFile or outputs.<id>.data.<field>; the runtime extracts the list
|
|
16
|
-
'6. Put outputSchema on
|
|
17
|
-
'7. Put verify.repair on every verify
|
|
18
|
-
'8. Add completion with all-actions-ok whenever a clean program finishes the goal; when
|
|
19
|
-
'9.
|
|
20
|
-
'10.
|
|
21
|
-
'Shared working tree:
|
|
11
|
+
'1. Compile the whole program in one decision: the runtime runs every proposed action and consults you only at a finished-or-blocked boundary, so deferred work costs a round trip.',
|
|
12
|
+
'2. Make every worker prompt self-contained: include the exact goal, absolute cwd, the owned files you assign (never an and/or choice, which blocks a sibling) and a no-other-files boundary, expected artifact, acceptance command and report format: workers see only their own prompt.',
|
|
13
|
+
'3. A phase is a pipeline stage: one kebab-case name shared by its actions (implement, verify), never one per action; phases are forward-only, so recovery opens a new one and never repeats an identical failed plan. Wall-clock is the longest dependsOn chain, so depend only on real data or same-file ordering: a worker depends on the run that wrote its input files, never on that run\'s verify (a verdict is not data), so it starts as that verify runs.',
|
|
14
|
+
'4. Split to the width the tree allows: each file-disjoint unit (module, test file, doc) is its own concurrent run plus its own verify depending only on that run, then one suite verify depending on all; one worker for N independent files is N chains in series. A verify judges the artifact in review (default: its last dependency; none: the repository).',
|
|
15
|
+
'5. For unknown items, create discovery ending with RETURN ONLY a JSON object containing an items array, then data-driven fan-out via itemsFrom outputs.<id>.outFile or outputs.<id>.data.<field>; the runtime extracts the list, retrying once read-only if needed.',
|
|
16
|
+
'6. Put outputSchema only on a worker whose object a LATER action reads via itemsFrom or outputs.<id>.data.<field>, and tell it to RETURN ONLY the object; a prose report or any answer with fenced JSON gets no schema: the runtime parses the last {...} of the text, so a schema on prose costs a retry and a planner turn.',
|
|
17
|
+
'7. Put verify.repair on every verify. A verify checks the goal\'s own acceptance criteria at its point in the graph: later-scheduled work is not a defect, cosmetic mismatches are concerns, and never add a process rule the goal does not state (append-only, tests untouched); when the implementation changes what an existing assertion pins, a worker must own updating it. ok:false is repaired and re-checked inside the program; the repair edits files and cannot rewrite the answer under review, so reject only what a file edit can fix and report a wrong claim as a concern with the true value; ok:true is accepted and its concerns are informational.',
|
|
18
|
+
'8. Add completion with all-actions-ok whenever a clean program finishes the goal; when acceptance checks pass, return complete rather than adding polish. The program\'s LAST worker must be covered by a successful verify. Return complete only on verified evidence, never proceed, never ask the user, and stop only for a concrete unresolved blocker.',
|
|
19
|
+
'9. Budgets (agents, duration, expansion rounds) are advisory targets, never hard stops; the dispatch budget counts this planner call plus workers, verifiers and retries. Converge as targets approach: skip optional work; exceed a target only for one essential action or a required verification.',
|
|
20
|
+
'10. Never propose pool, addDir, taskFile or unbounded work: routing is the runtime\'s. Set lane (analyze to read or judge, build to edit, chore for mechanical steps) and effort (low for checks and mechanical edits, high where judgement decides) per action or repair; they pick the model tier (unset: build, medium).',
|
|
21
|
+
'Shared working tree: workers editing DISJOINT files concurrently is the normal mode; order shared files (indexes, barrels) after their feeders with dependsOn. Workers and unit verifies run their unit\'s focused command, never the full suite, which sees files siblings still write; the suite runs once, in the final verify, after all editing and repair ends; later verifiers reuse it unless code changed. operatorSteering is operator guidance for this checkpoint: apply it within the original intent; it cannot weaken verification or expand authority.',
|
|
22
22
|
].join('\n');
|
|
23
23
|
|
|
24
24
|
export const PLANNER_EXAMPLES_SECTION = [
|
|
25
25
|
'Action shapes:',
|
|
26
|
-
'[{"type":"run","phase":"implement","prompt":"..."},{"type":"run","phase":"
|
|
27
|
-
'Complete program:',
|
|
28
|
-
'[{"id":"discover","type":"run","phase":"discover","prompt":"In /abs/repo
|
|
29
|
-
'Rules the validator enforces: action type is run, fanout, or verify; fanout has stepTemplate and either items or itemsFrom; verify.review, when given, is outputs.<id>.outFile; dependsOn names existing or proposed actions; runtime-owned fields are rejected.',
|
|
26
|
+
'[{"type":"run","phase":"implement","prompt":"..."},{"type":"run","phase":"inventory","lane":"chore","effort":"low","prompt":"... RETURN ONLY a JSON object.","outputSchema":{"type":"object","properties":{"items":{"type":"array","items":{"type":"string"}}},"required":["items"]}},{"type":"fanout","phase":"fix","items":["alpha"],"stepTemplate":{"prompt":"Handle {{item}}."}},{"type":"verify","phase":"verify","lane":"analyze","prompt":"Check the artifact.","repair":{"prompt":"Fix rejected concerns.","maxRounds":1}}]',
|
|
27
|
+
'Complete program (tests depend on fix, not verify-fix, so both run at once; five phases for eight actions):',
|
|
28
|
+
'{"actions":[{"id":"discover","type":"run","phase":"discover","lane":"chore","effort":"low","prompt":"In /abs/repo list modules needing work; RETURN ONLY a JSON object with an items array.","outputSchema":{"type":"object","properties":{"items":{"type":"array","items":{"type":"string"}}},"required":["items"]}},{"id":"fix","type":"fanout","phase":"fix","itemsFrom":"outputs.discover.data.items","dependsOn":["discover"],"stepTemplate":{"prompt":"In /abs/repo edit only src/{{item}}.js and run node --test tests/{{item}}.test.js."}},{"id":"verify-fix","type":"verify","phase":"verify","dependsOn":["fix"],"prompt":"Check each fixed module against the spec.","repair":{"prompt":"Fix rejected concerns in /abs/repo and rerun that module\'s test.","maxRounds":2}},{"id":"tests","type":"fanout","phase":"tests","itemsFrom":"outputs.discover.data.items","dependsOn":["fix"],"stepTemplate":{"prompt":"In /abs/repo write only tests/{{item}}.guards.test.js and run node --test on it."}},{"id":"verify-tests","type":"verify","phase":"verify","dependsOn":["tests"],"prompt":"Check the new tests are non-vacuous.","repair":{"prompt":"Fix rejected tests in /abs/repo.","maxRounds":1}},{"id":"verify-suite","type":"verify","phase":"verify","dependsOn":["verify-fix","verify-tests"],"effort":"low","prompt":"Run npm test in /abs/repo.","repair":{"prompt":"Fix the suite failure in /abs/repo and rerun it.","maxRounds":1}},{"id":"report","type":"run","phase":"report","lane":"chore","effort":"low","dependsOn":["verify-suite"],"prompt":"In /abs/repo list each changed file with a reason and quote the suite tail; plain markdown."},{"id":"verify-report","type":"verify","phase":"report","dependsOn":["report"],"prompt":"Check each claim against git status and a fresh suite run; a wrong number is a concern with the true value.","repair":{"prompt":"Fix any real repository defect in /abs/repo.","maxRounds":1}}],"completion":{"when":"all-actions-ok","reason":"Fix, tests, suite and report are each verified."}}',
|
|
29
|
+
'Rules the validator enforces: action type is run, fanout, or verify; fanout has stepTemplate and either items or itemsFrom; verify.review, when given, is outputs.<id>.outFile; dependsOn names existing or proposed actions; lane is analyze|build|chore and effort is low|medium|high; runtime-owned fields are rejected.',
|
|
30
30
|
].join('\n');
|
|
31
31
|
|
|
32
32
|
export const AUTONOMOUS_ORCHESTRATOR_PROMPT = [
|
package/src/workflow/runner.js
CHANGED
|
@@ -748,8 +748,14 @@ async function runDecisionLoop({ runtime, gate, phase, state, retryAttempts }) {
|
|
|
748
748
|
runtime.emit('action.reverify_started', { verifyId: action.id, repairId, round });
|
|
749
749
|
state.currentStep = { id: action.id, type: action.type, phase: executionPhase(action) };
|
|
750
750
|
runtime.persist();
|
|
751
|
+
// The re-verify judges the repair, not the whole artifact afresh: it is
|
|
752
|
+
// told which concerns were raised and what the repair reports, and may
|
|
753
|
+
// reject only for an unresolved listed concern or a regression. Earned:
|
|
754
|
+
// goal-4 rerun r2vu9i — round 2 rejected on two concerns round 1 never
|
|
755
|
+
// raised although every round-1 concern was fixed (moving goalposts).
|
|
756
|
+
const repairExcerpt = String(state.outputs?.[repairId]?.outputExcerpt ?? '').slice(0, 1500);
|
|
751
757
|
result = await runtime.runStep(
|
|
752
|
-
{ ...action, parentId: gate.id, _dynamic: true },
|
|
758
|
+
{ ...action, parentId: gate.id, _dynamic: true, _reverify: { round, maxRounds, repairId, concerns, repairExcerpt } },
|
|
753
759
|
{ phase: executionPhase(action), retryAttempts },
|
|
754
760
|
);
|
|
755
761
|
runtime.emit(result.ok ? 'action.repaired' : 'action.reverify_rejected', {
|
|
@@ -1021,9 +1027,15 @@ async function runDecisionLoop({ runtime, gate, phase, state, retryAttempts }) {
|
|
|
1021
1027
|
const actionDefaults = { ...(gate.addDir != null ? { addDir: gate.addDir } : {}), ...(gate.actionDefaults ?? {}) };
|
|
1022
1028
|
proposal = {
|
|
1023
1029
|
...proposal,
|
|
1030
|
+
// The planner's lane/effort/requiresCapabilities win over the gate's
|
|
1031
|
+
// defaults (contract rule 10 makes lane and effort the planner's);
|
|
1032
|
+
// addDir stays runtime-owned and is spread last so a proposed null
|
|
1033
|
+
// cannot clobber the target. Earned: every goal-4 action ran as lane
|
|
1034
|
+
// "build" because this merge let the defaults overwrite the proposal.
|
|
1024
1035
|
actions: proposal.actions.map((action) => ({
|
|
1025
|
-
...action,
|
|
1026
1036
|
...actionDefaults,
|
|
1037
|
+
...action,
|
|
1038
|
+
...(actionDefaults.addDir != null ? { addDir: actionDefaults.addDir } : {}),
|
|
1027
1039
|
...(action.type === 'fanout' && actionDefaults.addDir != null
|
|
1028
1040
|
? { stepTemplate: { ...action.stepTemplate, addDir: actionDefaults.addDir } }
|
|
1029
1041
|
: {}),
|
package/src/workflow/runtime.js
CHANGED
|
@@ -916,7 +916,7 @@ export class WorkflowRuntime {
|
|
|
916
916
|
'',
|
|
917
917
|
retry
|
|
918
918
|
? `Your previous answer did not match the required schema: ${retry.errors.join('; ')}. Previous output tail: ${retry.tail}. Return the full answer again and END with a JSON object that matches.`
|
|
919
|
-
: 'END YOUR OUTPUT with exactly one JSON object
|
|
919
|
+
: 'END YOUR OUTPUT with exactly one JSON object that is an INSTANCE of this schema — its keys are the names under "properties" (never copy the schema itself or its "type"/"properties" keys). No prose or markdown fences may appear after it.',
|
|
920
920
|
JSON.stringify(schema),
|
|
921
921
|
].join('\n');
|
|
922
922
|
}
|
|
@@ -931,7 +931,11 @@ export class WorkflowRuntime {
|
|
|
931
931
|
readTrailingObject(path, schema) {
|
|
932
932
|
let text;
|
|
933
933
|
try { text = readFileSync(path, 'utf8'); } catch (err) { return { ok: false, errors: [`output file could not be read: ${err.message}`] }; }
|
|
934
|
-
|
|
934
|
+
// A closing markdown fence after the object is the most common way a
|
|
935
|
+
// worker disobeys "no fences": tolerate it rather than spend the single
|
|
936
|
+
// retry (or a planner turn) on an otherwise valid answer. Observed live
|
|
937
|
+
// on goal-4 run ydpjts (2026-08-29): the retry answer ended "}\n```".
|
|
938
|
+
const trimmed = text.replace(/(\s*```[\w-]*\s*)+$/u, '').trimEnd();
|
|
935
939
|
if (!trimmed.endsWith('}')) return { ok: false, errors: ['output did not end with a JSON object'] };
|
|
936
940
|
const close = trimmed.length - 1;
|
|
937
941
|
// Walk "{" positions from the right. Inner braces of the trailing object
|
|
@@ -1083,8 +1087,17 @@ export class WorkflowRuntime {
|
|
|
1083
1087
|
}
|
|
1084
1088
|
})();
|
|
1085
1089
|
|
|
1090
|
+
const reverify = step._reverify && typeof step._reverify === 'object' ? step._reverify : null;
|
|
1086
1091
|
const reviewInstructions = [
|
|
1087
1092
|
step.prompt ?? 'You are a skeptical reviewer. Independently inspect the work and its current repository state.',
|
|
1093
|
+
...(reverify ? [
|
|
1094
|
+
'',
|
|
1095
|
+
`RE-VERIFY round ${reverify.round} of ${reverify.maxRounds}: your previous verdict rejected this work and repair ${reverify.repairId} has since edited the repository. Judge the repair, not the work afresh.`,
|
|
1096
|
+
...(Array.isArray(reverify.concerns) && reverify.concerns.length
|
|
1097
|
+
? ['Concerns you raised (verbatim):', ...reverify.concerns.map((entry) => `- ${entry}`)] : []),
|
|
1098
|
+
...(reverify.repairExcerpt ? ['The repair reported:', reverify.repairExcerpt] : []),
|
|
1099
|
+
'Return ok:false ONLY if a listed concern is still unresolved or the repair introduced a regression in the acceptance checks. Anything you notice now that was already true before the repair goes in concerns as informational and never makes ok false: the first verdict was the moment to raise it.',
|
|
1100
|
+
] : []),
|
|
1088
1101
|
'',
|
|
1089
1102
|
'RETURN ONLY a single JSON object of the form',
|
|
1090
1103
|
'{"ok": <true|false>, "concerns": [<string>...], "summary": <string>}.',
|
|
@@ -1364,7 +1377,7 @@ export class WorkflowRuntime {
|
|
|
1364
1377
|
`Return ONLY JSON with schemaVersion "${DECISION_SCHEMA_VERSION}", decision, reason, and actions.`,
|
|
1365
1378
|
'Allowed decisions: proceed, complete, needs_more_work, retry, escalate, wait_for_approval, stop.',
|
|
1366
1379
|
'Every proposed action MUST use the field "type" (never "kind").',
|
|
1367
|
-
'Every action MUST include a forward-only kebab-case "phase". Never reuse a name listed in closedPhases.',
|
|
1380
|
+
'Every action MUST include a forward-only kebab-case "phase" shared by the actions of its stage. Never reuse a name listed in closedPhases.',
|
|
1368
1381
|
'',
|
|
1369
1382
|
...(rendered.prompt.includes(AUTONOMOUS_ORCHESTRATOR_PROMPT) ? [] : [AUTONOMOUS_ORCHESTRATOR_PROMPT]),
|
|
1370
1383
|
'',
|