bullswarm 0.20.0 → 0.21.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +59 -1
- package/README.md +39 -7
- package/connectors/opencode2.json +2 -2
- package/docs/experiments/2026-08-29-dogfood-bullswarm-builds-bullswarm.md +162 -0
- package/docs/integration-audit-2026-08-31.md +411 -0
- package/package.json +1 -1
- package/skill/SKILL.md +11 -3
- package/skill/references/operations.md +7 -2
- package/src/delegate.js +120 -3
- package/src/help.js +8 -5
- package/src/lib/opencode-kaihk.js +16 -4
- package/src/workflow/dashboard.js +510 -137
- package/src/workflow/runner.js +8 -7
- package/src/workflow/status.js +2 -1
- package/src/workflow/tui.js +30 -7
package/CHANGELOG.md
CHANGED
|
@@ -1,12 +1,70 @@
|
|
|
1
1
|
# bullswarm changelog
|
|
2
2
|
|
|
3
|
+
## 0.21.0 — unified TUI shell and LLM-first delegation
|
|
4
|
+
|
|
5
|
+
- The interactive workflow viewer is now one application shell instead of
|
|
6
|
+
several screens with their own rules. The workflows list, a run, a phase and
|
|
7
|
+
an agent are four depths of one hierarchy: every screen carries the same
|
|
8
|
+
persistent breadcrumb at the top (`Workflows › hdtdxs · timeline-segments ›
|
|
9
|
+
Verify › verify-renderer`), which drops its deepest segments first when the
|
|
10
|
+
terminal is too narrow to hold the whole path. All four depths share key
|
|
11
|
+
bindings generated from a single key-map definition: Up/Down (or k/j) move
|
|
12
|
+
within the current level, Enter and Right (or l) go one level in, Esc and
|
|
13
|
+
Left (or h) go one level out, and Tab/Shift+Tab jump to the next or previous
|
|
14
|
+
workflow, re-entering the sibling at the same depth when the equivalent phase
|
|
15
|
+
exists. The drill-down layout is uniform too — a left sidebar listing the
|
|
16
|
+
current level beside a right pane previewing the highlighted item — at every
|
|
17
|
+
depth and on narrow terminals as well, including the run list, which was a
|
|
18
|
+
full-width table with no preview pane before.
|
|
19
|
+
|
|
20
|
+
- Workflow timelines now render in phase-segmented sections with continued
|
|
21
|
+
headers for interleaved phases, grouped Preflight scout/planner milestones,
|
|
22
|
+
elapsed or running phase state, and consistent desktop, narrow, and scroll
|
|
23
|
+
continuation behavior without per-line phase prefixes. Phase-completion
|
|
24
|
+
summary rows are retained, so rows such as `└─✓ completed 4/4` remain visible
|
|
25
|
+
beneath their phase headers.
|
|
26
|
+
|
|
27
|
+
- Non-interactive `workflow tui <run-id>` output now includes the same static
|
|
28
|
+
segmented timeline as the interactive viewer, alongside the historical detail
|
|
29
|
+
tree, so real command output can be used to inspect and verify the layout.
|
|
30
|
+
|
|
31
|
+
- Documentation now describes concerns as an attribute of a completed
|
|
32
|
+
outcome — `outcome.concerns` on a delivered result — rather than
|
|
33
|
+
presenting `completed_with_concerns` as its own terminal status to handle
|
|
34
|
+
separately from `completed`. Run records that carry the
|
|
35
|
+
`completed_with_concerns` status value, including legacy runs recorded
|
|
36
|
+
before this framing, remain fully readable: `runs result`, the TUI, and
|
|
37
|
+
`workflow watch` still read it exactly like `completed` — a delivered
|
|
38
|
+
result with concerns to review, never a failure.
|
|
39
|
+
|
|
40
|
+
- Delegation classification now starts with deterministic signals and, in
|
|
41
|
+
automatic execution, lets an LLM refine the choice between a single delegate
|
|
42
|
+
and a workflow. `--dry-run` performs that same bounded low-effort
|
|
43
|
+
classification request (one analyze-lane, low-effort dispatch) before
|
|
44
|
+
printing the plan, so the preview matches what a live run would decide — it
|
|
45
|
+
never dispatches the work itself. `--classify deterministic` bypasses that
|
|
46
|
+
refinement and remains the instant, no-dispatch preview; `--classify llm`
|
|
47
|
+
requires the refinement and fails if no usable LLM decision is available.
|
|
48
|
+
An explicit `--mode single|workflow` remains the caller's choice and bypasses
|
|
49
|
+
automatic LLM classification.
|
|
50
|
+
|
|
51
|
+
- OpenCode connector portability: `connectors/opencode2.json` no longer
|
|
52
|
+
hardcodes `--model kaihk/gpt-5.6-luna` in `spawn.cmd`. A plain OpenCode
|
|
53
|
+
install with no KaiHK provider configured now dispatches with OpenCode's
|
|
54
|
+
own default model instead of failing to resolve a KaiHK-only model.
|
|
55
|
+
`src/lib/opencode-kaihk.js` still injects the explicit
|
|
56
|
+
`--model <providerId>/gpt-5.6-luna` for each discovered KaiHK provider, so
|
|
57
|
+
the primary `opencode2` pool and any extra `opencode2:<id>` pools keep
|
|
58
|
+
dispatching with their pinned per-provider model exactly as before.
|
|
59
|
+
|
|
3
60
|
## 0.20.0 — common agent delegation entry point
|
|
4
61
|
|
|
5
62
|
- `/bullswarm` and `bullswarm delegate` now give agents one transparent entry
|
|
6
63
|
point for arbitrary self-contained tasks: classify the request as one bounded
|
|
7
64
|
delegate or an autonomous workflow, show the reason and conceptual plan, then
|
|
8
65
|
execute the selected engine. Explicit mode and lane overrides remain
|
|
9
|
-
available, and `--dry-run --json` exposes the decision
|
|
66
|
+
available, and `--dry-run --json` exposes the decision after its bounded
|
|
67
|
+
classification request without dispatching the work itself.
|
|
10
68
|
- Workflow decisions persist the suggested conceptual plan alongside the
|
|
11
69
|
original intent, while the packaged skill keeps the common path concise and
|
|
12
70
|
moves operational detail into a focused reference.
|
package/README.md
CHANGED
|
@@ -73,7 +73,7 @@ bullswarm pools # meter state, pace position, quarantine status
|
|
|
73
73
|
bullswarm strategy refresh --apply --yes # approve capability-aware tier autopilot
|
|
74
74
|
bullswarm delegate --cwd ~/some-repo --prompt "Explain the parser" # one agent
|
|
75
75
|
bullswarm delegate --cwd ~/some-repo --prompt "Audit all commands, fix help, and independently verify" # workflow
|
|
76
|
-
bullswarm delegate --dry-run --json --cwd ~/some-repo --prompt "Your task" #
|
|
76
|
+
bullswarm delegate --dry-run --json --cwd ~/some-repo --prompt "Your task" # bounded classification + decision/plan; no work dispatch
|
|
77
77
|
bullswarm run --lane analyze --add-dir ~/some-repo --task-file /tmp/t.md --json
|
|
78
78
|
bullswarm run --lane analyze --add-dir ~/some-repo --prompt "Inspect the parser" --json
|
|
79
79
|
bullswarm workflow goal "Fix the failing tests and verify the change" --cwd ~/some-repo
|
|
@@ -94,6 +94,22 @@ bullswarm health # re-judge saved outputs; catch gate failures
|
|
|
94
94
|
| `doctor` | Machine-readable readiness report; self-heals on first call |
|
|
95
95
|
| `workflow` | Start an autonomous goal, or run / validate / draft / inspect explicit workflows and their live instances. |
|
|
96
96
|
|
|
97
|
+
### Delegate classification
|
|
98
|
+
|
|
99
|
+
With the default `--mode auto`, `delegate` first uses deterministic task
|
|
100
|
+
signals, then uses an LLM to refine the choice between a single delegate and a
|
|
101
|
+
workflow during execution. If that optional refinement is unavailable or
|
|
102
|
+
unusable, automatic mode uses the deterministic decision.
|
|
103
|
+
|
|
104
|
+
Use `--classify deterministic` to bypass the LLM refinement — this is the
|
|
105
|
+
instant, no-dispatch preview. Use `--classify llm` when an LLM decision is
|
|
106
|
+
required: the command fails if it cannot obtain a usable one. In automatic
|
|
107
|
+
mode, `--dry-run` still performs that same bounded low-effort classification
|
|
108
|
+
request (one analyze-lane, low-effort dispatch) and prints the resulting
|
|
109
|
+
decision — it never dispatches the work itself. An explicit `--mode single` or
|
|
110
|
+
`--mode workflow` is the caller's decision and bypasses automatic LLM
|
|
111
|
+
classification.
|
|
112
|
+
|
|
97
113
|
Discover and validate workflow definitions without executing them:
|
|
98
114
|
|
|
99
115
|
```bash
|
|
@@ -213,7 +229,18 @@ bullswarm workflow goal "Implement and verify the change" --cwd . \
|
|
|
213
229
|
The worker lock covers the scout, ordinary runs, fan-out items, repairs,
|
|
214
230
|
re-verification, and runtime extraction helpers. A pool that cannot guarantee
|
|
215
231
|
the requested model is ineligible rather than silently substituting another
|
|
216
|
-
model.
|
|
232
|
+
model.
|
|
233
|
+
|
|
234
|
+
The `opencode2` connector itself does not require a KaiHK provider: its base
|
|
235
|
+
spawn command carries no hardcoded model, so a plain OpenCode installation
|
|
236
|
+
dispatches with OpenCode's own configured default. When
|
|
237
|
+
`~/.config/opencode/opencode.json` has one or more KaiHK providers configured,
|
|
238
|
+
Bullswarm discovers them and pins an explicit `--model <providerId>/gpt-5.6-luna`
|
|
239
|
+
per provider — the first as the primary `opencode2` pool, each additional one
|
|
240
|
+
as its own `opencode2:<id>` pool — which is what the `--worker-model
|
|
241
|
+
kaihk/gpt-5.6-luna` example above locks onto.
|
|
242
|
+
|
|
243
|
+
`--max-agents` and `--max-workflow-seconds` are
|
|
217
244
|
advisory planning targets; `--max-expansion-rounds` is also an advisory
|
|
218
245
|
convergence target. Hard structural safeguards are adjusted with
|
|
219
246
|
`--max-actions` and `--max-items-per-expansion`.
|
|
@@ -449,11 +476,16 @@ rather than discarding the run as a blanket failure. Delegates have no
|
|
|
449
476
|
implicit wall-clock timeout; set a step's `timeoutSec` (or direct-run
|
|
450
477
|
`--timeout`) only when an operator explicitly wants a hard termination timer.
|
|
451
478
|
|
|
452
|
-
An autonomous `complete` remains strictly verified. A planner `stop`
|
|
453
|
-
|
|
454
|
-
verification concerns and the stopping reason
|
|
455
|
-
|
|
456
|
-
|
|
479
|
+
An autonomous `complete` remains strictly verified. A planner `stop` still
|
|
480
|
+
delivers a completed outcome when a useful delivery exists: unresolved
|
|
481
|
+
verification concerns and the stopping reason ride along as `outcome.concerns`
|
|
482
|
+
and `outcome.reason`, attributes of that completed outcome rather than a
|
|
483
|
+
separate terminal status. `stop` produces `blocked` only when no useful
|
|
484
|
+
delivery exists. `workflow runs result` treats the completed outcome as ready
|
|
485
|
+
while reporting `verified:false`. The status value `completed_with_concerns`
|
|
486
|
+
still appears on some runs — including legacy ones recorded before this
|
|
487
|
+
framing — and every consumer reads it exactly like `completed`: a delivered
|
|
488
|
+
result with concerns to review, never a failure.
|
|
457
489
|
|
|
458
490
|
The planner returns versioned JSON. It may propose `needs_more_work` with
|
|
459
491
|
bounded `run`, inline-`fanout`, or `verify` actions. The deterministic runtime
|
|
@@ -3,7 +3,7 @@
|
|
|
3
3
|
"bin": "opencode",
|
|
4
4
|
"configDirs": ["~/.config/opencode"],
|
|
5
5
|
"spawn": {
|
|
6
|
-
"cmd": ["opencode", "run", "--auto", "
|
|
6
|
+
"cmd": ["opencode", "run", "--auto", "{taskFile}"],
|
|
7
7
|
"cwdMode": "pwd",
|
|
8
8
|
"$comment-cwdMode": "QUIRK: resolves its project from $PWD, not the spawn cwd. The watcher MUST set env.PWD and spawn with cwd inside the target repo, or it will silently analyse the wrong repository and answer confidently about it."
|
|
9
9
|
},
|
|
@@ -21,7 +21,7 @@
|
|
|
21
21
|
{ "match": { "path": "type", "equals": "text" }, "path": "part.text", "mode": "concat", "separator": "\n" }
|
|
22
22
|
]
|
|
23
23
|
},
|
|
24
|
-
"$comment-auto": "--auto is required for headless workflow dispatch: task files live under ~/.bullswarm, outside the target repo, and OpenCode otherwise pauses for an interactive permission approval. --model
|
|
24
|
+
"$comment-auto": "--auto is required for headless workflow dispatch: task files live under ~/.bullswarm, outside the target repo, and OpenCode otherwise pauses for an interactive permission approval. The base connector intentionally leaves --model unset so a plain OpenCode installation uses its own default; opencode-kaihk.js adds --model <providerId>/gpt-5.6-luna only for discovered KaiHK providers, after which modelSelection can replace it for assignments and step locks.",
|
|
25
25
|
"$comment-exit1": "known failure mode: writes a complete correct answer, then dies with a Console-sync auth error and exit 1. The verdict sets contentUsableDespiteExit instead of discarding the work.",
|
|
26
26
|
"meter": { "type": "none" },
|
|
27
27
|
"costRank": 1,
|
|
@@ -389,3 +389,165 @@ defect), every repair prompt says to edit only the reviewed work's files and nev
|
|
|
389
389
|
requires one owner per file including an existing test the change breaks, and the validator line states run-wide id
|
|
390
390
|
uniqueness. Direction from the user: "unless it is completely nonsense or unable to finish I don't see a reason to
|
|
391
391
|
reject so easily". Claim to test on the next rerun: none of the three `8ebi8a` rejection reasons can produce ok:false.
|
|
392
|
+
|
|
393
|
+
## Run `5cvj72` — bullswarm 0.20.0 builds bullswarm's next release (2026-08-31, real repo, installed runtime)
|
|
394
|
+
|
|
395
|
+
Goal: (R1) remove the hardcoded `kaihk/gpt-5.6-luna` from `connectors/opencode2.json` while keeping KaiHK per-provider
|
|
396
|
+
model injection, (R2) LLM-refined delegation classification on top of the deterministic guess, (R3) docs, (R4) full
|
|
397
|
+
suite green. Launched through `bullswarm delegate` itself (dry-run preview → executed pinned `--mode=workflow`).
|
|
398
|
+
|
|
399
|
+
Result: **13 min 58 s** (838 s), 1 planner turn (73 s, 8.7 %), 12 dispatches, parallelism 1.69, max 3 concurrent,
|
|
400
|
+
9-action program (3 implementers ∥ → 3 unit verifies covering R1–R3 → suite verify covering R1–R4), auto-completed
|
|
401
|
+
`completed_with_concerns` (1 informational concern), 388/388 tests — verified independently by the operator.
|
|
402
|
+
|
|
403
|
+
Reliability events, all self-healed: `update-documentation` on claude-code/sonnet-5 failed the output substance gate
|
|
404
|
+
("announcement without substance") → escalated to codex/gpt-5.6-terra, succeeded. `verify-classifier` rejected once —
|
|
405
|
+
legitimately under the lenient bar (R2 demanded BOTH reasons in the decision; `refineDecision()` dropped the
|
|
406
|
+
deterministic one; the focused tests themselves passed 18/18) — one repair round fixed it, re-verify accepted. The
|
|
407
|
+
coverage-evidence flip fired only alongside that real concern; no false rejection this run.
|
|
408
|
+
|
|
409
|
+
Routing: soft `preferredConcurrency:1` spillover worked as designed — 9/12 dispatches on `opencode2 kaihk-2/gpt-5.6-luna`
|
|
410
|
+
(the all-tier assignment), while parallel siblings spilled to claude-code (fable-5 for `implement-classifier`,
|
|
411
|
+
sonnet-5 for the failed docs attempt) and codex (terra). Note: the per-pool assignment pins the model only on its own
|
|
412
|
+
pool; spillover pools use their connector default, so a luna-only run needs either per-pool assignments or no
|
|
413
|
+
concurrent siblings.
|
|
414
|
+
|
|
415
|
+
Post-run verification by the operator: repo connector via a temp home — KaiHK off ⇒ `["opencode","run","--auto",
|
|
416
|
+
"{taskFile}"]` (no `--model`); KaiHK on ⇒ `--model kaihk/gpt-5.6-luna`, `kaihk-2/…`, `kaihk-3/…` injected per provider.
|
|
417
|
+
Live `--classify=llm` end-to-end: decision came back `source: llm-classifier` with both reasons and a sensible verdict.
|
|
418
|
+
Two observations recorded, not fixed: (a) `--dry-run` never consults the LLM, so the canonical skill flow
|
|
419
|
+
(dry-run preview → pinned re-execute) exercises only the deterministic classifier; (b) the substance gate rejected a
|
|
420
|
+
correct 16-char answer ("bullswarm 0.20.0") from the tiny smoke delegation — legitimately terse outputs still fail the
|
|
421
|
+
40/80-char floors. Installed homes keep the old connector until a setup upgrade copies the new file.
|
|
422
|
+
|
|
423
|
+
## Run p3jbha — always-LLM classification + cross-agent integration audit (2026-08-31)
|
|
424
|
+
|
|
425
|
+
Goal (4 numbered requirements): make the LLM the deciding classifier for every auto-mode
|
|
426
|
+
`delegate` call including `--dry-run` (deterministic stays as the pre-pass hint fed into the
|
|
427
|
+
LLM prompt); make LLM fallback visible in the envelope instead of silent; update every doc
|
|
428
|
+
that claimed dry-run was deterministic-only; write a read-only cross-agent integration audit
|
|
429
|
+
for Claude/Codex/Grok. Launched through the **repo binary at 8908ef8** (installed 0.20.0
|
|
430
|
+
predates the LLM classifier) via the canonical flow: `delegate --dry-run` preview → pinned
|
|
431
|
+
`--mode=workflow`. Before launch: `claude-fable-5` added to `strategy exclude-model`
|
|
432
|
+
(recorded no-Fable pref; last run's spillover violation), and three dangling
|
|
433
|
+
`bullswarm.broken-20260830` symlinks removed from all three CLIs' skill dirs.
|
|
434
|
+
|
|
435
|
+
**Preview accuracy specimen:** the deterministic classifier scored the goal correctly
|
|
436
|
+
(workflow, score 7) but the phrase "read-only cross-agent integration audit" tripped
|
|
437
|
+
`READ_ONLY_LABEL_RE`, so `hasMutationIntent` returned false and the suggested plan dropped
|
|
438
|
+
its Execute phase entirely (Inspect→Verify→Deliver for a mostly-code-edit task). A live
|
|
439
|
+
one-regex misfire — exactly the case for LLM-decided classification.
|
|
440
|
+
|
|
441
|
+
**Outcome: completed_with_concerns (verified: true), auto-completed.** Wall 1,292 s
|
|
442
|
+
(21m32s), attempt-busy 2,050 s, parallelism 1.59, max 3 concurrent, zero quota wait.
|
|
443
|
+
14 dispatches: 2 planner turns (99 s total = 7.7%) + 12 worker attempts. Diff: 6 files
|
|
444
|
+
+72/−25 plus the new 411-line audit doc. Independently re-verified by the operator:
|
|
445
|
+
`npm test` 389/389; live `--dry-run` auto → `source: llm-classifier` in 23.6 s with the
|
|
446
|
+
deterministic sub-object preserved; `--dry-run --classify=deterministic` → 0.12 s, no
|
|
447
|
+
dispatch.
|
|
448
|
+
|
|
449
|
+
**Routing:** implement/verify work on opencode2 (kaihk-2/gpt-5.6-luna); docs on
|
|
450
|
+
claude-code/claude-sonnet-5; the audit on claude-code/claude-opus-5 (822 s, the critical
|
|
451
|
+
path). The only "fable" string in the run state is the exclusion entry itself — the
|
|
452
|
+
mitigation held.
|
|
453
|
+
|
|
454
|
+
**Reliability tally — 3 ok:false verdicts, 1 genuinely earned:**
|
|
455
|
+
1. `verify-integration-audit` round 1: legitimate — the audit really ended with the
|
|
456
|
+
command appendix and lacked the required recommendations section; repair added §7
|
|
457
|
+
(five evidence-linked recommendations).
|
|
458
|
+
2. `verify-documentation` round 1: cross-ownership overreach — it rejected R3's docs
|
|
459
|
+
because R4's `docs/integration-audit-2026-08-31.md` (owned by a *different, still
|
|
460
|
+
running* action, 822 s) did not exist yet. Under the runtime's lenient acceptance
|
|
461
|
+
standard, later-scheduled work and other actions' files are concerns, not rejections.
|
|
462
|
+
3. `verify-integration-audit` re-verify after repair: **factually false** — it claimed
|
|
463
|
+
"no concrete recommendations appear at the end" and cited lines 384–411 as the final
|
|
464
|
+
section, while the repaired file (mtime 05:01:14Z, before the verifier started at
|
|
465
|
+
05:01:50Z) held §7 Recommendations at lines 345–383. The verifier anchored on its
|
|
466
|
+
previous verdict instead of re-reading. This exhausted maxRounds=1, failed the action,
|
|
467
|
+
blocked `verify-suite`, and forced planner turn 2 — which proposed a fresh
|
|
468
|
+
`verify-completion` that passed all four requirements with line-level evidence, then
|
|
469
|
+
auto-completed. Recovery cost ≈ 2 min of wall time.
|
|
470
|
+
|
|
471
|
+
**Observations recorded, not yet fixed:** (a) re-verify verdict anchoring — the reverify
|
|
472
|
+
prompt could require the verifier to re-read the changed files and address the repair's
|
|
473
|
+
report before repeating a rejection; (b) requirement-coverage entanglement keeps making
|
|
474
|
+
verifiers judge files other actions own (second run in a row); (c) the run's own audit
|
|
475
|
+
deliverable found a real product defect: `integrate status` computes skill-link identity
|
|
476
|
+
against the *invoking checkout's* path, so any other valid Bullswarm install reports
|
|
477
|
+
`conflict`/exit 1 and `integrate install` refuses to repair it — stale skill symlinks are
|
|
478
|
+
sticky until removed by hand (see docs/integration-audit-2026-08-31.md §2, §7).
|
|
479
|
+
|
|
480
|
+
## Run wxfwda — retire completed_with_concerns + two dashboard truthfulness fixes (2026-08-31)
|
|
481
|
+
|
|
482
|
+
User directive: "having concern is not a problem of bullswarm itself but part of the agent
|
|
483
|
+
lifecycle, I do not think we need to formalize it as a feature" — plus two screenshot
|
|
484
|
+
defects from viewing run p3jbha: permanent ✗ phase marks on a delivered run, and
|
|
485
|
+
"Final Verification · 1/1 complete" over a pane saying "Not started yet" for the
|
|
486
|
+
never-dispatched verify-suite. Launched through the repo binary at f11e544; the preview was
|
|
487
|
+
the **first live canonical-flow use of the always-LLM classifier** — `source:
|
|
488
|
+
llm-classifier` in 20.6 s, agreeing with the deterministic hint (workflow, score 7).
|
|
489
|
+
|
|
490
|
+
**Outcome: auto-completed, verified: true.** Wall 1,541 s (25m41s), busy 3,110 s,
|
|
491
|
+
**parallelism 2.02** (best of the series), max 4 concurrent, planner 80 s (5.2%),
|
|
492
|
+
15 attempts, zero quota wait. Diff: 11 files +444/−67. Operator-verified: `npm test`
|
|
493
|
+
**394/394** (5 new tests), and both defects proven fixed by rendering the real
|
|
494
|
+
wf-mtgr56l1-167281 run dir with the new code — phase 6 renders `✓ Audit Verification 2/2`,
|
|
495
|
+
phase 7 renders `⊘ Final Verification 0/1` with pane "⊘ verify-suite · never dispatched ·
|
|
496
|
+
blocked by verify-integration-audit", planner panel "! Completed with 5 concerns" from the
|
|
497
|
+
outcome envelope.
|
|
498
|
+
|
|
499
|
+
**Design as landed:** new runs always terminate `completed` (or blocked/failed/…); the
|
|
500
|
+
qualification lives in `outcome` (`verified`, `bestEffort`, `concerns`, new
|
|
501
|
+
`qualification: 'verified'|'qualified'`). The stage `delivered_with_concerns` and event
|
|
502
|
+
`run.completed_with_concerns` are gone for new runs. `completed_with_concerns` remains
|
|
503
|
+
parse-only for legacy run dirs (the `budget_exhausted` precedent), rendered through the same
|
|
504
|
+
outcome-driven sentences. Dashboard invariants: a delivered run's phase list carries no
|
|
505
|
+
failure marks (attempt rows keep true history; failed/blocked/interrupted runs keep ✗);
|
|
506
|
+
a never-dispatched dependency-blocked action gets `⊘`, is excluded from "N/N complete",
|
|
507
|
+
and its pane names the failed dependency (`outputs[id].dependencyBlocked`, which the real
|
|
508
|
+
runner already writes).
|
|
509
|
+
|
|
510
|
+
**Reliability tally:** 1 rejection (verify-watcher-compatibility round 1) — cross-ownership
|
|
511
|
+
again: its concerns *praised* the reviewed work (10/10 focused tests, correctly refused to
|
|
512
|
+
touch unowned files) and complained about a missing `// legacy runs` comment in status.js,
|
|
513
|
+
a file the action did not own. Repaired in 36 s, re-verified ok. Third run in a row where
|
|
514
|
+
the only rejections judge files outside the reviewed work's ownership — the
|
|
515
|
+
requirement-coverage entanglement follow-up is now clearly the top reliability fix.
|
|
516
|
+
This run itself was labeled `completed_with_concerns` by the pre-change runtime it ran on —
|
|
517
|
+
expected artifact, not a failed fix; the label class it removes dies with the next release.
|
|
518
|
+
|
|
519
|
+
## Run hdtdxs — phase-segmented timeline (2026-08-31)
|
|
520
|
+
|
|
521
|
+
User picked the segmented layout from a three-way mockup (headers on phase boundaries,
|
|
522
|
+
`· continued` on interleave, per-line `[Phase: …]` prefixes dropped, chronology preserved).
|
|
523
|
+
Launched through the repo binary at 5064c0f; preview `source: llm-classifier` in 16.6 s.
|
|
524
|
+
**First run executing on the post-removal runtime: it terminated plain `completed`
|
|
525
|
+
(qualification: qualified, 1 concern) — live validation of the wxfwda status collapse.**
|
|
526
|
+
|
|
527
|
+
**Outcome: auto-completed, verified: true.** Wall 2,290 s (38m10s), busy 3,038 s,
|
|
528
|
+
parallelism 1.33, max 2 concurrent, planner 214 s across 4 checkpoints + 1 correction
|
|
529
|
+
(9.3%), 23 attempts. Diff: 3 files +440/−49. Operator-verified: `npm test` **400/400**
|
|
530
|
+
(6 new tests) and `node --test tests/workflow-dashboard.test.js` 32/32 — which makes the
|
|
531
|
+
run's single recorded concern (a focused-test failure on the narrow-width continuation
|
|
532
|
+
header) stale by the time of handoff; the repair that closed verify-final-acceptance fixed
|
|
533
|
+
it. Rendering both real run dirs shows 0 `[Phase:` prefixes and correct headers:
|
|
534
|
+
`── Suite ─── 1m08s ──`, `── Verify · continued ─── 16m13s ──`,
|
|
535
|
+
`── Final Acceptance Verify ─── 6m58s ──`; scrolled viewports re-emit a continuation
|
|
536
|
+
header. Non-interactive `workflow tui <runId>` now prints the same segmented timeline
|
|
537
|
+
(added mid-run after a verifier rejected the render evidence as unverifiable).
|
|
538
|
+
|
|
539
|
+
**Reliability tally — the best verifier showing of the series: 5 rejections, ALL
|
|
540
|
+
legitimate.** verify-renderer r1 (3 stale tests genuinely failing under the acceptance
|
|
541
|
+
command), r2 (narrow variant genuinely still emitted `[Phase:` prefixes), verify-tests r1
|
|
542
|
+
(CHANGELOG entry genuinely missing — this one cascaded: 8 downstream actions
|
|
543
|
+
dependency-blocked across two recovery attempts before planner turn 4 landed the fix),
|
|
544
|
+
verify-final-changelog r1 (the goal's required documentation choice genuinely absent),
|
|
545
|
+
verify-final-acceptance r1 (render evidence genuinely invalid — non-interactive tui had no
|
|
546
|
+
timeline; the repair added it). Zero cross-ownership rejections for the first time in four
|
|
547
|
+
runs. Cost of the cascade: 2 extra planner turns + 1 planner correction (it tried to reuse
|
|
548
|
+
a finished phase name), ~8 min of wall. Routing clean: opencode2/luna everywhere except
|
|
549
|
+
add-tests on claude-code (856 s); the only "fable" string in state is the exclusion entry.
|
|
550
|
+
|
|
551
|
+
**Cosmetic observations for a later pass:** an empty `── Planner · continued ──` header can
|
|
552
|
+
render with no rows beneath it when its events fall outside the viewport; the top status
|
|
553
|
+
line now reads `· done` for a plain completed run (new wording from the status collapse).
|