bullswarm 0.37.3 → 0.38.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/AGENTS.md +45 -52
- package/CHANGELOG.md +45 -0
- package/docs/design/0.38.0-removals.md +123 -0
- package/docs/guide/observing.md +2 -2
- package/docs/guide/playbook.md +13 -11
- package/docs/guide/routing.md +17 -21
- package/docs/guide/workflows.md +33 -74
- package/docs/reference/cli.md +36 -113
- package/docs/reference/program.md +5 -6
- package/docs/reference/result.md +15 -12
- package/package.json +1 -1
- package/skill/SKILL.md +3 -4
- package/skill/references/operations.md +113 -236
- package/skill/references/patterns.md +5 -5
- package/skill/references/program.md +4 -2
- package/src/help.js +92 -207
- package/src/lib/cli-flags.js +28 -22
- package/src/workflow/attempt-bytes.js +5 -7
- package/src/workflow/caller-planner.js +9 -135
- package/src/workflow/cli-capabilities.js +4 -9
- package/src/workflow/cli-goal-document.js +9 -24
- package/src/workflow/cli-goal.js +22 -63
- package/src/workflow/cli-plan.js +50 -420
- package/src/workflow/cli-program-checks.js +17 -13
- package/src/workflow/cli-run-lookup.js +61 -5
- package/src/workflow/cli-run-verbs.js +36 -27
- package/src/workflow/cli-step-verbs.js +4 -4
- package/src/workflow/cli-steps.js +23 -4
- package/src/workflow/cli.js +9 -4
- package/src/workflow/goal.js +0 -64
- package/src/workflow/home-model.js +1 -1
- package/src/workflow/kernel-resume.js +1 -16
- package/src/workflow/legacy-verification.js +514 -0
- package/src/workflow/run-control.js +5 -32
- package/src/workflow/run-model.js +1 -1
- package/src/workflow/run-view.js +1 -1
- package/src/workflow/runs-cli.js +13 -13
- package/src/workflow/step-prompts.js +3 -89
- package/src/workflow/v2-dispatch.js +187 -535
- package/src/workflow/v2-outcome.js +1 -1
- package/src/workflow/v2-planner.js +2 -113
- package/src/workflow/v2-revision.js +1 -1
- package/src/workflow/v2-runtime.js +80 -725
- package/src/workflow/v2-state.js +3 -3
- package/src/workflow/watch-cli.js +16 -12
- package/src/workflow/workflow-flags.js +29 -2
- package/src/workflow/verify-rounds.js +0 -1228
package/AGENTS.md
CHANGED
|
@@ -14,31 +14,27 @@ The core is being redesigned around facts-only mechanics, four mandatory
|
|
|
14
14
|
principles, caller-chosen options and a pattern library, for any kind of work
|
|
15
15
|
rather than code only. The design, the decisions taken and the staged build
|
|
16
16
|
plan are in `docs/design/redesign-mechanics-principles-options.md`, and draft
|
|
17
|
-
pattern cards are in `docs/design/patterns/`.
|
|
18
|
-
|
|
19
|
-
|
|
20
|
-
Stage 1 (step vocabulary) has landed: program-mode steps may state a role and a
|
|
21
|
-
deliverable, each kind belongs to one role and keeps its exact routing, and
|
|
22
|
-
the no-op gate is now 'declared deliverable not produced' (failure kind
|
|
23
|
-
`not-produced`), measured over the whole step. Stage 2 (evidence v1) has landed:
|
|
24
|
-
program steps may declare command and schema `evidence` that the kernel runs
|
|
25
|
-
after the worker, a failure is `failed-evidence` with one same-pool retry, and
|
|
26
|
-
finished steps in new runs are labelled `proven by …` or `finished · unproven`.
|
|
27
|
-
Stage 3 (failure rule and routing constraints) has landed: one automatic retry
|
|
28
|
-
per step, then the caller; a usage limit goes straight to the caller, from a
|
|
29
|
-
step, the dispatched planner or the preflight scout alike; the
|
|
30
|
-
needs-you block with `step rerun --avoid` and `step accept`; the per-step
|
|
31
|
-
`route`; `verifyRounds` counts fixes (default 1); reviews are placed only by
|
|
32
|
-
route. Runs started earlier keep their rules (`features.json`).
|
|
33
|
-
0.37.0 (program v3, the generic model) has landed: a step is a run, and a
|
|
17
|
+
pattern cards are in `docs/design/patterns/`. Each release updates this file.
|
|
18
|
+
|
|
19
|
+
0.37.0 (program v3, the generic model) landed: a step is a run, and a
|
|
34
20
|
workflow composes steps, phases, gates and loops. `bullswarm run` is a
|
|
35
|
-
one-step workflow.
|
|
36
|
-
|
|
37
|
-
|
|
38
|
-
|
|
39
|
-
|
|
40
|
-
|
|
41
|
-
|
|
21
|
+
one-step workflow. A step passes by facts (clean exit, deliverable produced,
|
|
22
|
+
evidence passed, answer matching its schema), a gate waits for
|
|
23
|
+
`workflow continue`, a loop reruns its steps until one condition holds (at
|
|
24
|
+
most 5 rounds), and new work is added with `workflow add`, never by editing
|
|
25
|
+
the run's steps. v3 reports facts per step and has no requirement IDs,
|
|
26
|
+
`evidenceFor` or `verifyRounds`.
|
|
27
|
+
|
|
28
|
+
0.38.0 (the removals release, `docs/design/0.38.0-removals.md`) deleted the
|
|
29
|
+
mechanisms v3 replaced: the preflight scout, the dispatched planner
|
|
30
|
+
(`--orchestrator`), the caller-planner gap turn (`plan show`/`plan submit`),
|
|
31
|
+
whole-plan revise (`plan export`/`plan revise`), the v2 repair loop, the
|
|
32
|
+
kernel-written digest and review tasks, and the old dispatch rules. New
|
|
33
|
+
programs are `bullswarm.workflow.program.v3` only; a v2 program is refused
|
|
34
|
+
before a run folder exists. Only a run marked `programFormat: 3` is driven.
|
|
35
|
+
Every other saved run (v2, stage 1-3, legacy) is view-only: it stays
|
|
36
|
+
listable, showable and countable, and every driving command refuses it
|
|
37
|
+
before doing anything. Their readers and stored formats are unchanged.
|
|
42
38
|
|
|
43
39
|
## Non-negotiable doctrine
|
|
44
40
|
|
|
@@ -76,34 +72,31 @@ and saved runs keep running and replaying as before.
|
|
|
76
72
|
pool inside a run; only a transient rate limit backs off on the same pool,
|
|
77
73
|
at most twice (20 s, then 60 s, or a named wait of at most 2 minutes), then
|
|
78
74
|
goes to the caller.
|
|
79
|
-
Only a failed step's dependents wait.
|
|
80
|
-
|
|
81
|
-
`
|
|
82
|
-
|
|
83
|
-
|
|
84
|
-
|
|
85
|
-
|
|
86
|
-
`
|
|
87
|
-
|
|
88
|
-
|
|
89
|
-
|
|
90
|
-
|
|
91
|
-
|
|
92
|
-
|
|
93
|
-
|
|
94
|
-
|
|
95
|
-
|
|
96
|
-
|
|
97
|
-
|
|
98
|
-
|
|
99
|
-
(
|
|
100
|
-
|
|
101
|
-
|
|
102
|
-
|
|
103
|
-
|
|
104
|
-
preserve their original semantics (`features.json`).
|
|
105
|
-
8. Historical authored-graph runs remain visible as read-only `legacy` rows.
|
|
106
|
-
Their executor was removed in 0.27.0; driving commands fail closed before
|
|
75
|
+
Only a failed step's dependents wait.
|
|
76
|
+
6. Review is a caller option, recorded as a fact. A check is an ordinary
|
|
77
|
+
step with an `answer` and/or `evidence`, and it passes by facts only.
|
|
78
|
+
Where it runs is the caller's choice through `route` (`independentOf`,
|
|
79
|
+
`providers`, `pools`); Bullswarm never moves a review on its own, and
|
|
80
|
+
records who reviewed (pool, model, provider) and whether that provider
|
|
81
|
+
also wrote the work. A caller's `step accept` is recorded as evidence
|
|
82
|
+
`choice`. Saved v2 runs keep showing their requirement verdicts
|
|
83
|
+
(`evidenceFor`, `verified`) as they were recorded.
|
|
84
|
+
7. New goal workflows are caller-planned v3 programs in a shared workspace.
|
|
85
|
+
`bullswarm workflow goal --program` executes the graph; the caller writes
|
|
86
|
+
the plan (or a step whose answer is a list of steps, appended with
|
|
87
|
+
`workflow add --from-answer`). File territories are advisory scheduling
|
|
88
|
+
hints. A check is an ordinary step, and fixing until it passes is a loop
|
|
89
|
+
the caller declares (`loops`, `until` one condition, `maxRounds` 1-5); a
|
|
90
|
+
loop out of rounds or a gate waits for the caller (`workflow continue`),
|
|
91
|
+
and a failed step goes to the caller through the watcher's needs-you
|
|
92
|
+
block: rerun elsewhere, add steps (`workflow add`), take over, or accept
|
|
93
|
+
anyway. v3 validate refuses `defaults.verifyRounds`. `--isolation` opts
|
|
94
|
+
into strict per-worker worktrees. A saved run that is not v3 is view-only
|
|
95
|
+
(0.38.0): its original semantics are kept for display (`features.json`),
|
|
96
|
+
and nothing drives it again.
|
|
97
|
+
8. Historical authored-graph runs remain visible as read-only `legacy` rows,
|
|
98
|
+
and `runs show`/`result`/`watch` print a bounded summary of them. Their
|
|
99
|
+
executor was removed in 0.27.0; driving commands fail closed before
|
|
107
100
|
dispatch and historical run directories remain untouched.
|
|
108
101
|
|
|
109
102
|
## Development
|
package/CHANGELOG.md
CHANGED
|
@@ -2,6 +2,51 @@
|
|
|
2
2
|
|
|
3
3
|
## Unreleased
|
|
4
4
|
|
|
5
|
+
## 0.38.0 — The removals release: v3 only, saved runs view-only, less code
|
|
6
|
+
|
|
7
|
+
- removed: the mechanisms program v3 replaced (0.37.0) no longer run. The
|
|
8
|
+
preflight scout (`--scout`, `--no-scout`), the dispatched Workflow Planner
|
|
9
|
+
(`--orchestrator`, `--orchestrator-model`, `--orchestrator-strict`,
|
|
10
|
+
`--strict-orchestrator`, `--suggested-plan`, `--planner-reasoning`), the
|
|
11
|
+
caller-planner gap turn (`workflow plan show`, `workflow plan submit`),
|
|
12
|
+
whole-plan revise (`workflow plan export`, `workflow plan revise`),
|
|
13
|
+
`workflow plan contract --v2`, the v2 repair loop (`defaults.verifyRounds`),
|
|
14
|
+
the kernel-written digest and review tasks, and goal requirement
|
|
15
|
+
extraction are gone. For one release each removed flag and verb exits 2
|
|
16
|
+
with one sentence naming its replacement: a first step your later steps
|
|
17
|
+
depend on instead of the scout; a program you write, or a step whose answer
|
|
18
|
+
is a list of steps appended with `workflow add --from-answer`, instead of
|
|
19
|
+
the planner; `workflow add` and `workflow step rerun` instead of revise.
|
|
20
|
+
- v2 programs: a `bullswarm.workflow.program.v2` is refused by
|
|
21
|
+
`workflow goal --program` and `workflow plan validate` with exit 2, before
|
|
22
|
+
any run folder exists: `bullswarm.workflow.program.v2 is no longer accepted
|
|
23
|
+
for a new run; write a program.v3 (bullswarm workflow plan contract) and
|
|
24
|
+
check it (bullswarm workflow plan validate --program <file.json>)`.
|
|
25
|
+
- saved runs: only a run marked `programFormat: 3` is driven. A run an
|
|
26
|
+
earlier Bullswarm started (v2 and every earlier format, legacy included)
|
|
27
|
+
is view-only: `workflow resume`, `goal --resume`, `pause`, `steer`,
|
|
28
|
+
`step rerun`/`accept`/`restart` and the removed plan verbs exit 2 with
|
|
29
|
+
`run <id> was started by an earlier Bullswarm and is view-only; start a new
|
|
30
|
+
run: …` and change nothing (`workflow add` gives the same sentence with
|
|
31
|
+
exit 1); the kernel refuses to resume one too. `workflow cancel` still finalizes a v2 run left live. Every saved run
|
|
32
|
+
stays listed, shown and counted as before (`runs list --all`, `show`,
|
|
33
|
+
`result`, `--summary`, `watch`, the dashboard, `stats`, History), and a
|
|
34
|
+
saved repair-loop run keeps its recorded `next` text.
|
|
35
|
+
- legacy runs: `workflow runs show`, `runs result`, `watch` and `tui` on a
|
|
36
|
+
run from before 0.27 print a bounded read-only summary built from its saved
|
|
37
|
+
files instead of refusing; minutes or cost the files do not hold stay
|
|
38
|
+
unknown.
|
|
39
|
+
- dispatch: every step follows the one failure rule: one automatic retry,
|
|
40
|
+
then the caller; a usage limit goes straight to the caller; a transient
|
|
41
|
+
rate limit backs off on its own pool only. The rules older runs used are
|
|
42
|
+
gone with those runs' execution.
|
|
43
|
+
- `workflow capabilities` no longer lists `actionRoles` or the dispatched
|
|
44
|
+
planner mode; roles and kinds stay only as the reader of saved v2 runs.
|
|
45
|
+
- The text judge (`judgeContent`) stays: `pools probe`, `health` and
|
|
46
|
+
`contentUsableDespiteExit` in `bullswarm run --json` still use it.
|
|
47
|
+
- dispatch: a worker a signal stopped (the kernel's SIGTERM, a timeout) is
|
|
48
|
+
recorded as `interrupted`, not as a `process` failure, on v3 steps too.
|
|
49
|
+
|
|
5
50
|
## 0.37.3 — Price-band model picks, tier-level reasoning, per-step model
|
|
6
51
|
|
|
7
52
|
- strategy: a connector can opt into a price band (`priceBand: true`; Codex
|
|
@@ -0,0 +1,123 @@
|
|
|
1
|
+
# 0.38.0 removals: one live program, frozen readers
|
|
2
|
+
|
|
3
|
+
## Verdict
|
|
4
|
+
|
|
5
|
+
0.38.0 can remove the special-purpose mechanisms by deleting **execution paths only**. Two decisions make that safe: new v2 programs are refused (D2), and every run that is not a v3 run becomes view-only (D1). After that no code path can reach the preflight scout, the dispatched planner, the repair loop, whole-plan revise, the digest kind or the old dispatch rules for a *new* action. The readers that show and count saved runs stay byte-for-byte, and so do the stored formats (D5).
|
|
6
|
+
|
|
7
|
+
Three facts from the code limit how far this goes:
|
|
8
|
+
|
|
9
|
+
- The saved-run reader re-runs the full v2 validator. `validateV2DurableState` replays every stored v2 program revision through `validateActionProgram`, with its kinds, roles, requirements and `verifyRounds` (`src/workflow/v2-state.js:788-830`). So `action-validator.js` and `step-vocabulary.js` stay as the reader of old runs. Roles and kinds stop being **rules for new work**; the code that parses them does not go.
|
|
10
|
+
- Every v3 run is stored in the v2 envelope. A v3 launch is a caller-planner turn: `applyInitialCallerProgram` validates it with `validateV2PlannerResponse`, writes `candidate-workflow-planner-turn-N.json` and calls `acceptCallerPlannerResponse` (`src/workflow/v2-runtime.js:592-609`). v3 steps also store derived v2 fields: an implicit requirement `goal`, `affects`, and `role: "act"` for an outward deliverable (`src/workflow/program-v3.js:5-14`, `src/workflow/program-v3.js:36-39`, `src/workflow/program-v3.js:296-302`). Changing that stored form would need a second reader for 0.37.x v3 runs, so 0.38.0 keeps it.
|
|
11
|
+
- The text judge is not only a v2 leftover. `bullswarm pools probe` calls `watchOnce` with no output validator (`src/provider-cli.js:782-784`), and the `contentUsableDespiteExit` fact that `bullswarm run --json` prints is computed by `judgeContent` when a v3 step with no answer exits non-zero (`src/lib/watch.js:503`, `src/lib/watch.js:651-666`, `src/workflow/answers.js:118-120`, `src/workflow/run-verdict.js:4-7`, `src/workflow/run-verdict.js:110`). So `src/lib/verify.js` stays in 0.38.0 (R10, D6).
|
|
12
|
+
|
|
13
|
+
## Measured starting point (d932f4f, branch v38/removals = 0.37.2)
|
|
14
|
+
|
|
15
|
+
Re-measured in round 2 of this design, on 2026-09-30:
|
|
16
|
+
|
|
17
|
+
| Measure | Result |
|
|
18
|
+
|---|---|
|
|
19
|
+
| Source lines, `find src -name '*.js' -o -name '*.mjs' \| xargs wc -l` | **75,440** (167 files) |
|
|
20
|
+
| `npm test` | **2,623 / 2,623 pass**, 0 fail |
|
|
21
|
+
| Replay, saved corpus vs `replay-baseline-v0.36.0.json` | **340 runs, 0 differ** |
|
|
22
|
+
| Replay, QA corpus vs `replay-baseline-v0.36.0-qa.json` | **6 runs, 0 differ** |
|
|
23
|
+
| Replay, scratch copy of the 9 finished v3 runs a 0.37.x home holds, captured twice | **9 runs, 0 differ**; the checks state, revise, result, summary, handback and proof each ran on all 9 |
|
|
24
|
+
| Stats, `stats-totals.mjs` on the frozen snapshot, this checkout read twice | **0 differences** |
|
|
25
|
+
|
|
26
|
+
**The gates are blind to v3 runs.** The saved replay corpus holds 0 `features.json` files, so none of its runs is v3. The stats snapshot is pinned to 2026-09-28T07:44Z, before 0.37.0 shipped, and holds 0 runs marked `programFormat: 3`. A 0.37.x home holds 10 v3 runs: 9 finished and this design run, which is still live. 0.38.0 changes the code v3 runs go through, so slice S0 must freeze a v3 corpus and a new stats capture before any source slice starts.
|
|
27
|
+
|
|
28
|
+
The replay tool imports named exports **by file path** (`replay3.mjs:140-148`): `v2-state.js`, `v2-revision.js` (`exportV2Plan`, `normalizeRevisionInput`, `planV2Revision`), `execution-policy.js`, `v2-outcome.js`, `v2-cancellation.js`, `rollup.js`, `run-features.js`, `short-id.js`, `history-view.js`. The stats tool does the same for `rollup.js`, `stats-model.js`, `budget-model.js`, `history.js` and `lib/tasks.js` (`stats-totals.mjs:42-46`). A missing export shows up as a coverage difference, not a pass (`replay3.mjs:23-24`). **Rule: every export either tool loads stays at its current path.** One of them loses its last product importer in 0.38.0: `normalizeRevisionInput` is used only by `plan revise` (`src/workflow/cli-plan.js:468`). `tests/unused-exports.test.js` fails on an export only tests read (`tests/unused-exports.test.js:7-10`), so S2 adds it to that test's kept-on-purpose list (`TEST_SEAMS`, `tests/unused-exports.test.js:29-36`) with the reason "loaded by the replay tool".
|
|
29
|
+
|
|
30
|
+
## Removal map
|
|
31
|
+
|
|
32
|
+
The "Keep" column lists the reader that must survive. "Delete" means the execution path only.
|
|
33
|
+
|
|
34
|
+
| ID | Delete (execution) | Replaced in v3 by | Keep (reader) and risk |
|
|
35
|
+
|---|---|---|---|
|
|
36
|
+
| R1 requirements and `verified` | Goal requirement extraction: `compactRequirement`, `REQUIREMENT_GRANULARITY_HINT` and `extractGoalRequirements` (`src/workflow/goal.js:60-121`), called only for a v2 goal document (`src/workflow/cli-goal-document.js:31-35`, `src/workflow/cli-goal-document.js:83`) and by `plan contract --v2` advice (`src/workflow/cli-plan.js:6`, `src/workflow/cli-plan.js:135-140`). The review task `buildEvidenceTask` (`src/workflow/step-prompts.js:231`) and its call (`src/workflow/v2-runtime.js:859-863`). The requirement-driven repair functions are listed under R6. | A check is an ordinary step with an `answer` and/or `evidence`; the step passes by facts only (`src/workflow/answers.js:13-16`, `src/workflow/program-v3.js:164-201`); the printed v3 result drops `verified` and the implicit requirement (`src/workflow/runs-cli.js:40-47`). | Keep the ledger reader (`src/workflow/v2-state.js:7`), `verified` in rollups and History (`src/workflow/rollup.js:90-103`, `src/workflow/rollup.js:159-160`, `src/workflow/history.js:254-288`) and the v2 proof lines (`src/workflow/v2-outcome.js:5-9`). v3 keeps storing its one implicit, non-mandatory requirement (D5): `programV3AcceptanceIssues` requires it (`src/workflow/program-v3.js:497-505`) and `bullswarm run` creates it (`src/lib/run-step.js:153-160`). |
|
|
37
|
+
| R2 roles and kinds as rules | The kind rules left in dispatch: the `digest`/`integration` exemptions of the deliverable check (`src/workflow/v2-dispatch.js:564-572`); the role catalog that `workflow capabilities` prints as `actionRoles` (`src/workflow/cli-capabilities.js:12`, `src/workflow/cli-capabilities.js:35`) and its source `v2RoleCatalog` (`src/workflow/v2-planner.js:485`). | `lane`, `effort`, `route`, `deliverable` and `retry` on the step (`src/workflow/program-v3.js:46-49`, `src/workflow/program-v3.js:287-302`); v3 already refuses `kind`/`role` as input (`src/workflow/program-v3.js:279-282`). | The outward no-repeat rule stays as it is: it reads `roleOf(action) === 'act'` (`src/workflow/v2-dispatch.js:1720-1722`, `src/workflow/v2-dispatch.js:1781-1798`), and a v3 outward step stores `role: 'act'` (`src/workflow/program-v3.js:301`), which `roleOf` returns first (`src/workflow/step-vocabulary.js:86-88`). `action-validator.js` (kind table `src/workflow/action-validator.js:23-34`, fields `src/workflow/action-validator.js:46-53`) and `step-vocabulary.js` stay whole: stored v2 validation calls them (`src/workflow/v2-state.js:807-830`), and so does v3 validation (`src/workflow/program-v3.js:392-419`). Deleting a table row changes replay `state` results. |
|
|
38
|
+
| R3 digest kind | The kernel-written digest task `buildDigestTask` (`src/workflow/step-prompts.js:182-209`), its runtime branch (`src/workflow/v2-runtime.js:859-863`), its byte-ledger option (`src/workflow/attempt-bytes.js:9-11`, `src/workflow/attempt-bytes.js:29-32`) and its dispatch exemptions (`src/workflow/v2-dispatch.js:564-572`). | An ordinary step whose prompt the caller writes, with `dependsOn` and an `answer` or deliverable (`src/workflow/program-v3.js:287-302`). | The `digest` row in the validator (reader) and the `Digest` label for old steps (`src/workflow/run-model.js:692`). |
|
|
39
|
+
| R4 preflight scout | `--scout`/`--no-scout` (`src/workflow/cli-goal.js:39-70`); `runScout` (`src/workflow/v2-runtime.js:447-570`); the scout prompt and unit ids (`src/workflow/goal.js:124-160`), imported by the runtime (`src/workflow/v2-runtime.js:29`); scout crash recovery on resume (`src/workflow/kernel-resume.js:117-121`); the scout and planner limit-stop restart (`src/workflow/run-control.js:178-214`). | A first step or phase the caller declares, with later steps depending on it (`src/workflow/program-v3.js:341-352`). | `state.preflight.scout` stays in the stored-state validator and in views (`src/workflow/v2-state.js:605-613`, `src/workflow/run-view.js:550`, `src/workflow/watch-cli.js:448-464`); scout attempts stay counted (`src/workflow/rollup.js:126-128`, `src/workflow/metrics.js:561-568`). |
|
|
40
|
+
| R5 dispatched planner (`--orchestrator`) | The `--orchestrator*`, `--suggested-plan` and `--planner-reasoning` flags (`src/workflow/cli-goal.js:39-70`, `src/workflow/cli-goal.js:152-155`); planner dispatch `runPlanner` (`src/workflow/v2-runtime.js:611-777`); in `v2-planner.js`, everything from `readPlannerCandidate` to `createV2PlannerRequest` (`src/workflow/v2-planner.js:205-630`: planner context, contract rules, response shapes and examples, prompt, v2 contract, request) plus `plannerCorrectionRequest` and `buildPlannerPreflight` (`src/workflow/v2-planner.js:705-734`); the gap-turn caller handshake: `plan show`/`plan submit` (`src/workflow/cli-plan.js:83-88`, `src/workflow/cli-plan.js:226-366`) and, in `caller-planner.js`, `callerPlannerSubmitCommand` and everything after `acceptCallerPlannerResponse` (`src/workflow/caller-planner.js:25`, `src/workflow/caller-planner.js:67-167`); the awaiting-planner wait (`src/workflow/v2-runtime.js:1780-1782`). | A step whose `answer` is a list of steps, appended by the caller with `workflow add --from-answer` (`src/workflow/cli-steps.js:188-205`, `src/workflow/cli-steps.js:209-264`). | Keep the v3 launch path: `workspacePathIssues`, `validateV2PlannerResponse`, `normalizeCallerPlannerResponse` and `applyV2PlannerResponse` (`src/workflow/v2-planner.js:88`, `src/workflow/v2-planner.js:144`, `src/workflow/v2-planner.js:631`, `src/workflow/v2-planner.js:670`), called from `src/workflow/cli-program-checks.js:93-104`; `applyInitialCallerProgram` and `acceptCallerPlannerResponse` (`src/workflow/v2-runtime.js:592-609`, `src/workflow/caller-planner.js:34`). Planner attempts stay counted (`src/workflow/rollup.js:126-128`). |
|
|
41
|
+
| R6 v2 repair loop (`verifyRounds`) | The execution half of `verify-rounds.js`: 50 of its 92 top-level declarations, about 700 of its 1,228 lines. They are `createVerifyLoop`, `emptyRound`, `roundOfVerifyStep`, `roundOfRepair`, `openFirstRound`, `failingRequirements`, `rememberedPending`, `loopFailingRequirements`, `judgingVerifySteps`, `narrowedFailingRequirements`, `repairableRequirements`, `nextLoopStep`, `discoveryItems`, `roundKind`, `closeRound`, `freeActionId`, `highestEffort`, `affectingSteps`, `canonical`, `inheritedRepairEvidence`, `planRepairStep`, `repairInheritedPaths`, `repairChangedFiles`, `namesFile` and its regexes, `repairNarrowed`, `recheckSet`, `planVerifyStep`, `openNextRound`, `applyRevisionVerifyRounds`, `listText`, `repairHandoffBlock`, `repairBrief`, `roundBrief`, `repairedReports`, `ensureVerifyLoop`, `applyRevisionLoopBudget` and the constants only they read (`src/workflow/verify-rounds.js:35-37`, `src/workflow/verify-rounds.js:39-46`). Also their callers: the runtime loop driver (`src/workflow/v2-runtime.js:1463-1635`, apart from `partialFromFailedSteps` at `src/workflow/v2-runtime.js:1615-1619`, which is general and stays), the revision budget hooks (`src/workflow/run-control.js:113`, `src/workflow/v2-runtime.js:1453`), the launch hook (`src/workflow/caller-planner.js:23`) and the repair prompts (`src/workflow/step-prompts.js:7`). The hooks are no-ops for v3 already: `ensureVerifyLoop` returns at once for a v3 program (`src/workflow/verify-rounds.js:1211-1215`), and `applyRevisionLoopBudget` does nothing for a run with no loop record and existing steps (`src/workflow/verify-rounds.js:1221-1225`). | A caller-declared loop of ordinary steps, `until` one condition, `maxRounds` 1-5 (`src/workflow/program-v3.js:321-337`, `src/workflow/gates-loops.js:149-174`, `src/workflow/gates-loops.js:402-445`). | The read side moves whole into `legacy-verification.js`: the full dependency closure of every export a reader imports, 42 declarations and about 500 lines. It contains the constants `LEGACY_ROUNDS_MAX` and `FIX_ROUNDS_MAX` (`src/workflow/verify-rounds.js:32-33`), `VERIFY_LOOP_STOPS` (`src/workflow/verify-rounds.js:38`) and `NOT_JUDGED_STATUS` (`src/workflow/verify-rounds.js:48`); `loopRounds`, `clampRounds` (`src/workflow/verify-rounds.js:60-69`); `loopOf` through `kernelRepairActionIds` (`src/workflow/verify-rounds.js:91-113`); `notJudgedRequirements` (`src/workflow/verify-rounds.js:147-162`); `requirementAcceptances` and `actAffectedRequirements` (`src/workflow/verify-rounds.js:201-235`); `currentEvidenceRecords` and `latestJudgment` (`src/workflow/verify-rounds.js:298-317`); `ancestorsOf` (`src/workflow/verify-rounds.js:617-636`); `revisedVerifyRounds` (`src/workflow/verify-rounds.js:686-695`); `clip` (`src/workflow/verify-rounds.js:705-712`); `lastSucceededAttempt` (`src/workflow/verify-rounds.js:756-770`); and `loopStageLabel` through `verifyLoopResult` (`src/workflow/verify-rounds.js:881-1210`). The readers and what they import: the Run page's `kernelRepairActionIds` and `loopStageLabel` (`src/workflow/run-model.js:11`), `NOT_JUDGED_STATUS` and `verifyLoopResult` (`src/workflow/v2-outcome.js:5`), `VERIFY_LOOP_STOPS` (`src/workflow/v2-state.js:11`), `loopVerdictText` (`src/workflow/run-view.js:59`), `verifyRoundLabel` (`src/workflow/home-model.js:25`), and `revisedVerifyRounds`, which replay's `revise` check reaches through `planV2Revision` (`src/workflow/v2-revision.js:39`). Tests read the file by path too (`tests/workflow-vocabulary-docs.test.js:272`, `tests/workflow-vocabulary-docs.test.js:1154`) and import `callerDecision` from it (`tests/workflow-step-accept.test.js:16`). |
|
|
42
|
+
| R7 whole-plan revise | The `plan export`/`plan revise` CLI verbs, with `printRevisionChanges`, `parseIdList` and `loadProgramRun` (`src/workflow/cli-plan.js:367-556`). A v3 run already refuses revise, except for reruns (`src/workflow/revision-v3.js:18-19`), and `step rerun` covers reruns, so once v2 runs are view-only the verbs have no target. `plan export` is read-only, but it prints the document `plan revise` takes, and its `next` is the revise command (`src/workflow/cli-plan.js:416`). The kernel's queued-revision application (`src/workflow/v2-runtime.js:1426-1462`) is **not** removed: `workflow add` sends an append-only request through it (`src/workflow/revision-v3.js:61-70`, `src/workflow/cli-steps.js:258-260`), and so do `step rerun`/`accept` (`src/workflow/cli-step-verbs.js:257`, `src/workflow/cli-step-verbs.js:333`) via `reviseV2Program` (`src/workflow/run-control.js:252`). | `workflow add` appends; `step rerun`/`accept` and `workflow continue` keep their own purposes (`src/workflow/revision-v3.js:61-96`, `src/workflow/gates-loops.js:464-521`). | Keep the pure `exportV2Plan`/`normalizeRevisionInput`/`planV2Revision` in `v2-revision.js` (`src/workflow/v2-revision.js:79`, `src/workflow/v2-revision.js:321-343`). Replay's `revise` check calls them (`replay3.mjs:141`), and `workflow add` uses the same module (`src/workflow/cli-steps.js:33`, `src/workflow/cli-steps.js:247`). |
|
|
43
|
+
| R8 authoring new v2 programs | v2 admission at `goal --program` and `plan validate` (`src/workflow/cli-goal.js:190-225`, `src/workflow/cli-plan.js:153-205`), and `plan contract --v2` (`src/workflow/cli-plan.js:116-146`), whose builder `buildV2PlannerContract` goes with R5 (`src/workflow/v2-planner.js:498`). | v3 only; the v3 contract already names the v2-only fields (`src/workflow/contract-v3.js:117`), and `plan contract` prints v3 by default (`src/workflow/cli-plan.js:129-134`). | Stored v2 program validation stays (`src/workflow/v2-state.js:788-830`). |
|
|
44
|
+
| R9 old dispatch rules (wave U's core) | The `failureRule`-off sides of dispatch: the unmarked failure classifier choice and limits-only lane (`src/workflow/v2-dispatch.js:1033-1035`), the old-rules retry loop (`src/workflow/v2-dispatch.js:1192-1253`), the pinned evidence retry (`src/workflow/v2-dispatch.js:1724`), and the other `!failureRule` arms (`src/workflow/v2-dispatch.js:1885`, `src/workflow/v2-dispatch.js:2226`). Every run launched since stage 3 writes `failureRule: 1` (`src/workflow/v2-runtime.js:243`, `src/workflow/run-features.js:16-18`), so after D1 no driven run takes the old branch. | One rule: one automatic retry, then the caller; a usage limit goes straight to the caller. | `classifyFailure` stays: the marked classifier calls it (`src/workflow/v2-dispatch.js:123-124`). `runFeatureFlags` stays for readers (`src/workflow/run-features.js:26-34`): the outcome and needs-you code still branch on it to show old runs. Tests that launch a run the way an earlier stage did (`src/workflow/v2-runtime.js:239-243`) lose their subject and are deleted. |
|
|
45
|
+
| R10 text judgement (**not removed in 0.38.0**) | Nothing. The `judgeContent` fallback in `watchOnce` (`src/lib/watch.js:520-531`, `src/lib/watch.js:605-612`) is dead for workflow steps after D1, because every v3 step passes a validator (`src/workflow/answers.js:131`, `src/workflow/v2-runtime.js:959`). It is still live for `pools probe` (`src/provider-cli.js:782-784`), for `contentUsableDespiteExit` on a v3 step with no answer (`src/lib/watch.js:651-666`) and for `bullswarm health` (`src/cli.js:318`, `src/cli.js:335-370`). | Answers, evidence and the deliverable already decide v3 steps; the no-output/no-changes flags stay as facts (`src/workflow/answers.js:160-175`). | Removing it changes two public outputs, which D6 decides. The 80-character rule is `src/lib/verify.js:19` and `src/lib/verify.js:162-197`. |
|
|
46
|
+
|
|
47
|
+
## Compatibility contract
|
|
48
|
+
|
|
49
|
+
- **Which runs are drivable.** A run is drivable only when its marker says `programFormat: 3` (`isProgramV3Run`, `src/workflow/run-features.js:57-69`). Everything else is view-only: v2 runs, stage-1/2/3 runs, and legacy runs (no `state.v2`, `src/workflow/short-id.js:27-29`). No file is rewritten. `features.json` is written once, at launch (`src/workflow/run-features.js:1-4`).
|
|
50
|
+
- **One CLI guard.** Today's refusal helper, `legacyRunRefusal` (`src/workflow/cli-run-lookup.js:28-36`), guards pause, cancel, resume, steer, step, action, `plan export`/`plan revise` and the dashboard (`src/workflow/cli-run-verbs.js:27`, `src/workflow/cli-run-verbs.js:61`, `src/workflow/cli-run-verbs.js:181`, `src/workflow/cli-run-verbs.js:239`, `src/workflow/cli-step-verbs.js:452`, `src/workflow/cli-inspect.js:197`, `src/workflow/cli-plan.js:384`, `src/workflow/cli-plan.js:454`, `src/workflow/cli.js:88-93`). Add one sibling, `drivableRunRefusal(token, opts, verb)`, for the driving verbs: resume, `goal --resume`, pause, steer, and step rerun/accept/restart. `workflow add` already refuses a non-v3 run, but its message points at `plan export`/`plan revise`, which 0.38.0 removes, so the message must change (`src/workflow/cli-steps.js:221`). The guard exits 2 with: `run <id> was started by an earlier Bullswarm and is view-only; start a new run: bullswarm workflow goal "<goal>" --cwd <run folder> --program <file.json>`. `workflow cancel` still finalizes a live v2 run (legacy runs have no kernel, and keep today's refusal).
|
|
51
|
+
- **The kernel refuses too.** Today the kernel resumes any run with a v2 `goal.json` and `state.json` (`src/workflow/v2-runtime.js:176-210`). S4 adds the same refusal after `assertV2Resume`, so a direct kernel start cannot drive an old run either.
|
|
52
|
+
- **Viewable.** `runs list --all`, `show`, `result`, `--summary`, `watch` and the dashboard Runs/Run/Step pages keep rendering v2 runs through today's code. For legacy runs, `show`/`result`/`watch`/`tui` refuse today (`src/workflow/runs-cli.js:283-291`, `src/workflow/runs-cli.js:362-368`, `src/workflow/watch-cli.js:1165-1168`). D3 turns those into a bounded read-only summary built from `readLegacyRunFacts` (`src/workflow/rollup.js:328-401`); unknown minutes or cost stay unknown. `legacyRunLine` itself stays byte-identical, because replay's `refusal` check hashes it (`replay3.mjs:147`).
|
|
53
|
+
- **Old hints stay as they were.** A saved repair-loop run's `next` text names `plan export`/`plan revise` and `step rerun`/`accept` (`src/workflow/verify-rounds.js:1084-1093`, `src/workflow/verify-rounds.js:1164-1179`). It stays byte-identical, because replay hashes the summary and handback. The removed `plan export`/`plan revise`/`plan show`/`plan submit` answer with exit 2 and the view-only sentence rather than a bare usage error (D8).
|
|
54
|
+
- **Countable.** No change to `metrics.js`, `metrics-legacy.js`, `rollup.js`, `history.js`, `stats-model.js` or `budget-model.js`. Scout, planner and repair attempts keep counting (`src/workflow/rollup.js:114-160`, `src/workflow/metrics.js:561-568`, `src/workflow/metrics-legacy.js:72-124`, `src/workflow/history.js:215-288`).
|
|
55
|
+
- **New v2 programs are refused, not translated.** A translator would have to invent meanings for `evidenceFor`, requirements and `verifyRounds`. The v2 validator accepts them (`src/workflow/action-validator.js:46-67`) and v3 refuses them (`src/workflow/program-v3.js:279-282`, `src/workflow/contract-v3.js:117`). The refusal happens before a run folder exists (`src/workflow/cli-program-checks.js:93-104`) and exits 2 with: `bullswarm.workflow.program.v2 is no longer accepted for a new run; write a program.v3 (bullswarm workflow plan contract) and check it (bullswarm workflow plan validate --program <file.json>)`. `plan validate` gives the same answer.
|
|
56
|
+
- **Replay.** Saved and QA corpora: 0 differences against `replay-baseline-v0.36.0*.json`, apart from the named D3 differences. v3 corpus: 0 differences against the S0 capture.
|
|
57
|
+
- **Stats.** 0 differences against the S0 capture, on both snapshots.
|
|
58
|
+
|
|
59
|
+
## Parked wave U (patch at 29eb31c)
|
|
60
|
+
|
|
61
|
+
The patch makes earlier runs view-only, adds a `{"rules": 1}` marker, and gives dispatch one rule for steps, planner and scout (`wave-u-29eb31c.patch:4-17`). By `git apply --numstat`, its source bulk is `v2-dispatch.js` (+189/−546), `v2-runtime.js` (+104/−114), `workflow/cli.js` (+97/−48) and `run-features.js` (+68/−16). Its critic's three changes, findings (a), (b) and (c) in the redesign plan's release section, all concern the planner and the scout.
|
|
62
|
+
|
|
63
|
+
- **Moot:** the planner and scout halves of the one rule, since R4 and R5 delete those dispatch sites. So are findings (a), shared crash/schema retry for planner and scout, (b), a planner or scout stopped by a nearly spent pool reading as schema/process, and (c), a planner gate retry, together with the restart path (b) changed (`src/workflow/run-control.js:178-214`). The `rules` marker is also moot. Its view-only boundary was "no stage-3 marker", but 0.38.0's boundary is `programFormat: 3`, which already exists (`src/workflow/run-features.js:57-69`). A second version key would be exactly the temporary machinery that scoping rule 3 forbids. The `quota.js`, `watch.js` and `_schema.json` hunks only reword comments about runs "started by an earlier version" (`wave-u-29eb31c.patch:575-648`).
|
|
64
|
+
- **Keep, re-implemented on 0.37.2:** the view-only decision and its exit-2 verb list (S1), the kernel's own resume refusal (S4), and the step half of "one dispatch rule" (R9, S3). With v3-only execution, deleting the `failureRule`-off branches is safe, and it is the patch's largest deletion. Its new test file, `tests/workflow-view-only-runs.test.js` (`wave-u-29eb31c.patch:6304-6584`), is a starting point for S1's test.
|
|
65
|
+
- **Why not cherry-pick:** `git apply --check` of the patch on d932f4f fails on 21 files, among them `src/workflow/cli.js`, `v2-dispatch.js`, `v2-runtime.js`, `run-features.js` and `v2-state.js`. 0.37.0 split the old single `cli.js`: today `src/workflow/cli.js` is 150 lines, `wfStep` lives in `src/workflow/cli-step-verbs.js:444` and `wfAction` in `src/workflow/cli-inspect.js:189`.
|
|
66
|
+
|
|
67
|
+
## Owner decisions
|
|
68
|
+
|
|
69
|
+
| # | Question | Options | Recommendation |
|
|
70
|
+
|---|---|---|---|
|
|
71
|
+
| D1 | What happens to saved v2 and stage-1/2/3 runs? | (a) Keep a second live executor for them. (b) View-only; cancel still finalizes. | (b). It is the precondition for R4-R7 and R9; with (a) nothing can be deleted. |
|
|
72
|
+
| D2 | A v2 program submitted to 0.38.0? | (a) Translate it to v3. (b) Refuse with the v3 pointer. | (b). A translator would guess at requirements and repair rounds, and would be new code in a removals release. |
|
|
73
|
+
| D3 | Legacy (pre-0.27) runs under `show`/`result`/`watch`/`tui`? | (a) Keep today's one-line refusal. (b) A bounded read-only summary from `readLegacyRunFacts`. | (b). "Viewable" names these verbs (5b). The approved differences are named per verb; `legacyRunLine` stays unchanged. |
|
|
74
|
+
| D4 | `plan export` and the replay `revise` probe? | (a) Delete the CLI verbs and keep the pure helpers for replay. (b) Delete both and accept losing the `revise` check. | (a). The helpers are shared with `workflow add` anyway; `normalizeRevisionInput` goes on the unused-exports kept list. |
|
|
75
|
+
| D5 | The stored v3 form (derived v2 fields, implicit requirement, v2 envelope, planner turn file)? | (a) Keep it byte-identical in 0.38.0. (b) Drop the derived fields now and add a reader for 0.37.x v3 runs. | (a). (b) adds a reader and touches every v3 view; do it in a later release that owns a v3 state schema. |
|
|
76
|
+
| D6 | The text judge (`judgeContent`, `src/lib/verify.js`), which 0.37.0 stopped using as a verdict for v3 steps but which still sets the `pools probe` verdict, `contentUsableDespiteExit` and `health`'s re-judge? | (a) Keep it in 0.38.0. (b) Remove it: the probe passes a validator, as the free-model probe already does (`src/lib/probe.js:138-145`); `contentUsableDespiteExit` becomes "exit non-zero with non-empty output"; `health` loses its re-judge. | (a). All three are information, not a step verdict, so the approved deletion of 3.4 is already met for steps. (b) changes a key `bullswarm run --json` promises to keep (`src/workflow/run-verdict.js:4-7`), and no gate measures it. Revisit with a measured case. |
|
|
77
|
+
| D7 | The gate corpus? | (a) Keep today's corpora (no v3 runs). (b) S0 freezes a v3 corpus and a fresh stats snapshot first. | (b). Without it, 0.38.0's changes to v3 paths go unmeasured. |
|
|
78
|
+
| D8 | Removed verbs (`plan show/submit/export/revise`, `--scout`, `--orchestrator*`)? | (a) Drop them: the CLI's generic unknown-verb/unknown-flag errors. (b) Exit 2 with one sentence naming the v3 replacement (`workflow add`, a first step, `step rerun`). | (b) for one release, in the flag and verb tables that already exist, since saved hints still name them. Remove the sentences in 0.39.0. |
|
|
79
|
+
|
|
80
|
+
## Build slices
|
|
81
|
+
|
|
82
|
+
Each slice is one worker, and files are exact ownership. Rules for every slice:
|
|
83
|
+
|
|
84
|
+
- It deletes an export only when every file that imports it is in its list. An export it leaves with no product importer is either deleted in the same slice (its file is listed) or added to the kept list in `tests/unused-exports.test.js` with its reason; `npm test` stays green at every slice.
|
|
85
|
+
- It edits only its listed tests. A failure in any other test file is reported to the integrator, not fixed.
|
|
86
|
+
- Every source slice runs `npm test` and the three replay compares, and reports its line delta.
|
|
87
|
+
- A file may appear in two slices only when one depends on the other.
|
|
88
|
+
|
|
89
|
+
The chain is sequential:
|
|
90
|
+
|
|
91
|
+
```
|
|
92
|
+
S0 -> S1 -> S2 -> S3 -> S4 -> S5 -> S6
|
|
93
|
+
```
|
|
94
|
+
|
|
95
|
+
S2 and S3 share no file, but they do not run in parallel: wave U's one-rule change touched the goal and caller-planner tests that S2 rewrites (`wave-u-29eb31c.patch:20-62`). No two slices run at the same time, so no two parallel slices share a file.
|
|
96
|
+
|
|
97
|
+
| Slice | Depends on | Files | Removes | Checks |
|
|
98
|
+
|---|---|---|---|---|
|
|
99
|
+
| S0 freeze gates (coordinator, outside the repo) | none | `replay-baseline-0.37.2-v3.json`, `stats-baseline-0.37.2.json` and a frozen v3 corpus beside the tools | nothing | The v3 corpus holds the finished 0.37.x v3 runs plus fresh 0.37.2 runs covering: one-step run, gate, loop that passes in round 2, `add --from-answer`, answer, evidence, outward step, cancel. Capture each twice: 0 differences. A fresh stats snapshot of a home that holds v3 runs, captured twice: 0 differences. |
|
|
100
|
+
| S1 view-only guard and legacy summaries | S0 | `src/workflow/cli-run-lookup.js`, `src/workflow/cli-run-verbs.js`, `src/workflow/cli-step-verbs.js`, `src/workflow/cli-inspect.js`, `src/workflow/cli.js`, `src/workflow/runs-cli.js`, `src/workflow/watch-cli.js`, `tests/workflow-view-only-runs.test.js` (new), `tests/workflow-runs.test.js`, `tests/workflow-watch-cli.test.js` | Driving verbs on non-v3 runs; the legacy refusals of show/result/watch/tui (D3) | Exit 2 plus the message for each driving verb on a v2 and a legacy fixture; cancel still finalizes a live v2 run; show/result/watch on legacy print the summary; `node --test` on the three test files; replay saved/QA/v3 = only the D3 differences. |
|
|
101
|
+
| S2 CLI surface | S1 | `src/workflow/cli-goal.js`, `src/workflow/cli-goal-document.js`, `src/workflow/cli-plan.js`, `src/workflow/cli-program-checks.js`, `src/workflow/workflow-flags.js`, `src/workflow/cli-capabilities.js`, `src/workflow/cli-steps.js` (the refusal message only), `src/help.js`, `src/workflow/v2-planner.js` (CLI-only exports), `src/workflow/caller-planner.js` (the gap-turn handshake), `src/workflow/goal.js` (requirement extraction), `tests/workflow-goal.test.js`, `tests/workflow-program.test.js`, `tests/workflow-v2-caller-planner.test.js`, `tests/workflow-deliverable-cli.test.js`, `tests/workflow-vocabulary-docs.test.js`, `tests/help.test.js`, `tests/unused-exports.test.js` | R8 (v2 admission, `contract --v2`, `buildV2PlannerContract`); the R4/R5 flags; `plan show/submit/export/revise` and the handshake functions only they call (`readCallerPlannerRequest`, `submitCallerPlannerResponse`, `callerPlannerSubmitCommand`, `createV2PlannerRequest`); R1 goal requirement extraction (`goal.js:60-121`, `compactV2Requirements`); the role catalog (`v2RoleCatalog`, `actionRoles`); the drivable guard on `goal --resume`; D8 sentences; `normalizeRevisionInput` onto the kept list | A v2 program makes `goal --program` and `plan validate` exit 2 before any run folder exists; a v3 validate plus launch smoke passes; removed flags and verbs exit 2 with the D8 sentence; help has no `--orchestrator`/`--scout`/`plan revise`; replay all three corpora unchanged. |
|
|
102
|
+
| S3 one dispatch rule | S2 | `src/workflow/v2-dispatch.js`, `tests/workflow-v2-dispatch.test.js`, `tests/workflow-digest.test.js`, `tests/workflow-deliverable-runtime.test.js`, `tests/workflow-deliverable-earlier-work.test.js`, `tests/workflow-evidence-runtime.test.js`, `tests/workflow-run-features.test.js`, `tests/workflow-step-rerun.test.js`, `tests/workflow-v2-runtime.test.js`, `tests/workflow-verify-rounds-kernel.test.js` | R9 `failureRule`-off branches; the R2/R3 `digest`/`integration` exemptions. The `roleOf` outward rule stays. | Dispatch tests for: clean exit, not-produced, failed evidence, schema, one retry, usage limit, 429 backoff, outward no-repeat, and a caller-written digest-style step; replay all three corpora. |
|
|
103
|
+
| S4 kernel removals | S3 | `src/workflow/v2-runtime.js`, `src/workflow/kernel-resume.js`, `src/workflow/run-control.js`, `src/workflow/caller-planner.js`, `src/workflow/v2-planner.js`, `src/workflow/goal.js`, `src/workflow/step-prompts.js`, `src/workflow/attempt-bytes.js`, `src/workflow/verify-rounds.js` (execution half), `tests/workflow-v2-runtime.test.js`, `tests/workflow-v2-planner.test.js`, `tests/workflow-verify-rounds-kernel.test.js`, `tests/workflow-verify-rounds.test.js`, `tests/workflow-v2-live-revisions.test.js`, `tests/workflow-step-restart.test.js`, `tests/unused-exports.test.js` | R4 scout; R5 planner dispatch, context, prompt and awaiting; R6 the loop driver, its hooks and the 50 execution declarations of `verify-rounds.js`; R1 review task; R3 digest task and byte option. Adds the kernel refusal to resume a non-v3 run. The queued-revision path stays, because `workflow add` and `step rerun`/`accept` use it. | Runtime tests for: v3 one-step, multi-step, `add`, gate plus `continue`, loop round 2, usage limit, kernel restart, and resume refusal on a v2 fixture; `verify-rounds.js` then holds exactly the 42 read-side declarations; replay all three corpora; stats 0 differences. |
|
|
104
|
+
| S5 legacy verification reader | S4 | `src/workflow/legacy-verification.js` (new, by `git mv` of `verify-rounds.js`), `src/workflow/v2-outcome.js`, `src/workflow/v2-state.js`, `src/workflow/v2-revision.js`, `src/workflow/run-view.js`, `src/workflow/run-model.js`, `src/workflow/home-model.js`, `tests/workflow-verify-rounds.test.js`, `tests/workflow-step-accept.test.js`, `tests/workflow-vocabulary-docs.test.js`, `tests/workflow-v2-outcome.test.js`, `tests/workflow-run-view.test.js`, `tests/workflow-v2-handback.test.js` | R6: `verify-rounds.js` as a name; its read side, unchanged, becomes `legacy-verification.js`; the six product importers and three tests point at the new path | No file names `verify-rounds.js`; the replay `summary`/`handback`/`proof`/`revise` checks are unchanged on all corpora; Run/Home view tests pass; the Run page of a saved repair-loop run renders the same. |
|
|
105
|
+
| S6 integration and release gate | S5 | `AGENTS.md`, `CHANGELOG.md`, `docs/guide/workflows.md`, `docs/guide/routing.md`, `docs/reference/program.md`, `docs/reference/cli.md`, `skill/SKILL.md`, `skill/references/program.md`, `skill/references/operations.md`, `tests/unused-exports.test.js` | Doctrine and docs for v3-only authoring and view-only runs; exports left dead by S1-S5 | Full `npm test` in two time zones; replay saved/QA vs `replay-baseline-v0.36.0*.json` and v3 vs S0; `stats-totals.mjs --diff` vs S0 = 0; `runs --all --json` identical to S0; privacy scan; final line count. |
|
|
106
|
+
|
|
107
|
+
## Expected source change
|
|
108
|
+
|
|
109
|
+
Measured sizes of the main files (lines): `v2-dispatch.js` 2,271, `v2-runtime.js` 1,959, `v2-outcome.js` 1,634, `v2-state.js` 1,460, `verify-rounds.js` 1,228, `action-validator.js` 881, `v2-planner.js` 734, `v2-revision.js` 588, `cli-plan.js` 556, `caller-planner.js` 167, `goal.js` 160.
|
|
110
|
+
|
|
111
|
+
The **forecast** adds up the declaration ranges cited above. Only the dispatch figure comes from wave U, and help text is a guess. Net about **−2,300 to −3,000 lines**, ending at roughly 72,400-73,100:
|
|
112
|
+
|
|
113
|
+
- `verify-rounds.js` execution half: about −700. The read side, about 500 lines, moves.
|
|
114
|
+
- `v2-runtime.js` scout (124), planner (167), loop driver (about 165) and the digest/review branches: about −480.
|
|
115
|
+
- `v2-planner.js` 205-630 and 705-734: about −450.
|
|
116
|
+
- `cli-plan.js` show, submit, export, revise and `--v2`: about −370.
|
|
117
|
+
- `v2-dispatch.js` `failureRule`-off branches: about −350 (wave U measured +189/−546 on the pre-split file).
|
|
118
|
+
- `caller-planner.js` handshake (about −110) and `goal.js` 60-160 (about −100).
|
|
119
|
+
- `cli-goal.js`, `cli-goal-document.js`, `cli-capabilities.js`, `workflow-flags.js`, `step-prompts.js`, `run-control.js`, `kernel-resume.js`: about −190.
|
|
120
|
+
- `help.js` text: about −100 (a guess).
|
|
121
|
+
- Added: the guard, legacy summaries, kernel refusal and D8 sentences, about +150.
|
|
122
|
+
|
|
123
|
+
The v2 validator, the state reader and the text judge stay (R2, R10, D5, D6), and they cap the saving. S6 reports the measured total from the same `find`/`wc` command. The gates decide: code that compatibility needs stays, even if the forecast is missed.
|
package/docs/guide/observing.md
CHANGED
|
@@ -618,9 +618,9 @@ bullswarm workflow watch ab12cd --verbose
|
|
|
618
618
|
|
|
619
619
|
A usage-limit failure prints whether or not `--verbose` is given: a `⚠ <actionId> usage limit on <pool>` line, with `back at <time>` in it when the watch knows when the pool is back. In runs started by this version the line ends `back to you`: a usage limit (a spent 5-hour or weekly window, or no credit left) ends the step, nothing waits for the pool or moves the step, and a needs-you block follows. The attempt's event carries no return time, so the line reads `⚠ <actionId> usage limit on <pool> · back to you`; the needs-you block that follows carries the time instead, as its `back at <time>` line, when it is known. A run from an earlier version reads `retrying on another pool`, then prints a `↺ <actionId> now on <pool> · <model>` line once the retry lands there. A saved attempt that still promised a retry reads `no retry spent`.
|
|
620
620
|
|
|
621
|
-
|
|
621
|
+
0.38.0 removed the preflight scout and the dispatched planner, so these lines come only from a run saved by 0.37.x, which `watch` still replays. In such a run a usage limit, a rate limit still there after its backoff, or no free pool stopped the scout, and the watch prints `⚠ preflight scout stopped · <label> on <pool> · back at <time>`, where `<label>` is `out of quota`, `rate limited` or `no eligible pool` and `<pool>` reads `no pool` when none was picked; `back at` is left out when no return time is known. When the run had its caller's program the line ends `· the run continues without its report` and the steps ran; without one the run finished and its `reason:` starts `the preflight scout stopped on a usage limit` (or `the preflight scout stopped: no pool free`). A dispatched planner that stopped the same way prints `✗ planner stopped · <label> on <pool> · back at <time>`, with the same labels, `no pool` and `back at` rules, and the run finished with `the workflow planner stopped on a usage limit: …`; such a run is view-only now, so `workflow resume` refuses it and a new run with your own program v3 carries the work on. The planner's or scout's schema correction, and the one retry on the same pool it got when no other pool could run it, never went back to a pool that has become nearly spent: the correction moved to another free pool, and with none free it stopped the same way, with that pool's `nearly spent` reason. A planner turn that failed for any other reason, and every planner failure in a run from before 0.37.x, prints `× planning attempt rejected · <why>`. Before either stop line, a scout or planner attempt that hit a usage limit prints its own usage-limit line, which ends `no retry left` there (`⚠ preflight-scout usage limit on <pool> · no retry left`) and is followed by no needs-you block.
|
|
622
622
|
|
|
623
|
-
When a step comes back to you the watch prints a needs-you block. With `--jsonl` the same block is one object whose `options` hold the commands: `rerunElsewhere` (`step rerun <id> <step> --avoid <pool>`) when another pool could run the step, else `retryHere` (`step rerun <id> <step>`); `waitForIt` (`after <time>: step rerun <id> <step>`) in a run started by this version whenever the step's return time is known (after a usage limit, a rate limit that named a longer wait, or with no pool free); `changeStep` (`plan export …`, edit it, then `plan revise
|
|
623
|
+
When a step comes back to you the watch prints a needs-you block. With `--jsonl` the same block is one object whose `options` hold the commands: `rerunElsewhere` (`step rerun <id> <step> --avoid <pool>`) when another pool could run the step, else `retryHere` (`step rerun <id> <step>`); `waitForIt` (`after <time>: step rerun <id> <step>`) in a run started by this version whenever the step's return time is known (after a usage limit, a rate limit that named a longer wait, or with no pool free); `addSteps` (`workflow add <id> --steps part.json`, then `workflow wait <id> <added ids>`) in a v3 run, or, in a saved v2 run, `changeStep` (`plan export …`, edit it, then `plan revise …`, as it was written; 0.38.0 removed those verbs and the run is view-only); `takeOver`; and `acceptAnyway`. The same object then carries `backAt`, the time printed as `back at <time>`. A try that followed a rate-limit backoff on the same pool reads `· after a rate-limit backoff`. A review block adds `otherChecks` when a second check also failed a requirement, one entry per check with its own rerun and accept commands. A step rerun or accept that reopens a cancelled run prints `run reopened from cancelled by a plan revision`, followed by `· not run again (act step, may have acted): <steps>` when an `act` step stays cancelled. A finished run that `workflow resume` reopened prints `run reopened from <status> by workflow resume`.
|
|
624
624
|
|
|
625
625
|
## Watch until trouble
|
|
626
626
|
|
package/docs/guide/playbook.md
CHANGED
|
@@ -68,22 +68,24 @@ read-only Run, Step, and usage facts inside the agent session.
|
|
|
68
68
|
Watching is not controlling. Quitting the dashboard leaves the workflow
|
|
69
69
|
running, and a completed process is not automatically a verified outcome.
|
|
70
70
|
|
|
71
|
-
## 5. Steer or
|
|
71
|
+
## 5. Steer or extend deliberately
|
|
72
72
|
|
|
73
|
-
|
|
74
|
-
|
|
73
|
+
Steering queues guidance on a running run: `watch` prints it, and you act on
|
|
74
|
+
it. No planner reads it. To change what the run does, add steps to it, or run
|
|
75
|
+
a finished step again:
|
|
75
76
|
|
|
76
77
|
```bash
|
|
77
78
|
bullswarm workflow steer <runId> --message "Keep the public output backward compatible"
|
|
78
|
-
bullswarm workflow
|
|
79
|
-
bullswarm workflow
|
|
80
|
-
--summary "Re-run acceptance after the compatibility fix"
|
|
79
|
+
bullswarm workflow add <runId> --steps compat-fix.json
|
|
80
|
+
bullswarm workflow step rerun <runId> acceptance
|
|
81
81
|
```
|
|
82
82
|
|
|
83
|
-
|
|
84
|
-
|
|
85
|
-
|
|
86
|
-
|
|
83
|
+
Add steps when the evidence reveals missing work or a missing check; a loop
|
|
84
|
+
you declared in the program repeats its steps until its condition holds.
|
|
85
|
+
Resume when the program is still right and a retryable mechanical failure
|
|
86
|
+
stopped a step. Do not treat a failed acceptance check as a reason to repeat
|
|
87
|
+
the same unchecked program. (`plan export` and `plan revise` were removed in
|
|
88
|
+
0.38.0; a run an earlier Bullswarm started is view-only.)
|
|
87
89
|
|
|
88
90
|
## 6. Sign off from evidence
|
|
89
91
|
|
|
@@ -114,7 +116,7 @@ same ongoing goal in the same directory unless you explicitly pass `--again`.
|
|
|
114
116
|
The main agent should retain planning and review because it holds the user's
|
|
115
117
|
intent, current conversation, and cross-step trade-offs. Let workers own clear
|
|
116
118
|
execution territories and independent checks; let the main agent reconcile
|
|
117
|
-
their evidence,
|
|
119
|
+
their evidence, extend the program when needed, and make the final call.
|
|
118
120
|
|
|
119
121
|
Next, see [Workflows](/guide/workflows) for the complete program lifecycle and
|
|
120
122
|
[Observing runs](/guide/observing) for the dashboard and watcher reference.
|
package/docs/guide/routing.md
CHANGED
|
@@ -50,7 +50,7 @@ Pace compares a pool with itself: **surplus = elapsed% − used%** of its own su
|
|
|
50
50
|
|
|
51
51
|
## The 5-hour window
|
|
52
52
|
|
|
53
|
-
The rolling 5-hour window never paces — it only protects the last mile. A recorded reading at 100% (`BURST_BLOCK_PCT`) means the pool is at its limit until the window resets, and the same holds for a weekly or monthly reading at 100%: such a pool is never picked. The 75% (`FIVE_HOUR_NEAR_LIMIT_PCT`) line is no longer a cutoff: when another eligible pool is behind pace, a near-limit pool is ordered after it; when no such alternative exists, the near-limit pool remains selectable. A forecast above 100% is also selectable, ordered last. If the provider refuses the run at the wall, that attempt ends as a usage limit: in a workflow started by this version nothing retries it and the step comes back to you, and a single `bullswarm run` exits 1.
|
|
53
|
+
The rolling 5-hour window never paces — it only protects the last mile. A recorded reading at 100% (`BURST_BLOCK_PCT`) means the pool is at its limit until the window resets, and the same holds for a weekly or monthly reading at 100%: such a pool is never picked. The 75% (`FIVE_HOUR_NEAR_LIMIT_PCT`) line is no longer a cutoff: when another eligible pool is behind pace, a near-limit pool is ordered after it; when no such alternative exists, the near-limit pool remains selectable. A forecast above 100% is also selectable, ordered last. If the provider refuses the run at the wall, that attempt ends as a usage limit: in a workflow started by this version nothing retries it and the step comes back to you, and a single `bullswarm run` exits 1. A workflow started by an earlier version moved that attempt to another pool; such a run is view-only since 0.38.0.
|
|
54
54
|
|
|
55
55
|
The change follows the 2026-09-10 observation that `claude-code:acme` was at 81% with 23 minutes left (92.3% of its 5-hour window elapsed). The old guard sent a high-tier task to another account, while 34% of acme's weekly quota expired in the remaining 13% of that week. The last-mile rule lets the task use that quota; the risk it takes is one attempt that may stop at the wall and come back to the caller.
|
|
56
56
|
|
|
@@ -114,7 +114,7 @@ otherwise it says `no retry left`.
|
|
|
114
114
|
In new runs, Bullswarm does not move a review away from a writer on its own.
|
|
115
115
|
Place it with the step's optional `route`; `route` is a hard filter before
|
|
116
116
|
quota pacing. Use `independentOf` to avoid providers that worked on named
|
|
117
|
-
upstream steps
|
|
117
|
+
upstream steps. Use
|
|
118
118
|
`providers.use` / `providers.avoid` to select provider families, or
|
|
119
119
|
`pools.use` / `pools.avoid` for exact pool ids. Accounts served by one provider
|
|
120
120
|
count as one family. A route that leaves no free pool sends the step back to
|
|
@@ -122,15 +122,15 @@ you at once, as no eligible pool or with each pool's reason; it never waits.
|
|
|
122
122
|
|
|
123
123
|
```json
|
|
124
124
|
{
|
|
125
|
-
"schemaVersion": "bullswarm.workflow.program.
|
|
126
|
-
"
|
|
127
|
-
{ "id": "write-docs", "
|
|
128
|
-
{ "id": "review-docs", "
|
|
125
|
+
"schemaVersion": "bullswarm.workflow.program.v3",
|
|
126
|
+
"steps": [
|
|
127
|
+
{ "id": "write-docs", "lane": "build", "files": ["README.md"], "prompt": "Write README.md." },
|
|
128
|
+
{ "id": "review-docs", "dependsOn": ["write-docs"], "route": { "independentOf": ["write-docs"] }, "prompt": "Check README.md and report what is wrong with it." }
|
|
129
129
|
]
|
|
130
130
|
}
|
|
131
131
|
```
|
|
132
132
|
|
|
133
|
-
Saved runs
|
|
133
|
+
Saved runs show the former automatic writer avoidance, including its independence
|
|
134
134
|
tie-breaker and urgency waiver. A selected gate retry says `pinned to <pool>
|
|
135
135
|
(the same pool (gate retry))`; a manual restart says `pinned to <pool> (step
|
|
136
136
|
restart)`. Route constraints appear as `route: <summary>` in the reason.
|
|
@@ -147,7 +147,7 @@ Before this, a pinned evidence step read `evidence step: only the writer pool co
|
|
|
147
147
|
|
|
148
148
|
## Expiring-soon urgency
|
|
149
149
|
|
|
150
|
-
A pool whose pacing window resets within 24 hours (weekly) or 3 days (monthly) is ranked on urgency — its surplus divided by the fraction of the window still to run — instead of on the surplus alone. While any urgent pool can still spend its quota, it is the only one selectable, which is how a pool with two hours left beats a pool with three days left. A pool forecast at or above 95% of its pacing window *and* ahead of the window's own clock (forecast above the elapsed share) is `draining` and goes last; a pool at 96% with 98% of its month gone is spending at its own pace, not draining, and keeps its quota in play until the reset. In a workflow started by this version a `draining` pool does not go last: it is never given a step, the dispatched planner or the preflight scout, even as the only pool left, unless you pinned it. The work goes to another pool that can run it, or comes back to you with `<pool> nearly spent (forecast <n>%) until <time>` in its `why`.
|
|
150
|
+
A pool whose pacing window resets within 24 hours (weekly) or 3 days (monthly) is ranked on urgency — its surplus divided by the fraction of the window still to run — instead of on the surplus alone. While any urgent pool can still spend its quota, it is the only one selectable, which is how a pool with two hours left beats a pool with three days left. A pool forecast at or above 95% of its pacing window *and* ahead of the window's own clock (forecast above the elapsed share) is `draining` and goes last; a pool at 96% with 98% of its month gone is spending at its own pace, not draining, and keeps its quota in play until the reset. In a workflow started by this version a `draining` pool does not go last: it is never given a step, the dispatched planner or the preflight scout, even as the only pool left, unless you pinned it. (0.38.0 removed the dispatched planner and the preflight scout, so that part applies to runs saved by 0.37.x.) The work goes to another pool that can run it, or comes back to you with `<pool> nearly spent (forecast <n>%) until <time>` in its `why`.
|
|
151
151
|
|
|
152
152
|
## In-flight load
|
|
153
153
|
|
|
@@ -202,21 +202,17 @@ finds no free pool keeps its own failure and ends its `why` with `· no retry:
|
|
|
202
202
|
<pool> <reason>; …`. When another pool that can run the step is free, routing
|
|
203
203
|
picks it as usual.
|
|
204
204
|
|
|
205
|
-
|
|
206
|
-
|
|
207
|
-
|
|
208
|
-
|
|
209
|
-
finishes with `the workflow planner stopped on a usage limit: …` (or `the
|
|
210
|
-
preflight scout stopped on a usage limit: …`) and your call: after its `back
|
|
211
|
-
at` time, `workflow resume` runs it again. A scout that ran before your own
|
|
212
|
-
program lets the run go on without its report. A sign-in
|
|
213
|
-
failure, a provider error or a worker that died at start still moves the
|
|
214
|
-
planner or the scout to another pool.
|
|
205
|
+
A run saved by 0.37.x may end with `the workflow planner stopped on a usage
|
|
206
|
+
limit: …` or `the preflight scout stopped on a usage limit: …`. 0.38.0 removed
|
|
207
|
+
the dispatched planner and the preflight scout, and such a run is view-only:
|
|
208
|
+
start a new run with a program you write instead of resuming it.
|
|
215
209
|
|
|
216
210
|
A single `bullswarm run` makes one attempt: a usage limit ends it with exit 1,
|
|
217
|
-
with no retry and no move. Workflows started by an earlier version
|
|
218
|
-
rules: a usage limit
|
|
219
|
-
the same pool and then
|
|
211
|
+
with no retry and no move. Workflows started by an earlier version followed
|
|
212
|
+
older rules: a usage limit moved the attempt to another pool, and a throttle
|
|
213
|
+
retried the same pool and then moved, as below. Since 0.38.0 such a run is
|
|
214
|
+
view-only, so those rules only explain what its saved record shows; every run
|
|
215
|
+
Bullswarm drives follows the rules for a workflow started by this version.
|
|
220
216
|
|
|
221
217
|
## Throttles and exhausted windows
|
|
222
218
|
|