bullswarm 0.35.0 → 0.35.2
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +321 -13
- package/GOAL.md +1 -1
- package/data/openrouter-benchmarks.json +10786 -10813
- package/docs/audits/2026-09-09-codebase-audit.md +11 -11
- package/docs/design/cost-0.34.0/README.md +1 -1
- package/docs/design/cost-0.34.0/meter-precision-study.md +19 -19
- package/docs/design/dashboard-prototype.html +20 -20
- package/docs/design/dashboard-split/README.md +30 -30
- package/docs/design/mod-step-v2/README.md +29 -0
- package/docs/design/mod-step-v2/failed-detail-120.txt +26 -0
- package/docs/design/mod-step-v2/failed-detail-55.txt +32 -0
- package/docs/design/mod-step-v2/failed-overview-120.txt +26 -0
- package/docs/design/mod-step-v2/failed-overview-55.txt +32 -0
- package/docs/design/mod-step-v2/finished-detail-120.txt +355 -0
- package/docs/design/mod-step-v2/finished-detail-55.txt +604 -0
- package/docs/design/mod-step-v2/finished-overview-120.txt +63 -0
- package/docs/design/mod-step-v2/finished-overview-55.txt +54 -0
- package/docs/design/mod-step-v2/running-detail-120.txt +24 -0
- package/docs/design/mod-step-v2/running-detail-55.txt +28 -0
- package/docs/design/mod-step-v2/running-overview-120.txt +24 -0
- package/docs/design/mod-step-v2/running-overview-55.txt +28 -0
- package/docs/design/mod-step-v2/usage-no-run-after.txt +19 -0
- package/docs/design/mod-step-v2/usage-no-run-before.txt +16 -0
- package/docs/design/pricing-0.35.2/auto-reprice-study.md +233 -0
- package/docs/design/pricing-0.35.2/retention-and-capture-study.md +178 -0
- package/docs/design/prototype-frames/README.md +1 -2
- package/docs/design/prototype-frames/history-120.tagged.txt +18 -18
- package/docs/design/prototype-frames/history-120.txt +18 -18
- package/docs/design/prototype-frames/history-55.tagged.txt +10 -10
- package/docs/design/prototype-frames/history-55.txt +10 -10
- package/docs/design/prototype-frames/home-120.tagged.txt +4 -4
- package/docs/design/prototype-frames/home-120.txt +4 -4
- package/docs/design/prototype-frames/home-55.tagged.txt +2 -2
- package/docs/design/prototype-frames/home-55.txt +2 -2
- package/docs/design/prototype-frames/run-120.tagged.txt +7 -7
- package/docs/design/prototype-frames/run-120.txt +7 -7
- package/docs/design/prototype-frames/run-55.tagged.txt +5 -5
- package/docs/design/prototype-frames/run-55.txt +5 -5
- package/docs/design/prototype-frames/stats-projects-120.tagged.txt +2 -2
- package/docs/design/prototype-frames/stats-projects-120.txt +2 -2
- package/docs/design/prototype-frames/stats-projects-55.tagged.txt +2 -2
- package/docs/design/prototype-frames/stats-projects-55.txt +2 -2
- package/docs/design/prototype-frames/step-120.tagged.txt +6 -6
- package/docs/design/prototype-frames/step-120.txt +6 -6
- package/docs/design/prototype-frames/step-55.tagged.txt +5 -5
- package/docs/design/prototype-frames/step-55.txt +5 -5
- package/docs/design/prototype-frames-0.33.1/budget-120-v2.txt +2 -2
- package/docs/design/prototype-frames-0.33.1/budget-120-v3.txt +2 -2
- package/docs/design/prototype-frames-0.33.1/budget-120.txt +4 -4
- package/docs/design/prototype-frames-0.33.1/budget-55-v2.txt +2 -2
- package/docs/design/prototype-frames-0.33.1/budget-55-v3.txt +2 -2
- package/docs/design/prototype-frames-0.33.1/budget-55.txt +2 -2
- package/docs/design/prototype-frames-0.33.1/budget-README.md +3 -3
- package/docs/design/prototype-frames-0.33.1/home-today-120-v2.txt +4 -4
- package/docs/design/prototype-frames-0.33.1/home-today-120.txt +5 -5
- package/docs/design/prototype-frames-0.33.1/home-today-55-v2.txt +3 -3
- package/docs/design/prototype-frames-0.33.1/home-today-55.txt +6 -6
- package/docs/design/prototype-frames-0.33.1/home-today-README.md +10 -10
- package/docs/design/providers-0.35.2/CHANGELOG-draft.md +27 -0
- package/docs/design/providers-0.35.2/README.md +172 -0
- package/docs/design/stats-frames-0.33.2/spending-120.txt +2 -2
- package/docs/design/stats-frames-0.33.2/spending-55.txt +4 -4
- package/docs/design/stats-refactor-0.33.2.md +10 -10
- package/docs/design/step-economy-0.35.2/README.md +431 -0
- package/docs/design/step-page-0.35.0/README.md +31 -31
- package/docs/design/step-page-0.35.0/frames/real-failed-120.txt +6 -6
- package/docs/design/step-page-0.35.0/frames/real-failed-200.txt +6 -6
- package/docs/design/step-page-0.35.0/frames/real-failed-55.txt +7 -7
- package/docs/design/step-page-0.35.0/frames/real-finished-120.txt +2 -2
- package/docs/design/step-page-0.35.0/frames/real-finished-200.txt +2 -2
- package/docs/design/step-page-0.35.0/frames/real-finished-55.txt +3 -3
- package/docs/design/step-page-0.35.0/frames/real-running-120.txt +6 -6
- package/docs/design/step-page-0.35.0/frames/real-running-200.txt +6 -6
- package/docs/design/step-page-0.35.0/frames/rendered-failed-120.txt +1 -1
- package/docs/design/step-page-0.35.0/frames/rendered-failed-200.txt +1 -1
- package/docs/design/step-page-0.35.0/frames/rendered-finished-120.txt +1 -1
- package/docs/design/step-page-0.35.0/frames/rendered-finished-200.txt +1 -1
- package/docs/design/step-page-0.35.0/frames/rendered-running-120.txt +10 -10
- package/docs/design/step-page-0.35.0/frames/rendered-running-200.txt +10 -10
- package/docs/design/step-page-0.35.0/frames/rendered-running-55.txt +9 -9
- package/docs/design/step-page-0.35.0/frames/running-120.txt +1 -1
- package/docs/design/step-page-0.35.0/frames/running-200.txt +3 -3
- package/docs/design/step-page-0.35.0/frames/running-55.txt +1 -1
- package/docs/design/tidy-0.35.1/README.md +296 -0
- package/docs/design/tidy-0.35.1/frames/colour/real-home-200.txt +60 -0
- package/docs/design/tidy-0.35.1/frames/colour/real-home-55.txt +60 -0
- package/docs/design/tidy-0.35.1/frames/colour/real-run-finished-200.txt +60 -0
- package/docs/design/tidy-0.35.1/frames/colour/real-run-finished-55.txt +60 -0
- package/docs/design/tidy-0.35.1/frames/colour/real-run-running-200.txt +60 -0
- package/docs/design/tidy-0.35.1/frames/colour/real-run-running-55.txt +60 -0
- package/docs/design/tidy-0.35.1/frames/colour/real-step-detail-finished-200.txt +60 -0
- package/docs/design/tidy-0.35.1/frames/colour/real-step-detail-finished-55.txt +60 -0
- package/docs/design/tidy-0.35.1/frames/colour/real-step-overview-finished-200.txt +60 -0
- package/docs/design/tidy-0.35.1/frames/colour/real-step-overview-finished-55.txt +60 -0
- package/docs/design/tidy-0.35.1/frames/colour/real-step-overview-running-200.txt +60 -0
- package/docs/design/tidy-0.35.1/frames/colour/real-step-overview-running-55.txt +60 -0
- package/docs/design/tidy-0.35.1/frames/colour/real-task-200.txt +60 -0
- package/docs/design/tidy-0.35.1/frames/colour/real-task-55.txt +60 -0
- package/docs/design/tidy-0.35.1/frames/home-120.txt +16 -0
- package/docs/design/tidy-0.35.1/frames/home-200.txt +21 -0
- package/docs/design/tidy-0.35.1/frames/home-55.txt +26 -0
- package/docs/design/tidy-0.35.1/frames/real-home-120.txt +60 -0
- package/docs/design/tidy-0.35.1/frames/real-home-200.txt +60 -0
- package/docs/design/tidy-0.35.1/frames/real-home-55.txt +60 -0
- package/docs/design/tidy-0.35.1/frames/real-run-finished-120.txt +60 -0
- package/docs/design/tidy-0.35.1/frames/real-run-finished-200.txt +60 -0
- package/docs/design/tidy-0.35.1/frames/real-run-finished-55.txt +60 -0
- package/docs/design/tidy-0.35.1/frames/real-run-running-120.txt +60 -0
- package/docs/design/tidy-0.35.1/frames/real-run-running-200.txt +60 -0
- package/docs/design/tidy-0.35.1/frames/real-run-running-55.txt +60 -0
- package/docs/design/tidy-0.35.1/frames/real-stats-model-120.txt +60 -0
- package/docs/design/tidy-0.35.1/frames/real-stats-model-200.txt +60 -0
- package/docs/design/tidy-0.35.1/frames/real-stats-model-55.txt +60 -0
- package/docs/design/tidy-0.35.1/frames/real-stats-spending-120.txt +60 -0
- package/docs/design/tidy-0.35.1/frames/real-stats-spending-200.txt +60 -0
- package/docs/design/tidy-0.35.1/frames/real-stats-spending-55.txt +60 -0
- package/docs/design/tidy-0.35.1/frames/real-step-detail-failed-120.txt +60 -0
- package/docs/design/tidy-0.35.1/frames/real-step-detail-failed-200.txt +60 -0
- package/docs/design/tidy-0.35.1/frames/real-step-detail-failed-55.txt +60 -0
- package/docs/design/tidy-0.35.1/frames/real-step-detail-finished-120.txt +60 -0
- package/docs/design/tidy-0.35.1/frames/real-step-detail-finished-200.txt +60 -0
- package/docs/design/tidy-0.35.1/frames/real-step-detail-finished-55.txt +60 -0
- package/docs/design/tidy-0.35.1/frames/real-step-detail-running-120.txt +60 -0
- package/docs/design/tidy-0.35.1/frames/real-step-detail-running-200.txt +60 -0
- package/docs/design/tidy-0.35.1/frames/real-step-detail-running-55.txt +60 -0
- package/docs/design/tidy-0.35.1/frames/real-step-overview-failed-120.txt +60 -0
- package/docs/design/tidy-0.35.1/frames/real-step-overview-failed-200.txt +60 -0
- package/docs/design/tidy-0.35.1/frames/real-step-overview-failed-55.txt +60 -0
- package/docs/design/tidy-0.35.1/frames/real-step-overview-finished-120.txt +60 -0
- package/docs/design/tidy-0.35.1/frames/real-step-overview-finished-200.txt +60 -0
- package/docs/design/tidy-0.35.1/frames/real-step-overview-finished-55.txt +60 -0
- package/docs/design/tidy-0.35.1/frames/real-step-overview-running-120.txt +60 -0
- package/docs/design/tidy-0.35.1/frames/real-step-overview-running-200.txt +60 -0
- package/docs/design/tidy-0.35.1/frames/real-step-overview-running-55.txt +60 -0
- package/docs/design/tidy-0.35.1/frames/real-task-120.txt +60 -0
- package/docs/design/tidy-0.35.1/frames/real-task-200.txt +60 -0
- package/docs/design/tidy-0.35.1/frames/real-task-55.txt +60 -0
- package/docs/design/tidy-0.35.1/frames/rendered-failed-120.txt +30 -0
- package/docs/design/tidy-0.35.1/frames/rendered-failed-200.txt +28 -0
- package/docs/design/tidy-0.35.1/frames/rendered-failed-55.txt +30 -0
- package/docs/design/tidy-0.35.1/frames/rendered-finished-120.txt +30 -0
- package/docs/design/tidy-0.35.1/frames/rendered-finished-200.txt +28 -0
- package/docs/design/tidy-0.35.1/frames/rendered-finished-55.txt +30 -0
- package/docs/design/tidy-0.35.1/frames/rendered-running-120.txt +30 -0
- package/docs/design/tidy-0.35.1/frames/rendered-running-200.txt +28 -0
- package/docs/design/tidy-0.35.1/frames/rendered-running-55.txt +30 -0
- package/docs/design/tidy-0.35.1/frames/run-120.txt +13 -0
- package/docs/design/tidy-0.35.1/frames/run-200.txt +16 -0
- package/docs/design/tidy-0.35.1/frames/run-55.txt +21 -0
- package/docs/design/tidy-0.35.1/frames/step-detail-120.txt +30 -0
- package/docs/design/tidy-0.35.1/frames/step-detail-200.txt +29 -0
- package/docs/design/tidy-0.35.1/frames/step-detail-55.txt +46 -0
- package/docs/design/tidy-0.35.1/frames/step-overview-120.txt +28 -0
- package/docs/design/tidy-0.35.1/frames/step-overview-200.txt +29 -0
- package/docs/design/tidy-0.35.1/frames/step-overview-55.txt +30 -0
- package/docs/design/tidy-0.35.1/frames/task-120.txt +20 -0
- package/docs/design/tidy-0.35.1/frames/task-200.txt +21 -0
- package/docs/design/tidy-0.35.1/frames/task-55.txt +22 -0
- package/docs/design/tidy-0.35.1/run-v2/README.md +58 -0
- package/docs/design/tidy-0.35.1/run-v2/after-200.txt +38 -0
- package/docs/design/tidy-0.35.1/run-v2/after-55.txt +49 -0
- package/docs/design/tidy-0.35.1/run-v2/before-200.txt +50 -0
- package/docs/design/tidy-0.35.1/run-v2/before-55.txt +60 -0
- package/docs/design/tidy-0.35.1/step-v2/0.35.2-before-detail-finished-200.txt +60 -0
- package/docs/design/tidy-0.35.1/step-v2/0.35.2-before-detail-finished-55.txt +60 -0
- package/docs/design/tidy-0.35.1/step-v2/0.35.2-before-overview-finished-200.txt +38 -0
- package/docs/design/tidy-0.35.1/step-v2/0.35.2-before-overview-finished-55.txt +71 -0
- package/docs/design/tidy-0.35.1/step-v2/0.35.2-detail-finished-200.txt +60 -0
- package/docs/design/tidy-0.35.1/step-v2/0.35.2-detail-finished-55.txt +60 -0
- package/docs/design/tidy-0.35.1/step-v2/0.35.2-detail-tool-open-200.txt +60 -0
- package/docs/design/tidy-0.35.1/step-v2/0.35.2-detail-tool-open-55.txt +60 -0
- package/docs/design/tidy-0.35.1/step-v2/0.35.2-overview-finished-200.txt +33 -0
- package/docs/design/tidy-0.35.1/step-v2/0.35.2-overview-finished-55.txt +42 -0
- package/docs/design/tidy-0.35.1/step-v2/0.35.2-overview-running-200.txt +25 -0
- package/docs/design/tidy-0.35.1/step-v2/0.35.2-overview-running-55.txt +42 -0
- package/docs/design/tidy-0.35.1/step-v2/README.md +124 -0
- package/docs/design/tidy-0.35.1/step-v2/desktop-detail.txt +44 -0
- package/docs/design/tidy-0.35.1/step-v2/desktop-finished.txt +38 -0
- package/docs/design/tidy-0.35.1/step-v2/desktop-running.txt +24 -0
- package/docs/design/tidy-0.35.1/step-v2/phone-detail.txt +81 -0
- package/docs/design/tidy-0.35.1/step-v2/phone-finished.txt +71 -0
- package/docs/design/tidy-0.35.1/step-v2/phone-running.txt +46 -0
- package/docs/dynamic-workflow-handoff.md +2 -2
- package/docs/dynamic-workflow-v2-execution-plan.md +1 -1
- package/docs/experiments/2026-08-28-trending-ai-autonomy.md +31 -61
- package/docs/experiments/2026-08-29-dogfood-bullswarm-builds-bullswarm.md +3 -3
- package/docs/experiments/2026-09-06-caller-planner-evaluation.md +3 -3
- package/docs/guide/concepts.md +1 -1
- package/docs/guide/cost.md +128 -3
- package/docs/guide/observing.md +354 -125
- package/docs/guide/routing.md +13 -11
- package/docs/handoff-2026-09-13-custom-provider-config.md +2 -2
- package/docs/integration-audit-2026-08-31.md +32 -32
- package/docs/integrations/claude-code.md +1 -1
- package/docs/integrations/issue-watcher.md +3 -3
- package/docs/plans/0.33.1.goal.txt +1 -1
- package/docs/plans/0.33.1.program.json +13 -13
- package/docs/plans/0.33.2-stats.goal.txt +1 -1
- package/docs/plans/0.33.2-stats.program.json +7 -7
- package/docs/plans/cost-0.34.0.program.json +10 -10
- package/docs/plans/cost-audit.goal.txt +4 -4
- package/docs/plans/cost-audit.program.json +8 -8
- package/docs/plans/dashboard-0.33-fidelity.goal.txt +1 -1
- package/docs/plans/dashboard-0.33-fidelity.md +5 -5
- package/docs/plans/dashboard-0.33-fidelity.program.json +10 -10
- package/docs/plans/dashboard-0.33-fidelity.rev2.json +13 -13
- package/docs/plans/dashboard-0.33-review.goal.txt +8 -8
- package/docs/plans/dashboard-0.33-review.program.json +6 -6
- package/docs/plans/dashboard-0.33.goal.txt +1 -1
- package/docs/plans/dashboard-0.33.program.json +10 -10
- package/docs/plans/dashboard-split.program.json +5 -5
- package/docs/plans/fidelity.program.json +2 -2
- package/docs/plans/handoff.goal.txt +1 -1
- package/docs/plans/handoff.program.json +5 -5
- package/docs/plans/handoff.rev2.json +5 -5
- package/docs/plans/handoff.rev3.json +6 -6
- package/docs/plans/meter-ledger.program.json +6 -6
- package/docs/plans/providers-0.35.2.goal.txt +11 -0
- package/docs/plans/providers-0.35.2.program.json +160 -0
- package/docs/plans/step-page-0.35.0-build.goal.txt +1 -1
- package/docs/plans/step-page-0.35.0-build.program.json +7 -7
- package/docs/plans/step-page-0.35.0-design.goal.txt +2 -2
- package/docs/plans/step-page-0.35.0-design.program.json +3 -3
- package/docs/plans/tidy-0.35.1-colour.goal.txt +8 -0
- package/docs/plans/tidy-0.35.1-colour.program.json +88 -0
- package/docs/plans/tidy-0.35.1.goal.txt +11 -0
- package/docs/plans/tidy-0.35.1.program.json +302 -0
- package/docs/reference/cli.md +151 -7
- package/docs/reference/configuration.md +36 -1
- package/docs/reference/program.md +31 -1
- package/docs/reference/providers.md +80 -2
- package/docs/reference/result.md +78 -1
- package/docs/studies/cost-audit-2026-09-18/budget-page-audit.md +50 -51
- package/docs/studies/cost-audit-2026-09-18/claude-actual-vs-recorded.md +14 -14
- package/docs/studies/cost-audit-2026-09-18/cost-fix-plan.md +5 -5
- package/docs/studies/cost-audit-2026-09-18/estimator-audit.md +13 -13
- package/docs/studies/cost-audit-2026-09-18/grok-codex-actual-vs-recorded.md +38 -39
- package/docs/studies/cost-audit-2026-09-18/step-timeline-120.txt +2 -2
- package/docs/studies/cost-audit-2026-09-18/step-timeline-55.txt +2 -2
- package/docs/studies/cost-audit-2026-09-18/step-timeline-design.md +8 -8
- package/docs/studies/portal-token-diet.md +24 -24
- package/docs/workflow-agent-usability-audit-2026-08-27.md +4 -4
- package/mods/bullswarm/.claude-plugin/plugin.json +1 -1
- package/mods/bullswarm/README.md +4 -4
- package/mods/bullswarm/hooks/pane.tsx +120 -97
- package/mods/bullswarm/hooks/register.ts +62 -2
- package/mods/bullswarm/hooks/runs.ts +17 -1
- package/mods/bullswarm/hooks/step.ts +386 -0
- package/package.json +1 -1
- package/providers/contrib/command-code/connector.json +48 -3
- package/providers/contrib/command-code/provider.mjs +22 -1
- package/providers/contrib/opencode/connector.json +16 -0
- package/providers/contrib/opencode/provider.mjs +15 -0
- package/skill/SKILL.md +112 -21
- package/skill/references/operations.md +162 -17
- package/skill/references/program.md +28 -1
- package/src/cli.js +141 -17
- package/src/help.js +178 -16
- package/src/home-cli.js +170 -18
- package/src/lib/assignments.js +3 -1
- package/src/lib/cli-flags.js +9 -2
- package/src/lib/glyphs.js +2 -0
- package/src/lib/prices.js +16 -5
- package/src/lib/quota.js +460 -13
- package/src/lib/retention.js +541 -0
- package/src/lib/route.js +93 -20
- package/src/lib/stale.js +457 -0
- package/src/lib/state.js +141 -6
- package/src/lib/subscription-cost.js +6 -4
- package/src/lib/tasks.js +3 -0
- package/src/lib/transcripts/claude-code.js +146 -30
- package/src/lib/transcripts/codex.js +148 -30
- package/src/lib/transcripts/command-code.js +417 -0
- package/src/lib/transcripts/index.js +142 -32
- package/src/lib/transcripts/indexing.js +134 -0
- package/src/lib/transcripts/opencode.js +568 -0
- package/src/lib/usage-basis.js +44 -0
- package/src/lib/watch.js +404 -72
- package/src/meters/framework.js +1 -1
- package/src/provider-cli.js +13 -0
- package/src/providers/_schema.json +10 -2
- package/src/providers/claude-code/connector.json +15 -1
- package/src/providers/claude-code/provider.mjs +9 -1
- package/src/providers/codex/connector.json +17 -13
- package/src/providers/codex/provider.mjs +9 -1
- package/src/providers/grok/connector.json +35 -1
- package/src/providers/grok/provider.mjs +9 -1
- package/src/setup.js +8 -0
- package/src/workflow/action-validator.js +54 -7
- package/src/workflow/budget-model.js +71 -9
- package/src/workflow/budget-view.js +18 -0
- package/src/workflow/cli.js +125 -3
- package/src/workflow/dash-kit.js +49 -34
- package/src/workflow/dashboard.js +535 -98
- package/src/workflow/history-view.js +125 -35
- package/src/workflow/home-model.js +627 -8
- package/src/workflow/home-view.js +793 -118
- package/src/workflow/reconcile.js +832 -0
- package/src/workflow/reprice.js +220 -49
- package/src/workflow/rollup.js +112 -9
- package/src/workflow/run-model.js +525 -8
- package/src/workflow/run-view.js +933 -110
- package/src/workflow/runs-view.js +14 -4
- package/src/workflow/spend-facts.js +152 -0
- package/src/workflow/stat-kit.js +189 -14
- package/src/workflow/stats-model.js +101 -20
- package/src/workflow/stats-view.js +259 -38
- package/src/workflow/step-model.js +1775 -64
- package/src/workflow/step-view.js +1161 -472
- package/src/workflow/task-step.js +346 -0
- package/src/workflow/time-box.js +276 -0
- package/src/workflow/usage-preference.js +22 -0
- package/src/workflow/usage-view.js +2 -1
- package/src/workflow/v2-dispatch.js +148 -6
- package/src/workflow/v2-outcome.js +153 -11
- package/src/workflow/v2-planner.js +17 -4
- package/src/workflow/v2-revision.js +4 -1
- package/src/workflow/v2-runtime.js +370 -19
- package/src/workflow/v2-state.js +153 -2
- package/src/workflow/verify-rounds.js +828 -0
- package/src/workflow/watch-cli.js +206 -17
- package/docs/design/owner-review-2026-09-18/budget-page-current-0.33.1.png +0 -0
- package/docs/design/owner-review-2026-09-18/budget.jpeg +0 -0
- package/docs/design/owner-review-2026-09-18/history-unknown-project.jpeg +0 -0
- package/docs/design/owner-review-2026-09-18/home-licence-today-missing-codex-0.33.1.png +0 -0
- package/docs/design/owner-review-2026-09-18/home-phone-landscape.jpeg +0 -0
- package/docs/design/owner-review-2026-09-18/home-phone-portrait.jpeg +0 -0
- package/docs/design/owner-review-2026-09-18/round2-budget-phone.jpeg +0 -0
- package/docs/design/owner-review-2026-09-18/round2-trends-30d-horizontal-split.png +0 -0
- package/docs/design/owner-review-2026-09-18/round2-trends-desktop.png +0 -0
- package/docs/design/owner-review-2026-09-18/round2-trends-phone.png +0 -0
- package/docs/design/owner-review-2026-09-18/run-page-budget.jpeg +0 -0
- package/docs/design/owner-review-2026-09-18/runs-list.jpeg +0 -0
- package/docs/design/owner-review-2026-09-18/stats-pools.jpeg +0 -0
- package/docs/design/owner-review-2026-09-18/stats-trends-spent.jpeg +0 -0
- package/docs/design/owner-review-2026-09-19/stats-overview-current.png +0 -0
- package/docs/design/prototype-shots/Screenshot 2026-09-17 at 9.14.40/342/200/257AM.png +0 -0
- package/docs/design/prototype-shots/Screenshot 2026-09-17 at 9.15.23/342/200/257AM.png +0 -0
- package/docs/design/prototype-shots/Screenshot 2026-09-17 at 9.16.07/342/200/257AM.png +0 -0
- package/docs/design/prototype-shots/Screenshot 2026-09-17 at 9.16.18/342/200/257AM.png +0 -0
- package/docs/design/stats-frames-0.33.2/budget-120.png +0 -0
- package/docs/design/stats-frames-0.33.2/budget-55.png +0 -0
- package/docs/design/stats-frames-0.33.2/stats-spending-120.png +0 -0
- package/docs/studies/cost-audit-2026-09-18/owner-budget-page-2026-09-18.png +0 -0
package/CHANGELOG.md
CHANGED
|
@@ -1,5 +1,313 @@
|
|
|
1
1
|
# bullswarm changelog
|
|
2
2
|
|
|
3
|
+
## 0.35.2 — the Step page reads like a transcript; prices without a command, honest totals, a home that stops growing
|
|
4
|
+
|
|
5
|
+
- step: detail is now a scrollable transcript of every turn in order: the full
|
|
6
|
+
response followed by one row per command or tool call; opening a turn or
|
|
7
|
+
tool row still reaches every captured event field, argument, result and exit.
|
|
8
|
+
- step: overview opens on the latest ten turns on desktop and five on phones,
|
|
9
|
+
newest at the bottom, with one `turns 1–N · … · click for detail` line for
|
|
10
|
+
earlier work; a followed live window slides until the reader moves it.
|
|
11
|
+
- step: `overview · detail` sits in the activity heading, right after the
|
|
12
|
+
heading word, each word a click target; `v`, the footer and Help name the
|
|
13
|
+
same views, and standalone tasks use the same page.
|
|
14
|
+
- step: clicking a turn head toggles it exactly like Enter, and hover lights
|
|
15
|
+
only the turn's text rather than its padding or desktop side column.
|
|
16
|
+
- run: each timeline attempt row opens Step on that exact attempt; step and
|
|
17
|
+
phase rows that name no attempt continue to open the latest one.
|
|
18
|
+
- grok: `tool_call_update` captures merge into the `tool_call` with the same id,
|
|
19
|
+
so a call has one named row and one count instead of extra nameless `agent`
|
|
20
|
+
rows, while every update remains reachable from transcript detail.
|
|
21
|
+
- providers: every shipped connector declares `eventStream.toolKinds`, mapping
|
|
22
|
+
its tool vocabulary to command, read, search, edit or other; the Step model
|
|
23
|
+
contains no provider-specific tool-name table.
|
|
24
|
+
- design: the Step v2 record adds transcript, latest-turn-window and toggle
|
|
25
|
+
rules with before/after frames at 55 and 200 columns; committed real text and
|
|
26
|
+
colour frames are regenerated from scrubbed fixtures.
|
|
27
|
+
- mod: the Claude Code Step pane draws the same v2 header, turns, result, task,
|
|
28
|
+
cost, phone order, colours, transcript toggle and latest-turn window from
|
|
29
|
+
`action show --json`, for workflow steps and standalone tasks alike.
|
|
30
|
+
- mod: Usage opens with `u` or a click even with no workflow running, and back
|
|
31
|
+
returns to idle; validated temporary chunks carry large Step records within
|
|
32
|
+
hook limits and are removed immediately.
|
|
33
|
+
- reprice: an incremental reconciler prices every terminal attempt that is
|
|
34
|
+
still unknown or a bytes/4 estimate, with no user command. The kernel prices
|
|
35
|
+
its own run at quiet boundaries and before the result and rollup are
|
|
36
|
+
written; a single `bullswarm run` whose usage is unmeasured starts a detached
|
|
37
|
+
pass 30 s after it ends (only when the pool's provider keeps transcripts);
|
|
38
|
+
the dashboard starts one detached, throttled pass (one per 5 minutes, one at
|
|
39
|
+
a time) after its first paint and says `pricing N older records…` while it
|
|
40
|
+
runs. A per-attempt ledger in `pricing/reconcile.json` retries recent
|
|
41
|
+
attempts after 1 and 10 minutes, gives older ones one try, and reopens an
|
|
42
|
+
attempt only when a new or grown transcript covers its window, so a second
|
|
43
|
+
pass does nothing. The automatic pass never makes an attempt worse: no match
|
|
44
|
+
or an ambiguous match leaves it as it was. `workflow reprice` stays the
|
|
45
|
+
manual full backfill and gains `--incremental` (with `--trigger`,
|
|
46
|
+
`--transcript-home`, `--delay-ms`). On a copy of a real home the first pass
|
|
47
|
+
priced 165 of 281 eligible attempts; the second took milliseconds and wrote
|
|
48
|
+
nothing.
|
|
49
|
+
- transcripts: Codex and Claude matching tries the session id, then the task
|
|
50
|
+
text (the task-file path or exact text in the first user message), then cwd
|
|
51
|
+
+ time window; more than one candidate stays unknown, and the cwd/time step
|
|
52
|
+
skips transcripts that quote a different task file. On the same copy the
|
|
53
|
+
formerly ambiguous attempts resolved by task text: codex 100, claude-code 19,
|
|
54
|
+
claude-code:acme 14. Indexes store prompt paths and hashes, never the text.
|
|
55
|
+
Single tasks are priced too; a task's window now ends at its `endedAt`.
|
|
56
|
+
- capture: each attempt records an immutable `capture` block when its worker
|
|
57
|
+
exits — provider session id and its source, model, exclusive token classes,
|
|
58
|
+
provider-reported cost, exit code and signal — before the end meter read and
|
|
59
|
+
the transcript lookup, and persists it at once. Later pricing may upgrade an
|
|
60
|
+
estimate but never downgrades provider-reported usage. Planner and scout
|
|
61
|
+
attempts are captured too.
|
|
62
|
+
- grok: the final `end` event is decoded (session id, token classes, cost),
|
|
63
|
+
proven on a real grok 1.0.13 stream checked in scrubbed as
|
|
64
|
+
`tests/fixtures/stream/grok-capture.jsonl`; `"rate limit"` is no longer a
|
|
65
|
+
grok auth signature, so a grok 429 is a throttle, not a 10-minute auth pause.
|
|
66
|
+
- command-code: launched without `--no-session`, so its session transcript
|
|
67
|
+
exists afterwards and the reader matches it by session id.
|
|
68
|
+
- money: every money total on Home, Runs, Run, Stats and Budget that includes
|
|
69
|
+
unmeasured attempts reads `at least $X · N unmeasured`; a scope with no
|
|
70
|
+
recorded amount reads `api unknown`. Budget prints a `spent …` line per pool,
|
|
71
|
+
and the Stats hover, axis and coverage note use the same words. One month
|
|
72
|
+
length, `DAYS_PER_MONTH = 30.4375` (365.25/12), is used everywhere.
|
|
73
|
+
- home: the window-share cell shows the measured share from
|
|
74
|
+
`calibration/<pool>.json` (`5.0% measured`) when samples attribute a window
|
|
75
|
+
drop to today's runs, the labelled pace estimate only when none do, and a
|
|
76
|
+
dash otherwise. The recent list shows finished runs only — a reopened run
|
|
77
|
+
leaves it — and each row's mark is the Run page header's mark; a
|
|
78
|
+
non-terminal run's index row no longer carries a finish time.
|
|
79
|
+
- retention: `state.json.retention` `{ "enabled": true, "workspacesDays": 7 }`
|
|
80
|
+
removes the `workspaces/` copies inside terminal runs older than the limit,
|
|
81
|
+
in a detached background sweep started by the kernel, watch completion and
|
|
82
|
+
the dashboard (at most every 6 hours, one at a time, only when a run holds a
|
|
83
|
+
workspace copy). Records, reports, streams, task/out markdown, interrupted
|
|
84
|
+
runs and any leased run are never touched. `bullswarm home prune
|
|
85
|
+
[--dry-run|--yes] [--days n]` lists or removes the same set with bytes, and
|
|
86
|
+
`bullswarm home status` shows the policy, bytes on disk and the last prune
|
|
87
|
+
and reprice results.
|
|
88
|
+
- pauses: a limit notice pauses a pool for quota only on proof — the pool's
|
|
89
|
+
own meter at 95% or more on a running window, or a provider line that says a
|
|
90
|
+
usage window is spent and names its reset. Everything else, such as
|
|
91
|
+
`Rate limit exceeded. Please wait a moment and try again.`, is a transient
|
|
92
|
+
throttle: the same pool is retried after 20 s and 60 s (or the wait it
|
|
93
|
+
named), then the attempt moves on, and the pool is never paused. `pools`
|
|
94
|
+
prints `PAUSED until <time> · <proof> · provider: "<line>" · meter then: … ·
|
|
95
|
+
lift now: bullswarm pools resume <pool>`; `bullswarm pools resume <pool>`
|
|
96
|
+
lifts a pause, and `bullswarm strategy set-pausing off|on` turns automatic
|
|
97
|
+
pausing off or back on.
|
|
98
|
+
- pauses: `bullswarm strategy set-pausing off` now stops every automatic
|
|
99
|
+
pause, not only quota — auth, the credential-group siblings an auth pause
|
|
100
|
+
benches with it, and the soft bench a second strike writes — so with the
|
|
101
|
+
switch off nothing is taken out of service by a command's own judgement;
|
|
102
|
+
routing still reads meters and a failed attempt still moves to another pool.
|
|
103
|
+
It is stored as `strategy.pausing: "off"`. `pools` opens with `automatic
|
|
104
|
+
pausing: off`, and `pools resume` still lifts a pause that was already in
|
|
105
|
+
place. Quota and auth signatures are matched against the provider's own
|
|
106
|
+
error channel only — its stderr, the events it flags as errors and its
|
|
107
|
+
terminal `result` record — never an assistant's reply or a tool result: on
|
|
108
|
+
2026-09-21 a pool was paused for a sentence an agent wrote about
|
|
109
|
+
`usage_credits_required` in its own report.
|
|
110
|
+
- routing: an evidence step may run on the pool that wrote the judged work
|
|
111
|
+
when that pool is urgent; the reason then says `independence waived: <pool>
|
|
112
|
+
resets in <clock>`. Independence is judged by model family, so
|
|
113
|
+
`claude-code` and `claude-code:acme` are the same writer; with nothing
|
|
114
|
+
urgent, independence stays the tie-breaker.
|
|
115
|
+
- watch: `bullswarm workflow watch <run> --until outcome|trouble` prints only
|
|
116
|
+
trouble lines (failed, rejected, paused, stalled, stale, steering) and the
|
|
117
|
+
outcome; `trouble` exits at the first one with a `next:` relaunch line. A
|
|
118
|
+
running attempt gets a stale score (quiet with no command running, no file
|
|
119
|
+
change while commands continue, the same command repeated, wall time over 3×
|
|
120
|
+
the expected minutes) and one `⚠ <step> looks stale: <reasons>` line.
|
|
121
|
+
`bullswarm workflow step restart <run> <step> [--pool <pool>]` stops that
|
|
122
|
+
attempt and requeues the step with its durable handoff; nothing restarts on
|
|
123
|
+
its own. The packaged skill teaches one background `--until trouble` watch
|
|
124
|
+
per run and one tool call per wake.
|
|
125
|
+
- providers: OpenCode and Command Code expose the same durable-transcript
|
|
126
|
+
reader contract as the first-class providers. OpenCode reads its SQLite
|
|
127
|
+
sessions read-only; Command Code reads persisted project JSONL only when a
|
|
128
|
+
session transcript exists and carries usage.
|
|
129
|
+
- reprice: transcript lookup follows the loaded provider registry, so
|
|
130
|
+
`opencode2*`, `opencode2:kaihk-*`, and `command-code` pools reach their
|
|
131
|
+
provider-owned readers without a hard-coded provider list. Ambiguous,
|
|
132
|
+
missing, and checkpoint-only records stay unknown rather than becoming
|
|
133
|
+
zero-cost attempts.
|
|
134
|
+
- command-code: the 116 historical attempts studied for this release used
|
|
135
|
+
`--no-session`, so their checkpoints are documented as non-recoverable
|
|
136
|
+
history.
|
|
137
|
+
- pricing: public model cards are retained only with a source and date (the
|
|
138
|
+
cards were checked 2026-09-20). `kaihk/*` and `opencode/union-alpha` relay
|
|
139
|
+
identifiers have no public card, so the underlying OpenAI card is not
|
|
140
|
+
substituted; observed Command Code models use the cited Command Code card.
|
|
141
|
+
- evidence: live OpenCode and Command Code streams captured on 2026-09-20 are
|
|
142
|
+
checked in as `tests/fixtures/streams/opencode-hello.jsonl` and
|
|
143
|
+
`tests/fixtures/streams/command-code-hello.jsonl`; each contrib connector
|
|
144
|
+
claims the `eventStream.usage` rules those captures prove.
|
|
145
|
+
- tests: detached reprice and prune children never recreate a home that was
|
|
146
|
+
deleted under them, and the OpenCode read-only test runs everywhere against
|
|
147
|
+
the exported rows instead of skipping without a local database.
|
|
148
|
+
- time box: every work and evidence step's task ends with a soft time box: the
|
|
149
|
+
box in minutes, the start clock, a wrap-up point at 70% of the box, and an
|
|
150
|
+
invitation to stop and report `## Done`, `## Not done` and `## Suggested next
|
|
151
|
+
step`. The box is the action's `timeBox` (whole minutes, 0-240), else
|
|
152
|
+
`defaults.timeBox`, else 1.5 × the median wall minutes of this home's
|
|
153
|
+
succeeded attempts for the pool and kind (5 or more attempts), else for the
|
|
154
|
+
kind, else 20, rounded to 5 and kept within 10-60; `opencode` attempts never
|
|
155
|
+
feed it, and `timeBox: 0` leaves the paragraph out. It is a guide: timeouts,
|
|
156
|
+
stall detection, cancellation and routing are unchanged and nothing stops at
|
|
157
|
+
the box. On the fixture home the (codex, implement) pair has a median of
|
|
158
|
+
18.86 minutes over 80 attempts, so its box is 30. In an experiment on one
|
|
159
|
+
task (two runs per arm) a 15-minute box ran 10m54s and 10m03s against 23m30s
|
|
160
|
+
and 48m44s without one, with honest partial reports: a direction, not a
|
|
161
|
+
measurement.
|
|
162
|
+
- early return: a work step whose `## Not done` lists items still succeeds and
|
|
163
|
+
records `returnedEarly` with the count and the items. The Step page header and
|
|
164
|
+
the Run timeline row read `returned early · N not done`, the Step page shows
|
|
165
|
+
`box 20m · ran 34m` when an attempt ran past its box, `workflow watch` prints
|
|
166
|
+
`◐ <step> returned early · N not done`, and the items reach the verifiers with
|
|
167
|
+
the rest of the evidence.
|
|
168
|
+
- verify rounds: a failed verify no longer waits for the caller. The kernel runs
|
|
169
|
+
a bounded repair loop of at most 3 verify rounds (`defaults.verifyRounds`
|
|
170
|
+
1-3, default 3, 1 = the old single round): round 1 judges every requirement,
|
|
171
|
+
each failure starts a kernel `repair-<n>` step built from the verifier's
|
|
172
|
+
evidence, the not-done items and the handoffs of the steps that affect it,
|
|
173
|
+
round 2 re-checks the failures and looks for regressions and the same defect
|
|
174
|
+
elsewhere, and round 3 is final closure. A requirement that passed carries
|
|
175
|
+
forward and is judged again only when a repair changed a file its evidence
|
|
176
|
+
names. There is never a fourth round and a program still cannot declare a
|
|
177
|
+
repair step; the planner contract allows `timeBox` and says the kernel adds
|
|
178
|
+
`repair-<n>` and `verify-round-<n>`.
|
|
179
|
+
- verify rounds: the run ends `completed · verified`, or `completed · not
|
|
180
|
+
verified · verify rounds 3/3` with a caller-decision block (each requirement
|
|
181
|
+
still failing, its latest evidence and one suggested next step) in
|
|
182
|
+
`workflow runs result <run> --json --summary`. The result also lists each
|
|
183
|
+
verify round and each repair with its wall minutes, pool and cost. The Run
|
|
184
|
+
page, Home, Runs and `workflow watch` show each round and repair; the events
|
|
185
|
+
are `workflow.verify-round` and `workflow.repair`. A plan revision during the
|
|
186
|
+
loop is still accepted and never adds or refunds a round; its
|
|
187
|
+
`defaults.verifyRounds` sets the cap for the rest of the run, and a revision
|
|
188
|
+
that changes only the cap is accepted.
|
|
189
|
+
- docs: the skill, its operations and program references, the program and result
|
|
190
|
+
references and `docs/design/step-economy-0.35.2/README.md` teach `timeBox`,
|
|
191
|
+
`verifyRounds`, the early-return report and the caller-decision block; the
|
|
192
|
+
skill's "Observe and judge" no longer tells the caller to hand-add a fix step
|
|
193
|
+
for an ordinary failed check.
|
|
194
|
+
|
|
195
|
+
## 0.35.1 — a calmer dashboard with honest active time
|
|
196
|
+
|
|
197
|
+
- durations: Home, Runs, Run, Step, and Stats use the union-based
|
|
198
|
+
`minutes.active` duration; `minutes.span` remains secondary, worker-minutes
|
|
199
|
+
still sum attempt clocks, and `workflow reprice` corrects stored terminal
|
|
200
|
+
minute fields.
|
|
201
|
+
- home: today leads with up to three active-or-most-recent run cards, while the
|
|
202
|
+
licence block uses one plain-word `pool · worker-minutes · weekly share · API
|
|
203
|
+
· subscription` row per pool and keeps Running and Budget visible.
|
|
204
|
+
- home: from 120 columns the three cards sit side by side, with the licence
|
|
205
|
+
block beneath them up to 159 columns and to their right from 160; below 120
|
|
206
|
+
they stack full width with the licence block beneath.
|
|
207
|
+
- home: the last-7-days band and the five-row recent list are back below
|
|
208
|
+
`budget · this week`, with the `Last 30 days`/`All time` toggle, the spend
|
|
209
|
+
chart, the by-pool/by-model/by-project lists, and the summary lines; the full
|
|
210
|
+
run list still lives only on Runs.
|
|
211
|
+
- run: the Run page now follows the Step grammar — `done of total` with running
|
|
212
|
+
and waiting ids, active-of-span clocks, a per-pool attempt mix, a phone glyph
|
|
213
|
+
strip (`p` opens phase boxes), and a live block backed by the selected
|
|
214
|
+
attempt's latest stream turn.
|
|
215
|
+
- run: spend is an honest partial subtotal (`at least` when attempts are still
|
|
216
|
+
running or unmeasured, `≈` for estimates) with measured/estimated/running/
|
|
217
|
+
unmeasured coverage and per-pool API splits; the timeline keeps phase rules
|
|
218
|
+
and one routed attempt row per worker, folds middle phases, and removes
|
|
219
|
+
licence bars, `so far`, ETA, and started/completed filler rows.
|
|
220
|
+
- money: a run, day, pool or period whose attempts were only partly priced now
|
|
221
|
+
shows the recorded subtotal marked `≈` with its coverage instead of reading
|
|
222
|
+
as unrecorded; strict totals stay strict and a subtotal is never summed into
|
|
223
|
+
one.
|
|
224
|
+
- stats: `Median run` and `Longest run` fall back per record to the recorded
|
|
225
|
+
span and say so, and the Spending chart fills its panel on the desktop grid
|
|
226
|
+
instead of drawing six narrow bars beside it.
|
|
227
|
+
- charts: a column chart sizes its tick gutter to the widest label the axis
|
|
228
|
+
will actually print, so a money axis reads `≈$160.00` rather than `≈$160.…`.
|
|
229
|
+
- step: the page is one header and four blocks — turns, result, task, cost. The
|
|
230
|
+
header says the verdict, the purpose, `pool · model · effort · reasoning`, one
|
|
231
|
+
clock (span only when it differs) and the route sentence once; the overview
|
|
232
|
+
prints one row per turn with only its non-zero counts (`35 commands`, `no
|
|
233
|
+
tools`), the counts at the end of the response on the desk and on their own
|
|
234
|
+
row on the phone, `Enter` expands a turn's full response and its tool rows
|
|
235
|
+
with each clock and measured duration, and the last turn points at the report
|
|
236
|
+
with `→ the report, shown under result`.
|
|
237
|
+
- step: the result card reads the report itself — its first lines, the diff's
|
|
238
|
+
changed paths, the step's `Shared-file requests` when its report has that
|
|
239
|
+
heading, and only the artifacts it left (`files` names the run directory once,
|
|
240
|
+
then `task · out · stream (N events) · diff`); the task card shows the
|
|
241
|
+
author's own prompt with `owns`/`after`/`affects` and the kernel wrapper's
|
|
242
|
+
size (`Enter on task: full text 2.4 KB · kernel wrapper 5.1 KB`); cost is two
|
|
243
|
+
plain-word rows (API rate and `<pool> plan`) with token classes under the
|
|
244
|
+
amount, `≈` for estimates, `—` with its reason for unknowns, and a closing
|
|
245
|
+
`measured from the … transcript` line.
|
|
246
|
+
- step: at 160 columns the activity takes the left column and result, task and
|
|
247
|
+
cost stack on the right so a finished step fits one screen; 120 narrows the
|
|
248
|
+
right column to 40; below that the page stacks result → activity → task →
|
|
249
|
+
cost, with `now` leading while the step runs. The footer carries the keys
|
|
250
|
+
once (`v detail (every event)`, `t filter`), and a live attempt says when its
|
|
251
|
+
cost will be measured instead of printing a guess.
|
|
252
|
+
- tasks: a single `bullswarm run` task uses the same Step model and view through
|
|
253
|
+
an adapter, preserving recorded fields and leaving unavailable effort,
|
|
254
|
+
verification, usage, money, and activity unavailable.
|
|
255
|
+
- step: an expanded turn shows its newest three tool rows with a
|
|
256
|
+
`↑ N earlier commands · Space page up` fold; `Space`, `PageUp` and `PageDown`
|
|
257
|
+
walk that window back through the older rows, and opening, collapsing or
|
|
258
|
+
changing the turn resets it.
|
|
259
|
+
- step: a tool that is still running keeps a spinner and its elapsed time, stays
|
|
260
|
+
the last expanded row and is counted in the collapsed turn's counts; a
|
|
261
|
+
sub-second tool prints no duration rather than `0s`, and a capture with no
|
|
262
|
+
summary falls back to its event kind instead of `summary unavailable`.
|
|
263
|
+
- step: the activity reader reads a capture's `.tail` file as well as its head,
|
|
264
|
+
merges the truncation marker and drops the overlap, so a long attempt's newest
|
|
265
|
+
events are visible instead of ending at the head cap.
|
|
266
|
+
- step: `t` cycles the activity lens turns → tools → errors → all and `f`
|
|
267
|
+
follows the tail on this page; Fleet keeps `f` everywhere else.
|
|
268
|
+
- codex: a `file_change` item keeps its path and its raw `changes` array, so the
|
|
269
|
+
Step page can name the kind and path of each change (capped at three) instead
|
|
270
|
+
of printing an empty summary.
|
|
271
|
+
- durations: every clock is h/m/s — a step or phase past an hour reads `2h04m`,
|
|
272
|
+
not `124m09s`, on the Run page, the Step page and the timeline alike.
|
|
273
|
+
- run: the `live` and `spend` columns are one aligned band at 120 columns and
|
|
274
|
+
wider; the running phase is never folded into the `↑ phases a–b` summary; and
|
|
275
|
+
the page keys and the wheel scroll the page body, so a timeline longer than
|
|
276
|
+
the terminal can be read.
|
|
277
|
+
- run: a paused run's live block still says `bullswarm workflow resume <id>
|
|
278
|
+
continues it` rather than leaving `last finished` to read as "any moment now".
|
|
279
|
+
- mod: `workflow tui --overview` keeps the goal preview and draws its milestone
|
|
280
|
+
rows flush to the pane border, the two shapes the Claude mod's pane parses.
|
|
281
|
+
- colour: the Step page, the Run page and the single-task page paint one meaning
|
|
282
|
+
per colour — green for work that went well (`✓`, `succeeded`, `verified`),
|
|
283
|
+
amber for work running now (`▶`, `●`, the spinner, `following ●`), red for
|
|
284
|
+
failures (`✗`, `failed`, an error count above zero), dim for the supporting
|
|
285
|
+
detail (clocks, counts, row labels, basis phrases, block rule dashes, footer
|
|
286
|
+
hints, pending marks), bold for identity and every money amount, and each pool
|
|
287
|
+
in the colour its bar has on Home. Response, task, report and command text
|
|
288
|
+
stay plain, the selected row is inverse, and in ASCII mode only bold, dim and
|
|
289
|
+
inverse remain.
|
|
290
|
+
- colour: the frames the records are reviewed against are rendered twice —
|
|
291
|
+
`docs/design/tidy-0.35.1/frames/real-*.txt` as text and
|
|
292
|
+
`docs/design/tidy-0.35.1/frames/colour/` with the escape codes kept, from the
|
|
293
|
+
one snapshot render (`scripts/render-tidy-0.35.1-frames.mjs --colour`), so a
|
|
294
|
+
reviewer can grep for a colour instead of trusting a screenshot.
|
|
295
|
+
- tasks: the single-task page names the task the way the Runs list does —
|
|
296
|
+
`<lane> task · <8-char id>` (`analyze task · 3155fb3c`) — instead of the full
|
|
297
|
+
UUID or a six-character prefix.
|
|
298
|
+
- home: a card's durations use the same h/m/s clock as the Step and Run pages
|
|
299
|
+
(`58m37s active of 58m38s`, `1h38m active of 1h39m`), never decimal minutes
|
|
300
|
+
like `active 725.95m`.
|
|
301
|
+
- runs: the page opens with the cursor on its first row — the first active run
|
|
302
|
+
when one is running — drawn inverse, so `Enter` opens the highlighted row, and
|
|
303
|
+
`Up` on the first row stays there.
|
|
304
|
+
- tests: the real-data tests and the frame script now run from a scrubbed
|
|
305
|
+
in-repo fixture, `tests/fixtures/home-351` (2.8 MB), built by
|
|
306
|
+
`scripts/build-test-home.mjs`: every id, time, status, pool, model, token
|
|
307
|
+
count, cost and event kind is as recorded, while goals, prompts, reports,
|
|
308
|
+
responses, commands and paths are marked placeholder text and project names
|
|
309
|
+
are aliases, so the suite needs no copy of anyone's home.
|
|
310
|
+
|
|
3
311
|
## 0.35.0 — the Step page tells the whole story
|
|
4
312
|
|
|
5
313
|
- step: the Step page leads with identity and verdict, the selected attempt,
|
|
@@ -62,7 +370,7 @@
|
|
|
62
370
|
colour codes, so every bar came out grey while the legend beside it stayed
|
|
63
371
|
coloured — a chart titled "by pool" painted all five pools identically.
|
|
64
372
|
- stats: no two series in one chart share a hue. Colours were hashed per name,
|
|
65
|
-
which put `claude-code:
|
|
373
|
+
which put `claude-code:acme` and `codex` on the same blue.
|
|
66
374
|
|
|
67
375
|
## 0.33.1 — say what a number is, or say you do not know
|
|
68
376
|
|
|
@@ -191,8 +499,8 @@
|
|
|
191
499
|
price exactly as it spends a declared one and labels which it is —
|
|
192
500
|
`detected` or `declared` — so an operator's own figure is never confused with
|
|
193
501
|
one bullswarm inferred. The plan table is keyed by provider, not by pool, so a
|
|
194
|
-
discovered per-account pool such as `claude-code:
|
|
195
|
-
`claude-code`'s plans; it used to ask for `claude-code:
|
|
502
|
+
discovered per-account pool such as `claude-code:acme` resolves against
|
|
503
|
+
`claude-code`'s plans; it used to ask for `claude-code:acme`'s own plans and
|
|
196
504
|
find nothing, which left exactly the pools with one login per account
|
|
197
505
|
without a price.
|
|
198
506
|
- Mod: the pane's `usage` button works with no run in flight — it renders the
|
|
@@ -258,7 +566,7 @@ drawn from the real pool figures of 2026-09-18.
|
|
|
258
566
|
- dashboard: the phone `Home` no longer hides data behind its width. The
|
|
259
567
|
`budget · this week` block draws every pool the desktop draws — four today —
|
|
260
568
|
or ends with `+N more` when the rows genuinely do not fit, and its names are
|
|
261
|
-
wide enough that `claude-code:
|
|
569
|
+
wide enough that `claude-code:acme` and `claude-code` read as two pools
|
|
262
570
|
rather than one truncated one. The by-pool, by-model and by-project lists
|
|
263
571
|
keep at least their top three rows (`+N more` for the rest) and always paint
|
|
264
572
|
the whole percentage, never `24…`. The summary figures sit one per line, so
|
|
@@ -833,7 +1141,7 @@ drawn from the real pool figures of 2026-09-18.
|
|
|
833
1141
|
from a reset the operator declares —
|
|
834
1142
|
`bullswarm strategy set-subscription <pool> --resets-at <iso|unknown>`.
|
|
835
1143
|
The Relay token API stopped returning `expires_at` for the `1` and `Moham`
|
|
836
|
-
wallets on 2026-09-03 (the
|
|
1144
|
+
wallets on 2026-09-03 (the project-n log holds 28 dated snapshots between
|
|
837
1145
|
2026-08-31 and 2026-09-03, then only nulls), so `opencode2` and
|
|
838
1146
|
`opencode2:relay-3` printed `unmetered` at 73.8% and 10.7% of their $50
|
|
839
1147
|
wallets and ranked as neutral. With a declared reset the used% stays the
|
|
@@ -866,8 +1174,8 @@ drawn from the real pool figures of 2026-09-18.
|
|
|
866
1174
|
— its effective surplus divided by the fraction of the window still to run —
|
|
867
1175
|
ahead of every pool whose window is not about to close, instead of on the
|
|
868
1176
|
surplus alone. At 2026-09-11 12:26 HKT the medium lane went to
|
|
869
|
-
`claude-code:
|
|
870
|
-
`grok` (+13.8 points, 2h02m and 1.2% left, urgency ~1150 against
|
|
1177
|
+
`claude-code:acme` (+22.9 points, 13h33m and 8.1% of its week left) over
|
|
1178
|
+
`grok` (+13.8 points, 2h02m and 1.2% left, urgency ~1150 against acme's
|
|
871
1179
|
~283), and grok's points expired unspent two hours later; the owner had been
|
|
872
1180
|
pinning grok by hand for such runs. Three states for an expiring pool:
|
|
873
1181
|
`urgent` (surplus still to spend and a pacing forecast — the reading plus
|
|
@@ -888,17 +1196,17 @@ drawn from the real pool figures of 2026-09-18.
|
|
|
888
1196
|
- routing: the 5-hour near-limit line is now clock-relative. A pool is
|
|
889
1197
|
deprioritized only when its 5h forecast is at/above 75% AND ahead of the
|
|
890
1198
|
share of the 5h window that has already elapsed. On 2026-09-10 at 22:19Z a
|
|
891
|
-
high-tier integrator skipped `claude-code:
|
|
1199
|
+
high-tier integrator skipped `claude-code:acme` (81% used, 23 minutes to the
|
|
892
1200
|
reset — 92.3% of the window elapsed, 88.1% projected) and
|
|
893
|
-
`claude-code:
|
|
894
|
-
account already ahead of its weekly pace, while
|
|
1201
|
+
`claude-code:initech` (75.3% projected, 85.7% elapsed) and went to the one
|
|
1202
|
+
account already ahead of its weekly pace, while acme still held 34% of its
|
|
895
1203
|
weekly quota unspent with 13% of the week left to spend it. Both pools now keep the lane:
|
|
896
1204
|
88.1% with 23 minutes left is a pool spending at the clock's pace, not a pool
|
|
897
1205
|
about to hit a wall. Pools with no `resets_at`, an unparsable one, or a reset
|
|
898
1206
|
already past keep the fixed 75% line, and the 90% burst gate is unchanged —
|
|
899
1207
|
it ignores the clock. The routing reason and the `candidates[]` rows say
|
|
900
1208
|
which case applied: `5h used 81% -> 88.1% projected, under the clock (92.3%
|
|
901
|
-
elapsed)`, `skipped near 5h limit (projected): claude-code:
|
|
1209
|
+
elapsed)`, `skipped near 5h limit (projected): claude-code:acme 88.1% (20.0%
|
|
902
1210
|
elapsed)`, and a new `fiveHourElapsedPct` field. `bullswarm pools` shows the
|
|
903
1211
|
same clock: `5h=81% (92% elapsed)`.
|
|
904
1212
|
- routing: 5-hour spend is clipped at the reset. Quota spent after the window
|
|
@@ -1604,7 +1912,7 @@ drawn from the real pool figures of 2026-09-18.
|
|
|
1604
1912
|
|
|
1605
1913
|
- All of it is visible after the fact. The routing reason names the in-flight
|
|
1606
1914
|
counts and projections that moved the pick (`5h used 30% -> 41% projected, 2
|
|
1607
|
-
in flight`, `skipped near 5h limit (projected):
|
|
1915
|
+
in flight`, `skipped near 5h limit (projected): acme 76%`, `forecast-gated
|
|
1608
1916
|
at/above 90%: …`, `preferred over busier: …`), every candidate row carries
|
|
1609
1917
|
`pace`, `effectiveSurplus`, `inflight`, `projectedFiveHourPct`,
|
|
1610
1918
|
`forecastFiveHourPct`, `projectedWeeklyPct`, `ratePerMinute`,
|
|
@@ -2660,7 +2968,7 @@ Adopts the driving mechanics of Claude Code's dynamic workflow (documented in
|
|
|
2660
2968
|
former 900-second planner/action timeout defaults.
|
|
2661
2969
|
- Fixed adaptive completion policy, current-action metadata, provider routing
|
|
2662
2970
|
history, usage aggregation, latest-worker verification, and truthful partial
|
|
2663
|
-
token/cost accounting found during the
|
|
2971
|
+
token/cost accounting found during the project-b battle test.
|
|
2664
2972
|
- Added connector-owned native JSONL event adapters for Codex, Claude, Grok,
|
|
2665
2973
|
Command Code, and OpenCode. Workflows now retain and display the latest three
|
|
2666
2974
|
semantic shell/read/edit/write/response actions for every active agent.
|
package/GOAL.md
CHANGED
|
@@ -2,7 +2,7 @@
|
|
|
2
2
|
|
|
3
3
|
> Historical (2026-08-21): accurate when written; see CHANGELOG for what changed since.
|
|
4
4
|
|
|
5
|
-
**Status:** HISTORICAL PROTOTYPE CHARTER · **Owner:**
|
|
5
|
+
**Status:** HISTORICAL PROTOTYPE CHARTER · **Owner:** dev · **Created:** 2026-08-21
|
|
6
6
|
|
|
7
7
|
## One sentence
|
|
8
8
|
|