bullswarm 0.38.7 → 0.38.8

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (85) hide show
  1. package/CHANGELOG.md +20 -0
  2. package/README.md +6 -1
  3. package/data/openrouter-benchmarks.json +11775 -11739
  4. package/docs/design/runs-table-0.35.4/frames/runs-120.txt +54 -54
  5. package/docs/design/runs-table-0.35.4/frames/runs-200.txt +54 -54
  6. package/docs/design/runs-table-0.35.4/frames/runs-55.txt +54 -54
  7. package/docs/design/tidy-0.35.1/frames/colour/real-home-200.txt +17 -17
  8. package/docs/design/tidy-0.35.1/frames/colour/real-home-55.txt +11 -11
  9. package/docs/design/tidy-0.35.1/frames/colour/real-run-finished-200.txt +3 -3
  10. package/docs/design/tidy-0.35.1/frames/colour/real-run-finished-55.txt +3 -3
  11. package/docs/design/tidy-0.35.1/frames/colour/real-run-running-200.txt +3 -3
  12. package/docs/design/tidy-0.35.1/frames/colour/real-run-running-55.txt +3 -3
  13. package/docs/design/tidy-0.35.1/frames/colour/real-step-detail-finished-200.txt +1 -1
  14. package/docs/design/tidy-0.35.1/frames/colour/real-step-overview-finished-200.txt +1 -1
  15. package/docs/design/tidy-0.35.1/frames/colour/real-step-overview-finished-55.txt +1 -1
  16. package/docs/design/tidy-0.35.1/frames/colour/real-task-200.txt +1 -1
  17. package/docs/design/tidy-0.35.1/frames/colour/real-task-55.txt +1 -1
  18. package/docs/design/tidy-0.35.1/frames/real-home-120.txt +18 -18
  19. package/docs/design/tidy-0.35.1/frames/real-home-200.txt +17 -17
  20. package/docs/design/tidy-0.35.1/frames/real-home-55.txt +11 -11
  21. package/docs/design/tidy-0.35.1/frames/real-run-finished-120.txt +3 -3
  22. package/docs/design/tidy-0.35.1/frames/real-run-finished-200.txt +3 -3
  23. package/docs/design/tidy-0.35.1/frames/real-run-finished-55.txt +3 -3
  24. package/docs/design/tidy-0.35.1/frames/real-run-running-120.txt +3 -3
  25. package/docs/design/tidy-0.35.1/frames/real-run-running-200.txt +3 -3
  26. package/docs/design/tidy-0.35.1/frames/real-run-running-55.txt +3 -3
  27. package/docs/design/tidy-0.35.1/frames/real-stats-model-120.txt +8 -8
  28. package/docs/design/tidy-0.35.1/frames/real-stats-model-200.txt +8 -8
  29. package/docs/design/tidy-0.35.1/frames/real-stats-model-55.txt +8 -8
  30. package/docs/design/tidy-0.35.1/frames/real-stats-spending-120.txt +8 -8
  31. package/docs/design/tidy-0.35.1/frames/real-stats-spending-200.txt +8 -8
  32. package/docs/design/tidy-0.35.1/frames/real-stats-spending-55.txt +7 -7
  33. package/docs/design/tidy-0.35.1/frames/real-step-detail-failed-120.txt +1 -1
  34. package/docs/design/tidy-0.35.1/frames/real-step-detail-failed-200.txt +1 -1
  35. package/docs/design/tidy-0.35.1/frames/real-step-detail-finished-120.txt +1 -1
  36. package/docs/design/tidy-0.35.1/frames/real-step-detail-finished-200.txt +1 -1
  37. package/docs/design/tidy-0.35.1/frames/real-step-overview-failed-120.txt +1 -1
  38. package/docs/design/tidy-0.35.1/frames/real-step-overview-failed-200.txt +1 -1
  39. package/docs/design/tidy-0.35.1/frames/real-step-overview-failed-55.txt +1 -1
  40. package/docs/design/tidy-0.35.1/frames/real-step-overview-finished-120.txt +1 -1
  41. package/docs/design/tidy-0.35.1/frames/real-step-overview-finished-200.txt +1 -1
  42. package/docs/design/tidy-0.35.1/frames/real-step-overview-finished-55.txt +1 -1
  43. package/docs/design/tidy-0.35.1/frames/real-task-120.txt +1 -1
  44. package/docs/design/tidy-0.35.1/frames/real-task-200.txt +1 -1
  45. package/docs/design/tidy-0.35.1/frames/real-task-55.txt +1 -1
  46. package/docs/design/tidy-0.35.1/step-v2/0.35.2-detail-finished-200.txt +1 -1
  47. package/docs/design/tidy-0.35.1/step-v2/0.35.2-detail-tool-open-200.txt +1 -1
  48. package/docs/design/tidy-0.35.1/step-v2/0.35.2-overview-finished-200.txt +1 -1
  49. package/docs/design/tidy-0.35.1/step-v2/0.35.2-overview-finished-55.txt +1 -1
  50. package/docs/guide/index.md +4 -0
  51. package/mcp/server.mjs +4 -1
  52. package/package.json +1 -1
  53. package/skill/SKILL.md +1 -0
  54. package/src/lib/rate-card-check.js +66 -0
  55. package/src/lib/usage-basis.js +10 -12
  56. package/src/providers/claude-code/connector-history.json +3 -1
  57. package/src/providers/claude-code/connector.json +1 -0
  58. package/src/providers/grok/connector.json +1 -0
  59. package/src/workflow/budget-model.js +18 -1
  60. package/src/workflow/budget-view.js +10 -5
  61. package/src/workflow/dashboard-nav.js +4 -2
  62. package/src/workflow/dashboard-page-budget.js +5 -1
  63. package/src/workflow/dashboard-page-fleet.js +1 -1
  64. package/src/workflow/dashboard-page-help.js +5 -4
  65. package/src/workflow/dashboard.js +37 -2
  66. package/src/workflow/fleet-view.js +20 -5
  67. package/src/workflow/history-view.js +37 -9
  68. package/src/workflow/home-model.js +5 -5
  69. package/src/workflow/home-view.js +57 -30
  70. package/src/workflow/legacy-verification.js +2 -3
  71. package/src/workflow/rollup.js +41 -1
  72. package/src/workflow/run-counts.js +2 -1
  73. package/src/workflow/run-model.js +7 -5
  74. package/src/workflow/run-view.js +53 -45
  75. package/src/workflow/spend-facts.js +17 -28
  76. package/src/workflow/stat-kit.js +4 -4
  77. package/src/workflow/stats-model.js +43 -2
  78. package/src/workflow/stats-view.js +33 -7
  79. package/src/workflow/step-model-basis-words.js +5 -3
  80. package/src/workflow/step-model-cost.js +17 -7
  81. package/src/workflow/step-model-presentation.js +7 -2
  82. package/src/workflow/step-view.js +10 -5
  83. package/src/workflow/usage-view.js +47 -2
  84. package/src/workflow/v2-outcome.js +2 -1
  85. package/src/workflow/v3-display.js +1 -1
package/CHANGELOG.md CHANGED
@@ -2,6 +2,26 @@
2
2
 
3
3
  ## Unreleased
4
4
 
5
+ ## 0.38.8 — dashboard money says only what its source supports; proof counts on v3 runs
6
+
7
+ - dashboard: money says only as much as its source supports, in fewer words. An API price built from the provider's own token counts — reported, or summed from its transcript — prints plainly (`$4.01`, was `≈$4.01`): exact arithmetic at the dated rate card, checked on 2026-10-02 against all 251 priced codex attempts and 163 of 167 Claude attempts' own reported cost. `~` stays on prices estimated from text size. A total that leaves attempts out reads `≥$9.52` everywhere (was `at least $9.52 api · 6 unmeasured` or `≈`); the Runs day header and the Run page's spend rule (`25 of 26 attempts priced`) keep the counts. Every plan amount is a share of an account-wide meter, so it always carries `≈`, and a meter that did not move shows `—` (was `≈ $0.000`).
8
+ - dashboard: a missing price is named, not a dash. A finished run with nothing priced reads `not priced` on its Home card and `unpriced` in the Runs list; the Step page says `not priced · no rate card` or `not priced · no token counts`. The Run page's rows read `API price` and `plan share`; the Home legend is `≈ ~ estimate · ≥ partly priced`, the chart footnote `66 unpriced`, Stats' note `Money · 83 of 149 attempts priced`, the period sentence `Your 15 runs in this period: ≥$299.87 at API prices`, and a Budget pool with no runs `spent —` (was `api unknown`).
9
+ - dashboard: the Run header and timeline say `1 attempt` and `1 phase` (was `1 attempts`).
10
+ - dashboard: v3 runs report proof, not a verified share. Run records carry a `proof` count (proven, answer checked, accepted by choice, unproven, the words of the end-of-run proof line); runs indexed before this read it once from their own state and result, and nothing on disk is rewritten. Home says `Steps: 30 proven · 50 answer checked · 4 accepted by choice · 19 unproven`; Stats' outcome panel says `Proven 30/103 steps` and `All proven 23/58 runs`, and its Model panel counts fully proven runs. The old `verified 0 (0%)` came from the one v2 run left in the period. A Run page header shows the run's proof line, and a Runs row's ✓ is amber when a step was unproven or accepted by choice.
11
+ - dashboard: one run count. Home reads `Runs: 56 · 43 one-step · 13 workflows` and Stats `56 runs (43 one-step · 13 workflows)`, matching the `Your 56 runs` sentence (the figure used to lead with the one-step count alone).
12
+ - dashboard: a step whose check failed and that the caller accepted reads `accepted by choice` on its Step page header and result, not `succeeded`, and its timeline row keeps the ✗ with `check failed · accepted by choice`.
13
+ - dashboard: the Runs list's clock is the finish the row is filed and ordered by (a run still going shows its start); it showed the start, so a run begun at 23:48 sat under the next day, out of order. Under 120 columns the goal keeps 30% of the row (the project narrows to 10 and `3/3 steps` to `3/3` first): at 80 it kept 9 cells.
14
+ - dashboard: Budget's header says how fresh the newest meter is (`meters read just now · 1 older`) and an old reading is named on its own pool (`codex · meter read 14h ago`); the header used the oldest pool's age, so one stale meter made every fresh one look 14 hours old.
15
+ - dashboard: Fleet shows the router's pick right now for each tier (`next pick → claude-code:w · build lane · most-behind capable pool…`), the pick `bullswarm run --dry-run` makes, with no probe or ledger entry; it is read only while Fleet is open, once a minute.
16
+ - dashboard: Home's running rule counts the runs waiting on you (`running · 1 needs you`).
17
+ - dashboard: on a v3 run, `o` and `v` say the run has no planner instead of opening the v2 planner panel (`Session · pending` on a finished run). Help matches the keys: Step's `a · o · p` are listed, `t` is the Run phase list under 100 columns, the spend chart opens Stats › Spending (there is no Trends tab), the old rebinding note is gone, and the help button reads `? help`.
18
+ - dashboard: the Step page's right column grows between 120 and 160 columns (53 at 140, was 40); Stats' licence panel says `reset 17 Oct`, so the shared label column no longer cuts every pool name; Home's recent rows show one money phrase (`at least $9.25 · 1 unmeasured`, the phrase alone under 100 columns) instead of a cut pair; Stats' licence reset date is the local day, as Budget says it; Home's pace word is `fast` as on Budget (JSON `paceWord` unchanged); `1 attempt with a meter reading`; a lower bound of nothing reads `at least $0`; a skipped gate reads `skipped · condition not met: review.passed is false`.
19
+ - dashboard: a pool's slice of a Stats spend bar is its known subtotal when one of its attempts went unpriced, as the bar itself is; its whole slice was dropped and drawn as `unallocated` (2 Oct: $41.87 of $45.97).
20
+ - pricing: `claude-sonnet-5-5` has a price row ($2 input, $2.50 / $4 cache writes, $0.20 cache reads, $10 output per million, from Anthropic's pricing page on 2026-10-02); 34 attempts in a week had recorded no API-equivalent price although Claude Code reported one. `bullswarm workflow reprice --apply` prices the runs already recorded. A new test pins every dispatched Claude model's row to the published rates.
21
+ - pricing: `scripts/check-rate-cards.mjs` compares each attempt's rate-card price with the cost its CLI reported, read only. On the owner's home it agrees for 161 of 165 Claude attempts (a resumed attempt is compared with its session's total, which is what Claude Code reports); grok's CLI reports 0.34x the published xAI price, so Bullswarm keeps the published one.
22
+ - mcp: the server exits only after its last reply has been written out. A reply larger than a pipe's 64 KB buffer (`bullswarm_pools` crossed it with this release's price rows) was cut off before its end when the client had already closed its input, so the call never answered.
23
+ - docs: README, the guide and the skill say briefly that issues and pull requests are welcome, optional, public, and that an agent asks its human first.
24
+
5
25
  ## 0.38.7 — honest reviews by default; dead sign-ins and instant refusals no longer cost a retry
6
26
 
7
27
  - pick: a pool whose latest meter read was refused for its sign-in (HTTP 401, or a 403 its provider calls an auth error) is out of the pick with the reason `sign-in failed (meter read 401)`, however old its cached numbers are, and `bullswarm pools` and the strategy preview show it as not ready (no tier's route). Nothing is remembered beyond that reading: when its hold ends the next pick reads again, and a successful read makes the pool pickable. A 429 or a network error on the meter read changes nothing.
package/README.md CHANGED
@@ -116,5 +116,10 @@ Each provider describes how to launch its CLI, read its quota, list its models a
116
116
  - [Authoring workflows](https://bulls-work.github.io/bullswarm/guide/workflows)
117
117
  - [The dashboard](https://bulls-work.github.io/bullswarm/guide/gallery) and [observing runs](https://bulls-work.github.io/bullswarm/guide/observing)
118
118
  - [Provider reference](https://bulls-work.github.io/bullswarm/reference/providers)
119
- - Contributing: [open an issue](https://github.com/Bulls-Work/bullswarm/issues) or [submit a pull request](https://github.com/Bulls-Work/bullswarm/pulls)
120
119
  - [MIT licence](LICENSE)
120
+
121
+ ## Reporting a problem or contributing
122
+
123
+ Bullswarm is open source and gets better through the cases people hit in real use. If you are a person or an agent and you find a real bug or gap in Bullswarm itself, an [issue](https://github.com/Bulls-Work/bullswarm/issues) or a [pull request](https://github.com/Bulls-Work/bullswarm/pulls) is welcome. Nobody is expected to send one.
124
+
125
+ If you are an agent, check with the person you work for before you post anything. Issues on this repo are public, so leave out task text, file paths, repo or project names, account names, keys and anything else from the person's work. The Bullswarm version, the command, and the error line are usually enough.