bullswarm 0.35.1 → 0.35.3

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (304) hide show
  1. package/CHANGELOG.md +223 -13
  2. package/GOAL.md +1 -1
  3. package/data/openrouter-benchmarks.json +8463 -8463
  4. package/docs/audits/2026-09-09-codebase-audit.md +11 -11
  5. package/docs/design/cost-0.34.0/README.md +1 -1
  6. package/docs/design/cost-0.34.0/meter-precision-study.md +19 -19
  7. package/docs/design/dashboard-prototype.html +20 -20
  8. package/docs/design/dashboard-split/README.md +30 -30
  9. package/docs/design/mod-step-v2/README.md +29 -0
  10. package/docs/design/mod-step-v2/failed-detail-120.txt +26 -0
  11. package/docs/design/mod-step-v2/failed-detail-55.txt +32 -0
  12. package/docs/design/mod-step-v2/failed-overview-120.txt +26 -0
  13. package/docs/design/mod-step-v2/failed-overview-55.txt +32 -0
  14. package/docs/design/mod-step-v2/finished-detail-120.txt +355 -0
  15. package/docs/design/mod-step-v2/finished-detail-55.txt +604 -0
  16. package/docs/design/mod-step-v2/finished-overview-120.txt +63 -0
  17. package/docs/design/mod-step-v2/finished-overview-55.txt +54 -0
  18. package/docs/design/mod-step-v2/running-detail-120.txt +24 -0
  19. package/docs/design/mod-step-v2/running-detail-55.txt +28 -0
  20. package/docs/design/mod-step-v2/running-overview-120.txt +24 -0
  21. package/docs/design/mod-step-v2/running-overview-55.txt +28 -0
  22. package/docs/design/mod-step-v2/usage-no-run-after.txt +19 -0
  23. package/docs/design/mod-step-v2/usage-no-run-before.txt +16 -0
  24. package/docs/design/pricing-0.35.2/auto-reprice-study.md +233 -0
  25. package/docs/design/pricing-0.35.2/retention-and-capture-study.md +178 -0
  26. package/docs/design/prototype-frames/README.md +1 -2
  27. package/docs/design/prototype-frames/history-120.tagged.txt +18 -18
  28. package/docs/design/prototype-frames/history-120.txt +18 -18
  29. package/docs/design/prototype-frames/history-55.tagged.txt +10 -10
  30. package/docs/design/prototype-frames/history-55.txt +10 -10
  31. package/docs/design/prototype-frames/home-120.tagged.txt +4 -4
  32. package/docs/design/prototype-frames/home-120.txt +4 -4
  33. package/docs/design/prototype-frames/home-55.tagged.txt +2 -2
  34. package/docs/design/prototype-frames/home-55.txt +2 -2
  35. package/docs/design/prototype-frames/run-120.tagged.txt +7 -7
  36. package/docs/design/prototype-frames/run-120.txt +7 -7
  37. package/docs/design/prototype-frames/run-55.tagged.txt +5 -5
  38. package/docs/design/prototype-frames/run-55.txt +5 -5
  39. package/docs/design/prototype-frames/stats-projects-120.tagged.txt +2 -2
  40. package/docs/design/prototype-frames/stats-projects-120.txt +2 -2
  41. package/docs/design/prototype-frames/stats-projects-55.tagged.txt +2 -2
  42. package/docs/design/prototype-frames/stats-projects-55.txt +2 -2
  43. package/docs/design/prototype-frames/step-120.tagged.txt +6 -6
  44. package/docs/design/prototype-frames/step-120.txt +6 -6
  45. package/docs/design/prototype-frames/step-55.tagged.txt +5 -5
  46. package/docs/design/prototype-frames/step-55.txt +5 -5
  47. package/docs/design/prototype-frames-0.33.1/budget-120-v2.txt +2 -2
  48. package/docs/design/prototype-frames-0.33.1/budget-120-v3.txt +2 -2
  49. package/docs/design/prototype-frames-0.33.1/budget-120.txt +4 -4
  50. package/docs/design/prototype-frames-0.33.1/budget-55-v2.txt +2 -2
  51. package/docs/design/prototype-frames-0.33.1/budget-55-v3.txt +2 -2
  52. package/docs/design/prototype-frames-0.33.1/budget-55.txt +2 -2
  53. package/docs/design/prototype-frames-0.33.1/budget-README.md +3 -3
  54. package/docs/design/prototype-frames-0.33.1/home-today-120-v2.txt +4 -4
  55. package/docs/design/prototype-frames-0.33.1/home-today-120.txt +5 -5
  56. package/docs/design/prototype-frames-0.33.1/home-today-55-v2.txt +3 -3
  57. package/docs/design/prototype-frames-0.33.1/home-today-55.txt +6 -6
  58. package/docs/design/prototype-frames-0.33.1/home-today-README.md +10 -10
  59. package/docs/design/providers-0.35.2/CHANGELOG-draft.md +27 -0
  60. package/docs/design/providers-0.35.2/README.md +172 -0
  61. package/docs/design/stats-frames-0.33.2/spending-120.txt +2 -2
  62. package/docs/design/stats-frames-0.33.2/spending-55.txt +4 -4
  63. package/docs/design/stats-refactor-0.33.2.md +10 -10
  64. package/docs/design/step-economy-0.35.2/README.md +431 -0
  65. package/docs/design/step-page-0.35.0/README.md +31 -31
  66. package/docs/design/step-page-0.35.0/frames/real-failed-120.txt +6 -6
  67. package/docs/design/step-page-0.35.0/frames/real-failed-200.txt +6 -6
  68. package/docs/design/step-page-0.35.0/frames/real-failed-55.txt +7 -7
  69. package/docs/design/step-page-0.35.0/frames/real-finished-120.txt +2 -2
  70. package/docs/design/step-page-0.35.0/frames/real-finished-200.txt +2 -2
  71. package/docs/design/step-page-0.35.0/frames/real-finished-55.txt +3 -3
  72. package/docs/design/step-page-0.35.0/frames/real-running-120.txt +6 -6
  73. package/docs/design/step-page-0.35.0/frames/real-running-200.txt +6 -6
  74. package/docs/design/step-page-0.35.0/frames/running-120.txt +1 -1
  75. package/docs/design/step-page-0.35.0/frames/running-200.txt +3 -3
  76. package/docs/design/step-page-0.35.0/frames/running-55.txt +1 -1
  77. package/docs/design/tidy-0.35.1/README.md +1 -1
  78. package/docs/design/tidy-0.35.1/frames/colour/real-home-200.txt +43 -43
  79. package/docs/design/tidy-0.35.1/frames/colour/real-home-55.txt +30 -30
  80. package/docs/design/tidy-0.35.1/frames/colour/real-run-finished-200.txt +5 -5
  81. package/docs/design/tidy-0.35.1/frames/colour/real-run-finished-55.txt +5 -5
  82. package/docs/design/tidy-0.35.1/frames/colour/real-run-running-200.txt +5 -5
  83. package/docs/design/tidy-0.35.1/frames/colour/real-run-running-55.txt +5 -5
  84. package/docs/design/tidy-0.35.1/frames/colour/real-step-detail-finished-200.txt +55 -55
  85. package/docs/design/tidy-0.35.1/frames/colour/real-step-detail-finished-55.txt +42 -42
  86. package/docs/design/tidy-0.35.1/frames/colour/real-step-overview-finished-200.txt +30 -30
  87. package/docs/design/tidy-0.35.1/frames/colour/real-step-overview-finished-55.txt +33 -33
  88. package/docs/design/tidy-0.35.1/frames/colour/real-step-overview-running-200.txt +3 -3
  89. package/docs/design/tidy-0.35.1/frames/colour/real-step-overview-running-55.txt +6 -6
  90. package/docs/design/tidy-0.35.1/frames/colour/real-task-200.txt +2 -2
  91. package/docs/design/tidy-0.35.1/frames/colour/real-task-55.txt +1 -1
  92. package/docs/design/tidy-0.35.1/frames/home-120.txt +1 -1
  93. package/docs/design/tidy-0.35.1/frames/home-200.txt +1 -1
  94. package/docs/design/tidy-0.35.1/frames/home-55.txt +1 -1
  95. package/docs/design/tidy-0.35.1/frames/real-home-120.txt +44 -44
  96. package/docs/design/tidy-0.35.1/frames/real-home-200.txt +43 -43
  97. package/docs/design/tidy-0.35.1/frames/real-home-55.txt +29 -29
  98. package/docs/design/tidy-0.35.1/frames/real-run-finished-120.txt +5 -5
  99. package/docs/design/tidy-0.35.1/frames/real-run-finished-200.txt +5 -5
  100. package/docs/design/tidy-0.35.1/frames/real-run-finished-55.txt +5 -5
  101. package/docs/design/tidy-0.35.1/frames/real-run-running-120.txt +5 -5
  102. package/docs/design/tidy-0.35.1/frames/real-run-running-200.txt +5 -5
  103. package/docs/design/tidy-0.35.1/frames/real-run-running-55.txt +5 -5
  104. package/docs/design/tidy-0.35.1/frames/real-stats-model-120.txt +1 -1
  105. package/docs/design/tidy-0.35.1/frames/real-stats-model-200.txt +1 -1
  106. package/docs/design/tidy-0.35.1/frames/real-stats-model-55.txt +1 -1
  107. package/docs/design/tidy-0.35.1/frames/real-stats-spending-120.txt +15 -15
  108. package/docs/design/tidy-0.35.1/frames/real-stats-spending-200.txt +14 -14
  109. package/docs/design/tidy-0.35.1/frames/real-stats-spending-55.txt +17 -17
  110. package/docs/design/tidy-0.35.1/frames/real-step-detail-failed-120.txt +55 -55
  111. package/docs/design/tidy-0.35.1/frames/real-step-detail-failed-200.txt +55 -55
  112. package/docs/design/tidy-0.35.1/frames/real-step-detail-failed-55.txt +40 -40
  113. package/docs/design/tidy-0.35.1/frames/real-step-detail-finished-120.txt +55 -55
  114. package/docs/design/tidy-0.35.1/frames/real-step-detail-finished-200.txt +55 -55
  115. package/docs/design/tidy-0.35.1/frames/real-step-detail-finished-55.txt +42 -42
  116. package/docs/design/tidy-0.35.1/frames/real-step-detail-running-120.txt +54 -54
  117. package/docs/design/tidy-0.35.1/frames/real-step-detail-running-200.txt +55 -55
  118. package/docs/design/tidy-0.35.1/frames/real-step-detail-running-55.txt +51 -51
  119. package/docs/design/tidy-0.35.1/frames/real-step-overview-failed-120.txt +2 -2
  120. package/docs/design/tidy-0.35.1/frames/real-step-overview-failed-200.txt +2 -2
  121. package/docs/design/tidy-0.35.1/frames/real-step-overview-failed-55.txt +1 -1
  122. package/docs/design/tidy-0.35.1/frames/real-step-overview-finished-120.txt +30 -30
  123. package/docs/design/tidy-0.35.1/frames/real-step-overview-finished-200.txt +30 -30
  124. package/docs/design/tidy-0.35.1/frames/real-step-overview-finished-55.txt +32 -32
  125. package/docs/design/tidy-0.35.1/frames/real-step-overview-running-120.txt +2 -2
  126. package/docs/design/tidy-0.35.1/frames/real-step-overview-running-200.txt +2 -2
  127. package/docs/design/tidy-0.35.1/frames/real-step-overview-running-55.txt +5 -5
  128. package/docs/design/tidy-0.35.1/frames/real-task-120.txt +2 -2
  129. package/docs/design/tidy-0.35.1/frames/real-task-200.txt +2 -2
  130. package/docs/design/tidy-0.35.1/frames/real-task-55.txt +1 -1
  131. package/docs/design/tidy-0.35.1/run-v2/README.md +1 -1
  132. package/docs/design/tidy-0.35.1/run-v2/after-200.txt +4 -4
  133. package/docs/design/tidy-0.35.1/run-v2/after-55.txt +2 -2
  134. package/docs/design/tidy-0.35.1/run-v2/before-200.txt +2 -2
  135. package/docs/design/tidy-0.35.1/run-v2/before-55.txt +2 -2
  136. package/docs/design/tidy-0.35.1/step-v2/0.35.2-before-detail-finished-200.txt +60 -0
  137. package/docs/design/tidy-0.35.1/step-v2/0.35.2-before-detail-finished-55.txt +60 -0
  138. package/docs/design/tidy-0.35.1/step-v2/0.35.2-before-overview-finished-200.txt +38 -0
  139. package/docs/design/tidy-0.35.1/step-v2/0.35.2-before-overview-finished-55.txt +71 -0
  140. package/docs/design/tidy-0.35.1/step-v2/0.35.2-detail-finished-200.txt +60 -0
  141. package/docs/design/tidy-0.35.1/step-v2/0.35.2-detail-finished-55.txt +60 -0
  142. package/docs/design/tidy-0.35.1/step-v2/0.35.2-detail-tool-open-200.txt +60 -0
  143. package/docs/design/tidy-0.35.1/step-v2/0.35.2-detail-tool-open-55.txt +60 -0
  144. package/docs/design/tidy-0.35.1/step-v2/0.35.2-overview-finished-200.txt +33 -0
  145. package/docs/design/tidy-0.35.1/step-v2/0.35.2-overview-finished-55.txt +42 -0
  146. package/docs/design/tidy-0.35.1/step-v2/0.35.2-overview-running-200.txt +25 -0
  147. package/docs/design/tidy-0.35.1/step-v2/0.35.2-overview-running-55.txt +42 -0
  148. package/docs/design/tidy-0.35.1/step-v2/README.md +55 -2
  149. package/docs/dynamic-workflow-handoff.md +2 -2
  150. package/docs/dynamic-workflow-v2-execution-plan.md +1 -1
  151. package/docs/experiments/2026-08-28-trending-ai-autonomy.md +31 -61
  152. package/docs/experiments/2026-08-29-dogfood-bullswarm-builds-bullswarm.md +3 -3
  153. package/docs/experiments/2026-09-06-caller-planner-evaluation.md +3 -3
  154. package/docs/guide/concepts.md +1 -1
  155. package/docs/guide/cost.md +73 -11
  156. package/docs/guide/observing.md +175 -38
  157. package/docs/guide/routing.md +13 -11
  158. package/docs/handoff-2026-09-13-custom-provider-config.md +2 -2
  159. package/docs/integration-audit-2026-08-31.md +32 -32
  160. package/docs/integrations/claude-code.md +1 -1
  161. package/docs/integrations/issue-watcher.md +3 -3
  162. package/docs/plans/0.33.1.goal.txt +1 -1
  163. package/docs/plans/0.33.1.program.json +13 -13
  164. package/docs/plans/0.33.2-stats.goal.txt +1 -1
  165. package/docs/plans/0.33.2-stats.program.json +7 -7
  166. package/docs/plans/cost-0.34.0.program.json +10 -10
  167. package/docs/plans/cost-audit.goal.txt +4 -4
  168. package/docs/plans/cost-audit.program.json +8 -8
  169. package/docs/plans/dashboard-0.33-fidelity.goal.txt +1 -1
  170. package/docs/plans/dashboard-0.33-fidelity.md +5 -5
  171. package/docs/plans/dashboard-0.33-fidelity.program.json +10 -10
  172. package/docs/plans/dashboard-0.33-fidelity.rev2.json +13 -13
  173. package/docs/plans/dashboard-0.33-review.goal.txt +8 -8
  174. package/docs/plans/dashboard-0.33-review.program.json +6 -6
  175. package/docs/plans/dashboard-0.33.goal.txt +1 -1
  176. package/docs/plans/dashboard-0.33.program.json +10 -10
  177. package/docs/plans/dashboard-split.program.json +5 -5
  178. package/docs/plans/fidelity.program.json +2 -2
  179. package/docs/plans/handoff.goal.txt +1 -1
  180. package/docs/plans/handoff.program.json +5 -5
  181. package/docs/plans/handoff.rev2.json +5 -5
  182. package/docs/plans/handoff.rev3.json +6 -6
  183. package/docs/plans/meter-ledger.program.json +6 -6
  184. package/docs/plans/providers-0.35.2.goal.txt +11 -0
  185. package/docs/plans/providers-0.35.2.program.json +160 -0
  186. package/docs/plans/step-page-0.35.0-build.goal.txt +1 -1
  187. package/docs/plans/step-page-0.35.0-build.program.json +7 -7
  188. package/docs/plans/step-page-0.35.0-design.goal.txt +2 -2
  189. package/docs/plans/step-page-0.35.0-design.program.json +3 -3
  190. package/docs/plans/tidy-0.35.1-colour.program.json +4 -4
  191. package/docs/plans/tidy-0.35.1.goal.txt +1 -1
  192. package/docs/plans/tidy-0.35.1.program.json +11 -11
  193. package/docs/reference/cli.md +151 -7
  194. package/docs/reference/configuration.md +36 -1
  195. package/docs/reference/program.md +31 -1
  196. package/docs/reference/providers.md +80 -2
  197. package/docs/reference/result.md +78 -1
  198. package/docs/studies/cost-audit-2026-09-18/budget-page-audit.md +50 -51
  199. package/docs/studies/cost-audit-2026-09-18/claude-actual-vs-recorded.md +14 -14
  200. package/docs/studies/cost-audit-2026-09-18/cost-fix-plan.md +5 -5
  201. package/docs/studies/cost-audit-2026-09-18/estimator-audit.md +13 -13
  202. package/docs/studies/cost-audit-2026-09-18/grok-codex-actual-vs-recorded.md +38 -39
  203. package/docs/studies/cost-audit-2026-09-18/step-timeline-120.txt +2 -2
  204. package/docs/studies/cost-audit-2026-09-18/step-timeline-55.txt +2 -2
  205. package/docs/studies/cost-audit-2026-09-18/step-timeline-design.md +8 -8
  206. package/docs/studies/portal-token-diet.md +24 -24
  207. package/docs/workflow-agent-usability-audit-2026-08-27.md +4 -4
  208. package/mods/bullswarm/.claude-plugin/plugin.json +1 -1
  209. package/mods/bullswarm/README.md +4 -4
  210. package/mods/bullswarm/hooks/pane.tsx +72 -44
  211. package/mods/bullswarm/hooks/register.ts +62 -2
  212. package/mods/bullswarm/hooks/step.ts +286 -374
  213. package/package.json +1 -1
  214. package/providers/contrib/command-code/connector.json +48 -3
  215. package/providers/contrib/command-code/provider.mjs +22 -1
  216. package/providers/contrib/opencode/connector.json +16 -0
  217. package/providers/contrib/opencode/provider.mjs +15 -0
  218. package/skill/SKILL.md +112 -21
  219. package/skill/references/operations.md +162 -17
  220. package/skill/references/program.md +28 -1
  221. package/src/cli.js +137 -17
  222. package/src/help.js +178 -16
  223. package/src/home-cli.js +170 -18
  224. package/src/lib/cli-flags.js +9 -2
  225. package/src/lib/glyphs.js +2 -0
  226. package/src/lib/prices.js +16 -5
  227. package/src/lib/quota.js +460 -13
  228. package/src/lib/retention.js +541 -0
  229. package/src/lib/route.js +93 -20
  230. package/src/lib/stale.js +457 -0
  231. package/src/lib/state.js +141 -6
  232. package/src/lib/subscription-cost.js +6 -4
  233. package/src/lib/tasks.js +24 -0
  234. package/src/lib/transcripts/claude-code.js +146 -30
  235. package/src/lib/transcripts/codex.js +148 -30
  236. package/src/lib/transcripts/command-code.js +417 -0
  237. package/src/lib/transcripts/index.js +142 -32
  238. package/src/lib/transcripts/indexing.js +134 -0
  239. package/src/lib/transcripts/opencode.js +568 -0
  240. package/src/lib/watch.js +372 -70
  241. package/src/meters/framework.js +1 -1
  242. package/src/provider-cli.js +13 -0
  243. package/src/providers/_schema.json +10 -2
  244. package/src/providers/claude-code/connector.json +15 -1
  245. package/src/providers/claude-code/provider.mjs +9 -1
  246. package/src/providers/codex/connector.json +9 -1
  247. package/src/providers/codex/provider.mjs +9 -1
  248. package/src/providers/grok/connector.json +35 -1
  249. package/src/providers/grok/provider.mjs +9 -1
  250. package/src/setup.js +8 -0
  251. package/src/workflow/action-validator.js +54 -7
  252. package/src/workflow/budget-model.js +71 -9
  253. package/src/workflow/budget-view.js +18 -0
  254. package/src/workflow/cli.js +117 -3
  255. package/src/workflow/dashboard.js +201 -39
  256. package/src/workflow/history-view.js +107 -23
  257. package/src/workflow/history.js +11 -2
  258. package/src/workflow/home-model.js +182 -37
  259. package/src/workflow/home-view.js +643 -279
  260. package/src/workflow/reconcile.js +832 -0
  261. package/src/workflow/reprice.js +167 -45
  262. package/src/workflow/rollup.js +5 -2
  263. package/src/workflow/run-model.js +46 -6
  264. package/src/workflow/run-view.js +116 -32
  265. package/src/workflow/runs-view.js +5 -3
  266. package/src/workflow/spend-facts.js +152 -0
  267. package/src/workflow/stat-kit.js +6 -3
  268. package/src/workflow/stats-view.js +36 -19
  269. package/src/workflow/step-model.js +230 -72
  270. package/src/workflow/step-view.js +402 -117
  271. package/src/workflow/time-box.js +276 -0
  272. package/src/workflow/usage-preference.js +22 -0
  273. package/src/workflow/usage-view.js +2 -1
  274. package/src/workflow/v2-dispatch.js +148 -6
  275. package/src/workflow/v2-outcome.js +153 -11
  276. package/src/workflow/v2-planner.js +17 -4
  277. package/src/workflow/v2-revision.js +4 -1
  278. package/src/workflow/v2-runtime.js +348 -18
  279. package/src/workflow/v2-state.js +153 -2
  280. package/src/workflow/verify-rounds.js +828 -0
  281. package/src/workflow/watch-cli.js +206 -17
  282. package/docs/design/owner-review-2026-09-18/budget-page-current-0.33.1.png +0 -0
  283. package/docs/design/owner-review-2026-09-18/budget.jpeg +0 -0
  284. package/docs/design/owner-review-2026-09-18/history-unknown-project.jpeg +0 -0
  285. package/docs/design/owner-review-2026-09-18/home-licence-today-missing-codex-0.33.1.png +0 -0
  286. package/docs/design/owner-review-2026-09-18/home-phone-landscape.jpeg +0 -0
  287. package/docs/design/owner-review-2026-09-18/home-phone-portrait.jpeg +0 -0
  288. package/docs/design/owner-review-2026-09-18/round2-budget-phone.jpeg +0 -0
  289. package/docs/design/owner-review-2026-09-18/round2-trends-30d-horizontal-split.png +0 -0
  290. package/docs/design/owner-review-2026-09-18/round2-trends-desktop.png +0 -0
  291. package/docs/design/owner-review-2026-09-18/round2-trends-phone.png +0 -0
  292. package/docs/design/owner-review-2026-09-18/run-page-budget.jpeg +0 -0
  293. package/docs/design/owner-review-2026-09-18/runs-list.jpeg +0 -0
  294. package/docs/design/owner-review-2026-09-18/stats-pools.jpeg +0 -0
  295. package/docs/design/owner-review-2026-09-18/stats-trends-spent.jpeg +0 -0
  296. package/docs/design/owner-review-2026-09-19/stats-overview-current.png +0 -0
  297. package/docs/design/prototype-shots/Screenshot 2026-09-17 at 9.14.40/342/200/257AM.png +0 -0
  298. package/docs/design/prototype-shots/Screenshot 2026-09-17 at 9.15.23/342/200/257AM.png +0 -0
  299. package/docs/design/prototype-shots/Screenshot 2026-09-17 at 9.16.07/342/200/257AM.png +0 -0
  300. package/docs/design/prototype-shots/Screenshot 2026-09-17 at 9.16.18/342/200/257AM.png +0 -0
  301. package/docs/design/stats-frames-0.33.2/budget-120.png +0 -0
  302. package/docs/design/stats-frames-0.33.2/budget-55.png +0 -0
  303. package/docs/design/stats-frames-0.33.2/stats-spending-120.png +0 -0
  304. package/docs/studies/cost-audit-2026-09-18/owner-budget-page-2026-09-18.png +0 -0
package/CHANGELOG.md CHANGED
@@ -1,5 +1,215 @@
1
1
  # bullswarm changelog
2
2
 
3
+ ## 0.35.3 — whole right edges in Ghostty and herdr, Enter opens the row you are on, named phases
4
+
5
+ - dashboard: every row is erased before it is painted instead of after, so a
6
+ row that fills the terminal keeps its last cell in Ghostty and herdr (both
7
+ cleared it, turning `50%` into `50` and `16` into `1` on Home's right edge).
8
+ - home: from 110 columns the spent-per-day chart grows to the height of the
9
+ stacked breakdowns beside it; the summary band sizes `Spent` to its text and
10
+ wraps at ` · ` rather than cutting a figure; the closing sentence wraps too.
11
+ - runs: Enter and a click open exactly the row the cursor is on. Tasks recorded
12
+ before the task ledger have no id, and all of them shared one empty key, so
13
+ Enter on a workflow row could open one of those tasks instead.
14
+ - run: phase rules name the phase: `Phase 2 · Build · time-box · docs`
15
+ (Build, Integrate, Verify, Digest, Design, Repair from the steps' kinds,
16
+ joined with ` + ` when a phase mixes them); the phase number and name are
17
+ never cut, step names shorten first.
18
+ - home: day keys no longer build an `Intl.DateTimeFormat` per call to learn
19
+ the zone (it is cached against `TZ`); a Home paint takes about 40% less time.
20
+
21
+ ## 0.35.2 — the Step page reads like a transcript; prices without a command, honest totals, a home that stops growing
22
+
23
+ - step: detail is now a scrollable transcript of every turn in order: the full
24
+ response followed by one row per command or tool call; opening a turn or
25
+ tool row still reaches every captured event field, argument, result and exit.
26
+ - step: overview opens on the latest ten turns on desktop and five on phones,
27
+ newest at the bottom, with one `turns 1–N · … · click for detail` line for
28
+ earlier work; a followed live window slides until the reader moves it.
29
+ - step: `overview · detail` sits in the activity heading, right after the
30
+ heading word, each word a click target; `v`, the footer and Help name the
31
+ same views, and standalone tasks use the same page.
32
+ - step: clicking a turn head toggles it exactly like Enter, and hover lights
33
+ only the turn's text rather than its padding or desktop side column.
34
+ - run: each timeline attempt row opens Step on that exact attempt; step and
35
+ phase rows that name no attempt continue to open the latest one.
36
+ - grok: `tool_call_update` captures merge into the `tool_call` with the same id,
37
+ so a call has one named row and one count instead of extra nameless `agent`
38
+ rows, while every update remains reachable from transcript detail.
39
+ - providers: every shipped connector declares `eventStream.toolKinds`, mapping
40
+ its tool vocabulary to command, read, search, edit or other; the Step model
41
+ contains no provider-specific tool-name table.
42
+ - design: the Step v2 record adds transcript, latest-turn-window and toggle
43
+ rules with before/after frames at 55 and 200 columns; committed real text and
44
+ colour frames are regenerated from scrubbed fixtures.
45
+ - mod: the Claude Code Step pane draws the same v2 header, turns, result, task,
46
+ cost, phone order, colours, transcript toggle and latest-turn window from
47
+ `action show --json`, for workflow steps and standalone tasks alike.
48
+ - mod: Usage opens with `u` or a click even with no workflow running, and back
49
+ returns to idle; validated temporary chunks carry large Step records within
50
+ hook limits and are removed immediately.
51
+ - reprice: an incremental reconciler prices every terminal attempt that is
52
+ still unknown or a bytes/4 estimate, with no user command. The kernel prices
53
+ its own run at quiet boundaries and before the result and rollup are
54
+ written; a single `bullswarm run` whose usage is unmeasured starts a detached
55
+ pass 30 s after it ends (only when the pool's provider keeps transcripts);
56
+ the dashboard starts one detached, throttled pass (one per 5 minutes, one at
57
+ a time) after its first paint and says `pricing N older records…` while it
58
+ runs. A per-attempt ledger in `pricing/reconcile.json` retries recent
59
+ attempts after 1 and 10 minutes, gives older ones one try, and reopens an
60
+ attempt only when a new or grown transcript covers its window, so a second
61
+ pass does nothing. The automatic pass never makes an attempt worse: no match
62
+ or an ambiguous match leaves it as it was. `workflow reprice` stays the
63
+ manual full backfill and gains `--incremental` (with `--trigger`,
64
+ `--transcript-home`, `--delay-ms`). On a copy of a real home the first pass
65
+ priced 165 of 281 eligible attempts; the second took milliseconds and wrote
66
+ nothing.
67
+ - transcripts: Codex and Claude matching tries the session id, then the task
68
+ text (the task-file path or exact text in the first user message), then cwd
69
+ + time window; more than one candidate stays unknown, and the cwd/time step
70
+ skips transcripts that quote a different task file. On the same copy the
71
+ formerly ambiguous attempts resolved by task text: codex 100, claude-code 19,
72
+ claude-code:acme 14. Indexes store prompt paths and hashes, never the text.
73
+ Single tasks are priced too; a task's window now ends at its `endedAt`.
74
+ - capture: each attempt records an immutable `capture` block when its worker
75
+ exits — provider session id and its source, model, exclusive token classes,
76
+ provider-reported cost, exit code and signal — before the end meter read and
77
+ the transcript lookup, and persists it at once. Later pricing may upgrade an
78
+ estimate but never downgrades provider-reported usage. Planner and scout
79
+ attempts are captured too.
80
+ - grok: the final `end` event is decoded (session id, token classes, cost),
81
+ proven on a real grok 1.0.13 stream checked in scrubbed as
82
+ `tests/fixtures/stream/grok-capture.jsonl`; `"rate limit"` is no longer a
83
+ grok auth signature, so a grok 429 is a throttle, not a 10-minute auth pause.
84
+ - command-code: launched without `--no-session`, so its session transcript
85
+ exists afterwards and the reader matches it by session id.
86
+ - money: every money total on Home, Runs, Run, Stats and Budget that includes
87
+ unmeasured attempts reads `at least $X · N unmeasured`; a scope with no
88
+ recorded amount reads `api unknown`. Budget prints a `spent …` line per pool,
89
+ and the Stats hover, axis and coverage note use the same words. One month
90
+ length, `DAYS_PER_MONTH = 30.4375` (365.25/12), is used everywhere.
91
+ - home: the window-share cell shows the measured share from
92
+ `calibration/<pool>.json` (`5.0% measured`) when samples attribute a window
93
+ drop to today's runs, the labelled pace estimate only when none do, and a
94
+ dash otherwise. The recent list shows finished runs only — a reopened run
95
+ leaves it — and each row's mark is the Run page header's mark; a
96
+ non-terminal run's index row no longer carries a finish time.
97
+ - retention: `state.json.retention` `{ "enabled": true, "workspacesDays": 7 }`
98
+ removes the `workspaces/` copies inside terminal runs older than the limit,
99
+ in a detached background sweep started by the kernel, watch completion and
100
+ the dashboard (at most every 6 hours, one at a time, only when a run holds a
101
+ workspace copy). Records, reports, streams, task/out markdown, interrupted
102
+ runs and any leased run are never touched. `bullswarm home prune
103
+ [--dry-run|--yes] [--days n]` lists or removes the same set with bytes, and
104
+ `bullswarm home status` shows the policy, bytes on disk and the last prune
105
+ and reprice results.
106
+ - pauses: a limit notice pauses a pool for quota only on proof — the pool's
107
+ own meter at 95% or more on a running window, or a provider line that says a
108
+ usage window is spent and names its reset. Everything else, such as
109
+ `Rate limit exceeded. Please wait a moment and try again.`, is a transient
110
+ throttle: the same pool is retried after 20 s and 60 s (or the wait it
111
+ named), then the attempt moves on, and the pool is never paused. `pools`
112
+ prints `PAUSED until <time> · <proof> · provider: "<line>" · meter then: … ·
113
+ lift now: bullswarm pools resume <pool>`; `bullswarm pools resume <pool>`
114
+ lifts a pause, and `bullswarm strategy set-pausing off|on` turns automatic
115
+ pausing off or back on.
116
+ - pauses: `bullswarm strategy set-pausing off` now stops every automatic
117
+ pause, not only quota — auth, the credential-group siblings an auth pause
118
+ benches with it, and the soft bench a second strike writes — so with the
119
+ switch off nothing is taken out of service by a command's own judgement;
120
+ routing still reads meters and a failed attempt still moves to another pool.
121
+ It is stored as `strategy.pausing: "off"`. `pools` opens with `automatic
122
+ pausing: off`, and `pools resume` still lifts a pause that was already in
123
+ place. Quota and auth signatures are matched against the provider's own
124
+ error channel only — its stderr, the events it flags as errors and its
125
+ terminal `result` record — never an assistant's reply or a tool result: on
126
+ 2026-09-21 a pool was paused for a sentence an agent wrote about
127
+ `usage_credits_required` in its own report.
128
+ - routing: an evidence step may run on the pool that wrote the judged work
129
+ when that pool is urgent; the reason then says `independence waived: <pool>
130
+ resets in <clock>`. Independence is judged by model family, so
131
+ `claude-code` and `claude-code:acme` are the same writer; with nothing
132
+ urgent, independence stays the tie-breaker.
133
+ - watch: `bullswarm workflow watch <run> --until outcome|trouble` prints only
134
+ trouble lines (failed, rejected, paused, stalled, stale, steering) and the
135
+ outcome; `trouble` exits at the first one with a `next:` relaunch line. A
136
+ running attempt gets a stale score (quiet with no command running, no file
137
+ change while commands continue, the same command repeated, wall time over 3×
138
+ the expected minutes) and one `⚠ <step> looks stale: <reasons>` line.
139
+ `bullswarm workflow step restart <run> <step> [--pool <pool>]` stops that
140
+ attempt and requeues the step with its durable handoff; nothing restarts on
141
+ its own. The packaged skill teaches one background `--until trouble` watch
142
+ per run and one tool call per wake.
143
+ - providers: OpenCode and Command Code expose the same durable-transcript
144
+ reader contract as the first-class providers. OpenCode reads its SQLite
145
+ sessions read-only; Command Code reads persisted project JSONL only when a
146
+ session transcript exists and carries usage.
147
+ - reprice: transcript lookup follows the loaded provider registry, so
148
+ `opencode2*`, `opencode2:kaihk-*`, and `command-code` pools reach their
149
+ provider-owned readers without a hard-coded provider list. Ambiguous,
150
+ missing, and checkpoint-only records stay unknown rather than becoming
151
+ zero-cost attempts.
152
+ - command-code: the 116 historical attempts studied for this release used
153
+ `--no-session`, so their checkpoints are documented as non-recoverable
154
+ history.
155
+ - pricing: public model cards are retained only with a source and date (the
156
+ cards were checked 2026-09-20). `kaihk/*` and `opencode/union-alpha` relay
157
+ identifiers have no public card, so the underlying OpenAI card is not
158
+ substituted; observed Command Code models use the cited Command Code card.
159
+ - evidence: live OpenCode and Command Code streams captured on 2026-09-20 are
160
+ checked in as `tests/fixtures/streams/opencode-hello.jsonl` and
161
+ `tests/fixtures/streams/command-code-hello.jsonl`; each contrib connector
162
+ claims the `eventStream.usage` rules those captures prove.
163
+ - tests: detached reprice and prune children never recreate a home that was
164
+ deleted under them, and the OpenCode read-only test runs everywhere against
165
+ the exported rows instead of skipping without a local database.
166
+ - time box: every work and evidence step's task ends with a soft time box: the
167
+ box in minutes, the start clock, a wrap-up point at 70% of the box, and an
168
+ invitation to stop and report `## Done`, `## Not done` and `## Suggested next
169
+ step`. The box is the action's `timeBox` (whole minutes, 0-240), else
170
+ `defaults.timeBox`, else 1.5 × the median wall minutes of this home's
171
+ succeeded attempts for the pool and kind (5 or more attempts), else for the
172
+ kind, else 20, rounded to 5 and kept within 10-60; `opencode` attempts never
173
+ feed it, and `timeBox: 0` leaves the paragraph out. It is a guide: timeouts,
174
+ stall detection, cancellation and routing are unchanged and nothing stops at
175
+ the box. On the fixture home the (codex, implement) pair has a median of
176
+ 18.86 minutes over 80 attempts, so its box is 30. In an experiment on one
177
+ task (two runs per arm) a 15-minute box ran 10m54s and 10m03s against 23m30s
178
+ and 48m44s without one, with honest partial reports: a direction, not a
179
+ measurement.
180
+ - early return: a work step whose `## Not done` lists items still succeeds and
181
+ records `returnedEarly` with the count and the items. The Step page header and
182
+ the Run timeline row read `returned early · N not done`, the Step page shows
183
+ `box 20m · ran 34m` when an attempt ran past its box, `workflow watch` prints
184
+ `◐ <step> returned early · N not done`, and the items reach the verifiers with
185
+ the rest of the evidence.
186
+ - verify rounds: a failed verify no longer waits for the caller. The kernel runs
187
+ a bounded repair loop of at most 3 verify rounds (`defaults.verifyRounds`
188
+ 1-3, default 3, 1 = the old single round): round 1 judges every requirement,
189
+ each failure starts a kernel `repair-<n>` step built from the verifier's
190
+ evidence, the not-done items and the handoffs of the steps that affect it,
191
+ round 2 re-checks the failures and looks for regressions and the same defect
192
+ elsewhere, and round 3 is final closure. A requirement that passed carries
193
+ forward and is judged again only when a repair changed a file its evidence
194
+ names. There is never a fourth round and a program still cannot declare a
195
+ repair step; the planner contract allows `timeBox` and says the kernel adds
196
+ `repair-<n>` and `verify-round-<n>`.
197
+ - verify rounds: the run ends `completed · verified`, or `completed · not
198
+ verified · verify rounds 3/3` with a caller-decision block (each requirement
199
+ still failing, its latest evidence and one suggested next step) in
200
+ `workflow runs result <run> --json --summary`. The result also lists each
201
+ verify round and each repair with its wall minutes, pool and cost. The Run
202
+ page, Home, Runs and `workflow watch` show each round and repair; the events
203
+ are `workflow.verify-round` and `workflow.repair`. A plan revision during the
204
+ loop is still accepted and never adds or refunds a round; its
205
+ `defaults.verifyRounds` sets the cap for the rest of the run, and a revision
206
+ that changes only the cap is accepted.
207
+ - docs: the skill, its operations and program references, the program and result
208
+ references and `docs/design/step-economy-0.35.2/README.md` teach `timeBox`,
209
+ `verifyRounds`, the early-return report and the caller-decision block; the
210
+ skill's "Observe and judge" no longer tells the caller to hand-add a fix step
211
+ for an ordinary failed check.
212
+
3
213
  ## 0.35.1 — a calmer dashboard with honest active time
4
214
 
5
215
  - durations: Home, Runs, Run, Step, and Stats use the union-based
@@ -178,7 +388,7 @@
178
388
  colour codes, so every bar came out grey while the legend beside it stayed
179
389
  coloured — a chart titled "by pool" painted all five pools identically.
180
390
  - stats: no two series in one chart share a hue. Colours were hashed per name,
181
- which put `claude-code:wati` and `codex` on the same blue.
391
+ which put `claude-code:acme` and `codex` on the same blue.
182
392
 
183
393
  ## 0.33.1 — say what a number is, or say you do not know
184
394
 
@@ -307,8 +517,8 @@
307
517
  price exactly as it spends a declared one and labels which it is —
308
518
  `detected` or `declared` — so an operator's own figure is never confused with
309
519
  one bullswarm inferred. The plan table is keyed by provider, not by pool, so a
310
- discovered per-account pool such as `claude-code:wati` resolves against
311
- `claude-code`'s plans; it used to ask for `claude-code:wati`'s own plans and
520
+ discovered per-account pool such as `claude-code:acme` resolves against
521
+ `claude-code`'s plans; it used to ask for `claude-code:acme`'s own plans and
312
522
  find nothing, which left exactly the pools with one login per account
313
523
  without a price.
314
524
  - Mod: the pane's `usage` button works with no run in flight — it renders the
@@ -374,7 +584,7 @@ drawn from the real pool figures of 2026-09-18.
374
584
  - dashboard: the phone `Home` no longer hides data behind its width. The
375
585
  `budget · this week` block draws every pool the desktop draws — four today —
376
586
  or ends with `+N more` when the rows genuinely do not fit, and its names are
377
- wide enough that `claude-code:wati` and `claude-code` read as two pools
587
+ wide enough that `claude-code:acme` and `claude-code` read as two pools
378
588
  rather than one truncated one. The by-pool, by-model and by-project lists
379
589
  keep at least their top three rows (`+N more` for the rest) and always paint
380
590
  the whole percentage, never `24…`. The summary figures sit one per line, so
@@ -949,7 +1159,7 @@ drawn from the real pool figures of 2026-09-18.
949
1159
  from a reset the operator declares —
950
1160
  `bullswarm strategy set-subscription <pool> --resets-at <iso|unknown>`.
951
1161
  The Relay token API stopped returning `expires_at` for the `1` and `Moham`
952
- wallets on 2026-09-03 (the fleetlens log holds 28 dated snapshots between
1162
+ wallets on 2026-09-03 (the project-n log holds 28 dated snapshots between
953
1163
  2026-08-31 and 2026-09-03, then only nulls), so `opencode2` and
954
1164
  `opencode2:relay-3` printed `unmetered` at 73.8% and 10.7% of their $50
955
1165
  wallets and ranked as neutral. With a declared reset the used% stays the
@@ -982,8 +1192,8 @@ drawn from the real pool figures of 2026-09-18.
982
1192
  — its effective surplus divided by the fraction of the window still to run —
983
1193
  ahead of every pool whose window is not about to close, instead of on the
984
1194
  surplus alone. At 2026-09-11 12:26 HKT the medium lane went to
985
- `claude-code:wati` (+22.9 points, 13h33m and 8.1% of its week left) over
986
- `grok` (+13.8 points, 2h02m and 1.2% left, urgency ~1150 against wati's
1195
+ `claude-code:acme` (+22.9 points, 13h33m and 8.1% of its week left) over
1196
+ `grok` (+13.8 points, 2h02m and 1.2% left, urgency ~1150 against acme's
987
1197
  ~283), and grok's points expired unspent two hours later; the owner had been
988
1198
  pinning grok by hand for such runs. Three states for an expiring pool:
989
1199
  `urgent` (surplus still to spend and a pacing forecast — the reading plus
@@ -1004,17 +1214,17 @@ drawn from the real pool figures of 2026-09-18.
1004
1214
  - routing: the 5-hour near-limit line is now clock-relative. A pool is
1005
1215
  deprioritized only when its 5h forecast is at/above 75% AND ahead of the
1006
1216
  share of the 5h window that has already elapsed. On 2026-09-10 at 22:19Z a
1007
- high-tier integrator skipped `claude-code:wati` (81% used, 23 minutes to the
1217
+ high-tier integrator skipped `claude-code:acme` (81% used, 23 minutes to the
1008
1218
  reset — 92.3% of the window elapsed, 88.1% projected) and
1009
- `claude-code:petsona` (75.3% projected, 85.7% elapsed) and went to the one
1010
- account already ahead of its weekly pace, while wati still held 34% of its
1219
+ `claude-code:initech` (75.3% projected, 85.7% elapsed) and went to the one
1220
+ account already ahead of its weekly pace, while acme still held 34% of its
1011
1221
  weekly quota unspent with 13% of the week left to spend it. Both pools now keep the lane:
1012
1222
  88.1% with 23 minutes left is a pool spending at the clock's pace, not a pool
1013
1223
  about to hit a wall. Pools with no `resets_at`, an unparsable one, or a reset
1014
1224
  already past keep the fixed 75% line, and the 90% burst gate is unchanged —
1015
1225
  it ignores the clock. The routing reason and the `candidates[]` rows say
1016
1226
  which case applied: `5h used 81% -> 88.1% projected, under the clock (92.3%
1017
- elapsed)`, `skipped near 5h limit (projected): claude-code:wati 88.1% (20.0%
1227
+ elapsed)`, `skipped near 5h limit (projected): claude-code:acme 88.1% (20.0%
1018
1228
  elapsed)`, and a new `fiveHourElapsedPct` field. `bullswarm pools` shows the
1019
1229
  same clock: `5h=81% (92% elapsed)`.
1020
1230
  - routing: 5-hour spend is clipped at the reset. Quota spent after the window
@@ -1720,7 +1930,7 @@ drawn from the real pool figures of 2026-09-18.
1720
1930
 
1721
1931
  - All of it is visible after the fact. The routing reason names the in-flight
1722
1932
  counts and projections that moved the pick (`5h used 30% -> 41% projected, 2
1723
- in flight`, `skipped near 5h limit (projected): wati 76%`, `forecast-gated
1933
+ in flight`, `skipped near 5h limit (projected): acme 76%`, `forecast-gated
1724
1934
  at/above 90%: …`, `preferred over busier: …`), every candidate row carries
1725
1935
  `pace`, `effectiveSurplus`, `inflight`, `projectedFiveHourPct`,
1726
1936
  `forecastFiveHourPct`, `projectedWeeklyPct`, `ratePerMinute`,
@@ -2776,7 +2986,7 @@ Adopts the driving mechanics of Claude Code's dynamic workflow (documented in
2776
2986
  former 900-second planner/action timeout defaults.
2777
2987
  - Fixed adaptive completion policy, current-action metadata, provider routing
2778
2988
  history, usage aggregation, latest-worker verification, and truthful partial
2779
- token/cost accounting found during the Kipwise battle test.
2989
+ token/cost accounting found during the project-b battle test.
2780
2990
  - Added connector-owned native JSONL event adapters for Codex, Claude, Grok,
2781
2991
  Command Code, and OpenCode. Workflows now retain and display the latest three
2782
2992
  semantic shell/read/edit/write/response actions for every active agent.
package/GOAL.md CHANGED
@@ -2,7 +2,7 @@
2
2
 
3
3
  > Historical (2026-08-21): accurate when written; see CHANGELOG for what changed since.
4
4
 
5
- **Status:** HISTORICAL PROTOTYPE CHARTER · **Owner:** cowcow02 · **Created:** 2026-08-21
5
+ **Status:** HISTORICAL PROTOTYPE CHARTER · **Owner:** dev · **Created:** 2026-08-21
6
6
 
7
7
  ## One sentence
8
8