@iowarp/clio-coder 0.4.2 → 0.4.4

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (523) hide show
  1. package/CHANGELOG.md +98 -0
  2. package/CONTRIBUTING.md +86 -19
  3. package/README.md +35 -6
  4. package/dist/{acp-TMDQZDIG.js → acp-WNAYYF4F.js} +12 -13
  5. package/dist/{agents-5N5NG3XG.js → agents-3OKXHLOI.js} +60 -57
  6. package/dist/assets/codewiki.json +1 -1
  7. package/dist/{auth-Z5CCBXKQ.js → auth-VKNNMGPU.js} +21 -19
  8. package/dist/{builtins-K6TNDT24.js → builtins-WGALA46I.js} +9 -4
  9. package/dist/{chunk-XE3PCIXH.js → chunk-23L32XTI.js} +12 -9
  10. package/dist/{chunk-I64IFBLB.js → chunk-25QBEXRS.js} +18 -11
  11. package/dist/{chunk-CDNVLKUX.js → chunk-26QSH3EJ.js} +13 -7
  12. package/dist/{chunk-QQLGQY2A.js → chunk-2ASED4PZ.js} +22 -22
  13. package/dist/{chunk-MCEPRMZW.js → chunk-2CU2H6KE.js} +2 -2
  14. package/dist/chunk-2DSOYNFC.js +108 -0
  15. package/dist/{chunk-72GZI5EV.js → chunk-2JH2WHGE.js} +2 -2
  16. package/dist/{chunk-I66EAJFY.js → chunk-2WZ546HR.js} +267 -232
  17. package/dist/{chunk-O3YUNJZ2.js → chunk-2ZSONWVL.js} +82 -25
  18. package/dist/{chunk-2NHR3NAY.js → chunk-36CT5VVL.js} +331 -42
  19. package/dist/{chunk-2X4RYJTJ.js → chunk-3GY4F45V.js} +3 -3
  20. package/dist/{chunk-ZW55JB7N.js → chunk-3ODX73FK.js} +4 -6
  21. package/dist/{chunk-PBP4B7XR.js → chunk-3UNOLWNZ.js} +3 -3
  22. package/dist/{chunk-4JDLP6ZS.js → chunk-3UUXNFEX.js} +14 -10
  23. package/dist/{chunk-RKSR6VSF.js → chunk-4IUZQIJ3.js} +29 -1
  24. package/dist/{chunk-ZW4HH5JJ.js → chunk-4M6Z5QVF.js} +6 -6
  25. package/dist/{chunk-K6BSR66V.js → chunk-4NSRCOYP.js} +4 -1
  26. package/dist/{chunk-M2DAX4F6.js → chunk-4WR7VSYB.js} +2 -2
  27. package/dist/{chunk-FSP7CMNU.js → chunk-54X7T7DK.js} +61 -6
  28. package/dist/{chunk-54ODD65L.js → chunk-5636DCO5.js} +4 -4
  29. package/dist/chunk-57XXR6DR.js +3763 -0
  30. package/dist/{chunk-3KIPBMUA.js → chunk-5ICU3EUH.js} +2 -2
  31. package/dist/{chunk-77QIVUZB.js → chunk-5MEZN6CB.js} +4 -4
  32. package/dist/{chunk-O42A54GG.js → chunk-5OIVVPHF.js} +2 -2
  33. package/dist/{chunk-YJISEZKC.js → chunk-5TUB6SLS.js} +6 -6
  34. package/dist/{chunk-IMXMHHMQ.js → chunk-6OSVSQL5.js} +341 -57
  35. package/dist/{chunk-Q4XWMHX6.js → chunk-6PAZTBPA.js} +14 -2
  36. package/dist/{chunk-FVDGR2ZL.js → chunk-6Q3CYFD3.js} +112 -39
  37. package/dist/{chunk-IDNA72AH.js → chunk-6QOTUPRG.js} +155 -36
  38. package/dist/{chunk-X7IARSHT.js → chunk-6UINWWS6.js} +16 -10
  39. package/dist/{chunk-CYZW7JHJ.js → chunk-72YIHOZQ.js} +9 -9
  40. package/dist/{chunk-IKSLQ4XV.js → chunk-75W7L2E2.js} +752 -861
  41. package/dist/{chunk-CRFOIAX3.js → chunk-7UGL4MB5.js} +6 -6
  42. package/dist/{chunk-HIICAHCJ.js → chunk-AUPNRN7C.js} +2 -2
  43. package/dist/{chunk-7BHIY2MW.js → chunk-BJVFZO5U.js} +8 -50
  44. package/dist/{chunk-B74PXLU7.js → chunk-CUSRQKPU.js} +65 -3
  45. package/dist/chunk-DQOVN6KV.js +386 -0
  46. package/dist/{chunk-E7GT7O5N.js → chunk-DT3LWJOB.js} +7 -4
  47. package/dist/chunk-DXKJURES.js +671 -0
  48. package/dist/{chunk-JBCS7CRR.js → chunk-EL24TAU4.js} +10 -10
  49. package/dist/{chunk-TPEQIQIE.js → chunk-ELWDPP3Y.js} +8 -8
  50. package/dist/{chunk-NDINPTJ4.js → chunk-ELZVTCGV.js} +5 -4
  51. package/dist/chunk-EXLD33WO.js +381 -0
  52. package/dist/chunk-FEAXX7B6.js +101 -0
  53. package/dist/{chunk-RLYRBIYQ.js → chunk-FFUPXJC4.js} +90 -331
  54. package/dist/{chunk-5PFYMY2V.js → chunk-FTMGRKEF.js} +2 -2
  55. package/dist/{chunk-34BHNEE3.js → chunk-GHS5EBTQ.js} +58 -7
  56. package/dist/{chunk-DYHAXKHD.js → chunk-GWZNEVM2.js} +12 -8
  57. package/dist/chunk-GX5WYQO4.js +59 -0
  58. package/dist/chunk-GYV6VZOC.js +26 -0
  59. package/dist/{chunk-MQXIVJ35.js → chunk-HAXOFFRH.js} +5 -5
  60. package/dist/chunk-I2DWJ4GM.js +390 -0
  61. package/dist/{chunk-TXOTCRLG.js → chunk-I5FWO7L5.js} +5 -5
  62. package/dist/{chunk-WNP7O5WZ.js → chunk-ID64D7PE.js} +4 -4
  63. package/dist/{chunk-XQRY4DTA.js → chunk-IGLP3ODT.js} +10 -10
  64. package/dist/chunk-IRXAATOX.js +539 -0
  65. package/dist/chunk-IXIY2H4R.js +44 -0
  66. package/dist/{chunk-SSEYRH53.js → chunk-IZXGRF7P.js} +92 -147
  67. package/dist/{chunk-5TSRNF4G.js → chunk-JCI2ROMZ.js} +164 -6
  68. package/dist/{chunk-JWJGP5DQ.js → chunk-JEIYHLOR.js} +7 -7
  69. package/dist/{chunk-F2I26BDK.js → chunk-JQLNNIKT.js} +4 -4
  70. package/dist/{chunk-BYMNWQ7O.js → chunk-JSD46VO2.js} +315 -63
  71. package/dist/{chunk-AK5XEFVZ.js → chunk-JT2RFCC5.js} +64 -14
  72. package/dist/{chunk-MCMZMDAC.js → chunk-K6T2ZAMZ.js} +168 -6
  73. package/dist/{chunk-PGF63K6I.js → chunk-KFV5L5SK.js} +73 -4
  74. package/dist/{chunk-IHXBNWMM.js → chunk-KXDSS5WJ.js} +7 -3
  75. package/dist/{chunk-PJX3WQUQ.js → chunk-LLXSDWXS.js} +3 -3
  76. package/dist/{chunk-DZAW46HP.js → chunk-LTIKRKFL.js} +3 -3
  77. package/dist/{chunk-DZEK6CJN.js → chunk-N56KALIC.js} +21 -21
  78. package/dist/{chunk-B7HM5Z7T.js → chunk-NAI6ZFCY.js} +9 -5
  79. package/dist/{chunk-I66ZTYNP.js → chunk-NRO2BJRH.js} +2656 -2213
  80. package/dist/{chunk-ZGNYYXQ6.js → chunk-NXIMQY5W.js} +3 -3
  81. package/dist/{chunk-IKOZFYBN.js → chunk-NXYCB2VD.js} +149 -106
  82. package/dist/{chunk-CWVRRIEI.js → chunk-NZMNUPZZ.js} +2 -2
  83. package/dist/{chunk-462T4EGZ.js → chunk-O5CVSAG5.js} +2 -2
  84. package/dist/chunk-ODGTEFFI.js +50 -0
  85. package/dist/{chunk-3F7VUY77.js → chunk-OEJSLEPW.js} +2 -2
  86. package/dist/{chunk-KKOJXO6R.js → chunk-OMQNJVKW.js} +4 -2
  87. package/dist/{chunk-5KW52TEP.js → chunk-Q4WO54TA.js} +132 -77
  88. package/dist/{chunk-W6NIE6OW.js → chunk-QUFRYSWI.js} +13 -7
  89. package/dist/{chunk-42FMPA75.js → chunk-QZWQA4DE.js} +2 -2
  90. package/dist/chunk-R6Q67RJH.js +134 -0
  91. package/dist/{chunk-W4YEMFBX.js → chunk-RAY4OVGZ.js} +3 -3
  92. package/dist/{chunk-ZNT2M6TG.js → chunk-RQCKCSRL.js} +17 -17
  93. package/dist/{chunk-LJID3DYZ.js → chunk-RXTN6AKH.js} +3 -3
  94. package/dist/{chunk-P75RZCJW.js → chunk-RZDWV63N.js} +3 -3
  95. package/dist/{chunk-UH632ZYL.js → chunk-S6PYF2XF.js} +2 -2
  96. package/dist/{chunk-HJWWJ6IL.js → chunk-TOIVGRUX.js} +17 -5
  97. package/dist/{chunk-HLAFFSEK.js → chunk-TQAHXW6Y.js} +2 -2
  98. package/dist/{chunk-JIEGK6UF.js → chunk-U6TMQNSI.js} +48 -4
  99. package/dist/{chunk-2HFQNRV3.js → chunk-UEPWCCTY.js} +12 -12
  100. package/dist/chunk-UOIZ7DA4.js +41 -0
  101. package/dist/{chunk-UH347SHR.js → chunk-USR47QNF.js} +11 -11
  102. package/dist/{chunk-AZ4WMN4W.js → chunk-V6HJFQZE.js} +2 -2
  103. package/dist/chunk-V76WTFTW.js +318 -0
  104. package/dist/{chunk-NMJXSHBJ.js → chunk-W54I7H25.js} +2 -2
  105. package/dist/{chunk-KPXDY6QF.js → chunk-XRZT5WY5.js} +2 -2
  106. package/dist/{chunk-UBRFI4HS.js → chunk-XULDXHTN.js} +142 -50
  107. package/dist/chunk-XXYSBZIQ.js +283 -0
  108. package/dist/{chunk-HKMD33FO.js → chunk-Y55JBDO5.js} +405 -122
  109. package/dist/{chunk-XOXV5GKE.js → chunk-YD5GIKET.js} +17 -8
  110. package/dist/{chunk-XGDPUNND.js → chunk-YECAMM3D.js} +2 -2
  111. package/dist/{chunk-BO7Y52RY.js → chunk-YNFKXPEC.js} +7 -7
  112. package/dist/{chunk-VXMFAE2W.js → chunk-YPC6ZR5L.js} +19 -6
  113. package/dist/{chunk-M2WXEHER.js → chunk-ZA4VCIGV.js} +2 -2
  114. package/dist/cli/index.js +42 -40
  115. package/dist/{clio-7VB377CC.js → clio-QLICPCF5.js} +7 -7
  116. package/dist/{code-nav-YVLCYA7V.js → code-nav-IJR2DBPR.js} +9 -9
  117. package/dist/{components-UBWCQSRW.js → components-2TGAI2RC.js} +5 -6
  118. package/dist/{config-4HVOS65E.js → config-IUA6OYNS.js} +88 -81
  119. package/dist/{configure-PIWO7B24.js → configure-VEPX4NMX.js} +26 -25
  120. package/dist/{context-KQYIWPWT.js → context-2DKHWH2T.js} +60 -45
  121. package/dist/{context-IYEHL3WQ.js → context-4MPR7WKB.js} +78 -69
  122. package/dist/{context-N6ZE3LGJ.js → context-BOYF5EJM.js} +15 -11
  123. package/dist/{context-clear-G4OGZJDS.js → context-clear-S4ZJCQUX.js} +73 -65
  124. package/dist/context-map-COB37XXN.js +505 -0
  125. package/dist/{context-working-set-BWLF6LJP.js → context-working-set-3I3FYX6Y.js} +18 -17
  126. package/dist/detail-A7JAVSIG.js +98 -0
  127. package/dist/{dispatch-runner-2QQAITS3.js → dispatch-runner-RJ5I2F2O.js} +99 -75
  128. package/dist/{docs-PD3EXDKU.js → docs-SPOV3BAN.js} +3 -5
  129. package/dist/{doctor-LHBD36VU.js → doctor-DKICC2SN.js} +71 -48
  130. package/dist/{eval-C45FYRJ6.js → eval-OQOQUDHK.js} +308 -146
  131. package/dist/{eval-inventory-6DEJPLBF.js → eval-inventory-Y6QRFOH5.js} +4 -4
  132. package/dist/{evidence-6SHONYAF.js → evidence-4DQ25GUQ.js} +79 -175
  133. package/dist/evidence-4F5USFKH.js +208 -0
  134. package/dist/{evolve-KRKMV72X.js → evolve-GSS52E5J.js} +71 -67
  135. package/dist/{extensions-KPZ2UHBB.js → extensions-G7MFLYHT.js} +8 -9
  136. package/dist/{fleet-IVTCKDHT.js → fleet-Q37YHHAQ.js} +126 -118
  137. package/dist/{fleet-commands-EDWL3IT7.js → fleet-commands-G7E4N7SM.js} +16 -13
  138. package/dist/{fleet-decisions-YP3YEFGK.js → fleet-decisions-O7M6QBA2.js} +9 -8
  139. package/dist/{fleet-graph-ZFWKHY2M.js → fleet-graph-TOUBW6OW.js} +21 -20
  140. package/dist/{fleet-inspect-FVUNCBML.js → fleet-inspect-SW33JJNI.js} +65 -60
  141. package/dist/{fleet-preflight-UN5XED4R.js → fleet-preflight-CV2655TW.js} +4 -5
  142. package/dist/{fleet-validate-XOWC4HSX.js → fleet-validate-KESZX2YH.js} +25 -24
  143. package/dist/{fleet-verify-UN3SODEL.js → fleet-verify-UQPMTVE3.js} +66 -61
  144. package/dist/{fleet-view-TWHJKCN6.js → fleet-view-MN2VG4MR.js} +65 -60
  145. package/dist/{init-T2QORQ3Y.js → init-PXEXQSBF.js} +90 -82
  146. package/dist/{interop-IN5I2A66.js → interop-ZG5T62U3.js} +12 -13
  147. package/dist/inventory-C26CFDRR.js +101 -0
  148. package/dist/{library-LSCATDLZ.js → library-B2W4N74O.js} +29 -29
  149. package/dist/{memory-HYOKAGGJ.js → memory-YCANYS5A.js} +73 -69
  150. package/dist/{models-2GPMFYCM.js → models-GERTU3YI.js} +51 -47
  151. package/dist/{monitor-E4ASVUJH.js → monitor-CPNIUULB.js} +74 -67
  152. package/dist/{orchestrator-DDMPR3PY.js → orchestrator-J4BSH4WQ.js} +1288 -1606
  153. package/dist/{panes-E3RUXOW5.js → panes-BOHAEGYC.js} +4 -4
  154. package/dist/{panes-IXKLOKA2.js → panes-NXSLDQZ2.js} +10 -11
  155. package/dist/{paths-L7LGY6RN.js → paths-VSUWNC22.js} +6 -7
  156. package/dist/reset-TNWTB5LU.js +343 -0
  157. package/dist/{resources-OTRSN34L.js → resources-4PXNMD5G.js} +29 -22
  158. package/dist/{run-5DEYH5QK.js → run-D6XJ34CN.js} +132 -132
  159. package/dist/{share-IHWTLO3M.js → share-2NWMJJEE.js} +27 -27
  160. package/dist/{skills-IYMXMKW4.js → skills-KR7WON5G.js} +40 -33
  161. package/dist/{skills-eval-DROHSJAR.js → skills-eval-O2ZNOLDS.js} +81 -77
  162. package/dist/{skills-inventory-D7X4L4ZX.js → skills-inventory-ZZOUBK7O.js} +23 -22
  163. package/dist/{slash-commands-QBM7UZ3B.js → slash-commands-ZXPJD64J.js} +47 -37
  164. package/dist/{steer-Z5DO23FJ.js → steer-XA25PSCS.js} +4 -4
  165. package/dist/{support-U7QOWY26.js → support-7EMVWYG2.js} +6 -6
  166. package/dist/{targets-P2FUC4IL.js → targets-OMH2XCSN.js} +50 -49
  167. package/dist/tasks-IPAGMEIX.js +36 -0
  168. package/dist/{terminal-lease-YREJ3JX2.js → terminal-lease-C2J3JYRE.js} +4 -4
  169. package/dist/{tools-5B7RO6MV.js → tools-EFFEAIDP.js} +8 -9
  170. package/dist/{trace-YMGMUM6A.js → trace-FXMXUZUF.js} +7 -7
  171. package/dist/uninstall-HALS6BLF.js +407 -0
  172. package/dist/upgrade-MS72RJEP.js +306 -0
  173. package/dist/{usage-ME5MPXGX.js → usage-NHG6MCJM.js} +162 -108
  174. package/dist/{verifiers-BVZ7IWOO.js → verifiers-7AUNVXDY.js} +155 -22
  175. package/dist/{verify-5K7ZKQFC.js → verify-FWYGPKMR.js} +14 -12
  176. package/dist/{web-fetch-MPARV2K7.js → web-fetch-V4FKSDAV.js} +4 -4
  177. package/dist/{wiki-generate-F5W5QTYY.js → wiki-generate-743CIGJW.js} +99 -90
  178. package/dist/{with-panes-BYOJCLAM.js → with-panes-BDQEWBRT.js} +10 -10
  179. package/dist/worker/entry.js +72 -68
  180. package/docs/README.md +3 -2
  181. package/docs/architecture/acp.md +17 -0
  182. package/docs/architecture/artifact-placement.md +1 -0
  183. package/docs/architecture/artifact-versions.md +2 -2
  184. package/docs/architecture/context-engine.md +4 -0
  185. package/docs/architecture/dispatch-typed-intent.md +1 -1
  186. package/docs/architecture/evidence-and-memory.md +1 -1
  187. package/docs/architecture/middleware-and-components.md +1 -1
  188. package/docs/architecture/model-catalog.md +21 -10
  189. package/docs/architecture/observability.md +19 -2
  190. package/docs/architecture/prompt-envelope-and-tools.md +17 -5
  191. package/docs/architecture/provider-adapter-cookbook.md +63 -0
  192. package/docs/architecture/safety-model.md +25 -22
  193. package/docs/architecture/tui-design.md +1 -1
  194. package/docs/guide/built-in-agents.md +25 -11
  195. package/docs/guide/commands-and-modes.md +18 -3
  196. package/docs/guide/configuration-and-targets.md +100 -10
  197. package/docs/guide/configuration-reference.md +17 -7
  198. package/docs/guide/environment-variables.md +4 -2
  199. package/docs/guide/installation-and-lifecycle.md +37 -4
  200. package/docs/guide/proactive-memory.md +66 -55
  201. package/docs/guide/skills-marketplace.md +18 -0
  202. package/docs/guide/tool-usage.md +78 -3
  203. package/docs/history/config-knobs-audit.md +2 -2
  204. package/docs/process/development-pipeline.md +40 -2
  205. package/docs/process/eval-runner.md +67 -3
  206. package/docs/process/git-commit-provenance.md +15 -0
  207. package/docs/process/release-cut-checklist.md +207 -0
  208. package/docs/process/scientific-validation.md +18 -17
  209. package/evals/behavioral-machinery-support.ts +1 -0
  210. package/evals/behavioral-machinery.yaml +1 -1
  211. package/evals/behavioral-model.yaml +3 -2
  212. package/package.json +2 -2
  213. package/skills/README.md +7 -5
  214. package/skills/coding/ast-grep/SKILL.md +101 -30
  215. package/skills/coding/ast-grep/evals.md +26 -0
  216. package/skills/coding/coding-standards/SKILL.md +47 -2
  217. package/skills/coding/coding-standards/evals.md +23 -0
  218. package/skills/coding/prototype/SKILL.md +87 -28
  219. package/skills/coding/prototype/evals.md +19 -0
  220. package/skills/coding/tdd/SKILL.md +80 -53
  221. package/skills/coding/tdd/evals.md +20 -0
  222. package/skills/context/context-handoff/SKILL.md +43 -2
  223. package/skills/context/context-handoff/evals.md +44 -0
  224. package/skills/context/context-prime/SKILL.md +45 -15
  225. package/skills/context/context-prime/evals.md +45 -0
  226. package/skills/git/branch-closeout/SKILL.md +132 -0
  227. package/skills/git/branch-closeout/evals.md +133 -0
  228. package/skills/git/branch-closeout/references/closeout-checklist.md +81 -0
  229. package/skills/git/file-ticket/SKILL.md +77 -63
  230. package/skills/git/file-ticket/assets/issue-template.md +22 -0
  231. package/skills/git/file-ticket/evals.md +31 -26
  232. package/skills/git/file-ticket/references/issue-discovery.md +49 -0
  233. package/skills/git/fix-issue/SKILL.md +87 -64
  234. package/skills/git/fix-issue/evals.md +35 -31
  235. package/skills/git/fix-issue/references/diagnosis-and-rca.md +46 -0
  236. package/skills/git/resolve-merge-conflicts/SKILL.md +100 -51
  237. package/skills/git/resolve-merge-conflicts/evals.md +52 -25
  238. package/skills/git/resolve-merge-conflicts/references/conflict-matrix.md +126 -0
  239. package/skills/git/ship/SKILL.md +103 -67
  240. package/skills/git/ship/assets/pr-template.md +21 -0
  241. package/skills/git/ship/evals.md +44 -28
  242. package/skills/git/ship/references/remote-and-branch-policy.md +62 -0
  243. package/skills/git/worktree-create/SKILL.md +80 -50
  244. package/skills/git/worktree-create/evals.md +40 -33
  245. package/skills/git/worktree-create/references/worktree-setup.md +62 -66
  246. package/skills/git/worktree-merge/SKILL.md +112 -65
  247. package/skills/git/worktree-merge/evals.md +42 -34
  248. package/skills/git/worktree-merge/references/merge-strategies.md +52 -0
  249. package/skills/planning/archify/SKILL.md +196 -0
  250. package/skills/planning/archify/evals.md +65 -0
  251. package/skills/planning/architecture/SKILL.md +61 -12
  252. package/skills/planning/architecture/evals.md +65 -0
  253. package/skills/planning/backlog/SKILL.md +130 -14
  254. package/skills/planning/backlog/evals.md +142 -0
  255. package/skills/planning/prd/SKILL.md +47 -6
  256. package/skills/planning/prd/evals.md +54 -0
  257. package/skills/planning/product-intent/SKILL.md +57 -2
  258. package/skills/planning/product-intent/evals.md +70 -0
  259. package/skills/planning/tech-spec/SKILL.md +53 -2
  260. package/skills/planning/tech-spec/evals.md +73 -0
  261. package/skills/registry.yaml +58 -50
  262. package/skills/remote.yaml +13 -0
  263. package/skills/research/arxiv-literature/SKILL.md +76 -18
  264. package/skills/research/arxiv-literature/evals.md +50 -0
  265. package/skills/research/experiment-protocol/SKILL.md +20 -1
  266. package/skills/research/experiment-protocol/evals.md +23 -0
  267. package/skills/research/scientific-debugging/SKILL.md +24 -1
  268. package/skills/research/scientific-debugging/evals.md +18 -0
  269. package/skills/research/scientific-modernization/SKILL.md +26 -1
  270. package/skills/research/scientific-modernization/evals.md +27 -0
  271. package/skills/skill-marketplace.json +63 -28
  272. package/skills/workflow/cut-it/SKILL.md +64 -5
  273. package/skills/workflow/cut-it/evals.md +101 -0
  274. package/skills/workflow/design-council/SKILL.md +112 -27
  275. package/skills/workflow/design-council/evals.md +161 -0
  276. package/skills/workflow/grill-me/SKILL.md +85 -10
  277. package/skills/workflow/grill-me/evals.md +153 -0
  278. package/skills/workflow/workflow-distiller/SKILL.md +76 -17
  279. package/skills/workflow/workflow-distiller/evals.md +118 -0
  280. package/src/cli/args.ts +0 -8
  281. package/src/cli/configure-interop.ts +105 -13
  282. package/src/cli/configure-oauth.ts +57 -0
  283. package/src/cli/configure-onboarding.ts +980 -0
  284. package/src/cli/configure-target.ts +594 -0
  285. package/src/cli/configure.ts +1084 -529
  286. package/src/cli/context-map.ts +114 -0
  287. package/src/cli/context.ts +4 -0
  288. package/src/cli/doctor-state-size.ts +1 -12
  289. package/src/cli/doctor-validation-contract.ts +28 -0
  290. package/src/cli/doctor.ts +5 -0
  291. package/src/cli/evidence-detail.ts +1 -75
  292. package/src/cli/evidence-inventory.ts +1 -167
  293. package/src/cli/index.ts +3 -0
  294. package/src/cli/lifecycle-presenter.ts +436 -0
  295. package/src/cli/models.ts +10 -2
  296. package/src/cli/modes/print.ts +5 -1
  297. package/src/cli/reset.ts +228 -106
  298. package/src/cli/run.ts +7 -4
  299. package/src/cli/select.ts +664 -0
  300. package/src/cli/skills.ts +9 -2
  301. package/src/cli/targets.ts +3 -0
  302. package/src/cli/tasks.ts +84 -0
  303. package/src/cli/uninstall.ts +233 -165
  304. package/src/cli/upgrade.ts +210 -150
  305. package/src/cli/usage.ts +92 -27
  306. package/src/cli/validate-model.ts +3 -3
  307. package/src/cli/verifiers.ts +147 -1
  308. package/src/cli/wiki-generate.ts +1 -0
  309. package/src/core/commit-attribution.ts +41 -1
  310. package/src/core/config.ts +56 -0
  311. package/src/core/external-diagnostic.ts +44 -0
  312. package/src/core/gateway-routing.ts +157 -0
  313. package/src/core/git-commit-attribution.ts +46 -3
  314. package/src/core/run-overrides.ts +0 -5
  315. package/src/core/safe-exec.ts +17 -2
  316. package/src/core/skill-activation.ts +92 -2
  317. package/src/core/tool-names.ts +5 -2
  318. package/src/domains/agents/builtins/architect.md +1 -1
  319. package/src/domains/agents/builtins/coder.md +1 -1
  320. package/src/domains/agents/builtins/documenter.md +1 -1
  321. package/src/domains/agents/builtins/git-master.md +1 -1
  322. package/src/domains/agents/builtins/provenance.md +7 -7
  323. package/src/domains/agents/builtins/tester.md +1 -1
  324. package/src/domains/agents/builtins/verifier.md +2 -2
  325. package/src/domains/agents/builtins/wiki-writer.md +4 -3
  326. package/src/domains/agents/builtins/world-knowledge.md +31 -0
  327. package/src/domains/agents/catalog.ts +1 -1
  328. package/src/domains/agents/result-contract.ts +70 -0
  329. package/src/domains/context/extension.ts +31 -7
  330. package/src/domains/context/refresh.ts +3 -0
  331. package/src/domains/context/wiki/frontmatter.ts +5 -2
  332. package/src/domains/context/wiki/generate.ts +6 -0
  333. package/src/domains/context/wiki/map-seed.ts +589 -0
  334. package/src/domains/context/wiki/plan.ts +2 -2
  335. package/src/domains/context/wiki/prompts.ts +43 -0
  336. package/src/domains/dispatch/active-route-planner.ts +4 -0
  337. package/src/domains/dispatch/admission.ts +29 -0
  338. package/src/domains/dispatch/agent-candidates.ts +10 -0
  339. package/src/domains/dispatch/budget-envelope.ts +86 -1
  340. package/src/domains/dispatch/capability-match.ts +1 -0
  341. package/src/domains/dispatch/capacity-lease.ts +17 -0
  342. package/src/domains/dispatch/code-step.ts +11 -4
  343. package/src/domains/dispatch/contract.ts +42 -8
  344. package/src/domains/dispatch/execution-scheduler.ts +2 -0
  345. package/src/domains/dispatch/extension.ts +184 -56
  346. package/src/domains/dispatch/fleet-commit-attribution.ts +5 -0
  347. package/src/domains/dispatch/fleet-run.ts +1 -0
  348. package/src/domains/dispatch/host-verification.ts +114 -13
  349. package/src/domains/dispatch/intent.ts +28 -18
  350. package/src/domains/dispatch/orphan-recovery.ts +2 -0
  351. package/src/domains/dispatch/receipt-integrity.ts +4 -0
  352. package/src/domains/dispatch/reservation-store.ts +5 -3
  353. package/src/domains/dispatch/state.ts +15 -2
  354. package/src/domains/dispatch/types.ts +20 -2
  355. package/src/domains/dispatch/worker-model-metadata.ts +38 -0
  356. package/src/domains/eval/metrics/call-ledger-stream.ts +34 -11
  357. package/src/domains/eval/metrics/token-stream.ts +201 -31
  358. package/src/domains/eval/metrics/tracked.ts +40 -4
  359. package/src/domains/eval/runners/clio-run.ts +17 -11
  360. package/src/domains/eval/runners/context-index.ts +2 -7
  361. package/src/domains/eval/runners/context-init.ts +3 -6
  362. package/src/domains/eval/runners/external-command.ts +29 -11
  363. package/src/domains/eval/schema/suite.ts +28 -0
  364. package/src/domains/eval/schema/verdict.ts +2 -2
  365. package/src/domains/eval/suites/resolve.ts +13 -1
  366. package/src/domains/eval/suites/run.ts +24 -3
  367. package/src/domains/evidence/build.ts +102 -15
  368. package/src/domains/evidence/detail.ts +69 -0
  369. package/src/domains/evidence/eval.ts +13 -1
  370. package/src/domains/evidence/finish-contract-map.ts +5 -1
  371. package/src/domains/evidence/inventory.ts +167 -0
  372. package/src/domains/evidence/store.ts +16 -0
  373. package/src/domains/evidence/types.ts +12 -0
  374. package/src/domains/extensions/resources.ts +7 -0
  375. package/src/domains/interop/registry.ts +6 -2
  376. package/src/domains/interop/types.ts +4 -0
  377. package/src/domains/lifecycle/migrations/index.ts +4 -0
  378. package/src/domains/memory/task-memory-policy.ts +70 -26
  379. package/src/domains/memory/task-memory-telemetry.ts +1 -0
  380. package/src/domains/middleware/index.ts +0 -1
  381. package/src/domains/middleware/marketplace-offer.ts +22 -35
  382. package/src/domains/middleware/memory-intervention.ts +127 -32
  383. package/src/domains/middleware/memory-step-endpoint.ts +3 -2
  384. package/src/domains/middleware/runtime.ts +7 -3
  385. package/src/domains/middleware/skills-reminder.ts +31 -2
  386. package/src/domains/mux/detect.ts +3 -6
  387. package/src/domains/observability/accountability.ts +15 -1
  388. package/src/domains/observability/compaction-usage.ts +118 -0
  389. package/src/domains/observability/contract.ts +52 -7
  390. package/src/domains/observability/cost.ts +1 -1
  391. package/src/domains/observability/evidence-index.ts +10 -0
  392. package/src/domains/observability/extension.ts +15 -6
  393. package/src/domains/observability/out-of-turn-usage.ts +52 -21
  394. package/src/domains/observability/projection.ts +394 -45
  395. package/src/{interactive → domains/observability}/worker-progress.ts +3 -3
  396. package/src/domains/prompts/fragments/operating/contract.md +2 -0
  397. package/src/domains/prompts/fragments/wiki/page.md +8 -0
  398. package/src/domains/providers/contract.ts +4 -1
  399. package/src/domains/providers/extension.ts +40 -9
  400. package/src/domains/providers/model-capabilities.ts +9 -0
  401. package/src/domains/providers/model-discovery.ts +3 -4
  402. package/src/domains/providers/model-runtime-capabilities.ts +15 -5
  403. package/src/domains/providers/models/local-models/clio-coder-local-coding-targets.yaml +48 -26
  404. package/src/domains/providers/runtimes/antigravity/antigravity-code.ts +225 -45
  405. package/src/domains/providers/runtimes/claude/claude-code.ts +9 -0
  406. package/src/domains/providers/runtimes/common/lmstudio-http.ts +6 -2
  407. package/src/domains/providers/runtimes/common/local-synth.ts +2 -0
  408. package/src/domains/providers/runtimes/protocol/litellm.ts +119 -29
  409. package/src/domains/providers/support.ts +11 -5
  410. package/src/domains/providers/target-model-cache.ts +25 -2
  411. package/src/domains/providers/types/capability-flags.ts +2 -0
  412. package/src/domains/providers/types/runtime-descriptor.ts +20 -1
  413. package/src/domains/providers/types/target-descriptor.ts +19 -0
  414. package/src/domains/resources/index.ts +3 -0
  415. package/src/domains/resources/skills/install.ts +72 -7
  416. package/src/domains/resources/skills/loader.ts +7 -0
  417. package/src/domains/resources/skills/marketplace.ts +63 -11
  418. package/src/domains/safety/action-classifier.ts +7 -0
  419. package/src/domains/safety/autonomy.ts +15 -0
  420. package/src/domains/safety/default-path-policy.ts +2 -0
  421. package/src/domains/safety/finish-contract-registration.ts +29 -14
  422. package/src/domains/safety/finish-contract.ts +252 -40
  423. package/src/domains/safety/index.ts +21 -1
  424. package/src/domains/safety/path-policy.ts +1 -1
  425. package/src/domains/safety/policy-engine.ts +60 -17
  426. package/src/domains/safety/protected-artifacts.ts +191 -88
  427. package/src/domains/safety/rigor.ts +53 -39
  428. package/src/domains/safety/run-effects.ts +2 -22
  429. package/src/domains/safety/skill-authority.ts +55 -0
  430. package/src/domains/safety/validation-contract.ts +388 -0
  431. package/src/domains/session/archive-readers.ts +10 -1
  432. package/src/domains/session/compaction/compact.ts +72 -22
  433. package/src/domains/session/decision-board.ts +101 -2
  434. package/src/domains/session/entries.ts +50 -7
  435. package/src/domains/session/extension.ts +4 -4
  436. package/src/domains/session/handoff.ts +2 -1
  437. package/src/domains/session/manager.ts +2 -3
  438. package/src/domains/session/task-board.ts +14 -1
  439. package/src/domains/session/tree/fork.ts +1 -2
  440. package/src/domains/session/tree/navigator.ts +1 -1
  441. package/src/domains/session/usage.ts +3 -3
  442. package/src/domains/user-tasks/acceptance.ts +56 -0
  443. package/src/domains/user-tasks/active-acceptance.ts +40 -0
  444. package/src/domains/user-tasks/store.ts +34 -3
  445. package/src/engine/acp/adapter.ts +24 -6
  446. package/src/engine/acp/server.ts +21 -4
  447. package/src/engine/acp/transport.ts +53 -8
  448. package/src/engine/acp/types.ts +4 -0
  449. package/src/engine/agent.ts +13 -3
  450. package/src/engine/ai.ts +26 -8
  451. package/src/engine/antigravity/subprocess-runtime.ts +386 -120
  452. package/src/engine/api-registry.ts +3 -0
  453. package/src/engine/apis/ollama-native.ts +15 -0
  454. package/src/engine/apis/openai-completions.ts +117 -14
  455. package/src/engine/claude/subprocess-runtime.ts +107 -60
  456. package/src/engine/external-subprocess.ts +122 -6
  457. package/src/entry/background-model-metadata.ts +18 -0
  458. package/src/entry/compaction-prompt.ts +57 -0
  459. package/src/entry/orchestrator.ts +416 -218
  460. package/src/entry/task-memory-lifecycle.ts +35 -0
  461. package/src/interactive/chat-loop-messages.ts +13 -4
  462. package/src/interactive/chat-loop.ts +65 -2
  463. package/src/interactive/chat-renderer.ts +1 -0
  464. package/src/interactive/cost-overlay.ts +26 -2
  465. package/src/interactive/dispatch-board.ts +46 -717
  466. package/src/interactive/fleet-run-preview.ts +2 -1
  467. package/src/interactive/interactive-application.ts +3 -2
  468. package/src/interactive/interactive-presentation.ts +55 -12
  469. package/src/interactive/interactive-slash-runtime.ts +4 -2
  470. package/src/interactive/oracle.ts +5 -2
  471. package/src/interactive/overlays/fleet-run-approval.ts +3 -2
  472. package/src/interactive/overlays/message-picker.ts +2 -2
  473. package/src/interactive/overlays/settings.ts +2 -2
  474. package/src/interactive/overlays/tree-selector.ts +2 -2
  475. package/src/interactive/renderers/branch-summary.ts +1 -1
  476. package/src/interactive/renderers/worker-entry.ts +32 -0
  477. package/src/interactive/slash-autocomplete.ts +4 -6
  478. package/src/interactive/slash-commands.ts +49 -45
  479. package/src/interactive/slash-spec.ts +28 -0
  480. package/src/interactive/theme/labels.ts +19 -13
  481. package/src/interactive/turn-context.ts +9 -5
  482. package/src/interactive/turn-recovery.ts +8 -0
  483. package/src/interactive/turn-runtime.ts +27 -11
  484. package/src/interactive/turn-state.ts +7 -0
  485. package/src/interactive/view/artifacts.ts +2 -0
  486. package/src/interactive/worker-receipts.ts +1 -0
  487. package/src/interactive/worker-stream.ts +13 -4
  488. package/src/tools/bootstrap.ts +4 -0
  489. package/src/tools/builtin-tool-catalog.ts +31 -0
  490. package/src/tools/compete-worktrees.ts +7 -1
  491. package/src/tools/context/index.ts +30 -9
  492. package/src/tools/core-bootstrap.ts +16 -0
  493. package/src/tools/decide.ts +136 -0
  494. package/src/tools/dispatch-admission.ts +21 -0
  495. package/src/tools/dispatch-arguments.ts +1 -0
  496. package/src/tools/dispatch-event-text.ts +10 -0
  497. package/src/tools/dispatch-plan.ts +6 -2
  498. package/src/tools/dispatch-runner.ts +27 -1
  499. package/src/tools/dispatch-types.ts +6 -0
  500. package/src/tools/evidence.ts +96 -0
  501. package/src/tools/limitation.ts +76 -0
  502. package/src/tools/policy.ts +9 -0
  503. package/src/tools/presentation.ts +3 -0
  504. package/src/tools/registry.ts +11 -5
  505. package/src/tools/result-shaping.ts +17 -5
  506. package/src/tools/task-worktree.ts +13 -3
  507. package/src/tools/tasks.ts +10 -1
  508. package/src/tools/verify/authoring.ts +170 -83
  509. package/src/tools/verify/catalog.ts +122 -5
  510. package/src/tools/verify/index.ts +2 -1
  511. package/src/tools/verify/numeric.ts +298 -0
  512. package/src/tools/verify/perf.ts +143 -0
  513. package/src/tools/verify/scripts.ts +229 -2
  514. package/src/tools/worker-evidence.ts +3 -1
  515. package/src/worker/spec-contract.ts +4 -0
  516. package/dist/chunk-2Z2IKEXI.js +0 -1554
  517. package/dist/chunk-RVG5JXAL.js +0 -41
  518. package/dist/chunk-T56WDKA5.js +0 -183
  519. package/dist/chunk-VN3SHNBN.js +0 -313
  520. package/dist/chunk-VPKWYKEY.js +0 -169
  521. package/dist/reset-OAQP3W4O.js +0 -230
  522. package/dist/uninstall-N34PCTGJ.js +0 -331
  523. package/dist/upgrade-PXK3S2YM.js +0 -325
@@ -121,7 +121,7 @@ malformed response, or telemetry failure is silent and never blocks a tool.
121
121
  - `context.memory.cadenceToolCalls` (default `10`): Minimum completed-tool interval between background interventions.
122
122
  - `context.memory.trajectorySteps` (default `8`): Completed tool-trajectory window analyzed during background evaluation.
123
123
  - `context.memory.maxOutputTokens` (default `2000`): Bounds the rendered memory-bank context and the ordinary policy-model completion. An always-on-thinking model receives additional reasoning headroom, at least `4,000` tokens when its model cap permits, so it can still reach the strict envelope.
124
- - `context.memory.timeoutMs` (default `60000`): Wall-clock limit for one background memory-policy request. The step is detached, so this deadline never delays a turn, but it does hold a request slot on a real inference endpoint that your own turns and dispatched workers queue against. Raise it only after inspecting the timeout and hit-rate evidence for the selected route.
124
+ - `context.memory.timeoutMs` (default `60000`): Wall-clock limit for one background memory-policy step, including its one permitted chat fallback attempt. The step is detached, so this deadline never delays a turn, but it does hold a request slot on a real inference endpoint that your own turns and dispatched workers queue against. Raise it only after inspecting the timeout and hit-rate evidence for the selected route.
125
125
 
126
126
  ## Trigger semantics
127
127
 
@@ -254,38 +254,48 @@ unexamined one:
254
254
  would not have been spent. The current 60-second deadline is the source-backed
255
255
  bound on what one optional call may hold a shared local server for, not a
256
256
  prediction of when a route answers.
257
- - A step that would run on the endpoint the chat target is streaming against is
258
- skipped with reason `endpoint_busy`, and the skip is recorded. On a single-slot
259
- llama.cpp router the alternative is queueing behind the operator's own decoding
260
- or evicting the resident model, and neither is a cost an optional call may
261
- impose.
262
-
263
- ### What a background target costs on a shared local server
264
-
265
- If `context.memory.target` names the same server as `chat.target`, that
266
- server's slots are shared. On a llama.cpp router started with `--parallel 1`
267
- there is exactly one, and the memory step and the operator's turn contend for it.
268
-
269
- The consequence is worth stating plainly: a shared endpoint suppresses the model
270
- tier rather than merely delaying it. A step is started from the `turn_end` hook,
271
- which fires inside the streaming run at `agent_end`
272
- (`src/interactive/turn-runtime.ts`), while the chat loop still holds its
273
- foreground registration on that endpoint; the loop releases the hold afterwards,
274
- in the `finally` around the run (`src/interactive/chat-loop.ts`). Every boundary
275
- therefore finds the endpoint busy and records `dropped`/`endpoint_busy`. That is
276
- the intended trade: an optional call may not take the one slot the operator's own
277
- turn is using, and it may not make the server swap the resident model out. The
278
- `/memory` step list and `steps.jsonl` say so on every boundary, so the tier is
279
- visibly declining rather than quietly idle.
280
-
281
- The second mechanism is an `expected cold` stamp, for the case where a step did
282
- run on the chat endpoint. Its prompt is a trajectory rather than the chat prefix,
283
- so the next turn's prefill is expected to be cold; `/context` names
284
- `background_memory` as the reason instead of reporting an unexplained cold
285
- prefix.
286
-
287
- Pointing the background role at a second machine avoids both effects and is the
288
- arrangement the tier is designed for.
257
+ - A step whose known endpoint occupancy exhausts the resolved request capacity is
258
+ skipped with reason `endpoint_busy`. The same gateway URL alone does not imply
259
+ a single request slot or a single physical model server.
260
+
261
+ ### Dedicated routing, chat fallback and endpoint capacity
262
+
263
+ Clio prefers the explicitly configured memory target and model. If that route is
264
+ known unavailable (missing target/runtime, down target, absent model in a known
265
+ catalog, or a model reported unloaded/loading), Clio selects the active chat
266
+ route for that step. A runtime client error on the dedicated route permits one
267
+ chat attempt within the original remaining deadline. It does not retry after a
268
+ timeout, cancellation, session/branch switch, malformed envelope, or a model's
269
+ explicit silence. The fallback never falls back again. Both routes unavailable
270
+ produce a visible `client_error` outcome. An unset memory role remains rules-only;
271
+ chat fallback does not enable model-based memory by default.
272
+
273
+ A runtime notice names the selected chat fallback and its reason. Routing stays
274
+ session-local and does not edit saved settings. Each attempted call records its
275
+ own known usage under its actual target/model and emits its own telemetry outcome;
276
+ failed dedicated usage is retained alongside fallback usage. The final policy
277
+ result describes the final attempt, without merging different provider identities.
278
+ A generation change discards late content while preserving the original usage owner.
279
+
280
+ Capacity follows the same existing evidence as dispatch: explicit target
281
+ `maxConcurrentRequests`, current discovery, a valid discovery prior, then the
282
+ conservative one-slot default for a local-native runtime. Known occupancy includes
283
+ this process's foreground requests, active dispatch leases and held reservation
284
+ waves. A two-slot endpoint with one foreground request can admit memory; a full
285
+ one-slot endpoint skips it. Known dedicated saturation does not itself initiate
286
+ chat fallback. Larger declared capacities work without a new constant.
287
+
288
+ LiteLLM is a gateway protocol, so Clio does not invent a local one-slot limit for
289
+ its URL. Distinct model routes such as dynamo, mini and zbook can share that URL;
290
+ the gateway owns their physical routing and backend residency. An explicit or
291
+ observed endpoint-wide bound still applies when present. Clio's process-local
292
+ foreground holds and observed dispatch state are not a global scheduler for every
293
+ client using the gateway. Slot holds are released on success, failure and abort.
294
+
295
+ The existing `background_memory` expected-cold stamp remains keyed by endpoint
296
+ URL. It records a possible shared-endpoint cache disturbance, including after a
297
+ failed request, rather than proving that a different model behind the gateway
298
+ actually evicted the chat prefix.
289
299
 
290
300
  ## Choosing a background model
291
301
 
@@ -294,9 +304,11 @@ does not need to be clever. A small non-reasoning model is the right choice, and
294
304
  Clio always requests the memory route with thinking off. Version 2 therefore
295
305
  has no configurable memory thinking-level key.
296
306
 
297
- That request reaches the wire wherever the runtime carries a thinking control:
307
+ The off request depends on the resolved runtime and model metadata:
298
308
  llama.cpp reads `chat_template_kwargs.enable_thinking`, and LM Studio reads
299
- `reasoning_effort`, where `none` is the off value.
309
+ `reasoning_effort`, with the off value selected by the model family. A gateway
310
+ needs a recognized, unanimous upstream runtime declaration; unknown metadata
311
+ does not establish a dialect. Requesting off does not prove server compliance.
300
312
 
301
313
  A model that reasons anyway still works. Some genuinely cannot be silenced, and
302
314
  the catalog records those as always-on so the level reads `forced` rather than
@@ -306,12 +318,11 @@ blocks are discarded and only the envelope is kept, and the output budget is
306
318
  sized to let a reasoning preamble run its course first. The cost is latency,
307
319
  which the detached step absorbs.
308
320
 
309
- One configuration is refused rather than degraded. If the background role names
310
- the same target and model as the orchestrator, and that model reasons, the LLM
311
- memory tier stays off and memory runs on its free deterministic tier. A single
312
- reasoning model already driving chat, workers, and shadow agents cannot also
313
- deliberate over memory steps without contending with the work the operator
314
- actually asked for.
321
+ The active chat model may also serve memory, including a reasoning model, when
322
+ request capacity permits. Memory still requests thinking off and uses its own
323
+ bounded envelope and output budget. A dedicated small model is preferred for
324
+ latency and resource use; sharing a model is a capacity decision rather than an
325
+ unconditional refusal.
315
326
 
316
327
  This is a mix-and-match plane, not a local-only one. The background role resolves
317
328
  through the same target machinery as every other role, so the useful shapes are:
@@ -356,30 +367,30 @@ independent of the chat target and the fleet default. A running session owns
356
367
  its routing snapshot, while the saved
357
368
  selection becomes the default for new sessions.
358
369
 
359
- The reference topology separates chat from memory work:
370
+ An example gateway topology keeps the three backend routes distinct:
360
371
 
361
372
  | Role | Target | Runtime and endpoint | Model | Capacity |
362
373
  | --- | --- | --- | --- | --- |
363
- | Chat | `dynamo` | LM Studio at `192.168.86.143:1234` | `qwen3.8-27b-dynamo` | Reported by the LM Studio probe |
364
- | Background memory | `mini` | llama.cpp router at `192.168.86.141:8080` | `ornith1.5-35b-moe` | 4 parallel slots, 262,144 context tokens per slot |
374
+ | Chat and memory fallback | `dynamo` | LiteLLM at `http://gateway:4000` | `dynamo/qwen3.8-27b` | Gateway-owned unless an explicit or observed endpoint bound is available |
375
+ | Preferred background memory | `zbook` | Same LiteLLM gateway | `zbook/ornith-1.5-35b-a3b` | Same capacity evidence rules |
376
+ | Alternative memory candidate | `mini` | Same LiteLLM gateway | `mini/ornith1.5-35b-moe` | Same capacity evidence rules |
365
377
 
366
- The corresponding saved role selection is:
378
+ The corresponding role selection is:
367
379
 
368
380
  ```yaml
369
381
  chat:
370
382
  target: dynamo
371
- model: qwen3.8-27b-dynamo
383
+ model: dynamo/qwen3.8-27b
372
384
  context:
373
385
  memory:
374
- target: mini
375
- model: ornith1.5-35b-moe
386
+ target: zbook
387
+ model: zbook/ornith-1.5-35b-a3b
376
388
  ```
377
389
 
378
- A smaller, efficient model is the intended shape for this role. The current
379
- reference uses a separate multi-slot endpoint so memory does not compete with
380
- the LM Studio chat stream. Capability is not the only constraint; latency and
381
- endpoint contention still matter, and the detached step above bounds their
382
- effect on the operator's turn.
390
+ These are example configured target/model IDs, not built-in routes. Compare a
391
+ slow dedicated model with an alternative using actual step latency, valid bank
392
+ writes and known usage; the active chat route remains the fallback. Do not create
393
+ aliases or change endpoint URLs just to bypass capacity accounting.
383
394
 
384
395
  The deadline is a bound on what an optional call may hold that server for, not a
385
396
  figure sized to capture the tail. The shipped `60000` is a bounded compromise,
@@ -489,7 +500,7 @@ clear the prior heap bank before the new session can observe it.
489
500
 
490
501
  ## Telemetry
491
502
 
492
- Each completed memory step appends one content-free record to:
503
+ Each completed memory-policy attempt appends one content-free record to:
493
504
 
494
505
  ```text
495
506
  <stateDir>/memory/steps.jsonl
@@ -512,7 +523,7 @@ folds out of this file.
512
523
  `dropped` is the one outcome that ran no step. It has two causes, separated by
513
524
  the row's reason: `step_in_flight` means the boundary triggered while an earlier
514
525
  step still held the single in-flight slot, and `endpoint_busy` means the step
515
- would have called the endpoint the chat target was streaming against. Both cost
526
+ found the known endpoint request capacity exhausted. Both cost
516
527
  no tokens and no latency, both leave their triggers pending for the next free
517
528
  boundary, and neither replaces the operator-visible last decision. Counting
518
529
  `dropped` rows against `llm` rows over a session is how a starved cadence becomes
@@ -5,6 +5,14 @@
5
5
 
6
6
  The Skills Hub (`/skill`) shows project skills, user skills, and the marketplace. Every marketplace row comes from the same local lookup that `clio-coder skills install <name>` and `/skill <name>` resolve through, so the hub lists nothing it cannot install.
7
7
 
8
+ ## Operator ownership of installed skills
9
+
10
+ The active project `.clio-coder/skills/` tree and the resolved user `<configDir>/skills/` tree are operator-owned. Main-agent and worker tool admissions refuse writes, edits, artifacts and recognized shell mutations to either tree, including ancestor deletion and current symlink aliases. This boundary stays active at every autonomy level, even when the general default path policy is disabled. Reads and loading already installed skills retain their existing rules.
11
+
12
+ Draft a proposed skill outside these active trees, for example in `draft-skills/`. The operator installs or updates it through `clio-coder skills install`, `skills update`, `skills sync`, `library add --yes`, the Skills Hub, `/skill <name>`, or an explicitly accepted marketplace offer. Even `full-auto` must wait for that offer's bound operator answer; a task match alone no longer installs a skill. Model shell calls to recognizable Clio skill installation, update and sync commands are refused. Confirmed `library add --yes` commands are also reserved for operators, including additions of other resource kinds that may install skill dependencies. An unconfirmed `library add` only prints a plan and remains available, as do inventory, search, inspection and validation commands.
13
+
14
+ This is a tool-admission boundary, not an operating-system sandbox. Shell inspection covers literal paths and recognized command forms, respecting comments, quoted words and redirection operands when identifying Clio commands. Quoted search patterns remain read-only, and literal operator filenames do not hide later mutation operands. It cannot prove the effects of arbitrary scripts, dynamically constructed paths, aliases, or concurrent filesystem changes. Keep shell execution supervised when stronger confinement is required.
15
+
8
16
  ## Where marketplace rows come from
9
17
 
10
18
  There is one source, `discoverMarketplaceSkills()` in `src/domains/resources/skills/marketplace.ts`, and it reads two kinds of real local data:
@@ -62,6 +70,16 @@ creation. Extension resource roots and share archives are documented in
62
70
  [extensions-and-sharing.md](extensions-and-sharing.md); this page owns the TUI
63
71
  Hub and marketplace behavior.
64
72
 
73
+ ## Remote entries and overlays
74
+
75
+ Some skills are worth carrying in the catalog without vendoring their content. `skills/remote.yaml` lists them: each entry names a skill, its category, a `sourceUrl` that must be a GitHub tree URL at a pinned tag, an `overlay` package inside the catalog, and an optional `exclude` list of upstream top-level members. `npm run skills:pin` publishes such an entry into `skills/skill-marketplace.json` with `origin: "remote"`, the upstream URL as its `sourceUrl`, and the `overlay` and `exclude` fields attached. The overlay's `SKILL.md` is pinned in `skills/registry.yaml` like every other catalog skill, so `npm run skills:check` fails when it drifts.
76
+
77
+ `archify` is the worked example. Its entry points at `https://github.com/tt-a1i/archify/tree/v2.16.0/archify`, overlays `skills/planning/archify`, and excludes `test` and `package-lock.json`. Running `clio-coder skills install archify --project` clones that tag, drops the excluded members, copies the overlay over the clone so Clio's wrapper `SKILL.md` replaces the upstream one, validates the shaped tree, and swaps it into `.clio-coder/skills/archify/` with the usual provenance stamps. The renderer, its schemas, and its brand-mark notices come from upstream at install time and never enter the npm tarball. The wrapper omits upstream's update-awareness step on purpose: that step performs a network request during a chat turn, and no Clio chat turn depends on the network.
78
+
79
+ Discovery treats the overlay folder as part of the remote entry rather than as a skill of its own, so a bare-name install never lands the wrapper without the renderer it wraps. `clio-coder skills update` recovers the same overlay and exclude list from the marketplace, so an update refetches the pinned upstream and re-applies the wrapper instead of replacing it.
80
+
81
+ Two operational notes. A copy dropped into a project-scope `.claude/skills/archify` is a compat import and stays untrusted until `integrations.projectResources.trustProjectImports` is on; install through Clio to get a trusted, provenance-stamped copy. And since the skill's authoring loop is several `bash` calls (validate, deliver, verify), `auto-edit` is the sensible permission posture for a mapping session; full manual approval works but prompts on every command.
82
+
65
83
  ## Publishing a skill
66
84
 
67
85
  Add a directory under `skills/<category>/<name>/` (or `skills/<name>/`) in the repo containing a `SKILL.md` with `name` and `description` frontmatter. The directory name must match `[A-Za-z0-9][A-Za-z0-9._-]*`. Run `npm run skills:pin` to republish `skills/skill-marketplace.json`, which is the index consumers point `CLIO_CODER_SKILL_MARKETPLACE_INDEX` at or copy to `<configDir>/skill-marketplace.json`. Scientific and niche coding domains are the marketplace's focus; see the existing `skills/` tree for the house format.
@@ -311,7 +311,7 @@ Arguments:
311
311
  Projects may commit a versioned executable catalog at `.clio-coder/verifiers.yaml`:
312
312
 
313
313
  ```yaml
314
- version: 1
314
+ version: 2
315
315
  checks:
316
316
  - id: rust-workspace
317
317
  description: Run the Rust workspace tests
@@ -319,9 +319,29 @@ checks:
319
319
  cwd: .
320
320
  timeoutMs: 600000
321
321
  tags: [rust, test]
322
+ - id: grid-metadata
323
+ description: Compare the regional grid statistics against the reference
324
+ kind: numeric-compare
325
+ command: [python, tools/grid_stats.py, out/region_west.nc]
326
+ reference: tests/reference/region_west.json
327
+ tolerance: { relative: 1.0e-6, ulp: 4 }
328
+ cwd: .
329
+ timeoutMs: 120000
330
+ tags: [scientific, netcdf]
331
+ - id: solver-time
332
+ description: Keep the solver inside its wall-time budget
333
+ kind: perf-budget
334
+ command: [python, tools/solve.py, --small]
335
+ baseline: .clio-coder/baselines/solver-time.json
336
+ tolerance: { relative: 0.25 }
337
+ cwd: .
338
+ timeoutMs: 600000
339
+ tags: [scientific, performance]
322
340
  ```
323
341
 
324
- Version 1 is strict. Every root and check field shown above is required, unknown fields fail, and duplicate IDs fail. A project ID uses lowercase letters, digits, `.`, `_`, `:`, or `-`, begins with a letter or digit, and is at most 64 UTF-8 bytes. `frontend` is reserved. Descriptions are trimmed single-line text capped at 512 bytes. `command` is a nonempty argv array with at most 64 entries and 4096 bytes per entry. A shell command string is invalid, and explicit shell executables such as `sh`, `bash`, `pwsh`, and `cmd` are rejected. `cwd` is a repository-relative existing directory capped at 512 bytes; absolute paths, `..` escapes, and symbolic-link escapes fail. `timeoutMs` is a positive integer capped at 900000. A check may carry at most 16 distinct lowercase tags of at most 32 bytes each. The whole file is capped at 262144 bytes and may contain at most 128 checks. YAML aliases are disabled.
342
+ Every check has a `kind`, absent or `command` by default. A version 1 file still loads and every check there is `kind: command`; the kind fields require `version: 2`. `kind: command` reads the exit code. `kind: numeric-compare` runs the command, parses its stdout as a JSON object of `string -> number | number[]`, and judges it against `reference` (a repository-relative JSON file of the same shape) under `tolerance`, which names at least one of `relative`, `absolute`, or `ulp`; a value passes only when every named tolerance holds, a key missing on either side fails with the key named, arrays compare elementwise and fail on length mismatch, and `NaN` or infinity fails. `kind: perf-budget` runs the command and judges the wall time the harness measured against either `budget: {wallTimeMs, tolerance?: {relative}}` or `baseline`, a repository-relative JSON `{wallTimeMs}` that `clio-coder verifiers baseline <id>` records from one clean run, with an optional `tolerance: {relative}` of headroom over it. Exactly one of `budget` and `baseline` is present. A command that exits non-zero, times out, or is aborted fails before any judgement. Both kinds record a structured `report` on the `verify` result details and on the host-verification check of a dispatch receipt (per-key worst deviation and the failed tolerance, or measured time, effective budget, and ratio); a failing judgement is a check failure, not a new evidence category.
343
+
344
+ Version 2 keeps version 1's strictness. Every root and check field shown above is required, unknown fields fail, and duplicate IDs fail. A project ID uses lowercase letters, digits, `.`, `_`, `:`, or `-`, begins with a letter or digit, and is at most 64 UTF-8 bytes. `frontend` is reserved. Descriptions are trimmed single-line text capped at 512 bytes. `command` is a nonempty argv array with at most 64 entries and 4096 bytes per entry. A shell command string is invalid, and explicit shell executables such as `sh`, `bash`, `pwsh`, and `cmd` are rejected. `cwd` is a repository-relative existing directory capped at 512 bytes; absolute paths, `..` escapes, and symbolic-link escapes fail. `timeoutMs` is a positive integer capped at 900000. A check may carry at most 16 distinct lowercase tags of at most 32 bytes each. The whole file is capped at 262144 bytes and may contain at most 128 checks. YAML aliases are disabled.
325
345
 
326
346
  Provider IDs share one namespace. If a catalog ID collides with a discovered package script, listing and execution fail and identify both source files. Catalog parsing also fails closed before any package or project check runs.
327
347
 
@@ -348,8 +368,11 @@ clio-coder verifiers author
348
368
  clio-coder verifiers author --exclude cmake-build-debug --rename go-test=go-suite
349
369
  clio-coder verifiers author --dry-run go-suite --yes
350
370
  clio-coder verifiers validate
371
+ clio-coder verifiers baseline solver-time
351
372
  ```
352
373
 
374
+ `author` also lists one incomplete `numeric-compare` check for every validation-contract artifact that declares `numerical_tolerances`, with the reference and tolerance filled in and the exact `verifiers add` line to complete; the command is the operator's to supply, so the proposal never enters the catalog on its own. `baseline <id>` runs a `perf-budget` check once and writes its wall time to the check's `baseline` path; a failing or timed-out command records nothing.
375
+
353
376
  `validate` reads the committed file with the same parser used by `verify()`. `dry-run <id>` is an explicit request to execute one admitted check through the production `verify` path. `author --dry-run <id> --yes` writes only after confirmation and starts the selected dry run only after the write is accepted by production discovery.
354
377
 
355
378
  Later changes use the same preview and confirmation boundary. `edit` preserves the ID unless `rename` is requested. Renames and additions reject collisions with catalog IDs and active package-script IDs. Removals state that the deleted command will no longer be executable through catalog authority. Generated IDs are stable for a stable ordered signal set; a collision receives the first available deterministic `-2`, `-3`, and later suffix.
@@ -362,7 +385,7 @@ clio-coder verifiers rename validate-grid validate-regional-grid --yes
362
385
  clio-coder verifiers remove validate-regional-grid --yes
363
386
  ```
364
387
 
365
- The `add` command is the explicit path for an unsupported or ambiguous project. `--command` must be a JSON argv array, so manual entry still cannot turn a shell command string into executable catalog authority.
388
+ The `add` command is the explicit path for an unsupported or ambiguous project. `--command` must be a JSON argv array, so manual entry still cannot turn a shell command string into executable catalog authority. `--kind numeric-compare` takes `--reference <path>` and `--tolerance '<json>'`; `--kind perf-budget` takes either `--budget-ms <n>` with optional `--budget-relative <r>` or `--baseline <path>` with optional `--tolerance '{"relative": r}'`. `edit` accepts the same options to change a check's kind.
366
389
 
367
390
  `verify(check="frontend", path=<file>)` validates an HTML, CSS, or JavaScript artifact without shell access. The path must stay inside the workspace root and end in `.html`, `.htm`, `.css`, `.js`, `.mjs`, or `.cjs`. Checks per type: HTML tag balance (comment-aware, HTML5 optional end tags honored), inline and referenced script syntax (classic scripts parsed in-process, modules via `node --check`), inline and linked CSS brace/string/comment balance, local script and stylesheet references resolved and existence-checked (external and root-relative references are skipped), and an optional headless browser load. `browser="auto"` warns when no chromium/chrome/edge executable is on PATH, `"required"` fails, `"off"` skips. Each check reports pass, warn, fail, or skip; any fail makes the whole result an error. `details = {action: "verify", check: "frontend", path, browserMode, status, checks}`.
368
391
 
@@ -627,6 +650,58 @@ panes(action="open", preset="logs")
627
650
  panes(action="close", target="all")
628
651
  ```
629
652
 
653
+ ## evidence: inspect canonical evidence and trust status
654
+
655
+ Reads evidence bundles as JSON. Source: `src/tools/evidence.ts`. Read class; sequential, because `run` mode may materialize a bundle under Clio's data directory. It shares the inventory and trust projections behind `clio-coder evidence inventory` and `clio-coder evidence inspect`, so the model and the operator read the same record.
656
+
657
+ Arguments:
658
+
659
+ - `mode` (required). `list`, `inspect`, or `run`.
660
+ - `id` (required for `inspect`). An evidence bundle id.
661
+ - `runId` (required for `run`). A dispatch run id; the bundle is built first when none exists.
662
+
663
+ `list` returns the bounded newest-first inventory: provenance, tags, totals, and a worst-run trust verdict per bundle. `inspect` returns the bundle overview, the per-run trust axes and verdict, the gate decisions, and the findings. `run` resolves `run-<runId>` and builds the bundle when it is absent; a run with no ledger row is reported absent with `artifactAbsent: true` in the details. Results are capped at 16KB, and a truncated result stays valid JSON with a `preview`. Provenance requires this tool and Verifier may use it.
664
+
665
+ ```text
666
+ evidence(mode="list")
667
+ evidence(mode="inspect", id="run-r-42")
668
+ evidence(mode="run", runId="r-42")
669
+ ```
670
+
671
+ ## limitation: record what a turn could not verify
672
+
673
+ Records a typed limitation receipt for the finish contract. Source: `src/tools/limitation.ts`. Read class; parallel. The tool is pure: it touches no filesystem and runs no shell, so the successful receipt in the session ledger is its whole effect.
674
+
675
+ Arguments:
676
+
677
+ - `scope` (required). What could not be verified, in one sentence.
678
+ - `reason` (required). `no-runner`, `blocked`, `out-of-scope`, `environment`, or `other`.
679
+ - `paths` (optional). Repository-relative paths left unverified.
680
+
681
+ Call it once, before the final reply, when files changed and validation could not run. The finish contract accepts a successful `limitation` receipt inside the same window as the mutation scan in place of validation evidence. A rejected call (empty scope, unknown reason) leaves no receipt and does not count, and the assistant's prose never does. The six mutating recipes carry the tool and the operating contract tells the model to call it; see [the finish gate](../architecture/safety-model.md#the-finish-gate-and-re-prompt-behavior).
682
+
683
+ ```text
684
+ limitation(scope="CUDA kernels changed but no GPU is available here", reason="environment", paths=["src/kernels/solve.cu"])
685
+ ```
686
+
687
+ ## decide: record a design decision
688
+
689
+ Appends the model's own design choice to the session decision board beside operator `ask_user` answers. Source: `src/tools/decide.ts`. Read class; sequential, so two decisions in one batch cannot race the supersede lookup. The call succeeds only in a session with a decision board; a worker's call is refused.
690
+
691
+ Arguments:
692
+
693
+ - `key` (required). Stable kebab-case name, at most 64 bytes.
694
+ - `value` (required). The option chosen, at most 512 bytes.
695
+ - `alternatives` (required). One to six rejected options, at most 256 bytes each.
696
+ - `rationale` (required). Why the choice won, at most 1024 bytes.
697
+ - `label` (optional). Short title, at most 128 bytes.
698
+
699
+ The call appends one `decisionLedger` entry with `origin: "agent"` and returns the decision ref `<interviewId>/<key>`. A repeat key supersedes the earlier agent decision with the new rationale as its correction; an operator decision with the same key is never overwritten and the call fails. Dispatch seals every active ref onto the run request, envelope, and receipt, and Clio-controlled commits carry one `Clio-Decision:` trailer per ref; see [commit provenance](../process/git-commit-provenance.md).
700
+
701
+ ```text
702
+ decide(key="cache-key-shape", value="capability tuple", alternatives=["node id"], rationale="matches the existing buckets and survives fleet changes", label="Cache key")
703
+ ```
704
+
630
705
  ## ask_user: host-owned operator interviews
631
706
 
632
707
  Runs a host-owned interactive interview or single-question prompt with the operator, recording decisions and/or free-form answers. Source: `src/tools/ask-user.ts`. Read class; sequential.
@@ -34,7 +34,7 @@ Every runtime-tunable value needs both halves; the pair is one knob, not two. Th
34
34
  | `CLIO_CODER_MAX_TOOL_CALLS` | 50 | `src/engine/loop-guard.ts` → `src/engine/worker-runtime.ts` | Worker lifetime tool-call cap for a dispatched run. Different axis than the orchestrator budget despite the near-identical name. |
35
35
  | `CLIO_CODER_MAX_DISPATCH_RUNS` | 1000 | `src/domains/dispatch/state.ts` | Dispatch run-ledger retention cap. |
36
36
  | `CLIO_CODER_MAX_CONTEXT_TOKENS` | unset | `src/domains/providers/runtime-resolution.ts` | Context-window override for local runtimes. Also set internally by `clio-coder run --max-context-tokens` (see §6). |
37
- | `CLIO_CODER_KV_CACHE_MODE` | unset | retired | KV-cache quantization mode. Also set internally by the former `clio-coder run --kv-cache-mode` path. |
37
+ | `CLIO_CODER_KV_CACHE_MODE` | unset | retired | KV-cache quantization mode. Also set internally by the former `clio-coder run` path. |
38
38
  | `CLIO_CODER_SAMPLING_OVERRIDES` | unset | `src/engine/apis/sampling-overrides.ts` | JSON sampling-parameter override. Set internally by print-mode sampling flags. |
39
39
  | `CLIO_CODER_READ_MAX_BYTES` | 51200 (50 KB) | `src/tools/read.ts` | Per-call byte cap for the read tool. |
40
40
  | `CLIO_CODER_OBSERVATION_TURN_BUDGET_BYTES` | 196608 (192 KB) | `src/tools/observation.ts` | Shared per-turn byte pool across all observation tools. |
@@ -93,7 +93,7 @@ All default off; all enabled with `1`.
93
93
 
94
94
  ## 6. CLI flags that bridge through env vars (Pre-consolidated State)
95
95
 
96
- `clio-coder run --max-context-tokens` and `--kv-cache-mode` (`src/cli/run.ts:143-234`) and the print-mode sampling flags (`src/cli/modes/print.ts:306-333`) do not plumb their values through function arguments. They mutate `process.env` (`CLIO_CODER_MAX_CONTEXT_TOKENS`, `CLIO_CODER_KV_CACHE_MODE`, `CLIO_CODER_SAMPLING_OVERRIDES`), run the command, then restore the previous value in a `finally`. The env var is the transport between the CLI layer and deep engine code.
96
+ `clio-coder run --max-context-tokens` (`src/cli/run.ts:143-234`) and the print-mode sampling flags (`src/cli/modes/print.ts:306-333`) do not plumb their values through function arguments. They mutate `process.env` (`CLIO_CODER_MAX_CONTEXT_TOKENS`, `CLIO_CODER_KV_CACHE_MODE`, `CLIO_CODER_SAMPLING_OVERRIDES`), run the command, then restore the previous value in a `finally`. The env var is the transport between the CLI layer and deep engine code.
97
97
 
98
98
  ## 7. Script- and benchmark-only vars (Pre-consolidated State)
99
99
 
@@ -20,12 +20,13 @@ under pressure; git mechanics alone never justify a stage.
20
20
  | --- | --- | --- |
21
21
  | 1. File | [`file-ticket`](../../skills/git/file-ticket/) | A labeled GitHub issue with evidence and acceptance criteria |
22
22
  | 2. Fix | [`fix-issue`](../../skills/git/fix-issue/) | An uncommitted, verified change where failing tests preceded the fix, self-reviewed against the issue's acceptance criteria |
23
- | 3. Ship | [`ship`](../../skills/git/ship/) | An atomic conventional commit referencing the issue (`fixes #N`), a gated push, and an open PR; merge is a human decision |
23
+ | 3. Ship | [`ship`](../../skills/git/ship/) | An atomic conventional commit referencing the issue (`fixes #N`); contributors push it to their fork and open a PR, while maintainer work stays local for gated integration; merge is a human decision |
24
24
 
25
- Releases follow [release-cut-checklist.md](../history/release-cut-checklist.md) as a
25
+ Releases follow [release-cut-checklist.md](release-cut-checklist.md) as a
26
26
  human-gated checklist, not a skill. Worktrees
27
27
  ([`worktree-create`](../../skills/git/worktree-create/),
28
28
  [`worktree-merge`](../../skills/git/worktree-merge/)),
29
+ [`branch-closeout`](../../skills/git/branch-closeout/),
29
30
  [`resolve-merge-conflicts`](../../skills/git/resolve-merge-conflicts/), and
30
31
  [`tdd`](../../skills/coding/tdd/) are à-la-carte tools reached for when the
31
32
  situation calls for them, not stages every change passes through. An RCA
@@ -34,6 +35,38 @@ hard bugs, not a mandatory toll booth. Batch ticket creation from a PRD
34
35
  bypasses stage 1 and uses [`backlog`](../../skills/planning/backlog/)
35
36
  instead; everything downstream is identical.
36
37
 
38
+ ## Closeout
39
+
40
+ A merged PR is not operationally finished until its local scaffolding is
41
+ closed. The reusable [`branch-closeout`](../../skills/git/branch-closeout/) skill automates this verification and teardown safely. After the human merge decision:
42
+
43
+ 1. Fetch and prune, confirm the PR's merged state, and identify the resulting
44
+ commit on `origin/main`. Direct ancestry proves an ordinary merge; a squash
45
+ or cherry-pick needs the PR-to-result evidence because commit identity and
46
+ patch identity can both change during integration.
47
+ 2. Inspect every associated worktree for tracked changes, untracked files, and
48
+ ignored state that carries evidence rather than rebuildable output. Remove
49
+ it through `git worktree remove`; forcing removal requires explicit approval
50
+ to discard what remains.
51
+ 3. Delete the local source and integration branches. For a contributor PR,
52
+ delete the merged branch from the contributor's fork. The canonical
53
+ repository never hosts topic, integration, or release-candidate branches.
54
+ 4. Turn unfinished experimental findings into an issue with evidence and a
55
+ next decision. Do not use indefinite `work/`, `wip/`, `keep/`, or temporary
56
+ tags as a substitute for backlog state.
57
+ 5. Report the remaining worktrees, local branches, stashes, local-only tags,
58
+ and canonical remote heads. The expected canonical head set is exactly
59
+ `refs/heads/main`; every survivor needs an owner and purpose.
60
+
61
+ Maintainer release candidates are local-only and use a compact branch name
62
+ that cannot collide with their tag: branch `v043`, tag `v0.4.3`. Gate the exact
63
+ candidate, require fetched `origin/main` to be its ancestor, fast-forward local
64
+ `main`, fetch again, and push only `refs/heads/main:refs/heads/main` with
65
+ explicit authorization. After CI passes, push only the fully qualified
66
+ annotated tag. Once the tag's peeled commit equals the reviewed commit on
67
+ `main` and the release succeeds, delete the local candidate branch. Published
68
+ dotted release tags are immutable history and are never cleanup targets.
69
+
37
70
  ## Inheriting a Pi release
38
71
 
39
72
  Pi dependency upgrades use a fixed five-step review so that upstream fixes
@@ -73,6 +106,11 @@ There is no committed weighted-shard or special serial-lane runner. Keep timing
73
106
  claims within the focused contract that owns them, and use the full `npm run ci`
74
107
  gate before handoff.
75
108
 
109
+ The release smoke script `scripts/smoke-real-home.sh` (invoked via
110
+ `npm run smoke:real-home`) tests booting the built CLI binary against a copy of
111
+ the operator settings in a scratch home. An optional `--strict` flag makes doctor
112
+ exit 1 fail the smoke run on failing rows instead of tolerating fleet state.
113
+
76
114
  ## Issue conventions
77
115
 
78
116
  - **Title**: conventional tag plus imperative summary (`fix: memory overlay
@@ -139,7 +139,37 @@ Metrics collected during runs can be validated automatically using the `verify.a
139
139
  * `eq` (equal)
140
140
  * `neq` (not equal)
141
141
 
142
- Metrics that can be validated include `tokens.input`, `tokens.output`, `tokens.total`, `latency.wallMs`, `tools.totalCalls`, `tools.failed`, `tools.blocked`, `verifier.exitCode`, and `result.pass`.
142
+ Metrics that can be validated include `tokens.input`, `tokens.output`, `tokens.total`, `latency.wallMs`, `tools.totalCalls`, `tools.failed`, `tools.blocked`, `verifier.exitCode`, `result.pass`, and the `provider.*` metrics below. Each `verify.assertions` condition must hold; an unavailable metric fails closed.
143
+
144
+ ### Optional Provider-Health Gates
145
+
146
+ A task that recovers from a provider error can still pass its task checks by default. Provider health is a separate, opt-in requirement. To require observed provider events and no observed terminal errors, add these assertions to the task:
147
+
148
+ ```yaml
149
+ verify:
150
+ assertions:
151
+ - metric: "provider.measured"
152
+ op: "eq"
153
+ value: true
154
+ - metric: "provider.stopReason.error"
155
+ op: "eq"
156
+ value: 0
157
+ ```
158
+
159
+ Suite-level `thresholds.fail` uses the opposite condition: a matching condition is a failure. The corresponding hard gate checks each run as follows:
160
+
161
+ ```yaml
162
+ thresholds:
163
+ fail:
164
+ - metric: "provider.measured"
165
+ op: "eq"
166
+ value: false
167
+ - metric: "provider.stopReason.error"
168
+ op: "gt"
169
+ value: 0
170
+ ```
171
+
172
+ Unavailable metrics also fail closed in a hard threshold. `thresholds.informational` records findings without changing exit status. These examples reject an observed error followed by a successful recovery while leaving the default ungated task-pass policy unchanged. To reject any observed retry start or assistant abort as well, add conditions on `provider.retryStarted` or `provider.stopReason.aborted` with the same assertion-versus-failure polarity.
143
173
 
144
174
  ---
145
175
 
@@ -175,14 +205,44 @@ Under the hood, these are parsed and wrapped into a Suite v2 adapter with:
175
205
 
176
206
  Clio maintains two distinct token accounting streams with different provenances. These accounts are never merged, reconciled, or treated as interchangeable:
177
207
 
178
- 1. **`tokens.*` (Wire Streaming)**: Folded live off stdout from assistant `message_end` events watched by `token-stream.ts` / `createStreamInvariantFold`. This represents usage reported by the provider for assistant messages watched over the wire. On surfaces without stdout streaming (such as `clio-coder fleet run --json`), `tokens.measured` is `false`.
208
+ 1. **`tokens.*` (Wire Streaming)**: Folded live off stdout from assistant `message_end` events. These totals include all known usage on errored calls as well as successful calls; recovery never subtracts earlier spend. Only finite, nonnegative usage facts are admitted. On surfaces without the relevant stdout events (such as `clio-coder fleet run --json`), `tokens.measured` is `false`.
179
209
  2. **`receiptUsage.*` (Journal Receipts)**: Summed from an evaluation item's run journal. Every attempt writes a receipt carrying token counts and USD cost authenticated against its own ledger envelope.
180
210
 
181
211
  ### Fail-Closed Reporting
182
212
  Both accounting streams report unmeasured state with no counts at all rather than a numeric zero. Reporting zero for an unmeasured run would falsely claim the run cost nothing. On an unmeasured run, `tokens.total` resolves to `null` and fails closed on metric threshold comparisons.
183
213
 
214
+ An errored call must carry at least one positive reported token, reasoning, or cost fact before its usage is considered observed. Reasoning-only or cost-only observations do not establish ordinary token totals. A stream containing only errored calls with missing or synthetic all-zero usage remains `tokens.measured: false`. The current event shape cannot distinguish synthetic all-zero failures from genuinely reported zero usage, so it cannot establish measured zero spending in either case. Partial positive usage remains included as known spend. On failed calls, adapters can also initialize individual absent fields to zero; those zeros remain unattributed and make coverage incomplete. The existing inclusive numeric fields are known subtotals, so their zeros do not prove complete zero spending when failed usage is incomplete.
215
+
216
+ ### Provider Observations and Failed-Call Share
217
+
218
+ The `provider.*` metrics describe events observed on live stdout, folded before diagnostic output is truncated. Native runs and multi-command external runners retain these observations from their executed commands. They do not enumerate SDK-internal retries or network attempts that were never emitted. Filtered or opaque output can leave provider health unobserved even when the process exits successfully or a receipt reports task success.
219
+
220
+ | Metric | Meaning |
221
+ | --- | --- |
222
+ | `provider.measured` | Whether an assistant terminal reason or a counted retry phase was observed. With no such observations this is `false`, and provider counters are absent. It does not certify complete provider coverage. |
223
+ | `provider.stopReason.stop`, `provider.stopReason.toolUse`, `provider.stopReason.length` | Counts of these terminal reasons on assistant `message_end` events. |
224
+ | `provider.stopReason.error`, `provider.stopReason.aborted`, `provider.stopReason.other` | Separate counts of errored, aborted, and other observed terminal reasons. An unrecognized terminal reason goes into `other`. Partial updates and repeated messages in `turn_end` or `agent_end` do not add counts. |
225
+ | `provider.retryScheduled` | Observed `scheduled` phases: planned retries, including ones cancelled before execution. |
226
+ | `provider.retryStarted` | Observed `retrying` phases: retry execution starts. Repeated `waiting` countdown frames do not count as attempts. |
227
+ | `provider.retryCancelled`, `provider.retryExhausted`, `provider.retryRecovered` | Counts of the corresponding observed phases. They describe retry-chain outcomes and do not fabricate additional assistant calls. Attempt numbers can restart for each chain. |
228
+ | `provider.errorUsageObservedCalls` | Errored calls with at least one positive reported token, reasoning, or cost fact. |
229
+ | `provider.errorUsageUnobservedCalls` | Errored calls with no positive reported token, reasoning, or cost fact, including absent or all-zero usage. |
230
+ | `provider.errorUsageIncompleteCalls` | Errored calls with unobserved usage, incomplete token fields, or ambiguous normalized zero fields. This can overlap `errorUsageObservedCalls` when only part of the usage is known. |
231
+ | `provider.errorCostUnobservedCalls` | Errored calls without a positive cost fact. Zero or absent cost does not prove that an error was free, including when some token usage is known. |
232
+ | `provider.errorTokens.input`, `provider.errorTokens.output`, `provider.errorTokens.total`, `provider.errorTokens.cacheRead`, `provider.errorTokens.cacheWrite` | Known positive token subtotals for errored calls. Absent or ambiguous zero fields remain unattributed; when total usage is absent, a total can still be summed from known token fields and remains incomplete. |
233
+ | `provider.errorCostUsd` | Known positive cost subtotal from errored calls' stream usage objects. Cost can come from adapter pricing; it is not independently certified provider billing. |
234
+ | `provider.errorReasoningTokens`, `provider.errorReasoningUnobservedCalls` | Known failed-call reasoning subtotal and calls without attributable reported reasoning. Reasoning can overlap output, so it is never added to ordinary token totals. |
235
+
236
+ The failed-call share covers `stopReason: error`; aborted calls remain separately labeled. Failed-share amounts and usage-coverage counters appear only after an errored terminal message is observed. These share metrics supplement the inclusive `tokens.*` totals without changing the summary token shape, receipt schema, or verdict schema. Missing or partial failed usage makes them known subtotals, not a complete amount to subtract from total spend. The native runner's `cost.usd` can use receipt evidence, so equality with the stream-based `provider.errorCostUsd` is not guaranteed.
237
+
238
+ Positive reported reasoning uses the normalized `usage.reasoning` field, `reasoning_tokens`, and supported nested provider-detail fields. The root `reasoningTokens` alias can be an adapter estimate without a provenance marker, so it is left unattributed rather than promoted to reported provider usage. Normalized zero reasoning is also unattributed because adapters can fill it when provider detail is absent. A reasoning-only failure is observed but has incomplete ordinary-token coverage; no output or total is inferred from it.
239
+
240
+ These observations also do not reconcile the separate `trackedMetrics` ledger selection (#276). Tracked metrics prefer durable assistant-call facts when available, retain durable compaction and tool records, and otherwise fall back to stream calls. Artifacts expose source counts and warnings, but partial or mixed ledgers can omit stream-only calls and fork-inherited history remains unreconciled. Neither those tracked values nor the new provider counters prove complete run accounting; failed-compaction usage is retained separately in the out-of-turn usage ledger and usage report, and is not included by this eval fold.
241
+
184
242
  ---
185
243
 
244
+ Full reconciliation across session, stdout, fork and out-of-turn evidence is deferred to v0.4.5 or later. Version 0.4.3 does not add an automatic rejection of cost or efficiency comparisons merely because those sources are partial or mixed. Matching source counts do not prove complete coverage or shared call identity. Existing missing-metric, serving-configuration and execution-envelope comparison gates still apply.
245
+
186
246
  ## Eval Artifact Format (v4)
187
247
 
188
248
  Evaluation artifacts use format version 4 (`EvalArtifactV4`). Summary token metrics report `measuredRuns` out of total `runs`:
@@ -297,7 +357,7 @@ Eleven numbers plus a reason histogram, each carrying the source it came from. `
297
357
  | `generatedTokens` | ledger |
298
358
  | `reasoningTokens` | receipt; nullable, because absent and zero are different claims |
299
359
  | `toolCalls`, `toolErrors` | ledger when present, otherwise receipt |
300
- | `ttftMsFirstCall` | ledger |
360
+ | `ttftMsFirstCall` | ledger; nullable when first-call timing is absent |
301
361
  | `wallClockMs` | receipt |
302
362
  | `contextTokensAtEnd` | ledger |
303
363
  | `compactions` | ledger |
@@ -305,6 +365,10 @@ Eleven numbers plus a reason histogram, each carrying the source it came from. `
305
365
 
306
366
  A dispatched worker's receipt reports `sessionId: null` and writes no session archive, which is why the ledger source exists at all: the runner folds structured usage, backend timing, cache, and monotonic TTFT facts out of the worker's `message_end` events. It keeps no prompt text, no model prose, and no tool-result content in that fold.
307
367
 
368
+ First-call TTFT uses the earliest recorded assistant-call timestamp across the selected ledgers; equal timestamps retain their observed order. Missing or invalid timing, or an invalid timestamp that prevents ordering the calls, yields `{ value: null, source: "estimated" }`. A measured zero remains `{ value: 0, source: "ledger" }`. Native session timing starts at each stream invocation and includes the provider's response-header wait. Stdout-only fallback timing starts at the provider's `message_start`, which can arrive after headers; it requires first output, but is not complete request latency and must not be compared as equivalent to native invocation timing. A completion alone supplies neither a zero duration nor a first-token measurement. Verdict v1 consumers must accept nullable TTFT. Historical numeric values, including estimated zeros and native spans that omitted the pre-header wait, remain readable and are not rewritten.
369
+
370
+ When a stream message lacks a valid timestamp, its ledger payload marks `timestampEstimated: true` beside the legacy ISO placeholder. This leaves first-call chronology unmeasured while preserving any observed per-call monotonic timing.
371
+
308
372
  ### Scenario aggregates
309
373
 
310
374
  `aggregates` groups verdicts by `scenarioId`, sets `k` to the trial count, and records `passAtK` (any trial passed) and `passPowK` (every trial passed). Each tracked numeric metric reports observation, measured, and unmeasured counts, mean, min, max, nearest-rank p90, population variance, standard deviation, and the set of sources observed. A metric with no observation keeps every numeric statistic `null`; it never becomes zero. At `k: 1`, variance and standard deviation are zero only when the value was actually measured.
@@ -66,6 +66,21 @@ Clio-Evidence: receipt-v20/sha256:<64-character digest>
66
66
  Clio does not invent, shorten, or add an unrelated digest. The role trailers do
67
67
  not depend on this optional line.
68
68
 
69
+ A commit also names the decisions it was made under, one trailer per active
70
+ decision on the session decision board (an `ask_user` answer or a design
71
+ choice the agent recorded with `decide`):
72
+
73
+ ```text
74
+ Clio-Decision: <interviewId>/<key>
75
+ ```
76
+
77
+ The refs come from the sealed receipt's `decisionRefs` at the fleet seam and
78
+ from the live board at the session seam. They are sorted, capped at 32, added
79
+ once, and only refs of the `<id>/<key>` shape are written, where the key is an
80
+ operator `snake_case` key or an agent `kebab-case` key. A decision trailer
81
+ records rationale provenance; it is not evidence that the decision was correct
82
+ or that its work was validated.
83
+
69
84
  ## Commit paths and hooks
70
85
 
71
86
  The deterministic SDLC fleet attributes its plan, code, and documentation