@iowarp/clio-coder 0.4.2 → 0.4.3

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (369) hide show
  1. package/CHANGELOG.md +35 -0
  2. package/CONTRIBUTING.md +86 -19
  3. package/README.md +35 -6
  4. package/dist/{acp-TMDQZDIG.js → acp-H2NGRPWO.js} +11 -11
  5. package/dist/{agents-5N5NG3XG.js → agents-TL5LLUQP.js} +54 -53
  6. package/dist/assets/codewiki.json +1 -1
  7. package/dist/{auth-Z5CCBXKQ.js → auth-E5SW4HMS.js} +19 -16
  8. package/dist/{builtins-K6TNDT24.js → builtins-IA7V7FUC.js} +9 -4
  9. package/dist/{chunk-ZW4HH5JJ.js → chunk-2APPQIER.js} +6 -6
  10. package/dist/{chunk-72GZI5EV.js → chunk-2JH2WHGE.js} +2 -2
  11. package/dist/{chunk-UBRFI4HS.js → chunk-2UG5F4C5.js} +127 -47
  12. package/dist/{chunk-UH632ZYL.js → chunk-2UH2KFUP.js} +2 -2
  13. package/dist/{chunk-3F7VUY77.js → chunk-2VIKGWFZ.js} +2 -2
  14. package/dist/{chunk-I66EAJFY.js → chunk-2WZ546HR.js} +267 -232
  15. package/dist/{chunk-RKSR6VSF.js → chunk-4IUZQIJ3.js} +29 -1
  16. package/dist/{chunk-2X4RYJTJ.js → chunk-4UVU7BJ5.js} +2 -2
  17. package/dist/{chunk-M2DAX4F6.js → chunk-4WR7VSYB.js} +2 -2
  18. package/dist/{chunk-5KW52TEP.js → chunk-54CBCGIR.js} +5 -5
  19. package/dist/{chunk-3KIPBMUA.js → chunk-5ICU3EUH.js} +2 -2
  20. package/dist/{chunk-77QIVUZB.js → chunk-5MEZN6CB.js} +4 -4
  21. package/dist/{chunk-O42A54GG.js → chunk-5OIVVPHF.js} +2 -2
  22. package/dist/{chunk-LJID3DYZ.js → chunk-64I3JVYM.js} +2 -2
  23. package/dist/{chunk-4JDLP6ZS.js → chunk-6PTFB5VS.js} +7 -7
  24. package/dist/{chunk-VN3SHNBN.js → chunk-7DICMOS6.js} +2 -2
  25. package/dist/{chunk-HLAFFSEK.js → chunk-7DRAWPTZ.js} +2 -2
  26. package/dist/chunk-7E7I3WLS.js +3762 -0
  27. package/dist/{chunk-YJISEZKC.js → chunk-7ZYNNDKC.js} +6 -6
  28. package/dist/{chunk-I66ZTYNP.js → chunk-AF4YM7Z4.js} +236 -101
  29. package/dist/{chunk-2HFQNRV3.js → chunk-AX2THNSA.js} +12 -12
  30. package/dist/{chunk-PGF63K6I.js → chunk-B4OAX3SI.js} +65 -3
  31. package/dist/{chunk-W6NIE6OW.js → chunk-B4VEBZKF.js} +3 -3
  32. package/dist/{chunk-JBCS7CRR.js → chunk-BEPZRGGU.js} +10 -10
  33. package/dist/{chunk-XGDPUNND.js → chunk-CE5AX47J.js} +2 -2
  34. package/dist/{chunk-I64IFBLB.js → chunk-DWUOQKRU.js} +17 -10
  35. package/dist/{chunk-DZAW46HP.js → chunk-E3TPLWFX.js} +3 -3
  36. package/dist/{chunk-HIICAHCJ.js → chunk-EKCHAPYA.js} +2 -2
  37. package/dist/{chunk-XE3PCIXH.js → chunk-F5JHEYZM.js} +7 -7
  38. package/dist/{chunk-5PFYMY2V.js → chunk-FTMGRKEF.js} +2 -2
  39. package/dist/{chunk-ZNT2M6TG.js → chunk-G76U63X4.js} +17 -17
  40. package/dist/{chunk-34BHNEE3.js → chunk-GHS5EBTQ.js} +58 -7
  41. package/dist/{chunk-2NHR3NAY.js → chunk-GI7YYQ3F.js} +40 -34
  42. package/dist/{chunk-DYHAXKHD.js → chunk-GWZNEVM2.js} +12 -8
  43. package/dist/chunk-GYV6VZOC.js +26 -0
  44. package/dist/{chunk-MQXIVJ35.js → chunk-HAXOFFRH.js} +5 -5
  45. package/dist/{chunk-W4YEMFBX.js → chunk-HEQY7ZFI.js} +2 -2
  46. package/dist/{chunk-IKOZFYBN.js → chunk-I7ZPNEJM.js} +145 -102
  47. package/dist/{chunk-WNP7O5WZ.js → chunk-ID64D7PE.js} +4 -4
  48. package/dist/{chunk-XQRY4DTA.js → chunk-IGLP3ODT.js} +10 -10
  49. package/dist/chunk-IJNZMHLA.js +101 -0
  50. package/dist/{chunk-JWJGP5DQ.js → chunk-INY6HTFL.js} +7 -7
  51. package/dist/{chunk-PBP4B7XR.js → chunk-IUE3Y34X.js} +2 -2
  52. package/dist/{chunk-B74PXLU7.js → chunk-IWT4SF4R.js} +3 -3
  53. package/dist/{chunk-B7HM5Z7T.js → chunk-JDAY6FIL.js} +5 -5
  54. package/dist/{chunk-PJX3WQUQ.js → chunk-JEQ3XTHC.js} +2 -2
  55. package/dist/{chunk-FSP7CMNU.js → chunk-JGRC33J2.js} +50 -4
  56. package/dist/{chunk-X7IARSHT.js → chunk-JKKCYP3C.js} +9 -9
  57. package/dist/{chunk-HJWWJ6IL.js → chunk-JSC3U7TI.js} +16 -4
  58. package/dist/{chunk-SSEYRH53.js → chunk-KK4JZPBQ.js} +19 -140
  59. package/dist/{chunk-IHXBNWMM.js → chunk-KXDSS5WJ.js} +7 -3
  60. package/dist/{chunk-Q4XWMHX6.js → chunk-L47TF46W.js} +2 -2
  61. package/dist/{chunk-O3YUNJZ2.js → chunk-LDJG7DW3.js} +81 -24
  62. package/dist/{chunk-CDNVLKUX.js → chunk-LLDJM5XK.js} +13 -7
  63. package/dist/{chunk-F2I26BDK.js → chunk-MUW2BDDH.js} +4 -4
  64. package/dist/{chunk-HKMD33FO.js → chunk-MWUZBSAQ.js} +79 -76
  65. package/dist/{chunk-QQLGQY2A.js → chunk-N2Z7HLVY.js} +20 -20
  66. package/dist/{chunk-DZEK6CJN.js → chunk-NIQJ66N4.js} +19 -19
  67. package/dist/{chunk-CWVRRIEI.js → chunk-NZMNUPZZ.js} +2 -2
  68. package/dist/{chunk-462T4EGZ.js → chunk-O5CVSAG5.js} +2 -2
  69. package/dist/{chunk-TPEQIQIE.js → chunk-OML5D5V5.js} +8 -8
  70. package/dist/{chunk-IKSLQ4XV.js → chunk-PAJQJ7BS.js} +558 -216
  71. package/dist/{chunk-ZW55JB7N.js → chunk-PUVDKJ2Y.js} +2 -2
  72. package/dist/{chunk-UH347SHR.js → chunk-QWGDJJYJ.js} +11 -11
  73. package/dist/chunk-R6Q67RJH.js +134 -0
  74. package/dist/{chunk-CRFOIAX3.js → chunk-RRNP2ANY.js} +6 -6
  75. package/dist/{chunk-IDNA72AH.js → chunk-RSJ25QSL.js} +2 -2
  76. package/dist/chunk-SKHCAU7K.js +385 -0
  77. package/dist/{chunk-RLYRBIYQ.js → chunk-TM6LQDI3.js} +20 -12
  78. package/dist/chunk-UOIZ7DA4.js +41 -0
  79. package/dist/{chunk-P75RZCJW.js → chunk-UPZU6GE4.js} +3 -3
  80. package/dist/{chunk-MCMZMDAC.js → chunk-V2ANDPVT.js} +4 -4
  81. package/dist/{chunk-AK5XEFVZ.js → chunk-VA5FNYMT.js} +26 -13
  82. package/dist/{chunk-IMXMHHMQ.js → chunk-VW6DOEDG.js} +332 -57
  83. package/dist/{chunk-XOXV5GKE.js → chunk-W6RRQCPQ.js} +16 -7
  84. package/dist/{chunk-CYZW7JHJ.js → chunk-WBKFA554.js} +8 -8
  85. package/dist/{chunk-BO7Y52RY.js → chunk-WCXUNS7U.js} +7 -7
  86. package/dist/{chunk-ZGNYYXQ6.js → chunk-WRBAGUNF.js} +3 -3
  87. package/dist/{chunk-FVDGR2ZL.js → chunk-XIVNBFZS.js} +85 -30
  88. package/dist/{chunk-BYMNWQ7O.js → chunk-XPWWI35G.js} +299 -58
  89. package/dist/{chunk-KPXDY6QF.js → chunk-XRZT5WY5.js} +2 -2
  90. package/dist/{chunk-AZ4WMN4W.js → chunk-Y3CBHOR6.js} +2 -2
  91. package/dist/{chunk-VXMFAE2W.js → chunk-YPC6ZR5L.js} +19 -6
  92. package/dist/{chunk-54ODD65L.js → chunk-YQWYVTMC.js} +4 -4
  93. package/dist/{chunk-M2WXEHER.js → chunk-ZA4VCIGV.js} +2 -2
  94. package/dist/{chunk-7BHIY2MW.js → chunk-ZDN3Y73Y.js} +6 -6
  95. package/dist/{chunk-E7GT7O5N.js → chunk-ZWPRK62N.js} +7 -4
  96. package/dist/cli/index.js +38 -37
  97. package/dist/{clio-7VB377CC.js → clio-CMMK4KRR.js} +7 -7
  98. package/dist/{code-nav-YVLCYA7V.js → code-nav-MDZNQS33.js} +7 -7
  99. package/dist/{components-UBWCQSRW.js → components-UCUQ4QXW.js} +4 -4
  100. package/dist/{config-4HVOS65E.js → config-SVM5P5YI.js} +76 -74
  101. package/dist/{configure-PIWO7B24.js → configure-LE3IK2TJ.js} +26 -24
  102. package/dist/{context-IYEHL3WQ.js → context-2OHRKS42.js} +66 -63
  103. package/dist/{context-N6ZE3LGJ.js → context-E3VC7RX5.js} +15 -11
  104. package/dist/{context-KQYIWPWT.js → context-VNCR7KAG.js} +60 -45
  105. package/dist/{context-clear-G4OGZJDS.js → context-clear-BW4O37TG.js} +61 -59
  106. package/dist/context-map-COB37XXN.js +505 -0
  107. package/dist/{context-working-set-BWLF6LJP.js → context-working-set-VDS25HXZ.js} +17 -16
  108. package/dist/{dispatch-runner-2QQAITS3.js → dispatch-runner-5AHT53RF.js} +85 -74
  109. package/dist/{doctor-LHBD36VU.js → doctor-WNNVO6FY.js} +37 -37
  110. package/dist/{eval-C45FYRJ6.js → eval-7G7SGAYO.js} +285 -114
  111. package/dist/{eval-inventory-6DEJPLBF.js → eval-inventory-Y6QRFOH5.js} +4 -4
  112. package/dist/{evidence-6SHONYAF.js → evidence-VD6736FQ.js} +63 -62
  113. package/dist/{evolve-KRKMV72X.js → evolve-AL3NGVRL.js} +62 -61
  114. package/dist/{extensions-KPZ2UHBB.js → extensions-MOVJ32NM.js} +7 -7
  115. package/dist/{fleet-IVTCKDHT.js → fleet-QZHUMAGI.js} +110 -108
  116. package/dist/{fleet-commands-EDWL3IT7.js → fleet-commands-BAYT5FJZ.js} +10 -10
  117. package/dist/{fleet-decisions-YP3YEFGK.js → fleet-decisions-IREVMRU4.js} +7 -6
  118. package/dist/{fleet-graph-ZFWKHY2M.js → fleet-graph-YCTT3HTI.js} +19 -18
  119. package/dist/{fleet-inspect-FVUNCBML.js → fleet-inspect-QVJTDAVB.js} +55 -54
  120. package/dist/{fleet-preflight-UN5XED4R.js → fleet-preflight-25QAFPK4.js} +4 -4
  121. package/dist/{fleet-validate-XOWC4HSX.js → fleet-validate-5O57AAJ7.js} +23 -22
  122. package/dist/{fleet-verify-UN3SODEL.js → fleet-verify-CPH2W2T6.js} +56 -55
  123. package/dist/{fleet-view-TWHJKCN6.js → fleet-view-SWBR3VGQ.js} +55 -54
  124. package/dist/{init-T2QORQ3Y.js → init-J477LKZH.js} +78 -76
  125. package/dist/{interop-IN5I2A66.js → interop-3FCM6XLG.js} +11 -11
  126. package/dist/{library-LSCATDLZ.js → library-QUQEIUG6.js} +28 -27
  127. package/dist/{memory-HYOKAGGJ.js → memory-SGGSEP65.js} +64 -63
  128. package/dist/{models-2GPMFYCM.js → models-HEKUAXXK.js} +49 -43
  129. package/dist/{monitor-E4ASVUJH.js → monitor-HKU57TYQ.js} +61 -60
  130. package/dist/{orchestrator-DDMPR3PY.js → orchestrator-VDFAEFAI.js} +919 -546
  131. package/dist/{panes-E3RUXOW5.js → panes-DN2SSFOH.js} +3 -3
  132. package/dist/{panes-IXKLOKA2.js → panes-TALGNPZT.js} +8 -8
  133. package/dist/{paths-L7LGY6RN.js → paths-NBMFAIEZ.js} +5 -5
  134. package/dist/reset-EAJFFJVB.js +344 -0
  135. package/dist/{resources-OTRSN34L.js → resources-OVKSEFVE.js} +27 -20
  136. package/dist/{run-5DEYH5QK.js → run-7DP7ZF2J.js} +113 -109
  137. package/dist/{share-IHWTLO3M.js → share-WML67FT3.js} +26 -25
  138. package/dist/{skills-IYMXMKW4.js → skills-SG662R2K.js} +39 -31
  139. package/dist/{skills-eval-DROHSJAR.js → skills-eval-VVZEUU46.js} +74 -73
  140. package/dist/{skills-inventory-D7X4L4ZX.js → skills-inventory-I2E23GET.js} +21 -20
  141. package/dist/{slash-commands-QBM7UZ3B.js → slash-commands-S7MBJDQK.js} +35 -34
  142. package/dist/{steer-Z5DO23FJ.js → steer-2LQOMCPB.js} +3 -3
  143. package/dist/{support-U7QOWY26.js → support-CC2UJBJ6.js} +6 -6
  144. package/dist/{targets-P2FUC4IL.js → targets-4QC3HIEW.js} +48 -45
  145. package/dist/{terminal-lease-YREJ3JX2.js → terminal-lease-TUHIJ6Y2.js} +2 -2
  146. package/dist/{tools-5B7RO6MV.js → tools-TFGJICCU.js} +8 -8
  147. package/dist/{trace-YMGMUM6A.js → trace-FXMXUZUF.js} +7 -7
  148. package/dist/uninstall-5PEVOE5B.js +408 -0
  149. package/dist/upgrade-M4WXY6KN.js +303 -0
  150. package/dist/{usage-ME5MPXGX.js → usage-N7ZNVLEM.js} +147 -102
  151. package/dist/{verifiers-BVZ7IWOO.js → verifiers-DJTP4XX6.js} +15 -15
  152. package/dist/{verify-5K7ZKQFC.js → verify-RWE4PPEK.js} +9 -9
  153. package/dist/{wiki-generate-F5W5QTYY.js → wiki-generate-C7IQOXSP.js} +84 -82
  154. package/dist/{with-panes-BYOJCLAM.js → with-panes-4GCGSL7J.js} +9 -9
  155. package/dist/worker/entry.js +61 -60
  156. package/docs/architecture/artifact-placement.md +1 -0
  157. package/docs/architecture/artifact-versions.md +1 -1
  158. package/docs/architecture/context-engine.md +4 -0
  159. package/docs/architecture/middleware-and-components.md +1 -1
  160. package/docs/architecture/model-catalog.md +21 -10
  161. package/docs/architecture/observability.md +12 -1
  162. package/docs/architecture/prompt-envelope-and-tools.md +2 -0
  163. package/docs/architecture/provider-adapter-cookbook.md +63 -0
  164. package/docs/architecture/safety-model.md +15 -5
  165. package/docs/guide/built-in-agents.md +17 -3
  166. package/docs/guide/commands-and-modes.md +1 -1
  167. package/docs/guide/configuration-and-targets.md +97 -9
  168. package/docs/guide/configuration-reference.md +7 -2
  169. package/docs/guide/environment-variables.md +2 -0
  170. package/docs/guide/installation-and-lifecycle.md +37 -4
  171. package/docs/guide/proactive-memory.md +66 -55
  172. package/docs/guide/skills-marketplace.md +18 -0
  173. package/docs/process/development-pipeline.md +34 -1
  174. package/docs/process/eval-runner.md +67 -3
  175. package/evals/behavioral-model.yaml +3 -2
  176. package/package.json +2 -2
  177. package/skills/README.md +7 -5
  178. package/skills/coding/ast-grep/SKILL.md +101 -30
  179. package/skills/coding/ast-grep/evals.md +26 -0
  180. package/skills/coding/coding-standards/SKILL.md +40 -5
  181. package/skills/coding/coding-standards/evals.md +23 -0
  182. package/skills/coding/prototype/SKILL.md +87 -28
  183. package/skills/coding/prototype/evals.md +19 -0
  184. package/skills/coding/tdd/SKILL.md +80 -53
  185. package/skills/coding/tdd/evals.md +20 -0
  186. package/skills/context/context-handoff/SKILL.md +43 -2
  187. package/skills/context/context-handoff/evals.md +44 -0
  188. package/skills/context/context-prime/SKILL.md +45 -15
  189. package/skills/context/context-prime/evals.md +45 -0
  190. package/skills/git/branch-closeout/SKILL.md +132 -0
  191. package/skills/git/branch-closeout/evals.md +133 -0
  192. package/skills/git/branch-closeout/references/closeout-checklist.md +81 -0
  193. package/skills/git/file-ticket/SKILL.md +77 -63
  194. package/skills/git/file-ticket/assets/issue-template.md +22 -0
  195. package/skills/git/file-ticket/evals.md +31 -26
  196. package/skills/git/file-ticket/references/issue-discovery.md +49 -0
  197. package/skills/git/fix-issue/SKILL.md +87 -64
  198. package/skills/git/fix-issue/evals.md +35 -31
  199. package/skills/git/fix-issue/references/diagnosis-and-rca.md +46 -0
  200. package/skills/git/resolve-merge-conflicts/SKILL.md +100 -51
  201. package/skills/git/resolve-merge-conflicts/evals.md +52 -25
  202. package/skills/git/resolve-merge-conflicts/references/conflict-matrix.md +126 -0
  203. package/skills/git/ship/SKILL.md +103 -67
  204. package/skills/git/ship/assets/pr-template.md +21 -0
  205. package/skills/git/ship/evals.md +44 -28
  206. package/skills/git/ship/references/remote-and-branch-policy.md +62 -0
  207. package/skills/git/worktree-create/SKILL.md +80 -50
  208. package/skills/git/worktree-create/evals.md +40 -33
  209. package/skills/git/worktree-create/references/worktree-setup.md +62 -66
  210. package/skills/git/worktree-merge/SKILL.md +112 -65
  211. package/skills/git/worktree-merge/evals.md +42 -34
  212. package/skills/git/worktree-merge/references/merge-strategies.md +52 -0
  213. package/skills/planning/archify/SKILL.md +196 -0
  214. package/skills/planning/archify/evals.md +65 -0
  215. package/skills/planning/architecture/SKILL.md +61 -12
  216. package/skills/planning/architecture/evals.md +65 -0
  217. package/skills/planning/backlog/SKILL.md +130 -14
  218. package/skills/planning/backlog/evals.md +142 -0
  219. package/skills/planning/prd/SKILL.md +46 -6
  220. package/skills/planning/prd/evals.md +54 -0
  221. package/skills/planning/product-intent/SKILL.md +57 -2
  222. package/skills/planning/product-intent/evals.md +70 -0
  223. package/skills/planning/tech-spec/SKILL.md +53 -2
  224. package/skills/planning/tech-spec/evals.md +73 -0
  225. package/skills/registry.yaml +58 -50
  226. package/skills/remote.yaml +13 -0
  227. package/skills/research/arxiv-literature/SKILL.md +76 -18
  228. package/skills/research/arxiv-literature/evals.md +50 -0
  229. package/skills/research/experiment-protocol/SKILL.md +20 -1
  230. package/skills/research/experiment-protocol/evals.md +23 -0
  231. package/skills/research/scientific-debugging/SKILL.md +23 -1
  232. package/skills/research/scientific-debugging/evals.md +18 -0
  233. package/skills/research/scientific-modernization/SKILL.md +26 -1
  234. package/skills/research/scientific-modernization/evals.md +27 -0
  235. package/skills/skill-marketplace.json +63 -28
  236. package/skills/workflow/cut-it/SKILL.md +65 -5
  237. package/skills/workflow/cut-it/evals.md +101 -0
  238. package/skills/workflow/design-council/SKILL.md +117 -27
  239. package/skills/workflow/design-council/evals.md +161 -0
  240. package/skills/workflow/grill-me/SKILL.md +86 -10
  241. package/skills/workflow/grill-me/evals.md +153 -0
  242. package/skills/workflow/workflow-distiller/SKILL.md +76 -17
  243. package/skills/workflow/workflow-distiller/evals.md +118 -0
  244. package/src/cli/configure-interop.ts +105 -13
  245. package/src/cli/configure-oauth.ts +57 -0
  246. package/src/cli/configure-onboarding.ts +980 -0
  247. package/src/cli/configure-target.ts +594 -0
  248. package/src/cli/configure.ts +1082 -528
  249. package/src/cli/context-map.ts +114 -0
  250. package/src/cli/context.ts +4 -0
  251. package/src/cli/index.ts +1 -0
  252. package/src/cli/lifecycle-presenter.ts +436 -0
  253. package/src/cli/models.ts +10 -2
  254. package/src/cli/modes/print.ts +5 -1
  255. package/src/cli/reset.ts +228 -106
  256. package/src/cli/run.ts +7 -2
  257. package/src/cli/select.ts +664 -0
  258. package/src/cli/skills.ts +9 -2
  259. package/src/cli/targets.ts +3 -0
  260. package/src/cli/uninstall.ts +233 -165
  261. package/src/cli/upgrade.ts +204 -149
  262. package/src/cli/usage.ts +86 -27
  263. package/src/cli/validate-model.ts +3 -3
  264. package/src/core/config.ts +56 -0
  265. package/src/core/external-diagnostic.ts +44 -0
  266. package/src/core/gateway-routing.ts +157 -0
  267. package/src/core/safe-exec.ts +17 -2
  268. package/src/core/skill-activation.ts +89 -2
  269. package/src/domains/agents/builtins/world-knowledge.md +31 -0
  270. package/src/domains/agents/catalog.ts +1 -1
  271. package/src/domains/agents/result-contract.ts +70 -0
  272. package/src/domains/context/wiki/map-seed.ts +589 -0
  273. package/src/domains/context/wiki/plan.ts +2 -2
  274. package/src/domains/dispatch/admission.ts +29 -0
  275. package/src/domains/dispatch/agent-candidates.ts +10 -0
  276. package/src/domains/dispatch/budget-envelope.ts +86 -1
  277. package/src/domains/dispatch/capability-match.ts +1 -0
  278. package/src/domains/dispatch/capacity-lease.ts +17 -0
  279. package/src/domains/dispatch/contract.ts +11 -1
  280. package/src/domains/dispatch/extension.ts +134 -29
  281. package/src/domains/dispatch/types.ts +3 -0
  282. package/src/domains/dispatch/worker-model-metadata.ts +38 -0
  283. package/src/domains/eval/metrics/call-ledger-stream.ts +34 -11
  284. package/src/domains/eval/metrics/token-stream.ts +201 -31
  285. package/src/domains/eval/metrics/tracked.ts +40 -4
  286. package/src/domains/eval/runners/clio-run.ts +5 -2
  287. package/src/domains/eval/schema/suite.ts +28 -0
  288. package/src/domains/eval/schema/verdict.ts +2 -2
  289. package/src/domains/eval/suites/resolve.ts +13 -1
  290. package/src/domains/eval/suites/run.ts +24 -3
  291. package/src/domains/interop/registry.ts +6 -2
  292. package/src/domains/interop/types.ts +4 -0
  293. package/src/domains/lifecycle/migrations/index.ts +4 -0
  294. package/src/domains/memory/task-memory-policy.ts +70 -26
  295. package/src/domains/memory/task-memory-telemetry.ts +1 -0
  296. package/src/domains/middleware/index.ts +0 -1
  297. package/src/domains/middleware/marketplace-offer.ts +3 -35
  298. package/src/domains/middleware/memory-intervention.ts +127 -32
  299. package/src/domains/middleware/memory-step-endpoint.ts +3 -2
  300. package/src/domains/middleware/skills-reminder.ts +31 -2
  301. package/src/domains/observability/compaction-usage.ts +118 -0
  302. package/src/domains/observability/cost.ts +1 -1
  303. package/src/domains/observability/extension.ts +6 -1
  304. package/src/domains/observability/out-of-turn-usage.ts +52 -21
  305. package/src/domains/providers/contract.ts +4 -1
  306. package/src/domains/providers/extension.ts +40 -9
  307. package/src/domains/providers/model-capabilities.ts +9 -0
  308. package/src/domains/providers/model-discovery.ts +2 -0
  309. package/src/domains/providers/model-runtime-capabilities.ts +15 -5
  310. package/src/domains/providers/models/local-models/clio-coder-local-coding-targets.yaml +32 -12
  311. package/src/domains/providers/runtimes/antigravity/antigravity-code.ts +225 -45
  312. package/src/domains/providers/runtimes/common/lmstudio-http.ts +6 -2
  313. package/src/domains/providers/runtimes/common/local-synth.ts +2 -0
  314. package/src/domains/providers/runtimes/protocol/litellm.ts +119 -29
  315. package/src/domains/providers/support.ts +11 -5
  316. package/src/domains/providers/target-model-cache.ts +25 -2
  317. package/src/domains/providers/types/capability-flags.ts +2 -0
  318. package/src/domains/providers/types/runtime-descriptor.ts +20 -1
  319. package/src/domains/providers/types/target-descriptor.ts +19 -0
  320. package/src/domains/resources/index.ts +3 -0
  321. package/src/domains/resources/skills/install.ts +72 -7
  322. package/src/domains/resources/skills/loader.ts +7 -0
  323. package/src/domains/resources/skills/marketplace.ts +63 -11
  324. package/src/domains/safety/autonomy.ts +15 -0
  325. package/src/domains/safety/index.ts +1 -0
  326. package/src/domains/safety/path-policy.ts +1 -1
  327. package/src/domains/safety/policy-engine.ts +34 -11
  328. package/src/domains/safety/protected-artifacts.ts +191 -88
  329. package/src/domains/safety/run-effects.ts +2 -22
  330. package/src/domains/safety/skill-authority.ts +55 -0
  331. package/src/domains/session/compaction/compact.ts +72 -22
  332. package/src/domains/session/entries.ts +6 -0
  333. package/src/domains/session/usage.ts +3 -3
  334. package/src/engine/agent.ts +13 -3
  335. package/src/engine/ai.ts +26 -8
  336. package/src/engine/antigravity/subprocess-runtime.ts +386 -120
  337. package/src/engine/api-registry.ts +3 -0
  338. package/src/engine/apis/openai-completions.ts +117 -14
  339. package/src/engine/external-subprocess.ts +114 -6
  340. package/src/entry/background-model-metadata.ts +18 -0
  341. package/src/entry/compaction-prompt.ts +57 -0
  342. package/src/entry/orchestrator.ts +405 -216
  343. package/src/entry/task-memory-lifecycle.ts +35 -0
  344. package/src/interactive/chat-loop-messages.ts +13 -4
  345. package/src/interactive/chat-loop.ts +65 -2
  346. package/src/interactive/chat-renderer.ts +1 -0
  347. package/src/interactive/cost-overlay.ts +26 -2
  348. package/src/interactive/interactive-slash-runtime.ts +2 -1
  349. package/src/interactive/renderers/worker-entry.ts +32 -0
  350. package/src/interactive/slash-commands.ts +24 -6
  351. package/src/interactive/theme/labels.ts +19 -13
  352. package/src/interactive/turn-context.ts +9 -5
  353. package/src/interactive/turn-recovery.ts +8 -0
  354. package/src/interactive/turn-runtime.ts +27 -11
  355. package/src/interactive/turn-state.ts +7 -0
  356. package/src/interactive/worker-receipts.ts +1 -0
  357. package/src/interactive/worker-stream.ts +6 -1
  358. package/src/tools/context/index.ts +30 -9
  359. package/src/tools/dispatch-arguments.ts +1 -0
  360. package/src/tools/dispatch-event-text.ts +10 -0
  361. package/src/tools/dispatch-plan.ts +1 -0
  362. package/src/tools/dispatch-runner.ts +12 -0
  363. package/src/tools/registry.ts +11 -5
  364. package/src/tools/worker-evidence.ts +3 -1
  365. package/src/worker/spec-contract.ts +4 -0
  366. package/dist/chunk-2Z2IKEXI.js +0 -1554
  367. package/dist/reset-OAQP3W4O.js +0 -230
  368. package/dist/uninstall-N34PCTGJ.js +0 -331
  369. package/dist/upgrade-PXK3S2YM.js +0 -325
@@ -0,0 +1,118 @@
1
+ /** Preserve spending from summary streams that did not produce a checkpoint. */
2
+
3
+ import type { CompactionCallObservation } from "../session/compaction/compact.js";
4
+ import { extractReasoningTokens } from "../session/context-accounting.js";
5
+ import type { BackgroundMemoryUsageSink } from "./background-memory-usage.js";
6
+ import { appendOutOfTurnUsageRow, type OutOfTurnUsage, type OutOfTurnUsageRow } from "./out-of-turn-usage.js";
7
+
8
+ export interface CompactionUsageOrigin {
9
+ stateDir: string;
10
+ sessionId: string;
11
+ repoIdentity: string;
12
+ target: string;
13
+ model: string;
14
+ }
15
+
16
+ /** Failed adapter zeros cannot distinguish absent usage from a reported zero. */
17
+ function observedUsage(call: CompactionCallObservation): OutOfTurnUsage {
18
+ const raw = isRecord(call.usage) ? call.usage : {};
19
+ const reading = (value: unknown): number | null =>
20
+ typeof value === "number" && Number.isFinite(value) && value >= 0 && (call.outcome === "success" || value > 0)
21
+ ? value
22
+ : null;
23
+ const cost = isRecord(raw.cost) ? reading(raw.cost.total) : null;
24
+ // Adapter cost is a price estimate, not independent provider billing evidence.
25
+ const costUsd = cost !== null && cost > 0 ? cost : null;
26
+ const reasoning = extractReasoningTokens({ ...raw, reasoningTokens: undefined });
27
+ return {
28
+ input: reading(raw.input),
29
+ output: reading(raw.output),
30
+ cacheRead: reading(raw.cacheRead),
31
+ cacheWrite: reading(raw.cacheWrite),
32
+ reasoning: reasoning !== null && reasoning > 0 ? reasoning : null,
33
+ totalTokens: reading(raw.totalTokens),
34
+ costUsd,
35
+ costProvenance: costUsd === null ? "unknown" : "estimated",
36
+ };
37
+ }
38
+
39
+ /**
40
+ * The checkpoint is the sole accounting source on success. Call this only when
41
+ * no checkpoint exists; each row is one actual stream invocation, not a retry.
42
+ * Missing amounts remain null on disk. Numeric live counters are known subtotals.
43
+ */
44
+ export function recordFailedCompactionCalls(
45
+ origin: CompactionUsageOrigin,
46
+ calls: ReadonlyArray<CompactionCallObservation>,
47
+ observability?: BackgroundMemoryUsageSink,
48
+ ): void {
49
+ for (const call of calls) {
50
+ const usage = observedUsage(call);
51
+ const row: OutOfTurnUsageRow = {
52
+ label: "failed-compaction",
53
+ callOutcome: call.outcome,
54
+ sessionId: origin.sessionId,
55
+ repoIdentity: origin.repoIdentity,
56
+ timestamp: call.timestamp,
57
+ target: origin.target,
58
+ attributedModelId: origin.model,
59
+ usage,
60
+ timing: { durationMs: call.durationMs },
61
+ };
62
+ appendOutOfTurnUsageRow(origin.stateDir, row, { required: true });
63
+ // Do not create a measured-zero live entry for a wholly unobserved call.
64
+ if (!Object.values(usage).some((value) => typeof value === "number" && value > 0)) continue;
65
+ observability?.recordTokens(
66
+ origin.target,
67
+ origin.model,
68
+ usage.totalTokens ?? 0,
69
+ usage.costUsd ?? 0,
70
+ {
71
+ input: usage.input ?? 0,
72
+ output: usage.output ?? 0,
73
+ cacheRead: usage.cacheRead ?? 0,
74
+ cacheWrite: usage.cacheWrite ?? 0,
75
+ reasoningTokens: usage.reasoning ?? 0,
76
+ totalTokens: usage.totalTokens ?? 0,
77
+ apiCalls: 1,
78
+ },
79
+ usage.costProvenance,
80
+ undefined,
81
+ "failed-compaction",
82
+ );
83
+ }
84
+ }
85
+
86
+ function isRecord(value: unknown): value is Record<string, unknown> {
87
+ return typeof value === "object" && value !== null && !Array.isArray(value);
88
+ }
89
+
90
+ const USAGE_FIELDS = ["input", "output", "cacheRead", "cacheWrite", "reasoning", "totalTokens", "costUsd"] as const;
91
+ type UsageField = (typeof USAGE_FIELDS)[number];
92
+
93
+ /** Known subtotals and missing-field coverage; zero missing values are never inferred. */
94
+ export function summarizeFailedCompactionUsage(rows: ReadonlyArray<OutOfTurnUsageRow>) {
95
+ const calls = rows.filter((row) => row.label === "failed-compaction");
96
+ const knownUsage = Object.fromEntries(USAGE_FIELDS.map((field) => [field, null])) as Record<UsageField, number | null>;
97
+ const erroredKnownUsage = { ...knownUsage };
98
+ const unobservedUsageCalls = Object.fromEntries(USAGE_FIELDS.map((field) => [field, 0])) as Record<UsageField, number>;
99
+ for (const row of calls) {
100
+ for (const field of USAGE_FIELDS) {
101
+ const value = row.usage[field];
102
+ if (value === null) unobservedUsageCalls[field] += 1;
103
+ else {
104
+ knownUsage[field] = (knownUsage[field] ?? 0) + value;
105
+ if (row.callOutcome === "error") erroredKnownUsage[field] = (erroredKnownUsage[field] ?? 0) + value;
106
+ }
107
+ }
108
+ }
109
+ return {
110
+ calls: calls.length,
111
+ successfulCalls: calls.filter((row) => row.callOutcome === "success").length,
112
+ erroredCalls: calls.filter((row) => row.callOutcome === "error").length,
113
+ abortedCalls: calls.filter((row) => row.callOutcome === "aborted").length,
114
+ knownUsage,
115
+ erroredKnownUsage,
116
+ unobservedUsageCalls,
117
+ };
118
+ }
@@ -137,7 +137,7 @@ export interface UsageBreakdown {
137
137
  * the usage surfaces separate these out so an operator can see that money was
138
138
  * spent beside the session rather than inside it.
139
139
  */
140
- export type CostEntryLabel = "side-question" | "handoff" | "prewarm" | "background-memory";
140
+ export type CostEntryLabel = "side-question" | "handoff" | "prewarm" | "background-memory" | "failed-compaction";
141
141
 
142
142
  export interface CostEntry {
143
143
  providerId: string;
@@ -37,6 +37,11 @@ type DispatchTerminalLike = {
37
37
  [K in keyof DispatchCompletedPayload]?: DispatchCompletedPayload[K] | undefined;
38
38
  };
39
39
 
40
+ /** Pre-admission failures have an announced run id but never create a run ledger. */
41
+ export function dispatchHasEvidenceLedger(payload: DispatchTerminalLike): boolean {
42
+ return payload.lineage !== undefined;
43
+ }
44
+
40
45
  function recordDispatchCost(
41
46
  telemetry: ReturnType<typeof createTelemetry>,
42
47
  cost: ReturnType<typeof createCostTracker>,
@@ -222,7 +227,7 @@ export function createObservabilityBundle(
222
227
  recordDispatchCost(telemetry, cost, payload);
223
228
  // A failed run is never a first-pass success; still build the
224
229
  // bundle so its failure-cause tags exist for the index.
225
- if (typeof payload.runId === "string" && payload.runId.length > 0) {
230
+ if (typeof payload.runId === "string" && payload.runId.length > 0 && dispatchHasEvidenceLedger(payload)) {
226
231
  trackBuild(payload.runId, false, payload.lineage?.attempt);
227
232
  }
228
233
  }),
@@ -27,7 +27,7 @@
27
27
  * skipped, never fatal.
28
28
  */
29
29
 
30
- import { closeSync, mkdirSync, openSync, readFileSync, writeSync } from "node:fs";
30
+ import { closeSync, fsyncSync, mkdirSync, openSync, readFileSync, writeSync } from "node:fs";
31
31
  import { dirname, join } from "node:path";
32
32
  import { safeResourceWrite } from "../../core/safe-resource-write.js";
33
33
  import { withStateFileLockSync } from "../../core/state-file-lock.js";
@@ -74,13 +74,13 @@ export interface OutOfTurnPromptCache {
74
74
 
75
75
  /** Provider-reported usage for one out-of-turn call. */
76
76
  export interface OutOfTurnUsage {
77
- input: number;
78
- output: number;
79
- cacheRead: number;
80
- cacheWrite: number;
81
- reasoning: number;
82
- totalTokens: number;
83
- costUsd: number;
77
+ input: number | null;
78
+ output: number | null;
79
+ cacheRead: number | null;
80
+ cacheWrite: number | null;
81
+ reasoning: number | null;
82
+ totalTokens: number | null;
83
+ costUsd: number | null;
84
84
  costProvenance: CostProvenance;
85
85
  }
86
86
 
@@ -97,6 +97,8 @@ export interface OutOfTurnUsageRow {
97
97
  target: string;
98
98
  attributedModelId: string;
99
99
  usage: OutOfTurnUsage;
100
+ /** New failed-compaction rows preserve unknown usage as null. Legacy rows are unchanged. */
101
+ callOutcome?: "success" | "error" | "aborted";
100
102
  /** Present when the caller measured the call. */
101
103
  timing?: OutOfTurnTiming;
102
104
  /** Present when the serving backend reported prefill facts. */
@@ -115,21 +117,31 @@ export function outOfTurnUsagePath(stateDir: string): string {
115
117
  let appendsSinceBoundCheck = BOUND_CHECK_INTERVAL;
116
118
 
117
119
  /**
118
- * Append one priced out-of-turn call. Never throws: this runs on the same path
119
- * that answers an operator's side question, and a failed bookkeeping write must
120
- * not take the answer down with it.
120
+ * Append one out-of-turn call. Optional bookkeeping keeps its existing diagnostic
121
+ * write policy. Required failed-compaction records flush the append and throw on
122
+ * a failed or incomplete write, so their caller cannot report recorded spending.
121
123
  */
122
- export function appendOutOfTurnUsageRow(stateDir: string, row: OutOfTurnUsageRow): void {
124
+ export function appendOutOfTurnUsageRow(
125
+ stateDir: string,
126
+ row: OutOfTurnUsageRow,
127
+ options: { required?: boolean } = {},
128
+ ): void {
123
129
  const path = outOfTurnUsagePath(stateDir);
124
130
  try {
125
131
  mkdirSync(dirname(path), { recursive: true });
126
132
  const fd = openSync(path, "a");
127
133
  try {
128
- writeSync(fd, `${JSON.stringify(row)}\n`);
134
+ const line = `${JSON.stringify(row)}\n`;
135
+ const written = writeSync(fd, line);
136
+ if (options.required) {
137
+ if (written !== Buffer.byteLength(line)) throw new Error("incomplete usage row write");
138
+ fsyncSync(fd);
139
+ }
129
140
  } finally {
130
141
  closeSync(fd);
131
142
  }
132
143
  } catch (error) {
144
+ if (options.required) throw new Error(`compaction usage row not written: ${messageOf(error)}`, { cause: error });
133
145
  process.stderr.write(`[clio-coder:usage] out-of-turn usage row not written: ${messageOf(error)}\n`);
134
146
  return;
135
147
  }
@@ -205,11 +217,25 @@ export function readOutOfTurnUsageRows(stateDir: string): OutOfTurnUsageReadResu
205
217
  function asOutOfTurnUsageRow(value: unknown): OutOfTurnUsageRow | null {
206
218
  if (!isRecord(value)) return null;
207
219
  const label = value.label;
208
- if (label !== "side-question" && label !== "handoff" && label !== "prewarm" && label !== "background-memory")
220
+ if (
221
+ label !== "side-question" &&
222
+ label !== "handoff" &&
223
+ label !== "prewarm" &&
224
+ label !== "background-memory" &&
225
+ label !== "failed-compaction"
226
+ )
209
227
  return null;
210
228
  if (typeof value.timestamp !== "string" || value.timestamp.length === 0) return null;
211
229
  if (!isRecord(value.usage)) return null;
212
230
  const usage = value.usage;
231
+ if (
232
+ label === "failed-compaction" &&
233
+ value.callOutcome !== "success" &&
234
+ value.callOutcome !== "error" &&
235
+ value.callOutcome !== "aborted"
236
+ )
237
+ return null;
238
+ const reading = label === "failed-compaction" ? nullableNumber : numberOr0;
213
239
  const timing = asTiming(value.timing);
214
240
  const promptCache = asPromptCache(value.promptCache);
215
241
  return {
@@ -223,15 +249,16 @@ function asOutOfTurnUsageRow(value: unknown): OutOfTurnUsageRow | null {
223
249
  ? value.attributedModelId
224
250
  : "unknown",
225
251
  usage: {
226
- input: numberOr0(usage.input),
227
- output: numberOr0(usage.output),
228
- cacheRead: numberOr0(usage.cacheRead),
229
- cacheWrite: numberOr0(usage.cacheWrite),
230
- reasoning: numberOr0(usage.reasoning),
231
- totalTokens: numberOr0(usage.totalTokens),
232
- costUsd: numberOr0(usage.costUsd),
252
+ input: reading(usage.input),
253
+ output: reading(usage.output),
254
+ cacheRead: reading(usage.cacheRead),
255
+ cacheWrite: reading(usage.cacheWrite),
256
+ reasoning: reading(usage.reasoning),
257
+ totalTokens: reading(usage.totalTokens),
258
+ costUsd: reading(usage.costUsd),
233
259
  costProvenance: asCostProvenance(usage.costProvenance),
234
260
  },
261
+ ...(label === "failed-compaction" ? { callOutcome: value.callOutcome as "success" | "error" | "aborted" } : {}),
235
262
  ...(timing === null ? {} : { timing }),
236
263
  ...(promptCache === null ? {} : { promptCache }),
237
264
  };
@@ -259,6 +286,10 @@ function asCostProvenance(value: unknown): CostProvenance {
259
286
  return value === "known" || value === "known_free" || value === "estimated" ? value : "unknown";
260
287
  }
261
288
 
289
+ function nullableNumber(value: unknown): number | null {
290
+ return typeof value === "number" && Number.isFinite(value) && value >= 0 ? value : null;
291
+ }
292
+
262
293
  function numberOr0(value: unknown): number {
263
294
  return typeof value === "number" && Number.isFinite(value) && value > 0 ? value : 0;
264
295
  }
@@ -59,6 +59,8 @@ export interface TargetStatus {
59
59
  probeSurfaces?: Readonly<ProbeSurfaceMap>;
60
60
  /** Ids returned by the last successful probeModels() call. */
61
61
  discoveredModels: ReadonlyArray<string>;
62
+ /** Human-readable live/cache labels keyed by exact model slug. */
63
+ discoveredModelLabels?: Readonly<Record<string, string>>;
62
64
  /**
63
65
  * Source for `discoveredModels`. `probe` means the target just returned a
64
66
  * live catalog, `cache` is a previously probed catalog preserved across a
@@ -93,7 +95,8 @@ export interface ProvidersContract {
93
95
  * `reasoning: false` skips an inference-based reasoning-capability probe while
94
96
  * retaining the target's liveness and model-catalog checks.
95
97
  */
96
- probeTarget(id: string, options?: { reasoning?: boolean }): Promise<TargetStatus | null>;
98
+ /** Optional cancellation covers auth and metadata; a cancelled probe does not publish health. */
99
+ probeTarget(id: string, options?: { reasoning?: boolean; signal?: AbortSignal }): Promise<TargetStatus | null>;
97
100
 
98
101
  /** Clear in-memory live connection state for a configured target. */
99
102
  disconnectTarget(id: string): TargetStatus | null;
@@ -156,6 +156,7 @@ function sameProbeIdentity(previous: TargetDescriptor, next: TargetDescriptor):
156
156
 
157
157
  export interface ProbeMerge {
158
158
  discoveredModels: string[];
159
+ discoveredModelLabels: Readonly<Record<string, string>>;
159
160
  discoveredModelsSource: "probe" | "cache" | "runtime" | "none";
160
161
  discoveredModelStates: NonNullable<TargetStatus["discoveredModelStates"]> | null;
161
162
  probeCapabilities: NonNullable<TargetStatus["probeCapabilities"]> | null;
@@ -179,6 +180,7 @@ function mergeProbeResult(
179
180
  probe: ProbeResult | null,
180
181
  previous: TargetStatus | undefined,
181
182
  cachedModels: ReadonlyArray<string> = [],
183
+ cachedModelLabels: Readonly<Record<string, string>> = {},
182
184
  ): ProbeMerge {
183
185
  const probeSucceeded = probe?.ok ?? false;
184
186
  const preservePrevious = !probeSucceeded && previous !== undefined && sameProbeIdentity(previous.target, target);
@@ -201,8 +203,14 @@ function mergeProbeResult(
201
203
  [],
202
204
  );
203
205
  const discoveredModelStates = probe?.modelStates ?? (preservePrevious ? previous.discoveredModelStates : null) ?? null;
206
+ const discoveredModelLabels =
207
+ probe?.modelLabels ??
208
+ (preservePrevious ? previous.discoveredModelLabels : undefined) ??
209
+ (cachedModels.length > 0 ? cachedModelLabels : undefined) ??
210
+ {};
204
211
  const merge: ProbeMerge = {
205
212
  discoveredModels,
213
+ discoveredModelLabels,
206
214
  discoveredModelsSource: discoveredModelsSource(probe, preservePrevious, previous, cachedModels, desc),
207
215
  discoveredModelStates,
208
216
  probeCapabilities,
@@ -225,14 +233,16 @@ function mergeProbeResult(
225
233
  * there is no live evidence either way.
226
234
  */
227
235
  function unservedDefaultModelReason(
228
- desc: Pick<RuntimeDescriptor, "id">,
236
+ desc: Pick<RuntimeDescriptor, "id" | "externalAgentLoop">,
229
237
  target: Pick<TargetDescriptor, "defaultModel">,
230
238
  merge: Pick<ProbeMerge, "discoveredModels" | "discoveredModelsSource" | "discoveredModelStates">,
231
239
  ): string | null {
232
240
  const model = target.defaultModel;
233
241
  if (!model) return null;
234
242
  if (merge.discoveredModelsSource !== "probe" || merge.discoveredModels.length === 0) return null;
235
- if (listKnownModelsForRuntime(desc.id).length > 0) return null;
243
+ if (listKnownModelsForRuntime(desc.id).length > 0 && desc.externalAgentLoop?.modelCatalog !== "live-authoritative") {
244
+ return null;
245
+ }
236
246
  if (merge.discoveredModels.includes(model)) return null;
237
247
  // LM Studio lists a loaded model under its instance id and keeps the model
238
248
  // key in the state map; the request path resolves either.
@@ -258,14 +268,22 @@ export function createProvidersBundle(context: DomainContext): DomainBundle<Prov
258
268
  return readSettings();
259
269
  }
260
270
 
261
- async function buildProbeContextForTarget(target: TargetDescriptor, desc: RuntimeDescriptor): Promise<ProbeContext> {
271
+ async function buildProbeContextForTarget(
272
+ target: TargetDescriptor,
273
+ desc: RuntimeDescriptor,
274
+ signal?: AbortSignal,
275
+ ): Promise<ProbeContext> {
262
276
  const probeCtx: ProbeContext = {
263
277
  credentialsPresent: credentialsPresent(),
264
278
  httpTimeoutMs: DEFAULT_PROBE_TIMEOUT_MS,
279
+ ...(signal ? { signal } : {}),
265
280
  };
266
281
  if (!targetRequiresAuth(target, desc)) return probeCtx;
267
282
  try {
268
- const resolution = await authStore.resolveForTarget(resolveAuthTarget(target, desc), { includeFallback: false });
283
+ const resolution = await authStore.resolveForTarget(resolveAuthTarget(target, desc), {
284
+ includeFallback: false,
285
+ ...(signal ? { signal } : {}),
286
+ });
269
287
  if (resolution.apiKey) probeCtx.authToken = resolution.apiKey;
270
288
  } catch {
271
289
  // Authenticated probes should still return a runtime-specific missing-auth error.
@@ -325,8 +343,9 @@ export function createProvidersBundle(context: DomainContext): DomainBundle<Prov
325
343
  return out;
326
344
  }
327
345
  const availability = availabilityFor(desc, target, authStatusFor);
328
- const cachedModels = readTargetModelSnapshot(target)?.models ?? [];
329
- const merge = mergeProbeResult(desc, target, probe, previous, cachedModels);
346
+ const cachedSnapshot = readTargetModelSnapshot(target);
347
+ const cachedModels = cachedSnapshot?.models ?? [];
348
+ const merge = mergeProbeResult(desc, target, probe, previous, cachedModels, cachedSnapshot?.modelLabels ?? {});
330
349
  const { capabilities, contextWindowProvenance } = capabilitiesFor(desc, target, merge, kb);
331
350
  const healthy = probe !== null ? probe.ok : null;
332
351
  const unservedDefault = probe?.ok ? unservedDefaultModelReason(desc, target, merge) : null;
@@ -353,6 +372,7 @@ export function createProvidersBundle(context: DomainContext): DomainBundle<Prov
353
372
  probeModelCapabilities: merge.probeModelCapabilities,
354
373
  probeModelId: merge.probeModelId,
355
374
  discoveredModels: merge.discoveredModels,
375
+ discoveredModelLabels: merge.discoveredModelLabels,
356
376
  discoveredModelsSource: merge.discoveredModelsSource,
357
377
  discoveredModelStates: merge.discoveredModelStates,
358
378
  };
@@ -364,8 +384,9 @@ export function createProvidersBundle(context: DomainContext): DomainBundle<Prov
364
384
  async function probeTargetInternal(
365
385
  target: TargetDescriptor,
366
386
  live: boolean,
367
- options?: { reasoning?: boolean },
387
+ options?: { reasoning?: boolean; signal?: AbortSignal },
368
388
  ): Promise<TargetStatus> {
389
+ options?.signal?.throwIfAborted();
369
390
  const previous = statuses.get(target.id);
370
391
  const desc = registry.get(target.runtime);
371
392
  if (!desc) {
@@ -379,13 +400,15 @@ export function createProvidersBundle(context: DomainContext): DomainBundle<Prov
379
400
  statuses.set(target.id, status);
380
401
  return status;
381
402
  }
382
- const probeCtx = await buildProbeContextForTarget(target, desc);
403
+ const probeCtx = await buildProbeContextForTarget(target, desc, options?.signal);
404
+ options?.signal?.throwIfAborted();
383
405
  let probeResult: ProbeResult;
384
406
  try {
385
407
  probeResult = await desc.probe(target, probeCtx);
386
408
  } catch (err) {
387
409
  probeResult = { ok: false, error: err instanceof Error ? err.message : String(err) };
388
410
  }
411
+ options?.signal?.throwIfAborted();
389
412
  if (probeResult.ok && typeof desc.probeModels === "function" && !probeResult.models) {
390
413
  try {
391
414
  const ids = await desc.probeModels(target, probeCtx);
@@ -394,6 +417,7 @@ export function createProvidersBundle(context: DomainContext): DomainBundle<Prov
394
417
  // model discovery is best-effort; keep probe as-is.
395
418
  }
396
419
  }
420
+ options?.signal?.throwIfAborted();
397
421
  if (probeResult.ok && options?.reasoning !== false && typeof desc.probeReasoning === "function") {
398
422
  const settings = readConfig();
399
423
  const orchestratorTarget = settings.chat.target === target.id ? settings.chat.model : null;
@@ -401,6 +425,7 @@ export function createProvidersBundle(context: DomainContext): DomainBundle<Prov
401
425
  if (candidateModelId) {
402
426
  try {
403
427
  const result = await desc.probeReasoning(target, candidateModelId, probeCtx);
428
+ options?.signal?.throwIfAborted();
404
429
  reasoningCache.set(reasoningCacheKey(target.id, candidateModelId), result.reasoning);
405
430
  const capabilityModelId = probeResult.capabilityModelId ?? null;
406
431
  if (capabilityModelId === null || capabilityModelId === candidateModelId) {
@@ -422,10 +447,16 @@ export function createProvidersBundle(context: DomainContext): DomainBundle<Prov
422
447
  }
423
448
  }
424
449
  }
450
+ // A cancelled role preparation must not publish failed/stale target state.
451
+ options?.signal?.throwIfAborted();
425
452
  const status = buildStatus(target, desc, probeResult, previous);
426
453
  statuses.set(target.id, status);
427
454
  if (probeResult.ok && probeResult.models !== undefined) {
428
- recordTargetModelSnapshot(target, probeResult.models);
455
+ recordTargetModelSnapshot(
456
+ target,
457
+ probeResult.models,
458
+ probeResult.modelLabels ? { modelLabels: probeResult.modelLabels } : {},
459
+ );
429
460
  }
430
461
  // Durable so the next process resolves this endpoint's real slot count
431
462
  // instead of the conservative default. Fire and forget: the probe's answer
@@ -14,6 +14,7 @@ export interface ModelCapabilityPatchTarget {
14
14
  contextWindow?: number;
15
15
  maxTokens?: number;
16
16
  reasoning?: boolean;
17
+ clioCoder?: Record<string, unknown>;
17
18
  }
18
19
 
19
20
  /**
@@ -30,6 +31,14 @@ export function applyModelCapabilityPatch<T extends ModelCapabilityPatchTarget>(
30
31
  if (typeof caps.contextWindow === "number") model.contextWindow = caps.contextWindow;
31
32
  if (typeof caps.maxTokens === "number") model.maxTokens = caps.maxTokens;
32
33
  if (typeof caps.reasoning === "boolean") model.reasoning = caps.reasoning;
34
+ // Refresh the probe-only control hint together with the capability snapshot.
35
+ // Missing metadata after a later probe must not retain an earlier route claim.
36
+ if (model.clioCoder || caps.thinkingControlRuntime !== undefined) {
37
+ const metadata = { ...model.clioCoder };
38
+ delete metadata.thinkingControlRuntime;
39
+ if (caps.thinkingControlRuntime !== undefined) metadata.thinkingControlRuntime = caps.thinkingControlRuntime;
40
+ model.clioCoder = metadata;
41
+ }
33
42
  return model;
34
43
  }
35
44
 
@@ -6,6 +6,7 @@ export type ProviderModelSource = "configured" | "live" | "catalog" | "default";
6
6
 
7
7
  export interface ProviderModelCandidate {
8
8
  id: string;
9
+ label?: string;
9
10
  source: ProviderModelSource;
10
11
  loadState?: string;
11
12
  loadStateDetail?: string;
@@ -100,6 +101,7 @@ export function modelCandidatesForStatus(status: TargetStatus): ProviderModelCan
100
101
  const state = status.discoveredModelStates?.[trimmed];
101
102
  out.push({
102
103
  id: trimmed,
104
+ ...(status.discoveredModelLabels?.[trimmed] ? { label: status.discoveredModelLabels[trimmed] } : {}),
103
105
  source,
104
106
  ...(state ? { loadState: state.state } : {}),
105
107
  ...(state?.detail ? { loadStateDetail: state.detail } : {}),
@@ -131,6 +131,7 @@ interface ClioRuntimeMetadata {
131
131
  lifecycle?: "user-managed" | "clio-coder-managed";
132
132
  gateway?: boolean;
133
133
  family?: string;
134
+ thinkingControlRuntime?: CapabilityFlags["thinkingControlRuntime"];
134
135
  quirks?: LocalModelQuirks;
135
136
  };
136
137
  /** Released in-memory metadata name accepted at plugin/runtime read boundaries. */
@@ -140,6 +141,7 @@ interface ClioRuntimeMetadata {
140
141
  lifecycle?: "user-managed" | "clio-coder-managed";
141
142
  gateway?: boolean;
142
143
  family?: string;
144
+ thinkingControlRuntime?: CapabilityFlags["thinkingControlRuntime"];
143
145
  quirks?: LocalModelQuirks;
144
146
  };
145
147
  compat?: {
@@ -381,9 +383,10 @@ export function effectiveThinkingLevel(
381
383
  const fallback = available[0] ?? "off";
382
384
  if (!configured) return fallback;
383
385
  if (available.includes(configured)) return configured;
384
- // Catalog effort maps intentionally stop at xhigh; `max` must clamp to that
385
- // supported ceiling instead of falling through to the generic low fallback.
386
- if (configured === "max" && available.includes("xhigh")) return "xhigh";
386
+ // A family with xhigh but no high keeps the released high alias at its
387
+ // supported ceiling; neither high nor max may fall through to low.
388
+ if ((configured === "max" || (configured === "high" && !available.includes("high"))) && available.includes("xhigh"))
389
+ return "xhigh";
387
390
  if ((configured === "high" || configured === "xhigh" || configured === "max") && available.includes("high")) {
388
391
  return "high";
389
392
  }
@@ -553,7 +556,7 @@ function resolveThinkingCapability(
553
556
  */
554
557
  const REASONING_EFFORT_ONLY_RUNTIMES: ReadonlySet<string> = new Set(["lmstudio"]);
555
558
 
556
- /** `none` is LM Studio's documented off value; on-off models have no finer dial than `low`. */
559
+ /** `none` is the observed LM Studio HTTP off value; on-off models have no finer dial than `low`. */
557
560
  function onOffReasoningEffort(thinkingActive: boolean): string {
558
561
  return thinkingActive ? "low" : "none";
559
562
  }
@@ -647,7 +650,12 @@ export function resolveModelRuntimeCapabilities(
647
650
  family,
648
651
  capabilities: input.capabilities,
649
652
  thinking,
650
- request: resolveRequestCapability(thinking, parser, input.runtimeId, quirks),
653
+ request: resolveRequestCapability(
654
+ thinking,
655
+ parser,
656
+ input.runtimeId === "litellm" ? (input.capabilities.thinkingControlRuntime ?? input.runtimeId) : input.runtimeId,
657
+ quirks,
658
+ ),
651
659
  response: {
652
660
  parser,
653
661
  stripTokenizerSentinels: true,
@@ -735,6 +743,7 @@ function thinkingFormatFromModelApi(api: Api): CapabilityFlags["thinkingFormat"]
735
743
 
736
744
  function capabilitiesFromModel(model: Model<Api> & ClioRuntimeMetadata): CapabilityFlags {
737
745
  const format = model.compat?.thinkingFormat ?? thinkingFormatFromModelApi(model.api);
746
+ const controlRuntime = (model.clioCoder ?? model.clio)?.thinkingControlRuntime;
738
747
  const caps: CapabilityFlags = {
739
748
  chat: true,
740
749
  tools: true,
@@ -746,6 +755,7 @@ function capabilitiesFromModel(model: Model<Api> & ClioRuntimeMetadata): Capabil
746
755
  fim: false,
747
756
  contextWindow: model.contextWindow,
748
757
  maxTokens: model.maxTokens,
758
+ ...(controlRuntime ? { thinkingControlRuntime: controlRuntime } : {}),
749
759
  };
750
760
  if (
751
761
  format === "qwen-chat-template" ||
@@ -561,8 +561,12 @@
561
561
  guidance: |
562
562
  Ornith 1.5's Qwen3.5 chat template exposes thinking as on or off
563
563
  through enable_thinking; intermediate levels coerce to on. LM Studio
564
- reads the same switch from reasoning_effort (none or low), which Clio
565
- sends for on-off models on that runtime.
564
+ 2026-09-03 exposed a split vocabulary on this exact Q4_K_M: native
565
+ discovery advertised {off,on}, while its OpenAI-compatible request
566
+ validator rejected literal `on` and accepted only effort-level values.
567
+ `low` ran only after a warning and fallback to on, so Clio now omits an
568
+ active override for this default-on model and sends `reasoning_effort:
569
+ none` only when disabling it.
566
570
  serving: "Ornith-1.5-35B-A3B follows the 1.0 card: 262K context, qwen3 reasoning parsing, qwen3_coder tool parsing, sampler temperature=0.6, top_p=0.95, top_k=20."
567
571
 
568
572
  - family: qwen3.6-27b
@@ -696,7 +700,7 @@
696
700
  "128gb-unified-vulkan": "128 GB unified-memory Vulkan class: IQ4_NL uses about 17.3 GB and fits ctx=262144 with parallel=4. Vulkan backends show much slower prefill than ROCm/CUDA on this family; prefer ROCm where available."
697
701
  runtimePreference:
698
702
  llamaCpp: "Primary verified path. OpenAI-compatible chat completions with --jinja, --reasoning on, --reasoning-effort medium as server defaults; reasoning_content and tool_calls both confirmed to parse correctly. A request can override to off with chat_template_kwargs.enable_thinking:false, or to a different active level with reasoning_effort in {low, medium, xhigh} (never minimal/high/max on the wire; the stock template raises a fatal jinja exception on those and Clio's landed capability resolver already clamps to this set)."
699
- lmstudioOpenaiCompat: "Clio sends reasoning_effort over LM Studio's OpenAI-compatible HTTP port and omits chat_template_kwargs because LM Studio does not use that field for this model. With no override, LM Studio defaults this model to thinking off, unlike raw llama.cpp's template default of thinking on/xhigh. reasoning_effort:'none' explicitly disables it."
703
+ lmstudioOpenaiCompat: "Clio sends reasoning_effort over LM Studio's OpenAI-compatible HTTP port and omits chat_template_kwargs because LM Studio does not use that field for this model. The upstream template defaults to thinking on/xhigh; an LM Studio deployment may configure another default. reasoning_effort:'none' explicitly disables it. Through LiteLLM, this control also requires allowed_openai_params:['reasoning_effort']; Clio selects the LM Studio spelling only when every deployment of that alias explicitly declares model_info.runtime: lm-studio."
700
704
  lmstudioNative: "Clio's LM Studio runtime uses the OpenAI-compatible HTTP chat adapter with native REST model discovery and residency management."
701
705
  llamaCpp:
702
706
  ctxSize: 131072
@@ -718,7 +722,6 @@
718
722
  effortByLevel:
719
723
  low: low
720
724
  medium: medium
721
- high: xhigh
722
725
  xhigh: xhigh
723
726
  guidance: |
724
727
  Official Qwen3.8-27B chat template (chat_template.jinja) validates
@@ -726,19 +729,32 @@
726
729
  fatal jinja exception (hard 500, not a graceful fallback) on anything
727
730
  else, including minimal/high/max. The 27B card gives no /think or
728
731
  /no_think soft switch for this model; unlike Qwen3, that idiom has no
729
- effect here. Off must go through chat_template_kwargs.enable_thinking:
730
- false, never through reasoning_effort. preserve_thinking should stay
732
+ effect here. The template defaults to xhigh with thinking enabled.
733
+ The picker offers off/low/medium/xhigh; released high and max choices
734
+ normalize to xhigh rather than reaching the strict template unchanged.
735
+ Template-level off uses chat_template_kwargs.enable_thinking:false.
736
+ preserve_thinking should stay
731
737
  on for an agentic harness (the template's own recommendation for
732
738
  agent scenarios; improves KV cache reuse). llama.cpp's own
733
739
  reasoning_effort:"none" also disables reasoning at the server layer
734
740
  and is a separate mechanism from the template's enable_thinking flag;
735
- both were verified independently on llama.cpp and on LM Studio's HTTP
736
- port.
741
+ the none sentinel is consumed before strict template effort validation.
742
+ LM Studio's HTTP port also accepts none. A 2026-09-05 same-route
743
+ LiteLLM probe showed none and the template flag alone still reasoned
744
+ because the generic OpenAI route filtered the effort parameter. Adding
745
+ allowed_openai_params:["reasoning_effort"] produced zero reasoning
746
+ twice; an allowed low control still reasoned. Clio forwards only an
747
+ effort resolved from the model/runtime contract, and never infers the
748
+ upstream control runtime from a host name, port, or model alias.
737
749
 
738
750
  Thinking off is verified for this family, and no floor is imposed on
739
- it. JetBrains report that Qwen3.8 without reasoning gets stuck in a
751
+ it. The official card and template were rechecked at revision
752
+ 1d4bf0f2ff6012fd82039f2fa52739d0dd7c60c0:
753
+ https://huggingface.co/Qwen/Qwen3.8-27B/blob/1d4bf0f2ff6012fd82039f2fa52739d0dd7c60c0/chat_template.jinja
754
+ JetBrains report that Qwen3.8 without reasoning gets stuck in a
740
755
  loop repeating the same tool call indefinitely under Junie. Clio's
741
- shipped default is orchestrator.thinkingLevel: off, which on llama.cpp
756
+ shipped default is chat.thinkingLevel: off, an explicit Clio preference
757
+ rather than the vendor's template default. On llama.cpp it
742
758
  sends chat_template_kwargs.enable_thinking:false and on LM Studio's
743
759
  OpenAI-compatible port sends reasoning_effort:"none". Clio 0.3.8 ran
744
760
  that configuration on both runtimes (llama.cpp build b226-2115b73d8 on
@@ -746,8 +762,12 @@
746
762
  2026-08-29 and 2026-08-30: an interactive build-and-resume session, a
747
763
  three-worker fleet step, and a ten-task eval suite totalling 116 model
748
764
  calls, none of which repeated a tool call into a loop. The catalog
749
- therefore imposes no thinking floor and coerces no level for this
750
- family. If a future build or quantization regresses into the loop
765
+ therefore imposes no thinking floor for this family. A provisional 2026-09-03 LM Studio A/B on the heavily quantized
766
+ `qwen3.8-27b-gsq-rco` alias corroborated the switch: four thinking-off
767
+ scenarios reported zero reasoning tokens and no repeated-tool loop,
768
+ versus visible reasoning at the active level. The host had LM Link
769
+ enabled, so this records behavior only—not hardware throughput or a
770
+ controlled quality claim. If a future build or quantization regresses into the loop
751
771
  JetBrains describe, the guards that catch it are Clio's own, not a
752
772
  catalog setting: the tool-prose-loop detector
753
773
  (src/interactive/tool-prose-loop.ts), which is armed on the