@iowarp/clio-coder 0.4.2 → 0.4.3

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (369) hide show
  1. package/CHANGELOG.md +35 -0
  2. package/CONTRIBUTING.md +86 -19
  3. package/README.md +35 -6
  4. package/dist/{acp-TMDQZDIG.js → acp-H2NGRPWO.js} +11 -11
  5. package/dist/{agents-5N5NG3XG.js → agents-TL5LLUQP.js} +54 -53
  6. package/dist/assets/codewiki.json +1 -1
  7. package/dist/{auth-Z5CCBXKQ.js → auth-E5SW4HMS.js} +19 -16
  8. package/dist/{builtins-K6TNDT24.js → builtins-IA7V7FUC.js} +9 -4
  9. package/dist/{chunk-ZW4HH5JJ.js → chunk-2APPQIER.js} +6 -6
  10. package/dist/{chunk-72GZI5EV.js → chunk-2JH2WHGE.js} +2 -2
  11. package/dist/{chunk-UBRFI4HS.js → chunk-2UG5F4C5.js} +127 -47
  12. package/dist/{chunk-UH632ZYL.js → chunk-2UH2KFUP.js} +2 -2
  13. package/dist/{chunk-3F7VUY77.js → chunk-2VIKGWFZ.js} +2 -2
  14. package/dist/{chunk-I66EAJFY.js → chunk-2WZ546HR.js} +267 -232
  15. package/dist/{chunk-RKSR6VSF.js → chunk-4IUZQIJ3.js} +29 -1
  16. package/dist/{chunk-2X4RYJTJ.js → chunk-4UVU7BJ5.js} +2 -2
  17. package/dist/{chunk-M2DAX4F6.js → chunk-4WR7VSYB.js} +2 -2
  18. package/dist/{chunk-5KW52TEP.js → chunk-54CBCGIR.js} +5 -5
  19. package/dist/{chunk-3KIPBMUA.js → chunk-5ICU3EUH.js} +2 -2
  20. package/dist/{chunk-77QIVUZB.js → chunk-5MEZN6CB.js} +4 -4
  21. package/dist/{chunk-O42A54GG.js → chunk-5OIVVPHF.js} +2 -2
  22. package/dist/{chunk-LJID3DYZ.js → chunk-64I3JVYM.js} +2 -2
  23. package/dist/{chunk-4JDLP6ZS.js → chunk-6PTFB5VS.js} +7 -7
  24. package/dist/{chunk-VN3SHNBN.js → chunk-7DICMOS6.js} +2 -2
  25. package/dist/{chunk-HLAFFSEK.js → chunk-7DRAWPTZ.js} +2 -2
  26. package/dist/chunk-7E7I3WLS.js +3762 -0
  27. package/dist/{chunk-YJISEZKC.js → chunk-7ZYNNDKC.js} +6 -6
  28. package/dist/{chunk-I66ZTYNP.js → chunk-AF4YM7Z4.js} +236 -101
  29. package/dist/{chunk-2HFQNRV3.js → chunk-AX2THNSA.js} +12 -12
  30. package/dist/{chunk-PGF63K6I.js → chunk-B4OAX3SI.js} +65 -3
  31. package/dist/{chunk-W6NIE6OW.js → chunk-B4VEBZKF.js} +3 -3
  32. package/dist/{chunk-JBCS7CRR.js → chunk-BEPZRGGU.js} +10 -10
  33. package/dist/{chunk-XGDPUNND.js → chunk-CE5AX47J.js} +2 -2
  34. package/dist/{chunk-I64IFBLB.js → chunk-DWUOQKRU.js} +17 -10
  35. package/dist/{chunk-DZAW46HP.js → chunk-E3TPLWFX.js} +3 -3
  36. package/dist/{chunk-HIICAHCJ.js → chunk-EKCHAPYA.js} +2 -2
  37. package/dist/{chunk-XE3PCIXH.js → chunk-F5JHEYZM.js} +7 -7
  38. package/dist/{chunk-5PFYMY2V.js → chunk-FTMGRKEF.js} +2 -2
  39. package/dist/{chunk-ZNT2M6TG.js → chunk-G76U63X4.js} +17 -17
  40. package/dist/{chunk-34BHNEE3.js → chunk-GHS5EBTQ.js} +58 -7
  41. package/dist/{chunk-2NHR3NAY.js → chunk-GI7YYQ3F.js} +40 -34
  42. package/dist/{chunk-DYHAXKHD.js → chunk-GWZNEVM2.js} +12 -8
  43. package/dist/chunk-GYV6VZOC.js +26 -0
  44. package/dist/{chunk-MQXIVJ35.js → chunk-HAXOFFRH.js} +5 -5
  45. package/dist/{chunk-W4YEMFBX.js → chunk-HEQY7ZFI.js} +2 -2
  46. package/dist/{chunk-IKOZFYBN.js → chunk-I7ZPNEJM.js} +145 -102
  47. package/dist/{chunk-WNP7O5WZ.js → chunk-ID64D7PE.js} +4 -4
  48. package/dist/{chunk-XQRY4DTA.js → chunk-IGLP3ODT.js} +10 -10
  49. package/dist/chunk-IJNZMHLA.js +101 -0
  50. package/dist/{chunk-JWJGP5DQ.js → chunk-INY6HTFL.js} +7 -7
  51. package/dist/{chunk-PBP4B7XR.js → chunk-IUE3Y34X.js} +2 -2
  52. package/dist/{chunk-B74PXLU7.js → chunk-IWT4SF4R.js} +3 -3
  53. package/dist/{chunk-B7HM5Z7T.js → chunk-JDAY6FIL.js} +5 -5
  54. package/dist/{chunk-PJX3WQUQ.js → chunk-JEQ3XTHC.js} +2 -2
  55. package/dist/{chunk-FSP7CMNU.js → chunk-JGRC33J2.js} +50 -4
  56. package/dist/{chunk-X7IARSHT.js → chunk-JKKCYP3C.js} +9 -9
  57. package/dist/{chunk-HJWWJ6IL.js → chunk-JSC3U7TI.js} +16 -4
  58. package/dist/{chunk-SSEYRH53.js → chunk-KK4JZPBQ.js} +19 -140
  59. package/dist/{chunk-IHXBNWMM.js → chunk-KXDSS5WJ.js} +7 -3
  60. package/dist/{chunk-Q4XWMHX6.js → chunk-L47TF46W.js} +2 -2
  61. package/dist/{chunk-O3YUNJZ2.js → chunk-LDJG7DW3.js} +81 -24
  62. package/dist/{chunk-CDNVLKUX.js → chunk-LLDJM5XK.js} +13 -7
  63. package/dist/{chunk-F2I26BDK.js → chunk-MUW2BDDH.js} +4 -4
  64. package/dist/{chunk-HKMD33FO.js → chunk-MWUZBSAQ.js} +79 -76
  65. package/dist/{chunk-QQLGQY2A.js → chunk-N2Z7HLVY.js} +20 -20
  66. package/dist/{chunk-DZEK6CJN.js → chunk-NIQJ66N4.js} +19 -19
  67. package/dist/{chunk-CWVRRIEI.js → chunk-NZMNUPZZ.js} +2 -2
  68. package/dist/{chunk-462T4EGZ.js → chunk-O5CVSAG5.js} +2 -2
  69. package/dist/{chunk-TPEQIQIE.js → chunk-OML5D5V5.js} +8 -8
  70. package/dist/{chunk-IKSLQ4XV.js → chunk-PAJQJ7BS.js} +558 -216
  71. package/dist/{chunk-ZW55JB7N.js → chunk-PUVDKJ2Y.js} +2 -2
  72. package/dist/{chunk-UH347SHR.js → chunk-QWGDJJYJ.js} +11 -11
  73. package/dist/chunk-R6Q67RJH.js +134 -0
  74. package/dist/{chunk-CRFOIAX3.js → chunk-RRNP2ANY.js} +6 -6
  75. package/dist/{chunk-IDNA72AH.js → chunk-RSJ25QSL.js} +2 -2
  76. package/dist/chunk-SKHCAU7K.js +385 -0
  77. package/dist/{chunk-RLYRBIYQ.js → chunk-TM6LQDI3.js} +20 -12
  78. package/dist/chunk-UOIZ7DA4.js +41 -0
  79. package/dist/{chunk-P75RZCJW.js → chunk-UPZU6GE4.js} +3 -3
  80. package/dist/{chunk-MCMZMDAC.js → chunk-V2ANDPVT.js} +4 -4
  81. package/dist/{chunk-AK5XEFVZ.js → chunk-VA5FNYMT.js} +26 -13
  82. package/dist/{chunk-IMXMHHMQ.js → chunk-VW6DOEDG.js} +332 -57
  83. package/dist/{chunk-XOXV5GKE.js → chunk-W6RRQCPQ.js} +16 -7
  84. package/dist/{chunk-CYZW7JHJ.js → chunk-WBKFA554.js} +8 -8
  85. package/dist/{chunk-BO7Y52RY.js → chunk-WCXUNS7U.js} +7 -7
  86. package/dist/{chunk-ZGNYYXQ6.js → chunk-WRBAGUNF.js} +3 -3
  87. package/dist/{chunk-FVDGR2ZL.js → chunk-XIVNBFZS.js} +85 -30
  88. package/dist/{chunk-BYMNWQ7O.js → chunk-XPWWI35G.js} +299 -58
  89. package/dist/{chunk-KPXDY6QF.js → chunk-XRZT5WY5.js} +2 -2
  90. package/dist/{chunk-AZ4WMN4W.js → chunk-Y3CBHOR6.js} +2 -2
  91. package/dist/{chunk-VXMFAE2W.js → chunk-YPC6ZR5L.js} +19 -6
  92. package/dist/{chunk-54ODD65L.js → chunk-YQWYVTMC.js} +4 -4
  93. package/dist/{chunk-M2WXEHER.js → chunk-ZA4VCIGV.js} +2 -2
  94. package/dist/{chunk-7BHIY2MW.js → chunk-ZDN3Y73Y.js} +6 -6
  95. package/dist/{chunk-E7GT7O5N.js → chunk-ZWPRK62N.js} +7 -4
  96. package/dist/cli/index.js +38 -37
  97. package/dist/{clio-7VB377CC.js → clio-CMMK4KRR.js} +7 -7
  98. package/dist/{code-nav-YVLCYA7V.js → code-nav-MDZNQS33.js} +7 -7
  99. package/dist/{components-UBWCQSRW.js → components-UCUQ4QXW.js} +4 -4
  100. package/dist/{config-4HVOS65E.js → config-SVM5P5YI.js} +76 -74
  101. package/dist/{configure-PIWO7B24.js → configure-LE3IK2TJ.js} +26 -24
  102. package/dist/{context-IYEHL3WQ.js → context-2OHRKS42.js} +66 -63
  103. package/dist/{context-N6ZE3LGJ.js → context-E3VC7RX5.js} +15 -11
  104. package/dist/{context-KQYIWPWT.js → context-VNCR7KAG.js} +60 -45
  105. package/dist/{context-clear-G4OGZJDS.js → context-clear-BW4O37TG.js} +61 -59
  106. package/dist/context-map-COB37XXN.js +505 -0
  107. package/dist/{context-working-set-BWLF6LJP.js → context-working-set-VDS25HXZ.js} +17 -16
  108. package/dist/{dispatch-runner-2QQAITS3.js → dispatch-runner-5AHT53RF.js} +85 -74
  109. package/dist/{doctor-LHBD36VU.js → doctor-WNNVO6FY.js} +37 -37
  110. package/dist/{eval-C45FYRJ6.js → eval-7G7SGAYO.js} +285 -114
  111. package/dist/{eval-inventory-6DEJPLBF.js → eval-inventory-Y6QRFOH5.js} +4 -4
  112. package/dist/{evidence-6SHONYAF.js → evidence-VD6736FQ.js} +63 -62
  113. package/dist/{evolve-KRKMV72X.js → evolve-AL3NGVRL.js} +62 -61
  114. package/dist/{extensions-KPZ2UHBB.js → extensions-MOVJ32NM.js} +7 -7
  115. package/dist/{fleet-IVTCKDHT.js → fleet-QZHUMAGI.js} +110 -108
  116. package/dist/{fleet-commands-EDWL3IT7.js → fleet-commands-BAYT5FJZ.js} +10 -10
  117. package/dist/{fleet-decisions-YP3YEFGK.js → fleet-decisions-IREVMRU4.js} +7 -6
  118. package/dist/{fleet-graph-ZFWKHY2M.js → fleet-graph-YCTT3HTI.js} +19 -18
  119. package/dist/{fleet-inspect-FVUNCBML.js → fleet-inspect-QVJTDAVB.js} +55 -54
  120. package/dist/{fleet-preflight-UN5XED4R.js → fleet-preflight-25QAFPK4.js} +4 -4
  121. package/dist/{fleet-validate-XOWC4HSX.js → fleet-validate-5O57AAJ7.js} +23 -22
  122. package/dist/{fleet-verify-UN3SODEL.js → fleet-verify-CPH2W2T6.js} +56 -55
  123. package/dist/{fleet-view-TWHJKCN6.js → fleet-view-SWBR3VGQ.js} +55 -54
  124. package/dist/{init-T2QORQ3Y.js → init-J477LKZH.js} +78 -76
  125. package/dist/{interop-IN5I2A66.js → interop-3FCM6XLG.js} +11 -11
  126. package/dist/{library-LSCATDLZ.js → library-QUQEIUG6.js} +28 -27
  127. package/dist/{memory-HYOKAGGJ.js → memory-SGGSEP65.js} +64 -63
  128. package/dist/{models-2GPMFYCM.js → models-HEKUAXXK.js} +49 -43
  129. package/dist/{monitor-E4ASVUJH.js → monitor-HKU57TYQ.js} +61 -60
  130. package/dist/{orchestrator-DDMPR3PY.js → orchestrator-VDFAEFAI.js} +919 -546
  131. package/dist/{panes-E3RUXOW5.js → panes-DN2SSFOH.js} +3 -3
  132. package/dist/{panes-IXKLOKA2.js → panes-TALGNPZT.js} +8 -8
  133. package/dist/{paths-L7LGY6RN.js → paths-NBMFAIEZ.js} +5 -5
  134. package/dist/reset-EAJFFJVB.js +344 -0
  135. package/dist/{resources-OTRSN34L.js → resources-OVKSEFVE.js} +27 -20
  136. package/dist/{run-5DEYH5QK.js → run-7DP7ZF2J.js} +113 -109
  137. package/dist/{share-IHWTLO3M.js → share-WML67FT3.js} +26 -25
  138. package/dist/{skills-IYMXMKW4.js → skills-SG662R2K.js} +39 -31
  139. package/dist/{skills-eval-DROHSJAR.js → skills-eval-VVZEUU46.js} +74 -73
  140. package/dist/{skills-inventory-D7X4L4ZX.js → skills-inventory-I2E23GET.js} +21 -20
  141. package/dist/{slash-commands-QBM7UZ3B.js → slash-commands-S7MBJDQK.js} +35 -34
  142. package/dist/{steer-Z5DO23FJ.js → steer-2LQOMCPB.js} +3 -3
  143. package/dist/{support-U7QOWY26.js → support-CC2UJBJ6.js} +6 -6
  144. package/dist/{targets-P2FUC4IL.js → targets-4QC3HIEW.js} +48 -45
  145. package/dist/{terminal-lease-YREJ3JX2.js → terminal-lease-TUHIJ6Y2.js} +2 -2
  146. package/dist/{tools-5B7RO6MV.js → tools-TFGJICCU.js} +8 -8
  147. package/dist/{trace-YMGMUM6A.js → trace-FXMXUZUF.js} +7 -7
  148. package/dist/uninstall-5PEVOE5B.js +408 -0
  149. package/dist/upgrade-M4WXY6KN.js +303 -0
  150. package/dist/{usage-ME5MPXGX.js → usage-N7ZNVLEM.js} +147 -102
  151. package/dist/{verifiers-BVZ7IWOO.js → verifiers-DJTP4XX6.js} +15 -15
  152. package/dist/{verify-5K7ZKQFC.js → verify-RWE4PPEK.js} +9 -9
  153. package/dist/{wiki-generate-F5W5QTYY.js → wiki-generate-C7IQOXSP.js} +84 -82
  154. package/dist/{with-panes-BYOJCLAM.js → with-panes-4GCGSL7J.js} +9 -9
  155. package/dist/worker/entry.js +61 -60
  156. package/docs/architecture/artifact-placement.md +1 -0
  157. package/docs/architecture/artifact-versions.md +1 -1
  158. package/docs/architecture/context-engine.md +4 -0
  159. package/docs/architecture/middleware-and-components.md +1 -1
  160. package/docs/architecture/model-catalog.md +21 -10
  161. package/docs/architecture/observability.md +12 -1
  162. package/docs/architecture/prompt-envelope-and-tools.md +2 -0
  163. package/docs/architecture/provider-adapter-cookbook.md +63 -0
  164. package/docs/architecture/safety-model.md +15 -5
  165. package/docs/guide/built-in-agents.md +17 -3
  166. package/docs/guide/commands-and-modes.md +1 -1
  167. package/docs/guide/configuration-and-targets.md +97 -9
  168. package/docs/guide/configuration-reference.md +7 -2
  169. package/docs/guide/environment-variables.md +2 -0
  170. package/docs/guide/installation-and-lifecycle.md +37 -4
  171. package/docs/guide/proactive-memory.md +66 -55
  172. package/docs/guide/skills-marketplace.md +18 -0
  173. package/docs/process/development-pipeline.md +34 -1
  174. package/docs/process/eval-runner.md +67 -3
  175. package/evals/behavioral-model.yaml +3 -2
  176. package/package.json +2 -2
  177. package/skills/README.md +7 -5
  178. package/skills/coding/ast-grep/SKILL.md +101 -30
  179. package/skills/coding/ast-grep/evals.md +26 -0
  180. package/skills/coding/coding-standards/SKILL.md +40 -5
  181. package/skills/coding/coding-standards/evals.md +23 -0
  182. package/skills/coding/prototype/SKILL.md +87 -28
  183. package/skills/coding/prototype/evals.md +19 -0
  184. package/skills/coding/tdd/SKILL.md +80 -53
  185. package/skills/coding/tdd/evals.md +20 -0
  186. package/skills/context/context-handoff/SKILL.md +43 -2
  187. package/skills/context/context-handoff/evals.md +44 -0
  188. package/skills/context/context-prime/SKILL.md +45 -15
  189. package/skills/context/context-prime/evals.md +45 -0
  190. package/skills/git/branch-closeout/SKILL.md +132 -0
  191. package/skills/git/branch-closeout/evals.md +133 -0
  192. package/skills/git/branch-closeout/references/closeout-checklist.md +81 -0
  193. package/skills/git/file-ticket/SKILL.md +77 -63
  194. package/skills/git/file-ticket/assets/issue-template.md +22 -0
  195. package/skills/git/file-ticket/evals.md +31 -26
  196. package/skills/git/file-ticket/references/issue-discovery.md +49 -0
  197. package/skills/git/fix-issue/SKILL.md +87 -64
  198. package/skills/git/fix-issue/evals.md +35 -31
  199. package/skills/git/fix-issue/references/diagnosis-and-rca.md +46 -0
  200. package/skills/git/resolve-merge-conflicts/SKILL.md +100 -51
  201. package/skills/git/resolve-merge-conflicts/evals.md +52 -25
  202. package/skills/git/resolve-merge-conflicts/references/conflict-matrix.md +126 -0
  203. package/skills/git/ship/SKILL.md +103 -67
  204. package/skills/git/ship/assets/pr-template.md +21 -0
  205. package/skills/git/ship/evals.md +44 -28
  206. package/skills/git/ship/references/remote-and-branch-policy.md +62 -0
  207. package/skills/git/worktree-create/SKILL.md +80 -50
  208. package/skills/git/worktree-create/evals.md +40 -33
  209. package/skills/git/worktree-create/references/worktree-setup.md +62 -66
  210. package/skills/git/worktree-merge/SKILL.md +112 -65
  211. package/skills/git/worktree-merge/evals.md +42 -34
  212. package/skills/git/worktree-merge/references/merge-strategies.md +52 -0
  213. package/skills/planning/archify/SKILL.md +196 -0
  214. package/skills/planning/archify/evals.md +65 -0
  215. package/skills/planning/architecture/SKILL.md +61 -12
  216. package/skills/planning/architecture/evals.md +65 -0
  217. package/skills/planning/backlog/SKILL.md +130 -14
  218. package/skills/planning/backlog/evals.md +142 -0
  219. package/skills/planning/prd/SKILL.md +46 -6
  220. package/skills/planning/prd/evals.md +54 -0
  221. package/skills/planning/product-intent/SKILL.md +57 -2
  222. package/skills/planning/product-intent/evals.md +70 -0
  223. package/skills/planning/tech-spec/SKILL.md +53 -2
  224. package/skills/planning/tech-spec/evals.md +73 -0
  225. package/skills/registry.yaml +58 -50
  226. package/skills/remote.yaml +13 -0
  227. package/skills/research/arxiv-literature/SKILL.md +76 -18
  228. package/skills/research/arxiv-literature/evals.md +50 -0
  229. package/skills/research/experiment-protocol/SKILL.md +20 -1
  230. package/skills/research/experiment-protocol/evals.md +23 -0
  231. package/skills/research/scientific-debugging/SKILL.md +23 -1
  232. package/skills/research/scientific-debugging/evals.md +18 -0
  233. package/skills/research/scientific-modernization/SKILL.md +26 -1
  234. package/skills/research/scientific-modernization/evals.md +27 -0
  235. package/skills/skill-marketplace.json +63 -28
  236. package/skills/workflow/cut-it/SKILL.md +65 -5
  237. package/skills/workflow/cut-it/evals.md +101 -0
  238. package/skills/workflow/design-council/SKILL.md +117 -27
  239. package/skills/workflow/design-council/evals.md +161 -0
  240. package/skills/workflow/grill-me/SKILL.md +86 -10
  241. package/skills/workflow/grill-me/evals.md +153 -0
  242. package/skills/workflow/workflow-distiller/SKILL.md +76 -17
  243. package/skills/workflow/workflow-distiller/evals.md +118 -0
  244. package/src/cli/configure-interop.ts +105 -13
  245. package/src/cli/configure-oauth.ts +57 -0
  246. package/src/cli/configure-onboarding.ts +980 -0
  247. package/src/cli/configure-target.ts +594 -0
  248. package/src/cli/configure.ts +1082 -528
  249. package/src/cli/context-map.ts +114 -0
  250. package/src/cli/context.ts +4 -0
  251. package/src/cli/index.ts +1 -0
  252. package/src/cli/lifecycle-presenter.ts +436 -0
  253. package/src/cli/models.ts +10 -2
  254. package/src/cli/modes/print.ts +5 -1
  255. package/src/cli/reset.ts +228 -106
  256. package/src/cli/run.ts +7 -2
  257. package/src/cli/select.ts +664 -0
  258. package/src/cli/skills.ts +9 -2
  259. package/src/cli/targets.ts +3 -0
  260. package/src/cli/uninstall.ts +233 -165
  261. package/src/cli/upgrade.ts +204 -149
  262. package/src/cli/usage.ts +86 -27
  263. package/src/cli/validate-model.ts +3 -3
  264. package/src/core/config.ts +56 -0
  265. package/src/core/external-diagnostic.ts +44 -0
  266. package/src/core/gateway-routing.ts +157 -0
  267. package/src/core/safe-exec.ts +17 -2
  268. package/src/core/skill-activation.ts +89 -2
  269. package/src/domains/agents/builtins/world-knowledge.md +31 -0
  270. package/src/domains/agents/catalog.ts +1 -1
  271. package/src/domains/agents/result-contract.ts +70 -0
  272. package/src/domains/context/wiki/map-seed.ts +589 -0
  273. package/src/domains/context/wiki/plan.ts +2 -2
  274. package/src/domains/dispatch/admission.ts +29 -0
  275. package/src/domains/dispatch/agent-candidates.ts +10 -0
  276. package/src/domains/dispatch/budget-envelope.ts +86 -1
  277. package/src/domains/dispatch/capability-match.ts +1 -0
  278. package/src/domains/dispatch/capacity-lease.ts +17 -0
  279. package/src/domains/dispatch/contract.ts +11 -1
  280. package/src/domains/dispatch/extension.ts +134 -29
  281. package/src/domains/dispatch/types.ts +3 -0
  282. package/src/domains/dispatch/worker-model-metadata.ts +38 -0
  283. package/src/domains/eval/metrics/call-ledger-stream.ts +34 -11
  284. package/src/domains/eval/metrics/token-stream.ts +201 -31
  285. package/src/domains/eval/metrics/tracked.ts +40 -4
  286. package/src/domains/eval/runners/clio-run.ts +5 -2
  287. package/src/domains/eval/schema/suite.ts +28 -0
  288. package/src/domains/eval/schema/verdict.ts +2 -2
  289. package/src/domains/eval/suites/resolve.ts +13 -1
  290. package/src/domains/eval/suites/run.ts +24 -3
  291. package/src/domains/interop/registry.ts +6 -2
  292. package/src/domains/interop/types.ts +4 -0
  293. package/src/domains/lifecycle/migrations/index.ts +4 -0
  294. package/src/domains/memory/task-memory-policy.ts +70 -26
  295. package/src/domains/memory/task-memory-telemetry.ts +1 -0
  296. package/src/domains/middleware/index.ts +0 -1
  297. package/src/domains/middleware/marketplace-offer.ts +3 -35
  298. package/src/domains/middleware/memory-intervention.ts +127 -32
  299. package/src/domains/middleware/memory-step-endpoint.ts +3 -2
  300. package/src/domains/middleware/skills-reminder.ts +31 -2
  301. package/src/domains/observability/compaction-usage.ts +118 -0
  302. package/src/domains/observability/cost.ts +1 -1
  303. package/src/domains/observability/extension.ts +6 -1
  304. package/src/domains/observability/out-of-turn-usage.ts +52 -21
  305. package/src/domains/providers/contract.ts +4 -1
  306. package/src/domains/providers/extension.ts +40 -9
  307. package/src/domains/providers/model-capabilities.ts +9 -0
  308. package/src/domains/providers/model-discovery.ts +2 -0
  309. package/src/domains/providers/model-runtime-capabilities.ts +15 -5
  310. package/src/domains/providers/models/local-models/clio-coder-local-coding-targets.yaml +32 -12
  311. package/src/domains/providers/runtimes/antigravity/antigravity-code.ts +225 -45
  312. package/src/domains/providers/runtimes/common/lmstudio-http.ts +6 -2
  313. package/src/domains/providers/runtimes/common/local-synth.ts +2 -0
  314. package/src/domains/providers/runtimes/protocol/litellm.ts +119 -29
  315. package/src/domains/providers/support.ts +11 -5
  316. package/src/domains/providers/target-model-cache.ts +25 -2
  317. package/src/domains/providers/types/capability-flags.ts +2 -0
  318. package/src/domains/providers/types/runtime-descriptor.ts +20 -1
  319. package/src/domains/providers/types/target-descriptor.ts +19 -0
  320. package/src/domains/resources/index.ts +3 -0
  321. package/src/domains/resources/skills/install.ts +72 -7
  322. package/src/domains/resources/skills/loader.ts +7 -0
  323. package/src/domains/resources/skills/marketplace.ts +63 -11
  324. package/src/domains/safety/autonomy.ts +15 -0
  325. package/src/domains/safety/index.ts +1 -0
  326. package/src/domains/safety/path-policy.ts +1 -1
  327. package/src/domains/safety/policy-engine.ts +34 -11
  328. package/src/domains/safety/protected-artifacts.ts +191 -88
  329. package/src/domains/safety/run-effects.ts +2 -22
  330. package/src/domains/safety/skill-authority.ts +55 -0
  331. package/src/domains/session/compaction/compact.ts +72 -22
  332. package/src/domains/session/entries.ts +6 -0
  333. package/src/domains/session/usage.ts +3 -3
  334. package/src/engine/agent.ts +13 -3
  335. package/src/engine/ai.ts +26 -8
  336. package/src/engine/antigravity/subprocess-runtime.ts +386 -120
  337. package/src/engine/api-registry.ts +3 -0
  338. package/src/engine/apis/openai-completions.ts +117 -14
  339. package/src/engine/external-subprocess.ts +114 -6
  340. package/src/entry/background-model-metadata.ts +18 -0
  341. package/src/entry/compaction-prompt.ts +57 -0
  342. package/src/entry/orchestrator.ts +405 -216
  343. package/src/entry/task-memory-lifecycle.ts +35 -0
  344. package/src/interactive/chat-loop-messages.ts +13 -4
  345. package/src/interactive/chat-loop.ts +65 -2
  346. package/src/interactive/chat-renderer.ts +1 -0
  347. package/src/interactive/cost-overlay.ts +26 -2
  348. package/src/interactive/interactive-slash-runtime.ts +2 -1
  349. package/src/interactive/renderers/worker-entry.ts +32 -0
  350. package/src/interactive/slash-commands.ts +24 -6
  351. package/src/interactive/theme/labels.ts +19 -13
  352. package/src/interactive/turn-context.ts +9 -5
  353. package/src/interactive/turn-recovery.ts +8 -0
  354. package/src/interactive/turn-runtime.ts +27 -11
  355. package/src/interactive/turn-state.ts +7 -0
  356. package/src/interactive/worker-receipts.ts +1 -0
  357. package/src/interactive/worker-stream.ts +6 -1
  358. package/src/tools/context/index.ts +30 -9
  359. package/src/tools/dispatch-arguments.ts +1 -0
  360. package/src/tools/dispatch-event-text.ts +10 -0
  361. package/src/tools/dispatch-plan.ts +1 -0
  362. package/src/tools/dispatch-runner.ts +12 -0
  363. package/src/tools/registry.ts +11 -5
  364. package/src/tools/worker-evidence.ts +3 -1
  365. package/src/worker/spec-contract.ts +4 -0
  366. package/dist/chunk-2Z2IKEXI.js +0 -1554
  367. package/dist/reset-OAQP3W4O.js +0 -230
  368. package/dist/uninstall-N34PCTGJ.js +0 -331
  369. package/dist/upgrade-PXK3S2YM.js +0 -325
@@ -20,7 +20,11 @@ Configured `wireModels` and a target `defaultModel` remain selectable before a
20
20
  live catalog is known; Clio labels those rows as `configured` or `default`.
21
21
  Once a target returns a live catalog, that catalog is authoritative and models
22
22
  the runtime no longer reports stop resolving. Live probe discoveries are labeled
23
- `live` and carry load-state metadata when the runtime exposes it. This preserves
23
+ `live` and carry load-state metadata when the runtime exposes it. Runtime model
24
+ labels are separate metadata: the stable slug remains the wire identity while
25
+ `clio-coder models` and target status may show the human label beside it. Slugs,
26
+ labels, source, and freshness round-trip through the generic target model
27
+ snapshot; a cached label never replaces a live slug. This preserves
24
28
  operator-curated defaults while still letting runtime discovery take over after
25
29
  newly installed local models or newly entitled cloud models appear. Catalog YAML
26
30
  entries are loaded when the provider domain is built, so bundled or overlay
@@ -32,8 +36,9 @@ Live provider probes are the preferred source for loaded context and per-model m
32
36
  `probeCapabilitiesForModel` is the one exact-id selector. When a router serves several models, capability resolution queries `probeCapabilitiesForModel` to ensure probe data is extracted only from the `/v1/models` row keyed to its own exact wire model ID.
33
37
 
34
38
  Transient probe failures preserve the last-good catalog, load states,
35
- capabilities, and notes for the same target identity, but the target health is
36
- reported as down or unavailable with the probe error as the reason. Worker
39
+ labels, capabilities, and notes for the same target identity, but those model
40
+ rows are marked cached/stale and target health is reported as down or unavailable
41
+ with the probe error as the reason. Worker
37
42
  dispatch canonicalizes requested model ids against the live catalog when one is
38
43
  available, so a short alias can resolve to the canonical live id before the
39
44
  worker spec and receipt are written.
@@ -185,14 +190,19 @@ override the applicable setting.
185
190
  - **Ollama Native (`ollama-native`):** Ollama utilizes the native `thinking` field in the request and response payloads. The engine handles Ollama-specific effort levels and streams reasoning increments cleanly through the native thinking channel.
186
191
  - **LM Studio (`lmstudio`):** Chat uses the OpenAI-compatible `/v1/chat/completions` surface, including its `reasoning` stream field. Clio controls thinking only with `reasoning_effort` and never sends `chat_template_kwargs` to LM Studio. See <https://lmstudio.ai/docs/developer/openai-compat/chat-completions>.
187
192
  - **LiteLLM (`litellm`):** This is a gateway runtime, not an `openai-compat`
188
- alias. Discovery checks `/health/liveliness`, reads aliases and capability
193
+ alias. Discovery checks `/health/liveliness`, reads routed names and capability
189
194
  metadata from `/v1/model/info`, and records the physical deployment reported
190
- by `x-litellm-*` response headers. Defaults stay conservative when metadata is
191
- absent: tools, vision, and reasoning are not inferred. The runtime advertises
192
- standard `json_schema` structured output and treats schema-plus-tools as a
193
- conflict because the gateway alias does not identify one stable upstream.
194
- Residency is observe-only because LiteLLM owns loading and eviction behind the
195
- alias.
195
+ by `x-litellm-*` response headers. Deterministic gateways should publish one
196
+ `node/model` name per deployment; genuine multi-deployment aliases expose only
197
+ the capabilities guaranteed by every route and use the smallest unanimously
198
+ published context/output limits. Defaults stay conservative when metadata is
199
+ absent: tools, vision, reasoning, and structured output are not inferred.
200
+ Explicitly advertised schema support uses standard `json_schema` on the wire.
201
+ Gateway requests use no hidden OpenAI SDK retries, and LiteLLM failures bypass
202
+ Clio's interactive transient retry ladder so the operator can select another
203
+ route. Stable session ids, request tags, optional request-level timeouts, and
204
+ observed server retry/fallback headers remain supported. Residency is
205
+ observe-only because LiteLLM owns loading and eviction behind the route.
196
206
  - **OpenAI Completions (`openai-completions`):** The OpenAI-compatible completions provider preserves reasoning blocks within assistant messages. It replays thinking blocks via the `reasoning_content` parameter in the message history, ensuring that the model maintains its chain-of-thought across conversational turns without stripping the data.
197
207
  - **Anthropic OAuth / API (`anthropic-max`):** Uses the `anthropic-extended` thinking format. The engine supports Anthropic's native extended thinking block protocol, streaming thinking increments and outputting them wrapped appropriately or natively depending on target capabilities.
198
208
  - **Reasoning-Never Models (`thinking.mechanism: none`):** When a model is configured or cataloged with `thinking.mechanism: none`, it is treated as a reasoning-never model. For these models, Clio must not send any thinking fields or parameters in requests, must not replay thinking blocks, must not surface thinking events to the TUI, and must not preserve or log reasoning token usage in metrics.
@@ -206,6 +216,7 @@ Subscription models are registered and managed as standard HTTP/cloud targets:
206
216
  - **`openai-codex` (ChatGPT Plus/Pro OAuth):** Maps to catalog-backed Codex model ids surfaced by `clio-coder configure --list` and `clio-coder models` via a browser-minted subscription OAuth token, supporting complete chat, vision, and tool-use capabilities.
207
217
  - **`anthropic-max` (Claude Pro/Max OAuth):** Powers chat and workers using catalog-backed Claude model ids surfaced by `clio-coder configure --list` and `clio-coder models`. It relies on the engine's Anthropic OAuth provider. During auth initialization, it alerts the operator to usage-terms caveat via:
208
218
  `Connects with your Claude Pro/Max subscription via OAuth (the same path Claude Code uses). Using subscription credentials outside Anthropic's first-party apps may not align with their terms of service; enable at your own discretion.`
219
+ - **`antigravity-code` (experimental local delegation):** Is not an HTTP model provider and is never orchestrator-eligible. It invokes the operator's own authenticated official `agy` executable only for dispatch work, consumes structured `stream-json` results and token accounting, and discovers model slugs and labels from the non-generating JSON `models` command. Descriptor models are cold-start hints only; a successful target probe is authoritative for that account, including the disappearance of a former model.
209
220
 
210
221
  ---
211
222
 
@@ -151,7 +151,18 @@ A row has the following schema:
151
151
 
152
152
  `repoIdentity` is the same cwd hash the session ledger is filed under, which is what lets `usage report --repo <path>` select these rows with the hash it already computes for the ledgers.
153
153
 
154
- `label` is one of `side-question`, `handoff`, `prewarm`, or `background-memory`. The last two joined for the same reason as the first two: a prompt pre-warm and a proactive-memory step are provider calls the operator did not ask for and would otherwise never see, and neither appends anything to the session JSONL. A row may also carry `timing { durationMs }` and a `promptCache` block built from the backend's own prefill facts when the server reported them; a backend that reports no timings simply omits the block, as LM Studio's OpenAI-compatible port does.
154
+ `label` is one of `side-question`, `handoff`, `prewarm`, `background-memory`, or `failed-compaction`. Prompt pre-warm and proactive-memory calls are recorded here because neither appends an assistant call to the session JSONL; failed-compaction records preserve calls from an attempt that produced no checkpoint. A row may also carry `timing { durationMs }` and a `promptCache` block built from the backend's own prefill facts when the server reported them; a backend that reports no timings simply omits the block, as LM Studio's OpenAI-compatible port does.
155
+
156
+
157
+ Failed compaction attempts record one `failed-compaction` row per invoked summary stream when no checkpoint is produced. `callOutcome` distinguishes a completed first stream (`success`) from an `error` or `aborted` stream; a completed call can belong to an unsuccessful split-compaction attempt. The rows capture the originating session/repository and selected target/model before asynchronous work can switch context. Successful compactions keep usage solely on their checkpoint, and their live accounting uses the same selected route. Unset model controls retain the active chat route.
158
+
159
+ New failed-compaction rows preserve missing usage fields as `null`. Positive partial-response facts survive an error that resets missing fields to zero. Ambiguous failed zeros remain unknown, reasoning is separate from ordinary output/total tokens, and an absent total is not inferred. Positive adapter prices are labeled `estimated`; zero or missing pricing is `unknown`, not a free-call claim. Existing numeric rows and historical checkpoints remain readable without rewriting them.
160
+
161
+ `clio-coder usage report` includes these calls in its known subtotals, labels the failed-attempt count, and exposes `failedCompaction.knownUsage`, `erroredKnownUsage`, and per-field `unobservedUsageCalls` in the token and model JSON facts. A field with missing coverage and no known positive amount is `null`, including cost-only or wholly unobserved failures. Text output identifies incomplete subtotals. The live `/cost` view records positive known contributions under a failed-compaction label; its numeric token counters remain known subtotals. These figures do not certify provider billing or complete spending. The existing session-cost ceiling checks the numeric known sum, so unreported cost does not become an enforced complete-cost bound.
162
+
163
+ Eval tracked/stdout folds do not include the failed-compaction sidecar, and live `/cost` reseeding reads the session ledger rather than this store. A later usage report can therefore include retained failed-compaction amounts that those views omit. This consumer reconciliation is deferred to v0.4.5 or later; the retained known amounts must not be presented as complete cross-surface billing.
164
+
165
+ A failed or empty summary produces no checkpoint. Required failed-compaction usage appends are flushed and a write failure remains an explicit operation error; no model call is repeated to repair accounting. If checkpoint append throws, Clio checks that checkpoint's exact identity in the original ledger before choosing the sidecar: an already written checkpoint is not counted again, and proven absence permits the sidecar. An unreadable or malformed ledger that leaves persistence ambiguous fails visibly without a speculative duplicate write. This is the existing bounded usage store, not a new recovery store; its 1000-row retention and unknown telemetry limits still apply.
155
166
 
156
167
  ---
157
168
 
@@ -201,6 +201,8 @@ Compaction rewrites history, so the next turn on a local single-slot backend is
201
201
 
202
202
  Timing and cache behavior are persisted per API call, so a finished session can be inspected from its stored artifacts alone. Each assistant entry in the session ledger (`current.jsonl`, under the directory reported by `clio-coder paths`) carries `timing { ttftMs, apiMs }` and `promptCache { input, cacheRead, cacheWrite, backendVerdict }`, and the run's first persisted call also carries `expectedColdReasons`. Cache verdicts are `hot`, `partial`, `cold`, or `small`.
203
203
 
204
+ Native session timing uses a monotonic clock from each stream invocation, before the provider's response-header wait, to its first observed output (`ttftMs`) and completion (`apiMs`). A tool-loop continuation starts a new clock; no output leaves TTFT null, and a genuine rounded zero remains zero. Historical values are not rewritten and may omit the pre-header wait. Eval prefers these durable native call records. Its stdout-only fallback starts at the provider's `message_start` event, which can arrive after headers, so fallback spans are not complete request latency and must not be compared as equivalent measurements.
205
+
204
206
  For aggregate cost and token facts across sessions, use `clio-coder usage report --days <n>`. Inside the TUI, `/cost` shows session totals and `/context` opens the context-window ledger overlay.
205
207
 
206
208
  ## Self-documentation retrieval
@@ -170,6 +170,69 @@ inherited from the pinned Pi dependency; Clio no longer carries a separate
170
170
  contract test that reconstructs Pi's whole adaptive or budget payload.
171
171
 
172
172
 
173
+
174
+ ### 3.3 Thinking controls through LiteLLM
175
+
176
+ Dedicated memory and compaction roles read a cold LiteLLM target's metadata
177
+ before synthesizing the completion model. The read disables reasoning probes;
178
+ it does not run an extra inference request. Each selected target owns its probe
179
+ state, even when two targets share a gateway URL. A successful unknown or mixed
180
+ runtime declaration remains unknown and is not repeatedly probed for a preferred
181
+ answer. Memory includes discovery and auth in its existing generation/deadline
182
+ boundary; cancellation cannot launch a later completion or mark the endpoint
183
+ down. Fresh discovered output limits also bound its request. After metadata and auth,
184
+ memory rechecks actual endpoint occupancy immediately before registering its
185
+ inference hold. Late saturation stays a dropped `endpoint_busy` boundary with
186
+ no usage or cache-disturbance claim. Compaction checks
187
+ the originating session/branch after preparation and uses the existing simple
188
+ stream API with thinking off, rather than inferring an active level from the
189
+ model's reasoning capability.
190
+
191
+ Native worker admission also prepares the selected cold LiteLLM target before
192
+ freezing its capabilities and thinking controls into the worker specification
193
+ and receipt. Tool cancellation and the original admission deadline bound that
194
+ wait, including a delayed metadata response body. Failed preparation releases
195
+ the existing plan reservation and cannot launch a late worker or publish
196
+ cancelled health data. The approved target, model, endpoint and node remain
197
+ binding; changed route identity requires fresh admission. Metadata discovery
198
+ does not infer capacity or residency from another route sharing the gateway,
199
+ and does not bypass the existing capacity or approved tool-surface checks.
200
+
201
+
202
+ A gateway alias is not an upstream runtime identity. Clio consumes the optional
203
+ `model_info.runtime` deployment declaration from LiteLLM's `/v1/model/info` only
204
+ when every deployment of the alias names the same recognized control runtime:
205
+ `lm-studio` or `llama.cpp`. Missing, unknown, or mixed declarations produce no
206
+ runtime-specific control hint. The probe-only `thinkingControlRuntime` capability
207
+ travels through the existing main, background and worker model capability path;
208
+ `runtimeId`, authentication, the gateway URL and residency ownership stay LiteLLM.
209
+ Clio never infers this declaration from ports or model names and never loads or
210
+ unloads the gateway's upstream models.
211
+
212
+ The family still determines whether thinking is switchable and which active
213
+ levels exist. A declared LM Studio route receives `reasoning_effort: "none"` for
214
+ an effective off choice; a llama.cpp route uses its template switch. LiteLLM's
215
+ generic OpenAI adapter may silently filter a resolved effort for local model
216
+ names. Clio therefore adds `allowed_openai_params: ["reasoning_effort"]` only when
217
+ it sends that model/runtime's resolved `reasoning_effort`; unrelated parameters
218
+ and unknown off mechanisms are not newly allowed. This is a request control,
219
+ not a change to gateway configuration. See [LiteLLM parameter forwarding](https://docs.litellm.ai/docs/completion/drop_params).
220
+
221
+ [Qwen3.8-27B's pinned template](https://huggingface.co/Qwen/Qwen3.8-27B/blob/1d4bf0f2ff6012fd82039f2fa52739d0dd7c60c0/chat_template.jinja)
222
+ accepts active `low`, `medium`, and `xhigh`, with thinking enabled and `xhigh`
223
+ when the template receives no override. Clio shows off/low/medium/xhigh and maps
224
+ released high/max selections to xhigh. That vendor default does not overwrite
225
+ Clio's explicit chat preference or saved low setting. Disabling thinking does
226
+ not remove historical reasoning or discard reasoning a server actually returns.
227
+
228
+ Controlled gateway filtering tests cover discovery, capability transfer, actual
229
+ HTTP payloads, off/on switching, and built CLI persistence. Live same-route
230
+ probes on the selected LiteLLM deployment returned reasoning with the flag or
231
+ none alone, zero reasoning on two calls with none plus the explicit allowance,
232
+ and positive reasoning for an allowed low control. This verifies the measured
233
+ route and request contract; unknown or heterogeneous gateway aliases need their
234
+ own declared capabilities and acceptance.
235
+
173
236
  ---
174
237
 
175
238
  ## 4. Configuring Reasoning & Thinking Formats
@@ -129,11 +129,12 @@ The `/view` workspace category treats a recorded successful write as a durable f
129
129
 
130
130
  A `SKILL.md` may declare `allowed-tools` and `disallowed-tools`. The declaration is enforced at tool admission, between the safety net and the autonomy mapping, on every surface that activates skills (interactive turns, headless `clio-coder run` turns, and dispatched workers whose recipes declare skills).
131
131
 
132
- - **Window.** Narrowing arms when `context` (scope="skills") successfully loads the skill and lasts for the lifetime of the pending-skill policy: to the end of the current turn for the main agent, and to the end of the run for a worker. A later turn is unrestricted until a skill is requested and loaded again.
132
+ - **Window.** Narrowing arms when `context` (scope="skills") successfully loads the skill and lasts for the lifetime of the pending-skill policy. Interactively that is the session: the surface stays armed across the operator's later turns, because a multi-turn skill workflow is still the same workflow on the operator's next message. It ends when a different skill replaces it (the new skill's declaration replaces the old one, it is never merged into it), when the operator clears it with `/skill off`, or when the session ends. A worker keeps the run-scoped lifetime. Activation and clearing each emit one transcript line naming the armed skills.
133
+ - **Who activates.** At `read-only` and `suggest` only the operator activates a skill: a model `context(scope="skills", name=...)` call is refused and the model's move is the suggestion anchor. At `auto-edit` and `full-auto` the model activates an installed skill itself, under the same per-run policy `/skill` produces, because the operator has already chosen to let it act and narrowing can only subtract from the surface. The autonomy level is the whole opt-in; there is no frontmatter flag. A skill that is not installed stays operator-gated at every level, and the transcript line names who activated. This holds on every surface that resolves effective autonomy through the chat loop, which is all three: interactive turns, headless `clio-coder run --autonomy ...`, and ACP prompts (including a per-session level an ACP client sets). Dispatched workers are unaffected: a worker loads only the skills its recipe declares.
133
134
  - **Merge.** Denials win: a tool named in any loaded skill's `disallowed-tools` is blocked. Allow-narrowing applies only while every loaded skill declares `allowed-tools`; the merged surface is the union of those lists. A loaded skill that declares no `allowed-tools` keeps the full surface for its own workflow, which lifts the allow-narrowing (never the denials) for that window.
134
135
  - **Exemptions.** `context` (the remaining requested skills of the turn must still load) and `ask_user` (the escape hatch the block message points at) are always admitted.
135
136
  - **Direction.** Narrowing only blocks. It never grants a tool the safety net, damage-control rules, or autonomy mapping would refuse, and an out-of-surface call blocks terminally instead of parking for confirmation.
136
- - **Block message.** The rejection names the tool, the active skill(s), and the merged surface, and states the remediation: work within the declared surface, or use `ask_user` (when available) to hand the step to the operator. The audit row carries reason code `skill_surface`.
137
+ - **Block message.** The rejection names the tool, the active skill(s), the merged surface, and the lifetime that actually applies (session-scoped for a carried surface, turn/run-scoped otherwise), and states the remediation: work within the declared surface, or use `ask_user` (when available) to hand the step to the operator. The audit row carries reason code `skill_surface`.
137
138
 
138
139
  ---
139
140
 
@@ -236,7 +237,8 @@ Path-policy behavior:
236
237
  The default damage-control policy populates `noWritePaths` from the interop
237
238
  agent registry: `~/.claude/`, `.claude/`, `~/.codex/`, `.codex/`, `~/.config/opencode/`,
238
239
  `.opencode/`, `~/.gemini/`, `.gemini/`, `~/.copilot/`, `~/.cursor/`, `.cursor/`,
239
- `~/.antigravitycli/`, `.antigravitycli/`, `~/.agents/`, and `.agents/`. Clio never writes
240
+ `~/.gemini/antigravity-cli/`, `.gemini/antigravity-cli/`, the legacy
241
+ `~/.antigravitycli/` and `.antigravitycli/`, `~/.agents/`, and `.agents/`. Clio never writes
240
242
  into another coding agent's directory. It reads those roots for skills, prompts, and
241
243
  rule prose and has no reason to author them. A `write` or `edit` targeting any of
242
244
  these paths is refused at every posture including `auto-edit` and `full-auto`, with reason
@@ -290,14 +292,22 @@ Fleet dispatch is admitted only when the requested worker scope is a subset of t
290
292
 
291
293
  Dispatch workers can run the same HTTP or native runtimes as the orchestrator. Clio observes and governs those tool calls directly, so every worker run is subject to the same safety mapping and receipt accounting as an interactive turn.
292
294
 
293
- Three integration paths exist for driving Claude Code, ranging from fully enforced to advisory gating:
295
+ Three worker-runtime safety categories range from fully enforced to advisory gating:
294
296
 
295
297
  - **`claude-sdk` (Enforced Safety):** Drives [@anthropic-ai/claude-agent-sdk](https://www.npmjs.com/package/@anthropic-ai/claude-agent-sdk) directly. This is the **strong safety path** because Clio enforces tool gating before execution. Clio registers a `PreToolUse` hook (which fires for all tool uses, including auto-allowed reads) and wraps `canUseTool` for permission paths. Every tool request is mapped into a Clio tool/action class, evaluated by the safety net, and passed through the active autonomy matrix. Because a dispatched worker is noninteractive, any `ask` decision is resolved as a non-stall denial (`fleet.permissions.mode=deny` returns denial; `fleet.permissions.mode=fail` terminates the run with a permission-required code).
296
- - **`claude-code` (Subprocess Gating):** Drives `claude -p` as a subprocess. Because the CLI lacks a direct callback hook, Clio cannot evaluate each tool invocation. Instead, Clio maps the active autonomy level to the binary's command-line parameters (such as `--permission-mode` and tool allowlists). Unrecognized tools are gated by the subprocess runtime itself. Dispatch at autonomy `suggest` is refused outright (the same applies to `antigravity-code`): a subprocess cannot park a tool call for approval, so `suggest` has no honest mapping and the runner fails closed before launching the external CLI. A dangerous bypass (`--allow-dangerously-skip-permissions`) is only sent when autonomy is `full-auto` and `CLIO_CODER_ALLOW_EXTERNAL_FULL_ACCESS=1`, and it is never silent: the run's receipt records it (see the enforcement grades below) and evidence raises an external-bypass finding.
298
+ - **External CLI subprocesses:** `claude-code` drives `claude -p`; the experimental, dispatch-only `antigravity-code` runtime drives the operator's local `agy` through one literal stdin `stream-json` work order. Neither exposes a callback through which Clio can evaluate each tool invocation, so Clio maps autonomy onto each CLI's command-line controls. Antigravity launches always name an explicit mode and disable slash-command expansion rather than inheriting mutable interactive settings. Dispatch at autonomy `suggest` is refused outright: a subprocess cannot park a tool call for approval, so `suggest` has no honest mapping and the runner fails closed before launch. A dangerous bypass (`--allow-dangerously-skip-permissions` for Claude or `--dangerously-skip-permissions` for Antigravity) is sent only when autonomy is `full-auto` and `CLIO_CODER_ALLOW_EXTERNAL_FULL_ACCESS=1`; otherwise Antigravity full-auto is capped at `accept-edits`. A bypass is never silent: the run's receipt records it (see the enforcement grades below) and evidence raises an external-bypass finding. Clio sends an allowlisted child environment rather than its provider keys or external-full-access gate, bounds stdout/stderr and persisted diagnostics, validates the admitted workspace, and owns the deadline and cancellation. POSIX cancellation targets the process group with direct-child fallback and bounded SIGTERM-to-SIGKILL escalation; Windows uses the strongest honest direct-child termination available here.
297
299
  - **Claude Code over ACP (Advisory Gating):** Drives Zed's `@zed-industries/claude-code-acp` (or `@agentclientprotocol/claude-agent-acp`) bridge as an [Agent Client Protocol (ACP)](https://agentclientprotocol.com) delegation agent. Clio's ACP mediator intercepts tool calls and filters them against the safety net, but gating is ultimately **advisory** as Claude governs its own runtime execution. For strict, code-enforced per-tool safety, `claude-sdk` is preferred over ACP.
298
300
 
299
301
  All Claude Code runtimes rely on the user's existing CLI authentication and store no credentials in Clio.
300
302
 
303
+ External one-shot receipts also say what budget Clio can and cannot enforce. Clio
304
+ controls one process launch, its wall-clock deadline, cumulative output cap,
305
+ cancellation, and result-contract validation. Recipe per-tool calls, read reserve,
306
+ and synthesis numbers remain in the envelope but are explicitly
307
+ `unobserved-not-enforced`, because Antigravity owns its internal tools, network,
308
+ prompts, and approvals. Clio schedules no automatic retry of an external
309
+ generating agent loop.
310
+
301
311
  ### Autonomy enforcement grades
302
312
 
303
313
  How faithfully a runtime can honor the autonomy model is a recorded fact, not an assumption. Worker receipts carry an optional `autonomyEnforcement` block sealed into the integrity digest:
@@ -64,7 +64,8 @@ Internal orchestration helpers and internal process agents. They are hidden from
64
64
  | Agent ID | Primary tools | Purpose | Capability | Latency |
65
65
  | --- | --- | --- | --- | --- |
66
66
  | `scout` | read, grep, find, ls, context, code_nav, git, ledger | Broad repository reconnaissance with cited findings: orientation, structure and entry-point mapping, multi-file symbol hunting. | `read-only` | `fast` |
67
- | `researcher` | read, web_fetch, context, ledger | Researches external docs, standards, and papers for coding decisions. | `read-only` | `deep` |
67
+ | `researcher` | read, web_fetch, context, ledger | Extracts and compares concrete supplied URLs, standards, release notes, and papers through Clio-observed reads and URL retrieval. | `read-only` | `deep` |
68
+ | `world-knowledge` | optional web_fetch, read, context, ledger | Current open-world discovery, ecosystem comparison, broad external context, and an advisory second opinion; reports when discovery is unavailable. | `read-only` | `deep` |
68
69
  | `provenance` | read, grep, find, ls, git, ledger | Reads receipts, diffs, and telemetry for evidence-backed handoffs. | `read-only` | `balanced` |
69
70
  | `oracle` | read, grep, find, ls, code_nav, context, ledger | Shadow advisor behind `/oracle` that protects consistency with prior decisions and returns the strongest challenge to a question. | `read-only` | `deep` |
70
71
  | `context-bootstrap` | read, grep, find, ls, context, code_nav | Internal agent behind `clio-coder context init` that parses repository and returns CLIO-CODER.md payload. | `read-only` | `balanced` |
@@ -75,6 +76,19 @@ The builtin `architect` also serves as the default author for a version 5 fleet
75
76
 
76
77
  Grounding is checked against the run's own reads, not just against the file. The worker records the exact line span every successful read returned, and a cited line must fall inside one. A line that exists in the file but was never read fails, which is what stops an approximated or inferred line number from passing as observation. `grep` and `code_nav` hits are leads: read the file before citing what they point at.
77
78
 
79
+ The three discovery roles are deliberately non-overlapping. `scout` is
80
+ repository-only reconnaissance with live `path:line` grounding and never browses
81
+ external sources. `researcher` starts from concrete URLs or documents and uses
82
+ Clio-observed `read`/`web_fetch` calls to extract and compare them; `web_fetch` is
83
+ URL retrieval, not search. `world-knowledge` is for current open-world discovery,
84
+ ecosystem comparison, broad context, and an independent advisory opinion. Its
85
+ tools are all optional so it can run on a native Clio target or an opaque external
86
+ worker. A native target without discovery must use caller-supplied sources or say
87
+ discovery was unavailable. Its `world-knowledge-report` separates supported
88
+ facts and supplied source identifiers from synthesis, uncertainty, and follow-up
89
+ verification; it never fabricates citations. The capability class remains
90
+ `read-only` regardless of the caller's requested autonomy.
91
+
78
92
  `oracle` is the only shadow agent an operator reaches directly, and only through
79
93
  `/oracle <question>`. It never receives a forked transcript. `/oracle` packs a
80
94
  bounded digest instead and sends it as dispatch briefing data: the settled
@@ -181,9 +195,9 @@ To ensure security and proper boundary isolation, shadow and internal agents are
181
195
  In addition to standard HTTP targets and [Agent Client Protocol (ACP)](https://agentclientprotocol.com) delegation agents, Clio dispatches subagents to sanctioned subscription worker runtimes:
182
196
  - **`claude-sdk` (Claude Agent SDK):** Serves as a main worker runtime for driving fleet agents. It integrates with [@anthropic-ai/claude-agent-sdk](https://www.npmjs.com/package/@anthropic-ai/claude-agent-sdk) alongside Clio's native subagent workers (like a local [llama.cpp](https://github.com/ggerganov/llama.cpp), [Ollama](https://ollama.com), [LM Studio](https://lmstudio.ai), [vLLM](https://github.com/vllm-project/vllm), or [SGLang](https://github.com/sgl-project/sglang) fleet) to execute tasks under a Claude subscription. Every tool call is mediated by Clio (`canUseTool` plus a `PreToolUse` hook): the safety net and autonomy matrix apply, and the run's admitted tool surface, which is narrowed by any `tool_profile`, is enforced authoritatively. Consequently, an out-of-profile tool (for example `bash` under `minimal-local`) is denied even though the underlying preset offers it. The narrowed surface is also translated into the SDK's `disallowedTools` option as defense in depth. Because it routes tool calls through Clio safety, it behaves as a native worker.
183
197
  - **`claude-code` (Claude Subprocess):** Runs `claude -p` as a subprocess worker, mapping autonomy levels to the CLI's permission modes. It is a black box: tool calls run inside the `claude` process and are not routed through Clio's per-tool mediation, so Clio cannot enforce a per-tool profile on it. Dispatching a narrowing `tool_profile` (`minimal-local` or `science-local`) to this runtime is refused; use `full-agent` (or a native / `claude-sdk` worker) instead.
184
- - **`antigravity-code` (Antigravity CLI):** Runs an Antigravity CLI subprocess as a subscription worker target for fleet dispatch. Like `claude-code`, it is a black box with no per-tool mediation and no per-tool allowlist, so a narrowing `tool_profile` is refused rather than silently ignored.
198
+ - **`antigravity-code` (Antigravity CLI — experimental local delegation):** Runs the operator-installed and authenticated official `agy` command as a local external delegation worker. It is useful for a `world-knowledge` pass, a second opinion, or another bounded one-shot subtask; it is never an orchestrator or Gemini chat backend. Clio consumes agy's structured stream and live model catalog but cannot mediate individual tools, so a narrowing `tool_profile` is refused rather than silently ignored. The `world-knowledge` binding is permanently read-only.
185
199
 
186
- Agent budgets follow the same mediation boundary. Native workers and `claude-sdk` enforce canonical call counting, the canonical-`read` reserve, and the synthesis transition. Every valid Markdown recipe declares a budget, so recipe-based runs are refused on `claude-code` and `antigravity-code`; those subprocess targets can accept only work that reaches admission without an explicit recipe or request budget. Silently ignoring numeric bounds would be unsafe. Claude vendor aliases never appear in recipes or prompt authority and cannot reintroduce a canonical tool removed by admission.
200
+ Agent budgets follow the same mediation boundary. Native workers and `claude-sdk` enforce canonical call counting, the canonical-`read` reserve, and the synthesis transition. An opaque external loop instead receives an `external-one-shot` enforcement classification: Clio enforces one subprocess launch, its deadline, output cap, cancellation, and result-contract validation, while recording recipe per-tool numbers as `unobserved-not-enforced`. Receipts and status never label those internal per-tool limits enforced, and Clio never automatically retries a generating external-agent run. Claude vendor aliases never appear in recipes or prompt authority and cannot reintroduce a canonical tool removed by admission.
187
201
 
188
202
  Interactive TUI:
189
203
 
@@ -78,7 +78,7 @@ For process exit codes, stdout deliverable guarantees, and machine-readable JSON
78
78
  | `clio-coder extensions list\|discover\|install\|enable\|disable\|remove` | Manage installed extension packages and resource roots. `clio-coder ext` is an accepted alias. |
79
79
  | `clio-coder skills list\|search\|inspect\|validate\|install\|update\|sync\|eval` | Manage discovered skills, Clio-native skills, and local marketplace installs. |
80
80
  | `clio-coder docs [topic] [--no-open]` | Serve the interactive HTML docs of a source checkout on 127.0.0.1; the npm package ships the Markdown guides only. |
81
- | `clio-coder usage report [--repo <path>] [--days <n>] [--json]` | Cross-session usage facts and opportunities from the session and run ledgers. The window defaults to 30 days and the JSON schema is marked experimental. |
81
+ | `clio-coder usage report [--repo <path>] [--days <n>] [--json]` | Cross-session usage facts from session/run ledgers and retained out-of-turn calls, including known failed-compaction spending and missing coverage. The window defaults to 30 days and the JSON schema is marked experimental. |
82
82
  | `clio-coder dev share export --out <path> [--project\|--user\|--both] [--context] [--prompts] [--skills] [--settings] [--extensions]` | Export project context, prompts, skills, settings fragments, and extension bundles. |
83
83
  | `clio-coder dev share import <path> [--dry-run] [--force] [--project\|--user] [--json]` | Import a share archive with conflict reporting. |
84
84
  | `clio-coder dev share inspect <path> [--json]` | Inspect a share archive without importing it. |
@@ -3,7 +3,7 @@
3
3
  > **Visual blueprint:** The source checkout includes the complete
4
4
  > [Configuration, Targets, Runtimes, and Auth visual reference](https://github.com/iowarp/clio-coder/blob/main/docs/html/configuration_blueprint.html).
5
5
 
6
- Clio Coder is target-first: chat and fleet dispatch resolve through configured targets in `settings.yaml`, not through provider-specific ad hoc flags. Chat and print targets are HTTP and native engine-backed runtimes. Fleet dispatch can also target the sanctioned Claude Code subscription runtimes described below.
6
+ Clio Coder is target-first: chat and fleet dispatch resolve through configured targets in `settings.yaml`, not through provider-specific ad hoc flags. Chat and print targets are HTTP and native engine-backed runtimes. Fleet dispatch can also target sanctioned subscription and external-worker runtimes described below.
7
7
 
8
8
  Clio's engine is built on the pi SDK (see [docs/architecture/pi-boundary.md](../architecture/pi-boundary.md)). Broad provider/model support comes from engine-backed descriptors and from the generic `openai-compat` and `anthropic-compat` targets. Clio adds orchestration, local/native runtime ergonomics, target configuration, dispatch, safety, and receipts rather than creating a first-class descriptor for every provider.
9
9
 
@@ -72,6 +72,15 @@ Start one local runtime and register exactly one target first. Clio integrates w
72
72
  - **[llama.cpp](https://github.com/ggerganov/llama.cpp):** A minimal C/C++ implementation for local LLM inference. Target runtime ID: `llamacpp`.
73
73
  - **[vLLM](https://github.com/vllm-project/vllm):** A high-throughput and memory-efficient LLM serving engine. Target runtime ID: `vllm`.
74
74
  - **[SGLang](https://github.com/sgl-project/sglang):** A fast serving framework for large language models. Target runtime ID: `sglang`.
75
+ - **[LiteLLM](https://docs.litellm.ai):** An OpenAI-compatible gateway that publishes routed models across multiple inference endpoints. Target runtime ID: `litellm`.
76
+
77
+ First-run onboarding always completes a valid orchestrator before it offers any
78
+ worker-only colleague. If `agy` is already on `PATH`, the final optional step is
79
+ named **Antigravity CLI — experimental local delegation**. It runs only the
80
+ non-generating model-catalog probe, explains that Clio uses the operator's
81
+ existing local session without inspecting credentials, and can create and bind a
82
+ read-only `world-knowledge-external` profile. Skipping or failing that step does
83
+ not change the primary chat, background, or fleet-default pointers.
75
84
 
76
85
  Common local runtime IDs and default URLs are:
77
86
 
@@ -82,6 +91,7 @@ Common local runtime IDs and default URLs are:
82
91
  | llama.cpp server | `llamacpp` | `http://127.0.0.1:8080` |
83
92
  | vLLM | `vllm` | `http://127.0.0.1:8000` |
84
93
  | SGLang | `sglang` | `http://127.0.0.1:30000` |
94
+ | LiteLLM gateway | `litellm` | `http://127.0.0.1:4000` |
85
95
 
86
96
 
87
97
  Example registration:
@@ -230,11 +240,61 @@ integrations:
230
240
 
231
241
  Target capability overrides may include `chat`, `tools`, `toolCallFormat`, `reasoning`, `thinkingFormat`, `structuredOutputs`, `vision`, `audio`, `embeddings`, `rerank`, `fim`, `contextWindow`, and `maxTokens`.
232
242
 
243
+ ### LiteLLM gateways
244
+
245
+ Register LiteLLM with its first-class runtime so discovery reads
246
+ `/v1/model/info` rather than treating routed names as bare OpenAI models. For a
247
+ gateway where placement matters, publish one deterministic `node/model` name
248
+ per deployment instead of task categories such as `chat` or `code`:
249
+
250
+ ```yaml
251
+ targets:
252
+ - id: ai-gateway
253
+ runtime: litellm
254
+ url: http://gateway.example:4000
255
+ defaultModel: dynamo/qwen3.8-27b
256
+ auth:
257
+ apiKeyEnvVar: CLIO_AI_GATEWAY_KEY
258
+ litellm:
259
+ request:
260
+ tags: [homelab]
261
+ sendSessionId: true
262
+ # Optional proxy overrides. Keep retries at zero for physical routes.
263
+ timeoutSeconds: 300
264
+ streamTimeoutSeconds: 180
265
+ numRetries: 0
266
+ ```
267
+
268
+ Clio adds the `clio-coder` request tag and forwards its stable session id by
269
+ default. Explicit `x-litellm-*` headers on the target or request take
270
+ precedence. Set `sendSessionId: false` when the gateway must not correlate calls.
271
+ The timeout values and `numRetries` become LiteLLM request headers and override
272
+ the proxy's defaults only for this target. Clio disables the OpenAI SDK's
273
+ client-side retry layer on LiteLLM requests. A failed interactive LiteLLM call
274
+ also bypasses Clio's transient retry ladder even when `chat.retry` is enabled:
275
+ the provider error names the selected route and tells the operator to choose a
276
+ different one with `/model`. Configure `numRetries: 0` and no server fallback
277
+ maps when `node/model` is a placement guarantee. Context-overflow compaction is
278
+ still a local correction, not a route substitution.
279
+
280
+ When the gateway reports these fields, Clio records LiteLLM's selected model group, physical model,
281
+ sanitized upstream host, fallback/retry counts, and proxy timing in the session
282
+ and worker receipt. A nonzero fallback or retry is also announced in the live
283
+ transcript, making server-policy drift visible. Self-describing physical routes
284
+ are summarized in probe diagnostics instead of expanding into a giant list of
285
+ tautological mappings. If a genuine alias has multiple deployments, discovered
286
+ capabilities are the conservative intersection and numerical limits are the
287
+ smallest limits published by every deployment, so routing cannot select a
288
+ weaker backend than Clio planned for. If `/v1/model/info` omits a capability,
289
+ Clio does not invent it; that includes structured-output support.
290
+
291
+ Upstream `model_info.runtime` declarations may identify the thinking-control dialect of a routed model. Clio accepts only recognized declarations shared by every deployment of that alias; unknown or mixed declarations remain unguessed. A resolved LM Studio dialect uses its explicit off effort, while llama.cpp uses template controls. This metadata does not turn the target into a native management endpoint: LiteLLM continues to own authentication, routing, loading and eviction. Thinking off is a request to the selected runtime/model, not evidence that the server complied or that every route has been live-validated.
292
+
233
293
  ### `maxConcurrentRequests`
234
294
 
235
295
  `maxConcurrentRequests` is a per-target integer of at least 1, validated with the rest of the target block, and it is the operator's override for how many requests the inference endpoint behind that target can serve at once. It is not a settings-file default and has no shipped value, so it does not appear in the settings inventory below.
236
296
 
237
- Set it only when discovery is wrong. Clio resolves the limit in this order: this override; then a `parallelSlots` count cached on the target's probe result; then one slot for any other `local-native` runtime; then no bound at all for a cloud runtime, vLLM, or SGLang. llama.cpp discovery reads `total_slots` from the router's `/props`, falls back to the selected worker's `/props?model=<id>` when the router reports none, and falls back again to the `--parallel` argv on the selected `/v1/models` entry. LM Studio reads `config.parallel` off the loaded instance and otherwise reports one; Ollama reads `OLLAMA_NUM_PARALLEL` from the environment the Clio process can see and otherwise reports one.
297
+ Set it only when discovery is wrong. Clio resolves the limit in this order: this override; then a `parallelSlots` count cached on the target's probe result; then a persisted discovery result that is still fresh and matches the runtime; then one slot for any other `local-native` runtime; then no invented endpoint bound for a cloud runtime, LiteLLM, vLLM, or SGLang. llama.cpp discovery reads `total_slots` from the router's `/props`, falls back to the selected worker's `/props?model=<id>` when the router reports none, and falls back again to the `--parallel` argv on the selected `/v1/models` entry. LM Studio reads `config.parallel` off the loaded instance and otherwise reports one; Ollama reads `OLLAMA_NUM_PARALLEL` from the environment the Clio process can see and otherwise reports one.
238
298
 
239
299
  The limit is keyed on the endpoint rather than the target, so two targets pointed at the same normalized URL share it. Raising it above what the server will actually serve does not create capacity; it removes the refusal that would have told you the server was full. See [capacity-and-scheduling.md](../architecture/capacity-and-scheduling.md) for the admission model and the exact denial text.
240
300
 
@@ -612,6 +672,8 @@ This is the version-2 durable schema shipped in `DEFAULT_SETTINGS`. Validation i
612
672
  | `context.memory.maxOutputTokens` | `2000` | integer ≥ 1 | next turn |
613
673
  | `context.memory.timeoutMs` | `60000` | integer ≥ 1 | next turn |
614
674
 
675
+ The compaction and memory controls serve different roles. An unset `context.compaction.model` uses active chat; an explicit model must uniquely resolve to an available eligible summary route or fail visibly. `context.compaction.systemPrompt` is a nonempty UTF-8 prompt file, at most 65,536 bytes, read at compaction time and resolved relative to the session workspace. Configure `context.memory.target` and `context.memory.model` to opt into model-based memory; unset roles remain rules-only. Memory prefers its dedicated route and can use active chat when that route is unavailable and request capacity permits. Known dedicated saturation skips the step rather than initiating failover; neither routing choice edits saved settings.
676
+
615
677
  ### Safety
616
678
 
617
679
  The safety-limit leaves have no one-process `CLIO_CODER_*` overrides in the current schema. Resolution follows the normal settings stack, from session or project layers where supported through user `settings.yaml`, then the compiled default.
@@ -824,17 +886,37 @@ integrations:
824
886
  ```
825
887
  Then invoke it using `/delegate claude-code <task>`.
826
888
 
827
- ### 5. Google Antigravity CLI Runtime (Worker-Only)
889
+ ### 5. Antigravity CLI Experimental Local Delegation (Worker-Only)
828
890
 
829
- The `antigravity-code` runtime drives your local Google Antigravity CLI (`agy`) installation to execute subagent tasks. It runs the CLI as a subprocess using the `agy --print` command and maps Clio autonomy levels onto the CLI's permission flags.
891
+ The `antigravity-code` runtime is a local external delegation agent. It lets Clio ask an official Antigravity CLI (`agy`) installed and authenticated by the operator for research, world knowledge, a second opinion, or a bounded subtask. It is deliberately **not** a Gemini chat provider and can never become the Clio orchestrator.
830
892
 
831
- Google Antigravity supports a context window of up to 1,000,000 tokens and is suitable for large-context codebase reasoning. Because `agy` emits plain text without structured events, Clio cannot perform fine-grained tool call interception. Gating is applied coarsely: read-only runs pass both `--mode plan` (the no-change agent posture) and `--sandbox` (terminal restrictions). Full-auto passes `--dangerously-skip-permissions` only when the environment variable `CLIO_CODER_ALLOW_EXTERNAL_FULL_ACCESS=1` is explicitly set.
893
+ This integration is experimental and intended only for personal use on your own machine. Clio does not install Antigravity, initiate Google sign-in, copy credentials, or read the CLI's credential store. Install the [official Antigravity CLI](https://antigravity.google/docs/cli/overview), run `agy` yourself to sign in, and keep it current. Clio then starts that same local executable for an explicit delegation. The integration feature-probes the required structured catalog; if the installed CLI lacks it, Clio asks the operator to update `agy` rather than relying on a hard-coded minimum version.
832
894
 
833
- Configured targets use your existing local `agy` login and credentials. Supported model names include `Gemini 3.5 Flash (High)` as the default tier, `Gemini 3.5 Flash (Medium)`, `Gemini 3.5 Flash (Low)`, `Gemini 3.1 Pro (High)`, `Gemini 3.1 Pro (Low)`, `Claude Sonnet 4.6 (Thinking)`, `Claude Opus 4.6 (Thinking)`, and `GPT-OSS 120B (Medium)`.
895
+ Clio sends one literal prompt as a `stream-json` stdin record, closes stdin, and validates agy's init, delta, and single terminal-result sequence, including the opaque conversation identifier and provider-reported token counts. Prompt text never appears in argv and slash-command expansion is disabled. `clio-coder targets --probe --target <id>` runs the non-generating `agy --output-format json models` command and distinguishes missing CLI, sign-in required, CLI update required, unavailable live catalog, and configured-model disappearance. A successful account catalog is authoritative. The five built-in slugs are cold-start hints only; live `{id,label}` rows and last-good cached labels are displayed without replacing the stable model slug, and cached fallback is marked stale.
834
896
 
835
- **Configuration Example:**
897
+ Antigravity remains an external agent loop: Clio cannot intercept each tool call. Every launch disables slash-command expansion and sets an explicit posture rather than inheriting mutable CLI defaults:
898
+
899
+ | Clio autonomy | Antigravity launch posture |
900
+ | --- | --- |
901
+ | `read-only` | `--mode plan --sandbox` |
902
+ | `auto-edit` | `--mode accept-edits` |
903
+ | `suggest` | Refused because a headless subprocess cannot pause for Clio approval |
904
+ | `full-auto` | Capped at `accept-edits` unless the external full-access gate is explicitly enabled |
905
+
906
+ Only `full-auto` together with `CLIO_CODER_ALLOW_EXTERNAL_FULL_ACCESS=1` passes `--dangerously-skip-permissions`. Treat that as an external safety bypass: agy, not Clio's tool registry, controls the resulting filesystem, shell, network, prompts, and approvals. The gate is never enabled by onboarding. A `world-knowledge` dispatch remains read-only even if the caller requests a stronger posture; use another agent for mutation work.
907
+
908
+ **Setup and verification:**
836
909
  ```bash
837
- clio-coder configure --id agy-worker --runtime antigravity-code --model "Gemini 3.5 Flash (High)"
910
+ # Install using Google's instructions, then authenticate directly in the CLI.
911
+ agy
912
+ agy models
913
+
914
+ # Use a live model slug, then create and bind a read-only research profile.
915
+ clio-coder configure --id agy-research --runtime antigravity-code \
916
+ --model gemini-3.8-flash-high \
917
+ --agent-profile world-knowledge-external \
918
+ --bind-agent world-knowledge
919
+ clio-coder targets --probe --target agy-research
838
920
  ```
839
921
 
840
922
 
@@ -851,6 +933,7 @@ Useful flags:
851
933
  | `--fleet-model <id>` | Model to save for fleet default. |
852
934
  | `--agent-profile <name>` | Save this target/model as a named fleet profile. |
853
935
  | `--agent-profile-model <id>` | Model to save for the named fleet profile. |
936
+ | `--bind-agent <agentId>` | Bind the agent to `--agent-profile` in the same configuration write; requires `--agent-profile`. |
854
937
  | `--api-key-env <VAR>` | Read API key from the environment at call time. |
855
938
  | `--api-key <literal>` | Store an API key in `credentials.yaml`. |
856
939
  | `--force` | Allow model/capability choices outside the local catalog guardrails. |
@@ -968,6 +1051,11 @@ Model rows combine:
968
1051
  2. runtime-discovered models from probes;
969
1052
  3. known models from bundled/provider catalogs.
970
1053
 
1054
+ Each row reports its stable model id, optional human label, and source
1055
+ (`live`, cached, configured/default, or catalog). When a live authoritative
1056
+ catalog exists, a disappeared configured model is unavailable rather than being
1057
+ resurrected by a descriptor hint.
1058
+
971
1059
  Capability badges in CLI output are compact:
972
1060
 
973
1061
  | Badge | Capability |
@@ -990,7 +1078,7 @@ Representative built-in runtime IDs:
990
1078
  | --- | --- |
991
1079
  | Protocol-compatible | `openai-compat`, `anthropic-compat` generic surfaces for additional OpenAI-compatible or Anthropic-compatible APIs, including APIs such as InceptionAI when configured with the appropriate base URL and credentials. |
992
1080
  | Cloud | `alcf`, `anthropic`, `bedrock`, `deepseek`, `google`, `groq`, `mistral`, `openai`, `openrouter` |
993
- | Subscription and worker harnesses | `openai-codex` for ChatGPT OAuth, `anthropic-max` for Anthropic OAuth, `claude-sdk` for Claude Agent SDK workers, `claude-code` for `claude -p` subprocess workers, and `antigravity-code` for `agy --print` subprocess workers |
1081
+ | Subscription and worker harnesses | `openai-codex` for ChatGPT OAuth, `anthropic-max` for Anthropic OAuth, `claude-sdk` for Claude Agent SDK workers, `claude-code` for `claude -p` subprocess workers, and `antigravity-code` for structured `agy` external delegation |
994
1082
  | Local native | `llamacpp`, `lmstudio`, `ollama-native`, `vllm`, `sglang`, `lemonade`, `lemonade-anthropic` |
995
1083
 
996
1084
  Some hidden aliases exist for backward compatibility or special surfaces; use `clio-coder configure --list --all` to see them.
@@ -29,8 +29,8 @@ Precedence, where several surfaces set the same value: a one-run CLI flag beats
29
29
  | `chat.target` | `null` | Saved default chat target id (configured id or null); a dangling id normalizes to null and clears `chat.model`; seeds session routing at launch, so applies next session. | `run --target` > session routing (`/model`, `/thinking`, Alt+J/K, `/settings` apply-this-session) > `.clio-coder/settings.local.yaml` > `.clio-coder/settings.yaml` > user settings.yaml > default |
30
30
  | `chat.thinkingLevel` | `low` | Saved default reasoning effort for the chat model: `off`, `minimal`, `low`, `medium`, `high`, `xhigh`, `max`, clamped to what the model supports at request time; applies next session. | `run --thinking` > session routing (`/model`, `/thinking`, Alt+J/K, `/settings` apply-this-session) > `.clio-coder/settings.local.yaml` > `.clio-coder/settings.yaml` > user settings.yaml > default |
31
31
  | `context.compaction.auto` | `true` | Master switch for the pre-request compaction trigger when context pressure crosses `threshold` (boolean); manual `/context compact` still works when off; applies next turn. | session override (`/settings` apply-this-session) > `.clio-coder/settings.local.yaml` > `.clio-coder/settings.yaml` > user settings.yaml > default |
32
- | `context.compaction.model` | | Optional `provider/model` pattern resolved for the summary model when compaction runs; absent means the chat target summarizes; applies next turn. | |
33
- | `context.compaction.systemPrompt` | | Optional path to a prompt-override file read at compaction time to replace the built-in summary prompt; applies next turn. | |
32
+ | `context.compaction.model` | | Optional model pattern (prefer exact `target/model`) resolved when compaction runs. Must uniquely match an available HTTP chat model; invalid, ambiguous, unavailable, and worker-only selections fail explicitly. Absent means the current chat route summarizes. | |
33
+ | `context.compaction.systemPrompt` | | Optional UTF-8 prompt file read on every compaction to replace the built-in system prompt. Relative paths use the current session workspace (`cwd`); absolute paths are supported. Must be a nonempty regular file of at most 65,536 bytes; read failures are explicit. `/context compact` focus instructions remain separate. | |
34
34
  | `context.compaction.threshold` | `0.8` | Context pressure (estimated tokens over the context window, 0 through 1) at which eviction and then LLM summary compaction run; applies next turn. | session override (`/settings` apply-this-session) > `.clio-coder/settings.local.yaml` > `.clio-coder/settings.yaml` > user settings.yaml > default |
35
35
  | `context.toolResultMaxBytes` | `65536` | Maximum bytes returned from one tool result before the complete text spills to the session scratch file (integer >= 4096); applies to the next tool result. The unchanged `safety.limits.observationBytesPerTurn` default is 196608 bytes, so three full-size results consume the shared pool and a fourth finds it filled. | session override (`/settings` apply-this-session) > `.clio-coder/settings.local.yaml` > `.clio-coder/settings.yaml` > user settings.yaml > default |
36
36
  | `context.memory.cadenceToolCalls` | `10` | Tool calls between proactive memory interventions inside a turn (integer >= 2); applies next turn. | session override (`/settings` apply-this-session) > `.clio-coder/settings.local.yaml` > `.clio-coder/settings.yaml` > user settings.yaml > default |
@@ -165,6 +165,11 @@ Precedence, where several surfaces set the same value: a one-run CLI flag beats
165
165
  | `targets[].lmstudio.request.draftModel` | | Sends `draft_model` on OpenAI-compatible chat requests for speculative decoding (model id string). | |
166
166
  | `targets[].lmstudio.request.reasoning` | | How `reasoning_effort` is sent: `auto` maps the active thinking level, `off` sends `none`, `on` sends `low`, `low`/`medium`/`high` are literal but clamped to the efforts the model advertises. | |
167
167
  | `targets[].lmstudio.request.ttlSeconds` | | Sends `ttl` on chat requests so LM Studio auto-evicts the model after this idle time in seconds (integer >= 1). | |
168
+ | `targets[].litellm.request.numRetries` | | Optional LiteLLM router retry override sent as `x-litellm-num-retries` (integer >= 0); use `0` for deterministic physical routes. The OpenAI SDK and Clio interactive transient-retry layers remain disabled for LiteLLM failures. | |
169
+ | `targets[].litellm.request.sendSessionId` | `true` | Whether Clio forwards its stable session id as `x-litellm-session-id` (boolean); applies when the next LiteLLM agent runtime is created. | |
170
+ | `targets[].litellm.request.streamTimeoutSeconds` | | Optional LiteLLM streaming timeout override sent as `x-litellm-stream-timeout` (number from 0.001 through 86400); omit it to keep gateway policy. | |
171
+ | `targets[].litellm.request.tags` | `[clio-coder]` | Additional LiteLLM request tags (list of non-empty strings without commas); Clio always adds `clio-coder`. | |
172
+ | `targets[].litellm.request.timeoutSeconds` | | Optional LiteLLM request/upstream timeout override sent as `x-litellm-timeout` (number from 0.001 through 86400); omit it to keep gateway policy. | |
168
173
  | `targets[].maxConcurrentRequests` | | Explicit request-slot limit for this inference endpoint (integer >= 1); overrides live slot discovery and is shared by every target on the same normalized URL; applies next turn. | |
169
174
  | `targets[].pricing.cacheRead` | `0` | USD rate for cache-read tokens (number >= 0); 0 when absent. | |
170
175
  | `targets[].pricing.cacheWrite` | `0` | USD rate for cache-write tokens (number >= 0); 0 when absent. | |
@@ -113,6 +113,8 @@ Set by Clio for its own processes; not operator knobs.
113
113
  | --- | --- |
114
114
  | `CLIO_CODER_WORKER_FAUX` (+ `_MODEL`, `_TEXT`, `_STOP_REASON`, `_ERROR_MESSAGE`) | Fake worker model for tests (`src/engine/ai.ts`). |
115
115
  | `CLIO_CODER_TEST_UPGRADE_NO_NETWORK` | Skips npm install during upgrade tests (`src/cli/upgrade.ts`). |
116
+ | `CLIO_CODER_TEST_UPGRADE_AVAILABLE` | Sets mock available version for upgrade tests; `unreachable` stands for a registry that answered nothing (`src/cli/upgrade.ts`). |
117
+ | `CLIO_CODER_TEST_UPGRADE_FAIL` | Injects mock failures (`npm` or `migration`) for upgrade tests (`src/cli/upgrade.ts`). |
116
118
  | `CLIO_CODER_TEST_STAGE1_DELAY_MS`, `CLIO_CODER_TEST_STAGE1_FAIL` | `NODE_ENV=test`-only, bounded instant-shell interleaving and injected hydration failure seams for the built PTY acceptance suite (`src/cli/clio.ts`). |
117
119
  | `CLIO_CODER_REQUIRE_HOME_PREFIX` | Test guardrail: abort if resolved directories escape `CLIO_CODER_HOME` (`src/core/init.ts`). |
118
120
 
@@ -193,7 +193,7 @@ Runs a series of health sweeps across the environment:
193
193
  ### B. Upgrades (`clio-coder upgrade`)
194
194
  Refreshes state metadata and applies pending lifecycle migrations, which may update settings, state, or extension data.
195
195
  ```bash
196
- clio-coder upgrade [--dry-run] [--channel=<latest|beta|dev>] [--skip-migrations]
196
+ clio-coder upgrade [--dry-run] [--channel=<latest|beta|dev>] [--skip-migrations] [--json]
197
197
  ```
198
198
  The command detects the install method from the running binary. On a source
199
199
  checkout it never runs `npm install -g`: it performs its safe local duties
@@ -289,7 +289,7 @@ keyboard-facing text; all other target versions use the generic form
289
289
  ### C. System Resets (`clio-coder reset`)
290
290
  Selective recovery wipes:
291
291
  ```bash
292
- clio-coder reset [--state|--data|--cache|--auth|--config|--all] [--dry-run] [--force]
292
+ clio-coder reset [--state|--data|--cache|--auth|--config|--all] [--dry-run] [--force] [--json]
293
293
  ```
294
294
  Levels are combinable except `--all`. Each level clears exactly the root or file it names and nothing else, then bootstraps the missing structure again unless `--dry-run` is present. `--force` is required only for destructive execution.
295
295
 
@@ -311,7 +311,7 @@ new artifact is written into a root.
311
311
  removes all four roots (config, data, state, cache):
312
312
 
313
313
  ```bash
314
- clio-coder uninstall [--remove-binary] [--dry-run] [--force]
314
+ clio-coder uninstall [--remove-binary] [--keep-config] [--keep-data] [--dry-run] [--force] [--json]
315
315
  ```
316
316
 
317
317
  Preview first, then remove:
@@ -325,7 +325,9 @@ hash -r
325
325
  `--dry-run` prints the roots and the optional launcher action without changing
326
326
  anything, and enumerates the same resolved absolute paths the real run would
327
327
  remove. It prints binary-removal guidance for the active launcher, npm-global
328
- installs, npm links, and the local source symlink.
328
+ installs, npm links, and the local source symlink. Use `--keep-config` to preserve
329
+ `settings.yaml` and `credentials.yaml`, or `--keep-data` to preserve memory and
330
+ evidence.
329
331
 
330
332
  #### Per-project `.clio-coder/` directories
331
333
 
@@ -379,6 +381,37 @@ commands are idempotent, so the recovery is always the same: fix the permission
379
381
  or release the handle, then run the identical command again and it resumes from
380
382
  whatever is left. A partial delete never reports global success.
381
383
 
384
+ ### E. Interactive Configuration (`clio-coder configure`)
385
+ `clio-coder configure` provides a structured, multi-section configuration wizard
386
+ for model targets, runtime defaults, fleet limits, permissions, panes, and
387
+ integrations:
388
+
389
+ ```bash
390
+ clio-coder configure [--section <name>] [--json] [--interop] [--list] [--all]
391
+ ```
392
+
393
+ When run against an unconfigured home (or when launched interactively with no
394
+ model target), it routes immediately to the target category selection menu. Once
395
+ targets are configured, launching `clio-coder configure` presents the 8 top-level
396
+ runtime sections:
397
+
398
+ 1. **Targets & Auth**: manage providers, endpoints, credentials, and models.
399
+ 2. **Models & Thinking**: default models, thinking levels, model favorites, cycle set.
400
+ 3. **Chat Defaults**: smooth streaming, terminal progress, token limits, compaction.
401
+ 4. **Fleet**: concurrency, retries, tool call caps, worker timeouts, profiles.
402
+ 5. **Permissions & Autonomy**: autonomy level, worker permissions, cost limits, review watchdog.
403
+ 6. **Panes & Layout**: terminal panes capability, dock layout, TUI display mode, notifications.
404
+ 7. **Skills & Extensions**: trust project imports, external ACP agents, plugins, library sync.
405
+ 8. **Diagnostics**: version information, resolved directories, doctor check, raw settings inspection.
406
+
407
+ Every section header indicates the exact settings file path being modified
408
+ (`Source: ~/.config/clio-coder/settings.yaml`), prints the current active values,
409
+ and provides a safe back/exit option (`b` or `q`).
410
+
411
+ The `--section <name>` flag jumps directly into any section (e.g.,
412
+ `clio-coder configure --section models` or `clio-coder configure --section fleet`).
413
+ The `--json` flag emits the active settings in formatted JSON for scripting.
414
+
382
415
  ---
383
416
 
384
417
  ## 6. Residues Checklist for Manual Purging