@iowarp/clio-coder 0.4.2 → 0.4.3

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (369) hide show
  1. package/CHANGELOG.md +35 -0
  2. package/CONTRIBUTING.md +86 -19
  3. package/README.md +35 -6
  4. package/dist/{acp-TMDQZDIG.js → acp-H2NGRPWO.js} +11 -11
  5. package/dist/{agents-5N5NG3XG.js → agents-TL5LLUQP.js} +54 -53
  6. package/dist/assets/codewiki.json +1 -1
  7. package/dist/{auth-Z5CCBXKQ.js → auth-E5SW4HMS.js} +19 -16
  8. package/dist/{builtins-K6TNDT24.js → builtins-IA7V7FUC.js} +9 -4
  9. package/dist/{chunk-ZW4HH5JJ.js → chunk-2APPQIER.js} +6 -6
  10. package/dist/{chunk-72GZI5EV.js → chunk-2JH2WHGE.js} +2 -2
  11. package/dist/{chunk-UBRFI4HS.js → chunk-2UG5F4C5.js} +127 -47
  12. package/dist/{chunk-UH632ZYL.js → chunk-2UH2KFUP.js} +2 -2
  13. package/dist/{chunk-3F7VUY77.js → chunk-2VIKGWFZ.js} +2 -2
  14. package/dist/{chunk-I66EAJFY.js → chunk-2WZ546HR.js} +267 -232
  15. package/dist/{chunk-RKSR6VSF.js → chunk-4IUZQIJ3.js} +29 -1
  16. package/dist/{chunk-2X4RYJTJ.js → chunk-4UVU7BJ5.js} +2 -2
  17. package/dist/{chunk-M2DAX4F6.js → chunk-4WR7VSYB.js} +2 -2
  18. package/dist/{chunk-5KW52TEP.js → chunk-54CBCGIR.js} +5 -5
  19. package/dist/{chunk-3KIPBMUA.js → chunk-5ICU3EUH.js} +2 -2
  20. package/dist/{chunk-77QIVUZB.js → chunk-5MEZN6CB.js} +4 -4
  21. package/dist/{chunk-O42A54GG.js → chunk-5OIVVPHF.js} +2 -2
  22. package/dist/{chunk-LJID3DYZ.js → chunk-64I3JVYM.js} +2 -2
  23. package/dist/{chunk-4JDLP6ZS.js → chunk-6PTFB5VS.js} +7 -7
  24. package/dist/{chunk-VN3SHNBN.js → chunk-7DICMOS6.js} +2 -2
  25. package/dist/{chunk-HLAFFSEK.js → chunk-7DRAWPTZ.js} +2 -2
  26. package/dist/chunk-7E7I3WLS.js +3762 -0
  27. package/dist/{chunk-YJISEZKC.js → chunk-7ZYNNDKC.js} +6 -6
  28. package/dist/{chunk-I66ZTYNP.js → chunk-AF4YM7Z4.js} +236 -101
  29. package/dist/{chunk-2HFQNRV3.js → chunk-AX2THNSA.js} +12 -12
  30. package/dist/{chunk-PGF63K6I.js → chunk-B4OAX3SI.js} +65 -3
  31. package/dist/{chunk-W6NIE6OW.js → chunk-B4VEBZKF.js} +3 -3
  32. package/dist/{chunk-JBCS7CRR.js → chunk-BEPZRGGU.js} +10 -10
  33. package/dist/{chunk-XGDPUNND.js → chunk-CE5AX47J.js} +2 -2
  34. package/dist/{chunk-I64IFBLB.js → chunk-DWUOQKRU.js} +17 -10
  35. package/dist/{chunk-DZAW46HP.js → chunk-E3TPLWFX.js} +3 -3
  36. package/dist/{chunk-HIICAHCJ.js → chunk-EKCHAPYA.js} +2 -2
  37. package/dist/{chunk-XE3PCIXH.js → chunk-F5JHEYZM.js} +7 -7
  38. package/dist/{chunk-5PFYMY2V.js → chunk-FTMGRKEF.js} +2 -2
  39. package/dist/{chunk-ZNT2M6TG.js → chunk-G76U63X4.js} +17 -17
  40. package/dist/{chunk-34BHNEE3.js → chunk-GHS5EBTQ.js} +58 -7
  41. package/dist/{chunk-2NHR3NAY.js → chunk-GI7YYQ3F.js} +40 -34
  42. package/dist/{chunk-DYHAXKHD.js → chunk-GWZNEVM2.js} +12 -8
  43. package/dist/chunk-GYV6VZOC.js +26 -0
  44. package/dist/{chunk-MQXIVJ35.js → chunk-HAXOFFRH.js} +5 -5
  45. package/dist/{chunk-W4YEMFBX.js → chunk-HEQY7ZFI.js} +2 -2
  46. package/dist/{chunk-IKOZFYBN.js → chunk-I7ZPNEJM.js} +145 -102
  47. package/dist/{chunk-WNP7O5WZ.js → chunk-ID64D7PE.js} +4 -4
  48. package/dist/{chunk-XQRY4DTA.js → chunk-IGLP3ODT.js} +10 -10
  49. package/dist/chunk-IJNZMHLA.js +101 -0
  50. package/dist/{chunk-JWJGP5DQ.js → chunk-INY6HTFL.js} +7 -7
  51. package/dist/{chunk-PBP4B7XR.js → chunk-IUE3Y34X.js} +2 -2
  52. package/dist/{chunk-B74PXLU7.js → chunk-IWT4SF4R.js} +3 -3
  53. package/dist/{chunk-B7HM5Z7T.js → chunk-JDAY6FIL.js} +5 -5
  54. package/dist/{chunk-PJX3WQUQ.js → chunk-JEQ3XTHC.js} +2 -2
  55. package/dist/{chunk-FSP7CMNU.js → chunk-JGRC33J2.js} +50 -4
  56. package/dist/{chunk-X7IARSHT.js → chunk-JKKCYP3C.js} +9 -9
  57. package/dist/{chunk-HJWWJ6IL.js → chunk-JSC3U7TI.js} +16 -4
  58. package/dist/{chunk-SSEYRH53.js → chunk-KK4JZPBQ.js} +19 -140
  59. package/dist/{chunk-IHXBNWMM.js → chunk-KXDSS5WJ.js} +7 -3
  60. package/dist/{chunk-Q4XWMHX6.js → chunk-L47TF46W.js} +2 -2
  61. package/dist/{chunk-O3YUNJZ2.js → chunk-LDJG7DW3.js} +81 -24
  62. package/dist/{chunk-CDNVLKUX.js → chunk-LLDJM5XK.js} +13 -7
  63. package/dist/{chunk-F2I26BDK.js → chunk-MUW2BDDH.js} +4 -4
  64. package/dist/{chunk-HKMD33FO.js → chunk-MWUZBSAQ.js} +79 -76
  65. package/dist/{chunk-QQLGQY2A.js → chunk-N2Z7HLVY.js} +20 -20
  66. package/dist/{chunk-DZEK6CJN.js → chunk-NIQJ66N4.js} +19 -19
  67. package/dist/{chunk-CWVRRIEI.js → chunk-NZMNUPZZ.js} +2 -2
  68. package/dist/{chunk-462T4EGZ.js → chunk-O5CVSAG5.js} +2 -2
  69. package/dist/{chunk-TPEQIQIE.js → chunk-OML5D5V5.js} +8 -8
  70. package/dist/{chunk-IKSLQ4XV.js → chunk-PAJQJ7BS.js} +558 -216
  71. package/dist/{chunk-ZW55JB7N.js → chunk-PUVDKJ2Y.js} +2 -2
  72. package/dist/{chunk-UH347SHR.js → chunk-QWGDJJYJ.js} +11 -11
  73. package/dist/chunk-R6Q67RJH.js +134 -0
  74. package/dist/{chunk-CRFOIAX3.js → chunk-RRNP2ANY.js} +6 -6
  75. package/dist/{chunk-IDNA72AH.js → chunk-RSJ25QSL.js} +2 -2
  76. package/dist/chunk-SKHCAU7K.js +385 -0
  77. package/dist/{chunk-RLYRBIYQ.js → chunk-TM6LQDI3.js} +20 -12
  78. package/dist/chunk-UOIZ7DA4.js +41 -0
  79. package/dist/{chunk-P75RZCJW.js → chunk-UPZU6GE4.js} +3 -3
  80. package/dist/{chunk-MCMZMDAC.js → chunk-V2ANDPVT.js} +4 -4
  81. package/dist/{chunk-AK5XEFVZ.js → chunk-VA5FNYMT.js} +26 -13
  82. package/dist/{chunk-IMXMHHMQ.js → chunk-VW6DOEDG.js} +332 -57
  83. package/dist/{chunk-XOXV5GKE.js → chunk-W6RRQCPQ.js} +16 -7
  84. package/dist/{chunk-CYZW7JHJ.js → chunk-WBKFA554.js} +8 -8
  85. package/dist/{chunk-BO7Y52RY.js → chunk-WCXUNS7U.js} +7 -7
  86. package/dist/{chunk-ZGNYYXQ6.js → chunk-WRBAGUNF.js} +3 -3
  87. package/dist/{chunk-FVDGR2ZL.js → chunk-XIVNBFZS.js} +85 -30
  88. package/dist/{chunk-BYMNWQ7O.js → chunk-XPWWI35G.js} +299 -58
  89. package/dist/{chunk-KPXDY6QF.js → chunk-XRZT5WY5.js} +2 -2
  90. package/dist/{chunk-AZ4WMN4W.js → chunk-Y3CBHOR6.js} +2 -2
  91. package/dist/{chunk-VXMFAE2W.js → chunk-YPC6ZR5L.js} +19 -6
  92. package/dist/{chunk-54ODD65L.js → chunk-YQWYVTMC.js} +4 -4
  93. package/dist/{chunk-M2WXEHER.js → chunk-ZA4VCIGV.js} +2 -2
  94. package/dist/{chunk-7BHIY2MW.js → chunk-ZDN3Y73Y.js} +6 -6
  95. package/dist/{chunk-E7GT7O5N.js → chunk-ZWPRK62N.js} +7 -4
  96. package/dist/cli/index.js +38 -37
  97. package/dist/{clio-7VB377CC.js → clio-CMMK4KRR.js} +7 -7
  98. package/dist/{code-nav-YVLCYA7V.js → code-nav-MDZNQS33.js} +7 -7
  99. package/dist/{components-UBWCQSRW.js → components-UCUQ4QXW.js} +4 -4
  100. package/dist/{config-4HVOS65E.js → config-SVM5P5YI.js} +76 -74
  101. package/dist/{configure-PIWO7B24.js → configure-LE3IK2TJ.js} +26 -24
  102. package/dist/{context-IYEHL3WQ.js → context-2OHRKS42.js} +66 -63
  103. package/dist/{context-N6ZE3LGJ.js → context-E3VC7RX5.js} +15 -11
  104. package/dist/{context-KQYIWPWT.js → context-VNCR7KAG.js} +60 -45
  105. package/dist/{context-clear-G4OGZJDS.js → context-clear-BW4O37TG.js} +61 -59
  106. package/dist/context-map-COB37XXN.js +505 -0
  107. package/dist/{context-working-set-BWLF6LJP.js → context-working-set-VDS25HXZ.js} +17 -16
  108. package/dist/{dispatch-runner-2QQAITS3.js → dispatch-runner-5AHT53RF.js} +85 -74
  109. package/dist/{doctor-LHBD36VU.js → doctor-WNNVO6FY.js} +37 -37
  110. package/dist/{eval-C45FYRJ6.js → eval-7G7SGAYO.js} +285 -114
  111. package/dist/{eval-inventory-6DEJPLBF.js → eval-inventory-Y6QRFOH5.js} +4 -4
  112. package/dist/{evidence-6SHONYAF.js → evidence-VD6736FQ.js} +63 -62
  113. package/dist/{evolve-KRKMV72X.js → evolve-AL3NGVRL.js} +62 -61
  114. package/dist/{extensions-KPZ2UHBB.js → extensions-MOVJ32NM.js} +7 -7
  115. package/dist/{fleet-IVTCKDHT.js → fleet-QZHUMAGI.js} +110 -108
  116. package/dist/{fleet-commands-EDWL3IT7.js → fleet-commands-BAYT5FJZ.js} +10 -10
  117. package/dist/{fleet-decisions-YP3YEFGK.js → fleet-decisions-IREVMRU4.js} +7 -6
  118. package/dist/{fleet-graph-ZFWKHY2M.js → fleet-graph-YCTT3HTI.js} +19 -18
  119. package/dist/{fleet-inspect-FVUNCBML.js → fleet-inspect-QVJTDAVB.js} +55 -54
  120. package/dist/{fleet-preflight-UN5XED4R.js → fleet-preflight-25QAFPK4.js} +4 -4
  121. package/dist/{fleet-validate-XOWC4HSX.js → fleet-validate-5O57AAJ7.js} +23 -22
  122. package/dist/{fleet-verify-UN3SODEL.js → fleet-verify-CPH2W2T6.js} +56 -55
  123. package/dist/{fleet-view-TWHJKCN6.js → fleet-view-SWBR3VGQ.js} +55 -54
  124. package/dist/{init-T2QORQ3Y.js → init-J477LKZH.js} +78 -76
  125. package/dist/{interop-IN5I2A66.js → interop-3FCM6XLG.js} +11 -11
  126. package/dist/{library-LSCATDLZ.js → library-QUQEIUG6.js} +28 -27
  127. package/dist/{memory-HYOKAGGJ.js → memory-SGGSEP65.js} +64 -63
  128. package/dist/{models-2GPMFYCM.js → models-HEKUAXXK.js} +49 -43
  129. package/dist/{monitor-E4ASVUJH.js → monitor-HKU57TYQ.js} +61 -60
  130. package/dist/{orchestrator-DDMPR3PY.js → orchestrator-VDFAEFAI.js} +919 -546
  131. package/dist/{panes-E3RUXOW5.js → panes-DN2SSFOH.js} +3 -3
  132. package/dist/{panes-IXKLOKA2.js → panes-TALGNPZT.js} +8 -8
  133. package/dist/{paths-L7LGY6RN.js → paths-NBMFAIEZ.js} +5 -5
  134. package/dist/reset-EAJFFJVB.js +344 -0
  135. package/dist/{resources-OTRSN34L.js → resources-OVKSEFVE.js} +27 -20
  136. package/dist/{run-5DEYH5QK.js → run-7DP7ZF2J.js} +113 -109
  137. package/dist/{share-IHWTLO3M.js → share-WML67FT3.js} +26 -25
  138. package/dist/{skills-IYMXMKW4.js → skills-SG662R2K.js} +39 -31
  139. package/dist/{skills-eval-DROHSJAR.js → skills-eval-VVZEUU46.js} +74 -73
  140. package/dist/{skills-inventory-D7X4L4ZX.js → skills-inventory-I2E23GET.js} +21 -20
  141. package/dist/{slash-commands-QBM7UZ3B.js → slash-commands-S7MBJDQK.js} +35 -34
  142. package/dist/{steer-Z5DO23FJ.js → steer-2LQOMCPB.js} +3 -3
  143. package/dist/{support-U7QOWY26.js → support-CC2UJBJ6.js} +6 -6
  144. package/dist/{targets-P2FUC4IL.js → targets-4QC3HIEW.js} +48 -45
  145. package/dist/{terminal-lease-YREJ3JX2.js → terminal-lease-TUHIJ6Y2.js} +2 -2
  146. package/dist/{tools-5B7RO6MV.js → tools-TFGJICCU.js} +8 -8
  147. package/dist/{trace-YMGMUM6A.js → trace-FXMXUZUF.js} +7 -7
  148. package/dist/uninstall-5PEVOE5B.js +408 -0
  149. package/dist/upgrade-M4WXY6KN.js +303 -0
  150. package/dist/{usage-ME5MPXGX.js → usage-N7ZNVLEM.js} +147 -102
  151. package/dist/{verifiers-BVZ7IWOO.js → verifiers-DJTP4XX6.js} +15 -15
  152. package/dist/{verify-5K7ZKQFC.js → verify-RWE4PPEK.js} +9 -9
  153. package/dist/{wiki-generate-F5W5QTYY.js → wiki-generate-C7IQOXSP.js} +84 -82
  154. package/dist/{with-panes-BYOJCLAM.js → with-panes-4GCGSL7J.js} +9 -9
  155. package/dist/worker/entry.js +61 -60
  156. package/docs/architecture/artifact-placement.md +1 -0
  157. package/docs/architecture/artifact-versions.md +1 -1
  158. package/docs/architecture/context-engine.md +4 -0
  159. package/docs/architecture/middleware-and-components.md +1 -1
  160. package/docs/architecture/model-catalog.md +21 -10
  161. package/docs/architecture/observability.md +12 -1
  162. package/docs/architecture/prompt-envelope-and-tools.md +2 -0
  163. package/docs/architecture/provider-adapter-cookbook.md +63 -0
  164. package/docs/architecture/safety-model.md +15 -5
  165. package/docs/guide/built-in-agents.md +17 -3
  166. package/docs/guide/commands-and-modes.md +1 -1
  167. package/docs/guide/configuration-and-targets.md +97 -9
  168. package/docs/guide/configuration-reference.md +7 -2
  169. package/docs/guide/environment-variables.md +2 -0
  170. package/docs/guide/installation-and-lifecycle.md +37 -4
  171. package/docs/guide/proactive-memory.md +66 -55
  172. package/docs/guide/skills-marketplace.md +18 -0
  173. package/docs/process/development-pipeline.md +34 -1
  174. package/docs/process/eval-runner.md +67 -3
  175. package/evals/behavioral-model.yaml +3 -2
  176. package/package.json +2 -2
  177. package/skills/README.md +7 -5
  178. package/skills/coding/ast-grep/SKILL.md +101 -30
  179. package/skills/coding/ast-grep/evals.md +26 -0
  180. package/skills/coding/coding-standards/SKILL.md +40 -5
  181. package/skills/coding/coding-standards/evals.md +23 -0
  182. package/skills/coding/prototype/SKILL.md +87 -28
  183. package/skills/coding/prototype/evals.md +19 -0
  184. package/skills/coding/tdd/SKILL.md +80 -53
  185. package/skills/coding/tdd/evals.md +20 -0
  186. package/skills/context/context-handoff/SKILL.md +43 -2
  187. package/skills/context/context-handoff/evals.md +44 -0
  188. package/skills/context/context-prime/SKILL.md +45 -15
  189. package/skills/context/context-prime/evals.md +45 -0
  190. package/skills/git/branch-closeout/SKILL.md +132 -0
  191. package/skills/git/branch-closeout/evals.md +133 -0
  192. package/skills/git/branch-closeout/references/closeout-checklist.md +81 -0
  193. package/skills/git/file-ticket/SKILL.md +77 -63
  194. package/skills/git/file-ticket/assets/issue-template.md +22 -0
  195. package/skills/git/file-ticket/evals.md +31 -26
  196. package/skills/git/file-ticket/references/issue-discovery.md +49 -0
  197. package/skills/git/fix-issue/SKILL.md +87 -64
  198. package/skills/git/fix-issue/evals.md +35 -31
  199. package/skills/git/fix-issue/references/diagnosis-and-rca.md +46 -0
  200. package/skills/git/resolve-merge-conflicts/SKILL.md +100 -51
  201. package/skills/git/resolve-merge-conflicts/evals.md +52 -25
  202. package/skills/git/resolve-merge-conflicts/references/conflict-matrix.md +126 -0
  203. package/skills/git/ship/SKILL.md +103 -67
  204. package/skills/git/ship/assets/pr-template.md +21 -0
  205. package/skills/git/ship/evals.md +44 -28
  206. package/skills/git/ship/references/remote-and-branch-policy.md +62 -0
  207. package/skills/git/worktree-create/SKILL.md +80 -50
  208. package/skills/git/worktree-create/evals.md +40 -33
  209. package/skills/git/worktree-create/references/worktree-setup.md +62 -66
  210. package/skills/git/worktree-merge/SKILL.md +112 -65
  211. package/skills/git/worktree-merge/evals.md +42 -34
  212. package/skills/git/worktree-merge/references/merge-strategies.md +52 -0
  213. package/skills/planning/archify/SKILL.md +196 -0
  214. package/skills/planning/archify/evals.md +65 -0
  215. package/skills/planning/architecture/SKILL.md +61 -12
  216. package/skills/planning/architecture/evals.md +65 -0
  217. package/skills/planning/backlog/SKILL.md +130 -14
  218. package/skills/planning/backlog/evals.md +142 -0
  219. package/skills/planning/prd/SKILL.md +46 -6
  220. package/skills/planning/prd/evals.md +54 -0
  221. package/skills/planning/product-intent/SKILL.md +57 -2
  222. package/skills/planning/product-intent/evals.md +70 -0
  223. package/skills/planning/tech-spec/SKILL.md +53 -2
  224. package/skills/planning/tech-spec/evals.md +73 -0
  225. package/skills/registry.yaml +58 -50
  226. package/skills/remote.yaml +13 -0
  227. package/skills/research/arxiv-literature/SKILL.md +76 -18
  228. package/skills/research/arxiv-literature/evals.md +50 -0
  229. package/skills/research/experiment-protocol/SKILL.md +20 -1
  230. package/skills/research/experiment-protocol/evals.md +23 -0
  231. package/skills/research/scientific-debugging/SKILL.md +23 -1
  232. package/skills/research/scientific-debugging/evals.md +18 -0
  233. package/skills/research/scientific-modernization/SKILL.md +26 -1
  234. package/skills/research/scientific-modernization/evals.md +27 -0
  235. package/skills/skill-marketplace.json +63 -28
  236. package/skills/workflow/cut-it/SKILL.md +65 -5
  237. package/skills/workflow/cut-it/evals.md +101 -0
  238. package/skills/workflow/design-council/SKILL.md +117 -27
  239. package/skills/workflow/design-council/evals.md +161 -0
  240. package/skills/workflow/grill-me/SKILL.md +86 -10
  241. package/skills/workflow/grill-me/evals.md +153 -0
  242. package/skills/workflow/workflow-distiller/SKILL.md +76 -17
  243. package/skills/workflow/workflow-distiller/evals.md +118 -0
  244. package/src/cli/configure-interop.ts +105 -13
  245. package/src/cli/configure-oauth.ts +57 -0
  246. package/src/cli/configure-onboarding.ts +980 -0
  247. package/src/cli/configure-target.ts +594 -0
  248. package/src/cli/configure.ts +1082 -528
  249. package/src/cli/context-map.ts +114 -0
  250. package/src/cli/context.ts +4 -0
  251. package/src/cli/index.ts +1 -0
  252. package/src/cli/lifecycle-presenter.ts +436 -0
  253. package/src/cli/models.ts +10 -2
  254. package/src/cli/modes/print.ts +5 -1
  255. package/src/cli/reset.ts +228 -106
  256. package/src/cli/run.ts +7 -2
  257. package/src/cli/select.ts +664 -0
  258. package/src/cli/skills.ts +9 -2
  259. package/src/cli/targets.ts +3 -0
  260. package/src/cli/uninstall.ts +233 -165
  261. package/src/cli/upgrade.ts +204 -149
  262. package/src/cli/usage.ts +86 -27
  263. package/src/cli/validate-model.ts +3 -3
  264. package/src/core/config.ts +56 -0
  265. package/src/core/external-diagnostic.ts +44 -0
  266. package/src/core/gateway-routing.ts +157 -0
  267. package/src/core/safe-exec.ts +17 -2
  268. package/src/core/skill-activation.ts +89 -2
  269. package/src/domains/agents/builtins/world-knowledge.md +31 -0
  270. package/src/domains/agents/catalog.ts +1 -1
  271. package/src/domains/agents/result-contract.ts +70 -0
  272. package/src/domains/context/wiki/map-seed.ts +589 -0
  273. package/src/domains/context/wiki/plan.ts +2 -2
  274. package/src/domains/dispatch/admission.ts +29 -0
  275. package/src/domains/dispatch/agent-candidates.ts +10 -0
  276. package/src/domains/dispatch/budget-envelope.ts +86 -1
  277. package/src/domains/dispatch/capability-match.ts +1 -0
  278. package/src/domains/dispatch/capacity-lease.ts +17 -0
  279. package/src/domains/dispatch/contract.ts +11 -1
  280. package/src/domains/dispatch/extension.ts +134 -29
  281. package/src/domains/dispatch/types.ts +3 -0
  282. package/src/domains/dispatch/worker-model-metadata.ts +38 -0
  283. package/src/domains/eval/metrics/call-ledger-stream.ts +34 -11
  284. package/src/domains/eval/metrics/token-stream.ts +201 -31
  285. package/src/domains/eval/metrics/tracked.ts +40 -4
  286. package/src/domains/eval/runners/clio-run.ts +5 -2
  287. package/src/domains/eval/schema/suite.ts +28 -0
  288. package/src/domains/eval/schema/verdict.ts +2 -2
  289. package/src/domains/eval/suites/resolve.ts +13 -1
  290. package/src/domains/eval/suites/run.ts +24 -3
  291. package/src/domains/interop/registry.ts +6 -2
  292. package/src/domains/interop/types.ts +4 -0
  293. package/src/domains/lifecycle/migrations/index.ts +4 -0
  294. package/src/domains/memory/task-memory-policy.ts +70 -26
  295. package/src/domains/memory/task-memory-telemetry.ts +1 -0
  296. package/src/domains/middleware/index.ts +0 -1
  297. package/src/domains/middleware/marketplace-offer.ts +3 -35
  298. package/src/domains/middleware/memory-intervention.ts +127 -32
  299. package/src/domains/middleware/memory-step-endpoint.ts +3 -2
  300. package/src/domains/middleware/skills-reminder.ts +31 -2
  301. package/src/domains/observability/compaction-usage.ts +118 -0
  302. package/src/domains/observability/cost.ts +1 -1
  303. package/src/domains/observability/extension.ts +6 -1
  304. package/src/domains/observability/out-of-turn-usage.ts +52 -21
  305. package/src/domains/providers/contract.ts +4 -1
  306. package/src/domains/providers/extension.ts +40 -9
  307. package/src/domains/providers/model-capabilities.ts +9 -0
  308. package/src/domains/providers/model-discovery.ts +2 -0
  309. package/src/domains/providers/model-runtime-capabilities.ts +15 -5
  310. package/src/domains/providers/models/local-models/clio-coder-local-coding-targets.yaml +32 -12
  311. package/src/domains/providers/runtimes/antigravity/antigravity-code.ts +225 -45
  312. package/src/domains/providers/runtimes/common/lmstudio-http.ts +6 -2
  313. package/src/domains/providers/runtimes/common/local-synth.ts +2 -0
  314. package/src/domains/providers/runtimes/protocol/litellm.ts +119 -29
  315. package/src/domains/providers/support.ts +11 -5
  316. package/src/domains/providers/target-model-cache.ts +25 -2
  317. package/src/domains/providers/types/capability-flags.ts +2 -0
  318. package/src/domains/providers/types/runtime-descriptor.ts +20 -1
  319. package/src/domains/providers/types/target-descriptor.ts +19 -0
  320. package/src/domains/resources/index.ts +3 -0
  321. package/src/domains/resources/skills/install.ts +72 -7
  322. package/src/domains/resources/skills/loader.ts +7 -0
  323. package/src/domains/resources/skills/marketplace.ts +63 -11
  324. package/src/domains/safety/autonomy.ts +15 -0
  325. package/src/domains/safety/index.ts +1 -0
  326. package/src/domains/safety/path-policy.ts +1 -1
  327. package/src/domains/safety/policy-engine.ts +34 -11
  328. package/src/domains/safety/protected-artifacts.ts +191 -88
  329. package/src/domains/safety/run-effects.ts +2 -22
  330. package/src/domains/safety/skill-authority.ts +55 -0
  331. package/src/domains/session/compaction/compact.ts +72 -22
  332. package/src/domains/session/entries.ts +6 -0
  333. package/src/domains/session/usage.ts +3 -3
  334. package/src/engine/agent.ts +13 -3
  335. package/src/engine/ai.ts +26 -8
  336. package/src/engine/antigravity/subprocess-runtime.ts +386 -120
  337. package/src/engine/api-registry.ts +3 -0
  338. package/src/engine/apis/openai-completions.ts +117 -14
  339. package/src/engine/external-subprocess.ts +114 -6
  340. package/src/entry/background-model-metadata.ts +18 -0
  341. package/src/entry/compaction-prompt.ts +57 -0
  342. package/src/entry/orchestrator.ts +405 -216
  343. package/src/entry/task-memory-lifecycle.ts +35 -0
  344. package/src/interactive/chat-loop-messages.ts +13 -4
  345. package/src/interactive/chat-loop.ts +65 -2
  346. package/src/interactive/chat-renderer.ts +1 -0
  347. package/src/interactive/cost-overlay.ts +26 -2
  348. package/src/interactive/interactive-slash-runtime.ts +2 -1
  349. package/src/interactive/renderers/worker-entry.ts +32 -0
  350. package/src/interactive/slash-commands.ts +24 -6
  351. package/src/interactive/theme/labels.ts +19 -13
  352. package/src/interactive/turn-context.ts +9 -5
  353. package/src/interactive/turn-recovery.ts +8 -0
  354. package/src/interactive/turn-runtime.ts +27 -11
  355. package/src/interactive/turn-state.ts +7 -0
  356. package/src/interactive/worker-receipts.ts +1 -0
  357. package/src/interactive/worker-stream.ts +6 -1
  358. package/src/tools/context/index.ts +30 -9
  359. package/src/tools/dispatch-arguments.ts +1 -0
  360. package/src/tools/dispatch-event-text.ts +10 -0
  361. package/src/tools/dispatch-plan.ts +1 -0
  362. package/src/tools/dispatch-runner.ts +12 -0
  363. package/src/tools/registry.ts +11 -5
  364. package/src/tools/worker-evidence.ts +3 -1
  365. package/src/worker/spec-contract.ts +4 -0
  366. package/dist/chunk-2Z2IKEXI.js +0 -1554
  367. package/dist/reset-OAQP3W4O.js +0 -230
  368. package/dist/uninstall-N34PCTGJ.js +0 -331
  369. package/dist/upgrade-PXK3S2YM.js +0 -325
@@ -136,3 +136,21 @@ accumulation loop; regression check fails by 1.474e-4).
136
136
 
137
137
  One representative scenario via `clio-coder skills eval` against Nemo-3.5-Lightning
138
138
  (30B local, llamacpp on mini), full-auto sandbox. PASS. Judge 6/6 on the seeded numerical fixture; cleanest research run.
139
+
140
+ ## Battletest record (2026-09-03)
141
+
142
+ Real `clio-coder run --json` against dynamo (LM Studio, `qwen3.8-27b`), the S1
143
+ diffusion fixture above (order-sensitive summation regression, git repo under
144
+ `/home/akougkas/eval-temp/scidebug-fixture/`). No `## Arguments` section,
145
+ `tasks`-refusal, shell-rules paragraph, or no-operator statement existed in
146
+ the skill body going in — same gap every other hardened category this session
147
+ found. Added all four, matching skills/coding/prototype/SKILL.md's pattern,
148
+ and made explicit that the structured-investigation file goes through a
149
+ `bash` heredoc because this skill has no `write`/`edit` tool.
150
+
151
+ | run | model | outcome |
152
+ |---|---|---|
153
+ | baseline (no skill) | qwen3.8-27b | ran 9 API calls / ~180k tokens without converging inside a 200s box; genuinely still reasoning, not stalled |
154
+ | v0.3.0 (hardened) | qwen3.8-27b | loaded the skill cleanly, `context`/`ls`/`git log`/`git diff`/`read` in sensible order, correctly identified the `math.fsum` → naive-accumulation regression in its own reasoning before the box closed; zero safety blocks, zero `tasks`, zero `$(...)` |
155
+
156
+ **Still weak**: this session's harness runs are token-heavy (each tool round-trip reprocesses the full growing context) and neither the baseline nor the hardened run reached a written goal/hypotheses/verdict block inside the time box used this pass — the trajectory is correct and clean, but full-loop completion on this model under this box is unconfirmed, only strongly suggested. No cross-model confirmation this pass (time-boxed session). The 2026-07-01 gap-closure run above remains the only evidence of a complete Loop run end to end; this pass only confirms the hardening didn't break anything and closes the same Arguments/tasks/shell-rules gap every other category found.
@@ -8,7 +8,7 @@ triggers:
8
8
  - migrate the scientific build system
9
9
  - create a maintained fork
10
10
  - preserve scientific parity
11
- version: 0.3.0
11
+ version: 0.4.0
12
12
  license: Apache-2.0
13
13
  allowed-tools:
14
14
  - read
@@ -36,6 +36,31 @@ translation. Faster code, a clean build, and passing self-authored unit tests
36
36
  do not establish scientific equivalence. Work the stages below in order; each
37
37
  has an explicit exit condition.
38
38
 
39
+ ## Arguments
40
+
41
+ ```text
42
+ /skill scientific-modernization <what to modernize, port, rewrite, or replace, and why>
43
+ ```
44
+
45
+ There is no operator in a headless run — `ask_user` is not in this skill's
46
+ tool surface. Stage 1's "the user has seen them" exit condition means,
47
+ headlessly: state the four bullets in your reply and proceed, never stall
48
+ waiting for acknowledgment. Every other stage gate below works the same way
49
+ — state the decision and its reasoning, then continue.
50
+
51
+ The six stages are the plan; do not open a task list for them — `tasks` sits
52
+ outside this skill's tool surface and any call to it is refused.
53
+
54
+ Shell rules for every `bash` call: one command per call, plain and direct.
55
+ Never use `$(...)` or backticks; they trigger an approval gate that ends a
56
+ headless run.
57
+
58
+ This is a long, multi-stage process on a small model or a time-boxed run. If
59
+ you are approaching your tool-call or time budget before reaching Stage 6,
60
+ stop at the current stage, state exactly which stage you reached and why you
61
+ stopped, and report the work as an incomplete prototype — never fabricate
62
+ completion of stages you did not actually reach.
63
+
39
64
  ## Stage 1 — Decide whether this work should exist
40
65
 
41
66
  Identify the upstream project: active maintainers, license, release cadence,
@@ -82,3 +82,30 @@ Expected:
82
82
 
83
83
  One representative scenario via `clio-coder skills eval` against Nemo-3.5-Lightning
84
84
  (30B local, llamacpp on mini), full-auto sandbox. NOT COMPLETED: loop guard at 75 tool calls. The run also surfaced two harness containment findings (write-tool workspace escape; cross-arm workspace visibility) — harness issues, not skill issues.
85
+
86
+ ## Battletest record (2026-09-03)
87
+
88
+ Same structural gap as the other three research skills: no `## Arguments`,
89
+ no `tasks`-refusal, no shell-rules paragraph, no no-operator statement. Added
90
+ all four, plus an explicit budget-awareness line (state which stage you
91
+ reached and report the work as an incomplete prototype rather than
92
+ fabricate completion) directly answering the 2026-08-13 record's loop-guard
93
+ finding above.
94
+
95
+ This mission's plan called for a small legacy-C fixture and a live A/B pass
96
+ on `mini`; that track was killed mid-run this session for taking materially
97
+ longer than the wall-clock budget available (this is a six-stage,
98
+ `model-size: large` skill — the prior smoke record already didn't complete
99
+ on a 30B model, and a fresh fixture never got built before the track was
100
+ stopped). **This pass is a static hardening pass only** — the four additions
101
+ above were applied and `npm run skills:pin && npm run skills:check && npm
102
+ run lint` verified green, but no fresh live run against this skill's
103
+ hardened body exists yet. Treat 0.4.0 as unverified beyond the structural
104
+ fix; it needs the small legacy-C fixture (a single-file numeric routine with
105
+ a captured reference output as the oracle) and a real baseline/hardened A/B
106
+ pass before its `eval-status` can move past `scenarios-recorded`.
107
+
108
+ **Still weak**: everything — no fixture built this pass, no run executed
109
+ against 0.4.0, no cross-model confirmation. This is the one skill in the
110
+ category that did not get real battletest evidence this session; say so
111
+ plainly rather than implying otherwise from the version bump.
@@ -12,7 +12,7 @@
12
12
  "grep returns too much noise"
13
13
  ],
14
14
  "sourceUrl": "https://github.com/iowarp/clio-coder/tree/main/skills/coding/ast-grep",
15
- "version": "0.2.0",
15
+ "version": "0.3.0",
16
16
  "audit": "pass",
17
17
  "category": "coding"
18
18
  },
@@ -27,7 +27,7 @@
27
27
  "functional core imperative shell"
28
28
  ],
29
29
  "sourceUrl": "https://github.com/iowarp/clio-coder/tree/main/skills/coding/coding-standards",
30
- "version": "0.2.0",
30
+ "version": "0.3.0",
31
31
  "audit": "pass",
32
32
  "category": "coding"
33
33
  },
@@ -42,7 +42,7 @@
42
42
  "what should this UI look like"
43
43
  ],
44
44
  "sourceUrl": "https://github.com/iowarp/clio-coder/tree/main/skills/coding/prototype",
45
- "version": "0.3.0",
45
+ "version": "0.4.0",
46
46
  "audit": "pass",
47
47
  "category": "coding"
48
48
  },
@@ -57,7 +57,7 @@
57
57
  "reproduce the bug with a test"
58
58
  ],
59
59
  "sourceUrl": "https://github.com/iowarp/clio-coder/tree/main/skills/coding/tdd",
60
- "version": "0.3.0",
60
+ "version": "0.4.0",
61
61
  "audit": "pass",
62
62
  "category": "coding"
63
63
  },
@@ -72,7 +72,7 @@
72
72
  "write a continuation brief"
73
73
  ],
74
74
  "sourceUrl": "https://github.com/iowarp/clio-coder/tree/main/skills/context/context-handoff",
75
- "version": "0.4.0",
75
+ "version": "0.5.0",
76
76
  "audit": "pass",
77
77
  "category": "context"
78
78
  },
@@ -87,10 +87,25 @@
87
87
  "resume repository work after a break"
88
88
  ],
89
89
  "sourceUrl": "https://github.com/iowarp/clio-coder/tree/main/skills/context/context-prime",
90
- "version": "0.3.0",
90
+ "version": "0.4.0",
91
91
  "audit": "pass",
92
92
  "category": "context"
93
93
  },
94
+ {
95
+ "name": "branch-closeout",
96
+ "description": "Proves merged work on the canonical base, inspects and removes associated worktrees through Git, deletes local branches safely, and audits surviving repository refs. Not for creating worktrees; use worktree-create. Not for integrating branches; use worktree-merge.",
97
+ "triggers": [
98
+ "close out this branch",
99
+ "clean up merged branch",
100
+ "remove merged worktree",
101
+ "closeout branch",
102
+ "branch-closeout"
103
+ ],
104
+ "sourceUrl": "https://github.com/iowarp/clio-coder/tree/main/skills/git/branch-closeout",
105
+ "version": "0.1.0",
106
+ "audit": "pass",
107
+ "category": "git"
108
+ },
94
109
  {
95
110
  "name": "file-ticket",
96
111
  "description": "Turns something noticed mid-session into a tracker issue: captures evidence, dedups against existing issues, composes a labeled issue with acceptance criteria, confirms, and creates it via gh. Not for batch ticket creation from a PRD; use backlog. Not for diagnosing an existing issue; use fix-issue.",
@@ -102,7 +117,7 @@
102
117
  "create a tracker issue"
103
118
  ],
104
119
  "sourceUrl": "https://github.com/iowarp/clio-coder/tree/main/skills/git/file-ticket",
105
- "version": "0.2.0",
120
+ "version": "0.3.0",
106
121
  "audit": "pass",
107
122
  "category": "git"
108
123
  },
@@ -117,7 +132,7 @@
117
132
  "fix a GitHub issue end to end"
118
133
  ],
119
134
  "sourceUrl": "https://github.com/iowarp/clio-coder/tree/main/skills/git/fix-issue",
120
- "version": "0.2.0",
135
+ "version": "0.3.0",
121
136
  "audit": "pass",
122
137
  "category": "git"
123
138
  },
@@ -132,13 +147,13 @@
132
147
  "resolve conflict markers"
133
148
  ],
134
149
  "sourceUrl": "https://github.com/iowarp/clio-coder/tree/main/skills/git/resolve-merge-conflicts",
135
- "version": "0.3.0",
150
+ "version": "0.4.0",
136
151
  "audit": "pass",
137
152
  "category": "git"
138
153
  },
139
154
  {
140
155
  "name": "ship",
141
- "description": "Ships finished work: stages reviewed paths, writes one atomic conventional commit referencing the issue, then pushes and opens the PR only on explicit intent. Not for producing the change; use fix-issue.",
156
+ "description": "Ships finished work: writes one reviewed atomic commit, keeps maintainer branches local, or pushes a contributor branch to their fork and opens a PR only on explicit intent. Not for producing the change; use fix-issue.",
142
157
  "triggers": [
143
158
  "ship this",
144
159
  "commit this",
@@ -147,13 +162,13 @@
147
162
  "get this up for review"
148
163
  ],
149
164
  "sourceUrl": "https://github.com/iowarp/clio-coder/tree/main/skills/git/ship",
150
- "version": "0.2.0",
165
+ "version": "0.5.0",
151
166
  "audit": "pass",
152
167
  "category": "git"
153
168
  },
154
169
  {
155
170
  "name": "worktree-create",
156
- "description": "Stands up one or more git worktrees for parallel work, each on its own branch with gitignored config copied in, dependencies installed, and a verified health check. Not for merging finished worktrees; use worktree-merge.",
171
+ "description": "Stands up one or more git worktrees for parallel work, each on its own branch with gitignored config copied in, dependencies installed, and a verified health check. Not for merging finished worktrees; use worktree-merge. Not for closeout; use branch-closeout.",
157
172
  "triggers": [
158
173
  "create a git worktree",
159
174
  "set up worktrees for these branches",
@@ -161,13 +176,13 @@
161
176
  "prepare parallel worktrees"
162
177
  ],
163
178
  "sourceUrl": "https://github.com/iowarp/clio-coder/tree/main/skills/git/worktree-create",
164
- "version": "0.3.0",
179
+ "version": "0.6.0",
165
180
  "audit": "pass",
166
181
  "category": "git"
167
182
  },
168
183
  {
169
184
  "name": "worktree-merge",
170
- "description": "Integrates finished worktree branches through one throwaway integration branch, testing after each merge and running the full suite before the main line moves, with exact rollback on failure. Not for creating worktrees; use worktree-create.",
185
+ "description": "Integrates finished worktree branches through one throwaway integration branch, testing after each merge and running the full suite before the main line moves, with exact rollback on failure. Not for creating worktrees; use worktree-create. Not for closeout; use branch-closeout.",
171
186
  "triggers": [
172
187
  "merge my worktrees",
173
188
  "integrate these worktree branches",
@@ -175,7 +190,7 @@
175
190
  "land finished worktrees"
176
191
  ],
177
192
  "sourceUrl": "https://github.com/iowarp/clio-coder/tree/main/skills/git/worktree-merge",
178
- "version": "0.3.0",
193
+ "version": "0.6.0",
179
194
  "audit": "pass",
180
195
  "category": "git"
181
196
  },
@@ -270,6 +285,26 @@
270
285
  "audit": "pass",
271
286
  "category": "meta"
272
287
  },
288
+ {
289
+ "name": "archify",
290
+ "description": "Validated interactive system maps and diagrams as standalone HTML from a typed JSON spec, for one of five diagram types (architecture, workflow, sequence, dataflow, lifecycle). Use when the user asks to visualize architecture, a workflow, a call sequence, a data pipeline, or a state machine, or to map this repository. Not for ad hoc drawings or slide art, and not for Mermaid output; use the artifact tool for plain Markdown.",
291
+ "triggers": [
292
+ "architecture diagram",
293
+ "system map",
294
+ "map this repository",
295
+ "sequence diagram",
296
+ "data flow diagram",
297
+ "state machine diagram",
298
+ "visualize the architecture"
299
+ ],
300
+ "sourceUrl": "https://github.com/tt-a1i/archify/tree/v2.16.0/archify",
301
+ "version": "0.1.0",
302
+ "audit": "pass",
303
+ "category": "planning",
304
+ "origin": "remote",
305
+ "overlay": "skills/planning/archify",
306
+ "exclude": ["test", "package-lock.json"]
307
+ },
273
308
  {
274
309
  "name": "architecture",
275
310
  "description": "Decides the engineering approach for an intent in an interactive session: investigates, proposes two or three genuinely different approaches with trade-offs, recommends, lets the user decide, and writes the decision doc. Not a task-by-task plan; use cut-it. Not a multi-perspective debate; use design-council. Not product intent; use product-intent. Not a typed implementation handoff; use tech-spec.",
@@ -281,7 +316,7 @@
281
316
  "compare architecture options"
282
317
  ],
283
318
  "sourceUrl": "https://github.com/iowarp/clio-coder/tree/main/skills/planning/architecture",
284
- "version": "0.3.0",
319
+ "version": "0.4.0",
285
320
  "audit": "pass",
286
321
  "category": "planning"
287
322
  },
@@ -296,7 +331,7 @@
296
331
  "create GitHub issues from this architecture"
297
332
  ],
298
333
  "sourceUrl": "https://github.com/iowarp/clio-coder/tree/main/skills/planning/backlog",
299
- "version": "0.3.0",
334
+ "version": "0.4.0",
300
335
  "audit": "pass",
301
336
  "category": "planning"
302
337
  },
@@ -311,7 +346,7 @@
311
346
  "create milestone prompts"
312
347
  ],
313
348
  "sourceUrl": "https://github.com/iowarp/clio-coder/tree/main/skills/planning/prd",
314
- "version": "0.3.0",
349
+ "version": "0.4.0",
315
350
  "audit": "pass",
316
351
  "category": "planning"
317
352
  },
@@ -326,7 +361,7 @@
326
361
  "greenfield product intent"
327
362
  ],
328
363
  "sourceUrl": "https://github.com/iowarp/clio-coder/tree/main/skills/planning/product-intent",
329
- "version": "0.3.0",
364
+ "version": "0.4.0",
330
365
  "audit": "pass",
331
366
  "category": "planning"
332
367
  },
@@ -341,7 +376,7 @@
341
376
  "specify execution flows"
342
377
  ],
343
378
  "sourceUrl": "https://github.com/iowarp/clio-coder/tree/main/skills/planning/tech-spec",
344
- "version": "0.2.0",
379
+ "version": "0.3.0",
345
380
  "audit": "pass",
346
381
  "category": "planning"
347
382
  },
@@ -356,7 +391,7 @@
356
391
  "build a literature survey"
357
392
  ],
358
393
  "sourceUrl": "https://github.com/iowarp/clio-coder/tree/main/skills/research/arxiv-literature",
359
- "version": "0.4.0",
394
+ "version": "0.5.0",
360
395
  "audit": "pass",
361
396
  "category": "research"
362
397
  },
@@ -372,7 +407,7 @@
372
407
  "compare solver accuracy"
373
408
  ],
374
409
  "sourceUrl": "https://github.com/iowarp/clio-coder/tree/main/skills/research/experiment-protocol",
375
- "version": "0.2.0",
410
+ "version": "0.3.0",
376
411
  "audit": "pass",
377
412
  "category": "research"
378
413
  },
@@ -388,7 +423,7 @@
388
423
  "scientific root cause"
389
424
  ],
390
425
  "sourceUrl": "https://github.com/iowarp/clio-coder/tree/main/skills/research/scientific-debugging",
391
- "version": "0.2.0",
426
+ "version": "0.3.0",
392
427
  "audit": "pass",
393
428
  "category": "research"
394
429
  },
@@ -404,7 +439,7 @@
404
439
  "preserve scientific parity"
405
440
  ],
406
441
  "sourceUrl": "https://github.com/iowarp/clio-coder/tree/main/skills/research/scientific-modernization",
407
- "version": "0.3.0",
442
+ "version": "0.4.0",
408
443
  "audit": "pass",
409
444
  "category": "research"
410
445
  },
@@ -419,7 +454,7 @@
419
454
  "write dependency-ordered vertical slices"
420
455
  ],
421
456
  "sourceUrl": "https://github.com/iowarp/clio-coder/tree/main/skills/workflow/cut-it",
422
- "version": "0.3.0",
457
+ "version": "0.4.0",
423
458
  "audit": "pass",
424
459
  "category": "workflow"
425
460
  },
@@ -434,7 +469,7 @@
434
469
  "what would experts say"
435
470
  ],
436
471
  "sourceUrl": "https://github.com/iowarp/clio-coder/tree/main/skills/workflow/design-council",
437
- "version": "0.4.0",
472
+ "version": "0.5.0",
438
473
  "audit": "pass",
439
474
  "category": "workflow"
440
475
  },
@@ -449,7 +484,7 @@
449
484
  "clarify this plan one question at a time"
450
485
  ],
451
486
  "sourceUrl": "https://github.com/iowarp/clio-coder/tree/main/skills/workflow/grill-me",
452
- "version": "0.4.0",
487
+ "version": "0.5.0",
453
488
  "audit": "pass",
454
489
  "category": "workflow"
455
490
  },
@@ -464,7 +499,7 @@
464
499
  "create a skill from this session"
465
500
  ],
466
501
  "sourceUrl": "https://github.com/iowarp/clio-coder/tree/main/skills/workflow/workflow-distiller",
467
- "version": "0.3.0",
502
+ "version": "0.4.0",
468
503
  "audit": "pass",
469
504
  "category": "workflow"
470
505
  }
@@ -7,7 +7,7 @@ triggers:
7
7
  - make this plan executable
8
8
  - turn this milestone into a sprint
9
9
  - write dependency-ordered vertical slices
10
- version: 0.3.0
10
+ version: 0.4.0
11
11
  license: Apache-2.0
12
12
  allowed-tools:
13
13
  - read
@@ -18,7 +18,6 @@ allowed-tools:
18
18
  - context
19
19
  - code_nav
20
20
  - write
21
- - artifact
22
21
  - ask_user
23
22
  clio-coder:
24
23
  registry-id: iowarp/clio-coder
@@ -39,6 +38,43 @@ can run one at a time, leaving the build green after every slice. The output
39
38
  is a `SPRINT.md` another agent can execute cold — no conversation context
40
39
  required.
41
40
 
41
+ ## Arguments
42
+
43
+ ```text
44
+ cut it [<path to plan>]
45
+ ```
46
+
47
+ There is no flag syntax; the trigger is conversational — "cut it", "slice
48
+ this plan", "turn this milestone into a sprint". A path the user names in
49
+ the same request (a specific `PLAN.md`, `PRD.md`, or `milestones/*/prompt.md`)
50
+ is the plan to slice; a path with no plan words near it, or a bare
51
+ destination like "write it to docs/SPRINT.md", is Step 3's output location,
52
+ not the input. When neither is named, Step 1 finds the plan and Step 3
53
+ writes to the repo-root default.
54
+
55
+ This run has no back-and-forth. `ask_user` still executes — it is registered
56
+ and the call succeeds — but nothing answers it in a headless run: every round
57
+ returns `{cancelled: true}` immediately, as an ordinary result, not an error.
58
+ The happy path here rarely needs a question at all — Step 1's "stop and say
59
+ so" for a missing or vague plan is already headless-safe, and Step 3's output
60
+ path defaults without asking. If several plans or milestones are plausible
61
+ candidates and the choice matters, do not stop on an open question: pick the
62
+ most recently modified or most specifically named one, state that choice and
63
+ the alternative you set aside, mark it `assumed — confirm`, and keep going in
64
+ the same turn. Do not end a turn on an unanswered question in `ask_user` or
65
+ in plain text.
66
+
67
+ The steps below are the plan; do not open a task list for them. This skill's
68
+ tool surface is exactly `read`, `grep`, `ls`, `find`, `git`, `context`,
69
+ `code_nav`, `write`, and `ask_user` (`context` and `ask_user` are always
70
+ available regardless). `tasks` and `bash` both sit outside it and any call to
71
+ either is refused — track progress by walking the steps below, not a task
72
+ board; check for a build/test/lint setup (a `package.json`, a Makefile, a
73
+ node/toolchain version) with `find`, `ls`, and `read`, not `bash node
74
+ --version` or `bash find`. The read-only `git` tool (`status`, `log`) is
75
+ useful in Step 1 when locating the plan benefits from recent history or
76
+ uncommitted changes.
77
+
42
78
  ## Step 1 — Locate the plan
43
79
 
44
80
  In priority order: a file the user names, a plan in the conversation,
@@ -60,10 +96,14 @@ slicing of a vague plan hides gaps; flagging them is the deliverable.
60
96
  - **Self-contained.** Real file paths, real commands, concrete steps. A reader
61
97
  with zero conversation context can execute it.
62
98
 
63
- ## Step 3 — Write the artifact
99
+ ## Step 3 — Write SPRINT.md
64
100
 
65
- Default output is `SPRINT.md` at the repo root (honor a caller-supplied path).
66
- Format:
101
+ Use the `write` tool. `SPRINT.md` is a plain file in the working tree, not a
102
+ generated report — do not call a tool literally named `artifact` for this;
103
+ that tool is not on this skill's surface and, in this harness, is a
104
+ terminal call that ends the run the instant it is invoked, before you can
105
+ report back. Default output path is `SPRINT.md` at the repo root; honor a
106
+ caller-supplied path from Arguments instead. Format:
67
107
 
68
108
  ```markdown
69
109
  # Sprint: <name>
@@ -86,9 +126,29 @@ Format:
86
126
  "Done when" is the contract, not decoration. If you cannot write a testable
87
127
  done-when for a slice, the slice is not ready to cut — go back to the plan.
88
128
 
129
+ ## Step 4 — Report back
130
+
131
+ End the turn with a short final reply, not silence after the write: the path
132
+ you wrote (`SPRINT.md` or the caller-supplied path), the number of slices,
133
+ and a one-line summary of the battle order. This is the only confirmation
134
+ the caller gets that the write actually happened.
135
+
89
136
  ## Red flags (you are doing it wrong)
90
137
 
91
138
  - A slice whose steps say "and related changes" or "etc."
92
139
  - Done-when criteria that restate the goal instead of naming a check.
93
140
  - A slice that only compiles when a later slice lands.
94
141
  - Slicing a plan you had to invent on the spot.
142
+ - Calling a tool literally named `artifact` because Step 3 talks about "the
143
+ artifact" — that word here means "the deliverable document," not the
144
+ `artifact` tool. That tool is off this skill's surface and, in this
145
+ harness, terminates the run on the spot, writes to
146
+ `.clio-coder/artifacts/` instead of `SPRINT.md`, and skips Step 4 entirely.
147
+ Use `write`.
148
+ - Ending the run right after the write with no final reply — Step 4 is not
149
+ optional.
150
+ - Opening a `tasks` list for the steps above; `tasks` is refused.
151
+ - Reaching for `bash` (`node --version`, `find`, `ls -la`, ...) to check the
152
+ repo's toolchain or structure; `bash` is refused, use `find`/`ls`/`read`.
153
+ - Stopping to wait for an `ask_user` reply that headless runs never send;
154
+ see Arguments.
@@ -40,3 +40,104 @@ Expected:
40
40
 
41
41
  One representative scenario via `clio-coder skills eval` against Nemo-3.5-Lightning
42
42
  (30B local, llamacpp on mini), full-auto sandbox. PASS. Sliced the seeded PLAN.md; judge 4/4.
43
+
44
+ ## Battletest record (2026-09-03)
45
+
46
+ Fixture: `/home/akougkas/eval-temp/harness/test_cutit.py`. S1 reuses this
47
+ evals.md's own fixture text verbatim (`src/todos.js` stub + a concrete
48
+ three-feature `PLAN.md`); S2 is an empty repo with only a `README.md`; S3 is
49
+ a `PLAN.md` that says only "improve performance and clean up the code."
50
+ S1 is the primary grading fixture, scored on 12 checks against the raw
51
+ JSONL's tool-call/safety-block stream, the actual `SPRINT.md` written to
52
+ disk, and the reconstructed final assistant text: zero safety blocks, zero
53
+ real `artifact` tool calls, zero `tasks` calls, `SPRINT.md` exists with a
54
+ `## Battle order` and 2+ numbered slices, every slice carries all six
55
+ required fields (Goal/Depends on/Files/Steps/Done when/Out of scope), no
56
+ horizontal-layering red-flag language, slices trace to the plan's concrete
57
+ features, done-when blocks are command-shaped and testable, and the final
58
+ reply names the path and slice count. S2/S3 are graded on 5 checks each
59
+ (zero safety blocks, zero `artifact` calls, no `SPRINT.md` fabricated,
60
+ correctly flags the gap, recommends a next step). Primary model
61
+ `qwen3.8-27b` on `dynamo` (LM Studio); cross-model confirm on
62
+ `ornith1.5-35b-moe` on `mini` (llama.cpp), run against S1 and S2 both.
63
+
64
+ **The bug found reading the source, confirmed empirically first**: the
65
+ frozen skill (0.3.0) had `artifact` in `allowed-tools` and titled Step 3
66
+ "Write the artifact." `src/tools/artifact.ts` sets `terminate: true` on
67
+ every successful call — the run ends the instant the tool executes, with
68
+ no further LLM turn to confirm what happened. v1's run called `artifact`
69
+ with `kind: "plan"` and, by luck, an explicit `path: "SPRINT.md"` (the
70
+ model inferred this from "honor a caller-supplied path" even though no
71
+ caller supplied one) — so the file landed in the right place this time, but
72
+ the run still ended mid-sentence ("Writing the sprint:") with no
73
+ confirmation reply, and a separate `tasks` call (also off-surface) drew a
74
+ real safety block. Score 8/12: missing `reply_mentions_sprint_path`, one
75
+ real safety block. Had the model not guessed an explicit path, the default
76
+ would have been `.clio-coder/artifacts/PLAN.md` (kind defaults to `plan`,
77
+ see `core/artifact-paths.ts`) — wrong file, wrong location, same silent
78
+ termination. This is exactly the failure mode the planning category's
79
+ tech-spec baseline hit ("called `artifact` for an early exit... instead of
80
+ a spec").
81
+
82
+ | run | model | wall | turns | in / out tokens | safety blocks | score | outcome |
83
+ |---|---|---|---|---|---|---|---|
84
+ | baseline (no skill) | qwen3.8-27b | 76s | 7 | 73.8k / 7.1k | 5 (repeated `ls ".cl"` truncated-path retries, benign) | 5/12 | never invoked `/skill cut-it`; read the plan and module correctly, reasoned to genuinely good vertical slices with real done-when checks in its head, then hit a tool-call loop guard and delivered the entire sprint as **prose in the reply, never wrote `SPRINT.md`** — the exact gap this skill exists to close |
85
+ | v1 (frozen 0.3.0) | qwen3.8-27b | 115s | 6 | 68.6k / 11.3k | 1 real (`tasks` refused) | 8/12 | called the real `artifact` tool for Step 3 as titled; terminated the turn immediately after writing, mid-sentence, with no confirmation reply — the artifact-tool bug, confirmed |
86
+ | v2 (first hardened cut, 0.4.0) | qwen3.8-27b | 70s | 6 | 72.4k / 6.5k | 0 | 11/12 (12/12 after a grading-regex fix, see below) | used `write` correctly, used the `git` tool in Step 1, reported the path and slice count in the final reply; the one score miss was a test-harness regex that didn't handle a `**Done when** (fresh state...):` label followed by bulleted checks on the next lines — the actual done-when content was already command-shaped and testable, fixed in the harness, not the skill |
87
+ | v3 (bash-reflex found) | qwen3.8-27b | — | — | — | 2 (1 benign ENOENT, 1 real: `bash` refused) | 11/12 | reached for `bash` (`node --version`, `find`, `ls -la` chained with `&&`) to survey the toolchain even though `bash` was never in `allowed-tools` and nothing in the body named it explicitly — added the same explicit `bash`-refusal line the `tasks` refusal already had, plus a Red flags entry |
88
+ | v4 (final, stable) | qwen3.8-27b | 60s | 6 | 71.4k / 5.6k | 0 | **12/12** | clean run: `context` → `read`/`ls` → `git status` → `write`, self-contained final reply naming path, slice count, and battle order |
89
+ | final-s2 (no plan) | qwen3.8-27b | 26s | 4 | 43.2k / 2.0k | 0 | **5/5** | checked all three plan locations plus `git log`/`status`, correctly stopped with no `SPRINT.md` written, cited the skill's own red-flag language, recommended `grill-me` |
90
+ | final-s3 (vague plan) | qwen3.8-27b | 28s | 5 | 56.2k / 2.5k | 0 | **5/5** | read the one-line `PLAN.md`, explicitly invoked "the skill's own test" (a testable done-when), listed three concrete missing pieces (object/baseline/target for "performance", a definition of "clean", the absent codebase), stopped without writing `SPRINT.md` |
91
+ | final-mini (cross-model, S1) | ornith1.5-35b-moe (mini) | 42s | 6 | 10.8k / 3.2k | 0 | **12/12** | same clean shape on the second model family: `context` → `ls`/`read` → `write`, self-contained final reply |
92
+ | final-mini-s2 (cross-model, S2) | ornith1.5-35b-moe (mini) | 19s | 6 | 4.0k / 1.1k | 0 | **5/5** | correctly found nothing to slice, recommended `grill-me`, offered to slice immediately if pointed at a plan |
93
+
94
+ **Changes** (0.3.0 -> 0.4.0):
95
+
96
+ 1. **`artifact` removed from `allowed-tools`.** cut-it never needs the real
97
+ `artifact` tool — `SPRINT.md` is a plain file, always written with
98
+ `write`. This is the fix for the bug above.
99
+ 2. **Step 3 retitled** "Write SPRINT.md" (was "Write the artifact") and its
100
+ body now says explicitly "use the `write` tool" and names the
101
+ `artifact`-tool confusion directly, including what it actually does
102
+ wrong (terminal call, wrong default path under `.clio-coder/artifacts/`,
103
+ skips the report-back step).
104
+ 3. **New Step 4 — Report back**, an explicit final-reply requirement (path
105
+ written, slice count, one-line battle-order summary). Nothing in the
106
+ frozen skill told the model to confirm after writing; every hardened run
107
+ now does.
108
+ 4. **`## Arguments` contract**, the section the frozen skill never had:
109
+ conversational trigger syntax, how a named path splits between "the plan
110
+ to slice" and "where to write `SPRINT.md`", the no-operator/`ask_user`-
111
+ auto-cancels rule (adapted from `grill-me` 0.5.0's "this run has no
112
+ back-and-forth" framing — stated as a fact about the run, not gated on a
113
+ cancellation response), and — since cut-it's happy path rarely needs a
114
+ question at all — explicit guidance for the one place it plausibly might
115
+ (choosing among several candidate plans/milestones): pick the best one,
116
+ state the alternative, mark `assumed — confirm`, keep going.
117
+ 5. **`tasks` and `bash` explicitly named as refused**, in Arguments and Red
118
+ Flags, both found empirically: v1's `tasks` call (tracking its own
119
+ steps) and v3's `bash` call (`node --version`, `find`, `ls -la` chained
120
+ with `&&`, to survey the toolchain) were both real safety blocks despite
121
+ neither tool ever having been in `allowed-tools`.
122
+ 6. Three new Red Flags entries for the concrete failures observed: the
123
+ `artifact`-tool confusion, ending the run with no final reply, and the
124
+ `bash` reflex (`tasks` already had informal coverage, now explicit too).
125
+
126
+ **Still weak**: `code_nav` (in `allowed-tools`) was never exercised — this
127
+ fixture's grounding fit entirely in one small stub file, so `read`/`grep`
128
+ sufficed; a plan referencing a larger call graph might exercise it, none
129
+ was built here. The `ask_user`-unavailable path and the "several candidate
130
+ plans" disambiguation guidance in Arguments are reasoned prose, not
131
+ empirically run — no fixture here seeds multiple plausible plan files or a
132
+ scenario where the model actually reaches for `ask_user`; every hardened
133
+ run reasoned straight to the assumed-confirm default or never needed a
134
+ question at all. S3's "flag as vague" grading is phrase-matching against a
135
+ fixed term list (`vague`, `underspecified`, `insufficient`, ...) — a run
136
+ that flags the same gap in different words would under-score on a
137
+ technicality, though neither observed run did. The v2 grading-regex miss
138
+ (`done_when_testable`, fixed in the harness before v4) is a reminder that
139
+ this fixture's automated score can undercount a genuinely correct skill
140
+ output; the raw `SPRINT.md` files are worth spot-reading, not just the
141
+ score column. No timing was captured for v3 (an incremental re-run to
142
+ confirm the `bash` finding, superseded immediately by v4) — not a gap in
143
+ the skill's coverage, just an artifact of the iteration order.