@iowarp/clio-coder 0.4.2 → 0.4.3

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (369) hide show
  1. package/CHANGELOG.md +35 -0
  2. package/CONTRIBUTING.md +86 -19
  3. package/README.md +35 -6
  4. package/dist/{acp-TMDQZDIG.js → acp-H2NGRPWO.js} +11 -11
  5. package/dist/{agents-5N5NG3XG.js → agents-TL5LLUQP.js} +54 -53
  6. package/dist/assets/codewiki.json +1 -1
  7. package/dist/{auth-Z5CCBXKQ.js → auth-E5SW4HMS.js} +19 -16
  8. package/dist/{builtins-K6TNDT24.js → builtins-IA7V7FUC.js} +9 -4
  9. package/dist/{chunk-ZW4HH5JJ.js → chunk-2APPQIER.js} +6 -6
  10. package/dist/{chunk-72GZI5EV.js → chunk-2JH2WHGE.js} +2 -2
  11. package/dist/{chunk-UBRFI4HS.js → chunk-2UG5F4C5.js} +127 -47
  12. package/dist/{chunk-UH632ZYL.js → chunk-2UH2KFUP.js} +2 -2
  13. package/dist/{chunk-3F7VUY77.js → chunk-2VIKGWFZ.js} +2 -2
  14. package/dist/{chunk-I66EAJFY.js → chunk-2WZ546HR.js} +267 -232
  15. package/dist/{chunk-RKSR6VSF.js → chunk-4IUZQIJ3.js} +29 -1
  16. package/dist/{chunk-2X4RYJTJ.js → chunk-4UVU7BJ5.js} +2 -2
  17. package/dist/{chunk-M2DAX4F6.js → chunk-4WR7VSYB.js} +2 -2
  18. package/dist/{chunk-5KW52TEP.js → chunk-54CBCGIR.js} +5 -5
  19. package/dist/{chunk-3KIPBMUA.js → chunk-5ICU3EUH.js} +2 -2
  20. package/dist/{chunk-77QIVUZB.js → chunk-5MEZN6CB.js} +4 -4
  21. package/dist/{chunk-O42A54GG.js → chunk-5OIVVPHF.js} +2 -2
  22. package/dist/{chunk-LJID3DYZ.js → chunk-64I3JVYM.js} +2 -2
  23. package/dist/{chunk-4JDLP6ZS.js → chunk-6PTFB5VS.js} +7 -7
  24. package/dist/{chunk-VN3SHNBN.js → chunk-7DICMOS6.js} +2 -2
  25. package/dist/{chunk-HLAFFSEK.js → chunk-7DRAWPTZ.js} +2 -2
  26. package/dist/chunk-7E7I3WLS.js +3762 -0
  27. package/dist/{chunk-YJISEZKC.js → chunk-7ZYNNDKC.js} +6 -6
  28. package/dist/{chunk-I66ZTYNP.js → chunk-AF4YM7Z4.js} +236 -101
  29. package/dist/{chunk-2HFQNRV3.js → chunk-AX2THNSA.js} +12 -12
  30. package/dist/{chunk-PGF63K6I.js → chunk-B4OAX3SI.js} +65 -3
  31. package/dist/{chunk-W6NIE6OW.js → chunk-B4VEBZKF.js} +3 -3
  32. package/dist/{chunk-JBCS7CRR.js → chunk-BEPZRGGU.js} +10 -10
  33. package/dist/{chunk-XGDPUNND.js → chunk-CE5AX47J.js} +2 -2
  34. package/dist/{chunk-I64IFBLB.js → chunk-DWUOQKRU.js} +17 -10
  35. package/dist/{chunk-DZAW46HP.js → chunk-E3TPLWFX.js} +3 -3
  36. package/dist/{chunk-HIICAHCJ.js → chunk-EKCHAPYA.js} +2 -2
  37. package/dist/{chunk-XE3PCIXH.js → chunk-F5JHEYZM.js} +7 -7
  38. package/dist/{chunk-5PFYMY2V.js → chunk-FTMGRKEF.js} +2 -2
  39. package/dist/{chunk-ZNT2M6TG.js → chunk-G76U63X4.js} +17 -17
  40. package/dist/{chunk-34BHNEE3.js → chunk-GHS5EBTQ.js} +58 -7
  41. package/dist/{chunk-2NHR3NAY.js → chunk-GI7YYQ3F.js} +40 -34
  42. package/dist/{chunk-DYHAXKHD.js → chunk-GWZNEVM2.js} +12 -8
  43. package/dist/chunk-GYV6VZOC.js +26 -0
  44. package/dist/{chunk-MQXIVJ35.js → chunk-HAXOFFRH.js} +5 -5
  45. package/dist/{chunk-W4YEMFBX.js → chunk-HEQY7ZFI.js} +2 -2
  46. package/dist/{chunk-IKOZFYBN.js → chunk-I7ZPNEJM.js} +145 -102
  47. package/dist/{chunk-WNP7O5WZ.js → chunk-ID64D7PE.js} +4 -4
  48. package/dist/{chunk-XQRY4DTA.js → chunk-IGLP3ODT.js} +10 -10
  49. package/dist/chunk-IJNZMHLA.js +101 -0
  50. package/dist/{chunk-JWJGP5DQ.js → chunk-INY6HTFL.js} +7 -7
  51. package/dist/{chunk-PBP4B7XR.js → chunk-IUE3Y34X.js} +2 -2
  52. package/dist/{chunk-B74PXLU7.js → chunk-IWT4SF4R.js} +3 -3
  53. package/dist/{chunk-B7HM5Z7T.js → chunk-JDAY6FIL.js} +5 -5
  54. package/dist/{chunk-PJX3WQUQ.js → chunk-JEQ3XTHC.js} +2 -2
  55. package/dist/{chunk-FSP7CMNU.js → chunk-JGRC33J2.js} +50 -4
  56. package/dist/{chunk-X7IARSHT.js → chunk-JKKCYP3C.js} +9 -9
  57. package/dist/{chunk-HJWWJ6IL.js → chunk-JSC3U7TI.js} +16 -4
  58. package/dist/{chunk-SSEYRH53.js → chunk-KK4JZPBQ.js} +19 -140
  59. package/dist/{chunk-IHXBNWMM.js → chunk-KXDSS5WJ.js} +7 -3
  60. package/dist/{chunk-Q4XWMHX6.js → chunk-L47TF46W.js} +2 -2
  61. package/dist/{chunk-O3YUNJZ2.js → chunk-LDJG7DW3.js} +81 -24
  62. package/dist/{chunk-CDNVLKUX.js → chunk-LLDJM5XK.js} +13 -7
  63. package/dist/{chunk-F2I26BDK.js → chunk-MUW2BDDH.js} +4 -4
  64. package/dist/{chunk-HKMD33FO.js → chunk-MWUZBSAQ.js} +79 -76
  65. package/dist/{chunk-QQLGQY2A.js → chunk-N2Z7HLVY.js} +20 -20
  66. package/dist/{chunk-DZEK6CJN.js → chunk-NIQJ66N4.js} +19 -19
  67. package/dist/{chunk-CWVRRIEI.js → chunk-NZMNUPZZ.js} +2 -2
  68. package/dist/{chunk-462T4EGZ.js → chunk-O5CVSAG5.js} +2 -2
  69. package/dist/{chunk-TPEQIQIE.js → chunk-OML5D5V5.js} +8 -8
  70. package/dist/{chunk-IKSLQ4XV.js → chunk-PAJQJ7BS.js} +558 -216
  71. package/dist/{chunk-ZW55JB7N.js → chunk-PUVDKJ2Y.js} +2 -2
  72. package/dist/{chunk-UH347SHR.js → chunk-QWGDJJYJ.js} +11 -11
  73. package/dist/chunk-R6Q67RJH.js +134 -0
  74. package/dist/{chunk-CRFOIAX3.js → chunk-RRNP2ANY.js} +6 -6
  75. package/dist/{chunk-IDNA72AH.js → chunk-RSJ25QSL.js} +2 -2
  76. package/dist/chunk-SKHCAU7K.js +385 -0
  77. package/dist/{chunk-RLYRBIYQ.js → chunk-TM6LQDI3.js} +20 -12
  78. package/dist/chunk-UOIZ7DA4.js +41 -0
  79. package/dist/{chunk-P75RZCJW.js → chunk-UPZU6GE4.js} +3 -3
  80. package/dist/{chunk-MCMZMDAC.js → chunk-V2ANDPVT.js} +4 -4
  81. package/dist/{chunk-AK5XEFVZ.js → chunk-VA5FNYMT.js} +26 -13
  82. package/dist/{chunk-IMXMHHMQ.js → chunk-VW6DOEDG.js} +332 -57
  83. package/dist/{chunk-XOXV5GKE.js → chunk-W6RRQCPQ.js} +16 -7
  84. package/dist/{chunk-CYZW7JHJ.js → chunk-WBKFA554.js} +8 -8
  85. package/dist/{chunk-BO7Y52RY.js → chunk-WCXUNS7U.js} +7 -7
  86. package/dist/{chunk-ZGNYYXQ6.js → chunk-WRBAGUNF.js} +3 -3
  87. package/dist/{chunk-FVDGR2ZL.js → chunk-XIVNBFZS.js} +85 -30
  88. package/dist/{chunk-BYMNWQ7O.js → chunk-XPWWI35G.js} +299 -58
  89. package/dist/{chunk-KPXDY6QF.js → chunk-XRZT5WY5.js} +2 -2
  90. package/dist/{chunk-AZ4WMN4W.js → chunk-Y3CBHOR6.js} +2 -2
  91. package/dist/{chunk-VXMFAE2W.js → chunk-YPC6ZR5L.js} +19 -6
  92. package/dist/{chunk-54ODD65L.js → chunk-YQWYVTMC.js} +4 -4
  93. package/dist/{chunk-M2WXEHER.js → chunk-ZA4VCIGV.js} +2 -2
  94. package/dist/{chunk-7BHIY2MW.js → chunk-ZDN3Y73Y.js} +6 -6
  95. package/dist/{chunk-E7GT7O5N.js → chunk-ZWPRK62N.js} +7 -4
  96. package/dist/cli/index.js +38 -37
  97. package/dist/{clio-7VB377CC.js → clio-CMMK4KRR.js} +7 -7
  98. package/dist/{code-nav-YVLCYA7V.js → code-nav-MDZNQS33.js} +7 -7
  99. package/dist/{components-UBWCQSRW.js → components-UCUQ4QXW.js} +4 -4
  100. package/dist/{config-4HVOS65E.js → config-SVM5P5YI.js} +76 -74
  101. package/dist/{configure-PIWO7B24.js → configure-LE3IK2TJ.js} +26 -24
  102. package/dist/{context-IYEHL3WQ.js → context-2OHRKS42.js} +66 -63
  103. package/dist/{context-N6ZE3LGJ.js → context-E3VC7RX5.js} +15 -11
  104. package/dist/{context-KQYIWPWT.js → context-VNCR7KAG.js} +60 -45
  105. package/dist/{context-clear-G4OGZJDS.js → context-clear-BW4O37TG.js} +61 -59
  106. package/dist/context-map-COB37XXN.js +505 -0
  107. package/dist/{context-working-set-BWLF6LJP.js → context-working-set-VDS25HXZ.js} +17 -16
  108. package/dist/{dispatch-runner-2QQAITS3.js → dispatch-runner-5AHT53RF.js} +85 -74
  109. package/dist/{doctor-LHBD36VU.js → doctor-WNNVO6FY.js} +37 -37
  110. package/dist/{eval-C45FYRJ6.js → eval-7G7SGAYO.js} +285 -114
  111. package/dist/{eval-inventory-6DEJPLBF.js → eval-inventory-Y6QRFOH5.js} +4 -4
  112. package/dist/{evidence-6SHONYAF.js → evidence-VD6736FQ.js} +63 -62
  113. package/dist/{evolve-KRKMV72X.js → evolve-AL3NGVRL.js} +62 -61
  114. package/dist/{extensions-KPZ2UHBB.js → extensions-MOVJ32NM.js} +7 -7
  115. package/dist/{fleet-IVTCKDHT.js → fleet-QZHUMAGI.js} +110 -108
  116. package/dist/{fleet-commands-EDWL3IT7.js → fleet-commands-BAYT5FJZ.js} +10 -10
  117. package/dist/{fleet-decisions-YP3YEFGK.js → fleet-decisions-IREVMRU4.js} +7 -6
  118. package/dist/{fleet-graph-ZFWKHY2M.js → fleet-graph-YCTT3HTI.js} +19 -18
  119. package/dist/{fleet-inspect-FVUNCBML.js → fleet-inspect-QVJTDAVB.js} +55 -54
  120. package/dist/{fleet-preflight-UN5XED4R.js → fleet-preflight-25QAFPK4.js} +4 -4
  121. package/dist/{fleet-validate-XOWC4HSX.js → fleet-validate-5O57AAJ7.js} +23 -22
  122. package/dist/{fleet-verify-UN3SODEL.js → fleet-verify-CPH2W2T6.js} +56 -55
  123. package/dist/{fleet-view-TWHJKCN6.js → fleet-view-SWBR3VGQ.js} +55 -54
  124. package/dist/{init-T2QORQ3Y.js → init-J477LKZH.js} +78 -76
  125. package/dist/{interop-IN5I2A66.js → interop-3FCM6XLG.js} +11 -11
  126. package/dist/{library-LSCATDLZ.js → library-QUQEIUG6.js} +28 -27
  127. package/dist/{memory-HYOKAGGJ.js → memory-SGGSEP65.js} +64 -63
  128. package/dist/{models-2GPMFYCM.js → models-HEKUAXXK.js} +49 -43
  129. package/dist/{monitor-E4ASVUJH.js → monitor-HKU57TYQ.js} +61 -60
  130. package/dist/{orchestrator-DDMPR3PY.js → orchestrator-VDFAEFAI.js} +919 -546
  131. package/dist/{panes-E3RUXOW5.js → panes-DN2SSFOH.js} +3 -3
  132. package/dist/{panes-IXKLOKA2.js → panes-TALGNPZT.js} +8 -8
  133. package/dist/{paths-L7LGY6RN.js → paths-NBMFAIEZ.js} +5 -5
  134. package/dist/reset-EAJFFJVB.js +344 -0
  135. package/dist/{resources-OTRSN34L.js → resources-OVKSEFVE.js} +27 -20
  136. package/dist/{run-5DEYH5QK.js → run-7DP7ZF2J.js} +113 -109
  137. package/dist/{share-IHWTLO3M.js → share-WML67FT3.js} +26 -25
  138. package/dist/{skills-IYMXMKW4.js → skills-SG662R2K.js} +39 -31
  139. package/dist/{skills-eval-DROHSJAR.js → skills-eval-VVZEUU46.js} +74 -73
  140. package/dist/{skills-inventory-D7X4L4ZX.js → skills-inventory-I2E23GET.js} +21 -20
  141. package/dist/{slash-commands-QBM7UZ3B.js → slash-commands-S7MBJDQK.js} +35 -34
  142. package/dist/{steer-Z5DO23FJ.js → steer-2LQOMCPB.js} +3 -3
  143. package/dist/{support-U7QOWY26.js → support-CC2UJBJ6.js} +6 -6
  144. package/dist/{targets-P2FUC4IL.js → targets-4QC3HIEW.js} +48 -45
  145. package/dist/{terminal-lease-YREJ3JX2.js → terminal-lease-TUHIJ6Y2.js} +2 -2
  146. package/dist/{tools-5B7RO6MV.js → tools-TFGJICCU.js} +8 -8
  147. package/dist/{trace-YMGMUM6A.js → trace-FXMXUZUF.js} +7 -7
  148. package/dist/uninstall-5PEVOE5B.js +408 -0
  149. package/dist/upgrade-M4WXY6KN.js +303 -0
  150. package/dist/{usage-ME5MPXGX.js → usage-N7ZNVLEM.js} +147 -102
  151. package/dist/{verifiers-BVZ7IWOO.js → verifiers-DJTP4XX6.js} +15 -15
  152. package/dist/{verify-5K7ZKQFC.js → verify-RWE4PPEK.js} +9 -9
  153. package/dist/{wiki-generate-F5W5QTYY.js → wiki-generate-C7IQOXSP.js} +84 -82
  154. package/dist/{with-panes-BYOJCLAM.js → with-panes-4GCGSL7J.js} +9 -9
  155. package/dist/worker/entry.js +61 -60
  156. package/docs/architecture/artifact-placement.md +1 -0
  157. package/docs/architecture/artifact-versions.md +1 -1
  158. package/docs/architecture/context-engine.md +4 -0
  159. package/docs/architecture/middleware-and-components.md +1 -1
  160. package/docs/architecture/model-catalog.md +21 -10
  161. package/docs/architecture/observability.md +12 -1
  162. package/docs/architecture/prompt-envelope-and-tools.md +2 -0
  163. package/docs/architecture/provider-adapter-cookbook.md +63 -0
  164. package/docs/architecture/safety-model.md +15 -5
  165. package/docs/guide/built-in-agents.md +17 -3
  166. package/docs/guide/commands-and-modes.md +1 -1
  167. package/docs/guide/configuration-and-targets.md +97 -9
  168. package/docs/guide/configuration-reference.md +7 -2
  169. package/docs/guide/environment-variables.md +2 -0
  170. package/docs/guide/installation-and-lifecycle.md +37 -4
  171. package/docs/guide/proactive-memory.md +66 -55
  172. package/docs/guide/skills-marketplace.md +18 -0
  173. package/docs/process/development-pipeline.md +34 -1
  174. package/docs/process/eval-runner.md +67 -3
  175. package/evals/behavioral-model.yaml +3 -2
  176. package/package.json +2 -2
  177. package/skills/README.md +7 -5
  178. package/skills/coding/ast-grep/SKILL.md +101 -30
  179. package/skills/coding/ast-grep/evals.md +26 -0
  180. package/skills/coding/coding-standards/SKILL.md +40 -5
  181. package/skills/coding/coding-standards/evals.md +23 -0
  182. package/skills/coding/prototype/SKILL.md +87 -28
  183. package/skills/coding/prototype/evals.md +19 -0
  184. package/skills/coding/tdd/SKILL.md +80 -53
  185. package/skills/coding/tdd/evals.md +20 -0
  186. package/skills/context/context-handoff/SKILL.md +43 -2
  187. package/skills/context/context-handoff/evals.md +44 -0
  188. package/skills/context/context-prime/SKILL.md +45 -15
  189. package/skills/context/context-prime/evals.md +45 -0
  190. package/skills/git/branch-closeout/SKILL.md +132 -0
  191. package/skills/git/branch-closeout/evals.md +133 -0
  192. package/skills/git/branch-closeout/references/closeout-checklist.md +81 -0
  193. package/skills/git/file-ticket/SKILL.md +77 -63
  194. package/skills/git/file-ticket/assets/issue-template.md +22 -0
  195. package/skills/git/file-ticket/evals.md +31 -26
  196. package/skills/git/file-ticket/references/issue-discovery.md +49 -0
  197. package/skills/git/fix-issue/SKILL.md +87 -64
  198. package/skills/git/fix-issue/evals.md +35 -31
  199. package/skills/git/fix-issue/references/diagnosis-and-rca.md +46 -0
  200. package/skills/git/resolve-merge-conflicts/SKILL.md +100 -51
  201. package/skills/git/resolve-merge-conflicts/evals.md +52 -25
  202. package/skills/git/resolve-merge-conflicts/references/conflict-matrix.md +126 -0
  203. package/skills/git/ship/SKILL.md +103 -67
  204. package/skills/git/ship/assets/pr-template.md +21 -0
  205. package/skills/git/ship/evals.md +44 -28
  206. package/skills/git/ship/references/remote-and-branch-policy.md +62 -0
  207. package/skills/git/worktree-create/SKILL.md +80 -50
  208. package/skills/git/worktree-create/evals.md +40 -33
  209. package/skills/git/worktree-create/references/worktree-setup.md +62 -66
  210. package/skills/git/worktree-merge/SKILL.md +112 -65
  211. package/skills/git/worktree-merge/evals.md +42 -34
  212. package/skills/git/worktree-merge/references/merge-strategies.md +52 -0
  213. package/skills/planning/archify/SKILL.md +196 -0
  214. package/skills/planning/archify/evals.md +65 -0
  215. package/skills/planning/architecture/SKILL.md +61 -12
  216. package/skills/planning/architecture/evals.md +65 -0
  217. package/skills/planning/backlog/SKILL.md +130 -14
  218. package/skills/planning/backlog/evals.md +142 -0
  219. package/skills/planning/prd/SKILL.md +46 -6
  220. package/skills/planning/prd/evals.md +54 -0
  221. package/skills/planning/product-intent/SKILL.md +57 -2
  222. package/skills/planning/product-intent/evals.md +70 -0
  223. package/skills/planning/tech-spec/SKILL.md +53 -2
  224. package/skills/planning/tech-spec/evals.md +73 -0
  225. package/skills/registry.yaml +58 -50
  226. package/skills/remote.yaml +13 -0
  227. package/skills/research/arxiv-literature/SKILL.md +76 -18
  228. package/skills/research/arxiv-literature/evals.md +50 -0
  229. package/skills/research/experiment-protocol/SKILL.md +20 -1
  230. package/skills/research/experiment-protocol/evals.md +23 -0
  231. package/skills/research/scientific-debugging/SKILL.md +23 -1
  232. package/skills/research/scientific-debugging/evals.md +18 -0
  233. package/skills/research/scientific-modernization/SKILL.md +26 -1
  234. package/skills/research/scientific-modernization/evals.md +27 -0
  235. package/skills/skill-marketplace.json +63 -28
  236. package/skills/workflow/cut-it/SKILL.md +65 -5
  237. package/skills/workflow/cut-it/evals.md +101 -0
  238. package/skills/workflow/design-council/SKILL.md +117 -27
  239. package/skills/workflow/design-council/evals.md +161 -0
  240. package/skills/workflow/grill-me/SKILL.md +86 -10
  241. package/skills/workflow/grill-me/evals.md +153 -0
  242. package/skills/workflow/workflow-distiller/SKILL.md +76 -17
  243. package/skills/workflow/workflow-distiller/evals.md +118 -0
  244. package/src/cli/configure-interop.ts +105 -13
  245. package/src/cli/configure-oauth.ts +57 -0
  246. package/src/cli/configure-onboarding.ts +980 -0
  247. package/src/cli/configure-target.ts +594 -0
  248. package/src/cli/configure.ts +1082 -528
  249. package/src/cli/context-map.ts +114 -0
  250. package/src/cli/context.ts +4 -0
  251. package/src/cli/index.ts +1 -0
  252. package/src/cli/lifecycle-presenter.ts +436 -0
  253. package/src/cli/models.ts +10 -2
  254. package/src/cli/modes/print.ts +5 -1
  255. package/src/cli/reset.ts +228 -106
  256. package/src/cli/run.ts +7 -2
  257. package/src/cli/select.ts +664 -0
  258. package/src/cli/skills.ts +9 -2
  259. package/src/cli/targets.ts +3 -0
  260. package/src/cli/uninstall.ts +233 -165
  261. package/src/cli/upgrade.ts +204 -149
  262. package/src/cli/usage.ts +86 -27
  263. package/src/cli/validate-model.ts +3 -3
  264. package/src/core/config.ts +56 -0
  265. package/src/core/external-diagnostic.ts +44 -0
  266. package/src/core/gateway-routing.ts +157 -0
  267. package/src/core/safe-exec.ts +17 -2
  268. package/src/core/skill-activation.ts +89 -2
  269. package/src/domains/agents/builtins/world-knowledge.md +31 -0
  270. package/src/domains/agents/catalog.ts +1 -1
  271. package/src/domains/agents/result-contract.ts +70 -0
  272. package/src/domains/context/wiki/map-seed.ts +589 -0
  273. package/src/domains/context/wiki/plan.ts +2 -2
  274. package/src/domains/dispatch/admission.ts +29 -0
  275. package/src/domains/dispatch/agent-candidates.ts +10 -0
  276. package/src/domains/dispatch/budget-envelope.ts +86 -1
  277. package/src/domains/dispatch/capability-match.ts +1 -0
  278. package/src/domains/dispatch/capacity-lease.ts +17 -0
  279. package/src/domains/dispatch/contract.ts +11 -1
  280. package/src/domains/dispatch/extension.ts +134 -29
  281. package/src/domains/dispatch/types.ts +3 -0
  282. package/src/domains/dispatch/worker-model-metadata.ts +38 -0
  283. package/src/domains/eval/metrics/call-ledger-stream.ts +34 -11
  284. package/src/domains/eval/metrics/token-stream.ts +201 -31
  285. package/src/domains/eval/metrics/tracked.ts +40 -4
  286. package/src/domains/eval/runners/clio-run.ts +5 -2
  287. package/src/domains/eval/schema/suite.ts +28 -0
  288. package/src/domains/eval/schema/verdict.ts +2 -2
  289. package/src/domains/eval/suites/resolve.ts +13 -1
  290. package/src/domains/eval/suites/run.ts +24 -3
  291. package/src/domains/interop/registry.ts +6 -2
  292. package/src/domains/interop/types.ts +4 -0
  293. package/src/domains/lifecycle/migrations/index.ts +4 -0
  294. package/src/domains/memory/task-memory-policy.ts +70 -26
  295. package/src/domains/memory/task-memory-telemetry.ts +1 -0
  296. package/src/domains/middleware/index.ts +0 -1
  297. package/src/domains/middleware/marketplace-offer.ts +3 -35
  298. package/src/domains/middleware/memory-intervention.ts +127 -32
  299. package/src/domains/middleware/memory-step-endpoint.ts +3 -2
  300. package/src/domains/middleware/skills-reminder.ts +31 -2
  301. package/src/domains/observability/compaction-usage.ts +118 -0
  302. package/src/domains/observability/cost.ts +1 -1
  303. package/src/domains/observability/extension.ts +6 -1
  304. package/src/domains/observability/out-of-turn-usage.ts +52 -21
  305. package/src/domains/providers/contract.ts +4 -1
  306. package/src/domains/providers/extension.ts +40 -9
  307. package/src/domains/providers/model-capabilities.ts +9 -0
  308. package/src/domains/providers/model-discovery.ts +2 -0
  309. package/src/domains/providers/model-runtime-capabilities.ts +15 -5
  310. package/src/domains/providers/models/local-models/clio-coder-local-coding-targets.yaml +32 -12
  311. package/src/domains/providers/runtimes/antigravity/antigravity-code.ts +225 -45
  312. package/src/domains/providers/runtimes/common/lmstudio-http.ts +6 -2
  313. package/src/domains/providers/runtimes/common/local-synth.ts +2 -0
  314. package/src/domains/providers/runtimes/protocol/litellm.ts +119 -29
  315. package/src/domains/providers/support.ts +11 -5
  316. package/src/domains/providers/target-model-cache.ts +25 -2
  317. package/src/domains/providers/types/capability-flags.ts +2 -0
  318. package/src/domains/providers/types/runtime-descriptor.ts +20 -1
  319. package/src/domains/providers/types/target-descriptor.ts +19 -0
  320. package/src/domains/resources/index.ts +3 -0
  321. package/src/domains/resources/skills/install.ts +72 -7
  322. package/src/domains/resources/skills/loader.ts +7 -0
  323. package/src/domains/resources/skills/marketplace.ts +63 -11
  324. package/src/domains/safety/autonomy.ts +15 -0
  325. package/src/domains/safety/index.ts +1 -0
  326. package/src/domains/safety/path-policy.ts +1 -1
  327. package/src/domains/safety/policy-engine.ts +34 -11
  328. package/src/domains/safety/protected-artifacts.ts +191 -88
  329. package/src/domains/safety/run-effects.ts +2 -22
  330. package/src/domains/safety/skill-authority.ts +55 -0
  331. package/src/domains/session/compaction/compact.ts +72 -22
  332. package/src/domains/session/entries.ts +6 -0
  333. package/src/domains/session/usage.ts +3 -3
  334. package/src/engine/agent.ts +13 -3
  335. package/src/engine/ai.ts +26 -8
  336. package/src/engine/antigravity/subprocess-runtime.ts +386 -120
  337. package/src/engine/api-registry.ts +3 -0
  338. package/src/engine/apis/openai-completions.ts +117 -14
  339. package/src/engine/external-subprocess.ts +114 -6
  340. package/src/entry/background-model-metadata.ts +18 -0
  341. package/src/entry/compaction-prompt.ts +57 -0
  342. package/src/entry/orchestrator.ts +405 -216
  343. package/src/entry/task-memory-lifecycle.ts +35 -0
  344. package/src/interactive/chat-loop-messages.ts +13 -4
  345. package/src/interactive/chat-loop.ts +65 -2
  346. package/src/interactive/chat-renderer.ts +1 -0
  347. package/src/interactive/cost-overlay.ts +26 -2
  348. package/src/interactive/interactive-slash-runtime.ts +2 -1
  349. package/src/interactive/renderers/worker-entry.ts +32 -0
  350. package/src/interactive/slash-commands.ts +24 -6
  351. package/src/interactive/theme/labels.ts +19 -13
  352. package/src/interactive/turn-context.ts +9 -5
  353. package/src/interactive/turn-recovery.ts +8 -0
  354. package/src/interactive/turn-runtime.ts +27 -11
  355. package/src/interactive/turn-state.ts +7 -0
  356. package/src/interactive/worker-receipts.ts +1 -0
  357. package/src/interactive/worker-stream.ts +6 -1
  358. package/src/tools/context/index.ts +30 -9
  359. package/src/tools/dispatch-arguments.ts +1 -0
  360. package/src/tools/dispatch-event-text.ts +10 -0
  361. package/src/tools/dispatch-plan.ts +1 -0
  362. package/src/tools/dispatch-runner.ts +12 -0
  363. package/src/tools/registry.ts +11 -5
  364. package/src/tools/worker-evidence.ts +3 -1
  365. package/src/worker/spec-contract.ts +4 -0
  366. package/dist/chunk-2Z2IKEXI.js +0 -1554
  367. package/dist/reset-OAQP3W4O.js +0 -230
  368. package/dist/uninstall-N34PCTGJ.js +0 -331
  369. package/dist/upgrade-PXK3S2YM.js +0 -325
@@ -7,7 +7,7 @@ triggers:
7
7
  - code-shaped contracts
8
8
  - implementation-ready technical specification
9
9
  - specify execution flows
10
- version: 0.2.0
10
+ version: 0.3.0
11
11
  license: Apache-2.0
12
12
  disable-model-invocation: true
13
13
  allowed-tools:
@@ -42,6 +42,46 @@ TypeScript pseudocode plus end-to-end execution flows. Prose explains why;
42
42
  types and call stacks define what changes. Design only — never implement,
43
43
  and save a file only when the user asks; otherwise return the spec inline.
44
44
 
45
+ ## Arguments
46
+
47
+ ```text
48
+ /skill tech-spec <the change, in a few sentences, or a path to read first>
49
+ ```
50
+
51
+ - The text is the design problem: what's changing and why. A doc or file
52
+ path named in the request (a PRD, an architecture decision, a module) is
53
+ context to read, not more arguments — see "Load local context" below.
54
+ - Nothing is required beyond some text; a blank invocation falls straight
55
+ to Path B's first question rather than inventing a change to spec.
56
+ - **Output defaults to inline.** Write a file only when the request says
57
+ so explicitly — "save it", "write it to `<path>`", "put it in `docs/`".
58
+ Absent that, the finished spec is the reply itself: no file, in this run
59
+ or a prior one in the same session, gets created for it. This holds
60
+ regardless of which path below runs or how long the spec is — length is
61
+ never itself a reason to write a file.
62
+ - Disabled for model self-invocation and requires the `tdd` skill be
63
+ installed to reference in the TDD Test Plan section; both are frontmatter
64
+ facts, not something to explain to the user unless asked.
65
+
66
+ There is no operator in a headless run: `ask_user` either isn't registered
67
+ or nothing answers it, and a call that goes unanswered will not resolve
68
+ differently on a second try. In Path B (below), that means: state the
69
+ question, your recommendation grounded in the codebase and any docs read
70
+ (or the most defensible engineering default when nothing grounds it), and
71
+ the reasoning; adopt the recommendation; mark it `assumed — confirm`; move
72
+ to the next question. Run every question this way, end to end, not just
73
+ the first — the interview is the plan to execute, not an outline to
74
+ abbreviate because no one answered the opening question. Never invent a
75
+ fact or a codebase detail to back an assumption; anything genuinely
76
+ unknown becomes an Open Question in the spec, not a plausible guess. This
77
+ degrades the interview only — it never licenses writing a file that
78
+ wasn't asked for.
79
+
80
+ The steps below are the plan; do not open a task list for them. `tasks`
81
+ sits outside this skill's tool surface and any call to it is refused.
82
+ `bash` is also outside this skill's tool surface — verify what you wrote
83
+ with `grep`, `read`, and `find`, never `bash`.
84
+
45
85
  ## Choose the path
46
86
 
47
87
  - **Path A — convert context to spec**: the conversation, docs, or codebase
@@ -51,6 +91,8 @@ and save a file only when the user asks; otherwise return the spec inline.
51
91
  with a recommended answer per question (the grill-me posture); anything
52
92
  answerable by exploring the codebase is explored, not asked. When context
53
93
  suffices, run Path A. Never invent requirements to skip the interview.
94
+ See Arguments above for how a headless run carries every question
95
+ through instead of stalling on the first one.
54
96
 
55
97
  ## Path A
56
98
 
@@ -110,7 +152,8 @@ contracts, seams, call stacks, or the test plan for being hard):
110
152
 
111
153
  The spec follows the outline, every boundary has a typed contract or a
112
154
  stated reason it needs none, every behavior has a call stack, unknowns are
113
- open questions rather than invented design, and nothing was implemented.
155
+ open questions rather than invented design, nothing was implemented, and
156
+ no file was written unless the request asked for one.
114
157
 
115
158
  ## Red flags
116
159
 
@@ -119,3 +162,11 @@ open questions rather than invented design, and nothing was implemented.
119
162
  - Speculative seams no invariant, boundary, or test earns.
120
163
  - The same rule restated in three sections.
121
164
  - "While I'm here" implementation.
165
+ - Writing the spec to a file when nothing in the request asked for one —
166
+ the default output is always the inline reply.
167
+ - A Path B question left unanswered instead of run as the assumed-confirm
168
+ monologue, or an interview skipped straight into Path A without ever
169
+ asking the first question.
170
+ - Opening a task list for the steps above; `tasks` is refused. Reaching for
171
+ `bash` to grep or verify the spec; `bash` is not in this skill's tool
172
+ surface and the call is refused — use `grep`/`read`/`find`.
@@ -45,3 +45,76 @@ Expected:
45
45
 
46
46
  One representative scenario via `clio-coder skills eval` against Nemo-3.5-Lightning
47
47
  (30B local, llamacpp on mini), full-auto sandbox. PASS. Spec written and its claims exercised with node -e; judge 4/4.
48
+
49
+ ## Battletest record (2026-09-03)
50
+
51
+ Fixture: `/home/akougkas/eval-temp/harness/test_techspec.py`, continuing the
52
+ planning category's shared HPC log-triage domain (`product-intent` -> `prd`
53
+ -> `tech-spec`). Seeds a plausible root `PRD.md` (purpose, features,
54
+ out-of-scope, stack, integrations, data model, milestones, and the
55
+ always-on-vs-on-demand ingestion tension explicitly marked as *not this
56
+ document's decision*) plus the existing partial codebase (`src/scanner.py`:
57
+ a working `FailureEvent` + `scan_oom`, OOM only) and two sample dmesg logs
58
+ carrying real OOM/ECC/Xid line formats, inside a git repo. Three task
59
+ variants, one fixture:
60
+
61
+ - **base** (S1, Path A): "spec ECC + Xid detection and cross-signature
62
+ top-3 ranking" — sufficient context, no save request. Graded on 9 checks
63
+ against the *reconstructed final assistant text* (this skill's default
64
+ output is inline, not a file): all 11 outline-derived sections present,
65
+ >=5 domain grounding terms, >=2 materially different alternatives,
66
+ `FailureEvent` reused not respecced, nothing implemented (`scanner.py`
67
+ byte-identical to seed), the ingestion trade-off left unresolved, zero
68
+ safety blocks, no `tasks` call, and — the check this run exists to catch —
69
+ **no file written when nothing asked for one**.
70
+ - **save** (S1 variant, confirmation only, run once on the final version):
71
+ same task plus an explicit "save it to docs/tech-spec-log-triage.md" —
72
+ 10 checks, same 8 plus the file existing at exactly that path and no
73
+ other new file appearing.
74
+ - **thin** (S2, Path B, confirmation only, run once on the final version):
75
+ a genuinely vague "improve our failure detection, you'll need to ask me
76
+ stuff" request with 5 lighter checks — zero safety blocks, no silent
77
+ stall, no `tasks` call, and either a real `ask_user` exchange or the
78
+ assumed-confirm monologue (`assumed` + `confirm` both present), with a
79
+ real spec still produced.
80
+
81
+ | run | model | wall | turns | in / out tokens | safety blocks | score | outcome |
82
+ |---|---|---|---|---|---|---|---|
83
+ | baseline (no skill) | ornith-1.5-35b-a3b | 80s | 5 | 54.9k / 12.1k | 0 | 4/9 | never invoked `/skill tech-spec`; discovered the installed skill itself via `context(scope="skills")`, read its SKILL.md directly, then called `artifact` and terminated early with a `.clio-coder/artifacts/PLAN.md` instead of a spec — no alternatives, no sections, wrong output shape |
84
+ | v1 (frozen 0.2.0) | ornith-1.5-35b-a3b | 95s | 8 | 115.0k / 15.1k | 1 | 6/9 | ran Path A correctly and produced a genuinely strong spec (11/11 sections, 3 material alternatives, `FailureEvent` reused, nothing implemented) but opened a `tasks` plan (refused, safety block) and **wrote the spec to `docs/tech-spec-scanner-ecc-xid.md` without being asked to** — the exact Path-A/B default-output risk flagged going in |
85
+ | v2 (live 0.3.0) | ornith-1.5-35b-a3b | 68s | 5 | 57.2k / 10.4k | 0 | 9/9 | same spec quality, zero safety blocks, no `tasks` call, correctly returned inline with no file written; final text states explicitly "I did **not** write a file, since nothing in the request asked to save it" |
86
+ | v2 confirm — save | ornith-1.5-35b-a3b | 110s | 11 | 184.8k / 16.5k | 0 | 10/10 | explicit "save it to docs/tech-spec-log-triage.md" correctly produces exactly that file at that path, nothing else |
87
+ | v2 confirm — thin (Path B) | ornith-1.5-35b-a3b | 69s | 9 | 111.8k / 11.7k | 0 | 5/5 | correctly identified insufficient context, ran Path B, and carried all five scope decisions (S1-S5) through as an explicit assumed-confirm monologue headlessly instead of stalling or silently skipping to Path A |
88
+
89
+ **Changes**: (1) `## Arguments` contract with the slash-invocation syntax,
90
+ what's required vs. inferred, and — the section that mattered most here —
91
+ an explicit "output defaults to inline" rule stated as its own bullet
92
+ before the headless-monologue prose, so the fix for Path B's ask_user gap
93
+ can't be misread as license to always write a file; (2) the headless
94
+ no-operator paragraph, ported from `product-intent`/`prd`, applied to Path
95
+ B's grill-me interview: every question runs as state-question /
96
+ grounded-recommendation / reasoning / adopt / mark `assumed — confirm`,
97
+ end to end, not just the first one; (3) explicit `tasks` and `bash`
98
+ refusal lines — `tasks` was v1's only safety block; (4) `Done when` and
99
+ `Red flags` both gained a line naming the unrequested-file failure and the
100
+ unanswered-Path-B-question failure by name, plus the existing `tasks`/`bash`
101
+ refusal repeated as a red flag (matching `prd`'s and `product-intent`'s
102
+ pattern of naming the exact observed failure, not a generic reminder).
103
+ Version 0.2.0 -> 0.3.0.
104
+
105
+ **Still weak**: per this pass's coordinator note, no secondary-model
106
+ confirmation was run (qwen3.8-27b was skipped in favor of running one full
107
+ cycle on ornith-1.5-35b-a3b at speed, concurrently with a sibling agent
108
+ hardening `architecture` on `mini`); the fix is validated on one model
109
+ class only. The baseline's failure mode (discovering and improvising from
110
+ the installed skill file directly, without ever invoking it, then calling
111
+ `artifact` for an unrelated early exit) is a skill-selection/tool-scoping
112
+ gap this SKILL.md cannot fix from inside its own body. `code_nav` (in
113
+ `allowed-tools`) was never exercised — the fixture's one-file codebase
114
+ never needed it. `requires: [skill:tdd]` is a diagnostic-only reference in
115
+ this harness (unmet requires warn, never block `--skill`-path invocation);
116
+ the TDD Test Plan section reads fine without the `tdd` skill installed, but
117
+ that was not tested with `tdd` actually present to see if the reference
118
+ changes. Genuine unknowns (S3 from the original evals) were exercised only
119
+ incidentally via the Xid-severity and ECC-correctable open questions, not
120
+ as an isolated scenario.
@@ -6,54 +6,58 @@ skills:
6
6
  # ── coding ──
7
7
  - name: ast-grep
8
8
  path: coding/ast-grep
9
- version: 0.2.0
10
- sha256: d21a4b330d1488c43348eaee7fbeec24bb8a1d7c4e536059db02fd9377c62dfa
9
+ version: 0.3.0
10
+ sha256: 0a72f4c906303f550be78f80a927da7e24e5c2076f058ccafdec80dfbed275af
11
11
  - name: coding-standards
12
12
  path: coding/coding-standards
13
- version: 0.2.0
14
- sha256: da21ad373252575934ad484a2926e8c7827880c9d91b4e0c656fe92d71335063
13
+ version: 0.3.0
14
+ sha256: 3ee3481430591a8f8d041bd1c5fa078becd301d88eb49610e41d056494d3413d
15
15
  - name: prototype
16
16
  path: coding/prototype
17
- version: 0.3.0
18
- sha256: 151bd153752874d5970dd06bc15df14c0037631283d148195a8fa75c5a59a45b
17
+ version: 0.4.0
18
+ sha256: 0f82df386c2ee565b0e968216ece746b1a11b4acd79676812074c1a406b698e1
19
19
  - name: tdd
20
20
  path: coding/tdd
21
- version: 0.3.0
22
- sha256: f3a03655864d4981791218b085d09f8fe836d0795012cf4c8c07b430730abb2a
21
+ version: 0.4.0
22
+ sha256: 63b88f29424091a92a8d8c0cc94474ae7e78af76491aea17e6f54af66f32db2d
23
23
  # ── context ──
24
24
  - name: context-handoff
25
25
  path: context/context-handoff
26
- version: 0.4.0
27
- sha256: ba0568719d0b58bb3e1631eee72094ec0165bf2fb7f69e1e67d223ff50b364d8
26
+ version: 0.5.0
27
+ sha256: e69adb1533a6a850cb83babe7e580dd260e034781d7c35e65921453e3fa191a5
28
28
  - name: context-prime
29
29
  path: context/context-prime
30
- version: 0.3.0
31
- sha256: d8413255688f1697d40231df830ce3c8947ac2a551f3e6fbb0d10158c4fa8d0e
30
+ version: 0.4.0
31
+ sha256: 21587263297a7ee9e81a2d66f6fc800a1ee6db15f76de1e650563b71fdf19d8e
32
32
  # ── git ──
33
+ - name: branch-closeout
34
+ path: git/branch-closeout
35
+ version: 0.1.0
36
+ sha256: 3227f31b0693bb6428abb20a2d4f449aeba1d8e45a79cf0c2c7c333c30c10664
33
37
  - name: file-ticket
34
38
  path: git/file-ticket
35
- version: 0.2.0
36
- sha256: 07b5d996d459c6c8c99ea44f44421d418f4738075cccd0e4b759a6f8a7daf720
39
+ version: 0.3.0
40
+ sha256: 65029c2d9d545728750c8a713f92035bd6bdbfa9874de442df72bd059846cb7a
37
41
  - name: fix-issue
38
42
  path: git/fix-issue
39
- version: 0.2.0
40
- sha256: ec5e6a00056041494f9b100da8f0775455124aebc7000ed0edb7c9b8493f74c8
43
+ version: 0.3.0
44
+ sha256: 626e3aa8ca95bc603f2e0cdd2e7502aa96be85af6bfa92b635c39ef99d3845a5
41
45
  - name: resolve-merge-conflicts
42
46
  path: git/resolve-merge-conflicts
43
- version: 0.3.0
44
- sha256: a1df7dee85ef788f257f8b8fbf4c2564cac245c3010d3f60d3b62d721d185f5e
47
+ version: 0.4.0
48
+ sha256: 901f482aa2231552d64a721e0d1c58dc8f211347543b5a91fd2ab3c7e413c11c
45
49
  - name: ship
46
50
  path: git/ship
47
- version: 0.2.0
48
- sha256: 8a22f74e56a6e98c85a38051550d1e1a70c9ca5bb5cf271f4e6448f0f791e0ed
51
+ version: 0.5.0
52
+ sha256: 4d626718fdb8b67cf3f22d09ac5a228e5c8103befa1295a672587a6605f038b1
49
53
  - name: worktree-create
50
54
  path: git/worktree-create
51
- version: 0.3.0
52
- sha256: d40bb0bb9152c4f2e391de1285b6cbd2e67a45492e2cf865b1f82910e30a6bcc
55
+ version: 0.6.0
56
+ sha256: a4ae790d916a170977814b4334e14aef4f96fb304e3289f7beb5e31e8fb83881
53
57
  - name: worktree-merge
54
58
  path: git/worktree-merge
55
- version: 0.3.0
56
- sha256: bca803a2484b0eac90939377c81f9e66928242f36ac2d86cef9381b1ff997f3e
59
+ version: 0.6.0
60
+ sha256: 0fbe2297622ed955b2e4fb14306d75a51bd36db7eebf59e8e92fac1da5b6e64e
57
61
  # ── meta ──
58
62
  - name: clio-coder-dev
59
63
  path: meta/clio-coder-dev
@@ -80,57 +84,61 @@ skills:
80
84
  version: 0.3.0
81
85
  sha256: ba81e09412f26647a9384b07359201efb08949865c50934778a4ab3173d43096
82
86
  # ── planning ──
87
+ - name: archify
88
+ path: planning/archify
89
+ version: 0.1.0
90
+ sha256: d9e891eb5f3c27678de14165a0eb35d1eb2289be480e55009f9a57bcfe55be2d
83
91
  - name: architecture
84
92
  path: planning/architecture
85
- version: 0.3.0
86
- sha256: 0da15631b3a60946e90b63047047e0ee1d84b6705e4f9f5405036b87fc849454
93
+ version: 0.4.0
94
+ sha256: e9017412b6261492b987fe27f7573b72013a340c0d3eb0e7e568fd588b22a214
87
95
  - name: backlog
88
96
  path: planning/backlog
89
- version: 0.3.0
90
- sha256: af92244d2ba8723b656ce00d531df88850c7886915653751e7ba59957b5454be
97
+ version: 0.4.0
98
+ sha256: e634e2fb5035731e7daa2ef246ae4ab3ef5d792b0b604f4216dfe10335ee10c0
91
99
  - name: prd
92
100
  path: planning/prd
93
- version: 0.3.0
94
- sha256: 505460dcac23b762a87fd0f7f3aaa99806e3b0fbedeaa54f9c48608c49c16356
101
+ version: 0.4.0
102
+ sha256: 39fe417f495be4153541dd70c886cf773265636dc9fa1844cfc7012b71cbbec1
95
103
  - name: product-intent
96
104
  path: planning/product-intent
97
- version: 0.3.0
98
- sha256: d34300e6e2422fc1071e352f3e0f1b13a4e65dbcfd4eaf6dd9ef2b61f18d72ad
105
+ version: 0.4.0
106
+ sha256: d41689db9a8c9b2dad0cd630412bebaed5792f9d22d4edb18770e4b1b3308699
99
107
  - name: tech-spec
100
108
  path: planning/tech-spec
101
- version: 0.2.0
102
- sha256: 28b1b0a0a3e9b48b29f679ab2377ee66d719c9a670eed622a38ac69be8b72b2e
109
+ version: 0.3.0
110
+ sha256: 5f9add0d01e43feb6808b2eaeae9ac68075cf369cbfc793e020181a642d4d2c4
103
111
  # ── research ──
104
112
  - name: arxiv-literature
105
113
  path: research/arxiv-literature
106
- version: 0.4.0
107
- sha256: b0b1bb80dc8d15e52ab180ab4f5645859310b45d7ebdf87649e0e5a7f7bca987
114
+ version: 0.5.0
115
+ sha256: 0bc6b39d12934568804fc11b23219de4c6c2172e28c4bfedcee3104e2217aae6
108
116
  - name: experiment-protocol
109
117
  path: research/experiment-protocol
110
- version: 0.2.0
111
- sha256: 126e0f6dc3534a141978f2ccd1c6ea659df447beb1015bc92e41856d318584a2
118
+ version: 0.3.0
119
+ sha256: 235256cbdf44ebdd25d87cce60940e00ad1fca42a1b2b4e534bfc245f0d1135c
112
120
  - name: scientific-debugging
113
121
  path: research/scientific-debugging
114
- version: 0.2.0
115
- sha256: f2b2b6442740ef23270fad94f75ba95ce522da1877a468ba112cbaff68e5fa2a
122
+ version: 0.3.0
123
+ sha256: 9c29d70b4b82a440640b41665c23731b4112125915bde0e91016fc60aec67207
116
124
  - name: scientific-modernization
117
125
  path: research/scientific-modernization
118
- version: 0.3.0
119
- sha256: 66a99f38b38232735bb70d7924a85d0ab5addabfe3935cc82c0aaa3303d68773
126
+ version: 0.4.0
127
+ sha256: 491522d547fa0bddef19d2ec92bc35236b169787bded6ecc3a29dfec19892886
120
128
  # ── workflow ──
121
129
  - name: cut-it
122
130
  path: workflow/cut-it
123
- version: 0.3.0
124
- sha256: 6429a4725acf874379c1df7736dd95db07acd97ad5bd6f6226649370cd03d177
131
+ version: 0.4.0
132
+ sha256: d3742221f0ace1f1b62c07f5902030e478496059a58ef194e7cb3303ccdff05e
125
133
  - name: design-council
126
134
  path: workflow/design-council
127
- version: 0.4.0
128
- sha256: 1ecd53397ce97ec78294cc9e8c7291b20e2e7d740bcde8ea999739d359c91055
135
+ version: 0.5.0
136
+ sha256: c279a94a4cd66d46980f9a5024aedac3e673144fadc9d36b53ad87fc2ee63501
129
137
  - name: grill-me
130
138
  path: workflow/grill-me
131
- version: 0.4.0
132
- sha256: 9cc07b89a4a8a3ecd4fb73d2cd9f7893d272d26f592ea3eb9883ff44069a2cea
139
+ version: 0.5.0
140
+ sha256: f48687566a8419aa52f30f8234f8dcfd050eaefdfdcfa281723d3603f3e08fcc
133
141
  - name: workflow-distiller
134
142
  path: workflow/workflow-distiller
135
- version: 0.3.0
136
- sha256: 6aa6bb3088f0e4925946aaba3f217dfc5a6e4b8e31f2f051abeda599da91cd48
143
+ version: 0.4.0
144
+ sha256: dd7c312d6904f861fa105334c7c2f18c9d6e26a70a76874c85858cee90bede58
@@ -0,0 +1,13 @@
1
+ # Skills whose content lives in another repository at a pinned ref.
2
+ # Clio never vendors these. `npm run skills:pin` publishes each entry into
3
+ # skill-marketplace.json with the upstream tree as its sourceUrl, the catalog
4
+ # overlay whose files land on top of that tree at install, and the upstream
5
+ # top-level members the install drops. The overlay SKILL.md is pinned in
6
+ # registry.yaml like every other catalog skill.
7
+ version: 1
8
+ skills:
9
+ - name: archify
10
+ category: planning
11
+ sourceUrl: https://github.com/tt-a1i/archify/tree/v2.16.0/archify
12
+ overlay: skills/planning/archify
13
+ exclude: [test, package-lock.json]
@@ -7,7 +7,7 @@ triggers:
7
7
  - compare these papers
8
8
  - find recent research papers
9
9
  - build a literature survey
10
- version: 0.4.0
10
+ version: 0.5.0
11
11
  license: Apache-2.0
12
12
  allowed-tools:
13
13
  - web_fetch
@@ -16,7 +16,6 @@ allowed-tools:
16
16
  - grep
17
17
  - find
18
18
  - ls
19
- - artifact
20
19
  clio-coder:
21
20
  registry-id: iowarp/clio-coder
22
21
  source-url: https://github.com/iowarp/clio-coder/tree/main/skills/research/arxiv-literature
@@ -34,6 +33,20 @@ Find, summarize, or compare academic papers without flooding the main context
34
33
  window. Raw search results and paper text stay in a worker or get compressed
35
34
  immediately; only paper cards reach the user.
36
35
 
36
+ ## Arguments
37
+
38
+ ```text
39
+ /skill arxiv-literature <request>
40
+ ```
41
+
42
+ The request is free text: a paper URL/ID, a topic, or two or more IDs to
43
+ compare. There is no operator in a headless run — `ask_user` is not in this
44
+ skill's tool surface and nothing answers it. If the request is ambiguous (no
45
+ clear topic, an ID that doesn't resolve), state your best reading and
46
+ proceed; never stall waiting for clarification. `tasks` sits outside this
47
+ skill's tool surface and is refused; the steps below are the whole plan, do
48
+ not open a task list for them.
49
+
37
50
  ## Step 1 — Classify the request
38
51
 
39
52
  Pick exactly one:
@@ -45,29 +58,49 @@ Pick exactly one:
45
58
 
46
59
  ## Step 2 — Pick the vehicle
47
60
 
48
- - **Search, compare, survey**: dispatch the `researcher` shadow agent. Task
49
- prompt:
61
+ Default every request — single paper, search, compare, and survey alike to
62
+ `web_fetch` directly against arXiv:
50
63
 
51
- ```text
52
- Research arXiv literature for: <user goal>.
53
- Return only compact source-linked paper cards, comparison/synthesis,
54
- caveats, and read/skim/skip recommendations.
55
- ```
56
-
57
- - **Single paper, or dispatch unavailable**: use `web_fetch` directly.
58
- - Paper URL/ID: fetch the arXiv page. Clio normalizes it into structured
59
- metadata plus AlphaXiv enrichment when available.
60
- - Search query: fetch the arXiv Atom API; Clio compacts the XML into paper
61
- cards:
64
+ - Paper URL/ID: fetch the arXiv page. Clio normalizes it into structured
65
+ metadata plus AlphaXiv enrichment when available.
66
+ - Topic, compare, or survey: fetch the arXiv Atom API; Clio compacts the XML
67
+ into paper cards:
62
68
 
63
69
  ```text
64
70
  https://export.arxiv.org/api/query?search_query=all:QUERY&sortBy=submittedDate&sortOrder=descending&start=0&max_results=10
65
71
  ```
66
72
 
73
+ For a compare request, run one query per paper ID (`id_list=ID` instead of
74
+ `search_query`) or one broader query covering all of them — whichever stays
75
+ inside the fetch cap in Step 3.
76
+
67
77
  Useful categories for `search_query`: `cs.AI` (AI), `cs.LG` (ML), `cs.CL`
68
78
  (NLP/LLMs), `cs.CR` (security), `cs.SE` (software engineering), `cs.MA`
69
79
  (multi-agent), `cs.IR` (retrieval/RAG), `cs.CV` (vision), `cs.RO` (robotics).
70
80
 
81
+ **Only dispatch the `researcher` shadow agent when the user explicitly asks
82
+ for a deep or broad survey** ("survey the field", "don't just skim arXiv, go
83
+ wide") — never as the default for an ordinary search or compare. Left
84
+ unbounded, a dispatched worker has no arXiv-only restriction and no fetch
85
+ budget of its own, and will wander into Semantic Scholar, DBLP, OpenAlex, and
86
+ general web search, taking several minutes to return nothing useful. When you
87
+ do dispatch, state the same bound this skill uses directly, in the task
88
+ prompt itself:
89
+
90
+ ```text
91
+ Research arXiv literature for: <user goal>.
92
+ Search arXiv only (export.arxiv.org Atom API or arxiv.org paper pages) — do
93
+ not query Semantic Scholar, DBLP, OpenAlex, or general web search. One
94
+ round, at most two fetch attempts total (successes and failures both count).
95
+ On a timeout or HTTP error, do not retry with a different host, scheme, or
96
+ protocol — retry the identical URL at most once, then stop and report the
97
+ failure. Build cards from whatever you have; do not keep escalating.
98
+ Return only compact source-linked paper cards, comparison/synthesis,
99
+ caveats, and read/skim/skip recommendations.
100
+ ```
101
+
102
+ If dispatch is unavailable, fall back to the direct `web_fetch` path above.
103
+
71
104
  ## Step 3 — Return paper cards only
72
105
 
73
106
  Never paste raw Atom XML or full paper text into the response. Output format:
@@ -96,9 +129,14 @@ Never paste raw Atom XML or full paper text into the response. Output format:
96
129
  Done when every returned paper has a card with a working link and the
97
130
  recommendation section is filled in. Stop after one search round unless the
98
131
  user asks to go deeper; do not keep fetching to "be thorough". One round
99
- means at most two Atom API fetches: the initial query plus one refinement.
100
- Rewording the same query a third time is thrash build cards from what the
101
- first two returned.
132
+ means at most two Atom API fetch attempts total: the initial query plus one
133
+ refinement, or one retry of a failed fetch. Failed attempts (timeout, 429,
134
+ connection error) count against this cap the same as successful ones — a
135
+ timeout is not a free retry. Rewording the same query a third time, or
136
+ retrying through a different host/scheme (`http` vs `https`,
137
+ `export.arxiv.org` vs `arxiv.org/search`, an unofficial JSON mirror) after a
138
+ failure, is thrash: build cards from whatever the attempts returned, or
139
+ report the network failure plainly and stop.
102
140
 
103
141
  ## Gotchas
104
142
 
@@ -108,3 +146,23 @@ first two returned.
108
146
  - AlphaXiv is AI-generated enrichment: useful for scanning, never citable as
109
147
  authoritative.
110
148
  - Fetch or enrich only the top candidates, not every result.
149
+ - `export.arxiv.org` rate-limits (HTTP 429) under repeated hits; that is a
150
+ reason to stop at the fetch cap, not a reason to retry against a mirror.
151
+ - This skill's tool surface has no file-writing tool and no `artifact`. The
152
+ paper cards are the chat reply, never a document; do not go looking for a
153
+ way to save one.
154
+
155
+ ## Red flags
156
+
157
+ - Dispatching `researcher` for an ordinary search or compare "just in case"
158
+ it does a better job than a direct fetch — it is slower and unbounded by
159
+ default; reserve it for an explicit deep-survey ask.
160
+ - A dispatch task prompt without the arXiv-only, fetch-capped instruction —
161
+ that omission is what let a prior run wander into Semantic Scholar, DBLP,
162
+ and OpenAlex for minutes with nothing to show for it.
163
+ - Treating a timeout or 429 as free of the fetch cap and retrying against a
164
+ different host, scheme, or unofficial mirror instead of stopping.
165
+ - Raw Atom XML, a full abstract dump, or more than the top few candidates
166
+ reaching the final reply.
167
+ - Reaching for `artifact` or any write tool to "save" the result — it is not
168
+ in this skill's tool surface; the reply is the deliverable.
@@ -56,3 +56,53 @@ Expected:
56
56
 
57
57
  One representative scenario via `clio-coder skills eval` against Nemo-3.5-Lightning
58
58
  (30B local, llamacpp on mini), full-auto sandbox. PASS (driver FAIL overturned on transcript review): skill loaded, skill_surface blocked two curl attempts, retrieval ran through web_fetch on the Atom API per the designed dispatch-unavailable fallback. Judge misread the OR in bullet 1. Query thrash (8 fetches) led to the two-fetch cap now in the body.
59
+
60
+ ## Battletest record (2026-09-03)
61
+
62
+ S1 (topic search, speculative decoding) and S2 (single known paper,
63
+ `arxiv.org/abs/1706.03762`), `ornith1.5-35b-moe` on mini (llamacpp), `clio-coder
64
+ run --autonomy full-auto --json`, headless, real network access (no fixture).
65
+
66
+ | run | model | wall | turns | in / out tokens | safety blocks | outcome |
67
+ |---|---|---|---|---|---|---|
68
+ | baseline S1 (no skill) | ornith1.5-35b-moe | 286s (killed at timeout) | 9 | n/a (killed before `agent_end`) | 5 | never finished: fired raw `curl` via `bash` (permission-gated, refused) before falling back to `web_fetch`; hit real timeouts against `export.arxiv.org`; killed by the harness's 300s cap mid-turn with no cards produced |
69
+ | baseline S2 (no skill) | ornith1.5-35b-moe | 24s | 2 | 3.6k / 1.0k | 0 | fetched the real paper directly and wrote a good prose summary, but never labeled problem/method/evidence/limitation as distinct fields — the un-skilled gap S2 expects |
70
+ | v1 S1 (frozen v0.4.0, dispatch runaway) | ornith1.5-35b-moe | 286s (killed at timeout) | 6 | n/a (killed before `agent_end`) | 2 | opened a `tasks` plan (refused, outside surface), dispatched `researcher` twice (first dispatch flagged an absolute-path token in the briefing), then started re-fetching directly itself; killed by the harness's 300s cap before producing cards — the dispatch-runaway/no-cap bug reproduced live |
71
+ | v2 S1 (hardened v0.5.0) | ornith1.5-35b-moe | 109s | 4 | 2.9k / 3.9k | 2 (real `export.arxiv.org` 429 + timeout) | stayed on `web_fetch` only (no dispatch — topic search didn't ask for a "deep survey"), made exactly two fetch attempts against the identical URL per the tightened cap, hit a genuine rate limit then a timeout, **stopped at the cap**, explicitly refused to fabricate paper cards from invented IDs, and returned an honestly-labeled "established knowledge, not a fresh fetch" orientation instead — terminated cleanly on its own, no runaway |
72
+ | v2 S2 (hardened v0.5.0) | ornith1.5-35b-moe | 28s | 3 | 6.2k / 1.5k | 0 | direct `web_fetch` on the real paper page, full problem/method/evidence/limitation/relevance card, explicit read/skim/skip recommendation, AlphaXiv linked and labeled as enrichment, no `artifact` call — 5/5 on the S2 rubric |
73
+
74
+ Changes: removed `artifact` from `allowed-tools` — nothing in the procedure
75
+ ever called it, and it is a terminal tool that would end the run the moment
76
+ the model reached for it to "save" a result. Reworked Step 2 so `web_fetch`
77
+ directly against arXiv is the default vehicle for every request class
78
+ (single paper, search, compare, survey), and `dispatch` is reserved for an
79
+ explicitly requested deep/broad survey — the live probe that motivated this
80
+ pass showed a dispatched `researcher` wandering into Semantic Scholar, DBLP,
81
+ and OpenAlex with no arXiv-only restriction and no fetch budget, never
82
+ returning. When dispatch is used, its task-prompt template now states the
83
+ same arXiv-only, fetch-capped discipline the direct path uses, verbatim.
84
+ Tightened the Step 3 fetch cap to count failed attempts (timeout, 429,
85
+ connection error) against the same two-attempt budget as successes, and
86
+ banned escalating to a different host/scheme/mirror on failure — this closed
87
+ a real gap the v2 S1 run against a genuinely rate-limited `export.arxiv.org`
88
+ would otherwise have exploited (the same tightened wording is what let it
89
+ stop cleanly at 109s instead of retrying indefinitely). Added `## Arguments`
90
+ (free-text request, no operator headlessly, `tasks` refused) and a `##
91
+ Red flags` section naming the dispatch-runaway, uncapped-retry, and
92
+ artifact-reach failure modes actually observed.
93
+
94
+ Still weak: S3 (three-paper comparison) was not run against a real fixture
95
+ this pass — budget went to confirming the dispatch-runaway fix and the
96
+ fetch-cap fix on S1/S2, both of which reproduced live. The compare path's
97
+ "one query per ID or one broader query" guidance in Step 2 is new and
98
+ untested end-to-end. Both baseline and v1 runs for S1 were killed by the
99
+ harness's outer timeout rather than allowed to run to their own natural
100
+ (bad) conclusion — a longer timeout might show the old skill eventually
101
+ recovering, or might show it running further off scope; the fix (bounding
102
+ the fetch cap and defaulting off dispatch) is validated by the v2 behavior,
103
+ not by watching v1 fail for longer. `export.arxiv.org` rate-limited several
104
+ runs in this session from repeated hits in short succession; the 429s in the
105
+ v2 S1 run are a real external condition this pass ran into, not a fixture
106
+ simulation, but a quieter network day could produce a fully-populated
107
+ card set on the same prompt instead of the honest-failure path exercised
108
+ here — both are now handled, but only the failure path got a live rep.
@@ -8,7 +8,7 @@ triggers:
8
8
  - define numerical tolerances
9
9
  - reproduce these results
10
10
  - compare solver accuracy
11
- version: 0.2.0
11
+ version: 0.3.0
12
12
  license: Apache-2.0
13
13
  allowed-tools:
14
14
  - read
@@ -39,6 +39,25 @@ mode this protocol exists to prevent.
39
39
  Anti-trigger: if the question is "why is this output wrong", that is a
40
40
  diagnosis, not an experiment; use scientific-debugging.
41
41
 
42
+ ## Arguments
43
+
44
+ ```text
45
+ /skill experiment-protocol <what to benchmark, compare, or sweep>
46
+ ```
47
+
48
+ There is no operator in a headless run — `ask_user` is not in this skill's
49
+ tool surface. If a threshold, tolerance, or environment detail is unstated,
50
+ pick the most defensible default, record it as an explicit assumption in the
51
+ pre-registration, and proceed; never stall Phase 0 waiting for confirmation.
52
+
53
+ The phases below are the plan; do not open a task list for them — `tasks`
54
+ sits outside this skill's tool surface and any call to it is refused.
55
+
56
+ Shell rules for every `bash` call: one command per call, plain and direct.
57
+ Never use `$(...)` or backticks; they trigger an approval gate that ends a
58
+ headless run. Capture checksums and environment facts with direct calls
59
+ (`sha256sum data/mesh.h5`), never command substitution.
60
+
42
61
  ## Phase 0 - Pre-register
43
62
 
44
63
  Before any measurement, write the protocol into the repository validation
@@ -89,3 +89,26 @@ Prompt: "Make this kernel faster."
89
89
 
90
90
  One representative scenario via `clio-coder skills eval` against Nemo-3.5-Lightning
91
91
  (30B local, llamacpp on mini), full-auto sandbox. PASS. Pre-registration written before touching the seeded kernel; judge 5/5.
92
+
93
+ ## Battletest record (2026-09-03)
94
+
95
+ Real `clio-coder run --json` against dynamo (LM Studio, `qwen3.8-27b`), the S1
96
+ kernel.py fixture above, git repo under
97
+ `/home/akougkas/eval-temp/expprotocol-fixture/`. Same gap as every other
98
+ research skill going in: no `## Arguments`, no `tasks`-refusal, no shell-rules
99
+ paragraph, no no-operator statement. Added all four.
100
+
101
+ | run | model | outcome |
102
+ |---|---|---|
103
+ | v0.3.0 (hardened) | qwen3.8-27b | environment capture (`python3 --version`, `uname -a`, CPU model), `numpy` version check, `sha256sum` on the input array and the frozen baseline copy, `.clio-coder/validation.yaml` written with thresholds/tolerance semantics before any benchmark, a real 20-rep baseline vs. a vectorized candidate, a bit-exact accuracy check between them, then a correctness spot-check on a second slice-based candidate before timing it — 34 tool calls, zero safety blocks, zero `$(...)`, zero `tasks`, `write`/`edit` used correctly (in this skill's surface, unlike scientific-debugging) instead of a heredoc |
104
+
105
+ The run did not reach a final Phase 3 verdict/report inside the 280s box used
106
+ this pass (31 API calls, ~750k cumulative input tokens — the box closes on
107
+ context-reprocessing volume, not model slowness); everything observed up to
108
+ that point followed Phase 0-2 exactly as specified, including registering a
109
+ 100x stretch target, measuring ~13x, and correctly continuing to iterate
110
+ rather than declaring victory early.
111
+
112
+ **Still weak**: no observed run reaching a written Phase 3 verdict this pass
113
+ (same time-box cause as scientific-debugging's record above). No cross-model
114
+ confirmation this pass.
@@ -8,7 +8,7 @@ triggers:
8
8
  - diagnose NaNs
9
9
  - debug with falsifiable hypotheses
10
10
  - scientific root cause
11
- version: 0.2.0
11
+ version: 0.3.0
12
12
  license: Apache-2.0
13
13
  allowed-tools:
14
14
  - read
@@ -38,6 +38,28 @@ Anti-trigger: if the failure is a typo, a missing import, or an error message
38
38
  that names its own cause, fix it directly and skip this workflow. The loop
39
39
  below is for failures that survived the first obvious fix.
40
40
 
41
+ ## Arguments
42
+
43
+ ```text
44
+ /skill scientific-debugging <failure description>
45
+ ```
46
+
47
+ Everything after the skill name is the failure report: the observed wrong
48
+ behavior and whatever has already been tried. There is no operator in a
49
+ headless run — `ask_user` is not in this skill's tool surface. If the goal,
50
+ a fault-class split, or a ranking call is ambiguous, state your best reading
51
+ in Step 1 or Step 3 and proceed; never stall a step waiting for confirmation.
52
+
53
+ The Loop below is the plan; do not open a task list for it — `tasks` sits
54
+ outside this skill's tool surface and any call to it is refused.
55
+
56
+ Shell rules for every `bash` call: one command per call, plain and direct.
57
+ Never use `$(...)` or backticks; they trigger an approval gate that ends a
58
+ headless run. This skill has no `write`/`edit` tool — the structured-
59
+ investigation file in the Tiers section below is written with a `bash`
60
+ heredoc (`cat > file <<'EOF' ... EOF`), never through an edit tool that
61
+ isn't in this skill's surface.
62
+
41
63
  ## The Loop
42
64
 
43
65
  1. **Goal.** One sentence stating the observable "fixed" state. "The regression