@iowarp/clio-coder 0.4.1 → 0.4.3

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (604) hide show
  1. package/CHANGELOG.md +127 -0
  2. package/CONTRIBUTING.md +142 -52
  3. package/README.md +434 -473
  4. package/SECURITY.md +2 -1
  5. package/dist/{acp-ZILU3AUO.js → acp-H2NGRPWO.js} +12 -12
  6. package/dist/{agents-HYWGBGQR.js → agents-TL5LLUQP.js} +56 -55
  7. package/dist/assets/codewiki.json +1 -1
  8. package/dist/{auth-N3QT7CBO.js → auth-E5SW4HMS.js} +23 -21
  9. package/dist/builtins-IA7V7FUC.js +22 -0
  10. package/dist/{chunk-7RY5VZPH.js → chunk-2APPQIER.js} +8 -8
  11. package/dist/{chunk-72GZI5EV.js → chunk-2JH2WHGE.js} +2 -2
  12. package/dist/{chunk-JA5QWE4Z.js → chunk-2UG5F4C5.js} +1973 -1664
  13. package/dist/{chunk-5YHDIDBP.js → chunk-2UH2KFUP.js} +2 -2
  14. package/dist/{chunk-CTJ4RNAA.js → chunk-2VIKGWFZ.js} +2 -2
  15. package/dist/{chunk-I66EAJFY.js → chunk-2WZ546HR.js} +267 -232
  16. package/dist/{chunk-GIZNH63R.js → chunk-35MSIRKH.js} +9 -4
  17. package/dist/chunk-3EBYEESD.js +314 -0
  18. package/dist/{chunk-J5LZHVIT.js → chunk-3M6DQK6S.js} +113 -35
  19. package/dist/{chunk-RKSR6VSF.js → chunk-4IUZQIJ3.js} +29 -1
  20. package/dist/{chunk-6FN3E6KX.js → chunk-4O6MANBS.js} +2 -2
  21. package/dist/chunk-4UVU7BJ5.js +39 -0
  22. package/dist/{chunk-VKRH2TCS.js → chunk-4WR7VSYB.js} +2 -2
  23. package/dist/{chunk-BBTJOK6Y.js → chunk-54CBCGIR.js} +5 -5
  24. package/dist/{chunk-AP73CFDC.js → chunk-5ICU3EUH.js} +2 -2
  25. package/dist/chunk-5MEZN6CB.js +1334 -0
  26. package/dist/{chunk-O42A54GG.js → chunk-5OIVVPHF.js} +2 -2
  27. package/dist/{chunk-ABLSQ6JX.js → chunk-64I3JVYM.js} +8 -2
  28. package/dist/{chunk-AFKWHWXF.js → chunk-6PTFB5VS.js} +39 -22
  29. package/dist/{chunk-VN3SHNBN.js → chunk-7DICMOS6.js} +2 -2
  30. package/dist/chunk-7DRAWPTZ.js +360 -0
  31. package/dist/chunk-7E7I3WLS.js +3762 -0
  32. package/dist/{chunk-BJGUKIG4.js → chunk-7ZYNNDKC.js} +7 -7
  33. package/dist/{chunk-XKA2ICR3.js → chunk-AF4YM7Z4.js} +652 -252
  34. package/dist/{chunk-GVQJ5CCZ.js → chunk-AX2THNSA.js} +12 -12
  35. package/dist/{chunk-IG7BCQBA.js → chunk-B4OAX3SI.js} +65 -3
  36. package/dist/{chunk-TD3PGPQA.js → chunk-B4VEBZKF.js} +3 -3
  37. package/dist/{chunk-74YWRRU5.js → chunk-BEPZRGGU.js} +10 -10
  38. package/dist/{chunk-FEFIFZTL.js → chunk-CE5AX47J.js} +2 -2
  39. package/dist/{chunk-UAPGZHYC.js → chunk-DWUOQKRU.js} +25 -11
  40. package/dist/{chunk-THYWACCR.js → chunk-E3TPLWFX.js} +3 -3
  41. package/dist/{chunk-7EPLI7VL.js → chunk-EKCHAPYA.js} +2 -2
  42. package/dist/{chunk-HLW2MRKE.js → chunk-F4EKGO4N.js} +3 -1
  43. package/dist/{chunk-PJJ6MY27.js → chunk-F5JHEYZM.js} +7 -7
  44. package/dist/{chunk-6CCS4G3W.js → chunk-FTMGRKEF.js} +3 -3
  45. package/dist/{chunk-SINK3QR6.js → chunk-G76U63X4.js} +17 -17
  46. package/dist/{chunk-EIMVLWB3.js → chunk-GHS5EBTQ.js} +64 -9
  47. package/dist/{chunk-QMXC4JB7.js → chunk-GI7YYQ3F.js} +187 -1419
  48. package/dist/{chunk-TZSKNMZG.js → chunk-GTUD2WMY.js} +2 -1
  49. package/dist/{chunk-6HMJX2VU.js → chunk-GWZNEVM2.js} +44 -12
  50. package/dist/chunk-GYV6VZOC.js +26 -0
  51. package/dist/{chunk-MQXIVJ35.js → chunk-HAXOFFRH.js} +5 -5
  52. package/dist/{chunk-UXN6JT4W.js → chunk-HEQY7ZFI.js} +3 -3
  53. package/dist/{chunk-7PWAODYW.js → chunk-I7XBWTYH.js} +2 -2
  54. package/dist/{chunk-GCSMB2KY.js → chunk-I7ZPNEJM.js} +145 -102
  55. package/dist/{chunk-WNP7O5WZ.js → chunk-ID64D7PE.js} +4 -4
  56. package/dist/{chunk-QTFGO774.js → chunk-IGLP3ODT.js} +29 -16
  57. package/dist/chunk-IJNZMHLA.js +101 -0
  58. package/dist/{chunk-BDPT6GTK.js → chunk-INY6HTFL.js} +7 -7
  59. package/dist/{chunk-PBP4B7XR.js → chunk-IUE3Y34X.js} +2 -2
  60. package/dist/{chunk-6NJQITNH.js → chunk-IWT4SF4R.js} +6 -3
  61. package/dist/{chunk-R23Z6K6I.js → chunk-JDAY6FIL.js} +19 -19
  62. package/dist/chunk-JEQ3XTHC.js +42 -0
  63. package/dist/{chunk-FSP7CMNU.js → chunk-JGRC33J2.js} +50 -4
  64. package/dist/{chunk-TVH4ONAM.js → chunk-JKKCYP3C.js} +10 -10
  65. package/dist/{chunk-HJWWJ6IL.js → chunk-JSC3U7TI.js} +16 -4
  66. package/dist/{chunk-C537JADH.js → chunk-KK4JZPBQ.js} +19 -141
  67. package/dist/{chunk-K6BF4U2H.js → chunk-KKOJXO6R.js} +62 -14
  68. package/dist/{chunk-IHXBNWMM.js → chunk-KXDSS5WJ.js} +7 -3
  69. package/dist/{chunk-6DWBAZ5U.js → chunk-L47TF46W.js} +5 -7
  70. package/dist/{chunk-HUAS7ITX.js → chunk-LDJG7DW3.js} +91 -42
  71. package/dist/{chunk-CDNVLKUX.js → chunk-LLDJM5XK.js} +13 -7
  72. package/dist/{chunk-YPI3QQCF.js → chunk-MCEPRMZW.js} +2 -4
  73. package/dist/{chunk-Y4CAGMM6.js → chunk-MNJGS2IN.js} +5 -6
  74. package/dist/{chunk-VKFQTNDV.js → chunk-MUW2BDDH.js} +4 -4
  75. package/dist/{chunk-E67WX76H.js → chunk-MWUZBSAQ.js} +104 -152
  76. package/dist/{chunk-OJTRZGR3.js → chunk-N2Z7HLVY.js} +21 -21
  77. package/dist/{chunk-TVHHYFHE.js → chunk-NEDJ26B5.js} +2 -2
  78. package/dist/{chunk-FYUN5KZ3.js → chunk-NIQJ66N4.js} +21 -21
  79. package/dist/{chunk-U2WB7TZS.js → chunk-NMJXSHBJ.js} +97 -85
  80. package/dist/{chunk-CWVRRIEI.js → chunk-NZMNUPZZ.js} +2 -2
  81. package/dist/{chunk-VEGN6WIQ.js → chunk-O5CVSAG5.js} +3 -3
  82. package/dist/{chunk-MOPSG2X7.js → chunk-OML5D5V5.js} +8 -8
  83. package/dist/{chunk-2VG7KLYV.js → chunk-PAJQJ7BS.js} +5816 -3255
  84. package/dist/{chunk-ZW55JB7N.js → chunk-PUVDKJ2Y.js} +2 -2
  85. package/dist/{chunk-BTGG6BG2.js → chunk-QWGDJJYJ.js} +158 -19
  86. package/dist/chunk-R6Q67RJH.js +134 -0
  87. package/dist/{chunk-ZJLUDYFY.js → chunk-RRNP2ANY.js} +6 -6
  88. package/dist/{chunk-PVAMAVBB.js → chunk-RSJ25QSL.js} +102 -2
  89. package/dist/{chunk-NLFAQR7Z.js → chunk-S66XZJOF.js} +3 -23
  90. package/dist/chunk-SKHCAU7K.js +385 -0
  91. package/dist/chunk-SZAA6XDG.js +30 -0
  92. package/dist/{chunk-J4HBWF6Y.js → chunk-TM6LQDI3.js} +131 -28
  93. package/dist/chunk-UOIZ7DA4.js +41 -0
  94. package/dist/{chunk-MA3H6DM5.js → chunk-UPZU6GE4.js} +25 -3
  95. package/dist/{chunk-BWW4HLO4.js → chunk-UXCU4E3T.js} +8 -6
  96. package/dist/{chunk-N5UK64DP.js → chunk-V2ANDPVT.js} +4 -4
  97. package/dist/{chunk-AK5XEFVZ.js → chunk-VA5FNYMT.js} +26 -13
  98. package/dist/{chunk-6VC4OV3Z.js → chunk-VIA6RFQZ.js} +3 -11
  99. package/dist/{chunk-ZAZB4JMW.js → chunk-VKPAQYEB.js} +27 -8
  100. package/dist/{chunk-QKIFBZKT.js → chunk-VW6DOEDG.js} +497 -81
  101. package/dist/{chunk-SCYB3HA4.js → chunk-W6RRQCPQ.js} +63 -19
  102. package/dist/{chunk-2NM363SV.js → chunk-WBKFA554.js} +10 -10
  103. package/dist/{chunk-R32CLGZ6.js → chunk-WCXUNS7U.js} +82 -21
  104. package/dist/{chunk-GPPB3JBE.js → chunk-WRBAGUNF.js} +3 -3
  105. package/dist/{chunk-IXJT6DCX.js → chunk-XIVNBFZS.js} +85 -30
  106. package/dist/{chunk-UEDMSP56.js → chunk-XPWWI35G.js} +417 -201
  107. package/dist/chunk-XRZT5WY5.js +47 -0
  108. package/dist/{chunk-3QSOM6PA.js → chunk-Y3CBHOR6.js} +2 -2
  109. package/dist/{chunk-VXMFAE2W.js → chunk-YPC6ZR5L.js} +19 -6
  110. package/dist/{chunk-AKB4GYDL.js → chunk-YQWYVTMC.js} +5 -5
  111. package/dist/{chunk-6I5ILFOF.js → chunk-ZA4VCIGV.js} +3 -3
  112. package/dist/{chunk-7OBGU7UB.js → chunk-ZDN3Y73Y.js} +12 -18
  113. package/dist/{chunk-3I5NY75V.js → chunk-ZWPRK62N.js} +8 -5
  114. package/dist/cli/index.js +41 -39
  115. package/dist/{clio-IT3G3VQH.js → clio-CMMK4KRR.js} +9 -9
  116. package/dist/{code-nav-RK6S7F6E.js → code-nav-MDZNQS33.js} +89 -21
  117. package/dist/{components-UBWCQSRW.js → components-UCUQ4QXW.js} +4 -4
  118. package/dist/{config-3QZRWZJF.js → config-SVM5P5YI.js} +131 -84
  119. package/dist/{configure-FL7Y3KJF.js → configure-LE3IK2TJ.js} +28 -26
  120. package/dist/{context-5HE7ODYK.js → context-2OHRKS42.js} +69 -64
  121. package/dist/{context-KYQFRVDC.js → context-E3VC7RX5.js} +15 -11
  122. package/dist/{context-XNHL75JV.js → context-VNCR7KAG.js} +93 -65
  123. package/dist/{context-clear-N545L53A.js → context-clear-BW4O37TG.js} +64 -60
  124. package/dist/context-map-COB37XXN.js +505 -0
  125. package/dist/{context-working-set-QHKXSV2F.js → context-working-set-VDS25HXZ.js} +19 -18
  126. package/dist/{dispatch-runner-RGIE5PCT.js → dispatch-runner-5AHT53RF.js} +93 -82
  127. package/dist/{docs-5NAF6AU7.js → docs-PD3EXDKU.js} +21 -20
  128. package/dist/{doctor-ZGPEGHIP.js → doctor-WNNVO6FY.js} +48 -47
  129. package/dist/{eval-GXLL44RD.js → eval-7G7SGAYO.js} +287 -115
  130. package/dist/{eval-inventory-HBWSWQOK.js → eval-inventory-Y6QRFOH5.js} +4 -4
  131. package/dist/{evidence-HWLBRH3Q.js → evidence-VD6736FQ.js} +67 -64
  132. package/dist/{evolve-FTZBMNVW.js → evolve-AL3NGVRL.js} +65 -62
  133. package/dist/{extensions-VHRBEID7.js → extensions-MOVJ32NM.js} +9 -7
  134. package/dist/{fleet-CKZHJWZJ.js → fleet-QZHUMAGI.js} +114 -111
  135. package/dist/{fleet-commands-EXDXBMV6.js → fleet-commands-BAYT5FJZ.js} +10 -10
  136. package/dist/{fleet-decisions-OTHB6KRL.js → fleet-decisions-IREVMRU4.js} +7 -6
  137. package/dist/{fleet-graph-YTEZUCUT.js → fleet-graph-YCTT3HTI.js} +22 -19
  138. package/dist/{fleet-inspect-SS6YMDCK.js → fleet-inspect-QVJTDAVB.js} +58 -55
  139. package/dist/{fleet-preflight-PBY4VYOM.js → fleet-preflight-25QAFPK4.js} +4 -4
  140. package/dist/{fleet-validate-KMEM5L3S.js → fleet-validate-5O57AAJ7.js} +26 -23
  141. package/dist/{fleet-verify-QD5M7E7Q.js → fleet-verify-CPH2W2T6.js} +59 -56
  142. package/dist/{fleet-view-WAMJYNDT.js → fleet-view-SWBR3VGQ.js} +58 -55
  143. package/dist/{init-5XQRBOFV.js → init-J477LKZH.js} +82 -79
  144. package/dist/{interop-34TVO25M.js → interop-3FCM6XLG.js} +11 -11
  145. package/dist/{library-3QY6KF57.js → library-QUQEIUG6.js} +30 -27
  146. package/dist/{memory-L4UTIIIW.js → memory-SGGSEP65.js} +67 -64
  147. package/dist/{models-ZVX3QOWE.js → models-HEKUAXXK.js} +53 -46
  148. package/dist/{monitor-CEKVSYTS.js → monitor-HKU57TYQ.js} +63 -60
  149. package/dist/{orchestrator-77BAP6BC.js → orchestrator-VDFAEFAI.js} +1831 -1057
  150. package/dist/{panes-7STHOAUJ.js → panes-DN2SSFOH.js} +5 -5
  151. package/dist/{panes-SHAUIRXY.js → panes-TALGNPZT.js} +29 -14
  152. package/dist/{paths-L7LGY6RN.js → paths-NBMFAIEZ.js} +5 -5
  153. package/dist/reset-EAJFFJVB.js +344 -0
  154. package/dist/{resources-74GKTLSF.js → resources-OVKSEFVE.js} +29 -20
  155. package/dist/{run-HBAUJNNZ.js → run-7DP7ZF2J.js} +120 -115
  156. package/dist/{share-G3APVLVP.js → share-WML67FT3.js} +32 -27
  157. package/dist/{skills-35HHUKCR.js → skills-SG662R2K.js} +41 -31
  158. package/dist/{skills-eval-QN4HSHDC.js → skills-eval-VVZEUU46.js} +78 -77
  159. package/dist/{skills-inventory-J357J34F.js → skills-inventory-I2E23GET.js} +23 -20
  160. package/dist/{slash-commands-JZZCQA32.js → slash-commands-S7MBJDQK.js} +40 -36
  161. package/dist/{steer-XAVHJM22.js → steer-2LQOMCPB.js} +3 -3
  162. package/dist/{support-U7QOWY26.js → support-CC2UJBJ6.js} +6 -6
  163. package/dist/{targets-DSM6CY3M.js → targets-4QC3HIEW.js} +54 -54
  164. package/dist/{terminal-lease-JOPFUVEM.js → terminal-lease-TUHIJ6Y2.js} +5 -5
  165. package/dist/{tools-MKNWVPBH.js → tools-TFGJICCU.js} +10 -10
  166. package/dist/{trace-ECQ7TIYZ.js → trace-FXMXUZUF.js} +55 -7
  167. package/dist/uninstall-5PEVOE5B.js +408 -0
  168. package/dist/upgrade-M4WXY6KN.js +303 -0
  169. package/dist/{usage-X52N3IDJ.js → usage-N7ZNVLEM.js} +151 -104
  170. package/dist/{verifiers-EJTVVSMA.js → verifiers-DJTP4XX6.js} +15 -15
  171. package/dist/{verify-YJL6XET2.js → verify-RWE4PPEK.js} +9 -9
  172. package/dist/{web-fetch-MPIFL3LL.js → web-fetch-MPARV2K7.js} +2 -2
  173. package/dist/{wiki-generate-4NDZTQ4B.js → wiki-generate-C7IQOXSP.js} +89 -86
  174. package/dist/{with-panes-OBOBFIIR.js → with-panes-4GCGSL7J.js} +53 -257
  175. package/dist/worker/entry.js +90 -74
  176. package/docs/README.md +176 -81
  177. package/docs/{acp.md → architecture/acp.md} +36 -20
  178. package/docs/{alcf-provider.md → architecture/alcf-provider.md} +8 -5
  179. package/docs/{architecture.md → architecture/architecture.md} +43 -22
  180. package/docs/{artifact-placement.md → architecture/artifact-placement.md} +27 -23
  181. package/docs/architecture/artifact-versions.md +90 -0
  182. package/docs/{capacity-and-scheduling.md → architecture/capacity-and-scheduling.md} +26 -13
  183. package/docs/{context-engine.md → architecture/context-engine.md} +29 -25
  184. package/docs/{context-working-set.md → architecture/context-working-set.md} +13 -10
  185. package/docs/{dispatch-architecture-rationale.md → architecture/dispatch-architecture-rationale.md} +12 -9
  186. package/docs/{dispatch-typed-intent.md → architecture/dispatch-typed-intent.md} +68 -46
  187. package/docs/{evidence-and-memory.md → architecture/evidence-and-memory.md} +23 -16
  188. package/docs/{middleware-and-components.md → architecture/middleware-and-components.md} +11 -5
  189. package/docs/{model-catalog.md → architecture/model-catalog.md} +61 -27
  190. package/docs/{observability.md → architecture/observability.md} +38 -14
  191. package/docs/{pi-boundary.md → architecture/pi-boundary.md} +24 -11
  192. package/docs/{prompt-envelope-and-tools.md → architecture/prompt-envelope-and-tools.md} +57 -20
  193. package/docs/{provider-adapter-cookbook.md → architecture/provider-adapter-cookbook.md} +99 -25
  194. package/docs/{safety-model.md → architecture/safety-model.md} +35 -20
  195. package/docs/{session-lifecycle.md → architecture/session-lifecycle.md} +8 -5
  196. package/docs/architecture/time-conventions.md +125 -0
  197. package/docs/{trace-store.md → architecture/trace-store.md} +13 -5
  198. package/docs/{tui-design.md → architecture/tui-design.md} +13 -13
  199. package/docs/{worker-dispatch-mechanics.md → architecture/worker-dispatch-mechanics.md} +27 -30
  200. package/docs/{built-in-agents.md → guide/built-in-agents.md} +65 -35
  201. package/docs/{commands-and-modes.md → guide/commands-and-modes.md} +66 -61
  202. package/docs/{configuration-and-targets.md → guide/configuration-and-targets.md} +323 -297
  203. package/docs/guide/configuration-reference.md +1163 -0
  204. package/docs/{environment-variables.md → guide/environment-variables.md} +33 -28
  205. package/docs/{exit-codes-and-output.md → guide/exit-codes-and-output.md} +6 -3
  206. package/docs/{extensions-and-sharing.md → guide/extensions-and-sharing.md} +41 -14
  207. package/docs/{fleet-dispatch.md → guide/fleet-dispatch.md} +39 -43
  208. package/docs/{glossary.md → guide/glossary.md} +14 -11
  209. package/docs/{installation-and-lifecycle.md → guide/installation-and-lifecycle.md} +81 -17
  210. package/docs/guide/panes-and-files.md +290 -0
  211. package/docs/{proactive-memory.md → guide/proactive-memory.md} +131 -107
  212. package/docs/{resource-library.md → guide/resource-library.md} +13 -4
  213. package/docs/{skills-marketplace.md → guide/skills-marketplace.md} +25 -3
  214. package/docs/{tool-usage.md → guide/tool-usage.md} +87 -23
  215. package/docs/{troubleshooting.md → guide/troubleshooting.md} +9 -4
  216. package/docs/{config-knobs-audit.md → history/config-knobs-audit.md} +11 -11
  217. package/docs/{release-cut-checklist.md → history/release-cut-checklist.md} +29 -2
  218. package/docs/process/development-pipeline.md +152 -0
  219. package/docs/process/documentation-coverage.md +100 -0
  220. package/docs/process/documentation-guide.md +187 -0
  221. package/docs/{eval-runner.md → process/eval-runner.md} +108 -53
  222. package/docs/{evals-internal.md → process/evals-internal.md} +10 -10
  223. package/docs/{evolution.md → process/evolution.md} +2 -2
  224. package/docs/{fleet-demo-runbook.md → process/fleet-demo-runbook.md} +11 -7
  225. package/docs/{git-commit-provenance.md → process/git-commit-provenance.md} +11 -4
  226. package/docs/{performance-methodology.md → process/performance-methodology.md} +87 -69
  227. package/docs/{scientific-validation.md → process/scientific-validation.md} +4 -4
  228. package/evals/README.md +2 -2
  229. package/evals/behavioral-model.yaml +3 -2
  230. package/package.json +10 -8
  231. package/skills/README.md +52 -41
  232. package/skills/coding/ast-grep/SKILL.md +102 -31
  233. package/skills/coding/ast-grep/evals.md +26 -0
  234. package/skills/coding/coding-standards/SKILL.md +41 -6
  235. package/skills/coding/coding-standards/evals.md +23 -0
  236. package/skills/coding/prototype/SKILL.md +88 -29
  237. package/skills/coding/prototype/evals.md +19 -0
  238. package/skills/coding/tdd/SKILL.md +81 -54
  239. package/skills/coding/tdd/evals.md +20 -0
  240. package/skills/context/context-handoff/SKILL.md +44 -3
  241. package/skills/context/context-handoff/evals.md +44 -0
  242. package/skills/context/context-prime/SKILL.md +46 -16
  243. package/skills/context/context-prime/evals.md +45 -0
  244. package/skills/git/branch-closeout/SKILL.md +132 -0
  245. package/skills/git/branch-closeout/evals.md +133 -0
  246. package/skills/git/branch-closeout/references/closeout-checklist.md +81 -0
  247. package/skills/git/file-ticket/SKILL.md +78 -64
  248. package/skills/git/file-ticket/assets/issue-template.md +22 -0
  249. package/skills/git/file-ticket/evals.md +31 -26
  250. package/skills/git/file-ticket/references/issue-discovery.md +49 -0
  251. package/skills/git/fix-issue/SKILL.md +88 -65
  252. package/skills/git/fix-issue/evals.md +35 -31
  253. package/skills/git/fix-issue/references/diagnosis-and-rca.md +46 -0
  254. package/skills/git/resolve-merge-conflicts/SKILL.md +101 -52
  255. package/skills/git/resolve-merge-conflicts/evals.md +52 -25
  256. package/skills/git/resolve-merge-conflicts/references/conflict-matrix.md +126 -0
  257. package/skills/git/ship/SKILL.md +103 -67
  258. package/skills/git/ship/assets/pr-template.md +21 -0
  259. package/skills/git/ship/evals.md +44 -28
  260. package/skills/git/ship/references/remote-and-branch-policy.md +62 -0
  261. package/skills/git/worktree-create/SKILL.md +80 -50
  262. package/skills/git/worktree-create/evals.md +40 -33
  263. package/skills/git/worktree-create/references/worktree-setup.md +62 -66
  264. package/skills/git/worktree-merge/SKILL.md +112 -65
  265. package/skills/git/worktree-merge/evals.md +42 -34
  266. package/skills/git/worktree-merge/references/merge-strategies.md +52 -0
  267. package/skills/meta/clio-coder-dev/SKILL.md +9 -5
  268. package/skills/meta/clio-coder-dev/evals.md +3 -2
  269. package/skills/meta/clio-coder-test/SKILL.md +102 -95
  270. package/skills/meta/clio-coder-test/evals.md +9 -4
  271. package/skills/meta/clio-coder-test/references/harness.md +100 -124
  272. package/skills/meta/clio-coder-test/references/test-map.md +77 -50
  273. package/skills/meta/credentials/SKILL.md +2 -2
  274. package/skills/meta/find-skills/SKILL.md +2 -2
  275. package/skills/meta/herdr/SKILL.md +2 -2
  276. package/skills/meta/skill-craft/SKILL.md +22 -16
  277. package/skills/planning/archify/SKILL.md +196 -0
  278. package/skills/planning/archify/evals.md +65 -0
  279. package/skills/planning/architecture/SKILL.md +62 -13
  280. package/skills/planning/architecture/evals.md +65 -0
  281. package/skills/planning/backlog/SKILL.md +131 -15
  282. package/skills/planning/backlog/evals.md +142 -0
  283. package/skills/planning/prd/SKILL.md +47 -7
  284. package/skills/planning/prd/evals.md +54 -0
  285. package/skills/planning/product-intent/SKILL.md +58 -3
  286. package/skills/planning/product-intent/evals.md +70 -0
  287. package/skills/planning/tech-spec/SKILL.md +54 -3
  288. package/skills/planning/tech-spec/evals.md +73 -0
  289. package/skills/registry.yaml +70 -62
  290. package/skills/remote.yaml +13 -0
  291. package/skills/research/arxiv-literature/SKILL.md +77 -19
  292. package/skills/research/arxiv-literature/evals.md +50 -0
  293. package/skills/research/experiment-protocol/SKILL.md +21 -2
  294. package/skills/research/experiment-protocol/evals.md +23 -0
  295. package/skills/research/scientific-debugging/SKILL.md +24 -2
  296. package/skills/research/scientific-debugging/evals.md +18 -0
  297. package/skills/research/scientific-modernization/SKILL.md +27 -2
  298. package/skills/research/scientific-modernization/evals.md +27 -0
  299. package/skills/skill-marketplace.json +97 -62
  300. package/skills/workflow/cut-it/SKILL.md +66 -6
  301. package/skills/workflow/cut-it/evals.md +101 -0
  302. package/skills/workflow/design-council/SKILL.md +118 -28
  303. package/skills/workflow/design-council/evals.md +161 -0
  304. package/skills/workflow/grill-me/SKILL.md +87 -11
  305. package/skills/workflow/grill-me/evals.md +153 -0
  306. package/skills/workflow/workflow-distiller/SKILL.md +77 -18
  307. package/skills/workflow/workflow-distiller/evals.md +118 -0
  308. package/src/cli/args.ts +2 -2
  309. package/src/cli/bootstrap-generate.ts +1 -1
  310. package/src/cli/config-inspect.ts +65 -12
  311. package/src/cli/configure-interop.ts +105 -13
  312. package/src/cli/configure-oauth.ts +57 -0
  313. package/src/cli/configure-onboarding.ts +980 -0
  314. package/src/cli/configure-target.ts +594 -0
  315. package/src/cli/configure.ts +1082 -532
  316. package/src/cli/context-map.ts +114 -0
  317. package/src/cli/context.ts +4 -0
  318. package/src/cli/docs.ts +22 -14
  319. package/src/cli/doctor-naming.ts +5 -5
  320. package/src/cli/doctor-toolchain.ts +3 -3
  321. package/src/cli/eval.ts +1 -2
  322. package/src/cli/extensions.ts +2 -1
  323. package/src/cli/fleet.ts +1 -1
  324. package/src/cli/index.ts +3 -1
  325. package/src/cli/internal-dispatch.ts +3 -4
  326. package/src/cli/lifecycle-presenter.ts +436 -0
  327. package/src/cli/models.ts +10 -2
  328. package/src/cli/modes/print.ts +5 -1
  329. package/src/cli/panes.ts +19 -5
  330. package/src/cli/reset.ts +228 -106
  331. package/src/cli/run.ts +9 -4
  332. package/src/cli/select.ts +664 -0
  333. package/src/cli/share.ts +5 -1
  334. package/src/cli/skills-eval.ts +3 -3
  335. package/src/cli/skills.ts +9 -2
  336. package/src/cli/targets.ts +5 -6
  337. package/src/cli/trace.ts +55 -4
  338. package/src/cli/uninstall.ts +233 -165
  339. package/src/cli/upgrade.ts +204 -149
  340. package/src/cli/usage.ts +86 -27
  341. package/src/cli/validate-model.ts +3 -3
  342. package/src/cli/wiki-generate.ts +1 -1
  343. package/src/core/artifact-paths.ts +1 -1
  344. package/src/core/bash-exec.ts +131 -86
  345. package/src/core/bus-events.ts +51 -6
  346. package/src/core/config.ts +61 -1
  347. package/src/core/defaults.ts +7 -4
  348. package/src/core/dispatch-outcome.ts +16 -0
  349. package/src/core/external-diagnostic.ts +44 -0
  350. package/src/core/gateway-routing.ts +157 -0
  351. package/src/core/guardrails.ts +10 -49
  352. package/src/core/prompt-hint.ts +9 -0
  353. package/src/core/safe-exec.ts +17 -2
  354. package/src/core/skill-activation.ts +89 -2
  355. package/src/domains/agents/builtins/architect.md +2 -3
  356. package/src/domains/agents/builtins/coder.md +3 -2
  357. package/src/domains/agents/builtins/debugger.md +2 -2
  358. package/src/domains/agents/builtins/documenter.md +2 -2
  359. package/src/domains/agents/builtins/git-master.md +1 -1
  360. package/src/domains/agents/builtins/oracle.md +1 -1
  361. package/src/domains/agents/builtins/provenance.md +1 -1
  362. package/src/domains/agents/builtins/researcher.md +1 -1
  363. package/src/domains/agents/builtins/scout.md +1 -1
  364. package/src/domains/agents/builtins/tester.md +2 -2
  365. package/src/domains/agents/builtins/verifier.md +2 -2
  366. package/src/domains/agents/builtins/wiki-writer.md +1 -1
  367. package/src/domains/agents/builtins/world-knowledge.md +31 -0
  368. package/src/domains/agents/catalog.ts +13 -15
  369. package/src/domains/agents/contract.ts +2 -0
  370. package/src/domains/agents/extension.ts +23 -1
  371. package/src/domains/agents/result-contract.ts +70 -0
  372. package/src/domains/config/keybindings.ts +8 -0
  373. package/src/domains/context/extension.ts +0 -3
  374. package/src/domains/context/wiki/map-seed.ts +589 -0
  375. package/src/domains/context/wiki/plan.ts +2 -2
  376. package/src/domains/context/working-set/path-index.ts +1 -0
  377. package/src/domains/dispatch/admission.ts +29 -0
  378. package/src/domains/dispatch/agent-candidates.ts +10 -0
  379. package/src/domains/dispatch/budget-envelope.ts +86 -1
  380. package/src/domains/dispatch/capability-match.ts +11 -0
  381. package/src/domains/dispatch/capacity-lease.ts +17 -0
  382. package/src/domains/dispatch/contract.ts +11 -1
  383. package/src/domains/dispatch/extension.ts +237 -49
  384. package/src/domains/dispatch/host-verification.ts +435 -39
  385. package/src/domains/dispatch/intent-requirements.ts +10 -0
  386. package/src/domains/dispatch/intent.ts +18 -1
  387. package/src/domains/dispatch/path-scope.ts +235 -24
  388. package/src/domains/dispatch/run-event-journal.ts +4 -15
  389. package/src/domains/dispatch/state.ts +2 -3
  390. package/src/domains/dispatch/transport.ts +45 -21
  391. package/src/domains/dispatch/types.ts +58 -3
  392. package/src/domains/dispatch/worker-model-metadata.ts +38 -0
  393. package/src/domains/eval/artifacts/store.ts +5 -0
  394. package/src/domains/eval/metrics/call-ledger-stream.ts +34 -11
  395. package/src/domains/eval/metrics/token-stream.ts +201 -31
  396. package/src/domains/eval/metrics/tracked.ts +40 -4
  397. package/src/domains/eval/runners/clio-run.ts +5 -2
  398. package/src/domains/eval/schema/suite.ts +28 -0
  399. package/src/domains/eval/schema/verdict.ts +2 -2
  400. package/src/domains/eval/store.ts +8 -1
  401. package/src/domains/eval/suites/resolve.ts +13 -1
  402. package/src/domains/eval/suites/run.ts +24 -3
  403. package/src/domains/evidence/trust-status.ts +10 -1
  404. package/src/domains/extensions/contract.ts +15 -1
  405. package/src/domains/extensions/discovery.ts +238 -41
  406. package/src/domains/extensions/extension.ts +105 -6
  407. package/src/domains/extensions/index.ts +24 -0
  408. package/src/domains/extensions/integrity.ts +189 -0
  409. package/src/domains/extensions/manager.ts +17 -1
  410. package/src/domains/extensions/resource-path.ts +27 -0
  411. package/src/domains/extensions/resources.ts +18 -38
  412. package/src/domains/extensions/snapshot-store.ts +39 -0
  413. package/src/domains/extensions/snapshot.ts +180 -0
  414. package/src/domains/extensions/state.ts +385 -57
  415. package/src/domains/extensions/types.ts +118 -1
  416. package/src/domains/interop/registry.ts +6 -2
  417. package/src/domains/interop/types.ts +4 -0
  418. package/src/domains/lifecycle/migrations/2026-09-01-extension-install-digests.ts +27 -0
  419. package/src/domains/lifecycle/migrations/index.ts +6 -0
  420. package/src/domains/lifecycle/naming-resources.ts +19 -4
  421. package/src/domains/lifecycle/naming-yazi.ts +10 -5
  422. package/src/domains/memory/task-memory-policy.ts +70 -26
  423. package/src/domains/memory/task-memory-telemetry.ts +1 -0
  424. package/src/domains/middleware/contract.ts +26 -0
  425. package/src/domains/middleware/extension.ts +24 -24
  426. package/src/domains/middleware/hook-receipts.ts +27 -4
  427. package/src/domains/middleware/hooks-io.ts +65 -32
  428. package/src/domains/middleware/hooks.ts +64 -0
  429. package/src/domains/middleware/index.ts +28 -5
  430. package/src/domains/middleware/marketplace-offer.ts +3 -35
  431. package/src/domains/middleware/memory-intervention.ts +127 -32
  432. package/src/domains/middleware/memory-step-endpoint.ts +3 -2
  433. package/src/domains/middleware/registrations.ts +326 -0
  434. package/src/domains/middleware/runtime.ts +28 -0
  435. package/src/domains/middleware/skills-reminder.ts +31 -2
  436. package/src/domains/middleware/snapshot.ts +20 -7
  437. package/src/domains/mux/contract.ts +38 -0
  438. package/src/domains/mux/detect.ts +6 -13
  439. package/src/domains/mux/index.ts +1 -1
  440. package/src/domains/mux/operations.ts +44 -5
  441. package/src/domains/mux/yazi/assets/yazi.toml +2 -2
  442. package/src/domains/mux/yazi/session.ts +53 -4
  443. package/src/domains/mux/yazi/theme.ts +117 -17
  444. package/src/domains/observability/compaction-usage.ts +118 -0
  445. package/src/domains/observability/contract.ts +10 -11
  446. package/src/domains/observability/cost.ts +1 -1
  447. package/src/domains/observability/extension.ts +17 -4
  448. package/src/domains/observability/out-of-turn-usage.ts +52 -21
  449. package/src/domains/observability/projection.ts +14 -90
  450. package/src/domains/observability/trace-store.ts +43 -7
  451. package/src/domains/prompts/compiler.ts +73 -53
  452. package/src/domains/prompts/contract.ts +15 -3
  453. package/src/domains/prompts/extension.ts +97 -9
  454. package/src/domains/prompts/fragments/identity/clio-worker.md +1 -3
  455. package/src/domains/prompts/fragments/identity/clio.md +6 -12
  456. package/src/domains/prompts/fragments/identity/docs-routing.md +1 -2
  457. package/src/domains/prompts/fragments/identity/self-awareness.md +3 -11
  458. package/src/domains/prompts/fragments/operating/contract.md +7 -15
  459. package/src/domains/prompts/fragments/operating/delegation.md +32 -34
  460. package/src/domains/prompts/fragments/operating/skills.md +10 -24
  461. package/src/domains/prompts/fragments/operating/worker.md +1 -8
  462. package/src/domains/providers/contract.ts +4 -1
  463. package/src/domains/providers/extension.ts +40 -9
  464. package/src/domains/providers/index.ts +1 -1
  465. package/src/domains/providers/model-capabilities.ts +9 -0
  466. package/src/domains/providers/model-discovery.ts +2 -0
  467. package/src/domains/providers/model-runtime-capabilities.ts +99 -25
  468. package/src/domains/providers/models/local-models/clio-coder-local-coding-targets.yaml +699 -114
  469. package/src/domains/providers/runtime-resolution.ts +31 -0
  470. package/src/domains/providers/runtimes/antigravity/antigravity-code.ts +225 -45
  471. package/src/domains/providers/runtimes/common/lmstudio-http.ts +6 -2
  472. package/src/domains/providers/runtimes/common/local-synth.ts +2 -0
  473. package/src/domains/providers/runtimes/common/probe-helpers.ts +7 -2
  474. package/src/domains/providers/runtimes/local-native/llamacpp.ts +9 -1
  475. package/src/domains/providers/runtimes/protocol/litellm.ts +119 -29
  476. package/src/domains/providers/support.ts +11 -5
  477. package/src/domains/providers/target-model-cache.ts +25 -2
  478. package/src/domains/providers/types/capability-flags.ts +2 -0
  479. package/src/domains/providers/types/cost-provenance.ts +19 -0
  480. package/src/domains/providers/types/local-model-quirks.ts +85 -37
  481. package/src/domains/providers/types/runtime-descriptor.ts +20 -1
  482. package/src/domains/providers/types/target-descriptor.ts +19 -0
  483. package/src/domains/resources/index.ts +3 -0
  484. package/src/domains/resources/skills/install.ts +72 -7
  485. package/src/domains/resources/skills/loader.ts +23 -19
  486. package/src/domains/resources/skills/marketplace.ts +63 -11
  487. package/src/domains/safety/autonomy.ts +15 -0
  488. package/src/domains/safety/call-target.ts +1 -1
  489. package/src/domains/safety/index.ts +1 -0
  490. package/src/domains/safety/loop-detector.ts +7 -4
  491. package/src/domains/safety/path-policy.ts +1 -1
  492. package/src/domains/safety/policy-engine.ts +34 -11
  493. package/src/domains/safety/protected-artifacts.ts +191 -88
  494. package/src/domains/safety/run-effects.ts +2 -22
  495. package/src/domains/safety/skill-authority.ts +55 -0
  496. package/src/domains/session/compaction/compact.ts +72 -22
  497. package/src/domains/session/entries.ts +6 -0
  498. package/src/domains/session/task-board.ts +10 -9
  499. package/src/domains/session/usage.ts +3 -3
  500. package/src/domains/share/archive.ts +164 -7
  501. package/src/engine/acp/server.ts +62 -9
  502. package/src/engine/agent.ts +13 -3
  503. package/src/engine/ai.ts +26 -8
  504. package/src/engine/antigravity/subprocess-runtime.ts +386 -120
  505. package/src/engine/api-registry.ts +3 -0
  506. package/src/engine/apis/llamacpp-residency.ts +3 -4
  507. package/src/engine/apis/lmstudio.ts +3 -3
  508. package/src/engine/apis/ollama-native.ts +6 -6
  509. package/src/engine/apis/openai-completions.ts +145 -39
  510. package/src/engine/apis/output-budget.ts +8 -18
  511. package/src/engine/apis/residency.ts +8 -27
  512. package/src/engine/external-subprocess.ts +114 -6
  513. package/src/engine/gemma-channel-filter.ts +19 -0
  514. package/src/engine/loop-guard.ts +92 -12
  515. package/src/engine/worker-runtime.ts +40 -11
  516. package/src/engine/worker-tools.ts +3 -1
  517. package/src/entry/background-model-metadata.ts +18 -0
  518. package/src/entry/compaction-prompt.ts +57 -0
  519. package/src/entry/extension-hook-sources.ts +28 -0
  520. package/src/entry/extension-reload.ts +309 -0
  521. package/src/entry/orchestrator.ts +464 -251
  522. package/src/entry/task-memory-lifecycle.ts +35 -0
  523. package/src/interactive/application-controller.ts +2 -1
  524. package/src/interactive/bus-notices.ts +8 -1
  525. package/src/interactive/chat-loop-messages.ts +16 -17
  526. package/src/interactive/chat-loop.ts +75 -3
  527. package/src/interactive/chat-panel.ts +36 -13
  528. package/src/interactive/chat-renderer.ts +72 -7
  529. package/src/interactive/cost-overlay.ts +26 -2
  530. package/src/interactive/dispatch-board.ts +6 -11
  531. package/src/interactive/footer/widgets.ts +13 -0
  532. package/src/interactive/interactive-application.ts +39 -4
  533. package/src/interactive/interactive-input-runtime.ts +4 -0
  534. package/src/interactive/interactive-presentation.ts +2 -2
  535. package/src/interactive/interactive-slash-runtime.ts +4 -1
  536. package/src/interactive/overlays/extensions.ts +9 -1
  537. package/src/interactive/overlays/help-reference.ts +13 -0
  538. package/src/interactive/overlays/settings.ts +27 -16
  539. package/src/interactive/panes-runtime.ts +111 -35
  540. package/src/interactive/prompt-cache-identity.ts +88 -0
  541. package/src/interactive/renderers/worker-entry.ts +32 -0
  542. package/src/interactive/slash-commands.ts +153 -20
  543. package/src/interactive/stream-pacing-policy.ts +0 -23
  544. package/src/interactive/theme/labels.ts +19 -13
  545. package/src/interactive/turn-context.ts +39 -20
  546. package/src/interactive/turn-recovery.ts +8 -0
  547. package/src/interactive/turn-runtime.ts +27 -11
  548. package/src/interactive/turn-state.ts +7 -0
  549. package/src/interactive/worker-receipts.ts +1 -0
  550. package/src/interactive/worker-stream.ts +6 -1
  551. package/src/interactive/yazi-bridge.ts +60 -6
  552. package/src/tools/agent-tools.ts +30 -1
  553. package/src/tools/artifact.ts +2 -2
  554. package/src/tools/ask-user.ts +3 -3
  555. package/src/tools/bash.ts +1 -1
  556. package/src/tools/bootstrap.ts +4 -0
  557. package/src/tools/builtin-tool-catalog.ts +52 -22
  558. package/src/tools/codewiki/code-nav-surface.ts +6 -0
  559. package/src/tools/codewiki/code-nav.ts +99 -13
  560. package/src/tools/context/docs-engine.ts +20 -7
  561. package/src/tools/context/index.ts +59 -21
  562. package/src/tools/core-bootstrap.ts +28 -6
  563. package/src/tools/credential-present.ts +1 -2
  564. package/src/tools/dispatch-arguments.ts +6 -1
  565. package/src/tools/dispatch-event-text.ts +10 -0
  566. package/src/tools/dispatch-plan.ts +49 -4
  567. package/src/tools/dispatch-run-events.ts +1 -1
  568. package/src/tools/dispatch-runner.ts +12 -0
  569. package/src/tools/dispatch-schema.ts +338 -0
  570. package/src/tools/dispatch-types.ts +3 -0
  571. package/src/tools/dispatch.ts +9 -254
  572. package/src/tools/ledger.ts +3 -5
  573. package/src/tools/monitor-surface.ts +5 -13
  574. package/src/tools/observation.ts +4 -5
  575. package/src/tools/panes-surface.ts +4 -11
  576. package/src/tools/panes.ts +4 -2
  577. package/src/tools/policy.ts +15 -2
  578. package/src/tools/read.ts +5 -6
  579. package/src/tools/registry.ts +41 -12
  580. package/src/tools/result-shaping.ts +18 -14
  581. package/src/tools/steer-surface.ts +1 -1
  582. package/src/tools/tasks.ts +1 -1
  583. package/src/tools/truncate.ts +6 -5
  584. package/src/tools/verify/surface.ts +6 -12
  585. package/src/tools/web-fetch-surface.ts +1 -3
  586. package/src/tools/worker-evidence.ts +3 -1
  587. package/src/worker/spec-contract.ts +4 -0
  588. package/dist/builtins-UJLMOVOV.js +0 -17
  589. package/dist/chunk-5QIAJV2D.js +0 -48
  590. package/dist/chunk-JZWT5J3Y.js +0 -814
  591. package/dist/chunk-K7VKOLQQ.js +0 -15
  592. package/dist/chunk-PMZCIOCJ.js +0 -25
  593. package/dist/chunk-SUW5DORT.js +0 -819
  594. package/dist/chunk-UOV2BYIW.js +0 -107
  595. package/dist/chunk-WR6U3OVP.js +0 -45
  596. package/dist/chunk-Y45G3AXC.js +0 -1558
  597. package/dist/reset-EOLM7GVE.js +0 -230
  598. package/dist/uninstall-N34PCTGJ.js +0 -331
  599. package/dist/upgrade-H7TOM7YL.js +0 -323
  600. package/docs/artifact-versions.md +0 -67
  601. package/docs/development-pipeline.md +0 -121
  602. package/docs/documentation-coverage.md +0 -46
  603. package/docs/documentation-guide.md +0 -167
  604. package/docs/time-conventions.md +0 -101
@@ -1,11 +1,11 @@
1
1
  # Clio Coder Local Evaluation Runner
2
2
 
3
- > [!TIP]
4
- > **Interactive Spec Available:** An interactive task suite validator, subprocess execution simulator, and compare calculator is located at [docs/html/eval_blueprint.html](html/eval_blueprint.html) (Version: 0.4.0).
3
+ > **Visual blueprint:** The source checkout includes the complete
4
+ > [Clio Coder Local Evaluation Runner visual reference](https://github.com/iowarp/clio-coder/blob/main/docs/html/eval_blueprint.html).
5
5
 
6
6
  The local evaluation runner executes repository-local YAML task suites as deterministic subprocess checks. It is useful for comparing harness changes, prompts, tools, or local workflows.
7
7
 
8
- Source of truth: [src/domains/eval/](../src/domains/eval/) and [src/cli/eval.ts](../src/cli/eval.ts).
8
+ Source of truth: [src/domains/eval/](../../src/domains/eval/) and [src/cli/eval.ts](../../src/cli/eval.ts).
9
9
 
10
10
  ---
11
11
 
@@ -20,6 +20,7 @@ clio-coder eval run --task-file <tasks.yaml> [--repeat <n>] [--out <path>] [--cl
20
20
  clio-coder eval report <evalId> --format text|json|md|swe-jsonl|junit
21
21
  clio-coder eval compare <baselineEvalId> <candidateEvalId> [--metric <name>] [--format text|json|md|junit] [--allow-config-drift]
22
22
  clio-coder eval gate <candidateEvalId> --baseline <baselineEvalId> [--thresholds <file>]
23
+ clio-coder eval inventory --json
23
24
  ```
24
25
 
25
26
  ### Command Roles
@@ -33,6 +34,7 @@ clio-coder eval gate <candidateEvalId> --baseline <baselineEvalId> [--thresholds
33
34
  * `junit`: XML report for CI/CD integration.
34
35
  * **`compare`**: Compares two evaluation artifacts (baseline and candidate) by matching tasks.
35
36
  * **`gate`**: Compares candidate metrics against baseline and absolute thresholds. Correctness and safety regressions fail independently of informational budgets.
37
+ * **`inventory`**: Prints the fixed machine-readable inventory used by GUI hosts. It includes stored report identity, provenance, serving facts, accounting, and per-scenario outcomes without report attachments.
36
38
 
37
39
  Exit codes:
38
40
 
@@ -42,7 +44,7 @@ Exit codes:
42
44
  | `eval run` | `0` when all task repetitions pass | `1` when any task fails, `2` for invalid configs |
43
45
  | `eval report` | `0` when artifact loads | `1` if artifact cannot be read, `2` for invalid ID |
44
46
  | `eval compare` | `0` when both artifacts compare and the behavioral hard gate passes | `1` for a hard regression or unreadable artifact, `2` for invalid ID |
45
- | `eval gate` | `0` when correctness, safety, and hard threshold assertions pass | `1` for any hard failure, `2` for config/invalid ID errors |
47
+ | `eval gate` | `0` when correctness, safety, and hard threshold assertions pass | `1` for hard failures, unreadable inputs, and malformed threshold files; `2` for an invalid eval ID or usage error |
46
48
 
47
49
  ---
48
50
 
@@ -75,7 +77,7 @@ tasks:
75
77
  excludes:
76
78
  - "**/node_modules/**"
77
79
  runner:
78
- kind: "clio-run" # clio-run | context-index | context-init | external-command
80
+ kind: "clio-coder-run" # clio-coder-run | context-index | context-init | external-command
79
81
  prompt: "Optimize the FFT tolerance bounds in solver.ts"
80
82
  timeoutMs: 60000
81
83
  verify:
@@ -103,12 +105,12 @@ tasks:
103
105
  | --- | --- | --- |
104
106
  | `version` | - | Must equal `2`. |
105
107
  | `suite` | `id`, `title`, `visibility`, `description` | Metadata identifying the evaluation suite. |
106
- | `matrix` | `targets[]`, `repeats`, `dimensions[]` | Matrix of execution targets, repetition count, and the execution-envelope fields intentionally varied by the suite. |
107
- | `workspace` | `kind`, `path`, `url`, `commit`, `checkout`, `excludes` | Workspace strategy: `local` (run in-place), `git` (clone from URL), or `temp-copy` (isolated copy of a directory). |
108
- | `runner` | `kind`, `prompt`, `command`, `commands`, `args`, `timeoutMs` | Runner type: `clio-run` (starts Clio agent loop), `context-index` (runs indexer), `context-init` (initializes context), `external-command` (spawns subprocess). |
109
- | `behavioral` | `schema`, `corpus`, `execution`, `expectedBehavior`, `forbiddenBehavior`, `judge` | Optional `clio.eval.scenario.v1` behavioral contract. Rules name a closed category and a typed predicate over transcript, tool, receipt, or grader facts. |
108
+ | `matrix` | `targets[]`, `repeats`, `dimensions[]`, `maxCostUsd` | Matrix of execution targets, repetition count, execution-envelope fields intentionally varied by the suite, and an optional cumulative known-cost ceiling. |
109
+ | `workspace` | `kind`, `path`, `url`, `commit`, `checkout`, `excludes`, `setup` | Workspace strategy: `local` (run in-place), `git` (clone from URL), or `temp-copy` (isolated copy of a directory). Optional `setup` commands prepare the workspace before the runner starts. |
110
+ | `runner` | `kind`, `prompt`, `autonomy`, `agent`, `command`, `commands`, `args`, `timeoutMs` | Runner type: `clio-coder-run` (starts Clio's agent loop), `context-index` (runs the indexer), `context-init` (initializes context), or `external-command` (spawns a subprocess). `agent` selects a worker recipe and `autonomy` sets one-run headless authority. |
111
+ | `behavioral` | `schema`, `corpus`, `execution`, `expectedBehavior`, `forbiddenBehavior`, `judge` | Optional `clio-coder.eval.scenario.v1` behavioral contract. Rules name a closed category and a typed predicate over transcript, tool, receipt, or grader facts. |
110
112
  | `verify` | `commands`, `measure`, `assertions`, `forbidPaths` | Validation steps: shell commands, a task-outcome grader, metric assertions (e.g. `op: lt` for max token counts), and files/directories that must not be created or modified (`forbidPaths`). |
111
- | `metrics` | `collect` | List of metric names to compile for the evaluation runs. |
113
+ | `metrics` | `collect`, `readObservation` | Metric names to compile plus optional public allowlisted and decoy paths reduced to bounded read counters. Raw path strings do not enter behavioral facts. |
112
114
 
113
115
  ---
114
116
 
@@ -120,7 +122,7 @@ tasks:
120
122
  ---
121
123
 
122
124
  ## Runner Kinds
123
- * **`clio-run`**: Invokes the main Clio Coder agent loop with the task's prompt, tracing all tools.
125
+ * **`clio-coder-run`**: Invokes the main Clio Coder agent loop with the task's prompt, tracing all tools. The released `clio-run` spelling is accepted only as a legacy input alias and is normalized before validation; writers and new suites use `clio-coder-run`.
124
126
  * **`context-index`**: Triggers the context engine to build index structures (`codewiki`).
125
127
  * **`context-init`**: Initializes workspace files (such as generating `CLIO-CODER.md`).
126
128
  * **`external-command`**: Spawns an external command or sequence of commands in the task workspace.
@@ -137,7 +139,37 @@ Metrics collected during runs can be validated automatically using the `verify.a
137
139
  * `eq` (equal)
138
140
  * `neq` (not equal)
139
141
 
140
- Metrics that can be validated include `tokens.input`, `tokens.output`, `tokens.total`, `latency.wallMs`, `tools.totalCalls`, `tools.failed`, `tools.blocked`, `verifier.exitCode`, and `result.pass`.
142
+ Metrics that can be validated include `tokens.input`, `tokens.output`, `tokens.total`, `latency.wallMs`, `tools.totalCalls`, `tools.failed`, `tools.blocked`, `verifier.exitCode`, `result.pass`, and the `provider.*` metrics below. Each `verify.assertions` condition must hold; an unavailable metric fails closed.
143
+
144
+ ### Optional Provider-Health Gates
145
+
146
+ A task that recovers from a provider error can still pass its task checks by default. Provider health is a separate, opt-in requirement. To require observed provider events and no observed terminal errors, add these assertions to the task:
147
+
148
+ ```yaml
149
+ verify:
150
+ assertions:
151
+ - metric: "provider.measured"
152
+ op: "eq"
153
+ value: true
154
+ - metric: "provider.stopReason.error"
155
+ op: "eq"
156
+ value: 0
157
+ ```
158
+
159
+ Suite-level `thresholds.fail` uses the opposite condition: a matching condition is a failure. The corresponding hard gate checks each run as follows:
160
+
161
+ ```yaml
162
+ thresholds:
163
+ fail:
164
+ - metric: "provider.measured"
165
+ op: "eq"
166
+ value: false
167
+ - metric: "provider.stopReason.error"
168
+ op: "gt"
169
+ value: 0
170
+ ```
171
+
172
+ Unavailable metrics also fail closed in a hard threshold. `thresholds.informational` records findings without changing exit status. These examples reject an observed error followed by a successful recovery while leaving the default ungated task-pass policy unchanged. To reject any observed retry start or assistant abort as well, add conditions on `provider.retryStarted` or `provider.stopReason.aborted` with the same assertion-versus-failure polarity.
141
173
 
142
174
  ---
143
175
 
@@ -173,14 +205,44 @@ Under the hood, these are parsed and wrapped into a Suite v2 adapter with:
173
205
 
174
206
  Clio maintains two distinct token accounting streams with different provenances. These accounts are never merged, reconciled, or treated as interchangeable:
175
207
 
176
- 1. **`tokens.*` (Wire Streaming)**: Folded live off stdout from assistant `message_end` events watched by `token-stream.ts` / `createStreamInvariantFold`. This represents usage reported by the provider for assistant messages watched over the wire. On surfaces without stdout streaming (such as `clio-coder fleet run --json`), `tokens.measured` is `false`.
208
+ 1. **`tokens.*` (Wire Streaming)**: Folded live off stdout from assistant `message_end` events. These totals include all known usage on errored calls as well as successful calls; recovery never subtracts earlier spend. Only finite, nonnegative usage facts are admitted. On surfaces without the relevant stdout events (such as `clio-coder fleet run --json`), `tokens.measured` is `false`.
177
209
  2. **`receiptUsage.*` (Journal Receipts)**: Summed from an evaluation item's run journal. Every attempt writes a receipt carrying token counts and USD cost authenticated against its own ledger envelope.
178
210
 
179
211
  ### Fail-Closed Reporting
180
212
  Both accounting streams report unmeasured state with no counts at all rather than a numeric zero. Reporting zero for an unmeasured run would falsely claim the run cost nothing. On an unmeasured run, `tokens.total` resolves to `null` and fails closed on metric threshold comparisons.
181
213
 
214
+ An errored call must carry at least one positive reported token, reasoning, or cost fact before its usage is considered observed. Reasoning-only or cost-only observations do not establish ordinary token totals. A stream containing only errored calls with missing or synthetic all-zero usage remains `tokens.measured: false`. The current event shape cannot distinguish synthetic all-zero failures from genuinely reported zero usage, so it cannot establish measured zero spending in either case. Partial positive usage remains included as known spend. On failed calls, adapters can also initialize individual absent fields to zero; those zeros remain unattributed and make coverage incomplete. The existing inclusive numeric fields are known subtotals, so their zeros do not prove complete zero spending when failed usage is incomplete.
215
+
216
+ ### Provider Observations and Failed-Call Share
217
+
218
+ The `provider.*` metrics describe events observed on live stdout, folded before diagnostic output is truncated. Native runs and multi-command external runners retain these observations from their executed commands. They do not enumerate SDK-internal retries or network attempts that were never emitted. Filtered or opaque output can leave provider health unobserved even when the process exits successfully or a receipt reports task success.
219
+
220
+ | Metric | Meaning |
221
+ | --- | --- |
222
+ | `provider.measured` | Whether an assistant terminal reason or a counted retry phase was observed. With no such observations this is `false`, and provider counters are absent. It does not certify complete provider coverage. |
223
+ | `provider.stopReason.stop`, `provider.stopReason.toolUse`, `provider.stopReason.length` | Counts of these terminal reasons on assistant `message_end` events. |
224
+ | `provider.stopReason.error`, `provider.stopReason.aborted`, `provider.stopReason.other` | Separate counts of errored, aborted, and other observed terminal reasons. An unrecognized terminal reason goes into `other`. Partial updates and repeated messages in `turn_end` or `agent_end` do not add counts. |
225
+ | `provider.retryScheduled` | Observed `scheduled` phases: planned retries, including ones cancelled before execution. |
226
+ | `provider.retryStarted` | Observed `retrying` phases: retry execution starts. Repeated `waiting` countdown frames do not count as attempts. |
227
+ | `provider.retryCancelled`, `provider.retryExhausted`, `provider.retryRecovered` | Counts of the corresponding observed phases. They describe retry-chain outcomes and do not fabricate additional assistant calls. Attempt numbers can restart for each chain. |
228
+ | `provider.errorUsageObservedCalls` | Errored calls with at least one positive reported token, reasoning, or cost fact. |
229
+ | `provider.errorUsageUnobservedCalls` | Errored calls with no positive reported token, reasoning, or cost fact, including absent or all-zero usage. |
230
+ | `provider.errorUsageIncompleteCalls` | Errored calls with unobserved usage, incomplete token fields, or ambiguous normalized zero fields. This can overlap `errorUsageObservedCalls` when only part of the usage is known. |
231
+ | `provider.errorCostUnobservedCalls` | Errored calls without a positive cost fact. Zero or absent cost does not prove that an error was free, including when some token usage is known. |
232
+ | `provider.errorTokens.input`, `provider.errorTokens.output`, `provider.errorTokens.total`, `provider.errorTokens.cacheRead`, `provider.errorTokens.cacheWrite` | Known positive token subtotals for errored calls. Absent or ambiguous zero fields remain unattributed; when total usage is absent, a total can still be summed from known token fields and remains incomplete. |
233
+ | `provider.errorCostUsd` | Known positive cost subtotal from errored calls' stream usage objects. Cost can come from adapter pricing; it is not independently certified provider billing. |
234
+ | `provider.errorReasoningTokens`, `provider.errorReasoningUnobservedCalls` | Known failed-call reasoning subtotal and calls without attributable reported reasoning. Reasoning can overlap output, so it is never added to ordinary token totals. |
235
+
236
+ The failed-call share covers `stopReason: error`; aborted calls remain separately labeled. Failed-share amounts and usage-coverage counters appear only after an errored terminal message is observed. These share metrics supplement the inclusive `tokens.*` totals without changing the summary token shape, receipt schema, or verdict schema. Missing or partial failed usage makes them known subtotals, not a complete amount to subtract from total spend. The native runner's `cost.usd` can use receipt evidence, so equality with the stream-based `provider.errorCostUsd` is not guaranteed.
237
+
238
+ Positive reported reasoning uses the normalized `usage.reasoning` field, `reasoning_tokens`, and supported nested provider-detail fields. The root `reasoningTokens` alias can be an adapter estimate without a provenance marker, so it is left unattributed rather than promoted to reported provider usage. Normalized zero reasoning is also unattributed because adapters can fill it when provider detail is absent. A reasoning-only failure is observed but has incomplete ordinary-token coverage; no output or total is inferred from it.
239
+
240
+ These observations also do not reconcile the separate `trackedMetrics` ledger selection (#276). Tracked metrics prefer durable assistant-call facts when available, retain durable compaction and tool records, and otherwise fall back to stream calls. Artifacts expose source counts and warnings, but partial or mixed ledgers can omit stream-only calls and fork-inherited history remains unreconciled. Neither those tracked values nor the new provider counters prove complete run accounting; failed-compaction usage is retained separately in the out-of-turn usage ledger and usage report, and is not included by this eval fold.
241
+
182
242
  ---
183
243
 
244
+ Full reconciliation across session, stdout, fork and out-of-turn evidence is deferred to v0.4.5 or later. Version 0.4.3 does not add an automatic rejection of cost or efficiency comparisons merely because those sources are partial or mixed. Matching source counts do not prove complete coverage or shared call identity. Existing missing-metric, serving-configuration and execution-envelope comparison gates still apply.
245
+
184
246
  ## Eval Artifact Format (v4)
185
247
 
186
248
  Evaluation artifacts use format version 4 (`EvalArtifactV4`). Summary token metrics report `measuredRuns` out of total `runs`:
@@ -190,7 +252,7 @@ export interface EvalArtifactV4 {
190
252
  version: 4;
191
253
  evalId: string;
192
254
  suite: { id: string; hash: string };
193
- clio: EvalClioProvenance;
255
+ clioCoder: EvalClioProvenance;
194
256
  environment: EvalEnvironmentProvenance;
195
257
  matrix: { target: string; model: string | null; thinking: string | null };
196
258
  summary: EvalArtifactSummaryV4;
@@ -206,11 +268,11 @@ export interface EvalArtifactV4 {
206
268
 
207
269
  ## The verdict envelope
208
270
 
209
- Every result carries a strictly parsed `clio.eval.verdict.v1` envelope (`src/domains/eval/schema/verdict.ts`). Suite v2 results are adapted into it at one explicit boundary (`src/domains/eval/schema/adapter.ts`) rather than by widening the artifact version, because the envelope carries no information a v4 artifact cannot hold.
271
+ Every result carries a strictly parsed `clio-coder.eval.verdict.v1` envelope (`src/domains/eval/schema/verdict.ts`). Suite v2 results are adapted into it at one explicit boundary (`src/domains/eval/schema/adapter.ts`) rather than by widening the artifact version, because the envelope carries no information a v4 artifact cannot hold.
210
272
 
211
273
  ```json
212
274
  {
213
- "schema": "clio.eval.verdict.v1",
275
+ "schema": "clio-coder.eval.verdict.v1",
214
276
  "scenarioId": "latency-nonnegative",
215
277
  "trialIndex": 0,
216
278
  "outcome": "pass",
@@ -230,7 +292,7 @@ The envelope is fail-closed by construction. `outcome` is one of `pass`, `fail`,
230
292
 
231
293
  ### Behavioral scenario and verdict documents
232
294
 
233
- Behavioral evaluation is additive and does not change the persisted `clio.eval.verdict.v1` reader. A Suite v2 task may declare a `clio.eval.scenario.v1` block, and its Artifact v4 result then carries a sibling `clio.eval.behavior.v1` document whose `verdictRef` names the verdict schema, scenario id, and trial index. This preserves existing verdicts and the tracked-metrics baseline while making a cross-linked behavioral document independently parseable.
295
+ Behavioral evaluation is additive and does not change the persisted `clio-coder.eval.verdict.v1` reader. A Suite v2 task may declare a `clio-coder.eval.scenario.v1` block, and its Artifact v4 result then carries a sibling `clio-coder.eval.behavior.v1` document whose `verdictRef` names the verdict schema, scenario id, and trial index. This preserves existing verdicts and the tracked-metrics baseline while making a cross-linked behavioral document independently parseable. Readers normalize the released `clio.eval.*` identifiers for compatibility, but current writers emit only `clio-coder.eval.*` identifiers.
234
296
 
235
297
  The closed categories are `tool_choice`, `exploration`, `delegation`, `safety_comprehension`, `claim_grounding`, `denied_tool_recovery`, `completion_behavior`, and `task_correctness`. Each category result is exactly one of `satisfied`, `violated`, `unknown`, or `unmeasured`. The document outcome is `pass`, `behavioral_failure`, `unknown`, `unmeasured`, or `infrastructure_failure`; missing facts are never invented as successes, and an infrastructure failure cannot become a behavioral pass.
236
298
 
@@ -240,8 +302,8 @@ Suite execution adapts scalar run metrics into these observable facts at the Sui
240
302
 
241
303
  ### Public built-in behavioral corpus
242
304
 
243
- The repository ships corpus `public-built-in-behavior` version `1.0.0` under
244
- `benchmarks/eval/`. It contains no private prompts, endpoints, credentials, or
305
+ The source repository carries corpus `public-built-in-behavior` version `1.0.0`
306
+ under `evals/`. It contains no private prompts, endpoints, credentials, or
245
307
  mutable external dataset:
246
308
 
247
309
  - `behavioral-machinery.yaml` provides one positive and one adversarial
@@ -262,12 +324,15 @@ mutable external dataset:
262
324
  that the rules can reject observed model behavior rather than merely restate
263
325
  aggregate success counters.
264
326
 
265
- Build once, then run either focused suite from the repository root:
327
+ These are source-checkout workflows: the npm archive keeps the inputs for
328
+ inspection and reproducibility, but the deterministic TypeScript driver uses
329
+ the repository development toolchain. Build once, then run either focused
330
+ suite from the repository root:
266
331
 
267
332
  ```sh
268
- node dist/cli/index.js eval run --suite benchmarks/eval/behavioral-machinery.yaml --clio-coder-entry dist/cli/index.js
269
- node dist/cli/index.js eval run --suite benchmarks/eval/behavioral-model.yaml --target mini --clio-coder-entry dist/cli/index.js
270
- node dist/cli/index.js eval run --suite benchmarks/eval/behavioral-model-negative-control.yaml --target mini --clio-coder-entry dist/cli/index.js
333
+ node dist/cli/index.js eval run --suite evals/behavioral-machinery.yaml --clio-coder-entry dist/cli/index.js
334
+ node dist/cli/index.js eval run --suite evals/behavioral-model.yaml --target mini --clio-coder-entry dist/cli/index.js
335
+ node dist/cli/index.js eval run --suite evals/behavioral-model-negative-control.yaml --target mini --clio-coder-entry dist/cli/index.js
271
336
  ```
272
337
 
273
338
  The machinery tasks use the repository read-only and create only private
@@ -292,7 +357,7 @@ Eleven numbers plus a reason histogram, each carrying the source it came from. `
292
357
  | `generatedTokens` | ledger |
293
358
  | `reasoningTokens` | receipt; nullable, because absent and zero are different claims |
294
359
  | `toolCalls`, `toolErrors` | ledger when present, otherwise receipt |
295
- | `ttftMsFirstCall` | ledger |
360
+ | `ttftMsFirstCall` | ledger; nullable when first-call timing is absent |
296
361
  | `wallClockMs` | receipt |
297
362
  | `contextTokensAtEnd` | ledger |
298
363
  | `compactions` | ledger |
@@ -300,14 +365,18 @@ Eleven numbers plus a reason histogram, each carrying the source it came from. `
300
365
 
301
366
  A dispatched worker's receipt reports `sessionId: null` and writes no session archive, which is why the ledger source exists at all: the runner folds structured usage, backend timing, cache, and monotonic TTFT facts out of the worker's `message_end` events. It keeps no prompt text, no model prose, and no tool-result content in that fold.
302
367
 
368
+ First-call TTFT uses the earliest recorded assistant-call timestamp across the selected ledgers; equal timestamps retain their observed order. Missing or invalid timing, or an invalid timestamp that prevents ordering the calls, yields `{ value: null, source: "estimated" }`. A measured zero remains `{ value: 0, source: "ledger" }`. Native session timing starts at each stream invocation and includes the provider's response-header wait. Stdout-only fallback timing starts at the provider's `message_start`, which can arrive after headers; it requires first output, but is not complete request latency and must not be compared as equivalent to native invocation timing. A completion alone supplies neither a zero duration nor a first-token measurement. Verdict v1 consumers must accept nullable TTFT. Historical numeric values, including estimated zeros and native spans that omitted the pre-header wait, remain readable and are not rewritten.
369
+
370
+ When a stream message lacks a valid timestamp, its ledger payload marks `timestampEstimated: true` beside the legacy ISO placeholder. This leaves first-call chronology unmeasured while preserving any observed per-call monotonic timing.
371
+
303
372
  ### Scenario aggregates
304
373
 
305
374
  `aggregates` groups verdicts by `scenarioId`, sets `k` to the trial count, and records `passAtK` (any trial passed) and `passPowK` (every trial passed). Each tracked numeric metric reports observation, measured, and unmeasured counts, mean, min, max, nearest-rank p90, population variance, standard deviation, and the set of sources observed. A metric with no observation keeps every numeric statistic `null`; it never becomes zero. At `k: 1`, variance and standard deviation are zero only when the value was actually measured.
306
375
 
307
376
  ### Behavioral multi-metric results
308
377
 
309
- A result with a `clio.eval.behavior.v1` verdict also carries the additive
310
- `clio.eval.behavior.metrics.v1` projection. The projection binds the scenario
378
+ A result with a `clio-coder.eval.behavior.v1` verdict also carries the additive
379
+ `clio-coder.eval.behavior.metrics.v1` projection. The projection binds the scenario
311
380
  to its role and target/model envelope and records one `number | null`
312
381
  observation for each closed metric. The source travels beside every value:
313
382
 
@@ -356,8 +425,8 @@ is emitted as testcase output rather than a failed testcase.
356
425
  ### Execution-envelope provenance and comparability
357
426
 
358
427
  Every newly written behavioral result carries an additive
359
- `clio.eval.execution-envelope.v1` sibling. Artifact v4,
360
- `clio.eval.verdict.v1`, and `clio.eval.behavior.metrics.v1` retain their
428
+ `clio-coder.eval.execution-envelope.v1` sibling. Artifact v4,
429
+ `clio-coder.eval.verdict.v1`, and `clio-coder.eval.behavior.metrics.v1` retain their
361
430
  existing identities. The envelope records the selected prompt fragment ids,
362
431
  authored versions or `unversioned` marker, fragment content hashes, prompt
363
432
  composition hash, recipe id/version/fingerprint when a worker recipe applies,
@@ -382,31 +451,15 @@ metric means and variances. When the prompt or recipe identity changes, the
382
451
  generated evidence names each affected corpus scenario and role instead of
383
452
  hiding it behind an aggregate score.
384
453
 
385
- ### Checked behavioral release baseline
454
+ ### Reference behavioral baseline
386
455
 
387
- The checked deterministic baseline is
388
- `benchmarks/eval/behavioral-machinery-baseline.json`. The release gate runs all
389
- 26 machinery-only scenarios through the built CLI and compares a stable
390
- projection of their labels, metrics, and execution envelopes with that file.
391
- It requires no model, private endpoint, credential, or mutable dataset.
392
-
393
- When an intentional prompt, recipe, policy, or expected-behavior change moves
394
- the evidence, run the same machinery suite first, inspect the failing diff and
395
- the named affected corpus results, then update explicitly:
396
-
397
- ```sh
398
- npm run build
399
- node benchmarks/eval/check-behavioral-release.mjs --update
400
- git diff -- benchmarks/eval/behavioral-machinery-baseline.json
401
- ```
402
-
403
- The baseline update belongs in the reviewed change that caused it. Do not use
404
- the update command merely to make a red gate green. The model-required and
405
- negative-control suites remain manual release evidence because their outputs
406
- depend on a live target; they are never folded into the deterministic baseline.
407
- The projection excludes `latency.wallMs` because scheduler timing is not stable
408
- evidence. Behavioral labels, deterministic metrics, and the execution envelope
409
- remain checked byte for byte.
456
+ `evals/behavioral-machinery-baseline.json` is retained as reviewable reference
457
+ evidence for the machinery corpus. It is not a CI or release gate. Run the
458
+ current `evals/behavioral-machinery.yaml` through the built CLI when a prompt,
459
+ recipe, policy, or expected-behavior change needs a fresh measurement, inspect
460
+ the named scenario evidence, and update any retained baseline deliberately in
461
+ the reviewed change. Model-required and negative-control suites remain manual
462
+ measurements tied to their exact target and serving configuration.
410
463
 
411
464
  ### Hard thresholds and informational budgets
412
465
 
@@ -443,10 +496,12 @@ cannot offset a task or safety regression.
443
496
 
444
497
  ```text
445
498
  serving configuration drift; pass --allow-config-drift to compare these runs
446
- baseline serving: target=mini runtime=llamacpp model=... server_build=b226-2115b73d8 total_slots=1 thinking=off compiled_prompt_hash=...
499
+ baseline serving: target=mini runtime=llamacpp model=ornith1.5-35b-moe server_build=... total_slots=4 thinking=off compiled_prompt_hash=...
447
500
  candidate serving: ...
448
501
  ```
449
502
 
503
+ The current reference `mini` endpoint is the llama.cpp router at `192.168.86.141:8080`. It serves `ornith1.5-35b-moe` with four parallel slots and 262144 context tokens per slot. These deployment facts are reference topology, not defaults imposed on another target; retain the artifact's observed serving configuration with every comparison.
504
+
450
505
  `--allow-config-drift` proceeds and labels the comparison `config drift: allowed`. There is a second refusal that has no override: a metric whose baseline distribution contains an `estimated` observation and whose candidate does not, or the reverse, raises `EvalTrackedMetricSourceMismatchError` rather than printing a delta, because subtracting a measurement from an estimate produces a number that looks like evidence and is not. `--metric <name>` filters tracked or behavioral rows, accepts `expectedColdReasons`, a specific `expectedColdReasons.<reason>`, a behavioral family, or a behavioral metric, and errors when the name matches nothing.
451
506
 
452
507
  ---
@@ -1,7 +1,7 @@
1
1
  # Internal Eval Suites
2
2
 
3
- > [!TIP]
4
- > **Interactive Spec Available:** An interactive blueprint is available at [docs/html/evals_internal_blueprint.html](html/evals_internal_blueprint.html) (Version: 0.4.0).
3
+ > **Visual blueprint:** The source checkout includes the complete
4
+ > [Internal Eval Suites visual reference](https://github.com/iowarp/clio-coder/blob/main/docs/html/evals_internal_blueprint.html).
5
5
 
6
6
  Private suites should live outside this repository. Keep datasets, prompts,
7
7
  live fleet coordinates, calibration outputs, and raw run artifacts in a private
@@ -15,13 +15,13 @@ clio-coder eval run --suite <external-path> --clio-coder-entry dist/cli/index.js
15
15
  ```
16
16
 
17
17
  Use `--out <dir>` when the artifact should be written outside the default Clio
18
- data directory. Product eval artifacts and external benchmark campaigns are
19
- separate: public benchmark adapters live under `benchmarks/community/` and do
20
- not use the eval runner.
18
+ data directory. External benchmark campaigns should adapt their cases and
19
+ grader observations into the same eval engine while keeping private datasets,
20
+ credentials, endpoints, and raw artifacts outside this repository.
21
21
 
22
22
  The public behavioral corpus is the deliberate exception to the otherwise
23
23
  private Suite v2 data policy. Its reviewable, synthetic suites live under
24
- `benchmarks/eval/`: a model-free positive/adversarial authority pair for every
24
+ `evals/`: a model-free positive/adversarial authority pair for every
25
25
  built-in worker recipe, four tiny main-agent model scenarios covering all eight
26
26
  behavioral categories with event- and grader-derived facts, and an intentional
27
27
  decoy negative control. The model-free driver uses the shipped recipe catalog,
@@ -83,7 +83,7 @@ thinking level is measuring the server, not the change under test.
83
83
 
84
84
  The verdict envelope keeps its original `behavioral: null` field for compatibility.
85
85
  A suite that declares a versioned behavioral scenario records the result as a
86
- separate `clio.eval.behavior.v1` document on the Artifact v4 result, cross-linked
86
+ separate `clio-coder.eval.behavior.v1` document on the Artifact v4 result, cross-linked
87
87
  to the unchanged verdict identity. Its labels come only from bounded transcript,
88
88
  tool, receipt, or grader facts, never from an ungrounded judge paragraph. A run whose harness broke records
89
89
  `machinery: "infrastructure_failure"`, which the parser refuses to pair with a
@@ -210,7 +210,7 @@ tasks:
210
210
  - node_modules
211
211
  - dist
212
212
  runner:
213
- kind: clio-run
213
+ kind: clio-coder-run
214
214
  prompt: Fix the intentionally broken function so the local verifier passes.
215
215
  verify:
216
216
  commands:
@@ -271,7 +271,7 @@ tasks:
271
271
  - dist
272
272
  - .clio-coder
273
273
  runner:
274
- kind: clio-run
274
+ kind: clio-coder-run
275
275
  prompt: Summarize the repository purpose and make no file changes.
276
276
  verify:
277
277
  forbidPaths:
@@ -305,7 +305,7 @@ tasks:
305
305
  - dist
306
306
  - .clio-coder
307
307
  runner:
308
- kind: clio-run
308
+ kind: clio-coder-run
309
309
  prompt: Fix the failing unit test with the smallest source change.
310
310
  verify:
311
311
  commands:
@@ -1,7 +1,7 @@
1
1
  # Evolution and Change Manifests
2
2
 
3
- > [!TIP]
4
- > **Interactive Spec Available:** An interactive change manifest editor, authority risk assessor, and checklist workspace is located at [docs/html/evolution_blueprint.html](html/evolution_blueprint.html) (Version: 0.4.0).
3
+ > **Visual blueprint:** The source checkout includes the complete
4
+ > [Evolution and Change Manifests visual reference](https://github.com/iowarp/clio-coder/blob/main/docs/html/evolution_blueprint.html).
5
5
 
6
6
  Clio Coder uses change manifests to make harness changes reviewable, falsifiable, and rollback-friendly. CLIO stands for Context Layer for Input/Output, named for the Greek muse of history. A manifest is JSON, generated or checked with `clio-coder evolve manifest`, and should describe what changed, why, what evidence supports it, what could regress, how to validate it, and how to roll it back.
7
7
 
@@ -1,10 +1,13 @@
1
1
  # Fleet Demo Runbook
2
2
 
3
+ > **Visual blueprint:** The source checkout includes the complete
4
+ > [Fleet Demo Runbook visual reference](https://github.com/iowarp/clio-coder/blob/main/docs/html/fleet_demo_blueprint.html).
5
+
3
6
  A repeatable multi-node demonstration: one orchestrator drives a real
4
7
  CMake/C++ fix through a reviewer-gated dispatch across SSH nodes, and every
5
8
  worker's receipt (including the remote ones) verifies afterward. The steps
6
9
  are executable in order; this document doubles as the recording script.
7
- Background and reference: [fleet-dispatch.md](fleet-dispatch.md).
10
+ Background and reference: [fleet-dispatch.md](../guide/fleet-dispatch.md).
8
11
 
9
12
  ## Reference fabric
10
13
 
@@ -139,9 +142,9 @@ clio-coder evidence inspect <evidenceId>
139
142
  run ledger; a tampered or mismatched receipt fails the build with the field
140
143
  that diverged. The receipts of the remote runs verify on the orchestrator host because the
141
144
  ledger and receipts live on the shared filesystem. Current receipts use strict
142
- v16 and authenticate every current receipt and reconstructed-ledger field.
143
- Every other receipt version is rejected rather than reported as partial; the
144
- current binary has no historical receipt reader.
145
+ v20 and authenticate every current receipt and reconstructed-ledger field.
146
+ Lower versions are reported as retired and are never read as evidence or
147
+ migrated; malformed and future shapes fail verification.
145
148
 
146
149
  ## Provenance walkthrough: what a PI can verify from receipts alone
147
150
 
@@ -171,9 +174,10 @@ reconstruct:
171
174
  complete receipt schema and its stable ledger row. `clio-coder evidence build
172
175
  --run <id>` recomputes and cross-checks it; `verifyReceiptIntegrity` in
173
176
  `src/domains/dispatch/receipt-integrity.ts` is the reference
174
- implementation. Current receipts use v16 and every other version fails
175
- verification. Incompatible state must be archived or removed; it is never
176
- read as evidence through a compatibility verifier.
177
+ implementation. Current receipts use v20. Lower versions are reported as
178
+ retired, while malformed and future shapes fail verification. Incompatible
179
+ state may be archived for inspection, but it is never read as evidence through
180
+ a compatibility verifier.
177
181
 
178
182
  The walkthrough for an audience is three commands: `clio-coder evidence build
179
183
  --run <id>` (it verifies), open the receipt JSON (read `node`, `gate`,
@@ -1,13 +1,20 @@
1
1
  # Git Commit Provenance
2
2
 
3
+ > **Visual blueprint:** The source checkout includes the complete
4
+ > [Git Commit Provenance visual reference](https://github.com/iowarp/clio-coder/blob/main/docs/html/git_commit_provenance_blueprint.html).
5
+
3
6
  Clio Coder adds evidence-aware role trailers to commits created through Clio.
4
7
  The feature is enabled by default:
5
8
 
6
9
  ```yaml
7
- attribution:
8
- gitCommits: true
10
+ integrations:
11
+ git:
12
+ commitAttribution: true
9
13
  ```
10
14
 
15
+ The released `attribution.gitCommits` path is a migration alias. Current
16
+ settings files and writers use `integrations.git.commitAttribution`.
17
+
11
18
  Settings -> Advanced exposes the same switch as **Clio commit provenance**, with
12
19
  `enabled` and `disabled` values. A change applies immediately to subsequent
13
20
  commits in the session. When disabled, Clio leaves commit messages entirely
@@ -49,11 +56,11 @@ Co-authored-by: Clio Coder <clio-coder@iowarp.ai>
49
56
  Existing human trailers stay in place. A Clio trailer already present in any
50
57
  letter case is respected rather than repeated, line endings are normalized only
51
58
  while attribution is enabled, and repeated processing is idempotent. When a directly relevant
52
- receipt-v19 digest passes integrity verification, Clio may additionally add the
59
+ receipt-v20 digest passes integrity verification, Clio may additionally add the
53
60
  full digest:
54
61
 
55
62
  ```text
56
- Clio-Evidence: receipt-v19/sha256:<64-character digest>
63
+ Clio-Evidence: receipt-v20/sha256:<64-character digest>
57
64
  ```
58
65
 
59
66
  Clio does not invent, shorten, or add an unrelated digest. The role trailers do