@iowarp/clio-coder 0.4.1 → 0.4.2

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (437) hide show
  1. package/CHANGELOG.md +92 -0
  2. package/CONTRIBUTING.md +59 -36
  3. package/README.md +404 -472
  4. package/SECURITY.md +2 -1
  5. package/dist/{acp-ZILU3AUO.js → acp-TMDQZDIG.js} +7 -7
  6. package/dist/{agents-HYWGBGQR.js → agents-5N5NG3XG.js} +28 -28
  7. package/dist/assets/codewiki.json +1 -1
  8. package/dist/{auth-N3QT7CBO.js → auth-Z5CCBXKQ.js} +8 -9
  9. package/dist/{builtins-UJLMOVOV.js → builtins-K6TNDT24.js} +4 -4
  10. package/dist/{chunk-GVQJ5CCZ.js → chunk-2HFQNRV3.js} +7 -7
  11. package/dist/{chunk-QMXC4JB7.js → chunk-2NHR3NAY.js} +163 -1401
  12. package/dist/chunk-2X4RYJTJ.js +39 -0
  13. package/dist/{chunk-Y45G3AXC.js → chunk-2Z2IKEXI.js} +6 -10
  14. package/dist/{chunk-EIMVLWB3.js → chunk-34BHNEE3.js} +7 -3
  15. package/dist/{chunk-GIZNH63R.js → chunk-35MSIRKH.js} +9 -4
  16. package/dist/chunk-3EBYEESD.js +314 -0
  17. package/dist/{chunk-CTJ4RNAA.js → chunk-3F7VUY77.js} +2 -2
  18. package/dist/{chunk-AP73CFDC.js → chunk-3KIPBMUA.js} +2 -2
  19. package/dist/{chunk-J5LZHVIT.js → chunk-3M6DQK6S.js} +113 -35
  20. package/dist/{chunk-VEGN6WIQ.js → chunk-462T4EGZ.js} +2 -2
  21. package/dist/{chunk-AFKWHWXF.js → chunk-4JDLP6ZS.js} +33 -16
  22. package/dist/{chunk-6FN3E6KX.js → chunk-4O6MANBS.js} +2 -2
  23. package/dist/{chunk-AKB4GYDL.js → chunk-54ODD65L.js} +5 -5
  24. package/dist/{chunk-BBTJOK6Y.js → chunk-5KW52TEP.js} +3 -3
  25. package/dist/{chunk-6CCS4G3W.js → chunk-5PFYMY2V.js} +2 -2
  26. package/dist/chunk-77QIVUZB.js +1334 -0
  27. package/dist/{chunk-7OBGU7UB.js → chunk-7BHIY2MW.js} +7 -13
  28. package/dist/{chunk-3QSOM6PA.js → chunk-AZ4WMN4W.js} +2 -2
  29. package/dist/{chunk-6NJQITNH.js → chunk-B74PXLU7.js} +6 -3
  30. package/dist/{chunk-R23Z6K6I.js → chunk-B7HM5Z7T.js} +15 -15
  31. package/dist/{chunk-R32CLGZ6.js → chunk-BO7Y52RY.js} +81 -20
  32. package/dist/{chunk-UEDMSP56.js → chunk-BYMNWQ7O.js} +123 -148
  33. package/dist/{chunk-ZJLUDYFY.js → chunk-CRFOIAX3.js} +4 -4
  34. package/dist/{chunk-2NM363SV.js → chunk-CYZW7JHJ.js} +7 -7
  35. package/dist/{chunk-6HMJX2VU.js → chunk-DYHAXKHD.js} +38 -10
  36. package/dist/{chunk-THYWACCR.js → chunk-DZAW46HP.js} +3 -3
  37. package/dist/{chunk-FYUN5KZ3.js → chunk-DZEK6CJN.js} +17 -17
  38. package/dist/{chunk-3I5NY75V.js → chunk-E7GT7O5N.js} +5 -5
  39. package/dist/{chunk-VKFQTNDV.js → chunk-F2I26BDK.js} +4 -4
  40. package/dist/{chunk-HLW2MRKE.js → chunk-F4EKGO4N.js} +3 -1
  41. package/dist/{chunk-IXJT6DCX.js → chunk-FVDGR2ZL.js} +3 -3
  42. package/dist/{chunk-TZSKNMZG.js → chunk-GTUD2WMY.js} +2 -1
  43. package/dist/{chunk-7EPLI7VL.js → chunk-HIICAHCJ.js} +2 -2
  44. package/dist/{chunk-E67WX76H.js → chunk-HKMD33FO.js} +29 -80
  45. package/dist/chunk-HLAFFSEK.js +360 -0
  46. package/dist/{chunk-UAPGZHYC.js → chunk-I64IFBLB.js} +9 -2
  47. package/dist/{chunk-XKA2ICR3.js → chunk-I66ZTYNP.js} +440 -175
  48. package/dist/{chunk-7PWAODYW.js → chunk-I7XBWTYH.js} +2 -2
  49. package/dist/{chunk-PVAMAVBB.js → chunk-IDNA72AH.js} +102 -2
  50. package/dist/{chunk-GCSMB2KY.js → chunk-IKOZFYBN.js} +1 -1
  51. package/dist/{chunk-2VG7KLYV.js → chunk-IKSLQ4XV.js} +5460 -3241
  52. package/dist/{chunk-QKIFBZKT.js → chunk-IMXMHHMQ.js} +166 -25
  53. package/dist/{chunk-74YWRRU5.js → chunk-JBCS7CRR.js} +2 -2
  54. package/dist/{chunk-BDPT6GTK.js → chunk-JWJGP5DQ.js} +2 -2
  55. package/dist/{chunk-K6BF4U2H.js → chunk-KKOJXO6R.js} +62 -14
  56. package/dist/chunk-KPXDY6QF.js +47 -0
  57. package/dist/{chunk-ABLSQ6JX.js → chunk-LJID3DYZ.js} +7 -1
  58. package/dist/{chunk-VKRH2TCS.js → chunk-M2DAX4F6.js} +2 -2
  59. package/dist/{chunk-6I5ILFOF.js → chunk-M2WXEHER.js} +2 -2
  60. package/dist/{chunk-YPI3QQCF.js → chunk-MCEPRMZW.js} +2 -4
  61. package/dist/{chunk-N5UK64DP.js → chunk-MCMZMDAC.js} +2 -2
  62. package/dist/{chunk-Y4CAGMM6.js → chunk-MNJGS2IN.js} +5 -6
  63. package/dist/{chunk-TVHHYFHE.js → chunk-NEDJ26B5.js} +2 -2
  64. package/dist/{chunk-U2WB7TZS.js → chunk-NMJXSHBJ.js} +97 -85
  65. package/dist/{chunk-HUAS7ITX.js → chunk-O3YUNJZ2.js} +13 -21
  66. package/dist/{chunk-MA3H6DM5.js → chunk-P75RZCJW.js} +25 -3
  67. package/dist/{chunk-IG7BCQBA.js → chunk-PGF63K6I.js} +2 -2
  68. package/dist/chunk-PJX3WQUQ.js +42 -0
  69. package/dist/{chunk-6DWBAZ5U.js → chunk-Q4XWMHX6.js} +4 -6
  70. package/dist/{chunk-OJTRZGR3.js → chunk-QQLGQY2A.js} +8 -8
  71. package/dist/{chunk-J4HBWF6Y.js → chunk-RLYRBIYQ.js} +115 -20
  72. package/dist/{chunk-NLFAQR7Z.js → chunk-S66XZJOF.js} +3 -23
  73. package/dist/{chunk-C537JADH.js → chunk-SSEYRH53.js} +6 -7
  74. package/dist/chunk-SZAA6XDG.js +30 -0
  75. package/dist/{chunk-MOPSG2X7.js → chunk-TPEQIQIE.js} +6 -6
  76. package/dist/{chunk-JA5QWE4Z.js → chunk-UBRFI4HS.js} +1879 -1650
  77. package/dist/{chunk-BTGG6BG2.js → chunk-UH347SHR.js} +154 -15
  78. package/dist/{chunk-5YHDIDBP.js → chunk-UH632ZYL.js} +2 -2
  79. package/dist/{chunk-BWW4HLO4.js → chunk-UXCU4E3T.js} +8 -6
  80. package/dist/{chunk-6VC4OV3Z.js → chunk-VIA6RFQZ.js} +3 -11
  81. package/dist/{chunk-ZAZB4JMW.js → chunk-VKPAQYEB.js} +27 -8
  82. package/dist/{chunk-UXN6JT4W.js → chunk-W4YEMFBX.js} +2 -2
  83. package/dist/{chunk-TD3PGPQA.js → chunk-W6NIE6OW.js} +2 -2
  84. package/dist/{chunk-TVH4ONAM.js → chunk-X7IARSHT.js} +3 -3
  85. package/dist/{chunk-PJJ6MY27.js → chunk-XE3PCIXH.js} +3 -3
  86. package/dist/{chunk-FEFIFZTL.js → chunk-XGDPUNND.js} +2 -2
  87. package/dist/{chunk-SCYB3HA4.js → chunk-XOXV5GKE.js} +51 -16
  88. package/dist/{chunk-QTFGO774.js → chunk-XQRY4DTA.js} +24 -11
  89. package/dist/{chunk-BJGUKIG4.js → chunk-YJISEZKC.js} +2 -2
  90. package/dist/{chunk-GPPB3JBE.js → chunk-ZGNYYXQ6.js} +2 -2
  91. package/dist/{chunk-SINK3QR6.js → chunk-ZNT2M6TG.js} +7 -7
  92. package/dist/{chunk-7RY5VZPH.js → chunk-ZW4HH5JJ.js} +6 -6
  93. package/dist/cli/index.js +33 -32
  94. package/dist/{clio-IT3G3VQH.js → clio-7VB377CC.js} +7 -7
  95. package/dist/{code-nav-RK6S7F6E.js → code-nav-YVLCYA7V.js} +85 -17
  96. package/dist/{config-3QZRWZJF.js → config-4HVOS65E.js} +88 -43
  97. package/dist/{configure-FL7Y3KJF.js → configure-PIWO7B24.js} +10 -10
  98. package/dist/{context-5HE7ODYK.js → context-IYEHL3WQ.js} +33 -31
  99. package/dist/{context-XNHL75JV.js → context-KQYIWPWT.js} +47 -34
  100. package/dist/{context-KYQFRVDC.js → context-N6ZE3LGJ.js} +11 -11
  101. package/dist/{context-clear-N545L53A.js → context-clear-G4OGZJDS.js} +33 -31
  102. package/dist/{context-working-set-QHKXSV2F.js → context-working-set-BWLF6LJP.js} +7 -7
  103. package/dist/{dispatch-runner-RGIE5PCT.js → dispatch-runner-2QQAITS3.js} +38 -38
  104. package/dist/{docs-5NAF6AU7.js → docs-PD3EXDKU.js} +21 -20
  105. package/dist/{doctor-ZGPEGHIP.js → doctor-LHBD36VU.js} +23 -22
  106. package/dist/{eval-GXLL44RD.js → eval-C45FYRJ6.js} +21 -20
  107. package/dist/{eval-inventory-HBWSWQOK.js → eval-inventory-6DEJPLBF.js} +2 -2
  108. package/dist/{evidence-HWLBRH3Q.js → evidence-6SHONYAF.js} +30 -28
  109. package/dist/{evolve-FTZBMNVW.js → evolve-KRKMV72X.js} +30 -28
  110. package/dist/{extensions-VHRBEID7.js → extensions-KPZ2UHBB.js} +5 -3
  111. package/dist/{fleet-CKZHJWZJ.js → fleet-IVTCKDHT.js} +62 -61
  112. package/dist/{fleet-commands-EXDXBMV6.js → fleet-commands-EDWL3IT7.js} +5 -5
  113. package/dist/{fleet-decisions-OTHB6KRL.js → fleet-decisions-YP3YEFGK.js} +4 -4
  114. package/dist/{fleet-graph-YTEZUCUT.js → fleet-graph-ZFWKHY2M.js} +16 -14
  115. package/dist/{fleet-inspect-SS6YMDCK.js → fleet-inspect-FVUNCBML.js} +31 -29
  116. package/dist/{fleet-preflight-PBY4VYOM.js → fleet-preflight-UN5XED4R.js} +2 -2
  117. package/dist/{fleet-validate-KMEM5L3S.js → fleet-validate-XOWC4HSX.js} +17 -15
  118. package/dist/{fleet-verify-QD5M7E7Q.js → fleet-verify-UN3SODEL.js} +30 -28
  119. package/dist/{fleet-view-WAMJYNDT.js → fleet-view-TWHJKCN6.js} +31 -29
  120. package/dist/{init-5XQRBOFV.js → init-T2QORQ3Y.js} +50 -49
  121. package/dist/{interop-34TVO25M.js → interop-IN5I2A66.js} +5 -5
  122. package/dist/{library-3QY6KF57.js → library-LSCATDLZ.js} +15 -13
  123. package/dist/{memory-L4UTIIIW.js → memory-HYOKAGGJ.js} +31 -29
  124. package/dist/{models-ZVX3QOWE.js → models-2GPMFYCM.js} +22 -21
  125. package/dist/{monitor-CEKVSYTS.js → monitor-E4ASVUJH.js} +34 -32
  126. package/dist/{orchestrator-77BAP6BC.js → orchestrator-DDMPR3PY.js} +984 -583
  127. package/dist/{panes-7STHOAUJ.js → panes-E3RUXOW5.js} +4 -4
  128. package/dist/{panes-SHAUIRXY.js → panes-IXKLOKA2.js} +23 -8
  129. package/dist/{reset-EOLM7GVE.js → reset-OAQP3W4O.js} +4 -4
  130. package/dist/{resources-74GKTLSF.js → resources-OTRSN34L.js} +15 -13
  131. package/dist/{run-HBAUJNNZ.js → run-5DEYH5QK.js} +60 -59
  132. package/dist/{share-G3APVLVP.js → share-IHWTLO3M.js} +19 -15
  133. package/dist/{skills-35HHUKCR.js → skills-IYMXMKW4.js} +17 -15
  134. package/dist/{skills-eval-QN4HSHDC.js → skills-eval-DROHSJAR.js} +36 -36
  135. package/dist/{skills-inventory-J357J34F.js → skills-inventory-D7X4L4ZX.js} +15 -13
  136. package/dist/{slash-commands-JZZCQA32.js → slash-commands-QBM7UZ3B.js} +21 -18
  137. package/dist/{steer-XAVHJM22.js → steer-Z5DO23FJ.js} +2 -2
  138. package/dist/{targets-DSM6CY3M.js → targets-P2FUC4IL.js} +25 -28
  139. package/dist/{terminal-lease-JOPFUVEM.js → terminal-lease-YREJ3JX2.js} +5 -5
  140. package/dist/{tools-MKNWVPBH.js → tools-5B7RO6MV.js} +4 -4
  141. package/dist/{trace-ECQ7TIYZ.js → trace-YMGMUM6A.js} +55 -7
  142. package/dist/{upgrade-H7TOM7YL.js → upgrade-PXK3S2YM.js} +11 -9
  143. package/dist/{usage-X52N3IDJ.js → usage-ME5MPXGX.js} +36 -34
  144. package/dist/{verifiers-EJTVVSMA.js → verifiers-BVZ7IWOO.js} +5 -5
  145. package/dist/{verify-YJL6XET2.js → verify-5K7ZKQFC.js} +4 -4
  146. package/dist/{web-fetch-MPIFL3LL.js → web-fetch-MPARV2K7.js} +2 -2
  147. package/dist/{wiki-generate-4NDZTQ4B.js → wiki-generate-F5W5QTYY.js} +48 -47
  148. package/dist/{with-panes-OBOBFIIR.js → with-panes-BYOJCLAM.js} +51 -255
  149. package/dist/worker/entry.js +45 -30
  150. package/docs/README.md +176 -81
  151. package/docs/{acp.md → architecture/acp.md} +36 -20
  152. package/docs/{alcf-provider.md → architecture/alcf-provider.md} +8 -5
  153. package/docs/{architecture.md → architecture/architecture.md} +43 -22
  154. package/docs/{artifact-placement.md → architecture/artifact-placement.md} +26 -23
  155. package/docs/architecture/artifact-versions.md +90 -0
  156. package/docs/{capacity-and-scheduling.md → architecture/capacity-and-scheduling.md} +26 -13
  157. package/docs/{context-engine.md → architecture/context-engine.md} +25 -25
  158. package/docs/{context-working-set.md → architecture/context-working-set.md} +13 -10
  159. package/docs/{dispatch-architecture-rationale.md → architecture/dispatch-architecture-rationale.md} +12 -9
  160. package/docs/{dispatch-typed-intent.md → architecture/dispatch-typed-intent.md} +68 -46
  161. package/docs/{evidence-and-memory.md → architecture/evidence-and-memory.md} +23 -16
  162. package/docs/{middleware-and-components.md → architecture/middleware-and-components.md} +11 -5
  163. package/docs/{model-catalog.md → architecture/model-catalog.md} +40 -17
  164. package/docs/{observability.md → architecture/observability.md} +26 -13
  165. package/docs/{pi-boundary.md → architecture/pi-boundary.md} +24 -11
  166. package/docs/{prompt-envelope-and-tools.md → architecture/prompt-envelope-and-tools.md} +55 -20
  167. package/docs/{provider-adapter-cookbook.md → architecture/provider-adapter-cookbook.md} +35 -24
  168. package/docs/{safety-model.md → architecture/safety-model.md} +20 -15
  169. package/docs/{session-lifecycle.md → architecture/session-lifecycle.md} +8 -5
  170. package/docs/architecture/time-conventions.md +125 -0
  171. package/docs/{trace-store.md → architecture/trace-store.md} +13 -5
  172. package/docs/{tui-design.md → architecture/tui-design.md} +13 -13
  173. package/docs/{worker-dispatch-mechanics.md → architecture/worker-dispatch-mechanics.md} +27 -30
  174. package/docs/{built-in-agents.md → guide/built-in-agents.md} +50 -34
  175. package/docs/{commands-and-modes.md → guide/commands-and-modes.md} +65 -60
  176. package/docs/{configuration-and-targets.md → guide/configuration-and-targets.md} +227 -289
  177. package/docs/guide/configuration-reference.md +1158 -0
  178. package/docs/{environment-variables.md → guide/environment-variables.md} +31 -28
  179. package/docs/{exit-codes-and-output.md → guide/exit-codes-and-output.md} +6 -3
  180. package/docs/{extensions-and-sharing.md → guide/extensions-and-sharing.md} +41 -14
  181. package/docs/{fleet-dispatch.md → guide/fleet-dispatch.md} +39 -43
  182. package/docs/{glossary.md → guide/glossary.md} +14 -11
  183. package/docs/{installation-and-lifecycle.md → guide/installation-and-lifecycle.md} +44 -13
  184. package/docs/guide/panes-and-files.md +290 -0
  185. package/docs/{proactive-memory.md → guide/proactive-memory.md} +79 -66
  186. package/docs/{resource-library.md → guide/resource-library.md} +13 -4
  187. package/docs/{skills-marketplace.md → guide/skills-marketplace.md} +7 -3
  188. package/docs/{tool-usage.md → guide/tool-usage.md} +87 -23
  189. package/docs/{troubleshooting.md → guide/troubleshooting.md} +9 -4
  190. package/docs/{config-knobs-audit.md → history/config-knobs-audit.md} +11 -11
  191. package/docs/{release-cut-checklist.md → history/release-cut-checklist.md} +29 -2
  192. package/docs/{development-pipeline.md → process/development-pipeline.md} +24 -26
  193. package/docs/process/documentation-coverage.md +100 -0
  194. package/docs/process/documentation-guide.md +187 -0
  195. package/docs/{eval-runner.md → process/eval-runner.md} +41 -50
  196. package/docs/{evals-internal.md → process/evals-internal.md} +10 -10
  197. package/docs/{evolution.md → process/evolution.md} +2 -2
  198. package/docs/{fleet-demo-runbook.md → process/fleet-demo-runbook.md} +11 -7
  199. package/docs/{git-commit-provenance.md → process/git-commit-provenance.md} +11 -4
  200. package/docs/{performance-methodology.md → process/performance-methodology.md} +87 -69
  201. package/docs/{scientific-validation.md → process/scientific-validation.md} +4 -4
  202. package/evals/README.md +2 -2
  203. package/package.json +9 -7
  204. package/skills/README.md +46 -37
  205. package/skills/coding/ast-grep/SKILL.md +2 -2
  206. package/skills/coding/coding-standards/SKILL.md +2 -2
  207. package/skills/coding/prototype/SKILL.md +2 -2
  208. package/skills/coding/tdd/SKILL.md +2 -2
  209. package/skills/context/context-handoff/SKILL.md +2 -2
  210. package/skills/context/context-prime/SKILL.md +2 -2
  211. package/skills/git/file-ticket/SKILL.md +2 -2
  212. package/skills/git/fix-issue/SKILL.md +3 -3
  213. package/skills/git/resolve-merge-conflicts/SKILL.md +2 -2
  214. package/skills/git/ship/SKILL.md +2 -2
  215. package/skills/git/worktree-create/SKILL.md +2 -2
  216. package/skills/git/worktree-merge/SKILL.md +2 -2
  217. package/skills/meta/clio-coder-dev/SKILL.md +9 -5
  218. package/skills/meta/clio-coder-dev/evals.md +3 -2
  219. package/skills/meta/clio-coder-test/SKILL.md +102 -95
  220. package/skills/meta/clio-coder-test/evals.md +9 -4
  221. package/skills/meta/clio-coder-test/references/harness.md +100 -124
  222. package/skills/meta/clio-coder-test/references/test-map.md +77 -50
  223. package/skills/meta/credentials/SKILL.md +2 -2
  224. package/skills/meta/find-skills/SKILL.md +2 -2
  225. package/skills/meta/herdr/SKILL.md +2 -2
  226. package/skills/meta/skill-craft/SKILL.md +22 -16
  227. package/skills/planning/architecture/SKILL.md +2 -2
  228. package/skills/planning/backlog/SKILL.md +2 -2
  229. package/skills/planning/prd/SKILL.md +2 -2
  230. package/skills/planning/product-intent/SKILL.md +2 -2
  231. package/skills/planning/tech-spec/SKILL.md +2 -2
  232. package/skills/registry.yaml +62 -62
  233. package/skills/research/arxiv-literature/SKILL.md +2 -2
  234. package/skills/research/experiment-protocol/SKILL.md +2 -2
  235. package/skills/research/scientific-debugging/SKILL.md +2 -2
  236. package/skills/research/scientific-modernization/SKILL.md +2 -2
  237. package/skills/skill-marketplace.json +62 -62
  238. package/skills/workflow/cut-it/SKILL.md +2 -2
  239. package/skills/workflow/design-council/SKILL.md +2 -2
  240. package/skills/workflow/grill-me/SKILL.md +2 -2
  241. package/skills/workflow/workflow-distiller/SKILL.md +2 -2
  242. package/src/cli/args.ts +2 -2
  243. package/src/cli/bootstrap-generate.ts +1 -1
  244. package/src/cli/config-inspect.ts +65 -12
  245. package/src/cli/configure.ts +0 -4
  246. package/src/cli/docs.ts +22 -14
  247. package/src/cli/doctor-naming.ts +5 -5
  248. package/src/cli/doctor-toolchain.ts +3 -3
  249. package/src/cli/eval.ts +1 -2
  250. package/src/cli/extensions.ts +2 -1
  251. package/src/cli/fleet.ts +1 -1
  252. package/src/cli/index.ts +2 -1
  253. package/src/cli/internal-dispatch.ts +3 -4
  254. package/src/cli/panes.ts +19 -5
  255. package/src/cli/run.ts +2 -2
  256. package/src/cli/share.ts +5 -1
  257. package/src/cli/skills-eval.ts +3 -3
  258. package/src/cli/targets.ts +2 -6
  259. package/src/cli/trace.ts +55 -4
  260. package/src/cli/wiki-generate.ts +1 -1
  261. package/src/core/artifact-paths.ts +1 -1
  262. package/src/core/bash-exec.ts +131 -86
  263. package/src/core/bus-events.ts +51 -6
  264. package/src/core/config.ts +5 -1
  265. package/src/core/defaults.ts +7 -4
  266. package/src/core/dispatch-outcome.ts +16 -0
  267. package/src/core/guardrails.ts +10 -49
  268. package/src/core/prompt-hint.ts +9 -0
  269. package/src/domains/agents/builtins/architect.md +2 -3
  270. package/src/domains/agents/builtins/coder.md +3 -2
  271. package/src/domains/agents/builtins/debugger.md +2 -2
  272. package/src/domains/agents/builtins/documenter.md +2 -2
  273. package/src/domains/agents/builtins/git-master.md +1 -1
  274. package/src/domains/agents/builtins/oracle.md +1 -1
  275. package/src/domains/agents/builtins/provenance.md +1 -1
  276. package/src/domains/agents/builtins/researcher.md +1 -1
  277. package/src/domains/agents/builtins/scout.md +1 -1
  278. package/src/domains/agents/builtins/tester.md +2 -2
  279. package/src/domains/agents/builtins/verifier.md +2 -2
  280. package/src/domains/agents/builtins/wiki-writer.md +1 -1
  281. package/src/domains/agents/catalog.ts +12 -14
  282. package/src/domains/agents/contract.ts +2 -0
  283. package/src/domains/agents/extension.ts +23 -1
  284. package/src/domains/config/keybindings.ts +8 -0
  285. package/src/domains/context/extension.ts +0 -3
  286. package/src/domains/context/working-set/path-index.ts +1 -0
  287. package/src/domains/dispatch/capability-match.ts +10 -0
  288. package/src/domains/dispatch/extension.ts +105 -22
  289. package/src/domains/dispatch/host-verification.ts +435 -39
  290. package/src/domains/dispatch/intent-requirements.ts +10 -0
  291. package/src/domains/dispatch/intent.ts +18 -1
  292. package/src/domains/dispatch/path-scope.ts +235 -24
  293. package/src/domains/dispatch/run-event-journal.ts +4 -15
  294. package/src/domains/dispatch/state.ts +2 -3
  295. package/src/domains/dispatch/transport.ts +45 -21
  296. package/src/domains/dispatch/types.ts +55 -3
  297. package/src/domains/eval/artifacts/store.ts +5 -0
  298. package/src/domains/eval/store.ts +8 -1
  299. package/src/domains/evidence/trust-status.ts +10 -1
  300. package/src/domains/extensions/contract.ts +15 -1
  301. package/src/domains/extensions/discovery.ts +238 -41
  302. package/src/domains/extensions/extension.ts +105 -6
  303. package/src/domains/extensions/index.ts +24 -0
  304. package/src/domains/extensions/integrity.ts +189 -0
  305. package/src/domains/extensions/manager.ts +17 -1
  306. package/src/domains/extensions/resource-path.ts +27 -0
  307. package/src/domains/extensions/resources.ts +18 -38
  308. package/src/domains/extensions/snapshot-store.ts +39 -0
  309. package/src/domains/extensions/snapshot.ts +180 -0
  310. package/src/domains/extensions/state.ts +385 -57
  311. package/src/domains/extensions/types.ts +118 -1
  312. package/src/domains/lifecycle/migrations/2026-09-01-extension-install-digests.ts +27 -0
  313. package/src/domains/lifecycle/migrations/index.ts +2 -0
  314. package/src/domains/lifecycle/naming-resources.ts +19 -4
  315. package/src/domains/lifecycle/naming-yazi.ts +10 -5
  316. package/src/domains/middleware/contract.ts +26 -0
  317. package/src/domains/middleware/extension.ts +24 -24
  318. package/src/domains/middleware/hook-receipts.ts +27 -4
  319. package/src/domains/middleware/hooks-io.ts +65 -32
  320. package/src/domains/middleware/hooks.ts +64 -0
  321. package/src/domains/middleware/index.ts +28 -4
  322. package/src/domains/middleware/registrations.ts +326 -0
  323. package/src/domains/middleware/runtime.ts +28 -0
  324. package/src/domains/middleware/snapshot.ts +20 -7
  325. package/src/domains/mux/contract.ts +38 -0
  326. package/src/domains/mux/detect.ts +6 -13
  327. package/src/domains/mux/index.ts +1 -1
  328. package/src/domains/mux/operations.ts +44 -5
  329. package/src/domains/mux/yazi/assets/yazi.toml +2 -2
  330. package/src/domains/mux/yazi/session.ts +53 -4
  331. package/src/domains/mux/yazi/theme.ts +117 -17
  332. package/src/domains/observability/contract.ts +10 -11
  333. package/src/domains/observability/extension.ts +11 -3
  334. package/src/domains/observability/projection.ts +14 -90
  335. package/src/domains/observability/trace-store.ts +43 -7
  336. package/src/domains/prompts/compiler.ts +73 -53
  337. package/src/domains/prompts/contract.ts +15 -3
  338. package/src/domains/prompts/extension.ts +97 -9
  339. package/src/domains/prompts/fragments/identity/clio-worker.md +1 -3
  340. package/src/domains/prompts/fragments/identity/clio.md +6 -12
  341. package/src/domains/prompts/fragments/identity/docs-routing.md +1 -2
  342. package/src/domains/prompts/fragments/identity/self-awareness.md +3 -11
  343. package/src/domains/prompts/fragments/operating/contract.md +7 -15
  344. package/src/domains/prompts/fragments/operating/delegation.md +32 -34
  345. package/src/domains/prompts/fragments/operating/skills.md +10 -24
  346. package/src/domains/prompts/fragments/operating/worker.md +1 -8
  347. package/src/domains/providers/index.ts +1 -1
  348. package/src/domains/providers/model-runtime-capabilities.ts +85 -21
  349. package/src/domains/providers/models/local-models/clio-coder-local-coding-targets.yaml +669 -104
  350. package/src/domains/providers/runtime-resolution.ts +31 -0
  351. package/src/domains/providers/runtimes/common/probe-helpers.ts +7 -2
  352. package/src/domains/providers/runtimes/local-native/llamacpp.ts +9 -1
  353. package/src/domains/providers/types/cost-provenance.ts +19 -0
  354. package/src/domains/providers/types/local-model-quirks.ts +85 -37
  355. package/src/domains/resources/skills/loader.ts +16 -19
  356. package/src/domains/safety/call-target.ts +1 -1
  357. package/src/domains/safety/loop-detector.ts +7 -4
  358. package/src/domains/session/task-board.ts +10 -9
  359. package/src/domains/share/archive.ts +164 -7
  360. package/src/engine/acp/server.ts +62 -9
  361. package/src/engine/apis/llamacpp-residency.ts +3 -4
  362. package/src/engine/apis/lmstudio.ts +3 -3
  363. package/src/engine/apis/ollama-native.ts +6 -6
  364. package/src/engine/apis/openai-completions.ts +28 -25
  365. package/src/engine/apis/output-budget.ts +8 -18
  366. package/src/engine/apis/residency.ts +8 -27
  367. package/src/engine/gemma-channel-filter.ts +19 -0
  368. package/src/engine/loop-guard.ts +92 -12
  369. package/src/engine/worker-runtime.ts +40 -11
  370. package/src/engine/worker-tools.ts +3 -1
  371. package/src/entry/extension-hook-sources.ts +28 -0
  372. package/src/entry/extension-reload.ts +309 -0
  373. package/src/entry/orchestrator.ts +59 -35
  374. package/src/interactive/application-controller.ts +2 -1
  375. package/src/interactive/bus-notices.ts +8 -1
  376. package/src/interactive/chat-loop-messages.ts +3 -13
  377. package/src/interactive/chat-loop.ts +10 -1
  378. package/src/interactive/chat-panel.ts +36 -13
  379. package/src/interactive/chat-renderer.ts +71 -7
  380. package/src/interactive/dispatch-board.ts +6 -11
  381. package/src/interactive/footer/widgets.ts +13 -0
  382. package/src/interactive/interactive-application.ts +39 -4
  383. package/src/interactive/interactive-input-runtime.ts +4 -0
  384. package/src/interactive/interactive-presentation.ts +2 -2
  385. package/src/interactive/interactive-slash-runtime.ts +2 -0
  386. package/src/interactive/overlays/extensions.ts +9 -1
  387. package/src/interactive/overlays/help-reference.ts +13 -0
  388. package/src/interactive/overlays/settings.ts +27 -16
  389. package/src/interactive/panes-runtime.ts +111 -35
  390. package/src/interactive/prompt-cache-identity.ts +88 -0
  391. package/src/interactive/slash-commands.ts +129 -14
  392. package/src/interactive/stream-pacing-policy.ts +0 -23
  393. package/src/interactive/turn-context.ts +30 -15
  394. package/src/interactive/yazi-bridge.ts +60 -6
  395. package/src/tools/agent-tools.ts +30 -1
  396. package/src/tools/artifact.ts +2 -2
  397. package/src/tools/ask-user.ts +3 -3
  398. package/src/tools/bash.ts +1 -1
  399. package/src/tools/bootstrap.ts +4 -0
  400. package/src/tools/builtin-tool-catalog.ts +52 -22
  401. package/src/tools/codewiki/code-nav-surface.ts +6 -0
  402. package/src/tools/codewiki/code-nav.ts +99 -13
  403. package/src/tools/context/docs-engine.ts +20 -7
  404. package/src/tools/context/index.ts +29 -12
  405. package/src/tools/core-bootstrap.ts +28 -6
  406. package/src/tools/credential-present.ts +1 -2
  407. package/src/tools/dispatch-arguments.ts +5 -1
  408. package/src/tools/dispatch-plan.ts +48 -4
  409. package/src/tools/dispatch-run-events.ts +1 -1
  410. package/src/tools/dispatch-schema.ts +338 -0
  411. package/src/tools/dispatch-types.ts +3 -0
  412. package/src/tools/dispatch.ts +9 -254
  413. package/src/tools/ledger.ts +3 -5
  414. package/src/tools/monitor-surface.ts +5 -13
  415. package/src/tools/observation.ts +4 -5
  416. package/src/tools/panes-surface.ts +4 -11
  417. package/src/tools/panes.ts +4 -2
  418. package/src/tools/policy.ts +15 -2
  419. package/src/tools/read.ts +5 -6
  420. package/src/tools/registry.ts +30 -7
  421. package/src/tools/result-shaping.ts +18 -14
  422. package/src/tools/steer-surface.ts +1 -1
  423. package/src/tools/tasks.ts +1 -1
  424. package/src/tools/truncate.ts +6 -5
  425. package/src/tools/verify/surface.ts +6 -12
  426. package/src/tools/web-fetch-surface.ts +1 -3
  427. package/dist/chunk-5QIAJV2D.js +0 -48
  428. package/dist/chunk-JZWT5J3Y.js +0 -814
  429. package/dist/chunk-K7VKOLQQ.js +0 -15
  430. package/dist/chunk-PMZCIOCJ.js +0 -25
  431. package/dist/chunk-SUW5DORT.js +0 -819
  432. package/dist/chunk-UOV2BYIW.js +0 -107
  433. package/dist/chunk-WR6U3OVP.js +0 -45
  434. package/docs/artifact-versions.md +0 -67
  435. package/docs/documentation-coverage.md +0 -46
  436. package/docs/documentation-guide.md +0 -167
  437. package/docs/time-conventions.md +0 -101
@@ -1,11 +1,11 @@
1
1
  # Clio Coder Local Evaluation Runner
2
2
 
3
- > [!TIP]
4
- > **Interactive Spec Available:** An interactive task suite validator, subprocess execution simulator, and compare calculator is located at [docs/html/eval_blueprint.html](html/eval_blueprint.html) (Version: 0.4.0).
3
+ > **Visual blueprint:** The source checkout includes the complete
4
+ > [Clio Coder Local Evaluation Runner visual reference](https://github.com/iowarp/clio-coder/blob/main/docs/html/eval_blueprint.html).
5
5
 
6
6
  The local evaluation runner executes repository-local YAML task suites as deterministic subprocess checks. It is useful for comparing harness changes, prompts, tools, or local workflows.
7
7
 
8
- Source of truth: [src/domains/eval/](../src/domains/eval/) and [src/cli/eval.ts](../src/cli/eval.ts).
8
+ Source of truth: [src/domains/eval/](../../src/domains/eval/) and [src/cli/eval.ts](../../src/cli/eval.ts).
9
9
 
10
10
  ---
11
11
 
@@ -20,6 +20,7 @@ clio-coder eval run --task-file <tasks.yaml> [--repeat <n>] [--out <path>] [--cl
20
20
  clio-coder eval report <evalId> --format text|json|md|swe-jsonl|junit
21
21
  clio-coder eval compare <baselineEvalId> <candidateEvalId> [--metric <name>] [--format text|json|md|junit] [--allow-config-drift]
22
22
  clio-coder eval gate <candidateEvalId> --baseline <baselineEvalId> [--thresholds <file>]
23
+ clio-coder eval inventory --json
23
24
  ```
24
25
 
25
26
  ### Command Roles
@@ -33,6 +34,7 @@ clio-coder eval gate <candidateEvalId> --baseline <baselineEvalId> [--thresholds
33
34
  * `junit`: XML report for CI/CD integration.
34
35
  * **`compare`**: Compares two evaluation artifacts (baseline and candidate) by matching tasks.
35
36
  * **`gate`**: Compares candidate metrics against baseline and absolute thresholds. Correctness and safety regressions fail independently of informational budgets.
37
+ * **`inventory`**: Prints the fixed machine-readable inventory used by GUI hosts. It includes stored report identity, provenance, serving facts, accounting, and per-scenario outcomes without report attachments.
36
38
 
37
39
  Exit codes:
38
40
 
@@ -42,7 +44,7 @@ Exit codes:
42
44
  | `eval run` | `0` when all task repetitions pass | `1` when any task fails, `2` for invalid configs |
43
45
  | `eval report` | `0` when artifact loads | `1` if artifact cannot be read, `2` for invalid ID |
44
46
  | `eval compare` | `0` when both artifacts compare and the behavioral hard gate passes | `1` for a hard regression or unreadable artifact, `2` for invalid ID |
45
- | `eval gate` | `0` when correctness, safety, and hard threshold assertions pass | `1` for any hard failure, `2` for config/invalid ID errors |
47
+ | `eval gate` | `0` when correctness, safety, and hard threshold assertions pass | `1` for hard failures, unreadable inputs, and malformed threshold files; `2` for an invalid eval ID or usage error |
46
48
 
47
49
  ---
48
50
 
@@ -75,7 +77,7 @@ tasks:
75
77
  excludes:
76
78
  - "**/node_modules/**"
77
79
  runner:
78
- kind: "clio-run" # clio-run | context-index | context-init | external-command
80
+ kind: "clio-coder-run" # clio-coder-run | context-index | context-init | external-command
79
81
  prompt: "Optimize the FFT tolerance bounds in solver.ts"
80
82
  timeoutMs: 60000
81
83
  verify:
@@ -103,12 +105,12 @@ tasks:
103
105
  | --- | --- | --- |
104
106
  | `version` | - | Must equal `2`. |
105
107
  | `suite` | `id`, `title`, `visibility`, `description` | Metadata identifying the evaluation suite. |
106
- | `matrix` | `targets[]`, `repeats`, `dimensions[]` | Matrix of execution targets, repetition count, and the execution-envelope fields intentionally varied by the suite. |
107
- | `workspace` | `kind`, `path`, `url`, `commit`, `checkout`, `excludes` | Workspace strategy: `local` (run in-place), `git` (clone from URL), or `temp-copy` (isolated copy of a directory). |
108
- | `runner` | `kind`, `prompt`, `command`, `commands`, `args`, `timeoutMs` | Runner type: `clio-run` (starts Clio agent loop), `context-index` (runs indexer), `context-init` (initializes context), `external-command` (spawns subprocess). |
109
- | `behavioral` | `schema`, `corpus`, `execution`, `expectedBehavior`, `forbiddenBehavior`, `judge` | Optional `clio.eval.scenario.v1` behavioral contract. Rules name a closed category and a typed predicate over transcript, tool, receipt, or grader facts. |
108
+ | `matrix` | `targets[]`, `repeats`, `dimensions[]`, `maxCostUsd` | Matrix of execution targets, repetition count, execution-envelope fields intentionally varied by the suite, and an optional cumulative known-cost ceiling. |
109
+ | `workspace` | `kind`, `path`, `url`, `commit`, `checkout`, `excludes`, `setup` | Workspace strategy: `local` (run in-place), `git` (clone from URL), or `temp-copy` (isolated copy of a directory). Optional `setup` commands prepare the workspace before the runner starts. |
110
+ | `runner` | `kind`, `prompt`, `autonomy`, `agent`, `command`, `commands`, `args`, `timeoutMs` | Runner type: `clio-coder-run` (starts Clio's agent loop), `context-index` (runs the indexer), `context-init` (initializes context), or `external-command` (spawns a subprocess). `agent` selects a worker recipe and `autonomy` sets one-run headless authority. |
111
+ | `behavioral` | `schema`, `corpus`, `execution`, `expectedBehavior`, `forbiddenBehavior`, `judge` | Optional `clio-coder.eval.scenario.v1` behavioral contract. Rules name a closed category and a typed predicate over transcript, tool, receipt, or grader facts. |
110
112
  | `verify` | `commands`, `measure`, `assertions`, `forbidPaths` | Validation steps: shell commands, a task-outcome grader, metric assertions (e.g. `op: lt` for max token counts), and files/directories that must not be created or modified (`forbidPaths`). |
111
- | `metrics` | `collect` | List of metric names to compile for the evaluation runs. |
113
+ | `metrics` | `collect`, `readObservation` | Metric names to compile plus optional public allowlisted and decoy paths reduced to bounded read counters. Raw path strings do not enter behavioral facts. |
112
114
 
113
115
  ---
114
116
 
@@ -120,7 +122,7 @@ tasks:
120
122
  ---
121
123
 
122
124
  ## Runner Kinds
123
- * **`clio-run`**: Invokes the main Clio Coder agent loop with the task's prompt, tracing all tools.
125
+ * **`clio-coder-run`**: Invokes the main Clio Coder agent loop with the task's prompt, tracing all tools. The released `clio-run` spelling is accepted only as a legacy input alias and is normalized before validation; writers and new suites use `clio-coder-run`.
124
126
  * **`context-index`**: Triggers the context engine to build index structures (`codewiki`).
125
127
  * **`context-init`**: Initializes workspace files (such as generating `CLIO-CODER.md`).
126
128
  * **`external-command`**: Spawns an external command or sequence of commands in the task workspace.
@@ -190,7 +192,7 @@ export interface EvalArtifactV4 {
190
192
  version: 4;
191
193
  evalId: string;
192
194
  suite: { id: string; hash: string };
193
- clio: EvalClioProvenance;
195
+ clioCoder: EvalClioProvenance;
194
196
  environment: EvalEnvironmentProvenance;
195
197
  matrix: { target: string; model: string | null; thinking: string | null };
196
198
  summary: EvalArtifactSummaryV4;
@@ -206,11 +208,11 @@ export interface EvalArtifactV4 {
206
208
 
207
209
  ## The verdict envelope
208
210
 
209
- Every result carries a strictly parsed `clio.eval.verdict.v1` envelope (`src/domains/eval/schema/verdict.ts`). Suite v2 results are adapted into it at one explicit boundary (`src/domains/eval/schema/adapter.ts`) rather than by widening the artifact version, because the envelope carries no information a v4 artifact cannot hold.
211
+ Every result carries a strictly parsed `clio-coder.eval.verdict.v1` envelope (`src/domains/eval/schema/verdict.ts`). Suite v2 results are adapted into it at one explicit boundary (`src/domains/eval/schema/adapter.ts`) rather than by widening the artifact version, because the envelope carries no information a v4 artifact cannot hold.
210
212
 
211
213
  ```json
212
214
  {
213
- "schema": "clio.eval.verdict.v1",
215
+ "schema": "clio-coder.eval.verdict.v1",
214
216
  "scenarioId": "latency-nonnegative",
215
217
  "trialIndex": 0,
216
218
  "outcome": "pass",
@@ -230,7 +232,7 @@ The envelope is fail-closed by construction. `outcome` is one of `pass`, `fail`,
230
232
 
231
233
  ### Behavioral scenario and verdict documents
232
234
 
233
- Behavioral evaluation is additive and does not change the persisted `clio.eval.verdict.v1` reader. A Suite v2 task may declare a `clio.eval.scenario.v1` block, and its Artifact v4 result then carries a sibling `clio.eval.behavior.v1` document whose `verdictRef` names the verdict schema, scenario id, and trial index. This preserves existing verdicts and the tracked-metrics baseline while making a cross-linked behavioral document independently parseable.
235
+ Behavioral evaluation is additive and does not change the persisted `clio-coder.eval.verdict.v1` reader. A Suite v2 task may declare a `clio-coder.eval.scenario.v1` block, and its Artifact v4 result then carries a sibling `clio-coder.eval.behavior.v1` document whose `verdictRef` names the verdict schema, scenario id, and trial index. This preserves existing verdicts and the tracked-metrics baseline while making a cross-linked behavioral document independently parseable. Readers normalize the released `clio.eval.*` identifiers for compatibility, but current writers emit only `clio-coder.eval.*` identifiers.
234
236
 
235
237
  The closed categories are `tool_choice`, `exploration`, `delegation`, `safety_comprehension`, `claim_grounding`, `denied_tool_recovery`, `completion_behavior`, and `task_correctness`. Each category result is exactly one of `satisfied`, `violated`, `unknown`, or `unmeasured`. The document outcome is `pass`, `behavioral_failure`, `unknown`, `unmeasured`, or `infrastructure_failure`; missing facts are never invented as successes, and an infrastructure failure cannot become a behavioral pass.
236
238
 
@@ -240,8 +242,8 @@ Suite execution adapts scalar run metrics into these observable facts at the Sui
240
242
 
241
243
  ### Public built-in behavioral corpus
242
244
 
243
- The repository ships corpus `public-built-in-behavior` version `1.0.0` under
244
- `benchmarks/eval/`. It contains no private prompts, endpoints, credentials, or
245
+ The source repository carries corpus `public-built-in-behavior` version `1.0.0`
246
+ under `evals/`. It contains no private prompts, endpoints, credentials, or
245
247
  mutable external dataset:
246
248
 
247
249
  - `behavioral-machinery.yaml` provides one positive and one adversarial
@@ -262,12 +264,15 @@ mutable external dataset:
262
264
  that the rules can reject observed model behavior rather than merely restate
263
265
  aggregate success counters.
264
266
 
265
- Build once, then run either focused suite from the repository root:
267
+ These are source-checkout workflows: the npm archive keeps the inputs for
268
+ inspection and reproducibility, but the deterministic TypeScript driver uses
269
+ the repository development toolchain. Build once, then run either focused
270
+ suite from the repository root:
266
271
 
267
272
  ```sh
268
- node dist/cli/index.js eval run --suite benchmarks/eval/behavioral-machinery.yaml --clio-coder-entry dist/cli/index.js
269
- node dist/cli/index.js eval run --suite benchmarks/eval/behavioral-model.yaml --target mini --clio-coder-entry dist/cli/index.js
270
- node dist/cli/index.js eval run --suite benchmarks/eval/behavioral-model-negative-control.yaml --target mini --clio-coder-entry dist/cli/index.js
273
+ node dist/cli/index.js eval run --suite evals/behavioral-machinery.yaml --clio-coder-entry dist/cli/index.js
274
+ node dist/cli/index.js eval run --suite evals/behavioral-model.yaml --target mini --clio-coder-entry dist/cli/index.js
275
+ node dist/cli/index.js eval run --suite evals/behavioral-model-negative-control.yaml --target mini --clio-coder-entry dist/cli/index.js
271
276
  ```
272
277
 
273
278
  The machinery tasks use the repository read-only and create only private
@@ -306,8 +311,8 @@ A dispatched worker's receipt reports `sessionId: null` and writes no session ar
306
311
 
307
312
  ### Behavioral multi-metric results
308
313
 
309
- A result with a `clio.eval.behavior.v1` verdict also carries the additive
310
- `clio.eval.behavior.metrics.v1` projection. The projection binds the scenario
314
+ A result with a `clio-coder.eval.behavior.v1` verdict also carries the additive
315
+ `clio-coder.eval.behavior.metrics.v1` projection. The projection binds the scenario
311
316
  to its role and target/model envelope and records one `number | null`
312
317
  observation for each closed metric. The source travels beside every value:
313
318
 
@@ -356,8 +361,8 @@ is emitted as testcase output rather than a failed testcase.
356
361
  ### Execution-envelope provenance and comparability
357
362
 
358
363
  Every newly written behavioral result carries an additive
359
- `clio.eval.execution-envelope.v1` sibling. Artifact v4,
360
- `clio.eval.verdict.v1`, and `clio.eval.behavior.metrics.v1` retain their
364
+ `clio-coder.eval.execution-envelope.v1` sibling. Artifact v4,
365
+ `clio-coder.eval.verdict.v1`, and `clio-coder.eval.behavior.metrics.v1` retain their
361
366
  existing identities. The envelope records the selected prompt fragment ids,
362
367
  authored versions or `unversioned` marker, fragment content hashes, prompt
363
368
  composition hash, recipe id/version/fingerprint when a worker recipe applies,
@@ -382,31 +387,15 @@ metric means and variances. When the prompt or recipe identity changes, the
382
387
  generated evidence names each affected corpus scenario and role instead of
383
388
  hiding it behind an aggregate score.
384
389
 
385
- ### Checked behavioral release baseline
390
+ ### Reference behavioral baseline
386
391
 
387
- The checked deterministic baseline is
388
- `benchmarks/eval/behavioral-machinery-baseline.json`. The release gate runs all
389
- 26 machinery-only scenarios through the built CLI and compares a stable
390
- projection of their labels, metrics, and execution envelopes with that file.
391
- It requires no model, private endpoint, credential, or mutable dataset.
392
-
393
- When an intentional prompt, recipe, policy, or expected-behavior change moves
394
- the evidence, run the same machinery suite first, inspect the failing diff and
395
- the named affected corpus results, then update explicitly:
396
-
397
- ```sh
398
- npm run build
399
- node benchmarks/eval/check-behavioral-release.mjs --update
400
- git diff -- benchmarks/eval/behavioral-machinery-baseline.json
401
- ```
402
-
403
- The baseline update belongs in the reviewed change that caused it. Do not use
404
- the update command merely to make a red gate green. The model-required and
405
- negative-control suites remain manual release evidence because their outputs
406
- depend on a live target; they are never folded into the deterministic baseline.
407
- The projection excludes `latency.wallMs` because scheduler timing is not stable
408
- evidence. Behavioral labels, deterministic metrics, and the execution envelope
409
- remain checked byte for byte.
392
+ `evals/behavioral-machinery-baseline.json` is retained as reviewable reference
393
+ evidence for the machinery corpus. It is not a CI or release gate. Run the
394
+ current `evals/behavioral-machinery.yaml` through the built CLI when a prompt,
395
+ recipe, policy, or expected-behavior change needs a fresh measurement, inspect
396
+ the named scenario evidence, and update any retained baseline deliberately in
397
+ the reviewed change. Model-required and negative-control suites remain manual
398
+ measurements tied to their exact target and serving configuration.
410
399
 
411
400
  ### Hard thresholds and informational budgets
412
401
 
@@ -443,10 +432,12 @@ cannot offset a task or safety regression.
443
432
 
444
433
  ```text
445
434
  serving configuration drift; pass --allow-config-drift to compare these runs
446
- baseline serving: target=mini runtime=llamacpp model=... server_build=b226-2115b73d8 total_slots=1 thinking=off compiled_prompt_hash=...
435
+ baseline serving: target=mini runtime=llamacpp model=ornith1.5-35b-moe server_build=... total_slots=4 thinking=off compiled_prompt_hash=...
447
436
  candidate serving: ...
448
437
  ```
449
438
 
439
+ The current reference `mini` endpoint is the llama.cpp router at `192.168.86.141:8080`. It serves `ornith1.5-35b-moe` with four parallel slots and 262144 context tokens per slot. These deployment facts are reference topology, not defaults imposed on another target; retain the artifact's observed serving configuration with every comparison.
440
+
450
441
  `--allow-config-drift` proceeds and labels the comparison `config drift: allowed`. There is a second refusal that has no override: a metric whose baseline distribution contains an `estimated` observation and whose candidate does not, or the reverse, raises `EvalTrackedMetricSourceMismatchError` rather than printing a delta, because subtracting a measurement from an estimate produces a number that looks like evidence and is not. `--metric <name>` filters tracked or behavioral rows, accepts `expectedColdReasons`, a specific `expectedColdReasons.<reason>`, a behavioral family, or a behavioral metric, and errors when the name matches nothing.
451
442
 
452
443
  ---
@@ -1,7 +1,7 @@
1
1
  # Internal Eval Suites
2
2
 
3
- > [!TIP]
4
- > **Interactive Spec Available:** An interactive blueprint is available at [docs/html/evals_internal_blueprint.html](html/evals_internal_blueprint.html) (Version: 0.4.0).
3
+ > **Visual blueprint:** The source checkout includes the complete
4
+ > [Internal Eval Suites visual reference](https://github.com/iowarp/clio-coder/blob/main/docs/html/evals_internal_blueprint.html).
5
5
 
6
6
  Private suites should live outside this repository. Keep datasets, prompts,
7
7
  live fleet coordinates, calibration outputs, and raw run artifacts in a private
@@ -15,13 +15,13 @@ clio-coder eval run --suite <external-path> --clio-coder-entry dist/cli/index.js
15
15
  ```
16
16
 
17
17
  Use `--out <dir>` when the artifact should be written outside the default Clio
18
- data directory. Product eval artifacts and external benchmark campaigns are
19
- separate: public benchmark adapters live under `benchmarks/community/` and do
20
- not use the eval runner.
18
+ data directory. External benchmark campaigns should adapt their cases and
19
+ grader observations into the same eval engine while keeping private datasets,
20
+ credentials, endpoints, and raw artifacts outside this repository.
21
21
 
22
22
  The public behavioral corpus is the deliberate exception to the otherwise
23
23
  private Suite v2 data policy. Its reviewable, synthetic suites live under
24
- `benchmarks/eval/`: a model-free positive/adversarial authority pair for every
24
+ `evals/`: a model-free positive/adversarial authority pair for every
25
25
  built-in worker recipe, four tiny main-agent model scenarios covering all eight
26
26
  behavioral categories with event- and grader-derived facts, and an intentional
27
27
  decoy negative control. The model-free driver uses the shipped recipe catalog,
@@ -83,7 +83,7 @@ thinking level is measuring the server, not the change under test.
83
83
 
84
84
  The verdict envelope keeps its original `behavioral: null` field for compatibility.
85
85
  A suite that declares a versioned behavioral scenario records the result as a
86
- separate `clio.eval.behavior.v1` document on the Artifact v4 result, cross-linked
86
+ separate `clio-coder.eval.behavior.v1` document on the Artifact v4 result, cross-linked
87
87
  to the unchanged verdict identity. Its labels come only from bounded transcript,
88
88
  tool, receipt, or grader facts, never from an ungrounded judge paragraph. A run whose harness broke records
89
89
  `machinery: "infrastructure_failure"`, which the parser refuses to pair with a
@@ -210,7 +210,7 @@ tasks:
210
210
  - node_modules
211
211
  - dist
212
212
  runner:
213
- kind: clio-run
213
+ kind: clio-coder-run
214
214
  prompt: Fix the intentionally broken function so the local verifier passes.
215
215
  verify:
216
216
  commands:
@@ -271,7 +271,7 @@ tasks:
271
271
  - dist
272
272
  - .clio-coder
273
273
  runner:
274
- kind: clio-run
274
+ kind: clio-coder-run
275
275
  prompt: Summarize the repository purpose and make no file changes.
276
276
  verify:
277
277
  forbidPaths:
@@ -305,7 +305,7 @@ tasks:
305
305
  - dist
306
306
  - .clio-coder
307
307
  runner:
308
- kind: clio-run
308
+ kind: clio-coder-run
309
309
  prompt: Fix the failing unit test with the smallest source change.
310
310
  verify:
311
311
  commands:
@@ -1,7 +1,7 @@
1
1
  # Evolution and Change Manifests
2
2
 
3
- > [!TIP]
4
- > **Interactive Spec Available:** An interactive change manifest editor, authority risk assessor, and checklist workspace is located at [docs/html/evolution_blueprint.html](html/evolution_blueprint.html) (Version: 0.4.0).
3
+ > **Visual blueprint:** The source checkout includes the complete
4
+ > [Evolution and Change Manifests visual reference](https://github.com/iowarp/clio-coder/blob/main/docs/html/evolution_blueprint.html).
5
5
 
6
6
  Clio Coder uses change manifests to make harness changes reviewable, falsifiable, and rollback-friendly. CLIO stands for Context Layer for Input/Output, named for the Greek muse of history. A manifest is JSON, generated or checked with `clio-coder evolve manifest`, and should describe what changed, why, what evidence supports it, what could regress, how to validate it, and how to roll it back.
7
7
 
@@ -1,10 +1,13 @@
1
1
  # Fleet Demo Runbook
2
2
 
3
+ > **Visual blueprint:** The source checkout includes the complete
4
+ > [Fleet Demo Runbook visual reference](https://github.com/iowarp/clio-coder/blob/main/docs/html/fleet_demo_blueprint.html).
5
+
3
6
  A repeatable multi-node demonstration: one orchestrator drives a real
4
7
  CMake/C++ fix through a reviewer-gated dispatch across SSH nodes, and every
5
8
  worker's receipt (including the remote ones) verifies afterward. The steps
6
9
  are executable in order; this document doubles as the recording script.
7
- Background and reference: [fleet-dispatch.md](fleet-dispatch.md).
10
+ Background and reference: [fleet-dispatch.md](../guide/fleet-dispatch.md).
8
11
 
9
12
  ## Reference fabric
10
13
 
@@ -139,9 +142,9 @@ clio-coder evidence inspect <evidenceId>
139
142
  run ledger; a tampered or mismatched receipt fails the build with the field
140
143
  that diverged. The receipts of the remote runs verify on the orchestrator host because the
141
144
  ledger and receipts live on the shared filesystem. Current receipts use strict
142
- v16 and authenticate every current receipt and reconstructed-ledger field.
143
- Every other receipt version is rejected rather than reported as partial; the
144
- current binary has no historical receipt reader.
145
+ v20 and authenticate every current receipt and reconstructed-ledger field.
146
+ Lower versions are reported as retired and are never read as evidence or
147
+ migrated; malformed and future shapes fail verification.
145
148
 
146
149
  ## Provenance walkthrough: what a PI can verify from receipts alone
147
150
 
@@ -171,9 +174,10 @@ reconstruct:
171
174
  complete receipt schema and its stable ledger row. `clio-coder evidence build
172
175
  --run <id>` recomputes and cross-checks it; `verifyReceiptIntegrity` in
173
176
  `src/domains/dispatch/receipt-integrity.ts` is the reference
174
- implementation. Current receipts use v16 and every other version fails
175
- verification. Incompatible state must be archived or removed; it is never
176
- read as evidence through a compatibility verifier.
177
+ implementation. Current receipts use v20. Lower versions are reported as
178
+ retired, while malformed and future shapes fail verification. Incompatible
179
+ state may be archived for inspection, but it is never read as evidence through
180
+ a compatibility verifier.
177
181
 
178
182
  The walkthrough for an audience is three commands: `clio-coder evidence build
179
183
  --run <id>` (it verifies), open the receipt JSON (read `node`, `gate`,
@@ -1,13 +1,20 @@
1
1
  # Git Commit Provenance
2
2
 
3
+ > **Visual blueprint:** The source checkout includes the complete
4
+ > [Git Commit Provenance visual reference](https://github.com/iowarp/clio-coder/blob/main/docs/html/git_commit_provenance_blueprint.html).
5
+
3
6
  Clio Coder adds evidence-aware role trailers to commits created through Clio.
4
7
  The feature is enabled by default:
5
8
 
6
9
  ```yaml
7
- attribution:
8
- gitCommits: true
10
+ integrations:
11
+ git:
12
+ commitAttribution: true
9
13
  ```
10
14
 
15
+ The released `attribution.gitCommits` path is a migration alias. Current
16
+ settings files and writers use `integrations.git.commitAttribution`.
17
+
11
18
  Settings -> Advanced exposes the same switch as **Clio commit provenance**, with
12
19
  `enabled` and `disabled` values. A change applies immediately to subsequent
13
20
  commits in the session. When disabled, Clio leaves commit messages entirely
@@ -49,11 +56,11 @@ Co-authored-by: Clio Coder <clio-coder@iowarp.ai>
49
56
  Existing human trailers stay in place. A Clio trailer already present in any
50
57
  letter case is respected rather than repeated, line endings are normalized only
51
58
  while attribution is enabled, and repeated processing is idempotent. When a directly relevant
52
- receipt-v19 digest passes integrity verification, Clio may additionally add the
59
+ receipt-v20 digest passes integrity verification, Clio may additionally add the
53
60
  full digest:
54
61
 
55
62
  ```text
56
- Clio-Evidence: receipt-v19/sha256:<64-character digest>
63
+ Clio-Evidence: receipt-v20/sha256:<64-character digest>
57
64
  ```
58
65
 
59
66
  Clio does not invent, shorten, or add an unrelated digest. The role trailers do