@iowarp/clio-coder 0.4.1 → 0.4.2

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (437) hide show
  1. package/CHANGELOG.md +92 -0
  2. package/CONTRIBUTING.md +59 -36
  3. package/README.md +404 -472
  4. package/SECURITY.md +2 -1
  5. package/dist/{acp-ZILU3AUO.js → acp-TMDQZDIG.js} +7 -7
  6. package/dist/{agents-HYWGBGQR.js → agents-5N5NG3XG.js} +28 -28
  7. package/dist/assets/codewiki.json +1 -1
  8. package/dist/{auth-N3QT7CBO.js → auth-Z5CCBXKQ.js} +8 -9
  9. package/dist/{builtins-UJLMOVOV.js → builtins-K6TNDT24.js} +4 -4
  10. package/dist/{chunk-GVQJ5CCZ.js → chunk-2HFQNRV3.js} +7 -7
  11. package/dist/{chunk-QMXC4JB7.js → chunk-2NHR3NAY.js} +163 -1401
  12. package/dist/chunk-2X4RYJTJ.js +39 -0
  13. package/dist/{chunk-Y45G3AXC.js → chunk-2Z2IKEXI.js} +6 -10
  14. package/dist/{chunk-EIMVLWB3.js → chunk-34BHNEE3.js} +7 -3
  15. package/dist/{chunk-GIZNH63R.js → chunk-35MSIRKH.js} +9 -4
  16. package/dist/chunk-3EBYEESD.js +314 -0
  17. package/dist/{chunk-CTJ4RNAA.js → chunk-3F7VUY77.js} +2 -2
  18. package/dist/{chunk-AP73CFDC.js → chunk-3KIPBMUA.js} +2 -2
  19. package/dist/{chunk-J5LZHVIT.js → chunk-3M6DQK6S.js} +113 -35
  20. package/dist/{chunk-VEGN6WIQ.js → chunk-462T4EGZ.js} +2 -2
  21. package/dist/{chunk-AFKWHWXF.js → chunk-4JDLP6ZS.js} +33 -16
  22. package/dist/{chunk-6FN3E6KX.js → chunk-4O6MANBS.js} +2 -2
  23. package/dist/{chunk-AKB4GYDL.js → chunk-54ODD65L.js} +5 -5
  24. package/dist/{chunk-BBTJOK6Y.js → chunk-5KW52TEP.js} +3 -3
  25. package/dist/{chunk-6CCS4G3W.js → chunk-5PFYMY2V.js} +2 -2
  26. package/dist/chunk-77QIVUZB.js +1334 -0
  27. package/dist/{chunk-7OBGU7UB.js → chunk-7BHIY2MW.js} +7 -13
  28. package/dist/{chunk-3QSOM6PA.js → chunk-AZ4WMN4W.js} +2 -2
  29. package/dist/{chunk-6NJQITNH.js → chunk-B74PXLU7.js} +6 -3
  30. package/dist/{chunk-R23Z6K6I.js → chunk-B7HM5Z7T.js} +15 -15
  31. package/dist/{chunk-R32CLGZ6.js → chunk-BO7Y52RY.js} +81 -20
  32. package/dist/{chunk-UEDMSP56.js → chunk-BYMNWQ7O.js} +123 -148
  33. package/dist/{chunk-ZJLUDYFY.js → chunk-CRFOIAX3.js} +4 -4
  34. package/dist/{chunk-2NM363SV.js → chunk-CYZW7JHJ.js} +7 -7
  35. package/dist/{chunk-6HMJX2VU.js → chunk-DYHAXKHD.js} +38 -10
  36. package/dist/{chunk-THYWACCR.js → chunk-DZAW46HP.js} +3 -3
  37. package/dist/{chunk-FYUN5KZ3.js → chunk-DZEK6CJN.js} +17 -17
  38. package/dist/{chunk-3I5NY75V.js → chunk-E7GT7O5N.js} +5 -5
  39. package/dist/{chunk-VKFQTNDV.js → chunk-F2I26BDK.js} +4 -4
  40. package/dist/{chunk-HLW2MRKE.js → chunk-F4EKGO4N.js} +3 -1
  41. package/dist/{chunk-IXJT6DCX.js → chunk-FVDGR2ZL.js} +3 -3
  42. package/dist/{chunk-TZSKNMZG.js → chunk-GTUD2WMY.js} +2 -1
  43. package/dist/{chunk-7EPLI7VL.js → chunk-HIICAHCJ.js} +2 -2
  44. package/dist/{chunk-E67WX76H.js → chunk-HKMD33FO.js} +29 -80
  45. package/dist/chunk-HLAFFSEK.js +360 -0
  46. package/dist/{chunk-UAPGZHYC.js → chunk-I64IFBLB.js} +9 -2
  47. package/dist/{chunk-XKA2ICR3.js → chunk-I66ZTYNP.js} +440 -175
  48. package/dist/{chunk-7PWAODYW.js → chunk-I7XBWTYH.js} +2 -2
  49. package/dist/{chunk-PVAMAVBB.js → chunk-IDNA72AH.js} +102 -2
  50. package/dist/{chunk-GCSMB2KY.js → chunk-IKOZFYBN.js} +1 -1
  51. package/dist/{chunk-2VG7KLYV.js → chunk-IKSLQ4XV.js} +5460 -3241
  52. package/dist/{chunk-QKIFBZKT.js → chunk-IMXMHHMQ.js} +166 -25
  53. package/dist/{chunk-74YWRRU5.js → chunk-JBCS7CRR.js} +2 -2
  54. package/dist/{chunk-BDPT6GTK.js → chunk-JWJGP5DQ.js} +2 -2
  55. package/dist/{chunk-K6BF4U2H.js → chunk-KKOJXO6R.js} +62 -14
  56. package/dist/chunk-KPXDY6QF.js +47 -0
  57. package/dist/{chunk-ABLSQ6JX.js → chunk-LJID3DYZ.js} +7 -1
  58. package/dist/{chunk-VKRH2TCS.js → chunk-M2DAX4F6.js} +2 -2
  59. package/dist/{chunk-6I5ILFOF.js → chunk-M2WXEHER.js} +2 -2
  60. package/dist/{chunk-YPI3QQCF.js → chunk-MCEPRMZW.js} +2 -4
  61. package/dist/{chunk-N5UK64DP.js → chunk-MCMZMDAC.js} +2 -2
  62. package/dist/{chunk-Y4CAGMM6.js → chunk-MNJGS2IN.js} +5 -6
  63. package/dist/{chunk-TVHHYFHE.js → chunk-NEDJ26B5.js} +2 -2
  64. package/dist/{chunk-U2WB7TZS.js → chunk-NMJXSHBJ.js} +97 -85
  65. package/dist/{chunk-HUAS7ITX.js → chunk-O3YUNJZ2.js} +13 -21
  66. package/dist/{chunk-MA3H6DM5.js → chunk-P75RZCJW.js} +25 -3
  67. package/dist/{chunk-IG7BCQBA.js → chunk-PGF63K6I.js} +2 -2
  68. package/dist/chunk-PJX3WQUQ.js +42 -0
  69. package/dist/{chunk-6DWBAZ5U.js → chunk-Q4XWMHX6.js} +4 -6
  70. package/dist/{chunk-OJTRZGR3.js → chunk-QQLGQY2A.js} +8 -8
  71. package/dist/{chunk-J4HBWF6Y.js → chunk-RLYRBIYQ.js} +115 -20
  72. package/dist/{chunk-NLFAQR7Z.js → chunk-S66XZJOF.js} +3 -23
  73. package/dist/{chunk-C537JADH.js → chunk-SSEYRH53.js} +6 -7
  74. package/dist/chunk-SZAA6XDG.js +30 -0
  75. package/dist/{chunk-MOPSG2X7.js → chunk-TPEQIQIE.js} +6 -6
  76. package/dist/{chunk-JA5QWE4Z.js → chunk-UBRFI4HS.js} +1879 -1650
  77. package/dist/{chunk-BTGG6BG2.js → chunk-UH347SHR.js} +154 -15
  78. package/dist/{chunk-5YHDIDBP.js → chunk-UH632ZYL.js} +2 -2
  79. package/dist/{chunk-BWW4HLO4.js → chunk-UXCU4E3T.js} +8 -6
  80. package/dist/{chunk-6VC4OV3Z.js → chunk-VIA6RFQZ.js} +3 -11
  81. package/dist/{chunk-ZAZB4JMW.js → chunk-VKPAQYEB.js} +27 -8
  82. package/dist/{chunk-UXN6JT4W.js → chunk-W4YEMFBX.js} +2 -2
  83. package/dist/{chunk-TD3PGPQA.js → chunk-W6NIE6OW.js} +2 -2
  84. package/dist/{chunk-TVH4ONAM.js → chunk-X7IARSHT.js} +3 -3
  85. package/dist/{chunk-PJJ6MY27.js → chunk-XE3PCIXH.js} +3 -3
  86. package/dist/{chunk-FEFIFZTL.js → chunk-XGDPUNND.js} +2 -2
  87. package/dist/{chunk-SCYB3HA4.js → chunk-XOXV5GKE.js} +51 -16
  88. package/dist/{chunk-QTFGO774.js → chunk-XQRY4DTA.js} +24 -11
  89. package/dist/{chunk-BJGUKIG4.js → chunk-YJISEZKC.js} +2 -2
  90. package/dist/{chunk-GPPB3JBE.js → chunk-ZGNYYXQ6.js} +2 -2
  91. package/dist/{chunk-SINK3QR6.js → chunk-ZNT2M6TG.js} +7 -7
  92. package/dist/{chunk-7RY5VZPH.js → chunk-ZW4HH5JJ.js} +6 -6
  93. package/dist/cli/index.js +33 -32
  94. package/dist/{clio-IT3G3VQH.js → clio-7VB377CC.js} +7 -7
  95. package/dist/{code-nav-RK6S7F6E.js → code-nav-YVLCYA7V.js} +85 -17
  96. package/dist/{config-3QZRWZJF.js → config-4HVOS65E.js} +88 -43
  97. package/dist/{configure-FL7Y3KJF.js → configure-PIWO7B24.js} +10 -10
  98. package/dist/{context-5HE7ODYK.js → context-IYEHL3WQ.js} +33 -31
  99. package/dist/{context-XNHL75JV.js → context-KQYIWPWT.js} +47 -34
  100. package/dist/{context-KYQFRVDC.js → context-N6ZE3LGJ.js} +11 -11
  101. package/dist/{context-clear-N545L53A.js → context-clear-G4OGZJDS.js} +33 -31
  102. package/dist/{context-working-set-QHKXSV2F.js → context-working-set-BWLF6LJP.js} +7 -7
  103. package/dist/{dispatch-runner-RGIE5PCT.js → dispatch-runner-2QQAITS3.js} +38 -38
  104. package/dist/{docs-5NAF6AU7.js → docs-PD3EXDKU.js} +21 -20
  105. package/dist/{doctor-ZGPEGHIP.js → doctor-LHBD36VU.js} +23 -22
  106. package/dist/{eval-GXLL44RD.js → eval-C45FYRJ6.js} +21 -20
  107. package/dist/{eval-inventory-HBWSWQOK.js → eval-inventory-6DEJPLBF.js} +2 -2
  108. package/dist/{evidence-HWLBRH3Q.js → evidence-6SHONYAF.js} +30 -28
  109. package/dist/{evolve-FTZBMNVW.js → evolve-KRKMV72X.js} +30 -28
  110. package/dist/{extensions-VHRBEID7.js → extensions-KPZ2UHBB.js} +5 -3
  111. package/dist/{fleet-CKZHJWZJ.js → fleet-IVTCKDHT.js} +62 -61
  112. package/dist/{fleet-commands-EXDXBMV6.js → fleet-commands-EDWL3IT7.js} +5 -5
  113. package/dist/{fleet-decisions-OTHB6KRL.js → fleet-decisions-YP3YEFGK.js} +4 -4
  114. package/dist/{fleet-graph-YTEZUCUT.js → fleet-graph-ZFWKHY2M.js} +16 -14
  115. package/dist/{fleet-inspect-SS6YMDCK.js → fleet-inspect-FVUNCBML.js} +31 -29
  116. package/dist/{fleet-preflight-PBY4VYOM.js → fleet-preflight-UN5XED4R.js} +2 -2
  117. package/dist/{fleet-validate-KMEM5L3S.js → fleet-validate-XOWC4HSX.js} +17 -15
  118. package/dist/{fleet-verify-QD5M7E7Q.js → fleet-verify-UN3SODEL.js} +30 -28
  119. package/dist/{fleet-view-WAMJYNDT.js → fleet-view-TWHJKCN6.js} +31 -29
  120. package/dist/{init-5XQRBOFV.js → init-T2QORQ3Y.js} +50 -49
  121. package/dist/{interop-34TVO25M.js → interop-IN5I2A66.js} +5 -5
  122. package/dist/{library-3QY6KF57.js → library-LSCATDLZ.js} +15 -13
  123. package/dist/{memory-L4UTIIIW.js → memory-HYOKAGGJ.js} +31 -29
  124. package/dist/{models-ZVX3QOWE.js → models-2GPMFYCM.js} +22 -21
  125. package/dist/{monitor-CEKVSYTS.js → monitor-E4ASVUJH.js} +34 -32
  126. package/dist/{orchestrator-77BAP6BC.js → orchestrator-DDMPR3PY.js} +984 -583
  127. package/dist/{panes-7STHOAUJ.js → panes-E3RUXOW5.js} +4 -4
  128. package/dist/{panes-SHAUIRXY.js → panes-IXKLOKA2.js} +23 -8
  129. package/dist/{reset-EOLM7GVE.js → reset-OAQP3W4O.js} +4 -4
  130. package/dist/{resources-74GKTLSF.js → resources-OTRSN34L.js} +15 -13
  131. package/dist/{run-HBAUJNNZ.js → run-5DEYH5QK.js} +60 -59
  132. package/dist/{share-G3APVLVP.js → share-IHWTLO3M.js} +19 -15
  133. package/dist/{skills-35HHUKCR.js → skills-IYMXMKW4.js} +17 -15
  134. package/dist/{skills-eval-QN4HSHDC.js → skills-eval-DROHSJAR.js} +36 -36
  135. package/dist/{skills-inventory-J357J34F.js → skills-inventory-D7X4L4ZX.js} +15 -13
  136. package/dist/{slash-commands-JZZCQA32.js → slash-commands-QBM7UZ3B.js} +21 -18
  137. package/dist/{steer-XAVHJM22.js → steer-Z5DO23FJ.js} +2 -2
  138. package/dist/{targets-DSM6CY3M.js → targets-P2FUC4IL.js} +25 -28
  139. package/dist/{terminal-lease-JOPFUVEM.js → terminal-lease-YREJ3JX2.js} +5 -5
  140. package/dist/{tools-MKNWVPBH.js → tools-5B7RO6MV.js} +4 -4
  141. package/dist/{trace-ECQ7TIYZ.js → trace-YMGMUM6A.js} +55 -7
  142. package/dist/{upgrade-H7TOM7YL.js → upgrade-PXK3S2YM.js} +11 -9
  143. package/dist/{usage-X52N3IDJ.js → usage-ME5MPXGX.js} +36 -34
  144. package/dist/{verifiers-EJTVVSMA.js → verifiers-BVZ7IWOO.js} +5 -5
  145. package/dist/{verify-YJL6XET2.js → verify-5K7ZKQFC.js} +4 -4
  146. package/dist/{web-fetch-MPIFL3LL.js → web-fetch-MPARV2K7.js} +2 -2
  147. package/dist/{wiki-generate-4NDZTQ4B.js → wiki-generate-F5W5QTYY.js} +48 -47
  148. package/dist/{with-panes-OBOBFIIR.js → with-panes-BYOJCLAM.js} +51 -255
  149. package/dist/worker/entry.js +45 -30
  150. package/docs/README.md +176 -81
  151. package/docs/{acp.md → architecture/acp.md} +36 -20
  152. package/docs/{alcf-provider.md → architecture/alcf-provider.md} +8 -5
  153. package/docs/{architecture.md → architecture/architecture.md} +43 -22
  154. package/docs/{artifact-placement.md → architecture/artifact-placement.md} +26 -23
  155. package/docs/architecture/artifact-versions.md +90 -0
  156. package/docs/{capacity-and-scheduling.md → architecture/capacity-and-scheduling.md} +26 -13
  157. package/docs/{context-engine.md → architecture/context-engine.md} +25 -25
  158. package/docs/{context-working-set.md → architecture/context-working-set.md} +13 -10
  159. package/docs/{dispatch-architecture-rationale.md → architecture/dispatch-architecture-rationale.md} +12 -9
  160. package/docs/{dispatch-typed-intent.md → architecture/dispatch-typed-intent.md} +68 -46
  161. package/docs/{evidence-and-memory.md → architecture/evidence-and-memory.md} +23 -16
  162. package/docs/{middleware-and-components.md → architecture/middleware-and-components.md} +11 -5
  163. package/docs/{model-catalog.md → architecture/model-catalog.md} +40 -17
  164. package/docs/{observability.md → architecture/observability.md} +26 -13
  165. package/docs/{pi-boundary.md → architecture/pi-boundary.md} +24 -11
  166. package/docs/{prompt-envelope-and-tools.md → architecture/prompt-envelope-and-tools.md} +55 -20
  167. package/docs/{provider-adapter-cookbook.md → architecture/provider-adapter-cookbook.md} +35 -24
  168. package/docs/{safety-model.md → architecture/safety-model.md} +20 -15
  169. package/docs/{session-lifecycle.md → architecture/session-lifecycle.md} +8 -5
  170. package/docs/architecture/time-conventions.md +125 -0
  171. package/docs/{trace-store.md → architecture/trace-store.md} +13 -5
  172. package/docs/{tui-design.md → architecture/tui-design.md} +13 -13
  173. package/docs/{worker-dispatch-mechanics.md → architecture/worker-dispatch-mechanics.md} +27 -30
  174. package/docs/{built-in-agents.md → guide/built-in-agents.md} +50 -34
  175. package/docs/{commands-and-modes.md → guide/commands-and-modes.md} +65 -60
  176. package/docs/{configuration-and-targets.md → guide/configuration-and-targets.md} +227 -289
  177. package/docs/guide/configuration-reference.md +1158 -0
  178. package/docs/{environment-variables.md → guide/environment-variables.md} +31 -28
  179. package/docs/{exit-codes-and-output.md → guide/exit-codes-and-output.md} +6 -3
  180. package/docs/{extensions-and-sharing.md → guide/extensions-and-sharing.md} +41 -14
  181. package/docs/{fleet-dispatch.md → guide/fleet-dispatch.md} +39 -43
  182. package/docs/{glossary.md → guide/glossary.md} +14 -11
  183. package/docs/{installation-and-lifecycle.md → guide/installation-and-lifecycle.md} +44 -13
  184. package/docs/guide/panes-and-files.md +290 -0
  185. package/docs/{proactive-memory.md → guide/proactive-memory.md} +79 -66
  186. package/docs/{resource-library.md → guide/resource-library.md} +13 -4
  187. package/docs/{skills-marketplace.md → guide/skills-marketplace.md} +7 -3
  188. package/docs/{tool-usage.md → guide/tool-usage.md} +87 -23
  189. package/docs/{troubleshooting.md → guide/troubleshooting.md} +9 -4
  190. package/docs/{config-knobs-audit.md → history/config-knobs-audit.md} +11 -11
  191. package/docs/{release-cut-checklist.md → history/release-cut-checklist.md} +29 -2
  192. package/docs/{development-pipeline.md → process/development-pipeline.md} +24 -26
  193. package/docs/process/documentation-coverage.md +100 -0
  194. package/docs/process/documentation-guide.md +187 -0
  195. package/docs/{eval-runner.md → process/eval-runner.md} +41 -50
  196. package/docs/{evals-internal.md → process/evals-internal.md} +10 -10
  197. package/docs/{evolution.md → process/evolution.md} +2 -2
  198. package/docs/{fleet-demo-runbook.md → process/fleet-demo-runbook.md} +11 -7
  199. package/docs/{git-commit-provenance.md → process/git-commit-provenance.md} +11 -4
  200. package/docs/{performance-methodology.md → process/performance-methodology.md} +87 -69
  201. package/docs/{scientific-validation.md → process/scientific-validation.md} +4 -4
  202. package/evals/README.md +2 -2
  203. package/package.json +9 -7
  204. package/skills/README.md +46 -37
  205. package/skills/coding/ast-grep/SKILL.md +2 -2
  206. package/skills/coding/coding-standards/SKILL.md +2 -2
  207. package/skills/coding/prototype/SKILL.md +2 -2
  208. package/skills/coding/tdd/SKILL.md +2 -2
  209. package/skills/context/context-handoff/SKILL.md +2 -2
  210. package/skills/context/context-prime/SKILL.md +2 -2
  211. package/skills/git/file-ticket/SKILL.md +2 -2
  212. package/skills/git/fix-issue/SKILL.md +3 -3
  213. package/skills/git/resolve-merge-conflicts/SKILL.md +2 -2
  214. package/skills/git/ship/SKILL.md +2 -2
  215. package/skills/git/worktree-create/SKILL.md +2 -2
  216. package/skills/git/worktree-merge/SKILL.md +2 -2
  217. package/skills/meta/clio-coder-dev/SKILL.md +9 -5
  218. package/skills/meta/clio-coder-dev/evals.md +3 -2
  219. package/skills/meta/clio-coder-test/SKILL.md +102 -95
  220. package/skills/meta/clio-coder-test/evals.md +9 -4
  221. package/skills/meta/clio-coder-test/references/harness.md +100 -124
  222. package/skills/meta/clio-coder-test/references/test-map.md +77 -50
  223. package/skills/meta/credentials/SKILL.md +2 -2
  224. package/skills/meta/find-skills/SKILL.md +2 -2
  225. package/skills/meta/herdr/SKILL.md +2 -2
  226. package/skills/meta/skill-craft/SKILL.md +22 -16
  227. package/skills/planning/architecture/SKILL.md +2 -2
  228. package/skills/planning/backlog/SKILL.md +2 -2
  229. package/skills/planning/prd/SKILL.md +2 -2
  230. package/skills/planning/product-intent/SKILL.md +2 -2
  231. package/skills/planning/tech-spec/SKILL.md +2 -2
  232. package/skills/registry.yaml +62 -62
  233. package/skills/research/arxiv-literature/SKILL.md +2 -2
  234. package/skills/research/experiment-protocol/SKILL.md +2 -2
  235. package/skills/research/scientific-debugging/SKILL.md +2 -2
  236. package/skills/research/scientific-modernization/SKILL.md +2 -2
  237. package/skills/skill-marketplace.json +62 -62
  238. package/skills/workflow/cut-it/SKILL.md +2 -2
  239. package/skills/workflow/design-council/SKILL.md +2 -2
  240. package/skills/workflow/grill-me/SKILL.md +2 -2
  241. package/skills/workflow/workflow-distiller/SKILL.md +2 -2
  242. package/src/cli/args.ts +2 -2
  243. package/src/cli/bootstrap-generate.ts +1 -1
  244. package/src/cli/config-inspect.ts +65 -12
  245. package/src/cli/configure.ts +0 -4
  246. package/src/cli/docs.ts +22 -14
  247. package/src/cli/doctor-naming.ts +5 -5
  248. package/src/cli/doctor-toolchain.ts +3 -3
  249. package/src/cli/eval.ts +1 -2
  250. package/src/cli/extensions.ts +2 -1
  251. package/src/cli/fleet.ts +1 -1
  252. package/src/cli/index.ts +2 -1
  253. package/src/cli/internal-dispatch.ts +3 -4
  254. package/src/cli/panes.ts +19 -5
  255. package/src/cli/run.ts +2 -2
  256. package/src/cli/share.ts +5 -1
  257. package/src/cli/skills-eval.ts +3 -3
  258. package/src/cli/targets.ts +2 -6
  259. package/src/cli/trace.ts +55 -4
  260. package/src/cli/wiki-generate.ts +1 -1
  261. package/src/core/artifact-paths.ts +1 -1
  262. package/src/core/bash-exec.ts +131 -86
  263. package/src/core/bus-events.ts +51 -6
  264. package/src/core/config.ts +5 -1
  265. package/src/core/defaults.ts +7 -4
  266. package/src/core/dispatch-outcome.ts +16 -0
  267. package/src/core/guardrails.ts +10 -49
  268. package/src/core/prompt-hint.ts +9 -0
  269. package/src/domains/agents/builtins/architect.md +2 -3
  270. package/src/domains/agents/builtins/coder.md +3 -2
  271. package/src/domains/agents/builtins/debugger.md +2 -2
  272. package/src/domains/agents/builtins/documenter.md +2 -2
  273. package/src/domains/agents/builtins/git-master.md +1 -1
  274. package/src/domains/agents/builtins/oracle.md +1 -1
  275. package/src/domains/agents/builtins/provenance.md +1 -1
  276. package/src/domains/agents/builtins/researcher.md +1 -1
  277. package/src/domains/agents/builtins/scout.md +1 -1
  278. package/src/domains/agents/builtins/tester.md +2 -2
  279. package/src/domains/agents/builtins/verifier.md +2 -2
  280. package/src/domains/agents/builtins/wiki-writer.md +1 -1
  281. package/src/domains/agents/catalog.ts +12 -14
  282. package/src/domains/agents/contract.ts +2 -0
  283. package/src/domains/agents/extension.ts +23 -1
  284. package/src/domains/config/keybindings.ts +8 -0
  285. package/src/domains/context/extension.ts +0 -3
  286. package/src/domains/context/working-set/path-index.ts +1 -0
  287. package/src/domains/dispatch/capability-match.ts +10 -0
  288. package/src/domains/dispatch/extension.ts +105 -22
  289. package/src/domains/dispatch/host-verification.ts +435 -39
  290. package/src/domains/dispatch/intent-requirements.ts +10 -0
  291. package/src/domains/dispatch/intent.ts +18 -1
  292. package/src/domains/dispatch/path-scope.ts +235 -24
  293. package/src/domains/dispatch/run-event-journal.ts +4 -15
  294. package/src/domains/dispatch/state.ts +2 -3
  295. package/src/domains/dispatch/transport.ts +45 -21
  296. package/src/domains/dispatch/types.ts +55 -3
  297. package/src/domains/eval/artifacts/store.ts +5 -0
  298. package/src/domains/eval/store.ts +8 -1
  299. package/src/domains/evidence/trust-status.ts +10 -1
  300. package/src/domains/extensions/contract.ts +15 -1
  301. package/src/domains/extensions/discovery.ts +238 -41
  302. package/src/domains/extensions/extension.ts +105 -6
  303. package/src/domains/extensions/index.ts +24 -0
  304. package/src/domains/extensions/integrity.ts +189 -0
  305. package/src/domains/extensions/manager.ts +17 -1
  306. package/src/domains/extensions/resource-path.ts +27 -0
  307. package/src/domains/extensions/resources.ts +18 -38
  308. package/src/domains/extensions/snapshot-store.ts +39 -0
  309. package/src/domains/extensions/snapshot.ts +180 -0
  310. package/src/domains/extensions/state.ts +385 -57
  311. package/src/domains/extensions/types.ts +118 -1
  312. package/src/domains/lifecycle/migrations/2026-09-01-extension-install-digests.ts +27 -0
  313. package/src/domains/lifecycle/migrations/index.ts +2 -0
  314. package/src/domains/lifecycle/naming-resources.ts +19 -4
  315. package/src/domains/lifecycle/naming-yazi.ts +10 -5
  316. package/src/domains/middleware/contract.ts +26 -0
  317. package/src/domains/middleware/extension.ts +24 -24
  318. package/src/domains/middleware/hook-receipts.ts +27 -4
  319. package/src/domains/middleware/hooks-io.ts +65 -32
  320. package/src/domains/middleware/hooks.ts +64 -0
  321. package/src/domains/middleware/index.ts +28 -4
  322. package/src/domains/middleware/registrations.ts +326 -0
  323. package/src/domains/middleware/runtime.ts +28 -0
  324. package/src/domains/middleware/snapshot.ts +20 -7
  325. package/src/domains/mux/contract.ts +38 -0
  326. package/src/domains/mux/detect.ts +6 -13
  327. package/src/domains/mux/index.ts +1 -1
  328. package/src/domains/mux/operations.ts +44 -5
  329. package/src/domains/mux/yazi/assets/yazi.toml +2 -2
  330. package/src/domains/mux/yazi/session.ts +53 -4
  331. package/src/domains/mux/yazi/theme.ts +117 -17
  332. package/src/domains/observability/contract.ts +10 -11
  333. package/src/domains/observability/extension.ts +11 -3
  334. package/src/domains/observability/projection.ts +14 -90
  335. package/src/domains/observability/trace-store.ts +43 -7
  336. package/src/domains/prompts/compiler.ts +73 -53
  337. package/src/domains/prompts/contract.ts +15 -3
  338. package/src/domains/prompts/extension.ts +97 -9
  339. package/src/domains/prompts/fragments/identity/clio-worker.md +1 -3
  340. package/src/domains/prompts/fragments/identity/clio.md +6 -12
  341. package/src/domains/prompts/fragments/identity/docs-routing.md +1 -2
  342. package/src/domains/prompts/fragments/identity/self-awareness.md +3 -11
  343. package/src/domains/prompts/fragments/operating/contract.md +7 -15
  344. package/src/domains/prompts/fragments/operating/delegation.md +32 -34
  345. package/src/domains/prompts/fragments/operating/skills.md +10 -24
  346. package/src/domains/prompts/fragments/operating/worker.md +1 -8
  347. package/src/domains/providers/index.ts +1 -1
  348. package/src/domains/providers/model-runtime-capabilities.ts +85 -21
  349. package/src/domains/providers/models/local-models/clio-coder-local-coding-targets.yaml +669 -104
  350. package/src/domains/providers/runtime-resolution.ts +31 -0
  351. package/src/domains/providers/runtimes/common/probe-helpers.ts +7 -2
  352. package/src/domains/providers/runtimes/local-native/llamacpp.ts +9 -1
  353. package/src/domains/providers/types/cost-provenance.ts +19 -0
  354. package/src/domains/providers/types/local-model-quirks.ts +85 -37
  355. package/src/domains/resources/skills/loader.ts +16 -19
  356. package/src/domains/safety/call-target.ts +1 -1
  357. package/src/domains/safety/loop-detector.ts +7 -4
  358. package/src/domains/session/task-board.ts +10 -9
  359. package/src/domains/share/archive.ts +164 -7
  360. package/src/engine/acp/server.ts +62 -9
  361. package/src/engine/apis/llamacpp-residency.ts +3 -4
  362. package/src/engine/apis/lmstudio.ts +3 -3
  363. package/src/engine/apis/ollama-native.ts +6 -6
  364. package/src/engine/apis/openai-completions.ts +28 -25
  365. package/src/engine/apis/output-budget.ts +8 -18
  366. package/src/engine/apis/residency.ts +8 -27
  367. package/src/engine/gemma-channel-filter.ts +19 -0
  368. package/src/engine/loop-guard.ts +92 -12
  369. package/src/engine/worker-runtime.ts +40 -11
  370. package/src/engine/worker-tools.ts +3 -1
  371. package/src/entry/extension-hook-sources.ts +28 -0
  372. package/src/entry/extension-reload.ts +309 -0
  373. package/src/entry/orchestrator.ts +59 -35
  374. package/src/interactive/application-controller.ts +2 -1
  375. package/src/interactive/bus-notices.ts +8 -1
  376. package/src/interactive/chat-loop-messages.ts +3 -13
  377. package/src/interactive/chat-loop.ts +10 -1
  378. package/src/interactive/chat-panel.ts +36 -13
  379. package/src/interactive/chat-renderer.ts +71 -7
  380. package/src/interactive/dispatch-board.ts +6 -11
  381. package/src/interactive/footer/widgets.ts +13 -0
  382. package/src/interactive/interactive-application.ts +39 -4
  383. package/src/interactive/interactive-input-runtime.ts +4 -0
  384. package/src/interactive/interactive-presentation.ts +2 -2
  385. package/src/interactive/interactive-slash-runtime.ts +2 -0
  386. package/src/interactive/overlays/extensions.ts +9 -1
  387. package/src/interactive/overlays/help-reference.ts +13 -0
  388. package/src/interactive/overlays/settings.ts +27 -16
  389. package/src/interactive/panes-runtime.ts +111 -35
  390. package/src/interactive/prompt-cache-identity.ts +88 -0
  391. package/src/interactive/slash-commands.ts +129 -14
  392. package/src/interactive/stream-pacing-policy.ts +0 -23
  393. package/src/interactive/turn-context.ts +30 -15
  394. package/src/interactive/yazi-bridge.ts +60 -6
  395. package/src/tools/agent-tools.ts +30 -1
  396. package/src/tools/artifact.ts +2 -2
  397. package/src/tools/ask-user.ts +3 -3
  398. package/src/tools/bash.ts +1 -1
  399. package/src/tools/bootstrap.ts +4 -0
  400. package/src/tools/builtin-tool-catalog.ts +52 -22
  401. package/src/tools/codewiki/code-nav-surface.ts +6 -0
  402. package/src/tools/codewiki/code-nav.ts +99 -13
  403. package/src/tools/context/docs-engine.ts +20 -7
  404. package/src/tools/context/index.ts +29 -12
  405. package/src/tools/core-bootstrap.ts +28 -6
  406. package/src/tools/credential-present.ts +1 -2
  407. package/src/tools/dispatch-arguments.ts +5 -1
  408. package/src/tools/dispatch-plan.ts +48 -4
  409. package/src/tools/dispatch-run-events.ts +1 -1
  410. package/src/tools/dispatch-schema.ts +338 -0
  411. package/src/tools/dispatch-types.ts +3 -0
  412. package/src/tools/dispatch.ts +9 -254
  413. package/src/tools/ledger.ts +3 -5
  414. package/src/tools/monitor-surface.ts +5 -13
  415. package/src/tools/observation.ts +4 -5
  416. package/src/tools/panes-surface.ts +4 -11
  417. package/src/tools/panes.ts +4 -2
  418. package/src/tools/policy.ts +15 -2
  419. package/src/tools/read.ts +5 -6
  420. package/src/tools/registry.ts +30 -7
  421. package/src/tools/result-shaping.ts +18 -14
  422. package/src/tools/steer-surface.ts +1 -1
  423. package/src/tools/tasks.ts +1 -1
  424. package/src/tools/truncate.ts +6 -5
  425. package/src/tools/verify/surface.ts +6 -12
  426. package/src/tools/web-fetch-surface.ts +1 -3
  427. package/dist/chunk-5QIAJV2D.js +0 -48
  428. package/dist/chunk-JZWT5J3Y.js +0 -814
  429. package/dist/chunk-K7VKOLQQ.js +0 -15
  430. package/dist/chunk-PMZCIOCJ.js +0 -25
  431. package/dist/chunk-SUW5DORT.js +0 -819
  432. package/dist/chunk-UOV2BYIW.js +0 -107
  433. package/dist/chunk-WR6U3OVP.js +0 -45
  434. package/docs/artifact-versions.md +0 -67
  435. package/docs/documentation-coverage.md +0 -46
  436. package/docs/documentation-guide.md +0 -167
  437. package/docs/time-conventions.md +0 -101
package/CHANGELOG.md CHANGED
@@ -2,6 +2,98 @@
2
2
 
3
3
  All notable changes to Clio Coder are documented in this file. The format follows [Keep a Changelog](https://keepachangelog.com/en/1.1.0/) and versions follow Semantic Versioning; pre-1.0 minor releases may include incompatible changes.
4
4
 
5
+ ## 0.4.2 - 2026-09-02
6
+
7
+ ### Added
8
+ - Accountability reaches ACP. The observability extension publishes `accountability.evidenceReady` on the bus once per run whose evidence bundle and index row landed, with the payload spread from the same `ObservabilityRunEvidence` object it hands the projection, and the ACP server forwards it as the opt-in `accountability.evidenceReady` kind (`{runId,evidenceId,firstPassSuccess,findingCount,tags}`, `terminal:true`) through the shared `forwardEvent` sender. The v0.4.2 telemetry audit found first-pass success and finding counts reached the TUI, `usage`, `evidence-detail`, and the trace viewer but no ACP event, not even opt-in (#271). Bundle prose never crosses; a failed build sends nothing. Contract: `tests/contracts/acp-evidence-ready-event.test.ts`.
9
+ - `clio-coder trace code-steps <rootId> [--json]` reads back the deterministic code-step records every fleet code step writes to `<stateDir>/code-steps/<rootId>/`. The v0.4.2 telemetry audit found the writer had no reader anywhere (#270): each run recorded argv, cwd, env names, exit code, duration, output digest, and artifact paths that nothing surfaced. The store stays as it is; the command is the missing read surface, oldest first, verbatim under `--json`, empty state on an unknown root. Contract: `tests/contracts/trace-code-steps.test.ts`.
10
+ - `npm run smoke:real-home` (`scripts/smoke-real-home.sh`) boots the built binary against a copy of the operator's own settings. It copies `~/.config/clio-coder/settings.yaml` (and `credentials.yaml` when present) into a scratch `CLIO_CODER_HOME`, runs `doctor` and one read-only headless turn there, fails on a doctor crash or deprecation warning, a turn that exits non-zero, a tool policy drift refusal, or a turn with no `agent_end` event, prints doctor's own failing rows for the operator, and deletes the scratch home. The 0.4.2 release smoke ran only with isolated homes and missed two bugs a real settings file exposed on first launch (the raised read cap refusing to boot, and the legacy skill metadata warning), both fixed in this release. The release checklist names it beside the isolated-home turn.
11
+ - The files pane is a Clio Coder surface with its own verb and key. `/files` toggles a file view docked below the session and moves the keyboard into it; the same command or `Alt+E` (`clio-coder.files.toggle`, with the `Ctrl+G e` leader fallback) closes it and hands the keyboard back. `/files pick` and `/panes open files --once` borrow the pane for one selection. A pick lands in the composer as an `@file` mention and returns focus to the prompt; a pane herdr reports gone counts as closed, so the next toggle opens rather than trying to close nothing; reopening an open pane leaves any zoom that hides it and focuses it. Outside a pane host `/files` still runs the one-shot full-screen pick. The `panes` tool, `/panes open`, and the keybinding all reach one shared controller (`PanesOperations.files`). Verified end to end in a herdr 0.8.2 session on Linux x64: open, pick with `Ctrl+Y`, toggle close, key toggle, one-shot pick, and a clean `/quit` with the dock gone. Contract: `tests/contracts/panes-files.test.ts`.
12
+ - `clio-coder panes theme` prints Clio's theme tokens as a herdr `[theme.custom]` block. herdr styles its own chrome from its config file and offers no per-pane styling on the socket, so Clio prints the block for the operator to paste rather than editing another program's configuration. The files pane itself is themed from the same tokens at open time, now across the engine's manager, mode, status, which-key, pick, input, completion, task, help, and notification surfaces instead of five keys, with every emitted document parsed as TOML in the contract test.
13
+ - `docs/guide/panes-and-files.md` is the operator page for panes and the files pane from a clean machine: install, the two settings, what `doctor` prints at each stage, every command and key, what `/quit` closes, the theme boundary, and the exact messages with their remedies. Every quoted output is from a run of the built binary in a herdr session on this release; `docs/html/panes_files_blueprint.html` is its visual counterpart and the parity check holds both to the same command facts.
14
+ - Installed extensions are now projected into one immutable, generation-numbered snapshot per session. The extensions domain publishes nothing when it starts; the composition root builds and validates boot generation 1 and its user-hook registrations, then publishes both with adjacent reference assignments. Every in-process consumer of extension skill, prompt, agent, and fleet roots and of extension `hooks.yaml` declarations reads that paired generation instead of re-listing and re-hashing every installed tree on each load. `/resources extensions reload` is the only in-session trigger and uses the same paired publication path, so a turn observes either the previous resources with the previous hooks or the new resources with the new hooks and never a mixture. A rejected build, a stale candidate, or a re-entrant call publishes neither side, leaves the active generation untouched, and reports bounded diagnostics. Prompt fragments, cached agent recipes, and the session prompt refresh only when the committed content digest changed.
15
+ - Extension `hooks.yaml` declarations are now registered from the exact bytes hashed during install-digest verification and captured into the snapshot. A `hooks.yaml` rewritten after verification is never reopened; the tree fails verification on the next generation and contributes no hooks. Hook sources and hook receipts carry a structured `extension` object with the package provenance (id, scope, canonical root, manifest digest, content digest), the declarations digest, and the admitting generation, replacing the loose `extensionScope` and `installedContentDigest` fields. `hook-receipts.json` stays version 1 and older records still parse.
16
+ - Middleware registrations declared by user hooks are owned as one unit under a strictly increasing generation. A replacement for an older or equal generation is refused, a late disposer from a superseded generation is a no-op, builtin and host registration ids are never taken by an owned set, and an in-flight asynchronous hook evaluation finishes against the registration list it started with. Collisions surface as a `registration_conflict` middleware diagnostic and interactive notice.
17
+ - `docs/configuration-reference.md` is the audited inventory of every switch that changes what Clio Coder does at the 0.4.2 cut: 158 settings keys, 102 environment variables, 310 CLI flags across 33 commands, 60 `.clio-coder/` file keys, 22 recipe and 4 fragment frontmatter keys, 97 tool arguments on `dispatch`, `bash`, `context`, and `verify`, and 77 model knowledge-base tags, each with its default, what it controls, and what beats what when several surfaces set the same value. The audit behind it found 168 of them documented nowhere, 4 read by nothing, 13 with a second spelling, and 1 documented default that disagreed with the code; the entries below apply the result. Bounding constants are listed as code-owned invariants, not configuration.
18
+ - The documentation tree is organized by audience (`docs/guide/` for operators, `docs/architecture/` for contributors, `docs/process/` for how the project is run, `docs/history/` for dated records) with `docs/README.md` as the hub, and every one of the 51 Markdown pages has one HTML blueprint under `docs/html/` served by `clio-coder docs`. A source-first audit pinned to `ff56ea3e` found 42 pages with at least one claim that disagreed with the code (defaults, flag names, removed knobs, file paths) and corrected all of them; the blueprints were then synchronized from the corrected Markdown, and `docs/process/documentation-coverage.md` records the per-page result. The package ships the whole Markdown tree (never `docs/html`), `context(scope=docs)` walks the subdirectories, and the `docs-parity` hygiene check fails lint when a page and its blueprint disagree on the commands and environment variables they name or when either side is missing.
19
+ - Local model knowledge-base families and the quirks schema now declare static and thinking-level keyed chat-template kwargs (#267). Runtimes forward the merged kwargs object, preserving numeric values as numbers and letting existing thinking-control switches win on key collisions. NVIDIA Nemotron 3.5 Lightning declares static `force_nonempty_content`, Nemotron 3 Nano Omni declares numeric `reasoning_budget`, and Muse Glimmer declares level-keyed `reasoning_strength` while documenting LM Studio as unsupported.
20
+
21
+ ### Changed
22
+ - `/quit` in a pane host now says what it left behind. The policy is decided (#272): docks (the files pane, the workers watch pane) close with the session, utility panes opened with `/panes open shell|logs|<argv>` stay, and `/quit` prints one line after the terminal is restored naming each pane it left, the `/panes close all` that would have taken them along, and the `herdr pane close` that closes them now. Nothing is printed when only docks were open. Contract: `tests/contracts/panes-quit-left-behind.test.ts`.
23
+ - Tool results now use a configurable `context.toolResultMaxBytes` ceiling with a 65,536-byte default and a 4,096-byte minimum. Session overrides and configuration reloads apply to the next result, while overflow still preserves the complete text in the session scratch file and the separate 196,608-byte turn guardrail remains unchanged.
24
+ - Chat output now defaults to each model's advertised maximum instead of a fixed 32,768-token ceiling. Models without an advertised cap retain a bounded 32,768-token fallback, explicit settings still win, and the Qwen3.6 27B and 35B catalog caps now match the vendor's documented 81,920-token maximum.
25
+ - The files pane no longer names its engine anywhere an operator reads. The preset is `files` (`/panes open files`; `yazi` still parses as an alias so saved habits and scripts keep working), the pane label is `files`, the composer notices say "the files pane", the doctor rows read `files pane profile` and `naming files profile`, the Settings labels read Files pane, and the help overlay gains a "panes & files" topic. The engine's name survives exactly where it is a fact the operator acts on: the registry id in `clio-coder tools install yazi`, `tools status yazi`, and the `interface.panes.files.*` settings the reference already used. The model's `panes` tool enum carries `files`, `logs`, `shell`.
26
+ - `/panes open logs` and `/panes open shell` say why the pane layer is missing and how to get it, repeating detection's reason (`HERDR_ENV is not 1, so Clio is not running inside a pane host`) and the `clio-coder --with-panes` remedy, where they previously said only that the layer was unavailable. An empty logs preset names the journal root it watched instead of "has nothing to show yet". A second `/panes open shell` or `logs` focuses the pane that is already open rather than splitting a second one, and reports `focused pane` so the difference is visible.
27
+ - Local model knowledge now projects only engine-consumed sampling and thinking quirks. Removed structured KV cache recommendations and unused per-profile output caps; KV serving facts remain in `llamaCpp`, `measuredUnder`, and serving notes, while `capabilities.maxTokens` remains the authoritative output cap. No request default changes. Catalog authors should use those surviving provenance and capability fields instead of the five removed no-op tags.
28
+ - Model residency now resolves only from `targets[].lifecycle`. Removed the duplicate `CLIO_CODER_RESIDENCY` process switch. Local targets remain managed by default, and SSH nodes remain observe by default through WorkerSpec lifecycle projection. Operators should replace process overrides with `targets[].lifecycle` or `fleet.nodes[].residency`; an explicit target `user-managed` opt-out remains authoritative.
29
+ - Smooth-stream pacing and third-party project resource trust now resolve only from `interface.smoothStreaming` and `integrations.projectResources.trustProjectImports`. Removed both duplicate process environment overrides. `off` remains the smooth-stream default, and `false` remains the project-import trust default. Operators should replace environment overrides with those settings keys.
30
+ - Guardrail limits and the run event journal now use only their version 2 settings paths under `safety.limits` and `fleet`. Removed the seven duplicate environment overrides, and the existing settings defaults remain in force. Operators should replace environment overrides with the corresponding `safety.limits`, `fleet.limits`, and `fleet.history` keys. Session-only changes and configuration hot reload now refresh the process-local projections, and the orchestrator reads the current turn budget on each attempt.
31
+ - Removed the deprecated `CLIO_CODER_MAX_RUNS` and `CLIO_CODER_TRUST_PROJECT_SKILLS` environment spellings, the pre-fleet `--worker*` command aliases, the headless `run --runtime` alias, and eval's `--clio-entry` alias. No default behavior changes. Operators should use `fleet.history.maxRuns`, `integrations.projectResources.trustProjectImports`, and the canonical flags named by each command's help; old spellings now reject or fall through without compatibility warnings.
32
+ - The `dispatch` tool schema serializes its `intent` and `budget` object schemas once, under `$defs`, and references them by JSON pointer from the top level and from every task; a task object now carries only what varies per task (`task`, `briefing`, `agent`, `budget`, `target`, `model`, `node`, `worktree`, `intent`, `gate`), with `persona`, `tool_profile`, `cwd`, and `apply` inherited from the batch defaults (an item that still sends one is honored). The schema went from 9,863 to 7,959 characters; under Qwen3.8-27B on dynamo the `dispatch` tool went from 2,884 to 2,357 tokens and the full-capability first turn from 9,501 to 8,923, and on the Ornith 1.5 tokenizer the schema went from 2,190 to 1,782 tokens with the per-task object alone from 635 to 201. TypeBox's validator and pi's tool-argument validator both resolve the pointers, and `tests/contracts/dispatch-admission.test.ts` still admits a batch whose tasks declare their own intent and budget.
33
+ - Resource loads no longer re-hash installed extension trees on every skills, prompts, agents, or fleets read. An extension installed, enabled, disabled, or removed by another process while a session runs is not visible to that session until `/resources extensions reload` or a restart; previously it could appear on the next load without any explicit action. `clio-coder config inspect` reports hook and extension entries with `reload:reload` instead of `reload:restart`.
34
+ - Updated the Pi engine SDK libraries (`pi-ai`, `pi-agent-core`, `pi-tui`) from 0.84.0 to 0.84.4. The agent loop now runs Clio's post-tool continuation guard only when another assistant turn is about to start, so a terminating tool batch no longer triggers the guard after its final result; Clio's guard was already continuation-only and needed no change. Inherited fixes include OpenAI-compatible reasoning replay and thinking-signature serialization, Anthropic server-side refusal fallback pricing, optional tool arguments sent as `null` being treated as omitted, fullscreen transcript search and half-page/line scrolling actions, the SSH-aware `PI_TUI_ESC_TIMEOUT` escape window, and lower alternate-screen per-frame allocation. While the fullscreen search overlay is focused, `ctrl+g` advances the match and the Clio leader chord is unavailable; Clio's own bindings do not otherwise overlap the new defaults. New engine lifecycle contracts lock these behaviors.
35
+ - The delegation threshold now opens the Delegation section as a count taken before the first edit, with the dispatch call shape beside it: two or more independent file-scoped changes, or any repository-wide exploration however small the repository looks, means one `dispatch` call with `tasks` (one per change, `agent` coder, `mode` parallel, `intent` naming each file) or a `scout` dispatch before any repo-wide grep or read, and the parent keeps synthesis and validation. The Fleet block no longer carries the threshold. On the round-2 drive with Qwen3.8-27B on LM Studio at temperature 0 the bare Fleet-block threshold lost to inertia on every run (two-changes 0 of 2 dispatched, reconnaissance 0 of 2 dispatched scout); with the count up front and the call shape next to it, two-changes dispatched both workers in one parallel call on 6 of 6 runs and reconnaissance dispatched scout before any repo-wide read on 5 of 6, the miss being the run before "however small the repository looks" was added. On the final build the parent stayed off both assigned files on 2 of 2 two-changes runs, spot-checked the receipts, and ran the tests itself. Re-measured after the merge with v0.4.2's prompt hardening and the `dispatch` schema slim: two-changes 2 of 2 dispatched with the parent off both assigned files, reconnaissance 2 of 2 scout-first, inventory and docs-first 2 of 2, bad-argument recovery 2 of 2 without a same-shape retry, and the skill suggestion fired on 2 of 2 runs at 8,995 first-turn tokens.
36
+ - `ledger` is off the session tool surface and on every worker registry. The session never binds an agent-ledger port, so on that surface the tool could only answer "no ledger" while its schema cost 444 tokens of every first turn; `registerCoreTools` takes `includeLedgerTools`, the session leaves it off, and the worker registry sets it whether or not a port is bound, because the orchestrator admits `ledger` for every batch member and signs that surface. An interim build that gated the worker side on the port refused every batch worker with `Worker attestation rejected: tool surface drift`; `tests/contracts/worker-attestation-surface.test.ts` pins both sides.
37
+ - The compiled prompt reads the tool surface from the frozen name list only. `toolSurfaceHasTool` in `src/domains/prompts/compiler.ts` counted a registry hint as evidence that its tool was attached, so a stale hint could render the Skills or Delegation passage for a tool the model could not call; hints now render only for tools on that list and the surface gates read the same list.
38
+ - The `context(scope="skills")` listing opens with a label instead of a second copy of the suggest-and-continue protocol; the recency anchor at the bottom, the line literal models act on, is unchanged. Every catalog skill's description is now one lead sentence saying what the skill does plus its "Not for X; use Y" routing clauses, with the trigger phrases kept in `triggers`: the 31 descriptions went from 13,116 to 7,741 characters in the listing. `skills/README.md` and `skill-craft` state the split, and every catalog version took the minor bump the versioning policy requires for a description change.
39
+ - Recipes name `code_nav` only as "when it is among your tools": routine non-Scout dispatch removes the tool, so "when a codewiki exists, prefer `code_nav`" invited a call the worker could not make.
40
+ - Three dispatch rejections that cost a live orchestrator a round each now either read the intent or say what to do. `intent.read_roots` or `write_roots` entries of `.` or `./` name the repository root, which an empty scope already means, so they are read as that instead of refused; a briefing that quotes an import specifier such as `./parser.js` infers `parser.js` instead of failing the whole dispatch for a dot segment (a `..` token is still refused as an escape); and the `verification_check_undeclared` and empty-`gate` errors say what an entry is (a check id `verify()` lists, never a shell command) and offer the move that always works, omitting it and running the check after the receipt. The orchestrator that put `git diff src/formatter.ts` in `verification` had retried the same shape until the loop guard disabled tool calls for the turn.
41
+ - The model knowledge base gains an `ornith-1.5` family with an on-off thinking mechanism. The inherited `ornith` entry classified every Ornith checkpoint as always-on from a 2026-08-08 LM Studio measurement of Ornith 1.0, so a worker dispatched with thinking off still produced a full reasoning trace (1,061 of 1,518 output tokens on one debugger receipt). Measured 2026-09-02 on llama.cpp: `enable_thinking: false` answered in 4 completion tokens with no `reasoning_content`, `true` spent 79.
42
+ - Seven more ids the mini llama.cpp router serves now resolve to a family whose thinking mechanism was measured on that runtime. Six had no family at all, and what that cost depended on how the router launched them. `gemma4-26b-moe` and `gemma4-31b-dense` run under `--reasoning off` and `muse-30b-dense` under no `--reasoning` flag at all, so the probe reported no reasoning, the resolver returned mechanism `none`, pinned the dial to off, and stripped every thinking field from the payload. `nemo3.5-30b-moe`, `nemotron3-30b-moe-omni` and `thinkingcap-27b-dense-q4` run under `--reasoning on`, so the probe reported reasoning and the resolver already guessed `on-off` and already put `enable_thinking` on the wire; what they gain here is a measured classification instead of a guess, plus sampling defaults, guidance and provenance. `qwopus3.6-35b-moe` matched the `qwopus3.6-27b-v1-preview` distill through a bare `qwopus3.6` pattern and inherited a budget-tokens dial that emits nothing on llama.cpp. Measured 2026-09-02 on a one-line arithmetic prompt, completion tokens at `chat_template_kwargs.enable_thinking` false against true: `gemma4-26b-moe` 3/99, `gemma4-31b-dense` 3/68, `nemo3.5-30b-moe` 3/120, `nemotron3-30b-moe-omni` 4/42, `thinkingcap-27b-dense-q4` 3/130, `qwopus3.6-35b-moe` 3/130, every one of them on-off; `muse-30b-dense` gave 91 tokens with 332 characters of reasoning at false against 90 with 312 at true, so it is always-on and its strength scale has no off level to reach. `reasoning_effort` moved none of the seven on llama.cpp, so `nemotron-3-nano-omni-30b-a3b-reasoning` drops its budget-tokens dial and `qwopus3.6-35b-a3b-coder` drops its four-level effort map, both for on-off, and both keep their earlier LM Studio findings in `measuredUnder`. New families cover NVIDIA Nemotron 3.5 Lightning, Meta Muse Glimmer 30B, and BottleCapAI ThinkingCap-Qwen3.6-27B. The bare `qwopus3.6` pattern is deleted rather than reordered, because the matcher ranks by pattern length and ignores file order. The `thinkingcap-27b-dense-q6` the issue names was measured on the same date and then deleted from mini, so the family is keyed to the q4 the router still serves (`tests/contracts/local-model-family-resolution.test.ts`).
43
+ - The Gemma channel filter is keyed on the wire model id rather than on the resolved family, so it survives a Gemma id gaining a catalog entry. `capabilityFamily` returns the knowledge-base family whenever the catalog matches an id, and every Gemma entry names a build (`gemma4-26b-a4b`, `gemma-4-31b-it-qat-mtp`), never the literal `gemma-4` the gate compared against, so the filter had been running only for Gemma ids the catalog did not know. Naming `gemma4-26b-moe` and `gemma4-31b-dense` in the catalog would have turned it off for the two the mini router serves, and a llama-server that leaves the thought inline in `content` would have put `<|channel>own-think` and the private thought into visible assistant text and into turn history.
44
+ - The prompt-optimization recording proxy keeps llama.cpp's `timings` block (`prompt_n`, `cache_n`, `prompt_ms`) beside each captured request and keeps the tail of a long stream rather than its head, and `envelope-run.ts` takes `--runtime` so a llama.cpp target can be measured. Development instruments only; nothing under `scripts/` ships.
45
+ - The first turn of a full-capability session on this repository measured 12,950 input tokens under Qwen3.8-27B before this change, 69 percent of it tool schemas and 26 percent the `dispatch` schema alone, with the same rules restated across the Delegation fragment, the Fleet block, the Tool Contract, and the `dispatch` and `monitor` descriptions. The prompt now says each rule once: the Delegation fragment keeps receipt and spot-check discipline, the Fleet block keeps routing (including that `agent:"auto"` is a fallback, which no longer also appears in the Tool Contract), the Skills fragment is a short suggest-and-continue pointer since the first-turn `[Skills]` reminder is the channel local models act on, the identity block drops the biography and keeps the name, the vendor denial, the safety line, and the installed paths, and the operating contract drops the posture-toggle and report-artifact sentences. Every `dispatch` field keeps one discriminating sentence and the tool description states only the call shape; `monitor`, `steer`, `panes`, `ask_user`, `verify`, `bash`, `credential_present`, `web_fetch`, and `ledger` lose restated policy from their schemas. Recipe descriptions now lead with a job sentence that fits the 64-character roster line, so the Fleet section no longer ends nine of eleven lines in `...`.
46
+ - Tool schemas reach the provider without TypeBox 1.x's string-keyed markers. `Type.Unsafe` and `Type.Optional` stamp `"~unsafe": null` and `"~optional": true` on every enum field, and unlike the older symbol keys those survive `JSON.stringify`, so 44 such keys rode along on the default surface. `wireParameterSchema` in `src/tools/agent-tools.ts` hands the agent loop a copy with every `~`-prefixed key removed; `Value.Check` validates identically with and without them and the registry keeps the original schema object.
47
+ - Rebuilt the product README and operator documentation around the current CLI, settings-v2 contract, package layout, and experimental boundaries so new users and repository agents have one concise, accurate route into the product.
48
+ - Compacted the main and worker system prompts by removing duplicated routing, skill, retrieval, and worker-task prose while preserving Clio identity, autonomy and safety rules, Fleet receipts and evidence, permission routing, local-model guidance, result contracts, and stable section order.
49
+ - The `dispatch` tool schema is composed once per session from the fleet the session starts with (`src/tools/dispatch-schema.ts`). The council fields (`roster`, `members`, `synthesis`, `rounds`) are advertised only when `fleet.rosters` names a roster, the compete fields (`candidates`, `judge`, `apply_winner`) only when the fleet has more than one distinct target/model route, and the adaptive-routing fields (`routing.posture`, `minimumQuality`, `locality`, `failover`) only when `fleet.adaptiveRouting` activates a role or posture; the hard bounds (`maxCostUsd`, `deadlineMs`, `requiredCapabilities`) and the `mode` enum entries for the advertised modes stay. Admission reads every field whether or not it was advertised, so a caller that sends a hidden one is honored. `intent.write_roots` and `intent.expected_outputs` carry one-sentence descriptions saying they are paths under a write root, not prose: the interim two-changes batch (r4-mid) paid 3 and 1 dispatch rejections (`intent_outputs_outside_write_roots`, `verification_check_undeclared`) before its parallel call was admitted; the final batch (r4-final) dispatched both tasks in one parallel call with per-task intent on the first try in 2/2 runs, and the parent edited neither assigned file. On the one-route, no-roster sandbox under Qwen3.8-27B on dynamo the `dispatch` tool went from 2,357 to 1,800 tokens (1,766 before the two descriptions) and the full-capability first turn from 8,924 to 8,365 (r4-final; the same tree measured 8,323 on mini under Ornith 1.5); `tests/contracts/dispatch-schema.test.ts` pins the composition and the open schema.
50
+ - The `coder` recipe says to make each read and verification call once. Round-3 receipts on Ornith 1.5 showed every coder run blocked five times by the identical-call loop guard (repeated `code_nav` and `git diff`) while using 17 to 30 of its 50-call budget; the final batch's four coder runs used 3, 5, 9, and 17 calls with 0, 0, 1, and 2 blocks (r4-final against 30/5, 19/5, 27/5, 17/5 in r3-final). Budgets and result contracts stay as they were: no builtin recipe reached its budget or fired contract repair in any measured run, so the guard, not the budget, was the binding knob. The worker prompt pin moves from 5,179 to 5,335 characters.
51
+ - Thinking "off" now reaches LM Studio for effort-level families (qwen3.8 and its relatives) as `reasoning_effort: "none"`. LM Studio ignores `chat_template_kwargs.enable_thinking` for these models, so every session configured with `chat.thinkingLevel: off` on that runtime had been reasoning anyway (37k reasoning tokens in one 25-minute interactive turn), and every round-2 to round-4 main-agent measurement on dynamo ran with reasoning on. Measured on a one-line prompt: `enable_thinking:false` left 63 reasoning tokens, identical to no override; `reasoning_effort:"none"` produced 0 and a shorter rendered prompt. llama.cpp keeps the template flag alone, which it honors. The `chat.thinkingLevel` default moves from `off` to `low` in the same release (see the next entry), so a fresh home keeps the reasoning it had been getting by accident (`tests/contracts/thinking-off-wire.test.ts`).
52
+ - `chat.thinkingLevel` defaults to `low` instead of `off`. Every measurement before the wire fix ran with reasoning on, so `off` had never actually been the shipped behavior on LM Studio. With reasoning truly off, the same three-issue interactive exercise on Qwen3.8-27B re-emitted one batch of eleven read-only calls five times and finished nothing in its first run, and even with the loop guard fixed the model plans visibly worse than it does with a low effort budget. `fleet.default.thinkingLevel` stays `off`: workers run bounded recipes where the round-4 receipts showed no benefit. Existing settings files that name a level are unaffected.
53
+ - The identical-call loop detector retains the last 48 attempts instead of a 30-second window or the last four, and the loop guard counts repeats from the turn's last successful write or edit. A session with thinking off re-emitted the same batch of eleven read-only bash calls five times, every call admitted, and ran to the 60-call soft budget with nothing done; a batch re-emitted a third time is now blocked on its first call, while a check rerun after an edit stays a fresh call (`tests/contracts/loop-detector.test.ts`, `tests/contracts/loop-guard-epoch.test.ts`).
54
+ - A worker confined to `write_roots` is told so: the Declared Result Requirements block says the run has no bash or verify tool, that the host runs the declared checks after it finishes, and to write the code and tests and report; the `write_roots` schema description and the Delegation section say the same to the coordinator, which now puts checks in `verification` instead of the task text. Before this, three parallel coders on Ornith 1.5 told to "run the tests" spent 40, 23, and 15 `code_nav` calls looking for a way to run them and made zero source edits; with the sentence they finished in 10, 18, and 12 calls with every file edited (`tests/contracts/intent-requirements.test.ts`).
55
+ - A typed intent that declares no `write_roots` and no `expected_outputs` outranks the prose task classifier for a read-only recipe, so a scout survey that mentions writing a failing test is admitted instead of refused as a change-class task. An intent with a write root still refuses (`tests/contracts/dispatch-admission.test.ts`).
56
+ - The `[dispatch scope]` notice names only tokens with a source or document extension or a trailing separator, capped at twelve; a live three-task dispatch had printed 27 omitted "paths" per task, most of them `0.5`, `4/10`, `v24.9`, `e.g`, and method names like `Store.fromText`.
57
+ - `tasks done` on a task that was never started records the start and the completion together and names the implicit start in its notes; the evidence note stays mandatory. A live session had spent six start/done pairs of pure ceremony closing a board whose work was already done (`tests/contracts/task-board-done.test.ts`).
58
+ - The chat panel renders a `Suggested skill: /skill <name>` line wherever the model wrote it as the suggestion row, with the rest of the message as the answer. Round-4 skill batches fired the line 2/2 per batch but opened the reply with it 1/2, 1/2, and 0/2 across three wordings, so the harness recognizes it instead (`tests/contracts/rendering-invariants.test.ts`).
59
+ - The compact footer row names the live context window after the percent (`ctx ▰▰▱▱ 4.6% of 262.1k`) at 72 columns and wider, so a session on a 1M window and one on 128k no longer read the same. The window comes from the context ledger when one is bound and from the target capabilities otherwise; an unknown window shows the percent alone (`tests/contracts/footer-context-window.test.ts`).
60
+ - Self-awareness names the live settings file and state directory of the resolved home rather than the XDG defaults, and the tool-inventory line says `dispatch(list:true)` answers which target and model run the session and its workers. Asked which worker would run a coder dispatch and where that is configured, the first exercise run answered from the recipe alone; the third named the fleet default target and model from the isolated home's settings.yaml. The `artifact` description says to close the task board before writing, because the artifact ends the turn. Main prompt 10,313 -> 10,504 chars (`tests/contracts/compact-prompt-contracts.test.ts`).
61
+
62
+ ### Fixed
63
+ - The boot deprecation for a skill file that still carries `clio:` frontmatter names the file, and `clio-coder doctor` finds it. The warning read `'clio: skill metadata' is a deprecated Clio Coder identifier` with no path, while the doctor's `naming resources` row said installed skills were canonical, because `inspectSkillMetadataNaming` scanned only `<config>/skills` and `.clio-coder/skills` and the loader also reads every interop agent's user and project compatibility root (`.claude/skills`, `.codex/skills`, `.agents/skills`, and the rest). The operator's copy was a gitignored `.claude/skills/file-ticket/SKILL.md` in the project. The warning now reads `clio: skill metadata in <path>`, once per file, and the doctor scans the loader's roots and lists each offending path with the rename to make (`tests/contracts/settings-migration.test.ts`).
64
+ - The engine reads runtime metadata from `model.clioCoder` again. The 2026-09-01 naming migration moved the key the synthesizer writes from `model.clio` to `model.clioCoder` (`src/domains/providers/runtimes/common/local-synth.ts`) and updated the providers domain, but `src/engine/apis/openai-completions.ts`, `src/engine/apis/lmstudio.ts`, `src/engine/apis/ollama-native.ts`, and the prompt compiler in `src/interactive/turn-context.ts` kept reading `clio`, so for every model built since the migration the engine saw no metadata at all: LM Studio targets kept receiving the `chat_template_kwargs` map the runtime ignores and never ran residency, TTL, draft-model or advertised-effort handling; family sampling profiles never reached the request; llama.cpp requests carried no `cache_prompt` and skipped residency; llama.cpp and LM Studio backend timings were not attributed; Ollama residency and quirks did nothing; and the family's thinking guidance never rendered into the Runtime prompt block. Reproduced through the OpenAI-compatible fixture with a model shaped exactly as the synthesizer shapes it, and the contract suite had not caught it because the one engine test that builds such a model had been updated to the new key without the reader following. `tests/contracts/engine-model-metadata-key.test.ts` pins the LM Studio wire (no `chat_template_kwargs`, `reasoning_effort` from the dial), the sampling profile on the request, and llama.cpp's `cache_prompt`.
65
+ - Thinking off reaches LM Studio as `reasoning_effort: "none"` for every mechanism, not only effort-level families (#268). Measured 2026-09-02 on dynamo at temperature 0 on a one-line prompt: `chat_template_kwargs.enable_thinking: false` left 26 reasoning tokens on qwen3.8-27b, 52 on gemma-4-26b-a4b-it and 105 on nvidia-nemotron-3.5-lightning-30b-a3b, while `reasoning_effort: "none"` produced 0 on each. The resolver's on-off branch already carried the spelling; the budget-tokens branch did not, and the wire only reached it through the LM Studio payload composer that the metadata-key regression above had switched off. `resolveRequestCapability` now sets `none` for any inactive mechanism other than `none` and `always-on` on that runtime, and llama.cpp keeps the template flag alone. A family's own chat-template kwargs (#267) are no longer merged into a request the runtime drops: on LM Studio the resolver lists them under `request.undeliverableChatTemplateKwargs` with whether the family entry marks them `lmstudio: unsupported`, and runtime resolution prints one `chat-template-kwargs-undeliverable` warning per target and model, so a Nemotron 3.5 Lightning session on dynamo says it runs without `force_nonempty_content` instead of finding the key missing from the wire. `tests/contracts/thinking-off-wire.test.ts` covers on-off, budget-tokens and the shipped Gemma 4 and Nemotron 3.5 entries on both runtimes; `tests/contracts/chat-template-kwargs-diagnostic.test.ts` drives the warning through `resolveRuntimeTarget`.
66
+ - gpt-oss reasoning from LM Studio lands in thinking content (#269). LM Studio streams the Harmony analysis channel as the OpenAI-compatible `reasoning` delta field, with `reasoning_content` absent (measured 2026-09-02 on dynamo, gpt-oss-20b and gpt-oss-120b). The OpenAI-completions stream reader already accepts `reasoning_content`, `reasoning`, and `reasoning_text`, so the text was captured; what was missing was a contract holding that and a family entry naming the field per runtime. The `openai-gpt-oss` entry now states `reasoning_content` on llama.cpp and `reasoning` on LM Studio, and `tests/contracts/lmstudio-reasoning-field.test.ts` streams all three spellings through the engine against the fixture and checks the text stays out of the answer and the Harmony effort goes out as `reasoning_effort` with no `chat_template_kwargs`.
67
+ - `interface.panes.enabled: embedded` no longer turns every pane off. The settings overlay offered `embedded` beside `auto` and `off`, and picking it made boot print "embedded mode is not implemented yet ... this session has no panes at all" while the same operator inside herdr would have had guest panes under `auto`. Embedded mode is still unimplemented; the rung now runs the guest detection ladder and logs that it did, so the setting an operator chose because they wanted panes gives them panes (`tests/contracts/pane-remedies.test.ts`).
68
+ - A settings file that raises `safety.limits.readBytesPerCall` above the 50 KiB default no longer refuses to start. The read tool's result-size policy cap was computed once at module import from the guardrail default, while the drift check at registration compared it against the cap installed from settings, so `clio-coder` exited with `tool policy drift: tool read policy cap 53248B sits below self cap + slack (67584B)` before the first prompt. The cap is now recomputed when the tool is registered (`tests/contracts/read-policy-cap-follows-guardrail.test.ts`).
69
+ - `docs/architecture/artifact-versions.md` gives `<stateDir>/runs.json` (the dispatch run ledger, `RunEnvelope` in `src/domains/dispatch/state.ts`) its own registry row instead of only naming it as deliberately excluded. It has at least nine reader call sites across the eval, evidence, CLI, and TUI domains and genuinely has no version field on `RunEnvelope` to check, so the row follows this document's own existing pattern for a real but unversioned artifact (`Current Version: unversioned JSON array`, matching the Out-of-turn Usage Ledger and Library Pins rows already there) rather than inventing a version number that does not exist. Also corrected, in the same paragraph: the doc named a `detached-batch` store it deliberately excludes from its registry, but the actual on-disk file is `batches.json` (`src/domains/dispatch/batch-store.ts:55`); the sentence now names the real filename. Mirrored both changes into `docs/html/artifact_versions_blueprint.html`. Docs-only; no code or behavior changed, and `check-hygiene.ts`'s `docs-parity` check passes.
70
+ - The trace database's `runs` table now carries an explicit `source` column (`'dispatch'` or `'session'`) instead of leaving a dispatch run and an interactive session turn indistinguishable except through the undocumented sentinel `assignment_id = "session"`. `docs/architecture/trace-store.md` named that sentinel but never called it a discriminator a reader should rely on. The column is additive, the same way `processes.host`/`processes.birth_token` were added before it: `TraceStore` (the writer) backfills it in place from the sentinel on next open, with no schema-version bump, matching this database's existing migration precedent (`ensureProcessOwnerColumns`). Because `TraceReader` and the trace-viewer's `ViewerDatabase` both open read-only and can reach a database no writer has touched yet (the viewer opened fresh against a dormant install being the realistic case), both now detect a missing column via `PRAGMA table_info` and derive the same value from the sentinel at query time rather than showing an absent field. `clio-coder trace runs`' text-mode table gains a SOURCE column; its `--json` output and the trace-viewer API carried the field automatically once the column exists, since both already `SELECT *`. `tests/contracts/trace-store-run-source.test.ts` and `apps/trace-viewer/tests/server.test.mjs` cover the read-only-before-migration, writer-then-reader, and fresh-database cases.
71
+ - The observability projection's run summary and the interactive dispatch board no longer maintain two copies of the same DispatchCompleted/DispatchFailed field mapping. Both independently subscribed to the same three bus channels and built their own summary; `resolveFailedStatus`, mapping a failure `reason` to a terminal status, existed as two byte-for-byte identical function bodies in `src/domains/observability/projection.ts` and `src/interactive/dispatch-board.ts`, kept in sync only by a comment asking the next editor to remember to. Re-verified the two sides were not actually a wholesale duplicate before touching anything: the dispatch board's `DispatchBoardRow` carries dozens of fields (budget envelope, gate/council state, trust projection, context-window meter, retry and failover tracking) the observability summary has no concept of, driven by additional bus channels (`DispatchEnqueued`, `RunAborted`, assignment/attempt events) the summary never reads, so making the board read the projection instead of its own subscription, the audit's proposed direction, would have thrown away real state; consolidation was scoped to the two fields the discrepancy actually named. The status mapping is now `resolveDispatchFailureStatus` in the new `src/core/dispatch-outcome.ts`, imported by both. Cost provenance had a real, live divergence, not just a duplication risk: the projection validated a payload's `costProvenance` against the closed four-value set and kept the previous value otherwise, while the board accepted any truthy string via `payload.costProvenance ?? "unknown"`; both now call the new `resolveCostProvenance` in `src/domains/providers/types/cost-provenance.ts`, which applies the stricter, correct behavior on both sides. Fixing `applyTerminalTokens` to use the projection's own existing `num()` helper (already used elsewhere in the same file, just not here) for `costUsd`/`tokenCount`/`inputTokenCount`/`outputTokenCount`/`reasoningTokenCount` closes a related gap in the same code: those fields previously passed a bare `typeof === "number"` check that accepted `NaN`/`Infinity`, unlike the board's already-`Number.isFinite`-checked equivalent, so a malformed terminal payload could silently corrupt the projection's running total while the board correctly rejected it. `tests/contracts/dispatch-outcome-provenance.test.ts` pins the two shared resolvers directly; `tests/contracts/observability-run-summary.test.ts` drives the projection through a real bus end to end, including the NaN/Infinity and out-of-set-provenance cases.
72
+ - The observability projection's notice ring (`src/domains/observability/projection.ts`) no longer classifies seven notice kinds nothing ever read. `runtime`, `middleware`, `safety`, `loop`, `tool-budget`, `context`, and `budget` notices were built from the same bus channels the interactive layer's own live notice pipeline (`bus-notices.ts` and `interactive-event-projection.ts`) already classifies independently, richer and with real behavior attached (loop-guard and tool-budget notices can cancel the active turn, which the projection's pure classifier cannot do), for the toast surface a session actually shows; the projection's copies had no reader anywhere and could silently disagree with what the operator saw. Verified one kind is a genuine, live exception: `evidence`-kind notices (evidence-build failures) are read by the Dispatch Board's per-run evidence-failure-reason lookup and have no equivalent in the interactive notice pipeline, so that kind, `pushNotice`, and the notice ring itself are unchanged; `ObservabilityNotice.kind` narrows to `"evidence"` and `.ref` narrows to `{ runId }`, its only populated field. `tests/contracts/observability-notices.test.ts` pins that the seven removed channels now produce nothing and that evidence-build-failure notices still round-trip.
73
+ - `clio-coder config inspect` now reads the durable hook-execution receipt log it has claimed to read since the log's own doc comment was written. `hook-receipts.json` (`src/domains/middleware/hook-receipts.ts`, written on every user-defined hook execution) had zero readers anywhere in the repository, and `config inspect` only ever read hook *configuration* sources (`capturedHookSourcesFor`), never execution receipts; a doc comment at the top of the writer named `config inspect` as the reader regardless. The customization graph now carries one `hook-receipts` entry (count, capacity, an outcome tally, and the most recent receipt) sourced from `readPersistedHookReceipts`, a new reader that loads the persisted snapshot back for a process that never held the running session's in-memory ring. One rolled-up entry rather than one per receipt: the graph explains configuration provenance, and a couple hundred execution rows would swamp that rather than answer it. `tests/contracts/hook-receipts-inspect.test.ts` covers both the never-persisted and the populated case.
74
+ - `skills-eval` and `clio-coder eval run` no longer resolve to the same artifact file when handed the same eval id. Both `writeEvalArtifact` (version-1 `EvalRunArtifact`, `src/domains/eval/store.ts`) and `writeEvalArtifactV4` (version-4 `EvalArtifactV4`, `src/domains/eval/artifacts/store.ts`) computed `<dataDir>/evals/<evalId>.json`, so one writer could silently overwrite the other's differently-shaped file; `eval inventory`'s directory listing was also miscounting every on-disk `skills-eval` artifact as an unreadable retired shape, since it reads only version 4. The legacy version-1 writer now nests under `<dataDir>/evals/skills-eval/<evalId>.json`; the documented version-4 location (`docs/architecture/artifact-versions.md`'s Eval Artifact row) is unchanged. `tests/contracts/eval-artifact-namespacing.test.ts` writes both shapes under one shared id in both orders and reads both back intact.
75
+ - The files pane engine's per-session transport files (`.stream`, `.chooser`, `.cwd` under `<cache>/yazi/sessions/`) were never removed, so every open left three files behind for the life of the install. They are removed when the session ends, and anything older than a day is swept on the next open, which also cleans up installs that accumulated them before this release. The vendored engine configuration also loses a `title_format` key the pinned engine does not read.
76
+ - The llama.cpp probe no longer loads a model to read its slot count. A router answers `/props?model=<id>` by loading that model, and with `--models-max 1` that evicts whatever is resident, so every probe of a fleet target's default model unloaded the chat model and the next turn prefilled from zero (the round-3 resume run on mini prefilled 11,601 tokens with `cache_n` 0 even though the resumed request was byte-identical through the tool results). The router's model list already carries the worker's flags for an unloaded model, so the worker's props are read only when the router reports it resident; `tests/contracts/llamacpp-router-probe.test.ts` pins both cases. Re-measured on the same two-turn run: the resumed turn's first request prefills 90 new tokens against 11,525 cached in 0.66 s, and the turn's wall time went from 31 s to 5.9 s.
77
+ - A synthesis-locked worker round whose reply was tool-call markup only gets one re-prompt, delivered as a paired synthetic tool exchange, before the fallback notice stands as the run's whole output; when a result contract is active, that round costs the re-prompt rather than one of the contract's bounded repair slots. Locked rounds remove the tool surface on OpenAI-family runtimes and send `tool_choice: none` on Anthropic, whose API rejects a history carrying tool_use blocks when no tools are defined. Measured on mini with a two-call recipe that locks every run, the re-prompt recovered every Ornith 1.5 markup round and 2 of 4 on Qwen3.8-27B.
78
+ - Headless `run --autonomy <level>` takes effect. The flag was keyed by the bare word in the session overrides, which `setAtPath` wrote as a top-level key nothing reads, so a run in a home saved at auto-edit compiled "Autonomy: auto-edit" and parked a piped bash command for approval; the override is now keyed `safety.autonomy`, the path the effective view, the prompt compiler, and admission all read. `tests/smoke/cli-core.test.ts` runs the same home with and without the flag.
79
+ - A resumed session replays a parallel tool batch's results in the order the assistant issued the calls. The ledger records a result when its tool finishes while the live loop sends the batch by call index, so a two-call turn whose second call finished first replayed as [b, a] after being sent as [a, b], and the provider prefix cache missed from that message for the rest of the session (the round-2 two-turn run on mini swapped messages 3 and 4). Results are staged until the next non-result message and released in call order; an orphan keeps its ledger position after the known ones. The visible transcript still renders the ledger as recorded.
80
+ - `intent.verification: [{ check: "none" }]`, the shape a model writes to say "no verification", normalizes to an empty list instead of a refused dispatch round, unless the project declares a check by that name.
81
+ - A briefing that quotes an import specifier with a leading `../` no longer fails the whole dispatch (#266). Round-5 exercise run 3 lost a coder dispatch to `legacy_scope_path_malformed: briefing contains a malformed path token '../src/store.js'` after the orchestrator copied `import { Store } from "../src/store.js"` out of `test/store.test.ts`, then retried without the quote. The quote's origin never reaches prose inference, which is handed the task and briefing text and nothing else, so a leading `../` run is stripped and the remainder probed against the dispatch root: `../src/store.js` infers `src/store.js` when that file exists, and is otherwise dropped from inference instead of rejecting the call. That costs one `existsSync` per `../`-leading token (3 microseconds on a hit and 2 on a miss over 5,000 warm calls) and no stat for the ordinary token, which the grammar matches dozens of per dispatch. **A leading run that names nothing is dropped rather than refused**, including `../../../etc/passwd`, which is a deliberate reading of the issue rather than an oversight: a run's origin is unknowable at this layer, so "never reject the dispatch" and "still refuse a token whose `..` segments leave the repository root" cannot both hold for the same input, and only the first is decidable without an origin. Neither outcome widens authority, since an inferred path selects project rules but never becomes a write boundary and an anchored remainder carries no `..` to escape with, and neither is silent: both the drop and the rewrite are named per token in a `[dispatch scope]` notice, on the stderr diagnostic, and on the `dispatch_plan` approval artifact, which now probes the dispatch's own `cwd` rather than the orchestrator's process directory so the plan an operator approves resolves the same scope the run does. A `..` or `.` outside the leading run (`src/../../b.ts`, `../src/./b.ts`) walks back out of a segment it already anchored and keeps the refusal, and typed `intent.read_roots`, `write_roots`, and `relevant_paths` still refuse any `..` with `intent_path_escapes_root`.
82
+ - A fresh home no longer bounds every llama.cpp endpoint to one slot until someone runs `targets --probe`. The dispatch domain probes each configured endpoint whose bound resolved to the local-native default once when it starts, in the background and without the inference-based reasoning check, and the count lands in the provider statuses and the durable slot store where admission already looks.
83
+ - Compiled-prompt reuse now keys every resolved runtime input and the exact provider-facing attached-tool schema bytes, while handbook context, project rules, operator profile, workspace facts, and repository awareness are snapshotted per session and refreshed only by configuration invalidation or a new session; unrelated cache misses no longer admit opportunistic disk drift.
84
+ - Tool guidance now shares one deterministic role-aware normalizer, renders Marketplace installation ownership once, and remains absent for unavailable tools and roles. Bundled `code_nav source=clio` keeps its source in continuation guidance and resolves indexed paths under the installed or explicitly overridden package root.
85
+ - Extension manifests once again permit the documented omission of `resources` while strictly validating any value that is present. Upgrade now backs up and atomically adds verified content digests to valid pre-digest install records without changing their source, timestamp, or disabled state; invalid legacy trees remain visible but inactive. Corrupt extension state stays fail-closed but no longer traps operators: forced reinstall and removal preserve corrupt state and unverifiable package bytes before recovery. Share-imported extension packages now pass whole-tree preflight and the canonical transactional installer, so a successful import always has a verified install record before resources activate.
86
+ - Extension admission now rejects malformed manifests, invalid or changed installed trees, escaping links, hard-linked files, and special resource entries before they can contribute resources or hooks; canonical package discovery deduplicates aliases while retaining failures as deterministic diagnostics.
87
+ - Documentation that disagreed with the code: `integrations.externalAgents.defaults.toolGovernance` defaults to `clio-coder-policy` (the page said `clio-policy`); `CLIO_CODER_FORCE_COMPACT` compacts before every interactive turn while set, not once; `CLIO_CODER_TRUST_PROJECT_RESOURCES` can only enable trust, never revoke a setting that grants it; `CLIO_CODER_TIMING` prints only on the bannered non-interactive boot; `CLIO_CODER_RESIDENCY` also accepts `0`, `false`, `user`, and `user-managed` and is exported to SSH workers; `CLIO_CODER_RUN_JOURNAL`, `CLIO_CODER_EVAL_RUNNER_STDOUT_FILE`, and `CLIO_CODER_YAZI_PICK_TOKEN` have rows.
88
+ - `src/tools/dispatch-schema.ts` carried one NUL byte since round 4 and diffed as binary; stripped.
89
+ - A parallel dispatch batch runs each declared host check once, after every live member of the wave has finished, and charges a failing check by write-root coverage across the whole wave. The checks used to run per worker on the shared checkout while siblings were still editing it, so one worker's broken test sealed `host_verification_rejected` on all three receipts of a three-task batch (round-5 interactive exercise, 2026-09-02, kvlog with workers on mini `ornith1.5-35b-moe`). Exculpation always requires positive evidence: a member is cleared only when the failing check named a path and every named path falls inside its own boundary or inside a live sibling's, and a member that declared the check but no `intent.write_roots` ran with no write confinement at all, so it is charged unconditionally. A failing check that names no path, or names one no member claims, still rejects every member that declared it. A cleared member seals the new `hostVerification.status: "not_implicated"` rather than `verified`, because its `checks` array still carries the non-zero exit code and `verified` is what the trust surface, the board, and `worker evidence` all read as "the declared checks passed"; the run is not failed by it and no validator speaks for it. The receipt also seals `hostVerification.strategy: "batch-settled"` and the per-check attribution, whose `basis` names the weakest evidence behind the charge; a single-task dispatch omits both and its receipt bytes are unchanged. The barrier waits only on members that already hold a capacity lease and never on one still queued for capacity, so a batch larger than `fleet.concurrency` cannot deadlock on its own parked leases, and each later admission wave forms its own barrier instead of racing. The queued member of such a batch is admitted once the live members settle rather than when the first of them finishes, so a batch whose first member finishes inside the 60 s admission deadline and whose last does not now expires that member's admission and aborts the batch; releasing the lease before the barrier would readmit a writer into the checkout the settlement is about to judge, so that wait stands. Taking the workspace fingerprint once on the settled tree also stops the verification memo key churning between siblings (`tests/contracts/host-verification-batch.test.ts`).
90
+
91
+ - ACP's per-prompt usage accumulator (`AcpServerUsage`, `src/engine/acp/server.ts:206`) carried five discrete token fields and never a combined total or a dollar figure, while the TUI and the CLI `--json`/`--json-events` dispatch stream both carry `totalTokens` and `costUsd` for the same run; an editor connected over ACP had to reconstruct the total itself and had no way at all to show cost. `mergeUsage` now reads `usage.cost.total` off the same `message_end`/`agent_end` message objects `sumRunUsage` already reads for the TUI and json-stream paths, so ACP's dollar figure is never re-derived through a separate pricing calculation, and `totalTokens` follows the same explicit-value-or-summed-fallback rule `sumRunUsage` uses. Both fields ride in the existing `clio-coder/usage` `_meta` key on the `session/prompt` result. `tests/contracts/acp-usage-meta.test.ts` pins the accumulation, including the malformed-cost and explicit-total-wins cases.
92
+
93
+ ### Removed
94
+ - `CLIO_CODER_RESUME_SESSION_ID` and `CLIO_CODER_BOOTSTRAP_GENERATE_CHILD` are no longer read: nothing has set the first since headless `--session` and `--continue` replaced the self-restart, and the bootstrap scout that the second guarded against now runs as an internal dispatch. The `budgetTokens` key in seven prompt fragments' frontmatter went with them; the loader reads only `id`, `version`, `description`, and `dynamic`. `TERM_PROGRAM` is still read for OSC 9 notification routing but no longer copied into the stream-pacing environment it never consulted.
95
+ - `CLIO_CODER_SYNTHESIS_LOCK` is no longer read, and a synthesis-locked worker round always removes the tool surface on OpenAI-family runtimes. Its `tool-choice` mode kept the schemas and sent `tool_choice: none` instead, which preserved the locked round's prefix (350 to 900 new tokens against 1,200 to 1,800 for the strip, roughly half the prefill time), but it never beat the strip on any family measured on mini with a two-call recipe that locks every run, losing on two of the three and tying on the third. Lost-result counts, strip against tool-choice: Qwen3.8-27B 1 of 5 against 2 of 5, Ornith 1.5 35B 0 of 5 against 0 of 5, Ornith 1.5 9B 2 of 5 against 3 of 5. Under tool-choice the model called a tool anyway in 4 of 5 Qwen3.8-27B runs and 2 of 5 Ornith 1.5 35B runs, with no such count taken on Ornith 1.5 9B. Setting the variable now does nothing.
96
+
5
97
  ## 0.4.1 - 2026-09-01
6
98
 
7
99
  This grew beyond the bug-fix-only patch originally planned. v0.4.1 is a full release led by the version-2 `settings.yaml` contract, its automatic migration, and a smaller grammar-driven slash-command surface. It adds editor and marketplace workflows; fixes the release-blocking configuration, CLI, TUI, and Workbench failures found in final testing; replaces the oversized test and CI machinery with one fast deterministic gate; consolidates evaluation around the shipping eval domain and `evals/` reference suites; and moves machine-facing names into the `clio-coder` namespace without changing the product persona.
package/CONTRIBUTING.md CHANGED
@@ -38,20 +38,24 @@ Release gate (for maintainers before tags or release artifacts):
38
38
  npm run ci:release
39
39
  ```
40
40
 
41
- Live LLM smoke validation (manual/opt-in):
41
+ Live provider validation (manual/opt-in, after `npm run build`):
42
42
 
43
43
  ```bash
44
- npm run live:smoke -- --target <configured-target-id>
44
+ node dist/cli/index.js run \
45
+ --target <configured-target-id> \
46
+ --autonomy read-only \
47
+ "Reply with exactly: CLIO_LIVE_OK"
45
48
  ```
46
49
 
47
50
  ## Testing conventions
48
51
 
49
- CLI-facing contract tests drive the built binary through the child-process
50
- harness in `tests/harness/spawn.ts`: `makeScratchHome()` gives the run an
51
- isolated `CLIO_CODER_HOME`, and `runCli(args, { env, cwd })` spawns `dist/cli` and
52
- returns its captured `stdout`, `stderr`, and exit code. Rebuild `dist/` with
53
- `npm run build` after changing CLI source, since these tests exercise the
54
- built output.
52
+ Contract tests import `src/` directly through tsx. The test scripts preload
53
+ `tests/harness/tmp-root.ts`, which gives the run one guarded temporary root,
54
+ and stateful tests use the helpers in `tests/harness/scratch-env.ts` to isolate
55
+ Clio's data, config, state, and cache directories. Smoke tests exercise the
56
+ built `dist/cli/index.js`; their files own the process drivers needed for each
57
+ boundary. Rebuild `dist/` after changing CLI or entry-point source before
58
+ running a focused smoke test.
55
59
 
56
60
  Do not assert a CLI subcommand's output by capturing `process.stdout.write`
57
61
  in-process. In-process stdout capture fights the node:test spec reporter:
@@ -64,12 +68,15 @@ output in-process while the reporter runs.
64
68
  ## Releasing
65
69
 
66
70
  Releases are cut from a tag. The GitHub release is created by CI; the npm
67
- publish is a manual maintainer step. The ordered procedure for a cut lives in
68
- [docs/release-cut-checklist.md](docs/release-cut-checklist.md).
69
-
70
- 1. Bump `version` in `package.json` and retitle the top section of
71
- `CHANGELOG.md` to the version being cut. `scripts/check-release.mjs` fails
72
- when the two disagree or the heading still says `Unreleased`.
71
+ publish is a manual maintainer step. The current procedure is the sequence
72
+ below together with `.github/workflows/release.yml` and
73
+ `scripts/check-release.mjs`. The
74
+ [v0.4.1 release-cut checklist](docs/history/release-cut-checklist.md) is a
75
+ historical record, not a reusable current checklist.
76
+
77
+ 1. During development, keep the top changelog section at `## Unreleased` and
78
+ bump `version` in `package.json` when opening the release branch. Before the
79
+ cut, retitle that section `## <version> - YYYY-MM-DD`.
73
80
  2. Run `npm run ci:release`. It runs the full `ci` gate, then
74
81
  `scripts/check-release.mjs`, which verifies the built `dist/` and audits
75
82
  the exact npm package contents.
@@ -82,10 +89,13 @@ publish is a manual maintainer step. The ordered procedure for a cut lives in
82
89
  release with the tarball attached and the version's `CHANGELOG.md` section
83
90
  as the body. It does not publish to npm.
84
91
  6. A maintainer publishes from the tagged commit with `npm publish`;
85
- `prepublishOnly` runs the same `ci:release` gate first.
92
+ `prepublishOnly` runs the same `ci:release` gate in release mode first.
86
93
 
87
94
  What `scripts/check-release.mjs` enforces, and how to respond when it fails:
88
95
 
96
+ - Development branches may open with `## Unreleased`. Exact version tags and
97
+ `npm publish` require `## <version> - YYYY-MM-DD`, so unfinished notes cannot
98
+ enter an immutable artifact.
89
99
  - Only `dist/cli/index.js` and `dist/worker/entry.js` carry a shebang. The
90
100
  shebang comes from the hashbang line in each entry source file. Never add
91
101
  a tsup `banner`; it would stamp every chunk in `dist/`.
@@ -98,7 +108,7 @@ What `scripts/check-release.mjs` enforces, and how to respond when it fails:
98
108
  both package.json `files` and the required list in `check-release.mjs`.
99
109
  The double bookkeeping is deliberate: neither edit can silently drop a
100
110
  resource the CLI needs at runtime.
101
- - Size budgets: 15 MB tarball, 40 MB unpacked, set in `check-release.mjs`.
111
+ - Size budgets: 10 MB tarball, 50 MB unpacked, set in `check-release.mjs`.
102
112
  They are a tripwire for packaging defects such as a leaked `node_modules`
103
113
  or a doubled `dist/`, not a diet. If a legitimate change exceeds them,
104
114
  raise the budget in the same PR with a justification, never as a drive-by.
@@ -123,18 +133,26 @@ that needs them runs.
123
133
 
124
134
  ## Architecture Invariants
125
135
 
126
- The boundary checker enforces these:
127
-
128
- - Engine boundary: only `src/engine/**` value-imports pi SDK packages
129
- (`@earendil-works/pi-*`, pinned in `package.json`).
130
- - Worker isolation: `src/worker/**` value-imports only the worker-safe
131
- provider runtime rehydration modules under `src/domains/providers/**`;
132
- all other worker domain imports must be type-only.
133
- - Domain independence: cross-domain flows go through `SafeEventBus`.
136
+ The boundary checker enforces these six rules:
137
+
138
+ - Engine boundary: only `src/engine/**` imports the
139
+ `@earendil-works/pi-*` packages, including type-only imports.
140
+ - Worker isolation: `src/worker/**` may value-import only the declared
141
+ provider runtime rehydration seams under `src/domains/providers/**`; all
142
+ other worker imports from domains must be type-only.
143
+ - Domain independence: one domain never imports another domain's
144
+ `extension.ts`; cross-domain behavior uses public contracts and event buses.
145
+ - Tool substrate: `src/tools/**` never imports `src/interactive/**`.
146
+ - Entry-point composition: `src/interactive/turn-*.ts` and `chat-loop.ts` never
147
+ import `src/entry/**`.
148
+ - Stage 0 closure: external value importers enter the protected instant-shell
149
+ graph only through declared seams, and those seams may not create an
150
+ undeclared edge back into the closure.
134
151
 
135
152
  The checker runs as part of `npm run lint` (`scripts/check-hygiene.ts` imports
136
153
  `tests/boundaries/check-boundaries.ts`), so a boundary violation fails the
137
- same lint every PR runs.
154
+ same lint every PR runs. The full definitions and exceptions live in
155
+ [Architecture](docs/architecture/architecture.md#boundary-invariants).
138
156
 
139
157
  ## Branches
140
158
 
@@ -206,19 +224,24 @@ Agents should:
206
224
  `skills/` is the curated skills marketplace: maintainer-approved `SKILL.md`
207
225
  guides, distinct from the runtime skills any user can drop into a discovery
208
226
  root. It is not itself a discovery root, so nothing here auto-loads; skills
209
- activate only via `clio-coder skills install <name>`.
227
+ activate after `clio-coder skills install <name>` or from an explicit
228
+ `clio-coder --skill skills/<category>/<name>/SKILL.md` development path.
210
229
 
211
230
  To propose a skill:
212
231
 
213
- 1. Add `skills/<name>/SKILL.md`. Follow the `superpowers:writing-skills`
214
- methodology and Anthropic's skill-authoring guidance: a trigger-rich
215
- `description` (third person, "Use when ..."), one excellent example, and
216
- progressive disclosure (push heavy reference into `references/`).
217
- 2. Include the provenance frontmatter (`registry-id`, `source-url`, `version`,
218
- `license`) and ship an `evals.md` with the baseline scenarios you tested.
219
- 3. Verify locally: `clio-coder skills validate skills/<name>/SKILL.md`, then
220
- `clio-coder skills install <name>` and `clio-coder skills list`.
221
- 4. Open a PR. A maintainer reviews against the rubric, then sets `audit: pass`
222
- and the `version` to approve it for the catalog.
232
+ 1. Add `skills/<category>/<name>/SKILL.md`. Follow the local
233
+ [`skill-craft`](skills/meta/skill-craft/) guidance: put trigger phrases in
234
+ `triggers`, keep `description` to the job and explicit routing boundaries,
235
+ and move conditional detail into `references/`.
236
+ 2. Include the core `name`, `description`, `version`, and `license` fields plus
237
+ a nested `clio-coder:` block with `registry-id`, `source-url`, `provenance`,
238
+ and `eval-status`. Ship an `evals.md` with the baseline scenarios.
239
+ 3. Verify locally with
240
+ `clio-coder skills validate skills/<category>/<name>/SKILL.md`.
241
+ 4. Install the candidate by name and confirm it appears with
242
+ `clio-coder skills list`.
243
+ 5. Open a PR. A maintainer reviews against the rubric, sets `audit: pass`,
244
+ approves the catalog version, then regenerates and checks the catalog with
245
+ `npm run skills:pin` and `npm run skills:check`.
223
246
 
224
247
  Full catalog conventions and install options: [skills/README.md](skills/README.md).