@iowarp/clio-coder 0.3.8 → 0.3.9

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (333) hide show
  1. package/CHANGELOG.md +49 -0
  2. package/README.md +7 -3
  3. package/dist/{acp-U67UHUK2.js → acp-7LOELQFP.js} +6 -6
  4. package/dist/{agents-YU6SGALZ.js → agents-FIBG2SHA.js} +27 -25
  5. package/dist/assets/codewiki.json +1 -1
  6. package/dist/{auth-5ZPJOIVG.js → auth-OI4LIH2I.js} +11 -12
  7. package/dist/{builtins-C6JMZVV6.js → builtins-AD25UL3C.js} +5 -5
  8. package/dist/{chunk-5FR74PWO.js → chunk-2JDWVJND.js} +2 -2
  9. package/dist/chunk-3DPEIQKN.js +113 -0
  10. package/dist/{chunk-4SPRNWDE.js → chunk-3DUR4WUA.js} +15 -15
  11. package/dist/{chunk-VHN4MY6O.js → chunk-3MRC2YSQ.js} +2 -2
  12. package/dist/{chunk-5DHKRSMQ.js → chunk-3UUY7R3Z.js} +11 -7
  13. package/dist/{chunk-IGWKHNIQ.js → chunk-3V5AYSEQ.js} +8 -8
  14. package/dist/{chunk-FHJEP5SW.js → chunk-465CC7FK.js} +8 -5
  15. package/dist/{chunk-5WIGXA4T.js → chunk-47CMYGET.js} +111 -4
  16. package/dist/{chunk-A3WNZD3P.js → chunk-4H6ULJ3H.js} +67 -21
  17. package/dist/{chunk-TB5666IT.js → chunk-4LJX2PUC.js} +3 -3
  18. package/dist/{chunk-XWSF374K.js → chunk-56KB5IJP.js} +2 -2
  19. package/dist/{chunk-SPULKLCF.js → chunk-5DQRIYDZ.js} +2 -2
  20. package/dist/{chunk-2HEJ2F35.js → chunk-5HFBWUMU.js} +20 -8
  21. package/dist/{chunk-WNIJTQQK.js → chunk-5PVQ4SRS.js} +78 -6
  22. package/dist/{chunk-DYIM5TJT.js → chunk-5QKCQQ3E.js} +262 -6
  23. package/dist/{chunk-TYPGUK6W.js → chunk-5T7RBWN2.js} +111 -5
  24. package/dist/{chunk-IIZWH4XA.js → chunk-774ILSRL.js} +2 -2
  25. package/dist/chunk-7C6RYZGQ.js +391 -0
  26. package/dist/{chunk-TANS5ZJS.js → chunk-AD7Y7STJ.js} +3 -3
  27. package/dist/{chunk-RWSI4YD7.js → chunk-AEYBF3TB.js} +33 -12
  28. package/dist/{chunk-DGSYXYMX.js → chunk-AMKHQW3C.js} +2 -2
  29. package/dist/{chunk-VWZOAB7K.js → chunk-B5XRQOLB.js} +7 -7
  30. package/dist/{chunk-WXY7KU3G.js → chunk-BVDVID7E.js} +2 -2
  31. package/dist/{chunk-PMDBGQSJ.js → chunk-CA42X6KT.js} +2 -2
  32. package/dist/{chunk-HLE42MG7.js → chunk-D73KXYPF.js} +3 -3
  33. package/dist/{chunk-5Q2VVUKB.js → chunk-DG4M6ZUE.js} +3 -3
  34. package/dist/{chunk-ME6CCNFO.js → chunk-EBOC7MT3.js} +6 -6
  35. package/dist/{chunk-MXKJU4JB.js → chunk-ECUO3KDP.js} +48 -7
  36. package/dist/{chunk-26LEYJZH.js → chunk-FALJGAWU.js} +2 -2
  37. package/dist/{chunk-GWS3VEIW.js → chunk-FWDFM5ZU.js} +24 -3
  38. package/dist/{chunk-7RGZWPB6.js → chunk-GAYUJ7LE.js} +67 -13
  39. package/dist/{chunk-VCBR6CU7.js → chunk-HAY4ZE2P.js} +2 -2
  40. package/dist/{chunk-FBVTI2TJ.js → chunk-HCBCAYZU.js} +11 -130
  41. package/dist/{chunk-WSB3FPX7.js → chunk-HJB5IUKP.js} +32 -136
  42. package/dist/{chunk-J3YUBZWY.js → chunk-HKO36JWF.js} +33 -5
  43. package/dist/{chunk-E77JEWSD.js → chunk-HPCTNZM2.js} +6 -36
  44. package/dist/{chunk-FJ3H4MN5.js → chunk-IHKBWSXF.js} +2 -2
  45. package/dist/{chunk-NMPKI6XL.js → chunk-JEQQR47K.js} +37 -18
  46. package/dist/{chunk-3BINW3FP.js → chunk-KV2AOLDF.js} +24 -4
  47. package/dist/{chunk-YS5VLNH5.js → chunk-LXPJXFM5.js} +7 -7
  48. package/dist/{chunk-7RFXX52T.js → chunk-MIX5N5AC.js} +271 -41
  49. package/dist/{chunk-K4XHGFR5.js → chunk-MLOK6ZOS.js} +1297 -218
  50. package/dist/{chunk-ZNLWCMVZ.js → chunk-MV2VUEJC.js} +2 -2
  51. package/dist/{chunk-MVVUPGPW.js → chunk-MXHC5QYU.js} +6 -6
  52. package/dist/{chunk-5H3GB5BO.js → chunk-N3PBVRTZ.js} +4 -382
  53. package/dist/{chunk-2HFZQUHL.js → chunk-N5XKWMDW.js} +17 -7
  54. package/dist/{chunk-5C3AQNDW.js → chunk-NNNWO6F2.js} +124 -36
  55. package/dist/{chunk-GU2UIAFZ.js → chunk-NQ6UCCOD.js} +3 -3
  56. package/dist/{chunk-N22QMJKY.js → chunk-NZU6YDNV.js} +4 -4
  57. package/dist/{chunk-ZVJ5BLO2.js → chunk-O6I4CIEU.js} +151 -13
  58. package/dist/{chunk-XN3L4EYL.js → chunk-OEDBCISO.js} +2 -2
  59. package/dist/{chunk-VAWNZU7Z.js → chunk-P3JGPQFL.js} +2 -2
  60. package/dist/{chunk-IJ7RPIYJ.js → chunk-PNY46YEY.js} +20 -3
  61. package/dist/{chunk-U6MBIEMB.js → chunk-PZ4I4JE2.js} +56 -35
  62. package/dist/{chunk-GPIEI3LY.js → chunk-QQ7EKM72.js} +2 -2
  63. package/dist/{chunk-WLFILSD5.js → chunk-R7LNVMCS.js} +66 -28
  64. package/dist/{chunk-GOXNB3AO.js → chunk-RAPCMZL4.js} +75 -4
  65. package/dist/chunk-RKKLTLYB.js +45 -0
  66. package/dist/{chunk-TT36MB5S.js → chunk-RKRLDWD3.js} +3 -1
  67. package/dist/{chunk-TTHACPOM.js → chunk-S4COXYBG.js} +456 -18
  68. package/dist/{chunk-WWCZ5F23.js → chunk-T3Z6VAAF.js} +69 -10
  69. package/dist/{chunk-GN57SG4G.js → chunk-TD7UE2L5.js} +9 -7
  70. package/dist/{chunk-TLQJPP24.js → chunk-TEO2TLVN.js} +523 -322
  71. package/dist/{chunk-FCSXB6T2.js → chunk-UOSL25KY.js} +14 -2
  72. package/dist/{chunk-KTYTFRMB.js → chunk-VKBMFOYV.js} +17 -15
  73. package/dist/{chunk-PT7HYKEM.js → chunk-VO2LKSTM.js} +2 -2
  74. package/dist/{chunk-P43ETTHK.js → chunk-VPTUJU4P.js} +2 -2
  75. package/dist/{chunk-WEH5XRJQ.js → chunk-WIE7ZOSW.js} +2 -2
  76. package/dist/{chunk-EMYUUSFG.js → chunk-WXCJ7VME.js} +5 -5
  77. package/dist/{chunk-4DGYLA73.js → chunk-XDOQXGFO.js} +22 -7
  78. package/dist/{chunk-JOZYP4GM.js → chunk-YKOFT37S.js} +5 -5
  79. package/dist/chunk-YSEHGPCT.js +127 -0
  80. package/dist/{chunk-HFSBBKSQ.js → chunk-YW7UVM5V.js} +138 -3
  81. package/dist/cli/index.js +27 -27
  82. package/dist/{clio-QVTYJ57A.js → clio-LT5V7SSZ.js} +6 -6
  83. package/dist/{code-nav-FGGFIE7L.js → code-nav-LMW275PA.js} +4 -4
  84. package/dist/{config-LW5IJFQN.js → config-RXS5T3JT.js} +71 -43
  85. package/dist/{configure-7XIZCOU4.js → configure-2WYWSCSD.js} +14 -15
  86. package/dist/{context-Y6Y7QPR6.js → context-I3BTOTCS.js} +12 -12
  87. package/dist/{context-L3WL3X7K.js → context-MVOORGMF.js} +34 -33
  88. package/dist/{context-N52ZA626.js → context-PALKKQYL.js} +20 -20
  89. package/dist/{context-clear-MBQRLSDQ.js → context-clear-N2WOYZ2K.js} +34 -33
  90. package/dist/{context-index-HVMFQHK3.js → context-index-HNG3MOME.js} +2 -2
  91. package/dist/{context-working-set-GS6DSO7F.js → context-working-set-MIEVECVZ.js} +10 -11
  92. package/dist/{dispatch-runner-22ZCNOM3.js → dispatch-runner-VVA4SRRH.js} +34 -33
  93. package/dist/doctor-TWBWFK5V.js +165 -0
  94. package/dist/{eval-BEC2WHDA.js → eval-IJ5VEZDJ.js} +2016 -142
  95. package/dist/{evidence-REJUMSKM.js → evidence-L5APPXNV.js} +29 -28
  96. package/dist/{evolve-PY5ZBA5K.js → evolve-RGNKFJ52.js} +29 -28
  97. package/dist/{extensions-HVKU65YU.js → extensions-7WYWUX5A.js} +9 -3
  98. package/dist/{fleet-7WZEWRFA.js → fleet-6CNVBZZP.js} +87 -55
  99. package/dist/{fleet-commands-UVHWM76J.js → fleet-commands-L2SXSYEI.js} +6 -6
  100. package/dist/{fleet-graph-6ULH7PES.js → fleet-graph-2J3OOIPO.js} +14 -12
  101. package/dist/{fleet-preflight-J53T6CCE.js → fleet-preflight-CZRJ4JP5.js} +3 -4
  102. package/dist/{fleet-validate-72PC4SLA.js → fleet-validate-C5RI6DP7.js} +16 -15
  103. package/dist/{init-OG3TPGQG.js → init-VBN2ACVA.js} +50 -48
  104. package/dist/{library-CNTMPLRF.js → library-JHGUMLY2.js} +13 -11
  105. package/dist/{memory-6IS7F275.js → memory-K4OQIYWG.js} +31 -30
  106. package/dist/{models-ENRJDA5W.js → models-2NCZUWDD.js} +23 -22
  107. package/dist/{monitor-XLDVO7TN.js → monitor-MMVTJABD.js} +35 -34
  108. package/dist/{orchestrator-6KSPYRHA.js → orchestrator-ZKBPCHW6.js} +1627 -294
  109. package/dist/{reset-RZ4ER727.js → reset-DD5JGOY3.js} +3 -3
  110. package/dist/{run-Y2CNK5RU.js → run-QEGNX7FL.js} +56 -55
  111. package/dist/{share-A55GYP6Z.js → share-JKD3BQMW.js} +13 -11
  112. package/dist/{skills-ALC5J6AT.js → skills-LMQIKDOZ.js} +14 -12
  113. package/dist/{skills-eval-JPBEBYQU.js → skills-eval-I7X2774U.js} +34 -32
  114. package/dist/{steer-GGWFUJUD.js → steer-CF5TDANS.js} +3 -3
  115. package/dist/{support-MIETYA5E.js → support-I7LOJLIF.js} +4 -4
  116. package/dist/{targets-VGNXIR3S.js → targets-RUSR6B5Z.js} +58 -32
  117. package/dist/{terminal-lease-WOBR64YA.js → terminal-lease-QYVORFR4.js} +6 -4
  118. package/dist/{trace-PNCASAXC.js → trace-ODOQIVIW.js} +61 -6
  119. package/dist/{upgrade-FUSUAGHR.js → upgrade-XANW3FXB.js} +17 -16
  120. package/dist/{usage-N4MKVHKD.js → usage-4H7ZRXQT.js} +88 -44
  121. package/dist/{verifiers-YAWOJ3H2.js → verifiers-UZXNBZEB.js} +6 -6
  122. package/dist/{verify-LTDHYBGY.js → verify-BVKWTNDL.js} +5 -5
  123. package/dist/{wiki-generate-6M7GHTBJ.js → wiki-generate-MY7WV2QI.js} +49 -47
  124. package/dist/worker/entry.js +29 -33
  125. package/docs/alcf-provider.md +1 -1
  126. package/docs/architecture.md +1 -1
  127. package/docs/artifact-versions.md +6 -1
  128. package/docs/built-in-agents.md +1 -1
  129. package/docs/capacity-and-scheduling.md +23 -2
  130. package/docs/commands-and-modes.md +1 -1
  131. package/docs/configuration-and-targets.md +30 -3
  132. package/docs/context-engine.md +63 -4
  133. package/docs/documentation-coverage.md +3 -3
  134. package/docs/documentation-guide.md +1 -1
  135. package/docs/environment-variables.md +2 -0
  136. package/docs/eval-runner.md +262 -11
  137. package/docs/evals-internal.md +72 -2
  138. package/docs/evidence-and-memory.md +11 -10
  139. package/docs/evolution.md +1 -1
  140. package/docs/extensions-and-sharing.md +3 -1
  141. package/docs/fleet-dispatch.md +4 -4
  142. package/docs/installation-and-lifecycle.md +1 -1
  143. package/docs/middleware-and-components.md +1 -1
  144. package/docs/model-catalog.md +1 -1
  145. package/docs/observability.md +53 -2
  146. package/docs/proactive-memory.md +127 -14
  147. package/docs/prompt-envelope-and-tools.md +19 -1
  148. package/docs/provider-adapter-cookbook.md +1 -1
  149. package/docs/release-cut-checklist.md +19 -3
  150. package/docs/safety-model.md +1 -1
  151. package/docs/scientific-validation.md +1 -1
  152. package/docs/skills-marketplace.md +1 -1
  153. package/docs/tool-usage.md +1 -1
  154. package/docs/trace-store.md +1 -1
  155. package/docs/troubleshooting.md +87 -0
  156. package/docs/tui-design.md +1 -1
  157. package/docs/worker-dispatch-mechanics.md +1 -1
  158. package/package.json +2 -1
  159. package/src/cli/agents.ts +1 -1
  160. package/src/cli/config-inspect.ts +33 -6
  161. package/src/cli/config.ts +1 -1
  162. package/src/cli/doctor-state-size.ts +82 -0
  163. package/src/cli/doctor.ts +3 -1
  164. package/src/cli/eval.ts +80 -16
  165. package/src/cli/extensions.ts +5 -1
  166. package/src/cli/fleet.ts +32 -3
  167. package/src/cli/targets.ts +44 -13
  168. package/src/cli/trace.ts +63 -4
  169. package/src/cli/usage.ts +63 -14
  170. package/src/core/bus-events.ts +29 -1
  171. package/src/core/cache-telemetry.ts +42 -0
  172. package/src/core/config.ts +18 -0
  173. package/src/core/defaults.ts +36 -6
  174. package/src/core/endpoint-key.ts +27 -0
  175. package/src/core/residency-target-key.ts +25 -0
  176. package/src/core/response-schema.ts +36 -2
  177. package/src/domains/config/classify.ts +3 -0
  178. package/src/domains/context/codewiki/coordinator.ts +12 -4
  179. package/src/domains/dispatch/admission.ts +40 -3
  180. package/src/domains/dispatch/capacity-lease.ts +98 -9
  181. package/src/domains/dispatch/contract.ts +11 -0
  182. package/src/domains/dispatch/execution-plan.ts +44 -4
  183. package/src/domains/dispatch/extension.ts +166 -42
  184. package/src/domains/dispatch/fleet-run.ts +23 -3
  185. package/src/domains/dispatch/heartbeat.ts +32 -8
  186. package/src/domains/dispatch/index.ts +3 -0
  187. package/src/domains/dispatch/orphan-recovery.ts +5 -0
  188. package/src/domains/dispatch/reservation-store.ts +116 -8
  189. package/src/domains/dispatch/state.ts +4 -0
  190. package/src/domains/dispatch/worker-spawn.ts +25 -11
  191. package/src/domains/dispatch/write-boundary-enforcer.ts +20 -3
  192. package/src/domains/dispatch/write-boundary.ts +62 -1
  193. package/src/domains/eval/artifacts/store.ts +62 -0
  194. package/src/domains/eval/compare/behavioral.ts +224 -0
  195. package/src/domains/eval/compare/compare.ts +355 -2
  196. package/src/domains/eval/compare/envelope.ts +128 -0
  197. package/src/domains/eval/compare/gates.ts +24 -6
  198. package/src/domains/eval/compare/thresholds.ts +30 -3
  199. package/src/domains/eval/execution-provenance.ts +240 -0
  200. package/src/domains/eval/metrics/aggregate.ts +136 -0
  201. package/src/domains/eval/metrics/call-ledger-stream.ts +112 -0
  202. package/src/domains/eval/metrics/tracked.ts +413 -0
  203. package/src/domains/eval/provenance.ts +117 -0
  204. package/src/domains/eval/reports/comparison.ts +128 -0
  205. package/src/domains/eval/reports/junit.ts +17 -3
  206. package/src/domains/eval/reports/markdown.ts +3 -3
  207. package/src/domains/eval/reports/text.ts +14 -0
  208. package/src/domains/eval/run-compare.ts +20 -0
  209. package/src/domains/eval/runners/clio-run.ts +127 -0
  210. package/src/domains/eval/runners/external-command.ts +28 -3
  211. package/src/domains/eval/schema/adapter.ts +111 -0
  212. package/src/domains/eval/schema/artifact.ts +20 -0
  213. package/src/domains/eval/schema/behavioral-metrics.ts +204 -0
  214. package/src/domains/eval/schema/behavioral.ts +520 -0
  215. package/src/domains/eval/schema/execution-envelope.ts +194 -0
  216. package/src/domains/eval/schema/serving.ts +74 -0
  217. package/src/domains/eval/schema/suite.ts +38 -8
  218. package/src/domains/eval/schema/validate.ts +58 -3
  219. package/src/domains/eval/schema/verdict.ts +237 -0
  220. package/src/domains/eval/suites/resolve.ts +2 -0
  221. package/src/domains/eval/suites/run.ts +264 -33
  222. package/src/domains/eval/verifiers/command.ts +2 -1
  223. package/src/domains/eval/workspaces/temp-copy.ts +145 -13
  224. package/src/domains/evidence/build.ts +2 -13
  225. package/src/domains/evidence/eval.ts +2 -12
  226. package/src/domains/evidence/findings-markdown.ts +33 -0
  227. package/src/domains/evidence/run-trust.ts +7 -113
  228. package/src/domains/evidence/trust-projection.ts +2 -2
  229. package/src/domains/extensions/compatibility.ts +285 -0
  230. package/src/domains/extensions/discovery.ts +38 -3
  231. package/src/domains/extensions/resources.ts +1 -1
  232. package/src/domains/extensions/state.ts +12 -3
  233. package/src/domains/extensions/types.ts +2 -0
  234. package/src/domains/lifecycle/doctor.ts +69 -1
  235. package/src/domains/memory/index.ts +14 -0
  236. package/src/domains/memory/task-bank-promotion.ts +64 -0
  237. package/src/domains/memory/task-memory-policy.ts +77 -8
  238. package/src/domains/memory/task-memory-spend.ts +131 -0
  239. package/src/domains/memory/task-memory-status.ts +7 -0
  240. package/src/domains/memory/task-memory-telemetry.ts +2 -0
  241. package/src/domains/middleware/index.ts +1 -0
  242. package/src/domains/middleware/memory-intervention.ts +69 -5
  243. package/src/domains/middleware/memory-step-endpoint.ts +71 -0
  244. package/src/domains/observability/background-memory-usage.ts +140 -0
  245. package/src/domains/observability/cost.ts +1 -1
  246. package/src/domains/observability/index.ts +7 -0
  247. package/src/domains/observability/out-of-turn-usage.ts +51 -2
  248. package/src/domains/observability/trace-store.ts +192 -2
  249. package/src/domains/prompts/compiler.ts +100 -13
  250. package/src/domains/providers/endpoint-capacity.ts +96 -0
  251. package/src/domains/providers/index.ts +10 -0
  252. package/src/domains/providers/models/local-models/clio-local-coding-targets.yaml +243 -1
  253. package/src/domains/providers/runtime-resolution.ts +8 -1
  254. package/src/domains/providers/runtimes/common/probe-helpers.ts +31 -9
  255. package/src/domains/providers/runtimes/local-native/llamacpp-anthropic.ts +1 -1
  256. package/src/domains/providers/runtimes/local-native/llamacpp-completion.ts +1 -1
  257. package/src/domains/providers/runtimes/local-native/llamacpp-embed.ts +1 -1
  258. package/src/domains/providers/runtimes/local-native/llamacpp-rerank.ts +1 -1
  259. package/src/domains/providers/runtimes/local-native/llamacpp.ts +4 -1
  260. package/src/domains/providers/runtimes/local-native/lmstudio.ts +4 -1
  261. package/src/domains/providers/runtimes/local-native/ollama-native.ts +6 -1
  262. package/src/domains/providers/types/capability-flags.ts +2 -0
  263. package/src/domains/providers/types/target-descriptor.ts +2 -0
  264. package/src/domains/resources/prompts/loader.ts +95 -33
  265. package/src/domains/safety/call-target.ts +52 -0
  266. package/src/domains/safety/run-effects.ts +35 -4
  267. package/src/domains/session/context-accounting.ts +52 -1
  268. package/src/domains/session/context-ledger.ts +37 -13
  269. package/src/domains/session/index.ts +6 -0
  270. package/src/domains/session/prompt-cache.ts +140 -0
  271. package/src/domains/session/prompt-manifest.ts +42 -0
  272. package/src/engine/acp/adapter.ts +18 -3
  273. package/src/engine/ai.ts +35 -0
  274. package/src/engine/apis/llamacpp-residency.ts +55 -3
  275. package/src/engine/apis/lmstudio.ts +25 -5
  276. package/src/engine/apis/ollama-native.ts +2 -1
  277. package/src/engine/apis/openai-completions.ts +80 -17
  278. package/src/engine/apis/residency-lock.ts +3 -1
  279. package/src/engine/apis/residency.ts +34 -1
  280. package/src/engine/provider-payload.ts +29 -1
  281. package/src/entry/orchestrator.ts +176 -30
  282. package/src/interactive/chat-loop-messages.ts +26 -7
  283. package/src/interactive/chat-loop.ts +318 -41
  284. package/src/interactive/chat-panel.ts +62 -8
  285. package/src/interactive/clio-editor.ts +45 -8
  286. package/src/interactive/context-activity.ts +5 -1
  287. package/src/interactive/context-meter.ts +1 -1
  288. package/src/interactive/context-overlay.ts +40 -10
  289. package/src/interactive/cost-overlay.ts +64 -6
  290. package/src/interactive/dispatch-board.ts +84 -12
  291. package/src/interactive/fleet-run-preview.ts +41 -15
  292. package/src/interactive/handoff-round.ts +41 -2
  293. package/src/interactive/interactive-application.ts +24 -1
  294. package/src/interactive/interactive-input-runtime.ts +8 -0
  295. package/src/interactive/interactive-presentation.ts +4 -0
  296. package/src/interactive/interactive-shell.ts +20 -17
  297. package/src/interactive/interactive-slash-runtime.ts +27 -4
  298. package/src/interactive/memory-overlay.ts +8 -0
  299. package/src/interactive/mutation-preview.ts +295 -0
  300. package/src/interactive/overlay-general-openers.ts +16 -0
  301. package/src/interactive/overlay-key-routing.ts +38 -0
  302. package/src/interactive/overlay-lifecycle.ts +38 -5
  303. package/src/interactive/overlay-permission-lifecycle.ts +22 -2
  304. package/src/interactive/overlay-session-lifecycle.ts +73 -9
  305. package/src/interactive/overlays/ask-user.ts +91 -19
  306. package/src/interactive/overlays/help-reference.ts +4 -0
  307. package/src/interactive/overlays/prompts.ts +11 -1
  308. package/src/interactive/overlays/settings.ts +35 -1
  309. package/src/interactive/permission-hint.ts +34 -2
  310. package/src/interactive/permission-overlay.ts +159 -9
  311. package/src/interactive/prewarm.ts +197 -0
  312. package/src/interactive/render-trace.ts +162 -15
  313. package/src/interactive/renderers/tool-execution.ts +4 -0
  314. package/src/interactive/side-question.ts +58 -1
  315. package/src/interactive/status/controller.ts +11 -0
  316. package/src/interactive/status/state-machine.ts +54 -2
  317. package/src/interactive/status/types.ts +7 -0
  318. package/src/interactive/terminal-lease.ts +2 -0
  319. package/src/interactive/turn-context.ts +299 -31
  320. package/src/interactive/turn-persistence.ts +14 -4
  321. package/src/interactive/turn-prewarm.ts +364 -0
  322. package/src/interactive/turn-queues.ts +7 -4
  323. package/src/interactive/turn-runtime.ts +8 -1
  324. package/src/interactive/turn-state.ts +23 -0
  325. package/src/interactive/view/view-overlay.ts +28 -3
  326. package/src/tools/ask-user.ts +43 -2
  327. package/src/tools/dispatch-plan.ts +17 -9
  328. package/src/tools/dispatch-scout.ts +1 -1
  329. package/src/tools/registry.ts +16 -0
  330. package/dist/chunk-AOCYTWAV.js +0 -449
  331. package/dist/chunk-HWUFFB6L.js +0 -83
  332. package/dist/chunk-R346GLFC.js +0 -31
  333. package/dist/doctor-M7YEDGAE.js +0 -91
@@ -1,7 +1,7 @@
1
1
  # Clio Coder Local Evaluation Runner
2
2
 
3
3
  > [!TIP]
4
- > **Interactive Spec Available:** An interactive task suite validator, subprocess execution simulator, and compare calculator is located at [docs/html/eval_blueprint.html](html/eval_blueprint.html) (Version: 0.3.8).
4
+ > **Interactive Spec Available:** An interactive task suite validator, subprocess execution simulator, and compare calculator is located at [docs/html/eval_blueprint.html](html/eval_blueprint.html) (Version: 0.3.9).
5
5
 
6
6
  The local evaluation runner executes repository-local YAML task suites as deterministic subprocess checks. It is useful for comparing harness changes, prompts, tools, or local workflows.
7
7
 
@@ -15,10 +15,10 @@ The CLI commands under `clio-coder eval` support running, validating, reporting,
15
15
 
16
16
  ```bash
17
17
  clio-coder eval validate --suite <suite.yaml>
18
- clio-coder eval run --suite <suite.yaml> [--target <id>] [--model <id>] [--out <path>] [--clio-coder-entry <path>]
18
+ clio-coder eval run --suite <suite.yaml> [--trials <n>] [--target <id>] [--model <id>] [--out <path>] [--clio-coder-entry <path>]
19
19
  clio-coder eval run --task-file <tasks.yaml> [--repeat <n>] [--out <path>] [--clio-coder-entry <path>]
20
20
  clio-coder eval report <evalId> --format text|json|md|swe-jsonl|junit
21
- clio-coder eval compare <baselineEvalId> <candidateEvalId>
21
+ clio-coder eval compare <baselineEvalId> <candidateEvalId> [--metric <name>] [--format text|json|md|junit] [--allow-config-drift]
22
22
  clio-coder eval gate <candidateEvalId> --baseline <baselineEvalId> [--thresholds <file>]
23
23
  ```
24
24
 
@@ -32,7 +32,7 @@ clio-coder eval gate <candidateEvalId> --baseline <baselineEvalId> [--thresholds
32
32
  * `swe-jsonl`: Standardized JSONL format representing task runs (e.g. for SWE-bench comparisons).
33
33
  * `junit`: XML report for CI/CD integration.
34
34
  * **`compare`**: Compares two evaluation artifacts (baseline and candidate) by matching tasks.
35
- * **`gate`**: Compares candidate metrics against baseline or absolute thresholds, exiting non-zero if assertions fail (useful for PR gating).
35
+ * **`gate`**: Compares candidate metrics against baseline and absolute thresholds. Correctness and safety regressions fail independently of informational budgets.
36
36
 
37
37
  Exit codes:
38
38
 
@@ -41,8 +41,8 @@ Exit codes:
41
41
  | `eval validate` | `0` when validation passes | `2` for validation issues |
42
42
  | `eval run` | `0` when all task repetitions pass | `1` when any task fails, `2` for invalid configs |
43
43
  | `eval report` | `0` when artifact loads | `1` if artifact cannot be read, `2` for invalid ID |
44
- | `eval compare` | `0` when both artifacts load and compare succeeds | `1` if artifacts cannot be read, `2` for invalid ID |
45
- | `eval gate` | `0` when all threshold assertions pass | `1` if assertions fail, `2` for config/invalid ID errors |
44
+ | `eval compare` | `0` when both artifacts compare and the behavioral hard gate passes | `1` for a hard regression or unreadable artifact, `2` for invalid ID |
45
+ | `eval gate` | `0` when correctness, safety, and hard threshold assertions pass | `1` for any hard failure, `2` for config/invalid ID errors |
46
46
 
47
47
  ---
48
48
 
@@ -103,10 +103,11 @@ tasks:
103
103
  | --- | --- | --- |
104
104
  | `version` | - | Must equal `2`. |
105
105
  | `suite` | `id`, `title`, `visibility`, `description` | Metadata identifying the evaluation suite. |
106
- | `matrix` | `targets[]`, `repeats` | Matrix of execution targets (specifying model and thinking flags) and the repetition count. |
106
+ | `matrix` | `targets[]`, `repeats`, `dimensions[]` | Matrix of execution targets, repetition count, and the execution-envelope fields intentionally varied by the suite. |
107
107
  | `workspace` | `kind`, `path`, `url`, `commit`, `checkout`, `excludes` | Workspace strategy: `local` (run in-place), `git` (clone from URL), or `temp-copy` (isolated copy of a directory). |
108
108
  | `runner` | `kind`, `prompt`, `command`, `commands`, `args`, `timeoutMs` | Runner type: `clio-run` (starts Clio agent loop), `context-index` (runs indexer), `context-init` (initializes context), `external-command` (spawns subprocess). |
109
- | `verify` | `commands`, `assertions`, `forbidPaths` | Validation steps: shell commands, metric assertions (e.g. `op: lt` for max token counts), and files/directories that must not be created or modified (`forbidPaths`). |
109
+ | `behavioral` | `schema`, `corpus`, `execution`, `expectedBehavior`, `forbiddenBehavior`, `judge` | Optional `clio.eval.scenario.v1` behavioral contract. Rules name a closed category and a typed predicate over transcript, tool, receipt, or grader facts. |
110
+ | `verify` | `commands`, `measure`, `assertions`, `forbidPaths` | Validation steps: shell commands, a task-outcome grader, metric assertions (e.g. `op: lt` for max token counts), and files/directories that must not be created or modified (`forbidPaths`). |
110
111
  | `metrics` | `collect` | List of metric names to compile for the evaluation runs. |
111
112
 
112
113
  ---
@@ -114,7 +115,7 @@ tasks:
114
115
  ## Workspace Kinds
115
116
  * **`local`**: Executes the task directly in the specified local path.
116
117
  * **`git`**: Clones the repository from `url`, checks out the specified `commit` or `checkout` ref, and runs there.
117
- * **`temp-copy`**: Copies the directory at `path` to a temporary workspace location before running. This prevents side-effects from polluting other task runs.
118
+ * **`temp-copy`**: Copies the directory at `path` to a temporary workspace location immediately before the matrix item runs and removes it afterward. In a Git checkout, the copy contains exactly tracked files plus untracked files not excluded by Git ignore rules (`git ls-files --cached --others --exclude-standard`), with `excludes` applied afterward. Outside Git it retains the recursive directory copy. This prevents side-effects from polluting other task runs without copying ignored datasets or build trees.
118
119
 
119
120
  ---
120
121
 
@@ -194,12 +195,262 @@ export interface EvalArtifactV4 {
194
195
  matrix: { target: string; model: string | null; thinking: string | null };
195
196
  summary: EvalArtifactSummaryV4;
196
197
  results: EvalArtifactResultV4[];
198
+ servingConfiguration?: EvalServingConfigurationV1;
199
+ aggregates?: EvalScenarioAggregateV1[];
197
200
  }
198
201
  ```
199
202
 
203
+ `servingConfiguration` and `aggregates` are additive. The v4 reader still accepts an artifact that omits them, and each result's `verdict` is optional for the same reason, so an artifact written before this release loads unchanged.
204
+
200
205
  ---
201
206
 
202
- ## Task Outcome Measurement (`verify.measure`)
207
+ ## The verdict envelope
208
+
209
+ Every result carries a strictly parsed `clio.eval.verdict.v1` envelope (`src/domains/eval/schema/verdict.ts`). Suite v2 results are adapted into it at one explicit boundary (`src/domains/eval/schema/adapter.ts`) rather than by widening the artifact version, because the envelope carries no information a v4 artifact cannot hold.
210
+
211
+ ```json
212
+ {
213
+ "schema": "clio.eval.verdict.v1",
214
+ "scenarioId": "latency-nonnegative",
215
+ "trialIndex": 0,
216
+ "outcome": "pass",
217
+ "machinery": "ok",
218
+ "reason": null,
219
+ "trackedMetrics": { "...": "see below" },
220
+ "behavioral": null,
221
+ "evidence": {
222
+ "assignmentId": "qcy5rfopdrfw",
223
+ "terminalReceiptDigest": "d85a3ad4f8ae...",
224
+ "graderExitCode": 0
225
+ }
226
+ }
227
+ ```
228
+
229
+ The envelope is fail-closed by construction. `outcome` is one of `pass`, `fail`, or `unmeasured`; `machinery` is `ok` or `infrastructure_failure`; `reason` is null for a pass or unmeasured outcome and names the rule or failure class for every failure; the original `behavioral` reservation remains exactly `null`; and an envelope claiming both `infrastructure_failure` and `pass` is rejected at parse rather than recorded. A run whose harness broke therefore cannot be read as a model that succeeded. Behavioral results use the separately versioned sibling document below rather than changing this persisted schema.
230
+
231
+ ### Behavioral scenario and verdict documents
232
+
233
+ Behavioral evaluation is additive and does not change the persisted `clio.eval.verdict.v1` reader. A Suite v2 task may declare a `clio.eval.scenario.v1` block, and its Artifact v4 result then carries a sibling `clio.eval.behavior.v1` document whose `verdictRef` names the verdict schema, scenario id, and trial index. This preserves existing verdicts and the tracked-metrics baseline while making a cross-linked behavioral document independently parseable.
234
+
235
+ The closed categories are `tool_choice`, `exploration`, `delegation`, `safety_comprehension`, `claim_grounding`, `denied_tool_recovery`, `completion_behavior`, and `task_correctness`. Each category result is exactly one of `satisfied`, `violated`, `unknown`, or `unmeasured`. The document outcome is `pass`, `behavioral_failure`, `unknown`, `unmeasured`, or `infrastructure_failure`; missing facts are never invented as successes, and an infrastructure failure cannot become a behavioral pass.
236
+
237
+ Expected and forbidden rules contain typed predicates over facts sourced from `transcript`, `tool`, `receipt`, or `grader`. Facts cite a locator, SHA-256 digest, and optional bounded excerpt. The parser caps rules, facts, evidence per category, ids, and explanations. Before judging, facts and unavailable sources are sorted into a canonical representation and hashed as `judgeInputDigest`, so input order cannot change the judge result. Duplicate or conflicting facts, missing categories, malformed evidence, contradictory outcomes, and a behavioral document that references a different result are refused.
238
+
239
+ Suite execution adapts scalar run metrics into these observable facts at the Suite v2 to Artifact v4 boundary. A declared no-tool target leaves tool-dependent rules `unmeasured`, while an available evidence source that omits a required fact produces `unknown`. Categories a role-specific scenario does not claim to measure remain `unmeasured`; they are not numeric zero and do not silently satisfy a rule.
240
+
241
+ ### Public built-in behavioral corpus
242
+
243
+ The repository ships corpus `public-built-in-behavior` version `1.0.0` under
244
+ `benchmarks/eval/`. It contains no private prompts, endpoints, credentials, or
245
+ mutable external dataset:
246
+
247
+ - `behavioral-machinery.yaml` provides one positive and one adversarial
248
+ machinery-only check for each of the 13 shipped built-in worker recipes. Its
249
+ deterministic driver loads the production recipe catalog, admits a real
250
+ dispatch through the production gate, runs a scripted worker, and verifies
251
+ the sealed receipt and result-contract outcome. The 26 scenarios require no
252
+ model; they do not infer behavior by grepping recipe frontmatter.
253
+ - `behavioral-model.yaml` provides four isolated main-agent scenarios on the
254
+ `mini` target: a focused edit, adversarial scope control, required
255
+ delegation, and recovery after Bash is denied. Together they cover all eight
256
+ behavioral categories with per-tool call and blocked-call counts, distinct
257
+ and allowlisted read-path counts, declared decoy hits, and grader-emitted
258
+ claim-support and completion facts.
259
+ - `behavioral-model-negative-control.yaml` intentionally reads a declared
260
+ decoy. A healthy run solves its literal task while recording
261
+ `behavioral_failure` with violated exploration and safety labels, proving
262
+ that the rules can reject observed model behavior rather than merely restate
263
+ aggregate success counters.
264
+
265
+ Build once, then run either focused suite from the repository root:
266
+
267
+ ```sh
268
+ node dist/cli/index.js eval run --suite benchmarks/eval/behavioral-machinery.yaml --clio-coder-entry dist/cli/index.js
269
+ node dist/cli/index.js eval run --suite benchmarks/eval/behavioral-model.yaml --target mini --clio-coder-entry dist/cli/index.js
270
+ node dist/cli/index.js eval run --suite benchmarks/eval/behavioral-model-negative-control.yaml --target mini --clio-coder-entry dist/cli/index.js
271
+ ```
272
+
273
+ The machinery tasks use the repository read-only and create only private
274
+ scratch state under `TMPDIR`; model tasks use a fresh `temp-copy` workspace and
275
+ remove it after the matrix item settles. The machinery suite is the fast
276
+ admission, worker, and receipt contract. The model suite is the live behavioral
277
+ measurement: keep its Artifact v4 output as evidence for the exact target and
278
+ serving configuration that ran, rather than treating one observed model result
279
+ as a universal guarantee. Behavioral facts and their evidence store only
280
+ bounded read counters, not path strings. As with other eval runs, the artifact's
281
+ bounded diagnostic stdout may retain the underlying tool event stream.
282
+
283
+ ### `trackedMetrics`
284
+
285
+ Eleven numbers plus a reason histogram, each carrying the source it came from. `source` is `ledger` (the per-call ledger folded from the worker's own JSON stream), `receipt` (the sealed run receipt), or `estimated`, and `estimated` is what a missing observation is marked as rather than being silently counted as measured.
286
+
287
+ | Metric | Usual source |
288
+ | --- | --- |
289
+ | `modelCalls` | ledger |
290
+ | `uncachedPrefillTokens` | ledger, from `promptCache.backend` |
291
+ | `cacheReadTokens` | ledger, from `promptCache.backend`, falling back to pi-ai cache reads |
292
+ | `generatedTokens` | ledger |
293
+ | `reasoningTokens` | receipt; nullable, because absent and zero are different claims |
294
+ | `toolCalls`, `toolErrors` | ledger when present, otherwise receipt |
295
+ | `ttftMsFirstCall` | ledger |
296
+ | `wallClockMs` | receipt |
297
+ | `contextTokensAtEnd` | ledger |
298
+ | `compactions` | ledger |
299
+ | `expectedColdReasons` | ledger, one sourced count per reason |
300
+
301
+ A dispatched worker's receipt reports `sessionId: null` and writes no session archive, which is why the ledger source exists at all: the runner folds structured usage, backend timing, cache, and monotonic TTFT facts out of the worker's `message_end` events. It keeps no prompt text, no model prose, and no tool-result content in that fold.
302
+
303
+ ### Scenario aggregates
304
+
305
+ `aggregates` groups verdicts by `scenarioId`, sets `k` to the trial count, and records `passAtK` (any trial passed) and `passPowK` (every trial passed). Each tracked numeric metric reports observation, measured, and unmeasured counts, mean, min, max, nearest-rank p90, population variance, standard deviation, and the set of sources observed. A metric with no observation keeps every numeric statistic `null`; it never becomes zero. At `k: 1`, variance and standard deviation are zero only when the value was actually measured.
306
+
307
+ ### Behavioral multi-metric results
308
+
309
+ A result with a `clio.eval.behavior.v1` verdict also carries the additive
310
+ `clio.eval.behavior.metrics.v1` projection. The projection binds the scenario
311
+ to its role and target/model envelope and records one `number | null`
312
+ observation for each closed metric. The source travels beside every value:
313
+
314
+ | Family | Metric | Direction | Gate | Source |
315
+ |---|---|---|---|---|
316
+ | correctness | `correctness.taskSolved` | higher | hard | grader |
317
+ | safety | `safety.violations` | lower | hard | behavioral label |
318
+ | behavior | `behavior.labelViolations` | lower | informational | behavioral labels |
319
+ | efficiency | `efficiency.toolCalls` | lower | informational | terminal tool events |
320
+ | exploration | `exploration.unnecessaryReads` | lower | informational | read observation counters |
321
+ | delegation | `delegation.quality` | higher | informational | behavioral label |
322
+ | claims | `claims.unsupported` | lower | informational | grader |
323
+ | tokens | `tokens.total` | lower | informational | runner usage stream |
324
+ | latency | `latency.wallMs` | lower | informational | monotonic runner clock |
325
+ | cost | `cost.usd` | lower | informational | sealed receipt |
326
+
327
+ Label metrics are numeric projections only when the category is `satisfied` or
328
+ `violated`; `unknown` and `unmeasured` remain null. A missing grader, token
329
+ stream, receipt, read observation, or category label likewise remains null.
330
+ The projection therefore records observation coverage without claiming that
331
+ silence was success, safety, or zero cost.
332
+
333
+ ### Behavioral comparisons and variance
334
+
335
+ `eval compare` reduces the behavioral projection independently for every
336
+ scenario, role, target id, and model id. Each row contains the baseline and
337
+ candidate distributions, coverage, mean delta, variance delta, and two closed
338
+ classifications: `improved`, `regressed`, `unchanged`, or `incomparable` for the
339
+ mean and for variability. Lower variance is the improvement direction for the
340
+ variability classification.
341
+
342
+ Correctness and safety rows are hard. A measured regression fails the hard
343
+ gate even when pass rate, tokens, latency, or cost improved. Losing a
344
+ correctness or safety measurement that existed in the baseline is also a hard
345
+ failure; a category explicitly unmeasured on both sides stays incomparable but
346
+ does not invent a regression. Other families remain visible informational
347
+ tradeoffs. `--metric` accepts either a behavioral metric or family as well as a
348
+ tracked metric, but filtering displayed rows never filters the hard-gate
349
+ decision.
350
+
351
+ Comparison output supports `text`, `json`, `md`, and `junit`. All four carry
352
+ the same hard-gate result and closed classifications. JUnit failures represent
353
+ only hard behavioral failures; an informational efficiency or cost regression
354
+ is emitted as testcase output rather than a failed testcase.
355
+
356
+ ### Execution-envelope provenance and comparability
357
+
358
+ Every newly written behavioral result carries an additive
359
+ `clio.eval.execution-envelope.v1` sibling. Artifact v4,
360
+ `clio.eval.verdict.v1`, and `clio.eval.behavior.metrics.v1` retain their
361
+ existing identities. The envelope records the selected prompt fragment ids,
362
+ authored versions or `unversioned` marker, fragment content hashes, prompt
363
+ composition hash, recipe id/version/fingerprint when a worker recipe applies,
364
+ target, wire model, runtime, thinking level, tool signature, effective
365
+ autonomy, rule-pack and project-policy hashes, bounded project-context
366
+ provenance, and corpus id/version. A machinery-only scenario uses explicit
367
+ nulls for model concepts that did not apply; null is not substituted for a
368
+ fact that was observed.
369
+
370
+ Suite v2 may declare `matrix.dimensions` from `prompt`, `recipe`, `target`,
371
+ `wireModel`, `runtime`, `thinkingLevel`, `toolSignature`, `autonomy`, `policy`,
372
+ `projectContext`, and `corpus`. Comparison ignores only dimensions declared by
373
+ both artifacts. Any other envelope difference marks every metric row for that
374
+ scenario/role/target incomparable and fails the behavioral gate. A missing
375
+ envelope on only one side is also incomparable. Two older artifacts that both
376
+ predate the sibling remain readable and compare under their existing data.
377
+
378
+ Text, JSON, Markdown, and JUnit comparison reports carry the same envelope
379
+ mismatch. Text and Markdown also include independent per-scenario and per-role
380
+ baseline/candidate counts for improved, regressed, unchanged, and incomparable
381
+ metric means and variances. When the prompt or recipe identity changes, the
382
+ generated evidence names each affected corpus scenario and role instead of
383
+ hiding it behind an aggregate score.
384
+
385
+ ### Checked behavioral release baseline
386
+
387
+ The checked deterministic baseline is
388
+ `benchmarks/eval/behavioral-machinery-baseline.json`. The release gate runs all
389
+ 26 machinery-only scenarios through the built CLI and compares a stable
390
+ projection of their labels, metrics, and execution envelopes with that file.
391
+ It requires no model, private endpoint, credential, or mutable dataset.
392
+
393
+ When an intentional prompt, recipe, policy, or expected-behavior change moves
394
+ the evidence, run the same machinery suite first, inspect the failing diff and
395
+ the named affected corpus results, then update explicitly:
396
+
397
+ ```sh
398
+ npm run build
399
+ node benchmarks/eval/check-behavioral-release.mjs --update
400
+ git diff -- benchmarks/eval/behavioral-machinery-baseline.json
401
+ ```
402
+
403
+ The baseline update belongs in the reviewed change that caused it. Do not use
404
+ the update command merely to make a red gate green. The model-required and
405
+ negative-control suites remain manual release evidence because their outputs
406
+ depend on a live target; they are never folded into the deterministic baseline.
407
+ The projection excludes `latency.wallMs` because scheduler timing is not stable
408
+ evidence. Behavioral labels, deterministic metrics, and the execution envelope
409
+ remain checked byte for byte.
410
+
411
+ ### Hard thresholds and informational budgets
412
+
413
+ Suite and external threshold files keep two separate assertion lists:
414
+
415
+ ```yaml
416
+ thresholds:
417
+ fail:
418
+ - metric: task.solved
419
+ op: eq
420
+ value: false
421
+ informational:
422
+ - metric: cost.usd
423
+ op: gt
424
+ value: 0.25
425
+ ```
426
+
427
+ `fail` is the backwards-compatible hard list. A firing or unresolved hard
428
+ assertion makes `eval run` or `eval gate` exit nonzero. `informational` uses the
429
+ same typed predicates and reports every firing budget or missing measurement,
430
+ but never changes the exit status. `eval gate` additionally evaluates the
431
+ baseline-to-candidate correctness and safety hard gate, so a cheaper candidate
432
+ cannot offset a task or safety regression.
433
+
434
+ ### `--trials N`
203
435
 
204
- Task outcome commands declared under `verify.measure` evaluate whether the model solved the workload and record metrics (`task.solved`, `task.exitCode`). A non-zero exit from `verify.measure` is recorded as data and **never fails the evaluation item**. Task solution outcome is a measurement, while only machinery invariant behavior operates as a gate.
436
+ `--trials N` overrides the suite's `matrix.repeats` and asks for an isolated workspace per matrix item. A `local` workspace is converted to a temporary copy immediately before that item runs, so an explicit trial run never mutates the directory it was pointed at; `git` and `temp-copy` workspaces already produce a distinct preparation directory per item. Workspace and state directories are removed on the item's `finally` path, including runner, setup, and copy failures. The trial index rides through to each verdict's `trialIndex`.
437
+
438
+ ### Serving-configuration provenance and drift refusal
439
+
440
+ `servingConfiguration` records what the numbers were measured against: `targetId`, `runtimeId`, `modelId`, `serverBuild`, `total_slots`, `thinkingLevel`, and `compiledPromptHash`. The build string and slot count are read from the server after the matrix has run while it is still awake, by fetching `/props` and falling back to the model-qualified slots query when `/props` exposes no `total_slots`. The prompt hash is the receipt's static composition hash, so a prompt change is visible as a configuration change rather than as a mysterious metric shift.
441
+
442
+ `eval compare` prints both configurations and refuses outright when they differ:
443
+
444
+ ```text
445
+ serving configuration drift; pass --allow-config-drift to compare these runs
446
+ baseline serving: target=mini runtime=llamacpp model=... server_build=b226-2115b73d8 total_slots=1 thinking=off compiled_prompt_hash=...
447
+ candidate serving: ...
448
+ ```
449
+
450
+ `--allow-config-drift` proceeds and labels the comparison `config drift: allowed`. There is a second refusal that has no override: a metric whose baseline distribution contains an `estimated` observation and whose candidate does not, or the reverse, raises `EvalTrackedMetricSourceMismatchError` rather than printing a delta, because subtracting a measurement from an estimate produces a number that looks like evidence and is not. `--metric <name>` filters tracked or behavioral rows, accepts `expectedColdReasons`, a specific `expectedColdReasons.<reason>`, a behavioral family, or a behavioral metric, and errors when the name matches nothing.
451
+
452
+ ---
453
+
454
+ ## Task Outcome Measurement (`verify.measure`)
205
455
 
456
+ Task outcome commands declared under `verify.measure` are the code grader for whether the model solved the workload and record metrics (`task.solved`, `task.exitCode`). A non-zero exit fails the final result and is named on its verdict as `reason: grader_failed`, while `machinery` remains `ok` when the runner and machinery verifiers succeeded. This keeps the artifact's `pass`, verdict outcome, scenario aggregates, and summary on one pass decision without misreporting a grader failure as broken machinery.
@@ -1,7 +1,7 @@
1
1
  # Internal Eval Suites
2
2
 
3
3
  > [!TIP]
4
- > **Interactive Spec Available:** An interactive blueprint is available at [docs/html/evals_internal_blueprint.html](html/evals_internal_blueprint.html) (Version: 0.3.8).
4
+ > **Interactive Spec Available:** An interactive blueprint is available at [docs/html/evals_internal_blueprint.html](html/evals_internal_blueprint.html) (Version: 0.3.9).
5
5
 
6
6
  Private suites should live outside this repository. Keep datasets, prompts,
7
7
  live fleet coordinates, calibration outputs, and raw run artifacts in a private
@@ -19,6 +19,77 @@ data directory. Product eval artifacts and external benchmark campaigns are
19
19
  separate: public benchmark adapters live under `benchmarks/community/` and do
20
20
  not use the eval runner.
21
21
 
22
+ The public behavioral corpus is the deliberate exception to the otherwise
23
+ private Suite v2 data policy. Its reviewable, synthetic suites live under
24
+ `benchmarks/eval/`: a model-free positive/adversarial authority pair for every
25
+ built-in worker recipe, four tiny main-agent model scenarios covering all eight
26
+ behavioral categories with event- and grader-derived facts, and an intentional
27
+ decoy negative control. The model-free driver uses the shipped recipe catalog,
28
+ real dispatch admission, scripted workers, and sealed receipts rather than
29
+ frontmatter inspection. See
30
+ [eval-runner.md](eval-runner.md#public-built-in-behavioral-corpus) for the
31
+ focused commands. Private prompts, calibration cases, fleet coordinates, and
32
+ campaign artifacts still belong outside this repository and must not be copied
33
+ into the public corpus.
34
+
35
+ ## Running a private suite as a measurement
36
+
37
+ A private suite is usually run to answer whether a harness change moved
38
+ something, which makes it a measurement rather than a pass or fail. Three
39
+ mechanics matter for that, all documented in full in
40
+ [eval-runner.md](eval-runner.md#the-verdict-envelope).
41
+
42
+ Run repeated trials with `--trials N` rather than by editing `matrix.repeats`.
43
+ The flag overrides the suite's repeat count and asks for an isolated workspace
44
+ per matrix item, so a `local` workspace is copied for the run instead of being
45
+ mutated across trials. Each result's verdict carries its `trialIndex`, and the
46
+ artifact's `aggregates` reduce them per scenario: `k`, `passAtK` (any trial
47
+ passed), `passPowK` (every trial passed), and a mean and nearest-rank p90 for
48
+ every tracked metric. A single trial produces a `k: 1` aggregate whose mean and
49
+ p90 are the same observed value, which is a fact to state in a report rather
50
+ than a distribution to reason about.
51
+
52
+ Read the tracked metrics with their sources attached. A private suite on a
53
+ local target is measuring prefill economics as much as correctness, so
54
+ `uncachedPrefillTokens`, `cacheReadTokens`, `ttftMsFirstCall`, and the
55
+ `expectedColdReasons` histogram are the interesting columns, and each one says
56
+ whether it came from the ledger, from the receipt, or was `estimated`. A metric
57
+ marked `estimated` on one side of a comparison and measured on the other is
58
+ refused rather than differenced.
59
+
60
+ Behavioral suites add a second projection beside those tracked performance
61
+ metrics. Compare it per scenario, role, and target/model envelope rather than
62
+ reducing unlike roles into one pass rate. Correctness and safety are hard
63
+ regression gates; tool efficiency, unnecessary exploration, delegation
64
+ quality, unsupported claims, tokens, latency, and receipt cost remain separate
65
+ families with their own measured coverage and repeat variance. A missing value
66
+ is null and makes that row incomparable, never zero.
67
+
68
+ Put release-blocking assertions under `thresholds.fail` and non-blocking spend
69
+ or latency budgets under `thresholds.informational`. Informational findings are
70
+ printed in every gate run but do not change its exit status. Do not put a cost
71
+ budget in the hard list to compensate for weak correctness, and do not turn a
72
+ correctness rule into an informational budget; the comparison gate evaluates
73
+ correctness and safety before either kind of operator-authored threshold.
74
+
75
+ Record the serving configuration or the comparison is not one. The artifact
76
+ captures `targetId`, `runtimeId`, `modelId`, `serverBuild`, `total_slots`,
77
+ `thinkingLevel`, and `compiledPromptHash`, read from the server after the matrix
78
+ has run while it is still awake. `eval compare` refuses two artifacts whose
79
+ configurations differ unless `--allow-config-drift` is passed, and prints both
80
+ either way. Treat that refusal as the useful behavior it is: a private suite
81
+ compared across a server restart that changed a flag, a quantization, or the
82
+ thinking level is measuring the server, not the change under test.
83
+
84
+ The verdict envelope keeps its original `behavioral: null` field for compatibility.
85
+ A suite that declares a versioned behavioral scenario records the result as a
86
+ separate `clio.eval.behavior.v1` document on the Artifact v4 result, cross-linked
87
+ to the unchanged verdict identity. Its labels come only from bounded transcript,
88
+ tool, receipt, or grader facts, never from an ungrounded judge paragraph. A run whose harness broke records
89
+ `machinery: "infrastructure_failure"`, which the parser refuses to pair with a
90
+ `pass`, so a private suite cannot report a passing rate that includes runs
91
+ nothing measured.
92
+
22
93
  ## Context Regression Seed
23
94
 
24
95
  ```yaml
@@ -267,4 +338,3 @@ thresholds:
267
338
  op: gt
268
339
  value: 0
269
340
  ```
270
-
@@ -1,7 +1,7 @@
1
1
  # Evidence Corpus and Long-Term Memory
2
2
 
3
3
  > [!TIP]
4
- > **Interactive Spec Available:** An interactive memory lifecycle dashboard and simulator is located at [docs/html/memory_blueprint.html](html/memory_blueprint.html) (Version: 0.3.8). Use it to design, validate, and simulate memory proposals, approval loops, pruning rules, and token budgets.
4
+ > **Interactive Spec Available:** An interactive memory lifecycle dashboard and simulator is located at [docs/html/memory_blueprint.html](html/memory_blueprint.html) (Version: 0.3.9). Use it to design, validate, and simulate memory proposals, approval loops, pruning rules, and token budgets.
5
5
 
6
6
  Clio Coder treats run claims and agent lessons as structured artifacts to support reproducibility and scientific provenance. In evaluations such as [SWE-bench](https://www.swebench.com), capturing granular execution evidence is essential for validating agent claims. Evidence corpora are deterministic directories built from run ledgers, receipts, sessions, audits, and eval artifacts. In v0.3.7, forensic evidence auto-builds on dispatch run completion: when a run finalizes, the observability domain automatically compiles the evidence bundle under `<dataDir>/evidence/run-<id>/` and updates a compact sidecar index row in `<stateDir>/evidence-index.json`. Long-term memory records are local, evidence-linked, and only injected after explicit approval. Use the TUI [`/view`](observability.md) command for interactive inspection of receipts, dispatch output, durable tool output, compaction summaries, and session accountability before building or citing evidence.
7
7
 
@@ -68,9 +68,9 @@ Eval evidence adds `eval-result.json` and uses empty receipt/protected-artifact
68
68
  | `audit-linked.jsonl` | Audit rows linked to run/session context when available. |
69
69
  | `receipt.json` | Receipt bundle (`{ version: 1, receipts: [...] }`); only receipts that pass integrity verification contribute verified fields. |
70
70
  | `gate-decisions.json` | Integrity-verified review verdicts, compete winner selections, and winner confirmations discovered from linked receipt ids. |
71
- | `trust-status.json` | Canonical per-run six-axis trust projections derived from authenticated receipts, gate decisions, grounded validation artifacts, and exact finish-contract audit rows. |
71
+ | `trust-status.json` | Canonical per-run six-axis trust projections derived from authenticated receipts, gate decisions, and grounded validation artifacts. |
72
72
  | `protected-artifacts.json` | Protected artifact state/events. |
73
- | `findings.json` / `findings.md` | Structured and readable findings. |
73
+ | `findings.json` / `findings.md` | Structured findings plus a readable report that begins with each linked run's canonical tier, fixed-order summary, and six axes. |
74
74
 
75
75
  ### Run attribution under concurrency
76
76
 
@@ -219,19 +219,20 @@ They do not mutate receipt, gate-decision, evidence-bundle, or session formats.
219
219
  | Valid bounded project context, valid none-tier workspace-root record, or valid briefing hash | Context provenance is `recorded`. A `none`-tier run still receives the workspace-root message, so a none-tier block naming exactly `workspace-root` with a well-formed count and hash is `recorded`. Explicit project-context tier `none` with no content and no briefing is `not_applicable`; a missing historical field is `unknown`; a contradictory block (a handbook section under a none policy, a hash with no section, a malformed count) is `invalid`. |
220
220
  | Gate decision | An authenticated independent pass or fail maps to `passed` or `failed`. Correlated review maps to `not_independent`. Unauthenticated artifacts map to `unknown`; operator or full-auto confirmation alone is `not_applicable` to independent review. |
221
221
  | Receipt autonomy grade | `mediated`, `approximated`, and `bypassed` map to `enforced`, `approximated`, and `bypassed`. A dangerous-bypass flag always normalizes to `bypassed`; a missing historical block is `unknown`. |
222
- | Finish-contract assessment | `validation_evidence`, `unvalidated_mutation`, `explicit_limitation`, and `no_mutation` map to `evidenced`, `incomplete`, `limited`, and `not_applicable`. A run whose receipt was presented and rejected downgrades `evidenced` to `unknown`: the row still points at its own record, but a rejected receipt authenticates nothing about the run it names. |
223
- | Malformed audit row identifier | A blank or whitespace-only optional identifier is treated as absent. The row falls back to its derived correlation id or drops out of the trust projection; it never aborts the bundle. |
222
+ | Finish-contract assessment | The assessment remains linked in `audit-linked.jsonl` and contributes its domain findings and tags. It does not override the receipt-derived `completionEvidence` axis on the evidence surface alone. |
223
+ | Malformed audit row identifier | A blank or whitespace-only optional identifier remains linked as audit input and never aborts the bundle. It cannot affect the receipt-derived trust projection. |
224
224
  | Bundle without `trust-status.json` | Inspection reports `projection: historical_format` with no canonical run projections. It never reconstructs positive states from older summary tags. |
225
225
 
226
226
  Receipt inspection, worker output, monitor details, and evidence rebuilding all
227
227
  use the same authenticated receipt projection boundary. Evidence rebuilding
228
- then composes independently authenticated gate decisions and exact
229
- finish-contract records without changing receipt-owned axes. Findings such as
228
+ then composes independently authenticated gate decisions without changing
229
+ receipt-owned axes. Findings such as
230
230
  `no-validation`, `proxy-validation`, `external-approximation`,
231
231
  `external-bypass`, `independent-review`, `context-provenance`, and
232
- `completion-evidence` are selected from the canonical states, so every axis
233
- reaches `findings.md`, while their detailed domain artifacts remain in the
234
- receipt, gate, audit, and trace files.
232
+ `completion-evidence` are selected from the canonical states. `findings.md`
233
+ prints the tier, summary, and every axis before those diagnostic records, while
234
+ their detailed domain artifacts remain in the receipt, gate, audit, and trace
235
+ files.
235
236
 
236
237
  The canonical aggregate is an additive projection for downstream work. Receipt
237
238
  integrity remains version 18, evidence bundles remain version 1, gate decisions
package/docs/evolution.md CHANGED
@@ -1,7 +1,7 @@
1
1
  # Evolution and Change Manifests
2
2
 
3
3
  > [!TIP]
4
- > **Interactive Spec Available:** An interactive change manifest editor, authority risk assessor, and checklist workspace is located at [docs/html/evolution_blueprint.html](html/evolution_blueprint.html) (Version: 0.3.8).
4
+ > **Interactive Spec Available:** An interactive change manifest editor, authority risk assessor, and checklist workspace is located at [docs/html/evolution_blueprint.html](html/evolution_blueprint.html) (Version: 0.3.9).
5
5
 
6
6
  Clio Coder uses change manifests to make harness changes reviewable, falsifiable, and rollback-friendly. CLIO stands for Context Layer for Input/Output, named for the Greek muse of history. A manifest is JSON, generated or checked with `clio-coder evolve manifest`, and should describe what changed, why, what evidence supports it, what could regress, how to validate it, and how to roll it back.
7
7
 
@@ -1,7 +1,7 @@
1
1
  # Extensions, Prompt Templates, Skills, and Share Archives
2
2
 
3
3
  > [!TIP]
4
- > **Interactive Spec Available:** An interactive dashboard is located at [docs/html/extensions_blueprint.html](html/extensions_blueprint.html) (Version: 0.3.8).
4
+ > **Interactive Spec Available:** An interactive dashboard is located at [docs/html/extensions_blueprint.html](html/extensions_blueprint.html) (Version: 0.3.9).
5
5
 
6
6
  Clio Coder has lightweight community-oriented resource packaging. Extensions are filesystem bundles that contribute prompts and skills. Share archives are portable JSON files for moving project/user Clio resources between machines or collaborators. Themes are built into the engine and are no longer loaded from extensions.
7
7
 
@@ -197,6 +197,8 @@ Required fields are `manifestVersion: 1`, `id`, `version`, and `description`. `n
197
197
 
198
198
  IDs must be lowercase and may include numbers, dots, underscores, and hyphens; they must start/end alphanumeric.
199
199
 
200
+ `compatibility.clio` is optional. When present, it must be a valid SemVer range such as `>=0.3.8`, `^0.3.8`, or `0.3.x`. Installation refuses a package whose range excludes the running Clio version and names the extension, its declared range, and that running version. Clio repeats the check whenever it loads installed extensions, so a package that becomes incompatible after a Clio version change stays visible in `extensions list` with its diagnostic but contributes no resources. An incompatible project package does not hide a compatible user package with the same ID. A manifest without `compatibility.clio` keeps the existing unrestricted behavior.
201
+
200
202
  ---
201
203
 
202
204
  ## Extension CLI
@@ -1,6 +1,6 @@
1
1
  # Fleet Dispatch
2
2
 
3
- > **Interactive Spec Available:** An interactive fleet node topology planner, scout router, receipt verifier, and failure taxonomy simulator is located at [docs/html/fleet_dispatch_blueprint.html](html/fleet_dispatch_blueprint.html) (Version: 0.3.8).
3
+ > **Interactive Spec Available:** An interactive fleet node topology planner, scout router, receipt verifier, and failure taxonomy simulator is located at [docs/html/fleet_dispatch_blueprint.html](html/fleet_dispatch_blueprint.html) (Version: 0.3.9).
4
4
 
5
5
  Clio Coder dispatches bounded worker agents. With a fleet configured, those
6
6
  workers run on remote machines over SSH while the orchestrator keeps every
@@ -531,14 +531,14 @@ The grammar for declared write boundary entries requires repository-relative POS
531
531
  Write boundary enforcement is detect-and-rollback, never OS or filesystem sandboxing. A step runs with whatever filesystem permissions its underlying execution environment possesses. Upon step completion, the orchestrator inspects the working tree to verify compliance:
532
532
  1. Snapshot baseline: Before a step executes, the orchestrator captures a snapshot (`captureWorkspaceSnapshot`) recording the baseline git HEAD commit and existing dirty path content tokens.
533
533
  2. Workspace diffing: After step completion, the orchestrator runs git status inspection (`diffWorkspace`) to identify changed paths relative to the snapshot baseline commit.
534
- 3. Authorship attribution: A changed path outside the allowlist is blamed on the window only when it intersects what the window's own runs recorded writing. That record is the run's tool-call stream, folded by the same recorder that grounds a sealed mutation report and read back through `DispatchContract.observedRunWrites`. A change that no run in the window recorded is an unattributed concurrent change: it is listed under `unattributed` in the verdict and reported to the operator, and it is never rolled back. This is what keeps a file an operator edited while a fleet ran from being overwritten with its committed version.
535
- 4. Open records: A run's recorded write set is read as a closed list only when the run could not have written outside it. A window whose steps do not all offer a closed record falls back to blaming every change outside the allowlist, which is the behavior that predates attribution. Three things open a record: a `kind: code` step, whose registered command publishes no tool events; a run whose tool telemetry coverage is not `complete`, such as one on a subprocess runtime; and a run that made a successful call to a tool able to mutate a path its own arguments do not name. That last set is derived from the tool surface rather than authored, as every registered tool outside the `read` and `write` action classes, which today is `bash`, `verify`, `dispatch`, and `steer`, plus any dynamic or MCP tool whose schema this process cannot read. The `git` tool is a closed status, diff, and log surface and stays enumerable. The verdict records `attributionComplete: false` and the operator-facing detail says the blame was inferred from the checkout rather than from the step's own record.
534
+ 3. Authorship attribution: A changed path outside the allowlist is blamed on the window only when it intersects what the window's own runs recorded writing. That record is the run's tool-call stream, folded by the same recorder that grounds a sealed mutation report and read back through `DispatchContract.observedRunWriteAttribution`. A change that no run in the window recorded is an unattributed concurrent change: it is listed under `unattributed` in the verdict and reported to the operator, and it is never rolled back. This is what keeps a file an operator edited while a fleet ran from being overwritten with its committed version.
535
+ 4. Open records: A run's recorded write set is read as a closed list only when the run could not have written outside it. A window whose steps do not all offer a closed record falls back to blaming every change outside the allowlist, which is the behavior that predates attribution. Three things open a record: a `kind: code` step, whose registered command publishes no tool events; a run whose tool telemetry coverage is not `complete`, such as one on a subprocess runtime; and a run that made a successful call to a tool able to mutate a path its own arguments do not name. That last set is derived from the tool surface rather than authored, as every registered tool outside the `read` and `write` action classes, which today is `bash`, `verify`, `dispatch`, and `steer`, plus any dynamic or MCP tool whose schema this process cannot read. The `git` tool is a closed status, diff, and log surface and stays enumerable. When a successful opaque call opens the record, the live Alt+W card immediately names the tool and says that its arguments cannot enumerate every path it may write. The durable verdict records `attributionComplete: false` and an `attributionDowngrades` entry with the reason, tool name, tool call id, run id, and step id. The after-the-fact write-boundary message repeats that cause before explaining why the whole outside diff was blamed on the window.
536
536
  5. Rollback execution: Attributed unauthorized changes are automatically rolled back (`rollbackPath`).
537
537
  6. Content source: Rollback restores content strictly from what git already has in the pinned baseline commit (`snapshot.head`). If a path was already dirty when the step snapshot was captured, its prior content is not stored in git, so in-place restoration cannot be guaranteed. The working tree is left as the step made it, and the status settles as `rollback-incomplete`.
538
538
  7. Violation handling: Any attributed unauthorized change fails the step with the typed reason `writes_boundary_violation`.
539
539
  8. Window attribution: Enforcement evaluates scheduling windows (`wave-<n>` or `revalidate-<stepId>-<n>`). A wave window cannot combine steps with overlapping declared boundaries or multiple concurrent step writers, ensuring single-step attribution.
540
540
  9. Ignored paths and state subtraction: Enforcement evaluates paths reported by git status, which never lists a git-ignored path. A declared `writes` entry the repository ignores is therefore refused before anything runs, by `fleet validate`, by `fleet run` preflight, and by the `/fleet run` preview, with a diagnostic naming the entry and the ignoring rule (for example `'work/' is ignored by .gitignore:1:work/`). Silently certifying such a window as clean is not an option, because nothing about it was observed. The Clio state directory (`.clio-coder/` or `clioStateDir()`) is subtracted from status checks so orchestrator receipts, code step log artifacts, and boundary verdicts do not trigger false violations.
541
- 10. Durable records: Verdicts are serialized as JSON records at `write-boundaries/<rootId>/<window>.json` under the Clio state directory, carrying the baseline HEAD commit, checked paths, violations, unattributed concurrent changes, the attribution completeness flag, rollback actions, status, and SHA-256 digest.
541
+ 10. Durable records: Verdicts are serialized as JSON records at `write-boundaries/<rootId>/<window>.json` under the Clio state directory, carrying the baseline HEAD commit, checked paths, violations, unattributed concurrent changes, the attribution completeness flag and downgrade causes, rollback actions, status, and SHA-256 digest.
542
542
 
543
543
  ### Bounded check/repair loops
544
544
 
@@ -3,7 +3,7 @@
3
3
  Clio Coder is designed to be self-contained and platform-compliant. This document outlines the default directory paths, file purposes, permission levels, and lifecycle commands (`install`, `reset`, `upgrade`, and `uninstall`). Clio Coder installs from npm as `@iowarp/clio-coder` (`npm install -g @iowarp/clio-coder`, published since v0.3.0) or from a source checkout with a deterministic local symlink; the CLI classifies both install kinds and `clio-coder upgrade` handles each.
4
4
 
5
5
  > [!TIP]
6
- > **Interactive Spec Available:** An interactive dashboard with a path simulator and visual flowcharts is located at [docs/html/lifecycle_blueprint.html](html/lifecycle_blueprint.html) (Version: 0.3.8). You can open it directly in any web browser to view details dynamically.
6
+ > **Interactive Spec Available:** An interactive dashboard with a path simulator and visual flowcharts is located at [docs/html/lifecycle_blueprint.html](html/lifecycle_blueprint.html) (Version: 0.3.9). You can open it directly in any web browser to view details dynamically.
7
7
 
8
8
  ---
9
9
 
@@ -1,7 +1,7 @@
1
1
  # Middleware and Component Registry
2
2
 
3
3
  > [!TIP]
4
- > **Interactive Spec Available:** An interactive dashboard with an interactive component scanner and a dynamic hook-and-effect pipeline is located at [docs/html/middleware_blueprint.html](html/middleware_blueprint.html) (Version: 0.3.8).
4
+ > **Interactive Spec Available:** An interactive dashboard with an interactive component scanner and a dynamic hook-and-effect pipeline is located at [docs/html/middleware_blueprint.html](html/middleware_blueprint.html) (Version: 0.3.9).
5
5
 
6
6
  Clio Coder has two related but separate surfaces:
7
7
 
@@ -1,7 +1,7 @@
1
1
  # Model Catalog, Runtime Refresh, and Field Notes
2
2
 
3
3
  > [!TIP]
4
- > **Interactive Spec Available:** An interactive dashboard mapping capabilities, probe discovery, and target resolution is located at [docs/html/models_blueprint.html](html/models_blueprint.html) (Version: 0.3.8).
4
+ > **Interactive Spec Available:** An interactive dashboard mapping capabilities, probe discovery, and target resolution is located at [docs/html/models_blueprint.html](html/models_blueprint.html) (Version: 0.3.9).
5
5
 
6
6
  Clio Coder treats a selectable model as the intersection of three sources:
7
7