@iowarp/clio-coder 0.3.4 → 0.3.7

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (376) hide show
  1. package/CHANGELOG.md +65 -2
  2. package/CONTRIBUTING.md +6 -6
  3. package/README.md +16 -5
  4. package/dist/{acp-S5R4RR5B.js → acp-SK4MD6MM.js} +11 -11
  5. package/dist/{agents-P6DMMVZY.js → agents-2FN2K6ME.js} +33 -26
  6. package/dist/assets/codewiki.json +1 -1
  7. package/dist/{auth-2XCZLPKS.js → auth-QIYZWM5I.js} +15 -15
  8. package/dist/{chunk-EKMEHE4H.js → chunk-33YXPOE3.js} +2 -3
  9. package/dist/{chunk-VAWWTKDP.js → chunk-3HAPLH5M.js} +11 -11
  10. package/dist/{chunk-YCWGATWI.js → chunk-465YSENW.js} +3 -3
  11. package/dist/{chunk-4OC57DA6.js → chunk-4DGYLA73.js} +53 -2
  12. package/dist/{chunk-UZHIZC5S.js → chunk-4DWFMQDR.js} +61 -76
  13. package/dist/{chunk-ZWMF7253.js → chunk-5C3AQNDW.js} +328 -9
  14. package/dist/{chunk-22NAGB7X.js → chunk-5C77SEEY.js} +5 -94
  15. package/dist/{chunk-WPQLXFOZ.js → chunk-5FR74PWO.js} +3 -2
  16. package/dist/{chunk-35MKKU5R.js → chunk-5UJ6ECTS.js} +18 -10
  17. package/dist/{chunk-BRXQQJFP.js → chunk-6M7VS3J3.js} +571 -50
  18. package/dist/{chunk-QQK64KLB.js → chunk-6TUKSZVF.js} +141 -23
  19. package/dist/{chunk-N4CZJQRK.js → chunk-AB4XIIVB.js} +8 -6
  20. package/dist/{chunk-KRPY7NTG.js → chunk-BMWK7ZIZ.js} +14 -20
  21. package/dist/{chunk-4BPJXDWC.js → chunk-C4JBQ5SR.js} +30 -14
  22. package/dist/{chunk-ZYKPLLNQ.js → chunk-CEYBNUGC.js} +821 -83
  23. package/dist/{chunk-VEZEGCGW.js → chunk-D4MDIG46.js} +20 -18
  24. package/dist/chunk-DJNLUABN.js +843 -0
  25. package/dist/{chunk-BP4OYD6A.js → chunk-DMD2AGVS.js} +21 -2
  26. package/dist/{chunk-KOHPCX4K.js → chunk-DOOEX22V.js} +2 -2
  27. package/dist/chunk-DQA7QLMD.js +123 -0
  28. package/dist/chunk-DR52UMZW.js +21 -0
  29. package/dist/{chunk-3HZ5RWN2.js → chunk-EBEFWSGL.js} +9 -7
  30. package/dist/{chunk-EDRHSCIE.js → chunk-EELBMBT6.js} +128 -13
  31. package/dist/{chunk-HV5X7OR2.js → chunk-EOOQZZDE.js} +16 -14
  32. package/dist/{chunk-WR67VIZY.js → chunk-FOT2FX5J.js} +63 -5
  33. package/dist/{chunk-BPGS2WCQ.js → chunk-GEYXPTRF.js} +2 -1
  34. package/dist/{chunk-FYYLNIL5.js → chunk-GH5622CP.js} +2 -2
  35. package/dist/{chunk-BEY543CS.js → chunk-GOXNB3AO.js} +5 -2
  36. package/dist/chunk-GWS3VEIW.js +195 -0
  37. package/dist/{chunk-G4BMMOKF.js → chunk-HVDIIIQW.js} +2 -2
  38. package/dist/chunk-HWUFFB6L.js +83 -0
  39. package/dist/{chunk-X6COSD2O.js → chunk-J7PIKKWC.js} +8 -436
  40. package/dist/{chunk-NILBFAPG.js → chunk-JNXPYBB4.js} +2 -2
  41. package/dist/{chunk-4VP4KH3K.js → chunk-JRIO5UD2.js} +4 -4
  42. package/dist/{chunk-K6WL7QZT.js → chunk-JTSEDYVQ.js} +7 -7
  43. package/dist/{chunk-QKMUKYO7.js → chunk-KCMKRQX4.js} +236 -84
  44. package/dist/chunk-KZ2H5X4G.js +1026 -0
  45. package/dist/{chunk-A2GZF7DC.js → chunk-LADCF22A.js} +13 -13
  46. package/dist/chunk-LCGCVYZ4.js +57 -0
  47. package/dist/chunk-M4AKACEO.js +382 -0
  48. package/dist/{chunk-POHLU5DW.js → chunk-M6L6IDJG.js} +3 -3
  49. package/dist/{chunk-4JUF2NNX.js → chunk-MXI6J5JF.js} +7 -7
  50. package/dist/{chunk-X4RCMKVQ.js → chunk-NDINPTJ4.js} +2 -2
  51. package/dist/{chunk-TTNYS3EA.js → chunk-OB5HIGJY.js} +1 -1
  52. package/dist/{chunk-7RXG6QRZ.js → chunk-OBMAI2DP.js} +61 -840
  53. package/dist/{chunk-5M54SPOL.js → chunk-ODFEOB4F.js} +161 -5
  54. package/dist/chunk-PD3MESLB.js +242 -0
  55. package/dist/{chunk-ED4KHGC3.js → chunk-PPAMZ32Z.js} +9 -2
  56. package/dist/{chunk-VMNQ6OZA.js → chunk-QCTRSGHQ.js} +963 -786
  57. package/dist/chunk-RVG5JXAL.js +41 -0
  58. package/dist/{chunk-RD5U66HV.js → chunk-SROCI7ZU.js} +7 -7
  59. package/dist/{chunk-MFFY33HR.js → chunk-THKY7CD7.js} +466 -205
  60. package/dist/{chunk-34475P3I.js → chunk-TSHXZTOQ.js} +5 -4
  61. package/dist/{chunk-PCZJO5TI.js → chunk-UFQ3F4FW.js} +13 -178
  62. package/dist/{chunk-AD2SYQYC.js → chunk-UHXRNZ2J.js} +121 -3
  63. package/dist/chunk-UND3GU2L.js +103 -0
  64. package/dist/{chunk-QQL5RT5M.js → chunk-UUANF5CR.js} +2323 -2114
  65. package/dist/{chunk-VJWL6YS5.js → chunk-UUVG37B4.js} +2 -2
  66. package/dist/chunk-UVDSQ6LW.js +472 -0
  67. package/dist/{chunk-QWU7ZBO7.js → chunk-VQNODYQ4.js} +215 -56
  68. package/dist/chunk-VREKEFLL.js +37 -0
  69. package/dist/{chunk-2TZWSW76.js → chunk-WHGPSPT5.js} +2 -2
  70. package/dist/{chunk-TW3WDMVS.js → chunk-WHJYKASB.js} +2 -2
  71. package/dist/{chunk-MEQ45TQ4.js → chunk-WJHBC77E.js} +21 -7
  72. package/dist/{chunk-HXG4IURW.js → chunk-X2KV5FXT.js} +2 -2
  73. package/dist/{chunk-YHZX5GEU.js → chunk-XAKHZX5N.js} +2 -2
  74. package/dist/{chunk-2LZI5CAG.js → chunk-XEGB6BCN.js} +228 -36
  75. package/dist/{chunk-E25LMLRW.js → chunk-YD734TPH.js} +2 -2
  76. package/dist/{verifiers-4UUM6TEE.js → chunk-YTYFXUI3.js} +121 -372
  77. package/dist/{chunk-3JLKSKD7.js → chunk-ZGH7FGS5.js} +17 -7
  78. package/dist/{chunk-VSNATDE6.js → chunk-ZZMN5OM4.js} +2 -2
  79. package/dist/cli/index.js +34 -32
  80. package/dist/{clio-J5JIOIDS.js → clio-WBVQEBKO.js} +7 -7
  81. package/dist/{code-nav-AXCXSBHX.js → code-nav-FGGFIE7L.js} +7 -7
  82. package/dist/codewiki/build-worker.js +4 -4
  83. package/dist/{components-KELWS457.js → components-F7OEATSO.js} +5 -5
  84. package/dist/{config-OEBMIN2U.js → config-TRBL3RCF.js} +48 -41
  85. package/dist/{configure-PUQOSIXQ.js → configure-OLCVPHNM.js} +17 -17
  86. package/dist/{context-URSXPBCK.js → context-MJIJ6GOX.js} +12 -12
  87. package/dist/{context-EKDCKUUZ.js → context-WFPKQSM6.js} +26 -9
  88. package/dist/{context-MGSE4Z2T.js → context-XEWE3MOJ.js} +44 -37
  89. package/dist/{context-clear-KDAJRNUK.js → context-clear-KNOS2JPB.js} +44 -37
  90. package/dist/{context-index-BZ4UYMTC.js → context-index-SSR5ECNE.js} +3 -3
  91. package/dist/{context-working-set-SBKMPPI2.js → context-working-set-EUXAZI6N.js} +14 -13
  92. package/dist/{dispatch-runner-MSWN72NK.js → dispatch-runner-B7MTOVKL.js} +321 -60
  93. package/dist/{docs-2C2LTVT2.js → docs-FLJTIDSE.js} +5 -5
  94. package/dist/{doctor-7BSE27PJ.js → doctor-RN4YKO2X.js} +15 -15
  95. package/dist/{eval-IZGDOO4H.js → eval-RUBJVSNQ.js} +52 -236
  96. package/dist/{evidence-SR7WXB5B.js → evidence-JZNBUOQZ.js} +39 -33
  97. package/dist/{evolve-K7VE2CBX.js → evolve-FJVC4KKI.js} +39 -33
  98. package/dist/{extensions-QVDOHDGJ.js → extensions-IQL36S7K.js} +5 -5
  99. package/dist/{fleet-7XMJNQNF.js → fleet-BDKYJFCP.js} +243 -370
  100. package/dist/fleet-commands-ZFIWZSB3.js +70 -0
  101. package/dist/fleet-graph-Y6HPXIVF.js +125 -0
  102. package/dist/fleet-new-RDVJLHHH.js +48 -0
  103. package/dist/{fleet-preflight-AQNAH644.js → fleet-preflight-BHSNPBMH.js} +2 -2
  104. package/dist/fleet-validate-BIYREGIK.js +79 -0
  105. package/dist/{init-JGNPAYXT.js → init-LQUB5COQ.js} +57 -48
  106. package/dist/library-NJAHIGG4.js +217 -0
  107. package/dist/memory-OG6HOYKM.js +472 -0
  108. package/dist/{models-ZMMLFJNN.js → models-5ZG5XY7J.js} +23 -22
  109. package/dist/{monitor-2F3T5KHP.js → monitor-TJ7AMTGB.js} +69 -35
  110. package/dist/{orchestrator-ORHT43JB.js → orchestrator-WZYB54DM.js} +4868 -1189
  111. package/dist/{paths-UXLN5YYZ.js → paths-XUC7GS6E.js} +5 -5
  112. package/dist/{reset-NXGTYNUO.js → reset-PXQT45IY.js} +8 -8
  113. package/dist/{run-RF4WJGMT.js → run-FQ74YF62.js} +82 -62
  114. package/dist/{share-UT3W6E4M.js → share-FW7SVCL3.js} +34 -10
  115. package/dist/{skills-PSACKC5Q.js → skills-7E7IRB3R.js} +25 -9
  116. package/dist/{skills-eval-WJSI55RZ.js → skills-eval-LI75W6OK.js} +43 -35
  117. package/dist/{targets-PIIRAOYS.js → targets-4CIFKCTW.js} +27 -24
  118. package/dist/{terminal-lease-ULWXWNVY.js → terminal-lease-WUZY7ZV5.js} +5 -4
  119. package/dist/{uninstall-FZCQCDKC.js → uninstall-7FV7IP4E.js} +5 -5
  120. package/dist/{upgrade-346TZ6AV.js → upgrade-K2HVIVMQ.js} +21 -20
  121. package/dist/{usage-6KKXR32N.js → usage-GTZELZQX.js} +159 -59
  122. package/dist/verifiers-RLAHT27O.js +336 -0
  123. package/dist/{verify-X5HDROLA.js → verify-BX3BRKH5.js} +7 -6
  124. package/dist/{wiki-generate-7STOCIFZ.js → wiki-generate-ASIFASCN.js} +58 -48
  125. package/dist/worker/entry.js +98 -84
  126. package/dist/{workspace-G4ZWUIPR.js → workspace-ZJ6BFM3Q.js} +4 -4
  127. package/docs/README.md +4 -3
  128. package/docs/acp.md +1 -1
  129. package/docs/alcf-provider.md +1 -1
  130. package/docs/architecture.md +2 -2
  131. package/docs/artifact-placement.md +1 -2
  132. package/docs/artifact-versions.md +10 -6
  133. package/docs/built-in-agents.md +26 -2
  134. package/docs/capacity-and-scheduling.md +1 -1
  135. package/docs/commands-and-modes.md +90 -8
  136. package/docs/configuration-and-targets.md +90 -2
  137. package/docs/context-engine.md +4 -2
  138. package/docs/context-working-set.md +4 -4
  139. package/docs/development-pipeline.md +1 -1
  140. package/docs/dispatch-architecture-rationale.md +1 -1
  141. package/docs/documentation-coverage.md +4 -4
  142. package/docs/documentation-guide.md +4 -4
  143. package/docs/eval-runner.md +1 -1
  144. package/docs/evals-internal.md +4 -45
  145. package/docs/evidence-and-memory.md +70 -10
  146. package/docs/evolution.md +1 -1
  147. package/docs/exit-codes-and-output.md +4 -1
  148. package/docs/extensions-and-sharing.md +6 -2
  149. package/docs/fleet-demo-runbook.md +2 -2
  150. package/docs/fleet-dispatch.md +224 -11
  151. package/docs/git-commit-provenance.md +2 -2
  152. package/docs/glossary.md +1 -1
  153. package/docs/installation-and-lifecycle.md +2 -2
  154. package/docs/middleware-and-components.md +20 -2
  155. package/docs/model-catalog.md +1 -1
  156. package/docs/observability.md +55 -8
  157. package/docs/proactive-memory.md +26 -16
  158. package/docs/prompt-envelope-and-tools.md +4 -2
  159. package/docs/provider-adapter-cookbook.md +1 -1
  160. package/docs/release-cut-checklist.md +83 -65
  161. package/docs/resource-library.md +59 -0
  162. package/docs/safety-model.md +29 -7
  163. package/docs/scientific-validation.md +3 -3
  164. package/docs/session-lifecycle.md +37 -1
  165. package/docs/skills-marketplace.md +16 -3
  166. package/docs/tool-usage.md +14 -7
  167. package/docs/trace-store.md +1 -1
  168. package/docs/troubleshooting.md +1 -1
  169. package/docs/tui-design.md +38 -4
  170. package/docs/worker-dispatch-mechanics.md +3 -3
  171. package/package.json +7 -4
  172. package/src/cli/agents.ts +2 -3
  173. package/src/cli/argv.ts +14 -1
  174. package/src/cli/fleet-commands.ts +37 -0
  175. package/src/cli/fleet-graph.ts +102 -0
  176. package/src/cli/fleet-new.ts +36 -0
  177. package/src/cli/fleet-preflight.ts +121 -0
  178. package/src/cli/fleet-validate.ts +30 -0
  179. package/src/cli/fleet.ts +188 -335
  180. package/src/cli/index.ts +4 -2
  181. package/src/cli/library.ts +190 -0
  182. package/src/cli/memory.ts +272 -10
  183. package/src/cli/modes/json-stream.ts +2 -2
  184. package/src/cli/modes/print.ts +12 -1
  185. package/src/cli/run.ts +22 -2
  186. package/src/cli/share.ts +13 -1
  187. package/src/cli/targets.ts +12 -3
  188. package/src/cli/usage.ts +160 -20
  189. package/src/core/bus-events.ts +7 -0
  190. package/src/core/commit-attribution.ts +4 -4
  191. package/src/core/config.ts +130 -0
  192. package/src/core/defaults.ts +81 -0
  193. package/src/core/response-model-id.ts +134 -0
  194. package/src/core/toml.ts +62 -0
  195. package/src/core/workspace-files.ts +0 -1
  196. package/src/domains/agents/builtins/architect.md +2 -1
  197. package/src/domains/agents/builtins/oracle.md +33 -0
  198. package/src/domains/agents/catalog.ts +18 -5
  199. package/src/domains/agents/fleet-contract.ts +278 -16
  200. package/src/domains/agents/index.ts +14 -0
  201. package/src/domains/agents/recipe.ts +54 -14
  202. package/src/domains/agents/result-contract.ts +242 -5
  203. package/src/domains/config/classify.ts +4 -0
  204. package/src/domains/context/bootstrap.ts +36 -27
  205. package/src/domains/context/project-metadata.ts +19 -63
  206. package/src/domains/context/prompt-context.ts +8 -0
  207. package/src/domains/context/working-set/policies/index.ts +3 -4
  208. package/src/domains/dispatch/active-route-planner.ts +14 -0
  209. package/src/domains/dispatch/backoff.ts +2 -1
  210. package/src/domains/dispatch/budget-envelope.ts +396 -0
  211. package/src/domains/dispatch/capability-match.ts +1 -0
  212. package/src/domains/dispatch/checkout-writer-lease.ts +175 -0
  213. package/src/domains/dispatch/contract.ts +36 -0
  214. package/src/domains/dispatch/delegation-plan.ts +167 -0
  215. package/src/domains/dispatch/execution-plan.ts +76 -5
  216. package/src/domains/dispatch/execution-role.ts +3 -1
  217. package/src/domains/dispatch/execution-scheduler.ts +183 -67
  218. package/src/domains/dispatch/extension.ts +339 -36
  219. package/src/domains/dispatch/fleet-gate.ts +14 -0
  220. package/src/domains/dispatch/fleet-plan.ts +63 -3
  221. package/src/domains/dispatch/fleet-run.ts +737 -0
  222. package/src/domains/dispatch/gate-role-prompts.ts +9 -0
  223. package/src/domains/dispatch/host-verification.ts +178 -0
  224. package/src/domains/dispatch/index.ts +38 -0
  225. package/src/domains/dispatch/intent.ts +159 -0
  226. package/src/domains/dispatch/orphan-recovery.ts +1 -0
  227. package/src/domains/dispatch/receipt-integrity.ts +12 -4
  228. package/src/domains/dispatch/state.ts +37 -3
  229. package/src/domains/dispatch/types.ts +61 -9
  230. package/src/domains/dispatch/validation.ts +80 -6
  231. package/src/domains/dispatch/worker-spawn.ts +14 -3
  232. package/src/domains/eval/metrics/evidence.ts +0 -116
  233. package/src/domains/eval/metrics/invariants.ts +1 -1
  234. package/src/domains/eval/runners/clio-run.ts +1 -10
  235. package/src/domains/eval/runners/external-command.ts +2 -29
  236. package/src/domains/eval/schema/suite.ts +0 -7
  237. package/src/domains/eval/suites/run.ts +1 -7
  238. package/src/domains/evidence/trust-status.ts +10 -1
  239. package/src/domains/memory/index.ts +22 -0
  240. package/src/domains/memory/operations.ts +58 -1
  241. package/src/domains/memory/promotion.ts +281 -0
  242. package/src/domains/memory/prompt-section.ts +25 -5
  243. package/src/domains/memory/proposal.ts +51 -7
  244. package/src/domains/memory/task-bank.ts +3 -2
  245. package/src/domains/memory/task-memory-handoff.ts +181 -24
  246. package/src/domains/memory/task-memory-policy.ts +3 -1
  247. package/src/domains/memory/types.ts +37 -0
  248. package/src/domains/memory/validate.ts +178 -0
  249. package/src/domains/middleware/index.ts +15 -0
  250. package/src/domains/middleware/memory-intervention.ts +35 -25
  251. package/src/domains/middleware/runtime.ts +6 -0
  252. package/src/domains/middleware/skills-reminder.ts +19 -4
  253. package/src/domains/middleware/stalled-turn.ts +43 -1
  254. package/src/domains/middleware/types.ts +10 -0
  255. package/src/domains/middleware/watchdog.ts +281 -0
  256. package/src/domains/observability/contract.ts +9 -2
  257. package/src/domains/observability/cost.ts +31 -4
  258. package/src/domains/observability/extension.ts +2 -2
  259. package/src/domains/observability/index.ts +10 -0
  260. package/src/domains/observability/out-of-turn-usage.ts +223 -0
  261. package/src/domains/providers/index.ts +3 -0
  262. package/src/domains/providers/model-discovery.ts +9 -0
  263. package/src/domains/providers/runtime-resolution.ts +38 -1
  264. package/src/domains/providers/runtimes/common/probe-helpers.ts +97 -16
  265. package/src/domains/providers/types/context-window-slots.ts +18 -0
  266. package/src/domains/providers/types/runtime-descriptor.ts +3 -1
  267. package/src/domains/resources/index.ts +20 -0
  268. package/src/domains/resources/library.ts +326 -0
  269. package/src/domains/resources/skills/marketplace.ts +37 -12
  270. package/src/domains/safety/call-target.ts +211 -14
  271. package/src/domains/safety/decision-presentation.ts +268 -0
  272. package/src/domains/safety/redaction.ts +73 -0
  273. package/src/domains/session/context-ledger.ts +10 -1
  274. package/src/domains/session/decision-board.ts +4 -0
  275. package/src/domains/session/entries.ts +3 -0
  276. package/src/domains/session/handoff.ts +629 -0
  277. package/src/domains/session/history.ts +68 -19
  278. package/src/domains/session/usage.ts +24 -7
  279. package/src/domains/share/archive.ts +67 -2
  280. package/src/engine/acp/event-mapper.ts +7 -0
  281. package/src/engine/acp/server.ts +29 -2
  282. package/src/engine/apis/lmstudio.ts +25 -4
  283. package/src/engine/apis/openai-completions.ts +147 -22
  284. package/src/engine/claude/sdk-runtime.ts +8 -2
  285. package/src/engine/claude/tool-safety.ts +13 -0
  286. package/src/engine/loop-guard.ts +27 -3
  287. package/src/engine/worker-events.ts +4 -3
  288. package/src/engine/worker-runtime.ts +59 -54
  289. package/src/entry/orchestrator.ts +55 -1
  290. package/src/interactive/bus-notices.ts +26 -0
  291. package/src/interactive/chat-loop-messages.ts +22 -0
  292. package/src/interactive/chat-loop.ts +248 -1
  293. package/src/interactive/chat-renderer.ts +41 -3
  294. package/src/interactive/clio-editor.ts +44 -7
  295. package/src/interactive/context-overlay.ts +43 -5
  296. package/src/interactive/cost-overlay.ts +70 -11
  297. package/src/interactive/council-dispatch.ts +30 -0
  298. package/src/interactive/council-grid.ts +213 -0
  299. package/src/interactive/council.ts +99 -0
  300. package/src/interactive/dispatch-board.ts +471 -50
  301. package/src/interactive/fleet-run-preview.ts +307 -0
  302. package/src/interactive/footer/notifications.ts +219 -0
  303. package/src/interactive/footer/widgets.ts +13 -0
  304. package/src/interactive/handoff-round.ts +56 -0
  305. package/src/interactive/interactive-application.ts +49 -2
  306. package/src/interactive/interactive-event-projection.ts +9 -1
  307. package/src/interactive/interactive-input-runtime.ts +11 -1
  308. package/src/interactive/interactive-presentation.ts +11 -1
  309. package/src/interactive/interactive-slash-runtime.ts +52 -2
  310. package/src/interactive/interactive-subscriptions.ts +14 -2
  311. package/src/interactive/memory-overlay.ts +89 -4
  312. package/src/interactive/oracle.ts +179 -0
  313. package/src/interactive/overlay-ask-user-lifecycle.ts +7 -1
  314. package/src/interactive/overlay-frame.ts +5 -2
  315. package/src/interactive/overlay-general-openers.ts +230 -2
  316. package/src/interactive/overlay-key-routing.ts +58 -2
  317. package/src/interactive/overlay-lifecycle.ts +52 -5
  318. package/src/interactive/overlay-permission-lifecycle.ts +33 -8
  319. package/src/interactive/overlay-resource-openers.ts +11 -3
  320. package/src/interactive/overlay-session-lifecycle.ts +234 -2
  321. package/src/interactive/overlay-transitions.ts +11 -0
  322. package/src/interactive/overlays/ask-user.ts +74 -30
  323. package/src/interactive/overlays/decisions.ts +3 -1
  324. package/src/interactive/overlays/fleet-run-approval.ts +208 -0
  325. package/src/interactive/overlays/handoff-review.ts +185 -0
  326. package/src/interactive/overlays/library-install-confirm.ts +151 -0
  327. package/src/interactive/overlays/list-overlay.ts +168 -2
  328. package/src/interactive/overlays/settings.ts +101 -4
  329. package/src/interactive/overlays/side-question.ts +139 -0
  330. package/src/interactive/overlays/skills-hub.ts +401 -15
  331. package/src/interactive/permission-hint.ts +35 -0
  332. package/src/interactive/permission-overlay.ts +95 -45
  333. package/src/interactive/renderers/tool-execution.ts +19 -49
  334. package/src/interactive/session-last-turn.ts +8 -1
  335. package/src/interactive/session-usage-reseed.ts +36 -10
  336. package/src/interactive/side-question.ts +171 -0
  337. package/src/interactive/slash-commands.ts +434 -7
  338. package/src/interactive/slash-spec.ts +19 -6
  339. package/src/interactive/status/summary.ts +5 -0
  340. package/src/interactive/status/types.ts +5 -0
  341. package/src/interactive/terminal-lease.ts +1 -0
  342. package/src/interactive/theme/tokens.ts +30 -0
  343. package/src/interactive/turn-context.ts +96 -23
  344. package/src/interactive/turn-middleware.ts +16 -1
  345. package/src/interactive/turn-runtime.ts +37 -8
  346. package/src/interactive/turn-state.ts +3 -0
  347. package/src/interactive/watchdog-run.ts +75 -0
  348. package/src/interactive/worker-progress.ts +440 -0
  349. package/src/interactive/worker-share.ts +56 -1
  350. package/src/interactive/worker-stream.ts +58 -110
  351. package/src/tools/agent-tools.ts +28 -3
  352. package/src/tools/ask-user.ts +21 -1
  353. package/src/tools/bootstrap.ts +3 -0
  354. package/src/tools/compete-worktrees.ts +13 -79
  355. package/src/tools/context/index.ts +2 -2
  356. package/src/tools/dispatch-admission.ts +242 -8
  357. package/src/tools/dispatch-arguments.ts +65 -1
  358. package/src/tools/dispatch-event-text.ts +19 -0
  359. package/src/tools/dispatch-plan.ts +136 -6
  360. package/src/tools/dispatch-runner.ts +319 -13
  361. package/src/tools/dispatch-types.ts +20 -1
  362. package/src/tools/dispatch.ts +96 -3
  363. package/src/tools/monitor.ts +31 -0
  364. package/src/tools/profiles.ts +18 -4
  365. package/src/tools/registry.ts +15 -5
  366. package/src/tools/result-disposition.ts +156 -0
  367. package/src/tools/result-shaping.ts +59 -1
  368. package/src/tools/task-worktree.ts +238 -0
  369. package/src/tools/verify/authoring.ts +116 -55
  370. package/src/tools/verify/scripts.ts +62 -0
  371. package/src/tools/worker-evidence.ts +21 -1
  372. package/src/worker/spec-contract.ts +44 -3
  373. package/dist/chunk-EFADSJET.js +0 -18
  374. package/dist/chunk-HC4CLZ2Y.js +0 -68
  375. package/dist/memory-4ALKDJ4Q.js +0 -246
  376. package/src/domains/eval/metrics/chaos-stream.ts +0 -93
@@ -1,7 +1,7 @@
1
1
  # Commands and Modes
2
2
 
3
3
  > [!TIP]
4
- > **Interactive Spec Available:** An interactive dashboard is located at [docs/html/commands_blueprint.html](html/commands_blueprint.html) (Version: 0.3.4).
4
+ > **Interactive Spec Available:** An interactive dashboard is located at [docs/html/commands_blueprint.html](html/commands_blueprint.html) (Version: 0.3.7).
5
5
 
6
6
 
7
7
  Clio Coder is a terminal-first alpha harness. This page keeps the command
@@ -56,7 +56,7 @@ For process exit codes, stdout deliverable guarantees, and machine-readable JSON
56
56
  | `clio-coder dev components diff --from <a> --to <b> [--json]` | Compare component snapshots. |
57
57
  | `clio-coder evidence build\|inspect\|list` | Build and inspect deterministic evidence artifacts. |
58
58
  | `clio-coder eval validate\|run\|report\|compare\|gate` | Validate, run, report, compare, and gate local evaluation suites (Suite v2). |
59
- | `clio-coder memory list\|propose\|approve\|reject\|prune` | Manage scoped, evidence-linked memory records. |
59
+ | `clio-coder memory list\|propose\|promote\|approve\|reject\|prune` | Manage scoped, evidence-linked memory records. |
60
60
  | `clio-coder trace runs [--db PATH] [--limit N] [--json]` | List runs recorded in the durable trace mirror beside the ledger. |
61
61
  | `clio-coder trace phases <runId> [--db PATH]` | Show one run's recorded phases. |
62
62
  | `clio-coder trace tail <runId> [--follow] [--db PATH]` | Tail one run's recorded events; `--follow` streams as they land. |
@@ -145,20 +145,24 @@ The registry table below lists the available interactive slash commands. On a ba
145
145
  | `/quit` | `/quit` | Exit Clio Coder |
146
146
  | `/help` | `/help [query]` | Open the interactive help center showing commands and keys |
147
147
  | `/skill` | `/skill [name] [task]` | Open the Skills Hub or invoke a skill |
148
+ | `/library` | `/library [kind]` | Open the Skills Hub on a resource library tab |
148
149
  | `/prompts` | `/prompts` | List prompt templates |
149
150
  | `/extensions` | `/extensions` | List installed extensions |
150
151
  | `/interop` | `/interop` | Review other coding agents detected on this machine |
151
152
  | `/share` | `/share [runId] \| /share export <path> \| /share import [--dry-run] [--force] <path>` | Share a worker result with the main agent, or export and import Clio archives |
152
153
  | `/run` | `/run [--agent-profile <profile>] [--runtime <runtimeId>] [--target <id>] [--model <id>] [--thinking <level>] [--tool-profile <minimal-local\|science-local\|full-agent>] [--require <cap>] [--share] <agent> <task>` | Run a fleet agent |
153
154
  | `/delegate` | `/delegate [--share] <agent-id> <task>` | Run an ACP delegation agent |
155
+ | `/btw` | `/btw <question>` | Ask a side question that never enters the session transcript |
156
+ | `/oracle` | `/oracle <question>` | Ask a read-only advisor to challenge a question against this session's settled decisions |
157
+ | `/council` | `/council [--roster <name>] [--rounds <n>] [--synthesis <judge\|vote\|none>] <task>` | Ask a roster of read-only members the same task, with an optional vote or judge synthesis |
154
158
  | `/agents` | `/agents` | List Clio agents and ACP delegation agents |
155
159
  | `/targets` | `/targets` | Open Settings → Targets: health, use, connect, probe, remove |
156
160
  | `/cost` | `/cost` | Show session token and cost totals |
157
161
  | `/context` | `/context compact [instructions] \| /context recall <ref> \| /context init \| /context refresh \| /context reset` | Context hub: window overlay plus compact, recall, init, refresh, and reset |
158
- | `/fleet` | `/fleet` | Open Settings → Fleet: defaults, profiles, agent bindings, nodes |
162
+ | `/fleet` | `/fleet run [--var <key=value>] <name>` | Open Settings → Fleet, or run a fleet contract with an approval preview |
159
163
  | `/decisions` | `/decisions` | Show settled interview decisions and operator revisions |
160
164
  | `/tasks` | `/tasks add <text> \| /tasks hand <id> \| /tasks done <id> \| /tasks drop <id>` | Show the session board or manage project operator tasks |
161
- | `/memory` | `/memory seed` | Inspect task memory or seed it from the newest handoff |
165
+ | `/memory` | `/memory seed` | Inspect, promote, or seed task memory |
162
166
  | `/view` | `/view [filter] \| /view verify <runId>` | Browse session artifacts and verify receipts |
163
167
  | `/thinking` | `/thinking [level]` | Set the chat thinking level, or open Settings → Orchestrator |
164
168
  | `/output` | `/output [verbosity]` | Set transcript detail (minimal, default, verbose), or open Settings → Terminal |
@@ -167,6 +171,7 @@ The registry table below lists the available interactive slash commands. On a ba
167
171
  | `/settings` | `/settings [section]` | Open interactive settings |
168
172
  | `/resume` | `/resume` | Resume a past session |
169
173
  | `/new` | `/new` | Start a fresh session |
174
+ | `/handoff` | `/handoff <goal>` | Hand this session's working state to a fresh session for a stated goal |
170
175
  | `/tree` | `/tree` | Open session tree navigator |
171
176
  | `/fork` | `/fork` | Fork from an assistant turn |
172
177
  | `/export` | `/export [path]` | Export a self-contained HTML transcript by default; a `.md` path writes Markdown |
@@ -190,6 +195,72 @@ There are no slash-command aliases. `/context compact`, `/quit`, `/model`,
190
195
  spellings stay errors that name `/help` instead of guessing which operation the
191
196
  operator intended.
192
197
 
198
+ `/btw <question>` runs one model round beside the session and renders the answer
199
+ in an overlay. It sends the same compiled message history the next turn would
200
+ send, as read-only input, under a short system instruction saying this is a side
201
+ question, with no tools. Nothing about the round is appended: not the session
202
+ JSONL, not the transcript panel, not the context ledger, not the task board. That
203
+ is the point of it. A fleet run briefs its workers from the transcript, so a
204
+ question the operator asks to orient themselves mid-run would otherwise become
205
+ context every worker inherits. Esc closes the overlay, and cancels the round if it
206
+ is still streaming. `/btw` during an in-flight turn is refused with a notice
207
+ rather than queued, because a side question answered after the run it was asked
208
+ during has already missed its moment. The round's token usage still shows in
209
+ `/cost`, labeled as a side question, because it was a real call and cost real
210
+ money; it is deliberately not counted as a turn.
211
+
212
+ `/council [--roster <name>] [--rounds <n>] [--synthesis judge|vote|none] <task>`
213
+ asks a roster of two to five read-only members the same task and puts the group on
214
+ the Fleet Runs board as one card. It owns no dispatch path of its own: the command
215
+ builds dispatch-tool arguments and admits them through the tool registry, so a
216
+ supervised autonomy level parks the call and the approval overlay names every
217
+ member's label, target, model, node, round count, and synthesis mode before
218
+ anything runs. Members are pinned to read-only autonomy and the council tool
219
+ surface by admission, exactly as they are for a council the model asks for.
220
+
221
+ `--roster` names a `workers.rosters` entry. Without it the command takes
222
+ `workers.rosters.default` when that roster exists, and with neither it refuses
223
+ and names the setting to declare. A roster that is the only one configured is
224
+ still not the default: seating a council from whichever roster happens to be
225
+ present would run models the operator never chose. `--rounds` accepts one to
226
+ three and `--synthesis` accepts `judge`, `vote`, or `none`, which are the tool's
227
+ own bounds, enforced where the operator typed them so a council is never refused
228
+ after its plan has already been shown. `/council` during an in-flight turn is
229
+ refused with a notice rather than queued, for the same reason `/fleet run` is: an
230
+ approved plan describes the workspace as it stands. Nothing the members produce
231
+ enters the main agent's context until an operator runs `/share`.
232
+
233
+ `/handoff <goal>` carries this session's working state into a fresh session for a
234
+ goal the operator states. The goal is required and gated: a goal shorter than 12
235
+ characters is refused, and so is one of a small stoplist of non-goals such as
236
+ "continue", "next", or "resume". Both refusals name the rule they enforce, because
237
+ "keep going" is exactly the instruction a handoff exists to replace.
238
+
239
+ One model round then runs on the same out-of-turn seam `/btw` uses. It reads the
240
+ compiled message history the next turn would send, sends no tools, and answers
241
+ with JSON validated against a fixed response schema of decisions, facts, files,
242
+ commands, and open questions. Every list and every string is bounded; output over
243
+ a bound is truncated with a visible marker and the document names each bound that
244
+ fired, so nothing is cut silently and an over-eager answer is never a refusal.
245
+
246
+ Every file path the model names is checked against this session's read ledger and
247
+ never against the filesystem. Paths the session did not touch are dropped and
248
+ listed in the document under their own heading so the operator can see what the
249
+ model invented. Extracted decisions are merged with the session's settled decision
250
+ board, and the board wins. The result is one Markdown document opened for review:
251
+ Enter accepts it, `e` hands it to `$EDITOR`, and Esc cancels the whole handoff with
252
+ nothing written anywhere.
253
+
254
+ On accept, Clio mints a new session, writes the reviewed document into it as
255
+ bounded data labelled as a handoff from the old session id, and replays the old
256
+ session's skill activations so loaded skills carry forward. The document is never
257
+ written as a fabricated user turn. The old session is left untouched apart from one
258
+ terminal note recording the target session id. A handoff is a session operation
259
+ throughout: it writes no memory promotion candidate and never calls the task-memory
260
+ bank. `/handoff` during an in-flight turn is refused with a notice rather than
261
+ queued, because a document summarizing a session that is still moving would be
262
+ wrong by the time it was read.
263
+
193
264
  The `/resume` picker accepts Page Up and Page Down to move by its 12 visible rows. Arrow keys continue to move one session at a time, and typing continues to filter the list.
194
265
 
195
266
  Only active commands run. Typing anything command-shaped that the registry does
@@ -224,7 +295,7 @@ Configuration lives in one place: the `/settings` overlay. `/settings <section>`
224
295
 
225
296
  Settings → Targets presents an operational console table (`HEALTH`, `ID`, `ROLES`, `RUNTIME`, `LATENCY`) with an in-place action/detail drawer for URL, default model, last probe error, and reachability. `Enter` opens actions for `Use` (switches active chat target and rebases model), `Connect` (runs the API-key or OAuth flow then probes), `Probe`, and `Remove` (with preflight analysis of affected routes/profiles). Probing runs live when the overlay opens or when explicitly requested. Target creation is initiated via `clio-coder targets add`.
226
297
 
227
- Settings → Fleet is an entity workbench organized with dim group headers (`Defaults`, `Profiles`, `Agent routes`, `Placement`). Dispatched worker defaults and profile rows render as compact summaries (`fast-local node-a/example-coder-model high auto`), drilling into fields (`target`, `model`, `thinkingLevel`, `node`) on `Enter`. Profile removal is a named destructive action with affected-route preflight. Running and retrying dispatches live in the `Alt+W` Fleet Runs board, which also steers and cancels them.
298
+ Settings → Fleet is an entity workbench organized with dim group headers (`Defaults`, `Profiles`, `Agent routes`, `Placement`). Dispatched worker defaults and profile rows render as compact summaries (`fast-local node-a/example-coder-model high auto`), drilling into fields (`target`, `model`, `thinkingLevel`, `node`) on `Enter`. Profile removal is a named destructive action with affected-route preflight. Running and retrying dispatches live in the `Alt+W` Fleet Runs board, which also steers and cancels them. `Enter` opens the selected run's worker detail: the phase, the running call with its redacted action descriptor, and the bounded tail of the worker's own prose.
228
299
 
229
300
  `/run` and `/delegate` put the worker's answer on screen. Both echo the typed
230
301
  line dim above the block, then stream the run into the transcript as an attributed
@@ -256,6 +327,15 @@ operator steering whose run id names a receipt it can read, so a model that
256
327
  never dispatched the run does not discard it as unattributed output. A turn
257
328
  that only relays a shared note does not trip the unbacked-worker-claim
258
329
  advisory.
330
+ A council run shares as a council. `/share <synthesis runId>` brings the whole
331
+ `council-report` in as one bounded block: every final-round member's answer under
332
+ its roster label, each with its verdict when it declared one, then the synthesis
333
+ line naming the mode, the verdict, the tally, and the judge run when there was
334
+ one. `/share <member runId>` brings that one member's answer in under its roster
335
+ label, so a single voice never reaches the main agent as an unattributed one. A
336
+ synthesis run whose sealed text does not parse as a report is shared verbatim
337
+ rather than dropped, because the operator named that run.
338
+
259
339
  `/new` resets the transcript and the pool bare `/share` draws from, so a run
260
340
  from the previous session cannot be shared into the new one. Worker tool
261
341
  arguments never cross at all: the transcript carries tool names only, the same
@@ -350,7 +430,7 @@ editor reserves and can be rebound through `settings.yaml.keybindings`.
350
430
  | `Alt+U` | Toggle the footer dashboard between compact (quiet 2-zone) and expanded (4-zone urgency) layouts. |
351
431
  | `Alt+L` | Open the model and targets selector. |
352
432
  | `Alt+J` / `Alt+K` | Cycle forward / backward through the scoped model set (when empty, displays a notice directing the operator to `/scoped-models`). |
353
- | `Alt+W` | Toggle the Fleet Runs board (task, run ID, live telemetry, retry, and terminal history). |
433
+ | `Alt+W` | Toggle the Fleet Runs board (task, run ID, live telemetry, retry, and terminal history). Inside it, `Enter` opens the selected run's live worker detail, `s` steers, and `x` cancels. |
354
434
  | `Alt+B` | Open the composite session and operator task board (`/tasks`). Approved application-boundary override of editor word-back. |
355
435
  | `Alt+D` | Open the settled interview decision board (`/decisions`). Approved application-boundary override of editor word-delete. |
356
436
  | `Alt+S` / `Ctrl+Alt+B` | Convert an active attached dispatch to a detached background batch. |
@@ -415,7 +495,9 @@ Tool and command execution is governed by:
415
495
  - **Safety Net:** Granular rule packs loaded from `damage-control-rules.yaml`, project policies, and protected artifact paths; always on, identical at every autonomy level.
416
496
  - **Autonomy Mapping:** Once the net passes a call, the level decides whether it runs, asks, or is denied. See [safety-model.md](safety-model.md) for the full matrix.
417
497
 
418
- When an action asks for confirmation, whether from a safety-net rail or from the autonomy level, the call parks and the TUI displays a queued permission dialog whose `Asked by:` line names the asking axis. The operator can approve or deny that single action without changing the level.
498
+ When an action asks for confirmation, whether from a safety-net rail or from the autonomy level, the call parks and three surfaces say so at once. The transcript row reads `⏸ awaiting approval` with `action ·`, `axis ·`, and `target ·` lines under it; the footer phase pill reads `⏸ confirm`; and a consequence-tier dialog opens with the tool, target, action, authenticated requester, one-shot authority, reversibility, and deny and stop effects. Titles distinguish workspace authority, outward consequences, safety-net confirmation, system changes, and worker escalations. The dialog sits at bottom center with five rows reserved for the composer and footer, and it re-anchors on resize. The composer rail switches to `CONFIRM` and repeats the keys while the prompt owns the keyboard.
499
+
500
+ The keys are the same on both surfaces: `Enter` allows this one call, `Esc` denies it, and `s` denies it and stops the turn so nothing asks again. `Enter` allows only from an empty composer. While the composer holds a draft, the habitual send key does nothing, the rail and the dialog footer read `[Backspace] clear draft` instead of `[Enter] allow`, and only the deletion keys (`Backspace`, `Delete`, `Ctrl+U`, `Ctrl+W`, `Ctrl+K`) reach the editor until the draft is gone. Every other key is swallowed. A call that parks while another overlay holds the screen is announced with an `[approval]` notice and the dialog opens as soon as that overlay closes; the dialog lays itself out for any terminal width, so no width is too narrow for it. Approving or denying never changes the level.
419
501
 
420
502
  Notice vocabulary, one prefix per mechanism: `[safety-net]` for level-independent blocks, `[approval]` for parked calls, `[autonomy]` for read-only denials, and `[middleware]` for hook diagnostics.
421
503
 
@@ -465,7 +547,7 @@ to execute through the existing engine worker path, the sanctioned Claude Code w
465
547
  | --- | --- |
466
548
  | `npm run ci` | Local and GitHub PR gate: typecheck, lint, skills pin check, build, the deterministic test suite, and the trace-viewer suite. |
467
549
  | `npm run ci:release` | Maintainer release gate: `npm run ci`, then the `check-release` dist and packaging audit. |
468
- | `npm run live:smoke -- --target <id>` | One real headless turn against a configured target. Add `--delegation` for the `opencode` and `copilot` ACP agents. The other operator-run drivers (`live:recon`, `live:fleet-dispatch`, `live:tui`) are listed in `benchmarks/internal/README.md`. |
550
+ | `npm run live:smoke -- --target <id>` | One real headless turn against a configured target. Add `--delegation` for the `opencode` and `copilot` ACP agents. The other operator-run drivers (`live:fleet-dispatch`, `live:tui`, `live:home`) are listed in `benchmarks/internal/README.md`. |
469
551
  | `npm run typecheck` | Strict TypeScript pass. |
470
552
  | `npm run lint` | Biome checks plus `scripts/check-hygiene.ts`, which runs the boundary invariants, the skills pin check, and the README and docs drift rules. |
471
553
  | `npm run test` | Contract and smoke tests through the sharded runner. |
@@ -1,7 +1,7 @@
1
1
  # Configuration, Targets, Runtimes, and Auth
2
2
 
3
3
  > [!TIP]
4
- > **Interactive Spec Available:** An interactive configuration validator, target resolver, and CLI command generator is located at [docs/html/configuration_blueprint.html](html/configuration_blueprint.html) (Version: 0.3.4).
4
+ > **Interactive Spec Available:** An interactive configuration validator, target resolver, and CLI command generator is located at [docs/html/configuration_blueprint.html](html/configuration_blueprint.html) (Version: 0.3.7).
5
5
 
6
6
  Clio Coder is target-first: chat and fleet dispatch resolve through configured targets in `settings.yaml`, not through provider-specific ad hoc flags. Chat and print targets are HTTP and native engine-backed runtimes. Fleet dispatch can also target the sanctioned Claude Code subscription runtimes described below.
7
7
 
@@ -31,6 +31,8 @@ Default config file:
31
31
 
32
32
  Role contents: config holds user-authored files (settings, credentials, agents, skills, prompts, extensions, runtimes); data holds durable artifacts (memory, evidence, evals); state holds machine-produced session state (sessions, audit, receipts, runs.json, recent-models.json, install.json, interop.json, interviews, scratch); cache holds disposable derived files.
33
33
 
34
+ The `library` settings block configures the private resource catalog. `library.catalog` is an optional path and defaults to `<configDir>/library.yaml`. `library.remote` is an optional git remote URL, and the catalog repository must name that git remote `library`. `library.sync` defaults to `false`, which makes sync and push refuse before spawning git. `library.confirmedRemote` is written by `clio-coder library remote confirm <url>` and must exactly match `library.remote` before sync or push can run. Confirmation sets both values when `library.remote` is unset and refuses a differing configured URL with `library_remote_mismatch`. See [resource-library.md](resource-library.md).
35
+
34
36
  `clio-coder paths --json` prints the resolved directories and is the single source of truth for scripts.
35
37
 
36
38
  ---
@@ -174,6 +176,17 @@ workers:
174
176
  model: your-model-id
175
177
  thinkingLevel: off
176
178
  profiles: {}
179
+ rosters:
180
+ design:
181
+ members:
182
+ - label: local-a
183
+ target: local-lmstudio
184
+ model: your-model-id
185
+ thinking: medium
186
+ color: accent
187
+ - label: local-b
188
+ target: local-vllm
189
+ color: "#5ba8ff"
177
190
  agentBindings: {}
178
191
  maxRetries: 2
179
192
  onPermission: deny
@@ -208,6 +221,11 @@ terminal:
208
221
  tuiMode: regular # regular terminal scrollback or fullscreen sticky layout
209
222
  fullscreenScrollbar: auto # hidden, auto, or always in fullscreen mode
210
223
  smoothStreaming: off # off, conservative auto, or explicit on
224
+ notify: false # content-free desktop notification, interactive TTY only
225
+ watchdog:
226
+ enabled: false # opt-in read-only review of every mutating turn
227
+ # target: local-lmstudio # route the review at a cheap model
228
+ # cadenceToolCalls: 20 # also review every N tool calls inside a turn
211
229
  skills:
212
230
  trustProjectCompatRoots: false
213
231
  delegation:
@@ -306,6 +324,17 @@ LM Studio can require bearer authentication for its HTTP APIs
306
324
 
307
325
  A model id on an LM Studio target is resolved against that host's loaded instances. A key with a loaded instance is never sent bare (which would JIT-load a second copy). An instance id reported loaded by two configured LM Studio targets on different hosts is an LM Link peer projection. When a bare model key is requested and multiple instances of it are loaded, Clio selects an instance in this order: the target's configured `defaultModel`, then an instance not cross-listed by another configured LM Studio target, and finally the first loaded instance. This behavior tracks issue #113.
308
326
 
327
+ When the selected instance is also loaded on a peer, a request may be answered by that peer (#185). Clio separates the requested model id, the response observation, and the model id used for accounting. Every new assistant ledger entry carries `responseModelIdObservation` in one of these explicit shapes:
328
+
329
+ | State | Meaning | Accounting attribution |
330
+ | --- | --- | --- |
331
+ | `{ "state": "reported", "reportedModelId": "<id>" }` | Clio observed an OpenAI-compatible event stream and the provider reported a model id. | The reported id. |
332
+ | `{ "state": "not-reported" }` | Clio observed the event stream and it contained no model id. | `unknown`, because the provider did not identify the responding model. |
333
+ | `{ "state": "not-observed" }` | This provider path did not expose response model-id presence to the stream tap. | A differing `responseModel` when available, otherwise the requested model id. |
334
+ | `{ "state": "legacy-difference-only", "differingModelId": "<id>" }` or the same shape with `null` | The ledger predates #193 and recorded only whether the response `model` differed from the request. This state is produced while reading historical rows; new rows do not write it. | The historical differing id when available, otherwise the requested model id. |
335
+
336
+ The adapter retains `responseModel` as the differing response id because providers outside the stream tap still supply that fact. `clio-coder usage report` emits `attributedModelId`, `requestedModelIds`, and `responseModelIdObservationCounts`. Its text table and the `/cost` overlay use the labels `attributed model`, `requested model ids`, and `response model id observation`; requested ids are printed as ids rather than as `same`. The footer's last-turn line uses `response model id observation <state>`, with the id after `reported` or a historical `legacy difference-only` state. Dispatch receipt `upstreamResponses` entries carry `requestedModelId`, `responseModelIdObservation`, `differingResponseModelId`, and `providerResponseId`. The peer warning is said once per process per distinct fact (target, requested id, resolved instance, peer set), not once per turn.
337
+
309
338
 
310
339
  Prompt-template overrides, system prompts, GPU-offload ratios, KV-cache quantization, parallel slots,
311
340
  context checkpoints, and speculative-decoding variants are not writable through this Clio settings
@@ -456,7 +485,8 @@ The Settings Center organizes all configuration under four non-selectable group
456
485
  | **RUNTIME** | Budget (`budget`) | `budget.sessionCeilingUsd`, `defaults.maxTokens`, and `budget.concurrency` (restart required). |
457
486
  | **RUNTIME** | Compaction (`compaction`) | `compaction.auto`, `compaction.threshold`, and `compaction.excludeLastTurns`. |
458
487
  | **RUNTIME** | Retry (`retry`) | `retry.enabled`, `retry.maxRetries`, `retry.baseDelayMs`, and `retry.maxDelayMs`. |
459
- | **EXPERIENCE** | Terminal (`terminal`) | `terminal.showTerminalProgress`, `terminal.outputVerbosity` (`minimal`, `default`, `verbose`), `terminal.tuiMode` (`regular`, `fullscreen`), `terminal.fullscreenScrollbar` (`hidden`, `auto`, `always`), `terminal.smoothStreaming` (`off`, `auto`, `on`), and `theme`. |
488
+ | **EXPERIENCE** | Terminal (`terminal`) | `terminal.showTerminalProgress`, `terminal.outputVerbosity` (`minimal`, `default`, `verbose`), `terminal.tuiMode` (`regular`, `fullscreen`), `terminal.fullscreenScrollbar` (`hidden`, `auto`, `always`), `terminal.smoothStreaming` (`off`, `auto`, `on`), `terminal.notify`, and `theme`. |
489
+ | **EXPERIENCE** | Watchdog (`watchdog`) | `watchdog.enabled`, `watchdog.target`, and `watchdog.cadenceToolCalls`. The two optional keys are editable text rows that render their absence as `(session target)` and `(turn end only)`; submitting an empty value removes the key from `settings.yaml` rather than storing a blank. |
460
490
  | **EXPERIENCE** | Advanced (`advanced`) | `runtimePlugins`, `attribution.gitCommits`, `compaction.model`, `compaction.systemPrompt`, `delegation.defaults.connectTimeoutMs`, `delegation.defaults.turnTimeoutMs`, `delegation.defaults.permissionTimeoutMs`, `keybindings`, and `delegation.agents`. |
461
491
 
462
492
  `retry.streamStallMs` has no Settings Center row; edit it in `settings.yaml`.
@@ -506,6 +536,10 @@ Label to config path mapping:
506
536
  | TUI mode | `terminal.tuiMode` (`regular` or `fullscreen`, restart required) |
507
537
  | Fullscreen scrollbar | `terminal.fullscreenScrollbar` (`hidden`, `auto`, or `always`, restart required) |
508
538
  | Smooth streaming | `terminal.smoothStreaming` (`off`, `auto`, or `on`, live) |
539
+ | Desktop notifications | `terminal.notify` |
540
+ | Turn-end watchdog | `watchdog.enabled` |
541
+ | Watchdog target | `watchdog.target` (blank clears the key) |
542
+ | Watchdog cadence (tools) | `watchdog.cadenceToolCalls` (integer ≥ 1; blank clears the key) |
509
543
  | Theme | `theme` |
510
544
  | Runtime plugins | `runtimePlugins` |
511
545
  | Clio commit provenance | `attribution.gitCommits` (`enabled` or `disabled`, live) |
@@ -544,6 +578,15 @@ These are saved defaults, not a live control surface. See [Live routing vs saved
544
578
 
545
579
  ### Safety and worker policy
546
580
 
581
+ `workers.rosters.<name>.members` defines council membership beside
582
+ `workers.profiles`. Every member accepts `label`, `target`, and the optional
583
+ keys `model`, `thinking`, and `color`. Labels must match
584
+ `[a-z][a-z0-9_-]{0,31}` and must be unique inside the roster. A roster contains
585
+ two to five members. Colors accept a theme token such as `accent`, `success`,
586
+ or `reason`, or a six-digit hexadecimal value such as `#5ba8ff`. Unknown roster
587
+ and member keys are rejected during configuration load. The existing settings
588
+ watcher validates and publishes roster changes with every other hot reload.
589
+
547
590
  | Key | Default | Validation | When it applies |
548
591
  | --- | --- | --- | --- |
549
592
  | `autonomy` | `auto-edit` | `read-only`, `suggest`, `auto-edit`, `full-auto` | immediately |
@@ -553,8 +596,13 @@ These are saved defaults, not a live control surface. See [Live routing vs saved
553
596
  | `workers.maxRetries` | `2` | integer ≥ 0 | next dispatch |
554
597
  | `workers.resilienceCooldownMs` | `15000` | integer ≥ 0 | next dispatch |
555
598
  | `workers.profiles` | `{}` | map of profile name to a target/model/thinking choice | next dispatch |
599
+ | `workers.rosters` | `{}` | map of roster name to 2 to 5 council members | next dispatch |
556
600
  | `workers.agentBindings` | `{}` | map of agent id to a key present in `workers.profiles` | next dispatch |
557
601
  | `skills.trustProjectCompatRoots` | `false` | boolean | restart |
602
+ | `library.catalog` | `null` | string or null | immediately |
603
+ | `library.remote` | `null` | string or null | immediately |
604
+ | `library.confirmedRemote` | `null` | string or null | immediately |
605
+ | `library.sync` | `false` | boolean | immediately |
558
606
 
559
607
  ### Git commit provenance
560
608
 
@@ -612,6 +660,34 @@ Generic provider and transport errors are classified by transient retry rules, i
612
660
  | `memory.intervention.maxTokens` | `400` | integer ≥ 1 | next turn |
613
661
  | `memory.intervention.timeoutMs` | `180000` | integer ≥ 1 | next turn |
614
662
 
663
+ ### Turn-end watchdog
664
+
665
+ | Key | Default | Validation | When it applies |
666
+ | --- | --- | --- | --- |
667
+ | `watchdog.enabled` | `false` | boolean | immediately |
668
+ | `watchdog.target` | unset | non-empty target id | immediately |
669
+ | `watchdog.cadenceToolCalls` | unset | integer ≥ 1 | immediately |
670
+
671
+ The watchdog is off by default because it spends one worker run per mutating
672
+ turn. With `enabled: true`, a turn that changed the tree is handed to one
673
+ read-only `verifier` run briefed with the turn's coalesced diff and the task
674
+ board's current scope. Its blockers become one transcript notice naming the
675
+ count and the first three failed checks, and nothing else: it never follows up,
676
+ never queues a turn, and never mutates. A passing report emits nothing at all. A
677
+ turn with no file mutations never fires it.
678
+
679
+ `watchdog.target` routes the run at a named target, which is how a cheap local
680
+ model reviews turns run on a subscription route; unset, the run takes the
681
+ session's active target. `watchdog.cadenceToolCalls: N` additionally fires the
682
+ watchdog after every N tool calls inside a turn, with the same diff-and-scope
683
+ briefing, so mid-turn scope drift is visible before the turn ends. At most one
684
+ watchdog run is in flight at a time; a trigger that arrives while one is running
685
+ is dropped and counted rather than queued. Headless and ACP runs never fire the
686
+ watchdog regardless of the setting, because neither has an operator reading a
687
+ transcript. The block has its own Settings Center section under EXPERIENCE ›
688
+ Watchdog; clearing the target or the cadence row removes that key from
689
+ `settings.yaml` rather than writing an empty value.
690
+
615
691
  ### Delegation
616
692
 
617
693
  | Key | Default | Validation | When it applies |
@@ -634,10 +710,22 @@ Generic provider and transport errors are classified by transient retry rules, i
634
710
  | `terminal.tuiMode` | `regular` | `regular`, `fullscreen` | restart |
635
711
  | `terminal.fullscreenScrollbar` | `auto` | `hidden`, `auto`, `always` | restart |
636
712
  | `terminal.smoothStreaming` | `off` | `off`, `auto`, `on` | immediately |
713
+ | `terminal.notify` | `false` | boolean | immediately |
637
714
  | `modelSelector.favorites` | `[]` | list of strings | immediately |
638
715
  | `modelSelector.recentLimit` | `12` | integer ≥ 1 | immediately |
639
716
  | `keybindings` | `{}` | map of binding id to a key string or list of them | restart |
640
717
 
718
+ `terminal.notify` turns on a content-free desktop notification for the three
719
+ moments an operator is waiting: a turn ends, a detached fleet batch settles, and
720
+ a worker permission or `ask_user` request parks. The payload is fixed. The title
721
+ is always `clio-coder` and the body comes from a closed vocabulary (`turn
722
+ finished`, `batch <shortId> settled`, `approval needed`), so no prompt text, file
723
+ path, or model output ever leaves the process in a notification. Clio emits OSC
724
+ 777 by default and OSC 9 on iTerm2, Windows Terminal, and ConEmu, never both for
725
+ one event. Headless, ACP, and non-TTY runs never emit one regardless of the
726
+ setting. The knob has a Settings Center row under EXPERIENCE › Terminal,
727
+ labeled `Desktop notifications`.
728
+
641
729
  Recently selected models are runtime state and live in `recent-models.json` under the state directory, not here. A `state.recentModels` key in `settings.yaml` is an unknown-key error.
642
730
 
643
731
  ### Structural and catalog keys
@@ -1,7 +1,7 @@
1
1
  # Context Engine
2
2
 
3
3
  > [!TIP]
4
- > **Interactive Spec Available:** An interactive dashboard is located at [docs/html/context_blueprint.html](html/context_blueprint.html) (Version: 0.3.4).
4
+ > **Interactive Spec Available:** An interactive dashboard is located at [docs/html/context_blueprint.html](html/context_blueprint.html) (Version: 0.3.7).
5
5
 
6
6
  Clio Coder tracks context pressure, records per-turn snapshots, and protects the provider context with bounded tool results plus single-threshold compaction.
7
7
 
@@ -19,6 +19,8 @@ Local-native runtimes use a recommended minimum desired window of 128,000 tokens
19
19
 
20
20
  The `/context` overlay states which layer answered, next to the token total: `loaded`, `probed`, `configured`, `declared`, or `assumed`.
21
21
 
22
+ A probed llama.cpp window is the share one request gets, not the server's total. llama.cpp splits `--ctx-size` evenly across `--parallel` slots unless `--kv-unified` is set, so a server started with `--ctx-size 786432 --parallel 4 --no-kv-unified` admits 196,608 tokens per request, and that is the figure autocompact and the meter plan against. The probe reads the flags (long and short forms, `-c`, `-np`, `-kvu`, and the last of `--kv-unified` or `--no-kv-unified` given) off the router's per-model status, keeps the split on the model's discovery state, and `/context` prints the derivation next to the share: `196,608 (786,432 / 4 slots)`. `clio-coder targets` does the same in its `ctx` note for the target's default model and adds a probe note naming the flags.
23
+
22
24
  ## Token accounting and snapshots
23
25
 
24
26
  The estimator in `context-accounting.ts` uses a four-characters-per-token family for hot-path accounting. It estimates system prompt, tools, messages, pending input, and runtime categories without calling a model tokenizer on every TUI refresh.
@@ -190,7 +192,7 @@ In Git workspaces, the indexer uses the same visible file set across full builds
190
192
  incremental updates, fingerprints, and project profiles: tracked files plus
191
193
  untracked, unignored work in progress. It excludes symlinks, submodule gitlinks,
192
194
  generated output, scratch space, and local-state directories such as `.git`,
193
- `.clio-coder`, `.superpowers`, `.codex`, `.claude`, `.clio-coder-benchmark`, `node_modules`,
195
+ `.clio-coder`, `.superpowers`, `.codex`, `.claude`, `node_modules`,
194
196
  `dist`, `build`, `coverage`, virtualenvs, `target`, and `vendor`. Non-Git
195
197
  workspaces use a bounded filesystem walk with the same directory exclusions.
196
198
  Source coverage spans TypeScript, JavaScript, Python, Rust, Go, C, C++, CUDA
@@ -5,7 +5,7 @@ The working set is the part of the session ledger the model actually receives on
5
5
  Source of truth is `src/domains/context/working-set/` (`contract.ts`, `fold.ts`, `project.ts`, `marker.ts`, `protect.ts`, `engine.ts`, `recall.ts`, `policies/`), the ledger records in `src/domains/session/entries.ts`, and the compaction stage in `src/interactive/turn-context.ts` (`runAutoCompact`).
6
6
 
7
7
  > [!WARNING]
8
- > This is an experimental community alpha surface. The default policy is `structural-v1`, chosen from the replay tables under `benchmarks/results/context-replay/`. `age-horizon` reproduces the selection Clio made before this layer existed and stays available.
8
+ > This is an experimental community alpha surface. The default policy is `structural-v1`; `age-horizon` reproduces the selection Clio made before this layer existed and stays available.
9
9
 
10
10
  ## Vocabulary
11
11
 
@@ -170,7 +170,7 @@ context:
170
170
 
171
171
  ## What the operator sees
172
172
 
173
- - **`/context` overlay.** A working-set section under the category legend: the policy that produced the most recent event, evicted item count, evicted tokens, event count, recall count, and churn. Evicted tokens render as one line after the legend rather than as a meter category, because they are outside the window rather than a slice of it.
173
+ - **`/context` overlay.** A working-set section under the category legend: the configured policy with its state (`policy structural-v1 · no events yet` until the first event, `disabled` when `context.workingSet.enabled` is off, and `(last event by <policy>)` when the setting changed after an event), evicted item count, evicted tokens, event count, recall count, and churn. Evicted tokens render as one line after the legend rather than as a meter category, because they are outside the window rather than a slice of it.
174
174
  - **Transcript.** An evicted tool row keeps its full body and gains a dim `evicted · <reason>` tag. The transcript shows the ledger, never the projection, so `/resume`, `/tree`, `/fork`, and the HTML export are unaffected by eviction.
175
175
  - **`/context recall <ref>`.** Prints the ref, why it was evicted, the token count, and the offload pointer when there is one, followed by the original body. Transcript only.
176
176
  - **Prompt cache line.** Every applied event stamps `working_set_evict` on the next assistant entry's `promptCache.expectedColdReasons`. When the last settled run came back cold for that reason, the overlay adds `last cold turn: working-set eviction (expected)` and drops the shell-reused-but-backend-cold warning, because the cold turn is explained rather than surprising.
@@ -181,14 +181,14 @@ context:
181
181
  These are tracked follow-ups, not available behavior:
182
182
 
183
183
  - **Auto-readmission.** Nothing brings an evicted body back on its own. There are no path fingerprints and no registry of what the model is likely to need next.
184
- - **Cost model and deferred scheduling.** Pressure is the only trigger. There is no break-even horizon, no deferred eviction plan, and no piggybacking beyond the fact that the working-set stage already runs first inside `runAutoCompact`.
184
+ - **Cost model and deferred scheduling.** Pressure is the only trigger, and it is `compaction.threshold`, not `target`. The replay tables price every applied event by the cold prefix it re-prefills (about 29k tokens per event at a 64k budget), and batching from the threshold down to the target is what keeps one event per cycle; a trigger at the target would make every turn above 60% with one newly redundant read an event of its own, and no row in the sweep shows fewer summaries in return. There is no break-even horizon, no deferred eviction plan, and no piggybacking beyond the fact that the working-set stage already runs first inside `runAutoCompact`.
185
185
  - **Intra-turn eviction.** Eviction runs before a request is sent. A single turn whose tool results overflow the window is handled by the observation envelope's caps and by summary compaction, not by this layer.
186
186
  - **Worker runtimes.** Dispatched workers replay their own ledgers without the working-set stage.
187
187
  - **Digests.** A marker carries tool, size, and a first-line preview. The generated summaries from #165 are not embedded in it.
188
188
 
189
189
  ## See also
190
190
 
191
- - `clio-coder context replay --sessions <path>...` replays Clio ledgers, and `--synthetic <ids>` replays the seeded procedural corpora, through the same fold, projection, and policy code with `none`, `random`, and `oracle` controls; `clio-coder context working-set --session <id|path>` prints one session's fold and path index. Both are described under [Working-set replay](commands-and-modes.md#working-set-replay), and the committed tables with the default-policy rule are under `benchmarks/results/context-replay/`.
191
+ - `clio-coder context replay --sessions <path>...` replays Clio ledgers, and `--synthetic <ids>` replays the seeded procedural corpora, through the same fold, projection, and policy code with `none`, `random`, and `oracle` controls; `clio-coder context working-set --session <id|path>` prints one session's fold and path index. Both are described under [Working-set replay](commands-and-modes.md#working-set-replay). Generated replay tables are local artifacts rather than versioned benchmark results.
192
192
  - [context-engine.md](context-engine.md) for context window resolution, token accounting, and how this stage sits ahead of summary compaction.
193
193
  - [session-lifecycle.md](session-lifecycle.md) for the ledger format, active-path lineage, and branching.
194
194
  - [glossary.md](glossary.md) for the one-line definitions of these terms.
@@ -80,7 +80,7 @@ New `area:*` labels are proposed in an issue, not created ad hoc.
80
80
 
81
81
  ## Milestones are releases
82
82
 
83
- Each open milestone is the next version (`v0.3.4`, `v0.4.0`). Triage means
83
+ Each open milestone is the next version (`v0.3.7`, `v0.4.0`). Triage means
84
84
  assigning an issue to a milestone or explicitly leaving it in the backlog.
85
85
  A release cut requires every issue in its milestone to be closed
86
86
  or bumped; the milestone closes when the tag is published.
@@ -28,7 +28,7 @@ split would use. They cross them.
28
28
  | Write-boundary attribution is per scheduling *window*, so the compiler refuses a wave with two writers | scheduling, write boundaries, plan compilation | `execution-plan.ts`, `write-boundary.ts` |
29
29
  | A loop's later nodes are `unneeded`, decided by the scheduler, not the plan | plan compilation, scheduling, receipts | `fleet-plan.ts`, `execution-scheduler.ts` |
30
30
  | Staleness revalidation re-runs a verification a later workspace step invalidated | scheduling, plan compilation, code steps | `execution-scheduler.ts` |
31
- | Receipt integrity v15 seals normalized routing intent | routing, receipts | `receipt-integrity.ts`, `routing-intent.ts` |
31
+ | Receipt integrity v16 seals normalized routing intent | routing, receipts | `receipt-integrity.ts`, `routing-intent.ts` |
32
32
 
33
33
  The write-boundary and loop rows are the sharpest. Both are properties of a
34
34
  *wave*, which is a scheduling concept computed by the plan compiler and enforced
@@ -1,6 +1,6 @@
1
1
  # Clio Coder Documentation Coverage Matrix
2
2
 
3
- This matrix maps every top-level directory in `src/` and every domain directory under `src/domains/` to its authoritative documentation page. It records coverage status (`documented`, `partial`, `undocumented`), missing concepts, and key source contracts for `v0.3.4`.
3
+ This matrix maps every top-level directory in `src/` and every domain directory under `src/domains/` to its authoritative documentation page. It records coverage status (`documented`, `partial`, `undocumented`), missing concepts, and key source contracts for `v0.3.7`.
4
4
 
5
5
  ## Coverage Matrix
6
6
 
@@ -19,8 +19,8 @@ This matrix maps every top-level directory in `src/` and every domain directory
19
19
  | `src/domains/components/` | Component scanning, snapshots, hashing, diffing | [middleware-and-components.md](middleware-and-components.md) | `documented` | Documented in active component snapshot and middleware guide. |
20
20
  | `src/domains/config/` | Configuration contracts, file watcher, keybinding definitions, setting classifiers | [configuration-and-targets.md](configuration-and-targets.md), [commands-and-modes.md](commands-and-modes.md) | `documented` | Documented in configuration targets and command/keybinding reference. |
21
21
  | `src/domains/context/` | `CLIO-CODER.md` bootstrap, codewiki generation, prompt context assembly, project rules, non-destructive working-set eviction (`age-horizon` and `structural-v1` policies, protection predicates, path index, byte-stable markers, recall by ref) | [context-engine.md](context-engine.md), [context-working-set.md](context-working-set.md) | `documented` | Context window, token accounting, and the three compaction mechanisms in the engine reference; the working-set layer has its own guide covering the vocabulary, both ledger record kinds and format v4, the marker contract, both policies with their rule order, recall semantics, and the operator surfaces. |
22
- | `src/domains/dispatch/` | Fleet orchestration, assignment store, batch tracker, admission, route planner, receipt integrity v15 | [fleet-dispatch.md](fleet-dispatch.md), [dispatch-architecture-rationale.md](dispatch-architecture-rationale.md), [worker-dispatch-mechanics.md](worker-dispatch-mechanics.md) | `documented` | Multi-node fleet dispatch, admission invariants, and receipt verification fully documented. |
23
- | `src/domains/eval/` | Suite v2 YAML schema, eval runner, metrics, reporters, workspace sandboxing | [eval-runner.md](eval-runner.md), [evals-internal.md](evals-internal.md) | `documented` | Documented in eval runner and soak benchmark guides. |
22
+ | `src/domains/dispatch/` | Fleet orchestration, assignment store, batch tracker, admission, route planner, receipt integrity v16 | [fleet-dispatch.md](fleet-dispatch.md), [dispatch-architecture-rationale.md](dispatch-architecture-rationale.md), [worker-dispatch-mechanics.md](worker-dispatch-mechanics.md) | `documented` | Multi-node fleet dispatch, admission invariants, and receipt verification fully documented. |
23
+ | `src/domains/eval/` | Suite v2 YAML schema, eval runner, metrics, reporters, workspace sandboxing | [eval-runner.md](eval-runner.md), [evals-internal.md](evals-internal.md) | `documented` | Product evals are documented independently from external benchmarks. |
24
24
  | `src/domains/evidence/` | Evidence bundles, findings taxonomy, provenance store, failure attribution | [evidence-and-memory.md](evidence-and-memory.md) | `documented` | Documented in evidence directory structures and memory retrieval guide. |
25
25
  | `src/domains/evolution/` | Falsifiable Change Manifest JSON templates and `clio-coder evolve` self-edit gates | [evolution.md](evolution.md) | `documented` | Documented in evolution manifest reference and mutation validation rules. |
26
26
  | `src/domains/extensions/` | Extension manifest schemas, resource roots, portable share archives | [extensions-and-sharing.md](extensions-and-sharing.md) | `documented` | Documented in extensions and sharing guide. |
@@ -35,7 +35,7 @@ This matrix maps every top-level directory in `src/` and every domain directory
35
35
  | `src/domains/scheduling/` | Capacity lease acquisition, heartbeats, expiry, cross-process locks, cluster scheduling | [capacity-and-scheduling.md](capacity-and-scheduling.md), [fleet-dispatch.md](fleet-dispatch.md) | `documented` | Dedicated capacity leasing, heartbeat TTL, and cross-process lock reference. |
36
36
  | `src/domains/session/` | Session ledger format v4, tree branching (`/tree`), `/fork`, `/resume`, checkpoints, protected-artifact journal | [session-lifecycle.md](session-lifecycle.md), [context-working-set.md](context-working-set.md) | `documented` | Dedicated session lifecycle guide covering branching, journal, and recovery; the `contextEviction` and `contextRecall` records added at format v4 are specified in the working-set guide. |
37
37
  | `src/domains/share/` | Portable share archive bundles, manifest verification, import/export flows | [extensions-and-sharing.md](extensions-and-sharing.md) | `documented` | Share archives and portable bundle formats documented in extensions guide. |
38
- | `src/domains/webhook/` | Empty directory | None (Inert) | `inert` | Directory contains no active modules or exports in v0.3.4. |
38
+ | `src/domains/webhook/` | Empty directory | None (Inert) | `inert` | Directory contains no active modules or exports in v0.3.7. |
39
39
 
40
40
  ## Cross-Cutting Reference Guides
41
41
 
@@ -1,7 +1,7 @@
1
1
  # Documentation Standards and Codebase Alignment
2
2
 
3
3
  > [!TIP]
4
- > **Interactive Spec Available:** An interactive documentation link linter, phrasing/claim evaluator, and alignment portal is located at [docs/html/documentation_blueprint.html](html/documentation_blueprint.html) (Version: 0.3.4).
4
+ > **Interactive Spec Available:** An interactive documentation link linter, phrasing/claim evaluator, and alignment portal is located at [docs/html/documentation_blueprint.html](html/documentation_blueprint.html) (Version: 0.3.7).
5
5
 
6
6
  Clio Coder is an experimental community alpha. Documentation should help contributors and early users work from the source of truth without overstating maturity. When docs drift, prefer the current source and tests over older prose or aspirational roadmap notes.
7
7
 
@@ -46,10 +46,10 @@ Classify claims clearly:
46
46
  | [alcf-provider.md](alcf-provider.md) | `src/domains/providers/runtimes/cloud/alcf.ts`, `src/engine/alcf-oauth.ts` | Globus PKCE OAuth, openAuthStorage(), Sophia vLLM, Metis API, chatTemplateKwargsUnsupported. |
47
47
  | [environment-variables.md](environment-variables.md) | `src/core/guardrails.ts`, `src/core/xdg.ts`, `src/domains/providers/knowledge-base-path.ts` | Comprehensive env var matrix: guardrail overrides, directory layout (CLIO_CODER_HOME), debug toggles, and internal plumbing. |
48
48
  | [built-in-agents.md](built-in-agents.md) | `src/domains/agents/**`, `src/domains/agents/builtins/*.md`, `src/domains/dispatch/**` | Builtin agent recipes, discovery roots, frontmatter schema, fleet contract shadowing (`.clio-coder/fleets/<name>.md`), active route automation. |
49
- | [fleet-dispatch.md](fleet-dispatch.md) | `src/domains/dispatch/**` | Multi-node SSH dispatch: process-safe admission, capacity leases, Contract v4 write boundaries (detect-and-rollback), bounded check/repair loops (`loop_bound_exhausted`), deterministic code steps, attestation, receipts v15. |
49
+ | [fleet-dispatch.md](fleet-dispatch.md) | `src/domains/dispatch/**` | Multi-node SSH dispatch: process-safe admission, capacity leases, Contract v4 write boundaries (detect-and-rollback), bounded check/repair loops (`loop_bound_exhausted`), deterministic code steps, attestation, receipts v16. |
50
50
  | [capacity-and-scheduling.md](capacity-and-scheduling.md) | `src/domains/scheduling/**`, `src/domains/dispatch/capacity-lease.ts`, `src/domains/dispatch/reservation-store.ts` | Multi-process capacity leases (`dispatch-admission.json`), heartbeat TTLs, cross-process transaction locks (`dispatch-admission.json.lock`), and cluster drain controls. |
51
51
  | [worker-dispatch-mechanics.md](worker-dispatch-mechanics.md) | `src/worker/**` | NDJSON parent-child socket protocols, control/bulk lane demuxing, watchdog timers, worker attestation (13 protocol fields), permission parking, exit codes. |
52
- | [fleet-demo-runbook.md](fleet-demo-runbook.md) | `src/domains/dispatch/**` | Multi-node fleet demo: SSH setup, C++ build/repair workflow, reviewer gates, receipt verification v15. |
52
+ | [fleet-demo-runbook.md](fleet-demo-runbook.md) | `src/domains/dispatch/**` | Multi-node fleet demo: SSH setup, C++ build/repair workflow, reviewer gates, receipt verification v16. |
53
53
  | [session-lifecycle.md](session-lifecycle.md) | `src/engine/session.ts`, `src/domains/session/**` | Session lifecycle, on-disk ledger format v4 (`current.jsonl`), tree branching (`tree.json`), active-path lineage selection, `/fork`, `/resume`, checkpoints, and write-ahead protected-artifact journal. |
54
54
  | [acp.md](acp.md) | `src/engine/acp/**`, `src/cli/acp.ts` | Agent Client Protocol (ACP) server over stdio, tool mediation, non-stall permission handling, timeout bounds, and error taxonomy. |
55
55
  | [artifact-versions.md](artifact-versions.md) | `src/domains/dispatch/receipt-integrity.ts`, `src/engine/session.ts`, `src/worker/spec-contract.ts`, `src/domains/agents/fleet-contract.ts`, `src/domains/eval/schema/`, `src/domains/observability/trace-store.ts` | Version registry and migration policies for all 9 serialized artifact schemas across Clio Coder. |
@@ -65,7 +65,7 @@ Classify claims clearly:
65
65
  | [proactive-memory.md](proactive-memory.md) | `src/domains/memory/**` | Proactive task memory architecture, session task bank, intervention rules, and handoff carrying. |
66
66
  | [trace-store.md](trace-store.md) | `src/cli/trace.ts`, `src/domains/observability/trace-store.ts` | WAL SQLite trace mirror database schema, rowid cursor queries, rebuildability, 6 `clio-coder trace` subcommands (`runs`, `phases`, `tail`, `procs`, read-only `sql` SELECT, `ui`). |
67
67
  | [eval-runner.md](eval-runner.md) | `src/domains/eval/**`, `src/cli/eval.ts` | Local YAML eval tasks, dual token accountings (`tokens.*` wire vs `receiptUsage.*` journal), fail-closed null totals, EvalArtifactV4 format, `verify.measure` task outcome recording. |
68
- | [evals-internal.md](evals-internal.md) | `src/domains/eval/**`, `benchmarks/soak/**` | Private context index determinism, target smoke matrices, soak machinery benchmark suite (4 suites: `clio-soak`, `clio-soak-boundary`, `clio-soak-chaos`, `clio-soak-loop`). |
68
+ | [evals-internal.md](evals-internal.md) | `src/domains/eval/**` | Private context index determinism and target smoke matrices. External model benchmarks are documented under `benchmarks/`. |
69
69
  | [extensions-and-sharing.md](extensions-and-sharing.md) | `src/domains/extensions/**`, `src/domains/resources/**`, `src/domains/share/**`, `src/cli/extensions.ts`, `src/cli/share.ts` | Prompt and skill resources, extension manifests, portable share archives. |
70
70
  | [skills-marketplace.md](skills-marketplace.md) | `src/interactive/overlays/skills-hub.ts`, `src/domains/resources/skills/marketplace.ts` | Skills Hub marketplace discovery through the install resolver, empty state, install actions, publishing flow. |
71
71
  | [model-catalog.md](model-catalog.md) | `src/domains/providers/catalog.ts`, `src/domains/providers/models/**`, `src/domains/providers/probe/**`, `src/domains/providers/model-capabilities.ts` | Model catalog, live probes (`--offline` toggle), exact-id selector `probeCapabilitiesForModel`, field-note promotion. |
@@ -1,7 +1,7 @@
1
1
  # Clio Coder Local Evaluation Runner
2
2
 
3
3
  > [!TIP]
4
- > **Interactive Spec Available:** An interactive task suite validator, subprocess execution simulator, and compare calculator is located at [docs/html/eval_blueprint.html](html/eval_blueprint.html) (Version: 0.3.4).
4
+ > **Interactive Spec Available:** An interactive task suite validator, subprocess execution simulator, and compare calculator is located at [docs/html/eval_blueprint.html](html/eval_blueprint.html) (Version: 0.3.7).
5
5
 
6
6
  The local evaluation runner executes repository-local YAML task suites as deterministic subprocess checks. It is useful for comparing harness changes, prompts, tools, or local workflows.
7
7
 
@@ -1,7 +1,7 @@
1
1
  # Internal Eval Suites
2
2
 
3
3
  > [!TIP]
4
- > **Interactive Spec Available:** Interactive blueprints are available for internal evaluation suites at [docs/html/evals_internal_blueprint.html](html/evals_internal_blueprint.html) and soak benchmark suites at [docs/html/soak_blueprint.html](html/soak_blueprint.html) (Version: 0.3.4).
4
+ > **Interactive Spec Available:** An interactive blueprint is available at [docs/html/evals_internal_blueprint.html](html/evals_internal_blueprint.html) (Version: 0.3.7).
5
5
 
6
6
  Private suites should live outside this repository. Keep datasets, prompts,
7
7
  live fleet coordinates, calibration outputs, and raw run artifacts in a private
@@ -15,9 +15,9 @@ clio-coder eval run --suite <external-path> --clio-coder-entry dist/cli/index.js
15
15
  ```
16
16
 
17
17
  Use `--out <dir>` when the artifact should be written outside the default Clio
18
- data directory. Public summaries can be copied into
19
- `benchmarks/results/<suite>/<run-id>/` only after they have been sanitized down
20
- to `manifest.json` and `summary.json`.
18
+ data directory. Product eval artifacts and external benchmark campaigns are
19
+ separate: public benchmark adapters live under `benchmarks/community/` and do
20
+ not use the eval runner.
21
21
 
22
22
  ## Context Regression Seed
23
23
 
@@ -268,44 +268,3 @@ thresholds:
268
268
  value: 0
269
269
  ```
270
270
 
271
- ---
272
-
273
- ## Soak Benchmark Suite
274
-
275
- The soak benchmark suite located under [`benchmarks/soak/`](../benchmarks/soak/) measures Clio's own machinery performance, integrity, and structural invariant promises under load. Unlike standard evaluation suites, the soak suite evaluates the reliability of Clio rather than model capability. A weak model that fails to solve the workload still passes the suite if Clio's machinery behaves correctly; a strong model fails the suite if Clio fails to seal a receipt, cannot authenticate a receipt, or violates a system invariant.
276
-
277
- Every suite runs through the product's own eval runner against a configured
278
- target; there is no separate soak runner:
279
-
280
- ```bash
281
- npm run build
282
- clio-coder eval run --suite benchmarks/soak/clio-soak.yaml \
283
- --target <id> --model <wireId> --clio-coder-entry dist/cli/index.js
284
- ```
285
-
286
- `tests/contracts/eval-soak-suite.test.ts` loads all four files in CI and drives
287
- `clio-soak.yaml` against a stub that seals receipts on purpose, so the gate is
288
- known to fail when sealing fails; the model runs themselves are operator-run.
289
-
290
- The soak suite comprises four specialized suite files:
291
-
292
- ### 1. Machinery Under Load (`clio-soak.yaml`)
293
- Evaluates the same task workload across two execution surfaces: the headless main-agent surface (`clio-run`) and a dispatched worker surface (`agent: coder`). It tests single-file bugs, multi-file bugs, and compaction continuity across restarts.
294
- - **Surface Differences**: Main-agent tasks verify session ledger continuity (`ledger.formatVersion`, `ledger.toolPairsUnmatched`, `ledger.assistantBetweenCallAndResult`), while dispatch worker tasks verify process group cleanup (`process.orphanedChildren == 0`).
295
- - **Compaction Continuity**: Verifies that compaction summaries are present (`continuity.compactionSummaryPresent`) and that pre-compaction facts are preserved (`continuity.answeredFromPreCompaction`).
296
- - **Suite-Wide Gates**: Gates on `receipt.sealed`, `receipt.integrityValid`, `receipt.outcomeMatchesExit`, `tokens.measured`, `stream.cumulativeSnapshots == 0`, `stream.usageDoubleCounted == false`, and `stream.segmentUsageMatchesMessages == true`.
297
-
298
- ### 2. Per-Step Write Boundaries (`clio-soak-boundary.yaml`)
299
- Validates write boundary enforcement across steps without model participation. Enforcement is strictly detect-and-rollback and is never sandboxing.
300
- - `write-boundary.rolled-back`: Verifies clean detection of allowlist violations (`writes_boundary_violation`), git-level file restoration, and sealed verdict generation (`boundary.violationsRolledBack == 1`, `boundary.rollbackIncomplete == 0`).
301
- - `write-boundary.rollback-incomplete`: Tests honest failure reporting when a path was dirty prior to snapshot taking. prior bytes exist only in the overwritten tree, so rollback leaves the tree unchanged and records incomplete rollback (`boundary.rollbackIncomplete == 1`, `boundary.violationsRolledBack == 0`).
302
-
303
- ### 3. Fault Injection Chaos (`clio-soak-chaos.yaml`)
304
- Evaluates system resilience against process signals.
305
- - `chaos.sigint-mid-tool`: Prompts Clio for a long-running bash tool call and injects `SIGINT` once the subprocess initializes. Asserts exit code `130`, confirms no orphaned children remain (`process.orphanedChildren == 0`), and verifies receipt sealing, receipt integrity, and provider token reporting.
306
-
307
- ### 4. Bounded Loops (`clio-soak-loop.yaml`)
308
- Validates iteration bounds and receipt accounting for fleet loops (`bounded-loop.fleet`).
309
- - **Loop Bounds**: Asserts that verification attempts do not exceed declared limits (`loop.attemptsSpent <= 3`), recovery attempts seal individual receipts (`loop.receiptsMatchRepairs == true`), and unneeded nodes report as `unneeded` rather than skipped or failed (`loop.skippedNodes == 0`).
310
- - **Two Token Accountings**: Distinguishes `tokens.*` (folded live off wire stdout by `createStreamInvariantFold`) from `receiptUsage.*` (journal receipts sealed and authenticated against ledger envelopes). On fleet runs, wire streaming is absent (`tokens.measured == false`), while journal receipts provide authenticated usage (`receiptUsage.measured == true`).
311
-