@iowarp/clio-coder 0.4.2 → 0.4.4

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (523) hide show
  1. package/CHANGELOG.md +98 -0
  2. package/CONTRIBUTING.md +86 -19
  3. package/README.md +35 -6
  4. package/dist/{acp-TMDQZDIG.js → acp-WNAYYF4F.js} +12 -13
  5. package/dist/{agents-5N5NG3XG.js → agents-3OKXHLOI.js} +60 -57
  6. package/dist/assets/codewiki.json +1 -1
  7. package/dist/{auth-Z5CCBXKQ.js → auth-VKNNMGPU.js} +21 -19
  8. package/dist/{builtins-K6TNDT24.js → builtins-WGALA46I.js} +9 -4
  9. package/dist/{chunk-XE3PCIXH.js → chunk-23L32XTI.js} +12 -9
  10. package/dist/{chunk-I64IFBLB.js → chunk-25QBEXRS.js} +18 -11
  11. package/dist/{chunk-CDNVLKUX.js → chunk-26QSH3EJ.js} +13 -7
  12. package/dist/{chunk-QQLGQY2A.js → chunk-2ASED4PZ.js} +22 -22
  13. package/dist/{chunk-MCEPRMZW.js → chunk-2CU2H6KE.js} +2 -2
  14. package/dist/chunk-2DSOYNFC.js +108 -0
  15. package/dist/{chunk-72GZI5EV.js → chunk-2JH2WHGE.js} +2 -2
  16. package/dist/{chunk-I66EAJFY.js → chunk-2WZ546HR.js} +267 -232
  17. package/dist/{chunk-O3YUNJZ2.js → chunk-2ZSONWVL.js} +82 -25
  18. package/dist/{chunk-2NHR3NAY.js → chunk-36CT5VVL.js} +331 -42
  19. package/dist/{chunk-2X4RYJTJ.js → chunk-3GY4F45V.js} +3 -3
  20. package/dist/{chunk-ZW55JB7N.js → chunk-3ODX73FK.js} +4 -6
  21. package/dist/{chunk-PBP4B7XR.js → chunk-3UNOLWNZ.js} +3 -3
  22. package/dist/{chunk-4JDLP6ZS.js → chunk-3UUXNFEX.js} +14 -10
  23. package/dist/{chunk-RKSR6VSF.js → chunk-4IUZQIJ3.js} +29 -1
  24. package/dist/{chunk-ZW4HH5JJ.js → chunk-4M6Z5QVF.js} +6 -6
  25. package/dist/{chunk-K6BSR66V.js → chunk-4NSRCOYP.js} +4 -1
  26. package/dist/{chunk-M2DAX4F6.js → chunk-4WR7VSYB.js} +2 -2
  27. package/dist/{chunk-FSP7CMNU.js → chunk-54X7T7DK.js} +61 -6
  28. package/dist/{chunk-54ODD65L.js → chunk-5636DCO5.js} +4 -4
  29. package/dist/chunk-57XXR6DR.js +3763 -0
  30. package/dist/{chunk-3KIPBMUA.js → chunk-5ICU3EUH.js} +2 -2
  31. package/dist/{chunk-77QIVUZB.js → chunk-5MEZN6CB.js} +4 -4
  32. package/dist/{chunk-O42A54GG.js → chunk-5OIVVPHF.js} +2 -2
  33. package/dist/{chunk-YJISEZKC.js → chunk-5TUB6SLS.js} +6 -6
  34. package/dist/{chunk-IMXMHHMQ.js → chunk-6OSVSQL5.js} +341 -57
  35. package/dist/{chunk-Q4XWMHX6.js → chunk-6PAZTBPA.js} +14 -2
  36. package/dist/{chunk-FVDGR2ZL.js → chunk-6Q3CYFD3.js} +112 -39
  37. package/dist/{chunk-IDNA72AH.js → chunk-6QOTUPRG.js} +155 -36
  38. package/dist/{chunk-X7IARSHT.js → chunk-6UINWWS6.js} +16 -10
  39. package/dist/{chunk-CYZW7JHJ.js → chunk-72YIHOZQ.js} +9 -9
  40. package/dist/{chunk-IKSLQ4XV.js → chunk-75W7L2E2.js} +752 -861
  41. package/dist/{chunk-CRFOIAX3.js → chunk-7UGL4MB5.js} +6 -6
  42. package/dist/{chunk-HIICAHCJ.js → chunk-AUPNRN7C.js} +2 -2
  43. package/dist/{chunk-7BHIY2MW.js → chunk-BJVFZO5U.js} +8 -50
  44. package/dist/{chunk-B74PXLU7.js → chunk-CUSRQKPU.js} +65 -3
  45. package/dist/chunk-DQOVN6KV.js +386 -0
  46. package/dist/{chunk-E7GT7O5N.js → chunk-DT3LWJOB.js} +7 -4
  47. package/dist/chunk-DXKJURES.js +671 -0
  48. package/dist/{chunk-JBCS7CRR.js → chunk-EL24TAU4.js} +10 -10
  49. package/dist/{chunk-TPEQIQIE.js → chunk-ELWDPP3Y.js} +8 -8
  50. package/dist/{chunk-NDINPTJ4.js → chunk-ELZVTCGV.js} +5 -4
  51. package/dist/chunk-EXLD33WO.js +381 -0
  52. package/dist/chunk-FEAXX7B6.js +101 -0
  53. package/dist/{chunk-RLYRBIYQ.js → chunk-FFUPXJC4.js} +90 -331
  54. package/dist/{chunk-5PFYMY2V.js → chunk-FTMGRKEF.js} +2 -2
  55. package/dist/{chunk-34BHNEE3.js → chunk-GHS5EBTQ.js} +58 -7
  56. package/dist/{chunk-DYHAXKHD.js → chunk-GWZNEVM2.js} +12 -8
  57. package/dist/chunk-GX5WYQO4.js +59 -0
  58. package/dist/chunk-GYV6VZOC.js +26 -0
  59. package/dist/{chunk-MQXIVJ35.js → chunk-HAXOFFRH.js} +5 -5
  60. package/dist/chunk-I2DWJ4GM.js +390 -0
  61. package/dist/{chunk-TXOTCRLG.js → chunk-I5FWO7L5.js} +5 -5
  62. package/dist/{chunk-WNP7O5WZ.js → chunk-ID64D7PE.js} +4 -4
  63. package/dist/{chunk-XQRY4DTA.js → chunk-IGLP3ODT.js} +10 -10
  64. package/dist/chunk-IRXAATOX.js +539 -0
  65. package/dist/chunk-IXIY2H4R.js +44 -0
  66. package/dist/{chunk-SSEYRH53.js → chunk-IZXGRF7P.js} +92 -147
  67. package/dist/{chunk-5TSRNF4G.js → chunk-JCI2ROMZ.js} +164 -6
  68. package/dist/{chunk-JWJGP5DQ.js → chunk-JEIYHLOR.js} +7 -7
  69. package/dist/{chunk-F2I26BDK.js → chunk-JQLNNIKT.js} +4 -4
  70. package/dist/{chunk-BYMNWQ7O.js → chunk-JSD46VO2.js} +315 -63
  71. package/dist/{chunk-AK5XEFVZ.js → chunk-JT2RFCC5.js} +64 -14
  72. package/dist/{chunk-MCMZMDAC.js → chunk-K6T2ZAMZ.js} +168 -6
  73. package/dist/{chunk-PGF63K6I.js → chunk-KFV5L5SK.js} +73 -4
  74. package/dist/{chunk-IHXBNWMM.js → chunk-KXDSS5WJ.js} +7 -3
  75. package/dist/{chunk-PJX3WQUQ.js → chunk-LLXSDWXS.js} +3 -3
  76. package/dist/{chunk-DZAW46HP.js → chunk-LTIKRKFL.js} +3 -3
  77. package/dist/{chunk-DZEK6CJN.js → chunk-N56KALIC.js} +21 -21
  78. package/dist/{chunk-B7HM5Z7T.js → chunk-NAI6ZFCY.js} +9 -5
  79. package/dist/{chunk-I66ZTYNP.js → chunk-NRO2BJRH.js} +2656 -2213
  80. package/dist/{chunk-ZGNYYXQ6.js → chunk-NXIMQY5W.js} +3 -3
  81. package/dist/{chunk-IKOZFYBN.js → chunk-NXYCB2VD.js} +149 -106
  82. package/dist/{chunk-CWVRRIEI.js → chunk-NZMNUPZZ.js} +2 -2
  83. package/dist/{chunk-462T4EGZ.js → chunk-O5CVSAG5.js} +2 -2
  84. package/dist/chunk-ODGTEFFI.js +50 -0
  85. package/dist/{chunk-3F7VUY77.js → chunk-OEJSLEPW.js} +2 -2
  86. package/dist/{chunk-KKOJXO6R.js → chunk-OMQNJVKW.js} +4 -2
  87. package/dist/{chunk-5KW52TEP.js → chunk-Q4WO54TA.js} +132 -77
  88. package/dist/{chunk-W6NIE6OW.js → chunk-QUFRYSWI.js} +13 -7
  89. package/dist/{chunk-42FMPA75.js → chunk-QZWQA4DE.js} +2 -2
  90. package/dist/chunk-R6Q67RJH.js +134 -0
  91. package/dist/{chunk-W4YEMFBX.js → chunk-RAY4OVGZ.js} +3 -3
  92. package/dist/{chunk-ZNT2M6TG.js → chunk-RQCKCSRL.js} +17 -17
  93. package/dist/{chunk-LJID3DYZ.js → chunk-RXTN6AKH.js} +3 -3
  94. package/dist/{chunk-P75RZCJW.js → chunk-RZDWV63N.js} +3 -3
  95. package/dist/{chunk-UH632ZYL.js → chunk-S6PYF2XF.js} +2 -2
  96. package/dist/{chunk-HJWWJ6IL.js → chunk-TOIVGRUX.js} +17 -5
  97. package/dist/{chunk-HLAFFSEK.js → chunk-TQAHXW6Y.js} +2 -2
  98. package/dist/{chunk-JIEGK6UF.js → chunk-U6TMQNSI.js} +48 -4
  99. package/dist/{chunk-2HFQNRV3.js → chunk-UEPWCCTY.js} +12 -12
  100. package/dist/chunk-UOIZ7DA4.js +41 -0
  101. package/dist/{chunk-UH347SHR.js → chunk-USR47QNF.js} +11 -11
  102. package/dist/{chunk-AZ4WMN4W.js → chunk-V6HJFQZE.js} +2 -2
  103. package/dist/chunk-V76WTFTW.js +318 -0
  104. package/dist/{chunk-NMJXSHBJ.js → chunk-W54I7H25.js} +2 -2
  105. package/dist/{chunk-KPXDY6QF.js → chunk-XRZT5WY5.js} +2 -2
  106. package/dist/{chunk-UBRFI4HS.js → chunk-XULDXHTN.js} +142 -50
  107. package/dist/chunk-XXYSBZIQ.js +283 -0
  108. package/dist/{chunk-HKMD33FO.js → chunk-Y55JBDO5.js} +405 -122
  109. package/dist/{chunk-XOXV5GKE.js → chunk-YD5GIKET.js} +17 -8
  110. package/dist/{chunk-XGDPUNND.js → chunk-YECAMM3D.js} +2 -2
  111. package/dist/{chunk-BO7Y52RY.js → chunk-YNFKXPEC.js} +7 -7
  112. package/dist/{chunk-VXMFAE2W.js → chunk-YPC6ZR5L.js} +19 -6
  113. package/dist/{chunk-M2WXEHER.js → chunk-ZA4VCIGV.js} +2 -2
  114. package/dist/cli/index.js +42 -40
  115. package/dist/{clio-7VB377CC.js → clio-QLICPCF5.js} +7 -7
  116. package/dist/{code-nav-YVLCYA7V.js → code-nav-IJR2DBPR.js} +9 -9
  117. package/dist/{components-UBWCQSRW.js → components-2TGAI2RC.js} +5 -6
  118. package/dist/{config-4HVOS65E.js → config-IUA6OYNS.js} +88 -81
  119. package/dist/{configure-PIWO7B24.js → configure-VEPX4NMX.js} +26 -25
  120. package/dist/{context-KQYIWPWT.js → context-2DKHWH2T.js} +60 -45
  121. package/dist/{context-IYEHL3WQ.js → context-4MPR7WKB.js} +78 -69
  122. package/dist/{context-N6ZE3LGJ.js → context-BOYF5EJM.js} +15 -11
  123. package/dist/{context-clear-G4OGZJDS.js → context-clear-S4ZJCQUX.js} +73 -65
  124. package/dist/context-map-COB37XXN.js +505 -0
  125. package/dist/{context-working-set-BWLF6LJP.js → context-working-set-3I3FYX6Y.js} +18 -17
  126. package/dist/detail-A7JAVSIG.js +98 -0
  127. package/dist/{dispatch-runner-2QQAITS3.js → dispatch-runner-RJ5I2F2O.js} +99 -75
  128. package/dist/{docs-PD3EXDKU.js → docs-SPOV3BAN.js} +3 -5
  129. package/dist/{doctor-LHBD36VU.js → doctor-DKICC2SN.js} +71 -48
  130. package/dist/{eval-C45FYRJ6.js → eval-OQOQUDHK.js} +308 -146
  131. package/dist/{eval-inventory-6DEJPLBF.js → eval-inventory-Y6QRFOH5.js} +4 -4
  132. package/dist/{evidence-6SHONYAF.js → evidence-4DQ25GUQ.js} +79 -175
  133. package/dist/evidence-4F5USFKH.js +208 -0
  134. package/dist/{evolve-KRKMV72X.js → evolve-GSS52E5J.js} +71 -67
  135. package/dist/{extensions-KPZ2UHBB.js → extensions-G7MFLYHT.js} +8 -9
  136. package/dist/{fleet-IVTCKDHT.js → fleet-Q37YHHAQ.js} +126 -118
  137. package/dist/{fleet-commands-EDWL3IT7.js → fleet-commands-G7E4N7SM.js} +16 -13
  138. package/dist/{fleet-decisions-YP3YEFGK.js → fleet-decisions-O7M6QBA2.js} +9 -8
  139. package/dist/{fleet-graph-ZFWKHY2M.js → fleet-graph-TOUBW6OW.js} +21 -20
  140. package/dist/{fleet-inspect-FVUNCBML.js → fleet-inspect-SW33JJNI.js} +65 -60
  141. package/dist/{fleet-preflight-UN5XED4R.js → fleet-preflight-CV2655TW.js} +4 -5
  142. package/dist/{fleet-validate-XOWC4HSX.js → fleet-validate-KESZX2YH.js} +25 -24
  143. package/dist/{fleet-verify-UN3SODEL.js → fleet-verify-UQPMTVE3.js} +66 -61
  144. package/dist/{fleet-view-TWHJKCN6.js → fleet-view-MN2VG4MR.js} +65 -60
  145. package/dist/{init-T2QORQ3Y.js → init-PXEXQSBF.js} +90 -82
  146. package/dist/{interop-IN5I2A66.js → interop-ZG5T62U3.js} +12 -13
  147. package/dist/inventory-C26CFDRR.js +101 -0
  148. package/dist/{library-LSCATDLZ.js → library-B2W4N74O.js} +29 -29
  149. package/dist/{memory-HYOKAGGJ.js → memory-YCANYS5A.js} +73 -69
  150. package/dist/{models-2GPMFYCM.js → models-GERTU3YI.js} +51 -47
  151. package/dist/{monitor-E4ASVUJH.js → monitor-CPNIUULB.js} +74 -67
  152. package/dist/{orchestrator-DDMPR3PY.js → orchestrator-J4BSH4WQ.js} +1288 -1606
  153. package/dist/{panes-E3RUXOW5.js → panes-BOHAEGYC.js} +4 -4
  154. package/dist/{panes-IXKLOKA2.js → panes-NXSLDQZ2.js} +10 -11
  155. package/dist/{paths-L7LGY6RN.js → paths-VSUWNC22.js} +6 -7
  156. package/dist/reset-TNWTB5LU.js +343 -0
  157. package/dist/{resources-OTRSN34L.js → resources-4PXNMD5G.js} +29 -22
  158. package/dist/{run-5DEYH5QK.js → run-D6XJ34CN.js} +132 -132
  159. package/dist/{share-IHWTLO3M.js → share-2NWMJJEE.js} +27 -27
  160. package/dist/{skills-IYMXMKW4.js → skills-KR7WON5G.js} +40 -33
  161. package/dist/{skills-eval-DROHSJAR.js → skills-eval-O2ZNOLDS.js} +81 -77
  162. package/dist/{skills-inventory-D7X4L4ZX.js → skills-inventory-ZZOUBK7O.js} +23 -22
  163. package/dist/{slash-commands-QBM7UZ3B.js → slash-commands-ZXPJD64J.js} +47 -37
  164. package/dist/{steer-Z5DO23FJ.js → steer-XA25PSCS.js} +4 -4
  165. package/dist/{support-U7QOWY26.js → support-7EMVWYG2.js} +6 -6
  166. package/dist/{targets-P2FUC4IL.js → targets-OMH2XCSN.js} +50 -49
  167. package/dist/tasks-IPAGMEIX.js +36 -0
  168. package/dist/{terminal-lease-YREJ3JX2.js → terminal-lease-C2J3JYRE.js} +4 -4
  169. package/dist/{tools-5B7RO6MV.js → tools-EFFEAIDP.js} +8 -9
  170. package/dist/{trace-YMGMUM6A.js → trace-FXMXUZUF.js} +7 -7
  171. package/dist/uninstall-HALS6BLF.js +407 -0
  172. package/dist/upgrade-MS72RJEP.js +306 -0
  173. package/dist/{usage-ME5MPXGX.js → usage-NHG6MCJM.js} +162 -108
  174. package/dist/{verifiers-BVZ7IWOO.js → verifiers-7AUNVXDY.js} +155 -22
  175. package/dist/{verify-5K7ZKQFC.js → verify-FWYGPKMR.js} +14 -12
  176. package/dist/{web-fetch-MPARV2K7.js → web-fetch-V4FKSDAV.js} +4 -4
  177. package/dist/{wiki-generate-F5W5QTYY.js → wiki-generate-743CIGJW.js} +99 -90
  178. package/dist/{with-panes-BYOJCLAM.js → with-panes-BDQEWBRT.js} +10 -10
  179. package/dist/worker/entry.js +72 -68
  180. package/docs/README.md +3 -2
  181. package/docs/architecture/acp.md +17 -0
  182. package/docs/architecture/artifact-placement.md +1 -0
  183. package/docs/architecture/artifact-versions.md +2 -2
  184. package/docs/architecture/context-engine.md +4 -0
  185. package/docs/architecture/dispatch-typed-intent.md +1 -1
  186. package/docs/architecture/evidence-and-memory.md +1 -1
  187. package/docs/architecture/middleware-and-components.md +1 -1
  188. package/docs/architecture/model-catalog.md +21 -10
  189. package/docs/architecture/observability.md +19 -2
  190. package/docs/architecture/prompt-envelope-and-tools.md +17 -5
  191. package/docs/architecture/provider-adapter-cookbook.md +63 -0
  192. package/docs/architecture/safety-model.md +25 -22
  193. package/docs/architecture/tui-design.md +1 -1
  194. package/docs/guide/built-in-agents.md +25 -11
  195. package/docs/guide/commands-and-modes.md +18 -3
  196. package/docs/guide/configuration-and-targets.md +100 -10
  197. package/docs/guide/configuration-reference.md +17 -7
  198. package/docs/guide/environment-variables.md +4 -2
  199. package/docs/guide/installation-and-lifecycle.md +37 -4
  200. package/docs/guide/proactive-memory.md +66 -55
  201. package/docs/guide/skills-marketplace.md +18 -0
  202. package/docs/guide/tool-usage.md +78 -3
  203. package/docs/history/config-knobs-audit.md +2 -2
  204. package/docs/process/development-pipeline.md +40 -2
  205. package/docs/process/eval-runner.md +67 -3
  206. package/docs/process/git-commit-provenance.md +15 -0
  207. package/docs/process/release-cut-checklist.md +207 -0
  208. package/docs/process/scientific-validation.md +18 -17
  209. package/evals/behavioral-machinery-support.ts +1 -0
  210. package/evals/behavioral-machinery.yaml +1 -1
  211. package/evals/behavioral-model.yaml +3 -2
  212. package/package.json +2 -2
  213. package/skills/README.md +7 -5
  214. package/skills/coding/ast-grep/SKILL.md +101 -30
  215. package/skills/coding/ast-grep/evals.md +26 -0
  216. package/skills/coding/coding-standards/SKILL.md +47 -2
  217. package/skills/coding/coding-standards/evals.md +23 -0
  218. package/skills/coding/prototype/SKILL.md +87 -28
  219. package/skills/coding/prototype/evals.md +19 -0
  220. package/skills/coding/tdd/SKILL.md +80 -53
  221. package/skills/coding/tdd/evals.md +20 -0
  222. package/skills/context/context-handoff/SKILL.md +43 -2
  223. package/skills/context/context-handoff/evals.md +44 -0
  224. package/skills/context/context-prime/SKILL.md +45 -15
  225. package/skills/context/context-prime/evals.md +45 -0
  226. package/skills/git/branch-closeout/SKILL.md +132 -0
  227. package/skills/git/branch-closeout/evals.md +133 -0
  228. package/skills/git/branch-closeout/references/closeout-checklist.md +81 -0
  229. package/skills/git/file-ticket/SKILL.md +77 -63
  230. package/skills/git/file-ticket/assets/issue-template.md +22 -0
  231. package/skills/git/file-ticket/evals.md +31 -26
  232. package/skills/git/file-ticket/references/issue-discovery.md +49 -0
  233. package/skills/git/fix-issue/SKILL.md +87 -64
  234. package/skills/git/fix-issue/evals.md +35 -31
  235. package/skills/git/fix-issue/references/diagnosis-and-rca.md +46 -0
  236. package/skills/git/resolve-merge-conflicts/SKILL.md +100 -51
  237. package/skills/git/resolve-merge-conflicts/evals.md +52 -25
  238. package/skills/git/resolve-merge-conflicts/references/conflict-matrix.md +126 -0
  239. package/skills/git/ship/SKILL.md +103 -67
  240. package/skills/git/ship/assets/pr-template.md +21 -0
  241. package/skills/git/ship/evals.md +44 -28
  242. package/skills/git/ship/references/remote-and-branch-policy.md +62 -0
  243. package/skills/git/worktree-create/SKILL.md +80 -50
  244. package/skills/git/worktree-create/evals.md +40 -33
  245. package/skills/git/worktree-create/references/worktree-setup.md +62 -66
  246. package/skills/git/worktree-merge/SKILL.md +112 -65
  247. package/skills/git/worktree-merge/evals.md +42 -34
  248. package/skills/git/worktree-merge/references/merge-strategies.md +52 -0
  249. package/skills/planning/archify/SKILL.md +196 -0
  250. package/skills/planning/archify/evals.md +65 -0
  251. package/skills/planning/architecture/SKILL.md +61 -12
  252. package/skills/planning/architecture/evals.md +65 -0
  253. package/skills/planning/backlog/SKILL.md +130 -14
  254. package/skills/planning/backlog/evals.md +142 -0
  255. package/skills/planning/prd/SKILL.md +47 -6
  256. package/skills/planning/prd/evals.md +54 -0
  257. package/skills/planning/product-intent/SKILL.md +57 -2
  258. package/skills/planning/product-intent/evals.md +70 -0
  259. package/skills/planning/tech-spec/SKILL.md +53 -2
  260. package/skills/planning/tech-spec/evals.md +73 -0
  261. package/skills/registry.yaml +58 -50
  262. package/skills/remote.yaml +13 -0
  263. package/skills/research/arxiv-literature/SKILL.md +76 -18
  264. package/skills/research/arxiv-literature/evals.md +50 -0
  265. package/skills/research/experiment-protocol/SKILL.md +20 -1
  266. package/skills/research/experiment-protocol/evals.md +23 -0
  267. package/skills/research/scientific-debugging/SKILL.md +24 -1
  268. package/skills/research/scientific-debugging/evals.md +18 -0
  269. package/skills/research/scientific-modernization/SKILL.md +26 -1
  270. package/skills/research/scientific-modernization/evals.md +27 -0
  271. package/skills/skill-marketplace.json +63 -28
  272. package/skills/workflow/cut-it/SKILL.md +64 -5
  273. package/skills/workflow/cut-it/evals.md +101 -0
  274. package/skills/workflow/design-council/SKILL.md +112 -27
  275. package/skills/workflow/design-council/evals.md +161 -0
  276. package/skills/workflow/grill-me/SKILL.md +85 -10
  277. package/skills/workflow/grill-me/evals.md +153 -0
  278. package/skills/workflow/workflow-distiller/SKILL.md +76 -17
  279. package/skills/workflow/workflow-distiller/evals.md +118 -0
  280. package/src/cli/args.ts +0 -8
  281. package/src/cli/configure-interop.ts +105 -13
  282. package/src/cli/configure-oauth.ts +57 -0
  283. package/src/cli/configure-onboarding.ts +980 -0
  284. package/src/cli/configure-target.ts +594 -0
  285. package/src/cli/configure.ts +1084 -529
  286. package/src/cli/context-map.ts +114 -0
  287. package/src/cli/context.ts +4 -0
  288. package/src/cli/doctor-state-size.ts +1 -12
  289. package/src/cli/doctor-validation-contract.ts +28 -0
  290. package/src/cli/doctor.ts +5 -0
  291. package/src/cli/evidence-detail.ts +1 -75
  292. package/src/cli/evidence-inventory.ts +1 -167
  293. package/src/cli/index.ts +3 -0
  294. package/src/cli/lifecycle-presenter.ts +436 -0
  295. package/src/cli/models.ts +10 -2
  296. package/src/cli/modes/print.ts +5 -1
  297. package/src/cli/reset.ts +228 -106
  298. package/src/cli/run.ts +7 -4
  299. package/src/cli/select.ts +664 -0
  300. package/src/cli/skills.ts +9 -2
  301. package/src/cli/targets.ts +3 -0
  302. package/src/cli/tasks.ts +84 -0
  303. package/src/cli/uninstall.ts +233 -165
  304. package/src/cli/upgrade.ts +210 -150
  305. package/src/cli/usage.ts +92 -27
  306. package/src/cli/validate-model.ts +3 -3
  307. package/src/cli/verifiers.ts +147 -1
  308. package/src/cli/wiki-generate.ts +1 -0
  309. package/src/core/commit-attribution.ts +41 -1
  310. package/src/core/config.ts +56 -0
  311. package/src/core/external-diagnostic.ts +44 -0
  312. package/src/core/gateway-routing.ts +157 -0
  313. package/src/core/git-commit-attribution.ts +46 -3
  314. package/src/core/run-overrides.ts +0 -5
  315. package/src/core/safe-exec.ts +17 -2
  316. package/src/core/skill-activation.ts +92 -2
  317. package/src/core/tool-names.ts +5 -2
  318. package/src/domains/agents/builtins/architect.md +1 -1
  319. package/src/domains/agents/builtins/coder.md +1 -1
  320. package/src/domains/agents/builtins/documenter.md +1 -1
  321. package/src/domains/agents/builtins/git-master.md +1 -1
  322. package/src/domains/agents/builtins/provenance.md +7 -7
  323. package/src/domains/agents/builtins/tester.md +1 -1
  324. package/src/domains/agents/builtins/verifier.md +2 -2
  325. package/src/domains/agents/builtins/wiki-writer.md +4 -3
  326. package/src/domains/agents/builtins/world-knowledge.md +31 -0
  327. package/src/domains/agents/catalog.ts +1 -1
  328. package/src/domains/agents/result-contract.ts +70 -0
  329. package/src/domains/context/extension.ts +31 -7
  330. package/src/domains/context/refresh.ts +3 -0
  331. package/src/domains/context/wiki/frontmatter.ts +5 -2
  332. package/src/domains/context/wiki/generate.ts +6 -0
  333. package/src/domains/context/wiki/map-seed.ts +589 -0
  334. package/src/domains/context/wiki/plan.ts +2 -2
  335. package/src/domains/context/wiki/prompts.ts +43 -0
  336. package/src/domains/dispatch/active-route-planner.ts +4 -0
  337. package/src/domains/dispatch/admission.ts +29 -0
  338. package/src/domains/dispatch/agent-candidates.ts +10 -0
  339. package/src/domains/dispatch/budget-envelope.ts +86 -1
  340. package/src/domains/dispatch/capability-match.ts +1 -0
  341. package/src/domains/dispatch/capacity-lease.ts +17 -0
  342. package/src/domains/dispatch/code-step.ts +11 -4
  343. package/src/domains/dispatch/contract.ts +42 -8
  344. package/src/domains/dispatch/execution-scheduler.ts +2 -0
  345. package/src/domains/dispatch/extension.ts +184 -56
  346. package/src/domains/dispatch/fleet-commit-attribution.ts +5 -0
  347. package/src/domains/dispatch/fleet-run.ts +1 -0
  348. package/src/domains/dispatch/host-verification.ts +114 -13
  349. package/src/domains/dispatch/intent.ts +28 -18
  350. package/src/domains/dispatch/orphan-recovery.ts +2 -0
  351. package/src/domains/dispatch/receipt-integrity.ts +4 -0
  352. package/src/domains/dispatch/reservation-store.ts +5 -3
  353. package/src/domains/dispatch/state.ts +15 -2
  354. package/src/domains/dispatch/types.ts +20 -2
  355. package/src/domains/dispatch/worker-model-metadata.ts +38 -0
  356. package/src/domains/eval/metrics/call-ledger-stream.ts +34 -11
  357. package/src/domains/eval/metrics/token-stream.ts +201 -31
  358. package/src/domains/eval/metrics/tracked.ts +40 -4
  359. package/src/domains/eval/runners/clio-run.ts +17 -11
  360. package/src/domains/eval/runners/context-index.ts +2 -7
  361. package/src/domains/eval/runners/context-init.ts +3 -6
  362. package/src/domains/eval/runners/external-command.ts +29 -11
  363. package/src/domains/eval/schema/suite.ts +28 -0
  364. package/src/domains/eval/schema/verdict.ts +2 -2
  365. package/src/domains/eval/suites/resolve.ts +13 -1
  366. package/src/domains/eval/suites/run.ts +24 -3
  367. package/src/domains/evidence/build.ts +102 -15
  368. package/src/domains/evidence/detail.ts +69 -0
  369. package/src/domains/evidence/eval.ts +13 -1
  370. package/src/domains/evidence/finish-contract-map.ts +5 -1
  371. package/src/domains/evidence/inventory.ts +167 -0
  372. package/src/domains/evidence/store.ts +16 -0
  373. package/src/domains/evidence/types.ts +12 -0
  374. package/src/domains/extensions/resources.ts +7 -0
  375. package/src/domains/interop/registry.ts +6 -2
  376. package/src/domains/interop/types.ts +4 -0
  377. package/src/domains/lifecycle/migrations/index.ts +4 -0
  378. package/src/domains/memory/task-memory-policy.ts +70 -26
  379. package/src/domains/memory/task-memory-telemetry.ts +1 -0
  380. package/src/domains/middleware/index.ts +0 -1
  381. package/src/domains/middleware/marketplace-offer.ts +22 -35
  382. package/src/domains/middleware/memory-intervention.ts +127 -32
  383. package/src/domains/middleware/memory-step-endpoint.ts +3 -2
  384. package/src/domains/middleware/runtime.ts +7 -3
  385. package/src/domains/middleware/skills-reminder.ts +31 -2
  386. package/src/domains/mux/detect.ts +3 -6
  387. package/src/domains/observability/accountability.ts +15 -1
  388. package/src/domains/observability/compaction-usage.ts +118 -0
  389. package/src/domains/observability/contract.ts +52 -7
  390. package/src/domains/observability/cost.ts +1 -1
  391. package/src/domains/observability/evidence-index.ts +10 -0
  392. package/src/domains/observability/extension.ts +15 -6
  393. package/src/domains/observability/out-of-turn-usage.ts +52 -21
  394. package/src/domains/observability/projection.ts +394 -45
  395. package/src/{interactive → domains/observability}/worker-progress.ts +3 -3
  396. package/src/domains/prompts/fragments/operating/contract.md +2 -0
  397. package/src/domains/prompts/fragments/wiki/page.md +8 -0
  398. package/src/domains/providers/contract.ts +4 -1
  399. package/src/domains/providers/extension.ts +40 -9
  400. package/src/domains/providers/model-capabilities.ts +9 -0
  401. package/src/domains/providers/model-discovery.ts +3 -4
  402. package/src/domains/providers/model-runtime-capabilities.ts +15 -5
  403. package/src/domains/providers/models/local-models/clio-coder-local-coding-targets.yaml +48 -26
  404. package/src/domains/providers/runtimes/antigravity/antigravity-code.ts +225 -45
  405. package/src/domains/providers/runtimes/claude/claude-code.ts +9 -0
  406. package/src/domains/providers/runtimes/common/lmstudio-http.ts +6 -2
  407. package/src/domains/providers/runtimes/common/local-synth.ts +2 -0
  408. package/src/domains/providers/runtimes/protocol/litellm.ts +119 -29
  409. package/src/domains/providers/support.ts +11 -5
  410. package/src/domains/providers/target-model-cache.ts +25 -2
  411. package/src/domains/providers/types/capability-flags.ts +2 -0
  412. package/src/domains/providers/types/runtime-descriptor.ts +20 -1
  413. package/src/domains/providers/types/target-descriptor.ts +19 -0
  414. package/src/domains/resources/index.ts +3 -0
  415. package/src/domains/resources/skills/install.ts +72 -7
  416. package/src/domains/resources/skills/loader.ts +7 -0
  417. package/src/domains/resources/skills/marketplace.ts +63 -11
  418. package/src/domains/safety/action-classifier.ts +7 -0
  419. package/src/domains/safety/autonomy.ts +15 -0
  420. package/src/domains/safety/default-path-policy.ts +2 -0
  421. package/src/domains/safety/finish-contract-registration.ts +29 -14
  422. package/src/domains/safety/finish-contract.ts +252 -40
  423. package/src/domains/safety/index.ts +21 -1
  424. package/src/domains/safety/path-policy.ts +1 -1
  425. package/src/domains/safety/policy-engine.ts +60 -17
  426. package/src/domains/safety/protected-artifacts.ts +191 -88
  427. package/src/domains/safety/rigor.ts +53 -39
  428. package/src/domains/safety/run-effects.ts +2 -22
  429. package/src/domains/safety/skill-authority.ts +55 -0
  430. package/src/domains/safety/validation-contract.ts +388 -0
  431. package/src/domains/session/archive-readers.ts +10 -1
  432. package/src/domains/session/compaction/compact.ts +72 -22
  433. package/src/domains/session/decision-board.ts +101 -2
  434. package/src/domains/session/entries.ts +50 -7
  435. package/src/domains/session/extension.ts +4 -4
  436. package/src/domains/session/handoff.ts +2 -1
  437. package/src/domains/session/manager.ts +2 -3
  438. package/src/domains/session/task-board.ts +14 -1
  439. package/src/domains/session/tree/fork.ts +1 -2
  440. package/src/domains/session/tree/navigator.ts +1 -1
  441. package/src/domains/session/usage.ts +3 -3
  442. package/src/domains/user-tasks/acceptance.ts +56 -0
  443. package/src/domains/user-tasks/active-acceptance.ts +40 -0
  444. package/src/domains/user-tasks/store.ts +34 -3
  445. package/src/engine/acp/adapter.ts +24 -6
  446. package/src/engine/acp/server.ts +21 -4
  447. package/src/engine/acp/transport.ts +53 -8
  448. package/src/engine/acp/types.ts +4 -0
  449. package/src/engine/agent.ts +13 -3
  450. package/src/engine/ai.ts +26 -8
  451. package/src/engine/antigravity/subprocess-runtime.ts +386 -120
  452. package/src/engine/api-registry.ts +3 -0
  453. package/src/engine/apis/ollama-native.ts +15 -0
  454. package/src/engine/apis/openai-completions.ts +117 -14
  455. package/src/engine/claude/subprocess-runtime.ts +107 -60
  456. package/src/engine/external-subprocess.ts +122 -6
  457. package/src/entry/background-model-metadata.ts +18 -0
  458. package/src/entry/compaction-prompt.ts +57 -0
  459. package/src/entry/orchestrator.ts +416 -218
  460. package/src/entry/task-memory-lifecycle.ts +35 -0
  461. package/src/interactive/chat-loop-messages.ts +13 -4
  462. package/src/interactive/chat-loop.ts +65 -2
  463. package/src/interactive/chat-renderer.ts +1 -0
  464. package/src/interactive/cost-overlay.ts +26 -2
  465. package/src/interactive/dispatch-board.ts +46 -717
  466. package/src/interactive/fleet-run-preview.ts +2 -1
  467. package/src/interactive/interactive-application.ts +3 -2
  468. package/src/interactive/interactive-presentation.ts +55 -12
  469. package/src/interactive/interactive-slash-runtime.ts +4 -2
  470. package/src/interactive/oracle.ts +5 -2
  471. package/src/interactive/overlays/fleet-run-approval.ts +3 -2
  472. package/src/interactive/overlays/message-picker.ts +2 -2
  473. package/src/interactive/overlays/settings.ts +2 -2
  474. package/src/interactive/overlays/tree-selector.ts +2 -2
  475. package/src/interactive/renderers/branch-summary.ts +1 -1
  476. package/src/interactive/renderers/worker-entry.ts +32 -0
  477. package/src/interactive/slash-autocomplete.ts +4 -6
  478. package/src/interactive/slash-commands.ts +49 -45
  479. package/src/interactive/slash-spec.ts +28 -0
  480. package/src/interactive/theme/labels.ts +19 -13
  481. package/src/interactive/turn-context.ts +9 -5
  482. package/src/interactive/turn-recovery.ts +8 -0
  483. package/src/interactive/turn-runtime.ts +27 -11
  484. package/src/interactive/turn-state.ts +7 -0
  485. package/src/interactive/view/artifacts.ts +2 -0
  486. package/src/interactive/worker-receipts.ts +1 -0
  487. package/src/interactive/worker-stream.ts +13 -4
  488. package/src/tools/bootstrap.ts +4 -0
  489. package/src/tools/builtin-tool-catalog.ts +31 -0
  490. package/src/tools/compete-worktrees.ts +7 -1
  491. package/src/tools/context/index.ts +30 -9
  492. package/src/tools/core-bootstrap.ts +16 -0
  493. package/src/tools/decide.ts +136 -0
  494. package/src/tools/dispatch-admission.ts +21 -0
  495. package/src/tools/dispatch-arguments.ts +1 -0
  496. package/src/tools/dispatch-event-text.ts +10 -0
  497. package/src/tools/dispatch-plan.ts +6 -2
  498. package/src/tools/dispatch-runner.ts +27 -1
  499. package/src/tools/dispatch-types.ts +6 -0
  500. package/src/tools/evidence.ts +96 -0
  501. package/src/tools/limitation.ts +76 -0
  502. package/src/tools/policy.ts +9 -0
  503. package/src/tools/presentation.ts +3 -0
  504. package/src/tools/registry.ts +11 -5
  505. package/src/tools/result-shaping.ts +17 -5
  506. package/src/tools/task-worktree.ts +13 -3
  507. package/src/tools/tasks.ts +10 -1
  508. package/src/tools/verify/authoring.ts +170 -83
  509. package/src/tools/verify/catalog.ts +122 -5
  510. package/src/tools/verify/index.ts +2 -1
  511. package/src/tools/verify/numeric.ts +298 -0
  512. package/src/tools/verify/perf.ts +143 -0
  513. package/src/tools/verify/scripts.ts +229 -2
  514. package/src/tools/worker-evidence.ts +3 -1
  515. package/src/worker/spec-contract.ts +4 -0
  516. package/dist/chunk-2Z2IKEXI.js +0 -1554
  517. package/dist/chunk-RVG5JXAL.js +0 -41
  518. package/dist/chunk-T56WDKA5.js +0 -183
  519. package/dist/chunk-VN3SHNBN.js +0 -313
  520. package/dist/chunk-VPKWYKEY.js +0 -169
  521. package/dist/reset-OAQP3W4O.js +0 -230
  522. package/dist/uninstall-N34PCTGJ.js +0 -331
  523. package/dist/upgrade-PXK3S2YM.js +0 -325
@@ -7,7 +7,7 @@ triggers:
7
7
  - get multiple expert perspectives
8
8
  - weigh the architecture options
9
9
  - what would experts say
10
- version: 0.4.0
10
+ version: 0.6.0
11
11
  license: Apache-2.0
12
12
  allowed-tools:
13
13
  - dispatch
@@ -17,12 +17,13 @@ allowed-tools:
17
17
  - ls
18
18
  - context
19
19
  - code_nav
20
+ - ask_user
20
21
  clio-coder:
21
22
  registry-id: iowarp/clio-coder
22
23
  source-url: https://github.com/iowarp/clio-coder/tree/main/skills/workflow/design-council
23
24
  audit: pass
24
25
  provenance: designed
25
- eval-status: scenarios-recorded
26
+ eval-status: smoke-checked
26
27
  model-size: large
27
28
  agents:
28
29
  - scout
@@ -36,22 +37,90 @@ Run a bounded multi-perspective debate on a real design decision. The council
36
37
  surfaces the crux of a disagreement before code commits to one side. It is not
37
38
  a ritual: if experts would agree, do not convene it.
38
39
 
40
+ ## Arguments
41
+
42
+ ```text
43
+ convene a design council on <decision>
44
+ ```
45
+
46
+ There is no flag syntax; the trigger is conversational — "convene a design
47
+ council on X", "debate this design", "get multiple expert perspectives".
48
+ Whatever the user names is the decision. A referenced file, doc, or repo path
49
+ in the same request is Step 0/1's grounding to read first, not a separate
50
+ argument.
51
+
52
+ **Headless is the enforced default, not a suggestion.** A council is several
53
+ worker runs; nobody is waiting between rounds in a headless run, and every
54
+ round that actually runs costs real wall-clock time on top of the
55
+ orchestrator's own turns. `ask_user` is not registered in a headless run,
56
+ so the single question Step 1 asks is refused as an unregistered tool
57
+ rather than answered. Treat that refusal (or
58
+ skip the call and reason from this paragraph directly — one is not more valid
59
+ than the other) as the fixed answer **quick mode: exactly three perspectives,
60
+ exactly one round (Positions) plus synthesis.** Do not compose four or five
61
+ perspectives headlessly and do not run a Responses or Convergence round
62
+ headlessly, no matter how contested the topic looks — four/five perspectives
63
+ and multi-round debate are for a live session with an operator who actually
64
+ asked for the deeper pass. This is the one rule the skill's own prior smoke
65
+ history says a model will not reliably self-infer from prose alone, so treat
66
+ the number 3 and the number 1 as hard, not as defaults to raise if the topic
67
+ seems to deserve more.
68
+
69
+ **Dispatch call shape.** Compose the round's perspectives, then make exactly
70
+ one `dispatch` call with all of them in `tasks` and `mode="parallel"`. If
71
+ that call comes back admission-denied for endpoint or target capacity (a
72
+ single local model instance commonly allows only one concurrent worker, so a
73
+ 3-task parallel wave can be denied outright rather than queued), retry the
74
+ identical `tasks` batch in one dispatch call with `mode="sequential"` instead
75
+ — the tool runs them one after another itself. Never split a round into
76
+ several separate one-task `dispatch` calls made one at a time waiting on each
77
+ result before deciding the next; that is the serial-perspectives pattern that
78
+ produced the round-trip cost the skill's own timeout history is about, and it
79
+ does not fix the capacity problem the parallel call already reported. Do not
80
+ call `dispatch(list:true)` to probe capacity first — it answers nothing about
81
+ concurrency and only spends a call.
82
+
83
+ Declare `intent: {read_roots: [...], relevant_paths: [...]}` with paths
84
+ relative to the repo root on every dispatch call instead of pasting an
85
+ absolute path into `task`/`briefing` prose (a config value, a mount point).
86
+ An absolute path token in briefing/task text with no declared `intent` is
87
+ rejected as `legacy_scope_path_absolute`.
88
+
89
+ A dispatch call's own synchronous result already carries every worker's
90
+ output — do not follow it with a `bash`/`read` pass over the receipt file on
91
+ disk to re-read what you already have. The steps below are the plan; do not
92
+ open a `tasks` list for them. This skill's tool surface is exactly
93
+ `dispatch`, `read`, `grep`, `find`, `ls`, `context`, `code_nav`, and
94
+ `ask_user` (`context` and `ask_user` are always available regardless).
95
+ `tasks` and `bash` both sit outside it and any call to either is refused —
96
+ locate files with `find`/`ls`, not `bash find`/`bash ls`; inspect a receipt
97
+ with `read`, not `bash cat`. This skill never writes: `write` and `artifact`
98
+ are not on its surface, so the synthesis in Step 4 is chat output, never a
99
+ file.
100
+
39
101
  ## Step 0 — Check the question is contested
40
102
 
41
103
  Before composing anyone, ask: would credible experts actually disagree on the
42
104
  answer? If every perspective you can imagine picks the same option and differs
43
- only in caveats, stop here. Say the council is not needed, give the consensus
44
- answer with the caveats attached, and end.
105
+ only in caveats, stop here — do not dispatch anything. Say the council is not
106
+ needed, give the consensus answer with the caveats attached, and end. Do not
107
+ dispatch a round "just to confirm" a consensus call you already reached —
108
+ that is the ritual this step exists to skip, and it still costs the same
109
+ wall-clock time and worker slots as a real debate. Trust this self-check the
110
+ same way you trust the rest of your own reasoning; a council you convened to
111
+ double-check yourself is not more rigorous than the judgment behind it.
45
112
 
46
113
  ## Step 1 — Compose perspectives
47
114
 
48
115
  Derive perspectives from the topic itself, never from a generic role menu.
49
- Three is the default and the right number for almost every decision. Go to
50
- four or five only when the decision genuinely has that many independent
51
- stances, and never headless: each perspective is a worker run, and a model
52
- that dispatches the round serially instead of in parallel turns five
53
- perspectives into five sequential runs. If you are running without a user to
54
- wait on you, use three perspectives and one round.
116
+ On a genuinely contested topic, call `ask_user` once with `mode:
117
+ "single_question"` asking whether this should be a quick pass (three
118
+ perspectives, one round) or the full debate (up to five perspectives, up to
119
+ three rounds). See Arguments for what a headless run does with that call.
120
+ In a live session where the user answers, honor the requested depth. Compose
121
+ three perspectives by default; go to four or five only in that live full-
122
+ debate case, and only when the decision genuinely has that many independent
123
+ stances.
55
124
 
56
125
  Each perspective gets:
57
126
 
@@ -74,11 +143,13 @@ perspective from the read-only recipes in the live catalog:
74
143
  - `researcher`: stance leaning on external docs, standards, or papers.
75
144
  - `provenance`: stance arguing from runtime evidence and receipts.
76
145
 
77
- Run one round's perspectives in parallel: one `dispatch` call with the round's
78
- task prompts in `tasks` and `mode="parallel"`. Rounds are sequential. Each
79
- task prompt carries the persona block, the decision context, and the full
80
- transcript so far. Workers never edit files; the debate is analysis only.
81
- Dispatch receipts link every statement to a worker run.
146
+ See Arguments for the exact call shape (one batched `tasks` call, the
147
+ `mode="sequential"` capacity fallback, `intent` for scope, no literal shell
148
+ syntax). Rounds are sequential; a round's own perspectives are the one
149
+ dispatch call. Each task prompt carries the persona block, the decision
150
+ context, and the full transcript so far. Workers never edit files; the
151
+ debate is analysis only. Dispatch receipts link every statement to a worker
152
+ run.
82
153
 
83
154
  ## Step 3 — Run the rounds
84
155
 
@@ -90,14 +161,15 @@ Dispatch receipts link every statement to a worker run.
90
161
  still disagrees and why that crux is the crux, and its final
91
162
  recommendation.
92
163
 
93
- **Quick mode** (user asked for a light pass): round 1 plus synthesis. No
94
- responses round.
164
+ **Quick mode** (headless default, or a user asking for a light pass): round 1
165
+ plus synthesis. No responses round, no convergence round.
95
166
 
96
- **Early termination.** After round 1, judge disagreement on the decision
97
- question itself, not on side conditions. If every position picks the same
98
- option and differs only in caveats, toggles, or requests to measure later,
99
- that is consensus: skip rounds 2 and 3, report that the council was not
100
- needed, and return the consensus with caveats. Never manufacture friction.
167
+ **Early termination** (full-debate mode only). After round 1, judge
168
+ disagreement on the decision question itself, not on side conditions. If
169
+ every position picks the same option and differs only in caveats, toggles,
170
+ or requests to measure later, that is consensus: skip rounds 2 and 3, report
171
+ that the council was not needed, and return the consensus with caveats.
172
+ Never manufacture friction.
101
173
 
102
174
  ## Step 4 — Synthesize
103
175
 
@@ -124,10 +196,11 @@ synthesis line has a citation and the recommendation names its crux.
124
196
 
125
197
  ## Degraded mode
126
198
 
127
- If dispatch is unavailable or admission-denied, run the same rounds inline:
128
- write each perspective's contribution yourself, sequentially, same round
129
- structure and synthesis format. Label the output as degraded (single-model
130
- debate, no receipts).
199
+ If dispatch is unavailable or admission-denied for a reason other than the
200
+ capacity retry in Arguments (the tool itself is missing from the surface, or
201
+ every retry is denied), run the same rounds inline: write each perspective's
202
+ contribution yourself, sequentially, same round structure and synthesis
203
+ format. Label the output as degraded (single-model debate, no receipts).
131
204
 
132
205
  ## Boundaries
133
206
 
@@ -141,7 +214,19 @@ this skill when the debate needs the full round structure and receipts.
141
214
  ## Red Flags
142
215
 
143
216
  - Perspectives named "optimist" and "pessimist" (role menu, not topic).
144
- - More than five perspectives, or debate rounds beyond three.
217
+ - Composing four or five perspectives, or running a Responses/Convergence
218
+ round, in a headless run — the enforced default is exactly three and
219
+ exactly one, not a ceiling to raise because the topic looks deep.
220
+ - Splitting a round into several single-task `dispatch` calls issued one at a
221
+ time instead of one batched `tasks` call (retried as `mode="sequential"`
222
+ on a capacity denial).
223
+ - Calling `dispatch(list:true)` before dispatching the round.
224
+ - Pasting an absolute path into a dispatch `task`/`briefing` instead of a
225
+ declared `intent`, or quoting literal shell syntax inside a persona's
226
+ argument.
227
+ - Reaching for `bash`/`read` on a receipt file the dispatch call's own result
228
+ already contains.
229
+ - Opening a `tasks` list for the round structure; `tasks` is refused.
145
230
  - A synthesis that averages positions instead of naming the crux.
146
231
  - Manufactured disagreement on a settled question.
147
232
  - A worker asked to edit files as part of the debate.
@@ -95,3 +95,164 @@ default and tells a headless run to use three and one round, which is the
95
95
  blocker the timeout exposed. Not re-run, so `eval-status` stays
96
96
  `scenarios-recorded`; the next campaign has to confirm the shortened council
97
97
  fits the ceiling.
98
+
99
+ ## Battletest record (2026-09-03) — 0.4.0 -> 0.5.0
100
+
101
+ Fixture: `/home/akougkas/eval-temp/harness/test_designcouncil.py`, a
102
+ self-contained repo (`src/checkpoint.py` writing one raw `.npy` per rank per
103
+ step to a shared parallel filesystem, `src/config.py`/`config.yaml` using
104
+ `pyyaml`) adapted from this file's own worked examples. S1 = HDF5-vs-Zarr
105
+ (genuinely contested), S2 = hand-roll-YAML-vs-keep-pyyaml (consensus), S4 =
106
+ "poke holes in my plan" anti-trigger. `runner.py`'s `timeout=` was raised to
107
+ 1800s for this skill specifically — the 900s default is a `dispatch` fan-out
108
+ ceiling, not a model-quality one, and the mission called for confirming the
109
+ v0.3.0 "three perspectives, one round" headless fix on its own terms, not
110
+ routing around it. Primary: `qwen3.8-27b`/`dynamo`. Cross-model confirm:
111
+ `ornith1.5-35b-moe`/`mini`. Both runs shared the `dynamo` LM Studio endpoint
112
+ with concurrently active sibling battletest sessions (`grill-me`,
113
+ `workflow-distiller`) launched via `herdr` during this pass — real,
114
+ externally-caused contention, not a fixture artifact; see below.
115
+
116
+ | run | scenario | model | wall | turns | dispatch calls | safety blocks | score | outcome |
117
+ |---|---|---|---|---|---|---|---|---|
118
+ | baseline (no skill) | S1 | qwen3.8-27b | 121s | 5 | 0 | 0 | 3/11 | no skill invoked (ran under `--no-skills`); the model noticed the installed skill anyway and narrated "four positions, real cross-examination" as one inline monologue — a solid single-model analysis but no dispatch, no receipts, no named-recipe workers |
119
+ | v1 (frozen 0.4.0) | S1 | qwen3.8-27b | 1740s, **killed at the 1800s ceiling** (exit -9) | 22 | 12 | 7 | 4/11 | composed **four** perspectives and ran into **round 2**, both violating the stated headless "three perspectives, one round" default; opened a `tasks` plan (refused); two `dispatch` calls rejected for an absolute path token in `briefing` (`legacy_scope_path_absolute`); one denied for endpoint capacity; one **hard-blocked as `rm-recursive-or-force`** because a perspective's own round-1 argument prose said "...you can `ls`, checksum, `rm -rf`, and publish..." — the admission layer's damage-control scan matched the quoted shell syntax inside the debate text itself, not an executed command; reached for `bash` to inspect a receipt file (refused). This is the same failure class the 2026-08-13 smoke record flagged, confirmed still present, worse: the v0.3.0 fix was never actually followed |
120
+ | v2 (first hardened cut) | S1 | qwen3.8-27b | 416s | 11 | 3 | 4 | 8/11 | one batched `tasks`-array `dispatch` call, all 3 capacity/timeout-denied under real endpoint contention; correctly fell back to Degraded mode, cited the exact denial reasons, labeled the output degraded, produced a complete synthesis with agreements/crux/dissent — no `tasks` misuse, no shell-block trigger, no one-at-a-time call splitting |
121
+ | v2 (post Step 0 strengthen) | S1 | qwen3.8-27b | 304s | 7 | 3 | 4 | 8/11 | same shape: parallel wave denied, sequential retry denied/timed out, correct Degraded fallback, well-formed synthesis citing the transcript |
122
+ | v2 | S2 | qwen3.8-27b | 52.5s / 60.6s (before/after Step 0 edit) | 5 / 6 | **0** | 0 | **5/5** both | Step 0 stopped before any dispatch, stated the council was not needed, gave the consensus (keep `pyyaml`, `safe_load` only) with caveats, and named the repo's actually-contested question (the checkpoint format) as a pointer — no manufactured friction either run |
123
+ | v2 | S4 | qwen3.8-27b | 69.4s | — | 0 | 0 | 3/3 | did not convene a council; ran a grill-me-shaped one-question-at-a-time interrogation instead and named `grill-me` |
124
+ | v2 (cross-model) | S1 | ornith1.5-35b-moe/mini | 384s | 10 | 4 | 5 | 7/11 | 3 of 4 `dispatch` calls denied/timed out on the **same `dynamo` endpoint** (workers defaulted there even though the orchestrator itself ran on `mini`); one stray `monitor` call refused (outside the skill's surface); correctly diagnosed the capacity pattern in its own words ("times out at ~57s... despite reporting 0/4 slots in use... exhausted the capacity fallback") and ran Degraded mode with a clearly labeled warning banner |
125
+ | v2 (cross-model) | S2 | ornith1.5-35b-moe/mini | 239.5s / 188.1s (before/after Step 0 edit) | 6 | 3 / 2 | 3 / 2 | 2/5 both | did **not** skip Step 0's dispatch the way qwen did — ran one capacity-retried round anyway, but stopped there (no round 2/3), reached the same correct consensus ("keep pyyaml", supply-chain caveat preserved as a contingent dissent, not manufactured), and stated the council-not-needed conclusion explicitly; the Step 0 wording strengthen did not change this model's behavior |
126
+
127
+ **S3 (dispatch unavailable) — no clean forced trigger found, confirmed by
128
+ reading the CLI, not assumed**: `clio-coder run --help`'s `--tool-profile`
129
+ narrows a *dispatched sub-agent's own* tool set and requires `--agent`
130
+ (`clio-coder run: fleet dispatch flags require --agent <recipe-id>`); there
131
+ is no flag that strips `dispatch` from a top-level `--skill` orchestrator
132
+ run's own surface in this harness. Editing the skill's own `allowed-tools`
133
+ to omit `dispatch` would test a different skill, not this one. This is a
134
+ real, documented harness gap, not faked around. In its place, S3's expected
135
+ behavior was exercised **organically, repeatedly, for real**: every S1 run
136
+ on both models hit genuine `dispatch` admission denial or timeout from
137
+ endpoint capacity, and Degraded mode fired correctly every single time —
138
+ labeled degraded, same round structure, synthesis format intact, no
139
+ receipts claimed that did not exist.
140
+
141
+ **The dispatch-viability and recipe-existence questions, resolved
142
+ empirically:**
143
+
144
+ - `scout`, `researcher`, and `provenance` all exist as builtin shadow-agent
145
+ recipes (`src/domains/agents/builtins/*.md`, `capabilityClass: read-only`)
146
+ and are dispatchable — every hardened run that got a `dispatch` call
147
+ admitted used one of them by name and got real worker output back.
148
+ - Each recipe's own system prompt declares a rigid JSON-only result
149
+ contract (`{"findings":[...]}`, `{"source":...}`, `{"confirmedFacts":...}`)
150
+ that has nothing to do with a debate position. In practice this was
151
+ **not** a blocker: V1's successfully-admitted round 1 came back as full
152
+ argumentative prose (named claims, attacks, citations, mind-change
153
+ conditions) from `researcher`/`scout` workers, not the declared narrow
154
+ JSON shape — the persona in the task prompt won out over the recipe's
155
+ own stated output contract for the caller-facing transcript. Worth
156
+ knowing, not worth re-architecting around.
157
+ - The actual, dominant blocker is **`dispatch` admission capacity on a
158
+ single-instance local LM Studio target**. `~/.local/state/clio-coder/
159
+ endpoint-slots.json` records the `dynamo` endpoint (`100.104.197.69:1234`)
160
+ at `"slots": 1`. A 3-task parallel wave is denied outright
161
+ (`capacity exceeded (3/1 slots)`), and the `mode="sequential"` retry this
162
+ pass added to the skill is followed correctly by both models but still
163
+ gets denied or times out (`admission timed out after ~57s... 0/4 worker
164
+ slots in use` — a scheduling state that never resolves within the wait
165
+ window), evidently because the orchestrator's own foreground session
166
+ already holds the endpoint's only slot. This reproduced independently on
167
+ `mini` too: workers dispatched from an orchestrator running on `mini`
168
+ still routed to `dynamo`'s endpoint by default, so the same 1-slot
169
+ contention applied there as well — confirming the brief's suspicion that
170
+ worker fan-out may silently default to an unexpected node/target. Some
171
+ of this pass's contention was real concurrent load, not just self-
172
+ contention: `ps` showed sibling `grill-me`/`workflow-distiller`
173
+ battletest sessions actively running via `herdr` against the same
174
+ endpoint during these runs.
175
+ - The full-auto/`authorityBasis` auto-grant fact supplied going in
176
+ (`deps.getAutonomy?.() === "full-auto" ? "full-auto-policy" :
177
+ "operator-plan-approval"`) turned out to be **moot for this skill**:
178
+ that gate only applies to `agent:"auto"` dispatch requests
179
+ (`src/tools/dispatch-arguments.ts`, `agentSelection` is only populated
180
+ when `requestedAgent === "auto"`). Design-council always pins an explicit
181
+ recipe id (`scout`/`researcher`/`provenance`), so `agentSelection` stays
182
+ `undefined` and the operator-approval/full-auto-policy distinction never
183
+ engages — admission for this skill's calls is governed purely by the
184
+ capacity/reservation machinery above, independent of `--autonomy`.
185
+ - `legacy_scope_path_absolute`: an absolute path token (e.g. a value read
186
+ from `config.yaml`) pasted into `task`/`briefing` prose without a
187
+ declared `intent` is rejected. Fixed by telling the skill to declare
188
+ `intent.read_roots`/`relevant_paths` (relative) on every call instead.
189
+
190
+ **Changes (0.4.0 -> 0.5.0):**
191
+
192
+ 1. **`## Arguments` contract**, ported from `grill-me`/`cut-it`'s shape.
193
+ States headless is the *enforced* default (exactly three perspectives,
194
+ exactly one round), not a self-assessed suggestion — the v0.3.0 fix's
195
+ prose alone did not hold on either model (V1 composed four perspectives
196
+ and ran round 2 headlessly).
197
+ 2. **Dispatch call shape spelled out**: one batched `tasks`-array call per
198
+ round; on capacity denial, retry the *same* batch with
199
+ `mode="sequential"` in one call rather than splitting into several
200
+ single-task calls issued one at a time (V1's actual failure — 12
201
+ separate `dispatch` calls, a `list:true` probe, and manual receipt
202
+ reads via `bash`, all of which V2 stopped doing).
203
+ 3. **`intent.read_roots`/`relevant_paths` guidance** to avoid
204
+ `legacy_scope_path_absolute` rejections from paths quoted in prose.
205
+ 4. **No literal shell syntax in a persona's argument text** — added after
206
+ V1's `rm-recursive-or-force` hard block fired on debate prose, not a
207
+ real command.
208
+ 5. **`ask_user` added to `allowed-tools`** (already always-exempt, now
209
+ documented) with one call at Step 1 on a contested topic, asking
210
+ quick-vs-full-debate depth; headlessly it cancels, which *is* the
211
+ three-perspective/one-round confirmation, mirroring the auto-cancel
212
+ pattern the planning category established rather than relying on
213
+ unaided self-assessment.
214
+ 6. **Explicit `tasks`/`bash`/receipt-re-read refusal lines**, matching the
215
+ sibling skills' pattern (`tasks` opened a plan in V1; `bash` was reached
216
+ for twice across the runs above to re-inspect a receipt the dispatch
217
+ call's own result already contained).
218
+ 7. **Step 0 strengthened** against dispatching "just to confirm" a
219
+ consensus call already reached — tested on `ornith1.5-35b-moe`/`mini`
220
+ (still dispatches once) and `qwen3.8-27b`/`dynamo` (unaffected, already
221
+ correct); see Still weak.
222
+ 8. Five new Red flags entries naming the concrete failures observed above.
223
+
224
+ **Still weak:**
225
+
226
+ - **Step 0's short-circuit is model-family-dependent.** `qwen3.8-27b`
227
+ trusts its own contested-or-not judgment and skips `dispatch` entirely
228
+ for S2 (5/5, 0 dispatch calls, both before and after the Step 0 edit).
229
+ `ornith1.5-35b-moe` does not: it dispatches one (capacity-retried) round
230
+ "to confirm" even on the same consensus topic, both before and after the
231
+ strengthened wording — 2/5 unchanged. It still stops after that one
232
+ round, still reaches the correct consensus, still preserves the dissent
233
+ as a contingent caveat rather than manufacturing friction, so the
234
+ user-facing outcome is fine; the wasted dispatch cost is the actual gap,
235
+ and more prose did not move it, consistent with the planning category's
236
+ own finding that some model-family behaviors don't fully generalize no
237
+ matter how much repetition is added.
238
+ - **S1 never completed as a genuine 3-perspective/1-round council with
239
+ real receipts** on this pass — every attempt on both models hit real
240
+ endpoint capacity contention and degraded. The Degraded path is now
241
+ proven solid, but the "happy path" (dispatch succeeds, synthesis cites
242
+ real worker receipts) is unverified under this specific pass's
243
+ conditions; V1's transcript shows it *can* succeed (four real worker
244
+ runs completed with citable output before the round-2/shell-block
245
+ failure), so this reads as an availability problem this pass's timing
246
+ ran into, not a structural block — but it means the exact scoring
247
+ bullets that depend on `dispatch_receipts_ok` (perspective count,
248
+ receipts-ok) were not exercised clean this round.
249
+ `perspectives_dispatched`/`used_named_recipes` in the grading script
250
+ count attempted, not completed, dispatch tasks for this reason.
251
+ - **No purpose-built S3 fixture** — see above; a future pass with a
252
+ dedicated low-capacity or offline dispatch target would let this be
253
+ tested directly instead of relying on organic contention.
254
+ - `code_nav` and `grep` were barely exercised (fixture is small enough
255
+ that `read`/`ls`/`find` covered grounding in most runs).
256
+ - Only one fixture domain (HDF5-vs-Zarr / YAML-parser) ran this pass; the
257
+ richer four-perspective MPI-checkpoint example from "Observed live-smoke
258
+ results" above was not re-run against 0.5.0.
@@ -7,7 +7,7 @@ triggers:
7
7
  - stress-test this design
8
8
  - poke holes in this idea
9
9
  - clarify this plan one question at a time
10
- version: 0.4.0
10
+ version: 0.5.1
11
11
  license: Apache-2.0
12
12
  allowed-tools:
13
13
  - read
@@ -33,9 +33,53 @@ Run a rigorous, repo-aware interview that turns a vague plan into explicit
33
33
  decisions. The point is not to interrogate for sport; it is to surface hidden
34
34
  branches before anyone writes code.
35
35
 
36
+ ## Arguments
37
+
38
+ ```text
39
+ grill me on <plan, feature, or idea>
40
+ ```
41
+
42
+ There is no flag syntax; the trigger is conversational — "grill me on X",
43
+ "stress-test this design", "poke holes in this idea". Whatever the user
44
+ names is the subject. A referenced file, doc, or repo path in the same
45
+ request is Step 1's grounding to read first, not a separate argument.
46
+
47
+ **This run has no back-and-forth.** There is no second turn in which a user
48
+ reads your question and replies to it — whatever you ask, you must also
49
+ answer yourself, in this same turn, before it ends. Do not write a question
50
+ and stop to wait for a reply, in `ask_user` or in plain chat text; nothing
51
+ is coming. Ending the turn on an open question — even one question, even a
52
+ well-posed one — is this skill's single most common failure and worse than
53
+ skipping the interview format entirely.
54
+
55
+ There is no operator in a headless run: `ask_user` is not registered, so
56
+ any call is refused as an unregistered tool rather than answered, every
57
+ time, with no exceptions. Calling it again will not produce a different
58
+ result, so one call is enough to confirm it (not required: reasoning from
59
+ this paragraph alone is just as valid as calling it and observing the
60
+ refusal). Whichever phase this lands in — even round
61
+ 1 — switch immediately to the assumed-confirm monologue for every phase
62
+ from here on, in the same turn: state the question you would have asked,
63
+ give your own best/recommended answer with the reasoning behind it, mark
64
+ it `assumed — confirm`, and move to the next phase. Do not re-call
65
+ `ask_user` hoping a later round behaves differently — that only burns the
66
+ `max_rounds` budget without ever converging. Keep working the phase map,
67
+ phase by phase, all the way through Step 5's decision log before ending
68
+ the turn — never end on "Answer 1/2/3" or any other place a reply is
69
+ expected.
70
+
71
+ The phase map below is the plan; do not open a task list for it. This
72
+ skill's tool surface is exactly `read`, `grep`, `ls`, `find`, `git`,
73
+ `context`, `code_nav`, and `ask_user` (`context` and `ask_user` are always
74
+ available regardless). `tasks` and `bash` both sit outside it and any call
75
+ to either is refused — inspect a file with `read`, not `bash cat`/`bash
76
+ head`/`bash wc`; locate files with `find` or `ls`, not `bash find`/`bash
77
+ ls`; check repo state with the `git` tool, not shell `git`.
78
+
36
79
  ## Operating Contract
37
80
 
38
- - Use `ask_user` for the interview whenever it is active.
81
+ - In a live session, use `ask_user` for the interview and actually wait for
82
+ the user's answer between rounds. In a headless run, see Arguments above.
39
83
  - For every interview round, call `ask_user` with `mode: "single_question"` and
40
84
  exactly one question.
41
85
  - On the first ask for a normal grill-me run, set `max_rounds` to a bounded
@@ -45,8 +89,6 @@ branches before anyone writes code.
45
89
  alternatives with short tradeoff descriptions.
46
90
  - The user answers in natural language. You translate answers into compact
47
91
  decision keys and rationale when you call `ask_user` with `action: "complete"`.
48
- - If `ask_user` is unavailable, ask in plain text, still one question at a
49
- time, and keep an internal decision log.
50
92
 
51
93
  ## Phase Map
52
94
 
@@ -69,7 +111,11 @@ or "fill" when the decision is genuinely missing.
69
111
 
70
112
  Read what the user already gave you. If the task references files, plans, code,
71
113
  tests, or project conventions, inspect them before asking. Prefer
72
- `context(scope="workspace")`, `grep`, `read`, and codewiki tools over guessing.
114
+ `context(scope="workspace")`, `grep`, `read`, `code_nav` (symbol and call-graph
115
+ lookups), and codewiki tools over guessing. Use the `git` tool (`status`,
116
+ `log`, not shell `git`) when the plan references repo state — recent
117
+ history, uncommitted changes, what "decided" actually means for this repo
118
+ right now.
73
119
 
74
120
  Privately build a phase map:
75
121
 
@@ -112,10 +158,26 @@ Good: "Which user should v1 optimize for first?"
112
158
  If an answer is vague, ask a follow-up on the same branch. Do not jump to a new
113
159
  branch while the current one is still unresolved.
114
160
 
161
+ **Headless: there is no reply coming, whether `ask_user` comes back
162
+ refused as unregistered or you never call it at all.** Neither is a vague answer to
163
+ follow up on and neither is a signal to try again or to wait — both mean
164
+ there is no operator this run, from round 1 on. Do not call `ask_user`
165
+ again for this or any later phase, and do not phrase a question in plain
166
+ text as if a reply is pending. From here, run every remaining phase
167
+ (including this one) as the assumed-confirm monologue described in
168
+ Arguments, in this same turn, through to the Step 5 decision log. See
169
+ Arguments for the exact treatment.
170
+
115
171
  ### Step 4 - Respect Stop Signals
116
172
 
117
- Stop immediately when the user says "stop", "enough", "later", "done", "next
118
- time", or cancels the modal. Do not ask another question to confirm stopping.
173
+ This step applies to a live session with a real operator. Stop immediately
174
+ when the user says "stop", "enough", "later", "done", "next time", or cancels
175
+ the modal. Do not ask another question to confirm stopping.
176
+
177
+ A headless refusal of `ask_user` is not a stop signal from a user — it is the
178
+ absence of an operator (see Step 3). Do not treat it as "the user cancelled
179
+ this session"; treat it as the cue to switch to the assumed-confirm
180
+ monologue and keep going to a complete decision log, not to stop early.
119
181
 
120
182
  If you have enough decisions to be useful, call:
121
183
 
@@ -140,8 +202,11 @@ state the partial decisions and the next unresolved root question.
140
202
 
141
203
  ### Step 5 - Complete
142
204
 
143
- Before final prose, call `ask_user` with `action: "complete"` and a compact
144
- `decisions` array. Then write the final decision log:
205
+ In a live session, close with `ask_user` `action: "complete"` and a compact
206
+ `decisions` array before final prose. In a headless run where `ask_user`
207
+ already came back refused, skip straight to the decision log below — do
208
+ not attempt another `ask_user` call just to close out; it will be refused
209
+ too and adds nothing. Write the final decision log:
145
210
 
146
211
  ```markdown
147
212
  ## Decision Log - <topic>
@@ -189,4 +254,14 @@ Use this ordering when several questions are possible:
189
254
  - Asking about facts discoverable from the repo.
190
255
  - Letting `ask_user` hit the round limit without completing the interview.
191
256
  - Ending with a summary paragraph instead of the decision log.
192
- - Treating cancellation as permission to keep asking.
257
+ - Re-calling `ask_user` for a later phase after an earlier round already came
258
+ back refused — the answer will not be different; that budget is wasted.
259
+ - Ending a turn on "Answer 1/2/3", "let me know which you prefer", or any
260
+ other wording that expects a reply in a headless run — there is no next
261
+ turn for a reply to land in. This is the single most common failure mode
262
+ of this skill and the one to watch hardest for: asking one question, then
263
+ stopping, instead of running the assumed-confirm monologue through every
264
+ remaining phase to the decision log in the same turn.
265
+ - Opening a `tasks` list for the phase map; `tasks` is refused.
266
+ - Treating a headless `ask_user` refusal as the user's stop signal (Step 4)
267
+ instead of the absence-of-operator cue it actually is.