dsh-aris-panel 0.1.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (442) hide show
  1. package/LICENSE +21 -0
  2. package/README.md +98 -0
  3. package/README_CN.md +87 -0
  4. package/dsh/checkout.patch.yml +38 -0
  5. package/dsh/client.js +634 -0
  6. package/dsh/cordis.patch.yml +44 -0
  7. package/dsh/index.mjs +76 -0
  8. package/dsh/run-status.mjs +182 -0
  9. package/dsh/scope-limits.mjs +50 -0
  10. package/dsh/workbench.mjs +291 -0
  11. package/mcp-servers/claude-review/README.md +93 -0
  12. package/mcp-servers/claude-review/run_with_claude_aws.sh +49 -0
  13. package/mcp-servers/claude-review/server.py +718 -0
  14. package/mcp-servers/codex-image2/README.md +65 -0
  15. package/mcp-servers/codex-image2/server.py +893 -0
  16. package/mcp-servers/feishu-bridge/requirements.txt +1 -0
  17. package/mcp-servers/feishu-bridge/server.py +240 -0
  18. package/mcp-servers/gemini-review/README.md +171 -0
  19. package/mcp-servers/gemini-review/server.py +1856 -0
  20. package/mcp-servers/llm-chat/requirements.txt +1 -0
  21. package/mcp-servers/llm-chat/server.py +664 -0
  22. package/mcp-servers/manual-review/README.md +133 -0
  23. package/mcp-servers/manual-review/server.py +910 -0
  24. package/mcp-servers/manual-review/ui.html +279 -0
  25. package/mcp-servers/minimax-chat/requirements.txt +1 -0
  26. package/mcp-servers/minimax-chat/server.py +381 -0
  27. package/package.json +51 -0
  28. package/skills/ablation-planner/SKILL.md +123 -0
  29. package/skills/alphaxiv/SKILL.md +196 -0
  30. package/skills/analyze-results/SKILL.md +46 -0
  31. package/skills/arxiv/SKILL.md +248 -0
  32. package/skills/auto-paper-improvement-loop/SKILL.md +651 -0
  33. package/skills/auto-review-loop/SKILL.md +1137 -0
  34. package/skills/auto-review-loop-llm/SKILL.md +259 -0
  35. package/skills/auto-review-loop-minimax/SKILL.md +302 -0
  36. package/skills/citation-audit/SKILL.md +502 -0
  37. package/skills/claims-drafting/SKILL.md +227 -0
  38. package/skills/comm-lit-review/SKILL.md +297 -0
  39. package/skills/deepxiv/SKILL.md +263 -0
  40. package/skills/dse-loop/SKILL.md +296 -0
  41. package/skills/embodiment-description/SKILL.md +129 -0
  42. package/skills/exa-search/SKILL.md +205 -0
  43. package/skills/experiment-audit/SKILL.md +311 -0
  44. package/skills/experiment-bridge/SKILL.md +376 -0
  45. package/skills/experiment-plan/SKILL.md +249 -0
  46. package/skills/experiment-queue/SKILL.md +431 -0
  47. package/skills/experiment-queue/scripts/build_manifest.py +142 -0
  48. package/skills/experiment-queue/scripts/queue_manager.py +433 -0
  49. package/skills/feishu-notify/SKILL.md +156 -0
  50. package/skills/figure-description/SKILL.md +138 -0
  51. package/skills/figure-spec/SKILL.md +262 -0
  52. package/skills/figure-spec/scripts/figure_renderer.py +799 -0
  53. package/skills/formula-derivation/SKILL.md +280 -0
  54. package/skills/gemini-search/SKILL.md +231 -0
  55. package/skills/grant-proposal/SKILL.md +698 -0
  56. package/skills/idea-creator/SKILL.md +542 -0
  57. package/skills/idea-discovery/SKILL.md +521 -0
  58. package/skills/idea-discovery-robot/SKILL.md +363 -0
  59. package/skills/integrity-forensics/SKILL.md +284 -0
  60. package/skills/interview-cheatsheet/SKILL.md +245 -0
  61. package/skills/invention-structuring/SKILL.md +188 -0
  62. package/skills/jurisdiction-format/SKILL.md +192 -0
  63. package/skills/kill-argument/SKILL.md +437 -0
  64. package/skills/mermaid-diagram/SKILL.md +419 -0
  65. package/skills/meta-apply/SKILL.md +141 -0
  66. package/skills/meta-optimize/SKILL.md +437 -0
  67. package/skills/monitor-experiment/SKILL.md +140 -0
  68. package/skills/novelty-check/SKILL.md +101 -0
  69. package/skills/openalex/SKILL.md +237 -0
  70. package/skills/overleaf-sync/SKILL.md +220 -0
  71. package/skills/paper-claim-audit/SKILL.md +348 -0
  72. package/skills/paper-compile/SKILL.md +266 -0
  73. package/skills/paper-figure/SKILL.md +312 -0
  74. package/skills/paper-illustration/SKILL.md +736 -0
  75. package/skills/paper-illustration-image2/SKILL.md +391 -0
  76. package/skills/paper-illustration-image2/scripts/paper_illustration_image2.py +255 -0
  77. package/skills/paper-plan/SKILL.md +386 -0
  78. package/skills/paper-poster/SKILL.md +19 -0
  79. package/skills/paper-poster-html/DESIGN_FINAL.md +176 -0
  80. package/skills/paper-poster-html/IMPLEMENTATION_CONVENTIONS.md +161 -0
  81. package/skills/paper-poster-html/LICENSES/posterly-MIT.txt +21 -0
  82. package/skills/paper-poster-html/NOTICE.md +57 -0
  83. package/skills/paper-poster-html/SKILL.md +323 -0
  84. package/skills/paper-poster-html/scripts/_posterly/__init__.py +0 -0
  85. package/skills/paper-poster-html/scripts/_posterly/canvas.py +200 -0
  86. package/skills/paper-poster-html/scripts/_posterly/measure.py +588 -0
  87. package/skills/paper-poster-html/scripts/_posterly/polish.py +498 -0
  88. package/skills/paper-poster-html/scripts/_posterly/preflight.py +489 -0
  89. package/skills/paper-poster-html/scripts/_posterly/render.py +215 -0
  90. package/skills/paper-poster-html/scripts/_posterly/textutil.py +16 -0
  91. package/skills/paper-poster-html/scripts/_posterly/verify_final.py +171 -0
  92. package/skills/paper-poster-html/scripts/asset_check.py +897 -0
  93. package/skills/paper-poster-html/scripts/extract_pdf_figures.py +666 -0
  94. package/skills/paper-poster-html/scripts/poster_check.py +251 -0
  95. package/skills/paper-poster-html/scripts/preprocess_figures.py +238 -0
  96. package/skills/paper-poster-html/scripts/render_preview.py +217 -0
  97. package/skills/paper-poster-html/scripts/run_gates.py +556 -0
  98. package/skills/paper-poster-html/scripts/style_check.py +1324 -0
  99. package/skills/paper-poster-html/templates/COMPONENTS.md +462 -0
  100. package/skills/paper-poster-html/templates/README.md +170 -0
  101. package/skills/paper-poster-html/templates/landscape_4col.html +1032 -0
  102. package/skills/paper-poster-html/templates/landscape_hero.html +1046 -0
  103. package/skills/paper-poster-html/templates/portrait_2col.html +947 -0
  104. package/skills/paper-poster-html/templates/tokens/acl.json +9 -0
  105. package/skills/paper-poster-html/templates/tokens/cvpr.json +9 -0
  106. package/skills/paper-poster-html/templates/tokens/generic.json +9 -0
  107. package/skills/paper-poster-html/templates/tokens/iclr.json +9 -0
  108. package/skills/paper-poster-html/templates/tokens/icml.json +9 -0
  109. package/skills/paper-poster-html/templates/tokens/neurips.json +9 -0
  110. package/skills/paper-slides/SKILL.md +635 -0
  111. package/skills/paper-talk/SKILL.md +381 -0
  112. package/skills/paper-write/SKILL.md +604 -0
  113. package/skills/paper-write/templates/IEEEtran.bst +2409 -0
  114. package/skills/paper-write/templates/IEEEtran.cls +6347 -0
  115. package/skills/paper-write/templates/iclr2026.tex +84 -0
  116. package/skills/paper-write/templates/icml2025.tex +87 -0
  117. package/skills/paper-write/templates/ieee_conference.tex +89 -0
  118. package/skills/paper-write/templates/ieee_journal.tex +93 -0
  119. package/skills/paper-write/templates/math_commands.tex +48 -0
  120. package/skills/paper-write/templates/neurips2025.tex +80 -0
  121. package/skills/paper-writing/SKILL.md +916 -0
  122. package/skills/patent-novelty-check/SKILL.md +153 -0
  123. package/skills/patent-pipeline/SKILL.md +344 -0
  124. package/skills/patent-review/SKILL.md +203 -0
  125. package/skills/pixel-art/SKILL.md +137 -0
  126. package/skills/prior-art-search/SKILL.md +146 -0
  127. package/skills/proof-checker/SKILL.md +866 -0
  128. package/skills/proof-orchestrator/NOTICE.md +24 -0
  129. package/skills/proof-orchestrator/SKILL.md +254 -0
  130. package/skills/proof-orchestrator/references/audit-output-contract.md +126 -0
  131. package/skills/proof-orchestrator/references/deepseek-routing.md +74 -0
  132. package/skills/proof-orchestrator/references/dispatch-prompts.md +227 -0
  133. package/skills/proof-orchestrator/references/notation-audit.md +135 -0
  134. package/skills/proof-orchestrator/references/proof-audit-rubric.md +70 -0
  135. package/skills/proof-orchestrator/references/stress-tests.md +38 -0
  136. package/skills/proof-writer/SKILL.md +223 -0
  137. package/skills/qzcli/SKILL.md +324 -0
  138. package/skills/rebuttal/SKILL.md +376 -0
  139. package/skills/render-html/SKILL.md +316 -0
  140. package/skills/render-html/scripts/render_html.py +1006 -0
  141. package/skills/render-html/scripts/templates/academic.html +703 -0
  142. package/skills/render-html/scripts/templates/dashboard.html +333 -0
  143. package/skills/research-lit/SKILL.md +756 -0
  144. package/skills/research-pipeline/SKILL.md +384 -0
  145. package/skills/research-refine/SKILL.md +770 -0
  146. package/skills/research-refine-pipeline/SKILL.md +186 -0
  147. package/skills/research-review/SKILL.md +198 -0
  148. package/skills/research-wiki/SKILL.md +461 -0
  149. package/skills/resubmit-pipeline/SKILL.md +447 -0
  150. package/skills/result-to-claim/SKILL.md +311 -0
  151. package/skills/run-experiment/SKILL.md +313 -0
  152. package/skills/semantic-scholar/SKILL.md +236 -0
  153. package/skills/serverless-modal/SKILL.md +335 -0
  154. package/skills/shared-references/acceptance-gate.md +324 -0
  155. package/skills/shared-references/assurance-contract.md +248 -0
  156. package/skills/shared-references/capture-antipatterns.md +78 -0
  157. package/skills/shared-references/citation-discipline.md +583 -0
  158. package/skills/shared-references/compute-env-contract.md +163 -0
  159. package/skills/shared-references/effort-contract.md +183 -0
  160. package/skills/shared-references/evidence-precheck.md +65 -0
  161. package/skills/shared-references/experiment-integrity.md +49 -0
  162. package/skills/shared-references/external-cadence.md +326 -0
  163. package/skills/shared-references/fan-out-pattern.md +366 -0
  164. package/skills/shared-references/injection-hygiene.md +127 -0
  165. package/skills/shared-references/integration-contract.md +461 -0
  166. package/skills/shared-references/output-composition.md +93 -0
  167. package/skills/shared-references/output-language.md +45 -0
  168. package/skills/shared-references/output-manifest.md +49 -0
  169. package/skills/shared-references/output-versioning.md +111 -0
  170. package/skills/shared-references/patent-format-cn.md +199 -0
  171. package/skills/shared-references/patent-format-ep.md +173 -0
  172. package/skills/shared-references/patent-format-us.md +161 -0
  173. package/skills/shared-references/patent-writing-principles.md +197 -0
  174. package/skills/shared-references/prior-art-databases.md +141 -0
  175. package/skills/shared-references/resumable-runs.md +109 -0
  176. package/skills/shared-references/review-scope-limits.md +81 -0
  177. package/skills/shared-references/review-tracing.md +391 -0
  178. package/skills/shared-references/reviewer-independence.md +79 -0
  179. package/skills/shared-references/reviewer-routing.md +852 -0
  180. package/skills/shared-references/skill-governance.md +104 -0
  181. package/skills/shared-references/taste-calibration.md +85 -0
  182. package/skills/shared-references/venue-checklists.md +114 -0
  183. package/skills/shared-references/wiki-helper-resolution.md +134 -0
  184. package/skills/shared-references/writing-principles.md +525 -0
  185. package/skills/skills-codex/README.md +102 -0
  186. package/skills/skills-codex/README_CN.md +100 -0
  187. package/skills/skills-codex/ablation-planner/SKILL.md +126 -0
  188. package/skills/skills-codex/alphaxiv/SKILL.md +186 -0
  189. package/skills/skills-codex/analyze-results/SKILL.md +45 -0
  190. package/skills/skills-codex/arxiv/SKILL.md +210 -0
  191. package/skills/skills-codex/auto-paper-improvement-loop/SKILL.md +574 -0
  192. package/skills/skills-codex/auto-review-loop/SKILL.md +500 -0
  193. package/skills/skills-codex/auto-review-loop-llm/SKILL.md +247 -0
  194. package/skills/skills-codex/auto-review-loop-minimax/SKILL.md +290 -0
  195. package/skills/skills-codex/citation-audit/SKILL.md +504 -0
  196. package/skills/skills-codex/claims-drafting/SKILL.md +239 -0
  197. package/skills/skills-codex/comm-lit-review/SKILL.md +299 -0
  198. package/skills/skills-codex/comm-lit-review/references/domain-taxonomy.md +57 -0
  199. package/skills/skills-codex/comm-lit-review/references/output-template.md +37 -0
  200. package/skills/skills-codex/comm-lit-review/references/source-policy.md +99 -0
  201. package/skills/skills-codex/comm-lit-review/references/venue-tiering.md +112 -0
  202. package/skills/skills-codex/deepxiv/SKILL.md +142 -0
  203. package/skills/skills-codex/dse-loop/SKILL.md +285 -0
  204. package/skills/skills-codex/embodiment-description/SKILL.md +129 -0
  205. package/skills/skills-codex/exa-search/SKILL.md +192 -0
  206. package/skills/skills-codex/experiment-audit/SKILL.md +286 -0
  207. package/skills/skills-codex/experiment-bridge/SKILL.md +356 -0
  208. package/skills/skills-codex/experiment-plan/SKILL.md +249 -0
  209. package/skills/skills-codex/experiment-queue/SKILL.md +401 -0
  210. package/skills/skills-codex/feishu-notify/SKILL.md +155 -0
  211. package/skills/skills-codex/figure-description/SKILL.md +138 -0
  212. package/skills/skills-codex/figure-spec/SKILL.md +252 -0
  213. package/skills/skills-codex/formula-derivation/SKILL.md +280 -0
  214. package/skills/skills-codex/gemini-search/SKILL.md +205 -0
  215. package/skills/skills-codex/grant-proposal/SKILL.md +626 -0
  216. package/skills/skills-codex/idea-creator/SKILL.md +405 -0
  217. package/skills/skills-codex/idea-discovery/SKILL.md +475 -0
  218. package/skills/skills-codex/idea-discovery-robot/SKILL.md +362 -0
  219. package/skills/skills-codex/integrity-forensics/SKILL.md +106 -0
  220. package/skills/skills-codex/interview-cheatsheet/SKILL.md +245 -0
  221. package/skills/skills-codex/invention-structuring/SKILL.md +188 -0
  222. package/skills/skills-codex/jurisdiction-format/SKILL.md +192 -0
  223. package/skills/skills-codex/kill-argument/SKILL.md +403 -0
  224. package/skills/skills-codex/mermaid-diagram/SKILL.md +379 -0
  225. package/skills/skills-codex/meta-apply/SKILL.md +154 -0
  226. package/skills/skills-codex/meta-optimize/SKILL.md +348 -0
  227. package/skills/skills-codex/monitor-experiment/SKILL.md +98 -0
  228. package/skills/skills-codex/novelty-check/SKILL.md +89 -0
  229. package/skills/skills-codex/openalex/SKILL.md +228 -0
  230. package/skills/skills-codex/overleaf-sync/SKILL.md +220 -0
  231. package/skills/skills-codex/paper-claim-audit/SKILL.md +350 -0
  232. package/skills/skills-codex/paper-compile/SKILL.md +253 -0
  233. package/skills/skills-codex/paper-figure/SKILL.md +311 -0
  234. package/skills/skills-codex/paper-illustration/SKILL.md +690 -0
  235. package/skills/skills-codex/paper-illustration-image2/SKILL.md +383 -0
  236. package/skills/skills-codex/paper-illustration-image2/scripts/paper_illustration_image2.py +255 -0
  237. package/skills/skills-codex/paper-plan/SKILL.md +278 -0
  238. package/skills/skills-codex/paper-poster/SKILL.md +19 -0
  239. package/skills/skills-codex/paper-poster-html/SKILL.md +377 -0
  240. package/skills/skills-codex/paper-slides/SKILL.md +571 -0
  241. package/skills/skills-codex/paper-talk/SKILL.md +381 -0
  242. package/skills/skills-codex/paper-write/SKILL.md +411 -0
  243. package/skills/skills-codex/paper-write/templates/IEEEtran.bst +2409 -0
  244. package/skills/skills-codex/paper-write/templates/IEEEtran.cls +6347 -0
  245. package/skills/skills-codex/paper-write/templates/aaai2026.bst +1493 -0
  246. package/skills/skills-codex/paper-write/templates/aaai2026.sty +315 -0
  247. package/skills/skills-codex/paper-write/templates/aaai2026.tex +952 -0
  248. package/skills/skills-codex/paper-write/templates/acl.sty +312 -0
  249. package/skills/skills-codex/paper-write/templates/acl2026.tex +377 -0
  250. package/skills/skills-codex/paper-write/templates/acl_natbib.bst +1940 -0
  251. package/skills/skills-codex/paper-write/templates/acm.bst +3081 -0
  252. package/skills/skills-codex/paper-write/templates/acm_mm2026.tex +204 -0
  253. package/skills/skills-codex/paper-write/templates/acmart.cls +3520 -0
  254. package/skills/skills-codex/paper-write/templates/cvpr.bst +1448 -0
  255. package/skills/skills-codex/paper-write/templates/cvpr.sty +508 -0
  256. package/skills/skills-codex/paper-write/templates/cvpr2026.tex +63 -0
  257. package/skills/skills-codex/paper-write/templates/iclr2026.tex +84 -0
  258. package/skills/skills-codex/paper-write/templates/iclr2026_conference.bst +1440 -0
  259. package/skills/skills-codex/paper-write/templates/iclr2026_conference.sty +246 -0
  260. package/skills/skills-codex/paper-write/templates/icml2026.sty +767 -0
  261. package/skills/skills-codex/paper-write/templates/icml2026.tex +662 -0
  262. package/skills/skills-codex/paper-write/templates/ieee_conference.tex +89 -0
  263. package/skills/skills-codex/paper-write/templates/ieee_journal.tex +93 -0
  264. package/skills/skills-codex/paper-write/templates/math_commands.tex +48 -0
  265. package/skills/skills-codex/paper-write/templates/neurips2026.tex +493 -0
  266. package/skills/skills-codex/paper-write/templates/neurips_2026.sty +437 -0
  267. package/skills/skills-codex/paper-writing/SKILL.md +731 -0
  268. package/skills/skills-codex/patent-novelty-check/SKILL.md +153 -0
  269. package/skills/skills-codex/patent-pipeline/SKILL.md +344 -0
  270. package/skills/skills-codex/patent-review/SKILL.md +202 -0
  271. package/skills/skills-codex/pixel-art/SKILL.md +139 -0
  272. package/skills/skills-codex/prior-art-search/SKILL.md +146 -0
  273. package/skills/skills-codex/proof-checker/SKILL.md +554 -0
  274. package/skills/skills-codex/proof-orchestrator/SKILL.md +260 -0
  275. package/skills/skills-codex/proof-orchestrator/references/audit-output-contract.md +126 -0
  276. package/skills/skills-codex/proof-orchestrator/references/deepseek-routing.md +76 -0
  277. package/skills/skills-codex/proof-orchestrator/references/dispatch-prompts.md +227 -0
  278. package/skills/skills-codex/proof-orchestrator/references/notation-audit.md +135 -0
  279. package/skills/skills-codex/proof-orchestrator/references/proof-audit-rubric.md +70 -0
  280. package/skills/skills-codex/proof-orchestrator/references/stress-tests.md +38 -0
  281. package/skills/skills-codex/proof-writer/SKILL.md +222 -0
  282. package/skills/skills-codex/qzcli/SKILL.md +324 -0
  283. package/skills/skills-codex/rebuttal/SKILL.md +305 -0
  284. package/skills/skills-codex/render-html/SKILL.md +305 -0
  285. package/skills/skills-codex/render-html/scripts/__pycache__/render_html.cpython-314.pyc +0 -0
  286. package/skills/skills-codex/render-html/scripts/render_html.py +909 -0
  287. package/skills/skills-codex/render-html/scripts/templates/academic.html +342 -0
  288. package/skills/skills-codex/render-html/scripts/templates/dashboard.html +333 -0
  289. package/skills/skills-codex/research-lit/SKILL.md +464 -0
  290. package/skills/skills-codex/research-pipeline/SKILL.md +340 -0
  291. package/skills/skills-codex/research-refine/SKILL.md +721 -0
  292. package/skills/skills-codex/research-refine-pipeline/SKILL.md +186 -0
  293. package/skills/skills-codex/research-review/SKILL.md +135 -0
  294. package/skills/skills-codex/research-wiki/SKILL.md +421 -0
  295. package/skills/skills-codex/resubmit-pipeline/SKILL.md +444 -0
  296. package/skills/skills-codex/result-to-claim/SKILL.md +246 -0
  297. package/skills/skills-codex/run-experiment/SKILL.md +236 -0
  298. package/skills/skills-codex/semantic-scholar/SKILL.md +219 -0
  299. package/skills/skills-codex/serverless-modal/SKILL.md +335 -0
  300. package/skills/skills-codex/shared-references/acceptance-gate.md +336 -0
  301. package/skills/skills-codex/shared-references/assurance-contract.md +139 -0
  302. package/skills/skills-codex/shared-references/capture-antipatterns.md +84 -0
  303. package/skills/skills-codex/shared-references/citation-discipline.md +452 -0
  304. package/skills/skills-codex/shared-references/compute-env-contract.md +163 -0
  305. package/skills/skills-codex/shared-references/effort-contract.md +143 -0
  306. package/skills/skills-codex/shared-references/evidence-precheck.md +73 -0
  307. package/skills/skills-codex/shared-references/experiment-integrity.md +49 -0
  308. package/skills/skills-codex/shared-references/external-cadence.md +334 -0
  309. package/skills/skills-codex/shared-references/fan-out-pattern.md +375 -0
  310. package/skills/skills-codex/shared-references/injection-hygiene.md +132 -0
  311. package/skills/skills-codex/shared-references/integration-contract.md +372 -0
  312. package/skills/skills-codex/shared-references/output-composition.md +98 -0
  313. package/skills/skills-codex/shared-references/output-language.md +45 -0
  314. package/skills/skills-codex/shared-references/output-manifest.md +40 -0
  315. package/skills/skills-codex/shared-references/output-versioning.md +111 -0
  316. package/skills/skills-codex/shared-references/patent-format-cn.md +199 -0
  317. package/skills/skills-codex/shared-references/patent-format-ep.md +173 -0
  318. package/skills/skills-codex/shared-references/patent-format-us.md +161 -0
  319. package/skills/skills-codex/shared-references/patent-writing-principles.md +197 -0
  320. package/skills/skills-codex/shared-references/prior-art-databases.md +141 -0
  321. package/skills/skills-codex/shared-references/resumable-runs.md +125 -0
  322. package/skills/skills-codex/shared-references/review-scope-limits.md +81 -0
  323. package/skills/skills-codex/shared-references/review-tracing.md +144 -0
  324. package/skills/skills-codex/shared-references/reviewer-independence.md +66 -0
  325. package/skills/skills-codex/shared-references/reviewer-routing.md +128 -0
  326. package/skills/skills-codex/shared-references/skill-governance.md +119 -0
  327. package/skills/skills-codex/shared-references/taste-calibration.md +90 -0
  328. package/skills/skills-codex/shared-references/venue-checklists.md +73 -0
  329. package/skills/skills-codex/shared-references/wiki-helper-resolution.md +69 -0
  330. package/skills/skills-codex/shared-references/writing-principles.md +525 -0
  331. package/skills/skills-codex/slides-polish/SKILL.md +563 -0
  332. package/skills/skills-codex/specification-writing/SKILL.md +211 -0
  333. package/skills/skills-codex/system-profile/SKILL.md +103 -0
  334. package/skills/skills-codex/training-check/SKILL.md +83 -0
  335. package/skills/skills-codex/vast-gpu/SKILL.md +394 -0
  336. package/skills/skills-codex/web-debug-search/SKILL.md +334 -0
  337. package/skills/skills-codex/wiki-enrich/SKILL.md +255 -0
  338. package/skills/skills-codex/writing-systems-papers/SKILL.md +184 -0
  339. package/skills/skills-codex-claude-review/README.md +79 -0
  340. package/skills/skills-codex-claude-review/README_CN.md +78 -0
  341. package/skills/skills-codex-claude-review/auto-paper-improvement-loop/SKILL.md +581 -0
  342. package/skills/skills-codex-claude-review/auto-review-loop/SKILL.md +510 -0
  343. package/skills/skills-codex-claude-review/novelty-check/SKILL.md +102 -0
  344. package/skills/skills-codex-claude-review/paper-figure/SKILL.md +319 -0
  345. package/skills/skills-codex-claude-review/paper-plan/SKILL.md +287 -0
  346. package/skills/skills-codex-claude-review/paper-write/SKILL.md +420 -0
  347. package/skills/skills-codex-claude-review/research-refine/SKILL.md +732 -0
  348. package/skills/skills-codex-claude-review/research-review/SKILL.md +149 -0
  349. package/skills/skills-codex-gemini-review/README.md +176 -0
  350. package/skills/skills-codex-gemini-review/README_CN.md +175 -0
  351. package/skills/skills-codex-gemini-review/auto-paper-improvement-loop/SKILL.md +331 -0
  352. package/skills/skills-codex-gemini-review/auto-review-loop/SKILL.md +304 -0
  353. package/skills/skills-codex-gemini-review/grant-proposal/SKILL.md +630 -0
  354. package/skills/skills-codex-gemini-review/idea-creator/SKILL.md +263 -0
  355. package/skills/skills-codex-gemini-review/idea-discovery/SKILL.md +275 -0
  356. package/skills/skills-codex-gemini-review/idea-discovery-robot/SKILL.md +365 -0
  357. package/skills/skills-codex-gemini-review/novelty-check/SKILL.md +92 -0
  358. package/skills/skills-codex-gemini-review/paper-figure/SKILL.md +289 -0
  359. package/skills/skills-codex-gemini-review/paper-plan/SKILL.md +265 -0
  360. package/skills/skills-codex-gemini-review/paper-poster-html/SKILL.md +102 -0
  361. package/skills/skills-codex-gemini-review/paper-slides/SKILL.md +582 -0
  362. package/skills/skills-codex-gemini-review/paper-write/SKILL.md +346 -0
  363. package/skills/skills-codex-gemini-review/paper-writing/SKILL.md +312 -0
  364. package/skills/skills-codex-gemini-review/research-refine/SKILL.md +674 -0
  365. package/skills/skills-codex-gemini-review/research-review/SKILL.md +112 -0
  366. package/skills/slides-polish/SKILL.md +565 -0
  367. package/skills/specification-writing/SKILL.md +211 -0
  368. package/skills/system-profile/SKILL.md +103 -0
  369. package/skills/training-check/SKILL.md +132 -0
  370. package/skills/vast-gpu/SKILL.md +394 -0
  371. package/skills/web-debug-search/SKILL.md +334 -0
  372. package/skills/wiki-enrich/SKILL.md +257 -0
  373. package/skills/writing-systems-papers/SKILL.md +184 -0
  374. package/templates/CLAUDE_MD_TEMPLATE.md +29 -0
  375. package/templates/EXPERIMENT_LOG_TEMPLATE.md +47 -0
  376. package/templates/EXPERIMENT_PLAN_TEMPLATE.md +51 -0
  377. package/templates/EXPERIMENT_PLAN_TEMPLATE_CN.md +53 -0
  378. package/templates/FINDINGS_TEMPLATE.md +52 -0
  379. package/templates/IDEA_CANDIDATES_TEMPLATE.md +47 -0
  380. package/templates/IDEA_CANDIDATES_TEMPLATE_CN.md +47 -0
  381. package/templates/INVENTION_BRIEF_TEMPLATE.md +87 -0
  382. package/templates/MANIFEST_TEMPLATE.md +7 -0
  383. package/templates/NARRATIVE_REPORT_TEMPLATE.md +49 -0
  384. package/templates/PAPER_PLAN_TEMPLATE.md +47 -0
  385. package/templates/PATENT_CLAIMS_TEMPLATE.md +78 -0
  386. package/templates/PATENT_SPECIFICATION_TEMPLATE.md +68 -0
  387. package/templates/README.md +57 -0
  388. package/templates/RESEARCH_BRIEF_TEMPLATE.md +35 -0
  389. package/templates/RESEARCH_BRIEF_TEMPLATE_CN.md +41 -0
  390. package/templates/RESEARCH_CONTRACT_TEMPLATE.md +60 -0
  391. package/templates/claude-hooks/corpus_write_guard.json +16 -0
  392. package/templates/claude-hooks/corpus_write_guard.py +85 -0
  393. package/templates/claude-hooks/meta_logging.json +74 -0
  394. package/templates/gitignore-trace.txt +3 -0
  395. package/tools/__pycache__/check_skills_inventory.cpython-314.pyc +0 -0
  396. package/tools/arxiv_fetch.py +311 -0
  397. package/tools/capture_filter.py +126 -0
  398. package/tools/check_skills_inventory.py +273 -0
  399. package/tools/convert_skills_to_llm_chat.py +282 -0
  400. package/tools/copilot_native_evidence.py +818 -0
  401. package/tools/deepxiv_fetch.py +213 -0
  402. package/tools/evidence_check.py +212 -0
  403. package/tools/exa_search.py +425 -0
  404. package/tools/experiment_queue/README.md +118 -0
  405. package/tools/experiment_queue/build_manifest.py +44 -0
  406. package/tools/experiment_queue/queue_manager.py +44 -0
  407. package/tools/extract_paper_style.py +560 -0
  408. package/tools/figure_renderer.py +69 -0
  409. package/tools/forensics_gate.py +669 -0
  410. package/tools/generate_codex_claude_review_overrides.py +299 -0
  411. package/tools/idea_discovery_gate.py +256 -0
  412. package/tools/install_aris.ps1 +1372 -0
  413. package/tools/install_aris.sh +1370 -0
  414. package/tools/install_aris_codex.sh +1023 -0
  415. package/tools/install_aris_copilot.sh +1052 -0
  416. package/tools/iteration_log.py +143 -0
  417. package/tools/lint_skills_helpers.sh +84 -0
  418. package/tools/meta_opt/check_ready.sh +80 -0
  419. package/tools/meta_opt/log_event.sh +91 -0
  420. package/tools/meta_opt/trigger_eval.py +280 -0
  421. package/tools/meta_opt/trigger_evals.sample.json +28 -0
  422. package/tools/openalex_fetch.py +326 -0
  423. package/tools/overleaf_audit.sh +104 -0
  424. package/tools/overleaf_setup.sh +150 -0
  425. package/tools/paper_illustration_image2.py +62 -0
  426. package/tools/provenance.py +294 -0
  427. package/tools/research_wiki.py +1720 -0
  428. package/tools/review_gate.py +502 -0
  429. package/tools/run_state.py +399 -0
  430. package/tools/save_trace.sh +477 -0
  431. package/tools/semantic_scholar_fetch.py +438 -0
  432. package/tools/skill-groups.tsv +116 -0
  433. package/tools/skill_picker.py +238 -0
  434. package/tools/smart_update.ps1 +521 -0
  435. package/tools/smart_update.sh +591 -0
  436. package/tools/smart_update_codex.sh +419 -0
  437. package/tools/smart_update_copilot.sh +605 -0
  438. package/tools/threat_scan.py +222 -0
  439. package/tools/verify_paper_audits.sh +487 -0
  440. package/tools/verify_papers.py +613 -0
  441. package/tools/verify_wiki_coverage.sh +176 -0
  442. package/tools/watchdog.py +485 -0
@@ -0,0 +1,437 @@
1
+ ---
2
+ name: meta-optimize
3
+ description: "Analyze ARIS usage logs and propose optimizations to SKILL.md files, reviewer prompts, and workflow defaults. Outer-loop harness optimization inspired by Meta-Harness (Lee et al., 2026). Use when user says \"优化技能\", \"meta optimize\", \"improve skills\", \"分析使用记录\", or wants to optimize ARIS's own harness components based on accumulated experience."
4
+ argument-hint: "[target-skill-or-all]"
5
+ allowed-tools: Bash(*), Read, Grep, Glob, mcp__codex__codex, mcp__codex__codex-reply
6
+ ---
7
+
8
+ # Meta-Optimize: Outer-Loop Harness Optimization for ARIS
9
+
10
+ Analyze accumulated usage logs and propose optimizations for: **$ARGUMENTS**
11
+
12
+ ## Privilege boundary — this skill is a READ-ONLY PRODUCER
13
+
14
+ meta-optimize **proposes**; it does not **land**. The mutation of the skill corpus
15
+ is the exclusive job of a separate, human-invoked skill: [`/meta-apply`](../meta-apply/SKILL.md).
16
+ This split is structural, not advisory — it is why a missed instruction cannot let
17
+ this loop apply its own patch (the self-acquittal failure mode):
18
+
19
+ - **No `Write`/`Edit` tool.** This skill cannot edit a SKILL.md / shared-reference /
20
+ any corpus file with the frictionless mutators. Its only outputs are the REPORT and
21
+ staged patch files, written under `.aris/meta/` (a scratch area, never the corpus).
22
+ - **No apply step.** There is no in-skill "apply the patch" path (see Step 6). The
23
+ producer ends by *staging* approved patches for `/meta-apply`; a human must then
24
+ invoke `/meta-apply` to land them. That human action is the landing gate.
25
+ - **`Bash` writes to the corpus are filtered, not impossible — be honest about the
26
+ layers.** What IS fully closed: the *accidental / in-flow* self-acquittal — this skill
27
+ has no `Write`/`Edit` and no apply step, so an honest run cannot slip into editing the
28
+ corpus. Defense-in-depth: install the
29
+ [`corpus_write_guard`](../../templates/claude-hooks/corpus_write_guard.json) PreToolUse
30
+ hook (like `meta_logging.json`), which DENIES the common Bash shell-writes (`>`, `tee`,
31
+ `sed -i`, `cp`/`mv`, `touch`, `open(...,'w')`) to corpus paths. **This is a blacklist,
32
+ NOT a complete sandbox** — a *deliberately* obscured Bash write (`git apply`, `patch`,
33
+ `$var`/absolute paths, language file APIs) is not all caught. **Full structural
34
+ prevention requires either removing this skill's `Bash` or an FS sandbox** — over-built
35
+ for a not-yet-load-bearing producer, so deferred to when the gate carries real
36
+ auto-modification volume (a brick-3 trigger). The intended backstop against a deliberate
37
+ write is **detection, not prevention** — a corpus change with no valid/current
38
+ `provenance` stamp (content-hash mismatch) *would be* catchable in a pre-push integrity
39
+ check — but that verifier is **NOT yet built** (`provenance.py` has `content_hash` but no
40
+ integrity-check subcommand, and no pre-push hook runs one). So today the deliberate-write
41
+ case is neither prevented nor actively detected; track the integrity verifier as a
42
+ follow-up before this producer goes load-bearing. Its legitimate Bash writes go only to
43
+ `.aris/meta/`.
44
+
45
+ See [`shared-references/acceptance-gate.md`](../shared-references/acceptance-gate.md):
46
+ a loop can DRIVE (propose, review) same-model, but the ACQUITTAL that lands a change
47
+ must be cross-model (Step 4 jury) **and** the landing must be a separate human-gated
48
+ act (`/meta-apply`).
49
+
50
+ ## Context
51
+
52
+ ARIS is a **research harness** — a system of skills, bridges, workflows, and artifact contracts that wraps around LLMs to orchestrate research. This skill implements a prototype **outer loop** that observes how the harness is used and proposes improvements to the harness itself (not to the research artifacts it produces).
53
+
54
+ Inspired by Meta-Harness (Lee et al., 2026): the key insight is that harness design matters as much as model weights, and harness engineering can be partially automated by logging execution traces and using them to guide improvements.
55
+
56
+ ## What This Skill Optimizes (Harness Components)
57
+
58
+ | Component | Example | Optimizable? |
59
+ |-----------|---------|:---:|
60
+ | SKILL.md prompts | Reviewer instructions, quality gates, step descriptions | Yes |
61
+ | Default parameters | `difficulty: medium`, `MAX_ROUNDS: 4`, `threshold: 6/10` | Yes |
62
+ | Convergence rules | When to stop the review loop, retry counts | Yes |
63
+ | Workflow ordering | Skill chain sequence within a workflow | Yes |
64
+ | Artifact schemas | What fields go in EXPERIMENT_LOG.md, idea-stage/IDEA_REPORT.md | Cautious |
65
+ | MCP bridge config | Which reviewer model, routing rules | No (infra) |
66
+
67
+ **Not optimized**: The research artifacts themselves (papers, code, experiments). That's what the regular workflows do.
68
+
69
+ ## Prerequisites
70
+
71
+ 1. **Logging must be active.** Copy `templates/claude-hooks/meta_logging.json` into your project's `.claude/settings.json` (or merge the hooks section).
72
+ 2. **Sufficient data.** At least 5 complete workflow runs logged in `.aris/meta/events.jsonl`. The skill will check and warn if insufficient.
73
+
74
+ ## Workflow
75
+
76
+ ### Step 0: Check Data Availability
77
+
78
+ ```bash
79
+ EVENTS_FILE=".aris/meta/events.jsonl"
80
+ if [ ! -f "$EVENTS_FILE" ]; then
81
+ echo "ERROR: No event log found at $EVENTS_FILE"
82
+ echo "Enable logging first: copy templates/claude-hooks/meta_logging.json into .claude/settings.json"
83
+ exit 1
84
+ fi
85
+
86
+ EVENT_COUNT=$(wc -l < "$EVENTS_FILE")
87
+ SKILL_INVOCATIONS=$(grep -c '"skill_invoke"' "$EVENTS_FILE" || echo 0)
88
+ SESSIONS=$(grep -c '"session_start"' "$EVENTS_FILE" || echo 0)
89
+
90
+ echo "📊 Event log: $EVENT_COUNT events, $SKILL_INVOCATIONS skill invocations, $SESSIONS sessions"
91
+
92
+ if [ "$SKILL_INVOCATIONS" -lt 5 ]; then
93
+ echo "⚠️ Insufficient data (<5 skill invocations). Continue using ARIS normally and re-run later."
94
+ exit 0
95
+ fi
96
+
97
+ # Bottleneck succession: what did the LAST cycle say was the limiting stage?
98
+ BOTTLENECK_LOG=".aris/meta/bottleneck_log.jsonl"
99
+ if [ -f "$BOTTLENECK_LOG" ]; then
100
+ echo "🧭 Prior cycle's bottleneck: $(tail -1 "$BOTTLENECK_LOG")"
101
+ fi
102
+ ```
103
+
104
+ If a prior bottleneck entry exists, open the report (Step 5) by stating whether
105
+ that named bottleneck was **resolved** (and by which landed patches) and what it
106
+ has now **moved to** — bottleneck SUCCESSION, not just existence, is the signal
107
+ this ledger exists to carry.
108
+
109
+ ### Step 1: Analyze Usage Patterns
110
+
111
+ Read `.aris/meta/events.jsonl` and compute:
112
+
113
+ **Frequency analysis:**
114
+ - Which skills are invoked most often?
115
+ - Which slash commands do users type most?
116
+ - What parameter overrides are most common? (These suggest bad defaults.)
117
+
118
+ **Failure analysis:**
119
+ - Which tools fail most often? In which skills?
120
+ - What error patterns repeat? (OOM, import, compilation, timeout)
121
+ - How many auto-debug retries per workflow run?
122
+
123
+ **Convergence analysis (for auto-review-loop):**
124
+ - Average rounds to reach threshold
125
+ - Score trajectory shape (fast improvement? plateau? oscillation?)
126
+ - Which review round catches the most critical issues?
127
+ - Do users override difficulty mid-run?
128
+
129
+ **Human intervention analysis:**
130
+ - Where do users interrupt with manual prompts during workflows?
131
+ - What manual corrections do users make most? (These indicate skill gaps.)
132
+
133
+ **Model-delta analysis (harness diet):**
134
+ - Has the session model (`session_start` events' `model` field) or the pinned
135
+ reviewer model changed since a skill's SKILL.md was last touched?
136
+ (`git log -1 --format=%cs -- skills/<skill>/SKILL.md` vs the model-bump date.)
137
+ - A model bump is a **trigger to re-read, not evidence by itself**. For each
138
+ reasoning-scaffolding step or worked example in that SKILL.md, a deletion
139
+ proposal must cite TARGET-SPECIFIC evidence that the new model no longer
140
+ needs it: a capability-specific release note, or repeated observed behavior
141
+ in the event log (e.g. zero failures/interventions in the guarded step since
142
+ the bump). "The model got newer" alone never justifies a deletion.
143
+ - **Never deletion candidates**, regardless of model: privilege boundaries,
144
+ acceptance/review gates, corpus- and provenance-integrity rules, output
145
+ contracts, and safety checks. The diet targets model-compensation scaffolding
146
+ only — a capability the new model has natively is pure overhead (context
147
+ weight, drift surface, reading cost). A harness that only ever grows is a
148
+ harness nobody is re-reading.
149
+
150
+ **Trigger-rate analysis (optional, measured — not from the event log):**
151
+ - The event log shows which skills were USED, not which were WANTED-but-omitted
152
+ — the omission failure mode (Claude Code passing over the right skill when the
153
+ installed list is long) is invisible to it. `tools/meta_opt/trigger_eval.py`
154
+ measures it directly: `claude -p` probes with paraphrased-intent queries run
155
+ from a neutral cwd (so the realistic long installed corpus is loaded), scored
156
+ as trigger / confusion(→which skill) / miss.
157
+ - Run it when a specific skill is suspected of under- or mis-triggering, or as a
158
+ before/after check around a description edit:
159
+ `python3 tools/meta_opt/trigger_eval.py --eval-file tools/meta_opt/trigger_evals.sample.json --skills <name> --samples 2`
160
+ - The **confusion matrix is the signal**, not just the rate: a query that keeps
161
+ landing on a sibling skill means the two descriptions overlap on that intent —
162
+ the fix is disambiguation, not "make the description pushier".
163
+ - **Measure-only, evidence not verdict.** A low trigger rate is an INPUT to a
164
+ Step-2 proposal (which lands only via `/meta-apply`), never a self-applied
165
+ description rewrite. Trigger rate is model-dependent, so compare like with
166
+ like (record the probe model) and treat it as a proxy — it measures selection
167
+ under a query set, not the full long-list omission problem.
168
+
169
+ Present findings as a structured summary table.
170
+
171
+ ### Step 1.5: Name the Current Bottleneck
172
+
173
+ Synthesize the Step-1 analyses into **one sentence naming the single
174
+ most-limiting pipeline stage right now** — e.g. "planning", "verification
175
+ quality", "experiment execution reliability", "writing polish" — with the
176
+ supporting evidence. The bottleneck always moves: when coding stops being the
177
+ constraint, planning becomes it; when planning is solved, verification; when
178
+ verification is automated, taste. This step exists to make the CURRENT
179
+ constraint visible, so Step 2's ranked table reads as sub-fixes for one named
180
+ constraint instead of scattered tweaks.
181
+
182
+ Append the verdict to the append-only ledger `.aris/meta/bottleneck_log.jsonl`
183
+ (same never-mutate discipline as `.aris/runs/<run_id>.iterations.jsonl`):
184
+
185
+ ```bash
186
+ mkdir -p .aris/meta
187
+ # json.dumps, NOT hand-interpolated shell strings: bottleneck/evidence are
188
+ # natural language — a stray quote must not break the JSONL (or the shell).
189
+ python3 - <<'PY'
190
+ import json, datetime
191
+ entry = {
192
+ "ts": datetime.datetime.now().astimezone().isoformat(timespec="seconds"),
193
+ "cycle": 3,
194
+ "bottleneck": "verification quality",
195
+ "evidence": "review rounds plateau at 6/10 while tool failures are rare",
196
+ "top_patch_ids": ["P1", "P2"],
197
+ }
198
+ with open(".aris/meta/bottleneck_log.jsonl", "a", encoding="utf-8") as fh:
199
+ fh.write(json.dumps(entry, ensure_ascii=False) + "\n")
200
+ PY
201
+ ```
202
+
203
+ Never edit or delete prior lines — succession history is the point.
204
+
205
+ ### Step 2: Identify Optimization Targets
206
+
207
+ Based on Step 1, rank optimization opportunities by expected impact:
208
+
209
+ ```markdown
210
+ ## Optimization Opportunities (ranked)
211
+
212
+ | # | Target | Signal | Proposed Change | Expected Impact |
213
+ |---|--------|--------|-----------------|-----------------|
214
+ | 1 | auto-review-loop default threshold | Users override to 7/10 in 60% of runs | Change default from 6/10 to 7/10 | Fewer manual overrides |
215
+ | 2 | experiment-bridge retry count | 40% of runs hit max retries on OOM | Add OOM-specific recovery (reduce batch size) | Fewer failed experiments |
216
+ | 3 | paper-write de-AI patterns | Users manually fix "delve" in 80% of runs | Add "delve" to default watchword list | Fewer manual edits |
217
+ | 4 | experiment-bridge Phase-2 hand-holding steps | Model bump (session_start model changed); scaffold untouched since 2 generations ago; zero tool_failures in the steps it guards | **DELETE steps N–M — the new model does this unprompted** | Smaller harness, less drift surface |
218
+ ```
219
+
220
+ The Proposed-Change column is explicitly allowed to be a **deletion** — "DELETE
221
+ step N, new model does this for free" is a first-class optimization, ranked by
222
+ the same impact logic as additions.
223
+
224
+ If `$ARGUMENTS` specifies a target skill, focus analysis on that skill only.
225
+ If `$ARGUMENTS` is empty or "all", analyze all skills with sufficient data.
226
+
227
+ ### Step 3: Generate Patch Proposals
228
+
229
+ For each optimization target, generate a concrete diff:
230
+
231
+ ```diff
232
+ --- a/skills/auto-review-loop/SKILL.md
233
+ +++ b/skills/auto-review-loop/SKILL.md
234
+ @@ -15,7 +15,7 @@
235
+ ## Constants
236
+
237
+ -- **SCORE_THRESHOLD = 6** — Minimum review score to accept.
238
+ +- **SCORE_THRESHOLD = 7** — Minimum review score to accept. (Raised based on usage data: 60% of users overrode to 7+.)
239
+ ```
240
+
241
+ **Rules for patch generation:**
242
+ - One patch per optimization target
243
+ - Each patch must include a comment explaining WHY (with data from the log)
244
+ - Patches must be minimal — change only what the data supports
245
+ - Never change artifact schemas or MCP bridge config in v1
246
+ - Never change behavior that would break existing user workflows
247
+ - **Anti-self-poisoning screen** (see [`shared-references/capture-antipatterns.md`](../shared-references/capture-antipatterns.md)):
248
+ run a proposed patch's rationale through `tools/capture_filter.py` (resolve via
249
+ the canonical chain). NEVER propose a change that encodes a **negative
250
+ tool-capability claim** ("codex can't…", "gemini is broken") or a **one-off /
251
+ transient failure** as a durable rule — those harden into self-cited refusals.
252
+ Encode the *fix / the flag needed / the workaround*, not "X can't do Y".
253
+
254
+ ### Step 4: Cross-Model Review of Patches (ADVISORY pre-screen)
255
+
256
+ > This review is **advisory** — it sharpens the Step-5 REPORT so the human can decide
257
+ > what to stage. It is **not** the landing verdict. The binding cross-model jury runs
258
+ > later, at landing, inside [`/meta-apply`](../meta-apply/SKILL.md), on the actual staged
259
+ > diff (a producer-relayed verdict would be forgeable). Record this result as
260
+ > `advisory_screen` only.
261
+
262
+ Send each patch to GPT-5.6-Sol xhigh for adversarial review:
263
+
264
+ ```
265
+ mcp__codex__codex:
266
+ model: gpt-5.6-sol
267
+ config: {"model_reasoning_effort": "xhigh"}
268
+ prompt: |
269
+ You are reviewing a proposed optimization to an ARIS SKILL.md file.
270
+
271
+ ## Original Skill (relevant section)
272
+ [paste original]
273
+
274
+ ## Proposed Patch
275
+ [paste diff]
276
+
277
+ ## Evidence from Usage Log
278
+ [paste summary stats]
279
+
280
+ Review this patch:
281
+ 1. Does the evidence support the change?
282
+ 2. Could this change hurt other use cases?
283
+ 3. Is the change minimal and safe?
284
+ 4. Score 1-10: should this be applied?
285
+
286
+ If score < 7, explain what additional evidence would be needed.
287
+
288
+ === SCOPE LIMITS (these bound what you PROPOSE, never what you look for) ===
289
+ Report anything that is actually wrong here — including a rare-looking case, if
290
+ this repo actually produces it. Then keep the fix in scope:
291
+ 1. This is a RESEARCH-WORKFLOW tool, not a security paper. Verification is
292
+ welcome; over-defense is not. Assume a cooperating operator on their own
293
+ machine — a malicious local user is NOT in the threat model.
294
+ 2. Do NOT propose SHA / hash / content-fingerprint / digest-binding schemes.
295
+ Reporting a real defect in hashing code that already exists is fine.
296
+ 3. NO speculative machinery: do not add feature flags, migration frameworks,
297
+ compat layers, wrappers, pins, or similar mechanisms unless evidence shows
298
+ a current repo defect they fix or an explicit existing invariant they must
299
+ preserve. "Load-bearing", "compatibility", and "not scaffolding" are labels,
300
+ not evidence. Point to the failing path/artifact or invariant, and check the
301
+ proposal's factual premises, such as whether a named package version exists.
302
+ 4. NO corner-case obsession: exotic encodings, symlink races, RTL text and
303
+ millisecond races are out of scope unless you can show the case arises here.
304
+ 5. Where a rubric or checklist is genuinely needed, do not over-mechanize
305
+ judgement. A clear sentence a human reads beats a scored table nobody
306
+ maintains.
307
+ Exception: code that runs remote commands, starts a network service, or installs
308
+ an MCP server runs on the user's machine with their credentials — trust-boundary
309
+ findings there are in scope and the default is strict.
310
+ Say plainly when something is correct. Do not manufacture findings.
311
+ ```
312
+
313
+ ### Step 5: Present Results
314
+
315
+ Output a structured report:
316
+
317
+ ```markdown
318
+ # ARIS Meta-Optimization Report
319
+
320
+ **Date**: [today]
321
+ **Data**: [N] events, [M] skill invocations, [K] sessions
322
+ **Target**: [skill name or "all"]
323
+
324
+ ## Current Bottleneck
325
+
326
+ **[one-phrase name]** — [one-line evidence]. Prior cycle's bottleneck: [name —
327
+ resolved by <patch ids> / unresolved / first recorded cycle]. (Ledger:
328
+ `.aris/meta/bottleneck_log.jsonl`)
329
+
330
+ ## Proposed Changes
331
+
332
+ ### Change 1: [title]
333
+ - **Target**: [skill/file:line]
334
+ - **Signal**: [what the data shows]
335
+ - **Patch**: [diff]
336
+ - **Reviewer Score**: [X/10]
337
+ - **Reviewer Notes**: [summary]
338
+ - **Status**: ✅ Recommended / ⚠️ Needs more data / ❌ Rejected
339
+
340
+ ### Change 2: ...
341
+
342
+ ## Changes NOT Made (insufficient evidence)
343
+ - [pattern observed but too few samples]
344
+
345
+ ## Recommendations
346
+ - [ ] Apply Change 1 (reviewer approved)
347
+ - [ ] Collect more data for Change 3 (need N more runs)
348
+ - [ ] Consider manual review of Change 2
349
+
350
+ ## Next Steps
351
+ This skill only **proposes**. To land changes: tell me which to stage, then run
352
+ `/meta-apply` (a separate, human-invoked applier that re-checks the cross-model
353
+ verdict before mutating anything). meta-optimize never applies.
354
+ ```
355
+
356
+ ### Step 6: Stage approved patches for `/meta-apply` (NO in-skill apply)
357
+
358
+ This skill does **not** apply anything. After the user has read the Step-5 REPORT and
359
+ indicated which changes to land, **stage** them for the privileged applier:
360
+
361
+ 1. For each approved change `N`, write its unified diff to
362
+ `.aris/meta/pending/<NN>_<skill>.diff` and append a row to
363
+ `.aris/meta/pending/manifest.jsonl`:
364
+ `{patch: "<NN>_<skill>.diff", target: "<corpus path>", author_model: "<executor>",
365
+ advisory_screen: "pass|kill", advisory_reason: "<one line>"}`.
366
+ The `advisory_screen` (your Step-4 codex pre-review) is **advisory only** — it helps
367
+ the human read the REPORT. It is **NOT** the landing verdict and `/meta-apply` does not
368
+ trust it: a producer-written verdict would be forgeable. The binding cross-model jury
369
+ runs **at landing, inside `/meta-apply`,** on the actual staged diff.
370
+ 2. Tell the user: *"Staged M patches. Run `/meta-apply` to judge & land them."*
371
+
372
+ The backup → **fresh jury-at-landing** → apply → **provenance stamp** → log all happen
373
+ inside [`/meta-apply`](../meta-apply/SKILL.md). meta-optimize never touches the corpus and
374
+ never produces the acquittal.
375
+
376
+ **Never apply in this skill. Landing is `/meta-apply` + a fresh jury + a human, always.**
377
+
378
+ ## Key Rules
379
+
380
+ - **Log-driven, not speculative.** Every proposed change must cite specific data from the event log. No "I think this would be better."
381
+ - **Minimal patches.** Change one thing at a time. Don't rewrite entire skills — the one sanctioned large edit is a scaffolding **deletion** backed by TARGET-SPECIFIC model-delta evidence (a capability-specific release note, or repeated post-bump event-log behavior showing the scaffold is unused — the model name changing is a trigger to look, never sufficient evidence). Privilege boundaries, acceptance gates, corpus/provenance rules, output contracts, and safety checks are never deletion candidates. Deletions go through the same review + approval gates as everything else.
382
+ - **Reviewer-gated.** Every patch goes through cross-model review before recommendation.
383
+ - **Reversible.** Always back up before applying. Always log what changed.
384
+ - **User-approved.** Never auto-apply. Present, explain, let the user decide.
385
+ - **Honest about uncertainty.** If the data is insufficient, say so. Don't optimize on noise.
386
+ - **Portable.** Optimizations should improve the skill for all users, not just one user's style. If a change seems user-specific, flag it.
387
+
388
+ ## Event Schema Reference
389
+
390
+ The log at `.aris/meta/events.jsonl` contains JSONL records with these shapes:
391
+
392
+ ```jsonl
393
+ {"ts":"...","session":"...","event":"skill_invoke","skill":"auto-review-loop","args":"difficulty: hard"}
394
+ {"ts":"...","session":"...","event":"PostToolUse","tool":"Bash","input_summary":"pdflatex main.tex"}
395
+ {"ts":"...","session":"...","event":"codex_call","tool":"mcp__codex__codex","input_summary":"review..."}
396
+ {"ts":"...","session":"...","event":"tool_failure","tool":"Bash","input_summary":"python train.py"}
397
+ {"ts":"...","session":"...","event":"slash_command","command":"/auto-review-loop","args":""}
398
+ {"ts":"...","session":"...","event":"user_prompt","prompt_preview":"change difficulty to hard"}
399
+ {"ts":"...","session":"...","event":"session_start","source":"startup","model":"claude-opus-4-6"}
400
+ {"ts":"...","session":"...","event":"session_end"}
401
+ ```
402
+
403
+ ## Triggering
404
+
405
+ This skill is NOT part of the standard W1→W1.5→W2→W3→W4 pipeline. It is a **maintenance workflow** with three trigger mechanisms:
406
+
407
+ 1. **Passive logging** (always on): Claude Code hooks record events to `.aris/meta/events.jsonl` automatically during normal usage. Zero user effort.
408
+
409
+ 2. **Automatic readiness check** (SessionEnd hook): When a Claude Code session ends, `check_ready.sh` counts skill invocations since the last `/meta-optimize` run. If ≥5 new invocations have accumulated, it prints a reminder:
410
+ ```
411
+ 📊 ARIS has logged 8 skill runs since last optimization. Run /meta-optimize to check for improvement opportunities.
412
+ ```
413
+ It ALSO fires — regardless of invocation count — when the session model has
414
+ changed since the last optimize (compared against `.aris/meta/.last_optimize_model`):
415
+ ```
416
+ 🔁 Model changed since last optimization (claude-opus-4-6 → claude-opus-4-8). Run /meta-optimize — a model bump makes existing scaffolding a deletion candidate (harness diet).
417
+ ```
418
+ Both are **suggestions only** — they do not auto-run optimization.
419
+
420
+ 3. **Manual trigger**: User runs `/meta-optimize` when they see the reminder or whenever they want.
421
+
422
+ **After each `/meta-optimize` run**, the skill writes the current timestamp to `.aris/meta/.last_optimize` and the current session model (latest `session_start` event's `model` field) to `.aris/meta/.last_optimize_model`, so the readiness check can detect both new usage and model bumps.
423
+
424
+ ## Acknowledgements
425
+
426
+ Inspired by [Meta-Harness](https://arxiv.org/abs/2603.28052) (Lee et al., 2026) — end-to-end optimization of model harnesses via filesystem-based experience access and agentic code search.
427
+
428
+ ## Output Protocols
429
+
430
+ > Follow these shared protocols for all output files:
431
+ > - **[Output Versioning Protocol](../shared-references/output-versioning.md)** — write timestamped file first, then copy to fixed name
432
+ > - **[Output Manifest Protocol](../shared-references/output-manifest.md)** — log every output to MANIFEST.md
433
+ > - **[Output Language Protocol](../shared-references/output-language.md)** — respect the project's language setting
434
+
435
+ ## Review Tracing
436
+
437
+ After each `mcp__codex__codex` or `mcp__codex__codex-reply` reviewer call, save the trace following `shared-references/review-tracing.md` (Policy C — forensic; never silently skip). Use `save_trace.sh` (resolved per the chain in `shared-references/integration-contract.md` §2) or write files directly to `.aris/traces/<skill>/<date>_run<NN>/`. Respect the `--- trace:` parameter (default: `full`).
@@ -0,0 +1,140 @@
1
+ ---
2
+ name: monitor-experiment
3
+ description: Monitor running experiments, check progress, collect results. Use when user says "check results", "is it done", "monitor", or wants experiment output.
4
+ argument-hint: "[server-alias or screen-name]"
5
+ allowed-tools: Bash(ssh *), Bash(echo *), Read, Write, Edit
6
+ ---
7
+
8
+ # Monitor Experiment Results
9
+
10
+ > ⏱ **External cadence is appropriate here.** This skill waits on an external
11
+ > fact (job completion / progress), so it is a natural `/loop` / `CronCreate`
12
+ > surface: the wake reads status and self-judges only **machine-checkable**
13
+ > completion (exit code, file exists, epoch logged) — never quality. This is
14
+ > the additive external-wait shape in
15
+ > [`shared-references/external-cadence.md`](../shared-references/external-cadence.md).
16
+ > If a scheduled wait here ends in a verdict step (e.g. then audit results),
17
+ > run that verdict **once** after the wait clears — not re-entered per tick.
18
+
19
+ Monitor: $ARGUMENTS
20
+
21
+ ## Workflow
22
+
23
+ ### Step 1: Check What's Running
24
+
25
+ **SSH server:**
26
+ ```bash
27
+ ssh <server> "screen -ls"
28
+ ```
29
+
30
+ **Vast.ai instance** (read `ssh_host`, `ssh_port` from `vast-instances.json`):
31
+ ```bash
32
+ ssh -p <PORT> root@<HOST> "screen -ls"
33
+ ```
34
+
35
+ Also check vast.ai instance status:
36
+ ```bash
37
+ vastai show instances
38
+ ```
39
+
40
+ **Modal** (when `gpu: modal` in CLAUDE.md):
41
+ ```bash
42
+ modal app list # List running/recent apps
43
+ modal app logs <app> # Stream logs from a running app
44
+ ```
45
+ Modal apps auto-terminate when done — if it's not in the list, it already finished. Check results via `modal volume ls <volume>` or local output.
46
+
47
+ ### Step 2: Collect Output from Each Screen
48
+ For each screen session, capture the last N lines:
49
+ ```bash
50
+ ssh <server> "screen -S <name> -X hardcopy /tmp/screen_<name>.txt && tail -50 /tmp/screen_<name>.txt"
51
+ ```
52
+
53
+ If hardcopy fails, check for log files or tee output.
54
+
55
+ ### Step 3: Check for JSON Result Files
56
+ ```bash
57
+ ssh <server> "ls -lt <results_dir>/*.json 2>/dev/null | head -20"
58
+ ```
59
+
60
+ If JSON results exist, fetch and parse them:
61
+ ```bash
62
+ ssh <server> "cat <results_dir>/<latest>.json"
63
+ ```
64
+
65
+ ### Step 3.5: Pull W&B Metrics (when `wandb: true` in CLAUDE.md)
66
+
67
+ **Skip this step entirely if `wandb` is not set or is `false` in CLAUDE.md.**
68
+
69
+ Pull training curves and metrics from Weights & Biases via Python API:
70
+
71
+ ```bash
72
+ # List recent runs in the project
73
+ ssh <server> "python3 -c \"
74
+ import wandb
75
+ api = wandb.Api()
76
+ runs = api.runs('<entity>/<project>', per_page=10)
77
+ for r in runs:
78
+ print(f'{r.id} {r.state} {r.name} {r.summary.get(\"eval/loss\", \"N/A\")}')
79
+ \""
80
+
81
+ # Pull specific metrics from a run (last 50 steps)
82
+ ssh <server> "python3 -c \"
83
+ import wandb, json
84
+ api = wandb.Api()
85
+ run = api.run('<entity>/<project>/<run_id>')
86
+ history = list(run.scan_history(keys=['train/loss', 'eval/loss', 'eval/ppl', 'train/lr'], page_size=50))
87
+ print(json.dumps(history[-10:], indent=2))
88
+ \""
89
+
90
+ # Pull run summary (final metrics)
91
+ ssh <server> "python3 -c \"
92
+ import wandb, json
93
+ api = wandb.Api()
94
+ run = api.run('<entity>/<project>/<run_id>')
95
+ print(json.dumps(dict(run.summary), indent=2, default=str))
96
+ \""
97
+ ```
98
+
99
+ **What to extract:**
100
+ - **Training loss curve** — is it converging? diverging? plateauing?
101
+ - **Eval metrics** — loss, PPL, accuracy at latest checkpoint
102
+ - **Learning rate** — is the schedule behaving as expected?
103
+ - **GPU memory** — any OOM risk?
104
+ - **Run status** — running / finished / crashed?
105
+
106
+ **W&B dashboard link** (include in summary for user):
107
+ ```
108
+ https://wandb.ai/<entity>/<project>/runs/<run_id>
109
+ ```
110
+
111
+ > This gives the auto-review-loop richer signal than just screen output — training dynamics, loss curves, and metric trends over time.
112
+
113
+ ### Step 4: Summarize Results
114
+
115
+ Present results in a comparison table:
116
+ ```
117
+ | Experiment | Metric | Delta vs Baseline | Status |
118
+ |-----------|--------|-------------------|--------|
119
+ | Baseline | X.XX | — | done |
120
+ | Method A | X.XX | +Y.Y | done |
121
+ ```
122
+
123
+ ### Step 5: Interpret
124
+ - Compare against known baselines
125
+ - Flag unexpected results (negative delta, NaN, divergence)
126
+ - Suggest next steps based on findings
127
+
128
+ ### Step 6: Feishu Notification (if configured)
129
+
130
+ After results are collected, check `~/.claude/feishu.json`:
131
+ - Send `experiment_done` notification: results summary table, delta vs baseline
132
+ - If config absent or mode `"off"`: skip entirely (no-op)
133
+
134
+ ## Key Rules
135
+ - Always show raw numbers before interpretation
136
+ - Compare against the correct baseline (same config)
137
+ - Note if experiments are still running (check progress bars, iteration counts)
138
+ - If results look wrong, check training logs for errors before concluding
139
+ - **Vast.ai cost awareness**: When monitoring vast.ai instances, report the running cost (hours * $/hr from `vast-instances.json`). If all experiments on an instance are done, remind the user to run `/vast-gpu destroy <instance_id>` to stop billing
140
+ - **Modal cost awareness**: Modal auto-scales to zero — no idle billing. When reporting results from Modal runs, note the actual execution time and estimated cost (time * $/hr from the GPU tier used). No cleanup action needed