dsh-aris-panel 0.1.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (442) hide show
  1. package/LICENSE +21 -0
  2. package/README.md +98 -0
  3. package/README_CN.md +87 -0
  4. package/dsh/checkout.patch.yml +38 -0
  5. package/dsh/client.js +634 -0
  6. package/dsh/cordis.patch.yml +44 -0
  7. package/dsh/index.mjs +76 -0
  8. package/dsh/run-status.mjs +182 -0
  9. package/dsh/scope-limits.mjs +50 -0
  10. package/dsh/workbench.mjs +291 -0
  11. package/mcp-servers/claude-review/README.md +93 -0
  12. package/mcp-servers/claude-review/run_with_claude_aws.sh +49 -0
  13. package/mcp-servers/claude-review/server.py +718 -0
  14. package/mcp-servers/codex-image2/README.md +65 -0
  15. package/mcp-servers/codex-image2/server.py +893 -0
  16. package/mcp-servers/feishu-bridge/requirements.txt +1 -0
  17. package/mcp-servers/feishu-bridge/server.py +240 -0
  18. package/mcp-servers/gemini-review/README.md +171 -0
  19. package/mcp-servers/gemini-review/server.py +1856 -0
  20. package/mcp-servers/llm-chat/requirements.txt +1 -0
  21. package/mcp-servers/llm-chat/server.py +664 -0
  22. package/mcp-servers/manual-review/README.md +133 -0
  23. package/mcp-servers/manual-review/server.py +910 -0
  24. package/mcp-servers/manual-review/ui.html +279 -0
  25. package/mcp-servers/minimax-chat/requirements.txt +1 -0
  26. package/mcp-servers/minimax-chat/server.py +381 -0
  27. package/package.json +51 -0
  28. package/skills/ablation-planner/SKILL.md +123 -0
  29. package/skills/alphaxiv/SKILL.md +196 -0
  30. package/skills/analyze-results/SKILL.md +46 -0
  31. package/skills/arxiv/SKILL.md +248 -0
  32. package/skills/auto-paper-improvement-loop/SKILL.md +651 -0
  33. package/skills/auto-review-loop/SKILL.md +1137 -0
  34. package/skills/auto-review-loop-llm/SKILL.md +259 -0
  35. package/skills/auto-review-loop-minimax/SKILL.md +302 -0
  36. package/skills/citation-audit/SKILL.md +502 -0
  37. package/skills/claims-drafting/SKILL.md +227 -0
  38. package/skills/comm-lit-review/SKILL.md +297 -0
  39. package/skills/deepxiv/SKILL.md +263 -0
  40. package/skills/dse-loop/SKILL.md +296 -0
  41. package/skills/embodiment-description/SKILL.md +129 -0
  42. package/skills/exa-search/SKILL.md +205 -0
  43. package/skills/experiment-audit/SKILL.md +311 -0
  44. package/skills/experiment-bridge/SKILL.md +376 -0
  45. package/skills/experiment-plan/SKILL.md +249 -0
  46. package/skills/experiment-queue/SKILL.md +431 -0
  47. package/skills/experiment-queue/scripts/build_manifest.py +142 -0
  48. package/skills/experiment-queue/scripts/queue_manager.py +433 -0
  49. package/skills/feishu-notify/SKILL.md +156 -0
  50. package/skills/figure-description/SKILL.md +138 -0
  51. package/skills/figure-spec/SKILL.md +262 -0
  52. package/skills/figure-spec/scripts/figure_renderer.py +799 -0
  53. package/skills/formula-derivation/SKILL.md +280 -0
  54. package/skills/gemini-search/SKILL.md +231 -0
  55. package/skills/grant-proposal/SKILL.md +698 -0
  56. package/skills/idea-creator/SKILL.md +542 -0
  57. package/skills/idea-discovery/SKILL.md +521 -0
  58. package/skills/idea-discovery-robot/SKILL.md +363 -0
  59. package/skills/integrity-forensics/SKILL.md +284 -0
  60. package/skills/interview-cheatsheet/SKILL.md +245 -0
  61. package/skills/invention-structuring/SKILL.md +188 -0
  62. package/skills/jurisdiction-format/SKILL.md +192 -0
  63. package/skills/kill-argument/SKILL.md +437 -0
  64. package/skills/mermaid-diagram/SKILL.md +419 -0
  65. package/skills/meta-apply/SKILL.md +141 -0
  66. package/skills/meta-optimize/SKILL.md +437 -0
  67. package/skills/monitor-experiment/SKILL.md +140 -0
  68. package/skills/novelty-check/SKILL.md +101 -0
  69. package/skills/openalex/SKILL.md +237 -0
  70. package/skills/overleaf-sync/SKILL.md +220 -0
  71. package/skills/paper-claim-audit/SKILL.md +348 -0
  72. package/skills/paper-compile/SKILL.md +266 -0
  73. package/skills/paper-figure/SKILL.md +312 -0
  74. package/skills/paper-illustration/SKILL.md +736 -0
  75. package/skills/paper-illustration-image2/SKILL.md +391 -0
  76. package/skills/paper-illustration-image2/scripts/paper_illustration_image2.py +255 -0
  77. package/skills/paper-plan/SKILL.md +386 -0
  78. package/skills/paper-poster/SKILL.md +19 -0
  79. package/skills/paper-poster-html/DESIGN_FINAL.md +176 -0
  80. package/skills/paper-poster-html/IMPLEMENTATION_CONVENTIONS.md +161 -0
  81. package/skills/paper-poster-html/LICENSES/posterly-MIT.txt +21 -0
  82. package/skills/paper-poster-html/NOTICE.md +57 -0
  83. package/skills/paper-poster-html/SKILL.md +323 -0
  84. package/skills/paper-poster-html/scripts/_posterly/__init__.py +0 -0
  85. package/skills/paper-poster-html/scripts/_posterly/canvas.py +200 -0
  86. package/skills/paper-poster-html/scripts/_posterly/measure.py +588 -0
  87. package/skills/paper-poster-html/scripts/_posterly/polish.py +498 -0
  88. package/skills/paper-poster-html/scripts/_posterly/preflight.py +489 -0
  89. package/skills/paper-poster-html/scripts/_posterly/render.py +215 -0
  90. package/skills/paper-poster-html/scripts/_posterly/textutil.py +16 -0
  91. package/skills/paper-poster-html/scripts/_posterly/verify_final.py +171 -0
  92. package/skills/paper-poster-html/scripts/asset_check.py +897 -0
  93. package/skills/paper-poster-html/scripts/extract_pdf_figures.py +666 -0
  94. package/skills/paper-poster-html/scripts/poster_check.py +251 -0
  95. package/skills/paper-poster-html/scripts/preprocess_figures.py +238 -0
  96. package/skills/paper-poster-html/scripts/render_preview.py +217 -0
  97. package/skills/paper-poster-html/scripts/run_gates.py +556 -0
  98. package/skills/paper-poster-html/scripts/style_check.py +1324 -0
  99. package/skills/paper-poster-html/templates/COMPONENTS.md +462 -0
  100. package/skills/paper-poster-html/templates/README.md +170 -0
  101. package/skills/paper-poster-html/templates/landscape_4col.html +1032 -0
  102. package/skills/paper-poster-html/templates/landscape_hero.html +1046 -0
  103. package/skills/paper-poster-html/templates/portrait_2col.html +947 -0
  104. package/skills/paper-poster-html/templates/tokens/acl.json +9 -0
  105. package/skills/paper-poster-html/templates/tokens/cvpr.json +9 -0
  106. package/skills/paper-poster-html/templates/tokens/generic.json +9 -0
  107. package/skills/paper-poster-html/templates/tokens/iclr.json +9 -0
  108. package/skills/paper-poster-html/templates/tokens/icml.json +9 -0
  109. package/skills/paper-poster-html/templates/tokens/neurips.json +9 -0
  110. package/skills/paper-slides/SKILL.md +635 -0
  111. package/skills/paper-talk/SKILL.md +381 -0
  112. package/skills/paper-write/SKILL.md +604 -0
  113. package/skills/paper-write/templates/IEEEtran.bst +2409 -0
  114. package/skills/paper-write/templates/IEEEtran.cls +6347 -0
  115. package/skills/paper-write/templates/iclr2026.tex +84 -0
  116. package/skills/paper-write/templates/icml2025.tex +87 -0
  117. package/skills/paper-write/templates/ieee_conference.tex +89 -0
  118. package/skills/paper-write/templates/ieee_journal.tex +93 -0
  119. package/skills/paper-write/templates/math_commands.tex +48 -0
  120. package/skills/paper-write/templates/neurips2025.tex +80 -0
  121. package/skills/paper-writing/SKILL.md +916 -0
  122. package/skills/patent-novelty-check/SKILL.md +153 -0
  123. package/skills/patent-pipeline/SKILL.md +344 -0
  124. package/skills/patent-review/SKILL.md +203 -0
  125. package/skills/pixel-art/SKILL.md +137 -0
  126. package/skills/prior-art-search/SKILL.md +146 -0
  127. package/skills/proof-checker/SKILL.md +866 -0
  128. package/skills/proof-orchestrator/NOTICE.md +24 -0
  129. package/skills/proof-orchestrator/SKILL.md +254 -0
  130. package/skills/proof-orchestrator/references/audit-output-contract.md +126 -0
  131. package/skills/proof-orchestrator/references/deepseek-routing.md +74 -0
  132. package/skills/proof-orchestrator/references/dispatch-prompts.md +227 -0
  133. package/skills/proof-orchestrator/references/notation-audit.md +135 -0
  134. package/skills/proof-orchestrator/references/proof-audit-rubric.md +70 -0
  135. package/skills/proof-orchestrator/references/stress-tests.md +38 -0
  136. package/skills/proof-writer/SKILL.md +223 -0
  137. package/skills/qzcli/SKILL.md +324 -0
  138. package/skills/rebuttal/SKILL.md +376 -0
  139. package/skills/render-html/SKILL.md +316 -0
  140. package/skills/render-html/scripts/render_html.py +1006 -0
  141. package/skills/render-html/scripts/templates/academic.html +703 -0
  142. package/skills/render-html/scripts/templates/dashboard.html +333 -0
  143. package/skills/research-lit/SKILL.md +756 -0
  144. package/skills/research-pipeline/SKILL.md +384 -0
  145. package/skills/research-refine/SKILL.md +770 -0
  146. package/skills/research-refine-pipeline/SKILL.md +186 -0
  147. package/skills/research-review/SKILL.md +198 -0
  148. package/skills/research-wiki/SKILL.md +461 -0
  149. package/skills/resubmit-pipeline/SKILL.md +447 -0
  150. package/skills/result-to-claim/SKILL.md +311 -0
  151. package/skills/run-experiment/SKILL.md +313 -0
  152. package/skills/semantic-scholar/SKILL.md +236 -0
  153. package/skills/serverless-modal/SKILL.md +335 -0
  154. package/skills/shared-references/acceptance-gate.md +324 -0
  155. package/skills/shared-references/assurance-contract.md +248 -0
  156. package/skills/shared-references/capture-antipatterns.md +78 -0
  157. package/skills/shared-references/citation-discipline.md +583 -0
  158. package/skills/shared-references/compute-env-contract.md +163 -0
  159. package/skills/shared-references/effort-contract.md +183 -0
  160. package/skills/shared-references/evidence-precheck.md +65 -0
  161. package/skills/shared-references/experiment-integrity.md +49 -0
  162. package/skills/shared-references/external-cadence.md +326 -0
  163. package/skills/shared-references/fan-out-pattern.md +366 -0
  164. package/skills/shared-references/injection-hygiene.md +127 -0
  165. package/skills/shared-references/integration-contract.md +461 -0
  166. package/skills/shared-references/output-composition.md +93 -0
  167. package/skills/shared-references/output-language.md +45 -0
  168. package/skills/shared-references/output-manifest.md +49 -0
  169. package/skills/shared-references/output-versioning.md +111 -0
  170. package/skills/shared-references/patent-format-cn.md +199 -0
  171. package/skills/shared-references/patent-format-ep.md +173 -0
  172. package/skills/shared-references/patent-format-us.md +161 -0
  173. package/skills/shared-references/patent-writing-principles.md +197 -0
  174. package/skills/shared-references/prior-art-databases.md +141 -0
  175. package/skills/shared-references/resumable-runs.md +109 -0
  176. package/skills/shared-references/review-scope-limits.md +81 -0
  177. package/skills/shared-references/review-tracing.md +391 -0
  178. package/skills/shared-references/reviewer-independence.md +79 -0
  179. package/skills/shared-references/reviewer-routing.md +852 -0
  180. package/skills/shared-references/skill-governance.md +104 -0
  181. package/skills/shared-references/taste-calibration.md +85 -0
  182. package/skills/shared-references/venue-checklists.md +114 -0
  183. package/skills/shared-references/wiki-helper-resolution.md +134 -0
  184. package/skills/shared-references/writing-principles.md +525 -0
  185. package/skills/skills-codex/README.md +102 -0
  186. package/skills/skills-codex/README_CN.md +100 -0
  187. package/skills/skills-codex/ablation-planner/SKILL.md +126 -0
  188. package/skills/skills-codex/alphaxiv/SKILL.md +186 -0
  189. package/skills/skills-codex/analyze-results/SKILL.md +45 -0
  190. package/skills/skills-codex/arxiv/SKILL.md +210 -0
  191. package/skills/skills-codex/auto-paper-improvement-loop/SKILL.md +574 -0
  192. package/skills/skills-codex/auto-review-loop/SKILL.md +500 -0
  193. package/skills/skills-codex/auto-review-loop-llm/SKILL.md +247 -0
  194. package/skills/skills-codex/auto-review-loop-minimax/SKILL.md +290 -0
  195. package/skills/skills-codex/citation-audit/SKILL.md +504 -0
  196. package/skills/skills-codex/claims-drafting/SKILL.md +239 -0
  197. package/skills/skills-codex/comm-lit-review/SKILL.md +299 -0
  198. package/skills/skills-codex/comm-lit-review/references/domain-taxonomy.md +57 -0
  199. package/skills/skills-codex/comm-lit-review/references/output-template.md +37 -0
  200. package/skills/skills-codex/comm-lit-review/references/source-policy.md +99 -0
  201. package/skills/skills-codex/comm-lit-review/references/venue-tiering.md +112 -0
  202. package/skills/skills-codex/deepxiv/SKILL.md +142 -0
  203. package/skills/skills-codex/dse-loop/SKILL.md +285 -0
  204. package/skills/skills-codex/embodiment-description/SKILL.md +129 -0
  205. package/skills/skills-codex/exa-search/SKILL.md +192 -0
  206. package/skills/skills-codex/experiment-audit/SKILL.md +286 -0
  207. package/skills/skills-codex/experiment-bridge/SKILL.md +356 -0
  208. package/skills/skills-codex/experiment-plan/SKILL.md +249 -0
  209. package/skills/skills-codex/experiment-queue/SKILL.md +401 -0
  210. package/skills/skills-codex/feishu-notify/SKILL.md +155 -0
  211. package/skills/skills-codex/figure-description/SKILL.md +138 -0
  212. package/skills/skills-codex/figure-spec/SKILL.md +252 -0
  213. package/skills/skills-codex/formula-derivation/SKILL.md +280 -0
  214. package/skills/skills-codex/gemini-search/SKILL.md +205 -0
  215. package/skills/skills-codex/grant-proposal/SKILL.md +626 -0
  216. package/skills/skills-codex/idea-creator/SKILL.md +405 -0
  217. package/skills/skills-codex/idea-discovery/SKILL.md +475 -0
  218. package/skills/skills-codex/idea-discovery-robot/SKILL.md +362 -0
  219. package/skills/skills-codex/integrity-forensics/SKILL.md +106 -0
  220. package/skills/skills-codex/interview-cheatsheet/SKILL.md +245 -0
  221. package/skills/skills-codex/invention-structuring/SKILL.md +188 -0
  222. package/skills/skills-codex/jurisdiction-format/SKILL.md +192 -0
  223. package/skills/skills-codex/kill-argument/SKILL.md +403 -0
  224. package/skills/skills-codex/mermaid-diagram/SKILL.md +379 -0
  225. package/skills/skills-codex/meta-apply/SKILL.md +154 -0
  226. package/skills/skills-codex/meta-optimize/SKILL.md +348 -0
  227. package/skills/skills-codex/monitor-experiment/SKILL.md +98 -0
  228. package/skills/skills-codex/novelty-check/SKILL.md +89 -0
  229. package/skills/skills-codex/openalex/SKILL.md +228 -0
  230. package/skills/skills-codex/overleaf-sync/SKILL.md +220 -0
  231. package/skills/skills-codex/paper-claim-audit/SKILL.md +350 -0
  232. package/skills/skills-codex/paper-compile/SKILL.md +253 -0
  233. package/skills/skills-codex/paper-figure/SKILL.md +311 -0
  234. package/skills/skills-codex/paper-illustration/SKILL.md +690 -0
  235. package/skills/skills-codex/paper-illustration-image2/SKILL.md +383 -0
  236. package/skills/skills-codex/paper-illustration-image2/scripts/paper_illustration_image2.py +255 -0
  237. package/skills/skills-codex/paper-plan/SKILL.md +278 -0
  238. package/skills/skills-codex/paper-poster/SKILL.md +19 -0
  239. package/skills/skills-codex/paper-poster-html/SKILL.md +377 -0
  240. package/skills/skills-codex/paper-slides/SKILL.md +571 -0
  241. package/skills/skills-codex/paper-talk/SKILL.md +381 -0
  242. package/skills/skills-codex/paper-write/SKILL.md +411 -0
  243. package/skills/skills-codex/paper-write/templates/IEEEtran.bst +2409 -0
  244. package/skills/skills-codex/paper-write/templates/IEEEtran.cls +6347 -0
  245. package/skills/skills-codex/paper-write/templates/aaai2026.bst +1493 -0
  246. package/skills/skills-codex/paper-write/templates/aaai2026.sty +315 -0
  247. package/skills/skills-codex/paper-write/templates/aaai2026.tex +952 -0
  248. package/skills/skills-codex/paper-write/templates/acl.sty +312 -0
  249. package/skills/skills-codex/paper-write/templates/acl2026.tex +377 -0
  250. package/skills/skills-codex/paper-write/templates/acl_natbib.bst +1940 -0
  251. package/skills/skills-codex/paper-write/templates/acm.bst +3081 -0
  252. package/skills/skills-codex/paper-write/templates/acm_mm2026.tex +204 -0
  253. package/skills/skills-codex/paper-write/templates/acmart.cls +3520 -0
  254. package/skills/skills-codex/paper-write/templates/cvpr.bst +1448 -0
  255. package/skills/skills-codex/paper-write/templates/cvpr.sty +508 -0
  256. package/skills/skills-codex/paper-write/templates/cvpr2026.tex +63 -0
  257. package/skills/skills-codex/paper-write/templates/iclr2026.tex +84 -0
  258. package/skills/skills-codex/paper-write/templates/iclr2026_conference.bst +1440 -0
  259. package/skills/skills-codex/paper-write/templates/iclr2026_conference.sty +246 -0
  260. package/skills/skills-codex/paper-write/templates/icml2026.sty +767 -0
  261. package/skills/skills-codex/paper-write/templates/icml2026.tex +662 -0
  262. package/skills/skills-codex/paper-write/templates/ieee_conference.tex +89 -0
  263. package/skills/skills-codex/paper-write/templates/ieee_journal.tex +93 -0
  264. package/skills/skills-codex/paper-write/templates/math_commands.tex +48 -0
  265. package/skills/skills-codex/paper-write/templates/neurips2026.tex +493 -0
  266. package/skills/skills-codex/paper-write/templates/neurips_2026.sty +437 -0
  267. package/skills/skills-codex/paper-writing/SKILL.md +731 -0
  268. package/skills/skills-codex/patent-novelty-check/SKILL.md +153 -0
  269. package/skills/skills-codex/patent-pipeline/SKILL.md +344 -0
  270. package/skills/skills-codex/patent-review/SKILL.md +202 -0
  271. package/skills/skills-codex/pixel-art/SKILL.md +139 -0
  272. package/skills/skills-codex/prior-art-search/SKILL.md +146 -0
  273. package/skills/skills-codex/proof-checker/SKILL.md +554 -0
  274. package/skills/skills-codex/proof-orchestrator/SKILL.md +260 -0
  275. package/skills/skills-codex/proof-orchestrator/references/audit-output-contract.md +126 -0
  276. package/skills/skills-codex/proof-orchestrator/references/deepseek-routing.md +76 -0
  277. package/skills/skills-codex/proof-orchestrator/references/dispatch-prompts.md +227 -0
  278. package/skills/skills-codex/proof-orchestrator/references/notation-audit.md +135 -0
  279. package/skills/skills-codex/proof-orchestrator/references/proof-audit-rubric.md +70 -0
  280. package/skills/skills-codex/proof-orchestrator/references/stress-tests.md +38 -0
  281. package/skills/skills-codex/proof-writer/SKILL.md +222 -0
  282. package/skills/skills-codex/qzcli/SKILL.md +324 -0
  283. package/skills/skills-codex/rebuttal/SKILL.md +305 -0
  284. package/skills/skills-codex/render-html/SKILL.md +305 -0
  285. package/skills/skills-codex/render-html/scripts/__pycache__/render_html.cpython-314.pyc +0 -0
  286. package/skills/skills-codex/render-html/scripts/render_html.py +909 -0
  287. package/skills/skills-codex/render-html/scripts/templates/academic.html +342 -0
  288. package/skills/skills-codex/render-html/scripts/templates/dashboard.html +333 -0
  289. package/skills/skills-codex/research-lit/SKILL.md +464 -0
  290. package/skills/skills-codex/research-pipeline/SKILL.md +340 -0
  291. package/skills/skills-codex/research-refine/SKILL.md +721 -0
  292. package/skills/skills-codex/research-refine-pipeline/SKILL.md +186 -0
  293. package/skills/skills-codex/research-review/SKILL.md +135 -0
  294. package/skills/skills-codex/research-wiki/SKILL.md +421 -0
  295. package/skills/skills-codex/resubmit-pipeline/SKILL.md +444 -0
  296. package/skills/skills-codex/result-to-claim/SKILL.md +246 -0
  297. package/skills/skills-codex/run-experiment/SKILL.md +236 -0
  298. package/skills/skills-codex/semantic-scholar/SKILL.md +219 -0
  299. package/skills/skills-codex/serverless-modal/SKILL.md +335 -0
  300. package/skills/skills-codex/shared-references/acceptance-gate.md +336 -0
  301. package/skills/skills-codex/shared-references/assurance-contract.md +139 -0
  302. package/skills/skills-codex/shared-references/capture-antipatterns.md +84 -0
  303. package/skills/skills-codex/shared-references/citation-discipline.md +452 -0
  304. package/skills/skills-codex/shared-references/compute-env-contract.md +163 -0
  305. package/skills/skills-codex/shared-references/effort-contract.md +143 -0
  306. package/skills/skills-codex/shared-references/evidence-precheck.md +73 -0
  307. package/skills/skills-codex/shared-references/experiment-integrity.md +49 -0
  308. package/skills/skills-codex/shared-references/external-cadence.md +334 -0
  309. package/skills/skills-codex/shared-references/fan-out-pattern.md +375 -0
  310. package/skills/skills-codex/shared-references/injection-hygiene.md +132 -0
  311. package/skills/skills-codex/shared-references/integration-contract.md +372 -0
  312. package/skills/skills-codex/shared-references/output-composition.md +98 -0
  313. package/skills/skills-codex/shared-references/output-language.md +45 -0
  314. package/skills/skills-codex/shared-references/output-manifest.md +40 -0
  315. package/skills/skills-codex/shared-references/output-versioning.md +111 -0
  316. package/skills/skills-codex/shared-references/patent-format-cn.md +199 -0
  317. package/skills/skills-codex/shared-references/patent-format-ep.md +173 -0
  318. package/skills/skills-codex/shared-references/patent-format-us.md +161 -0
  319. package/skills/skills-codex/shared-references/patent-writing-principles.md +197 -0
  320. package/skills/skills-codex/shared-references/prior-art-databases.md +141 -0
  321. package/skills/skills-codex/shared-references/resumable-runs.md +125 -0
  322. package/skills/skills-codex/shared-references/review-scope-limits.md +81 -0
  323. package/skills/skills-codex/shared-references/review-tracing.md +144 -0
  324. package/skills/skills-codex/shared-references/reviewer-independence.md +66 -0
  325. package/skills/skills-codex/shared-references/reviewer-routing.md +128 -0
  326. package/skills/skills-codex/shared-references/skill-governance.md +119 -0
  327. package/skills/skills-codex/shared-references/taste-calibration.md +90 -0
  328. package/skills/skills-codex/shared-references/venue-checklists.md +73 -0
  329. package/skills/skills-codex/shared-references/wiki-helper-resolution.md +69 -0
  330. package/skills/skills-codex/shared-references/writing-principles.md +525 -0
  331. package/skills/skills-codex/slides-polish/SKILL.md +563 -0
  332. package/skills/skills-codex/specification-writing/SKILL.md +211 -0
  333. package/skills/skills-codex/system-profile/SKILL.md +103 -0
  334. package/skills/skills-codex/training-check/SKILL.md +83 -0
  335. package/skills/skills-codex/vast-gpu/SKILL.md +394 -0
  336. package/skills/skills-codex/web-debug-search/SKILL.md +334 -0
  337. package/skills/skills-codex/wiki-enrich/SKILL.md +255 -0
  338. package/skills/skills-codex/writing-systems-papers/SKILL.md +184 -0
  339. package/skills/skills-codex-claude-review/README.md +79 -0
  340. package/skills/skills-codex-claude-review/README_CN.md +78 -0
  341. package/skills/skills-codex-claude-review/auto-paper-improvement-loop/SKILL.md +581 -0
  342. package/skills/skills-codex-claude-review/auto-review-loop/SKILL.md +510 -0
  343. package/skills/skills-codex-claude-review/novelty-check/SKILL.md +102 -0
  344. package/skills/skills-codex-claude-review/paper-figure/SKILL.md +319 -0
  345. package/skills/skills-codex-claude-review/paper-plan/SKILL.md +287 -0
  346. package/skills/skills-codex-claude-review/paper-write/SKILL.md +420 -0
  347. package/skills/skills-codex-claude-review/research-refine/SKILL.md +732 -0
  348. package/skills/skills-codex-claude-review/research-review/SKILL.md +149 -0
  349. package/skills/skills-codex-gemini-review/README.md +176 -0
  350. package/skills/skills-codex-gemini-review/README_CN.md +175 -0
  351. package/skills/skills-codex-gemini-review/auto-paper-improvement-loop/SKILL.md +331 -0
  352. package/skills/skills-codex-gemini-review/auto-review-loop/SKILL.md +304 -0
  353. package/skills/skills-codex-gemini-review/grant-proposal/SKILL.md +630 -0
  354. package/skills/skills-codex-gemini-review/idea-creator/SKILL.md +263 -0
  355. package/skills/skills-codex-gemini-review/idea-discovery/SKILL.md +275 -0
  356. package/skills/skills-codex-gemini-review/idea-discovery-robot/SKILL.md +365 -0
  357. package/skills/skills-codex-gemini-review/novelty-check/SKILL.md +92 -0
  358. package/skills/skills-codex-gemini-review/paper-figure/SKILL.md +289 -0
  359. package/skills/skills-codex-gemini-review/paper-plan/SKILL.md +265 -0
  360. package/skills/skills-codex-gemini-review/paper-poster-html/SKILL.md +102 -0
  361. package/skills/skills-codex-gemini-review/paper-slides/SKILL.md +582 -0
  362. package/skills/skills-codex-gemini-review/paper-write/SKILL.md +346 -0
  363. package/skills/skills-codex-gemini-review/paper-writing/SKILL.md +312 -0
  364. package/skills/skills-codex-gemini-review/research-refine/SKILL.md +674 -0
  365. package/skills/skills-codex-gemini-review/research-review/SKILL.md +112 -0
  366. package/skills/slides-polish/SKILL.md +565 -0
  367. package/skills/specification-writing/SKILL.md +211 -0
  368. package/skills/system-profile/SKILL.md +103 -0
  369. package/skills/training-check/SKILL.md +132 -0
  370. package/skills/vast-gpu/SKILL.md +394 -0
  371. package/skills/web-debug-search/SKILL.md +334 -0
  372. package/skills/wiki-enrich/SKILL.md +257 -0
  373. package/skills/writing-systems-papers/SKILL.md +184 -0
  374. package/templates/CLAUDE_MD_TEMPLATE.md +29 -0
  375. package/templates/EXPERIMENT_LOG_TEMPLATE.md +47 -0
  376. package/templates/EXPERIMENT_PLAN_TEMPLATE.md +51 -0
  377. package/templates/EXPERIMENT_PLAN_TEMPLATE_CN.md +53 -0
  378. package/templates/FINDINGS_TEMPLATE.md +52 -0
  379. package/templates/IDEA_CANDIDATES_TEMPLATE.md +47 -0
  380. package/templates/IDEA_CANDIDATES_TEMPLATE_CN.md +47 -0
  381. package/templates/INVENTION_BRIEF_TEMPLATE.md +87 -0
  382. package/templates/MANIFEST_TEMPLATE.md +7 -0
  383. package/templates/NARRATIVE_REPORT_TEMPLATE.md +49 -0
  384. package/templates/PAPER_PLAN_TEMPLATE.md +47 -0
  385. package/templates/PATENT_CLAIMS_TEMPLATE.md +78 -0
  386. package/templates/PATENT_SPECIFICATION_TEMPLATE.md +68 -0
  387. package/templates/README.md +57 -0
  388. package/templates/RESEARCH_BRIEF_TEMPLATE.md +35 -0
  389. package/templates/RESEARCH_BRIEF_TEMPLATE_CN.md +41 -0
  390. package/templates/RESEARCH_CONTRACT_TEMPLATE.md +60 -0
  391. package/templates/claude-hooks/corpus_write_guard.json +16 -0
  392. package/templates/claude-hooks/corpus_write_guard.py +85 -0
  393. package/templates/claude-hooks/meta_logging.json +74 -0
  394. package/templates/gitignore-trace.txt +3 -0
  395. package/tools/__pycache__/check_skills_inventory.cpython-314.pyc +0 -0
  396. package/tools/arxiv_fetch.py +311 -0
  397. package/tools/capture_filter.py +126 -0
  398. package/tools/check_skills_inventory.py +273 -0
  399. package/tools/convert_skills_to_llm_chat.py +282 -0
  400. package/tools/copilot_native_evidence.py +818 -0
  401. package/tools/deepxiv_fetch.py +213 -0
  402. package/tools/evidence_check.py +212 -0
  403. package/tools/exa_search.py +425 -0
  404. package/tools/experiment_queue/README.md +118 -0
  405. package/tools/experiment_queue/build_manifest.py +44 -0
  406. package/tools/experiment_queue/queue_manager.py +44 -0
  407. package/tools/extract_paper_style.py +560 -0
  408. package/tools/figure_renderer.py +69 -0
  409. package/tools/forensics_gate.py +669 -0
  410. package/tools/generate_codex_claude_review_overrides.py +299 -0
  411. package/tools/idea_discovery_gate.py +256 -0
  412. package/tools/install_aris.ps1 +1372 -0
  413. package/tools/install_aris.sh +1370 -0
  414. package/tools/install_aris_codex.sh +1023 -0
  415. package/tools/install_aris_copilot.sh +1052 -0
  416. package/tools/iteration_log.py +143 -0
  417. package/tools/lint_skills_helpers.sh +84 -0
  418. package/tools/meta_opt/check_ready.sh +80 -0
  419. package/tools/meta_opt/log_event.sh +91 -0
  420. package/tools/meta_opt/trigger_eval.py +280 -0
  421. package/tools/meta_opt/trigger_evals.sample.json +28 -0
  422. package/tools/openalex_fetch.py +326 -0
  423. package/tools/overleaf_audit.sh +104 -0
  424. package/tools/overleaf_setup.sh +150 -0
  425. package/tools/paper_illustration_image2.py +62 -0
  426. package/tools/provenance.py +294 -0
  427. package/tools/research_wiki.py +1720 -0
  428. package/tools/review_gate.py +502 -0
  429. package/tools/run_state.py +399 -0
  430. package/tools/save_trace.sh +477 -0
  431. package/tools/semantic_scholar_fetch.py +438 -0
  432. package/tools/skill-groups.tsv +116 -0
  433. package/tools/skill_picker.py +238 -0
  434. package/tools/smart_update.ps1 +521 -0
  435. package/tools/smart_update.sh +591 -0
  436. package/tools/smart_update_codex.sh +419 -0
  437. package/tools/smart_update_copilot.sh +605 -0
  438. package/tools/threat_scan.py +222 -0
  439. package/tools/verify_paper_audits.sh +487 -0
  440. package/tools/verify_papers.py +613 -0
  441. package/tools/verify_wiki_coverage.sh +176 -0
  442. package/tools/watchdog.py +485 -0
@@ -0,0 +1,163 @@
1
+ # Compute Environment Contract
2
+
3
+ > One declarative spec for what an environment IS; per-provider knowledge for
4
+ > how it gets built HERE; a content-hash ledger so "did the env change?" has a
5
+ > mechanical answer; and a three-tier validation ladder whose top tier is a
6
+ > fresh agent following the skill's own doc verbatim. Adapted from Anthropic's
7
+ > Claude Science `compute-env-setup` skill (Apache-2.0); de-coupled from its
8
+ > proprietary `host.*` runtime — everything here runs on plain bash + SSH +
9
+ > subagents.
10
+
11
+ Every ARIS compute skill (`/run-experiment`, `/experiment-queue`,
12
+ `/serverless-modal`, `/vast-gpu`, `/qzcli`) needs the same three things for a
13
+ job: a software stack (exact versions, often with load-bearing install order),
14
+ possibly large weights placed where the tool looks, and a resource shape. What
15
+ varies per provider is only HOW those materialize. Without a shared contract,
16
+ each skill re-encodes provider quirks and every "environment is ready" claim is
17
+ vibes. The classic failure this prevents: agent says "env ready", the overnight
18
+ run dies at `import flash_attn`, 8 GPUs idle until morning.
19
+
20
+ ## 1. Provider shapes — recognize, don't choose
21
+
22
+ You are rarely choosing a shape; you are recognizing which one this provider
23
+ already is. The shape determines what "build", "register", and "resolve" mean.
24
+
25
+ | Shape | ARIS examples | Build = | Env name resolves to |
26
+ |---|---|---|---|
27
+ | **Direct SSH host** (conda/venv) | personal GPU boxes, lab servers | YOU are the renderer: `conda create -n <name> python=<X>`, then run `pip_phases` in order | the conda env name itself (`conda run -n <name> …`) |
28
+ | **Scheduler cluster** (Slurm/PBS; Qizhi-like platforms) | `/qzcli` targets | `module load` or a container image built OFF-cluster and pulled (compute nodes often have **no internet** — pre-stage everything) | scheduler directives + container path in shared scratch (mind purge windows) |
29
+ | **Managed API** (serverless) | `/serverless-modal`, Vast.ai templates | the provider's image definition (Modal `Image`, Vast template) — render the same spec into it | the provider's opaque image ref, recorded in the ledger |
30
+
31
+ ## 2. The declarative spec (write WHAT once; render per provider)
32
+
33
+ ```yaml
34
+ # env-spec: one dict per environment, portable across shapes
35
+ base: "cuda12.8 + python3.10" # FROM-image / conda create versions
36
+ system_pkgs: [git, tmux] # apt in a container; conda-forge subset on no-root hosts
37
+ pip_phases: # ORDERED list of lists — each inner list = ONE pip call
38
+ - [torch==2.8.0] # phase 1 first, so later packages
39
+ - [flash-attn --no-build-isolation] # can't drag torch to a wrong wheel
40
+ - [transformers, peft, accelerate]
41
+ env: {HF_ENDPOINT: "...", OMP_NUM_THREADS: "<tier.cpus>"}
42
+ run_commands: [] # escape-hatch shell (RUN / %post / plain SSH)
43
+ weight_dirs: {chai: {path: /scratch/weights/chai, source: "tool's own loader", gated: false}}
44
+ smoke: # probes that run INSIDE the env on every shape
45
+ import_names: [torch, flash_attn]
46
+ gpu_tests: # each = {cmd, expect}; expect is the witness regex
47
+ - cmd: "python -c 'import torch;torch.manual_seed(0);x=torch.randn(8,8,device=\"cuda\");print(\"WITNESS\", (x@x).shape, torch.cuda.get_device_name())'"
48
+ expect: "^WITNESS torch.Size"
49
+ cli_checks: [nvidia-smi]
50
+ ```
51
+
52
+ - **`pip_phases` ordering IS the fix** for every "package A drags B to the
53
+ wrong version" problem: each phase is its own pip invocation, and pip leaves
54
+ an already-satisfied requirement alone unless asked to upgrade. Pin the
55
+ fought-over package in an EARLIER phase than the fighter.
56
+ - A clean spec renders unchanged through every renderer. If you find yourself
57
+ adding a field only one backend understands, that field belongs in the
58
+ provider's ledger entry, not the spec.
59
+ - **Weights**: small (<~500 MB) and read by every job → bake into the env at
60
+ build time. Large with a cache env var → persistent scratch + point the var
61
+ there. Populate with the **tool's own loader** (hand-curled layouts miss
62
+ marker files), then verify from the tool's perspective: run the real
63
+ entrypoint once against the staged dir and `du -sh` every subdir — 0 B means
64
+ a swallowed download error.
65
+
66
+ ## 3. The environment ledger (content-hash = mechanical staleness)
67
+
68
+ Per provider, keep an append-friendly `.aris/compute/<provider>.md` (or the
69
+ project's existing server-notes file). One block per env, keyed by a content
70
+ hash of the spec, computed over an EXACT canonical form so two agents can
71
+ never hash the same spec differently: parse the spec file, re-serialize as
72
+ JSON with sorted keys and no whitespace, sha256, first 8 hex chars —
73
+
74
+ ```bash
75
+ # spec stored as YAML (env-spec.yaml); requires PyYAML. If PyYAML is absent,
76
+ # store the spec as JSON instead and drop the yaml import — same pipeline.
77
+ python3 -c 'import sys,json,hashlib,yaml; \
78
+ s=json.dumps(yaml.safe_load(open(sys.argv[1])),sort_keys=True,separators=(",",":")); \
79
+ print(hashlib.sha256(s.encode()).hexdigest()[:8])' env-spec.yaml
80
+ ```
81
+
82
+ Key order, comments, indentation, and trailing whitespace in the source file
83
+ do NOT affect the hash — only the parsed content does:
84
+
85
+ ```
86
+ ### env: dllm@a3f9c2e1
87
+ how: conda env "dllm" on <host> # or: modal image ref / .sif path + partition
88
+ tier: {cpus: 8, mem_gib: 64, gpus: 1}
89
+ weights: HF_HOME=/scratch/hf (24 GB; purge-window 30d)
90
+ validated: 2026-07-02 (witness + agent-follows-doc clean)
91
+ gotcha: <any diagnosis-table row hit on THIS provider>
92
+ ```
93
+
94
+ Spec changed → hash changes → **cache miss**: the ledger entry no longer
95
+ matches and the env must be rebuilt (or a new block added). Spec unchanged →
96
+ warm-reuse without rebuilding or re-validating tier 1–2. This turns "I think
97
+ the env is the same as last week" into a string comparison. Note `.aris/` is
98
+ gitignored by convention — the ledger is **project-local and uncommitted** by
99
+ default (like `.aris/traces/`). If you want committed, git-blameable history,
100
+ keep the ledger blocks in the project's tracked server-notes file instead;
101
+ the block format is the contract, not the path.
102
+
103
+ ## 4. Validation — three tiers; the gap between them is where debugging lives
104
+
105
+ 1. **Import works** — `python -c "import <pkg>"` exits 0. Necessary, cheap,
106
+ catches almost nothing interesting.
107
+ 2. **Kernel-dispatch witness** — a tiny SEEDED forward pass that prints a
108
+ sentinel line (output shape + device name + non-emptiness). Catches "torch
109
+ sees the GPU but the kernel was compiled for an older SM", "the compiled
110
+ extension's `.so` isn't on the loader path", "inference writes to a
111
+ read-only cache". Keep the witness command in the spec's `smoke.gpu_tests`
112
+ with an `expect:` regex so the SAME probe runs on every backend. Cheap —
113
+ run on every build.
114
+ 3. **Agent-follows-doc** — the validation that actually matters and the one
115
+ that's easy to skip. Spawn a FRESH subagent that gets ONLY: the compute
116
+ skill's doc, the provider's ledger entry, and the documented invocation.
117
+ It must run the invocation **verbatim** — no improvisation, no fixing —
118
+ and report every point where the doc's claim and reality diverge. This is
119
+ where you find the doc says `--ligand` but the flag is
120
+ `--ligand_description`, or the weights path exists but lacks the completion
121
+ marker the tool checks. The author agent cannot self-certify its own doc
122
+ (it walks through on hidden knowledge the doc never wrote down — same
123
+ principle as `acceptance-gate.md`: the writer never acquits its own
124
+ artifact); the fresh agent's stuck-point IS the doc's lie. Expensive —
125
+ reserve for the two moments doc and env can drift: **after any env rebuild
126
+ or doc edit, and before declaring an env ready**.
127
+
128
+ ## 5. Diagnosis table (symptom → layer → fix)
129
+
130
+ When a documented invocation fails, don't patch reflexively — ask which LAYER
131
+ is wrong: spec, build, weights, resolution, or doc. Grep-able rows (container
132
+ rows apply only to container shapes):
133
+
134
+ | Symptom | Layer | Fix |
135
+ |---|---|---|
136
+ | `no kernel image is available for execution` | build/spec | torch compiled for older SM than this GPU — record `sm_range` in the ledger and route jobs; rebuild only if no compatible hardware |
137
+ | `ModuleNotFoundError` for a package not in the spec | spec | a `--no-deps` install skipped a runtime dep — read the package's `pyproject.toml` and add an explicit phase |
138
+ | Wrong torch/numpy version after install | spec | a later package's pin won — add a `force-reinstall --no-deps` snap-back phase after it |
139
+ | `ImportError: libfoo.so: cannot open shared object` | build | compiled `.so` not on loader path — `find` it, add its dir to `LD_LIBRARY_PATH` |
140
+ | Tool re-downloads despite populated weights | weights | `du -sh $CACHE_VAR` first: 0 B = swallowed error; non-zero = tool checks a marker file, stage that too |
141
+ | `OSError: Read-only file system` under cache var | weights (container) | tool writes locks next to weights on an RO mount — symlink blobs into writable `/tmp` cache |
142
+ | 80-way thread storm on a 4-CPU allocation | exec | `os.cpu_count()` returns the HOST's cores — export `OMP/MKL/OPENBLAS_NUM_THREADS=<tier.cpus>` on every backend |
143
+ | First job slow, every later job equally slow | build | expensive precompute runs at job time in a non-persistent workdir — run it once at build time |
144
+ | Job COMPLETED but output dir empty | exec | the wrapper writing the completion marker never ran — often `#!/bin/bash` on a runtime that only ships `/bin/sh` |
145
+
146
+ Hit a row on a specific provider → append symptom + fix to that provider's
147
+ ledger `gotcha:` line, so the next agent doesn't rediscover it.
148
+
149
+ ## How compute skills use this
150
+
151
+ - **Before building**: read the provider's ledger. The env — or a near-match
152
+ to extend — may already exist; an unchanged hash means skip the rebuild.
153
+ - **When building**: write the spec first (§2), render it for the shape (§1),
154
+ run tier-1/2 validation (§4), append the ledger block (§3).
155
+ - **Before declaring ready** (and after any rebuild/doc edit): run the
156
+ agent-follows-doc pass (§4.3).
157
+ - **On failure**: diagnosis table (§5) before patching; record provider-true
158
+ gotchas in the ledger.
159
+
160
+ Attribution: the spec/ledger/three-tier-validation design is adapted from
161
+ Anthropic's Claude Science `compute-env-setup` skill (Apache-2.0, re-hosted by
162
+ HughYau/AcademicForge); this document ports it off the proprietary `host.*`
163
+ runtime onto plain bash + SSH + ARIS subagents.
@@ -0,0 +1,183 @@
1
+ # Effort Contract
2
+
3
+ ## Overview
4
+
5
+ Every ARIS skill accepts an optional `effort` parameter that controls how much work the system does. This affects breadth, depth, iterations, and coverage — but **never** the quality of cross-model review.
6
+
7
+ > Design stance: the unattended *procedure* (gather → reason → act → verify → repeat) is the engineered artifact, not any single prompt — after Karpathy's "write the loop, not the prompt" (LOOPS.md, *Field Notes on Agents That Run for Days*).
8
+
9
+ ```
10
+ /any-skill "args" — effort: lite | balanced | max | beast
11
+ ```
12
+
13
+ Default: `balanced` (current behavior, zero change for existing users).
14
+
15
+ ## Hard Invariants (NEVER changed by effort)
16
+
17
+ | Setting | Value | Why |
18
+ |---------|-------|-----|
19
+ | Codex reasoning_effort | **≥ xhigh** (deep-audit skills run `ultra` — tier table in `reviewer-routing.md`) | Reviewer quality is non-negotiable. `effort` never moves the reviewer tier in either direction — and ARIS `— effort: max` is NOT Codex `model_reasoning_effort: max` (different axes: pipeline workload vs reviewer reasoning depth) |
20
+ | DBLP/CrossRef citations | **on** | Citation integrity is non-negotiable |
21
+ | Reviewer independence | **on** | Cross-model protocol is non-negotiable |
22
+ | Experiment integrity | **on** | Fraud prevention is non-negotiable |
23
+ | Sanity check | **on** | Safety is non-negotiable |
24
+ | **Mandatory audit emission** | **always** | At `assurance: submission`, every mandatory audit emits a verdict (PASS/WARN/FAIL/NOT_APPLICABLE/BLOCKED/ERROR). Silent skip is forbidden. See `assurance-contract.md`. |
25
+ | AUTO_PROCEED | **user decides** | Orthogonal to effort |
26
+ | difficulty | **user decides** | Orthogonal to effort |
27
+ | `assurance` | **derived from `effort`** (see Assurance Axis below) | Audit strictness is a separate axis from depth |
28
+
29
+ ## Four Levels
30
+
31
+ ### `lite` (~0.4x tokens)
32
+ For budget-constrained users or quick explorations. Minimum viable depth.
33
+ Implies `assurance: draft` (see below).
34
+
35
+ ### `balanced` (1x tokens) — DEFAULT
36
+ Current ARIS behavior. What existing users get today. No breakage.
37
+ Implies `assurance: draft`.
38
+
39
+ ### `max` (~2.5x tokens)
40
+ Go deeper than defaults. More papers, more ideas, more rounds, more detail.
41
+ Implies `assurance: submission` — mandatory audits are load-bearing.
42
+
43
+ ### `beast` (~5-8x tokens)
44
+ No budget limit. Every knob to maximum. For top-venue submission sprints.
45
+ Implies `assurance: submission` — mandatory audits are load-bearing and the
46
+ final report is tagged `submission-ready` only when the verifier agrees.
47
+
48
+ ## Assurance Axis (separate concern from `effort`)
49
+
50
+ Audit strictness lives on a second axis, `assurance`. Full contract:
51
+ **`shared-references/assurance-contract.md`**.
52
+
53
+ ```
54
+ — assurance: draft | submission
55
+ ```
56
+
57
+ Default mapping (if `assurance` not given explicitly):
58
+
59
+ | `effort` | implied `assurance` | Behavior |
60
+ |----------|---------------------|----------|
61
+ | `lite` | `draft` | Audits run only if content detector matches; silent skip allowed |
62
+ | `balanced` | `draft` | Same as lite — current behavior, zero breakage |
63
+ | `max` | `submission` | Every mandatory audit emits a verdict; verifier blocks Final Report on FAIL/BLOCKED/ERROR/STALE |
64
+ | `beast` | `submission` | Same as max + final report tagged `submission-ready` |
65
+
66
+ User can override independently:
67
+ - `— effort: balanced, assurance: submission` → normal depth, strict audits
68
+ - `— effort: beast, assurance: draft` → maximum depth, no audit gate (legal but discouraged for real submissions)
69
+
70
+ **Why split the axes?** Historically `effort: beast` did not enforce audits — phases like `/proof-checker`, `/paper-claim-audit`, `/citation-audit` were gated by content detectors that allowed silent skip. A user reported `effort: beast` produced a "draft-quality" paper with all three submission gates skipped. The split makes audit strictness independently verifiable and stops conflating "do more work" with "be more rigorous."
71
+
72
+ ## Per-Skill Profiles
73
+
74
+ ### Discovery & Planning
75
+
76
+ | Skill | Dimension | lite | balanced | max | beast |
77
+ |-------|-----------|------|----------|-----|-------|
78
+ | research-lit | papers found | 6-8 | 10-15 | 18-25 | 40-50 |
79
+ | research-lit | query variants | 2 | 5 | 8 | 15+ |
80
+ | research-lit | deep reads | 3 | 5-8 | 8 | 15+ |
81
+ | idea-creator | ideas generated | 4-6 | 8-12 | 12-16 | 20-30 |
82
+ | idea-creator | pilots | 1-2 | 2-3 | 3-4 | 5-6 |
83
+ | novelty-check | claims checked | 2-3 | 3-4 | 4-6 | all |
84
+ | novelty-check | closest works | top-3 | top-5 | top-8 | top-10+ |
85
+ | research-refine | max rounds | 3 | 5 | 7 | 10+ |
86
+ | research-refine | papers considered | 8 | 15 | 24 | 30+ |
87
+ | experiment-plan | core experiments | 3 | 5 | 7 | 10+ |
88
+ | experiment-plan | seeds | 1 | 3 | 5 | 5 |
89
+ | experiment-plan | baseline families | 2 | 3 | 4 | 5+ |
90
+
91
+ ### Execution
92
+
93
+ | Skill | Dimension | lite | balanced | max | beast |
94
+ |-------|-----------|------|----------|-----|-------|
95
+ | experiment-bridge | scope | sanity + main | main + basic ablation | + top ablation + robustness | full suite + cross-validation |
96
+ | run-experiment | launches | smoke + main | smoke + multi-seed | + dry run + manifest | full config + multi-GPU parallel |
97
+ | monitor-experiment | depth | latest log | log + JSON | + W&B + anomaly | real-time + auto-alert + trend |
98
+ | analyze-results | findings | 3 | 5 | 8 | full-dimensional + stat tests |
99
+ | ablation-planner | ablations | 2-3 | 4-5 | 6-8 | 10+ |
100
+
101
+ ### Review
102
+
103
+ | Skill | Dimension | lite | balanced | max | beast |
104
+ |-------|-----------|------|----------|-----|-------|
105
+ | auto-review-loop | max rounds | 2 | 3-4 | 6 | 8+ (until converged) |
106
+ | auto-review-loop | fixes per round | 1-2 | 3-4 | 4-6 | all actionable |
107
+ | research-review | passes | 1 | 1 + follow-up | 1 + 2 follow-ups | 2 independent + cross-compare |
108
+ | experiment-audit | depth | skip | basic 4 checks | full 6 checks | line-by-line + reproduce |
109
+
110
+ ### Writing & Rebuttal
111
+
112
+ | Skill | Dimension | lite | balanced | max | beast |
113
+ |-------|-----------|------|----------|-----|-------|
114
+ | paper-plan | outline reviews | 0 | 1 | 2 | 3 |
115
+ | paper-plan | citations/section | 2-3 | 4-5 | 5-8 | 8+ |
116
+ | paper-figure | caption reviews | 1 | 1 | 2 | 3 |
117
+ | paper-write | abstract variants | 1 | 1 | 2 | 3 |
118
+ | paper-write | related work depth | shallow | standard | deep | exhaustive |
119
+ | paper-compile | fix attempts | 2 | 3 | 4 | until zero warnings |
120
+ | auto-paper-improvement | rounds | 1 | 2 | 3 | 5 |
121
+ | paper-illustration | render iterations | 2 | 3 | 5 | 7 |
122
+ | rebuttal | draft rounds | 1 | 2 | 3 | 5 |
123
+ | rebuttal | stress tests | 0-1 | 1 | 2 | 3 |
124
+
125
+ ## How to Read Effort in a Skill
126
+
127
+ Add this to the Constants section of each skill:
128
+
129
+ ```markdown
130
+ ## Constants
131
+
132
+ - **EFFORT = `balanced`** — Work intensity. Options: `lite`, `balanced`, `max`, `beast`. Override: `— effort: max`
133
+ ```
134
+
135
+ Then adjust numeric constants based on effort level. Example:
136
+
137
+ ```
138
+ Parse $ARGUMENTS for `— effort:` directive.
139
+ If not specified, default to `balanced`.
140
+
141
+ Adjust constants:
142
+ if effort == "lite": MAX_PAPERS = 8, MAX_IDEAS = 6, MAX_ROUNDS = 2
143
+ if effort == "balanced": MAX_PAPERS = 15, MAX_IDEAS = 12, MAX_ROUNDS = 4
144
+ if effort == "max": MAX_PAPERS = 25, MAX_IDEAS = 16, MAX_ROUNDS = 6
145
+ if effort == "beast": MAX_PAPERS = 50, MAX_IDEAS = 30, MAX_ROUNDS = 8
146
+ ```
147
+
148
+ ## Transparency
149
+
150
+ Every skill should print its effort configuration at the start:
151
+
152
+ ```
153
+ ⚡ [effort: max] papers=25, ideas=16, rounds=6 | Codex: tier per reviewer-routing.md (floor xhigh)
154
+ ```
155
+
156
+ ## Precedence
157
+
158
+ ```
159
+ explicit concrete knob (e.g., review_rounds: 2)
160
+ > explicit dimension override
161
+ > overall effort level
162
+ > skill default (balanced)
163
+ ```
164
+
165
+ Example: `— effort: beast, review_rounds: 3` → everything beast except review capped at 3.
166
+
167
+ For the `assurance` axis, precedence is independent:
168
+ ```
169
+ explicit `assurance: ...` directive
170
+ > effort-implied default (lite/balanced → draft, max/beast → submission)
171
+ > skill default (draft)
172
+ ```
173
+
174
+ Example: `— effort: balanced, assurance: submission` → normal depth knobs but submission-gate audit enforcement.
175
+
176
+ ## Token Cost Estimation
177
+
178
+ | Level | LLM tokens | GPU/wall-clock | Best for |
179
+ |-------|-----------|----------------|----------|
180
+ | lite | ~0.4x | ~0.5x | Quick exploration, budget users |
181
+ | balanced | 1x | 1x | Normal research workflow |
182
+ | max | ~2.5x | ~2x | Serious submission prep |
183
+ | beast | ~5-8x | ~3-4x | Top-venue final sprint |
@@ -0,0 +1,65 @@
1
+ # Evidence Pre-check
2
+
3
+ ARIS's claim audits (`/result-to-claim`, `/experiment-audit`, `/paper-claim-audit`)
4
+ spend a cross-model (codex/gemini) call to judge whether a claim is supported. The
5
+ cheapest, most common integrity failure is *hallucinated evidence*: a claim cites
6
+ a number + a source file, and the file doesn't exist or the number isn't in it.
7
+ You should not need a model call to catch that.
8
+
9
+ ## Two stages — and `verified` ≠ `correct`
10
+
11
+ ```
12
+ stage 1 tools/evidence_check.py deterministic · no model · fail-closed
13
+ catches HALLUCINATION — cited path missing, or cited value not in source.
14
+ stage 2 the cross-model jury codex/gemini
15
+ catches WRONG-BUT-REAL — the number IS in the file, but it doesn't
16
+ support the claim.
17
+ ```
18
+
19
+ A `verified` from stage 1 means **only that the cited evidence exists** — never
20
+ that the claim holds. Existence is execution-completeness (deterministic / safe
21
+ same-model); *support* is a quality verdict that stays with the cross-model jury
22
+ (`acceptance-gate.md`: the pre-check DRIVES a gate, it cannot ACQUIT a claim).
23
+ This is the reconcile pattern — a model's self-report cross-checked against
24
+ mechanical ground-truth (adapted from Hermes's curator reconcile-classifier),
25
+ made into a cheap pre-gate that catches hallucination *before* the jury runs and
26
+ spares the codex call on fabricated evidence.
27
+
28
+ ## Conservative by design
29
+
30
+ The pre-check favors **false-negative over false-positive**: when in doubt it
31
+ returns not-verified and lets the jury decide — it must never emit a false
32
+ `verified`. A pure number is matched by **numeric-token equality** (so `73.2`
33
+ matches `73.20` but `73` does NOT match `73.5`); a non-numeric value by
34
+ normalized substring.
35
+
36
+ ## Where ARIS uses it
37
+
38
+ - **`/result-to-claim`** Step 1.5: parse each claim's cited `(value, source)`,
39
+ run the batch pre-check, and **before the codex judgment** mark any claim whose
40
+ evidence is `path_missing` / `value_not_found` as **unsupported — evidence not
41
+ found**, and pass the per-claim pre-check status into the codex prompt so the
42
+ jury sees which claims have verified vs hallucinated evidence.
43
+ - **To extend:** `/experiment-audit` (the "phantom results" check is exactly
44
+ this) and `/paper-claim-audit` (every reported number → its result file).
45
+
46
+ ## API / CLI
47
+
48
+ ```
49
+ from evidence_check import check_claim, check_batch
50
+ check_claim(value, source, root=".") # -> {status: verified|path_missing|value_not_found, ...}
51
+ check_batch([{value, source, id?}, ...], root) # -> {results:[...], summary:{status: n}}
52
+ ```
53
+ ```
54
+ python3 tools/evidence_check.py <root> --value 73.2 --source results/eval.json # exit 0 verified
55
+ python3 tools/evidence_check.py <root> --batch claims.json # exit 1 if any claim hallucinated
56
+ ```
57
+
58
+ ## Cross-references
59
+ - `acceptance-gate.md` — the pre-check is the deterministic DRIVE; the jury is the
60
+ ACQUIT. `verified` is existence (execution-completeness), not correctness.
61
+ - `reviewer-independence.md` — the jury still reads the artifacts itself; the
62
+ pre-check only flags which claims have evidence to read, never pre-digests the
63
+ verdict.
64
+ - `experiment-integrity.md` — fabricated/phantom results are exactly what stage 1
65
+ catches deterministically before stage 2.
@@ -0,0 +1,49 @@
1
+ # Experiment Integrity Protocol
2
+
3
+ ## Core Principle
4
+
5
+ **The model that writes experiment code must NOT be the model that judges experiment integrity.** This is the same principle as reviewer-independence, applied to experiments.
6
+
7
+ ## Prohibited Patterns
8
+
9
+ ### 1. Fake Ground Truth
10
+ - ❌ Creating synthetic "reference" from model outputs and comparing against it
11
+ - ❌ Using baseline model outputs as ground truth
12
+ - ❌ Generating pseudo-GT that is structurally similar to predictions
13
+ - ✅ Using dataset-provided ground truth
14
+ - ✅ Using official evaluation scripts when available
15
+ - ✅ Proxy evaluation is allowed IF explicitly labeled as `synthetic_proxy`
16
+
17
+ ### 2. Score Normalization Fraud
18
+ - ❌ Dividing metrics by max/min of model's own output to get 0.99+
19
+ - ❌ Rescaling scores to hide poor performance
20
+ - ✅ Standard normalization (e.g., min-max across ALL methods including baselines)
21
+ - ✅ Reporting raw and normalized scores side by side
22
+
23
+ ### 3. Phantom Results
24
+ - ❌ Claiming results from files that don't exist
25
+ - ❌ Referencing metrics from functions that are never called
26
+ - ❌ Reporting TRACKER status as DONE when it's still TODO
27
+ - ✅ Every claimed number must trace to an actual output file
28
+
29
+ ### 4. Insufficient Scope
30
+ - ❌ Reporting 2-scene pilot as "comprehensive evaluation"
31
+ - ❌ Using words like "robust", "extensive", "across settings" for tiny experiments
32
+ - ✅ Honestly label scope: "pilot (N=2)", "preliminary", "limited evaluation"
33
+ - ✅ State exact scope: N scenes, N seeds, N configurations
34
+
35
+ ## Evaluation Types (must be declared)
36
+
37
+ | Type | Label | What it means | Claim ceiling |
38
+ |------|-------|---------------|---------------|
39
+ | Real GT | `real_gt` | Dataset-provided ground truth | Full performance claims |
40
+ | Synthetic proxy | `synthetic_proxy` | Model-generated reference | "Proxy consistency" only |
41
+ | Self-supervised | `self_supervised_proxy` | No GT by design | Relative improvement only |
42
+ | Simulation | `simulation_only` | Simulated environment | "In simulation" qualifier |
43
+ | Human eval | `human_eval` | Human judges | Subject to inter-rater stats |
44
+
45
+ ## Who Checks
46
+
47
+ The **reviewer model** (different family from executor) performs integrity checks via `/experiment-audit`. The executor collects file paths; the reviewer reads code and results directly.
48
+
49
+ **Never let the executor judge its own experiment integrity.**