dsh-aris-panel 0.1.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (442) hide show
  1. package/LICENSE +21 -0
  2. package/README.md +98 -0
  3. package/README_CN.md +87 -0
  4. package/dsh/checkout.patch.yml +38 -0
  5. package/dsh/client.js +634 -0
  6. package/dsh/cordis.patch.yml +44 -0
  7. package/dsh/index.mjs +76 -0
  8. package/dsh/run-status.mjs +182 -0
  9. package/dsh/scope-limits.mjs +50 -0
  10. package/dsh/workbench.mjs +291 -0
  11. package/mcp-servers/claude-review/README.md +93 -0
  12. package/mcp-servers/claude-review/run_with_claude_aws.sh +49 -0
  13. package/mcp-servers/claude-review/server.py +718 -0
  14. package/mcp-servers/codex-image2/README.md +65 -0
  15. package/mcp-servers/codex-image2/server.py +893 -0
  16. package/mcp-servers/feishu-bridge/requirements.txt +1 -0
  17. package/mcp-servers/feishu-bridge/server.py +240 -0
  18. package/mcp-servers/gemini-review/README.md +171 -0
  19. package/mcp-servers/gemini-review/server.py +1856 -0
  20. package/mcp-servers/llm-chat/requirements.txt +1 -0
  21. package/mcp-servers/llm-chat/server.py +664 -0
  22. package/mcp-servers/manual-review/README.md +133 -0
  23. package/mcp-servers/manual-review/server.py +910 -0
  24. package/mcp-servers/manual-review/ui.html +279 -0
  25. package/mcp-servers/minimax-chat/requirements.txt +1 -0
  26. package/mcp-servers/minimax-chat/server.py +381 -0
  27. package/package.json +51 -0
  28. package/skills/ablation-planner/SKILL.md +123 -0
  29. package/skills/alphaxiv/SKILL.md +196 -0
  30. package/skills/analyze-results/SKILL.md +46 -0
  31. package/skills/arxiv/SKILL.md +248 -0
  32. package/skills/auto-paper-improvement-loop/SKILL.md +651 -0
  33. package/skills/auto-review-loop/SKILL.md +1137 -0
  34. package/skills/auto-review-loop-llm/SKILL.md +259 -0
  35. package/skills/auto-review-loop-minimax/SKILL.md +302 -0
  36. package/skills/citation-audit/SKILL.md +502 -0
  37. package/skills/claims-drafting/SKILL.md +227 -0
  38. package/skills/comm-lit-review/SKILL.md +297 -0
  39. package/skills/deepxiv/SKILL.md +263 -0
  40. package/skills/dse-loop/SKILL.md +296 -0
  41. package/skills/embodiment-description/SKILL.md +129 -0
  42. package/skills/exa-search/SKILL.md +205 -0
  43. package/skills/experiment-audit/SKILL.md +311 -0
  44. package/skills/experiment-bridge/SKILL.md +376 -0
  45. package/skills/experiment-plan/SKILL.md +249 -0
  46. package/skills/experiment-queue/SKILL.md +431 -0
  47. package/skills/experiment-queue/scripts/build_manifest.py +142 -0
  48. package/skills/experiment-queue/scripts/queue_manager.py +433 -0
  49. package/skills/feishu-notify/SKILL.md +156 -0
  50. package/skills/figure-description/SKILL.md +138 -0
  51. package/skills/figure-spec/SKILL.md +262 -0
  52. package/skills/figure-spec/scripts/figure_renderer.py +799 -0
  53. package/skills/formula-derivation/SKILL.md +280 -0
  54. package/skills/gemini-search/SKILL.md +231 -0
  55. package/skills/grant-proposal/SKILL.md +698 -0
  56. package/skills/idea-creator/SKILL.md +542 -0
  57. package/skills/idea-discovery/SKILL.md +521 -0
  58. package/skills/idea-discovery-robot/SKILL.md +363 -0
  59. package/skills/integrity-forensics/SKILL.md +284 -0
  60. package/skills/interview-cheatsheet/SKILL.md +245 -0
  61. package/skills/invention-structuring/SKILL.md +188 -0
  62. package/skills/jurisdiction-format/SKILL.md +192 -0
  63. package/skills/kill-argument/SKILL.md +437 -0
  64. package/skills/mermaid-diagram/SKILL.md +419 -0
  65. package/skills/meta-apply/SKILL.md +141 -0
  66. package/skills/meta-optimize/SKILL.md +437 -0
  67. package/skills/monitor-experiment/SKILL.md +140 -0
  68. package/skills/novelty-check/SKILL.md +101 -0
  69. package/skills/openalex/SKILL.md +237 -0
  70. package/skills/overleaf-sync/SKILL.md +220 -0
  71. package/skills/paper-claim-audit/SKILL.md +348 -0
  72. package/skills/paper-compile/SKILL.md +266 -0
  73. package/skills/paper-figure/SKILL.md +312 -0
  74. package/skills/paper-illustration/SKILL.md +736 -0
  75. package/skills/paper-illustration-image2/SKILL.md +391 -0
  76. package/skills/paper-illustration-image2/scripts/paper_illustration_image2.py +255 -0
  77. package/skills/paper-plan/SKILL.md +386 -0
  78. package/skills/paper-poster/SKILL.md +19 -0
  79. package/skills/paper-poster-html/DESIGN_FINAL.md +176 -0
  80. package/skills/paper-poster-html/IMPLEMENTATION_CONVENTIONS.md +161 -0
  81. package/skills/paper-poster-html/LICENSES/posterly-MIT.txt +21 -0
  82. package/skills/paper-poster-html/NOTICE.md +57 -0
  83. package/skills/paper-poster-html/SKILL.md +323 -0
  84. package/skills/paper-poster-html/scripts/_posterly/__init__.py +0 -0
  85. package/skills/paper-poster-html/scripts/_posterly/canvas.py +200 -0
  86. package/skills/paper-poster-html/scripts/_posterly/measure.py +588 -0
  87. package/skills/paper-poster-html/scripts/_posterly/polish.py +498 -0
  88. package/skills/paper-poster-html/scripts/_posterly/preflight.py +489 -0
  89. package/skills/paper-poster-html/scripts/_posterly/render.py +215 -0
  90. package/skills/paper-poster-html/scripts/_posterly/textutil.py +16 -0
  91. package/skills/paper-poster-html/scripts/_posterly/verify_final.py +171 -0
  92. package/skills/paper-poster-html/scripts/asset_check.py +897 -0
  93. package/skills/paper-poster-html/scripts/extract_pdf_figures.py +666 -0
  94. package/skills/paper-poster-html/scripts/poster_check.py +251 -0
  95. package/skills/paper-poster-html/scripts/preprocess_figures.py +238 -0
  96. package/skills/paper-poster-html/scripts/render_preview.py +217 -0
  97. package/skills/paper-poster-html/scripts/run_gates.py +556 -0
  98. package/skills/paper-poster-html/scripts/style_check.py +1324 -0
  99. package/skills/paper-poster-html/templates/COMPONENTS.md +462 -0
  100. package/skills/paper-poster-html/templates/README.md +170 -0
  101. package/skills/paper-poster-html/templates/landscape_4col.html +1032 -0
  102. package/skills/paper-poster-html/templates/landscape_hero.html +1046 -0
  103. package/skills/paper-poster-html/templates/portrait_2col.html +947 -0
  104. package/skills/paper-poster-html/templates/tokens/acl.json +9 -0
  105. package/skills/paper-poster-html/templates/tokens/cvpr.json +9 -0
  106. package/skills/paper-poster-html/templates/tokens/generic.json +9 -0
  107. package/skills/paper-poster-html/templates/tokens/iclr.json +9 -0
  108. package/skills/paper-poster-html/templates/tokens/icml.json +9 -0
  109. package/skills/paper-poster-html/templates/tokens/neurips.json +9 -0
  110. package/skills/paper-slides/SKILL.md +635 -0
  111. package/skills/paper-talk/SKILL.md +381 -0
  112. package/skills/paper-write/SKILL.md +604 -0
  113. package/skills/paper-write/templates/IEEEtran.bst +2409 -0
  114. package/skills/paper-write/templates/IEEEtran.cls +6347 -0
  115. package/skills/paper-write/templates/iclr2026.tex +84 -0
  116. package/skills/paper-write/templates/icml2025.tex +87 -0
  117. package/skills/paper-write/templates/ieee_conference.tex +89 -0
  118. package/skills/paper-write/templates/ieee_journal.tex +93 -0
  119. package/skills/paper-write/templates/math_commands.tex +48 -0
  120. package/skills/paper-write/templates/neurips2025.tex +80 -0
  121. package/skills/paper-writing/SKILL.md +916 -0
  122. package/skills/patent-novelty-check/SKILL.md +153 -0
  123. package/skills/patent-pipeline/SKILL.md +344 -0
  124. package/skills/patent-review/SKILL.md +203 -0
  125. package/skills/pixel-art/SKILL.md +137 -0
  126. package/skills/prior-art-search/SKILL.md +146 -0
  127. package/skills/proof-checker/SKILL.md +866 -0
  128. package/skills/proof-orchestrator/NOTICE.md +24 -0
  129. package/skills/proof-orchestrator/SKILL.md +254 -0
  130. package/skills/proof-orchestrator/references/audit-output-contract.md +126 -0
  131. package/skills/proof-orchestrator/references/deepseek-routing.md +74 -0
  132. package/skills/proof-orchestrator/references/dispatch-prompts.md +227 -0
  133. package/skills/proof-orchestrator/references/notation-audit.md +135 -0
  134. package/skills/proof-orchestrator/references/proof-audit-rubric.md +70 -0
  135. package/skills/proof-orchestrator/references/stress-tests.md +38 -0
  136. package/skills/proof-writer/SKILL.md +223 -0
  137. package/skills/qzcli/SKILL.md +324 -0
  138. package/skills/rebuttal/SKILL.md +376 -0
  139. package/skills/render-html/SKILL.md +316 -0
  140. package/skills/render-html/scripts/render_html.py +1006 -0
  141. package/skills/render-html/scripts/templates/academic.html +703 -0
  142. package/skills/render-html/scripts/templates/dashboard.html +333 -0
  143. package/skills/research-lit/SKILL.md +756 -0
  144. package/skills/research-pipeline/SKILL.md +384 -0
  145. package/skills/research-refine/SKILL.md +770 -0
  146. package/skills/research-refine-pipeline/SKILL.md +186 -0
  147. package/skills/research-review/SKILL.md +198 -0
  148. package/skills/research-wiki/SKILL.md +461 -0
  149. package/skills/resubmit-pipeline/SKILL.md +447 -0
  150. package/skills/result-to-claim/SKILL.md +311 -0
  151. package/skills/run-experiment/SKILL.md +313 -0
  152. package/skills/semantic-scholar/SKILL.md +236 -0
  153. package/skills/serverless-modal/SKILL.md +335 -0
  154. package/skills/shared-references/acceptance-gate.md +324 -0
  155. package/skills/shared-references/assurance-contract.md +248 -0
  156. package/skills/shared-references/capture-antipatterns.md +78 -0
  157. package/skills/shared-references/citation-discipline.md +583 -0
  158. package/skills/shared-references/compute-env-contract.md +163 -0
  159. package/skills/shared-references/effort-contract.md +183 -0
  160. package/skills/shared-references/evidence-precheck.md +65 -0
  161. package/skills/shared-references/experiment-integrity.md +49 -0
  162. package/skills/shared-references/external-cadence.md +326 -0
  163. package/skills/shared-references/fan-out-pattern.md +366 -0
  164. package/skills/shared-references/injection-hygiene.md +127 -0
  165. package/skills/shared-references/integration-contract.md +461 -0
  166. package/skills/shared-references/output-composition.md +93 -0
  167. package/skills/shared-references/output-language.md +45 -0
  168. package/skills/shared-references/output-manifest.md +49 -0
  169. package/skills/shared-references/output-versioning.md +111 -0
  170. package/skills/shared-references/patent-format-cn.md +199 -0
  171. package/skills/shared-references/patent-format-ep.md +173 -0
  172. package/skills/shared-references/patent-format-us.md +161 -0
  173. package/skills/shared-references/patent-writing-principles.md +197 -0
  174. package/skills/shared-references/prior-art-databases.md +141 -0
  175. package/skills/shared-references/resumable-runs.md +109 -0
  176. package/skills/shared-references/review-scope-limits.md +81 -0
  177. package/skills/shared-references/review-tracing.md +391 -0
  178. package/skills/shared-references/reviewer-independence.md +79 -0
  179. package/skills/shared-references/reviewer-routing.md +852 -0
  180. package/skills/shared-references/skill-governance.md +104 -0
  181. package/skills/shared-references/taste-calibration.md +85 -0
  182. package/skills/shared-references/venue-checklists.md +114 -0
  183. package/skills/shared-references/wiki-helper-resolution.md +134 -0
  184. package/skills/shared-references/writing-principles.md +525 -0
  185. package/skills/skills-codex/README.md +102 -0
  186. package/skills/skills-codex/README_CN.md +100 -0
  187. package/skills/skills-codex/ablation-planner/SKILL.md +126 -0
  188. package/skills/skills-codex/alphaxiv/SKILL.md +186 -0
  189. package/skills/skills-codex/analyze-results/SKILL.md +45 -0
  190. package/skills/skills-codex/arxiv/SKILL.md +210 -0
  191. package/skills/skills-codex/auto-paper-improvement-loop/SKILL.md +574 -0
  192. package/skills/skills-codex/auto-review-loop/SKILL.md +500 -0
  193. package/skills/skills-codex/auto-review-loop-llm/SKILL.md +247 -0
  194. package/skills/skills-codex/auto-review-loop-minimax/SKILL.md +290 -0
  195. package/skills/skills-codex/citation-audit/SKILL.md +504 -0
  196. package/skills/skills-codex/claims-drafting/SKILL.md +239 -0
  197. package/skills/skills-codex/comm-lit-review/SKILL.md +299 -0
  198. package/skills/skills-codex/comm-lit-review/references/domain-taxonomy.md +57 -0
  199. package/skills/skills-codex/comm-lit-review/references/output-template.md +37 -0
  200. package/skills/skills-codex/comm-lit-review/references/source-policy.md +99 -0
  201. package/skills/skills-codex/comm-lit-review/references/venue-tiering.md +112 -0
  202. package/skills/skills-codex/deepxiv/SKILL.md +142 -0
  203. package/skills/skills-codex/dse-loop/SKILL.md +285 -0
  204. package/skills/skills-codex/embodiment-description/SKILL.md +129 -0
  205. package/skills/skills-codex/exa-search/SKILL.md +192 -0
  206. package/skills/skills-codex/experiment-audit/SKILL.md +286 -0
  207. package/skills/skills-codex/experiment-bridge/SKILL.md +356 -0
  208. package/skills/skills-codex/experiment-plan/SKILL.md +249 -0
  209. package/skills/skills-codex/experiment-queue/SKILL.md +401 -0
  210. package/skills/skills-codex/feishu-notify/SKILL.md +155 -0
  211. package/skills/skills-codex/figure-description/SKILL.md +138 -0
  212. package/skills/skills-codex/figure-spec/SKILL.md +252 -0
  213. package/skills/skills-codex/formula-derivation/SKILL.md +280 -0
  214. package/skills/skills-codex/gemini-search/SKILL.md +205 -0
  215. package/skills/skills-codex/grant-proposal/SKILL.md +626 -0
  216. package/skills/skills-codex/idea-creator/SKILL.md +405 -0
  217. package/skills/skills-codex/idea-discovery/SKILL.md +475 -0
  218. package/skills/skills-codex/idea-discovery-robot/SKILL.md +362 -0
  219. package/skills/skills-codex/integrity-forensics/SKILL.md +106 -0
  220. package/skills/skills-codex/interview-cheatsheet/SKILL.md +245 -0
  221. package/skills/skills-codex/invention-structuring/SKILL.md +188 -0
  222. package/skills/skills-codex/jurisdiction-format/SKILL.md +192 -0
  223. package/skills/skills-codex/kill-argument/SKILL.md +403 -0
  224. package/skills/skills-codex/mermaid-diagram/SKILL.md +379 -0
  225. package/skills/skills-codex/meta-apply/SKILL.md +154 -0
  226. package/skills/skills-codex/meta-optimize/SKILL.md +348 -0
  227. package/skills/skills-codex/monitor-experiment/SKILL.md +98 -0
  228. package/skills/skills-codex/novelty-check/SKILL.md +89 -0
  229. package/skills/skills-codex/openalex/SKILL.md +228 -0
  230. package/skills/skills-codex/overleaf-sync/SKILL.md +220 -0
  231. package/skills/skills-codex/paper-claim-audit/SKILL.md +350 -0
  232. package/skills/skills-codex/paper-compile/SKILL.md +253 -0
  233. package/skills/skills-codex/paper-figure/SKILL.md +311 -0
  234. package/skills/skills-codex/paper-illustration/SKILL.md +690 -0
  235. package/skills/skills-codex/paper-illustration-image2/SKILL.md +383 -0
  236. package/skills/skills-codex/paper-illustration-image2/scripts/paper_illustration_image2.py +255 -0
  237. package/skills/skills-codex/paper-plan/SKILL.md +278 -0
  238. package/skills/skills-codex/paper-poster/SKILL.md +19 -0
  239. package/skills/skills-codex/paper-poster-html/SKILL.md +377 -0
  240. package/skills/skills-codex/paper-slides/SKILL.md +571 -0
  241. package/skills/skills-codex/paper-talk/SKILL.md +381 -0
  242. package/skills/skills-codex/paper-write/SKILL.md +411 -0
  243. package/skills/skills-codex/paper-write/templates/IEEEtran.bst +2409 -0
  244. package/skills/skills-codex/paper-write/templates/IEEEtran.cls +6347 -0
  245. package/skills/skills-codex/paper-write/templates/aaai2026.bst +1493 -0
  246. package/skills/skills-codex/paper-write/templates/aaai2026.sty +315 -0
  247. package/skills/skills-codex/paper-write/templates/aaai2026.tex +952 -0
  248. package/skills/skills-codex/paper-write/templates/acl.sty +312 -0
  249. package/skills/skills-codex/paper-write/templates/acl2026.tex +377 -0
  250. package/skills/skills-codex/paper-write/templates/acl_natbib.bst +1940 -0
  251. package/skills/skills-codex/paper-write/templates/acm.bst +3081 -0
  252. package/skills/skills-codex/paper-write/templates/acm_mm2026.tex +204 -0
  253. package/skills/skills-codex/paper-write/templates/acmart.cls +3520 -0
  254. package/skills/skills-codex/paper-write/templates/cvpr.bst +1448 -0
  255. package/skills/skills-codex/paper-write/templates/cvpr.sty +508 -0
  256. package/skills/skills-codex/paper-write/templates/cvpr2026.tex +63 -0
  257. package/skills/skills-codex/paper-write/templates/iclr2026.tex +84 -0
  258. package/skills/skills-codex/paper-write/templates/iclr2026_conference.bst +1440 -0
  259. package/skills/skills-codex/paper-write/templates/iclr2026_conference.sty +246 -0
  260. package/skills/skills-codex/paper-write/templates/icml2026.sty +767 -0
  261. package/skills/skills-codex/paper-write/templates/icml2026.tex +662 -0
  262. package/skills/skills-codex/paper-write/templates/ieee_conference.tex +89 -0
  263. package/skills/skills-codex/paper-write/templates/ieee_journal.tex +93 -0
  264. package/skills/skills-codex/paper-write/templates/math_commands.tex +48 -0
  265. package/skills/skills-codex/paper-write/templates/neurips2026.tex +493 -0
  266. package/skills/skills-codex/paper-write/templates/neurips_2026.sty +437 -0
  267. package/skills/skills-codex/paper-writing/SKILL.md +731 -0
  268. package/skills/skills-codex/patent-novelty-check/SKILL.md +153 -0
  269. package/skills/skills-codex/patent-pipeline/SKILL.md +344 -0
  270. package/skills/skills-codex/patent-review/SKILL.md +202 -0
  271. package/skills/skills-codex/pixel-art/SKILL.md +139 -0
  272. package/skills/skills-codex/prior-art-search/SKILL.md +146 -0
  273. package/skills/skills-codex/proof-checker/SKILL.md +554 -0
  274. package/skills/skills-codex/proof-orchestrator/SKILL.md +260 -0
  275. package/skills/skills-codex/proof-orchestrator/references/audit-output-contract.md +126 -0
  276. package/skills/skills-codex/proof-orchestrator/references/deepseek-routing.md +76 -0
  277. package/skills/skills-codex/proof-orchestrator/references/dispatch-prompts.md +227 -0
  278. package/skills/skills-codex/proof-orchestrator/references/notation-audit.md +135 -0
  279. package/skills/skills-codex/proof-orchestrator/references/proof-audit-rubric.md +70 -0
  280. package/skills/skills-codex/proof-orchestrator/references/stress-tests.md +38 -0
  281. package/skills/skills-codex/proof-writer/SKILL.md +222 -0
  282. package/skills/skills-codex/qzcli/SKILL.md +324 -0
  283. package/skills/skills-codex/rebuttal/SKILL.md +305 -0
  284. package/skills/skills-codex/render-html/SKILL.md +305 -0
  285. package/skills/skills-codex/render-html/scripts/__pycache__/render_html.cpython-314.pyc +0 -0
  286. package/skills/skills-codex/render-html/scripts/render_html.py +909 -0
  287. package/skills/skills-codex/render-html/scripts/templates/academic.html +342 -0
  288. package/skills/skills-codex/render-html/scripts/templates/dashboard.html +333 -0
  289. package/skills/skills-codex/research-lit/SKILL.md +464 -0
  290. package/skills/skills-codex/research-pipeline/SKILL.md +340 -0
  291. package/skills/skills-codex/research-refine/SKILL.md +721 -0
  292. package/skills/skills-codex/research-refine-pipeline/SKILL.md +186 -0
  293. package/skills/skills-codex/research-review/SKILL.md +135 -0
  294. package/skills/skills-codex/research-wiki/SKILL.md +421 -0
  295. package/skills/skills-codex/resubmit-pipeline/SKILL.md +444 -0
  296. package/skills/skills-codex/result-to-claim/SKILL.md +246 -0
  297. package/skills/skills-codex/run-experiment/SKILL.md +236 -0
  298. package/skills/skills-codex/semantic-scholar/SKILL.md +219 -0
  299. package/skills/skills-codex/serverless-modal/SKILL.md +335 -0
  300. package/skills/skills-codex/shared-references/acceptance-gate.md +336 -0
  301. package/skills/skills-codex/shared-references/assurance-contract.md +139 -0
  302. package/skills/skills-codex/shared-references/capture-antipatterns.md +84 -0
  303. package/skills/skills-codex/shared-references/citation-discipline.md +452 -0
  304. package/skills/skills-codex/shared-references/compute-env-contract.md +163 -0
  305. package/skills/skills-codex/shared-references/effort-contract.md +143 -0
  306. package/skills/skills-codex/shared-references/evidence-precheck.md +73 -0
  307. package/skills/skills-codex/shared-references/experiment-integrity.md +49 -0
  308. package/skills/skills-codex/shared-references/external-cadence.md +334 -0
  309. package/skills/skills-codex/shared-references/fan-out-pattern.md +375 -0
  310. package/skills/skills-codex/shared-references/injection-hygiene.md +132 -0
  311. package/skills/skills-codex/shared-references/integration-contract.md +372 -0
  312. package/skills/skills-codex/shared-references/output-composition.md +98 -0
  313. package/skills/skills-codex/shared-references/output-language.md +45 -0
  314. package/skills/skills-codex/shared-references/output-manifest.md +40 -0
  315. package/skills/skills-codex/shared-references/output-versioning.md +111 -0
  316. package/skills/skills-codex/shared-references/patent-format-cn.md +199 -0
  317. package/skills/skills-codex/shared-references/patent-format-ep.md +173 -0
  318. package/skills/skills-codex/shared-references/patent-format-us.md +161 -0
  319. package/skills/skills-codex/shared-references/patent-writing-principles.md +197 -0
  320. package/skills/skills-codex/shared-references/prior-art-databases.md +141 -0
  321. package/skills/skills-codex/shared-references/resumable-runs.md +125 -0
  322. package/skills/skills-codex/shared-references/review-scope-limits.md +81 -0
  323. package/skills/skills-codex/shared-references/review-tracing.md +144 -0
  324. package/skills/skills-codex/shared-references/reviewer-independence.md +66 -0
  325. package/skills/skills-codex/shared-references/reviewer-routing.md +128 -0
  326. package/skills/skills-codex/shared-references/skill-governance.md +119 -0
  327. package/skills/skills-codex/shared-references/taste-calibration.md +90 -0
  328. package/skills/skills-codex/shared-references/venue-checklists.md +73 -0
  329. package/skills/skills-codex/shared-references/wiki-helper-resolution.md +69 -0
  330. package/skills/skills-codex/shared-references/writing-principles.md +525 -0
  331. package/skills/skills-codex/slides-polish/SKILL.md +563 -0
  332. package/skills/skills-codex/specification-writing/SKILL.md +211 -0
  333. package/skills/skills-codex/system-profile/SKILL.md +103 -0
  334. package/skills/skills-codex/training-check/SKILL.md +83 -0
  335. package/skills/skills-codex/vast-gpu/SKILL.md +394 -0
  336. package/skills/skills-codex/web-debug-search/SKILL.md +334 -0
  337. package/skills/skills-codex/wiki-enrich/SKILL.md +255 -0
  338. package/skills/skills-codex/writing-systems-papers/SKILL.md +184 -0
  339. package/skills/skills-codex-claude-review/README.md +79 -0
  340. package/skills/skills-codex-claude-review/README_CN.md +78 -0
  341. package/skills/skills-codex-claude-review/auto-paper-improvement-loop/SKILL.md +581 -0
  342. package/skills/skills-codex-claude-review/auto-review-loop/SKILL.md +510 -0
  343. package/skills/skills-codex-claude-review/novelty-check/SKILL.md +102 -0
  344. package/skills/skills-codex-claude-review/paper-figure/SKILL.md +319 -0
  345. package/skills/skills-codex-claude-review/paper-plan/SKILL.md +287 -0
  346. package/skills/skills-codex-claude-review/paper-write/SKILL.md +420 -0
  347. package/skills/skills-codex-claude-review/research-refine/SKILL.md +732 -0
  348. package/skills/skills-codex-claude-review/research-review/SKILL.md +149 -0
  349. package/skills/skills-codex-gemini-review/README.md +176 -0
  350. package/skills/skills-codex-gemini-review/README_CN.md +175 -0
  351. package/skills/skills-codex-gemini-review/auto-paper-improvement-loop/SKILL.md +331 -0
  352. package/skills/skills-codex-gemini-review/auto-review-loop/SKILL.md +304 -0
  353. package/skills/skills-codex-gemini-review/grant-proposal/SKILL.md +630 -0
  354. package/skills/skills-codex-gemini-review/idea-creator/SKILL.md +263 -0
  355. package/skills/skills-codex-gemini-review/idea-discovery/SKILL.md +275 -0
  356. package/skills/skills-codex-gemini-review/idea-discovery-robot/SKILL.md +365 -0
  357. package/skills/skills-codex-gemini-review/novelty-check/SKILL.md +92 -0
  358. package/skills/skills-codex-gemini-review/paper-figure/SKILL.md +289 -0
  359. package/skills/skills-codex-gemini-review/paper-plan/SKILL.md +265 -0
  360. package/skills/skills-codex-gemini-review/paper-poster-html/SKILL.md +102 -0
  361. package/skills/skills-codex-gemini-review/paper-slides/SKILL.md +582 -0
  362. package/skills/skills-codex-gemini-review/paper-write/SKILL.md +346 -0
  363. package/skills/skills-codex-gemini-review/paper-writing/SKILL.md +312 -0
  364. package/skills/skills-codex-gemini-review/research-refine/SKILL.md +674 -0
  365. package/skills/skills-codex-gemini-review/research-review/SKILL.md +112 -0
  366. package/skills/slides-polish/SKILL.md +565 -0
  367. package/skills/specification-writing/SKILL.md +211 -0
  368. package/skills/system-profile/SKILL.md +103 -0
  369. package/skills/training-check/SKILL.md +132 -0
  370. package/skills/vast-gpu/SKILL.md +394 -0
  371. package/skills/web-debug-search/SKILL.md +334 -0
  372. package/skills/wiki-enrich/SKILL.md +257 -0
  373. package/skills/writing-systems-papers/SKILL.md +184 -0
  374. package/templates/CLAUDE_MD_TEMPLATE.md +29 -0
  375. package/templates/EXPERIMENT_LOG_TEMPLATE.md +47 -0
  376. package/templates/EXPERIMENT_PLAN_TEMPLATE.md +51 -0
  377. package/templates/EXPERIMENT_PLAN_TEMPLATE_CN.md +53 -0
  378. package/templates/FINDINGS_TEMPLATE.md +52 -0
  379. package/templates/IDEA_CANDIDATES_TEMPLATE.md +47 -0
  380. package/templates/IDEA_CANDIDATES_TEMPLATE_CN.md +47 -0
  381. package/templates/INVENTION_BRIEF_TEMPLATE.md +87 -0
  382. package/templates/MANIFEST_TEMPLATE.md +7 -0
  383. package/templates/NARRATIVE_REPORT_TEMPLATE.md +49 -0
  384. package/templates/PAPER_PLAN_TEMPLATE.md +47 -0
  385. package/templates/PATENT_CLAIMS_TEMPLATE.md +78 -0
  386. package/templates/PATENT_SPECIFICATION_TEMPLATE.md +68 -0
  387. package/templates/README.md +57 -0
  388. package/templates/RESEARCH_BRIEF_TEMPLATE.md +35 -0
  389. package/templates/RESEARCH_BRIEF_TEMPLATE_CN.md +41 -0
  390. package/templates/RESEARCH_CONTRACT_TEMPLATE.md +60 -0
  391. package/templates/claude-hooks/corpus_write_guard.json +16 -0
  392. package/templates/claude-hooks/corpus_write_guard.py +85 -0
  393. package/templates/claude-hooks/meta_logging.json +74 -0
  394. package/templates/gitignore-trace.txt +3 -0
  395. package/tools/__pycache__/check_skills_inventory.cpython-314.pyc +0 -0
  396. package/tools/arxiv_fetch.py +311 -0
  397. package/tools/capture_filter.py +126 -0
  398. package/tools/check_skills_inventory.py +273 -0
  399. package/tools/convert_skills_to_llm_chat.py +282 -0
  400. package/tools/copilot_native_evidence.py +818 -0
  401. package/tools/deepxiv_fetch.py +213 -0
  402. package/tools/evidence_check.py +212 -0
  403. package/tools/exa_search.py +425 -0
  404. package/tools/experiment_queue/README.md +118 -0
  405. package/tools/experiment_queue/build_manifest.py +44 -0
  406. package/tools/experiment_queue/queue_manager.py +44 -0
  407. package/tools/extract_paper_style.py +560 -0
  408. package/tools/figure_renderer.py +69 -0
  409. package/tools/forensics_gate.py +669 -0
  410. package/tools/generate_codex_claude_review_overrides.py +299 -0
  411. package/tools/idea_discovery_gate.py +256 -0
  412. package/tools/install_aris.ps1 +1372 -0
  413. package/tools/install_aris.sh +1370 -0
  414. package/tools/install_aris_codex.sh +1023 -0
  415. package/tools/install_aris_copilot.sh +1052 -0
  416. package/tools/iteration_log.py +143 -0
  417. package/tools/lint_skills_helpers.sh +84 -0
  418. package/tools/meta_opt/check_ready.sh +80 -0
  419. package/tools/meta_opt/log_event.sh +91 -0
  420. package/tools/meta_opt/trigger_eval.py +280 -0
  421. package/tools/meta_opt/trigger_evals.sample.json +28 -0
  422. package/tools/openalex_fetch.py +326 -0
  423. package/tools/overleaf_audit.sh +104 -0
  424. package/tools/overleaf_setup.sh +150 -0
  425. package/tools/paper_illustration_image2.py +62 -0
  426. package/tools/provenance.py +294 -0
  427. package/tools/research_wiki.py +1720 -0
  428. package/tools/review_gate.py +502 -0
  429. package/tools/run_state.py +399 -0
  430. package/tools/save_trace.sh +477 -0
  431. package/tools/semantic_scholar_fetch.py +438 -0
  432. package/tools/skill-groups.tsv +116 -0
  433. package/tools/skill_picker.py +238 -0
  434. package/tools/smart_update.ps1 +521 -0
  435. package/tools/smart_update.sh +591 -0
  436. package/tools/smart_update_codex.sh +419 -0
  437. package/tools/smart_update_copilot.sh +605 -0
  438. package/tools/threat_scan.py +222 -0
  439. package/tools/verify_paper_audits.sh +487 -0
  440. package/tools/verify_papers.py +613 -0
  441. package/tools/verify_wiki_coverage.sh +176 -0
  442. package/tools/watchdog.py +485 -0
@@ -0,0 +1,324 @@
1
+ # Acceptance-Gate Provenance
2
+
3
+ ## Core Principle
4
+
5
+ **An autonomous loop's STOP/ACCEPT gate determines whether the loop is
6
+ same-family-safe. The thing being judged at that gate — not the loop's
7
+ subject matter, not how many agents ran — decides whether Claude may
8
+ judge it.**
9
+
10
+ ARIS has loops that keep working until a condition is met: `/auto-review-loop`,
11
+ `/dse-loop`, the `/experiment-bridge` auto-debug cycle, the
12
+ `/auto-paper-improvement-loop`, and any future "keep going until X"
13
+ skill. Every such loop terminates on a gate it evaluates each iteration:
14
+ "are we done yet?" That gate is where same-family self-acquittal sneaks
15
+ in. The loop body can be all Claude; the **gate** is what this contract
16
+ governs.
17
+
18
+ This is `reviewer-independence.md` and `experiment-integrity.md` applied
19
+ to the temporal/iterative case: those two cover single-shot review and
20
+ single-shot experiment judging; this one covers the *recurring verdict*
21
+ a loop makes on itself, round after round, with no human in between.
22
+
23
+ One-liner, and the whole doc in seven words:
24
+
25
+ > **A goal/loop can DRIVE; it cannot ACQUIT.**
26
+
27
+ The loop may freely *drive* itself toward a target — schedule the next
28
+ config, recompile, re-run the failed job, spawn ten search branches.
29
+ What it may not do is *acquit* its own work — declare the paper good,
30
+ the proof valid, the claim supported, the idea novel, the review
31
+ satisfied. Acquittal is a cross-model act.
32
+
33
+ ## The two gate types
34
+
35
+ Classify **every** stop/accept gate of a loop as exactly one of these.
36
+ There is no third bucket; if a gate seems to be both, it is two gates
37
+ and you split it (see "Compound gates" below).
38
+
39
+ ### Type-A — EXECUTION / OBJECTIVE gate
40
+
41
+ A machine-checkable or externally-observable signal of *what happened*,
42
+ with no judgment of *merit*. Claude **MAY** self-judge Type-A gates —
43
+ it is execution bookkeeping, not a verdict.
44
+
45
+ A gate is Type-A iff a non-LLM process (a shell exit code, a stat on the
46
+ filesystem, a counter, a parser reading a benchmark's own output) could
47
+ in principle answer it with the same answer Claude gives.
48
+
49
+ - ✅ exit code == 0
50
+ - ✅ `figures/result.png` exists / `paper/main.pdf` compiled (LaTeX returned 0)
51
+ - ✅ N/N jobs finished (queue drained)
52
+ - ✅ test suite passed (pytest exit 0)
53
+ - ✅ the reviewer **was invoked** (a `codex` thread returned, a JSON verdict file exists)
54
+ - ✅ all checklist items were **attempted** (each row touched)
55
+ - ✅ no `NaN` in the loss log / training reached `max_steps`
56
+ - ✅ the benchmark harness emitted a number and it parsed
57
+ - ✅ PATIENCE/TIMEOUT/MAX_ROUNDS budget exhausted (a counter hit its bound)
58
+
59
+ Type-A gates are *coverage and completion* facts. Claude self-judging
60
+ "did the audit run?" is fine; Claude self-judging "did the audit pass?"
61
+ is not (that's Type-B).
62
+
63
+ ### Type-B — QUALITY / CORRECTNESS / ACCEPTANCE gate
64
+
65
+ A judgment of *merit, correctness, or sufficiency*. Claude must
66
+ **NEVER** self-judge a Type-B gate — it requires a **different model
67
+ family** (per `reviewer-routing.md`: `codex` default, `oracle-pro` on
68
+ request, or `manual` **only when** the human routes the prompt to a
69
+ genuinely non-Claude model and records which one). This is the
70
+ cross-model invariant, applied to the loop's terminating verdict.
71
+
72
+ - ❌ "the paper is good" / "submission-ready"
73
+ - ❌ "the proof is valid" / "the gap is closed"
74
+ - ❌ "the claim is supported by the results"
75
+ - ❌ "the idea is novel"
76
+ - ❌ "the review is satisfied" / "the weaknesses are addressed"
77
+ - ❌ "score >= 6" — when *Claude* assigned the score
78
+ - ❌ "this config is good enough to publish" / "the result is strong"
79
+ - ❌ "the rebuttal answers the reviewer"
80
+ - ❌ "the fix is correct" (as opposed to "the fix made the test pass" — that's Type-A)
81
+
82
+ A Type-B gate, left to the executor, is the loop quietly grading its own
83
+ homework every round and stopping the moment it likes the grade. The
84
+ fact that it ran a hundred iterations does not launder the verdict: a
85
+ hundred rounds of Claude-judging-Claude is still one model family.
86
+
87
+ ### The dividing question
88
+
89
+ > *Could a dumb script with no taste answer this gate?*
90
+ >
91
+ > **Yes → Type-A** (Claude may self-judge — it's bookkeeping).
92
+ > **No, it needs taste / correctness / domain judgment → Type-B** (route to a different model family).
93
+
94
+ "The PDF compiled" needs no taste — Type-A. "The PDF is a good paper"
95
+ is nothing *but* taste — Type-B. "The job exited 0" — Type-A. "The job's
96
+ output is the right answer" — Type-B.
97
+
98
+ ## Compound gates: split, don't average
99
+
100
+ Many natural-language stop conditions secretly bundle an A-part and a
101
+ B-part. `/auto-review-loop`'s real condition is *"score >= 6 AND verdict
102
+ contains 'ready'"* evaluated each round — but the **score and the
103
+ verdict both come from the cross-model reviewer**, so the A-part Claude
104
+ owns is only "did round N's reviewer return?" and "is round < MAX_ROUNDS?".
105
+
106
+ When you meet a compound gate, decompose it:
107
+
108
+ ```
109
+ STOP when "the paper is submission-ready"
110
+ ├─ A: all 3 audits were invoked and emitted JSON → Claude self-checks
111
+ ├─ A: verify_paper_audits.sh exit code == 0 → external process, Claude reads it
112
+ └─ B: "the paper is actually good enough to submit" → cross-model verdict
113
+ ```
114
+
115
+ Never collapse a compound gate to its A-part and call the loop safe. The
116
+ B-part doesn't disappear because it's inconvenient; it gets *routed*.
117
+
118
+ ## Decision procedure (for any new autonomous loop)
119
+
120
+ When you author or review a "keep working until X" skill:
121
+
122
+ 1. **Enumerate every stop/accept gate.** Not just the headline one —
123
+ the early-exit on convergence, the PATIENCE bail-out, the
124
+ per-iteration "is this round done?" check, the final "are we
125
+ finished?" check. Write them down.
126
+
127
+ 2. **Classify each gate A or B** using the dividing question. If it's
128
+ compound, split it (above) and classify the parts.
129
+
130
+ 3. **For every Type-A gate:** Claude may self-judge. Prefer an
131
+ *external* check where one exists (read an exit code, stat a file,
132
+ read a counter) over an LLM "I believe it finished" — Type-A is
133
+ exactly the place where a cheap deterministic check beats a vibe.
134
+
135
+ 4. **For every Type-B gate:** route it to a cross-model verdict per
136
+ `reviewer-routing.md` (default `mcp__codex__codex` at
137
+ `reasoning_effort: xhigh`; `oracle-pro` on request; `manual` only if
138
+ the routed model is verifiably non-Claude and recorded — otherwise it
139
+ is same-family self-acquittal in disguise). Pass file paths, not
140
+ summaries (`reviewer-independence.md`).
141
+ The loop **continues or stops on the reviewer's verdict**, not on
142
+ Claude's reading of it. Save the verdict as an artifact
143
+ (`integration-contract.md` §3) so a third party can confirm the
144
+ acquittal was external.
145
+
146
+ 5. **State the provenance in the SKILL.** One line: "STOP gate = Type-B,
147
+ routed to codex." A reviewer of the SKILL should be able to find,
148
+ for each terminating condition, which model family signs off.
149
+
150
+ 6. **Refuse the anti-pattern:** a loop whose continue/stop decision reads
151
+ an LLM-produced quality verdict that the **same** model family
152
+ (Claude) produced. That is self-acquittal regardless of how the
153
+ prompt is phrased.
154
+
155
+ Rule of thumb: **if removing the cross-model reviewer would still let
156
+ the loop decide to stop, the loop is self-acquitting.** A safe Type-B
157
+ loop is *designed* (by this contract) so that removing the external
158
+ family's verdict leaves it unable to terminate-accept — a design rule
159
+ the skill author enforces, not an automatic structural property.
160
+
161
+ ## ARIS loops mapped to the taxonomy
162
+
163
+ The codebase **already** follows this rule. This section makes the
164
+ implicit pattern explicit and operational for the next loop someone
165
+ writes.
166
+
167
+ | Loop | Headline stop gate | Type | Who acquits | Status |
168
+ |---|---|---|---|---|
169
+ | `/dse-loop` | objective metric converged / TIMEOUT / PATIENCE | A | benchmark harness emits the number; Claude reads & compares to budget | ✅ safe same-model |
170
+ | `/experiment-bridge` auto-debug | "did it run / did it converge" (exit 0, no NaN, training started) | A | exit codes, log parse | ✅ safe same-model |
171
+ | `/run-experiment`, `/experiment-queue` retry | job finished / OOM-retry exhausted / N jobs done | A | scheduler + exit codes | ✅ safe same-model |
172
+ | `/auto-review-loop` | score >= 6 AND verdict "ready", per round | B | **codex** assigns score & verdict | ✅ already cross-model |
173
+ | `/auto-paper-improvement-loop` | "review satisfied" (2 rounds) | B | **codex (GPT xhigh)** review | ✅ already cross-model |
174
+ | `/result-to-claim` | `claim_supported ∈ {yes,partial,no}` + `integrity_status` | B | **codex** judges results vs claims | ✅ cross-model |
175
+ | `/kill-argument` | rejection memo → defense, residual issues | B | two fresh **codex** threads | ✅ cross-model |
176
+ | `/proof-checker` | each gap closed, per round | B | **codex** re-reviews each round | ✅ cross-model |
177
+ | `/experiment-audit` | integrity verdict (fake GT, normalization fraud) | B | **codex** audits the eval code | ✅ cross-model |
178
+ | `/paper-claim-audit` | every number matches result files | B | fresh zero-context **cross-model** reviewer | ✅ cross-model |
179
+ | `/citation-audit` | every entry real & in-context | B | fresh **cross-model** reviewer | ✅ cross-model |
180
+ | `/paper-writing` Phase 6 (submission) | `verify_paper_audits.sh` exit 0 | A (gate) **wrapping** B (the audits) | external verifier reads cross-model JSON | ✅ A-gate over B-verdicts |
181
+
182
+ > 📌 The `/auto-review-loop` row reflects the skill's stop logic: `score >= 6`
183
+ > AND verdict contains "ready"/"almost", evaluated each round. (Its `Constants`
184
+ > block previously stated this with `OR` and a stale verdict vocabulary — an
185
+ > internal inconsistency now reconciled to the `AND` form the Phase-E stop
186
+ > check actually uses, in `auto-review-loop` and its `-llm`/`-minimax`
187
+ > siblings.) The acquittal is **codex's** score+verdict, so the Type-B
188
+ > classification is unchanged.
189
+
190
+ Two patterns to notice:
191
+
192
+ - **The execution loops (dse, auto-debug, queue) are Type-A all the way
193
+ down** — "did it run / did it converge" is a fact a harness reports.
194
+ They are *correctly* allowed to self-acquit, because there is nothing
195
+ of merit being judged: a converged number from a real simulator is an
196
+ observation, not an opinion. (The moment someone adds *"...and the
197
+ result is good enough to claim"* to a dse stop condition, that clause
198
+ is Type-B and must route out — see the dse caveat below.)
199
+
200
+ - **Every quality/correctness loop already routes its acquittal to
201
+ codex.** Nothing here is new behavior; the doc names the rule the
202
+ codebase converged on so the next author doesn't have to rediscover it
203
+ by getting reviewed.
204
+
205
+ ### The dse-loop caveat (objective ≠ acceptance)
206
+
207
+ `/dse-loop` optimizes a metric the benchmark *itself* produces (cycles,
208
+ area, coverage). "Config B beats config A on the harness's own number"
209
+ is Type-A — a parser, not Claude, owns it. But two adjacent judgments are
210
+ Type-B and must NOT be folded into the loop's self-acquittal:
211
+
212
+ - "this config is **good enough to ship/publish**" — sufficiency verdict.
213
+ - "the benchmark/metric **is the right thing to optimize** / the result
214
+ **generalizes**" — correctness-of-framing verdict.
215
+
216
+ So dse may self-terminate on *"best config found within budget"* (A), but
217
+ the claim *"and this is a publishable result"* leaves the loop and goes
218
+ through `/result-to-claim` (B). Driving the search is in-family; acquitting
219
+ the science is not.
220
+
221
+ ## Tie to fan-out: breadth is same-family; the jury is not
222
+
223
+ `fan-out-pattern.md` describes skill-layer fan-out — spawning multiple
224
+ agents for breadth (parallel search branches, per-section drafting,
225
+ per-entry citation checks). Fan-out interacts with this contract in
226
+ exactly one dangerous way:
227
+
228
+ **Same-family breadth is fine for Type-A coverage. It is NEVER a Type-B
229
+ jury.**
230
+
231
+ - ✅ Ten Claude branches each *attempting* a different search query, then
232
+ unioning hits — Type-A coverage (did we look broadly?). Self-judged
233
+ fine.
234
+ - ✅ N Claudes each drafting a section, a Type-A "all sections drafted"
235
+ completion check.
236
+ - ❌ N Claude reviewers each scoring the paper, then taking the
237
+ **majority/average as the accept verdict.** This *feels* like a jury
238
+ — independent voters! — but it is correlated same-family blindness
239
+ wearing a jury costume. N agreeing Claudes share the same training
240
+ priors and the same blind spots; their agreement is evidence of
241
+ shared bias, not of correctness. A Type-B verdict needs a **different
242
+ family**, not a *bigger N of the same one*.
243
+
244
+ > **Known failure mode:** "We ran the review 5× and all 5 said accept,
245
+ > so it's robust." Five draws from one distribution is one opinion with
246
+ > error bars, not five opinions. The cross-model invariant is about
247
+ > *family diversity*, not *sample count*. Fan-out scales breadth and
248
+ > Type-A coverage; it can never substitute for the one cross-family
249
+ > acquittal a Type-B gate requires.
250
+
251
+ Fan-out and this contract compose cleanly: fan-out (same family) does
252
+ the broad *driving*; the loop always funnels into the identical
253
+ cross-model *acquittal* at the Type-B gate. Breadth degrades gracefully
254
+ across runtimes (fewer parallel agents = slower, not unsafe); the
255
+ acquittal does not degrade — it is always the cross-family verdict, or
256
+ the loop is unsafe.
257
+
258
+ ## Required components (for a loop to claim same-family-safe)
259
+
260
+ A loop is same-family-safe iff **all** hold:
261
+
262
+ 1. **Every stop/accept gate is classified** A or B in the SKILL (compound
263
+ gates split).
264
+ 2. **Every Type-B gate routes to a cross-model verdict** per
265
+ `reviewer-routing.md`; the loop's continue/stop reads *that* verdict,
266
+ not a Claude re-judgment of it.
267
+ 3. **The cross-model verdict is an artifact** (`integration-contract.md`
268
+ §3) — a JSON/file a third party can inspect to confirm the acquittal
269
+ was external.
270
+ 4. **No same-family majority is treated as a Type-B jury** — fan-out
271
+ breadth never substitutes for cross-family acquittal.
272
+ 5. **Type-A self-judgment prefers an external check** (exit code, stat,
273
+ counter) over an LLM "I think it's done" wherever one exists.
274
+
275
+ If any fails, the loop can self-acquit and is **not** same-family-safe —
276
+ regardless of how many rounds it runs or how confident it sounds.
277
+
278
+ ## Anti-patterns to refuse in review
279
+
280
+ - **"The loop decides when it's good enough."** Good-enough is Type-B;
281
+ the loop may decide when it's *done running*, not when it's *good*.
282
+ - **"We re-review until it passes."** Fine — but *who* says it passed? If
283
+ the answer is Claude, the loop is self-acquitting.
284
+ - **"N agreeing agents = consensus."** Same-family agreement is correlated
285
+ blindness, not a jury (see fan-out section).
286
+ - **"It converged, so it's correct."** Convergence is Type-A
287
+ (it stopped moving); correctness/sufficiency is Type-B.
288
+ - **"Score >= 6, so stop."** Only safe if a *different family* assigned
289
+ the score. Claude scoring Claude and stopping at 6 is self-acquittal.
290
+ - **`/loop` wrapping an internal semantic loop.** External cadence
291
+ (`/loop`) is additive only for external-world waits (GPU done?
292
+ overnight heartbeat?). Wrapping ARIS's internal semantic loops with a
293
+ timer breaks `threadId` continuity and re-runs Type-B verdicts on a
294
+ clock instead of on the reviewer's turn — noise at best, a corrupted
295
+ acquittal at worst. Keep external cadence outside the acceptance gate.
296
+
297
+ ## Epistemic status of a PASS
298
+
299
+ A cross-model PASS is a **heterogeneous second opinion**, not external ground truth. Its
300
+ value is specific and bounded: a reviewer from a different model family breaks *correlated*
301
+ blind spots — the executor's own failure modes it cannot see in itself — so a PASS means
302
+ "a differently-built model, reading the artifact cold, did not find the flaw the author
303
+ would miss." It does **not** mean the work is correct, novel, publishable, or that a venue
304
+ will accept it. Same-family review (Claude judging Claude) does not even clear that bar,
305
+ which is why the jury must be cross-family.
306
+
307
+ Treat a PASS as the strongest *automatable* heterogeneous quality check this framework has, then keep the human in the loop for
308
+ what no in-framework verdict can supply: updated literature, venue taste, and ground truth.
309
+ A green gate lowers risk; it does not transfer accountability.
310
+
311
+ ## See Also
312
+
313
+ - `reviewer-independence.md` — the single-shot form: executor never
314
+ filters the reviewer's inputs. Type-B gates inherit this in full.
315
+ - `experiment-integrity.md` — the experiment form: the model that writes
316
+ experiment code must not judge its integrity. `/experiment-audit`'s
317
+ Type-B verdict is the loop instance of this rule.
318
+ - `reviewer-routing.md` — where Type-B gates send their verdict (codex
319
+ default, oracle-pro on request, manual only with a verified non-Claude
320
+ target).
321
+ - `fan-out-pattern.md` — breadth via same-family spawn; this doc's
322
+ fan-out section is the guardrail that keeps breadth out of the jury box.
323
+ - `integration-contract.md` §3 — the cross-model verdict must leave an
324
+ inspectable artifact.
@@ -0,0 +1,248 @@
1
+ # Assurance Contract
2
+
3
+ ARIS audits emit machine-readable verdicts. The `assurance` axis decides whether
4
+ those verdicts are advisory (draft mode) or load-bearing gates (submission mode).
5
+ This contract is referenced by `paper-writing`, `paper-claim-audit`, `citation-audit`,
6
+ `proof-checker`, and the external verifier (canonical name `verify_paper_audits.sh`;
7
+ callers resolve the actual path via `integration-contract.md` §2).
8
+
9
+ ## Why a separate axis from `effort`
10
+
11
+ Historically `effort` (lite/balanced/max/beast) was conflated with audit strictness.
12
+ The result: `effort: beast` did not guarantee mandatory audits ran — phases were
13
+ gated by content detectors (e.g. `if \begin{theorem} exists`) and could silently
14
+ skip. A user reported `effort: beast` produced a "draft-quality" paper with all
15
+ three submission-gate audits skipped.
16
+
17
+ The fix is to split the concerns:
18
+
19
+ | Axis | Controls | Default |
20
+ |------|----------|---------|
21
+ | `effort` | depth/cost (papers, rounds, ideation) | `balanced` |
22
+ | `assurance` | audit strictness — silent-skip-allowed vs verdict-required | derived from `effort` (see mapping) |
23
+
24
+ Override either independently: `— effort: balanced, assurance: submission` is
25
+ legal and means "normal depth, but every audit must emit a verdict before
26
+ finalization."
27
+
28
+ ## Assurance Levels
29
+
30
+ ### `draft` — current behavior, no breakage
31
+ - Audits run only if their content detector matches.
32
+ - Silent skip allowed.
33
+ - `paper-writing` Phase 6 produces a final report regardless.
34
+ - For: rapid iteration, exploratory drafts, early-stage research.
35
+
36
+ ### `submission` — load-bearing audits
37
+ - All mandatory audits **must** emit a verdict (one of the 6 below).
38
+ - Silent skip is **forbidden**.
39
+ - `paper-writing` Phase 6 invokes `verify_paper_audits.sh` (resolved per
40
+ `integration-contract.md` §2); non-zero exit blocks Final Report.
41
+ - The Final Report tags itself `submission-ready: yes/no` based on verifier output.
42
+ - For: conference / journal submission, anything you'd put your name on.
43
+
44
+ ## Default Mapping (derived if `assurance` not given)
45
+
46
+ | `effort` | implied `assurance` |
47
+ |----------|---------------------|
48
+ | `lite` | `draft` |
49
+ | `balanced` | `draft` |
50
+ | `max` | `submission` |
51
+ | `beast` | `submission` |
52
+
53
+ This means a user passing only `— effort: beast` automatically gets full audit
54
+ enforcement — matching their intent ("turn everything up"). Users wanting
55
+ strict audits at lower depth pass `— assurance: submission` explicitly.
56
+
57
+ ## Verdict State Machine
58
+
59
+ Every mandatory audit must emit exactly one of these — never silent skip:
60
+
61
+ | Verdict | Meaning | Audit ran? | Submission-blocking? |
62
+ |---------|---------|-----------|----------------------|
63
+ | `PASS` | All checks passed | Yes | No |
64
+ | `WARN` | Issues found, none disqualifying | Yes | No |
65
+ | `FAIL` | Disqualifying issues found | Yes | **Yes** |
66
+ | `NOT_APPLICABLE` | Detector negative; nothing to audit (e.g., no theorems in paper, no `\cite`s, no numeric claims) | Audit phase ran, child audit invocation may have been skipped | No |
67
+ | `BLOCKED` | Audit should apply but prerequisites are missing or unsupported (e.g., paper has numeric claims but no `results/` directory; paper cites references but `.bib` missing) | Could not complete | **Yes** |
68
+ | `ERROR` | Audit invocation failed (network, timeout, malformed reviewer output) | Attempted but errored | **Yes** at submission |
69
+
70
+ ### Why `NOT_APPLICABLE` is not the same as `SKIP`
71
+
72
+ `NOT_APPLICABLE` means **the audit phase ran**, the detector returned negative,
73
+ and a verdict artifact was written documenting "we checked, there's nothing to
74
+ verify." This is verifiable from outside the LLM — the artifact file exists.
75
+
76
+ A silent skip leaves no record. There's no way to distinguish "we checked and
77
+ there was nothing" from "we forgot." This contract makes that distinction
78
+ mandatory.
79
+
80
+ ### Why `BLOCKED` is more dangerous than `NOT_APPLICABLE`
81
+
82
+ `BLOCKED` means the audit *should* have run but cannot. Example: a paper claims
83
+ `accuracy = 89.2%` but has no `results/` directory to verify against. That's not
84
+ "nothing to audit" — that's "we cannot verify a load-bearing claim." Treating
85
+ this as `SKIP` masks the danger; `BLOCKED` surfaces it and blocks submission.
86
+
87
+ ## Required Audit Artifact Schema
88
+
89
+ Every mandatory audit must write a JSON artifact (and may also write a
90
+ human-readable Markdown sibling). The JSON must contain at minimum:
91
+
92
+ ```json
93
+ {
94
+ "audit_skill": "paper-claim-audit", // citation-audit, proof-checker, etc.
95
+ "verdict": "PASS", // one of the 6 above
96
+ "reason_code": "all_numbers_match", // skill-specific short string
97
+ "summary": "Verified 23 numeric claims against 4 result files; no mismatches.",
98
+ "audited_input_hashes": {
99
+ "main.tex": "sha256:a3f8...",
100
+ "sections/5.evidence.tex": "sha256:b2d1...",
101
+ "/Users/me/project/results/run_2026_04_19.json": "sha256:c9e4..."
102
+ },
103
+ "trace_path": ".aris/traces/paper-claim-audit/2026-04-21_run01/",
104
+ "thread_id": "019dae73-fc12-4ab8-...",
105
+ "executor_model": "claude-opus-4-8",
106
+ "executor_family": "anthropic",
107
+ "reviewer_model": "gpt-5.6-sol",
108
+ "reviewer_family": "openai",
109
+ "review_independence": "cross-family",
110
+ "acceptance_status": "accepted",
111
+ "reviewer_reasoning": "xhigh",
112
+ "generated_at": "2026-04-21T14:23:01Z",
113
+ "details": {
114
+ // skill-specific structured data
115
+ }
116
+ }
117
+ ```
118
+
119
+ Field semantics:
120
+
121
+ - **`audited_input_hashes`** — SHA256 of every file the audit consumed.
122
+ - Keys are **paths relative to the paper directory** (the argument
123
+ passed to `verify_paper_audits.sh`) for files inside it, or
124
+ **absolute paths** for files outside it (e.g. `../results/run.json`
125
+ is legal but `/Users/me/project/results/run.json` is more portable).
126
+ Do NOT prefix in-paper files with `paper/` — the verifier already
127
+ resolves relative to the paper dir and `paper/paper/main.tex` will
128
+ false-fail. The verifier rehashes the current files and flags `STALE`
129
+ if any hash changed since the audit ran. (User edited `main.tex`
130
+ after running `paper-claim-audit`? The next verifier run will catch it.)
131
+ - **`trace_path`** — directory containing the full reviewer prompt + response
132
+ pair, per `review-tracing.md`. Required for mandatory audits — not optional.
133
+ - **`thread_id` or `agent_id`** — durable reviewer handle. MCP routes use a
134
+ thread ID; Codex `spawn_agent` routes use an agent ID. At least one is required.
135
+ - **`reviewer_model`** + **`reviewer_reasoning`** — proves cross-family review
136
+ invariant was honored.
137
+ - **`review_independence`** — `same-family`, `cross-family`, or `deterministic`.
138
+ `deterministic` is valid ONLY for what a process can actually decide
139
+ (compilation, schema validity, hash freshness, test suites) — the audit
140
+ aggregator REJECTS a deterministic label on the four semantic paper audits
141
+ (proof / claims / citations / attack), which only a cross-family model
142
+ review can accept.
143
+ **`acceptance_status`** is `provisional` for same-family review and
144
+ `accepted` for cross-family/deterministic review; neither field rewrites the
145
+ substantive verdict.
146
+ - **`generated_at`** — UTC ISO-8601 timestamp.
147
+
148
+ ## Verifier Contract
149
+
150
+ `verify_paper_audits.sh <paper-dir>` (canonical name; resolved per
151
+ `integration-contract.md` §2) is the single source of truth for
152
+ "are mandatory audits complete and current?" It must:
153
+
154
+ 1. Locate the paper-writing manifest (which mandatory audits applied this run).
155
+ 2. For each, check artifact JSON exists at expected path.
156
+ 3. Validate artifact JSON against required-fields schema (above).
157
+ 4. Verify `verdict` is one of the 6 allowed values.
158
+ 5. Recompute SHA256 of every file in `audited_input_hashes`; flag `STALE` if any
159
+ mismatches.
160
+ 6. Verify `trace_path` exists and is non-empty.
161
+ 7. Output a structured JSON report and exit 0 (all green) or 1 (any FAIL /
162
+ BLOCKED / ERROR / STALE / missing artifact).
163
+
164
+ The report also emits `overall_assurance`: `blocked` for a blocking condition,
165
+ `provisional` when green artifacts include same-family or legacy-unspecified
166
+ review, and `accepted` only when all green artifacts are cross-family or
167
+ deterministic. Provisional remains exit 0 but must never be presented as
168
+ submission-ready yes.
169
+
170
+ Phase 6 of `paper-writing` invokes the verifier; at `assurance: submission`,
171
+ non-zero exit blocks Final Report generation.
172
+
173
+ ## Subskill Contract: "Always Emit, Never Block"
174
+
175
+ Child audit skills (`paper-claim-audit`, `citation-audit`, `proof-checker`)
176
+ follow this contract:
177
+
178
+ - **Always emit a verdict artifact**, even on detector-negative or error paths.
179
+ - **Never block** the parent's flow themselves — they only emit verdicts.
180
+ - **The parent skill** (`paper-writing` Phase 6 + verifier) decides whether a
181
+ given verdict blocks finalization. This decision lives in *one* place
182
+ (`assurance` axis + verifier), not duplicated across child skills.
183
+
184
+ Earlier wording in `paper-claim-audit` and `citation-audit` (e.g., "audit is
185
+ advisory, never blocking") referred to this division of labor — but conflicted
186
+ with `paper-writing`'s declaration that they were "mandatory submission gates."
187
+ This contract resolves the conflict: child = always emit; parent = decides
188
+ blocking based on assurance level.
189
+
190
+ ## Examples
191
+
192
+ ### Theory paper, beast effort
193
+ ```
194
+ — effort: beast (implies assurance: submission)
195
+ ```
196
+ - `proof-checker` runs, audits theorems → `PASS` or `WARN` or `FAIL`
197
+ - `paper-claim-audit` runs, finds numbers → `PASS`
198
+ - `citation-audit` runs, audits refs → `PASS`
199
+ - Verifier: all green
200
+ - Final Report: `submission-ready: yes`
201
+
202
+ ### Position paper (no theorems, no numbers, no experiments), beast effort
203
+ ```
204
+ — effort: beast (implies assurance: submission)
205
+ ```
206
+ - `proof-checker` invoked → no theorems found → emits `NOT_APPLICABLE`
207
+ - `paper-claim-audit` invoked → no numeric claims → emits `NOT_APPLICABLE`
208
+ - `citation-audit` invoked → audits refs → `PASS`
209
+ - Verifier: all green (NOT_APPLICABLE is not blocking)
210
+ - Final Report: `submission-ready: yes` with note "no theorems / no numeric claims to audit"
211
+
212
+ ### Empirical paper missing raw results, beast effort
213
+ ```
214
+ — effort: beast
215
+ ```
216
+ - `proof-checker` → `NOT_APPLICABLE`
217
+ - `paper-claim-audit` invoked → finds claims like `accuracy = 89.2%` but
218
+ `results/` is empty → emits `BLOCKED` with reason_code `no_raw_evidence`
219
+ - `citation-audit` → `PASS`
220
+ - Verifier: exit 1 (BLOCKED is submission-blocking)
221
+ - Final Report: **refuses to finalize**; surfaces "Mandatory audit BLOCKED:
222
+ paper-claim-audit cannot verify numeric claims — no raw result files found.
223
+ Add results/ or downgrade to `— assurance: draft`."
224
+
225
+ ### Stale audit (user edited paper after running audits)
226
+ - User runs `/paper-writing` at beast → all audits PASS, files written
227
+ - User edits `sec/5.evidence.tex` to change a number
228
+ - User reruns the verifier (or re-finalizes)
229
+ - Verifier rehashes → `audited_input_hashes` mismatch → `STALE` flag → exit 1
230
+ - Final Report: refuses; instructs user to rerun `paper-claim-audit` and
231
+ `citation-audit` before re-finalizing.
232
+
233
+ ## Backward Compatibility
234
+
235
+ - Users on `effort: balanced` (default) get `assurance: draft` — **identical
236
+ current behavior, no breakage**.
237
+ - Users explicitly using `effort: max` or `effort: beast` automatically get
238
+ `assurance: submission` — matching their intent.
239
+ - Users wanting the old "beast = depth only, no audit enforcement" can pass
240
+ `— effort: beast, assurance: draft` (explicit override). This combination is
241
+ legal but discouraged for actual submissions.
242
+
243
+ ## See Also
244
+
245
+ - `effort-contract.md` — depth/cost axis (separate concern)
246
+ - `review-tracing.md` — trace artifact protocol (referenced by `trace_path`)
247
+ - `reviewer-independence.md` — cross-model review invariant
248
+ - `tools/verify_paper_audits.sh` — external verifier implementation
@@ -0,0 +1,78 @@
1
+ # Capture Anti-patterns (anti-self-poisoning)
2
+
3
+ When ARIS captures *durable* knowledge — a research-wiki idea / claim / experiment
4
+ node, a `/meta-optimize` SKILL.md proposal — it must not store **operational
5
+ noise** that later hardens into a self-cited falsehood. This is the failure mode
6
+ Hermes's self-improvement loop hit and patched with a hand-written "Do NOT
7
+ capture" list: negative tool-capability claims that *"harden into refusals the
8
+ agent cites against itself for months after the actual problem was fixed."*
9
+ ARIS's research-wiki "failed ideas → anti-repeat memory" is the GOOD inverse (a
10
+ class-level *research* finding worth remembering); this is the blocklist for the
11
+ BAD kind (transient *operational* state masquerading as a durable fact).
12
+
13
+ ## The four anti-patterns — do NOT capture
14
+
15
+ | class | example (do NOT store) | store INSTEAD |
16
+ |-------|------------------------|---------------|
17
+ | **env-specific failure** | "pip failed: No module named torch", "command not found" | the fix / the missing dependency / the correct config |
18
+ | **transient error** | "got a 429", "CUDA OOM", "connection refused" | nothing — it self-resolves; or the retry/backoff that worked |
19
+ | **negative tool-capability claim** | "codex can't handle long files", "gemini is broken", "don't use oracle" | the workaround, or "needs flag X" — never "tool can't do Y" |
20
+ | **single-instance narrative** | "in run 47 the loss spiked at step 300" | only the *class-level* rule it implies, if any ("LR > 3e-4 diverges on this model") |
21
+
22
+ The cardinal rule: **store *how to fix* / *what config is missing* / *the
23
+ workaround*, never *"X can't do Y"*.** A negative capability claim about your own
24
+ tooling is the most dangerous capture — it gets loaded into every future session
25
+ and the agent cites it against itself long after the real cause is gone.
26
+
27
+ ## Mechanical vs judgment
28
+
29
+ - **Mechanical** (deterministic, `tools/capture_filter.py`): the unambiguous
30
+ classes — raw error output (`No module named`, `command not found`,
31
+ `ModuleNotFoundError`, `Permission denied`), transient errors (rate-limit / OOM
32
+ / network), and explicitly-broken-tool phrasing anchored on ARIS infrastructure
33
+ nouns (codex / gemini / oracle / the reviewer / the MCP / the CLI …).
34
+ - **Judgment** (this doc): the single-instance-narrative class, and any operational
35
+ note dressed up as a finding. The agent applies this when deciding what to persist.
36
+
37
+ The mechanical filter is **deliberately conservative**: it does NOT flag
38
+ legitimate *research* findings about a model/method ("the model can't generalize
39
+ to OOD", "our method fails on long sequences") — it targets ARIS's own *tooling*
40
+ being declared broken, and raw error text. False negatives are fine (the jury
41
+ still judges); a flagged note just goes to manual review / gets rewritten.
42
+
43
+ ## The asymmetry (acceptance-gate.md)
44
+
45
+ This filter may **REJECT a capture same-model** — it is a mechanical safety screen,
46
+ low risk, and same-model is always allowed to *reject*. But anything that **passes**
47
+ the filter and would become a **load-bearing** skill/claim still goes to the
48
+ **cross-model jury** before it is trusted. Same-model is fine to reject; it is
49
+ never enough to *accept* into the load-bearing set.
50
+
51
+ ## Helper
52
+
53
+ ```
54
+ from capture_filter import screen, reason_detail
55
+ screen(text) # -> [reason, ...] ([] = clean); reason ∈ {env_failure, transient_error, negative_tool_claim}
56
+ ```
57
+ ```
58
+ python3 tools/capture_filter.py <file|-> # exit 1 + reasons if anti-pattern found
59
+ ```
60
+
61
+ ## Where ARIS uses it
62
+ - **`/research-wiki`** (and `/idea-creator` Phase-3 annotations): screen an
63
+ idea/claim/experiment note before persisting it; if flagged, rewrite to the
64
+ fix or drop it — don't let operational noise become a durable node.
65
+ - **`/meta-optimize`**: screen the rationale of a proposed SKILL.md change; never
66
+ propose a change that encodes a negative tool-capability claim or a one-off
67
+ failure as a durable rule.
68
+
69
+ ## Cross-references
70
+ - `acceptance-gate.md` — the reject/accept asymmetry: same-model may reject, only
71
+ cross-model may accept into the load-bearing set.
72
+ - `evidence-precheck.md` / `injection-hygiene.md` — sibling deterministic
73
+ pre-gates feeding the cross-model jury.
74
+
75
+ > Anti-pattern taxonomy adapted from NousResearch/hermes-agent's background-review
76
+ > "Do NOT capture" list (MIT). ARIS's increment: Hermes patches self-poisoning with
77
+ > more self-judged prose; ARIS adds the deterministic screen + the cross-model
78
+ > acceptance gate on anything that survives it.