dsh-aris-panel 0.1.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (442) hide show
  1. package/LICENSE +21 -0
  2. package/README.md +98 -0
  3. package/README_CN.md +87 -0
  4. package/dsh/checkout.patch.yml +38 -0
  5. package/dsh/client.js +634 -0
  6. package/dsh/cordis.patch.yml +44 -0
  7. package/dsh/index.mjs +76 -0
  8. package/dsh/run-status.mjs +182 -0
  9. package/dsh/scope-limits.mjs +50 -0
  10. package/dsh/workbench.mjs +291 -0
  11. package/mcp-servers/claude-review/README.md +93 -0
  12. package/mcp-servers/claude-review/run_with_claude_aws.sh +49 -0
  13. package/mcp-servers/claude-review/server.py +718 -0
  14. package/mcp-servers/codex-image2/README.md +65 -0
  15. package/mcp-servers/codex-image2/server.py +893 -0
  16. package/mcp-servers/feishu-bridge/requirements.txt +1 -0
  17. package/mcp-servers/feishu-bridge/server.py +240 -0
  18. package/mcp-servers/gemini-review/README.md +171 -0
  19. package/mcp-servers/gemini-review/server.py +1856 -0
  20. package/mcp-servers/llm-chat/requirements.txt +1 -0
  21. package/mcp-servers/llm-chat/server.py +664 -0
  22. package/mcp-servers/manual-review/README.md +133 -0
  23. package/mcp-servers/manual-review/server.py +910 -0
  24. package/mcp-servers/manual-review/ui.html +279 -0
  25. package/mcp-servers/minimax-chat/requirements.txt +1 -0
  26. package/mcp-servers/minimax-chat/server.py +381 -0
  27. package/package.json +51 -0
  28. package/skills/ablation-planner/SKILL.md +123 -0
  29. package/skills/alphaxiv/SKILL.md +196 -0
  30. package/skills/analyze-results/SKILL.md +46 -0
  31. package/skills/arxiv/SKILL.md +248 -0
  32. package/skills/auto-paper-improvement-loop/SKILL.md +651 -0
  33. package/skills/auto-review-loop/SKILL.md +1137 -0
  34. package/skills/auto-review-loop-llm/SKILL.md +259 -0
  35. package/skills/auto-review-loop-minimax/SKILL.md +302 -0
  36. package/skills/citation-audit/SKILL.md +502 -0
  37. package/skills/claims-drafting/SKILL.md +227 -0
  38. package/skills/comm-lit-review/SKILL.md +297 -0
  39. package/skills/deepxiv/SKILL.md +263 -0
  40. package/skills/dse-loop/SKILL.md +296 -0
  41. package/skills/embodiment-description/SKILL.md +129 -0
  42. package/skills/exa-search/SKILL.md +205 -0
  43. package/skills/experiment-audit/SKILL.md +311 -0
  44. package/skills/experiment-bridge/SKILL.md +376 -0
  45. package/skills/experiment-plan/SKILL.md +249 -0
  46. package/skills/experiment-queue/SKILL.md +431 -0
  47. package/skills/experiment-queue/scripts/build_manifest.py +142 -0
  48. package/skills/experiment-queue/scripts/queue_manager.py +433 -0
  49. package/skills/feishu-notify/SKILL.md +156 -0
  50. package/skills/figure-description/SKILL.md +138 -0
  51. package/skills/figure-spec/SKILL.md +262 -0
  52. package/skills/figure-spec/scripts/figure_renderer.py +799 -0
  53. package/skills/formula-derivation/SKILL.md +280 -0
  54. package/skills/gemini-search/SKILL.md +231 -0
  55. package/skills/grant-proposal/SKILL.md +698 -0
  56. package/skills/idea-creator/SKILL.md +542 -0
  57. package/skills/idea-discovery/SKILL.md +521 -0
  58. package/skills/idea-discovery-robot/SKILL.md +363 -0
  59. package/skills/integrity-forensics/SKILL.md +284 -0
  60. package/skills/interview-cheatsheet/SKILL.md +245 -0
  61. package/skills/invention-structuring/SKILL.md +188 -0
  62. package/skills/jurisdiction-format/SKILL.md +192 -0
  63. package/skills/kill-argument/SKILL.md +437 -0
  64. package/skills/mermaid-diagram/SKILL.md +419 -0
  65. package/skills/meta-apply/SKILL.md +141 -0
  66. package/skills/meta-optimize/SKILL.md +437 -0
  67. package/skills/monitor-experiment/SKILL.md +140 -0
  68. package/skills/novelty-check/SKILL.md +101 -0
  69. package/skills/openalex/SKILL.md +237 -0
  70. package/skills/overleaf-sync/SKILL.md +220 -0
  71. package/skills/paper-claim-audit/SKILL.md +348 -0
  72. package/skills/paper-compile/SKILL.md +266 -0
  73. package/skills/paper-figure/SKILL.md +312 -0
  74. package/skills/paper-illustration/SKILL.md +736 -0
  75. package/skills/paper-illustration-image2/SKILL.md +391 -0
  76. package/skills/paper-illustration-image2/scripts/paper_illustration_image2.py +255 -0
  77. package/skills/paper-plan/SKILL.md +386 -0
  78. package/skills/paper-poster/SKILL.md +19 -0
  79. package/skills/paper-poster-html/DESIGN_FINAL.md +176 -0
  80. package/skills/paper-poster-html/IMPLEMENTATION_CONVENTIONS.md +161 -0
  81. package/skills/paper-poster-html/LICENSES/posterly-MIT.txt +21 -0
  82. package/skills/paper-poster-html/NOTICE.md +57 -0
  83. package/skills/paper-poster-html/SKILL.md +323 -0
  84. package/skills/paper-poster-html/scripts/_posterly/__init__.py +0 -0
  85. package/skills/paper-poster-html/scripts/_posterly/canvas.py +200 -0
  86. package/skills/paper-poster-html/scripts/_posterly/measure.py +588 -0
  87. package/skills/paper-poster-html/scripts/_posterly/polish.py +498 -0
  88. package/skills/paper-poster-html/scripts/_posterly/preflight.py +489 -0
  89. package/skills/paper-poster-html/scripts/_posterly/render.py +215 -0
  90. package/skills/paper-poster-html/scripts/_posterly/textutil.py +16 -0
  91. package/skills/paper-poster-html/scripts/_posterly/verify_final.py +171 -0
  92. package/skills/paper-poster-html/scripts/asset_check.py +897 -0
  93. package/skills/paper-poster-html/scripts/extract_pdf_figures.py +666 -0
  94. package/skills/paper-poster-html/scripts/poster_check.py +251 -0
  95. package/skills/paper-poster-html/scripts/preprocess_figures.py +238 -0
  96. package/skills/paper-poster-html/scripts/render_preview.py +217 -0
  97. package/skills/paper-poster-html/scripts/run_gates.py +556 -0
  98. package/skills/paper-poster-html/scripts/style_check.py +1324 -0
  99. package/skills/paper-poster-html/templates/COMPONENTS.md +462 -0
  100. package/skills/paper-poster-html/templates/README.md +170 -0
  101. package/skills/paper-poster-html/templates/landscape_4col.html +1032 -0
  102. package/skills/paper-poster-html/templates/landscape_hero.html +1046 -0
  103. package/skills/paper-poster-html/templates/portrait_2col.html +947 -0
  104. package/skills/paper-poster-html/templates/tokens/acl.json +9 -0
  105. package/skills/paper-poster-html/templates/tokens/cvpr.json +9 -0
  106. package/skills/paper-poster-html/templates/tokens/generic.json +9 -0
  107. package/skills/paper-poster-html/templates/tokens/iclr.json +9 -0
  108. package/skills/paper-poster-html/templates/tokens/icml.json +9 -0
  109. package/skills/paper-poster-html/templates/tokens/neurips.json +9 -0
  110. package/skills/paper-slides/SKILL.md +635 -0
  111. package/skills/paper-talk/SKILL.md +381 -0
  112. package/skills/paper-write/SKILL.md +604 -0
  113. package/skills/paper-write/templates/IEEEtran.bst +2409 -0
  114. package/skills/paper-write/templates/IEEEtran.cls +6347 -0
  115. package/skills/paper-write/templates/iclr2026.tex +84 -0
  116. package/skills/paper-write/templates/icml2025.tex +87 -0
  117. package/skills/paper-write/templates/ieee_conference.tex +89 -0
  118. package/skills/paper-write/templates/ieee_journal.tex +93 -0
  119. package/skills/paper-write/templates/math_commands.tex +48 -0
  120. package/skills/paper-write/templates/neurips2025.tex +80 -0
  121. package/skills/paper-writing/SKILL.md +916 -0
  122. package/skills/patent-novelty-check/SKILL.md +153 -0
  123. package/skills/patent-pipeline/SKILL.md +344 -0
  124. package/skills/patent-review/SKILL.md +203 -0
  125. package/skills/pixel-art/SKILL.md +137 -0
  126. package/skills/prior-art-search/SKILL.md +146 -0
  127. package/skills/proof-checker/SKILL.md +866 -0
  128. package/skills/proof-orchestrator/NOTICE.md +24 -0
  129. package/skills/proof-orchestrator/SKILL.md +254 -0
  130. package/skills/proof-orchestrator/references/audit-output-contract.md +126 -0
  131. package/skills/proof-orchestrator/references/deepseek-routing.md +74 -0
  132. package/skills/proof-orchestrator/references/dispatch-prompts.md +227 -0
  133. package/skills/proof-orchestrator/references/notation-audit.md +135 -0
  134. package/skills/proof-orchestrator/references/proof-audit-rubric.md +70 -0
  135. package/skills/proof-orchestrator/references/stress-tests.md +38 -0
  136. package/skills/proof-writer/SKILL.md +223 -0
  137. package/skills/qzcli/SKILL.md +324 -0
  138. package/skills/rebuttal/SKILL.md +376 -0
  139. package/skills/render-html/SKILL.md +316 -0
  140. package/skills/render-html/scripts/render_html.py +1006 -0
  141. package/skills/render-html/scripts/templates/academic.html +703 -0
  142. package/skills/render-html/scripts/templates/dashboard.html +333 -0
  143. package/skills/research-lit/SKILL.md +756 -0
  144. package/skills/research-pipeline/SKILL.md +384 -0
  145. package/skills/research-refine/SKILL.md +770 -0
  146. package/skills/research-refine-pipeline/SKILL.md +186 -0
  147. package/skills/research-review/SKILL.md +198 -0
  148. package/skills/research-wiki/SKILL.md +461 -0
  149. package/skills/resubmit-pipeline/SKILL.md +447 -0
  150. package/skills/result-to-claim/SKILL.md +311 -0
  151. package/skills/run-experiment/SKILL.md +313 -0
  152. package/skills/semantic-scholar/SKILL.md +236 -0
  153. package/skills/serverless-modal/SKILL.md +335 -0
  154. package/skills/shared-references/acceptance-gate.md +324 -0
  155. package/skills/shared-references/assurance-contract.md +248 -0
  156. package/skills/shared-references/capture-antipatterns.md +78 -0
  157. package/skills/shared-references/citation-discipline.md +583 -0
  158. package/skills/shared-references/compute-env-contract.md +163 -0
  159. package/skills/shared-references/effort-contract.md +183 -0
  160. package/skills/shared-references/evidence-precheck.md +65 -0
  161. package/skills/shared-references/experiment-integrity.md +49 -0
  162. package/skills/shared-references/external-cadence.md +326 -0
  163. package/skills/shared-references/fan-out-pattern.md +366 -0
  164. package/skills/shared-references/injection-hygiene.md +127 -0
  165. package/skills/shared-references/integration-contract.md +461 -0
  166. package/skills/shared-references/output-composition.md +93 -0
  167. package/skills/shared-references/output-language.md +45 -0
  168. package/skills/shared-references/output-manifest.md +49 -0
  169. package/skills/shared-references/output-versioning.md +111 -0
  170. package/skills/shared-references/patent-format-cn.md +199 -0
  171. package/skills/shared-references/patent-format-ep.md +173 -0
  172. package/skills/shared-references/patent-format-us.md +161 -0
  173. package/skills/shared-references/patent-writing-principles.md +197 -0
  174. package/skills/shared-references/prior-art-databases.md +141 -0
  175. package/skills/shared-references/resumable-runs.md +109 -0
  176. package/skills/shared-references/review-scope-limits.md +81 -0
  177. package/skills/shared-references/review-tracing.md +391 -0
  178. package/skills/shared-references/reviewer-independence.md +79 -0
  179. package/skills/shared-references/reviewer-routing.md +852 -0
  180. package/skills/shared-references/skill-governance.md +104 -0
  181. package/skills/shared-references/taste-calibration.md +85 -0
  182. package/skills/shared-references/venue-checklists.md +114 -0
  183. package/skills/shared-references/wiki-helper-resolution.md +134 -0
  184. package/skills/shared-references/writing-principles.md +525 -0
  185. package/skills/skills-codex/README.md +102 -0
  186. package/skills/skills-codex/README_CN.md +100 -0
  187. package/skills/skills-codex/ablation-planner/SKILL.md +126 -0
  188. package/skills/skills-codex/alphaxiv/SKILL.md +186 -0
  189. package/skills/skills-codex/analyze-results/SKILL.md +45 -0
  190. package/skills/skills-codex/arxiv/SKILL.md +210 -0
  191. package/skills/skills-codex/auto-paper-improvement-loop/SKILL.md +574 -0
  192. package/skills/skills-codex/auto-review-loop/SKILL.md +500 -0
  193. package/skills/skills-codex/auto-review-loop-llm/SKILL.md +247 -0
  194. package/skills/skills-codex/auto-review-loop-minimax/SKILL.md +290 -0
  195. package/skills/skills-codex/citation-audit/SKILL.md +504 -0
  196. package/skills/skills-codex/claims-drafting/SKILL.md +239 -0
  197. package/skills/skills-codex/comm-lit-review/SKILL.md +299 -0
  198. package/skills/skills-codex/comm-lit-review/references/domain-taxonomy.md +57 -0
  199. package/skills/skills-codex/comm-lit-review/references/output-template.md +37 -0
  200. package/skills/skills-codex/comm-lit-review/references/source-policy.md +99 -0
  201. package/skills/skills-codex/comm-lit-review/references/venue-tiering.md +112 -0
  202. package/skills/skills-codex/deepxiv/SKILL.md +142 -0
  203. package/skills/skills-codex/dse-loop/SKILL.md +285 -0
  204. package/skills/skills-codex/embodiment-description/SKILL.md +129 -0
  205. package/skills/skills-codex/exa-search/SKILL.md +192 -0
  206. package/skills/skills-codex/experiment-audit/SKILL.md +286 -0
  207. package/skills/skills-codex/experiment-bridge/SKILL.md +356 -0
  208. package/skills/skills-codex/experiment-plan/SKILL.md +249 -0
  209. package/skills/skills-codex/experiment-queue/SKILL.md +401 -0
  210. package/skills/skills-codex/feishu-notify/SKILL.md +155 -0
  211. package/skills/skills-codex/figure-description/SKILL.md +138 -0
  212. package/skills/skills-codex/figure-spec/SKILL.md +252 -0
  213. package/skills/skills-codex/formula-derivation/SKILL.md +280 -0
  214. package/skills/skills-codex/gemini-search/SKILL.md +205 -0
  215. package/skills/skills-codex/grant-proposal/SKILL.md +626 -0
  216. package/skills/skills-codex/idea-creator/SKILL.md +405 -0
  217. package/skills/skills-codex/idea-discovery/SKILL.md +475 -0
  218. package/skills/skills-codex/idea-discovery-robot/SKILL.md +362 -0
  219. package/skills/skills-codex/integrity-forensics/SKILL.md +106 -0
  220. package/skills/skills-codex/interview-cheatsheet/SKILL.md +245 -0
  221. package/skills/skills-codex/invention-structuring/SKILL.md +188 -0
  222. package/skills/skills-codex/jurisdiction-format/SKILL.md +192 -0
  223. package/skills/skills-codex/kill-argument/SKILL.md +403 -0
  224. package/skills/skills-codex/mermaid-diagram/SKILL.md +379 -0
  225. package/skills/skills-codex/meta-apply/SKILL.md +154 -0
  226. package/skills/skills-codex/meta-optimize/SKILL.md +348 -0
  227. package/skills/skills-codex/monitor-experiment/SKILL.md +98 -0
  228. package/skills/skills-codex/novelty-check/SKILL.md +89 -0
  229. package/skills/skills-codex/openalex/SKILL.md +228 -0
  230. package/skills/skills-codex/overleaf-sync/SKILL.md +220 -0
  231. package/skills/skills-codex/paper-claim-audit/SKILL.md +350 -0
  232. package/skills/skills-codex/paper-compile/SKILL.md +253 -0
  233. package/skills/skills-codex/paper-figure/SKILL.md +311 -0
  234. package/skills/skills-codex/paper-illustration/SKILL.md +690 -0
  235. package/skills/skills-codex/paper-illustration-image2/SKILL.md +383 -0
  236. package/skills/skills-codex/paper-illustration-image2/scripts/paper_illustration_image2.py +255 -0
  237. package/skills/skills-codex/paper-plan/SKILL.md +278 -0
  238. package/skills/skills-codex/paper-poster/SKILL.md +19 -0
  239. package/skills/skills-codex/paper-poster-html/SKILL.md +377 -0
  240. package/skills/skills-codex/paper-slides/SKILL.md +571 -0
  241. package/skills/skills-codex/paper-talk/SKILL.md +381 -0
  242. package/skills/skills-codex/paper-write/SKILL.md +411 -0
  243. package/skills/skills-codex/paper-write/templates/IEEEtran.bst +2409 -0
  244. package/skills/skills-codex/paper-write/templates/IEEEtran.cls +6347 -0
  245. package/skills/skills-codex/paper-write/templates/aaai2026.bst +1493 -0
  246. package/skills/skills-codex/paper-write/templates/aaai2026.sty +315 -0
  247. package/skills/skills-codex/paper-write/templates/aaai2026.tex +952 -0
  248. package/skills/skills-codex/paper-write/templates/acl.sty +312 -0
  249. package/skills/skills-codex/paper-write/templates/acl2026.tex +377 -0
  250. package/skills/skills-codex/paper-write/templates/acl_natbib.bst +1940 -0
  251. package/skills/skills-codex/paper-write/templates/acm.bst +3081 -0
  252. package/skills/skills-codex/paper-write/templates/acm_mm2026.tex +204 -0
  253. package/skills/skills-codex/paper-write/templates/acmart.cls +3520 -0
  254. package/skills/skills-codex/paper-write/templates/cvpr.bst +1448 -0
  255. package/skills/skills-codex/paper-write/templates/cvpr.sty +508 -0
  256. package/skills/skills-codex/paper-write/templates/cvpr2026.tex +63 -0
  257. package/skills/skills-codex/paper-write/templates/iclr2026.tex +84 -0
  258. package/skills/skills-codex/paper-write/templates/iclr2026_conference.bst +1440 -0
  259. package/skills/skills-codex/paper-write/templates/iclr2026_conference.sty +246 -0
  260. package/skills/skills-codex/paper-write/templates/icml2026.sty +767 -0
  261. package/skills/skills-codex/paper-write/templates/icml2026.tex +662 -0
  262. package/skills/skills-codex/paper-write/templates/ieee_conference.tex +89 -0
  263. package/skills/skills-codex/paper-write/templates/ieee_journal.tex +93 -0
  264. package/skills/skills-codex/paper-write/templates/math_commands.tex +48 -0
  265. package/skills/skills-codex/paper-write/templates/neurips2026.tex +493 -0
  266. package/skills/skills-codex/paper-write/templates/neurips_2026.sty +437 -0
  267. package/skills/skills-codex/paper-writing/SKILL.md +731 -0
  268. package/skills/skills-codex/patent-novelty-check/SKILL.md +153 -0
  269. package/skills/skills-codex/patent-pipeline/SKILL.md +344 -0
  270. package/skills/skills-codex/patent-review/SKILL.md +202 -0
  271. package/skills/skills-codex/pixel-art/SKILL.md +139 -0
  272. package/skills/skills-codex/prior-art-search/SKILL.md +146 -0
  273. package/skills/skills-codex/proof-checker/SKILL.md +554 -0
  274. package/skills/skills-codex/proof-orchestrator/SKILL.md +260 -0
  275. package/skills/skills-codex/proof-orchestrator/references/audit-output-contract.md +126 -0
  276. package/skills/skills-codex/proof-orchestrator/references/deepseek-routing.md +76 -0
  277. package/skills/skills-codex/proof-orchestrator/references/dispatch-prompts.md +227 -0
  278. package/skills/skills-codex/proof-orchestrator/references/notation-audit.md +135 -0
  279. package/skills/skills-codex/proof-orchestrator/references/proof-audit-rubric.md +70 -0
  280. package/skills/skills-codex/proof-orchestrator/references/stress-tests.md +38 -0
  281. package/skills/skills-codex/proof-writer/SKILL.md +222 -0
  282. package/skills/skills-codex/qzcli/SKILL.md +324 -0
  283. package/skills/skills-codex/rebuttal/SKILL.md +305 -0
  284. package/skills/skills-codex/render-html/SKILL.md +305 -0
  285. package/skills/skills-codex/render-html/scripts/__pycache__/render_html.cpython-314.pyc +0 -0
  286. package/skills/skills-codex/render-html/scripts/render_html.py +909 -0
  287. package/skills/skills-codex/render-html/scripts/templates/academic.html +342 -0
  288. package/skills/skills-codex/render-html/scripts/templates/dashboard.html +333 -0
  289. package/skills/skills-codex/research-lit/SKILL.md +464 -0
  290. package/skills/skills-codex/research-pipeline/SKILL.md +340 -0
  291. package/skills/skills-codex/research-refine/SKILL.md +721 -0
  292. package/skills/skills-codex/research-refine-pipeline/SKILL.md +186 -0
  293. package/skills/skills-codex/research-review/SKILL.md +135 -0
  294. package/skills/skills-codex/research-wiki/SKILL.md +421 -0
  295. package/skills/skills-codex/resubmit-pipeline/SKILL.md +444 -0
  296. package/skills/skills-codex/result-to-claim/SKILL.md +246 -0
  297. package/skills/skills-codex/run-experiment/SKILL.md +236 -0
  298. package/skills/skills-codex/semantic-scholar/SKILL.md +219 -0
  299. package/skills/skills-codex/serverless-modal/SKILL.md +335 -0
  300. package/skills/skills-codex/shared-references/acceptance-gate.md +336 -0
  301. package/skills/skills-codex/shared-references/assurance-contract.md +139 -0
  302. package/skills/skills-codex/shared-references/capture-antipatterns.md +84 -0
  303. package/skills/skills-codex/shared-references/citation-discipline.md +452 -0
  304. package/skills/skills-codex/shared-references/compute-env-contract.md +163 -0
  305. package/skills/skills-codex/shared-references/effort-contract.md +143 -0
  306. package/skills/skills-codex/shared-references/evidence-precheck.md +73 -0
  307. package/skills/skills-codex/shared-references/experiment-integrity.md +49 -0
  308. package/skills/skills-codex/shared-references/external-cadence.md +334 -0
  309. package/skills/skills-codex/shared-references/fan-out-pattern.md +375 -0
  310. package/skills/skills-codex/shared-references/injection-hygiene.md +132 -0
  311. package/skills/skills-codex/shared-references/integration-contract.md +372 -0
  312. package/skills/skills-codex/shared-references/output-composition.md +98 -0
  313. package/skills/skills-codex/shared-references/output-language.md +45 -0
  314. package/skills/skills-codex/shared-references/output-manifest.md +40 -0
  315. package/skills/skills-codex/shared-references/output-versioning.md +111 -0
  316. package/skills/skills-codex/shared-references/patent-format-cn.md +199 -0
  317. package/skills/skills-codex/shared-references/patent-format-ep.md +173 -0
  318. package/skills/skills-codex/shared-references/patent-format-us.md +161 -0
  319. package/skills/skills-codex/shared-references/patent-writing-principles.md +197 -0
  320. package/skills/skills-codex/shared-references/prior-art-databases.md +141 -0
  321. package/skills/skills-codex/shared-references/resumable-runs.md +125 -0
  322. package/skills/skills-codex/shared-references/review-scope-limits.md +81 -0
  323. package/skills/skills-codex/shared-references/review-tracing.md +144 -0
  324. package/skills/skills-codex/shared-references/reviewer-independence.md +66 -0
  325. package/skills/skills-codex/shared-references/reviewer-routing.md +128 -0
  326. package/skills/skills-codex/shared-references/skill-governance.md +119 -0
  327. package/skills/skills-codex/shared-references/taste-calibration.md +90 -0
  328. package/skills/skills-codex/shared-references/venue-checklists.md +73 -0
  329. package/skills/skills-codex/shared-references/wiki-helper-resolution.md +69 -0
  330. package/skills/skills-codex/shared-references/writing-principles.md +525 -0
  331. package/skills/skills-codex/slides-polish/SKILL.md +563 -0
  332. package/skills/skills-codex/specification-writing/SKILL.md +211 -0
  333. package/skills/skills-codex/system-profile/SKILL.md +103 -0
  334. package/skills/skills-codex/training-check/SKILL.md +83 -0
  335. package/skills/skills-codex/vast-gpu/SKILL.md +394 -0
  336. package/skills/skills-codex/web-debug-search/SKILL.md +334 -0
  337. package/skills/skills-codex/wiki-enrich/SKILL.md +255 -0
  338. package/skills/skills-codex/writing-systems-papers/SKILL.md +184 -0
  339. package/skills/skills-codex-claude-review/README.md +79 -0
  340. package/skills/skills-codex-claude-review/README_CN.md +78 -0
  341. package/skills/skills-codex-claude-review/auto-paper-improvement-loop/SKILL.md +581 -0
  342. package/skills/skills-codex-claude-review/auto-review-loop/SKILL.md +510 -0
  343. package/skills/skills-codex-claude-review/novelty-check/SKILL.md +102 -0
  344. package/skills/skills-codex-claude-review/paper-figure/SKILL.md +319 -0
  345. package/skills/skills-codex-claude-review/paper-plan/SKILL.md +287 -0
  346. package/skills/skills-codex-claude-review/paper-write/SKILL.md +420 -0
  347. package/skills/skills-codex-claude-review/research-refine/SKILL.md +732 -0
  348. package/skills/skills-codex-claude-review/research-review/SKILL.md +149 -0
  349. package/skills/skills-codex-gemini-review/README.md +176 -0
  350. package/skills/skills-codex-gemini-review/README_CN.md +175 -0
  351. package/skills/skills-codex-gemini-review/auto-paper-improvement-loop/SKILL.md +331 -0
  352. package/skills/skills-codex-gemini-review/auto-review-loop/SKILL.md +304 -0
  353. package/skills/skills-codex-gemini-review/grant-proposal/SKILL.md +630 -0
  354. package/skills/skills-codex-gemini-review/idea-creator/SKILL.md +263 -0
  355. package/skills/skills-codex-gemini-review/idea-discovery/SKILL.md +275 -0
  356. package/skills/skills-codex-gemini-review/idea-discovery-robot/SKILL.md +365 -0
  357. package/skills/skills-codex-gemini-review/novelty-check/SKILL.md +92 -0
  358. package/skills/skills-codex-gemini-review/paper-figure/SKILL.md +289 -0
  359. package/skills/skills-codex-gemini-review/paper-plan/SKILL.md +265 -0
  360. package/skills/skills-codex-gemini-review/paper-poster-html/SKILL.md +102 -0
  361. package/skills/skills-codex-gemini-review/paper-slides/SKILL.md +582 -0
  362. package/skills/skills-codex-gemini-review/paper-write/SKILL.md +346 -0
  363. package/skills/skills-codex-gemini-review/paper-writing/SKILL.md +312 -0
  364. package/skills/skills-codex-gemini-review/research-refine/SKILL.md +674 -0
  365. package/skills/skills-codex-gemini-review/research-review/SKILL.md +112 -0
  366. package/skills/slides-polish/SKILL.md +565 -0
  367. package/skills/specification-writing/SKILL.md +211 -0
  368. package/skills/system-profile/SKILL.md +103 -0
  369. package/skills/training-check/SKILL.md +132 -0
  370. package/skills/vast-gpu/SKILL.md +394 -0
  371. package/skills/web-debug-search/SKILL.md +334 -0
  372. package/skills/wiki-enrich/SKILL.md +257 -0
  373. package/skills/writing-systems-papers/SKILL.md +184 -0
  374. package/templates/CLAUDE_MD_TEMPLATE.md +29 -0
  375. package/templates/EXPERIMENT_LOG_TEMPLATE.md +47 -0
  376. package/templates/EXPERIMENT_PLAN_TEMPLATE.md +51 -0
  377. package/templates/EXPERIMENT_PLAN_TEMPLATE_CN.md +53 -0
  378. package/templates/FINDINGS_TEMPLATE.md +52 -0
  379. package/templates/IDEA_CANDIDATES_TEMPLATE.md +47 -0
  380. package/templates/IDEA_CANDIDATES_TEMPLATE_CN.md +47 -0
  381. package/templates/INVENTION_BRIEF_TEMPLATE.md +87 -0
  382. package/templates/MANIFEST_TEMPLATE.md +7 -0
  383. package/templates/NARRATIVE_REPORT_TEMPLATE.md +49 -0
  384. package/templates/PAPER_PLAN_TEMPLATE.md +47 -0
  385. package/templates/PATENT_CLAIMS_TEMPLATE.md +78 -0
  386. package/templates/PATENT_SPECIFICATION_TEMPLATE.md +68 -0
  387. package/templates/README.md +57 -0
  388. package/templates/RESEARCH_BRIEF_TEMPLATE.md +35 -0
  389. package/templates/RESEARCH_BRIEF_TEMPLATE_CN.md +41 -0
  390. package/templates/RESEARCH_CONTRACT_TEMPLATE.md +60 -0
  391. package/templates/claude-hooks/corpus_write_guard.json +16 -0
  392. package/templates/claude-hooks/corpus_write_guard.py +85 -0
  393. package/templates/claude-hooks/meta_logging.json +74 -0
  394. package/templates/gitignore-trace.txt +3 -0
  395. package/tools/__pycache__/check_skills_inventory.cpython-314.pyc +0 -0
  396. package/tools/arxiv_fetch.py +311 -0
  397. package/tools/capture_filter.py +126 -0
  398. package/tools/check_skills_inventory.py +273 -0
  399. package/tools/convert_skills_to_llm_chat.py +282 -0
  400. package/tools/copilot_native_evidence.py +818 -0
  401. package/tools/deepxiv_fetch.py +213 -0
  402. package/tools/evidence_check.py +212 -0
  403. package/tools/exa_search.py +425 -0
  404. package/tools/experiment_queue/README.md +118 -0
  405. package/tools/experiment_queue/build_manifest.py +44 -0
  406. package/tools/experiment_queue/queue_manager.py +44 -0
  407. package/tools/extract_paper_style.py +560 -0
  408. package/tools/figure_renderer.py +69 -0
  409. package/tools/forensics_gate.py +669 -0
  410. package/tools/generate_codex_claude_review_overrides.py +299 -0
  411. package/tools/idea_discovery_gate.py +256 -0
  412. package/tools/install_aris.ps1 +1372 -0
  413. package/tools/install_aris.sh +1370 -0
  414. package/tools/install_aris_codex.sh +1023 -0
  415. package/tools/install_aris_copilot.sh +1052 -0
  416. package/tools/iteration_log.py +143 -0
  417. package/tools/lint_skills_helpers.sh +84 -0
  418. package/tools/meta_opt/check_ready.sh +80 -0
  419. package/tools/meta_opt/log_event.sh +91 -0
  420. package/tools/meta_opt/trigger_eval.py +280 -0
  421. package/tools/meta_opt/trigger_evals.sample.json +28 -0
  422. package/tools/openalex_fetch.py +326 -0
  423. package/tools/overleaf_audit.sh +104 -0
  424. package/tools/overleaf_setup.sh +150 -0
  425. package/tools/paper_illustration_image2.py +62 -0
  426. package/tools/provenance.py +294 -0
  427. package/tools/research_wiki.py +1720 -0
  428. package/tools/review_gate.py +502 -0
  429. package/tools/run_state.py +399 -0
  430. package/tools/save_trace.sh +477 -0
  431. package/tools/semantic_scholar_fetch.py +438 -0
  432. package/tools/skill-groups.tsv +116 -0
  433. package/tools/skill_picker.py +238 -0
  434. package/tools/smart_update.ps1 +521 -0
  435. package/tools/smart_update.sh +591 -0
  436. package/tools/smart_update_codex.sh +419 -0
  437. package/tools/smart_update_copilot.sh +605 -0
  438. package/tools/threat_scan.py +222 -0
  439. package/tools/verify_paper_audits.sh +487 -0
  440. package/tools/verify_papers.py +613 -0
  441. package/tools/verify_wiki_coverage.sh +176 -0
  442. package/tools/watchdog.py +485 -0
@@ -0,0 +1,866 @@
1
+ ---
2
+ name: proof-checker
3
+ description: Rigorous mathematical proof verification and fixing workflow. Reads a LaTeX proof, identifies gaps via cross-model review (external reviewer backend, ultra reasoning), fixes each gap with full derivations, re-reviews, and generates an audit report. Use when user says "检查证明", "verify proof", "proof check", "审证明", "check this proof", or wants rigorous mathematical verification of a theory paper.
4
+ argument-hint: "[path-to-tex-file or proof-description] [--deep-fix] [--restatement-check]"
5
+ allowed-tools: Bash(*), Read, Grep, Glob, Write, Edit, Agent, mcp__codex__codex, mcp__codex__codex-reply, mcp__manual_review__review, mcp__manual_review__review_reply
6
+ ---
7
+
8
+ # Proof Checker: Rigorous Mathematical Verification & Fixing
9
+
10
+ > 🔒 **Do not wrap this skill in `/loop`, `/schedule`, or `CronCreate`.** It is
11
+ > verdict-bearing — it judges proof validity across rounds, threading the
12
+ > reviewer's memory from Phase 1 → Phase 3 via `codex-reply` so the reviewer can
13
+ > check whether a fix actually closed the gap it flagged. An external timer
14
+ > re-enters from the top each tick, starting a fresh thread and losing that
15
+ > memory. Schedule the *external wait that precedes it*, not the verdict. See
16
+ > [`shared-references/external-cadence.md`](../shared-references/external-cadence.md).
17
+
18
+ Systematically verify a mathematical proof via cross-model adversarial review, fix identified gaps, re-review until convergence, and generate a detailed audit report with proof-obligation accounting.
19
+
20
+ ## Context: $ARGUMENTS
21
+
22
+ ## Constants
23
+
24
+ - MAX_REVIEW_ROUNDS = 3
25
+ - REVIEWER_MODEL = `gpt-5.6-sol` — Default model for the Codex backend, reasoning effort `ultra` (deep-audit tier; capability fallback `gpt-5.6-sol`+`xhigh` → `gpt-5.5`+`xhigh` per `shared-references/reviewer-routing.md`, capability errors only — never below `xhigh`). Manual backend uses a model the user chooses, **but it must be a non-Claude model ARIS can classify** (OpenAI, Google, DeepSeek, Moonshot/Kimi, Qwen) — the executor is Claude, so routing the proof review into any Claude product makes Claude judge Claude and voids the cross-model invariant (see `shared-references/reviewer-routing.md`).
26
+ - **REVIEWER_BACKEND = `codex`** — Default: Codex MCP (ultra). Override with `— reviewer: oracle-pro` for Oracle MCP, or `— reviewer: manual` for Manual Review MCP. If manual-review MCP is unavailable, stop and print the install command; do not fall back to Codex. See `shared-references/reviewer-routing.md`.
27
+
28
+ ## Reviewer Calling Convention
29
+
30
+ When calling the reviewer, branch on REVIEWER_BACKEND:
31
+
32
+ **If REVIEWER_BACKEND = `codex`:**
33
+ Use `mcp__codex__codex` for new review threads
34
+ (`model: gpt-5.6-sol`, `config: {"model_reasoning_effort": "ultra"}`).
35
+ Use `mcp__codex__codex-reply` for follow-up rounds (reuse threadId).
36
+
37
+ **If REVIEWER_BACKEND = `manual`:**
38
+ Use `mcp__manual_review__review` for new review threads with:
39
+ prompt: [exact same prompt that would go to Codex]
40
+ config: {"model_reasoning_effort": "xhigh", "executor_model": "<actual executor model>", "require_reviewer_model": true}
41
+ Save the returned `threadId`.
42
+ Use `mcp__manual_review__review_reply` for follow-up rounds with:
43
+ threadId: [saved manual-review threadId]
44
+ prompt: [follow-up prompt]
45
+ config: {"model_reasoning_effort": "xhigh", "executor_model": "<actual executor model>", "require_reviewer_model": true}
46
+
47
+ Prompt fidelity: the manual prompt must be exactly the same text that Codex would receive.
48
+ Review tracing applies equally to both backends.
49
+
50
+ - AUDIT_DOC: `PROOF_AUDIT.md` at the paper directory root, alongside `main.tex` (cumulative log; when invoked via `/paper-writing`, this is `paper/PROOF_AUDIT.md`)
51
+ - REPORT_TEX: `proof_audit_report.tex` (formal before/after PDF)
52
+ - STATE_FILE: `PROOF_CHECK_STATE.json` (for recovery)
53
+ - SKELETON_DOC: `PROOF_SKELETON.md` (micro-claim inventory)
54
+ - **RENDER_HTML = true** — When `true` (default), auto-render `PROOF_AUDIT.md` to HTML at workflow end via `/render-html`. Uses **full Codex review gate** (audit-class artifact — math-heavy content; render-fidelity check protects against MathJax breakage and matches the skill's cross-model audit invariant). Set `false` to skip, or pass `— render html: false`.
55
+
56
+ ### Acceptance Gate (objective, replaces subjective scoring)
57
+
58
+ The proof passes when ALL of the following hold:
59
+ 1. Zero open FATAL or CRITICAL issues
60
+ 2. Every theorem/lemma has: (i) explicit hypotheses, (ii) proof with all interchanges justified, (iii) every application discharges hypotheses in the ledger
61
+ 3. All big-O/Θ/o statements have declared parameter dependence and uniformity scope
62
+ 4. Counterexample pass executed on all key lemmas (log candidates even if none found)
63
+
64
+ ## Issue Taxonomy (20 categories, 4 groups)
65
+
66
+ ### Group A: Logic & Proof Structure
67
+
68
+ | Category | Description | Example |
69
+ |----------|-------------|---------|
70
+ | **UNJUSTIFIED_ASSERTION** | Claim stated without proof or reference | "The Hessian splits into Gram blocks" |
71
+ | **UNPROVEN_SUBCLAIM** | "Clearly" / "it follows" hides a nontrivial lemma | "By symmetry, the cross-terms vanish" without checking |
72
+ | **QUANTIFIER_ERROR** | Wrong order ∀/∃, missing "for sufficiently small κ" | "For all π, there exists ε" vs "there exists ε for all π" |
73
+ | **IMPLICATION_REVERSAL** | Uses (A⇒B) as (B⇒A), or claims equivalence with only one direction | |
74
+ | **CASE_INCOMPLETE** | Misses boundary/degenerate cases | Singular covariance, zero weight, non-unique argmin |
75
+ | **CIRCULAR_DEPENDENCY** | Lemma uses theorem that depends on it | |
76
+ | **LOGICAL_GAP** | A step is not justified by what precedes it | B=Θ(1) → β_K=0 without analyzing W |
77
+
78
+ ### Group B: Analysis & Measure Theory
79
+
80
+ | Category | Description | Example |
81
+ |----------|-------------|---------|
82
+ | **ILLEGAL_INTERCHANGE** | Swaps limit/expectation/derivative/integral without DCT/MCT/Fubini | Differentiating under E without domination |
83
+ | **NONUNIFORM_CONVERGENCE** | Pointwise convergence used as uniform | sup and limit swapped |
84
+ | **MISSING_DOMINATION** | DCT cited but no dominating function given | |
85
+ | **INTEGRABILITY_GAP** | Uses E|X|^p without proving/assuming finite moments | |
86
+ | **REGULARITY_GAP** | Differentiability/Lipschitz/convexity used but not established | |
87
+ | **STOCHASTIC_MODE_CONFUSION** | Mixes a.s./in prob./in L²/in expectation | |
88
+
89
+ ### Group C: Model & Parameter Tracking
90
+
91
+ | Category | Description | Example |
92
+ |----------|-------------|---------|
93
+ | **MISSING_DERIVATION** | A quantity is used but never derived from the model | Risk functional with undefined B, W |
94
+ | **HIDDEN_ASSUMPTION** | Proof silently uses a condition not in the theorem | Gaussianity assumed but not stated |
95
+ | **INSUFFICIENT_ASSUMPTION** | Hypotheses too weak for proof (counterexample exists) | Moment conditions admitting 2-point distributions |
96
+ | **DIMENSION_TRACKING** | Parameter dependence (d, n, K, ...) not explicit | d enters only through κ |
97
+ | **NORMALIZATION_MISMATCH** | Coordinate/scaling conventions inconsistent | Rescaled vs raw coordinates |
98
+ | **CONSTANT_DEPENDENCE_HIDDEN** | "C" depends on d,n,K but treated as universal | |
99
+
100
+ ### Group D: Scope & Claims
101
+
102
+ | Category | Description | Example |
103
+ |----------|-------------|---------|
104
+ | **SCOPE_OVERCLAIM** | Conclusion stated more broadly than proof supports | "β_K=0" with only generic overlap |
105
+ | **REFERENCE_MISMATCH** | Cited theorem's hypotheses not verified at point of use | |
106
+
107
+ ## Two-Axis Severity System
108
+
109
+ ### Axis A — Proof Status (what is wrong)
110
+
111
+ | Status | Meaning |
112
+ |--------|---------|
113
+ | **INVALID** | Statement false as written (counterexample exists or contradiction) |
114
+ | **UNJUSTIFIED** | Could be true, but current proof does not establish it |
115
+ | **UNDERSTATED** | True only after strengthening assumptions |
116
+ | **OVERSTATED** | True only after weakening conclusion / adding qualifiers |
117
+ | **UNCLEAR** | Ambiguous notation / definition drift (not wrong per se) |
118
+
119
+ ### Axis B — Impact (how much breaks)
120
+
121
+ | Impact | Meaning |
122
+ |--------|---------|
123
+ | **GLOBAL** | Breaks main theorem or core dependency chain |
124
+ | **LOCAL** | Affects a side result but not the main theorem |
125
+ | **COSMETIC** | Exposition only |
126
+
127
+ ### Severity Labels (derived)
128
+
129
+ | Label | Definition |
130
+ |-------|------------|
131
+ | **FATAL** | INVALID + GLOBAL |
132
+ | **CRITICAL** | (INVALID + LOCAL) or (UNJUSTIFIED + GLOBAL) |
133
+ | **MAJOR** | (UNJUSTIFIED + LOCAL) or (UNDERSTATED/OVERSTATED + GLOBAL) |
134
+ | **MINOR** | Clarity / notation / dimension bookkeeping that doesn't change claims |
135
+
136
+ ## Side-Condition Checklists for Common Theorems
137
+
138
+ When the proof invokes any of the following, require explicit verification of ALL listed conditions:
139
+
140
+ | Theorem | Required Conditions |
141
+ |---------|-------------------|
142
+ | **DCT** (Dominated Convergence) | Pointwise a.e. convergence + integrable dominating function |
143
+ | **MCT** (Monotone Convergence) | Monotone increasing + non-negative |
144
+ | **Fubini/Tonelli** | Product measurability + integrability (Fubini) or non-negative (Tonelli) |
145
+ | **Leibniz integral rule** | Continuity of integrand + dominating function for derivative |
146
+ | **Implicit Function Theorem** | Continuous differentiability + non-singular Jacobian |
147
+ | **Taylor with remainder** | Sufficient differentiability + remainder form (Lagrange/integral) |
148
+ | **Jensen's inequality** | Convexity of function + integrability |
149
+ | **Cauchy-Schwarz** | Correct inner product space + integrability of both factors |
150
+ | **Weyl/Davis-Kahan** | Symmetry/Hermiticity + perturbation bound conditions |
151
+ | **Analytic continuation** | Domain connectivity + identity theorem conditions |
152
+ | **WLOG reduction** | Invariance under claimed symmetry + reduction is reversible |
153
+
154
+ ## Workflow
155
+
156
+ ### Phase 0: Preparation
157
+
158
+ 1. **Locate the proof**: Find the main `.tex` file(s).
159
+ 2. **Read the entire proof**: Extract list of all theorems/lemmas/propositions/corollaries/definitions/assumptions.
160
+ 3. **Read reference materials**: Reference papers, prior results.
161
+ 4. **Build a section map**: Structured list with line numbers and key claims.
162
+ 5. **Identify the main theorem**: Central result, assumptions, claims.
163
+
164
+ ### Phase 0.5: Proof-Obligation Ledger
165
+
166
+ > **Fan-out (Tier-aware) — build the ledger in parallel; never judge in
167
+ > parallel.** For a large multi-theorem paper, ledger *construction* is breadth
168
+ > over independent sections. **Tier 1** (Workflow): spawn one Claude subagent
169
+ > per section/theorem to extract that unit's symbols, assumptions, micro-claims,
170
+ > and local quantified statements, each returning a structured ledger fragment.
171
+ > **Tier 2**: the same subagents via the Agent tool. **Tier 3**: walk the
172
+ > sections sequentially. This follows
173
+ > [`shared-references/fan-out-pattern.md`](../shared-references/fan-out-pattern.md).
174
+ >
175
+ > Two hard rules:
176
+ > 1. **The shards EXTRACT, they do not ADJUDICATE.** Building the ledger
177
+ > (inventorying obligations, typing symbols, restating with explicit
178
+ > quantifiers) is structural extraction. Whether a proof step is *valid* —
179
+ > whether an obligation is actually discharged — is a Type-B correctness
180
+ > verdict reserved for the cross-model jury in Phase 1 / Phase 3 (codex or
181
+ > manual, `ultra`). A Claude shard MUST NOT mark a micro-claim "proved" or
182
+ > "sound"; it only records the obligation and where the paper claims to
183
+ > discharge it. See [`acceptance-gate.md`](../shared-references/acceptance-gate.md)
184
+ > — the loop may self-verify *that the ledger is complete*, never *that the
185
+ > proofs are correct*.
186
+ > - **This governs the ledger spec wording below.** Where the artifacts say
187
+ > "WHERE each is verified", "or mark UNVERIFIED", or "where conditions are
188
+ > proven", a shard records a **location pointer** (`file:line` the paper
189
+ > claims discharge) — never its own judgment that the discharge is
190
+ > mathematically valid. A shard's `UNVERIFIED` means *"the paper cites no
191
+ > discharge location"*, NOT *"the shard checked the math and it fails"*.
192
+ > Soundness is the jury's verdict, not the shard's.
193
+ >
194
+ > **Shard output** (extraction schema, per
195
+ > [`fan-out-pattern.md`](../shared-references/fan-out-pattern.md)): each shard
196
+ > returns `{shard_id: "<section/theorem id>", entries: [...]}` — the typed
197
+ > ledger items (symbols, assumptions, micro-claims, canonical statements,
198
+ > limit-order facts) for that unit, each carrying its canonical id (e.g.
199
+ > `MC-17`, the symbol name) as `dedup_key`. Never prose-only; never a validity
200
+ > verdict field.
201
+ > 2. **Global artifacts are a barrier, computed on the merged ledger, not
202
+ > per-shard.** The Dependency DAG and its cycle detection (incl. semantic
203
+ > circularity), and cross-section symbol-type consistency, require the whole
204
+ > paper in view. Merge all shard fragments first, then compute these on the
205
+ > union — a per-shard DAG would miss exactly the cross-section cycles this
206
+ > phase exists to catch.
207
+
208
+ Build formal accounting artifacts. Save to `PROOF_SKELETON.md`:
209
+
210
+ #### 1. Dependency DAG
211
+ Nodes = Definitions / Assumptions / Lemmas / Theorems. Edges = "uses". **Detect cycles** (including semantic circularity where Lemma A uses a corollary that quietly depends on A).
212
+
213
+ #### 2. Assumption Ledger
214
+ For each theorem/lemma, list every hypothesis with WHERE each is verified — i.e. the **location pointer** the paper claims discharges it (`file:line`), not a judgment that the discharge is valid; mark "UNVERIFIED" when the paper cites no discharge location (not when you believe the math fails — that is the jury's call). Track **usage-minimal assumption sets** — which assumptions were actually used vs merely stated.
215
+
216
+ #### 3. Typed Symbol Table
217
+ Each symbol must have a **type signature**:
218
+ ```
219
+ κ : scalar ∈ (0,1), depends on (d, α_t, Σ, μ)
220
+ u* : vector ∈ ℝ^d, u* = C^{-1}m
221
+ B^even : matrix ∈ ℝ^{(L+1)×(L+1)}, symmetric PSD
222
+ Ψ_v : function ℝ → ℝ, analytic in (ζ,κ), parity determined by v
223
+ ```
224
+ Flag any symbol whose meaning changes or whose type is inconsistent across uses.
225
+
226
+ #### 4. Canonical Quantified Statements
227
+ For each theorem/lemma, rewrite the statement with **explicit quantifiers, domains, and limit order**:
228
+ ```
229
+ ∀K ≥ 3, ∀π ∈ Π_K^{ms,∘} \ E_K, ∃κ_0 > 0 such that ∀κ ∈ (0, κ_0):
230
+ h_act^{(K,π)} = Θ(κ^{α_K^act}) [uniform in π on compact subsets]
231
+ ```
232
+ If you cannot restate a theorem this precisely, mark it **UNCLEAR — needs disambiguation**.
233
+
234
+ #### 5. Micro-Claim Inventory
235
+ Every nontrivial step becomes a numbered micro-claim in **sequent form**:
236
+ ```
237
+ MC-17: Context: [Lemma 3.1, κ < κ_0, Z_κ has bounded moments up to order 2m+2]
238
+ ⊢ Goal: P̂_0 is positive definite
239
+ Rule: monomials linearly independent on support of continuous distribution
240
+ Side-conditions: positive density near origin — claimed discharge: §B.2 (paper argues via GMM weak convergence; validity is the jury's call, not the shard's)
241
+ ```
242
+ Each micro-claim has: justification rule name + required conditions + where conditions are proven (a location pointer to where the paper claims to discharge them, not a validity judgment).
243
+
244
+ #### 6. Limit-Order Map
245
+ Track every asymptotic statement's **limit order and uniformity scope**:
246
+ ```
247
+ h_act = Θ(κ^α) [as κ→0, uniform in π on compact subsets of Π_K, for fixed K]
248
+ τ_act ~ (b/a)n [as n→∞, for fixed κ,K,π with x_K ≪ 1]
249
+ ```
250
+ Flag any statement where limit order is ambiguous or uniformity is unclear.
251
+
252
+ ### Phase 1: First Review (reviewer backend, ultra reasoning)
253
+
254
+ Submit the **complete proof content** with the checklist below, using the selected backend.
255
+
256
+ For `codex`, call `mcp__codex__codex` and always pin `model: gpt-5.6-sol` + `config: {"model_reasoning_effort": "ultra"}` (deep-audit tier). For `manual`, call `mcp__manual_review__review` with the identity-bearing config from the Reviewer Calling Convention above — `model`, `sandbox` and `cwd` are Codex-only.
257
+
258
+ Use this exact prompt for both backends:
259
+
260
+ ```
261
+ You are performing a rigorous mathematical proof review. For EVERY theorem,
262
+ lemma, and proposition, check ALL of the following:
263
+
264
+ ## MANDATORY CHECKS
265
+
266
+ A. DEFINITIONS: List any symbol whose meaning is ambiguous or changes.
267
+ B. HYPOTHESIS DISCHARGE: For each lemma/theorem APPLICATION (not statement),
268
+ list each hypothesis and whether it was verified, with location.
269
+ C. INEQUALITY AUDIT: For each inequality chain, verify direction, missing
270
+ absolute values, missing conditions (convexity, PSD, integrability).
271
+ D. INTERCHANGE AUDIT: Flag every limit/derivative/expectation/integral
272
+ interchange. State which theorem justifies it (DCT/MCT/Fubini/Leibniz)
273
+ and which conditions are verified/missing.
274
+ E. PROBABILITY MODE: Track whether claims are a.s./in prob./in expectation/
275
+ w.h.p. Ensure transitions are justified.
276
+ F. UNIFORMITY & CONSTANTS: For every O(·), o(·), Θ(·), ≲, state whether
277
+ it is uniform over all parameters. List hidden parameter dependence.
278
+ G. EDGE/DEGENERATE CASES: Attempt to break each key lemma with a 1D,
279
+ low-rank, or extreme-parameter construction.
280
+ H. DEPENDENCY CONSISTENCY: Detect cycles or forward references to unproven
281
+ results.
282
+
283
+ ## OUTPUT FORMAT (per issue)
284
+ For each issue found, provide:
285
+ - id: sequential number
286
+ - status: INVALID / UNJUSTIFIED / UNDERSTATED / OVERSTATED / UNCLEAR
287
+ - impact: GLOBAL / LOCAL / COSMETIC
288
+ - category: [from taxonomy]
289
+ - location: section/equation/line
290
+ - statement: what the proof claims
291
+ - why_invalid: why this is wrong or unjustified
292
+ - counterexample: YES (describe) / NO / CANDIDATE (describe attempt)
293
+ - affects: which downstream results break if this is wrong
294
+ - minimal_fix: how to fix it
295
+
296
+ [FULL PROOF CONTENT HERE]
297
+ ```
298
+
299
+ #### Phase 1 addendum — `--deep-fix` opt-in
300
+
301
+ If the user passed `--deep-fix` on invocation, append the following block to the reviewer prompt **after** the OUTPUT FORMAT block above (do **not** modify the original block; the new fields are additive). Default invocations skip this block entirely and emit the original output schema unchanged.
302
+
303
+ ```
304
+ ## DEEP-FIX OUTPUT (opt-in, only when --deep-fix is set)
305
+
306
+ For EACH issue listed above, additionally provide a `deep_fix_plan`
307
+ that is repair-grade — sufficient for an executor to apply the fix
308
+ in one Edit pass without spawning a follow-up review thread:
309
+
310
+ - issue_id: same as the issue id above
311
+ - corrected_statement: the theorem/lemma statement as it should
312
+ read after the fix, with explicit quantifiers, regime conditions,
313
+ and uniformity scope (LaTeX, paste-ready)
314
+ - changed_equations: list of {before: <LaTeX>, after: <LaTeX>}
315
+ pairs for each equation that needs replacement
316
+ - downstream_labels: list of \label{...} keys whose statements or
317
+ proofs depend on this fix and must be re-checked or rewritten
318
+ - minimal_tex_patch_plan: ordered list of concrete edits, each as
319
+ {file: <path>, anchor_old: <unique LaTeX snippet to find>,
320
+ replacement_new: <LaTeX to insert>}; the executor will pass
321
+ these directly to its file-editing tool
322
+ - closure_tests: 2-5 sanity checks the executor must run after
323
+ applying the fix (e.g., "verify constant_dependence_diff matches
324
+ computed value", "limit case γ→0 reduces to identity",
325
+ "dimension count matches before/after")
326
+
327
+ ## ALGEBRA / TYPE SANITY PASS (opt-in, only when --deep-fix is set)
328
+
329
+ If any issue invokes Schur test, Young's inequality, Cauchy-Schwarz,
330
+ Hölder, quadratic form, operator norm, or power counting, the
331
+ deep_fix_plan for that issue MUST also include an `algebra_sanity`
332
+ object:
333
+
334
+ - dimension_table: map of {symbol: type_signature}, e.g.
335
+ {"K(i,α)": "scalar ≥ 0",
336
+ "‖K‖_{2→2}": "scalar ≥ 0",
337
+ "Σ_i V_i^rem": "scalar quadratic in w"}
338
+ - power_count: number of times each operator-norm or Schur factor
339
+ appears on each side; flag mismatch as INVALID
340
+ - zero_coupling_check: evaluate the expression at γ=0 (or the
341
+ analogous degenerate point); confirm it reduces to the expected
342
+ identity / vanishing case
343
+ - constant_dependence_diff: list of constants whose dependence on
344
+ (d, K, n, ...) changes between BEFORE and AFTER, with the new
345
+ explicit dependence written out
346
+
347
+ Be precise. The executor will apply this plan literally; vague
348
+ prose ("strengthen the bound", "redo the Schur step") is not
349
+ acceptable in deep-fix mode. If you cannot produce a precise plan
350
+ for an issue, omit that issue's deep-fix block and signal the
351
+ deep-fix path is unavailable — do NOT emit a vague plan, and do
352
+ NOT add a deep-fix-only category (e.g. UNCLEAR_DEEP_FIX) into the
353
+ standard issue list, since that contaminates default-call output.
354
+ ```
355
+ A verdict-bearing manual response MUST begin with
356
+ `Reviewer-Model: <exact-model-id>` — pass the model THIS session is actually
357
+ running as in `executor_model`. Missing, unknown, or same-family identity
358
+ cannot acquit; emit `REVIEW_UNAVAILABLE` rather than guessing. If the executor
359
+ model cannot be named, manual review's cross-family claim is unprovable — say
360
+ so in the report instead of asserting it.
361
+
362
+
363
+ **Save the threadId.** Parse into structured issue list. Write to `PROOF_AUDIT.md`.
364
+
365
+ ### Phase 1.5: Counterexample Red Team
366
+
367
+ For each CRITICAL or MAJOR issue, and for every key lemma that introduces:
368
+ - a new inequality bound
369
+ - an identifiability/uniqueness claim
370
+ - a curvature/PSD/strong convexity assertion
371
+ - a uniform-in-parameter claim
372
+ - a convergence mode upgrade (pointwise → uniform, in prob → w.h.p.)
373
+
374
+ Systematically attempt to construct counterexamples using:
375
+
376
+ | Strategy | Description |
377
+ |----------|-------------|
378
+ | **Dimensional collapse** | Set d=1 or 2, K=2, n small |
379
+ | **Degeneracy** | Singular covariance, tiny weight, overlapping means, identical components |
380
+ | **Extremal distributions** | Two-point ±a, bounded non-subGaussian, heavy tails |
381
+ | **Adversarial parameter scaling** | Pick parameters making neglected terms dominate |
382
+ | **Numeric falsification** | Translate lemma to a function, brute-force optimize over small domain |
383
+
384
+ **Rule**: Label "counterexample found" ONLY if algebraically verified. Otherwise log as "candidate counterexample — needs verification."
385
+
386
+ Record all attempts (successful or not) in `PROOF_AUDIT.md`.
387
+
388
+ ### Phase 2: Fix Implementation
389
+
390
+ For each issue, ordered by severity (FATAL → CRITICAL → MAJOR → MINOR):
391
+
392
+ #### Step 2a: Choose fix strategy
393
+ For each issue, explicitly choose one of:
394
+ - **ADD_DERIVATION**: Write missing proof steps
395
+ - **STRENGTHEN_ASSUMPTION**: Add conditions to theorem statement
396
+ - **WEAKEN_CLAIM**: Reduce conclusion scope
397
+ - **ADD_REFERENCE**: Cite known result + verify its conditions apply
398
+
399
+ Log this choice — it is a scope-changing decision when it alters theorem statements.
400
+
401
+ #### Step 2b: Derive the fix mathematically
402
+ - Complete mathematical derivation, not just a claim
403
+ - If new proposition/lemma needed, write in full theorem-proof style
404
+
405
+ #### Step 2c: Implement in LaTeX
406
+ - Edit the `.tex` file
407
+ - Preserve existing `\label` references where possible
408
+
409
+ #### Step 2d: Record the fix
410
+ ```markdown
411
+ ### Fix N: [SHORT TITLE]
412
+ **Issue**: [id] [CATEGORY] — [description]
413
+ **Severity**: FATAL / CRITICAL / MAJOR / MINOR
414
+ **Status**: INVALID / UNJUSTIFIED / UNDERSTATED / OVERSTATED
415
+ **Impact**: GLOBAL / LOCAL / COSMETIC
416
+ **Fix strategy**: ADD_DERIVATION / STRENGTHEN_ASSUMPTION / WEAKEN_CLAIM / ADD_REFERENCE
417
+ **Location**: Section X, Lines Y-Z
418
+
419
+ **BEFORE**: [what the proof originally did]
420
+ **WHY WRONG**: [mathematical problem, with counterexample if applicable]
421
+ **AFTER**: [what the fix does]
422
+ **KEY EQUATION**: [central new equation]
423
+ **PROOF OBLIGATIONS ADDED**: [new conditions/lemmas introduced]
424
+ **DOWNSTREAM EFFECTS**: [which results now need re-checking]
425
+ ```
426
+
427
+ #### Step 2e: Compile check
428
+ ```bash
429
+ pdflatex -interaction=nonstopmode <file>.tex 2>&1 | grep -E "Error|Warning|undefined"
430
+ ```
431
+
432
+ ### Phase 3: Re-Review (reviewer backend, ultra reasoning)
433
+
434
+ Continue with the selected backend. For `codex`, use `mcp__codex__codex-reply` with the saved threadId. For `manual`, use `mcp__manual_review__review_reply` with the saved threadId. Include fix summaries. Request the same mandatory checklist.
435
+
436
+ Check acceptance gate. If not met, repeat Phases 2-3 (up to MAX_REVIEW_ROUNDS).
437
+
438
+ ### Phase 3.5: Global Closure & Independent Verification
439
+
440
+ #### Global closure checks
441
+ After all fixes, verify the proof as a whole:
442
+ - **Statement–conclusion match**: Does the proof end with EXACTLY what the theorem claims (quantifiers, constants, uniformity)?
443
+ - **All obligations discharged**: Every node in the obligation DAG is proven or explicitly assumed (and the theorem statement includes it).
444
+ - **Case analysis coverage**: Cases partition the domain AND include boundary/degenerate cases.
445
+ - **Induction correctness** (if applicable): Base case, inductive step, correct use of IH, induction measure strictly decreases.
446
+ - **WLOG reductions**: Each "without loss of generality" spawns a micro-claim proving the reduction is lossless.
447
+ - **No silent assumption strengthening**: Any fix that strengthened assumptions has propagated to the main theorem statement.
448
+
449
+ #### Independent second review for FATAL/CRITICAL fixes
450
+ For any fix that resolved a FATAL or CRITICAL issue, submit the **fixed section alone** (without showing the previous critique) to a **fresh reviewer thread** using the selected backend. Do NOT use a reply tool — this step must be blind.
451
+
452
+ *For codex:* start a fresh `mcp__codex__codex` thread.
453
+ *For manual:* start a fresh `mcp__manual_review__review` thread.
454
+
455
+ The blind review prompt:
456
+
457
+ ```
458
+ [Codex:]
459
+ mcp__codex__codex:
460
+ model: gpt-5.6-sol
461
+ config: {"model_reasoning_effort": "ultra"}
462
+ prompt: |
463
+ Blind review of the following proof section. You have NOT seen any prior
464
+ review or discussion. Check every step for correctness, hidden assumptions,
465
+ illegal interchanges, and counterexamples.
466
+ [FIXED SECTION ONLY]
467
+ ```
468
+
469
+ If the blind reviewer finds new issues, re-enter Phase 2.
470
+
471
+ #### Regression proof-audit
472
+ After fixes, re-run:
473
+ - DAG acyclicity check (no new cycles introduced)
474
+ - Counterexample suite on all DOWNSTREAM lemmas of modified results
475
+ - Assumption-delta report: what became stronger/weaker due to fixes?
476
+
477
+ ### Phase 3.6: Theorem Restatement Regression (opt-in)
478
+
479
+ **Default**: skipped. Existing callers see no change.
480
+
481
+ **Opt-in**: pass `--restatement-check` on invocation. The skill then runs a cross-location consistency pass after Phase 3.5 (Global Closure) and before Phase 3.9 (Unrecoverable Protocol).
482
+
483
+ This phase catches a specific class of bugs that Phase 3.5's "Statement-conclusion match" check does NOT catch: **drift between the canonical theorem statement and its restatements elsewhere in the paper** (summary tables, "Key Contributions" / "Summary" sections, abstract, discussion, captions). Common drift patterns observed in practice:
484
+
485
+ - Main theorem makes a partly-conditional claim (e.g., "regime A unconditional, regime B conditional on assumption X") but a later summary cites it as the fully unconditional version.
486
+ - Main theorem κ exponent is `O(d²K²)` but a constants table writes `O(K²)`.
487
+ - Main theorem regime condition is squared envelope, but a remark elsewhere still shows the first-order envelope.
488
+ - Restatement quietly drops a quantifier ("for $n \ge n_0$") that the proof relied on.
489
+
490
+ #### Algorithm
491
+
492
+ 1. **Build canonical statement table.** Scan all `*.tex` files in the paper for `\begin{theorem}` / `\begin{lemma}` / `\begin{proposition}` / `\begin{corollary}` blocks with a `\label{...}`. For each: record `(label, full_statement_text, file:line_range)`.
493
+
494
+ 2. **Collect restatement candidates.** For each canonical label `thm:foo`:
495
+ - Every `\Cref{thm:foo}` / `\ref{thm:foo}` / `\cref{thm:foo}` invocation, with the surrounding 2 sentences as candidate restatement context.
496
+ - Rows in tables (`\begin{tabular}` … `\end{tabular}`) that mention the label, the theorem's informal name, or its constants.
497
+ - Bullets in "Key Contributions" / "Summary" / "Main Results" lists.
498
+ - Sentences in the abstract and introduction that paraphrase the theorem.
499
+
500
+ 3. **Normalized diff.** For each (canonical_statement, restatement_context) pair:
501
+ - Strip `\,` `\;` `\!` `~`, normalize whitespace, normalize math-mode delimiters (`$..$`, `\(..\)`, display vs inline).
502
+ - Detect drift signatures (one or more):
503
+ - **conditional_loss** — canonical says "under \Cref{ass:X}" or "for w=1"; restatement omits the conditional.
504
+ - **scope_change** — big-O exponent or rate differs (`O(K^2)` vs `O(d^2K^2)`; `√n` vs `n`).
505
+ - **quantifier_loss** — quantifier present canonically (e.g. "for $n \ge n_0$", "for sufficiently small $\gamma$") absent in restatement.
506
+ - **regime_envelope_change** — first-order leakage `Cγ/(1-Δγ)` vs squared envelope `(Cγ/(1-Δγ))²` (or analogous).
507
+ - **constant_change** — different numeric constant or different parameter dependence stated.
508
+ - **variable_rename** — same role in argument played by differently-named symbol with no explicit alias.
509
+
510
+ 4. **Emit findings.** Each detected drift becomes an entry in `details.restatement_drift` (see "Submission Artifact Emission" below). Severity defaults to **MAJOR** (UNDERSTATED/OVERSTATED + GLOBAL); reviewer may downgrade to MINOR if drift is purely cosmetic (e.g. `\,` placement) and upgrade to CRITICAL only if the restatement is used downstream as if it were the canonical (e.g. another proof cites the restated version).
511
+
512
+ #### What this phase does NOT do
513
+ - It does **not** fix drift automatically. The output is advisory; the executor or a follow-up `--deep-fix` run handles the rewrite.
514
+ - It does **not** alter `details.issues`; restatement drift is reported as a sibling field, so existing consumers reading `issues[]` see no schema change.
515
+ - It does **not** alter the top-level `verdict` decision rule. A paper with non-empty `restatement_drift` may still emit `PASS` if all proof obligations are otherwise discharged; the drift is independent of proof correctness. (The reviewer may at its discretion mirror a **CRITICAL**-severity drift into the `issues` list as a regular issue — typically when the restated version is used downstream as if it were canonical — but that is a per-issue judgment, never an automatic rule. MAJOR or MINOR drift never mirrors into `issues`.)
516
+
517
+ #### Failure mode
518
+ If `--restatement-check` is set but the cross-location scan cannot complete, emit `details.restatement_drift: []` plus `details.restatement_check_status: "unavailable"` with a one-line note explaining why. Verifier gates and downstream skills MUST treat `"unavailable"` identically to the field being absent: not blocking. Phases 1 / 1.5 / 2 / 3 / 3.5 / 3.9 / 4 / 5 still run normally. Cases that trigger this fallback include:
519
+ - Unreadable `.tex` (parser / encoding error on a file that contains a theorem block).
520
+ - Ambiguous label resolution (e.g., the same `\label{thm:foo}` appears more than once with no clear canonical pick).
521
+ - **No labeled canonical theorem-like block found** (the algorithm only inspects `\begin{theorem|lemma|proposition|corollary}` blocks with an explicit `\label{...}`; if there is no such block, there is nothing to compare restatements against).
522
+
523
+ ### Phase 3.9: Unrecoverable Proof Protocol
524
+
525
+ If acceptance gate is not met after MAX_REVIEW_ROUNDS, output a **Proof Unrecoverable Report**:
526
+ 1. Minimal set of blocking FATAL/CRITICAL issues that could not be resolved
527
+ 2. Salvage options ranked: (a) weaken claim, (b) strengthen assumptions, (c) add missing lemmas, (d) restructure argument
528
+ 3. Which parts of the proof are likely still reusable
529
+ 4. Recommended next steps for the author
530
+
531
+ Do NOT silently declare success. The report must be honest.
532
+
533
+ ### Phase 4: Audit Report Generation
534
+
535
+ Generate `proof_audit_report.tex` with:
536
+
537
+ 1. **Overview table**: All issues with two-axis severity, category, fix strategy, status
538
+ 2. **Before/After logic chain**: Red (BEFORE) → Green (AFTER) comparison
539
+ 3. **For each fix**: original proof → why wrong → counterexample (if any) → complete derivation → remaining subtleties
540
+ 4. **Proof-obligation diff**: What was unverified before, what is verified now
541
+ 5. **Summary**: Now proven / still assumed / open problems
542
+ 6. **Colored boxes**: BEFORE (red), AFTER (green), WHY WRONG (orange), KEY INSIGHT (blue), WARNING (yellow)
543
+
544
+ Compile: `pdflatex proof_audit_report.tex && pdflatex proof_audit_report.tex`
545
+
546
+ ### Phase 5: State Persistence
547
+
548
+ Write `PROOF_CHECK_STATE.json`:
549
+ ```json
550
+ {
551
+ "status": "completed",
552
+ "rounds": 2,
553
+ "threadId": "...",
554
+ "fatal_fixed": 0,
555
+ "critical_fixed": 3,
556
+ "major_fixed": 2,
557
+ "minor_fixed": 1,
558
+ "counterexamples_found": 1,
559
+ "counterexample_candidates": 2,
560
+ "acceptance_gate": "PASS",
561
+ "timestamp": "..."
562
+ }
563
+ ```
564
+
565
+ ### Phase 5.5: Research Wiki Claim Ledger (additive; only if a wiki is active)
566
+
567
+ If — and only if — a `research-wiki/` exists, persist each top-level
568
+ theorem/headline as a **claim node** so the wiki's PROVE/JUDGE ledger records what
569
+ was proven and with what honesty. This is the **birth point** for wiki claim nodes
570
+ (`claims/<slug>.md`). It is a **detect-only record, never a verdict**: it never
571
+ changes the audit's `verdict`/`reason_code`, never blocks, and is skipped entirely
572
+ when `verdict == NOT_APPLICABLE` (no theorems) or no wiki is found.
573
+
574
+ > The claim's `status` is the **PROOF axis only** (`verified` / `sound-modulo-imports`
575
+ > / `refuted` / `unproven` / `drafted` / `retracted`). Empirical experiment support is
576
+ > a **separate axis** carried by `supports` / `invalidates` *edges* from
577
+ > `/result-to-claim` — those words are NEVER written into this `status` field (the
578
+ > `research_wiki.py` validator rejects them).
579
+
580
+ Resolve the helper via the **canonical resolver** (integration-contract §2). Check
581
+ `NOT_APPLICABLE` first (in cwd, where the audit ran), then run from the project root
582
+ (the wiki sits at root; a paper audit may run from a `paper/` subdir). Every guard
583
+ warn-and-skips — the audit is already complete and is never affected:
584
+
585
+ ```bash
586
+ grep -q '"verdict"[[:space:]]*:[[:space:]]*"NOT_APPLICABLE"' PROOF_AUDIT.json 2>/dev/null \
587
+ && { echo "verdict NOT_APPLICABLE (no theorems); skipping claim ledger" >&2; exit 0; }
588
+ cd "$(git rev-parse --show-toplevel 2>/dev/null || pwd)" || exit 0
589
+ [ -d research-wiki ] || { echo "no research-wiki/ at project root; skipping claim ledger (audit complete)" >&2; exit 0; }
590
+ if [ -z "${ARIS_REPO:-}" ] && [ -f .aris/installed-skills.txt ]; then
591
+ ARIS_REPO=$(awk -F'\t' '$1=="repo_root"{print $2; exit}' .aris/installed-skills.txt 2>/dev/null) || true
592
+ fi
593
+ if [ -z "${ARIS_REPO:-}" ] && [ -f "$HOME/.aris/repo" ]; then
594
+ ARIS_REPO=$(cat "$HOME/.aris/repo" 2>/dev/null) || true
595
+ fi
596
+ WIKI_SCRIPT=".aris/tools/research_wiki.py"
597
+ [ -f "$WIKI_SCRIPT" ] || WIKI_SCRIPT="tools/research_wiki.py"
598
+ [ -f "$WIKI_SCRIPT" ] || { [ -n "${ARIS_REPO:-}" ] && WIKI_SCRIPT="$ARIS_REPO/tools/research_wiki.py"; }
599
+ [ -f "$WIKI_SCRIPT" ] || { echo "WARN: research_wiki.py not resolved; skipping claim ledger (audit unaffected)" >&2; exit 0; }
600
+ ```
601
+
602
+ For each top-level theorem/headline in `PROOF_SKELETON.md` (the main theorem + the
603
+ key lemmas the headline depends on — NOT every micro-claim), map the audit outcome
604
+ to an **honest** claim status:
605
+
606
+ | Audit outcome (`PROOF_AUDIT.json` verdict + Phase 1.5 counterexamples) | `--status` |
607
+ |---|---|
608
+ | `verdict=PASS` / `all_proofs_complete`, no flagged imports | `verified` |
609
+ | proof closes but rests on a flagged `[unverified-axiom]` / imported result | `sound-modulo-imports` |
610
+ | counterexample found (Phase 1.5) **or** reviewer judged the statement false | `refuted` |
611
+ | open gap (unresolved FATAL/CRITICAL / UNJUSTIFIED, no counterexample) | `unproven` |
612
+
613
+ Never invent a status: a plain proof gap is `unproven` — never `refuted` (which
614
+ asserts falsity) and never `verified`/`drafted`. Then record (idempotent):
615
+
616
+ ```bash
617
+ python3 "$WIKI_SCRIPT" add_claim research-wiki/ \
618
+ --slug "<stable-theorem-id, e.g. thm-main-ub>" --name "<theorem headline>" \
619
+ --status "<mapped status>" --provenance "<trace_path from PROOF_AUDIT.json>" \
620
+ --statement "<canonical theorem statement>" \
621
+ --scope "<what it does NOT say; any flagged imports>" \
622
+ --evidence "<PROOF_AUDIT.json verdict + counterexample / obligation pointers>" \
623
+ --update-on-exist \
624
+ || echo "WARN: add_claim failed for <slug> (audit unaffected; fix wiki/status and re-run)" >&2
625
+ ```
626
+
627
+ `--update-on-exist` lets a re-audit refresh a claim's status (an `unproven` claim
628
+ becomes `verified` once the gap closes). `--provenance` is the honesty receipt (the
629
+ `trace_path`); never omit it.
630
+
631
+ ## Deep-Fix Mode (opt-in)
632
+
633
+ **Default**: disabled. The Phase 1 reviewer emits issues with `minimal_fix` (a 1-2 sentence pointer); existing callers see no change.
634
+
635
+ **Opt-in**: pass `--deep-fix` on invocation. The Phase 1 reviewer prompt is **augmented** (not replaced) with the "DEEP-FIX OUTPUT" and "ALGEBRA / TYPE SANITY PASS" blocks above, so the reviewer also returns a `deep_fix_plan` per issue: corrected statement, changed equations, downstream label list, minimal LaTeX patch plan, and closure tests. Issues invoking Schur / Young / Cauchy-Schwarz / Hölder / quadratic forms / operator norms / power counting additionally carry an `algebra_sanity` block (dimension table + power count + zero-coupling check + constant-dependence diff).
636
+
637
+ ### Why opt-in
638
+ The default `minimal_fix` prose is intentionally short — it suits the common case where the executor wants high-level pointers and will derive the patch separately. Forcing `deep_fix_plan` on every run would (a) inflate every reviewer call by 2-5×, (b) trigger re-review thrash for issues the executor has already decided to weaken or defer, and (c) change the shape of `details.issues` for every existing caller. The flag preserves zero behavior change for default invocations while letting the executor request repair-grade output when the fix is going to be applied immediately.
639
+
640
+ ### Effect when enabled
641
+ - The Phase 1 reviewer prompt is augmented with the deep-fix and algebra-sanity blocks; nothing in the original mandatory checklist or output format is removed.
642
+ - `PROOF_AUDIT.json` `details` gains a sibling field `deep_fix_plans` (parallel to `details.issues`); see "Submission Artifact Emission" below.
643
+ - The top-level `verdict`, `reason_code`, and `summary` are **unchanged in shape and decision rule**: deep-fix output is advisory tooling for the executor, not a verdict-altering signal.
644
+ - Verifier gates and downstream skills (`paper-writing` Phase 6, `verify_paper_audits.sh`) MUST treat absence of `deep_fix_plans` as the only valid default state and MUST NOT block on its presence or content.
645
+
646
+ ### When opt-in is appropriate
647
+ - The executor intends to apply the fix in the same session and wants to skip a follow-up "give me a concrete patch" thread.
648
+ - A previous default-mode run identified a CRITICAL or MAJOR issue whose `minimal_fix` was too vague to act on (e.g., "redo the Schur step" without specifying the corrected operator-norm bound).
649
+ - Algebra-heavy proofs with Schur / quadratic-form / operator-norm steps where the reviewer's first pass has consistently produced under-specified fixes.
650
+
651
+ ### Failure modes
652
+
653
+ A deep-fix-only failure must never contaminate the default proof-check output. All of the following paths emit `details.deep_fix_status: "unavailable"` + `details.deep_fix_plans: []` (with a one-line note in `details.deep_fix_note`) and leave the standard issue list, top-level verdict, reason_code, and summary unchanged:
654
+
655
+ - The reviewer refuses to produce a repair-grade plan because the fix would require choices the reviewer is unwilling to make.
656
+ - The Phase 1 reviewer call returns truncated or malformed deep-fix output (parse failure on the augmented section).
657
+ - The augmented Phase 1 call times out before producing the deep-fix block, but otherwise returned a valid normal proof review.
658
+
659
+ Verifier gates MUST treat `unavailable` identically to the field being absent: not blocking. Do **not** add a `UNCLEAR_DEEP_FIX` (or any deep-fix-only) entry into `details.issues`, since `details.issues` is the default schema's issue list and adding deep-fix-specific failures to it would change default behavior for callers without the flag.
660
+
661
+ If the augmented Phase 1 call fails so badly that the normal proof review cannot be recovered (e.g., the reviewer thread itself errored), retry once with the unaugmented prompt **only when the error proves the call never executed** (schema/validation or explicit capability error per the fallback chain in `reviewer-routing.md`); on timeout, rate-limit, transport, or server errors do NOT blind-retry (the review may have run — double-running double-bills), fall through directly to the existing reviewer-failure path that maps to the top-level `ERROR` verdict.
662
+
663
+ ## Key Rules
664
+
665
+ ### Mathematical rigor
666
+ - **Never accept a proof step on faith**. "Clearly" / "it follows" / "by standard arguments" are red flags — each must spawn a micro-claim.
667
+ - **Hypothesis discharge**: Every time a lemma is APPLIED, verify EACH of its hypotheses at that point. Use the side-condition checklists above.
668
+ - **Interchange discipline**: Every swap of limit/expectation/derivative/integral must cite a theorem (DCT/MCT/Fubini/Leibniz) and verify its conditions with explicit dominating function or integrability proof.
669
+ - **Uniformity discipline**: Every O(·)/Θ(·) must declare what parameters it is uniform over. "O(1)" that secretly depends on d,n,K is a CONSTANT_DEPENDENCE_HIDDEN issue.
670
+ - **Quantifier discipline**: Check ∀/∃ order. "For sufficiently small κ" must specify: does κ₀ depend on K? On π? On d?
671
+ - **Counterexample-first**: Before trying to fix a gap, first try to break it.
672
+ - **WLOG prohibition**: Every "without loss of generality" must have an explicit micro-claim proving the reduction. No free WLOGs.
673
+ - **No silent assumption strengthening**: Any fix that adds conditions must propagate to the theorem statement.
674
+
675
+ ### Cross-model protocol
676
+ - **Executor analyzes, reviewer critiques**: Claude reads proof, formulates questions, implements fixes. The external reviewer provides adversarial review.
677
+ - **Reviewer reasoning always ultra** (deep-audit tier): never below `xhigh` — only the capability fallback chain in `reviewer-routing.md` may step down, and only on explicit capability errors.
678
+ - **Send full content**: Don't summarize — send actual math for line-by-line checking.
679
+ - **Preserve threadId within a single run**: Use the appropriate reply tool (`mcp__codex__codex-reply` or `mcp__manual_review__review_reply`) for Phase 3 follow-up rounds within the same top-level `/proof-checker` invocation, so the reviewer keeps prior-issue context when judging whether a fix closed the gap. Across separate top-level invocations, always start a fresh thread (see "Thread independence" below).
680
+
681
+ ### Fix quality
682
+ - **Minimal fixes**: Fix exactly what's broken, nothing more.
683
+ - **Full derivation**: Every fix includes complete mathematical argument.
684
+ - **Explicit scope decisions**: Each fix is tagged ADD_DERIVATION / STRENGTHEN_ASSUMPTION / WEAKEN_CLAIM / ADD_REFERENCE.
685
+ - **Compile after each fix**: LaTeX must compile cleanly.
686
+
687
+ ### Scope honesty
688
+ - **Don't overclaim**: If a fix makes a result conditional, say so.
689
+ - **Separate "proven" from "assumed"**: The audit report has an explicit section for this.
690
+ - **Log open problems**: Issues requiring future work are listed, not hidden.
691
+
692
+ ### Opt-in flag discipline
693
+ - **Deep-fix is opt-in only**: never auto-enable; never block on `deep_fix_plans` content; existing callers must observe identical reviewer output and identical JSON schema if they do not pass `--deep-fix`.
694
+ - **Reviewer prompt augmentation is additive**: the deep-fix block is appended to the Phase 1 prompt, not substituted for any part of it. The original mandatory checklist (A-H) and original per-issue OUTPUT FORMAT remain in place verbatim.
695
+ - **Restatement check is opt-in only**: Phase 3.6 runs only when `--restatement-check` is set; existing callers must observe identical reviewer output and identical JSON schema if they do not pass the flag.
696
+ - **No Phase reordering**: enabling Phase 3.6 inserts it strictly between 3.5 and 3.9; it does not skip any other phase or change their semantics.
697
+ - **No verdict crosstalk**: neither deep-fix output nor `restatement_drift` ever alters top-level `verdict` or `reason_code`. A paper with non-empty drift or with deep-fix plans may still pass; a paper with FAIL verdict stays FAIL whether or not either flag was set.
698
+
699
+ ## Output Files
700
+
701
+ | File | Content | When |
702
+ |------|---------|------|
703
+ | `PROOF_SKELETON.md` | Dependency DAG + assumption ledger + micro-claims | Phase 0.5 |
704
+ | `PROOF_AUDIT.md` | Cumulative round-by-round audit log | Updated each round |
705
+ | `PROOF_AUDIT.json` | Machine-readable submission verdict (see below) | Always emitted |
706
+ | `proof_audit_report.tex/.pdf` | Formal before/after report | Phase 4 |
707
+ | `PROOF_CHECK_STATE.json` | State for recovery | Phase 5 |
708
+ | `PROOF_AUDIT.html` (+ `.review.json` sidecar) | Single-file HTML view of `PROOF_AUDIT.md` auto-rendered via `/render-html "PROOF_AUDIT.md" --json "PROOF_AUDIT.json"`. **Non-blocking** — if `/render-html` fails the audit still counts as complete; `PROOF_AUDIT.{md,json}` are the canonical outputs. | Workflow end (when `RENDER_HTML = true`, default) |
709
+
710
+ When `--restatement-check` is set, `PROOF_AUDIT.json` additionally carries `details.restatement_drift` and `details.restatement_check_status`; both fields are omitted when the flag is unset. See "Submission Artifact Emission" below.
711
+
712
+ ## Submission Artifact Emission
713
+
714
+ This skill **always** writes `PROOF_AUDIT.json` at the paper directory
715
+ root (i.e. `paper/PROOF_AUDIT.json` when invoked from `/paper-writing`
716
+ with paper-dir `paper/`; `<your-paper-dir>/PROOF_AUDIT.json` when invoked
717
+ standalone), regardless of caller or whether the paper contains theorems.
718
+ A paper with no `\begin{theorem}` / `\begin{lemma}` / `\begin{proof}` emits
719
+ verdict `NOT_APPLICABLE`; silent skip is forbidden. `paper-writing`
720
+ Phase 6 and `verify_paper_audits.sh` both rely on this artifact
721
+ existing at `<paper-dir>/PROOF_AUDIT.json`.
722
+
723
+ The artifact conforms to the schema in `shared-references/assurance-contract.md`:
724
+
725
+ ```json
726
+ {
727
+ "audit_skill": "proof-checker",
728
+ "verdict": "PASS | WARN | FAIL | NOT_APPLICABLE | BLOCKED | ERROR",
729
+ "reason_code": "all_proofs_complete | minor_gaps | critical_gap | no_theorems | ...",
730
+ "summary": "One-line human-readable verdict summary.",
731
+ "audited_input_hashes": {
732
+ "main.tex": "sha256:...",
733
+ "sections/4.theory.tex": "sha256:..."
734
+ },
735
+ "trace_path": ".aris/traces/proof-checker/<date>_run<NN>/",
736
+ "thread_id": "<codex mcp thread id>",
737
+ "reviewer_model": "<resolved — the model that actually ran (target: gpt-5.6-sol)>",
738
+ "reviewer_reasoning": "<resolved — the effort that actually ran (target: ultra)>",
739
+ "generated_at": "<UTC ISO-8601>",
740
+ "details": {
741
+ "theorems_audited": <int>,
742
+ "issues": [ { "id": "T1-H3", "severity": "FATAL|CRITICAL|MAJOR|MINOR",
743
+ "category": "quantifier|domination|...",
744
+ "location": "sections/4.theory.tex:L182",
745
+ "note": "..." }, ... ]
746
+ }
747
+ }
748
+ ```
749
+
750
+ ### Optional: `details.deep_fix_plans` (only when `--deep-fix` is set)
751
+
752
+ ```json
753
+ "details": {
754
+ ...
755
+ "deep_fix_plans": [
756
+ {
757
+ "issue_id": "T1-H3",
758
+ "corrected_statement": "<LaTeX, paste-ready>",
759
+ "changed_equations": [{"before": "<LaTeX>", "after": "<LaTeX>"}, ...],
760
+ "downstream_labels": ["thm:convergence", "cor:minimax", ...],
761
+ "minimal_tex_patch_plan": [
762
+ {"file": "sections/4.theory.tex",
763
+ "anchor_old": "<unique LaTeX snippet>",
764
+ "replacement_new": "<LaTeX to insert>"},
765
+ ...
766
+ ],
767
+ "closure_tests": [
768
+ "verify constant_dependence_diff matches computed value",
769
+ "limit case γ→0 reduces to identity",
770
+ ...
771
+ ],
772
+ "algebra_sanity": {
773
+ "dimension_table": {"<symbol>": "<type_signature>", ...},
774
+ "power_count": "<one-line check>",
775
+ "zero_coupling_check": "<one-line check>",
776
+ "constant_dependence_diff": "<before vs after>"
777
+ }
778
+ }
779
+ ],
780
+ "deep_fix_status": "ok" | "unavailable"
781
+ }
782
+ ```
783
+
784
+ Field semantics:
785
+ - Both `deep_fix_plans` and `deep_fix_status` are **omitted entirely** when the flag is not set. The default schema does not include either key.
786
+ - When the flag is set and reviewer returns well-formed plans, `deep_fix_status` is `"ok"` and `deep_fix_plans` mirrors `details.issues` one-to-one (each plan referenced by `issue_id`); `algebra_sanity` is present only for issues invoking Schur / Young / Cauchy-Schwarz / Hölder / quadratic-form / operator-norm / power-counting steps.
787
+ - When the flag is set but reviewer output is malformed or truncated, `deep_fix_status` is `"unavailable"` and `deep_fix_plans` is `[]`. Downstream consumers MUST treat `"unavailable"` identically to the field being absent: not blocking.
788
+ - Downstream consumers MUST treat absence of either field as the only valid default state and MUST NOT raise on missing.
789
+ - `deep_fix_plans` is advisory tooling for the executor; `verify_paper_audits.sh` and `paper-writing` Phase 6 do not block on its content or shape.
790
+
791
+ ### Optional: `details.restatement_drift` (only when `--restatement-check` is set)
792
+
793
+ ```json
794
+ "details": {
795
+ ...
796
+ "restatement_drift": [
797
+ {
798
+ "label": "thm:main",
799
+ "canonical_location": "sections/2.setup.tex:117",
800
+ "restatement_location": "sections/appendix.tex:1614",
801
+ "drift_type": "conditional_loss" | "scope_change" | "quantifier_loss" |
802
+ "regime_envelope_change" | "constant_change" | "variable_rename",
803
+ "canonical_excerpt": "<short LaTeX>",
804
+ "restatement_excerpt": "<short LaTeX>",
805
+ "severity": "MAJOR" | "MINOR" | "CRITICAL",
806
+ "note": "..."
807
+ }
808
+ ],
809
+ "restatement_check_status": "ok" | "unavailable"
810
+ }
811
+ ```
812
+
813
+ Field semantics:
814
+ - Both `restatement_drift` and `restatement_check_status` are **omitted entirely** when the flag is not set. The default schema does not include either key.
815
+ - When the flag is set and the cross-location scan completes, `restatement_check_status` is `"ok"` and `restatement_drift` lists detected pairs (possibly empty if none found).
816
+ - When the scan fails (per "Failure mode" in Phase 3.6), `restatement_check_status` is `"unavailable"` and `restatement_drift` is `[]`.
817
+ - Downstream consumers MUST treat absence or `"unavailable"` identically as the default state and MUST NOT raise on missing.
818
+ - `restatement_drift` does **not** alter `details.issues` and does **not** change the top-level `verdict` decision rule. The reviewer may at its discretion mirror a CRITICAL-severity drift into `details.issues` as a regular issue, but that is a per-issue judgment, not an automatic schema-driven rule.
819
+
820
+ ### `audited_input_hashes` scope
821
+
822
+ Hash the **declared input set** actually reviewed — the theorem-bearing
823
+ `.tex` files passed into this invocation — not a repo-wide union and not
824
+ the reviewer's self-reported opened subset. The external verifier rehashes
825
+ these entries; any mismatch flags `STALE`.
826
+
827
+ **Path convention** (must match `verify_paper_audits.sh`): keys are
828
+ **paths relative to the paper directory** (no `paper/` prefix — the
829
+ verifier resolves relative to the paper dir; prefixing produces
830
+ `paper/paper/...` and false-fails as STALE). Use **absolute paths** for
831
+ files outside the paper dir.
832
+
833
+ ### Verdict decision table
834
+
835
+ | Input state | Verdict | `reason_code` example |
836
+ |-------------------------------------------------------|------------------|-----------------------|
837
+ | No theorems / lemmas / proofs in paper | `NOT_APPLICABLE` | `no_theorems` |
838
+ | Theorems present but referenced files unreadable | `BLOCKED` | `source_unreadable` |
839
+ | All proof obligations discharged, no gaps | `PASS` | `all_proofs_complete` |
840
+ | Only MINOR issues (notation / exposition) | `WARN` | `minor_gaps` |
841
+ | Any FATAL or CRITICAL issue (logic gap, wrong claim) | `FAIL` | `critical_gap` |
842
+ | Reviewer invocation failed (network / malformed) | `ERROR` | `reviewer_error` |
843
+
844
+ MAJOR issues alone map to `WARN` or `FAIL` at the reviewer's discretion and
845
+ must carry an explicit justification in `summary` + `details.issues`.
846
+
847
+ ### Thread independence
848
+
849
+ Every **top-level** `/proof-checker` invocation starts a fresh reviewer thread. For codex this is `mcp__codex__codex`; for manual this is `mcp__manual_review__review`. Do not reuse a saved threadId across separate invocations of this skill. Within a single top-level invocation, the appropriate reply tool (`mcp__codex__codex-reply` or `mcp__manual_review__review_reply`) threads the Phase 3 follow-up rounds — the reviewer needs prior-issue context to judge whether a fix actually closed the gap, and the Phase 1→3 flow above explicitly relies on this. The Phase 3.5 "Independent second review for FATAL/CRITICAL fixes" sub-step is the deliberate exception inside a single run: it must spawn a fresh thread so the blind reviewer has no exposure to the original critique.
850
+
851
+ Do not accept prior audit outputs (PAPER_CLAIM_AUDIT, CITATION_AUDIT, EXPERIMENT_LOG) as input across separate invocations — the cross-run freshness is what preserves reviewer independence per `shared-references/reviewer-independence.md`.
852
+
853
+ This skill never blocks by itself; `paper-writing` Phase 6 plus the
854
+ verifier decide whether the verdict blocks finalization based on the
855
+ `assurance` level.
856
+
857
+ ## Example Invocations
858
+
859
+ ```
860
+ /proof-checker "neurips_2025.tex"
861
+ /proof-checker "check the GMM generalization proof, focus on dimension dependence"
862
+ /proof-checker "verify proof in paper.tex — difficulty: nightmare"
863
+ /proof-checker "paper/main.tex --deep-fix" # opt-in: ask reviewer to also emit repair-grade deep_fix_plans
864
+ /proof-checker "paper/main.tex --restatement-check" # opt-in: run Phase 3.6 to detect cross-location theorem-statement drift
865
+ /proof-checker "paper/main.tex --deep-fix --restatement-check" # both opt-ins, independent
866
+ ```