dsh-aris-panel 0.1.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/LICENSE +21 -0
- package/README.md +98 -0
- package/README_CN.md +87 -0
- package/dsh/checkout.patch.yml +38 -0
- package/dsh/client.js +634 -0
- package/dsh/cordis.patch.yml +44 -0
- package/dsh/index.mjs +76 -0
- package/dsh/run-status.mjs +182 -0
- package/dsh/scope-limits.mjs +50 -0
- package/dsh/workbench.mjs +291 -0
- package/mcp-servers/claude-review/README.md +93 -0
- package/mcp-servers/claude-review/run_with_claude_aws.sh +49 -0
- package/mcp-servers/claude-review/server.py +718 -0
- package/mcp-servers/codex-image2/README.md +65 -0
- package/mcp-servers/codex-image2/server.py +893 -0
- package/mcp-servers/feishu-bridge/requirements.txt +1 -0
- package/mcp-servers/feishu-bridge/server.py +240 -0
- package/mcp-servers/gemini-review/README.md +171 -0
- package/mcp-servers/gemini-review/server.py +1856 -0
- package/mcp-servers/llm-chat/requirements.txt +1 -0
- package/mcp-servers/llm-chat/server.py +664 -0
- package/mcp-servers/manual-review/README.md +133 -0
- package/mcp-servers/manual-review/server.py +910 -0
- package/mcp-servers/manual-review/ui.html +279 -0
- package/mcp-servers/minimax-chat/requirements.txt +1 -0
- package/mcp-servers/minimax-chat/server.py +381 -0
- package/package.json +51 -0
- package/skills/ablation-planner/SKILL.md +123 -0
- package/skills/alphaxiv/SKILL.md +196 -0
- package/skills/analyze-results/SKILL.md +46 -0
- package/skills/arxiv/SKILL.md +248 -0
- package/skills/auto-paper-improvement-loop/SKILL.md +651 -0
- package/skills/auto-review-loop/SKILL.md +1137 -0
- package/skills/auto-review-loop-llm/SKILL.md +259 -0
- package/skills/auto-review-loop-minimax/SKILL.md +302 -0
- package/skills/citation-audit/SKILL.md +502 -0
- package/skills/claims-drafting/SKILL.md +227 -0
- package/skills/comm-lit-review/SKILL.md +297 -0
- package/skills/deepxiv/SKILL.md +263 -0
- package/skills/dse-loop/SKILL.md +296 -0
- package/skills/embodiment-description/SKILL.md +129 -0
- package/skills/exa-search/SKILL.md +205 -0
- package/skills/experiment-audit/SKILL.md +311 -0
- package/skills/experiment-bridge/SKILL.md +376 -0
- package/skills/experiment-plan/SKILL.md +249 -0
- package/skills/experiment-queue/SKILL.md +431 -0
- package/skills/experiment-queue/scripts/build_manifest.py +142 -0
- package/skills/experiment-queue/scripts/queue_manager.py +433 -0
- package/skills/feishu-notify/SKILL.md +156 -0
- package/skills/figure-description/SKILL.md +138 -0
- package/skills/figure-spec/SKILL.md +262 -0
- package/skills/figure-spec/scripts/figure_renderer.py +799 -0
- package/skills/formula-derivation/SKILL.md +280 -0
- package/skills/gemini-search/SKILL.md +231 -0
- package/skills/grant-proposal/SKILL.md +698 -0
- package/skills/idea-creator/SKILL.md +542 -0
- package/skills/idea-discovery/SKILL.md +521 -0
- package/skills/idea-discovery-robot/SKILL.md +363 -0
- package/skills/integrity-forensics/SKILL.md +284 -0
- package/skills/interview-cheatsheet/SKILL.md +245 -0
- package/skills/invention-structuring/SKILL.md +188 -0
- package/skills/jurisdiction-format/SKILL.md +192 -0
- package/skills/kill-argument/SKILL.md +437 -0
- package/skills/mermaid-diagram/SKILL.md +419 -0
- package/skills/meta-apply/SKILL.md +141 -0
- package/skills/meta-optimize/SKILL.md +437 -0
- package/skills/monitor-experiment/SKILL.md +140 -0
- package/skills/novelty-check/SKILL.md +101 -0
- package/skills/openalex/SKILL.md +237 -0
- package/skills/overleaf-sync/SKILL.md +220 -0
- package/skills/paper-claim-audit/SKILL.md +348 -0
- package/skills/paper-compile/SKILL.md +266 -0
- package/skills/paper-figure/SKILL.md +312 -0
- package/skills/paper-illustration/SKILL.md +736 -0
- package/skills/paper-illustration-image2/SKILL.md +391 -0
- package/skills/paper-illustration-image2/scripts/paper_illustration_image2.py +255 -0
- package/skills/paper-plan/SKILL.md +386 -0
- package/skills/paper-poster/SKILL.md +19 -0
- package/skills/paper-poster-html/DESIGN_FINAL.md +176 -0
- package/skills/paper-poster-html/IMPLEMENTATION_CONVENTIONS.md +161 -0
- package/skills/paper-poster-html/LICENSES/posterly-MIT.txt +21 -0
- package/skills/paper-poster-html/NOTICE.md +57 -0
- package/skills/paper-poster-html/SKILL.md +323 -0
- package/skills/paper-poster-html/scripts/_posterly/__init__.py +0 -0
- package/skills/paper-poster-html/scripts/_posterly/canvas.py +200 -0
- package/skills/paper-poster-html/scripts/_posterly/measure.py +588 -0
- package/skills/paper-poster-html/scripts/_posterly/polish.py +498 -0
- package/skills/paper-poster-html/scripts/_posterly/preflight.py +489 -0
- package/skills/paper-poster-html/scripts/_posterly/render.py +215 -0
- package/skills/paper-poster-html/scripts/_posterly/textutil.py +16 -0
- package/skills/paper-poster-html/scripts/_posterly/verify_final.py +171 -0
- package/skills/paper-poster-html/scripts/asset_check.py +897 -0
- package/skills/paper-poster-html/scripts/extract_pdf_figures.py +666 -0
- package/skills/paper-poster-html/scripts/poster_check.py +251 -0
- package/skills/paper-poster-html/scripts/preprocess_figures.py +238 -0
- package/skills/paper-poster-html/scripts/render_preview.py +217 -0
- package/skills/paper-poster-html/scripts/run_gates.py +556 -0
- package/skills/paper-poster-html/scripts/style_check.py +1324 -0
- package/skills/paper-poster-html/templates/COMPONENTS.md +462 -0
- package/skills/paper-poster-html/templates/README.md +170 -0
- package/skills/paper-poster-html/templates/landscape_4col.html +1032 -0
- package/skills/paper-poster-html/templates/landscape_hero.html +1046 -0
- package/skills/paper-poster-html/templates/portrait_2col.html +947 -0
- package/skills/paper-poster-html/templates/tokens/acl.json +9 -0
- package/skills/paper-poster-html/templates/tokens/cvpr.json +9 -0
- package/skills/paper-poster-html/templates/tokens/generic.json +9 -0
- package/skills/paper-poster-html/templates/tokens/iclr.json +9 -0
- package/skills/paper-poster-html/templates/tokens/icml.json +9 -0
- package/skills/paper-poster-html/templates/tokens/neurips.json +9 -0
- package/skills/paper-slides/SKILL.md +635 -0
- package/skills/paper-talk/SKILL.md +381 -0
- package/skills/paper-write/SKILL.md +604 -0
- package/skills/paper-write/templates/IEEEtran.bst +2409 -0
- package/skills/paper-write/templates/IEEEtran.cls +6347 -0
- package/skills/paper-write/templates/iclr2026.tex +84 -0
- package/skills/paper-write/templates/icml2025.tex +87 -0
- package/skills/paper-write/templates/ieee_conference.tex +89 -0
- package/skills/paper-write/templates/ieee_journal.tex +93 -0
- package/skills/paper-write/templates/math_commands.tex +48 -0
- package/skills/paper-write/templates/neurips2025.tex +80 -0
- package/skills/paper-writing/SKILL.md +916 -0
- package/skills/patent-novelty-check/SKILL.md +153 -0
- package/skills/patent-pipeline/SKILL.md +344 -0
- package/skills/patent-review/SKILL.md +203 -0
- package/skills/pixel-art/SKILL.md +137 -0
- package/skills/prior-art-search/SKILL.md +146 -0
- package/skills/proof-checker/SKILL.md +866 -0
- package/skills/proof-orchestrator/NOTICE.md +24 -0
- package/skills/proof-orchestrator/SKILL.md +254 -0
- package/skills/proof-orchestrator/references/audit-output-contract.md +126 -0
- package/skills/proof-orchestrator/references/deepseek-routing.md +74 -0
- package/skills/proof-orchestrator/references/dispatch-prompts.md +227 -0
- package/skills/proof-orchestrator/references/notation-audit.md +135 -0
- package/skills/proof-orchestrator/references/proof-audit-rubric.md +70 -0
- package/skills/proof-orchestrator/references/stress-tests.md +38 -0
- package/skills/proof-writer/SKILL.md +223 -0
- package/skills/qzcli/SKILL.md +324 -0
- package/skills/rebuttal/SKILL.md +376 -0
- package/skills/render-html/SKILL.md +316 -0
- package/skills/render-html/scripts/render_html.py +1006 -0
- package/skills/render-html/scripts/templates/academic.html +703 -0
- package/skills/render-html/scripts/templates/dashboard.html +333 -0
- package/skills/research-lit/SKILL.md +756 -0
- package/skills/research-pipeline/SKILL.md +384 -0
- package/skills/research-refine/SKILL.md +770 -0
- package/skills/research-refine-pipeline/SKILL.md +186 -0
- package/skills/research-review/SKILL.md +198 -0
- package/skills/research-wiki/SKILL.md +461 -0
- package/skills/resubmit-pipeline/SKILL.md +447 -0
- package/skills/result-to-claim/SKILL.md +311 -0
- package/skills/run-experiment/SKILL.md +313 -0
- package/skills/semantic-scholar/SKILL.md +236 -0
- package/skills/serverless-modal/SKILL.md +335 -0
- package/skills/shared-references/acceptance-gate.md +324 -0
- package/skills/shared-references/assurance-contract.md +248 -0
- package/skills/shared-references/capture-antipatterns.md +78 -0
- package/skills/shared-references/citation-discipline.md +583 -0
- package/skills/shared-references/compute-env-contract.md +163 -0
- package/skills/shared-references/effort-contract.md +183 -0
- package/skills/shared-references/evidence-precheck.md +65 -0
- package/skills/shared-references/experiment-integrity.md +49 -0
- package/skills/shared-references/external-cadence.md +326 -0
- package/skills/shared-references/fan-out-pattern.md +366 -0
- package/skills/shared-references/injection-hygiene.md +127 -0
- package/skills/shared-references/integration-contract.md +461 -0
- package/skills/shared-references/output-composition.md +93 -0
- package/skills/shared-references/output-language.md +45 -0
- package/skills/shared-references/output-manifest.md +49 -0
- package/skills/shared-references/output-versioning.md +111 -0
- package/skills/shared-references/patent-format-cn.md +199 -0
- package/skills/shared-references/patent-format-ep.md +173 -0
- package/skills/shared-references/patent-format-us.md +161 -0
- package/skills/shared-references/patent-writing-principles.md +197 -0
- package/skills/shared-references/prior-art-databases.md +141 -0
- package/skills/shared-references/resumable-runs.md +109 -0
- package/skills/shared-references/review-scope-limits.md +81 -0
- package/skills/shared-references/review-tracing.md +391 -0
- package/skills/shared-references/reviewer-independence.md +79 -0
- package/skills/shared-references/reviewer-routing.md +852 -0
- package/skills/shared-references/skill-governance.md +104 -0
- package/skills/shared-references/taste-calibration.md +85 -0
- package/skills/shared-references/venue-checklists.md +114 -0
- package/skills/shared-references/wiki-helper-resolution.md +134 -0
- package/skills/shared-references/writing-principles.md +525 -0
- package/skills/skills-codex/README.md +102 -0
- package/skills/skills-codex/README_CN.md +100 -0
- package/skills/skills-codex/ablation-planner/SKILL.md +126 -0
- package/skills/skills-codex/alphaxiv/SKILL.md +186 -0
- package/skills/skills-codex/analyze-results/SKILL.md +45 -0
- package/skills/skills-codex/arxiv/SKILL.md +210 -0
- package/skills/skills-codex/auto-paper-improvement-loop/SKILL.md +574 -0
- package/skills/skills-codex/auto-review-loop/SKILL.md +500 -0
- package/skills/skills-codex/auto-review-loop-llm/SKILL.md +247 -0
- package/skills/skills-codex/auto-review-loop-minimax/SKILL.md +290 -0
- package/skills/skills-codex/citation-audit/SKILL.md +504 -0
- package/skills/skills-codex/claims-drafting/SKILL.md +239 -0
- package/skills/skills-codex/comm-lit-review/SKILL.md +299 -0
- package/skills/skills-codex/comm-lit-review/references/domain-taxonomy.md +57 -0
- package/skills/skills-codex/comm-lit-review/references/output-template.md +37 -0
- package/skills/skills-codex/comm-lit-review/references/source-policy.md +99 -0
- package/skills/skills-codex/comm-lit-review/references/venue-tiering.md +112 -0
- package/skills/skills-codex/deepxiv/SKILL.md +142 -0
- package/skills/skills-codex/dse-loop/SKILL.md +285 -0
- package/skills/skills-codex/embodiment-description/SKILL.md +129 -0
- package/skills/skills-codex/exa-search/SKILL.md +192 -0
- package/skills/skills-codex/experiment-audit/SKILL.md +286 -0
- package/skills/skills-codex/experiment-bridge/SKILL.md +356 -0
- package/skills/skills-codex/experiment-plan/SKILL.md +249 -0
- package/skills/skills-codex/experiment-queue/SKILL.md +401 -0
- package/skills/skills-codex/feishu-notify/SKILL.md +155 -0
- package/skills/skills-codex/figure-description/SKILL.md +138 -0
- package/skills/skills-codex/figure-spec/SKILL.md +252 -0
- package/skills/skills-codex/formula-derivation/SKILL.md +280 -0
- package/skills/skills-codex/gemini-search/SKILL.md +205 -0
- package/skills/skills-codex/grant-proposal/SKILL.md +626 -0
- package/skills/skills-codex/idea-creator/SKILL.md +405 -0
- package/skills/skills-codex/idea-discovery/SKILL.md +475 -0
- package/skills/skills-codex/idea-discovery-robot/SKILL.md +362 -0
- package/skills/skills-codex/integrity-forensics/SKILL.md +106 -0
- package/skills/skills-codex/interview-cheatsheet/SKILL.md +245 -0
- package/skills/skills-codex/invention-structuring/SKILL.md +188 -0
- package/skills/skills-codex/jurisdiction-format/SKILL.md +192 -0
- package/skills/skills-codex/kill-argument/SKILL.md +403 -0
- package/skills/skills-codex/mermaid-diagram/SKILL.md +379 -0
- package/skills/skills-codex/meta-apply/SKILL.md +154 -0
- package/skills/skills-codex/meta-optimize/SKILL.md +348 -0
- package/skills/skills-codex/monitor-experiment/SKILL.md +98 -0
- package/skills/skills-codex/novelty-check/SKILL.md +89 -0
- package/skills/skills-codex/openalex/SKILL.md +228 -0
- package/skills/skills-codex/overleaf-sync/SKILL.md +220 -0
- package/skills/skills-codex/paper-claim-audit/SKILL.md +350 -0
- package/skills/skills-codex/paper-compile/SKILL.md +253 -0
- package/skills/skills-codex/paper-figure/SKILL.md +311 -0
- package/skills/skills-codex/paper-illustration/SKILL.md +690 -0
- package/skills/skills-codex/paper-illustration-image2/SKILL.md +383 -0
- package/skills/skills-codex/paper-illustration-image2/scripts/paper_illustration_image2.py +255 -0
- package/skills/skills-codex/paper-plan/SKILL.md +278 -0
- package/skills/skills-codex/paper-poster/SKILL.md +19 -0
- package/skills/skills-codex/paper-poster-html/SKILL.md +377 -0
- package/skills/skills-codex/paper-slides/SKILL.md +571 -0
- package/skills/skills-codex/paper-talk/SKILL.md +381 -0
- package/skills/skills-codex/paper-write/SKILL.md +411 -0
- package/skills/skills-codex/paper-write/templates/IEEEtran.bst +2409 -0
- package/skills/skills-codex/paper-write/templates/IEEEtran.cls +6347 -0
- package/skills/skills-codex/paper-write/templates/aaai2026.bst +1493 -0
- package/skills/skills-codex/paper-write/templates/aaai2026.sty +315 -0
- package/skills/skills-codex/paper-write/templates/aaai2026.tex +952 -0
- package/skills/skills-codex/paper-write/templates/acl.sty +312 -0
- package/skills/skills-codex/paper-write/templates/acl2026.tex +377 -0
- package/skills/skills-codex/paper-write/templates/acl_natbib.bst +1940 -0
- package/skills/skills-codex/paper-write/templates/acm.bst +3081 -0
- package/skills/skills-codex/paper-write/templates/acm_mm2026.tex +204 -0
- package/skills/skills-codex/paper-write/templates/acmart.cls +3520 -0
- package/skills/skills-codex/paper-write/templates/cvpr.bst +1448 -0
- package/skills/skills-codex/paper-write/templates/cvpr.sty +508 -0
- package/skills/skills-codex/paper-write/templates/cvpr2026.tex +63 -0
- package/skills/skills-codex/paper-write/templates/iclr2026.tex +84 -0
- package/skills/skills-codex/paper-write/templates/iclr2026_conference.bst +1440 -0
- package/skills/skills-codex/paper-write/templates/iclr2026_conference.sty +246 -0
- package/skills/skills-codex/paper-write/templates/icml2026.sty +767 -0
- package/skills/skills-codex/paper-write/templates/icml2026.tex +662 -0
- package/skills/skills-codex/paper-write/templates/ieee_conference.tex +89 -0
- package/skills/skills-codex/paper-write/templates/ieee_journal.tex +93 -0
- package/skills/skills-codex/paper-write/templates/math_commands.tex +48 -0
- package/skills/skills-codex/paper-write/templates/neurips2026.tex +493 -0
- package/skills/skills-codex/paper-write/templates/neurips_2026.sty +437 -0
- package/skills/skills-codex/paper-writing/SKILL.md +731 -0
- package/skills/skills-codex/patent-novelty-check/SKILL.md +153 -0
- package/skills/skills-codex/patent-pipeline/SKILL.md +344 -0
- package/skills/skills-codex/patent-review/SKILL.md +202 -0
- package/skills/skills-codex/pixel-art/SKILL.md +139 -0
- package/skills/skills-codex/prior-art-search/SKILL.md +146 -0
- package/skills/skills-codex/proof-checker/SKILL.md +554 -0
- package/skills/skills-codex/proof-orchestrator/SKILL.md +260 -0
- package/skills/skills-codex/proof-orchestrator/references/audit-output-contract.md +126 -0
- package/skills/skills-codex/proof-orchestrator/references/deepseek-routing.md +76 -0
- package/skills/skills-codex/proof-orchestrator/references/dispatch-prompts.md +227 -0
- package/skills/skills-codex/proof-orchestrator/references/notation-audit.md +135 -0
- package/skills/skills-codex/proof-orchestrator/references/proof-audit-rubric.md +70 -0
- package/skills/skills-codex/proof-orchestrator/references/stress-tests.md +38 -0
- package/skills/skills-codex/proof-writer/SKILL.md +222 -0
- package/skills/skills-codex/qzcli/SKILL.md +324 -0
- package/skills/skills-codex/rebuttal/SKILL.md +305 -0
- package/skills/skills-codex/render-html/SKILL.md +305 -0
- package/skills/skills-codex/render-html/scripts/__pycache__/render_html.cpython-314.pyc +0 -0
- package/skills/skills-codex/render-html/scripts/render_html.py +909 -0
- package/skills/skills-codex/render-html/scripts/templates/academic.html +342 -0
- package/skills/skills-codex/render-html/scripts/templates/dashboard.html +333 -0
- package/skills/skills-codex/research-lit/SKILL.md +464 -0
- package/skills/skills-codex/research-pipeline/SKILL.md +340 -0
- package/skills/skills-codex/research-refine/SKILL.md +721 -0
- package/skills/skills-codex/research-refine-pipeline/SKILL.md +186 -0
- package/skills/skills-codex/research-review/SKILL.md +135 -0
- package/skills/skills-codex/research-wiki/SKILL.md +421 -0
- package/skills/skills-codex/resubmit-pipeline/SKILL.md +444 -0
- package/skills/skills-codex/result-to-claim/SKILL.md +246 -0
- package/skills/skills-codex/run-experiment/SKILL.md +236 -0
- package/skills/skills-codex/semantic-scholar/SKILL.md +219 -0
- package/skills/skills-codex/serverless-modal/SKILL.md +335 -0
- package/skills/skills-codex/shared-references/acceptance-gate.md +336 -0
- package/skills/skills-codex/shared-references/assurance-contract.md +139 -0
- package/skills/skills-codex/shared-references/capture-antipatterns.md +84 -0
- package/skills/skills-codex/shared-references/citation-discipline.md +452 -0
- package/skills/skills-codex/shared-references/compute-env-contract.md +163 -0
- package/skills/skills-codex/shared-references/effort-contract.md +143 -0
- package/skills/skills-codex/shared-references/evidence-precheck.md +73 -0
- package/skills/skills-codex/shared-references/experiment-integrity.md +49 -0
- package/skills/skills-codex/shared-references/external-cadence.md +334 -0
- package/skills/skills-codex/shared-references/fan-out-pattern.md +375 -0
- package/skills/skills-codex/shared-references/injection-hygiene.md +132 -0
- package/skills/skills-codex/shared-references/integration-contract.md +372 -0
- package/skills/skills-codex/shared-references/output-composition.md +98 -0
- package/skills/skills-codex/shared-references/output-language.md +45 -0
- package/skills/skills-codex/shared-references/output-manifest.md +40 -0
- package/skills/skills-codex/shared-references/output-versioning.md +111 -0
- package/skills/skills-codex/shared-references/patent-format-cn.md +199 -0
- package/skills/skills-codex/shared-references/patent-format-ep.md +173 -0
- package/skills/skills-codex/shared-references/patent-format-us.md +161 -0
- package/skills/skills-codex/shared-references/patent-writing-principles.md +197 -0
- package/skills/skills-codex/shared-references/prior-art-databases.md +141 -0
- package/skills/skills-codex/shared-references/resumable-runs.md +125 -0
- package/skills/skills-codex/shared-references/review-scope-limits.md +81 -0
- package/skills/skills-codex/shared-references/review-tracing.md +144 -0
- package/skills/skills-codex/shared-references/reviewer-independence.md +66 -0
- package/skills/skills-codex/shared-references/reviewer-routing.md +128 -0
- package/skills/skills-codex/shared-references/skill-governance.md +119 -0
- package/skills/skills-codex/shared-references/taste-calibration.md +90 -0
- package/skills/skills-codex/shared-references/venue-checklists.md +73 -0
- package/skills/skills-codex/shared-references/wiki-helper-resolution.md +69 -0
- package/skills/skills-codex/shared-references/writing-principles.md +525 -0
- package/skills/skills-codex/slides-polish/SKILL.md +563 -0
- package/skills/skills-codex/specification-writing/SKILL.md +211 -0
- package/skills/skills-codex/system-profile/SKILL.md +103 -0
- package/skills/skills-codex/training-check/SKILL.md +83 -0
- package/skills/skills-codex/vast-gpu/SKILL.md +394 -0
- package/skills/skills-codex/web-debug-search/SKILL.md +334 -0
- package/skills/skills-codex/wiki-enrich/SKILL.md +255 -0
- package/skills/skills-codex/writing-systems-papers/SKILL.md +184 -0
- package/skills/skills-codex-claude-review/README.md +79 -0
- package/skills/skills-codex-claude-review/README_CN.md +78 -0
- package/skills/skills-codex-claude-review/auto-paper-improvement-loop/SKILL.md +581 -0
- package/skills/skills-codex-claude-review/auto-review-loop/SKILL.md +510 -0
- package/skills/skills-codex-claude-review/novelty-check/SKILL.md +102 -0
- package/skills/skills-codex-claude-review/paper-figure/SKILL.md +319 -0
- package/skills/skills-codex-claude-review/paper-plan/SKILL.md +287 -0
- package/skills/skills-codex-claude-review/paper-write/SKILL.md +420 -0
- package/skills/skills-codex-claude-review/research-refine/SKILL.md +732 -0
- package/skills/skills-codex-claude-review/research-review/SKILL.md +149 -0
- package/skills/skills-codex-gemini-review/README.md +176 -0
- package/skills/skills-codex-gemini-review/README_CN.md +175 -0
- package/skills/skills-codex-gemini-review/auto-paper-improvement-loop/SKILL.md +331 -0
- package/skills/skills-codex-gemini-review/auto-review-loop/SKILL.md +304 -0
- package/skills/skills-codex-gemini-review/grant-proposal/SKILL.md +630 -0
- package/skills/skills-codex-gemini-review/idea-creator/SKILL.md +263 -0
- package/skills/skills-codex-gemini-review/idea-discovery/SKILL.md +275 -0
- package/skills/skills-codex-gemini-review/idea-discovery-robot/SKILL.md +365 -0
- package/skills/skills-codex-gemini-review/novelty-check/SKILL.md +92 -0
- package/skills/skills-codex-gemini-review/paper-figure/SKILL.md +289 -0
- package/skills/skills-codex-gemini-review/paper-plan/SKILL.md +265 -0
- package/skills/skills-codex-gemini-review/paper-poster-html/SKILL.md +102 -0
- package/skills/skills-codex-gemini-review/paper-slides/SKILL.md +582 -0
- package/skills/skills-codex-gemini-review/paper-write/SKILL.md +346 -0
- package/skills/skills-codex-gemini-review/paper-writing/SKILL.md +312 -0
- package/skills/skills-codex-gemini-review/research-refine/SKILL.md +674 -0
- package/skills/skills-codex-gemini-review/research-review/SKILL.md +112 -0
- package/skills/slides-polish/SKILL.md +565 -0
- package/skills/specification-writing/SKILL.md +211 -0
- package/skills/system-profile/SKILL.md +103 -0
- package/skills/training-check/SKILL.md +132 -0
- package/skills/vast-gpu/SKILL.md +394 -0
- package/skills/web-debug-search/SKILL.md +334 -0
- package/skills/wiki-enrich/SKILL.md +257 -0
- package/skills/writing-systems-papers/SKILL.md +184 -0
- package/templates/CLAUDE_MD_TEMPLATE.md +29 -0
- package/templates/EXPERIMENT_LOG_TEMPLATE.md +47 -0
- package/templates/EXPERIMENT_PLAN_TEMPLATE.md +51 -0
- package/templates/EXPERIMENT_PLAN_TEMPLATE_CN.md +53 -0
- package/templates/FINDINGS_TEMPLATE.md +52 -0
- package/templates/IDEA_CANDIDATES_TEMPLATE.md +47 -0
- package/templates/IDEA_CANDIDATES_TEMPLATE_CN.md +47 -0
- package/templates/INVENTION_BRIEF_TEMPLATE.md +87 -0
- package/templates/MANIFEST_TEMPLATE.md +7 -0
- package/templates/NARRATIVE_REPORT_TEMPLATE.md +49 -0
- package/templates/PAPER_PLAN_TEMPLATE.md +47 -0
- package/templates/PATENT_CLAIMS_TEMPLATE.md +78 -0
- package/templates/PATENT_SPECIFICATION_TEMPLATE.md +68 -0
- package/templates/README.md +57 -0
- package/templates/RESEARCH_BRIEF_TEMPLATE.md +35 -0
- package/templates/RESEARCH_BRIEF_TEMPLATE_CN.md +41 -0
- package/templates/RESEARCH_CONTRACT_TEMPLATE.md +60 -0
- package/templates/claude-hooks/corpus_write_guard.json +16 -0
- package/templates/claude-hooks/corpus_write_guard.py +85 -0
- package/templates/claude-hooks/meta_logging.json +74 -0
- package/templates/gitignore-trace.txt +3 -0
- package/tools/__pycache__/check_skills_inventory.cpython-314.pyc +0 -0
- package/tools/arxiv_fetch.py +311 -0
- package/tools/capture_filter.py +126 -0
- package/tools/check_skills_inventory.py +273 -0
- package/tools/convert_skills_to_llm_chat.py +282 -0
- package/tools/copilot_native_evidence.py +818 -0
- package/tools/deepxiv_fetch.py +213 -0
- package/tools/evidence_check.py +212 -0
- package/tools/exa_search.py +425 -0
- package/tools/experiment_queue/README.md +118 -0
- package/tools/experiment_queue/build_manifest.py +44 -0
- package/tools/experiment_queue/queue_manager.py +44 -0
- package/tools/extract_paper_style.py +560 -0
- package/tools/figure_renderer.py +69 -0
- package/tools/forensics_gate.py +669 -0
- package/tools/generate_codex_claude_review_overrides.py +299 -0
- package/tools/idea_discovery_gate.py +256 -0
- package/tools/install_aris.ps1 +1372 -0
- package/tools/install_aris.sh +1370 -0
- package/tools/install_aris_codex.sh +1023 -0
- package/tools/install_aris_copilot.sh +1052 -0
- package/tools/iteration_log.py +143 -0
- package/tools/lint_skills_helpers.sh +84 -0
- package/tools/meta_opt/check_ready.sh +80 -0
- package/tools/meta_opt/log_event.sh +91 -0
- package/tools/meta_opt/trigger_eval.py +280 -0
- package/tools/meta_opt/trigger_evals.sample.json +28 -0
- package/tools/openalex_fetch.py +326 -0
- package/tools/overleaf_audit.sh +104 -0
- package/tools/overleaf_setup.sh +150 -0
- package/tools/paper_illustration_image2.py +62 -0
- package/tools/provenance.py +294 -0
- package/tools/research_wiki.py +1720 -0
- package/tools/review_gate.py +502 -0
- package/tools/run_state.py +399 -0
- package/tools/save_trace.sh +477 -0
- package/tools/semantic_scholar_fetch.py +438 -0
- package/tools/skill-groups.tsv +116 -0
- package/tools/skill_picker.py +238 -0
- package/tools/smart_update.ps1 +521 -0
- package/tools/smart_update.sh +591 -0
- package/tools/smart_update_codex.sh +419 -0
- package/tools/smart_update_copilot.sh +605 -0
- package/tools/threat_scan.py +222 -0
- package/tools/verify_paper_audits.sh +487 -0
- package/tools/verify_papers.py +613 -0
- package/tools/verify_wiki_coverage.sh +176 -0
- package/tools/watchdog.py +485 -0
|
@@ -0,0 +1,331 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: "auto-paper-improvement-loop"
|
|
3
|
+
description: "Autonomously improve a generated paper via Gemini review through gemini-review MCP → implement fixes → recompile, for 2 rounds. Use when user says \"改论文\", \"improve paper\", \"论文润色循环\", \"auto improve\", or wants to iteratively polish a generated paper."
|
|
4
|
+
---
|
|
5
|
+
|
|
6
|
+
> Override for Codex users who want **Gemini**, not a second Codex agent, to act as the reviewer. Install this package **after** `skills/skills-codex/*`.
|
|
7
|
+
|
|
8
|
+
# Auto Paper Improvement Loop: Review → Fix → Recompile
|
|
9
|
+
|
|
10
|
+
> **Gemini overlay assurance:** `review_independence: cross-family` and `acceptance_status: accepted`.
|
|
11
|
+
|
|
12
|
+
Autonomously improve the paper at: **$ARGUMENTS**
|
|
13
|
+
|
|
14
|
+
## Context
|
|
15
|
+
|
|
16
|
+
This skill is designed to run **after** Workflow 3 (`/paper-plan` → `/paper-figure` → `/paper-write` → `/paper-compile`). It takes a compiled paper and iteratively improves it through external LLM review.
|
|
17
|
+
|
|
18
|
+
Unlike `/auto-review-loop` (which iterates on **research** — running experiments, collecting data, rewriting narrative), this skill iterates on **paper writing quality** — fixing theoretical inconsistencies, softening overclaims, adding missing content, and improving presentation.
|
|
19
|
+
|
|
20
|
+
## Constants
|
|
21
|
+
|
|
22
|
+
- **MAX_ROUNDS = 2** — Two rounds of review→fix→recompile. Empirically, Round 1 catches structural issues (4→6/10), Round 2 catches remaining presentation issues (6→7/10). Diminishing returns beyond 2 rounds for writing-only improvements.
|
|
23
|
+
- **REVIEWER_MODEL = `gemini-review`** — Gemini reviewer invoked through the local `gemini-review` MCP bridge. Set `GEMINI_REVIEW_MODEL` if you need a specific Gemini model override.
|
|
24
|
+
- **REVIEW_LOG = `PAPER_IMPROVEMENT_LOG.md`** — Cumulative log of all rounds, stored in paper directory.
|
|
25
|
+
- **HUMAN_CHECKPOINT = false** — When `true`, pause after each round's review and present score + weaknesses to the user. The user can approve fixes, provide custom modification instructions, skip specific fixes, or stop early. When `false` (default), runs fully autonomously.
|
|
26
|
+
|
|
27
|
+
> 💡 Override: `/auto-paper-improvement-loop "paper/" — human checkpoint: true`
|
|
28
|
+
|
|
29
|
+
## Inputs
|
|
30
|
+
|
|
31
|
+
1. **Compiled paper** — `paper/main.pdf` + LaTeX source files
|
|
32
|
+
2. **All section `.tex` files** — concatenated for review prompt
|
|
33
|
+
|
|
34
|
+
## State Persistence (Compact Recovery)
|
|
35
|
+
|
|
36
|
+
If the context window fills up mid-loop, Codex auto-compacts. To recover, this skill writes `PAPER_IMPROVEMENT_STATE.json` after each round:
|
|
37
|
+
|
|
38
|
+
```json
|
|
39
|
+
{
|
|
40
|
+
"current_round": 1,
|
|
41
|
+
"thread_id": "019ce736-...",
|
|
42
|
+
"last_score": 6,
|
|
43
|
+
"status": "in_progress",
|
|
44
|
+
"timestamp": "2026-03-13T21:00:00"
|
|
45
|
+
}
|
|
46
|
+
```
|
|
47
|
+
|
|
48
|
+
**On startup**: if `PAPER_IMPROVEMENT_STATE.json` exists with `"status": "in_progress"` AND `timestamp` is within 24 hours, read it + `PAPER_IMPROVEMENT_LOG.md` to recover context, then resume from the next round. Otherwise (file absent, `"status": "completed"`, or older than 24 hours), start fresh.
|
|
49
|
+
|
|
50
|
+
**After each round**: overwrite the state file. **On completion**: set `"status": "completed"`.
|
|
51
|
+
|
|
52
|
+
## Workflow
|
|
53
|
+
|
|
54
|
+
### Step 0: Preserve Original
|
|
55
|
+
|
|
56
|
+
```bash
|
|
57
|
+
cp paper/main.pdf paper/main_round0_original.pdf
|
|
58
|
+
```
|
|
59
|
+
|
|
60
|
+
### Step 1: Collect Paper Text
|
|
61
|
+
|
|
62
|
+
Concatenate all section files into a single text block for the review prompt:
|
|
63
|
+
|
|
64
|
+
```bash
|
|
65
|
+
# Collect all sections in order
|
|
66
|
+
for f in paper/sections/*.tex; do
|
|
67
|
+
echo "% === $(basename $f) ==="
|
|
68
|
+
cat "$f"
|
|
69
|
+
done > /tmp/paper_full_text.txt
|
|
70
|
+
```
|
|
71
|
+
|
|
72
|
+
### Step 2: Round 1 Review
|
|
73
|
+
|
|
74
|
+
Send the full paper text to Gemini review:
|
|
75
|
+
|
|
76
|
+
```
|
|
77
|
+
mcp__gemini-review__review_start:
|
|
78
|
+
prompt: |
|
|
79
|
+
You are reviewing a [VENUE] paper. Please provide a detailed, structured review.
|
|
80
|
+
|
|
81
|
+
## Full Paper Text:
|
|
82
|
+
[paste concatenated sections]
|
|
83
|
+
|
|
84
|
+
## Review Instructions
|
|
85
|
+
Please act as a senior ML reviewer ([VENUE] level). Provide:
|
|
86
|
+
1. **Overall Score** (1-10, where 6 = weak accept, 7 = accept)
|
|
87
|
+
2. **Summary** (2-3 sentences)
|
|
88
|
+
3. **Strengths** (bullet list, ranked)
|
|
89
|
+
4. **Weaknesses** (bullet list, ranked: CRITICAL > MAJOR > MINOR)
|
|
90
|
+
5. **For each CRITICAL/MAJOR weakness**: A specific, actionable fix
|
|
91
|
+
6. **Missing References** (if any)
|
|
92
|
+
7. **Verdict**: Ready for submission? Yes / Almost / No
|
|
93
|
+
|
|
94
|
+
Focus on: theoretical rigor, claims vs evidence alignment, writing clarity,
|
|
95
|
+
self-containedness, notation consistency.
|
|
96
|
+
```
|
|
97
|
+
|
|
98
|
+
After this start call, immediately save the returned `jobId` and poll `mcp__gemini-review__review_status` with a bounded `waitSeconds` until `done=true`. Treat the completed status payload's `response` as the reviewer output, and save the completed `threadId` for any follow-up round.
|
|
99
|
+
|
|
100
|
+
Save the returned `jobId`, poll `mcp__gemini-review__review_status` until `done=true`, then save the completed `threadId` for Round 2.
|
|
101
|
+
|
|
102
|
+
### Step 2b: Human Checkpoint (if enabled)
|
|
103
|
+
|
|
104
|
+
**Skip if `HUMAN_CHECKPOINT = false`.**
|
|
105
|
+
|
|
106
|
+
Present the review results and wait for user input:
|
|
107
|
+
|
|
108
|
+
```
|
|
109
|
+
📋 Round 1 review complete.
|
|
110
|
+
|
|
111
|
+
Score: X/10 — [verdict]
|
|
112
|
+
Key weaknesses (by severity):
|
|
113
|
+
1. [CRITICAL] ...
|
|
114
|
+
2. [MAJOR] ...
|
|
115
|
+
3. [MINOR] ...
|
|
116
|
+
|
|
117
|
+
Reply "go" to implement all fixes, give custom instructions, "skip 2" to skip specific fixes, or "stop" to end.
|
|
118
|
+
```
|
|
119
|
+
|
|
120
|
+
Parse user response same as `/auto-review-loop`: approve / custom instructions / skip / stop.
|
|
121
|
+
|
|
122
|
+
### Step 3: Implement Round 1 Fixes
|
|
123
|
+
|
|
124
|
+
Parse the review and implement fixes by severity:
|
|
125
|
+
|
|
126
|
+
**Priority order:**
|
|
127
|
+
1. CRITICAL fixes (assumption mismatches, internal contradictions)
|
|
128
|
+
2. MAJOR fixes (overclaims, missing content, notation issues)
|
|
129
|
+
3. MINOR fixes (if time permits)
|
|
130
|
+
|
|
131
|
+
**Common fix patterns:**
|
|
132
|
+
|
|
133
|
+
| Issue | Fix Pattern |
|
|
134
|
+
|-------|-------------|
|
|
135
|
+
| Assumption-model mismatch | Rewrite assumption to match the model, add formal proposition bridging the gap |
|
|
136
|
+
| Overclaims | Soften language: "validate" → "demonstrate practical relevance", "comparable" → "qualitatively competitive" |
|
|
137
|
+
| Missing metrics | Add quantitative table with honest parameter counts and caveats |
|
|
138
|
+
| Theorem not self-contained | Add "Interpretation" paragraph listing all dependencies |
|
|
139
|
+
| Notation confusion | Rename conflicting symbols globally, add Notation paragraph |
|
|
140
|
+
| Missing references | Add to `references.bib`, cite in appropriate locations |
|
|
141
|
+
| Theory-practice gap | Explicitly frame theory as idealized; add synthetic validation subsection |
|
|
142
|
+
|
|
143
|
+
### Step 4: Recompile Round 1
|
|
144
|
+
|
|
145
|
+
```bash
|
|
146
|
+
cd paper && latexmk -C && latexmk -pdf -interaction=nonstopmode -halt-on-error main.tex
|
|
147
|
+
cp main.pdf main_round1.pdf
|
|
148
|
+
```
|
|
149
|
+
|
|
150
|
+
Verify: 0 undefined references, 0 undefined citations.
|
|
151
|
+
|
|
152
|
+
### Step 5: Round 2 Review
|
|
153
|
+
|
|
154
|
+
Use `mcp__gemini-review__review_reply_start` with the saved completed `threadId`:
|
|
155
|
+
|
|
156
|
+
```
|
|
157
|
+
mcp__gemini-review__review_reply_start:
|
|
158
|
+
threadId: [saved from Round 1]
|
|
159
|
+
prompt: |
|
|
160
|
+
[Round 2 update]
|
|
161
|
+
|
|
162
|
+
Since your last review, we have implemented:
|
|
163
|
+
1. [Fix 1]: [description]
|
|
164
|
+
2. [Fix 2]: [description]
|
|
165
|
+
...
|
|
166
|
+
|
|
167
|
+
Please re-score and re-assess. Same format:
|
|
168
|
+
Score, Summary, Strengths, Weaknesses, Actionable fixes, Verdict.
|
|
169
|
+
```
|
|
170
|
+
|
|
171
|
+
After this start call, immediately save the returned `jobId` and poll `mcp__gemini-review__review_status` with a bounded `waitSeconds` until `done=true`. Treat the completed status payload's `response` as the reviewer output, and save the completed `threadId` for any follow-up round.
|
|
172
|
+
|
|
173
|
+
### Step 5b: Human Checkpoint (if enabled)
|
|
174
|
+
|
|
175
|
+
**Skip if `HUMAN_CHECKPOINT = false`.** Same as Step 2b — present Round 2 review, wait for user input.
|
|
176
|
+
|
|
177
|
+
### Step 6: Implement Round 2 Fixes
|
|
178
|
+
|
|
179
|
+
Same process as Step 3. Typical Round 2 fixes:
|
|
180
|
+
- Add controlled synthetic experiments validating theory
|
|
181
|
+
- Further soften any remaining overclaims
|
|
182
|
+
- Formalize informal arguments (e.g., truncation → formal proposition)
|
|
183
|
+
- Strengthen limitations section
|
|
184
|
+
|
|
185
|
+
### Step 7: Recompile Round 2
|
|
186
|
+
|
|
187
|
+
```bash
|
|
188
|
+
cd paper && latexmk -C && latexmk -pdf -interaction=nonstopmode -halt-on-error main.tex
|
|
189
|
+
cp main.pdf main_round2.pdf
|
|
190
|
+
```
|
|
191
|
+
|
|
192
|
+
### Step 8: Format Check
|
|
193
|
+
|
|
194
|
+
After the final recompilation, run a format compliance check:
|
|
195
|
+
|
|
196
|
+
```bash
|
|
197
|
+
# 1. Page count vs venue limit
|
|
198
|
+
PAGES=$(pdfinfo paper/main.pdf | grep Pages | awk '{print $2}')
|
|
199
|
+
echo "Pages: $PAGES (limit: 9 main body for ICLR/NeurIPS)"
|
|
200
|
+
|
|
201
|
+
# 2. Overfull hbox warnings (content exceeding margins)
|
|
202
|
+
OVERFULL=$(grep -c "Overfull" paper/main.log 2>/dev/null || echo 0)
|
|
203
|
+
echo "Overfull hbox warnings: $OVERFULL"
|
|
204
|
+
grep "Overfull" paper/main.log 2>/dev/null | head -10
|
|
205
|
+
|
|
206
|
+
# 3. Underfull hbox warnings (loose spacing)
|
|
207
|
+
UNDERFULL=$(grep -c "Underfull" paper/main.log 2>/dev/null || echo 0)
|
|
208
|
+
echo "Underfull hbox warnings: $UNDERFULL"
|
|
209
|
+
|
|
210
|
+
# 4. Bad boxes summary
|
|
211
|
+
grep -c "badness" paper/main.log 2>/dev/null || echo "0 badness warnings"
|
|
212
|
+
```
|
|
213
|
+
|
|
214
|
+
**Auto-fix patterns:**
|
|
215
|
+
|
|
216
|
+
| Issue | Fix |
|
|
217
|
+
|-------|-----|
|
|
218
|
+
| Overfull hbox in equation | Wrap in `\resizebox` or split with `\split`/`aligned` |
|
|
219
|
+
| Overfull hbox in table | Reduce font (`\small`/`\footnotesize`) or use `\resizebox{\linewidth}{!}{...}` |
|
|
220
|
+
| Overfull hbox in text | Rephrase sentence or add `\allowbreak` / `\-` hints |
|
|
221
|
+
| Over page limit | Move content to appendix, compress tables, reduce figure sizes |
|
|
222
|
+
| Underfull hbox (loose) | Rephrase for better line filling or add `\looseness=-1` |
|
|
223
|
+
|
|
224
|
+
If any overfull hbox > 10pt is found, fix it and recompile before documenting.
|
|
225
|
+
|
|
226
|
+
### Step 9: Document Results
|
|
227
|
+
|
|
228
|
+
Create `PAPER_IMPROVEMENT_LOG.md` in the paper directory:
|
|
229
|
+
|
|
230
|
+
```markdown
|
|
231
|
+
# Paper Improvement Log
|
|
232
|
+
|
|
233
|
+
## Score Progression
|
|
234
|
+
|
|
235
|
+
| Round | Score | Verdict | Key Changes |
|
|
236
|
+
|-------|-------|---------|-------------|
|
|
237
|
+
| Round 0 (original) | X/10 | No/Almost/Yes | Baseline |
|
|
238
|
+
| Round 1 | Y/10 | No/Almost/Yes | [summary of fixes] |
|
|
239
|
+
| Round 2 | Z/10 | No/Almost/Yes | [summary of fixes] |
|
|
240
|
+
|
|
241
|
+
## Round 1 Review & Fixes
|
|
242
|
+
|
|
243
|
+
<details>
|
|
244
|
+
<summary>Gemini Review (Round 1)</summary>
|
|
245
|
+
|
|
246
|
+
[Full raw review text, verbatim]
|
|
247
|
+
|
|
248
|
+
</details>
|
|
249
|
+
|
|
250
|
+
### Fixes Implemented
|
|
251
|
+
1. [Fix description]
|
|
252
|
+
2. [Fix description]
|
|
253
|
+
...
|
|
254
|
+
|
|
255
|
+
## Round 2 Review & Fixes
|
|
256
|
+
|
|
257
|
+
<details>
|
|
258
|
+
<summary>Gemini Review (Round 2)</summary>
|
|
259
|
+
|
|
260
|
+
[Full raw review text, verbatim]
|
|
261
|
+
|
|
262
|
+
</details>
|
|
263
|
+
|
|
264
|
+
### Fixes Implemented
|
|
265
|
+
1. [Fix description]
|
|
266
|
+
2. [Fix description]
|
|
267
|
+
...
|
|
268
|
+
|
|
269
|
+
## PDFs
|
|
270
|
+
- `main_round0_original.pdf` — Original generated paper
|
|
271
|
+
- `main_round1.pdf` — After Round 1 fixes
|
|
272
|
+
- `main_round2.pdf` — Final version after Round 2 fixes
|
|
273
|
+
```
|
|
274
|
+
|
|
275
|
+
### Step 9: Summary
|
|
276
|
+
|
|
277
|
+
Report to user:
|
|
278
|
+
- Score progression table
|
|
279
|
+
- Number of CRITICAL/MAJOR/MINOR issues fixed per round
|
|
280
|
+
- Final page count
|
|
281
|
+
- Remaining issues (if any)
|
|
282
|
+
|
|
283
|
+
### Feishu Notification (if configured)
|
|
284
|
+
|
|
285
|
+
After each round's review AND at final completion, check `~/.codex/feishu.json`:
|
|
286
|
+
- **After each round**: Send `review_scored` — "Round N: X/10 — [key changes]"
|
|
287
|
+
- **After final round**: Send `pipeline_done` — score progression table + final page count
|
|
288
|
+
- If config absent or mode `"off"`: skip entirely (no-op)
|
|
289
|
+
|
|
290
|
+
## Output
|
|
291
|
+
|
|
292
|
+
```
|
|
293
|
+
paper/
|
|
294
|
+
├── main_round0_original.pdf # Original
|
|
295
|
+
├── main_round1.pdf # After Round 1
|
|
296
|
+
├── main_round2.pdf # After Round 2 (final)
|
|
297
|
+
├── main.pdf # = main_round2.pdf
|
|
298
|
+
└── PAPER_IMPROVEMENT_LOG.md # Full review log with scores
|
|
299
|
+
```
|
|
300
|
+
|
|
301
|
+
## Output Protocols
|
|
302
|
+
|
|
303
|
+
> Follow these shared protocols for all output files:
|
|
304
|
+
> - **[Output Versioning Protocol](../../shared-references/output-versioning.md)** — write timestamped file first, then copy to fixed name
|
|
305
|
+
> - **[Output Manifest Protocol](../../shared-references/output-manifest.md)** — log every output to MANIFEST.md
|
|
306
|
+
> - **[Output Language Protocol](../../shared-references/output-language.md)** — respect the project's language setting
|
|
307
|
+
|
|
308
|
+
## Key Rules
|
|
309
|
+
|
|
310
|
+
- **Large file handling**: If the Write tool fails due to file size, immediately retry using Bash (`cat << 'EOF' > file`) to write in chunks. Do NOT ask the user for permission — just do it silently.
|
|
311
|
+
|
|
312
|
+
- **Preserve all PDF versions** — user needs to compare progression
|
|
313
|
+
- **Save FULL raw review text** — do not summarize or truncate Gemini reviewer responses
|
|
314
|
+
- **Use `mcp__gemini-review__review_reply_start` plus `mcp__gemini-review__review_status`** for Round 2 to maintain conversation context
|
|
315
|
+
- **Always recompile after fixes** — verify 0 errors before proceeding
|
|
316
|
+
- **Do not fabricate experimental results** — synthetic validation must describe methodology, not invent numbers
|
|
317
|
+
- **Respect the paper's claims** — soften overclaims rather than adding unsupported new claims
|
|
318
|
+
- **Global consistency** — when renaming notation or softening claims, check ALL files (abstract, intro, method, experiments, theory sections, conclusion, tables, figure captions)
|
|
319
|
+
|
|
320
|
+
## Typical Score Progression
|
|
321
|
+
|
|
322
|
+
Based on end-to-end testing on a real theory-paper run:
|
|
323
|
+
|
|
324
|
+
| Round | Score | Key Improvements |
|
|
325
|
+
|-------|-------|-----------------|
|
|
326
|
+
| Round 0 | 4/10 (content) | Baseline: assumption-model mismatch, overclaims, notation issues |
|
|
327
|
+
| Round 1 | 6/10 (content) | Fixed assumptions, softened claims, added interpretation, renamed notation |
|
|
328
|
+
| Round 2 | 7/10 (content) | Added synthetic validation, formal truncation proposition, stronger limitations |
|
|
329
|
+
| Round 3 | 5→8.5/10 (format) | Removed hero fig, appendix, compressed conclusion, fixed overfull hbox |
|
|
330
|
+
|
|
331
|
+
**+4.5 points across 3 rounds** (2 content + 1 format) is typical for a well-structured but rough first draft. Final state at submission: clean overfull-hbox count and venue-format-compliant length.
|
|
@@ -0,0 +1,304 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: "auto-review-loop"
|
|
3
|
+
description: "Autonomous multi-round research review loop. Repeatedly reviews using Gemini via gemini-review MCP, implements fixes, and re-reviews until positive assessment or max rounds reached. Use when user says \"auto review loop\", \"review until it passes\", or wants autonomous iterative improvement."
|
|
4
|
+
---
|
|
5
|
+
|
|
6
|
+
> Override for Codex users who want **Gemini**, not a second Codex agent, to act as the reviewer. Install this package **after** `skills/skills-codex/*`.
|
|
7
|
+
|
|
8
|
+
# Auto Review Loop: Autonomous Research Improvement
|
|
9
|
+
|
|
10
|
+
> **Gemini overlay assurance:** `review_independence: cross-family` and `acceptance_status: accepted`.
|
|
11
|
+
|
|
12
|
+
Autonomously iterate: review → implement fixes → re-review, until the external reviewer gives a positive assessment or MAX_ROUNDS is reached.
|
|
13
|
+
|
|
14
|
+
## Context: $ARGUMENTS
|
|
15
|
+
|
|
16
|
+
## Constants
|
|
17
|
+
|
|
18
|
+
- MAX_ROUNDS = 4
|
|
19
|
+
- POSITIVE_THRESHOLD: score >= 6/10 AND verdict ∈ {"ready", "almost"} — both must hold, matching the operative STOP CONDITION below. Verdict vocabulary is {"ready", "almost", "not ready"}. (Earlier wording used "or" + a stale verdict set; the AND form is authoritative.)
|
|
20
|
+
- REVIEW_DOC: `review-stage/AUTO_REVIEW.md` (cumulative log) *(fall back to `./AUTO_REVIEW.md` for legacy projects)*
|
|
21
|
+
- **OUTPUT_DIR = `review-stage/`** — Directory for review output files.
|
|
22
|
+
- **REVIEWER_MODEL = `gemini-review`** — Gemini reviewer invoked through the local `gemini-review` MCP bridge. Set `GEMINI_REVIEW_MODEL` if you need a specific Gemini model override.
|
|
23
|
+
- **HUMAN_CHECKPOINT = false** — When `true`, pause after each round's review (Phase B) and present the score + weaknesses to the user. Wait for user input before proceeding to Phase C. The user can: approve the suggested fixes, provide custom modification instructions, skip specific fixes, or stop the loop early. When `false` (default), the loop runs fully autonomously.
|
|
24
|
+
|
|
25
|
+
> 💡 Override: `/auto-review-loop "topic" — human checkpoint: true`
|
|
26
|
+
|
|
27
|
+
## State Persistence (Compact Recovery)
|
|
28
|
+
|
|
29
|
+
Long-running loops may hit the context window limit, triggering automatic compaction. To survive this, persist state to `review-stage/REVIEW_STATE.json` after each round:
|
|
30
|
+
|
|
31
|
+
```json
|
|
32
|
+
{
|
|
33
|
+
"round": 2,
|
|
34
|
+
"thread_id": "019cd392-...",
|
|
35
|
+
"status": "in_progress",
|
|
36
|
+
"last_score": 5.0,
|
|
37
|
+
"last_verdict": "not ready",
|
|
38
|
+
"pending_experiments": ["screen_name_1"],
|
|
39
|
+
"timestamp": "2026-03-13T21:00:00"
|
|
40
|
+
}
|
|
41
|
+
```
|
|
42
|
+
|
|
43
|
+
**Write this file at the end of every Phase E** (after documenting the round). Overwrite each time — only the latest state matters.
|
|
44
|
+
|
|
45
|
+
**On completion** (positive assessment or max rounds), set `"status": "completed"` so future invocations don't accidentally resume a finished loop.
|
|
46
|
+
|
|
47
|
+
## Workflow
|
|
48
|
+
|
|
49
|
+
### Initialization
|
|
50
|
+
|
|
51
|
+
1. **Check for `review-stage/REVIEW_STATE.json`** *(fall back to `./REVIEW_STATE.json` if not found — legacy path)*:
|
|
52
|
+
- If neither path exists: **fresh start** (normal case, identical to behavior before this feature existed)
|
|
53
|
+
- If it exists AND `status` is `"completed"`: **fresh start** (previous loop finished normally)
|
|
54
|
+
- If it exists AND `status` is `"in_progress"` AND `timestamp` is older than 24 hours: **fresh start** (stale state from a killed/abandoned run — delete the file and start over)
|
|
55
|
+
- If it exists AND `status` is `"in_progress"` AND `timestamp` is within 24 hours: **resume**
|
|
56
|
+
- Read the state file to recover `round`, `thread_id`, `last_score`, `pending_experiments`
|
|
57
|
+
- Read `review-stage/AUTO_REVIEW.md` to restore full context of prior rounds *(fall back to `./AUTO_REVIEW.md`)*
|
|
58
|
+
- If `pending_experiments` is non-empty, check if they have completed (e.g., check screen sessions)
|
|
59
|
+
- Resume from the next round (round = saved round + 1)
|
|
60
|
+
- Log: "Recovered from context compaction. Resuming at Round N."
|
|
61
|
+
2. Read project narrative documents, memory files, and any prior review documents
|
|
62
|
+
3. Read recent experiment results (check output directories, logs)
|
|
63
|
+
4. Identify current weaknesses and open TODOs from prior reviews
|
|
64
|
+
5. Initialize round counter = 1 (unless recovered from state file)
|
|
65
|
+
6. Create/update `review-stage/AUTO_REVIEW.md` with header and timestamp
|
|
66
|
+
|
|
67
|
+
### Loop (repeat up to MAX_ROUNDS)
|
|
68
|
+
|
|
69
|
+
#### Phase A: Review
|
|
70
|
+
|
|
71
|
+
Send comprehensive context to the external reviewer:
|
|
72
|
+
|
|
73
|
+
```
|
|
74
|
+
mcp__gemini-review__review_start:
|
|
75
|
+
prompt: |
|
|
76
|
+
[Round N/MAX_ROUNDS of autonomous review loop]
|
|
77
|
+
|
|
78
|
+
[Full research context: claims, methods, results, known weaknesses]
|
|
79
|
+
[Changes since last round, if any]
|
|
80
|
+
|
|
81
|
+
Please act as a senior ML reviewer (NeurIPS/ICML level).
|
|
82
|
+
|
|
83
|
+
1. Score this work 1-10 for a top venue
|
|
84
|
+
2. List remaining critical weaknesses (ranked by severity)
|
|
85
|
+
3. For each weakness, specify the MINIMUM fix (experiment, analysis, or reframing)
|
|
86
|
+
4. State clearly: is this READY for submission? Yes/No/Almost
|
|
87
|
+
|
|
88
|
+
Be brutally honest. If the work is ready, say so clearly.
|
|
89
|
+
|
|
90
|
+
=== SCOPE LIMITS (these bound what you PROPOSE, never what you look for) ===
|
|
91
|
+
Report anything that is actually wrong here — including a rare-looking case, if
|
|
92
|
+
this repo actually produces it. Then keep the fix in scope:
|
|
93
|
+
1. This is a RESEARCH-WORKFLOW tool, not a security paper. Verification is
|
|
94
|
+
welcome; over-defense is not. Assume a cooperating operator on their own
|
|
95
|
+
machine — a malicious local user is NOT in the threat model.
|
|
96
|
+
2. Do NOT propose SHA / hash / content-fingerprint / digest-binding schemes.
|
|
97
|
+
Reporting a real defect in hashing code that already exists is fine.
|
|
98
|
+
3. NO speculative machinery: do not add feature flags, migration frameworks,
|
|
99
|
+
compat layers, wrappers, pins, or similar mechanisms unless evidence shows
|
|
100
|
+
a current repo defect they fix or an explicit existing invariant they must
|
|
101
|
+
preserve. "Load-bearing", "compatibility", and "not scaffolding" are labels,
|
|
102
|
+
not evidence. Point to the failing path/artifact or invariant, and check the
|
|
103
|
+
proposal's factual premises, such as whether a named package version exists.
|
|
104
|
+
4. NO corner-case obsession: exotic encodings, symlink races, RTL text and
|
|
105
|
+
millisecond races are out of scope unless you can show the case arises here.
|
|
106
|
+
5. Where a rubric or checklist is genuinely needed, do not over-mechanize
|
|
107
|
+
judgement. A clear sentence a human reads beats a scored table nobody
|
|
108
|
+
maintains.
|
|
109
|
+
Exception: code that runs remote commands, starts a network service, or installs
|
|
110
|
+
an MCP server runs on the user's machine with their credentials — trust-boundary
|
|
111
|
+
findings there are in scope and the default is strict.
|
|
112
|
+
Say plainly when something is correct. Do not manufacture findings.
|
|
113
|
+
```
|
|
114
|
+
|
|
115
|
+
After this start call, immediately save the returned `jobId` and poll `mcp__gemini-review__review_status` with a bounded `waitSeconds` until `done=true`. Treat the completed status payload's `response` as the reviewer output, and save the completed `threadId` for any follow-up round.
|
|
116
|
+
|
|
117
|
+
If this is round 2+, use `mcp__gemini-review__review_reply_start` with the saved completed `threadId`, then poll `mcp__gemini-review__review_status` with the returned `jobId` until `done=true` to maintain continuity.
|
|
118
|
+
|
|
119
|
+
#### Phase B: Parse Assessment
|
|
120
|
+
|
|
121
|
+
**CRITICAL: Save the FULL raw response** from the external reviewer verbatim (store in a variable for Phase E). Do NOT discard or summarize — the raw text is the primary record.
|
|
122
|
+
|
|
123
|
+
Then extract structured fields:
|
|
124
|
+
- **Score** (numeric 1-10)
|
|
125
|
+
- **Verdict** ("ready" / "almost" / "not ready")
|
|
126
|
+
- **Action items** (ranked list of fixes)
|
|
127
|
+
|
|
128
|
+
**STOP CONDITION**: If score >= 6 AND verdict ∈ {"ready", "almost"} (exact match — "not ready" does NOT qualify) → stop loop, document final state.
|
|
129
|
+
|
|
130
|
+
#### Human Checkpoint (if enabled)
|
|
131
|
+
|
|
132
|
+
**Skip this step entirely if `HUMAN_CHECKPOINT = false`.**
|
|
133
|
+
|
|
134
|
+
When `HUMAN_CHECKPOINT = true`, present the review results and wait for user input:
|
|
135
|
+
|
|
136
|
+
```
|
|
137
|
+
📋 Round N/MAX_ROUNDS review complete.
|
|
138
|
+
|
|
139
|
+
Score: X/10 — [verdict]
|
|
140
|
+
Top weaknesses:
|
|
141
|
+
1. [weakness 1]
|
|
142
|
+
2. [weakness 2]
|
|
143
|
+
3. [weakness 3]
|
|
144
|
+
|
|
145
|
+
Suggested fixes:
|
|
146
|
+
1. [fix 1]
|
|
147
|
+
2. [fix 2]
|
|
148
|
+
3. [fix 3]
|
|
149
|
+
|
|
150
|
+
Options:
|
|
151
|
+
- Reply "go" or "continue" → implement all suggested fixes
|
|
152
|
+
- Reply with custom instructions → implement your modifications instead
|
|
153
|
+
- Reply "skip 2" → skip fix #2, implement the rest
|
|
154
|
+
- Reply "stop" → end the loop, document current state
|
|
155
|
+
```
|
|
156
|
+
|
|
157
|
+
Wait for the user's response. Parse their input:
|
|
158
|
+
- **Approval** ("go", "continue", "ok", "proceed"): proceed to Phase C with all suggested fixes
|
|
159
|
+
- **Custom instructions** (any other text): treat as additional/replacement guidance for Phase C. Merge with reviewer suggestions where appropriate
|
|
160
|
+
- **Skip specific fixes** ("skip 1,3"): remove those fixes from the action list
|
|
161
|
+
- **Stop** ("stop", "enough", "done"): terminate the loop, jump to Termination
|
|
162
|
+
|
|
163
|
+
#### Feishu Notification (if configured)
|
|
164
|
+
|
|
165
|
+
After parsing the score, check if `~/.codex/feishu.json` exists and mode is not `"off"`:
|
|
166
|
+
- Send a `review_scored` notification: "Round N: X/10 — [verdict]" with top 3 weaknesses
|
|
167
|
+
- If **interactive** mode and verdict is "almost": send as checkpoint, wait for user reply on whether to continue or stop
|
|
168
|
+
- If config absent or mode off: skip entirely (no-op)
|
|
169
|
+
|
|
170
|
+
#### Phase C: Implement Fixes (if not stopping)
|
|
171
|
+
|
|
172
|
+
For each action item (highest priority first):
|
|
173
|
+
|
|
174
|
+
1. **Code changes**: Write/modify experiment scripts, model code, analysis scripts
|
|
175
|
+
2. **Run experiments**: Deploy to GPU server via SSH + screen/tmux
|
|
176
|
+
3. **Analysis**: Run evaluation, collect results, update figures/tables
|
|
177
|
+
4. **Documentation**: Update project notes and review document
|
|
178
|
+
|
|
179
|
+
Prioritization rules:
|
|
180
|
+
- Skip fixes requiring excessive compute (flag for manual follow-up)
|
|
181
|
+
- Skip fixes requiring external data/models not available
|
|
182
|
+
- Prefer reframing/analysis over new experiments when both address the concern
|
|
183
|
+
- Always implement metric additions (cheap, high impact)
|
|
184
|
+
|
|
185
|
+
#### Phase D: Wait for Results
|
|
186
|
+
|
|
187
|
+
If experiments were launched:
|
|
188
|
+
- Monitor remote sessions for completion
|
|
189
|
+
- Collect results from output files and logs
|
|
190
|
+
|
|
191
|
+
#### Phase E: Document Round
|
|
192
|
+
|
|
193
|
+
Append to `review-stage/AUTO_REVIEW.md`:
|
|
194
|
+
|
|
195
|
+
```markdown
|
|
196
|
+
## Round N (timestamp)
|
|
197
|
+
|
|
198
|
+
### Assessment (Summary)
|
|
199
|
+
- Score: X/10
|
|
200
|
+
- Verdict: [ready/almost/not ready]
|
|
201
|
+
- Key criticisms: [bullet list]
|
|
202
|
+
|
|
203
|
+
### Reviewer Raw Response
|
|
204
|
+
|
|
205
|
+
<details>
|
|
206
|
+
<summary>Click to expand full reviewer response</summary>
|
|
207
|
+
|
|
208
|
+
[Paste the COMPLETE raw response from the external reviewer here — verbatim, unedited.
|
|
209
|
+
This is the authoritative record. Do NOT truncate or paraphrase.]
|
|
210
|
+
|
|
211
|
+
</details>
|
|
212
|
+
|
|
213
|
+
### Actions Taken
|
|
214
|
+
- [what was implemented/changed]
|
|
215
|
+
|
|
216
|
+
### Results
|
|
217
|
+
- [experiment outcomes, if any]
|
|
218
|
+
|
|
219
|
+
### Status
|
|
220
|
+
- [continuing to round N+1 / stopping]
|
|
221
|
+
```
|
|
222
|
+
|
|
223
|
+
**Write `review-stage/REVIEW_STATE.json`** with current round, agent id, score, verdict, and any pending experiments.
|
|
224
|
+
|
|
225
|
+
Increment round counter → back to Phase A.
|
|
226
|
+
|
|
227
|
+
### Termination
|
|
228
|
+
|
|
229
|
+
When loop ends (positive assessment or max rounds):
|
|
230
|
+
|
|
231
|
+
1. Update `review-stage/REVIEW_STATE.json` with `"status": "completed"`
|
|
232
|
+
2. Write final summary to `review-stage/AUTO_REVIEW.md`
|
|
233
|
+
3. Update project notes with conclusions
|
|
234
|
+
4. If stopped at max rounds without positive assessment:
|
|
235
|
+
- List remaining blockers
|
|
236
|
+
- Estimate effort needed for each
|
|
237
|
+
- Suggest whether to continue manually or pivot
|
|
238
|
+
5. **Feishu notification** (if configured): Send `pipeline_done` with final score progression table
|
|
239
|
+
|
|
240
|
+
## Output Protocols
|
|
241
|
+
|
|
242
|
+
> Follow these shared protocols for all output files:
|
|
243
|
+
> - **[Output Versioning Protocol](../../shared-references/output-versioning.md)** — write timestamped file first, then copy to fixed name
|
|
244
|
+
> - **[Output Manifest Protocol](../../shared-references/output-manifest.md)** — log every output to MANIFEST.md
|
|
245
|
+
> - **[Output Language Protocol](../../shared-references/output-language.md)** — respect the project's language setting
|
|
246
|
+
|
|
247
|
+
## Key Rules
|
|
248
|
+
|
|
249
|
+
- **Large file handling**: If the Write tool fails due to file size, immediately retry using Bash (`cat << 'EOF' > file`) to write in chunks. Do NOT ask the user for permission — just do it silently.
|
|
250
|
+
|
|
251
|
+
- Always ask the Gemini reviewer for strict, high-rigor feedback.
|
|
252
|
+
- Save the completed `threadId` from the first `mcp__gemini-review__review_status` result, then use `mcp__gemini-review__review_reply_start` plus `mcp__gemini-review__review_status` for subsequent rounds
|
|
253
|
+
- Be honest — include negative results and failed experiments
|
|
254
|
+
- Do NOT hide weaknesses to game a positive score
|
|
255
|
+
- Implement fixes BEFORE re-reviewing (don't just promise to fix)
|
|
256
|
+
- If an experiment takes > 30 minutes, launch it and continue with other fixes while waiting
|
|
257
|
+
- Document EVERYTHING — the review log should be self-contained
|
|
258
|
+
- Update project notes after each round, not just at the end
|
|
259
|
+
|
|
260
|
+
## Prompt Template for Round 2+
|
|
261
|
+
|
|
262
|
+
```
|
|
263
|
+
mcp__gemini-review__review_reply_start:
|
|
264
|
+
threadId: [saved from round 1]
|
|
265
|
+
prompt: |
|
|
266
|
+
[Round N update]
|
|
267
|
+
|
|
268
|
+
Since your last review, we have:
|
|
269
|
+
1. [Action 1]: [result]
|
|
270
|
+
2. [Action 2]: [result]
|
|
271
|
+
3. [Action 3]: [result]
|
|
272
|
+
|
|
273
|
+
Updated results table:
|
|
274
|
+
[paste metrics]
|
|
275
|
+
|
|
276
|
+
Please re-score and re-assess. Are the remaining concerns addressed?
|
|
277
|
+
Same format: Score, Verdict, Remaining Weaknesses, Minimum Fixes.
|
|
278
|
+
|
|
279
|
+
=== SCOPE LIMITS (these bound what you PROPOSE, never what you look for) ===
|
|
280
|
+
Report anything that is actually wrong here — including a rare-looking case, if
|
|
281
|
+
this repo actually produces it. Then keep the fix in scope:
|
|
282
|
+
1. This is a RESEARCH-WORKFLOW tool, not a security paper. Verification is
|
|
283
|
+
welcome; over-defense is not. Assume a cooperating operator on their own
|
|
284
|
+
machine — a malicious local user is NOT in the threat model.
|
|
285
|
+
2. Do NOT propose SHA / hash / content-fingerprint / digest-binding schemes.
|
|
286
|
+
Reporting a real defect in hashing code that already exists is fine.
|
|
287
|
+
3. NO speculative machinery: do not add feature flags, migration frameworks,
|
|
288
|
+
compat layers, wrappers, pins, or similar mechanisms unless evidence shows
|
|
289
|
+
a current repo defect they fix or an explicit existing invariant they must
|
|
290
|
+
preserve. "Load-bearing", "compatibility", and "not scaffolding" are labels,
|
|
291
|
+
not evidence. Point to the failing path/artifact or invariant, and check the
|
|
292
|
+
proposal's factual premises, such as whether a named package version exists.
|
|
293
|
+
4. NO corner-case obsession: exotic encodings, symlink races, RTL text and
|
|
294
|
+
millisecond races are out of scope unless you can show the case arises here.
|
|
295
|
+
5. Where a rubric or checklist is genuinely needed, do not over-mechanize
|
|
296
|
+
judgement. A clear sentence a human reads beats a scored table nobody
|
|
297
|
+
maintains.
|
|
298
|
+
Exception: code that runs remote commands, starts a network service, or installs
|
|
299
|
+
an MCP server runs on the user's machine with their credentials — trust-boundary
|
|
300
|
+
findings there are in scope and the default is strict.
|
|
301
|
+
Say plainly when something is correct. Do not manufacture findings.
|
|
302
|
+
```
|
|
303
|
+
|
|
304
|
+
After this start call, immediately save the returned `jobId` and poll `mcp__gemini-review__review_status` with a bounded `waitSeconds` until `done=true`. Treat the completed status payload's `response` as the reviewer output, and save the completed `threadId` for any follow-up round.
|