dsh-aris-panel 0.1.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/LICENSE +21 -0
- package/README.md +98 -0
- package/README_CN.md +87 -0
- package/dsh/checkout.patch.yml +38 -0
- package/dsh/client.js +634 -0
- package/dsh/cordis.patch.yml +44 -0
- package/dsh/index.mjs +76 -0
- package/dsh/run-status.mjs +182 -0
- package/dsh/scope-limits.mjs +50 -0
- package/dsh/workbench.mjs +291 -0
- package/mcp-servers/claude-review/README.md +93 -0
- package/mcp-servers/claude-review/run_with_claude_aws.sh +49 -0
- package/mcp-servers/claude-review/server.py +718 -0
- package/mcp-servers/codex-image2/README.md +65 -0
- package/mcp-servers/codex-image2/server.py +893 -0
- package/mcp-servers/feishu-bridge/requirements.txt +1 -0
- package/mcp-servers/feishu-bridge/server.py +240 -0
- package/mcp-servers/gemini-review/README.md +171 -0
- package/mcp-servers/gemini-review/server.py +1856 -0
- package/mcp-servers/llm-chat/requirements.txt +1 -0
- package/mcp-servers/llm-chat/server.py +664 -0
- package/mcp-servers/manual-review/README.md +133 -0
- package/mcp-servers/manual-review/server.py +910 -0
- package/mcp-servers/manual-review/ui.html +279 -0
- package/mcp-servers/minimax-chat/requirements.txt +1 -0
- package/mcp-servers/minimax-chat/server.py +381 -0
- package/package.json +51 -0
- package/skills/ablation-planner/SKILL.md +123 -0
- package/skills/alphaxiv/SKILL.md +196 -0
- package/skills/analyze-results/SKILL.md +46 -0
- package/skills/arxiv/SKILL.md +248 -0
- package/skills/auto-paper-improvement-loop/SKILL.md +651 -0
- package/skills/auto-review-loop/SKILL.md +1137 -0
- package/skills/auto-review-loop-llm/SKILL.md +259 -0
- package/skills/auto-review-loop-minimax/SKILL.md +302 -0
- package/skills/citation-audit/SKILL.md +502 -0
- package/skills/claims-drafting/SKILL.md +227 -0
- package/skills/comm-lit-review/SKILL.md +297 -0
- package/skills/deepxiv/SKILL.md +263 -0
- package/skills/dse-loop/SKILL.md +296 -0
- package/skills/embodiment-description/SKILL.md +129 -0
- package/skills/exa-search/SKILL.md +205 -0
- package/skills/experiment-audit/SKILL.md +311 -0
- package/skills/experiment-bridge/SKILL.md +376 -0
- package/skills/experiment-plan/SKILL.md +249 -0
- package/skills/experiment-queue/SKILL.md +431 -0
- package/skills/experiment-queue/scripts/build_manifest.py +142 -0
- package/skills/experiment-queue/scripts/queue_manager.py +433 -0
- package/skills/feishu-notify/SKILL.md +156 -0
- package/skills/figure-description/SKILL.md +138 -0
- package/skills/figure-spec/SKILL.md +262 -0
- package/skills/figure-spec/scripts/figure_renderer.py +799 -0
- package/skills/formula-derivation/SKILL.md +280 -0
- package/skills/gemini-search/SKILL.md +231 -0
- package/skills/grant-proposal/SKILL.md +698 -0
- package/skills/idea-creator/SKILL.md +542 -0
- package/skills/idea-discovery/SKILL.md +521 -0
- package/skills/idea-discovery-robot/SKILL.md +363 -0
- package/skills/integrity-forensics/SKILL.md +284 -0
- package/skills/interview-cheatsheet/SKILL.md +245 -0
- package/skills/invention-structuring/SKILL.md +188 -0
- package/skills/jurisdiction-format/SKILL.md +192 -0
- package/skills/kill-argument/SKILL.md +437 -0
- package/skills/mermaid-diagram/SKILL.md +419 -0
- package/skills/meta-apply/SKILL.md +141 -0
- package/skills/meta-optimize/SKILL.md +437 -0
- package/skills/monitor-experiment/SKILL.md +140 -0
- package/skills/novelty-check/SKILL.md +101 -0
- package/skills/openalex/SKILL.md +237 -0
- package/skills/overleaf-sync/SKILL.md +220 -0
- package/skills/paper-claim-audit/SKILL.md +348 -0
- package/skills/paper-compile/SKILL.md +266 -0
- package/skills/paper-figure/SKILL.md +312 -0
- package/skills/paper-illustration/SKILL.md +736 -0
- package/skills/paper-illustration-image2/SKILL.md +391 -0
- package/skills/paper-illustration-image2/scripts/paper_illustration_image2.py +255 -0
- package/skills/paper-plan/SKILL.md +386 -0
- package/skills/paper-poster/SKILL.md +19 -0
- package/skills/paper-poster-html/DESIGN_FINAL.md +176 -0
- package/skills/paper-poster-html/IMPLEMENTATION_CONVENTIONS.md +161 -0
- package/skills/paper-poster-html/LICENSES/posterly-MIT.txt +21 -0
- package/skills/paper-poster-html/NOTICE.md +57 -0
- package/skills/paper-poster-html/SKILL.md +323 -0
- package/skills/paper-poster-html/scripts/_posterly/__init__.py +0 -0
- package/skills/paper-poster-html/scripts/_posterly/canvas.py +200 -0
- package/skills/paper-poster-html/scripts/_posterly/measure.py +588 -0
- package/skills/paper-poster-html/scripts/_posterly/polish.py +498 -0
- package/skills/paper-poster-html/scripts/_posterly/preflight.py +489 -0
- package/skills/paper-poster-html/scripts/_posterly/render.py +215 -0
- package/skills/paper-poster-html/scripts/_posterly/textutil.py +16 -0
- package/skills/paper-poster-html/scripts/_posterly/verify_final.py +171 -0
- package/skills/paper-poster-html/scripts/asset_check.py +897 -0
- package/skills/paper-poster-html/scripts/extract_pdf_figures.py +666 -0
- package/skills/paper-poster-html/scripts/poster_check.py +251 -0
- package/skills/paper-poster-html/scripts/preprocess_figures.py +238 -0
- package/skills/paper-poster-html/scripts/render_preview.py +217 -0
- package/skills/paper-poster-html/scripts/run_gates.py +556 -0
- package/skills/paper-poster-html/scripts/style_check.py +1324 -0
- package/skills/paper-poster-html/templates/COMPONENTS.md +462 -0
- package/skills/paper-poster-html/templates/README.md +170 -0
- package/skills/paper-poster-html/templates/landscape_4col.html +1032 -0
- package/skills/paper-poster-html/templates/landscape_hero.html +1046 -0
- package/skills/paper-poster-html/templates/portrait_2col.html +947 -0
- package/skills/paper-poster-html/templates/tokens/acl.json +9 -0
- package/skills/paper-poster-html/templates/tokens/cvpr.json +9 -0
- package/skills/paper-poster-html/templates/tokens/generic.json +9 -0
- package/skills/paper-poster-html/templates/tokens/iclr.json +9 -0
- package/skills/paper-poster-html/templates/tokens/icml.json +9 -0
- package/skills/paper-poster-html/templates/tokens/neurips.json +9 -0
- package/skills/paper-slides/SKILL.md +635 -0
- package/skills/paper-talk/SKILL.md +381 -0
- package/skills/paper-write/SKILL.md +604 -0
- package/skills/paper-write/templates/IEEEtran.bst +2409 -0
- package/skills/paper-write/templates/IEEEtran.cls +6347 -0
- package/skills/paper-write/templates/iclr2026.tex +84 -0
- package/skills/paper-write/templates/icml2025.tex +87 -0
- package/skills/paper-write/templates/ieee_conference.tex +89 -0
- package/skills/paper-write/templates/ieee_journal.tex +93 -0
- package/skills/paper-write/templates/math_commands.tex +48 -0
- package/skills/paper-write/templates/neurips2025.tex +80 -0
- package/skills/paper-writing/SKILL.md +916 -0
- package/skills/patent-novelty-check/SKILL.md +153 -0
- package/skills/patent-pipeline/SKILL.md +344 -0
- package/skills/patent-review/SKILL.md +203 -0
- package/skills/pixel-art/SKILL.md +137 -0
- package/skills/prior-art-search/SKILL.md +146 -0
- package/skills/proof-checker/SKILL.md +866 -0
- package/skills/proof-orchestrator/NOTICE.md +24 -0
- package/skills/proof-orchestrator/SKILL.md +254 -0
- package/skills/proof-orchestrator/references/audit-output-contract.md +126 -0
- package/skills/proof-orchestrator/references/deepseek-routing.md +74 -0
- package/skills/proof-orchestrator/references/dispatch-prompts.md +227 -0
- package/skills/proof-orchestrator/references/notation-audit.md +135 -0
- package/skills/proof-orchestrator/references/proof-audit-rubric.md +70 -0
- package/skills/proof-orchestrator/references/stress-tests.md +38 -0
- package/skills/proof-writer/SKILL.md +223 -0
- package/skills/qzcli/SKILL.md +324 -0
- package/skills/rebuttal/SKILL.md +376 -0
- package/skills/render-html/SKILL.md +316 -0
- package/skills/render-html/scripts/render_html.py +1006 -0
- package/skills/render-html/scripts/templates/academic.html +703 -0
- package/skills/render-html/scripts/templates/dashboard.html +333 -0
- package/skills/research-lit/SKILL.md +756 -0
- package/skills/research-pipeline/SKILL.md +384 -0
- package/skills/research-refine/SKILL.md +770 -0
- package/skills/research-refine-pipeline/SKILL.md +186 -0
- package/skills/research-review/SKILL.md +198 -0
- package/skills/research-wiki/SKILL.md +461 -0
- package/skills/resubmit-pipeline/SKILL.md +447 -0
- package/skills/result-to-claim/SKILL.md +311 -0
- package/skills/run-experiment/SKILL.md +313 -0
- package/skills/semantic-scholar/SKILL.md +236 -0
- package/skills/serverless-modal/SKILL.md +335 -0
- package/skills/shared-references/acceptance-gate.md +324 -0
- package/skills/shared-references/assurance-contract.md +248 -0
- package/skills/shared-references/capture-antipatterns.md +78 -0
- package/skills/shared-references/citation-discipline.md +583 -0
- package/skills/shared-references/compute-env-contract.md +163 -0
- package/skills/shared-references/effort-contract.md +183 -0
- package/skills/shared-references/evidence-precheck.md +65 -0
- package/skills/shared-references/experiment-integrity.md +49 -0
- package/skills/shared-references/external-cadence.md +326 -0
- package/skills/shared-references/fan-out-pattern.md +366 -0
- package/skills/shared-references/injection-hygiene.md +127 -0
- package/skills/shared-references/integration-contract.md +461 -0
- package/skills/shared-references/output-composition.md +93 -0
- package/skills/shared-references/output-language.md +45 -0
- package/skills/shared-references/output-manifest.md +49 -0
- package/skills/shared-references/output-versioning.md +111 -0
- package/skills/shared-references/patent-format-cn.md +199 -0
- package/skills/shared-references/patent-format-ep.md +173 -0
- package/skills/shared-references/patent-format-us.md +161 -0
- package/skills/shared-references/patent-writing-principles.md +197 -0
- package/skills/shared-references/prior-art-databases.md +141 -0
- package/skills/shared-references/resumable-runs.md +109 -0
- package/skills/shared-references/review-scope-limits.md +81 -0
- package/skills/shared-references/review-tracing.md +391 -0
- package/skills/shared-references/reviewer-independence.md +79 -0
- package/skills/shared-references/reviewer-routing.md +852 -0
- package/skills/shared-references/skill-governance.md +104 -0
- package/skills/shared-references/taste-calibration.md +85 -0
- package/skills/shared-references/venue-checklists.md +114 -0
- package/skills/shared-references/wiki-helper-resolution.md +134 -0
- package/skills/shared-references/writing-principles.md +525 -0
- package/skills/skills-codex/README.md +102 -0
- package/skills/skills-codex/README_CN.md +100 -0
- package/skills/skills-codex/ablation-planner/SKILL.md +126 -0
- package/skills/skills-codex/alphaxiv/SKILL.md +186 -0
- package/skills/skills-codex/analyze-results/SKILL.md +45 -0
- package/skills/skills-codex/arxiv/SKILL.md +210 -0
- package/skills/skills-codex/auto-paper-improvement-loop/SKILL.md +574 -0
- package/skills/skills-codex/auto-review-loop/SKILL.md +500 -0
- package/skills/skills-codex/auto-review-loop-llm/SKILL.md +247 -0
- package/skills/skills-codex/auto-review-loop-minimax/SKILL.md +290 -0
- package/skills/skills-codex/citation-audit/SKILL.md +504 -0
- package/skills/skills-codex/claims-drafting/SKILL.md +239 -0
- package/skills/skills-codex/comm-lit-review/SKILL.md +299 -0
- package/skills/skills-codex/comm-lit-review/references/domain-taxonomy.md +57 -0
- package/skills/skills-codex/comm-lit-review/references/output-template.md +37 -0
- package/skills/skills-codex/comm-lit-review/references/source-policy.md +99 -0
- package/skills/skills-codex/comm-lit-review/references/venue-tiering.md +112 -0
- package/skills/skills-codex/deepxiv/SKILL.md +142 -0
- package/skills/skills-codex/dse-loop/SKILL.md +285 -0
- package/skills/skills-codex/embodiment-description/SKILL.md +129 -0
- package/skills/skills-codex/exa-search/SKILL.md +192 -0
- package/skills/skills-codex/experiment-audit/SKILL.md +286 -0
- package/skills/skills-codex/experiment-bridge/SKILL.md +356 -0
- package/skills/skills-codex/experiment-plan/SKILL.md +249 -0
- package/skills/skills-codex/experiment-queue/SKILL.md +401 -0
- package/skills/skills-codex/feishu-notify/SKILL.md +155 -0
- package/skills/skills-codex/figure-description/SKILL.md +138 -0
- package/skills/skills-codex/figure-spec/SKILL.md +252 -0
- package/skills/skills-codex/formula-derivation/SKILL.md +280 -0
- package/skills/skills-codex/gemini-search/SKILL.md +205 -0
- package/skills/skills-codex/grant-proposal/SKILL.md +626 -0
- package/skills/skills-codex/idea-creator/SKILL.md +405 -0
- package/skills/skills-codex/idea-discovery/SKILL.md +475 -0
- package/skills/skills-codex/idea-discovery-robot/SKILL.md +362 -0
- package/skills/skills-codex/integrity-forensics/SKILL.md +106 -0
- package/skills/skills-codex/interview-cheatsheet/SKILL.md +245 -0
- package/skills/skills-codex/invention-structuring/SKILL.md +188 -0
- package/skills/skills-codex/jurisdiction-format/SKILL.md +192 -0
- package/skills/skills-codex/kill-argument/SKILL.md +403 -0
- package/skills/skills-codex/mermaid-diagram/SKILL.md +379 -0
- package/skills/skills-codex/meta-apply/SKILL.md +154 -0
- package/skills/skills-codex/meta-optimize/SKILL.md +348 -0
- package/skills/skills-codex/monitor-experiment/SKILL.md +98 -0
- package/skills/skills-codex/novelty-check/SKILL.md +89 -0
- package/skills/skills-codex/openalex/SKILL.md +228 -0
- package/skills/skills-codex/overleaf-sync/SKILL.md +220 -0
- package/skills/skills-codex/paper-claim-audit/SKILL.md +350 -0
- package/skills/skills-codex/paper-compile/SKILL.md +253 -0
- package/skills/skills-codex/paper-figure/SKILL.md +311 -0
- package/skills/skills-codex/paper-illustration/SKILL.md +690 -0
- package/skills/skills-codex/paper-illustration-image2/SKILL.md +383 -0
- package/skills/skills-codex/paper-illustration-image2/scripts/paper_illustration_image2.py +255 -0
- package/skills/skills-codex/paper-plan/SKILL.md +278 -0
- package/skills/skills-codex/paper-poster/SKILL.md +19 -0
- package/skills/skills-codex/paper-poster-html/SKILL.md +377 -0
- package/skills/skills-codex/paper-slides/SKILL.md +571 -0
- package/skills/skills-codex/paper-talk/SKILL.md +381 -0
- package/skills/skills-codex/paper-write/SKILL.md +411 -0
- package/skills/skills-codex/paper-write/templates/IEEEtran.bst +2409 -0
- package/skills/skills-codex/paper-write/templates/IEEEtran.cls +6347 -0
- package/skills/skills-codex/paper-write/templates/aaai2026.bst +1493 -0
- package/skills/skills-codex/paper-write/templates/aaai2026.sty +315 -0
- package/skills/skills-codex/paper-write/templates/aaai2026.tex +952 -0
- package/skills/skills-codex/paper-write/templates/acl.sty +312 -0
- package/skills/skills-codex/paper-write/templates/acl2026.tex +377 -0
- package/skills/skills-codex/paper-write/templates/acl_natbib.bst +1940 -0
- package/skills/skills-codex/paper-write/templates/acm.bst +3081 -0
- package/skills/skills-codex/paper-write/templates/acm_mm2026.tex +204 -0
- package/skills/skills-codex/paper-write/templates/acmart.cls +3520 -0
- package/skills/skills-codex/paper-write/templates/cvpr.bst +1448 -0
- package/skills/skills-codex/paper-write/templates/cvpr.sty +508 -0
- package/skills/skills-codex/paper-write/templates/cvpr2026.tex +63 -0
- package/skills/skills-codex/paper-write/templates/iclr2026.tex +84 -0
- package/skills/skills-codex/paper-write/templates/iclr2026_conference.bst +1440 -0
- package/skills/skills-codex/paper-write/templates/iclr2026_conference.sty +246 -0
- package/skills/skills-codex/paper-write/templates/icml2026.sty +767 -0
- package/skills/skills-codex/paper-write/templates/icml2026.tex +662 -0
- package/skills/skills-codex/paper-write/templates/ieee_conference.tex +89 -0
- package/skills/skills-codex/paper-write/templates/ieee_journal.tex +93 -0
- package/skills/skills-codex/paper-write/templates/math_commands.tex +48 -0
- package/skills/skills-codex/paper-write/templates/neurips2026.tex +493 -0
- package/skills/skills-codex/paper-write/templates/neurips_2026.sty +437 -0
- package/skills/skills-codex/paper-writing/SKILL.md +731 -0
- package/skills/skills-codex/patent-novelty-check/SKILL.md +153 -0
- package/skills/skills-codex/patent-pipeline/SKILL.md +344 -0
- package/skills/skills-codex/patent-review/SKILL.md +202 -0
- package/skills/skills-codex/pixel-art/SKILL.md +139 -0
- package/skills/skills-codex/prior-art-search/SKILL.md +146 -0
- package/skills/skills-codex/proof-checker/SKILL.md +554 -0
- package/skills/skills-codex/proof-orchestrator/SKILL.md +260 -0
- package/skills/skills-codex/proof-orchestrator/references/audit-output-contract.md +126 -0
- package/skills/skills-codex/proof-orchestrator/references/deepseek-routing.md +76 -0
- package/skills/skills-codex/proof-orchestrator/references/dispatch-prompts.md +227 -0
- package/skills/skills-codex/proof-orchestrator/references/notation-audit.md +135 -0
- package/skills/skills-codex/proof-orchestrator/references/proof-audit-rubric.md +70 -0
- package/skills/skills-codex/proof-orchestrator/references/stress-tests.md +38 -0
- package/skills/skills-codex/proof-writer/SKILL.md +222 -0
- package/skills/skills-codex/qzcli/SKILL.md +324 -0
- package/skills/skills-codex/rebuttal/SKILL.md +305 -0
- package/skills/skills-codex/render-html/SKILL.md +305 -0
- package/skills/skills-codex/render-html/scripts/__pycache__/render_html.cpython-314.pyc +0 -0
- package/skills/skills-codex/render-html/scripts/render_html.py +909 -0
- package/skills/skills-codex/render-html/scripts/templates/academic.html +342 -0
- package/skills/skills-codex/render-html/scripts/templates/dashboard.html +333 -0
- package/skills/skills-codex/research-lit/SKILL.md +464 -0
- package/skills/skills-codex/research-pipeline/SKILL.md +340 -0
- package/skills/skills-codex/research-refine/SKILL.md +721 -0
- package/skills/skills-codex/research-refine-pipeline/SKILL.md +186 -0
- package/skills/skills-codex/research-review/SKILL.md +135 -0
- package/skills/skills-codex/research-wiki/SKILL.md +421 -0
- package/skills/skills-codex/resubmit-pipeline/SKILL.md +444 -0
- package/skills/skills-codex/result-to-claim/SKILL.md +246 -0
- package/skills/skills-codex/run-experiment/SKILL.md +236 -0
- package/skills/skills-codex/semantic-scholar/SKILL.md +219 -0
- package/skills/skills-codex/serverless-modal/SKILL.md +335 -0
- package/skills/skills-codex/shared-references/acceptance-gate.md +336 -0
- package/skills/skills-codex/shared-references/assurance-contract.md +139 -0
- package/skills/skills-codex/shared-references/capture-antipatterns.md +84 -0
- package/skills/skills-codex/shared-references/citation-discipline.md +452 -0
- package/skills/skills-codex/shared-references/compute-env-contract.md +163 -0
- package/skills/skills-codex/shared-references/effort-contract.md +143 -0
- package/skills/skills-codex/shared-references/evidence-precheck.md +73 -0
- package/skills/skills-codex/shared-references/experiment-integrity.md +49 -0
- package/skills/skills-codex/shared-references/external-cadence.md +334 -0
- package/skills/skills-codex/shared-references/fan-out-pattern.md +375 -0
- package/skills/skills-codex/shared-references/injection-hygiene.md +132 -0
- package/skills/skills-codex/shared-references/integration-contract.md +372 -0
- package/skills/skills-codex/shared-references/output-composition.md +98 -0
- package/skills/skills-codex/shared-references/output-language.md +45 -0
- package/skills/skills-codex/shared-references/output-manifest.md +40 -0
- package/skills/skills-codex/shared-references/output-versioning.md +111 -0
- package/skills/skills-codex/shared-references/patent-format-cn.md +199 -0
- package/skills/skills-codex/shared-references/patent-format-ep.md +173 -0
- package/skills/skills-codex/shared-references/patent-format-us.md +161 -0
- package/skills/skills-codex/shared-references/patent-writing-principles.md +197 -0
- package/skills/skills-codex/shared-references/prior-art-databases.md +141 -0
- package/skills/skills-codex/shared-references/resumable-runs.md +125 -0
- package/skills/skills-codex/shared-references/review-scope-limits.md +81 -0
- package/skills/skills-codex/shared-references/review-tracing.md +144 -0
- package/skills/skills-codex/shared-references/reviewer-independence.md +66 -0
- package/skills/skills-codex/shared-references/reviewer-routing.md +128 -0
- package/skills/skills-codex/shared-references/skill-governance.md +119 -0
- package/skills/skills-codex/shared-references/taste-calibration.md +90 -0
- package/skills/skills-codex/shared-references/venue-checklists.md +73 -0
- package/skills/skills-codex/shared-references/wiki-helper-resolution.md +69 -0
- package/skills/skills-codex/shared-references/writing-principles.md +525 -0
- package/skills/skills-codex/slides-polish/SKILL.md +563 -0
- package/skills/skills-codex/specification-writing/SKILL.md +211 -0
- package/skills/skills-codex/system-profile/SKILL.md +103 -0
- package/skills/skills-codex/training-check/SKILL.md +83 -0
- package/skills/skills-codex/vast-gpu/SKILL.md +394 -0
- package/skills/skills-codex/web-debug-search/SKILL.md +334 -0
- package/skills/skills-codex/wiki-enrich/SKILL.md +255 -0
- package/skills/skills-codex/writing-systems-papers/SKILL.md +184 -0
- package/skills/skills-codex-claude-review/README.md +79 -0
- package/skills/skills-codex-claude-review/README_CN.md +78 -0
- package/skills/skills-codex-claude-review/auto-paper-improvement-loop/SKILL.md +581 -0
- package/skills/skills-codex-claude-review/auto-review-loop/SKILL.md +510 -0
- package/skills/skills-codex-claude-review/novelty-check/SKILL.md +102 -0
- package/skills/skills-codex-claude-review/paper-figure/SKILL.md +319 -0
- package/skills/skills-codex-claude-review/paper-plan/SKILL.md +287 -0
- package/skills/skills-codex-claude-review/paper-write/SKILL.md +420 -0
- package/skills/skills-codex-claude-review/research-refine/SKILL.md +732 -0
- package/skills/skills-codex-claude-review/research-review/SKILL.md +149 -0
- package/skills/skills-codex-gemini-review/README.md +176 -0
- package/skills/skills-codex-gemini-review/README_CN.md +175 -0
- package/skills/skills-codex-gemini-review/auto-paper-improvement-loop/SKILL.md +331 -0
- package/skills/skills-codex-gemini-review/auto-review-loop/SKILL.md +304 -0
- package/skills/skills-codex-gemini-review/grant-proposal/SKILL.md +630 -0
- package/skills/skills-codex-gemini-review/idea-creator/SKILL.md +263 -0
- package/skills/skills-codex-gemini-review/idea-discovery/SKILL.md +275 -0
- package/skills/skills-codex-gemini-review/idea-discovery-robot/SKILL.md +365 -0
- package/skills/skills-codex-gemini-review/novelty-check/SKILL.md +92 -0
- package/skills/skills-codex-gemini-review/paper-figure/SKILL.md +289 -0
- package/skills/skills-codex-gemini-review/paper-plan/SKILL.md +265 -0
- package/skills/skills-codex-gemini-review/paper-poster-html/SKILL.md +102 -0
- package/skills/skills-codex-gemini-review/paper-slides/SKILL.md +582 -0
- package/skills/skills-codex-gemini-review/paper-write/SKILL.md +346 -0
- package/skills/skills-codex-gemini-review/paper-writing/SKILL.md +312 -0
- package/skills/skills-codex-gemini-review/research-refine/SKILL.md +674 -0
- package/skills/skills-codex-gemini-review/research-review/SKILL.md +112 -0
- package/skills/slides-polish/SKILL.md +565 -0
- package/skills/specification-writing/SKILL.md +211 -0
- package/skills/system-profile/SKILL.md +103 -0
- package/skills/training-check/SKILL.md +132 -0
- package/skills/vast-gpu/SKILL.md +394 -0
- package/skills/web-debug-search/SKILL.md +334 -0
- package/skills/wiki-enrich/SKILL.md +257 -0
- package/skills/writing-systems-papers/SKILL.md +184 -0
- package/templates/CLAUDE_MD_TEMPLATE.md +29 -0
- package/templates/EXPERIMENT_LOG_TEMPLATE.md +47 -0
- package/templates/EXPERIMENT_PLAN_TEMPLATE.md +51 -0
- package/templates/EXPERIMENT_PLAN_TEMPLATE_CN.md +53 -0
- package/templates/FINDINGS_TEMPLATE.md +52 -0
- package/templates/IDEA_CANDIDATES_TEMPLATE.md +47 -0
- package/templates/IDEA_CANDIDATES_TEMPLATE_CN.md +47 -0
- package/templates/INVENTION_BRIEF_TEMPLATE.md +87 -0
- package/templates/MANIFEST_TEMPLATE.md +7 -0
- package/templates/NARRATIVE_REPORT_TEMPLATE.md +49 -0
- package/templates/PAPER_PLAN_TEMPLATE.md +47 -0
- package/templates/PATENT_CLAIMS_TEMPLATE.md +78 -0
- package/templates/PATENT_SPECIFICATION_TEMPLATE.md +68 -0
- package/templates/README.md +57 -0
- package/templates/RESEARCH_BRIEF_TEMPLATE.md +35 -0
- package/templates/RESEARCH_BRIEF_TEMPLATE_CN.md +41 -0
- package/templates/RESEARCH_CONTRACT_TEMPLATE.md +60 -0
- package/templates/claude-hooks/corpus_write_guard.json +16 -0
- package/templates/claude-hooks/corpus_write_guard.py +85 -0
- package/templates/claude-hooks/meta_logging.json +74 -0
- package/templates/gitignore-trace.txt +3 -0
- package/tools/__pycache__/check_skills_inventory.cpython-314.pyc +0 -0
- package/tools/arxiv_fetch.py +311 -0
- package/tools/capture_filter.py +126 -0
- package/tools/check_skills_inventory.py +273 -0
- package/tools/convert_skills_to_llm_chat.py +282 -0
- package/tools/copilot_native_evidence.py +818 -0
- package/tools/deepxiv_fetch.py +213 -0
- package/tools/evidence_check.py +212 -0
- package/tools/exa_search.py +425 -0
- package/tools/experiment_queue/README.md +118 -0
- package/tools/experiment_queue/build_manifest.py +44 -0
- package/tools/experiment_queue/queue_manager.py +44 -0
- package/tools/extract_paper_style.py +560 -0
- package/tools/figure_renderer.py +69 -0
- package/tools/forensics_gate.py +669 -0
- package/tools/generate_codex_claude_review_overrides.py +299 -0
- package/tools/idea_discovery_gate.py +256 -0
- package/tools/install_aris.ps1 +1372 -0
- package/tools/install_aris.sh +1370 -0
- package/tools/install_aris_codex.sh +1023 -0
- package/tools/install_aris_copilot.sh +1052 -0
- package/tools/iteration_log.py +143 -0
- package/tools/lint_skills_helpers.sh +84 -0
- package/tools/meta_opt/check_ready.sh +80 -0
- package/tools/meta_opt/log_event.sh +91 -0
- package/tools/meta_opt/trigger_eval.py +280 -0
- package/tools/meta_opt/trigger_evals.sample.json +28 -0
- package/tools/openalex_fetch.py +326 -0
- package/tools/overleaf_audit.sh +104 -0
- package/tools/overleaf_setup.sh +150 -0
- package/tools/paper_illustration_image2.py +62 -0
- package/tools/provenance.py +294 -0
- package/tools/research_wiki.py +1720 -0
- package/tools/review_gate.py +502 -0
- package/tools/run_state.py +399 -0
- package/tools/save_trace.sh +477 -0
- package/tools/semantic_scholar_fetch.py +438 -0
- package/tools/skill-groups.tsv +116 -0
- package/tools/skill_picker.py +238 -0
- package/tools/smart_update.ps1 +521 -0
- package/tools/smart_update.sh +591 -0
- package/tools/smart_update_codex.sh +419 -0
- package/tools/smart_update_copilot.sh +605 -0
- package/tools/threat_scan.py +222 -0
- package/tools/verify_paper_audits.sh +487 -0
- package/tools/verify_papers.py +613 -0
- package/tools/verify_wiki_coverage.sh +176 -0
- package/tools/watchdog.py +485 -0
|
@@ -0,0 +1,163 @@
|
|
|
1
|
+
# Compute Environment Contract
|
|
2
|
+
|
|
3
|
+
> One declarative spec for what an environment IS; per-provider knowledge for
|
|
4
|
+
> how it gets built HERE; a content-hash ledger so "did the env change?" has a
|
|
5
|
+
> mechanical answer; and a three-tier validation ladder whose top tier is a
|
|
6
|
+
> fresh agent following the skill's own doc verbatim. Adapted from Anthropic's
|
|
7
|
+
> Claude Science `compute-env-setup` skill (Apache-2.0); de-coupled from its
|
|
8
|
+
> proprietary `host.*` runtime — everything here runs on plain bash + SSH +
|
|
9
|
+
> subagents.
|
|
10
|
+
|
|
11
|
+
Every ARIS compute skill (`/run-experiment`, `/experiment-queue`,
|
|
12
|
+
`/serverless-modal`, `/vast-gpu`, `/qzcli`) needs the same three things for a
|
|
13
|
+
job: a software stack (exact versions, often with load-bearing install order),
|
|
14
|
+
possibly large weights placed where the tool looks, and a resource shape. What
|
|
15
|
+
varies per provider is only HOW those materialize. Without a shared contract,
|
|
16
|
+
each skill re-encodes provider quirks and every "environment is ready" claim is
|
|
17
|
+
vibes. The classic failure this prevents: agent says "env ready", the overnight
|
|
18
|
+
run dies at `import flash_attn`, 8 GPUs idle until morning.
|
|
19
|
+
|
|
20
|
+
## 1. Provider shapes — recognize, don't choose
|
|
21
|
+
|
|
22
|
+
You are rarely choosing a shape; you are recognizing which one this provider
|
|
23
|
+
already is. The shape determines what "build", "register", and "resolve" mean.
|
|
24
|
+
|
|
25
|
+
| Shape | ARIS examples | Build = | Env name resolves to |
|
|
26
|
+
|---|---|---|---|
|
|
27
|
+
| **Direct SSH host** (conda/venv) | personal GPU boxes, lab servers | YOU are the renderer: `conda create -n <name> python=<X>`, then run `pip_phases` in order | the conda env name itself (`conda run -n <name> …`) |
|
|
28
|
+
| **Scheduler cluster** (Slurm/PBS; Qizhi-like platforms) | `/qzcli` targets | `module load` or a container image built OFF-cluster and pulled (compute nodes often have **no internet** — pre-stage everything) | scheduler directives + container path in shared scratch (mind purge windows) |
|
|
29
|
+
| **Managed API** (serverless) | `/serverless-modal`, Vast.ai templates | the provider's image definition (Modal `Image`, Vast template) — render the same spec into it | the provider's opaque image ref, recorded in the ledger |
|
|
30
|
+
|
|
31
|
+
## 2. The declarative spec (write WHAT once; render per provider)
|
|
32
|
+
|
|
33
|
+
```yaml
|
|
34
|
+
# env-spec: one dict per environment, portable across shapes
|
|
35
|
+
base: "cuda12.8 + python3.10" # FROM-image / conda create versions
|
|
36
|
+
system_pkgs: [git, tmux] # apt in a container; conda-forge subset on no-root hosts
|
|
37
|
+
pip_phases: # ORDERED list of lists — each inner list = ONE pip call
|
|
38
|
+
- [torch==2.8.0] # phase 1 first, so later packages
|
|
39
|
+
- [flash-attn --no-build-isolation] # can't drag torch to a wrong wheel
|
|
40
|
+
- [transformers, peft, accelerate]
|
|
41
|
+
env: {HF_ENDPOINT: "...", OMP_NUM_THREADS: "<tier.cpus>"}
|
|
42
|
+
run_commands: [] # escape-hatch shell (RUN / %post / plain SSH)
|
|
43
|
+
weight_dirs: {chai: {path: /scratch/weights/chai, source: "tool's own loader", gated: false}}
|
|
44
|
+
smoke: # probes that run INSIDE the env on every shape
|
|
45
|
+
import_names: [torch, flash_attn]
|
|
46
|
+
gpu_tests: # each = {cmd, expect}; expect is the witness regex
|
|
47
|
+
- cmd: "python -c 'import torch;torch.manual_seed(0);x=torch.randn(8,8,device=\"cuda\");print(\"WITNESS\", (x@x).shape, torch.cuda.get_device_name())'"
|
|
48
|
+
expect: "^WITNESS torch.Size"
|
|
49
|
+
cli_checks: [nvidia-smi]
|
|
50
|
+
```
|
|
51
|
+
|
|
52
|
+
- **`pip_phases` ordering IS the fix** for every "package A drags B to the
|
|
53
|
+
wrong version" problem: each phase is its own pip invocation, and pip leaves
|
|
54
|
+
an already-satisfied requirement alone unless asked to upgrade. Pin the
|
|
55
|
+
fought-over package in an EARLIER phase than the fighter.
|
|
56
|
+
- A clean spec renders unchanged through every renderer. If you find yourself
|
|
57
|
+
adding a field only one backend understands, that field belongs in the
|
|
58
|
+
provider's ledger entry, not the spec.
|
|
59
|
+
- **Weights**: small (<~500 MB) and read by every job → bake into the env at
|
|
60
|
+
build time. Large with a cache env var → persistent scratch + point the var
|
|
61
|
+
there. Populate with the **tool's own loader** (hand-curled layouts miss
|
|
62
|
+
marker files), then verify from the tool's perspective: run the real
|
|
63
|
+
entrypoint once against the staged dir and `du -sh` every subdir — 0 B means
|
|
64
|
+
a swallowed download error.
|
|
65
|
+
|
|
66
|
+
## 3. The environment ledger (content-hash = mechanical staleness)
|
|
67
|
+
|
|
68
|
+
Per provider, keep an append-friendly `.aris/compute/<provider>.md` (or the
|
|
69
|
+
project's existing server-notes file). One block per env, keyed by a content
|
|
70
|
+
hash of the spec, computed over an EXACT canonical form so two agents can
|
|
71
|
+
never hash the same spec differently: parse the spec file, re-serialize as
|
|
72
|
+
JSON with sorted keys and no whitespace, sha256, first 8 hex chars —
|
|
73
|
+
|
|
74
|
+
```bash
|
|
75
|
+
# spec stored as YAML (env-spec.yaml); requires PyYAML. If PyYAML is absent,
|
|
76
|
+
# store the spec as JSON instead and drop the yaml import — same pipeline.
|
|
77
|
+
python3 -c 'import sys,json,hashlib,yaml; \
|
|
78
|
+
s=json.dumps(yaml.safe_load(open(sys.argv[1])),sort_keys=True,separators=(",",":")); \
|
|
79
|
+
print(hashlib.sha256(s.encode()).hexdigest()[:8])' env-spec.yaml
|
|
80
|
+
```
|
|
81
|
+
|
|
82
|
+
Key order, comments, indentation, and trailing whitespace in the source file
|
|
83
|
+
do NOT affect the hash — only the parsed content does:
|
|
84
|
+
|
|
85
|
+
```
|
|
86
|
+
### env: dllm@a3f9c2e1
|
|
87
|
+
how: conda env "dllm" on <host> # or: modal image ref / .sif path + partition
|
|
88
|
+
tier: {cpus: 8, mem_gib: 64, gpus: 1}
|
|
89
|
+
weights: HF_HOME=/scratch/hf (24 GB; purge-window 30d)
|
|
90
|
+
validated: 2026-07-02 (witness + agent-follows-doc clean)
|
|
91
|
+
gotcha: <any diagnosis-table row hit on THIS provider>
|
|
92
|
+
```
|
|
93
|
+
|
|
94
|
+
Spec changed → hash changes → **cache miss**: the ledger entry no longer
|
|
95
|
+
matches and the env must be rebuilt (or a new block added). Spec unchanged →
|
|
96
|
+
warm-reuse without rebuilding or re-validating tier 1–2. This turns "I think
|
|
97
|
+
the env is the same as last week" into a string comparison. Note `.aris/` is
|
|
98
|
+
gitignored by convention — the ledger is **project-local and uncommitted** by
|
|
99
|
+
default (like `.aris/traces/`). If you want committed, git-blameable history,
|
|
100
|
+
keep the ledger blocks in the project's tracked server-notes file instead;
|
|
101
|
+
the block format is the contract, not the path.
|
|
102
|
+
|
|
103
|
+
## 4. Validation — three tiers; the gap between them is where debugging lives
|
|
104
|
+
|
|
105
|
+
1. **Import works** — `python -c "import <pkg>"` exits 0. Necessary, cheap,
|
|
106
|
+
catches almost nothing interesting.
|
|
107
|
+
2. **Kernel-dispatch witness** — a tiny SEEDED forward pass that prints a
|
|
108
|
+
sentinel line (output shape + device name + non-emptiness). Catches "torch
|
|
109
|
+
sees the GPU but the kernel was compiled for an older SM", "the compiled
|
|
110
|
+
extension's `.so` isn't on the loader path", "inference writes to a
|
|
111
|
+
read-only cache". Keep the witness command in the spec's `smoke.gpu_tests`
|
|
112
|
+
with an `expect:` regex so the SAME probe runs on every backend. Cheap —
|
|
113
|
+
run on every build.
|
|
114
|
+
3. **Agent-follows-doc** — the validation that actually matters and the one
|
|
115
|
+
that's easy to skip. Spawn a FRESH subagent that gets ONLY: the compute
|
|
116
|
+
skill's doc, the provider's ledger entry, and the documented invocation.
|
|
117
|
+
It must run the invocation **verbatim** — no improvisation, no fixing —
|
|
118
|
+
and report every point where the doc's claim and reality diverge. This is
|
|
119
|
+
where you find the doc says `--ligand` but the flag is
|
|
120
|
+
`--ligand_description`, or the weights path exists but lacks the completion
|
|
121
|
+
marker the tool checks. The author agent cannot self-certify its own doc
|
|
122
|
+
(it walks through on hidden knowledge the doc never wrote down — same
|
|
123
|
+
principle as `acceptance-gate.md`: the writer never acquits its own
|
|
124
|
+
artifact); the fresh agent's stuck-point IS the doc's lie. Expensive —
|
|
125
|
+
reserve for the two moments doc and env can drift: **after any env rebuild
|
|
126
|
+
or doc edit, and before declaring an env ready**.
|
|
127
|
+
|
|
128
|
+
## 5. Diagnosis table (symptom → layer → fix)
|
|
129
|
+
|
|
130
|
+
When a documented invocation fails, don't patch reflexively — ask which LAYER
|
|
131
|
+
is wrong: spec, build, weights, resolution, or doc. Grep-able rows (container
|
|
132
|
+
rows apply only to container shapes):
|
|
133
|
+
|
|
134
|
+
| Symptom | Layer | Fix |
|
|
135
|
+
|---|---|---|
|
|
136
|
+
| `no kernel image is available for execution` | build/spec | torch compiled for older SM than this GPU — record `sm_range` in the ledger and route jobs; rebuild only if no compatible hardware |
|
|
137
|
+
| `ModuleNotFoundError` for a package not in the spec | spec | a `--no-deps` install skipped a runtime dep — read the package's `pyproject.toml` and add an explicit phase |
|
|
138
|
+
| Wrong torch/numpy version after install | spec | a later package's pin won — add a `force-reinstall --no-deps` snap-back phase after it |
|
|
139
|
+
| `ImportError: libfoo.so: cannot open shared object` | build | compiled `.so` not on loader path — `find` it, add its dir to `LD_LIBRARY_PATH` |
|
|
140
|
+
| Tool re-downloads despite populated weights | weights | `du -sh $CACHE_VAR` first: 0 B = swallowed error; non-zero = tool checks a marker file, stage that too |
|
|
141
|
+
| `OSError: Read-only file system` under cache var | weights (container) | tool writes locks next to weights on an RO mount — symlink blobs into writable `/tmp` cache |
|
|
142
|
+
| 80-way thread storm on a 4-CPU allocation | exec | `os.cpu_count()` returns the HOST's cores — export `OMP/MKL/OPENBLAS_NUM_THREADS=<tier.cpus>` on every backend |
|
|
143
|
+
| First job slow, every later job equally slow | build | expensive precompute runs at job time in a non-persistent workdir — run it once at build time |
|
|
144
|
+
| Job COMPLETED but output dir empty | exec | the wrapper writing the completion marker never ran — often `#!/bin/bash` on a runtime that only ships `/bin/sh` |
|
|
145
|
+
|
|
146
|
+
Hit a row on a specific provider → append symptom + fix to that provider's
|
|
147
|
+
ledger `gotcha:` line, so the next agent doesn't rediscover it.
|
|
148
|
+
|
|
149
|
+
## How compute skills use this
|
|
150
|
+
|
|
151
|
+
- **Before building**: read the provider's ledger. The env — or a near-match
|
|
152
|
+
to extend — may already exist; an unchanged hash means skip the rebuild.
|
|
153
|
+
- **When building**: write the spec first (§2), render it for the shape (§1),
|
|
154
|
+
run tier-1/2 validation (§4), append the ledger block (§3).
|
|
155
|
+
- **Before declaring ready** (and after any rebuild/doc edit): run the
|
|
156
|
+
agent-follows-doc pass (§4.3).
|
|
157
|
+
- **On failure**: diagnosis table (§5) before patching; record provider-true
|
|
158
|
+
gotchas in the ledger.
|
|
159
|
+
|
|
160
|
+
Attribution: the spec/ledger/three-tier-validation design is adapted from
|
|
161
|
+
Anthropic's Claude Science `compute-env-setup` skill (Apache-2.0, re-hosted by
|
|
162
|
+
HughYau/AcademicForge); this document ports it off the proprietary `host.*`
|
|
163
|
+
runtime onto plain bash + SSH + ARIS subagents.
|
|
@@ -0,0 +1,183 @@
|
|
|
1
|
+
# Effort Contract
|
|
2
|
+
|
|
3
|
+
## Overview
|
|
4
|
+
|
|
5
|
+
Every ARIS skill accepts an optional `effort` parameter that controls how much work the system does. This affects breadth, depth, iterations, and coverage — but **never** the quality of cross-model review.
|
|
6
|
+
|
|
7
|
+
> Design stance: the unattended *procedure* (gather → reason → act → verify → repeat) is the engineered artifact, not any single prompt — after Karpathy's "write the loop, not the prompt" (LOOPS.md, *Field Notes on Agents That Run for Days*).
|
|
8
|
+
|
|
9
|
+
```
|
|
10
|
+
/any-skill "args" — effort: lite | balanced | max | beast
|
|
11
|
+
```
|
|
12
|
+
|
|
13
|
+
Default: `balanced` (current behavior, zero change for existing users).
|
|
14
|
+
|
|
15
|
+
## Hard Invariants (NEVER changed by effort)
|
|
16
|
+
|
|
17
|
+
| Setting | Value | Why |
|
|
18
|
+
|---------|-------|-----|
|
|
19
|
+
| Codex reasoning_effort | **≥ xhigh** (deep-audit skills run `ultra` — tier table in `reviewer-routing.md`) | Reviewer quality is non-negotiable. `effort` never moves the reviewer tier in either direction — and ARIS `— effort: max` is NOT Codex `model_reasoning_effort: max` (different axes: pipeline workload vs reviewer reasoning depth) |
|
|
20
|
+
| DBLP/CrossRef citations | **on** | Citation integrity is non-negotiable |
|
|
21
|
+
| Reviewer independence | **on** | Cross-model protocol is non-negotiable |
|
|
22
|
+
| Experiment integrity | **on** | Fraud prevention is non-negotiable |
|
|
23
|
+
| Sanity check | **on** | Safety is non-negotiable |
|
|
24
|
+
| **Mandatory audit emission** | **always** | At `assurance: submission`, every mandatory audit emits a verdict (PASS/WARN/FAIL/NOT_APPLICABLE/BLOCKED/ERROR). Silent skip is forbidden. See `assurance-contract.md`. |
|
|
25
|
+
| AUTO_PROCEED | **user decides** | Orthogonal to effort |
|
|
26
|
+
| difficulty | **user decides** | Orthogonal to effort |
|
|
27
|
+
| `assurance` | **derived from `effort`** (see Assurance Axis below) | Audit strictness is a separate axis from depth |
|
|
28
|
+
|
|
29
|
+
## Four Levels
|
|
30
|
+
|
|
31
|
+
### `lite` (~0.4x tokens)
|
|
32
|
+
For budget-constrained users or quick explorations. Minimum viable depth.
|
|
33
|
+
Implies `assurance: draft` (see below).
|
|
34
|
+
|
|
35
|
+
### `balanced` (1x tokens) — DEFAULT
|
|
36
|
+
Current ARIS behavior. What existing users get today. No breakage.
|
|
37
|
+
Implies `assurance: draft`.
|
|
38
|
+
|
|
39
|
+
### `max` (~2.5x tokens)
|
|
40
|
+
Go deeper than defaults. More papers, more ideas, more rounds, more detail.
|
|
41
|
+
Implies `assurance: submission` — mandatory audits are load-bearing.
|
|
42
|
+
|
|
43
|
+
### `beast` (~5-8x tokens)
|
|
44
|
+
No budget limit. Every knob to maximum. For top-venue submission sprints.
|
|
45
|
+
Implies `assurance: submission` — mandatory audits are load-bearing and the
|
|
46
|
+
final report is tagged `submission-ready` only when the verifier agrees.
|
|
47
|
+
|
|
48
|
+
## Assurance Axis (separate concern from `effort`)
|
|
49
|
+
|
|
50
|
+
Audit strictness lives on a second axis, `assurance`. Full contract:
|
|
51
|
+
**`shared-references/assurance-contract.md`**.
|
|
52
|
+
|
|
53
|
+
```
|
|
54
|
+
— assurance: draft | submission
|
|
55
|
+
```
|
|
56
|
+
|
|
57
|
+
Default mapping (if `assurance` not given explicitly):
|
|
58
|
+
|
|
59
|
+
| `effort` | implied `assurance` | Behavior |
|
|
60
|
+
|----------|---------------------|----------|
|
|
61
|
+
| `lite` | `draft` | Audits run only if content detector matches; silent skip allowed |
|
|
62
|
+
| `balanced` | `draft` | Same as lite — current behavior, zero breakage |
|
|
63
|
+
| `max` | `submission` | Every mandatory audit emits a verdict; verifier blocks Final Report on FAIL/BLOCKED/ERROR/STALE |
|
|
64
|
+
| `beast` | `submission` | Same as max + final report tagged `submission-ready` |
|
|
65
|
+
|
|
66
|
+
User can override independently:
|
|
67
|
+
- `— effort: balanced, assurance: submission` → normal depth, strict audits
|
|
68
|
+
- `— effort: beast, assurance: draft` → maximum depth, no audit gate (legal but discouraged for real submissions)
|
|
69
|
+
|
|
70
|
+
**Why split the axes?** Historically `effort: beast` did not enforce audits — phases like `/proof-checker`, `/paper-claim-audit`, `/citation-audit` were gated by content detectors that allowed silent skip. A user reported `effort: beast` produced a "draft-quality" paper with all three submission gates skipped. The split makes audit strictness independently verifiable and stops conflating "do more work" with "be more rigorous."
|
|
71
|
+
|
|
72
|
+
## Per-Skill Profiles
|
|
73
|
+
|
|
74
|
+
### Discovery & Planning
|
|
75
|
+
|
|
76
|
+
| Skill | Dimension | lite | balanced | max | beast |
|
|
77
|
+
|-------|-----------|------|----------|-----|-------|
|
|
78
|
+
| research-lit | papers found | 6-8 | 10-15 | 18-25 | 40-50 |
|
|
79
|
+
| research-lit | query variants | 2 | 5 | 8 | 15+ |
|
|
80
|
+
| research-lit | deep reads | 3 | 5-8 | 8 | 15+ |
|
|
81
|
+
| idea-creator | ideas generated | 4-6 | 8-12 | 12-16 | 20-30 |
|
|
82
|
+
| idea-creator | pilots | 1-2 | 2-3 | 3-4 | 5-6 |
|
|
83
|
+
| novelty-check | claims checked | 2-3 | 3-4 | 4-6 | all |
|
|
84
|
+
| novelty-check | closest works | top-3 | top-5 | top-8 | top-10+ |
|
|
85
|
+
| research-refine | max rounds | 3 | 5 | 7 | 10+ |
|
|
86
|
+
| research-refine | papers considered | 8 | 15 | 24 | 30+ |
|
|
87
|
+
| experiment-plan | core experiments | 3 | 5 | 7 | 10+ |
|
|
88
|
+
| experiment-plan | seeds | 1 | 3 | 5 | 5 |
|
|
89
|
+
| experiment-plan | baseline families | 2 | 3 | 4 | 5+ |
|
|
90
|
+
|
|
91
|
+
### Execution
|
|
92
|
+
|
|
93
|
+
| Skill | Dimension | lite | balanced | max | beast |
|
|
94
|
+
|-------|-----------|------|----------|-----|-------|
|
|
95
|
+
| experiment-bridge | scope | sanity + main | main + basic ablation | + top ablation + robustness | full suite + cross-validation |
|
|
96
|
+
| run-experiment | launches | smoke + main | smoke + multi-seed | + dry run + manifest | full config + multi-GPU parallel |
|
|
97
|
+
| monitor-experiment | depth | latest log | log + JSON | + W&B + anomaly | real-time + auto-alert + trend |
|
|
98
|
+
| analyze-results | findings | 3 | 5 | 8 | full-dimensional + stat tests |
|
|
99
|
+
| ablation-planner | ablations | 2-3 | 4-5 | 6-8 | 10+ |
|
|
100
|
+
|
|
101
|
+
### Review
|
|
102
|
+
|
|
103
|
+
| Skill | Dimension | lite | balanced | max | beast |
|
|
104
|
+
|-------|-----------|------|----------|-----|-------|
|
|
105
|
+
| auto-review-loop | max rounds | 2 | 3-4 | 6 | 8+ (until converged) |
|
|
106
|
+
| auto-review-loop | fixes per round | 1-2 | 3-4 | 4-6 | all actionable |
|
|
107
|
+
| research-review | passes | 1 | 1 + follow-up | 1 + 2 follow-ups | 2 independent + cross-compare |
|
|
108
|
+
| experiment-audit | depth | skip | basic 4 checks | full 6 checks | line-by-line + reproduce |
|
|
109
|
+
|
|
110
|
+
### Writing & Rebuttal
|
|
111
|
+
|
|
112
|
+
| Skill | Dimension | lite | balanced | max | beast |
|
|
113
|
+
|-------|-----------|------|----------|-----|-------|
|
|
114
|
+
| paper-plan | outline reviews | 0 | 1 | 2 | 3 |
|
|
115
|
+
| paper-plan | citations/section | 2-3 | 4-5 | 5-8 | 8+ |
|
|
116
|
+
| paper-figure | caption reviews | 1 | 1 | 2 | 3 |
|
|
117
|
+
| paper-write | abstract variants | 1 | 1 | 2 | 3 |
|
|
118
|
+
| paper-write | related work depth | shallow | standard | deep | exhaustive |
|
|
119
|
+
| paper-compile | fix attempts | 2 | 3 | 4 | until zero warnings |
|
|
120
|
+
| auto-paper-improvement | rounds | 1 | 2 | 3 | 5 |
|
|
121
|
+
| paper-illustration | render iterations | 2 | 3 | 5 | 7 |
|
|
122
|
+
| rebuttal | draft rounds | 1 | 2 | 3 | 5 |
|
|
123
|
+
| rebuttal | stress tests | 0-1 | 1 | 2 | 3 |
|
|
124
|
+
|
|
125
|
+
## How to Read Effort in a Skill
|
|
126
|
+
|
|
127
|
+
Add this to the Constants section of each skill:
|
|
128
|
+
|
|
129
|
+
```markdown
|
|
130
|
+
## Constants
|
|
131
|
+
|
|
132
|
+
- **EFFORT = `balanced`** — Work intensity. Options: `lite`, `balanced`, `max`, `beast`. Override: `— effort: max`
|
|
133
|
+
```
|
|
134
|
+
|
|
135
|
+
Then adjust numeric constants based on effort level. Example:
|
|
136
|
+
|
|
137
|
+
```
|
|
138
|
+
Parse $ARGUMENTS for `— effort:` directive.
|
|
139
|
+
If not specified, default to `balanced`.
|
|
140
|
+
|
|
141
|
+
Adjust constants:
|
|
142
|
+
if effort == "lite": MAX_PAPERS = 8, MAX_IDEAS = 6, MAX_ROUNDS = 2
|
|
143
|
+
if effort == "balanced": MAX_PAPERS = 15, MAX_IDEAS = 12, MAX_ROUNDS = 4
|
|
144
|
+
if effort == "max": MAX_PAPERS = 25, MAX_IDEAS = 16, MAX_ROUNDS = 6
|
|
145
|
+
if effort == "beast": MAX_PAPERS = 50, MAX_IDEAS = 30, MAX_ROUNDS = 8
|
|
146
|
+
```
|
|
147
|
+
|
|
148
|
+
## Transparency
|
|
149
|
+
|
|
150
|
+
Every skill should print its effort configuration at the start:
|
|
151
|
+
|
|
152
|
+
```
|
|
153
|
+
⚡ [effort: max] papers=25, ideas=16, rounds=6 | Codex: tier per reviewer-routing.md (floor xhigh)
|
|
154
|
+
```
|
|
155
|
+
|
|
156
|
+
## Precedence
|
|
157
|
+
|
|
158
|
+
```
|
|
159
|
+
explicit concrete knob (e.g., review_rounds: 2)
|
|
160
|
+
> explicit dimension override
|
|
161
|
+
> overall effort level
|
|
162
|
+
> skill default (balanced)
|
|
163
|
+
```
|
|
164
|
+
|
|
165
|
+
Example: `— effort: beast, review_rounds: 3` → everything beast except review capped at 3.
|
|
166
|
+
|
|
167
|
+
For the `assurance` axis, precedence is independent:
|
|
168
|
+
```
|
|
169
|
+
explicit `assurance: ...` directive
|
|
170
|
+
> effort-implied default (lite/balanced → draft, max/beast → submission)
|
|
171
|
+
> skill default (draft)
|
|
172
|
+
```
|
|
173
|
+
|
|
174
|
+
Example: `— effort: balanced, assurance: submission` → normal depth knobs but submission-gate audit enforcement.
|
|
175
|
+
|
|
176
|
+
## Token Cost Estimation
|
|
177
|
+
|
|
178
|
+
| Level | LLM tokens | GPU/wall-clock | Best for |
|
|
179
|
+
|-------|-----------|----------------|----------|
|
|
180
|
+
| lite | ~0.4x | ~0.5x | Quick exploration, budget users |
|
|
181
|
+
| balanced | 1x | 1x | Normal research workflow |
|
|
182
|
+
| max | ~2.5x | ~2x | Serious submission prep |
|
|
183
|
+
| beast | ~5-8x | ~3-4x | Top-venue final sprint |
|
|
@@ -0,0 +1,65 @@
|
|
|
1
|
+
# Evidence Pre-check
|
|
2
|
+
|
|
3
|
+
ARIS's claim audits (`/result-to-claim`, `/experiment-audit`, `/paper-claim-audit`)
|
|
4
|
+
spend a cross-model (codex/gemini) call to judge whether a claim is supported. The
|
|
5
|
+
cheapest, most common integrity failure is *hallucinated evidence*: a claim cites
|
|
6
|
+
a number + a source file, and the file doesn't exist or the number isn't in it.
|
|
7
|
+
You should not need a model call to catch that.
|
|
8
|
+
|
|
9
|
+
## Two stages — and `verified` ≠ `correct`
|
|
10
|
+
|
|
11
|
+
```
|
|
12
|
+
stage 1 tools/evidence_check.py deterministic · no model · fail-closed
|
|
13
|
+
catches HALLUCINATION — cited path missing, or cited value not in source.
|
|
14
|
+
stage 2 the cross-model jury codex/gemini
|
|
15
|
+
catches WRONG-BUT-REAL — the number IS in the file, but it doesn't
|
|
16
|
+
support the claim.
|
|
17
|
+
```
|
|
18
|
+
|
|
19
|
+
A `verified` from stage 1 means **only that the cited evidence exists** — never
|
|
20
|
+
that the claim holds. Existence is execution-completeness (deterministic / safe
|
|
21
|
+
same-model); *support* is a quality verdict that stays with the cross-model jury
|
|
22
|
+
(`acceptance-gate.md`: the pre-check DRIVES a gate, it cannot ACQUIT a claim).
|
|
23
|
+
This is the reconcile pattern — a model's self-report cross-checked against
|
|
24
|
+
mechanical ground-truth (adapted from Hermes's curator reconcile-classifier),
|
|
25
|
+
made into a cheap pre-gate that catches hallucination *before* the jury runs and
|
|
26
|
+
spares the codex call on fabricated evidence.
|
|
27
|
+
|
|
28
|
+
## Conservative by design
|
|
29
|
+
|
|
30
|
+
The pre-check favors **false-negative over false-positive**: when in doubt it
|
|
31
|
+
returns not-verified and lets the jury decide — it must never emit a false
|
|
32
|
+
`verified`. A pure number is matched by **numeric-token equality** (so `73.2`
|
|
33
|
+
matches `73.20` but `73` does NOT match `73.5`); a non-numeric value by
|
|
34
|
+
normalized substring.
|
|
35
|
+
|
|
36
|
+
## Where ARIS uses it
|
|
37
|
+
|
|
38
|
+
- **`/result-to-claim`** Step 1.5: parse each claim's cited `(value, source)`,
|
|
39
|
+
run the batch pre-check, and **before the codex judgment** mark any claim whose
|
|
40
|
+
evidence is `path_missing` / `value_not_found` as **unsupported — evidence not
|
|
41
|
+
found**, and pass the per-claim pre-check status into the codex prompt so the
|
|
42
|
+
jury sees which claims have verified vs hallucinated evidence.
|
|
43
|
+
- **To extend:** `/experiment-audit` (the "phantom results" check is exactly
|
|
44
|
+
this) and `/paper-claim-audit` (every reported number → its result file).
|
|
45
|
+
|
|
46
|
+
## API / CLI
|
|
47
|
+
|
|
48
|
+
```
|
|
49
|
+
from evidence_check import check_claim, check_batch
|
|
50
|
+
check_claim(value, source, root=".") # -> {status: verified|path_missing|value_not_found, ...}
|
|
51
|
+
check_batch([{value, source, id?}, ...], root) # -> {results:[...], summary:{status: n}}
|
|
52
|
+
```
|
|
53
|
+
```
|
|
54
|
+
python3 tools/evidence_check.py <root> --value 73.2 --source results/eval.json # exit 0 verified
|
|
55
|
+
python3 tools/evidence_check.py <root> --batch claims.json # exit 1 if any claim hallucinated
|
|
56
|
+
```
|
|
57
|
+
|
|
58
|
+
## Cross-references
|
|
59
|
+
- `acceptance-gate.md` — the pre-check is the deterministic DRIVE; the jury is the
|
|
60
|
+
ACQUIT. `verified` is existence (execution-completeness), not correctness.
|
|
61
|
+
- `reviewer-independence.md` — the jury still reads the artifacts itself; the
|
|
62
|
+
pre-check only flags which claims have evidence to read, never pre-digests the
|
|
63
|
+
verdict.
|
|
64
|
+
- `experiment-integrity.md` — fabricated/phantom results are exactly what stage 1
|
|
65
|
+
catches deterministically before stage 2.
|
|
@@ -0,0 +1,49 @@
|
|
|
1
|
+
# Experiment Integrity Protocol
|
|
2
|
+
|
|
3
|
+
## Core Principle
|
|
4
|
+
|
|
5
|
+
**The model that writes experiment code must NOT be the model that judges experiment integrity.** This is the same principle as reviewer-independence, applied to experiments.
|
|
6
|
+
|
|
7
|
+
## Prohibited Patterns
|
|
8
|
+
|
|
9
|
+
### 1. Fake Ground Truth
|
|
10
|
+
- ❌ Creating synthetic "reference" from model outputs and comparing against it
|
|
11
|
+
- ❌ Using baseline model outputs as ground truth
|
|
12
|
+
- ❌ Generating pseudo-GT that is structurally similar to predictions
|
|
13
|
+
- ✅ Using dataset-provided ground truth
|
|
14
|
+
- ✅ Using official evaluation scripts when available
|
|
15
|
+
- ✅ Proxy evaluation is allowed IF explicitly labeled as `synthetic_proxy`
|
|
16
|
+
|
|
17
|
+
### 2. Score Normalization Fraud
|
|
18
|
+
- ❌ Dividing metrics by max/min of model's own output to get 0.99+
|
|
19
|
+
- ❌ Rescaling scores to hide poor performance
|
|
20
|
+
- ✅ Standard normalization (e.g., min-max across ALL methods including baselines)
|
|
21
|
+
- ✅ Reporting raw and normalized scores side by side
|
|
22
|
+
|
|
23
|
+
### 3. Phantom Results
|
|
24
|
+
- ❌ Claiming results from files that don't exist
|
|
25
|
+
- ❌ Referencing metrics from functions that are never called
|
|
26
|
+
- ❌ Reporting TRACKER status as DONE when it's still TODO
|
|
27
|
+
- ✅ Every claimed number must trace to an actual output file
|
|
28
|
+
|
|
29
|
+
### 4. Insufficient Scope
|
|
30
|
+
- ❌ Reporting 2-scene pilot as "comprehensive evaluation"
|
|
31
|
+
- ❌ Using words like "robust", "extensive", "across settings" for tiny experiments
|
|
32
|
+
- ✅ Honestly label scope: "pilot (N=2)", "preliminary", "limited evaluation"
|
|
33
|
+
- ✅ State exact scope: N scenes, N seeds, N configurations
|
|
34
|
+
|
|
35
|
+
## Evaluation Types (must be declared)
|
|
36
|
+
|
|
37
|
+
| Type | Label | What it means | Claim ceiling |
|
|
38
|
+
|------|-------|---------------|---------------|
|
|
39
|
+
| Real GT | `real_gt` | Dataset-provided ground truth | Full performance claims |
|
|
40
|
+
| Synthetic proxy | `synthetic_proxy` | Model-generated reference | "Proxy consistency" only |
|
|
41
|
+
| Self-supervised | `self_supervised_proxy` | No GT by design | Relative improvement only |
|
|
42
|
+
| Simulation | `simulation_only` | Simulated environment | "In simulation" qualifier |
|
|
43
|
+
| Human eval | `human_eval` | Human judges | Subject to inter-rater stats |
|
|
44
|
+
|
|
45
|
+
## Who Checks
|
|
46
|
+
|
|
47
|
+
The **reviewer model** (different family from executor) performs integrity checks via `/experiment-audit`. The executor collects file paths; the reviewer reads code and results directly.
|
|
48
|
+
|
|
49
|
+
**Never let the executor judge its own experiment integrity.**
|