dsh-aris-panel 0.1.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/LICENSE +21 -0
- package/README.md +98 -0
- package/README_CN.md +87 -0
- package/dsh/checkout.patch.yml +38 -0
- package/dsh/client.js +634 -0
- package/dsh/cordis.patch.yml +44 -0
- package/dsh/index.mjs +76 -0
- package/dsh/run-status.mjs +182 -0
- package/dsh/scope-limits.mjs +50 -0
- package/dsh/workbench.mjs +291 -0
- package/mcp-servers/claude-review/README.md +93 -0
- package/mcp-servers/claude-review/run_with_claude_aws.sh +49 -0
- package/mcp-servers/claude-review/server.py +718 -0
- package/mcp-servers/codex-image2/README.md +65 -0
- package/mcp-servers/codex-image2/server.py +893 -0
- package/mcp-servers/feishu-bridge/requirements.txt +1 -0
- package/mcp-servers/feishu-bridge/server.py +240 -0
- package/mcp-servers/gemini-review/README.md +171 -0
- package/mcp-servers/gemini-review/server.py +1856 -0
- package/mcp-servers/llm-chat/requirements.txt +1 -0
- package/mcp-servers/llm-chat/server.py +664 -0
- package/mcp-servers/manual-review/README.md +133 -0
- package/mcp-servers/manual-review/server.py +910 -0
- package/mcp-servers/manual-review/ui.html +279 -0
- package/mcp-servers/minimax-chat/requirements.txt +1 -0
- package/mcp-servers/minimax-chat/server.py +381 -0
- package/package.json +51 -0
- package/skills/ablation-planner/SKILL.md +123 -0
- package/skills/alphaxiv/SKILL.md +196 -0
- package/skills/analyze-results/SKILL.md +46 -0
- package/skills/arxiv/SKILL.md +248 -0
- package/skills/auto-paper-improvement-loop/SKILL.md +651 -0
- package/skills/auto-review-loop/SKILL.md +1137 -0
- package/skills/auto-review-loop-llm/SKILL.md +259 -0
- package/skills/auto-review-loop-minimax/SKILL.md +302 -0
- package/skills/citation-audit/SKILL.md +502 -0
- package/skills/claims-drafting/SKILL.md +227 -0
- package/skills/comm-lit-review/SKILL.md +297 -0
- package/skills/deepxiv/SKILL.md +263 -0
- package/skills/dse-loop/SKILL.md +296 -0
- package/skills/embodiment-description/SKILL.md +129 -0
- package/skills/exa-search/SKILL.md +205 -0
- package/skills/experiment-audit/SKILL.md +311 -0
- package/skills/experiment-bridge/SKILL.md +376 -0
- package/skills/experiment-plan/SKILL.md +249 -0
- package/skills/experiment-queue/SKILL.md +431 -0
- package/skills/experiment-queue/scripts/build_manifest.py +142 -0
- package/skills/experiment-queue/scripts/queue_manager.py +433 -0
- package/skills/feishu-notify/SKILL.md +156 -0
- package/skills/figure-description/SKILL.md +138 -0
- package/skills/figure-spec/SKILL.md +262 -0
- package/skills/figure-spec/scripts/figure_renderer.py +799 -0
- package/skills/formula-derivation/SKILL.md +280 -0
- package/skills/gemini-search/SKILL.md +231 -0
- package/skills/grant-proposal/SKILL.md +698 -0
- package/skills/idea-creator/SKILL.md +542 -0
- package/skills/idea-discovery/SKILL.md +521 -0
- package/skills/idea-discovery-robot/SKILL.md +363 -0
- package/skills/integrity-forensics/SKILL.md +284 -0
- package/skills/interview-cheatsheet/SKILL.md +245 -0
- package/skills/invention-structuring/SKILL.md +188 -0
- package/skills/jurisdiction-format/SKILL.md +192 -0
- package/skills/kill-argument/SKILL.md +437 -0
- package/skills/mermaid-diagram/SKILL.md +419 -0
- package/skills/meta-apply/SKILL.md +141 -0
- package/skills/meta-optimize/SKILL.md +437 -0
- package/skills/monitor-experiment/SKILL.md +140 -0
- package/skills/novelty-check/SKILL.md +101 -0
- package/skills/openalex/SKILL.md +237 -0
- package/skills/overleaf-sync/SKILL.md +220 -0
- package/skills/paper-claim-audit/SKILL.md +348 -0
- package/skills/paper-compile/SKILL.md +266 -0
- package/skills/paper-figure/SKILL.md +312 -0
- package/skills/paper-illustration/SKILL.md +736 -0
- package/skills/paper-illustration-image2/SKILL.md +391 -0
- package/skills/paper-illustration-image2/scripts/paper_illustration_image2.py +255 -0
- package/skills/paper-plan/SKILL.md +386 -0
- package/skills/paper-poster/SKILL.md +19 -0
- package/skills/paper-poster-html/DESIGN_FINAL.md +176 -0
- package/skills/paper-poster-html/IMPLEMENTATION_CONVENTIONS.md +161 -0
- package/skills/paper-poster-html/LICENSES/posterly-MIT.txt +21 -0
- package/skills/paper-poster-html/NOTICE.md +57 -0
- package/skills/paper-poster-html/SKILL.md +323 -0
- package/skills/paper-poster-html/scripts/_posterly/__init__.py +0 -0
- package/skills/paper-poster-html/scripts/_posterly/canvas.py +200 -0
- package/skills/paper-poster-html/scripts/_posterly/measure.py +588 -0
- package/skills/paper-poster-html/scripts/_posterly/polish.py +498 -0
- package/skills/paper-poster-html/scripts/_posterly/preflight.py +489 -0
- package/skills/paper-poster-html/scripts/_posterly/render.py +215 -0
- package/skills/paper-poster-html/scripts/_posterly/textutil.py +16 -0
- package/skills/paper-poster-html/scripts/_posterly/verify_final.py +171 -0
- package/skills/paper-poster-html/scripts/asset_check.py +897 -0
- package/skills/paper-poster-html/scripts/extract_pdf_figures.py +666 -0
- package/skills/paper-poster-html/scripts/poster_check.py +251 -0
- package/skills/paper-poster-html/scripts/preprocess_figures.py +238 -0
- package/skills/paper-poster-html/scripts/render_preview.py +217 -0
- package/skills/paper-poster-html/scripts/run_gates.py +556 -0
- package/skills/paper-poster-html/scripts/style_check.py +1324 -0
- package/skills/paper-poster-html/templates/COMPONENTS.md +462 -0
- package/skills/paper-poster-html/templates/README.md +170 -0
- package/skills/paper-poster-html/templates/landscape_4col.html +1032 -0
- package/skills/paper-poster-html/templates/landscape_hero.html +1046 -0
- package/skills/paper-poster-html/templates/portrait_2col.html +947 -0
- package/skills/paper-poster-html/templates/tokens/acl.json +9 -0
- package/skills/paper-poster-html/templates/tokens/cvpr.json +9 -0
- package/skills/paper-poster-html/templates/tokens/generic.json +9 -0
- package/skills/paper-poster-html/templates/tokens/iclr.json +9 -0
- package/skills/paper-poster-html/templates/tokens/icml.json +9 -0
- package/skills/paper-poster-html/templates/tokens/neurips.json +9 -0
- package/skills/paper-slides/SKILL.md +635 -0
- package/skills/paper-talk/SKILL.md +381 -0
- package/skills/paper-write/SKILL.md +604 -0
- package/skills/paper-write/templates/IEEEtran.bst +2409 -0
- package/skills/paper-write/templates/IEEEtran.cls +6347 -0
- package/skills/paper-write/templates/iclr2026.tex +84 -0
- package/skills/paper-write/templates/icml2025.tex +87 -0
- package/skills/paper-write/templates/ieee_conference.tex +89 -0
- package/skills/paper-write/templates/ieee_journal.tex +93 -0
- package/skills/paper-write/templates/math_commands.tex +48 -0
- package/skills/paper-write/templates/neurips2025.tex +80 -0
- package/skills/paper-writing/SKILL.md +916 -0
- package/skills/patent-novelty-check/SKILL.md +153 -0
- package/skills/patent-pipeline/SKILL.md +344 -0
- package/skills/patent-review/SKILL.md +203 -0
- package/skills/pixel-art/SKILL.md +137 -0
- package/skills/prior-art-search/SKILL.md +146 -0
- package/skills/proof-checker/SKILL.md +866 -0
- package/skills/proof-orchestrator/NOTICE.md +24 -0
- package/skills/proof-orchestrator/SKILL.md +254 -0
- package/skills/proof-orchestrator/references/audit-output-contract.md +126 -0
- package/skills/proof-orchestrator/references/deepseek-routing.md +74 -0
- package/skills/proof-orchestrator/references/dispatch-prompts.md +227 -0
- package/skills/proof-orchestrator/references/notation-audit.md +135 -0
- package/skills/proof-orchestrator/references/proof-audit-rubric.md +70 -0
- package/skills/proof-orchestrator/references/stress-tests.md +38 -0
- package/skills/proof-writer/SKILL.md +223 -0
- package/skills/qzcli/SKILL.md +324 -0
- package/skills/rebuttal/SKILL.md +376 -0
- package/skills/render-html/SKILL.md +316 -0
- package/skills/render-html/scripts/render_html.py +1006 -0
- package/skills/render-html/scripts/templates/academic.html +703 -0
- package/skills/render-html/scripts/templates/dashboard.html +333 -0
- package/skills/research-lit/SKILL.md +756 -0
- package/skills/research-pipeline/SKILL.md +384 -0
- package/skills/research-refine/SKILL.md +770 -0
- package/skills/research-refine-pipeline/SKILL.md +186 -0
- package/skills/research-review/SKILL.md +198 -0
- package/skills/research-wiki/SKILL.md +461 -0
- package/skills/resubmit-pipeline/SKILL.md +447 -0
- package/skills/result-to-claim/SKILL.md +311 -0
- package/skills/run-experiment/SKILL.md +313 -0
- package/skills/semantic-scholar/SKILL.md +236 -0
- package/skills/serverless-modal/SKILL.md +335 -0
- package/skills/shared-references/acceptance-gate.md +324 -0
- package/skills/shared-references/assurance-contract.md +248 -0
- package/skills/shared-references/capture-antipatterns.md +78 -0
- package/skills/shared-references/citation-discipline.md +583 -0
- package/skills/shared-references/compute-env-contract.md +163 -0
- package/skills/shared-references/effort-contract.md +183 -0
- package/skills/shared-references/evidence-precheck.md +65 -0
- package/skills/shared-references/experiment-integrity.md +49 -0
- package/skills/shared-references/external-cadence.md +326 -0
- package/skills/shared-references/fan-out-pattern.md +366 -0
- package/skills/shared-references/injection-hygiene.md +127 -0
- package/skills/shared-references/integration-contract.md +461 -0
- package/skills/shared-references/output-composition.md +93 -0
- package/skills/shared-references/output-language.md +45 -0
- package/skills/shared-references/output-manifest.md +49 -0
- package/skills/shared-references/output-versioning.md +111 -0
- package/skills/shared-references/patent-format-cn.md +199 -0
- package/skills/shared-references/patent-format-ep.md +173 -0
- package/skills/shared-references/patent-format-us.md +161 -0
- package/skills/shared-references/patent-writing-principles.md +197 -0
- package/skills/shared-references/prior-art-databases.md +141 -0
- package/skills/shared-references/resumable-runs.md +109 -0
- package/skills/shared-references/review-scope-limits.md +81 -0
- package/skills/shared-references/review-tracing.md +391 -0
- package/skills/shared-references/reviewer-independence.md +79 -0
- package/skills/shared-references/reviewer-routing.md +852 -0
- package/skills/shared-references/skill-governance.md +104 -0
- package/skills/shared-references/taste-calibration.md +85 -0
- package/skills/shared-references/venue-checklists.md +114 -0
- package/skills/shared-references/wiki-helper-resolution.md +134 -0
- package/skills/shared-references/writing-principles.md +525 -0
- package/skills/skills-codex/README.md +102 -0
- package/skills/skills-codex/README_CN.md +100 -0
- package/skills/skills-codex/ablation-planner/SKILL.md +126 -0
- package/skills/skills-codex/alphaxiv/SKILL.md +186 -0
- package/skills/skills-codex/analyze-results/SKILL.md +45 -0
- package/skills/skills-codex/arxiv/SKILL.md +210 -0
- package/skills/skills-codex/auto-paper-improvement-loop/SKILL.md +574 -0
- package/skills/skills-codex/auto-review-loop/SKILL.md +500 -0
- package/skills/skills-codex/auto-review-loop-llm/SKILL.md +247 -0
- package/skills/skills-codex/auto-review-loop-minimax/SKILL.md +290 -0
- package/skills/skills-codex/citation-audit/SKILL.md +504 -0
- package/skills/skills-codex/claims-drafting/SKILL.md +239 -0
- package/skills/skills-codex/comm-lit-review/SKILL.md +299 -0
- package/skills/skills-codex/comm-lit-review/references/domain-taxonomy.md +57 -0
- package/skills/skills-codex/comm-lit-review/references/output-template.md +37 -0
- package/skills/skills-codex/comm-lit-review/references/source-policy.md +99 -0
- package/skills/skills-codex/comm-lit-review/references/venue-tiering.md +112 -0
- package/skills/skills-codex/deepxiv/SKILL.md +142 -0
- package/skills/skills-codex/dse-loop/SKILL.md +285 -0
- package/skills/skills-codex/embodiment-description/SKILL.md +129 -0
- package/skills/skills-codex/exa-search/SKILL.md +192 -0
- package/skills/skills-codex/experiment-audit/SKILL.md +286 -0
- package/skills/skills-codex/experiment-bridge/SKILL.md +356 -0
- package/skills/skills-codex/experiment-plan/SKILL.md +249 -0
- package/skills/skills-codex/experiment-queue/SKILL.md +401 -0
- package/skills/skills-codex/feishu-notify/SKILL.md +155 -0
- package/skills/skills-codex/figure-description/SKILL.md +138 -0
- package/skills/skills-codex/figure-spec/SKILL.md +252 -0
- package/skills/skills-codex/formula-derivation/SKILL.md +280 -0
- package/skills/skills-codex/gemini-search/SKILL.md +205 -0
- package/skills/skills-codex/grant-proposal/SKILL.md +626 -0
- package/skills/skills-codex/idea-creator/SKILL.md +405 -0
- package/skills/skills-codex/idea-discovery/SKILL.md +475 -0
- package/skills/skills-codex/idea-discovery-robot/SKILL.md +362 -0
- package/skills/skills-codex/integrity-forensics/SKILL.md +106 -0
- package/skills/skills-codex/interview-cheatsheet/SKILL.md +245 -0
- package/skills/skills-codex/invention-structuring/SKILL.md +188 -0
- package/skills/skills-codex/jurisdiction-format/SKILL.md +192 -0
- package/skills/skills-codex/kill-argument/SKILL.md +403 -0
- package/skills/skills-codex/mermaid-diagram/SKILL.md +379 -0
- package/skills/skills-codex/meta-apply/SKILL.md +154 -0
- package/skills/skills-codex/meta-optimize/SKILL.md +348 -0
- package/skills/skills-codex/monitor-experiment/SKILL.md +98 -0
- package/skills/skills-codex/novelty-check/SKILL.md +89 -0
- package/skills/skills-codex/openalex/SKILL.md +228 -0
- package/skills/skills-codex/overleaf-sync/SKILL.md +220 -0
- package/skills/skills-codex/paper-claim-audit/SKILL.md +350 -0
- package/skills/skills-codex/paper-compile/SKILL.md +253 -0
- package/skills/skills-codex/paper-figure/SKILL.md +311 -0
- package/skills/skills-codex/paper-illustration/SKILL.md +690 -0
- package/skills/skills-codex/paper-illustration-image2/SKILL.md +383 -0
- package/skills/skills-codex/paper-illustration-image2/scripts/paper_illustration_image2.py +255 -0
- package/skills/skills-codex/paper-plan/SKILL.md +278 -0
- package/skills/skills-codex/paper-poster/SKILL.md +19 -0
- package/skills/skills-codex/paper-poster-html/SKILL.md +377 -0
- package/skills/skills-codex/paper-slides/SKILL.md +571 -0
- package/skills/skills-codex/paper-talk/SKILL.md +381 -0
- package/skills/skills-codex/paper-write/SKILL.md +411 -0
- package/skills/skills-codex/paper-write/templates/IEEEtran.bst +2409 -0
- package/skills/skills-codex/paper-write/templates/IEEEtran.cls +6347 -0
- package/skills/skills-codex/paper-write/templates/aaai2026.bst +1493 -0
- package/skills/skills-codex/paper-write/templates/aaai2026.sty +315 -0
- package/skills/skills-codex/paper-write/templates/aaai2026.tex +952 -0
- package/skills/skills-codex/paper-write/templates/acl.sty +312 -0
- package/skills/skills-codex/paper-write/templates/acl2026.tex +377 -0
- package/skills/skills-codex/paper-write/templates/acl_natbib.bst +1940 -0
- package/skills/skills-codex/paper-write/templates/acm.bst +3081 -0
- package/skills/skills-codex/paper-write/templates/acm_mm2026.tex +204 -0
- package/skills/skills-codex/paper-write/templates/acmart.cls +3520 -0
- package/skills/skills-codex/paper-write/templates/cvpr.bst +1448 -0
- package/skills/skills-codex/paper-write/templates/cvpr.sty +508 -0
- package/skills/skills-codex/paper-write/templates/cvpr2026.tex +63 -0
- package/skills/skills-codex/paper-write/templates/iclr2026.tex +84 -0
- package/skills/skills-codex/paper-write/templates/iclr2026_conference.bst +1440 -0
- package/skills/skills-codex/paper-write/templates/iclr2026_conference.sty +246 -0
- package/skills/skills-codex/paper-write/templates/icml2026.sty +767 -0
- package/skills/skills-codex/paper-write/templates/icml2026.tex +662 -0
- package/skills/skills-codex/paper-write/templates/ieee_conference.tex +89 -0
- package/skills/skills-codex/paper-write/templates/ieee_journal.tex +93 -0
- package/skills/skills-codex/paper-write/templates/math_commands.tex +48 -0
- package/skills/skills-codex/paper-write/templates/neurips2026.tex +493 -0
- package/skills/skills-codex/paper-write/templates/neurips_2026.sty +437 -0
- package/skills/skills-codex/paper-writing/SKILL.md +731 -0
- package/skills/skills-codex/patent-novelty-check/SKILL.md +153 -0
- package/skills/skills-codex/patent-pipeline/SKILL.md +344 -0
- package/skills/skills-codex/patent-review/SKILL.md +202 -0
- package/skills/skills-codex/pixel-art/SKILL.md +139 -0
- package/skills/skills-codex/prior-art-search/SKILL.md +146 -0
- package/skills/skills-codex/proof-checker/SKILL.md +554 -0
- package/skills/skills-codex/proof-orchestrator/SKILL.md +260 -0
- package/skills/skills-codex/proof-orchestrator/references/audit-output-contract.md +126 -0
- package/skills/skills-codex/proof-orchestrator/references/deepseek-routing.md +76 -0
- package/skills/skills-codex/proof-orchestrator/references/dispatch-prompts.md +227 -0
- package/skills/skills-codex/proof-orchestrator/references/notation-audit.md +135 -0
- package/skills/skills-codex/proof-orchestrator/references/proof-audit-rubric.md +70 -0
- package/skills/skills-codex/proof-orchestrator/references/stress-tests.md +38 -0
- package/skills/skills-codex/proof-writer/SKILL.md +222 -0
- package/skills/skills-codex/qzcli/SKILL.md +324 -0
- package/skills/skills-codex/rebuttal/SKILL.md +305 -0
- package/skills/skills-codex/render-html/SKILL.md +305 -0
- package/skills/skills-codex/render-html/scripts/__pycache__/render_html.cpython-314.pyc +0 -0
- package/skills/skills-codex/render-html/scripts/render_html.py +909 -0
- package/skills/skills-codex/render-html/scripts/templates/academic.html +342 -0
- package/skills/skills-codex/render-html/scripts/templates/dashboard.html +333 -0
- package/skills/skills-codex/research-lit/SKILL.md +464 -0
- package/skills/skills-codex/research-pipeline/SKILL.md +340 -0
- package/skills/skills-codex/research-refine/SKILL.md +721 -0
- package/skills/skills-codex/research-refine-pipeline/SKILL.md +186 -0
- package/skills/skills-codex/research-review/SKILL.md +135 -0
- package/skills/skills-codex/research-wiki/SKILL.md +421 -0
- package/skills/skills-codex/resubmit-pipeline/SKILL.md +444 -0
- package/skills/skills-codex/result-to-claim/SKILL.md +246 -0
- package/skills/skills-codex/run-experiment/SKILL.md +236 -0
- package/skills/skills-codex/semantic-scholar/SKILL.md +219 -0
- package/skills/skills-codex/serverless-modal/SKILL.md +335 -0
- package/skills/skills-codex/shared-references/acceptance-gate.md +336 -0
- package/skills/skills-codex/shared-references/assurance-contract.md +139 -0
- package/skills/skills-codex/shared-references/capture-antipatterns.md +84 -0
- package/skills/skills-codex/shared-references/citation-discipline.md +452 -0
- package/skills/skills-codex/shared-references/compute-env-contract.md +163 -0
- package/skills/skills-codex/shared-references/effort-contract.md +143 -0
- package/skills/skills-codex/shared-references/evidence-precheck.md +73 -0
- package/skills/skills-codex/shared-references/experiment-integrity.md +49 -0
- package/skills/skills-codex/shared-references/external-cadence.md +334 -0
- package/skills/skills-codex/shared-references/fan-out-pattern.md +375 -0
- package/skills/skills-codex/shared-references/injection-hygiene.md +132 -0
- package/skills/skills-codex/shared-references/integration-contract.md +372 -0
- package/skills/skills-codex/shared-references/output-composition.md +98 -0
- package/skills/skills-codex/shared-references/output-language.md +45 -0
- package/skills/skills-codex/shared-references/output-manifest.md +40 -0
- package/skills/skills-codex/shared-references/output-versioning.md +111 -0
- package/skills/skills-codex/shared-references/patent-format-cn.md +199 -0
- package/skills/skills-codex/shared-references/patent-format-ep.md +173 -0
- package/skills/skills-codex/shared-references/patent-format-us.md +161 -0
- package/skills/skills-codex/shared-references/patent-writing-principles.md +197 -0
- package/skills/skills-codex/shared-references/prior-art-databases.md +141 -0
- package/skills/skills-codex/shared-references/resumable-runs.md +125 -0
- package/skills/skills-codex/shared-references/review-scope-limits.md +81 -0
- package/skills/skills-codex/shared-references/review-tracing.md +144 -0
- package/skills/skills-codex/shared-references/reviewer-independence.md +66 -0
- package/skills/skills-codex/shared-references/reviewer-routing.md +128 -0
- package/skills/skills-codex/shared-references/skill-governance.md +119 -0
- package/skills/skills-codex/shared-references/taste-calibration.md +90 -0
- package/skills/skills-codex/shared-references/venue-checklists.md +73 -0
- package/skills/skills-codex/shared-references/wiki-helper-resolution.md +69 -0
- package/skills/skills-codex/shared-references/writing-principles.md +525 -0
- package/skills/skills-codex/slides-polish/SKILL.md +563 -0
- package/skills/skills-codex/specification-writing/SKILL.md +211 -0
- package/skills/skills-codex/system-profile/SKILL.md +103 -0
- package/skills/skills-codex/training-check/SKILL.md +83 -0
- package/skills/skills-codex/vast-gpu/SKILL.md +394 -0
- package/skills/skills-codex/web-debug-search/SKILL.md +334 -0
- package/skills/skills-codex/wiki-enrich/SKILL.md +255 -0
- package/skills/skills-codex/writing-systems-papers/SKILL.md +184 -0
- package/skills/skills-codex-claude-review/README.md +79 -0
- package/skills/skills-codex-claude-review/README_CN.md +78 -0
- package/skills/skills-codex-claude-review/auto-paper-improvement-loop/SKILL.md +581 -0
- package/skills/skills-codex-claude-review/auto-review-loop/SKILL.md +510 -0
- package/skills/skills-codex-claude-review/novelty-check/SKILL.md +102 -0
- package/skills/skills-codex-claude-review/paper-figure/SKILL.md +319 -0
- package/skills/skills-codex-claude-review/paper-plan/SKILL.md +287 -0
- package/skills/skills-codex-claude-review/paper-write/SKILL.md +420 -0
- package/skills/skills-codex-claude-review/research-refine/SKILL.md +732 -0
- package/skills/skills-codex-claude-review/research-review/SKILL.md +149 -0
- package/skills/skills-codex-gemini-review/README.md +176 -0
- package/skills/skills-codex-gemini-review/README_CN.md +175 -0
- package/skills/skills-codex-gemini-review/auto-paper-improvement-loop/SKILL.md +331 -0
- package/skills/skills-codex-gemini-review/auto-review-loop/SKILL.md +304 -0
- package/skills/skills-codex-gemini-review/grant-proposal/SKILL.md +630 -0
- package/skills/skills-codex-gemini-review/idea-creator/SKILL.md +263 -0
- package/skills/skills-codex-gemini-review/idea-discovery/SKILL.md +275 -0
- package/skills/skills-codex-gemini-review/idea-discovery-robot/SKILL.md +365 -0
- package/skills/skills-codex-gemini-review/novelty-check/SKILL.md +92 -0
- package/skills/skills-codex-gemini-review/paper-figure/SKILL.md +289 -0
- package/skills/skills-codex-gemini-review/paper-plan/SKILL.md +265 -0
- package/skills/skills-codex-gemini-review/paper-poster-html/SKILL.md +102 -0
- package/skills/skills-codex-gemini-review/paper-slides/SKILL.md +582 -0
- package/skills/skills-codex-gemini-review/paper-write/SKILL.md +346 -0
- package/skills/skills-codex-gemini-review/paper-writing/SKILL.md +312 -0
- package/skills/skills-codex-gemini-review/research-refine/SKILL.md +674 -0
- package/skills/skills-codex-gemini-review/research-review/SKILL.md +112 -0
- package/skills/slides-polish/SKILL.md +565 -0
- package/skills/specification-writing/SKILL.md +211 -0
- package/skills/system-profile/SKILL.md +103 -0
- package/skills/training-check/SKILL.md +132 -0
- package/skills/vast-gpu/SKILL.md +394 -0
- package/skills/web-debug-search/SKILL.md +334 -0
- package/skills/wiki-enrich/SKILL.md +257 -0
- package/skills/writing-systems-papers/SKILL.md +184 -0
- package/templates/CLAUDE_MD_TEMPLATE.md +29 -0
- package/templates/EXPERIMENT_LOG_TEMPLATE.md +47 -0
- package/templates/EXPERIMENT_PLAN_TEMPLATE.md +51 -0
- package/templates/EXPERIMENT_PLAN_TEMPLATE_CN.md +53 -0
- package/templates/FINDINGS_TEMPLATE.md +52 -0
- package/templates/IDEA_CANDIDATES_TEMPLATE.md +47 -0
- package/templates/IDEA_CANDIDATES_TEMPLATE_CN.md +47 -0
- package/templates/INVENTION_BRIEF_TEMPLATE.md +87 -0
- package/templates/MANIFEST_TEMPLATE.md +7 -0
- package/templates/NARRATIVE_REPORT_TEMPLATE.md +49 -0
- package/templates/PAPER_PLAN_TEMPLATE.md +47 -0
- package/templates/PATENT_CLAIMS_TEMPLATE.md +78 -0
- package/templates/PATENT_SPECIFICATION_TEMPLATE.md +68 -0
- package/templates/README.md +57 -0
- package/templates/RESEARCH_BRIEF_TEMPLATE.md +35 -0
- package/templates/RESEARCH_BRIEF_TEMPLATE_CN.md +41 -0
- package/templates/RESEARCH_CONTRACT_TEMPLATE.md +60 -0
- package/templates/claude-hooks/corpus_write_guard.json +16 -0
- package/templates/claude-hooks/corpus_write_guard.py +85 -0
- package/templates/claude-hooks/meta_logging.json +74 -0
- package/templates/gitignore-trace.txt +3 -0
- package/tools/__pycache__/check_skills_inventory.cpython-314.pyc +0 -0
- package/tools/arxiv_fetch.py +311 -0
- package/tools/capture_filter.py +126 -0
- package/tools/check_skills_inventory.py +273 -0
- package/tools/convert_skills_to_llm_chat.py +282 -0
- package/tools/copilot_native_evidence.py +818 -0
- package/tools/deepxiv_fetch.py +213 -0
- package/tools/evidence_check.py +212 -0
- package/tools/exa_search.py +425 -0
- package/tools/experiment_queue/README.md +118 -0
- package/tools/experiment_queue/build_manifest.py +44 -0
- package/tools/experiment_queue/queue_manager.py +44 -0
- package/tools/extract_paper_style.py +560 -0
- package/tools/figure_renderer.py +69 -0
- package/tools/forensics_gate.py +669 -0
- package/tools/generate_codex_claude_review_overrides.py +299 -0
- package/tools/idea_discovery_gate.py +256 -0
- package/tools/install_aris.ps1 +1372 -0
- package/tools/install_aris.sh +1370 -0
- package/tools/install_aris_codex.sh +1023 -0
- package/tools/install_aris_copilot.sh +1052 -0
- package/tools/iteration_log.py +143 -0
- package/tools/lint_skills_helpers.sh +84 -0
- package/tools/meta_opt/check_ready.sh +80 -0
- package/tools/meta_opt/log_event.sh +91 -0
- package/tools/meta_opt/trigger_eval.py +280 -0
- package/tools/meta_opt/trigger_evals.sample.json +28 -0
- package/tools/openalex_fetch.py +326 -0
- package/tools/overleaf_audit.sh +104 -0
- package/tools/overleaf_setup.sh +150 -0
- package/tools/paper_illustration_image2.py +62 -0
- package/tools/provenance.py +294 -0
- package/tools/research_wiki.py +1720 -0
- package/tools/review_gate.py +502 -0
- package/tools/run_state.py +399 -0
- package/tools/save_trace.sh +477 -0
- package/tools/semantic_scholar_fetch.py +438 -0
- package/tools/skill-groups.tsv +116 -0
- package/tools/skill_picker.py +238 -0
- package/tools/smart_update.ps1 +521 -0
- package/tools/smart_update.sh +591 -0
- package/tools/smart_update_codex.sh +419 -0
- package/tools/smart_update_copilot.sh +605 -0
- package/tools/threat_scan.py +222 -0
- package/tools/verify_paper_audits.sh +487 -0
- package/tools/verify_papers.py +613 -0
- package/tools/verify_wiki_coverage.sh +176 -0
- package/tools/watchdog.py +485 -0
|
@@ -0,0 +1,144 @@
|
|
|
1
|
+
# Review Tracing Protocol
|
|
2
|
+
|
|
3
|
+
## Purpose
|
|
4
|
+
|
|
5
|
+
Save full prompt/response pairs for every reviewer call, enabling:
|
|
6
|
+
|
|
7
|
+
- reviewer-independence audit
|
|
8
|
+
- reproducibility across follow-up reviewer turns
|
|
9
|
+
- meta-optimization from real review traces
|
|
10
|
+
|
|
11
|
+
## When to Trace
|
|
12
|
+
|
|
13
|
+
Trace every Codex reviewer call that serves a critique, scoring, claim-verification, experiment-audit, or patch-gating function.
|
|
14
|
+
|
|
15
|
+
This includes:
|
|
16
|
+
|
|
17
|
+
- `spawn_agent` reviewer calls
|
|
18
|
+
- `send_input` reviewer continuations
|
|
19
|
+
- optional overlay reviewer routes
|
|
20
|
+
- adversarial reviewer calls used for stress tests
|
|
21
|
+
|
|
22
|
+
Do not trace purely informational agent calls that are not acting as reviewers.
|
|
23
|
+
|
|
24
|
+
## How to Trace
|
|
25
|
+
|
|
26
|
+
After each reviewer call, save the trace using `save_trace.sh`,
|
|
27
|
+
resolved through the canonical helper chain (see
|
|
28
|
+
`integration-contract.md` §2 — failure policy C, "forensic helper").
|
|
29
|
+
A Codex-side SKILL must NOT hard-code `tools/save_trace.sh`; instead
|
|
30
|
+
it resolves `$TRACE_HELPER` via the chain and either invokes the
|
|
31
|
+
helper or writes trace artifacts directly per the schemas below. If
|
|
32
|
+
the resolver returns the empty string, write the four files inline
|
|
33
|
+
— do not silently skip unless `--- trace: off` was requested.
|
|
34
|
+
|
|
35
|
+
## Trace Directory
|
|
36
|
+
|
|
37
|
+
```text
|
|
38
|
+
.aris/traces/<skill-name>/<YYYY-MM-DD>_run<NN>/
|
|
39
|
+
run.meta.json
|
|
40
|
+
001-<purpose>.request.json
|
|
41
|
+
001-<purpose>.response.md
|
|
42
|
+
001-<purpose>.meta.json
|
|
43
|
+
002-<purpose>.request.json
|
|
44
|
+
...
|
|
45
|
+
```
|
|
46
|
+
|
|
47
|
+
## File Schemas
|
|
48
|
+
|
|
49
|
+
`run.meta.json`:
|
|
50
|
+
|
|
51
|
+
```json
|
|
52
|
+
{
|
|
53
|
+
"skill": "auto-review-loop",
|
|
54
|
+
"run_id": "2026-04-15_run01",
|
|
55
|
+
"started_at": "2026-04-15T14:30:00+08:00",
|
|
56
|
+
"executor": "codex",
|
|
57
|
+
"executor_model": "gpt-5.6-sol",
|
|
58
|
+
"executor_family": "openai",
|
|
59
|
+
"review_independence": "same-family",
|
|
60
|
+
"acceptance_status": "provisional",
|
|
61
|
+
"project_dir": "/path/to/project"
|
|
62
|
+
}
|
|
63
|
+
```
|
|
64
|
+
|
|
65
|
+
`NNN-<purpose>.request.json`:
|
|
66
|
+
|
|
67
|
+
```json
|
|
68
|
+
{
|
|
69
|
+
"call_number": 1,
|
|
70
|
+
"purpose": "round-1-review",
|
|
71
|
+
"timestamp": "2026-04-15T14:31:00+08:00",
|
|
72
|
+
"tool": "spawn_agent",
|
|
73
|
+
"model": "gpt-5.6-sol",
|
|
74
|
+
"reasoning_effort": "xhigh",
|
|
75
|
+
"files_referenced": ["paper/sections/3_method.tex", "results/table1.csv"],
|
|
76
|
+
"prompt": "<full prompt text>"
|
|
77
|
+
}
|
|
78
|
+
```
|
|
79
|
+
|
|
80
|
+
`NNN-<purpose>.response.md` stores the full reviewer response verbatim.
|
|
81
|
+
|
|
82
|
+
`NNN-<purpose>.meta.json`:
|
|
83
|
+
|
|
84
|
+
```json
|
|
85
|
+
{
|
|
86
|
+
"call_number": 1,
|
|
87
|
+
"purpose": "round-1-review",
|
|
88
|
+
"timestamp": "2026-04-15T14:33:00+08:00",
|
|
89
|
+
"agent_id": "019d8fe0-b25d-...",
|
|
90
|
+
"model": "gpt-5.6-sol",
|
|
91
|
+
"reviewer_family": "openai",
|
|
92
|
+
"review_independence": "same-family",
|
|
93
|
+
"acceptance_status": "provisional",
|
|
94
|
+
"duration_ms": 142000,
|
|
95
|
+
"status": "ok"
|
|
96
|
+
}
|
|
97
|
+
```
|
|
98
|
+
|
|
99
|
+
For a Claude/Gemini overlay write `review_independence: cross-family` and
|
|
100
|
+
`acceptance_status: accepted`. For a deterministic verifier write
|
|
101
|
+
`review_independence: deterministic` and `acceptance_status: accepted`. These
|
|
102
|
+
fields describe the reviewer route; they do not rewrite the substantive verdict.
|
|
103
|
+
|
|
104
|
+
## Configuration
|
|
105
|
+
|
|
106
|
+
Respect inline parameter `--- trace: off | meta | full`:
|
|
107
|
+
|
|
108
|
+
- `full` default: save full prompt and full response
|
|
109
|
+
- `meta`: save metadata only
|
|
110
|
+
- `off`: disable tracing
|
|
111
|
+
|
|
112
|
+
## Events
|
|
113
|
+
|
|
114
|
+
After writing a trace, append a compact event to `.aris/meta/events.jsonl`:
|
|
115
|
+
|
|
116
|
+
```json
|
|
117
|
+
{"event":"review_trace","skill":"auto-review-loop","purpose":"round-1-review","agent_id":"...","trace_path":".aris/traces/auto-review-loop/2026-04-15_run01/","status":"ok"}
|
|
118
|
+
```
|
|
119
|
+
|
|
120
|
+
## Debugging With Traces
|
|
121
|
+
|
|
122
|
+
Traces are not only audit evidence — they are the **first place to look when a
|
|
123
|
+
verdict is surprising**: a score regresses round-to-round, two reviewer agents
|
|
124
|
+
disagree, or a claim check contradicts an earlier round. Before re-invoking the
|
|
125
|
+
reviewer for "a better answer", read the raw transcript and find the moment its
|
|
126
|
+
judgment actually changed:
|
|
127
|
+
|
|
128
|
+
```bash
|
|
129
|
+
skill=auto-review-loop run=2026-04-15_run01
|
|
130
|
+
diff ".aris/traces/$skill/$run/002-round-2.response.md" \
|
|
131
|
+
".aris/traces/$skill/$run/003-round-3.response.md"
|
|
132
|
+
grep -En 'however|but|concern|missing|cannot' \
|
|
133
|
+
".aris/traces/$skill/$run/003-round-3.response.md"
|
|
134
|
+
```
|
|
135
|
+
|
|
136
|
+
The paragraph where the assessment changed **is** the causal explanation for the
|
|
137
|
+
divergence — cite it, don't guess. Re-running the reviewer without reading the
|
|
138
|
+
trace is tuning by vibe: you get a new opinion, not an explanation. Same muscle
|
|
139
|
+
as reading a stack trace before retrying a failed run — the trace is just
|
|
140
|
+
written in English, and most of it is the reviewer talking to itself.
|
|
141
|
+
|
|
142
|
+
## Privacy
|
|
143
|
+
|
|
144
|
+
`.aris/traces/` is project-local and should not be committed. Use `--- trace: off` for strict confidentiality projects.
|
|
@@ -0,0 +1,66 @@
|
|
|
1
|
+
# Reviewer Independence Protocol
|
|
2
|
+
|
|
3
|
+
## Core Principle
|
|
4
|
+
|
|
5
|
+
The reviewer must judge primary artifacts directly. The executor can define the review role, scope, and file list, but must not pre-digest the content into a preferred narrative.
|
|
6
|
+
|
|
7
|
+
## What Can Be Passed
|
|
8
|
+
|
|
9
|
+
- reviewer role or persona
|
|
10
|
+
- review objective
|
|
11
|
+
- absolute file paths
|
|
12
|
+
- structural metadata such as section count or venue
|
|
13
|
+
- concrete output schema
|
|
14
|
+
|
|
15
|
+
## What Must Not Be Passed
|
|
16
|
+
|
|
17
|
+
- executor summaries of file contents
|
|
18
|
+
- executor interpretations of results
|
|
19
|
+
- executor recommendations about what the reviewer should conclude
|
|
20
|
+
- "what changed since last round" narratives unless the skill explicitly requires diff-focused follow-up
|
|
21
|
+
- leading or coaching questions
|
|
22
|
+
|
|
23
|
+
## Correct Pattern
|
|
24
|
+
|
|
25
|
+
```text
|
|
26
|
+
spawn_agent:
|
|
27
|
+
model: gpt-5.6-sol
|
|
28
|
+
reasoning_effort: xhigh
|
|
29
|
+
message: |
|
|
30
|
+
Review the project as a senior ML reviewer.
|
|
31
|
+
|
|
32
|
+
Files to read directly:
|
|
33
|
+
- /path/to/PROPOSAL.md
|
|
34
|
+
- /path/to/EXPERIMENT_LOG.md
|
|
35
|
+
- /path/to/paper/main.tex
|
|
36
|
+
- /path/to/src/
|
|
37
|
+
```
|
|
38
|
+
|
|
39
|
+
## Incorrect Pattern
|
|
40
|
+
|
|
41
|
+
```text
|
|
42
|
+
spawn_agent:
|
|
43
|
+
model: gpt-5.6-sol
|
|
44
|
+
reasoning_effort: xhigh
|
|
45
|
+
message: |
|
|
46
|
+
The main contribution is a new loss function that improves by 15%.
|
|
47
|
+
I think the weak point is the ablation.
|
|
48
|
+
Please confirm this is publishable.
|
|
49
|
+
```
|
|
50
|
+
|
|
51
|
+
## Multi-Round Follow-Up
|
|
52
|
+
|
|
53
|
+
When a skill uses multi-round review, reuse the same reviewer id with `send_input`, but still avoid injecting executor conclusions. Pass revised artifacts or targeted follow-up requests, not spin.
|
|
54
|
+
|
|
55
|
+
## Applies To
|
|
56
|
+
|
|
57
|
+
This protocol applies to all cross-agent review calls in `skills/skills-codex/`, including:
|
|
58
|
+
|
|
59
|
+
- `research-review`
|
|
60
|
+
- `auto-review-loop`
|
|
61
|
+
- `paper-plan`
|
|
62
|
+
- `paper-write`
|
|
63
|
+
- `paper-figure`
|
|
64
|
+
- `rebuttal`
|
|
65
|
+
- `meta-optimize`
|
|
66
|
+
- any skill that launches a reviewer via `spawn_agent` or continues one via `send_input`
|
|
@@ -0,0 +1,128 @@
|
|
|
1
|
+
# Reviewer Routing
|
|
2
|
+
|
|
3
|
+
## Default Reviewer Contract
|
|
4
|
+
|
|
5
|
+
All reviewer-heavy Codex base skills use the same default contract:
|
|
6
|
+
|
|
7
|
+
- executor: current Codex main agent
|
|
8
|
+
- reviewer: second Codex reviewer, model `gpt-5.6-sol` (GPT-5.6-Sol)
|
|
9
|
+
- reasoning effort: **two tiers** (since 2026-07-10; `ultra`/`max` need codex-cli ≥ 0.144.1) —
|
|
10
|
+
**deep-audit** skills use `ultra` (`proof-checker`, `kill-argument` core threads, `research-review`,
|
|
11
|
+
`experiment-audit`, `paper-claim-audit`, `result-to-claim`, `meta-apply`); **every other**
|
|
12
|
+
reviewer call uses `xhigh` (multi-round loops and per-item fan-outs stay `xhigh` — a
|
|
13
|
+
follow-up `send_input` cannot change model/effort, and per-item `ultra` multiplies cost)
|
|
14
|
+
- round 1: `spawn_agent`
|
|
15
|
+
- follow-up rounds: `send_input`
|
|
16
|
+
|
|
17
|
+
This is the base default for `skills/skills-codex/`. No ARIS `— effort:` level or unrelated parameter changes the tier (ARIS `— effort: max` ≠ `reasoning_effort: max` — pipeline workload vs reviewer reasoning are different axes).
|
|
18
|
+
|
|
19
|
+
**Capability fallback (first spawn of each tier only):** if `spawn_agent` errors explicitly on the effort enum (older codex-cli — applies only to the deep tier's `ultra`; `xhigh` predates 0.144.1), retry `reasoning_effort: xhigh`; if it errors explicitly on the model being unknown/unavailable to this account, retry `model: gpt-5.5` + `xhigh`. NEVER downgrade on timeout / rate-limit / auth / transport / server / context-length errors (risk of double-running). Never run a verdict-bearing review below `xhigh`; if no allowed pair works, report `REVIEW_UNAVAILABLE` — never substitute the executor's own judgment.
|
|
20
|
+
|
|
21
|
+
> ⚠️ **Same-family by default — provisional, never accepted.** The executor here
|
|
22
|
+
> is Codex (GPT family) and the reviewer is a fresh Codex agent from the same
|
|
23
|
+
> family. Its substantive PASS/WARN/FAIL may drive revisions, terminate a loop,
|
|
24
|
+
> and advance a resumable phase, but every positive result records:
|
|
25
|
+
>
|
|
26
|
+
> ```yaml
|
|
27
|
+
> review_independence: same-family
|
|
28
|
+
> acceptance_status: provisional
|
|
29
|
+
> ```
|
|
30
|
+
>
|
|
31
|
+
> It must never be described as cross-model acceptance. Install the
|
|
32
|
+
> **`skills-codex-claude-review`** or **`skills-codex-gemini-review`** overlay
|
|
33
|
+
> for `review_independence: cross-family` and `acceptance_status: accepted`.
|
|
34
|
+
> A deterministic verifier may also record accepted. `oracle-pro` is GPT family,
|
|
35
|
+
> so it remains provisional for a Codex executor.
|
|
36
|
+
|
|
37
|
+
## Default Pattern
|
|
38
|
+
|
|
39
|
+
Single-round review:
|
|
40
|
+
|
|
41
|
+
```text
|
|
42
|
+
spawn_agent:
|
|
43
|
+
model: gpt-5.6-sol
|
|
44
|
+
reasoning_effort: xhigh # deep-audit skills: ultra (see tier table above)
|
|
45
|
+
message: |
|
|
46
|
+
[role + task]
|
|
47
|
+
Read the listed files directly.
|
|
48
|
+
```
|
|
49
|
+
|
|
50
|
+
Multi-round review:
|
|
51
|
+
|
|
52
|
+
```text
|
|
53
|
+
spawn_agent:
|
|
54
|
+
model: gpt-5.6-sol
|
|
55
|
+
reasoning_effort: xhigh # deep-audit skills: ultra (see tier table above)
|
|
56
|
+
message: |
|
|
57
|
+
[initial review prompt]
|
|
58
|
+
```
|
|
59
|
+
|
|
60
|
+
Save the returned reviewer id, then continue with:
|
|
61
|
+
|
|
62
|
+
```text
|
|
63
|
+
send_input:
|
|
64
|
+
target: <saved reviewer id>
|
|
65
|
+
message: |
|
|
66
|
+
[follow-up materials only]
|
|
67
|
+
```
|
|
68
|
+
|
|
69
|
+
## Oracle Pro Override
|
|
70
|
+
|
|
71
|
+
When the user explicitly passes `--reviewer: oracle-pro`, switch only the reviewer route:
|
|
72
|
+
|
|
73
|
+
- default reviewer remains Codex at the call's declared tier (deep-audit: ultra / regular: xhigh) if no reviewer is specified
|
|
74
|
+
- `oracle-pro` is optional, not the base default
|
|
75
|
+
|
|
76
|
+
Routing rule:
|
|
77
|
+
|
|
78
|
+
```text
|
|
79
|
+
If reviewer is omitted or reviewer=codex:
|
|
80
|
+
use spawn_agent / send_input with the Codex reviewer at the call's declared tier
|
|
81
|
+
|
|
82
|
+
If reviewer=oracle-pro:
|
|
83
|
+
check Oracle MCP availability
|
|
84
|
+
if available:
|
|
85
|
+
call mcp__oracle__consult with model gpt-5.5-pro
|
|
86
|
+
if unavailable:
|
|
87
|
+
print a clear warning
|
|
88
|
+
fall back to the default Codex reviewer at the call's declared tier
|
|
89
|
+
```
|
|
90
|
+
|
|
91
|
+
## Invariants
|
|
92
|
+
|
|
93
|
+
- Base skills do not use the legacy Codex MCP thread path as the default reviewer route.
|
|
94
|
+
- Reviewer independence still applies: pass file paths and task framing, not executor summaries.
|
|
95
|
+
- Overlay packages may replace only the reviewer route.
|
|
96
|
+
- Overlay packages do not change executor semantics.
|
|
97
|
+
- Every trace and audit artifact records `review_independence` and
|
|
98
|
+
`acceptance_status`; missing metadata is treated as provisional.
|
|
99
|
+
- If `spawn_agent` is unavailable or fails, emit `BLOCKED` /
|
|
100
|
+
`REVIEW_UNAVAILABLE`; never fabricate a provisional PASS.
|
|
101
|
+
- Do not wrap verdict-bearing skills in `/loop`, cron, or wall-clock retries.
|
|
102
|
+
Schedule only external-world waits, then invoke the reviewer once after the
|
|
103
|
+
artifact changes. See `external-cadence.md`.
|
|
104
|
+
- Browser-based Oracle review is acceptable for one-shot stress tests, not ideal for tight multi-round loops.
|
|
105
|
+
|
|
106
|
+
## Copilot CLI reviewer behavior in the main skill set
|
|
107
|
+
|
|
108
|
+
The main `skills/shared-references/reviewer-routing.md` defaults
|
|
109
|
+
`/auto-review-loop` to Copilot CLI's native complementary `rubber-duck`
|
|
110
|
+
subagent when a host-session marker binds. Its stop gate requires revalidated
|
|
111
|
+
native lifecycle/model/response evidence and a known cross-family pair. The
|
|
112
|
+
older `--reviewer: copilot` custom-agent subprocess remains an explicit
|
|
113
|
+
compatibility drive mode and still needs a Codex/manual finalizer.
|
|
114
|
+
|
|
115
|
+
**This routing applies to the main skills at `skills/`, not this Codex-mirror
|
|
116
|
+
pack.** Here `spawn_agent` remains the native reviewer. See the main
|
|
117
|
+
[`reviewer-routing.md`](../../../skills/shared-references/reviewer-routing.md#copilot-cli-native-rubber-duck-default-for-auto-review-loop)
|
|
118
|
+
for the Copilot contract.
|
|
119
|
+
|
|
120
|
+
## Skills That Commonly Benefit From `oracle-pro`
|
|
121
|
+
|
|
122
|
+
- `research-review`
|
|
123
|
+
- `auto-review-loop`
|
|
124
|
+
- `experiment-audit`
|
|
125
|
+
- `proof-checker`
|
|
126
|
+
- `rebuttal`
|
|
127
|
+
- `idea-creator`
|
|
128
|
+
- `research-lit`
|
|
@@ -0,0 +1,119 @@
|
|
|
1
|
+
# Skill Governance: Provenance-as-Authorization
|
|
2
|
+
|
|
3
|
+
> **Codex mirror adaptation (normative).** A Codex-authored change reviewed by a
|
|
4
|
+
> fresh Codex agent uses `provenance.py stamp-provisional`. The receipt preserves
|
|
5
|
+
> author/reviewer family, verdict id and content hash but does not make the
|
|
6
|
+
> artifact auto-curatable. Cross-family overlays or deterministic verifiers use
|
|
7
|
+
> strict `stamp` and may grant accepted authorization.
|
|
8
|
+
|
|
9
|
+
When ARIS **auto-writes** a durable artifact — a meta-optimize patch to a SKILL.md,
|
|
10
|
+
an auto-curated research-wiki node, a machine-proposed reviewer prompt — two
|
|
11
|
+
questions must have a recorded, checkable answer *before* that artifact is allowed
|
|
12
|
+
to influence future runs:
|
|
13
|
+
|
|
14
|
+
1. **Who authored it, and who acquitted it?** (and were they different model families?)
|
|
15
|
+
2. **Is auto-curation even allowed to touch this file?** (or is it a hand-written
|
|
16
|
+
canonical skill / a user's own note, which automation must never rewrite?)
|
|
17
|
+
|
|
18
|
+
`tools/provenance.py` is the primitive that answers both. This is the
|
|
19
|
+
**provenance-as-authorization boundary**: a provenance record is not just metadata,
|
|
20
|
+
it *is* the authorization to auto-curate.
|
|
21
|
+
|
|
22
|
+
## The rule
|
|
23
|
+
|
|
24
|
+
- **Auto-curation may ONLY touch artifacts where `is_auto_curatable(path)` is True.**
|
|
25
|
+
An auto-authored artifact carries a `.provenance.json` sidecar with
|
|
26
|
+
`created_by == "aris-auto"`. Canonical hand-written skills and user notes have no
|
|
27
|
+
such record — `is_auto_curatable` returns False — and are **off-limits** to any
|
|
28
|
+
automated rewrite/delete. (A loop can author new machine artifacts; it must not
|
|
29
|
+
silently edit human ones.)
|
|
30
|
+
|
|
31
|
+
- **A provenance record cannot be same-family.** `stamp()` calls
|
|
32
|
+
`assert_cross_family(author, reviewer)` and **refuses (raises)** if the author and
|
|
33
|
+
the reviewer are the same model family — e.g. `claude`+`sonnet`, or `gpt-5.5`+`codex`,
|
|
34
|
+
or the trap `gpt-5.5`+`oracle-pro` (oracle routes to a GPT-Pro tier, so it is the
|
|
35
|
+
**openai** family, not a separate one). You therefore *cannot produce* a valid
|
|
36
|
+
authorization for a self-acquitted artifact. This is the structural form of the
|
|
37
|
+
cross-model invariant from [`reviewer-independence.md`](reviewer-independence.md)
|
|
38
|
+
and the acceptance boundary from [`acceptance-gate.md`](acceptance-gate.md):
|
|
39
|
+
**a loop can DRIVE, it cannot ACQUIT itself.**
|
|
40
|
+
|
|
41
|
+
- **A provisional receipt may be same-family but grants no authorization.**
|
|
42
|
+
`stamp_provisional()` requires the author/reviewer families to match and writes
|
|
43
|
+
`review_independence: same-family`, `acceptance_status: provisional`.
|
|
44
|
+
`is_auto_authored()` remains true as an identity fact, while
|
|
45
|
+
`is_auto_curatable()` is false. This is the Codex-only review path.
|
|
46
|
+
|
|
47
|
+
- **A deterministic verifier is a valid reviewer.** A reviewer named
|
|
48
|
+
`deterministic:<verifier>` (e.g. `deterministic:evidence_check`,
|
|
49
|
+
`deterministic:pytest`) passes the gate regardless of the author — a process is
|
|
50
|
+
not a model family. This is the same Type-A escape hatch as in
|
|
51
|
+
[`acceptance-gate.md`](acceptance-gate.md): an execution-completeness / mechanical
|
|
52
|
+
check is safe same-model. Use it when the acquittal is a passing test or a
|
|
53
|
+
deterministic pre-check (see [`evidence-precheck.md`](evidence-precheck.md)), not a
|
|
54
|
+
semantic judgement.
|
|
55
|
+
|
|
56
|
+
- **Unknown family fails closed.** If either name maps to no known family,
|
|
57
|
+
`assert_cross_family` raises rather than guessing — you must use a recognized
|
|
58
|
+
reviewer or a deterministic verifier.
|
|
59
|
+
|
|
60
|
+
## The record
|
|
61
|
+
|
|
62
|
+
Strict `stamp(target, author_model, reviewer_model, verdict_id)` writes accepted
|
|
63
|
+
authorization; `stamp_provisional(...)` writes the same receipt shape with
|
|
64
|
+
provisional review metadata.
|
|
65
|
+
|
|
66
|
+
```json
|
|
67
|
+
{
|
|
68
|
+
"created_by": "aris-auto",
|
|
69
|
+
"author_model": "claude-opus-4-8",
|
|
70
|
+
"author_family": "anthropic",
|
|
71
|
+
"reviewer_model": "gpt-5.5",
|
|
72
|
+
"reviewer_family": "openai",
|
|
73
|
+
"verdict_id": "codex_thread_abc123",
|
|
74
|
+
"content_hash": "<sha256 of the artifact>",
|
|
75
|
+
"stamped_at": "2026-05-30T00:00:00Z"
|
|
76
|
+
}
|
|
77
|
+
```
|
|
78
|
+
|
|
79
|
+
- `verdict_id` is the **traceable** acquittal — the codex thread id, the oracle
|
|
80
|
+
session id, or for a deterministic reviewer the verifier report path / sha. It is
|
|
81
|
+
required (empty → refused), so every authorization points back to an auditable
|
|
82
|
+
review.
|
|
83
|
+
- `content_hash` is tamper-evidence: if the artifact is later edited by hand, the
|
|
84
|
+
hash no longer matches the record, and an accepted re-stamp (= a fresh
|
|
85
|
+
cross-family or deterministic review) is required before auto-curation may
|
|
86
|
+
treat it as machine-owned again.
|
|
87
|
+
|
|
88
|
+
## How skills use it
|
|
89
|
+
|
|
90
|
+
- **meta-optimize** — when an auto-patch lands (Step 6), `stamp()` the changed
|
|
91
|
+
SKILL.md with `author_model` = the executor that drafted the patch and
|
|
92
|
+
`reviewer_model` = the codex/oracle reviewer that scored it (Step 4). The stamp
|
|
93
|
+
*refusing* is the last line of defense: if the patch was never cross-model
|
|
94
|
+
reviewed, there is no valid `verdict_id`/family pair to stamp, so it cannot be
|
|
95
|
+
recorded as authorized.
|
|
96
|
+
- **research-wiki** — auto-curated nodes (machine-merged, machine-pruned) get a
|
|
97
|
+
provenance stamp; user-written and import-from-paper nodes do not, so a future
|
|
98
|
+
auto-curator can tell which nodes it is allowed to rewrite.
|
|
99
|
+
- **acceptance-gate** — the cross-family assertion here is the *enforcement* of the
|
|
100
|
+
Type-B (quality/correctness) rule: a quality acquittal of an auto-authored
|
|
101
|
+
artifact must be cross-model, and the provenance record is the proof that it was.
|
|
102
|
+
|
|
103
|
+
## Why (the Hermes contrast)
|
|
104
|
+
|
|
105
|
+
Hermes-agent (NousResearch, MIT) has the *shape* — a `skill_provenance` marker and
|
|
106
|
+
a `created_by` tag — but its cross-model curator is **optional config that defaults
|
|
107
|
+
to the same chat model**, so by default one model writes, judges, and consumes its
|
|
108
|
+
own skills. That is exactly the self-poisoning failure
|
|
109
|
+
[`capture-antipatterns.md`](capture-antipatterns.md) guards against, one level up:
|
|
110
|
+
there it is a *captured claim* that hardens into a self-cited falsehood; here it is
|
|
111
|
+
a *captured skill* that an unreviewed loop grants itself authority to keep. ARIS's
|
|
112
|
+
increment is to make cross-family a **structural assertion on the stamp's content**:
|
|
113
|
+
`assert_cross_family` genuinely raises, so you cannot produce a *valid* same-family stamp
|
|
114
|
+
(not a config flag you can flip). The scope is precise — this governs *what a stamp can
|
|
115
|
+
say*, not *whether a write calls `stamp()` at all*. Whether a corpus mutation is stamped
|
|
116
|
+
is governed by `/meta-apply`'s landing procedure; an edit that skips the stamp leaves no
|
|
117
|
+
provenance — that case is **detectable, not impossible** (a missing / stale-hash stamp is
|
|
118
|
+
what a pre-push integrity check would catch — that verifier is a follow-up, see
|
|
119
|
+
meta-optimize's note).
|
|
@@ -0,0 +1,90 @@
|
|
|
1
|
+
# Taste Calibration Protocol
|
|
2
|
+
|
|
3
|
+
> **Codex mirror adaptation (normative).** The weighted axes, human-curated
|
|
4
|
+
> anchors and GAP output contract are unchanged. A calibrated score from a fresh
|
|
5
|
+
> Codex reviewer may drive revisions and yield a provisional result; it is not
|
|
6
|
+
> cross-family acceptance unless an overlay supplies the reviewer.
|
|
7
|
+
|
|
8
|
+
> Subjective quality is gradable **if you write the taste down and anchor the
|
|
9
|
+
> scale**. The model will not invent taste; it will only converge toward the
|
|
10
|
+
> taste you described — so the whole game is (1) a weighted rubric worth
|
|
11
|
+
> converging to, and (2) reference anchors that pin what "good" and "slop"
|
|
12
|
+
> actually look like. (After Karpathy's LOOPS.md VI, "score the subjective".)
|
|
13
|
+
|
|
14
|
+
Use this protocol whenever a skill grades an artifact on axes that are matters
|
|
15
|
+
of judgment — visual design, writing elegance, proposal quality — rather than
|
|
16
|
+
machine-checkable facts. It layers ON TOP of any deterministic gates the skill
|
|
17
|
+
already has (hard caps, measurement gates); it never replaces them.
|
|
18
|
+
|
|
19
|
+
## 1. Named axes with explicit numeric weights
|
|
20
|
+
|
|
21
|
+
Define 2–7 named axes and give each an explicit weight; weights sum to 1.0.
|
|
22
|
+
Model the table on `research-refine`'s working precedent (its Phase 2 uses
|
|
23
|
+
15/25/25/15/10/5/5% across seven axes):
|
|
24
|
+
|
|
25
|
+
```markdown
|
|
26
|
+
| Axis | Weight | What it measures |
|
|
27
|
+
|---------------|:------:|---------------------------------------------------|
|
|
28
|
+
| Design | 0.35 | hierarchy, spacing, restraint, gestalt |
|
|
29
|
+
| Originality | 0.15 | distinct voice vs template sameness |
|
|
30
|
+
| Craft | 0.30 | detail quality: typography, alignment, math, figs |
|
|
31
|
+
| Functionality | 0.20 | does it do its job (readable at distance, ...) |
|
|
32
|
+
```
|
|
33
|
+
|
|
34
|
+
Score each axis 0–1 (or 1–10 rescaled). Composite = Σ weightᵢ · axisᵢ. A single
|
|
35
|
+
holistic number without named weighted axes is not this protocol.
|
|
36
|
+
|
|
37
|
+
## 2. Calibrate on reference anchors BEFORE scoring the target
|
|
38
|
+
|
|
39
|
+
The grader first scores **3 known-good and 3 known-bad reference exemplars** on
|
|
40
|
+
the same axes, so the scale is anchored to concrete artifacts instead of the
|
|
41
|
+
grader's free-floating prior.
|
|
42
|
+
|
|
43
|
+
- References are **pre-existing, human-curated files** — the executor never
|
|
44
|
+
selects, generates, or searches for anchors itself (an executor-picked anchor
|
|
45
|
+
set just smuggles the free-floating prior back in). The invoking skill
|
|
46
|
+
supplies the paths — convention: `<skill-dir>/references/good/` and
|
|
47
|
+
`<skill-dir>/references/bad/` (images, PDFs, or text artifacts of the same
|
|
48
|
+
kind as the target). Project-local references may override the skill-local
|
|
49
|
+
set.
|
|
50
|
+
- The grader is TOLD which set is which ("these three are good, these three are
|
|
51
|
+
slop") — calibration is about anchoring the scale, not blind classification.
|
|
52
|
+
- Sanity check: if the calibrated scores don't separate the sets (a "bad"
|
|
53
|
+
exemplar scores at or above a "good" one on the composite), the rubric is
|
|
54
|
+
broken — fix the rubric before trusting any target score.
|
|
55
|
+
|
|
56
|
+
**Graceful degradation:** if no reference sets exist, proceed with the weighted
|
|
57
|
+
rubric alone, but the output MUST carry `calibration: none` so a downstream
|
|
58
|
+
reader never mistakes an unanchored score for an anchored one. Do not fabricate
|
|
59
|
+
or hallucinate reference scores.
|
|
60
|
+
|
|
61
|
+
## 3. Output contract
|
|
62
|
+
|
|
63
|
+
```
|
|
64
|
+
COMPOSITE: 0.xx (weighted; also give per-axis scores)
|
|
65
|
+
CALIBRATION: anchored | none
|
|
66
|
+
GAP: <one mandatory paragraph naming WHICH reference exemplar(s) the target
|
|
67
|
+
falls short of or exceeds, on WHICH axes, and why — "0.71 because the
|
|
68
|
+
figure hierarchy matches good/poster_B but the typography is closer to
|
|
69
|
+
bad/poster_A's crowding" — never just a number>
|
|
70
|
+
```
|
|
71
|
+
|
|
72
|
+
The GAP paragraph is what makes the score actionable: converging toward the
|
|
73
|
+
described taste requires knowing where the artifact sits relative to the
|
|
74
|
+
anchors, not just its scalar.
|
|
75
|
+
|
|
76
|
+
## 4. Interaction with existing gates and the jury
|
|
77
|
+
|
|
78
|
+
- **Deterministic caps stay hard floors.** A calibrated composite never
|
|
79
|
+
overrides a measurement gate or critical cap (e.g. paper-poster-html's
|
|
80
|
+
"< 2 real figures → ≤ 3"). Compute caps first; the composite lives under
|
|
81
|
+
them.
|
|
82
|
+
- **Calibration ≠ acquittal.** A taste score produced by the executor's own
|
|
83
|
+
model family may DRIVE the fix loop (rank issues, decide what to patch next)
|
|
84
|
+
but can never acquit: wherever a skill's acceptance requires a cross-model
|
|
85
|
+
verdict, that requirement is unchanged. (Per `acceptance-gate.md`'s taxonomy
|
|
86
|
+
a model-assigned score is a semantic judgment — calibration narrows its
|
|
87
|
+
variance; it does not make it machine-checkable.)
|
|
88
|
+
- **Rubric drift is meta-optimize's job.** If users repeatedly override
|
|
89
|
+
calibrated scores, the rubric or the anchors are wrong — that's an event-log
|
|
90
|
+
signal for `/meta-optimize`, not a reason to hand-tweak scores per run.
|
|
@@ -0,0 +1,73 @@
|
|
|
1
|
+
# Venue Checklists for ICLR, NeurIPS, and ICML
|
|
2
|
+
|
|
3
|
+
Use this reference near the end of `paper-plan` and during the final checks in `paper-write`.
|
|
4
|
+
|
|
5
|
+
## When to Read
|
|
6
|
+
|
|
7
|
+
- Read once when setting the target venue.
|
|
8
|
+
- Read again before locking the outline.
|
|
9
|
+
- Read again during final submission-readiness checks.
|
|
10
|
+
|
|
11
|
+
## Universal Requirements
|
|
12
|
+
|
|
13
|
+
Across these venues, the following are usually expected:
|
|
14
|
+
|
|
15
|
+
- anonymous submission unless preparing a camera-ready version,
|
|
16
|
+
- references and appendices outside the main page budget,
|
|
17
|
+
- enough experimental detail for reproduction,
|
|
18
|
+
- honest limitations and scope boundaries,
|
|
19
|
+
- clear mapping from claims to evidence.
|
|
20
|
+
|
|
21
|
+
## NeurIPS
|
|
22
|
+
|
|
23
|
+
Planning implications:
|
|
24
|
+
|
|
25
|
+
- The paper checklist is mandatory.
|
|
26
|
+
- Claims in the Abstract and Introduction must align with the actual evidence.
|
|
27
|
+
- The paper should discuss limitations honestly.
|
|
28
|
+
- Reproducibility details, hyperparameters, data access, and compute usage should be documented.
|
|
29
|
+
- Statistical reporting should specify error bars, number of runs, and how uncertainty is computed.
|
|
30
|
+
|
|
31
|
+
Final-check implications:
|
|
32
|
+
|
|
33
|
+
- Confirm the paper checklist is complete.
|
|
34
|
+
- Ensure limitations, reproducibility details, and compute reporting exist somewhere appropriate.
|
|
35
|
+
- Verify theory papers include assumptions and full proofs in the main paper or appendix.
|
|
36
|
+
|
|
37
|
+
## ICML
|
|
38
|
+
|
|
39
|
+
Planning implications:
|
|
40
|
+
|
|
41
|
+
- The paper must budget space for an ICML-style Broader Impact statement.
|
|
42
|
+
- Reproducibility expectations are strong: data splits, hyperparameters, search ranges, and compute should be documented.
|
|
43
|
+
- Statistical reporting should state whether uncertainty uses standard deviation, standard error, or confidence intervals.
|
|
44
|
+
|
|
45
|
+
Final-check implications:
|
|
46
|
+
|
|
47
|
+
- Ensure the Broader Impact statement is present in the expected location.
|
|
48
|
+
- Confirm anonymization is strict: no author names, acknowledgments, grant IDs, or self-identifying repository links.
|
|
49
|
+
- Verify experimental details are detailed enough for replication.
|
|
50
|
+
|
|
51
|
+
## ICLR
|
|
52
|
+
|
|
53
|
+
Planning implications:
|
|
54
|
+
|
|
55
|
+
- Reproducibility and ethics statements are often recommended even if not always mandatory.
|
|
56
|
+
- If LLMs materially contributed to ideation or writing to the point of authorship-like contribution, plan a disclosure section or appendix note.
|
|
57
|
+
- Keep the story front-loaded because ICLR reviewers often judge quickly from the early pages.
|
|
58
|
+
|
|
59
|
+
Final-check implications:
|
|
60
|
+
|
|
61
|
+
- Decide whether LLM disclosure is required for this project.
|
|
62
|
+
- Confirm the paper includes enough reproducibility guidance, code/data availability information, and limitations discussion.
|
|
63
|
+
- Check that the contribution is already clear by the end of the Introduction.
|
|
64
|
+
|
|
65
|
+
## Minimal Submission Checklist
|
|
66
|
+
|
|
67
|
+
Before submission, verify:
|
|
68
|
+
|
|
69
|
+
- the venue-specific required sections are present,
|
|
70
|
+
- the page budget is satisfied for the main body,
|
|
71
|
+
- the contribution bullets do not overclaim,
|
|
72
|
+
- citations, figures, tables, and references are internally consistent,
|
|
73
|
+
- the PDF is anonymized and ready for reviewer consumption.
|