dsh-aris-panel 0.1.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/LICENSE +21 -0
- package/README.md +98 -0
- package/README_CN.md +87 -0
- package/dsh/checkout.patch.yml +38 -0
- package/dsh/client.js +634 -0
- package/dsh/cordis.patch.yml +44 -0
- package/dsh/index.mjs +76 -0
- package/dsh/run-status.mjs +182 -0
- package/dsh/scope-limits.mjs +50 -0
- package/dsh/workbench.mjs +291 -0
- package/mcp-servers/claude-review/README.md +93 -0
- package/mcp-servers/claude-review/run_with_claude_aws.sh +49 -0
- package/mcp-servers/claude-review/server.py +718 -0
- package/mcp-servers/codex-image2/README.md +65 -0
- package/mcp-servers/codex-image2/server.py +893 -0
- package/mcp-servers/feishu-bridge/requirements.txt +1 -0
- package/mcp-servers/feishu-bridge/server.py +240 -0
- package/mcp-servers/gemini-review/README.md +171 -0
- package/mcp-servers/gemini-review/server.py +1856 -0
- package/mcp-servers/llm-chat/requirements.txt +1 -0
- package/mcp-servers/llm-chat/server.py +664 -0
- package/mcp-servers/manual-review/README.md +133 -0
- package/mcp-servers/manual-review/server.py +910 -0
- package/mcp-servers/manual-review/ui.html +279 -0
- package/mcp-servers/minimax-chat/requirements.txt +1 -0
- package/mcp-servers/minimax-chat/server.py +381 -0
- package/package.json +51 -0
- package/skills/ablation-planner/SKILL.md +123 -0
- package/skills/alphaxiv/SKILL.md +196 -0
- package/skills/analyze-results/SKILL.md +46 -0
- package/skills/arxiv/SKILL.md +248 -0
- package/skills/auto-paper-improvement-loop/SKILL.md +651 -0
- package/skills/auto-review-loop/SKILL.md +1137 -0
- package/skills/auto-review-loop-llm/SKILL.md +259 -0
- package/skills/auto-review-loop-minimax/SKILL.md +302 -0
- package/skills/citation-audit/SKILL.md +502 -0
- package/skills/claims-drafting/SKILL.md +227 -0
- package/skills/comm-lit-review/SKILL.md +297 -0
- package/skills/deepxiv/SKILL.md +263 -0
- package/skills/dse-loop/SKILL.md +296 -0
- package/skills/embodiment-description/SKILL.md +129 -0
- package/skills/exa-search/SKILL.md +205 -0
- package/skills/experiment-audit/SKILL.md +311 -0
- package/skills/experiment-bridge/SKILL.md +376 -0
- package/skills/experiment-plan/SKILL.md +249 -0
- package/skills/experiment-queue/SKILL.md +431 -0
- package/skills/experiment-queue/scripts/build_manifest.py +142 -0
- package/skills/experiment-queue/scripts/queue_manager.py +433 -0
- package/skills/feishu-notify/SKILL.md +156 -0
- package/skills/figure-description/SKILL.md +138 -0
- package/skills/figure-spec/SKILL.md +262 -0
- package/skills/figure-spec/scripts/figure_renderer.py +799 -0
- package/skills/formula-derivation/SKILL.md +280 -0
- package/skills/gemini-search/SKILL.md +231 -0
- package/skills/grant-proposal/SKILL.md +698 -0
- package/skills/idea-creator/SKILL.md +542 -0
- package/skills/idea-discovery/SKILL.md +521 -0
- package/skills/idea-discovery-robot/SKILL.md +363 -0
- package/skills/integrity-forensics/SKILL.md +284 -0
- package/skills/interview-cheatsheet/SKILL.md +245 -0
- package/skills/invention-structuring/SKILL.md +188 -0
- package/skills/jurisdiction-format/SKILL.md +192 -0
- package/skills/kill-argument/SKILL.md +437 -0
- package/skills/mermaid-diagram/SKILL.md +419 -0
- package/skills/meta-apply/SKILL.md +141 -0
- package/skills/meta-optimize/SKILL.md +437 -0
- package/skills/monitor-experiment/SKILL.md +140 -0
- package/skills/novelty-check/SKILL.md +101 -0
- package/skills/openalex/SKILL.md +237 -0
- package/skills/overleaf-sync/SKILL.md +220 -0
- package/skills/paper-claim-audit/SKILL.md +348 -0
- package/skills/paper-compile/SKILL.md +266 -0
- package/skills/paper-figure/SKILL.md +312 -0
- package/skills/paper-illustration/SKILL.md +736 -0
- package/skills/paper-illustration-image2/SKILL.md +391 -0
- package/skills/paper-illustration-image2/scripts/paper_illustration_image2.py +255 -0
- package/skills/paper-plan/SKILL.md +386 -0
- package/skills/paper-poster/SKILL.md +19 -0
- package/skills/paper-poster-html/DESIGN_FINAL.md +176 -0
- package/skills/paper-poster-html/IMPLEMENTATION_CONVENTIONS.md +161 -0
- package/skills/paper-poster-html/LICENSES/posterly-MIT.txt +21 -0
- package/skills/paper-poster-html/NOTICE.md +57 -0
- package/skills/paper-poster-html/SKILL.md +323 -0
- package/skills/paper-poster-html/scripts/_posterly/__init__.py +0 -0
- package/skills/paper-poster-html/scripts/_posterly/canvas.py +200 -0
- package/skills/paper-poster-html/scripts/_posterly/measure.py +588 -0
- package/skills/paper-poster-html/scripts/_posterly/polish.py +498 -0
- package/skills/paper-poster-html/scripts/_posterly/preflight.py +489 -0
- package/skills/paper-poster-html/scripts/_posterly/render.py +215 -0
- package/skills/paper-poster-html/scripts/_posterly/textutil.py +16 -0
- package/skills/paper-poster-html/scripts/_posterly/verify_final.py +171 -0
- package/skills/paper-poster-html/scripts/asset_check.py +897 -0
- package/skills/paper-poster-html/scripts/extract_pdf_figures.py +666 -0
- package/skills/paper-poster-html/scripts/poster_check.py +251 -0
- package/skills/paper-poster-html/scripts/preprocess_figures.py +238 -0
- package/skills/paper-poster-html/scripts/render_preview.py +217 -0
- package/skills/paper-poster-html/scripts/run_gates.py +556 -0
- package/skills/paper-poster-html/scripts/style_check.py +1324 -0
- package/skills/paper-poster-html/templates/COMPONENTS.md +462 -0
- package/skills/paper-poster-html/templates/README.md +170 -0
- package/skills/paper-poster-html/templates/landscape_4col.html +1032 -0
- package/skills/paper-poster-html/templates/landscape_hero.html +1046 -0
- package/skills/paper-poster-html/templates/portrait_2col.html +947 -0
- package/skills/paper-poster-html/templates/tokens/acl.json +9 -0
- package/skills/paper-poster-html/templates/tokens/cvpr.json +9 -0
- package/skills/paper-poster-html/templates/tokens/generic.json +9 -0
- package/skills/paper-poster-html/templates/tokens/iclr.json +9 -0
- package/skills/paper-poster-html/templates/tokens/icml.json +9 -0
- package/skills/paper-poster-html/templates/tokens/neurips.json +9 -0
- package/skills/paper-slides/SKILL.md +635 -0
- package/skills/paper-talk/SKILL.md +381 -0
- package/skills/paper-write/SKILL.md +604 -0
- package/skills/paper-write/templates/IEEEtran.bst +2409 -0
- package/skills/paper-write/templates/IEEEtran.cls +6347 -0
- package/skills/paper-write/templates/iclr2026.tex +84 -0
- package/skills/paper-write/templates/icml2025.tex +87 -0
- package/skills/paper-write/templates/ieee_conference.tex +89 -0
- package/skills/paper-write/templates/ieee_journal.tex +93 -0
- package/skills/paper-write/templates/math_commands.tex +48 -0
- package/skills/paper-write/templates/neurips2025.tex +80 -0
- package/skills/paper-writing/SKILL.md +916 -0
- package/skills/patent-novelty-check/SKILL.md +153 -0
- package/skills/patent-pipeline/SKILL.md +344 -0
- package/skills/patent-review/SKILL.md +203 -0
- package/skills/pixel-art/SKILL.md +137 -0
- package/skills/prior-art-search/SKILL.md +146 -0
- package/skills/proof-checker/SKILL.md +866 -0
- package/skills/proof-orchestrator/NOTICE.md +24 -0
- package/skills/proof-orchestrator/SKILL.md +254 -0
- package/skills/proof-orchestrator/references/audit-output-contract.md +126 -0
- package/skills/proof-orchestrator/references/deepseek-routing.md +74 -0
- package/skills/proof-orchestrator/references/dispatch-prompts.md +227 -0
- package/skills/proof-orchestrator/references/notation-audit.md +135 -0
- package/skills/proof-orchestrator/references/proof-audit-rubric.md +70 -0
- package/skills/proof-orchestrator/references/stress-tests.md +38 -0
- package/skills/proof-writer/SKILL.md +223 -0
- package/skills/qzcli/SKILL.md +324 -0
- package/skills/rebuttal/SKILL.md +376 -0
- package/skills/render-html/SKILL.md +316 -0
- package/skills/render-html/scripts/render_html.py +1006 -0
- package/skills/render-html/scripts/templates/academic.html +703 -0
- package/skills/render-html/scripts/templates/dashboard.html +333 -0
- package/skills/research-lit/SKILL.md +756 -0
- package/skills/research-pipeline/SKILL.md +384 -0
- package/skills/research-refine/SKILL.md +770 -0
- package/skills/research-refine-pipeline/SKILL.md +186 -0
- package/skills/research-review/SKILL.md +198 -0
- package/skills/research-wiki/SKILL.md +461 -0
- package/skills/resubmit-pipeline/SKILL.md +447 -0
- package/skills/result-to-claim/SKILL.md +311 -0
- package/skills/run-experiment/SKILL.md +313 -0
- package/skills/semantic-scholar/SKILL.md +236 -0
- package/skills/serverless-modal/SKILL.md +335 -0
- package/skills/shared-references/acceptance-gate.md +324 -0
- package/skills/shared-references/assurance-contract.md +248 -0
- package/skills/shared-references/capture-antipatterns.md +78 -0
- package/skills/shared-references/citation-discipline.md +583 -0
- package/skills/shared-references/compute-env-contract.md +163 -0
- package/skills/shared-references/effort-contract.md +183 -0
- package/skills/shared-references/evidence-precheck.md +65 -0
- package/skills/shared-references/experiment-integrity.md +49 -0
- package/skills/shared-references/external-cadence.md +326 -0
- package/skills/shared-references/fan-out-pattern.md +366 -0
- package/skills/shared-references/injection-hygiene.md +127 -0
- package/skills/shared-references/integration-contract.md +461 -0
- package/skills/shared-references/output-composition.md +93 -0
- package/skills/shared-references/output-language.md +45 -0
- package/skills/shared-references/output-manifest.md +49 -0
- package/skills/shared-references/output-versioning.md +111 -0
- package/skills/shared-references/patent-format-cn.md +199 -0
- package/skills/shared-references/patent-format-ep.md +173 -0
- package/skills/shared-references/patent-format-us.md +161 -0
- package/skills/shared-references/patent-writing-principles.md +197 -0
- package/skills/shared-references/prior-art-databases.md +141 -0
- package/skills/shared-references/resumable-runs.md +109 -0
- package/skills/shared-references/review-scope-limits.md +81 -0
- package/skills/shared-references/review-tracing.md +391 -0
- package/skills/shared-references/reviewer-independence.md +79 -0
- package/skills/shared-references/reviewer-routing.md +852 -0
- package/skills/shared-references/skill-governance.md +104 -0
- package/skills/shared-references/taste-calibration.md +85 -0
- package/skills/shared-references/venue-checklists.md +114 -0
- package/skills/shared-references/wiki-helper-resolution.md +134 -0
- package/skills/shared-references/writing-principles.md +525 -0
- package/skills/skills-codex/README.md +102 -0
- package/skills/skills-codex/README_CN.md +100 -0
- package/skills/skills-codex/ablation-planner/SKILL.md +126 -0
- package/skills/skills-codex/alphaxiv/SKILL.md +186 -0
- package/skills/skills-codex/analyze-results/SKILL.md +45 -0
- package/skills/skills-codex/arxiv/SKILL.md +210 -0
- package/skills/skills-codex/auto-paper-improvement-loop/SKILL.md +574 -0
- package/skills/skills-codex/auto-review-loop/SKILL.md +500 -0
- package/skills/skills-codex/auto-review-loop-llm/SKILL.md +247 -0
- package/skills/skills-codex/auto-review-loop-minimax/SKILL.md +290 -0
- package/skills/skills-codex/citation-audit/SKILL.md +504 -0
- package/skills/skills-codex/claims-drafting/SKILL.md +239 -0
- package/skills/skills-codex/comm-lit-review/SKILL.md +299 -0
- package/skills/skills-codex/comm-lit-review/references/domain-taxonomy.md +57 -0
- package/skills/skills-codex/comm-lit-review/references/output-template.md +37 -0
- package/skills/skills-codex/comm-lit-review/references/source-policy.md +99 -0
- package/skills/skills-codex/comm-lit-review/references/venue-tiering.md +112 -0
- package/skills/skills-codex/deepxiv/SKILL.md +142 -0
- package/skills/skills-codex/dse-loop/SKILL.md +285 -0
- package/skills/skills-codex/embodiment-description/SKILL.md +129 -0
- package/skills/skills-codex/exa-search/SKILL.md +192 -0
- package/skills/skills-codex/experiment-audit/SKILL.md +286 -0
- package/skills/skills-codex/experiment-bridge/SKILL.md +356 -0
- package/skills/skills-codex/experiment-plan/SKILL.md +249 -0
- package/skills/skills-codex/experiment-queue/SKILL.md +401 -0
- package/skills/skills-codex/feishu-notify/SKILL.md +155 -0
- package/skills/skills-codex/figure-description/SKILL.md +138 -0
- package/skills/skills-codex/figure-spec/SKILL.md +252 -0
- package/skills/skills-codex/formula-derivation/SKILL.md +280 -0
- package/skills/skills-codex/gemini-search/SKILL.md +205 -0
- package/skills/skills-codex/grant-proposal/SKILL.md +626 -0
- package/skills/skills-codex/idea-creator/SKILL.md +405 -0
- package/skills/skills-codex/idea-discovery/SKILL.md +475 -0
- package/skills/skills-codex/idea-discovery-robot/SKILL.md +362 -0
- package/skills/skills-codex/integrity-forensics/SKILL.md +106 -0
- package/skills/skills-codex/interview-cheatsheet/SKILL.md +245 -0
- package/skills/skills-codex/invention-structuring/SKILL.md +188 -0
- package/skills/skills-codex/jurisdiction-format/SKILL.md +192 -0
- package/skills/skills-codex/kill-argument/SKILL.md +403 -0
- package/skills/skills-codex/mermaid-diagram/SKILL.md +379 -0
- package/skills/skills-codex/meta-apply/SKILL.md +154 -0
- package/skills/skills-codex/meta-optimize/SKILL.md +348 -0
- package/skills/skills-codex/monitor-experiment/SKILL.md +98 -0
- package/skills/skills-codex/novelty-check/SKILL.md +89 -0
- package/skills/skills-codex/openalex/SKILL.md +228 -0
- package/skills/skills-codex/overleaf-sync/SKILL.md +220 -0
- package/skills/skills-codex/paper-claim-audit/SKILL.md +350 -0
- package/skills/skills-codex/paper-compile/SKILL.md +253 -0
- package/skills/skills-codex/paper-figure/SKILL.md +311 -0
- package/skills/skills-codex/paper-illustration/SKILL.md +690 -0
- package/skills/skills-codex/paper-illustration-image2/SKILL.md +383 -0
- package/skills/skills-codex/paper-illustration-image2/scripts/paper_illustration_image2.py +255 -0
- package/skills/skills-codex/paper-plan/SKILL.md +278 -0
- package/skills/skills-codex/paper-poster/SKILL.md +19 -0
- package/skills/skills-codex/paper-poster-html/SKILL.md +377 -0
- package/skills/skills-codex/paper-slides/SKILL.md +571 -0
- package/skills/skills-codex/paper-talk/SKILL.md +381 -0
- package/skills/skills-codex/paper-write/SKILL.md +411 -0
- package/skills/skills-codex/paper-write/templates/IEEEtran.bst +2409 -0
- package/skills/skills-codex/paper-write/templates/IEEEtran.cls +6347 -0
- package/skills/skills-codex/paper-write/templates/aaai2026.bst +1493 -0
- package/skills/skills-codex/paper-write/templates/aaai2026.sty +315 -0
- package/skills/skills-codex/paper-write/templates/aaai2026.tex +952 -0
- package/skills/skills-codex/paper-write/templates/acl.sty +312 -0
- package/skills/skills-codex/paper-write/templates/acl2026.tex +377 -0
- package/skills/skills-codex/paper-write/templates/acl_natbib.bst +1940 -0
- package/skills/skills-codex/paper-write/templates/acm.bst +3081 -0
- package/skills/skills-codex/paper-write/templates/acm_mm2026.tex +204 -0
- package/skills/skills-codex/paper-write/templates/acmart.cls +3520 -0
- package/skills/skills-codex/paper-write/templates/cvpr.bst +1448 -0
- package/skills/skills-codex/paper-write/templates/cvpr.sty +508 -0
- package/skills/skills-codex/paper-write/templates/cvpr2026.tex +63 -0
- package/skills/skills-codex/paper-write/templates/iclr2026.tex +84 -0
- package/skills/skills-codex/paper-write/templates/iclr2026_conference.bst +1440 -0
- package/skills/skills-codex/paper-write/templates/iclr2026_conference.sty +246 -0
- package/skills/skills-codex/paper-write/templates/icml2026.sty +767 -0
- package/skills/skills-codex/paper-write/templates/icml2026.tex +662 -0
- package/skills/skills-codex/paper-write/templates/ieee_conference.tex +89 -0
- package/skills/skills-codex/paper-write/templates/ieee_journal.tex +93 -0
- package/skills/skills-codex/paper-write/templates/math_commands.tex +48 -0
- package/skills/skills-codex/paper-write/templates/neurips2026.tex +493 -0
- package/skills/skills-codex/paper-write/templates/neurips_2026.sty +437 -0
- package/skills/skills-codex/paper-writing/SKILL.md +731 -0
- package/skills/skills-codex/patent-novelty-check/SKILL.md +153 -0
- package/skills/skills-codex/patent-pipeline/SKILL.md +344 -0
- package/skills/skills-codex/patent-review/SKILL.md +202 -0
- package/skills/skills-codex/pixel-art/SKILL.md +139 -0
- package/skills/skills-codex/prior-art-search/SKILL.md +146 -0
- package/skills/skills-codex/proof-checker/SKILL.md +554 -0
- package/skills/skills-codex/proof-orchestrator/SKILL.md +260 -0
- package/skills/skills-codex/proof-orchestrator/references/audit-output-contract.md +126 -0
- package/skills/skills-codex/proof-orchestrator/references/deepseek-routing.md +76 -0
- package/skills/skills-codex/proof-orchestrator/references/dispatch-prompts.md +227 -0
- package/skills/skills-codex/proof-orchestrator/references/notation-audit.md +135 -0
- package/skills/skills-codex/proof-orchestrator/references/proof-audit-rubric.md +70 -0
- package/skills/skills-codex/proof-orchestrator/references/stress-tests.md +38 -0
- package/skills/skills-codex/proof-writer/SKILL.md +222 -0
- package/skills/skills-codex/qzcli/SKILL.md +324 -0
- package/skills/skills-codex/rebuttal/SKILL.md +305 -0
- package/skills/skills-codex/render-html/SKILL.md +305 -0
- package/skills/skills-codex/render-html/scripts/__pycache__/render_html.cpython-314.pyc +0 -0
- package/skills/skills-codex/render-html/scripts/render_html.py +909 -0
- package/skills/skills-codex/render-html/scripts/templates/academic.html +342 -0
- package/skills/skills-codex/render-html/scripts/templates/dashboard.html +333 -0
- package/skills/skills-codex/research-lit/SKILL.md +464 -0
- package/skills/skills-codex/research-pipeline/SKILL.md +340 -0
- package/skills/skills-codex/research-refine/SKILL.md +721 -0
- package/skills/skills-codex/research-refine-pipeline/SKILL.md +186 -0
- package/skills/skills-codex/research-review/SKILL.md +135 -0
- package/skills/skills-codex/research-wiki/SKILL.md +421 -0
- package/skills/skills-codex/resubmit-pipeline/SKILL.md +444 -0
- package/skills/skills-codex/result-to-claim/SKILL.md +246 -0
- package/skills/skills-codex/run-experiment/SKILL.md +236 -0
- package/skills/skills-codex/semantic-scholar/SKILL.md +219 -0
- package/skills/skills-codex/serverless-modal/SKILL.md +335 -0
- package/skills/skills-codex/shared-references/acceptance-gate.md +336 -0
- package/skills/skills-codex/shared-references/assurance-contract.md +139 -0
- package/skills/skills-codex/shared-references/capture-antipatterns.md +84 -0
- package/skills/skills-codex/shared-references/citation-discipline.md +452 -0
- package/skills/skills-codex/shared-references/compute-env-contract.md +163 -0
- package/skills/skills-codex/shared-references/effort-contract.md +143 -0
- package/skills/skills-codex/shared-references/evidence-precheck.md +73 -0
- package/skills/skills-codex/shared-references/experiment-integrity.md +49 -0
- package/skills/skills-codex/shared-references/external-cadence.md +334 -0
- package/skills/skills-codex/shared-references/fan-out-pattern.md +375 -0
- package/skills/skills-codex/shared-references/injection-hygiene.md +132 -0
- package/skills/skills-codex/shared-references/integration-contract.md +372 -0
- package/skills/skills-codex/shared-references/output-composition.md +98 -0
- package/skills/skills-codex/shared-references/output-language.md +45 -0
- package/skills/skills-codex/shared-references/output-manifest.md +40 -0
- package/skills/skills-codex/shared-references/output-versioning.md +111 -0
- package/skills/skills-codex/shared-references/patent-format-cn.md +199 -0
- package/skills/skills-codex/shared-references/patent-format-ep.md +173 -0
- package/skills/skills-codex/shared-references/patent-format-us.md +161 -0
- package/skills/skills-codex/shared-references/patent-writing-principles.md +197 -0
- package/skills/skills-codex/shared-references/prior-art-databases.md +141 -0
- package/skills/skills-codex/shared-references/resumable-runs.md +125 -0
- package/skills/skills-codex/shared-references/review-scope-limits.md +81 -0
- package/skills/skills-codex/shared-references/review-tracing.md +144 -0
- package/skills/skills-codex/shared-references/reviewer-independence.md +66 -0
- package/skills/skills-codex/shared-references/reviewer-routing.md +128 -0
- package/skills/skills-codex/shared-references/skill-governance.md +119 -0
- package/skills/skills-codex/shared-references/taste-calibration.md +90 -0
- package/skills/skills-codex/shared-references/venue-checklists.md +73 -0
- package/skills/skills-codex/shared-references/wiki-helper-resolution.md +69 -0
- package/skills/skills-codex/shared-references/writing-principles.md +525 -0
- package/skills/skills-codex/slides-polish/SKILL.md +563 -0
- package/skills/skills-codex/specification-writing/SKILL.md +211 -0
- package/skills/skills-codex/system-profile/SKILL.md +103 -0
- package/skills/skills-codex/training-check/SKILL.md +83 -0
- package/skills/skills-codex/vast-gpu/SKILL.md +394 -0
- package/skills/skills-codex/web-debug-search/SKILL.md +334 -0
- package/skills/skills-codex/wiki-enrich/SKILL.md +255 -0
- package/skills/skills-codex/writing-systems-papers/SKILL.md +184 -0
- package/skills/skills-codex-claude-review/README.md +79 -0
- package/skills/skills-codex-claude-review/README_CN.md +78 -0
- package/skills/skills-codex-claude-review/auto-paper-improvement-loop/SKILL.md +581 -0
- package/skills/skills-codex-claude-review/auto-review-loop/SKILL.md +510 -0
- package/skills/skills-codex-claude-review/novelty-check/SKILL.md +102 -0
- package/skills/skills-codex-claude-review/paper-figure/SKILL.md +319 -0
- package/skills/skills-codex-claude-review/paper-plan/SKILL.md +287 -0
- package/skills/skills-codex-claude-review/paper-write/SKILL.md +420 -0
- package/skills/skills-codex-claude-review/research-refine/SKILL.md +732 -0
- package/skills/skills-codex-claude-review/research-review/SKILL.md +149 -0
- package/skills/skills-codex-gemini-review/README.md +176 -0
- package/skills/skills-codex-gemini-review/README_CN.md +175 -0
- package/skills/skills-codex-gemini-review/auto-paper-improvement-loop/SKILL.md +331 -0
- package/skills/skills-codex-gemini-review/auto-review-loop/SKILL.md +304 -0
- package/skills/skills-codex-gemini-review/grant-proposal/SKILL.md +630 -0
- package/skills/skills-codex-gemini-review/idea-creator/SKILL.md +263 -0
- package/skills/skills-codex-gemini-review/idea-discovery/SKILL.md +275 -0
- package/skills/skills-codex-gemini-review/idea-discovery-robot/SKILL.md +365 -0
- package/skills/skills-codex-gemini-review/novelty-check/SKILL.md +92 -0
- package/skills/skills-codex-gemini-review/paper-figure/SKILL.md +289 -0
- package/skills/skills-codex-gemini-review/paper-plan/SKILL.md +265 -0
- package/skills/skills-codex-gemini-review/paper-poster-html/SKILL.md +102 -0
- package/skills/skills-codex-gemini-review/paper-slides/SKILL.md +582 -0
- package/skills/skills-codex-gemini-review/paper-write/SKILL.md +346 -0
- package/skills/skills-codex-gemini-review/paper-writing/SKILL.md +312 -0
- package/skills/skills-codex-gemini-review/research-refine/SKILL.md +674 -0
- package/skills/skills-codex-gemini-review/research-review/SKILL.md +112 -0
- package/skills/slides-polish/SKILL.md +565 -0
- package/skills/specification-writing/SKILL.md +211 -0
- package/skills/system-profile/SKILL.md +103 -0
- package/skills/training-check/SKILL.md +132 -0
- package/skills/vast-gpu/SKILL.md +394 -0
- package/skills/web-debug-search/SKILL.md +334 -0
- package/skills/wiki-enrich/SKILL.md +257 -0
- package/skills/writing-systems-papers/SKILL.md +184 -0
- package/templates/CLAUDE_MD_TEMPLATE.md +29 -0
- package/templates/EXPERIMENT_LOG_TEMPLATE.md +47 -0
- package/templates/EXPERIMENT_PLAN_TEMPLATE.md +51 -0
- package/templates/EXPERIMENT_PLAN_TEMPLATE_CN.md +53 -0
- package/templates/FINDINGS_TEMPLATE.md +52 -0
- package/templates/IDEA_CANDIDATES_TEMPLATE.md +47 -0
- package/templates/IDEA_CANDIDATES_TEMPLATE_CN.md +47 -0
- package/templates/INVENTION_BRIEF_TEMPLATE.md +87 -0
- package/templates/MANIFEST_TEMPLATE.md +7 -0
- package/templates/NARRATIVE_REPORT_TEMPLATE.md +49 -0
- package/templates/PAPER_PLAN_TEMPLATE.md +47 -0
- package/templates/PATENT_CLAIMS_TEMPLATE.md +78 -0
- package/templates/PATENT_SPECIFICATION_TEMPLATE.md +68 -0
- package/templates/README.md +57 -0
- package/templates/RESEARCH_BRIEF_TEMPLATE.md +35 -0
- package/templates/RESEARCH_BRIEF_TEMPLATE_CN.md +41 -0
- package/templates/RESEARCH_CONTRACT_TEMPLATE.md +60 -0
- package/templates/claude-hooks/corpus_write_guard.json +16 -0
- package/templates/claude-hooks/corpus_write_guard.py +85 -0
- package/templates/claude-hooks/meta_logging.json +74 -0
- package/templates/gitignore-trace.txt +3 -0
- package/tools/__pycache__/check_skills_inventory.cpython-314.pyc +0 -0
- package/tools/arxiv_fetch.py +311 -0
- package/tools/capture_filter.py +126 -0
- package/tools/check_skills_inventory.py +273 -0
- package/tools/convert_skills_to_llm_chat.py +282 -0
- package/tools/copilot_native_evidence.py +818 -0
- package/tools/deepxiv_fetch.py +213 -0
- package/tools/evidence_check.py +212 -0
- package/tools/exa_search.py +425 -0
- package/tools/experiment_queue/README.md +118 -0
- package/tools/experiment_queue/build_manifest.py +44 -0
- package/tools/experiment_queue/queue_manager.py +44 -0
- package/tools/extract_paper_style.py +560 -0
- package/tools/figure_renderer.py +69 -0
- package/tools/forensics_gate.py +669 -0
- package/tools/generate_codex_claude_review_overrides.py +299 -0
- package/tools/idea_discovery_gate.py +256 -0
- package/tools/install_aris.ps1 +1372 -0
- package/tools/install_aris.sh +1370 -0
- package/tools/install_aris_codex.sh +1023 -0
- package/tools/install_aris_copilot.sh +1052 -0
- package/tools/iteration_log.py +143 -0
- package/tools/lint_skills_helpers.sh +84 -0
- package/tools/meta_opt/check_ready.sh +80 -0
- package/tools/meta_opt/log_event.sh +91 -0
- package/tools/meta_opt/trigger_eval.py +280 -0
- package/tools/meta_opt/trigger_evals.sample.json +28 -0
- package/tools/openalex_fetch.py +326 -0
- package/tools/overleaf_audit.sh +104 -0
- package/tools/overleaf_setup.sh +150 -0
- package/tools/paper_illustration_image2.py +62 -0
- package/tools/provenance.py +294 -0
- package/tools/research_wiki.py +1720 -0
- package/tools/review_gate.py +502 -0
- package/tools/run_state.py +399 -0
- package/tools/save_trace.sh +477 -0
- package/tools/semantic_scholar_fetch.py +438 -0
- package/tools/skill-groups.tsv +116 -0
- package/tools/skill_picker.py +238 -0
- package/tools/smart_update.ps1 +521 -0
- package/tools/smart_update.sh +591 -0
- package/tools/smart_update_codex.sh +419 -0
- package/tools/smart_update_copilot.sh +605 -0
- package/tools/threat_scan.py +222 -0
- package/tools/verify_paper_audits.sh +487 -0
- package/tools/verify_papers.py +613 -0
- package/tools/verify_wiki_coverage.sh +176 -0
- package/tools/watchdog.py +485 -0
|
@@ -0,0 +1,324 @@
|
|
|
1
|
+
# Acceptance-Gate Provenance
|
|
2
|
+
|
|
3
|
+
## Core Principle
|
|
4
|
+
|
|
5
|
+
**An autonomous loop's STOP/ACCEPT gate determines whether the loop is
|
|
6
|
+
same-family-safe. The thing being judged at that gate — not the loop's
|
|
7
|
+
subject matter, not how many agents ran — decides whether Claude may
|
|
8
|
+
judge it.**
|
|
9
|
+
|
|
10
|
+
ARIS has loops that keep working until a condition is met: `/auto-review-loop`,
|
|
11
|
+
`/dse-loop`, the `/experiment-bridge` auto-debug cycle, the
|
|
12
|
+
`/auto-paper-improvement-loop`, and any future "keep going until X"
|
|
13
|
+
skill. Every such loop terminates on a gate it evaluates each iteration:
|
|
14
|
+
"are we done yet?" That gate is where same-family self-acquittal sneaks
|
|
15
|
+
in. The loop body can be all Claude; the **gate** is what this contract
|
|
16
|
+
governs.
|
|
17
|
+
|
|
18
|
+
This is `reviewer-independence.md` and `experiment-integrity.md` applied
|
|
19
|
+
to the temporal/iterative case: those two cover single-shot review and
|
|
20
|
+
single-shot experiment judging; this one covers the *recurring verdict*
|
|
21
|
+
a loop makes on itself, round after round, with no human in between.
|
|
22
|
+
|
|
23
|
+
One-liner, and the whole doc in seven words:
|
|
24
|
+
|
|
25
|
+
> **A goal/loop can DRIVE; it cannot ACQUIT.**
|
|
26
|
+
|
|
27
|
+
The loop may freely *drive* itself toward a target — schedule the next
|
|
28
|
+
config, recompile, re-run the failed job, spawn ten search branches.
|
|
29
|
+
What it may not do is *acquit* its own work — declare the paper good,
|
|
30
|
+
the proof valid, the claim supported, the idea novel, the review
|
|
31
|
+
satisfied. Acquittal is a cross-model act.
|
|
32
|
+
|
|
33
|
+
## The two gate types
|
|
34
|
+
|
|
35
|
+
Classify **every** stop/accept gate of a loop as exactly one of these.
|
|
36
|
+
There is no third bucket; if a gate seems to be both, it is two gates
|
|
37
|
+
and you split it (see "Compound gates" below).
|
|
38
|
+
|
|
39
|
+
### Type-A — EXECUTION / OBJECTIVE gate
|
|
40
|
+
|
|
41
|
+
A machine-checkable or externally-observable signal of *what happened*,
|
|
42
|
+
with no judgment of *merit*. Claude **MAY** self-judge Type-A gates —
|
|
43
|
+
it is execution bookkeeping, not a verdict.
|
|
44
|
+
|
|
45
|
+
A gate is Type-A iff a non-LLM process (a shell exit code, a stat on the
|
|
46
|
+
filesystem, a counter, a parser reading a benchmark's own output) could
|
|
47
|
+
in principle answer it with the same answer Claude gives.
|
|
48
|
+
|
|
49
|
+
- ✅ exit code == 0
|
|
50
|
+
- ✅ `figures/result.png` exists / `paper/main.pdf` compiled (LaTeX returned 0)
|
|
51
|
+
- ✅ N/N jobs finished (queue drained)
|
|
52
|
+
- ✅ test suite passed (pytest exit 0)
|
|
53
|
+
- ✅ the reviewer **was invoked** (a `codex` thread returned, a JSON verdict file exists)
|
|
54
|
+
- ✅ all checklist items were **attempted** (each row touched)
|
|
55
|
+
- ✅ no `NaN` in the loss log / training reached `max_steps`
|
|
56
|
+
- ✅ the benchmark harness emitted a number and it parsed
|
|
57
|
+
- ✅ PATIENCE/TIMEOUT/MAX_ROUNDS budget exhausted (a counter hit its bound)
|
|
58
|
+
|
|
59
|
+
Type-A gates are *coverage and completion* facts. Claude self-judging
|
|
60
|
+
"did the audit run?" is fine; Claude self-judging "did the audit pass?"
|
|
61
|
+
is not (that's Type-B).
|
|
62
|
+
|
|
63
|
+
### Type-B — QUALITY / CORRECTNESS / ACCEPTANCE gate
|
|
64
|
+
|
|
65
|
+
A judgment of *merit, correctness, or sufficiency*. Claude must
|
|
66
|
+
**NEVER** self-judge a Type-B gate — it requires a **different model
|
|
67
|
+
family** (per `reviewer-routing.md`: `codex` default, `oracle-pro` on
|
|
68
|
+
request, or `manual` **only when** the human routes the prompt to a
|
|
69
|
+
genuinely non-Claude model and records which one). This is the
|
|
70
|
+
cross-model invariant, applied to the loop's terminating verdict.
|
|
71
|
+
|
|
72
|
+
- ❌ "the paper is good" / "submission-ready"
|
|
73
|
+
- ❌ "the proof is valid" / "the gap is closed"
|
|
74
|
+
- ❌ "the claim is supported by the results"
|
|
75
|
+
- ❌ "the idea is novel"
|
|
76
|
+
- ❌ "the review is satisfied" / "the weaknesses are addressed"
|
|
77
|
+
- ❌ "score >= 6" — when *Claude* assigned the score
|
|
78
|
+
- ❌ "this config is good enough to publish" / "the result is strong"
|
|
79
|
+
- ❌ "the rebuttal answers the reviewer"
|
|
80
|
+
- ❌ "the fix is correct" (as opposed to "the fix made the test pass" — that's Type-A)
|
|
81
|
+
|
|
82
|
+
A Type-B gate, left to the executor, is the loop quietly grading its own
|
|
83
|
+
homework every round and stopping the moment it likes the grade. The
|
|
84
|
+
fact that it ran a hundred iterations does not launder the verdict: a
|
|
85
|
+
hundred rounds of Claude-judging-Claude is still one model family.
|
|
86
|
+
|
|
87
|
+
### The dividing question
|
|
88
|
+
|
|
89
|
+
> *Could a dumb script with no taste answer this gate?*
|
|
90
|
+
>
|
|
91
|
+
> **Yes → Type-A** (Claude may self-judge — it's bookkeeping).
|
|
92
|
+
> **No, it needs taste / correctness / domain judgment → Type-B** (route to a different model family).
|
|
93
|
+
|
|
94
|
+
"The PDF compiled" needs no taste — Type-A. "The PDF is a good paper"
|
|
95
|
+
is nothing *but* taste — Type-B. "The job exited 0" — Type-A. "The job's
|
|
96
|
+
output is the right answer" — Type-B.
|
|
97
|
+
|
|
98
|
+
## Compound gates: split, don't average
|
|
99
|
+
|
|
100
|
+
Many natural-language stop conditions secretly bundle an A-part and a
|
|
101
|
+
B-part. `/auto-review-loop`'s real condition is *"score >= 6 AND verdict
|
|
102
|
+
contains 'ready'"* evaluated each round — but the **score and the
|
|
103
|
+
verdict both come from the cross-model reviewer**, so the A-part Claude
|
|
104
|
+
owns is only "did round N's reviewer return?" and "is round < MAX_ROUNDS?".
|
|
105
|
+
|
|
106
|
+
When you meet a compound gate, decompose it:
|
|
107
|
+
|
|
108
|
+
```
|
|
109
|
+
STOP when "the paper is submission-ready"
|
|
110
|
+
├─ A: all 3 audits were invoked and emitted JSON → Claude self-checks
|
|
111
|
+
├─ A: verify_paper_audits.sh exit code == 0 → external process, Claude reads it
|
|
112
|
+
└─ B: "the paper is actually good enough to submit" → cross-model verdict
|
|
113
|
+
```
|
|
114
|
+
|
|
115
|
+
Never collapse a compound gate to its A-part and call the loop safe. The
|
|
116
|
+
B-part doesn't disappear because it's inconvenient; it gets *routed*.
|
|
117
|
+
|
|
118
|
+
## Decision procedure (for any new autonomous loop)
|
|
119
|
+
|
|
120
|
+
When you author or review a "keep working until X" skill:
|
|
121
|
+
|
|
122
|
+
1. **Enumerate every stop/accept gate.** Not just the headline one —
|
|
123
|
+
the early-exit on convergence, the PATIENCE bail-out, the
|
|
124
|
+
per-iteration "is this round done?" check, the final "are we
|
|
125
|
+
finished?" check. Write them down.
|
|
126
|
+
|
|
127
|
+
2. **Classify each gate A or B** using the dividing question. If it's
|
|
128
|
+
compound, split it (above) and classify the parts.
|
|
129
|
+
|
|
130
|
+
3. **For every Type-A gate:** Claude may self-judge. Prefer an
|
|
131
|
+
*external* check where one exists (read an exit code, stat a file,
|
|
132
|
+
read a counter) over an LLM "I believe it finished" — Type-A is
|
|
133
|
+
exactly the place where a cheap deterministic check beats a vibe.
|
|
134
|
+
|
|
135
|
+
4. **For every Type-B gate:** route it to a cross-model verdict per
|
|
136
|
+
`reviewer-routing.md` (default `mcp__codex__codex` at
|
|
137
|
+
`reasoning_effort: xhigh`; `oracle-pro` on request; `manual` only if
|
|
138
|
+
the routed model is verifiably non-Claude and recorded — otherwise it
|
|
139
|
+
is same-family self-acquittal in disguise). Pass file paths, not
|
|
140
|
+
summaries (`reviewer-independence.md`).
|
|
141
|
+
The loop **continues or stops on the reviewer's verdict**, not on
|
|
142
|
+
Claude's reading of it. Save the verdict as an artifact
|
|
143
|
+
(`integration-contract.md` §3) so a third party can confirm the
|
|
144
|
+
acquittal was external.
|
|
145
|
+
|
|
146
|
+
5. **State the provenance in the SKILL.** One line: "STOP gate = Type-B,
|
|
147
|
+
routed to codex." A reviewer of the SKILL should be able to find,
|
|
148
|
+
for each terminating condition, which model family signs off.
|
|
149
|
+
|
|
150
|
+
6. **Refuse the anti-pattern:** a loop whose continue/stop decision reads
|
|
151
|
+
an LLM-produced quality verdict that the **same** model family
|
|
152
|
+
(Claude) produced. That is self-acquittal regardless of how the
|
|
153
|
+
prompt is phrased.
|
|
154
|
+
|
|
155
|
+
Rule of thumb: **if removing the cross-model reviewer would still let
|
|
156
|
+
the loop decide to stop, the loop is self-acquitting.** A safe Type-B
|
|
157
|
+
loop is *designed* (by this contract) so that removing the external
|
|
158
|
+
family's verdict leaves it unable to terminate-accept — a design rule
|
|
159
|
+
the skill author enforces, not an automatic structural property.
|
|
160
|
+
|
|
161
|
+
## ARIS loops mapped to the taxonomy
|
|
162
|
+
|
|
163
|
+
The codebase **already** follows this rule. This section makes the
|
|
164
|
+
implicit pattern explicit and operational for the next loop someone
|
|
165
|
+
writes.
|
|
166
|
+
|
|
167
|
+
| Loop | Headline stop gate | Type | Who acquits | Status |
|
|
168
|
+
|---|---|---|---|---|
|
|
169
|
+
| `/dse-loop` | objective metric converged / TIMEOUT / PATIENCE | A | benchmark harness emits the number; Claude reads & compares to budget | ✅ safe same-model |
|
|
170
|
+
| `/experiment-bridge` auto-debug | "did it run / did it converge" (exit 0, no NaN, training started) | A | exit codes, log parse | ✅ safe same-model |
|
|
171
|
+
| `/run-experiment`, `/experiment-queue` retry | job finished / OOM-retry exhausted / N jobs done | A | scheduler + exit codes | ✅ safe same-model |
|
|
172
|
+
| `/auto-review-loop` | score >= 6 AND verdict "ready", per round | B | **codex** assigns score & verdict | ✅ already cross-model |
|
|
173
|
+
| `/auto-paper-improvement-loop` | "review satisfied" (2 rounds) | B | **codex (GPT xhigh)** review | ✅ already cross-model |
|
|
174
|
+
| `/result-to-claim` | `claim_supported ∈ {yes,partial,no}` + `integrity_status` | B | **codex** judges results vs claims | ✅ cross-model |
|
|
175
|
+
| `/kill-argument` | rejection memo → defense, residual issues | B | two fresh **codex** threads | ✅ cross-model |
|
|
176
|
+
| `/proof-checker` | each gap closed, per round | B | **codex** re-reviews each round | ✅ cross-model |
|
|
177
|
+
| `/experiment-audit` | integrity verdict (fake GT, normalization fraud) | B | **codex** audits the eval code | ✅ cross-model |
|
|
178
|
+
| `/paper-claim-audit` | every number matches result files | B | fresh zero-context **cross-model** reviewer | ✅ cross-model |
|
|
179
|
+
| `/citation-audit` | every entry real & in-context | B | fresh **cross-model** reviewer | ✅ cross-model |
|
|
180
|
+
| `/paper-writing` Phase 6 (submission) | `verify_paper_audits.sh` exit 0 | A (gate) **wrapping** B (the audits) | external verifier reads cross-model JSON | ✅ A-gate over B-verdicts |
|
|
181
|
+
|
|
182
|
+
> 📌 The `/auto-review-loop` row reflects the skill's stop logic: `score >= 6`
|
|
183
|
+
> AND verdict contains "ready"/"almost", evaluated each round. (Its `Constants`
|
|
184
|
+
> block previously stated this with `OR` and a stale verdict vocabulary — an
|
|
185
|
+
> internal inconsistency now reconciled to the `AND` form the Phase-E stop
|
|
186
|
+
> check actually uses, in `auto-review-loop` and its `-llm`/`-minimax`
|
|
187
|
+
> siblings.) The acquittal is **codex's** score+verdict, so the Type-B
|
|
188
|
+
> classification is unchanged.
|
|
189
|
+
|
|
190
|
+
Two patterns to notice:
|
|
191
|
+
|
|
192
|
+
- **The execution loops (dse, auto-debug, queue) are Type-A all the way
|
|
193
|
+
down** — "did it run / did it converge" is a fact a harness reports.
|
|
194
|
+
They are *correctly* allowed to self-acquit, because there is nothing
|
|
195
|
+
of merit being judged: a converged number from a real simulator is an
|
|
196
|
+
observation, not an opinion. (The moment someone adds *"...and the
|
|
197
|
+
result is good enough to claim"* to a dse stop condition, that clause
|
|
198
|
+
is Type-B and must route out — see the dse caveat below.)
|
|
199
|
+
|
|
200
|
+
- **Every quality/correctness loop already routes its acquittal to
|
|
201
|
+
codex.** Nothing here is new behavior; the doc names the rule the
|
|
202
|
+
codebase converged on so the next author doesn't have to rediscover it
|
|
203
|
+
by getting reviewed.
|
|
204
|
+
|
|
205
|
+
### The dse-loop caveat (objective ≠ acceptance)
|
|
206
|
+
|
|
207
|
+
`/dse-loop` optimizes a metric the benchmark *itself* produces (cycles,
|
|
208
|
+
area, coverage). "Config B beats config A on the harness's own number"
|
|
209
|
+
is Type-A — a parser, not Claude, owns it. But two adjacent judgments are
|
|
210
|
+
Type-B and must NOT be folded into the loop's self-acquittal:
|
|
211
|
+
|
|
212
|
+
- "this config is **good enough to ship/publish**" — sufficiency verdict.
|
|
213
|
+
- "the benchmark/metric **is the right thing to optimize** / the result
|
|
214
|
+
**generalizes**" — correctness-of-framing verdict.
|
|
215
|
+
|
|
216
|
+
So dse may self-terminate on *"best config found within budget"* (A), but
|
|
217
|
+
the claim *"and this is a publishable result"* leaves the loop and goes
|
|
218
|
+
through `/result-to-claim` (B). Driving the search is in-family; acquitting
|
|
219
|
+
the science is not.
|
|
220
|
+
|
|
221
|
+
## Tie to fan-out: breadth is same-family; the jury is not
|
|
222
|
+
|
|
223
|
+
`fan-out-pattern.md` describes skill-layer fan-out — spawning multiple
|
|
224
|
+
agents for breadth (parallel search branches, per-section drafting,
|
|
225
|
+
per-entry citation checks). Fan-out interacts with this contract in
|
|
226
|
+
exactly one dangerous way:
|
|
227
|
+
|
|
228
|
+
**Same-family breadth is fine for Type-A coverage. It is NEVER a Type-B
|
|
229
|
+
jury.**
|
|
230
|
+
|
|
231
|
+
- ✅ Ten Claude branches each *attempting* a different search query, then
|
|
232
|
+
unioning hits — Type-A coverage (did we look broadly?). Self-judged
|
|
233
|
+
fine.
|
|
234
|
+
- ✅ N Claudes each drafting a section, a Type-A "all sections drafted"
|
|
235
|
+
completion check.
|
|
236
|
+
- ❌ N Claude reviewers each scoring the paper, then taking the
|
|
237
|
+
**majority/average as the accept verdict.** This *feels* like a jury
|
|
238
|
+
— independent voters! — but it is correlated same-family blindness
|
|
239
|
+
wearing a jury costume. N agreeing Claudes share the same training
|
|
240
|
+
priors and the same blind spots; their agreement is evidence of
|
|
241
|
+
shared bias, not of correctness. A Type-B verdict needs a **different
|
|
242
|
+
family**, not a *bigger N of the same one*.
|
|
243
|
+
|
|
244
|
+
> **Known failure mode:** "We ran the review 5× and all 5 said accept,
|
|
245
|
+
> so it's robust." Five draws from one distribution is one opinion with
|
|
246
|
+
> error bars, not five opinions. The cross-model invariant is about
|
|
247
|
+
> *family diversity*, not *sample count*. Fan-out scales breadth and
|
|
248
|
+
> Type-A coverage; it can never substitute for the one cross-family
|
|
249
|
+
> acquittal a Type-B gate requires.
|
|
250
|
+
|
|
251
|
+
Fan-out and this contract compose cleanly: fan-out (same family) does
|
|
252
|
+
the broad *driving*; the loop always funnels into the identical
|
|
253
|
+
cross-model *acquittal* at the Type-B gate. Breadth degrades gracefully
|
|
254
|
+
across runtimes (fewer parallel agents = slower, not unsafe); the
|
|
255
|
+
acquittal does not degrade — it is always the cross-family verdict, or
|
|
256
|
+
the loop is unsafe.
|
|
257
|
+
|
|
258
|
+
## Required components (for a loop to claim same-family-safe)
|
|
259
|
+
|
|
260
|
+
A loop is same-family-safe iff **all** hold:
|
|
261
|
+
|
|
262
|
+
1. **Every stop/accept gate is classified** A or B in the SKILL (compound
|
|
263
|
+
gates split).
|
|
264
|
+
2. **Every Type-B gate routes to a cross-model verdict** per
|
|
265
|
+
`reviewer-routing.md`; the loop's continue/stop reads *that* verdict,
|
|
266
|
+
not a Claude re-judgment of it.
|
|
267
|
+
3. **The cross-model verdict is an artifact** (`integration-contract.md`
|
|
268
|
+
§3) — a JSON/file a third party can inspect to confirm the acquittal
|
|
269
|
+
was external.
|
|
270
|
+
4. **No same-family majority is treated as a Type-B jury** — fan-out
|
|
271
|
+
breadth never substitutes for cross-family acquittal.
|
|
272
|
+
5. **Type-A self-judgment prefers an external check** (exit code, stat,
|
|
273
|
+
counter) over an LLM "I think it's done" wherever one exists.
|
|
274
|
+
|
|
275
|
+
If any fails, the loop can self-acquit and is **not** same-family-safe —
|
|
276
|
+
regardless of how many rounds it runs or how confident it sounds.
|
|
277
|
+
|
|
278
|
+
## Anti-patterns to refuse in review
|
|
279
|
+
|
|
280
|
+
- **"The loop decides when it's good enough."** Good-enough is Type-B;
|
|
281
|
+
the loop may decide when it's *done running*, not when it's *good*.
|
|
282
|
+
- **"We re-review until it passes."** Fine — but *who* says it passed? If
|
|
283
|
+
the answer is Claude, the loop is self-acquitting.
|
|
284
|
+
- **"N agreeing agents = consensus."** Same-family agreement is correlated
|
|
285
|
+
blindness, not a jury (see fan-out section).
|
|
286
|
+
- **"It converged, so it's correct."** Convergence is Type-A
|
|
287
|
+
(it stopped moving); correctness/sufficiency is Type-B.
|
|
288
|
+
- **"Score >= 6, so stop."** Only safe if a *different family* assigned
|
|
289
|
+
the score. Claude scoring Claude and stopping at 6 is self-acquittal.
|
|
290
|
+
- **`/loop` wrapping an internal semantic loop.** External cadence
|
|
291
|
+
(`/loop`) is additive only for external-world waits (GPU done?
|
|
292
|
+
overnight heartbeat?). Wrapping ARIS's internal semantic loops with a
|
|
293
|
+
timer breaks `threadId` continuity and re-runs Type-B verdicts on a
|
|
294
|
+
clock instead of on the reviewer's turn — noise at best, a corrupted
|
|
295
|
+
acquittal at worst. Keep external cadence outside the acceptance gate.
|
|
296
|
+
|
|
297
|
+
## Epistemic status of a PASS
|
|
298
|
+
|
|
299
|
+
A cross-model PASS is a **heterogeneous second opinion**, not external ground truth. Its
|
|
300
|
+
value is specific and bounded: a reviewer from a different model family breaks *correlated*
|
|
301
|
+
blind spots — the executor's own failure modes it cannot see in itself — so a PASS means
|
|
302
|
+
"a differently-built model, reading the artifact cold, did not find the flaw the author
|
|
303
|
+
would miss." It does **not** mean the work is correct, novel, publishable, or that a venue
|
|
304
|
+
will accept it. Same-family review (Claude judging Claude) does not even clear that bar,
|
|
305
|
+
which is why the jury must be cross-family.
|
|
306
|
+
|
|
307
|
+
Treat a PASS as the strongest *automatable* heterogeneous quality check this framework has, then keep the human in the loop for
|
|
308
|
+
what no in-framework verdict can supply: updated literature, venue taste, and ground truth.
|
|
309
|
+
A green gate lowers risk; it does not transfer accountability.
|
|
310
|
+
|
|
311
|
+
## See Also
|
|
312
|
+
|
|
313
|
+
- `reviewer-independence.md` — the single-shot form: executor never
|
|
314
|
+
filters the reviewer's inputs. Type-B gates inherit this in full.
|
|
315
|
+
- `experiment-integrity.md` — the experiment form: the model that writes
|
|
316
|
+
experiment code must not judge its integrity. `/experiment-audit`'s
|
|
317
|
+
Type-B verdict is the loop instance of this rule.
|
|
318
|
+
- `reviewer-routing.md` — where Type-B gates send their verdict (codex
|
|
319
|
+
default, oracle-pro on request, manual only with a verified non-Claude
|
|
320
|
+
target).
|
|
321
|
+
- `fan-out-pattern.md` — breadth via same-family spawn; this doc's
|
|
322
|
+
fan-out section is the guardrail that keeps breadth out of the jury box.
|
|
323
|
+
- `integration-contract.md` §3 — the cross-model verdict must leave an
|
|
324
|
+
inspectable artifact.
|
|
@@ -0,0 +1,248 @@
|
|
|
1
|
+
# Assurance Contract
|
|
2
|
+
|
|
3
|
+
ARIS audits emit machine-readable verdicts. The `assurance` axis decides whether
|
|
4
|
+
those verdicts are advisory (draft mode) or load-bearing gates (submission mode).
|
|
5
|
+
This contract is referenced by `paper-writing`, `paper-claim-audit`, `citation-audit`,
|
|
6
|
+
`proof-checker`, and the external verifier (canonical name `verify_paper_audits.sh`;
|
|
7
|
+
callers resolve the actual path via `integration-contract.md` §2).
|
|
8
|
+
|
|
9
|
+
## Why a separate axis from `effort`
|
|
10
|
+
|
|
11
|
+
Historically `effort` (lite/balanced/max/beast) was conflated with audit strictness.
|
|
12
|
+
The result: `effort: beast` did not guarantee mandatory audits ran — phases were
|
|
13
|
+
gated by content detectors (e.g. `if \begin{theorem} exists`) and could silently
|
|
14
|
+
skip. A user reported `effort: beast` produced a "draft-quality" paper with all
|
|
15
|
+
three submission-gate audits skipped.
|
|
16
|
+
|
|
17
|
+
The fix is to split the concerns:
|
|
18
|
+
|
|
19
|
+
| Axis | Controls | Default |
|
|
20
|
+
|------|----------|---------|
|
|
21
|
+
| `effort` | depth/cost (papers, rounds, ideation) | `balanced` |
|
|
22
|
+
| `assurance` | audit strictness — silent-skip-allowed vs verdict-required | derived from `effort` (see mapping) |
|
|
23
|
+
|
|
24
|
+
Override either independently: `— effort: balanced, assurance: submission` is
|
|
25
|
+
legal and means "normal depth, but every audit must emit a verdict before
|
|
26
|
+
finalization."
|
|
27
|
+
|
|
28
|
+
## Assurance Levels
|
|
29
|
+
|
|
30
|
+
### `draft` — current behavior, no breakage
|
|
31
|
+
- Audits run only if their content detector matches.
|
|
32
|
+
- Silent skip allowed.
|
|
33
|
+
- `paper-writing` Phase 6 produces a final report regardless.
|
|
34
|
+
- For: rapid iteration, exploratory drafts, early-stage research.
|
|
35
|
+
|
|
36
|
+
### `submission` — load-bearing audits
|
|
37
|
+
- All mandatory audits **must** emit a verdict (one of the 6 below).
|
|
38
|
+
- Silent skip is **forbidden**.
|
|
39
|
+
- `paper-writing` Phase 6 invokes `verify_paper_audits.sh` (resolved per
|
|
40
|
+
`integration-contract.md` §2); non-zero exit blocks Final Report.
|
|
41
|
+
- The Final Report tags itself `submission-ready: yes/no` based on verifier output.
|
|
42
|
+
- For: conference / journal submission, anything you'd put your name on.
|
|
43
|
+
|
|
44
|
+
## Default Mapping (derived if `assurance` not given)
|
|
45
|
+
|
|
46
|
+
| `effort` | implied `assurance` |
|
|
47
|
+
|----------|---------------------|
|
|
48
|
+
| `lite` | `draft` |
|
|
49
|
+
| `balanced` | `draft` |
|
|
50
|
+
| `max` | `submission` |
|
|
51
|
+
| `beast` | `submission` |
|
|
52
|
+
|
|
53
|
+
This means a user passing only `— effort: beast` automatically gets full audit
|
|
54
|
+
enforcement — matching their intent ("turn everything up"). Users wanting
|
|
55
|
+
strict audits at lower depth pass `— assurance: submission` explicitly.
|
|
56
|
+
|
|
57
|
+
## Verdict State Machine
|
|
58
|
+
|
|
59
|
+
Every mandatory audit must emit exactly one of these — never silent skip:
|
|
60
|
+
|
|
61
|
+
| Verdict | Meaning | Audit ran? | Submission-blocking? |
|
|
62
|
+
|---------|---------|-----------|----------------------|
|
|
63
|
+
| `PASS` | All checks passed | Yes | No |
|
|
64
|
+
| `WARN` | Issues found, none disqualifying | Yes | No |
|
|
65
|
+
| `FAIL` | Disqualifying issues found | Yes | **Yes** |
|
|
66
|
+
| `NOT_APPLICABLE` | Detector negative; nothing to audit (e.g., no theorems in paper, no `\cite`s, no numeric claims) | Audit phase ran, child audit invocation may have been skipped | No |
|
|
67
|
+
| `BLOCKED` | Audit should apply but prerequisites are missing or unsupported (e.g., paper has numeric claims but no `results/` directory; paper cites references but `.bib` missing) | Could not complete | **Yes** |
|
|
68
|
+
| `ERROR` | Audit invocation failed (network, timeout, malformed reviewer output) | Attempted but errored | **Yes** at submission |
|
|
69
|
+
|
|
70
|
+
### Why `NOT_APPLICABLE` is not the same as `SKIP`
|
|
71
|
+
|
|
72
|
+
`NOT_APPLICABLE` means **the audit phase ran**, the detector returned negative,
|
|
73
|
+
and a verdict artifact was written documenting "we checked, there's nothing to
|
|
74
|
+
verify." This is verifiable from outside the LLM — the artifact file exists.
|
|
75
|
+
|
|
76
|
+
A silent skip leaves no record. There's no way to distinguish "we checked and
|
|
77
|
+
there was nothing" from "we forgot." This contract makes that distinction
|
|
78
|
+
mandatory.
|
|
79
|
+
|
|
80
|
+
### Why `BLOCKED` is more dangerous than `NOT_APPLICABLE`
|
|
81
|
+
|
|
82
|
+
`BLOCKED` means the audit *should* have run but cannot. Example: a paper claims
|
|
83
|
+
`accuracy = 89.2%` but has no `results/` directory to verify against. That's not
|
|
84
|
+
"nothing to audit" — that's "we cannot verify a load-bearing claim." Treating
|
|
85
|
+
this as `SKIP` masks the danger; `BLOCKED` surfaces it and blocks submission.
|
|
86
|
+
|
|
87
|
+
## Required Audit Artifact Schema
|
|
88
|
+
|
|
89
|
+
Every mandatory audit must write a JSON artifact (and may also write a
|
|
90
|
+
human-readable Markdown sibling). The JSON must contain at minimum:
|
|
91
|
+
|
|
92
|
+
```json
|
|
93
|
+
{
|
|
94
|
+
"audit_skill": "paper-claim-audit", // citation-audit, proof-checker, etc.
|
|
95
|
+
"verdict": "PASS", // one of the 6 above
|
|
96
|
+
"reason_code": "all_numbers_match", // skill-specific short string
|
|
97
|
+
"summary": "Verified 23 numeric claims against 4 result files; no mismatches.",
|
|
98
|
+
"audited_input_hashes": {
|
|
99
|
+
"main.tex": "sha256:a3f8...",
|
|
100
|
+
"sections/5.evidence.tex": "sha256:b2d1...",
|
|
101
|
+
"/Users/me/project/results/run_2026_04_19.json": "sha256:c9e4..."
|
|
102
|
+
},
|
|
103
|
+
"trace_path": ".aris/traces/paper-claim-audit/2026-04-21_run01/",
|
|
104
|
+
"thread_id": "019dae73-fc12-4ab8-...",
|
|
105
|
+
"executor_model": "claude-opus-4-8",
|
|
106
|
+
"executor_family": "anthropic",
|
|
107
|
+
"reviewer_model": "gpt-5.6-sol",
|
|
108
|
+
"reviewer_family": "openai",
|
|
109
|
+
"review_independence": "cross-family",
|
|
110
|
+
"acceptance_status": "accepted",
|
|
111
|
+
"reviewer_reasoning": "xhigh",
|
|
112
|
+
"generated_at": "2026-04-21T14:23:01Z",
|
|
113
|
+
"details": {
|
|
114
|
+
// skill-specific structured data
|
|
115
|
+
}
|
|
116
|
+
}
|
|
117
|
+
```
|
|
118
|
+
|
|
119
|
+
Field semantics:
|
|
120
|
+
|
|
121
|
+
- **`audited_input_hashes`** — SHA256 of every file the audit consumed.
|
|
122
|
+
- Keys are **paths relative to the paper directory** (the argument
|
|
123
|
+
passed to `verify_paper_audits.sh`) for files inside it, or
|
|
124
|
+
**absolute paths** for files outside it (e.g. `../results/run.json`
|
|
125
|
+
is legal but `/Users/me/project/results/run.json` is more portable).
|
|
126
|
+
Do NOT prefix in-paper files with `paper/` — the verifier already
|
|
127
|
+
resolves relative to the paper dir and `paper/paper/main.tex` will
|
|
128
|
+
false-fail. The verifier rehashes the current files and flags `STALE`
|
|
129
|
+
if any hash changed since the audit ran. (User edited `main.tex`
|
|
130
|
+
after running `paper-claim-audit`? The next verifier run will catch it.)
|
|
131
|
+
- **`trace_path`** — directory containing the full reviewer prompt + response
|
|
132
|
+
pair, per `review-tracing.md`. Required for mandatory audits — not optional.
|
|
133
|
+
- **`thread_id` or `agent_id`** — durable reviewer handle. MCP routes use a
|
|
134
|
+
thread ID; Codex `spawn_agent` routes use an agent ID. At least one is required.
|
|
135
|
+
- **`reviewer_model`** + **`reviewer_reasoning`** — proves cross-family review
|
|
136
|
+
invariant was honored.
|
|
137
|
+
- **`review_independence`** — `same-family`, `cross-family`, or `deterministic`.
|
|
138
|
+
`deterministic` is valid ONLY for what a process can actually decide
|
|
139
|
+
(compilation, schema validity, hash freshness, test suites) — the audit
|
|
140
|
+
aggregator REJECTS a deterministic label on the four semantic paper audits
|
|
141
|
+
(proof / claims / citations / attack), which only a cross-family model
|
|
142
|
+
review can accept.
|
|
143
|
+
**`acceptance_status`** is `provisional` for same-family review and
|
|
144
|
+
`accepted` for cross-family/deterministic review; neither field rewrites the
|
|
145
|
+
substantive verdict.
|
|
146
|
+
- **`generated_at`** — UTC ISO-8601 timestamp.
|
|
147
|
+
|
|
148
|
+
## Verifier Contract
|
|
149
|
+
|
|
150
|
+
`verify_paper_audits.sh <paper-dir>` (canonical name; resolved per
|
|
151
|
+
`integration-contract.md` §2) is the single source of truth for
|
|
152
|
+
"are mandatory audits complete and current?" It must:
|
|
153
|
+
|
|
154
|
+
1. Locate the paper-writing manifest (which mandatory audits applied this run).
|
|
155
|
+
2. For each, check artifact JSON exists at expected path.
|
|
156
|
+
3. Validate artifact JSON against required-fields schema (above).
|
|
157
|
+
4. Verify `verdict` is one of the 6 allowed values.
|
|
158
|
+
5. Recompute SHA256 of every file in `audited_input_hashes`; flag `STALE` if any
|
|
159
|
+
mismatches.
|
|
160
|
+
6. Verify `trace_path` exists and is non-empty.
|
|
161
|
+
7. Output a structured JSON report and exit 0 (all green) or 1 (any FAIL /
|
|
162
|
+
BLOCKED / ERROR / STALE / missing artifact).
|
|
163
|
+
|
|
164
|
+
The report also emits `overall_assurance`: `blocked` for a blocking condition,
|
|
165
|
+
`provisional` when green artifacts include same-family or legacy-unspecified
|
|
166
|
+
review, and `accepted` only when all green artifacts are cross-family or
|
|
167
|
+
deterministic. Provisional remains exit 0 but must never be presented as
|
|
168
|
+
submission-ready yes.
|
|
169
|
+
|
|
170
|
+
Phase 6 of `paper-writing` invokes the verifier; at `assurance: submission`,
|
|
171
|
+
non-zero exit blocks Final Report generation.
|
|
172
|
+
|
|
173
|
+
## Subskill Contract: "Always Emit, Never Block"
|
|
174
|
+
|
|
175
|
+
Child audit skills (`paper-claim-audit`, `citation-audit`, `proof-checker`)
|
|
176
|
+
follow this contract:
|
|
177
|
+
|
|
178
|
+
- **Always emit a verdict artifact**, even on detector-negative or error paths.
|
|
179
|
+
- **Never block** the parent's flow themselves — they only emit verdicts.
|
|
180
|
+
- **The parent skill** (`paper-writing` Phase 6 + verifier) decides whether a
|
|
181
|
+
given verdict blocks finalization. This decision lives in *one* place
|
|
182
|
+
(`assurance` axis + verifier), not duplicated across child skills.
|
|
183
|
+
|
|
184
|
+
Earlier wording in `paper-claim-audit` and `citation-audit` (e.g., "audit is
|
|
185
|
+
advisory, never blocking") referred to this division of labor — but conflicted
|
|
186
|
+
with `paper-writing`'s declaration that they were "mandatory submission gates."
|
|
187
|
+
This contract resolves the conflict: child = always emit; parent = decides
|
|
188
|
+
blocking based on assurance level.
|
|
189
|
+
|
|
190
|
+
## Examples
|
|
191
|
+
|
|
192
|
+
### Theory paper, beast effort
|
|
193
|
+
```
|
|
194
|
+
— effort: beast (implies assurance: submission)
|
|
195
|
+
```
|
|
196
|
+
- `proof-checker` runs, audits theorems → `PASS` or `WARN` or `FAIL`
|
|
197
|
+
- `paper-claim-audit` runs, finds numbers → `PASS`
|
|
198
|
+
- `citation-audit` runs, audits refs → `PASS`
|
|
199
|
+
- Verifier: all green
|
|
200
|
+
- Final Report: `submission-ready: yes`
|
|
201
|
+
|
|
202
|
+
### Position paper (no theorems, no numbers, no experiments), beast effort
|
|
203
|
+
```
|
|
204
|
+
— effort: beast (implies assurance: submission)
|
|
205
|
+
```
|
|
206
|
+
- `proof-checker` invoked → no theorems found → emits `NOT_APPLICABLE`
|
|
207
|
+
- `paper-claim-audit` invoked → no numeric claims → emits `NOT_APPLICABLE`
|
|
208
|
+
- `citation-audit` invoked → audits refs → `PASS`
|
|
209
|
+
- Verifier: all green (NOT_APPLICABLE is not blocking)
|
|
210
|
+
- Final Report: `submission-ready: yes` with note "no theorems / no numeric claims to audit"
|
|
211
|
+
|
|
212
|
+
### Empirical paper missing raw results, beast effort
|
|
213
|
+
```
|
|
214
|
+
— effort: beast
|
|
215
|
+
```
|
|
216
|
+
- `proof-checker` → `NOT_APPLICABLE`
|
|
217
|
+
- `paper-claim-audit` invoked → finds claims like `accuracy = 89.2%` but
|
|
218
|
+
`results/` is empty → emits `BLOCKED` with reason_code `no_raw_evidence`
|
|
219
|
+
- `citation-audit` → `PASS`
|
|
220
|
+
- Verifier: exit 1 (BLOCKED is submission-blocking)
|
|
221
|
+
- Final Report: **refuses to finalize**; surfaces "Mandatory audit BLOCKED:
|
|
222
|
+
paper-claim-audit cannot verify numeric claims — no raw result files found.
|
|
223
|
+
Add results/ or downgrade to `— assurance: draft`."
|
|
224
|
+
|
|
225
|
+
### Stale audit (user edited paper after running audits)
|
|
226
|
+
- User runs `/paper-writing` at beast → all audits PASS, files written
|
|
227
|
+
- User edits `sec/5.evidence.tex` to change a number
|
|
228
|
+
- User reruns the verifier (or re-finalizes)
|
|
229
|
+
- Verifier rehashes → `audited_input_hashes` mismatch → `STALE` flag → exit 1
|
|
230
|
+
- Final Report: refuses; instructs user to rerun `paper-claim-audit` and
|
|
231
|
+
`citation-audit` before re-finalizing.
|
|
232
|
+
|
|
233
|
+
## Backward Compatibility
|
|
234
|
+
|
|
235
|
+
- Users on `effort: balanced` (default) get `assurance: draft` — **identical
|
|
236
|
+
current behavior, no breakage**.
|
|
237
|
+
- Users explicitly using `effort: max` or `effort: beast` automatically get
|
|
238
|
+
`assurance: submission` — matching their intent.
|
|
239
|
+
- Users wanting the old "beast = depth only, no audit enforcement" can pass
|
|
240
|
+
`— effort: beast, assurance: draft` (explicit override). This combination is
|
|
241
|
+
legal but discouraged for actual submissions.
|
|
242
|
+
|
|
243
|
+
## See Also
|
|
244
|
+
|
|
245
|
+
- `effort-contract.md` — depth/cost axis (separate concern)
|
|
246
|
+
- `review-tracing.md` — trace artifact protocol (referenced by `trace_path`)
|
|
247
|
+
- `reviewer-independence.md` — cross-model review invariant
|
|
248
|
+
- `tools/verify_paper_audits.sh` — external verifier implementation
|
|
@@ -0,0 +1,78 @@
|
|
|
1
|
+
# Capture Anti-patterns (anti-self-poisoning)
|
|
2
|
+
|
|
3
|
+
When ARIS captures *durable* knowledge — a research-wiki idea / claim / experiment
|
|
4
|
+
node, a `/meta-optimize` SKILL.md proposal — it must not store **operational
|
|
5
|
+
noise** that later hardens into a self-cited falsehood. This is the failure mode
|
|
6
|
+
Hermes's self-improvement loop hit and patched with a hand-written "Do NOT
|
|
7
|
+
capture" list: negative tool-capability claims that *"harden into refusals the
|
|
8
|
+
agent cites against itself for months after the actual problem was fixed."*
|
|
9
|
+
ARIS's research-wiki "failed ideas → anti-repeat memory" is the GOOD inverse (a
|
|
10
|
+
class-level *research* finding worth remembering); this is the blocklist for the
|
|
11
|
+
BAD kind (transient *operational* state masquerading as a durable fact).
|
|
12
|
+
|
|
13
|
+
## The four anti-patterns — do NOT capture
|
|
14
|
+
|
|
15
|
+
| class | example (do NOT store) | store INSTEAD |
|
|
16
|
+
|-------|------------------------|---------------|
|
|
17
|
+
| **env-specific failure** | "pip failed: No module named torch", "command not found" | the fix / the missing dependency / the correct config |
|
|
18
|
+
| **transient error** | "got a 429", "CUDA OOM", "connection refused" | nothing — it self-resolves; or the retry/backoff that worked |
|
|
19
|
+
| **negative tool-capability claim** | "codex can't handle long files", "gemini is broken", "don't use oracle" | the workaround, or "needs flag X" — never "tool can't do Y" |
|
|
20
|
+
| **single-instance narrative** | "in run 47 the loss spiked at step 300" | only the *class-level* rule it implies, if any ("LR > 3e-4 diverges on this model") |
|
|
21
|
+
|
|
22
|
+
The cardinal rule: **store *how to fix* / *what config is missing* / *the
|
|
23
|
+
workaround*, never *"X can't do Y"*.** A negative capability claim about your own
|
|
24
|
+
tooling is the most dangerous capture — it gets loaded into every future session
|
|
25
|
+
and the agent cites it against itself long after the real cause is gone.
|
|
26
|
+
|
|
27
|
+
## Mechanical vs judgment
|
|
28
|
+
|
|
29
|
+
- **Mechanical** (deterministic, `tools/capture_filter.py`): the unambiguous
|
|
30
|
+
classes — raw error output (`No module named`, `command not found`,
|
|
31
|
+
`ModuleNotFoundError`, `Permission denied`), transient errors (rate-limit / OOM
|
|
32
|
+
/ network), and explicitly-broken-tool phrasing anchored on ARIS infrastructure
|
|
33
|
+
nouns (codex / gemini / oracle / the reviewer / the MCP / the CLI …).
|
|
34
|
+
- **Judgment** (this doc): the single-instance-narrative class, and any operational
|
|
35
|
+
note dressed up as a finding. The agent applies this when deciding what to persist.
|
|
36
|
+
|
|
37
|
+
The mechanical filter is **deliberately conservative**: it does NOT flag
|
|
38
|
+
legitimate *research* findings about a model/method ("the model can't generalize
|
|
39
|
+
to OOD", "our method fails on long sequences") — it targets ARIS's own *tooling*
|
|
40
|
+
being declared broken, and raw error text. False negatives are fine (the jury
|
|
41
|
+
still judges); a flagged note just goes to manual review / gets rewritten.
|
|
42
|
+
|
|
43
|
+
## The asymmetry (acceptance-gate.md)
|
|
44
|
+
|
|
45
|
+
This filter may **REJECT a capture same-model** — it is a mechanical safety screen,
|
|
46
|
+
low risk, and same-model is always allowed to *reject*. But anything that **passes**
|
|
47
|
+
the filter and would become a **load-bearing** skill/claim still goes to the
|
|
48
|
+
**cross-model jury** before it is trusted. Same-model is fine to reject; it is
|
|
49
|
+
never enough to *accept* into the load-bearing set.
|
|
50
|
+
|
|
51
|
+
## Helper
|
|
52
|
+
|
|
53
|
+
```
|
|
54
|
+
from capture_filter import screen, reason_detail
|
|
55
|
+
screen(text) # -> [reason, ...] ([] = clean); reason ∈ {env_failure, transient_error, negative_tool_claim}
|
|
56
|
+
```
|
|
57
|
+
```
|
|
58
|
+
python3 tools/capture_filter.py <file|-> # exit 1 + reasons if anti-pattern found
|
|
59
|
+
```
|
|
60
|
+
|
|
61
|
+
## Where ARIS uses it
|
|
62
|
+
- **`/research-wiki`** (and `/idea-creator` Phase-3 annotations): screen an
|
|
63
|
+
idea/claim/experiment note before persisting it; if flagged, rewrite to the
|
|
64
|
+
fix or drop it — don't let operational noise become a durable node.
|
|
65
|
+
- **`/meta-optimize`**: screen the rationale of a proposed SKILL.md change; never
|
|
66
|
+
propose a change that encodes a negative tool-capability claim or a one-off
|
|
67
|
+
failure as a durable rule.
|
|
68
|
+
|
|
69
|
+
## Cross-references
|
|
70
|
+
- `acceptance-gate.md` — the reject/accept asymmetry: same-model may reject, only
|
|
71
|
+
cross-model may accept into the load-bearing set.
|
|
72
|
+
- `evidence-precheck.md` / `injection-hygiene.md` — sibling deterministic
|
|
73
|
+
pre-gates feeding the cross-model jury.
|
|
74
|
+
|
|
75
|
+
> Anti-pattern taxonomy adapted from NousResearch/hermes-agent's background-review
|
|
76
|
+
> "Do NOT capture" list (MIT). ARIS's increment: Hermes patches self-poisoning with
|
|
77
|
+
> more self-judged prose; ARIS adds the deterministic screen + the cross-model
|
|
78
|
+
> acceptance gate on anything that survives it.
|