paperlint 2.0.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/.github/dependabot.yml +72 -0
- package/.github/workflows/ci.yml +297 -0
- package/.github/workflows/dependabot-automerge.yml +70 -0
- package/.github/workflows/pr-title.yml +59 -0
- package/.github/workflows/release.yml +54 -0
- package/CLAUDE.md +598 -0
- package/CONTRIBUTING.md +159 -0
- package/LICENSE +21 -0
- package/README.md +240 -0
- package/action.harness.mjs +287 -0
- package/action.mutations.mjs +162 -0
- package/action.yml +138 -0
- package/bin/rpp.mjs +43 -0
- package/dist/action-ref.d.ts +12 -0
- package/dist/action-ref.d.ts.map +1 -0
- package/dist/action-ref.js +16 -0
- package/dist/action-ref.js.map +1 -0
- package/dist/adapters/banal/failure.d.ts +73 -0
- package/dist/adapters/banal/failure.d.ts.map +1 -0
- package/dist/adapters/banal/failure.js +58 -0
- package/dist/adapters/banal/failure.js.map +1 -0
- package/dist/adapters/banal/index.d.ts +17 -0
- package/dist/adapters/banal/index.d.ts.map +1 -0
- package/dist/adapters/banal/index.js +56 -0
- package/dist/adapters/banal/index.js.map +1 -0
- package/dist/adapters/banal/install.d.ts +26 -0
- package/dist/adapters/banal/install.d.ts.map +1 -0
- package/dist/adapters/banal/install.js +15 -0
- package/dist/adapters/banal/install.js.map +1 -0
- package/dist/adapters/banal/invocation.d.ts +48 -0
- package/dist/adapters/banal/invocation.d.ts.map +1 -0
- package/dist/adapters/banal/invocation.js +43 -0
- package/dist/adapters/banal/invocation.js.map +1 -0
- package/dist/adapters/banal/locate.d.ts +50 -0
- package/dist/adapters/banal/locate.d.ts.map +1 -0
- package/dist/adapters/banal/locate.js +34 -0
- package/dist/adapters/banal/locate.js.map +1 -0
- package/dist/adapters/banal/output.d.ts +27 -0
- package/dist/adapters/banal/output.d.ts.map +1 -0
- package/dist/adapters/banal/output.js +112 -0
- package/dist/adapters/banal/output.js.map +1 -0
- package/dist/adapters/banal/pin.d.ts +19 -0
- package/dist/adapters/banal/pin.d.ts.map +1 -0
- package/dist/adapters/banal/pin.js +15 -0
- package/dist/adapters/banal/pin.js.map +1 -0
- package/dist/adapters/banal/probe.d.ts +12 -0
- package/dist/adapters/banal/probe.d.ts.map +1 -0
- package/dist/adapters/banal/probe.js +27 -0
- package/dist/adapters/banal/probe.js.map +1 -0
- package/dist/adapters/banal/run.d.ts +89 -0
- package/dist/adapters/banal/run.d.ts.map +1 -0
- package/dist/adapters/banal/run.js +104 -0
- package/dist/adapters/banal/run.js.map +1 -0
- package/dist/adapters/banal/settings.d.ts +18 -0
- package/dist/adapters/banal/settings.d.ts.map +1 -0
- package/dist/adapters/banal/settings.js +29 -0
- package/dist/adapters/banal/settings.js.map +1 -0
- package/dist/adapters/banal/xml.d.ts +48 -0
- package/dist/adapters/banal/xml.d.ts.map +1 -0
- package/dist/adapters/banal/xml.js +67 -0
- package/dist/adapters/banal/xml.js.map +1 -0
- package/dist/adapters/curl/download.io.d.ts +14 -0
- package/dist/adapters/curl/download.io.d.ts.map +1 -0
- package/dist/adapters/curl/download.io.js +69 -0
- package/dist/adapters/curl/download.io.js.map +1 -0
- package/dist/adapters/curl/index.d.ts +6 -0
- package/dist/adapters/curl/index.d.ts.map +1 -0
- package/dist/adapters/curl/index.js +6 -0
- package/dist/adapters/curl/index.js.map +1 -0
- package/dist/adapters/memory/index.d.ts +43 -0
- package/dist/adapters/memory/index.d.ts.map +1 -0
- package/dist/adapters/memory/index.js +79 -0
- package/dist/adapters/memory/index.js.map +1 -0
- package/dist/adapters/node/files.io.d.ts +3 -0
- package/dist/adapters/node/files.io.d.ts.map +1 -0
- package/dist/adapters/node/files.io.js +31 -0
- package/dist/adapters/node/files.io.js.map +1 -0
- package/dist/adapters/node/host.io.d.ts +3 -0
- package/dist/adapters/node/host.io.d.ts.map +1 -0
- package/dist/adapters/node/host.io.js +14 -0
- package/dist/adapters/node/host.io.js.map +1 -0
- package/dist/adapters/node/index.d.ts +25 -0
- package/dist/adapters/node/index.d.ts.map +1 -0
- package/dist/adapters/node/index.js +14 -0
- package/dist/adapters/node/index.js.map +1 -0
- package/dist/adapters/node/process.io.d.ts +14 -0
- package/dist/adapters/node/process.io.d.ts.map +1 -0
- package/dist/adapters/node/process.io.js +41 -0
- package/dist/adapters/node/process.io.js.map +1 -0
- package/dist/adapters/node/workspace.io.d.ts +4 -0
- package/dist/adapters/node/workspace.io.d.ts.map +1 -0
- package/dist/adapters/node/workspace.io.js +33 -0
- package/dist/adapters/node/workspace.io.js.map +1 -0
- package/dist/adapters/pdfjs/fill.d.ts +42 -0
- package/dist/adapters/pdfjs/fill.d.ts.map +1 -0
- package/dist/adapters/pdfjs/fill.js +91 -0
- package/dist/adapters/pdfjs/fill.js.map +1 -0
- package/dist/build-engine.d.ts +48 -0
- package/dist/build-engine.d.ts.map +1 -0
- package/dist/build-engine.js +148 -0
- package/dist/build-engine.js.map +1 -0
- package/dist/build.d.ts +163 -0
- package/dist/build.d.ts.map +1 -0
- package/dist/build.js +575 -0
- package/dist/build.js.map +1 -0
- package/dist/cli.d.ts +151 -0
- package/dist/cli.d.ts.map +1 -0
- package/dist/cli.js +951 -0
- package/dist/cli.js.map +1 -0
- package/dist/doctor.d.ts +42 -0
- package/dist/doctor.d.ts.map +1 -0
- package/dist/doctor.js +280 -0
- package/dist/doctor.js.map +1 -0
- package/dist/domain/geometry.d.ts +71 -0
- package/dist/domain/geometry.d.ts.map +1 -0
- package/dist/domain/geometry.js +35 -0
- package/dist/domain/geometry.js.map +1 -0
- package/dist/domain/host.d.ts +16 -0
- package/dist/domain/host.d.ts.map +1 -0
- package/dist/domain/host.js +8 -0
- package/dist/domain/host.js.map +1 -0
- package/dist/domain/page-layout.d.ts +34 -0
- package/dist/domain/page-layout.d.ts.map +1 -0
- package/dist/domain/page-layout.js +8 -0
- package/dist/domain/page-layout.js.map +1 -0
- package/dist/domain/paths.d.ts +5 -0
- package/dist/domain/paths.d.ts.map +1 -0
- package/dist/domain/paths.js +2 -0
- package/dist/domain/paths.js.map +1 -0
- package/dist/domain/result.d.ts +23 -0
- package/dist/domain/result.d.ts.map +1 -0
- package/dist/domain/result.js +10 -0
- package/dist/domain/result.js.map +1 -0
- package/dist/domain/sha256.d.ts +7 -0
- package/dist/domain/sha256.d.ts.map +1 -0
- package/dist/domain/sha256.js +14 -0
- package/dist/domain/sha256.js.map +1 -0
- package/dist/domain/text.d.ts +6 -0
- package/dist/domain/text.d.ts.map +1 -0
- package/dist/domain/text.js +7 -0
- package/dist/domain/text.js.map +1 -0
- package/dist/engine.d.ts +93 -0
- package/dist/engine.d.ts.map +1 -0
- package/dist/engine.js +119 -0
- package/dist/engine.js.map +1 -0
- package/dist/exit-code.d.ts +22 -0
- package/dist/exit-code.d.ts.map +1 -0
- package/dist/exit-code.js +10 -0
- package/dist/exit-code.js.map +1 -0
- package/dist/facts-file.d.ts +96 -0
- package/dist/facts-file.d.ts.map +1 -0
- package/dist/facts-file.js +134 -0
- package/dist/facts-file.js.map +1 -0
- package/dist/hooks-settings.d.ts +141 -0
- package/dist/hooks-settings.d.ts.map +1 -0
- package/dist/hooks-settings.js +306 -0
- package/dist/hooks-settings.js.map +1 -0
- package/dist/init.d.ts +201 -0
- package/dist/init.d.ts.map +1 -0
- package/dist/init.js +579 -0
- package/dist/init.js.map +1 -0
- package/dist/latex-log.d.ts +80 -0
- package/dist/latex-log.d.ts.map +1 -0
- package/dist/latex-log.js +187 -0
- package/dist/latex-log.js.map +1 -0
- package/dist/latex-loop.d.ts +129 -0
- package/dist/latex-loop.d.ts.map +1 -0
- package/dist/latex-loop.js +113 -0
- package/dist/latex-loop.js.map +1 -0
- package/dist/link-skills.d.ts +51 -0
- package/dist/link-skills.d.ts.map +1 -0
- package/dist/link-skills.js +199 -0
- package/dist/link-skills.js.map +1 -0
- package/dist/new-paper.d.ts +48 -0
- package/dist/new-paper.d.ts.map +1 -0
- package/dist/new-paper.js +110 -0
- package/dist/new-paper.js.map +1 -0
- package/dist/pdf-facts.d.ts +44 -0
- package/dist/pdf-facts.d.ts.map +1 -0
- package/dist/pdf-facts.js +239 -0
- package/dist/pdf-facts.js.map +1 -0
- package/dist/pdf-geometry.d.ts +170 -0
- package/dist/pdf-geometry.d.ts.map +1 -0
- package/dist/pdf-geometry.js +158 -0
- package/dist/pdf-geometry.js.map +1 -0
- package/dist/ports/download.d.ts +9 -0
- package/dist/ports/download.d.ts.map +1 -0
- package/dist/ports/download.js +2 -0
- package/dist/ports/download.js.map +1 -0
- package/dist/ports/files.d.ts +11 -0
- package/dist/ports/files.d.ts.map +1 -0
- package/dist/ports/files.js +2 -0
- package/dist/ports/files.js.map +1 -0
- package/dist/ports/measure-geometry.d.ts +8 -0
- package/dist/ports/measure-geometry.d.ts.map +1 -0
- package/dist/ports/measure-geometry.js +2 -0
- package/dist/ports/measure-geometry.js.map +1 -0
- package/dist/ports/process.d.ts +45 -0
- package/dist/ports/process.d.ts.map +1 -0
- package/dist/ports/process.js +2 -0
- package/dist/ports/process.js.map +1 -0
- package/dist/ports/tool-installer.d.ts +29 -0
- package/dist/ports/tool-installer.d.ts.map +1 -0
- package/dist/ports/tool-installer.js +2 -0
- package/dist/ports/tool-installer.js.map +1 -0
- package/dist/ports/workspace.d.ts +18 -0
- package/dist/ports/workspace.d.ts.map +1 -0
- package/dist/ports/workspace.js +2 -0
- package/dist/ports/workspace.js.map +1 -0
- package/dist/rules-config.d.ts +34 -0
- package/dist/rules-config.d.ts.map +1 -0
- package/dist/rules-config.js +132 -0
- package/dist/rules-config.js.map +1 -0
- package/dist/structure.d.ts +34 -0
- package/dist/structure.d.ts.map +1 -0
- package/dist/structure.js +149 -0
- package/dist/structure.js.map +1 -0
- package/dist/tex-requirements.d.ts +43 -0
- package/dist/tex-requirements.d.ts.map +1 -0
- package/dist/tex-requirements.js +127 -0
- package/dist/tex-requirements.js.map +1 -0
- package/dist/toolchain.d.ts +159 -0
- package/dist/toolchain.d.ts.map +1 -0
- package/dist/toolchain.js +542 -0
- package/dist/toolchain.js.map +1 -0
- package/dist/types.d.ts +110 -0
- package/dist/types.d.ts.map +1 -0
- package/dist/types.js +2 -0
- package/dist/types.js.map +1 -0
- package/docs/configuration.md +235 -0
- package/docs/e2e.md +152 -0
- package/docs/incidents.md +59 -0
- package/docs/install.md +170 -0
- package/docs/optional-rules.md +107 -0
- package/docs/package-shape-options.md +262 -0
- package/docs/prior-art/README.md +76 -0
- package/docs/prior-art/blocking-vs-advisory.md +83 -0
- package/docs/prior-art/content-delivery.md +124 -0
- package/docs/prior-art/multi-mode-tools.md +106 -0
- package/docs/prior-art/nondeterministic-checks.md +99 -0
- package/docs/prior-art/package-location.md +422 -0
- package/docs/prior-art/paper-folder-scaffolding.md +538 -0
- package/docs/prior-art/readme-structure.md +69 -0
- package/docs/prior-art/repro/README.md +92 -0
- package/docs/prior-art/repro/claim1-allowedtools.mjs +66 -0
- package/docs/prior-art/repro/claim1-at2.mjs +40 -0
- package/docs/prior-art/repro/claim1-crosschannel.mjs +54 -0
- package/docs/prior-art/repro/claim1-frontmatter.mjs +76 -0
- package/docs/prior-art/repro/claim1-hook-payload-reporter.mjs +10 -0
- package/docs/prior-art/repro/claim1-plugin-frontmatter.mjs +27 -0
- package/docs/prior-art/repro/claim1-plugin-skill.mjs +52 -0
- package/docs/prior-art/repro/claim1-project-skill.mjs +81 -0
- package/docs/prior-art/repro/claim2-marketplace-flat-asclaimed.json +1 -0
- package/docs/prior-art/repro/claim2-marketplace-negative-control.json +1 -0
- package/docs/prior-art/repro/claim2-marketplace-nested-exact.json +9 -0
- package/docs/prior-art/repro/claim2-marketplace-nested-noversion.json +9 -0
- package/docs/prior-art/repro/claim2-marketplace-nested-range.json +1 -0
- package/docs/prior-art/repro/claim3-imports.mjs +50 -0
- package/docs/prior-art/repro/claim4-find-package-json.mjs +8 -0
- package/docs/prior-art/repro/claim4-package-dir.mjs +39 -0
- package/docs/prior-art/repro/claim4-parent-arg.mjs +17 -0
- package/docs/prior-art/repro/claim4-resolve-apis.mjs +21 -0
- package/docs/prior-art/repro/claim4-setup-consumers.mjs +45 -0
- package/docs/prior-art/repro/claim4-yarn-pnp.mjs +70 -0
- package/docs/prior-art/repro/claim5-bin-launch.mjs +39 -0
- package/docs/prior-art/repro/claim5-exports-mutation.mjs +57 -0
- package/docs/prior-art/repro/claim5-resolved-location-and-bin.mjs +33 -0
- package/docs/prior-art/repro/claim6-candidate-ambiguity.mjs +17 -0
- package/docs/prior-art/repro/claim6-doc-path-candidates.mjs +27 -0
- package/docs/prior-art/test-tooling.md +131 -0
- package/docs/rules.md +58 -0
- package/docs/texlive-install-decision.md +230 -0
- package/docs/toolchain.md +152 -0
- package/eslint-rules/doc-fields.harness.mjs +336 -0
- package/eslint-rules/doc-fields.mjs +186 -0
- package/eslint-rules/doc-fields.mutations.mjs +96 -0
- package/eslint-rules/install-path-literals.harness.mjs +121 -0
- package/eslint-rules/install-path-literals.mjs +108 -0
- package/eslint-rules/install-path-literals.mutations.mjs +62 -0
- package/eslint-rules/latex-language.harness.mjs +599 -0
- package/eslint-rules/latex-language.mjs +591 -0
- package/eslint-rules/latex-language.mutations.mjs +196 -0
- package/eslint-rules/paper-research-question.harness.mjs +146 -0
- package/eslint-rules/paper-research-question.mjs +180 -0
- package/eslint-rules/paper-research-question.mutations.mjs +127 -0
- package/eslint-rules/paper-stages.harness.mjs +356 -0
- package/eslint-rules/paper-stages.mjs +455 -0
- package/eslint-rules/paper-stages.mutations.mjs +157 -0
- package/eslint-rules/paper-typography.harness.mjs +291 -0
- package/eslint-rules/paper-typography.mjs +313 -0
- package/eslint-rules/paper-typography.mutations.mjs +131 -0
- package/eslint-rules/papers.harness.mjs +259 -0
- package/eslint-rules/papers.mjs +166 -0
- package/eslint-rules/papers.mutations.mjs +186 -0
- package/eslint-rules/pdf-last-page-balance.harness.mjs +206 -0
- package/eslint-rules/pdf-last-page-balance.mjs +208 -0
- package/eslint-rules/review-findings-cause.harness.mjs +228 -0
- package/eslint-rules/review-findings-cause.mjs +135 -0
- package/eslint-rules/review-findings-cause.mutations.mjs +72 -0
- package/eslint-rules/temp-root-realpath.harness.mjs +176 -0
- package/eslint-rules/temp-root-realpath.mjs +129 -0
- package/eslint-rules/temp-root-realpath.mutations.mjs +99 -0
- package/eslint-rules/tex-build.harness.mjs +753 -0
- package/eslint-rules/tex-build.mjs +322 -0
- package/eslint-rules/tex-build.mutations.mjs +258 -0
- package/eslint.config.mjs +521 -0
- package/fixtures/build-e2e/acmart/PIPELINE-STATUS.md +3 -0
- package/fixtures/build-e2e/acmart/paper.tex +11 -0
- package/fixtures/build-e2e/acmart/venue.json +1 -0
- package/fixtures/build-e2e/broken/PIPELINE-STATUS.md +3 -0
- package/fixtures/build-e2e/broken/paper.tex +7 -0
- package/fixtures/build-e2e/cite/PIPELINE-STATUS.md +3 -0
- package/fixtures/build-e2e/cite/build.sh +5 -0
- package/fixtures/build-e2e/cite/paper.tex +10 -0
- package/fixtures/build-e2e/cite/refs.bib +9 -0
- package/fixtures/build-e2e/empty/PIPELINE-STATUS.md +3 -0
- package/fixtures/build-e2e/empty/paper.tex +6 -0
- package/fixtures/build-e2e/fallback/PIPELINE-STATUS.md +3 -0
- package/fixtures/build-e2e/fallback/paper.tex +11 -0
- package/fixtures/build-e2e/guards/PIPELINE-STATUS.md +3 -0
- package/fixtures/build-e2e/guards/paper.tex +10 -0
- package/fixtures/build-e2e/no-source/PIPELINE-STATUS.md +3 -0
- package/fixtures/build-e2e/unbalanced/PIPELINE-STATUS.md +3 -0
- package/fixtures/build-e2e/unbalanced/paper.tex +28 -0
- package/fixtures/build-e2e/unbalanced/refs.bib +269 -0
- package/fixtures/install-path-literals/clean.fixture.mjs +3 -0
- package/fixtures/install-path-literals/clean.md +15 -0
- package/fixtures/install-path-literals/defect.fixture.mjs +3 -0
- package/fixtures/install-path-literals/defect.md +14 -0
- package/fixtures/latex-language/clean.tex +50 -0
- package/fixtures/latex-language/defect.tex +52 -0
- package/fixtures/paper-research-question/comment-only/PIPELINE-STATUS.md +9 -0
- package/fixtures/paper-research-question/comment-only/paper.tex +7 -0
- package/fixtures/paper-research-question/declared-not-in-paper/PIPELINE-STATUS.md +10 -0
- package/fixtures/paper-research-question/declared-not-in-paper/paper.tex +6 -0
- package/fixtures/paper-research-question/draft/PIPELINE-STATUS.md +6 -0
- package/fixtures/paper-research-question/draft/paper.tex +2 -0
- package/fixtures/paper-research-question/markdown-no-rq/PIPELINE-STATUS.md +9 -0
- package/fixtures/paper-research-question/markdown-no-rq/paper.md +4 -0
- package/fixtures/paper-research-question/shipped-no-rq/PIPELINE-STATUS.md +12 -0
- package/fixtures/paper-research-question/shipped-no-rq/paper.tex +3 -0
- package/fixtures/paper-research-question/shipped-with-rq/PIPELINE-STATUS.md +10 -0
- package/fixtures/paper-research-question/shipped-with-rq/paper.tex +2 -0
- package/fixtures/paper-stages/authors-ran/PIPELINE-STATUS.md +16 -0
- package/fixtures/paper-stages/marker-in-prose/PIPELINE-STATUS.md +17 -0
- package/fixtures/paper-stages/nofile/PIPELINE-STATUS.md +8 -0
- package/fixtures/paper-stages/noheader/PIPELINE-STATUS.md +1 -0
- package/fixtures/paper-stages/noheader/versions/2026-07-22-submitted.pdf +0 -0
- package/fixtures/paper-stages/nothing/PIPELINE-STATUS.md +3 -0
- package/fixtures/paper-stages/ok/PIPELINE-STATUS.md +9 -0
- package/fixtures/paper-stages/ok/versions/2026-07-22-submitted.pdf +0 -0
- package/fixtures/paper-stages/stale/PIPELINE-STATUS.md +1 -0
- package/fixtures/paper-stages/stale/versions/2026-07-22-submitted.STALE-WRONG-FILE.pdf +0 -0
- package/fixtures/paper-stages/twice/PIPELINE-STATUS.md +14 -0
- package/fixtures/paper-stages/twice/versions/2026-08-06-submitted.pdf +0 -0
- package/fixtures/paper-stages/twice/versions/2026-10-24-submitted.pdf +0 -0
- package/fixtures/paper-stages/undeclared/PIPELINE-STATUS.md +8 -0
- package/fixtures/paper-stages/undeclared/versions/2026-07-22-submitted.pdf +0 -0
- package/fixtures/paper-stages/undeclared/versions/2026-08-29-camera-ready.pdf +0 -0
- package/fixtures/paper-stages/wrongsize/PIPELINE-STATUS.md +8 -0
- package/fixtures/paper-stages/wrongsize/versions/2026-07-22-submitted.pdf +0 -0
- package/fixtures/paper-typography/clean-paper/paper.tex +29 -0
- package/fixtures/paper-typography/messy-paper/paper.tex +27 -0
- package/fixtures/pdf-facts/README.md +22 -0
- package/fixtures/pdf-facts/corrupt-font.pdf +0 -0
- package/fixtures/pdf-facts/encrypted.pdf +0 -0
- package/fixtures/pdf-facts/hidden-text.pdf +0 -0
- package/fixtures/pdf-facts/hidden-text.tex +28 -0
- package/fixtures/pdf-facts/t3-all.pdf +0 -0
- package/fixtures/pdf-facts/t3-all.tex +8 -0
- package/fixtures/pdf-facts/t3-mixed.pdf +0 -0
- package/fixtures/pdf-facts/t3-mixed.tex +9 -0
- package/fixtures/pdf-facts/ttf.pdf +2240 -1
- package/fixtures/pdf-facts/ttf.tex +6 -0
- package/fixtures/real-markdown-paper/baseline.json +24 -0
- package/fixtures/real-markdown-paper/baseline.mjs +48 -0
- package/fixtures/render-paper/build-clean.sh +25 -0
- package/fixtures/render-paper/build-defect.sh +15 -0
- package/fixtures/review-findings-cause/clean.md +17 -0
- package/fixtures/review-findings-cause/defect.md +14 -0
- package/fixtures/review-findings-cause/old-debt.md +14 -0
- package/fixtures/review-findings-cause/quiet-in-fence.md +16 -0
- package/fixtures/tex-build/clean.tex +21 -0
- package/fixtures/tex-build/defect.tex +24 -0
- package/fixtures/tex-build/frontmatter-clean.tex +25 -0
- package/fixtures/tex-build/frontmatter-defect.tex +23 -0
- package/fixtures/toolchain-mirror/catalog.txt +5 -0
- package/fixtures/toolchain-mirror/install-tl +27 -0
- package/fixtures/toolchain-mirror/release-texlive.txt +3 -0
- package/fixtures/toolchain-mirror/release-year +1 -0
- package/fixtures/toolchain-mirror/stub-kpsewhich +8 -0
- package/fixtures/toolchain-mirror/stub-pdflatex +3 -0
- package/fixtures/toolchain-mirror/stub-tlmgr +44 -0
- package/hooks/hooks.harness.mjs +713 -0
- package/hooks/hooks.mutations.mjs +337 -0
- package/hooks/paper-edit-guard.hook.d.mts +13 -0
- package/hooks/paper-edit-guard.hook.mjs +457 -0
- package/hooks/paper-skills-nudge.hook.mjs +136 -0
- package/hooks/paper-status-gates.hook.mjs +156 -0
- package/hooks/paper-status-gates.sh +91 -0
- package/lib/agent-cli-version.harness.mjs +165 -0
- package/lib/agent-cli-version.mjs +106 -0
- package/lib/agent-cli-version.mutations.mjs +109 -0
- package/lib/markdown.mjs +386 -0
- package/lib/mutation-driver.harness.mjs +227 -0
- package/lib/mutation-driver.mjs +397 -0
- package/lib/mutation-driver.mutations.mjs +68 -0
- package/lib/paper-config.d.mts +34 -0
- package/lib/paper-config.harness.mjs +286 -0
- package/lib/paper-config.mjs +142 -0
- package/lib/paper-config.mutations.mjs +143 -0
- package/lib/skill-checks.mjs +701 -0
- package/lib/skill-corpus.mjs +403 -0
- package/lib/skill-eval-fixture.mjs +63 -0
- package/lib/skill-eval-kit.mjs +257 -0
- package/lib/skill-trigger-cases.harness.mjs +170 -0
- package/lib/skill-trigger-cases.mjs +446 -0
- package/lib/skill-trigger-cases.mutations.mjs +65 -0
- package/lib/trigger-ledger.mjs +215 -0
- package/package.json +97 -0
- package/plugin/.claude-plugin/plugin.json +8 -0
- package/plugin/hooks/hooks.json +30 -0
- package/scripts/check.harness.mjs +177 -0
- package/scripts/check.mjs +239 -0
- package/scripts/check.mutations.mjs +110 -0
- package/scripts/eslint-report-guard.mjs +82 -0
- package/scripts/exclusive.mjs +138 -0
- package/scripts/harness-api.frozen.json +76 -0
- package/scripts/harness-api.test.ts +175 -0
- package/scripts/layer-legacy-frozen.d.mts +28 -0
- package/scripts/layer-legacy-frozen.mjs +152 -0
- package/scripts/layer-legacy-frozen.test.ts +115 -0
- package/scripts/layer-legacy.frozen.json +50 -0
- package/scripts/mutation-batteries-frozen.harness.mjs +204 -0
- package/scripts/mutation-batteries-frozen.mjs +238 -0
- package/scripts/mutation-batteries.frozen.json +117 -0
- package/scripts/release-config.test.ts +90 -0
- package/scripts/rules-are-content-only.harness.mjs +113 -0
- package/scripts/rules-are-content-only.mjs +138 -0
- package/scripts/rules-are-content-only.mutations.mjs +81 -0
- package/scripts/rules-see-files.harness.mjs +115 -0
- package/scripts/rules-see-files.mjs +99 -0
- package/scripts/rules-see-files.mutations.mjs +131 -0
- package/scripts/run-mutations.mjs +100 -0
- package/scripts/semantic-release-plugins.d.ts +16 -0
- package/skills/README.md +15 -0
- package/skills/analyze-sibling-paper/SKILL.md +170 -0
- package/skills/analyze-sibling-paper/SKILL.md.spec.ts +186 -0
- package/skills/analyze-sibling-paper/analyze-sibling-paper.eval.mjs +19 -0
- package/skills/analyze-sibling-paper/analyze-sibling-paper.harness.mjs +23 -0
- package/skills/argument-arc/SKILL.md +177 -0
- package/skills/argument-arc/SKILL.md.spec.ts +192 -0
- package/skills/argument-arc/argument-arc.eval.mjs +19 -0
- package/skills/argument-arc/argument-arc.harness.mjs +23 -0
- package/skills/build-benchmark/SKILL.md +213 -0
- package/skills/build-benchmark/SKILL.md.spec.ts +220 -0
- package/skills/build-benchmark/build-benchmark.eval.mjs +19 -0
- package/skills/build-benchmark/build-benchmark.harness.mjs +23 -0
- package/skills/build-benchmark/references/adversarial-cold-repro.md +68 -0
- package/skills/camera-ready/SKILL.md +148 -0
- package/skills/camera-ready/SKILL.md.spec.ts +164 -0
- package/skills/camera-ready/camera-ready.eval.mjs +19 -0
- package/skills/camera-ready/camera-ready.harness.mjs +23 -0
- package/skills/cold-read-diff/SKILL.md +160 -0
- package/skills/cold-read-diff/SKILL.md.spec.ts +166 -0
- package/skills/cold-read-diff/cold-read-diff.eval.mjs +19 -0
- package/skills/cold-read-diff/cold-read-diff.harness.mjs +23 -0
- package/skills/draft-paper/SKILL.md +152 -0
- package/skills/draft-paper/SKILL.md.spec.ts +169 -0
- package/skills/draft-paper/draft-paper.eval.mjs +19 -0
- package/skills/draft-paper/draft-paper.harness.mjs +23 -0
- package/skills/extend-paper/SKILL.md +99 -0
- package/skills/extend-paper/SKILL.md.spec.ts +116 -0
- package/skills/extend-paper/extend-paper.eval.mjs +19 -0
- package/skills/extend-paper/extend-paper.harness.mjs +23 -0
- package/skills/find-venue/SKILL.md +128 -0
- package/skills/find-venue/SKILL.md.spec.ts +145 -0
- package/skills/find-venue/find-venue.eval.mjs +19 -0
- package/skills/find-venue/find-venue.harness.mjs +23 -0
- package/skills/grade-paper-writing/SKILL.md +436 -0
- package/skills/grade-paper-writing/SKILL.md.spec.ts +453 -0
- package/skills/grade-paper-writing/fixtures/control_gopen.txt +1 -0
- package/skills/grade-paper-writing/fixtures/control_human_paper.txt +1 -0
- package/skills/grade-paper-writing/fixtures/rewrite.txt +1 -0
- package/skills/grade-paper-writing/fixtures/specimen.txt +1 -0
- package/skills/grade-paper-writing/fixtures/structure-checks.md +22 -0
- package/skills/grade-paper-writing/grade-paper-writing.eval.mjs +19 -0
- package/skills/grade-paper-writing/grade-paper-writing.harness.mjs +23 -0
- package/skills/grade-paper-writing/prose-lint.mjs +713 -0
- package/skills/harden-paper/SKILL.md +318 -0
- package/skills/harden-paper/SKILL.md.spec.ts +336 -0
- package/skills/harden-paper/check-numbers.sh +33 -0
- package/skills/harden-paper/check-release-claims.sh +35 -0
- package/skills/harden-paper/fixtures/uncited-assertions-sample.md +43 -0
- package/skills/harden-paper/fixtures/uncited-assertions-sample.tex +77 -0
- package/skills/harden-paper/harden-paper.eval.mjs +19 -0
- package/skills/harden-paper/harden-paper.harness.mjs +23 -0
- package/skills/map-prior-work/SKILL.md +211 -0
- package/skills/map-prior-work/SKILL.md.spec.ts +227 -0
- package/skills/map-prior-work/map-prior-work.eval.mjs +19 -0
- package/skills/map-prior-work/map-prior-work.harness.mjs +23 -0
- package/skills/osf-artifact-upload/SKILL.md +52 -0
- package/skills/osf-artifact-upload/SKILL.md.spec.ts +59 -0
- package/skills/osf-artifact-upload/osf-artifact-upload.eval.mjs +22 -0
- package/skills/osf-artifact-upload/osf-artifact-upload.harness.mjs +103 -0
- package/skills/paper-adversarial-review/SKILL.md +126 -0
- package/skills/paper-adversarial-review/SKILL.md.spec.ts +142 -0
- package/skills/paper-adversarial-review/paper-adversarial-review.eval.mjs +19 -0
- package/skills/paper-adversarial-review/paper-adversarial-review.harness.mjs +23 -0
- package/skills/paper-pipeline/PIPELINE-MAP.md +371 -0
- package/skills/paper-pipeline/SKILL.md +499 -0
- package/skills/paper-pipeline/SKILL.md.spec.ts +517 -0
- package/skills/paper-pipeline/description-language.eval.mjs +347 -0
- package/skills/paper-pipeline/framing-vs-vocabulary.eval.mjs +891 -0
- package/skills/paper-pipeline/grade-paper-writing-ablation.eval.mjs +1254 -0
- package/skills/paper-pipeline/paper-pipeline.eval.mjs +22 -0
- package/skills/paper-pipeline/paper-pipeline.harness.mjs +143 -0
- package/skills/paper-pipeline/pipeline-firing.baseline.json +270 -0
- package/skills/paper-pipeline/pipeline-firing.eval.mjs +664 -0
- package/skills/paper-pipeline/pipeline-language.eval.mjs +672 -0
- package/skills/paper-pipeline/references/acceptance-gate.md +329 -0
- package/skills/paper-pipeline/references/acl-venue-rules.md +142 -0
- package/skills/paper-pipeline/references/anonymization.md +68 -0
- package/skills/paper-pipeline/references/artifact-checklist.md +93 -0
- package/skills/paper-pipeline/references/body-vs-appendix.md +97 -0
- package/skills/paper-pipeline/references/credit-criteria.md +69 -0
- package/skills/paper-pipeline/references/occupancy-2026-08-06-probe/README.md +35 -0
- package/skills/paper-pipeline/references/occupancy-2026-08-06-probe/run_retext.mjs +24 -0
- package/skills/paper-pipeline/references/occupancy-2026-08-06-probe/sentences.txt +11 -0
- package/skills/paper-pipeline/references/occupancy-2026-08-06-probe/test_sentences.py +25 -0
- package/skills/paper-pipeline/references/occupancy-2026-08-06-prose-checkers.md +538 -0
- package/skills/paper-pipeline/references/occupancy-2026-08-06-reproducible-tooling.md +431 -0
- package/skills/paper-pipeline/references/occupancy-2026-08-06-staleness-and-orchestration.md +592 -0
- package/skills/paper-pipeline/references/pipeline-status-template.md +162 -0
- package/skills/paper-pipeline/references/review-ratchet.md +36 -0
- package/skills/paper-pipeline/references/sweep-2026-08-09-ideal-pipeline.md +585 -0
- package/skills/paper-pipeline/references/writing-craft.md +448 -0
- package/skills/paper-pipeline/repro/2026-08-07-description-language-control.log +63 -0
- package/skills/paper-pipeline/repro/2026-08-07-fork-check.log +52 -0
- package/skills/paper-pipeline/repro/2026-08-07-fork-check2.log +33 -0
- package/skills/paper-pipeline/repro/2026-08-07-language-eval-pilot.log +33 -0
- package/skills/paper-pipeline/repro/2026-08-07-language-eval-raw.log +166 -0
- package/skills/paper-pipeline/repro/2026-08-08-framing-vs-vocabulary-oracle.json +338 -0
- package/skills/paper-pipeline/repro/2026-08-08-framing-vs-vocabulary-oracle.log +118 -0
- package/skills/paper-pipeline/repro/2026-08-08-framing-vs-vocabulary-raw.log +245 -0
- package/skills/paper-pipeline/repro/2026-08-08-framing-vs-vocabulary.json +776 -0
- package/skills/paper-pipeline/repro/2026-08-08-grade-paper-writing-ablation-A6-oracle.log +53 -0
- package/skills/paper-pipeline/repro/2026-08-08-grade-paper-writing-ablation-A6-raw.log +89 -0
- package/skills/paper-pipeline/repro/2026-08-08-grade-paper-writing-ablation-oracle.json +450 -0
- package/skills/paper-pipeline/repro/2026-08-08-grade-paper-writing-ablation-oracle.log +136 -0
- package/skills/paper-pipeline/repro/2026-08-08-grade-paper-writing-ablation-raw.log +242 -0
- package/skills/paper-pipeline/repro/2026-08-08-grade-paper-writing-ablation-setupdiff.log +59 -0
- package/skills/paper-pipeline/repro/2026-08-08-grade-paper-writing-ablation.json +1032 -0
- package/skills/paper-pipeline/repro/2026-08-08-parent-replication-gpw.json +139 -0
- package/skills/paper-pipeline/repro/2026-08-08-parent-replication-gpw.log +98 -0
- package/skills/paper-pipeline/repro/2026-08-08-parent-replication.mjs +92 -0
- package/skills/paper-pipeline/repro/README.md +129 -0
- package/skills/paper-pipeline/repro/analyze-language-eval.py +116 -0
- package/skills/paper-pipeline/scripts/README.md +344 -0
- package/skills/paper-pipeline/scripts/announce.mjs +67 -0
- package/skills/paper-pipeline/scripts/artifact-coverage.harness.mjs +496 -0
- package/skills/paper-pipeline/scripts/artifact-coverage.mjs +397 -0
- package/skills/paper-pipeline/scripts/artifact-coverage.mutations.mjs +218 -0
- package/skills/paper-pipeline/scripts/check-provenance.mjs +184 -0
- package/skills/paper-pipeline/scripts/consumer.d.mts +32 -0
- package/skills/paper-pipeline/scripts/consumer.harness.mjs +562 -0
- package/skills/paper-pipeline/scripts/consumer.mjs +535 -0
- package/skills/paper-pipeline/scripts/consumer.mutations.mjs +190 -0
- package/skills/paper-pipeline/scripts/extract-ref-facts.harness.mjs +457 -0
- package/skills/paper-pipeline/scripts/extract-ref-facts.mjs +656 -0
- package/skills/paper-pipeline/scripts/extract-ref-facts.mutations.mjs +54 -0
- package/skills/paper-pipeline/scripts/fixtures/clean/PIPELINE-STATUS.md +51 -0
- package/skills/paper-pipeline/scripts/fixtures/dirty/PIPELINE-STATUS.md +52 -0
- package/skills/paper-pipeline/scripts/fixtures/dirty/paper.md +6 -0
- package/skills/paper-pipeline/scripts/fixtures/real-bib/refs.bib +153 -0
- package/skills/paper-pipeline/scripts/generated-code.harness.mjs +466 -0
- package/skills/paper-pipeline/scripts/generated-code.mjs +338 -0
- package/skills/paper-pipeline/scripts/generated-code.mutations.mjs +254 -0
- package/skills/paper-pipeline/scripts/ledger.mjs +623 -0
- package/skills/paper-pipeline/scripts/ledger.selftest.mjs +286 -0
- package/skills/paper-pipeline/scripts/pipeline-check.harness.mjs +389 -0
- package/skills/paper-pipeline/scripts/pipeline-check.mjs +737 -0
- package/skills/paper-pipeline/scripts/pipeline-check.mutations.mjs +54 -0
- package/skills/paper-pipeline/scripts/pipeline-edges.mjs +169 -0
- package/skills/paper-pipeline/scripts/population-map.harness.mjs +178 -0
- package/skills/paper-pipeline/scripts/population-map.mjs +181 -0
- package/skills/paper-pipeline/scripts/population-map.mutations.mjs +65 -0
- package/skills/paper-pipeline/scripts/population-map.selftest.mjs +122 -0
- package/skills/paper-pipeline/scripts/provenance.harness.mjs +240 -0
- package/skills/paper-pipeline/scripts/provenance.mutations.mjs +59 -0
- package/skills/paper-pipeline/scripts/round-diff.harness.mjs +881 -0
- package/skills/paper-pipeline/scripts/round-diff.mjs +576 -0
- package/skills/paper-pipeline/scripts/round-diff.mutations.mjs +276 -0
- package/skills/paper-pipeline/scripts/run-mechanical.mjs +633 -0
- package/skills/paper-pipeline/scripts/status.mjs +295 -0
- package/skills/paper-status/SKILL.md +183 -0
- package/skills/paper-status/SKILL.md.spec.ts +190 -0
- package/skills/paper-status/paper-status.eval.mjs +22 -0
- package/skills/paper-status/paper-status.harness.mjs +25 -0
- package/skills/pc-panel-review/SKILL.md +263 -0
- package/skills/pc-panel-review/SKILL.md.spec.ts +280 -0
- package/skills/pc-panel-review/pc-panel-review.eval.mjs +19 -0
- package/skills/pc-panel-review/pc-panel-review.harness.mjs +23 -0
- package/skills/plan-paper-timeline/SKILL.md +182 -0
- package/skills/plan-paper-timeline/SKILL.md.spec.ts +200 -0
- package/skills/plan-paper-timeline/fixtures/fake-google-calendar.mjs +239 -0
- package/skills/plan-paper-timeline/plan-paper-timeline.effects.harness.mjs +431 -0
- package/skills/plan-paper-timeline/plan-paper-timeline.effects.mutations.mjs +65 -0
- package/skills/plan-paper-timeline/plan-paper-timeline.eval.mjs +19 -0
- package/skills/plan-paper-timeline/plan-paper-timeline.harness.mjs +23 -0
- package/skills/render-paper/SKILL.md +159 -0
- package/skills/render-paper/SKILL.md.spec.ts +166 -0
- package/skills/render-paper/check-render.sh +419 -0
- package/skills/render-paper/checkers-requirements.txt +55 -0
- package/skills/render-paper/ensure-checkers.sh +69 -0
- package/skills/render-paper/extract-pdf-facts.harness.mjs +166 -0
- package/skills/render-paper/extract-pdf-facts.mjs +144 -0
- package/skills/render-paper/render-paper.eval.mjs +19 -0
- package/skills/render-paper/render-paper.harness.mjs +339 -0
- package/skills/research-ideate/SKILL.md +136 -0
- package/skills/research-ideate/SKILL.md.spec.ts +152 -0
- package/skills/research-ideate/research-ideate.eval.mjs +19 -0
- package/skills/research-ideate/research-ideate.harness.mjs +23 -0
- package/skills/skill-contract.mutations.mjs +179 -0
- package/skills/study-accepted-papers/SKILL.md +206 -0
- package/skills/study-accepted-papers/SKILL.md.spec.ts +223 -0
- package/skills/study-accepted-papers/study-accepted-papers.eval.mjs +19 -0
- package/skills/study-accepted-papers/study-accepted-papers.harness.mjs +23 -0
- package/skills/submit-paper/SKILL.md +182 -0
- package/skills/submit-paper/SKILL.md.spec.ts +199 -0
- package/skills/submit-paper/check-deanon.sh +149 -0
- package/skills/submit-paper/references/publishers/acm.md +92 -0
- package/skills/submit-paper/references/venues/agenticdev.jsonc +108 -0
- package/skills/submit-paper/references/venues/agenticdev.md +139 -0
- package/skills/submit-paper/references/venues/agenticdev.tex +19 -0
- package/skills/submit-paper/references/venues/aisec.jsonc +101 -0
- package/skills/submit-paper/references/venues/aisec.md +105 -0
- package/skills/submit-paper/references/venues/paper-guards.tex +41 -0
- package/skills/submit-paper/references/venues/realm.jsonc +81 -0
- package/skills/submit-paper/references/venues/realm.md +155 -0
- package/skills/submit-paper/references/venues/tex-base.jsonc +50 -0
- package/skills/submit-paper/references/venues/venue-profile.schema.json +74 -0
- package/skills/submit-paper/submit-paper.eval.mjs +19 -0
- package/skills/submit-paper/submit-paper.harness.mjs +23 -0
- package/skills/sweep-design-space/SKILL.md +269 -0
- package/skills/sweep-design-space/SKILL.md.spec.ts +285 -0
- package/skills/sweep-design-space/sweep-design-space.eval.mjs +19 -0
- package/skills/sweep-design-space/sweep-design-space.harness.mjs +23 -0
- package/skills/tighten-paper/SKILL.md +368 -0
- package/skills/tighten-paper/SKILL.md.spec.ts +384 -0
- package/skills/tighten-paper/structure.mjs +371 -0
- package/skills/tighten-paper/tighten-paper.eval.mjs +19 -0
- package/skills/tighten-paper/tighten-paper.harness.mjs +23 -0
- package/skills/verify-citations/SKILL.md +328 -0
- package/skills/verify-citations/SKILL.md.spec.ts +345 -0
- package/skills/verify-citations/scripts/bib-authors.mjs +479 -0
- package/skills/verify-citations/scripts/bib-authors.test.mjs +175 -0
- package/skills/verify-citations/scripts/verify-cites.mjs +1108 -0
- package/skills/verify-citations/scripts/verify-cites.test.mjs +735 -0
- package/skills/verify-citations/verify-citations.eval.mjs +19 -0
- package/skills/verify-citations/verify-citations.harness.mjs +23 -0
- package/src/CLAUDE.md +51 -0
- package/src/action-ref.test.ts +26 -0
- package/src/action-ref.ts +15 -0
- package/src/adapters/banal/failure.test.ts +63 -0
- package/src/adapters/banal/failure.ts +118 -0
- package/src/adapters/banal/index.test.ts +119 -0
- package/src/adapters/banal/index.ts +100 -0
- package/src/adapters/banal/install.test.ts +20 -0
- package/src/adapters/banal/install.ts +41 -0
- package/src/adapters/banal/invocation.test.ts +74 -0
- package/src/adapters/banal/invocation.ts +95 -0
- package/src/adapters/banal/locate.test.ts +52 -0
- package/src/adapters/banal/locate.ts +84 -0
- package/src/adapters/banal/output.test.ts +140 -0
- package/src/adapters/banal/output.ts +141 -0
- package/src/adapters/banal/pin.ts +30 -0
- package/src/adapters/banal/probe.ts +35 -0
- package/src/adapters/banal/run.test.ts +191 -0
- package/src/adapters/banal/run.ts +244 -0
- package/src/adapters/banal/settings.test.ts +31 -0
- package/src/adapters/banal/settings.ts +55 -0
- package/src/adapters/banal/xml.test.ts +111 -0
- package/src/adapters/banal/xml.ts +112 -0
- package/src/adapters/curl/download.io.ts +73 -0
- package/src/adapters/curl/download.test.ts +55 -0
- package/src/adapters/curl/index.ts +5 -0
- package/src/adapters/memory/index.ts +131 -0
- package/src/adapters/node/files.io.ts +39 -0
- package/src/adapters/node/files.test.ts +28 -0
- package/src/adapters/node/host.io.ts +15 -0
- package/src/adapters/node/index.ts +36 -0
- package/src/adapters/node/process.io.ts +49 -0
- package/src/adapters/node/process.test.ts +46 -0
- package/src/adapters/node/workspace.io.ts +40 -0
- package/src/adapters/node/workspace.test.ts +58 -0
- package/src/adapters/pdfjs/fill.test.ts +111 -0
- package/src/adapters/pdfjs/fill.ts +141 -0
- package/src/build-engine.harness.mjs +314 -0
- package/src/build-engine.ts +219 -0
- package/src/build.harness.mjs +631 -0
- package/src/build.mutations.mjs +195 -0
- package/src/build.ts +793 -0
- package/src/cli.harness.mjs +2007 -0
- package/src/cli.mutations.mjs +448 -0
- package/src/cli.ts +1189 -0
- package/src/doctor.harness.mjs +396 -0
- package/src/doctor.mutations.mjs +175 -0
- package/src/doctor.ts +356 -0
- package/src/domain/geometry.ts +108 -0
- package/src/domain/host.ts +23 -0
- package/src/domain/page-layout.ts +32 -0
- package/src/domain/paths.ts +5 -0
- package/src/domain/result.test.ts +26 -0
- package/src/domain/result.ts +29 -0
- package/src/domain/sha256.test.ts +12 -0
- package/src/domain/sha256.ts +21 -0
- package/src/domain/text.ts +11 -0
- package/src/engine.harness.mjs +252 -0
- package/src/engine.ts +176 -0
- package/src/exit-code.test.ts +21 -0
- package/src/exit-code.ts +38 -0
- package/src/facts-file.test.ts +240 -0
- package/src/facts-file.ts +241 -0
- package/src/hooks-settings.harness.mjs +386 -0
- package/src/hooks-settings.mutations.mjs +116 -0
- package/src/hooks-settings.ts +434 -0
- package/src/init.ts +900 -0
- package/src/latex-log.harness.mjs +226 -0
- package/src/latex-log.ts +234 -0
- package/src/latex-loop.harness.mjs +449 -0
- package/src/latex-loop.ts +211 -0
- package/src/link-skills.harness.mjs +273 -0
- package/src/link-skills.mutations.mjs +136 -0
- package/src/link-skills.ts +258 -0
- package/src/new-paper.harness.mjs +216 -0
- package/src/new-paper.mutations.mjs +79 -0
- package/src/new-paper.ts +158 -0
- package/src/pdf-facts.harness.mjs +188 -0
- package/src/pdf-facts.ts +327 -0
- package/src/pdf-geometry.harness.mjs +254 -0
- package/src/pdf-geometry.ts +300 -0
- package/src/ports/download.ts +10 -0
- package/src/ports/files.ts +11 -0
- package/src/ports/measure-geometry.ts +8 -0
- package/src/ports/process.ts +46 -0
- package/src/ports/tool-installer.ts +33 -0
- package/src/ports/workspace.ts +20 -0
- package/src/rules-config.harness.mjs +114 -0
- package/src/rules-config.ts +178 -0
- package/src/structure.harness.mjs +179 -0
- package/src/structure.mutations.mjs +83 -0
- package/src/structure.ts +166 -0
- package/src/tex-requirements.harness.mjs +238 -0
- package/src/tex-requirements.ts +181 -0
- package/src/toolchain.harness.mjs +651 -0
- package/src/toolchain.ts +755 -0
- package/src/types.ts +106 -0
- package/templates/paper/PIPELINE-STATUS.md +72 -0
- package/templates/paper/paper.md +4 -0
- package/templates/paper/paper.tex +8 -0
- package/tsconfig.json +23 -0
|
@@ -0,0 +1,177 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: argument-arc
|
|
3
|
+
description: Build or repair the paper's argument architecture — the one-sentence-per-section outline, the bottom-up inevitability pass, and the name/number budget. Run it when the reader says the paper throws ideas at them, when a structural objection repeats, or before any large rewrite. Not a prose or length skill.
|
|
4
|
+
allowed-tools: [Read, Write, Grep, Glob, Agent, Bash(node .claude/skills/paper-pipeline/scripts/announce.mjs:*), Bash(node .claude/skills/paper-pipeline/scripts/ledger.mjs:*)]
|
|
5
|
+
---
|
|
6
|
+
|
|
7
|
+
<!-- vigiles:sha256:05a1eaa79bb596d5 compiled from skills/argument-arc/SKILL.md.spec.ts -->
|
|
8
|
+
|
|
9
|
+
# argument-arc — does the paper carry the reader to one conclusion
|
|
10
|
+
|
|
11
|
+
> **Which structure skill?** `tighten-paper` = length, sag, what to cut. `grade-paper-writing` =
|
|
12
|
+
> sentences, jargon, where a reader stalls. **`argument-arc` (you are here) = the load-bearing order
|
|
13
|
+
> of ideas.** The first two assume the argument is right and the delivery is wrong. This one asks
|
|
14
|
+
> whether the argument exists. Run it *before* them — cutting words inside a broken arc is how a
|
|
15
|
+
> session burns a day and ships the same paper.
|
|
16
|
+
|
|
17
|
+
## Run me
|
|
18
|
+
|
|
19
|
+
🔴 FIRST, before any other step:
|
|
20
|
+
|
|
21
|
+
```
|
|
22
|
+
node .claude/skills/paper-pipeline/scripts/announce.mjs argument-arc <paper-dir>
|
|
23
|
+
```
|
|
24
|
+
|
|
25
|
+
An advisory pass cannot be observed failing — silence is both its error state and its normal
|
|
26
|
+
state — so starting is an event, and events get written down.
|
|
27
|
+
|
|
28
|
+
The failure it exists to catch has a signature: every paragraph is defensible, every number is real,
|
|
29
|
+
and the reader still finishes the section unable to say what it was for. That is not a prose problem
|
|
30
|
+
and no amount of rewriting sentences fixes it.
|
|
31
|
+
|
|
32
|
+
## Two modes — and one of them runs BEFORE the study exists
|
|
33
|
+
|
|
34
|
+
| Mode | Where it runs | Scorecard row | What it produces |
|
|
35
|
+
|---|---|---|---|
|
|
36
|
+
| **frame** | SETUP, before any run | **`frame`** | **One paragraph: what will this paper claim, and what would have to be true for that claim to hold?** |
|
|
37
|
+
| **rebuild** | the LOOP, on a draft whose framing is wrong | `arc` | A new arc, built from scratch through steps 1–4 |
|
|
38
|
+
|
|
39
|
+
### Frame mode (SETUP) — one paragraph, before the first run
|
|
40
|
+
|
|
41
|
+
Write it before the study is designed. Not an abstract, not an outline — **a paragraph saying what
|
|
42
|
+
the paper will claim** and what would have to be true for that claim to hold.
|
|
43
|
+
|
|
44
|
+
**Why it is here and not later:** the frame decides which results matter. A study designed without a
|
|
45
|
+
stated claim measures what the harness can already see, and then the claim gets fitted to whatever
|
|
46
|
+
came out — which is how a run gets thrown away. That is not hypothetical: it was observed on
|
|
47
|
+
`the reference paper`, and it is the reason this skill has a rebuild mode at all. **Reframing after the
|
|
48
|
+
data is collected costs the data. Costs a paragraph; saves a study.**
|
|
49
|
+
|
|
50
|
+
`build-benchmark` reads this paragraph before designing, and refuses to proceed if `frame` is empty.
|
|
51
|
+
`pipeline-check.mjs` reports a study that ran with no stated claim.
|
|
52
|
+
|
|
53
|
+
*Added 2026-08-03 alongside the replacement of the numbered stage list — this skill previously existed
|
|
54
|
+
only as a repair, which meant nobody stated the claim while stating it was still cheap.*
|
|
55
|
+
|
|
56
|
+
## When the arc pass fires (row `arc`, in the LOOP)
|
|
57
|
+
|
|
58
|
+
- 🔴 **A structural objection repeats.** Not "this is unclear" but *"it throws ideas at me"*,
|
|
59
|
+
*"I can't hold this in my head"*, *"where is this going"*. **The second time you hear the same
|
|
60
|
+
objection, stop editing and run this.** The third time means you already ignored the second.
|
|
61
|
+
- Before a rewrite touching more than one section.
|
|
62
|
+
- After a claim dies. When a refutation pass kills a load-bearing claim, the arc built on it is
|
|
63
|
+
usually dead too, and it will not announce itself — the sections still read fine one at a time.
|
|
64
|
+
- Before `tighten-paper` / `grade-paper-writing` / `pc-panel-review` on any draft whose thesis has
|
|
65
|
+
changed since those gates last ran.
|
|
66
|
+
|
|
67
|
+
## How to run it
|
|
68
|
+
|
|
69
|
+
### 1. Write the conclusion first, in one sentence, in the author's own words
|
|
70
|
+
Not the abstract. The sentence you want the reader thinking as they close the paper. If it takes two
|
|
71
|
+
sentences, the paper has two papers in it and that is the finding.
|
|
72
|
+
|
|
73
|
+
Ask the author for it if there is any doubt. A conclusion you inferred is a conclusion you will
|
|
74
|
+
defend against them.
|
|
75
|
+
|
|
76
|
+
### 2. One sentence per section: what does the reader carry out
|
|
77
|
+
Build the whole outline as a flat list, section by section, each line answering **only** *what does
|
|
78
|
+
the reader now believe that they did not believe before this section*. Not "what this section
|
|
79
|
+
covers" — coverage is a table of contents and it hides the defect.
|
|
80
|
+
|
|
81
|
+
Then read the list on its own, without the paper:
|
|
82
|
+
- **Two lines carrying the same belief** → the sections merge. No exceptions; "but they use different
|
|
83
|
+
evidence" means one section with two pieces of evidence.
|
|
84
|
+
- **A line you cannot write** → that section has no job. It goes to the artifact or it goes away.
|
|
85
|
+
- **A line that is a topic, not a belief** ("we survey related work") → rewrite it as a belief or
|
|
86
|
+
admit the section is filler.
|
|
87
|
+
|
|
88
|
+
### 3. The inevitability pass — bottom to top
|
|
89
|
+
Start at the conclusion and walk **upward**, asking at each section: *does this make the conclusion
|
|
90
|
+
harder to escape?* A section that merely supports the conclusion is not enough; the test is whether
|
|
91
|
+
removing it lets the reader wriggle out.
|
|
92
|
+
|
|
93
|
+
This direction matters. Top-down you will narrate the paper you wrote. Bottom-up you find the
|
|
94
|
+
sections that are true, interesting and load-bearing for nothing.
|
|
95
|
+
|
|
96
|
+
### 4. Count the names
|
|
97
|
+
Count every named system, benchmark, corpus and acronym in the body. For each, ask: **must the reader
|
|
98
|
+
remember this later?** If the answer is no, it becomes an unnamed instance ("one corpus study",
|
|
99
|
+
"a trace monitor") or a citation, and the name moves to the artifact or a footnote.
|
|
100
|
+
|
|
101
|
+
Report the count. A reader who says they cannot hold the paper in their head is usually telling you
|
|
102
|
+
this number, not the word count. There is no correct threshold — but a body carrying more names than
|
|
103
|
+
sections is worth arguing about.
|
|
104
|
+
|
|
105
|
+
### 5. 🔴 Show the outline to the author BEFORE editing anything
|
|
106
|
+
The deliverable of this skill is the outline and the verdict, **not a rewritten paper**. Hand over:
|
|
107
|
+
the one-sentence conclusion · the section list · which sections merge, move, die · the name count ·
|
|
108
|
+
the single biggest arc defect in one sentence.
|
|
109
|
+
|
|
110
|
+
Editing before this is agreed is how five patch rounds happen. The author asked for an architecture;
|
|
111
|
+
returning a diff answers a question they did not ask.
|
|
112
|
+
|
|
113
|
+
## The rebuild mode — when the frame itself is wrong
|
|
114
|
+
|
|
115
|
+
Every other gate in the pipeline assumes a settled draft and improves it. None has a mode for *the
|
|
116
|
+
framing is wrong, start the arc over*, which is why gates feel inappropriate exactly when the paper
|
|
117
|
+
most needs help. This skill does:
|
|
118
|
+
|
|
119
|
+
- **Do not defend the current form because it exists.** The question is *"if all this work had been
|
|
120
|
+
done yesterday by someone else, would I choose this shape?"* Time already spent is not evidence.
|
|
121
|
+
- **Salvage explicitly, in writing:** which measurements survive the reframe untouched, which need
|
|
122
|
+
re-analysis, which die with the old frame. Measurements usually survive; framing paragraphs rarely do.
|
|
123
|
+
- **A deadline bounds how much you rebuild. It is never an argument that the current shape is right.**
|
|
124
|
+
If the author says the paper is *weak*, that is a quality judgment, not procrastination, and citing
|
|
125
|
+
the calendar against it is a category error.
|
|
126
|
+
- **Rebuild produces a new outline through steps 1–4 above, and it too gets shown before editing.**
|
|
127
|
+
|
|
128
|
+
## Record the verdict
|
|
129
|
+
|
|
130
|
+
🔴 LAST step, once the deliverable exists:
|
|
131
|
+
|
|
132
|
+
```
|
|
133
|
+
node .claude/skills/paper-pipeline/scripts/ledger.mjs record argument-arc <paper-dir> FINDING <count> <report-path>
|
|
134
|
+
node .claude/skills/paper-pipeline/scripts/ledger.mjs record argument-arc <paper-dir> ABSTAINED <reason> "<one line>"
|
|
135
|
+
```
|
|
136
|
+
|
|
137
|
+
🔴 **There is no PASS.** "The arc carries the reader" is not a value this pass can write down; it is
|
|
138
|
+
what a reader may conclude from the absence of findings, and the concluding is theirs.
|
|
139
|
+
|
|
140
|
+
**FINDING** — the arc is broken; `<count>` is the number of sections the new outline moves, merges
|
|
141
|
+
or deletes, and `<report-path>` is that outline. Add `--blocking` when there is no conclusion
|
|
142
|
+
sentence to build an arc from: that refusal is a finding about the draft, not an aborted run, and
|
|
143
|
+
`<report-path>` is the one-line statement of what is missing.
|
|
144
|
+
**ABSTAINED** — `no-witness`: the outline was read end to end and nothing moved. `blocked`: there
|
|
145
|
+
was no outline to read at all.
|
|
146
|
+
|
|
147
|
+
An `ABSTAINED no-witness` on a draft the author has twice called a pile is a missed finding, and
|
|
148
|
+
`status.mjs` will start asking about a check that has only ever abstained.
|
|
149
|
+
|
|
150
|
+
## Rules
|
|
151
|
+
|
|
152
|
+
- **The arc is judged on the outline, never on the draft.** If you find yourself rereading paragraphs
|
|
153
|
+
to decide whether a section belongs, you are grading prose again.
|
|
154
|
+
- **Merge is the default remedy, deletion the second.** Most arc defects are two sections doing one
|
|
155
|
+
job, not a section doing nothing.
|
|
156
|
+
- **Never answer a structural objection with a local edit.** Moving a paragraph and reporting "done"
|
|
157
|
+
is the exact failure this skill was built from.
|
|
158
|
+
- **The author's confusion is data about the paper, not about the author.** If they ask twice what a
|
|
159
|
+
section is for, the section is the problem.
|
|
160
|
+
- **Say what the reader must hold in their head at each point.** If that set only grows, the arc is
|
|
161
|
+
a pile.
|
|
162
|
+
|
|
163
|
+
## Compose with
|
|
164
|
+
|
|
165
|
+
- **Before** `tighten-paper` (length/sag) and `grade-paper-writing` (sentences/stalls) — both assume
|
|
166
|
+
the arc holds. Their verdicts are the gate inputs `pc-panel-review` demands.
|
|
167
|
+
- **After** any refutation pass that killed a load-bearing claim (see the paper's `CLAIMS.md`).
|
|
168
|
+
- `sweep-design-space` when the outline shows the paper has no distinctive move to make — that is a
|
|
169
|
+
design problem, not an arc problem.
|
|
170
|
+
|
|
171
|
+
## Provenance
|
|
172
|
+
|
|
173
|
+
Built 2026-07-30 after a full-day rewrite of `the reference paper` in which the author said the same
|
|
174
|
+
thing five times — *"does not guide the reader sequentially through the ideas, but simply throws a bunch of
|
|
175
|
+
different ideas, references, and benchmarks at them"* — and got five local edits in reply. The pipeline had a skill
|
|
176
|
+
for length, a skill for sentences and a skill for defects; the question *does this paper carry the
|
|
177
|
+
reader to one conclusion* belonged to nobody, and that is the one that failed.
|
|
@@ -0,0 +1,192 @@
|
|
|
1
|
+
// Compiled to SKILL.md by `vigiles compile`. Edit THIS file, never the markdown.
|
|
2
|
+
//
|
|
3
|
+
// Adopted 2026-08-17 as part of the second batch. The body is the previous SKILL.md
|
|
4
|
+
// VERBATIM, so the compiled diff shows only what the compiler adds. No `disallowedTools`
|
|
5
|
+
// fence yet: the field exists on `SkillSpec` as of the branch `claude/skill-disallowed-tools`
|
|
6
|
+
// but is not in a release `mine` installs, so adding it here would not compile.
|
|
7
|
+
import { experimental_skill } from "vigiles/spec";
|
|
8
|
+
|
|
9
|
+
export default experimental_skill({
|
|
10
|
+
name: "argument-arc",
|
|
11
|
+
description:
|
|
12
|
+
"Build or repair the paper's argument architecture — the one-sentence-per-section outline, the bottom-up inevitability pass, and the name/number budget. Run it when the reader says the paper throws ideas at them, when a structural objection repeats, or before any large rewrite. Not a prose or length skill.",
|
|
13
|
+
tools: [
|
|
14
|
+
"Read",
|
|
15
|
+
"Write",
|
|
16
|
+
"Grep",
|
|
17
|
+
"Glob",
|
|
18
|
+
"Agent",
|
|
19
|
+
"Bash(node .claude/skills/paper-pipeline/scripts/announce.mjs:*)",
|
|
20
|
+
"Bash(node .claude/skills/paper-pipeline/scripts/ledger.mjs:*)",
|
|
21
|
+
],
|
|
22
|
+
body: `
|
|
23
|
+
# argument-arc — does the paper carry the reader to one conclusion
|
|
24
|
+
|
|
25
|
+
> **Which structure skill?** \`tighten-paper\` = length, sag, what to cut. \`grade-paper-writing\` =
|
|
26
|
+
> sentences, jargon, where a reader stalls. **\`argument-arc\` (you are here) = the load-bearing order
|
|
27
|
+
> of ideas.** The first two assume the argument is right and the delivery is wrong. This one asks
|
|
28
|
+
> whether the argument exists. Run it *before* them — cutting words inside a broken arc is how a
|
|
29
|
+
> session burns a day and ships the same paper.
|
|
30
|
+
|
|
31
|
+
## Run me
|
|
32
|
+
|
|
33
|
+
🔴 FIRST, before any other step:
|
|
34
|
+
|
|
35
|
+
\`\`\`
|
|
36
|
+
node .claude/skills/paper-pipeline/scripts/announce.mjs argument-arc <paper-dir>
|
|
37
|
+
\`\`\`
|
|
38
|
+
|
|
39
|
+
An advisory pass cannot be observed failing — silence is both its error state and its normal
|
|
40
|
+
state — so starting is an event, and events get written down.
|
|
41
|
+
|
|
42
|
+
The failure it exists to catch has a signature: every paragraph is defensible, every number is real,
|
|
43
|
+
and the reader still finishes the section unable to say what it was for. That is not a prose problem
|
|
44
|
+
and no amount of rewriting sentences fixes it.
|
|
45
|
+
|
|
46
|
+
## Two modes — and one of them runs BEFORE the study exists
|
|
47
|
+
|
|
48
|
+
| Mode | Where it runs | Scorecard row | What it produces |
|
|
49
|
+
|---|---|---|---|
|
|
50
|
+
| **frame** | SETUP, before any run | **\`frame\`** | **One paragraph: what will this paper claim, and what would have to be true for that claim to hold?** |
|
|
51
|
+
| **rebuild** | the LOOP, on a draft whose framing is wrong | \`arc\` | A new arc, built from scratch through steps 1–4 |
|
|
52
|
+
|
|
53
|
+
### Frame mode (SETUP) — one paragraph, before the first run
|
|
54
|
+
|
|
55
|
+
Write it before the study is designed. Not an abstract, not an outline — **a paragraph saying what
|
|
56
|
+
the paper will claim** and what would have to be true for that claim to hold.
|
|
57
|
+
|
|
58
|
+
**Why it is here and not later:** the frame decides which results matter. A study designed without a
|
|
59
|
+
stated claim measures what the harness can already see, and then the claim gets fitted to whatever
|
|
60
|
+
came out — which is how a run gets thrown away. That is not hypothetical: it was observed on
|
|
61
|
+
\`the reference paper\`, and it is the reason this skill has a rebuild mode at all. **Reframing after the
|
|
62
|
+
data is collected costs the data. Costs a paragraph; saves a study.**
|
|
63
|
+
|
|
64
|
+
\`build-benchmark\` reads this paragraph before designing, and refuses to proceed if \`frame\` is empty.
|
|
65
|
+
\`pipeline-check.mjs\` reports a study that ran with no stated claim.
|
|
66
|
+
|
|
67
|
+
*Added 2026-08-03 alongside the replacement of the numbered stage list — this skill previously existed
|
|
68
|
+
only as a repair, which meant nobody stated the claim while stating it was still cheap.*
|
|
69
|
+
|
|
70
|
+
## When the arc pass fires (row \`arc\`, in the LOOP)
|
|
71
|
+
|
|
72
|
+
- 🔴 **A structural objection repeats.** Not "this is unclear" but *"it throws ideas at me"*,
|
|
73
|
+
*"I can't hold this in my head"*, *"where is this going"*. **The second time you hear the same
|
|
74
|
+
objection, stop editing and run this.** The third time means you already ignored the second.
|
|
75
|
+
- Before a rewrite touching more than one section.
|
|
76
|
+
- After a claim dies. When a refutation pass kills a load-bearing claim, the arc built on it is
|
|
77
|
+
usually dead too, and it will not announce itself — the sections still read fine one at a time.
|
|
78
|
+
- Before \`tighten-paper\` / \`grade-paper-writing\` / \`pc-panel-review\` on any draft whose thesis has
|
|
79
|
+
changed since those gates last ran.
|
|
80
|
+
|
|
81
|
+
## How to run it
|
|
82
|
+
|
|
83
|
+
### 1. Write the conclusion first, in one sentence, in the author's own words
|
|
84
|
+
Not the abstract. The sentence you want the reader thinking as they close the paper. If it takes two
|
|
85
|
+
sentences, the paper has two papers in it and that is the finding.
|
|
86
|
+
|
|
87
|
+
Ask the author for it if there is any doubt. A conclusion you inferred is a conclusion you will
|
|
88
|
+
defend against them.
|
|
89
|
+
|
|
90
|
+
### 2. One sentence per section: what does the reader carry out
|
|
91
|
+
Build the whole outline as a flat list, section by section, each line answering **only** *what does
|
|
92
|
+
the reader now believe that they did not believe before this section*. Not "what this section
|
|
93
|
+
covers" — coverage is a table of contents and it hides the defect.
|
|
94
|
+
|
|
95
|
+
Then read the list on its own, without the paper:
|
|
96
|
+
- **Two lines carrying the same belief** → the sections merge. No exceptions; "but they use different
|
|
97
|
+
evidence" means one section with two pieces of evidence.
|
|
98
|
+
- **A line you cannot write** → that section has no job. It goes to the artifact or it goes away.
|
|
99
|
+
- **A line that is a topic, not a belief** ("we survey related work") → rewrite it as a belief or
|
|
100
|
+
admit the section is filler.
|
|
101
|
+
|
|
102
|
+
### 3. The inevitability pass — bottom to top
|
|
103
|
+
Start at the conclusion and walk **upward**, asking at each section: *does this make the conclusion
|
|
104
|
+
harder to escape?* A section that merely supports the conclusion is not enough; the test is whether
|
|
105
|
+
removing it lets the reader wriggle out.
|
|
106
|
+
|
|
107
|
+
This direction matters. Top-down you will narrate the paper you wrote. Bottom-up you find the
|
|
108
|
+
sections that are true, interesting and load-bearing for nothing.
|
|
109
|
+
|
|
110
|
+
### 4. Count the names
|
|
111
|
+
Count every named system, benchmark, corpus and acronym in the body. For each, ask: **must the reader
|
|
112
|
+
remember this later?** If the answer is no, it becomes an unnamed instance ("one corpus study",
|
|
113
|
+
"a trace monitor") or a citation, and the name moves to the artifact or a footnote.
|
|
114
|
+
|
|
115
|
+
Report the count. A reader who says they cannot hold the paper in their head is usually telling you
|
|
116
|
+
this number, not the word count. There is no correct threshold — but a body carrying more names than
|
|
117
|
+
sections is worth arguing about.
|
|
118
|
+
|
|
119
|
+
### 5. 🔴 Show the outline to the author BEFORE editing anything
|
|
120
|
+
The deliverable of this skill is the outline and the verdict, **not a rewritten paper**. Hand over:
|
|
121
|
+
the one-sentence conclusion · the section list · which sections merge, move, die · the name count ·
|
|
122
|
+
the single biggest arc defect in one sentence.
|
|
123
|
+
|
|
124
|
+
Editing before this is agreed is how five patch rounds happen. The author asked for an architecture;
|
|
125
|
+
returning a diff answers a question they did not ask.
|
|
126
|
+
|
|
127
|
+
## The rebuild mode — when the frame itself is wrong
|
|
128
|
+
|
|
129
|
+
Every other gate in the pipeline assumes a settled draft and improves it. None has a mode for *the
|
|
130
|
+
framing is wrong, start the arc over*, which is why gates feel inappropriate exactly when the paper
|
|
131
|
+
most needs help. This skill does:
|
|
132
|
+
|
|
133
|
+
- **Do not defend the current form because it exists.** The question is *"if all this work had been
|
|
134
|
+
done yesterday by someone else, would I choose this shape?"* Time already spent is not evidence.
|
|
135
|
+
- **Salvage explicitly, in writing:** which measurements survive the reframe untouched, which need
|
|
136
|
+
re-analysis, which die with the old frame. Measurements usually survive; framing paragraphs rarely do.
|
|
137
|
+
- **A deadline bounds how much you rebuild. It is never an argument that the current shape is right.**
|
|
138
|
+
If the author says the paper is *weak*, that is a quality judgment, not procrastination, and citing
|
|
139
|
+
the calendar against it is a category error.
|
|
140
|
+
- **Rebuild produces a new outline through steps 1–4 above, and it too gets shown before editing.**
|
|
141
|
+
|
|
142
|
+
## Record the verdict
|
|
143
|
+
|
|
144
|
+
🔴 LAST step, once the deliverable exists:
|
|
145
|
+
|
|
146
|
+
\`\`\`
|
|
147
|
+
node .claude/skills/paper-pipeline/scripts/ledger.mjs record argument-arc <paper-dir> FINDING <count> <report-path>
|
|
148
|
+
node .claude/skills/paper-pipeline/scripts/ledger.mjs record argument-arc <paper-dir> ABSTAINED <reason> "<one line>"
|
|
149
|
+
\`\`\`
|
|
150
|
+
|
|
151
|
+
🔴 **There is no PASS.** "The arc carries the reader" is not a value this pass can write down; it is
|
|
152
|
+
what a reader may conclude from the absence of findings, and the concluding is theirs.
|
|
153
|
+
|
|
154
|
+
**FINDING** — the arc is broken; \`<count>\` is the number of sections the new outline moves, merges
|
|
155
|
+
or deletes, and \`<report-path>\` is that outline. Add \`--blocking\` when there is no conclusion
|
|
156
|
+
sentence to build an arc from: that refusal is a finding about the draft, not an aborted run, and
|
|
157
|
+
\`<report-path>\` is the one-line statement of what is missing.
|
|
158
|
+
**ABSTAINED** — \`no-witness\`: the outline was read end to end and nothing moved. \`blocked\`: there
|
|
159
|
+
was no outline to read at all.
|
|
160
|
+
|
|
161
|
+
An \`ABSTAINED no-witness\` on a draft the author has twice called a pile is a missed finding, and
|
|
162
|
+
\`status.mjs\` will start asking about a check that has only ever abstained.
|
|
163
|
+
|
|
164
|
+
## Rules
|
|
165
|
+
|
|
166
|
+
- **The arc is judged on the outline, never on the draft.** If you find yourself rereading paragraphs
|
|
167
|
+
to decide whether a section belongs, you are grading prose again.
|
|
168
|
+
- **Merge is the default remedy, deletion the second.** Most arc defects are two sections doing one
|
|
169
|
+
job, not a section doing nothing.
|
|
170
|
+
- **Never answer a structural objection with a local edit.** Moving a paragraph and reporting "done"
|
|
171
|
+
is the exact failure this skill was built from.
|
|
172
|
+
- **The author's confusion is data about the paper, not about the author.** If they ask twice what a
|
|
173
|
+
section is for, the section is the problem.
|
|
174
|
+
- **Say what the reader must hold in their head at each point.** If that set only grows, the arc is
|
|
175
|
+
a pile.
|
|
176
|
+
|
|
177
|
+
## Compose with
|
|
178
|
+
|
|
179
|
+
- **Before** \`tighten-paper\` (length/sag) and \`grade-paper-writing\` (sentences/stalls) — both assume
|
|
180
|
+
the arc holds. Their verdicts are the gate inputs \`pc-panel-review\` demands.
|
|
181
|
+
- **After** any refutation pass that killed a load-bearing claim (see the paper's \`CLAIMS.md\`).
|
|
182
|
+
- \`sweep-design-space\` when the outline shows the paper has no distinctive move to make — that is a
|
|
183
|
+
design problem, not an arc problem.
|
|
184
|
+
|
|
185
|
+
## Provenance
|
|
186
|
+
|
|
187
|
+
Built 2026-07-30 after a full-day rewrite of \`the reference paper\` in which the author said the same
|
|
188
|
+
thing five times — *"does not guide the reader sequentially through the ideas, but simply throws a bunch of
|
|
189
|
+
different ideas, references, and benchmarks at them"* — and got five local edits in reply. The pipeline had a skill
|
|
190
|
+
for length, a skill for sentences and a skill for defects; the question *does this paper carry the
|
|
191
|
+
reader to one conclusion* belonged to nobody, and that is the one that failed.`,
|
|
192
|
+
});
|
|
@@ -0,0 +1,19 @@
|
|
|
1
|
+
/**
|
|
2
|
+
* argument-arc — the PAID tier: does this skill's description actually fire?
|
|
3
|
+
*
|
|
4
|
+
* COLOCATED ON PURPOSE (vigiles decides coverage by placement as of 2026-08-11).
|
|
5
|
+
* The prompts live in `.claude/lib/skill-trigger-cases.mjs` so all 21 cases
|
|
6
|
+
* are reviewed as one table where collisions between siblings are visible;
|
|
7
|
+
* copying them here would recreate the drift that rule exists to prevent.
|
|
8
|
+
*
|
|
9
|
+
* Measures recall (fires on its own territory) AND precision (stays quiet on a
|
|
10
|
+
* colliding sibling's territory), against the REAL `.claude` harness so the skill
|
|
11
|
+
* competes with every other installed description — an isolated run overstates
|
|
12
|
+
* recall and understates false positives.
|
|
13
|
+
*
|
|
14
|
+
* Costs money; not CI.
|
|
15
|
+
* node .claude/skills/argument-arc/argument-arc.eval.mjs [trials]
|
|
16
|
+
*/
|
|
17
|
+
import { runSkillTriggerEval } from "../../lib/skill-eval-kit.mjs";
|
|
18
|
+
|
|
19
|
+
await runSkillTriggerEval("argument-arc");
|
|
@@ -0,0 +1,23 @@
|
|
|
1
|
+
/**
|
|
2
|
+
* argument-arc — the free, deterministic tier. No model, no network.
|
|
3
|
+
*
|
|
4
|
+
* COLOCATED ON PURPOSE. vigiles decides coverage by PLACEMENT as of 2026-08-11:
|
|
5
|
+
* a test that merely names a surface no longer counts, because that tier was
|
|
6
|
+
* crediting surfaces nothing touched. So each skill needs a file inside its own
|
|
7
|
+
* directory — this one.
|
|
8
|
+
*
|
|
9
|
+
* The assertions live in `.claude/lib/skill-checks.mjs` and are CALLED here with
|
|
10
|
+
* this skill's name. They are not copied: 22 copies of the same checks is the drift that
|
|
11
|
+
* module exists to avoid. (Until 2026-08-11 this was an env-var side channel into a
|
|
12
|
+
* 614-line file named after no surface; it is a function call now.)
|
|
13
|
+
*
|
|
14
|
+
* What this proves: this skill's frontmatter parses as strict YAML, its declared
|
|
15
|
+
* tool contract is sane, its pipeline wiring points at scripts that exist, and it
|
|
16
|
+
* announces/records under ITS OWN identity rather than a sibling's.
|
|
17
|
+
*
|
|
18
|
+
* What it does NOT prove: that the skill fires, or that its guidance produces a
|
|
19
|
+
* good result. Those need a real model — see `argument-arc.eval.mjs`.
|
|
20
|
+
*/
|
|
21
|
+
import { checkSkill } from "../../lib/skill-checks.mjs";
|
|
22
|
+
|
|
23
|
+
await checkSkill("argument-arc");
|
|
@@ -0,0 +1,213 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: build-benchmark
|
|
3
|
+
description: Design and run the empirical study behind a measurement paper, and ship a reviewer-proof reproduction artifact. Covers the study design (invert the pitfall you're critiquing), honest statistics (paired/Welch t, CIs, Bonferroni, small-n spread as a result, not noise), a structural bound that outlives the specific artifacts tested, and a self-checking artifact that recomputes every headline number and exits non-zero on drift. Use when the paper's contribution is a way to MEASURE something and you're building the evidence + the thing reviewers will run. Compose with research-ideate (upstream), draft-paper, pc-panel-review (its artifact-runner executes this), submit-paper.
|
|
4
|
+
context: fork
|
|
5
|
+
allowed-tools: [Read, Write, Edit, Grep, Glob, Bash, Agent]
|
|
6
|
+
---
|
|
7
|
+
|
|
8
|
+
<!-- vigiles:sha256:067ada6f8aa44d5b compiled from skills/build-benchmark/SKILL.md.spec.ts -->
|
|
9
|
+
|
|
10
|
+
# build-benchmark — the study + the artifact reviewers can run
|
|
11
|
+
|
|
12
|
+
## Run me
|
|
13
|
+
|
|
14
|
+
🔴 FIRST, before any other step:
|
|
15
|
+
|
|
16
|
+
```
|
|
17
|
+
node .claude/skills/paper-pipeline/scripts/announce.mjs build-benchmark <paper-dir>
|
|
18
|
+
```
|
|
19
|
+
|
|
20
|
+
An advisory pass cannot be observed failing — silence is both its error state and its normal
|
|
21
|
+
state — so starting is an event, and events get written down.
|
|
22
|
+
|
|
23
|
+
A measurement paper is only as strong as the artifact a reviewer can `cd` into and re-run. This skill
|
|
24
|
+
covers both halves: the study design that makes the finding true, and the self-checking artifact that
|
|
25
|
+
makes it *verifiable*. The contribution is the **method**, not the one tool you happened to test — build
|
|
26
|
+
so both survive review.
|
|
27
|
+
|
|
28
|
+
## 0. 🔴 READ THE STATED CLAIM FIRST — and refuse to design without it
|
|
29
|
+
|
|
30
|
+
**Before designing anything, open `<paper-dir>/PIPELINE-STATUS.md` and read row `frame`** — the one
|
|
31
|
+
paragraph saying what this paper will claim (written by `argument-arc` in **frame mode**, in SETUP).
|
|
32
|
+
|
|
33
|
+
🔴 **If `frame` is empty, stop. Do not design the study.** Go run `argument-arc` frame mode, get the
|
|
34
|
+
paragraph, then come back. This is a refusal, not a recommendation:
|
|
35
|
+
|
|
36
|
+
- **The frame decides which results matter.** A study designed without a stated claim measures
|
|
37
|
+
whatever the available harness can already see, and the claim gets fitted afterwards to whatever
|
|
38
|
+
came out. That is backwards, and it is expensive in exactly one direction — **reframing after the
|
|
39
|
+
data is collected is how runs get thrown away**, observed on `compile-rules-2026`.
|
|
40
|
+
- **The claim is what makes §1 answerable.** "What would make the critiqued pitfall impossible here?"
|
|
41
|
+
has no answer until you have written down what you are claiming.
|
|
42
|
+
- **It costs a paragraph and it saves a study.** There is no version of this trade that favours
|
|
43
|
+
starting the runs first.
|
|
44
|
+
|
|
45
|
+
Cross-check the claim against the prevention rule in the `CLAUDE.md` beside the papers: if the claim is
|
|
46
|
+
about *preventing* something, write the one sentence saying **what physically stops the bad outcome in
|
|
47
|
+
the experimental arm**. "There is different text in it" means you are measuring persuasion, not
|
|
48
|
+
prevention, and the design is wrong before a single run.
|
|
49
|
+
|
|
50
|
+
## 1. Design that inverts the pitfall you're critiquing
|
|
51
|
+
|
|
52
|
+
The strongest measurement papers are a corrected version of the mistake they name. Pick the design that
|
|
53
|
+
is the *inverse* of the flaw:
|
|
54
|
+
|
|
55
|
+
- **Critiquing a proxy metric?** Measure the real thing, gated on correctness. The cost study is a
|
|
56
|
+
**cost-aware, correctness-gated A/B**: two arms × two tools, every run scored pass/fail first, then
|
|
57
|
+
the *dollar* bill compared — because the whole point was that the token-proxy and the dollar disagree.
|
|
58
|
+
Never let a cheaper-but-broken run count as a win; correctness is the gate before cost is even read.
|
|
59
|
+
- **Measuring something that runs untrusted/destructive code?** **Parse, don't execute.** The guard eval
|
|
60
|
+
is a **SAFE STATIC evaluation**: transcribe each guard's predicate by hand and reason over it against a
|
|
61
|
+
disaster battery — never run the guard or the attack. See Safety below; this is non-negotiable.
|
|
62
|
+
|
|
63
|
+
Write the design as the answer to "what would make the critiqued pitfall impossible here?"
|
|
64
|
+
|
|
65
|
+
## 2. Honest statistics — the part reviewers attack first
|
|
66
|
+
|
|
67
|
+
- **Paired where the design is paired; Welch where variances differ.** Use a **paired-t** when the same
|
|
68
|
+
task is run under both arms (blocks task difficulty); use **Welch's t** for unequal-variance unpaired
|
|
69
|
+
comparisons. Do NOT label a test "paired" unless the pairing is real — a mislabeled test is a blocker.
|
|
70
|
+
- **Report CIs and k/n, never bare percentages over tiny n.** "2/10 guards covered" beats "20% coverage";
|
|
71
|
+
"median 2/10 (range 0–5)" beats a single mean. A percentage over n=10 hides the n.
|
|
72
|
+
- **Report the SOURCE CONCENTRATION of a mined corpus — the single-source share is a number, not a
|
|
73
|
+
caveat.** For any corpus mined from several repositories / trackers / vendors, compute and print
|
|
74
|
+
**n per source and the max single-source share**, and put it in the paper (a table row or one
|
|
75
|
+
sentence), not only in Threats. A corpus advertised as covering *k* sources but dominated by one is
|
|
76
|
+
a top reason a mined-corpus paper gets rejected, and it is invisible to every other check — the
|
|
77
|
+
totals, the percentages and the internal consistency all pass, because the cut was simply never
|
|
78
|
+
made. **Mechanical leg:** the artifact emits `source, n, share` for the corpus and the top share
|
|
79
|
+
appears in the paper. If the top source exceeds ~50%, say so in the abstract's scope or narrow the
|
|
80
|
+
claimed population to what you actually sampled.
|
|
81
|
+
- **Inter-rater agreement is void if the rater was trained by, and scored against, the codebook's
|
|
82
|
+
author.** Agreement with the person who wrote the scheme measures *trainability*, not construct
|
|
83
|
+
reliability — it is the human form of the failure this suite already names in code (a checker and
|
|
84
|
+
its self-test written in one pass by one model). Requirements, all mechanical: (i) reference labels
|
|
85
|
+
come from a rater who did **not** author the codebook and did **not** see the hypothesis; (ii)
|
|
86
|
+
report the agreement **denominator as a share of the full corpus** — "κ=0.93 on 69 of 547 (12.6%)",
|
|
87
|
+
never a bare κ; (iii) use a coefficient that matches the design — Cohen's κ is single-label, so
|
|
88
|
+
multi-label coding needs per-label κ or Krippendorff's α; (iv) treat **κ = 1.00 as a red flag to
|
|
89
|
+
investigate, not a result to report** — perfect agreement on a many-category scheme usually means
|
|
90
|
+
the validation set was easy or calibration was de-facto joint coding.
|
|
91
|
+
- **Correct for the family.** Multiple comparisons across a family of tasks/tools → **Bonferroni** (or
|
|
92
|
+
state the correction you used). Report it; don't p-hack the one significant cell.
|
|
93
|
+
- **Small-n spread is a first-class result, not noise.** If five trials of the same task swing wildly,
|
|
94
|
+
that variance IS the finding (the thing under test is unstable) — report it, don't average it away.
|
|
95
|
+
- Repeated trials: fix trial count up front (e.g. per-task × trials × arms × tools), report the full n,
|
|
96
|
+
and treat every run — including failures — as data.
|
|
97
|
+
|
|
98
|
+
## 3. A structural bound that outlives the artifacts
|
|
99
|
+
|
|
100
|
+
Numbers about today's tools rot. A **mechanism or structural bound** doesn't. Give the paper one claim
|
|
101
|
+
that holds regardless of which specific tool/version you measured:
|
|
102
|
+
|
|
103
|
+
- the cost study's bound: token-efficiency and dollar cost are decoupled by pricing structure, so a
|
|
104
|
+
token-optimizing tool *cannot* be assumed to cut the bill — a property of the pricing, not the tool.
|
|
105
|
+
- the guard eval's axes: **mutation-evasion** (does a trivial rephrase of the attack slip the guard?),
|
|
106
|
+
**held-out generalization** (does coverage transfer to commands the guard wasn't written for?), and
|
|
107
|
+
**per-step causal ablation** (remove one guard step, measure the coverage delta — which step actually
|
|
108
|
+
does the work). These are properties of the *defense class*, reusable against the next guard set.
|
|
109
|
+
|
|
110
|
+
State the bound explicitly; it's what makes the paper a benchmark and not a product review.
|
|
111
|
+
|
|
112
|
+
## 4. Build a SELF-CHECKING artifact
|
|
113
|
+
|
|
114
|
+
This is the single biggest accept-probability lever, and pc-panel-review's artifact-runner WILL execute
|
|
115
|
+
it. Do **not** restate the requirements here — build to the canonical checklist:
|
|
116
|
+
|
|
117
|
+
→ **`paper-pipeline/references/artifact-checklist.md`** (self-recompute from raw data, exit non-zero on
|
|
118
|
+
drift, stdlib-only, no network, a README that maps each paper-number to where it prints, LICENSE, honest
|
|
119
|
+
badge scope, de-anon scanned).
|
|
120
|
+
|
|
121
|
+
The load-bearing property: each script **recomputes** every headline number from raw data and
|
|
122
|
+
**self-asserts** — exit non-zero the moment a cell drifts from the paper (`reproduce.py` on the cost
|
|
123
|
+
study; `evaluate.py`/`mutate.py`/`reproduce.mjs`/`ablation.mjs` on the guard eval). Derive from base
|
|
124
|
+
fields, never echo a stored/precomputed field — a reviewer calls that circular and they're right. For
|
|
125
|
+
de-anonymization before hosting, follow **`paper-pipeline/references/anonymization.md`**; host per
|
|
126
|
+
`submit-paper` (OSF anonymized view-only link; automation at `papers/osf_upload.py`).
|
|
127
|
+
|
|
128
|
+
## Safety — never execute untrusted or destructive commands to measure them
|
|
129
|
+
|
|
130
|
+
If the object of study is code that could delete, exfiltrate, or run attacker-controlled input, you
|
|
131
|
+
**transcribe and parse its predicate** — you never execute it, and you never run the attack against it.
|
|
132
|
+
The guard eval scores 46 real command-guards by reading each guard's logic against a 10-command disaster
|
|
133
|
+
battery statically; nothing in the battery is ever run. A benchmark that has to detonate the payload to
|
|
134
|
+
score it is a liability, not evidence. Same rule inside the artifact: no network, no shelling out to the
|
|
135
|
+
thing under test.
|
|
136
|
+
|
|
137
|
+
## Validate the transcription against real-code execution, not just a blind re-derivation
|
|
138
|
+
|
|
139
|
+
When you score third-party code by **transcribing its logic** (e.g. a guard's regexes) into your own
|
|
140
|
+
evaluator, a *blind re-derivation* (a second person re-reads the code and re-transcribes it) only checks
|
|
141
|
+
transcription-vs-human-reading — it does **not** check transcription-vs-real-behavior. There's a cheap,
|
|
142
|
+
zero-execution-risk way to close that gap: **run the actual scraped code against its REAL agent input
|
|
143
|
+
contract** and reconcile its live verdict with your static prediction.
|
|
144
|
+
|
|
145
|
+
- **Feed the real contract, not a bare string.** For Claude Code PreToolUse hooks, the command arrives as
|
|
146
|
+
a **JSON envelope on stdin** with the command at `.tool_input.command`. Feed each battery/benign command
|
|
147
|
+
as a **string inside that envelope** to the real guard's stdin and read its verdict (exit 2 = block; or
|
|
148
|
+
stdout JSON `permissionDecision` deny/ask/block). The command is **never executed** — the guard only
|
|
149
|
+
inspects the string.
|
|
150
|
+
- **SAFETY: sandbox + stub-PATH first.** Run inside a throwaway sandbox with a **stub PATH** (fake
|
|
151
|
+
`rm`/`dd`/`git`/`curl`/… → `exit 0`) as defense-in-depth, so even a guard that shells out can't do
|
|
152
|
+
damage. Invoke the real interpreters by absolute path.
|
|
153
|
+
- **Reconcile per code×input pair.** Every mismatch between live verdict and static prediction is a
|
|
154
|
+
**finding** — either a transcription nit to fix, or a real behavior your static model structurally
|
|
155
|
+
cannot capture.
|
|
156
|
+
|
|
157
|
+
**Concrete payoff (the motivating example):** this method surfaced a **"whole-stdin confound"** — a
|
|
158
|
+
widely-copied guard that greps its **entire raw stdin** rather than the extracted command. The JSON
|
|
159
|
+
envelope's trailing bytes defeated its anchored `rm …/$` rule, so it **FAILED TO BLOCK `rm -rf /`** live,
|
|
160
|
+
even though it "blocked" the bare string. A pure transcription/static model can't see this; only
|
|
161
|
+
real-code-through-real-contract execution reveals it.
|
|
162
|
+
|
|
163
|
+
Before submitting, verify the artifact you built with the reviewer-side reproduction protocol:
|
|
164
|
+
**`references/adversarial-cold-repro.md`** (cold re-run, three-way reconcile, robustness attacks).
|
|
165
|
+
|
|
166
|
+
## Record the verdict
|
|
167
|
+
|
|
168
|
+
🔴 LAST step, once the deliverable exists:
|
|
169
|
+
|
|
170
|
+
```
|
|
171
|
+
node .claude/skills/paper-pipeline/scripts/ledger.mjs record build-benchmark <paper-dir> FINDING <count> <report-path>
|
|
172
|
+
node .claude/skills/paper-pipeline/scripts/ledger.mjs record build-benchmark <paper-dir> ABSTAINED <reason> "<one line>"
|
|
173
|
+
```
|
|
174
|
+
|
|
175
|
+
🔴 **There is no PASS.** An artifact that recomputes every headline number and exits 0 has produced
|
|
176
|
+
no finding; it has not certified the study. Record the absence, do not name it a success.
|
|
177
|
+
|
|
178
|
+
**FINDING** — `<count>` numbers did not reproduce, or reproduce only with caveats; `<report-path>`
|
|
179
|
+
is the artifact output. Add `--blocking` for the §0 refusal: `frame` is empty, so the study was not
|
|
180
|
+
designed.
|
|
181
|
+
**ABSTAINED** — `no-witness`: the artifact ran and every number reproduced. `input-missing`: the
|
|
182
|
+
raw data is not in the repo. `crashed`: the harness itself died.
|
|
183
|
+
|
|
184
|
+
🔴 That §0 refusal is a **finding, not an aborted run** — it is a fact about the work, so it is a
|
|
185
|
+
FINDING and not an abstention. Record it and stop. A study that was never designed leaves exactly
|
|
186
|
+
the same silence as one that is still running, and the difference costs a week to rediscover.
|
|
187
|
+
|
|
188
|
+
🔴 **THREE MECHANICAL CHECKS FILE UNDER THIS SKILL** and each has its own ledger row:
|
|
189
|
+
`build-benchmark/check-provenance`, `build-benchmark/arm-permutation`,
|
|
190
|
+
`build-benchmark/delivered-pdf` (see `.claude/skills/paper-pipeline/scripts/run-mechanical.mjs`). Until 2026-08-10 the
|
|
191
|
+
ledger keyed on the SKILL, so a clean run of one erased a finding of another from every derived
|
|
192
|
+
view — which is why two of the three sat unwired. This block records the JUDGEMENT pass, under the
|
|
193
|
+
bare key `build-benchmark`; never file a mechanical result here.
|
|
194
|
+
|
|
195
|
+
## Compose with
|
|
196
|
+
- **argument-arc (frame mode)** — 🔴 **hard input.** Owns row `frame`, the stated claim. No `frame`, no study.
|
|
197
|
+
- **research-ideate** (upstream) — supplies the contribution and the pitfall to invert; this skill turns
|
|
198
|
+
it into evidence.
|
|
199
|
+
- **draft-paper** — the study's stats, bound, and artifact-number table feed Methods/Results/Availability.
|
|
200
|
+
- **pc-panel-review** — its artifact-runner reviewer `cd`s in and re-executes every harness; build so
|
|
201
|
+
that pass is deterministic.
|
|
202
|
+
- **submit-paper** — hosts the scrubbed artifact anonymously and links it in Availability.
|
|
203
|
+
|
|
204
|
+
## Provenance
|
|
205
|
+
- **AgenticDev 2026 @ ASE — "Measuring the Wrong Number":** a cost-aware, correctness-gated A/B harness
|
|
206
|
+
over 140 runs (7 tasks × 5 trials × 2 arms × 2 tools); Welch t + paired-t CIs + Bonferroni; the finding
|
|
207
|
+
that token-efficiency tools don't cut the dollar bill. Artifact = a stdlib-only `reproduce.py` that
|
|
208
|
+
recomputes every headline number from raw data and exits non-zero on drift.
|
|
209
|
+
- **AISec 2026 @ ACM CCS — "Safety Theater":** a 10-command disaster battery against 46 real
|
|
210
|
+
command-guards, SAFE STATIC evaluation (transcribe each predicate, never execute), median coverage
|
|
211
|
+
2/10, plus mutation-evasion, held-out generalization, and per-step causal ablation axes. Artifact =
|
|
212
|
+
`evaluate.py`/`mutate.py`/`reproduce.mjs`/`ablation.mjs`, each self-asserting on run. Both artifacts
|
|
213
|
+
are hosted anonymized on OSF.
|