pi-dev-team 0.2.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/LICENSE +21 -0
- package/PORTING.md +134 -0
- package/README.md +207 -0
- package/UPSTREAM.json +64 -0
- package/agents/Explore.md +15 -0
- package/agents/a11y-review.md +118 -0
- package/agents/adr-author.md +70 -0
- package/agents/ai-provenance-review.md +120 -0
- package/agents/angular-reactivity-review.md +95 -0
- package/agents/arch-review.md +135 -0
- package/agents/architect.md +78 -0
- package/agents/autoship-batch-proposer.md +69 -0
- package/agents/claude-setup-review.md +136 -0
- package/agents/codebase-recon.md +184 -0
- package/agents/component-architecture-review.md +119 -0
- package/agents/concurrency-review.md +109 -0
- package/agents/correctness-review.md +290 -0
- package/agents/data-flow-tracer.md +120 -0
- package/agents/doc-review.md +165 -0
- package/agents/domain-review.md +136 -0
- package/agents/general-purpose.md +10 -0
- package/agents/gherkin-quality-critic.md +113 -0
- package/agents/js-fp-review.md +114 -0
- package/agents/mutation-kill.md +684 -0
- package/agents/naming-review.md +142 -0
- package/agents/orchestrator.md +339 -0
- package/agents/performance-review.md +105 -0
- package/agents/plan-review-acceptance.md +115 -0
- package/agents/plan-review-design.md +90 -0
- package/agents/plan-review-parallelization.md +84 -0
- package/agents/plan-review-strategic.md +96 -0
- package/agents/plan-review-ux.md +110 -0
- package/agents/platform-engineer.md +64 -0
- package/agents/product-manager.md +68 -0
- package/agents/progress-guardian.md +79 -0
- package/agents/qa-engineer.md +289 -0
- package/agents/quality-reviewer.md +132 -0
- package/agents/react-reactivity-review.md +102 -0
- package/agents/refactor-opportunity-review.md +128 -0
- package/agents/security-engineer.md +60 -0
- package/agents/security-review.md +218 -0
- package/agents/session-analysis.md +95 -0
- package/agents/software-engineer.md +105 -0
- package/agents/spec-compliance-review.md +100 -0
- package/agents/spec-reviewer.md +114 -0
- package/agents/structure-review.md +146 -0
- package/agents/tech-writer.md +84 -0
- package/agents/test-review.md +246 -0
- package/agents/test-smell-review.md +188 -0
- package/agents/token-efficiency-review.md +139 -0
- package/agents/ui-ux-designer.md +54 -0
- package/agents/vue-reactivity-review.md +95 -0
- package/bin/__pycache__/claudecpython-314.pyc +0 -0
- package/bin/claude +258 -0
- package/docs/upstream/.pages +1 -0
- package/docs/upstream/CHANGELOG.md +2586 -0
- package/docs/upstream/README.md +155 -0
- package/docs/upstream/agent-architecture.md +214 -0
- package/docs/upstream/agent_info.md +187 -0
- package/docs/upstream/artifact-migration.md +124 -0
- package/docs/upstream/code-intelligence-nudge.md +149 -0
- package/docs/upstream/code-review-process.md +294 -0
- package/docs/upstream/concurrent-use.md +73 -0
- package/docs/upstream/context-management.md +111 -0
- package/docs/upstream/developer-notes.md +280 -0
- package/docs/upstream/diagrams/architecture-overview.svg +101 -0
- package/docs/upstream/diagrams/review-dispatch.svg +139 -0
- package/docs/upstream/diagrams/team-agents.svg +128 -0
- package/docs/upstream/diagrams/test-improve-flow.svg +166 -0
- package/docs/upstream/diagrams/workflow-linear.svg +66 -0
- package/docs/upstream/diagrams/workflow-three-phase.svg +200 -0
- package/docs/upstream/eval-maintenance.md +95 -0
- package/docs/upstream/eval-running-guide.md +147 -0
- package/docs/upstream/eval-system.md +291 -0
- package/docs/upstream/session-review-oss-complements.md +75 -0
- package/docs/upstream/session-review.md +212 -0
- package/docs/upstream/skills.md +188 -0
- package/docs/upstream/team-structure.md +21 -0
- package/docs/upstream/telemetry-ci-access.md +129 -0
- package/docs/upstream/telemetry-repo-security.md +120 -0
- package/docs/upstream/test-evaluation.md +277 -0
- package/docs/upstream/test-improve.md +154 -0
- package/docs/upstream/triage-workflow.md +282 -0
- package/docs/upstream/workflows.md +289 -0
- package/extensions/dev-team/index.ts +539 -0
- package/extensions/dev-team/lib/agents.ts +272 -0
- package/extensions/dev-team/lib/ai-credits.ts +92 -0
- package/extensions/dev-team/lib/autocompact.ts +81 -0
- package/extensions/dev-team/lib/child-run.ts +102 -0
- package/extensions/dev-team/lib/config.ts +236 -0
- package/extensions/dev-team/lib/gh-command.ts +103 -0
- package/extensions/dev-team/lib/github-style.ts +307 -0
- package/extensions/dev-team/lib/hooks.ts +350 -0
- package/extensions/dev-team/lib/metrics.ts +115 -0
- package/extensions/dev-team/lib/safe-read.ts +49 -0
- package/extensions/dev-team/lib/session-files.ts +57 -0
- package/extensions/dev-team/lib/session-spend.ts +123 -0
- package/extensions/dev-team/lib/shell-scan.ts +205 -0
- package/extensions/dev-team/lib/skills.ts +213 -0
- package/extensions/dev-team/lib/subagent-render.ts +245 -0
- package/extensions/dev-team/lib/subagent-types.ts +164 -0
- package/extensions/dev-team/lib/subagent.ts +596 -0
- package/extensions/dev-team/lib/terminal-text.ts +54 -0
- package/extensions/dev-team/lib/tools-misc.ts +152 -0
- package/extensions/dev-team/lib/transcript.ts +110 -0
- package/extensions/dev-team/lib/trust.ts +52 -0
- package/extensions/dev-team/lib/usage-breakdown.ts +176 -0
- package/extensions/dev-team/lib/usage-chart.ts +153 -0
- package/extensions/dev-team/lib/usage-command.ts +107 -0
- package/extensions/dev-team/lib/usage-history.ts +203 -0
- package/extensions/dev-team/lib/usage-render.ts +225 -0
- package/extensions/dev-team/lib/usage-split-bar.ts +127 -0
- package/extensions/dev-team/lib/usage-state.ts +116 -0
- package/extensions/dev-team/lib/usage-text.ts +159 -0
- package/extensions/dev-team/lib/usage-view.ts +109 -0
- package/hooks/__pycache__/refactor_test_freeze_guard.cpython-314.pyc +0 -0
- package/hooks/agent_dispatch_ledger.py +190 -0
- package/hooks/autocompact_setup_nudge.py +99 -0
- package/hooks/bash_retry_guard.py +228 -0
- package/hooks/boundary_events_write_guard.py +352 -0
- package/hooks/code_intelligence_nudge.py +293 -0
- package/hooks/code_intelligence_turn_mark.py +317 -0
- package/hooks/codegraph_bootstrap.py +139 -0
- package/hooks/contract_version_guard.py +362 -0
- package/hooks/cost_meter.py +106 -0
- package/hooks/destructive-commands.json +62 -0
- package/hooks/destructive_guard.py +477 -0
- package/hooks/eval_compliance_check.py +440 -0
- package/hooks/guards.json +17 -0
- package/hooks/hooks.json +323 -0
- package/hooks/internal_double_gate.py +296 -0
- package/hooks/js_fp_review.py +212 -0
- package/hooks/knowledge_index.py +119 -0
- package/hooks/lib/__pycache__/artifact_paths.cpython-314.pyc +0 -0
- package/hooks/lib/__pycache__/atomic_state.cpython-314.pyc +0 -0
- package/hooks/lib/__pycache__/autocompact_config.cpython-314.pyc +0 -0
- package/hooks/lib/__pycache__/boundary_events.cpython-314.pyc +0 -0
- package/hooks/lib/__pycache__/doc_classification.cpython-314.pyc +0 -0
- package/hooks/lib/__pycache__/gh_pr_create_detect.cpython-314.pyc +0 -0
- package/hooks/lib/__pycache__/git_safe_diff.cpython-314.pyc +0 -0
- package/hooks/lib/__pycache__/instrument_log.cpython-314.pyc +0 -0
- package/hooks/lib/__pycache__/metrics_query.cpython-314.pyc +0 -0
- package/hooks/lib/__pycache__/plugin_version.cpython-314.pyc +0 -0
- package/hooks/lib/__pycache__/pre_commit_doc_classifier.cpython-314.pyc +0 -0
- package/hooks/lib/__pycache__/review_agent_registry.cpython-314.pyc +0 -0
- package/hooks/lib/__pycache__/review_gate_corroboration.cpython-314.pyc +0 -0
- package/hooks/lib/__pycache__/review_gate_hash.cpython-314.pyc +0 -0
- package/hooks/lib/__pycache__/review_verdicts.cpython-314.pyc +0 -0
- package/hooks/lib/__pycache__/stdin_json.cpython-314.pyc +0 -0
- package/hooks/lib/__pycache__/stryker_invocation.cpython-314.pyc +0 -0
- package/hooks/lib/__pycache__/telemetry_consent.cpython-314.pyc +0 -0
- package/hooks/lib/__pycache__/test_file_classify.cpython-314.pyc +0 -0
- package/hooks/lib/__pycache__/token_efficiency_limits.cpython-314.pyc +0 -0
- package/hooks/lib/__pycache__/verify_guard_state.cpython-314.pyc +0 -0
- package/hooks/lib/__pycache__/xunit_v3_operator_gate.cpython-314.pyc +0 -0
- package/hooks/lib/agent_skill_hints.py +74 -0
- package/hooks/lib/artifact_paths.py +263 -0
- package/hooks/lib/atomic_state.py +557 -0
- package/hooks/lib/autocompact_config.py +103 -0
- package/hooks/lib/autoship_log.py +106 -0
- package/hooks/lib/banned_scripts_policy.py +51 -0
- package/hooks/lib/boundary_events.py +436 -0
- package/hooks/lib/build_knowledge_index.py +504 -0
- package/hooks/lib/build_skills_index.py +361 -0
- package/hooks/lib/build_state.py +116 -0
- package/hooks/lib/classify_ship_outcome.py +126 -0
- package/hooks/lib/config_changelog_schema.py +115 -0
- package/hooks/lib/cost_meter.py +955 -0
- package/hooks/lib/doc_classification.py +116 -0
- package/hooks/lib/gh_pr_create_detect.py +136 -0
- package/hooks/lib/git_safe_diff.py +123 -0
- package/hooks/lib/instrument_log.py +66 -0
- package/hooks/lib/iteration_journal_gate.py +197 -0
- package/hooks/lib/knowledge_index_paths.py +88 -0
- package/hooks/lib/mcp_json_repowise.py +177 -0
- package/hooks/lib/metrics_query.py +202 -0
- package/hooks/lib/minimal_yaml.py +434 -0
- package/hooks/lib/plugin_version.py +142 -0
- package/hooks/lib/pre_commit_detect.py +537 -0
- package/hooks/lib/pre_commit_doc_classifier.py +126 -0
- package/hooks/lib/pricing.py +118 -0
- package/hooks/lib/report_pdf.py +371 -0
- package/hooks/lib/review_agent_registry.py +142 -0
- package/hooks/lib/review_dispatch_ledger.py +101 -0
- package/hooks/lib/review_gate_corroboration.py +521 -0
- package/hooks/lib/review_gate_hash.py +252 -0
- package/hooks/lib/review_gate_normalized_hash.py +1115 -0
- package/hooks/lib/review_verdicts.py +301 -0
- package/hooks/lib/run_report.py +160 -0
- package/hooks/lib/skill_categories.yaml +125 -0
- package/hooks/lib/stdin_json.py +57 -0
- package/hooks/lib/stryker_invocation.py +102 -0
- package/hooks/lib/telemetry_consent.py +41 -0
- package/hooks/lib/telemetry_report.py +108 -0
- package/hooks/lib/test_file_classify.py +160 -0
- package/hooks/lib/token_efficiency_limits.py +51 -0
- package/hooks/lib/turn_identity.py +77 -0
- package/hooks/lib/verify_guard_state.py +110 -0
- package/hooks/lib/workflow_state.py +206 -0
- package/hooks/lib/xunit_v3_operator_gate.py +596 -0
- package/hooks/mcp_json_repowise_nudge.py +74 -0
- package/hooks/mutation_adapters/__init__.py +7 -0
- package/hooks/mutation_adapters/__pycache__/__init__.cpython-314.pyc +0 -0
- package/hooks/mutation_adapters/__pycache__/lib.cpython-314.pyc +0 -0
- package/hooks/mutation_adapters/__pycache__/mutmut.cpython-314.pyc +0 -0
- package/hooks/mutation_adapters/__pycache__/pitest.cpython-314.pyc +0 -0
- package/hooks/mutation_adapters/__pycache__/stryker.cpython-314.pyc +0 -0
- package/hooks/mutation_adapters/__pycache__/stryker_net.cpython-314.pyc +0 -0
- package/hooks/mutation_adapters/lib.py +478 -0
- package/hooks/mutation_adapters/mutmut.py +188 -0
- package/hooks/mutation_adapters/pitest.py +266 -0
- package/hooks/mutation_adapters/stryker.py +157 -0
- package/hooks/mutation_adapters/stryker_net.py +264 -0
- package/hooks/mutation_gate.py +193 -0
- package/hooks/mutation_testing_smoke_gate.py +371 -0
- package/hooks/pending_review_notify.py +121 -0
- package/hooks/phase_marker.py +138 -0
- package/hooks/post_compact_state_reinject.py +180 -0
- package/hooks/post_format.py +115 -0
- package/hooks/pre_commit_knowledge_index.py +128 -0
- package/hooks/pre_commit_review.py +66 -0
- package/hooks/pre_pr_review.py +694 -0
- package/hooks/pre_tool_guard.py +405 -0
- package/hooks/py.sh +73 -0
- package/hooks/refactor-bash-write-patterns.json +29 -0
- package/hooks/refactor_test_bash_guard.py +253 -0
- package/hooks/refactor_test_freeze_guard.py +139 -0
- package/hooks/refactor_test_revert_guard.py +186 -0
- package/hooks/repo_review_nudge.py +287 -0
- package/hooks/review_verdict_recorder.py +464 -0
- package/hooks/scan_bash_command_for_banned_scripts.py +428 -0
- package/hooks/scan_worktree_for_banned_scripts.py +238 -0
- package/hooks/session_learning_trigger.py +248 -0
- package/hooks/skills_index.py +126 -0
- package/hooks/stryker_xunit_shim_guard.py +571 -0
- package/hooks/subagent_completion_guard.py +309 -0
- package/hooks/subagent_skill_context.py +139 -0
- package/hooks/task_completion_metrics.py +216 -0
- package/hooks/tdd_guard.py +229 -0
- package/hooks/telemetry.py +341 -0
- package/hooks/token_efficiency_review.py +194 -0
- package/hooks/verify_guard.py +183 -0
- package/hooks/verify_guard_edit_marker.py +73 -0
- package/hooks/version_check.py +173 -0
- package/knowledge/accepted-risks-schema.md +98 -0
- package/knowledge/adr-decision-criteria.md +64 -0
- package/knowledge/adversarial-review-protocol.md +139 -0
- package/knowledge/agent-registry.md +228 -0
- package/knowledge/agent-review-methodology.md +80 -0
- package/knowledge/ai-friendly-repo-guidelines.md +67 -0
- package/knowledge/architecture-assessment.md +96 -0
- package/knowledge/artifact-lifecycle.md +57 -0
- package/knowledge/cd-maturity-model.md +82 -0
- package/knowledge/cd-test-architecture.md +190 -0
- package/knowledge/ci-cd-file-scope.md +24 -0
- package/knowledge/codegraph-vs-graphify.md +192 -0
- package/knowledge/component-test-patterns.md +139 -0
- package/knowledge/database-change-management.md +80 -0
- package/knowledge/database-test-patterns.md +79 -0
- package/knowledge/decision-defaults.md +88 -0
- package/knowledge/dependency-breaking-techniques.md +116 -0
- package/knowledge/deployment-pipeline.md +86 -0
- package/knowledge/design-smells.md +122 -0
- package/knowledge/directory-enumeration.md +38 -0
- package/knowledge/domain-modeling.md +123 -0
- package/knowledge/evidence-bundle.md +90 -0
- package/knowledge/exploratory-testing-field-guide.md +122 -0
- package/knowledge/failure-routing.md +28 -0
- package/knowledge/fixture-construction.md +56 -0
- package/knowledge/frontend-component-architecture.md +139 -0
- package/knowledge/gherkin-quality-review-dispatch.md +135 -0
- package/knowledge/index.json +6766 -0
- package/knowledge/internal-collaborator-doubling.md +101 -0
- package/knowledge/legacy-test-strategy.md +71 -0
- package/knowledge/long-run-waiting.md +66 -0
- package/knowledge/microservice-testing.md +71 -0
- package/knowledge/model-pricing.json +23 -0
- package/knowledge/mutation-score-formulas.md +60 -0
- package/knowledge/object-calisthenics.md +147 -0
- package/knowledge/oracle-provenance.md +94 -0
- package/knowledge/orchestrator-script-implementation.md +185 -0
- package/knowledge/owasp-detection.md +148 -0
- package/knowledge/plan-review-rubric.md +56 -0
- package/knowledge/proxy-connectivity.md +62 -0
- package/knowledge/reactive-effect-patterns.md +73 -0
- package/knowledge/recon-inventory-excludes.txt +32 -0
- package/knowledge/references/bdd-value-guide.md +61 -0
- package/knowledge/references/csharp-http-client-testing.md +264 -0
- package/knowledge/release-strategies.md +74 -0
- package/knowledge/report-output-location.md +117 -0
- package/knowledge/report-pdf-integration.md +63 -0
- package/knowledge/report-print.css +129 -0
- package/knowledge/report-template.md +114 -0
- package/knowledge/report-to-pdf.md +69 -0
- package/knowledge/request-processing-flow.md +63 -0
- package/knowledge/result-verification.md +52 -0
- package/knowledge/review-agent-output-contract.md +121 -0
- package/knowledge/review-lens-classification.md +113 -0
- package/knowledge/review-rubric.md +62 -0
- package/knowledge/review-template.md +104 -0
- package/knowledge/rule-fixtures/A02.insecure-random-js/negative.js +1 -0
- package/knowledge/rule-fixtures/A02.insecure-random-js/positive.js +1 -0
- package/knowledge/rule-fixtures/A02.weak-hashing-md5/negative.py +1 -0
- package/knowledge/rule-fixtures/A02.weak-hashing-md5/positive.py +1 -0
- package/knowledge/rule-fixtures/A03.command-injection/negative.js +1 -0
- package/knowledge/rule-fixtures/A03.command-injection/positive.js +1 -0
- package/knowledge/rule-fixtures/A03.sql-injection/negative.js +1 -0
- package/knowledge/rule-fixtures/A03.sql-injection/positive.js +1 -0
- package/knowledge/rule-fixtures/A03.xss-innerhtml/negative.js +1 -0
- package/knowledge/rule-fixtures/A03.xss-innerhtml/positive.js +1 -0
- package/knowledge/rule-fixtures/A05.cors-wildcard/negative.js +1 -0
- package/knowledge/rule-fixtures/A05.cors-wildcard/positive.js +1 -0
- package/knowledge/rule-fixtures/A05.default-credentials/negative.js +1 -0
- package/knowledge/rule-fixtures/A05.default-credentials/positive.js +1 -0
- package/knowledge/rule-fixtures/A07.jwt-alg-none/negative.js +1 -0
- package/knowledge/rule-fixtures/A07.jwt-alg-none/positive.js +1 -0
- package/knowledge/rule-fixtures/A08.binary-formatter/negative.cs +1 -0
- package/knowledge/rule-fixtures/A08.binary-formatter/positive.cs +1 -0
- package/knowledge/rule-fixtures/A08.js-eval/negative.js +1 -0
- package/knowledge/rule-fixtures/A08.js-eval/positive.js +1 -0
- package/knowledge/rule-fixtures/A08.object-input-stream/negative.java +1 -0
- package/knowledge/rule-fixtures/A08.object-input-stream/positive.java +1 -0
- package/knowledge/schemas/disposition-register-v1.json +65 -0
- package/knowledge/schemas/recon-envelope-v1.json +198 -0
- package/knowledge/schemas/unified-finding-v1.json +72 -0
- package/knowledge/security-primitives-contract.md +301 -0
- package/knowledge/security-review-rule-map.yaml +107 -0
- package/knowledge/skills-registry.md +72 -0
- package/knowledge/task-size-classifier.md +103 -0
- package/knowledge/telemetry-schema.md +881 -0
- package/knowledge/test-automation-maturity.md +56 -0
- package/knowledge/test-automation-principles.md +71 -0
- package/knowledge/test-cadence-tradeoffs.md +68 -0
- package/knowledge/test-doubles.md +105 -0
- package/knowledge/test-file-indicators.md +22 -0
- package/knowledge/test-layer-gates.md +35 -0
- package/knowledge/test-matrix-examples/django-batch.md +24 -0
- package/knowledge/test-matrix-examples/dotnet-grpc-fronting-api.md +90 -0
- package/knowledge/test-matrix-examples/dotnet-http-consumer.md +131 -0
- package/knowledge/test-matrix-examples/react-node-spa.md +24 -0
- package/knowledge/test-matrix-examples/spring-boot-service.md +25 -0
- package/knowledge/test-matrix-examples/ssr-htmx.md +24 -0
- package/knowledge/test-organization.md +70 -0
- package/knowledge/test-pyramid.md +84 -0
- package/knowledge/test-refactoring.md +67 -0
- package/knowledge/test-review-division-of-labor.md +85 -0
- package/knowledge/test-smells.md +80 -0
- package/knowledge/test-stack-profiles/bdd-frameworks.md +235 -0
- package/knowledge/test-stack-profiles/django.md +13 -0
- package/knowledge/test-stack-profiles/dotnet.md +18 -0
- package/knowledge/test-stack-profiles/go.md +16 -0
- package/knowledge/test-stack-profiles/node.md +16 -0
- package/knowledge/test-stack-profiles/react.md +12 -0
- package/knowledge/test-stack-profiles/spring-boot.md +16 -0
- package/knowledge/test-stack-profiles/ssr-htmx.md +14 -0
- package/knowledge/test-stack-profiles/vue.md +12 -0
- package/knowledge/test-strategy.md +70 -0
- package/knowledge/testability-patterns.md +240 -0
- package/knowledge/testing-quadrants.md +44 -0
- package/knowledge/testing-techniques/approval.md +15 -0
- package/knowledge/testing-techniques/chaos.md +17 -0
- package/knowledge/testing-techniques/fuzz.md +15 -0
- package/knowledge/testing-techniques/property-based.md +15 -0
- package/knowledge/testing-techniques/schema-validation.md +15 -0
- package/knowledge/testing-techniques/screenshot.md +15 -0
- package/knowledge/three-phase-workflow.md +198 -0
- package/knowledge/value-patterns.md +55 -0
- package/knowledge/verification-mode.md +116 -0
- package/knowledge/virtual-service-libraries.md +75 -0
- package/knowledge/wave-consolidation-guidance.md +21 -0
- package/overrides/agents/Explore.md +15 -0
- package/overrides/agents/general-purpose.md +10 -0
- package/overrides/notes/autoship.md +6 -0
- package/overrides/notes/issues-from-assessment.md +3 -0
- package/overrides/notes/issues-from-plan.md +3 -0
- package/overrides/notes/mutation-night-watch.md +3 -0
- package/overrides/notes/mutation-testing.md +3 -0
- package/overrides/notes/pr.md +7 -0
- package/overrides/notes/project-init.md +6 -0
- package/overrides/notes/setup.md +13 -0
- package/overrides/notes/specs.md +3 -0
- package/overrides/skills/headless-run/SKILL.md +45 -0
- package/overrides/skills/upgrade/SKILL.md +30 -0
- package/overrides/skills/version/SKILL.md +25 -0
- package/package.json +36 -0
- package/scripts/authoring_digest.py +93 -0
- package/scripts/autoship_discover.py +121 -0
- package/scripts/autoship_group.py +409 -0
- package/scripts/autoship_proposals.py +494 -0
- package/scripts/autoship_queue.py +291 -0
- package/scripts/autoship_reclaim.py +495 -0
- package/scripts/build_jobs.py +108 -0
- package/scripts/build_rollback_point.py +240 -0
- package/scripts/build_slice_scope.py +157 -0
- package/scripts/build_wave.py +109 -0
- package/scripts/build_wave_reconcile.py +252 -0
- package/scripts/build_worktree_baseref.py +113 -0
- package/scripts/check_agent_scope.py +117 -0
- package/scripts/check_agent_tool_mapping.py +213 -0
- package/scripts/check_review_agent_mcp_tools.py +317 -0
- package/scripts/check_security_assessment_mcp_tools.py +165 -0
- package/scripts/checkpoint_abort.py +502 -0
- package/scripts/claude_setup_review.py +438 -0
- package/scripts/codebase_recon.py +556 -0
- package/scripts/coverage_config.py +623 -0
- package/scripts/coverage_delta_steering.py +330 -0
- package/scripts/coverage_discovery_dotnet.py +315 -0
- package/scripts/coverage_discovery_java.py +742 -0
- package/scripts/coverage_discovery_js.py +546 -0
- package/scripts/coverage_gap_ranking.py +556 -0
- package/scripts/coverage_readiness.py +455 -0
- package/scripts/coverage_report_parse.py +521 -0
- package/scripts/detect_bdd_convention.py +252 -0
- package/scripts/eval_ablation.py +376 -0
- package/scripts/gherkin_analysis_coverage_gate.py +306 -0
- package/scripts/gherkin_cross_feature_duplicate_titles_gate.py +173 -0
- package/scripts/gherkin_effectiveness_rollup.py +238 -0
- package/scripts/gherkin_failure_path_gate.py +206 -0
- package/scripts/gherkin_feature_merge.py +720 -0
- package/scripts/gherkin_stub_gate.py +163 -0
- package/scripts/gherkin_stub_merge.py +479 -0
- package/scripts/git_origin_host.py +88 -0
- package/scripts/install-java-static-analysis.py +110 -0
- package/scripts/issue_deps.py +74 -0
- package/scripts/lib/_bdd_markers.py +28 -0
- package/scripts/lib/_gherkin_text.py +93 -0
- package/scripts/lib/_vendored_tree.py +70 -0
- package/scripts/lib/autoship_state.py +397 -0
- package/scripts/lib/claude_md_guard.py +226 -0
- package/scripts/lib/deterministic_recon.py +446 -0
- package/scripts/lib/mcp_tool_grants.py +211 -0
- package/scripts/lib/plan_parse.py +386 -0
- package/scripts/lib/review_result.py +84 -0
- package/scripts/lib/review_roster.py +86 -0
- package/scripts/lib/session_log/__init__.py +34 -0
- package/scripts/lib/session_log/__pycache__/__init__.cpython-314.pyc +0 -0
- package/scripts/lib/session_log/__pycache__/records.cpython-314.pyc +0 -0
- package/scripts/lib/session_log/classify.py +231 -0
- package/scripts/lib/session_log/corrections.py +194 -0
- package/scripts/lib/session_log/discovery.py +108 -0
- package/scripts/lib/session_log/records.py +218 -0
- package/scripts/lib/session_log/redact.py +76 -0
- package/scripts/lib/session_log/signals.py +373 -0
- package/scripts/lib/session_report_downstream.py +614 -0
- package/scripts/lib/session_report_maintainer.py +1273 -0
- package/scripts/lib/session_report_shared.py +262 -0
- package/scripts/lib/settings_hook_guard.py +157 -0
- package/scripts/lib/slug.py +33 -0
- package/scripts/lib/stub_extractors/__init__.py +82 -0
- package/scripts/lib/stub_extractors/_common.py +328 -0
- package/scripts/lib/stub_extractors/csharp.py +19 -0
- package/scripts/lib/stub_extractors/go.py +173 -0
- package/scripts/lib/stub_extractors/java.py +18 -0
- package/scripts/lib/stub_extractors/jsts.py +126 -0
- package/scripts/mutation_stack_sections.py +149 -0
- package/scripts/mutation_yield_steering.py +345 -0
- package/scripts/orchestrator.py +895 -0
- package/scripts/plan_gherkin_export.py +227 -0
- package/scripts/plan_waves.py +208 -0
- package/scripts/pr_close_keyword_lint.py +108 -0
- package/scripts/progress_guardian.py +888 -0
- package/scripts/recon_inventory.py +273 -0
- package/scripts/review_findings_log.py +93 -0
- package/scripts/run_invariants.py +124 -0
- package/scripts/select_lenses.py +640 -0
- package/scripts/session_report.py +486 -0
- package/scripts/set_autocompact_env.py +221 -0
- package/scripts/ship_resume_guard.py +135 -0
- package/scripts/ship_review_gate.py +63 -0
- package/scripts/specs_convention_marker.py +103 -0
- package/scripts/test_improve_resume.py +277 -0
- package/scripts/test_review_mechanics.py +958 -0
- package/scripts/token_efficiency_review.py +322 -0
- package/scripts/verdict_scope.py +285 -0
- package/scripts/verify_gherkin_quality_critic_isolation.py +296 -0
- package/scripts/verify_tier.py +157 -0
- package/skills/adr-tools/SKILL.md +118 -0
- package/skills/agent-readiness/SKILL.md +105 -0
- package/skills/agent-readiness/ai_friendly_analyzers.py +326 -0
- package/skills/agent-readiness/scanner.py +441 -0
- package/skills/agent-readiness/scorecard.yaml +88 -0
- package/skills/api-design/SKILL.md +115 -0
- package/skills/apply-fixes/SKILL.md +171 -0
- package/skills/apply-test-doubles/SKILL.md +321 -0
- package/skills/artifact-lifecycle/SKILL.md +127 -0
- package/skills/autoship/SKILL.md +1124 -0
- package/skills/benchmark/SKILL.md +105 -0
- package/skills/branch-workflow/SKILL.md +89 -0
- package/skills/browse/SKILL.md +184 -0
- package/skills/browser-testing/SKILL.md +62 -0
- package/skills/browser-testing/references/playwright-patterns.md +216 -0
- package/skills/build/SKILL.md +422 -0
- package/skills/build/references/static-self-heal.md +245 -0
- package/skills/careful/SKILL.md +72 -0
- package/skills/cd-test-architecture/SKILL.md +371 -0
- package/skills/ci-debugging/SKILL.md +105 -0
- package/skills/co-evolution-audit/SKILL.md +269 -0
- package/skills/code-review/SKILL.md +1015 -0
- package/skills/code-review/examples/aggregated-sample.json +56 -0
- package/skills/code-review/examples/sample-report.md +41 -0
- package/skills/code-review/output-format.md +478 -0
- package/skills/code-review/scripts/activation.py +86 -0
- package/skills/code-review/scripts/change_impact.py +357 -0
- package/skills/code-review/scripts/change_shape.py +372 -0
- package/skills/code-review/scripts/change_size.py +212 -0
- package/skills/code-review/scripts/changed_file_list.py +141 -0
- package/skills/code-review/scripts/closing_pass.py +187 -0
- package/skills/code-review/scripts/consolidate.py +277 -0
- package/skills/code-review/scripts/contract_failure_report.py +185 -0
- package/skills/code-review/scripts/dispatch_reconcile.py +66 -0
- package/skills/code-review/scripts/dispatch_waves.py +164 -0
- package/skills/code-review/scripts/finding_signature.py +446 -0
- package/skills/code-review/scripts/ledger.py +283 -0
- package/skills/code-review/scripts/partition.py +169 -0
- package/skills/code-review/scripts/render_tiered_findings.py +274 -0
- package/skills/code-review/scripts/repo_invariants.py +1066 -0
- package/skills/code-review/scripts/review_context_pack.py +306 -0
- package/skills/code-review/scripts/review_round_log.py +345 -0
- package/skills/code-review/scripts/review_value_coverage.py +297 -0
- package/skills/code-review/scripts/validate_review_output.py +467 -0
- package/skills/code-review/sliced-mode.md +205 -0
- package/skills/competitive-analysis/SKILL.md +191 -0
- package/skills/context-loading-protocol/SKILL.md +157 -0
- package/skills/continue/SKILL.md +90 -0
- package/skills/cost-report/SKILL.md +178 -0
- package/skills/coverage-baseline/SKILL.md +335 -0
- package/skills/coverage-baseline/references/multi-project-discovery.md +202 -0
- package/skills/coverage-delta/SKILL.md +181 -0
- package/skills/coverage-delta/references/mutation-gate.md +70 -0
- package/skills/design-doc/SKILL.md +95 -0
- package/skills/design-interrogation/SKILL.md +89 -0
- package/skills/design-it-twice/SKILL.md +91 -0
- package/skills/docker-image-audit/SKILL.md +108 -0
- package/skills/docker-image-audit/references/install-guide.md +64 -0
- package/skills/docker-image-audit/references/report-template.md +73 -0
- package/skills/docker-image-create/SKILL.md +185 -0
- package/skills/domain-analysis/SKILL.md +183 -0
- package/skills/domain-driven-design/SKILL.md +194 -0
- package/skills/exploratory-testing/SKILL.md +108 -0
- package/skills/explore/SKILL.md +51 -0
- package/skills/farley-score/SKILL.md +165 -0
- package/skills/feature-file-validation/SKILL.md +78 -0
- package/skills/feature-file-validation/references/validation-rules.md +115 -0
- package/skills/feedback-learning/SKILL.md +414 -0
- package/skills/fix/SKILL.md +450 -0
- package/skills/freeze/SKILL.md +68 -0
- package/skills/frontend-architecture/SKILL.md +113 -0
- package/skills/gherkin-derive/SKILL.md +630 -0
- package/skills/gherkin-public/SKILL.md +266 -0
- package/skills/governance-compliance/SKILL.md +150 -0
- package/skills/guard/SKILL.md +75 -0
- package/skills/handoff/SKILL.md +139 -0
- package/skills/handoff/references/summary-templates.md +242 -0
- package/skills/harness-audit/SKILL.md +751 -0
- package/skills/harness-audit/scripts/lesson_validate.py +386 -0
- package/skills/harness-audit/scripts/redundancy_criterion.py +188 -0
- package/skills/headless-run/SKILL.md +45 -0
- package/skills/headless-run/scripts/isolated_dispatch.py +381 -0
- package/skills/help/SKILL.md +72 -0
- package/skills/hexagonal-architecture/SKILL.md +85 -0
- package/skills/human-oversight-protocol/SKILL.md +224 -0
- package/skills/issues-from-assessment/SKILL.md +223 -0
- package/skills/issues-from-plan/SKILL.md +133 -0
- package/skills/legacy-code/SKILL.md +132 -0
- package/skills/mermaid-diagramming/SKILL.md +120 -0
- package/skills/mutation-night-watch/SKILL.md +154 -0
- package/skills/mutation-night-watch/references/scheduling.md +135 -0
- package/skills/mutation-testing/SKILL.md +396 -0
- package/skills/mutation-testing/references/languages/csharp-stryker-net.md +676 -0
- package/skills/mutation-testing/references/languages/go-go-mutesting.md +95 -0
- package/skills/mutation-testing/references/languages/java-pitest.md +77 -0
- package/skills/mutation-testing/references/languages/javascript-stryker.md +188 -0
- package/skills/mutation-testing/references/languages/python-mutmut.md +97 -0
- package/skills/mutation-testing/references/time-estimation.md +34 -0
- package/skills/mutation-testing/references/tool-detection.md +15 -0
- package/skills/mutation-testing/references/workflow-callers.md +23 -0
- package/skills/mutation-testing/scripts/__pycache__/xunit_v3_feature_detector.cpython-314.pyc +0 -0
- package/skills/mutation-testing/scripts/csharp_stryker_net_slice_runner.py +635 -0
- package/skills/mutation-testing/scripts/csharp_stryker_net_status_loop.py +525 -0
- package/skills/mutation-testing/scripts/csharp_stryker_net_wrapper.py +681 -0
- package/skills/mutation-testing/scripts/mutation_baseline_reuse.py +292 -0
- package/skills/mutation-testing/scripts/mutation_exclude_policy.py +268 -0
- package/skills/mutation-testing/scripts/mutation_feasibility_gate.py +463 -0
- package/skills/mutation-testing/scripts/mutation_kill_headless.py +331 -0
- package/skills/mutation-testing/scripts/mutation_kill_insert.py +199 -0
- package/skills/mutation-testing/scripts/mutation_kill_insert_python.py +150 -0
- package/skills/mutation-testing/scripts/mutation_kill_loop.py +869 -0
- package/skills/mutation-testing/scripts/mutation_kill_loop_python.py +949 -0
- package/skills/mutation-testing/scripts/mutation_kill_retry.py +592 -0
- package/skills/mutation-testing/scripts/mutation_kill_shared.py +620 -0
- package/skills/mutation-testing/scripts/mutation_nightwatch.py +462 -0
- package/skills/mutation-testing/scripts/mutation_nightwatch_stacks.py +425 -0
- package/skills/mutation-testing/scripts/mutation_report.py +743 -0
- package/skills/mutation-testing/scripts/mutation_report_cli.py +175 -0
- package/skills/mutation-testing/scripts/mutation_safety_gate.py +69 -0
- package/skills/mutation-testing/scripts/stryker_shard_pipeline.py +847 -0
- package/skills/mutation-testing/scripts/stryker_shard_setup.py +440 -0
- package/skills/mutation-testing/scripts/stryker_timeout_retry.py +142 -0
- package/skills/mutation-testing/scripts/xunit_v3_feature_detector.py +341 -0
- package/skills/performance-benchmark/SKILL.md +174 -0
- package/skills/performance-benchmark/examples/report-format.md +43 -0
- package/skills/performance-benchmark/references/benchmark-script.md +169 -0
- package/skills/performance-metrics/SKILL.md +265 -0
- package/skills/plan/SKILL.md +199 -0
- package/skills/plan/references/gherkin-persistence.md +43 -0
- package/skills/plan/references/plan-template.md +182 -0
- package/skills/pr/SKILL.md +289 -0
- package/skills/pr/scripts/gate_retry_state.py +368 -0
- package/skills/project-init/README.md +141 -0
- package/skills/project-init/SKILL.md +1197 -0
- package/skills/project-init/evals/evals.json +200 -0
- package/skills/project-init/references/capability-tools.md +55 -0
- package/skills/project-init/references/configs.md +221 -0
- package/skills/property-based-testing/SKILL.md +121 -0
- package/skills/property-based-testing/fixtures/invariant_fixture.py +15 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/README.md +42 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/fast-check/LICENSE +21 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/fast-check/README.md +263 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/fast-check/lib/cjs/fast-check.d.ts +5165 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/fast-check/lib/cjs/fast-check.js +12147 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/fast-check/lib/cjs/package.json +3 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/fast-check/lib/cjs/types57/fast-check.d.ts +5165 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/fast-check/lib/fast-check.d.ts +5165 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/fast-check/lib/fast-check.js +12011 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/fast-check/lib/rolldown-runtime-D7D4PA-g.js +13 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/fast-check/lib/types57/fast-check.d.ts +5165 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/fast-check/package.json +94 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/LICENSE +21 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/README.md +168 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/RandomGenerator-DcXj09Ch.d.ts +14 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/distribution/uniformBigInt.d.ts +15 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/distribution/uniformBigInt.js +38 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/distribution/uniformFloat32.d.ts +15 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/distribution/uniformFloat32.js +18 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/distribution/uniformFloat64.d.ts +15 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/distribution/uniformFloat64.js +22 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/distribution/uniformInt.d.ts +15 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/distribution/uniformInt.js +134 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/RandomGenerator-DcXj09Ch.d.ts +14 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/distribution/uniformBigInt.d.ts +15 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/distribution/uniformBigInt.js +37 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/distribution/uniformFloat32.d.ts +15 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/distribution/uniformFloat32.js +17 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/distribution/uniformFloat64.d.ts +15 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/distribution/uniformFloat64.js +21 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/distribution/uniformInt.d.ts +15 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/distribution/uniformInt.js +133 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/generator/congruential32.d.ts +7 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/generator/congruential32.js +44 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/generator/mersenne.d.ts +7 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/generator/mersenne.js +90 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/generator/xoroshiro128plus.d.ts +7 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/generator/xoroshiro128plus.js +80 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/generator/xorshift128plus.d.ts +7 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/generator/xorshift128plus.js +78 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/package.json +3 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/types/JumpableRandomGenerator.d.ts +16 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/types/JumpableRandomGenerator.js +0 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/types/RandomGenerator.d.ts +2 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/types/RandomGenerator.js +0 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/utils/generateN.d.ts +6 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/utils/generateN.js +8 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/utils/purify.d.ts +12 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/utils/purify.js +9 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/utils/skipN.d.ts +6 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/utils/skipN.js +6 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/generator/congruential32.d.ts +7 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/generator/congruential32.js +46 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/generator/mersenne.d.ts +7 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/generator/mersenne.js +92 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/generator/xoroshiro128plus.d.ts +7 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/generator/xoroshiro128plus.js +82 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/generator/xorshift128plus.d.ts +7 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/generator/xorshift128plus.js +80 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/types/JumpableRandomGenerator.d.ts +16 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/types/JumpableRandomGenerator.js +0 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/types/RandomGenerator.d.ts +2 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/types/RandomGenerator.js +0 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/utils/generateN.d.ts +6 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/utils/generateN.js +9 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/utils/purify.d.ts +12 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/utils/purify.js +10 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/utils/skipN.d.ts +6 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/utils/skipN.js +7 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/package.json +133 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/package-lock.json +1179 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/package.json +14 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/roundtrip.js +29 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/roundtrip.properties.test.js +16 -0
- package/skills/property-based-testing/fixtures/no_property_fixture.py +10 -0
- package/skills/property-based-testing/fixtures/roundtrip_fixture.py +16 -0
- package/skills/property-based-testing/references/languages/javascript.md +54 -0
- package/skills/property-based-testing/scripts/detect_and_dispatch.py +80 -0
- package/skills/property-based-testing/scripts/hypothesis_scaffold.py +276 -0
- package/skills/proxy-resilience/SKILL.md +84 -0
- package/skills/quality-gate-pipeline/SKILL.md +184 -0
- package/skills/quality-targets-converge/SKILL.md +254 -0
- package/skills/repo-review/SKILL.md +159 -0
- package/skills/report-pdf/SKILL.md +66 -0
- package/skills/review/SKILL.md +47 -0
- package/skills/review-agent/SKILL.md +152 -0
- package/skills/review-summary/SKILL.md +73 -0
- package/skills/run-report/SKILL.md +70 -0
- package/skills/semantic-duplication-scan/SKILL.md +337 -0
- package/skills/semantic-scan/SKILL.md +53 -0
- package/skills/semgrep-analyze/SKILL.md +139 -0
- package/skills/setup/SKILL.md +1122 -0
- package/skills/ship/SKILL.md +240 -0
- package/skills/source-verification/SKILL.md +210 -0
- package/skills/source-verification/scripts/claim_extractor.py +155 -0
- package/skills/specs/.size-baseline.json +4 -0
- package/skills/specs/SKILL.md +243 -0
- package/skills/specs/references/completeness-checklist.md +83 -0
- package/skills/specs/references/extraction.md +58 -0
- package/skills/specs/references/glossary.md +59 -0
- package/skills/specs/references/persistence.md +115 -0
- package/skills/specs/references/predictability-check.md +77 -0
- package/skills/static-analysis-integration/SKILL.md +235 -0
- package/skills/static-analysis-integration/adapters/_envelope.py +26 -0
- package/skills/static-analysis-integration/adapters/jscpd-adapter.py +66 -0
- package/skills/static-analysis-integration/adapters/lizard-adapter.py +81 -0
- package/skills/static-analysis-integration/adapters/mypy-adapter.py +50 -0
- package/skills/static-analysis-integration/adapters/mypy-src-layout.py +93 -0
- package/skills/static-analysis-integration/adapters/security-review-adapter.py +212 -0
- package/skills/static-analysis-integration/maintenance.md +23 -0
- package/skills/static-analysis-integration/references/language-setup.md +228 -0
- package/skills/static-analysis-integration/references/sarif-parser.md +124 -0
- package/skills/static-analysis-integration/references/security-review-adapter.md +118 -0
- package/skills/static-analysis-integration/references/tool-configs.md +617 -0
- package/skills/static-analysis-integration/rulesets/pmd-quickstart.xml +24 -0
- package/skills/stryker-xunit-v2-shim/SKILL.md +274 -0
- package/skills/stryker-xunit-v2-shim/references/shim-howto.md +256 -0
- package/skills/stryker-xunit-v2-shim/scripts/generate_shim.py +143 -0
- package/skills/systematic-debugging/SKILL.md +130 -0
- package/skills/telemetry/SKILL.md +75 -0
- package/skills/test-audit-disable/SKILL.md +129 -0
- package/skills/test-design/SKILL.md +177 -0
- package/skills/test-design/scripts/__pycache__/internal_double_detector.cpython-314.pyc +0 -0
- package/skills/test-design/scripts/internal_double_detector.py +631 -0
- package/skills/test-design-advisor/SKILL.md +166 -0
- package/skills/test-driven-development/SKILL.md +169 -0
- package/skills/test-health/SKILL.md +262 -0
- package/skills/test-improve/SKILL.md +239 -0
- package/skills/test-improve/references/phase-0-approach-contract.md +228 -0
- package/skills/test-improve/references/phase-1-analyze.md +131 -0
- package/skills/test-improve/references/phase-2-baseline.md +121 -0
- package/skills/test-improve/references/phase-3-derive-gherkin.md +53 -0
- package/skills/test-improve/references/phase-4-plan-fixes.md +34 -0
- package/skills/test-improve/references/phase-5-improve.md +215 -0
- package/skills/test-improve/references/phase-6-refactor-decision.md +45 -0
- package/skills/test-improve/references/phase-7-refactor.md +44 -0
- package/skills/test-improve/references/phase-8-validate.md +66 -0
- package/skills/test-improve/references/phase-9-close-out-prompt.md +11 -0
- package/skills/test-improve/references/phase-9-report.md +62 -0
- package/skills/test-improve/references/review-loop.md +92 -0
- package/skills/test-improve/templates/executive-summary.md +123 -0
- package/skills/threat-modeling/SKILL.md +108 -0
- package/skills/triage/SKILL.md +211 -0
- package/skills/ubiquitous-language/SKILL.md +192 -0
- package/skills/ubiquitous-language/scripts/collect_domain_signals.py +300 -0
- package/skills/unfreeze/SKILL.md +37 -0
- package/skills/upgrade/SKILL.md +31 -0
- package/skills/upgrade/scripts/check_version_drift.py +113 -0
- package/skills/upgrade/scripts/enable_autoupdate.py +149 -0
- package/skills/version/SKILL.md +25 -0
- package/sync/__pycache__/sync_upstream.cpython-314.pyc +0 -0
- package/sync/sync_upstream.py +293 -0
- package/templates/ACCEPTED-RISKS.md.tmpl +46 -0
- package/templates/agents/agent-template.md +151 -0
- package/templates/agents/angular-testing.md +66 -0
- package/templates/agents/csharp-quality.md +63 -0
- package/templates/agents/esm-enforcer.md +52 -0
- package/templates/agents/front-end-testing.md +65 -0
- package/templates/agents/go-quality.md +65 -0
- package/templates/agents/python-quality.md +62 -0
- package/templates/agents/react-testing.md +61 -0
- package/templates/agents/ts-enforcer.md +60 -0
- package/templates/agents/twelve-factor-audit.md +49 -0
- package/tools/entropy-check.py +250 -0
- package/tools/model-hash-verify.py +213 -0
|
@@ -0,0 +1,1015 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: code-review
|
|
3
|
+
description: >-
|
|
4
|
+
Run all enabled review agents against target files. Use this whenever the
|
|
5
|
+
user asks for a code review, wants feedback on their code, says "review my
|
|
6
|
+
code", "check this before I PR", "what's wrong with this", "run the
|
|
7
|
+
agents", or has just finished implementing a feature. Use proactively
|
|
8
|
+
before commits and pull requests.
|
|
9
|
+
argument-hint: >-
|
|
10
|
+
[--agent <name>] [--since <ref>] [--path <dir>] [--all] [--json]
|
|
11
|
+
[--expand <finding-id>|all]
|
|
12
|
+
[--internal] [--force --reason "<text>"]
|
|
13
|
+
[--static-analysis|--no-static-analysis] [--init-risks] [--background]
|
|
14
|
+
[--pdf]
|
|
15
|
+
user-invocable: true
|
|
16
|
+
allowed-tools: >-
|
|
17
|
+
Read, Write, Edit, Grep, Glob, AskUserQuestion, Agent,
|
|
18
|
+
Bash(git diff *), Bash(npx *), Bash(npm run *),
|
|
19
|
+
Bash(pnpm *), Bash(yarn *), Bash(tsc *), Bash(eslint *),
|
|
20
|
+
Bash(git log *), Bash(gh run *), Bash(semgrep *),
|
|
21
|
+
Bash(gitleaks *), Bash(lizard *), Bash(jscpd *),
|
|
22
|
+
Bash(ruff *), Bash(mypy *), Skill(review-agent *)
|
|
23
|
+
---
|
|
24
|
+
|
|
25
|
+
# Code Review
|
|
26
|
+
|
|
27
|
+
**The review-agent panel is the primary quality gate** (Rec 5,
|
|
28
|
+
`docs/experiments/RECOMMENDATIONS.md`). The review-agent lens — SRP,
|
|
29
|
+
complexity, coupling, duplication — was the only quality axis that separated
|
|
30
|
+
workflow arms in the experiment line. Coverage and mutation scores saturate
|
|
31
|
+
near-identically across every workflow shape and must **never** be used to
|
|
32
|
+
rank workflow quality: the losing big-batch and split arms posted *higher*
|
|
33
|
+
mutation scores (0.93–0.98) than the two winners (0.80–0.86). A higher
|
|
34
|
+
coverage or mutation number is not evidence that code — or the workflow that
|
|
35
|
+
produced it — is better. (The deterministic static-analysis pre-pass below is
|
|
36
|
+
a different, complementary axis: mechanical findings cleared before the
|
|
37
|
+
semantic panel runs, not a metric competing with it.)
|
|
38
|
+
|
|
39
|
+
Role: orchestrator. Route work to review agents; do not review code yourself. Pass each agent's `model:`/`effort:` frontmatter as declared when dispatching — the harness resolves both fields natively before dispatch, per Model/Effort Resolution in `agents/orchestrator.md` (ADR 0026).
|
|
40
|
+
|
|
41
|
+
Output templates and JSON schemas: [`output-format.md`](output-format.md). Example report: [`examples/sample-report.md`](examples/sample-report.md).
|
|
42
|
+
|
|
43
|
+
## Orchestrator constraints
|
|
44
|
+
|
|
45
|
+
**MUST — confirm agent-dispatch capability before anything else in this skill (issue #1461).** Before attempting to dispatch ANY review agent (Step 4), you MUST confirm the `Agent` (or `Task`) tool is actually present and available in your current toolset. If it is not present: **STOP.** Do not proceed with a self-applied, inline, or checklist-based review of any kind as a substitute for independent dispatch — an orchestrator applying the review agents' checklists itself is not a review, it is self-certification, and it defeats the entire purpose of this gate. Do not write `.pr-review-passed` under any circumstance in this state. Instead, report to the user/operator plainly: code review cannot run in this environment because no agent-dispatch capability (`Agent`/`Task` tool) is available; name exactly what's missing; and state that the PR gate cannot be satisfied until `/code-review` is re-run from a session that has that capability. This is a hard requirement, not a preference — "should dispatch agents" is not sufficient; a missing `Agent`/`Task` tool always halts this skill before Step 2.
|
|
46
|
+
|
|
47
|
+
1. **Do not review code yourself.** Delegate all semantic analysis to review agents.
|
|
48
|
+
2. **Minimize context per agent.** Pass only what each agent's `Context needs` field requires.
|
|
49
|
+
3. **Route to the right model.** Each agent's `model:`/`effort:` frontmatter declares its model alias and reasoning effort; the harness resolves both fields natively before dispatch, per `agents/orchestrator.md` → Model/Effort Resolution (ADR 0026). Do not override the frontmatter value.
|
|
50
|
+
4. **Run deterministic gates first.** Lint, type-check, secret scan are cheaper than AI. Stop if they fail.
|
|
51
|
+
5. **Return structured results.** Aggregate agent JSON; do not add your own findings.
|
|
52
|
+
6. **Be concise.** Tables and JSON, no preambles, no filler.
|
|
53
|
+
|
|
54
|
+
## Parse Arguments
|
|
55
|
+
|
|
56
|
+
Arguments: $ARGUMENTS
|
|
57
|
+
|
|
58
|
+
| Flag | Behavior |
|
|
59
|
+
| --- | --- |
|
|
60
|
+
| `--agent <name>` | Run only the named agent (delegates to `/review-agent`) |
|
|
61
|
+
| `--since <ref>` | Review files changed since the ref — see step 1 for the exact command (the `-c diff.relative=false -c core.quotePath=false` overrides there are load-bearing, not cosmetic) |
|
|
62
|
+
| `--path <dir>` | Review only files in this directory |
|
|
63
|
+
| `--all` | Force full-repository review even when uncommitted changes exist |
|
|
64
|
+
| `--slice <N>` | Engage sliced large-repo review explicitly, capping each slice at N files (module-aligned) at any repo size. `N` must be a positive integer. See [`sliced-mode.md`](sliced-mode.md). |
|
|
65
|
+
| `--resume` | Resume a sliced run — skip slices whose section artifact already exists on disk. See [`sliced-mode.md`](sliced-mode.md). |
|
|
66
|
+
| `--no-slice` | Escape hatch — force the legacy single-pass review even on a large full-repo scope that would otherwise auto-engage sliced mode. |
|
|
67
|
+
| `--json` | Output aggregated JSON to **stdout** instead of prose. Contractually non-interactive (for CI): never prompts; defaults to report-only (no code modified). |
|
|
68
|
+
| `--expand <finding-id>|all` | Prose-mode only (step 7): render Tier-2 (full message + suggested fix) for the named finding-id, or for every finding with `all`, after the Tier-1 report — see step 7. A no-op under `--json` (see step 7's `--json` branch). **Only meaningful within the SAME run that computed the ids** — pass it alongside `--since`/`--path`/etc. in one invocation once you already know a specific id, e.g. because the operating Claude session read the prior Tier-1 output and is now re-invoking this skill with the same scope plus `--expand <id>` still in the same conversation; that path never re-dispatches anything beyond what the scope would have dispatched anyway. A cold, separate `/code-review --expand <id>` run with no memory of where that id came from IS a full re-dispatch of the panel (steps 1-6 run in full, same as any other invocation) and the id is not guaranteed to still exist or mean the same finding — `render_tiered_findings.py`'s own docstring says ids are not stable across runs. `--expand` never triggers a SECOND panel dispatch on top of an already-running one; it only changes step 7's rendering of the one panel a given invocation already ran. |
|
|
69
|
+
| `--pdf` | After the durable report is written, also render it to a sibling PDF via `hooks/lib/report_pdf.py`. See `knowledge/report-pdf-integration.md`. No-op with a message when no report file is written (`--json` or `--internal`); under `--json`, that status goes to **stderr** so stdout stays pure JSON. Additive: never changes the review's own output or exit status. |
|
|
70
|
+
| `--internal` | This is an orchestrator-internal dispatch (`/build`'s Step 6 backstop review, `/test-improve`'s Phase 4/5 end-of-phase review loop) — skip the `.dev-team-reports/code-review.md` report write in step 7. Orthogonal to `--json`: `--internal` alone still runs the prose/fix-loop path; both sanctioned callers use `--internal` without `--json` specifically to keep the fix loop. `/build` and `/test-improve` are the only sanctioned callers of this flag today — see `knowledge/report-output-location.md` for `/ship`'s deliberate exception (writes the report by default, no `--internal`). |
|
|
71
|
+
| `--init-risks` | Scaffold `ACCEPTED-RISKS.md` from `templates/ACCEPTED-RISKS.md.tmpl` if absent. Exits non-zero without overwriting if present. Schema: `knowledge/accepted-risks-schema.md`. |
|
|
72
|
+
| `--force` | Skip pre-flight gates **and the documentation-only short-circuit** (forces a full review of doc-only changes). **Requires `--reason "<text>"`** — logged to `.claude/metrics/override-audit.jsonl`. |
|
|
73
|
+
| `--reason "<text>"` | Override justification (required with `--force`) |
|
|
74
|
+
| `--static-analysis` / `--no-static-analysis` | Force on/off the static analysis pre-pass (Semgrep, ESLint, TypeScript, Ruff, mypy). Auto-enabled when tools are detected. |
|
|
75
|
+
| `--background` | Drift review mode — review default branch for documentation, naming, and structural drift. Runs doc-review, arch-review, naming-review, structure-review only. Skips pre-flight gates. |
|
|
76
|
+
| (no flags) | **Auto-scope**: review uncommitted changes if any exist, otherwise full repository |
|
|
77
|
+
|
|
78
|
+
## Progress tracking
|
|
79
|
+
|
|
80
|
+
```text
|
|
81
|
+
- [ ] Target files determined
|
|
82
|
+
- [ ] Documentation-only check (short-circuit if all docs)
|
|
83
|
+
- [ ] Pre-flight gates passed
|
|
84
|
+
- [ ] Static analysis pre-pass (if enabled)
|
|
85
|
+
- [ ] Agents loaded and filtered
|
|
86
|
+
- [ ] All agents executed
|
|
87
|
+
- [ ] Results aggregated
|
|
88
|
+
- [ ] User asked: fix or report only?
|
|
89
|
+
- [ ] Review-fix loop (if user chose fix, up to 5 iterations)
|
|
90
|
+
- [ ] Report generated
|
|
91
|
+
- [ ] Correction prompts saved
|
|
92
|
+
- [ ] Pre-commit gate file written (if auto-scoped to uncommitted changes)
|
|
93
|
+
```
|
|
94
|
+
|
|
95
|
+
## Steps
|
|
96
|
+
|
|
97
|
+
### 1. Determine target files
|
|
98
|
+
|
|
99
|
+
Priority order:
|
|
100
|
+
|
|
101
|
+
1. `--path <dir>` — files in that directory (exclude node_modules, .git, dist, build, coverage)
|
|
102
|
+
2. `--since <ref>` — `git -c diff.relative=false -c core.quotePath=false diff --name-only <ref>...HEAD`
|
|
103
|
+
3. `--all` — all source files
|
|
104
|
+
4. **Auto-scope** (no flags): run `git -c diff.relative=false -c core.quotePath=false diff --name-only` + `git -c diff.relative=false -c core.quotePath=false diff --cached --name-only`, combine and dedupe. If non-empty, review those files. If empty, review the full repository. The explicit `-c diff.relative=false` matters here (#1461 fourth security re-review): a repo/global `diff.relative=true` config would otherwise silently scope this listing to the invocation's cwd, and `review_gate_hash()`/`_staged_names()` (which pin the same override) would then hash/gate a broader staged patch than what was actually reviewed. `-c core.quotePath=false` (#1733) keeps this listing byte-identical to step 3's `changed_file_list.py` input for the same ref/scope — without it, a non-ASCII path would arrive C-quoted here but raw there, and `select_lenses.py`'s `--added` membership test (an exact string comparison) would silently fail to match it.
|
|
105
|
+
|
|
106
|
+
**Stage auto-scoped changes now, before anything else (#1461).** When the auto-scope path found a non-empty file set, `git add` those files immediately — before pre-flight gates, static analysis, or any agent dispatch — so the staged content's hash is fixed from this point through step 9's gate write. This is not cosmetic: `agent_dispatch_ledger.py` stamps each review-agent dispatch's `subject_hash` with `review_gate_hash()` at **dispatch time** (step 4). If staging happened only at step 9 (after dispatch) as previously documented, the dispatch-time hash and the gate-write-time hash would differ whenever the auto-scope target was unstaged — the common case — and every genuine dispatch would silently fail to corroborate the gate, forcing a hard block on a fully legitimate review. Staging here, before dispatch, is what makes step 9's hash and the dispatch ledger's `subject_hash` the same value. An unstaged working-tree edit after this point does **not** by itself change the staged hash (`review_gate_hash()` hashes `git diff --cached`, not the working tree) — step 6a's fix loop explicitly re-stages (`git add`) each iteration's fixes for exactly this reason; see that step for how corroboration is re-established after a fix loop runs.
|
|
107
|
+
|
|
108
|
+
**Never `Read` a directory path directly to enumerate its contents** — `Read` on a directory throws `EISDIR` (the same hazard step 3 avoids for agent-roster enumeration). This applies to `--path <dir>`, `--all`, and the full-repository fallback alike: always list files with `Glob` (e.g. `Glob("<dir>/**/*")`), never a bare `Read` on the directory itself. See `${CLAUDE_PLUGIN_ROOT}/knowledge/directory-enumeration.md` for the shared rule.
|
|
109
|
+
|
|
110
|
+
**Scope validation** (full-repo paths only):
|
|
111
|
+
|
|
112
|
+
| File count | Action |
|
|
113
|
+
| --- | --- |
|
|
114
|
+
| ≤200 | Proceed |
|
|
115
|
+
| 201–500 | Warn: "Reviewing {N} files — consider `--path` to narrow scope." Proceed. |
|
|
116
|
+
| >500 | **Auto-engage sliced mode** (large-repo review) unless `--no-slice`. |
|
|
117
|
+
|
|
118
|
+
**Sliced large-repo review.** On a full-repo scope exceeding the >500 tier (or
|
|
119
|
+
whenever `--slice <N>` is passed), **auto-engage sliced mode**: run the sliced
|
|
120
|
+
path in [`sliced-mode.md`](sliced-mode.md) instead of steps 4–9 below. That file
|
|
121
|
+
owns the full activation precedence (via `scripts/activation.py`), partitioning,
|
|
122
|
+
per-slice panels, persist-and-drop, `--resume`, and cross-slice consolidation —
|
|
123
|
+
not restated here. `--no-slice` forces the legacy single-pass review (steps 2–9)
|
|
124
|
+
even past the threshold; Exactly at 500 files does not auto-engage.
|
|
125
|
+
**Non-full-repo scope** (`--path`, `--since`, auto-scoped uncommitted changes)
|
|
126
|
+
**never** auto-engages, regardless of file count — the review proceeds exactly
|
|
127
|
+
as before this feature. Sliced mode is **report-only** (no interactive fix loop).
|
|
128
|
+
|
|
129
|
+
**Documentation-only short-circuit.** After the target set is known, classify each file. A file is **documentation** when it matches a doc type or path:
|
|
130
|
+
|
|
131
|
+
- extension `.md`, `.mdx`, `.markdown`, `.rst`, `.txt`, `.adoc`
|
|
132
|
+
- any path under a `docs/` directory
|
|
133
|
+
- a root doc: `README*`, `CHANGELOG*`, `CONTRIBUTING*`, `LICENSE*`, `NOTICE*`, `AUTHORS*`, `CODE_OF_CONDUCT*`
|
|
134
|
+
|
|
135
|
+
…**except functional Claude-config markdown, which is never documentation** (it drives agent/skill/command behavior and must be reviewed): any path containing a `.claude/` segment, or under `agents/`, `skills/`, `prompts/`, `knowledge/`, or `templates/agents/`. Treat `CLAUDE.md` and `AGENTS.md` as functional config too, not documentation.
|
|
136
|
+
|
|
137
|
+
If **every** target file is documentation, short-circuit:
|
|
138
|
+
|
|
139
|
+
1. Emit: `Documentation-only changeset ({N} files) — skipping code review. Re-run with --force --reason "<text>" to review anyway.`
|
|
140
|
+
2. If the review was auto-scoped to uncommitted changes or scoped via `--since <base>` (issue #1904 Bug 2b — same extension as step 9's own gate condition, and for the same reason: `/pr`'s only path to `gh pr create` reviews via `--since <base>`), write the `.pr-review-passed` gate file (per step 9) so `hooks/pre_pr_review.py` allows the next `gh pr create`. **Contemporaneously** (before or immediately after that write), record the doc-only exemption as an explicit, auditable boundary event — the `.pr-review-passed` gate's dispatch-ledger corroboration (#1461, #1886) reads this event, bound to the gate's own hash, to let the doc-only path stay exempt from agent-dispatch evidence without being a silent, unaccountable code-path skip:
|
|
141
|
+
```bash
|
|
142
|
+
HASH=$(python3 "${CLAUDE_PLUGIN_ROOT}/hooks/lib/review_gate_hash.py" --branch-diff)
|
|
143
|
+
mkdir -p .claude/memory && echo "$HASH" > .claude/memory/.pr-review-passed
|
|
144
|
+
python3 "${CLAUDE_PLUGIN_ROOT}/hooks/lib/boundary_events.py" --event doc-only --subject-hash "$HASH"
|
|
145
|
+
```
|
|
146
|
+
3. In `--json` mode, emit `{"status": "skipped", "reason": "documentation-only", "files": [<list>]}` instead.
|
|
147
|
+
4. **Stop.** Do not run pre-flight gates, static analysis, or any agent.
|
|
148
|
+
|
|
149
|
+
**Bypass:** the short-circuit does **not** apply with `--force` (with `--reason`), `--agent <name>`, or `--background` (drift review always inspects docs).
|
|
150
|
+
|
|
151
|
+
### 1b. Check for institutional context
|
|
152
|
+
|
|
153
|
+
If `REVIEW-CONTEXT.md` exists at the repo root, read it and pass its contents to every agent in step 4, prefixed with: "Institutional context provided for this review:". This file is optional.
|
|
154
|
+
|
|
155
|
+
### 1c. Probe for optional MCP tools
|
|
156
|
+
|
|
157
|
+
| Tool | Check | Use |
|
|
158
|
+
| --- | --- | --- |
|
|
159
|
+
| RoslynMCP | `get_code_metrics` / `search_symbols` available | C# metrics, compiler diagnostics |
|
|
160
|
+
| CodeGraph | `.codegraph/` present / `mcp__codegraph__codegraph_explore` available | Verified structural skeletons, resolved callers/callees/impact |
|
|
161
|
+
| Repowise | `get_context` / `get_symbol` / `search_codebase` / `get_risk` available | Verified file/symbol context + modification-risk lookups |
|
|
162
|
+
| Documentation MCP | wiki/docs search available | Architecture docs |
|
|
163
|
+
| Semgrep | `which semgrep` | SAST context for security-review |
|
|
164
|
+
|
|
165
|
+
Pass availability info to each agent so they can use enhanced tools or fall back to Glob/Grep/Read. All read-only review agents grant these MCP tools; see [`knowledge/codegraph-vs-graphify.md`](../../knowledge/codegraph-vs-graphify.md) for tool selection and the fallback contract. Include availability in the final report per `knowledge/review-template.md`.
|
|
166
|
+
|
|
167
|
+
### 2. Pre-flight gates
|
|
168
|
+
|
|
169
|
+
Skip entirely if `--background`. If `--force` without `--reason`, halt:
|
|
170
|
+
|
|
171
|
+
```
|
|
172
|
+
ERROR: --force requires --reason "<justification>".
|
|
173
|
+
```
|
|
174
|
+
|
|
175
|
+
If `--force` with `--reason`, append an entry to `.claude/metrics/override-audit.jsonl` per the schema in [`output-format.md`](output-format.md#override-audit-log-entry-step-2---force-path), then proceed to step 3.
|
|
176
|
+
|
|
177
|
+
Otherwise run these in sequence (stop on first failure):
|
|
178
|
+
|
|
179
|
+
1. **Lint**: `npx eslint` (or project lint command) on target files.
|
|
180
|
+
2. **Type check**: `npx tsc --noEmit` if `tsconfig.json` exists.
|
|
181
|
+
3. **Secret scan** (#1977). Prefer the purpose-built scanner when it is installed, and keep the grep as the zero-dependency fallback — this gate is the one place a committed credential must hard-stop the review, and it was checking a single regex while a real secrets scanner already ran one step later in the pre-pass:
|
|
182
|
+
- If `command -v gitleaks` succeeds, run it with the **canonical invocation** in [`skills/static-analysis-integration/references/tool-configs.md`](../static-analysis-integration/references/tool-configs.md) § gitleaks — one documented command, not a variant of it (`--no-verify` is what keeps it fully offline). Any finding on a target file → **fail the gate**. Report the rule id and `file:line` only; **never echo the matched secret value** into the report or transcript. Record that gitleaks ran, and **do not run it again in step 2b** — same "do not run Semgrep twice" rule.
|
|
183
|
+
- Otherwise (gitleaks absent): grep target files for the runnable pattern in [`knowledge/owasp-detection.md`](../../knowledge/owasp-detection.md) § Hardcoded-key pattern (the fenced code block, not the table row — table cells escape `|` as `\|`, a literal pipe rather than alternation). Note in the report that the fallback ran, so a reader can tell "no secrets found by gitleaks" apart from "no secrets found by one regex".
|
|
184
|
+
4. **Semgrep SAST**: `semgrep scan --config auto --quiet --json` on target files if installed. ERROR-severity → fail. WARNING-severity → continue, include in report. Save findings for security-review context.
|
|
185
|
+
5. **Pipeline-red check**: `gh run list --branch $(git branch --show-current) --limit 1 --json conclusion -q '.[0].conclusion'` if `gh` is available. If the last CI run failed, warn: "Pipeline is red. Fix CI before adding new code. Use `--force` to override."
|
|
186
|
+
|
|
187
|
+
Skip any gate silently if its tool is unavailable.
|
|
188
|
+
|
|
189
|
+
### 2b. Static analysis pre-pass
|
|
190
|
+
|
|
191
|
+
Skip if `--no-static-analysis` or `--background`.
|
|
192
|
+
|
|
193
|
+
Follow the detection, execution, and deduplication procedure in [`skills/static-analysis-integration/SKILL.md`](../static-analysis-integration/SKILL.md). Output is structured findings injected into agent context in step 4. **This step does not gate execution** — it collects context only.
|
|
194
|
+
|
|
195
|
+
**Growing this registry is a rule, not a discretion (#1981).** When any review agent reports the same mechanically-checkable finding class for the **second** time, and the check is expressible as a deterministic script, it becomes a `CHECKS` entry in `scripts/repo_invariants.py` **in the same PR that fixes the finding**. Two occurrences make it a class; converting there turns an unbounded stream of re-derivations into one bounded conversion. The root `CLAUDE.md` Working Rules state the same rule — it applies to this repo's own development, and the mechanism ships for downstream projects to use the same way.
|
|
196
|
+
|
|
197
|
+
**Repo-specific invariant pre-pass (#1608).** Also run:
|
|
198
|
+
|
|
199
|
+
```bash
|
|
200
|
+
python3 "$CLAUDE_PLUGIN_ROOT/skills/code-review/scripts/repo_invariants.py" --files <target files>
|
|
201
|
+
```
|
|
202
|
+
|
|
203
|
+
It checks a small, growable list of this repo's own "every X should have
|
|
204
|
+
exactly one corresponding Y" invariants — mechanically checkable facts a full
|
|
205
|
+
agent panel would otherwise re-derive independently, once per agent, every
|
|
206
|
+
round. Its `findings` array merges into step 4's static-analysis context using
|
|
207
|
+
the same envelope and the same "detected by static analysis — do not
|
|
208
|
+
re-report, focus on semantic concerns" framing. Expand `CHECKS` in that script
|
|
209
|
+
as more rediscovered-N-times cases turn up; this step never needs to change to
|
|
210
|
+
pick up a new check.
|
|
211
|
+
|
|
212
|
+
**Internal-collaborator-doubling pre-pass (#2130).** Also run:
|
|
213
|
+
|
|
214
|
+
```bash
|
|
215
|
+
python3 "$CLAUDE_PLUGIN_ROOT/skills/test-design/scripts/internal_double_detector.py" . --files <target files> --json
|
|
216
|
+
```
|
|
217
|
+
|
|
218
|
+
Detects an unwaived (or malformed/invalid-blocker-waiver) double of a
|
|
219
|
+
project first-party collaborator, per
|
|
220
|
+
`${CLAUDE_PLUGIN_ROOT}/knowledge/internal-collaborator-doubling.md`. Same
|
|
221
|
+
`<target files>` list as the `repo_invariants.py` block above — one
|
|
222
|
+
scoping mechanism for both tools in this step, not two. Its `findings`
|
|
223
|
+
array merges into step 4's static-analysis context using the same envelope
|
|
224
|
+
and the same "detected by static analysis — do not re-report, focus on
|
|
225
|
+
semantic concerns" framing. This step does not gate on the finding — it
|
|
226
|
+
collects context only, exactly like every other check in this step; the
|
|
227
|
+
actual gate is the `hooks/internal_double_gate.py` PreToolUse hook plus
|
|
228
|
+
the required `"Plugin content & hooks"` CI check (#2128). `test-review` and
|
|
229
|
+
`test-smell-review` cite this pre-pass's finding rather than re-deriving
|
|
230
|
+
it when it's present (see
|
|
231
|
+
`${CLAUDE_PLUGIN_ROOT}/knowledge/test-review-division-of-labor.md`).
|
|
232
|
+
|
|
233
|
+
**Test-review mechanical pre-phase (#2169).** Also run, for each test file in `<target files>`:
|
|
234
|
+
|
|
235
|
+
```bash
|
|
236
|
+
python3 "$CLAUDE_PLUGIN_ROOT/scripts/test_review_mechanics.py" . <file>
|
|
237
|
+
```
|
|
238
|
+
|
|
239
|
+
Only relevant when `test-review` is in the dispatched lens set for this
|
|
240
|
+
round; skip entirely otherwise. Unlike the two pre-passes above, this one
|
|
241
|
+
runs **once per file** rather than once over the whole `<target files>`
|
|
242
|
+
list, because `test-review.md`'s own Phase 0 (`agents/test-review.md` →
|
|
243
|
+
Protocol) needs each file's own `mechanicalFail`/findings result supplied as
|
|
244
|
+
that file's context — the agent has no `Bash` tool and never runs this
|
|
245
|
+
script itself. Keep the per-file results keyed by file path when assembling
|
|
246
|
+
step 4's context so each file's `test-review` dispatch gets its own result,
|
|
247
|
+
not the whole batch's. Pass each result to `test-review` as its Phase 0
|
|
248
|
+
input using `agents/test-review.md`'s own framing ("detected by static
|
|
249
|
+
analysis, do not re-derive" — the agent still reports it as this file's own
|
|
250
|
+
finding when `mechanicalFail` is true, per that file's Phase 0 bullets) —
|
|
251
|
+
**not** the generic "detected by static analysis — do not re-report, focus
|
|
252
|
+
on semantic concerns" envelope the two pre-passes above use for every other
|
|
253
|
+
agent. That generic framing is correct for `repo_invariants.py`/
|
|
254
|
+
`internal_double_detector.py`'s findings, which every dispatched agent
|
|
255
|
+
receives as already-covered context to fold silently into a semantic
|
|
256
|
+
review; it would be wrong here, since `test-review` is this pre-pass's
|
|
257
|
+
sole intended reporter, not one of several agents absorbing someone else's
|
|
258
|
+
finding.
|
|
259
|
+
|
|
260
|
+
**Pass `--files` (#1629).** Several checks are scoped to the changeset,
|
|
261
|
+
because the conventions they enforce are "required going forward, do not
|
|
262
|
+
retrofit" (`evals/README.md`'s `_calibration` rule is the motivating case).
|
|
263
|
+
Without `--files` those checks stay silent rather than reporting the ~150
|
|
264
|
+
pre-existing findings the conventions explicitly do not require fixing. The
|
|
265
|
+
`--all` flag exists for deliberate backlog triage and must **not** be used
|
|
266
|
+
here.
|
|
267
|
+
|
|
268
|
+
**Authoring-time ordering (#1629).** When *writing* fixtures or agent files,
|
|
269
|
+
run this same command at edit time, before the first panel dispatches — same
|
|
270
|
+
command, earlier. Of #1619's 8 follow-up rounds, at least 4 were triggered by
|
|
271
|
+
defect classes these deterministic checks catch, plus factually wrong
|
|
272
|
+
runtime-semantics claims that `evals/README.md`'s **executable-claims
|
|
273
|
+
convention** requires verifying by execution at authoring time. A claim the
|
|
274
|
+
author has already run is a claim the panel reviews as evidence rather than
|
|
275
|
+
adjudicates from scratch.
|
|
276
|
+
|
|
277
|
+
If Semgrep already ran in the pre-flight gate, reuse those findings. Do not run Semgrep twice.
|
|
278
|
+
|
|
279
|
+
### 3. Determine enabled agents
|
|
280
|
+
|
|
281
|
+
If `--background`: run only `doc-review`, `arch-review`, `naming-review`, `structure-review`. Skip all others.
|
|
282
|
+
|
|
283
|
+
Otherwise read the roster from the **Review Agents** section of `knowledge/agent-registry.md` — each row names an agent and its `agents/<name>.md` file. **Never `Read` the bare `agents/` directory** (it throws `EISDIR`); if you must confirm files on disk, list them with `Glob("agents/*.md")`, never a directory `Read` (see `${CLAUDE_PLUGIN_ROOT}/knowledge/directory-enumeration.md`). All are enabled by default.
|
|
284
|
+
|
|
285
|
+
**Agent eligibility is resolved by `select_lenses.py` (#1523).** For a diff-scoped run (auto-scope or `--since <ref>`) compute the changed-file list first — the same helper step 4 reuses for the `project-structure` context payload, so this is one computation feeding two consumers, not two ways to derive the same fact (#1733, #1734). **Always** carry the same `-c diff.relative=false -c core.quotePath=false` overrides step 1's own listing uses — omitting them here would let a repo/global `diff.relative=true` (or a non-ASCII path under default `core.quotePath`) desync this list from step 1's `--files`, silently zeroing every `--added` membership match below:
|
|
286
|
+
|
|
287
|
+
```bash
|
|
288
|
+
set -o pipefail # a pipeline's status is its LAST command's without this —
|
|
289
|
+
# changed_file_list.py succeeds trivially on empty stdin, so
|
|
290
|
+
# an upstream git failure would otherwise pass silently.
|
|
291
|
+
|
|
292
|
+
# Auto-scope (uncommitted changes):
|
|
293
|
+
CHANGED_JSON=$({ git -c diff.relative=false -c core.quotePath=false diff --name-status; \
|
|
294
|
+
git -c diff.relative=false -c core.quotePath=false diff --cached --name-status; } \
|
|
295
|
+
| python3 "$CLAUDE_PLUGIN_ROOT/skills/code-review/scripts/changed_file_list.py" --name-status-from -) \
|
|
296
|
+
|| { echo "ERROR: failed to compute the changed-file list" >&2; exit 1; }
|
|
297
|
+
|
|
298
|
+
# --since <ref>:
|
|
299
|
+
CHANGED_JSON=$(git -c diff.relative=false -c core.quotePath=false diff --name-status <ref>...HEAD \
|
|
300
|
+
| python3 "$CLAUDE_PLUGIN_ROOT/skills/code-review/scripts/changed_file_list.py" --name-status-from -) \
|
|
301
|
+
|| { echo "ERROR: failed to compute the changed-file list" >&2; exit 1; }
|
|
302
|
+
```
|
|
303
|
+
|
|
304
|
+
`$CHANGED_JSON` holds `{"files": [{"path", "status"}, ...], "added": [...]}`. Extract both lists into their own variables **first, with an explicit failure check** — a process substitution's own exit status is invisible to the command it feeds, so if the extraction silently produced nothing this step is where that must be caught, not left for `select_lenses.py` to (indistinguishably) treat as "nothing changed":
|
|
305
|
+
|
|
306
|
+
```bash
|
|
307
|
+
FILES_LIST=$(printf '%s' "$CHANGED_JSON" | python3 -c 'import json, sys; print("\n".join(f["path"] for f in json.load(sys.stdin)["files"]))') \
|
|
308
|
+
|| { echo "ERROR: failed to extract file list from CHANGED_JSON" >&2; exit 1; }
|
|
309
|
+
ADDED_LIST=$(printf '%s' "$CHANGED_JSON" | python3 -c 'import json, sys; print("\n".join(json.load(sys.stdin)["added"]))') \
|
|
310
|
+
|| { echo "ERROR: failed to extract added-file list from CHANGED_JSON" >&2; exit 1; }
|
|
311
|
+
```
|
|
312
|
+
|
|
313
|
+
Now run, feeding those two variables — not a separately-interpolated `<target files>` placeholder — to `select_lenses.py` via `--files-from`/`--added-from` process substitution:
|
|
314
|
+
|
|
315
|
+
```bash
|
|
316
|
+
python3 "${CLAUDE_PLUGIN_ROOT}/scripts/select_lenses.py" \
|
|
317
|
+
--files-from <(printf '%s\n' "$FILES_LIST") \
|
|
318
|
+
--added-from <(printf '%s\n' "$ADDED_LIST")
|
|
319
|
+
```
|
|
320
|
+
|
|
321
|
+
Deriving both from the same quoted `$CHANGED_JSON` variable — rather than re-interpolating individual paths as shell words — is what actually closes the injection surface; `--files-from`/`--added-from` on their own only fix the two hazards specific to `select_lenses.py`'s **own** argv parsing (a path beginning with `-` reinterpreted as a flag; word-splitting on an unquoted space) and do not by themselves protect a caller that builds their input by shell-interpolating untrusted path text some other way. For `--path`/`--all`/the full-repository fallback, where there is no diff and therefore no `$CHANGED_JSON`, this skill still passes the target-file list as plain `--files <target files>` argv, matching every other file-list-consuming script call in this skill (`change_shape.py`, `change_size.py`, `closing_pass.py`) — narrowing that broader, pre-existing pattern is a separate initiative, not part of this fix.
|
|
322
|
+
|
|
323
|
+
Always pass `--added-from` for a diff-scoped run, **even when `added` is `[]`** — an empty process substitution still supplies an explicit empty set (narrows away any added-only lens), whereas omitting the flag entirely reverts to the fail-safe fallback (matches an added-only `Scope:` like a plain glob list). Omit both `--files-from` and `--added-from` only for `--path`/`--all`/the full-repository fallback.
|
|
324
|
+
|
|
325
|
+
Take its `lenses` array as the Scope-eligible roster, and **surface its `warnings`** in the review output — a bare agent name means that agent is missing its `Scope:` declaration and was included include-biased; `unnarrowed-added-only:<name>` (#1733) means an added-only lens was kept un-narrowed (matched like a plain glob list) because this run supplied no `--added`/`--added-from`; `skipped-non-executable:<name>` (#1923) means a `Scope: always` lens on the resolver's own `NON_EXECUTABLE_SKIP_ELIGIBLE` allowlist (`correctness-review` today) was dropped from `lenses` because every changed file matched a docs/config/asset/lockfile pattern that lens's own `## Skip` clause already covers — this is a deliberate cost optimization, not a coverage gap, so treat it as informational rather than `fail`-equivalent, distinct from the two shapes below; `unreadable-registry:<file>` means the roster could not be read at all; `unreadable-files-from:<path>`/`unreadable-added-from:<path>` mean the named `--files-from`/`--added-from` source could not be read — **treat either as equivalent to a `fail` status** for this run (an unreadable source is not "nothing changed") rather than proceeding as if the (now-truncated) file list were complete. Never silently drop any of these shapes from the report. The resolver reads each review agent's body-level `Scope:` declaration — `Scope: always` (eligible for any non-empty changeset), a glob list (eligible only when at least one target file matches a declared glob), `Scope: added-only` + globs (eligible only when a target file matching a declared glob was newly *added* — `component-architecture-review`'s dual-placement rule, #1733: unconditional in `/repo-review`, added-only here), `Scope: test-files` (eligible only when a changed file is a test file, resolved against `knowledge/test-file-indicators.md`'s single shared encoding rather than a glob list, which could not express `test_*.py`, `__tests__/`, or the C#/Java annotation indicators — `test-smell-review` declares this, #1978; note `test-review` deliberately stays `Scope: always`, because its coverage-gap check must see production diffs that add code *without* a matching test), or `Scope: on-demand` (never eligible for this per-diff roster at all — `token-efficiency-review`, `ai-provenance-review`, and `claude-setup-review` declare this; they are repo-wide drift/trend metrics dispatched instead by the whole-tree `/repo-review` command, #1735, and `refactor-opportunity-review` declares it too, #1976, dispatched by name at `/build`'s slice review checkpoint where its post-GREEN charter actually applies). `Scope:` is a body declaration, not frontmatter (`agent-contract.json`). This is the single source of truth shared with `/build`'s inline checkpoints: adding or changing an agent's trigger scope needs only an edit to that agent's own body — zero edits to this skill. (The framework-reactivity agents react/vue/angular are **not** in the resolver's roster; they are governed by the manifest rule below.)
|
|
326
|
+
|
|
327
|
+
**Framework-specific reactivity review** — dispatch based on the project's dependency manifest (`package.json` etc.):
|
|
328
|
+
|
|
329
|
+
- React (`react` / `react-dom` in deps): include `react-reactivity-review` scoped to `.jsx`/`.tsx` and React-importing `.js`/`.ts` files
|
|
330
|
+
- Vue (`vue` in deps): include `vue-reactivity-review` scoped to `.vue` and Vue-importing `.js`/`.ts` files
|
|
331
|
+
- Angular (`@angular/core` in deps): include `angular-reactivity-review` scoped to `*.component.ts`, `*.component.html`, `*.service.ts`, and general `.ts` files
|
|
332
|
+
|
|
333
|
+
If `review-config.json` exists at the repo root, honor its per-agent `"enabled": false` flags.
|
|
334
|
+
|
|
335
|
+
**Change-shape gate for low-yield lenses (#1254).** After the eligible roster is
|
|
336
|
+
known, drop the two low-yield code lenses (`performance-review`,
|
|
337
|
+
`correctness-review`) when the changeset has **no runtime surface** — every
|
|
338
|
+
target file is documentation or config, so those lenses would only no-op. Decide
|
|
339
|
+
deterministically with the shared helper (not by eyeballing the file list):
|
|
340
|
+
|
|
341
|
+
```bash
|
|
342
|
+
python3 "$CLAUDE_PLUGIN_ROOT/skills/code-review/scripts/change_shape.py" --files <target files>
|
|
343
|
+
```
|
|
344
|
+
|
|
345
|
+
It prints `{"hasRuntimeSurface": <bool>, "isTestOnly": <bool>, "isProseOnly":
|
|
346
|
+
<bool>, "skipLenses": [...]}`. When `skipLenses` is non-empty, exclude those
|
|
347
|
+
agents from this run and note the skip in the report (they were gated by
|
|
348
|
+
change shape, not by `Scope:`).
|
|
349
|
+
|
|
350
|
+
`isProseOnly` (#2104) reports a third, independent property: every changed
|
|
351
|
+
file is `.md`/`.mdx`. Unlike `hasRuntimeSurface`, this makes **no** exception
|
|
352
|
+
for functional Claude-config markdown (`agents/`, `skills/`, `knowledge/`,
|
|
353
|
+
`.claude/`, …) — that markdown drives agent behavior, so `performance-review`
|
|
354
|
+
and `correctness-review` still apply to it, but a `.md` file cannot exhibit an
|
|
355
|
+
injection/auth/data-exposure vulnerability, a domain-boundary leak, a
|
|
356
|
+
test-coverage gap, or a resource leak/N+1 query regardless of whether it also
|
|
357
|
+
happens to be functional config. When `isProseOnly` is true, `skipLenses`
|
|
358
|
+
additionally drops `security-review`, `domain-review`, `test-review`, and
|
|
359
|
+
`performance-review` — each of those four agents' own `## Skip` clause
|
|
360
|
+
already self-reports skip on a documentation-only target, so keeping them in
|
|
361
|
+
the roster either pays for a self-reported skip or, worse, produces an
|
|
362
|
+
ungrounded finding stretched to fit the lens (the motivating case: a 9-agent
|
|
363
|
+
panel dispatched against a single-file skill-markdown diff produced
|
|
364
|
+
elaborate security/domain/test/performance framing for what were really
|
|
365
|
+
prose nits). `correctness-review`, `spec-compliance-review`, `doc-review`,
|
|
366
|
+
`structure-review`, `naming-review`, and `arch-review` stay in the roster —
|
|
367
|
+
they meaningfully review markdown-as-instructions. This is narrower than
|
|
368
|
+
`select_lenses.py`'s own `NON_EXECUTABLE_SKIP_ELIGIBLE` allowlist, which
|
|
369
|
+
considered and rejected filtering `security-review`/`domain-review` for its
|
|
370
|
+
broader "non-executable" category (docs **and** config/lockfiles/assets) —
|
|
371
|
+
see that module's comment. This gate never widens to config, so that
|
|
372
|
+
rejection does not apply here.
|
|
373
|
+
|
|
374
|
+
`isTestOnly` (#1964) reports a second, independent property: every changed file
|
|
375
|
+
is *provably* a test file (`knowledge/test-file-indicators.md`). It currently
|
|
376
|
+
**gates nothing** — `change_shape.py`'s `TEST_ONLY_SKIP_LENSES` ships empty —
|
|
377
|
+
and exists so `/build` can stamp `diff_shape` on its `review-value.jsonl` rows
|
|
378
|
+
and `/harness-audit` can split per-lens outcomes by it. Populating that list is
|
|
379
|
+
a separate, per-lens decision that must cite the measured split, exactly as the
|
|
380
|
+
architectural-impact gate requires for widening `GATED_LENSES`. Read the field
|
|
381
|
+
for telemetry; do not narrow a roster on it until it does gate something. The gate is **fail-safe**: any
|
|
382
|
+
file it cannot prove is doc/config (source, an unknown extension, or functional
|
|
383
|
+
Claude-config markdown under `agents/`, `skills/`, `knowledge/`, `.claude/`, …)
|
|
384
|
+
counts as runtime surface and keeps every lens. This never fires on a pure-docs
|
|
385
|
+
changeset — that is already handled earlier by the documentation-only
|
|
386
|
+
short-circuit; this gate covers the doc/config-**mixed** and config-only diffs
|
|
387
|
+
the short-circuit does not. Bypassed by `--force` and by `--agent <name>` (an
|
|
388
|
+
explicit single-agent request always runs that agent).
|
|
389
|
+
|
|
390
|
+
**Change-size gate for small changesets (#1339).** After `Scope:` eligibility
|
|
391
|
+
and the change-shape gate above have both been applied, apply this gate —
|
|
392
|
+
never before, and never in a way that re-adds an agent either already removed.
|
|
393
|
+
It narrows the `Scope: always` roster by diff *size* rather than file *type*:
|
|
394
|
+
the pre-PR hook (`hooks/pre_pr_review.py`, #1886) requires a `.pr-review-passed`
|
|
395
|
+
hash match **and** (#1461, floor lowered to 1 by #2147) >= 1 distinct, recent,
|
|
396
|
+
registered review-agent dispatch recorded in the dispatch ledger — so this
|
|
397
|
+
gate must never narrow `keepAgents` below 1, and today's four-agent floor
|
|
398
|
+
(`security-review`, `correctness-review`, `spec-compliance-review`,
|
|
399
|
+
`doc-review`) clears that with room to spare. Which specific agents to keep at
|
|
400
|
+
a given diff size remains this step's decision, not the hook's — the hook
|
|
401
|
+
only enforces the *count*
|
|
402
|
+
floor, never which agents satisfy it.
|
|
403
|
+
|
|
404
|
+
**Applies only to diff-scoped reviews** — auto-scoped uncommitted changes, or
|
|
405
|
+
`--since <ref>`. `--path`, `--all`, and the full-repository fallback review
|
|
406
|
+
complete files, not a diff, so this gate never engages for those scopes
|
|
407
|
+
(existing eligibility unchanged).
|
|
408
|
+
|
|
409
|
+
Compute the numstat lines and feed them to the shared helper — for auto-scope,
|
|
410
|
+
union unstaged and staged the same way step 1 unions `--name-only`:
|
|
411
|
+
|
|
412
|
+
```bash
|
|
413
|
+
# Auto-scope (uncommitted changes):
|
|
414
|
+
{ git diff --numstat; git diff --cached --numstat; } | python3 "$CLAUDE_PLUGIN_ROOT/skills/code-review/scripts/change_size.py" --numstat-from -
|
|
415
|
+
|
|
416
|
+
# --since <ref>:
|
|
417
|
+
git diff --numstat <ref>...HEAD | python3 "$CLAUDE_PLUGIN_ROOT/skills/code-review/scripts/change_size.py" --numstat-from -
|
|
418
|
+
```
|
|
419
|
+
|
|
420
|
+
It prints `{"filesChanged": <int>, "addedLines": <int>, "qualifiesForFastPath":
|
|
421
|
+
<bool>, "keepAgents": [...]}`. When `qualifiesForFastPath` is `true`, drop
|
|
422
|
+
every `Scope: always` agent **not** in `keepAgents` (today: `security-review`,
|
|
423
|
+
`correctness-review`, `spec-compliance-review`, `doc-review` — the four lenses
|
|
424
|
+
that stay meaningful at any diff size; the rest are code-quality-at-scale
|
|
425
|
+
concerns a diff this small essentially cannot exhibit meaningfully) and note
|
|
426
|
+
the drop in the report (gated by change size, not by `Scope:`).
|
|
427
|
+
`Scope:`-glob-matched agents are unaffected — they already run only against
|
|
428
|
+
matching file types, so a diff this small already narrows their incremental
|
|
429
|
+
cost to near-zero. The gate is **fail-safe**: any `git diff --numstat` error,
|
|
430
|
+
binary-file marker, or unparseable line disqualifies the run (full panel), as
|
|
431
|
+
does any file under `hooks/` or `skills/code-review/` (the enforcement
|
|
432
|
+
machinery and this gate's own orchestration) — a change there is exactly the
|
|
433
|
+
case where a cheap, self-certifying review is a problem, so it never qualifies
|
|
434
|
+
for the shortcut it defines, regardless of size. Bypassed by `--force` and by
|
|
435
|
+
`--agent <name>`, matching the change-shape gate's bypass list.
|
|
436
|
+
|
|
437
|
+
**Diff-signal gate for structural and concurrency lenses.** Apply this
|
|
438
|
+
**third**, after `Scope:` eligibility, the change-shape gate, and the
|
|
439
|
+
change-size gate — never before, and never to re-add an agent an earlier gate
|
|
440
|
+
already removed. It narrows by *what the diff's content proves is absent*
|
|
441
|
+
rather than by file type or diff size, and it gates two lenses:
|
|
442
|
+
`arch-review` (structural signals) and `concurrency-review` (concurrency
|
|
443
|
+
primitives, #1975).
|
|
444
|
+
|
|
445
|
+
`arch-review` is `Scope: always` and opus-tier, so it runs on every non-empty
|
|
446
|
+
changeset — including diffs that cannot exhibit what it looks for. Its scope
|
|
447
|
+
is ADR compliance, layer-boundary violations, dependency direction, and
|
|
448
|
+
pattern consistency: all properties of *structure*. A diff that adds a guard
|
|
449
|
+
clause inside an existing function, with no import change, no added/moved/
|
|
450
|
+
deleted file, no manifest edit, and no public-interface change, has moved no
|
|
451
|
+
boundary for it to evaluate. Decide deterministically:
|
|
452
|
+
|
|
453
|
+
```bash
|
|
454
|
+
# Auto-scope (uncommitted changes):
|
|
455
|
+
{ git -c diff.relative=false diff --no-color; git -c diff.relative=false diff --cached --no-color; } \
|
|
456
|
+
| python3 "$CLAUDE_PLUGIN_ROOT/skills/code-review/scripts/change_impact.py" --files <target files>
|
|
457
|
+
|
|
458
|
+
# --since <ref>:
|
|
459
|
+
git -c diff.relative=false diff --no-color <ref>...HEAD \
|
|
460
|
+
| python3 "$CLAUDE_PLUGIN_ROOT/skills/code-review/scripts/change_impact.py" --files <target files>
|
|
461
|
+
```
|
|
462
|
+
|
|
463
|
+
It prints `{"signals": [...], "hasArchitecturalImpact": <bool>, "skipLenses":
|
|
464
|
+
[...], "reason": <str|null>}`. Exclude any agent in `skipLenses` and note the
|
|
465
|
+
skip in the report (gated by diff signal, not by `Scope:`). The seven signals
|
|
466
|
+
are `structure` (file added/deleted/renamed), `dependency` (an import/require
|
|
467
|
+
line added or removed), `manifest`, `infra`, `interface` (a public/exported
|
|
468
|
+
symbol declaration added or removed), `adr`, and `concurrency` (a concurrency
|
|
469
|
+
primitive added, removed, **or visible in a hunk's context lines** —
|
|
470
|
+
async/await, threads, locks, channels, atomics; removing synchronization is a
|
|
471
|
+
concurrency change exactly as much as adding it, and the context lines are
|
|
472
|
+
what keep the lens on a body-only edit inside an already-locked block, which
|
|
473
|
+
carries no primitive on its own changed line). `hasArchitecturalImpact` reports only the first six: a diff whose
|
|
474
|
+
sole signal is `concurrency` has moved no boundary, so it keeps
|
|
475
|
+
`concurrency-review` while `arch-review` still drops.
|
|
476
|
+
|
|
477
|
+
The gate is **fail-safe and include-biased**: an unparseable diff, an empty
|
|
478
|
+
diff, or any file it cannot classify all count as impact and keep every lens.
|
|
479
|
+
It can only remove a lens it can prove has nothing to look at. Bypassed by
|
|
480
|
+
`--force` and `--agent <name>`, matching the other two gates' bypass list.
|
|
481
|
+
|
|
482
|
+
**Only these two lenses are gated, deliberately.** Both pass the same test —
|
|
483
|
+
the lens's subject must be *provably absent* from the diff, not merely
|
|
484
|
+
unlikely to appear in it: `arch-review` reviews structure, and
|
|
485
|
+
`concurrency-review` reviews races, async ordering, idempotency, and
|
|
486
|
+
shared-state safety, none of which can exist where no concurrency primitive
|
|
487
|
+
does. `domain-review`
|
|
488
|
+
is the obvious next candidate and is excluded on purpose: its scope covers
|
|
489
|
+
"business logic placement", and putting business logic into a controller
|
|
490
|
+
method body is a real violation introduced by a *body-only* edit with no
|
|
491
|
+
structural signal — exactly the diff shape this gate skips. Widen
|
|
492
|
+
`GATED_LENSES` from #1624's measured per-agent data, not from intuition about
|
|
493
|
+
which lens probably no-ops. Same evidence-first discipline
|
|
494
|
+
`knowledge/verification-mode.md` applies to tier-down opt-ins.
|
|
495
|
+
|
|
496
|
+
### 4. Run each enabled agent
|
|
497
|
+
|
|
498
|
+
**Dispatch-capability gate (re-confirm here, not just at the top of this file — issue #1461).** Before spawning anything below, re-verify the `Agent`/`Task` tool is present in this toolset. If it is not, STOP per the Orchestrator constraints above — do not fall back to reviewing the files yourself, inline, as a stand-in for the panel; report the missing capability and halt the run before any agent is spawned.
|
|
499
|
+
|
|
500
|
+
**Dispatch batching — bounded dispatch waves (issue #1752).** A real run that spawned all 16 eligible agents as parallel `Agent` calls in one message lost its last 6 to `[Tool result missing due to internal error]` — see `dispatch_waves.py`'s module docstring for the full incident account; not restated here to avoid two copies drifting apart. Before spawning, compute the wave split deterministically instead of guessing a safe batch size by eye:
|
|
501
|
+
|
|
502
|
+
```bash
|
|
503
|
+
sh "$CLAUDE_PLUGIN_ROOT/hooks/py.sh" "$CLAUDE_PLUGIN_ROOT/skills/code-review/scripts/dispatch_waves.py" --agents "<comma-separated eligible agent names, cheap-first order as select_lenses.py returned them, filtered by the change-shape/change-size/change-impact gates but not re-sorted>"
|
|
504
|
+
```
|
|
505
|
+
|
|
506
|
+
Prints `{"maxParallel": N, "waves": [[...], [...]]}` — `maxParallel` defaults to **10**, overridable with `DEV_TEAM_MAX_PARALLEL_REVIEW_AGENTS` (see the script's own docstring for the exact fallback rule; don't re-derive it here). Dispatch **exactly the waves the script printed, in that order** as parallel subagents in a single message per wave using the Agent tool — exactly as before, just bounded per message — waiting for each wave to fully return before dispatching the next, and for the last wave before aggregating. A roster no larger than `maxParallel` is always a single wave; nothing changes from today's behavior in that case.
|
|
507
|
+
|
|
508
|
+
**Optional: shared context pack (#2006, opt-in — off by default).** `scripts/review_context_pack.py` can prepare the panel's file context **once** — changed-file list, diff, and complete line-numbered file bodies — so each lens reads one prepared artifact instead of opening the same changed files itself.
|
|
509
|
+
|
|
510
|
+
**Do not use it unless the caller explicitly opts in** (`DEV_TEAM_REVIEW_CONTEXT_PACK=on`). The default dispatch path is the per-agent context payload described below, unchanged.
|
|
511
|
+
|
|
512
|
+
Why it is off by default: [ADR 0034](../../../../docs/adr/0034-do-not-build-shared-context-pre-pass-for-duplicate-full-file-reads-1611.md) declined exactly this pre-pass after measuring duplicate full-file reads at **0.38%–4.86% of a round's total input spend, median 0.8%** (#1618). A later re-measurement with the same tool (`scripts/measure_full_file_duplication.py`) put it at **0.22%**. The 4.31x figure sometimes quoted for this is a *read-volume ratio*, not a share of spend — a different denominator, and not the one this decision turns on. The pack ships full file bodies rather than the structural skeleton ADR 0034 warned would degrade line-level lenses, so it carries no known quality risk; it simply has not been shown to pay for itself. #2024 tracks measuring panels of >= 8 agents, which is the shape that could change the answer.
|
|
513
|
+
|
|
514
|
+
When opted in:
|
|
515
|
+
|
|
516
|
+
```bash
|
|
517
|
+
git -c core.quotePath=false diff --name-status <base> \
|
|
518
|
+
| sh "$CLAUDE_PLUGIN_ROOT/hooks/py.sh" "$CLAUDE_PLUGIN_ROOT/skills/code-review/scripts/review_context_pack.py" \
|
|
519
|
+
--name-status-from - --base <base>
|
|
520
|
+
```
|
|
521
|
+
|
|
522
|
+
Prints a manifest naming the pack path, its byte size, and `files_omitted`. Pass the pack path to each dispatched agent, and observe both rules:
|
|
523
|
+
|
|
524
|
+
- **The pack narrows repeated reads, not the review.** A lens may still open anything not in the pack — a caller in an unchanged file, a sibling module. Never instruct an agent to treat the pack as the complete world.
|
|
525
|
+
- **`files_omitted` is not optional to relay.** When the manifest reports omissions (a file over the per-file cap, a binary, a body that would exhaust the budget), name those paths in each agent's prompt and tell it to open them directly. The pack body says so too, but a silently skipped file is a coverage hole that reads as a clean review.
|
|
526
|
+
|
|
527
|
+
- **File scope**: pass only files matching each agent's declared scope. Skip the agent if no files match.
|
|
528
|
+
- **Ledger-scoped dispatch (#2167).** Once every agent's File scope above is known, consult the per-lens verdict ledger before building any dispatch prompt — a repeat review must not pay to re-derive an outcome an exact `(lens, file, content)` match already recorded:
|
|
529
|
+
```bash
|
|
530
|
+
python3 "$CLAUDE_PLUGIN_ROOT/scripts/verdict_scope.py" --root . --lens-files '<JSON: {"<agent>": [<its File-scope files from the bullet above>], ...} for every agent surviving step 3''s gates>'
|
|
531
|
+
```
|
|
532
|
+
Prints `{"toDispatch": {<agent>: [...]}, "skipped": {<agent>: [{"file": ..., "verdict": {...}}]}, "fullySkippedLenses": [...]}`.
|
|
533
|
+
- Narrow each agent's File scope to `toDispatch[<agent>]` before building its dispatch prompt and Scope marker (below) — its ledger-cleared files are simply removed from both, never added to. An agent named in `fullySkippedLenses` gets exactly the same treatment the File-scope bullet above already describes for "no files match": do not dispatch it this round, and leave its wave slot unused — no change to how `dispatch_waves.py` is invoked above.
|
|
534
|
+
- **Report loudly, never silently.** A second `/code-review` over an unchanged target set must dispatch zero lenses AND say why: the report (step 7) and `--json` output name every ledger-skipped `(lens, file)` pair together with the matched row's `ts`/`file_content_hash`/`plugin_version` (`verdict_scope.py`'s own `skipped[<agent>]` entries carry this verbatim). A run whose entire roster lands in `fullySkippedLenses` states so explicitly ("N lenses skipped via ledger evidence, 0 dispatched this run") — an empty per-agent table with no such note is indistinguishable from a broken run and must never appear.
|
|
535
|
+
- **Fail closed.** `verdict_scope.py` skips a file only on an exact `(lens, file_path, current content hash)` match whose most-recent ledger row is `outcome: "pass"`. A missing/unreadable ledger, a hash it couldn't compute (deleted/unreadable/oversized file), a `findings`-outcome row, or a stale `plugin_version` row (`hooks/lib/review_verdicts.py`'s own `_is_usable_version`) all dispatch that file normally — see that script's module docstring for the exhaustive list. This step only ever narrows the File scope computed above; it never widens it.
|
|
536
|
+
- **The PR gate needs no change here.** `hooks/pre_pr_review.py`'s dispatch-ledger corroboration (step 9, below) is keyed on the branch/staged-diff's own `subject_hash`, independent of this per-file ledger: an unchanged target set that fully ledger-skips also has an unchanged `subject_hash`, so the same dispatch-ledger row a prior, genuine dispatch already stamped for that hash still corroborates the gate — "reviewed, nothing to do" and "never reviewed" stay distinguishable through that pre-existing mechanism, not a new one this step would have to invent. Whenever at least one file's content actually changed, at least one real dispatch happens for it and stamps a fresh row for the new `subject_hash`, exactly as before this step existed.
|
|
537
|
+
- **`overall` (step 5) is unaffected by which files were skipped here.** A fully ledger-skipped agent contributes no *new* finding (its last verdict for this exact content was already `pass`), so it can never turn `overall` to `warn`/`fail` on its own — `--json` computes the same `overall` a from-scratch full-panel run over this exact content would.
|
|
538
|
+
- **Scope marker (#2166)**: append one structured, single-line marker to every dispatch prompt, listing the exact files passed under File scope above (as narrowed by Ledger-scoped dispatch, immediately above), comma-separated: `Files in scope for this review: <path>, <path>, ...`. This is metadata for `hooks/review_verdict_recorder.py` (#2166 Step 2.3), the `SubagentStop` verdict recorder that parses it back out of the transcript.
|
|
539
|
+
- **Context payload** (controlled by the agent's `Context needs`):
|
|
540
|
+
- `diff-only` → diff output only (for auto-scope or `--since` only)
|
|
541
|
+
- `full-file` → complete files
|
|
542
|
+
- `project-structure` → full files + directory tree + the changed-file list (path + change type — `A`/`M`/`D`/`R`/`C`) computed in step 3 via `changed_file_list.py`, for diff-scoped runs (auto-scope or `--since`). Omit the changed-file list for `--path`/`--all`/full-repository scope — there is no diff to describe. Every `project-structure` agent (`arch-review`, `doc-review`, `domain-review`) has no Bash grant and must never invoke `git` itself to "see what changed" — that call is denied and surfaces as a spurious error (#1734); this payload is what makes that unnecessary.
|
|
543
|
+
- When reviewing full repository (clean auto-scope, `--all`, or `--path`), always pass full files.
|
|
544
|
+
- **Model**: pass each agent's declared `model:`/`effort:` frontmatter. The harness resolves both fields natively before dispatch, per `agents/orchestrator.md` → Model/Effort Resolution (ADR 0026).
|
|
545
|
+
- **Static analysis context**: if step 2b produced findings, inject into every agent's prompt using the format in `skills/static-analysis-integration/SKILL.md`: "These issues were detected by static analysis. Do not re-report them. Focus on semantic concerns."
|
|
546
|
+
- **Per-agent output**: the shared contract in [`knowledge/review-agent-output-contract.md`](../../knowledge/review-agent-output-contract.md), wrapped with `agentName`/`modelTier` (full aggregation shape in `output-format.md`).
|
|
547
|
+
|
|
548
|
+
**Graph-assisted review**: pass tool availability to **all read-only review agents** — the structural lenses (`arch-review`, `component-architecture-review`, `structure-review`, `domain-review`) benefit most from resolved call graphs, but every lens gains cheaper verified reads — so they may consult the index for impact/dependency context before flagging findings. Tool selection and the fallback contract are the same as step 1c above; see [`knowledge/codegraph-vs-graphify.md`](../../knowledge/codegraph-vs-graphify.md).
|
|
549
|
+
|
|
550
|
+
**Dispatch failure handling — retry once, never drop silently (issue #1752).** After **each wave** returns, check every agent dispatched **in that wave** (not the full eligible roster — a later wave hasn't dispatched yet) for a valid per-agent result matching [`review-agent-output-contract.md`](../../knowledge/review-agent-output-contract.md).
|
|
551
|
+
|
|
552
|
+
**Contract-valid is a per-agent classification, computed, never eyeballed (issue #1998).** Session-report analysis found 18.2% of review-agent outputs silently discarded because "does this look like the contract?" was a judgment call with no record of what it decided or why. For **each** agent that returned *something* this wave (not the ones that errored out with no return at all — those go straight to `dispatch_reconcile.py`'s missing set), write its raw final-turn text with the **Write tool** — never a shell heredoc, `echo`, or `printf`, which would interpolate agent-produced text into a command line — to a file named `.claude/memory/contract-raw-<agent>.md` (the `.md` suffix matters: only that suffix, not an arbitrary filename under `.claude/memory/`, is covered by this repo's own `.gitignore`; project-local), then classify it deterministically before deciding whether it belongs in `--returned` below. Delete the file once classification completes — including when `validate_review_output.py` itself errors out — so unredacted raw agent output never lingers on disk longer than necessary:
|
|
553
|
+
|
|
554
|
+
```bash
|
|
555
|
+
sh "$CLAUDE_PLUGIN_ROOT/hooks/py.sh" "$CLAUDE_PLUGIN_ROOT/skills/code-review/scripts/validate_review_output.py" --agent "<name>" --file "<path to that agent's raw output>"
|
|
556
|
+
```
|
|
557
|
+
|
|
558
|
+
Exit 0 (`"valid": true`) means include this agent's name in `dispatch_reconcile.py`'s `--returned` list below and proceed to parse its JSON normally. Exit 1 means it does **not** go in `--returned` — the script has already appended a diagnostic (agent name, a secret-redacted prefix of the raw output, and the specific validation error, classified by `shape` — see [`knowledge/telemetry-schema.md`](../../knowledge/telemetry-schema.md)'s `contract-failures.jsonl` section for the closed set of shapes, not re-enumerated here to avoid the two copies drifting apart) to `.claude/metrics/contract-failures.jsonl`, so this is no longer a silent drop. A `schema-drift` or `malformed-json` failure additionally carries `extraction`, naming which extraction path (`clean`/`fenced`/`prose-preamble`) recovered the object before it failed, so that information survives even when the object arrived fenced or behind a prose preamble. Carry the printed `shape`/`extraction`/`error` forward into this agent's step 4 item 3 `dispatchFailures` entry below (as their own fields, not concatenated into `error`) instead of the generic `"output that doesn't parse against the contract"` text, so the report names the actual reason. `skills/code-review/scripts/contract_failure_report.py --json` joins that log against `agent_dispatch_ledger.py`'s dispatch counts to report a real per-agent contract-parse failure *rate* — consult it before citing any per-lens `$/finding` figure, which divides by this same denominator.
|
|
559
|
+
|
|
560
|
+
Now compute the dispatched-vs-returned coverage check deterministically over the resulting `--returned` set instead of eyeballing the two lists:
|
|
561
|
+
|
|
562
|
+
```bash
|
|
563
|
+
sh "$CLAUDE_PLUGIN_ROOT/hooks/py.sh" "$CLAUDE_PLUGIN_ROOT/skills/code-review/scripts/dispatch_reconcile.py" --dispatched "<this wave's dispatched agent names, in the order dispatch_waves.py listed them>" --returned "<this wave's contract-valid agent names>"
|
|
564
|
+
```
|
|
565
|
+
|
|
566
|
+
Both flags are required — pass an empty string (`--returned ""`) when no agent in the wave returned a contract-valid result (the whole-wave-loss case #1752 exists for). Prints `{"missing": [...]}` — every name it returns is a **dispatch failure**: a call that came back as `[Tool result missing due to internal error]`, with no `agentId`, or with output that doesn't parse against the contract per the deterministic check above, so it never produced a contract-valid return. Distinct from `skip` (agent had nothing to review this run — [`review-agent-output-contract.md`](../../knowledge/review-agent-output-contract.md#status-values)) and from `fail` (agent ran and found errors); it means the lens never actually ran.
|
|
567
|
+
|
|
568
|
+
1. Retry each failed agent **exactly once**, individually — same prompt, model, context payload, and file scope as the original call, dispatched on its own (not re-batched with the rest of that wave).
|
|
569
|
+
2. Run the retry's raw output through `validate_review_output.py` the same way as the first attempt above before deciding it succeeded. If it comes back valid, use its result and continue as normal — this never shows up as a failure in the final report. A recovered dispatch — one that fails once but succeeds on its single retry — never reaches the dispatch-failure emission point in step 3 below: no `dispatch-failure` boundary event is ever emitted for it.
|
|
570
|
+
3. If the retry also fails contract validation, do **not** proceed as if that lens's coverage were complete:
|
|
571
|
+
- Carry it into step 5's aggregation as a `dispatchFailures` entry (`{agentName, attempts: 2, error, shape, extraction}`), using the retry attempt's `validate_review_output.py` `shape`/`extraction`/`error` when the retry returned unparseable output — the same specific reason already logged to `contract-failures.jsonl` — or the transport error text (`"Tool result missing due to internal error"`, etc.) in `error` with `shape`/`extraction` both `null` when it never returned at all. The `dispatchFailures` key itself is always present in `--json` output (an empty array when there are none, per `output-format.md`); the prose report's `## Dispatch Failures` section renders only when the array is non-empty, and is never omitted in that case because "the rest of the panel passed."
|
|
572
|
+
- At this same moment — the point an unrecovered failure is determined — also emit a `dispatch-failure` boundary event, bound to the `subject_hash` in effect for this dispatch (the same `branch_diff_gate_hash(default_base_ref(cwd), cwd)` value this file's doc-only/single-agent exemption calls already compute, per #1886) — so `hooks/pre_pr_review.py`'s own `_dispatch_failure_verdict` can find it at `gh pr create` time:
|
|
573
|
+
```bash
|
|
574
|
+
HASH=$(python3 "${CLAUDE_PLUGIN_ROOT}/hooks/lib/review_gate_hash.py" --branch-diff)
|
|
575
|
+
python3 "${CLAUDE_PLUGIN_ROOT}/hooks/lib/boundary_events.py" --event dispatch-failure --agent "<name>" --subject-hash "$HASH"
|
|
576
|
+
```
|
|
577
|
+
- Treat it as fail-equivalent for step 9's gate-write condition — the same treatment step 3's `unreadable-registry`/`unreadable-files-from` handling already gets: a lens that never ran is a coverage gap, not a passing result, so `.pr-review-passed` must not be written while any dispatch failure is outstanding.
|
|
578
|
+
- State plainly, in both prose and `--json` output, which agent(s) failed twice and the error text — a missing lens must always be visible, never inferred from a shorter-than-expected agent table.
|
|
579
|
+
|
|
580
|
+
### 5. Aggregate results
|
|
581
|
+
|
|
582
|
+
**Fold in dispatch failures first (issue #1752).** Before scoring or suppression, add every step 4 `dispatchFailures` entry (agents that failed dispatch, then failed their single retry) to the aggregation. They are not agent results — they carry no `issues[]` and never enter ACCEPTED-RISKS suppression or health scoring — but they are never dropped either: carry the full `dispatchFailures` list through to the report (step 7) and the `--json` object (`output-format.md`) unchanged, and remember it for step 9's gate condition.
|
|
583
|
+
|
|
584
|
+
**A non-empty `dispatchFailures` forces `overall: "fail"`, unconditionally (issue #1752).** This is not the same rule as step 9's gate-blocking condition below — it belongs here, in the aggregate itself, because step 9 (and its gate) is **skipped entirely under `--json`** (step 7), while `overall` is the one field every `--json` caller reads. `/pr --json` (the sole such caller) checks only `overall`/`status` before proceeding to open a PR; without this rule here, a lens that failed dispatch twice could sit invisibly behind an `overall: "pass"` computed only from the agents that did return, and `/pr` would open the PR anyway — the exact silent-coverage-gap failure mode #1752 exists to close, just reached through a different caller than the interactive gate. Apply this override after health scoring computes what `overall` would otherwise be, so it always wins regardless of the per-agent severity mix.
|
|
585
|
+
|
|
586
|
+
**Fold in ledger skips the same way (#2167).** Carry step 4's `verdict_scope.py` `skipped` map through to the report and the `--json` object as `ledgerSkipped` (`output-format.md`), unchanged and never dropped, even when it is every agent in the roster. Unlike `dispatchFailures`, a ledger skip does **not** force `overall` toward `fail`/`warn` — it means a lens already reached `pass` on this exact content, not that it never ran — so score `overall` from the agents that actually dispatched this round exactly as if the ledger-skipped ones were absent from the roster entirely (never absent from the *report*, only from the score). A roster that is ENTIRELY ledger-skipped therefore reports `overall: "pass"` with zero `agents[]` entries and a non-empty `ledgerSkipped` — the report and summary text must say so explicitly (see step 4's "report loudly" rule) so that state is never mistaken for an empty, broken, or unreviewed run.
|
|
587
|
+
|
|
588
|
+
#### 5a. Apply ACCEPTED-RISKS.md
|
|
589
|
+
|
|
590
|
+
If `ACCEPTED-RISKS.md` exists at the repo root, parse its `rules:` YAML frontmatter per `knowledge/accepted-risks-schema.md`. For each finding, check rules in declaration order; the first match suppresses and emits one audit entry:
|
|
591
|
+
|
|
592
|
+
```
|
|
593
|
+
SUPPRESSED: <file>:<line> [<rule_id>] by ACCEPTED-RISKS rule <rule.id>
|
|
594
|
+
```
|
|
595
|
+
|
|
596
|
+
- Expired rules become inert: stop suppressing, emit a WARN naming the rule and owner, list in an Expiry Report section.
|
|
597
|
+
- Rules with `broad: true` (wildcard `rule_id` or multi-file globs) emit an informational notice for auditor attention.
|
|
598
|
+
- Schema-invalid rules fail the run with a parse error naming the rule id.
|
|
599
|
+
|
|
600
|
+
Suppressed findings are removed from scoring, listed under "Suppressed by ACCEPTED-RISKS" in the report (grouped by rule id), and bypass the fix loop.
|
|
601
|
+
|
|
602
|
+
#### 5b. Health scoring
|
|
603
|
+
|
|
604
|
+
Read `knowledge/review-rubric.md` for the formula. Compute the overall health score; security failures auto-escalate to 🔴.
|
|
605
|
+
|
|
606
|
+
Classify each issue by actionability:
|
|
607
|
+
|
|
608
|
+
| Severity | Confidence | Actionable? |
|
|
609
|
+
| --- | --- | --- |
|
|
610
|
+
| error or warning | high or medium | **Yes** — auto-apply |
|
|
611
|
+
| error or warning | none | No — report only (human judgment) |
|
|
612
|
+
| suggestion | any | No — report only |
|
|
613
|
+
|
|
614
|
+
**Actionable issues** drive the fix loop.
|
|
615
|
+
|
|
616
|
+
#### 5b-i. Record round 1 (#1624)
|
|
617
|
+
|
|
618
|
+
The initial panel is **round 1**. Append its row to
|
|
619
|
+
`.claude/metrics/review-value.jsonl` now, before any fix is applied — this
|
|
620
|
+
stream is what makes #1623's "is this churn or value?" question answerable at
|
|
621
|
+
all, and a row written only on the happy path would bias every derived metric:
|
|
622
|
+
|
|
623
|
+
> **Reading this stream later.** `scripts/review_value_coverage.py` reconciles
|
|
624
|
+
> these rows against `agent_dispatch_ledger`'s deterministic dispatch records
|
|
625
|
+
> and rules on whether the sample can support a per-lens pruning decision
|
|
626
|
+
> (`no-data` / `unverifiable` / `undercollected` / `insufficient` / `biased` /
|
|
627
|
+
> `usable`). Both writers of this stream are triggered by agent instruction
|
|
628
|
+
> rather than by mechanism, so rows skew toward rounds that found something
|
|
629
|
+
> (#2019). `/harness-audit` step 4 consults it before citing any per-lens
|
|
630
|
+
> value; nothing in this skill needs to run it.
|
|
631
|
+
|
|
632
|
+
```bash
|
|
633
|
+
python3 "$CLAUDE_PLUGIN_ROOT/skills/code-review/scripts/review_round_log.py" \
|
|
634
|
+
--round 1 --agents "<comma-separated agents dispatched>" \
|
|
635
|
+
--findings <path-to-this-round's-findings.json> \
|
|
636
|
+
--purpose discovery --outcome "<fixed|no-op|escalated>"
|
|
637
|
+
```
|
|
638
|
+
|
|
639
|
+
Round 1 never passes `--fix-diff`: it has no preceding fix, so its
|
|
640
|
+
`fix_provenance_new` is `0` by definition. The script writes counts, agent
|
|
641
|
+
names, and enum values only — never file paths, code, or finding text.
|
|
642
|
+
Full schema: `knowledge/telemetry-schema.md` § `review-value.jsonl`.
|
|
643
|
+
|
|
644
|
+
Every later round records itself the same way from step 6a — see that step's
|
|
645
|
+
"Record each round" item for the `--fix-diff` argument that turns
|
|
646
|
+
`fix_provenance_new` into the "the previous fix introduced this" signal.
|
|
647
|
+
|
|
648
|
+
#### 5c. Consolidate cross-agent findings
|
|
649
|
+
|
|
650
|
+
When multiple agents flag the same `file:line`, emit one `topFindings` entry: `severity` = the single **highest** enum for that finding, `agents` = an array of the reporting agents (e.g. `["structure-review", "correctness-review"]`). Never pack multiple values into `severity` or any agent scalar — no slash- or comma-joined strings. Every scalar field stays single-valued; multi-agent attribution lives only in the `agents: []` array. Schema: [`output-format.md`](output-format.md#aggregated-json-result---json-flag).
|
|
651
|
+
|
|
652
|
+
**Dedup across agents, not just across identical lines — prose only, never the `topFindings` array itself.** The `topFindings` JSON array keeps the existing exact `file:line` dedup key unchanged — one entry per distinct `file:line`, matching `output-format.md`'s contract and `scripts/consolidate.py`'s sliced-mode dedup key. The instruction below governs only how findings are *described in the human-facing prose summary/report*: when writing that prose, collapse any two findings — from different agents, even at slightly different lines — that describe the same underlying defect into a single description; do not restate the same defect twice in prose just because two agents (or two nearby lines) reported it.
|
|
653
|
+
|
|
654
|
+
**Condensation cap.** Condense each surviving finding to ≤ 3 lines per finding before final synthesis output — the essential defect description and fix, not each agent's full reasoning. Applies only to the human-facing summary/report; `topFindings` entries keep their full `message`/`suggestedFix` text unchanged.
|
|
655
|
+
|
|
656
|
+
### 6. Present findings and ask for direction
|
|
657
|
+
|
|
658
|
+
If zero actionable issues, skip to step 7.
|
|
659
|
+
|
|
660
|
+
Otherwise present the Review Findings prompt (template: [`output-format.md`](output-format.md#review-findings-prompt-interactive--step-6)) and ask: **"Fix these issues automatically, or save as report only?"**
|
|
661
|
+
|
|
662
|
+
- "Fix" / "apply" / "yes" → step 6a
|
|
663
|
+
- "Report" / "no" / "don't fix" → step 7 (no code modified)
|
|
664
|
+
|
|
665
|
+
**Exception — non-interactive mode**: skip this prompt when the run is non-interactive.
|
|
666
|
+
|
|
667
|
+
- (a) If `--json` (or `--yes`), **default to report only** — proceed to step 7 and emit the aggregated JSON; **never modify code** without an explicit caller opt-in. `--json` is contractually non-interactive (CI-safe): it never blocks on this prompt.
|
|
668
|
+
- (b) If running inside `/build`, `/pr`, or `/test-improve`, proceed to the fix loop. The caller owns the human gate (the orchestrator's Phase 3 approval for `/build`; the pre-PR confirmation for `/pr`; for `/test-improve`, the Phase 3 Story-set approval gating entry to Phase 4 and the `[r]evise/[w]aive/[q]uit` prompt raised after 2 failed iterations of its own end-of-phase review loop — see `../test-improve/SKILL.md`'s Phase 4/5 "End-of-phase review loop" sections).
|
|
669
|
+
|
|
670
|
+
### 6a. Review-fix loop
|
|
671
|
+
|
|
672
|
+
```
|
|
673
|
+
iteration = 1
|
|
674
|
+
MAX_ITERATIONS = 5
|
|
675
|
+
|
|
676
|
+
while actionable_issues > 0 AND iteration ≤ MAX_ITERATIONS:
|
|
677
|
+
1. Apply fixes for all actionable issues (file-by-file, top-to-bottom by line)
|
|
678
|
+
2. After each iteration's fixes, run the project's test suite.
|
|
679
|
+
If tests fail, revert the last fix that broke them and mark the
|
|
680
|
+
issue [auto-fix failed — human review required].
|
|
681
|
+
3. **When the review was auto-scoped to uncommitted changes**, stage the
|
|
682
|
+
fixes just applied (`git add` the modified files) — an Edit/Write only
|
|
683
|
+
touches the working tree, it does not change `git diff --cached`, so
|
|
684
|
+
without this the fixes would never reach the eventual commit (#1461
|
|
685
|
+
security re-review: an earlier draft's step 1 claimed a working-tree
|
|
686
|
+
edit "naturally" changes the staged hash — false for
|
|
687
|
+
`sha256(git diff --cached)`, and it silently dropped every fix-loop
|
|
688
|
+
iteration's output from the final commit). For `--path`/`--all`
|
|
689
|
+
scopes, leave the index untouched — no gate is ever written for those
|
|
690
|
+
scopes, so staging here would only mutate the operator's index
|
|
691
|
+
unasked, for no corroboration benefit. **`--since <base>` also writes
|
|
692
|
+
a gate file (issue #1904 Bug 2b — no longer "the only scope", per step
|
|
693
|
+
9's own extended condition), but has no staging concept to mirror
|
|
694
|
+
here at all**: its content is already-committed history
|
|
695
|
+
(`base...HEAD`), so a fix applied mid-loop would need a fresh COMMIT,
|
|
696
|
+
not a `git add`, to change what the eventual `--branch-diff` hash
|
|
697
|
+
covers — a disclosed gap in this loop's mechanics for that scope, not
|
|
698
|
+
fixed here; the closing pass below stays scoped to the auto-scope
|
|
699
|
+
staging model for the same reason.
|
|
700
|
+
3b. **Deterministic-first triage (#1610) — language-agnostic, not
|
|
701
|
+
Python-specific.** Before re-dispatching an agent to re-verify a fix,
|
|
702
|
+
check whether the fix already qualifies for a cheaper, deterministic
|
|
703
|
+
close: (a) it is a pure rename/mechanical edit (docstring correction,
|
|
704
|
+
import fix, identifier rename), (b) **whichever language-appropriate
|
|
705
|
+
lint/type-check tool(s) step 2b's static-analysis pre-pass already
|
|
706
|
+
detected and ran for this repo** — Tier 1 in
|
|
707
|
+
`skills/static-analysis-integration/references/tool-configs.md`
|
|
708
|
+
(semgrep + ruff/mypy for Python, pmd for Java/Kotlin, ESLint/tsc for
|
|
709
|
+
JS/TS, `dotnet format`/`dotnet build` for C#, gofmt/`go vet` for Go,
|
|
710
|
+
etc. — whatever the target project's own stack is, never assume
|
|
711
|
+
Python) — plus the full test suite already ran clean in step 2, and
|
|
712
|
+
(c) the specific claim needing verification is itself checkable by a
|
|
713
|
+
targeted `grep`/diff (e.g. "every occurrence was renamed, no
|
|
714
|
+
partial/mangled identifiers", "the removed import has no remaining
|
|
715
|
+
references"). When all three hold, run that deterministic check now
|
|
716
|
+
and mark the issue resolved on a pass — do not spend a re-dispatch
|
|
717
|
+
confirming what the language's own lint/test/grep tooling already
|
|
718
|
+
proved. Escalate to the normal per-agent re-dispatch (step 4) whenever
|
|
719
|
+
any condition fails to hold, or the check itself can't fully close the
|
|
720
|
+
question (e.g. judging whether a restored docstring's *prose* is
|
|
721
|
+
accurate needs semantic reading, not a grep). This is a triage habit,
|
|
722
|
+
not a gate: it only ever *removes* work from step 4, never adds new
|
|
723
|
+
issues or skips a fix that genuinely needs judgment. The same triage
|
|
724
|
+
applies to ad-hoc fix-verification inside `/build`'s inline review
|
|
725
|
+
checkpoints (`../build/SKILL.md` sub-steps 4/6) — one shared habit,
|
|
726
|
+
not a duplicated checklist.
|
|
727
|
+
4. Re-run only the agents whose remaining actionable issues were not
|
|
728
|
+
already closed by step 3b's deterministic triage, **in verification
|
|
729
|
+
mode** (#1628) — pass the finding, the fix diff hunks ± ~20 lines,
|
|
730
|
+
and the agent's lens definition, NOT the full target file set, and
|
|
731
|
+
grant the mandatory `insufficient-context` escape. Resolve each
|
|
732
|
+
agent's verification tier with `python3
|
|
733
|
+
"$CLAUDE_PLUGIN_ROOT/scripts/verify_tier.py" --agent <name>`. Full
|
|
734
|
+
contract: [`knowledge/verification-mode.md`](../../knowledge/verification-mode.md).
|
|
735
|
+
Carry forward statuses of agents that passed.
|
|
736
|
+
5. Re-aggregate. Reclassify remaining issues.
|
|
737
|
+
5a. **Classify the round against the ledger (#1625).** Run the round
|
|
738
|
+
ledger (below). It decides new-vs-carried by finding signature and
|
|
739
|
+
returns `terminate`/`reason` — honor it: `converged` and `round-cap`
|
|
740
|
+
both leave this loop.
|
|
741
|
+
5b. **Record this round (#1624).** Append one row per re-dispatch round
|
|
742
|
+
to `.claude/metrics/review-value.jsonl`, passing THIS iteration's fix
|
|
743
|
+
diff so `fix_provenance_new` can be computed (see below).
|
|
744
|
+
6. iteration += 1
|
|
745
|
+
|
|
746
|
+
if iteration > MAX_ITERATIONS AND actionable_issues > 0:
|
|
747
|
+
escalate to human with remaining issues
|
|
748
|
+
```
|
|
749
|
+
|
|
750
|
+
**Round ledger and termination rules (#1625).** The loop above has an
|
|
751
|
+
iteration cap but, on its own, no notion of finding *identity* across rounds
|
|
752
|
+
— so nothing detects "this round found only residue from the last fix" or
|
|
753
|
+
"we are churning". Classify every round's findings with the shared helper:
|
|
754
|
+
|
|
755
|
+
```bash
|
|
756
|
+
# RUN_ID identifies this changeset; the ledger is discarded if it belongs to
|
|
757
|
+
# a different one. Any stable digest of the target set works.
|
|
758
|
+
RUN_ID=$(printf '%s\n' <target files> | sort | sha256sum | cut -c1-16)
|
|
759
|
+
python3 "$CLAUDE_PLUGIN_ROOT/skills/code-review/scripts/finding_signature.py" \
|
|
760
|
+
--round <N> --findings <this-round's-findings.json> \
|
|
761
|
+
--run-id "$RUN_ID" \
|
|
762
|
+
--state .claude/memory/review-round-state.json
|
|
763
|
+
```
|
|
764
|
+
|
|
765
|
+
It prints `{"round", "new", "carried", "actionable_new", "ledger_reset",
|
|
766
|
+
"terminate", "reason"}` and maintains the durable ledger at
|
|
767
|
+
`.claude/memory/review-round-state.json`, so `/continue` can resume a review
|
|
768
|
+
mid-loop and step 6a's `--carried` count for #1624 comes from the same
|
|
769
|
+
source rather than being re-derived from memory.
|
|
770
|
+
|
|
771
|
+
**Ledger lifecycle — the ledger must never leak across runs.** Reusing a
|
|
772
|
+
ledger built for a different changeset would misclassify a genuinely new
|
|
773
|
+
finding as "carried" (silently skipping a round that should have run) and
|
|
774
|
+
inflate the round counter toward the cap on unrelated history. Four reset
|
|
775
|
+
triggers, in precedence order, all handled by the script:
|
|
776
|
+
|
|
777
|
+
| Trigger | When |
|
|
778
|
+
| --- | --- |
|
|
779
|
+
| `--reset` | Explicit, caller-forced |
|
|
780
|
+
| Round 1 | The initial panel **is** the start of a new run by definition. `/code-review` always calls round 1 first, so an abandoned ledger can never leak into the next review — no caller bookkeeping needed |
|
|
781
|
+
| `run-id-mismatch` | The stored ledger was built for different target files. Catches a resume that legitimately starts at round ≥ 2 against a different changeset |
|
|
782
|
+
| `stale-state` | The run started more than 24h ago — abandoned residue |
|
|
783
|
+
|
|
784
|
+
This ledger covers rounds *within* one `/code-review` invocation only — a
|
|
785
|
+
caller that re-invokes `/code-review` across separate command runs is the
|
|
786
|
+
place that scopes those *separate* invocations to the fix-diff instead of
|
|
787
|
+
restarting at round 1 every time. `/pr`'s own gate-retry loop
|
|
788
|
+
(`../pr/SKILL.md` step 2) is that caller today.
|
|
789
|
+
|
|
790
|
+
Reported as `ledger_reset` on every call (`null` when a stored ledger was
|
|
791
|
+
legitimately resumed). The script fails **toward** a reset: an unreadable or
|
|
792
|
+
malformed state file starts fresh. Starting fresh costs at most one extra
|
|
793
|
+
round; reusing a wrong ledger silently skips one. A finding's signature is
|
|
794
|
+
`(agent, file, category, normalized message)` with the line compared at
|
|
795
|
+
±3 rather than hashed — see the script's own docstring for why the line is
|
|
796
|
+
deliberately outside the hash.
|
|
797
|
+
|
|
798
|
+
Three termination rules, evaluated at each round boundary, **first match
|
|
799
|
+
wins** — the helper implements all three, this text is the contract:
|
|
800
|
+
|
|
801
|
+
| Rule | Trigger | Effect |
|
|
802
|
+
| --- | --- | --- |
|
|
803
|
+
| **Hard round cap** | `round >= 4` (initial panel + 3) | Escalate to human, attaching the round ledger as evidence — which rounds found what, with fix provenance. Same posture as the existing `MAX_ITERATIONS` escalation, and step 9 treats it the same way (no gate write). Applied to #1619's case study, this alone would have surfaced the churn at round 4 instead of round 9. |
|
|
804
|
+
| **Severity floor** (rounds ≥ 2) | No *new-signature* finding is `error`/`warning` at `high`/`medium` confidence | Converged — leave the loop. Suggestion-tier and low-confidence findings from round ≥ 2 still go to `corrections/` and the report: **logged, never chased.** Round 1 is unaffected — its actionability is step 5b's table, not this floor. |
|
|
805
|
+
| **Loop-until-dry** | A round produces zero new-signature findings clearing the floor | Converged. Carried signatures that survived a fix attempt are already covered by the existing "same issues persist → escalate" exit; they are not a reason to keep going here. |
|
|
806
|
+
|
|
807
|
+
The same three rules govern `/build`'s inline checkpoint fix loops
|
|
808
|
+
(`../build/SKILL.md` sub-steps 4/6) — one shared statement, one shared
|
|
809
|
+
implementation, not a duplicated table.
|
|
810
|
+
|
|
811
|
+
**Record each round (#1624).** The initial panel was round 1 (step 5b-i);
|
|
812
|
+
each fix-loop iteration's re-dispatch set is one further round. Capture the
|
|
813
|
+
iteration's fix diff **before** re-staging (item 3) so the row can attribute
|
|
814
|
+
this round's new findings to the previous round's fix:
|
|
815
|
+
|
|
816
|
+
```bash
|
|
817
|
+
# Item 1 applied fixes; capture them as a diff, then (item 3) `git add` them.
|
|
818
|
+
git -c diff.relative=false diff --no-color > "$FIX_DIFF"
|
|
819
|
+
# …after item 5's re-aggregation:
|
|
820
|
+
python3 "$CLAUDE_PLUGIN_ROOT/skills/code-review/scripts/review_round_log.py" \
|
|
821
|
+
--round <N> --agents "<agents re-dispatched this round>" \
|
|
822
|
+
--findings <this-round's-NEW-findings.json> \
|
|
823
|
+
--carried <count of findings carried over from the prior round> \
|
|
824
|
+
--purpose "<discovery|verification|closing>" \
|
|
825
|
+
--outcome "<fixed|no-op|escalated>" \
|
|
826
|
+
--fix-diff "$FIX_DIFF"
|
|
827
|
+
```
|
|
828
|
+
|
|
829
|
+
`fix_provenance_new` — how many of this round's new findings land inside the
|
|
830
|
+
line ranges the previous round's fix touched — is the judgment-free "the fix
|
|
831
|
+
introduced it" signal #1623 asks for. It is interval math over the diff, not
|
|
832
|
+
an LLM call: a round whose new error/warning findings **all** carry
|
|
833
|
+
provenance is churn by construction. `--purpose` distinguishes a discovery
|
|
834
|
+
panel from a fix-verification re-dispatch and from the gate-closing pass, so
|
|
835
|
+
per-agent cost can be split by purpose rather than lumped into one dispatch
|
|
836
|
+
count. Derived metrics (churn ratio, per-agent discovery-vs-verification
|
|
837
|
+
split, gate recidivism) are computed by `/harness-audit` — see its Step 4a.
|
|
838
|
+
|
|
839
|
+
**Closing pass — re-establishing dispatch-ledger corroboration after the loop (#1461, narrowed by #1626, floor lowered to 1 by #2147; auto-scope only — same condition as item 3 above).** Step 3's `git add` changes the staged content's hash, so `agent_dispatch_ledger.py` stamps each iteration's re-dispatched agents (step 4) with that NEW hash — not step 4 (the outer, pre-loop)'s original dispatch hash, and not an earlier iteration's hash either. Step 9's gate write needs **>= 1 distinct dispatch whose `subject_hash` equals the FINAL staged content's hash** (the one actually committed). Since floor 1 is normally satisfied by construction (the loop only re-stages when it applied at least one fix, and that fixer re-dispatches against the final content in the same iteration), the closing pass below degenerates to "just the fixer(s), self-verifying" in the common case — it still exists, unconditionally, for the cases that don't already clear the floor on their own: the escape hatch (scope grew mid-loop) and a `fixed_by_agents` set that ends up empty despite a loop iteration having run.
|
|
840
|
+
|
|
841
|
+
This used to be satisfied by re-dispatching the **full** original panel, which made a one-line fix cost an 18-agent round. **Unconditionally, after any loop iteration ran** (i.e. any fix was applied and re-staged) — not only when the count looks short, since that count isn't something to reason about from memory — run a **closing pass** instead. Compose it deterministically, don't pick the set by hand:
|
|
842
|
+
|
|
843
|
+
```bash
|
|
844
|
+
python3 "$CLAUDE_PLUGIN_ROOT/skills/code-review/scripts/closing_pass.py" \
|
|
845
|
+
--fixed-by "<agents whose findings were fixed during the loop>" \
|
|
846
|
+
--roster "@<select_lenses.py output>" --panel "<the round-1 panel>" \
|
|
847
|
+
--panel-files "<files the round-1 panel targeted>" \
|
|
848
|
+
--fix-delta-files "<files the cumulative fix touched>"
|
|
849
|
+
```
|
|
850
|
+
|
|
851
|
+
It prints `{"agents", "scope", "escape_hatch", "reason", "topped_up"}`. Dispatch exactly the `agents` it returns, and record them with `dispatch_purpose: "closing"` (#1624) so the cost effect is measurable.
|
|
852
|
+
|
|
853
|
+
- **Composition**: every agent whose findings were fixed during the loop (each verifies its own fixes at the final hash), plus — only if that set has fewer than 1 distinct agent — a cheap-first top-up from the resolver's eligible roster until 1 distinct registered agent has dispatched at the final hash.
|
|
854
|
+
- **Why this is sound**: 1 is `pre_pr_review.py`'s `_MIN_DISTINCT_DISPATCHES` (#2147; lowered from 2), so the gate's corroboration floor is satisfied **by construction** — no hook change, no exemption event, no ledger change. The threat model #1461 closed (self-certification without dispatch) is untouched: these are genuine dispatches carrying real review authority over the only content that changed since full-panel coverage. A drift test pins the script's constant to the hook's, so raising the gate's floor can never silently under-compose this pass.
|
|
855
|
+
- **Scope**: the closing pass reviews the **cumulative fix delta** — the diff between what the round-1 panel reviewed and the final staged content — with the round ledger's fixed findings as context. Not the whole changeset: the panel's round-1 coverage of unchanged content is still valid; only the fix delta is unreviewed.
|
|
856
|
+
- **Escape hatch**: when the fix delta touches files outside the original panel's target set (scope grew mid-loop), the script returns `escape_hatch: true` and `scope: "full-changeset"` — fall back to the full re-dispatch. This is a set comparison of two file lists, not a judgment call.
|
|
857
|
+
|
|
858
|
+
**This pass is a real review, not a rubber stamp**: closing-pass agents keep full authority. If any reports an actionable issue, treat it exactly like any other iteration — re-enter this loop (subject to `MAX_ITERATIONS` and #1625's round cap) rather than proceeding to step 7. What #1626 changed is only *how many* agents re-read *how much* content; never whether their findings count. If the iteration limit or the round cap is reached with issues still outstanding, follow the existing "escalate to human" exit condition below — step 9's gate-write condition explicitly excludes this case (treat it as if overall status were `fail` for that one purpose, even if every outstanding issue is only `warning`-severity), so an escalation is never silently overridden by a passing gate write. A corroboration pass whose findings carry no consequence would be exactly the "dispatch trivial calls purely to clear the gate" abuse `pre_commit_review.py`'s own module docstring names as the residual risk this mechanism does NOT protect against.
|
|
859
|
+
|
|
860
|
+
**Exit conditions**:
|
|
861
|
+
|
|
862
|
+
| Condition | Action |
|
|
863
|
+
| --- | --- |
|
|
864
|
+
| Zero actionable issues | Exit → step 7 |
|
|
865
|
+
| Round ledger returns `converged` (#1625) | Exit → step 7. A clean convergence, not an escalation: no new-signature finding cleared the severity floor, so step 9's gate write proceeds normally |
|
|
866
|
+
| Round ledger returns `round-cap` (#1625) | Exit → escalate with the round ledger attached. Treated exactly like the iteration-limit row below for step 9's gate-write condition |
|
|
867
|
+
| Iteration limit (5) | Exit → escalate (#1461: step 9 treats this as `fail` for its gate-write condition, even if remaining issues are only `warning`-severity) |
|
|
868
|
+
| Same issues persist | Exit → escalate — not converging (same #1461 step 9 treatment as the iteration-limit row: this is also an escalation with actionable issues outstanding, not a quiet exit) |
|
|
869
|
+
| Tests fail after fix and revert | Mark issue human-required; continue |
|
|
870
|
+
|
|
871
|
+
The round cap (4) binds before `MAX_ITERATIONS` (5) in practice: the cap
|
|
872
|
+
counts total dispatch rounds including the initial panel, the iteration
|
|
873
|
+
limit counts fix-loop passes only. Both remain — the cap is the churn
|
|
874
|
+
control, the iteration limit the original backstop.
|
|
875
|
+
|
|
876
|
+
**Record the escalation state for step 7, not only step 9 (issue #1880).**
|
|
877
|
+
Whichever exit condition above was hit — `round-cap`, iteration limit, or
|
|
878
|
+
"same issues persist" — is an **escalation**; `converged` (including the
|
|
879
|
+
zero-actionable-issues and round-ledger-`converged` rows) is not. Carry that
|
|
880
|
+
boolean (escalated vs. converged) forward out of this loop: step 9 already
|
|
881
|
+
consults it for the `.pr-review-passed` gate-write condition, and step 7 now
|
|
882
|
+
also consults it — under `--json`, where step 9 never runs at all — to force
|
|
883
|
+
`overall: "fail"` in the emitted JSON object per the parallel rule in
|
|
884
|
+
[`output-format.md`](output-format.md#aggregated-json-result---json-flag).
|
|
885
|
+
Without carrying this state to step 7, an escalated review with only
|
|
886
|
+
warning-severity issues remaining would emit `overall: "warn"` in `--json`
|
|
887
|
+
mode and a caller like `/pr`'s internal `--json` call would never see the
|
|
888
|
+
escalation.
|
|
889
|
+
|
|
890
|
+
Track each iteration for the report — template in [`output-format.md`](output-format.md#review-fix-loop-iteration-log-step-6a-iv).
|
|
891
|
+
|
|
892
|
+
### 7. Generate report
|
|
893
|
+
|
|
894
|
+
**Output paths.** All file artifacts (`./corrections/*.json`, `.claude/memory/.pr-review-passed`) are repo-relative to the target repository's working directory (the cwd `/code-review` was invoked in). Never prepend a scratchpad, sandbox, or session root onto an already-absolute path, and never join two absolute paths. `--json` prints to **stdout** and writes no file.
|
|
895
|
+
|
|
896
|
+
Read `knowledge/review-template.md` for the structure.
|
|
897
|
+
|
|
898
|
+
**If `--json`: the JSON object is the ONLY thing printed to stdout for this run — non-negotiable, not model discretion.** Emit the aggregated JSON object per the schema in [`output-format.md`](output-format.md#aggregated-json-result---json-flag) to **stdout**, write no report file, and **skip step 8 in this run, regardless of how many issues were found or whether any are actionable.** There is no fallback to prose, and no `corrections/` persistence, in `--json` mode — ever. (`/pr`'s `--json` call already only reads this JSON object's `overall`/`status` field, so this loses nothing a caller depends on.)
|
|
899
|
+
|
|
900
|
+
**Step 9 is NOT skipped by `--json` (issue #1904 Bug 2b) — emitting `--json` output and writing the PR-time gate file are orthogonal concerns.** `/pr`'s only path to `gh pr create` (`skills/pr/SKILL.md` step 2.4) invokes `/code-review --since "$BASE" --json` — so a review that is BOTH `--json` AND scoped via `--since <base>` is exactly the shape that must reach step 9, or `.claude/memory/.pr-review-passed` is never written on the only path that actually opens a PR, leaving `PR_GATE_BYPASS_REASON` as the only way to ever open one (the "gate that cannot fail is worse than no gate" anti-pattern this repo's own root `CLAUDE.md` names explicitly). After emitting the JSON object above, continue to step 9 unconditionally — its own scope condition already narrows correctly (a no-op for `--path`/`--all`/full-repository scope, same as before). Anything step 9 itself produces (boundary events, file writes) goes to disk or stderr, never stdout — stdout must stay pure JSON.
|
|
901
|
+
|
|
902
|
+
**When step 6a ran, consult its escalation state before computing `overall` here (issue #1880).** If step 6a exited via escalation (round-cap, iteration limit, or "same issues persist" — see that step's "Record the escalation state for step 7" note), force `overall: "fail"` in this JSON object, exactly like the `dispatchFailures` override — apply it after the totals-based computation so it always wins. This is the same rule stated in [`output-format.md`](output-format.md#aggregated-json-result---json-flag); it is restated here because step 9, where this escalation previously only mattered for the `.pr-review-passed` gate, never runs under `--json`. A clean `converged` exit (or a run that never entered the fix loop at all — zero actionable issues) does not trigger this override.
|
|
903
|
+
|
|
904
|
+
**A sentence describing the JSON is not the JSON.** A completed run whose final text reads like "Aggregated JSON emitted to stdout per `--json` contract; run stops here" — with no `{...}` object actually present anywhere in that text — is a contract violation, not compliance, even though it correctly stopped rather than proceeding further. The literal final output of the turn must be the JSON object itself, not a narration of having produced it. If the next action being considered is a summary sentence announcing that the JSON was (or is about to be) emitted, that is the signal to emit the actual object instead — there is no valid end state for a `--json` run that consists of prose alone.
|
|
905
|
+
|
|
906
|
+
Otherwise (no `--json`): emit the prose summary using the Code Review Summary template in [`output-format.md`](output-format.md#code-review-summary-report-step-7-prose-mode). For that template's per-finding listing, render this round's aggregated finding list (the same list already assembled for the `--json` branch above and for step 8 — not re-derived) with `render_tiered_findings.py` (#2170) instead of listing each finding's full message inline:
|
|
907
|
+
|
|
908
|
+
```bash
|
|
909
|
+
python3 "$CLAUDE_PLUGIN_ROOT/skills/code-review/scripts/render_tiered_findings.py" --findings <path-to-this-round's-finding-list.json> [--expand <finding-id>|all]
|
|
910
|
+
```
|
|
911
|
+
|
|
912
|
+
Pass `--expand` through exactly as the caller supplied it (omit the flag entirely when the caller did not pass one): Tier-1 lines plus the expansion hint by default; the matching Tier-2 block(s) appended after the Tier-1 report when `--expand` was given. An unknown `--expand` id: relay the script's non-zero exit and "finding-id not found" message to the user rather than silently rendering nothing or crashing. Append the iteration table.
|
|
913
|
+
|
|
914
|
+
**Scope of this wiring: the prose-mode path only.** `--json` (this step's branch above) and `./corrections/*.json` (step 8) already read and write the full finding objects independently of this rendering path — neither branch calls `render_tiered_findings.py`, and this change does not touch either of them. In particular, **`--expand` is a no-op under `--json`**: the `--json` branch above is unconditional ("the JSON object is the ONLY thing printed to stdout... non-negotiable") and must never call `render_tiered_findings.py`, so under `--json` there is nothing for `--expand` to act on. This is enforced structurally — by the `--json` branch never reaching the tiered-rendering code path described here — not by a check inside `render_tiered_findings.py` or inside the `--json` branch itself.
|
|
915
|
+
|
|
916
|
+
**Write the durable report (skip when `--internal`).** See
|
|
917
|
+
`knowledge/report-output-location.md` for the shared write-scope convention
|
|
918
|
+
this step follows. When `--internal`
|
|
919
|
+
was **not** passed, write the identical prose summary to
|
|
920
|
+
`.dev-team-reports/code-review.md` in the target repository's working
|
|
921
|
+
directory (creating the directory if absent), overwriting any existing
|
|
922
|
+
file at that path — write it even when the review found zero issues. Print
|
|
923
|
+
one confirmation line: `Report written: .dev-team-reports/code-review.md`,
|
|
924
|
+
or `Report written: .dev-team-reports/code-review.md (replaced previous
|
|
925
|
+
run)` when a file already existed at that path. If the write fails
|
|
926
|
+
(permission/read-only): report `Cannot write
|
|
927
|
+
.dev-team-reports/code-review.md: <error>` to chat and continue unaffected —
|
|
928
|
+
the write failure is non-fatal. When `--internal` **was** passed, skip this
|
|
929
|
+
write entirely (the fix loop and every other prose-mode behavior above are
|
|
930
|
+
unaffected — `--internal` only suppresses this one write). Then continue to
|
|
931
|
+
step 8.
|
|
932
|
+
|
|
933
|
+
**`--pdf` (additive, after the write).** When `--pdf` was passed and a report
|
|
934
|
+
file **was** written this run, render it to a sibling PDF per
|
|
935
|
+
`knowledge/report-pdf-integration.md`:
|
|
936
|
+
|
|
937
|
+
```bash
|
|
938
|
+
sh "$CLAUDE_PLUGIN_ROOT/hooks/py.sh" "$CLAUDE_PLUGIN_ROOT/hooks/lib/report_pdf.py" .dev-team-reports/code-review.md
|
|
939
|
+
```
|
|
940
|
+
|
|
941
|
+
Surface the module's `Rendering PDF via <engine>…` and result lines. When no
|
|
942
|
+
report file was written this run (`--json` or `--internal`), `--pdf` is a
|
|
943
|
+
no-op: state `--pdf: no report file was written this run, nothing to render.`
|
|
944
|
+
and do nothing else. Under `--json`, emit that no-op line (and any render
|
|
945
|
+
status) to **stderr** so stdout stays valid JSON. `--pdf` never alters the
|
|
946
|
+
review's own output or exit status — a missing engine or render error is
|
|
947
|
+
non-fatal.
|
|
948
|
+
|
|
949
|
+
### 8. Save correction prompts for remaining issues
|
|
950
|
+
|
|
951
|
+
**Skip this entire step if `--json` was set.** Step 7 already skips this step for `--json` mode; corrections are never written to disk in `--json` mode. (Step 9, unlike this step, is NOT skipped by `--json` — see that step's own condition.)
|
|
952
|
+
|
|
953
|
+
For issues NOT auto-fixed (confidence: none, auto-fix failed, or suggestions), generate one correction prompt per issue using the Correction prompt schema in [`output-format.md`](output-format.md#correction-prompt-json). Save to `./corrections/` **in the target repository's working directory** (the cwd `/code-review` was invoked in). Write all output artifacts only to these repo-relative paths — never prepend a scratchpad, sandbox, or session root, and never join two absolute paths. These can be addressed manually or via `/apply-fixes`.
|
|
954
|
+
|
|
955
|
+
### 9. Write pre-commit gate file
|
|
956
|
+
|
|
957
|
+
**This step is NOT skipped by `--json` (issue #1904 Bug 2b) — see step 7's own note.** It applies whenever the scope condition below holds, `--json` or not; step 7 only skips step **8** for `--json`.
|
|
958
|
+
|
|
959
|
+
**Dispatch failures block the gate (issue #1752).** If step 5's `dispatchFailures` list is non-empty — any agent that failed dispatch and then failed its single retry — do not write `.pr-review-passed`: `.pr-review-passed` must not be written while any dispatch failure is outstanding, regardless of the overall status computed from the agents that did return. The same rationale as step 3's `unreadable-registry` treatment: a lens that never ran is a coverage gap, not a passing result, so this condition is checked **before** the status check below, not folded into it. This prose rule now has a mechanical backstop (issue #1763, carried forward to the PR-time gate by #1886): `hooks/pre_pr_review.py`'s own `_dispatch_failure_verdict` independently vetoes the gate at `gh pr create` time when a `dispatch-failure` boundary event (emitted at Step 4, above) is on record for the current branch-diff content — so a bug or a future caller that skips this step's condition, or writes `.pr-review-passed` directly, still can't silently bypass it on this path. Two disclosed limits, neither exploitable today but both worth naming rather than silently assuming away (#1763 security review):
|
|
960
|
+
|
|
961
|
+
- The backstop is inert on [`sliced-mode.md`](sliced-mode.md)'s path: sliced mode auto-engages only on a full-repo scope with nothing staged, and never writes `.pr-review-passed` at all — it replaces this step entirely, report-only — so the backstop's inertness there holds unconditionally, regardless of what `branch_diff_gate_hash()` evaluates to for that scope.
|
|
962
|
+
- The backstop queries only the CURRENT branch-diff hash. A dispatch failure recorded against an earlier hash — e.g. one orphaned by a later commit landing on the same branch, or by step 6a's fix loop, before that step's own condition (above) is (mis)evaluated — is not queried here. `hooks/pre_pr_review.py` has no analogue of the retired cosmetic-delta carry-forward mechanism (deliberately dropped, per that module's own docstring) to union multiple hash bindings, so a branch-diff change since a real dispatch failure silently drops that failure's veto power until a fresh review re-emits it against the new hash.
|
|
963
|
+
|
|
964
|
+
**Scope condition, extended by issue #1904 Bug 2b.** If the review was auto-scoped to uncommitted changes **OR scoped via `--since <base>`** — and the overall status is `pass` or `warn` **and step 6a did not exit with actionable issues outstanding** — whether via the iteration limit or the "not converging" exit, both of which are escalations, per that step's Exit conditions table (regardless of whether those outstanding issues are only `warning`-severity — either escalation overrides `warn` for this condition specifically, since escalating and then writing a passing gate anyway would silently defeat the escalation) — write `.pr-review-passed` to `.claude/memory/` so `hooks/pre_pr_review.py` allows the next `gh pr create` (#1886). Use the **shared gate-hash helper**, in its `--branch-diff` mode, so the writer and the pre-PR hook compute the hash identically — it hashes the branch's diff against its base (`git diff <base>...HEAD`), not the staged patch, so a commit landed on the branch after this write invalidates the gate:
|
|
965
|
+
|
|
966
|
+
**Why `--since <base>` belongs here (closing the gap Bug 2b's premise names):** `/pr`'s only path to `gh pr create` (`skills/pr/SKILL.md` step 2.4) invokes `/code-review --since "$BASE" --json` — before this fix, this condition fired ONLY for auto-scoped uncommitted changes, so `.claude/memory/.pr-review-passed` was NEVER written on the one real path that opens a PR, making `PR_GATE_BYPASS_REASON` the only way to ever open one (the "gate that cannot fail is worse than no gate" anti-pattern named in this repo's own root `CLAUDE.md`). This also resolves the hash-timing-mismatch the auto-scope path still has (see the "Known limitation" note below): `/pr`'s step 1 requires a CLEAN working tree before invoking `--since`, so the hash computed here — AFTER those commits already landed — is computed against exactly the same content `hooks/pre_pr_review.py` recomputes at `gh pr create` time; there is no "staged while uncommitted" content this write could omit.
|
|
967
|
+
|
|
968
|
+
```bash
|
|
969
|
+
HASH=$(python3 "${CLAUDE_PLUGIN_ROOT}/hooks/lib/review_gate_hash.py" --branch-diff)
|
|
970
|
+
mkdir -p .claude/memory && printf '%s\n' "$HASH" > .claude/memory/.pr-review-passed
|
|
971
|
+
```
|
|
972
|
+
|
|
973
|
+
Single line, unlike the retired `.review-passed`'s two-line format —
|
|
974
|
+
`hooks/pre_pr_review.py`'s own `_stored_gate_hash()` reads only the first
|
|
975
|
+
line. There is no normalization-invariant second line here: the
|
|
976
|
+
cosmetic-delta carry-forward mechanism that line existed for (#1627) was
|
|
977
|
+
deliberately dropped for this gate (per `pre_pr_review.py`'s own docstring),
|
|
978
|
+
since it fires once, at PR-creation time, rather than at every commit — the
|
|
979
|
+
friction that mechanism relieved does not arise here.
|
|
980
|
+
|
|
981
|
+
**Known limitation, narrowed by issue #1904 Bug 2b (was #1886 follow-up).**
|
|
982
|
+
`agent_dispatch_ledger.py` stamps a dispatch's `subject_hash` with
|
|
983
|
+
`review_gate_hash()` (the staged `--cached` diff) whenever something IS
|
|
984
|
+
staged, but falls back to `branch_diff_gate_hash(default_base_ref(cwd),
|
|
985
|
+
cwd)` — the SAME content domain this step writes and
|
|
986
|
+
`hooks/pre_pr_review.py` checks — whenever nothing is staged (see that
|
|
987
|
+
hook's own module docstring for the fallback's rationale). For a `--since
|
|
988
|
+
<base>`-scoped review, nothing is EVER staged (`/pr`'s step 1 requires a
|
|
989
|
+
clean working tree first), so every dispatch during that review stamps the
|
|
990
|
+
branch-diff hash directly — the write above and every corroborating
|
|
991
|
+
dispatch now agree on one content domain for this mode, closing the gap for
|
|
992
|
+
the shape #1886 identified it in.
|
|
993
|
+
|
|
994
|
+
The residual gap that remains is narrower: on an **auto-scoped
|
|
995
|
+
uncommitted-changes** review with multiple separate review-and-commit
|
|
996
|
+
cycles on the same branch, a dispatch's staged-diff `subject_hash` and this
|
|
997
|
+
step's branch-diff hash are mathematically identical only in the common
|
|
998
|
+
single-commit-then-PR shape (a branch cut from its base, reviewed once
|
|
999
|
+
while staged, committed, then a PR opened immediately) — exactly as before.
|
|
1000
|
+
On a branch with multiple such cycles, the branch-diff hash written here
|
|
1001
|
+
will not match an EARLIER cycle's dispatch `subject_hash`, and
|
|
1002
|
+
`hooks/pre_pr_review.py` correctly fails closed at `gh pr create` time,
|
|
1003
|
+
requiring a fresh `/code-review` run against the branch's current diff (or
|
|
1004
|
+
a `--since <base>`-scoped re-review, which now closes cleanly per the
|
|
1005
|
+
paragraph above) before opening the PR.
|
|
1006
|
+
|
|
1007
|
+
**If `--agent <name>` was used** (a sanctioned single-agent review — it deliberately dispatches exactly 1 agent, which now clears the dispatch-ledger gate's `>= 1` distinct-dispatch floor on its own since #2147 lowered it from 2; the explicit exemption event below is consequently no longer load-bearing for this case, but is still written for an unambiguous, explicit audit trail rather than relying on the ordinary count path to imply "this was a sanctioned single-agent review" after the fact), record that as an explicit, auditable exemption event bound to this same hash **contemporaneously** with the write above — same pattern as the doc-only short-circuit's exemption event (step 1a):
|
|
1008
|
+
|
|
1009
|
+
```bash
|
|
1010
|
+
python3 "${CLAUDE_PLUGIN_ROOT}/hooks/lib/boundary_events.py" --event single-agent --subject-hash "$HASH"
|
|
1011
|
+
```
|
|
1012
|
+
|
|
1013
|
+
This step only runs when the review was auto-scoped to uncommitted changes or scoped via `--since <base>` (see the gate condition above). Do not `git add` a different file set, and do not recompute `$HASH` against different content, at this point: staging or hashing something other than what this run actually reviewed would write a gate hash unrelated to the review that produced it.
|
|
1014
|
+
|
|
1015
|
+
If overall status is `fail`, do **not** write the gate file — `hooks/pre_pr_review.py` will keep blocking `gh pr create` until issues are resolved and the review re-run.
|