pi-dev-team 0.2.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/LICENSE +21 -0
- package/PORTING.md +134 -0
- package/README.md +207 -0
- package/UPSTREAM.json +64 -0
- package/agents/Explore.md +15 -0
- package/agents/a11y-review.md +118 -0
- package/agents/adr-author.md +70 -0
- package/agents/ai-provenance-review.md +120 -0
- package/agents/angular-reactivity-review.md +95 -0
- package/agents/arch-review.md +135 -0
- package/agents/architect.md +78 -0
- package/agents/autoship-batch-proposer.md +69 -0
- package/agents/claude-setup-review.md +136 -0
- package/agents/codebase-recon.md +184 -0
- package/agents/component-architecture-review.md +119 -0
- package/agents/concurrency-review.md +109 -0
- package/agents/correctness-review.md +290 -0
- package/agents/data-flow-tracer.md +120 -0
- package/agents/doc-review.md +165 -0
- package/agents/domain-review.md +136 -0
- package/agents/general-purpose.md +10 -0
- package/agents/gherkin-quality-critic.md +113 -0
- package/agents/js-fp-review.md +114 -0
- package/agents/mutation-kill.md +684 -0
- package/agents/naming-review.md +142 -0
- package/agents/orchestrator.md +339 -0
- package/agents/performance-review.md +105 -0
- package/agents/plan-review-acceptance.md +115 -0
- package/agents/plan-review-design.md +90 -0
- package/agents/plan-review-parallelization.md +84 -0
- package/agents/plan-review-strategic.md +96 -0
- package/agents/plan-review-ux.md +110 -0
- package/agents/platform-engineer.md +64 -0
- package/agents/product-manager.md +68 -0
- package/agents/progress-guardian.md +79 -0
- package/agents/qa-engineer.md +289 -0
- package/agents/quality-reviewer.md +132 -0
- package/agents/react-reactivity-review.md +102 -0
- package/agents/refactor-opportunity-review.md +128 -0
- package/agents/security-engineer.md +60 -0
- package/agents/security-review.md +218 -0
- package/agents/session-analysis.md +95 -0
- package/agents/software-engineer.md +105 -0
- package/agents/spec-compliance-review.md +100 -0
- package/agents/spec-reviewer.md +114 -0
- package/agents/structure-review.md +146 -0
- package/agents/tech-writer.md +84 -0
- package/agents/test-review.md +246 -0
- package/agents/test-smell-review.md +188 -0
- package/agents/token-efficiency-review.md +139 -0
- package/agents/ui-ux-designer.md +54 -0
- package/agents/vue-reactivity-review.md +95 -0
- package/bin/__pycache__/claudecpython-314.pyc +0 -0
- package/bin/claude +258 -0
- package/docs/upstream/.pages +1 -0
- package/docs/upstream/CHANGELOG.md +2586 -0
- package/docs/upstream/README.md +155 -0
- package/docs/upstream/agent-architecture.md +214 -0
- package/docs/upstream/agent_info.md +187 -0
- package/docs/upstream/artifact-migration.md +124 -0
- package/docs/upstream/code-intelligence-nudge.md +149 -0
- package/docs/upstream/code-review-process.md +294 -0
- package/docs/upstream/concurrent-use.md +73 -0
- package/docs/upstream/context-management.md +111 -0
- package/docs/upstream/developer-notes.md +280 -0
- package/docs/upstream/diagrams/architecture-overview.svg +101 -0
- package/docs/upstream/diagrams/review-dispatch.svg +139 -0
- package/docs/upstream/diagrams/team-agents.svg +128 -0
- package/docs/upstream/diagrams/test-improve-flow.svg +166 -0
- package/docs/upstream/diagrams/workflow-linear.svg +66 -0
- package/docs/upstream/diagrams/workflow-three-phase.svg +200 -0
- package/docs/upstream/eval-maintenance.md +95 -0
- package/docs/upstream/eval-running-guide.md +147 -0
- package/docs/upstream/eval-system.md +291 -0
- package/docs/upstream/session-review-oss-complements.md +75 -0
- package/docs/upstream/session-review.md +212 -0
- package/docs/upstream/skills.md +188 -0
- package/docs/upstream/team-structure.md +21 -0
- package/docs/upstream/telemetry-ci-access.md +129 -0
- package/docs/upstream/telemetry-repo-security.md +120 -0
- package/docs/upstream/test-evaluation.md +277 -0
- package/docs/upstream/test-improve.md +154 -0
- package/docs/upstream/triage-workflow.md +282 -0
- package/docs/upstream/workflows.md +289 -0
- package/extensions/dev-team/index.ts +539 -0
- package/extensions/dev-team/lib/agents.ts +272 -0
- package/extensions/dev-team/lib/ai-credits.ts +92 -0
- package/extensions/dev-team/lib/autocompact.ts +81 -0
- package/extensions/dev-team/lib/child-run.ts +102 -0
- package/extensions/dev-team/lib/config.ts +236 -0
- package/extensions/dev-team/lib/gh-command.ts +103 -0
- package/extensions/dev-team/lib/github-style.ts +307 -0
- package/extensions/dev-team/lib/hooks.ts +350 -0
- package/extensions/dev-team/lib/metrics.ts +115 -0
- package/extensions/dev-team/lib/safe-read.ts +49 -0
- package/extensions/dev-team/lib/session-files.ts +57 -0
- package/extensions/dev-team/lib/session-spend.ts +123 -0
- package/extensions/dev-team/lib/shell-scan.ts +205 -0
- package/extensions/dev-team/lib/skills.ts +213 -0
- package/extensions/dev-team/lib/subagent-render.ts +245 -0
- package/extensions/dev-team/lib/subagent-types.ts +164 -0
- package/extensions/dev-team/lib/subagent.ts +596 -0
- package/extensions/dev-team/lib/terminal-text.ts +54 -0
- package/extensions/dev-team/lib/tools-misc.ts +152 -0
- package/extensions/dev-team/lib/transcript.ts +110 -0
- package/extensions/dev-team/lib/trust.ts +52 -0
- package/extensions/dev-team/lib/usage-breakdown.ts +176 -0
- package/extensions/dev-team/lib/usage-chart.ts +153 -0
- package/extensions/dev-team/lib/usage-command.ts +107 -0
- package/extensions/dev-team/lib/usage-history.ts +203 -0
- package/extensions/dev-team/lib/usage-render.ts +225 -0
- package/extensions/dev-team/lib/usage-split-bar.ts +127 -0
- package/extensions/dev-team/lib/usage-state.ts +116 -0
- package/extensions/dev-team/lib/usage-text.ts +159 -0
- package/extensions/dev-team/lib/usage-view.ts +109 -0
- package/hooks/__pycache__/refactor_test_freeze_guard.cpython-314.pyc +0 -0
- package/hooks/agent_dispatch_ledger.py +190 -0
- package/hooks/autocompact_setup_nudge.py +99 -0
- package/hooks/bash_retry_guard.py +228 -0
- package/hooks/boundary_events_write_guard.py +352 -0
- package/hooks/code_intelligence_nudge.py +293 -0
- package/hooks/code_intelligence_turn_mark.py +317 -0
- package/hooks/codegraph_bootstrap.py +139 -0
- package/hooks/contract_version_guard.py +362 -0
- package/hooks/cost_meter.py +106 -0
- package/hooks/destructive-commands.json +62 -0
- package/hooks/destructive_guard.py +477 -0
- package/hooks/eval_compliance_check.py +440 -0
- package/hooks/guards.json +17 -0
- package/hooks/hooks.json +323 -0
- package/hooks/internal_double_gate.py +296 -0
- package/hooks/js_fp_review.py +212 -0
- package/hooks/knowledge_index.py +119 -0
- package/hooks/lib/__pycache__/artifact_paths.cpython-314.pyc +0 -0
- package/hooks/lib/__pycache__/atomic_state.cpython-314.pyc +0 -0
- package/hooks/lib/__pycache__/autocompact_config.cpython-314.pyc +0 -0
- package/hooks/lib/__pycache__/boundary_events.cpython-314.pyc +0 -0
- package/hooks/lib/__pycache__/doc_classification.cpython-314.pyc +0 -0
- package/hooks/lib/__pycache__/gh_pr_create_detect.cpython-314.pyc +0 -0
- package/hooks/lib/__pycache__/git_safe_diff.cpython-314.pyc +0 -0
- package/hooks/lib/__pycache__/instrument_log.cpython-314.pyc +0 -0
- package/hooks/lib/__pycache__/metrics_query.cpython-314.pyc +0 -0
- package/hooks/lib/__pycache__/plugin_version.cpython-314.pyc +0 -0
- package/hooks/lib/__pycache__/pre_commit_doc_classifier.cpython-314.pyc +0 -0
- package/hooks/lib/__pycache__/review_agent_registry.cpython-314.pyc +0 -0
- package/hooks/lib/__pycache__/review_gate_corroboration.cpython-314.pyc +0 -0
- package/hooks/lib/__pycache__/review_gate_hash.cpython-314.pyc +0 -0
- package/hooks/lib/__pycache__/review_verdicts.cpython-314.pyc +0 -0
- package/hooks/lib/__pycache__/stdin_json.cpython-314.pyc +0 -0
- package/hooks/lib/__pycache__/stryker_invocation.cpython-314.pyc +0 -0
- package/hooks/lib/__pycache__/telemetry_consent.cpython-314.pyc +0 -0
- package/hooks/lib/__pycache__/test_file_classify.cpython-314.pyc +0 -0
- package/hooks/lib/__pycache__/token_efficiency_limits.cpython-314.pyc +0 -0
- package/hooks/lib/__pycache__/verify_guard_state.cpython-314.pyc +0 -0
- package/hooks/lib/__pycache__/xunit_v3_operator_gate.cpython-314.pyc +0 -0
- package/hooks/lib/agent_skill_hints.py +74 -0
- package/hooks/lib/artifact_paths.py +263 -0
- package/hooks/lib/atomic_state.py +557 -0
- package/hooks/lib/autocompact_config.py +103 -0
- package/hooks/lib/autoship_log.py +106 -0
- package/hooks/lib/banned_scripts_policy.py +51 -0
- package/hooks/lib/boundary_events.py +436 -0
- package/hooks/lib/build_knowledge_index.py +504 -0
- package/hooks/lib/build_skills_index.py +361 -0
- package/hooks/lib/build_state.py +116 -0
- package/hooks/lib/classify_ship_outcome.py +126 -0
- package/hooks/lib/config_changelog_schema.py +115 -0
- package/hooks/lib/cost_meter.py +955 -0
- package/hooks/lib/doc_classification.py +116 -0
- package/hooks/lib/gh_pr_create_detect.py +136 -0
- package/hooks/lib/git_safe_diff.py +123 -0
- package/hooks/lib/instrument_log.py +66 -0
- package/hooks/lib/iteration_journal_gate.py +197 -0
- package/hooks/lib/knowledge_index_paths.py +88 -0
- package/hooks/lib/mcp_json_repowise.py +177 -0
- package/hooks/lib/metrics_query.py +202 -0
- package/hooks/lib/minimal_yaml.py +434 -0
- package/hooks/lib/plugin_version.py +142 -0
- package/hooks/lib/pre_commit_detect.py +537 -0
- package/hooks/lib/pre_commit_doc_classifier.py +126 -0
- package/hooks/lib/pricing.py +118 -0
- package/hooks/lib/report_pdf.py +371 -0
- package/hooks/lib/review_agent_registry.py +142 -0
- package/hooks/lib/review_dispatch_ledger.py +101 -0
- package/hooks/lib/review_gate_corroboration.py +521 -0
- package/hooks/lib/review_gate_hash.py +252 -0
- package/hooks/lib/review_gate_normalized_hash.py +1115 -0
- package/hooks/lib/review_verdicts.py +301 -0
- package/hooks/lib/run_report.py +160 -0
- package/hooks/lib/skill_categories.yaml +125 -0
- package/hooks/lib/stdin_json.py +57 -0
- package/hooks/lib/stryker_invocation.py +102 -0
- package/hooks/lib/telemetry_consent.py +41 -0
- package/hooks/lib/telemetry_report.py +108 -0
- package/hooks/lib/test_file_classify.py +160 -0
- package/hooks/lib/token_efficiency_limits.py +51 -0
- package/hooks/lib/turn_identity.py +77 -0
- package/hooks/lib/verify_guard_state.py +110 -0
- package/hooks/lib/workflow_state.py +206 -0
- package/hooks/lib/xunit_v3_operator_gate.py +596 -0
- package/hooks/mcp_json_repowise_nudge.py +74 -0
- package/hooks/mutation_adapters/__init__.py +7 -0
- package/hooks/mutation_adapters/__pycache__/__init__.cpython-314.pyc +0 -0
- package/hooks/mutation_adapters/__pycache__/lib.cpython-314.pyc +0 -0
- package/hooks/mutation_adapters/__pycache__/mutmut.cpython-314.pyc +0 -0
- package/hooks/mutation_adapters/__pycache__/pitest.cpython-314.pyc +0 -0
- package/hooks/mutation_adapters/__pycache__/stryker.cpython-314.pyc +0 -0
- package/hooks/mutation_adapters/__pycache__/stryker_net.cpython-314.pyc +0 -0
- package/hooks/mutation_adapters/lib.py +478 -0
- package/hooks/mutation_adapters/mutmut.py +188 -0
- package/hooks/mutation_adapters/pitest.py +266 -0
- package/hooks/mutation_adapters/stryker.py +157 -0
- package/hooks/mutation_adapters/stryker_net.py +264 -0
- package/hooks/mutation_gate.py +193 -0
- package/hooks/mutation_testing_smoke_gate.py +371 -0
- package/hooks/pending_review_notify.py +121 -0
- package/hooks/phase_marker.py +138 -0
- package/hooks/post_compact_state_reinject.py +180 -0
- package/hooks/post_format.py +115 -0
- package/hooks/pre_commit_knowledge_index.py +128 -0
- package/hooks/pre_commit_review.py +66 -0
- package/hooks/pre_pr_review.py +694 -0
- package/hooks/pre_tool_guard.py +405 -0
- package/hooks/py.sh +73 -0
- package/hooks/refactor-bash-write-patterns.json +29 -0
- package/hooks/refactor_test_bash_guard.py +253 -0
- package/hooks/refactor_test_freeze_guard.py +139 -0
- package/hooks/refactor_test_revert_guard.py +186 -0
- package/hooks/repo_review_nudge.py +287 -0
- package/hooks/review_verdict_recorder.py +464 -0
- package/hooks/scan_bash_command_for_banned_scripts.py +428 -0
- package/hooks/scan_worktree_for_banned_scripts.py +238 -0
- package/hooks/session_learning_trigger.py +248 -0
- package/hooks/skills_index.py +126 -0
- package/hooks/stryker_xunit_shim_guard.py +571 -0
- package/hooks/subagent_completion_guard.py +309 -0
- package/hooks/subagent_skill_context.py +139 -0
- package/hooks/task_completion_metrics.py +216 -0
- package/hooks/tdd_guard.py +229 -0
- package/hooks/telemetry.py +341 -0
- package/hooks/token_efficiency_review.py +194 -0
- package/hooks/verify_guard.py +183 -0
- package/hooks/verify_guard_edit_marker.py +73 -0
- package/hooks/version_check.py +173 -0
- package/knowledge/accepted-risks-schema.md +98 -0
- package/knowledge/adr-decision-criteria.md +64 -0
- package/knowledge/adversarial-review-protocol.md +139 -0
- package/knowledge/agent-registry.md +228 -0
- package/knowledge/agent-review-methodology.md +80 -0
- package/knowledge/ai-friendly-repo-guidelines.md +67 -0
- package/knowledge/architecture-assessment.md +96 -0
- package/knowledge/artifact-lifecycle.md +57 -0
- package/knowledge/cd-maturity-model.md +82 -0
- package/knowledge/cd-test-architecture.md +190 -0
- package/knowledge/ci-cd-file-scope.md +24 -0
- package/knowledge/codegraph-vs-graphify.md +192 -0
- package/knowledge/component-test-patterns.md +139 -0
- package/knowledge/database-change-management.md +80 -0
- package/knowledge/database-test-patterns.md +79 -0
- package/knowledge/decision-defaults.md +88 -0
- package/knowledge/dependency-breaking-techniques.md +116 -0
- package/knowledge/deployment-pipeline.md +86 -0
- package/knowledge/design-smells.md +122 -0
- package/knowledge/directory-enumeration.md +38 -0
- package/knowledge/domain-modeling.md +123 -0
- package/knowledge/evidence-bundle.md +90 -0
- package/knowledge/exploratory-testing-field-guide.md +122 -0
- package/knowledge/failure-routing.md +28 -0
- package/knowledge/fixture-construction.md +56 -0
- package/knowledge/frontend-component-architecture.md +139 -0
- package/knowledge/gherkin-quality-review-dispatch.md +135 -0
- package/knowledge/index.json +6766 -0
- package/knowledge/internal-collaborator-doubling.md +101 -0
- package/knowledge/legacy-test-strategy.md +71 -0
- package/knowledge/long-run-waiting.md +66 -0
- package/knowledge/microservice-testing.md +71 -0
- package/knowledge/model-pricing.json +23 -0
- package/knowledge/mutation-score-formulas.md +60 -0
- package/knowledge/object-calisthenics.md +147 -0
- package/knowledge/oracle-provenance.md +94 -0
- package/knowledge/orchestrator-script-implementation.md +185 -0
- package/knowledge/owasp-detection.md +148 -0
- package/knowledge/plan-review-rubric.md +56 -0
- package/knowledge/proxy-connectivity.md +62 -0
- package/knowledge/reactive-effect-patterns.md +73 -0
- package/knowledge/recon-inventory-excludes.txt +32 -0
- package/knowledge/references/bdd-value-guide.md +61 -0
- package/knowledge/references/csharp-http-client-testing.md +264 -0
- package/knowledge/release-strategies.md +74 -0
- package/knowledge/report-output-location.md +117 -0
- package/knowledge/report-pdf-integration.md +63 -0
- package/knowledge/report-print.css +129 -0
- package/knowledge/report-template.md +114 -0
- package/knowledge/report-to-pdf.md +69 -0
- package/knowledge/request-processing-flow.md +63 -0
- package/knowledge/result-verification.md +52 -0
- package/knowledge/review-agent-output-contract.md +121 -0
- package/knowledge/review-lens-classification.md +113 -0
- package/knowledge/review-rubric.md +62 -0
- package/knowledge/review-template.md +104 -0
- package/knowledge/rule-fixtures/A02.insecure-random-js/negative.js +1 -0
- package/knowledge/rule-fixtures/A02.insecure-random-js/positive.js +1 -0
- package/knowledge/rule-fixtures/A02.weak-hashing-md5/negative.py +1 -0
- package/knowledge/rule-fixtures/A02.weak-hashing-md5/positive.py +1 -0
- package/knowledge/rule-fixtures/A03.command-injection/negative.js +1 -0
- package/knowledge/rule-fixtures/A03.command-injection/positive.js +1 -0
- package/knowledge/rule-fixtures/A03.sql-injection/negative.js +1 -0
- package/knowledge/rule-fixtures/A03.sql-injection/positive.js +1 -0
- package/knowledge/rule-fixtures/A03.xss-innerhtml/negative.js +1 -0
- package/knowledge/rule-fixtures/A03.xss-innerhtml/positive.js +1 -0
- package/knowledge/rule-fixtures/A05.cors-wildcard/negative.js +1 -0
- package/knowledge/rule-fixtures/A05.cors-wildcard/positive.js +1 -0
- package/knowledge/rule-fixtures/A05.default-credentials/negative.js +1 -0
- package/knowledge/rule-fixtures/A05.default-credentials/positive.js +1 -0
- package/knowledge/rule-fixtures/A07.jwt-alg-none/negative.js +1 -0
- package/knowledge/rule-fixtures/A07.jwt-alg-none/positive.js +1 -0
- package/knowledge/rule-fixtures/A08.binary-formatter/negative.cs +1 -0
- package/knowledge/rule-fixtures/A08.binary-formatter/positive.cs +1 -0
- package/knowledge/rule-fixtures/A08.js-eval/negative.js +1 -0
- package/knowledge/rule-fixtures/A08.js-eval/positive.js +1 -0
- package/knowledge/rule-fixtures/A08.object-input-stream/negative.java +1 -0
- package/knowledge/rule-fixtures/A08.object-input-stream/positive.java +1 -0
- package/knowledge/schemas/disposition-register-v1.json +65 -0
- package/knowledge/schemas/recon-envelope-v1.json +198 -0
- package/knowledge/schemas/unified-finding-v1.json +72 -0
- package/knowledge/security-primitives-contract.md +301 -0
- package/knowledge/security-review-rule-map.yaml +107 -0
- package/knowledge/skills-registry.md +72 -0
- package/knowledge/task-size-classifier.md +103 -0
- package/knowledge/telemetry-schema.md +881 -0
- package/knowledge/test-automation-maturity.md +56 -0
- package/knowledge/test-automation-principles.md +71 -0
- package/knowledge/test-cadence-tradeoffs.md +68 -0
- package/knowledge/test-doubles.md +105 -0
- package/knowledge/test-file-indicators.md +22 -0
- package/knowledge/test-layer-gates.md +35 -0
- package/knowledge/test-matrix-examples/django-batch.md +24 -0
- package/knowledge/test-matrix-examples/dotnet-grpc-fronting-api.md +90 -0
- package/knowledge/test-matrix-examples/dotnet-http-consumer.md +131 -0
- package/knowledge/test-matrix-examples/react-node-spa.md +24 -0
- package/knowledge/test-matrix-examples/spring-boot-service.md +25 -0
- package/knowledge/test-matrix-examples/ssr-htmx.md +24 -0
- package/knowledge/test-organization.md +70 -0
- package/knowledge/test-pyramid.md +84 -0
- package/knowledge/test-refactoring.md +67 -0
- package/knowledge/test-review-division-of-labor.md +85 -0
- package/knowledge/test-smells.md +80 -0
- package/knowledge/test-stack-profiles/bdd-frameworks.md +235 -0
- package/knowledge/test-stack-profiles/django.md +13 -0
- package/knowledge/test-stack-profiles/dotnet.md +18 -0
- package/knowledge/test-stack-profiles/go.md +16 -0
- package/knowledge/test-stack-profiles/node.md +16 -0
- package/knowledge/test-stack-profiles/react.md +12 -0
- package/knowledge/test-stack-profiles/spring-boot.md +16 -0
- package/knowledge/test-stack-profiles/ssr-htmx.md +14 -0
- package/knowledge/test-stack-profiles/vue.md +12 -0
- package/knowledge/test-strategy.md +70 -0
- package/knowledge/testability-patterns.md +240 -0
- package/knowledge/testing-quadrants.md +44 -0
- package/knowledge/testing-techniques/approval.md +15 -0
- package/knowledge/testing-techniques/chaos.md +17 -0
- package/knowledge/testing-techniques/fuzz.md +15 -0
- package/knowledge/testing-techniques/property-based.md +15 -0
- package/knowledge/testing-techniques/schema-validation.md +15 -0
- package/knowledge/testing-techniques/screenshot.md +15 -0
- package/knowledge/three-phase-workflow.md +198 -0
- package/knowledge/value-patterns.md +55 -0
- package/knowledge/verification-mode.md +116 -0
- package/knowledge/virtual-service-libraries.md +75 -0
- package/knowledge/wave-consolidation-guidance.md +21 -0
- package/overrides/agents/Explore.md +15 -0
- package/overrides/agents/general-purpose.md +10 -0
- package/overrides/notes/autoship.md +6 -0
- package/overrides/notes/issues-from-assessment.md +3 -0
- package/overrides/notes/issues-from-plan.md +3 -0
- package/overrides/notes/mutation-night-watch.md +3 -0
- package/overrides/notes/mutation-testing.md +3 -0
- package/overrides/notes/pr.md +7 -0
- package/overrides/notes/project-init.md +6 -0
- package/overrides/notes/setup.md +13 -0
- package/overrides/notes/specs.md +3 -0
- package/overrides/skills/headless-run/SKILL.md +45 -0
- package/overrides/skills/upgrade/SKILL.md +30 -0
- package/overrides/skills/version/SKILL.md +25 -0
- package/package.json +36 -0
- package/scripts/authoring_digest.py +93 -0
- package/scripts/autoship_discover.py +121 -0
- package/scripts/autoship_group.py +409 -0
- package/scripts/autoship_proposals.py +494 -0
- package/scripts/autoship_queue.py +291 -0
- package/scripts/autoship_reclaim.py +495 -0
- package/scripts/build_jobs.py +108 -0
- package/scripts/build_rollback_point.py +240 -0
- package/scripts/build_slice_scope.py +157 -0
- package/scripts/build_wave.py +109 -0
- package/scripts/build_wave_reconcile.py +252 -0
- package/scripts/build_worktree_baseref.py +113 -0
- package/scripts/check_agent_scope.py +117 -0
- package/scripts/check_agent_tool_mapping.py +213 -0
- package/scripts/check_review_agent_mcp_tools.py +317 -0
- package/scripts/check_security_assessment_mcp_tools.py +165 -0
- package/scripts/checkpoint_abort.py +502 -0
- package/scripts/claude_setup_review.py +438 -0
- package/scripts/codebase_recon.py +556 -0
- package/scripts/coverage_config.py +623 -0
- package/scripts/coverage_delta_steering.py +330 -0
- package/scripts/coverage_discovery_dotnet.py +315 -0
- package/scripts/coverage_discovery_java.py +742 -0
- package/scripts/coverage_discovery_js.py +546 -0
- package/scripts/coverage_gap_ranking.py +556 -0
- package/scripts/coverage_readiness.py +455 -0
- package/scripts/coverage_report_parse.py +521 -0
- package/scripts/detect_bdd_convention.py +252 -0
- package/scripts/eval_ablation.py +376 -0
- package/scripts/gherkin_analysis_coverage_gate.py +306 -0
- package/scripts/gherkin_cross_feature_duplicate_titles_gate.py +173 -0
- package/scripts/gherkin_effectiveness_rollup.py +238 -0
- package/scripts/gherkin_failure_path_gate.py +206 -0
- package/scripts/gherkin_feature_merge.py +720 -0
- package/scripts/gherkin_stub_gate.py +163 -0
- package/scripts/gherkin_stub_merge.py +479 -0
- package/scripts/git_origin_host.py +88 -0
- package/scripts/install-java-static-analysis.py +110 -0
- package/scripts/issue_deps.py +74 -0
- package/scripts/lib/_bdd_markers.py +28 -0
- package/scripts/lib/_gherkin_text.py +93 -0
- package/scripts/lib/_vendored_tree.py +70 -0
- package/scripts/lib/autoship_state.py +397 -0
- package/scripts/lib/claude_md_guard.py +226 -0
- package/scripts/lib/deterministic_recon.py +446 -0
- package/scripts/lib/mcp_tool_grants.py +211 -0
- package/scripts/lib/plan_parse.py +386 -0
- package/scripts/lib/review_result.py +84 -0
- package/scripts/lib/review_roster.py +86 -0
- package/scripts/lib/session_log/__init__.py +34 -0
- package/scripts/lib/session_log/__pycache__/__init__.cpython-314.pyc +0 -0
- package/scripts/lib/session_log/__pycache__/records.cpython-314.pyc +0 -0
- package/scripts/lib/session_log/classify.py +231 -0
- package/scripts/lib/session_log/corrections.py +194 -0
- package/scripts/lib/session_log/discovery.py +108 -0
- package/scripts/lib/session_log/records.py +218 -0
- package/scripts/lib/session_log/redact.py +76 -0
- package/scripts/lib/session_log/signals.py +373 -0
- package/scripts/lib/session_report_downstream.py +614 -0
- package/scripts/lib/session_report_maintainer.py +1273 -0
- package/scripts/lib/session_report_shared.py +262 -0
- package/scripts/lib/settings_hook_guard.py +157 -0
- package/scripts/lib/slug.py +33 -0
- package/scripts/lib/stub_extractors/__init__.py +82 -0
- package/scripts/lib/stub_extractors/_common.py +328 -0
- package/scripts/lib/stub_extractors/csharp.py +19 -0
- package/scripts/lib/stub_extractors/go.py +173 -0
- package/scripts/lib/stub_extractors/java.py +18 -0
- package/scripts/lib/stub_extractors/jsts.py +126 -0
- package/scripts/mutation_stack_sections.py +149 -0
- package/scripts/mutation_yield_steering.py +345 -0
- package/scripts/orchestrator.py +895 -0
- package/scripts/plan_gherkin_export.py +227 -0
- package/scripts/plan_waves.py +208 -0
- package/scripts/pr_close_keyword_lint.py +108 -0
- package/scripts/progress_guardian.py +888 -0
- package/scripts/recon_inventory.py +273 -0
- package/scripts/review_findings_log.py +93 -0
- package/scripts/run_invariants.py +124 -0
- package/scripts/select_lenses.py +640 -0
- package/scripts/session_report.py +486 -0
- package/scripts/set_autocompact_env.py +221 -0
- package/scripts/ship_resume_guard.py +135 -0
- package/scripts/ship_review_gate.py +63 -0
- package/scripts/specs_convention_marker.py +103 -0
- package/scripts/test_improve_resume.py +277 -0
- package/scripts/test_review_mechanics.py +958 -0
- package/scripts/token_efficiency_review.py +322 -0
- package/scripts/verdict_scope.py +285 -0
- package/scripts/verify_gherkin_quality_critic_isolation.py +296 -0
- package/scripts/verify_tier.py +157 -0
- package/skills/adr-tools/SKILL.md +118 -0
- package/skills/agent-readiness/SKILL.md +105 -0
- package/skills/agent-readiness/ai_friendly_analyzers.py +326 -0
- package/skills/agent-readiness/scanner.py +441 -0
- package/skills/agent-readiness/scorecard.yaml +88 -0
- package/skills/api-design/SKILL.md +115 -0
- package/skills/apply-fixes/SKILL.md +171 -0
- package/skills/apply-test-doubles/SKILL.md +321 -0
- package/skills/artifact-lifecycle/SKILL.md +127 -0
- package/skills/autoship/SKILL.md +1124 -0
- package/skills/benchmark/SKILL.md +105 -0
- package/skills/branch-workflow/SKILL.md +89 -0
- package/skills/browse/SKILL.md +184 -0
- package/skills/browser-testing/SKILL.md +62 -0
- package/skills/browser-testing/references/playwright-patterns.md +216 -0
- package/skills/build/SKILL.md +422 -0
- package/skills/build/references/static-self-heal.md +245 -0
- package/skills/careful/SKILL.md +72 -0
- package/skills/cd-test-architecture/SKILL.md +371 -0
- package/skills/ci-debugging/SKILL.md +105 -0
- package/skills/co-evolution-audit/SKILL.md +269 -0
- package/skills/code-review/SKILL.md +1015 -0
- package/skills/code-review/examples/aggregated-sample.json +56 -0
- package/skills/code-review/examples/sample-report.md +41 -0
- package/skills/code-review/output-format.md +478 -0
- package/skills/code-review/scripts/activation.py +86 -0
- package/skills/code-review/scripts/change_impact.py +357 -0
- package/skills/code-review/scripts/change_shape.py +372 -0
- package/skills/code-review/scripts/change_size.py +212 -0
- package/skills/code-review/scripts/changed_file_list.py +141 -0
- package/skills/code-review/scripts/closing_pass.py +187 -0
- package/skills/code-review/scripts/consolidate.py +277 -0
- package/skills/code-review/scripts/contract_failure_report.py +185 -0
- package/skills/code-review/scripts/dispatch_reconcile.py +66 -0
- package/skills/code-review/scripts/dispatch_waves.py +164 -0
- package/skills/code-review/scripts/finding_signature.py +446 -0
- package/skills/code-review/scripts/ledger.py +283 -0
- package/skills/code-review/scripts/partition.py +169 -0
- package/skills/code-review/scripts/render_tiered_findings.py +274 -0
- package/skills/code-review/scripts/repo_invariants.py +1066 -0
- package/skills/code-review/scripts/review_context_pack.py +306 -0
- package/skills/code-review/scripts/review_round_log.py +345 -0
- package/skills/code-review/scripts/review_value_coverage.py +297 -0
- package/skills/code-review/scripts/validate_review_output.py +467 -0
- package/skills/code-review/sliced-mode.md +205 -0
- package/skills/competitive-analysis/SKILL.md +191 -0
- package/skills/context-loading-protocol/SKILL.md +157 -0
- package/skills/continue/SKILL.md +90 -0
- package/skills/cost-report/SKILL.md +178 -0
- package/skills/coverage-baseline/SKILL.md +335 -0
- package/skills/coverage-baseline/references/multi-project-discovery.md +202 -0
- package/skills/coverage-delta/SKILL.md +181 -0
- package/skills/coverage-delta/references/mutation-gate.md +70 -0
- package/skills/design-doc/SKILL.md +95 -0
- package/skills/design-interrogation/SKILL.md +89 -0
- package/skills/design-it-twice/SKILL.md +91 -0
- package/skills/docker-image-audit/SKILL.md +108 -0
- package/skills/docker-image-audit/references/install-guide.md +64 -0
- package/skills/docker-image-audit/references/report-template.md +73 -0
- package/skills/docker-image-create/SKILL.md +185 -0
- package/skills/domain-analysis/SKILL.md +183 -0
- package/skills/domain-driven-design/SKILL.md +194 -0
- package/skills/exploratory-testing/SKILL.md +108 -0
- package/skills/explore/SKILL.md +51 -0
- package/skills/farley-score/SKILL.md +165 -0
- package/skills/feature-file-validation/SKILL.md +78 -0
- package/skills/feature-file-validation/references/validation-rules.md +115 -0
- package/skills/feedback-learning/SKILL.md +414 -0
- package/skills/fix/SKILL.md +450 -0
- package/skills/freeze/SKILL.md +68 -0
- package/skills/frontend-architecture/SKILL.md +113 -0
- package/skills/gherkin-derive/SKILL.md +630 -0
- package/skills/gherkin-public/SKILL.md +266 -0
- package/skills/governance-compliance/SKILL.md +150 -0
- package/skills/guard/SKILL.md +75 -0
- package/skills/handoff/SKILL.md +139 -0
- package/skills/handoff/references/summary-templates.md +242 -0
- package/skills/harness-audit/SKILL.md +751 -0
- package/skills/harness-audit/scripts/lesson_validate.py +386 -0
- package/skills/harness-audit/scripts/redundancy_criterion.py +188 -0
- package/skills/headless-run/SKILL.md +45 -0
- package/skills/headless-run/scripts/isolated_dispatch.py +381 -0
- package/skills/help/SKILL.md +72 -0
- package/skills/hexagonal-architecture/SKILL.md +85 -0
- package/skills/human-oversight-protocol/SKILL.md +224 -0
- package/skills/issues-from-assessment/SKILL.md +223 -0
- package/skills/issues-from-plan/SKILL.md +133 -0
- package/skills/legacy-code/SKILL.md +132 -0
- package/skills/mermaid-diagramming/SKILL.md +120 -0
- package/skills/mutation-night-watch/SKILL.md +154 -0
- package/skills/mutation-night-watch/references/scheduling.md +135 -0
- package/skills/mutation-testing/SKILL.md +396 -0
- package/skills/mutation-testing/references/languages/csharp-stryker-net.md +676 -0
- package/skills/mutation-testing/references/languages/go-go-mutesting.md +95 -0
- package/skills/mutation-testing/references/languages/java-pitest.md +77 -0
- package/skills/mutation-testing/references/languages/javascript-stryker.md +188 -0
- package/skills/mutation-testing/references/languages/python-mutmut.md +97 -0
- package/skills/mutation-testing/references/time-estimation.md +34 -0
- package/skills/mutation-testing/references/tool-detection.md +15 -0
- package/skills/mutation-testing/references/workflow-callers.md +23 -0
- package/skills/mutation-testing/scripts/__pycache__/xunit_v3_feature_detector.cpython-314.pyc +0 -0
- package/skills/mutation-testing/scripts/csharp_stryker_net_slice_runner.py +635 -0
- package/skills/mutation-testing/scripts/csharp_stryker_net_status_loop.py +525 -0
- package/skills/mutation-testing/scripts/csharp_stryker_net_wrapper.py +681 -0
- package/skills/mutation-testing/scripts/mutation_baseline_reuse.py +292 -0
- package/skills/mutation-testing/scripts/mutation_exclude_policy.py +268 -0
- package/skills/mutation-testing/scripts/mutation_feasibility_gate.py +463 -0
- package/skills/mutation-testing/scripts/mutation_kill_headless.py +331 -0
- package/skills/mutation-testing/scripts/mutation_kill_insert.py +199 -0
- package/skills/mutation-testing/scripts/mutation_kill_insert_python.py +150 -0
- package/skills/mutation-testing/scripts/mutation_kill_loop.py +869 -0
- package/skills/mutation-testing/scripts/mutation_kill_loop_python.py +949 -0
- package/skills/mutation-testing/scripts/mutation_kill_retry.py +592 -0
- package/skills/mutation-testing/scripts/mutation_kill_shared.py +620 -0
- package/skills/mutation-testing/scripts/mutation_nightwatch.py +462 -0
- package/skills/mutation-testing/scripts/mutation_nightwatch_stacks.py +425 -0
- package/skills/mutation-testing/scripts/mutation_report.py +743 -0
- package/skills/mutation-testing/scripts/mutation_report_cli.py +175 -0
- package/skills/mutation-testing/scripts/mutation_safety_gate.py +69 -0
- package/skills/mutation-testing/scripts/stryker_shard_pipeline.py +847 -0
- package/skills/mutation-testing/scripts/stryker_shard_setup.py +440 -0
- package/skills/mutation-testing/scripts/stryker_timeout_retry.py +142 -0
- package/skills/mutation-testing/scripts/xunit_v3_feature_detector.py +341 -0
- package/skills/performance-benchmark/SKILL.md +174 -0
- package/skills/performance-benchmark/examples/report-format.md +43 -0
- package/skills/performance-benchmark/references/benchmark-script.md +169 -0
- package/skills/performance-metrics/SKILL.md +265 -0
- package/skills/plan/SKILL.md +199 -0
- package/skills/plan/references/gherkin-persistence.md +43 -0
- package/skills/plan/references/plan-template.md +182 -0
- package/skills/pr/SKILL.md +289 -0
- package/skills/pr/scripts/gate_retry_state.py +368 -0
- package/skills/project-init/README.md +141 -0
- package/skills/project-init/SKILL.md +1197 -0
- package/skills/project-init/evals/evals.json +200 -0
- package/skills/project-init/references/capability-tools.md +55 -0
- package/skills/project-init/references/configs.md +221 -0
- package/skills/property-based-testing/SKILL.md +121 -0
- package/skills/property-based-testing/fixtures/invariant_fixture.py +15 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/README.md +42 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/fast-check/LICENSE +21 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/fast-check/README.md +263 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/fast-check/lib/cjs/fast-check.d.ts +5165 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/fast-check/lib/cjs/fast-check.js +12147 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/fast-check/lib/cjs/package.json +3 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/fast-check/lib/cjs/types57/fast-check.d.ts +5165 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/fast-check/lib/fast-check.d.ts +5165 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/fast-check/lib/fast-check.js +12011 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/fast-check/lib/rolldown-runtime-D7D4PA-g.js +13 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/fast-check/lib/types57/fast-check.d.ts +5165 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/fast-check/package.json +94 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/LICENSE +21 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/README.md +168 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/RandomGenerator-DcXj09Ch.d.ts +14 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/distribution/uniformBigInt.d.ts +15 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/distribution/uniformBigInt.js +38 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/distribution/uniformFloat32.d.ts +15 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/distribution/uniformFloat32.js +18 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/distribution/uniformFloat64.d.ts +15 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/distribution/uniformFloat64.js +22 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/distribution/uniformInt.d.ts +15 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/distribution/uniformInt.js +134 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/RandomGenerator-DcXj09Ch.d.ts +14 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/distribution/uniformBigInt.d.ts +15 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/distribution/uniformBigInt.js +37 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/distribution/uniformFloat32.d.ts +15 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/distribution/uniformFloat32.js +17 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/distribution/uniformFloat64.d.ts +15 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/distribution/uniformFloat64.js +21 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/distribution/uniformInt.d.ts +15 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/distribution/uniformInt.js +133 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/generator/congruential32.d.ts +7 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/generator/congruential32.js +44 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/generator/mersenne.d.ts +7 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/generator/mersenne.js +90 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/generator/xoroshiro128plus.d.ts +7 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/generator/xoroshiro128plus.js +80 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/generator/xorshift128plus.d.ts +7 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/generator/xorshift128plus.js +78 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/package.json +3 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/types/JumpableRandomGenerator.d.ts +16 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/types/JumpableRandomGenerator.js +0 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/types/RandomGenerator.d.ts +2 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/types/RandomGenerator.js +0 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/utils/generateN.d.ts +6 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/utils/generateN.js +8 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/utils/purify.d.ts +12 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/utils/purify.js +9 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/utils/skipN.d.ts +6 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/utils/skipN.js +6 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/generator/congruential32.d.ts +7 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/generator/congruential32.js +46 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/generator/mersenne.d.ts +7 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/generator/mersenne.js +92 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/generator/xoroshiro128plus.d.ts +7 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/generator/xoroshiro128plus.js +82 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/generator/xorshift128plus.d.ts +7 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/generator/xorshift128plus.js +80 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/types/JumpableRandomGenerator.d.ts +16 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/types/JumpableRandomGenerator.js +0 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/types/RandomGenerator.d.ts +2 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/types/RandomGenerator.js +0 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/utils/generateN.d.ts +6 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/utils/generateN.js +9 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/utils/purify.d.ts +12 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/utils/purify.js +10 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/utils/skipN.d.ts +6 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/utils/skipN.js +7 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/package.json +133 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/package-lock.json +1179 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/package.json +14 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/roundtrip.js +29 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/roundtrip.properties.test.js +16 -0
- package/skills/property-based-testing/fixtures/no_property_fixture.py +10 -0
- package/skills/property-based-testing/fixtures/roundtrip_fixture.py +16 -0
- package/skills/property-based-testing/references/languages/javascript.md +54 -0
- package/skills/property-based-testing/scripts/detect_and_dispatch.py +80 -0
- package/skills/property-based-testing/scripts/hypothesis_scaffold.py +276 -0
- package/skills/proxy-resilience/SKILL.md +84 -0
- package/skills/quality-gate-pipeline/SKILL.md +184 -0
- package/skills/quality-targets-converge/SKILL.md +254 -0
- package/skills/repo-review/SKILL.md +159 -0
- package/skills/report-pdf/SKILL.md +66 -0
- package/skills/review/SKILL.md +47 -0
- package/skills/review-agent/SKILL.md +152 -0
- package/skills/review-summary/SKILL.md +73 -0
- package/skills/run-report/SKILL.md +70 -0
- package/skills/semantic-duplication-scan/SKILL.md +337 -0
- package/skills/semantic-scan/SKILL.md +53 -0
- package/skills/semgrep-analyze/SKILL.md +139 -0
- package/skills/setup/SKILL.md +1122 -0
- package/skills/ship/SKILL.md +240 -0
- package/skills/source-verification/SKILL.md +210 -0
- package/skills/source-verification/scripts/claim_extractor.py +155 -0
- package/skills/specs/.size-baseline.json +4 -0
- package/skills/specs/SKILL.md +243 -0
- package/skills/specs/references/completeness-checklist.md +83 -0
- package/skills/specs/references/extraction.md +58 -0
- package/skills/specs/references/glossary.md +59 -0
- package/skills/specs/references/persistence.md +115 -0
- package/skills/specs/references/predictability-check.md +77 -0
- package/skills/static-analysis-integration/SKILL.md +235 -0
- package/skills/static-analysis-integration/adapters/_envelope.py +26 -0
- package/skills/static-analysis-integration/adapters/jscpd-adapter.py +66 -0
- package/skills/static-analysis-integration/adapters/lizard-adapter.py +81 -0
- package/skills/static-analysis-integration/adapters/mypy-adapter.py +50 -0
- package/skills/static-analysis-integration/adapters/mypy-src-layout.py +93 -0
- package/skills/static-analysis-integration/adapters/security-review-adapter.py +212 -0
- package/skills/static-analysis-integration/maintenance.md +23 -0
- package/skills/static-analysis-integration/references/language-setup.md +228 -0
- package/skills/static-analysis-integration/references/sarif-parser.md +124 -0
- package/skills/static-analysis-integration/references/security-review-adapter.md +118 -0
- package/skills/static-analysis-integration/references/tool-configs.md +617 -0
- package/skills/static-analysis-integration/rulesets/pmd-quickstart.xml +24 -0
- package/skills/stryker-xunit-v2-shim/SKILL.md +274 -0
- package/skills/stryker-xunit-v2-shim/references/shim-howto.md +256 -0
- package/skills/stryker-xunit-v2-shim/scripts/generate_shim.py +143 -0
- package/skills/systematic-debugging/SKILL.md +130 -0
- package/skills/telemetry/SKILL.md +75 -0
- package/skills/test-audit-disable/SKILL.md +129 -0
- package/skills/test-design/SKILL.md +177 -0
- package/skills/test-design/scripts/__pycache__/internal_double_detector.cpython-314.pyc +0 -0
- package/skills/test-design/scripts/internal_double_detector.py +631 -0
- package/skills/test-design-advisor/SKILL.md +166 -0
- package/skills/test-driven-development/SKILL.md +169 -0
- package/skills/test-health/SKILL.md +262 -0
- package/skills/test-improve/SKILL.md +239 -0
- package/skills/test-improve/references/phase-0-approach-contract.md +228 -0
- package/skills/test-improve/references/phase-1-analyze.md +131 -0
- package/skills/test-improve/references/phase-2-baseline.md +121 -0
- package/skills/test-improve/references/phase-3-derive-gherkin.md +53 -0
- package/skills/test-improve/references/phase-4-plan-fixes.md +34 -0
- package/skills/test-improve/references/phase-5-improve.md +215 -0
- package/skills/test-improve/references/phase-6-refactor-decision.md +45 -0
- package/skills/test-improve/references/phase-7-refactor.md +44 -0
- package/skills/test-improve/references/phase-8-validate.md +66 -0
- package/skills/test-improve/references/phase-9-close-out-prompt.md +11 -0
- package/skills/test-improve/references/phase-9-report.md +62 -0
- package/skills/test-improve/references/review-loop.md +92 -0
- package/skills/test-improve/templates/executive-summary.md +123 -0
- package/skills/threat-modeling/SKILL.md +108 -0
- package/skills/triage/SKILL.md +211 -0
- package/skills/ubiquitous-language/SKILL.md +192 -0
- package/skills/ubiquitous-language/scripts/collect_domain_signals.py +300 -0
- package/skills/unfreeze/SKILL.md +37 -0
- package/skills/upgrade/SKILL.md +31 -0
- package/skills/upgrade/scripts/check_version_drift.py +113 -0
- package/skills/upgrade/scripts/enable_autoupdate.py +149 -0
- package/skills/version/SKILL.md +25 -0
- package/sync/__pycache__/sync_upstream.cpython-314.pyc +0 -0
- package/sync/sync_upstream.py +293 -0
- package/templates/ACCEPTED-RISKS.md.tmpl +46 -0
- package/templates/agents/agent-template.md +151 -0
- package/templates/agents/angular-testing.md +66 -0
- package/templates/agents/csharp-quality.md +63 -0
- package/templates/agents/esm-enforcer.md +52 -0
- package/templates/agents/front-end-testing.md +65 -0
- package/templates/agents/go-quality.md +65 -0
- package/templates/agents/python-quality.md +62 -0
- package/templates/agents/react-testing.md +61 -0
- package/templates/agents/ts-enforcer.md +60 -0
- package/templates/agents/twelve-factor-audit.md +49 -0
- package/tools/entropy-check.py +250 -0
- package/tools/model-hash-verify.py +213 -0
|
@@ -0,0 +1,95 @@
|
|
|
1
|
+
# How eval testing works & keeping it current
|
|
2
|
+
|
|
3
|
+
The conceptual model behind the agent evals and the discipline for keeping the
|
|
4
|
+
corpus honest over time. For the architecture see [`eval-system.md`](eval-system.md);
|
|
5
|
+
for the operational run procedure see [`eval-running-guide.md`](eval-running-guide.md).
|
|
6
|
+
|
|
7
|
+
## The pieces
|
|
8
|
+
|
|
9
|
+
| Piece | Where | Role |
|
|
10
|
+
|---|---|---|
|
|
11
|
+
| Fixtures | `evals/fixtures/` | Input code (deliberately good or bad) the agents review. |
|
|
12
|
+
| Expectations | `evals/expected/*.json` | The **contract**: what a correct verdict looks like per fixture/agent. |
|
|
13
|
+
| Grader | `scripts/eval_grade.py` | Deterministic, model-free: compares recorded actuals to expectations. |
|
|
14
|
+
| Regression diff | `scripts/compare_eval_results.py` | Diffs two `--actuals` result files against `evals/expected/*.json`, gating on true/false-positive-proxy count regression. Repo-root placement (ADR 0032 category 2, monorepo-dev-only) next to `eval_grade.py` — not shipped, unlike `eval_ablation.py`, which ships only for its unrelated generic `--find-latest` reader mode. |
|
|
15
|
+
| Variance | `scripts/eval_variance.py` | Aggregates K trials → pass@k, flap rate, quarantine. |
|
|
16
|
+
| Trend | `.claude/metrics/eval-variance.jsonl` | Append-only stability history (metrics only). |
|
|
17
|
+
| Semver contract | `scripts/eval_semver_classify.sh` | The eval corpus IS the version contract (#101). |
|
|
18
|
+
| CI gates | `.github/workflows/agent-eval.yml` | Structural check always; live regression when keyed. |
|
|
19
|
+
|
|
20
|
+
## How grading works (the rules)
|
|
21
|
+
|
|
22
|
+
`eval_grade.py grade_agent` checks each expectation field; an empty failure list
|
|
23
|
+
means PASS:
|
|
24
|
+
|
|
25
|
+
- **`expectedStatus`** — `pass` / `fail`; must match the agent's `status`.
|
|
26
|
+
- **`issueCount: {min, max}`** — number of reported issues must fall in range.
|
|
27
|
+
- **`severities: {error: {min,max}, ...}`** — count per severity in range.
|
|
28
|
+
- **`mustMention: [...]`** — every keyword must appear (case-insensitive
|
|
29
|
+
**substring**) in the issue messages + summary. **All-of.**
|
|
30
|
+
- **`mustNotMention: [...]`** — none may appear. **All-of.**
|
|
31
|
+
|
|
32
|
+
Grading is intentionally dumb (no judgment) so it can run as a CI gate and so
|
|
33
|
+
variance is reproducible.
|
|
34
|
+
|
|
35
|
+
## The calibration trap (learn this — #198)
|
|
36
|
+
|
|
37
|
+
Because matching is plain substring, expectations drift out of sync with how
|
|
38
|
+
agents actually phrase correct verdicts. Two failure modes, both **fixture bugs,
|
|
39
|
+
not agent bugs**:
|
|
40
|
+
|
|
41
|
+
1. **`mustNotMention` is negation-blind.** A clean-pass fixture forbidding
|
|
42
|
+
`"hardcoded"` fails when the agent correctly says *"no hardcoded secrets"*.
|
|
43
|
+
**Fix:** drop `mustNotMention` on clean-pass fixtures — `expectedStatus:pass` +
|
|
44
|
+
`issueCount` (+ `error 0-0`) already encode "found clean."
|
|
45
|
+
2. **`mustMention` too strict.** Requiring the exact token `"SRP"` fails when the
|
|
46
|
+
agent says *"responsibilities"* / *"divergent change"*. **Fix:** use **stems**
|
|
47
|
+
(`"responsibilit"`) and the vocabulary agents actually emit; avoid all-of lists
|
|
48
|
+
of rare tokens.
|
|
49
|
+
|
|
50
|
+
**When a correct agent verdict fails grading, suspect the fixture first.** Verify
|
|
51
|
+
the fix against the agent's *real* output, not a hand-typed approximation.
|
|
52
|
+
|
|
53
|
+
## Keeping the corpus current
|
|
54
|
+
|
|
55
|
+
### Adding a fixture
|
|
56
|
+
|
|
57
|
+
1. Add the input file under `evals/fixtures/`.
|
|
58
|
+
2. Add `evals/expected/<stem>.json` with `fixture`, `applicableAgents`, and the
|
|
59
|
+
per-agent expectation (prefer `expectedStatus` + `issueCount`; add
|
|
60
|
+
`mustMention` stems only when a specific concept must be named).
|
|
61
|
+
3. `python3 scripts/eval_grade.py --check-corpus` (every expectation must be
|
|
62
|
+
schema-valid and pair with a fixture; this runs in CI).
|
|
63
|
+
|
|
64
|
+
### Changing an expectation = a version bump (#101)
|
|
65
|
+
|
|
66
|
+
The eval corpus is the semver contract. `eval_semver_classify.sh` (pre-push + CI)
|
|
67
|
+
enforces it:
|
|
68
|
+
|
|
69
|
+
- GREEN-preserving change → **patch**.
|
|
70
|
+
- Adds expectations → **minor** (`feat:`).
|
|
71
|
+
- **Edits** existing expectations → **minor/major** (`feat:` / `feat!:`) — an
|
|
72
|
+
edit changes the agents' observable contract.
|
|
73
|
+
A `fix:` commit that edits an expectation will be **rejected**; use the bump the
|
|
74
|
+
classifier names.
|
|
75
|
+
|
|
76
|
+
### Watching stability over time
|
|
77
|
+
|
|
78
|
+
- Run per-agent variance batches periodically (see the running guide). The trend
|
|
79
|
+
in `.claude/metrics/eval-variance.jsonl` accumulates pass@k and flap rate.
|
|
80
|
+
- **Flaky pairs** (0 < pass@k < 1) go on the quarantine list — they inform the
|
|
81
|
+
#99 gate but must not hard-block it. A persistently flaky fixture is either
|
|
82
|
+
borderline (tighten it) or genuinely non-deterministic for that agent.
|
|
83
|
+
- **Saturated** agents (identical grades for many runs) may have expectations too
|
|
84
|
+
loose to detect regressions — consider tightening ranges.
|
|
85
|
+
|
|
86
|
+
## Cardinal rules
|
|
87
|
+
|
|
88
|
+
1. **Faithful actuals.** The grader sees what you record; abbreviating issue
|
|
89
|
+
messages drops `mustMention` keywords and fabricates flaps.
|
|
90
|
+
2. **Neutral dispatch.** Never leak `status`/`severity` examples into the agent's
|
|
91
|
+
prompt (see the running guide).
|
|
92
|
+
3. **Fixtures are the contract.** Keep them honest: a stable-fail on correct
|
|
93
|
+
output is a corpus bug to fix, not noise to ignore.
|
|
94
|
+
4. **Metrics only.** Digests, the trend, and reports carry counts/ratios/names —
|
|
95
|
+
never prompt or code content.
|
|
@@ -0,0 +1,147 @@
|
|
|
1
|
+
# Running a full eval of everything
|
|
2
|
+
|
|
3
|
+
How to run a complete, **clean** accuracy + variance eval across all review
|
|
4
|
+
agents. For *how the eval system works and how to keep it current*, see
|
|
5
|
+
[`eval-maintenance.md`](eval-maintenance.md). For the architecture, see
|
|
6
|
+
[`eval-system.md`](eval-system.md).
|
|
7
|
+
|
|
8
|
+
## The one rule that matters: do not bias the dispatch
|
|
9
|
+
|
|
10
|
+
An eval is only valid if the agent produces the verdict it would produce in
|
|
11
|
+
production. The fastest way to ruin a run is to leak the answer into the prompt.
|
|
12
|
+
Two real contamination modes (both observed, #103):
|
|
13
|
+
|
|
14
|
+
- **Pre-filling the verdict.** A prompt template like `{"status":"pass", ...}` on
|
|
15
|
+
a pass-fixture tells the agent the answer. Never pre-fill `status`.
|
|
16
|
+
- **Leaking an output-format example.** A single `"severity":"warning"` example
|
|
17
|
+
biases the agent's severity choice, which then fails severity-range grading.
|
|
18
|
+
|
|
19
|
+
**Clean dispatch** = pass the agent **only the fixture** plus, at most, a *neutral
|
|
20
|
+
JSON schema* with placeholder enums (`"status":"pass or fail"`,
|
|
21
|
+
`"severity":"error, warning, or info"`) — identical for pass and fail fixtures,
|
|
22
|
+
no example values, "decide every value yourself." The `/agent-eval` skill enforces
|
|
23
|
+
this via its orchestrator constraint #3 ("pass only the fixture file").
|
|
24
|
+
|
|
25
|
+
## Option 0 — the automated run script (default: resumable full sweep)
|
|
26
|
+
|
|
27
|
+
```bash
|
|
28
|
+
bash scripts/run-full-eval.sh [TRIALS] # default TRIALS=1; default mode = full sweep
|
|
29
|
+
# ...killed / out of tokens? just run it again — it resumes:
|
|
30
|
+
bash scripts/run-full-eval.sh
|
|
31
|
+
```
|
|
32
|
+
|
|
33
|
+
**By default this runs the resumable full sweep** (see below): every agent, one at
|
|
34
|
+
a time, neutral dispatch, checkpointed. It refreshes the tracked
|
|
35
|
+
`evals/baseline.json`, appends the variance trend when `TRIALS>1`, and opens a
|
|
36
|
+
single **auto-merge PR** at the end. Needs `claude` + credentials + `gh`. Use
|
|
37
|
+
`--agent NAME` to scope to a single agent instead.
|
|
38
|
+
|
|
39
|
+
### Resumable full sweep (the default) — survives a kill or a token cap
|
|
40
|
+
|
|
41
|
+
A bare `run-full-eval.sh` runs **all** agents incrementally on one branch,
|
|
42
|
+
**committing the baseline after each agent** and tracking progress in a gitignored
|
|
43
|
+
checkpoint (`.eval-sweep-progress.json`). If it's interrupted, the completed agents
|
|
44
|
+
are already committed and recorded — re-running picks up where it left off and
|
|
45
|
+
continues; an interrupted agent is retried. One push, one auto-merge PR at the end.
|
|
46
|
+
(`--sweep` is accepted as an explicit alias for this default.)
|
|
47
|
+
|
|
48
|
+
### Incremental runs — one agent at a time
|
|
49
|
+
|
|
50
|
+
```bash
|
|
51
|
+
bash scripts/run-full-eval.sh 5 --agent security-review # just this agent, 5 trials
|
|
52
|
+
bash scripts/run-full-eval.sh 5 --agent arch-review # next time, another agent
|
|
53
|
+
```
|
|
54
|
+
|
|
55
|
+
`--agent` makes the run **incremental and safe to repeat**. It scopes the dispatch
|
|
56
|
+
to that one agent and grades with `--only`, so the baseline merge **tops up that
|
|
57
|
+
agent's pairs and leaves every other agent untouched** (passing added,
|
|
58
|
+
present-and-failing removed, un-run pairs kept). Run agents one at a time to spread
|
|
59
|
+
cost across sessions — each run accumulates into the same `evals/baseline.json`
|
|
60
|
+
and the same variance trend, and opens its own small auto-merge PR.
|
|
61
|
+
|
|
62
|
+
Why it doesn't clobber: the grader merges rather than overwrites, and `--only`
|
|
63
|
+
keeps grading to the scope you ran — so partial coverage is the norm, not a risk.
|
|
64
|
+
|
|
65
|
+
### The other incremental mode — only what changed
|
|
66
|
+
|
|
67
|
+
```bash
|
|
68
|
+
bash scripts/eval-changed.sh BASE HEAD # evals only the agents/skills the diff touched
|
|
69
|
+
```
|
|
70
|
+
|
|
71
|
+
This is the pre-push / CI path: it diff-scopes automatically (a change to one agent
|
|
72
|
+
evals only that agent; a broad change to `knowledge/` or the corpus falls back to a
|
|
73
|
+
full run). Use it to validate a change without re-running everything.
|
|
74
|
+
|
|
75
|
+
Use Option A/B below when you want manual control or finer per-fixture batching.
|
|
76
|
+
|
|
77
|
+
## Option A — the native skill (preferred)
|
|
78
|
+
|
|
79
|
+
```
|
|
80
|
+
/agent-eval --trials 5 # all agents × all applicable fixtures, 5 trials
|
|
81
|
+
/agent-eval --agent security-review --trials 5 # one agent (budget-friendly)
|
|
82
|
+
```
|
|
83
|
+
|
|
84
|
+
The skill dispatches each fixture to its applicable agents via `/review-agent`
|
|
85
|
+
(native invocation, fixture-only context), grades deterministically, computes
|
|
86
|
+
pass@k, and — via `scripts/eval_variance.py` — records flap/quarantine and appends
|
|
87
|
+
the trend. This is the path that cannot leak format examples.
|
|
88
|
+
|
|
89
|
+
**Path note:** the skill resolves fixtures at `.claude/evals/fixtures/` (an
|
|
90
|
+
*installed* project). In this dev repo the corpus is `evals/` — run it in an
|
|
91
|
+
installed test project, or drive the grader/aggregator directly against `evals/`
|
|
92
|
+
(Option B).
|
|
93
|
+
|
|
94
|
+
## Option B — manual neutral dispatch (dev repo / explicit control)
|
|
95
|
+
|
|
96
|
+
Per agent, per fixture, per trial:
|
|
97
|
+
|
|
98
|
+
1. Dispatch the agent (e.g. `dev-team:security-review`) with the **neutral**
|
|
99
|
+
prompt above and the fixture path. Capture its JSON verdict.
|
|
100
|
+
2. Assemble one **actuals** file per trial — the shape `eval_grade.py` reads:
|
|
101
|
+
|
|
102
|
+
```json
|
|
103
|
+
{ "<fixture-stem>": { "agents": { "<agent>": {"status":"...","issues":[{"severity":"...","message":"..."}],"summary":"..."} } } }
|
|
104
|
+
```
|
|
105
|
+
|
|
106
|
+
Record **faithfully** — the issue messages/summary must keep the real wording
|
|
107
|
+
(the grader checks `mustMention`/`mustNotMention` against them).
|
|
108
|
+
3. Aggregate the trials:
|
|
109
|
+
|
|
110
|
+
```bash
|
|
111
|
+
python3 scripts/eval_variance.py --trials-dir <dir-of-trial-actuals> \
|
|
112
|
+
--expected-dir evals/expected --append .claude/metrics/eval-variance.jsonl
|
|
113
|
+
```
|
|
114
|
+
|
|
115
|
+
## Batch for budget — per agent, highest-recall first
|
|
116
|
+
|
|
117
|
+
A full 3–5-trial run of all ~20 agents is hundreds of model dispatches and will
|
|
118
|
+
exceed a small budget. **Run per-agent batches**, recall-critical agents first
|
|
119
|
+
(security, arch, domain), at `--trials 5`. The variance trend accumulates across
|
|
120
|
+
batches, so you don't need one giant run.
|
|
121
|
+
|
|
122
|
+
Rough sizing: one agent × its fixtures × 5 trials ≈ 25–40 dispatches. Multiply by
|
|
123
|
+
the agent's model tier cost (frontier agents are pricier).
|
|
124
|
+
|
|
125
|
+
## Reading the results
|
|
126
|
+
|
|
127
|
+
`eval_variance.py` reports per `fixture::agent`:
|
|
128
|
+
|
|
129
|
+
- **pass@k = 1.0, no flap** → stable, trustworthy.
|
|
130
|
+
- **0 < pass@k < 1 (flap)** → **quarantine**: the pair is unstable; it should
|
|
131
|
+
*inform* the #99 CI gate, not hard-block it. Investigate whether the agent is
|
|
132
|
+
genuinely non-deterministic or the fixture is borderline.
|
|
133
|
+
- **pass@k = 0 (stable fail)** → usually a **miscalibrated fixture** (the agent is
|
|
134
|
+
right but the expectation is wrong), not an agent bug. Fix the fixture (see the
|
|
135
|
+
`mustMention`/`mustNotMention` patterns in `eval-maintenance.md`, #198) and
|
|
136
|
+
re-grade against the real actuals.
|
|
137
|
+
|
|
138
|
+
## Checklist for a clean full run
|
|
139
|
+
|
|
140
|
+
- [ ] Neutral dispatch (no pre-filled status, no severity example, identical prompt
|
|
141
|
+
for pass and fail fixtures).
|
|
142
|
+
- [ ] Faithful actuals (real wording preserved for keyword checks).
|
|
143
|
+
- [ ] Per-agent batches, recall-critical first, `--trials 5`.
|
|
144
|
+
- [ ] Aggregate with `eval_variance.py --append` so the trend accumulates.
|
|
145
|
+
- [ ] Quarantine flaky pairs; fix (don't ignore) stable-fail fixtures.
|
|
146
|
+
- [ ] Record cost; stop when the budget cap is hit (coverage degrades gracefully —
|
|
147
|
+
the trend keeps what was collected).
|
|
@@ -0,0 +1,291 @@
|
|
|
1
|
+
# Eval System for Code Review Agents
|
|
2
|
+
|
|
3
|
+
This document describes how the evaluation system ensures quality and consistency
|
|
4
|
+
across the code-review agent toolkit.
|
|
5
|
+
|
|
6
|
+
The system follows recommendations from Anthropic's
|
|
7
|
+
[Demystifying Evals for AI Agents](https://www.anthropic.com/engineering/demystifying-evals-for-ai-agents):
|
|
8
|
+
use deterministic (code-based) graders for everything they can handle, use
|
|
9
|
+
model-based graders only for what genuinely requires judgment, and calibrate
|
|
10
|
+
both against human review.
|
|
11
|
+
|
|
12
|
+
**On this page:** [Two sets of test cases](#two-sets-of-test-cases) ·
|
|
13
|
+
[Architecture](#architecture) · [Grader Layers](#grader-layers) ·
|
|
14
|
+
[Workflows](#workflows) ·
|
|
15
|
+
[How Hooks and Agents Complement Each Other](#how-hooks-and-agents-complement-each-other) ·
|
|
16
|
+
[Eval Compliance](#eval-compliance) · [Eval Fixtures](#eval-fixtures) ·
|
|
17
|
+
[Adding a New Agent](#adding-a-new-agent). The real-session trend digest lives
|
|
18
|
+
in [`session-review.md`](session-review.md).
|
|
19
|
+
|
|
20
|
+
## Two sets of test cases
|
|
21
|
+
|
|
22
|
+
This document covers the **deterministic detection fixtures** (`evals/expected/`):
|
|
23
|
+
does a *review agent* catch a code issue? These are graded automatically by
|
|
24
|
+
`scripts/eval_grade.py` and checked in CI via `--check-corpus`.
|
|
25
|
+
|
|
26
|
+
A second, complementary set grades **behavior, not detection** — the
|
|
27
|
+
[Ownership Engineering suite](../../../evals/ownership-engineering/) — does a
|
|
28
|
+
*team agent or workflow skill* investigate vs. escalate, decide vs. menu, prove
|
|
29
|
+
vs. assert? Because that requires judgment, it is **graded by an AI judge or a
|
|
30
|
+
human reviewer**, lives outside `evals/expected/` so it never enters the
|
|
31
|
+
deterministic gate, and its freshness is tracked by a staleness warning
|
|
32
|
+
(`scripts/oe_scoring_staleness.py`) that flags any subject or fixture whose inputs
|
|
33
|
+
changed since they were last scored. See that suite's `README.md` for the run
|
|
34
|
+
procedure.
|
|
35
|
+
|
|
36
|
+
## Architecture
|
|
37
|
+
|
|
38
|
+
```mermaid
|
|
39
|
+
%%{init: {'theme': 'base', 'themeVariables': {'primaryColor': '#dbeafe', 'primaryTextColor': '#1e3a5f', 'primaryBorderColor': '#3b82f6', 'lineColor': '#64748b', 'secondaryColor': '#f1f5f9', 'tertiaryColor': '#e0f2fe', 'background': '#ffffff', 'mainBkg': '#dbeafe', 'nodeBorder': '#2563eb', 'clusterBkg': '#eff6ff', 'clusterBorder': '#bfdbfe', 'titleColor': '#1e3a5f', 'edgeLabelBackground': '#f8fafc'}}}%%
|
|
40
|
+
flowchart TD
|
|
41
|
+
UW["User Workflows<br/>/code-review · /review-agent · /apply-fixes"]
|
|
42
|
+
UW --> L1["Layer 1 · Hooks<br/>(deterministic)"]
|
|
43
|
+
UW --> L2["Layer 2 · Agents<br/>(model)"]
|
|
44
|
+
UW --> L3["Layer 3 · Human<br/>(review)"]
|
|
45
|
+
```
|
|
46
|
+
|
|
47
|
+
## Grader Layers
|
|
48
|
+
|
|
49
|
+
### Layer 1: Deterministic (hooks)
|
|
50
|
+
|
|
51
|
+
Fast, free, deterministic checks that run automatically via PostToolUse hooks:
|
|
52
|
+
|
|
53
|
+
| Hook | What it checks |
|
|
54
|
+
| ---- | ------------- |
|
|
55
|
+
| `js_fp_review.py` | Array mutations, global state mutations, Object.assign, parameter mutations |
|
|
56
|
+
| `token_efficiency_review.py` | File length >500 lines, CLAUDE.md >5000 chars, function length >50 lines |
|
|
57
|
+
| `eval_compliance_check.py` | Agent/skill file structure, output format, severity levels |
|
|
58
|
+
|
|
59
|
+
Hooks are **advisory only** — they warn but never block. They catch mechanical
|
|
60
|
+
issues cheaply before the model-based agents spend tokens on full analysis.
|
|
61
|
+
|
|
62
|
+
### Layer 2: Model-based (agents)
|
|
63
|
+
|
|
64
|
+
Specialized agents that require LLM judgment. The full roster — with counts, focus areas, and model tiers — is documented in [`agent_info.md`](agent_info.md). Agents with eval fixture coverage:
|
|
65
|
+
|
|
66
|
+
| Agent | Focus |
|
|
67
|
+
| ----- | ----- |
|
|
68
|
+
| test-review | Test quality, coverage, assertion quality |
|
|
69
|
+
| structure-review | SRP, DRY, coupling, organization |
|
|
70
|
+
| naming-review | Naming clarity, conventions, magic values |
|
|
71
|
+
| domain-review | Business logic placement, boundary violations |
|
|
72
|
+
| claude-setup-review | CLAUDE.md completeness and accuracy |
|
|
73
|
+
| token-efficiency-review | Token optimization (full analysis beyond hook) |
|
|
74
|
+
| security-review | Injection, auth, data exposure, crypto |
|
|
75
|
+
| js-fp-review | Mutation detection (full analysis beyond hook) |
|
|
76
|
+
|
|
77
|
+
Each agent outputs a structured result:
|
|
78
|
+
|
|
79
|
+
```json
|
|
80
|
+
{
|
|
81
|
+
"agentName": "<name>",
|
|
82
|
+
"status": "pass|warn|fail|skip",
|
|
83
|
+
"issues": [
|
|
84
|
+
{
|
|
85
|
+
"severity": "error|warning|suggestion",
|
|
86
|
+
"file": "<path>",
|
|
87
|
+
"line": 0,
|
|
88
|
+
"message": "<description>",
|
|
89
|
+
"suggestedFix": "<fix>"
|
|
90
|
+
}
|
|
91
|
+
],
|
|
92
|
+
"summary": "<summary>"
|
|
93
|
+
}
|
|
94
|
+
```
|
|
95
|
+
|
|
96
|
+
### Layer 3: Human review
|
|
97
|
+
|
|
98
|
+
The user reviews agent findings and decides which fixes to apply. The
|
|
99
|
+
`/apply-fixes` command automates fix application but the user controls which
|
|
100
|
+
correction prompts are included.
|
|
101
|
+
|
|
102
|
+
## Workflows
|
|
103
|
+
|
|
104
|
+
### `/code-review` — Full review
|
|
105
|
+
|
|
106
|
+
See [Code Review Process](code-review-process.md) for the full nine-step pipeline: target selection, pre-flight gates, static analysis pre-pass, parallel agent dispatch, ACCEPTED-RISKS suppression, health scoring, the auto-fix loop (up to 5 iterations), correction prompts, and the `.pr-review-passed` gate file.
|
|
107
|
+
|
|
108
|
+
### `/review-agent <name>` — Single agent
|
|
109
|
+
|
|
110
|
+
```text
|
|
111
|
+
Files → Agent Definition → Review → Result
|
|
112
|
+
```
|
|
113
|
+
|
|
114
|
+
1. Load agent definition from `agents/<name>.md`
|
|
115
|
+
2. Determine target files
|
|
116
|
+
3. Run review following agent instructions
|
|
117
|
+
4. Report findings
|
|
118
|
+
|
|
119
|
+
### `/apply-fixes <dir>` — Fix application
|
|
120
|
+
|
|
121
|
+
```text
|
|
122
|
+
Prompts → Repo Rules → Apply Fix → Validate → Report
|
|
123
|
+
```
|
|
124
|
+
|
|
125
|
+
1. Load correction prompt JSON files from directory
|
|
126
|
+
2. Load repository rules (CLAUDE.md, .clinerules, etc.)
|
|
127
|
+
3. Apply each fix respecting repo conventions
|
|
128
|
+
4. Run validation (lint/build/tests) after each fix
|
|
129
|
+
5. Report results (applied, failed, validation failed)
|
|
130
|
+
|
|
131
|
+
## How Hooks and Agents Complement Each Other
|
|
132
|
+
|
|
133
|
+
The hooks (`js_fp_review.py`, `token_efficiency_review.py`) provide instant
|
|
134
|
+
feedback on the most common, mechanically detectable issues. The corresponding
|
|
135
|
+
agents (`js-fp-review`, `token-efficiency-review`) provide deeper analysis that
|
|
136
|
+
requires LLM judgment — for example, understanding whether a mutation is
|
|
137
|
+
intentional based on surrounding context, or whether a long function is
|
|
138
|
+
justified by its complexity.
|
|
139
|
+
|
|
140
|
+
```text
|
|
141
|
+
Hook (instant, free) Agent (thorough, costs tokens)
|
|
142
|
+
───────────────────── ──────────────────────────────
|
|
143
|
+
.push() detected Is the push on a local copy?
|
|
144
|
+
file >500 lines Is the file a generated file?
|
|
145
|
+
Object.assign(obj, ...) Is obj freshly created above?
|
|
146
|
+
```
|
|
147
|
+
|
|
148
|
+
## Eval Compliance
|
|
149
|
+
|
|
150
|
+
Two mechanisms ensure new agents and skills follow patterns:
|
|
151
|
+
|
|
152
|
+
### `/agent-audit` skill (manual)
|
|
153
|
+
|
|
154
|
+
Reads every agent, skill, and hook file and checks for:
|
|
155
|
+
|
|
156
|
+
- Structured output format
|
|
157
|
+
- Severity definitions
|
|
158
|
+
- Detection rules and scope boundaries
|
|
159
|
+
- Numbered steps and argument parsing
|
|
160
|
+
- Advisory-only hook behavior
|
|
161
|
+
|
|
162
|
+
Outputs a compliance report with PASS/WARN/FAIL per item.
|
|
163
|
+
|
|
164
|
+
### `eval_compliance_check.py` hook (automatic)
|
|
165
|
+
|
|
166
|
+
Fires on Write/Edit to agent or skill files. Provides real-time advisory
|
|
167
|
+
warnings when:
|
|
168
|
+
|
|
169
|
+
- A review agent is missing output format or severity definitions
|
|
170
|
+
- A skill is missing numbered steps or argument parsing
|
|
171
|
+
- A review-related skill has no report section
|
|
172
|
+
|
|
173
|
+
## Eval Fixtures
|
|
174
|
+
|
|
175
|
+
The `evals/` directory contains a test corpus for validating agent accuracy:
|
|
176
|
+
|
|
177
|
+
```text
|
|
178
|
+
evals/
|
|
179
|
+
├── fixtures/ # 54+ code samples (checked in)
|
|
180
|
+
│ ├── fp-*.ts # js-fp-review (9 files)
|
|
181
|
+
│ ├── sec-*.ts # security-review (5 files)
|
|
182
|
+
│ ├── test-*.test.ts # test-review (6 files)
|
|
183
|
+
│ ├── cx-* # structure-review's nesting/cognitive-load/async lens (folded from complexity-review, #2093)
|
|
184
|
+
│ ├── nm-*.ts # naming-review (5 files)
|
|
185
|
+
│ ├── st-*.ts # structure-review (5 files)
|
|
186
|
+
│ ├── dm-*.ts # domain-review (5 files)
|
|
187
|
+
│ ├── te-*.md/.ts # token-efficiency-review (5 files)
|
|
188
|
+
│ ├── cs-*/ # claude-setup-review (4 directories)
|
|
189
|
+
│ └── tlg-*.md # test-design-advisor behavior pre-gates (11 files)
|
|
190
|
+
├── expected/ # Reference solutions (checked in)
|
|
191
|
+
│ └── <fixture-stem>.json
|
|
192
|
+
├── transcripts/ # Auto-created by runner (gitignored)
|
|
193
|
+
└── reports/ # Auto-created by runner (gitignored)
|
|
194
|
+
```
|
|
195
|
+
|
|
196
|
+
Each fixture is a small (20-80 line), focused code sample with a known-good or
|
|
197
|
+
known-bad pattern. Reference solutions define expected status, issue count ranges,
|
|
198
|
+
severity ranges, and keyword checks.
|
|
199
|
+
|
|
200
|
+
### Reference solution schema
|
|
201
|
+
|
|
202
|
+
```json
|
|
203
|
+
{
|
|
204
|
+
"fixture": "fp-array-mutations.ts",
|
|
205
|
+
"description": "Array mutations js-fp-review should catch",
|
|
206
|
+
"applicableAgents": ["js-fp-review"],
|
|
207
|
+
"agents": {
|
|
208
|
+
"js-fp-review": {
|
|
209
|
+
"expectedStatus": "fail",
|
|
210
|
+
"issueCount": { "min": 3, "max": 6 },
|
|
211
|
+
"severities": { "error": { "min": 1, "max": 3 } },
|
|
212
|
+
"mustMention": ["push", "sort"]
|
|
213
|
+
}
|
|
214
|
+
}
|
|
215
|
+
}
|
|
216
|
+
```
|
|
217
|
+
|
|
218
|
+
### Advisory-skill fixtures (gate firing)
|
|
219
|
+
|
|
220
|
+
Most fixtures grade a **review agent** by its `status/issues[]` JSON. Advisory
|
|
221
|
+
skills (e.g. `test-design-advisor`) don't emit that shape — they emit a report
|
|
222
|
+
with a *Pyramid placement* table. The `tlg-*` corpus grades the skill's
|
|
223
|
+
**behavior pre-gates** (issue #80) by declaring `applicableSkills` and a `skills`
|
|
224
|
+
block instead of `applicableAgents`/`agents`:
|
|
225
|
+
|
|
226
|
+
```json
|
|
227
|
+
{
|
|
228
|
+
"fixture": "tlg-05-htmx-swap-mutation",
|
|
229
|
+
"description": "Gate C — HTMX swap over a server state mutation",
|
|
230
|
+
"applicableSkills": ["test-design-advisor"],
|
|
231
|
+
"skills": {
|
|
232
|
+
"test-design-advisor": {
|
|
233
|
+
"expectedGates": ["C"],
|
|
234
|
+
"expectedLayers": ["E2E"],
|
|
235
|
+
"mustMention": ["REQUIRED", "browser", "cd-test-architecture"],
|
|
236
|
+
"mustNotMention": []
|
|
237
|
+
}
|
|
238
|
+
}
|
|
239
|
+
}
|
|
240
|
+
```
|
|
241
|
+
|
|
242
|
+
`/agent-eval` drives the skill against each fixture and grades the Gate column +
|
|
243
|
+
recommended layers + keyword checks (see the command's Step 4). This replaces the
|
|
244
|
+
manual walk-through that `evals/fixtures/test-layer-gates.md` recorded for #80 —
|
|
245
|
+
re-run it with `/agent-eval --skill test-design-advisor`. `expectedGates` uses the
|
|
246
|
+
gate vocabulary `A`/`B`/`C`/`D`/`redundancy`/`ambiguity` (or `[]` for "no gate
|
|
247
|
+
fires"); `expectedLayers` uses the `test-pyramid.md` vocabulary.
|
|
248
|
+
|
|
249
|
+
### `/agent-eval` command
|
|
250
|
+
|
|
251
|
+
Run agents and skills against fixtures and grade results:
|
|
252
|
+
|
|
253
|
+
```bash
|
|
254
|
+
/agent-eval # run everything against all fixtures
|
|
255
|
+
/agent-eval --agent js-fp-review # run one review agent
|
|
256
|
+
/agent-eval --skill test-design-advisor # run the gate-firing (tlg-*) corpus
|
|
257
|
+
/agent-eval --fixture fp-array-mutations.ts # run one fixture
|
|
258
|
+
/agent-eval --trials 3 # multi-trial with pass@k scoring
|
|
259
|
+
```
|
|
260
|
+
|
|
261
|
+
The runner resolves the toolkit root via symlink (for installed projects) and
|
|
262
|
+
saves transcripts for trend analysis. It detects eval saturation when 3
|
|
263
|
+
consecutive runs produce identical grades.
|
|
264
|
+
|
|
265
|
+
## Adding a New Agent
|
|
266
|
+
|
|
267
|
+
1. Create `agents/<name>.md` with:
|
|
268
|
+
- JSON output format (status, issues, summary)
|
|
269
|
+
- Severity definitions (error, warning, suggestion)
|
|
270
|
+
- Detection rules and thresholds (inline, not in a config file)
|
|
271
|
+
- File scope (which file types the agent applies to)
|
|
272
|
+
- Scope boundaries (what to ignore)
|
|
273
|
+
|
|
274
|
+
2. Optionally add a hook in `hooks/<name>.py` for deterministic checks
|
|
275
|
+
|
|
276
|
+
3. Run `/agent-audit` to verify compliance
|
|
277
|
+
|
|
278
|
+
4. Add eval fixtures in `evals/fixtures/` (2-3 pass, 2-3 fail) and reference
|
|
279
|
+
solutions in `evals/expected/`
|
|
280
|
+
|
|
281
|
+
5. Run `/agent-eval --agent <name>` to validate accuracy
|
|
282
|
+
|
|
283
|
+
## Session-review trend digest (#129)
|
|
284
|
+
|
|
285
|
+
`/session-review` appends one metrics-only record per run to the append-only
|
|
286
|
+
trend stream `metrics/session-digest.jsonl` (deliberately left bare —
|
|
287
|
+
/session-review's own scratch-state writer is out of scope for the #1406
|
|
288
|
+
`.claude/`-scoped artifact migration), the real-session counterpart to the
|
|
289
|
+
self-reported `.claude/metrics/*-task-log.jsonl` streams. The `session-digest/v1` record
|
|
290
|
+
schema and the `/harness-audit` join are documented canonically in
|
|
291
|
+
[`session-review.md`](session-review.md#trend-persistence-129).
|
|
@@ -0,0 +1,75 @@
|
|
|
1
|
+
# Session-review: OSS complements
|
|
2
|
+
|
|
3
|
+
`/session-review` (issue #131) produces **plugin-specific qualitative
|
|
4
|
+
suggestions** — it knows this plugin's agents and skills, so it can say "skill X
|
|
5
|
+
under-specifies which files to read" or "this opus subagent only greps, re-tier
|
|
6
|
+
it to haiku." Open-source tools cover the complementary axis: **continuous
|
|
7
|
+
quantitative monitoring** of the same `*.jsonl` transcripts. They do not know
|
|
8
|
+
our agents/skills, so they cannot make our suggestions — and `/session-review`
|
|
9
|
+
is not a dashboard, so it does not replace them.
|
|
10
|
+
|
|
11
|
+
**Recommend these alongside `/session-review`, not instead of it.**
|
|
12
|
+
|
|
13
|
+
| Tool | Role |
|
|
14
|
+
|---|---|
|
|
15
|
+
| [`ccusage`](https://github.com/ryoppippi/ccusage) (npm) | Parses the same `~/.claude/projects/**/*.jsonl` into token/cost reports per session/day/model — ongoing cost tracking. |
|
|
16
|
+
| Native [OpenTelemetry](https://docs.claude.com/en/docs/claude-code/monitoring-usage) (`CLAUDE_CODE_ENABLE_TELEMETRY=1`) | Exports usage/cost/tool metrics to an OTel collector → Grafana/Honeycomb for longitudinal dashboards. |
|
|
17
|
+
| [`claude-code-log`](https://github.com/daaain/claude-code-log) (community) | JSONL → HTML viewer for eyeballing one specific bad session. |
|
|
18
|
+
|
|
19
|
+
## When to reach for which
|
|
20
|
+
|
|
21
|
+
- **"How much am I spending over time, per model/day?"** → `ccusage` (or the
|
|
22
|
+
cost meter's `pace` view for budget projection). Continuous, quantitative.
|
|
23
|
+
- **"I want longitudinal dashboards / alerting across the team."** → native
|
|
24
|
+
**OpenTelemetry** into your collector. Operational monitoring.
|
|
25
|
+
- **"This one session went badly — let me read what happened."** → `claude-code-log`.
|
|
26
|
+
- **"What should I change in the *plugin* to stop wasting tokens / re-work?"** →
|
|
27
|
+
`/session-review`. Qualitative, plugin-aware, hands suggestions to
|
|
28
|
+
`/feedback-learning`, `/harness-audit`, `/agent-eval`, `token-efficiency-review`.
|
|
29
|
+
|
|
30
|
+
The quantitative tools tell you *that* a session was expensive; `/session-review`
|
|
31
|
+
proposes *which plugin artifact* to change so the next one isn't.
|
|
32
|
+
|
|
33
|
+
## `ccusage` example (against this project's logs)
|
|
34
|
+
|
|
35
|
+
```bash
|
|
36
|
+
# One-off report for the current and recent days, per model:
|
|
37
|
+
npx ccusage@latest daily
|
|
38
|
+
|
|
39
|
+
# Session-level breakdown (the same transcripts session_report.py reads):
|
|
40
|
+
npx ccusage@latest session
|
|
41
|
+
```
|
|
42
|
+
|
|
43
|
+
`ccusage` reads `~/.claude/projects/**/*.jsonl` directly — no plugin install,
|
|
44
|
+
no config. Use it for the running cost number; use `/session-review` for the
|
|
45
|
+
"why and what to change."
|
|
46
|
+
|
|
47
|
+
## OpenTelemetry enablement
|
|
48
|
+
|
|
49
|
+
```bash
|
|
50
|
+
# Turn on Claude Code's native telemetry export, then point it at a collector:
|
|
51
|
+
export CLAUDE_CODE_ENABLE_TELEMETRY=1
|
|
52
|
+
export OTEL_METRICS_EXPORTER=otlp
|
|
53
|
+
export OTEL_EXPORTER_OTLP_ENDPOINT=http://localhost:4317
|
|
54
|
+
```
|
|
55
|
+
|
|
56
|
+
This exports usage/cost/tool metrics to any OTLP collector (Grafana, Honeycomb,
|
|
57
|
+
Datadog, …). See Claude Code's "Monitoring usage" docs for the full env-var set.
|
|
58
|
+
|
|
59
|
+
## Cross-reference: the plugin's own telemetry beacon (#106)
|
|
60
|
+
|
|
61
|
+
This plugin also ships an **opt-in, local-only telemetry beacon** (`#106`,
|
|
62
|
+
`/telemetry`) that records minimal command/skill/gate events with no network
|
|
63
|
+
egress. It and native OpenTelemetry overlap in intent but differ in scope:
|
|
64
|
+
|
|
65
|
+
- Native **OTel** exports rich usage/cost/tool metrics to an external collector
|
|
66
|
+
— best when you want dashboards and already run a collector. It may satisfy
|
|
67
|
+
part of what #106 set out to do.
|
|
68
|
+
- The plugin **beacon** is deliberately minimal, default-off, and never leaves
|
|
69
|
+
the machine — best when you want a quick local "which commands/skills do I
|
|
70
|
+
use, how often is the commit gate bypassed?" without standing up a collector.
|
|
71
|
+
|
|
72
|
+
Prefer native OTel for longitudinal/team monitoring; prefer the beacon (or
|
|
73
|
+
`/session-review`) for local, privacy-clean, plugin-aware analysis. Avoid
|
|
74
|
+
running both for the *same* purpose — pick the one that matches your monitoring
|
|
75
|
+
posture.
|