pi-dev-team 0.2.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/LICENSE +21 -0
- package/PORTING.md +134 -0
- package/README.md +207 -0
- package/UPSTREAM.json +64 -0
- package/agents/Explore.md +15 -0
- package/agents/a11y-review.md +118 -0
- package/agents/adr-author.md +70 -0
- package/agents/ai-provenance-review.md +120 -0
- package/agents/angular-reactivity-review.md +95 -0
- package/agents/arch-review.md +135 -0
- package/agents/architect.md +78 -0
- package/agents/autoship-batch-proposer.md +69 -0
- package/agents/claude-setup-review.md +136 -0
- package/agents/codebase-recon.md +184 -0
- package/agents/component-architecture-review.md +119 -0
- package/agents/concurrency-review.md +109 -0
- package/agents/correctness-review.md +290 -0
- package/agents/data-flow-tracer.md +120 -0
- package/agents/doc-review.md +165 -0
- package/agents/domain-review.md +136 -0
- package/agents/general-purpose.md +10 -0
- package/agents/gherkin-quality-critic.md +113 -0
- package/agents/js-fp-review.md +114 -0
- package/agents/mutation-kill.md +684 -0
- package/agents/naming-review.md +142 -0
- package/agents/orchestrator.md +339 -0
- package/agents/performance-review.md +105 -0
- package/agents/plan-review-acceptance.md +115 -0
- package/agents/plan-review-design.md +90 -0
- package/agents/plan-review-parallelization.md +84 -0
- package/agents/plan-review-strategic.md +96 -0
- package/agents/plan-review-ux.md +110 -0
- package/agents/platform-engineer.md +64 -0
- package/agents/product-manager.md +68 -0
- package/agents/progress-guardian.md +79 -0
- package/agents/qa-engineer.md +289 -0
- package/agents/quality-reviewer.md +132 -0
- package/agents/react-reactivity-review.md +102 -0
- package/agents/refactor-opportunity-review.md +128 -0
- package/agents/security-engineer.md +60 -0
- package/agents/security-review.md +218 -0
- package/agents/session-analysis.md +95 -0
- package/agents/software-engineer.md +105 -0
- package/agents/spec-compliance-review.md +100 -0
- package/agents/spec-reviewer.md +114 -0
- package/agents/structure-review.md +146 -0
- package/agents/tech-writer.md +84 -0
- package/agents/test-review.md +246 -0
- package/agents/test-smell-review.md +188 -0
- package/agents/token-efficiency-review.md +139 -0
- package/agents/ui-ux-designer.md +54 -0
- package/agents/vue-reactivity-review.md +95 -0
- package/bin/__pycache__/claudecpython-314.pyc +0 -0
- package/bin/claude +258 -0
- package/docs/upstream/.pages +1 -0
- package/docs/upstream/CHANGELOG.md +2586 -0
- package/docs/upstream/README.md +155 -0
- package/docs/upstream/agent-architecture.md +214 -0
- package/docs/upstream/agent_info.md +187 -0
- package/docs/upstream/artifact-migration.md +124 -0
- package/docs/upstream/code-intelligence-nudge.md +149 -0
- package/docs/upstream/code-review-process.md +294 -0
- package/docs/upstream/concurrent-use.md +73 -0
- package/docs/upstream/context-management.md +111 -0
- package/docs/upstream/developer-notes.md +280 -0
- package/docs/upstream/diagrams/architecture-overview.svg +101 -0
- package/docs/upstream/diagrams/review-dispatch.svg +139 -0
- package/docs/upstream/diagrams/team-agents.svg +128 -0
- package/docs/upstream/diagrams/test-improve-flow.svg +166 -0
- package/docs/upstream/diagrams/workflow-linear.svg +66 -0
- package/docs/upstream/diagrams/workflow-three-phase.svg +200 -0
- package/docs/upstream/eval-maintenance.md +95 -0
- package/docs/upstream/eval-running-guide.md +147 -0
- package/docs/upstream/eval-system.md +291 -0
- package/docs/upstream/session-review-oss-complements.md +75 -0
- package/docs/upstream/session-review.md +212 -0
- package/docs/upstream/skills.md +188 -0
- package/docs/upstream/team-structure.md +21 -0
- package/docs/upstream/telemetry-ci-access.md +129 -0
- package/docs/upstream/telemetry-repo-security.md +120 -0
- package/docs/upstream/test-evaluation.md +277 -0
- package/docs/upstream/test-improve.md +154 -0
- package/docs/upstream/triage-workflow.md +282 -0
- package/docs/upstream/workflows.md +289 -0
- package/extensions/dev-team/index.ts +539 -0
- package/extensions/dev-team/lib/agents.ts +272 -0
- package/extensions/dev-team/lib/ai-credits.ts +92 -0
- package/extensions/dev-team/lib/autocompact.ts +81 -0
- package/extensions/dev-team/lib/child-run.ts +102 -0
- package/extensions/dev-team/lib/config.ts +236 -0
- package/extensions/dev-team/lib/gh-command.ts +103 -0
- package/extensions/dev-team/lib/github-style.ts +307 -0
- package/extensions/dev-team/lib/hooks.ts +350 -0
- package/extensions/dev-team/lib/metrics.ts +115 -0
- package/extensions/dev-team/lib/safe-read.ts +49 -0
- package/extensions/dev-team/lib/session-files.ts +57 -0
- package/extensions/dev-team/lib/session-spend.ts +123 -0
- package/extensions/dev-team/lib/shell-scan.ts +205 -0
- package/extensions/dev-team/lib/skills.ts +213 -0
- package/extensions/dev-team/lib/subagent-render.ts +245 -0
- package/extensions/dev-team/lib/subagent-types.ts +164 -0
- package/extensions/dev-team/lib/subagent.ts +596 -0
- package/extensions/dev-team/lib/terminal-text.ts +54 -0
- package/extensions/dev-team/lib/tools-misc.ts +152 -0
- package/extensions/dev-team/lib/transcript.ts +110 -0
- package/extensions/dev-team/lib/trust.ts +52 -0
- package/extensions/dev-team/lib/usage-breakdown.ts +176 -0
- package/extensions/dev-team/lib/usage-chart.ts +153 -0
- package/extensions/dev-team/lib/usage-command.ts +107 -0
- package/extensions/dev-team/lib/usage-history.ts +203 -0
- package/extensions/dev-team/lib/usage-render.ts +225 -0
- package/extensions/dev-team/lib/usage-split-bar.ts +127 -0
- package/extensions/dev-team/lib/usage-state.ts +116 -0
- package/extensions/dev-team/lib/usage-text.ts +159 -0
- package/extensions/dev-team/lib/usage-view.ts +109 -0
- package/hooks/__pycache__/refactor_test_freeze_guard.cpython-314.pyc +0 -0
- package/hooks/agent_dispatch_ledger.py +190 -0
- package/hooks/autocompact_setup_nudge.py +99 -0
- package/hooks/bash_retry_guard.py +228 -0
- package/hooks/boundary_events_write_guard.py +352 -0
- package/hooks/code_intelligence_nudge.py +293 -0
- package/hooks/code_intelligence_turn_mark.py +317 -0
- package/hooks/codegraph_bootstrap.py +139 -0
- package/hooks/contract_version_guard.py +362 -0
- package/hooks/cost_meter.py +106 -0
- package/hooks/destructive-commands.json +62 -0
- package/hooks/destructive_guard.py +477 -0
- package/hooks/eval_compliance_check.py +440 -0
- package/hooks/guards.json +17 -0
- package/hooks/hooks.json +323 -0
- package/hooks/internal_double_gate.py +296 -0
- package/hooks/js_fp_review.py +212 -0
- package/hooks/knowledge_index.py +119 -0
- package/hooks/lib/__pycache__/artifact_paths.cpython-314.pyc +0 -0
- package/hooks/lib/__pycache__/atomic_state.cpython-314.pyc +0 -0
- package/hooks/lib/__pycache__/autocompact_config.cpython-314.pyc +0 -0
- package/hooks/lib/__pycache__/boundary_events.cpython-314.pyc +0 -0
- package/hooks/lib/__pycache__/doc_classification.cpython-314.pyc +0 -0
- package/hooks/lib/__pycache__/gh_pr_create_detect.cpython-314.pyc +0 -0
- package/hooks/lib/__pycache__/git_safe_diff.cpython-314.pyc +0 -0
- package/hooks/lib/__pycache__/instrument_log.cpython-314.pyc +0 -0
- package/hooks/lib/__pycache__/metrics_query.cpython-314.pyc +0 -0
- package/hooks/lib/__pycache__/plugin_version.cpython-314.pyc +0 -0
- package/hooks/lib/__pycache__/pre_commit_doc_classifier.cpython-314.pyc +0 -0
- package/hooks/lib/__pycache__/review_agent_registry.cpython-314.pyc +0 -0
- package/hooks/lib/__pycache__/review_gate_corroboration.cpython-314.pyc +0 -0
- package/hooks/lib/__pycache__/review_gate_hash.cpython-314.pyc +0 -0
- package/hooks/lib/__pycache__/review_verdicts.cpython-314.pyc +0 -0
- package/hooks/lib/__pycache__/stdin_json.cpython-314.pyc +0 -0
- package/hooks/lib/__pycache__/stryker_invocation.cpython-314.pyc +0 -0
- package/hooks/lib/__pycache__/telemetry_consent.cpython-314.pyc +0 -0
- package/hooks/lib/__pycache__/test_file_classify.cpython-314.pyc +0 -0
- package/hooks/lib/__pycache__/token_efficiency_limits.cpython-314.pyc +0 -0
- package/hooks/lib/__pycache__/verify_guard_state.cpython-314.pyc +0 -0
- package/hooks/lib/__pycache__/xunit_v3_operator_gate.cpython-314.pyc +0 -0
- package/hooks/lib/agent_skill_hints.py +74 -0
- package/hooks/lib/artifact_paths.py +263 -0
- package/hooks/lib/atomic_state.py +557 -0
- package/hooks/lib/autocompact_config.py +103 -0
- package/hooks/lib/autoship_log.py +106 -0
- package/hooks/lib/banned_scripts_policy.py +51 -0
- package/hooks/lib/boundary_events.py +436 -0
- package/hooks/lib/build_knowledge_index.py +504 -0
- package/hooks/lib/build_skills_index.py +361 -0
- package/hooks/lib/build_state.py +116 -0
- package/hooks/lib/classify_ship_outcome.py +126 -0
- package/hooks/lib/config_changelog_schema.py +115 -0
- package/hooks/lib/cost_meter.py +955 -0
- package/hooks/lib/doc_classification.py +116 -0
- package/hooks/lib/gh_pr_create_detect.py +136 -0
- package/hooks/lib/git_safe_diff.py +123 -0
- package/hooks/lib/instrument_log.py +66 -0
- package/hooks/lib/iteration_journal_gate.py +197 -0
- package/hooks/lib/knowledge_index_paths.py +88 -0
- package/hooks/lib/mcp_json_repowise.py +177 -0
- package/hooks/lib/metrics_query.py +202 -0
- package/hooks/lib/minimal_yaml.py +434 -0
- package/hooks/lib/plugin_version.py +142 -0
- package/hooks/lib/pre_commit_detect.py +537 -0
- package/hooks/lib/pre_commit_doc_classifier.py +126 -0
- package/hooks/lib/pricing.py +118 -0
- package/hooks/lib/report_pdf.py +371 -0
- package/hooks/lib/review_agent_registry.py +142 -0
- package/hooks/lib/review_dispatch_ledger.py +101 -0
- package/hooks/lib/review_gate_corroboration.py +521 -0
- package/hooks/lib/review_gate_hash.py +252 -0
- package/hooks/lib/review_gate_normalized_hash.py +1115 -0
- package/hooks/lib/review_verdicts.py +301 -0
- package/hooks/lib/run_report.py +160 -0
- package/hooks/lib/skill_categories.yaml +125 -0
- package/hooks/lib/stdin_json.py +57 -0
- package/hooks/lib/stryker_invocation.py +102 -0
- package/hooks/lib/telemetry_consent.py +41 -0
- package/hooks/lib/telemetry_report.py +108 -0
- package/hooks/lib/test_file_classify.py +160 -0
- package/hooks/lib/token_efficiency_limits.py +51 -0
- package/hooks/lib/turn_identity.py +77 -0
- package/hooks/lib/verify_guard_state.py +110 -0
- package/hooks/lib/workflow_state.py +206 -0
- package/hooks/lib/xunit_v3_operator_gate.py +596 -0
- package/hooks/mcp_json_repowise_nudge.py +74 -0
- package/hooks/mutation_adapters/__init__.py +7 -0
- package/hooks/mutation_adapters/__pycache__/__init__.cpython-314.pyc +0 -0
- package/hooks/mutation_adapters/__pycache__/lib.cpython-314.pyc +0 -0
- package/hooks/mutation_adapters/__pycache__/mutmut.cpython-314.pyc +0 -0
- package/hooks/mutation_adapters/__pycache__/pitest.cpython-314.pyc +0 -0
- package/hooks/mutation_adapters/__pycache__/stryker.cpython-314.pyc +0 -0
- package/hooks/mutation_adapters/__pycache__/stryker_net.cpython-314.pyc +0 -0
- package/hooks/mutation_adapters/lib.py +478 -0
- package/hooks/mutation_adapters/mutmut.py +188 -0
- package/hooks/mutation_adapters/pitest.py +266 -0
- package/hooks/mutation_adapters/stryker.py +157 -0
- package/hooks/mutation_adapters/stryker_net.py +264 -0
- package/hooks/mutation_gate.py +193 -0
- package/hooks/mutation_testing_smoke_gate.py +371 -0
- package/hooks/pending_review_notify.py +121 -0
- package/hooks/phase_marker.py +138 -0
- package/hooks/post_compact_state_reinject.py +180 -0
- package/hooks/post_format.py +115 -0
- package/hooks/pre_commit_knowledge_index.py +128 -0
- package/hooks/pre_commit_review.py +66 -0
- package/hooks/pre_pr_review.py +694 -0
- package/hooks/pre_tool_guard.py +405 -0
- package/hooks/py.sh +73 -0
- package/hooks/refactor-bash-write-patterns.json +29 -0
- package/hooks/refactor_test_bash_guard.py +253 -0
- package/hooks/refactor_test_freeze_guard.py +139 -0
- package/hooks/refactor_test_revert_guard.py +186 -0
- package/hooks/repo_review_nudge.py +287 -0
- package/hooks/review_verdict_recorder.py +464 -0
- package/hooks/scan_bash_command_for_banned_scripts.py +428 -0
- package/hooks/scan_worktree_for_banned_scripts.py +238 -0
- package/hooks/session_learning_trigger.py +248 -0
- package/hooks/skills_index.py +126 -0
- package/hooks/stryker_xunit_shim_guard.py +571 -0
- package/hooks/subagent_completion_guard.py +309 -0
- package/hooks/subagent_skill_context.py +139 -0
- package/hooks/task_completion_metrics.py +216 -0
- package/hooks/tdd_guard.py +229 -0
- package/hooks/telemetry.py +341 -0
- package/hooks/token_efficiency_review.py +194 -0
- package/hooks/verify_guard.py +183 -0
- package/hooks/verify_guard_edit_marker.py +73 -0
- package/hooks/version_check.py +173 -0
- package/knowledge/accepted-risks-schema.md +98 -0
- package/knowledge/adr-decision-criteria.md +64 -0
- package/knowledge/adversarial-review-protocol.md +139 -0
- package/knowledge/agent-registry.md +228 -0
- package/knowledge/agent-review-methodology.md +80 -0
- package/knowledge/ai-friendly-repo-guidelines.md +67 -0
- package/knowledge/architecture-assessment.md +96 -0
- package/knowledge/artifact-lifecycle.md +57 -0
- package/knowledge/cd-maturity-model.md +82 -0
- package/knowledge/cd-test-architecture.md +190 -0
- package/knowledge/ci-cd-file-scope.md +24 -0
- package/knowledge/codegraph-vs-graphify.md +192 -0
- package/knowledge/component-test-patterns.md +139 -0
- package/knowledge/database-change-management.md +80 -0
- package/knowledge/database-test-patterns.md +79 -0
- package/knowledge/decision-defaults.md +88 -0
- package/knowledge/dependency-breaking-techniques.md +116 -0
- package/knowledge/deployment-pipeline.md +86 -0
- package/knowledge/design-smells.md +122 -0
- package/knowledge/directory-enumeration.md +38 -0
- package/knowledge/domain-modeling.md +123 -0
- package/knowledge/evidence-bundle.md +90 -0
- package/knowledge/exploratory-testing-field-guide.md +122 -0
- package/knowledge/failure-routing.md +28 -0
- package/knowledge/fixture-construction.md +56 -0
- package/knowledge/frontend-component-architecture.md +139 -0
- package/knowledge/gherkin-quality-review-dispatch.md +135 -0
- package/knowledge/index.json +6766 -0
- package/knowledge/internal-collaborator-doubling.md +101 -0
- package/knowledge/legacy-test-strategy.md +71 -0
- package/knowledge/long-run-waiting.md +66 -0
- package/knowledge/microservice-testing.md +71 -0
- package/knowledge/model-pricing.json +23 -0
- package/knowledge/mutation-score-formulas.md +60 -0
- package/knowledge/object-calisthenics.md +147 -0
- package/knowledge/oracle-provenance.md +94 -0
- package/knowledge/orchestrator-script-implementation.md +185 -0
- package/knowledge/owasp-detection.md +148 -0
- package/knowledge/plan-review-rubric.md +56 -0
- package/knowledge/proxy-connectivity.md +62 -0
- package/knowledge/reactive-effect-patterns.md +73 -0
- package/knowledge/recon-inventory-excludes.txt +32 -0
- package/knowledge/references/bdd-value-guide.md +61 -0
- package/knowledge/references/csharp-http-client-testing.md +264 -0
- package/knowledge/release-strategies.md +74 -0
- package/knowledge/report-output-location.md +117 -0
- package/knowledge/report-pdf-integration.md +63 -0
- package/knowledge/report-print.css +129 -0
- package/knowledge/report-template.md +114 -0
- package/knowledge/report-to-pdf.md +69 -0
- package/knowledge/request-processing-flow.md +63 -0
- package/knowledge/result-verification.md +52 -0
- package/knowledge/review-agent-output-contract.md +121 -0
- package/knowledge/review-lens-classification.md +113 -0
- package/knowledge/review-rubric.md +62 -0
- package/knowledge/review-template.md +104 -0
- package/knowledge/rule-fixtures/A02.insecure-random-js/negative.js +1 -0
- package/knowledge/rule-fixtures/A02.insecure-random-js/positive.js +1 -0
- package/knowledge/rule-fixtures/A02.weak-hashing-md5/negative.py +1 -0
- package/knowledge/rule-fixtures/A02.weak-hashing-md5/positive.py +1 -0
- package/knowledge/rule-fixtures/A03.command-injection/negative.js +1 -0
- package/knowledge/rule-fixtures/A03.command-injection/positive.js +1 -0
- package/knowledge/rule-fixtures/A03.sql-injection/negative.js +1 -0
- package/knowledge/rule-fixtures/A03.sql-injection/positive.js +1 -0
- package/knowledge/rule-fixtures/A03.xss-innerhtml/negative.js +1 -0
- package/knowledge/rule-fixtures/A03.xss-innerhtml/positive.js +1 -0
- package/knowledge/rule-fixtures/A05.cors-wildcard/negative.js +1 -0
- package/knowledge/rule-fixtures/A05.cors-wildcard/positive.js +1 -0
- package/knowledge/rule-fixtures/A05.default-credentials/negative.js +1 -0
- package/knowledge/rule-fixtures/A05.default-credentials/positive.js +1 -0
- package/knowledge/rule-fixtures/A07.jwt-alg-none/negative.js +1 -0
- package/knowledge/rule-fixtures/A07.jwt-alg-none/positive.js +1 -0
- package/knowledge/rule-fixtures/A08.binary-formatter/negative.cs +1 -0
- package/knowledge/rule-fixtures/A08.binary-formatter/positive.cs +1 -0
- package/knowledge/rule-fixtures/A08.js-eval/negative.js +1 -0
- package/knowledge/rule-fixtures/A08.js-eval/positive.js +1 -0
- package/knowledge/rule-fixtures/A08.object-input-stream/negative.java +1 -0
- package/knowledge/rule-fixtures/A08.object-input-stream/positive.java +1 -0
- package/knowledge/schemas/disposition-register-v1.json +65 -0
- package/knowledge/schemas/recon-envelope-v1.json +198 -0
- package/knowledge/schemas/unified-finding-v1.json +72 -0
- package/knowledge/security-primitives-contract.md +301 -0
- package/knowledge/security-review-rule-map.yaml +107 -0
- package/knowledge/skills-registry.md +72 -0
- package/knowledge/task-size-classifier.md +103 -0
- package/knowledge/telemetry-schema.md +881 -0
- package/knowledge/test-automation-maturity.md +56 -0
- package/knowledge/test-automation-principles.md +71 -0
- package/knowledge/test-cadence-tradeoffs.md +68 -0
- package/knowledge/test-doubles.md +105 -0
- package/knowledge/test-file-indicators.md +22 -0
- package/knowledge/test-layer-gates.md +35 -0
- package/knowledge/test-matrix-examples/django-batch.md +24 -0
- package/knowledge/test-matrix-examples/dotnet-grpc-fronting-api.md +90 -0
- package/knowledge/test-matrix-examples/dotnet-http-consumer.md +131 -0
- package/knowledge/test-matrix-examples/react-node-spa.md +24 -0
- package/knowledge/test-matrix-examples/spring-boot-service.md +25 -0
- package/knowledge/test-matrix-examples/ssr-htmx.md +24 -0
- package/knowledge/test-organization.md +70 -0
- package/knowledge/test-pyramid.md +84 -0
- package/knowledge/test-refactoring.md +67 -0
- package/knowledge/test-review-division-of-labor.md +85 -0
- package/knowledge/test-smells.md +80 -0
- package/knowledge/test-stack-profiles/bdd-frameworks.md +235 -0
- package/knowledge/test-stack-profiles/django.md +13 -0
- package/knowledge/test-stack-profiles/dotnet.md +18 -0
- package/knowledge/test-stack-profiles/go.md +16 -0
- package/knowledge/test-stack-profiles/node.md +16 -0
- package/knowledge/test-stack-profiles/react.md +12 -0
- package/knowledge/test-stack-profiles/spring-boot.md +16 -0
- package/knowledge/test-stack-profiles/ssr-htmx.md +14 -0
- package/knowledge/test-stack-profiles/vue.md +12 -0
- package/knowledge/test-strategy.md +70 -0
- package/knowledge/testability-patterns.md +240 -0
- package/knowledge/testing-quadrants.md +44 -0
- package/knowledge/testing-techniques/approval.md +15 -0
- package/knowledge/testing-techniques/chaos.md +17 -0
- package/knowledge/testing-techniques/fuzz.md +15 -0
- package/knowledge/testing-techniques/property-based.md +15 -0
- package/knowledge/testing-techniques/schema-validation.md +15 -0
- package/knowledge/testing-techniques/screenshot.md +15 -0
- package/knowledge/three-phase-workflow.md +198 -0
- package/knowledge/value-patterns.md +55 -0
- package/knowledge/verification-mode.md +116 -0
- package/knowledge/virtual-service-libraries.md +75 -0
- package/knowledge/wave-consolidation-guidance.md +21 -0
- package/overrides/agents/Explore.md +15 -0
- package/overrides/agents/general-purpose.md +10 -0
- package/overrides/notes/autoship.md +6 -0
- package/overrides/notes/issues-from-assessment.md +3 -0
- package/overrides/notes/issues-from-plan.md +3 -0
- package/overrides/notes/mutation-night-watch.md +3 -0
- package/overrides/notes/mutation-testing.md +3 -0
- package/overrides/notes/pr.md +7 -0
- package/overrides/notes/project-init.md +6 -0
- package/overrides/notes/setup.md +13 -0
- package/overrides/notes/specs.md +3 -0
- package/overrides/skills/headless-run/SKILL.md +45 -0
- package/overrides/skills/upgrade/SKILL.md +30 -0
- package/overrides/skills/version/SKILL.md +25 -0
- package/package.json +36 -0
- package/scripts/authoring_digest.py +93 -0
- package/scripts/autoship_discover.py +121 -0
- package/scripts/autoship_group.py +409 -0
- package/scripts/autoship_proposals.py +494 -0
- package/scripts/autoship_queue.py +291 -0
- package/scripts/autoship_reclaim.py +495 -0
- package/scripts/build_jobs.py +108 -0
- package/scripts/build_rollback_point.py +240 -0
- package/scripts/build_slice_scope.py +157 -0
- package/scripts/build_wave.py +109 -0
- package/scripts/build_wave_reconcile.py +252 -0
- package/scripts/build_worktree_baseref.py +113 -0
- package/scripts/check_agent_scope.py +117 -0
- package/scripts/check_agent_tool_mapping.py +213 -0
- package/scripts/check_review_agent_mcp_tools.py +317 -0
- package/scripts/check_security_assessment_mcp_tools.py +165 -0
- package/scripts/checkpoint_abort.py +502 -0
- package/scripts/claude_setup_review.py +438 -0
- package/scripts/codebase_recon.py +556 -0
- package/scripts/coverage_config.py +623 -0
- package/scripts/coverage_delta_steering.py +330 -0
- package/scripts/coverage_discovery_dotnet.py +315 -0
- package/scripts/coverage_discovery_java.py +742 -0
- package/scripts/coverage_discovery_js.py +546 -0
- package/scripts/coverage_gap_ranking.py +556 -0
- package/scripts/coverage_readiness.py +455 -0
- package/scripts/coverage_report_parse.py +521 -0
- package/scripts/detect_bdd_convention.py +252 -0
- package/scripts/eval_ablation.py +376 -0
- package/scripts/gherkin_analysis_coverage_gate.py +306 -0
- package/scripts/gherkin_cross_feature_duplicate_titles_gate.py +173 -0
- package/scripts/gherkin_effectiveness_rollup.py +238 -0
- package/scripts/gherkin_failure_path_gate.py +206 -0
- package/scripts/gherkin_feature_merge.py +720 -0
- package/scripts/gherkin_stub_gate.py +163 -0
- package/scripts/gherkin_stub_merge.py +479 -0
- package/scripts/git_origin_host.py +88 -0
- package/scripts/install-java-static-analysis.py +110 -0
- package/scripts/issue_deps.py +74 -0
- package/scripts/lib/_bdd_markers.py +28 -0
- package/scripts/lib/_gherkin_text.py +93 -0
- package/scripts/lib/_vendored_tree.py +70 -0
- package/scripts/lib/autoship_state.py +397 -0
- package/scripts/lib/claude_md_guard.py +226 -0
- package/scripts/lib/deterministic_recon.py +446 -0
- package/scripts/lib/mcp_tool_grants.py +211 -0
- package/scripts/lib/plan_parse.py +386 -0
- package/scripts/lib/review_result.py +84 -0
- package/scripts/lib/review_roster.py +86 -0
- package/scripts/lib/session_log/__init__.py +34 -0
- package/scripts/lib/session_log/__pycache__/__init__.cpython-314.pyc +0 -0
- package/scripts/lib/session_log/__pycache__/records.cpython-314.pyc +0 -0
- package/scripts/lib/session_log/classify.py +231 -0
- package/scripts/lib/session_log/corrections.py +194 -0
- package/scripts/lib/session_log/discovery.py +108 -0
- package/scripts/lib/session_log/records.py +218 -0
- package/scripts/lib/session_log/redact.py +76 -0
- package/scripts/lib/session_log/signals.py +373 -0
- package/scripts/lib/session_report_downstream.py +614 -0
- package/scripts/lib/session_report_maintainer.py +1273 -0
- package/scripts/lib/session_report_shared.py +262 -0
- package/scripts/lib/settings_hook_guard.py +157 -0
- package/scripts/lib/slug.py +33 -0
- package/scripts/lib/stub_extractors/__init__.py +82 -0
- package/scripts/lib/stub_extractors/_common.py +328 -0
- package/scripts/lib/stub_extractors/csharp.py +19 -0
- package/scripts/lib/stub_extractors/go.py +173 -0
- package/scripts/lib/stub_extractors/java.py +18 -0
- package/scripts/lib/stub_extractors/jsts.py +126 -0
- package/scripts/mutation_stack_sections.py +149 -0
- package/scripts/mutation_yield_steering.py +345 -0
- package/scripts/orchestrator.py +895 -0
- package/scripts/plan_gherkin_export.py +227 -0
- package/scripts/plan_waves.py +208 -0
- package/scripts/pr_close_keyword_lint.py +108 -0
- package/scripts/progress_guardian.py +888 -0
- package/scripts/recon_inventory.py +273 -0
- package/scripts/review_findings_log.py +93 -0
- package/scripts/run_invariants.py +124 -0
- package/scripts/select_lenses.py +640 -0
- package/scripts/session_report.py +486 -0
- package/scripts/set_autocompact_env.py +221 -0
- package/scripts/ship_resume_guard.py +135 -0
- package/scripts/ship_review_gate.py +63 -0
- package/scripts/specs_convention_marker.py +103 -0
- package/scripts/test_improve_resume.py +277 -0
- package/scripts/test_review_mechanics.py +958 -0
- package/scripts/token_efficiency_review.py +322 -0
- package/scripts/verdict_scope.py +285 -0
- package/scripts/verify_gherkin_quality_critic_isolation.py +296 -0
- package/scripts/verify_tier.py +157 -0
- package/skills/adr-tools/SKILL.md +118 -0
- package/skills/agent-readiness/SKILL.md +105 -0
- package/skills/agent-readiness/ai_friendly_analyzers.py +326 -0
- package/skills/agent-readiness/scanner.py +441 -0
- package/skills/agent-readiness/scorecard.yaml +88 -0
- package/skills/api-design/SKILL.md +115 -0
- package/skills/apply-fixes/SKILL.md +171 -0
- package/skills/apply-test-doubles/SKILL.md +321 -0
- package/skills/artifact-lifecycle/SKILL.md +127 -0
- package/skills/autoship/SKILL.md +1124 -0
- package/skills/benchmark/SKILL.md +105 -0
- package/skills/branch-workflow/SKILL.md +89 -0
- package/skills/browse/SKILL.md +184 -0
- package/skills/browser-testing/SKILL.md +62 -0
- package/skills/browser-testing/references/playwright-patterns.md +216 -0
- package/skills/build/SKILL.md +422 -0
- package/skills/build/references/static-self-heal.md +245 -0
- package/skills/careful/SKILL.md +72 -0
- package/skills/cd-test-architecture/SKILL.md +371 -0
- package/skills/ci-debugging/SKILL.md +105 -0
- package/skills/co-evolution-audit/SKILL.md +269 -0
- package/skills/code-review/SKILL.md +1015 -0
- package/skills/code-review/examples/aggregated-sample.json +56 -0
- package/skills/code-review/examples/sample-report.md +41 -0
- package/skills/code-review/output-format.md +478 -0
- package/skills/code-review/scripts/activation.py +86 -0
- package/skills/code-review/scripts/change_impact.py +357 -0
- package/skills/code-review/scripts/change_shape.py +372 -0
- package/skills/code-review/scripts/change_size.py +212 -0
- package/skills/code-review/scripts/changed_file_list.py +141 -0
- package/skills/code-review/scripts/closing_pass.py +187 -0
- package/skills/code-review/scripts/consolidate.py +277 -0
- package/skills/code-review/scripts/contract_failure_report.py +185 -0
- package/skills/code-review/scripts/dispatch_reconcile.py +66 -0
- package/skills/code-review/scripts/dispatch_waves.py +164 -0
- package/skills/code-review/scripts/finding_signature.py +446 -0
- package/skills/code-review/scripts/ledger.py +283 -0
- package/skills/code-review/scripts/partition.py +169 -0
- package/skills/code-review/scripts/render_tiered_findings.py +274 -0
- package/skills/code-review/scripts/repo_invariants.py +1066 -0
- package/skills/code-review/scripts/review_context_pack.py +306 -0
- package/skills/code-review/scripts/review_round_log.py +345 -0
- package/skills/code-review/scripts/review_value_coverage.py +297 -0
- package/skills/code-review/scripts/validate_review_output.py +467 -0
- package/skills/code-review/sliced-mode.md +205 -0
- package/skills/competitive-analysis/SKILL.md +191 -0
- package/skills/context-loading-protocol/SKILL.md +157 -0
- package/skills/continue/SKILL.md +90 -0
- package/skills/cost-report/SKILL.md +178 -0
- package/skills/coverage-baseline/SKILL.md +335 -0
- package/skills/coverage-baseline/references/multi-project-discovery.md +202 -0
- package/skills/coverage-delta/SKILL.md +181 -0
- package/skills/coverage-delta/references/mutation-gate.md +70 -0
- package/skills/design-doc/SKILL.md +95 -0
- package/skills/design-interrogation/SKILL.md +89 -0
- package/skills/design-it-twice/SKILL.md +91 -0
- package/skills/docker-image-audit/SKILL.md +108 -0
- package/skills/docker-image-audit/references/install-guide.md +64 -0
- package/skills/docker-image-audit/references/report-template.md +73 -0
- package/skills/docker-image-create/SKILL.md +185 -0
- package/skills/domain-analysis/SKILL.md +183 -0
- package/skills/domain-driven-design/SKILL.md +194 -0
- package/skills/exploratory-testing/SKILL.md +108 -0
- package/skills/explore/SKILL.md +51 -0
- package/skills/farley-score/SKILL.md +165 -0
- package/skills/feature-file-validation/SKILL.md +78 -0
- package/skills/feature-file-validation/references/validation-rules.md +115 -0
- package/skills/feedback-learning/SKILL.md +414 -0
- package/skills/fix/SKILL.md +450 -0
- package/skills/freeze/SKILL.md +68 -0
- package/skills/frontend-architecture/SKILL.md +113 -0
- package/skills/gherkin-derive/SKILL.md +630 -0
- package/skills/gherkin-public/SKILL.md +266 -0
- package/skills/governance-compliance/SKILL.md +150 -0
- package/skills/guard/SKILL.md +75 -0
- package/skills/handoff/SKILL.md +139 -0
- package/skills/handoff/references/summary-templates.md +242 -0
- package/skills/harness-audit/SKILL.md +751 -0
- package/skills/harness-audit/scripts/lesson_validate.py +386 -0
- package/skills/harness-audit/scripts/redundancy_criterion.py +188 -0
- package/skills/headless-run/SKILL.md +45 -0
- package/skills/headless-run/scripts/isolated_dispatch.py +381 -0
- package/skills/help/SKILL.md +72 -0
- package/skills/hexagonal-architecture/SKILL.md +85 -0
- package/skills/human-oversight-protocol/SKILL.md +224 -0
- package/skills/issues-from-assessment/SKILL.md +223 -0
- package/skills/issues-from-plan/SKILL.md +133 -0
- package/skills/legacy-code/SKILL.md +132 -0
- package/skills/mermaid-diagramming/SKILL.md +120 -0
- package/skills/mutation-night-watch/SKILL.md +154 -0
- package/skills/mutation-night-watch/references/scheduling.md +135 -0
- package/skills/mutation-testing/SKILL.md +396 -0
- package/skills/mutation-testing/references/languages/csharp-stryker-net.md +676 -0
- package/skills/mutation-testing/references/languages/go-go-mutesting.md +95 -0
- package/skills/mutation-testing/references/languages/java-pitest.md +77 -0
- package/skills/mutation-testing/references/languages/javascript-stryker.md +188 -0
- package/skills/mutation-testing/references/languages/python-mutmut.md +97 -0
- package/skills/mutation-testing/references/time-estimation.md +34 -0
- package/skills/mutation-testing/references/tool-detection.md +15 -0
- package/skills/mutation-testing/references/workflow-callers.md +23 -0
- package/skills/mutation-testing/scripts/__pycache__/xunit_v3_feature_detector.cpython-314.pyc +0 -0
- package/skills/mutation-testing/scripts/csharp_stryker_net_slice_runner.py +635 -0
- package/skills/mutation-testing/scripts/csharp_stryker_net_status_loop.py +525 -0
- package/skills/mutation-testing/scripts/csharp_stryker_net_wrapper.py +681 -0
- package/skills/mutation-testing/scripts/mutation_baseline_reuse.py +292 -0
- package/skills/mutation-testing/scripts/mutation_exclude_policy.py +268 -0
- package/skills/mutation-testing/scripts/mutation_feasibility_gate.py +463 -0
- package/skills/mutation-testing/scripts/mutation_kill_headless.py +331 -0
- package/skills/mutation-testing/scripts/mutation_kill_insert.py +199 -0
- package/skills/mutation-testing/scripts/mutation_kill_insert_python.py +150 -0
- package/skills/mutation-testing/scripts/mutation_kill_loop.py +869 -0
- package/skills/mutation-testing/scripts/mutation_kill_loop_python.py +949 -0
- package/skills/mutation-testing/scripts/mutation_kill_retry.py +592 -0
- package/skills/mutation-testing/scripts/mutation_kill_shared.py +620 -0
- package/skills/mutation-testing/scripts/mutation_nightwatch.py +462 -0
- package/skills/mutation-testing/scripts/mutation_nightwatch_stacks.py +425 -0
- package/skills/mutation-testing/scripts/mutation_report.py +743 -0
- package/skills/mutation-testing/scripts/mutation_report_cli.py +175 -0
- package/skills/mutation-testing/scripts/mutation_safety_gate.py +69 -0
- package/skills/mutation-testing/scripts/stryker_shard_pipeline.py +847 -0
- package/skills/mutation-testing/scripts/stryker_shard_setup.py +440 -0
- package/skills/mutation-testing/scripts/stryker_timeout_retry.py +142 -0
- package/skills/mutation-testing/scripts/xunit_v3_feature_detector.py +341 -0
- package/skills/performance-benchmark/SKILL.md +174 -0
- package/skills/performance-benchmark/examples/report-format.md +43 -0
- package/skills/performance-benchmark/references/benchmark-script.md +169 -0
- package/skills/performance-metrics/SKILL.md +265 -0
- package/skills/plan/SKILL.md +199 -0
- package/skills/plan/references/gherkin-persistence.md +43 -0
- package/skills/plan/references/plan-template.md +182 -0
- package/skills/pr/SKILL.md +289 -0
- package/skills/pr/scripts/gate_retry_state.py +368 -0
- package/skills/project-init/README.md +141 -0
- package/skills/project-init/SKILL.md +1197 -0
- package/skills/project-init/evals/evals.json +200 -0
- package/skills/project-init/references/capability-tools.md +55 -0
- package/skills/project-init/references/configs.md +221 -0
- package/skills/property-based-testing/SKILL.md +121 -0
- package/skills/property-based-testing/fixtures/invariant_fixture.py +15 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/README.md +42 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/fast-check/LICENSE +21 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/fast-check/README.md +263 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/fast-check/lib/cjs/fast-check.d.ts +5165 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/fast-check/lib/cjs/fast-check.js +12147 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/fast-check/lib/cjs/package.json +3 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/fast-check/lib/cjs/types57/fast-check.d.ts +5165 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/fast-check/lib/fast-check.d.ts +5165 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/fast-check/lib/fast-check.js +12011 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/fast-check/lib/rolldown-runtime-D7D4PA-g.js +13 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/fast-check/lib/types57/fast-check.d.ts +5165 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/fast-check/package.json +94 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/LICENSE +21 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/README.md +168 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/RandomGenerator-DcXj09Ch.d.ts +14 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/distribution/uniformBigInt.d.ts +15 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/distribution/uniformBigInt.js +38 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/distribution/uniformFloat32.d.ts +15 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/distribution/uniformFloat32.js +18 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/distribution/uniformFloat64.d.ts +15 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/distribution/uniformFloat64.js +22 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/distribution/uniformInt.d.ts +15 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/distribution/uniformInt.js +134 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/RandomGenerator-DcXj09Ch.d.ts +14 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/distribution/uniformBigInt.d.ts +15 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/distribution/uniformBigInt.js +37 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/distribution/uniformFloat32.d.ts +15 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/distribution/uniformFloat32.js +17 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/distribution/uniformFloat64.d.ts +15 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/distribution/uniformFloat64.js +21 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/distribution/uniformInt.d.ts +15 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/distribution/uniformInt.js +133 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/generator/congruential32.d.ts +7 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/generator/congruential32.js +44 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/generator/mersenne.d.ts +7 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/generator/mersenne.js +90 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/generator/xoroshiro128plus.d.ts +7 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/generator/xoroshiro128plus.js +80 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/generator/xorshift128plus.d.ts +7 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/generator/xorshift128plus.js +78 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/package.json +3 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/types/JumpableRandomGenerator.d.ts +16 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/types/JumpableRandomGenerator.js +0 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/types/RandomGenerator.d.ts +2 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/types/RandomGenerator.js +0 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/utils/generateN.d.ts +6 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/utils/generateN.js +8 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/utils/purify.d.ts +12 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/utils/purify.js +9 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/utils/skipN.d.ts +6 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/utils/skipN.js +6 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/generator/congruential32.d.ts +7 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/generator/congruential32.js +46 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/generator/mersenne.d.ts +7 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/generator/mersenne.js +92 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/generator/xoroshiro128plus.d.ts +7 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/generator/xoroshiro128plus.js +82 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/generator/xorshift128plus.d.ts +7 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/generator/xorshift128plus.js +80 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/types/JumpableRandomGenerator.d.ts +16 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/types/JumpableRandomGenerator.js +0 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/types/RandomGenerator.d.ts +2 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/types/RandomGenerator.js +0 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/utils/generateN.d.ts +6 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/utils/generateN.js +9 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/utils/purify.d.ts +12 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/utils/purify.js +10 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/utils/skipN.d.ts +6 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/utils/skipN.js +7 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/package.json +133 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/package-lock.json +1179 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/package.json +14 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/roundtrip.js +29 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/roundtrip.properties.test.js +16 -0
- package/skills/property-based-testing/fixtures/no_property_fixture.py +10 -0
- package/skills/property-based-testing/fixtures/roundtrip_fixture.py +16 -0
- package/skills/property-based-testing/references/languages/javascript.md +54 -0
- package/skills/property-based-testing/scripts/detect_and_dispatch.py +80 -0
- package/skills/property-based-testing/scripts/hypothesis_scaffold.py +276 -0
- package/skills/proxy-resilience/SKILL.md +84 -0
- package/skills/quality-gate-pipeline/SKILL.md +184 -0
- package/skills/quality-targets-converge/SKILL.md +254 -0
- package/skills/repo-review/SKILL.md +159 -0
- package/skills/report-pdf/SKILL.md +66 -0
- package/skills/review/SKILL.md +47 -0
- package/skills/review-agent/SKILL.md +152 -0
- package/skills/review-summary/SKILL.md +73 -0
- package/skills/run-report/SKILL.md +70 -0
- package/skills/semantic-duplication-scan/SKILL.md +337 -0
- package/skills/semantic-scan/SKILL.md +53 -0
- package/skills/semgrep-analyze/SKILL.md +139 -0
- package/skills/setup/SKILL.md +1122 -0
- package/skills/ship/SKILL.md +240 -0
- package/skills/source-verification/SKILL.md +210 -0
- package/skills/source-verification/scripts/claim_extractor.py +155 -0
- package/skills/specs/.size-baseline.json +4 -0
- package/skills/specs/SKILL.md +243 -0
- package/skills/specs/references/completeness-checklist.md +83 -0
- package/skills/specs/references/extraction.md +58 -0
- package/skills/specs/references/glossary.md +59 -0
- package/skills/specs/references/persistence.md +115 -0
- package/skills/specs/references/predictability-check.md +77 -0
- package/skills/static-analysis-integration/SKILL.md +235 -0
- package/skills/static-analysis-integration/adapters/_envelope.py +26 -0
- package/skills/static-analysis-integration/adapters/jscpd-adapter.py +66 -0
- package/skills/static-analysis-integration/adapters/lizard-adapter.py +81 -0
- package/skills/static-analysis-integration/adapters/mypy-adapter.py +50 -0
- package/skills/static-analysis-integration/adapters/mypy-src-layout.py +93 -0
- package/skills/static-analysis-integration/adapters/security-review-adapter.py +212 -0
- package/skills/static-analysis-integration/maintenance.md +23 -0
- package/skills/static-analysis-integration/references/language-setup.md +228 -0
- package/skills/static-analysis-integration/references/sarif-parser.md +124 -0
- package/skills/static-analysis-integration/references/security-review-adapter.md +118 -0
- package/skills/static-analysis-integration/references/tool-configs.md +617 -0
- package/skills/static-analysis-integration/rulesets/pmd-quickstart.xml +24 -0
- package/skills/stryker-xunit-v2-shim/SKILL.md +274 -0
- package/skills/stryker-xunit-v2-shim/references/shim-howto.md +256 -0
- package/skills/stryker-xunit-v2-shim/scripts/generate_shim.py +143 -0
- package/skills/systematic-debugging/SKILL.md +130 -0
- package/skills/telemetry/SKILL.md +75 -0
- package/skills/test-audit-disable/SKILL.md +129 -0
- package/skills/test-design/SKILL.md +177 -0
- package/skills/test-design/scripts/__pycache__/internal_double_detector.cpython-314.pyc +0 -0
- package/skills/test-design/scripts/internal_double_detector.py +631 -0
- package/skills/test-design-advisor/SKILL.md +166 -0
- package/skills/test-driven-development/SKILL.md +169 -0
- package/skills/test-health/SKILL.md +262 -0
- package/skills/test-improve/SKILL.md +239 -0
- package/skills/test-improve/references/phase-0-approach-contract.md +228 -0
- package/skills/test-improve/references/phase-1-analyze.md +131 -0
- package/skills/test-improve/references/phase-2-baseline.md +121 -0
- package/skills/test-improve/references/phase-3-derive-gherkin.md +53 -0
- package/skills/test-improve/references/phase-4-plan-fixes.md +34 -0
- package/skills/test-improve/references/phase-5-improve.md +215 -0
- package/skills/test-improve/references/phase-6-refactor-decision.md +45 -0
- package/skills/test-improve/references/phase-7-refactor.md +44 -0
- package/skills/test-improve/references/phase-8-validate.md +66 -0
- package/skills/test-improve/references/phase-9-close-out-prompt.md +11 -0
- package/skills/test-improve/references/phase-9-report.md +62 -0
- package/skills/test-improve/references/review-loop.md +92 -0
- package/skills/test-improve/templates/executive-summary.md +123 -0
- package/skills/threat-modeling/SKILL.md +108 -0
- package/skills/triage/SKILL.md +211 -0
- package/skills/ubiquitous-language/SKILL.md +192 -0
- package/skills/ubiquitous-language/scripts/collect_domain_signals.py +300 -0
- package/skills/unfreeze/SKILL.md +37 -0
- package/skills/upgrade/SKILL.md +31 -0
- package/skills/upgrade/scripts/check_version_drift.py +113 -0
- package/skills/upgrade/scripts/enable_autoupdate.py +149 -0
- package/skills/version/SKILL.md +25 -0
- package/sync/__pycache__/sync_upstream.cpython-314.pyc +0 -0
- package/sync/sync_upstream.py +293 -0
- package/templates/ACCEPTED-RISKS.md.tmpl +46 -0
- package/templates/agents/agent-template.md +151 -0
- package/templates/agents/angular-testing.md +66 -0
- package/templates/agents/csharp-quality.md +63 -0
- package/templates/agents/esm-enforcer.md +52 -0
- package/templates/agents/front-end-testing.md +65 -0
- package/templates/agents/go-quality.md +65 -0
- package/templates/agents/python-quality.md +62 -0
- package/templates/agents/react-testing.md +61 -0
- package/templates/agents/ts-enforcer.md +60 -0
- package/templates/agents/twelve-factor-audit.md +49 -0
- package/tools/entropy-check.py +250 -0
- package/tools/model-hash-verify.py +213 -0
|
@@ -0,0 +1,166 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: test-design-advisor
|
|
3
|
+
description: Worker skill — assess testability, recommend the right test-pyramid layer and test-double strategy, and propose a behavior-preserving refactor sequence to make hard-to-test code testable. Invoked by /test-design and /test-health; not user-invocable. For a user-facing entry point use /test-design (add --advise --path <file> for forward-design on a single unit).
|
|
4
|
+
role: worker
|
|
5
|
+
---
|
|
6
|
+
|
|
7
|
+
# Test Design Advisor
|
|
8
|
+
|
|
9
|
+
## Overview
|
|
10
|
+
|
|
11
|
+
An **advisory** skill: it recommends how to test code and how to make untestable code testable. It does not write tests or refactor code — it produces a design the human (or `/build`) then implements. Use it before writing a test suite for an untested or hard-to-test module, or when a test is hard to write and you suspect the design is the cause.
|
|
12
|
+
|
|
13
|
+
Grounded in these knowledge references: `knowledge/test-smells.md`, `knowledge/test-doubles.md`, `knowledge/test-pyramid.md` (layers + shapes), `knowledge/test-layer-gates.md` for behavior pre-gates, `knowledge/microservice-testing.md`, `knowledge/testability-patterns.md` for production-code seams, and `knowledge/test-strategy.md` for fixture and SUT-interaction strategy. For the xUnit pattern families it grounds in `knowledge/fixture-construction.md` (how the fixture is built/disposed), `knowledge/value-patterns.md` (Literal/Derived/Generated Value + Dummy Object for test data), `knowledge/result-verification.md` (assertion patterns), `knowledge/test-organization.md` (Four-Phase + suite structure), `knowledge/test-refactoring.md` (goals/principles + the test-side refactoring catalog), and `knowledge/test-automation-principles.md` (the goals/principles a recommendation must honor — the rubric behind every "why"). When the target is untested legacy code, it grounds the get-under-test-first sequence in `knowledge/dependency-breaking-techniques.md` (behavior-preserving seams) and `knowledge/legacy-test-strategy.md` (effect/pinch reasoning for where to place the tests). On match it overlays `knowledge/testing-techniques/` (specialized techniques), resolves tools from `knowledge/test-stack-profiles/<stack>.md`, and may adapt a worked template from `knowledge/test-matrix-examples/`.
|
|
14
|
+
|
|
15
|
+
## Constraints
|
|
16
|
+
|
|
17
|
+
- Advisory only. Do not edit production code or write test files — output a recommendation.
|
|
18
|
+
- A hard-to-test design is a production-code problem. Recommend the seam (constructor injection, interface extraction), never a test workaround (reflection, `InternalsVisibleTo`, mocking concrete classes).
|
|
19
|
+
- Prefer the lowest test-pyramid layer that can verify the behavior; prefer state verification and the simplest double.
|
|
20
|
+
- Refactor sequences must be behavior-preserving and start with characterization tests when the code is currently untested.
|
|
21
|
+
- Be concise: tables and ordered steps, not prose. No restating the source material — cite the knowledge file.
|
|
22
|
+
- **Altitude boundary.** When a gate mandates application-level E2E/browser architecture, flag the seam (`→ cd-test-architecture`) and defer the harness/pipeline design to the `cd-test-architecture` skill — do not design the E2E harness here.
|
|
23
|
+
- **Vocabulary (MinimumCD).** Use the six MinimumCD test types defined in `knowledge/cd-test-architecture.md` § *The Six Test Types*: **static analysis / unit / component / contract / integration / E2E**. **"Contract test"** is the primary term — use it. When the codebase or external context calls it a **"narrow integration test"**, gloss it once in the same sentence: `contract test (also called narrow integration test)`. Never use "narrow integration" alone. When the codebase uses different names for layers (e.g. "service test", "API test", "scenario test"), emit a **Terminology mapping** table at the top of the report and then use the MinimumCD term consistently from that point.
|
|
24
|
+
- **Test type definitions.** When the report mentions any of {unit, component, contract, integration, E2E, static analysis, sociable unit, solitary unit, resilience}, define each term on first use — either inline as a one-line gloss or in a small **Test type definitions used in this report** block at the top. Definitions come verbatim from `knowledge/cd-test-architecture.md` § *The Six Test Types*.
|
|
25
|
+
- **The pyramid is a cost heuristic, not a target shape** — canonical rule in `knowledge/cd-test-architecture.md#the-pyramid-is-a-cost-heuristic-not-a-target-shape`. No "current shape vs recommended shape" tables and no per-layer target counts; place each behavior at the lowest layer that can verify it.
|
|
26
|
+
- **E2E justification gate** — canonical in `knowledge/cd-test-architecture.md#the-e2e-justification-gate`. Recommend E2E only when all four conditions hold, and surface the four-condition verdict in the E2E justification table (Output).
|
|
27
|
+
|
|
28
|
+
## Parse Arguments
|
|
29
|
+
|
|
30
|
+
Arguments: target file(s), module, or a description of the code to test. If no target is given, ask for one. Detect language **and framework/stack** (from manifests — `package.json`, `build.gradle`/`pom.xml`, `*.csproj`, `go.mod`, `requirements.txt`/`pyproject.toml`, and frontend deps like react/vue/htmx) so tool resolution can pick the right `test-stack-profiles/<stack>.md`. Detect whether the target crosses independently-deployable service boundaries (load `microservice-testing.md` only if so).
|
|
31
|
+
|
|
32
|
+
## Steps
|
|
33
|
+
|
|
34
|
+
### 1. Assess testability
|
|
35
|
+
|
|
36
|
+
Read the target. For each unit, determine whether it can be constructed and driven through its public API with controlled inputs. Use the decision flow in `knowledge/testability-patterns.md`. Record blockers: static factories/singletons, new-ed-up dependencies, hidden global/clock/RNG access, concrete-class coupling, private logic with no public path.
|
|
37
|
+
|
|
38
|
+
**Graph-assisted testability assessment.** Prefer CodeGraph/Repowise over raw `Grep`/`Read` for finding the unit's collaborators, its callers, and blast radius before recommending a seam — see [`knowledge/codegraph-vs-graphify.md`](../../knowledge/codegraph-vs-graphify.md) for tool selection and the fallback contract.
|
|
39
|
+
|
|
40
|
+
### 1b. Behavior pre-gates (escalate the layer by failure mode)
|
|
41
|
+
|
|
42
|
+
Before pyramid placement, run the gates in `knowledge/test-layer-gates.md`. They **escalate upward only** — never lower the Step 2 pick; silent when none fire:
|
|
43
|
+
|
|
44
|
+
- **Gate A** user-facing dynamic → E2E alongside lower layers (state cost + amortization)
|
|
45
|
+
- **Gate B** bug fix → regression at the discovery layer (no escalation above it)
|
|
46
|
+
- **Gate C** HTMX/Alpine/Turbo swap → browser test REQUIRED for the seam
|
|
47
|
+
- **Gate D** visual artifact → approval/screenshot (with a reference) or manual
|
|
48
|
+
|
|
49
|
+
If dynamic-ness is ambiguous, state your assumption and ask **once** (batch ambiguities; offer "treat all as dynamic") — never escalate silently. When a gate mandates app-level E2E, flag it `→ cd-test-architecture` and defer the harness design there.
|
|
50
|
+
|
|
51
|
+
### 2. Place each behavior on the pyramid
|
|
52
|
+
|
|
53
|
+
Using `knowledge/test-pyramid.md`, assign each behavior to the lowest layer that can meaningfully verify it (unit / component / contract / integration / E2E — MinimumCD names per the Vocabulary constraint). Flag anything currently mis-layered. For service boundaries, apply contract testing per `knowledge/microservice-testing.md` instead of E2E. For any behavior placed at integration or E2E, the E2E justification gate (Constraints) MUST be satisfied and the four-condition verdict surfaced in the report.
|
|
54
|
+
|
|
55
|
+
**Two-direction justification.** The placement table's *Why this layer* column carries a two-direction justification: when the pick is unit or component, explain why a higher layer would be redundant or cupcake-shaped duplication; when the pick is integration or E2E, explain why a contract or component test cannot cover the behavior. The advisor must articulate the trade-off rather than pattern-match a layer.
|
|
56
|
+
|
|
57
|
+
**No target counts.** Do NOT emit a "current shape vs recommended shape" table or any per-layer target count. Per-behavior placement is the only valid output for layer recommendations (see the Constraints: *The pyramid is a cost heuristic, not a target shape*).
|
|
58
|
+
|
|
59
|
+
**Redundancy check (business-critical only).** After placement, for any behavior determined business-critical (labelled, or confirm by asking), if it is covered at only one layer, flag it and name a second layer with a different failure mode (catches/misses table in `test-layer-gates.md`) plus a concrete recommendation.
|
|
60
|
+
|
|
61
|
+
**Gate-column output schema.** In the *Pyramid placement* table the `Gate` column uses: `—` (no gate), `↑<layer>` (escalated), `→ cd-test-architecture` (E2E architecture deferred). When multiple gates fire, union the layers (no duplicate) and list each gate's cost once, with reconciliation guidance if the amortization advice differs.
|
|
62
|
+
|
|
63
|
+
**Tool resolution (by detected stack).** After the layer is fixed, name the concrete tool. If the detected stack has a profile in `knowledge/test-stack-profiles/<stack>.md`, read it and fill the `Tool` column from its layer→tool map (e.g. Spring integration → MockMvc; SSR frontend → JSDOM+MSW, **not** a browser unit runner). If no profile matches, proceed with stack-agnostic guidance and **name the missing profile** in the report — never block on it.
|
|
64
|
+
|
|
65
|
+
**Few-shot templates.** When building a multi-behavior matrix for a common stack, you *may* load a worked example from `knowledge/test-matrix-examples/` as a template and adapt its rows — never copy it verbatim.
|
|
66
|
+
|
|
67
|
+
### 2b. Specialized technique overlay (on trigger match)
|
|
68
|
+
|
|
69
|
+
Some behaviors break in a way no pyramid layer addresses. Consult `knowledge/testing-techniques/` and add a technique **only when its trigger matches** — silent otherwise. Each overlay note states the technique and its maintenance cost; it complements (does not replace) the layer pick.
|
|
70
|
+
|
|
71
|
+
| Trigger | Overlay | File |
|
|
72
|
+
|---------|---------|------|
|
|
73
|
+
| Invariant/law holds for all inputs | property-based | `property-based.md` |
|
|
74
|
+
| Parses untrusted/unstructured input | fuzz | `fuzz.md` |
|
|
75
|
+
| Large text/structured artifact vs a reference | approval | `approval.md` |
|
|
76
|
+
| Visual/CSS/layout fidelity matters | screenshot | `screenshot.md` |
|
|
77
|
+
| Payload governed by a declared schema (OpenAPI/JSON Schema/Avro) | schema-validation | `schema-validation.md` |
|
|
78
|
+
| Resilience claim under dependency failure | chaos | `chaos.md` |
|
|
79
|
+
|
|
80
|
+
**Exclusion.** Do **not** add a contract-testing technique here. Consumer↔provider agreement across an owned service boundary routes to `knowledge/microservice-testing.md` (CDC) — never double-route.
|
|
81
|
+
|
|
82
|
+
### 3. Choose doubles
|
|
83
|
+
|
|
84
|
+
For each collaborator at each test, recommend the simplest double using the decision flow in `knowledge/test-doubles.md` (dummy/stub/spy/mock/fake) and whether to verify by state or behavior. Default to state verification + stub/fake; reserve mock/spy for true side-effect boundaries.
|
|
85
|
+
|
|
86
|
+
### 3b. Recommend fixture and interaction strategy
|
|
87
|
+
|
|
88
|
+
**Under `/test-design`**, the smell agent has already named the smell and its remedy family (via `remedyFamily`, per `knowledge/test-review-division-of-labor.md` § "test-smell-review ↔ test-design-advisor — remedy division"); this step supplies the specific remedy pattern from that family and its per-behavior application. When invoked standalone, this step supplies both the smell framing and the pattern.
|
|
89
|
+
|
|
90
|
+
Using `knowledge/test-strategy.md`, recommend per test group: fixture design + lifecycle (default Minimal + Fresh; escalate to Immutable Shared → Shared only under measured speed pressure), how the test is driven (scripted default; data-driven when variation is purely data), and SUT interaction (front-door by default; Layer Test for layered code; Back Door Manipulation only when the front door obscures intent). Flag any reliance on a mutable Shared Fixture as an Interacting-Tests risk.
|
|
91
|
+
|
|
92
|
+
For the construction **mechanics** that realize that strategy, recommend a specific pattern from `knowledge/fixture-construction.md` — Creation Method / Test Data Builder / Object Mother for building, the right setup location, and **Automated Teardown** for persistent fixtures — to fix fixture smells (Mystery Guest, General Fixture, Irrelevant Information, Test Code Duplication) at the root rather than only naming them.
|
|
93
|
+
|
|
94
|
+
### 3c. Recommend verification and test structure
|
|
95
|
+
|
|
96
|
+
**Under `/test-design`**, the smell agent has already named the smell and its remedy family (via `remedyFamily`, per `knowledge/test-review-division-of-labor.md` § "test-smell-review ↔ test-design-advisor — remedy division"); this step supplies the specific remedy pattern from that family (result-verification and test-organization patterns) and its per-behavior application.
|
|
97
|
+
|
|
98
|
+
Using `knowledge/result-verification.md`, recommend the assertion pattern that fixes verification smells: **Expected Object** for field-by-field clutter, **Custom Assertion / Verification Method** for repeated complex comparisons or poor failure messages, **Guard Assertion** before a precondition-dependent assertion, **Delta Assertion** against a baseline for shared/persistent fixtures — and verify one logical condition per test.
|
|
99
|
+
|
|
100
|
+
Using `knowledge/test-organization.md`, recommend test/suite structure: make the **Four-Phase Test** (Setup → Exercise → Verify → Teardown) visible, group **Testcase Class per Class / Feature / Fixture** when setups diverge, share via a **Test Utility Method / Helper** (composition) over a **Testcase Superclass** (inheritance), and collapse data-only duplication into a **Parameterized Test**.
|
|
101
|
+
|
|
102
|
+
### 4. Propose a behavior-preserving refactor sequence (only if blockers or test smells exist)
|
|
103
|
+
|
|
104
|
+
If Step 1 found blockers in **production** code, produce an ordered sequence that makes it testable without changing behavior:
|
|
105
|
+
|
|
106
|
+
1. Add characterization tests around current behavior (if untested) — pin existing behavior first.
|
|
107
|
+
2. Introduce the seam (the specific pattern from `testability-patterns.md`).
|
|
108
|
+
3. Write the now-possible tests at the layer from Step 2.
|
|
109
|
+
4. Refactor under green.
|
|
110
|
+
|
|
111
|
+
When the smell is in the **test** itself (not production code), name the specific behavior-preserving move from `knowledge/test-refactoring.md` that removes the smell and shifts the test toward the violated goal/principle — e.g. *Inline Mystery Guest* → Fresh Fixture via Creation Method, *Replace General Fixture with Minimal Fixture*, *Introduce Expected Object*, *Extract Custom Assertion*, *Split Test*. Test refactorings are behavior-preserving and **characterization-first** when the target is untested.
|
|
112
|
+
|
|
113
|
+
Each step names the pattern and the exact change required.
|
|
114
|
+
|
|
115
|
+
### 5. Report
|
|
116
|
+
|
|
117
|
+
Write the recommendation (see Output). Keep it actionable — every recommendation maps to a concrete next edit.
|
|
118
|
+
|
|
119
|
+
## Output
|
|
120
|
+
|
|
121
|
+
A concise advisory report (to chat for a single unit, or to `.dev-team-reports/test-design-<target>.md` for a module):
|
|
122
|
+
|
|
123
|
+
```markdown
|
|
124
|
+
## Test Design — <target>
|
|
125
|
+
|
|
126
|
+
### Test type definitions used in this report
|
|
127
|
+
<one-line glosses for every MinimumCD term used below; verbatim from
|
|
128
|
+
`knowledge/cd-test-architecture.md` § The Six Test Types>
|
|
129
|
+
|
|
130
|
+
### Terminology mapping (only if the codebase uses non-MinimumCD names)
|
|
131
|
+
| Local name | MinimumCD term |
|
|
132
|
+
|
|
133
|
+
### Testability
|
|
134
|
+
| Unit | Testable as-is? | Blocker | Seam (testability-patterns.md) |
|
|
135
|
+
|
|
136
|
+
### Pyramid placement
|
|
137
|
+
| Behavior | Layer | Gate | Tool | Why this layer (not the one above or below) |
|
|
138
|
+
|
|
139
|
+
(`Tool` from the stack profile, or `—` / "no profile: <stack>" when none matches.
|
|
140
|
+
*Why this layer* MUST be two-direction: unit/component picks justify why higher
|
|
141
|
+
would be redundant; integration/E2E picks justify why contract/component
|
|
142
|
+
cannot cover the behavior.)
|
|
143
|
+
|
|
144
|
+
### E2E justification (only when a behavior is placed at E2E)
|
|
145
|
+
| Behavior | (1) Contract test ruled out — why | (2) Component test ruled out — why | (3) Resilience test ruled out — why | (4) User journey + multi-component rationale | Pipeline stage |
|
|
146
|
+
|
|
147
|
+
### Technique overlay (only if a trigger fired)
|
|
148
|
+
| Behavior | Technique | Cost note |
|
|
149
|
+
|
|
150
|
+
### Double strategy
|
|
151
|
+
| Test | Collaborator | Double | Verify by |
|
|
152
|
+
|
|
153
|
+
### Refactor sequence (if blockers)
|
|
154
|
+
1. <characterization tests> → 2. <seam> → 3. <tests> → 4. <refactor>
|
|
155
|
+
|
|
156
|
+
### Next edit
|
|
157
|
+
<the single concrete first action>
|
|
158
|
+
```
|
|
159
|
+
|
|
160
|
+
**Do NOT emit a "current shape vs recommended shape" table or any per-layer target count.** The pyramid is a cost heuristic; per-behavior placement is the only valid layer output. See Constraints.
|
|
161
|
+
|
|
162
|
+
## Integration
|
|
163
|
+
|
|
164
|
+
- Pairs with the `test-smell-review` agent (which *detects* smells) and the `farley-score` skill (which *scores* an existing suite). This skill *designs* tests forward.
|
|
165
|
+
- For application-level test architecture (CD pipeline alignment, deterministic config-free CI gate, per-component UI/service/batch patterns), defer to the `cd-test-architecture` skill and its knowledge files (`cd-test-architecture.md`, `component-test-patterns.md`). This skill stays at unit/module altitude.
|
|
166
|
+
- Hand the refactor sequence to `/plan` or `/build` for TDD implementation. This skill stops at the design.
|
|
@@ -0,0 +1,169 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: test-driven-development
|
|
3
|
+
description: Advisory reference for the Classic RED-GREEN-REFACTOR TDD discipline with hard gates — not a build cadence toggle. The plugin's single build cadence is Code-First Small Batches (docs/experiments/RECOMMENDATIONS.md Rec 3); /build does not dispatch into this skill. Use on explicit user request when someone wants test-first discipline for the code being written, or when reviewing code to verify TDD discipline was followed by hand.
|
|
4
|
+
role: worker
|
|
5
|
+
user-invocable: true
|
|
6
|
+
---
|
|
7
|
+
|
|
8
|
+
# Test-Driven Development
|
|
9
|
+
|
|
10
|
+
## Overview
|
|
11
|
+
|
|
12
|
+
Enforces strict RED-GREEN-REFACTOR discipline with verifiable gates. LLMs are especially prone to skipping tests or writing them after implementation — this skill exists because that tendency produces code that looks tested but isn't actually validated.
|
|
13
|
+
|
|
14
|
+
**Positioning:** this is an **advisory methodology reference**, not a build cadence. `/build` runs a single cadence — Code-First Small Batches: implement one behavior, write its test in the same cycle, refactor on every green (`docs/experiments/RECOMMENDATIONS.md` Rec 3: test-first ordering by itself buys nothing; the per-green refactoring is the mechanism, and it is mandatory). `/build` never dispatches into this skill automatically. Use it only when a human explicitly asks for test-first discipline outside of `/build`, or when auditing existing code for TDD discipline after the fact. When applied, every rule below still holds — including the Iron Law and its hard gates.
|
|
15
|
+
|
|
16
|
+
**Defect fixes are not covered by this advisory status.** General feature development under `/build` uses Code-First Small Batches and never requires the full RED-GREEN-REFACTOR cycle in this skill. Fixing a bug is different and is **mandatory, not optional**: always reproduce the defect with a failing test before writing the fix. That hard gate lives in `skills/systematic-debugging/SKILL.md` Phase 4, not here — this skill remains advisory-only for new-feature construction.
|
|
17
|
+
|
|
18
|
+
## Iron Law
|
|
19
|
+
|
|
20
|
+
**No production code without a failing test first.** If you didn't watch the test fail, you don't know if it tests the right thing. Code written before tests must be deleted and reimplemented from the test — no exceptions.
|
|
21
|
+
|
|
22
|
+
## Constraints
|
|
23
|
+
- Do not write implementation code without a failing test first
|
|
24
|
+
- Do not move to the next unit of work until all tests pass
|
|
25
|
+
- Do not skip the refactor step — it's where design quality happens
|
|
26
|
+
- Do not rationalize exceptions to the cycle (see Rationalization Prevention below)
|
|
27
|
+
- Before doubling any collaborator, check it against `${CLAUDE_PLUGIN_ROOT}/knowledge/internal-collaborator-doubling.md#the-three-blockers-exhaustive`'s blocker table
|
|
28
|
+
|
|
29
|
+
## The Cycle
|
|
30
|
+
|
|
31
|
+
Each unit of work follows three phases with hard gates between them:
|
|
32
|
+
|
|
33
|
+
### 1. RED — Write a failing test
|
|
34
|
+
- Write the smallest test that describes the next behavior
|
|
35
|
+
- Use real code, not mocks, whenever avoidable
|
|
36
|
+
- Run the test suite — **the new test must fail**
|
|
37
|
+
- **Hard gate**: paste the failing test output. No output = no proceeding.
|
|
38
|
+
- Verify the failure is for the expected reason (missing feature, not a typo or import error)
|
|
39
|
+
- If the test passes without new code, the behavior already exists — pick a different test
|
|
40
|
+
|
|
41
|
+
### 2. GREEN — Make it pass
|
|
42
|
+
- Write the minimum implementation to make the failing test pass
|
|
43
|
+
- Run the test suite — **all tests must pass** with no errors or warnings
|
|
44
|
+
- **Hard gate**: paste the passing test output. No output = no proceeding.
|
|
45
|
+
- Do not add behavior beyond what the test requires
|
|
46
|
+
- Do not refactor yet
|
|
47
|
+
|
|
48
|
+
### 3. REFACTOR — Review then clean up
|
|
49
|
+
|
|
50
|
+
After GREEN, dispatch review agents **before** making any structural changes.
|
|
51
|
+
|
|
52
|
+
#### 3a. Post-GREEN review
|
|
53
|
+
|
|
54
|
+
1. **Always** — run `refactor-opportunity-review` on all changed files
|
|
55
|
+
2. **When non-trivial** (any changed function exceeds 20 lines, nesting exceeds 3 levels, or more than one function was modified) — also run `structure-review` on the changed files
|
|
56
|
+
3. Collect all `error` and `warning` findings from both agents — these are the refactor candidates for this cycle
|
|
57
|
+
|
|
58
|
+
#### 3b. Bounded fix loop (max 3 iterations)
|
|
59
|
+
|
|
60
|
+
Repeat until no `error`/`warning` findings remain, or 3 iterations have elapsed:
|
|
61
|
+
|
|
62
|
+
1. Apply the highest-severity finding — smallest behavior-preserving change possible
|
|
63
|
+
2. Run the test suite — **all tests must still pass**
|
|
64
|
+
3. If tests break: undo the change immediately; mark the finding as untouchable; move to the next finding
|
|
65
|
+
4. If tests pass: accept the change; move to the next finding
|
|
66
|
+
|
|
67
|
+
After 3 iterations, stop. Remaining findings are noted for the next cycle or escalated to the human.
|
|
68
|
+
|
|
69
|
+
#### 3c. Verify green
|
|
70
|
+
|
|
71
|
+
Run the full test suite one final time after the fix loop:
|
|
72
|
+
|
|
73
|
+
- **All tests must still pass** before leaving REFACTOR
|
|
74
|
+
- Do not proceed to the next RED cycle with any test red
|
|
75
|
+
|
|
76
|
+
Then return to RED for the next behavior.
|
|
77
|
+
|
|
78
|
+
## Rationalization Prevention
|
|
79
|
+
|
|
80
|
+
LLMs generate plausible excuses for skipping TDD. These are the common ones and why they're wrong:
|
|
81
|
+
|
|
82
|
+
| Excuse | Reality |
|
|
83
|
+
|--------|---------|
|
|
84
|
+
| "I'll add tests after the implementation" | You won't. And if you do, you'll write tests that pass by definition — they test what you wrote, not what should work. |
|
|
85
|
+
| "This is too simple to test" | Simple code breaks too. Testing takes 30 seconds. The one-line change that caused the most expensive bug looked simple too. |
|
|
86
|
+
| "Writing the test first would be slower" | TDD is faster than debugging. It catches errors at the cheapest possible moment. |
|
|
87
|
+
| "I need to see the implementation shape first" | That's called a spike. Do the spike, throw it away, then TDD the real implementation. |
|
|
88
|
+
| "The test framework isn't set up yet" | Set it up. That's the first task, not a reason to skip testing. |
|
|
89
|
+
| "I'm just refactoring, not adding behavior" | Then existing tests should pass throughout. If there are no existing tests, write characterization tests first. |
|
|
90
|
+
| "This is glue code / config / boilerplate" | Glue code that breaks takes down the system. If it can break, it needs a test. |
|
|
91
|
+
| "I already tested it manually" | Manual testing lacks systematic, re-runnable verification. It doesn't cover edge cases and you re-test every change. |
|
|
92
|
+
| "Deleting my existing code is wasteful" | Sunk cost fallacy. Unverified code is technical debt, not an asset. |
|
|
93
|
+
| "Let me keep my code as a reference and write tests first" | You'll adapt it instead of TDD-ing. That becomes testing-after with extra steps. |
|
|
94
|
+
| "The test is hard to write — I'll come back to it" | Hard-to-test code is hard-to-use code. The test is telling you the design needs work. Listen to it. |
|
|
95
|
+
| "TDD slows me down / I'm being pragmatic" | TDD is the pragmatic choice. Truly pragmatic means test-first because debugging costs more than testing. |
|
|
96
|
+
|
|
97
|
+
If you catch yourself composing an excuse not on this list, it's still an excuse. Write the test first.
|
|
98
|
+
|
|
99
|
+
## Red Flags Requiring Restart
|
|
100
|
+
|
|
101
|
+
Stop immediately and restart from RED if you notice:
|
|
102
|
+
- Writing implementation code before tests
|
|
103
|
+
- Adding tests after implementation
|
|
104
|
+
- Tests passing immediately without new implementation (testing existing behavior)
|
|
105
|
+
- Tests deferred to "later"
|
|
106
|
+
- Any rationalization beginning with "just this once"
|
|
107
|
+
- Manual testing claims replacing automated verification
|
|
108
|
+
- "Keep as reference" or "adapt existing code" language
|
|
109
|
+
- Sunk cost justifications for keeping pre-test code
|
|
110
|
+
|
|
111
|
+
**Response**: Delete the code written without tests. Start over with RED.
|
|
112
|
+
|
|
113
|
+
**A failing test you can't explain is a debugging task, not a restart.** Do not delete code to escape an unexplained failure. Enter [Systematic Debugging](../systematic-debugging/SKILL.md) (reproduce → root cause); only once you understand *why* it failed do you decide whether the fix is code or a restart from RED.
|
|
114
|
+
|
|
115
|
+
## Verification Checklist
|
|
116
|
+
|
|
117
|
+
Before completing a unit of work:
|
|
118
|
+
- [ ] Every new function/method has a test
|
|
119
|
+
- [ ] Each test was watched failing before implementation
|
|
120
|
+
- [ ] Each failure occurred for the expected reason (missing feature, not typo)
|
|
121
|
+
- [ ] Minimal code written to pass each test
|
|
122
|
+
- [ ] The whole suite passing with clean output (no errors, no warnings) — a red test anywhere is failure, not just the ones this change touched; "pre-existing / not my diff" does not clear it
|
|
123
|
+
- [ ] Tests use real code (mocks only when unavoidable)
|
|
124
|
+
- [ ] Edge cases and error conditions covered
|
|
125
|
+
|
|
126
|
+
Missing any checkbox = TDD was skipped. Restart from RED.
|
|
127
|
+
|
|
128
|
+
## Exception Permissions
|
|
129
|
+
|
|
130
|
+
Ask your human partner before skipping TDD for:
|
|
131
|
+
- Throwaway prototypes (spike-and-discard)
|
|
132
|
+
- Generated code (scaffolding tools, codegen output)
|
|
133
|
+
- Configuration files with no behavioral logic
|
|
134
|
+
|
|
135
|
+
Even with permission, document the exception.
|
|
136
|
+
|
|
137
|
+
## Anti-Pattern: Horizontal Slicing
|
|
138
|
+
|
|
139
|
+
**DO NOT write all tests first, then all implementation.** This is "horizontal slicing" — treating RED as "write all tests" and GREEN as "write all code."
|
|
140
|
+
|
|
141
|
+
This produces bad tests:
|
|
142
|
+
- Tests written in bulk test *imagined* behavior, not *actual* behavior
|
|
143
|
+
- You end up testing the *shape* of things (data structures, signatures) rather than user-facing behavior
|
|
144
|
+
- Tests become insensitive to real changes — passing when behavior breaks, failing when behavior is fine
|
|
145
|
+
- You outrun your headlights, committing to test structure before understanding the implementation
|
|
146
|
+
|
|
147
|
+
**Correct approach: vertical slices via tracer bullets.** One test → one implementation → repeat. Each test responds to what you learned from the previous cycle.
|
|
148
|
+
|
|
149
|
+
```
|
|
150
|
+
WRONG (horizontal):
|
|
151
|
+
RED: test1, test2, test3, test4, test5
|
|
152
|
+
GREEN: impl1, impl2, impl3, impl4, impl5
|
|
153
|
+
|
|
154
|
+
RIGHT (vertical / tracer bullet):
|
|
155
|
+
RED→GREEN: test1→impl1
|
|
156
|
+
RED→GREEN: test2→impl2
|
|
157
|
+
RED→GREEN: test3→impl3
|
|
158
|
+
```
|
|
159
|
+
|
|
160
|
+
The first vertical slice is the **tracer bullet** — it proves the path works end-to-end before you invest in breadth. If the tracer bullet reveals a bad design assumption, you've wasted one cycle, not five.
|
|
161
|
+
|
|
162
|
+
## Integration with Phases
|
|
163
|
+
|
|
164
|
+
- **Phase 2 (Plan)**: Test strategy is part of the plan — identify what tests will be written for each unit
|
|
165
|
+
- **Phase 3 (Implement)**: Every unit of work follows RED-GREEN-REFACTOR. The inline review checkpoint runs after GREEN, not during RED.
|
|
166
|
+
- **Acceptance tests**: Feature file scenarios (Gherkin) define the outer loop. TDD operates within each scenario's implementation.
|
|
167
|
+
|
|
168
|
+
## Output
|
|
169
|
+
Verified RED-GREEN-REFACTOR cycle evidence: failing test output, passing test output, and refactored code with passing tests for each unit of work.
|
|
@@ -0,0 +1,262 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: test-health
|
|
3
|
+
description: Project-wide test-strategy audit — derive the suite's shape and shape-vs-architecture fit, map coverage to the Agile Testing Quadrants, roll up coverage + mutation health, flag flaky tests and automation maturity, and produce an ordered improvement plan. Delegates CD-determinism + pipeline assessment to cd-test-architecture. Use when the user says "audit our tests", "how healthy is our test suite", "test strategy review", or runs /test-health. Advisory — writes a report, does not edit.
|
|
4
|
+
role: worker
|
|
5
|
+
user-invocable: true
|
|
6
|
+
argument-hint: "[--path <dir>] [--pdf] [--no-mutation]"
|
|
7
|
+
---
|
|
8
|
+
|
|
9
|
+
# Test Health
|
|
10
|
+
|
|
11
|
+
Role: worker. This command produces a strategic test-health report — it does
|
|
12
|
+
not edit code or tests; fixes go to `/apply-fixes`, refactors to `/plan` /
|
|
13
|
+
`/build`.
|
|
14
|
+
|
|
15
|
+
## Overview
|
|
16
|
+
|
|
17
|
+
An **advisory, project-wide** skill: it produces the *strategic-health* view of a test suite that a team needs periodically — the suite's **shape** vs. its architecture, **Agile Testing Quadrant** coverage, **coverage + mutation** health rolled up to ROI, flaky-test management, and **automation maturity** — then an ordered improvement plan. It complements, and does not duplicate, `cd-test-architecture`: that skill owns the CD-determinism + pipeline-placement assessment, which this skill **delegates to** rather than re-deriving.
|
|
18
|
+
|
|
19
|
+
Grounded in: `knowledge/testing-quadrants.md`, `knowledge/test-pyramid.md` (shapes + shape↔architecture fit), `knowledge/test-automation-maturity.md`, `knowledge/test-smells.md` (project smells / flakiness), and `knowledge/test-automation-principles.md` (the goals/principles that frame *why* a project smell hurts — e.g. Developers Not Writing Tests, Frequent Debugging → lost Defect Localization). It calls the `cd-test-architecture`, `/test-design`, and `mutation-testing` skills and folds their results into the strategic rollup.
|
|
20
|
+
|
|
21
|
+
## Constraints
|
|
22
|
+
|
|
23
|
+
- **Advisory only.** Write a report; do not edit code or tests. Hand fixes to `/apply-fixes`, refactors to `/plan` / `/build`.
|
|
24
|
+
- **Delegate, don't re-derive.** The architecture/pipeline section comes from `cd-test-architecture` — summarize its output, never restate or contradict its CD-determinism findings.
|
|
25
|
+
- **Strategic altitude.** This is a suite-level diagnostic. Per-file findings belong to `test-review` / `test-smell-review`; per-unit design belongs to `test-design-advisor`. Point to them; don't reproduce them.
|
|
26
|
+
- **No scoring reinvention.** Quantitative quality scoring and per-file design findings come from `/test-design` (Farley Score + test-review / test-smell-review) — consume them; summarize the themes and link to its report, don't re-derive or reproduce the per-file table.
|
|
27
|
+
- **Be concise.** One report; findings as tables, each item mapped to a concrete next move. No restating the knowledge files — cite them.
|
|
28
|
+
|
|
29
|
+
## Parse Arguments
|
|
30
|
+
|
|
31
|
+
Target repo/subtree path (default: cwd). Detect the test runner, coverage tool, and CI config from manifests and `.github/`/`.gitlab-ci.yml`/etc.
|
|
32
|
+
|
|
33
|
+
Optional:
|
|
34
|
+
|
|
35
|
+
- `--path <dir>`: audit only this subtree (default: cwd).
|
|
36
|
+
- `--pdf`: also render the report to a sibling PDF (Step 9).
|
|
37
|
+
- `--no-mutation`: **skip the `mutation-testing` invocation in Step 5** and
|
|
38
|
+
render the report's mutation column as `not enabled for this run`. Everything
|
|
39
|
+
else in Step 5 (the `/test-design` dispatch and its roll-up) is unchanged.
|
|
40
|
+
Default **off** — an unflagged `/test-health` runs mutation ROI exactly as
|
|
41
|
+
before, because mutation health is a real part of the standalone strategic
|
|
42
|
+
audit. The flag exists for a caller that has already decided mutation work is
|
|
43
|
+
out of scope for its run and would otherwise pay for a measurement it then
|
|
44
|
+
discards: `/test-improve` Phase 1 passes it when `phase-0.md` recorded
|
|
45
|
+
mutation mode `off` (#1961). A skip is **not** a waiver and not a failure —
|
|
46
|
+
the target was never in scope for this run.
|
|
47
|
+
|
|
48
|
+
## Steps
|
|
49
|
+
|
|
50
|
+
### 1. Trivial-suite short-circuit
|
|
51
|
+
|
|
52
|
+
If the suite is tiny (few test files), shows no shape pathology, and follows clear conventions, **stop here** and return a one-paragraph summary ("suite is small and healthy; nothing structural to fix; revisit when it grows") instead of the full diagnostic.
|
|
53
|
+
|
|
54
|
+
### 2. Derive the test shape + architecture fit
|
|
55
|
+
|
|
56
|
+
Inventory tests by layer (unit / integration / component / contract / E2E). Derive the actual **shape** and compare it to the shape the architecture *should* produce, using the *Other shapes* + *Shape ↔ architecture fit* tables in `test-pyramid.md`. Report the mismatch (e.g. tall pyramid over thin-glue code, or ice-cream cone), not the silhouette alone.
|
|
57
|
+
|
|
58
|
+
**Graph-assisted inventory.** Prefer CodeGraph/Repowise over raw `Grep` for mapping test files to the architecture layers they exercise — see [`knowledge/codegraph-vs-graphify.md`](../../knowledge/codegraph-vs-graphify.md) for tool selection and the fallback contract.
|
|
59
|
+
|
|
60
|
+
### 3. Quadrant coverage
|
|
61
|
+
|
|
62
|
+
Classify coverage across the four quadrants (`testing-quadrants.md`) as strong / thin / empty, and for each gap name the **business impact** of leaving it empty (e.g. empty Q3 → no human catches confusing flows; empty Q4 → non-functional failures reach prod).
|
|
63
|
+
|
|
64
|
+
**Gherkin scenarios as a Q2 signal.** Run `python3
|
|
65
|
+
${CLAUDE_PLUGIN_ROOT}/scripts/detect_bdd_convention.py` (the same detection
|
|
66
|
+
`/gherkin-derive` and `/plan` already use) to locate the project's
|
|
67
|
+
`.feature` directory. When a directory is reported, enumerate its scenario
|
|
68
|
+
titles (`Scenario:` / `Scenario Outline:` lines) and count them as an
|
|
69
|
+
additional Q2 (business-facing) signal alongside any existing
|
|
70
|
+
business-facing tests — a documented scenario counts toward Q2 strength
|
|
71
|
+
whether or not it is yet bound to a runnable test; binding status feeds
|
|
72
|
+
Step 7's gap classification, not this step's strong/thin/empty call. When
|
|
73
|
+
`detect_bdd_convention.py` reports no signal, Q2's classification is
|
|
74
|
+
unaffected — it falls back to the existing business-facing-test-based
|
|
75
|
+
signal only, with no error and no placeholder text.
|
|
76
|
+
|
|
77
|
+
### 4. Delegate architecture + pipeline
|
|
78
|
+
|
|
79
|
+
Invoke `cd-test-architecture` on the target. Summarize its findings (which tests can't run in a clean pre-merge gate, target architecture, migration path) in one section — **do not re-derive**.
|
|
80
|
+
|
|
81
|
+
### 5. Test-design + mutation health (ROI)
|
|
82
|
+
|
|
83
|
+
Invoke `/test-design` on the target, passing the same scope this run was
|
|
84
|
+
invoked with — the dispatch must be explicit so a subtree audit never
|
|
85
|
+
inherits a whole-repo Farley Score:
|
|
86
|
+
|
|
87
|
+
- `/test-health --path <dir>` → dispatch `/test-design --path <dir>`.
|
|
88
|
+
- `/test-health` (unscoped) → dispatch `/test-design` with no scope flag.
|
|
89
|
+
|
|
90
|
+
Consume its results: the scope-labelled **Farley Score** (`(all tests)` when
|
|
91
|
+
unscoped, `(under <dir>)` when passing `--path`, or the empty-scope note
|
|
92
|
+
`no in-scope test files` when the in-scope set is empty), the dominant
|
|
93
|
+
`test-review` / `test-smell-review` themes (weak assertions,
|
|
94
|
+
non-determinism, fixture/structure smells, testability blockers), and the
|
|
95
|
+
advisor's testability verdicts. Then invoke `mutation-testing` on the
|
|
96
|
+
**critical-logic** modules only (not the whole repo — that's the ROI
|
|
97
|
+
framing).
|
|
98
|
+
|
|
99
|
+
**`--no-mutation` skips that invocation entirely (#1961).** When the flag is
|
|
100
|
+
set, do not run `mutation-testing`, do not estimate a mutation figure to stand
|
|
101
|
+
in for it, and render the report's mutation content as `not enabled for this
|
|
102
|
+
run`. The rest of this step is unaffected — `/test-design` still runs and its
|
|
103
|
+
themes still feed the ordered plan (Step 8), which simply carries no
|
|
104
|
+
mutation-hotspot input for this run. Without the flag (the default), invoke
|
|
105
|
+
`mutation-testing` as above.
|
|
106
|
+
|
|
107
|
+
Roll both up: where is coverage high but mutation-weak
|
|
108
|
+
(assertions that don't catch bugs)? Where do test-design smells
|
|
109
|
+
concentrate? Where is critical logic under-covered? Prioritize by risk,
|
|
110
|
+
not by raw %. Both feed the ordered plan (Step 7) — summarize the themes
|
|
111
|
+
and link to the `/test-design` report for per-file detail; do not
|
|
112
|
+
reproduce it.
|
|
113
|
+
|
|
114
|
+
### 6. Flaky-test + automation maturity
|
|
115
|
+
|
|
116
|
+
Flag flakiness signals (`test-smells.md` project/behavior smells: order-dependence, unstubbed clock/RNG, real I/O at unit level) and a management recommendation (quarantine + fix, don't `retry`). Assess automation maturity with `test-automation-maturity.md`: report the rung and the single-point-of-change metric, scaled by suite size (graduated thresholds).
|
|
117
|
+
|
|
118
|
+
### 7. Classify gaps + recommend removals
|
|
119
|
+
|
|
120
|
+
Classify every gap the audit surfaces into one of three **actionable** classes, plus one **non-actionable** class reserved for behavior that doesn't exist yet (see Gherkin gaps below), so the improvement plan (Step 8) only ever plans work that delivers signal:
|
|
121
|
+
|
|
122
|
+
| Class | Meaning | Action |
|
|
123
|
+
| --- | --- | --- |
|
|
124
|
+
| `NO_REFACTOR` | A test can be added against the code as it stands | Plan it |
|
|
125
|
+
| `REFACTOR_REQUIRED` | Production code needs a testability change before a meaningful test is possible | Plan the production-code change first |
|
|
126
|
+
| `LOW_VALUE` | Technically feasible but delivers no signal — **skip, never plan** | List for removal, not for work |
|
|
127
|
+
| `NOT_IMPLEMENTED` | The scenario's behavior doesn't exist in production code at all — not a testability gap | Feature-gap call-out for the product backlog, never a test-improve target |
|
|
128
|
+
|
|
129
|
+
A finding is `LOW_VALUE` only when **all three** hold:
|
|
130
|
+
|
|
131
|
+
1. **No branching logic** — trivial getters/setters, pass-through constructors, framework boilerplate, or auto-generated code.
|
|
132
|
+
2. **No observable outcome** — the only assertion possible is that a mock was called.
|
|
133
|
+
3. **Coverage already provided** — a higher-layer test already exercises the same path.
|
|
134
|
+
|
|
135
|
+
For existing tests that meet all three criteria, emit a **Recommended removals** table: the redundant test, the higher-layer test that already covers it, and a one-line rationale. These are the suite's `LOW_VALUE` tests — keeping them costs maintenance for no defect-localization gain.
|
|
136
|
+
|
|
137
|
+
**Gherkin gaps (tag `gherkin-gap`).** A documented Gherkin scenario surfaced
|
|
138
|
+
in Step 3 with no bound step-definition (`bdd-runner` mode) or cited xUnit
|
|
139
|
+
test (`xunit-with-annotations` mode) classifies `NO_REFACTOR` when its
|
|
140
|
+
behavior already exists in production code and a test can be added against
|
|
141
|
+
it as-is. It classifies `REFACTOR_REQUIRED` when the behavior exists but
|
|
142
|
+
needs a testability seam first (interface extraction, DI point, virtual
|
|
143
|
+
method promotion) — the same kind of seam-only change `/test-improve`'s
|
|
144
|
+
Phase 7 is scoped to perform.
|
|
145
|
+
|
|
146
|
+
**When the scenario's behavior doesn't exist yet in production code at
|
|
147
|
+
all, classify it `NOT_IMPLEMENTED` instead — neither `NO_REFACTOR` nor
|
|
148
|
+
`REFACTOR_REQUIRED` applies.** Phase 7 accepts seam introductions only —
|
|
149
|
+
implementing new behavior is explicitly out of its scope — so routing a
|
|
150
|
+
`NOT_IMPLEMENTED` scenario through `REFACTOR_REQUIRED` would dead-end there
|
|
151
|
+
with no seam to introduce. Tag it `gherkin-gap` as usual, mark it
|
|
152
|
+
`NOT_IMPLEMENTED` in the Gap classification table's Class column, and carry
|
|
153
|
+
it into Step 8's ordered improvement plan as a feature-gap call-out for the
|
|
154
|
+
product backlog — never as a refactor-for-testability item, and never
|
|
155
|
+
written as a Phase-5 Story or deferred to Phase 7 by `/test-improve`.
|
|
156
|
+
|
|
157
|
+
**Discriminator: judge existence against the `Then` outcome, not the
|
|
158
|
+
`Given`/`When` setup.** If no production code path can produce the asserted
|
|
159
|
+
outcome, the behavior does not exist even when the surrounding
|
|
160
|
+
`Given`/`When` code does. Split a partially-implemented Scenario Outline
|
|
161
|
+
per example row — some rows can be `REFACTOR_REQUIRED` while others are
|
|
162
|
+
`NOT_IMPLEMENTED`. When existence is genuinely ambiguous, default to
|
|
163
|
+
`NOT_IMPLEMENTED`: an over-classified feature gap is inert and visible in
|
|
164
|
+
the report, whereas an over-classified `REFACTOR_REQUIRED` dead-ends at
|
|
165
|
+
Phase 7 with no seam to introduce.
|
|
166
|
+
|
|
167
|
+
Use this exact wording for the finding's Meaning/Action for the
|
|
168
|
+
`NO_REFACTOR`/`REFACTOR_REQUIRED` cases:
|
|
169
|
+
|
|
170
|
+
> **Meaning:** "documented Gherkin scenario '<title>' has no bound step-definition or cited test yet." **Action:** "implement the missing binding/test before treating this area as low-risk."
|
|
171
|
+
|
|
172
|
+
For the `NOT_IMPLEMENTED` case, use this wording instead — it must never
|
|
173
|
+
tell the operator to add a binding/test, since there is no behavior yet to
|
|
174
|
+
bind one to:
|
|
175
|
+
|
|
176
|
+
> **Meaning:** "documented Gherkin scenario '<title>' describes behavior that does not exist in production code yet." **Action:** "implement the behavior via the product backlog; the binding/test follows once it exists — this is not a /test-improve target."
|
|
177
|
+
|
|
178
|
+
A Gherkin gap can never qualify as `LOW_VALUE`: a documented scenario has an
|
|
179
|
+
observable outcome by construction (its own `Then` steps assert one), so it
|
|
180
|
+
always fails criterion 2 above. `NOT_IMPLEMENTED` and `LOW_VALUE` are
|
|
181
|
+
mutually exclusive for the same reason `LOW_VALUE` never applies to a
|
|
182
|
+
Gherkin gap.
|
|
183
|
+
|
|
184
|
+
### 8. Ordered improvement plan
|
|
185
|
+
|
|
186
|
+
Produce a risk-ordered, incremental plan — each item a concrete next move (which layer to add, which shape to correct, which quadrant to fill, which abstraction to extract, which weak-assertion or smell cluster to fix), driven by the test-design themes and mutation hotspots from Step 5. `LOW_VALUE` findings never appear here — they live only in the Recommended removals table.
|
|
187
|
+
|
|
188
|
+
`NOT_IMPLEMENTED` findings never appear here either — no test-improvement move can close a scenario whose behavior doesn't exist yet, so it can never satisfy "a concrete next move." List them instead in a dedicated **Feature gaps (product backlog)** call-out beneath the ordered plan, and never count them in the risk ordering.
|
|
189
|
+
|
|
190
|
+
### 9. Report
|
|
191
|
+
|
|
192
|
+
Write `.dev-team-reports/test-health-<date>.md`.
|
|
193
|
+
|
|
194
|
+
When `--pdf` was passed, render that report to a sibling PDF per
|
|
195
|
+
`knowledge/report-pdf-integration.md` (additive; non-fatal if no engine):
|
|
196
|
+
|
|
197
|
+
```bash
|
|
198
|
+
sh "$CLAUDE_PLUGIN_ROOT/hooks/py.sh" "$CLAUDE_PLUGIN_ROOT/hooks/lib/report_pdf.py" .dev-team-reports/test-health-<date>.md
|
|
199
|
+
```
|
|
200
|
+
|
|
201
|
+
## Output
|
|
202
|
+
|
|
203
|
+
For the header block and closing Provenance section, follow
|
|
204
|
+
`knowledge/report-template.md`; the sections below are this skill's own
|
|
205
|
+
body.
|
|
206
|
+
|
|
207
|
+
```markdown
|
|
208
|
+
# Test Health
|
|
209
|
+
|
|
210
|
+
**Date**: <ISO 8601>
|
|
211
|
+
**Target**: <repo>
|
|
212
|
+
**Tool versions**: <coverage tool, mutation tool versions — or _Not applicable — <reason>._>
|
|
213
|
+
**Scope**: <full repo | --path <dir>>
|
|
214
|
+
|
|
215
|
+
## Test Health — <repo> (<date>)
|
|
216
|
+
|
|
217
|
+
**Shape**: <derived> · **Expected for this architecture**: <expected> · **Fit**: <match|mismatch + why>
|
|
218
|
+
|
|
219
|
+
### Quadrant coverage
|
|
220
|
+
| Quadrant | Status | Gap impact |
|
|
221
|
+
|
|
222
|
+
### Architecture & pipeline (via cd-test-architecture)
|
|
223
|
+
<one-paragraph summary + link to its report>
|
|
224
|
+
|
|
225
|
+
### Test-design & mutation health (via /test-design + mutation-testing)
|
|
226
|
+
<When `--no-mutation` was passed, this section's mutation content reads
|
|
227
|
+
`not enabled for this run` — never a number, never a waiver.>
|
|
228
|
+
<Farley Score — render the scope-labelled value from `/test-design`
|
|
229
|
+
verbatim (`(all tests)` when this run is unscoped, `(under <dir>)` when
|
|
230
|
+
`--path` is set), or the literal `no in-scope test files` when the in-scope
|
|
231
|
+
set was empty; do not synthesize a number. Top test-design themes ·
|
|
232
|
+
mutation ROI hotspots · under-covered critical logic>
|
|
233
|
+
|
|
234
|
+
### Gap classification
|
|
235
|
+
| Gap | Class (NO_REFACTOR / REFACTOR_REQUIRED / LOW_VALUE / NOT_IMPLEMENTED) | Note |
|
|
236
|
+
|
|
237
|
+
### Recommended removals (LOW_VALUE existing tests)
|
|
238
|
+
| Test to remove | Covering test | Rationale |
|
|
239
|
+
|
|
240
|
+
### Flakiness & automation maturity
|
|
241
|
+
<flaky signals + management rec · maturity rung · single-point-of-change metric>
|
|
242
|
+
|
|
243
|
+
### Improvement plan (ordered)
|
|
244
|
+
1. <highest-leverage move> …
|
|
245
|
+
|
|
246
|
+
### Feature gaps (product backlog)
|
|
247
|
+
<NOT_IMPLEMENTED gherkin-gap findings only — never a numbered plan item;
|
|
248
|
+
omit this section entirely when there are none>
|
|
249
|
+
- <gap> — <one-line rationale for why the behavior doesn't exist yet>…
|
|
250
|
+
|
|
251
|
+
## Provenance
|
|
252
|
+
|
|
253
|
+
- Repository: `<repo path>`
|
|
254
|
+
- Branch / SHA: `<branch>` / `<sha>`
|
|
255
|
+
- Run parameters: `<flags — e.g. --path <dir>>`
|
|
256
|
+
- `dev-team` plugin version: `<plugin_version>`
|
|
257
|
+
```
|
|
258
|
+
|
|
259
|
+
## Integration
|
|
260
|
+
|
|
261
|
+
- **Front door** for periodic test-strategy review; the unified entry point that runs `cd-test-architecture` + `/test-design` + `mutation-testing` and rolls their results into one strategic view.
|
|
262
|
+
- `/test-design` runs inside this flow (Step 5) and also stands alone for a focused per-file review. For *forward* design of a specific module, use `test-design-advisor`. This skill is the strategic rollup that consumes their output.
|