pi-dev-team 0.2.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/LICENSE +21 -0
- package/PORTING.md +134 -0
- package/README.md +207 -0
- package/UPSTREAM.json +64 -0
- package/agents/Explore.md +15 -0
- package/agents/a11y-review.md +118 -0
- package/agents/adr-author.md +70 -0
- package/agents/ai-provenance-review.md +120 -0
- package/agents/angular-reactivity-review.md +95 -0
- package/agents/arch-review.md +135 -0
- package/agents/architect.md +78 -0
- package/agents/autoship-batch-proposer.md +69 -0
- package/agents/claude-setup-review.md +136 -0
- package/agents/codebase-recon.md +184 -0
- package/agents/component-architecture-review.md +119 -0
- package/agents/concurrency-review.md +109 -0
- package/agents/correctness-review.md +290 -0
- package/agents/data-flow-tracer.md +120 -0
- package/agents/doc-review.md +165 -0
- package/agents/domain-review.md +136 -0
- package/agents/general-purpose.md +10 -0
- package/agents/gherkin-quality-critic.md +113 -0
- package/agents/js-fp-review.md +114 -0
- package/agents/mutation-kill.md +684 -0
- package/agents/naming-review.md +142 -0
- package/agents/orchestrator.md +339 -0
- package/agents/performance-review.md +105 -0
- package/agents/plan-review-acceptance.md +115 -0
- package/agents/plan-review-design.md +90 -0
- package/agents/plan-review-parallelization.md +84 -0
- package/agents/plan-review-strategic.md +96 -0
- package/agents/plan-review-ux.md +110 -0
- package/agents/platform-engineer.md +64 -0
- package/agents/product-manager.md +68 -0
- package/agents/progress-guardian.md +79 -0
- package/agents/qa-engineer.md +289 -0
- package/agents/quality-reviewer.md +132 -0
- package/agents/react-reactivity-review.md +102 -0
- package/agents/refactor-opportunity-review.md +128 -0
- package/agents/security-engineer.md +60 -0
- package/agents/security-review.md +218 -0
- package/agents/session-analysis.md +95 -0
- package/agents/software-engineer.md +105 -0
- package/agents/spec-compliance-review.md +100 -0
- package/agents/spec-reviewer.md +114 -0
- package/agents/structure-review.md +146 -0
- package/agents/tech-writer.md +84 -0
- package/agents/test-review.md +246 -0
- package/agents/test-smell-review.md +188 -0
- package/agents/token-efficiency-review.md +139 -0
- package/agents/ui-ux-designer.md +54 -0
- package/agents/vue-reactivity-review.md +95 -0
- package/bin/__pycache__/claudecpython-314.pyc +0 -0
- package/bin/claude +258 -0
- package/docs/upstream/.pages +1 -0
- package/docs/upstream/CHANGELOG.md +2586 -0
- package/docs/upstream/README.md +155 -0
- package/docs/upstream/agent-architecture.md +214 -0
- package/docs/upstream/agent_info.md +187 -0
- package/docs/upstream/artifact-migration.md +124 -0
- package/docs/upstream/code-intelligence-nudge.md +149 -0
- package/docs/upstream/code-review-process.md +294 -0
- package/docs/upstream/concurrent-use.md +73 -0
- package/docs/upstream/context-management.md +111 -0
- package/docs/upstream/developer-notes.md +280 -0
- package/docs/upstream/diagrams/architecture-overview.svg +101 -0
- package/docs/upstream/diagrams/review-dispatch.svg +139 -0
- package/docs/upstream/diagrams/team-agents.svg +128 -0
- package/docs/upstream/diagrams/test-improve-flow.svg +166 -0
- package/docs/upstream/diagrams/workflow-linear.svg +66 -0
- package/docs/upstream/diagrams/workflow-three-phase.svg +200 -0
- package/docs/upstream/eval-maintenance.md +95 -0
- package/docs/upstream/eval-running-guide.md +147 -0
- package/docs/upstream/eval-system.md +291 -0
- package/docs/upstream/session-review-oss-complements.md +75 -0
- package/docs/upstream/session-review.md +212 -0
- package/docs/upstream/skills.md +188 -0
- package/docs/upstream/team-structure.md +21 -0
- package/docs/upstream/telemetry-ci-access.md +129 -0
- package/docs/upstream/telemetry-repo-security.md +120 -0
- package/docs/upstream/test-evaluation.md +277 -0
- package/docs/upstream/test-improve.md +154 -0
- package/docs/upstream/triage-workflow.md +282 -0
- package/docs/upstream/workflows.md +289 -0
- package/extensions/dev-team/index.ts +539 -0
- package/extensions/dev-team/lib/agents.ts +272 -0
- package/extensions/dev-team/lib/ai-credits.ts +92 -0
- package/extensions/dev-team/lib/autocompact.ts +81 -0
- package/extensions/dev-team/lib/child-run.ts +102 -0
- package/extensions/dev-team/lib/config.ts +236 -0
- package/extensions/dev-team/lib/gh-command.ts +103 -0
- package/extensions/dev-team/lib/github-style.ts +307 -0
- package/extensions/dev-team/lib/hooks.ts +350 -0
- package/extensions/dev-team/lib/metrics.ts +115 -0
- package/extensions/dev-team/lib/safe-read.ts +49 -0
- package/extensions/dev-team/lib/session-files.ts +57 -0
- package/extensions/dev-team/lib/session-spend.ts +123 -0
- package/extensions/dev-team/lib/shell-scan.ts +205 -0
- package/extensions/dev-team/lib/skills.ts +213 -0
- package/extensions/dev-team/lib/subagent-render.ts +245 -0
- package/extensions/dev-team/lib/subagent-types.ts +164 -0
- package/extensions/dev-team/lib/subagent.ts +596 -0
- package/extensions/dev-team/lib/terminal-text.ts +54 -0
- package/extensions/dev-team/lib/tools-misc.ts +152 -0
- package/extensions/dev-team/lib/transcript.ts +110 -0
- package/extensions/dev-team/lib/trust.ts +52 -0
- package/extensions/dev-team/lib/usage-breakdown.ts +176 -0
- package/extensions/dev-team/lib/usage-chart.ts +153 -0
- package/extensions/dev-team/lib/usage-command.ts +107 -0
- package/extensions/dev-team/lib/usage-history.ts +203 -0
- package/extensions/dev-team/lib/usage-render.ts +225 -0
- package/extensions/dev-team/lib/usage-split-bar.ts +127 -0
- package/extensions/dev-team/lib/usage-state.ts +116 -0
- package/extensions/dev-team/lib/usage-text.ts +159 -0
- package/extensions/dev-team/lib/usage-view.ts +109 -0
- package/hooks/__pycache__/refactor_test_freeze_guard.cpython-314.pyc +0 -0
- package/hooks/agent_dispatch_ledger.py +190 -0
- package/hooks/autocompact_setup_nudge.py +99 -0
- package/hooks/bash_retry_guard.py +228 -0
- package/hooks/boundary_events_write_guard.py +352 -0
- package/hooks/code_intelligence_nudge.py +293 -0
- package/hooks/code_intelligence_turn_mark.py +317 -0
- package/hooks/codegraph_bootstrap.py +139 -0
- package/hooks/contract_version_guard.py +362 -0
- package/hooks/cost_meter.py +106 -0
- package/hooks/destructive-commands.json +62 -0
- package/hooks/destructive_guard.py +477 -0
- package/hooks/eval_compliance_check.py +440 -0
- package/hooks/guards.json +17 -0
- package/hooks/hooks.json +323 -0
- package/hooks/internal_double_gate.py +296 -0
- package/hooks/js_fp_review.py +212 -0
- package/hooks/knowledge_index.py +119 -0
- package/hooks/lib/__pycache__/artifact_paths.cpython-314.pyc +0 -0
- package/hooks/lib/__pycache__/atomic_state.cpython-314.pyc +0 -0
- package/hooks/lib/__pycache__/autocompact_config.cpython-314.pyc +0 -0
- package/hooks/lib/__pycache__/boundary_events.cpython-314.pyc +0 -0
- package/hooks/lib/__pycache__/doc_classification.cpython-314.pyc +0 -0
- package/hooks/lib/__pycache__/gh_pr_create_detect.cpython-314.pyc +0 -0
- package/hooks/lib/__pycache__/git_safe_diff.cpython-314.pyc +0 -0
- package/hooks/lib/__pycache__/instrument_log.cpython-314.pyc +0 -0
- package/hooks/lib/__pycache__/metrics_query.cpython-314.pyc +0 -0
- package/hooks/lib/__pycache__/plugin_version.cpython-314.pyc +0 -0
- package/hooks/lib/__pycache__/pre_commit_doc_classifier.cpython-314.pyc +0 -0
- package/hooks/lib/__pycache__/review_agent_registry.cpython-314.pyc +0 -0
- package/hooks/lib/__pycache__/review_gate_corroboration.cpython-314.pyc +0 -0
- package/hooks/lib/__pycache__/review_gate_hash.cpython-314.pyc +0 -0
- package/hooks/lib/__pycache__/review_verdicts.cpython-314.pyc +0 -0
- package/hooks/lib/__pycache__/stdin_json.cpython-314.pyc +0 -0
- package/hooks/lib/__pycache__/stryker_invocation.cpython-314.pyc +0 -0
- package/hooks/lib/__pycache__/telemetry_consent.cpython-314.pyc +0 -0
- package/hooks/lib/__pycache__/test_file_classify.cpython-314.pyc +0 -0
- package/hooks/lib/__pycache__/token_efficiency_limits.cpython-314.pyc +0 -0
- package/hooks/lib/__pycache__/verify_guard_state.cpython-314.pyc +0 -0
- package/hooks/lib/__pycache__/xunit_v3_operator_gate.cpython-314.pyc +0 -0
- package/hooks/lib/agent_skill_hints.py +74 -0
- package/hooks/lib/artifact_paths.py +263 -0
- package/hooks/lib/atomic_state.py +557 -0
- package/hooks/lib/autocompact_config.py +103 -0
- package/hooks/lib/autoship_log.py +106 -0
- package/hooks/lib/banned_scripts_policy.py +51 -0
- package/hooks/lib/boundary_events.py +436 -0
- package/hooks/lib/build_knowledge_index.py +504 -0
- package/hooks/lib/build_skills_index.py +361 -0
- package/hooks/lib/build_state.py +116 -0
- package/hooks/lib/classify_ship_outcome.py +126 -0
- package/hooks/lib/config_changelog_schema.py +115 -0
- package/hooks/lib/cost_meter.py +955 -0
- package/hooks/lib/doc_classification.py +116 -0
- package/hooks/lib/gh_pr_create_detect.py +136 -0
- package/hooks/lib/git_safe_diff.py +123 -0
- package/hooks/lib/instrument_log.py +66 -0
- package/hooks/lib/iteration_journal_gate.py +197 -0
- package/hooks/lib/knowledge_index_paths.py +88 -0
- package/hooks/lib/mcp_json_repowise.py +177 -0
- package/hooks/lib/metrics_query.py +202 -0
- package/hooks/lib/minimal_yaml.py +434 -0
- package/hooks/lib/plugin_version.py +142 -0
- package/hooks/lib/pre_commit_detect.py +537 -0
- package/hooks/lib/pre_commit_doc_classifier.py +126 -0
- package/hooks/lib/pricing.py +118 -0
- package/hooks/lib/report_pdf.py +371 -0
- package/hooks/lib/review_agent_registry.py +142 -0
- package/hooks/lib/review_dispatch_ledger.py +101 -0
- package/hooks/lib/review_gate_corroboration.py +521 -0
- package/hooks/lib/review_gate_hash.py +252 -0
- package/hooks/lib/review_gate_normalized_hash.py +1115 -0
- package/hooks/lib/review_verdicts.py +301 -0
- package/hooks/lib/run_report.py +160 -0
- package/hooks/lib/skill_categories.yaml +125 -0
- package/hooks/lib/stdin_json.py +57 -0
- package/hooks/lib/stryker_invocation.py +102 -0
- package/hooks/lib/telemetry_consent.py +41 -0
- package/hooks/lib/telemetry_report.py +108 -0
- package/hooks/lib/test_file_classify.py +160 -0
- package/hooks/lib/token_efficiency_limits.py +51 -0
- package/hooks/lib/turn_identity.py +77 -0
- package/hooks/lib/verify_guard_state.py +110 -0
- package/hooks/lib/workflow_state.py +206 -0
- package/hooks/lib/xunit_v3_operator_gate.py +596 -0
- package/hooks/mcp_json_repowise_nudge.py +74 -0
- package/hooks/mutation_adapters/__init__.py +7 -0
- package/hooks/mutation_adapters/__pycache__/__init__.cpython-314.pyc +0 -0
- package/hooks/mutation_adapters/__pycache__/lib.cpython-314.pyc +0 -0
- package/hooks/mutation_adapters/__pycache__/mutmut.cpython-314.pyc +0 -0
- package/hooks/mutation_adapters/__pycache__/pitest.cpython-314.pyc +0 -0
- package/hooks/mutation_adapters/__pycache__/stryker.cpython-314.pyc +0 -0
- package/hooks/mutation_adapters/__pycache__/stryker_net.cpython-314.pyc +0 -0
- package/hooks/mutation_adapters/lib.py +478 -0
- package/hooks/mutation_adapters/mutmut.py +188 -0
- package/hooks/mutation_adapters/pitest.py +266 -0
- package/hooks/mutation_adapters/stryker.py +157 -0
- package/hooks/mutation_adapters/stryker_net.py +264 -0
- package/hooks/mutation_gate.py +193 -0
- package/hooks/mutation_testing_smoke_gate.py +371 -0
- package/hooks/pending_review_notify.py +121 -0
- package/hooks/phase_marker.py +138 -0
- package/hooks/post_compact_state_reinject.py +180 -0
- package/hooks/post_format.py +115 -0
- package/hooks/pre_commit_knowledge_index.py +128 -0
- package/hooks/pre_commit_review.py +66 -0
- package/hooks/pre_pr_review.py +694 -0
- package/hooks/pre_tool_guard.py +405 -0
- package/hooks/py.sh +73 -0
- package/hooks/refactor-bash-write-patterns.json +29 -0
- package/hooks/refactor_test_bash_guard.py +253 -0
- package/hooks/refactor_test_freeze_guard.py +139 -0
- package/hooks/refactor_test_revert_guard.py +186 -0
- package/hooks/repo_review_nudge.py +287 -0
- package/hooks/review_verdict_recorder.py +464 -0
- package/hooks/scan_bash_command_for_banned_scripts.py +428 -0
- package/hooks/scan_worktree_for_banned_scripts.py +238 -0
- package/hooks/session_learning_trigger.py +248 -0
- package/hooks/skills_index.py +126 -0
- package/hooks/stryker_xunit_shim_guard.py +571 -0
- package/hooks/subagent_completion_guard.py +309 -0
- package/hooks/subagent_skill_context.py +139 -0
- package/hooks/task_completion_metrics.py +216 -0
- package/hooks/tdd_guard.py +229 -0
- package/hooks/telemetry.py +341 -0
- package/hooks/token_efficiency_review.py +194 -0
- package/hooks/verify_guard.py +183 -0
- package/hooks/verify_guard_edit_marker.py +73 -0
- package/hooks/version_check.py +173 -0
- package/knowledge/accepted-risks-schema.md +98 -0
- package/knowledge/adr-decision-criteria.md +64 -0
- package/knowledge/adversarial-review-protocol.md +139 -0
- package/knowledge/agent-registry.md +228 -0
- package/knowledge/agent-review-methodology.md +80 -0
- package/knowledge/ai-friendly-repo-guidelines.md +67 -0
- package/knowledge/architecture-assessment.md +96 -0
- package/knowledge/artifact-lifecycle.md +57 -0
- package/knowledge/cd-maturity-model.md +82 -0
- package/knowledge/cd-test-architecture.md +190 -0
- package/knowledge/ci-cd-file-scope.md +24 -0
- package/knowledge/codegraph-vs-graphify.md +192 -0
- package/knowledge/component-test-patterns.md +139 -0
- package/knowledge/database-change-management.md +80 -0
- package/knowledge/database-test-patterns.md +79 -0
- package/knowledge/decision-defaults.md +88 -0
- package/knowledge/dependency-breaking-techniques.md +116 -0
- package/knowledge/deployment-pipeline.md +86 -0
- package/knowledge/design-smells.md +122 -0
- package/knowledge/directory-enumeration.md +38 -0
- package/knowledge/domain-modeling.md +123 -0
- package/knowledge/evidence-bundle.md +90 -0
- package/knowledge/exploratory-testing-field-guide.md +122 -0
- package/knowledge/failure-routing.md +28 -0
- package/knowledge/fixture-construction.md +56 -0
- package/knowledge/frontend-component-architecture.md +139 -0
- package/knowledge/gherkin-quality-review-dispatch.md +135 -0
- package/knowledge/index.json +6766 -0
- package/knowledge/internal-collaborator-doubling.md +101 -0
- package/knowledge/legacy-test-strategy.md +71 -0
- package/knowledge/long-run-waiting.md +66 -0
- package/knowledge/microservice-testing.md +71 -0
- package/knowledge/model-pricing.json +23 -0
- package/knowledge/mutation-score-formulas.md +60 -0
- package/knowledge/object-calisthenics.md +147 -0
- package/knowledge/oracle-provenance.md +94 -0
- package/knowledge/orchestrator-script-implementation.md +185 -0
- package/knowledge/owasp-detection.md +148 -0
- package/knowledge/plan-review-rubric.md +56 -0
- package/knowledge/proxy-connectivity.md +62 -0
- package/knowledge/reactive-effect-patterns.md +73 -0
- package/knowledge/recon-inventory-excludes.txt +32 -0
- package/knowledge/references/bdd-value-guide.md +61 -0
- package/knowledge/references/csharp-http-client-testing.md +264 -0
- package/knowledge/release-strategies.md +74 -0
- package/knowledge/report-output-location.md +117 -0
- package/knowledge/report-pdf-integration.md +63 -0
- package/knowledge/report-print.css +129 -0
- package/knowledge/report-template.md +114 -0
- package/knowledge/report-to-pdf.md +69 -0
- package/knowledge/request-processing-flow.md +63 -0
- package/knowledge/result-verification.md +52 -0
- package/knowledge/review-agent-output-contract.md +121 -0
- package/knowledge/review-lens-classification.md +113 -0
- package/knowledge/review-rubric.md +62 -0
- package/knowledge/review-template.md +104 -0
- package/knowledge/rule-fixtures/A02.insecure-random-js/negative.js +1 -0
- package/knowledge/rule-fixtures/A02.insecure-random-js/positive.js +1 -0
- package/knowledge/rule-fixtures/A02.weak-hashing-md5/negative.py +1 -0
- package/knowledge/rule-fixtures/A02.weak-hashing-md5/positive.py +1 -0
- package/knowledge/rule-fixtures/A03.command-injection/negative.js +1 -0
- package/knowledge/rule-fixtures/A03.command-injection/positive.js +1 -0
- package/knowledge/rule-fixtures/A03.sql-injection/negative.js +1 -0
- package/knowledge/rule-fixtures/A03.sql-injection/positive.js +1 -0
- package/knowledge/rule-fixtures/A03.xss-innerhtml/negative.js +1 -0
- package/knowledge/rule-fixtures/A03.xss-innerhtml/positive.js +1 -0
- package/knowledge/rule-fixtures/A05.cors-wildcard/negative.js +1 -0
- package/knowledge/rule-fixtures/A05.cors-wildcard/positive.js +1 -0
- package/knowledge/rule-fixtures/A05.default-credentials/negative.js +1 -0
- package/knowledge/rule-fixtures/A05.default-credentials/positive.js +1 -0
- package/knowledge/rule-fixtures/A07.jwt-alg-none/negative.js +1 -0
- package/knowledge/rule-fixtures/A07.jwt-alg-none/positive.js +1 -0
- package/knowledge/rule-fixtures/A08.binary-formatter/negative.cs +1 -0
- package/knowledge/rule-fixtures/A08.binary-formatter/positive.cs +1 -0
- package/knowledge/rule-fixtures/A08.js-eval/negative.js +1 -0
- package/knowledge/rule-fixtures/A08.js-eval/positive.js +1 -0
- package/knowledge/rule-fixtures/A08.object-input-stream/negative.java +1 -0
- package/knowledge/rule-fixtures/A08.object-input-stream/positive.java +1 -0
- package/knowledge/schemas/disposition-register-v1.json +65 -0
- package/knowledge/schemas/recon-envelope-v1.json +198 -0
- package/knowledge/schemas/unified-finding-v1.json +72 -0
- package/knowledge/security-primitives-contract.md +301 -0
- package/knowledge/security-review-rule-map.yaml +107 -0
- package/knowledge/skills-registry.md +72 -0
- package/knowledge/task-size-classifier.md +103 -0
- package/knowledge/telemetry-schema.md +881 -0
- package/knowledge/test-automation-maturity.md +56 -0
- package/knowledge/test-automation-principles.md +71 -0
- package/knowledge/test-cadence-tradeoffs.md +68 -0
- package/knowledge/test-doubles.md +105 -0
- package/knowledge/test-file-indicators.md +22 -0
- package/knowledge/test-layer-gates.md +35 -0
- package/knowledge/test-matrix-examples/django-batch.md +24 -0
- package/knowledge/test-matrix-examples/dotnet-grpc-fronting-api.md +90 -0
- package/knowledge/test-matrix-examples/dotnet-http-consumer.md +131 -0
- package/knowledge/test-matrix-examples/react-node-spa.md +24 -0
- package/knowledge/test-matrix-examples/spring-boot-service.md +25 -0
- package/knowledge/test-matrix-examples/ssr-htmx.md +24 -0
- package/knowledge/test-organization.md +70 -0
- package/knowledge/test-pyramid.md +84 -0
- package/knowledge/test-refactoring.md +67 -0
- package/knowledge/test-review-division-of-labor.md +85 -0
- package/knowledge/test-smells.md +80 -0
- package/knowledge/test-stack-profiles/bdd-frameworks.md +235 -0
- package/knowledge/test-stack-profiles/django.md +13 -0
- package/knowledge/test-stack-profiles/dotnet.md +18 -0
- package/knowledge/test-stack-profiles/go.md +16 -0
- package/knowledge/test-stack-profiles/node.md +16 -0
- package/knowledge/test-stack-profiles/react.md +12 -0
- package/knowledge/test-stack-profiles/spring-boot.md +16 -0
- package/knowledge/test-stack-profiles/ssr-htmx.md +14 -0
- package/knowledge/test-stack-profiles/vue.md +12 -0
- package/knowledge/test-strategy.md +70 -0
- package/knowledge/testability-patterns.md +240 -0
- package/knowledge/testing-quadrants.md +44 -0
- package/knowledge/testing-techniques/approval.md +15 -0
- package/knowledge/testing-techniques/chaos.md +17 -0
- package/knowledge/testing-techniques/fuzz.md +15 -0
- package/knowledge/testing-techniques/property-based.md +15 -0
- package/knowledge/testing-techniques/schema-validation.md +15 -0
- package/knowledge/testing-techniques/screenshot.md +15 -0
- package/knowledge/three-phase-workflow.md +198 -0
- package/knowledge/value-patterns.md +55 -0
- package/knowledge/verification-mode.md +116 -0
- package/knowledge/virtual-service-libraries.md +75 -0
- package/knowledge/wave-consolidation-guidance.md +21 -0
- package/overrides/agents/Explore.md +15 -0
- package/overrides/agents/general-purpose.md +10 -0
- package/overrides/notes/autoship.md +6 -0
- package/overrides/notes/issues-from-assessment.md +3 -0
- package/overrides/notes/issues-from-plan.md +3 -0
- package/overrides/notes/mutation-night-watch.md +3 -0
- package/overrides/notes/mutation-testing.md +3 -0
- package/overrides/notes/pr.md +7 -0
- package/overrides/notes/project-init.md +6 -0
- package/overrides/notes/setup.md +13 -0
- package/overrides/notes/specs.md +3 -0
- package/overrides/skills/headless-run/SKILL.md +45 -0
- package/overrides/skills/upgrade/SKILL.md +30 -0
- package/overrides/skills/version/SKILL.md +25 -0
- package/package.json +36 -0
- package/scripts/authoring_digest.py +93 -0
- package/scripts/autoship_discover.py +121 -0
- package/scripts/autoship_group.py +409 -0
- package/scripts/autoship_proposals.py +494 -0
- package/scripts/autoship_queue.py +291 -0
- package/scripts/autoship_reclaim.py +495 -0
- package/scripts/build_jobs.py +108 -0
- package/scripts/build_rollback_point.py +240 -0
- package/scripts/build_slice_scope.py +157 -0
- package/scripts/build_wave.py +109 -0
- package/scripts/build_wave_reconcile.py +252 -0
- package/scripts/build_worktree_baseref.py +113 -0
- package/scripts/check_agent_scope.py +117 -0
- package/scripts/check_agent_tool_mapping.py +213 -0
- package/scripts/check_review_agent_mcp_tools.py +317 -0
- package/scripts/check_security_assessment_mcp_tools.py +165 -0
- package/scripts/checkpoint_abort.py +502 -0
- package/scripts/claude_setup_review.py +438 -0
- package/scripts/codebase_recon.py +556 -0
- package/scripts/coverage_config.py +623 -0
- package/scripts/coverage_delta_steering.py +330 -0
- package/scripts/coverage_discovery_dotnet.py +315 -0
- package/scripts/coverage_discovery_java.py +742 -0
- package/scripts/coverage_discovery_js.py +546 -0
- package/scripts/coverage_gap_ranking.py +556 -0
- package/scripts/coverage_readiness.py +455 -0
- package/scripts/coverage_report_parse.py +521 -0
- package/scripts/detect_bdd_convention.py +252 -0
- package/scripts/eval_ablation.py +376 -0
- package/scripts/gherkin_analysis_coverage_gate.py +306 -0
- package/scripts/gherkin_cross_feature_duplicate_titles_gate.py +173 -0
- package/scripts/gherkin_effectiveness_rollup.py +238 -0
- package/scripts/gherkin_failure_path_gate.py +206 -0
- package/scripts/gherkin_feature_merge.py +720 -0
- package/scripts/gherkin_stub_gate.py +163 -0
- package/scripts/gherkin_stub_merge.py +479 -0
- package/scripts/git_origin_host.py +88 -0
- package/scripts/install-java-static-analysis.py +110 -0
- package/scripts/issue_deps.py +74 -0
- package/scripts/lib/_bdd_markers.py +28 -0
- package/scripts/lib/_gherkin_text.py +93 -0
- package/scripts/lib/_vendored_tree.py +70 -0
- package/scripts/lib/autoship_state.py +397 -0
- package/scripts/lib/claude_md_guard.py +226 -0
- package/scripts/lib/deterministic_recon.py +446 -0
- package/scripts/lib/mcp_tool_grants.py +211 -0
- package/scripts/lib/plan_parse.py +386 -0
- package/scripts/lib/review_result.py +84 -0
- package/scripts/lib/review_roster.py +86 -0
- package/scripts/lib/session_log/__init__.py +34 -0
- package/scripts/lib/session_log/__pycache__/__init__.cpython-314.pyc +0 -0
- package/scripts/lib/session_log/__pycache__/records.cpython-314.pyc +0 -0
- package/scripts/lib/session_log/classify.py +231 -0
- package/scripts/lib/session_log/corrections.py +194 -0
- package/scripts/lib/session_log/discovery.py +108 -0
- package/scripts/lib/session_log/records.py +218 -0
- package/scripts/lib/session_log/redact.py +76 -0
- package/scripts/lib/session_log/signals.py +373 -0
- package/scripts/lib/session_report_downstream.py +614 -0
- package/scripts/lib/session_report_maintainer.py +1273 -0
- package/scripts/lib/session_report_shared.py +262 -0
- package/scripts/lib/settings_hook_guard.py +157 -0
- package/scripts/lib/slug.py +33 -0
- package/scripts/lib/stub_extractors/__init__.py +82 -0
- package/scripts/lib/stub_extractors/_common.py +328 -0
- package/scripts/lib/stub_extractors/csharp.py +19 -0
- package/scripts/lib/stub_extractors/go.py +173 -0
- package/scripts/lib/stub_extractors/java.py +18 -0
- package/scripts/lib/stub_extractors/jsts.py +126 -0
- package/scripts/mutation_stack_sections.py +149 -0
- package/scripts/mutation_yield_steering.py +345 -0
- package/scripts/orchestrator.py +895 -0
- package/scripts/plan_gherkin_export.py +227 -0
- package/scripts/plan_waves.py +208 -0
- package/scripts/pr_close_keyword_lint.py +108 -0
- package/scripts/progress_guardian.py +888 -0
- package/scripts/recon_inventory.py +273 -0
- package/scripts/review_findings_log.py +93 -0
- package/scripts/run_invariants.py +124 -0
- package/scripts/select_lenses.py +640 -0
- package/scripts/session_report.py +486 -0
- package/scripts/set_autocompact_env.py +221 -0
- package/scripts/ship_resume_guard.py +135 -0
- package/scripts/ship_review_gate.py +63 -0
- package/scripts/specs_convention_marker.py +103 -0
- package/scripts/test_improve_resume.py +277 -0
- package/scripts/test_review_mechanics.py +958 -0
- package/scripts/token_efficiency_review.py +322 -0
- package/scripts/verdict_scope.py +285 -0
- package/scripts/verify_gherkin_quality_critic_isolation.py +296 -0
- package/scripts/verify_tier.py +157 -0
- package/skills/adr-tools/SKILL.md +118 -0
- package/skills/agent-readiness/SKILL.md +105 -0
- package/skills/agent-readiness/ai_friendly_analyzers.py +326 -0
- package/skills/agent-readiness/scanner.py +441 -0
- package/skills/agent-readiness/scorecard.yaml +88 -0
- package/skills/api-design/SKILL.md +115 -0
- package/skills/apply-fixes/SKILL.md +171 -0
- package/skills/apply-test-doubles/SKILL.md +321 -0
- package/skills/artifact-lifecycle/SKILL.md +127 -0
- package/skills/autoship/SKILL.md +1124 -0
- package/skills/benchmark/SKILL.md +105 -0
- package/skills/branch-workflow/SKILL.md +89 -0
- package/skills/browse/SKILL.md +184 -0
- package/skills/browser-testing/SKILL.md +62 -0
- package/skills/browser-testing/references/playwright-patterns.md +216 -0
- package/skills/build/SKILL.md +422 -0
- package/skills/build/references/static-self-heal.md +245 -0
- package/skills/careful/SKILL.md +72 -0
- package/skills/cd-test-architecture/SKILL.md +371 -0
- package/skills/ci-debugging/SKILL.md +105 -0
- package/skills/co-evolution-audit/SKILL.md +269 -0
- package/skills/code-review/SKILL.md +1015 -0
- package/skills/code-review/examples/aggregated-sample.json +56 -0
- package/skills/code-review/examples/sample-report.md +41 -0
- package/skills/code-review/output-format.md +478 -0
- package/skills/code-review/scripts/activation.py +86 -0
- package/skills/code-review/scripts/change_impact.py +357 -0
- package/skills/code-review/scripts/change_shape.py +372 -0
- package/skills/code-review/scripts/change_size.py +212 -0
- package/skills/code-review/scripts/changed_file_list.py +141 -0
- package/skills/code-review/scripts/closing_pass.py +187 -0
- package/skills/code-review/scripts/consolidate.py +277 -0
- package/skills/code-review/scripts/contract_failure_report.py +185 -0
- package/skills/code-review/scripts/dispatch_reconcile.py +66 -0
- package/skills/code-review/scripts/dispatch_waves.py +164 -0
- package/skills/code-review/scripts/finding_signature.py +446 -0
- package/skills/code-review/scripts/ledger.py +283 -0
- package/skills/code-review/scripts/partition.py +169 -0
- package/skills/code-review/scripts/render_tiered_findings.py +274 -0
- package/skills/code-review/scripts/repo_invariants.py +1066 -0
- package/skills/code-review/scripts/review_context_pack.py +306 -0
- package/skills/code-review/scripts/review_round_log.py +345 -0
- package/skills/code-review/scripts/review_value_coverage.py +297 -0
- package/skills/code-review/scripts/validate_review_output.py +467 -0
- package/skills/code-review/sliced-mode.md +205 -0
- package/skills/competitive-analysis/SKILL.md +191 -0
- package/skills/context-loading-protocol/SKILL.md +157 -0
- package/skills/continue/SKILL.md +90 -0
- package/skills/cost-report/SKILL.md +178 -0
- package/skills/coverage-baseline/SKILL.md +335 -0
- package/skills/coverage-baseline/references/multi-project-discovery.md +202 -0
- package/skills/coverage-delta/SKILL.md +181 -0
- package/skills/coverage-delta/references/mutation-gate.md +70 -0
- package/skills/design-doc/SKILL.md +95 -0
- package/skills/design-interrogation/SKILL.md +89 -0
- package/skills/design-it-twice/SKILL.md +91 -0
- package/skills/docker-image-audit/SKILL.md +108 -0
- package/skills/docker-image-audit/references/install-guide.md +64 -0
- package/skills/docker-image-audit/references/report-template.md +73 -0
- package/skills/docker-image-create/SKILL.md +185 -0
- package/skills/domain-analysis/SKILL.md +183 -0
- package/skills/domain-driven-design/SKILL.md +194 -0
- package/skills/exploratory-testing/SKILL.md +108 -0
- package/skills/explore/SKILL.md +51 -0
- package/skills/farley-score/SKILL.md +165 -0
- package/skills/feature-file-validation/SKILL.md +78 -0
- package/skills/feature-file-validation/references/validation-rules.md +115 -0
- package/skills/feedback-learning/SKILL.md +414 -0
- package/skills/fix/SKILL.md +450 -0
- package/skills/freeze/SKILL.md +68 -0
- package/skills/frontend-architecture/SKILL.md +113 -0
- package/skills/gherkin-derive/SKILL.md +630 -0
- package/skills/gherkin-public/SKILL.md +266 -0
- package/skills/governance-compliance/SKILL.md +150 -0
- package/skills/guard/SKILL.md +75 -0
- package/skills/handoff/SKILL.md +139 -0
- package/skills/handoff/references/summary-templates.md +242 -0
- package/skills/harness-audit/SKILL.md +751 -0
- package/skills/harness-audit/scripts/lesson_validate.py +386 -0
- package/skills/harness-audit/scripts/redundancy_criterion.py +188 -0
- package/skills/headless-run/SKILL.md +45 -0
- package/skills/headless-run/scripts/isolated_dispatch.py +381 -0
- package/skills/help/SKILL.md +72 -0
- package/skills/hexagonal-architecture/SKILL.md +85 -0
- package/skills/human-oversight-protocol/SKILL.md +224 -0
- package/skills/issues-from-assessment/SKILL.md +223 -0
- package/skills/issues-from-plan/SKILL.md +133 -0
- package/skills/legacy-code/SKILL.md +132 -0
- package/skills/mermaid-diagramming/SKILL.md +120 -0
- package/skills/mutation-night-watch/SKILL.md +154 -0
- package/skills/mutation-night-watch/references/scheduling.md +135 -0
- package/skills/mutation-testing/SKILL.md +396 -0
- package/skills/mutation-testing/references/languages/csharp-stryker-net.md +676 -0
- package/skills/mutation-testing/references/languages/go-go-mutesting.md +95 -0
- package/skills/mutation-testing/references/languages/java-pitest.md +77 -0
- package/skills/mutation-testing/references/languages/javascript-stryker.md +188 -0
- package/skills/mutation-testing/references/languages/python-mutmut.md +97 -0
- package/skills/mutation-testing/references/time-estimation.md +34 -0
- package/skills/mutation-testing/references/tool-detection.md +15 -0
- package/skills/mutation-testing/references/workflow-callers.md +23 -0
- package/skills/mutation-testing/scripts/__pycache__/xunit_v3_feature_detector.cpython-314.pyc +0 -0
- package/skills/mutation-testing/scripts/csharp_stryker_net_slice_runner.py +635 -0
- package/skills/mutation-testing/scripts/csharp_stryker_net_status_loop.py +525 -0
- package/skills/mutation-testing/scripts/csharp_stryker_net_wrapper.py +681 -0
- package/skills/mutation-testing/scripts/mutation_baseline_reuse.py +292 -0
- package/skills/mutation-testing/scripts/mutation_exclude_policy.py +268 -0
- package/skills/mutation-testing/scripts/mutation_feasibility_gate.py +463 -0
- package/skills/mutation-testing/scripts/mutation_kill_headless.py +331 -0
- package/skills/mutation-testing/scripts/mutation_kill_insert.py +199 -0
- package/skills/mutation-testing/scripts/mutation_kill_insert_python.py +150 -0
- package/skills/mutation-testing/scripts/mutation_kill_loop.py +869 -0
- package/skills/mutation-testing/scripts/mutation_kill_loop_python.py +949 -0
- package/skills/mutation-testing/scripts/mutation_kill_retry.py +592 -0
- package/skills/mutation-testing/scripts/mutation_kill_shared.py +620 -0
- package/skills/mutation-testing/scripts/mutation_nightwatch.py +462 -0
- package/skills/mutation-testing/scripts/mutation_nightwatch_stacks.py +425 -0
- package/skills/mutation-testing/scripts/mutation_report.py +743 -0
- package/skills/mutation-testing/scripts/mutation_report_cli.py +175 -0
- package/skills/mutation-testing/scripts/mutation_safety_gate.py +69 -0
- package/skills/mutation-testing/scripts/stryker_shard_pipeline.py +847 -0
- package/skills/mutation-testing/scripts/stryker_shard_setup.py +440 -0
- package/skills/mutation-testing/scripts/stryker_timeout_retry.py +142 -0
- package/skills/mutation-testing/scripts/xunit_v3_feature_detector.py +341 -0
- package/skills/performance-benchmark/SKILL.md +174 -0
- package/skills/performance-benchmark/examples/report-format.md +43 -0
- package/skills/performance-benchmark/references/benchmark-script.md +169 -0
- package/skills/performance-metrics/SKILL.md +265 -0
- package/skills/plan/SKILL.md +199 -0
- package/skills/plan/references/gherkin-persistence.md +43 -0
- package/skills/plan/references/plan-template.md +182 -0
- package/skills/pr/SKILL.md +289 -0
- package/skills/pr/scripts/gate_retry_state.py +368 -0
- package/skills/project-init/README.md +141 -0
- package/skills/project-init/SKILL.md +1197 -0
- package/skills/project-init/evals/evals.json +200 -0
- package/skills/project-init/references/capability-tools.md +55 -0
- package/skills/project-init/references/configs.md +221 -0
- package/skills/property-based-testing/SKILL.md +121 -0
- package/skills/property-based-testing/fixtures/invariant_fixture.py +15 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/README.md +42 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/fast-check/LICENSE +21 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/fast-check/README.md +263 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/fast-check/lib/cjs/fast-check.d.ts +5165 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/fast-check/lib/cjs/fast-check.js +12147 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/fast-check/lib/cjs/package.json +3 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/fast-check/lib/cjs/types57/fast-check.d.ts +5165 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/fast-check/lib/fast-check.d.ts +5165 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/fast-check/lib/fast-check.js +12011 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/fast-check/lib/rolldown-runtime-D7D4PA-g.js +13 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/fast-check/lib/types57/fast-check.d.ts +5165 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/fast-check/package.json +94 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/LICENSE +21 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/README.md +168 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/RandomGenerator-DcXj09Ch.d.ts +14 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/distribution/uniformBigInt.d.ts +15 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/distribution/uniformBigInt.js +38 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/distribution/uniformFloat32.d.ts +15 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/distribution/uniformFloat32.js +18 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/distribution/uniformFloat64.d.ts +15 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/distribution/uniformFloat64.js +22 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/distribution/uniformInt.d.ts +15 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/distribution/uniformInt.js +134 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/RandomGenerator-DcXj09Ch.d.ts +14 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/distribution/uniformBigInt.d.ts +15 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/distribution/uniformBigInt.js +37 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/distribution/uniformFloat32.d.ts +15 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/distribution/uniformFloat32.js +17 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/distribution/uniformFloat64.d.ts +15 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/distribution/uniformFloat64.js +21 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/distribution/uniformInt.d.ts +15 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/distribution/uniformInt.js +133 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/generator/congruential32.d.ts +7 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/generator/congruential32.js +44 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/generator/mersenne.d.ts +7 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/generator/mersenne.js +90 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/generator/xoroshiro128plus.d.ts +7 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/generator/xoroshiro128plus.js +80 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/generator/xorshift128plus.d.ts +7 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/generator/xorshift128plus.js +78 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/package.json +3 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/types/JumpableRandomGenerator.d.ts +16 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/types/JumpableRandomGenerator.js +0 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/types/RandomGenerator.d.ts +2 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/types/RandomGenerator.js +0 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/utils/generateN.d.ts +6 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/utils/generateN.js +8 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/utils/purify.d.ts +12 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/utils/purify.js +9 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/utils/skipN.d.ts +6 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/utils/skipN.js +6 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/generator/congruential32.d.ts +7 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/generator/congruential32.js +46 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/generator/mersenne.d.ts +7 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/generator/mersenne.js +92 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/generator/xoroshiro128plus.d.ts +7 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/generator/xoroshiro128plus.js +82 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/generator/xorshift128plus.d.ts +7 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/generator/xorshift128plus.js +80 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/types/JumpableRandomGenerator.d.ts +16 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/types/JumpableRandomGenerator.js +0 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/types/RandomGenerator.d.ts +2 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/types/RandomGenerator.js +0 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/utils/generateN.d.ts +6 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/utils/generateN.js +9 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/utils/purify.d.ts +12 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/utils/purify.js +10 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/utils/skipN.d.ts +6 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/utils/skipN.js +7 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/package.json +133 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/package-lock.json +1179 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/package.json +14 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/roundtrip.js +29 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/roundtrip.properties.test.js +16 -0
- package/skills/property-based-testing/fixtures/no_property_fixture.py +10 -0
- package/skills/property-based-testing/fixtures/roundtrip_fixture.py +16 -0
- package/skills/property-based-testing/references/languages/javascript.md +54 -0
- package/skills/property-based-testing/scripts/detect_and_dispatch.py +80 -0
- package/skills/property-based-testing/scripts/hypothesis_scaffold.py +276 -0
- package/skills/proxy-resilience/SKILL.md +84 -0
- package/skills/quality-gate-pipeline/SKILL.md +184 -0
- package/skills/quality-targets-converge/SKILL.md +254 -0
- package/skills/repo-review/SKILL.md +159 -0
- package/skills/report-pdf/SKILL.md +66 -0
- package/skills/review/SKILL.md +47 -0
- package/skills/review-agent/SKILL.md +152 -0
- package/skills/review-summary/SKILL.md +73 -0
- package/skills/run-report/SKILL.md +70 -0
- package/skills/semantic-duplication-scan/SKILL.md +337 -0
- package/skills/semantic-scan/SKILL.md +53 -0
- package/skills/semgrep-analyze/SKILL.md +139 -0
- package/skills/setup/SKILL.md +1122 -0
- package/skills/ship/SKILL.md +240 -0
- package/skills/source-verification/SKILL.md +210 -0
- package/skills/source-verification/scripts/claim_extractor.py +155 -0
- package/skills/specs/.size-baseline.json +4 -0
- package/skills/specs/SKILL.md +243 -0
- package/skills/specs/references/completeness-checklist.md +83 -0
- package/skills/specs/references/extraction.md +58 -0
- package/skills/specs/references/glossary.md +59 -0
- package/skills/specs/references/persistence.md +115 -0
- package/skills/specs/references/predictability-check.md +77 -0
- package/skills/static-analysis-integration/SKILL.md +235 -0
- package/skills/static-analysis-integration/adapters/_envelope.py +26 -0
- package/skills/static-analysis-integration/adapters/jscpd-adapter.py +66 -0
- package/skills/static-analysis-integration/adapters/lizard-adapter.py +81 -0
- package/skills/static-analysis-integration/adapters/mypy-adapter.py +50 -0
- package/skills/static-analysis-integration/adapters/mypy-src-layout.py +93 -0
- package/skills/static-analysis-integration/adapters/security-review-adapter.py +212 -0
- package/skills/static-analysis-integration/maintenance.md +23 -0
- package/skills/static-analysis-integration/references/language-setup.md +228 -0
- package/skills/static-analysis-integration/references/sarif-parser.md +124 -0
- package/skills/static-analysis-integration/references/security-review-adapter.md +118 -0
- package/skills/static-analysis-integration/references/tool-configs.md +617 -0
- package/skills/static-analysis-integration/rulesets/pmd-quickstart.xml +24 -0
- package/skills/stryker-xunit-v2-shim/SKILL.md +274 -0
- package/skills/stryker-xunit-v2-shim/references/shim-howto.md +256 -0
- package/skills/stryker-xunit-v2-shim/scripts/generate_shim.py +143 -0
- package/skills/systematic-debugging/SKILL.md +130 -0
- package/skills/telemetry/SKILL.md +75 -0
- package/skills/test-audit-disable/SKILL.md +129 -0
- package/skills/test-design/SKILL.md +177 -0
- package/skills/test-design/scripts/__pycache__/internal_double_detector.cpython-314.pyc +0 -0
- package/skills/test-design/scripts/internal_double_detector.py +631 -0
- package/skills/test-design-advisor/SKILL.md +166 -0
- package/skills/test-driven-development/SKILL.md +169 -0
- package/skills/test-health/SKILL.md +262 -0
- package/skills/test-improve/SKILL.md +239 -0
- package/skills/test-improve/references/phase-0-approach-contract.md +228 -0
- package/skills/test-improve/references/phase-1-analyze.md +131 -0
- package/skills/test-improve/references/phase-2-baseline.md +121 -0
- package/skills/test-improve/references/phase-3-derive-gherkin.md +53 -0
- package/skills/test-improve/references/phase-4-plan-fixes.md +34 -0
- package/skills/test-improve/references/phase-5-improve.md +215 -0
- package/skills/test-improve/references/phase-6-refactor-decision.md +45 -0
- package/skills/test-improve/references/phase-7-refactor.md +44 -0
- package/skills/test-improve/references/phase-8-validate.md +66 -0
- package/skills/test-improve/references/phase-9-close-out-prompt.md +11 -0
- package/skills/test-improve/references/phase-9-report.md +62 -0
- package/skills/test-improve/references/review-loop.md +92 -0
- package/skills/test-improve/templates/executive-summary.md +123 -0
- package/skills/threat-modeling/SKILL.md +108 -0
- package/skills/triage/SKILL.md +211 -0
- package/skills/ubiquitous-language/SKILL.md +192 -0
- package/skills/ubiquitous-language/scripts/collect_domain_signals.py +300 -0
- package/skills/unfreeze/SKILL.md +37 -0
- package/skills/upgrade/SKILL.md +31 -0
- package/skills/upgrade/scripts/check_version_drift.py +113 -0
- package/skills/upgrade/scripts/enable_autoupdate.py +149 -0
- package/skills/version/SKILL.md +25 -0
- package/sync/__pycache__/sync_upstream.cpython-314.pyc +0 -0
- package/sync/sync_upstream.py +293 -0
- package/templates/ACCEPTED-RISKS.md.tmpl +46 -0
- package/templates/agents/agent-template.md +151 -0
- package/templates/agents/angular-testing.md +66 -0
- package/templates/agents/csharp-quality.md +63 -0
- package/templates/agents/esm-enforcer.md +52 -0
- package/templates/agents/front-end-testing.md +65 -0
- package/templates/agents/go-quality.md +65 -0
- package/templates/agents/python-quality.md +62 -0
- package/templates/agents/react-testing.md +61 -0
- package/templates/agents/ts-enforcer.md +60 -0
- package/templates/agents/twelve-factor-audit.md +49 -0
- package/tools/entropy-check.py +250 -0
- package/tools/model-hash-verify.py +213 -0
|
@@ -0,0 +1,422 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: build
|
|
3
|
+
description: >-
|
|
4
|
+
Execute an approved implementation plan in small per-behavior batches.
|
|
5
|
+
Reads the plan, implements each step one behavior at a time in the
|
|
6
|
+
Code-First Small Batches cadence with a refactor on every green, runs
|
|
7
|
+
inline review checkpoints, and produces verification evidence. Use when
|
|
8
|
+
the user says "build this", "implement the plan", "start building", or
|
|
9
|
+
after /plan has been approved.
|
|
10
|
+
argument-hint: "[--plan <path>] [--yes] [--backstop-review=skip]"
|
|
11
|
+
user-invocable: true
|
|
12
|
+
allowed-tools: Read, Write, Edit, Glob, Grep, Bash, Agent, AskUserQuestion
|
|
13
|
+
---
|
|
14
|
+
|
|
15
|
+
# Build
|
|
16
|
+
|
|
17
|
+
Role: orchestrator. This command implements an approved plan — it does not create plans or specs.
|
|
18
|
+
|
|
19
|
+
You have been invoked with the `/build` command.
|
|
20
|
+
|
|
21
|
+
## Orchestrator constraints
|
|
22
|
+
|
|
23
|
+
**MUST — confirm agent-dispatch capability before any review-agent dispatch (issue #1461).** This skill dispatches review agents at three points: Step 3 (`spec-compliance-review` acceptance-criteria gate), sub-step 4 (per-step review checkpoint for `complex` steps), and sub-step 6 (slice review checkpoint). Before any of those dispatches, you MUST confirm the `Agent` (or `Task`) tool is actually present and available in your current toolset. If it is not: **STOP** at that point — do not self-apply the reviewer's checklist inline as a substitute for independent dispatch, and do not mark the gate/checkpoint passed on that basis. Report to the user/operator plainly: this review checkpoint cannot run in this environment because no agent-dispatch capability (`Agent`/`Task` tool) is available; name exactly what's missing; and state that the build cannot proceed past this gate (nor reach the Step 6 `/code-review` backstop, which independently enforces the same rule) until it is run from a session with that capability. This is a hard requirement, not a preference.
|
|
24
|
+
|
|
25
|
+
1. **Follow the plan exactly.** If the plan is wrong, stop and ask the user — do not deviate silently.
|
|
26
|
+
2. **Every step is small per-behavior batches, one agent.** The cadence is **Code-First Small Batches** — each behavior follows IMPLEMENT → TEST → REFACTOR (the statistically separated winner of the workflow experiments; `docs/experiments/RECOMMENDATIONS.md` Rec 3). There is no cadence to resolve and no per-plan opt-in (ADR 0017) — Code-First Small Batches is the only cycle a step follows. One agent writes the code, the test, and the refactor for a unit of work; the refactor runs on **every** green — never deferred to an end-of-build pass, never skipped, never made conditional on task size or complexity (Rec 4); tests are frozen during REFACTOR; and big-batch shapes are prohibited — never all the code then all the tests, never all the tests then all the code.
|
|
27
|
+
3. **Incremental.** Each step must leave the codebase in a working, committable state.
|
|
28
|
+
4. **Verification evidence required.** Paste fresh test output before claiming a step is done.
|
|
29
|
+
5. **Review checkpoints (granularity scales with complexity).** Run inline review (static self-heal pass first, then spec-compliance, then quality agents) per step for `complex` steps; batch `standard`/`trivial` steps into one review at the slice boundary. When `~/.claude/telemetry.json` consent is enabled (and `DEV_TEAM_REVIEW_VALUE` is not `off`), record each checkpoint's find/fix/no-op outcome to `.claude/metrics/review-value.jsonl`. The final `/code-review` is the backstop.
|
|
30
|
+
6. **Be concise.** Report step status and verification evidence, no narration.
|
|
31
|
+
7. **Diagnose before retry.** When any bash command run during the build fails (a script, a test run, a build-tooling invocation), read the exit code and error text and state a one-line cause hypothesis before correcting and re-issuing it. Never re-run a failed command unmodified. If the shallow diagnosis reveals a real defect rather than a shell-level mistake (bad path, typo, missing arg), escalate to Systematic Debugging per the Escalation section below.
|
|
32
|
+
|
|
33
|
+
## Parse Arguments
|
|
34
|
+
|
|
35
|
+
Arguments: $ARGUMENTS
|
|
36
|
+
|
|
37
|
+
- `--plan <path>`: Path to the plan file. If omitted, search `plans/` for the most recently modified plan with status `approved`.
|
|
38
|
+
- `--yes`: Auto-approve the build's approval gates (steps 2 and 3) without prompting (non-interactive opt-in).
|
|
39
|
+
- `--backstop-review=skip`: Suppress **only** Step 6's backstop `/code-review` (#1962). Legal **only** when the caller has its own review pass over a diff that strictly contains this build's — an enclosing orchestrator's end-of-phase review loop, not a human running `/build` directly. Default **off**: an unflagged `/build` runs Step 6 exactly as before. Everything else is untouched — the inline checkpoints (sub-steps 4/6), the full-suite run (Step 5), runtime verification (4.9), invariants (4.10), and the Farley Score (Step 7) all still run, so this narrows one duplicated layer rather than trading review for speed. When set, print an audit line into the build output naming the enclosing reviewer: `Step-6 backstop review skipped (--backstop-review=skip) — enclosing reviewer: <name>.` Never set it on a `/build` whose diff nothing else reviews.
|
|
40
|
+
|
|
41
|
+
**Interactivity.** The run is **non-interactive** when any of these hold: `--yes` was passed, `DEV_TEAM_AUTO_APPROVE=1` is set, or `DEV_TEAM_INTERACTIVE` is not `1` (pi sets it only when a human UI is attached; a tool shell never has a TTY, so `test -t 0` must not be used — the headless/CI/automation case). The approval gates in steps 2 and 3 use this: interactive runs prompt exactly as before; non-interactive runs auto-proceed and **record the bypass in the build output** rather than hanging.
|
|
42
|
+
|
|
43
|
+
## Steps
|
|
44
|
+
|
|
45
|
+
### 0. Context loading
|
|
46
|
+
|
|
47
|
+
Invoke the [Context Loading Protocol](../context-loading-protocol/SKILL.md) at the start of this task. It decides which agents and skills to load for the current phase, and sets the context budget before any implementation work begins. This is a lightweight read — it does not add agents to context; it decides the load order so context stays under the 40% ceiling throughout the build.
|
|
48
|
+
|
|
49
|
+
### 1. Find the plan
|
|
50
|
+
|
|
51
|
+
If `--plan` was provided, read that file. Otherwise, search `plans/` for `.md` files and find the most recently modified one with `**Status**: approved`. If no approved plan is found, tell the user: "No approved plan found. Run `/plan` first, then approve it."
|
|
52
|
+
|
|
53
|
+
### 1.5. JS project bootstrap gate
|
|
54
|
+
|
|
55
|
+
Before any implementation step runs, make sure a JS-flavored plan has a project to build in. A fresh JS project with no `package.json` will fail the first test run (no test runner, no scripts), so bootstrap it first.
|
|
56
|
+
|
|
57
|
+
1. **Check for `package.json`** in the working directory. If it exists, **skip this gate silently** and continue to Step 2.
|
|
58
|
+
2. **Check the plan file for JS/TS signals** — any of `.js`, `.mjs`, `.ts`, `.jsx`, `.tsx`, `node`, `npm`, `vitest`, `jest`, `eslint`. If none are present, the plan is non-JS: **skip this gate silently** and continue to Step 2.
|
|
59
|
+
3. **Bootstrap** (only when `package.json` is absent **and** the plan is JS-flavored): print exactly one line — `No package.json found for a JS plan — bootstrapping with project-init.` — then invoke the `project-init` skill.
|
|
60
|
+
4. **Halt on failure.** If `project-init` fails, **stop `/build`** and report the failure — do not proceed to Step 2 or any implementation step.
|
|
61
|
+
|
|
62
|
+
The user sees no more than one line of output before `project-init` runs, and nothing at all when the gate is skipped.
|
|
63
|
+
|
|
64
|
+
### 2. Verify plan status
|
|
65
|
+
|
|
66
|
+
Read the plan file. If the status is not `approved`:
|
|
67
|
+
|
|
68
|
+
- **Interactive** → ask the user: "This plan has status '<status>'. Approve it before building, or continue anyway?"
|
|
69
|
+
- **Non-interactive** (see Parse Arguments) → do **not** block. Auto-approve and continue, and print an explicit audit line into the build output: `Auto-approved plan status '<status>' (non-interactive) — no human gate. Trigger: <--yes | DEV_TEAM_AUTO_APPROVE=1 | no TTY>.`
|
|
70
|
+
|
|
71
|
+
Either path appends an `approval` entry to `.claude/metrics/config-changelog.jsonl` per the
|
|
72
|
+
[human-oversight-protocol § Audit trail](../human-oversight-protocol/SKILL.md#audit-trail)
|
|
73
|
+
schema — `proposed` states the plan status being approved, `evidence_shown` points at
|
|
74
|
+
the plan file, `risks_surfaced` is `[]` unless the plan's status itself signals a risk
|
|
75
|
+
(e.g. resuming an `in-progress` plan). The non-interactive path writes the same three
|
|
76
|
+
fields; `description` names the bypass trigger.
|
|
77
|
+
|
|
78
|
+
### 3. Verify acceptance criteria (gate)
|
|
79
|
+
|
|
80
|
+
**Dispatch-capability gate (re-confirm here — issue #1461).** Before dispatching `spec-compliance-review` below, re-verify the `Agent`/`Task` tool is present. If it is not, STOP per the Orchestrator constraints above — do not self-verify the criteria inline; report the missing capability and halt before implementation begins.
|
|
81
|
+
|
|
82
|
+
Before implementation begins, dispatch the `spec-compliance-review` agent (by `subagent_type`) in **criteria verification mode**: pass it the plan's acceptance criteria and per-step test expectations, and ask it to evaluate the criteria themselves (see below) rather than a diff.
|
|
83
|
+
|
|
84
|
+
The reviewer evaluates each criterion for:
|
|
85
|
+
|
|
86
|
+
- **Specificity**: Could two developers independently verify this criterion and agree on pass/fail?
|
|
87
|
+
- **Testability**: Can this criterion be validated with a test or observable output?
|
|
88
|
+
- **Completeness**: Are edge cases and error conditions addressed?
|
|
89
|
+
|
|
90
|
+
If any criteria are flagged:
|
|
91
|
+
|
|
92
|
+
1. Present the findings to the user with the reviewer's suggested improvements
|
|
93
|
+
2. **Interactive** → Ask: "Revise these criteria before building, or proceed anyway?"
|
|
94
|
+
- If the user overrides, log the override in the build output and continue
|
|
95
|
+
- If the user revises, update the plan file and re-verify
|
|
96
|
+
3. **Non-interactive** (see Parse Arguments) → do **not** block. Proceed and record the bypass in the build output: `Acceptance-criteria gate auto-passed with N flagged criterion(s) (non-interactive) — no human gate. Trigger: <--yes | DEV_TEAM_AUTO_APPROVE=1 | no TTY>.` Include the flagged findings in the record so the bypass is auditable.
|
|
97
|
+
|
|
98
|
+
Whichever path is taken (proceed, revise, or override), append an `approval` entry to
|
|
99
|
+
`.claude/metrics/config-changelog.jsonl` per the [human-oversight-protocol § Audit trail](../human-oversight-protocol/SKILL.md#audit-trail)
|
|
100
|
+
schema — `proposed` is the acceptance-criteria set under review, `evidence_shown`
|
|
101
|
+
points at the plan file (and, for an interactive override, the reviewer's findings if
|
|
102
|
+
written to `.claude/memory/`), and `risks_surfaced` lists the flagged criteria (`[]` if none
|
|
103
|
+
were flagged). An interactive `override` (user overrides the reviewer's findings) is
|
|
104
|
+
logged as `type: "override"` instead, with `proposed` recording the reviewer's
|
|
105
|
+
flagged concern and `description` recording the user's decision to proceed anyway.
|
|
106
|
+
|
|
107
|
+
### 4. Implement each step
|
|
108
|
+
|
|
109
|
+
Work the plan **wave by wave** (the plan's `## Parallelization` section, derived by `scripts/plan_waves.py`). Within a wave, independent slices may build concurrently; across waves a barrier holds the next wave until the current one reconciles green.
|
|
110
|
+
|
|
111
|
+
**Base-ref check (top-level session, before any subagent dispatch).** Worktree subagents (`isolation: "worktree"`) must branch from the caller's local HEAD, not `origin/<default>`, so the `docs/specs/<slug>.md` (when the spec was file-persisted — see `/specs`' GitHub-issue persistence path for the alternative, which leaves no local file to miss) and `plans/<slug>.md` files `/ship` just produced are visible to them (issue #553). This is controlled by Claude Code's `worktree.baseRef` setting, which is honored only at project (`.claude/settings.json`) or user (`~/.claude/settings.json`) scope — **not** at plugin or project-local scope (`docs/spikes/worktree-baseref-head-spike.md`). `/build` cannot set this on the user's behalf, so it runs a read-only detect-and-warn **in this top-level `/build` session, before any subagent dispatch**, so the warning is visible in the human-facing transcript:
|
|
112
|
+
|
|
113
|
+
```bash
|
|
114
|
+
python3 "${CLAUDE_PLUGIN_ROOT}/scripts/build_worktree_baseref.py" detect # prints head|fresh|unset|unknown
|
|
115
|
+
```
|
|
116
|
+
|
|
117
|
+
- **`head`** → no warning. Proceed.
|
|
118
|
+
- **`fresh`**, **`unset`**, or **`unknown`** (detection failed — e.g. `jq` unavailable; treated fail-safe) → unless `DEV_TEAM_WORKTREE_BASE_FRESH=1` is set, print a loud warning naming the exact file to edit and a paste-ready snippet, then continue:
|
|
119
|
+
|
|
120
|
+
```
|
|
121
|
+
⚠ worktree.baseRef is not "head" (detected: <value>) — worktree subagents
|
|
122
|
+
will branch from origin/<default>, not your current HEAD. Any
|
|
123
|
+
uncommitted-to-origin spec/plan/WIP files will be invisible to them.
|
|
124
|
+
|
|
125
|
+
Add this to .claude/settings.json (project) or ~/.claude/settings.json
|
|
126
|
+
(user) — plugin and project-local settings.json are NOT honored for
|
|
127
|
+
this key:
|
|
128
|
+
|
|
129
|
+
{ "worktree": { "baseRef": "head" } }
|
|
130
|
+
|
|
131
|
+
To keep fresh-from-origin worktrees deliberately, set
|
|
132
|
+
DEV_TEAM_WORKTREE_BASE_FRESH=1 to silence this warning.
|
|
133
|
+
```
|
|
134
|
+
|
|
135
|
+
If detection returned `unknown`, the warning additionally states that `worktree.baseRef could not be detected`.
|
|
136
|
+
|
|
137
|
+
**`/build` never mutates a settings file.** The check is read-only end to end — it never writes `.claude/settings.json`, `~/.claude/settings.json`, or any other settings file. There is nothing to restore and no crash-recovery surface.
|
|
138
|
+
|
|
139
|
+
**Resolve the wave schedule and concurrency first:**
|
|
140
|
+
|
|
141
|
+
```bash
|
|
142
|
+
python3 "${CLAUDE_PLUGIN_ROOT}/scripts/build_wave.py" <plan-file> # ordered waves + members
|
|
143
|
+
python3 "${CLAUDE_PLUGIN_ROOT}/scripts/build_jobs.py" --wave-width <W> [--jobs N] # effective concurrency
|
|
144
|
+
```
|
|
145
|
+
|
|
146
|
+
`build_jobs.py` resolves `min(--jobs, DEV_TEAM_MAX_PARALLEL_BUILDS, wave width)`. Parallel fan-out is **opt-in**: when neither `--jobs` nor `DEV_TEAM_MAX_PARALLEL_BUILDS` is set, the **default is sequential** (effective 1) — fan-out never *saves* tokens, it trades them for wall-clock (#1515). Opt in with `--jobs N` (bounded only by wave width) or `DEV_TEAM_MAX_PARALLEL_BUILDS` (an explicit env value is honored verbatim, never re-capped); non-positive/non-integer values clamp to 1. **Sequential fallback:** when effective concurrency is **1** (the unset default, a fully-dependent plan, `--jobs 1`, or max 1), build slices one at a time in a single worktree in dependency order — **no worktree fan-out, no reconcile step** (today's behavior exactly).
|
|
147
|
+
|
|
148
|
+
**Concurrent dispatch (effective concurrency > 1):**
|
|
149
|
+
|
|
150
|
+
1. Dispatch each independent slice in the wave to its **own** git worktree (`isolation: "worktree"`), up to the effective concurrency. Each slice's changes stay isolated until reconcile, and each slice still runs its full per-behavior cycle and inline review gates.
|
|
151
|
+
2. **Report the concrete level and cost** — name the slice count and the resulting multiplier, e.g. *"building wave 2 — 5 slices concurrently; ~5× wall-clock speedup, but burns your token budget faster — roughly 5× for 5 concurrent slices."* The cost is reported, never auto-throttled: cap it yourself with `--jobs N` or `DEV_TEAM_MAX_PARALLEL_BUILDS` if the burn rate is too high (#1170).
|
|
152
|
+
3. **Barrier + reconcile** once the wave's slices finish: `build_wave_reconcile.py --into <integration> --base <ref> --test-cmd "<full suite>" <slice-branch>...` merges them order-independently and gates on the full suite before any next-wave slice starts.
|
|
153
|
+
4. **Loud halt, never silent:**
|
|
154
|
+
- A **failing slice** → exit non-zero naming it, list which same-wave slices succeeded and where their (preserved) worktrees are, print the resume command, and start no next-wave slice. Resume rebuilds only the incomplete slice.
|
|
155
|
+
- A **reconcile conflict** (two same-wave diffs touch one file) → exit non-zero naming the file, pick no side, start no next-wave slice.
|
|
156
|
+
|
|
157
|
+
**Slice dispatch bookkeeping (issue #865).** Before a slice's first step begins:
|
|
158
|
+
|
|
159
|
+
1. **Freeze scope (opt-in only).** Check whether the plan opts into scope enforcement: `python3 ${CLAUDE_PLUGIN_ROOT}/scripts/build_slice_scope.py enabled <plan-file>` (exit 0 = engaged). Declaring slice-level `**Files:**` alone never freezes anything — only a `**Scope enforcement:** freeze` metadata line does (Ambiguity Log Q1). When engaged **and** the dispatching slice declares `**Files:**`, run `python3 ${CLAUDE_PLUGIN_ROOT}/scripts/build_slice_scope.py engage <plan-file> --slice <id> --hooks-dir <worktree>/.claude/hooks` — this writes `.claude/hooks/freeze-state.json` with `allowed_patterns` set to the slice's declared paths plus the fixed bookkeeping allowlist (the plan file, `.claude/memory/**`, `.claude/metrics/**`, and the AC3-exempt `metrics/verify-log.jsonl`), so `hooks/pre_tool_guard.py` blocks any Write/Edit outside that scope without also blocking `/build`'s own progress writes. The `--hooks-dir` value must point at `<worktree>/.claude/hooks`, not `<worktree>/hooks` — `hooks/pre_tool_guard.py` reads the freeze state via `hooks/lib/artifact_paths.py`'s per-repo `.claude/hooks/freeze-state.json` convention (issue #1890), and a mismatched write location means the freeze is silently never enforced. Clear it at slice completion (sub-step 5 below): `python3 ${CLAUDE_PLUGIN_ROOT}/scripts/build_slice_scope.py clear --hooks-dir <worktree>/.claude/hooks`.
|
|
160
|
+
2. **Rollback point.** When the slice declares `**Rollback point:**`, resolve the symbolic value to a concrete SHA and record it: `python3 ${CLAUDE_PLUGIN_ROOT}/scripts/build_rollback_point.py resolve --symbolic <value> --repo <worktree> --slice-start <HEAD-at-dispatch> --wave-start <wave-start-ref> --plan-start <plan-start-ref>`, then `python3 ${CLAUDE_PLUGIN_ROOT}/scripts/build_rollback_point.py record --path .claude/memory/build-rollback.json --slice <id> --symbolic <value> --sha <resolved-sha>`. This is the boundary a dead-end escalation (issue #864) names verbatim: "revert to `<sha>` (`<symbolic>`)" — retrieve it with the script's `get` subcommand. A slice without `Rollback point` records nothing.
|
|
161
|
+
|
|
162
|
+
For each step within a slice, dispatch the `software-engineer` agent (by `subagent_type`) scoped to a single unit of work. Pass it its step **and the slice's Gherkin scenario(s)** — the scenarios are the behavioral contract the step's test must satisfy — plus an explicit constraint: **the design is settled; do not design.** The agent implements the step exactly as planned, one behavior at a time, per the per-behavior cycle below — it does not revisit or re-derive design decisions the plan already made. **Require its step report to state assumptions explicitly**: when the plan under-specifies a detail that doesn't rise to an escalation (exact error-message wording, an unspecified boundary condition, a choice between two equally plan-consistent shapes), the agent records the decision and its basis in the report rather than resolving it silently — `/pr`'s "Decisions & assumptions" section (`skills/pr/SKILL.md`) collects these from each step's report.
|
|
163
|
+
|
|
164
|
+
**Authoring digest (#2209).** Append the step's write-time checklist to the dispatch prompt: `python3 "${CLAUDE_PLUGIN_ROOT}/scripts/authoring_digest.py" --files <the step's planned files>` — it resolves the same `select_lenses.py` lenses the review checkpoints use and prints only their `## Authoring checklist` bullets (~500 tokens for a typical diff, empty when no lens has one). Paste the output verbatim under a `## Authoring checklist (write-time)` heading; an empty digest adds nothing. Re-use the same call for review-fix correction prompts, scoped to the files being fixed. Planned files come from the step's `**Files:**` (or the slice's); when neither is declared, skip the digest rather than guessing.
|
|
165
|
+
|
|
166
|
+
Within the per-behavior mini-cycle below, repeated Write/Edit calls can race a `PostToolUse` hook that rewrites files (e.g., a formatter): an `Edit` failing on a stale `old_string` is expected to self-correct by re-`Read`ing the file before the next `Edit` attempt, not to escalate immediately.
|
|
167
|
+
|
|
168
|
+
**Phase-state bookkeeping (guard input).** `/build` owns `.claude/memory/build-phase.json` as mechanical step bookkeeping: write `{"phase": "<implement|test|refactor>", "step": "<N.M>", "written_at": "<ISO8601>", "test_files_staged": [], "plan_path": "<repo-relative path of the plan file>"}` at **each** phase transition. At step completion do **not** delete the file: rewrite it with `"phase": "between-steps"`, `"step": "<next N.M>"`, a fresh `written_at`, `"test_files_staged": []` and the same `plan_path`, so a compaction between steps still restores the plan path and next step (the guards ignore any phase other than `refactor`). Clear the file only when the plan's last step completes ("cleared" = absent, empty, or `{}`); the 4-hour staleness rule still bounds a crashed run. `plan_path` lets the post-compaction re-injection hook find the active plan without globbing; the one shared reader is `hooks/lib/build_state.py`, pinned to this schema by `tests/hooks/test_build_state_contract.py`. At the **TEST → REFACTOR transition**, additionally stage the step's test files — `git add` them, including new/untracked ones — and record their paths in `test_files_staged`: the index becomes the refactor baseline the `refactor_test_freeze_guard` / `refactor_test_revert_guard` hooks enforce the tests-frozen invariant against. **Standalone-dispatch fallback:** when `software-engineer` is dispatched directly in an isolated worktree and no `/build` session is writing this record for it (the record is absent, or its `phase` is not `refactor`, at REFACTOR entry), the dispatched agent writes `.claude/memory/build-phase.json` itself before entering REFACTOR — entering REFACTOR with no phase record present silently disables the tests-frozen guard (`refactor_test_freeze_guard.py` treats an absent/non-`refactor` record as "nothing to enforce").
|
|
169
|
+
|
|
170
|
+
Work each step **one behavior at a time** — never all the code then all the tests, never all the tests then all the code:
|
|
171
|
+
|
|
172
|
+
1. **First phase — IMPLEMENT.** Implement exactly one behavior from the step — no cleanup, no behavior beyond what the step requires. Apply the software engineer's [Per-Edit Authoring Discipline](../../agents/software-engineer.md#per-edit-authoring-discipline) checklist (Surgical Changes, Simplicity First, Think Before Coding) at this phase, not deferred to review.
|
|
173
|
+
|
|
174
|
+
**Explore before editing.** Before writing the behavior's code, prefer CodeGraph/Repowise over a raw Grep/Read sweep — one call returns the relevant source plus its callers and blast radius, which is what the Per-Edit Authoring Discipline checklist needs to size the change correctly. See [`knowledge/codegraph-vs-graphify.md`](../../knowledge/codegraph-vs-graphify.md) for tool selection and the fallback contract.
|
|
175
|
+
2. **Second phase — TEST.** Write the test covering the behavior's slice scenario, immediately after the code. Before writing a test double for this behavior, check the collaborator against `${CLAUDE_PLUGIN_ROOT}/knowledge/internal-collaborator-doubling.md#the-three-blockers-exhaustive`'s blocker table. Run the full test suite. **Hard gate: all tests must pass — paste the passing output.** Do NOT proceed to REFACTOR without pasted passing output.
|
|
176
|
+
|
|
177
|
+
**Before each repair iteration** (here and in the review-fix loop, sub-step 4), read `${CLAUDE_PLUGIN_ROOT}/knowledge/failure-routing.md` and classify the failing output/exit code by its regex table — deterministic pattern match only, no LLM call, no extra dispatch. Follow the matched route (inline fix / systematic-debugging / test-generation / security-engineer dispatch / human arbitration); `unclassified` falls through to the generic loop below, unchanged. A route switch spends from the same iteration budget — it never resets or raises the cap.
|
|
178
|
+
|
|
179
|
+
**2a. Repair loop on failure — failure-signature dead-end detection (issue #864).** Whichever route the classification above sends the failure down, repair it in place rather than handing back a bare failure:
|
|
180
|
+
|
|
181
|
+
- **Compute a failure signature after every repair iteration** (an edit followed by a re-run): the pair of (1) the sorted, deduplicated set of failing test identifiers, using the runner's native IDs (pytest node IDs, jest/vitest full test names, `go test` names, etc.), and (2) the error class per failing test (assertion failure vs. exception type vs. compile/collection error, e.g. `AssertionError`, `TypeError`, `SyntaxError`).
|
|
182
|
+
- **Normalize before comparing.** Strip volatile output first — timestamps, durations, memory addresses, temp paths, PIDs/ports, random seeds — so two runs identical except for that noise produce the same signature. Never compare raw output.
|
|
183
|
+
- **Track signatures as in-context iteration state** — a small per-step table (iteration → signature) held for the duration of this repair loop. This is not a `.claude/memory/` file; the durable record on dead-end is the checkpoint commit below (plus the existing `.claude/memory/build-escalation-<plan-slug>.md` record on a non-interactive halt).
|
|
184
|
+
- **Two identical consecutive signatures is a dead-end.** If iteration N+1's normalized signature equals iteration N's, stop — do not dispatch a third attempt against the unchanged signature.
|
|
185
|
+
- **A changed signature continues repair normally.** Fewer or different failing tests, or a different error class, is progress: keep repairing, and restart the dead-end comparison from the new signature. **No new iteration cap** — the review loop's 5-iteration cap (sub-step 4) is untouched; this is a no-progress cutoff, not a count cap, so a repair loop that keeps changing its signature may run as long as it keeps progressing. A route switch (per the classification above) spends from this same budget — it never resets or raises it.
|
|
186
|
+
- **On dead-end, commit a checkpoint before escalating.** Commit the current working tree as-is (no per-iteration snapshots in v1) on the working branch — **never `main`** — with a conventional message explicitly marked as a dead-end checkpoint, e.g. `chore(build): dead-end checkpoint — step <N>, <M> tests still failing`. If an earlier iteration was strictly better than the current one, name that regression in the escalation rather than reverting to it.
|
|
187
|
+
- **Escalate with the best candidate, not a bare failure**, stating all three: (a) **improved** — tests that were failing at repair start and now pass, (b) **remaining** — the current (unchanged) failing signature, (c) the **checkpoint commit ref**.
|
|
188
|
+
- **Cite the architecture-questioning rule at 3+ failed attempts.** Count every repair iteration that ended with a real edit and a re-run that failed to reach green (regardless of whether its signature changed) as one failed fix attempt. When 3 or more distinct fix attempts have failed by the time the dead-end fires, the escalation must explicitly cite [Systematic Debugging](../systematic-debugging/SKILL.md)'s rule: "After 3+ failed fix attempts, question the architecture — stop patching."
|
|
189
|
+
- **This is a hard stop, matching the Escalation section below**: leave plan status unchanged and never proceed to `/pr` over the unresolved escalation. A red checkpoint commit is never presented as done.
|
|
190
|
+
- **Out of scope / unchanged**: `hooks/verify_guard.py` is not modified and continues to own the separate, syntactic case — the same verify command re-run with zero intervening edits. This repair loop fires only when edits *do* happen but the failure signature doesn't change.
|
|
191
|
+
|
|
192
|
+
3. **REFACTOR (every green, never skipped).** Clean up structure, naming, duplication without changing behavior. Runs in **every** per-behavior cycle: never deferred to an end-of-build pass, never made conditional on task size or complexity (`docs/experiments/RECOMMENDATIONS.md` Rec 4 — deleting just this step erased the cadence's changeability advantage entirely). **Tests are frozen for the phase** — a refactor must never change a test (enforced by the freeze/revert guards; recovery: return to the TEST phase, change the test there, re-verify green, re-enter REFACTOR). Run tests again — they must still pass. If tests break, undo and try a smaller change. A no-op refactor (nothing worth changing, stated in one line) satisfies the phase — the mandate is the check on every green, not a diff — and any refactor made stays within the code the step touched; adjacent-file cleanups are follow-ups, not refactors.
|
|
193
|
+
|
|
194
|
+
**Self-verification, mandatory (#2107).** Before this step may be marked done (sub-step 5), also run the project's lint and type-check tools — whichever apply to its stack (`ruff`/ESLint/`tsc`/`pmd`/`dotnet build`, etc.) — when the project has them, and confirm they pass; paste the output alongside the test run. This closes a real gap the TEST phase's test-only hard gate leaves open: a green suite proves behavior, not that the REFACTOR pass above left no unused import, type error, or style violation behind — those otherwise surface only downstream, in Step 6's backstop review, after the step already reads `[x]`. Same bar `quality-gate-pipeline`'s Phase 2 ("Required Evidence") already states for every completion claim; this is that bar made explicit and mandatory at the one moment in this cadence it was previously only implied. Skip only what the project genuinely has none of (state which, in one line) — never skip because it "should still be clean."
|
|
195
|
+
4. **Inline review checkpoint — granularity scales with complexity.** *Where* the checkpoint runs depends on the step's **Complexity** classification (review *depth* still scales too):
|
|
196
|
+
- **trivial**: Skip inline review. The final `/code-review` (step 6) covers all modified files.
|
|
197
|
+
- **standard**: **Defer** review to the slice boundary (sub-step 6) — do not review now. Track the step's changed files so the slice checkpoint reviews them in one batch. Per-step review on standard steps is N near-identical passes where one at slice end largely does the same work, and the final `/code-review` (step 6) remains the backstop. This is the batching win — fewer review dispatches per multi-step slice at bounded quality risk.
|
|
198
|
+
- **complex**: **Dispatch-capability gate (re-confirm here — issue #1461):** before this review dispatches, re-verify the `Agent`/`Task` tool is present. If it is not, STOP per the Orchestrator constraints above — do not self-apply this checkpoint's checklist inline; report the missing capability and halt rather than marking the step done. Otherwise, review **now, per step** — smaller blast radius per fix. Run the static self-heal pass to completion first — pass, or cap-and-escalate, per `references/static-self-heal.md` — then `/review-agent spec-compliance-review --internal`, then **the quality lenses the resolver selects for this step's changed files** — `python3 ${CLAUDE_PLUGIN_ROOT}/scripts/select_lenses.py --files <step's changed files>` returns the applicable lenses (see the resolver note after sub-step 6) — dispatched **cheap-first (non-opus lenses before the opus-tier `security-review`/`domain-review`/`arch-review`)**, with the review-fix loop (up to 5 iterations per `${CLAUDE_PLUGIN_ROOT}/knowledge/three-phase-workflow.md#review-loop`) — including its deterministic-first re-verification triage (`../code-review/SKILL.md` step 6a, #1610): before re-dispatching an agent to confirm a fix, prefer a close via whichever language-appropriate lint/type-check tool(s) applies to this project's own stack (not just Python's `ruff` — ESLint/`tsc` for JS/TS, `pmd` for Java, `dotnet format`/`dotnet build` for C#, etc.) plus the test suite/`grep`, when the fix is mechanical and the claim is deterministically checkable, and escalate to an agent only when semantic judgment is genuinely needed. **Ledger-scoped dispatch (#2167).** Before dispatching those lenses, narrow their file sets and append the dispatch-prompt scope marker — see the shared **Ledger-scoped dispatch (#2167)** rule stated once after sub-step 6 below (applies to both checkpoint fix loops). **Abort check (#2168).** Before dispatching the remaining opus-tier lenses, run `checkpoint_abort.py`'s abort check — see the shared **Abort check (#2168)** rule stated once after sub-step 6 below (applies to both checkpoint fix loops). Before each review-fix iteration, classify the finding/failure via `${CLAUDE_PLUGIN_ROOT}/knowledge/failure-routing.md` and follow its route (see the TEST-phase note above) — a security-finding class dispatches security-engineer, a reviewer-conflict class routes to human arbitration, `unclassified` stays in the generic loop. Escalate to user if the loop doesn't converge. Then **record the checkpoint outcome** (sub-step 7).
|
|
199
|
+
- If no complexity is specified, default to **standard**.
|
|
200
|
+
- **UI changes (any complexity)**: After the relevant review passes (per-step for complex, at the slice checkpoint for standard), run browser verification via `/browse` in automated smoke test mode. Skip with warning if the dev server is not running. See `${CLAUDE_PLUGIN_ROOT}/knowledge/three-phase-workflow.md#phase-3-implement` Stage 3.
|
|
201
|
+
5. **Mark step done** — Use the Edit tool to update the plan file's `## Build Progress` section on disk:
|
|
202
|
+
- **Do not flip a checkbox on a "should work"/"should be fixed" impression.** The self-verification evidence required in sub-step 3 above (tests, and lint/type-check where the project has them) must be fresh from this session — not recalled from earlier in the conversation, not assumed — before this bullet may fire. This is the same completion-claim bar `quality-gate-pipeline`'s Phase 2 states generally, made mandatory at this specific step-completion moment (#2107).
|
|
203
|
+
- Change `- [ ] Step N.M: <title>` to `- [x] Step N.M: <title>` for the completed step.
|
|
204
|
+
- When every step under a slice is `[x]`, that is not the same as the slice being done — check off the parent `- [ ] Slice N: <title>` only after sub-steps 4.9 (runtime verification) and 4.10 (invariants) both pass, if applicable; a slice with no runtime surface and no declared invariants has nothing further to wait on and may be checked off once its steps and review checkpoint(s) are done.
|
|
205
|
+
- After all slices are `[x]`, change `**Status**: approved` to `**Status**: in-progress`.
|
|
206
|
+
- **Phase-state handoff.** Rewrite `.claude/memory/build-phase.json` as the `between-steps` record for the next unchecked step (see Phase-state bookkeeping above); delete it instead when this was the plan's last step.
|
|
207
|
+
- This disk write is the durable commit. If a `/clear` occurs, `/continue` reads `## Build Progress` to determine the resume point without needing conversation history.
|
|
208
|
+
- **Clear freeze scope (issue #865).** When every step under the slice is `[x]` and freeze was engaged for it (dispatch bookkeeping above), run `python3 ${CLAUDE_PLUGIN_ROOT}/scripts/build_slice_scope.py clear --hooks-dir <worktree>/.claude/hooks` before starting the next slice. A slice that never engaged freeze has nothing to clear.
|
|
209
|
+
6. **Slice review checkpoint (batched).** **Dispatch-capability gate (re-confirm here — issue #1461):** before this checkpoint dispatches, re-verify the `Agent`/`Task` tool is present. If it is not, STOP per the Orchestrator constraints above — do not self-apply the batched checkpoint's checklist inline; report the missing capability and halt rather than checking off the slice. Otherwise, when every step under the current slice is `[x]` **and** the slice had any deferred `standard` (or unspecified) steps, run **one** review pass over the slice's accumulated changed files: the static self-heal pass first (`references/static-self-heal.md`), then `/review-agent spec-compliance-review --internal`, then **the quality lenses the resolver selects** — `python3 ${CLAUDE_PLUGIN_ROOT}/scripts/select_lenses.py --files <slice's accumulated changed files>` — dispatched **cheap-first (non-opus before opus-tier)**, **plus `refactor-opportunity-review` dispatched by name** (#1976). That lens declares `Scope: on-demand`, so the resolver no longer returns it; this once-per-slice dispatch is its post-GREEN home, and it is the ONLY place `/build` runs it — do not also dispatch it per behavior at the REFACTOR phase (that would spend more, not less, than the per-diff panel slot it replaced) and do not treat the resolver's silence as "this lens was dropped". Its subject is a slice's accumulated shape — semantic vs. structural duplication across everything the slice touched — which is visible here and not in any single behavior's diff. **Ledger-scoped dispatch (#2167).** Before dispatching, narrow file sets and append the dispatch-prompt scope marker — see the shared rule below (applies to both checkpoint fix loops above). **Abort check (#2168).** Before dispatching the remaining opus-tier lenses, run `checkpoint_abort.py`'s abort check on the cheap-tier lenses' (from the cheap-first dispatch above) results — see the shared rule below (applies to both checkpoint fix loops above). Either way, apply the same review-fix loop (up to 5 iterations; escalate if it doesn't converge). `trivial`-only and all-`complex` slices have nothing to batch — skip this pass. Then **record the checkpoint outcome** (sub-step 7).
|
|
210
|
+
|
|
211
|
+
**Abort check (#2168) — applies to both checkpoint fix loops above (sub-steps 4 and 6).** Once the cheap-tier lenses return, invoke `checkpoint_abort.py` (`python3 ${CLAUDE_PLUGIN_ROOT}/scripts/checkpoint_abort.py --cheap-results-from <cheap-tier results> --lenses <the resolver's full ordered lens list>`) before dispatching the remaining opus-tier lenses. On `aborted: true`: skip dispatching the deferred opus-tier lenses this round, run the review-fix loop against the cheap-tier findings only, and once it converges, dispatch the deferred lenses and merge findings into the round's finding set via `python3 ${CLAUDE_PLUGIN_ROOT}/scripts/checkpoint_abort.py --mode merge --from <JSON: {"existing": [...], "new": [...]}>` (`checkpoint_abort.merge_findings(existing, new)`'s CLI form) — the same dedup-by-`(agent, file, line, severity, message)` function Step 1.1 already proved order-independent. **`existing` here is the cheap-tier findings STILL OUTSTANDING after the fix loop converged (whatever the fix loop's own re-classification left as unresolved — typically none, if the fix loop actually fixed the triggering finding), never the original pre-fix-loop cheap-tier result set.** The triggering finding is `error`/`high` by construction (`checkpoint_abort.py`'s `QUALIFYING_SEVERITY`/`QUALIFYING_CONFIDENCE`), so passing the unfixed original set as `existing` would keep it in the merged set regardless of the fix loop's outcome, and `--mode outcome` below would then report `blocked` on every aborted round no matter what the deferred lenses found (backstop review, #2168) — the fix loop's own result is what `existing` must reflect. `new` is the deferred lenses' findings, dispatched fresh at the fix loop's final content. This checkpoint's report must name every deferred lens and the triggering finding (`checkpoint_abort.py`'s `deferredLenses`/`triggeringFinding`/`triggeringAgent`). The round's pass/blocked outcome always comes from calling `python3 ${CLAUDE_PLUGIN_ROOT}/scripts/checkpoint_abort.py --mode outcome --from <JSON: {"aborted": ..., "redispatched": ..., "findings": [the merged set above]}>` (`compute_round_outcome(aborted, redispatched, findings)`'s CLI form) — never independently reimplemented — so a round that aborted and never re-dispatched its deferred lenses cannot report a pass. On `aborted: false`, dispatch every opus-tier lens normally.
|
|
212
|
+
|
|
213
|
+
**Finding log (#2211).** After each lens returns at a checkpoint, record every finding (fail-open; never blocks the phase): `python3 "${CLAUDE_PLUGIN_ROOT}/scripts/review_findings_log.py" append --lens <lens> --category <finding category> --severity <severity> --iteration <fix-loop iteration> [--fixed]` (`--fixed` once a later iteration clears it). `review_findings_log.py report` ranks `(lens, category)` by frequency and first-pass-fix rate — the input for promoting or pruning authoring-checklist bullets.
|
|
214
|
+
|
|
215
|
+
**Ledger-scoped dispatch (#2167) — applies to both checkpoint fix loops above (sub-steps 4 and 6).** Two additions, both needed together — the first is what makes the second possible:
|
|
216
|
+
|
|
217
|
+
1. **Scope marker.** Append the same structured, single-line marker `../code-review/SKILL.md` step 4 documents to every checkpoint dispatch prompt here — `Files in scope for this review: <path>, <path>, ...` (the literal `hooks/lib/review_verdicts.SCOPE_MARKER_PREFIX`), listing exactly the files passed to that dispatch. Without this, `hooks/review_verdict_recorder.py` (the `SubagentStop` hook that writes `.claude/metrics/review-verdicts.jsonl`) has no scope to parse for a checkpoint dispatch and records nothing for it — today's actual gap: only `/code-review`'s own step 4 emits this marker, so a lens sub-step 4/6 clears never gets a ledger row at all, and Step 6's backstop has nothing to skip against no matter how many checkpoints already reviewed the same content.
|
|
218
|
+
2. **Consult before dispatch.** Before dispatching the resolver-selected lenses (the cheap-first list `select_lenses.py` returned above), narrow them against the ledger:
|
|
219
|
+
```bash
|
|
220
|
+
python3 "${CLAUDE_PLUGIN_ROOT}/scripts/verdict_scope.py" --root <worktree> --lens-files '<JSON: {"<lens>": [<this checkpoint's changed files>], ...} for every lens the resolver returned>'
|
|
221
|
+
```
|
|
222
|
+
Dispatch each lens only against `toDispatch[<lens>]`; a lens named in `fullySkippedLenses` is dropped from this checkpoint's dispatch entirely (nothing left to review, so no fix-loop iteration for it either). Name every ledger-skipped `(lens, file)` pair — and its matched row — in this checkpoint's own report, the same "report loudly, never silently" rule `../code-review/SKILL.md` states for its own copy of this mechanism; that file's copy also states the exhaustive fail-closed list (`verdict_scope.py`'s own module docstring), not repeated here to avoid two independently-drifting copies.
|
|
223
|
+
|
|
224
|
+
**Why this closes the Step 6 backstop gap, with no Step-6-specific code.** Step 6 runs `/code-review --internal`, which performs this exact same consult (`../code-review/SKILL.md` step 4's own "Ledger-scoped dispatch" bullet) against whatever `.claude/metrics/review-verdicts.jsonl` rows already exist. Once sub-steps 4/6 above are writing rows too (via the scope marker addition), a file/lens pair a checkpoint already cleared already has a `pass` row by the time the backstop's own consult runs — so a slice whose steps were all reviewed at sub-step 4 has nothing left for the backstop to re-review, without the backstop needing to know anything about checkpoints at all. `--backstop-review=skip` is unaffected either way — it still suppresses the whole Step 6 dispatch, ledger-scoped or not.
|
|
225
|
+
|
|
226
|
+
**Verification-mode re-dispatch (#1628) — applies to both checkpoint fix loops above (sub-steps 4 and 6).** When a checkpoint re-dispatches an agent to CONFIRM a fix (rather than to discover new problems), send the narrowed verification payload — the finding, the fix diff hunks ± ~20 lines, and the agent's lens definition, never the full file set — with the mandatory `insufficient-context` escape, and resolve the agent's verification tier with `python3 "$CLAUDE_PLUGIN_ROOT/scripts/verify_tier.py" --agent <name>`. The payload contract, the escape's escalation path, and the declared-never-inferred tier-down rule are stated once in [`../../knowledge/verification-mode.md`](../../knowledge/verification-mode.md) and are not restated here.
|
|
227
|
+
|
|
228
|
+
**Round-ledger termination rules (#1625) — applies to both checkpoint fix loops above (sub-steps 4 and 6).** The 5-iteration cap is a backstop, not a churn control: it counts passes, not finding *identity*, so a loop that keeps re-finding residue from its own fixes looks identical to one making progress. Classify each checkpoint round with the shared helper and honor its verdict:
|
|
229
|
+
|
|
230
|
+
```bash
|
|
231
|
+
# RUN_ID scopes the ledger to THIS checkpoint. Each checkpoint is its own
|
|
232
|
+
# run — without it, one step's signatures would suppress the next step's
|
|
233
|
+
# genuinely-new findings as "carried".
|
|
234
|
+
python3 "$CLAUDE_PLUGIN_ROOT/skills/code-review/scripts/finding_signature.py" \
|
|
235
|
+
--round <N> --findings <this-round's-findings.json> \
|
|
236
|
+
--run-id "<plan>-slice<S>-step<N.M>" \
|
|
237
|
+
--state .claude/memory/review-round-state.json
|
|
238
|
+
```
|
|
239
|
+
|
|
240
|
+
The script resets the ledger on round 1, on a `--run-id` mismatch, and on
|
|
241
|
+
a state file older than 24h — see [`../code-review/SKILL.md`](../code-review/SKILL.md) step 6a's ledger-lifecycle table. Pass a `--run-id` that is unique per checkpoint; the step/slice identity above is sufficient.
|
|
242
|
+
|
|
243
|
+
**Architectural-impact gate.** The resolver's `lenses` list is narrowed
|
|
244
|
+
further for structural lenses: pipe the checkpoint's diff through
|
|
245
|
+
`skills/code-review/scripts/change_impact.py --files <changed files>` and
|
|
246
|
+
drop any agent in its `skipLenses`. A step whose diff changes no import,
|
|
247
|
+
adds/moves/deletes no file, edits no manifest or infra file, and changes
|
|
248
|
+
no public symbol has moved no boundary for `arch-review` to evaluate. The
|
|
249
|
+
gate is fail-safe — anything it cannot classify keeps every lens. Contract
|
|
250
|
+
and rationale: `../code-review/SKILL.md` step 3.
|
|
251
|
+
|
|
252
|
+
The three rules — hard round cap at 4, severity floor from round 2 (only `error`/`warning` at `high`/`medium` confidence justifies another round; suggestion-tier findings are logged, never chased), and loop-until-dry — are stated once in [`../code-review/SKILL.md`](../code-review/SKILL.md) step 6a and implemented once in that script. They are not restated here: this is the same contract, same implementation, reached from a different caller. A `round-cap` verdict escalates to the user with the ledger attached, exactly like a non-converging loop.
|
|
253
|
+
|
|
254
|
+
**Resolver note (`select_lenses.py`, #1516).** The resolver reads each review agent's `Scope:` declaration and returns only the lenses whose domain matches the changed files, so a backend-only diff does not dispatch frontend/UI lenses (a11y, js-fp, component-architecture, svelte) that would only no-op — while the `Scope: always` lenses (correctness, security, structure, spec-compliance, domain, arch, …) run on every checkpoint — with one exception (#1923): `correctness-review` additionally drops out when every changed file is non-executable (docs/config/assets/lockfiles, never functional Claude-config markdown), since it self-skips that diff shape per its own `## Skip` clause anyway; the resolver removes it from `lenses` before dispatch rather than paying for that self-reported skip, and records the drop as a `skipped-non-executable:correctness-review` entry in `warnings`. No other `Scope: always` lens is affected by this. It prints `{"lenses":[...cheap-first...],"warnings":[...]}`; dispatch the `lenses` in order and surface `warnings` in the checkpoint's own findings/telemetry the same way `/code-review` does (see `../code-review/SKILL.md`'s resolver-eligibility section for the full warning-shape list). The final `/code-review` (step 6) remains the backstop and applies its own selection independently. **Scope note (#2168):** `/code-review` step 4's parallel bounded-wave dispatch (`dispatch_waves.py`) computes and fires whole waves sized only by `maxParallel`, with no cheap-tier/opus-tier ordering boundary within a wave, so this `checkpoint_abort.py` abort-on-cheap-blocker optimization is **not** ported there — it is scoped to `/build`'s own checkpoints (sub-steps 4 and 6) only.
|
|
255
|
+
7. **Record review value (#348).** For **each** checkpoint that runs (per-step `complex` in sub-step 4, and per-slice in sub-step 6), check `~/.claude/telemetry.json` consent first — if consent is not enabled, skip this step entirely (no file is written). When consent is enabled, append one JSON line to `.claude/metrics/review-value.jsonl` capturing whether review actually changed anything — counts and outcomes only, never code or file content (consistent with the cost meter's privacy boundary). Schema in `performance-metrics`:
|
|
256
|
+
|
|
257
|
+
```json
|
|
258
|
+
{"timestamp":"<ISO8601>","plan":"<plan-file>","slice":"<N>","step":"<N.M or all>","checkpoint":"step|slice|backstop","complexity":"standard|complex","source":"build-checkpoint|build-backstop","diff_shape":"test-only|mixed","agents_run":["spec-compliance-review","..."],"issues_found":0,"severity_breakdown":{"errors":0,"warnings":0,"suggestions":0},"issues_fixed":0,"fix_iterations":0,"outcome":"no-op|fixed|escalated|skipped"}
|
|
259
|
+
```
|
|
260
|
+
|
|
261
|
+
Step 6's backstop row (#1962) uses this same schema with
|
|
262
|
+
`source: "build-backstop"` and `checkpoint: "backstop"`; `outcome:
|
|
263
|
+
"skipped"` is reserved for a backstop suppressed by
|
|
264
|
+
`--backstop-review=skip` and never appears on a checkpoint row.
|
|
265
|
+
|
|
266
|
+
**`diff_shape` (#1964)** records the *shape* of the diff this checkpoint
|
|
267
|
+
reviewed, so per-lens outcomes can be split by it. Do not eyeball the file
|
|
268
|
+
list — read it from the same deterministic helper the gates use:
|
|
269
|
+
|
|
270
|
+
```bash
|
|
271
|
+
python3 "$CLAUDE_PLUGIN_ROOT/skills/code-review/scripts/change_shape.py" --files <this checkpoint's changed files>
|
|
272
|
+
```
|
|
273
|
+
|
|
274
|
+
Its `isTestOnly` field decides the value: `true` → `"test-only"`, `false`
|
|
275
|
+
→ `"mixed"`. The helper is include-biased (any file not *provably* a test
|
|
276
|
+
makes the answer `false`), so `test-only` is never over-claimed. This is
|
|
277
|
+
the measurement that decides whether any lens may eventually be gated out
|
|
278
|
+
of a test-only diff — it gates nothing today.
|
|
279
|
+
|
|
280
|
+
`outcome` is `no-op` when the checkpoint passed clean (found nothing), `fixed` when it found and auto-fixed actionable issues, `escalated` when the loop didn't converge. `severity_breakdown` splits `issues_found` by severity (`errors`/`warnings`/`suggestions`, the same enum as `/code-review`), so `/harness-audit` Step 3 can flag a lens producing mostly minor findings — the three counts must sum to `issues_found` (#1256). This is the sensor that tells a build where review caught a real defect from one where every loop passed no-op — it turns the pipeline's "value untested" into "value measured" and feeds the plan/step tiering decisions. Disable with `DEV_TEAM_REVIEW_VALUE=off`.
|
|
281
|
+
|
|
282
|
+
### 4.9. Verify runtime behavior before the slice is done (issue #727)
|
|
283
|
+
|
|
284
|
+
A "done" step that only passed its own tests is not the same as a feature that works — a red suite catches structural regressions, not "it fails the first time someone actually runs it." Once a slice's steps are all `[x]` (sub-step 5) and its review checkpoint(s) have run (sub-steps 4/6), decide whether the slice has a runtime surface to exercise **before the slice may be marked `[x]` complete**:
|
|
285
|
+
|
|
286
|
+
1. **Classify the slice's changed files**, per `knowledge/test-file-indicators.md`. If every changed file is a test file, or the rest are docs/config only (no source or runtime file changed), there is nothing to exercise at runtime — record `outcome: "skipped"` with a `reason` (below) and continue.
|
|
287
|
+
2. **Otherwise, exercise the change end-to-end** using the project's own test/verification tooling — its test runner, checker scripts, or a direct invocation of the changed entry point (CLI command, API call, script run) — scoped to the slice's changed runtime files, before the slice's checkbox is flipped to `[x]`. This is a pattern, not a named command: there is no `/verify` skill shipped by this plugin, so pick whatever the project already uses to run/exercise the affected surface for real (e.g. its integration test suite, a smoke-test script, or manually invoking the changed function/endpoint/command). This generalizes the UI-only `/browse` smoke test (sub-step 4's UI bullet) into a universal completion criterion: APIs, CLIs, bots, and background jobs get the same "did this actually run" check UI changes already get.
|
|
288
|
+
3. **Not bypassable by `--yes`, `DEV_TEAM_AUTO_APPROVE=1`, or no-TTY.** Contrast with the approval gates in Steps 2–3: those bypass a human judgment call when no human is present. This gate needs no human judgment — the agent runs the verification itself — so non-interactive mode never skips it. There is no override flag for this step.
|
|
289
|
+
4. **A failed verification run is a failing test.** Per Step 5's "Quality ownership" language: do not mark the slice `[x]` or the plan `implemented`. Enter [Systematic Debugging](../systematic-debugging/SKILL.md), find the root cause, fix it, and re-run the verification before proceeding — never silently override.
|
|
290
|
+
5. **Record the outcome.** Append exactly one JSON line per slice with a runtime surface to `metrics/verify-log.jsonl`, schema modeled on `.claude/metrics/review-value.jsonl` (sub-step 7):
|
|
291
|
+
|
|
292
|
+
```json
|
|
293
|
+
{"timestamp":"<ISO8601>","plan":"<plan-file>","slice":"<N>","branch":"<branch>","files":["<changed runtime file>","..."],"outcome":"ran|skipped|failed-then-fixed","reason":"<set when outcome is skipped>"}
|
|
294
|
+
```
|
|
295
|
+
|
|
296
|
+
`outcome` is `ran` (the verification ran and passed), `skipped` (no runtime surface in the diff — `reason` states why, e.g. `"tests-only"` or `"docs-only"`), or `failed-then-fixed` (the verification failed at least once before the fix landed). `python3 ${CLAUDE_PLUGIN_ROOT}/scripts/progress_guardian.py --pre-pr` reads this log: a branch with runtime-surface changes and no matching entry fails the pre-PR gate the same way an incomplete step or a missing commit does.
|
|
297
|
+
|
|
298
|
+
### 4.10. Run slice invariants (issue #865)
|
|
299
|
+
|
|
300
|
+
When the slice declares `**Invariants:**`, run them **after** the slice's own suite is green (sub-step 5) and its review checkpoint(s) have run (sub-steps 4/6) — invariants check what must stay green *beyond* the slice's new acceptance tests, so they gate on top of everything else, not instead of it:
|
|
301
|
+
|
|
302
|
+
```bash
|
|
303
|
+
python3 "${CLAUDE_PLUGIN_ROOT}/scripts/run_invariants.py" --plan <plan-file> --slice <id> --repo <worktree>
|
|
304
|
+
```
|
|
305
|
+
|
|
306
|
+
A non-zero exit **fails the slice gate exactly like a red test** — fix it or escalate (Escalation section below), never step over it, and never flip the slice checkbox to `[x]` until it's green. A slice with no `Invariants` line runs its gate unchanged (the script itself no-ops with "No invariants declared" — nothing to enforce). Invariant commands run as-is from the repo root; the plan author owns their portability, same trust model as the plan's own test commands.
|
|
307
|
+
|
|
308
|
+
### 5. Run full test suite
|
|
309
|
+
|
|
310
|
+
After all steps are complete, delete `.claude/memory/build-phase.json` if a `between-steps` record remains, then run the full test suite. Paste the output as final verification evidence.
|
|
311
|
+
|
|
312
|
+
**Quality ownership — the whole suite must be green, not just this branch's tests.** A failing test is a failing test regardless of whether this change caused it: a red suite blocks `/pr` even when the failure pre-dates the branch. Do not wave a failure past as "pre-existing / unrelated." Either fix it (enter [Systematic Debugging](../systematic-debugging/SKILL.md) for the root cause), or — if it is genuinely out of scope — explicitly surface and triage it (`/triage` a record or quarantine it with a reason) and report the suite as **not green**. Never proceed to `/pr` on red by attributing the failure to someone else's change.
|
|
313
|
+
|
|
314
|
+
### 6. Run code review
|
|
315
|
+
|
|
316
|
+
Run `/code-review --internal` against all files modified during the build —
|
|
317
|
+
deliberately not `--json`, to keep the review-fix loop running per
|
|
318
|
+
`/code-review`'s own step-6 exception (b).
|
|
319
|
+
|
|
320
|
+
**Record the backstop's value (#1962).** This pass is the build's most
|
|
321
|
+
duplicated review layer — every file it reviews was already seen by an inline
|
|
322
|
+
checkpoint (sub-steps 4/6) — but until now nothing measured whether it earned
|
|
323
|
+
its cost, because `review-value.jsonl` recorded only checkpoint rows. Subject
|
|
324
|
+
to the same `~/.claude/telemetry.json` consent check and `DEV_TEAM_REVIEW_VALUE`
|
|
325
|
+
kill switch as sub-step 7, append one row for this pass using the identical
|
|
326
|
+
schema, with `source: "build-backstop"`, `checkpoint: "backstop"`, and
|
|
327
|
+
`step: "all"`. That row is what turns "is the backstop redundant?" from an
|
|
328
|
+
argument into a measurement.
|
|
329
|
+
|
|
330
|
+
**`--backstop-review=skip` suppresses this step (#1962).** When the flag is
|
|
331
|
+
set: do not dispatch the panel, print the audit line from Parse Arguments
|
|
332
|
+
naming the enclosing reviewer, and append a row with `outcome: "skipped"` so
|
|
333
|
+
the suppression is visible in the same stream as the runs it replaces — never
|
|
334
|
+
a silent absence. Proceed to Step 7. The flag reaches only this step; a build
|
|
335
|
+
that skips the backstop still cannot reach `/pr` on red (Step 5) or on an
|
|
336
|
+
unresolved escalation.
|
|
337
|
+
|
|
338
|
+
**The flag is a caller's assertion, not a shortcut this skill may take on its
|
|
339
|
+
own initiative.** `/build` never infers it — an orchestrator passes it, and
|
|
340
|
+
only when its own review pass covers a superset of this build's diff.
|
|
341
|
+
|
|
342
|
+
### 7. Final test quality score (branch)
|
|
343
|
+
|
|
344
|
+
Produce a Farley Score for the tests written on this branch — the last quality signal before `/pr`.
|
|
345
|
+
|
|
346
|
+
1. Resolve the branch base: `git merge-base HEAD origin/HEAD` (fall back to `origin/main`, then `main`, `master`, `develop`).
|
|
347
|
+
**Degenerate-base check (issue #916).** The fallback chain assumes at least one candidate ref sits meaningfully behind HEAD. That's false in a single-branch/no-remote repo — there's no `origin` at all (the `merge-base` call fails outright) or every commit landed directly on the fallback branch itself (e.g. `master`), so `git merge-base HEAD master` resolves to HEAD. Treat **base == HEAD, or every candidate ref unresolvable,** as a resolution failure, not a valid base — silently continuing to sub-step 2 diffs HEAD against itself and reports a false "no tests written":
|
|
348
|
+
- **Fall back to the plan's recorded plan-start commit** — the same anchor a `Rollback point: plan-start` slice already resolved against (issue #865): `python3 ${CLAUDE_PLUGIN_ROOT}/scripts/build_rollback_point.py get-by-symbolic --path .claude/memory/build-rollback.json --symbolic plan-start --repo <worktree> --ancestor-of HEAD`. `.claude/memory/build-rollback.json` is a flat store never cleared between plans sharing a worktree, so `--repo`/`--ancestor-of` matter: they reject a stale `plan-start` entry left behind by an unrelated earlier build (its SHA won't be an ancestor of this branch's HEAD) rather than trusting the first match blindly. If a qualifying entry is found, use its `sha` as `<base>` for sub-step 2 — this is the "found" case below.
|
|
349
|
+
- **If no qualifying plan-start rollback point was found** (exit 1: no slice in this build declared `Rollback point: plan-start`, or the only recorded entry failed the ancestry check as stale), do not proceed to sub-step 2 as if nothing were wrong. Print an explicit warning — `Branch-base resolution degraded: origin/HEAD, origin/main, main, master, develop all resolved to HEAD or failed, and no plan-start rollback point is recorded — the Farley Score step below cannot distinguish "no tests written" from a resolution failure.` — then continue to sub-step 2 with the degraded (== HEAD) base anyway, so sub-step 3 knows to treat an empty diff as inconclusive rather than clean. This is the "not found" case sub-step 3 branches on below — distinct from the "found" case immediately above, where the plan-start SHA is a trustworthy base and an empty diff against it is genuine evidence of no tests written.
|
|
350
|
+
2. List the branch's changed test files: `git diff --name-only <base>...HEAD`, keeping only test files (indicators in `knowledge/test-file-indicators.md` — `*.test.*` / `*.spec.*` / `__tests__/`, xUnit/JUnit attributes, `.feature` files).
|
|
351
|
+
3. If no test files changed on the branch, print one line — `No tests written on this branch — skipping Farley Score.` — and continue to Step 8. **Exception:** when sub-step 1's degenerate-base check hit the "not found" case (no qualifying plan-start rollback point, base left at HEAD), print that sub-step's degraded-resolution warning instead of this clean line — an empty diff off an unresolved HEAD-equals-base is a resolution failure, not evidence of a clean branch. The "found" case (a validated plan-start SHA used as base) is a real base, so a genuine empty diff against it prints the ordinary clean line above. Either way, continue to Step 8.
|
|
352
|
+
4. Otherwise invoke the `farley-score` skill scoped to those files. Present the suite-level Farley Score, rating, and distribution as the final pre-PR signal. This is **informational** — a low score does not block `/pr`, but surface it so the user can decide whether to revise before opening the PR.
|
|
353
|
+
|
|
354
|
+
### 7.5. Assemble the evidence bundle
|
|
355
|
+
|
|
356
|
+
Before the completion report, assemble a structured evidence bundle per
|
|
357
|
+
`${CLAUDE_PLUGIN_ROOT}/knowledge/evidence-bundle.md` — **no new checks, no
|
|
358
|
+
re-execution**; it renders data this run already produced:
|
|
359
|
+
|
|
360
|
+
- **Checks run**: the Step 5 full-suite command + result, the Step 6
|
|
361
|
+
`/code-review` status, the Step 7 Farley Score command/output (or its
|
|
362
|
+
skip line when no tests changed).
|
|
363
|
+
- **Scope notes**: review agents dispatched vs. skipped across the build's
|
|
364
|
+
checkpoints (sub-steps 4/6), and any gate reported "not applicable."
|
|
365
|
+
- **Untested regions**: read `baseline-coverage.json` / `coverage-history.json`
|
|
366
|
+
if present (from `/coverage-baseline` / `/coverage-delta`); otherwise state
|
|
367
|
+
"not measured — no coverage tool detected."
|
|
368
|
+
- **Residual risks**: derived-first from this run's `.claude/metrics/review-value.jsonl`
|
|
369
|
+
entries with `outcome: "escalated"`, any non-interactive gate-bypass audit
|
|
370
|
+
lines printed in Steps 2–3, and any `failed-then-fixed` runtime-verification
|
|
371
|
+
entries in `metrics/verify-log.jsonl`. "None identified" only when all of
|
|
372
|
+
those are empty.
|
|
373
|
+
|
|
374
|
+
Follow the degradation rule: every one of the four section headers appears in
|
|
375
|
+
the completion report even when a section has nothing to show — it states why.
|
|
376
|
+
|
|
377
|
+
### 8. Update plan status
|
|
378
|
+
|
|
379
|
+
Use the Edit tool to change `**Status**: in-progress` to `**Status**: implemented` in the plan file. Briefly confirm completion, report the branch Farley Score, include the Step 7.5 evidence bundle in the completion report, and direct the user to `/pr`.
|
|
380
|
+
|
|
381
|
+
### 9. Learning loop
|
|
382
|
+
|
|
383
|
+
Invoke the [Feedback & Learning](../feedback-learning/SKILL.md) skill at task completion to capture any correction turns from this session. If the user used correction language during the build (e.g. "no, actually", "revert", "that's wrong", "stop doing X"), record the pattern so it can become an instruction rule. If no corrections occurred, this step is a no-op — invoke and it will report nothing to capture.
|
|
384
|
+
|
|
385
|
+
## Escalation
|
|
386
|
+
|
|
387
|
+
A failure is a debugging task first, not a hand-back. Before escalating any test, review, or bash/command failure, run a [Systematic Debugging](../systematic-debugging/SKILL.md) pass — reproduce, find the root cause, state it in one sentence — and escalate **with that diagnosis**, never just an attempt count.
|
|
388
|
+
|
|
389
|
+
Stop and ask the user when:
|
|
390
|
+
|
|
391
|
+
- A test still fails *after systematic debugging has identified the root cause* and the fix needs a decision you can't make (e.g. it requires changing the spec or the architecture)
|
|
392
|
+
- The plan requires architectural decisions not covered by the plan
|
|
393
|
+
- A review checkpoint fails after 2 correction iterations *and* the root cause is understood but unresolvable within scope
|
|
394
|
+
- You discover the plan is incomplete or contradictory
|
|
395
|
+
- **A step's required behavior cannot be tested in isolation** — the plan has an architectural gap (e.g. the step's contract depends on scaffolding no earlier step produced)
|
|
396
|
+
- **A step's dependency was not produced by a prior step that was supposed to produce it** (most likely in a same-wave or cross-wave parallel build, `isolation: "worktree"`) — stop and escalate rather than stubbing or guessing the missing interface inline, which would silently expand the step's scope beyond what it was dispatched to do
|
|
397
|
+
- The `verify_guard.py` hook blocks a verify command (`[BLOCK]` on a test/lint/build re-run) — this is the deterministic signal that the same command has run repeatedly with no intervening code change, i.e. a stuck loop rather than a progressing per-behavior cycle. Run the Systematic Debugging pass above instead of retrying the command again, and escalate with the diagnosis if it's still unresolvable in scope.
|
|
398
|
+
- **The step-4 repair loop hits a failure-signature dead-end** (issue #864): two consecutive repair iterations produce the same normalized failure signature (failing test IDs + error class, volatile output stripped). This is a hard stop, not another auto-approval: commit the current working tree as a checkpoint on the working branch — never `main` — with a conventional message explicitly marked as a dead-end checkpoint (e.g. `chore(build): dead-end checkpoint — step <N>, <M> tests still failing`), then escalate stating (a) **improved** — tests that went failing → passing since repair start, (b) **remaining** — the current failing signature, (c) the **checkpoint commit ref**. If 3 or more distinct fix attempts have failed, the escalation must also cite [Systematic Debugging](../systematic-debugging/SKILL.md)'s "3+ failed fix attempts → question the architecture" rule. Leave plan status unchanged and never proceed to `/pr` over this escalation.
|
|
399
|
+
|
|
400
|
+
**Non-interactive runs: an escalation is a hard stop, not another auto-approval.**
|
|
401
|
+
The approval gates in Steps 2–3 auto-proceed because they bypass a judgment call the
|
|
402
|
+
human delegated by going non-interactive; an escalation exists because the agent hit
|
|
403
|
+
something it cannot safely decide — that authority was never delegated. When any
|
|
404
|
+
condition above fires and no user can answer (`--yes`, `DEV_TEAM_AUTO_APPROVE=1`, or
|
|
405
|
+
no TTY): write the escalation (trigger, one-sentence root-cause diagnosis, options
|
|
406
|
+
considered) to `.claude/memory/build-escalation-<plan-slug>.md`, leave the plan status
|
|
407
|
+
unchanged, report the halt as the build result, and stop. Never resolve an escalation
|
|
408
|
+
by picking an interpretation and continuing, and never proceed to `/pr` over an
|
|
409
|
+
unresolved escalation.
|
|
410
|
+
|
|
411
|
+
## Integration
|
|
412
|
+
|
|
413
|
+
- `/specs` produces the intent, architecture, and acceptance-criteria artifacts that inform the plan
|
|
414
|
+
- `/plan` decomposes the feature into slices, authors each slice's Gherkin, and produces the plan this command executes
|
|
415
|
+
- Sub-step 4.9 exercises each runtime-surface slice end-to-end, using the project's own test/verification tooling scoped to the diff, before it may be marked done (issue #727) — this is a pattern to follow, not a named `/verify` command
|
|
416
|
+
- `/code-review` runs the full review suite after implementation
|
|
417
|
+
- `farley-score` scores the branch's tests (Farley Score) as the final pre-PR quality signal
|
|
418
|
+
- `${CLAUDE_PLUGIN_ROOT}/knowledge/evidence-bundle.md` defines the structured evidence bundle assembled in Step 7.5 and surfaced in the Step 8 completion report
|
|
419
|
+
- `/pr` creates the pull request after a successful build, assembling its own evidence bundle independently (no handoff file)
|
|
420
|
+
- `/continue` can resume a partially completed build across sessions
|
|
421
|
+
- `python3 ${CLAUDE_PLUGIN_ROOT}/scripts/progress_guardian.py --plan <plan-file>` validates step completion and commit discipline at each step boundary; `--pre-pr` also fails closed when runtime-surface changes have no matching `metrics/verify-log.jsonl` entry (issue #727), and warns (never fails) on out-of-scope edits against declared slice `Files` (issue #865)
|
|
422
|
+
- `${CLAUDE_PLUGIN_ROOT}/scripts/build_slice_scope.py`, `${CLAUDE_PLUGIN_ROOT}/scripts/build_rollback_point.py`, and `${CLAUDE_PLUGIN_ROOT}/scripts/run_invariants.py` implement the plan-as-contract fields (issue #865): opt-in freeze scope, rollback-point resolution/recording, and the slice invariants gate, respectively
|