pi-dev-team 0.2.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/LICENSE +21 -0
- package/PORTING.md +134 -0
- package/README.md +207 -0
- package/UPSTREAM.json +64 -0
- package/agents/Explore.md +15 -0
- package/agents/a11y-review.md +118 -0
- package/agents/adr-author.md +70 -0
- package/agents/ai-provenance-review.md +120 -0
- package/agents/angular-reactivity-review.md +95 -0
- package/agents/arch-review.md +135 -0
- package/agents/architect.md +78 -0
- package/agents/autoship-batch-proposer.md +69 -0
- package/agents/claude-setup-review.md +136 -0
- package/agents/codebase-recon.md +184 -0
- package/agents/component-architecture-review.md +119 -0
- package/agents/concurrency-review.md +109 -0
- package/agents/correctness-review.md +290 -0
- package/agents/data-flow-tracer.md +120 -0
- package/agents/doc-review.md +165 -0
- package/agents/domain-review.md +136 -0
- package/agents/general-purpose.md +10 -0
- package/agents/gherkin-quality-critic.md +113 -0
- package/agents/js-fp-review.md +114 -0
- package/agents/mutation-kill.md +684 -0
- package/agents/naming-review.md +142 -0
- package/agents/orchestrator.md +339 -0
- package/agents/performance-review.md +105 -0
- package/agents/plan-review-acceptance.md +115 -0
- package/agents/plan-review-design.md +90 -0
- package/agents/plan-review-parallelization.md +84 -0
- package/agents/plan-review-strategic.md +96 -0
- package/agents/plan-review-ux.md +110 -0
- package/agents/platform-engineer.md +64 -0
- package/agents/product-manager.md +68 -0
- package/agents/progress-guardian.md +79 -0
- package/agents/qa-engineer.md +289 -0
- package/agents/quality-reviewer.md +132 -0
- package/agents/react-reactivity-review.md +102 -0
- package/agents/refactor-opportunity-review.md +128 -0
- package/agents/security-engineer.md +60 -0
- package/agents/security-review.md +218 -0
- package/agents/session-analysis.md +95 -0
- package/agents/software-engineer.md +105 -0
- package/agents/spec-compliance-review.md +100 -0
- package/agents/spec-reviewer.md +114 -0
- package/agents/structure-review.md +146 -0
- package/agents/tech-writer.md +84 -0
- package/agents/test-review.md +246 -0
- package/agents/test-smell-review.md +188 -0
- package/agents/token-efficiency-review.md +139 -0
- package/agents/ui-ux-designer.md +54 -0
- package/agents/vue-reactivity-review.md +95 -0
- package/bin/__pycache__/claudecpython-314.pyc +0 -0
- package/bin/claude +258 -0
- package/docs/upstream/.pages +1 -0
- package/docs/upstream/CHANGELOG.md +2586 -0
- package/docs/upstream/README.md +155 -0
- package/docs/upstream/agent-architecture.md +214 -0
- package/docs/upstream/agent_info.md +187 -0
- package/docs/upstream/artifact-migration.md +124 -0
- package/docs/upstream/code-intelligence-nudge.md +149 -0
- package/docs/upstream/code-review-process.md +294 -0
- package/docs/upstream/concurrent-use.md +73 -0
- package/docs/upstream/context-management.md +111 -0
- package/docs/upstream/developer-notes.md +280 -0
- package/docs/upstream/diagrams/architecture-overview.svg +101 -0
- package/docs/upstream/diagrams/review-dispatch.svg +139 -0
- package/docs/upstream/diagrams/team-agents.svg +128 -0
- package/docs/upstream/diagrams/test-improve-flow.svg +166 -0
- package/docs/upstream/diagrams/workflow-linear.svg +66 -0
- package/docs/upstream/diagrams/workflow-three-phase.svg +200 -0
- package/docs/upstream/eval-maintenance.md +95 -0
- package/docs/upstream/eval-running-guide.md +147 -0
- package/docs/upstream/eval-system.md +291 -0
- package/docs/upstream/session-review-oss-complements.md +75 -0
- package/docs/upstream/session-review.md +212 -0
- package/docs/upstream/skills.md +188 -0
- package/docs/upstream/team-structure.md +21 -0
- package/docs/upstream/telemetry-ci-access.md +129 -0
- package/docs/upstream/telemetry-repo-security.md +120 -0
- package/docs/upstream/test-evaluation.md +277 -0
- package/docs/upstream/test-improve.md +154 -0
- package/docs/upstream/triage-workflow.md +282 -0
- package/docs/upstream/workflows.md +289 -0
- package/extensions/dev-team/index.ts +539 -0
- package/extensions/dev-team/lib/agents.ts +272 -0
- package/extensions/dev-team/lib/ai-credits.ts +92 -0
- package/extensions/dev-team/lib/autocompact.ts +81 -0
- package/extensions/dev-team/lib/child-run.ts +102 -0
- package/extensions/dev-team/lib/config.ts +236 -0
- package/extensions/dev-team/lib/gh-command.ts +103 -0
- package/extensions/dev-team/lib/github-style.ts +307 -0
- package/extensions/dev-team/lib/hooks.ts +350 -0
- package/extensions/dev-team/lib/metrics.ts +115 -0
- package/extensions/dev-team/lib/safe-read.ts +49 -0
- package/extensions/dev-team/lib/session-files.ts +57 -0
- package/extensions/dev-team/lib/session-spend.ts +123 -0
- package/extensions/dev-team/lib/shell-scan.ts +205 -0
- package/extensions/dev-team/lib/skills.ts +213 -0
- package/extensions/dev-team/lib/subagent-render.ts +245 -0
- package/extensions/dev-team/lib/subagent-types.ts +164 -0
- package/extensions/dev-team/lib/subagent.ts +596 -0
- package/extensions/dev-team/lib/terminal-text.ts +54 -0
- package/extensions/dev-team/lib/tools-misc.ts +152 -0
- package/extensions/dev-team/lib/transcript.ts +110 -0
- package/extensions/dev-team/lib/trust.ts +52 -0
- package/extensions/dev-team/lib/usage-breakdown.ts +176 -0
- package/extensions/dev-team/lib/usage-chart.ts +153 -0
- package/extensions/dev-team/lib/usage-command.ts +107 -0
- package/extensions/dev-team/lib/usage-history.ts +203 -0
- package/extensions/dev-team/lib/usage-render.ts +225 -0
- package/extensions/dev-team/lib/usage-split-bar.ts +127 -0
- package/extensions/dev-team/lib/usage-state.ts +116 -0
- package/extensions/dev-team/lib/usage-text.ts +159 -0
- package/extensions/dev-team/lib/usage-view.ts +109 -0
- package/hooks/__pycache__/refactor_test_freeze_guard.cpython-314.pyc +0 -0
- package/hooks/agent_dispatch_ledger.py +190 -0
- package/hooks/autocompact_setup_nudge.py +99 -0
- package/hooks/bash_retry_guard.py +228 -0
- package/hooks/boundary_events_write_guard.py +352 -0
- package/hooks/code_intelligence_nudge.py +293 -0
- package/hooks/code_intelligence_turn_mark.py +317 -0
- package/hooks/codegraph_bootstrap.py +139 -0
- package/hooks/contract_version_guard.py +362 -0
- package/hooks/cost_meter.py +106 -0
- package/hooks/destructive-commands.json +62 -0
- package/hooks/destructive_guard.py +477 -0
- package/hooks/eval_compliance_check.py +440 -0
- package/hooks/guards.json +17 -0
- package/hooks/hooks.json +323 -0
- package/hooks/internal_double_gate.py +296 -0
- package/hooks/js_fp_review.py +212 -0
- package/hooks/knowledge_index.py +119 -0
- package/hooks/lib/__pycache__/artifact_paths.cpython-314.pyc +0 -0
- package/hooks/lib/__pycache__/atomic_state.cpython-314.pyc +0 -0
- package/hooks/lib/__pycache__/autocompact_config.cpython-314.pyc +0 -0
- package/hooks/lib/__pycache__/boundary_events.cpython-314.pyc +0 -0
- package/hooks/lib/__pycache__/doc_classification.cpython-314.pyc +0 -0
- package/hooks/lib/__pycache__/gh_pr_create_detect.cpython-314.pyc +0 -0
- package/hooks/lib/__pycache__/git_safe_diff.cpython-314.pyc +0 -0
- package/hooks/lib/__pycache__/instrument_log.cpython-314.pyc +0 -0
- package/hooks/lib/__pycache__/metrics_query.cpython-314.pyc +0 -0
- package/hooks/lib/__pycache__/plugin_version.cpython-314.pyc +0 -0
- package/hooks/lib/__pycache__/pre_commit_doc_classifier.cpython-314.pyc +0 -0
- package/hooks/lib/__pycache__/review_agent_registry.cpython-314.pyc +0 -0
- package/hooks/lib/__pycache__/review_gate_corroboration.cpython-314.pyc +0 -0
- package/hooks/lib/__pycache__/review_gate_hash.cpython-314.pyc +0 -0
- package/hooks/lib/__pycache__/review_verdicts.cpython-314.pyc +0 -0
- package/hooks/lib/__pycache__/stdin_json.cpython-314.pyc +0 -0
- package/hooks/lib/__pycache__/stryker_invocation.cpython-314.pyc +0 -0
- package/hooks/lib/__pycache__/telemetry_consent.cpython-314.pyc +0 -0
- package/hooks/lib/__pycache__/test_file_classify.cpython-314.pyc +0 -0
- package/hooks/lib/__pycache__/token_efficiency_limits.cpython-314.pyc +0 -0
- package/hooks/lib/__pycache__/verify_guard_state.cpython-314.pyc +0 -0
- package/hooks/lib/__pycache__/xunit_v3_operator_gate.cpython-314.pyc +0 -0
- package/hooks/lib/agent_skill_hints.py +74 -0
- package/hooks/lib/artifact_paths.py +263 -0
- package/hooks/lib/atomic_state.py +557 -0
- package/hooks/lib/autocompact_config.py +103 -0
- package/hooks/lib/autoship_log.py +106 -0
- package/hooks/lib/banned_scripts_policy.py +51 -0
- package/hooks/lib/boundary_events.py +436 -0
- package/hooks/lib/build_knowledge_index.py +504 -0
- package/hooks/lib/build_skills_index.py +361 -0
- package/hooks/lib/build_state.py +116 -0
- package/hooks/lib/classify_ship_outcome.py +126 -0
- package/hooks/lib/config_changelog_schema.py +115 -0
- package/hooks/lib/cost_meter.py +955 -0
- package/hooks/lib/doc_classification.py +116 -0
- package/hooks/lib/gh_pr_create_detect.py +136 -0
- package/hooks/lib/git_safe_diff.py +123 -0
- package/hooks/lib/instrument_log.py +66 -0
- package/hooks/lib/iteration_journal_gate.py +197 -0
- package/hooks/lib/knowledge_index_paths.py +88 -0
- package/hooks/lib/mcp_json_repowise.py +177 -0
- package/hooks/lib/metrics_query.py +202 -0
- package/hooks/lib/minimal_yaml.py +434 -0
- package/hooks/lib/plugin_version.py +142 -0
- package/hooks/lib/pre_commit_detect.py +537 -0
- package/hooks/lib/pre_commit_doc_classifier.py +126 -0
- package/hooks/lib/pricing.py +118 -0
- package/hooks/lib/report_pdf.py +371 -0
- package/hooks/lib/review_agent_registry.py +142 -0
- package/hooks/lib/review_dispatch_ledger.py +101 -0
- package/hooks/lib/review_gate_corroboration.py +521 -0
- package/hooks/lib/review_gate_hash.py +252 -0
- package/hooks/lib/review_gate_normalized_hash.py +1115 -0
- package/hooks/lib/review_verdicts.py +301 -0
- package/hooks/lib/run_report.py +160 -0
- package/hooks/lib/skill_categories.yaml +125 -0
- package/hooks/lib/stdin_json.py +57 -0
- package/hooks/lib/stryker_invocation.py +102 -0
- package/hooks/lib/telemetry_consent.py +41 -0
- package/hooks/lib/telemetry_report.py +108 -0
- package/hooks/lib/test_file_classify.py +160 -0
- package/hooks/lib/token_efficiency_limits.py +51 -0
- package/hooks/lib/turn_identity.py +77 -0
- package/hooks/lib/verify_guard_state.py +110 -0
- package/hooks/lib/workflow_state.py +206 -0
- package/hooks/lib/xunit_v3_operator_gate.py +596 -0
- package/hooks/mcp_json_repowise_nudge.py +74 -0
- package/hooks/mutation_adapters/__init__.py +7 -0
- package/hooks/mutation_adapters/__pycache__/__init__.cpython-314.pyc +0 -0
- package/hooks/mutation_adapters/__pycache__/lib.cpython-314.pyc +0 -0
- package/hooks/mutation_adapters/__pycache__/mutmut.cpython-314.pyc +0 -0
- package/hooks/mutation_adapters/__pycache__/pitest.cpython-314.pyc +0 -0
- package/hooks/mutation_adapters/__pycache__/stryker.cpython-314.pyc +0 -0
- package/hooks/mutation_adapters/__pycache__/stryker_net.cpython-314.pyc +0 -0
- package/hooks/mutation_adapters/lib.py +478 -0
- package/hooks/mutation_adapters/mutmut.py +188 -0
- package/hooks/mutation_adapters/pitest.py +266 -0
- package/hooks/mutation_adapters/stryker.py +157 -0
- package/hooks/mutation_adapters/stryker_net.py +264 -0
- package/hooks/mutation_gate.py +193 -0
- package/hooks/mutation_testing_smoke_gate.py +371 -0
- package/hooks/pending_review_notify.py +121 -0
- package/hooks/phase_marker.py +138 -0
- package/hooks/post_compact_state_reinject.py +180 -0
- package/hooks/post_format.py +115 -0
- package/hooks/pre_commit_knowledge_index.py +128 -0
- package/hooks/pre_commit_review.py +66 -0
- package/hooks/pre_pr_review.py +694 -0
- package/hooks/pre_tool_guard.py +405 -0
- package/hooks/py.sh +73 -0
- package/hooks/refactor-bash-write-patterns.json +29 -0
- package/hooks/refactor_test_bash_guard.py +253 -0
- package/hooks/refactor_test_freeze_guard.py +139 -0
- package/hooks/refactor_test_revert_guard.py +186 -0
- package/hooks/repo_review_nudge.py +287 -0
- package/hooks/review_verdict_recorder.py +464 -0
- package/hooks/scan_bash_command_for_banned_scripts.py +428 -0
- package/hooks/scan_worktree_for_banned_scripts.py +238 -0
- package/hooks/session_learning_trigger.py +248 -0
- package/hooks/skills_index.py +126 -0
- package/hooks/stryker_xunit_shim_guard.py +571 -0
- package/hooks/subagent_completion_guard.py +309 -0
- package/hooks/subagent_skill_context.py +139 -0
- package/hooks/task_completion_metrics.py +216 -0
- package/hooks/tdd_guard.py +229 -0
- package/hooks/telemetry.py +341 -0
- package/hooks/token_efficiency_review.py +194 -0
- package/hooks/verify_guard.py +183 -0
- package/hooks/verify_guard_edit_marker.py +73 -0
- package/hooks/version_check.py +173 -0
- package/knowledge/accepted-risks-schema.md +98 -0
- package/knowledge/adr-decision-criteria.md +64 -0
- package/knowledge/adversarial-review-protocol.md +139 -0
- package/knowledge/agent-registry.md +228 -0
- package/knowledge/agent-review-methodology.md +80 -0
- package/knowledge/ai-friendly-repo-guidelines.md +67 -0
- package/knowledge/architecture-assessment.md +96 -0
- package/knowledge/artifact-lifecycle.md +57 -0
- package/knowledge/cd-maturity-model.md +82 -0
- package/knowledge/cd-test-architecture.md +190 -0
- package/knowledge/ci-cd-file-scope.md +24 -0
- package/knowledge/codegraph-vs-graphify.md +192 -0
- package/knowledge/component-test-patterns.md +139 -0
- package/knowledge/database-change-management.md +80 -0
- package/knowledge/database-test-patterns.md +79 -0
- package/knowledge/decision-defaults.md +88 -0
- package/knowledge/dependency-breaking-techniques.md +116 -0
- package/knowledge/deployment-pipeline.md +86 -0
- package/knowledge/design-smells.md +122 -0
- package/knowledge/directory-enumeration.md +38 -0
- package/knowledge/domain-modeling.md +123 -0
- package/knowledge/evidence-bundle.md +90 -0
- package/knowledge/exploratory-testing-field-guide.md +122 -0
- package/knowledge/failure-routing.md +28 -0
- package/knowledge/fixture-construction.md +56 -0
- package/knowledge/frontend-component-architecture.md +139 -0
- package/knowledge/gherkin-quality-review-dispatch.md +135 -0
- package/knowledge/index.json +6766 -0
- package/knowledge/internal-collaborator-doubling.md +101 -0
- package/knowledge/legacy-test-strategy.md +71 -0
- package/knowledge/long-run-waiting.md +66 -0
- package/knowledge/microservice-testing.md +71 -0
- package/knowledge/model-pricing.json +23 -0
- package/knowledge/mutation-score-formulas.md +60 -0
- package/knowledge/object-calisthenics.md +147 -0
- package/knowledge/oracle-provenance.md +94 -0
- package/knowledge/orchestrator-script-implementation.md +185 -0
- package/knowledge/owasp-detection.md +148 -0
- package/knowledge/plan-review-rubric.md +56 -0
- package/knowledge/proxy-connectivity.md +62 -0
- package/knowledge/reactive-effect-patterns.md +73 -0
- package/knowledge/recon-inventory-excludes.txt +32 -0
- package/knowledge/references/bdd-value-guide.md +61 -0
- package/knowledge/references/csharp-http-client-testing.md +264 -0
- package/knowledge/release-strategies.md +74 -0
- package/knowledge/report-output-location.md +117 -0
- package/knowledge/report-pdf-integration.md +63 -0
- package/knowledge/report-print.css +129 -0
- package/knowledge/report-template.md +114 -0
- package/knowledge/report-to-pdf.md +69 -0
- package/knowledge/request-processing-flow.md +63 -0
- package/knowledge/result-verification.md +52 -0
- package/knowledge/review-agent-output-contract.md +121 -0
- package/knowledge/review-lens-classification.md +113 -0
- package/knowledge/review-rubric.md +62 -0
- package/knowledge/review-template.md +104 -0
- package/knowledge/rule-fixtures/A02.insecure-random-js/negative.js +1 -0
- package/knowledge/rule-fixtures/A02.insecure-random-js/positive.js +1 -0
- package/knowledge/rule-fixtures/A02.weak-hashing-md5/negative.py +1 -0
- package/knowledge/rule-fixtures/A02.weak-hashing-md5/positive.py +1 -0
- package/knowledge/rule-fixtures/A03.command-injection/negative.js +1 -0
- package/knowledge/rule-fixtures/A03.command-injection/positive.js +1 -0
- package/knowledge/rule-fixtures/A03.sql-injection/negative.js +1 -0
- package/knowledge/rule-fixtures/A03.sql-injection/positive.js +1 -0
- package/knowledge/rule-fixtures/A03.xss-innerhtml/negative.js +1 -0
- package/knowledge/rule-fixtures/A03.xss-innerhtml/positive.js +1 -0
- package/knowledge/rule-fixtures/A05.cors-wildcard/negative.js +1 -0
- package/knowledge/rule-fixtures/A05.cors-wildcard/positive.js +1 -0
- package/knowledge/rule-fixtures/A05.default-credentials/negative.js +1 -0
- package/knowledge/rule-fixtures/A05.default-credentials/positive.js +1 -0
- package/knowledge/rule-fixtures/A07.jwt-alg-none/negative.js +1 -0
- package/knowledge/rule-fixtures/A07.jwt-alg-none/positive.js +1 -0
- package/knowledge/rule-fixtures/A08.binary-formatter/negative.cs +1 -0
- package/knowledge/rule-fixtures/A08.binary-formatter/positive.cs +1 -0
- package/knowledge/rule-fixtures/A08.js-eval/negative.js +1 -0
- package/knowledge/rule-fixtures/A08.js-eval/positive.js +1 -0
- package/knowledge/rule-fixtures/A08.object-input-stream/negative.java +1 -0
- package/knowledge/rule-fixtures/A08.object-input-stream/positive.java +1 -0
- package/knowledge/schemas/disposition-register-v1.json +65 -0
- package/knowledge/schemas/recon-envelope-v1.json +198 -0
- package/knowledge/schemas/unified-finding-v1.json +72 -0
- package/knowledge/security-primitives-contract.md +301 -0
- package/knowledge/security-review-rule-map.yaml +107 -0
- package/knowledge/skills-registry.md +72 -0
- package/knowledge/task-size-classifier.md +103 -0
- package/knowledge/telemetry-schema.md +881 -0
- package/knowledge/test-automation-maturity.md +56 -0
- package/knowledge/test-automation-principles.md +71 -0
- package/knowledge/test-cadence-tradeoffs.md +68 -0
- package/knowledge/test-doubles.md +105 -0
- package/knowledge/test-file-indicators.md +22 -0
- package/knowledge/test-layer-gates.md +35 -0
- package/knowledge/test-matrix-examples/django-batch.md +24 -0
- package/knowledge/test-matrix-examples/dotnet-grpc-fronting-api.md +90 -0
- package/knowledge/test-matrix-examples/dotnet-http-consumer.md +131 -0
- package/knowledge/test-matrix-examples/react-node-spa.md +24 -0
- package/knowledge/test-matrix-examples/spring-boot-service.md +25 -0
- package/knowledge/test-matrix-examples/ssr-htmx.md +24 -0
- package/knowledge/test-organization.md +70 -0
- package/knowledge/test-pyramid.md +84 -0
- package/knowledge/test-refactoring.md +67 -0
- package/knowledge/test-review-division-of-labor.md +85 -0
- package/knowledge/test-smells.md +80 -0
- package/knowledge/test-stack-profiles/bdd-frameworks.md +235 -0
- package/knowledge/test-stack-profiles/django.md +13 -0
- package/knowledge/test-stack-profiles/dotnet.md +18 -0
- package/knowledge/test-stack-profiles/go.md +16 -0
- package/knowledge/test-stack-profiles/node.md +16 -0
- package/knowledge/test-stack-profiles/react.md +12 -0
- package/knowledge/test-stack-profiles/spring-boot.md +16 -0
- package/knowledge/test-stack-profiles/ssr-htmx.md +14 -0
- package/knowledge/test-stack-profiles/vue.md +12 -0
- package/knowledge/test-strategy.md +70 -0
- package/knowledge/testability-patterns.md +240 -0
- package/knowledge/testing-quadrants.md +44 -0
- package/knowledge/testing-techniques/approval.md +15 -0
- package/knowledge/testing-techniques/chaos.md +17 -0
- package/knowledge/testing-techniques/fuzz.md +15 -0
- package/knowledge/testing-techniques/property-based.md +15 -0
- package/knowledge/testing-techniques/schema-validation.md +15 -0
- package/knowledge/testing-techniques/screenshot.md +15 -0
- package/knowledge/three-phase-workflow.md +198 -0
- package/knowledge/value-patterns.md +55 -0
- package/knowledge/verification-mode.md +116 -0
- package/knowledge/virtual-service-libraries.md +75 -0
- package/knowledge/wave-consolidation-guidance.md +21 -0
- package/overrides/agents/Explore.md +15 -0
- package/overrides/agents/general-purpose.md +10 -0
- package/overrides/notes/autoship.md +6 -0
- package/overrides/notes/issues-from-assessment.md +3 -0
- package/overrides/notes/issues-from-plan.md +3 -0
- package/overrides/notes/mutation-night-watch.md +3 -0
- package/overrides/notes/mutation-testing.md +3 -0
- package/overrides/notes/pr.md +7 -0
- package/overrides/notes/project-init.md +6 -0
- package/overrides/notes/setup.md +13 -0
- package/overrides/notes/specs.md +3 -0
- package/overrides/skills/headless-run/SKILL.md +45 -0
- package/overrides/skills/upgrade/SKILL.md +30 -0
- package/overrides/skills/version/SKILL.md +25 -0
- package/package.json +36 -0
- package/scripts/authoring_digest.py +93 -0
- package/scripts/autoship_discover.py +121 -0
- package/scripts/autoship_group.py +409 -0
- package/scripts/autoship_proposals.py +494 -0
- package/scripts/autoship_queue.py +291 -0
- package/scripts/autoship_reclaim.py +495 -0
- package/scripts/build_jobs.py +108 -0
- package/scripts/build_rollback_point.py +240 -0
- package/scripts/build_slice_scope.py +157 -0
- package/scripts/build_wave.py +109 -0
- package/scripts/build_wave_reconcile.py +252 -0
- package/scripts/build_worktree_baseref.py +113 -0
- package/scripts/check_agent_scope.py +117 -0
- package/scripts/check_agent_tool_mapping.py +213 -0
- package/scripts/check_review_agent_mcp_tools.py +317 -0
- package/scripts/check_security_assessment_mcp_tools.py +165 -0
- package/scripts/checkpoint_abort.py +502 -0
- package/scripts/claude_setup_review.py +438 -0
- package/scripts/codebase_recon.py +556 -0
- package/scripts/coverage_config.py +623 -0
- package/scripts/coverage_delta_steering.py +330 -0
- package/scripts/coverage_discovery_dotnet.py +315 -0
- package/scripts/coverage_discovery_java.py +742 -0
- package/scripts/coverage_discovery_js.py +546 -0
- package/scripts/coverage_gap_ranking.py +556 -0
- package/scripts/coverage_readiness.py +455 -0
- package/scripts/coverage_report_parse.py +521 -0
- package/scripts/detect_bdd_convention.py +252 -0
- package/scripts/eval_ablation.py +376 -0
- package/scripts/gherkin_analysis_coverage_gate.py +306 -0
- package/scripts/gherkin_cross_feature_duplicate_titles_gate.py +173 -0
- package/scripts/gherkin_effectiveness_rollup.py +238 -0
- package/scripts/gherkin_failure_path_gate.py +206 -0
- package/scripts/gherkin_feature_merge.py +720 -0
- package/scripts/gherkin_stub_gate.py +163 -0
- package/scripts/gherkin_stub_merge.py +479 -0
- package/scripts/git_origin_host.py +88 -0
- package/scripts/install-java-static-analysis.py +110 -0
- package/scripts/issue_deps.py +74 -0
- package/scripts/lib/_bdd_markers.py +28 -0
- package/scripts/lib/_gherkin_text.py +93 -0
- package/scripts/lib/_vendored_tree.py +70 -0
- package/scripts/lib/autoship_state.py +397 -0
- package/scripts/lib/claude_md_guard.py +226 -0
- package/scripts/lib/deterministic_recon.py +446 -0
- package/scripts/lib/mcp_tool_grants.py +211 -0
- package/scripts/lib/plan_parse.py +386 -0
- package/scripts/lib/review_result.py +84 -0
- package/scripts/lib/review_roster.py +86 -0
- package/scripts/lib/session_log/__init__.py +34 -0
- package/scripts/lib/session_log/__pycache__/__init__.cpython-314.pyc +0 -0
- package/scripts/lib/session_log/__pycache__/records.cpython-314.pyc +0 -0
- package/scripts/lib/session_log/classify.py +231 -0
- package/scripts/lib/session_log/corrections.py +194 -0
- package/scripts/lib/session_log/discovery.py +108 -0
- package/scripts/lib/session_log/records.py +218 -0
- package/scripts/lib/session_log/redact.py +76 -0
- package/scripts/lib/session_log/signals.py +373 -0
- package/scripts/lib/session_report_downstream.py +614 -0
- package/scripts/lib/session_report_maintainer.py +1273 -0
- package/scripts/lib/session_report_shared.py +262 -0
- package/scripts/lib/settings_hook_guard.py +157 -0
- package/scripts/lib/slug.py +33 -0
- package/scripts/lib/stub_extractors/__init__.py +82 -0
- package/scripts/lib/stub_extractors/_common.py +328 -0
- package/scripts/lib/stub_extractors/csharp.py +19 -0
- package/scripts/lib/stub_extractors/go.py +173 -0
- package/scripts/lib/stub_extractors/java.py +18 -0
- package/scripts/lib/stub_extractors/jsts.py +126 -0
- package/scripts/mutation_stack_sections.py +149 -0
- package/scripts/mutation_yield_steering.py +345 -0
- package/scripts/orchestrator.py +895 -0
- package/scripts/plan_gherkin_export.py +227 -0
- package/scripts/plan_waves.py +208 -0
- package/scripts/pr_close_keyword_lint.py +108 -0
- package/scripts/progress_guardian.py +888 -0
- package/scripts/recon_inventory.py +273 -0
- package/scripts/review_findings_log.py +93 -0
- package/scripts/run_invariants.py +124 -0
- package/scripts/select_lenses.py +640 -0
- package/scripts/session_report.py +486 -0
- package/scripts/set_autocompact_env.py +221 -0
- package/scripts/ship_resume_guard.py +135 -0
- package/scripts/ship_review_gate.py +63 -0
- package/scripts/specs_convention_marker.py +103 -0
- package/scripts/test_improve_resume.py +277 -0
- package/scripts/test_review_mechanics.py +958 -0
- package/scripts/token_efficiency_review.py +322 -0
- package/scripts/verdict_scope.py +285 -0
- package/scripts/verify_gherkin_quality_critic_isolation.py +296 -0
- package/scripts/verify_tier.py +157 -0
- package/skills/adr-tools/SKILL.md +118 -0
- package/skills/agent-readiness/SKILL.md +105 -0
- package/skills/agent-readiness/ai_friendly_analyzers.py +326 -0
- package/skills/agent-readiness/scanner.py +441 -0
- package/skills/agent-readiness/scorecard.yaml +88 -0
- package/skills/api-design/SKILL.md +115 -0
- package/skills/apply-fixes/SKILL.md +171 -0
- package/skills/apply-test-doubles/SKILL.md +321 -0
- package/skills/artifact-lifecycle/SKILL.md +127 -0
- package/skills/autoship/SKILL.md +1124 -0
- package/skills/benchmark/SKILL.md +105 -0
- package/skills/branch-workflow/SKILL.md +89 -0
- package/skills/browse/SKILL.md +184 -0
- package/skills/browser-testing/SKILL.md +62 -0
- package/skills/browser-testing/references/playwright-patterns.md +216 -0
- package/skills/build/SKILL.md +422 -0
- package/skills/build/references/static-self-heal.md +245 -0
- package/skills/careful/SKILL.md +72 -0
- package/skills/cd-test-architecture/SKILL.md +371 -0
- package/skills/ci-debugging/SKILL.md +105 -0
- package/skills/co-evolution-audit/SKILL.md +269 -0
- package/skills/code-review/SKILL.md +1015 -0
- package/skills/code-review/examples/aggregated-sample.json +56 -0
- package/skills/code-review/examples/sample-report.md +41 -0
- package/skills/code-review/output-format.md +478 -0
- package/skills/code-review/scripts/activation.py +86 -0
- package/skills/code-review/scripts/change_impact.py +357 -0
- package/skills/code-review/scripts/change_shape.py +372 -0
- package/skills/code-review/scripts/change_size.py +212 -0
- package/skills/code-review/scripts/changed_file_list.py +141 -0
- package/skills/code-review/scripts/closing_pass.py +187 -0
- package/skills/code-review/scripts/consolidate.py +277 -0
- package/skills/code-review/scripts/contract_failure_report.py +185 -0
- package/skills/code-review/scripts/dispatch_reconcile.py +66 -0
- package/skills/code-review/scripts/dispatch_waves.py +164 -0
- package/skills/code-review/scripts/finding_signature.py +446 -0
- package/skills/code-review/scripts/ledger.py +283 -0
- package/skills/code-review/scripts/partition.py +169 -0
- package/skills/code-review/scripts/render_tiered_findings.py +274 -0
- package/skills/code-review/scripts/repo_invariants.py +1066 -0
- package/skills/code-review/scripts/review_context_pack.py +306 -0
- package/skills/code-review/scripts/review_round_log.py +345 -0
- package/skills/code-review/scripts/review_value_coverage.py +297 -0
- package/skills/code-review/scripts/validate_review_output.py +467 -0
- package/skills/code-review/sliced-mode.md +205 -0
- package/skills/competitive-analysis/SKILL.md +191 -0
- package/skills/context-loading-protocol/SKILL.md +157 -0
- package/skills/continue/SKILL.md +90 -0
- package/skills/cost-report/SKILL.md +178 -0
- package/skills/coverage-baseline/SKILL.md +335 -0
- package/skills/coverage-baseline/references/multi-project-discovery.md +202 -0
- package/skills/coverage-delta/SKILL.md +181 -0
- package/skills/coverage-delta/references/mutation-gate.md +70 -0
- package/skills/design-doc/SKILL.md +95 -0
- package/skills/design-interrogation/SKILL.md +89 -0
- package/skills/design-it-twice/SKILL.md +91 -0
- package/skills/docker-image-audit/SKILL.md +108 -0
- package/skills/docker-image-audit/references/install-guide.md +64 -0
- package/skills/docker-image-audit/references/report-template.md +73 -0
- package/skills/docker-image-create/SKILL.md +185 -0
- package/skills/domain-analysis/SKILL.md +183 -0
- package/skills/domain-driven-design/SKILL.md +194 -0
- package/skills/exploratory-testing/SKILL.md +108 -0
- package/skills/explore/SKILL.md +51 -0
- package/skills/farley-score/SKILL.md +165 -0
- package/skills/feature-file-validation/SKILL.md +78 -0
- package/skills/feature-file-validation/references/validation-rules.md +115 -0
- package/skills/feedback-learning/SKILL.md +414 -0
- package/skills/fix/SKILL.md +450 -0
- package/skills/freeze/SKILL.md +68 -0
- package/skills/frontend-architecture/SKILL.md +113 -0
- package/skills/gherkin-derive/SKILL.md +630 -0
- package/skills/gherkin-public/SKILL.md +266 -0
- package/skills/governance-compliance/SKILL.md +150 -0
- package/skills/guard/SKILL.md +75 -0
- package/skills/handoff/SKILL.md +139 -0
- package/skills/handoff/references/summary-templates.md +242 -0
- package/skills/harness-audit/SKILL.md +751 -0
- package/skills/harness-audit/scripts/lesson_validate.py +386 -0
- package/skills/harness-audit/scripts/redundancy_criterion.py +188 -0
- package/skills/headless-run/SKILL.md +45 -0
- package/skills/headless-run/scripts/isolated_dispatch.py +381 -0
- package/skills/help/SKILL.md +72 -0
- package/skills/hexagonal-architecture/SKILL.md +85 -0
- package/skills/human-oversight-protocol/SKILL.md +224 -0
- package/skills/issues-from-assessment/SKILL.md +223 -0
- package/skills/issues-from-plan/SKILL.md +133 -0
- package/skills/legacy-code/SKILL.md +132 -0
- package/skills/mermaid-diagramming/SKILL.md +120 -0
- package/skills/mutation-night-watch/SKILL.md +154 -0
- package/skills/mutation-night-watch/references/scheduling.md +135 -0
- package/skills/mutation-testing/SKILL.md +396 -0
- package/skills/mutation-testing/references/languages/csharp-stryker-net.md +676 -0
- package/skills/mutation-testing/references/languages/go-go-mutesting.md +95 -0
- package/skills/mutation-testing/references/languages/java-pitest.md +77 -0
- package/skills/mutation-testing/references/languages/javascript-stryker.md +188 -0
- package/skills/mutation-testing/references/languages/python-mutmut.md +97 -0
- package/skills/mutation-testing/references/time-estimation.md +34 -0
- package/skills/mutation-testing/references/tool-detection.md +15 -0
- package/skills/mutation-testing/references/workflow-callers.md +23 -0
- package/skills/mutation-testing/scripts/__pycache__/xunit_v3_feature_detector.cpython-314.pyc +0 -0
- package/skills/mutation-testing/scripts/csharp_stryker_net_slice_runner.py +635 -0
- package/skills/mutation-testing/scripts/csharp_stryker_net_status_loop.py +525 -0
- package/skills/mutation-testing/scripts/csharp_stryker_net_wrapper.py +681 -0
- package/skills/mutation-testing/scripts/mutation_baseline_reuse.py +292 -0
- package/skills/mutation-testing/scripts/mutation_exclude_policy.py +268 -0
- package/skills/mutation-testing/scripts/mutation_feasibility_gate.py +463 -0
- package/skills/mutation-testing/scripts/mutation_kill_headless.py +331 -0
- package/skills/mutation-testing/scripts/mutation_kill_insert.py +199 -0
- package/skills/mutation-testing/scripts/mutation_kill_insert_python.py +150 -0
- package/skills/mutation-testing/scripts/mutation_kill_loop.py +869 -0
- package/skills/mutation-testing/scripts/mutation_kill_loop_python.py +949 -0
- package/skills/mutation-testing/scripts/mutation_kill_retry.py +592 -0
- package/skills/mutation-testing/scripts/mutation_kill_shared.py +620 -0
- package/skills/mutation-testing/scripts/mutation_nightwatch.py +462 -0
- package/skills/mutation-testing/scripts/mutation_nightwatch_stacks.py +425 -0
- package/skills/mutation-testing/scripts/mutation_report.py +743 -0
- package/skills/mutation-testing/scripts/mutation_report_cli.py +175 -0
- package/skills/mutation-testing/scripts/mutation_safety_gate.py +69 -0
- package/skills/mutation-testing/scripts/stryker_shard_pipeline.py +847 -0
- package/skills/mutation-testing/scripts/stryker_shard_setup.py +440 -0
- package/skills/mutation-testing/scripts/stryker_timeout_retry.py +142 -0
- package/skills/mutation-testing/scripts/xunit_v3_feature_detector.py +341 -0
- package/skills/performance-benchmark/SKILL.md +174 -0
- package/skills/performance-benchmark/examples/report-format.md +43 -0
- package/skills/performance-benchmark/references/benchmark-script.md +169 -0
- package/skills/performance-metrics/SKILL.md +265 -0
- package/skills/plan/SKILL.md +199 -0
- package/skills/plan/references/gherkin-persistence.md +43 -0
- package/skills/plan/references/plan-template.md +182 -0
- package/skills/pr/SKILL.md +289 -0
- package/skills/pr/scripts/gate_retry_state.py +368 -0
- package/skills/project-init/README.md +141 -0
- package/skills/project-init/SKILL.md +1197 -0
- package/skills/project-init/evals/evals.json +200 -0
- package/skills/project-init/references/capability-tools.md +55 -0
- package/skills/project-init/references/configs.md +221 -0
- package/skills/property-based-testing/SKILL.md +121 -0
- package/skills/property-based-testing/fixtures/invariant_fixture.py +15 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/README.md +42 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/fast-check/LICENSE +21 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/fast-check/README.md +263 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/fast-check/lib/cjs/fast-check.d.ts +5165 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/fast-check/lib/cjs/fast-check.js +12147 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/fast-check/lib/cjs/package.json +3 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/fast-check/lib/cjs/types57/fast-check.d.ts +5165 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/fast-check/lib/fast-check.d.ts +5165 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/fast-check/lib/fast-check.js +12011 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/fast-check/lib/rolldown-runtime-D7D4PA-g.js +13 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/fast-check/lib/types57/fast-check.d.ts +5165 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/fast-check/package.json +94 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/LICENSE +21 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/README.md +168 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/RandomGenerator-DcXj09Ch.d.ts +14 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/distribution/uniformBigInt.d.ts +15 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/distribution/uniformBigInt.js +38 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/distribution/uniformFloat32.d.ts +15 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/distribution/uniformFloat32.js +18 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/distribution/uniformFloat64.d.ts +15 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/distribution/uniformFloat64.js +22 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/distribution/uniformInt.d.ts +15 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/distribution/uniformInt.js +134 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/RandomGenerator-DcXj09Ch.d.ts +14 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/distribution/uniformBigInt.d.ts +15 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/distribution/uniformBigInt.js +37 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/distribution/uniformFloat32.d.ts +15 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/distribution/uniformFloat32.js +17 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/distribution/uniformFloat64.d.ts +15 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/distribution/uniformFloat64.js +21 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/distribution/uniformInt.d.ts +15 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/distribution/uniformInt.js +133 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/generator/congruential32.d.ts +7 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/generator/congruential32.js +44 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/generator/mersenne.d.ts +7 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/generator/mersenne.js +90 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/generator/xoroshiro128plus.d.ts +7 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/generator/xoroshiro128plus.js +80 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/generator/xorshift128plus.d.ts +7 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/generator/xorshift128plus.js +78 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/package.json +3 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/types/JumpableRandomGenerator.d.ts +16 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/types/JumpableRandomGenerator.js +0 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/types/RandomGenerator.d.ts +2 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/types/RandomGenerator.js +0 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/utils/generateN.d.ts +6 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/utils/generateN.js +8 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/utils/purify.d.ts +12 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/utils/purify.js +9 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/utils/skipN.d.ts +6 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/utils/skipN.js +6 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/generator/congruential32.d.ts +7 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/generator/congruential32.js +46 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/generator/mersenne.d.ts +7 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/generator/mersenne.js +92 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/generator/xoroshiro128plus.d.ts +7 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/generator/xoroshiro128plus.js +82 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/generator/xorshift128plus.d.ts +7 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/generator/xorshift128plus.js +80 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/types/JumpableRandomGenerator.d.ts +16 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/types/JumpableRandomGenerator.js +0 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/types/RandomGenerator.d.ts +2 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/types/RandomGenerator.js +0 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/utils/generateN.d.ts +6 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/utils/generateN.js +9 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/utils/purify.d.ts +12 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/utils/purify.js +10 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/utils/skipN.d.ts +6 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/utils/skipN.js +7 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/package.json +133 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/package-lock.json +1179 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/package.json +14 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/roundtrip.js +29 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/roundtrip.properties.test.js +16 -0
- package/skills/property-based-testing/fixtures/no_property_fixture.py +10 -0
- package/skills/property-based-testing/fixtures/roundtrip_fixture.py +16 -0
- package/skills/property-based-testing/references/languages/javascript.md +54 -0
- package/skills/property-based-testing/scripts/detect_and_dispatch.py +80 -0
- package/skills/property-based-testing/scripts/hypothesis_scaffold.py +276 -0
- package/skills/proxy-resilience/SKILL.md +84 -0
- package/skills/quality-gate-pipeline/SKILL.md +184 -0
- package/skills/quality-targets-converge/SKILL.md +254 -0
- package/skills/repo-review/SKILL.md +159 -0
- package/skills/report-pdf/SKILL.md +66 -0
- package/skills/review/SKILL.md +47 -0
- package/skills/review-agent/SKILL.md +152 -0
- package/skills/review-summary/SKILL.md +73 -0
- package/skills/run-report/SKILL.md +70 -0
- package/skills/semantic-duplication-scan/SKILL.md +337 -0
- package/skills/semantic-scan/SKILL.md +53 -0
- package/skills/semgrep-analyze/SKILL.md +139 -0
- package/skills/setup/SKILL.md +1122 -0
- package/skills/ship/SKILL.md +240 -0
- package/skills/source-verification/SKILL.md +210 -0
- package/skills/source-verification/scripts/claim_extractor.py +155 -0
- package/skills/specs/.size-baseline.json +4 -0
- package/skills/specs/SKILL.md +243 -0
- package/skills/specs/references/completeness-checklist.md +83 -0
- package/skills/specs/references/extraction.md +58 -0
- package/skills/specs/references/glossary.md +59 -0
- package/skills/specs/references/persistence.md +115 -0
- package/skills/specs/references/predictability-check.md +77 -0
- package/skills/static-analysis-integration/SKILL.md +235 -0
- package/skills/static-analysis-integration/adapters/_envelope.py +26 -0
- package/skills/static-analysis-integration/adapters/jscpd-adapter.py +66 -0
- package/skills/static-analysis-integration/adapters/lizard-adapter.py +81 -0
- package/skills/static-analysis-integration/adapters/mypy-adapter.py +50 -0
- package/skills/static-analysis-integration/adapters/mypy-src-layout.py +93 -0
- package/skills/static-analysis-integration/adapters/security-review-adapter.py +212 -0
- package/skills/static-analysis-integration/maintenance.md +23 -0
- package/skills/static-analysis-integration/references/language-setup.md +228 -0
- package/skills/static-analysis-integration/references/sarif-parser.md +124 -0
- package/skills/static-analysis-integration/references/security-review-adapter.md +118 -0
- package/skills/static-analysis-integration/references/tool-configs.md +617 -0
- package/skills/static-analysis-integration/rulesets/pmd-quickstart.xml +24 -0
- package/skills/stryker-xunit-v2-shim/SKILL.md +274 -0
- package/skills/stryker-xunit-v2-shim/references/shim-howto.md +256 -0
- package/skills/stryker-xunit-v2-shim/scripts/generate_shim.py +143 -0
- package/skills/systematic-debugging/SKILL.md +130 -0
- package/skills/telemetry/SKILL.md +75 -0
- package/skills/test-audit-disable/SKILL.md +129 -0
- package/skills/test-design/SKILL.md +177 -0
- package/skills/test-design/scripts/__pycache__/internal_double_detector.cpython-314.pyc +0 -0
- package/skills/test-design/scripts/internal_double_detector.py +631 -0
- package/skills/test-design-advisor/SKILL.md +166 -0
- package/skills/test-driven-development/SKILL.md +169 -0
- package/skills/test-health/SKILL.md +262 -0
- package/skills/test-improve/SKILL.md +239 -0
- package/skills/test-improve/references/phase-0-approach-contract.md +228 -0
- package/skills/test-improve/references/phase-1-analyze.md +131 -0
- package/skills/test-improve/references/phase-2-baseline.md +121 -0
- package/skills/test-improve/references/phase-3-derive-gherkin.md +53 -0
- package/skills/test-improve/references/phase-4-plan-fixes.md +34 -0
- package/skills/test-improve/references/phase-5-improve.md +215 -0
- package/skills/test-improve/references/phase-6-refactor-decision.md +45 -0
- package/skills/test-improve/references/phase-7-refactor.md +44 -0
- package/skills/test-improve/references/phase-8-validate.md +66 -0
- package/skills/test-improve/references/phase-9-close-out-prompt.md +11 -0
- package/skills/test-improve/references/phase-9-report.md +62 -0
- package/skills/test-improve/references/review-loop.md +92 -0
- package/skills/test-improve/templates/executive-summary.md +123 -0
- package/skills/threat-modeling/SKILL.md +108 -0
- package/skills/triage/SKILL.md +211 -0
- package/skills/ubiquitous-language/SKILL.md +192 -0
- package/skills/ubiquitous-language/scripts/collect_domain_signals.py +300 -0
- package/skills/unfreeze/SKILL.md +37 -0
- package/skills/upgrade/SKILL.md +31 -0
- package/skills/upgrade/scripts/check_version_drift.py +113 -0
- package/skills/upgrade/scripts/enable_autoupdate.py +149 -0
- package/skills/version/SKILL.md +25 -0
- package/sync/__pycache__/sync_upstream.cpython-314.pyc +0 -0
- package/sync/sync_upstream.py +293 -0
- package/templates/ACCEPTED-RISKS.md.tmpl +46 -0
- package/templates/agents/agent-template.md +151 -0
- package/templates/agents/angular-testing.md +66 -0
- package/templates/agents/csharp-quality.md +63 -0
- package/templates/agents/esm-enforcer.md +52 -0
- package/templates/agents/front-end-testing.md +65 -0
- package/templates/agents/go-quality.md +65 -0
- package/templates/agents/python-quality.md +62 -0
- package/templates/agents/react-testing.md +61 -0
- package/templates/agents/ts-enforcer.md +60 -0
- package/templates/agents/twelve-factor-audit.md +49 -0
- package/tools/entropy-check.py +250 -0
- package/tools/model-hash-verify.py +213 -0
|
@@ -0,0 +1,169 @@
|
|
|
1
|
+
# Playwright Benchmark Script Template
|
|
2
|
+
|
|
3
|
+
## Full Collection Script
|
|
4
|
+
|
|
5
|
+
This script template collects all performance metrics for a single page. The `/benchmark` command generates and runs a version of this script for each target URL.
|
|
6
|
+
|
|
7
|
+
```javascript
|
|
8
|
+
const { chromium } = require('playwright');
|
|
9
|
+
|
|
10
|
+
async function benchmark(url, options = {}) {
|
|
11
|
+
const { runs = 3, device = 'desktop', throttle = false } = options;
|
|
12
|
+
const results = [];
|
|
13
|
+
|
|
14
|
+
for (let i = 0; i < runs; i++) {
|
|
15
|
+
const browser = await chromium.launch({
|
|
16
|
+
args: ['--disable-gpu', '--disable-extensions', '--no-sandbox']
|
|
17
|
+
});
|
|
18
|
+
|
|
19
|
+
const context = await browser.newContext({
|
|
20
|
+
viewport: device === 'mobile'
|
|
21
|
+
? { width: 375, height: 812 }
|
|
22
|
+
: { width: 1280, height: 720 },
|
|
23
|
+
userAgent: device === 'mobile'
|
|
24
|
+
? 'Mozilla/5.0 (iPhone; CPU iPhone OS 16_0 like Mac OS X) Mobile'
|
|
25
|
+
: undefined
|
|
26
|
+
});
|
|
27
|
+
|
|
28
|
+
if (throttle) {
|
|
29
|
+
const cdp = await context.newCDPSession(await context.newPage());
|
|
30
|
+
await cdp.send('Network.emulateNetworkConditions', {
|
|
31
|
+
offline: false,
|
|
32
|
+
downloadThroughput: 1.5 * 1024 * 1024 / 8, // 1.5 Mbps
|
|
33
|
+
uploadThroughput: 750 * 1024 / 8, // 750 Kbps
|
|
34
|
+
latency: 40
|
|
35
|
+
});
|
|
36
|
+
await cdp.send('Emulation.setCPUThrottlingRate', { rate: 4 });
|
|
37
|
+
}
|
|
38
|
+
|
|
39
|
+
const page = await context.newPage();
|
|
40
|
+
|
|
41
|
+
// Inject Performance Observer before navigation
|
|
42
|
+
await page.addInitScript(() => {
|
|
43
|
+
window.__perfMetrics = {};
|
|
44
|
+
new PerformanceObserver((list) => {
|
|
45
|
+
for (const entry of list.getEntries()) {
|
|
46
|
+
if (entry.entryType === 'largest-contentful-paint') {
|
|
47
|
+
window.__perfMetrics.LCP = entry.startTime;
|
|
48
|
+
}
|
|
49
|
+
if (entry.entryType === 'paint' && entry.name === 'first-contentful-paint') {
|
|
50
|
+
window.__perfMetrics.FCP = entry.startTime;
|
|
51
|
+
}
|
|
52
|
+
if (entry.entryType === 'layout-shift' && !entry.hadRecentInput) {
|
|
53
|
+
window.__perfMetrics.CLS = (window.__perfMetrics.CLS || 0) + entry.value;
|
|
54
|
+
}
|
|
55
|
+
}
|
|
56
|
+
}).observe({ type: 'largest-contentful-paint', buffered: true });
|
|
57
|
+
|
|
58
|
+
new PerformanceObserver((list) => {
|
|
59
|
+
for (const entry of list.getEntries()) {
|
|
60
|
+
if (entry.name === 'first-contentful-paint') {
|
|
61
|
+
window.__perfMetrics.FCP = entry.startTime;
|
|
62
|
+
}
|
|
63
|
+
}
|
|
64
|
+
}).observe({ type: 'paint', buffered: true });
|
|
65
|
+
|
|
66
|
+
new PerformanceObserver((list) => {
|
|
67
|
+
for (const entry of list.getEntries()) {
|
|
68
|
+
if (!entry.hadRecentInput) {
|
|
69
|
+
window.__perfMetrics.CLS = (window.__perfMetrics.CLS || 0) + entry.value;
|
|
70
|
+
}
|
|
71
|
+
}
|
|
72
|
+
}).observe({ type: 'layout-shift', buffered: true });
|
|
73
|
+
});
|
|
74
|
+
|
|
75
|
+
// Navigate and wait for network idle
|
|
76
|
+
await page.goto(url, { waitUntil: 'networkidle', timeout: 30000 });
|
|
77
|
+
|
|
78
|
+
// Wait for metrics to stabilize
|
|
79
|
+
await page.waitForTimeout(2000);
|
|
80
|
+
|
|
81
|
+
// Collect all metrics
|
|
82
|
+
const metrics = await page.evaluate(() => {
|
|
83
|
+
const timing = performance.timing;
|
|
84
|
+
const resources = performance.getEntriesByType('resource');
|
|
85
|
+
|
|
86
|
+
const jsResources = resources.filter(r => r.name.endsWith('.js') || r.initiatorType === 'script');
|
|
87
|
+
const cssResources = resources.filter(r => r.name.endsWith('.css') || r.initiatorType === 'link');
|
|
88
|
+
const imgResources = resources.filter(r =>
|
|
89
|
+
r.initiatorType === 'img' || /\.(png|jpg|jpeg|gif|svg|webp|avif)/.test(r.name)
|
|
90
|
+
);
|
|
91
|
+
|
|
92
|
+
return {
|
|
93
|
+
vitals: {
|
|
94
|
+
LCP: window.__perfMetrics.LCP || null,
|
|
95
|
+
FCP: window.__perfMetrics.FCP || null,
|
|
96
|
+
CLS: window.__perfMetrics.CLS || 0,
|
|
97
|
+
TTFB: timing.responseStart - timing.navigationStart,
|
|
98
|
+
domInteractive: timing.domInteractive - timing.navigationStart,
|
|
99
|
+
loadComplete: timing.loadEventEnd - timing.navigationStart
|
|
100
|
+
},
|
|
101
|
+
resources: {
|
|
102
|
+
totalTransferSize: resources.reduce((sum, r) => sum + (r.transferSize || 0), 0),
|
|
103
|
+
jsSize: jsResources.reduce((sum, r) => sum + (r.transferSize || 0), 0),
|
|
104
|
+
cssSize: cssResources.reduce((sum, r) => sum + (r.transferSize || 0), 0),
|
|
105
|
+
imageSize: imgResources.reduce((sum, r) => sum + (r.transferSize || 0), 0),
|
|
106
|
+
requestCount: resources.length,
|
|
107
|
+
largestResource: resources.reduce((max, r) =>
|
|
108
|
+
(r.transferSize || 0) > (max.size || 0)
|
|
109
|
+
? { url: r.name, size: r.transferSize }
|
|
110
|
+
: max,
|
|
111
|
+
{ url: '', size: 0 }
|
|
112
|
+
)
|
|
113
|
+
}
|
|
114
|
+
};
|
|
115
|
+
});
|
|
116
|
+
|
|
117
|
+
results.push(metrics);
|
|
118
|
+
await browser.close();
|
|
119
|
+
}
|
|
120
|
+
|
|
121
|
+
return results;
|
|
122
|
+
}
|
|
123
|
+
|
|
124
|
+
// Compute median across runs
|
|
125
|
+
function median(values) {
|
|
126
|
+
const sorted = [...values].sort((a, b) => a - b);
|
|
127
|
+
const mid = Math.floor(sorted.length / 2);
|
|
128
|
+
return sorted.length % 2 ? sorted[mid] : (sorted[mid - 1] + sorted[mid]) / 2;
|
|
129
|
+
}
|
|
130
|
+
|
|
131
|
+
function p95(values) {
|
|
132
|
+
const sorted = [...values].sort((a, b) => a - b);
|
|
133
|
+
return sorted[Math.ceil(sorted.length * 0.95) - 1];
|
|
134
|
+
}
|
|
135
|
+
|
|
136
|
+
module.exports = { benchmark, median, p95 };
|
|
137
|
+
```
|
|
138
|
+
|
|
139
|
+
## Usage in /benchmark Command
|
|
140
|
+
|
|
141
|
+
The command generates a runner script that:
|
|
142
|
+
|
|
143
|
+
1. Imports the template functions
|
|
144
|
+
2. Calls `benchmark(url, options)` for each target URL
|
|
145
|
+
3. Computes median/p95 across runs
|
|
146
|
+
4. Compares against baseline (if exists)
|
|
147
|
+
5. Checks against budget (if exists)
|
|
148
|
+
6. Outputs structured JSON
|
|
149
|
+
|
|
150
|
+
## CPU Throttling Profiles
|
|
151
|
+
|
|
152
|
+
| Profile | CPU Rate | Network | Use Case |
|
|
153
|
+
|---------|----------|---------|----------|
|
|
154
|
+
| Desktop (default) | 1x | No throttle | Standard desktop experience |
|
|
155
|
+
| Mobile (`--mobile`) | 4x | 1.5 Mbps / 40ms | Mobile device simulation |
|
|
156
|
+
| Slow 3G (`--3g`) | 4x | 400 Kbps / 400ms | Worst-case mobile |
|
|
157
|
+
|
|
158
|
+
## Console Error Capture
|
|
159
|
+
|
|
160
|
+
The script also captures console errors during page load — a console error during benchmark indicates a broken page, not just a slow one:
|
|
161
|
+
|
|
162
|
+
```javascript
|
|
163
|
+
const consoleErrors = [];
|
|
164
|
+
page.on('console', msg => {
|
|
165
|
+
if (msg.type() === 'error') consoleErrors.push(msg.text());
|
|
166
|
+
});
|
|
167
|
+
```
|
|
168
|
+
|
|
169
|
+
Console errors are included in the output as `"errors": [...]` for the report.
|
|
@@ -0,0 +1,265 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: performance-metrics
|
|
3
|
+
description: Log task completion data to .claude/metrics/. Use at the end of every task to record tokens, cost, agents used, rework cycles, and hallucination events. Also use for periodic reporting to identify efficiency and quality trends.
|
|
4
|
+
role: worker
|
|
5
|
+
user-invocable: true
|
|
6
|
+
---
|
|
7
|
+
|
|
8
|
+
# Performance Metrics
|
|
9
|
+
|
|
10
|
+
## Overview
|
|
11
|
+
|
|
12
|
+
Schema and procedures for capturing performance data in `.claude/metrics/`. Metrics enable evidence-based evaluation of agent effectiveness, cost efficiency, and quality outcomes.
|
|
13
|
+
|
|
14
|
+
## Constraints
|
|
15
|
+
|
|
16
|
+
- Never log credentials, API keys, or PII in metric entries
|
|
17
|
+
- Log entries are append-only; do not modify or delete existing JSONL records
|
|
18
|
+
- Log at task completion, not mid-task; mid-task state belongs in `.claude/memory/` progress files
|
|
19
|
+
- Use the defined JSONL schema; do not invent new top-level fields without updating the reference
|
|
20
|
+
|
|
21
|
+
## Metric Categories
|
|
22
|
+
|
|
23
|
+
> **Targets discipline.** Per CLAUDE.md → "Claims discipline", a numeric target
|
|
24
|
+
> may only ship if an instrument measures it. The instrumented metrics below cite
|
|
25
|
+
> their sensor; the rest read "Aspirational — no sensor yet" until one exists
|
|
26
|
+
> (tracked by #102 cost metering and #106 telemetry). Do not reintroduce bare
|
|
27
|
+
> numeric targets — `tests/docs/prose_honesty_test.bats` enforces this across all
|
|
28
|
+
> shipped prose.
|
|
29
|
+
|
|
30
|
+
### Efficiency Metrics
|
|
31
|
+
|
|
32
|
+
| Metric | Description | Target |
|
|
33
|
+
| --- | --- | --- |
|
|
34
|
+
| Task completion time | Wall-clock time from request to delivery | Track trend, no fixed target |
|
|
35
|
+
| Token usage per task | Total input + output tokens consumed | Minimize for comparable quality |
|
|
36
|
+
| Agent loading overhead | Tokens spent on agent/skill file reads | Aspirational — no sensor yet |
|
|
37
|
+
| Context summarization frequency | How often summarization triggers per task | Aspirational — no sensor yet |
|
|
38
|
+
|
|
39
|
+
### Quality Metrics
|
|
40
|
+
|
|
41
|
+
| Metric | Description | Target |
|
|
42
|
+
| --- | --- | --- |
|
|
43
|
+
| First-pass acceptance rate | Tasks accepted without rework | Aspirational — no sensor yet |
|
|
44
|
+
| Rework count | Number of revision cycles per task | Aspirational — no sensor yet |
|
|
45
|
+
| Hallucination incidents | Outputs containing fabricated information | Aspirational — no sensor yet |
|
|
46
|
+
| Accuracy score | Correctness of structured data extraction | Aspirational — no sensor yet |
|
|
47
|
+
| Test coverage | Percentage of code covered by generated tests | Track per project |
|
|
48
|
+
|
|
49
|
+
### Cost Metrics
|
|
50
|
+
|
|
51
|
+
| Metric | Description | Target |
|
|
52
|
+
| --- | --- | --- |
|
|
53
|
+
| Cost per task | Total API cost (input + output tokens at rate) | Track trend |
|
|
54
|
+
| LLM routing ratio | Percentage of tasks routed to each LLM | Track distribution |
|
|
55
|
+
| Selective loading savings | Tokens saved vs. loading all agents | Aspirational — no sensor yet |
|
|
56
|
+
|
|
57
|
+
## Log Format
|
|
58
|
+
|
|
59
|
+
Metrics are stored in `.claude/metrics/` as JSONL files (one JSON object per line).
|
|
60
|
+
|
|
61
|
+
### File Naming
|
|
62
|
+
|
|
63
|
+
```
|
|
64
|
+
.claude/metrics/{date}-task-log.jsonl
|
|
65
|
+
```
|
|
66
|
+
|
|
67
|
+
Example: `.claude/metrics/2026-02-20-task-log.jsonl`
|
|
68
|
+
|
|
69
|
+
### Task Completion Entry
|
|
70
|
+
|
|
71
|
+
Logged at the end of each task:
|
|
72
|
+
|
|
73
|
+
```json
|
|
74
|
+
{
|
|
75
|
+
"timestamp": "2026-02-20T14:30:00Z",
|
|
76
|
+
"task_id": "unique-id",
|
|
77
|
+
"task_type": "implementation",
|
|
78
|
+
"task_description": "Build REST API for user authentication",
|
|
79
|
+
"agents_used": ["software-engineer", "architect"],
|
|
80
|
+
"skills_used": ["hexagonal-architecture"],
|
|
81
|
+
"tokens": {
|
|
82
|
+
"input": 12500,
|
|
83
|
+
"output": 3200,
|
|
84
|
+
"total": 15700
|
|
85
|
+
},
|
|
86
|
+
"cost_usd": 0.043,
|
|
87
|
+
"llm": "opus",
|
|
88
|
+
"handoffs": 0,
|
|
89
|
+
"phases": 2,
|
|
90
|
+
"rework_cycles": 1,
|
|
91
|
+
"accepted": true,
|
|
92
|
+
"hallucination_detected": false,
|
|
93
|
+
"duration_seconds": 180
|
|
94
|
+
}
|
|
95
|
+
```
|
|
96
|
+
|
|
97
|
+
### Field Reference
|
|
98
|
+
|
|
99
|
+
| Field | Type | Description |
|
|
100
|
+
| --- | --- | --- |
|
|
101
|
+
| `timestamp` | string | ISO 8601 completion time |
|
|
102
|
+
| `task_id` | string | Unique identifier for this task |
|
|
103
|
+
| `task_type` | string | `implementation`, `design`, `bugfix`, `testing`, `documentation`, `analysis` |
|
|
104
|
+
| `task_description` | string | Brief description of the task |
|
|
105
|
+
| `agents_used` | string[] | Agent names that were loaded |
|
|
106
|
+
| `skills_used` | string[] | Skill names that were loaded |
|
|
107
|
+
| `tokens.input` | number | Input tokens from API usage field |
|
|
108
|
+
| `tokens.output` | number | Output tokens from API usage field |
|
|
109
|
+
| `tokens.total` | number | Sum of input + output |
|
|
110
|
+
| `cost_usd` | number | Estimated cost based on token rates |
|
|
111
|
+
| `llm` | string | Model ID used |
|
|
112
|
+
| `handoffs` | number | Times the handoff skill was triggered (either mode) |
|
|
113
|
+
| `phases` | number | Number of loading phases |
|
|
114
|
+
| `rework_cycles` | number | Number of revision cycles |
|
|
115
|
+
| `accepted` | boolean | Whether the user accepted the output |
|
|
116
|
+
| `hallucination_detected` | boolean | Whether a hallucination was flagged |
|
|
117
|
+
| `duration_seconds` | number | Wall-clock seconds from start to delivery |
|
|
118
|
+
|
|
119
|
+
### Cost Metering Entry (#102, #1094)
|
|
120
|
+
|
|
121
|
+
The `Stop`/`SubagentStop` hook (`hooks/cost_meter.py` →
|
|
122
|
+
`hooks/lib/cost_meter.py record`) appends one entry per fire to
|
|
123
|
+
`.claude/metrics/cost-metering.jsonl` — the running per-session token/cost summary
|
|
124
|
+
parsed from the transcript. Full schema reference:
|
|
125
|
+
`knowledge/telemetry-schema.md`.
|
|
126
|
+
|
|
127
|
+
```json
|
|
128
|
+
{
|
|
129
|
+
"timestamp": "2026-07-18T14:30:00Z",
|
|
130
|
+
"transcript": "session.jsonl",
|
|
131
|
+
"total": {"input_tokens": 18000, "output_tokens": 3500, "cost_usd": 0.16, "messages": 12},
|
|
132
|
+
"by_model": {"<model-id>": {"cost_usd": 0.16, "input_tokens": 18000, "output_tokens": 3500}},
|
|
133
|
+
"by_thread": {"main": {"cost_usd": 0.10, "input_tokens": 10000, "output_tokens": 2000},
|
|
134
|
+
"subagent": {"cost_usd": 0.06, "input_tokens": 8000, "output_tokens": 1500}},
|
|
135
|
+
"by_agent_type": {"main": {"cost_usd": 0.10, "input_tokens": 10000, "output_tokens": 2000},
|
|
136
|
+
"security-review": {"cost_usd": 0.06, "input_tokens": 8000, "output_tokens": 1500}}
|
|
137
|
+
}
|
|
138
|
+
```
|
|
139
|
+
|
|
140
|
+
| Field | Type | Description |
|
|
141
|
+
| --- | --- | --- |
|
|
142
|
+
| `total` | object | Session-cumulative token counts, `cost_usd`, `messages` |
|
|
143
|
+
| `by_model` | object | Per-model slim breakdown (`cost_usd`, `input_tokens`, `output_tokens`) |
|
|
144
|
+
| `by_thread` | object | `main` vs `subagent`, from the native `isSidechain` flag |
|
|
145
|
+
| `by_agent_type` | object | `main` for main-loop turns; sidechain turns keyed by subagent type via the harness-recorded `attributionAgent` field or the Task-dispatch `subagent_type`/`agentId` join; unmappable sidechain spend lands in `unattributed` (#1094) |
|
|
146
|
+
|
|
147
|
+
**Privacy:** token counts, dollar amounts, model identifiers, and thread/
|
|
148
|
+
agent-type identifiers only — never prompt text, code, file paths, or tool
|
|
149
|
+
payloads. Attribution reads only fields the Claude Code harness itself writes
|
|
150
|
+
to the transcript; the meter never guesses (see #170 for the buckets removed
|
|
151
|
+
because the harness records no signal for them). Disable with
|
|
152
|
+
`DEV_TEAM_COST_METER=off`. Report it with `/cost-report`.
|
|
153
|
+
|
|
154
|
+
### Review Value Entry (#348)
|
|
155
|
+
|
|
156
|
+
`/build` appends one entry per **inline review checkpoint** to
|
|
157
|
+
`.claude/metrics/review-value.jsonl` so the pipeline's review overhead becomes
|
|
158
|
+
*measurable* — distinguishing a build where review caught and fixed a real defect
|
|
159
|
+
from one where every loop passed no-op. This is the sensor that lets the plan/step
|
|
160
|
+
tiering (the `/plan` plan-tier and `/build` per-step complexity routing) be
|
|
161
|
+
right-sized with evidence rather than guessed.
|
|
162
|
+
|
|
163
|
+
```json
|
|
164
|
+
{
|
|
165
|
+
"timestamp": "2026-06-22T14:30:00Z",
|
|
166
|
+
"plan": "plans/add-auth.md",
|
|
167
|
+
"slice": "2",
|
|
168
|
+
"step": "all",
|
|
169
|
+
"checkpoint": "slice",
|
|
170
|
+
"complexity": "standard",
|
|
171
|
+
"source": "build-checkpoint",
|
|
172
|
+
"agents_run": ["spec-compliance-review", "security-review"],
|
|
173
|
+
"issues_found": 1,
|
|
174
|
+
"severity_breakdown": {"errors": 1, "warnings": 0, "suggestions": 0},
|
|
175
|
+
"issues_fixed": 1,
|
|
176
|
+
"fix_iterations": 1,
|
|
177
|
+
"outcome": "fixed"
|
|
178
|
+
}
|
|
179
|
+
```
|
|
180
|
+
|
|
181
|
+
| Field | Type | Description |
|
|
182
|
+
| --- | --- | --- |
|
|
183
|
+
| `checkpoint` | string | `step` (per-step `complex` review), `slice` (batched slice-boundary review), or `backstop` (the Step-6 backstop pass, #1962) |
|
|
184
|
+
| `source` | string | Row provenance — `build-checkpoint` (fix-applying `/build` inline checkpoint, the default when absent), `build-backstop` (fix-applying `/build` Step-6 backstop, #1962), or `code-review` (read-only standalone review). `/harness-audit` Step 4 excludes `code-review` rows from fix-rate drop-candidate logic (#1257); `build-backstop` rows are fix-applying and stay in it, and additionally drive the backstop-redundancy readout that gates `--backstop-review=skip` |
|
|
185
|
+
| `outcome` | string | `no-op`, `fixed`, `escalated`, or — backstop rows only — `skipped` when `--backstop-review=skip` suppressed the pass. A `skipped` row records the suppression and is excluded from every rate (#1962) |
|
|
186
|
+
| `diff_shape` | string | `test-only` or `mixed`, from `change_shape.py`'s `isTestOnly` (never eyeballed). Splits per-lens outcomes by diff shape so a test-only lens gate can be justified by measurement rather than intuition (#1964) |
|
|
187
|
+
| `step` | string | `N.M` for a per-step checkpoint; `all` for a batched slice checkpoint |
|
|
188
|
+
| `agents_run` | string[] | Review agents and static-analysis lane tools run at this checkpoint |
|
|
189
|
+
| `issues_found` | number | Actionable issues the checkpoint surfaced (semantic review + static lanes) |
|
|
190
|
+
| `severity_breakdown` | object | `{errors, warnings, suggestions}` counts (same enum as `/code-review`); the three sum to `issues_found`. Lets `/harness-audit` Step 3 flag low-value (mostly-minor) lenses (#1256) |
|
|
191
|
+
| `issues_fixed` | number | Of those, how many were auto-fixed |
|
|
192
|
+
| `fix_iterations` | number | Review-fix and static self-heal fix-loop iterations consumed |
|
|
193
|
+
| `outcome` | string | `no-op` (passed clean), `fixed` (found + fixed), `escalated` (a fix loop didn't converge — including a static lane capping out) |
|
|
194
|
+
|
|
195
|
+
**Privacy:** counts and outcomes only — never prompt text, code, or file content,
|
|
196
|
+
consistent with the cost meter's privacy boundary. Disable with
|
|
197
|
+
`DEV_TEAM_REVIEW_VALUE=off`. Report it with `/cost-report` (its "review value"
|
|
198
|
+
section).
|
|
199
|
+
|
|
200
|
+
### Verify Log Entry (#727)
|
|
201
|
+
|
|
202
|
+
`/build` appends one entry per **slice with a runtime surface** to
|
|
203
|
+
`metrics/verify-log.jsonl` (sub-step 4.9) — evidence that the project's own
|
|
204
|
+
test/verification tooling actually exercised the change end-to-end before the
|
|
205
|
+
slice was marked done, or was explicitly skipped because the diff had no
|
|
206
|
+
runtime surface to drive. This is the sensor that closes the gap the #727
|
|
207
|
+
investigation found: a "done" feature that fails the first time it's really
|
|
208
|
+
run means runtime verification was either skipped or never mandated in the
|
|
209
|
+
first place.
|
|
210
|
+
|
|
211
|
+
```json
|
|
212
|
+
{
|
|
213
|
+
"timestamp": "2026-07-02T14:30:00Z",
|
|
214
|
+
"plan": "plans/add-auth.md",
|
|
215
|
+
"slice": "2",
|
|
216
|
+
"branch": "feat/add-auth",
|
|
217
|
+
"files": ["src/api/auth.py"],
|
|
218
|
+
"outcome": "ran",
|
|
219
|
+
"reason": null
|
|
220
|
+
}
|
|
221
|
+
```
|
|
222
|
+
|
|
223
|
+
| Field | Type | Description |
|
|
224
|
+
| --- | --- | --- |
|
|
225
|
+
| `branch` | string | Current branch name — `${CLAUDE_PLUGIN_ROOT}/scripts/progress_guardian.py --pre-pr` matches on this |
|
|
226
|
+
| `files` | string[] | The slice's changed runtime files the verification was scoped to |
|
|
227
|
+
| `outcome` | string | `ran` (the verification ran and passed), `skipped` (no runtime surface — see `reason`), or `failed-then-fixed` (the verification failed at least once before the fix landed) |
|
|
228
|
+
| `reason` | string \| null | Required when `outcome` is `skipped` (e.g. `"tests-only"`, `"docs-only"`); `null` otherwise |
|
|
229
|
+
|
|
230
|
+
**Not disableable.** Unlike Review Value logging (`DEV_TEAM_REVIEW_VALUE=off`),
|
|
231
|
+
there is no env var to turn this off — `${CLAUDE_PLUGIN_ROOT}/scripts/progress_guardian.py --pre-pr`
|
|
232
|
+
fails closed on a branch with runtime-surface changes and no matching entry,
|
|
233
|
+
and that gate is a correctness control, not a metrics-collection nicety.
|
|
234
|
+
|
|
235
|
+
## When to Log
|
|
236
|
+
|
|
237
|
+
> **Task completion entries are written automatically** by `hooks/task_completion_metrics.py`
|
|
238
|
+
> on every `Stop` / `SubagentStop` event — skills no longer need a manual
|
|
239
|
+
> "log this on completion" step. To surface task-specific values (task type,
|
|
240
|
+
> agents used, hallucination flag, rework count, defect count, or a config
|
|
241
|
+
> change), populate `.claude/session-metrics.json` before the session ends; the
|
|
242
|
+
> hook reads and clears it. Leave the file absent for a minimal heartbeat entry.
|
|
243
|
+
|
|
244
|
+
| Event | Action |
|
|
245
|
+
| --- | --- |
|
|
246
|
+
| Task completed | Logged automatically by `hooks/task_completion_metrics.py` (no manual step needed) |
|
|
247
|
+
| `/build` inline review checkpoint | Append a Review Value entry to `.claude/metrics/review-value.jsonl` (#348) |
|
|
248
|
+
| `/build` slice with a runtime surface | Append a Verify Log entry to `metrics/verify-log.jsonl` (#727) |
|
|
249
|
+
| Configuration change | Add `config_change` to `.claude/session-metrics.json`; hook writes `.claude/metrics/config-changelog.jsonl` |
|
|
250
|
+
| Hallucination detected | Set `hallucination_detected: true` in `.claude/session-metrics.json` |
|
|
251
|
+
| Context summarization triggered | Increment counter in current task entry |
|
|
252
|
+
|
|
253
|
+
## Output
|
|
254
|
+
|
|
255
|
+
JSONL log entries written to `.claude/metrics/` and/or a summary report of metric trends. Be concise — report anomalies and trend signals; omit entries within normal range.
|
|
256
|
+
|
|
257
|
+
## Reporting
|
|
258
|
+
|
|
259
|
+
Periodically review metrics to identify patterns:
|
|
260
|
+
|
|
261
|
+
1. **Weekly**: Review task completion entries for rework trends and hallucination rate
|
|
262
|
+
2. **Monthly**: Aggregate cost metrics and LLM routing distribution
|
|
263
|
+
3. **Per-project**: Compare first-pass acceptance rate across task types
|
|
264
|
+
|
|
265
|
+
Summaries can be written to `.dev-team-reports/` for historical reference.
|
|
@@ -0,0 +1,199 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: plan
|
|
3
|
+
description: >-
|
|
4
|
+
Create a structured implementation plan with goal, acceptance criteria,
|
|
5
|
+
incremental Code-First Small Batches steps, and a pre-PR quality gate. Use this for tasks that
|
|
6
|
+
need a plan but not the full three-phase orchestration, or when the user
|
|
7
|
+
says "plan this", "make a plan", "break this down", or "how should I
|
|
8
|
+
implement this".
|
|
9
|
+
argument-hint: "<task-description> [--output <path>] [--yes] [--spec-issue <url>]"
|
|
10
|
+
user-invocable: true
|
|
11
|
+
allowed-tools: Read, Write, Glob, Grep, Bash(mkdir *), Bash(date *), Bash(git branch *), Bash(test *), Bash(gh issue view *), AskUserQuestion
|
|
12
|
+
---
|
|
13
|
+
|
|
14
|
+
# Plan
|
|
15
|
+
|
|
16
|
+
Role: orchestrator. This command creates a structured plan — it does not implement anything.
|
|
17
|
+
|
|
18
|
+
You have been invoked with the `/plan` command.
|
|
19
|
+
|
|
20
|
+
## Orchestrator constraints
|
|
21
|
+
|
|
22
|
+
**MUST — confirm agent-dispatch capability before Step 5b (issue #1461).** Step 5b below dispatches the five `plan-review-*` critics as parallel sub-agents. Before that dispatch, you MUST confirm the `Agent` (or `Task`) tool is actually present and available in your current toolset. If it is not: **STOP.** Do not self-apply the plan-review critics' checklists inline as a substitute for independent dispatch, and do not present the plan as reviewed or approved on that basis — a self-certified pass is not a review. Report to the user/operator plainly: plan review cannot run in this environment because no agent-dispatch capability (`Agent`/`Task` tool) is available; name exactly what's missing; and state that `/plan` cannot complete Step 5b, and the plan cannot reach the human approval gate (Step 6), until it is re-run from a session with that capability. This is a hard requirement, not a preference.
|
|
23
|
+
|
|
24
|
+
1. **Do not implement.** Produce only the plan. No code, no scaffolding, no file edits beyond the plan file itself. One narrow carve-out: **after approval**, derived `.feature` files are written via the export script `plan_gherkin_export.py` (step 6) — never before approval, never by hand.
|
|
25
|
+
2. **Every step is one behavior with a full cycle.** The cadence is Code-First Small Batches — each step follows IMPLEMENT → TEST → REFACTOR (`docs/experiments/RECOMMENDATIONS.md` Rec 3). The refactor runs in every cycle.
|
|
26
|
+
3. **Incremental.** Each step must leave the codebase in a working, committable state.
|
|
27
|
+
4. **Human approval required.** Present the plan for approval before any implementation begins.
|
|
28
|
+
5. **Be concise.** The plan is the artifact; keep chat to decisions and gaps.
|
|
29
|
+
6. **State your approach stance.** For any high-reversal-cost axis in `${CLAUDE_PLUGIN_ROOT}/knowledge/decision-defaults.md` the task touches (replace-vs-merge, format fidelity, migrate-vs-edit-stub, auto-merge-vs-direct, scope), state the chosen stance explicitly in the plan so it is visible — and correctable — at the human gate.
|
|
30
|
+
|
|
31
|
+
## Parse Arguments
|
|
32
|
+
|
|
33
|
+
Arguments: $ARGUMENTS
|
|
34
|
+
|
|
35
|
+
- Positional: task description (required)
|
|
36
|
+
- `--output <path>`: Write plan to a specific path. Default: `plans/<slugified-task>.md`
|
|
37
|
+
- `--yes`: Auto-approve the plan without prompting (non-interactive opt-in; see step 6).
|
|
38
|
+
- `--spec-issue <url>`: Treat the given GitHub issue as the spec source, bypassing file discovery in step 1 — supplied by `/specs`' GitHub-issue persistence branch, or usable directly by an operator pointing `/plan` at an existing spec issue.
|
|
39
|
+
|
|
40
|
+
## Steps
|
|
41
|
+
|
|
42
|
+
### 1. Check for spec artifacts
|
|
43
|
+
|
|
44
|
+
**If `--spec-issue <url>` was supplied**, skip file discovery entirely: run `gh issue view <url> --json title,body` and extract Intent Description, Architecture Specification, Acceptance Criteria, and Ambiguity Log directly from the issue body — the same headings `/specs`' GitHub-issue persistence branch writes (Ambiguity Log is one more than the file path's three-artifact minimum, so ambiguities `/specs` already resolved with the human stay visible here). This narrowly scoped, read-only call is warranted directly in this step — unlike step 6's "offer GitHub issues" section, which delegates issue *creation* to `/issues-from-plan` — because this step runs before any plan content exists, so there is nothing yet to delegate to. If `gh issue view <url>` exits non-zero (deleted issue, bad URL, expired auth, network down), report the failure and cause to chat — do not silently proceed as if unsupplied — and fall back to the file search below. **Otherwise**, search for specification artifacts produced by `/specs` — look for files matching `docs/specs/**` or `specs/**` related to the task. Check for the three artifacts: Intent Description, Architecture Specification, and Acceptance Criteria. The spec does **not** contain Gherkin — authoring the behavioral scenarios is this command's job. If no spec artifacts are found (by either path), ask the user: "No specification artifacts found for this task. Run `/specs` first to produce them, or continue planning without specs?"
|
|
45
|
+
|
|
46
|
+
**Non-interactive** (per step 6's interactivity rule): do not prompt — log
|
|
47
|
+
`No spec artifacts found — continuing without specs (non-interactive).`, proceed, and
|
|
48
|
+
record the absence in the plan's `## Risks & Open Questions` so it reaches the PR's
|
|
49
|
+
Decisions & Assumptions section.
|
|
50
|
+
|
|
51
|
+
If the user chooses to continue without specs, proceed. Otherwise, stop and let them run `/specs` first.
|
|
52
|
+
|
|
53
|
+
### 2. Understand the task and cut the slices
|
|
54
|
+
|
|
55
|
+
Read relevant code and context to understand what needs to change. Keep exploration focused — this is planning, not research. If the task is complex enough to need deep research, suggest `/design-doc` instead. If spec artifacts exist, use them as the primary source for goals, constraints, and acceptance criteria. Prefer CodeGraph (`codegraph_explore`/`codegraph_impact`) or Repowise (`get_context`/`search_codebase`) over raw `Grep`/`Read` for this — cheaper blast-radius scoping when either is present; fall back to `Read`/`Grep`/`Glob` when not.
|
|
56
|
+
|
|
57
|
+
Then **decompose the feature into vertical slices**. A slice is a vertically deliverable increment — independently testable and, ideally, independently shippable. Sequence slices so trunk stays releasable at every step: land incomplete behavior behind a feature toggle or an abstraction, and order any data change as expand-before-contract rather than a single breaking migration (see `knowledge/release-strategies.md` and `knowledge/database-change-management.md`). For each slice, **author the Gherkin scenario(s)** that define its observable behavior. This is where the behavioral contract is written; the spec only described the change and its goals.
|
|
58
|
+
|
|
59
|
+
When authoring each slice's Gherkin, cover:
|
|
60
|
+
|
|
61
|
+
- **Happy path** — the primary success behavior.
|
|
62
|
+
- **Negative cases** — invalid, unauthorized, missing, or malformed input.
|
|
63
|
+
- **Edge cases** — empty collections, boundary values, concurrent access, idempotency.
|
|
64
|
+
- **Error scenarios** — specify observable error behavior, not just "should fail".
|
|
65
|
+
|
|
66
|
+
Keep scenarios implementation-independent (no databases, selectors, or internal data structures in step text) and deterministic. Every acceptance criterion from the spec must be covered by at least one scenario across the slices. Each step traces back to one or more scenarios in its slice.
|
|
67
|
+
|
|
68
|
+
**Depth-audit finding (issue #1450): no behavior change here.** `/gherkin-derive` was made to mandatorily analyze controllers, handlers, services, domain logic, workflows, validation rules, error handling, and business processes in depth, because it derives scenarios directly from a codebase with no prior scoping. This step does not: slice authoring here is deliberately spec-scoped (the spec's Intent/Architecture/Acceptance Criteria are the primary source — step 2 above) and exploration-bounded ("this is planning, not research" — step 2 above). Widening this step into full controller/service/domain-logic discovery would contradict its own designed scope and duplicate work the spec phase already did.
|
|
69
|
+
|
|
70
|
+
**Filter `LOW_VALUE` findings out of the work streams.** A gap classified `LOW_VALUE` by `/specs` or `/test-health` (no branching logic, no observable outcome, coverage already provided by a higher-layer test) never becomes a slice or a step. List such findings in a `## Skipped (low value)` section with their one-line rationale so the decision is visible — they document why the work was *not* planned, not deferred work to revisit. **The plan gate rejects any slice or step classified `LOW_VALUE`**: if a `LOW_VALUE` finding has leaked into a work stream, move it to the Skipped section before the human gate.
|
|
71
|
+
|
|
72
|
+
### 3. Create the plan
|
|
73
|
+
|
|
74
|
+
Write the plan file using the structure in `references/plan-template.md` (goal,
|
|
75
|
+
acceptance criteria, per-slice Gherkin + Code-First Small Batches steps, parallelization DAG, complexity
|
|
76
|
+
classification, pre-PR gate, skipped-findings, risks, and the machine-parseable
|
|
77
|
+
Build Progress section).
|
|
78
|
+
|
|
79
|
+
#### Resolve the Gherkin persistence decision
|
|
80
|
+
|
|
81
|
+
Follow `references/gherkin-persistence.md`: honor a re-run's recorded
|
|
82
|
+
decision, otherwise detect the project's BDD convention
|
|
83
|
+
(`scripts/detect_bdd_convention.py`, conservative precedence), prompt once —
|
|
84
|
+
or log a skip when non-interactive — and record + echo the resulting
|
|
85
|
+
`**Gherkin persistence**:` metadata line (destination directory only). No
|
|
86
|
+
`.feature` file is written before approval.
|
|
87
|
+
|
|
88
|
+
### 4. Create the plans directory
|
|
89
|
+
|
|
90
|
+
Create `plans/` if it doesn't exist. When writing the plan file, populate the `## Build Progress` section by copying slice and step titles from `## Slices`. These are the checkboxes `/build` will update on disk as each step completes — a slice is checked off once all its steps are. Acceptance Criteria live only in the top-level `## Acceptance Criteria` section as PR-checklist material; the operator ticks each one at PR review time after behaviorally verifying it, not during build.
|
|
91
|
+
|
|
92
|
+
Then derive the waves — never hand-author them:
|
|
93
|
+
|
|
94
|
+
```bash
|
|
95
|
+
python3 "${CLAUDE_PLUGIN_ROOT}/scripts/plan_waves.py" <plan-file>
|
|
96
|
+
```
|
|
97
|
+
|
|
98
|
+
Render the `## Parallelization` Mermaid DAG + wave table and the wave-grouped
|
|
99
|
+
`## Build Progress` from its JSON (`waves`, per-slice `wave`). If it exits non-zero
|
|
100
|
+
(cycle, missing `Depends-on`, or unknown reference) or reports a `collisions` entry,
|
|
101
|
+
fix the plan and re-run before the human gate — those defeat safe concurrent build.
|
|
102
|
+
|
|
103
|
+
**`scope_mismatches` (#865) is fix-or-acknowledge, not fix-before-gate.** For a slice
|
|
104
|
+
declaring slice-level `**Files:**`, this JSON array compares it against the union of
|
|
105
|
+
the slice's per-step `**Files**:` lines (`under_declared`/`over_declared`, `[]` when
|
|
106
|
+
matching or undeclared). Unlike `collisions`, don't block on it — surface it at the
|
|
107
|
+
human gate (step 6); if the author proceeds unreconciled, record it in `## Risks &
|
|
108
|
+
Open Questions`.
|
|
109
|
+
|
|
110
|
+
### 5. Run plan review personas
|
|
111
|
+
|
|
112
|
+
Before presenting to the user, dispatch the plan review personas in parallel as sub-agents. Each critically challenges the plan from a different perspective. **The reviewer set scales to plan complexity** — a one-function plan does not pay the same review ceremony as a complex feature (the fixed-overhead cost the TDD experiment surfaced; see `docs/experiments/01-final-results.md`).
|
|
113
|
+
|
|
114
|
+
#### 5a. Classify the plan tier
|
|
115
|
+
|
|
116
|
+
Derive a **plan tier** from objective signals already on hand — the same `trivial | standard | complex` vocabulary `/build` uses for per-step review depth, so the concept is consistent across the pipeline. Inputs: the slice count and wave structure from the `scripts/plan_waves.py` JSON, the file count, the per-step Complexity ratings, and whether the plan takes a stance on any high-reversal-cost axis in `knowledge/decision-defaults.md`.
|
|
117
|
+
|
|
118
|
+
| Tier | Signals | Reviewers |
|
|
119
|
+
| ------ | --------- | ----------- |
|
|
120
|
+
| `trivial` | 1 slice, ≤ 2 files, no `complex` step, touches no high-reversal-cost decision axis | **Acceptance Test Critic only** (1) |
|
|
121
|
+
| `standard` | anything between — e.g. a single slice with a few files, or a small multi-slice plan within existing patterns | **Acceptance Test Critic + Design & Architecture Critic**, plus **UX Critic** if the plan has a user-facing/UI surface, plus **Parallelization Critic** if slice count > 1 (2–4) |
|
|
122
|
+
| `complex` | > 1 wave, ≥ 4 slices, any `complex` step, a security-sensitive/cross-cutting change, or a **non-default** stance on a high-reversal-cost decision axis, or the axis was **contested** at the `/ship` gate | **all 5** |
|
|
123
|
+
|
|
124
|
+
Every `/ship`-driven plan states a stance on the axes in `knowledge/decision-defaults.md` — merely restating the default is not, by itself, a `complex` signal (treating it as one would defeat the tier system's own review-scaling goal). "Contested" means a recorded objection to the stance, e.g. a note in the plan's `## Risks & Open Questions` section — not an unrecorded verbal disagreement.
|
|
125
|
+
|
|
126
|
+
**Worked example** (`Integration: auto-merge vs. direct-to-trunk` axis): a plan stating "open a PR and use auto-merge gated on green checks" (the documented default) does not trigger `complex` on this signal alone. A plan stating "merge directly to trunk, bypassing the PR gate" (a non-default stance) does trigger `complex`, as does a plan stating the default stance where a reviewer's recorded objection challenges it.
|
|
127
|
+
|
|
128
|
+
When in doubt, classify up (standard rather than trivial, complex rather than standard).
|
|
129
|
+
|
|
130
|
+
**Parallelization Critic gate (all tiers):** the Parallelization Critic only finds same-wave file collisions and disjoint-file coupling — a **single-slice plan has no waves to parallelize and no same-wave collisions by construction**, so it is a guaranteed no-op there. Run it **only when slice count > 1**, regardless of tier, and log the skip (`Parallelization Critic skipped — single-slice plan`) when it is omitted.
|
|
131
|
+
|
|
132
|
+
#### 5b. Dispatch the selected reviewers
|
|
133
|
+
|
|
134
|
+
Each persona is a registered agent — dispatch by `subagent_type` like any
|
|
135
|
+
other agent. The harness reads each one's own `model:`/`effort:`
|
|
136
|
+
frontmatter natively; there is no dispatch-time override to pass and no
|
|
137
|
+
plugin-side resolution step.
|
|
138
|
+
|
|
139
|
+
| Reviewer | Agent | Focus |
|
|
140
|
+
| ---------- | ---------- | ------- |
|
|
141
|
+
| Acceptance Test Critic | `plan-review-acceptance` | Per-slice Gherkin quality (determinism, isolation, implementation-independence), scenario gaps, error paths, criteria coverage, step traceability |
|
|
142
|
+
| Design & Architecture Critic | `plan-review-design` | Coupling, abstractions, structural risks, pattern adherence |
|
|
143
|
+
| UX Critic | `plan-review-ux` | User journey, error UX, cognitive load, accessibility |
|
|
144
|
+
| Strategic Critic | `plan-review-strategic` | Problem fit, scope, slice boundaries, risk, opportunity cost |
|
|
145
|
+
| Parallelization Critic | `plan-review-parallelization` | Same-wave independence: file-overlap collisions (from `scripts/plan_waves.py`), disjoint-file behavioral coupling, residual cycles/mis-layering |
|
|
146
|
+
|
|
147
|
+
Pass each reviewer the full plan content. Also pass the Parallelization Critic the `scripts/plan_waves.py` JSON for this plan (its `collisions` array is the deterministic input). Each returns a structured verdict (`approve` or `needs-revision`) with issues. The Acceptance Test Critic is the gate for the scenarios authored in step 2 — it validates the per-slice Gherkin the same way `feature-file-validation` would, so no separate scenario-review pass is needed before the human gate. It is the one reviewer that always runs (every tier). A `needs-revision` from the Parallelization Critic triggers plan revision (re-wave the colliding slices) before the human sees the plan.
|
|
148
|
+
|
|
149
|
+
**If any reviewer returns `needs-revision`**: Address all `blocker` issues by revising the plan. Re-run only the reviewers that flagged blockers. Repeat until all pass (max 2 iterations — escalate to user if still failing).
|
|
150
|
+
|
|
151
|
+
**After all pass**: Append a `## Plan Review Summary` section to the plan file with the aggregated findings (warnings and observations from the dispatched reviewers). **Record the chosen tier and the reviewer set at the top of that section** (e.g. `Plan tier: standard — reviewers: Acceptance, Design, Parallelization (UX skipped — no UI surface)`) so the scaling decision is visible and auditable.
|
|
152
|
+
|
|
153
|
+
### 6. Present for approval
|
|
154
|
+
|
|
155
|
+
**Gate check before presenting.** Verify the `## Skipped (low value)` section captures every `LOW_VALUE` finding — no slice or step may carry one. If a `LOW_VALUE` finding leaked into a work stream, move it to the Skipped section before the plan reaches the human.
|
|
156
|
+
|
|
157
|
+
**Surface any `scope_mismatches` (#865)** alongside the review summary, same visibility as `collisions` — never blocking. Unreconciled on approval → append it to `## Risks & Open Questions`.
|
|
158
|
+
|
|
159
|
+
**First determine interactivity.** The run is **non-interactive** when any of these hold: `--yes` was passed, `DEV_TEAM_AUTO_APPROVE=1` is set in the environment, or `DEV_TEAM_INTERACTIVE` is not `1` (pi sets it only when a human UI is attached; a tool shell never has a TTY, so `test -t 0` must not be used — the headless/CI/automation case). Otherwise it is **interactive**. This is the same non-interactive principle the GitHub-issue prompt below already follows; the approval gate now follows it too, so a headless `/plan`→`/build` run never hangs waiting for input.
|
|
160
|
+
|
|
161
|
+
- **Interactive** (unchanged from prior behavior) → Display the plan and the review summary. Ask: "Approve this plan to begin implementation, or suggest changes?" Mark the plan status as `approved` once the user confirms. If the user requests changes, update the plan and re-present. This is the Phase 2→3 gate: append an `approval` entry to `.claude/metrics/config-changelog.jsonl` per the [human-oversight-protocol § Audit trail](../human-oversight-protocol/SKILL.md#audit-trail) schema — `proposed` names the plan (e.g. "Approve plan for <task>"), `evidence_shown` points at the plan file (and any spec artifact), `risks_surfaced` lists the plan's `## Risks & Open Questions` items (`[]` if none).
|
|
162
|
+
- **Non-interactive** → do **not** prompt or block. Auto-approve: set `**Status**: approved` and append an explicit audit record to the plan so the bypass is **never silent** — add an `## Approval` section reading: `Auto-approved (non-interactive) at <date> — no human review gate. Trigger: <--yes | DEV_TEAM_AUTO_APPROVE=1 | no TTY>.` Then continue, appending the same three-field `.claude/metrics/config-changelog.jsonl` entry; `description` states the bypass trigger.
|
|
163
|
+
|
|
164
|
+
#### Post-approval: persist the Gherkin (`.feature` export)
|
|
165
|
+
|
|
166
|
+
If the plan's recorded `**Gherkin persistence**:` decision is a destination
|
|
167
|
+
directory, run the export **from the repo root** (destinations resolve against
|
|
168
|
+
the invocation cwd):
|
|
169
|
+
|
|
170
|
+
```bash
|
|
171
|
+
python3 "${CLAUDE_PLUGIN_ROOT}/scripts/plan_gherkin_export.py" <plan-file>
|
|
172
|
+
```
|
|
173
|
+
|
|
174
|
+
Show its summary (destination, files written, overwritten, stale removed) to
|
|
175
|
+
the operator. The `<dir>/<plan-slug>/` subdirectory is tool-owned — anything
|
|
176
|
+
inside is derived and overwritable; files outside it are never touched.
|
|
177
|
+
A non-zero exit is a failure: report it with the script's stderr — never claim
|
|
178
|
+
success on a failed export. `plan-file-only` decisions skip cleanly (the
|
|
179
|
+
script no-ops with a note).
|
|
180
|
+
|
|
181
|
+
#### Post-approval: offer GitHub issues (GitHub origin only)
|
|
182
|
+
|
|
183
|
+
After approval, classify the origin remote — **only** offer issue creation on an actual GitHub host:
|
|
184
|
+
|
|
185
|
+
```bash
|
|
186
|
+
python3 "${CLAUDE_PLUGIN_ROOT}/scripts/git_origin_host.py"
|
|
187
|
+
```
|
|
188
|
+
|
|
189
|
+
- **`github`** → prompt **once**, showing the count: *"Open 1 parent issue and N linked slice issues from this plan? [y/N]"* (N = number of slices). The default is **No**. Invoke `/issues-from-plan` **only on explicit `y`**; on No (or anything else), create nothing and continue.
|
|
190
|
+
- **`other`** (non-GitHub host, including lookalikes like `notgithub.example.com`) or **`none`** (no origin) → **no prompt**; continue silently.
|
|
191
|
+
- **Non-interactive** (`/plan` run without a usable TTY) → do **not** prompt or block: log *"skipping the GitHub issue prompt (non-interactive)"* and continue.
|
|
192
|
+
|
|
193
|
+
Never create issues without an explicit `y` on an interactive GitHub-origin prompt.
|
|
194
|
+
|
|
195
|
+
## Integration
|
|
196
|
+
|
|
197
|
+
- The progress-guardian agent tracks step completion against this plan
|
|
198
|
+
- `/continue` reads active plans to resume work
|
|
199
|
+
- The orchestrator's Phase 2 produces plans in this same format for larger tasks
|