pi-dev-team 0.2.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/LICENSE +21 -0
- package/PORTING.md +134 -0
- package/README.md +207 -0
- package/UPSTREAM.json +64 -0
- package/agents/Explore.md +15 -0
- package/agents/a11y-review.md +118 -0
- package/agents/adr-author.md +70 -0
- package/agents/ai-provenance-review.md +120 -0
- package/agents/angular-reactivity-review.md +95 -0
- package/agents/arch-review.md +135 -0
- package/agents/architect.md +78 -0
- package/agents/autoship-batch-proposer.md +69 -0
- package/agents/claude-setup-review.md +136 -0
- package/agents/codebase-recon.md +184 -0
- package/agents/component-architecture-review.md +119 -0
- package/agents/concurrency-review.md +109 -0
- package/agents/correctness-review.md +290 -0
- package/agents/data-flow-tracer.md +120 -0
- package/agents/doc-review.md +165 -0
- package/agents/domain-review.md +136 -0
- package/agents/general-purpose.md +10 -0
- package/agents/gherkin-quality-critic.md +113 -0
- package/agents/js-fp-review.md +114 -0
- package/agents/mutation-kill.md +684 -0
- package/agents/naming-review.md +142 -0
- package/agents/orchestrator.md +339 -0
- package/agents/performance-review.md +105 -0
- package/agents/plan-review-acceptance.md +115 -0
- package/agents/plan-review-design.md +90 -0
- package/agents/plan-review-parallelization.md +84 -0
- package/agents/plan-review-strategic.md +96 -0
- package/agents/plan-review-ux.md +110 -0
- package/agents/platform-engineer.md +64 -0
- package/agents/product-manager.md +68 -0
- package/agents/progress-guardian.md +79 -0
- package/agents/qa-engineer.md +289 -0
- package/agents/quality-reviewer.md +132 -0
- package/agents/react-reactivity-review.md +102 -0
- package/agents/refactor-opportunity-review.md +128 -0
- package/agents/security-engineer.md +60 -0
- package/agents/security-review.md +218 -0
- package/agents/session-analysis.md +95 -0
- package/agents/software-engineer.md +105 -0
- package/agents/spec-compliance-review.md +100 -0
- package/agents/spec-reviewer.md +114 -0
- package/agents/structure-review.md +146 -0
- package/agents/tech-writer.md +84 -0
- package/agents/test-review.md +246 -0
- package/agents/test-smell-review.md +188 -0
- package/agents/token-efficiency-review.md +139 -0
- package/agents/ui-ux-designer.md +54 -0
- package/agents/vue-reactivity-review.md +95 -0
- package/bin/__pycache__/claudecpython-314.pyc +0 -0
- package/bin/claude +258 -0
- package/docs/upstream/.pages +1 -0
- package/docs/upstream/CHANGELOG.md +2586 -0
- package/docs/upstream/README.md +155 -0
- package/docs/upstream/agent-architecture.md +214 -0
- package/docs/upstream/agent_info.md +187 -0
- package/docs/upstream/artifact-migration.md +124 -0
- package/docs/upstream/code-intelligence-nudge.md +149 -0
- package/docs/upstream/code-review-process.md +294 -0
- package/docs/upstream/concurrent-use.md +73 -0
- package/docs/upstream/context-management.md +111 -0
- package/docs/upstream/developer-notes.md +280 -0
- package/docs/upstream/diagrams/architecture-overview.svg +101 -0
- package/docs/upstream/diagrams/review-dispatch.svg +139 -0
- package/docs/upstream/diagrams/team-agents.svg +128 -0
- package/docs/upstream/diagrams/test-improve-flow.svg +166 -0
- package/docs/upstream/diagrams/workflow-linear.svg +66 -0
- package/docs/upstream/diagrams/workflow-three-phase.svg +200 -0
- package/docs/upstream/eval-maintenance.md +95 -0
- package/docs/upstream/eval-running-guide.md +147 -0
- package/docs/upstream/eval-system.md +291 -0
- package/docs/upstream/session-review-oss-complements.md +75 -0
- package/docs/upstream/session-review.md +212 -0
- package/docs/upstream/skills.md +188 -0
- package/docs/upstream/team-structure.md +21 -0
- package/docs/upstream/telemetry-ci-access.md +129 -0
- package/docs/upstream/telemetry-repo-security.md +120 -0
- package/docs/upstream/test-evaluation.md +277 -0
- package/docs/upstream/test-improve.md +154 -0
- package/docs/upstream/triage-workflow.md +282 -0
- package/docs/upstream/workflows.md +289 -0
- package/extensions/dev-team/index.ts +539 -0
- package/extensions/dev-team/lib/agents.ts +272 -0
- package/extensions/dev-team/lib/ai-credits.ts +92 -0
- package/extensions/dev-team/lib/autocompact.ts +81 -0
- package/extensions/dev-team/lib/child-run.ts +102 -0
- package/extensions/dev-team/lib/config.ts +236 -0
- package/extensions/dev-team/lib/gh-command.ts +103 -0
- package/extensions/dev-team/lib/github-style.ts +307 -0
- package/extensions/dev-team/lib/hooks.ts +350 -0
- package/extensions/dev-team/lib/metrics.ts +115 -0
- package/extensions/dev-team/lib/safe-read.ts +49 -0
- package/extensions/dev-team/lib/session-files.ts +57 -0
- package/extensions/dev-team/lib/session-spend.ts +123 -0
- package/extensions/dev-team/lib/shell-scan.ts +205 -0
- package/extensions/dev-team/lib/skills.ts +213 -0
- package/extensions/dev-team/lib/subagent-render.ts +245 -0
- package/extensions/dev-team/lib/subagent-types.ts +164 -0
- package/extensions/dev-team/lib/subagent.ts +596 -0
- package/extensions/dev-team/lib/terminal-text.ts +54 -0
- package/extensions/dev-team/lib/tools-misc.ts +152 -0
- package/extensions/dev-team/lib/transcript.ts +110 -0
- package/extensions/dev-team/lib/trust.ts +52 -0
- package/extensions/dev-team/lib/usage-breakdown.ts +176 -0
- package/extensions/dev-team/lib/usage-chart.ts +153 -0
- package/extensions/dev-team/lib/usage-command.ts +107 -0
- package/extensions/dev-team/lib/usage-history.ts +203 -0
- package/extensions/dev-team/lib/usage-render.ts +225 -0
- package/extensions/dev-team/lib/usage-split-bar.ts +127 -0
- package/extensions/dev-team/lib/usage-state.ts +116 -0
- package/extensions/dev-team/lib/usage-text.ts +159 -0
- package/extensions/dev-team/lib/usage-view.ts +109 -0
- package/hooks/__pycache__/refactor_test_freeze_guard.cpython-314.pyc +0 -0
- package/hooks/agent_dispatch_ledger.py +190 -0
- package/hooks/autocompact_setup_nudge.py +99 -0
- package/hooks/bash_retry_guard.py +228 -0
- package/hooks/boundary_events_write_guard.py +352 -0
- package/hooks/code_intelligence_nudge.py +293 -0
- package/hooks/code_intelligence_turn_mark.py +317 -0
- package/hooks/codegraph_bootstrap.py +139 -0
- package/hooks/contract_version_guard.py +362 -0
- package/hooks/cost_meter.py +106 -0
- package/hooks/destructive-commands.json +62 -0
- package/hooks/destructive_guard.py +477 -0
- package/hooks/eval_compliance_check.py +440 -0
- package/hooks/guards.json +17 -0
- package/hooks/hooks.json +323 -0
- package/hooks/internal_double_gate.py +296 -0
- package/hooks/js_fp_review.py +212 -0
- package/hooks/knowledge_index.py +119 -0
- package/hooks/lib/__pycache__/artifact_paths.cpython-314.pyc +0 -0
- package/hooks/lib/__pycache__/atomic_state.cpython-314.pyc +0 -0
- package/hooks/lib/__pycache__/autocompact_config.cpython-314.pyc +0 -0
- package/hooks/lib/__pycache__/boundary_events.cpython-314.pyc +0 -0
- package/hooks/lib/__pycache__/doc_classification.cpython-314.pyc +0 -0
- package/hooks/lib/__pycache__/gh_pr_create_detect.cpython-314.pyc +0 -0
- package/hooks/lib/__pycache__/git_safe_diff.cpython-314.pyc +0 -0
- package/hooks/lib/__pycache__/instrument_log.cpython-314.pyc +0 -0
- package/hooks/lib/__pycache__/metrics_query.cpython-314.pyc +0 -0
- package/hooks/lib/__pycache__/plugin_version.cpython-314.pyc +0 -0
- package/hooks/lib/__pycache__/pre_commit_doc_classifier.cpython-314.pyc +0 -0
- package/hooks/lib/__pycache__/review_agent_registry.cpython-314.pyc +0 -0
- package/hooks/lib/__pycache__/review_gate_corroboration.cpython-314.pyc +0 -0
- package/hooks/lib/__pycache__/review_gate_hash.cpython-314.pyc +0 -0
- package/hooks/lib/__pycache__/review_verdicts.cpython-314.pyc +0 -0
- package/hooks/lib/__pycache__/stdin_json.cpython-314.pyc +0 -0
- package/hooks/lib/__pycache__/stryker_invocation.cpython-314.pyc +0 -0
- package/hooks/lib/__pycache__/telemetry_consent.cpython-314.pyc +0 -0
- package/hooks/lib/__pycache__/test_file_classify.cpython-314.pyc +0 -0
- package/hooks/lib/__pycache__/token_efficiency_limits.cpython-314.pyc +0 -0
- package/hooks/lib/__pycache__/verify_guard_state.cpython-314.pyc +0 -0
- package/hooks/lib/__pycache__/xunit_v3_operator_gate.cpython-314.pyc +0 -0
- package/hooks/lib/agent_skill_hints.py +74 -0
- package/hooks/lib/artifact_paths.py +263 -0
- package/hooks/lib/atomic_state.py +557 -0
- package/hooks/lib/autocompact_config.py +103 -0
- package/hooks/lib/autoship_log.py +106 -0
- package/hooks/lib/banned_scripts_policy.py +51 -0
- package/hooks/lib/boundary_events.py +436 -0
- package/hooks/lib/build_knowledge_index.py +504 -0
- package/hooks/lib/build_skills_index.py +361 -0
- package/hooks/lib/build_state.py +116 -0
- package/hooks/lib/classify_ship_outcome.py +126 -0
- package/hooks/lib/config_changelog_schema.py +115 -0
- package/hooks/lib/cost_meter.py +955 -0
- package/hooks/lib/doc_classification.py +116 -0
- package/hooks/lib/gh_pr_create_detect.py +136 -0
- package/hooks/lib/git_safe_diff.py +123 -0
- package/hooks/lib/instrument_log.py +66 -0
- package/hooks/lib/iteration_journal_gate.py +197 -0
- package/hooks/lib/knowledge_index_paths.py +88 -0
- package/hooks/lib/mcp_json_repowise.py +177 -0
- package/hooks/lib/metrics_query.py +202 -0
- package/hooks/lib/minimal_yaml.py +434 -0
- package/hooks/lib/plugin_version.py +142 -0
- package/hooks/lib/pre_commit_detect.py +537 -0
- package/hooks/lib/pre_commit_doc_classifier.py +126 -0
- package/hooks/lib/pricing.py +118 -0
- package/hooks/lib/report_pdf.py +371 -0
- package/hooks/lib/review_agent_registry.py +142 -0
- package/hooks/lib/review_dispatch_ledger.py +101 -0
- package/hooks/lib/review_gate_corroboration.py +521 -0
- package/hooks/lib/review_gate_hash.py +252 -0
- package/hooks/lib/review_gate_normalized_hash.py +1115 -0
- package/hooks/lib/review_verdicts.py +301 -0
- package/hooks/lib/run_report.py +160 -0
- package/hooks/lib/skill_categories.yaml +125 -0
- package/hooks/lib/stdin_json.py +57 -0
- package/hooks/lib/stryker_invocation.py +102 -0
- package/hooks/lib/telemetry_consent.py +41 -0
- package/hooks/lib/telemetry_report.py +108 -0
- package/hooks/lib/test_file_classify.py +160 -0
- package/hooks/lib/token_efficiency_limits.py +51 -0
- package/hooks/lib/turn_identity.py +77 -0
- package/hooks/lib/verify_guard_state.py +110 -0
- package/hooks/lib/workflow_state.py +206 -0
- package/hooks/lib/xunit_v3_operator_gate.py +596 -0
- package/hooks/mcp_json_repowise_nudge.py +74 -0
- package/hooks/mutation_adapters/__init__.py +7 -0
- package/hooks/mutation_adapters/__pycache__/__init__.cpython-314.pyc +0 -0
- package/hooks/mutation_adapters/__pycache__/lib.cpython-314.pyc +0 -0
- package/hooks/mutation_adapters/__pycache__/mutmut.cpython-314.pyc +0 -0
- package/hooks/mutation_adapters/__pycache__/pitest.cpython-314.pyc +0 -0
- package/hooks/mutation_adapters/__pycache__/stryker.cpython-314.pyc +0 -0
- package/hooks/mutation_adapters/__pycache__/stryker_net.cpython-314.pyc +0 -0
- package/hooks/mutation_adapters/lib.py +478 -0
- package/hooks/mutation_adapters/mutmut.py +188 -0
- package/hooks/mutation_adapters/pitest.py +266 -0
- package/hooks/mutation_adapters/stryker.py +157 -0
- package/hooks/mutation_adapters/stryker_net.py +264 -0
- package/hooks/mutation_gate.py +193 -0
- package/hooks/mutation_testing_smoke_gate.py +371 -0
- package/hooks/pending_review_notify.py +121 -0
- package/hooks/phase_marker.py +138 -0
- package/hooks/post_compact_state_reinject.py +180 -0
- package/hooks/post_format.py +115 -0
- package/hooks/pre_commit_knowledge_index.py +128 -0
- package/hooks/pre_commit_review.py +66 -0
- package/hooks/pre_pr_review.py +694 -0
- package/hooks/pre_tool_guard.py +405 -0
- package/hooks/py.sh +73 -0
- package/hooks/refactor-bash-write-patterns.json +29 -0
- package/hooks/refactor_test_bash_guard.py +253 -0
- package/hooks/refactor_test_freeze_guard.py +139 -0
- package/hooks/refactor_test_revert_guard.py +186 -0
- package/hooks/repo_review_nudge.py +287 -0
- package/hooks/review_verdict_recorder.py +464 -0
- package/hooks/scan_bash_command_for_banned_scripts.py +428 -0
- package/hooks/scan_worktree_for_banned_scripts.py +238 -0
- package/hooks/session_learning_trigger.py +248 -0
- package/hooks/skills_index.py +126 -0
- package/hooks/stryker_xunit_shim_guard.py +571 -0
- package/hooks/subagent_completion_guard.py +309 -0
- package/hooks/subagent_skill_context.py +139 -0
- package/hooks/task_completion_metrics.py +216 -0
- package/hooks/tdd_guard.py +229 -0
- package/hooks/telemetry.py +341 -0
- package/hooks/token_efficiency_review.py +194 -0
- package/hooks/verify_guard.py +183 -0
- package/hooks/verify_guard_edit_marker.py +73 -0
- package/hooks/version_check.py +173 -0
- package/knowledge/accepted-risks-schema.md +98 -0
- package/knowledge/adr-decision-criteria.md +64 -0
- package/knowledge/adversarial-review-protocol.md +139 -0
- package/knowledge/agent-registry.md +228 -0
- package/knowledge/agent-review-methodology.md +80 -0
- package/knowledge/ai-friendly-repo-guidelines.md +67 -0
- package/knowledge/architecture-assessment.md +96 -0
- package/knowledge/artifact-lifecycle.md +57 -0
- package/knowledge/cd-maturity-model.md +82 -0
- package/knowledge/cd-test-architecture.md +190 -0
- package/knowledge/ci-cd-file-scope.md +24 -0
- package/knowledge/codegraph-vs-graphify.md +192 -0
- package/knowledge/component-test-patterns.md +139 -0
- package/knowledge/database-change-management.md +80 -0
- package/knowledge/database-test-patterns.md +79 -0
- package/knowledge/decision-defaults.md +88 -0
- package/knowledge/dependency-breaking-techniques.md +116 -0
- package/knowledge/deployment-pipeline.md +86 -0
- package/knowledge/design-smells.md +122 -0
- package/knowledge/directory-enumeration.md +38 -0
- package/knowledge/domain-modeling.md +123 -0
- package/knowledge/evidence-bundle.md +90 -0
- package/knowledge/exploratory-testing-field-guide.md +122 -0
- package/knowledge/failure-routing.md +28 -0
- package/knowledge/fixture-construction.md +56 -0
- package/knowledge/frontend-component-architecture.md +139 -0
- package/knowledge/gherkin-quality-review-dispatch.md +135 -0
- package/knowledge/index.json +6766 -0
- package/knowledge/internal-collaborator-doubling.md +101 -0
- package/knowledge/legacy-test-strategy.md +71 -0
- package/knowledge/long-run-waiting.md +66 -0
- package/knowledge/microservice-testing.md +71 -0
- package/knowledge/model-pricing.json +23 -0
- package/knowledge/mutation-score-formulas.md +60 -0
- package/knowledge/object-calisthenics.md +147 -0
- package/knowledge/oracle-provenance.md +94 -0
- package/knowledge/orchestrator-script-implementation.md +185 -0
- package/knowledge/owasp-detection.md +148 -0
- package/knowledge/plan-review-rubric.md +56 -0
- package/knowledge/proxy-connectivity.md +62 -0
- package/knowledge/reactive-effect-patterns.md +73 -0
- package/knowledge/recon-inventory-excludes.txt +32 -0
- package/knowledge/references/bdd-value-guide.md +61 -0
- package/knowledge/references/csharp-http-client-testing.md +264 -0
- package/knowledge/release-strategies.md +74 -0
- package/knowledge/report-output-location.md +117 -0
- package/knowledge/report-pdf-integration.md +63 -0
- package/knowledge/report-print.css +129 -0
- package/knowledge/report-template.md +114 -0
- package/knowledge/report-to-pdf.md +69 -0
- package/knowledge/request-processing-flow.md +63 -0
- package/knowledge/result-verification.md +52 -0
- package/knowledge/review-agent-output-contract.md +121 -0
- package/knowledge/review-lens-classification.md +113 -0
- package/knowledge/review-rubric.md +62 -0
- package/knowledge/review-template.md +104 -0
- package/knowledge/rule-fixtures/A02.insecure-random-js/negative.js +1 -0
- package/knowledge/rule-fixtures/A02.insecure-random-js/positive.js +1 -0
- package/knowledge/rule-fixtures/A02.weak-hashing-md5/negative.py +1 -0
- package/knowledge/rule-fixtures/A02.weak-hashing-md5/positive.py +1 -0
- package/knowledge/rule-fixtures/A03.command-injection/negative.js +1 -0
- package/knowledge/rule-fixtures/A03.command-injection/positive.js +1 -0
- package/knowledge/rule-fixtures/A03.sql-injection/negative.js +1 -0
- package/knowledge/rule-fixtures/A03.sql-injection/positive.js +1 -0
- package/knowledge/rule-fixtures/A03.xss-innerhtml/negative.js +1 -0
- package/knowledge/rule-fixtures/A03.xss-innerhtml/positive.js +1 -0
- package/knowledge/rule-fixtures/A05.cors-wildcard/negative.js +1 -0
- package/knowledge/rule-fixtures/A05.cors-wildcard/positive.js +1 -0
- package/knowledge/rule-fixtures/A05.default-credentials/negative.js +1 -0
- package/knowledge/rule-fixtures/A05.default-credentials/positive.js +1 -0
- package/knowledge/rule-fixtures/A07.jwt-alg-none/negative.js +1 -0
- package/knowledge/rule-fixtures/A07.jwt-alg-none/positive.js +1 -0
- package/knowledge/rule-fixtures/A08.binary-formatter/negative.cs +1 -0
- package/knowledge/rule-fixtures/A08.binary-formatter/positive.cs +1 -0
- package/knowledge/rule-fixtures/A08.js-eval/negative.js +1 -0
- package/knowledge/rule-fixtures/A08.js-eval/positive.js +1 -0
- package/knowledge/rule-fixtures/A08.object-input-stream/negative.java +1 -0
- package/knowledge/rule-fixtures/A08.object-input-stream/positive.java +1 -0
- package/knowledge/schemas/disposition-register-v1.json +65 -0
- package/knowledge/schemas/recon-envelope-v1.json +198 -0
- package/knowledge/schemas/unified-finding-v1.json +72 -0
- package/knowledge/security-primitives-contract.md +301 -0
- package/knowledge/security-review-rule-map.yaml +107 -0
- package/knowledge/skills-registry.md +72 -0
- package/knowledge/task-size-classifier.md +103 -0
- package/knowledge/telemetry-schema.md +881 -0
- package/knowledge/test-automation-maturity.md +56 -0
- package/knowledge/test-automation-principles.md +71 -0
- package/knowledge/test-cadence-tradeoffs.md +68 -0
- package/knowledge/test-doubles.md +105 -0
- package/knowledge/test-file-indicators.md +22 -0
- package/knowledge/test-layer-gates.md +35 -0
- package/knowledge/test-matrix-examples/django-batch.md +24 -0
- package/knowledge/test-matrix-examples/dotnet-grpc-fronting-api.md +90 -0
- package/knowledge/test-matrix-examples/dotnet-http-consumer.md +131 -0
- package/knowledge/test-matrix-examples/react-node-spa.md +24 -0
- package/knowledge/test-matrix-examples/spring-boot-service.md +25 -0
- package/knowledge/test-matrix-examples/ssr-htmx.md +24 -0
- package/knowledge/test-organization.md +70 -0
- package/knowledge/test-pyramid.md +84 -0
- package/knowledge/test-refactoring.md +67 -0
- package/knowledge/test-review-division-of-labor.md +85 -0
- package/knowledge/test-smells.md +80 -0
- package/knowledge/test-stack-profiles/bdd-frameworks.md +235 -0
- package/knowledge/test-stack-profiles/django.md +13 -0
- package/knowledge/test-stack-profiles/dotnet.md +18 -0
- package/knowledge/test-stack-profiles/go.md +16 -0
- package/knowledge/test-stack-profiles/node.md +16 -0
- package/knowledge/test-stack-profiles/react.md +12 -0
- package/knowledge/test-stack-profiles/spring-boot.md +16 -0
- package/knowledge/test-stack-profiles/ssr-htmx.md +14 -0
- package/knowledge/test-stack-profiles/vue.md +12 -0
- package/knowledge/test-strategy.md +70 -0
- package/knowledge/testability-patterns.md +240 -0
- package/knowledge/testing-quadrants.md +44 -0
- package/knowledge/testing-techniques/approval.md +15 -0
- package/knowledge/testing-techniques/chaos.md +17 -0
- package/knowledge/testing-techniques/fuzz.md +15 -0
- package/knowledge/testing-techniques/property-based.md +15 -0
- package/knowledge/testing-techniques/schema-validation.md +15 -0
- package/knowledge/testing-techniques/screenshot.md +15 -0
- package/knowledge/three-phase-workflow.md +198 -0
- package/knowledge/value-patterns.md +55 -0
- package/knowledge/verification-mode.md +116 -0
- package/knowledge/virtual-service-libraries.md +75 -0
- package/knowledge/wave-consolidation-guidance.md +21 -0
- package/overrides/agents/Explore.md +15 -0
- package/overrides/agents/general-purpose.md +10 -0
- package/overrides/notes/autoship.md +6 -0
- package/overrides/notes/issues-from-assessment.md +3 -0
- package/overrides/notes/issues-from-plan.md +3 -0
- package/overrides/notes/mutation-night-watch.md +3 -0
- package/overrides/notes/mutation-testing.md +3 -0
- package/overrides/notes/pr.md +7 -0
- package/overrides/notes/project-init.md +6 -0
- package/overrides/notes/setup.md +13 -0
- package/overrides/notes/specs.md +3 -0
- package/overrides/skills/headless-run/SKILL.md +45 -0
- package/overrides/skills/upgrade/SKILL.md +30 -0
- package/overrides/skills/version/SKILL.md +25 -0
- package/package.json +36 -0
- package/scripts/authoring_digest.py +93 -0
- package/scripts/autoship_discover.py +121 -0
- package/scripts/autoship_group.py +409 -0
- package/scripts/autoship_proposals.py +494 -0
- package/scripts/autoship_queue.py +291 -0
- package/scripts/autoship_reclaim.py +495 -0
- package/scripts/build_jobs.py +108 -0
- package/scripts/build_rollback_point.py +240 -0
- package/scripts/build_slice_scope.py +157 -0
- package/scripts/build_wave.py +109 -0
- package/scripts/build_wave_reconcile.py +252 -0
- package/scripts/build_worktree_baseref.py +113 -0
- package/scripts/check_agent_scope.py +117 -0
- package/scripts/check_agent_tool_mapping.py +213 -0
- package/scripts/check_review_agent_mcp_tools.py +317 -0
- package/scripts/check_security_assessment_mcp_tools.py +165 -0
- package/scripts/checkpoint_abort.py +502 -0
- package/scripts/claude_setup_review.py +438 -0
- package/scripts/codebase_recon.py +556 -0
- package/scripts/coverage_config.py +623 -0
- package/scripts/coverage_delta_steering.py +330 -0
- package/scripts/coverage_discovery_dotnet.py +315 -0
- package/scripts/coverage_discovery_java.py +742 -0
- package/scripts/coverage_discovery_js.py +546 -0
- package/scripts/coverage_gap_ranking.py +556 -0
- package/scripts/coverage_readiness.py +455 -0
- package/scripts/coverage_report_parse.py +521 -0
- package/scripts/detect_bdd_convention.py +252 -0
- package/scripts/eval_ablation.py +376 -0
- package/scripts/gherkin_analysis_coverage_gate.py +306 -0
- package/scripts/gherkin_cross_feature_duplicate_titles_gate.py +173 -0
- package/scripts/gherkin_effectiveness_rollup.py +238 -0
- package/scripts/gherkin_failure_path_gate.py +206 -0
- package/scripts/gherkin_feature_merge.py +720 -0
- package/scripts/gherkin_stub_gate.py +163 -0
- package/scripts/gherkin_stub_merge.py +479 -0
- package/scripts/git_origin_host.py +88 -0
- package/scripts/install-java-static-analysis.py +110 -0
- package/scripts/issue_deps.py +74 -0
- package/scripts/lib/_bdd_markers.py +28 -0
- package/scripts/lib/_gherkin_text.py +93 -0
- package/scripts/lib/_vendored_tree.py +70 -0
- package/scripts/lib/autoship_state.py +397 -0
- package/scripts/lib/claude_md_guard.py +226 -0
- package/scripts/lib/deterministic_recon.py +446 -0
- package/scripts/lib/mcp_tool_grants.py +211 -0
- package/scripts/lib/plan_parse.py +386 -0
- package/scripts/lib/review_result.py +84 -0
- package/scripts/lib/review_roster.py +86 -0
- package/scripts/lib/session_log/__init__.py +34 -0
- package/scripts/lib/session_log/__pycache__/__init__.cpython-314.pyc +0 -0
- package/scripts/lib/session_log/__pycache__/records.cpython-314.pyc +0 -0
- package/scripts/lib/session_log/classify.py +231 -0
- package/scripts/lib/session_log/corrections.py +194 -0
- package/scripts/lib/session_log/discovery.py +108 -0
- package/scripts/lib/session_log/records.py +218 -0
- package/scripts/lib/session_log/redact.py +76 -0
- package/scripts/lib/session_log/signals.py +373 -0
- package/scripts/lib/session_report_downstream.py +614 -0
- package/scripts/lib/session_report_maintainer.py +1273 -0
- package/scripts/lib/session_report_shared.py +262 -0
- package/scripts/lib/settings_hook_guard.py +157 -0
- package/scripts/lib/slug.py +33 -0
- package/scripts/lib/stub_extractors/__init__.py +82 -0
- package/scripts/lib/stub_extractors/_common.py +328 -0
- package/scripts/lib/stub_extractors/csharp.py +19 -0
- package/scripts/lib/stub_extractors/go.py +173 -0
- package/scripts/lib/stub_extractors/java.py +18 -0
- package/scripts/lib/stub_extractors/jsts.py +126 -0
- package/scripts/mutation_stack_sections.py +149 -0
- package/scripts/mutation_yield_steering.py +345 -0
- package/scripts/orchestrator.py +895 -0
- package/scripts/plan_gherkin_export.py +227 -0
- package/scripts/plan_waves.py +208 -0
- package/scripts/pr_close_keyword_lint.py +108 -0
- package/scripts/progress_guardian.py +888 -0
- package/scripts/recon_inventory.py +273 -0
- package/scripts/review_findings_log.py +93 -0
- package/scripts/run_invariants.py +124 -0
- package/scripts/select_lenses.py +640 -0
- package/scripts/session_report.py +486 -0
- package/scripts/set_autocompact_env.py +221 -0
- package/scripts/ship_resume_guard.py +135 -0
- package/scripts/ship_review_gate.py +63 -0
- package/scripts/specs_convention_marker.py +103 -0
- package/scripts/test_improve_resume.py +277 -0
- package/scripts/test_review_mechanics.py +958 -0
- package/scripts/token_efficiency_review.py +322 -0
- package/scripts/verdict_scope.py +285 -0
- package/scripts/verify_gherkin_quality_critic_isolation.py +296 -0
- package/scripts/verify_tier.py +157 -0
- package/skills/adr-tools/SKILL.md +118 -0
- package/skills/agent-readiness/SKILL.md +105 -0
- package/skills/agent-readiness/ai_friendly_analyzers.py +326 -0
- package/skills/agent-readiness/scanner.py +441 -0
- package/skills/agent-readiness/scorecard.yaml +88 -0
- package/skills/api-design/SKILL.md +115 -0
- package/skills/apply-fixes/SKILL.md +171 -0
- package/skills/apply-test-doubles/SKILL.md +321 -0
- package/skills/artifact-lifecycle/SKILL.md +127 -0
- package/skills/autoship/SKILL.md +1124 -0
- package/skills/benchmark/SKILL.md +105 -0
- package/skills/branch-workflow/SKILL.md +89 -0
- package/skills/browse/SKILL.md +184 -0
- package/skills/browser-testing/SKILL.md +62 -0
- package/skills/browser-testing/references/playwright-patterns.md +216 -0
- package/skills/build/SKILL.md +422 -0
- package/skills/build/references/static-self-heal.md +245 -0
- package/skills/careful/SKILL.md +72 -0
- package/skills/cd-test-architecture/SKILL.md +371 -0
- package/skills/ci-debugging/SKILL.md +105 -0
- package/skills/co-evolution-audit/SKILL.md +269 -0
- package/skills/code-review/SKILL.md +1015 -0
- package/skills/code-review/examples/aggregated-sample.json +56 -0
- package/skills/code-review/examples/sample-report.md +41 -0
- package/skills/code-review/output-format.md +478 -0
- package/skills/code-review/scripts/activation.py +86 -0
- package/skills/code-review/scripts/change_impact.py +357 -0
- package/skills/code-review/scripts/change_shape.py +372 -0
- package/skills/code-review/scripts/change_size.py +212 -0
- package/skills/code-review/scripts/changed_file_list.py +141 -0
- package/skills/code-review/scripts/closing_pass.py +187 -0
- package/skills/code-review/scripts/consolidate.py +277 -0
- package/skills/code-review/scripts/contract_failure_report.py +185 -0
- package/skills/code-review/scripts/dispatch_reconcile.py +66 -0
- package/skills/code-review/scripts/dispatch_waves.py +164 -0
- package/skills/code-review/scripts/finding_signature.py +446 -0
- package/skills/code-review/scripts/ledger.py +283 -0
- package/skills/code-review/scripts/partition.py +169 -0
- package/skills/code-review/scripts/render_tiered_findings.py +274 -0
- package/skills/code-review/scripts/repo_invariants.py +1066 -0
- package/skills/code-review/scripts/review_context_pack.py +306 -0
- package/skills/code-review/scripts/review_round_log.py +345 -0
- package/skills/code-review/scripts/review_value_coverage.py +297 -0
- package/skills/code-review/scripts/validate_review_output.py +467 -0
- package/skills/code-review/sliced-mode.md +205 -0
- package/skills/competitive-analysis/SKILL.md +191 -0
- package/skills/context-loading-protocol/SKILL.md +157 -0
- package/skills/continue/SKILL.md +90 -0
- package/skills/cost-report/SKILL.md +178 -0
- package/skills/coverage-baseline/SKILL.md +335 -0
- package/skills/coverage-baseline/references/multi-project-discovery.md +202 -0
- package/skills/coverage-delta/SKILL.md +181 -0
- package/skills/coverage-delta/references/mutation-gate.md +70 -0
- package/skills/design-doc/SKILL.md +95 -0
- package/skills/design-interrogation/SKILL.md +89 -0
- package/skills/design-it-twice/SKILL.md +91 -0
- package/skills/docker-image-audit/SKILL.md +108 -0
- package/skills/docker-image-audit/references/install-guide.md +64 -0
- package/skills/docker-image-audit/references/report-template.md +73 -0
- package/skills/docker-image-create/SKILL.md +185 -0
- package/skills/domain-analysis/SKILL.md +183 -0
- package/skills/domain-driven-design/SKILL.md +194 -0
- package/skills/exploratory-testing/SKILL.md +108 -0
- package/skills/explore/SKILL.md +51 -0
- package/skills/farley-score/SKILL.md +165 -0
- package/skills/feature-file-validation/SKILL.md +78 -0
- package/skills/feature-file-validation/references/validation-rules.md +115 -0
- package/skills/feedback-learning/SKILL.md +414 -0
- package/skills/fix/SKILL.md +450 -0
- package/skills/freeze/SKILL.md +68 -0
- package/skills/frontend-architecture/SKILL.md +113 -0
- package/skills/gherkin-derive/SKILL.md +630 -0
- package/skills/gherkin-public/SKILL.md +266 -0
- package/skills/governance-compliance/SKILL.md +150 -0
- package/skills/guard/SKILL.md +75 -0
- package/skills/handoff/SKILL.md +139 -0
- package/skills/handoff/references/summary-templates.md +242 -0
- package/skills/harness-audit/SKILL.md +751 -0
- package/skills/harness-audit/scripts/lesson_validate.py +386 -0
- package/skills/harness-audit/scripts/redundancy_criterion.py +188 -0
- package/skills/headless-run/SKILL.md +45 -0
- package/skills/headless-run/scripts/isolated_dispatch.py +381 -0
- package/skills/help/SKILL.md +72 -0
- package/skills/hexagonal-architecture/SKILL.md +85 -0
- package/skills/human-oversight-protocol/SKILL.md +224 -0
- package/skills/issues-from-assessment/SKILL.md +223 -0
- package/skills/issues-from-plan/SKILL.md +133 -0
- package/skills/legacy-code/SKILL.md +132 -0
- package/skills/mermaid-diagramming/SKILL.md +120 -0
- package/skills/mutation-night-watch/SKILL.md +154 -0
- package/skills/mutation-night-watch/references/scheduling.md +135 -0
- package/skills/mutation-testing/SKILL.md +396 -0
- package/skills/mutation-testing/references/languages/csharp-stryker-net.md +676 -0
- package/skills/mutation-testing/references/languages/go-go-mutesting.md +95 -0
- package/skills/mutation-testing/references/languages/java-pitest.md +77 -0
- package/skills/mutation-testing/references/languages/javascript-stryker.md +188 -0
- package/skills/mutation-testing/references/languages/python-mutmut.md +97 -0
- package/skills/mutation-testing/references/time-estimation.md +34 -0
- package/skills/mutation-testing/references/tool-detection.md +15 -0
- package/skills/mutation-testing/references/workflow-callers.md +23 -0
- package/skills/mutation-testing/scripts/__pycache__/xunit_v3_feature_detector.cpython-314.pyc +0 -0
- package/skills/mutation-testing/scripts/csharp_stryker_net_slice_runner.py +635 -0
- package/skills/mutation-testing/scripts/csharp_stryker_net_status_loop.py +525 -0
- package/skills/mutation-testing/scripts/csharp_stryker_net_wrapper.py +681 -0
- package/skills/mutation-testing/scripts/mutation_baseline_reuse.py +292 -0
- package/skills/mutation-testing/scripts/mutation_exclude_policy.py +268 -0
- package/skills/mutation-testing/scripts/mutation_feasibility_gate.py +463 -0
- package/skills/mutation-testing/scripts/mutation_kill_headless.py +331 -0
- package/skills/mutation-testing/scripts/mutation_kill_insert.py +199 -0
- package/skills/mutation-testing/scripts/mutation_kill_insert_python.py +150 -0
- package/skills/mutation-testing/scripts/mutation_kill_loop.py +869 -0
- package/skills/mutation-testing/scripts/mutation_kill_loop_python.py +949 -0
- package/skills/mutation-testing/scripts/mutation_kill_retry.py +592 -0
- package/skills/mutation-testing/scripts/mutation_kill_shared.py +620 -0
- package/skills/mutation-testing/scripts/mutation_nightwatch.py +462 -0
- package/skills/mutation-testing/scripts/mutation_nightwatch_stacks.py +425 -0
- package/skills/mutation-testing/scripts/mutation_report.py +743 -0
- package/skills/mutation-testing/scripts/mutation_report_cli.py +175 -0
- package/skills/mutation-testing/scripts/mutation_safety_gate.py +69 -0
- package/skills/mutation-testing/scripts/stryker_shard_pipeline.py +847 -0
- package/skills/mutation-testing/scripts/stryker_shard_setup.py +440 -0
- package/skills/mutation-testing/scripts/stryker_timeout_retry.py +142 -0
- package/skills/mutation-testing/scripts/xunit_v3_feature_detector.py +341 -0
- package/skills/performance-benchmark/SKILL.md +174 -0
- package/skills/performance-benchmark/examples/report-format.md +43 -0
- package/skills/performance-benchmark/references/benchmark-script.md +169 -0
- package/skills/performance-metrics/SKILL.md +265 -0
- package/skills/plan/SKILL.md +199 -0
- package/skills/plan/references/gherkin-persistence.md +43 -0
- package/skills/plan/references/plan-template.md +182 -0
- package/skills/pr/SKILL.md +289 -0
- package/skills/pr/scripts/gate_retry_state.py +368 -0
- package/skills/project-init/README.md +141 -0
- package/skills/project-init/SKILL.md +1197 -0
- package/skills/project-init/evals/evals.json +200 -0
- package/skills/project-init/references/capability-tools.md +55 -0
- package/skills/project-init/references/configs.md +221 -0
- package/skills/property-based-testing/SKILL.md +121 -0
- package/skills/property-based-testing/fixtures/invariant_fixture.py +15 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/README.md +42 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/fast-check/LICENSE +21 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/fast-check/README.md +263 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/fast-check/lib/cjs/fast-check.d.ts +5165 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/fast-check/lib/cjs/fast-check.js +12147 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/fast-check/lib/cjs/package.json +3 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/fast-check/lib/cjs/types57/fast-check.d.ts +5165 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/fast-check/lib/fast-check.d.ts +5165 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/fast-check/lib/fast-check.js +12011 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/fast-check/lib/rolldown-runtime-D7D4PA-g.js +13 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/fast-check/lib/types57/fast-check.d.ts +5165 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/fast-check/package.json +94 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/LICENSE +21 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/README.md +168 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/RandomGenerator-DcXj09Ch.d.ts +14 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/distribution/uniformBigInt.d.ts +15 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/distribution/uniformBigInt.js +38 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/distribution/uniformFloat32.d.ts +15 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/distribution/uniformFloat32.js +18 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/distribution/uniformFloat64.d.ts +15 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/distribution/uniformFloat64.js +22 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/distribution/uniformInt.d.ts +15 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/distribution/uniformInt.js +134 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/RandomGenerator-DcXj09Ch.d.ts +14 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/distribution/uniformBigInt.d.ts +15 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/distribution/uniformBigInt.js +37 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/distribution/uniformFloat32.d.ts +15 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/distribution/uniformFloat32.js +17 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/distribution/uniformFloat64.d.ts +15 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/distribution/uniformFloat64.js +21 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/distribution/uniformInt.d.ts +15 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/distribution/uniformInt.js +133 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/generator/congruential32.d.ts +7 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/generator/congruential32.js +44 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/generator/mersenne.d.ts +7 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/generator/mersenne.js +90 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/generator/xoroshiro128plus.d.ts +7 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/generator/xoroshiro128plus.js +80 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/generator/xorshift128plus.d.ts +7 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/generator/xorshift128plus.js +78 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/package.json +3 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/types/JumpableRandomGenerator.d.ts +16 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/types/JumpableRandomGenerator.js +0 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/types/RandomGenerator.d.ts +2 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/types/RandomGenerator.js +0 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/utils/generateN.d.ts +6 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/utils/generateN.js +8 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/utils/purify.d.ts +12 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/utils/purify.js +9 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/utils/skipN.d.ts +6 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/utils/skipN.js +6 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/generator/congruential32.d.ts +7 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/generator/congruential32.js +46 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/generator/mersenne.d.ts +7 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/generator/mersenne.js +92 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/generator/xoroshiro128plus.d.ts +7 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/generator/xoroshiro128plus.js +82 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/generator/xorshift128plus.d.ts +7 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/generator/xorshift128plus.js +80 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/types/JumpableRandomGenerator.d.ts +16 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/types/JumpableRandomGenerator.js +0 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/types/RandomGenerator.d.ts +2 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/types/RandomGenerator.js +0 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/utils/generateN.d.ts +6 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/utils/generateN.js +9 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/utils/purify.d.ts +12 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/utils/purify.js +10 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/utils/skipN.d.ts +6 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/utils/skipN.js +7 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/package.json +133 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/package-lock.json +1179 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/package.json +14 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/roundtrip.js +29 -0
- package/skills/property-based-testing/fixtures/js-roundtrip/roundtrip.properties.test.js +16 -0
- package/skills/property-based-testing/fixtures/no_property_fixture.py +10 -0
- package/skills/property-based-testing/fixtures/roundtrip_fixture.py +16 -0
- package/skills/property-based-testing/references/languages/javascript.md +54 -0
- package/skills/property-based-testing/scripts/detect_and_dispatch.py +80 -0
- package/skills/property-based-testing/scripts/hypothesis_scaffold.py +276 -0
- package/skills/proxy-resilience/SKILL.md +84 -0
- package/skills/quality-gate-pipeline/SKILL.md +184 -0
- package/skills/quality-targets-converge/SKILL.md +254 -0
- package/skills/repo-review/SKILL.md +159 -0
- package/skills/report-pdf/SKILL.md +66 -0
- package/skills/review/SKILL.md +47 -0
- package/skills/review-agent/SKILL.md +152 -0
- package/skills/review-summary/SKILL.md +73 -0
- package/skills/run-report/SKILL.md +70 -0
- package/skills/semantic-duplication-scan/SKILL.md +337 -0
- package/skills/semantic-scan/SKILL.md +53 -0
- package/skills/semgrep-analyze/SKILL.md +139 -0
- package/skills/setup/SKILL.md +1122 -0
- package/skills/ship/SKILL.md +240 -0
- package/skills/source-verification/SKILL.md +210 -0
- package/skills/source-verification/scripts/claim_extractor.py +155 -0
- package/skills/specs/.size-baseline.json +4 -0
- package/skills/specs/SKILL.md +243 -0
- package/skills/specs/references/completeness-checklist.md +83 -0
- package/skills/specs/references/extraction.md +58 -0
- package/skills/specs/references/glossary.md +59 -0
- package/skills/specs/references/persistence.md +115 -0
- package/skills/specs/references/predictability-check.md +77 -0
- package/skills/static-analysis-integration/SKILL.md +235 -0
- package/skills/static-analysis-integration/adapters/_envelope.py +26 -0
- package/skills/static-analysis-integration/adapters/jscpd-adapter.py +66 -0
- package/skills/static-analysis-integration/adapters/lizard-adapter.py +81 -0
- package/skills/static-analysis-integration/adapters/mypy-adapter.py +50 -0
- package/skills/static-analysis-integration/adapters/mypy-src-layout.py +93 -0
- package/skills/static-analysis-integration/adapters/security-review-adapter.py +212 -0
- package/skills/static-analysis-integration/maintenance.md +23 -0
- package/skills/static-analysis-integration/references/language-setup.md +228 -0
- package/skills/static-analysis-integration/references/sarif-parser.md +124 -0
- package/skills/static-analysis-integration/references/security-review-adapter.md +118 -0
- package/skills/static-analysis-integration/references/tool-configs.md +617 -0
- package/skills/static-analysis-integration/rulesets/pmd-quickstart.xml +24 -0
- package/skills/stryker-xunit-v2-shim/SKILL.md +274 -0
- package/skills/stryker-xunit-v2-shim/references/shim-howto.md +256 -0
- package/skills/stryker-xunit-v2-shim/scripts/generate_shim.py +143 -0
- package/skills/systematic-debugging/SKILL.md +130 -0
- package/skills/telemetry/SKILL.md +75 -0
- package/skills/test-audit-disable/SKILL.md +129 -0
- package/skills/test-design/SKILL.md +177 -0
- package/skills/test-design/scripts/__pycache__/internal_double_detector.cpython-314.pyc +0 -0
- package/skills/test-design/scripts/internal_double_detector.py +631 -0
- package/skills/test-design-advisor/SKILL.md +166 -0
- package/skills/test-driven-development/SKILL.md +169 -0
- package/skills/test-health/SKILL.md +262 -0
- package/skills/test-improve/SKILL.md +239 -0
- package/skills/test-improve/references/phase-0-approach-contract.md +228 -0
- package/skills/test-improve/references/phase-1-analyze.md +131 -0
- package/skills/test-improve/references/phase-2-baseline.md +121 -0
- package/skills/test-improve/references/phase-3-derive-gherkin.md +53 -0
- package/skills/test-improve/references/phase-4-plan-fixes.md +34 -0
- package/skills/test-improve/references/phase-5-improve.md +215 -0
- package/skills/test-improve/references/phase-6-refactor-decision.md +45 -0
- package/skills/test-improve/references/phase-7-refactor.md +44 -0
- package/skills/test-improve/references/phase-8-validate.md +66 -0
- package/skills/test-improve/references/phase-9-close-out-prompt.md +11 -0
- package/skills/test-improve/references/phase-9-report.md +62 -0
- package/skills/test-improve/references/review-loop.md +92 -0
- package/skills/test-improve/templates/executive-summary.md +123 -0
- package/skills/threat-modeling/SKILL.md +108 -0
- package/skills/triage/SKILL.md +211 -0
- package/skills/ubiquitous-language/SKILL.md +192 -0
- package/skills/ubiquitous-language/scripts/collect_domain_signals.py +300 -0
- package/skills/unfreeze/SKILL.md +37 -0
- package/skills/upgrade/SKILL.md +31 -0
- package/skills/upgrade/scripts/check_version_drift.py +113 -0
- package/skills/upgrade/scripts/enable_autoupdate.py +149 -0
- package/skills/version/SKILL.md +25 -0
- package/sync/__pycache__/sync_upstream.cpython-314.pyc +0 -0
- package/sync/sync_upstream.py +293 -0
- package/templates/ACCEPTED-RISKS.md.tmpl +46 -0
- package/templates/agents/agent-template.md +151 -0
- package/templates/agents/angular-testing.md +66 -0
- package/templates/agents/csharp-quality.md +63 -0
- package/templates/agents/esm-enforcer.md +52 -0
- package/templates/agents/front-end-testing.md +65 -0
- package/templates/agents/go-quality.md +65 -0
- package/templates/agents/python-quality.md +62 -0
- package/templates/agents/react-testing.md +61 -0
- package/templates/agents/ts-enforcer.md +60 -0
- package/templates/agents/twelve-factor-audit.md +49 -0
- package/tools/entropy-check.py +250 -0
- package/tools/model-hash-verify.py +213 -0
|
@@ -0,0 +1,881 @@
|
|
|
1
|
+
# Telemetry Schema Reference
|
|
2
|
+
|
|
3
|
+
Every `.claude/metrics/*.jsonl` and `.claude/metrics/*.json` file the dev-team plugin writes,
|
|
4
|
+
in one place, so `session-analysis`, `/session-review`, `/harness-audit`,
|
|
5
|
+
`/cost-report`, and future cross-machine aggregation (#178) compose against
|
|
6
|
+
stable, named schemas instead of reverse-engineering emitters.
|
|
7
|
+
|
|
8
|
+
**Privacy stance (non-negotiable, all streams):** rule IDs, counts, hashes,
|
|
9
|
+
and enums only — never command text, prompt text, file contents, or free-text
|
|
10
|
+
reasons beyond what a stream explicitly documents below as human-authored
|
|
11
|
+
(e.g. `config-changelog.jsonl`'s `description`, which is a deliberate,
|
|
12
|
+
human/agent-reviewed audit note, not incidental free text). Where a stream
|
|
13
|
+
predates this doc and already carries a `reason` field with freeform text
|
|
14
|
+
(e.g. `refactor-freeze.jsonl`'s internal-error diagnostics), that is existing,
|
|
15
|
+
unchanged precedent — not a new exception.
|
|
16
|
+
|
|
17
|
+
Each section below names: fields, types, emitter, consent gating, and
|
|
18
|
+
consumers. A companion test
|
|
19
|
+
(`tests/hooks/test_boundary_events.py::test_schema_doc_covers_all_metrics_paths`)
|
|
20
|
+
cross-checks every `.claude/metrics/*.jsonl` / `.claude/metrics/*.json` path string referenced
|
|
21
|
+
in shipped code against this doc's coverage and fails on omission.
|
|
22
|
+
|
|
23
|
+
---
|
|
24
|
+
|
|
25
|
+
## `boundary-events.jsonl`
|
|
26
|
+
|
|
27
|
+
**Added by #859.** The boundary-level (policy-gateway) channel: every guard
|
|
28
|
+
hook's block/warn/bypass decision, plus human-intervention keywords. Extended
|
|
29
|
+
by #906 with a fifth decision, `revert`, for hooks that don't warn or block
|
|
30
|
+
but actively correct state after the fact. Extended again by #1461 with a
|
|
31
|
+
sixth decision, `record` — a **non-verdict, observational** entry: it does
|
|
32
|
+
not block, warn, bypass, intervene, or revert anything, it merely notes that
|
|
33
|
+
a genuine, registered review-agent dispatch occurred. Emitted by
|
|
34
|
+
`hooks/agent_dispatch_ledger.py` on every `Agent`/`Task` dispatch whose
|
|
35
|
+
`subagent_type` is a real, registered `agents/*-review.md` name (never a
|
|
36
|
+
fabricated/unregistered one — those are never written to the ledger at all).
|
|
37
|
+
This is a **high-frequency** entry (one per genuine review-agent dispatch,
|
|
38
|
+
not a rare guard trip like the other five decisions) — a consumer counting
|
|
39
|
+
"policy decisions" or "guard verdicts" from this stream must explicitly
|
|
40
|
+
exclude `record` rows, or it will badly overcount routine dispatch activity
|
|
41
|
+
as gate verdicts. `hooks/pre_pr_review.py`'s `.pr-review-passed` gate (#1886;
|
|
42
|
+
formerly `hooks/pre_commit_review.py`'s `.review-passed` gate on `git
|
|
43
|
+
commit` — that hook is now a documented no-op) reads this stream (via
|
|
44
|
+
`hooks/lib/review_gate_corroboration.py`) to corroborate that a
|
|
45
|
+
hash-matching gate write was backed by real, independent Agent-tool
|
|
46
|
+
dispatch — see that module's own docstring for its fail-**closed** posture,
|
|
47
|
+
the deliberate opposite of this stream's own fail-open write side.
|
|
48
|
+
Extended again by #1763 with a seventh decision, `dispatch-failure` — also
|
|
49
|
+
**non-verdict, observational**, mirroring `record`'s precedent, but the
|
|
50
|
+
opposite polarity: it notes that a dispatched review agent still failed to
|
|
51
|
+
return a contract-valid result after one retry. Emitted via
|
|
52
|
+
`hooks/lib/boundary_events.py`'s CLI (`--event dispatch-failure --agent
|
|
53
|
+
<name> --subject-hash <hash>`) from `skills/code-review/SKILL.md` Step 4;
|
|
54
|
+
`<name>` is validated against the registered review-agent set at write time
|
|
55
|
+
(with the same plugin-prefix normalization as `record`) — an unregistered
|
|
56
|
+
name is silently not recorded. Unlike `record`, `dispatch-failure` is
|
|
57
|
+
consumed only as NEGATIVE evidence: the gate veto in
|
|
58
|
+
`hooks/lib/review_gate_corroboration.py` / `hooks/pre_pr_review.py` (#1886;
|
|
59
|
+
`_dispatch_failure_verdict`, #1763) treats it as a reason to reject a
|
|
60
|
+
`.pr-review-passed` write, never as corroboration for one. A forged/hand-run
|
|
61
|
+
`dispatch-failure` event can only
|
|
62
|
+
ever cause a false rejection, never a false pass — the opposite forgery
|
|
63
|
+
direction from `record`, which is why this decision (unlike `record`) is
|
|
64
|
+
safely reachable from the CLI's closed `--event` vocabulary.
|
|
65
|
+
|
|
66
|
+
**#2188** adds `SubagentStop` as a `tool` value and `subagent_completion_guard.py`
|
|
67
|
+
as an emitter, reusing the existing `record` decision (not a new one) for the
|
|
68
|
+
same non-verdict, observational reason #1461 introduced it: this hook never
|
|
69
|
+
blocks/warns/bypasses/intervenes/reverts anything — it classifies a
|
|
70
|
+
subagent's own transcript tail (`clean` | `empty-final-turn` |
|
|
71
|
+
`truncated-final-turn` | `unreadable`) and writes one `record` row only for
|
|
72
|
+
the two non-clean, explainable outcomes, with `matched_rule` set to the
|
|
73
|
+
classification itself (`empty-final-turn` or `truncated-final-turn`); the
|
|
74
|
+
common `clean` case and the unexplainable `unreadable` case write nothing.
|
|
75
|
+
|
|
76
|
+
**#2166 Fix #3** adds `review_verdict_recorder.py` as a `SubagentStop`
|
|
77
|
+
emitter, same `record` decision, same non-verdict posture: once this hook
|
|
78
|
+
has confirmed a dispatch's `subagent_type` IS a registered review lens (so
|
|
79
|
+
the dispatch SHOULD produce `review-verdicts.jsonl` rows), a degenerate exit
|
|
80
|
+
that would otherwise be silently indistinguishable from a legitimate no-op
|
|
81
|
+
instead writes one `record` row naming why, via `matched_rule` of
|
|
82
|
+
`missing-scope-marker` (the dispatch prompt's Step 2.1 scope marker is
|
|
83
|
+
missing or reformatted) or `unparseable-result` (the agent's final JSON
|
|
84
|
+
result couldn't be recovered, even by the tolerant extractor). The
|
|
85
|
+
PRE-resolution exits (unreadable transcript, unresolvable `subagent_type`,
|
|
86
|
+
an unregistered or registered-but-non-review `subagent_type`) stay silent —
|
|
87
|
+
those are legitimate no-ops, not degenerate states.
|
|
88
|
+
|
|
89
|
+
| Field | Type | Values / source |
|
|
90
|
+
| --- | --- | --- |
|
|
91
|
+
| `ts` | string | ISO-8601 UTC `%Y-%m-%dT%H:%M:%SZ` |
|
|
92
|
+
| `hook` | string | Emitting hook's module name, e.g. `destructive_guard`, `verify_guard`, `pre_pr_review` (the review-corroboration gate, #1886; `pre_commit_review` is now a documented no-op and emits nothing), `telemetry`, `agent_dispatch_ledger` — or `code-review` for the CLI-emitted events (`--event doc-only`/`single-agent`/`dispatch-failure`), which carry the invoking skill's name rather than a hook module name |
|
|
93
|
+
| `tool` | string | Hooked tool/event: `Bash`, `Write`, `Edit`, `Skill`, `Agent`, `UserPromptSubmit`, `SubagentStop` (#2188) |
|
|
94
|
+
| `decision` | string enum | `block` \| `warn` \| `bypass` \| `intervention` \| `revert` \| `record` \| `dispatch-failure` |
|
|
95
|
+
| `matched_rule` | string | Rule ID from a closed vocabulary (pattern ID, hook-defined constant, bypass flag name, intervention keyword, or — for `record`/`dispatch-failure` — the dispatched review-agent's registered name, or — for `subagent_completion_guard.py`'s `record` rows — `empty-final-turn`/`truncated-final-turn`, #2188, or — for `review_verdict_recorder.py`'s `record` rows — `missing-scope-marker`/`unparseable-result`, #2166 Fix #3, or — for `boundary_events_write_guard.py`'s `block` rows — `ledger-write-blocked`, #2171) — never free text |
|
|
96
|
+
| `plugin_version` | string | From `.claude-plugin/plugin.json` |
|
|
97
|
+
| `session_id` | string, optional | Opaque per-session ID, when present in the hook payload — enables joins with `session-digest.jsonl` |
|
|
98
|
+
| `subject_hash` | string, optional | `review_gate_hash()` value (#1461) binding this event to the staged content it corroborates. A hex digest, not free text |
|
|
99
|
+
| `subject_hash_normalized` | string, optional | `normalized_gate_hash()` value (#1627) — the same binding computed after doc-hunk and indentation normalization. Stamped by `agent_dispatch_ledger.py` alongside `subject_hash`, and read by the gate's cosmetic-delta carry-forward lens. Absent on events written before #1627, which therefore never match on the normalized path. **The digest ALGORITHM has changed twice since** — #1638 (heredoc-body marking plus whole-file `--unified=100000` context) and #1660/#1661/#1662/#1663 (further grammars, a hunk-start guard, and a payload cap), each of which changes the computed VALUE for any changeset touching an affected extension. An event recorded by an earlier plugin version therefore never matches after an upgrade. This fails closed — one lost carry-forward, one extra dispatch — and is not a correctness bug, but it is why a version bump can look like a spurious re-review |
|
|
100
|
+
|
|
101
|
+
**`cosmetic-delta-carry-forward` (#1627, historical).** `pre_commit_review.py`
|
|
102
|
+
used to emit this `bypass`-decision event every time the OLD commit-time gate
|
|
103
|
+
passed a commit whose raw staged hash mismatched but whose normalized hash
|
|
104
|
+
matched. #1886 moved the gate to `gh pr create` and deliberately did NOT
|
|
105
|
+
carry this lens forward — the friction it existed to relieve (a whitespace-only
|
|
106
|
+
re-stage forcing a fresh review-agent dispatch before the NEXT commit) was a
|
|
107
|
+
direct consequence of gating every commit; a gate that fires once, at
|
|
108
|
+
PR-creation time, against the branch's cumulative diff, does not have that
|
|
109
|
+
problem. `hooks/pre_pr_review.py` never emits this event. Existing rows in
|
|
110
|
+
`boundary-events.jsonl` from before the migration remain valid history.
|
|
111
|
+
|
|
112
|
+
- **Emitter:** `hooks/lib/boundary_events.py::emit_boundary_event()`, called from `destructive_guard.py`, `verify_guard.py`, `pre_pr_review.py` (#1886), `telemetry.py` (intervention keywords), `agent_dispatch_ledger.py` (decision `record`, #1461), `subagent_completion_guard.py` (decision `record`, `tool` `SubagentStop`, #2188), `review_verdict_recorder.py` (decision `record`, `tool` `SubagentStop`, `matched_rule` `missing-scope-marker`\|`unparseable-result`, #2166 Fix #3), `boundary_events_write_guard.py` (decision `block`, `tool` `Write`\|`Edit`\|`Bash`, `matched_rule` `ledger-write-blocked` — the PreToolUse guard blocking a direct Write/Edit/Bash write to this same ledger, #2171), the mechanically-adopted guards (`pre_tool_guard.py`, `bash_retry_guard.py`, `refactor_test_freeze_guard.py`, `refactor_test_bash_guard.py`, `refactor_test_revert_guard.py` (decision `revert`, #906), `contract_version_guard.py`, `mutation_testing_smoke_gate.py`, `mutation_gate.py`, `tdd_guard.py`), and `boundary_events.py`'s own CLI (`--event dispatch-failure`, decision `dispatch-failure`, #1763) invoked from `skills/code-review/SKILL.md` Step 4. `--event gate-ran --verdict {allow,block,errored}` (decision `record`, `matched_rule` of `gate-ran-<verdict>`, #2037) is invoked from the repo-root `.husky/pre-commit` git hook — the real, git-native pre-commit gate (distinct from `pre_pr_review.py`, a Claude-Code-level PreToolUse hook gating `gh pr create`) — at every exit point, success or failure alike, so `${CLAUDE_PLUGIN_ROOT}/scripts/session_report.py --profile maintainer` can correlate a commit-attempt Bash record against a nearby `gate_ran` event and classify the previously-unmeasured "the gate silently never ran" population (`gate_ran_absent`) apart from a genuine internal failure (`gate_ran_errored`). This event carries no `session_id` in practice — a real git hook has no Claude Code session_id to attach — so correlation is by time proximity, not session join; see `session_report.py`'s "gate-run correlation (#2037)" section.
|
|
113
|
+
- **Consent:** ALWAYS-ON — not gated by `DEV_TEAM_TELEMETRY`. Local-only, rule-IDs-only safety/accountability channel; no observability holes by design.
|
|
114
|
+
- **Fail-open:** every exception in the emit helper is swallowed — never changes the calling hook's exit code, stdout, or stderr.
|
|
115
|
+
- **Consumers:** `skills/session-review/SKILL.md`, `skills/harness-audit/SKILL.md`, `agents/session-analysis.md`, `skills/cost-report/`, `skills/run-report/SKILL.md` (#1167), `hooks/lib/review_gate_corroboration.py` (#1461 `record` rows; #1763 also reads `dispatch-failure` rows as negative evidence for the gate veto), future `agent-telemetry` cross-machine aggregation (#178).
|
|
116
|
+
|
|
117
|
+
---
|
|
118
|
+
|
|
119
|
+
## `review-verdicts.jsonl`
|
|
120
|
+
|
|
121
|
+
**Added by #2166** (plan: `plans/2164-verdict-ledger-writer.md`, Slice 2). A
|
|
122
|
+
**new, separate** store from `boundary-events.jsonl` (Decision 1) — not an
|
|
123
|
+
overload of that stream's `record` decision — because a per-file verdict
|
|
124
|
+
needs a real `file_path`, which `boundary_events.py`'s own "never write free
|
|
125
|
+
text ... file paths ... must never appear" invariant forbids. Records, per
|
|
126
|
+
genuine review-agent dispatch, an outcome (`pass` \| `findings`) bound to
|
|
127
|
+
`(lens, file_path, file_content_hash)` — a verdict about *this exact file
|
|
128
|
+
content*, not about any one diff, so it can be looked up again the next time
|
|
129
|
+
the same content recurs regardless of which diff produced it.
|
|
130
|
+
|
|
131
|
+
The recorder identifies which lens dispatched via the native
|
|
132
|
+
`attributionAgent` field the harness stamps on the subagent's own transcript
|
|
133
|
+
records (`hooks/lib/cost_meter.py`'s "Attribution dimensions" mechanism,
|
|
134
|
+
reused via `scripts/lib/session_log.records`), falling back to the
|
|
135
|
+
documented Task/Agent-dispatch join only when that field is absent. It reads
|
|
136
|
+
the in-scope file list from a structured marker
|
|
137
|
+
(`skills/code-review/SKILL.md` step 4: `Files in scope for this review:
|
|
138
|
+
<path>, ...`) in the dispatch prompt — the subagent transcript's own first
|
|
139
|
+
turn — and cross-references it against the agent's final JSON result's
|
|
140
|
+
`issues[].file` list (`knowledge/review-agent-output-contract.md`).
|
|
141
|
+
**Disclosed trust boundary (Decision 4a):** the in-scope list is the
|
|
142
|
+
orchestrating session's own declared scope, not independently re-verified
|
|
143
|
+
against any diff — the property this store adds is that a *real*
|
|
144
|
+
`SubagentStop` event occurred for a *registered* review agent, not
|
|
145
|
+
omniscient verification of review depth.
|
|
146
|
+
|
|
147
|
+
| Field | Type | Values / source |
|
|
148
|
+
| --- | --- | --- |
|
|
149
|
+
| `ts` | string | ISO-8601 UTC `%Y-%m-%dT%H:%M:%SZ` |
|
|
150
|
+
| `lens` | string | The dispatched review agent's registered name (e.g. `structure-review`), plugin-prefix-stripped |
|
|
151
|
+
| `file_path` | string | One file the dispatch prompt's scope marker declared in scope, in its canonical form: `cwd`-relative POSIX (forward-slash) path, not the raw form the scope marker carried — falls back to an absolute resolved POSIX path only when the file can't be expressed relative to `cwd` |
|
|
152
|
+
| `file_content_hash` | string | sha256 hex digest of `file_path`'s content at the time the recorder ran (current content, not the content at dispatch time) |
|
|
153
|
+
| `outcome` | string enum | `pass` \| `findings` — whether `file_path` appears in the agent's final `issues[]` |
|
|
154
|
+
| `plugin_version` | string | From `.claude-plugin/plugin.json` |
|
|
155
|
+
| `session_id` | string, optional | Opaque per-session ID, when present in the hook payload |
|
|
156
|
+
|
|
157
|
+
- **Emitter:** `hooks/review_verdict_recorder.py` (a `SubagentStop` hook) via `hooks/lib/review_verdicts.emit_review_verdict()`. No-op (zero rows) for any `subagent_type` outside `hooks/lib/review_agent_registry`'s closed set of registered `agents/*-review.md` names, and fail-open throughout (missing/unreadable transcript, unresolved `subagent_type`, a missing/reformatted scope marker, or an unparseable final JSON result all degrade to zero rows, never an exception); a single deleted/unreadable in-scope file is skipped without affecting the other rows.
|
|
158
|
+
- **Consent:** ALWAYS-ON — same posture as `boundary-events.jsonl` (Decision 2), not gated by `DEV_TEAM_TELEMETRY`/`~/.claude/telemetry.json`. Local-only, mechanical accountability data (lens/path/hash/outcome), no prose.
|
|
159
|
+
- **Fail-open:** every exception in `emit_review_verdict()` is swallowed — never changes the calling hook's exit code, stdout, or stderr. `hooks/lib/review_verdicts.load_verdicts()` mirrors this on the read side: an absent file, a corrupted line, or a stale `plugin_version` row all degrade to "no usable rows", never an exception.
|
|
160
|
+
- **Consumers:** none yet — this slice is deliberately writer-only (#2167 is the queued consumer slice).
|
|
161
|
+
|
|
162
|
+
---
|
|
163
|
+
|
|
164
|
+
## `telemetry.jsonl`
|
|
165
|
+
|
|
166
|
+
Opt-in usage beacon: which slash commands / skills get invoked, and whether
|
|
167
|
+
the pre-commit review gate fired or was bypassed.
|
|
168
|
+
|
|
169
|
+
| Field | Type | Values / source |
|
|
170
|
+
| --- | --- | --- |
|
|
171
|
+
| `ts` | string | ISO-8601 UTC |
|
|
172
|
+
| `event` | string enum | `command` \| `skill` \| `gate` |
|
|
173
|
+
| `name` | string | Grammar-matched slash-command name, skill name, or `pre-pr-review` (#1886 — the gate moved from `git commit` to `gh pr create`) |
|
|
174
|
+
| `outcome` | string | `invoked` \| `fired` \| `bypassed` |
|
|
175
|
+
| `plugin_version` | string | From `.claude-plugin/plugin.json` |
|
|
176
|
+
|
|
177
|
+
- **Emitter:** `hooks/telemetry.py::_emit()`. Written to `~/.claude/metrics/telemetry.jsonl` — home-scoped, out of the project entirely (#1405/#1406), never a project's own `metrics/`.
|
|
178
|
+
- **Consent:** opt-in — `~/.claude/telemetry.json` `{"enabled": true}`, home-scoped only. `DEV_TEAM_TELEMETRY` and a project-scoped `<cwd>/.claude/telemetry.json` are now inert (one-time-per-session stderr notice only, no effect on consent). Off by default; nothing recorded, nothing leaves the machine.
|
|
179
|
+
- **Consumers:** `skills/telemetry/SKILL.md`, `plugins/dev-team/scripts/session_report.py`.
|
|
180
|
+
|
|
181
|
+
---
|
|
182
|
+
|
|
183
|
+
## `cost-metering.jsonl`
|
|
184
|
+
|
|
185
|
+
Per-session token/cost summary, incrementally accumulated from the
|
|
186
|
+
transcript on each `Stop` hook fire.
|
|
187
|
+
|
|
188
|
+
| Field | Type | Values / source |
|
|
189
|
+
| --- | --- | --- |
|
|
190
|
+
| `timestamp` | string | ISO-8601 UTC |
|
|
191
|
+
| `transcript` | string | Transcript file basename (not full path) |
|
|
192
|
+
| `total` | object | Aggregated token counts + `cost_usd` + `messages` across the session |
|
|
193
|
+
| `by_model` | object | Per-model slim breakdown: `cost_usd`, `input_tokens`, `output_tokens` |
|
|
194
|
+
| `by_thread` | object | Per-thread slim breakdown, same shape as `by_model` |
|
|
195
|
+
| `by_agent_type` | object | Per-agent-type slim breakdown, same shape as `by_model`: `main` for main-loop turns; sidechain turns keyed by subagent type via `attributionAgent` or the Task-dispatch join; honest `unattributed` bucket when neither signal exists (#1094) |
|
|
196
|
+
|
|
197
|
+
- **Emitter:** `hooks/cost_meter.py` (wrapper) → `hooks/lib/cost_meter.py::cmd_record()`.
|
|
198
|
+
- **Consent:** gated by `telemetry_consent.is_enabled()` (`~/.claude/telemetry.json` `{"enabled": true}`, home-scoped) — no longer unconditional as of Slice 2 (#1406).
|
|
199
|
+
- **Consumers:** `skills/cost-report/SKILL.md`, `skills/harness-audit/SKILL.md`, `cmd_regression`/`cmd_pace` in the same library, `skills/run-report/SKILL.md` (#1167, best-effort only — see that skill's Join limitations section: this stream has no `session_id` field).
|
|
200
|
+
|
|
201
|
+
---
|
|
202
|
+
|
|
203
|
+
## `phase-markers.jsonl`
|
|
204
|
+
|
|
205
|
+
Per-phase context-pollution markers (#1520), one row appended at each `/handoff`
|
|
206
|
+
(a phase boundary). Distinct stream from `cost-metering.jsonl` — deliberately
|
|
207
|
+
kept out of that log's incremental `record` state so this additive dimension
|
|
208
|
+
never touches the security-sensitive hot path.
|
|
209
|
+
|
|
210
|
+
| Field | Type | Values / source |
|
|
211
|
+
| --- | --- | --- |
|
|
212
|
+
| `timestamp` | string | ISO-8601 UTC |
|
|
213
|
+
| `transcript` | string | Transcript file basename (not full path) |
|
|
214
|
+
| `phase` | string | Phase label — the first `/handoff` args token when sane, else `handoff` (or `unlabeled` for a direct library call) |
|
|
215
|
+
| `resident_tokens` | int | Main-loop context occupancy at the boundary: the most-recent non-sidechain turn's `input + cache_read + cache_creation` |
|
|
216
|
+
| `spent_output_cumulative` | int | Cumulative main-loop `output_tokens` across the session up to this boundary (monotonic; `phase-report` deltas it into per-phase spend) |
|
|
217
|
+
|
|
218
|
+
- **Emitter:** `hooks/phase_marker.py` (PostToolUse:Skill, filters to `handoff`) → `hooks/lib/cost_meter.py::cmd_phase_mark()`.
|
|
219
|
+
- **Consent:** gated by `telemetry_consent.is_enabled()`; shares the cost meter's `DEV_TEAM_COST_METER=off` opt-out.
|
|
220
|
+
- **Consumers:** `skills/cost-report/SKILL.md` (§ Context pollution, via `phase-report`), `skills/harness-audit/SKILL.md` (§ Analyze orchestration complexity).
|
|
221
|
+
|
|
222
|
+
---
|
|
223
|
+
|
|
224
|
+
## `artifact-usage.json`
|
|
225
|
+
|
|
226
|
+
Not JSONL — a single JSON object keyed by skill/agent name, upserted on
|
|
227
|
+
every invocation.
|
|
228
|
+
|
|
229
|
+
| Field | Type | Values / source |
|
|
230
|
+
| --- | --- | --- |
|
|
231
|
+
| `<skill_name>.use_count` | integer | Cumulative invocation count |
|
|
232
|
+
| `<skill_name>.last_used_at` | string | ISO-8601 UTC of the most recent invocation |
|
|
233
|
+
| `<skill_name>.lifecycle` | string | `active` (set on creation; other lifecycle states are assigned externally by `/artifact-lifecycle`) |
|
|
234
|
+
|
|
235
|
+
- **Emitter:** `hooks/telemetry.py::_upsert_artifact_usage()` (atomic rewrite via tempfile + `os.replace`). Written to `~/.claude/metrics/artifact-usage.json` — home-scoped, out of the project entirely (#1405/#1406), never a project's own `metrics/`.
|
|
236
|
+
- **Consent:** follows `telemetry.jsonl`'s opt-in gate (`~/.claude/telemetry.json` `{"enabled": true}`, home-scoped). The project-scoped explicit-off switch this section used to document (a project-level `.claude/telemetry.json` `{"enabled": false}` disabling usage tracking specifically) no longer exists — project-scoped `.claude/telemetry.json` is inert entirely, same as `telemetry.jsonl`'s row above.
|
|
237
|
+
- **Consumers:** `skills/artifact-lifecycle/SKILL.md`.
|
|
238
|
+
|
|
239
|
+
---
|
|
240
|
+
|
|
241
|
+
## `gate-bypass-audit.jsonl`
|
|
242
|
+
|
|
243
|
+
Accountability record for a bypass of the review-corroboration gate. #1886
|
|
244
|
+
moved the gate from `git commit` to `gh pr create`; this stream now carries
|
|
245
|
+
`hooks/pre_pr_review.py`'s `PR_GATE_BYPASS_REASON` bypasses.
|
|
246
|
+
`hooks/pre_commit_review.py` is now a documented no-op and no longer writes
|
|
247
|
+
to this stream — historical rows from before the migration
|
|
248
|
+
(`triggeredBy: "--no-verify"`/`"-n"`) remain valid history but no new ones
|
|
249
|
+
are produced.
|
|
250
|
+
|
|
251
|
+
| Field | Type | Values / source |
|
|
252
|
+
| --- | --- | --- |
|
|
253
|
+
| `timestamp` | string | ISO-8601 UTC |
|
|
254
|
+
| `branch` | string | Current git branch |
|
|
255
|
+
| `triggeredBy` | string | `PR_GATE_BYPASS_REASON` (current); `--no-verify`/`-n` (historical, pre-#1886) |
|
|
256
|
+
| `reason` | string | Value of `PR_GATE_BYPASS_REASON` (current) / `GATE_BYPASS_REASON` (historical) — human/agent-authored, required to be non-empty |
|
|
257
|
+
| `stagedFileCount` | integer | Count of files in the branch diff at bypass time (historical rows: staged files at commit time) |
|
|
258
|
+
| `pluginVersion` | string | From `.claude-plugin/plugin.json` |
|
|
259
|
+
|
|
260
|
+
- **Emitter:** `hooks/pre_pr_review.py::_record_bypass_audit()` (#1886). `hooks/pre_commit_review.py::_record_bypass_audit()` was the historical emitter, now removed along with the rest of that module's gating logic.
|
|
261
|
+
- **Consent:** unconditional — accountability record for an actively-chosen bypass, not passive usage telemetry.
|
|
262
|
+
- **Consumers:** `skills/code-review/SKILL.md`, `docs/code-review-process.md`.
|
|
263
|
+
|
|
264
|
+
---
|
|
265
|
+
|
|
266
|
+
## `gate-bypass.jsonl`
|
|
267
|
+
|
|
268
|
+
Accountability record for `MUTATION_SMOKE_GATE_SKIP=1` bypasses of the
|
|
269
|
+
mutation-testing smoke gate. Distinct stream from `gate-bypass-audit.jsonl`
|
|
270
|
+
above (different gate, different hook).
|
|
271
|
+
|
|
272
|
+
| Field | Type | Values / source |
|
|
273
|
+
| --- | --- | --- |
|
|
274
|
+
| `timestamp` | string | ISO-8601 UTC (`Z`-suffixed) |
|
|
275
|
+
| `hook` | string | Always `mutation-testing-smoke-gate` |
|
|
276
|
+
| `command_hash` | string | First 16 hex chars of `sha256(raw_command)` — the raw command is never logged |
|
|
277
|
+
| `cwd` | string | Payload cwd |
|
|
278
|
+
|
|
279
|
+
- **Emitter:** `hooks/mutation_testing_smoke_gate.py::log_bypass_audit()`.
|
|
280
|
+
- **Consent:** unconditional.
|
|
281
|
+
- **Consumers:** `skills/mutation-testing/SKILL.md`.
|
|
282
|
+
|
|
283
|
+
---
|
|
284
|
+
|
|
285
|
+
## `config-changelog.jsonl`
|
|
286
|
+
|
|
287
|
+
Audit trail for `/feedback-learning` config changes and human-oversight
|
|
288
|
+
protocol events (approval / override / pause / stop).
|
|
289
|
+
|
|
290
|
+
| Field | Type | Values / source |
|
|
291
|
+
| --- | --- | --- |
|
|
292
|
+
| `timestamp` | string | ISO-8601 UTC |
|
|
293
|
+
| `type` | string enum | `amend` \| `approval` \| `override` \| `pause` \| `stop` (feedback-learning change types, or oversight event types) |
|
|
294
|
+
| `trigger` | string | `user` (who/what triggered the change) |
|
|
295
|
+
| `description` | string | Human/agent-authored summary of what happened and why (deliberate audit note, not incidental free text) |
|
|
296
|
+
| `file_modified` | string, optional | Config file touched (feedback-learning changes) |
|
|
297
|
+
| `section_modified` | string, optional | Section within the file |
|
|
298
|
+
| `previous_value` / `new_value` | string, optional | Before/after values |
|
|
299
|
+
| `approved_by` | string, optional | Who approved the change |
|
|
300
|
+
|
|
301
|
+
- **Emitter:** `/feedback-learning` skill (model-authored append) and `/human-oversight-protocol` skill.
|
|
302
|
+
- **Consent:** unconditional (append-only governance record).
|
|
303
|
+
- **Consumers:** `skills/feedback-learning/SKILL.md`, `skills/human-oversight-protocol/SKILL.md`, `skills/governance-compliance/SKILL.md`.
|
|
304
|
+
|
|
305
|
+
---
|
|
306
|
+
|
|
307
|
+
## `session-digest.jsonl`
|
|
308
|
+
|
|
309
|
+
Trend digest from `/session-review` (backed by `${CLAUDE_PLUGIN_ROOT}/scripts/session_report.py --profile maintainer`):
|
|
310
|
+
aggregate counts only, no file names, prompts, command strings, or code.
|
|
311
|
+
|
|
312
|
+
| Field | Type | Values / source |
|
|
313
|
+
| --- | --- | --- |
|
|
314
|
+
| `recorded_at` | string | UTC ISO-8601 of the run |
|
|
315
|
+
| `plugin_version` | string | The `dev-team` plugin's `.claude-plugin/plugin.json` version active when this record was produced (`"unknown"` if the manifest couldn't be read), #1471. Lets consumers tell a friction already fixed in a newer version apart from one that's still current — see `--version-scope` below |
|
|
316
|
+
| `sessions`, `transcripts` | integer | How many sessions/transcripts the digest covered |
|
|
317
|
+
| `tokens` | object | Input/output/cache token totals |
|
|
318
|
+
| `cost_usd`, `cache_hit_ratio` | number | Session cost and cache-read efficiency |
|
|
319
|
+
| `rework` | object | `failed_edits`, `repeated_file_edits`, `retried_bash_commands`, `retried_bash_commands_by_skill`/`retried_bash_commands_by_agent` (#2110 — each retry attributed, at the moment it's detected, to whichever skill/agent is sticky-active; `retried_bash_commands` is the derived sum, never a second independent count), `repeated_verify_runs`, `permission_denials`, `compaction_events` |
|
|
320
|
+
| `accuracy` | object | `tool_calls`, `tool_error_rate`, `user_correction_turns`, `by_skill`/`by_agent` (correction counts, double-bucketed against whichever skill/agent is sticky-active), `correction_rate_by_skill`/`correction_rate_by_agent` (corrections per skill invocation / agent dispatch — absent for a never-invoked name, never a misleading `0.0`), `correction_causes` (#2013, below) |
|
|
321
|
+
| `utilization` | object | `skills_invoked`, `agents_invoked` (agent RUNS), `agent_dispatches` (Agent/Task tool calls), `never_observed_skills`, `never_observed_agents` |
|
|
322
|
+
|
|
323
|
+
**`accuracy.correction_causes` (#2013).** Deterministic cause data for every
|
|
324
|
+
detected correction turn — no model call, see
|
|
325
|
+
`plugins/dev-team/scripts/lib/session_log/corrections.py`'s module
|
|
326
|
+
docstring for the classifier. Present in BOTH `session-digest/v4` and
|
|
327
|
+
`downstream-session-report/v4` (a correction's producing component is
|
|
328
|
+
exactly the "which components generate the most corrections per dispatch"
|
|
329
|
+
question the issue exists to answer, and that question is as live for a
|
|
330
|
+
downstream user's own report as for this repo's own trend stream). Shape:
|
|
331
|
+
|
|
332
|
+
| Field | Type | Values |
|
|
333
|
+
| --- | --- | --- |
|
|
334
|
+
| `by_what` | object | Counts by `code-edit` \| `plan` \| `review-finding` \| `tool-choice` \| `factual-claim` \| `other` |
|
|
335
|
+
| `by_component` | object | Counts by `main-loop` or the single most-recently-dispatched skill/agent name (never double-bucketed, unlike `accuracy.by_skill`/`by_agent`) |
|
|
336
|
+
| `by_shape` | object | Counts by `reverted` \| `redirected` \| `narrowed-scope` \| `flagged-wrong` \| `not-what-asked` \| `ambiguous` |
|
|
337
|
+
| `ambiguous_share` | number | `by_shape["ambiguous"] / user_correction_turns` — the classifier's honest inference-share statistic, never hidden |
|
|
338
|
+
|
|
339
|
+
Never the correction text itself — only these four closed-vocabulary labels.
|
|
340
|
+
`test_session_log_corrections.py::test_classify_correction_never_leaks_correction_text`
|
|
341
|
+
pins this the same way
|
|
342
|
+
`test_session_report_golden.py::test_no_sentinel_leaks_in_either_golden`
|
|
343
|
+
pins the rest of this stream.
|
|
344
|
+
|
|
345
|
+
`session-digest/v2` (#1994) counts dispatched agents' own transcripts for the
|
|
346
|
+
first time, so token/tool-call/rework totals jump against v1, and
|
|
347
|
+
`retried_bash_commands` / `repeated_verify_runs` moved from a session-keyed to
|
|
348
|
+
a per-thread basis. Records from the two eras are not comparable; split on
|
|
349
|
+
`schema` before trending. `session-digest/v3` (#2046) is a schema-label-only
|
|
350
|
+
bump: it is emitted by the new unified `session_report.py --profile
|
|
351
|
+
maintainer` entry point, which every real consumer now runs (#2047) instead
|
|
352
|
+
of the earlier monorepo-only extractor (retired in #2048). No data-shape
|
|
353
|
+
change from v2 — a v2 and a v3 record trend together. `session-digest/v4`
|
|
354
|
+
(#2018) changes what `plugin_version` MEANS on the per-session
|
|
355
|
+
`session-sync/v4` records `--sync-out` writes (see "Version tagging" below)
|
|
356
|
+
— a v3 sync record and a v4 sync record are not comparable on that field,
|
|
357
|
+
though every other field is unchanged and they still trend together
|
|
358
|
+
otherwise. The single-shot digest/trend record keeps its v2/v3 meaning for
|
|
359
|
+
`plugin_version` (extraction-time) even under the v4 label — see below.
|
|
360
|
+
|
|
361
|
+
- **Emitter:** `/session-review` skill via `${CLAUDE_PLUGIN_ROOT}/scripts/session_report.py --profile maintainer`.
|
|
362
|
+
- **Consent:** unconditional (aggregate counts only, no file/prompt/command content).
|
|
363
|
+
- **Enforcement (#2045):** every name/label-shaped field this stream (and the
|
|
364
|
+
shipped `session_report.py --profile downstream` report) emits passes
|
|
365
|
+
through `plugins/dev-team/scripts/lib/session_log/redact.redact()` — the
|
|
366
|
+
one function both extractors route file basenames, project labels, skill
|
|
367
|
+
names, agent names, and model ids through before writing them out.
|
|
368
|
+
Previously this line's promise was a convention restated independently at
|
|
369
|
+
each call site; `redact()` is the single enforcement point, pinned by
|
|
370
|
+
`tests/scripts/test_session_report_golden.py::test_no_sentinel_leaks_in_either_golden`
|
|
371
|
+
against a corpus seeding real prompt text, source code, a full shell
|
|
372
|
+
command string, and absolute POSIX/Windows paths.
|
|
373
|
+
- **Consumers:** `skills/harness-audit/SKILL.md` (joins with self-reported task logs), `agents/session-analysis.md`.
|
|
374
|
+
- **Version tagging (#1471, revised #2018):** `plugin_version` is also carried on the per-session `session-sync/v1`-`v4` records synced by `--sync-out` (used by `--rollup`/`--escalate`/`--correlate`) and on the single-shot digest itself. These are no longer the same value stamped two ways — as of `session-sync/v4` (#2018) the two paths diverge deliberately:
|
|
375
|
+
- **Per-session `--sync-out` records (`session-sync/v4`)** — the path that runs unattended at every `SessionStart` (`.claude/ensure_session_archive.py`) and accumulates into the durable, cross-release archive — carry the SESSION's OWN version, resolved by `resolve_session_plugin_version()` from that project's `<cwd>/.claude/metrics/boundary-events.jsonl` (a stream stamped live, at hook-dispatch time, by `hooks/lib/boundary_events.py`). When a session's boundary-events carry more than one plugin_version (e.g. an in-session `/dev-team:upgrade`), the EARLIEST by `ts` wins — documented, not solved, since a session spanning two releases has no single correct answer. A session with no matching boundary-events record (never dispatched anything through a hook that stamps `session_id`) is tagged the explicit string `"unknown"` — never silently the extractor's own checked-out version, which is the exact mislabeling #2018 was filed to close (5,562 archived records, all `plugin_version: null`, none of them attributable to a release).
|
|
376
|
+
- **The single-shot digest/trend record** (`--transcript`/`--project-dir`, no `--sync-out`) is UNCHANGED behavior: it still reflects the plugin version active on the machine *at extraction time*. It can aggregate many sessions into one record, and a single field cannot honestly attribute a multi-session aggregate to one session's version — this is a deliberate scope limitation, not an oversight. In the common case (`/session-review` run against the current project, right after the session it's summarizing) extraction-time and session-time coincide anyway; the divergence #2018 closes is specific to `--sync-out`'s day/weeks-later, unattended, cross-session runs.
|
|
377
|
+
- **Version scoping (#1480):** `session_report.py --profile maintainer --rollup`/`--escalate`/`--correlate` accept `--version-scope {all,current-and-previous}` (default `all`, unbounded history). `current-and-previous` drops any record whose `plugin_version` isn't the currently-installed version or the version immediately before it *as observed in the digests being read* (there is no release-history lookup) — records with no `plugin_version` at all (pre-#1471 data), or tagged the explicit `"unknown"` string (#2018), are dropped too, since neither can be proven current. The result gains a `version_window` field (the concrete versions included; `[]` when unscoped). `/session-review` defaults to local-only; its `--cross-machine` opt-in always applies `current-and-previous` scoping (see its SKILL.md) — the skill itself never exposes an unscoped cross-machine mode. Unbounded history across every version remains available only via a direct `session_report.py --profile maintainer --rollup ... --version-scope all` invocation (the CLI default), outside the skill.
|
|
378
|
+
- **Version-filtered downstream report coverage (#2018):** `session_report.py --profile downstream --plugin-version VERSION` scopes the report to sessions whose project recorded `VERSION` in `boundary-events.jsonl` (`sessions_matching_plugin_version`, best-effort, same source as above) — sessions with no matching event are excluded from the report's own counts exactly as before. What's new: the report's top-level `version_filter_coverage` field (present, non-null, only when `--plugin-version` was passed) states `requested_version`, `sessions_considered` (every distinct session in the report's own since/until window, ignoring the version filter), `sessions_attributed` (how many of those matched), `sessions_attributed_other_version` (recorded a DIFFERENT known `plugin_version` — `_sessions_with_known_plugin_version` — the filter correctly excluding them, not a data gap), and `sessions_unattributed` (no resolvable `plugin_version` at all — genuinely missing data) — so an operator can tell "belongs to another release" apart from "this repo's own instrumentation never saw it," instead of one conflated exclusion count. The exclusion behavior itself is unchanged; this only makes it observable and precise.
|
|
379
|
+
|
|
380
|
+
---
|
|
381
|
+
|
|
382
|
+
## `review-value.jsonl`
|
|
383
|
+
|
|
384
|
+
Whether a `/build` inline review checkpoint actually changed anything —
|
|
385
|
+
counts and outcomes only, never code or file content.
|
|
386
|
+
|
|
387
|
+
| Field | Type | Values / source |
|
|
388
|
+
| --- | --- | --- |
|
|
389
|
+
| `timestamp` | string | ISO-8601 UTC |
|
|
390
|
+
| `plan` | string | Plan file path |
|
|
391
|
+
| `slice` | string | Slice number |
|
|
392
|
+
| `step` | string | Step number (`N.M`) or `all` |
|
|
393
|
+
| `checkpoint` | string enum | `step` \| `slice` \| `backstop` (the Step-6 backstop pass, #1962) |
|
|
394
|
+
| `complexity` | string enum | `standard` \| `complex` |
|
|
395
|
+
| `agents_run` | array of string | Review agents dispatched |
|
|
396
|
+
| `issues_found`, `issues_fixed`, `fix_iterations` | integer | Counts |
|
|
397
|
+
| `severity_breakdown` | object | `{errors, warnings, suggestions}` counts (same enum as `/code-review`); the three sum to `issues_found`. Lets `/harness-audit` Step 3 flag mostly-minor lenses (#1256). Absent on pre-#1256 rows |
|
|
398
|
+
| `source` | string enum | Row provenance: `build-checkpoint` (fix-applying `/build` inline checkpoint) \| `build-backstop` (fix-applying `/build` Step-6 backstop pass, #1962) \| `code-review` (read-only standalone review) \| `harness` (the `evals/code-review-benchmark/` replay harness — #2051, see below). **Absent = `build-checkpoint`** (back-compat). `/harness-audit` Step 4 excludes `code-review` rows from fix-rate drop-candidate logic (#1257); `build-backstop` rows are fix-applying and stay in it |
|
|
399
|
+
| `diff_shape` | string enum | Shape of the reviewed diff: `test-only` (every changed file provably a test per `knowledge/test-file-indicators.md`) \| `mixed` (anything else). Classified by `skills/code-review/scripts/change_shape.py`'s `isTestOnly`, never by eye; include-biased, so `test-only` is never over-claimed. Lets `/harness-audit` split per-lens outcomes by diff shape — the evidence a test-only lens gate waits on (#1964). Absent on pre-#1964 rows |
|
|
400
|
+
| `outcome` | string enum | `no-op` \| `fixed` \| `escalated` \| `skipped` (backstop only — suppressed by `--backstop-review=skip`; never counted in a rate, #1962) |
|
|
401
|
+
|
|
402
|
+
- **Emitter:** `/build` skill (model-authored append, sub-step 7) writes `source: "build-checkpoint"` for inline checkpoints and, from Step 6, `source: "build-backstop"` for the backstop pass (#1962). Disable with `DEV_TEAM_REVIEW_VALUE=off`.
|
|
403
|
+
- **Consent:** unconditional when enabled (no code/file content recorded).
|
|
404
|
+
- **Consumers:** `skills/cost-report/SKILL.md`, `skills/harness-audit/SKILL.md`.
|
|
405
|
+
- **Provenance (#1257):** fix-rate ROI is only meaningful for fix-applying rows. A read-only review that never applies fixes (`source: "code-review"`) always has `issues_fixed: 0`; Step 4 must not read that as a zero-value drop candidate — it reports finding-rate for those instead.
|
|
406
|
+
|
|
407
|
+
### Round rows — `source: "code-review"` with a `round` field (#1624)
|
|
408
|
+
|
|
409
|
+
`/code-review` appends one row **per dispatch round** (round 1 = the initial
|
|
410
|
+
panel; each fix-loop iteration's re-dispatch set is one further round), so
|
|
411
|
+
|
|
412
|
+
# 1623's "is this agent's dispatch frequency value or churn?" becomes
|
|
413
|
+
|
|
414
|
+
answerable. Written by `skills/code-review/scripts/review_round_log.py`.
|
|
415
|
+
|
|
416
|
+
| Field | Type | Values / source |
|
|
417
|
+
| --- | --- | --- |
|
|
418
|
+
| `timestamp` | string | ISO-8601 UTC |
|
|
419
|
+
| `source` | string enum | Always `code-review` for these rows |
|
|
420
|
+
| `round` | integer | 1 = initial panel; each fix-loop re-dispatch set increments. **Presence of this field is what distinguishes a round row** from the original whole-run `code-review` row |
|
|
421
|
+
| `agents_run` | array of string | Registered review agents dispatched this round (sorted, deduped) |
|
|
422
|
+
| `findings_new` | integer | Findings whose signature was not present in a prior round (signature identity: #1625). The round-row analogue of `issues_found` |
|
|
423
|
+
| `findings_carried` | integer | Prior-round signatures that survived this round's fix attempt |
|
|
424
|
+
| `severity_breakdown` | object | `{errors, warnings, suggestions}` over `findings_new`; same enum as the `/build` rows |
|
|
425
|
+
| `fix_provenance_new` | integer | How many of `findings_new` fall inside the line ranges the **previous** round's fix touched — the judgment-free "the fix introduced it" signal. Computed by unified-diff interval math, never by model judgment. Always `0` for `round: 1` (no preceding fix) |
|
|
426
|
+
| `dispatch_purpose` | string enum | `discovery` (a panel looking for new problems) \| `verification` (confirming a specific fix, #1628) \| `closing` (the scoped gate-closing pass, #1626) |
|
|
427
|
+
| `outcome` | string enum | `no-op` \| `fixed` \| `escalated` — same enum as the `/build` rows |
|
|
428
|
+
|
|
429
|
+
- **Emitter:** `/code-review` (steps 5b-i and 6a) via `review_round_log.py`.
|
|
430
|
+
- **Consent:** **unconditional** — written to `.claude/metrics/` like `boundary-events.jsonl`, *not* gated behind `~/.claude/telemetry.json` the way `/build`'s rows are. Rationale (#1624 design item 2): this is the same class of local, counts-only operational stream the commit gate itself already depends on, and consent-gating it would make #1623's success criteria depend on consent being enabled per dev machine. Rows carry counts, agent names, and enum values only — no file paths, code, or finding text.
|
|
431
|
+
- **Consumers:** `skills/harness-audit/SKILL.md` Step 4a (churn ratio, per-agent discovery-vs-verification split, gate recidivism).
|
|
432
|
+
- **Backstop rows (#1962).** `source: "build-backstop"` marks `/build`'s Step-6 pass — the one review layer whose files an inline checkpoint already reviewed in the same run. It is fix-applying (the `--internal` panel runs the review-fix loop), so it belongs in fix-rate analysis alongside `build-checkpoint`; what it exists to answer is whether that duplicated layer is ~all `no-op`, which is the evidence `/build`'s `--backstop-review=skip` flag waits on. `outcome: "skipped"` marks a backstop suppressed by that flag: recorded so the suppression is visible in the same stream, and excluded from every rate because it never ran.
|
|
433
|
+
- **Reconciling the `source` values.** `build-checkpoint` and `build-backstop` rows are fix-applying and carry `plan`/`slice`/`step`/`checkpoint`/`complexity`/`issues_found`/`issues_fixed`/`fix_iterations`. `code-review` rows are read-only; those with a `round` field use the round schema above. A consumer wanting "how many issues did this row surface" should read `(.issues_found // .findings_new)`, which covers all three shapes.
|
|
434
|
+
|
|
435
|
+
### Harness rows — `source: "harness"` (#2051)
|
|
436
|
+
|
|
437
|
+
Written by `evals/code-review-benchmark/runner.emit_review_value_rows()` —
|
|
438
|
+
the `/code-review` **replay harness** (#821), not a live session. One row
|
|
439
|
+
per lens per dispatch, written straight from the parsed `/code-review
|
|
440
|
+
--json` payload's `agents[]` list — by mechanism, not by agent instruction
|
|
441
|
+
— so every dispatched lens gets a row regardless of outcome, including a
|
|
442
|
+
lens that found nothing. This is the fix for the collection bias #2019/
|
|
443
|
+
#1512 documented in the live writers above.
|
|
444
|
+
|
|
445
|
+
| Field | Type | Values / source |
|
|
446
|
+
| --- | --- | --- |
|
|
447
|
+
| `agents_run` | array of string | Always exactly one lens name — one row per lens, not per dispatch batch |
|
|
448
|
+
| `issues_found` | integer | Count of that lens's issues this dispatch |
|
|
449
|
+
| `severity_breakdown` | object | `{errors, warnings, suggestions}`, same enum as the live rows |
|
|
450
|
+
| `outcome` | string | Always `"no-op"` — the harness is read-only and never applies a fix |
|
|
451
|
+
| `diff_shape` | string enum | `test-only` \| `mixed`, same classifier as the live rows; the recorded-diff adapter (below) is what actually supplies real `test-only` cases |
|
|
452
|
+
| `dataset` | string | `defects4j` \| `bugsjs` \| `recorded-diff` |
|
|
453
|
+
| `project`, `bug_id` | string | The benchmark case's identifiers (not a `/build` plan/slice/step — a structurally different key space) |
|
|
454
|
+
|
|
455
|
+
- **Emitter:** `runner.emit_review_value_rows()`, called from `runner.run_case()` and `runner.run_recorded_diff_case()` after a dispatch's `--json` payload parses successfully. Never called for an unparseable dispatch — that failure is already captured by the harness's own `skipped.jsonl`.
|
|
456
|
+
- **File location, deliberately not `.claude/metrics/`.** Rows land in the harness's own results directory (`evals/code-review-benchmark/results/review-value.jsonl` by default, `--results-dir` elsewhere) — never the live metrics tree. This is a structural guarantee against pooling, on top of the `source: "harness"` label itself: even a caller reading the wrong file could not accidentally merge harness rows into the live population, because they are not in the same file.
|
|
457
|
+
- **Consumers:** `skills/harness-audit/SKILL.md` §4b, which reads this stream as a separate, explicitly-labelled population and must never merge it into Step 3/4's live-row computations.
|
|
458
|
+
- **Recorded-diff adapter (#2051).** `evals/code-review-benchmark/adapters/recorded_diff_adapter.py` supplies `dataset: "recorded-diff"` cases from saved diffs (most usefully real `/test-improve` Phase-5 diffs) — the only source that can give this stream a genuine `diff_shape: "test-only"` row, since Defects4J/BugsJS are real production bug fixes and structurally cannot be test-only.
|
|
459
|
+
|
|
460
|
+
---
|
|
461
|
+
|
|
462
|
+
## `contract-failures.jsonl`
|
|
463
|
+
|
|
464
|
+
Diagnostic record for a review-agent output that fails the shared JSON
|
|
465
|
+
contract (`knowledge/review-agent-output-contract.md`) — the gap #1998
|
|
466
|
+
closes. Session-report analysis found 18.2% of review-agent outputs
|
|
467
|
+
discarded silently, with no record of which agent, what it returned, or
|
|
468
|
+
why it didn't parse; today's alternative to this stream is nothing.
|
|
469
|
+
|
|
470
|
+
| Field | Type | Values / source |
|
|
471
|
+
| --- | --- | --- |
|
|
472
|
+
| `timestamp` | string | ISO-8601 UTC |
|
|
473
|
+
| `agent` | string | Name of the review agent whose output failed validation |
|
|
474
|
+
| `shape` | string enum | `empty` \| `truncated` \| `malformed-json` \| `schema-drift` \| `not-json` — the closed set `validate_review_output.FAILURE_SHAPES` exports. `validate_review_output.py` also recognizes `clean`/`fenced`/`prose-preamble` (`SUCCESS_SHAPES`), but those three name *successful* extraction (the JSON was found and matched the contract) and so never appear as `shape` in a failure row — a successfully-extracted object that then fails schema validation is logged as `schema-drift`, not the extraction shape that found it. `malformed-json` is distinct from `truncated`: a balanced `{...}` object (unquoted keys, a trailing comma, a Python-repr dict) that still fails to parse is `malformed-json`; a `{` that never balances back to depth zero before EOF is `truncated` |
|
|
475
|
+
| `extraction` | string enum, nullable | Set whenever a JSON-shaped candidate was actually recovered before failing — i.e. `shape` is `schema-drift` or `malformed-json` — naming which of `clean`/`fenced`/`prose-preamble` recovered it, so that information survives the downgrade instead of being discarded. `null` for `empty`/`truncated`/`not-json` — no candidate was ever recovered for those |
|
|
476
|
+
| `error` | string | The specific validation/parse error (e.g. a `JSONDecodeError` message, or `status='ok' not one of [...]`), after the same secret-redaction pass as `raw_prefix` and capped at 256 characters — `_validate_schema` interpolates agent-controlled values into this string, so it needs the same two controls |
|
|
477
|
+
| `raw_prefix` | string | First 200 characters of the agent's raw output, **after** a secret-redaction pass (`validate_review_output._redact()`: this repo's canonical hardcoded-key pattern from `knowledge/owasp-detection.md`, plus common vendor token prefixes). Deliberate, capped exception to this file's "never incidental free text" default (mirroring `gate-bypass-audit.jsonl`'s `reason` field) — without seeing what was actually returned, the failure shapes #1998 exists to classify cannot be told apart. AI-authored review-agent output, which may quote repository source verbatim — including any secret present in the reviewed diff, since a lens's own job is to find and quote such things — so this is a transitive channel for repo content, not a claim that the text is free of it; the redaction pass is the actual control, the 200-char cap only bounds volume |
|
|
478
|
+
|
|
479
|
+
- **Emitter:** `skills/code-review/scripts/validate_review_output.py::log_failure()`, called once per non-contract-valid agent result during `/code-review` step 4's dispatch-failure handling.
|
|
480
|
+
- **Consent:** unconditional, matching `boundary-events.jsonl` — this is the same class of local, counts-and-diagnostics operational stream the commit gate itself already depends on.
|
|
481
|
+
- **Consumers:** `skills/code-review/scripts/contract_failure_report.py`, which joins this stream against `boundary-events.jsonl`'s dispatch counts to report a real per-agent failure *rate* (not just a count) for #1980/#1982 to read before citing any `$/finding` figure.
|
|
482
|
+
|
|
483
|
+
---
|
|
484
|
+
|
|
485
|
+
## `verify-log.jsonl`
|
|
486
|
+
|
|
487
|
+
Evidence that the project's own test/verification tooling actually exercised
|
|
488
|
+
the change end-to-end (or was legitimately skipped) before a `/build` slice
|
|
489
|
+
with a runtime surface was marked complete. Schema modeled on
|
|
490
|
+
`review-value.jsonl`.
|
|
491
|
+
|
|
492
|
+
| Field | Type | Values / source |
|
|
493
|
+
| --- | --- | --- |
|
|
494
|
+
| `timestamp` | string | ISO-8601 UTC |
|
|
495
|
+
| `plan` | string | Plan file path |
|
|
496
|
+
| `slice` | string | Slice number |
|
|
497
|
+
| `branch` | string | Current git branch |
|
|
498
|
+
| `files` | array of string | Changed runtime files in scope |
|
|
499
|
+
| `outcome` | string enum | `ran` \| `skipped` \| `failed-then-fixed` |
|
|
500
|
+
| `reason` | string, optional | Set when `outcome` is `skipped` (e.g. `"tests-only"`, `"docs-only"`) |
|
|
501
|
+
|
|
502
|
+
- **Emitter:** `/build` skill (model-authored append, sub-step 4.9).
|
|
503
|
+
- **Consent:** unconditional.
|
|
504
|
+
- **Consumers:** `${CLAUDE_PLUGIN_ROOT}/scripts/progress_guardian.py --pre-pr` (fails closed on a runtime-surface change with no matching entry), `skills/performance-metrics/SKILL.md`.
|
|
505
|
+
|
|
506
|
+
---
|
|
507
|
+
|
|
508
|
+
## `override-audit.jsonl`
|
|
509
|
+
|
|
510
|
+
Audit trail for `/code-review --force --reason "<text>"`, which skips all
|
|
511
|
+
gates and the documentation-only short-circuit.
|
|
512
|
+
|
|
513
|
+
| Field | Type | Values / source |
|
|
514
|
+
| --- | --- | --- |
|
|
515
|
+
| `timestamp` | string | ISO-8601 |
|
|
516
|
+
| `branch` | string | Current git branch |
|
|
517
|
+
| `triggeredBy` | string | Always `--force` |
|
|
518
|
+
| `reason` | string | Value of `--reason` (required, human/agent-authored) |
|
|
519
|
+
| `targetFiles` | array of string | Files the forced review targeted |
|
|
520
|
+
| `gatesSkipped` | array of string | e.g. `["lint", "type-check", "secret-scan", "semgrep", "pipeline-red"]` |
|
|
521
|
+
|
|
522
|
+
- **Emitter:** `/code-review` skill (model-authored append, step 2).
|
|
523
|
+
- **Consent:** unconditional.
|
|
524
|
+
- **Consumers:** `skills/code-review/SKILL.md`, `docs/code-review-process.md`.
|
|
525
|
+
|
|
526
|
+
---
|
|
527
|
+
|
|
528
|
+
## `eval-variance.jsonl`
|
|
529
|
+
|
|
530
|
+
Multi-trial pass@k stability trend for `/agent-eval` fixtures.
|
|
531
|
+
|
|
532
|
+
| Field | Type | Values / source |
|
|
533
|
+
| --- | --- | --- |
|
|
534
|
+
| `recorded_at` | string | ISO-8601 UTC |
|
|
535
|
+
| `schema` | string | `eval-variance/v1` |
|
|
536
|
+
| `trials` | integer | Number of trials in this run |
|
|
537
|
+
| `pairs_evaluated` | integer | Fixture/agent pairs evaluated |
|
|
538
|
+
| `flaky_count` | integer | Pairs that neither always passed nor always failed |
|
|
539
|
+
| `mean_pass_at_k` | number | Mean pass@k across evaluated agents |
|
|
540
|
+
|
|
541
|
+
- **Emitter:** `scripts/eval_variance.py --append`.
|
|
542
|
+
- **Consent:** unconditional (eval infra, not user-session telemetry).
|
|
543
|
+
- **Consumers:** `skills/agent-eval/SKILL.md`.
|
|
544
|
+
|
|
545
|
+
---
|
|
546
|
+
|
|
547
|
+
## `eval-ablation.jsonl`
|
|
548
|
+
|
|
549
|
+
Causal per-agent ablation evidence from `/agent-eval --ablation <agent>` (#868):
|
|
550
|
+
a controlled baseline-vs-ablated integration-tier delta (issues caught,
|
|
551
|
+
`testCommands` results, token cost), not accumulated usage data.
|
|
552
|
+
|
|
553
|
+
| Field | Type | Values / source |
|
|
554
|
+
| --- | --- | --- |
|
|
555
|
+
| `schema` | string | `eval-ablation/v1` |
|
|
556
|
+
| `recorded_at` | string | ISO-8601 UTC |
|
|
557
|
+
| `ablated_agent` | string | Target agent name |
|
|
558
|
+
| `fixtures` | array of strings | Integration fixtures exercised |
|
|
559
|
+
| `model` | string | Model version(s) used for orchestrator/builder dispatch — deltas are model-dependent, always recorded |
|
|
560
|
+
| `baseline` | object | `{issues_caught, test_commands: [{command, exit_code}], tokens, grade}` — full roster arm |
|
|
561
|
+
| `ablated` | object | Same shape as `baseline` — roster-minus-target-agent arm |
|
|
562
|
+
| `delta` | object | `{issues_caught, test_commands_passed, tokens}` (ablated − baseline) |
|
|
563
|
+
| `verdict` | string | e.g. `"no measured impact — supports drop"` / `"agent is load-bearing — retain"` / `"baseline failed — inconclusive"` |
|
|
564
|
+
|
|
565
|
+
- **Emitter:** `plugins/dev-team/scripts/eval_ablation.py --mode agent` (moved from `scripts/` in #1653).
|
|
566
|
+
- **Consent:** unconditional (eval infra, not user-session telemetry); opt-in/label-gated dispatch per the live-eval cost policy (#134) — the record is only ever written after an explicit operator-confirmed live run.
|
|
567
|
+
- **Consumers:** `skills/harness-audit/SKILL.md` (Step 3 drop-candidate recommendations cite the measured delta/verdict when a record exists).
|
|
568
|
+
|
|
569
|
+
---
|
|
570
|
+
|
|
571
|
+
## `refactor-freeze.jsonl`
|
|
572
|
+
|
|
573
|
+
Audit log for the tests-frozen-during-REFACTOR invariant (`#813`) — both the
|
|
574
|
+
enforcement decision and any fail-open diagnostic. Extended by `#906` with
|
|
575
|
+
`bash-freeze`, the preventive PreToolUse(Bash) sibling of `freeze`.
|
|
576
|
+
|
|
577
|
+
| Field | Type | Values / source |
|
|
578
|
+
| --- | --- | --- |
|
|
579
|
+
| `timestamp` | string | ISO-8601 |
|
|
580
|
+
| `hook` | string | `freeze` \| `bash-freeze` \| `revert` |
|
|
581
|
+
| `event` | string enum | `block` \| `fail-open` \| `revert` \| `remove` |
|
|
582
|
+
| `file` | string, optional | File path involved |
|
|
583
|
+
| `step` | string, optional | Plan step label |
|
|
584
|
+
| `reason` | string, optional | Fail-open diagnostic (existing precedent — internal-error text, not a rule ID; unchanged by #859) |
|
|
585
|
+
|
|
586
|
+
- **Emitter:** `hooks/refactor_test_freeze_guard.py::audit()`, `hooks/refactor_test_revert_guard.py` and `hooks/refactor_test_bash_guard.py` (both via the same `audit()` import).
|
|
587
|
+
- **Consent:** unconditional (fails open, audits itself).
|
|
588
|
+
- **Consumers:** none automated yet; inspected manually when the freeze invariant is investigated.
|
|
589
|
+
|
|
590
|
+
---
|
|
591
|
+
|
|
592
|
+
## `contract-version-guard-audit.jsonl`
|
|
593
|
+
|
|
594
|
+
Audit log for release-please's bypass of the security-primitives-contract
|
|
595
|
+
version-bump requirement.
|
|
596
|
+
|
|
597
|
+
| Field | Type | Values / source |
|
|
598
|
+
| --- | --- | --- |
|
|
599
|
+
| `ts` | string | ISO-8601 UTC |
|
|
600
|
+
| `bypass` | boolean | Always `true` |
|
|
601
|
+
| `reason` | string | Always `release-please-actor` |
|
|
602
|
+
| `github_actor` | string | `$GITHUB_ACTOR` env value |
|
|
603
|
+
| `git_email` | string | `$GIT_AUTHOR_EMAIL` env value |
|
|
604
|
+
|
|
605
|
+
- **Emitter:** `hooks/contract_version_guard.py::_log_bypass()`.
|
|
606
|
+
- **Consent:** unconditional.
|
|
607
|
+
- **Consumers:** none automated yet; CI-only diagnostic trail.
|
|
608
|
+
|
|
609
|
+
---
|
|
610
|
+
|
|
611
|
+
## `learning-loop-state.json`
|
|
612
|
+
|
|
613
|
+
Not JSONL — a single current-value JSON file: a counter gating when
|
|
614
|
+
`session_learning_trigger.py` dispatches background session analysis.
|
|
615
|
+
|
|
616
|
+
| Field | Type | Values / source |
|
|
617
|
+
|---|---|---|
|
|
618
|
+
| `counter` | integer | Turns since the last dispatch |
|
|
619
|
+
|
|
620
|
+
- **Emitter:** `hooks/session_learning_trigger.py::_write_state()`.
|
|
621
|
+
- **Consent:** unconditional (internal scheduling state, no content).
|
|
622
|
+
- **Consumers:** `hooks/session_learning_trigger.py` itself (read on next fire).
|
|
623
|
+
|
|
624
|
+
---
|
|
625
|
+
|
|
626
|
+
## `pending-review.jsonl`
|
|
627
|
+
|
|
628
|
+
Queued findings from the background session-analysis dispatch, before
|
|
629
|
+
`/session-review` consumes them.
|
|
630
|
+
|
|
631
|
+
| Field | Type | Values / source |
|
|
632
|
+
| --- | --- | --- |
|
|
633
|
+
| `queued_at` | string | ISO-8601 UTC |
|
|
634
|
+
| `source` | string | Always `session-learning-trigger` |
|
|
635
|
+
| `session_id` | string, optional | Session ID when available |
|
|
636
|
+
| `findings` | array of object | Each: `lever`, `evidence`, `target_artifact`, `proposed_change`, `route` |
|
|
637
|
+
|
|
638
|
+
- **Emitter:** background `claude --print` run dispatched by `hooks/session_learning_trigger.py::_dispatch_background_analysis()`, writing via `session-analysis` agent output.
|
|
639
|
+
- **Consent:** unconditional (dispatch happens automatically; content is model-authored analysis, not raw session data).
|
|
640
|
+
- **Consumers:** `/session-review` skill.
|
|
641
|
+
|
|
642
|
+
---
|
|
643
|
+
|
|
644
|
+
## `.claude/metrics/{date}-task-log.jsonl` (e.g. `2026-02-20-task-log.jsonl`)
|
|
645
|
+
|
|
646
|
+
Self-reported per-task completion log, one file per calendar date.
|
|
647
|
+
|
|
648
|
+
| Field | Type | Values / source |
|
|
649
|
+
| --- | --- | --- |
|
|
650
|
+
| `timestamp` | string | ISO-8601 |
|
|
651
|
+
| (task-specific fields) | — | Tokens, cost, agents used, rework cycles, hallucination events — see `skills/performance-metrics/SKILL.md` for the full field list |
|
|
652
|
+
|
|
653
|
+
- **Emitter:** `/performance-metrics` skill (model-authored append at task completion), via `hooks/task_completion_metrics.py`.
|
|
654
|
+
- **Consent:** gated by `telemetry_consent.is_enabled()` (`~/.claude/telemetry.json` `{"enabled": true}`, home-scoped) — no longer unconditional as of Slice 2 (#1406).
|
|
655
|
+
- **Consumers:** `skills/harness-audit/SKILL.md` (self-reported half of the harness-audit join, alongside `session-digest.jsonl`'s real-session half), `skills/governance-compliance/SKILL.md`.
|
|
656
|
+
|
|
657
|
+
---
|
|
658
|
+
|
|
659
|
+
## `gherkin-derive-effectiveness.jsonl`
|
|
660
|
+
|
|
661
|
+
Per-scenario roll-up correlating a `/gherkin-derive`-discovered surface with
|
|
662
|
+
whatever coverage/mutation-delta data the calling workflow already measured,
|
|
663
|
+
so there is a signal on whether BDD-derived scenarios track real
|
|
664
|
+
coverage/mutation movement (issue #1296). One record per scenario per
|
|
665
|
+
roll-up run — not deduplicated across runs, since coverage/mutation deltas
|
|
666
|
+
are re-measured every convergence iteration.
|
|
667
|
+
|
|
668
|
+
| Field | Type | Values / source |
|
|
669
|
+
| --- | --- | --- |
|
|
670
|
+
| `surface` | string, nullable | The discovered surface name/path from `gherkin.md`'s surface-inventory table |
|
|
671
|
+
| `discovery_source` | string, nullable | `openapi` \| `route` \| `test` \| `signature` (per `/gherkin-derive` Step 2), as recorded in the inventory |
|
|
672
|
+
| `provenance` | string, nullable | `specification` \| `characterization`, as recorded in the inventory |
|
|
673
|
+
| `binding_mode` | string, nullable | `none` \| `xunit-with-annotations` \| `bdd-runner` |
|
|
674
|
+
| `bound_story` | number or string, nullable | The Story/issue id from `gherkin-bindings.json`, when that file exists for the run (only produced by `/gherkin-public`) |
|
|
675
|
+
| `coverage_delta` | object, nullable | `{line_pct, branch_pct}` — workflow-level delta between the two coverage snapshots passed to the roll-up, not an isolated per-scenario attribution (no finer-grained mapping exists today) |
|
|
676
|
+
| `mutation_delta` | object, nullable | `{survivors_after_delta}` — workflow-level survivor-count delta, same caveat as `coverage_delta` |
|
|
677
|
+
|
|
678
|
+
- **Emitter:** `plugins/dev-team/scripts/gherkin_effectiveness_rollup.py`, invoked from `/quality-targets-converge` Step 6b after each convergence iteration's re-measure, when `gherkin.md` exists for the workflow slug.
|
|
679
|
+
- **Consent:** unconditional (derived metrics only; no prompt/file-content capture).
|
|
680
|
+
- **Consumers:** none yet — this is the roll-up a future `/harness-audit`-style review reads to compare BDD-derived vs. hand-written test effectiveness.
|
|
681
|
+
|
|
682
|
+
---
|
|
683
|
+
|
|
684
|
+
## Benefit-measurement streams (#2201)
|
|
685
|
+
|
|
686
|
+
**Added by #2201** (epic #2200, slice 0). Four observational JSONL streams that
|
|
687
|
+
make post-merge benefit numbers for epics #2164 and #2172 exist. Each row is
|
|
688
|
+
written fail-open by `hooks/lib/instrument_log.py` (never affects stdout, exit
|
|
689
|
+
code, or control flow) and carries `ts`, `plugin_version`, and, when the
|
|
690
|
+
emitter has one, `session_id`. They are **separate from `boundary-events.jsonl`**
|
|
691
|
+
on purpose: a `record` row there is read by the review-gate corroboration path,
|
|
692
|
+
so measurement rows must not share it. Counts, enums and lens/agent names only.
|
|
693
|
+
|
|
694
|
+
| Stream | Emitter | Fields | Answers |
|
|
695
|
+
|---|---|---|---|
|
|
696
|
+
| `subagent-stops.jsonl` | `hooks/subagent_completion_guard.py` (every `SubagentStop`) | `classification` (`clean` \| `empty-final-turn` \| `truncated-final-turn` \| `unreadable`) | The completion-guard divergence rate's denominator. `boundary-events.jsonl` only carries the two non-clean classes, so a rate was not computable before. |
|
|
697
|
+
| `skill-injection.jsonl` | `hooks/subagent_skill_context.py` (when a hint is injected) | `agent_type`, `skills` (list), `added_chars` | Injection overhead (`added_chars`) and the denominator for uptake; uptake itself is read from `Skill` tool calls in the subagent transcripts (`scripts/lib/session_log`). |
|
|
698
|
+
| `ledger-skips.jsonl` | `scripts/verdict_scope.py` (every CLI consult) | `candidate_pairs`, `skipped_pairs`, `fully_skipped_lenses` (list) | Realized delta-scoping skip rate = `skipped_pairs / candidate_pairs`. Previously only printed to stdout. |
|
|
699
|
+
| `checkpoint-aborts.jsonl` | `scripts/checkpoint_abort.py` | `mode: "abort"`: `aborted`, `triggering_agent`, `deferred_lenses`. `mode: "outcome"`: `aborted`, `redispatched`, `findings`, `blocking_findings`, `outcome` | Abort frequency, deferred-lens yield (`outcome` rows with `aborted` and `redispatched`), previously only printed. |
|
|
700
|
+
|
|
701
|
+
### Instrument audit (#2201)
|
|
702
|
+
|
|
703
|
+
Static audit of each instrument; the per-session confirmation the issue also
|
|
704
|
+
asks for (≥ 3 real sessions, IDs listed on #2200) must be run on the
|
|
705
|
+
maintainer's machine after these emitters ship.
|
|
706
|
+
|
|
707
|
+
| Instrument | Finding | Action |
|
|
708
|
+
|---|---|---|
|
|
709
|
+
| `review-verdicts.jsonl` | Rows are written only when the dispatch prompt carries the scope marker (`review_verdict_recorder.py`); a dispatch without it emits a `boundary-events.jsonl` `record` row with `matched_rule: "missing-scope-marker"`. | None; count those `boundary-events.jsonl` rows as the ledger-coverage gap. |
|
|
710
|
+
| `ledgerSkipped` / `fullySkippedLenses` | Returned on stdout only. | New `ledger-skips.jsonl`. |
|
|
711
|
+
| `checkpoint_abort.py` | Outcomes printed only. | New `checkpoint-aborts.jsonl`. |
|
|
712
|
+
| `subagent_skill_context.py` | No signal of injection or of a skill being loaded. | New `skill-injection.jsonl`; "loaded" is derived from subagent transcripts. |
|
|
713
|
+
| `subagent_completion_guard.py` | Emits `empty-final-turn` / `truncated-final-turn` to `boundary-events.jsonl` (`decision: "record"`); `clean`/`unreadable` silent. | New `subagent-stops.jsonl` with every classification. |
|
|
714
|
+
|
|
715
|
+
Consent gating: none beyond the existing per-project `.claude/metrics/`
|
|
716
|
+
location; rows are local files and are never transmitted.
|
|
717
|
+
|
|
718
|
+
---
|
|
719
|
+
|
|
720
|
+
## Adding a new stream
|
|
721
|
+
|
|
722
|
+
1. Name it `.claude/metrics/<name>.jsonl` (or `.json` for a single-current-value
|
|
723
|
+
file) — one stream per concern, matching existing precedent.
|
|
724
|
+
2. Append-only, compact JSON (`separators=(",", ":")`) + trailing newline for
|
|
725
|
+
JSONL streams.
|
|
726
|
+
3. Rule IDs / counts / enums only — never command text, prompt text, file
|
|
727
|
+
contents, or incidental free text.
|
|
728
|
+
4. Add a section to this file with the same shape as the ones above
|
|
729
|
+
(fields/types, emitter, consent gating, consumers) in the same PR that
|
|
730
|
+
introduces the emitter — the coverage test in
|
|
731
|
+
`tests/hooks/test_boundary_events.py` enforces this.
|
|
732
|
+
|
|
733
|
+
---
|
|
734
|
+
|
|
735
|
+
## `autoship-log.jsonl`
|
|
736
|
+
|
|
737
|
+
One record per `/autoship` dispatch-unit outcome (a solo issue or a batch),
|
|
738
|
+
plus one `round_summary` event per round. Every record is appended via the
|
|
739
|
+
shared `hooks/lib/autoship_log.py` appender (its `--json`/`--json-file`
|
|
740
|
+
CLI), which stamps `logged_at` regardless of which of the three shapes
|
|
741
|
+
below the caller passes it — the library itself is schema-agnostic; the
|
|
742
|
+
shape is entirely determined by `/autoship`'s own SKILL.md (Steps 3f/4).
|
|
743
|
+
|
|
744
|
+
**Solo entry** — one record per solo-dispatched issue:
|
|
745
|
+
|
|
746
|
+
| Field | Type | Values / source |
|
|
747
|
+
| --- | --- | --- |
|
|
748
|
+
| `logged_at` | string | ISO-8601 (UTC) — stamped by `autoship_log.py` |
|
|
749
|
+
| `round_id` | string | ISO-8601 timestamp generated once at round start (before Step 1) |
|
|
750
|
+
| `issue` | integer | The dispatched issue number |
|
|
751
|
+
| `status` | string enum | `shipped` \| `failed` \| `unrecognized` \| `blocked` |
|
|
752
|
+
| `blocked_reason` | string, nullable | The extracted stakeholder question (`blocked`), the Step-3d.1-synthesized classifier-verdict string (`failed`/`unrecognized`), or `null` for `shipped` |
|
|
753
|
+
|
|
754
|
+
**Batch entry** — one record per batch dispatch unit, never one per member
|
|
755
|
+
issue, applying identically across every outcome (`shipped`, `blocked`,
|
|
756
|
+
`failed`, and `unrecognized` alike):
|
|
757
|
+
|
|
758
|
+
| Field | Type | Values / source |
|
|
759
|
+
| --- | --- | --- |
|
|
760
|
+
| `logged_at` | string | ISO-8601 (UTC) — stamped by `autoship_log.py` |
|
|
761
|
+
| `round_id` | string | Same round-start timestamp as any solo entry logged the same round |
|
|
762
|
+
| `batch_id` | string | e.g. `grp-101` — `autoship_group.py`'s deterministic batch id |
|
|
763
|
+
| `issues` | array of integer | Every member issue number |
|
|
764
|
+
| `status` | string enum | `shipped` \| `failed` \| `unrecognized` \| `blocked` |
|
|
765
|
+
| `blocked_reason` | string, nullable | Same convention as the solo entry's field, applied to the whole batch |
|
|
766
|
+
|
|
767
|
+
**`round_summary` event** — one per round, appended after Step 3's
|
|
768
|
+
per-dispatch-unit loop ends (skipped in `--dry-run`):
|
|
769
|
+
|
|
770
|
+
| Field | Type | Values / source |
|
|
771
|
+
| --- | --- | --- |
|
|
772
|
+
| `logged_at` | string | ISO-8601 (UTC) — stamped by `autoship_log.py` |
|
|
773
|
+
| `round_id` | string | Same round-start timestamp as this round's solo/batch entries |
|
|
774
|
+
| `event` | string | Always `"round_summary"` — distinguishes this record from a solo/batch entry above |
|
|
775
|
+
| `processed_units` / `processed_issues` | integer | Dispatch units Step 3c actually dispatched this round / their total member-issue count |
|
|
776
|
+
| `discovered_units` / `discovered_issues` | integer | Every dispatch unit `autoship_queue.py` produced this round (`queue` + `deferred` combined) / their total member-issue count |
|
|
777
|
+
| `deferred_units` / `deferred_issues` | integer | Dispatch units left in `deferred` / their total member-issue count (a solo unit counts as 1) |
|
|
778
|
+
| `blocked_pending_confirmation_units` / `blocked_pending_confirmation_issues` | integer | Proposed batches Step 2c actually blocked pending human confirmation this round / their total member-issue count — always present, `0` when Step 2b/2c never ran or blocked nothing |
|
|
779
|
+
| `cost_usd` | number | Accumulated round cost |
|
|
780
|
+
| `status` | string enum | `complete` \| `cost_cap_reached` \| `dry_run` \| `no_eligible_issues` \| `no_unit_fits_cap` \| `blocked_pending_confirmation` |
|
|
781
|
+
|
|
782
|
+
- **Emitter:** `hooks/lib/autoship_log.py` called from the `/autoship` skill (Step 3f for solo/batch entries, Step 4 for the `round_summary` event).
|
|
783
|
+
- **Consent:** unconditional (cost/count aggregates and enum values only — no prompt text or file contents).
|
|
784
|
+
- **Consumers:** `/cost-report`, `/telemetry` (aggregate reporting).
|
|
785
|
+
|
|
786
|
+
---
|
|
787
|
+
|
|
788
|
+
## `workflow-states.jsonl`
|
|
789
|
+
|
|
790
|
+
**Added by #1166.** Event-sourced workflow lifecycle stream for orchestrated
|
|
791
|
+
flows (`/ship`, `/autoship`, `/build`): persists only state-*transition*
|
|
792
|
+
events. Current state and per-state dwell time are always **derived** by
|
|
793
|
+
replaying the stream for a given `session_id` — never stored — per the
|
|
794
|
+
event-sourcing discipline in the competitive analysis this issue is drawn
|
|
795
|
+
from. Canonical (informational, not enforced) lifecycle: `SPEC -> PLAN ->
|
|
796
|
+
BUILD -> REVIEW -> COMMIT -> PR`.
|
|
797
|
+
|
|
798
|
+
| Field | Type | Values / source |
|
|
799
|
+
| --- | --- | --- |
|
|
800
|
+
| `ts` | string | ISO-8601 UTC `%Y-%m-%dT%H:%M:%SZ` |
|
|
801
|
+
| `workflow` | string | Orchestrated flow name, e.g. `ship`, `autoship`, `build` |
|
|
802
|
+
| `prior_state` | string, optional (`null` for the initial transition) | State the workflow was in before this transition |
|
|
803
|
+
| `new_state` | string | State the workflow is entering |
|
|
804
|
+
| `plugin_version` | string | From `.claude-plugin/plugin.json` |
|
|
805
|
+
| `session_id` | string, optional | Opaque per-session ID — enables joins with `boundary-events.jsonl` and `cost-metering.jsonl` |
|
|
806
|
+
|
|
807
|
+
- **Emitter:** `hooks/lib/workflow_state.py::emit_state_transition()`, invoked via its `record` CLI subcommand as a model-authored append at each phase boundary in `/ship`, `/autoship`, and `/build` (same convention as `review-value.jsonl`/`verify-log.jsonl`).
|
|
808
|
+
- **Consent:** unconditional (workflow/state names + counts only — no prompt text or file contents).
|
|
809
|
+
- **Derivation:** `hooks/lib/workflow_state.py::derive_current_state()` and `compute_dwell_times()` (also exposed via the `report` CLI subcommand) replay a session's transitions — never a stored snapshot.
|
|
810
|
+
- **Consumers:** `skills/run-report/SKILL.md` (#1167), `skills/session-review/SKILL.md`, `skills/harness-audit/SKILL.md`, `skills/cost-report/SKILL.md`.
|
|
811
|
+
|
|
812
|
+
---
|
|
813
|
+
|
|
814
|
+
## `iteration-journal.jsonl`
|
|
815
|
+
|
|
816
|
+
**Added by #1168.** Hard per-iteration decision journal for the autonomous
|
|
817
|
+
`/autoship`/`/ship` loops: one entry per round/iteration recording what was
|
|
818
|
+
attempted, its outcome, and the next action — the accountability record an
|
|
819
|
+
autonomous run needs to be debuggable after the fact. Unlike
|
|
820
|
+
`workflow-states.jsonl`'s phase transitions, this stream is not derived; each
|
|
821
|
+
entry is a durable, once-written decision note.
|
|
822
|
+
|
|
823
|
+
| Field | Type | Values / source |
|
|
824
|
+
| --- | --- | --- |
|
|
825
|
+
| `ts` | string | ISO-8601 UTC `%Y-%m-%dT%H:%M:%SZ` |
|
|
826
|
+
| `round_id` | string | Identifier for the current round/iteration (`/autoship`'s round_id, or `/ship`'s issue identifier) |
|
|
827
|
+
| `attempted` | string | Short structured note — what was attempted this iteration (deliberate, agent-authored rationale, not incidental free text — same precedent as `config-changelog.jsonl`'s `description`) |
|
|
828
|
+
| `outcome` | string | Short structured note — what happened |
|
|
829
|
+
| `next_action` | string | Short structured note — what happens next |
|
|
830
|
+
| `plugin_version` | string | From `.claude-plugin/plugin.json` |
|
|
831
|
+
| `session_id` | string, optional | Opaque per-session ID — enables joins with `boundary-events.jsonl` / `cost-metering.jsonl` |
|
|
832
|
+
|
|
833
|
+
- **Emitter:** `hooks/lib/iteration_journal_gate.py::record_iteration_entry()`, invoked via its `record` CLI subcommand as a model-authored append in `/autoship`'s per-issue loop (Step 3) and `/ship`'s per-phase loop, before the corresponding `check` subcommand gates advancement.
|
|
834
|
+
- **Gate:** `hooks/lib/iteration_journal_gate.py::check_iteration_journal()` (`check` CLI subcommand) hard-blocks advancement to the next issue/iteration — exit 1 — unless >=1 entry exists for the current `round_id`; a block also emits a `boundary-events.jsonl` event (`hook: iteration_journal_gate`, `decision: block`, `matched_rule: iteration-journal-missing`). This is a skill-level check-before-advance (mirroring `verify-log.jsonl`'s `progress_guardian.py --pre-pr` pattern), not a `settings.json` PreToolUse/PostToolUse registration — `/autoship`'s and `/ship`'s loop advancement is model-authored control flow inside a skill, not a tool call the harness intercepts at a distinct boundary. Complements, does not replace, the advisory plan-step-keyed `progress-guardian` agent.
|
|
835
|
+
- **Consent:** unconditional (a deliberate per-iteration accountability record, not passive usage telemetry).
|
|
836
|
+
- **Consumers:** `skills/autoship/SKILL.md`, `skills/ship/SKILL.md`, joinable with `skills/run-report/SKILL.md` (#1167) via `round_id`/`session_id`.
|
|
837
|
+
|
|
838
|
+
---
|
|
839
|
+
|
|
840
|
+
## `xunit-v3-shim-decisions.json`
|
|
841
|
+
|
|
842
|
+
**Added by #1791.** Not JSONL — a single current-value JSON object keyed by test
|
|
843
|
+
project name, holding the operator's chosen remediation when xunit.v3
|
|
844
|
+
constructs block the Stryker v2 shim. This is the enforcement record, not
|
|
845
|
+
telemetry: `stryker_xunit_shim_guard.py` blocks every `dotnet-stryker` run
|
|
846
|
+
against a blocked project until an entry covering the current blocker set
|
|
847
|
+
exists, which is what makes the always-ask gate a guarantee rather than hook
|
|
848
|
+
stdout an agent may paraphrase or skip.
|
|
849
|
+
|
|
850
|
+
Each value:
|
|
851
|
+
|
|
852
|
+
| Field | Type | Values / source |
|
|
853
|
+
| --- | --- | --- |
|
|
854
|
+
| `project` | string | Test project name (the real test `.csproj` stem) |
|
|
855
|
+
| `choice` | string | `port` \| `exclude` \| `skip` \| `degrade` — the four documented remediations; any other value is rejected at write time |
|
|
856
|
+
| `fingerprint` | string, required | 16-hex digest over the blocker set's `file::construct` pairs (line numbers deliberately excluded). Scopes the decision to the blockers the operator actually saw; a mismatch — or an absent value, which would make the entry a blanket answer — re-asks |
|
|
857
|
+
| `files` | array of string | Flagged files the choice covers, project-relative |
|
|
858
|
+
| `note` | string, nullable | Operator rationale, when given |
|
|
859
|
+
| `recorded_at` | string | ISO-8601 UTC `%Y-%m-%dT%H:%M:%SZ` |
|
|
860
|
+
|
|
861
|
+
- **Emitter:** `hooks/lib/xunit_v3_operator_gate.py::record_decision()`, invoked
|
|
862
|
+
via its `record` CLI subcommand after the operator answers the gate.
|
|
863
|
+
- **Gate:** `hooks/lib/xunit_v3_operator_gate.py::decision_for()`, read by
|
|
864
|
+
`hooks/stryker_xunit_shim_guard.py` (PreToolUse on `Bash`). No covering entry
|
|
865
|
+
→ exit 2 with the operator question as the block body. Fails closed on every
|
|
866
|
+
axis: a fingerprint mismatch, an absent fingerprint, and a stored `choice`
|
|
867
|
+
outside the four all re-ask rather than letting a run proceed unasked.
|
|
868
|
+
- **Consent:** unconditional (an explicit operator decision record, not passive
|
|
869
|
+
usage telemetry).
|
|
870
|
+
- **Consumers:** `hooks/stryker_xunit_shim_guard.py`,
|
|
871
|
+
`skills/mutation-testing/scripts/mutation_feasibility_gate.py` (same question
|
|
872
|
+
payload), `skills/stryker-xunit-v2-shim/SKILL.md` Step 1a. No path override
|
|
873
|
+
(#1870 dropped `DEV_TEAM_XUNIT3_SHIM_DECISION_FILE` entirely — no legitimate
|
|
874
|
+
caller needs runtime relocation of the store).
|
|
875
|
+
- **Audit trail (#1870):** every `record_decision()` write and every
|
|
876
|
+
`decision_for()` honor (a stored decision covering the current question,
|
|
877
|
+
about to drive the guard's outcome) also emits a `boundary-events.jsonl`
|
|
878
|
+
entry — `matched_rule` of `xunit-v3-shim-decision-record-<choice>` or
|
|
879
|
+
`xunit-v3-shim-decision-honor-<choice>`, `subject_hash` bound to the
|
|
880
|
+
question's `fingerprint` — so a self-recorded choice is visible in the same
|
|
881
|
+
stream the review-gate corroboration mechanism is audited from.
|