pi-dev-team 0.2.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (780) hide show
  1. package/LICENSE +21 -0
  2. package/PORTING.md +134 -0
  3. package/README.md +207 -0
  4. package/UPSTREAM.json +64 -0
  5. package/agents/Explore.md +15 -0
  6. package/agents/a11y-review.md +118 -0
  7. package/agents/adr-author.md +70 -0
  8. package/agents/ai-provenance-review.md +120 -0
  9. package/agents/angular-reactivity-review.md +95 -0
  10. package/agents/arch-review.md +135 -0
  11. package/agents/architect.md +78 -0
  12. package/agents/autoship-batch-proposer.md +69 -0
  13. package/agents/claude-setup-review.md +136 -0
  14. package/agents/codebase-recon.md +184 -0
  15. package/agents/component-architecture-review.md +119 -0
  16. package/agents/concurrency-review.md +109 -0
  17. package/agents/correctness-review.md +290 -0
  18. package/agents/data-flow-tracer.md +120 -0
  19. package/agents/doc-review.md +165 -0
  20. package/agents/domain-review.md +136 -0
  21. package/agents/general-purpose.md +10 -0
  22. package/agents/gherkin-quality-critic.md +113 -0
  23. package/agents/js-fp-review.md +114 -0
  24. package/agents/mutation-kill.md +684 -0
  25. package/agents/naming-review.md +142 -0
  26. package/agents/orchestrator.md +339 -0
  27. package/agents/performance-review.md +105 -0
  28. package/agents/plan-review-acceptance.md +115 -0
  29. package/agents/plan-review-design.md +90 -0
  30. package/agents/plan-review-parallelization.md +84 -0
  31. package/agents/plan-review-strategic.md +96 -0
  32. package/agents/plan-review-ux.md +110 -0
  33. package/agents/platform-engineer.md +64 -0
  34. package/agents/product-manager.md +68 -0
  35. package/agents/progress-guardian.md +79 -0
  36. package/agents/qa-engineer.md +289 -0
  37. package/agents/quality-reviewer.md +132 -0
  38. package/agents/react-reactivity-review.md +102 -0
  39. package/agents/refactor-opportunity-review.md +128 -0
  40. package/agents/security-engineer.md +60 -0
  41. package/agents/security-review.md +218 -0
  42. package/agents/session-analysis.md +95 -0
  43. package/agents/software-engineer.md +105 -0
  44. package/agents/spec-compliance-review.md +100 -0
  45. package/agents/spec-reviewer.md +114 -0
  46. package/agents/structure-review.md +146 -0
  47. package/agents/tech-writer.md +84 -0
  48. package/agents/test-review.md +246 -0
  49. package/agents/test-smell-review.md +188 -0
  50. package/agents/token-efficiency-review.md +139 -0
  51. package/agents/ui-ux-designer.md +54 -0
  52. package/agents/vue-reactivity-review.md +95 -0
  53. package/bin/__pycache__/claudecpython-314.pyc +0 -0
  54. package/bin/claude +258 -0
  55. package/docs/upstream/.pages +1 -0
  56. package/docs/upstream/CHANGELOG.md +2586 -0
  57. package/docs/upstream/README.md +155 -0
  58. package/docs/upstream/agent-architecture.md +214 -0
  59. package/docs/upstream/agent_info.md +187 -0
  60. package/docs/upstream/artifact-migration.md +124 -0
  61. package/docs/upstream/code-intelligence-nudge.md +149 -0
  62. package/docs/upstream/code-review-process.md +294 -0
  63. package/docs/upstream/concurrent-use.md +73 -0
  64. package/docs/upstream/context-management.md +111 -0
  65. package/docs/upstream/developer-notes.md +280 -0
  66. package/docs/upstream/diagrams/architecture-overview.svg +101 -0
  67. package/docs/upstream/diagrams/review-dispatch.svg +139 -0
  68. package/docs/upstream/diagrams/team-agents.svg +128 -0
  69. package/docs/upstream/diagrams/test-improve-flow.svg +166 -0
  70. package/docs/upstream/diagrams/workflow-linear.svg +66 -0
  71. package/docs/upstream/diagrams/workflow-three-phase.svg +200 -0
  72. package/docs/upstream/eval-maintenance.md +95 -0
  73. package/docs/upstream/eval-running-guide.md +147 -0
  74. package/docs/upstream/eval-system.md +291 -0
  75. package/docs/upstream/session-review-oss-complements.md +75 -0
  76. package/docs/upstream/session-review.md +212 -0
  77. package/docs/upstream/skills.md +188 -0
  78. package/docs/upstream/team-structure.md +21 -0
  79. package/docs/upstream/telemetry-ci-access.md +129 -0
  80. package/docs/upstream/telemetry-repo-security.md +120 -0
  81. package/docs/upstream/test-evaluation.md +277 -0
  82. package/docs/upstream/test-improve.md +154 -0
  83. package/docs/upstream/triage-workflow.md +282 -0
  84. package/docs/upstream/workflows.md +289 -0
  85. package/extensions/dev-team/index.ts +539 -0
  86. package/extensions/dev-team/lib/agents.ts +272 -0
  87. package/extensions/dev-team/lib/ai-credits.ts +92 -0
  88. package/extensions/dev-team/lib/autocompact.ts +81 -0
  89. package/extensions/dev-team/lib/child-run.ts +102 -0
  90. package/extensions/dev-team/lib/config.ts +236 -0
  91. package/extensions/dev-team/lib/gh-command.ts +103 -0
  92. package/extensions/dev-team/lib/github-style.ts +307 -0
  93. package/extensions/dev-team/lib/hooks.ts +350 -0
  94. package/extensions/dev-team/lib/metrics.ts +115 -0
  95. package/extensions/dev-team/lib/safe-read.ts +49 -0
  96. package/extensions/dev-team/lib/session-files.ts +57 -0
  97. package/extensions/dev-team/lib/session-spend.ts +123 -0
  98. package/extensions/dev-team/lib/shell-scan.ts +205 -0
  99. package/extensions/dev-team/lib/skills.ts +213 -0
  100. package/extensions/dev-team/lib/subagent-render.ts +245 -0
  101. package/extensions/dev-team/lib/subagent-types.ts +164 -0
  102. package/extensions/dev-team/lib/subagent.ts +596 -0
  103. package/extensions/dev-team/lib/terminal-text.ts +54 -0
  104. package/extensions/dev-team/lib/tools-misc.ts +152 -0
  105. package/extensions/dev-team/lib/transcript.ts +110 -0
  106. package/extensions/dev-team/lib/trust.ts +52 -0
  107. package/extensions/dev-team/lib/usage-breakdown.ts +176 -0
  108. package/extensions/dev-team/lib/usage-chart.ts +153 -0
  109. package/extensions/dev-team/lib/usage-command.ts +107 -0
  110. package/extensions/dev-team/lib/usage-history.ts +203 -0
  111. package/extensions/dev-team/lib/usage-render.ts +225 -0
  112. package/extensions/dev-team/lib/usage-split-bar.ts +127 -0
  113. package/extensions/dev-team/lib/usage-state.ts +116 -0
  114. package/extensions/dev-team/lib/usage-text.ts +159 -0
  115. package/extensions/dev-team/lib/usage-view.ts +109 -0
  116. package/hooks/__pycache__/refactor_test_freeze_guard.cpython-314.pyc +0 -0
  117. package/hooks/agent_dispatch_ledger.py +190 -0
  118. package/hooks/autocompact_setup_nudge.py +99 -0
  119. package/hooks/bash_retry_guard.py +228 -0
  120. package/hooks/boundary_events_write_guard.py +352 -0
  121. package/hooks/code_intelligence_nudge.py +293 -0
  122. package/hooks/code_intelligence_turn_mark.py +317 -0
  123. package/hooks/codegraph_bootstrap.py +139 -0
  124. package/hooks/contract_version_guard.py +362 -0
  125. package/hooks/cost_meter.py +106 -0
  126. package/hooks/destructive-commands.json +62 -0
  127. package/hooks/destructive_guard.py +477 -0
  128. package/hooks/eval_compliance_check.py +440 -0
  129. package/hooks/guards.json +17 -0
  130. package/hooks/hooks.json +323 -0
  131. package/hooks/internal_double_gate.py +296 -0
  132. package/hooks/js_fp_review.py +212 -0
  133. package/hooks/knowledge_index.py +119 -0
  134. package/hooks/lib/__pycache__/artifact_paths.cpython-314.pyc +0 -0
  135. package/hooks/lib/__pycache__/atomic_state.cpython-314.pyc +0 -0
  136. package/hooks/lib/__pycache__/autocompact_config.cpython-314.pyc +0 -0
  137. package/hooks/lib/__pycache__/boundary_events.cpython-314.pyc +0 -0
  138. package/hooks/lib/__pycache__/doc_classification.cpython-314.pyc +0 -0
  139. package/hooks/lib/__pycache__/gh_pr_create_detect.cpython-314.pyc +0 -0
  140. package/hooks/lib/__pycache__/git_safe_diff.cpython-314.pyc +0 -0
  141. package/hooks/lib/__pycache__/instrument_log.cpython-314.pyc +0 -0
  142. package/hooks/lib/__pycache__/metrics_query.cpython-314.pyc +0 -0
  143. package/hooks/lib/__pycache__/plugin_version.cpython-314.pyc +0 -0
  144. package/hooks/lib/__pycache__/pre_commit_doc_classifier.cpython-314.pyc +0 -0
  145. package/hooks/lib/__pycache__/review_agent_registry.cpython-314.pyc +0 -0
  146. package/hooks/lib/__pycache__/review_gate_corroboration.cpython-314.pyc +0 -0
  147. package/hooks/lib/__pycache__/review_gate_hash.cpython-314.pyc +0 -0
  148. package/hooks/lib/__pycache__/review_verdicts.cpython-314.pyc +0 -0
  149. package/hooks/lib/__pycache__/stdin_json.cpython-314.pyc +0 -0
  150. package/hooks/lib/__pycache__/stryker_invocation.cpython-314.pyc +0 -0
  151. package/hooks/lib/__pycache__/telemetry_consent.cpython-314.pyc +0 -0
  152. package/hooks/lib/__pycache__/test_file_classify.cpython-314.pyc +0 -0
  153. package/hooks/lib/__pycache__/token_efficiency_limits.cpython-314.pyc +0 -0
  154. package/hooks/lib/__pycache__/verify_guard_state.cpython-314.pyc +0 -0
  155. package/hooks/lib/__pycache__/xunit_v3_operator_gate.cpython-314.pyc +0 -0
  156. package/hooks/lib/agent_skill_hints.py +74 -0
  157. package/hooks/lib/artifact_paths.py +263 -0
  158. package/hooks/lib/atomic_state.py +557 -0
  159. package/hooks/lib/autocompact_config.py +103 -0
  160. package/hooks/lib/autoship_log.py +106 -0
  161. package/hooks/lib/banned_scripts_policy.py +51 -0
  162. package/hooks/lib/boundary_events.py +436 -0
  163. package/hooks/lib/build_knowledge_index.py +504 -0
  164. package/hooks/lib/build_skills_index.py +361 -0
  165. package/hooks/lib/build_state.py +116 -0
  166. package/hooks/lib/classify_ship_outcome.py +126 -0
  167. package/hooks/lib/config_changelog_schema.py +115 -0
  168. package/hooks/lib/cost_meter.py +955 -0
  169. package/hooks/lib/doc_classification.py +116 -0
  170. package/hooks/lib/gh_pr_create_detect.py +136 -0
  171. package/hooks/lib/git_safe_diff.py +123 -0
  172. package/hooks/lib/instrument_log.py +66 -0
  173. package/hooks/lib/iteration_journal_gate.py +197 -0
  174. package/hooks/lib/knowledge_index_paths.py +88 -0
  175. package/hooks/lib/mcp_json_repowise.py +177 -0
  176. package/hooks/lib/metrics_query.py +202 -0
  177. package/hooks/lib/minimal_yaml.py +434 -0
  178. package/hooks/lib/plugin_version.py +142 -0
  179. package/hooks/lib/pre_commit_detect.py +537 -0
  180. package/hooks/lib/pre_commit_doc_classifier.py +126 -0
  181. package/hooks/lib/pricing.py +118 -0
  182. package/hooks/lib/report_pdf.py +371 -0
  183. package/hooks/lib/review_agent_registry.py +142 -0
  184. package/hooks/lib/review_dispatch_ledger.py +101 -0
  185. package/hooks/lib/review_gate_corroboration.py +521 -0
  186. package/hooks/lib/review_gate_hash.py +252 -0
  187. package/hooks/lib/review_gate_normalized_hash.py +1115 -0
  188. package/hooks/lib/review_verdicts.py +301 -0
  189. package/hooks/lib/run_report.py +160 -0
  190. package/hooks/lib/skill_categories.yaml +125 -0
  191. package/hooks/lib/stdin_json.py +57 -0
  192. package/hooks/lib/stryker_invocation.py +102 -0
  193. package/hooks/lib/telemetry_consent.py +41 -0
  194. package/hooks/lib/telemetry_report.py +108 -0
  195. package/hooks/lib/test_file_classify.py +160 -0
  196. package/hooks/lib/token_efficiency_limits.py +51 -0
  197. package/hooks/lib/turn_identity.py +77 -0
  198. package/hooks/lib/verify_guard_state.py +110 -0
  199. package/hooks/lib/workflow_state.py +206 -0
  200. package/hooks/lib/xunit_v3_operator_gate.py +596 -0
  201. package/hooks/mcp_json_repowise_nudge.py +74 -0
  202. package/hooks/mutation_adapters/__init__.py +7 -0
  203. package/hooks/mutation_adapters/__pycache__/__init__.cpython-314.pyc +0 -0
  204. package/hooks/mutation_adapters/__pycache__/lib.cpython-314.pyc +0 -0
  205. package/hooks/mutation_adapters/__pycache__/mutmut.cpython-314.pyc +0 -0
  206. package/hooks/mutation_adapters/__pycache__/pitest.cpython-314.pyc +0 -0
  207. package/hooks/mutation_adapters/__pycache__/stryker.cpython-314.pyc +0 -0
  208. package/hooks/mutation_adapters/__pycache__/stryker_net.cpython-314.pyc +0 -0
  209. package/hooks/mutation_adapters/lib.py +478 -0
  210. package/hooks/mutation_adapters/mutmut.py +188 -0
  211. package/hooks/mutation_adapters/pitest.py +266 -0
  212. package/hooks/mutation_adapters/stryker.py +157 -0
  213. package/hooks/mutation_adapters/stryker_net.py +264 -0
  214. package/hooks/mutation_gate.py +193 -0
  215. package/hooks/mutation_testing_smoke_gate.py +371 -0
  216. package/hooks/pending_review_notify.py +121 -0
  217. package/hooks/phase_marker.py +138 -0
  218. package/hooks/post_compact_state_reinject.py +180 -0
  219. package/hooks/post_format.py +115 -0
  220. package/hooks/pre_commit_knowledge_index.py +128 -0
  221. package/hooks/pre_commit_review.py +66 -0
  222. package/hooks/pre_pr_review.py +694 -0
  223. package/hooks/pre_tool_guard.py +405 -0
  224. package/hooks/py.sh +73 -0
  225. package/hooks/refactor-bash-write-patterns.json +29 -0
  226. package/hooks/refactor_test_bash_guard.py +253 -0
  227. package/hooks/refactor_test_freeze_guard.py +139 -0
  228. package/hooks/refactor_test_revert_guard.py +186 -0
  229. package/hooks/repo_review_nudge.py +287 -0
  230. package/hooks/review_verdict_recorder.py +464 -0
  231. package/hooks/scan_bash_command_for_banned_scripts.py +428 -0
  232. package/hooks/scan_worktree_for_banned_scripts.py +238 -0
  233. package/hooks/session_learning_trigger.py +248 -0
  234. package/hooks/skills_index.py +126 -0
  235. package/hooks/stryker_xunit_shim_guard.py +571 -0
  236. package/hooks/subagent_completion_guard.py +309 -0
  237. package/hooks/subagent_skill_context.py +139 -0
  238. package/hooks/task_completion_metrics.py +216 -0
  239. package/hooks/tdd_guard.py +229 -0
  240. package/hooks/telemetry.py +341 -0
  241. package/hooks/token_efficiency_review.py +194 -0
  242. package/hooks/verify_guard.py +183 -0
  243. package/hooks/verify_guard_edit_marker.py +73 -0
  244. package/hooks/version_check.py +173 -0
  245. package/knowledge/accepted-risks-schema.md +98 -0
  246. package/knowledge/adr-decision-criteria.md +64 -0
  247. package/knowledge/adversarial-review-protocol.md +139 -0
  248. package/knowledge/agent-registry.md +228 -0
  249. package/knowledge/agent-review-methodology.md +80 -0
  250. package/knowledge/ai-friendly-repo-guidelines.md +67 -0
  251. package/knowledge/architecture-assessment.md +96 -0
  252. package/knowledge/artifact-lifecycle.md +57 -0
  253. package/knowledge/cd-maturity-model.md +82 -0
  254. package/knowledge/cd-test-architecture.md +190 -0
  255. package/knowledge/ci-cd-file-scope.md +24 -0
  256. package/knowledge/codegraph-vs-graphify.md +192 -0
  257. package/knowledge/component-test-patterns.md +139 -0
  258. package/knowledge/database-change-management.md +80 -0
  259. package/knowledge/database-test-patterns.md +79 -0
  260. package/knowledge/decision-defaults.md +88 -0
  261. package/knowledge/dependency-breaking-techniques.md +116 -0
  262. package/knowledge/deployment-pipeline.md +86 -0
  263. package/knowledge/design-smells.md +122 -0
  264. package/knowledge/directory-enumeration.md +38 -0
  265. package/knowledge/domain-modeling.md +123 -0
  266. package/knowledge/evidence-bundle.md +90 -0
  267. package/knowledge/exploratory-testing-field-guide.md +122 -0
  268. package/knowledge/failure-routing.md +28 -0
  269. package/knowledge/fixture-construction.md +56 -0
  270. package/knowledge/frontend-component-architecture.md +139 -0
  271. package/knowledge/gherkin-quality-review-dispatch.md +135 -0
  272. package/knowledge/index.json +6766 -0
  273. package/knowledge/internal-collaborator-doubling.md +101 -0
  274. package/knowledge/legacy-test-strategy.md +71 -0
  275. package/knowledge/long-run-waiting.md +66 -0
  276. package/knowledge/microservice-testing.md +71 -0
  277. package/knowledge/model-pricing.json +23 -0
  278. package/knowledge/mutation-score-formulas.md +60 -0
  279. package/knowledge/object-calisthenics.md +147 -0
  280. package/knowledge/oracle-provenance.md +94 -0
  281. package/knowledge/orchestrator-script-implementation.md +185 -0
  282. package/knowledge/owasp-detection.md +148 -0
  283. package/knowledge/plan-review-rubric.md +56 -0
  284. package/knowledge/proxy-connectivity.md +62 -0
  285. package/knowledge/reactive-effect-patterns.md +73 -0
  286. package/knowledge/recon-inventory-excludes.txt +32 -0
  287. package/knowledge/references/bdd-value-guide.md +61 -0
  288. package/knowledge/references/csharp-http-client-testing.md +264 -0
  289. package/knowledge/release-strategies.md +74 -0
  290. package/knowledge/report-output-location.md +117 -0
  291. package/knowledge/report-pdf-integration.md +63 -0
  292. package/knowledge/report-print.css +129 -0
  293. package/knowledge/report-template.md +114 -0
  294. package/knowledge/report-to-pdf.md +69 -0
  295. package/knowledge/request-processing-flow.md +63 -0
  296. package/knowledge/result-verification.md +52 -0
  297. package/knowledge/review-agent-output-contract.md +121 -0
  298. package/knowledge/review-lens-classification.md +113 -0
  299. package/knowledge/review-rubric.md +62 -0
  300. package/knowledge/review-template.md +104 -0
  301. package/knowledge/rule-fixtures/A02.insecure-random-js/negative.js +1 -0
  302. package/knowledge/rule-fixtures/A02.insecure-random-js/positive.js +1 -0
  303. package/knowledge/rule-fixtures/A02.weak-hashing-md5/negative.py +1 -0
  304. package/knowledge/rule-fixtures/A02.weak-hashing-md5/positive.py +1 -0
  305. package/knowledge/rule-fixtures/A03.command-injection/negative.js +1 -0
  306. package/knowledge/rule-fixtures/A03.command-injection/positive.js +1 -0
  307. package/knowledge/rule-fixtures/A03.sql-injection/negative.js +1 -0
  308. package/knowledge/rule-fixtures/A03.sql-injection/positive.js +1 -0
  309. package/knowledge/rule-fixtures/A03.xss-innerhtml/negative.js +1 -0
  310. package/knowledge/rule-fixtures/A03.xss-innerhtml/positive.js +1 -0
  311. package/knowledge/rule-fixtures/A05.cors-wildcard/negative.js +1 -0
  312. package/knowledge/rule-fixtures/A05.cors-wildcard/positive.js +1 -0
  313. package/knowledge/rule-fixtures/A05.default-credentials/negative.js +1 -0
  314. package/knowledge/rule-fixtures/A05.default-credentials/positive.js +1 -0
  315. package/knowledge/rule-fixtures/A07.jwt-alg-none/negative.js +1 -0
  316. package/knowledge/rule-fixtures/A07.jwt-alg-none/positive.js +1 -0
  317. package/knowledge/rule-fixtures/A08.binary-formatter/negative.cs +1 -0
  318. package/knowledge/rule-fixtures/A08.binary-formatter/positive.cs +1 -0
  319. package/knowledge/rule-fixtures/A08.js-eval/negative.js +1 -0
  320. package/knowledge/rule-fixtures/A08.js-eval/positive.js +1 -0
  321. package/knowledge/rule-fixtures/A08.object-input-stream/negative.java +1 -0
  322. package/knowledge/rule-fixtures/A08.object-input-stream/positive.java +1 -0
  323. package/knowledge/schemas/disposition-register-v1.json +65 -0
  324. package/knowledge/schemas/recon-envelope-v1.json +198 -0
  325. package/knowledge/schemas/unified-finding-v1.json +72 -0
  326. package/knowledge/security-primitives-contract.md +301 -0
  327. package/knowledge/security-review-rule-map.yaml +107 -0
  328. package/knowledge/skills-registry.md +72 -0
  329. package/knowledge/task-size-classifier.md +103 -0
  330. package/knowledge/telemetry-schema.md +881 -0
  331. package/knowledge/test-automation-maturity.md +56 -0
  332. package/knowledge/test-automation-principles.md +71 -0
  333. package/knowledge/test-cadence-tradeoffs.md +68 -0
  334. package/knowledge/test-doubles.md +105 -0
  335. package/knowledge/test-file-indicators.md +22 -0
  336. package/knowledge/test-layer-gates.md +35 -0
  337. package/knowledge/test-matrix-examples/django-batch.md +24 -0
  338. package/knowledge/test-matrix-examples/dotnet-grpc-fronting-api.md +90 -0
  339. package/knowledge/test-matrix-examples/dotnet-http-consumer.md +131 -0
  340. package/knowledge/test-matrix-examples/react-node-spa.md +24 -0
  341. package/knowledge/test-matrix-examples/spring-boot-service.md +25 -0
  342. package/knowledge/test-matrix-examples/ssr-htmx.md +24 -0
  343. package/knowledge/test-organization.md +70 -0
  344. package/knowledge/test-pyramid.md +84 -0
  345. package/knowledge/test-refactoring.md +67 -0
  346. package/knowledge/test-review-division-of-labor.md +85 -0
  347. package/knowledge/test-smells.md +80 -0
  348. package/knowledge/test-stack-profiles/bdd-frameworks.md +235 -0
  349. package/knowledge/test-stack-profiles/django.md +13 -0
  350. package/knowledge/test-stack-profiles/dotnet.md +18 -0
  351. package/knowledge/test-stack-profiles/go.md +16 -0
  352. package/knowledge/test-stack-profiles/node.md +16 -0
  353. package/knowledge/test-stack-profiles/react.md +12 -0
  354. package/knowledge/test-stack-profiles/spring-boot.md +16 -0
  355. package/knowledge/test-stack-profiles/ssr-htmx.md +14 -0
  356. package/knowledge/test-stack-profiles/vue.md +12 -0
  357. package/knowledge/test-strategy.md +70 -0
  358. package/knowledge/testability-patterns.md +240 -0
  359. package/knowledge/testing-quadrants.md +44 -0
  360. package/knowledge/testing-techniques/approval.md +15 -0
  361. package/knowledge/testing-techniques/chaos.md +17 -0
  362. package/knowledge/testing-techniques/fuzz.md +15 -0
  363. package/knowledge/testing-techniques/property-based.md +15 -0
  364. package/knowledge/testing-techniques/schema-validation.md +15 -0
  365. package/knowledge/testing-techniques/screenshot.md +15 -0
  366. package/knowledge/three-phase-workflow.md +198 -0
  367. package/knowledge/value-patterns.md +55 -0
  368. package/knowledge/verification-mode.md +116 -0
  369. package/knowledge/virtual-service-libraries.md +75 -0
  370. package/knowledge/wave-consolidation-guidance.md +21 -0
  371. package/overrides/agents/Explore.md +15 -0
  372. package/overrides/agents/general-purpose.md +10 -0
  373. package/overrides/notes/autoship.md +6 -0
  374. package/overrides/notes/issues-from-assessment.md +3 -0
  375. package/overrides/notes/issues-from-plan.md +3 -0
  376. package/overrides/notes/mutation-night-watch.md +3 -0
  377. package/overrides/notes/mutation-testing.md +3 -0
  378. package/overrides/notes/pr.md +7 -0
  379. package/overrides/notes/project-init.md +6 -0
  380. package/overrides/notes/setup.md +13 -0
  381. package/overrides/notes/specs.md +3 -0
  382. package/overrides/skills/headless-run/SKILL.md +45 -0
  383. package/overrides/skills/upgrade/SKILL.md +30 -0
  384. package/overrides/skills/version/SKILL.md +25 -0
  385. package/package.json +36 -0
  386. package/scripts/authoring_digest.py +93 -0
  387. package/scripts/autoship_discover.py +121 -0
  388. package/scripts/autoship_group.py +409 -0
  389. package/scripts/autoship_proposals.py +494 -0
  390. package/scripts/autoship_queue.py +291 -0
  391. package/scripts/autoship_reclaim.py +495 -0
  392. package/scripts/build_jobs.py +108 -0
  393. package/scripts/build_rollback_point.py +240 -0
  394. package/scripts/build_slice_scope.py +157 -0
  395. package/scripts/build_wave.py +109 -0
  396. package/scripts/build_wave_reconcile.py +252 -0
  397. package/scripts/build_worktree_baseref.py +113 -0
  398. package/scripts/check_agent_scope.py +117 -0
  399. package/scripts/check_agent_tool_mapping.py +213 -0
  400. package/scripts/check_review_agent_mcp_tools.py +317 -0
  401. package/scripts/check_security_assessment_mcp_tools.py +165 -0
  402. package/scripts/checkpoint_abort.py +502 -0
  403. package/scripts/claude_setup_review.py +438 -0
  404. package/scripts/codebase_recon.py +556 -0
  405. package/scripts/coverage_config.py +623 -0
  406. package/scripts/coverage_delta_steering.py +330 -0
  407. package/scripts/coverage_discovery_dotnet.py +315 -0
  408. package/scripts/coverage_discovery_java.py +742 -0
  409. package/scripts/coverage_discovery_js.py +546 -0
  410. package/scripts/coverage_gap_ranking.py +556 -0
  411. package/scripts/coverage_readiness.py +455 -0
  412. package/scripts/coverage_report_parse.py +521 -0
  413. package/scripts/detect_bdd_convention.py +252 -0
  414. package/scripts/eval_ablation.py +376 -0
  415. package/scripts/gherkin_analysis_coverage_gate.py +306 -0
  416. package/scripts/gherkin_cross_feature_duplicate_titles_gate.py +173 -0
  417. package/scripts/gherkin_effectiveness_rollup.py +238 -0
  418. package/scripts/gherkin_failure_path_gate.py +206 -0
  419. package/scripts/gherkin_feature_merge.py +720 -0
  420. package/scripts/gherkin_stub_gate.py +163 -0
  421. package/scripts/gherkin_stub_merge.py +479 -0
  422. package/scripts/git_origin_host.py +88 -0
  423. package/scripts/install-java-static-analysis.py +110 -0
  424. package/scripts/issue_deps.py +74 -0
  425. package/scripts/lib/_bdd_markers.py +28 -0
  426. package/scripts/lib/_gherkin_text.py +93 -0
  427. package/scripts/lib/_vendored_tree.py +70 -0
  428. package/scripts/lib/autoship_state.py +397 -0
  429. package/scripts/lib/claude_md_guard.py +226 -0
  430. package/scripts/lib/deterministic_recon.py +446 -0
  431. package/scripts/lib/mcp_tool_grants.py +211 -0
  432. package/scripts/lib/plan_parse.py +386 -0
  433. package/scripts/lib/review_result.py +84 -0
  434. package/scripts/lib/review_roster.py +86 -0
  435. package/scripts/lib/session_log/__init__.py +34 -0
  436. package/scripts/lib/session_log/__pycache__/__init__.cpython-314.pyc +0 -0
  437. package/scripts/lib/session_log/__pycache__/records.cpython-314.pyc +0 -0
  438. package/scripts/lib/session_log/classify.py +231 -0
  439. package/scripts/lib/session_log/corrections.py +194 -0
  440. package/scripts/lib/session_log/discovery.py +108 -0
  441. package/scripts/lib/session_log/records.py +218 -0
  442. package/scripts/lib/session_log/redact.py +76 -0
  443. package/scripts/lib/session_log/signals.py +373 -0
  444. package/scripts/lib/session_report_downstream.py +614 -0
  445. package/scripts/lib/session_report_maintainer.py +1273 -0
  446. package/scripts/lib/session_report_shared.py +262 -0
  447. package/scripts/lib/settings_hook_guard.py +157 -0
  448. package/scripts/lib/slug.py +33 -0
  449. package/scripts/lib/stub_extractors/__init__.py +82 -0
  450. package/scripts/lib/stub_extractors/_common.py +328 -0
  451. package/scripts/lib/stub_extractors/csharp.py +19 -0
  452. package/scripts/lib/stub_extractors/go.py +173 -0
  453. package/scripts/lib/stub_extractors/java.py +18 -0
  454. package/scripts/lib/stub_extractors/jsts.py +126 -0
  455. package/scripts/mutation_stack_sections.py +149 -0
  456. package/scripts/mutation_yield_steering.py +345 -0
  457. package/scripts/orchestrator.py +895 -0
  458. package/scripts/plan_gherkin_export.py +227 -0
  459. package/scripts/plan_waves.py +208 -0
  460. package/scripts/pr_close_keyword_lint.py +108 -0
  461. package/scripts/progress_guardian.py +888 -0
  462. package/scripts/recon_inventory.py +273 -0
  463. package/scripts/review_findings_log.py +93 -0
  464. package/scripts/run_invariants.py +124 -0
  465. package/scripts/select_lenses.py +640 -0
  466. package/scripts/session_report.py +486 -0
  467. package/scripts/set_autocompact_env.py +221 -0
  468. package/scripts/ship_resume_guard.py +135 -0
  469. package/scripts/ship_review_gate.py +63 -0
  470. package/scripts/specs_convention_marker.py +103 -0
  471. package/scripts/test_improve_resume.py +277 -0
  472. package/scripts/test_review_mechanics.py +958 -0
  473. package/scripts/token_efficiency_review.py +322 -0
  474. package/scripts/verdict_scope.py +285 -0
  475. package/scripts/verify_gherkin_quality_critic_isolation.py +296 -0
  476. package/scripts/verify_tier.py +157 -0
  477. package/skills/adr-tools/SKILL.md +118 -0
  478. package/skills/agent-readiness/SKILL.md +105 -0
  479. package/skills/agent-readiness/ai_friendly_analyzers.py +326 -0
  480. package/skills/agent-readiness/scanner.py +441 -0
  481. package/skills/agent-readiness/scorecard.yaml +88 -0
  482. package/skills/api-design/SKILL.md +115 -0
  483. package/skills/apply-fixes/SKILL.md +171 -0
  484. package/skills/apply-test-doubles/SKILL.md +321 -0
  485. package/skills/artifact-lifecycle/SKILL.md +127 -0
  486. package/skills/autoship/SKILL.md +1124 -0
  487. package/skills/benchmark/SKILL.md +105 -0
  488. package/skills/branch-workflow/SKILL.md +89 -0
  489. package/skills/browse/SKILL.md +184 -0
  490. package/skills/browser-testing/SKILL.md +62 -0
  491. package/skills/browser-testing/references/playwright-patterns.md +216 -0
  492. package/skills/build/SKILL.md +422 -0
  493. package/skills/build/references/static-self-heal.md +245 -0
  494. package/skills/careful/SKILL.md +72 -0
  495. package/skills/cd-test-architecture/SKILL.md +371 -0
  496. package/skills/ci-debugging/SKILL.md +105 -0
  497. package/skills/co-evolution-audit/SKILL.md +269 -0
  498. package/skills/code-review/SKILL.md +1015 -0
  499. package/skills/code-review/examples/aggregated-sample.json +56 -0
  500. package/skills/code-review/examples/sample-report.md +41 -0
  501. package/skills/code-review/output-format.md +478 -0
  502. package/skills/code-review/scripts/activation.py +86 -0
  503. package/skills/code-review/scripts/change_impact.py +357 -0
  504. package/skills/code-review/scripts/change_shape.py +372 -0
  505. package/skills/code-review/scripts/change_size.py +212 -0
  506. package/skills/code-review/scripts/changed_file_list.py +141 -0
  507. package/skills/code-review/scripts/closing_pass.py +187 -0
  508. package/skills/code-review/scripts/consolidate.py +277 -0
  509. package/skills/code-review/scripts/contract_failure_report.py +185 -0
  510. package/skills/code-review/scripts/dispatch_reconcile.py +66 -0
  511. package/skills/code-review/scripts/dispatch_waves.py +164 -0
  512. package/skills/code-review/scripts/finding_signature.py +446 -0
  513. package/skills/code-review/scripts/ledger.py +283 -0
  514. package/skills/code-review/scripts/partition.py +169 -0
  515. package/skills/code-review/scripts/render_tiered_findings.py +274 -0
  516. package/skills/code-review/scripts/repo_invariants.py +1066 -0
  517. package/skills/code-review/scripts/review_context_pack.py +306 -0
  518. package/skills/code-review/scripts/review_round_log.py +345 -0
  519. package/skills/code-review/scripts/review_value_coverage.py +297 -0
  520. package/skills/code-review/scripts/validate_review_output.py +467 -0
  521. package/skills/code-review/sliced-mode.md +205 -0
  522. package/skills/competitive-analysis/SKILL.md +191 -0
  523. package/skills/context-loading-protocol/SKILL.md +157 -0
  524. package/skills/continue/SKILL.md +90 -0
  525. package/skills/cost-report/SKILL.md +178 -0
  526. package/skills/coverage-baseline/SKILL.md +335 -0
  527. package/skills/coverage-baseline/references/multi-project-discovery.md +202 -0
  528. package/skills/coverage-delta/SKILL.md +181 -0
  529. package/skills/coverage-delta/references/mutation-gate.md +70 -0
  530. package/skills/design-doc/SKILL.md +95 -0
  531. package/skills/design-interrogation/SKILL.md +89 -0
  532. package/skills/design-it-twice/SKILL.md +91 -0
  533. package/skills/docker-image-audit/SKILL.md +108 -0
  534. package/skills/docker-image-audit/references/install-guide.md +64 -0
  535. package/skills/docker-image-audit/references/report-template.md +73 -0
  536. package/skills/docker-image-create/SKILL.md +185 -0
  537. package/skills/domain-analysis/SKILL.md +183 -0
  538. package/skills/domain-driven-design/SKILL.md +194 -0
  539. package/skills/exploratory-testing/SKILL.md +108 -0
  540. package/skills/explore/SKILL.md +51 -0
  541. package/skills/farley-score/SKILL.md +165 -0
  542. package/skills/feature-file-validation/SKILL.md +78 -0
  543. package/skills/feature-file-validation/references/validation-rules.md +115 -0
  544. package/skills/feedback-learning/SKILL.md +414 -0
  545. package/skills/fix/SKILL.md +450 -0
  546. package/skills/freeze/SKILL.md +68 -0
  547. package/skills/frontend-architecture/SKILL.md +113 -0
  548. package/skills/gherkin-derive/SKILL.md +630 -0
  549. package/skills/gherkin-public/SKILL.md +266 -0
  550. package/skills/governance-compliance/SKILL.md +150 -0
  551. package/skills/guard/SKILL.md +75 -0
  552. package/skills/handoff/SKILL.md +139 -0
  553. package/skills/handoff/references/summary-templates.md +242 -0
  554. package/skills/harness-audit/SKILL.md +751 -0
  555. package/skills/harness-audit/scripts/lesson_validate.py +386 -0
  556. package/skills/harness-audit/scripts/redundancy_criterion.py +188 -0
  557. package/skills/headless-run/SKILL.md +45 -0
  558. package/skills/headless-run/scripts/isolated_dispatch.py +381 -0
  559. package/skills/help/SKILL.md +72 -0
  560. package/skills/hexagonal-architecture/SKILL.md +85 -0
  561. package/skills/human-oversight-protocol/SKILL.md +224 -0
  562. package/skills/issues-from-assessment/SKILL.md +223 -0
  563. package/skills/issues-from-plan/SKILL.md +133 -0
  564. package/skills/legacy-code/SKILL.md +132 -0
  565. package/skills/mermaid-diagramming/SKILL.md +120 -0
  566. package/skills/mutation-night-watch/SKILL.md +154 -0
  567. package/skills/mutation-night-watch/references/scheduling.md +135 -0
  568. package/skills/mutation-testing/SKILL.md +396 -0
  569. package/skills/mutation-testing/references/languages/csharp-stryker-net.md +676 -0
  570. package/skills/mutation-testing/references/languages/go-go-mutesting.md +95 -0
  571. package/skills/mutation-testing/references/languages/java-pitest.md +77 -0
  572. package/skills/mutation-testing/references/languages/javascript-stryker.md +188 -0
  573. package/skills/mutation-testing/references/languages/python-mutmut.md +97 -0
  574. package/skills/mutation-testing/references/time-estimation.md +34 -0
  575. package/skills/mutation-testing/references/tool-detection.md +15 -0
  576. package/skills/mutation-testing/references/workflow-callers.md +23 -0
  577. package/skills/mutation-testing/scripts/__pycache__/xunit_v3_feature_detector.cpython-314.pyc +0 -0
  578. package/skills/mutation-testing/scripts/csharp_stryker_net_slice_runner.py +635 -0
  579. package/skills/mutation-testing/scripts/csharp_stryker_net_status_loop.py +525 -0
  580. package/skills/mutation-testing/scripts/csharp_stryker_net_wrapper.py +681 -0
  581. package/skills/mutation-testing/scripts/mutation_baseline_reuse.py +292 -0
  582. package/skills/mutation-testing/scripts/mutation_exclude_policy.py +268 -0
  583. package/skills/mutation-testing/scripts/mutation_feasibility_gate.py +463 -0
  584. package/skills/mutation-testing/scripts/mutation_kill_headless.py +331 -0
  585. package/skills/mutation-testing/scripts/mutation_kill_insert.py +199 -0
  586. package/skills/mutation-testing/scripts/mutation_kill_insert_python.py +150 -0
  587. package/skills/mutation-testing/scripts/mutation_kill_loop.py +869 -0
  588. package/skills/mutation-testing/scripts/mutation_kill_loop_python.py +949 -0
  589. package/skills/mutation-testing/scripts/mutation_kill_retry.py +592 -0
  590. package/skills/mutation-testing/scripts/mutation_kill_shared.py +620 -0
  591. package/skills/mutation-testing/scripts/mutation_nightwatch.py +462 -0
  592. package/skills/mutation-testing/scripts/mutation_nightwatch_stacks.py +425 -0
  593. package/skills/mutation-testing/scripts/mutation_report.py +743 -0
  594. package/skills/mutation-testing/scripts/mutation_report_cli.py +175 -0
  595. package/skills/mutation-testing/scripts/mutation_safety_gate.py +69 -0
  596. package/skills/mutation-testing/scripts/stryker_shard_pipeline.py +847 -0
  597. package/skills/mutation-testing/scripts/stryker_shard_setup.py +440 -0
  598. package/skills/mutation-testing/scripts/stryker_timeout_retry.py +142 -0
  599. package/skills/mutation-testing/scripts/xunit_v3_feature_detector.py +341 -0
  600. package/skills/performance-benchmark/SKILL.md +174 -0
  601. package/skills/performance-benchmark/examples/report-format.md +43 -0
  602. package/skills/performance-benchmark/references/benchmark-script.md +169 -0
  603. package/skills/performance-metrics/SKILL.md +265 -0
  604. package/skills/plan/SKILL.md +199 -0
  605. package/skills/plan/references/gherkin-persistence.md +43 -0
  606. package/skills/plan/references/plan-template.md +182 -0
  607. package/skills/pr/SKILL.md +289 -0
  608. package/skills/pr/scripts/gate_retry_state.py +368 -0
  609. package/skills/project-init/README.md +141 -0
  610. package/skills/project-init/SKILL.md +1197 -0
  611. package/skills/project-init/evals/evals.json +200 -0
  612. package/skills/project-init/references/capability-tools.md +55 -0
  613. package/skills/project-init/references/configs.md +221 -0
  614. package/skills/property-based-testing/SKILL.md +121 -0
  615. package/skills/property-based-testing/fixtures/invariant_fixture.py +15 -0
  616. package/skills/property-based-testing/fixtures/js-roundtrip/README.md +42 -0
  617. package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/fast-check/LICENSE +21 -0
  618. package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/fast-check/README.md +263 -0
  619. package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/fast-check/lib/cjs/fast-check.d.ts +5165 -0
  620. package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/fast-check/lib/cjs/fast-check.js +12147 -0
  621. package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/fast-check/lib/cjs/package.json +3 -0
  622. package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/fast-check/lib/cjs/types57/fast-check.d.ts +5165 -0
  623. package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/fast-check/lib/fast-check.d.ts +5165 -0
  624. package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/fast-check/lib/fast-check.js +12011 -0
  625. package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/fast-check/lib/rolldown-runtime-D7D4PA-g.js +13 -0
  626. package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/fast-check/lib/types57/fast-check.d.ts +5165 -0
  627. package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/fast-check/package.json +94 -0
  628. package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/LICENSE +21 -0
  629. package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/README.md +168 -0
  630. package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/RandomGenerator-DcXj09Ch.d.ts +14 -0
  631. package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/distribution/uniformBigInt.d.ts +15 -0
  632. package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/distribution/uniformBigInt.js +38 -0
  633. package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/distribution/uniformFloat32.d.ts +15 -0
  634. package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/distribution/uniformFloat32.js +18 -0
  635. package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/distribution/uniformFloat64.d.ts +15 -0
  636. package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/distribution/uniformFloat64.js +22 -0
  637. package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/distribution/uniformInt.d.ts +15 -0
  638. package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/distribution/uniformInt.js +134 -0
  639. package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/RandomGenerator-DcXj09Ch.d.ts +14 -0
  640. package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/distribution/uniformBigInt.d.ts +15 -0
  641. package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/distribution/uniformBigInt.js +37 -0
  642. package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/distribution/uniformFloat32.d.ts +15 -0
  643. package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/distribution/uniformFloat32.js +17 -0
  644. package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/distribution/uniformFloat64.d.ts +15 -0
  645. package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/distribution/uniformFloat64.js +21 -0
  646. package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/distribution/uniformInt.d.ts +15 -0
  647. package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/distribution/uniformInt.js +133 -0
  648. package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/generator/congruential32.d.ts +7 -0
  649. package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/generator/congruential32.js +44 -0
  650. package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/generator/mersenne.d.ts +7 -0
  651. package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/generator/mersenne.js +90 -0
  652. package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/generator/xoroshiro128plus.d.ts +7 -0
  653. package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/generator/xoroshiro128plus.js +80 -0
  654. package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/generator/xorshift128plus.d.ts +7 -0
  655. package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/generator/xorshift128plus.js +78 -0
  656. package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/package.json +3 -0
  657. package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/types/JumpableRandomGenerator.d.ts +16 -0
  658. package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/types/JumpableRandomGenerator.js +0 -0
  659. package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/types/RandomGenerator.d.ts +2 -0
  660. package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/types/RandomGenerator.js +0 -0
  661. package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/utils/generateN.d.ts +6 -0
  662. package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/utils/generateN.js +8 -0
  663. package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/utils/purify.d.ts +12 -0
  664. package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/utils/purify.js +9 -0
  665. package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/utils/skipN.d.ts +6 -0
  666. package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/utils/skipN.js +6 -0
  667. package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/generator/congruential32.d.ts +7 -0
  668. package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/generator/congruential32.js +46 -0
  669. package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/generator/mersenne.d.ts +7 -0
  670. package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/generator/mersenne.js +92 -0
  671. package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/generator/xoroshiro128plus.d.ts +7 -0
  672. package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/generator/xoroshiro128plus.js +82 -0
  673. package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/generator/xorshift128plus.d.ts +7 -0
  674. package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/generator/xorshift128plus.js +80 -0
  675. package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/types/JumpableRandomGenerator.d.ts +16 -0
  676. package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/types/JumpableRandomGenerator.js +0 -0
  677. package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/types/RandomGenerator.d.ts +2 -0
  678. package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/types/RandomGenerator.js +0 -0
  679. package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/utils/generateN.d.ts +6 -0
  680. package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/utils/generateN.js +9 -0
  681. package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/utils/purify.d.ts +12 -0
  682. package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/utils/purify.js +10 -0
  683. package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/utils/skipN.d.ts +6 -0
  684. package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/utils/skipN.js +7 -0
  685. package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/package.json +133 -0
  686. package/skills/property-based-testing/fixtures/js-roundtrip/package-lock.json +1179 -0
  687. package/skills/property-based-testing/fixtures/js-roundtrip/package.json +14 -0
  688. package/skills/property-based-testing/fixtures/js-roundtrip/roundtrip.js +29 -0
  689. package/skills/property-based-testing/fixtures/js-roundtrip/roundtrip.properties.test.js +16 -0
  690. package/skills/property-based-testing/fixtures/no_property_fixture.py +10 -0
  691. package/skills/property-based-testing/fixtures/roundtrip_fixture.py +16 -0
  692. package/skills/property-based-testing/references/languages/javascript.md +54 -0
  693. package/skills/property-based-testing/scripts/detect_and_dispatch.py +80 -0
  694. package/skills/property-based-testing/scripts/hypothesis_scaffold.py +276 -0
  695. package/skills/proxy-resilience/SKILL.md +84 -0
  696. package/skills/quality-gate-pipeline/SKILL.md +184 -0
  697. package/skills/quality-targets-converge/SKILL.md +254 -0
  698. package/skills/repo-review/SKILL.md +159 -0
  699. package/skills/report-pdf/SKILL.md +66 -0
  700. package/skills/review/SKILL.md +47 -0
  701. package/skills/review-agent/SKILL.md +152 -0
  702. package/skills/review-summary/SKILL.md +73 -0
  703. package/skills/run-report/SKILL.md +70 -0
  704. package/skills/semantic-duplication-scan/SKILL.md +337 -0
  705. package/skills/semantic-scan/SKILL.md +53 -0
  706. package/skills/semgrep-analyze/SKILL.md +139 -0
  707. package/skills/setup/SKILL.md +1122 -0
  708. package/skills/ship/SKILL.md +240 -0
  709. package/skills/source-verification/SKILL.md +210 -0
  710. package/skills/source-verification/scripts/claim_extractor.py +155 -0
  711. package/skills/specs/.size-baseline.json +4 -0
  712. package/skills/specs/SKILL.md +243 -0
  713. package/skills/specs/references/completeness-checklist.md +83 -0
  714. package/skills/specs/references/extraction.md +58 -0
  715. package/skills/specs/references/glossary.md +59 -0
  716. package/skills/specs/references/persistence.md +115 -0
  717. package/skills/specs/references/predictability-check.md +77 -0
  718. package/skills/static-analysis-integration/SKILL.md +235 -0
  719. package/skills/static-analysis-integration/adapters/_envelope.py +26 -0
  720. package/skills/static-analysis-integration/adapters/jscpd-adapter.py +66 -0
  721. package/skills/static-analysis-integration/adapters/lizard-adapter.py +81 -0
  722. package/skills/static-analysis-integration/adapters/mypy-adapter.py +50 -0
  723. package/skills/static-analysis-integration/adapters/mypy-src-layout.py +93 -0
  724. package/skills/static-analysis-integration/adapters/security-review-adapter.py +212 -0
  725. package/skills/static-analysis-integration/maintenance.md +23 -0
  726. package/skills/static-analysis-integration/references/language-setup.md +228 -0
  727. package/skills/static-analysis-integration/references/sarif-parser.md +124 -0
  728. package/skills/static-analysis-integration/references/security-review-adapter.md +118 -0
  729. package/skills/static-analysis-integration/references/tool-configs.md +617 -0
  730. package/skills/static-analysis-integration/rulesets/pmd-quickstart.xml +24 -0
  731. package/skills/stryker-xunit-v2-shim/SKILL.md +274 -0
  732. package/skills/stryker-xunit-v2-shim/references/shim-howto.md +256 -0
  733. package/skills/stryker-xunit-v2-shim/scripts/generate_shim.py +143 -0
  734. package/skills/systematic-debugging/SKILL.md +130 -0
  735. package/skills/telemetry/SKILL.md +75 -0
  736. package/skills/test-audit-disable/SKILL.md +129 -0
  737. package/skills/test-design/SKILL.md +177 -0
  738. package/skills/test-design/scripts/__pycache__/internal_double_detector.cpython-314.pyc +0 -0
  739. package/skills/test-design/scripts/internal_double_detector.py +631 -0
  740. package/skills/test-design-advisor/SKILL.md +166 -0
  741. package/skills/test-driven-development/SKILL.md +169 -0
  742. package/skills/test-health/SKILL.md +262 -0
  743. package/skills/test-improve/SKILL.md +239 -0
  744. package/skills/test-improve/references/phase-0-approach-contract.md +228 -0
  745. package/skills/test-improve/references/phase-1-analyze.md +131 -0
  746. package/skills/test-improve/references/phase-2-baseline.md +121 -0
  747. package/skills/test-improve/references/phase-3-derive-gherkin.md +53 -0
  748. package/skills/test-improve/references/phase-4-plan-fixes.md +34 -0
  749. package/skills/test-improve/references/phase-5-improve.md +215 -0
  750. package/skills/test-improve/references/phase-6-refactor-decision.md +45 -0
  751. package/skills/test-improve/references/phase-7-refactor.md +44 -0
  752. package/skills/test-improve/references/phase-8-validate.md +66 -0
  753. package/skills/test-improve/references/phase-9-close-out-prompt.md +11 -0
  754. package/skills/test-improve/references/phase-9-report.md +62 -0
  755. package/skills/test-improve/references/review-loop.md +92 -0
  756. package/skills/test-improve/templates/executive-summary.md +123 -0
  757. package/skills/threat-modeling/SKILL.md +108 -0
  758. package/skills/triage/SKILL.md +211 -0
  759. package/skills/ubiquitous-language/SKILL.md +192 -0
  760. package/skills/ubiquitous-language/scripts/collect_domain_signals.py +300 -0
  761. package/skills/unfreeze/SKILL.md +37 -0
  762. package/skills/upgrade/SKILL.md +31 -0
  763. package/skills/upgrade/scripts/check_version_drift.py +113 -0
  764. package/skills/upgrade/scripts/enable_autoupdate.py +149 -0
  765. package/skills/version/SKILL.md +25 -0
  766. package/sync/__pycache__/sync_upstream.cpython-314.pyc +0 -0
  767. package/sync/sync_upstream.py +293 -0
  768. package/templates/ACCEPTED-RISKS.md.tmpl +46 -0
  769. package/templates/agents/agent-template.md +151 -0
  770. package/templates/agents/angular-testing.md +66 -0
  771. package/templates/agents/csharp-quality.md +63 -0
  772. package/templates/agents/esm-enforcer.md +52 -0
  773. package/templates/agents/front-end-testing.md +65 -0
  774. package/templates/agents/go-quality.md +65 -0
  775. package/templates/agents/python-quality.md +62 -0
  776. package/templates/agents/react-testing.md +61 -0
  777. package/templates/agents/ts-enforcer.md +60 -0
  778. package/templates/agents/twelve-factor-audit.md +49 -0
  779. package/tools/entropy-check.py +250 -0
  780. package/tools/model-hash-verify.py +213 -0
@@ -0,0 +1,881 @@
1
+ # Telemetry Schema Reference
2
+
3
+ Every `.claude/metrics/*.jsonl` and `.claude/metrics/*.json` file the dev-team plugin writes,
4
+ in one place, so `session-analysis`, `/session-review`, `/harness-audit`,
5
+ `/cost-report`, and future cross-machine aggregation (#178) compose against
6
+ stable, named schemas instead of reverse-engineering emitters.
7
+
8
+ **Privacy stance (non-negotiable, all streams):** rule IDs, counts, hashes,
9
+ and enums only — never command text, prompt text, file contents, or free-text
10
+ reasons beyond what a stream explicitly documents below as human-authored
11
+ (e.g. `config-changelog.jsonl`'s `description`, which is a deliberate,
12
+ human/agent-reviewed audit note, not incidental free text). Where a stream
13
+ predates this doc and already carries a `reason` field with freeform text
14
+ (e.g. `refactor-freeze.jsonl`'s internal-error diagnostics), that is existing,
15
+ unchanged precedent — not a new exception.
16
+
17
+ Each section below names: fields, types, emitter, consent gating, and
18
+ consumers. A companion test
19
+ (`tests/hooks/test_boundary_events.py::test_schema_doc_covers_all_metrics_paths`)
20
+ cross-checks every `.claude/metrics/*.jsonl` / `.claude/metrics/*.json` path string referenced
21
+ in shipped code against this doc's coverage and fails on omission.
22
+
23
+ ---
24
+
25
+ ## `boundary-events.jsonl`
26
+
27
+ **Added by #859.** The boundary-level (policy-gateway) channel: every guard
28
+ hook's block/warn/bypass decision, plus human-intervention keywords. Extended
29
+ by #906 with a fifth decision, `revert`, for hooks that don't warn or block
30
+ but actively correct state after the fact. Extended again by #1461 with a
31
+ sixth decision, `record` — a **non-verdict, observational** entry: it does
32
+ not block, warn, bypass, intervene, or revert anything, it merely notes that
33
+ a genuine, registered review-agent dispatch occurred. Emitted by
34
+ `hooks/agent_dispatch_ledger.py` on every `Agent`/`Task` dispatch whose
35
+ `subagent_type` is a real, registered `agents/*-review.md` name (never a
36
+ fabricated/unregistered one — those are never written to the ledger at all).
37
+ This is a **high-frequency** entry (one per genuine review-agent dispatch,
38
+ not a rare guard trip like the other five decisions) — a consumer counting
39
+ "policy decisions" or "guard verdicts" from this stream must explicitly
40
+ exclude `record` rows, or it will badly overcount routine dispatch activity
41
+ as gate verdicts. `hooks/pre_pr_review.py`'s `.pr-review-passed` gate (#1886;
42
+ formerly `hooks/pre_commit_review.py`'s `.review-passed` gate on `git
43
+ commit` — that hook is now a documented no-op) reads this stream (via
44
+ `hooks/lib/review_gate_corroboration.py`) to corroborate that a
45
+ hash-matching gate write was backed by real, independent Agent-tool
46
+ dispatch — see that module's own docstring for its fail-**closed** posture,
47
+ the deliberate opposite of this stream's own fail-open write side.
48
+ Extended again by #1763 with a seventh decision, `dispatch-failure` — also
49
+ **non-verdict, observational**, mirroring `record`'s precedent, but the
50
+ opposite polarity: it notes that a dispatched review agent still failed to
51
+ return a contract-valid result after one retry. Emitted via
52
+ `hooks/lib/boundary_events.py`'s CLI (`--event dispatch-failure --agent
53
+ <name> --subject-hash <hash>`) from `skills/code-review/SKILL.md` Step 4;
54
+ `<name>` is validated against the registered review-agent set at write time
55
+ (with the same plugin-prefix normalization as `record`) — an unregistered
56
+ name is silently not recorded. Unlike `record`, `dispatch-failure` is
57
+ consumed only as NEGATIVE evidence: the gate veto in
58
+ `hooks/lib/review_gate_corroboration.py` / `hooks/pre_pr_review.py` (#1886;
59
+ `_dispatch_failure_verdict`, #1763) treats it as a reason to reject a
60
+ `.pr-review-passed` write, never as corroboration for one. A forged/hand-run
61
+ `dispatch-failure` event can only
62
+ ever cause a false rejection, never a false pass — the opposite forgery
63
+ direction from `record`, which is why this decision (unlike `record`) is
64
+ safely reachable from the CLI's closed `--event` vocabulary.
65
+
66
+ **#2188** adds `SubagentStop` as a `tool` value and `subagent_completion_guard.py`
67
+ as an emitter, reusing the existing `record` decision (not a new one) for the
68
+ same non-verdict, observational reason #1461 introduced it: this hook never
69
+ blocks/warns/bypasses/intervenes/reverts anything — it classifies a
70
+ subagent's own transcript tail (`clean` | `empty-final-turn` |
71
+ `truncated-final-turn` | `unreadable`) and writes one `record` row only for
72
+ the two non-clean, explainable outcomes, with `matched_rule` set to the
73
+ classification itself (`empty-final-turn` or `truncated-final-turn`); the
74
+ common `clean` case and the unexplainable `unreadable` case write nothing.
75
+
76
+ **#2166 Fix #3** adds `review_verdict_recorder.py` as a `SubagentStop`
77
+ emitter, same `record` decision, same non-verdict posture: once this hook
78
+ has confirmed a dispatch's `subagent_type` IS a registered review lens (so
79
+ the dispatch SHOULD produce `review-verdicts.jsonl` rows), a degenerate exit
80
+ that would otherwise be silently indistinguishable from a legitimate no-op
81
+ instead writes one `record` row naming why, via `matched_rule` of
82
+ `missing-scope-marker` (the dispatch prompt's Step 2.1 scope marker is
83
+ missing or reformatted) or `unparseable-result` (the agent's final JSON
84
+ result couldn't be recovered, even by the tolerant extractor). The
85
+ PRE-resolution exits (unreadable transcript, unresolvable `subagent_type`,
86
+ an unregistered or registered-but-non-review `subagent_type`) stay silent —
87
+ those are legitimate no-ops, not degenerate states.
88
+
89
+ | Field | Type | Values / source |
90
+ | --- | --- | --- |
91
+ | `ts` | string | ISO-8601 UTC `%Y-%m-%dT%H:%M:%SZ` |
92
+ | `hook` | string | Emitting hook's module name, e.g. `destructive_guard`, `verify_guard`, `pre_pr_review` (the review-corroboration gate, #1886; `pre_commit_review` is now a documented no-op and emits nothing), `telemetry`, `agent_dispatch_ledger` — or `code-review` for the CLI-emitted events (`--event doc-only`/`single-agent`/`dispatch-failure`), which carry the invoking skill's name rather than a hook module name |
93
+ | `tool` | string | Hooked tool/event: `Bash`, `Write`, `Edit`, `Skill`, `Agent`, `UserPromptSubmit`, `SubagentStop` (#2188) |
94
+ | `decision` | string enum | `block` \| `warn` \| `bypass` \| `intervention` \| `revert` \| `record` \| `dispatch-failure` |
95
+ | `matched_rule` | string | Rule ID from a closed vocabulary (pattern ID, hook-defined constant, bypass flag name, intervention keyword, or — for `record`/`dispatch-failure` — the dispatched review-agent's registered name, or — for `subagent_completion_guard.py`'s `record` rows — `empty-final-turn`/`truncated-final-turn`, #2188, or — for `review_verdict_recorder.py`'s `record` rows — `missing-scope-marker`/`unparseable-result`, #2166 Fix #3, or — for `boundary_events_write_guard.py`'s `block` rows — `ledger-write-blocked`, #2171) — never free text |
96
+ | `plugin_version` | string | From `.claude-plugin/plugin.json` |
97
+ | `session_id` | string, optional | Opaque per-session ID, when present in the hook payload — enables joins with `session-digest.jsonl` |
98
+ | `subject_hash` | string, optional | `review_gate_hash()` value (#1461) binding this event to the staged content it corroborates. A hex digest, not free text |
99
+ | `subject_hash_normalized` | string, optional | `normalized_gate_hash()` value (#1627) — the same binding computed after doc-hunk and indentation normalization. Stamped by `agent_dispatch_ledger.py` alongside `subject_hash`, and read by the gate's cosmetic-delta carry-forward lens. Absent on events written before #1627, which therefore never match on the normalized path. **The digest ALGORITHM has changed twice since** — #1638 (heredoc-body marking plus whole-file `--unified=100000` context) and #1660/#1661/#1662/#1663 (further grammars, a hunk-start guard, and a payload cap), each of which changes the computed VALUE for any changeset touching an affected extension. An event recorded by an earlier plugin version therefore never matches after an upgrade. This fails closed — one lost carry-forward, one extra dispatch — and is not a correctness bug, but it is why a version bump can look like a spurious re-review |
100
+
101
+ **`cosmetic-delta-carry-forward` (#1627, historical).** `pre_commit_review.py`
102
+ used to emit this `bypass`-decision event every time the OLD commit-time gate
103
+ passed a commit whose raw staged hash mismatched but whose normalized hash
104
+ matched. #1886 moved the gate to `gh pr create` and deliberately did NOT
105
+ carry this lens forward — the friction it existed to relieve (a whitespace-only
106
+ re-stage forcing a fresh review-agent dispatch before the NEXT commit) was a
107
+ direct consequence of gating every commit; a gate that fires once, at
108
+ PR-creation time, against the branch's cumulative diff, does not have that
109
+ problem. `hooks/pre_pr_review.py` never emits this event. Existing rows in
110
+ `boundary-events.jsonl` from before the migration remain valid history.
111
+
112
+ - **Emitter:** `hooks/lib/boundary_events.py::emit_boundary_event()`, called from `destructive_guard.py`, `verify_guard.py`, `pre_pr_review.py` (#1886), `telemetry.py` (intervention keywords), `agent_dispatch_ledger.py` (decision `record`, #1461), `subagent_completion_guard.py` (decision `record`, `tool` `SubagentStop`, #2188), `review_verdict_recorder.py` (decision `record`, `tool` `SubagentStop`, `matched_rule` `missing-scope-marker`\|`unparseable-result`, #2166 Fix #3), `boundary_events_write_guard.py` (decision `block`, `tool` `Write`\|`Edit`\|`Bash`, `matched_rule` `ledger-write-blocked` — the PreToolUse guard blocking a direct Write/Edit/Bash write to this same ledger, #2171), the mechanically-adopted guards (`pre_tool_guard.py`, `bash_retry_guard.py`, `refactor_test_freeze_guard.py`, `refactor_test_bash_guard.py`, `refactor_test_revert_guard.py` (decision `revert`, #906), `contract_version_guard.py`, `mutation_testing_smoke_gate.py`, `mutation_gate.py`, `tdd_guard.py`), and `boundary_events.py`'s own CLI (`--event dispatch-failure`, decision `dispatch-failure`, #1763) invoked from `skills/code-review/SKILL.md` Step 4. `--event gate-ran --verdict {allow,block,errored}` (decision `record`, `matched_rule` of `gate-ran-<verdict>`, #2037) is invoked from the repo-root `.husky/pre-commit` git hook — the real, git-native pre-commit gate (distinct from `pre_pr_review.py`, a Claude-Code-level PreToolUse hook gating `gh pr create`) — at every exit point, success or failure alike, so `${CLAUDE_PLUGIN_ROOT}/scripts/session_report.py --profile maintainer` can correlate a commit-attempt Bash record against a nearby `gate_ran` event and classify the previously-unmeasured "the gate silently never ran" population (`gate_ran_absent`) apart from a genuine internal failure (`gate_ran_errored`). This event carries no `session_id` in practice — a real git hook has no Claude Code session_id to attach — so correlation is by time proximity, not session join; see `session_report.py`'s "gate-run correlation (#2037)" section.
113
+ - **Consent:** ALWAYS-ON — not gated by `DEV_TEAM_TELEMETRY`. Local-only, rule-IDs-only safety/accountability channel; no observability holes by design.
114
+ - **Fail-open:** every exception in the emit helper is swallowed — never changes the calling hook's exit code, stdout, or stderr.
115
+ - **Consumers:** `skills/session-review/SKILL.md`, `skills/harness-audit/SKILL.md`, `agents/session-analysis.md`, `skills/cost-report/`, `skills/run-report/SKILL.md` (#1167), `hooks/lib/review_gate_corroboration.py` (#1461 `record` rows; #1763 also reads `dispatch-failure` rows as negative evidence for the gate veto), future `agent-telemetry` cross-machine aggregation (#178).
116
+
117
+ ---
118
+
119
+ ## `review-verdicts.jsonl`
120
+
121
+ **Added by #2166** (plan: `plans/2164-verdict-ledger-writer.md`, Slice 2). A
122
+ **new, separate** store from `boundary-events.jsonl` (Decision 1) — not an
123
+ overload of that stream's `record` decision — because a per-file verdict
124
+ needs a real `file_path`, which `boundary_events.py`'s own "never write free
125
+ text ... file paths ... must never appear" invariant forbids. Records, per
126
+ genuine review-agent dispatch, an outcome (`pass` \| `findings`) bound to
127
+ `(lens, file_path, file_content_hash)` — a verdict about *this exact file
128
+ content*, not about any one diff, so it can be looked up again the next time
129
+ the same content recurs regardless of which diff produced it.
130
+
131
+ The recorder identifies which lens dispatched via the native
132
+ `attributionAgent` field the harness stamps on the subagent's own transcript
133
+ records (`hooks/lib/cost_meter.py`'s "Attribution dimensions" mechanism,
134
+ reused via `scripts/lib/session_log.records`), falling back to the
135
+ documented Task/Agent-dispatch join only when that field is absent. It reads
136
+ the in-scope file list from a structured marker
137
+ (`skills/code-review/SKILL.md` step 4: `Files in scope for this review:
138
+ <path>, ...`) in the dispatch prompt — the subagent transcript's own first
139
+ turn — and cross-references it against the agent's final JSON result's
140
+ `issues[].file` list (`knowledge/review-agent-output-contract.md`).
141
+ **Disclosed trust boundary (Decision 4a):** the in-scope list is the
142
+ orchestrating session's own declared scope, not independently re-verified
143
+ against any diff — the property this store adds is that a *real*
144
+ `SubagentStop` event occurred for a *registered* review agent, not
145
+ omniscient verification of review depth.
146
+
147
+ | Field | Type | Values / source |
148
+ | --- | --- | --- |
149
+ | `ts` | string | ISO-8601 UTC `%Y-%m-%dT%H:%M:%SZ` |
150
+ | `lens` | string | The dispatched review agent's registered name (e.g. `structure-review`), plugin-prefix-stripped |
151
+ | `file_path` | string | One file the dispatch prompt's scope marker declared in scope, in its canonical form: `cwd`-relative POSIX (forward-slash) path, not the raw form the scope marker carried — falls back to an absolute resolved POSIX path only when the file can't be expressed relative to `cwd` |
152
+ | `file_content_hash` | string | sha256 hex digest of `file_path`'s content at the time the recorder ran (current content, not the content at dispatch time) |
153
+ | `outcome` | string enum | `pass` \| `findings` — whether `file_path` appears in the agent's final `issues[]` |
154
+ | `plugin_version` | string | From `.claude-plugin/plugin.json` |
155
+ | `session_id` | string, optional | Opaque per-session ID, when present in the hook payload |
156
+
157
+ - **Emitter:** `hooks/review_verdict_recorder.py` (a `SubagentStop` hook) via `hooks/lib/review_verdicts.emit_review_verdict()`. No-op (zero rows) for any `subagent_type` outside `hooks/lib/review_agent_registry`'s closed set of registered `agents/*-review.md` names, and fail-open throughout (missing/unreadable transcript, unresolved `subagent_type`, a missing/reformatted scope marker, or an unparseable final JSON result all degrade to zero rows, never an exception); a single deleted/unreadable in-scope file is skipped without affecting the other rows.
158
+ - **Consent:** ALWAYS-ON — same posture as `boundary-events.jsonl` (Decision 2), not gated by `DEV_TEAM_TELEMETRY`/`~/.claude/telemetry.json`. Local-only, mechanical accountability data (lens/path/hash/outcome), no prose.
159
+ - **Fail-open:** every exception in `emit_review_verdict()` is swallowed — never changes the calling hook's exit code, stdout, or stderr. `hooks/lib/review_verdicts.load_verdicts()` mirrors this on the read side: an absent file, a corrupted line, or a stale `plugin_version` row all degrade to "no usable rows", never an exception.
160
+ - **Consumers:** none yet — this slice is deliberately writer-only (#2167 is the queued consumer slice).
161
+
162
+ ---
163
+
164
+ ## `telemetry.jsonl`
165
+
166
+ Opt-in usage beacon: which slash commands / skills get invoked, and whether
167
+ the pre-commit review gate fired or was bypassed.
168
+
169
+ | Field | Type | Values / source |
170
+ | --- | --- | --- |
171
+ | `ts` | string | ISO-8601 UTC |
172
+ | `event` | string enum | `command` \| `skill` \| `gate` |
173
+ | `name` | string | Grammar-matched slash-command name, skill name, or `pre-pr-review` (#1886 — the gate moved from `git commit` to `gh pr create`) |
174
+ | `outcome` | string | `invoked` \| `fired` \| `bypassed` |
175
+ | `plugin_version` | string | From `.claude-plugin/plugin.json` |
176
+
177
+ - **Emitter:** `hooks/telemetry.py::_emit()`. Written to `~/.claude/metrics/telemetry.jsonl` — home-scoped, out of the project entirely (#1405/#1406), never a project's own `metrics/`.
178
+ - **Consent:** opt-in — `~/.claude/telemetry.json` `{"enabled": true}`, home-scoped only. `DEV_TEAM_TELEMETRY` and a project-scoped `<cwd>/.claude/telemetry.json` are now inert (one-time-per-session stderr notice only, no effect on consent). Off by default; nothing recorded, nothing leaves the machine.
179
+ - **Consumers:** `skills/telemetry/SKILL.md`, `plugins/dev-team/scripts/session_report.py`.
180
+
181
+ ---
182
+
183
+ ## `cost-metering.jsonl`
184
+
185
+ Per-session token/cost summary, incrementally accumulated from the
186
+ transcript on each `Stop` hook fire.
187
+
188
+ | Field | Type | Values / source |
189
+ | --- | --- | --- |
190
+ | `timestamp` | string | ISO-8601 UTC |
191
+ | `transcript` | string | Transcript file basename (not full path) |
192
+ | `total` | object | Aggregated token counts + `cost_usd` + `messages` across the session |
193
+ | `by_model` | object | Per-model slim breakdown: `cost_usd`, `input_tokens`, `output_tokens` |
194
+ | `by_thread` | object | Per-thread slim breakdown, same shape as `by_model` |
195
+ | `by_agent_type` | object | Per-agent-type slim breakdown, same shape as `by_model`: `main` for main-loop turns; sidechain turns keyed by subagent type via `attributionAgent` or the Task-dispatch join; honest `unattributed` bucket when neither signal exists (#1094) |
196
+
197
+ - **Emitter:** `hooks/cost_meter.py` (wrapper) → `hooks/lib/cost_meter.py::cmd_record()`.
198
+ - **Consent:** gated by `telemetry_consent.is_enabled()` (`~/.claude/telemetry.json` `{"enabled": true}`, home-scoped) — no longer unconditional as of Slice 2 (#1406).
199
+ - **Consumers:** `skills/cost-report/SKILL.md`, `skills/harness-audit/SKILL.md`, `cmd_regression`/`cmd_pace` in the same library, `skills/run-report/SKILL.md` (#1167, best-effort only — see that skill's Join limitations section: this stream has no `session_id` field).
200
+
201
+ ---
202
+
203
+ ## `phase-markers.jsonl`
204
+
205
+ Per-phase context-pollution markers (#1520), one row appended at each `/handoff`
206
+ (a phase boundary). Distinct stream from `cost-metering.jsonl` — deliberately
207
+ kept out of that log's incremental `record` state so this additive dimension
208
+ never touches the security-sensitive hot path.
209
+
210
+ | Field | Type | Values / source |
211
+ | --- | --- | --- |
212
+ | `timestamp` | string | ISO-8601 UTC |
213
+ | `transcript` | string | Transcript file basename (not full path) |
214
+ | `phase` | string | Phase label — the first `/handoff` args token when sane, else `handoff` (or `unlabeled` for a direct library call) |
215
+ | `resident_tokens` | int | Main-loop context occupancy at the boundary: the most-recent non-sidechain turn's `input + cache_read + cache_creation` |
216
+ | `spent_output_cumulative` | int | Cumulative main-loop `output_tokens` across the session up to this boundary (monotonic; `phase-report` deltas it into per-phase spend) |
217
+
218
+ - **Emitter:** `hooks/phase_marker.py` (PostToolUse:Skill, filters to `handoff`) → `hooks/lib/cost_meter.py::cmd_phase_mark()`.
219
+ - **Consent:** gated by `telemetry_consent.is_enabled()`; shares the cost meter's `DEV_TEAM_COST_METER=off` opt-out.
220
+ - **Consumers:** `skills/cost-report/SKILL.md` (§ Context pollution, via `phase-report`), `skills/harness-audit/SKILL.md` (§ Analyze orchestration complexity).
221
+
222
+ ---
223
+
224
+ ## `artifact-usage.json`
225
+
226
+ Not JSONL — a single JSON object keyed by skill/agent name, upserted on
227
+ every invocation.
228
+
229
+ | Field | Type | Values / source |
230
+ | --- | --- | --- |
231
+ | `<skill_name>.use_count` | integer | Cumulative invocation count |
232
+ | `<skill_name>.last_used_at` | string | ISO-8601 UTC of the most recent invocation |
233
+ | `<skill_name>.lifecycle` | string | `active` (set on creation; other lifecycle states are assigned externally by `/artifact-lifecycle`) |
234
+
235
+ - **Emitter:** `hooks/telemetry.py::_upsert_artifact_usage()` (atomic rewrite via tempfile + `os.replace`). Written to `~/.claude/metrics/artifact-usage.json` — home-scoped, out of the project entirely (#1405/#1406), never a project's own `metrics/`.
236
+ - **Consent:** follows `telemetry.jsonl`'s opt-in gate (`~/.claude/telemetry.json` `{"enabled": true}`, home-scoped). The project-scoped explicit-off switch this section used to document (a project-level `.claude/telemetry.json` `{"enabled": false}` disabling usage tracking specifically) no longer exists — project-scoped `.claude/telemetry.json` is inert entirely, same as `telemetry.jsonl`'s row above.
237
+ - **Consumers:** `skills/artifact-lifecycle/SKILL.md`.
238
+
239
+ ---
240
+
241
+ ## `gate-bypass-audit.jsonl`
242
+
243
+ Accountability record for a bypass of the review-corroboration gate. #1886
244
+ moved the gate from `git commit` to `gh pr create`; this stream now carries
245
+ `hooks/pre_pr_review.py`'s `PR_GATE_BYPASS_REASON` bypasses.
246
+ `hooks/pre_commit_review.py` is now a documented no-op and no longer writes
247
+ to this stream — historical rows from before the migration
248
+ (`triggeredBy: "--no-verify"`/`"-n"`) remain valid history but no new ones
249
+ are produced.
250
+
251
+ | Field | Type | Values / source |
252
+ | --- | --- | --- |
253
+ | `timestamp` | string | ISO-8601 UTC |
254
+ | `branch` | string | Current git branch |
255
+ | `triggeredBy` | string | `PR_GATE_BYPASS_REASON` (current); `--no-verify`/`-n` (historical, pre-#1886) |
256
+ | `reason` | string | Value of `PR_GATE_BYPASS_REASON` (current) / `GATE_BYPASS_REASON` (historical) — human/agent-authored, required to be non-empty |
257
+ | `stagedFileCount` | integer | Count of files in the branch diff at bypass time (historical rows: staged files at commit time) |
258
+ | `pluginVersion` | string | From `.claude-plugin/plugin.json` |
259
+
260
+ - **Emitter:** `hooks/pre_pr_review.py::_record_bypass_audit()` (#1886). `hooks/pre_commit_review.py::_record_bypass_audit()` was the historical emitter, now removed along with the rest of that module's gating logic.
261
+ - **Consent:** unconditional — accountability record for an actively-chosen bypass, not passive usage telemetry.
262
+ - **Consumers:** `skills/code-review/SKILL.md`, `docs/code-review-process.md`.
263
+
264
+ ---
265
+
266
+ ## `gate-bypass.jsonl`
267
+
268
+ Accountability record for `MUTATION_SMOKE_GATE_SKIP=1` bypasses of the
269
+ mutation-testing smoke gate. Distinct stream from `gate-bypass-audit.jsonl`
270
+ above (different gate, different hook).
271
+
272
+ | Field | Type | Values / source |
273
+ | --- | --- | --- |
274
+ | `timestamp` | string | ISO-8601 UTC (`Z`-suffixed) |
275
+ | `hook` | string | Always `mutation-testing-smoke-gate` |
276
+ | `command_hash` | string | First 16 hex chars of `sha256(raw_command)` — the raw command is never logged |
277
+ | `cwd` | string | Payload cwd |
278
+
279
+ - **Emitter:** `hooks/mutation_testing_smoke_gate.py::log_bypass_audit()`.
280
+ - **Consent:** unconditional.
281
+ - **Consumers:** `skills/mutation-testing/SKILL.md`.
282
+
283
+ ---
284
+
285
+ ## `config-changelog.jsonl`
286
+
287
+ Audit trail for `/feedback-learning` config changes and human-oversight
288
+ protocol events (approval / override / pause / stop).
289
+
290
+ | Field | Type | Values / source |
291
+ | --- | --- | --- |
292
+ | `timestamp` | string | ISO-8601 UTC |
293
+ | `type` | string enum | `amend` \| `approval` \| `override` \| `pause` \| `stop` (feedback-learning change types, or oversight event types) |
294
+ | `trigger` | string | `user` (who/what triggered the change) |
295
+ | `description` | string | Human/agent-authored summary of what happened and why (deliberate audit note, not incidental free text) |
296
+ | `file_modified` | string, optional | Config file touched (feedback-learning changes) |
297
+ | `section_modified` | string, optional | Section within the file |
298
+ | `previous_value` / `new_value` | string, optional | Before/after values |
299
+ | `approved_by` | string, optional | Who approved the change |
300
+
301
+ - **Emitter:** `/feedback-learning` skill (model-authored append) and `/human-oversight-protocol` skill.
302
+ - **Consent:** unconditional (append-only governance record).
303
+ - **Consumers:** `skills/feedback-learning/SKILL.md`, `skills/human-oversight-protocol/SKILL.md`, `skills/governance-compliance/SKILL.md`.
304
+
305
+ ---
306
+
307
+ ## `session-digest.jsonl`
308
+
309
+ Trend digest from `/session-review` (backed by `${CLAUDE_PLUGIN_ROOT}/scripts/session_report.py --profile maintainer`):
310
+ aggregate counts only, no file names, prompts, command strings, or code.
311
+
312
+ | Field | Type | Values / source |
313
+ | --- | --- | --- |
314
+ | `recorded_at` | string | UTC ISO-8601 of the run |
315
+ | `plugin_version` | string | The `dev-team` plugin's `.claude-plugin/plugin.json` version active when this record was produced (`"unknown"` if the manifest couldn't be read), #1471. Lets consumers tell a friction already fixed in a newer version apart from one that's still current — see `--version-scope` below |
316
+ | `sessions`, `transcripts` | integer | How many sessions/transcripts the digest covered |
317
+ | `tokens` | object | Input/output/cache token totals |
318
+ | `cost_usd`, `cache_hit_ratio` | number | Session cost and cache-read efficiency |
319
+ | `rework` | object | `failed_edits`, `repeated_file_edits`, `retried_bash_commands`, `retried_bash_commands_by_skill`/`retried_bash_commands_by_agent` (#2110 — each retry attributed, at the moment it's detected, to whichever skill/agent is sticky-active; `retried_bash_commands` is the derived sum, never a second independent count), `repeated_verify_runs`, `permission_denials`, `compaction_events` |
320
+ | `accuracy` | object | `tool_calls`, `tool_error_rate`, `user_correction_turns`, `by_skill`/`by_agent` (correction counts, double-bucketed against whichever skill/agent is sticky-active), `correction_rate_by_skill`/`correction_rate_by_agent` (corrections per skill invocation / agent dispatch — absent for a never-invoked name, never a misleading `0.0`), `correction_causes` (#2013, below) |
321
+ | `utilization` | object | `skills_invoked`, `agents_invoked` (agent RUNS), `agent_dispatches` (Agent/Task tool calls), `never_observed_skills`, `never_observed_agents` |
322
+
323
+ **`accuracy.correction_causes` (#2013).** Deterministic cause data for every
324
+ detected correction turn — no model call, see
325
+ `plugins/dev-team/scripts/lib/session_log/corrections.py`'s module
326
+ docstring for the classifier. Present in BOTH `session-digest/v4` and
327
+ `downstream-session-report/v4` (a correction's producing component is
328
+ exactly the "which components generate the most corrections per dispatch"
329
+ question the issue exists to answer, and that question is as live for a
330
+ downstream user's own report as for this repo's own trend stream). Shape:
331
+
332
+ | Field | Type | Values |
333
+ | --- | --- | --- |
334
+ | `by_what` | object | Counts by `code-edit` \| `plan` \| `review-finding` \| `tool-choice` \| `factual-claim` \| `other` |
335
+ | `by_component` | object | Counts by `main-loop` or the single most-recently-dispatched skill/agent name (never double-bucketed, unlike `accuracy.by_skill`/`by_agent`) |
336
+ | `by_shape` | object | Counts by `reverted` \| `redirected` \| `narrowed-scope` \| `flagged-wrong` \| `not-what-asked` \| `ambiguous` |
337
+ | `ambiguous_share` | number | `by_shape["ambiguous"] / user_correction_turns` — the classifier's honest inference-share statistic, never hidden |
338
+
339
+ Never the correction text itself — only these four closed-vocabulary labels.
340
+ `test_session_log_corrections.py::test_classify_correction_never_leaks_correction_text`
341
+ pins this the same way
342
+ `test_session_report_golden.py::test_no_sentinel_leaks_in_either_golden`
343
+ pins the rest of this stream.
344
+
345
+ `session-digest/v2` (#1994) counts dispatched agents' own transcripts for the
346
+ first time, so token/tool-call/rework totals jump against v1, and
347
+ `retried_bash_commands` / `repeated_verify_runs` moved from a session-keyed to
348
+ a per-thread basis. Records from the two eras are not comparable; split on
349
+ `schema` before trending. `session-digest/v3` (#2046) is a schema-label-only
350
+ bump: it is emitted by the new unified `session_report.py --profile
351
+ maintainer` entry point, which every real consumer now runs (#2047) instead
352
+ of the earlier monorepo-only extractor (retired in #2048). No data-shape
353
+ change from v2 — a v2 and a v3 record trend together. `session-digest/v4`
354
+ (#2018) changes what `plugin_version` MEANS on the per-session
355
+ `session-sync/v4` records `--sync-out` writes (see "Version tagging" below)
356
+ — a v3 sync record and a v4 sync record are not comparable on that field,
357
+ though every other field is unchanged and they still trend together
358
+ otherwise. The single-shot digest/trend record keeps its v2/v3 meaning for
359
+ `plugin_version` (extraction-time) even under the v4 label — see below.
360
+
361
+ - **Emitter:** `/session-review` skill via `${CLAUDE_PLUGIN_ROOT}/scripts/session_report.py --profile maintainer`.
362
+ - **Consent:** unconditional (aggregate counts only, no file/prompt/command content).
363
+ - **Enforcement (#2045):** every name/label-shaped field this stream (and the
364
+ shipped `session_report.py --profile downstream` report) emits passes
365
+ through `plugins/dev-team/scripts/lib/session_log/redact.redact()` — the
366
+ one function both extractors route file basenames, project labels, skill
367
+ names, agent names, and model ids through before writing them out.
368
+ Previously this line's promise was a convention restated independently at
369
+ each call site; `redact()` is the single enforcement point, pinned by
370
+ `tests/scripts/test_session_report_golden.py::test_no_sentinel_leaks_in_either_golden`
371
+ against a corpus seeding real prompt text, source code, a full shell
372
+ command string, and absolute POSIX/Windows paths.
373
+ - **Consumers:** `skills/harness-audit/SKILL.md` (joins with self-reported task logs), `agents/session-analysis.md`.
374
+ - **Version tagging (#1471, revised #2018):** `plugin_version` is also carried on the per-session `session-sync/v1`-`v4` records synced by `--sync-out` (used by `--rollup`/`--escalate`/`--correlate`) and on the single-shot digest itself. These are no longer the same value stamped two ways — as of `session-sync/v4` (#2018) the two paths diverge deliberately:
375
+ - **Per-session `--sync-out` records (`session-sync/v4`)** — the path that runs unattended at every `SessionStart` (`.claude/ensure_session_archive.py`) and accumulates into the durable, cross-release archive — carry the SESSION's OWN version, resolved by `resolve_session_plugin_version()` from that project's `<cwd>/.claude/metrics/boundary-events.jsonl` (a stream stamped live, at hook-dispatch time, by `hooks/lib/boundary_events.py`). When a session's boundary-events carry more than one plugin_version (e.g. an in-session `/dev-team:upgrade`), the EARLIEST by `ts` wins — documented, not solved, since a session spanning two releases has no single correct answer. A session with no matching boundary-events record (never dispatched anything through a hook that stamps `session_id`) is tagged the explicit string `"unknown"` — never silently the extractor's own checked-out version, which is the exact mislabeling #2018 was filed to close (5,562 archived records, all `plugin_version: null`, none of them attributable to a release).
376
+ - **The single-shot digest/trend record** (`--transcript`/`--project-dir`, no `--sync-out`) is UNCHANGED behavior: it still reflects the plugin version active on the machine *at extraction time*. It can aggregate many sessions into one record, and a single field cannot honestly attribute a multi-session aggregate to one session's version — this is a deliberate scope limitation, not an oversight. In the common case (`/session-review` run against the current project, right after the session it's summarizing) extraction-time and session-time coincide anyway; the divergence #2018 closes is specific to `--sync-out`'s day/weeks-later, unattended, cross-session runs.
377
+ - **Version scoping (#1480):** `session_report.py --profile maintainer --rollup`/`--escalate`/`--correlate` accept `--version-scope {all,current-and-previous}` (default `all`, unbounded history). `current-and-previous` drops any record whose `plugin_version` isn't the currently-installed version or the version immediately before it *as observed in the digests being read* (there is no release-history lookup) — records with no `plugin_version` at all (pre-#1471 data), or tagged the explicit `"unknown"` string (#2018), are dropped too, since neither can be proven current. The result gains a `version_window` field (the concrete versions included; `[]` when unscoped). `/session-review` defaults to local-only; its `--cross-machine` opt-in always applies `current-and-previous` scoping (see its SKILL.md) — the skill itself never exposes an unscoped cross-machine mode. Unbounded history across every version remains available only via a direct `session_report.py --profile maintainer --rollup ... --version-scope all` invocation (the CLI default), outside the skill.
378
+ - **Version-filtered downstream report coverage (#2018):** `session_report.py --profile downstream --plugin-version VERSION` scopes the report to sessions whose project recorded `VERSION` in `boundary-events.jsonl` (`sessions_matching_plugin_version`, best-effort, same source as above) — sessions with no matching event are excluded from the report's own counts exactly as before. What's new: the report's top-level `version_filter_coverage` field (present, non-null, only when `--plugin-version` was passed) states `requested_version`, `sessions_considered` (every distinct session in the report's own since/until window, ignoring the version filter), `sessions_attributed` (how many of those matched), `sessions_attributed_other_version` (recorded a DIFFERENT known `plugin_version` — `_sessions_with_known_plugin_version` — the filter correctly excluding them, not a data gap), and `sessions_unattributed` (no resolvable `plugin_version` at all — genuinely missing data) — so an operator can tell "belongs to another release" apart from "this repo's own instrumentation never saw it," instead of one conflated exclusion count. The exclusion behavior itself is unchanged; this only makes it observable and precise.
379
+
380
+ ---
381
+
382
+ ## `review-value.jsonl`
383
+
384
+ Whether a `/build` inline review checkpoint actually changed anything —
385
+ counts and outcomes only, never code or file content.
386
+
387
+ | Field | Type | Values / source |
388
+ | --- | --- | --- |
389
+ | `timestamp` | string | ISO-8601 UTC |
390
+ | `plan` | string | Plan file path |
391
+ | `slice` | string | Slice number |
392
+ | `step` | string | Step number (`N.M`) or `all` |
393
+ | `checkpoint` | string enum | `step` \| `slice` \| `backstop` (the Step-6 backstop pass, #1962) |
394
+ | `complexity` | string enum | `standard` \| `complex` |
395
+ | `agents_run` | array of string | Review agents dispatched |
396
+ | `issues_found`, `issues_fixed`, `fix_iterations` | integer | Counts |
397
+ | `severity_breakdown` | object | `{errors, warnings, suggestions}` counts (same enum as `/code-review`); the three sum to `issues_found`. Lets `/harness-audit` Step 3 flag mostly-minor lenses (#1256). Absent on pre-#1256 rows |
398
+ | `source` | string enum | Row provenance: `build-checkpoint` (fix-applying `/build` inline checkpoint) \| `build-backstop` (fix-applying `/build` Step-6 backstop pass, #1962) \| `code-review` (read-only standalone review) \| `harness` (the `evals/code-review-benchmark/` replay harness — #2051, see below). **Absent = `build-checkpoint`** (back-compat). `/harness-audit` Step 4 excludes `code-review` rows from fix-rate drop-candidate logic (#1257); `build-backstop` rows are fix-applying and stay in it |
399
+ | `diff_shape` | string enum | Shape of the reviewed diff: `test-only` (every changed file provably a test per `knowledge/test-file-indicators.md`) \| `mixed` (anything else). Classified by `skills/code-review/scripts/change_shape.py`'s `isTestOnly`, never by eye; include-biased, so `test-only` is never over-claimed. Lets `/harness-audit` split per-lens outcomes by diff shape — the evidence a test-only lens gate waits on (#1964). Absent on pre-#1964 rows |
400
+ | `outcome` | string enum | `no-op` \| `fixed` \| `escalated` \| `skipped` (backstop only — suppressed by `--backstop-review=skip`; never counted in a rate, #1962) |
401
+
402
+ - **Emitter:** `/build` skill (model-authored append, sub-step 7) writes `source: "build-checkpoint"` for inline checkpoints and, from Step 6, `source: "build-backstop"` for the backstop pass (#1962). Disable with `DEV_TEAM_REVIEW_VALUE=off`.
403
+ - **Consent:** unconditional when enabled (no code/file content recorded).
404
+ - **Consumers:** `skills/cost-report/SKILL.md`, `skills/harness-audit/SKILL.md`.
405
+ - **Provenance (#1257):** fix-rate ROI is only meaningful for fix-applying rows. A read-only review that never applies fixes (`source: "code-review"`) always has `issues_fixed: 0`; Step 4 must not read that as a zero-value drop candidate — it reports finding-rate for those instead.
406
+
407
+ ### Round rows — `source: "code-review"` with a `round` field (#1624)
408
+
409
+ `/code-review` appends one row **per dispatch round** (round 1 = the initial
410
+ panel; each fix-loop iteration's re-dispatch set is one further round), so
411
+
412
+ # 1623's "is this agent's dispatch frequency value or churn?" becomes
413
+
414
+ answerable. Written by `skills/code-review/scripts/review_round_log.py`.
415
+
416
+ | Field | Type | Values / source |
417
+ | --- | --- | --- |
418
+ | `timestamp` | string | ISO-8601 UTC |
419
+ | `source` | string enum | Always `code-review` for these rows |
420
+ | `round` | integer | 1 = initial panel; each fix-loop re-dispatch set increments. **Presence of this field is what distinguishes a round row** from the original whole-run `code-review` row |
421
+ | `agents_run` | array of string | Registered review agents dispatched this round (sorted, deduped) |
422
+ | `findings_new` | integer | Findings whose signature was not present in a prior round (signature identity: #1625). The round-row analogue of `issues_found` |
423
+ | `findings_carried` | integer | Prior-round signatures that survived this round's fix attempt |
424
+ | `severity_breakdown` | object | `{errors, warnings, suggestions}` over `findings_new`; same enum as the `/build` rows |
425
+ | `fix_provenance_new` | integer | How many of `findings_new` fall inside the line ranges the **previous** round's fix touched — the judgment-free "the fix introduced it" signal. Computed by unified-diff interval math, never by model judgment. Always `0` for `round: 1` (no preceding fix) |
426
+ | `dispatch_purpose` | string enum | `discovery` (a panel looking for new problems) \| `verification` (confirming a specific fix, #1628) \| `closing` (the scoped gate-closing pass, #1626) |
427
+ | `outcome` | string enum | `no-op` \| `fixed` \| `escalated` — same enum as the `/build` rows |
428
+
429
+ - **Emitter:** `/code-review` (steps 5b-i and 6a) via `review_round_log.py`.
430
+ - **Consent:** **unconditional** — written to `.claude/metrics/` like `boundary-events.jsonl`, *not* gated behind `~/.claude/telemetry.json` the way `/build`'s rows are. Rationale (#1624 design item 2): this is the same class of local, counts-only operational stream the commit gate itself already depends on, and consent-gating it would make #1623's success criteria depend on consent being enabled per dev machine. Rows carry counts, agent names, and enum values only — no file paths, code, or finding text.
431
+ - **Consumers:** `skills/harness-audit/SKILL.md` Step 4a (churn ratio, per-agent discovery-vs-verification split, gate recidivism).
432
+ - **Backstop rows (#1962).** `source: "build-backstop"` marks `/build`'s Step-6 pass — the one review layer whose files an inline checkpoint already reviewed in the same run. It is fix-applying (the `--internal` panel runs the review-fix loop), so it belongs in fix-rate analysis alongside `build-checkpoint`; what it exists to answer is whether that duplicated layer is ~all `no-op`, which is the evidence `/build`'s `--backstop-review=skip` flag waits on. `outcome: "skipped"` marks a backstop suppressed by that flag: recorded so the suppression is visible in the same stream, and excluded from every rate because it never ran.
433
+ - **Reconciling the `source` values.** `build-checkpoint` and `build-backstop` rows are fix-applying and carry `plan`/`slice`/`step`/`checkpoint`/`complexity`/`issues_found`/`issues_fixed`/`fix_iterations`. `code-review` rows are read-only; those with a `round` field use the round schema above. A consumer wanting "how many issues did this row surface" should read `(.issues_found // .findings_new)`, which covers all three shapes.
434
+
435
+ ### Harness rows — `source: "harness"` (#2051)
436
+
437
+ Written by `evals/code-review-benchmark/runner.emit_review_value_rows()` —
438
+ the `/code-review` **replay harness** (#821), not a live session. One row
439
+ per lens per dispatch, written straight from the parsed `/code-review
440
+ --json` payload's `agents[]` list — by mechanism, not by agent instruction
441
+ — so every dispatched lens gets a row regardless of outcome, including a
442
+ lens that found nothing. This is the fix for the collection bias #2019/
443
+ #1512 documented in the live writers above.
444
+
445
+ | Field | Type | Values / source |
446
+ | --- | --- | --- |
447
+ | `agents_run` | array of string | Always exactly one lens name — one row per lens, not per dispatch batch |
448
+ | `issues_found` | integer | Count of that lens's issues this dispatch |
449
+ | `severity_breakdown` | object | `{errors, warnings, suggestions}`, same enum as the live rows |
450
+ | `outcome` | string | Always `"no-op"` — the harness is read-only and never applies a fix |
451
+ | `diff_shape` | string enum | `test-only` \| `mixed`, same classifier as the live rows; the recorded-diff adapter (below) is what actually supplies real `test-only` cases |
452
+ | `dataset` | string | `defects4j` \| `bugsjs` \| `recorded-diff` |
453
+ | `project`, `bug_id` | string | The benchmark case's identifiers (not a `/build` plan/slice/step — a structurally different key space) |
454
+
455
+ - **Emitter:** `runner.emit_review_value_rows()`, called from `runner.run_case()` and `runner.run_recorded_diff_case()` after a dispatch's `--json` payload parses successfully. Never called for an unparseable dispatch — that failure is already captured by the harness's own `skipped.jsonl`.
456
+ - **File location, deliberately not `.claude/metrics/`.** Rows land in the harness's own results directory (`evals/code-review-benchmark/results/review-value.jsonl` by default, `--results-dir` elsewhere) — never the live metrics tree. This is a structural guarantee against pooling, on top of the `source: "harness"` label itself: even a caller reading the wrong file could not accidentally merge harness rows into the live population, because they are not in the same file.
457
+ - **Consumers:** `skills/harness-audit/SKILL.md` §4b, which reads this stream as a separate, explicitly-labelled population and must never merge it into Step 3/4's live-row computations.
458
+ - **Recorded-diff adapter (#2051).** `evals/code-review-benchmark/adapters/recorded_diff_adapter.py` supplies `dataset: "recorded-diff"` cases from saved diffs (most usefully real `/test-improve` Phase-5 diffs) — the only source that can give this stream a genuine `diff_shape: "test-only"` row, since Defects4J/BugsJS are real production bug fixes and structurally cannot be test-only.
459
+
460
+ ---
461
+
462
+ ## `contract-failures.jsonl`
463
+
464
+ Diagnostic record for a review-agent output that fails the shared JSON
465
+ contract (`knowledge/review-agent-output-contract.md`) — the gap #1998
466
+ closes. Session-report analysis found 18.2% of review-agent outputs
467
+ discarded silently, with no record of which agent, what it returned, or
468
+ why it didn't parse; today's alternative to this stream is nothing.
469
+
470
+ | Field | Type | Values / source |
471
+ | --- | --- | --- |
472
+ | `timestamp` | string | ISO-8601 UTC |
473
+ | `agent` | string | Name of the review agent whose output failed validation |
474
+ | `shape` | string enum | `empty` \| `truncated` \| `malformed-json` \| `schema-drift` \| `not-json` — the closed set `validate_review_output.FAILURE_SHAPES` exports. `validate_review_output.py` also recognizes `clean`/`fenced`/`prose-preamble` (`SUCCESS_SHAPES`), but those three name *successful* extraction (the JSON was found and matched the contract) and so never appear as `shape` in a failure row — a successfully-extracted object that then fails schema validation is logged as `schema-drift`, not the extraction shape that found it. `malformed-json` is distinct from `truncated`: a balanced `{...}` object (unquoted keys, a trailing comma, a Python-repr dict) that still fails to parse is `malformed-json`; a `{` that never balances back to depth zero before EOF is `truncated` |
475
+ | `extraction` | string enum, nullable | Set whenever a JSON-shaped candidate was actually recovered before failing — i.e. `shape` is `schema-drift` or `malformed-json` — naming which of `clean`/`fenced`/`prose-preamble` recovered it, so that information survives the downgrade instead of being discarded. `null` for `empty`/`truncated`/`not-json` — no candidate was ever recovered for those |
476
+ | `error` | string | The specific validation/parse error (e.g. a `JSONDecodeError` message, or `status='ok' not one of [...]`), after the same secret-redaction pass as `raw_prefix` and capped at 256 characters — `_validate_schema` interpolates agent-controlled values into this string, so it needs the same two controls |
477
+ | `raw_prefix` | string | First 200 characters of the agent's raw output, **after** a secret-redaction pass (`validate_review_output._redact()`: this repo's canonical hardcoded-key pattern from `knowledge/owasp-detection.md`, plus common vendor token prefixes). Deliberate, capped exception to this file's "never incidental free text" default (mirroring `gate-bypass-audit.jsonl`'s `reason` field) — without seeing what was actually returned, the failure shapes #1998 exists to classify cannot be told apart. AI-authored review-agent output, which may quote repository source verbatim — including any secret present in the reviewed diff, since a lens's own job is to find and quote such things — so this is a transitive channel for repo content, not a claim that the text is free of it; the redaction pass is the actual control, the 200-char cap only bounds volume |
478
+
479
+ - **Emitter:** `skills/code-review/scripts/validate_review_output.py::log_failure()`, called once per non-contract-valid agent result during `/code-review` step 4's dispatch-failure handling.
480
+ - **Consent:** unconditional, matching `boundary-events.jsonl` — this is the same class of local, counts-and-diagnostics operational stream the commit gate itself already depends on.
481
+ - **Consumers:** `skills/code-review/scripts/contract_failure_report.py`, which joins this stream against `boundary-events.jsonl`'s dispatch counts to report a real per-agent failure *rate* (not just a count) for #1980/#1982 to read before citing any `$/finding` figure.
482
+
483
+ ---
484
+
485
+ ## `verify-log.jsonl`
486
+
487
+ Evidence that the project's own test/verification tooling actually exercised
488
+ the change end-to-end (or was legitimately skipped) before a `/build` slice
489
+ with a runtime surface was marked complete. Schema modeled on
490
+ `review-value.jsonl`.
491
+
492
+ | Field | Type | Values / source |
493
+ | --- | --- | --- |
494
+ | `timestamp` | string | ISO-8601 UTC |
495
+ | `plan` | string | Plan file path |
496
+ | `slice` | string | Slice number |
497
+ | `branch` | string | Current git branch |
498
+ | `files` | array of string | Changed runtime files in scope |
499
+ | `outcome` | string enum | `ran` \| `skipped` \| `failed-then-fixed` |
500
+ | `reason` | string, optional | Set when `outcome` is `skipped` (e.g. `"tests-only"`, `"docs-only"`) |
501
+
502
+ - **Emitter:** `/build` skill (model-authored append, sub-step 4.9).
503
+ - **Consent:** unconditional.
504
+ - **Consumers:** `${CLAUDE_PLUGIN_ROOT}/scripts/progress_guardian.py --pre-pr` (fails closed on a runtime-surface change with no matching entry), `skills/performance-metrics/SKILL.md`.
505
+
506
+ ---
507
+
508
+ ## `override-audit.jsonl`
509
+
510
+ Audit trail for `/code-review --force --reason "<text>"`, which skips all
511
+ gates and the documentation-only short-circuit.
512
+
513
+ | Field | Type | Values / source |
514
+ | --- | --- | --- |
515
+ | `timestamp` | string | ISO-8601 |
516
+ | `branch` | string | Current git branch |
517
+ | `triggeredBy` | string | Always `--force` |
518
+ | `reason` | string | Value of `--reason` (required, human/agent-authored) |
519
+ | `targetFiles` | array of string | Files the forced review targeted |
520
+ | `gatesSkipped` | array of string | e.g. `["lint", "type-check", "secret-scan", "semgrep", "pipeline-red"]` |
521
+
522
+ - **Emitter:** `/code-review` skill (model-authored append, step 2).
523
+ - **Consent:** unconditional.
524
+ - **Consumers:** `skills/code-review/SKILL.md`, `docs/code-review-process.md`.
525
+
526
+ ---
527
+
528
+ ## `eval-variance.jsonl`
529
+
530
+ Multi-trial pass@k stability trend for `/agent-eval` fixtures.
531
+
532
+ | Field | Type | Values / source |
533
+ | --- | --- | --- |
534
+ | `recorded_at` | string | ISO-8601 UTC |
535
+ | `schema` | string | `eval-variance/v1` |
536
+ | `trials` | integer | Number of trials in this run |
537
+ | `pairs_evaluated` | integer | Fixture/agent pairs evaluated |
538
+ | `flaky_count` | integer | Pairs that neither always passed nor always failed |
539
+ | `mean_pass_at_k` | number | Mean pass@k across evaluated agents |
540
+
541
+ - **Emitter:** `scripts/eval_variance.py --append`.
542
+ - **Consent:** unconditional (eval infra, not user-session telemetry).
543
+ - **Consumers:** `skills/agent-eval/SKILL.md`.
544
+
545
+ ---
546
+
547
+ ## `eval-ablation.jsonl`
548
+
549
+ Causal per-agent ablation evidence from `/agent-eval --ablation <agent>` (#868):
550
+ a controlled baseline-vs-ablated integration-tier delta (issues caught,
551
+ `testCommands` results, token cost), not accumulated usage data.
552
+
553
+ | Field | Type | Values / source |
554
+ | --- | --- | --- |
555
+ | `schema` | string | `eval-ablation/v1` |
556
+ | `recorded_at` | string | ISO-8601 UTC |
557
+ | `ablated_agent` | string | Target agent name |
558
+ | `fixtures` | array of strings | Integration fixtures exercised |
559
+ | `model` | string | Model version(s) used for orchestrator/builder dispatch — deltas are model-dependent, always recorded |
560
+ | `baseline` | object | `{issues_caught, test_commands: [{command, exit_code}], tokens, grade}` — full roster arm |
561
+ | `ablated` | object | Same shape as `baseline` — roster-minus-target-agent arm |
562
+ | `delta` | object | `{issues_caught, test_commands_passed, tokens}` (ablated − baseline) |
563
+ | `verdict` | string | e.g. `"no measured impact — supports drop"` / `"agent is load-bearing — retain"` / `"baseline failed — inconclusive"` |
564
+
565
+ - **Emitter:** `plugins/dev-team/scripts/eval_ablation.py --mode agent` (moved from `scripts/` in #1653).
566
+ - **Consent:** unconditional (eval infra, not user-session telemetry); opt-in/label-gated dispatch per the live-eval cost policy (#134) — the record is only ever written after an explicit operator-confirmed live run.
567
+ - **Consumers:** `skills/harness-audit/SKILL.md` (Step 3 drop-candidate recommendations cite the measured delta/verdict when a record exists).
568
+
569
+ ---
570
+
571
+ ## `refactor-freeze.jsonl`
572
+
573
+ Audit log for the tests-frozen-during-REFACTOR invariant (`#813`) — both the
574
+ enforcement decision and any fail-open diagnostic. Extended by `#906` with
575
+ `bash-freeze`, the preventive PreToolUse(Bash) sibling of `freeze`.
576
+
577
+ | Field | Type | Values / source |
578
+ | --- | --- | --- |
579
+ | `timestamp` | string | ISO-8601 |
580
+ | `hook` | string | `freeze` \| `bash-freeze` \| `revert` |
581
+ | `event` | string enum | `block` \| `fail-open` \| `revert` \| `remove` |
582
+ | `file` | string, optional | File path involved |
583
+ | `step` | string, optional | Plan step label |
584
+ | `reason` | string, optional | Fail-open diagnostic (existing precedent — internal-error text, not a rule ID; unchanged by #859) |
585
+
586
+ - **Emitter:** `hooks/refactor_test_freeze_guard.py::audit()`, `hooks/refactor_test_revert_guard.py` and `hooks/refactor_test_bash_guard.py` (both via the same `audit()` import).
587
+ - **Consent:** unconditional (fails open, audits itself).
588
+ - **Consumers:** none automated yet; inspected manually when the freeze invariant is investigated.
589
+
590
+ ---
591
+
592
+ ## `contract-version-guard-audit.jsonl`
593
+
594
+ Audit log for release-please's bypass of the security-primitives-contract
595
+ version-bump requirement.
596
+
597
+ | Field | Type | Values / source |
598
+ | --- | --- | --- |
599
+ | `ts` | string | ISO-8601 UTC |
600
+ | `bypass` | boolean | Always `true` |
601
+ | `reason` | string | Always `release-please-actor` |
602
+ | `github_actor` | string | `$GITHUB_ACTOR` env value |
603
+ | `git_email` | string | `$GIT_AUTHOR_EMAIL` env value |
604
+
605
+ - **Emitter:** `hooks/contract_version_guard.py::_log_bypass()`.
606
+ - **Consent:** unconditional.
607
+ - **Consumers:** none automated yet; CI-only diagnostic trail.
608
+
609
+ ---
610
+
611
+ ## `learning-loop-state.json`
612
+
613
+ Not JSONL — a single current-value JSON file: a counter gating when
614
+ `session_learning_trigger.py` dispatches background session analysis.
615
+
616
+ | Field | Type | Values / source |
617
+ |---|---|---|
618
+ | `counter` | integer | Turns since the last dispatch |
619
+
620
+ - **Emitter:** `hooks/session_learning_trigger.py::_write_state()`.
621
+ - **Consent:** unconditional (internal scheduling state, no content).
622
+ - **Consumers:** `hooks/session_learning_trigger.py` itself (read on next fire).
623
+
624
+ ---
625
+
626
+ ## `pending-review.jsonl`
627
+
628
+ Queued findings from the background session-analysis dispatch, before
629
+ `/session-review` consumes them.
630
+
631
+ | Field | Type | Values / source |
632
+ | --- | --- | --- |
633
+ | `queued_at` | string | ISO-8601 UTC |
634
+ | `source` | string | Always `session-learning-trigger` |
635
+ | `session_id` | string, optional | Session ID when available |
636
+ | `findings` | array of object | Each: `lever`, `evidence`, `target_artifact`, `proposed_change`, `route` |
637
+
638
+ - **Emitter:** background `claude --print` run dispatched by `hooks/session_learning_trigger.py::_dispatch_background_analysis()`, writing via `session-analysis` agent output.
639
+ - **Consent:** unconditional (dispatch happens automatically; content is model-authored analysis, not raw session data).
640
+ - **Consumers:** `/session-review` skill.
641
+
642
+ ---
643
+
644
+ ## `.claude/metrics/{date}-task-log.jsonl` (e.g. `2026-02-20-task-log.jsonl`)
645
+
646
+ Self-reported per-task completion log, one file per calendar date.
647
+
648
+ | Field | Type | Values / source |
649
+ | --- | --- | --- |
650
+ | `timestamp` | string | ISO-8601 |
651
+ | (task-specific fields) | — | Tokens, cost, agents used, rework cycles, hallucination events — see `skills/performance-metrics/SKILL.md` for the full field list |
652
+
653
+ - **Emitter:** `/performance-metrics` skill (model-authored append at task completion), via `hooks/task_completion_metrics.py`.
654
+ - **Consent:** gated by `telemetry_consent.is_enabled()` (`~/.claude/telemetry.json` `{"enabled": true}`, home-scoped) — no longer unconditional as of Slice 2 (#1406).
655
+ - **Consumers:** `skills/harness-audit/SKILL.md` (self-reported half of the harness-audit join, alongside `session-digest.jsonl`'s real-session half), `skills/governance-compliance/SKILL.md`.
656
+
657
+ ---
658
+
659
+ ## `gherkin-derive-effectiveness.jsonl`
660
+
661
+ Per-scenario roll-up correlating a `/gherkin-derive`-discovered surface with
662
+ whatever coverage/mutation-delta data the calling workflow already measured,
663
+ so there is a signal on whether BDD-derived scenarios track real
664
+ coverage/mutation movement (issue #1296). One record per scenario per
665
+ roll-up run — not deduplicated across runs, since coverage/mutation deltas
666
+ are re-measured every convergence iteration.
667
+
668
+ | Field | Type | Values / source |
669
+ | --- | --- | --- |
670
+ | `surface` | string, nullable | The discovered surface name/path from `gherkin.md`'s surface-inventory table |
671
+ | `discovery_source` | string, nullable | `openapi` \| `route` \| `test` \| `signature` (per `/gherkin-derive` Step 2), as recorded in the inventory |
672
+ | `provenance` | string, nullable | `specification` \| `characterization`, as recorded in the inventory |
673
+ | `binding_mode` | string, nullable | `none` \| `xunit-with-annotations` \| `bdd-runner` |
674
+ | `bound_story` | number or string, nullable | The Story/issue id from `gherkin-bindings.json`, when that file exists for the run (only produced by `/gherkin-public`) |
675
+ | `coverage_delta` | object, nullable | `{line_pct, branch_pct}` — workflow-level delta between the two coverage snapshots passed to the roll-up, not an isolated per-scenario attribution (no finer-grained mapping exists today) |
676
+ | `mutation_delta` | object, nullable | `{survivors_after_delta}` — workflow-level survivor-count delta, same caveat as `coverage_delta` |
677
+
678
+ - **Emitter:** `plugins/dev-team/scripts/gherkin_effectiveness_rollup.py`, invoked from `/quality-targets-converge` Step 6b after each convergence iteration's re-measure, when `gherkin.md` exists for the workflow slug.
679
+ - **Consent:** unconditional (derived metrics only; no prompt/file-content capture).
680
+ - **Consumers:** none yet — this is the roll-up a future `/harness-audit`-style review reads to compare BDD-derived vs. hand-written test effectiveness.
681
+
682
+ ---
683
+
684
+ ## Benefit-measurement streams (#2201)
685
+
686
+ **Added by #2201** (epic #2200, slice 0). Four observational JSONL streams that
687
+ make post-merge benefit numbers for epics #2164 and #2172 exist. Each row is
688
+ written fail-open by `hooks/lib/instrument_log.py` (never affects stdout, exit
689
+ code, or control flow) and carries `ts`, `plugin_version`, and, when the
690
+ emitter has one, `session_id`. They are **separate from `boundary-events.jsonl`**
691
+ on purpose: a `record` row there is read by the review-gate corroboration path,
692
+ so measurement rows must not share it. Counts, enums and lens/agent names only.
693
+
694
+ | Stream | Emitter | Fields | Answers |
695
+ |---|---|---|---|
696
+ | `subagent-stops.jsonl` | `hooks/subagent_completion_guard.py` (every `SubagentStop`) | `classification` (`clean` \| `empty-final-turn` \| `truncated-final-turn` \| `unreadable`) | The completion-guard divergence rate's denominator. `boundary-events.jsonl` only carries the two non-clean classes, so a rate was not computable before. |
697
+ | `skill-injection.jsonl` | `hooks/subagent_skill_context.py` (when a hint is injected) | `agent_type`, `skills` (list), `added_chars` | Injection overhead (`added_chars`) and the denominator for uptake; uptake itself is read from `Skill` tool calls in the subagent transcripts (`scripts/lib/session_log`). |
698
+ | `ledger-skips.jsonl` | `scripts/verdict_scope.py` (every CLI consult) | `candidate_pairs`, `skipped_pairs`, `fully_skipped_lenses` (list) | Realized delta-scoping skip rate = `skipped_pairs / candidate_pairs`. Previously only printed to stdout. |
699
+ | `checkpoint-aborts.jsonl` | `scripts/checkpoint_abort.py` | `mode: "abort"`: `aborted`, `triggering_agent`, `deferred_lenses`. `mode: "outcome"`: `aborted`, `redispatched`, `findings`, `blocking_findings`, `outcome` | Abort frequency, deferred-lens yield (`outcome` rows with `aborted` and `redispatched`), previously only printed. |
700
+
701
+ ### Instrument audit (#2201)
702
+
703
+ Static audit of each instrument; the per-session confirmation the issue also
704
+ asks for (≥ 3 real sessions, IDs listed on #2200) must be run on the
705
+ maintainer's machine after these emitters ship.
706
+
707
+ | Instrument | Finding | Action |
708
+ |---|---|---|
709
+ | `review-verdicts.jsonl` | Rows are written only when the dispatch prompt carries the scope marker (`review_verdict_recorder.py`); a dispatch without it emits a `boundary-events.jsonl` `record` row with `matched_rule: "missing-scope-marker"`. | None; count those `boundary-events.jsonl` rows as the ledger-coverage gap. |
710
+ | `ledgerSkipped` / `fullySkippedLenses` | Returned on stdout only. | New `ledger-skips.jsonl`. |
711
+ | `checkpoint_abort.py` | Outcomes printed only. | New `checkpoint-aborts.jsonl`. |
712
+ | `subagent_skill_context.py` | No signal of injection or of a skill being loaded. | New `skill-injection.jsonl`; "loaded" is derived from subagent transcripts. |
713
+ | `subagent_completion_guard.py` | Emits `empty-final-turn` / `truncated-final-turn` to `boundary-events.jsonl` (`decision: "record"`); `clean`/`unreadable` silent. | New `subagent-stops.jsonl` with every classification. |
714
+
715
+ Consent gating: none beyond the existing per-project `.claude/metrics/`
716
+ location; rows are local files and are never transmitted.
717
+
718
+ ---
719
+
720
+ ## Adding a new stream
721
+
722
+ 1. Name it `.claude/metrics/<name>.jsonl` (or `.json` for a single-current-value
723
+ file) — one stream per concern, matching existing precedent.
724
+ 2. Append-only, compact JSON (`separators=(",", ":")`) + trailing newline for
725
+ JSONL streams.
726
+ 3. Rule IDs / counts / enums only — never command text, prompt text, file
727
+ contents, or incidental free text.
728
+ 4. Add a section to this file with the same shape as the ones above
729
+ (fields/types, emitter, consent gating, consumers) in the same PR that
730
+ introduces the emitter — the coverage test in
731
+ `tests/hooks/test_boundary_events.py` enforces this.
732
+
733
+ ---
734
+
735
+ ## `autoship-log.jsonl`
736
+
737
+ One record per `/autoship` dispatch-unit outcome (a solo issue or a batch),
738
+ plus one `round_summary` event per round. Every record is appended via the
739
+ shared `hooks/lib/autoship_log.py` appender (its `--json`/`--json-file`
740
+ CLI), which stamps `logged_at` regardless of which of the three shapes
741
+ below the caller passes it — the library itself is schema-agnostic; the
742
+ shape is entirely determined by `/autoship`'s own SKILL.md (Steps 3f/4).
743
+
744
+ **Solo entry** — one record per solo-dispatched issue:
745
+
746
+ | Field | Type | Values / source |
747
+ | --- | --- | --- |
748
+ | `logged_at` | string | ISO-8601 (UTC) — stamped by `autoship_log.py` |
749
+ | `round_id` | string | ISO-8601 timestamp generated once at round start (before Step 1) |
750
+ | `issue` | integer | The dispatched issue number |
751
+ | `status` | string enum | `shipped` \| `failed` \| `unrecognized` \| `blocked` |
752
+ | `blocked_reason` | string, nullable | The extracted stakeholder question (`blocked`), the Step-3d.1-synthesized classifier-verdict string (`failed`/`unrecognized`), or `null` for `shipped` |
753
+
754
+ **Batch entry** — one record per batch dispatch unit, never one per member
755
+ issue, applying identically across every outcome (`shipped`, `blocked`,
756
+ `failed`, and `unrecognized` alike):
757
+
758
+ | Field | Type | Values / source |
759
+ | --- | --- | --- |
760
+ | `logged_at` | string | ISO-8601 (UTC) — stamped by `autoship_log.py` |
761
+ | `round_id` | string | Same round-start timestamp as any solo entry logged the same round |
762
+ | `batch_id` | string | e.g. `grp-101` — `autoship_group.py`'s deterministic batch id |
763
+ | `issues` | array of integer | Every member issue number |
764
+ | `status` | string enum | `shipped` \| `failed` \| `unrecognized` \| `blocked` |
765
+ | `blocked_reason` | string, nullable | Same convention as the solo entry's field, applied to the whole batch |
766
+
767
+ **`round_summary` event** — one per round, appended after Step 3's
768
+ per-dispatch-unit loop ends (skipped in `--dry-run`):
769
+
770
+ | Field | Type | Values / source |
771
+ | --- | --- | --- |
772
+ | `logged_at` | string | ISO-8601 (UTC) — stamped by `autoship_log.py` |
773
+ | `round_id` | string | Same round-start timestamp as this round's solo/batch entries |
774
+ | `event` | string | Always `"round_summary"` — distinguishes this record from a solo/batch entry above |
775
+ | `processed_units` / `processed_issues` | integer | Dispatch units Step 3c actually dispatched this round / their total member-issue count |
776
+ | `discovered_units` / `discovered_issues` | integer | Every dispatch unit `autoship_queue.py` produced this round (`queue` + `deferred` combined) / their total member-issue count |
777
+ | `deferred_units` / `deferred_issues` | integer | Dispatch units left in `deferred` / their total member-issue count (a solo unit counts as 1) |
778
+ | `blocked_pending_confirmation_units` / `blocked_pending_confirmation_issues` | integer | Proposed batches Step 2c actually blocked pending human confirmation this round / their total member-issue count — always present, `0` when Step 2b/2c never ran or blocked nothing |
779
+ | `cost_usd` | number | Accumulated round cost |
780
+ | `status` | string enum | `complete` \| `cost_cap_reached` \| `dry_run` \| `no_eligible_issues` \| `no_unit_fits_cap` \| `blocked_pending_confirmation` |
781
+
782
+ - **Emitter:** `hooks/lib/autoship_log.py` called from the `/autoship` skill (Step 3f for solo/batch entries, Step 4 for the `round_summary` event).
783
+ - **Consent:** unconditional (cost/count aggregates and enum values only — no prompt text or file contents).
784
+ - **Consumers:** `/cost-report`, `/telemetry` (aggregate reporting).
785
+
786
+ ---
787
+
788
+ ## `workflow-states.jsonl`
789
+
790
+ **Added by #1166.** Event-sourced workflow lifecycle stream for orchestrated
791
+ flows (`/ship`, `/autoship`, `/build`): persists only state-*transition*
792
+ events. Current state and per-state dwell time are always **derived** by
793
+ replaying the stream for a given `session_id` — never stored — per the
794
+ event-sourcing discipline in the competitive analysis this issue is drawn
795
+ from. Canonical (informational, not enforced) lifecycle: `SPEC -> PLAN ->
796
+ BUILD -> REVIEW -> COMMIT -> PR`.
797
+
798
+ | Field | Type | Values / source |
799
+ | --- | --- | --- |
800
+ | `ts` | string | ISO-8601 UTC `%Y-%m-%dT%H:%M:%SZ` |
801
+ | `workflow` | string | Orchestrated flow name, e.g. `ship`, `autoship`, `build` |
802
+ | `prior_state` | string, optional (`null` for the initial transition) | State the workflow was in before this transition |
803
+ | `new_state` | string | State the workflow is entering |
804
+ | `plugin_version` | string | From `.claude-plugin/plugin.json` |
805
+ | `session_id` | string, optional | Opaque per-session ID — enables joins with `boundary-events.jsonl` and `cost-metering.jsonl` |
806
+
807
+ - **Emitter:** `hooks/lib/workflow_state.py::emit_state_transition()`, invoked via its `record` CLI subcommand as a model-authored append at each phase boundary in `/ship`, `/autoship`, and `/build` (same convention as `review-value.jsonl`/`verify-log.jsonl`).
808
+ - **Consent:** unconditional (workflow/state names + counts only — no prompt text or file contents).
809
+ - **Derivation:** `hooks/lib/workflow_state.py::derive_current_state()` and `compute_dwell_times()` (also exposed via the `report` CLI subcommand) replay a session's transitions — never a stored snapshot.
810
+ - **Consumers:** `skills/run-report/SKILL.md` (#1167), `skills/session-review/SKILL.md`, `skills/harness-audit/SKILL.md`, `skills/cost-report/SKILL.md`.
811
+
812
+ ---
813
+
814
+ ## `iteration-journal.jsonl`
815
+
816
+ **Added by #1168.** Hard per-iteration decision journal for the autonomous
817
+ `/autoship`/`/ship` loops: one entry per round/iteration recording what was
818
+ attempted, its outcome, and the next action — the accountability record an
819
+ autonomous run needs to be debuggable after the fact. Unlike
820
+ `workflow-states.jsonl`'s phase transitions, this stream is not derived; each
821
+ entry is a durable, once-written decision note.
822
+
823
+ | Field | Type | Values / source |
824
+ | --- | --- | --- |
825
+ | `ts` | string | ISO-8601 UTC `%Y-%m-%dT%H:%M:%SZ` |
826
+ | `round_id` | string | Identifier for the current round/iteration (`/autoship`'s round_id, or `/ship`'s issue identifier) |
827
+ | `attempted` | string | Short structured note — what was attempted this iteration (deliberate, agent-authored rationale, not incidental free text — same precedent as `config-changelog.jsonl`'s `description`) |
828
+ | `outcome` | string | Short structured note — what happened |
829
+ | `next_action` | string | Short structured note — what happens next |
830
+ | `plugin_version` | string | From `.claude-plugin/plugin.json` |
831
+ | `session_id` | string, optional | Opaque per-session ID — enables joins with `boundary-events.jsonl` / `cost-metering.jsonl` |
832
+
833
+ - **Emitter:** `hooks/lib/iteration_journal_gate.py::record_iteration_entry()`, invoked via its `record` CLI subcommand as a model-authored append in `/autoship`'s per-issue loop (Step 3) and `/ship`'s per-phase loop, before the corresponding `check` subcommand gates advancement.
834
+ - **Gate:** `hooks/lib/iteration_journal_gate.py::check_iteration_journal()` (`check` CLI subcommand) hard-blocks advancement to the next issue/iteration — exit 1 — unless >=1 entry exists for the current `round_id`; a block also emits a `boundary-events.jsonl` event (`hook: iteration_journal_gate`, `decision: block`, `matched_rule: iteration-journal-missing`). This is a skill-level check-before-advance (mirroring `verify-log.jsonl`'s `progress_guardian.py --pre-pr` pattern), not a `settings.json` PreToolUse/PostToolUse registration — `/autoship`'s and `/ship`'s loop advancement is model-authored control flow inside a skill, not a tool call the harness intercepts at a distinct boundary. Complements, does not replace, the advisory plan-step-keyed `progress-guardian` agent.
835
+ - **Consent:** unconditional (a deliberate per-iteration accountability record, not passive usage telemetry).
836
+ - **Consumers:** `skills/autoship/SKILL.md`, `skills/ship/SKILL.md`, joinable with `skills/run-report/SKILL.md` (#1167) via `round_id`/`session_id`.
837
+
838
+ ---
839
+
840
+ ## `xunit-v3-shim-decisions.json`
841
+
842
+ **Added by #1791.** Not JSONL — a single current-value JSON object keyed by test
843
+ project name, holding the operator's chosen remediation when xunit.v3
844
+ constructs block the Stryker v2 shim. This is the enforcement record, not
845
+ telemetry: `stryker_xunit_shim_guard.py` blocks every `dotnet-stryker` run
846
+ against a blocked project until an entry covering the current blocker set
847
+ exists, which is what makes the always-ask gate a guarantee rather than hook
848
+ stdout an agent may paraphrase or skip.
849
+
850
+ Each value:
851
+
852
+ | Field | Type | Values / source |
853
+ | --- | --- | --- |
854
+ | `project` | string | Test project name (the real test `.csproj` stem) |
855
+ | `choice` | string | `port` \| `exclude` \| `skip` \| `degrade` — the four documented remediations; any other value is rejected at write time |
856
+ | `fingerprint` | string, required | 16-hex digest over the blocker set's `file::construct` pairs (line numbers deliberately excluded). Scopes the decision to the blockers the operator actually saw; a mismatch — or an absent value, which would make the entry a blanket answer — re-asks |
857
+ | `files` | array of string | Flagged files the choice covers, project-relative |
858
+ | `note` | string, nullable | Operator rationale, when given |
859
+ | `recorded_at` | string | ISO-8601 UTC `%Y-%m-%dT%H:%M:%SZ` |
860
+
861
+ - **Emitter:** `hooks/lib/xunit_v3_operator_gate.py::record_decision()`, invoked
862
+ via its `record` CLI subcommand after the operator answers the gate.
863
+ - **Gate:** `hooks/lib/xunit_v3_operator_gate.py::decision_for()`, read by
864
+ `hooks/stryker_xunit_shim_guard.py` (PreToolUse on `Bash`). No covering entry
865
+ → exit 2 with the operator question as the block body. Fails closed on every
866
+ axis: a fingerprint mismatch, an absent fingerprint, and a stored `choice`
867
+ outside the four all re-ask rather than letting a run proceed unasked.
868
+ - **Consent:** unconditional (an explicit operator decision record, not passive
869
+ usage telemetry).
870
+ - **Consumers:** `hooks/stryker_xunit_shim_guard.py`,
871
+ `skills/mutation-testing/scripts/mutation_feasibility_gate.py` (same question
872
+ payload), `skills/stryker-xunit-v2-shim/SKILL.md` Step 1a. No path override
873
+ (#1870 dropped `DEV_TEAM_XUNIT3_SHIM_DECISION_FILE` entirely — no legitimate
874
+ caller needs runtime relocation of the store).
875
+ - **Audit trail (#1870):** every `record_decision()` write and every
876
+ `decision_for()` honor (a stored decision covering the current question,
877
+ about to drive the guard's outcome) also emits a `boundary-events.jsonl`
878
+ entry — `matched_rule` of `xunit-v3-shim-decision-record-<choice>` or
879
+ `xunit-v3-shim-decision-honor-<choice>`, `subject_hash` bound to the
880
+ question's `fingerprint` — so a self-recorded choice is visible in the same
881
+ stream the review-gate corroboration mechanism is audited from.