pi-dev-team 0.2.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (780) hide show
  1. package/LICENSE +21 -0
  2. package/PORTING.md +134 -0
  3. package/README.md +207 -0
  4. package/UPSTREAM.json +64 -0
  5. package/agents/Explore.md +15 -0
  6. package/agents/a11y-review.md +118 -0
  7. package/agents/adr-author.md +70 -0
  8. package/agents/ai-provenance-review.md +120 -0
  9. package/agents/angular-reactivity-review.md +95 -0
  10. package/agents/arch-review.md +135 -0
  11. package/agents/architect.md +78 -0
  12. package/agents/autoship-batch-proposer.md +69 -0
  13. package/agents/claude-setup-review.md +136 -0
  14. package/agents/codebase-recon.md +184 -0
  15. package/agents/component-architecture-review.md +119 -0
  16. package/agents/concurrency-review.md +109 -0
  17. package/agents/correctness-review.md +290 -0
  18. package/agents/data-flow-tracer.md +120 -0
  19. package/agents/doc-review.md +165 -0
  20. package/agents/domain-review.md +136 -0
  21. package/agents/general-purpose.md +10 -0
  22. package/agents/gherkin-quality-critic.md +113 -0
  23. package/agents/js-fp-review.md +114 -0
  24. package/agents/mutation-kill.md +684 -0
  25. package/agents/naming-review.md +142 -0
  26. package/agents/orchestrator.md +339 -0
  27. package/agents/performance-review.md +105 -0
  28. package/agents/plan-review-acceptance.md +115 -0
  29. package/agents/plan-review-design.md +90 -0
  30. package/agents/plan-review-parallelization.md +84 -0
  31. package/agents/plan-review-strategic.md +96 -0
  32. package/agents/plan-review-ux.md +110 -0
  33. package/agents/platform-engineer.md +64 -0
  34. package/agents/product-manager.md +68 -0
  35. package/agents/progress-guardian.md +79 -0
  36. package/agents/qa-engineer.md +289 -0
  37. package/agents/quality-reviewer.md +132 -0
  38. package/agents/react-reactivity-review.md +102 -0
  39. package/agents/refactor-opportunity-review.md +128 -0
  40. package/agents/security-engineer.md +60 -0
  41. package/agents/security-review.md +218 -0
  42. package/agents/session-analysis.md +95 -0
  43. package/agents/software-engineer.md +105 -0
  44. package/agents/spec-compliance-review.md +100 -0
  45. package/agents/spec-reviewer.md +114 -0
  46. package/agents/structure-review.md +146 -0
  47. package/agents/tech-writer.md +84 -0
  48. package/agents/test-review.md +246 -0
  49. package/agents/test-smell-review.md +188 -0
  50. package/agents/token-efficiency-review.md +139 -0
  51. package/agents/ui-ux-designer.md +54 -0
  52. package/agents/vue-reactivity-review.md +95 -0
  53. package/bin/__pycache__/claudecpython-314.pyc +0 -0
  54. package/bin/claude +258 -0
  55. package/docs/upstream/.pages +1 -0
  56. package/docs/upstream/CHANGELOG.md +2586 -0
  57. package/docs/upstream/README.md +155 -0
  58. package/docs/upstream/agent-architecture.md +214 -0
  59. package/docs/upstream/agent_info.md +187 -0
  60. package/docs/upstream/artifact-migration.md +124 -0
  61. package/docs/upstream/code-intelligence-nudge.md +149 -0
  62. package/docs/upstream/code-review-process.md +294 -0
  63. package/docs/upstream/concurrent-use.md +73 -0
  64. package/docs/upstream/context-management.md +111 -0
  65. package/docs/upstream/developer-notes.md +280 -0
  66. package/docs/upstream/diagrams/architecture-overview.svg +101 -0
  67. package/docs/upstream/diagrams/review-dispatch.svg +139 -0
  68. package/docs/upstream/diagrams/team-agents.svg +128 -0
  69. package/docs/upstream/diagrams/test-improve-flow.svg +166 -0
  70. package/docs/upstream/diagrams/workflow-linear.svg +66 -0
  71. package/docs/upstream/diagrams/workflow-three-phase.svg +200 -0
  72. package/docs/upstream/eval-maintenance.md +95 -0
  73. package/docs/upstream/eval-running-guide.md +147 -0
  74. package/docs/upstream/eval-system.md +291 -0
  75. package/docs/upstream/session-review-oss-complements.md +75 -0
  76. package/docs/upstream/session-review.md +212 -0
  77. package/docs/upstream/skills.md +188 -0
  78. package/docs/upstream/team-structure.md +21 -0
  79. package/docs/upstream/telemetry-ci-access.md +129 -0
  80. package/docs/upstream/telemetry-repo-security.md +120 -0
  81. package/docs/upstream/test-evaluation.md +277 -0
  82. package/docs/upstream/test-improve.md +154 -0
  83. package/docs/upstream/triage-workflow.md +282 -0
  84. package/docs/upstream/workflows.md +289 -0
  85. package/extensions/dev-team/index.ts +539 -0
  86. package/extensions/dev-team/lib/agents.ts +272 -0
  87. package/extensions/dev-team/lib/ai-credits.ts +92 -0
  88. package/extensions/dev-team/lib/autocompact.ts +81 -0
  89. package/extensions/dev-team/lib/child-run.ts +102 -0
  90. package/extensions/dev-team/lib/config.ts +236 -0
  91. package/extensions/dev-team/lib/gh-command.ts +103 -0
  92. package/extensions/dev-team/lib/github-style.ts +307 -0
  93. package/extensions/dev-team/lib/hooks.ts +350 -0
  94. package/extensions/dev-team/lib/metrics.ts +115 -0
  95. package/extensions/dev-team/lib/safe-read.ts +49 -0
  96. package/extensions/dev-team/lib/session-files.ts +57 -0
  97. package/extensions/dev-team/lib/session-spend.ts +123 -0
  98. package/extensions/dev-team/lib/shell-scan.ts +205 -0
  99. package/extensions/dev-team/lib/skills.ts +213 -0
  100. package/extensions/dev-team/lib/subagent-render.ts +245 -0
  101. package/extensions/dev-team/lib/subagent-types.ts +164 -0
  102. package/extensions/dev-team/lib/subagent.ts +596 -0
  103. package/extensions/dev-team/lib/terminal-text.ts +54 -0
  104. package/extensions/dev-team/lib/tools-misc.ts +152 -0
  105. package/extensions/dev-team/lib/transcript.ts +110 -0
  106. package/extensions/dev-team/lib/trust.ts +52 -0
  107. package/extensions/dev-team/lib/usage-breakdown.ts +176 -0
  108. package/extensions/dev-team/lib/usage-chart.ts +153 -0
  109. package/extensions/dev-team/lib/usage-command.ts +107 -0
  110. package/extensions/dev-team/lib/usage-history.ts +203 -0
  111. package/extensions/dev-team/lib/usage-render.ts +225 -0
  112. package/extensions/dev-team/lib/usage-split-bar.ts +127 -0
  113. package/extensions/dev-team/lib/usage-state.ts +116 -0
  114. package/extensions/dev-team/lib/usage-text.ts +159 -0
  115. package/extensions/dev-team/lib/usage-view.ts +109 -0
  116. package/hooks/__pycache__/refactor_test_freeze_guard.cpython-314.pyc +0 -0
  117. package/hooks/agent_dispatch_ledger.py +190 -0
  118. package/hooks/autocompact_setup_nudge.py +99 -0
  119. package/hooks/bash_retry_guard.py +228 -0
  120. package/hooks/boundary_events_write_guard.py +352 -0
  121. package/hooks/code_intelligence_nudge.py +293 -0
  122. package/hooks/code_intelligence_turn_mark.py +317 -0
  123. package/hooks/codegraph_bootstrap.py +139 -0
  124. package/hooks/contract_version_guard.py +362 -0
  125. package/hooks/cost_meter.py +106 -0
  126. package/hooks/destructive-commands.json +62 -0
  127. package/hooks/destructive_guard.py +477 -0
  128. package/hooks/eval_compliance_check.py +440 -0
  129. package/hooks/guards.json +17 -0
  130. package/hooks/hooks.json +323 -0
  131. package/hooks/internal_double_gate.py +296 -0
  132. package/hooks/js_fp_review.py +212 -0
  133. package/hooks/knowledge_index.py +119 -0
  134. package/hooks/lib/__pycache__/artifact_paths.cpython-314.pyc +0 -0
  135. package/hooks/lib/__pycache__/atomic_state.cpython-314.pyc +0 -0
  136. package/hooks/lib/__pycache__/autocompact_config.cpython-314.pyc +0 -0
  137. package/hooks/lib/__pycache__/boundary_events.cpython-314.pyc +0 -0
  138. package/hooks/lib/__pycache__/doc_classification.cpython-314.pyc +0 -0
  139. package/hooks/lib/__pycache__/gh_pr_create_detect.cpython-314.pyc +0 -0
  140. package/hooks/lib/__pycache__/git_safe_diff.cpython-314.pyc +0 -0
  141. package/hooks/lib/__pycache__/instrument_log.cpython-314.pyc +0 -0
  142. package/hooks/lib/__pycache__/metrics_query.cpython-314.pyc +0 -0
  143. package/hooks/lib/__pycache__/plugin_version.cpython-314.pyc +0 -0
  144. package/hooks/lib/__pycache__/pre_commit_doc_classifier.cpython-314.pyc +0 -0
  145. package/hooks/lib/__pycache__/review_agent_registry.cpython-314.pyc +0 -0
  146. package/hooks/lib/__pycache__/review_gate_corroboration.cpython-314.pyc +0 -0
  147. package/hooks/lib/__pycache__/review_gate_hash.cpython-314.pyc +0 -0
  148. package/hooks/lib/__pycache__/review_verdicts.cpython-314.pyc +0 -0
  149. package/hooks/lib/__pycache__/stdin_json.cpython-314.pyc +0 -0
  150. package/hooks/lib/__pycache__/stryker_invocation.cpython-314.pyc +0 -0
  151. package/hooks/lib/__pycache__/telemetry_consent.cpython-314.pyc +0 -0
  152. package/hooks/lib/__pycache__/test_file_classify.cpython-314.pyc +0 -0
  153. package/hooks/lib/__pycache__/token_efficiency_limits.cpython-314.pyc +0 -0
  154. package/hooks/lib/__pycache__/verify_guard_state.cpython-314.pyc +0 -0
  155. package/hooks/lib/__pycache__/xunit_v3_operator_gate.cpython-314.pyc +0 -0
  156. package/hooks/lib/agent_skill_hints.py +74 -0
  157. package/hooks/lib/artifact_paths.py +263 -0
  158. package/hooks/lib/atomic_state.py +557 -0
  159. package/hooks/lib/autocompact_config.py +103 -0
  160. package/hooks/lib/autoship_log.py +106 -0
  161. package/hooks/lib/banned_scripts_policy.py +51 -0
  162. package/hooks/lib/boundary_events.py +436 -0
  163. package/hooks/lib/build_knowledge_index.py +504 -0
  164. package/hooks/lib/build_skills_index.py +361 -0
  165. package/hooks/lib/build_state.py +116 -0
  166. package/hooks/lib/classify_ship_outcome.py +126 -0
  167. package/hooks/lib/config_changelog_schema.py +115 -0
  168. package/hooks/lib/cost_meter.py +955 -0
  169. package/hooks/lib/doc_classification.py +116 -0
  170. package/hooks/lib/gh_pr_create_detect.py +136 -0
  171. package/hooks/lib/git_safe_diff.py +123 -0
  172. package/hooks/lib/instrument_log.py +66 -0
  173. package/hooks/lib/iteration_journal_gate.py +197 -0
  174. package/hooks/lib/knowledge_index_paths.py +88 -0
  175. package/hooks/lib/mcp_json_repowise.py +177 -0
  176. package/hooks/lib/metrics_query.py +202 -0
  177. package/hooks/lib/minimal_yaml.py +434 -0
  178. package/hooks/lib/plugin_version.py +142 -0
  179. package/hooks/lib/pre_commit_detect.py +537 -0
  180. package/hooks/lib/pre_commit_doc_classifier.py +126 -0
  181. package/hooks/lib/pricing.py +118 -0
  182. package/hooks/lib/report_pdf.py +371 -0
  183. package/hooks/lib/review_agent_registry.py +142 -0
  184. package/hooks/lib/review_dispatch_ledger.py +101 -0
  185. package/hooks/lib/review_gate_corroboration.py +521 -0
  186. package/hooks/lib/review_gate_hash.py +252 -0
  187. package/hooks/lib/review_gate_normalized_hash.py +1115 -0
  188. package/hooks/lib/review_verdicts.py +301 -0
  189. package/hooks/lib/run_report.py +160 -0
  190. package/hooks/lib/skill_categories.yaml +125 -0
  191. package/hooks/lib/stdin_json.py +57 -0
  192. package/hooks/lib/stryker_invocation.py +102 -0
  193. package/hooks/lib/telemetry_consent.py +41 -0
  194. package/hooks/lib/telemetry_report.py +108 -0
  195. package/hooks/lib/test_file_classify.py +160 -0
  196. package/hooks/lib/token_efficiency_limits.py +51 -0
  197. package/hooks/lib/turn_identity.py +77 -0
  198. package/hooks/lib/verify_guard_state.py +110 -0
  199. package/hooks/lib/workflow_state.py +206 -0
  200. package/hooks/lib/xunit_v3_operator_gate.py +596 -0
  201. package/hooks/mcp_json_repowise_nudge.py +74 -0
  202. package/hooks/mutation_adapters/__init__.py +7 -0
  203. package/hooks/mutation_adapters/__pycache__/__init__.cpython-314.pyc +0 -0
  204. package/hooks/mutation_adapters/__pycache__/lib.cpython-314.pyc +0 -0
  205. package/hooks/mutation_adapters/__pycache__/mutmut.cpython-314.pyc +0 -0
  206. package/hooks/mutation_adapters/__pycache__/pitest.cpython-314.pyc +0 -0
  207. package/hooks/mutation_adapters/__pycache__/stryker.cpython-314.pyc +0 -0
  208. package/hooks/mutation_adapters/__pycache__/stryker_net.cpython-314.pyc +0 -0
  209. package/hooks/mutation_adapters/lib.py +478 -0
  210. package/hooks/mutation_adapters/mutmut.py +188 -0
  211. package/hooks/mutation_adapters/pitest.py +266 -0
  212. package/hooks/mutation_adapters/stryker.py +157 -0
  213. package/hooks/mutation_adapters/stryker_net.py +264 -0
  214. package/hooks/mutation_gate.py +193 -0
  215. package/hooks/mutation_testing_smoke_gate.py +371 -0
  216. package/hooks/pending_review_notify.py +121 -0
  217. package/hooks/phase_marker.py +138 -0
  218. package/hooks/post_compact_state_reinject.py +180 -0
  219. package/hooks/post_format.py +115 -0
  220. package/hooks/pre_commit_knowledge_index.py +128 -0
  221. package/hooks/pre_commit_review.py +66 -0
  222. package/hooks/pre_pr_review.py +694 -0
  223. package/hooks/pre_tool_guard.py +405 -0
  224. package/hooks/py.sh +73 -0
  225. package/hooks/refactor-bash-write-patterns.json +29 -0
  226. package/hooks/refactor_test_bash_guard.py +253 -0
  227. package/hooks/refactor_test_freeze_guard.py +139 -0
  228. package/hooks/refactor_test_revert_guard.py +186 -0
  229. package/hooks/repo_review_nudge.py +287 -0
  230. package/hooks/review_verdict_recorder.py +464 -0
  231. package/hooks/scan_bash_command_for_banned_scripts.py +428 -0
  232. package/hooks/scan_worktree_for_banned_scripts.py +238 -0
  233. package/hooks/session_learning_trigger.py +248 -0
  234. package/hooks/skills_index.py +126 -0
  235. package/hooks/stryker_xunit_shim_guard.py +571 -0
  236. package/hooks/subagent_completion_guard.py +309 -0
  237. package/hooks/subagent_skill_context.py +139 -0
  238. package/hooks/task_completion_metrics.py +216 -0
  239. package/hooks/tdd_guard.py +229 -0
  240. package/hooks/telemetry.py +341 -0
  241. package/hooks/token_efficiency_review.py +194 -0
  242. package/hooks/verify_guard.py +183 -0
  243. package/hooks/verify_guard_edit_marker.py +73 -0
  244. package/hooks/version_check.py +173 -0
  245. package/knowledge/accepted-risks-schema.md +98 -0
  246. package/knowledge/adr-decision-criteria.md +64 -0
  247. package/knowledge/adversarial-review-protocol.md +139 -0
  248. package/knowledge/agent-registry.md +228 -0
  249. package/knowledge/agent-review-methodology.md +80 -0
  250. package/knowledge/ai-friendly-repo-guidelines.md +67 -0
  251. package/knowledge/architecture-assessment.md +96 -0
  252. package/knowledge/artifact-lifecycle.md +57 -0
  253. package/knowledge/cd-maturity-model.md +82 -0
  254. package/knowledge/cd-test-architecture.md +190 -0
  255. package/knowledge/ci-cd-file-scope.md +24 -0
  256. package/knowledge/codegraph-vs-graphify.md +192 -0
  257. package/knowledge/component-test-patterns.md +139 -0
  258. package/knowledge/database-change-management.md +80 -0
  259. package/knowledge/database-test-patterns.md +79 -0
  260. package/knowledge/decision-defaults.md +88 -0
  261. package/knowledge/dependency-breaking-techniques.md +116 -0
  262. package/knowledge/deployment-pipeline.md +86 -0
  263. package/knowledge/design-smells.md +122 -0
  264. package/knowledge/directory-enumeration.md +38 -0
  265. package/knowledge/domain-modeling.md +123 -0
  266. package/knowledge/evidence-bundle.md +90 -0
  267. package/knowledge/exploratory-testing-field-guide.md +122 -0
  268. package/knowledge/failure-routing.md +28 -0
  269. package/knowledge/fixture-construction.md +56 -0
  270. package/knowledge/frontend-component-architecture.md +139 -0
  271. package/knowledge/gherkin-quality-review-dispatch.md +135 -0
  272. package/knowledge/index.json +6766 -0
  273. package/knowledge/internal-collaborator-doubling.md +101 -0
  274. package/knowledge/legacy-test-strategy.md +71 -0
  275. package/knowledge/long-run-waiting.md +66 -0
  276. package/knowledge/microservice-testing.md +71 -0
  277. package/knowledge/model-pricing.json +23 -0
  278. package/knowledge/mutation-score-formulas.md +60 -0
  279. package/knowledge/object-calisthenics.md +147 -0
  280. package/knowledge/oracle-provenance.md +94 -0
  281. package/knowledge/orchestrator-script-implementation.md +185 -0
  282. package/knowledge/owasp-detection.md +148 -0
  283. package/knowledge/plan-review-rubric.md +56 -0
  284. package/knowledge/proxy-connectivity.md +62 -0
  285. package/knowledge/reactive-effect-patterns.md +73 -0
  286. package/knowledge/recon-inventory-excludes.txt +32 -0
  287. package/knowledge/references/bdd-value-guide.md +61 -0
  288. package/knowledge/references/csharp-http-client-testing.md +264 -0
  289. package/knowledge/release-strategies.md +74 -0
  290. package/knowledge/report-output-location.md +117 -0
  291. package/knowledge/report-pdf-integration.md +63 -0
  292. package/knowledge/report-print.css +129 -0
  293. package/knowledge/report-template.md +114 -0
  294. package/knowledge/report-to-pdf.md +69 -0
  295. package/knowledge/request-processing-flow.md +63 -0
  296. package/knowledge/result-verification.md +52 -0
  297. package/knowledge/review-agent-output-contract.md +121 -0
  298. package/knowledge/review-lens-classification.md +113 -0
  299. package/knowledge/review-rubric.md +62 -0
  300. package/knowledge/review-template.md +104 -0
  301. package/knowledge/rule-fixtures/A02.insecure-random-js/negative.js +1 -0
  302. package/knowledge/rule-fixtures/A02.insecure-random-js/positive.js +1 -0
  303. package/knowledge/rule-fixtures/A02.weak-hashing-md5/negative.py +1 -0
  304. package/knowledge/rule-fixtures/A02.weak-hashing-md5/positive.py +1 -0
  305. package/knowledge/rule-fixtures/A03.command-injection/negative.js +1 -0
  306. package/knowledge/rule-fixtures/A03.command-injection/positive.js +1 -0
  307. package/knowledge/rule-fixtures/A03.sql-injection/negative.js +1 -0
  308. package/knowledge/rule-fixtures/A03.sql-injection/positive.js +1 -0
  309. package/knowledge/rule-fixtures/A03.xss-innerhtml/negative.js +1 -0
  310. package/knowledge/rule-fixtures/A03.xss-innerhtml/positive.js +1 -0
  311. package/knowledge/rule-fixtures/A05.cors-wildcard/negative.js +1 -0
  312. package/knowledge/rule-fixtures/A05.cors-wildcard/positive.js +1 -0
  313. package/knowledge/rule-fixtures/A05.default-credentials/negative.js +1 -0
  314. package/knowledge/rule-fixtures/A05.default-credentials/positive.js +1 -0
  315. package/knowledge/rule-fixtures/A07.jwt-alg-none/negative.js +1 -0
  316. package/knowledge/rule-fixtures/A07.jwt-alg-none/positive.js +1 -0
  317. package/knowledge/rule-fixtures/A08.binary-formatter/negative.cs +1 -0
  318. package/knowledge/rule-fixtures/A08.binary-formatter/positive.cs +1 -0
  319. package/knowledge/rule-fixtures/A08.js-eval/negative.js +1 -0
  320. package/knowledge/rule-fixtures/A08.js-eval/positive.js +1 -0
  321. package/knowledge/rule-fixtures/A08.object-input-stream/negative.java +1 -0
  322. package/knowledge/rule-fixtures/A08.object-input-stream/positive.java +1 -0
  323. package/knowledge/schemas/disposition-register-v1.json +65 -0
  324. package/knowledge/schemas/recon-envelope-v1.json +198 -0
  325. package/knowledge/schemas/unified-finding-v1.json +72 -0
  326. package/knowledge/security-primitives-contract.md +301 -0
  327. package/knowledge/security-review-rule-map.yaml +107 -0
  328. package/knowledge/skills-registry.md +72 -0
  329. package/knowledge/task-size-classifier.md +103 -0
  330. package/knowledge/telemetry-schema.md +881 -0
  331. package/knowledge/test-automation-maturity.md +56 -0
  332. package/knowledge/test-automation-principles.md +71 -0
  333. package/knowledge/test-cadence-tradeoffs.md +68 -0
  334. package/knowledge/test-doubles.md +105 -0
  335. package/knowledge/test-file-indicators.md +22 -0
  336. package/knowledge/test-layer-gates.md +35 -0
  337. package/knowledge/test-matrix-examples/django-batch.md +24 -0
  338. package/knowledge/test-matrix-examples/dotnet-grpc-fronting-api.md +90 -0
  339. package/knowledge/test-matrix-examples/dotnet-http-consumer.md +131 -0
  340. package/knowledge/test-matrix-examples/react-node-spa.md +24 -0
  341. package/knowledge/test-matrix-examples/spring-boot-service.md +25 -0
  342. package/knowledge/test-matrix-examples/ssr-htmx.md +24 -0
  343. package/knowledge/test-organization.md +70 -0
  344. package/knowledge/test-pyramid.md +84 -0
  345. package/knowledge/test-refactoring.md +67 -0
  346. package/knowledge/test-review-division-of-labor.md +85 -0
  347. package/knowledge/test-smells.md +80 -0
  348. package/knowledge/test-stack-profiles/bdd-frameworks.md +235 -0
  349. package/knowledge/test-stack-profiles/django.md +13 -0
  350. package/knowledge/test-stack-profiles/dotnet.md +18 -0
  351. package/knowledge/test-stack-profiles/go.md +16 -0
  352. package/knowledge/test-stack-profiles/node.md +16 -0
  353. package/knowledge/test-stack-profiles/react.md +12 -0
  354. package/knowledge/test-stack-profiles/spring-boot.md +16 -0
  355. package/knowledge/test-stack-profiles/ssr-htmx.md +14 -0
  356. package/knowledge/test-stack-profiles/vue.md +12 -0
  357. package/knowledge/test-strategy.md +70 -0
  358. package/knowledge/testability-patterns.md +240 -0
  359. package/knowledge/testing-quadrants.md +44 -0
  360. package/knowledge/testing-techniques/approval.md +15 -0
  361. package/knowledge/testing-techniques/chaos.md +17 -0
  362. package/knowledge/testing-techniques/fuzz.md +15 -0
  363. package/knowledge/testing-techniques/property-based.md +15 -0
  364. package/knowledge/testing-techniques/schema-validation.md +15 -0
  365. package/knowledge/testing-techniques/screenshot.md +15 -0
  366. package/knowledge/three-phase-workflow.md +198 -0
  367. package/knowledge/value-patterns.md +55 -0
  368. package/knowledge/verification-mode.md +116 -0
  369. package/knowledge/virtual-service-libraries.md +75 -0
  370. package/knowledge/wave-consolidation-guidance.md +21 -0
  371. package/overrides/agents/Explore.md +15 -0
  372. package/overrides/agents/general-purpose.md +10 -0
  373. package/overrides/notes/autoship.md +6 -0
  374. package/overrides/notes/issues-from-assessment.md +3 -0
  375. package/overrides/notes/issues-from-plan.md +3 -0
  376. package/overrides/notes/mutation-night-watch.md +3 -0
  377. package/overrides/notes/mutation-testing.md +3 -0
  378. package/overrides/notes/pr.md +7 -0
  379. package/overrides/notes/project-init.md +6 -0
  380. package/overrides/notes/setup.md +13 -0
  381. package/overrides/notes/specs.md +3 -0
  382. package/overrides/skills/headless-run/SKILL.md +45 -0
  383. package/overrides/skills/upgrade/SKILL.md +30 -0
  384. package/overrides/skills/version/SKILL.md +25 -0
  385. package/package.json +36 -0
  386. package/scripts/authoring_digest.py +93 -0
  387. package/scripts/autoship_discover.py +121 -0
  388. package/scripts/autoship_group.py +409 -0
  389. package/scripts/autoship_proposals.py +494 -0
  390. package/scripts/autoship_queue.py +291 -0
  391. package/scripts/autoship_reclaim.py +495 -0
  392. package/scripts/build_jobs.py +108 -0
  393. package/scripts/build_rollback_point.py +240 -0
  394. package/scripts/build_slice_scope.py +157 -0
  395. package/scripts/build_wave.py +109 -0
  396. package/scripts/build_wave_reconcile.py +252 -0
  397. package/scripts/build_worktree_baseref.py +113 -0
  398. package/scripts/check_agent_scope.py +117 -0
  399. package/scripts/check_agent_tool_mapping.py +213 -0
  400. package/scripts/check_review_agent_mcp_tools.py +317 -0
  401. package/scripts/check_security_assessment_mcp_tools.py +165 -0
  402. package/scripts/checkpoint_abort.py +502 -0
  403. package/scripts/claude_setup_review.py +438 -0
  404. package/scripts/codebase_recon.py +556 -0
  405. package/scripts/coverage_config.py +623 -0
  406. package/scripts/coverage_delta_steering.py +330 -0
  407. package/scripts/coverage_discovery_dotnet.py +315 -0
  408. package/scripts/coverage_discovery_java.py +742 -0
  409. package/scripts/coverage_discovery_js.py +546 -0
  410. package/scripts/coverage_gap_ranking.py +556 -0
  411. package/scripts/coverage_readiness.py +455 -0
  412. package/scripts/coverage_report_parse.py +521 -0
  413. package/scripts/detect_bdd_convention.py +252 -0
  414. package/scripts/eval_ablation.py +376 -0
  415. package/scripts/gherkin_analysis_coverage_gate.py +306 -0
  416. package/scripts/gherkin_cross_feature_duplicate_titles_gate.py +173 -0
  417. package/scripts/gherkin_effectiveness_rollup.py +238 -0
  418. package/scripts/gherkin_failure_path_gate.py +206 -0
  419. package/scripts/gherkin_feature_merge.py +720 -0
  420. package/scripts/gherkin_stub_gate.py +163 -0
  421. package/scripts/gherkin_stub_merge.py +479 -0
  422. package/scripts/git_origin_host.py +88 -0
  423. package/scripts/install-java-static-analysis.py +110 -0
  424. package/scripts/issue_deps.py +74 -0
  425. package/scripts/lib/_bdd_markers.py +28 -0
  426. package/scripts/lib/_gherkin_text.py +93 -0
  427. package/scripts/lib/_vendored_tree.py +70 -0
  428. package/scripts/lib/autoship_state.py +397 -0
  429. package/scripts/lib/claude_md_guard.py +226 -0
  430. package/scripts/lib/deterministic_recon.py +446 -0
  431. package/scripts/lib/mcp_tool_grants.py +211 -0
  432. package/scripts/lib/plan_parse.py +386 -0
  433. package/scripts/lib/review_result.py +84 -0
  434. package/scripts/lib/review_roster.py +86 -0
  435. package/scripts/lib/session_log/__init__.py +34 -0
  436. package/scripts/lib/session_log/__pycache__/__init__.cpython-314.pyc +0 -0
  437. package/scripts/lib/session_log/__pycache__/records.cpython-314.pyc +0 -0
  438. package/scripts/lib/session_log/classify.py +231 -0
  439. package/scripts/lib/session_log/corrections.py +194 -0
  440. package/scripts/lib/session_log/discovery.py +108 -0
  441. package/scripts/lib/session_log/records.py +218 -0
  442. package/scripts/lib/session_log/redact.py +76 -0
  443. package/scripts/lib/session_log/signals.py +373 -0
  444. package/scripts/lib/session_report_downstream.py +614 -0
  445. package/scripts/lib/session_report_maintainer.py +1273 -0
  446. package/scripts/lib/session_report_shared.py +262 -0
  447. package/scripts/lib/settings_hook_guard.py +157 -0
  448. package/scripts/lib/slug.py +33 -0
  449. package/scripts/lib/stub_extractors/__init__.py +82 -0
  450. package/scripts/lib/stub_extractors/_common.py +328 -0
  451. package/scripts/lib/stub_extractors/csharp.py +19 -0
  452. package/scripts/lib/stub_extractors/go.py +173 -0
  453. package/scripts/lib/stub_extractors/java.py +18 -0
  454. package/scripts/lib/stub_extractors/jsts.py +126 -0
  455. package/scripts/mutation_stack_sections.py +149 -0
  456. package/scripts/mutation_yield_steering.py +345 -0
  457. package/scripts/orchestrator.py +895 -0
  458. package/scripts/plan_gherkin_export.py +227 -0
  459. package/scripts/plan_waves.py +208 -0
  460. package/scripts/pr_close_keyword_lint.py +108 -0
  461. package/scripts/progress_guardian.py +888 -0
  462. package/scripts/recon_inventory.py +273 -0
  463. package/scripts/review_findings_log.py +93 -0
  464. package/scripts/run_invariants.py +124 -0
  465. package/scripts/select_lenses.py +640 -0
  466. package/scripts/session_report.py +486 -0
  467. package/scripts/set_autocompact_env.py +221 -0
  468. package/scripts/ship_resume_guard.py +135 -0
  469. package/scripts/ship_review_gate.py +63 -0
  470. package/scripts/specs_convention_marker.py +103 -0
  471. package/scripts/test_improve_resume.py +277 -0
  472. package/scripts/test_review_mechanics.py +958 -0
  473. package/scripts/token_efficiency_review.py +322 -0
  474. package/scripts/verdict_scope.py +285 -0
  475. package/scripts/verify_gherkin_quality_critic_isolation.py +296 -0
  476. package/scripts/verify_tier.py +157 -0
  477. package/skills/adr-tools/SKILL.md +118 -0
  478. package/skills/agent-readiness/SKILL.md +105 -0
  479. package/skills/agent-readiness/ai_friendly_analyzers.py +326 -0
  480. package/skills/agent-readiness/scanner.py +441 -0
  481. package/skills/agent-readiness/scorecard.yaml +88 -0
  482. package/skills/api-design/SKILL.md +115 -0
  483. package/skills/apply-fixes/SKILL.md +171 -0
  484. package/skills/apply-test-doubles/SKILL.md +321 -0
  485. package/skills/artifact-lifecycle/SKILL.md +127 -0
  486. package/skills/autoship/SKILL.md +1124 -0
  487. package/skills/benchmark/SKILL.md +105 -0
  488. package/skills/branch-workflow/SKILL.md +89 -0
  489. package/skills/browse/SKILL.md +184 -0
  490. package/skills/browser-testing/SKILL.md +62 -0
  491. package/skills/browser-testing/references/playwright-patterns.md +216 -0
  492. package/skills/build/SKILL.md +422 -0
  493. package/skills/build/references/static-self-heal.md +245 -0
  494. package/skills/careful/SKILL.md +72 -0
  495. package/skills/cd-test-architecture/SKILL.md +371 -0
  496. package/skills/ci-debugging/SKILL.md +105 -0
  497. package/skills/co-evolution-audit/SKILL.md +269 -0
  498. package/skills/code-review/SKILL.md +1015 -0
  499. package/skills/code-review/examples/aggregated-sample.json +56 -0
  500. package/skills/code-review/examples/sample-report.md +41 -0
  501. package/skills/code-review/output-format.md +478 -0
  502. package/skills/code-review/scripts/activation.py +86 -0
  503. package/skills/code-review/scripts/change_impact.py +357 -0
  504. package/skills/code-review/scripts/change_shape.py +372 -0
  505. package/skills/code-review/scripts/change_size.py +212 -0
  506. package/skills/code-review/scripts/changed_file_list.py +141 -0
  507. package/skills/code-review/scripts/closing_pass.py +187 -0
  508. package/skills/code-review/scripts/consolidate.py +277 -0
  509. package/skills/code-review/scripts/contract_failure_report.py +185 -0
  510. package/skills/code-review/scripts/dispatch_reconcile.py +66 -0
  511. package/skills/code-review/scripts/dispatch_waves.py +164 -0
  512. package/skills/code-review/scripts/finding_signature.py +446 -0
  513. package/skills/code-review/scripts/ledger.py +283 -0
  514. package/skills/code-review/scripts/partition.py +169 -0
  515. package/skills/code-review/scripts/render_tiered_findings.py +274 -0
  516. package/skills/code-review/scripts/repo_invariants.py +1066 -0
  517. package/skills/code-review/scripts/review_context_pack.py +306 -0
  518. package/skills/code-review/scripts/review_round_log.py +345 -0
  519. package/skills/code-review/scripts/review_value_coverage.py +297 -0
  520. package/skills/code-review/scripts/validate_review_output.py +467 -0
  521. package/skills/code-review/sliced-mode.md +205 -0
  522. package/skills/competitive-analysis/SKILL.md +191 -0
  523. package/skills/context-loading-protocol/SKILL.md +157 -0
  524. package/skills/continue/SKILL.md +90 -0
  525. package/skills/cost-report/SKILL.md +178 -0
  526. package/skills/coverage-baseline/SKILL.md +335 -0
  527. package/skills/coverage-baseline/references/multi-project-discovery.md +202 -0
  528. package/skills/coverage-delta/SKILL.md +181 -0
  529. package/skills/coverage-delta/references/mutation-gate.md +70 -0
  530. package/skills/design-doc/SKILL.md +95 -0
  531. package/skills/design-interrogation/SKILL.md +89 -0
  532. package/skills/design-it-twice/SKILL.md +91 -0
  533. package/skills/docker-image-audit/SKILL.md +108 -0
  534. package/skills/docker-image-audit/references/install-guide.md +64 -0
  535. package/skills/docker-image-audit/references/report-template.md +73 -0
  536. package/skills/docker-image-create/SKILL.md +185 -0
  537. package/skills/domain-analysis/SKILL.md +183 -0
  538. package/skills/domain-driven-design/SKILL.md +194 -0
  539. package/skills/exploratory-testing/SKILL.md +108 -0
  540. package/skills/explore/SKILL.md +51 -0
  541. package/skills/farley-score/SKILL.md +165 -0
  542. package/skills/feature-file-validation/SKILL.md +78 -0
  543. package/skills/feature-file-validation/references/validation-rules.md +115 -0
  544. package/skills/feedback-learning/SKILL.md +414 -0
  545. package/skills/fix/SKILL.md +450 -0
  546. package/skills/freeze/SKILL.md +68 -0
  547. package/skills/frontend-architecture/SKILL.md +113 -0
  548. package/skills/gherkin-derive/SKILL.md +630 -0
  549. package/skills/gherkin-public/SKILL.md +266 -0
  550. package/skills/governance-compliance/SKILL.md +150 -0
  551. package/skills/guard/SKILL.md +75 -0
  552. package/skills/handoff/SKILL.md +139 -0
  553. package/skills/handoff/references/summary-templates.md +242 -0
  554. package/skills/harness-audit/SKILL.md +751 -0
  555. package/skills/harness-audit/scripts/lesson_validate.py +386 -0
  556. package/skills/harness-audit/scripts/redundancy_criterion.py +188 -0
  557. package/skills/headless-run/SKILL.md +45 -0
  558. package/skills/headless-run/scripts/isolated_dispatch.py +381 -0
  559. package/skills/help/SKILL.md +72 -0
  560. package/skills/hexagonal-architecture/SKILL.md +85 -0
  561. package/skills/human-oversight-protocol/SKILL.md +224 -0
  562. package/skills/issues-from-assessment/SKILL.md +223 -0
  563. package/skills/issues-from-plan/SKILL.md +133 -0
  564. package/skills/legacy-code/SKILL.md +132 -0
  565. package/skills/mermaid-diagramming/SKILL.md +120 -0
  566. package/skills/mutation-night-watch/SKILL.md +154 -0
  567. package/skills/mutation-night-watch/references/scheduling.md +135 -0
  568. package/skills/mutation-testing/SKILL.md +396 -0
  569. package/skills/mutation-testing/references/languages/csharp-stryker-net.md +676 -0
  570. package/skills/mutation-testing/references/languages/go-go-mutesting.md +95 -0
  571. package/skills/mutation-testing/references/languages/java-pitest.md +77 -0
  572. package/skills/mutation-testing/references/languages/javascript-stryker.md +188 -0
  573. package/skills/mutation-testing/references/languages/python-mutmut.md +97 -0
  574. package/skills/mutation-testing/references/time-estimation.md +34 -0
  575. package/skills/mutation-testing/references/tool-detection.md +15 -0
  576. package/skills/mutation-testing/references/workflow-callers.md +23 -0
  577. package/skills/mutation-testing/scripts/__pycache__/xunit_v3_feature_detector.cpython-314.pyc +0 -0
  578. package/skills/mutation-testing/scripts/csharp_stryker_net_slice_runner.py +635 -0
  579. package/skills/mutation-testing/scripts/csharp_stryker_net_status_loop.py +525 -0
  580. package/skills/mutation-testing/scripts/csharp_stryker_net_wrapper.py +681 -0
  581. package/skills/mutation-testing/scripts/mutation_baseline_reuse.py +292 -0
  582. package/skills/mutation-testing/scripts/mutation_exclude_policy.py +268 -0
  583. package/skills/mutation-testing/scripts/mutation_feasibility_gate.py +463 -0
  584. package/skills/mutation-testing/scripts/mutation_kill_headless.py +331 -0
  585. package/skills/mutation-testing/scripts/mutation_kill_insert.py +199 -0
  586. package/skills/mutation-testing/scripts/mutation_kill_insert_python.py +150 -0
  587. package/skills/mutation-testing/scripts/mutation_kill_loop.py +869 -0
  588. package/skills/mutation-testing/scripts/mutation_kill_loop_python.py +949 -0
  589. package/skills/mutation-testing/scripts/mutation_kill_retry.py +592 -0
  590. package/skills/mutation-testing/scripts/mutation_kill_shared.py +620 -0
  591. package/skills/mutation-testing/scripts/mutation_nightwatch.py +462 -0
  592. package/skills/mutation-testing/scripts/mutation_nightwatch_stacks.py +425 -0
  593. package/skills/mutation-testing/scripts/mutation_report.py +743 -0
  594. package/skills/mutation-testing/scripts/mutation_report_cli.py +175 -0
  595. package/skills/mutation-testing/scripts/mutation_safety_gate.py +69 -0
  596. package/skills/mutation-testing/scripts/stryker_shard_pipeline.py +847 -0
  597. package/skills/mutation-testing/scripts/stryker_shard_setup.py +440 -0
  598. package/skills/mutation-testing/scripts/stryker_timeout_retry.py +142 -0
  599. package/skills/mutation-testing/scripts/xunit_v3_feature_detector.py +341 -0
  600. package/skills/performance-benchmark/SKILL.md +174 -0
  601. package/skills/performance-benchmark/examples/report-format.md +43 -0
  602. package/skills/performance-benchmark/references/benchmark-script.md +169 -0
  603. package/skills/performance-metrics/SKILL.md +265 -0
  604. package/skills/plan/SKILL.md +199 -0
  605. package/skills/plan/references/gherkin-persistence.md +43 -0
  606. package/skills/plan/references/plan-template.md +182 -0
  607. package/skills/pr/SKILL.md +289 -0
  608. package/skills/pr/scripts/gate_retry_state.py +368 -0
  609. package/skills/project-init/README.md +141 -0
  610. package/skills/project-init/SKILL.md +1197 -0
  611. package/skills/project-init/evals/evals.json +200 -0
  612. package/skills/project-init/references/capability-tools.md +55 -0
  613. package/skills/project-init/references/configs.md +221 -0
  614. package/skills/property-based-testing/SKILL.md +121 -0
  615. package/skills/property-based-testing/fixtures/invariant_fixture.py +15 -0
  616. package/skills/property-based-testing/fixtures/js-roundtrip/README.md +42 -0
  617. package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/fast-check/LICENSE +21 -0
  618. package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/fast-check/README.md +263 -0
  619. package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/fast-check/lib/cjs/fast-check.d.ts +5165 -0
  620. package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/fast-check/lib/cjs/fast-check.js +12147 -0
  621. package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/fast-check/lib/cjs/package.json +3 -0
  622. package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/fast-check/lib/cjs/types57/fast-check.d.ts +5165 -0
  623. package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/fast-check/lib/fast-check.d.ts +5165 -0
  624. package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/fast-check/lib/fast-check.js +12011 -0
  625. package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/fast-check/lib/rolldown-runtime-D7D4PA-g.js +13 -0
  626. package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/fast-check/lib/types57/fast-check.d.ts +5165 -0
  627. package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/fast-check/package.json +94 -0
  628. package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/LICENSE +21 -0
  629. package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/README.md +168 -0
  630. package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/RandomGenerator-DcXj09Ch.d.ts +14 -0
  631. package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/distribution/uniformBigInt.d.ts +15 -0
  632. package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/distribution/uniformBigInt.js +38 -0
  633. package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/distribution/uniformFloat32.d.ts +15 -0
  634. package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/distribution/uniformFloat32.js +18 -0
  635. package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/distribution/uniformFloat64.d.ts +15 -0
  636. package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/distribution/uniformFloat64.js +22 -0
  637. package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/distribution/uniformInt.d.ts +15 -0
  638. package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/distribution/uniformInt.js +134 -0
  639. package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/RandomGenerator-DcXj09Ch.d.ts +14 -0
  640. package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/distribution/uniformBigInt.d.ts +15 -0
  641. package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/distribution/uniformBigInt.js +37 -0
  642. package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/distribution/uniformFloat32.d.ts +15 -0
  643. package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/distribution/uniformFloat32.js +17 -0
  644. package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/distribution/uniformFloat64.d.ts +15 -0
  645. package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/distribution/uniformFloat64.js +21 -0
  646. package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/distribution/uniformInt.d.ts +15 -0
  647. package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/distribution/uniformInt.js +133 -0
  648. package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/generator/congruential32.d.ts +7 -0
  649. package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/generator/congruential32.js +44 -0
  650. package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/generator/mersenne.d.ts +7 -0
  651. package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/generator/mersenne.js +90 -0
  652. package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/generator/xoroshiro128plus.d.ts +7 -0
  653. package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/generator/xoroshiro128plus.js +80 -0
  654. package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/generator/xorshift128plus.d.ts +7 -0
  655. package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/generator/xorshift128plus.js +78 -0
  656. package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/package.json +3 -0
  657. package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/types/JumpableRandomGenerator.d.ts +16 -0
  658. package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/types/JumpableRandomGenerator.js +0 -0
  659. package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/types/RandomGenerator.d.ts +2 -0
  660. package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/types/RandomGenerator.js +0 -0
  661. package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/utils/generateN.d.ts +6 -0
  662. package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/utils/generateN.js +8 -0
  663. package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/utils/purify.d.ts +12 -0
  664. package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/utils/purify.js +9 -0
  665. package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/utils/skipN.d.ts +6 -0
  666. package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/esm/utils/skipN.js +6 -0
  667. package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/generator/congruential32.d.ts +7 -0
  668. package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/generator/congruential32.js +46 -0
  669. package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/generator/mersenne.d.ts +7 -0
  670. package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/generator/mersenne.js +92 -0
  671. package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/generator/xoroshiro128plus.d.ts +7 -0
  672. package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/generator/xoroshiro128plus.js +82 -0
  673. package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/generator/xorshift128plus.d.ts +7 -0
  674. package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/generator/xorshift128plus.js +80 -0
  675. package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/types/JumpableRandomGenerator.d.ts +16 -0
  676. package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/types/JumpableRandomGenerator.js +0 -0
  677. package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/types/RandomGenerator.d.ts +2 -0
  678. package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/types/RandomGenerator.js +0 -0
  679. package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/utils/generateN.d.ts +6 -0
  680. package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/utils/generateN.js +9 -0
  681. package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/utils/purify.d.ts +12 -0
  682. package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/utils/purify.js +10 -0
  683. package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/utils/skipN.d.ts +6 -0
  684. package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/lib/utils/skipN.js +7 -0
  685. package/skills/property-based-testing/fixtures/js-roundtrip/node_modules/pure-rand/package.json +133 -0
  686. package/skills/property-based-testing/fixtures/js-roundtrip/package-lock.json +1179 -0
  687. package/skills/property-based-testing/fixtures/js-roundtrip/package.json +14 -0
  688. package/skills/property-based-testing/fixtures/js-roundtrip/roundtrip.js +29 -0
  689. package/skills/property-based-testing/fixtures/js-roundtrip/roundtrip.properties.test.js +16 -0
  690. package/skills/property-based-testing/fixtures/no_property_fixture.py +10 -0
  691. package/skills/property-based-testing/fixtures/roundtrip_fixture.py +16 -0
  692. package/skills/property-based-testing/references/languages/javascript.md +54 -0
  693. package/skills/property-based-testing/scripts/detect_and_dispatch.py +80 -0
  694. package/skills/property-based-testing/scripts/hypothesis_scaffold.py +276 -0
  695. package/skills/proxy-resilience/SKILL.md +84 -0
  696. package/skills/quality-gate-pipeline/SKILL.md +184 -0
  697. package/skills/quality-targets-converge/SKILL.md +254 -0
  698. package/skills/repo-review/SKILL.md +159 -0
  699. package/skills/report-pdf/SKILL.md +66 -0
  700. package/skills/review/SKILL.md +47 -0
  701. package/skills/review-agent/SKILL.md +152 -0
  702. package/skills/review-summary/SKILL.md +73 -0
  703. package/skills/run-report/SKILL.md +70 -0
  704. package/skills/semantic-duplication-scan/SKILL.md +337 -0
  705. package/skills/semantic-scan/SKILL.md +53 -0
  706. package/skills/semgrep-analyze/SKILL.md +139 -0
  707. package/skills/setup/SKILL.md +1122 -0
  708. package/skills/ship/SKILL.md +240 -0
  709. package/skills/source-verification/SKILL.md +210 -0
  710. package/skills/source-verification/scripts/claim_extractor.py +155 -0
  711. package/skills/specs/.size-baseline.json +4 -0
  712. package/skills/specs/SKILL.md +243 -0
  713. package/skills/specs/references/completeness-checklist.md +83 -0
  714. package/skills/specs/references/extraction.md +58 -0
  715. package/skills/specs/references/glossary.md +59 -0
  716. package/skills/specs/references/persistence.md +115 -0
  717. package/skills/specs/references/predictability-check.md +77 -0
  718. package/skills/static-analysis-integration/SKILL.md +235 -0
  719. package/skills/static-analysis-integration/adapters/_envelope.py +26 -0
  720. package/skills/static-analysis-integration/adapters/jscpd-adapter.py +66 -0
  721. package/skills/static-analysis-integration/adapters/lizard-adapter.py +81 -0
  722. package/skills/static-analysis-integration/adapters/mypy-adapter.py +50 -0
  723. package/skills/static-analysis-integration/adapters/mypy-src-layout.py +93 -0
  724. package/skills/static-analysis-integration/adapters/security-review-adapter.py +212 -0
  725. package/skills/static-analysis-integration/maintenance.md +23 -0
  726. package/skills/static-analysis-integration/references/language-setup.md +228 -0
  727. package/skills/static-analysis-integration/references/sarif-parser.md +124 -0
  728. package/skills/static-analysis-integration/references/security-review-adapter.md +118 -0
  729. package/skills/static-analysis-integration/references/tool-configs.md +617 -0
  730. package/skills/static-analysis-integration/rulesets/pmd-quickstart.xml +24 -0
  731. package/skills/stryker-xunit-v2-shim/SKILL.md +274 -0
  732. package/skills/stryker-xunit-v2-shim/references/shim-howto.md +256 -0
  733. package/skills/stryker-xunit-v2-shim/scripts/generate_shim.py +143 -0
  734. package/skills/systematic-debugging/SKILL.md +130 -0
  735. package/skills/telemetry/SKILL.md +75 -0
  736. package/skills/test-audit-disable/SKILL.md +129 -0
  737. package/skills/test-design/SKILL.md +177 -0
  738. package/skills/test-design/scripts/__pycache__/internal_double_detector.cpython-314.pyc +0 -0
  739. package/skills/test-design/scripts/internal_double_detector.py +631 -0
  740. package/skills/test-design-advisor/SKILL.md +166 -0
  741. package/skills/test-driven-development/SKILL.md +169 -0
  742. package/skills/test-health/SKILL.md +262 -0
  743. package/skills/test-improve/SKILL.md +239 -0
  744. package/skills/test-improve/references/phase-0-approach-contract.md +228 -0
  745. package/skills/test-improve/references/phase-1-analyze.md +131 -0
  746. package/skills/test-improve/references/phase-2-baseline.md +121 -0
  747. package/skills/test-improve/references/phase-3-derive-gherkin.md +53 -0
  748. package/skills/test-improve/references/phase-4-plan-fixes.md +34 -0
  749. package/skills/test-improve/references/phase-5-improve.md +215 -0
  750. package/skills/test-improve/references/phase-6-refactor-decision.md +45 -0
  751. package/skills/test-improve/references/phase-7-refactor.md +44 -0
  752. package/skills/test-improve/references/phase-8-validate.md +66 -0
  753. package/skills/test-improve/references/phase-9-close-out-prompt.md +11 -0
  754. package/skills/test-improve/references/phase-9-report.md +62 -0
  755. package/skills/test-improve/references/review-loop.md +92 -0
  756. package/skills/test-improve/templates/executive-summary.md +123 -0
  757. package/skills/threat-modeling/SKILL.md +108 -0
  758. package/skills/triage/SKILL.md +211 -0
  759. package/skills/ubiquitous-language/SKILL.md +192 -0
  760. package/skills/ubiquitous-language/scripts/collect_domain_signals.py +300 -0
  761. package/skills/unfreeze/SKILL.md +37 -0
  762. package/skills/upgrade/SKILL.md +31 -0
  763. package/skills/upgrade/scripts/check_version_drift.py +113 -0
  764. package/skills/upgrade/scripts/enable_autoupdate.py +149 -0
  765. package/skills/version/SKILL.md +25 -0
  766. package/sync/__pycache__/sync_upstream.cpython-314.pyc +0 -0
  767. package/sync/sync_upstream.py +293 -0
  768. package/templates/ACCEPTED-RISKS.md.tmpl +46 -0
  769. package/templates/agents/agent-template.md +151 -0
  770. package/templates/agents/angular-testing.md +66 -0
  771. package/templates/agents/csharp-quality.md +63 -0
  772. package/templates/agents/esm-enforcer.md +52 -0
  773. package/templates/agents/front-end-testing.md +65 -0
  774. package/templates/agents/go-quality.md +65 -0
  775. package/templates/agents/python-quality.md +62 -0
  776. package/templates/agents/react-testing.md +61 -0
  777. package/templates/agents/ts-enforcer.md +60 -0
  778. package/templates/agents/twelve-factor-audit.md +49 -0
  779. package/tools/entropy-check.py +250 -0
  780. package/tools/model-hash-verify.py +213 -0
@@ -0,0 +1,240 @@
1
+ # Testability Design Patterns
2
+
3
+ Reference file for test-review and structure-review agents. When flagging untestable code, use the patterns below to specify the required **production code change** — not a test workaround. Tests never hack around untestable designs; the design must change.
4
+
5
+ Core principle: if code can't be tested through its public API, the production code's design must change. Per Seemann: "The hallmark of a testable design is that it's also a good design."
6
+
7
+ ---
8
+
9
+ ## "I Can't Test This Class" Decision Flow
10
+
11
+ ```
12
+ Can I construct the object with the values I need for the test?
13
+ │
14
+ ├─ YES → Write the test using the public constructor.
15
+ │
16
+ ├─ NO: it has private-set properties with no public constructor
17
+ │ └─ Add a constructor that accepts values directly.
18
+ │ For 5+ parameters: apply Test Data Builder (Pattern 2).
19
+ │
20
+ ├─ NO: it only has a static factory that calls a real database/service
21
+ │ └─ Add a constructor alongside the factory.
22
+ │ The factory uses the constructor internally.
23
+ │
24
+ ├─ NO: the method I need to verify is protected/private
25
+ │ └─ Extract the logic into a public method or invoke through the
26
+ │ public behavior that calls it.
27
+ │
28
+ └─ NO: it depends on a concrete class I can't replace
29
+ └─ Extract an interface for what this class needs.
30
+ Mock the interface in tests, wire the concrete in production.
31
+
32
+ Each NO branch requires a production code change. That IS the work.
33
+ ```
34
+
35
+ > **Legacy code (no tests yet)?** Don't make the design change cold. First get the code under characterization tests using a *behavior-preserving* seam from [`dependency-breaking-techniques.md`](dependency-breaking-techniques.md), guided by effect/pinch reasoning in [`legacy-test-strategy.md`](legacy-test-strategy.md) — then refactor toward the target patterns below. The patterns here are the destination; those techniques are how you get there safely.
36
+
37
+ ---
38
+
39
+ ## Pattern 1: Constructor Injection (Replace Static Factories / Singletons)
40
+
41
+ **Problem**: A class creates its own dependencies internally — `new ConcreteService()`, `ServiceLocator.get(T)`, static calls. The production code can't be given test doubles.
42
+
43
+ **Solution**: Accept collaborators as constructor parameters. Static factories and singletons remain for production wiring; they delegate to the constructor internally.
44
+
45
+ ```
46
+ // BEFORE — untestable
47
+ class OrderProcessor:
48
+ def process(orderId):
49
+ db = Database.getInstance() // static singleton
50
+ logger = new FileLogger() // new-ed up
51
+ ...
52
+
53
+ // AFTER — injectable
54
+ class OrderProcessor:
55
+ constructor(db: IDatabase, logger: ILogger):
56
+ self.db = db
57
+ self.logger = logger
58
+
59
+ def process(orderId):
60
+ ...
61
+ ```
62
+
63
+ **When to apply**: class instantiates its own collaborators; class has no constructor that accepts its dependencies; production code works but tests can't isolate the unit.
64
+
65
+ ---
66
+
67
+ ## Pattern 2: Test Data Builder
68
+
69
+ **Problem**: Domain objects with many properties are painful to construct in every test. Most tests only care about 1-2 properties.
70
+
71
+ **Solution**: A builder in the test project with sensible defaults and a fluent API.
72
+
73
+ ```
74
+ // Builder lives in the test project
75
+ class OrderBuilder:
76
+ customerId = "TEST-CUST"
77
+ amount = 100
78
+ status = "pending"
79
+
80
+ withCustomer(id): self.customerId = id; return self
81
+ withAmount(amt): self.amount = amt; return self
82
+ withStatus(s): self.status = s; return self
83
+ build(): return Order(customerId, amount, status)
84
+
85
+ // In tests — only set what the test cares about
86
+ order = new OrderBuilder().withStatus("cancelled").build()
87
+ ```
88
+
89
+ **When to apply**: domain object has 5+ properties; multiple tests need variants of the same object; default values work for most tests.
90
+
91
+ ---
92
+
93
+ ## Pattern 3: Interface Extraction for Large Contexts
94
+
95
+ **Problem**: Classes depend on a large context object with many properties. Mocking the full context is impractical and fragile.
96
+
97
+ **Solution**: Extract an interface for what each consumer actually needs.
98
+
99
+ ```
100
+ // BEFORE — consumer takes the whole context
101
+ class FileProcessor:
102
+ constructor(context: ProcessingContext) // 40-field object
103
+
104
+ // AFTER — consumer takes only what it needs
105
+ interface IFileProcessorContext:
106
+ tempFilePath(): string
107
+ isReprocess: bool
108
+ fileNames: List<string>
109
+
110
+ class FileProcessor:
111
+ constructor(context: IFileProcessorContext)
112
+ ```
113
+
114
+ The full context class implements this interface. Tests mock only the interface.
115
+
116
+ **When to apply**: context has 30+ properties but any given consumer uses 5-10; creating a full context for each test is excessive.
117
+
118
+ ---
119
+
120
+ ## Pattern 4: Fake Data Generators (Test Data with Variation)
121
+
122
+ **Problem**: Tests need realistic but controlled test data across many scenarios.
123
+
124
+ **Solution**: A typed faker or factory in the test project that generates valid domain values.
125
+
126
+ ```
127
+ // Instead of magic literals
128
+ order = new Order(customerId="ABC123", amount=150.00, ...)
129
+
130
+ // Use a generator that makes intent clear
131
+ order = OrderFaker.pending()
132
+ order = OrderFaker.withAmount(5000)
133
+ orders = OrderFaker.generateBatch(100)
134
+ ```
135
+
136
+ **When to apply**: tests need multiple records with varied data; domain validation rules constrain acceptable values; tests should not use unexplained magic numbers.
137
+
138
+ ---
139
+
140
+ ## The Design-for-Testability seam family
141
+
142
+ Constructor Injection (Pattern 1) is the common case, but it is one of a named family of seams from *xUnit Test Patterns* Ch. 26. Pick by *how* the dependency reaches the SUT and *why* it's hard to test.
143
+
144
+ | Seam | The change | Use when | Cost |
145
+ | ------ | ----------- | ---------- | ------ |
146
+ | **Dependency Injection** (Pattern 1) | Pass collaborators in (constructor/setter) | You control construction and can thread the dependency through | Low; the default |
147
+ | **Dependency Lookup** | SUT asks a broker/service-locator for the collaborator; the test configures the broker | The DOC is buried deep and threading it through every caller would be messy (e.g. a Fake DB behind a service facade) | Medium; hides the dependency, but far easier to **retrofit onto legacy** code than DI |
148
+ | **Humble Object** | Extract logic out of a hard-to-instantiate shell into a plain, synchronously-testable object | Logic is trapped in a UI control, framework callback, thread, or transaction boundary | Medium; the shell becomes a thin adapter |
149
+ | **Test Hook** | Conditional `if (testing)` behavior baked into production code | **Last resort only** — nothing else can break the dependency | High; it *is* the Test Logic in Production smell |
150
+
151
+ ---
152
+
153
+ ## Pattern 5: Dependency Lookup (retrofit seam for legacy)
154
+
155
+ **Problem**: A collaborator is constructed deep inside the system and there's no clean path to pass a test double down from the test — wiring DI through every intermediate layer would be invasive.
156
+
157
+ **Solution**: The SUT requests its collaborator from a **component broker / service locator** instead of `new`-ing it. The test (or a Setup Decorator) configures the broker to hand back a double.
158
+
159
+ ```
160
+ // Production wiring registers the real implementation
161
+ ServiceRegistry.register(IClock, SystemClock)
162
+
163
+ // SUT looks up rather than receiving
164
+ class ExpiryService:
165
+ def isExpired(token):
166
+ clock = ServiceRegistry.get(IClock) // broker, not `new`
167
+ ...
168
+
169
+ // Test reconfigures the broker, no constructor threading needed
170
+ ServiceRegistry.register(IClock, FakeClock(at="2026-01-01"))
171
+ ```
172
+
173
+ **When to apply**: retrofitting tests onto legacy code where DI would touch too many call sites; a deep DOC (data-access layer, facade-backed service) you want to replace with a Fake for a whole test run. **Trade-off**: the dependency is no longer visible in the signature, and the broker is global state that each test must reset — keep DI the default and reach for lookup when threading is genuinely impractical.
174
+
175
+ ---
176
+
177
+ ## Pattern 6: Humble Object (rescue logic from untestable shells)
178
+
179
+ **Problem**: Meaningful logic lives inside an object that's expensive or impossible to instantiate in a unit test — a GUI control, a framework callback, an async worker/thread, a transaction-managing controller. The asynchronicity or framework coupling forces slow, nondeterministic tests, so the logic ends up untested.
180
+
181
+ **Solution**: Extract **all** the logic into a separate, plain object that's testable through synchronous calls. The original shell becomes a *humble* adapter that holds no logic — it just forwards framework calls to the testable object. The shell is so thin it needs no tests of its own.
182
+
183
+ ```
184
+ // BEFORE — logic trapped in a UI/framework/thread shell
185
+ class OrderView(FrameworkWidget):
186
+ def onSubmit(): // framework-invoked, hard to test
187
+ ...validation + pricing + state transitions...
188
+
189
+ // AFTER — Humble shell + testable presenter
190
+ class OrderPresenter: // plain object, fully unit-testable
191
+ def submit(input) -> ViewModel: ...the real logic...
192
+
193
+ class OrderView(FrameworkWidget):
194
+ def onSubmit(): // humble: no logic, just delegation
195
+ self.render(self.presenter.submit(self.readInput()))
196
+ ```
197
+
198
+ Named variations: **Humble Dialog** (UI logic → presenter), **Humble Executable** (logic out of a thread/active object), **Humble Transaction Controller** (the business method takes no responsibility for begin/commit, so a test can wrap it in a transaction and roll back — the prerequisite for Transaction Rollback Teardown in `database-test-patterns.md`).
199
+
200
+ **When to apply**: nontrivial logic in any component that's problematic to instantiate because it depends on a framework, a UI toolkit, a thread, or a transaction. This is the production-code answer to the *Minimize Untestable Code* principle.
201
+
202
+ ---
203
+
204
+ ## Pattern 7: Test Hook (last resort — prefer any seam above)
205
+
206
+ **Problem**: None of the above seams can be introduced (e.g. a closed dependency that can't be injected, looked up, or subclassed), yet the code must be brought under test.
207
+
208
+ **Solution**: A conditional in production code that alters behavior under test. **This is deliberately listed last because it *is* the Test Logic in Production smell** (`test-smells.md`): the tested path no longer equals the shipped path, and the hook can fail in production. Use only when no structural seam is possible, isolate the hook tightly, and treat its removal (by refactoring to a real seam) as outstanding debt.
209
+
210
+ ```
211
+ // AVOID unless nothing else works — and plan to delete it
212
+ def charge(amount):
213
+ if TEST_MODE: return FakeGatewayResult.ok() // test logic in production
214
+ return realGateway.charge(amount)
215
+ ```
216
+
217
+ Prefer, in order: Dependency Injection → Dependency Lookup → Test-Specific Subclass (override the seam method in a test subclass; see `test-doubles.md`) → Humble Object. A Test Hook is an admission that the design resisted all of them.
218
+
219
+ ---
220
+
221
+ ## Anti-Patterns Table
222
+
223
+ | Anti-Pattern | Why It's Wrong | Correct Alternative |
224
+ | --- | --- | --- |
225
+ | `InternalsVisibleTo` + `internal set` for tests | Weakens encapsulation for test convenience | Add a public constructor that accepts values |
226
+ | Reflection into private members as primary strategy — Java: `getDeclaredMethod`/`getDeclaredField` + `setAccessible(true)`, `Method.invoke` on a private/protected member; C#: `Type.GetMethod(..., BindingFlags.NonPublic \| BindingFlags.Instance)`, `Type.InvokeMember`; Python: `getattr`/`setattr`/`hasattr` on a name-mangled (`_ClassName__attr`) or underscore-prefixed attribute; JS/TS: bracket-notation into a `private`/non-exported member, or `Object.getOwnPropertyDescriptor`/`Object.defineProperty` to reach one | Fragile; breaks on rename; masks coupling — this is an architecture/encapsulation issue the test is reaching around, not a test-hygiene nit | Pick by shape of the code: extract the logic into a collaborator with its own public seam; relax visibility to package-private/internal only when a production collaborator in the same module/assembly independently needs it — never as a grant solely so the test can reach in (that recreates the `InternalsVisibleTo`/`@VisibleForTesting` rows above); or test the behavior through the existing public API (if already reachable) |
227
+ | Static test helper that mutates private state | Bypasses object invariants | Use Test Data Builder with public construction |
228
+ | Mocking concrete classes | Fragile; requires virtual/open methods; masks design issues | Extract interface; mock the interface |
229
+ | Tests configuring global/static state | Shared state causes order-dependent failures | Inject dependencies through constructors |
230
+ | Changing `private set` to `public set` for test access | Removes invariant protection | Add constructor parameter instead |
231
+ | `[InternalsVisibleTo]` / `@VisibleForTesting` as the primary testability mechanism | Couples test and production assemblies | Redesign the API surface so tests don't need access to internals |
232
+
233
+ ---
234
+
235
+ ## References
236
+
237
+ - Michael Feathers, *Working Effectively with Legacy Code* — seam insertion, Parameterize Constructor, Extract Interface
238
+ - Mark Seemann & Steven van Deursen, *Dependency Injection: Principles, Practices, and Patterns* — Pure DI, Composition Root
239
+ - Steve Freeman & Nat Pryce, *Growing Object-Oriented Software, Guided by Tests* — outside-in TDD, mock-roles-not-objects
240
+ - Vladimir Khorikov, *Unit Testing: Principles, Practices, and Patterns* — resilient vs. fragile tests
@@ -0,0 +1,44 @@
1
+ # Agile Testing Quadrants
2
+
3
+ Reference file for the `test-health` skill (project-wide audit) and `test-design-advisor`. The quadrants are a *coverage-completeness* lens — they answer "what **kind** of testing is missing?", orthogonal to `test-pyramid.md`'s "what **layer**?". Use them to find blind spots an all-unit suite can't see.
4
+
5
+ Source: Brian Marick's testing matrix, popularized by Lisa Crispin & Janet Gregory, *Agile Testing* / *More Agile Testing*. Language- and stack-agnostic.
6
+
7
+ Two axes: **business-facing ↔ technology-facing** (what the test speaks to) and **supporting the team ↔ critiquing the product** (does it guide building, or probe the finished thing).
8
+
9
+ ---
10
+
11
+ ## The Four Quadrants
12
+
13
+ | Q | Facing × Stance | Tests | Mode |
14
+ |---|-----------------|-------|------|
15
+ | **Q1** | technology · support | unit, component, integration | automated |
16
+ | **Q2** | business · support | functional / acceptance, story tests, BDD examples, prototypes | automated + manual |
17
+ | **Q3** | business · critique | exploratory, usability, UAT, alpha/beta | manual |
18
+ | **Q4** | technology · critique | performance, load, security, resilience, the "-ilities" | tool-driven |
19
+
20
+ Q1+Q2 **guide development** (write them first — they're specifications). Q3+Q4 **probe the built product** for what specifications miss.
21
+
22
+ ---
23
+
24
+ ## Reading coverage against the quadrants
25
+
26
+ For each quadrant, classify the suite as **strong / thin / empty**, then name the business impact of a gap:
27
+
28
+ | Quadrant weak/empty | What slips through | How to strengthen |
29
+ |---------------------|--------------------|-------------------|
30
+ | **Q1 thin** | logic/wiring regressions | push checks down the pyramid (`test-pyramid.md`) |
31
+ | **Q2 empty** | "built the wrong thing" — no shared definition of done | add acceptance/BDD examples before coding (see the `specs` skill) |
32
+ | **Q3 empty** | bugs only a human notices: confusing flows, broken edge journeys | charter exploratory sessions (`exploratory-testing-field-guide.md`) |
33
+ | **Q4 empty** | non-functional failures in prod: slow, insecure, falls over under load | add perf/security/resilience checks; route to `security-review`, `performance-review` |
34
+
35
+ ---
36
+
37
+ ## Anti-pattern: "Q3 as the gate"
38
+
39
+ Leaning on manual exploratory/UAT (Q3) as the *primary* safety net — instead of Q1/Q2 automation — is the quadrant form of the ice-cream cone (`test-pyramid.md`). Exploratory testing **finds new** problems; it must not be the regression net. If Q3 is the only thing catching regressions, the fix is more Q1/Q2, not more manual testing.
40
+
41
+ ## Boundaries
42
+
43
+ - A suite need not fill every quadrant equally — weight by risk. A pure library leans Q1/Q4; a user-facing app needs Q2/Q3. Flag *empty* quadrants where the product's risk clearly demands coverage, not arithmetic imbalance.
44
+ - This file classifies; it does not score. Quantitative suite scoring stays in `farley-score`.
@@ -0,0 +1,15 @@
1
+ # Approval testing
2
+
3
+ Overlay technique for `test-design-advisor`. Loaded only when the trigger matches.
4
+
5
+ **Trigger.** A behavior produces a large or structured **text artifact** — rendered HTML/Markdown, JSON/CSV export, generated code, a report, a log — and asserting field-by-field would be unreadable.
6
+
7
+ **What it is.** Capture the output once, have a human approve it as the *approved* (golden) file; the test then diffs current output against approved. A change surfaces as a diff to re-approve, not a rewritten assertion.
8
+
9
+ **When to use.** Output is wide, stable, and a diff is more legible than N equality asserts. Excellent for characterizing legacy output before refactoring.
10
+
11
+ **Trade-offs / cost.** Approved files must be reviewed on every intended change — rubber-stamping defeats the test. Non-deterministic fields (timestamps, ids, ordering) must be scrubbed/normalized first or the test flaps. Store approved files in VCS.
12
+
13
+ **Minimal shape.** `approve("invoice-html", renderInvoice(order))` → first run writes `invoice-html.approved`; later runs diff against it.
14
+
15
+ **Complements.** A verification *style* at unit/component/integration — not a new pyramid layer. For **visual/CSS** fidelity use `screenshot.md` instead; for **text** correctness, approval is cheaper. Tools: ApprovalTests, Verify, jest snapshots (treat snapshots as approval — review them).
@@ -0,0 +1,17 @@
1
+ # Chaos / resilience testing
2
+
3
+ Overlay technique for `test-design-advisor`. Loaded only when the trigger matches.
4
+
5
+ **Trigger.** A behavior makes a **resilience claim under dependency failure** — "retries on a 5xx", "degrades gracefully when the cache is down", "the circuit breaker opens", "recovers after the broker reconnects". The risk is in the *failure* path, which happy-path tests never exercise.
6
+
7
+ **What it is.** Deliberately inject failure — latency, errors, dropped connections, killed dependencies — and assert the system's degradation/recovery behavior matches the claim.
8
+
9
+ **When to use.** Distributed systems, anything with retries/timeouts/circuit-breakers/fallbacks, jobs that must survive a mid-run dependency outage.
10
+
11
+ **Scope split (important).** Inject failure at the **owned adapter** in a component test to verify retry/fallback *logic* deterministically (pre-merge). Reserve infrastructure-level chaos (kill a pod, partition the network) for **out-of-band/staging** — it is non-deterministic and never belongs in the pre-merge gate (`cd-test-architecture.md`).
12
+
13
+ **Trade-offs / cost.** Infra chaos is flaky and slow; needs a controlled environment and observability to read results. Start with adapter-level fault injection — most resilience bugs surface there cheaply.
14
+
15
+ **Minimal shape.** Stub the payment adapter to throw `Timeout` twice then succeed → assert two retries then success.
16
+
17
+ **Complements.** Component (adapter fault injection) and out-of-band (infra). Q4 in `testing-quadrants.md`. Tools: Toxiproxy, fault-injecting test doubles, Chaos Monkey / Litmus (infra).
@@ -0,0 +1,15 @@
1
+ # Fuzz testing
2
+
3
+ Overlay technique for `test-design-advisor`. Loaded only when the trigger matches.
4
+
5
+ **Trigger.** A behavior **parses or accepts untrusted/unstructured input** — a file/format parser, protocol decoder, deserializer, public API request body, anything at a trust boundary where malformed input is an attack surface.
6
+
7
+ **What it is.** Feed large volumes of malformed, random, or mutated input and assert the code **never crashes, hangs, or corrupts state** — it either handles or cleanly rejects. Coverage-guided fuzzers mutate toward new code paths.
8
+
9
+ **When to use.** Input parsers, decoders, file/upload handling, anything reachable by an attacker. The goal is robustness, not a specific output.
10
+
11
+ **Trade-offs / cost.** Findings are crashes/hangs, not "wrong answer" — pair with example/property tests for correctness. Needs a sanitizer or crash oracle to be useful; can be slow; corpus and seeds need maintenance.
12
+
13
+ **Minimal shape.** `fuzz(parseConfig)` runs mutated byte strings; any unhandled exception / OOM / timeout is a failure.
14
+
15
+ **Complements.** Sits at unit/integration on the parser; security-adjacent — cross-reference `security-review` for the trust-boundary finding. For *valid* inputs obeying a law, use `property-based.md`; for *declared schema* conformance, see `schema-validation.md`. Tools: libFuzzer/AFL, jazzer (Java), Atheris (Python), go-fuzz.
@@ -0,0 +1,15 @@
1
+ # Property-based testing
2
+
3
+ Overlay technique for `test-design-advisor`. Loaded only when the trigger matches.
4
+
5
+ **Trigger.** A behavior has an **invariant that holds for all inputs** — a roundtrip (`decode(encode(x)) == x`), a mathematical law (commutativity, idempotency, sum-preservation), an ordering/closure property — rather than a handful of known example pairs.
6
+
7
+ **What it is.** Instead of fixed examples, declare a property and a generator; the framework generates hundreds of random inputs, and on failure **shrinks** to the minimal counterexample.
8
+
9
+ **When to use.** Parsers/serializers, encoders, money/quantity math, sorting/merging, state machines, anything with an algebraic law. Pairs well with a few example tests for documentation.
10
+
11
+ **Trade-offs / cost.** Finding the right property is the hard part — a weak property passes vacuously. Generators for complex domain types take effort. Slower than example tests; seed failures so they reproduce.
12
+
13
+ **Minimal shape.** `forAll(integers, integers, (a,b) => add(a,b) === add(b,a))`.
14
+
15
+ **Complements.** A unit/integration *technique*, not a layer — apply it at the layer where the invariant lives. Tools: fast-check (JS), Hypothesis (Python), jqwik (Java), FsCheck (.NET). For *malformed/hostile* inputs rather than law-checking, see `fuzz.md`.
@@ -0,0 +1,15 @@
1
+ # Schema-validation testing
2
+
3
+ Overlay technique for `test-design-advisor`. Loaded only when the trigger matches.
4
+
5
+ **Trigger.** A behavior produces or consumes a payload governed by a **declared schema** — an OpenAPI/Swagger spec, JSON Schema, Avro/Protobuf, a GraphQL type. The risk is the payload **drifting from its own declared shape**.
6
+
7
+ **What it is.** Assert that real request/response payloads validate against the schema artifact — and, in reverse, that the schema matches what the code emits. The schema becomes an executable spec, not just documentation.
8
+
9
+ **When to use.** Public/partner APIs with a published spec, event payloads on a shared bus, config files with a schema, codegen boundaries. Catches "docs say one thing, code does another" before consumers do.
10
+
11
+ **Trade-offs / cost.** Validates *shape*, not *meaning* — a semantically wrong-but-well-formed payload still passes (cover that with unit/contract tests). The schema must be kept authoritative or the check rots.
12
+
13
+ **Minimal shape.** `expect(validate(openapi.paths['/orders'].post.response, actualBody)).toPass()`.
14
+
15
+ **Complements.** Integration/component on the boundary. **Not a substitute for contract testing** between two owned services — for consumer↔provider agreement route to `microservice-testing.md` (CDC). Schema-validation checks conformance to a *declared* schema; CDC checks two parties still *agree*. Cross-reference `security-review` for input-validation at trust boundaries. Tools: ajv, openapi-validator, Spectral, schemathesis.
@@ -0,0 +1,15 @@
1
+ # Screenshot / visual-regression testing
2
+
3
+ Overlay technique for `test-design-advisor`. Loaded only when the trigger matches.
4
+
5
+ **Trigger.** A behavior's correctness is **visual** — CSS layout, theming, charts/canvas, print/PDF rendering, responsive breakpoints — where the right markup can still *look* wrong and a DOM/text assertion can't catch it.
6
+
7
+ **What it is.** Render the component/page, capture an image, and diff it against an approved baseline image (approval testing for pixels). A visual change surfaces as an image diff to re-approve.
8
+
9
+ **When to use.** Design-system components, layout-critical pages, anything where appearance is the requirement. Pin a few representative states, not every page.
10
+
11
+ **Trade-offs / cost.** The most maintenance-heavy technique: baselines drift across OS/browser/font-rendering → false diffs. Mitigate with a consistent render environment (containerized/CI runner), tolerance thresholds, and masking dynamic regions. Baselines live in VCS and need review on every intended change.
12
+
13
+ **Minimal shape.** `await expect(page).toHaveScreenshot('cart-badge.png')`.
14
+
15
+ **Complements.** Component/E2E layer. Decision: **text/markup** correctness → `approval.md` (cheaper, stabler); **CSS/visual** fidelity → screenshot. Often the visual half of a Gate D / Gate C behavior (`test-layer-gates.md`). Tools: Playwright/Percy/Chromatic/Storybook test-runner, BackstopJS.
@@ -0,0 +1,198 @@
1
+ # Three-Phase Workflow
2
+
3
+ The orchestrator's phase reference, loaded on demand. `agents/orchestrator.md`
4
+ § Three-Phase Workflow carries the always-on part — the phase list, each phase's
5
+ goal and human gate, and the invariants that hold across all three. Everything
6
+ below is the detail a session needs *once it is actually running a phase*: the
7
+ persona rosters to dispatch, the conditional dispatch rules, the review
8
+ checkpoints, and the wave mechanics. Read the section for the phase you are in
9
+ via its anchor, not the whole file.
10
+
11
+ Notes on what `scripts/orchestrator.py` — the `Enforcement: script` deterministic
12
+ implementation — actually does, and where it diverges from the policy below, live
13
+ in `${CLAUDE_PLUGIN_ROOT}/knowledge/orchestrator-script-implementation.md`. They are needed when running or working on that script, not
14
+ when following the policy interactively.
15
+
16
+ Every non-trivial task follows three explicit phases. Each phase runs in minimal context, and a human review gate separates each phase. The output of each phase is a structured progress file written to `.claude/memory/` that onboards the next phase.
17
+
18
+ ## Phase 1: Research
19
+
20
+ - **Goal**: Understand how the system works, identify all relevant files, locate the problem or feature surface area
21
+ - **Agents**: `codebase-recon` (gated on RECON artifact freshness — see Codebase Recon dispatch below), `architect`, `data-flow-tracer` (always dispatched — see the Research persona roster in `${CLAUDE_PLUGIN_ROOT}/knowledge/orchestrator-script-implementation.md#phase-1-research`), `security-engineer` (conditional — see Security Engineer dispatch), plus the Orchestrator itself and further sub-agents for exploration as needed (context isolation — sub-agents search, read, and return concise findings so the parent context stays clean)
22
+ - **Output**: A research progress file with file paths, line numbers, data flows, and key findings
23
+ - **Design doc**: For non-trivial features (see Design Doc skill for criteria), produce a design document at `docs/specs/{feature-name}.md` with problem statement, proposed approach, alternatives, key decisions, and scope boundaries. The human approves the design doc as part of the research gate.
24
+ - **Human gate**: Human reviews the research findings and design doc before planning begins. Catching a misunderstanding here prevents hundreds of bad lines of code downstream.
25
+ - **Context**: Compact after this phase — write progress file, start fresh context for Phase 2
26
+
27
+ Script behavior and known gaps: `${CLAUDE_PLUGIN_ROOT}/knowledge/orchestrator-script-implementation.md#phase-1-research`.
28
+
29
+ ### Codebase Recon dispatch
30
+
31
+ Dispatch `codebase-recon` as a sub-agent **before any other exploration**,
32
+ at the start of Research, when no `.claude/memory/recon-<slug>.json`
33
+ artifact exists, or the existing one is more than 24 hours old (`<slug>` is
34
+ the repo basename) — skip the dispatch (silently) when a fresh artifact is
35
+ present. It returns entry points, dependency graph, security surface, and
36
+ git history in a structured artifact (`.claude/memory/recon-<slug>.json` /
37
+ `.claude/memory/recon-<slug>.md`, checked for freshness against the `.json`
38
+ half since that's the machine-readable artifact other agents consume)
39
+ intended to onboard the Architect and Security Engineer without those
40
+ agents needing to re-read the codebase themselves.
41
+
42
+ Script behavior and known gaps: `${CLAUDE_PLUGIN_ROOT}/knowledge/orchestrator-script-implementation.md#phase-1-research`.
43
+
44
+ ### Security Engineer dispatch
45
+
46
+ Dispatch `security-engineer` during Research when any of these signals is
47
+ present: the task touches authentication, authorization, cryptography,
48
+ session management, or secrets handling; it introduces a new external
49
+ integration or API surface; or the user explicitly asks to "threat model
50
+ this", "design this securely", or "what's the attack surface here" (these
51
+ three match `agents/security-engineer.md`'s own dispatch description) — or a
52
+ recent `/code-review` run's `security-review` produced a `fail` verdict with
53
+ high-severity findings on this area (orchestrator-owned: only the
54
+ orchestrator sees `/code-review` history). Its `effort: high` cost is only
55
+ justified on security-relevant work, so this dispatch stays conditional
56
+ rather than unconditional.
57
+
58
+ Script behavior and known gaps: `${CLAUDE_PLUGIN_ROOT}/knowledge/orchestrator-script-implementation.md#phase-1-research`.
59
+
60
+ ## Phase 2: Plan
61
+
62
+ - **Goal**: Specify every change to be made — files, snippets, test strategy, verification steps
63
+ - **Agents**: `product-manager`, `architect`, `qa-engineer` (core trio that drafts the plan — always dispatched, see the Plan persona roster in `${CLAUDE_PLUGIN_ROOT}/knowledge/orchestrator-script-implementation.md#phase-2-plan`), then the `plan-review-*` critics reviewing that draft (below)
64
+ - **Input**: Research progress file from Phase 1 + approved design doc (if produced in Phase 1)
65
+ - **Output**: An implementation plan with explicit file changes, test expectations, and acceptance criteria
66
+ - **Automated plan review**: Before the human gate, dispatch the plan review
67
+ personas in parallel as sub-agents — `plan-review-acceptance`,
68
+ `plan-review-design`, `plan-review-ux`, `plan-review-strategic`,
69
+ `plan-review-parallelization`. Each is a registered agent
70
+ (`agents/plan-review-<name>.md`); dispatch by `subagent_type` like any
71
+ other agent — the harness reads its `model:`/`effort:` frontmatter
72
+ natively, no dispatch-time override needed. The reviewer set scales to
73
+ plan tier and complexity; see the plan skill's
74
+ [Run plan review personas step](../skills/plan/SKILL.md#5-run-plan-review-personas)
75
+ for the tier classification (that table is the single source of truth —
76
+ do not re-duplicate the reviewer set here, it drifts).
77
+
78
+ Each returns a `verdict` of `approve` or `needs-revision`. If **any**
79
+ dispatched reviewer returns `needs-revision`, address the blocker issues
80
+ before presenting to the human. Aggregate all findings (including
81
+ warnings from approving reviewers) into the plan review summary.
82
+ - **Human gate**: Human reviews the plan and the aggregated review findings. This is the primary review artifact — 200 lines of plan is far more reviewable than 2,000 lines of code. If the plan is wrong, fix it here, not in code.
83
+ - **Design intent: no choice made during Implementation is meant to compensate for a weak plan.** A plan carrying an unresolved `needs-revision` verdict (a blocker, or 3+ warnings — 2+ for `plan-review-parallelization` — per `${CLAUDE_PLUGIN_ROOT}/knowledge/plan-review-rubric.md#verdict-rules`) is revised and re-reviewed before the human gate — see the plan skill's [Run plan review personas step](../skills/plan/SKILL.md#5-run-plan-review-personas) for that iteration cap and its escalation path — and never carried silently into Phase 3. Code-First Small Batches (Phase 3's sole cadence, per ADR 0017) is not a substitute for plan quality, it is what a *good* plan gets executed with. When a plan looks weak going into the human gate, the fix is another Phase 2 iteration, never a Phase 3 workaround. See `${CLAUDE_PLUGIN_ROOT}/knowledge/test-cadence-tradeoffs.md#the-decision-rule` for the evidence bar an alternative Phase 3 cadence has to clear before it changes this.
84
+ - **Context**: Compact after this phase — write progress file, start fresh context for Phase 3
85
+
86
+ Script behavior and known gaps: `${CLAUDE_PLUGIN_ROOT}/knowledge/orchestrator-script-implementation.md#phase-2-plan`.
87
+
88
+ ## Phase 3: Implement
89
+
90
+ - **Goal**: Execute the plan. Write code, run tests, verify at each step.
91
+ - **Agents**: Software Engineer (primary), QA Engineer (validation), others as needed
92
+ - **Input**: Plan progress file from Phase 2
93
+ - **Subagent dispatch**: Dispatch the `software-engineer` agent by `subagent_type` when dispatching implementation subagents, scoped to a single plan step — the harness reads its `model:`/`effort:` frontmatter natively, no dispatch-time override needed. For parallel implementation of independent units, prefer `isolation: "worktree"` on the Agent tool to give each subagent its own git worktree. Disjoint *final* file sets are not sufficient justification to skip it: two parallel subagents sharing one working directory can still observe each other's mid-edit, intermediate state — a shared production file transiting a broken half-edited form at the exact moment a sibling's own test run reads it — even when neither subagent's final diff ever touches the other's files (#1609). If worktree isolation is skipped for a given dispatch because its per-agent cost (~200-500ms + disk) isn't justified for that unit of work, treat any test failure one parallel subagent reports while siblings are still running as provisional: re-verify it once every parallel subagent in that batch has completed before treating it as a real regression, not a timing artifact. **Never instruct a dispatched parallel worker to dispatch its own review agent.** A subagent that calls the Agent tool itself does not receive that call's completion notification the way the top-level session does — notifications for a subagent's own dispatched children route to the top-level session instead, so a worker that dispatches a review and then stops to wait for it stalls indefinitely (#1881). This applies to any fan-out of independent units, not only the wave mechanism documented in § Wave-aware build dispatch: each parallel worker fixes/implements, tests, commits, and pushes only; review is dispatched once, from the top level, after every worker in the batch has finished — never once per worker. (This is why the Three-stage inline review below is described as something the orchestrator runs after each unit completes, not something a software-engineer subagent runs on itself.)
94
+ - **Cadence enforcement**: The Software Engineer follows the single per-behavior cadence for every unit — Code-First Small Batches (IMPLEMENT → TEST → REFACTOR), per `docs/experiments/RECOMMENDATIONS.md` Rec 3. The orchestrator verifies that each unit's output includes the cadence's verification evidence: green full-suite output. Defect fixes are the one exception — they follow `systematic-debugging`'s mandatory Phase 4 gate, which requires a failing test that reproduces the bug before any fix code is written.
95
+ - **Output**: Working code that passes all tests, acceptance criteria, and code review
96
+ - **Three-stage inline review**: After each discrete unit of work completes, run the deterministic static self-heal pass to pass-or-cap (`skills/build/references/static-self-heal.md`), then spec-compliance, then quality, then browser verification for UI changes:
97
+ 1. **Stage 1 — Spec compliance**: Dispatch the `spec-reviewer` agent by `subagent_type`. Does the code match the spec? If fail → fix before proceeding to Stage 2. (This is a distinct, narrower per-step check than the `spec-compliance-review` agent used as the first gate before the final `/code-review` — see § Inline review checkpoint below and `${CLAUDE_PLUGIN_ROOT}/knowledge/agent-registry.md#review-agents` for how the two differ.)
98
+ 2. **Stage 2 — Code quality**: Dispatch the `quality-reviewer` agent by `subagent_type` to run the standard **Inline Review Checkpoint** (see below). Is the code high quality?
99
+ 3. **Stage 3 — Browser verification (UI changes only)**: If the plan step involves UI components, run `/browse` in automated smoke test mode against the running dev server. Capture screenshots, verify rendering, and check basic interaction. If the dev server is not running, skip with a warning (do not fail). Timeout: 30 seconds. Failures enter the review loop (max 2 iterations). This stage is skipped for non-UI changes.
100
+ - **Final verify**: After all units complete and tests pass, run `/code-review` on all modified files:
101
+ - `fail` → Software Engineer addresses critical issues, re-run review
102
+ - `warn` → include findings in human gate summary
103
+ - `pass` → proceed to doc review
104
+ - **Doc review**: Before the human gate, invoke `dev-team:tech-writer` to review all documentation affected by the changes:
105
+ - Any behavioral or architectural change → check `docs/agent-architecture.md`, `README.md`
106
+ - Any configuration or tooling change → check `docs/agent-architecture.md` (Governance section)
107
+ - Any agent or skill change → check `CLAUDE.md`, `docs/agent_info.md`, `docs/team-structure.md`; regenerate `docs/skills.md` (generated — `hooks/lib/build_skills_index.py`)
108
+ - Tech-writer updates outdated sections and confirms all docs reflect current behavior before proceeding
109
+ - **Human gate**: Human reviews the final output. If the plan was good, implementation review is lightweight.
110
+ - **Context**: If implementation is large, compact mid-phase — update the plan progress file with completed steps and continue in a fresh context
111
+
112
+ Script behavior and known gaps: `${CLAUDE_PLUGIN_ROOT}/knowledge/orchestrator-script-implementation.md#phase-3-implement`.
113
+
114
+ ### Wave-aware build dispatch
115
+
116
+ During `/build`, the orchestrator executes the plan **wave by wave** (the plan's `## Parallelization` schedule from `scripts/plan_waves.py`):
117
+
118
+ 1. **Resolve** the wave schedule (`build_wave.py`) and the effective concurrency (`build_jobs.py` → `min(--jobs, DEV_TEAM_MAX_PARALLEL_BUILDS, wave width)`).
119
+ 2. **Dispatch** each independent slice in the wave to its own git worktree (`isolation: "worktree"`) up to that concurrency — each runs its full per-behavior cycle (Code-First Small Batches) + inline review in isolation.
120
+ 3. **Barrier + reconcile** (`build_wave_reconcile.py`): order-independently merge the wave's slice branches, gate on the full suite, and only then start the next wave. A failing slice or a reconcile conflict halts loudly (names the offender, preserves succeeded worktrees, prints the resume command) and starts no next-wave slice.
121
+
122
+ Effective concurrency 1 (fully-dependent plan, `--jobs 1`, or `DEV_TEAM_MAX_PARALLEL_BUILDS=1`) degrades to sequential single-worktree build with no fan-out or reconcile.
123
+
124
+ > Read a slice's status during a wave only from its structured result, never from a live transcript read into this orchestrating context — see `docs/agent-architecture.md` → Subagent status checks.
125
+
126
+ **`worktree.baseRef` prerequisite (issue #553).** Worktree fan-out only works when Claude Code's `worktree.baseRef` setting is `"head"` — otherwise each subagent worktree branches from `origin/<default>` and cannot see the caller's uncommitted-to-remote spec, plan, or prior-wave commits. Users must set this in **`.claude/settings.json`** (project scope) or **`~/.claude/settings.json`** (user scope); plugin-scope `plugins/<name>/settings.json` and project-local `.claude/settings.local.json` are **not** honored by 2.1.198's worktree isolation. `/build`'s Step 4 detect-and-warn surfaces the requirement loudly on every invocation until the user sets it (or opts out with `DEV_TEAM_WORKTREE_BASE_FRESH=1`). Full audit trail: `docs/spikes/worktree-baseref-head-spike.md`.
127
+
128
+ ### Review depth by complexity
129
+
130
+ Each plan step includes a **Complexity** classification that controls review depth:
131
+
132
+ | Complexity | Inline review behavior | Granularity |
133
+ |------------|----------------------|-------------|
134
+ | `trivial` | Skip inline review entirely. The final `/code-review` covers all files. | — |
135
+ | `standard` | Run spec-compliance + quality agents relevant to the change type (see table below). | **Batched at the slice boundary** — one pass over the slice's accumulated `standard`/`trivial` changes once all its steps are green, not per step. |
136
+ | `complex` | Run spec-compliance + full quality suite including high-effort agents (security-review, domain-review, arch-review). | **Per step** — smaller blast radius per fix. |
137
+
138
+ If a step has no complexity annotation, default to `standard`.
139
+
140
+ Each checkpoint that runs records a find/fix/no-op outcome to `.claude/metrics/review-value.jsonl` (#348) so the review overhead is measurable and the tiering can be evidence-based.
141
+
142
+ ### Inline review checkpoint
143
+
144
+ After each discrete unit of work classified as **standard** or **complex** (a function, a module, a feature slice — as defined in the Phase 2 plan):
145
+
146
+ **Step 1 — Select agents by what changed:**
147
+
148
+ | Changed | Agents to run |
149
+ |---|---|
150
+ | JS/TS functions | naming-review, js-fp-review |
151
+ | Test files | test-review |
152
+ | API surface / auth | security-review |
153
+ | Domain/business logic | domain-review |
154
+ | UI components | a11y-review, structure-review, component-architecture-review |
155
+ | Agent or command files | eval-compliance-check hook runs automatically; also run /agent-audit |
156
+ | Dockerfile or .dockerignore | docker-image-audit skill |
157
+ | Documentation files (.md) | doc-review |
158
+ | Architecture/dependency changes | arch-review |
159
+ | All changes | structure-review as a baseline |
160
+ | All changes (before quality review) | spec-compliance-review as first gate |
161
+
162
+ **Step 2 — Run selected agents in parallel** using the Agent tool by `subagent_type` — the harness reads each agent's `model:`/`effort:` frontmatter natively per `agents/orchestrator.md` § Model/Effort Resolution.
163
+
164
+ When the selection above would dispatch 5+ agents in one wave, note the coordination-cost signal and consider batching high-overlap lenses per `${CLAUDE_PLUGIN_ROOT}/knowledge/wave-consolidation-guidance.md#when-it-applies` — advisory only; dispatch still proceeds.
165
+
166
+ **Step 3 — Aggregate findings and apply Review Loop:**
167
+
168
+ - `pass` / `warn` → log findings in phase output, continue
169
+ - `fail` → enter the **Review Loop** below
170
+
171
+ ### Review loop
172
+
173
+ When any checkpoint agent returns `fail`:
174
+
175
+ 1. Classify issues by actionability (same criteria as `/code-review` step 5):
176
+ - **Actionable**: severity `error` or `warning` with confidence `high` or `medium`
177
+ - **Human-required**: confidence `none` — log and skip, do not attempt auto-fix
178
+ 2. For actionable issues, apply the minimal fix directly:
179
+ - Apply file-by-file, top-to-bottom by line number
180
+ - Run tests after each batch of fixes — revert and mark as human-required if tests break
181
+ 3. Re-run only the agents that reported actionable issues.
182
+ 4. Repeat up to **5 iterations** total (matching `/code-review` loop behavior).
183
+ 5. **Exit conditions**:
184
+ - Zero actionable issues remain → continue to next plan step
185
+ - Same issues persist after fix attempt → not converging, escalate
186
+ - Iteration limit reached (5) → escalate to human with:
187
+ - The original findings
188
+ - All fix attempts
189
+ - Remaining issues and recommended resolution path
190
+ 6. `warn` after any iteration is acceptable; document in phase output and continue.
191
+
192
+ ## Phase transitions
193
+
194
+ 1. Complete the current phase's work
195
+ 2. Write a structured progress file to `.claude/memory/` (see Context Summarization skill)
196
+ 3. Human reviews and approves before proceeding
197
+ 4. Start new context window for the next phase
198
+ 5. Load only the progress file + agents needed for the new phase