codegraph-brain 0.9.0__tar.gz → 0.14.0__tar.gz

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (320) hide show
  1. codegraph_brain-0.14.0/.gitattributes +6 -0
  2. codegraph_brain-0.14.0/.github/workflows/ci.yml +212 -0
  3. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/.gitignore +1 -0
  4. codegraph_brain-0.14.0/.pre-commit-config.yaml +45 -0
  5. codegraph_brain-0.14.0/.release-please-manifest.json +3 -0
  6. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/CHANGELOG.md +100 -0
  7. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/PKG-INFO +4 -2
  8. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/README.md +2 -0
  9. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/benchmarks/guardian/CURATION.md +153 -4
  10. codegraph_brain-0.14.0/benchmarks/guardian/calibration.jsonl +236 -0
  11. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/benchmarks/guardian/results.jsonl +118 -118
  12. codegraph_brain-0.14.0/benchmarks/martian/README.md +98 -0
  13. codegraph_brain-0.14.0/benchmarks/martian/cal_dot_com.json +267 -0
  14. codegraph_brain-0.14.0/benchmarks/martian/discourse.json +277 -0
  15. codegraph_brain-0.14.0/benchmarks/martian/grafana.json +187 -0
  16. codegraph_brain-0.14.0/benchmarks/martian/keycloak.json +214 -0
  17. codegraph_brain-0.14.0/benchmarks/martian/sentry.json +248 -0
  18. codegraph_brain-0.14.0/benchmarks/martian-judged.jsonl +109 -0
  19. codegraph_brain-0.14.0/benchmarks/martian-p3-judged-run1.jsonl +2 -0
  20. codegraph_brain-0.14.0/benchmarks/martian-p3-judged-run1.jsonl.corrupted-backup +11 -0
  21. codegraph_brain-0.14.0/benchmarks/martian-p3-run1.jsonl +19 -0
  22. codegraph_brain-0.14.0/benchmarks/martian-p3-run2.jsonl +19 -0
  23. codegraph_brain-0.14.0/benchmarks/martian-p3-run3.jsonl +7 -0
  24. codegraph_brain-0.14.0/benchmarks/martian-plan.json +1275 -0
  25. codegraph_brain-0.14.0/benchmarks/martian-repeat-judged.jsonl +6 -0
  26. codegraph_brain-0.14.0/benchmarks/martian-repeat-reviews.jsonl +6 -0
  27. codegraph_brain-0.14.0/benchmarks/martian-reviews.jsonl +64 -0
  28. codegraph_brain-0.14.0/docs/GUARDIAN_LOCAL_BENCH.md +157 -0
  29. codegraph_brain-0.14.0/docs/GUARDIAN_REMOTE_OLLAMA.md +105 -0
  30. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/docs/specs/2026-07-31-finder-bug-class-taxonomy.md +211 -0
  31. codegraph_brain-0.14.0/docs/specs/2026-08-01-aura-autoevolution-poc.md +840 -0
  32. codegraph_brain-0.14.0/docs/specs/2026-08-11-guardian-code-review-bench.md +1428 -0
  33. codegraph_brain-0.14.0/docs/specs/2026-08-14-review-fingerprint-design.md +545 -0
  34. codegraph_brain-0.14.0/docs/specs/plans/2026-08-14-review-fingerprint.md +1389 -0
  35. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/pyproject.toml +1 -1
  36. codegraph_brain-0.14.0/scripts/backfill_calibration_fingerprint.py +246 -0
  37. codegraph_brain-0.14.0/scripts/backfill_review_fingerprint.py +314 -0
  38. codegraph_brain-0.14.0/scripts/check_pytest_raises.py +152 -0
  39. codegraph_brain-0.14.0/scripts/colab_bench.sh +90 -0
  40. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/scripts/guardian_bench.py +68 -8
  41. codegraph_brain-0.14.0/scripts/guardian_calibrate.py +633 -0
  42. codegraph_brain-0.14.0/scripts/guardian_martian.py +1266 -0
  43. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/scripts/guardian_review.py +25 -2
  44. codegraph_brain-0.14.0/scripts/ollama_visitor.sh +107 -0
  45. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/src/cgis/api/mcp_server.py +9 -3
  46. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/src/cgis/cli.py +12 -10
  47. codegraph_brain-0.14.0/src/cgis/extractors/registry.py +93 -0
  48. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/src/cgis/extractors/typescript_extractor.py +5 -1
  49. codegraph_brain-0.14.0/src/cgis/guardian/axes.py +130 -0
  50. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/src/cgis/guardian/bench.py +36 -5
  51. codegraph_brain-0.14.0/src/cgis/guardian/calibrate.py +450 -0
  52. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/src/cgis/guardian/chunked.py +14 -21
  53. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/src/cgis/guardian/collector.py +28 -13
  54. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/src/cgis/guardian/core.py +30 -4
  55. codegraph_brain-0.14.0/src/cgis/guardian/findings.py +157 -0
  56. codegraph_brain-0.14.0/src/cgis/guardian/martian.py +799 -0
  57. codegraph_brain-0.14.0/src/cgis/guardian/metrics.py +221 -0
  58. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/src/cgis/guardian/prompts.py +85 -32
  59. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/src/cgis/guardian/providers/base.py +45 -0
  60. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/src/cgis/guardian/providers/gemini.py +4 -0
  61. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/src/cgis/guardian/providers/mistral.py +49 -2
  62. codegraph_brain-0.14.0/src/cgis/guardian/providers/ollama.py +211 -0
  63. codegraph_brain-0.14.0/src/cgis/guardian/review_fingerprint.py +288 -0
  64. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/src/cgis/guardian/runner.py +183 -8
  65. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/tests/unit/conftest.py +45 -1
  66. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/tests/unit/guardian_stubs.py +44 -0
  67. codegraph_brain-0.14.0/tests/unit/review_path_inventory.txt +43 -0
  68. codegraph_brain-0.14.0/tests/unit/test_backfill_calibration_fingerprint.py +745 -0
  69. codegraph_brain-0.14.0/tests/unit/test_backfill_review_fingerprint.py +454 -0
  70. codegraph_brain-0.14.0/tests/unit/test_check_pytest_raises.py +119 -0
  71. codegraph_brain-0.14.0/tests/unit/test_extractor_registry.py +184 -0
  72. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/tests/unit/test_fingerprint.py +3 -1
  73. codegraph_brain-0.14.0/tests/unit/test_guardian_axes.py +214 -0
  74. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/tests/unit/test_guardian_bench.py +44 -3
  75. codegraph_brain-0.14.0/tests/unit/test_guardian_bench_script.py +195 -0
  76. codegraph_brain-0.14.0/tests/unit/test_guardian_calibrate.py +503 -0
  77. codegraph_brain-0.14.0/tests/unit/test_guardian_calibrate_script.py +696 -0
  78. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/tests/unit/test_guardian_chunked.py +6 -4
  79. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/tests/unit/test_guardian_collector.py +93 -12
  80. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/tests/unit/test_guardian_core.py +41 -8
  81. codegraph_brain-0.14.0/tests/unit/test_guardian_martian.py +463 -0
  82. codegraph_brain-0.14.0/tests/unit/test_guardian_martian_script.py +2191 -0
  83. codegraph_brain-0.14.0/tests/unit/test_guardian_martian_union.py +347 -0
  84. codegraph_brain-0.14.0/tests/unit/test_guardian_metrics.py +435 -0
  85. codegraph_brain-0.14.0/tests/unit/test_guardian_mistral_sampling.py +213 -0
  86. codegraph_brain-0.14.0/tests/unit/test_guardian_ollama_sampling.py +254 -0
  87. codegraph_brain-0.14.0/tests/unit/test_guardian_ollama_truncation.py +204 -0
  88. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/tests/unit/test_guardian_providers.py +4 -0
  89. codegraph_brain-0.14.0/tests/unit/test_guardian_providers_name.py +128 -0
  90. codegraph_brain-0.14.0/tests/unit/test_guardian_review_script.py +56 -0
  91. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/tests/unit/test_guardian_runner.py +248 -0
  92. codegraph_brain-0.14.0/tests/unit/test_guardian_salvage.py +299 -0
  93. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/tests/unit/test_guardian_skeptic.py +3 -0
  94. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/tests/unit/test_mcp_server.py +19 -0
  95. codegraph_brain-0.14.0/tests/unit/test_review_fingerprint_closure.py +259 -0
  96. codegraph_brain-0.14.0/tests/unit/test_review_fingerprint_contract.py +211 -0
  97. codegraph_brain-0.14.0/tests/unit/test_review_fingerprint_digest.py +226 -0
  98. codegraph_brain-0.14.0/tests/unit/test_review_fingerprint_record.py +25 -0
  99. codegraph_brain-0.14.0/tests/unit/test_temperature_source.py +168 -0
  100. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/uv.lock +43 -46
  101. codegraph_brain-0.9.0/.github/workflows/ci.yml +0 -79
  102. codegraph_brain-0.9.0/.pre-commit-config.yaml +0 -27
  103. codegraph_brain-0.9.0/.release-please-manifest.json +0 -3
  104. codegraph_brain-0.9.0/src/cgis/guardian/findings.py +0 -74
  105. codegraph_brain-0.9.0/src/cgis/guardian/metrics.py +0 -108
  106. codegraph_brain-0.9.0/src/cgis/guardian/providers/ollama.py +0 -99
  107. codegraph_brain-0.9.0/tests/unit/test_guardian_metrics.py +0 -210
  108. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/.claude-plugin/marketplace.json +0 -0
  109. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/.github/workflows/autodoc.yml +0 -0
  110. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/.github/workflows/guardian.yml +0 -0
  111. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/.github/workflows/pr-title.yml +0 -0
  112. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/.github/workflows/release-please.yml +0 -0
  113. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/.python-version +0 -0
  114. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/CLAUDE.md +0 -0
  115. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/CONTRIBUTING.md +0 -0
  116. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/LICENSE +0 -0
  117. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/Makefile +0 -0
  118. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/PRIVACY.md +0 -0
  119. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/benchmarks/guardian/pr-122.yaml +0 -0
  120. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/benchmarks/guardian/pr-140.yaml +0 -0
  121. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/benchmarks/guardian/pr-141.yaml +0 -0
  122. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/benchmarks/guardian/pr-142.yaml +0 -0
  123. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/benchmarks/guardian/pr-143.yaml +0 -0
  124. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/benchmarks/guardian/pr-144.yaml +0 -0
  125. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/benchmarks/guardian/pr-278.yaml +0 -0
  126. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/benchmarks/guardian/pr-313.yaml +0 -0
  127. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/data/.gitkeep +0 -0
  128. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/docs/AUDIT.md +0 -0
  129. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/docs/CASE_STUDY.md +0 -0
  130. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/docs/architecture/HOW_IT_WORKS.md +0 -0
  131. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/docs/architecture/ONTOLOGY.md +0 -0
  132. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/docs/architecture/PATTERNS_AND_TRIADS.md +0 -0
  133. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/docs/architecture/SELF_PORTRAIT.md +0 -0
  134. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/docs/architecture/diagrams/pipeline_flow.mermaid +0 -0
  135. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/docs/architecture/health_badge.json +0 -0
  136. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/docs/assets/.gitignore +0 -0
  137. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/docs/assets/cgis-app-avatar.png +0 -0
  138. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/docs/assets/cgis-app-avatar.svg +0 -0
  139. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/docs/assets/generate_avatar.py +0 -0
  140. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/docs/examples/.gitkeep +0 -0
  141. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/docs/how-to/AGENT_ONBOARDING.md +0 -0
  142. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/docs/how-to/CLI_USAGE.md +0 -0
  143. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/docs/how-to/MCP_REFERENCE.md +0 -0
  144. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/docs/lab-notes/2026-06-11-chunked-review-negative-result.md +0 -0
  145. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/docs/ontology/.gitkeep +0 -0
  146. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/docs/ontology/core.yaml +0 -0
  147. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/docs/ontology/domains.yaml +0 -0
  148. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/docs/ontology/patterns.yaml +0 -0
  149. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/docs/ontology/tolerances.lock +0 -0
  150. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/docs/specs/2026-06-09-domain-pattern-fingerprint-design.md +0 -0
  151. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/docs/specs/2026-06-09-pattern-alphabet-motif-basis-design.md +0 -0
  152. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/docs/specs/2026-06-10-guardian-sprint-design.md +0 -0
  153. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/docs/specs/2026-06-11-fastapi-di-edges-design.md +0 -0
  154. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/docs/specs/2026-06-11-guardian-chunked-review-design.md +0 -0
  155. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/docs/specs/2026-06-11-guardian-chunker-design.md +0 -0
  156. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/docs/specs/2026-06-11-mcp-drift-validate-fqn-design.md +0 -0
  157. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/docs/specs/2026-06-11-resolver-split-design.md +0 -0
  158. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/docs/specs/2026-06-11-symbol-import-edges-design.md +0 -0
  159. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/docs/specs/2026-06-12-drift-empty-domains-design.md +0 -0
  160. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/docs/specs/2026-06-12-gate-semantics-design.md +0 -0
  161. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/docs/specs/2026-06-12-init-ontology-design.md +0 -0
  162. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/docs/specs/2026-06-12-release-please-ci-design.md +0 -0
  163. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/docs/specs/2026-06-13-suggest-packages-design.md +0 -0
  164. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/docs/specs/2026-06-13-tangle-anti-pattern-design.md +0 -0
  165. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/docs/specs/2026-07-29-guardian-skeptic-scoring-design.md +0 -0
  166. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/docs/specs/2026-07-30-cgis-fractal-design.md +0 -0
  167. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/docs/specs/2026-07-30-cgis-fractal-plan.md +0 -0
  168. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/docs/specs/2026-07-30-chunk-source-filter-design.md +0 -0
  169. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/docs/specs/2026-07-30-chunk-source-filter-plan.md +0 -0
  170. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/docs/specs/2026-07-30-genai-client-close-design.md +0 -0
  171. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/docs/specs/2026-07-30-genai-client-close-plan.md +0 -0
  172. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/docs/specs/2026-07-30-guardian-precision-bench-design.md +0 -0
  173. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/docs/specs/2026-07-30-guardian-precision-bench-plan.md +0 -0
  174. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/docs/specs/2026-07-30-guardian-timeout-retry-design.md +0 -0
  175. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/docs/specs/2026-07-30-guardian-timeout-retry-plan.md +0 -0
  176. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/docs/specs/BLUEPRINT.md +0 -0
  177. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/docs/specs/PRD.md +0 -0
  178. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/docs/specs/TDD.md +0 -0
  179. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/docs/specs/plans/2026-06-09-fingerprint-drift.md +0 -0
  180. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/docs/specs/plans/2026-06-10-guardian-context-skeptic-inline.md +0 -0
  181. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/docs/specs/plans/2026-06-10-guardian-structured-findings-bench.md +0 -0
  182. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/docs/specs/plans/2026-06-10-motif-basis-part-b.md +0 -0
  183. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/docs/specs/plans/2026-06-10-unified-pattern-alphabet.md +0 -0
  184. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/docs/specs/plans/2026-06-11-fastapi-di-edges.md +0 -0
  185. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/docs/specs/plans/2026-06-11-guardian-chunked-review.md +0 -0
  186. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/docs/specs/plans/2026-06-11-guardian-chunker.md +0 -0
  187. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/docs/specs/plans/2026-06-11-mcp-drift-validate-fqn.md +0 -0
  188. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/docs/specs/plans/2026-06-11-resolver-split.md +0 -0
  189. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/docs/specs/plans/2026-06-12-drift-empty-domains.md +0 -0
  190. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/docs/specs/plans/2026-06-12-gate-semantics.md +0 -0
  191. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/docs/specs/plans/2026-06-12-init-ontology.md +0 -0
  192. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/docs/specs/plans/2026-06-12-release-please-ci.md +0 -0
  193. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/docs/specs/plans/2026-06-12-symbol-import-edges.md +0 -0
  194. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/docs/specs/plans/2026-06-13-suggest-packages.md +0 -0
  195. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/docs/specs/plans/2026-07-29-guardian-skeptic-scoring.md +0 -0
  196. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/main.py +0 -0
  197. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/plugin/.claude-plugin/plugin.json +0 -0
  198. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/plugin/.mcp.json +0 -0
  199. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/plugin/README.md +0 -0
  200. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/plugin/skills/cgis/SKILL.md +0 -0
  201. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/plugin/skills/ingest/SKILL.md +0 -0
  202. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/release-please-config.json +0 -0
  203. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/scripts/gen_ideal_graph.py +0 -0
  204. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/scripts/generate_health.py +0 -0
  205. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/scripts/generate_mcp_ref.py +0 -0
  206. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/scripts/generate_schema_docs.py +0 -0
  207. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/scripts/inject_readme_graph.py +0 -0
  208. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/scripts/probe_closure_gap.py +0 -0
  209. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/scripts/probe_tier_ladder.py +0 -0
  210. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/src/cgis/__init__.py +0 -0
  211. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/src/cgis/__main__.py +0 -0
  212. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/src/cgis/api/.gitkeep +0 -0
  213. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/src/cgis/api/__init__.py +0 -0
  214. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/src/cgis/core/.gitkeep +0 -0
  215. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/src/cgis/core/models.py +0 -0
  216. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/src/cgis/extractors/.gitkeep +0 -0
  217. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/src/cgis/extractors/_python_ast.py +0 -0
  218. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/src/cgis/extractors/_python_classes.py +0 -0
  219. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/src/cgis/extractors/_python_functions.py +0 -0
  220. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/src/cgis/extractors/_python_imports.py +0 -0
  221. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/src/cgis/extractors/_python_types.py +0 -0
  222. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/src/cgis/extractors/base.py +0 -0
  223. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/src/cgis/extractors/python_extractor.py +0 -0
  224. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/src/cgis/guardian/__init__.py +0 -0
  225. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/src/cgis/guardian/chunker.py +0 -0
  226. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/src/cgis/guardian/diff_index.py +0 -0
  227. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/src/cgis/guardian/github_poster.py +0 -0
  228. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/src/cgis/guardian/providers/__init__.py +0 -0
  229. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/src/cgis/guardian/recording.py +0 -0
  230. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/src/cgis/guardian/render.py +0 -0
  231. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/src/cgis/guardian/skeptic.py +0 -0
  232. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/src/cgis/pipeline.py +0 -0
  233. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/src/cgis/py.typed +0 -0
  234. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/src/cgis/query/analysis/__init__.py +0 -0
  235. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/src/cgis/query/analysis/analyzer.py +0 -0
  236. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/src/cgis/query/analysis/anomaly.py +0 -0
  237. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/src/cgis/query/analysis/cohesion.py +0 -0
  238. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/src/cgis/query/analysis/health.py +0 -0
  239. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/src/cgis/query/analysis/suggest_service.py +0 -0
  240. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/src/cgis/query/context/__init__.py +0 -0
  241. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/src/cgis/query/context/audit.py +0 -0
  242. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/src/cgis/query/context/context_service.py +0 -0
  243. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/src/cgis/query/context/prompt.py +0 -0
  244. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/src/cgis/query/context/snippet.py +0 -0
  245. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/src/cgis/query/drift/__init__.py +0 -0
  246. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/src/cgis/query/drift/_scc.py +0 -0
  247. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/src/cgis/query/drift/drift.py +0 -0
  248. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/src/cgis/query/drift/drift_service.py +0 -0
  249. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/src/cgis/query/drift/fingerprint.py +0 -0
  250. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/src/cgis/query/drift/fractal.py +0 -0
  251. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/src/cgis/query/drift/ontology_init.py +0 -0
  252. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/src/cgis/query/drift/quotient.py +0 -0
  253. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/src/cgis/query/drift/triads.py +0 -0
  254. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/src/cgis/query/engine.py +0 -0
  255. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/src/cgis/query/fqn.py +0 -0
  256. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/src/cgis/query/render/__init__.py +0 -0
  257. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/src/cgis/query/render/graph_json.py +0 -0
  258. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/src/cgis/query/render/mermaid.py +0 -0
  259. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/src/cgis/query/render/metrics.py +0 -0
  260. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/src/cgis/resolver/.gitkeep +0 -0
  261. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/src/cgis/resolver/__init__.py +0 -0
  262. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/src/cgis/resolver/engine.py +0 -0
  263. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/src/cgis/resolver/indices.py +0 -0
  264. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/src/cgis/resolver/symbols.py +0 -0
  265. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/src/cgis/resolver/uplift.py +0 -0
  266. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/src/cgis/storage/.gitkeep +0 -0
  267. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/src/cgis/storage/sqlite_store.py +0 -0
  268. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/tests/integration/.gitkeep +0 -0
  269. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/tests/self_parsing/.gitkeep +0 -0
  270. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/tests/self_parsing/__init__.py +0 -0
  271. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/tests/self_parsing/conftest.py +0 -0
  272. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/tests/self_parsing/test_architecture.py +0 -0
  273. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/tests/self_parsing/test_drift.py +0 -0
  274. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/tests/self_parsing/test_fractal.py +0 -0
  275. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/tests/self_parsing/test_init_ontology_roundtrip.py +0 -0
  276. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/tests/self_parsing/test_self_parse.py +0 -0
  277. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/tests/self_parsing/test_self_parse_ts.py +0 -0
  278. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/tests/self_parsing/test_suggest.py +0 -0
  279. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/tests/unit/.gitkeep +0 -0
  280. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/tests/unit/test___main__.py +0 -0
  281. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/tests/unit/test_analyzer.py +0 -0
  282. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/tests/unit/test_audit.py +0 -0
  283. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/tests/unit/test_cli.py +0 -0
  284. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/tests/unit/test_cohesion.py +0 -0
  285. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/tests/unit/test_context_service.py +0 -0
  286. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/tests/unit/test_di_acceptance.py +0 -0
  287. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/tests/unit/test_drift.py +0 -0
  288. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/tests/unit/test_drift_service.py +0 -0
  289. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/tests/unit/test_fqn.py +0 -0
  290. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/tests/unit/test_fractal.py +0 -0
  291. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/tests/unit/test_gen_ideal_graph.py +0 -0
  292. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/tests/unit/test_generate_mcp_ref.py +0 -0
  293. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/tests/unit/test_graph_json.py +0 -0
  294. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/tests/unit/test_guardian_chunker.py +0 -0
  295. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/tests/unit/test_guardian_diff_index.py +0 -0
  296. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/tests/unit/test_guardian_findings.py +0 -0
  297. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/tests/unit/test_guardian_poster.py +0 -0
  298. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/tests/unit/test_guardian_recording.py +0 -0
  299. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/tests/unit/test_guardian_render.py +0 -0
  300. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/tests/unit/test_health_scorer.py +0 -0
  301. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/tests/unit/test_import_acceptance.py +0 -0
  302. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/tests/unit/test_mermaid.py +0 -0
  303. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/tests/unit/test_metrics.py +0 -0
  304. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/tests/unit/test_models.py +0 -0
  305. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/tests/unit/test_ontology_compliance.py +0 -0
  306. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/tests/unit/test_ontology_init.py +0 -0
  307. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/tests/unit/test_patterns_yaml.py +0 -0
  308. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/tests/unit/test_pipeline.py +0 -0
  309. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/tests/unit/test_prompt.py +0 -0
  310. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/tests/unit/test_python_extractor.py +0 -0
  311. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/tests/unit/test_quotient.py +0 -0
  312. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/tests/unit/test_resolver.py +0 -0
  313. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/tests/unit/test_resolver_indices.py +0 -0
  314. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/tests/unit/test_resolver_symbols.py +0 -0
  315. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/tests/unit/test_snippet.py +0 -0
  316. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/tests/unit/test_sqlite_store.py +0 -0
  317. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/tests/unit/test_suggest_service.py +0 -0
  318. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/tests/unit/test_triads.py +0 -0
  319. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/tests/unit/test_typescript_extractor.py +0 -0
  320. {codegraph_brain-0.9.0 → codegraph_brain-0.14.0}/tests/unit/test_uplift.py +0 -0
@@ -0,0 +1,6 @@
1
+ # The review fingerprint (#375) hashes file bytes, so a checkout that stores
2
+ # CRLF would produce a different digest for identical code — and the measured
3
+ # digest would disagree with the reconstructed one on the same commit. Pinned
4
+ # rather than normalised at hash time, because a normaliser can only fail by
5
+ # merging two reviewers that differ.
6
+ *.py text eol=lf
@@ -0,0 +1,212 @@
1
+ name: Continuous Integration
2
+
3
+ on:
4
+ push:
5
+ branches: [ main, dev, "feat/**", "fix/**", "docs/**" ]
6
+ pull_request:
7
+ branches: [ main ]
8
+
9
+ jobs:
10
+ python-ci:
11
+ name: Python Verification
12
+ runs-on: ubuntu-latest
13
+ # Job level, not step level: a step's own `env` block is not reliably visible
14
+ # to that same step's `if`, and a condition that silently never fires is a
15
+ # gate that looks like protection and is not.
16
+ env:
17
+ SONAR_TOKEN: ${{ secrets.SONAR_TOKEN }}
18
+ steps:
19
+ - name: Checkout Code
20
+ uses: actions/checkout@34e114876b0b11c390a56381ad16ebd13914f8d5 # v4
21
+ with:
22
+ # Full history: the scanner computes new-code metrics by diffing against
23
+ # the base branch, and a depth-1 clone makes it log
24
+ # WARN Could not find ref 'main' in refs/heads, refs/remotes…
25
+ # which silently degrades every new_* metric, new_coverage included.
26
+ fetch-depth: 0
27
+
28
+ - name: Setup uv
29
+ uses: astral-sh/setup-uv@caf0cab7a618c569241d31dcd442f54681755d39 # v3
30
+ with:
31
+ enable-cache: true
32
+
33
+ - name: Lockfile in step with pyproject
34
+ # This is the only step that validates the lockfile. `--frozen` below stops
35
+ # uv rewriting uv.lock, but it does not check that the lock is current —
36
+ # that is `--locked`. Keeping the check explicit gives a clearer failure
37
+ # than a sync error, and it still runs first so an unfrozen uv could never
38
+ # repair the lock ahead of it.
39
+ run: uv lock --check
40
+
41
+ - name: Install Dependencies
42
+ run: uv sync --frozen --group dev
43
+
44
+ - name: Ruff Format Check
45
+ run: uv run --frozen ruff format --check .
46
+
47
+ - name: Ruff Lint
48
+ run: uv run --frozen ruff check .
49
+
50
+ - name: Interrogate (≥90% docstring coverage)
51
+ run: uv run --frozen interrogate -v --fail-under 90 src/cgis
52
+
53
+ - name: Mypy (strict)
54
+ run: uv run --frozen mypy src
55
+
56
+ - name: Pytest
57
+ # xml alongside term-missing: the console form is for a human reading the
58
+ # log, the XML is the only form SonarCloud can import.
59
+ #
60
+ # `--cov=scripts` because `sonar.sources` below is `src,scripts`: Sonar
61
+ # counts every executable line in scripts/ as a line to cover whether or
62
+ # not the report mentions it, so omitting the directory does not exclude
63
+ # it from the coverage metric — it just reports all of it as uncovered.
64
+ # That alone put new-code coverage at 37.8% on a PR whose script was
65
+ # covered at 99% locally.
66
+ run: >
67
+ uv run --frozen pytest --cov=cgis --cov=scripts
68
+ --cov-report=term-missing --cov-report=xml
69
+
70
+ - name: Project version for Sonar
71
+ id: version
72
+ # Without this every analysis records the version as the literal string
73
+ # "not provided", which is both a lie in the history and, if the
74
+ # new-code period is "previous version", a baseline that can never move
75
+ # again — no VERSION event fires when the string never changes, so the
76
+ # window grows without bound and the gate quietly becomes a
77
+ # whole-project gate (#309).
78
+ #
79
+ # Validated before use: the value reaches the scanner's args, and a
80
+ # version string is not somewhere to discover that pyproject can carry
81
+ # arbitrary text.
82
+ if: env.SONAR_TOKEN != ''
83
+ run: |
84
+ version="$(uv version --short)"
85
+ case "$version" in
86
+ "" | *[!0-9A-Za-z.+-]* )
87
+ echo "Refusing unexpected project version: '$version'" >&2
88
+ exit 1 ;;
89
+ esac
90
+ echo "value=$version" >> "$GITHUB_OUTPUT"
91
+
92
+ - name: SonarQube Cloud scan
93
+ # Inert until SONAR_TOKEN exists, so this can land before the migration.
94
+ # Coverage cannot be imported by Automatic Analysis at all — see #317 —
95
+ # so the token must be added only AFTER Automatic Analysis is switched
96
+ # off in the SonarCloud UI, or the two modes conflict.
97
+ if: env.SONAR_TOKEN != ''
98
+ uses: SonarSource/sonarqube-scan-action@22918119ff8e1ca75a623e15c8296b6ea4fbe28f # v8.2.1
99
+ with:
100
+ # Passed as args rather than a committed sonar-project.properties:
101
+ # SonarCloud treats that file as a signal to disable Automatic
102
+ # Analysis, which would break the working setup the moment this
103
+ # merges rather than when the migration is actually performed.
104
+ # pythonsecurity:S8707 ("an LLM running this with faulty CLI arguments
105
+ # can escape file system restrictions") fires on any script that takes
106
+ # a path from argparse and reads or writes it. Every script here does,
107
+ # by definition — they are developer and CI tools whose arguments come
108
+ # from this repository's own workflows, not from untrusted input, and
109
+ # confining them to a root would break ad-hoc runs like
110
+ # `--out /tmp/scratch.jsonl`. main already carries four unresolved
111
+ # instances (guardian_bench.py, gen_ideal_graph.py x2, metrics.py) and
112
+ # has never acted on them; the gate only blocked once a PR made such
113
+ # lines *new*. This states that de-facto position explicitly.
114
+ #
115
+ # Deliberately scoped to scripts/. src/ keeps the rule — metrics.py:70
116
+ # is library code and its instance deserves a real answer (#342).
117
+ #
118
+ # pythonsecurity:S2083 ("do not construct the path from user-controlled
119
+ # data") is exempted for scripts/ too, but for a DIFFERENT reason than
120
+ # S8707 above, and the difference matters. S8707's argument is that
121
+ # confining these tools to a root would break legitimate ad-hoc runs.
122
+ # S2083 was raised against backfill_review_fingerprint.py, which does
123
+ # the opposite: `safe_corpus_path` resolves the argument, requires a
124
+ # .jsonl regular file, confines it under the repository root, raises
125
+ # otherwise, and returns the resolved path that `backfill()` then uses
126
+ # for every read and write — the original argument never reaches the
127
+ # filesystem again (#375).
128
+ #
129
+ # The finding survived that. It was reported first at the read, then at
130
+ # the write, unchanged by binding the check to the value used, because
131
+ # SonarCloud's Python taint engine does not model `Path.resolve()` plus
132
+ # `is_relative_to()` as a sanitiser. The remaining report is an engine
133
+ # limitation, not an unvalidated path. Verified rather than assumed:
134
+ # main carries no open S2083, and the one S8707 left in src/ is marked
135
+ # FALSE-POSITIVE by a human, so neither rule has a working precedent
136
+ # here to copy blindly.
137
+ #
138
+ # The alternative considered and rejected was taking a bare filename
139
+ # and building the path from a fixed base, which the engine would
140
+ # accept. It makes traversal inexpressible, but it changes a documented
141
+ # CLI to satisfy an analyser rather than to close a hole that is
142
+ # already closed. If a script ever takes a path it does NOT validate,
143
+ # this exemption hides it — that is the cost, and it is the reason the
144
+ # rule stays on for src/.
145
+ #
146
+ # `sonar.newCode.referenceBranch` overrides the server-side new-code
147
+ # definition for this analysis. The instance default is
148
+ # `previous_version` (confirmed: sonar.leak.period=previous_version,
149
+ # parentOrigin=INSTANCE), and with the project version unset until
150
+ # #352 it had nothing to anchor to and collapsed to *everything* —
151
+ # `new_lines` on main read 39257, i.e. the whole codebase was "new".
152
+ # A gate measuring the whole project is not the gate anyone intended,
153
+ # and it fails the day one legacy file dips.
154
+ #
155
+ # Reference branch says what the gate actually means: do not make this
156
+ # worse than main. The known consequence, accepted rather than
157
+ # discovered: on main's own analysis the reference is main, so its new
158
+ # code is empty and its gate goes quiet. PR analyses are unaffected
159
+ # either way — SonarCloud always scopes a PR to its diff against the
160
+ # target branch (#309).
161
+ #
162
+ # `.github` is in sources because leaving it out did not mean "no
163
+ # findings there", it meant "nobody is looking". Automatic Analysis
164
+ # used to scan it and reported 34 workflow vulnerabilities; the moment
165
+ # this scanner took over with `src,scripts`, SonarCloud closed all 34
166
+ # as FIXED. Nothing had been fixed. Silent closure in the one
167
+ # directory where supply-chain surface actually matters is worse than
168
+ # a red count (#309).
169
+ #
170
+ # The two workflow rules are then exempted — on evidence, not
171
+ # preference, which is the difference between this and the silent
172
+ # closure above:
173
+ #
174
+ # S8541 "omitting --no-build can execute setup scripts": the
175
+ # remediation is impossible here, verified by running it. This
176
+ # project installs itself as an editable distribution, so both
177
+ # `uv sync --frozen --no-build` and `uv run --frozen --no-build`
178
+ # exit with "Distribution `codegraph-brain==0.10.0 @ editable+.`
179
+ # can't be installed because it is marked as `--no-build` but has no
180
+ # binary distribution". Every third-party dependency ships a wheel,
181
+ # so the only source distribution `uv sync` builds is ours — which
182
+ # is not the untrusted code the rule is about. #309 called this
183
+ # "mechanical, pattern already established"; the established pattern
184
+ # (release-please.yml's `uvx --no-build --from twine==7.0.0`) works
185
+ # only because uvx runs a *foreign* tool, and does not transplant.
186
+ #
187
+ # S8544 "dependencies without locked resolved versions": untrue
188
+ # here. Every invocation passes --frozen against a committed
189
+ # uv.lock, and the "Lockfile in step with pyproject" step above
190
+ # fails the build if that lock is stale.
191
+ #
192
+ # Residual risk, stated because the glob cannot be narrower: a future
193
+ # `uvx` invocation would go unchecked by S8541, where it WOULD be both
194
+ # applicable and fixable. Keep passing --no-build to uvx by hand.
195
+ args: >
196
+ -Dsonar.projectKey=zaebee_codegraph-brain
197
+ -Dsonar.organization=zaebee
198
+ -Dsonar.projectVersion=${{ steps.version.outputs.value }}
199
+ -Dsonar.newCode.referenceBranch=main
200
+ -Dsonar.sources=src,scripts,.github
201
+ -Dsonar.tests=tests
202
+ -Dsonar.python.version=3.12
203
+ -Dsonar.python.coverage.reportPaths=coverage.xml
204
+ -Dsonar.issue.ignore.multicriteria=e1,e2,e3,e4
205
+ -Dsonar.issue.ignore.multicriteria.e1.ruleKey=pythonsecurity:S8707
206
+ -Dsonar.issue.ignore.multicriteria.e1.resourceKey=scripts/**/*.py
207
+ -Dsonar.issue.ignore.multicriteria.e4.ruleKey=pythonsecurity:S2083
208
+ -Dsonar.issue.ignore.multicriteria.e4.resourceKey=scripts/**/*.py
209
+ -Dsonar.issue.ignore.multicriteria.e2.ruleKey=githubactions:S8541
210
+ -Dsonar.issue.ignore.multicriteria.e2.resourceKey=.github/workflows/*.yml
211
+ -Dsonar.issue.ignore.multicriteria.e3.ruleKey=githubactions:S8544
212
+ -Dsonar.issue.ignore.multicriteria.e3.resourceKey=.github/workflows/*.yml
@@ -229,3 +229,4 @@ guardian_metrics.jsonl
229
229
  # SQLite WAL/SHM sidecars
230
230
  *.db-wal
231
231
  *.db-shm
232
+ .martian-workspace/
@@ -0,0 +1,45 @@
1
+ repos:
2
+ - repo: local
3
+ hooks:
4
+ # SonarCloud's python:S5778, enforced before the push rather than after it.
5
+ # Sonar only decorates a pull request, so the loop was write -> push ->
6
+ # get told -> fix -> forget, five times over. `pass_filenames: false`
7
+ # because the check is a repo-wide ratchet (a stale baseline entry is a
8
+ # failure too), which a staged-files-only view cannot see.
9
+ #
10
+ # This hook is the fast half, not the gate: CI does not run pre-commit,
11
+ # so `tests/unit/test_check_pytest_raises.py` is what actually blocks a
12
+ # merge. Both call the same `check()`.
13
+ - id: pytest-raises-single-call
14
+ name: one throwing call per pytest.raises block (S5778)
15
+ entry: python scripts/check_pytest_raises.py
16
+ language: system
17
+ pass_filenames: false
18
+ files: ^tests/.*\.py$
19
+
20
+ - repo: https://github.com/astral-sh/ruff-pre-commit
21
+ rev: v0.15.16
22
+ hooks:
23
+ - id: ruff
24
+ args: [--fix]
25
+ - id: ruff-format
26
+
27
+ - repo: https://github.com/pre-commit/mirrors-mypy
28
+ rev: v2.1.0
29
+ hooks:
30
+ - id: mypy
31
+ args: [--config-file=pyproject.toml]
32
+ files: ^src/
33
+ additional_dependencies:
34
+ - httpx>=0.27 # imported directly by guardian/providers/base.py (#275)
35
+ - pydantic>=2.13.4
36
+ - structlog>=25.5.0
37
+ - types-pyyaml>=6.0.12.20260518
38
+ - rich>=15.0.0
39
+ - typer>=0.26.7
40
+ - tree-sitter>=0.25.2
41
+ - tree-sitter-python>=0.25.0
42
+ - tree-sitter-typescript>=0.23.2
43
+ - pytest>=9.0.3
44
+ - "mcp[cli]>=2" # must match the pyproject bound
45
+ - duckdb>=1.1.0
@@ -0,0 +1,3 @@
1
+ {
2
+ ".": "0.14.0"
3
+ }
@@ -1,5 +1,105 @@
1
1
  # Changelog
2
2
 
3
+ ## [0.14.0](https://github.com/zaebee/codegraph-brain/compare/codegraph-brain-v0.13.0...codegraph-brain-v0.14.0) (2026-08-16)
4
+
5
+
6
+ ### Features
7
+
8
+ * **guardian:** give measured recall a reviewer it can name ([#390](https://github.com/zaebee/codegraph-brain/issues/390)) ([#391](https://github.com/zaebee/codegraph-brain/issues/391)) ([acc1887](https://github.com/zaebee/codegraph-brain/commit/acc18874b7ce67fa38b7dab41552838f9578516d))
9
+ * **guardian:** say whether a temperature was chosen or inherited ([#393](https://github.com/zaebee/codegraph-brain/issues/393)) ([#395](https://github.com/zaebee/codegraph-brain/issues/395)) ([1b2c3c3](https://github.com/zaebee/codegraph-brain/commit/1b2c3c3f1619f00d20c36fb315eaca3dce5661d9))
10
+
11
+
12
+ ### Bug Fixes
13
+
14
+ * **guardian:** close the deferred cleanups that could only narrow silently ([#385](https://github.com/zaebee/codegraph-brain/issues/385)) ([#386](https://github.com/zaebee/codegraph-brain/issues/386)) ([e7020ff](https://github.com/zaebee/codegraph-brain/commit/e7020ff237964a66728f0b437f13b70dc8bfe2f9))
15
+ * **guardian:** refuse a model name carrying whitespace ([#382](https://github.com/zaebee/codegraph-brain/issues/382)) ([#389](https://github.com/zaebee/codegraph-brain/issues/389)) ([45fb110](https://github.com/zaebee/codegraph-brain/commit/45fb110d3e440e858f57c28cd3e343e61218697c))
16
+
17
+
18
+ ### Documentation
19
+
20
+ * **bench:** a failed parse is not a draw ([#394](https://github.com/zaebee/codegraph-brain/issues/394)) ([a70b09b](https://github.com/zaebee/codegraph-brain/commit/a70b09bf1faf67ddefff2d7fdcf555aba51ced9f))
21
+ * **bench:** repeated rows are samples, not corrections ([#392](https://github.com/zaebee/codegraph-brain/issues/392)) ([e1fcc3f](https://github.com/zaebee/codegraph-brain/commit/e1fcc3f4728e61363d06b022b14bc66f45f3a031))
22
+
23
+ ## [0.13.0](https://github.com/zaebee/codegraph-brain/compare/codegraph-brain-v0.12.0...codegraph-brain-v0.13.0) (2026-08-15)
24
+
25
+
26
+ ### Features
27
+
28
+ * **guardian:** an identity for a reviewer, that a commit does not fragment ([#375](https://github.com/zaebee/codegraph-brain/issues/375)) ([#383](https://github.com/zaebee/codegraph-brain/issues/383)) ([e339a38](https://github.com/zaebee/codegraph-brain/commit/e339a38743ddfdaac1e7c8f0724281844d5a8b38))
29
+
30
+ ## [0.12.0](https://github.com/zaebee/codegraph-brain/compare/codegraph-brain-v0.11.0...codegraph-brain-v0.12.0) (2026-08-13)
31
+
32
+
33
+ ### Features
34
+
35
+ * **guardian:** an ablation arm, because G5 cannot separate graph from language ([#342](https://github.com/zaebee/codegraph-brain/issues/342)) ([#366](https://github.com/zaebee/codegraph-brain/issues/366)) ([e2374a7](https://github.com/zaebee/codegraph-brain/commit/e2374a72016413414241e71d3cc37a7ce251c590))
36
+ * **guardian:** bound the local finder's output, because it does not stop ([#246](https://github.com/zaebee/codegraph-brain/issues/246)) ([#381](https://github.com/zaebee/codegraph-brain/issues/381)) ([1e0b5e1](https://github.com/zaebee/codegraph-brain/commit/1e0b5e13332c697102a22838f90fed265c0fec67))
37
+ * **guardian:** expose judge concurrency, because Mistral needs it at 1 ([#342](https://github.com/zaebee/codegraph-brain/issues/342)) ([#364](https://github.com/zaebee/codegraph-brain/issues/364)) ([47392e8](https://github.com/zaebee/codegraph-brain/commit/47392e829e8653e6d145533a637342e8a46b4982))
38
+ * **guardian:** keep the valid prefix of a truncated finder response ([#248](https://github.com/zaebee/codegraph-brain/issues/248)) ([#377](https://github.com/zaebee/codegraph-brain/issues/377)) ([d34f01f](https://github.com/zaebee/codegraph-brain/commit/d34f01fa11208c0f8202579edc7e07dccb163ad4))
39
+ * **guardian:** record what the judge spent, per row ([#342](https://github.com/zaebee/codegraph-brain/issues/342)) ([#378](https://github.com/zaebee/codegraph-brain/issues/378)) ([c4eea31](https://github.com/zaebee/codegraph-brain/commit/c4eea31b12492d90dafb92b3afacfbe9a942bab5))
40
+ * **guardian:** refuse a review of a truncated prompt, and supply the missing visitor ([#248](https://github.com/zaebee/codegraph-brain/issues/248)) ([#371](https://github.com/zaebee/codegraph-brain/issues/371)) ([5f3e9b7](https://github.com/zaebee/codegraph-brain/commit/5f3e9b73bcc36aabe64d634e4aa95ea0f5bd7fb7))
41
+ * **guardian:** review --slice, so the registered population is a command ([#342](https://github.com/zaebee/codegraph-brain/issues/342)) ([#373](https://github.com/zaebee/codegraph-brain/issues/373)) ([4d1fe6a](https://github.com/zaebee/codegraph-brain/commit/4d1fe6a8073641cdd46a9d9af24a3c31abd70764))
42
+ * **guardian:** sampling reaches Ollama, instead of the chat template deciding ([#246](https://github.com/zaebee/codegraph-brain/issues/246)) ([#380](https://github.com/zaebee/codegraph-brain/issues/380)) ([14cb966](https://github.com/zaebee/codegraph-brain/commit/14cb9664f200aa9507f1f364f72dfb0de0981259))
43
+ * **guardian:** send the registered temperature, and survive a 429 ([#342](https://github.com/zaebee/codegraph-brain/issues/342)) ([#372](https://github.com/zaebee/codegraph-brain/issues/372)) ([3d7efd8](https://github.com/zaebee/codegraph-brain/commit/3d7efd82efe750283a43eb0d5cb7941f7f67ddd1))
44
+ * **guardian:** the union arm scores offline, and a unit bug made G8 unfailable ([#342](https://github.com/zaebee/codegraph-brain/issues/342)) ([#370](https://github.com/zaebee/codegraph-brain/issues/370)) ([dd290fe](https://github.com/zaebee/codegraph-brain/commit/dd290fef986ad0dcf0b62da9fa2824e0ace1b90d))
45
+
46
+
47
+ ### Bug Fixes
48
+
49
+ * **guardian:** a dead judge is not a reviewer that found nothing ([#342](https://github.com/zaebee/codegraph-brain/issues/342)) ([#376](https://github.com/zaebee/codegraph-brain/issues/376)) ([2772c30](https://github.com/zaebee/codegraph-brain/commit/2772c305e9d7ca70c31bdfc2b12297392768330c))
50
+ * **guardian:** a truncated review is not a review that found nothing ([#342](https://github.com/zaebee/codegraph-brain/issues/342)) ([#374](https://github.com/zaebee/codegraph-brain/issues/374)) ([9f86bb6](https://github.com/zaebee/codegraph-brain/commit/9f86bb63e297e2b8bdda6bf6214e5d1e1bbd9339))
51
+ * **guardian:** the ablation must skip PRs with no graph to remove ([#342](https://github.com/zaebee/codegraph-brain/issues/342)) ([#367](https://github.com/zaebee/codegraph-brain/issues/367)) ([f9c36f5](https://github.com/zaebee/codegraph-brain/commit/f9c36f5cae503d83002ad03957eb7bf0cf31d263))
52
+
53
+
54
+ ### Documentation
55
+
56
+ * **guardian:** bench a local model in a notebook, and retire a wrong lever ([#246](https://github.com/zaebee/codegraph-brain/issues/246)) ([#379](https://github.com/zaebee/codegraph-brain/issues/379)) ([419fcc1](https://github.com/zaebee/codegraph-brain/commit/419fcc1eb28fddd9a2c0cef814480d861e6d8c88))
57
+ * **spec:** Phase 2 results — G5 fails at +9.5 pp against a 10 pp gate ([#342](https://github.com/zaebee/codegraph-brain/issues/342)) ([#362](https://github.com/zaebee/codegraph-brain/issues/362)) ([9481a54](https://github.com/zaebee/codegraph-brain/commit/9481a54c02b60786e14e6317d3a71c8a4cee12a0))
58
+ * **spec:** Phase 3 registers the union arm, and the pilot says it is marginal ([#342](https://github.com/zaebee/codegraph-brain/issues/342)) ([#369](https://github.com/zaebee/codegraph-brain/issues/369)) ([bc9caf4](https://github.com/zaebee/codegraph-brain/commit/bc9caf43c16a2ce27a0075b14eb0766569ae8c2d))
59
+ * **spec:** R5 — the noise floor equals the effect, so Phase 2 cannot answer its question ([#342](https://github.com/zaebee/codegraph-brain/issues/342)) ([#368](https://github.com/zaebee/codegraph-brain/issues/368)) ([8540bf1](https://github.com/zaebee/codegraph-brain/commit/8540bf1178a8535849bdd2fadfa5473f42040eb8))
60
+ * **spec:** second judge — G5 fails under both, and the row becomes publishable ([#342](https://github.com/zaebee/codegraph-brain/issues/342)) ([#365](https://github.com/zaebee/codegraph-brain/issues/365)) ([62c49e4](https://github.com/zaebee/codegraph-brain/commit/62c49e4b36c0950df0c10cce26c09418674130eb))
61
+
62
+ ## [0.11.0](https://github.com/zaebee/codegraph-brain/compare/codegraph-brain-v0.10.0...codegraph-brain-v0.11.0) (2026-08-12)
63
+
64
+
65
+ ### Features
66
+
67
+ * **collector:** collect TypeScript context via a language registry ([#344](https://github.com/zaebee/codegraph-brain/issues/344)) ([#349](https://github.com/zaebee/codegraph-brain/issues/349)) ([f40f22f](https://github.com/zaebee/codegraph-brain/commit/f40f22fcfae57602fff24ae183bfc8977f246473)), closes [#342](https://github.com/zaebee/codegraph-brain/issues/342)
68
+ * **guardian:** Martian corpus layer, and the profiles are not what the spec said ([#342](https://github.com/zaebee/codegraph-brain/issues/342)) ([#355](https://github.com/zaebee/codegraph-brain/issues/355)) ([16f9640](https://github.com/zaebee/codegraph-brain/commit/16f96408e46bfe9d5544fa10dd4ea9d68b2e8f34))
69
+ * **guardian:** Phase 1 calibration harness — score recorded reviews with Martian's judge ([#342](https://github.com/zaebee/codegraph-brain/issues/342)) ([#346](https://github.com/zaebee/codegraph-brain/issues/346)) ([6d7771a](https://github.com/zaebee/codegraph-brain/commit/6d7771aad9b7bb1bfc7b26a0ed2c9d56b2715466))
70
+ * **guardian:** Phase 2 judge pass, and the first scored PR ([#342](https://github.com/zaebee/codegraph-brain/issues/342)) ([#359](https://github.com/zaebee/codegraph-brain/issues/359)) ([eaf6626](https://github.com/zaebee/codegraph-brain/commit/eaf6626ea473e0c0bd8f2468f4fc989d29a3decf))
71
+ * **guardian:** Phase 2 planner — resolve slices before spending anything ([#342](https://github.com/zaebee/codegraph-brain/issues/342)) ([#356](https://github.com/zaebee/codegraph-brain/issues/356)) ([75a278d](https://github.com/zaebee/codegraph-brain/commit/75a278d5caad07e3865c1fa95ed79c5f63e64640))
72
+ * **guardian:** Phase 2 report — G4/G5/G6, and G5 refuses a vacuous comparison ([#342](https://github.com/zaebee/codegraph-brain/issues/342)) ([#360](https://github.com/zaebee/codegraph-brain/issues/360)) ([14d0fc4](https://github.com/zaebee/codegraph-brain/commit/14d0fc4fb989c1ae0224e3367de30f1b202a701d))
73
+ * **guardian:** Phase 2 review step, and the first paid run ([#342](https://github.com/zaebee/codegraph-brain/issues/342)) ([#358](https://github.com/zaebee/codegraph-brain/issues/358)) ([933c229](https://github.com/zaebee/codegraph-brain/commit/933c2298b1051962c6860a99ea584b8ca6f204e5))
74
+ * **guardian:** Phase 2 workspace — prepare finds the bug that would have sunk G5 ([#342](https://github.com/zaebee/codegraph-brain/issues/342)) ([#357](https://github.com/zaebee/codegraph-brain/issues/357)) ([1ecd962](https://github.com/zaebee/codegraph-brain/commit/1ecd9629f46cab10b907dae285d0f58b0eef5e21))
75
+
76
+
77
+ ### Bug Fixes
78
+
79
+ * **guardian:** an ambiguous hit is a false positive ([#345](https://github.com/zaebee/codegraph-brain/issues/345)) ([#348](https://github.com/zaebee/codegraph-brain/issues/348)) ([fa37624](https://github.com/zaebee/codegraph-brain/commit/fa37624721a3a58a8b33cb01af71b7a7cb0bd34c)), closes [#342](https://github.com/zaebee/codegraph-brain/issues/342)
80
+ * **guardian:** give record_review a contract for the path it writes to ([#347](https://github.com/zaebee/codegraph-brain/issues/347)) ([#350](https://github.com/zaebee/codegraph-brain/issues/350)) ([af968b9](https://github.com/zaebee/codegraph-brain/commit/af968b90d931c8152b5c9372d343cec44c8281fe))
81
+
82
+
83
+ ### Documentation
84
+
85
+ * **spec:** auto-evolution PoC — unit of selection, fitness, first mutation gate ([#335](https://github.com/zaebee/codegraph-brain/issues/335)) ([#336](https://github.com/zaebee/codegraph-brain/issues/336)) ([b7359a7](https://github.com/zaebee/codegraph-brain/commit/b7359a78b3734ff5f23131b28b6923de6bc37707))
86
+ * **spec:** first live data falsifies two parts of the PoC gate ([#335](https://github.com/zaebee/codegraph-brain/issues/335)) ([#338](https://github.com/zaebee/codegraph-brain/issues/338)) ([a417e89](https://github.com/zaebee/codegraph-brain/commit/a417e8977505abd49ddf7c2642a7eeb5614d578a))
87
+ * **spec:** Guardian vs Martian Code Review Bench — Phase 1 calibration results ([#342](https://github.com/zaebee/codegraph-brain/issues/342)) ([7ca8434](https://github.com/zaebee/codegraph-brain/commit/7ca8434ff1363ef6c8c8d0cf97541e7e2039e098))
88
+ * **spec:** Phase 2 corpus reconnaissance and the amended G5 ([#342](https://github.com/zaebee/codegraph-brain/issues/342)) ([#354](https://github.com/zaebee/codegraph-brain/issues/354)) ([e63ec24](https://github.com/zaebee/codegraph-brain/commit/e63ec244f61988048a915e9c7564e897ab68d6cf))
89
+
90
+ ## [0.10.0](https://github.com/zaebee/codegraph-brain/compare/codegraph-brain-v0.9.0...codegraph-brain-v0.10.0) (2026-08-01)
91
+
92
+
93
+ ### Features
94
+
95
+ * **guardian:** kinship-grouped axis batches behind GUARDIAN_FEATURES=axes_paired ([#334](https://github.com/zaebee/codegraph-brain/issues/334)) ([7951dcb](https://github.com/zaebee/codegraph-brain/commit/7951dcb60a2fff2f232b144af76043a9d06d3eee))
96
+ * **guardian:** per-axis review fan-out behind GUARDIAN_FEATURES=axes ([#333](https://github.com/zaebee/codegraph-brain/issues/333)) ([ada3205](https://github.com/zaebee/codegraph-brain/commit/ada3205e54433356f3a1e9ef4944ed5475ed2d7c)), closes [#331](https://github.com/zaebee/codegraph-brain/issues/331)
97
+
98
+
99
+ ### Documentation
100
+
101
+ * **spec:** record the Arm A result — fails its gate, and shows the mechanism ([#330](https://github.com/zaebee/codegraph-brain/issues/330)) ([af8d543](https://github.com/zaebee/codegraph-brain/commit/af8d543fb656af233a961ef081bc8fed5f4331a0))
102
+
3
103
  ## [0.9.0](https://github.com/zaebee/codegraph-brain/compare/codegraph-brain-v0.8.0...codegraph-brain-v0.9.0) (2026-08-01)
4
104
 
5
105
 
@@ -1,6 +1,6 @@
1
- Metadata-Version: 2.4
1
+ Metadata-Version: 2.5
2
2
  Name: codegraph-brain
3
- Version: 0.9.0
3
+ Version: 0.14.0
4
4
  Summary: Semantic code graph for AI agents — deterministic FQN resolution, impact analysis and architectural drift gates, exposed over MCP.
5
5
  Project-URL: Homepage, https://github.com/zaebee/codegraph-brain
6
6
  Project-URL: Repository, https://github.com/zaebee/codegraph-brain
@@ -169,6 +169,8 @@ GUARDIAN_PROVIDER=ollama GUARDIAN_MODEL=qwen2.5-coder:14b \
169
169
  uv run python scripts/guardian_review.py --pr 123 --db graph.db --inline
170
170
  ```
171
171
 
172
+ No GPU on hand? **[Benchmark it on a notebook GPU →](docs/GUARDIAN_LOCAL_BENCH.md)** — free end to end, since the fixtures score without an LLM judge. Or **[point Guardian at a remote Ollama →](docs/GUARDIAN_REMOTE_OLLAMA.md)** — over an frp stcp tunnel, no public port, and a guard that refuses a review of a silently truncated prompt.
173
+
172
174
  ---
173
175
 
174
176
  ## 📈 Proof at Real Scale
@@ -119,6 +119,8 @@ GUARDIAN_PROVIDER=ollama GUARDIAN_MODEL=qwen2.5-coder:14b \
119
119
  uv run python scripts/guardian_review.py --pr 123 --db graph.db --inline
120
120
  ```
121
121
 
122
+ No GPU on hand? **[Benchmark it on a notebook GPU →](docs/GUARDIAN_LOCAL_BENCH.md)** — free end to end, since the fixtures score without an LLM judge. Or **[point Guardian at a remote Ollama →](docs/GUARDIAN_REMOTE_OLLAMA.md)** — over an frp stcp tunnel, no public port, and a guard that refuses a review of a silently truncated prompt.
123
+
122
124
  ---
123
125
 
124
126
  ## 📈 Proof at Real Scale
@@ -7,17 +7,166 @@ first review actually ran on; the final `refs/pull/N/head` contains the
7
7
  FIXES, so replaying it would make every fixed finding unfindable).
8
8
  Spec §3.1 erratum: `head` = review head, not final pull head.
9
9
 
10
+ ## What `ambiguous` does to the score (decided 2026-08-11, #345)
11
+
12
+ **An `ambiguous` hit counts as a false positive.** It is recorded apart from
13
+ `noise` — `MatchResult.ambiguous_hits`, `BenchScore.ambiguous_hits` — but that
14
+ separation is curation diagnostics, not a scoring adjustment. Reported
15
+ precision is `TP / (TP + noise + ambiguous_hits)`.
16
+
17
+ It used to be exempt from the precision denominator. Two findings retired that.
18
+
19
+ **It stopped the number describing the review.** The exemption was per *file*:
20
+ one entry on `drift.py` exempted every prediction in `drift.py`. On pr-142
21
+ every candidate landed on a file carrying one entry, the denominator emptied,
22
+ and `score()` returned the vacuous `1.0` reserved for "nothing wrong was said".
23
+ Two independent LLM judges scored those same reviews at 0.14 and 0.19. Ten of
24
+ 118 recorded reviews reported that vacuous 1.0, across three of the six PRs
25
+ with runs — normal, not exceptional. pr-144 shows the reach: its entries cover
26
+ `drift.py` and `triads.py`, the files carrying 3 of its 5 ground-truth
27
+ findings.
28
+
29
+ **And the exemption was never argued for.** The rule below routes style nits
30
+ here to keep them from depressing **recall** — but omitting them from
31
+ `findings` already does that, since recall divides by `len(findings)`. Removing
32
+ them from the precision denominator as well came along with the mechanism, and
33
+ it points the wrong way: guardian's precision rules forbid style nits, so
34
+ emitting one *is* a precision failure, and the exemption hid exactly the
35
+ failure those rules exist to catch.
36
+
37
+ Measured over the same 118 reviews against two judges (#342 Phase 1), the new
38
+ definition is also the one that tracks an independent scorer:
39
+
40
+ | | old (exempt) | new (strict) |
41
+ |---|---|---|
42
+ | Spearman ρ vs judge, all reviews | +0.72 / +0.71 | **+0.92 / +0.92** |
43
+ | ρ vs judge, non-empty reviews | +0.42 / +0.35 | **+0.76 / +0.72** |
44
+ | mean \|difference from judge\| | 0.25 / 0.26 | **0.10 / 0.10** |
45
+
46
+ What this costs, stated plainly: a genuinely debatable suggestion — pr-144's
47
+ three declined clip proposals, which gemini raised and then agreed to drop —
48
+ now counts against precision. That is the price of a number an external scorer
49
+ can be compared to, and `ambiguous_hits` keeps the fact visible.
50
+
51
+ `benchmarks/guardian/results.jsonl` is **not** rescored. Its `precision` was
52
+ correct under the policy in force when each row was written, and every row
53
+ carries `matched`, `noise` and `ambiguous_hits`, so either definition can be
54
+ re-derived from it. (Contrast `calibration.jsonl`, which *was* rewritten — that
55
+ was a scorer bug, wrong under its own stated algorithm, not a policy change.)
56
+
57
+ ## Repeated rows are samples, not corrections (2026-08-16, #390)
58
+
59
+ **Do not deduplicate any corpus in this repository by "one row per subject,
60
+ latest wins."** Every repeat here was paid for on purpose, and collapsing to the
61
+ newest draw silently discards the experiment it belongs to.
62
+
63
+ Written down because a downstream consumer adopted exactly that rule — reasoning
64
+ that a re-run is a correction, so counting both would let a reviewer improve its
65
+ record by re-running — and nothing in this repository contradicted it. The
66
+ reasoning is sound for a corpus of corrections. It is wrong for these.
67
+
68
+ ### `benchmarks/martian-*.jsonl` — 115 review rows, 83 distinct subjects
69
+
70
+ Keyed by subject and genome (`url`, `head_sha`, `review_fingerprint`,
71
+ `finder_model`, `had_graph`), 32 of 115 rows are repeats. Every one of them
72
+ comes from a registered sampling arm:
73
+
74
+ | repeat group | count | source |
75
+ |---|---|---|
76
+ | `p3-run1` + `p3-run2` | 12 | Phase 3 union arm |
77
+ | `p3-run1` + `p3-run2` + `p3-run3` | 7 | Phase 3 union arm |
78
+ | `reviews` + `repeat-reviews` | 6 | R5 repeat probe |
79
+
80
+ Phase 3 registered three runs at `temperature = 0.7` **because the runs must
81
+ differ** — see the spec, "Configuration under test": at temperature 0 "the three
82
+ runs collapse toward one and the union arm becomes identically equal to a single
83
+ run." Gate G8 is defined as `F₂(union) > F₂(mean of the 3 runs)`; a mean over
84
+ three draws is not computable from one of them.
85
+
86
+ The variance is not theoretical. **21 of the 25 repeat groups disagree on how
87
+ many findings the review produced**, with spreads including 8→14, 15→24, 26→38
88
+ and one 0→25 — the same reviewer, the same commit, the same configuration. Keep
89
+ one draw and the number you publish is a coin flip over that range.
90
+
91
+ ### `benchmarks/guardian/results.jsonl` — 118 scored rows, 53 distinct
92
+
93
+ Repeats here carry an explicit `run` index (0/1/2) for the same reason. A row is
94
+ identified by `pr` + `model` + `guardian_sha` + `run`, and dropping `run` averages
95
+ away the per-PR sampling noise that R5 was written to measure.
96
+
97
+ ### `benchmarks/guardian/calibration.jsonl` — 236 rows, 118 subjects
98
+
99
+ Exactly two rows per recorded review, one per judge (`gemini-2.5-flash` and
100
+ `mistral-medium-latest`). The pair *is* the measurement — inter-judge agreement
101
+ is the G3 statistic — so `row_key` alone is not a unique key here; `row_key` +
102
+ `judge_model` is.
103
+
104
+ ### If you do need one row per reviewer
105
+
106
+ Aggregate over the draws (mean, or union where the arm defines one); do not
107
+ select among them. Selecting by recency is the one choice guaranteed to be
108
+ uncorrelated with quality, and it deletes 28% of this corpus.
109
+
110
+ **Aggregate over the subject's *draws*, which is not the same set as its rows.**
111
+ Say what a draw is rather than which rows to exclude: an exclusion list is
112
+ open-ended and the next entry arrives after it has already been believed, while
113
+ a positive definition is closed.
114
+
115
+ **A draw is a review that completed and whose output was parsed** — a row with
116
+ `parse_failed: false`. That is the whole test; `error` is `None` on all 115 rows,
117
+ because a run that failed outright never reached a record.
118
+
119
+ **A `parse_failed` row is not a draw.** Its structured output could not be read,
120
+ so it carries zero findings, and averaging it in charges the model for a harness
121
+ failure. This is exactly the distinction #374 established ("a truncated review is
122
+ not a review that found nothing"); before it, `parse_failed` was recorded and
123
+ read by nothing.
124
+
125
+ **A row with zero findings and `parse_failed: false` *is* a draw.** Three exist
126
+ (PRs 6, 77754 and 107534, all gemini), and each spent real tokens — 106, 117 and
127
+ 124 completion tokens. The reviewer ran and said nothing, which is evidence about
128
+ the reviewer. Same output shape as the parse failure, opposite meaning; keeping
129
+ them and dropping the parse failure is the one combination that is right about
130
+ both. The `arm` field exists for the mirror of this hazard: a failure and a
131
+ controlled removal must not look alike in the record.
132
+
133
+ There is exactly **one** parse-failed row in the review corpora:
134
+ `martian-p3-run1.jsonl` for PR 11059 — the same row #374 was written about. One
135
+ row out of 115 sounds ignorable and is not, because of how the averaging works:
136
+
137
+ - PR 11059 has three rows. Dropping the parse-failed one leaves draws of `9/13`
138
+ and `8/8`, so the subject contributes mean `c = 8.5` over mean `n = 10.5` —
139
+ a rate of **81.0%**, which is *below* mistral/graph's pooled 86.6%.
140
+ - Averaging the parse failure in as a third draw keeps that 81.0% but shrinks the
141
+ subject to `c = 5.667` over `n = 7.0`. Its **weight** falls, not its rate.
142
+ - A below-average subject carrying less weight pulls the pooled figure **up**:
143
+ 86.6% → **86.9%**.
144
+
145
+ So including the harness failure makes the reviewer look better, which is the
146
+ direction that makes this easy to leave in. That is how the gap was found — a
147
+ downstream consumer computed 86.6%, this repository computed 86.9%, and the
148
+ disagreement had exactly one cause.
149
+
150
+ The same applies to the `error` rows in `results.jsonl`, for the same reason and
151
+ by the same rule: they are already excluded there because `is_scored` requires
152
+ `matched` and `precision`, which a failed run never gets.
153
+
10
154
  ## Curation policy
11
155
 
12
156
  - Style/idiom-only findings → `ambiguous` (guardian's PRECISION RULES forbid
13
- style nits; penalizing recall for them would measure the wrong thing).
157
+ style nits; penalizing recall for them would measure the wrong thing
158
+ keeping them out of GT `findings` is what achieves that, see above).
14
159
  - Review-dialogue resolutions (open questions answered during review) →
15
160
  `ambiguous` (not defects present in the diff).
16
161
  - Declined-with-reason suggestions → `ambiguous` (per spec §3.1).
17
162
  - Refuted claims → omitted from GT entirely, NOT `ambiguous`. A finding
18
- disproved by execution (the code cannot behave as claimed) is noise, and
19
- `ambiguous` exempts hits per FILE filing errors there would make the suite
20
- structurally unable to observe a precision failure (#279).
163
+ disproved by execution (the code cannot behave as claimed) is an error, not a
164
+ judgement call. The original reason was that `ambiguous` exempted hits per
165
+ file, so filing errors there made the suite structurally unable to observe a
166
+ precision failure (#279); since #345 removed the exemption the rule no longer
167
+ protects the score, but it still keeps `ambiguous_hits` meaning what it says —
168
+ "landed on ground we already marked debatable" — rather than becoming a
169
+ second name for noise.
21
170
  - `lines` required wherever the source provides line provenance (quality
22
171
  review I3: rangeless entries are gameable by file-level predictions).
23
172
  - Sonar quality-gate findings (cognitive complexity, float equality) stay in