agentgauge-harness 0.4.0__tar.gz

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (382) hide show
  1. agentgauge_harness-0.4.0/.claude/agents/executor.md +23 -0
  2. agentgauge_harness-0.4.0/.claude/agents/verifier.md +27 -0
  3. agentgauge_harness-0.4.0/.claude/operator-prompt.md +126 -0
  4. agentgauge_harness-0.4.0/.github/pull_request_template.md +5 -0
  5. agentgauge_harness-0.4.0/.github/workflows/ci.yml +88 -0
  6. agentgauge_harness-0.4.0/.github/workflows/release.yml +43 -0
  7. agentgauge_harness-0.4.0/.gitignore +89 -0
  8. agentgauge_harness-0.4.0/.pre-commit-config.yaml +27 -0
  9. agentgauge_harness-0.4.0/AUTONOMY.md +41 -0
  10. agentgauge_harness-0.4.0/CLAUDE.md +244 -0
  11. agentgauge_harness-0.4.0/LICENSE +190 -0
  12. agentgauge_harness-0.4.0/PKG-INFO +382 -0
  13. agentgauge_harness-0.4.0/PLAN.md +106 -0
  14. agentgauge_harness-0.4.0/PREDICTIVE_VALIDITY_HANDOFF.md +137 -0
  15. agentgauge_harness-0.4.0/README.md +345 -0
  16. agentgauge_harness-0.4.0/ROADMAP.md +56 -0
  17. agentgauge_harness-0.4.0/STATUS.md +1210 -0
  18. agentgauge_harness-0.4.0/TASKS.md +531 -0
  19. agentgauge_harness-0.4.0/agentgauge/__init__.py +3 -0
  20. agentgauge_harness-0.4.0/agentgauge/__main__.py +5 -0
  21. agentgauge_harness-0.4.0/agentgauge/_json.py +43 -0
  22. agentgauge_harness-0.4.0/agentgauge/ab_harness.py +197 -0
  23. agentgauge_harness-0.4.0/agentgauge/audit.py +334 -0
  24. agentgauge_harness-0.4.0/agentgauge/cli.py +1122 -0
  25. agentgauge_harness-0.4.0/agentgauge/client.py +112 -0
  26. agentgauge_harness-0.4.0/agentgauge/constraints.py +105 -0
  27. agentgauge_harness-0.4.0/agentgauge/exp1_classifier.py +191 -0
  28. agentgauge_harness-0.4.0/agentgauge/exp1_doc_density.py +172 -0
  29. agentgauge_harness-0.4.0/agentgauge/exp1_mirror.py +273 -0
  30. agentgauge_harness-0.4.0/agentgauge/exp1_tool_def_extractor.py +406 -0
  31. agentgauge_harness-0.4.0/agentgauge/fixer.py +927 -0
  32. agentgauge_harness-0.4.0/agentgauge/frontier.py +69 -0
  33. agentgauge_harness-0.4.0/agentgauge/frozen_protocol.py +131 -0
  34. agentgauge_harness-0.4.0/agentgauge/harness.py +894 -0
  35. agentgauge_harness-0.4.0/agentgauge/linter.py +607 -0
  36. agentgauge_harness-0.4.0/agentgauge/localizer.py +326 -0
  37. agentgauge_harness-0.4.0/agentgauge/providers.py +326 -0
  38. agentgauge_harness-0.4.0/agentgauge/q2a_harness.py +87 -0
  39. agentgauge_harness-0.4.0/agentgauge/report.py +146 -0
  40. agentgauge_harness-0.4.0/agentgauge/runner.py +108 -0
  41. agentgauge_harness-0.4.0/agentgauge/scorer.py +947 -0
  42. agentgauge_harness-0.4.0/agentgauge/tasks.py +48 -0
  43. agentgauge_harness-0.4.0/agentgauge_pivot_onepager.md +81 -0
  44. agentgauge_harness-0.4.0/agentready-spec.md +123 -0
  45. agentgauge_harness-0.4.0/docs/paper/evidence_table.md +221 -0
  46. agentgauge_harness-0.4.0/docs/paper/latex/README.md +21 -0
  47. agentgauge_harness-0.4.0/docs/paper/latex/abstract_body.tex +24 -0
  48. agentgauge_harness-0.4.0/docs/paper/latex/body_content.tex +1318 -0
  49. agentgauge_harness-0.4.0/docs/paper/latex/main.pdf +0 -0
  50. agentgauge_harness-0.4.0/docs/paper/latex/main.tex +49 -0
  51. agentgauge_harness-0.4.0/docs/paper/latex/references.bib +73 -0
  52. agentgauge_harness-0.4.0/docs/paper/paper.md +881 -0
  53. agentgauge_harness-0.4.0/docs/paper/repo_triage.md +344 -0
  54. agentgauge_harness-0.4.0/docs/paper/revision_changelog.md +422 -0
  55. agentgauge_harness-0.4.0/docs/paper/skeleton.md +118 -0
  56. agentgauge_harness-0.4.0/docs/paper/threats_to_validity.md +125 -0
  57. agentgauge_harness-0.4.0/docs/research/exp3_pre_registration.md +267 -0
  58. agentgauge_harness-0.4.0/docs/research/exp4_regime_map.md +308 -0
  59. agentgauge_harness-0.4.0/docs/research/frontier_t18_result.md +47 -0
  60. agentgauge_harness-0.4.0/docs/research/frozen_protocol.md +179 -0
  61. agentgauge_harness-0.4.0/docs/research/phase1-buyer-and-landscape.md +148 -0
  62. agentgauge_harness-0.4.0/evals/__init__.py +1 -0
  63. agentgauge_harness-0.4.0/evals/fixtures/__init__.py +1 -0
  64. agentgauge_harness-0.4.0/evals/fixtures/exp1_anchor_validation.json +96 -0
  65. agentgauge_harness-0.4.0/evals/fixtures/exp1_candidate_list.json +5333 -0
  66. agentgauge_harness-0.4.0/evals/fixtures/exp1_doc_density_scores.json +469 -0
  67. agentgauge_harness-0.4.0/evals/fixtures/exp1_exclusion_log.json +177 -0
  68. agentgauge_harness-0.4.0/evals/fixtures/exp1_family_candidates.json +936 -0
  69. agentgauge_harness-0.4.0/evals/fixtures/exp1_mirrors/AminForou-mcp-gsc.json +439 -0
  70. agentgauge_harness-0.4.0/evals/fixtures/exp1_mirrors/Dataojitori-nocturne_memory.json +181 -0
  71. agentgauge_harness-0.4.0/evals/fixtures/exp1_mirrors/LycheeMem-LycheeMem.json +38 -0
  72. agentgauge_harness-0.4.0/evals/fixtures/exp1_mirrors/aws-iam-mcp.json +647 -0
  73. agentgauge_harness-0.4.0/evals/fixtures/exp1_mirrors/blazickjp-arxiv-mcp-server.json +66 -0
  74. agentgauge_harness-0.4.0/evals/fixtures/exp1_mirrors/datalayer-jupyter-mcp-server.json +359 -0
  75. agentgauge_harness-0.4.0/evals/fixtures/exp1_mirrors/github-mcp.json +663 -0
  76. agentgauge_harness-0.4.0/evals/fixtures/exp1_mirrors/lucasastorian-llmwiki.json +293 -0
  77. agentgauge_harness-0.4.0/evals/fixtures/exp1_mirrors/mrexodia-ida-pro-mcp.json +1367 -0
  78. agentgauge_harness-0.4.0/evals/fixtures/exp1_mirrors/stefanoamorelli-sec-edgar-mcp.json +438 -0
  79. agentgauge_harness-0.4.0/evals/fixtures/exp1_mirrors/stickerdaniel-linkedin-mcp-server.json +420 -0
  80. agentgauge_harness-0.4.0/evals/fixtures/exp1_mirrors/taylorwilsdon-google_workspace_mcp.json +4878 -0
  81. agentgauge_harness-0.4.0/evals/fixtures/exp1_mirrors_manifest.json +118 -0
  82. agentgauge_harness-0.4.0/evals/fixtures/exp1_pre_registration.json +169 -0
  83. agentgauge_harness-0.4.0/evals/fixtures/exp1_registry_candidates.json +13781 -0
  84. agentgauge_harness-0.4.0/evals/fixtures/exp1_server_frame.json +296 -0
  85. agentgauge_harness-0.4.0/evals/fixtures/exp1_trial_AminForou-mcp-gsc_arm_a.json +296 -0
  86. agentgauge_harness-0.4.0/evals/fixtures/exp1_trial_AminForou-mcp-gsc_arm_b.json +182 -0
  87. agentgauge_harness-0.4.0/evals/fixtures/exp1_trial_batch_summary.json +62 -0
  88. agentgauge_harness-0.4.0/evals/fixtures/exp1_trial_datalayer-jupyter-mcp-server_arm_a.json +296 -0
  89. agentgauge_harness-0.4.0/evals/fixtures/exp1_trial_datalayer-jupyter-mcp-server_arm_b.json +182 -0
  90. agentgauge_harness-0.4.0/evals/fixtures/exp1_trial_lucasastorian-llmwiki_arm_a.json +296 -0
  91. agentgauge_harness-0.4.0/evals/fixtures/exp1_trial_mrexodia-ida-pro-mcp_arm_a.json +296 -0
  92. agentgauge_harness-0.4.0/evals/fixtures/exp1_trial_stefanoamorelli-sec-edgar-mcp_arm_a.json +296 -0
  93. agentgauge_harness-0.4.0/evals/fixtures/exp1_trial_stickerdaniel-linkedin-mcp-server_arm_a.json +296 -0
  94. agentgauge_harness-0.4.0/evals/fixtures/exp1_trial_taylorwilsdon-google_workspace_mcp_arm_a.json +296 -0
  95. agentgauge_harness-0.4.0/evals/fixtures/exp1_trial_taylorwilsdon-google_workspace_mcp_arm_b.json +182 -0
  96. agentgauge_harness-0.4.0/evals/fixtures/exp3_ground_truth.json +272 -0
  97. agentgauge_harness-0.4.0/evals/fixtures/exp3_localizer_graded_result.json +408 -0
  98. agentgauge_harness-0.4.0/evals/fixtures/exp3_localizer_result.json +383 -0
  99. agentgauge_harness-0.4.0/evals/fixtures/frontier_t18_step2_raw_calls.json +1 -0
  100. agentgauge_harness-0.4.0/evals/fixtures/frontier_t18_step2_result.json +1486 -0
  101. agentgauge_harness-0.4.0/evals/fixtures/p2a_arm_guardb_descriptions.json +50 -0
  102. agentgauge_harness-0.4.0/evals/fixtures/p2a_f2_retrieval_spec.json +57 -0
  103. agentgauge_harness-0.4.0/evals/fixtures/p2a_internal_proxy_catalog.py +872 -0
  104. agentgauge_harness-0.4.0/evals/fixtures/predictive_validity/__init__.py +1 -0
  105. agentgauge_harness-0.4.0/evals/fixtures/predictive_validity/blind_tasks.py +937 -0
  106. agentgauge_harness-0.4.0/evals/fixtures/predictive_validity/constraints.py +1934 -0
  107. agentgauge_harness-0.4.0/evals/fixtures/predictive_validity/manifest.py +271 -0
  108. agentgauge_harness-0.4.0/evals/fixtures/predictive_validity/phase3_mechanism_results.json +205 -0
  109. agentgauge_harness-0.4.0/evals/fixtures/predictive_validity/results_raw.json +45714 -0
  110. agentgauge_harness-0.4.0/evals/fixtures/predictive_validity/results_raw.ndjson +88 -0
  111. agentgauge_harness-0.4.0/evals/fixtures/predictive_validity/results_raw_INVALID_leaked_tasks.json +1585 -0
  112. agentgauge_harness-0.4.0/evals/fixtures/predictive_validity/results_raw_INVALID_leaked_tasks.ndjson +14 -0
  113. agentgauge_harness-0.4.0/evals/fixtures/predictive_validity/results_raw_PHASE2_binary_1trial.json +3652 -0
  114. agentgauge_harness-0.4.0/evals/fixtures/predictive_validity/results_raw_PHASE2_binary_1trial.ndjson +18 -0
  115. agentgauge_harness-0.4.0/evals/fixtures/predictive_validity/schema_consistency_results.json +8261 -0
  116. agentgauge_harness-0.4.0/evals/fixtures/q3_arm_f_body_descriptions.json +14 -0
  117. agentgauge_harness-0.4.0/evals/fixtures/q3_arm_f_doc_descriptions.json +14 -0
  118. agentgauge_harness-0.4.0/evals/fixtures/q3_catalog.py +204 -0
  119. agentgauge_harness-0.4.0/evals/fixtures/q4_arm_f_body_scoped_descriptions.json +14 -0
  120. agentgauge_harness-0.4.0/evals/fixtures/q4_arm_f_doc_scoped_descriptions.json +14 -0
  121. agentgauge_harness-0.4.0/evals/fixtures/q5_arm_f_doc_guarded_descriptions.json +14 -0
  122. agentgauge_harness-0.4.0/evals/fixtures/q6_arm_f_doc_guarded_descriptions.json +25 -0
  123. agentgauge_harness-0.4.0/evals/fixtures/q6_catalog.py +405 -0
  124. agentgauge_harness-0.4.0/evals/fixtures/rw1_arm_guardb_descriptions.json +23 -0
  125. agentgauge_harness-0.4.0/evals/fixtures/rw1_github_catalog.py +678 -0
  126. agentgauge_harness-0.4.0/evals/fixtures/rw2_arm_guardb_descriptions.json +31 -0
  127. agentgauge_harness-0.4.0/evals/fixtures/rw2_aws_iam_catalog.py +1125 -0
  128. agentgauge_harness-0.4.0/evals/fixtures/t17_tasks.py +69 -0
  129. agentgauge_harness-0.4.0/evals/fixtures/t18_arm_f_descriptions.json +62 -0
  130. agentgauge_harness-0.4.0/evals/fixtures/t18_arm_f_q2b_descriptions.json +62 -0
  131. agentgauge_harness-0.4.0/evals/fixtures/t18_catalog.py +521 -0
  132. agentgauge_harness-0.4.0/evals/fixtures/ty2_tasks.py +292 -0
  133. agentgauge_harness-0.4.0/evals/fixtures/ty_tasks.py +191 -0
  134. agentgauge_harness-0.4.0/evals/fixtures/v2_1_cross_model_validation.json +44 -0
  135. agentgauge_harness-0.4.0/evals/fixtures/v2_1_false_alarm_new_estimator.json +94 -0
  136. agentgauge_harness-0.4.0/evals/fixtures/v2_1_linter_recall_fix.json +470 -0
  137. agentgauge_harness-0.4.0/evals/fixtures/v2_1_llm_linter_baseline.json +765 -0
  138. agentgauge_harness-0.4.0/evals/fixtures/v2_1_mde_ablation.json +101 -0
  139. agentgauge_harness-0.4.0/evals/fixtures/v2_1_severity_gate.json +16 -0
  140. agentgauge_harness-0.4.0/evals/fixtures/v2_2_causal_chain.json +572 -0
  141. agentgauge_harness-0.4.0/evals/fixtures/v2_2_causal_chain_multimodel.json +1718 -0
  142. agentgauge_harness-0.4.0/evals/fixtures/v2_2_cross_model_full.json +44 -0
  143. agentgauge_harness-0.4.0/evals/fixtures/v2_2_cross_model_pooled.json +44 -0
  144. agentgauge_harness-0.4.0/evals/fixtures/v2_2_few_clusters_correction.json +125 -0
  145. agentgauge_harness-0.4.0/evals/fixtures/v2_2_optimal_allocation.json +247 -0
  146. agentgauge_harness-0.4.0/evals/fixtures/v2_3_advisory_audit.json +586 -0
  147. agentgauge_harness-0.4.0/evals/fixtures/v2_3_corrected_advisory_effect.json +407 -0
  148. agentgauge_harness-0.4.0/evals/fixtures/v2_4_blocking_remeasurement.json +1142 -0
  149. agentgauge_harness-0.4.0/evals/fixtures/v2_4_corpus/__init__.py +1 -0
  150. agentgauge_harness-0.4.0/evals/fixtures/v2_4_corpus/aws_s3_NOTES.md +88 -0
  151. agentgauge_harness-0.4.0/evals/fixtures/v2_4_corpus/aws_s3_fixture.py +282 -0
  152. agentgauge_harness-0.4.0/evals/fixtures/v2_4_corpus/docker_containers_NOTES.md +82 -0
  153. agentgauge_harness-0.4.0/evals/fixtures/v2_4_corpus/docker_containers_fixture.py +299 -0
  154. agentgauge_harness-0.4.0/evals/fixtures/v2_4_corpus/gcal_NOTES.md +77 -0
  155. agentgauge_harness-0.4.0/evals/fixtures/v2_4_corpus/gcal_fixture.py +239 -0
  156. agentgauge_harness-0.4.0/evals/fixtures/v2_4_corpus/github_issues_NOTES.md +52 -0
  157. agentgauge_harness-0.4.0/evals/fixtures/v2_4_corpus/github_issues_fixture.py +293 -0
  158. agentgauge_harness-0.4.0/evals/fixtures/v2_4_corpus/jira_issues_NOTES.md +64 -0
  159. agentgauge_harness-0.4.0/evals/fixtures/v2_4_corpus/jira_issues_fixture.py +292 -0
  160. agentgauge_harness-0.4.0/evals/fixtures/v2_4_corpus/k8s_workloads_NOTES.md +87 -0
  161. agentgauge_harness-0.4.0/evals/fixtures/v2_4_corpus/k8s_workloads_fixture.py +301 -0
  162. agentgauge_harness-0.4.0/evals/fixtures/v2_4_corpus/slack_messaging_NOTES.md +94 -0
  163. agentgauge_harness-0.4.0/evals/fixtures/v2_4_corpus/slack_messaging_fixture.py +197 -0
  164. agentgauge_harness-0.4.0/evals/fixtures/v2_4_corpus/spotify_playlists_NOTES.md +76 -0
  165. agentgauge_harness-0.4.0/evals/fixtures/v2_4_corpus/spotify_playlists_fixture.py +232 -0
  166. agentgauge_harness-0.4.0/evals/fixtures/v2_4_corpus/stripe_payments_NOTES.md +76 -0
  167. agentgauge_harness-0.4.0/evals/fixtures/v2_4_corpus/stripe_payments_fixture.py +208 -0
  168. agentgauge_harness-0.4.0/evals/fixtures/v2_4_corpus/twilio_messaging_NOTES.md +64 -0
  169. agentgauge_harness-0.4.0/evals/fixtures/v2_4_corpus/twilio_messaging_fixture.py +274 -0
  170. agentgauge_harness-0.4.0/evals/fixtures/v2_5_argument_degradation_live.jsonl +1518 -0
  171. agentgauge_harness-0.4.0/evals/fixtures/v2_5_argument_degradation_summary.json +50 -0
  172. agentgauge_harness-0.4.0/evals/fixtures/v2_defect_injection_results.json +1961 -0
  173. agentgauge_harness-0.4.0/evals/fixtures/v2_false_alarm_determinism.json +279 -0
  174. agentgauge_harness-0.4.0/evals/fixtures/v2_lint_baseline.json +20946 -0
  175. agentgauge_harness-0.4.0/evals/fixtures/v2_mde_continuous_crosscheck.json +50 -0
  176. agentgauge_harness-0.4.0/evals/fixtures/v2_mde_table.json +114 -0
  177. agentgauge_harness-0.4.0/evals/fixtures/v2_tool_definitions.json +20856 -0
  178. agentgauge_harness-0.4.0/evals/fixtures/v2_variance_structure.json +62 -0
  179. agentgauge_harness-0.4.0/examples/aws_s3_server.py +120 -0
  180. agentgauge_harness-0.4.0/examples/aws_s3_server_fixed.py +162 -0
  181. agentgauge_harness-0.4.0/examples/call_constraints_server.py +153 -0
  182. agentgauge_harness-0.4.0/examples/call_constraints_server_fixed.py +153 -0
  183. agentgauge_harness-0.4.0/examples/call_constraints_server_oracle.py +184 -0
  184. agentgauge_harness-0.4.0/examples/call_constraints_v2_server.py +137 -0
  185. agentgauge_harness-0.4.0/examples/call_constraints_v2_server_fixed.py +137 -0
  186. agentgauge_harness-0.4.0/examples/call_constraints_v2_server_oracle.py +180 -0
  187. agentgauge_harness-0.4.0/examples/confusable_server.py +266 -0
  188. agentgauge_harness-0.4.0/examples/confusable_server_fixed.py +266 -0
  189. agentgauge_harness-0.4.0/examples/confusable_server_oracle.py +288 -0
  190. agentgauge_harness-0.4.0/examples/docker_containers_server.py +120 -0
  191. agentgauge_harness-0.4.0/examples/docker_containers_server_fixed.py +168 -0
  192. agentgauge_harness-0.4.0/examples/echo_server.py +112 -0
  193. agentgauge_harness-0.4.0/examples/echo_server_fixed.py +114 -0
  194. agentgauge_harness-0.4.0/examples/echo_server_fixed_dqonly.py +112 -0
  195. agentgauge_harness-0.4.0/examples/exp1_AminForou_mcp_gsc_mirror.py +437 -0
  196. agentgauge_harness-0.4.0/examples/exp1_AminForou_mcp_gsc_mirror_oracle.py +437 -0
  197. agentgauge_harness-0.4.0/examples/exp1_Dataojitori_nocturne_memory_mirror.py +213 -0
  198. agentgauge_harness-0.4.0/examples/exp1_LycheeMem_LycheeMem_mirror.py +89 -0
  199. agentgauge_harness-0.4.0/examples/exp1_blazickjp_arxiv_mcp_server_mirror.py +113 -0
  200. agentgauge_harness-0.4.0/examples/exp1_datalayer_jupyter_mcp_server_mirror.py +369 -0
  201. agentgauge_harness-0.4.0/examples/exp1_datalayer_jupyter_mcp_server_mirror_oracle.py +369 -0
  202. agentgauge_harness-0.4.0/examples/exp1_dataojitori_nocturne_memory_mirror_fixed.py +213 -0
  203. agentgauge_harness-0.4.0/examples/exp1_lucasastorian_llmwiki_mirror.py +307 -0
  204. agentgauge_harness-0.4.0/examples/exp1_mrexodia_ida_pro_mcp_mirror.py +1279 -0
  205. agentgauge_harness-0.4.0/examples/exp1_stefanoamorelli_sec_edgar_mcp_mirror.py +441 -0
  206. agentgauge_harness-0.4.0/examples/exp1_stickerdaniel_linkedin_mcp_server_mirror.py +419 -0
  207. agentgauge_harness-0.4.0/examples/exp1_taylorwilsdon_google_workspace_mcp_mirror.py +4145 -0
  208. agentgauge_harness-0.4.0/examples/exp1_taylorwilsdon_google_workspace_mcp_mirror_oracle.py +4145 -0
  209. agentgauge_harness-0.4.0/examples/gcal_server.py +118 -0
  210. agentgauge_harness-0.4.0/examples/gcal_server_fixed.py +150 -0
  211. agentgauge_harness-0.4.0/examples/github_issues_server.py +119 -0
  212. agentgauge_harness-0.4.0/examples/github_issues_server_fixed.py +155 -0
  213. agentgauge_harness-0.4.0/examples/grounded_server.py +183 -0
  214. agentgauge_harness-0.4.0/examples/grounded_server_fixed.py +183 -0
  215. agentgauge_harness-0.4.0/examples/grounded_server_oracle.py +164 -0
  216. agentgauge_harness-0.4.0/examples/jira_issues_server.py +117 -0
  217. agentgauge_harness-0.4.0/examples/jira_issues_server_fixed.py +152 -0
  218. agentgauge_harness-0.4.0/examples/k8s_workloads_server.py +124 -0
  219. agentgauge_harness-0.4.0/examples/k8s_workloads_server_fixed.py +172 -0
  220. agentgauge_harness-0.4.0/examples/mediocre_server.py +253 -0
  221. agentgauge_harness-0.4.0/examples/mediocre_server_fixed.py +258 -0
  222. agentgauge_harness-0.4.0/examples/p2a_arm_a.py +56 -0
  223. agentgauge_harness-0.4.0/examples/p2a_arm_guardb.py +75 -0
  224. agentgauge_harness-0.4.0/examples/p2a_arm_oracle.py +58 -0
  225. agentgauge_harness-0.4.0/examples/p2a_internal_proxy_mirror.py +499 -0
  226. agentgauge_harness-0.4.0/examples/q3_arm_a.py +57 -0
  227. agentgauge_harness-0.4.0/examples/q3_arm_f_body.py +69 -0
  228. agentgauge_harness-0.4.0/examples/q3_arm_f_doc.py +69 -0
  229. agentgauge_harness-0.4.0/examples/q3_arm_o.py +56 -0
  230. agentgauge_harness-0.4.0/examples/q3_real_server.py +223 -0
  231. agentgauge_harness-0.4.0/examples/q3_real_server_fixed.py +223 -0
  232. agentgauge_harness-0.4.0/examples/q3_real_server_fixed_dqonly.py +223 -0
  233. agentgauge_harness-0.4.0/examples/q4_arm_f_body_scoped.py +70 -0
  234. agentgauge_harness-0.4.0/examples/q4_arm_f_doc_scoped.py +70 -0
  235. agentgauge_harness-0.4.0/examples/q5_arm_f_doc_guarded.py +71 -0
  236. agentgauge_harness-0.4.0/examples/q6_arm_a.py +58 -0
  237. agentgauge_harness-0.4.0/examples/q6_arm_f_doc_guarded.py +72 -0
  238. agentgauge_harness-0.4.0/examples/q6_real_server.py +423 -0
  239. agentgauge_harness-0.4.0/examples/rw1_arm_a.py +57 -0
  240. agentgauge_harness-0.4.0/examples/rw1_arm_guardb.py +74 -0
  241. agentgauge_harness-0.4.0/examples/rw1_arm_oracle.py +56 -0
  242. agentgauge_harness-0.4.0/examples/rw1_github_mirror.py +402 -0
  243. agentgauge_harness-0.4.0/examples/rw2_arm_a.py +52 -0
  244. agentgauge_harness-0.4.0/examples/rw2_arm_guardb.py +69 -0
  245. agentgauge_harness-0.4.0/examples/rw2_aws_iam_mirror.py +484 -0
  246. agentgauge_harness-0.4.0/examples/slack_messaging_server.py +116 -0
  247. agentgauge_harness-0.4.0/examples/slack_messaging_server_fixed.py +149 -0
  248. agentgauge_harness-0.4.0/examples/spotify_playlists_server.py +118 -0
  249. agentgauge_harness-0.4.0/examples/spotify_playlists_server_fixed.py +154 -0
  250. agentgauge_harness-0.4.0/examples/stripe_payments_server.py +119 -0
  251. agentgauge_harness-0.4.0/examples/stripe_payments_server_fixed.py +158 -0
  252. agentgauge_harness-0.4.0/examples/t18_fixer_server.py +69 -0
  253. agentgauge_harness-0.4.0/examples/t18_oracle_server.py +67 -0
  254. agentgauge_harness-0.4.0/examples/t18_q2b_server.py +81 -0
  255. agentgauge_harness-0.4.0/examples/t18_vague_server.py +66 -0
  256. agentgauge_harness-0.4.0/examples/twilio_messaging_server.py +119 -0
  257. agentgauge_harness-0.4.0/examples/twilio_messaging_server_fixed.py +169 -0
  258. agentgauge_harness-0.4.0/p2a_spec.md +133 -0
  259. agentgauge_harness-0.4.0/paper_framing_options.md +65 -0
  260. agentgauge_harness-0.4.0/pyproject.toml +76 -0
  261. agentgauge_harness-0.4.0/scripts/Dockerfile.agentgauge-agent +35 -0
  262. agentgauge_harness-0.4.0/scripts/_mutated_stdio_server.py +106 -0
  263. agentgauge_harness-0.4.0/scripts/agentgauge-agent-service.yaml +44 -0
  264. agentgauge_harness-0.4.0/scripts/agentgauge-judge-service.yaml +53 -0
  265. agentgauge_harness-0.4.0/scripts/build_fixed_fixtures.py +127 -0
  266. agentgauge_harness-0.4.0/scripts/build_fixed_fixtures_v2.py +97 -0
  267. agentgauge_harness-0.4.0/scripts/check_paper_latex_sync.py +68 -0
  268. agentgauge_harness-0.4.0/scripts/exp1_build_mirrors.py +278 -0
  269. agentgauge_harness-0.4.0/scripts/exp1_discover_registry.py +141 -0
  270. agentgauge_harness-0.4.0/scripts/exp1_discover_servers.py +255 -0
  271. agentgauge_harness-0.4.0/scripts/exp1_generate_mirror_server.py +168 -0
  272. agentgauge_harness-0.4.0/scripts/exp1_identify_families.py +102 -0
  273. agentgauge_harness-0.4.0/scripts/exp1_run_remaining_trials.py +374 -0
  274. agentgauge_harness-0.4.0/scripts/exp1_run_trial.py +294 -0
  275. agentgauge_harness-0.4.0/scripts/exp1_score_doc_density.py +158 -0
  276. agentgauge_harness-0.4.0/scripts/exp1_validate_anchors.py +168 -0
  277. agentgauge_harness-0.4.0/scripts/exp3_run_localizer.py +346 -0
  278. agentgauge_harness-0.4.0/scripts/generate_arm_f_descriptions.py +61 -0
  279. agentgauge_harness-0.4.0/scripts/generate_arm_f_descriptions_q2b.py +70 -0
  280. agentgauge_harness-0.4.0/scripts/generate_q3_descriptions.py +150 -0
  281. agentgauge_harness-0.4.0/scripts/generate_q4_descriptions.py +189 -0
  282. agentgauge_harness-0.4.0/scripts/generate_q5_descriptions.py +173 -0
  283. agentgauge_harness-0.4.0/scripts/generate_q6_descriptions.py +206 -0
  284. agentgauge_harness-0.4.0/scripts/mde_grid_v2_5.py +38 -0
  285. agentgauge_harness-0.4.0/scripts/p2a_f2_retrieval.py +452 -0
  286. agentgauge_harness-0.4.0/scripts/p2a_frontier_gate.py +337 -0
  287. agentgauge_harness-0.4.0/scripts/p2a_phase1_generate.py +188 -0
  288. agentgauge_harness-0.4.0/scripts/p2a_phase2_ab.py +600 -0
  289. agentgauge_harness-0.4.0/scripts/phase3_mechanism_test.py +228 -0
  290. agentgauge_harness-0.4.0/scripts/predictive_validity_analysis.py +536 -0
  291. agentgauge_harness-0.4.0/scripts/predictive_validity_study.py +440 -0
  292. agentgauge_harness-0.4.0/scripts/run_ab_experiment.py +207 -0
  293. agentgauge_harness-0.4.0/scripts/run_all_phase3_expansion_builds.py +82 -0
  294. agentgauge_harness-0.4.0/scripts/run_build_fixed_fixtures_via_gcp.py +30 -0
  295. agentgauge_harness-0.4.0/scripts/run_frontier_t18.py +452 -0
  296. agentgauge_harness-0.4.0/scripts/run_predictive_validity_via_gcp.py +32 -0
  297. agentgauge_harness-0.4.0/scripts/run_q2a_three_arm.py +424 -0
  298. agentgauge_harness-0.4.0/scripts/run_q2b_three_arm.py +404 -0
  299. agentgauge_harness-0.4.0/scripts/run_q3_four_arm.py +408 -0
  300. agentgauge_harness-0.4.0/scripts/run_q4_four_arm.py +461 -0
  301. agentgauge_harness-0.4.0/scripts/run_q5_four_arm.py +450 -0
  302. agentgauge_harness-0.4.0/scripts/run_q6_regression.py +508 -0
  303. agentgauge_harness-0.4.0/scripts/run_schema_consistency.py +98 -0
  304. agentgauge_harness-0.4.0/scripts/run_t17_oracle_ab.py +313 -0
  305. agentgauge_harness-0.4.0/scripts/run_t18_oracle_ab.py +333 -0
  306. agentgauge_harness-0.4.0/scripts/run_tx_experiment.py +392 -0
  307. agentgauge_harness-0.4.0/scripts/run_ty2_oracle_ab.py +420 -0
  308. agentgauge_harness-0.4.0/scripts/run_ty_oracle_ab.py +370 -0
  309. agentgauge_harness-0.4.0/scripts/rw1_part1_discoverability.py +235 -0
  310. agentgauge_harness-0.4.0/scripts/rw1_phase1_generate.py +194 -0
  311. agentgauge_harness-0.4.0/scripts/rw1_phase2_ab.py +431 -0
  312. agentgauge_harness-0.4.0/scripts/rw2_phase1_generate.py +213 -0
  313. agentgauge_harness-0.4.0/scripts/rw2_phase2_ab.py +469 -0
  314. agentgauge_harness-0.4.0/scripts/schema_consistency_checker.py +172 -0
  315. agentgauge_harness-0.4.0/scripts/v2_1_cross_model_validation.py +238 -0
  316. agentgauge_harness-0.4.0/scripts/v2_1_false_alarm_new_estimator.py +121 -0
  317. agentgauge_harness-0.4.0/scripts/v2_1_linter_recall_fix.py +176 -0
  318. agentgauge_harness-0.4.0/scripts/v2_1_llm_linter_baseline.py +220 -0
  319. agentgauge_harness-0.4.0/scripts/v2_1_mde_ablation.py +160 -0
  320. agentgauge_harness-0.4.0/scripts/v2_1_severity_gate_measurement.py +81 -0
  321. agentgauge_harness-0.4.0/scripts/v2_2_causal_chain.py +273 -0
  322. agentgauge_harness-0.4.0/scripts/v2_2_causal_chain_multimodel.py +258 -0
  323. agentgauge_harness-0.4.0/scripts/v2_2_cross_model_full.py +155 -0
  324. agentgauge_harness-0.4.0/scripts/v2_2_cross_model_pooled.py +183 -0
  325. agentgauge_harness-0.4.0/scripts/v2_2_few_clusters_correction.py +122 -0
  326. agentgauge_harness-0.4.0/scripts/v2_2_optimal_allocation.py +170 -0
  327. agentgauge_harness-0.4.0/scripts/v2_3_advisory_audit.py +155 -0
  328. agentgauge_harness-0.4.0/scripts/v2_3_corrected_advisory_effect.py +135 -0
  329. agentgauge_harness-0.4.0/scripts/v2_4_blocking_remeasurement.py +235 -0
  330. agentgauge_harness-0.4.0/scripts/v2_5_argument_degradation_live.py +348 -0
  331. agentgauge_harness-0.4.0/scripts/v2_5_argument_degradation_live_gcp.py +97 -0
  332. agentgauge_harness-0.4.0/scripts/v2_defect_injector.py +330 -0
  333. agentgauge_harness-0.4.0/scripts/v2_extract_tool_definitions.py +72 -0
  334. agentgauge_harness-0.4.0/scripts/v2_false_alarm_and_determinism.py +110 -0
  335. agentgauge_harness-0.4.0/scripts/v2_mde_continuous_crosscheck.py +118 -0
  336. agentgauge_harness-0.4.0/scripts/v2_mde_table.py +79 -0
  337. agentgauge_harness-0.4.0/scripts/v2_variance_structure.py +260 -0
  338. agentgauge_harness-0.4.0/scripts/verify.sh +41 -0
  339. agentgauge_harness-0.4.0/spec.md +91 -0
  340. agentgauge_harness-0.4.0/spec_ty2.md +326 -0
  341. agentgauge_harness-0.4.0/tests/__init__.py +0 -0
  342. agentgauge_harness-0.4.0/tests/test_ab_harness.py +250 -0
  343. agentgauge_harness-0.4.0/tests/test_audit.py +352 -0
  344. agentgauge_harness-0.4.0/tests/test_cli.py +295 -0
  345. agentgauge_harness-0.4.0/tests/test_client.py +65 -0
  346. agentgauge_harness-0.4.0/tests/test_constraints.py +66 -0
  347. agentgauge_harness-0.4.0/tests/test_discoverability.py +297 -0
  348. agentgauge_harness-0.4.0/tests/test_docs_manifest.py +209 -0
  349. agentgauge_harness-0.4.0/tests/test_error_legibility.py +294 -0
  350. agentgauge_harness-0.4.0/tests/test_exp1_classifier.py +385 -0
  351. agentgauge_harness-0.4.0/tests/test_exp1_doc_density.py +169 -0
  352. agentgauge_harness-0.4.0/tests/test_exp1_generate_mirror_server.py +77 -0
  353. agentgauge_harness-0.4.0/tests/test_exp1_mirror.py +561 -0
  354. agentgauge_harness-0.4.0/tests/test_exp1_pre_reg.py +70 -0
  355. agentgauge_harness-0.4.0/tests/test_exp1_tool_def_extractor.py +490 -0
  356. agentgauge_harness-0.4.0/tests/test_fixer.py +1494 -0
  357. agentgauge_harness-0.4.0/tests/test_frontier.py +398 -0
  358. agentgauge_harness-0.4.0/tests/test_frozen_protocol.py +129 -0
  359. agentgauge_harness-0.4.0/tests/test_harness.py +593 -0
  360. agentgauge_harness-0.4.0/tests/test_json.py +144 -0
  361. agentgauge_harness-0.4.0/tests/test_linter.py +365 -0
  362. agentgauge_harness-0.4.0/tests/test_localizer.py +306 -0
  363. agentgauge_harness-0.4.0/tests/test_p2a_catalog.py +178 -0
  364. agentgauge_harness-0.4.0/tests/test_predictive_validity_analysis.py +477 -0
  365. agentgauge_harness-0.4.0/tests/test_q2a_recovery.py +230 -0
  366. agentgauge_harness-0.4.0/tests/test_q3.py +290 -0
  367. agentgauge_harness-0.4.0/tests/test_q4_scoped.py +291 -0
  368. agentgauge_harness-0.4.0/tests/test_q5_guarded.py +192 -0
  369. agentgauge_harness-0.4.0/tests/test_q6_fixture.py +305 -0
  370. agentgauge_harness-0.4.0/tests/test_report.py +99 -0
  371. agentgauge_harness-0.4.0/tests/test_robustness.py +302 -0
  372. agentgauge_harness-0.4.0/tests/test_runner.py +307 -0
  373. agentgauge_harness-0.4.0/tests/test_rw1_github.py +436 -0
  374. agentgauge_harness-0.4.0/tests/test_rw2_aws_iam.py +440 -0
  375. agentgauge_harness-0.4.0/tests/test_scorer.py +222 -0
  376. agentgauge_harness-0.4.0/tests/test_t17_fixture.py +140 -0
  377. agentgauge_harness-0.4.0/tests/test_t18_fixture.py +141 -0
  378. agentgauge_harness-0.4.0/tests/test_tasks.py +91 -0
  379. agentgauge_harness-0.4.0/tests/test_ty2_fixture.py +285 -0
  380. agentgauge_harness-0.4.0/tests/test_ty_fixture.py +182 -0
  381. agentgauge_harness-0.4.0/tests/test_ux1.py +435 -0
  382. agentgauge_harness-0.4.0/uv.lock +1269 -0
@@ -0,0 +1,23 @@
1
+ ---
2
+ name: executor
3
+ description: Use this subagent to implement specific, well-scoped code changes. Invoke with a clear task description, target files, and any constraints. The subagent writes code, runs verification, and reports back. Do NOT invoke for planning, architecture decisions, or open-ended exploration.
4
+ tools: Read, Write, Edit, Glob, Grep, Bash
5
+ model: sonnet
6
+ ---
7
+
8
+ You are a senior software engineer focused on implementation. You receive a scoped task from the orchestrator and implement it cleanly.
9
+
10
+ Rules:
11
+ - Stay strictly within the scope assigned. If you discover the task is bigger than described, stop and report rather than expanding scope.
12
+ - Add type hints on every function, docstrings on non-trivial ones, and at least one unit test for any new non-trivial function.
13
+ - Run any verification commands the orchestrator specifies. If a command fails, fix and re-run - but if the same check fails 3 times in a row, stop and report with the error output.
14
+ - Commit in small, conventionally-named commits (feat:, fix:, test:, chore:, etc.). One concept per commit.
15
+ - Do NOT install new dependencies without asking. Surface the need in your report.
16
+ - Do NOT touch files outside the project directory.
17
+
18
+ Final report format:
19
+ - One-line summary of what was done
20
+ - List of commit messages
21
+ - Verification results (pass/fail per check)
22
+ - Any deviations from assigned scope, with reasoning
23
+ - Anything the orchestrator should know before the next task
@@ -0,0 +1,27 @@
1
+ ---
2
+ name: verifier
3
+ description: Use this subagent to run a project's verification commands (tests, linters, type checkers) and report structured pass/fail results. Invoke after any code-writing pass. Does not write code - only runs commands and reports.
4
+ tools: Read, Glob, Grep, Bash
5
+ model: sonnet
6
+ ---
7
+
8
+ You are a verification specialist. You run commands and report results. You do not write code, fix bugs, or interpret failures - the orchestrator handles that.
9
+
10
+ For each command in the list you are given:
11
+ 1. Run it in the project root.
12
+ 2. Capture exit code, the last 100 lines of combined stdout+stderr, and wall-clock duration.
13
+ 3. Classify as pass (exit 0) or fail (non-zero).
14
+
15
+ Always output in this format:
16
+
17
+ VERIFICATION RESULTS
18
+ ====================
19
+ [pass/fail] <name>: <command> (<duration>s)
20
+ Summary: <one-line summary>
21
+ Detail: <relevant excerpt of failure output, omit if passed>
22
+
23
+ [pass/fail] ...
24
+
25
+ OVERALL: <pass | fail>
26
+
27
+ You have no Write or Edit tools. Do not attempt to modify any files.
@@ -0,0 +1,126 @@
1
+ # AgentGauge — Autonomous Run Operator Prompt
2
+
3
+ You are a Claude Code cloud scheduled task running against the `agentgauge` repository.
4
+
5
+ ## Merge policy
6
+
7
+ **All PRs opened by this loop are DRAFT.** Never open a ready-to-merge PR autonomously.
8
+ Never call `gh pr merge`. The human reviews and merges every PR.
9
+
10
+ The following categories of task especially require human judgment beyond what a green test suite
11
+ can verify — but note that DRAFT is the rule for ALL tasks, not just these:
12
+
13
+ 1. **Touches the LLM judge** — changes to rubric prompts, scoring logic, calibration constants,
14
+ judge model selection, or blending weights. A green CI means the mock honored the contract,
15
+ not that the judge produces better scores on real inputs.
16
+
17
+ 2. **Generates fixes or actions against real servers** — any task that calls a live MCP server,
18
+ a real Ollama instance, or any external API as part of its primary operation.
19
+
20
+ 3. **Depends on real-model or real-network behavior** — tasks where the acceptance criteria
21
+ require measuring calibration, validating score bands, or comparing outputs across judge runs.
22
+ These require human review of measured results, not just a green test suite.
23
+
24
+ **The test suite is necessary but not sufficient.** Passing mocks prove the code path runs,
25
+ not that real-model calibration or real-server behavior is correct. All merges require human review.
26
+
27
+ ## Before you do anything
28
+
29
+ 1. Read `CLAUDE.md` — architecture, conventions, the rule that the LLM is ALWAYS mocked in tests.
30
+ 2. Read `AUTONOMY.md` — the hard rules you must never violate.
31
+ 3. Read `TASKS.md` — the backlog.
32
+
33
+ ## Pick your task
34
+
35
+ Select the **single highest-priority TODO item** in TASKS.md that has clear, testable acceptance
36
+ criteria.
37
+
38
+ **Before starting any work**, check whether an open `claude/*` PR already addresses that task:
39
+
40
+ ```bash
41
+ gh pr list --state open --search "<task-id>" --json number,title,headRefName
42
+ ```
43
+
44
+ For example, for T3:
45
+ ```bash
46
+ gh pr list --state open --search "T3" --json number,title,headRefName
47
+ ```
48
+
49
+ - If a matching open PR exists → **skip that item** and pick the next eligible TODO.
50
+ - If no next eligible TODO exists → stop here: write a short `BLOCKER.md` at repo root
51
+ explaining that all TODO items already have open PRs, commit it to a
52
+ `claude/blocker-<date>` branch, open a DRAFT PR titled "Blocker report: all TODOs in
53
+ review", and exit. **Never open a second PR for a task that already has one open.**
54
+
55
+ If no item qualifies at all (all are ambiguous, blocked, or already in review), **before writing
56
+ any BLOCKER.md**, check whether an open `claude/blocker-*` PR already exists:
57
+
58
+ ```bash
59
+ gh pr list --state open --search "Blocker report" --json number,title,headRefName
60
+ ```
61
+
62
+ - If an open blocker PR already exists → **do not open another one.** Instead, comment on the
63
+ existing PR with today's date and the current board state (e.g. `gh pr comment <number> --body
64
+ "Re-check <date>: board unchanged — TODO still empty, FUTURE/DEFERRED items still blocked."`),
65
+ then exit. One open blocker PR is enough; duplicate idle-path PRs add noise.
66
+ - If no open blocker PR exists → write `BLOCKER.md`, commit to `claude/blocker-<date>`, open a
67
+ DRAFT PR titled "Blocker report: all TODOs in review", and exit.
68
+
69
+ Do not invent work.
70
+
71
+ ## Implement
72
+
73
+ Use the orchestrator → executor → verifier pattern:
74
+ - Orchestrator (you): plan the change, break into steps, verify the plan against acceptance criteria.
75
+ - Executor subagent: implement the code changes only.
76
+ - Verifier subagent: confirm the change is correct and complete.
77
+
78
+ Branch name: `claude/<kebab-task-name>` (e.g., `claude/task-generator`).
79
+
80
+ ## Definition of done
81
+
82
+ Run `./scripts/verify.sh`. It exits 0 only if:
83
+ - `ruff check` passes
84
+ - `ruff format --check` passes
85
+ - `mypy` passes (if configured)
86
+ - All tests pass — no tests removed, no mocks weakened
87
+
88
+ ## Commit and PR
89
+
90
+ If `verify.sh` exits 0:
91
+
92
+ ### Step 1 — commit
93
+
94
+ - Commit with conventional-commit message: `feat(scope): description`
95
+ - **Do NOT include claude.ai session URLs in commit bodies.**
96
+ - Push branch: `git push origin claude/<task-name>`
97
+ - In TASKS.md on your branch, move the item from TODO to IN-REVIEW.
98
+ - Commit that TASKS.md update as a separate `chore: move <task> to IN-REVIEW` commit.
99
+
100
+ ### Step 2 — open PR
101
+
102
+ **All PRs are DRAFT.** Open every PR as draft regardless of task type:
103
+
104
+ ```bash
105
+ gh pr create --draft --title "..." --body "..."
106
+ ```
107
+
108
+ Do **not** call `gh pr merge` at all. Do **not** push directly to main.
109
+ Report the PR link. The human will review and merge.
110
+
111
+ If `verify.sh` does not exit 0:
112
+ - Do NOT commit
113
+ - Do NOT open a PR
114
+ - Write a blocker summary explaining what failed and why
115
+ - Exit
116
+
117
+ ## End of run report
118
+
119
+ Always end your run with a brief report (5-10 lines):
120
+ - Task selected
121
+ - Open-PR dedup check result (any skipped tasks and why)
122
+ - What was implemented
123
+ - `verify.sh` result (exit code + any relevant failure lines)
124
+ - PR link (if created)
125
+ - PR opened as DRAFT (confirm link)
126
+ - Any blockers or caveats
@@ -0,0 +1,5 @@
1
+ ## What & why (2–4 lines)
2
+ ## Changes (bulleted, grouped)
3
+ ## Testing (what was run, what it proves)
4
+ ## Screenshots (before/after — required for any visible UI change, else "n/a")
5
+ ## Risk & rollback (blast radius, revert plan)
@@ -0,0 +1,88 @@
1
+ name: CI
2
+
3
+ on:
4
+ push:
5
+ branches: [main]
6
+ pull_request:
7
+ branches: [main]
8
+
9
+ concurrency:
10
+ group: ${{ github.workflow }}-${{ github.ref }}
11
+ cancel-in-progress: true
12
+
13
+ jobs:
14
+ verify:
15
+ runs-on: ubuntu-latest
16
+ steps:
17
+ - uses: actions/checkout@v4
18
+
19
+ - uses: actions/setup-python@v5
20
+ with:
21
+ python-version: "3.11"
22
+
23
+ - name: Install uv
24
+ run: pip install uv --quiet
25
+
26
+ - name: Run verify.sh
27
+ run: bash scripts/verify.sh
28
+
29
+ hygiene:
30
+ runs-on: ubuntu-latest
31
+ steps:
32
+ - uses: actions/checkout@v4
33
+ with:
34
+ fetch-depth: 0
35
+
36
+ - name: Check commit hygiene — no claude.ai session URLs or Co-authored-by Claude trailers
37
+ run: |
38
+ set -euo pipefail
39
+ if [ "$GITHUB_EVENT_NAME" = "pull_request" ]; then
40
+ RANGE="origin/main..HEAD"
41
+ else
42
+ # Push to main: check only the newly pushed commit(s).
43
+ BEFORE="${{ github.event.before }}"
44
+ SHA="${{ github.sha }}"
45
+ if [ "$BEFORE" = "0000000000000000000000000000000000000000" ] || ! git cat-file -e "$BEFORE" 2>/dev/null; then
46
+ RANGE="HEAD^..HEAD"
47
+ else
48
+ RANGE="$BEFORE..$SHA"
49
+ fi
50
+ fi
51
+ MSGS=$(git log "$RANGE" --format='%B' 2>/dev/null || git log -1 --format='%B')
52
+ if echo "$MSGS" | grep -qF 'claude.ai/code/session'; then
53
+ echo "ERROR: commit message contains a claude.ai/code/session URL — remove it before merging"
54
+ exit 1
55
+ fi
56
+ if echo "$MSGS" | grep -iqF 'co-authored-by: claude'; then
57
+ echo "ERROR: commit message contains a Co-authored-by: Claude trailer — remove it before merging"
58
+ exit 1
59
+ fi
60
+ echo "OK: commit hygiene verified"
61
+ echo "RANGE=$RANGE" >> "$GITHUB_ENV"
62
+
63
+ - uses: actions/setup-python@v5
64
+ with:
65
+ python-version: "3.11"
66
+
67
+ - name: Check docs/paper/paper.md and docs/paper/latex/ stay in sync
68
+ run: python3 scripts/check_paper_latex_sync.py "$RANGE"
69
+
70
+ pip-audit:
71
+ runs-on: ubuntu-latest
72
+ steps:
73
+ - uses: actions/checkout@v4
74
+
75
+ - uses: actions/setup-python@v5
76
+ with:
77
+ python-version: "3.11"
78
+
79
+ - name: Install uv
80
+ run: pip install uv --quiet
81
+
82
+ - name: Sync dependencies
83
+ run: uv sync --extra dev
84
+
85
+ - name: Run pip-audit
86
+ # Non-blocking (T1 tier): findings are surfaced in logs but never fail the build.
87
+ continue-on-error: true
88
+ run: uv run --with pip-audit pip-audit
@@ -0,0 +1,43 @@
1
+ name: Release
2
+
3
+ on:
4
+ push:
5
+ tags:
6
+ - "v*"
7
+
8
+ concurrency:
9
+ group: ${{ github.workflow }}-${{ github.ref }}
10
+ cancel-in-progress: false
11
+
12
+ jobs:
13
+ publish:
14
+ runs-on: ubuntu-latest
15
+ environment: pypi
16
+ permissions:
17
+ id-token: write # required for PyPI Trusted Publishing (OIDC) -- no token/password used
18
+ contents: read
19
+
20
+ steps:
21
+ - uses: actions/checkout@v4
22
+ with:
23
+ # Explicit: build from the tag that triggered this workflow, never a
24
+ # branch head. For a `push: tags:` event, github.ref IS the tag ref
25
+ # (refs/tags/vX.Y.Z) -- stated explicitly here so this can't silently
26
+ # drift to a branch checkout if the trigger config ever changes.
27
+ ref: ${{ github.ref }}
28
+
29
+ - uses: actions/setup-python@v5
30
+ with:
31
+ python-version: "3.11"
32
+
33
+ - name: Install uv
34
+ run: pip install uv --quiet
35
+
36
+ - name: Build sdist + wheel
37
+ run: uv build
38
+
39
+ - name: twine check
40
+ run: uvx twine check dist/*.whl dist/*.tar.gz
41
+
42
+ - name: Publish to PyPI (Trusted Publishing -- OIDC, no token)
43
+ uses: pypa/gh-action-pypi-publish@release/v1
@@ -0,0 +1,89 @@
1
+ # Python
2
+ __pycache__/
3
+ *.py[cod]
4
+ *$py.class
5
+ *.so
6
+ .Python
7
+ build/
8
+ dist/
9
+ *.egg-info/
10
+ .eggs/
11
+ *.egg
12
+ MANIFEST
13
+
14
+ # Virtual environments
15
+ .venv/
16
+ venv/
17
+ env/
18
+ ENV/
19
+
20
+ # uv
21
+ .uv/
22
+
23
+ # Testing
24
+ .pytest_cache/
25
+ .coverage
26
+ coverage.xml
27
+ htmlcov/
28
+ *.coveragerc
29
+
30
+ # Mypy
31
+ .mypy_cache/
32
+ .dmypy.json
33
+ dmypy.json
34
+
35
+ # Ruff
36
+ .ruff_cache/
37
+
38
+ # IDE
39
+ .vscode/
40
+ .idea/
41
+ *.swp
42
+ *.swo
43
+
44
+ # OS
45
+ .DS_Store
46
+ Thumbs.db
47
+
48
+ # Secrets / env
49
+ .env
50
+ .env.*
51
+ !.env.example
52
+
53
+ # Distribution
54
+ *.whl
55
+ *.tar.gz
56
+
57
+ # Reports
58
+ reports/
59
+ *.html
60
+ !docs/**/*.html
61
+
62
+ # Worktrees
63
+ .claude/worktrees/
64
+
65
+ # Ad-hoc spot-check scripts (uncommitted per convention)
66
+ scripts/spot_check_discoverability_judge.py
67
+ scripts/validate_t5_docs_manifest.py
68
+ scripts/spot_check_fixer.py
69
+ scripts/pr32_body.txt
70
+ scripts/tx_breakdown_analysis.py
71
+ scripts/analyze_frontier_t18.py
72
+
73
+ # Run-output scratch files (experiment results, watchdog logs)
74
+ *_result.txt
75
+ *_results.txt
76
+ *_stderr.txt
77
+ *_watchdog.txt
78
+ *test_output.txt
79
+ ollama_conn_monitor.*
80
+ t18_watchdog.ps1
81
+ ty_t18_watchdog.txt
82
+
83
+ # LaTeX build artifacts (docs/paper/latex/main.pdf is committed deliberately)
84
+ docs/paper/latex/*.aux
85
+ docs/paper/latex/*.bbl
86
+ docs/paper/latex/*.log
87
+ docs/paper/latex/*.out
88
+ docs/paper/latex/*.blg
89
+ docs/paper/latex/*.toc
@@ -0,0 +1,27 @@
1
+ repos:
2
+ - repo: https://github.com/pre-commit/pre-commit-hooks
3
+ rev: v5.0.0
4
+ hooks:
5
+ - id: trailing-whitespace
6
+ - id: end-of-file-fixer
7
+ - id: check-yaml
8
+ - id: check-merge-conflict
9
+ - id: debug-statements
10
+ - id: check-added-large-files
11
+
12
+ - repo: https://github.com/astral-sh/ruff-pre-commit
13
+ rev: v0.15.15
14
+ hooks:
15
+ - id: ruff
16
+ args: [--fix]
17
+ - id: ruff-format
18
+
19
+ - repo: https://github.com/pre-commit/mirrors-mypy
20
+ rev: v2.1.0
21
+ hooks:
22
+ - id: mypy
23
+
24
+ - repo: https://github.com/gitleaks/gitleaks
25
+ rev: v8.29.1
26
+ hooks:
27
+ - id: gitleaks
@@ -0,0 +1,41 @@
1
+ # Autonomy Contract
2
+
3
+ This repo is driven by a Claude Code cloud scheduled task. Each run picks one
4
+ backlog item from TASKS.md, implements it on a `claude/<task>` branch, and opens a
5
+ draft PR for a human to review. These rules govern every autonomous run.
6
+
7
+ ## Hard rules
8
+
9
+ 1. **Only push to `claude/*` branches.** Never touch `main`, `chore/*`, `feat/*`, or any
10
+ branch not prefixed `claude/`.
11
+ 2. **Never modify CI, workflows, or secrets.** `.github/workflows/`, `.github/secrets`,
12
+ `pyproject.toml` CI config, and `scripts/verify.sh` are read-only unless the task
13
+ explicitly changes them.
14
+ 3. **Never delete, skip, or weaken tests to make `verify.sh` pass.** If tests fail, fix
15
+ the implementation — never remove the test. `pytest` must pass at full coverage.
16
+ 4. **Exactly one TASKS.md item per run.** Pick the single highest-priority TODO item
17
+ with clear, testable acceptance criteria. If none qualify, stop and write a blocker
18
+ report — do not invent work.
19
+ 5. **No dependency upgrades unless the task IS the upgrade.** Do not bump versions in
20
+ `pyproject.toml` as a side effect.
21
+ 6. **PRs are always DRAFT.** Never open a ready-to-merge PR autonomously.
22
+ 7. **If acceptance criteria are unclear, stop and report.** Do not guess at intent.
23
+ Write a `BLOCKER.md` at repo root explaining what is unclear and exit.
24
+
25
+ ## Workflow per run
26
+
27
+ 1. Read `CLAUDE.md` (architecture, conventions).
28
+ 2. Read `AUTONOMY.md` (this file).
29
+ 3. Read `TASKS.md` — pick the single highest-priority TODO with clear acceptance criteria.
30
+ 4. Branch: `claude/<kebab-task-name>`.
31
+ 5. Implement via orchestrator → executor → verifier pattern.
32
+ 6. Run `./scripts/verify.sh`. This is the definition of done.
33
+ 7. If green: commit, push branch, open DRAFT PR, move item to IN-REVIEW in TASKS.md.
34
+ 8. If not green: leave branch uncommitted, open no PR, write a blocker summary in the PR
35
+ description or a comment, do NOT move the item.
36
+ 9. End with a run report (what was done, verify.sh result, PR link if created).
37
+
38
+ ## Definition of done
39
+
40
+ `./scripts/verify.sh` exits 0. That means: ruff passes, mypy passes (if configured),
41
+ all tests pass with no mocks removed or tests deleted.