agentgauge-harness 0.4.0__tar.gz
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- agentgauge_harness-0.4.0/.claude/agents/executor.md +23 -0
- agentgauge_harness-0.4.0/.claude/agents/verifier.md +27 -0
- agentgauge_harness-0.4.0/.claude/operator-prompt.md +126 -0
- agentgauge_harness-0.4.0/.github/pull_request_template.md +5 -0
- agentgauge_harness-0.4.0/.github/workflows/ci.yml +88 -0
- agentgauge_harness-0.4.0/.github/workflows/release.yml +43 -0
- agentgauge_harness-0.4.0/.gitignore +89 -0
- agentgauge_harness-0.4.0/.pre-commit-config.yaml +27 -0
- agentgauge_harness-0.4.0/AUTONOMY.md +41 -0
- agentgauge_harness-0.4.0/CLAUDE.md +244 -0
- agentgauge_harness-0.4.0/LICENSE +190 -0
- agentgauge_harness-0.4.0/PKG-INFO +382 -0
- agentgauge_harness-0.4.0/PLAN.md +106 -0
- agentgauge_harness-0.4.0/PREDICTIVE_VALIDITY_HANDOFF.md +137 -0
- agentgauge_harness-0.4.0/README.md +345 -0
- agentgauge_harness-0.4.0/ROADMAP.md +56 -0
- agentgauge_harness-0.4.0/STATUS.md +1210 -0
- agentgauge_harness-0.4.0/TASKS.md +531 -0
- agentgauge_harness-0.4.0/agentgauge/__init__.py +3 -0
- agentgauge_harness-0.4.0/agentgauge/__main__.py +5 -0
- agentgauge_harness-0.4.0/agentgauge/_json.py +43 -0
- agentgauge_harness-0.4.0/agentgauge/ab_harness.py +197 -0
- agentgauge_harness-0.4.0/agentgauge/audit.py +334 -0
- agentgauge_harness-0.4.0/agentgauge/cli.py +1122 -0
- agentgauge_harness-0.4.0/agentgauge/client.py +112 -0
- agentgauge_harness-0.4.0/agentgauge/constraints.py +105 -0
- agentgauge_harness-0.4.0/agentgauge/exp1_classifier.py +191 -0
- agentgauge_harness-0.4.0/agentgauge/exp1_doc_density.py +172 -0
- agentgauge_harness-0.4.0/agentgauge/exp1_mirror.py +273 -0
- agentgauge_harness-0.4.0/agentgauge/exp1_tool_def_extractor.py +406 -0
- agentgauge_harness-0.4.0/agentgauge/fixer.py +927 -0
- agentgauge_harness-0.4.0/agentgauge/frontier.py +69 -0
- agentgauge_harness-0.4.0/agentgauge/frozen_protocol.py +131 -0
- agentgauge_harness-0.4.0/agentgauge/harness.py +894 -0
- agentgauge_harness-0.4.0/agentgauge/linter.py +607 -0
- agentgauge_harness-0.4.0/agentgauge/localizer.py +326 -0
- agentgauge_harness-0.4.0/agentgauge/providers.py +326 -0
- agentgauge_harness-0.4.0/agentgauge/q2a_harness.py +87 -0
- agentgauge_harness-0.4.0/agentgauge/report.py +146 -0
- agentgauge_harness-0.4.0/agentgauge/runner.py +108 -0
- agentgauge_harness-0.4.0/agentgauge/scorer.py +947 -0
- agentgauge_harness-0.4.0/agentgauge/tasks.py +48 -0
- agentgauge_harness-0.4.0/agentgauge_pivot_onepager.md +81 -0
- agentgauge_harness-0.4.0/agentready-spec.md +123 -0
- agentgauge_harness-0.4.0/docs/paper/evidence_table.md +221 -0
- agentgauge_harness-0.4.0/docs/paper/latex/README.md +21 -0
- agentgauge_harness-0.4.0/docs/paper/latex/abstract_body.tex +24 -0
- agentgauge_harness-0.4.0/docs/paper/latex/body_content.tex +1318 -0
- agentgauge_harness-0.4.0/docs/paper/latex/main.pdf +0 -0
- agentgauge_harness-0.4.0/docs/paper/latex/main.tex +49 -0
- agentgauge_harness-0.4.0/docs/paper/latex/references.bib +73 -0
- agentgauge_harness-0.4.0/docs/paper/paper.md +881 -0
- agentgauge_harness-0.4.0/docs/paper/repo_triage.md +344 -0
- agentgauge_harness-0.4.0/docs/paper/revision_changelog.md +422 -0
- agentgauge_harness-0.4.0/docs/paper/skeleton.md +118 -0
- agentgauge_harness-0.4.0/docs/paper/threats_to_validity.md +125 -0
- agentgauge_harness-0.4.0/docs/research/exp3_pre_registration.md +267 -0
- agentgauge_harness-0.4.0/docs/research/exp4_regime_map.md +308 -0
- agentgauge_harness-0.4.0/docs/research/frontier_t18_result.md +47 -0
- agentgauge_harness-0.4.0/docs/research/frozen_protocol.md +179 -0
- agentgauge_harness-0.4.0/docs/research/phase1-buyer-and-landscape.md +148 -0
- agentgauge_harness-0.4.0/evals/__init__.py +1 -0
- agentgauge_harness-0.4.0/evals/fixtures/__init__.py +1 -0
- agentgauge_harness-0.4.0/evals/fixtures/exp1_anchor_validation.json +96 -0
- agentgauge_harness-0.4.0/evals/fixtures/exp1_candidate_list.json +5333 -0
- agentgauge_harness-0.4.0/evals/fixtures/exp1_doc_density_scores.json +469 -0
- agentgauge_harness-0.4.0/evals/fixtures/exp1_exclusion_log.json +177 -0
- agentgauge_harness-0.4.0/evals/fixtures/exp1_family_candidates.json +936 -0
- agentgauge_harness-0.4.0/evals/fixtures/exp1_mirrors/AminForou-mcp-gsc.json +439 -0
- agentgauge_harness-0.4.0/evals/fixtures/exp1_mirrors/Dataojitori-nocturne_memory.json +181 -0
- agentgauge_harness-0.4.0/evals/fixtures/exp1_mirrors/LycheeMem-LycheeMem.json +38 -0
- agentgauge_harness-0.4.0/evals/fixtures/exp1_mirrors/aws-iam-mcp.json +647 -0
- agentgauge_harness-0.4.0/evals/fixtures/exp1_mirrors/blazickjp-arxiv-mcp-server.json +66 -0
- agentgauge_harness-0.4.0/evals/fixtures/exp1_mirrors/datalayer-jupyter-mcp-server.json +359 -0
- agentgauge_harness-0.4.0/evals/fixtures/exp1_mirrors/github-mcp.json +663 -0
- agentgauge_harness-0.4.0/evals/fixtures/exp1_mirrors/lucasastorian-llmwiki.json +293 -0
- agentgauge_harness-0.4.0/evals/fixtures/exp1_mirrors/mrexodia-ida-pro-mcp.json +1367 -0
- agentgauge_harness-0.4.0/evals/fixtures/exp1_mirrors/stefanoamorelli-sec-edgar-mcp.json +438 -0
- agentgauge_harness-0.4.0/evals/fixtures/exp1_mirrors/stickerdaniel-linkedin-mcp-server.json +420 -0
- agentgauge_harness-0.4.0/evals/fixtures/exp1_mirrors/taylorwilsdon-google_workspace_mcp.json +4878 -0
- agentgauge_harness-0.4.0/evals/fixtures/exp1_mirrors_manifest.json +118 -0
- agentgauge_harness-0.4.0/evals/fixtures/exp1_pre_registration.json +169 -0
- agentgauge_harness-0.4.0/evals/fixtures/exp1_registry_candidates.json +13781 -0
- agentgauge_harness-0.4.0/evals/fixtures/exp1_server_frame.json +296 -0
- agentgauge_harness-0.4.0/evals/fixtures/exp1_trial_AminForou-mcp-gsc_arm_a.json +296 -0
- agentgauge_harness-0.4.0/evals/fixtures/exp1_trial_AminForou-mcp-gsc_arm_b.json +182 -0
- agentgauge_harness-0.4.0/evals/fixtures/exp1_trial_batch_summary.json +62 -0
- agentgauge_harness-0.4.0/evals/fixtures/exp1_trial_datalayer-jupyter-mcp-server_arm_a.json +296 -0
- agentgauge_harness-0.4.0/evals/fixtures/exp1_trial_datalayer-jupyter-mcp-server_arm_b.json +182 -0
- agentgauge_harness-0.4.0/evals/fixtures/exp1_trial_lucasastorian-llmwiki_arm_a.json +296 -0
- agentgauge_harness-0.4.0/evals/fixtures/exp1_trial_mrexodia-ida-pro-mcp_arm_a.json +296 -0
- agentgauge_harness-0.4.0/evals/fixtures/exp1_trial_stefanoamorelli-sec-edgar-mcp_arm_a.json +296 -0
- agentgauge_harness-0.4.0/evals/fixtures/exp1_trial_stickerdaniel-linkedin-mcp-server_arm_a.json +296 -0
- agentgauge_harness-0.4.0/evals/fixtures/exp1_trial_taylorwilsdon-google_workspace_mcp_arm_a.json +296 -0
- agentgauge_harness-0.4.0/evals/fixtures/exp1_trial_taylorwilsdon-google_workspace_mcp_arm_b.json +182 -0
- agentgauge_harness-0.4.0/evals/fixtures/exp3_ground_truth.json +272 -0
- agentgauge_harness-0.4.0/evals/fixtures/exp3_localizer_graded_result.json +408 -0
- agentgauge_harness-0.4.0/evals/fixtures/exp3_localizer_result.json +383 -0
- agentgauge_harness-0.4.0/evals/fixtures/frontier_t18_step2_raw_calls.json +1 -0
- agentgauge_harness-0.4.0/evals/fixtures/frontier_t18_step2_result.json +1486 -0
- agentgauge_harness-0.4.0/evals/fixtures/p2a_arm_guardb_descriptions.json +50 -0
- agentgauge_harness-0.4.0/evals/fixtures/p2a_f2_retrieval_spec.json +57 -0
- agentgauge_harness-0.4.0/evals/fixtures/p2a_internal_proxy_catalog.py +872 -0
- agentgauge_harness-0.4.0/evals/fixtures/predictive_validity/__init__.py +1 -0
- agentgauge_harness-0.4.0/evals/fixtures/predictive_validity/blind_tasks.py +937 -0
- agentgauge_harness-0.4.0/evals/fixtures/predictive_validity/constraints.py +1934 -0
- agentgauge_harness-0.4.0/evals/fixtures/predictive_validity/manifest.py +271 -0
- agentgauge_harness-0.4.0/evals/fixtures/predictive_validity/phase3_mechanism_results.json +205 -0
- agentgauge_harness-0.4.0/evals/fixtures/predictive_validity/results_raw.json +45714 -0
- agentgauge_harness-0.4.0/evals/fixtures/predictive_validity/results_raw.ndjson +88 -0
- agentgauge_harness-0.4.0/evals/fixtures/predictive_validity/results_raw_INVALID_leaked_tasks.json +1585 -0
- agentgauge_harness-0.4.0/evals/fixtures/predictive_validity/results_raw_INVALID_leaked_tasks.ndjson +14 -0
- agentgauge_harness-0.4.0/evals/fixtures/predictive_validity/results_raw_PHASE2_binary_1trial.json +3652 -0
- agentgauge_harness-0.4.0/evals/fixtures/predictive_validity/results_raw_PHASE2_binary_1trial.ndjson +18 -0
- agentgauge_harness-0.4.0/evals/fixtures/predictive_validity/schema_consistency_results.json +8261 -0
- agentgauge_harness-0.4.0/evals/fixtures/q3_arm_f_body_descriptions.json +14 -0
- agentgauge_harness-0.4.0/evals/fixtures/q3_arm_f_doc_descriptions.json +14 -0
- agentgauge_harness-0.4.0/evals/fixtures/q3_catalog.py +204 -0
- agentgauge_harness-0.4.0/evals/fixtures/q4_arm_f_body_scoped_descriptions.json +14 -0
- agentgauge_harness-0.4.0/evals/fixtures/q4_arm_f_doc_scoped_descriptions.json +14 -0
- agentgauge_harness-0.4.0/evals/fixtures/q5_arm_f_doc_guarded_descriptions.json +14 -0
- agentgauge_harness-0.4.0/evals/fixtures/q6_arm_f_doc_guarded_descriptions.json +25 -0
- agentgauge_harness-0.4.0/evals/fixtures/q6_catalog.py +405 -0
- agentgauge_harness-0.4.0/evals/fixtures/rw1_arm_guardb_descriptions.json +23 -0
- agentgauge_harness-0.4.0/evals/fixtures/rw1_github_catalog.py +678 -0
- agentgauge_harness-0.4.0/evals/fixtures/rw2_arm_guardb_descriptions.json +31 -0
- agentgauge_harness-0.4.0/evals/fixtures/rw2_aws_iam_catalog.py +1125 -0
- agentgauge_harness-0.4.0/evals/fixtures/t17_tasks.py +69 -0
- agentgauge_harness-0.4.0/evals/fixtures/t18_arm_f_descriptions.json +62 -0
- agentgauge_harness-0.4.0/evals/fixtures/t18_arm_f_q2b_descriptions.json +62 -0
- agentgauge_harness-0.4.0/evals/fixtures/t18_catalog.py +521 -0
- agentgauge_harness-0.4.0/evals/fixtures/ty2_tasks.py +292 -0
- agentgauge_harness-0.4.0/evals/fixtures/ty_tasks.py +191 -0
- agentgauge_harness-0.4.0/evals/fixtures/v2_1_cross_model_validation.json +44 -0
- agentgauge_harness-0.4.0/evals/fixtures/v2_1_false_alarm_new_estimator.json +94 -0
- agentgauge_harness-0.4.0/evals/fixtures/v2_1_linter_recall_fix.json +470 -0
- agentgauge_harness-0.4.0/evals/fixtures/v2_1_llm_linter_baseline.json +765 -0
- agentgauge_harness-0.4.0/evals/fixtures/v2_1_mde_ablation.json +101 -0
- agentgauge_harness-0.4.0/evals/fixtures/v2_1_severity_gate.json +16 -0
- agentgauge_harness-0.4.0/evals/fixtures/v2_2_causal_chain.json +572 -0
- agentgauge_harness-0.4.0/evals/fixtures/v2_2_causal_chain_multimodel.json +1718 -0
- agentgauge_harness-0.4.0/evals/fixtures/v2_2_cross_model_full.json +44 -0
- agentgauge_harness-0.4.0/evals/fixtures/v2_2_cross_model_pooled.json +44 -0
- agentgauge_harness-0.4.0/evals/fixtures/v2_2_few_clusters_correction.json +125 -0
- agentgauge_harness-0.4.0/evals/fixtures/v2_2_optimal_allocation.json +247 -0
- agentgauge_harness-0.4.0/evals/fixtures/v2_3_advisory_audit.json +586 -0
- agentgauge_harness-0.4.0/evals/fixtures/v2_3_corrected_advisory_effect.json +407 -0
- agentgauge_harness-0.4.0/evals/fixtures/v2_4_blocking_remeasurement.json +1142 -0
- agentgauge_harness-0.4.0/evals/fixtures/v2_4_corpus/__init__.py +1 -0
- agentgauge_harness-0.4.0/evals/fixtures/v2_4_corpus/aws_s3_NOTES.md +88 -0
- agentgauge_harness-0.4.0/evals/fixtures/v2_4_corpus/aws_s3_fixture.py +282 -0
- agentgauge_harness-0.4.0/evals/fixtures/v2_4_corpus/docker_containers_NOTES.md +82 -0
- agentgauge_harness-0.4.0/evals/fixtures/v2_4_corpus/docker_containers_fixture.py +299 -0
- agentgauge_harness-0.4.0/evals/fixtures/v2_4_corpus/gcal_NOTES.md +77 -0
- agentgauge_harness-0.4.0/evals/fixtures/v2_4_corpus/gcal_fixture.py +239 -0
- agentgauge_harness-0.4.0/evals/fixtures/v2_4_corpus/github_issues_NOTES.md +52 -0
- agentgauge_harness-0.4.0/evals/fixtures/v2_4_corpus/github_issues_fixture.py +293 -0
- agentgauge_harness-0.4.0/evals/fixtures/v2_4_corpus/jira_issues_NOTES.md +64 -0
- agentgauge_harness-0.4.0/evals/fixtures/v2_4_corpus/jira_issues_fixture.py +292 -0
- agentgauge_harness-0.4.0/evals/fixtures/v2_4_corpus/k8s_workloads_NOTES.md +87 -0
- agentgauge_harness-0.4.0/evals/fixtures/v2_4_corpus/k8s_workloads_fixture.py +301 -0
- agentgauge_harness-0.4.0/evals/fixtures/v2_4_corpus/slack_messaging_NOTES.md +94 -0
- agentgauge_harness-0.4.0/evals/fixtures/v2_4_corpus/slack_messaging_fixture.py +197 -0
- agentgauge_harness-0.4.0/evals/fixtures/v2_4_corpus/spotify_playlists_NOTES.md +76 -0
- agentgauge_harness-0.4.0/evals/fixtures/v2_4_corpus/spotify_playlists_fixture.py +232 -0
- agentgauge_harness-0.4.0/evals/fixtures/v2_4_corpus/stripe_payments_NOTES.md +76 -0
- agentgauge_harness-0.4.0/evals/fixtures/v2_4_corpus/stripe_payments_fixture.py +208 -0
- agentgauge_harness-0.4.0/evals/fixtures/v2_4_corpus/twilio_messaging_NOTES.md +64 -0
- agentgauge_harness-0.4.0/evals/fixtures/v2_4_corpus/twilio_messaging_fixture.py +274 -0
- agentgauge_harness-0.4.0/evals/fixtures/v2_5_argument_degradation_live.jsonl +1518 -0
- agentgauge_harness-0.4.0/evals/fixtures/v2_5_argument_degradation_summary.json +50 -0
- agentgauge_harness-0.4.0/evals/fixtures/v2_defect_injection_results.json +1961 -0
- agentgauge_harness-0.4.0/evals/fixtures/v2_false_alarm_determinism.json +279 -0
- agentgauge_harness-0.4.0/evals/fixtures/v2_lint_baseline.json +20946 -0
- agentgauge_harness-0.4.0/evals/fixtures/v2_mde_continuous_crosscheck.json +50 -0
- agentgauge_harness-0.4.0/evals/fixtures/v2_mde_table.json +114 -0
- agentgauge_harness-0.4.0/evals/fixtures/v2_tool_definitions.json +20856 -0
- agentgauge_harness-0.4.0/evals/fixtures/v2_variance_structure.json +62 -0
- agentgauge_harness-0.4.0/examples/aws_s3_server.py +120 -0
- agentgauge_harness-0.4.0/examples/aws_s3_server_fixed.py +162 -0
- agentgauge_harness-0.4.0/examples/call_constraints_server.py +153 -0
- agentgauge_harness-0.4.0/examples/call_constraints_server_fixed.py +153 -0
- agentgauge_harness-0.4.0/examples/call_constraints_server_oracle.py +184 -0
- agentgauge_harness-0.4.0/examples/call_constraints_v2_server.py +137 -0
- agentgauge_harness-0.4.0/examples/call_constraints_v2_server_fixed.py +137 -0
- agentgauge_harness-0.4.0/examples/call_constraints_v2_server_oracle.py +180 -0
- agentgauge_harness-0.4.0/examples/confusable_server.py +266 -0
- agentgauge_harness-0.4.0/examples/confusable_server_fixed.py +266 -0
- agentgauge_harness-0.4.0/examples/confusable_server_oracle.py +288 -0
- agentgauge_harness-0.4.0/examples/docker_containers_server.py +120 -0
- agentgauge_harness-0.4.0/examples/docker_containers_server_fixed.py +168 -0
- agentgauge_harness-0.4.0/examples/echo_server.py +112 -0
- agentgauge_harness-0.4.0/examples/echo_server_fixed.py +114 -0
- agentgauge_harness-0.4.0/examples/echo_server_fixed_dqonly.py +112 -0
- agentgauge_harness-0.4.0/examples/exp1_AminForou_mcp_gsc_mirror.py +437 -0
- agentgauge_harness-0.4.0/examples/exp1_AminForou_mcp_gsc_mirror_oracle.py +437 -0
- agentgauge_harness-0.4.0/examples/exp1_Dataojitori_nocturne_memory_mirror.py +213 -0
- agentgauge_harness-0.4.0/examples/exp1_LycheeMem_LycheeMem_mirror.py +89 -0
- agentgauge_harness-0.4.0/examples/exp1_blazickjp_arxiv_mcp_server_mirror.py +113 -0
- agentgauge_harness-0.4.0/examples/exp1_datalayer_jupyter_mcp_server_mirror.py +369 -0
- agentgauge_harness-0.4.0/examples/exp1_datalayer_jupyter_mcp_server_mirror_oracle.py +369 -0
- agentgauge_harness-0.4.0/examples/exp1_dataojitori_nocturne_memory_mirror_fixed.py +213 -0
- agentgauge_harness-0.4.0/examples/exp1_lucasastorian_llmwiki_mirror.py +307 -0
- agentgauge_harness-0.4.0/examples/exp1_mrexodia_ida_pro_mcp_mirror.py +1279 -0
- agentgauge_harness-0.4.0/examples/exp1_stefanoamorelli_sec_edgar_mcp_mirror.py +441 -0
- agentgauge_harness-0.4.0/examples/exp1_stickerdaniel_linkedin_mcp_server_mirror.py +419 -0
- agentgauge_harness-0.4.0/examples/exp1_taylorwilsdon_google_workspace_mcp_mirror.py +4145 -0
- agentgauge_harness-0.4.0/examples/exp1_taylorwilsdon_google_workspace_mcp_mirror_oracle.py +4145 -0
- agentgauge_harness-0.4.0/examples/gcal_server.py +118 -0
- agentgauge_harness-0.4.0/examples/gcal_server_fixed.py +150 -0
- agentgauge_harness-0.4.0/examples/github_issues_server.py +119 -0
- agentgauge_harness-0.4.0/examples/github_issues_server_fixed.py +155 -0
- agentgauge_harness-0.4.0/examples/grounded_server.py +183 -0
- agentgauge_harness-0.4.0/examples/grounded_server_fixed.py +183 -0
- agentgauge_harness-0.4.0/examples/grounded_server_oracle.py +164 -0
- agentgauge_harness-0.4.0/examples/jira_issues_server.py +117 -0
- agentgauge_harness-0.4.0/examples/jira_issues_server_fixed.py +152 -0
- agentgauge_harness-0.4.0/examples/k8s_workloads_server.py +124 -0
- agentgauge_harness-0.4.0/examples/k8s_workloads_server_fixed.py +172 -0
- agentgauge_harness-0.4.0/examples/mediocre_server.py +253 -0
- agentgauge_harness-0.4.0/examples/mediocre_server_fixed.py +258 -0
- agentgauge_harness-0.4.0/examples/p2a_arm_a.py +56 -0
- agentgauge_harness-0.4.0/examples/p2a_arm_guardb.py +75 -0
- agentgauge_harness-0.4.0/examples/p2a_arm_oracle.py +58 -0
- agentgauge_harness-0.4.0/examples/p2a_internal_proxy_mirror.py +499 -0
- agentgauge_harness-0.4.0/examples/q3_arm_a.py +57 -0
- agentgauge_harness-0.4.0/examples/q3_arm_f_body.py +69 -0
- agentgauge_harness-0.4.0/examples/q3_arm_f_doc.py +69 -0
- agentgauge_harness-0.4.0/examples/q3_arm_o.py +56 -0
- agentgauge_harness-0.4.0/examples/q3_real_server.py +223 -0
- agentgauge_harness-0.4.0/examples/q3_real_server_fixed.py +223 -0
- agentgauge_harness-0.4.0/examples/q3_real_server_fixed_dqonly.py +223 -0
- agentgauge_harness-0.4.0/examples/q4_arm_f_body_scoped.py +70 -0
- agentgauge_harness-0.4.0/examples/q4_arm_f_doc_scoped.py +70 -0
- agentgauge_harness-0.4.0/examples/q5_arm_f_doc_guarded.py +71 -0
- agentgauge_harness-0.4.0/examples/q6_arm_a.py +58 -0
- agentgauge_harness-0.4.0/examples/q6_arm_f_doc_guarded.py +72 -0
- agentgauge_harness-0.4.0/examples/q6_real_server.py +423 -0
- agentgauge_harness-0.4.0/examples/rw1_arm_a.py +57 -0
- agentgauge_harness-0.4.0/examples/rw1_arm_guardb.py +74 -0
- agentgauge_harness-0.4.0/examples/rw1_arm_oracle.py +56 -0
- agentgauge_harness-0.4.0/examples/rw1_github_mirror.py +402 -0
- agentgauge_harness-0.4.0/examples/rw2_arm_a.py +52 -0
- agentgauge_harness-0.4.0/examples/rw2_arm_guardb.py +69 -0
- agentgauge_harness-0.4.0/examples/rw2_aws_iam_mirror.py +484 -0
- agentgauge_harness-0.4.0/examples/slack_messaging_server.py +116 -0
- agentgauge_harness-0.4.0/examples/slack_messaging_server_fixed.py +149 -0
- agentgauge_harness-0.4.0/examples/spotify_playlists_server.py +118 -0
- agentgauge_harness-0.4.0/examples/spotify_playlists_server_fixed.py +154 -0
- agentgauge_harness-0.4.0/examples/stripe_payments_server.py +119 -0
- agentgauge_harness-0.4.0/examples/stripe_payments_server_fixed.py +158 -0
- agentgauge_harness-0.4.0/examples/t18_fixer_server.py +69 -0
- agentgauge_harness-0.4.0/examples/t18_oracle_server.py +67 -0
- agentgauge_harness-0.4.0/examples/t18_q2b_server.py +81 -0
- agentgauge_harness-0.4.0/examples/t18_vague_server.py +66 -0
- agentgauge_harness-0.4.0/examples/twilio_messaging_server.py +119 -0
- agentgauge_harness-0.4.0/examples/twilio_messaging_server_fixed.py +169 -0
- agentgauge_harness-0.4.0/p2a_spec.md +133 -0
- agentgauge_harness-0.4.0/paper_framing_options.md +65 -0
- agentgauge_harness-0.4.0/pyproject.toml +76 -0
- agentgauge_harness-0.4.0/scripts/Dockerfile.agentgauge-agent +35 -0
- agentgauge_harness-0.4.0/scripts/_mutated_stdio_server.py +106 -0
- agentgauge_harness-0.4.0/scripts/agentgauge-agent-service.yaml +44 -0
- agentgauge_harness-0.4.0/scripts/agentgauge-judge-service.yaml +53 -0
- agentgauge_harness-0.4.0/scripts/build_fixed_fixtures.py +127 -0
- agentgauge_harness-0.4.0/scripts/build_fixed_fixtures_v2.py +97 -0
- agentgauge_harness-0.4.0/scripts/check_paper_latex_sync.py +68 -0
- agentgauge_harness-0.4.0/scripts/exp1_build_mirrors.py +278 -0
- agentgauge_harness-0.4.0/scripts/exp1_discover_registry.py +141 -0
- agentgauge_harness-0.4.0/scripts/exp1_discover_servers.py +255 -0
- agentgauge_harness-0.4.0/scripts/exp1_generate_mirror_server.py +168 -0
- agentgauge_harness-0.4.0/scripts/exp1_identify_families.py +102 -0
- agentgauge_harness-0.4.0/scripts/exp1_run_remaining_trials.py +374 -0
- agentgauge_harness-0.4.0/scripts/exp1_run_trial.py +294 -0
- agentgauge_harness-0.4.0/scripts/exp1_score_doc_density.py +158 -0
- agentgauge_harness-0.4.0/scripts/exp1_validate_anchors.py +168 -0
- agentgauge_harness-0.4.0/scripts/exp3_run_localizer.py +346 -0
- agentgauge_harness-0.4.0/scripts/generate_arm_f_descriptions.py +61 -0
- agentgauge_harness-0.4.0/scripts/generate_arm_f_descriptions_q2b.py +70 -0
- agentgauge_harness-0.4.0/scripts/generate_q3_descriptions.py +150 -0
- agentgauge_harness-0.4.0/scripts/generate_q4_descriptions.py +189 -0
- agentgauge_harness-0.4.0/scripts/generate_q5_descriptions.py +173 -0
- agentgauge_harness-0.4.0/scripts/generate_q6_descriptions.py +206 -0
- agentgauge_harness-0.4.0/scripts/mde_grid_v2_5.py +38 -0
- agentgauge_harness-0.4.0/scripts/p2a_f2_retrieval.py +452 -0
- agentgauge_harness-0.4.0/scripts/p2a_frontier_gate.py +337 -0
- agentgauge_harness-0.4.0/scripts/p2a_phase1_generate.py +188 -0
- agentgauge_harness-0.4.0/scripts/p2a_phase2_ab.py +600 -0
- agentgauge_harness-0.4.0/scripts/phase3_mechanism_test.py +228 -0
- agentgauge_harness-0.4.0/scripts/predictive_validity_analysis.py +536 -0
- agentgauge_harness-0.4.0/scripts/predictive_validity_study.py +440 -0
- agentgauge_harness-0.4.0/scripts/run_ab_experiment.py +207 -0
- agentgauge_harness-0.4.0/scripts/run_all_phase3_expansion_builds.py +82 -0
- agentgauge_harness-0.4.0/scripts/run_build_fixed_fixtures_via_gcp.py +30 -0
- agentgauge_harness-0.4.0/scripts/run_frontier_t18.py +452 -0
- agentgauge_harness-0.4.0/scripts/run_predictive_validity_via_gcp.py +32 -0
- agentgauge_harness-0.4.0/scripts/run_q2a_three_arm.py +424 -0
- agentgauge_harness-0.4.0/scripts/run_q2b_three_arm.py +404 -0
- agentgauge_harness-0.4.0/scripts/run_q3_four_arm.py +408 -0
- agentgauge_harness-0.4.0/scripts/run_q4_four_arm.py +461 -0
- agentgauge_harness-0.4.0/scripts/run_q5_four_arm.py +450 -0
- agentgauge_harness-0.4.0/scripts/run_q6_regression.py +508 -0
- agentgauge_harness-0.4.0/scripts/run_schema_consistency.py +98 -0
- agentgauge_harness-0.4.0/scripts/run_t17_oracle_ab.py +313 -0
- agentgauge_harness-0.4.0/scripts/run_t18_oracle_ab.py +333 -0
- agentgauge_harness-0.4.0/scripts/run_tx_experiment.py +392 -0
- agentgauge_harness-0.4.0/scripts/run_ty2_oracle_ab.py +420 -0
- agentgauge_harness-0.4.0/scripts/run_ty_oracle_ab.py +370 -0
- agentgauge_harness-0.4.0/scripts/rw1_part1_discoverability.py +235 -0
- agentgauge_harness-0.4.0/scripts/rw1_phase1_generate.py +194 -0
- agentgauge_harness-0.4.0/scripts/rw1_phase2_ab.py +431 -0
- agentgauge_harness-0.4.0/scripts/rw2_phase1_generate.py +213 -0
- agentgauge_harness-0.4.0/scripts/rw2_phase2_ab.py +469 -0
- agentgauge_harness-0.4.0/scripts/schema_consistency_checker.py +172 -0
- agentgauge_harness-0.4.0/scripts/v2_1_cross_model_validation.py +238 -0
- agentgauge_harness-0.4.0/scripts/v2_1_false_alarm_new_estimator.py +121 -0
- agentgauge_harness-0.4.0/scripts/v2_1_linter_recall_fix.py +176 -0
- agentgauge_harness-0.4.0/scripts/v2_1_llm_linter_baseline.py +220 -0
- agentgauge_harness-0.4.0/scripts/v2_1_mde_ablation.py +160 -0
- agentgauge_harness-0.4.0/scripts/v2_1_severity_gate_measurement.py +81 -0
- agentgauge_harness-0.4.0/scripts/v2_2_causal_chain.py +273 -0
- agentgauge_harness-0.4.0/scripts/v2_2_causal_chain_multimodel.py +258 -0
- agentgauge_harness-0.4.0/scripts/v2_2_cross_model_full.py +155 -0
- agentgauge_harness-0.4.0/scripts/v2_2_cross_model_pooled.py +183 -0
- agentgauge_harness-0.4.0/scripts/v2_2_few_clusters_correction.py +122 -0
- agentgauge_harness-0.4.0/scripts/v2_2_optimal_allocation.py +170 -0
- agentgauge_harness-0.4.0/scripts/v2_3_advisory_audit.py +155 -0
- agentgauge_harness-0.4.0/scripts/v2_3_corrected_advisory_effect.py +135 -0
- agentgauge_harness-0.4.0/scripts/v2_4_blocking_remeasurement.py +235 -0
- agentgauge_harness-0.4.0/scripts/v2_5_argument_degradation_live.py +348 -0
- agentgauge_harness-0.4.0/scripts/v2_5_argument_degradation_live_gcp.py +97 -0
- agentgauge_harness-0.4.0/scripts/v2_defect_injector.py +330 -0
- agentgauge_harness-0.4.0/scripts/v2_extract_tool_definitions.py +72 -0
- agentgauge_harness-0.4.0/scripts/v2_false_alarm_and_determinism.py +110 -0
- agentgauge_harness-0.4.0/scripts/v2_mde_continuous_crosscheck.py +118 -0
- agentgauge_harness-0.4.0/scripts/v2_mde_table.py +79 -0
- agentgauge_harness-0.4.0/scripts/v2_variance_structure.py +260 -0
- agentgauge_harness-0.4.0/scripts/verify.sh +41 -0
- agentgauge_harness-0.4.0/spec.md +91 -0
- agentgauge_harness-0.4.0/spec_ty2.md +326 -0
- agentgauge_harness-0.4.0/tests/__init__.py +0 -0
- agentgauge_harness-0.4.0/tests/test_ab_harness.py +250 -0
- agentgauge_harness-0.4.0/tests/test_audit.py +352 -0
- agentgauge_harness-0.4.0/tests/test_cli.py +295 -0
- agentgauge_harness-0.4.0/tests/test_client.py +65 -0
- agentgauge_harness-0.4.0/tests/test_constraints.py +66 -0
- agentgauge_harness-0.4.0/tests/test_discoverability.py +297 -0
- agentgauge_harness-0.4.0/tests/test_docs_manifest.py +209 -0
- agentgauge_harness-0.4.0/tests/test_error_legibility.py +294 -0
- agentgauge_harness-0.4.0/tests/test_exp1_classifier.py +385 -0
- agentgauge_harness-0.4.0/tests/test_exp1_doc_density.py +169 -0
- agentgauge_harness-0.4.0/tests/test_exp1_generate_mirror_server.py +77 -0
- agentgauge_harness-0.4.0/tests/test_exp1_mirror.py +561 -0
- agentgauge_harness-0.4.0/tests/test_exp1_pre_reg.py +70 -0
- agentgauge_harness-0.4.0/tests/test_exp1_tool_def_extractor.py +490 -0
- agentgauge_harness-0.4.0/tests/test_fixer.py +1494 -0
- agentgauge_harness-0.4.0/tests/test_frontier.py +398 -0
- agentgauge_harness-0.4.0/tests/test_frozen_protocol.py +129 -0
- agentgauge_harness-0.4.0/tests/test_harness.py +593 -0
- agentgauge_harness-0.4.0/tests/test_json.py +144 -0
- agentgauge_harness-0.4.0/tests/test_linter.py +365 -0
- agentgauge_harness-0.4.0/tests/test_localizer.py +306 -0
- agentgauge_harness-0.4.0/tests/test_p2a_catalog.py +178 -0
- agentgauge_harness-0.4.0/tests/test_predictive_validity_analysis.py +477 -0
- agentgauge_harness-0.4.0/tests/test_q2a_recovery.py +230 -0
- agentgauge_harness-0.4.0/tests/test_q3.py +290 -0
- agentgauge_harness-0.4.0/tests/test_q4_scoped.py +291 -0
- agentgauge_harness-0.4.0/tests/test_q5_guarded.py +192 -0
- agentgauge_harness-0.4.0/tests/test_q6_fixture.py +305 -0
- agentgauge_harness-0.4.0/tests/test_report.py +99 -0
- agentgauge_harness-0.4.0/tests/test_robustness.py +302 -0
- agentgauge_harness-0.4.0/tests/test_runner.py +307 -0
- agentgauge_harness-0.4.0/tests/test_rw1_github.py +436 -0
- agentgauge_harness-0.4.0/tests/test_rw2_aws_iam.py +440 -0
- agentgauge_harness-0.4.0/tests/test_scorer.py +222 -0
- agentgauge_harness-0.4.0/tests/test_t17_fixture.py +140 -0
- agentgauge_harness-0.4.0/tests/test_t18_fixture.py +141 -0
- agentgauge_harness-0.4.0/tests/test_tasks.py +91 -0
- agentgauge_harness-0.4.0/tests/test_ty2_fixture.py +285 -0
- agentgauge_harness-0.4.0/tests/test_ty_fixture.py +182 -0
- agentgauge_harness-0.4.0/tests/test_ux1.py +435 -0
- agentgauge_harness-0.4.0/uv.lock +1269 -0
|
@@ -0,0 +1,23 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: executor
|
|
3
|
+
description: Use this subagent to implement specific, well-scoped code changes. Invoke with a clear task description, target files, and any constraints. The subagent writes code, runs verification, and reports back. Do NOT invoke for planning, architecture decisions, or open-ended exploration.
|
|
4
|
+
tools: Read, Write, Edit, Glob, Grep, Bash
|
|
5
|
+
model: sonnet
|
|
6
|
+
---
|
|
7
|
+
|
|
8
|
+
You are a senior software engineer focused on implementation. You receive a scoped task from the orchestrator and implement it cleanly.
|
|
9
|
+
|
|
10
|
+
Rules:
|
|
11
|
+
- Stay strictly within the scope assigned. If you discover the task is bigger than described, stop and report rather than expanding scope.
|
|
12
|
+
- Add type hints on every function, docstrings on non-trivial ones, and at least one unit test for any new non-trivial function.
|
|
13
|
+
- Run any verification commands the orchestrator specifies. If a command fails, fix and re-run - but if the same check fails 3 times in a row, stop and report with the error output.
|
|
14
|
+
- Commit in small, conventionally-named commits (feat:, fix:, test:, chore:, etc.). One concept per commit.
|
|
15
|
+
- Do NOT install new dependencies without asking. Surface the need in your report.
|
|
16
|
+
- Do NOT touch files outside the project directory.
|
|
17
|
+
|
|
18
|
+
Final report format:
|
|
19
|
+
- One-line summary of what was done
|
|
20
|
+
- List of commit messages
|
|
21
|
+
- Verification results (pass/fail per check)
|
|
22
|
+
- Any deviations from assigned scope, with reasoning
|
|
23
|
+
- Anything the orchestrator should know before the next task
|
|
@@ -0,0 +1,27 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: verifier
|
|
3
|
+
description: Use this subagent to run a project's verification commands (tests, linters, type checkers) and report structured pass/fail results. Invoke after any code-writing pass. Does not write code - only runs commands and reports.
|
|
4
|
+
tools: Read, Glob, Grep, Bash
|
|
5
|
+
model: sonnet
|
|
6
|
+
---
|
|
7
|
+
|
|
8
|
+
You are a verification specialist. You run commands and report results. You do not write code, fix bugs, or interpret failures - the orchestrator handles that.
|
|
9
|
+
|
|
10
|
+
For each command in the list you are given:
|
|
11
|
+
1. Run it in the project root.
|
|
12
|
+
2. Capture exit code, the last 100 lines of combined stdout+stderr, and wall-clock duration.
|
|
13
|
+
3. Classify as pass (exit 0) or fail (non-zero).
|
|
14
|
+
|
|
15
|
+
Always output in this format:
|
|
16
|
+
|
|
17
|
+
VERIFICATION RESULTS
|
|
18
|
+
====================
|
|
19
|
+
[pass/fail] <name>: <command> (<duration>s)
|
|
20
|
+
Summary: <one-line summary>
|
|
21
|
+
Detail: <relevant excerpt of failure output, omit if passed>
|
|
22
|
+
|
|
23
|
+
[pass/fail] ...
|
|
24
|
+
|
|
25
|
+
OVERALL: <pass | fail>
|
|
26
|
+
|
|
27
|
+
You have no Write or Edit tools. Do not attempt to modify any files.
|
|
@@ -0,0 +1,126 @@
|
|
|
1
|
+
# AgentGauge — Autonomous Run Operator Prompt
|
|
2
|
+
|
|
3
|
+
You are a Claude Code cloud scheduled task running against the `agentgauge` repository.
|
|
4
|
+
|
|
5
|
+
## Merge policy
|
|
6
|
+
|
|
7
|
+
**All PRs opened by this loop are DRAFT.** Never open a ready-to-merge PR autonomously.
|
|
8
|
+
Never call `gh pr merge`. The human reviews and merges every PR.
|
|
9
|
+
|
|
10
|
+
The following categories of task especially require human judgment beyond what a green test suite
|
|
11
|
+
can verify — but note that DRAFT is the rule for ALL tasks, not just these:
|
|
12
|
+
|
|
13
|
+
1. **Touches the LLM judge** — changes to rubric prompts, scoring logic, calibration constants,
|
|
14
|
+
judge model selection, or blending weights. A green CI means the mock honored the contract,
|
|
15
|
+
not that the judge produces better scores on real inputs.
|
|
16
|
+
|
|
17
|
+
2. **Generates fixes or actions against real servers** — any task that calls a live MCP server,
|
|
18
|
+
a real Ollama instance, or any external API as part of its primary operation.
|
|
19
|
+
|
|
20
|
+
3. **Depends on real-model or real-network behavior** — tasks where the acceptance criteria
|
|
21
|
+
require measuring calibration, validating score bands, or comparing outputs across judge runs.
|
|
22
|
+
These require human review of measured results, not just a green test suite.
|
|
23
|
+
|
|
24
|
+
**The test suite is necessary but not sufficient.** Passing mocks prove the code path runs,
|
|
25
|
+
not that real-model calibration or real-server behavior is correct. All merges require human review.
|
|
26
|
+
|
|
27
|
+
## Before you do anything
|
|
28
|
+
|
|
29
|
+
1. Read `CLAUDE.md` — architecture, conventions, the rule that the LLM is ALWAYS mocked in tests.
|
|
30
|
+
2. Read `AUTONOMY.md` — the hard rules you must never violate.
|
|
31
|
+
3. Read `TASKS.md` — the backlog.
|
|
32
|
+
|
|
33
|
+
## Pick your task
|
|
34
|
+
|
|
35
|
+
Select the **single highest-priority TODO item** in TASKS.md that has clear, testable acceptance
|
|
36
|
+
criteria.
|
|
37
|
+
|
|
38
|
+
**Before starting any work**, check whether an open `claude/*` PR already addresses that task:
|
|
39
|
+
|
|
40
|
+
```bash
|
|
41
|
+
gh pr list --state open --search "<task-id>" --json number,title,headRefName
|
|
42
|
+
```
|
|
43
|
+
|
|
44
|
+
For example, for T3:
|
|
45
|
+
```bash
|
|
46
|
+
gh pr list --state open --search "T3" --json number,title,headRefName
|
|
47
|
+
```
|
|
48
|
+
|
|
49
|
+
- If a matching open PR exists → **skip that item** and pick the next eligible TODO.
|
|
50
|
+
- If no next eligible TODO exists → stop here: write a short `BLOCKER.md` at repo root
|
|
51
|
+
explaining that all TODO items already have open PRs, commit it to a
|
|
52
|
+
`claude/blocker-<date>` branch, open a DRAFT PR titled "Blocker report: all TODOs in
|
|
53
|
+
review", and exit. **Never open a second PR for a task that already has one open.**
|
|
54
|
+
|
|
55
|
+
If no item qualifies at all (all are ambiguous, blocked, or already in review), **before writing
|
|
56
|
+
any BLOCKER.md**, check whether an open `claude/blocker-*` PR already exists:
|
|
57
|
+
|
|
58
|
+
```bash
|
|
59
|
+
gh pr list --state open --search "Blocker report" --json number,title,headRefName
|
|
60
|
+
```
|
|
61
|
+
|
|
62
|
+
- If an open blocker PR already exists → **do not open another one.** Instead, comment on the
|
|
63
|
+
existing PR with today's date and the current board state (e.g. `gh pr comment <number> --body
|
|
64
|
+
"Re-check <date>: board unchanged — TODO still empty, FUTURE/DEFERRED items still blocked."`),
|
|
65
|
+
then exit. One open blocker PR is enough; duplicate idle-path PRs add noise.
|
|
66
|
+
- If no open blocker PR exists → write `BLOCKER.md`, commit to `claude/blocker-<date>`, open a
|
|
67
|
+
DRAFT PR titled "Blocker report: all TODOs in review", and exit.
|
|
68
|
+
|
|
69
|
+
Do not invent work.
|
|
70
|
+
|
|
71
|
+
## Implement
|
|
72
|
+
|
|
73
|
+
Use the orchestrator → executor → verifier pattern:
|
|
74
|
+
- Orchestrator (you): plan the change, break into steps, verify the plan against acceptance criteria.
|
|
75
|
+
- Executor subagent: implement the code changes only.
|
|
76
|
+
- Verifier subagent: confirm the change is correct and complete.
|
|
77
|
+
|
|
78
|
+
Branch name: `claude/<kebab-task-name>` (e.g., `claude/task-generator`).
|
|
79
|
+
|
|
80
|
+
## Definition of done
|
|
81
|
+
|
|
82
|
+
Run `./scripts/verify.sh`. It exits 0 only if:
|
|
83
|
+
- `ruff check` passes
|
|
84
|
+
- `ruff format --check` passes
|
|
85
|
+
- `mypy` passes (if configured)
|
|
86
|
+
- All tests pass — no tests removed, no mocks weakened
|
|
87
|
+
|
|
88
|
+
## Commit and PR
|
|
89
|
+
|
|
90
|
+
If `verify.sh` exits 0:
|
|
91
|
+
|
|
92
|
+
### Step 1 — commit
|
|
93
|
+
|
|
94
|
+
- Commit with conventional-commit message: `feat(scope): description`
|
|
95
|
+
- **Do NOT include claude.ai session URLs in commit bodies.**
|
|
96
|
+
- Push branch: `git push origin claude/<task-name>`
|
|
97
|
+
- In TASKS.md on your branch, move the item from TODO to IN-REVIEW.
|
|
98
|
+
- Commit that TASKS.md update as a separate `chore: move <task> to IN-REVIEW` commit.
|
|
99
|
+
|
|
100
|
+
### Step 2 — open PR
|
|
101
|
+
|
|
102
|
+
**All PRs are DRAFT.** Open every PR as draft regardless of task type:
|
|
103
|
+
|
|
104
|
+
```bash
|
|
105
|
+
gh pr create --draft --title "..." --body "..."
|
|
106
|
+
```
|
|
107
|
+
|
|
108
|
+
Do **not** call `gh pr merge` at all. Do **not** push directly to main.
|
|
109
|
+
Report the PR link. The human will review and merge.
|
|
110
|
+
|
|
111
|
+
If `verify.sh` does not exit 0:
|
|
112
|
+
- Do NOT commit
|
|
113
|
+
- Do NOT open a PR
|
|
114
|
+
- Write a blocker summary explaining what failed and why
|
|
115
|
+
- Exit
|
|
116
|
+
|
|
117
|
+
## End of run report
|
|
118
|
+
|
|
119
|
+
Always end your run with a brief report (5-10 lines):
|
|
120
|
+
- Task selected
|
|
121
|
+
- Open-PR dedup check result (any skipped tasks and why)
|
|
122
|
+
- What was implemented
|
|
123
|
+
- `verify.sh` result (exit code + any relevant failure lines)
|
|
124
|
+
- PR link (if created)
|
|
125
|
+
- PR opened as DRAFT (confirm link)
|
|
126
|
+
- Any blockers or caveats
|
|
@@ -0,0 +1,88 @@
|
|
|
1
|
+
name: CI
|
|
2
|
+
|
|
3
|
+
on:
|
|
4
|
+
push:
|
|
5
|
+
branches: [main]
|
|
6
|
+
pull_request:
|
|
7
|
+
branches: [main]
|
|
8
|
+
|
|
9
|
+
concurrency:
|
|
10
|
+
group: ${{ github.workflow }}-${{ github.ref }}
|
|
11
|
+
cancel-in-progress: true
|
|
12
|
+
|
|
13
|
+
jobs:
|
|
14
|
+
verify:
|
|
15
|
+
runs-on: ubuntu-latest
|
|
16
|
+
steps:
|
|
17
|
+
- uses: actions/checkout@v4
|
|
18
|
+
|
|
19
|
+
- uses: actions/setup-python@v5
|
|
20
|
+
with:
|
|
21
|
+
python-version: "3.11"
|
|
22
|
+
|
|
23
|
+
- name: Install uv
|
|
24
|
+
run: pip install uv --quiet
|
|
25
|
+
|
|
26
|
+
- name: Run verify.sh
|
|
27
|
+
run: bash scripts/verify.sh
|
|
28
|
+
|
|
29
|
+
hygiene:
|
|
30
|
+
runs-on: ubuntu-latest
|
|
31
|
+
steps:
|
|
32
|
+
- uses: actions/checkout@v4
|
|
33
|
+
with:
|
|
34
|
+
fetch-depth: 0
|
|
35
|
+
|
|
36
|
+
- name: Check commit hygiene — no claude.ai session URLs or Co-authored-by Claude trailers
|
|
37
|
+
run: |
|
|
38
|
+
set -euo pipefail
|
|
39
|
+
if [ "$GITHUB_EVENT_NAME" = "pull_request" ]; then
|
|
40
|
+
RANGE="origin/main..HEAD"
|
|
41
|
+
else
|
|
42
|
+
# Push to main: check only the newly pushed commit(s).
|
|
43
|
+
BEFORE="${{ github.event.before }}"
|
|
44
|
+
SHA="${{ github.sha }}"
|
|
45
|
+
if [ "$BEFORE" = "0000000000000000000000000000000000000000" ] || ! git cat-file -e "$BEFORE" 2>/dev/null; then
|
|
46
|
+
RANGE="HEAD^..HEAD"
|
|
47
|
+
else
|
|
48
|
+
RANGE="$BEFORE..$SHA"
|
|
49
|
+
fi
|
|
50
|
+
fi
|
|
51
|
+
MSGS=$(git log "$RANGE" --format='%B' 2>/dev/null || git log -1 --format='%B')
|
|
52
|
+
if echo "$MSGS" | grep -qF 'claude.ai/code/session'; then
|
|
53
|
+
echo "ERROR: commit message contains a claude.ai/code/session URL — remove it before merging"
|
|
54
|
+
exit 1
|
|
55
|
+
fi
|
|
56
|
+
if echo "$MSGS" | grep -iqF 'co-authored-by: claude'; then
|
|
57
|
+
echo "ERROR: commit message contains a Co-authored-by: Claude trailer — remove it before merging"
|
|
58
|
+
exit 1
|
|
59
|
+
fi
|
|
60
|
+
echo "OK: commit hygiene verified"
|
|
61
|
+
echo "RANGE=$RANGE" >> "$GITHUB_ENV"
|
|
62
|
+
|
|
63
|
+
- uses: actions/setup-python@v5
|
|
64
|
+
with:
|
|
65
|
+
python-version: "3.11"
|
|
66
|
+
|
|
67
|
+
- name: Check docs/paper/paper.md and docs/paper/latex/ stay in sync
|
|
68
|
+
run: python3 scripts/check_paper_latex_sync.py "$RANGE"
|
|
69
|
+
|
|
70
|
+
pip-audit:
|
|
71
|
+
runs-on: ubuntu-latest
|
|
72
|
+
steps:
|
|
73
|
+
- uses: actions/checkout@v4
|
|
74
|
+
|
|
75
|
+
- uses: actions/setup-python@v5
|
|
76
|
+
with:
|
|
77
|
+
python-version: "3.11"
|
|
78
|
+
|
|
79
|
+
- name: Install uv
|
|
80
|
+
run: pip install uv --quiet
|
|
81
|
+
|
|
82
|
+
- name: Sync dependencies
|
|
83
|
+
run: uv sync --extra dev
|
|
84
|
+
|
|
85
|
+
- name: Run pip-audit
|
|
86
|
+
# Non-blocking (T1 tier): findings are surfaced in logs but never fail the build.
|
|
87
|
+
continue-on-error: true
|
|
88
|
+
run: uv run --with pip-audit pip-audit
|
|
@@ -0,0 +1,43 @@
|
|
|
1
|
+
name: Release
|
|
2
|
+
|
|
3
|
+
on:
|
|
4
|
+
push:
|
|
5
|
+
tags:
|
|
6
|
+
- "v*"
|
|
7
|
+
|
|
8
|
+
concurrency:
|
|
9
|
+
group: ${{ github.workflow }}-${{ github.ref }}
|
|
10
|
+
cancel-in-progress: false
|
|
11
|
+
|
|
12
|
+
jobs:
|
|
13
|
+
publish:
|
|
14
|
+
runs-on: ubuntu-latest
|
|
15
|
+
environment: pypi
|
|
16
|
+
permissions:
|
|
17
|
+
id-token: write # required for PyPI Trusted Publishing (OIDC) -- no token/password used
|
|
18
|
+
contents: read
|
|
19
|
+
|
|
20
|
+
steps:
|
|
21
|
+
- uses: actions/checkout@v4
|
|
22
|
+
with:
|
|
23
|
+
# Explicit: build from the tag that triggered this workflow, never a
|
|
24
|
+
# branch head. For a `push: tags:` event, github.ref IS the tag ref
|
|
25
|
+
# (refs/tags/vX.Y.Z) -- stated explicitly here so this can't silently
|
|
26
|
+
# drift to a branch checkout if the trigger config ever changes.
|
|
27
|
+
ref: ${{ github.ref }}
|
|
28
|
+
|
|
29
|
+
- uses: actions/setup-python@v5
|
|
30
|
+
with:
|
|
31
|
+
python-version: "3.11"
|
|
32
|
+
|
|
33
|
+
- name: Install uv
|
|
34
|
+
run: pip install uv --quiet
|
|
35
|
+
|
|
36
|
+
- name: Build sdist + wheel
|
|
37
|
+
run: uv build
|
|
38
|
+
|
|
39
|
+
- name: twine check
|
|
40
|
+
run: uvx twine check dist/*.whl dist/*.tar.gz
|
|
41
|
+
|
|
42
|
+
- name: Publish to PyPI (Trusted Publishing -- OIDC, no token)
|
|
43
|
+
uses: pypa/gh-action-pypi-publish@release/v1
|
|
@@ -0,0 +1,89 @@
|
|
|
1
|
+
# Python
|
|
2
|
+
__pycache__/
|
|
3
|
+
*.py[cod]
|
|
4
|
+
*$py.class
|
|
5
|
+
*.so
|
|
6
|
+
.Python
|
|
7
|
+
build/
|
|
8
|
+
dist/
|
|
9
|
+
*.egg-info/
|
|
10
|
+
.eggs/
|
|
11
|
+
*.egg
|
|
12
|
+
MANIFEST
|
|
13
|
+
|
|
14
|
+
# Virtual environments
|
|
15
|
+
.venv/
|
|
16
|
+
venv/
|
|
17
|
+
env/
|
|
18
|
+
ENV/
|
|
19
|
+
|
|
20
|
+
# uv
|
|
21
|
+
.uv/
|
|
22
|
+
|
|
23
|
+
# Testing
|
|
24
|
+
.pytest_cache/
|
|
25
|
+
.coverage
|
|
26
|
+
coverage.xml
|
|
27
|
+
htmlcov/
|
|
28
|
+
*.coveragerc
|
|
29
|
+
|
|
30
|
+
# Mypy
|
|
31
|
+
.mypy_cache/
|
|
32
|
+
.dmypy.json
|
|
33
|
+
dmypy.json
|
|
34
|
+
|
|
35
|
+
# Ruff
|
|
36
|
+
.ruff_cache/
|
|
37
|
+
|
|
38
|
+
# IDE
|
|
39
|
+
.vscode/
|
|
40
|
+
.idea/
|
|
41
|
+
*.swp
|
|
42
|
+
*.swo
|
|
43
|
+
|
|
44
|
+
# OS
|
|
45
|
+
.DS_Store
|
|
46
|
+
Thumbs.db
|
|
47
|
+
|
|
48
|
+
# Secrets / env
|
|
49
|
+
.env
|
|
50
|
+
.env.*
|
|
51
|
+
!.env.example
|
|
52
|
+
|
|
53
|
+
# Distribution
|
|
54
|
+
*.whl
|
|
55
|
+
*.tar.gz
|
|
56
|
+
|
|
57
|
+
# Reports
|
|
58
|
+
reports/
|
|
59
|
+
*.html
|
|
60
|
+
!docs/**/*.html
|
|
61
|
+
|
|
62
|
+
# Worktrees
|
|
63
|
+
.claude/worktrees/
|
|
64
|
+
|
|
65
|
+
# Ad-hoc spot-check scripts (uncommitted per convention)
|
|
66
|
+
scripts/spot_check_discoverability_judge.py
|
|
67
|
+
scripts/validate_t5_docs_manifest.py
|
|
68
|
+
scripts/spot_check_fixer.py
|
|
69
|
+
scripts/pr32_body.txt
|
|
70
|
+
scripts/tx_breakdown_analysis.py
|
|
71
|
+
scripts/analyze_frontier_t18.py
|
|
72
|
+
|
|
73
|
+
# Run-output scratch files (experiment results, watchdog logs)
|
|
74
|
+
*_result.txt
|
|
75
|
+
*_results.txt
|
|
76
|
+
*_stderr.txt
|
|
77
|
+
*_watchdog.txt
|
|
78
|
+
*test_output.txt
|
|
79
|
+
ollama_conn_monitor.*
|
|
80
|
+
t18_watchdog.ps1
|
|
81
|
+
ty_t18_watchdog.txt
|
|
82
|
+
|
|
83
|
+
# LaTeX build artifacts (docs/paper/latex/main.pdf is committed deliberately)
|
|
84
|
+
docs/paper/latex/*.aux
|
|
85
|
+
docs/paper/latex/*.bbl
|
|
86
|
+
docs/paper/latex/*.log
|
|
87
|
+
docs/paper/latex/*.out
|
|
88
|
+
docs/paper/latex/*.blg
|
|
89
|
+
docs/paper/latex/*.toc
|
|
@@ -0,0 +1,27 @@
|
|
|
1
|
+
repos:
|
|
2
|
+
- repo: https://github.com/pre-commit/pre-commit-hooks
|
|
3
|
+
rev: v5.0.0
|
|
4
|
+
hooks:
|
|
5
|
+
- id: trailing-whitespace
|
|
6
|
+
- id: end-of-file-fixer
|
|
7
|
+
- id: check-yaml
|
|
8
|
+
- id: check-merge-conflict
|
|
9
|
+
- id: debug-statements
|
|
10
|
+
- id: check-added-large-files
|
|
11
|
+
|
|
12
|
+
- repo: https://github.com/astral-sh/ruff-pre-commit
|
|
13
|
+
rev: v0.15.15
|
|
14
|
+
hooks:
|
|
15
|
+
- id: ruff
|
|
16
|
+
args: [--fix]
|
|
17
|
+
- id: ruff-format
|
|
18
|
+
|
|
19
|
+
- repo: https://github.com/pre-commit/mirrors-mypy
|
|
20
|
+
rev: v2.1.0
|
|
21
|
+
hooks:
|
|
22
|
+
- id: mypy
|
|
23
|
+
|
|
24
|
+
- repo: https://github.com/gitleaks/gitleaks
|
|
25
|
+
rev: v8.29.1
|
|
26
|
+
hooks:
|
|
27
|
+
- id: gitleaks
|
|
@@ -0,0 +1,41 @@
|
|
|
1
|
+
# Autonomy Contract
|
|
2
|
+
|
|
3
|
+
This repo is driven by a Claude Code cloud scheduled task. Each run picks one
|
|
4
|
+
backlog item from TASKS.md, implements it on a `claude/<task>` branch, and opens a
|
|
5
|
+
draft PR for a human to review. These rules govern every autonomous run.
|
|
6
|
+
|
|
7
|
+
## Hard rules
|
|
8
|
+
|
|
9
|
+
1. **Only push to `claude/*` branches.** Never touch `main`, `chore/*`, `feat/*`, or any
|
|
10
|
+
branch not prefixed `claude/`.
|
|
11
|
+
2. **Never modify CI, workflows, or secrets.** `.github/workflows/`, `.github/secrets`,
|
|
12
|
+
`pyproject.toml` CI config, and `scripts/verify.sh` are read-only unless the task
|
|
13
|
+
explicitly changes them.
|
|
14
|
+
3. **Never delete, skip, or weaken tests to make `verify.sh` pass.** If tests fail, fix
|
|
15
|
+
the implementation — never remove the test. `pytest` must pass at full coverage.
|
|
16
|
+
4. **Exactly one TASKS.md item per run.** Pick the single highest-priority TODO item
|
|
17
|
+
with clear, testable acceptance criteria. If none qualify, stop and write a blocker
|
|
18
|
+
report — do not invent work.
|
|
19
|
+
5. **No dependency upgrades unless the task IS the upgrade.** Do not bump versions in
|
|
20
|
+
`pyproject.toml` as a side effect.
|
|
21
|
+
6. **PRs are always DRAFT.** Never open a ready-to-merge PR autonomously.
|
|
22
|
+
7. **If acceptance criteria are unclear, stop and report.** Do not guess at intent.
|
|
23
|
+
Write a `BLOCKER.md` at repo root explaining what is unclear and exit.
|
|
24
|
+
|
|
25
|
+
## Workflow per run
|
|
26
|
+
|
|
27
|
+
1. Read `CLAUDE.md` (architecture, conventions).
|
|
28
|
+
2. Read `AUTONOMY.md` (this file).
|
|
29
|
+
3. Read `TASKS.md` — pick the single highest-priority TODO with clear acceptance criteria.
|
|
30
|
+
4. Branch: `claude/<kebab-task-name>`.
|
|
31
|
+
5. Implement via orchestrator → executor → verifier pattern.
|
|
32
|
+
6. Run `./scripts/verify.sh`. This is the definition of done.
|
|
33
|
+
7. If green: commit, push branch, open DRAFT PR, move item to IN-REVIEW in TASKS.md.
|
|
34
|
+
8. If not green: leave branch uncommitted, open no PR, write a blocker summary in the PR
|
|
35
|
+
description or a comment, do NOT move the item.
|
|
36
|
+
9. End with a run report (what was done, verify.sh result, PR link if created).
|
|
37
|
+
|
|
38
|
+
## Definition of done
|
|
39
|
+
|
|
40
|
+
`./scripts/verify.sh` exits 0. That means: ruff passes, mypy passes (if configured),
|
|
41
|
+
all tests pass with no mocks removed or tests deleted.
|