outerloop-science 0.1.0.dev0__tar.gz → 0.1.0.dev2__tar.gz

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (169) hide show
  1. {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/.gitignore +3 -0
  2. {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/CHANGELOG.md +141 -0
  3. {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/CLAUDE.md +1 -1
  4. outerloop_science-0.1.0.dev2/CONTRIBUTING.md +89 -0
  5. {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/PKG-INFO +14 -10
  6. {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/README.md +11 -7
  7. outerloop_science-0.1.0.dev2/containers/README.md +8 -0
  8. outerloop_science-0.1.0.dev2/containers/agent-py312.def +43 -0
  9. {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/docs/design/architecture.md +6 -2
  10. {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/docs/design/dispatcher.md +19 -4
  11. {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/docs/design/github-app-auth.md +5 -5
  12. {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/docs/design/headline.md +1 -1
  13. {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/docs/design/onboarding.md +3 -2
  14. {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/docs/design/resident-tick.md +5 -5
  15. {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/docs/design/reviewer-infra.md +24 -0
  16. {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/docs/design/role-cli.md +4 -4
  17. {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/docs/install.md +75 -9
  18. {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/docs/roadmap.md +4 -1
  19. {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/docs/validation/author-syscalls.md +5 -5
  20. {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/pyproject.toml +14 -5
  21. {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/scripts/README.md +1 -1
  22. {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/scripts/install_codex.sh +2 -2
  23. {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/scripts/tick_chain.sbatch +41 -35
  24. {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/scripts/tick_deploy.sh +42 -38
  25. {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/scripts/tick_resident.sh +23 -20
  26. outerloop_science-0.1.0.dev2/src/outerloop/__init__.py +18 -0
  27. {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/src/outerloop/appauth.py +3 -3
  28. {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/src/outerloop/appmanifest.py +16 -11
  29. {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/src/outerloop/attempt.py +80 -52
  30. {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/src/outerloop/cli.py +170 -86
  31. {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/src/outerloop/climbboard.py +209 -4
  32. {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/src/outerloop/compute.py +84 -7
  33. {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/src/outerloop/contract.py +3 -2
  34. {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/src/outerloop/followup.py +37 -23
  35. {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/src/outerloop/github.py +4 -4
  36. {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/src/outerloop/harness.py +18 -7
  37. outerloop_science-0.1.0.dev2/src/outerloop/image.py +372 -0
  38. outerloop_science-0.1.0.dev2/src/outerloop/init.py +701 -0
  39. {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/src/outerloop/limits.py +1 -1
  40. {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/src/outerloop/measure.py +2 -2
  41. {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/src/outerloop/orchestrator.py +1 -1
  42. {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/src/outerloop/paths.py +14 -1
  43. {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/src/outerloop/review.py +1 -1
  44. {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/src/outerloop/steward.py +3 -3
  45. {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/src/outerloop/syscall.py +1 -0
  46. {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/src/outerloop/tick.py +400 -105
  47. outerloop_science-0.1.0.dev2/tests/conftest.py +50 -0
  48. {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/tests/test_attempt.py +42 -29
  49. {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/tests/test_bot_aliases.py +7 -7
  50. {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/tests/test_bot_login.py +5 -5
  51. {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/tests/test_climbboard.py +104 -0
  52. {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/tests/test_compute.py +64 -10
  53. outerloop_science-0.1.0.dev2/tests/test_env_bridge.py +86 -0
  54. {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/tests/test_followup.py +60 -0
  55. {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/tests/test_harness.py +35 -8
  56. outerloop_science-0.1.0.dev2/tests/test_image.py +294 -0
  57. outerloop_science-0.1.0.dev2/tests/test_init.py +763 -0
  58. {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/tests/test_local_compute.py +30 -6
  59. {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/tests/test_measure.py +1 -1
  60. {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/tests/test_start.py +214 -66
  61. {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/tests/test_sweep_git_locks.py +7 -0
  62. {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/tests/test_tick.py +569 -76
  63. {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/tests/test_tick_chain_successors.py +2 -2
  64. {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/tests/test_tick_resident.py +33 -32
  65. outerloop_science-0.1.0.dev2/tests/test_tiers.py +37 -0
  66. {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/tests/test_verify_agent.py +13 -11
  67. {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/uv.lock +140 -6
  68. outerloop_science-0.1.0.dev0/CONTRIBUTING.md +0 -75
  69. outerloop_science-0.1.0.dev0/src/outerloop/__init__.py +0 -18
  70. outerloop_science-0.1.0.dev0/src/outerloop/init.py +0 -313
  71. outerloop_science-0.1.0.dev0/tests/test_env_bridge.py +0 -61
  72. outerloop_science-0.1.0.dev0/tests/test_init.py +0 -242
  73. {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/.pre-commit-config.yaml +0 -0
  74. {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/.python-version +0 -0
  75. {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/LICENSE +0 -0
  76. {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/NOTICE +0 -0
  77. {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/RELEASING.md +0 -0
  78. {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/SECURITY.md +0 -0
  79. {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/docs/assets/icon-dark.svg +0 -0
  80. {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/docs/assets/icon-light.svg +0 -0
  81. {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/docs/assets/icon.svg +0 -0
  82. {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/docs/compute.md +0 -0
  83. {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/docs/contract.md +0 -0
  84. {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/docs/design/agent-substrate.md +0 -0
  85. {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/docs/design/consolidation.md +0 -0
  86. {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/docs/design/external.md +0 -0
  87. {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/docs/design/judge-placement.md +0 -0
  88. {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/docs/design/meta.md +0 -0
  89. {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/docs/design/orchestrator-verify.md +0 -0
  90. {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/docs/design/public-surface.md +0 -0
  91. {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/docs/design/research-lines.md +0 -0
  92. {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/docs/design/research-loop-buildout.md +0 -0
  93. {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/docs/design/research-loop.md +0 -0
  94. {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/docs/design/review-placement.md +0 -0
  95. {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/docs/design/roles.md +0 -0
  96. {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/docs/design/scaling.md +0 -0
  97. {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/docs/reviewer.md +0 -0
  98. {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/examples/review.yml +0 -0
  99. {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/scripts/install_hermes.sh +0 -0
  100. {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/scripts/requeue_moved_successors.sh +0 -0
  101. {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/scripts/setup_branch_protection.sh +0 -0
  102. {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/scripts/sweep_git_locks.sh +0 -0
  103. {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/src/outerloop/__main__.py +0 -0
  104. {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/src/outerloop/brief.py +0 -0
  105. {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/src/outerloop/contract_cli.py +0 -0
  106. {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/src/outerloop/disk.py +0 -0
  107. {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/src/outerloop/dispatch.py +0 -0
  108. {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/src/outerloop/housekeeping.py +0 -0
  109. {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/src/outerloop/intake.py +0 -0
  110. {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/src/outerloop/markers.py +0 -0
  111. {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/src/outerloop/panel.py +0 -0
  112. {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/src/outerloop/posting.py +0 -0
  113. {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/src/outerloop/progress.py +0 -0
  114. {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/src/outerloop/py.typed +0 -0
  115. {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/src/outerloop/review_agent.py +0 -0
  116. {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/src/outerloop/review_agent_cli.py +0 -0
  117. {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/src/outerloop/review_post_cli.py +0 -0
  118. {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/src/outerloop/review_summarize_cli.py +0 -0
  119. {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/src/outerloop/role_runner.py +0 -0
  120. {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/src/outerloop/roles.py +0 -0
  121. {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/src/outerloop/rolespec.py +0 -0
  122. {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/src/outerloop/runstate.py +0 -0
  123. {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/src/outerloop/style.py +0 -0
  124. {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/src/outerloop/syscall_cli.py +0 -0
  125. {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/src/outerloop/verifier.py +0 -0
  126. {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/src/outerloop/verify_agent.py +0 -0
  127. {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/src/outerloop/verify_agent_cli.py +0 -0
  128. {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/src/outerloop/verify_post_cli.py +0 -0
  129. {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/tests/test_appauth.py +0 -0
  130. {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/tests/test_appmanifest.py +0 -0
  131. {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/tests/test_brief.py +0 -0
  132. {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/tests/test_channel_dir.py +0 -0
  133. {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/tests/test_codex_harness.py +0 -0
  134. {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/tests/test_contract.py +0 -0
  135. {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/tests/test_contract_cli.py +0 -0
  136. {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/tests/test_contract_names.py +0 -0
  137. {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/tests/test_disk.py +0 -0
  138. {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/tests/test_dispatch.py +0 -0
  139. {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/tests/test_github.py +0 -0
  140. {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/tests/test_hardening.py +0 -0
  141. {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/tests/test_hermes_harness.py +0 -0
  142. {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/tests/test_housekeeping.py +0 -0
  143. {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/tests/test_import.py +0 -0
  144. {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/tests/test_intake.py +0 -0
  145. {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/tests/test_limits.py +0 -0
  146. {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/tests/test_markers.py +0 -0
  147. {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/tests/test_measure_and_decide.py +0 -0
  148. {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/tests/test_orchestrator.py +0 -0
  149. {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/tests/test_packaging.py +0 -0
  150. {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/tests/test_panel.py +0 -0
  151. {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/tests/test_paths.py +0 -0
  152. {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/tests/test_posting.py +0 -0
  153. {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/tests/test_progress.py +0 -0
  154. {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/tests/test_requeue_moved_successors.py +0 -0
  155. {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/tests/test_review.py +0 -0
  156. {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/tests/test_review_agent.py +0 -0
  157. {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/tests/test_review_agent_cli.py +0 -0
  158. {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/tests/test_review_hardening.py +0 -0
  159. {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/tests/test_review_policy.py +0 -0
  160. {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/tests/test_review_summarize.py +0 -0
  161. {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/tests/test_role_runner.py +0 -0
  162. {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/tests/test_rolespec.py +0 -0
  163. {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/tests/test_runstate.py +0 -0
  164. {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/tests/test_steward.py +0 -0
  165. {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/tests/test_syscall.py +0 -0
  166. {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/tests/test_syscall_cli.py +0 -0
  167. {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/tests/test_verifier.py +0 -0
  168. {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/tests/test_verify_agent_cli.py +0 -0
  169. {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/tests/test_version.py +0 -0
@@ -31,3 +31,6 @@ outputs/
31
31
 
32
32
  # Cloudflare wrangler local cache (setup page deploys)
33
33
  .wrangler/
34
+
35
+ # pytest-testmon's change-tracking database
36
+ .testmondata
@@ -6,6 +6,82 @@ Versions follow [SemVer](https://semver.org).
6
6
 
7
7
  ## [Unreleased]
8
8
 
9
+ ### Added
10
+
11
+ - `climb/status.json` carries the fleet's queue: the kernel's own Slurm jobs
12
+ (tick chain, sessions, wakes, evals, launches), each attributed to an agent,
13
+ with state, elapsed time, partition and submit time. Jobs on the account
14
+ that are not the kernel's never appear. The strip republishes when a job
15
+ appears, leaves, or changes state, not on elapsed drift.
16
+ - `outerloop init` asks for the author's model API key (hidden) and writes it to
17
+ `~/.config/outerloop/<backend>_key` (0600), or takes `--author-key-file` for an
18
+ existing file (checked to exist, stored absolute); the `.env` records it as
19
+ `OUTERLOOP_<BACKEND>_KEY_FILE`. The key lands in the config dir in use:
20
+ `~/.config/outerloop/`, or `~/.config/autoresearch/` on a machine set up
21
+ before the rename. The focused `--github-app` run still asks nothing about
22
+ the author. Every secret `init` writes (keys, PAT, App PEM and JSON, `.env`)
23
+ is now created 0600 in one step rather than written and then tightened.
24
+ - The agent image is published: `containers/agent-py312.def` is its recipe
25
+ (Ubuntu 24.04, Python 3.12, uv, git, build tools), the `build-image` workflow
26
+ builds and uploads it with a checksum to huggingface.co/outerloop-science/
27
+ agent-image on every recipe change, and `outerloop init` downloads it on
28
+ Linux when Apptainer is installed, so local mode runs contained by default
29
+ (`--image` names your own, `--no-image` opts out).
30
+ - `outerloop init` checks that Apptainer can actually run a container before
31
+ downloading the image (on Ubuntu 24.04 a hand-installed Apptainer is on PATH
32
+ but cannot create user namespaces), and when it cannot, prints the install
33
+ steps for the machine's distribution and continues uncontained. The image
34
+ download shows a progress bar with size, speed and ETA on a terminal, one
35
+ line per 10% otherwise, and confirms the checksum. `docs/install.md` gains
36
+ an "Installing Apptainer" section.
37
+ - Launch admission: with `OUTERLOOP_MAX_LAUNCH_GPUS` set to the per-user GPU
38
+ cap, author launches are submitted held and the tick releases them
39
+ oldest-first while the user's GPU jobs fit under the cap, cancels held
40
+ launches of runs that have ended, and stops releasing when Slurm parks a
41
+ released launch on a per-user reason. `squeue` rows on the board carry the
42
+ pending reason and GRES. Unset, launches queue as before.
43
+
44
+ ### Changed
45
+
46
+ - The board's queue is a card in the live strip, not a preformatted dump: a
47
+ full-width table with running jobs first, state pills, elapsed time,
48
+ partition (first of several, the rest counted, all on hover), and the
49
+ owning agent in its color; long job names are cut with an ellipsis and
50
+ shown whole on hover.
51
+ - The Claude author's key follows the same rule as Codex's: the file is
52
+ `~/.config/outerloop/claude_key` and the setting `OUTERLOOP_CLAUDE_KEY_FILE`.
53
+ The pre-rename `harness_key` file and `*_HARNESS_KEY_FILE` setting are still
54
+ read (the new name wins when both exist), so an existing setup keeps working.
55
+ - `AUTORESEARCH_TARGET` is required for in-review servicing (follow-ups, wakes,
56
+ the board). The kernel no longer falls back to the retired
57
+ `agentic-learning-ai-lab/autoresearch-pilot` repo when it is unset; the tick
58
+ logs the missing setting like any other.
59
+ - The test suite runs on all cores by default (pytest-xdist), about four times
60
+ faster locally; `pytest --testmon` runs only the tests affected by your edits
61
+ (pytest-testmon). Tier selection moved from `addopts` into `tests/conftest.py`
62
+ (a `-m` in `addopts` disables testmon), and a `serial` tier holds the tests
63
+ that inspect the process table: skipped while workers run, and `pytest -m
64
+ serial` turns workers off, so the tier always runs alone (CI's second step).
65
+ - The kernel reads `OUTERLOOP_*` everywhere: every internal read, log line,
66
+ chain script and test now uses the public names. The pre-rename
67
+ `AUTORESEARCH_*` names are still accepted for one release, bridged at the
68
+ process boundary, the `.env` file and the chain scripts, and will be
69
+ dropped in the release after 0.1. The review footer names `outerloop`, and
70
+ `outerloop --version` exists; `start`, `tick` and `init` describe themselves
71
+ and every flag in `--help`.
72
+ - The on-disk and queue names follow the rename: the resident tick job is
73
+ `outerloop-resident` (the per-cadence chain `outerloop-tick`), the local
74
+ state root defaults to `~/.outerloop`, the default image path is
75
+ `~/outerloop-images/agent-py312.sif`. `outerloop start` refuses while a
76
+ pre-rename `autoresearch-resident` is queued or running, since Slurm's
77
+ singleton serializes by name; cancel it first. An existing `~/.autoresearch`
78
+ or `~/autoresearch-images/` is used while the new path does not exist. A
79
+ chain-mode deployment (manual `tick_chain.sbatch`) must be stopped with
80
+ `scancel --name autoresearch-tick` before deploying this version. Every
81
+ honored pre-rename name (`~/.config/autoresearch/`, `.autoresearch.yaml`, the
82
+ `.autoresearch` channel dir, the old root, image and job names) is dropped in
83
+ the release after 0.1; the comment marker keeps recognizing old comments.
84
+
9
85
  ### Changed
10
86
 
11
87
  - The syscall channel dir in a workspace is now `.outerloop/` (the tool, the
@@ -51,9 +127,74 @@ Versions follow [SemVer](https://semver.org).
51
127
 
52
128
  ### Fixed
53
129
 
130
+ - `outerloop init --github-app` says what to do at each step: which page
131
+ opens, which button to click, where the code appears, how to install the
132
+ App on the repository, and that the last step checks write access. It also
133
+ asks whether to create the App under your account or an organization,
134
+ which before needed the undocumented `--org` flag.
135
+ - The App flow's final write check asked GitHub a question installation
136
+ tokens cannot answer, so it warned about a missing push permission on Apps
137
+ that could push. It now checks the installation's repository list and its
138
+ granted permissions.
139
+ - `init --github-app` records the App's login (`<slug>[bot]`) as
140
+ `OUTERLOOP_BOT_LOGIN`. Without it the kernel fell back to a built-in
141
+ default that is not the adopter's identity and did not recognize its own
142
+ pull requests.
143
+ - The local loop launches attempts without a container image. On a machine
144
+ with no Apptainer image, `AUTORESEARCH_COMPUTE=local` now brings up
145
+ servicing with an empty image, logs once that sessions run under the
146
+ harness's own sandbox and evaluations run bare on that machine, and passes
147
+ `--uncontained` to its jobs. The panel is off in that mode unless
148
+ `OUTERLOOP_PANEL_UNCONTAINED=1` opts in, and says so when it runs. Slurm
149
+ still requires the image, and so does a Codex author in every mode (#289).
150
+ - The PyPI description no longer refers to the lab.
151
+ - A base-synced re-measurement is finished instead of abandoned. The parked
152
+ stage recorded the session's local merge of the moved base as its parent,
153
+ a commit GitHub never saw, so the resume judged the PR head as "moved" on
154
+ every base-synced park and threw the landed result away. The stage now
155
+ records the PR head at park and the resume compares against that. An
156
+ abandon also releases the wake-attempt cap, so the "ask again" it invites
157
+ can be served.
158
+ - A follow-up re-measurement that has landed is now finished even when the
159
+ run has spent its wake-attempt cap. The sessions that sync a stale PR onto
160
+ the moved base and dispatch the measurement each bill an attempt, so a
161
+ measurement could complete with nobody left to push its number; the run then
162
+ idled with "follow-up attempts without progress" every tick. The finishing
163
+ follow-up has its own allowance of `MAX_WAKE_ATTEMPTS`, counted on the
164
+ stage, so a session that fails to finish is retried a bounded number of
165
+ times; the cap still holds for follow-ups that have nothing landed to finish.
54
166
  - `outerloop init` no longer overwrites an existing `~/.config/autoresearch/.env`
55
167
  silently: interactively it asks; with `--yes` it refuses unless `--force` is
56
168
  passed. The check runs before any GitHub App is created.
169
+ - Slurm account and partition are both optional (#300). `outerloop start`
170
+ no longer requires an account, `outerloop init` no longer insists on one,
171
+ the tick's in-review servicing and the attempt's resume and dispatch gates
172
+ need only the image, and every sbatch the kernel builds passes `--account`
173
+ and `--partition` only when set, so Slurm bills the caller's default
174
+ association and picks its default partition.
175
+ - A launch still queued past its deadline is no longer cancelled and woken as
176
+ "unschedulable" when Slurm's pending reason is a wait: a busy queue
177
+ (Priority, Resources), a reservation or dependency, or the account's own
178
+ per-user/per-account/group cap (`QOSMaxGRESPerUser` and kin). The sweep
179
+ extends the deadline by one slack window and logs the reasons; a queued job
180
+ keeps its priority age, where a re-launch would start over at the back.
181
+ Reasons that never clear (a dependency that cannot be satisfied, per-job
182
+ limits, an invalid account or QOS, a held job) still cancel and wake.
183
+ - `outerloop init` stops with exit 1 when the credential cannot open pull
184
+ requests on the target (401, 403, 404, or no write access) instead of
185
+ printing a warning and `next: outerloop start` (#285); a network failure stays
186
+ a warning. The PAT path records the token's login as `OUTERLOOP_BOT_LOGIN`,
187
+ and the tick warns when the login is unset and the lab default applies (#298).
188
+ - `init` locates the author's CLI (`claude`/`codex`) on PATH or in
189
+ `~/.local/bin`, records it as `OUTERLOOP_CLAUDE_BIN`/`OUTERLOOP_CODEX_BIN`, and
190
+ says how to install it when absent; `attempt --claude-bin` honors that
191
+ setting; `start` refuses to launch when a recorded binary is missing; a
192
+ `spawn-error` names the binary and the OS error (#294).
193
+ - Local jobs keep their combined stdout/stderr in `local_jobs/<id>.out` beside
194
+ the state file, and the log line names it when the job did not complete (#295).
195
+ - `cryptography` is a base dependency, so the GitHub App identity works from a
196
+ plain `pip install`; the `app-auth` extra is now a no-op (#293). The `tick`
197
+ subcommand's help no longer uses lab-internal vocabulary (#286).
57
198
 
58
199
  ### Added
59
200
 
@@ -8,7 +8,7 @@ Scaffold phase — `docs/roadmap.md` says what exists vs. planned;
8
8
 
9
9
  ```bash
10
10
  uv sync
11
- uv run pytest # slow/llm/slurm excluded by default
11
+ uv run pytest # all cores; --testmon = only tests affected by your edits
12
12
  uv run ruff check --fix . && uv run ruff format .
13
13
  uv run mypy
14
14
  uv run pre-commit run --all-files
@@ -0,0 +1,89 @@
1
+ # Contributing
2
+
3
+ Outerloop is developed in the open by the [Agentic Learning AI Lab](https://agenticlearning.ai)
4
+ at NYU, and its own agents are regular contributors to it. Contributions from
5
+ people are welcome too: bug reports, fixes, documentation, new compute
6
+ backends, and benchmark contracts for your own repos.
7
+
8
+ ## Before you start
9
+
10
+ - Look through the open issues. For anything larger than a small fix, open an
11
+ issue first so we can agree on the shape before you write code.
12
+ - Security problems go to the contact in [SECURITY.md](SECURITY.md), not to a
13
+ public issue.
14
+
15
+ ## Setup
16
+
17
+ ```bash
18
+ uv sync
19
+ uv run pre-commit install
20
+ uv run pytest
21
+ ```
22
+
23
+ ## Making a change
24
+
25
+ 1. Branch from `main`: `fix/<topic>` or `feat/<topic>`.
26
+ 2. Keep the change focused. Add or update tests, and a line under
27
+ `[Unreleased]` in `CHANGELOG.md` for anything a user would notice.
28
+ 3. Run the same gate CI runs before you push:
29
+
30
+ ```bash
31
+ uv run pre-commit run --all-files # lint, formatting, secret scan
32
+ uv lock --check
33
+ uv run mypy
34
+ uv run pytest && uv run pytest -m serial
35
+ ```
36
+
37
+ 4. Open a pull request against `main`. Say what changed, why, and how you
38
+ verified it.
39
+
40
+ ## What happens to your pull request
41
+
42
+ - CI must pass.
43
+ - One of our agents reviews it and posts its findings as a review, usually
44
+ within the hour; no time is guaranteed. It never approves or blocks; a
45
+ maintainer decides. Fix what is right and reply to what is not; findings
46
+ answered with a reason are fine. After you push fixes, a maintainer removes
47
+ and re-adds the `outerloop:review` label to run another round.
48
+ - The agent review runs only for branches in this repository. A pull request
49
+ from a fork gets CI but no agent review; a maintainer can push your branch
50
+ here to run one.
51
+ - A maintainer merges with a merge commit. We do not squash or rebase, so the
52
+ history of how a change came to be stays intact. Keep your branch current
53
+ with `git merge main`, and never force-push a branch someone else has seen.
54
+
55
+ ## Conventions
56
+
57
+ - Python 3.12, absolute imports (`from outerloop ...`), and code that works
58
+ from the installed wheel; no paths relative to a checkout.
59
+ - `ruff` for lint and formatting (line length 100); `mypy` clean.
60
+ - Dependencies go in the `pyproject.toml` group that owns them, then
61
+ `uv lock`; CI checks the lock.
62
+ - Docs are plain prose. Short sentences, no metaphors, no internal jargon.
63
+
64
+ ## Tests
65
+
66
+ | Tier | Marker | Runs |
67
+ | --- | --- | --- |
68
+ | unit | *(none)* | CI, every PR, on all cores |
69
+ | serial | `serial` | CI, second step: `pytest -m serial` (inspects the process table) |
70
+ | slow | `slow` | nightly or manual |
71
+ | llm | `llm` | manual only, paid APIs |
72
+ | slurm | `slurm` | manual only, needs a cluster |
73
+
74
+ While editing, `uv run pytest --testmon` runs only the tests affected by your
75
+ change; `uv run pytest -n0` runs everything serially.
76
+
77
+ ## Pull requests from the agents
78
+
79
+ Pull requests authored by `outerloop-science[bot]` are the system improving
80
+ the repositories it works on. They appear on those target repositories and are
81
+ skipped by the advisory reviewer by design. What gates them is each target's
82
+ own rules: its required human review, or, where a target's contract allows
83
+ automatic merging, the measurement gate and review panel that ran before the
84
+ PR opened. This document is about contributing to Outerloop, this repository.
85
+
86
+ ## License
87
+
88
+ By contributing you agree that your contributions are licensed under the
89
+ [Apache License 2.0](LICENSE), the same as the project.
@@ -1,7 +1,7 @@
1
1
  Metadata-Version: 2.5
2
2
  Name: outerloop-science
3
- Version: 0.1.0.dev0
4
- Summary: Autonomous research agent that co-develops the lab's benchmark-bearing repos
3
+ Version: 0.1.0.dev2
4
+ Summary: Autonomous research agents that improve the benchmark you point them at, one verified pull request at a time
5
5
  Project-URL: Homepage, https://outerloop.science
6
6
  Project-URL: Repository, https://github.com/outerloop-science/outerloop
7
7
  Project-URL: Changelog, https://github.com/outerloop-science/outerloop/blob/main/CHANGELOG.md
@@ -11,10 +11,10 @@ License-Expression: Apache-2.0
11
11
  License-File: LICENSE
12
12
  License-File: NOTICE
13
13
  Requires-Python: >=3.12
14
+ Requires-Dist: cryptography>=42
14
15
  Requires-Dist: pydantic>=2
15
16
  Requires-Dist: pyyaml>=6
16
17
  Provides-Extra: app-auth
17
- Requires-Dist: cryptography>=42; extra == 'app-auth'
18
18
  Description-Content-Type: text/markdown
19
19
 
20
20
  <picture>
@@ -56,18 +56,22 @@ themselves.
56
56
 
57
57
  ## Get started
58
58
 
59
- Three commands and two files: your model key and the contract. You need a repo
60
- with a benchmark command, a model API key (Anthropic by default), and a Slurm
61
- cluster or one machine with a GPU.
59
+ Three commands and one file. You need a repo with a benchmark command, an API
60
+ key for the model that will write the code, and a Slurm cluster or one machine
61
+ with a GPU.
62
62
 
63
63
  ```bash
64
64
  pip install outerloop-science
65
- outerloop init # where the loop runs, which repo, your GitHub bot; writes the config
65
+ outerloop init # where the loop runs, which repo, which model and its key, your GitHub identity
66
66
  ```
67
67
 
68
- Put your Anthropic API key in `~/.config/outerloop/harness_key`: one line,
69
- readable only by you (`chmod 600`). Then add one file, `.outerloop.yaml`, to
70
- the repo you want improved:
68
+ The wizard asks for a GitHub identity for the agents to open pull requests
69
+ as. Pick `app` and it walks you through creating a GitHub App, under your
70
+ account or under an organization you name, and installing it on the repo:
71
+ two browser pages and a code pasted back. It then checks that the App can
72
+ write the repo and tells you if it cannot. Pick `pat` if you already have a
73
+ token. It writes the config and the key files; nothing to edit by hand. Then
74
+ add one file, `.outerloop.yaml`, to the repo you want improved:
71
75
 
72
76
  ```yaml
73
77
  benchmarks:
@@ -37,18 +37,22 @@ themselves.
37
37
 
38
38
  ## Get started
39
39
 
40
- Three commands and two files: your model key and the contract. You need a repo
41
- with a benchmark command, a model API key (Anthropic by default), and a Slurm
42
- cluster or one machine with a GPU.
40
+ Three commands and one file. You need a repo with a benchmark command, an API
41
+ key for the model that will write the code, and a Slurm cluster or one machine
42
+ with a GPU.
43
43
 
44
44
  ```bash
45
45
  pip install outerloop-science
46
- outerloop init # where the loop runs, which repo, your GitHub bot; writes the config
46
+ outerloop init # where the loop runs, which repo, which model and its key, your GitHub identity
47
47
  ```
48
48
 
49
- Put your Anthropic API key in `~/.config/outerloop/harness_key`: one line,
50
- readable only by you (`chmod 600`). Then add one file, `.outerloop.yaml`, to
51
- the repo you want improved:
49
+ The wizard asks for a GitHub identity for the agents to open pull requests
50
+ as. Pick `app` and it walks you through creating a GitHub App, under your
51
+ account or under an organization you name, and installing it on the repo:
52
+ two browser pages and a code pasted back. It then checks that the App can
53
+ write the repo and tells you if it cannot. Pick `pat` if you already have a
54
+ token. It writes the config and the key files; nothing to edit by hand. Then
55
+ add one file, `.outerloop.yaml`, to the repo you want improved:
52
56
 
53
57
  ```yaml
54
58
  benchmarks:
@@ -0,0 +1,8 @@
1
+ # containers
2
+
3
+ `agent-py312.def` is the Apptainer recipe for the image the kernel runs sessions and
4
+ evaluations in. The `build-image` workflow builds it on every change to this directory and
5
+ on manual dispatch, and publishes the result to
6
+ [huggingface.co/outerloop-science/agent-image](https://huggingface.co/outerloop-science/agent-image)
7
+ with a checksum. `outerloop init` downloads it on Linux when Apptainer is installed; the
8
+ tick reads `~/outerloop-images/agent-py312.sif` unless `OUTERLOOP_IMAGE` says otherwise.
@@ -0,0 +1,43 @@
1
+ # The Outerloop agent image: what a contained session and a dispatched
2
+ # evaluation see. The kernel binds the workspace, a per-run home and the
3
+ # harness binary (claude/codex) into it; --nv adds the host's GPU driver for
4
+ # GPU benchmarks. So the image only has to provide the base a target's own
5
+ # `uv run ...` needs: Python 3.12, uv, git and a compiler.
6
+ #
7
+ # Built by .github/workflows/build-image.yml and published to
8
+ # https://huggingface.co/outerloop-science/agent-image; `outerloop init` pulls
9
+ # it on Linux. Build by hand: apptainer build agent-py312.sif containers/agent-py312.def
10
+
11
+ Bootstrap: docker
12
+ From: ubuntu:24.04
13
+
14
+ %post
15
+ set -eu
16
+ export DEBIAN_FRONTEND=noninteractive
17
+ apt-get update
18
+ apt-get install -y --no-install-recommends \
19
+ ca-certificates curl git git-lfs openssh-client \
20
+ build-essential pkg-config \
21
+ python3 python3-venv python3-dev \
22
+ libgomp1 libglib2.0-0 procps less unzip zip jq ripgrep
23
+ rm -rf /var/lib/apt/lists/*
24
+ # uv installs its own Pythons when a target pins one; the system 3.12 is the default
25
+ curl -LsSf https://astral.sh/uv/install.sh | env UV_INSTALL_DIR=/usr/local/bin INSTALLER_NO_MODIFY_PATH=1 sh
26
+ ln -sf /usr/bin/python3 /usr/local/bin/python
27
+ git lfs install --system
28
+ python3 --version && uv --version && git --version
29
+
30
+ %environment
31
+ export LC_ALL=C.UTF-8
32
+ export LANG=C.UTF-8
33
+ export PATH=/usr/local/bin:/usr/local/sbin:/usr/sbin:/usr/bin:/sbin:/bin
34
+
35
+ %labels
36
+ org.opencontainers.image.title outerloop agent image
37
+ org.opencontainers.image.source https://github.com/outerloop-science/outerloop
38
+ outerloop.recipe containers/agent-py312.def
39
+
40
+ %help
41
+ Outerloop agent image: Ubuntu 24.04, Python 3.12, uv, git, build tools.
42
+ The kernel runs sessions and evaluations in it with --containall --cleanenv,
43
+ binding only the workspace, a per-run home and the harness binary.
@@ -371,8 +371,12 @@ failure can strand a run:
371
371
  `deadline = submit_time + walltime + slack` (recomputed from `start_time`
372
372
  once the job starts, so late scheduling never truncates a healthy run).
373
373
  Past the deadline the sweep consults `sacct` and acts on what it finds:
374
- still PENDING → the experiment is unschedulable in practice; cancel it and
375
- wake the run with that fact. No record on a *successful* query → the job
374
+ still PENDING → ask squeue why: a wait (Priority/Resources, a reservation
375
+ or dependency, the account's own per-user/group cap such as
376
+ `QOSMaxGRESPerUser`) moves the deadline out by one slack window, since a
377
+ queued job keeps its priority age and a re-launch would start over; any
378
+ other reason means unschedulable in practice — cancel it and wake the
379
+ run with that fact. No record on a *successful* query → the job
376
380
  vanished; wake with that fact. Query failed or timed out → that is
377
381
  "Slurm unknown", never "job gone" — defer to the next tick, and only a
378
382
  sustained outage (its own alert via the watchdog) escalates. A fail-safe
@@ -66,7 +66,7 @@ orchestrator chooses per eval site, not per repo.
66
66
  launches alike: the JobSpec requests `--gpus=N`, and the jail adds `--nv` so
67
67
  the allocation is visible inside the container; nothing else about the
68
68
  containment changes. GPU jobs need their own lane, so the deployment names
69
- one — `AUTORESEARCH_GPU_PARTITION`, optionally `AUTORESEARCH_GPU_ACCOUNT`
69
+ one — `OUTERLOOP_GPU_PARTITION`, optionally `OUTERLOOP_GPU_ACCOUNT`
70
70
  (default: the CPU account) — and a benchmark with `gpus > 0` on a deployment
71
71
  without a lane is REFUSED at launch (both attempt lanes), never queued into
72
72
  evals that can never run. GPUs exist only on DISPATCHED jobs: the contract
@@ -76,6 +76,18 @@ benchmarks outright for now — a stewardship validates its rewrite in-job.
76
76
  The first target in this shape is the speedrun (`gpt-speedrun`: one H200
77
77
  per eval, ~3.5h).
78
78
 
79
+ **Launch admission.** A per-user GPU cap (a QOS's `MaxTRESPerUser`, 16 on
80
+ Torch) means the fleet can run only so many launches at once, whatever the
81
+ agent count: six 8-GPU launches queued against a 16-GPU cap left four pending
82
+ for a day on `QOSMaxGRESPerUser`. With `OUTERLOOP_MAX_LAUNCH_GPUS` set to that
83
+ cap, author launches are submitted HELD, and every tick `service_admission`
84
+ releases them oldest-first while the user's eligible GPU jobs (evals and
85
+ launches alike) fit under the cap; held launches whose run has since ended
86
+ are cancelled. A released launch that Slurm parks on a per-user reason means
87
+ the cap is set too high, and nothing more is released that tick. The wake's
88
+ `afterany` dependency is unchanged (held jobs are pending), and the sweep
89
+ treats the kernel's hold as a wait. Unset, launches queue as submitted.
90
+
79
91
  **Who decides the eval walltime.** The gate measures steps, not time, so
80
92
  walltime should bound only SPEND — and the attempt already has a spend
81
93
  budget. The contract's `eval_minutes` is the DEFAULT; the author may declare
@@ -234,9 +246,12 @@ dies -> the `eval-error` outcome at wake (ending `aborted`, exactly as the
234
246
  in-job eval failure maps today); a LOST wake -> the sweep's primary backup,
235
247
  which wakes any run whose job reads terminal after the grace period —
236
248
  plus GONE past the deadline (vanished) and PENDING past the deadline
237
- (cancel, then wake as unschedulable); a still-RUNNING job is deliberately
249
+ (when Slurm's reason is a wait — a busy queue, a reservation or dependency,
250
+ the account's own per-user/group cap — the deadline moves out by one slack
251
+ window; otherwise cancel, then wake as unschedulable); a still-RUNNING job is deliberately
238
252
  left alone, bounded by its own walltime; wakes that fire without producing
239
- progress -> `stuck` at MAX_WAKE_ATTEMPTS. A moved base during the wait is
253
+ progress -> `stuck` at MAX_WAKE_ATTEMPTS (a landed re-measure has its own
254
+ allowance of MAX_WAKE_ATTEMPTS finishing follow-ups, counted on the stage). A moved base during the wait is
240
255
  review's to handle — the sealed candidate publishes as-is (research-loop.md:
241
256
  a stale PR is a re-wake, never an orchestrator auto-merge). Worktree cleanup has a named owner at every exit: the
242
257
  wake's own finally (primary, as in-job measures do today), the sweep's
@@ -286,7 +301,7 @@ GPU benchmarks and portfolio climbs both multiply eval load; both were
286
301
  designed against this seam. Portfolio's N concurrent climbs become N
287
302
  waiting records with independent wakes — the serialization question stays
288
303
  in the picker, not the dispatcher. The tick's job-partition knobs
289
- (`AUTORESEARCH_JOB_PARTITION`) already route work jobs; per-benchmark
304
+ (`OUTERLOOP_JOB_PARTITION`) already route work jobs; per-benchmark
290
305
  partition/GPU allocation fields are a phase-3 contract-schema change
291
306
  (`Benchmark` rejects unknown fields today, deliberately), validated and
292
307
  clamped by operator ceilings like every other budget.
@@ -108,7 +108,7 @@ Swapping the live fleet's identity mid-campaign is an ops event:
108
108
 
109
109
  ## The flag
110
110
 
111
- `AUTORESEARCH_GITHUB_APP_FILE` names a JSON config —
111
+ `OUTERLOOP_GITHUB_APP_FILE` names a JSON config —
112
112
 
113
113
  ```json
114
114
  {"app_id": 1234, "installation_id": 5678,
@@ -117,13 +117,13 @@ Swapping the live fleet's identity mid-campaign is an ops event:
117
117
 
118
118
  — and rides the rails the PAT already rides: role CLIs default their
119
119
  `--github-app-file` from it (jobs inherit the tick's environment, the same
120
- way `AUTORESEARCH_AUTHOR_*` reaches them), and every bot-auth construction
120
+ way `OUTERLOOP_AUTHOR_*` reaches them), and every bot-auth construction
121
121
  goes through one factory, `appauth.resolve_bot_auth(pat_file, app_file)`:
122
122
  App provider when the config is set, PAT otherwise. Revert = unset the env
123
123
  var. The ids are not secrets; the key path inside keeps PAT-file custody
124
124
  (600, never committed).
125
125
 
126
- The identity flips WITH the credential: `AUTORESEARCH_BOT_LOGIN` names the
126
+ The identity flips WITH the credential: `OUTERLOOP_BOT_LOGIN` names the
127
127
  login the kernel now posts as (`<app-slug>[bot]`, e.g.
128
128
  `outerloop-autoresearch[bot]`), read by every role through one
129
129
  `github.bot_login_from_env()` default. Every own-comment filter, alarm-issue
@@ -135,10 +135,10 @@ Everything the kernel created BEFORE the flip carries the old login — the
135
135
  research-log issue, contract alarms, intake claims, its PRs. Recognition is
136
136
  keyed on login on purpose (the markers are public strings anyone can paste),
137
137
  so a flip must widen the set of our logins, not loosen the check:
138
- `AUTORESEARCH_BOT_ALIASES=agentic-learning-bot` names the former identity,
138
+ `OUTERLOOP_BOT_ALIASES=agentic-learning-bot` names the former identity,
139
139
  and every "is this ours" gate goes through `github.is_own_login`. The
140
140
  reusable review and verify workflows take the same value as a `bot_aliases`
141
- input (exported as `AUTORESEARCH_BOT_ALIASES`; the verifier's author gate
141
+ input (exported as `OUTERLOOP_BOT_ALIASES`; the verifier's author gate
142
142
  accepts an alias too), so a target repo passes `bot_login` = the App login and
143
143
  `bot_aliases` = the former account. Live lesson: without it, the first tick
144
144
  under the App claimed the kernel's own research-log issue as a research order.
@@ -61,7 +61,7 @@ monitoring agent + ~100 interventions; AIDE: in-process tree search). Ours is
61
61
  **decentralized on Slurm**: no resident daemon — a self-perpetuating chain of
62
62
  ~16 s stateless ticks; all state in typed records on the shared FS; every
63
63
  role (author session, eval, panel judge, wake) its own Slurm job — wake
64
- dispatch sits behind an operator flag (`AUTORESEARCH_DISPATCH_WAKE`, on in
64
+ dispatch sits behind an operator flag (`OUTERLOOP_DISPATCH_WAKE`, on in
65
65
  our deployment since 2026-08-20; retiring the dark-launch flag to
66
66
  default-on is a listed cleanup); crash or preemption anywhere heals
67
67
  through the sweep into honest endings.
@@ -34,8 +34,9 @@ non-interactive use):
34
34
  discovers the installation id itself (App JWT → `GET /app/installations`).
35
35
  A pre-existing PAT path is accepted as the alternative for orgs that
36
36
  prefer it.
37
- - **Collect the author key.** One path prompt per configured backend
38
- (`OUTERLOOP_HARNESS_KEY_FILE` / codex equivalent), mode-600 enforced.
37
+ - **Collect the author key.** A hidden paste for the configured backend,
38
+ written to `~/.config/outerloop/<backend>_key` (0600) and recorded as
39
+ `OUTERLOOP_<BACKEND>_KEY_FILE`; `--author-key-file` for an existing file.
39
40
  - **Write `~/.config/outerloop/.env`** — only keys the operator chose;
40
41
  the tick chain's allowlist is the contract for what matters.
41
42
  - **Doctor.** Re-run every check the tick already preflights, plus the
@@ -1,7 +1,7 @@
1
1
  # The resident tick
2
2
 
3
3
  *Design note, 2026-09-02. Status: built as the opt-in mode
4
- (`AUTORESEARCH_RESIDENT=1`, `scripts/tick_resident.sh`); the chain restarts of
4
+ (`OUTERLOOP_RESIDENT=1`, `scripts/tick_resident.sh`); the chain restarts of
5
5
  2026-09-02 are the evidence.*
6
6
 
7
7
  ## The problem
@@ -91,16 +91,16 @@ covers it.
91
91
 
92
92
  ## Rollout
93
93
 
94
- 1. Opt-in mode in `scripts/tick_chain.sbatch`: `AUTORESEARCH_RESIDENT=1` in
94
+ 1. Opt-in mode in `scripts/tick_chain.sbatch`: `OUTERLOOP_RESIDENT=1` in
95
95
  the chain's environment selects the loop; unset keeps the per-cadence
96
96
  chain. The walltime is fixed by Slurm before the script runs, so the
97
97
  resident chain is STARTED with an explicit walltime and its own job name:
98
98
 
99
99
  ```
100
- sbatch --time=360 --job-name=autoresearch-resident --dependency=singleton \
100
+ sbatch --time=360 --job-name=outerloop-resident --dependency=singleton \
101
101
  --account=… --partition=cpu_short \
102
- --export=ALL,AUTORESEARCH_RESIDENT=1,AUTORESEARCH_HOME=…,AUTORESEARCH_ROOT=…,\
103
- AUTORESEARCH_ACCOUNT=…,AUTORESEARCH_PARTITION=cpu_short,AUTORESEARCH_PAT_FILE=… \
102
+ --export=ALL,OUTERLOOP_RESIDENT=1,OUTERLOOP_HOME=…,OUTERLOOP_ROOT=…,\
103
+ OUTERLOOP_ACCOUNT=…,OUTERLOOP_PARTITION=cpu_short,OUTERLOOP_PAT_FILE=… \
104
104
  scripts/tick_chain.sbatch
105
105
  ```
106
106
 
@@ -334,3 +334,27 @@ the same infrastructure the fork-PR phase needs.
334
334
  Re-test on version bumps: whether Claude's `Read` opens `/proc`, a hermes
335
335
  read-only toolset, a codex sandbox that hides `/proc`, or a runner that grants
336
336
  PID namespaces would each rewrite this section.
337
+
338
+ ## Maintainers: how many review rounds
339
+
340
+ The advisory reviewer runs on open; a maintainer re-runs it after a fix by
341
+ removing and re-adding the `outerloop:review` label once the push has
342
+ settled (the workflow fires on the labeled event). Only the most recent round,
343
+ run against the head commit, counts; a quiet verdict on an older diff
344
+ authorizes nothing. Findings rejected on rationale get a reply on the thread.
345
+
346
+ Termination is judged, not literal, because an eager reviewer can always find
347
+ one more wording nit:
348
+
349
+ - code PRs iterate until a round yields no new medium-or-higher or
350
+ behavior-affecting finding; wording nits in an otherwise quiet round are
351
+ fixed or answered without another round;
352
+ - docs and process PRs get one round, its nits batched into one fix;
353
+ - hard cap of four rounds: still finding mediums by then means the change is
354
+ the problem; stop and rethink rather than cycle.
355
+
356
+ The habit exists because later rounds have repeatedly found real defects in
357
+ earlier rounds' own fixes. Bot-authored improvement PRs sit outside this gate:
358
+ the reviewer skips them by design, and their gate is the target repo's
359
+ required human review (the publish step arms auto-merge only when the target's
360
+ branch protection requires one).
@@ -40,7 +40,7 @@ The invariants, proven on #132/#133 and non-negotiable everywhere:
40
40
  in the same PR — never two ways to say the same thing.
41
41
  4. **Verbs are RoleSpec-gated.** The RoleSpec already caps each role's
42
42
  tools/scope/key; the CLI's live verbs are part of that cap. It is ONE tool
43
- (`.autoresearch/syscall`) whose verbs the brief exposes per role — an author
43
+ (`.outerloop/syscall`) whose verbs the brief exposes per role — an author
44
44
  uses `launch`/`sleep`, a judge uses `finding`/`conclude` — and a syscall is
45
45
  TYPED, so the kernel dispatches by type (a sleep parks + wakes; a verdict is
46
46
  read back). Every role runs the tool the same way: a shell in the jail. Roles
@@ -49,7 +49,7 @@ The invariants, proven on #132/#133 and non-negotiable everywhere:
49
49
 
50
50
  ## End-state
51
51
 
52
- One `.autoresearch/syscall` tool, its live verbs gated per role by RoleSpec:
52
+ One `.outerloop/syscall` tool, its live verbs gated per role by RoleSpec:
53
53
 
54
54
  | Role | Verbs | Win | Installed where |
55
55
  | --- | --- | --- | --- |
@@ -70,7 +70,7 @@ lands with tests, and deletes what it replaces.
70
70
  ### Phase 0 — finish Phase A (in flight, prerequisite)
71
71
 
72
72
  The wake for `author-sleep` parks: gather each launch's results from the run
73
- dir, deliver declared artifacts into `.autoresearch/results/<name>/`, resume
73
+ dir, deliver declared artifacts into `.outerloop/results/<name>/`, resume
74
74
  the **same session** through the climb's resume-entry (#129) with the results
75
75
  data-fenced, update the budget file. (The transitional `AUTHOR_SLEEP_WAKE_READY`
76
76
  constant and env-flag arming have since retired: enablement is contract-driven.) Plus
@@ -95,7 +95,7 @@ findings drafts the PR for a human. A submit costs the sleep it rides on.
95
95
 
96
96
  A verdict is not a second tool — it is a syscall TYPE on the one surface.
97
97
  Reviewer/verifier/panel stop emitting one schema-constrained final message.
98
- Instead: `python .autoresearch/syscall finding --file --line --confidence
98
+ Instead: `python .outerloop/syscall finding --file --line --confidence
99
99
  --summary --detail [--blocking] [--kind] [--category]` per finding, then
100
100
  `... conclude --notes ...` to commit — each call validated on the spot, the
101
101
  verdict assembled kernel-side (`type: "verdict"`), well-formed by construction.