outerloop-science 0.1.0.dev0__tar.gz → 0.1.0.dev2__tar.gz
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/.gitignore +3 -0
- {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/CHANGELOG.md +141 -0
- {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/CLAUDE.md +1 -1
- outerloop_science-0.1.0.dev2/CONTRIBUTING.md +89 -0
- {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/PKG-INFO +14 -10
- {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/README.md +11 -7
- outerloop_science-0.1.0.dev2/containers/README.md +8 -0
- outerloop_science-0.1.0.dev2/containers/agent-py312.def +43 -0
- {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/docs/design/architecture.md +6 -2
- {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/docs/design/dispatcher.md +19 -4
- {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/docs/design/github-app-auth.md +5 -5
- {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/docs/design/headline.md +1 -1
- {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/docs/design/onboarding.md +3 -2
- {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/docs/design/resident-tick.md +5 -5
- {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/docs/design/reviewer-infra.md +24 -0
- {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/docs/design/role-cli.md +4 -4
- {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/docs/install.md +75 -9
- {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/docs/roadmap.md +4 -1
- {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/docs/validation/author-syscalls.md +5 -5
- {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/pyproject.toml +14 -5
- {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/scripts/README.md +1 -1
- {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/scripts/install_codex.sh +2 -2
- {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/scripts/tick_chain.sbatch +41 -35
- {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/scripts/tick_deploy.sh +42 -38
- {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/scripts/tick_resident.sh +23 -20
- outerloop_science-0.1.0.dev2/src/outerloop/__init__.py +18 -0
- {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/src/outerloop/appauth.py +3 -3
- {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/src/outerloop/appmanifest.py +16 -11
- {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/src/outerloop/attempt.py +80 -52
- {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/src/outerloop/cli.py +170 -86
- {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/src/outerloop/climbboard.py +209 -4
- {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/src/outerloop/compute.py +84 -7
- {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/src/outerloop/contract.py +3 -2
- {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/src/outerloop/followup.py +37 -23
- {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/src/outerloop/github.py +4 -4
- {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/src/outerloop/harness.py +18 -7
- outerloop_science-0.1.0.dev2/src/outerloop/image.py +372 -0
- outerloop_science-0.1.0.dev2/src/outerloop/init.py +701 -0
- {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/src/outerloop/limits.py +1 -1
- {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/src/outerloop/measure.py +2 -2
- {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/src/outerloop/orchestrator.py +1 -1
- {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/src/outerloop/paths.py +14 -1
- {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/src/outerloop/review.py +1 -1
- {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/src/outerloop/steward.py +3 -3
- {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/src/outerloop/syscall.py +1 -0
- {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/src/outerloop/tick.py +400 -105
- outerloop_science-0.1.0.dev2/tests/conftest.py +50 -0
- {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/tests/test_attempt.py +42 -29
- {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/tests/test_bot_aliases.py +7 -7
- {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/tests/test_bot_login.py +5 -5
- {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/tests/test_climbboard.py +104 -0
- {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/tests/test_compute.py +64 -10
- outerloop_science-0.1.0.dev2/tests/test_env_bridge.py +86 -0
- {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/tests/test_followup.py +60 -0
- {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/tests/test_harness.py +35 -8
- outerloop_science-0.1.0.dev2/tests/test_image.py +294 -0
- outerloop_science-0.1.0.dev2/tests/test_init.py +763 -0
- {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/tests/test_local_compute.py +30 -6
- {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/tests/test_measure.py +1 -1
- {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/tests/test_start.py +214 -66
- {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/tests/test_sweep_git_locks.py +7 -0
- {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/tests/test_tick.py +569 -76
- {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/tests/test_tick_chain_successors.py +2 -2
- {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/tests/test_tick_resident.py +33 -32
- outerloop_science-0.1.0.dev2/tests/test_tiers.py +37 -0
- {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/tests/test_verify_agent.py +13 -11
- {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/uv.lock +140 -6
- outerloop_science-0.1.0.dev0/CONTRIBUTING.md +0 -75
- outerloop_science-0.1.0.dev0/src/outerloop/__init__.py +0 -18
- outerloop_science-0.1.0.dev0/src/outerloop/init.py +0 -313
- outerloop_science-0.1.0.dev0/tests/test_env_bridge.py +0 -61
- outerloop_science-0.1.0.dev0/tests/test_init.py +0 -242
- {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/.pre-commit-config.yaml +0 -0
- {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/.python-version +0 -0
- {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/LICENSE +0 -0
- {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/NOTICE +0 -0
- {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/RELEASING.md +0 -0
- {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/SECURITY.md +0 -0
- {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/docs/assets/icon-dark.svg +0 -0
- {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/docs/assets/icon-light.svg +0 -0
- {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/docs/assets/icon.svg +0 -0
- {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/docs/compute.md +0 -0
- {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/docs/contract.md +0 -0
- {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/docs/design/agent-substrate.md +0 -0
- {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/docs/design/consolidation.md +0 -0
- {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/docs/design/external.md +0 -0
- {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/docs/design/judge-placement.md +0 -0
- {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/docs/design/meta.md +0 -0
- {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/docs/design/orchestrator-verify.md +0 -0
- {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/docs/design/public-surface.md +0 -0
- {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/docs/design/research-lines.md +0 -0
- {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/docs/design/research-loop-buildout.md +0 -0
- {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/docs/design/research-loop.md +0 -0
- {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/docs/design/review-placement.md +0 -0
- {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/docs/design/roles.md +0 -0
- {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/docs/design/scaling.md +0 -0
- {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/docs/reviewer.md +0 -0
- {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/examples/review.yml +0 -0
- {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/scripts/install_hermes.sh +0 -0
- {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/scripts/requeue_moved_successors.sh +0 -0
- {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/scripts/setup_branch_protection.sh +0 -0
- {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/scripts/sweep_git_locks.sh +0 -0
- {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/src/outerloop/__main__.py +0 -0
- {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/src/outerloop/brief.py +0 -0
- {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/src/outerloop/contract_cli.py +0 -0
- {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/src/outerloop/disk.py +0 -0
- {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/src/outerloop/dispatch.py +0 -0
- {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/src/outerloop/housekeeping.py +0 -0
- {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/src/outerloop/intake.py +0 -0
- {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/src/outerloop/markers.py +0 -0
- {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/src/outerloop/panel.py +0 -0
- {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/src/outerloop/posting.py +0 -0
- {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/src/outerloop/progress.py +0 -0
- {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/src/outerloop/py.typed +0 -0
- {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/src/outerloop/review_agent.py +0 -0
- {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/src/outerloop/review_agent_cli.py +0 -0
- {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/src/outerloop/review_post_cli.py +0 -0
- {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/src/outerloop/review_summarize_cli.py +0 -0
- {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/src/outerloop/role_runner.py +0 -0
- {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/src/outerloop/roles.py +0 -0
- {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/src/outerloop/rolespec.py +0 -0
- {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/src/outerloop/runstate.py +0 -0
- {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/src/outerloop/style.py +0 -0
- {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/src/outerloop/syscall_cli.py +0 -0
- {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/src/outerloop/verifier.py +0 -0
- {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/src/outerloop/verify_agent.py +0 -0
- {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/src/outerloop/verify_agent_cli.py +0 -0
- {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/src/outerloop/verify_post_cli.py +0 -0
- {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/tests/test_appauth.py +0 -0
- {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/tests/test_appmanifest.py +0 -0
- {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/tests/test_brief.py +0 -0
- {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/tests/test_channel_dir.py +0 -0
- {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/tests/test_codex_harness.py +0 -0
- {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/tests/test_contract.py +0 -0
- {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/tests/test_contract_cli.py +0 -0
- {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/tests/test_contract_names.py +0 -0
- {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/tests/test_disk.py +0 -0
- {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/tests/test_dispatch.py +0 -0
- {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/tests/test_github.py +0 -0
- {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/tests/test_hardening.py +0 -0
- {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/tests/test_hermes_harness.py +0 -0
- {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/tests/test_housekeeping.py +0 -0
- {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/tests/test_import.py +0 -0
- {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/tests/test_intake.py +0 -0
- {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/tests/test_limits.py +0 -0
- {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/tests/test_markers.py +0 -0
- {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/tests/test_measure_and_decide.py +0 -0
- {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/tests/test_orchestrator.py +0 -0
- {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/tests/test_packaging.py +0 -0
- {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/tests/test_panel.py +0 -0
- {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/tests/test_paths.py +0 -0
- {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/tests/test_posting.py +0 -0
- {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/tests/test_progress.py +0 -0
- {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/tests/test_requeue_moved_successors.py +0 -0
- {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/tests/test_review.py +0 -0
- {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/tests/test_review_agent.py +0 -0
- {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/tests/test_review_agent_cli.py +0 -0
- {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/tests/test_review_hardening.py +0 -0
- {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/tests/test_review_policy.py +0 -0
- {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/tests/test_review_summarize.py +0 -0
- {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/tests/test_role_runner.py +0 -0
- {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/tests/test_rolespec.py +0 -0
- {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/tests/test_runstate.py +0 -0
- {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/tests/test_steward.py +0 -0
- {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/tests/test_syscall.py +0 -0
- {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/tests/test_syscall_cli.py +0 -0
- {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/tests/test_verifier.py +0 -0
- {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/tests/test_verify_agent_cli.py +0 -0
- {outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/tests/test_version.py +0 -0
|
@@ -6,6 +6,82 @@ Versions follow [SemVer](https://semver.org).
|
|
|
6
6
|
|
|
7
7
|
## [Unreleased]
|
|
8
8
|
|
|
9
|
+
### Added
|
|
10
|
+
|
|
11
|
+
- `climb/status.json` carries the fleet's queue: the kernel's own Slurm jobs
|
|
12
|
+
(tick chain, sessions, wakes, evals, launches), each attributed to an agent,
|
|
13
|
+
with state, elapsed time, partition and submit time. Jobs on the account
|
|
14
|
+
that are not the kernel's never appear. The strip republishes when a job
|
|
15
|
+
appears, leaves, or changes state, not on elapsed drift.
|
|
16
|
+
- `outerloop init` asks for the author's model API key (hidden) and writes it to
|
|
17
|
+
`~/.config/outerloop/<backend>_key` (0600), or takes `--author-key-file` for an
|
|
18
|
+
existing file (checked to exist, stored absolute); the `.env` records it as
|
|
19
|
+
`OUTERLOOP_<BACKEND>_KEY_FILE`. The key lands in the config dir in use:
|
|
20
|
+
`~/.config/outerloop/`, or `~/.config/autoresearch/` on a machine set up
|
|
21
|
+
before the rename. The focused `--github-app` run still asks nothing about
|
|
22
|
+
the author. Every secret `init` writes (keys, PAT, App PEM and JSON, `.env`)
|
|
23
|
+
is now created 0600 in one step rather than written and then tightened.
|
|
24
|
+
- The agent image is published: `containers/agent-py312.def` is its recipe
|
|
25
|
+
(Ubuntu 24.04, Python 3.12, uv, git, build tools), the `build-image` workflow
|
|
26
|
+
builds and uploads it with a checksum to huggingface.co/outerloop-science/
|
|
27
|
+
agent-image on every recipe change, and `outerloop init` downloads it on
|
|
28
|
+
Linux when Apptainer is installed, so local mode runs contained by default
|
|
29
|
+
(`--image` names your own, `--no-image` opts out).
|
|
30
|
+
- `outerloop init` checks that Apptainer can actually run a container before
|
|
31
|
+
downloading the image (on Ubuntu 24.04 a hand-installed Apptainer is on PATH
|
|
32
|
+
but cannot create user namespaces), and when it cannot, prints the install
|
|
33
|
+
steps for the machine's distribution and continues uncontained. The image
|
|
34
|
+
download shows a progress bar with size, speed and ETA on a terminal, one
|
|
35
|
+
line per 10% otherwise, and confirms the checksum. `docs/install.md` gains
|
|
36
|
+
an "Installing Apptainer" section.
|
|
37
|
+
- Launch admission: with `OUTERLOOP_MAX_LAUNCH_GPUS` set to the per-user GPU
|
|
38
|
+
cap, author launches are submitted held and the tick releases them
|
|
39
|
+
oldest-first while the user's GPU jobs fit under the cap, cancels held
|
|
40
|
+
launches of runs that have ended, and stops releasing when Slurm parks a
|
|
41
|
+
released launch on a per-user reason. `squeue` rows on the board carry the
|
|
42
|
+
pending reason and GRES. Unset, launches queue as before.
|
|
43
|
+
|
|
44
|
+
### Changed
|
|
45
|
+
|
|
46
|
+
- The board's queue is a card in the live strip, not a preformatted dump: a
|
|
47
|
+
full-width table with running jobs first, state pills, elapsed time,
|
|
48
|
+
partition (first of several, the rest counted, all on hover), and the
|
|
49
|
+
owning agent in its color; long job names are cut with an ellipsis and
|
|
50
|
+
shown whole on hover.
|
|
51
|
+
- The Claude author's key follows the same rule as Codex's: the file is
|
|
52
|
+
`~/.config/outerloop/claude_key` and the setting `OUTERLOOP_CLAUDE_KEY_FILE`.
|
|
53
|
+
The pre-rename `harness_key` file and `*_HARNESS_KEY_FILE` setting are still
|
|
54
|
+
read (the new name wins when both exist), so an existing setup keeps working.
|
|
55
|
+
- `AUTORESEARCH_TARGET` is required for in-review servicing (follow-ups, wakes,
|
|
56
|
+
the board). The kernel no longer falls back to the retired
|
|
57
|
+
`agentic-learning-ai-lab/autoresearch-pilot` repo when it is unset; the tick
|
|
58
|
+
logs the missing setting like any other.
|
|
59
|
+
- The test suite runs on all cores by default (pytest-xdist), about four times
|
|
60
|
+
faster locally; `pytest --testmon` runs only the tests affected by your edits
|
|
61
|
+
(pytest-testmon). Tier selection moved from `addopts` into `tests/conftest.py`
|
|
62
|
+
(a `-m` in `addopts` disables testmon), and a `serial` tier holds the tests
|
|
63
|
+
that inspect the process table: skipped while workers run, and `pytest -m
|
|
64
|
+
serial` turns workers off, so the tier always runs alone (CI's second step).
|
|
65
|
+
- The kernel reads `OUTERLOOP_*` everywhere: every internal read, log line,
|
|
66
|
+
chain script and test now uses the public names. The pre-rename
|
|
67
|
+
`AUTORESEARCH_*` names are still accepted for one release, bridged at the
|
|
68
|
+
process boundary, the `.env` file and the chain scripts, and will be
|
|
69
|
+
dropped in the release after 0.1. The review footer names `outerloop`, and
|
|
70
|
+
`outerloop --version` exists; `start`, `tick` and `init` describe themselves
|
|
71
|
+
and every flag in `--help`.
|
|
72
|
+
- The on-disk and queue names follow the rename: the resident tick job is
|
|
73
|
+
`outerloop-resident` (the per-cadence chain `outerloop-tick`), the local
|
|
74
|
+
state root defaults to `~/.outerloop`, the default image path is
|
|
75
|
+
`~/outerloop-images/agent-py312.sif`. `outerloop start` refuses while a
|
|
76
|
+
pre-rename `autoresearch-resident` is queued or running, since Slurm's
|
|
77
|
+
singleton serializes by name; cancel it first. An existing `~/.autoresearch`
|
|
78
|
+
or `~/autoresearch-images/` is used while the new path does not exist. A
|
|
79
|
+
chain-mode deployment (manual `tick_chain.sbatch`) must be stopped with
|
|
80
|
+
`scancel --name autoresearch-tick` before deploying this version. Every
|
|
81
|
+
honored pre-rename name (`~/.config/autoresearch/`, `.autoresearch.yaml`, the
|
|
82
|
+
`.autoresearch` channel dir, the old root, image and job names) is dropped in
|
|
83
|
+
the release after 0.1; the comment marker keeps recognizing old comments.
|
|
84
|
+
|
|
9
85
|
### Changed
|
|
10
86
|
|
|
11
87
|
- The syscall channel dir in a workspace is now `.outerloop/` (the tool, the
|
|
@@ -51,9 +127,74 @@ Versions follow [SemVer](https://semver.org).
|
|
|
51
127
|
|
|
52
128
|
### Fixed
|
|
53
129
|
|
|
130
|
+
- `outerloop init --github-app` says what to do at each step: which page
|
|
131
|
+
opens, which button to click, where the code appears, how to install the
|
|
132
|
+
App on the repository, and that the last step checks write access. It also
|
|
133
|
+
asks whether to create the App under your account or an organization,
|
|
134
|
+
which before needed the undocumented `--org` flag.
|
|
135
|
+
- The App flow's final write check asked GitHub a question installation
|
|
136
|
+
tokens cannot answer, so it warned about a missing push permission on Apps
|
|
137
|
+
that could push. It now checks the installation's repository list and its
|
|
138
|
+
granted permissions.
|
|
139
|
+
- `init --github-app` records the App's login (`<slug>[bot]`) as
|
|
140
|
+
`OUTERLOOP_BOT_LOGIN`. Without it the kernel fell back to a built-in
|
|
141
|
+
default that is not the adopter's identity and did not recognize its own
|
|
142
|
+
pull requests.
|
|
143
|
+
- The local loop launches attempts without a container image. On a machine
|
|
144
|
+
with no Apptainer image, `AUTORESEARCH_COMPUTE=local` now brings up
|
|
145
|
+
servicing with an empty image, logs once that sessions run under the
|
|
146
|
+
harness's own sandbox and evaluations run bare on that machine, and passes
|
|
147
|
+
`--uncontained` to its jobs. The panel is off in that mode unless
|
|
148
|
+
`OUTERLOOP_PANEL_UNCONTAINED=1` opts in, and says so when it runs. Slurm
|
|
149
|
+
still requires the image, and so does a Codex author in every mode (#289).
|
|
150
|
+
- The PyPI description no longer refers to the lab.
|
|
151
|
+
- A base-synced re-measurement is finished instead of abandoned. The parked
|
|
152
|
+
stage recorded the session's local merge of the moved base as its parent,
|
|
153
|
+
a commit GitHub never saw, so the resume judged the PR head as "moved" on
|
|
154
|
+
every base-synced park and threw the landed result away. The stage now
|
|
155
|
+
records the PR head at park and the resume compares against that. An
|
|
156
|
+
abandon also releases the wake-attempt cap, so the "ask again" it invites
|
|
157
|
+
can be served.
|
|
158
|
+
- A follow-up re-measurement that has landed is now finished even when the
|
|
159
|
+
run has spent its wake-attempt cap. The sessions that sync a stale PR onto
|
|
160
|
+
the moved base and dispatch the measurement each bill an attempt, so a
|
|
161
|
+
measurement could complete with nobody left to push its number; the run then
|
|
162
|
+
idled with "follow-up attempts without progress" every tick. The finishing
|
|
163
|
+
follow-up has its own allowance of `MAX_WAKE_ATTEMPTS`, counted on the
|
|
164
|
+
stage, so a session that fails to finish is retried a bounded number of
|
|
165
|
+
times; the cap still holds for follow-ups that have nothing landed to finish.
|
|
54
166
|
- `outerloop init` no longer overwrites an existing `~/.config/autoresearch/.env`
|
|
55
167
|
silently: interactively it asks; with `--yes` it refuses unless `--force` is
|
|
56
168
|
passed. The check runs before any GitHub App is created.
|
|
169
|
+
- Slurm account and partition are both optional (#300). `outerloop start`
|
|
170
|
+
no longer requires an account, `outerloop init` no longer insists on one,
|
|
171
|
+
the tick's in-review servicing and the attempt's resume and dispatch gates
|
|
172
|
+
need only the image, and every sbatch the kernel builds passes `--account`
|
|
173
|
+
and `--partition` only when set, so Slurm bills the caller's default
|
|
174
|
+
association and picks its default partition.
|
|
175
|
+
- A launch still queued past its deadline is no longer cancelled and woken as
|
|
176
|
+
"unschedulable" when Slurm's pending reason is a wait: a busy queue
|
|
177
|
+
(Priority, Resources), a reservation or dependency, or the account's own
|
|
178
|
+
per-user/per-account/group cap (`QOSMaxGRESPerUser` and kin). The sweep
|
|
179
|
+
extends the deadline by one slack window and logs the reasons; a queued job
|
|
180
|
+
keeps its priority age, where a re-launch would start over at the back.
|
|
181
|
+
Reasons that never clear (a dependency that cannot be satisfied, per-job
|
|
182
|
+
limits, an invalid account or QOS, a held job) still cancel and wake.
|
|
183
|
+
- `outerloop init` stops with exit 1 when the credential cannot open pull
|
|
184
|
+
requests on the target (401, 403, 404, or no write access) instead of
|
|
185
|
+
printing a warning and `next: outerloop start` (#285); a network failure stays
|
|
186
|
+
a warning. The PAT path records the token's login as `OUTERLOOP_BOT_LOGIN`,
|
|
187
|
+
and the tick warns when the login is unset and the lab default applies (#298).
|
|
188
|
+
- `init` locates the author's CLI (`claude`/`codex`) on PATH or in
|
|
189
|
+
`~/.local/bin`, records it as `OUTERLOOP_CLAUDE_BIN`/`OUTERLOOP_CODEX_BIN`, and
|
|
190
|
+
says how to install it when absent; `attempt --claude-bin` honors that
|
|
191
|
+
setting; `start` refuses to launch when a recorded binary is missing; a
|
|
192
|
+
`spawn-error` names the binary and the OS error (#294).
|
|
193
|
+
- Local jobs keep their combined stdout/stderr in `local_jobs/<id>.out` beside
|
|
194
|
+
the state file, and the log line names it when the job did not complete (#295).
|
|
195
|
+
- `cryptography` is a base dependency, so the GitHub App identity works from a
|
|
196
|
+
plain `pip install`; the `app-auth` extra is now a no-op (#293). The `tick`
|
|
197
|
+
subcommand's help no longer uses lab-internal vocabulary (#286).
|
|
57
198
|
|
|
58
199
|
### Added
|
|
59
200
|
|
|
@@ -8,7 +8,7 @@ Scaffold phase — `docs/roadmap.md` says what exists vs. planned;
|
|
|
8
8
|
|
|
9
9
|
```bash
|
|
10
10
|
uv sync
|
|
11
|
-
uv run pytest #
|
|
11
|
+
uv run pytest # all cores; --testmon = only tests affected by your edits
|
|
12
12
|
uv run ruff check --fix . && uv run ruff format .
|
|
13
13
|
uv run mypy
|
|
14
14
|
uv run pre-commit run --all-files
|
|
@@ -0,0 +1,89 @@
|
|
|
1
|
+
# Contributing
|
|
2
|
+
|
|
3
|
+
Outerloop is developed in the open by the [Agentic Learning AI Lab](https://agenticlearning.ai)
|
|
4
|
+
at NYU, and its own agents are regular contributors to it. Contributions from
|
|
5
|
+
people are welcome too: bug reports, fixes, documentation, new compute
|
|
6
|
+
backends, and benchmark contracts for your own repos.
|
|
7
|
+
|
|
8
|
+
## Before you start
|
|
9
|
+
|
|
10
|
+
- Look through the open issues. For anything larger than a small fix, open an
|
|
11
|
+
issue first so we can agree on the shape before you write code.
|
|
12
|
+
- Security problems go to the contact in [SECURITY.md](SECURITY.md), not to a
|
|
13
|
+
public issue.
|
|
14
|
+
|
|
15
|
+
## Setup
|
|
16
|
+
|
|
17
|
+
```bash
|
|
18
|
+
uv sync
|
|
19
|
+
uv run pre-commit install
|
|
20
|
+
uv run pytest
|
|
21
|
+
```
|
|
22
|
+
|
|
23
|
+
## Making a change
|
|
24
|
+
|
|
25
|
+
1. Branch from `main`: `fix/<topic>` or `feat/<topic>`.
|
|
26
|
+
2. Keep the change focused. Add or update tests, and a line under
|
|
27
|
+
`[Unreleased]` in `CHANGELOG.md` for anything a user would notice.
|
|
28
|
+
3. Run the same gate CI runs before you push:
|
|
29
|
+
|
|
30
|
+
```bash
|
|
31
|
+
uv run pre-commit run --all-files # lint, formatting, secret scan
|
|
32
|
+
uv lock --check
|
|
33
|
+
uv run mypy
|
|
34
|
+
uv run pytest && uv run pytest -m serial
|
|
35
|
+
```
|
|
36
|
+
|
|
37
|
+
4. Open a pull request against `main`. Say what changed, why, and how you
|
|
38
|
+
verified it.
|
|
39
|
+
|
|
40
|
+
## What happens to your pull request
|
|
41
|
+
|
|
42
|
+
- CI must pass.
|
|
43
|
+
- One of our agents reviews it and posts its findings as a review, usually
|
|
44
|
+
within the hour; no time is guaranteed. It never approves or blocks; a
|
|
45
|
+
maintainer decides. Fix what is right and reply to what is not; findings
|
|
46
|
+
answered with a reason are fine. After you push fixes, a maintainer removes
|
|
47
|
+
and re-adds the `outerloop:review` label to run another round.
|
|
48
|
+
- The agent review runs only for branches in this repository. A pull request
|
|
49
|
+
from a fork gets CI but no agent review; a maintainer can push your branch
|
|
50
|
+
here to run one.
|
|
51
|
+
- A maintainer merges with a merge commit. We do not squash or rebase, so the
|
|
52
|
+
history of how a change came to be stays intact. Keep your branch current
|
|
53
|
+
with `git merge main`, and never force-push a branch someone else has seen.
|
|
54
|
+
|
|
55
|
+
## Conventions
|
|
56
|
+
|
|
57
|
+
- Python 3.12, absolute imports (`from outerloop ...`), and code that works
|
|
58
|
+
from the installed wheel; no paths relative to a checkout.
|
|
59
|
+
- `ruff` for lint and formatting (line length 100); `mypy` clean.
|
|
60
|
+
- Dependencies go in the `pyproject.toml` group that owns them, then
|
|
61
|
+
`uv lock`; CI checks the lock.
|
|
62
|
+
- Docs are plain prose. Short sentences, no metaphors, no internal jargon.
|
|
63
|
+
|
|
64
|
+
## Tests
|
|
65
|
+
|
|
66
|
+
| Tier | Marker | Runs |
|
|
67
|
+
| --- | --- | --- |
|
|
68
|
+
| unit | *(none)* | CI, every PR, on all cores |
|
|
69
|
+
| serial | `serial` | CI, second step: `pytest -m serial` (inspects the process table) |
|
|
70
|
+
| slow | `slow` | nightly or manual |
|
|
71
|
+
| llm | `llm` | manual only, paid APIs |
|
|
72
|
+
| slurm | `slurm` | manual only, needs a cluster |
|
|
73
|
+
|
|
74
|
+
While editing, `uv run pytest --testmon` runs only the tests affected by your
|
|
75
|
+
change; `uv run pytest -n0` runs everything serially.
|
|
76
|
+
|
|
77
|
+
## Pull requests from the agents
|
|
78
|
+
|
|
79
|
+
Pull requests authored by `outerloop-science[bot]` are the system improving
|
|
80
|
+
the repositories it works on. They appear on those target repositories and are
|
|
81
|
+
skipped by the advisory reviewer by design. What gates them is each target's
|
|
82
|
+
own rules: its required human review, or, where a target's contract allows
|
|
83
|
+
automatic merging, the measurement gate and review panel that ran before the
|
|
84
|
+
PR opened. This document is about contributing to Outerloop, this repository.
|
|
85
|
+
|
|
86
|
+
## License
|
|
87
|
+
|
|
88
|
+
By contributing you agree that your contributions are licensed under the
|
|
89
|
+
[Apache License 2.0](LICENSE), the same as the project.
|
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
Metadata-Version: 2.5
|
|
2
2
|
Name: outerloop-science
|
|
3
|
-
Version: 0.1.0.
|
|
4
|
-
Summary: Autonomous research
|
|
3
|
+
Version: 0.1.0.dev2
|
|
4
|
+
Summary: Autonomous research agents that improve the benchmark you point them at, one verified pull request at a time
|
|
5
5
|
Project-URL: Homepage, https://outerloop.science
|
|
6
6
|
Project-URL: Repository, https://github.com/outerloop-science/outerloop
|
|
7
7
|
Project-URL: Changelog, https://github.com/outerloop-science/outerloop/blob/main/CHANGELOG.md
|
|
@@ -11,10 +11,10 @@ License-Expression: Apache-2.0
|
|
|
11
11
|
License-File: LICENSE
|
|
12
12
|
License-File: NOTICE
|
|
13
13
|
Requires-Python: >=3.12
|
|
14
|
+
Requires-Dist: cryptography>=42
|
|
14
15
|
Requires-Dist: pydantic>=2
|
|
15
16
|
Requires-Dist: pyyaml>=6
|
|
16
17
|
Provides-Extra: app-auth
|
|
17
|
-
Requires-Dist: cryptography>=42; extra == 'app-auth'
|
|
18
18
|
Description-Content-Type: text/markdown
|
|
19
19
|
|
|
20
20
|
<picture>
|
|
@@ -56,18 +56,22 @@ themselves.
|
|
|
56
56
|
|
|
57
57
|
## Get started
|
|
58
58
|
|
|
59
|
-
Three commands and
|
|
60
|
-
|
|
61
|
-
|
|
59
|
+
Three commands and one file. You need a repo with a benchmark command, an API
|
|
60
|
+
key for the model that will write the code, and a Slurm cluster or one machine
|
|
61
|
+
with a GPU.
|
|
62
62
|
|
|
63
63
|
```bash
|
|
64
64
|
pip install outerloop-science
|
|
65
|
-
outerloop init # where the loop runs, which repo,
|
|
65
|
+
outerloop init # where the loop runs, which repo, which model and its key, your GitHub identity
|
|
66
66
|
```
|
|
67
67
|
|
|
68
|
-
|
|
69
|
-
|
|
70
|
-
|
|
68
|
+
The wizard asks for a GitHub identity for the agents to open pull requests
|
|
69
|
+
as. Pick `app` and it walks you through creating a GitHub App, under your
|
|
70
|
+
account or under an organization you name, and installing it on the repo:
|
|
71
|
+
two browser pages and a code pasted back. It then checks that the App can
|
|
72
|
+
write the repo and tells you if it cannot. Pick `pat` if you already have a
|
|
73
|
+
token. It writes the config and the key files; nothing to edit by hand. Then
|
|
74
|
+
add one file, `.outerloop.yaml`, to the repo you want improved:
|
|
71
75
|
|
|
72
76
|
```yaml
|
|
73
77
|
benchmarks:
|
|
@@ -37,18 +37,22 @@ themselves.
|
|
|
37
37
|
|
|
38
38
|
## Get started
|
|
39
39
|
|
|
40
|
-
Three commands and
|
|
41
|
-
|
|
42
|
-
|
|
40
|
+
Three commands and one file. You need a repo with a benchmark command, an API
|
|
41
|
+
key for the model that will write the code, and a Slurm cluster or one machine
|
|
42
|
+
with a GPU.
|
|
43
43
|
|
|
44
44
|
```bash
|
|
45
45
|
pip install outerloop-science
|
|
46
|
-
outerloop init # where the loop runs, which repo,
|
|
46
|
+
outerloop init # where the loop runs, which repo, which model and its key, your GitHub identity
|
|
47
47
|
```
|
|
48
48
|
|
|
49
|
-
|
|
50
|
-
|
|
51
|
-
|
|
49
|
+
The wizard asks for a GitHub identity for the agents to open pull requests
|
|
50
|
+
as. Pick `app` and it walks you through creating a GitHub App, under your
|
|
51
|
+
account or under an organization you name, and installing it on the repo:
|
|
52
|
+
two browser pages and a code pasted back. It then checks that the App can
|
|
53
|
+
write the repo and tells you if it cannot. Pick `pat` if you already have a
|
|
54
|
+
token. It writes the config and the key files; nothing to edit by hand. Then
|
|
55
|
+
add one file, `.outerloop.yaml`, to the repo you want improved:
|
|
52
56
|
|
|
53
57
|
```yaml
|
|
54
58
|
benchmarks:
|
|
@@ -0,0 +1,8 @@
|
|
|
1
|
+
# containers
|
|
2
|
+
|
|
3
|
+
`agent-py312.def` is the Apptainer recipe for the image the kernel runs sessions and
|
|
4
|
+
evaluations in. The `build-image` workflow builds it on every change to this directory and
|
|
5
|
+
on manual dispatch, and publishes the result to
|
|
6
|
+
[huggingface.co/outerloop-science/agent-image](https://huggingface.co/outerloop-science/agent-image)
|
|
7
|
+
with a checksum. `outerloop init` downloads it on Linux when Apptainer is installed; the
|
|
8
|
+
tick reads `~/outerloop-images/agent-py312.sif` unless `OUTERLOOP_IMAGE` says otherwise.
|
|
@@ -0,0 +1,43 @@
|
|
|
1
|
+
# The Outerloop agent image: what a contained session and a dispatched
|
|
2
|
+
# evaluation see. The kernel binds the workspace, a per-run home and the
|
|
3
|
+
# harness binary (claude/codex) into it; --nv adds the host's GPU driver for
|
|
4
|
+
# GPU benchmarks. So the image only has to provide the base a target's own
|
|
5
|
+
# `uv run ...` needs: Python 3.12, uv, git and a compiler.
|
|
6
|
+
#
|
|
7
|
+
# Built by .github/workflows/build-image.yml and published to
|
|
8
|
+
# https://huggingface.co/outerloop-science/agent-image; `outerloop init` pulls
|
|
9
|
+
# it on Linux. Build by hand: apptainer build agent-py312.sif containers/agent-py312.def
|
|
10
|
+
|
|
11
|
+
Bootstrap: docker
|
|
12
|
+
From: ubuntu:24.04
|
|
13
|
+
|
|
14
|
+
%post
|
|
15
|
+
set -eu
|
|
16
|
+
export DEBIAN_FRONTEND=noninteractive
|
|
17
|
+
apt-get update
|
|
18
|
+
apt-get install -y --no-install-recommends \
|
|
19
|
+
ca-certificates curl git git-lfs openssh-client \
|
|
20
|
+
build-essential pkg-config \
|
|
21
|
+
python3 python3-venv python3-dev \
|
|
22
|
+
libgomp1 libglib2.0-0 procps less unzip zip jq ripgrep
|
|
23
|
+
rm -rf /var/lib/apt/lists/*
|
|
24
|
+
# uv installs its own Pythons when a target pins one; the system 3.12 is the default
|
|
25
|
+
curl -LsSf https://astral.sh/uv/install.sh | env UV_INSTALL_DIR=/usr/local/bin INSTALLER_NO_MODIFY_PATH=1 sh
|
|
26
|
+
ln -sf /usr/bin/python3 /usr/local/bin/python
|
|
27
|
+
git lfs install --system
|
|
28
|
+
python3 --version && uv --version && git --version
|
|
29
|
+
|
|
30
|
+
%environment
|
|
31
|
+
export LC_ALL=C.UTF-8
|
|
32
|
+
export LANG=C.UTF-8
|
|
33
|
+
export PATH=/usr/local/bin:/usr/local/sbin:/usr/sbin:/usr/bin:/sbin:/bin
|
|
34
|
+
|
|
35
|
+
%labels
|
|
36
|
+
org.opencontainers.image.title outerloop agent image
|
|
37
|
+
org.opencontainers.image.source https://github.com/outerloop-science/outerloop
|
|
38
|
+
outerloop.recipe containers/agent-py312.def
|
|
39
|
+
|
|
40
|
+
%help
|
|
41
|
+
Outerloop agent image: Ubuntu 24.04, Python 3.12, uv, git, build tools.
|
|
42
|
+
The kernel runs sessions and evaluations in it with --containall --cleanenv,
|
|
43
|
+
binding only the workspace, a per-run home and the harness binary.
|
|
@@ -371,8 +371,12 @@ failure can strand a run:
|
|
|
371
371
|
`deadline = submit_time + walltime + slack` (recomputed from `start_time`
|
|
372
372
|
once the job starts, so late scheduling never truncates a healthy run).
|
|
373
373
|
Past the deadline the sweep consults `sacct` and acts on what it finds:
|
|
374
|
-
still PENDING →
|
|
375
|
-
|
|
374
|
+
still PENDING → ask squeue why: a wait (Priority/Resources, a reservation
|
|
375
|
+
or dependency, the account's own per-user/group cap such as
|
|
376
|
+
`QOSMaxGRESPerUser`) moves the deadline out by one slack window, since a
|
|
377
|
+
queued job keeps its priority age and a re-launch would start over; any
|
|
378
|
+
other reason means unschedulable in practice — cancel it and wake the
|
|
379
|
+
run with that fact. No record on a *successful* query → the job
|
|
376
380
|
vanished; wake with that fact. Query failed or timed out → that is
|
|
377
381
|
"Slurm unknown", never "job gone" — defer to the next tick, and only a
|
|
378
382
|
sustained outage (its own alert via the watchdog) escalates. A fail-safe
|
|
@@ -66,7 +66,7 @@ orchestrator chooses per eval site, not per repo.
|
|
|
66
66
|
launches alike: the JobSpec requests `--gpus=N`, and the jail adds `--nv` so
|
|
67
67
|
the allocation is visible inside the container; nothing else about the
|
|
68
68
|
containment changes. GPU jobs need their own lane, so the deployment names
|
|
69
|
-
one — `
|
|
69
|
+
one — `OUTERLOOP_GPU_PARTITION`, optionally `OUTERLOOP_GPU_ACCOUNT`
|
|
70
70
|
(default: the CPU account) — and a benchmark with `gpus > 0` on a deployment
|
|
71
71
|
without a lane is REFUSED at launch (both attempt lanes), never queued into
|
|
72
72
|
evals that can never run. GPUs exist only on DISPATCHED jobs: the contract
|
|
@@ -76,6 +76,18 @@ benchmarks outright for now — a stewardship validates its rewrite in-job.
|
|
|
76
76
|
The first target in this shape is the speedrun (`gpt-speedrun`: one H200
|
|
77
77
|
per eval, ~3.5h).
|
|
78
78
|
|
|
79
|
+
**Launch admission.** A per-user GPU cap (a QOS's `MaxTRESPerUser`, 16 on
|
|
80
|
+
Torch) means the fleet can run only so many launches at once, whatever the
|
|
81
|
+
agent count: six 8-GPU launches queued against a 16-GPU cap left four pending
|
|
82
|
+
for a day on `QOSMaxGRESPerUser`. With `OUTERLOOP_MAX_LAUNCH_GPUS` set to that
|
|
83
|
+
cap, author launches are submitted HELD, and every tick `service_admission`
|
|
84
|
+
releases them oldest-first while the user's eligible GPU jobs (evals and
|
|
85
|
+
launches alike) fit under the cap; held launches whose run has since ended
|
|
86
|
+
are cancelled. A released launch that Slurm parks on a per-user reason means
|
|
87
|
+
the cap is set too high, and nothing more is released that tick. The wake's
|
|
88
|
+
`afterany` dependency is unchanged (held jobs are pending), and the sweep
|
|
89
|
+
treats the kernel's hold as a wait. Unset, launches queue as submitted.
|
|
90
|
+
|
|
79
91
|
**Who decides the eval walltime.** The gate measures steps, not time, so
|
|
80
92
|
walltime should bound only SPEND — and the attempt already has a spend
|
|
81
93
|
budget. The contract's `eval_minutes` is the DEFAULT; the author may declare
|
|
@@ -234,9 +246,12 @@ dies -> the `eval-error` outcome at wake (ending `aborted`, exactly as the
|
|
|
234
246
|
in-job eval failure maps today); a LOST wake -> the sweep's primary backup,
|
|
235
247
|
which wakes any run whose job reads terminal after the grace period —
|
|
236
248
|
plus GONE past the deadline (vanished) and PENDING past the deadline
|
|
237
|
-
(
|
|
249
|
+
(when Slurm's reason is a wait — a busy queue, a reservation or dependency,
|
|
250
|
+
the account's own per-user/group cap — the deadline moves out by one slack
|
|
251
|
+
window; otherwise cancel, then wake as unschedulable); a still-RUNNING job is deliberately
|
|
238
252
|
left alone, bounded by its own walltime; wakes that fire without producing
|
|
239
|
-
progress -> `stuck` at MAX_WAKE_ATTEMPTS
|
|
253
|
+
progress -> `stuck` at MAX_WAKE_ATTEMPTS (a landed re-measure has its own
|
|
254
|
+
allowance of MAX_WAKE_ATTEMPTS finishing follow-ups, counted on the stage). A moved base during the wait is
|
|
240
255
|
review's to handle — the sealed candidate publishes as-is (research-loop.md:
|
|
241
256
|
a stale PR is a re-wake, never an orchestrator auto-merge). Worktree cleanup has a named owner at every exit: the
|
|
242
257
|
wake's own finally (primary, as in-job measures do today), the sweep's
|
|
@@ -286,7 +301,7 @@ GPU benchmarks and portfolio climbs both multiply eval load; both were
|
|
|
286
301
|
designed against this seam. Portfolio's N concurrent climbs become N
|
|
287
302
|
waiting records with independent wakes — the serialization question stays
|
|
288
303
|
in the picker, not the dispatcher. The tick's job-partition knobs
|
|
289
|
-
(`
|
|
304
|
+
(`OUTERLOOP_JOB_PARTITION`) already route work jobs; per-benchmark
|
|
290
305
|
partition/GPU allocation fields are a phase-3 contract-schema change
|
|
291
306
|
(`Benchmark` rejects unknown fields today, deliberately), validated and
|
|
292
307
|
clamped by operator ceilings like every other budget.
|
{outerloop_science-0.1.0.dev0 → outerloop_science-0.1.0.dev2}/docs/design/github-app-auth.md
RENAMED
|
@@ -108,7 +108,7 @@ Swapping the live fleet's identity mid-campaign is an ops event:
|
|
|
108
108
|
|
|
109
109
|
## The flag
|
|
110
110
|
|
|
111
|
-
`
|
|
111
|
+
`OUTERLOOP_GITHUB_APP_FILE` names a JSON config —
|
|
112
112
|
|
|
113
113
|
```json
|
|
114
114
|
{"app_id": 1234, "installation_id": 5678,
|
|
@@ -117,13 +117,13 @@ Swapping the live fleet's identity mid-campaign is an ops event:
|
|
|
117
117
|
|
|
118
118
|
— and rides the rails the PAT already rides: role CLIs default their
|
|
119
119
|
`--github-app-file` from it (jobs inherit the tick's environment, the same
|
|
120
|
-
way `
|
|
120
|
+
way `OUTERLOOP_AUTHOR_*` reaches them), and every bot-auth construction
|
|
121
121
|
goes through one factory, `appauth.resolve_bot_auth(pat_file, app_file)`:
|
|
122
122
|
App provider when the config is set, PAT otherwise. Revert = unset the env
|
|
123
123
|
var. The ids are not secrets; the key path inside keeps PAT-file custody
|
|
124
124
|
(600, never committed).
|
|
125
125
|
|
|
126
|
-
The identity flips WITH the credential: `
|
|
126
|
+
The identity flips WITH the credential: `OUTERLOOP_BOT_LOGIN` names the
|
|
127
127
|
login the kernel now posts as (`<app-slug>[bot]`, e.g.
|
|
128
128
|
`outerloop-autoresearch[bot]`), read by every role through one
|
|
129
129
|
`github.bot_login_from_env()` default. Every own-comment filter, alarm-issue
|
|
@@ -135,10 +135,10 @@ Everything the kernel created BEFORE the flip carries the old login — the
|
|
|
135
135
|
research-log issue, contract alarms, intake claims, its PRs. Recognition is
|
|
136
136
|
keyed on login on purpose (the markers are public strings anyone can paste),
|
|
137
137
|
so a flip must widen the set of our logins, not loosen the check:
|
|
138
|
-
`
|
|
138
|
+
`OUTERLOOP_BOT_ALIASES=agentic-learning-bot` names the former identity,
|
|
139
139
|
and every "is this ours" gate goes through `github.is_own_login`. The
|
|
140
140
|
reusable review and verify workflows take the same value as a `bot_aliases`
|
|
141
|
-
input (exported as `
|
|
141
|
+
input (exported as `OUTERLOOP_BOT_ALIASES`; the verifier's author gate
|
|
142
142
|
accepts an alias too), so a target repo passes `bot_login` = the App login and
|
|
143
143
|
`bot_aliases` = the former account. Live lesson: without it, the first tick
|
|
144
144
|
under the App claimed the kernel's own research-log issue as a research order.
|
|
@@ -61,7 +61,7 @@ monitoring agent + ~100 interventions; AIDE: in-process tree search). Ours is
|
|
|
61
61
|
**decentralized on Slurm**: no resident daemon — a self-perpetuating chain of
|
|
62
62
|
~16 s stateless ticks; all state in typed records on the shared FS; every
|
|
63
63
|
role (author session, eval, panel judge, wake) its own Slurm job — wake
|
|
64
|
-
dispatch sits behind an operator flag (`
|
|
64
|
+
dispatch sits behind an operator flag (`OUTERLOOP_DISPATCH_WAKE`, on in
|
|
65
65
|
our deployment since 2026-08-20; retiring the dark-launch flag to
|
|
66
66
|
default-on is a listed cleanup); crash or preemption anywhere heals
|
|
67
67
|
through the sweep into honest endings.
|
|
@@ -34,8 +34,9 @@ non-interactive use):
|
|
|
34
34
|
discovers the installation id itself (App JWT → `GET /app/installations`).
|
|
35
35
|
A pre-existing PAT path is accepted as the alternative for orgs that
|
|
36
36
|
prefer it.
|
|
37
|
-
- **Collect the author key.**
|
|
38
|
-
|
|
37
|
+
- **Collect the author key.** A hidden paste for the configured backend,
|
|
38
|
+
written to `~/.config/outerloop/<backend>_key` (0600) and recorded as
|
|
39
|
+
`OUTERLOOP_<BACKEND>_KEY_FILE`; `--author-key-file` for an existing file.
|
|
39
40
|
- **Write `~/.config/outerloop/.env`** — only keys the operator chose;
|
|
40
41
|
the tick chain's allowlist is the contract for what matters.
|
|
41
42
|
- **Doctor.** Re-run every check the tick already preflights, plus the
|
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
# The resident tick
|
|
2
2
|
|
|
3
3
|
*Design note, 2026-09-02. Status: built as the opt-in mode
|
|
4
|
-
(`
|
|
4
|
+
(`OUTERLOOP_RESIDENT=1`, `scripts/tick_resident.sh`); the chain restarts of
|
|
5
5
|
2026-09-02 are the evidence.*
|
|
6
6
|
|
|
7
7
|
## The problem
|
|
@@ -91,16 +91,16 @@ covers it.
|
|
|
91
91
|
|
|
92
92
|
## Rollout
|
|
93
93
|
|
|
94
|
-
1. Opt-in mode in `scripts/tick_chain.sbatch`: `
|
|
94
|
+
1. Opt-in mode in `scripts/tick_chain.sbatch`: `OUTERLOOP_RESIDENT=1` in
|
|
95
95
|
the chain's environment selects the loop; unset keeps the per-cadence
|
|
96
96
|
chain. The walltime is fixed by Slurm before the script runs, so the
|
|
97
97
|
resident chain is STARTED with an explicit walltime and its own job name:
|
|
98
98
|
|
|
99
99
|
```
|
|
100
|
-
sbatch --time=360 --job-name=
|
|
100
|
+
sbatch --time=360 --job-name=outerloop-resident --dependency=singleton \
|
|
101
101
|
--account=… --partition=cpu_short \
|
|
102
|
-
--export=ALL,
|
|
103
|
-
|
|
102
|
+
--export=ALL,OUTERLOOP_RESIDENT=1,OUTERLOOP_HOME=…,OUTERLOOP_ROOT=…,\
|
|
103
|
+
OUTERLOOP_ACCOUNT=…,OUTERLOOP_PARTITION=cpu_short,OUTERLOOP_PAT_FILE=… \
|
|
104
104
|
scripts/tick_chain.sbatch
|
|
105
105
|
```
|
|
106
106
|
|
|
@@ -334,3 +334,27 @@ the same infrastructure the fork-PR phase needs.
|
|
|
334
334
|
Re-test on version bumps: whether Claude's `Read` opens `/proc`, a hermes
|
|
335
335
|
read-only toolset, a codex sandbox that hides `/proc`, or a runner that grants
|
|
336
336
|
PID namespaces would each rewrite this section.
|
|
337
|
+
|
|
338
|
+
## Maintainers: how many review rounds
|
|
339
|
+
|
|
340
|
+
The advisory reviewer runs on open; a maintainer re-runs it after a fix by
|
|
341
|
+
removing and re-adding the `outerloop:review` label once the push has
|
|
342
|
+
settled (the workflow fires on the labeled event). Only the most recent round,
|
|
343
|
+
run against the head commit, counts; a quiet verdict on an older diff
|
|
344
|
+
authorizes nothing. Findings rejected on rationale get a reply on the thread.
|
|
345
|
+
|
|
346
|
+
Termination is judged, not literal, because an eager reviewer can always find
|
|
347
|
+
one more wording nit:
|
|
348
|
+
|
|
349
|
+
- code PRs iterate until a round yields no new medium-or-higher or
|
|
350
|
+
behavior-affecting finding; wording nits in an otherwise quiet round are
|
|
351
|
+
fixed or answered without another round;
|
|
352
|
+
- docs and process PRs get one round, its nits batched into one fix;
|
|
353
|
+
- hard cap of four rounds: still finding mediums by then means the change is
|
|
354
|
+
the problem; stop and rethink rather than cycle.
|
|
355
|
+
|
|
356
|
+
The habit exists because later rounds have repeatedly found real defects in
|
|
357
|
+
earlier rounds' own fixes. Bot-authored improvement PRs sit outside this gate:
|
|
358
|
+
the reviewer skips them by design, and their gate is the target repo's
|
|
359
|
+
required human review (the publish step arms auto-merge only when the target's
|
|
360
|
+
branch protection requires one).
|
|
@@ -40,7 +40,7 @@ The invariants, proven on #132/#133 and non-negotiable everywhere:
|
|
|
40
40
|
in the same PR — never two ways to say the same thing.
|
|
41
41
|
4. **Verbs are RoleSpec-gated.** The RoleSpec already caps each role's
|
|
42
42
|
tools/scope/key; the CLI's live verbs are part of that cap. It is ONE tool
|
|
43
|
-
(`.
|
|
43
|
+
(`.outerloop/syscall`) whose verbs the brief exposes per role — an author
|
|
44
44
|
uses `launch`/`sleep`, a judge uses `finding`/`conclude` — and a syscall is
|
|
45
45
|
TYPED, so the kernel dispatches by type (a sleep parks + wakes; a verdict is
|
|
46
46
|
read back). Every role runs the tool the same way: a shell in the jail. Roles
|
|
@@ -49,7 +49,7 @@ The invariants, proven on #132/#133 and non-negotiable everywhere:
|
|
|
49
49
|
|
|
50
50
|
## End-state
|
|
51
51
|
|
|
52
|
-
One `.
|
|
52
|
+
One `.outerloop/syscall` tool, its live verbs gated per role by RoleSpec:
|
|
53
53
|
|
|
54
54
|
| Role | Verbs | Win | Installed where |
|
|
55
55
|
| --- | --- | --- | --- |
|
|
@@ -70,7 +70,7 @@ lands with tests, and deletes what it replaces.
|
|
|
70
70
|
### Phase 0 — finish Phase A (in flight, prerequisite)
|
|
71
71
|
|
|
72
72
|
The wake for `author-sleep` parks: gather each launch's results from the run
|
|
73
|
-
dir, deliver declared artifacts into `.
|
|
73
|
+
dir, deliver declared artifacts into `.outerloop/results/<name>/`, resume
|
|
74
74
|
the **same session** through the climb's resume-entry (#129) with the results
|
|
75
75
|
data-fenced, update the budget file. (The transitional `AUTHOR_SLEEP_WAKE_READY`
|
|
76
76
|
constant and env-flag arming have since retired: enablement is contract-driven.) Plus
|
|
@@ -95,7 +95,7 @@ findings drafts the PR for a human. A submit costs the sleep it rides on.
|
|
|
95
95
|
|
|
96
96
|
A verdict is not a second tool — it is a syscall TYPE on the one surface.
|
|
97
97
|
Reviewer/verifier/panel stop emitting one schema-constrained final message.
|
|
98
|
-
Instead: `python .
|
|
98
|
+
Instead: `python .outerloop/syscall finding --file --line --confidence
|
|
99
99
|
--summary --detail [--blocking] [--kind] [--category]` per finding, then
|
|
100
100
|
`... conclude --notes ...` to commit — each call validated on the spot, the
|
|
101
101
|
verdict assembled kernel-side (`type: "verdict"`), well-formed by construction.
|