nanoPyCodeAgent 0.7.0__tar.gz → 0.9.0__tar.gz
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- {nanopycodeagent-0.7.0 → nanopycodeagent-0.9.0}/.agents/skills/release/scripts/verify-release.sh +5 -5
- {nanopycodeagent-0.7.0 → nanopycodeagent-0.9.0}/.github/workflows/ci.yml +18 -0
- {nanopycodeagent-0.7.0 → nanopycodeagent-0.9.0}/.gitignore +3 -0
- {nanopycodeagent-0.7.0 → nanopycodeagent-0.9.0}/AGENTS.md +45 -1
- nanopycodeagent-0.9.0/CONTEXT.md +65 -0
- {nanopycodeagent-0.7.0 → nanopycodeagent-0.9.0}/PKG-INFO +21 -24
- {nanopycodeagent-0.7.0 → nanopycodeagent-0.9.0}/README.md +19 -23
- {nanopycodeagent-0.7.0 → nanopycodeagent-0.9.0}/README.zh-CN.md +15 -22
- nanopycodeagent-0.9.0/benchmarks/harbor/README.md +196 -0
- nanopycodeagent-0.9.0/benchmarks/harbor/README.zh-CN.md +158 -0
- nanopycodeagent-0.9.0/benchmarks/harbor/configs/qemu-alpine-ssh.json +21 -0
- nanopycodeagent-0.9.0/benchmarks/harbor/pyproject.toml +27 -0
- nanopycodeagent-0.9.0/benchmarks/harbor/reports/deepseek-official-20260930.md +86 -0
- nanopycodeagent-0.9.0/benchmarks/harbor/reports/dual-budget-20260929.md +246 -0
- nanopycodeagent-0.9.0/benchmarks/harbor/reports/stream-recovery-20260915.md +139 -0
- nanopycodeagent-0.9.0/benchmarks/harbor/results/tb21-deepseek-official-pilot20-20260930.json +387 -0
- nanopycodeagent-0.9.0/benchmarks/harbor/results/tb21-dualbudget-fixed-provider-20260929.json +1670 -0
- nanopycodeagent-0.9.0/benchmarks/harbor/results/tb21-exp1-turn100-20260926.json +120 -0
- nanopycodeagent-0.9.0/benchmarks/harbor/results/tb21-exp2-timebudget-20260926.json +127 -0
- nanopycodeagent-0.9.0/benchmarks/harbor/results/tb21-exp3-timebudget-tail-20260926.json +115 -0
- nanopycodeagent-0.9.0/benchmarks/harbor/results/tb21-flash0731-pilot20-20260906.json +1477 -0
- nanopycodeagent-0.9.0/benchmarks/harbor/results/tb21-generation-budget-65536-20260912.json +1449 -0
- nanopycodeagent-0.9.0/benchmarks/harbor/results/tb21-tool-input-recovery-20260914.json +3541 -0
- nanopycodeagent-0.9.0/benchmarks/harbor/results/tb21-v41flash-6task-20260925.json +125 -0
- nanopycodeagent-0.9.0/benchmarks/harbor/src/harbor_adapter/__init__.py +6 -0
- nanopycodeagent-0.9.0/benchmarks/harbor/src/harbor_adapter/__main__.py +182 -0
- nanopycodeagent-0.9.0/benchmarks/harbor/src/harbor_adapter/_verifier_runner.py +143 -0
- nanopycodeagent-0.9.0/benchmarks/harbor/src/harbor_adapter/adapter.py +332 -0
- nanopycodeagent-0.9.0/benchmarks/harbor/src/harbor_adapter/bootstrap.py +183 -0
- nanopycodeagent-0.9.0/benchmarks/harbor/src/harbor_adapter/verifier.py +150 -0
- nanopycodeagent-0.9.0/benchmarks/harbor/tests/fixtures/atif-journal-v1.jsonl +7 -0
- nanopycodeagent-0.9.0/benchmarks/harbor/tests/test_adapter.py +355 -0
- nanopycodeagent-0.9.0/benchmarks/harbor/tests/test_atif_compatibility.py +157 -0
- nanopycodeagent-0.9.0/benchmarks/harbor/tests/test_bootstrap.py +229 -0
- nanopycodeagent-0.9.0/benchmarks/harbor/tests/test_verifier.py +324 -0
- nanopycodeagent-0.9.0/benchmarks/harbor/tests/test_workflow.py +90 -0
- nanopycodeagent-0.9.0/benchmarks/harbor/uv.lock +2059 -0
- nanopycodeagent-0.9.0/docs/changelogs/0.8.x.md +67 -0
- nanopycodeagent-0.9.0/docs/changelogs/0.9.x.md +81 -0
- {nanopycodeagent-0.7.0 → nanopycodeagent-0.9.0}/docs/changelogs/README.md +3 -0
- nanopycodeagent-0.9.0/docs/dev_docs/README.md +41 -0
- nanopycodeagent-0.9.0/docs/dev_docs/en/event-journal-protocol-v1.md +549 -0
- nanopycodeagent-0.9.0/docs/dev_docs/en/event-journal-protocol-v2.md +81 -0
- nanopycodeagent-0.9.0/docs/dev_docs/en/event-journal-protocol-v3.md +61 -0
- nanopycodeagent-0.9.0/docs/dev_docs/en/event-journal-protocol-v4.md +102 -0
- nanopycodeagent-0.9.0/docs/dev_docs/zh-CN/event-journal-protocol-v1.md +452 -0
- nanopycodeagent-0.9.0/docs/dev_docs/zh-CN/event-journal-protocol-v2.md +61 -0
- nanopycodeagent-0.9.0/docs/dev_docs/zh-CN/event-journal-protocol-v3.md +50 -0
- nanopycodeagent-0.9.0/docs/dev_docs/zh-CN/event-journal-protocol-v4.md +80 -0
- {nanopycodeagent-0.7.0 → nanopycodeagent-0.9.0}/docs/dev_notes/README.md +6 -1
- nanopycodeagent-0.9.0/docs/dev_notes/en/0.8.x.md +293 -0
- nanopycodeagent-0.9.0/docs/dev_notes/en/0.9.x.md +1076 -0
- nanopycodeagent-0.9.0/docs/dev_notes/zh-CN/0.8.x.md +295 -0
- nanopycodeagent-0.9.0/docs/dev_notes/zh-CN/0.9.x.md +1451 -0
- {nanopycodeagent-0.7.0 → nanopycodeagent-0.9.0}/docs/research/README.md +10 -1
- nanopycodeagent-0.9.0/docs/research/en/agent_events_to_atif_examples.md +727 -0
- nanopycodeagent-0.9.0/docs/research/en/agent_output_and_trajectory.md +755 -0
- nanopycodeagent-0.9.0/docs/research/en/benchmark_headless_interface.md +259 -0
- nanopycodeagent-0.9.0/docs/research/en/code_agent_benchmark.md +513 -0
- nanopycodeagent-0.9.0/docs/research/en/openrouter_cost_accounting.md +335 -0
- nanopycodeagent-0.9.0/docs/research/en/openrouter_unified_protocol.md +392 -0
- nanopycodeagent-0.9.0/docs/research/zh-CN/agent_events_to_atif_examples.md +727 -0
- nanopycodeagent-0.9.0/docs/research/zh-CN/agent_output_and_trajectory.md +755 -0
- nanopycodeagent-0.9.0/docs/research/zh-CN/benchmark_headless_interface.md +259 -0
- nanopycodeagent-0.9.0/docs/research/zh-CN/code_agent_benchmark.md +513 -0
- nanopycodeagent-0.9.0/docs/research/zh-CN/openrouter_cost_accounting.md +335 -0
- nanopycodeagent-0.9.0/docs/research/zh-CN/openrouter_unified_protocol.md +392 -0
- nanopycodeagent-0.9.0/docs/user_docs/README.md +10 -0
- nanopycodeagent-0.9.0/docs/user_docs/en/cli_reference.md +210 -0
- nanopycodeagent-0.9.0/docs/user_docs/en/configuration.md +184 -0
- nanopycodeagent-0.9.0/docs/user_docs/zh-CN/cli_reference.md +176 -0
- nanopycodeagent-0.9.0/docs/user_docs/zh-CN/configuration.md +165 -0
- {nanopycodeagent-0.7.0 → nanopycodeagent-0.9.0}/pyproject.toml +1 -0
- nanopycodeagent-0.9.0/src/nanopycodeagent/__init__.py +3 -0
- nanopycodeagent-0.9.0/src/nanopycodeagent/agent.py +1006 -0
- nanopycodeagent-0.9.0/src/nanopycodeagent/atif.py +591 -0
- nanopycodeagent-0.9.0/src/nanopycodeagent/bash_tool.py +130 -0
- nanopycodeagent-0.9.0/src/nanopycodeagent/cli.py +172 -0
- nanopycodeagent-0.9.0/src/nanopycodeagent/cost.py +211 -0
- nanopycodeagent-0.9.0/src/nanopycodeagent/deadline.py +39 -0
- nanopycodeagent-0.9.0/src/nanopycodeagent/event_journal.py +809 -0
- {nanopycodeagent-0.7.0 → nanopycodeagent-0.9.0}/src/nanopycodeagent/settings.py +20 -0
- nanopycodeagent-0.9.0/src/nanopycodeagent/tool_validation.py +40 -0
- nanopycodeagent-0.9.0/src/nanopycodeagent/transport.py +18 -0
- {nanopycodeagent-0.7.0 → nanopycodeagent-0.9.0}/tests/conftest.py +4 -1
- {nanopycodeagent-0.7.0 → nanopycodeagent-0.9.0}/tests/helpers.py +55 -8
- {nanopycodeagent-0.7.0 → nanopycodeagent-0.9.0}/tests/test_agent.py +6 -5
- nanopycodeagent-0.9.0/tests/test_agent_events.py +455 -0
- nanopycodeagent-0.9.0/tests/test_atif.py +611 -0
- nanopycodeagent-0.9.0/tests/test_bash_tool.py +120 -0
- nanopycodeagent-0.9.0/tests/test_cli.py +367 -0
- nanopycodeagent-0.9.0/tests/test_cost.py +201 -0
- nanopycodeagent-0.9.0/tests/test_event_journal.py +541 -0
- nanopycodeagent-0.9.0/tests/test_generation_budget.py +117 -0
- nanopycodeagent-0.9.0/tests/test_stream_recovery.py +226 -0
- nanopycodeagent-0.9.0/tests/test_time_budget.py +377 -0
- nanopycodeagent-0.9.0/tests/test_tool_validation.py +228 -0
- nanopycodeagent-0.9.0/tests/test_truncation.py +200 -0
- {nanopycodeagent-0.7.0 → nanopycodeagent-0.9.0}/uv.lock +5 -1
- nanopycodeagent-0.7.0/src/nanopycodeagent/__init__.py +0 -5
- nanopycodeagent-0.7.0/src/nanopycodeagent/agent.py +0 -194
- nanopycodeagent-0.7.0/src/nanopycodeagent/bash_tool.py +0 -77
- nanopycodeagent-0.7.0/tests/test_bash_tool.py +0 -62
- {nanopycodeagent-0.7.0 → nanopycodeagent-0.9.0}/.agents/skills/land-pr/SKILL.md +0 -0
- {nanopycodeagent-0.7.0 → nanopycodeagent-0.9.0}/.agents/skills/release/SKILL.md +0 -0
- {nanopycodeagent-0.7.0 → nanopycodeagent-0.9.0}/.github/workflows/release.yml +0 -0
- {nanopycodeagent-0.7.0 → nanopycodeagent-0.9.0}/.ignore +0 -0
- {nanopycodeagent-0.7.0 → nanopycodeagent-0.9.0}/.python-version +0 -0
- {nanopycodeagent-0.7.0 → nanopycodeagent-0.9.0}/.zed/debug.json +0 -0
- {nanopycodeagent-0.7.0 → nanopycodeagent-0.9.0}/CLAUDE.md +0 -0
- {nanopycodeagent-0.7.0 → nanopycodeagent-0.9.0}/LICENSE +0 -0
- {nanopycodeagent-0.7.0 → nanopycodeagent-0.9.0}/docs/RELEASING.md +0 -0
- {nanopycodeagent-0.7.0 → nanopycodeagent-0.9.0}/docs/changelogs/0.1.x.md +0 -0
- {nanopycodeagent-0.7.0 → nanopycodeagent-0.9.0}/docs/changelogs/0.2.x.md +0 -0
- {nanopycodeagent-0.7.0 → nanopycodeagent-0.9.0}/docs/changelogs/0.3.x.md +0 -0
- {nanopycodeagent-0.7.0 → nanopycodeagent-0.9.0}/docs/changelogs/0.4.x.md +0 -0
- {nanopycodeagent-0.7.0 → nanopycodeagent-0.9.0}/docs/changelogs/0.5.x.md +0 -0
- {nanopycodeagent-0.7.0 → nanopycodeagent-0.9.0}/docs/changelogs/0.6.x.md +0 -0
- {nanopycodeagent-0.7.0 → nanopycodeagent-0.9.0}/docs/changelogs/0.7.x.md +0 -0
- {nanopycodeagent-0.7.0 → nanopycodeagent-0.9.0}/docs/dev_notes/en/0.1.x.md +0 -0
- {nanopycodeagent-0.7.0 → nanopycodeagent-0.9.0}/docs/dev_notes/en/0.2.x.md +0 -0
- {nanopycodeagent-0.7.0 → nanopycodeagent-0.9.0}/docs/dev_notes/en/0.3.x.md +0 -0
- {nanopycodeagent-0.7.0 → nanopycodeagent-0.9.0}/docs/dev_notes/en/0.4.x.md +0 -0
- {nanopycodeagent-0.7.0 → nanopycodeagent-0.9.0}/docs/dev_notes/en/0.5.x.md +0 -0
- {nanopycodeagent-0.7.0 → nanopycodeagent-0.9.0}/docs/dev_notes/en/0.6.x.md +0 -0
- {nanopycodeagent-0.7.0 → nanopycodeagent-0.9.0}/docs/dev_notes/en/0.7.x.md +0 -0
- {nanopycodeagent-0.7.0 → nanopycodeagent-0.9.0}/docs/dev_notes/zh-CN/0.1.x.md +0 -0
- {nanopycodeagent-0.7.0 → nanopycodeagent-0.9.0}/docs/dev_notes/zh-CN/0.2.x.md +0 -0
- {nanopycodeagent-0.7.0 → nanopycodeagent-0.9.0}/docs/dev_notes/zh-CN/0.3.x.md +0 -0
- {nanopycodeagent-0.7.0 → nanopycodeagent-0.9.0}/docs/dev_notes/zh-CN/0.4.x.md +0 -0
- {nanopycodeagent-0.7.0 → nanopycodeagent-0.9.0}/docs/dev_notes/zh-CN/0.5.x.md +0 -0
- {nanopycodeagent-0.7.0 → nanopycodeagent-0.9.0}/docs/dev_notes/zh-CN/0.6.x.md +0 -0
- {nanopycodeagent-0.7.0 → nanopycodeagent-0.9.0}/docs/dev_notes/zh-CN/0.7.x.md +0 -0
- {nanopycodeagent-0.7.0 → nanopycodeagent-0.9.0}/docs/research/en/agent_tools.md +0 -0
- {nanopycodeagent-0.7.0 → nanopycodeagent-0.9.0}/docs/research/en/edit_tool.md +0 -0
- {nanopycodeagent-0.7.0 → nanopycodeagent-0.9.0}/docs/research/en/read_tool.md +0 -0
- {nanopycodeagent-0.7.0 → nanopycodeagent-0.9.0}/docs/research/en/write_tool.md +0 -0
- {nanopycodeagent-0.7.0 → nanopycodeagent-0.9.0}/docs/research/zh-CN/agent_tools.md +0 -0
- {nanopycodeagent-0.7.0 → nanopycodeagent-0.9.0}/docs/research/zh-CN/edit_tool.md +0 -0
- {nanopycodeagent-0.7.0 → nanopycodeagent-0.9.0}/docs/research/zh-CN/read_tool.md +0 -0
- {nanopycodeagent-0.7.0 → nanopycodeagent-0.9.0}/docs/research/zh-CN/write_tool.md +0 -0
- {nanopycodeagent-0.7.0 → nanopycodeagent-0.9.0}/docs/superpowers/plans/2026-06-21-release-skills.md +0 -0
- {nanopycodeagent-0.7.0 → nanopycodeagent-0.9.0}/docs/superpowers/specs/2026-06-21-release-skill-design.md +0 -0
- {nanopycodeagent-0.7.0 → nanopycodeagent-0.9.0}/docs/superpowers/specs/2026-06-29-config-file-support-design.md +0 -0
- {nanopycodeagent-0.7.0 → nanopycodeagent-0.9.0}/src/nanopycodeagent/edit_tool.py +0 -0
- {nanopycodeagent-0.7.0 → nanopycodeagent-0.9.0}/src/nanopycodeagent/read_tool.py +0 -0
- {nanopycodeagent-0.7.0 → nanopycodeagent-0.9.0}/src/nanopycodeagent/terminal.py +0 -0
- {nanopycodeagent-0.7.0 → nanopycodeagent-0.9.0}/src/nanopycodeagent/write_tool.py +0 -0
- {nanopycodeagent-0.7.0 → nanopycodeagent-0.9.0}/tests/test_edit_tool.py +0 -0
- {nanopycodeagent-0.7.0 → nanopycodeagent-0.9.0}/tests/test_line_editing.py +0 -0
- {nanopycodeagent-0.7.0 → nanopycodeagent-0.9.0}/tests/test_read_tool.py +0 -0
- {nanopycodeagent-0.7.0 → nanopycodeagent-0.9.0}/tests/test_settings.py +0 -0
- {nanopycodeagent-0.7.0 → nanopycodeagent-0.9.0}/tests/test_terminal.py +0 -0
- {nanopycodeagent-0.7.0 → nanopycodeagent-0.9.0}/tests/test_write_tool.py +0 -0
{nanopycodeagent-0.7.0 → nanopycodeagent-0.9.0}/.agents/skills/release/scripts/verify-release.sh
RENAMED
|
@@ -8,8 +8,7 @@
|
|
|
8
8
|
#
|
|
9
9
|
# Usage: verify-release.sh <version|vVERSION> e.g. verify-release.sh 0.1.1
|
|
10
10
|
#
|
|
11
|
-
# Smoke test
|
|
12
|
-
# If the entry point ever gains required args/interaction, adjust step 3.
|
|
11
|
+
# Smoke test uses --version without a task and checks the installed version.
|
|
13
12
|
set -euo pipefail
|
|
14
13
|
|
|
15
14
|
PACKAGE="nanoPyCodeAgent"
|
|
@@ -63,11 +62,12 @@ printf 'OK: GitHub Release %s has wheel + sdist\n' "$tag"
|
|
|
63
62
|
# --default-index (not --index): --index only *adds* an index, and uv's
|
|
64
63
|
# first-index strategy would stop at a locally configured mirror that has the
|
|
65
64
|
# package but not the new version yet. --refresh: uv's cached simple-index
|
|
66
|
-
# metadata can be stale right after a publish.
|
|
67
|
-
#
|
|
65
|
+
# metadata can be stale right after a publish. --version exits before task input
|
|
66
|
+
# or credential validation; empty stdin without it is a command-line error.
|
|
68
67
|
log "Smoke-testing uvx ${PACKAGE}@${version}"
|
|
69
68
|
deadline=$(( $(date +%s) + POLL_TIMEOUT ))
|
|
70
|
-
until uvx --refresh --default-index "$PYPI_INDEX" --from "${PACKAGE}@${version}" "$PACKAGE" >/dev/null
|
|
69
|
+
until reported_version="$(uvx --refresh --default-index "$PYPI_INDEX" --from "${PACKAGE}@${version}" "$PACKAGE" --version 2>/dev/null </dev/null)" \
|
|
70
|
+
&& [[ "$reported_version" == "${PACKAGE} ${version}" ]]; do
|
|
71
71
|
(( $(date +%s) >= deadline )) && fail "uvx smoke test failed for ${PACKAGE}@${version} after ${POLL_TIMEOUT}s"
|
|
72
72
|
printf 'uvx not ready (index propagation?); retrying in %ss...\n' "$POLL_INTERVAL"; sleep "$POLL_INTERVAL"
|
|
73
73
|
done
|
|
@@ -39,3 +39,21 @@ jobs:
|
|
|
39
39
|
# --frozen installs exactly what uv.lock pins and fails if the lock has
|
|
40
40
|
# drifted from pyproject.toml, so CI tests the declared dependencies.
|
|
41
41
|
run: uv run --frozen pytest
|
|
42
|
+
- name: Test the httpx2 SDK transport
|
|
43
|
+
run: uv run --with anthropic==1.5.0 pytest
|
|
44
|
+
|
|
45
|
+
harbor:
|
|
46
|
+
name: Harbor workflow contracts
|
|
47
|
+
runs-on: ubuntu-latest
|
|
48
|
+
steps:
|
|
49
|
+
- uses: actions/checkout@v6
|
|
50
|
+
with:
|
|
51
|
+
fetch-depth: 0
|
|
52
|
+
- name: Install uv
|
|
53
|
+
uses: astral-sh/setup-uv@08807647e7069bb48b6ef5acd8ec9567f424441b # v8.1.0
|
|
54
|
+
with:
|
|
55
|
+
python-version: "3.13"
|
|
56
|
+
- name: Test cache integrity, preflight, and verifier retries
|
|
57
|
+
run: >-
|
|
58
|
+
uv run --frozen --project benchmarks/harbor pytest
|
|
59
|
+
-c benchmarks/harbor/pyproject.toml benchmarks/harbor/tests
|
|
@@ -19,6 +19,24 @@ the matching one:
|
|
|
19
19
|
|
|
20
20
|
See `docs/RELEASING.md` for the release process and prerequisites.
|
|
21
21
|
|
|
22
|
+
## Benchmarks
|
|
23
|
+
|
|
24
|
+
For new Harbor benchmark runs, use the standard workflow from the repository
|
|
25
|
+
root: `uv run --project benchmarks/harbor python -m harbor_adapter run --config
|
|
26
|
+
<config.json>`. See `benchmarks/harbor/README.md` for configuration, verified APT
|
|
27
|
+
cache import, and the `--install-only` dependency preflight option.
|
|
28
|
+
|
|
29
|
+
- Do not create per-job adapter copies or ad hoc bootstrap scripts. Add reusable
|
|
30
|
+
fixes and reviewed task-revision profiles to `benchmarks/harbor` with tests.
|
|
31
|
+
- Preserve task images, package sources, and official test scripts. Cache hits
|
|
32
|
+
must match the current package index's filename, size, and SHA256 digest.
|
|
33
|
+
- Run dependency preflight before model work when a task has a registered
|
|
34
|
+
profile. Record unsupported profiles explicitly; do not claim they passed.
|
|
35
|
+
- Keep setup failures and verifier errors separate from scored failures. Only
|
|
36
|
+
verifier timeouts may retry, at most three total attempts using the same agent
|
|
37
|
+
output and the task's original per-attempt time limit.
|
|
38
|
+
|
|
39
|
+
|
|
22
40
|
## Language
|
|
23
41
|
|
|
24
42
|
Write everything in **English** by default — source code (identifiers,
|
|
@@ -37,11 +55,15 @@ The **only** exception is documentation explicitly designated as Chinese:
|
|
|
37
55
|
Everything outside that list — including code under `src/`, this file, and all
|
|
38
56
|
other docs — is English.
|
|
39
57
|
|
|
58
|
+
Do not label a document as `hand-written` unless its contents are actually
|
|
59
|
+
authored and maintained by humans. For agent-authored source-of-truth documents,
|
|
60
|
+
use wording such as `Chinese source` instead.
|
|
61
|
+
|
|
40
62
|
### Bilingual research notes
|
|
41
63
|
|
|
42
64
|
Research notes are bilingual:
|
|
43
65
|
|
|
44
|
-
- `docs/research/zh-CN/` contains the
|
|
66
|
+
- `docs/research/zh-CN/` contains the Chinese sources of truth.
|
|
45
67
|
- `docs/research/en/` contains English versions generated from the Chinese
|
|
46
68
|
sources; regenerate the whole corresponding file instead of hand-editing it.
|
|
47
69
|
|
|
@@ -52,6 +74,28 @@ the English version is in sync, even when it has already been updated.
|
|
|
52
74
|
|
|
53
75
|
## Commits & PRs
|
|
54
76
|
|
|
77
|
+
### Branch and pull request workflow
|
|
78
|
+
|
|
79
|
+
All non-release changes MUST be made on a dedicated branch and merged into
|
|
80
|
+
`main` through a GitHub pull request.
|
|
81
|
+
|
|
82
|
+
- Before modifying repository files, agents MUST verify that the current branch
|
|
83
|
+
is not `main`. If it is `main`, create or switch to a dedicated branch first.
|
|
84
|
+
- Agents MUST NOT commit directly to `main` or push non-release commits to
|
|
85
|
+
`origin/main`.
|
|
86
|
+
- Open a pull request targeting `main`, and use the `land-pr` skill to merge it.
|
|
87
|
+
Do not merge without the skill's mandatory maintainer confirmation.
|
|
88
|
+
- If work was accidentally started on `main`, move the changes to a dedicated
|
|
89
|
+
branch before committing or pushing them.
|
|
90
|
+
|
|
91
|
+
The sole exception is the release workflow in `.agents/skills/release/SKILL.md`.
|
|
92
|
+
Only when the maintainer explicitly requests a release may that skill commit and
|
|
93
|
+
push its release-only changes directly to `main`, and only after its mandatory
|
|
94
|
+
confirmation gate. Documentation, chores, hotfixes, and all other work still
|
|
95
|
+
require a branch and pull request.
|
|
96
|
+
|
|
97
|
+
### Commit and PR title format
|
|
98
|
+
|
|
55
99
|
All agents working in this repo MUST follow
|
|
56
100
|
[Conventional Commits 1.0.0](https://www.conventionalcommits.org/en/v1.0.0/)
|
|
57
101
|
for **commit messages** and **PR titles** (squash-merge often uses the PR
|
|
@@ -0,0 +1,65 @@
|
|
|
1
|
+
# Code Agent Runtime
|
|
2
|
+
|
|
3
|
+
This glossary defines the runtime artifacts produced and maintained by a code agent so that public output, diagnostics, evaluation data, and resumable state remain distinct.
|
|
4
|
+
|
|
5
|
+
## Language
|
|
6
|
+
|
|
7
|
+
**Agent Run**:
|
|
8
|
+
A bounded execution that starts with an instruction and ends when the agent stops autonomous work.
|
|
9
|
+
_Avoid_: Session, turn
|
|
10
|
+
|
|
11
|
+
**Run Output**:
|
|
12
|
+
The public result or event protocol that one agent run emits to its caller.
|
|
13
|
+
_Avoid_: Trace, transcript
|
|
14
|
+
|
|
15
|
+
**Run Result Object**:
|
|
16
|
+
A single public object that summarizes how one Agent Run ended, including its final result and optional usage metadata.
|
|
17
|
+
_Avoid_: Trajectory, session export
|
|
18
|
+
|
|
19
|
+
**Public Event Stream**:
|
|
20
|
+
A caller-facing Run Output that exposes a stable, selected sequence of Agent Run events as they occur.
|
|
21
|
+
_Avoid_: Internal event bus, execution trace
|
|
22
|
+
|
|
23
|
+
**Source Record**:
|
|
24
|
+
An agent-specific record observed at a CLI, SDK, or protocol boundary and consumed by an adapter; it may already be a projection of the agent's internal state.
|
|
25
|
+
_Avoid_: Native event, journal entry
|
|
26
|
+
|
|
27
|
+
**Native Event**:
|
|
28
|
+
An agent-independent runtime fact accepted by the core directly or normalized from a Source Record.
|
|
29
|
+
_Avoid_: Source record, public event record, trajectory step
|
|
30
|
+
|
|
31
|
+
**Journal Entry**:
|
|
32
|
+
A Native Event together with the identity, ordering, and recording metadata required for durable history.
|
|
33
|
+
_Avoid_: Persistence envelope, trajectory step
|
|
34
|
+
|
|
35
|
+
**Event Journal**:
|
|
36
|
+
A durable append-only sequence of Journal Entries for an Agent Run from which derived artifacts can be rebuilt.
|
|
37
|
+
_Avoid_: Trajectory, session, public event stream
|
|
38
|
+
|
|
39
|
+
**Partial Content**:
|
|
40
|
+
An unfinished representation of a message, reasoning block, or tool input that is exposed before the logical content is complete.
|
|
41
|
+
_Avoid_: Final message, completed event
|
|
42
|
+
|
|
43
|
+
**Content Delta**:
|
|
44
|
+
The incremental fragment or change relative to content already emitted for the same logical item.
|
|
45
|
+
_Avoid_: Partial snapshot, final message
|
|
46
|
+
|
|
47
|
+
**Terminal Event**:
|
|
48
|
+
The final Public Event Stream record that explicitly states how an Agent Run ended.
|
|
49
|
+
_Avoid_: EOF, last message
|
|
50
|
+
|
|
51
|
+
**Execution Trace**:
|
|
52
|
+
A diagnostic record of runtime activity retained to explain failures, timing, and internal behavior.
|
|
53
|
+
_Avoid_: Run output, trajectory
|
|
54
|
+
|
|
55
|
+
**Trajectory**:
|
|
56
|
+
A task-scoped record of observations, actions, results, and outcome prepared for evaluation or offline analysis.
|
|
57
|
+
_Avoid_: Execution trace, session, transcript
|
|
58
|
+
|
|
59
|
+
**ATIF Trajectory**:
|
|
60
|
+
A Trajectory represented by the Agent Trajectory Interchange Format and exposed as the runtime's public trajectory contract.
|
|
61
|
+
_Avoid_: Event journal, native trajectory
|
|
62
|
+
|
|
63
|
+
**Session**:
|
|
64
|
+
Durable agent state that can span multiple runs and supports continuing or branching prior work.
|
|
65
|
+
_Avoid_: Agent run, trajectory, transcript
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
Metadata-Version: 2.5
|
|
2
2
|
Name: nanoPyCodeAgent
|
|
3
|
-
Version: 0.
|
|
3
|
+
Version: 0.9.0
|
|
4
4
|
Summary: A nano code agent built from scratch in pure Python.
|
|
5
5
|
Project-URL: Homepage, https://github.com/minixalpha/nanoPyCodeAgent
|
|
6
6
|
Project-URL: Repository, https://github.com/minixalpha/nanoPyCodeAgent
|
|
@@ -14,6 +14,7 @@ Classifier: Programming Language :: Python :: 3
|
|
|
14
14
|
Classifier: Programming Language :: Python :: 3.13
|
|
15
15
|
Requires-Python: >=3.13
|
|
16
16
|
Requires-Dist: anthropic>=0.112.0
|
|
17
|
+
Requires-Dist: httpx>=0.25.0
|
|
17
18
|
Description-Content-Type: text/markdown
|
|
18
19
|
|
|
19
20
|
# nanoPyCodeAgent
|
|
@@ -49,6 +50,20 @@ uv tool install nanoPyCodeAgent # or: pipx install nanoPyCodeAgent
|
|
|
49
50
|
nanoPyCodeAgent
|
|
50
51
|
```
|
|
51
52
|
|
|
53
|
+
#### Run one task and exit
|
|
54
|
+
|
|
55
|
+
Give it a task and it works through it on its own, with no prompt and nothing
|
|
56
|
+
to confirm — the shape a script or a benchmark harness needs:
|
|
57
|
+
|
|
58
|
+
```bash
|
|
59
|
+
nanoPyCodeAgent -p "add a --version flag and run the tests"
|
|
60
|
+
```
|
|
61
|
+
|
|
62
|
+
The task runs in the current directory and the command exits when the run
|
|
63
|
+
ends. See the [complete CLI reference](docs/user_docs/en/cli_reference.md) for
|
|
64
|
+
all task input methods, options, exit statuses, trajectory output, and Event
|
|
65
|
+
Journal behavior.
|
|
66
|
+
|
|
52
67
|
#### Run a branch or tagged version
|
|
53
68
|
|
|
54
69
|
Run an unreleased branch or a specific release tag straight from GitHub:
|
|
@@ -63,29 +78,11 @@ uvx --from "git+https://github.com/minixalpha/nanoPyCodeAgent@v0.1.0" nanoPyCode
|
|
|
63
78
|
|
|
64
79
|
### Configuration
|
|
65
80
|
|
|
66
|
-
|
|
67
|
-
|
|
68
|
-
|
|
69
|
-
|
|
70
|
-
|
|
71
|
-
The config file mirrors [Claude Code's settings](https://code.claude.com/docs/en/settings):
|
|
72
|
-
put the values under an `env` object. Empty or whitespace-only values are ignored.
|
|
73
|
-
|
|
74
|
-
```json
|
|
75
|
-
{
|
|
76
|
-
"env": {
|
|
77
|
-
"ANTHROPIC_API_KEY": "",
|
|
78
|
-
"ANTHROPIC_BASE_URL": "",
|
|
79
|
-
"ANTHROPIC_MODEL": ""
|
|
80
|
-
}
|
|
81
|
-
}
|
|
82
|
-
```
|
|
83
|
-
|
|
84
|
-
| Variable | Required | Default | Description |
|
|
85
|
-
| --- | --- | --- | --- |
|
|
86
|
-
| `ANTHROPIC_API_KEY` | Yes | — | Your Anthropic API key, or the key for a third-party / proxy service. |
|
|
87
|
-
| `ANTHROPIC_BASE_URL` | No | `https://api.anthropic.com` | Point the SDK at a non-official / proxy endpoint. Leave it unset to use the official API — an empty value breaks requests. |
|
|
88
|
-
| `ANTHROPIC_MODEL` | No | `claude-sonnet-4-6` | Override the model. An empty or whitespace-only value falls back to the default. |
|
|
81
|
+
Set `ANTHROPIC_API_KEY` or `ANTHROPIC_AUTH_TOKEN` before running the agent.
|
|
82
|
+
You can also configure the endpoint and model through environment variables or
|
|
83
|
+
`~/.nanoPyCodeAgent/settings.json`. See the [configuration reference](docs/user_docs/en/configuration.md)
|
|
84
|
+
for supported variables, defaults, file format, precedence, and empty-value
|
|
85
|
+
handling.
|
|
89
86
|
|
|
90
87
|
### How to Update
|
|
91
88
|
|
|
@@ -31,6 +31,20 @@ uv tool install nanoPyCodeAgent # or: pipx install nanoPyCodeAgent
|
|
|
31
31
|
nanoPyCodeAgent
|
|
32
32
|
```
|
|
33
33
|
|
|
34
|
+
#### Run one task and exit
|
|
35
|
+
|
|
36
|
+
Give it a task and it works through it on its own, with no prompt and nothing
|
|
37
|
+
to confirm — the shape a script or a benchmark harness needs:
|
|
38
|
+
|
|
39
|
+
```bash
|
|
40
|
+
nanoPyCodeAgent -p "add a --version flag and run the tests"
|
|
41
|
+
```
|
|
42
|
+
|
|
43
|
+
The task runs in the current directory and the command exits when the run
|
|
44
|
+
ends. See the [complete CLI reference](docs/user_docs/en/cli_reference.md) for
|
|
45
|
+
all task input methods, options, exit statuses, trajectory output, and Event
|
|
46
|
+
Journal behavior.
|
|
47
|
+
|
|
34
48
|
#### Run a branch or tagged version
|
|
35
49
|
|
|
36
50
|
Run an unreleased branch or a specific release tag straight from GitHub:
|
|
@@ -45,29 +59,11 @@ uvx --from "git+https://github.com/minixalpha/nanoPyCodeAgent@v0.1.0" nanoPyCode
|
|
|
45
59
|
|
|
46
60
|
### Configuration
|
|
47
61
|
|
|
48
|
-
|
|
49
|
-
|
|
50
|
-
|
|
51
|
-
|
|
52
|
-
|
|
53
|
-
The config file mirrors [Claude Code's settings](https://code.claude.com/docs/en/settings):
|
|
54
|
-
put the values under an `env` object. Empty or whitespace-only values are ignored.
|
|
55
|
-
|
|
56
|
-
```json
|
|
57
|
-
{
|
|
58
|
-
"env": {
|
|
59
|
-
"ANTHROPIC_API_KEY": "",
|
|
60
|
-
"ANTHROPIC_BASE_URL": "",
|
|
61
|
-
"ANTHROPIC_MODEL": ""
|
|
62
|
-
}
|
|
63
|
-
}
|
|
64
|
-
```
|
|
65
|
-
|
|
66
|
-
| Variable | Required | Default | Description |
|
|
67
|
-
| --- | --- | --- | --- |
|
|
68
|
-
| `ANTHROPIC_API_KEY` | Yes | — | Your Anthropic API key, or the key for a third-party / proxy service. |
|
|
69
|
-
| `ANTHROPIC_BASE_URL` | No | `https://api.anthropic.com` | Point the SDK at a non-official / proxy endpoint. Leave it unset to use the official API — an empty value breaks requests. |
|
|
70
|
-
| `ANTHROPIC_MODEL` | No | `claude-sonnet-4-6` | Override the model. An empty or whitespace-only value falls back to the default. |
|
|
62
|
+
Set `ANTHROPIC_API_KEY` or `ANTHROPIC_AUTH_TOKEN` before running the agent.
|
|
63
|
+
You can also configure the endpoint and model through environment variables or
|
|
64
|
+
`~/.nanoPyCodeAgent/settings.json`. See the [configuration reference](docs/user_docs/en/configuration.md)
|
|
65
|
+
for supported variables, defaults, file format, precedence, and empty-value
|
|
66
|
+
handling.
|
|
71
67
|
|
|
72
68
|
### How to Update
|
|
73
69
|
|
|
@@ -31,6 +31,18 @@ uv tool install nanoPyCodeAgent # 或: pipx install nanoPyCodeAgent
|
|
|
31
31
|
nanoPyCodeAgent
|
|
32
32
|
```
|
|
33
33
|
|
|
34
|
+
#### 一次性任务(非交互)
|
|
35
|
+
|
|
36
|
+
给它一个任务,它会自己做完退出——不给提示符,也不会停下来等确认,这正是脚本或
|
|
37
|
+
benchmark harness 需要的形态:
|
|
38
|
+
|
|
39
|
+
```bash
|
|
40
|
+
nanoPyCodeAgent -p "add a --version flag and run the tests"
|
|
41
|
+
```
|
|
42
|
+
|
|
43
|
+
任务在当前目录下执行,run 结束后命令随即退出。全部任务输入方式、选项、退出状态、
|
|
44
|
+
trajectory 输出和 Event Journal 行为请参阅[完整 CLI 参考](docs/user_docs/zh-CN/cli_reference.md)。
|
|
45
|
+
|
|
34
46
|
#### 运行某个分支或标签版本
|
|
35
47
|
|
|
36
48
|
直接从 GitHub 运行未发布的分支,或某个具体的发布标签:
|
|
@@ -45,28 +57,9 @@ uvx --from "git+https://github.com/minixalpha/nanoPyCodeAgent@v0.1.0" nanoPyCode
|
|
|
45
57
|
|
|
46
58
|
### 配置
|
|
47
59
|
|
|
48
|
-
|
|
49
|
-
`~/.nanoPyCodeAgent/settings.json
|
|
50
|
-
|
|
51
|
-
|
|
52
|
-
配置文件写法与 [Claude Code settings](https://code.claude.com/docs/en/settings)
|
|
53
|
-
一致:把值放在 `env` 对象下。空值或纯空白会被忽略。
|
|
54
|
-
|
|
55
|
-
```json
|
|
56
|
-
{
|
|
57
|
-
"env": {
|
|
58
|
-
"ANTHROPIC_API_KEY": "",
|
|
59
|
-
"ANTHROPIC_BASE_URL": "",
|
|
60
|
-
"ANTHROPIC_MODEL": ""
|
|
61
|
-
}
|
|
62
|
-
}
|
|
63
|
-
```
|
|
64
|
-
|
|
65
|
-
| 变量 | 是否必填 | 默认值 | 说明 |
|
|
66
|
-
| --- | --- | --- | --- |
|
|
67
|
-
| `ANTHROPIC_API_KEY` | 是 | 无 | 你的 Anthropic API key,或第三方/代理服务的 key。 |
|
|
68
|
-
| `ANTHROPIC_BASE_URL` | 否 | `https://api.anthropic.com` | 将 SDK 指向非官方/代理 endpoint。使用官方 API 时保持不设置;留空值会导致请求失败。 |
|
|
69
|
-
| `ANTHROPIC_MODEL` | 否 | `claude-sonnet-4-6` | 覆盖默认模型。空值或纯空白会回退到默认值。 |
|
|
60
|
+
运行前请设置 `ANTHROPIC_API_KEY` 或 `ANTHROPIC_AUTH_TOKEN`。你也可以通过
|
|
61
|
+
环境变量或 `~/.nanoPyCodeAgent/settings.json` 配置 endpoint 与模型。支持的变量、
|
|
62
|
+
默认值、文件格式、优先级和空值处理请参阅[配置参考](docs/user_docs/zh-CN/configuration.md)。
|
|
70
63
|
|
|
71
64
|
### 如何更新
|
|
72
65
|
|
|
@@ -0,0 +1,196 @@
|
|
|
1
|
+
# Harbor adapter
|
|
2
|
+
|
|
3
|
+
[English](README.md) | [简体中文](README.zh-CN.md)
|
|
4
|
+
|
|
5
|
+
This isolated workspace contains the nanoPyCodeAgent adapter for Terminal-Bench
|
|
6
|
+
and other benchmarks run by Harbor. It is development infrastructure, not part
|
|
7
|
+
of the end-user `nanoPyCodeAgent` package. Harbor is pinned to 0.21.0 in this
|
|
8
|
+
workspace's lockfile.
|
|
9
|
+
|
|
10
|
+
## Run a benchmark
|
|
11
|
+
|
|
12
|
+
Set the connection credentials used by the CLI. For a third-party or proxy
|
|
13
|
+
endpoint, configure the API key and base URL; select the model with Harbor's
|
|
14
|
+
`--model` option shown below:
|
|
15
|
+
|
|
16
|
+
```bash
|
|
17
|
+
export ANTHROPIC_API_KEY="..."
|
|
18
|
+
export ANTHROPIC_BASE_URL="https://gateway.example"
|
|
19
|
+
```
|
|
20
|
+
|
|
21
|
+
Run new benchmarks through the standard entry point, from the repository root:
|
|
22
|
+
|
|
23
|
+
```bash
|
|
24
|
+
uv run --project benchmarks/harbor python -m harbor_adapter run \
|
|
25
|
+
--config benchmarks/harbor/configs/qemu-alpine-ssh.json
|
|
26
|
+
```
|
|
27
|
+
|
|
28
|
+
This example pins the QEMU task revision and uses `deepseek/deepseek-flash`.
|
|
29
|
+
Set the endpoint and credentials for that model, or copy the configuration and
|
|
30
|
+
change `agents[].model_name`, task selection, and budgets for another run. These
|
|
31
|
+
are ordinary Harbor JSON configurations; the entry point supplies the standard
|
|
32
|
+
verifier and cache policy. It accepts the repository's `NanoPyCodeAgent` adapter
|
|
33
|
+
with budget kwargs; credentials belong in the environment.
|
|
34
|
+
|
|
35
|
+
The entry point builds the current working tree, including uncommitted changes,
|
|
36
|
+
into a wheel and checks its Python sources against the recorded snapshot. It
|
|
37
|
+
also snapshots the adapter, records source and wheel hashes, and uses that
|
|
38
|
+
snapshot throughout the run. Inputs and logs live in `jobs/<job-name>-input/`;
|
|
39
|
+
Harbor outputs remain in `jobs/<job-name>/`. Existing jobs are never overwritten.
|
|
40
|
+
`summary.json` records scores, setup/preflight status, cache use, and verifier
|
|
41
|
+
attempts. An infrastructure exception has no score; it is not converted to zero.
|
|
42
|
+
The command exits nonzero for execution errors or incomplete results, while a
|
|
43
|
+
completed score of zero is a valid benchmark result.
|
|
44
|
+
|
|
45
|
+
Use `--job-name <unique-name>` for a readable run name. Add `--prepare-only` to
|
|
46
|
+
build and record inputs without starting containers. Add `--install-only` to
|
|
47
|
+
exercise dependency installation and preflight without calling the model or
|
|
48
|
+
executing the official tests. A full run still performs setup and preflight in
|
|
49
|
+
its own fresh container; an install-only check does not resume into model work.
|
|
50
|
+
|
|
51
|
+
### Verified dependencies and preflight
|
|
52
|
+
|
|
53
|
+
The standard adapter automatically uses `.cache/nanopy-harbor/`, shared across
|
|
54
|
+
runs and ignored by Git. Override it with `NANOPY_HARBOR_CACHE` or the entry
|
|
55
|
+
point's `--cache-dir`. On APT systems it refreshes the package index, resolves a
|
|
56
|
+
SHA256 download plan, restores matching cached files, and downloads missing
|
|
57
|
+
ones. Successful downloads are verified and saved before installation, so
|
|
58
|
+
APT cleanup hooks cannot erase the shared copies. A different version or digest
|
|
59
|
+
does not reuse an old cached file. Corrupt files fail setup before model work.
|
|
60
|
+
Other package managers use Harbor's regular dependency installation.
|
|
61
|
+
|
|
62
|
+
To import previously verified downloads once:
|
|
63
|
+
|
|
64
|
+
```bash
|
|
65
|
+
uv run --project benchmarks/harbor python -m harbor_adapter import-apt \
|
|
66
|
+
/path/to/bootstrap-cache/manifest.json
|
|
67
|
+
```
|
|
68
|
+
|
|
69
|
+
The manifest is a JSON array of objects with `url`, `filename`, `size`, and
|
|
70
|
+
`sha256`, accompanied by `packages/<filename>` files. Import validates every
|
|
71
|
+
file. At use time, the current APT index must independently match each entry.
|
|
72
|
+
Copying the entire shared cache to another machine is also supported. A new
|
|
73
|
+
machine with no matching cache still needs the original downloads to be
|
|
74
|
+
available; this mechanism does not replace package sources or invent missing
|
|
75
|
+
versions.
|
|
76
|
+
|
|
77
|
+
Verifier dependency profiles are tied to reviewed task revisions in
|
|
78
|
+
`src/harbor_adapter/bootstrap.py`. The current QEMU profile installs `sshpass`
|
|
79
|
+
and warms the original verifier's uv 0.9.5, Python 3.13, pytest 8.4.1, and
|
|
80
|
+
pytest-json-ctrf 0.3.5 dependencies using `pytest --version`. It runs no task
|
|
81
|
+
assertions or solution code. Unknown QEMU revisions fail setup pending profile
|
|
82
|
+
review; tasks without a profile explicitly report `not_configured`. Add a
|
|
83
|
+
profile and regression tests when another task needs verifier preflight.
|
|
84
|
+
|
|
85
|
+
Every installation writes `agent/bootstrap.json`, including partial progress
|
|
86
|
+
on failure. Preflight failure stops the trial before any model call. The
|
|
87
|
+
official verifier script remains unchanged and may still access the network
|
|
88
|
+
when it runs; a successful preflight is not a guarantee against later outages.
|
|
89
|
+
Full standard runs enable the timeout retry policy below by default, with zero
|
|
90
|
+
whole-trial retries. CI checks these contracts without model calls or Docker.
|
|
91
|
+
|
|
92
|
+
### Direct Harbor invocation
|
|
93
|
+
|
|
94
|
+
For adapter development or comparison with published versions, the lower-level
|
|
95
|
+
Harbor interface remains available. Pin the agent installed in the container:
|
|
96
|
+
|
|
97
|
+
```bash
|
|
98
|
+
uv run --project benchmarks/harbor harbor run \
|
|
99
|
+
--task terminal-bench/openssl-selfsigned-cert \
|
|
100
|
+
--agent harbor_adapter:NanoPyCodeAgent \
|
|
101
|
+
--agent-kwarg git_ref=<commit-sha> \
|
|
102
|
+
--model anthropic/claude-sonnet-4-6 \
|
|
103
|
+
--verifier harbor_adapter:RetryingVerifier \
|
|
104
|
+
--verifier-timeout-multiplier 4 \
|
|
105
|
+
--env docker \
|
|
106
|
+
--n-concurrent 1 \
|
|
107
|
+
--n-attempts 1
|
|
108
|
+
```
|
|
109
|
+
|
|
110
|
+
Use `--agent-kwarg version=<released-version>` instead of `git_ref` to install a
|
|
111
|
+
published PyPI release. The two pins are mutually exclusive. If neither is
|
|
112
|
+
provided, the adapter installs the latest published release. For reproducible
|
|
113
|
+
benchmark results, always provide one of them. Prefer a full 40-character commit
|
|
114
|
+
SHA for `git_ref`; Harbor may parse an unquoted abbreviated SHA such as `83e6271`
|
|
115
|
+
as a number. To use an abbreviated revision, preserve its string type with
|
|
116
|
+
`--agent-kwarg 'git_ref="83e6271"'`.
|
|
117
|
+
|
|
118
|
+
The adapter sends the instruction through stdin, runs the agent in the task
|
|
119
|
+
container's current directory, and saves combined stdout/stderr to
|
|
120
|
+
`/logs/agent/nanopycodeagent.txt`. It uses the CLI's 50-turn default; override
|
|
121
|
+
that with `--agent-kwarg max_turns=20`.
|
|
122
|
+
|
|
123
|
+
Use `--agent-kwarg time_budget_seconds=N` to bound task work and remind the model
|
|
124
|
+
to finalize before either time or replies run out. Choose a budget inside each
|
|
125
|
+
task's native timeout, reserving at least 30 seconds for cost reconciliation
|
|
126
|
+
plus trajectory writing and harness overhead. For example, the 3600-second
|
|
127
|
+
targeted experiments use 3420 seconds; that value does not fit a 900-second task.
|
|
128
|
+
No time budget is inferred automatically by the adapter.
|
|
129
|
+
|
|
130
|
+
For the per-reply generation limit, pass `--agent-kwarg max_tokens=32768`.
|
|
131
|
+
This becomes `--max-tokens 32768` in the container and overrides the forwarded
|
|
132
|
+
`ANTHROPIC_MAX_TOKENS` environment variable. When omitted, the adapter sends no
|
|
133
|
+
token flag, so the installed agent's environment/settings/default applies
|
|
134
|
+
(32768 in the version introducing this option). Older releases require omitting
|
|
135
|
+
the new option. The effective budget is recorded in the startup log and in
|
|
136
|
+
`agent.extra.max_tokens` in the ATIF trajectory.
|
|
137
|
+
|
|
138
|
+
The adapter also asks the agent to write an ATIF-v1.7 trajectory directly to
|
|
139
|
+
`/logs/agent/trajectory.json`. Harbor collects that file as the trial's native
|
|
140
|
+
ATIF output and backfills prompt, completion, cache-token, and cost totals into
|
|
141
|
+
the agent result. Missing, invalid, or partial trajectories remain explicitly
|
|
142
|
+
diagnosed; unknown usage or cost is not reported as zero.
|
|
143
|
+
If any costs are estimated, `final_metrics.extra` records
|
|
144
|
+
`cost_is_estimated: true` and the estimated portion in `estimated_cost_usd`.
|
|
145
|
+
The adapter preserves both fields in the agent result's `metadata.trajectory`,
|
|
146
|
+
including for partial totals. Estimates with a generation ID remain eligible
|
|
147
|
+
for billing reconciliation; a successful lookup replaces the estimate.
|
|
148
|
+
|
|
149
|
+
By default, the adapter strips the first provider prefix from `--model` and
|
|
150
|
+
passes the result to nanoPyCodeAgent as `ANTHROPIC_MODEL`. Set
|
|
151
|
+
`ANTHROPIC_MODEL` only when a custom endpoint requires an actual model name that
|
|
152
|
+
differs from Harbor's `provider/model` identity; this explicit override takes
|
|
153
|
+
precedence. Harbor-native provider credentials and configured base URLs are
|
|
154
|
+
also normalized to the `ANTHROPIC_*` variables expected by nanoPyCodeAgent's
|
|
155
|
+
SDK.
|
|
156
|
+
|
|
157
|
+
## Retry verifier timeouts
|
|
158
|
+
|
|
159
|
+
The example above uses `harbor_adapter:RetryingVerifier` for single-step Linux
|
|
160
|
+
tasks. It runs the original test script up to **three times total**, retrying
|
|
161
|
+
only when the script exceeds the task's native `[verifier].timeout_sec`. A
|
|
162
|
+
900-second task therefore gets at most three 900-second attempts. A completed
|
|
163
|
+
verification stops immediately, including a reward of zero. Script errors,
|
|
164
|
+
missing or invalid rewards, and cancellation do not trigger another attempt.
|
|
165
|
+
|
|
166
|
+
`--verifier-timeout-multiplier 4` gives Harbor's outer watchdog room for all
|
|
167
|
+
three attempts plus test upload, process cleanup, and log collection. It does
|
|
168
|
+
not extend each test-script attempt. Conflicting timeout overrides or caps
|
|
169
|
+
that leave insufficient overall time are rejected. Use
|
|
170
|
+
`--verifier-kwarg max_attempts=2` for two total attempts, or `max_attempts=1` for
|
|
171
|
+
one; values above three are rejected.
|
|
172
|
+
|
|
173
|
+
The runner reuses the same container, agent output, and dependency caches;
|
|
174
|
+
the agent is run once. Test scripts must tolerate being rerun in that
|
|
175
|
+
environment; verifier-side filesystem changes are not rolled back. Timed-out
|
|
176
|
+
process groups and other members of their Linux session are killed before
|
|
177
|
+
retrying. Each attempt's logs and any reward files are archived under
|
|
178
|
+
`verifier/attempts/1/`, `2/`, and `3/`. The runner clears current verifier
|
|
179
|
+
outputs before the next attempt, so a stale reward cannot determine its score.
|
|
180
|
+
`verifier/retry-summary.json` records each attempt's status, duration, and exit
|
|
181
|
+
code when available. Exhausting the attempts still produces Harbor's
|
|
182
|
+
`VerifierTimeoutError`.
|
|
183
|
+
|
|
184
|
+
This verifier requires `python3` and `bash` in the container, which the
|
|
185
|
+
nanoPyCodeAgent adapter installs. These options apply to new runs; they cannot
|
|
186
|
+
restore a deleted container or missing task outputs from an old trial.
|
|
187
|
+
Harbor's ordinary `--max-retries` setting reruns the entire trial, including
|
|
188
|
+
the agent, and is independent of these verifier attempts.
|
|
189
|
+
|
|
190
|
+
## Test the adapter
|
|
191
|
+
|
|
192
|
+
```bash
|
|
193
|
+
uv run --project benchmarks/harbor pytest \
|
|
194
|
+
-c benchmarks/harbor/pyproject.toml \
|
|
195
|
+
benchmarks/harbor/tests
|
|
196
|
+
```
|