nanoPyCodeAgent 0.7.0__tar.gz → 0.9.0__tar.gz

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (154) hide show
  1. {nanopycodeagent-0.7.0 → nanopycodeagent-0.9.0}/.agents/skills/release/scripts/verify-release.sh +5 -5
  2. {nanopycodeagent-0.7.0 → nanopycodeagent-0.9.0}/.github/workflows/ci.yml +18 -0
  3. {nanopycodeagent-0.7.0 → nanopycodeagent-0.9.0}/.gitignore +3 -0
  4. {nanopycodeagent-0.7.0 → nanopycodeagent-0.9.0}/AGENTS.md +45 -1
  5. nanopycodeagent-0.9.0/CONTEXT.md +65 -0
  6. {nanopycodeagent-0.7.0 → nanopycodeagent-0.9.0}/PKG-INFO +21 -24
  7. {nanopycodeagent-0.7.0 → nanopycodeagent-0.9.0}/README.md +19 -23
  8. {nanopycodeagent-0.7.0 → nanopycodeagent-0.9.0}/README.zh-CN.md +15 -22
  9. nanopycodeagent-0.9.0/benchmarks/harbor/README.md +196 -0
  10. nanopycodeagent-0.9.0/benchmarks/harbor/README.zh-CN.md +158 -0
  11. nanopycodeagent-0.9.0/benchmarks/harbor/configs/qemu-alpine-ssh.json +21 -0
  12. nanopycodeagent-0.9.0/benchmarks/harbor/pyproject.toml +27 -0
  13. nanopycodeagent-0.9.0/benchmarks/harbor/reports/deepseek-official-20260930.md +86 -0
  14. nanopycodeagent-0.9.0/benchmarks/harbor/reports/dual-budget-20260929.md +246 -0
  15. nanopycodeagent-0.9.0/benchmarks/harbor/reports/stream-recovery-20260915.md +139 -0
  16. nanopycodeagent-0.9.0/benchmarks/harbor/results/tb21-deepseek-official-pilot20-20260930.json +387 -0
  17. nanopycodeagent-0.9.0/benchmarks/harbor/results/tb21-dualbudget-fixed-provider-20260929.json +1670 -0
  18. nanopycodeagent-0.9.0/benchmarks/harbor/results/tb21-exp1-turn100-20260926.json +120 -0
  19. nanopycodeagent-0.9.0/benchmarks/harbor/results/tb21-exp2-timebudget-20260926.json +127 -0
  20. nanopycodeagent-0.9.0/benchmarks/harbor/results/tb21-exp3-timebudget-tail-20260926.json +115 -0
  21. nanopycodeagent-0.9.0/benchmarks/harbor/results/tb21-flash0731-pilot20-20260906.json +1477 -0
  22. nanopycodeagent-0.9.0/benchmarks/harbor/results/tb21-generation-budget-65536-20260912.json +1449 -0
  23. nanopycodeagent-0.9.0/benchmarks/harbor/results/tb21-tool-input-recovery-20260914.json +3541 -0
  24. nanopycodeagent-0.9.0/benchmarks/harbor/results/tb21-v41flash-6task-20260925.json +125 -0
  25. nanopycodeagent-0.9.0/benchmarks/harbor/src/harbor_adapter/__init__.py +6 -0
  26. nanopycodeagent-0.9.0/benchmarks/harbor/src/harbor_adapter/__main__.py +182 -0
  27. nanopycodeagent-0.9.0/benchmarks/harbor/src/harbor_adapter/_verifier_runner.py +143 -0
  28. nanopycodeagent-0.9.0/benchmarks/harbor/src/harbor_adapter/adapter.py +332 -0
  29. nanopycodeagent-0.9.0/benchmarks/harbor/src/harbor_adapter/bootstrap.py +183 -0
  30. nanopycodeagent-0.9.0/benchmarks/harbor/src/harbor_adapter/verifier.py +150 -0
  31. nanopycodeagent-0.9.0/benchmarks/harbor/tests/fixtures/atif-journal-v1.jsonl +7 -0
  32. nanopycodeagent-0.9.0/benchmarks/harbor/tests/test_adapter.py +355 -0
  33. nanopycodeagent-0.9.0/benchmarks/harbor/tests/test_atif_compatibility.py +157 -0
  34. nanopycodeagent-0.9.0/benchmarks/harbor/tests/test_bootstrap.py +229 -0
  35. nanopycodeagent-0.9.0/benchmarks/harbor/tests/test_verifier.py +324 -0
  36. nanopycodeagent-0.9.0/benchmarks/harbor/tests/test_workflow.py +90 -0
  37. nanopycodeagent-0.9.0/benchmarks/harbor/uv.lock +2059 -0
  38. nanopycodeagent-0.9.0/docs/changelogs/0.8.x.md +67 -0
  39. nanopycodeagent-0.9.0/docs/changelogs/0.9.x.md +81 -0
  40. {nanopycodeagent-0.7.0 → nanopycodeagent-0.9.0}/docs/changelogs/README.md +3 -0
  41. nanopycodeagent-0.9.0/docs/dev_docs/README.md +41 -0
  42. nanopycodeagent-0.9.0/docs/dev_docs/en/event-journal-protocol-v1.md +549 -0
  43. nanopycodeagent-0.9.0/docs/dev_docs/en/event-journal-protocol-v2.md +81 -0
  44. nanopycodeagent-0.9.0/docs/dev_docs/en/event-journal-protocol-v3.md +61 -0
  45. nanopycodeagent-0.9.0/docs/dev_docs/en/event-journal-protocol-v4.md +102 -0
  46. nanopycodeagent-0.9.0/docs/dev_docs/zh-CN/event-journal-protocol-v1.md +452 -0
  47. nanopycodeagent-0.9.0/docs/dev_docs/zh-CN/event-journal-protocol-v2.md +61 -0
  48. nanopycodeagent-0.9.0/docs/dev_docs/zh-CN/event-journal-protocol-v3.md +50 -0
  49. nanopycodeagent-0.9.0/docs/dev_docs/zh-CN/event-journal-protocol-v4.md +80 -0
  50. {nanopycodeagent-0.7.0 → nanopycodeagent-0.9.0}/docs/dev_notes/README.md +6 -1
  51. nanopycodeagent-0.9.0/docs/dev_notes/en/0.8.x.md +293 -0
  52. nanopycodeagent-0.9.0/docs/dev_notes/en/0.9.x.md +1076 -0
  53. nanopycodeagent-0.9.0/docs/dev_notes/zh-CN/0.8.x.md +295 -0
  54. nanopycodeagent-0.9.0/docs/dev_notes/zh-CN/0.9.x.md +1451 -0
  55. {nanopycodeagent-0.7.0 → nanopycodeagent-0.9.0}/docs/research/README.md +10 -1
  56. nanopycodeagent-0.9.0/docs/research/en/agent_events_to_atif_examples.md +727 -0
  57. nanopycodeagent-0.9.0/docs/research/en/agent_output_and_trajectory.md +755 -0
  58. nanopycodeagent-0.9.0/docs/research/en/benchmark_headless_interface.md +259 -0
  59. nanopycodeagent-0.9.0/docs/research/en/code_agent_benchmark.md +513 -0
  60. nanopycodeagent-0.9.0/docs/research/en/openrouter_cost_accounting.md +335 -0
  61. nanopycodeagent-0.9.0/docs/research/en/openrouter_unified_protocol.md +392 -0
  62. nanopycodeagent-0.9.0/docs/research/zh-CN/agent_events_to_atif_examples.md +727 -0
  63. nanopycodeagent-0.9.0/docs/research/zh-CN/agent_output_and_trajectory.md +755 -0
  64. nanopycodeagent-0.9.0/docs/research/zh-CN/benchmark_headless_interface.md +259 -0
  65. nanopycodeagent-0.9.0/docs/research/zh-CN/code_agent_benchmark.md +513 -0
  66. nanopycodeagent-0.9.0/docs/research/zh-CN/openrouter_cost_accounting.md +335 -0
  67. nanopycodeagent-0.9.0/docs/research/zh-CN/openrouter_unified_protocol.md +392 -0
  68. nanopycodeagent-0.9.0/docs/user_docs/README.md +10 -0
  69. nanopycodeagent-0.9.0/docs/user_docs/en/cli_reference.md +210 -0
  70. nanopycodeagent-0.9.0/docs/user_docs/en/configuration.md +184 -0
  71. nanopycodeagent-0.9.0/docs/user_docs/zh-CN/cli_reference.md +176 -0
  72. nanopycodeagent-0.9.0/docs/user_docs/zh-CN/configuration.md +165 -0
  73. {nanopycodeagent-0.7.0 → nanopycodeagent-0.9.0}/pyproject.toml +1 -0
  74. nanopycodeagent-0.9.0/src/nanopycodeagent/__init__.py +3 -0
  75. nanopycodeagent-0.9.0/src/nanopycodeagent/agent.py +1006 -0
  76. nanopycodeagent-0.9.0/src/nanopycodeagent/atif.py +591 -0
  77. nanopycodeagent-0.9.0/src/nanopycodeagent/bash_tool.py +130 -0
  78. nanopycodeagent-0.9.0/src/nanopycodeagent/cli.py +172 -0
  79. nanopycodeagent-0.9.0/src/nanopycodeagent/cost.py +211 -0
  80. nanopycodeagent-0.9.0/src/nanopycodeagent/deadline.py +39 -0
  81. nanopycodeagent-0.9.0/src/nanopycodeagent/event_journal.py +809 -0
  82. {nanopycodeagent-0.7.0 → nanopycodeagent-0.9.0}/src/nanopycodeagent/settings.py +20 -0
  83. nanopycodeagent-0.9.0/src/nanopycodeagent/tool_validation.py +40 -0
  84. nanopycodeagent-0.9.0/src/nanopycodeagent/transport.py +18 -0
  85. {nanopycodeagent-0.7.0 → nanopycodeagent-0.9.0}/tests/conftest.py +4 -1
  86. {nanopycodeagent-0.7.0 → nanopycodeagent-0.9.0}/tests/helpers.py +55 -8
  87. {nanopycodeagent-0.7.0 → nanopycodeagent-0.9.0}/tests/test_agent.py +6 -5
  88. nanopycodeagent-0.9.0/tests/test_agent_events.py +455 -0
  89. nanopycodeagent-0.9.0/tests/test_atif.py +611 -0
  90. nanopycodeagent-0.9.0/tests/test_bash_tool.py +120 -0
  91. nanopycodeagent-0.9.0/tests/test_cli.py +367 -0
  92. nanopycodeagent-0.9.0/tests/test_cost.py +201 -0
  93. nanopycodeagent-0.9.0/tests/test_event_journal.py +541 -0
  94. nanopycodeagent-0.9.0/tests/test_generation_budget.py +117 -0
  95. nanopycodeagent-0.9.0/tests/test_stream_recovery.py +226 -0
  96. nanopycodeagent-0.9.0/tests/test_time_budget.py +377 -0
  97. nanopycodeagent-0.9.0/tests/test_tool_validation.py +228 -0
  98. nanopycodeagent-0.9.0/tests/test_truncation.py +200 -0
  99. {nanopycodeagent-0.7.0 → nanopycodeagent-0.9.0}/uv.lock +5 -1
  100. nanopycodeagent-0.7.0/src/nanopycodeagent/__init__.py +0 -5
  101. nanopycodeagent-0.7.0/src/nanopycodeagent/agent.py +0 -194
  102. nanopycodeagent-0.7.0/src/nanopycodeagent/bash_tool.py +0 -77
  103. nanopycodeagent-0.7.0/tests/test_bash_tool.py +0 -62
  104. {nanopycodeagent-0.7.0 → nanopycodeagent-0.9.0}/.agents/skills/land-pr/SKILL.md +0 -0
  105. {nanopycodeagent-0.7.0 → nanopycodeagent-0.9.0}/.agents/skills/release/SKILL.md +0 -0
  106. {nanopycodeagent-0.7.0 → nanopycodeagent-0.9.0}/.github/workflows/release.yml +0 -0
  107. {nanopycodeagent-0.7.0 → nanopycodeagent-0.9.0}/.ignore +0 -0
  108. {nanopycodeagent-0.7.0 → nanopycodeagent-0.9.0}/.python-version +0 -0
  109. {nanopycodeagent-0.7.0 → nanopycodeagent-0.9.0}/.zed/debug.json +0 -0
  110. {nanopycodeagent-0.7.0 → nanopycodeagent-0.9.0}/CLAUDE.md +0 -0
  111. {nanopycodeagent-0.7.0 → nanopycodeagent-0.9.0}/LICENSE +0 -0
  112. {nanopycodeagent-0.7.0 → nanopycodeagent-0.9.0}/docs/RELEASING.md +0 -0
  113. {nanopycodeagent-0.7.0 → nanopycodeagent-0.9.0}/docs/changelogs/0.1.x.md +0 -0
  114. {nanopycodeagent-0.7.0 → nanopycodeagent-0.9.0}/docs/changelogs/0.2.x.md +0 -0
  115. {nanopycodeagent-0.7.0 → nanopycodeagent-0.9.0}/docs/changelogs/0.3.x.md +0 -0
  116. {nanopycodeagent-0.7.0 → nanopycodeagent-0.9.0}/docs/changelogs/0.4.x.md +0 -0
  117. {nanopycodeagent-0.7.0 → nanopycodeagent-0.9.0}/docs/changelogs/0.5.x.md +0 -0
  118. {nanopycodeagent-0.7.0 → nanopycodeagent-0.9.0}/docs/changelogs/0.6.x.md +0 -0
  119. {nanopycodeagent-0.7.0 → nanopycodeagent-0.9.0}/docs/changelogs/0.7.x.md +0 -0
  120. {nanopycodeagent-0.7.0 → nanopycodeagent-0.9.0}/docs/dev_notes/en/0.1.x.md +0 -0
  121. {nanopycodeagent-0.7.0 → nanopycodeagent-0.9.0}/docs/dev_notes/en/0.2.x.md +0 -0
  122. {nanopycodeagent-0.7.0 → nanopycodeagent-0.9.0}/docs/dev_notes/en/0.3.x.md +0 -0
  123. {nanopycodeagent-0.7.0 → nanopycodeagent-0.9.0}/docs/dev_notes/en/0.4.x.md +0 -0
  124. {nanopycodeagent-0.7.0 → nanopycodeagent-0.9.0}/docs/dev_notes/en/0.5.x.md +0 -0
  125. {nanopycodeagent-0.7.0 → nanopycodeagent-0.9.0}/docs/dev_notes/en/0.6.x.md +0 -0
  126. {nanopycodeagent-0.7.0 → nanopycodeagent-0.9.0}/docs/dev_notes/en/0.7.x.md +0 -0
  127. {nanopycodeagent-0.7.0 → nanopycodeagent-0.9.0}/docs/dev_notes/zh-CN/0.1.x.md +0 -0
  128. {nanopycodeagent-0.7.0 → nanopycodeagent-0.9.0}/docs/dev_notes/zh-CN/0.2.x.md +0 -0
  129. {nanopycodeagent-0.7.0 → nanopycodeagent-0.9.0}/docs/dev_notes/zh-CN/0.3.x.md +0 -0
  130. {nanopycodeagent-0.7.0 → nanopycodeagent-0.9.0}/docs/dev_notes/zh-CN/0.4.x.md +0 -0
  131. {nanopycodeagent-0.7.0 → nanopycodeagent-0.9.0}/docs/dev_notes/zh-CN/0.5.x.md +0 -0
  132. {nanopycodeagent-0.7.0 → nanopycodeagent-0.9.0}/docs/dev_notes/zh-CN/0.6.x.md +0 -0
  133. {nanopycodeagent-0.7.0 → nanopycodeagent-0.9.0}/docs/dev_notes/zh-CN/0.7.x.md +0 -0
  134. {nanopycodeagent-0.7.0 → nanopycodeagent-0.9.0}/docs/research/en/agent_tools.md +0 -0
  135. {nanopycodeagent-0.7.0 → nanopycodeagent-0.9.0}/docs/research/en/edit_tool.md +0 -0
  136. {nanopycodeagent-0.7.0 → nanopycodeagent-0.9.0}/docs/research/en/read_tool.md +0 -0
  137. {nanopycodeagent-0.7.0 → nanopycodeagent-0.9.0}/docs/research/en/write_tool.md +0 -0
  138. {nanopycodeagent-0.7.0 → nanopycodeagent-0.9.0}/docs/research/zh-CN/agent_tools.md +0 -0
  139. {nanopycodeagent-0.7.0 → nanopycodeagent-0.9.0}/docs/research/zh-CN/edit_tool.md +0 -0
  140. {nanopycodeagent-0.7.0 → nanopycodeagent-0.9.0}/docs/research/zh-CN/read_tool.md +0 -0
  141. {nanopycodeagent-0.7.0 → nanopycodeagent-0.9.0}/docs/research/zh-CN/write_tool.md +0 -0
  142. {nanopycodeagent-0.7.0 → nanopycodeagent-0.9.0}/docs/superpowers/plans/2026-06-21-release-skills.md +0 -0
  143. {nanopycodeagent-0.7.0 → nanopycodeagent-0.9.0}/docs/superpowers/specs/2026-06-21-release-skill-design.md +0 -0
  144. {nanopycodeagent-0.7.0 → nanopycodeagent-0.9.0}/docs/superpowers/specs/2026-06-29-config-file-support-design.md +0 -0
  145. {nanopycodeagent-0.7.0 → nanopycodeagent-0.9.0}/src/nanopycodeagent/edit_tool.py +0 -0
  146. {nanopycodeagent-0.7.0 → nanopycodeagent-0.9.0}/src/nanopycodeagent/read_tool.py +0 -0
  147. {nanopycodeagent-0.7.0 → nanopycodeagent-0.9.0}/src/nanopycodeagent/terminal.py +0 -0
  148. {nanopycodeagent-0.7.0 → nanopycodeagent-0.9.0}/src/nanopycodeagent/write_tool.py +0 -0
  149. {nanopycodeagent-0.7.0 → nanopycodeagent-0.9.0}/tests/test_edit_tool.py +0 -0
  150. {nanopycodeagent-0.7.0 → nanopycodeagent-0.9.0}/tests/test_line_editing.py +0 -0
  151. {nanopycodeagent-0.7.0 → nanopycodeagent-0.9.0}/tests/test_read_tool.py +0 -0
  152. {nanopycodeagent-0.7.0 → nanopycodeagent-0.9.0}/tests/test_settings.py +0 -0
  153. {nanopycodeagent-0.7.0 → nanopycodeagent-0.9.0}/tests/test_terminal.py +0 -0
  154. {nanopycodeagent-0.7.0 → nanopycodeagent-0.9.0}/tests/test_write_tool.py +0 -0
@@ -8,8 +8,7 @@
8
8
  #
9
9
  # Usage: verify-release.sh <version|vVERSION> e.g. verify-release.sh 0.1.1
10
10
  #
11
- # Smoke test runs the entry point with no args (currently side-effect-free).
12
- # If the entry point ever gains required args/interaction, adjust step 3.
11
+ # Smoke test uses --version without a task and checks the installed version.
13
12
  set -euo pipefail
14
13
 
15
14
  PACKAGE="nanoPyCodeAgent"
@@ -63,11 +62,12 @@ printf 'OK: GitHub Release %s has wheel + sdist\n' "$tag"
63
62
  # --default-index (not --index): --index only *adds* an index, and uv's
64
63
  # first-index strategy would stop at a locally configured mirror that has the
65
64
  # package but not the new version yet. --refresh: uv's cached simple-index
66
- # metadata can be stale right after a publish. </dev/null: the entry point
67
- # reads stdin; EOF makes it exit immediately instead of hanging interactively.
65
+ # metadata can be stale right after a publish. --version exits before task input
66
+ # or credential validation; empty stdin without it is a command-line error.
68
67
  log "Smoke-testing uvx ${PACKAGE}@${version}"
69
68
  deadline=$(( $(date +%s) + POLL_TIMEOUT ))
70
- until uvx --refresh --default-index "$PYPI_INDEX" --from "${PACKAGE}@${version}" "$PACKAGE" >/dev/null 2>&1 </dev/null; do
69
+ until reported_version="$(uvx --refresh --default-index "$PYPI_INDEX" --from "${PACKAGE}@${version}" "$PACKAGE" --version 2>/dev/null </dev/null)" \
70
+ && [[ "$reported_version" == "${PACKAGE} ${version}" ]]; do
71
71
  (( $(date +%s) >= deadline )) && fail "uvx smoke test failed for ${PACKAGE}@${version} after ${POLL_TIMEOUT}s"
72
72
  printf 'uvx not ready (index propagation?); retrying in %ss...\n' "$POLL_INTERVAL"; sleep "$POLL_INTERVAL"
73
73
  done
@@ -39,3 +39,21 @@ jobs:
39
39
  # --frozen installs exactly what uv.lock pins and fails if the lock has
40
40
  # drifted from pyproject.toml, so CI tests the declared dependencies.
41
41
  run: uv run --frozen pytest
42
+ - name: Test the httpx2 SDK transport
43
+ run: uv run --with anthropic==1.5.0 pytest
44
+
45
+ harbor:
46
+ name: Harbor workflow contracts
47
+ runs-on: ubuntu-latest
48
+ steps:
49
+ - uses: actions/checkout@v6
50
+ with:
51
+ fetch-depth: 0
52
+ - name: Install uv
53
+ uses: astral-sh/setup-uv@08807647e7069bb48b6ef5acd8ec9567f424441b # v8.1.0
54
+ with:
55
+ python-version: "3.13"
56
+ - name: Test cache integrity, preflight, and verifier retries
57
+ run: >-
58
+ uv run --frozen --project benchmarks/harbor pytest
59
+ -c benchmarks/harbor/pyproject.toml benchmarks/harbor/tests
@@ -218,3 +218,6 @@ __marimo__/
218
218
  .streamlit/secrets.toml
219
219
 
220
220
  references/
221
+
222
+ # Harbor benchmark results
223
+ /jobs/
@@ -19,6 +19,24 @@ the matching one:
19
19
 
20
20
  See `docs/RELEASING.md` for the release process and prerequisites.
21
21
 
22
+ ## Benchmarks
23
+
24
+ For new Harbor benchmark runs, use the standard workflow from the repository
25
+ root: `uv run --project benchmarks/harbor python -m harbor_adapter run --config
26
+ <config.json>`. See `benchmarks/harbor/README.md` for configuration, verified APT
27
+ cache import, and the `--install-only` dependency preflight option.
28
+
29
+ - Do not create per-job adapter copies or ad hoc bootstrap scripts. Add reusable
30
+ fixes and reviewed task-revision profiles to `benchmarks/harbor` with tests.
31
+ - Preserve task images, package sources, and official test scripts. Cache hits
32
+ must match the current package index's filename, size, and SHA256 digest.
33
+ - Run dependency preflight before model work when a task has a registered
34
+ profile. Record unsupported profiles explicitly; do not claim they passed.
35
+ - Keep setup failures and verifier errors separate from scored failures. Only
36
+ verifier timeouts may retry, at most three total attempts using the same agent
37
+ output and the task's original per-attempt time limit.
38
+
39
+
22
40
  ## Language
23
41
 
24
42
  Write everything in **English** by default — source code (identifiers,
@@ -37,11 +55,15 @@ The **only** exception is documentation explicitly designated as Chinese:
37
55
  Everything outside that list — including code under `src/`, this file, and all
38
56
  other docs — is English.
39
57
 
58
+ Do not label a document as `hand-written` unless its contents are actually
59
+ authored and maintained by humans. For agent-authored source-of-truth documents,
60
+ use wording such as `Chinese source` instead.
61
+
40
62
  ### Bilingual research notes
41
63
 
42
64
  Research notes are bilingual:
43
65
 
44
- - `docs/research/zh-CN/` contains the hand-written Chinese sources of truth.
66
+ - `docs/research/zh-CN/` contains the Chinese sources of truth.
45
67
  - `docs/research/en/` contains English versions generated from the Chinese
46
68
  sources; regenerate the whole corresponding file instead of hand-editing it.
47
69
 
@@ -52,6 +74,28 @@ the English version is in sync, even when it has already been updated.
52
74
 
53
75
  ## Commits & PRs
54
76
 
77
+ ### Branch and pull request workflow
78
+
79
+ All non-release changes MUST be made on a dedicated branch and merged into
80
+ `main` through a GitHub pull request.
81
+
82
+ - Before modifying repository files, agents MUST verify that the current branch
83
+ is not `main`. If it is `main`, create or switch to a dedicated branch first.
84
+ - Agents MUST NOT commit directly to `main` or push non-release commits to
85
+ `origin/main`.
86
+ - Open a pull request targeting `main`, and use the `land-pr` skill to merge it.
87
+ Do not merge without the skill's mandatory maintainer confirmation.
88
+ - If work was accidentally started on `main`, move the changes to a dedicated
89
+ branch before committing or pushing them.
90
+
91
+ The sole exception is the release workflow in `.agents/skills/release/SKILL.md`.
92
+ Only when the maintainer explicitly requests a release may that skill commit and
93
+ push its release-only changes directly to `main`, and only after its mandatory
94
+ confirmation gate. Documentation, chores, hotfixes, and all other work still
95
+ require a branch and pull request.
96
+
97
+ ### Commit and PR title format
98
+
55
99
  All agents working in this repo MUST follow
56
100
  [Conventional Commits 1.0.0](https://www.conventionalcommits.org/en/v1.0.0/)
57
101
  for **commit messages** and **PR titles** (squash-merge often uses the PR
@@ -0,0 +1,65 @@
1
+ # Code Agent Runtime
2
+
3
+ This glossary defines the runtime artifacts produced and maintained by a code agent so that public output, diagnostics, evaluation data, and resumable state remain distinct.
4
+
5
+ ## Language
6
+
7
+ **Agent Run**:
8
+ A bounded execution that starts with an instruction and ends when the agent stops autonomous work.
9
+ _Avoid_: Session, turn
10
+
11
+ **Run Output**:
12
+ The public result or event protocol that one agent run emits to its caller.
13
+ _Avoid_: Trace, transcript
14
+
15
+ **Run Result Object**:
16
+ A single public object that summarizes how one Agent Run ended, including its final result and optional usage metadata.
17
+ _Avoid_: Trajectory, session export
18
+
19
+ **Public Event Stream**:
20
+ A caller-facing Run Output that exposes a stable, selected sequence of Agent Run events as they occur.
21
+ _Avoid_: Internal event bus, execution trace
22
+
23
+ **Source Record**:
24
+ An agent-specific record observed at a CLI, SDK, or protocol boundary and consumed by an adapter; it may already be a projection of the agent's internal state.
25
+ _Avoid_: Native event, journal entry
26
+
27
+ **Native Event**:
28
+ An agent-independent runtime fact accepted by the core directly or normalized from a Source Record.
29
+ _Avoid_: Source record, public event record, trajectory step
30
+
31
+ **Journal Entry**:
32
+ A Native Event together with the identity, ordering, and recording metadata required for durable history.
33
+ _Avoid_: Persistence envelope, trajectory step
34
+
35
+ **Event Journal**:
36
+ A durable append-only sequence of Journal Entries for an Agent Run from which derived artifacts can be rebuilt.
37
+ _Avoid_: Trajectory, session, public event stream
38
+
39
+ **Partial Content**:
40
+ An unfinished representation of a message, reasoning block, or tool input that is exposed before the logical content is complete.
41
+ _Avoid_: Final message, completed event
42
+
43
+ **Content Delta**:
44
+ The incremental fragment or change relative to content already emitted for the same logical item.
45
+ _Avoid_: Partial snapshot, final message
46
+
47
+ **Terminal Event**:
48
+ The final Public Event Stream record that explicitly states how an Agent Run ended.
49
+ _Avoid_: EOF, last message
50
+
51
+ **Execution Trace**:
52
+ A diagnostic record of runtime activity retained to explain failures, timing, and internal behavior.
53
+ _Avoid_: Run output, trajectory
54
+
55
+ **Trajectory**:
56
+ A task-scoped record of observations, actions, results, and outcome prepared for evaluation or offline analysis.
57
+ _Avoid_: Execution trace, session, transcript
58
+
59
+ **ATIF Trajectory**:
60
+ A Trajectory represented by the Agent Trajectory Interchange Format and exposed as the runtime's public trajectory contract.
61
+ _Avoid_: Event journal, native trajectory
62
+
63
+ **Session**:
64
+ Durable agent state that can span multiple runs and supports continuing or branching prior work.
65
+ _Avoid_: Agent run, trajectory, transcript
@@ -1,6 +1,6 @@
1
1
  Metadata-Version: 2.5
2
2
  Name: nanoPyCodeAgent
3
- Version: 0.7.0
3
+ Version: 0.9.0
4
4
  Summary: A nano code agent built from scratch in pure Python.
5
5
  Project-URL: Homepage, https://github.com/minixalpha/nanoPyCodeAgent
6
6
  Project-URL: Repository, https://github.com/minixalpha/nanoPyCodeAgent
@@ -14,6 +14,7 @@ Classifier: Programming Language :: Python :: 3
14
14
  Classifier: Programming Language :: Python :: 3.13
15
15
  Requires-Python: >=3.13
16
16
  Requires-Dist: anthropic>=0.112.0
17
+ Requires-Dist: httpx>=0.25.0
17
18
  Description-Content-Type: text/markdown
18
19
 
19
20
  # nanoPyCodeAgent
@@ -49,6 +50,20 @@ uv tool install nanoPyCodeAgent # or: pipx install nanoPyCodeAgent
49
50
  nanoPyCodeAgent
50
51
  ```
51
52
 
53
+ #### Run one task and exit
54
+
55
+ Give it a task and it works through it on its own, with no prompt and nothing
56
+ to confirm — the shape a script or a benchmark harness needs:
57
+
58
+ ```bash
59
+ nanoPyCodeAgent -p "add a --version flag and run the tests"
60
+ ```
61
+
62
+ The task runs in the current directory and the command exits when the run
63
+ ends. See the [complete CLI reference](docs/user_docs/en/cli_reference.md) for
64
+ all task input methods, options, exit statuses, trajectory output, and Event
65
+ Journal behavior.
66
+
52
67
  #### Run a branch or tagged version
53
68
 
54
69
  Run an unreleased branch or a specific release tag straight from GitHub:
@@ -63,29 +78,11 @@ uvx --from "git+https://github.com/minixalpha/nanoPyCodeAgent@v0.1.0" nanoPyCode
63
78
 
64
79
  ### Configuration
65
80
 
66
- Credentials and the model come from two sources: **environment variables** and
67
- an optional user-level config file at `~/.nanoPyCodeAgent/settings.json`.
68
- Environment variables take precedence — the config file only fills in keys you
69
- have not set in the environment.
70
-
71
- The config file mirrors [Claude Code's settings](https://code.claude.com/docs/en/settings):
72
- put the values under an `env` object. Empty or whitespace-only values are ignored.
73
-
74
- ```json
75
- {
76
- "env": {
77
- "ANTHROPIC_API_KEY": "",
78
- "ANTHROPIC_BASE_URL": "",
79
- "ANTHROPIC_MODEL": ""
80
- }
81
- }
82
- ```
83
-
84
- | Variable | Required | Default | Description |
85
- | --- | --- | --- | --- |
86
- | `ANTHROPIC_API_KEY` | Yes | — | Your Anthropic API key, or the key for a third-party / proxy service. |
87
- | `ANTHROPIC_BASE_URL` | No | `https://api.anthropic.com` | Point the SDK at a non-official / proxy endpoint. Leave it unset to use the official API — an empty value breaks requests. |
88
- | `ANTHROPIC_MODEL` | No | `claude-sonnet-4-6` | Override the model. An empty or whitespace-only value falls back to the default. |
81
+ Set `ANTHROPIC_API_KEY` or `ANTHROPIC_AUTH_TOKEN` before running the agent.
82
+ You can also configure the endpoint and model through environment variables or
83
+ `~/.nanoPyCodeAgent/settings.json`. See the [configuration reference](docs/user_docs/en/configuration.md)
84
+ for supported variables, defaults, file format, precedence, and empty-value
85
+ handling.
89
86
 
90
87
  ### How to Update
91
88
 
@@ -31,6 +31,20 @@ uv tool install nanoPyCodeAgent # or: pipx install nanoPyCodeAgent
31
31
  nanoPyCodeAgent
32
32
  ```
33
33
 
34
+ #### Run one task and exit
35
+
36
+ Give it a task and it works through it on its own, with no prompt and nothing
37
+ to confirm — the shape a script or a benchmark harness needs:
38
+
39
+ ```bash
40
+ nanoPyCodeAgent -p "add a --version flag and run the tests"
41
+ ```
42
+
43
+ The task runs in the current directory and the command exits when the run
44
+ ends. See the [complete CLI reference](docs/user_docs/en/cli_reference.md) for
45
+ all task input methods, options, exit statuses, trajectory output, and Event
46
+ Journal behavior.
47
+
34
48
  #### Run a branch or tagged version
35
49
 
36
50
  Run an unreleased branch or a specific release tag straight from GitHub:
@@ -45,29 +59,11 @@ uvx --from "git+https://github.com/minixalpha/nanoPyCodeAgent@v0.1.0" nanoPyCode
45
59
 
46
60
  ### Configuration
47
61
 
48
- Credentials and the model come from two sources: **environment variables** and
49
- an optional user-level config file at `~/.nanoPyCodeAgent/settings.json`.
50
- Environment variables take precedence — the config file only fills in keys you
51
- have not set in the environment.
52
-
53
- The config file mirrors [Claude Code's settings](https://code.claude.com/docs/en/settings):
54
- put the values under an `env` object. Empty or whitespace-only values are ignored.
55
-
56
- ```json
57
- {
58
- "env": {
59
- "ANTHROPIC_API_KEY": "",
60
- "ANTHROPIC_BASE_URL": "",
61
- "ANTHROPIC_MODEL": ""
62
- }
63
- }
64
- ```
65
-
66
- | Variable | Required | Default | Description |
67
- | --- | --- | --- | --- |
68
- | `ANTHROPIC_API_KEY` | Yes | — | Your Anthropic API key, or the key for a third-party / proxy service. |
69
- | `ANTHROPIC_BASE_URL` | No | `https://api.anthropic.com` | Point the SDK at a non-official / proxy endpoint. Leave it unset to use the official API — an empty value breaks requests. |
70
- | `ANTHROPIC_MODEL` | No | `claude-sonnet-4-6` | Override the model. An empty or whitespace-only value falls back to the default. |
62
+ Set `ANTHROPIC_API_KEY` or `ANTHROPIC_AUTH_TOKEN` before running the agent.
63
+ You can also configure the endpoint and model through environment variables or
64
+ `~/.nanoPyCodeAgent/settings.json`. See the [configuration reference](docs/user_docs/en/configuration.md)
65
+ for supported variables, defaults, file format, precedence, and empty-value
66
+ handling.
71
67
 
72
68
  ### How to Update
73
69
 
@@ -31,6 +31,18 @@ uv tool install nanoPyCodeAgent # 或: pipx install nanoPyCodeAgent
31
31
  nanoPyCodeAgent
32
32
  ```
33
33
 
34
+ #### 一次性任务(非交互)
35
+
36
+ 给它一个任务,它会自己做完退出——不给提示符,也不会停下来等确认,这正是脚本或
37
+ benchmark harness 需要的形态:
38
+
39
+ ```bash
40
+ nanoPyCodeAgent -p "add a --version flag and run the tests"
41
+ ```
42
+
43
+ 任务在当前目录下执行,run 结束后命令随即退出。全部任务输入方式、选项、退出状态、
44
+ trajectory 输出和 Event Journal 行为请参阅[完整 CLI 参考](docs/user_docs/zh-CN/cli_reference.md)。
45
+
34
46
  #### 运行某个分支或标签版本
35
47
 
36
48
  直接从 GitHub 运行未发布的分支,或某个具体的发布标签:
@@ -45,28 +57,9 @@ uvx --from "git+https://github.com/minixalpha/nanoPyCodeAgent@v0.1.0" nanoPyCode
45
57
 
46
58
  ### 配置
47
59
 
48
- 凭据与模型有两种配置来源:**环境变量**,以及可选的用户级配置文件
49
- `~/.nanoPyCodeAgent/settings.json`。环境变量优先级更高——配置文件只用于填补你
50
- 未在环境变量中设置的键。
51
-
52
- 配置文件写法与 [Claude Code settings](https://code.claude.com/docs/en/settings)
53
- 一致:把值放在 `env` 对象下。空值或纯空白会被忽略。
54
-
55
- ```json
56
- {
57
- "env": {
58
- "ANTHROPIC_API_KEY": "",
59
- "ANTHROPIC_BASE_URL": "",
60
- "ANTHROPIC_MODEL": ""
61
- }
62
- }
63
- ```
64
-
65
- | 变量 | 是否必填 | 默认值 | 说明 |
66
- | --- | --- | --- | --- |
67
- | `ANTHROPIC_API_KEY` | 是 | 无 | 你的 Anthropic API key,或第三方/代理服务的 key。 |
68
- | `ANTHROPIC_BASE_URL` | 否 | `https://api.anthropic.com` | 将 SDK 指向非官方/代理 endpoint。使用官方 API 时保持不设置;留空值会导致请求失败。 |
69
- | `ANTHROPIC_MODEL` | 否 | `claude-sonnet-4-6` | 覆盖默认模型。空值或纯空白会回退到默认值。 |
60
+ 运行前请设置 `ANTHROPIC_API_KEY` 或 `ANTHROPIC_AUTH_TOKEN`。你也可以通过
61
+ 环境变量或 `~/.nanoPyCodeAgent/settings.json` 配置 endpoint 与模型。支持的变量、
62
+ 默认值、文件格式、优先级和空值处理请参阅[配置参考](docs/user_docs/zh-CN/configuration.md)。
70
63
 
71
64
  ### 如何更新
72
65
 
@@ -0,0 +1,196 @@
1
+ # Harbor adapter
2
+
3
+ [English](README.md) | [简体中文](README.zh-CN.md)
4
+
5
+ This isolated workspace contains the nanoPyCodeAgent adapter for Terminal-Bench
6
+ and other benchmarks run by Harbor. It is development infrastructure, not part
7
+ of the end-user `nanoPyCodeAgent` package. Harbor is pinned to 0.21.0 in this
8
+ workspace's lockfile.
9
+
10
+ ## Run a benchmark
11
+
12
+ Set the connection credentials used by the CLI. For a third-party or proxy
13
+ endpoint, configure the API key and base URL; select the model with Harbor's
14
+ `--model` option shown below:
15
+
16
+ ```bash
17
+ export ANTHROPIC_API_KEY="..."
18
+ export ANTHROPIC_BASE_URL="https://gateway.example"
19
+ ```
20
+
21
+ Run new benchmarks through the standard entry point, from the repository root:
22
+
23
+ ```bash
24
+ uv run --project benchmarks/harbor python -m harbor_adapter run \
25
+ --config benchmarks/harbor/configs/qemu-alpine-ssh.json
26
+ ```
27
+
28
+ This example pins the QEMU task revision and uses `deepseek/deepseek-flash`.
29
+ Set the endpoint and credentials for that model, or copy the configuration and
30
+ change `agents[].model_name`, task selection, and budgets for another run. These
31
+ are ordinary Harbor JSON configurations; the entry point supplies the standard
32
+ verifier and cache policy. It accepts the repository's `NanoPyCodeAgent` adapter
33
+ with budget kwargs; credentials belong in the environment.
34
+
35
+ The entry point builds the current working tree, including uncommitted changes,
36
+ into a wheel and checks its Python sources against the recorded snapshot. It
37
+ also snapshots the adapter, records source and wheel hashes, and uses that
38
+ snapshot throughout the run. Inputs and logs live in `jobs/<job-name>-input/`;
39
+ Harbor outputs remain in `jobs/<job-name>/`. Existing jobs are never overwritten.
40
+ `summary.json` records scores, setup/preflight status, cache use, and verifier
41
+ attempts. An infrastructure exception has no score; it is not converted to zero.
42
+ The command exits nonzero for execution errors or incomplete results, while a
43
+ completed score of zero is a valid benchmark result.
44
+
45
+ Use `--job-name <unique-name>` for a readable run name. Add `--prepare-only` to
46
+ build and record inputs without starting containers. Add `--install-only` to
47
+ exercise dependency installation and preflight without calling the model or
48
+ executing the official tests. A full run still performs setup and preflight in
49
+ its own fresh container; an install-only check does not resume into model work.
50
+
51
+ ### Verified dependencies and preflight
52
+
53
+ The standard adapter automatically uses `.cache/nanopy-harbor/`, shared across
54
+ runs and ignored by Git. Override it with `NANOPY_HARBOR_CACHE` or the entry
55
+ point's `--cache-dir`. On APT systems it refreshes the package index, resolves a
56
+ SHA256 download plan, restores matching cached files, and downloads missing
57
+ ones. Successful downloads are verified and saved before installation, so
58
+ APT cleanup hooks cannot erase the shared copies. A different version or digest
59
+ does not reuse an old cached file. Corrupt files fail setup before model work.
60
+ Other package managers use Harbor's regular dependency installation.
61
+
62
+ To import previously verified downloads once:
63
+
64
+ ```bash
65
+ uv run --project benchmarks/harbor python -m harbor_adapter import-apt \
66
+ /path/to/bootstrap-cache/manifest.json
67
+ ```
68
+
69
+ The manifest is a JSON array of objects with `url`, `filename`, `size`, and
70
+ `sha256`, accompanied by `packages/<filename>` files. Import validates every
71
+ file. At use time, the current APT index must independently match each entry.
72
+ Copying the entire shared cache to another machine is also supported. A new
73
+ machine with no matching cache still needs the original downloads to be
74
+ available; this mechanism does not replace package sources or invent missing
75
+ versions.
76
+
77
+ Verifier dependency profiles are tied to reviewed task revisions in
78
+ `src/harbor_adapter/bootstrap.py`. The current QEMU profile installs `sshpass`
79
+ and warms the original verifier's uv 0.9.5, Python 3.13, pytest 8.4.1, and
80
+ pytest-json-ctrf 0.3.5 dependencies using `pytest --version`. It runs no task
81
+ assertions or solution code. Unknown QEMU revisions fail setup pending profile
82
+ review; tasks without a profile explicitly report `not_configured`. Add a
83
+ profile and regression tests when another task needs verifier preflight.
84
+
85
+ Every installation writes `agent/bootstrap.json`, including partial progress
86
+ on failure. Preflight failure stops the trial before any model call. The
87
+ official verifier script remains unchanged and may still access the network
88
+ when it runs; a successful preflight is not a guarantee against later outages.
89
+ Full standard runs enable the timeout retry policy below by default, with zero
90
+ whole-trial retries. CI checks these contracts without model calls or Docker.
91
+
92
+ ### Direct Harbor invocation
93
+
94
+ For adapter development or comparison with published versions, the lower-level
95
+ Harbor interface remains available. Pin the agent installed in the container:
96
+
97
+ ```bash
98
+ uv run --project benchmarks/harbor harbor run \
99
+ --task terminal-bench/openssl-selfsigned-cert \
100
+ --agent harbor_adapter:NanoPyCodeAgent \
101
+ --agent-kwarg git_ref=<commit-sha> \
102
+ --model anthropic/claude-sonnet-4-6 \
103
+ --verifier harbor_adapter:RetryingVerifier \
104
+ --verifier-timeout-multiplier 4 \
105
+ --env docker \
106
+ --n-concurrent 1 \
107
+ --n-attempts 1
108
+ ```
109
+
110
+ Use `--agent-kwarg version=<released-version>` instead of `git_ref` to install a
111
+ published PyPI release. The two pins are mutually exclusive. If neither is
112
+ provided, the adapter installs the latest published release. For reproducible
113
+ benchmark results, always provide one of them. Prefer a full 40-character commit
114
+ SHA for `git_ref`; Harbor may parse an unquoted abbreviated SHA such as `83e6271`
115
+ as a number. To use an abbreviated revision, preserve its string type with
116
+ `--agent-kwarg 'git_ref="83e6271"'`.
117
+
118
+ The adapter sends the instruction through stdin, runs the agent in the task
119
+ container's current directory, and saves combined stdout/stderr to
120
+ `/logs/agent/nanopycodeagent.txt`. It uses the CLI's 50-turn default; override
121
+ that with `--agent-kwarg max_turns=20`.
122
+
123
+ Use `--agent-kwarg time_budget_seconds=N` to bound task work and remind the model
124
+ to finalize before either time or replies run out. Choose a budget inside each
125
+ task's native timeout, reserving at least 30 seconds for cost reconciliation
126
+ plus trajectory writing and harness overhead. For example, the 3600-second
127
+ targeted experiments use 3420 seconds; that value does not fit a 900-second task.
128
+ No time budget is inferred automatically by the adapter.
129
+
130
+ For the per-reply generation limit, pass `--agent-kwarg max_tokens=32768`.
131
+ This becomes `--max-tokens 32768` in the container and overrides the forwarded
132
+ `ANTHROPIC_MAX_TOKENS` environment variable. When omitted, the adapter sends no
133
+ token flag, so the installed agent's environment/settings/default applies
134
+ (32768 in the version introducing this option). Older releases require omitting
135
+ the new option. The effective budget is recorded in the startup log and in
136
+ `agent.extra.max_tokens` in the ATIF trajectory.
137
+
138
+ The adapter also asks the agent to write an ATIF-v1.7 trajectory directly to
139
+ `/logs/agent/trajectory.json`. Harbor collects that file as the trial's native
140
+ ATIF output and backfills prompt, completion, cache-token, and cost totals into
141
+ the agent result. Missing, invalid, or partial trajectories remain explicitly
142
+ diagnosed; unknown usage or cost is not reported as zero.
143
+ If any costs are estimated, `final_metrics.extra` records
144
+ `cost_is_estimated: true` and the estimated portion in `estimated_cost_usd`.
145
+ The adapter preserves both fields in the agent result's `metadata.trajectory`,
146
+ including for partial totals. Estimates with a generation ID remain eligible
147
+ for billing reconciliation; a successful lookup replaces the estimate.
148
+
149
+ By default, the adapter strips the first provider prefix from `--model` and
150
+ passes the result to nanoPyCodeAgent as `ANTHROPIC_MODEL`. Set
151
+ `ANTHROPIC_MODEL` only when a custom endpoint requires an actual model name that
152
+ differs from Harbor's `provider/model` identity; this explicit override takes
153
+ precedence. Harbor-native provider credentials and configured base URLs are
154
+ also normalized to the `ANTHROPIC_*` variables expected by nanoPyCodeAgent's
155
+ SDK.
156
+
157
+ ## Retry verifier timeouts
158
+
159
+ The example above uses `harbor_adapter:RetryingVerifier` for single-step Linux
160
+ tasks. It runs the original test script up to **three times total**, retrying
161
+ only when the script exceeds the task's native `[verifier].timeout_sec`. A
162
+ 900-second task therefore gets at most three 900-second attempts. A completed
163
+ verification stops immediately, including a reward of zero. Script errors,
164
+ missing or invalid rewards, and cancellation do not trigger another attempt.
165
+
166
+ `--verifier-timeout-multiplier 4` gives Harbor's outer watchdog room for all
167
+ three attempts plus test upload, process cleanup, and log collection. It does
168
+ not extend each test-script attempt. Conflicting timeout overrides or caps
169
+ that leave insufficient overall time are rejected. Use
170
+ `--verifier-kwarg max_attempts=2` for two total attempts, or `max_attempts=1` for
171
+ one; values above three are rejected.
172
+
173
+ The runner reuses the same container, agent output, and dependency caches;
174
+ the agent is run once. Test scripts must tolerate being rerun in that
175
+ environment; verifier-side filesystem changes are not rolled back. Timed-out
176
+ process groups and other members of their Linux session are killed before
177
+ retrying. Each attempt's logs and any reward files are archived under
178
+ `verifier/attempts/1/`, `2/`, and `3/`. The runner clears current verifier
179
+ outputs before the next attempt, so a stale reward cannot determine its score.
180
+ `verifier/retry-summary.json` records each attempt's status, duration, and exit
181
+ code when available. Exhausting the attempts still produces Harbor's
182
+ `VerifierTimeoutError`.
183
+
184
+ This verifier requires `python3` and `bash` in the container, which the
185
+ nanoPyCodeAgent adapter installs. These options apply to new runs; they cannot
186
+ restore a deleted container or missing task outputs from an old trial.
187
+ Harbor's ordinary `--max-retries` setting reruns the entire trial, including
188
+ the agent, and is independent of these verifier attempts.
189
+
190
+ ## Test the adapter
191
+
192
+ ```bash
193
+ uv run --project benchmarks/harbor pytest \
194
+ -c benchmarks/harbor/pyproject.toml \
195
+ benchmarks/harbor/tests
196
+ ```