pyattacker 0.2.0__tar.gz → 0.3.0__tar.gz
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- {pyattacker-0.2.0 → pyattacker-0.3.0}/.github/workflows/ci.yml +12 -0
- {pyattacker-0.2.0 → pyattacker-0.3.0}/.github/workflows/release.yml +88 -9
- {pyattacker-0.2.0 → pyattacker-0.3.0}/CHANGELOG.md +211 -2
- {pyattacker-0.2.0 → pyattacker-0.3.0}/PKG-INFO +132 -11
- {pyattacker-0.2.0 → pyattacker-0.3.0}/README.md +131 -10
- pyattacker-0.3.0/README.zh-CN.md +417 -0
- {pyattacker-0.2.0 → pyattacker-0.3.0}/docs/benchmark.md +2 -0
- {pyattacker-0.2.0 → pyattacker-0.3.0}/docs/cli.md +95 -10
- {pyattacker-0.2.0 → pyattacker-0.3.0}/docs/design.md +427 -18
- {pyattacker-0.2.0 → pyattacker-0.3.0}/docs/reference.md +680 -31
- {pyattacker-0.2.0 → pyattacker-0.3.0}/docs/releasing.md +32 -12
- {pyattacker-0.2.0 → pyattacker-0.3.0}/docs/tutorial.md +475 -17
- pyattacker-0.3.0/docs/zh-CN/benchmark.md +208 -0
- pyattacker-0.3.0/docs/zh-CN/cli.md +283 -0
- pyattacker-0.3.0/docs/zh-CN/design.md +702 -0
- pyattacker-0.3.0/docs/zh-CN/reference.md +2149 -0
- pyattacker-0.3.0/docs/zh-CN/releasing.md +122 -0
- pyattacker-0.3.0/docs/zh-CN/tutorial.md +1862 -0
- {pyattacker-0.2.0 → pyattacker-0.3.0}/examples/llm_eval/README.md +10 -0
- pyattacker-0.3.0/examples/llm_eval/README.zh-CN.md +102 -0
- {pyattacker-0.2.0 → pyattacker-0.3.0}/examples/plugin_package/README.md +2 -0
- pyattacker-0.3.0/examples/plugin_package/README.zh-CN.md +22 -0
- {pyattacker-0.2.0 → pyattacker-0.3.0}/pyproject.toml +11 -2
- {pyattacker-0.2.0 → pyattacker-0.3.0}/src/pyattacker/__init__.py +21 -2
- {pyattacker-0.2.0 → pyattacker-0.3.0}/src/pyattacker/artifact.py +14 -1
- {pyattacker-0.2.0 → pyattacker-0.3.0}/src/pyattacker/cli.py +37 -1
- {pyattacker-0.2.0 → pyattacker-0.3.0}/src/pyattacker/declarative.py +17 -4
- {pyattacker-0.2.0 → pyattacker-0.3.0}/src/pyattacker/errors.py +48 -0
- {pyattacker-0.2.0 → pyattacker-0.3.0}/src/pyattacker/export.py +6 -2
- pyattacker-0.3.0/src/pyattacker/handoff.py +499 -0
- pyattacker-0.3.0/src/pyattacker/history.py +146 -0
- {pyattacker-0.2.0 → pyattacker-0.3.0}/src/pyattacker/merge.py +97 -11
- {pyattacker-0.2.0 → pyattacker-0.3.0}/src/pyattacker/monitor.py +4 -1
- {pyattacker-0.2.0 → pyattacker-0.3.0}/src/pyattacker/pipeline.py +73 -6
- pyattacker-0.3.0/src/pyattacker/runner.py +2584 -0
- {pyattacker-0.2.0 → pyattacker-0.3.0}/src/pyattacker/server.py +35 -19
- {pyattacker-0.2.0 → pyattacker-0.3.0}/src/pyattacker/store/__init__.py +18 -0
- {pyattacker-0.2.0 → pyattacker-0.3.0}/src/pyattacker/store/base.py +219 -7
- {pyattacker-0.2.0 → pyattacker-0.3.0}/src/pyattacker/store/memory.py +288 -17
- {pyattacker-0.2.0 → pyattacker-0.3.0}/src/pyattacker/store/sqlite.py +513 -30
- pyattacker-0.3.0/src/pyattacker/store/visits.py +352 -0
- {pyattacker-0.2.0 → pyattacker-0.3.0}/src/pyattacker/store/writebehind.py +101 -5
- {pyattacker-0.2.0 → pyattacker-0.3.0}/src/pyattacker/task.py +4 -0
- {pyattacker-0.2.0 → pyattacker-0.3.0}/src/pyattacker/tasks/__init__.py +16 -0
- pyattacker-0.3.0/tests/test_backward.py +1459 -0
- {pyattacker-0.2.0 → pyattacker-0.3.0}/tests/test_cli.py +292 -0
- {pyattacker-0.2.0 → pyattacker-0.3.0}/tests/test_docs_examples.py +10 -5
- pyattacker-0.3.0/tests/test_docs_facts.py +126 -0
- pyattacker-0.3.0/tests/test_docs_i18n.py +222 -0
- {pyattacker-0.2.0 → pyattacker-0.3.0}/tests/test_export.py +245 -1
- pyattacker-0.3.0/tests/test_handoff.py +1480 -0
- {pyattacker-0.2.0 → pyattacker-0.3.0}/tests/test_packaging.py +17 -2
- pyattacker-0.3.0/tests/test_resume_cursor.py +659 -0
- {pyattacker-0.2.0 → pyattacker-0.3.0}/tests/test_store.py +102 -0
- pyattacker-0.3.0/tests/test_worker_liveness.py +425 -0
- {pyattacker-0.2.0 → pyattacker-0.3.0}/uv.lock +1 -1
- pyattacker-0.2.0/assets/logo/icon.png +0 -0
- pyattacker-0.2.0/assets/logo/title.png +0 -0
- pyattacker-0.2.0/src/pyattacker/runner.py +0 -1441
- {pyattacker-0.2.0 → pyattacker-0.3.0}/.gitignore +0 -0
- {pyattacker-0.2.0 → pyattacker-0.3.0}/LICENSE +0 -0
- {pyattacker-0.2.0 → pyattacker-0.3.0}/examples/data.jsonl +0 -0
- {pyattacker-0.2.0 → pyattacker-0.3.0}/examples/llm_eval/__init__.py +0 -0
- {pyattacker-0.2.0 → pyattacker-0.3.0}/examples/llm_eval/backend.py +0 -0
- {pyattacker-0.2.0 → pyattacker-0.3.0}/examples/llm_eval/demo.py +0 -0
- {pyattacker-0.2.0 → pyattacker-0.3.0}/examples/llm_eval/pipelines.py +0 -0
- {pyattacker-0.2.0 → pyattacker-0.3.0}/examples/plugin_package/pa_demo_plugin/__init__.py +0 -0
- {pyattacker-0.2.0 → pyattacker-0.3.0}/examples/plugin_package/pa_demo_plugin/algorithms.py +0 -0
- {pyattacker-0.2.0 → pyattacker-0.3.0}/examples/plugin_package/pa_demo_plugin/codecs.py +0 -0
- {pyattacker-0.2.0 → pyattacker-0.3.0}/examples/plugin_package/pa_demo_plugin/tasks.py +0 -0
- {pyattacker-0.2.0 → pyattacker-0.3.0}/examples/plugin_package/pyproject.toml +0 -0
- {pyattacker-0.2.0 → pyattacker-0.3.0}/examples/plugin_tasks.yaml +0 -0
- {pyattacker-0.2.0 → pyattacker-0.3.0}/examples/qa_eval.yaml +0 -0
- {pyattacker-0.2.0 → pyattacker-0.3.0}/examples/quickstart.py +0 -0
- {pyattacker-0.2.0 → pyattacker-0.3.0}/examples/sharded.py +0 -0
- {pyattacker-0.2.0 → pyattacker-0.3.0}/src/pyattacker/__main__.py +0 -0
- {pyattacker-0.2.0 → pyattacker-0.3.0}/src/pyattacker/algorithm.py +0 -0
- {pyattacker-0.2.0 → pyattacker-0.3.0}/src/pyattacker/backends.py +0 -0
- {pyattacker-0.2.0 → pyattacker-0.3.0}/src/pyattacker/benchmark/__init__.py +0 -0
- {pyattacker-0.2.0 → pyattacker-0.3.0}/src/pyattacker/benchmark/clock.py +0 -0
- {pyattacker-0.2.0 → pyattacker-0.3.0}/src/pyattacker/benchmark/harness.py +0 -0
- {pyattacker-0.2.0 → pyattacker-0.3.0}/src/pyattacker/benchmark/metrics.py +0 -0
- {pyattacker-0.2.0 → pyattacker-0.3.0}/src/pyattacker/benchmark/provider.py +0 -0
- {pyattacker-0.2.0 → pyattacker-0.3.0}/src/pyattacker/benchmark/report.py +0 -0
- {pyattacker-0.2.0 → pyattacker-0.3.0}/src/pyattacker/benchmark/scenario.py +0 -0
- {pyattacker-0.2.0 → pyattacker-0.3.0}/src/pyattacker/plugins.py +0 -0
- {pyattacker-0.2.0 → pyattacker-0.3.0}/src/pyattacker/resource.py +0 -0
- {pyattacker-0.2.0 → pyattacker-0.3.0}/src/pyattacker/scheduler.py +0 -0
- {pyattacker-0.2.0 → pyattacker-0.3.0}/src/pyattacker/shard.py +0 -0
- {pyattacker-0.2.0 → pyattacker-0.3.0}/tests/helpers.py +0 -0
- {pyattacker-0.2.0 → pyattacker-0.3.0}/tests/test_algorithms.py +0 -0
- {pyattacker-0.2.0 → pyattacker-0.3.0}/tests/test_artifact.py +0 -0
- {pyattacker-0.2.0 → pyattacker-0.3.0}/tests/test_backends.py +0 -0
- {pyattacker-0.2.0 → pyattacker-0.3.0}/tests/test_benchmark_cli.py +0 -0
- {pyattacker-0.2.0 → pyattacker-0.3.0}/tests/test_benchmark_clock.py +0 -0
- {pyattacker-0.2.0 → pyattacker-0.3.0}/tests/test_benchmark_harness.py +0 -0
- {pyattacker-0.2.0 → pyattacker-0.3.0}/tests/test_benchmark_provider.py +0 -0
- {pyattacker-0.2.0 → pyattacker-0.3.0}/tests/test_benchmark_report.py +0 -0
- {pyattacker-0.2.0 → pyattacker-0.3.0}/tests/test_benchmark_scenario.py +0 -0
- {pyattacker-0.2.0 → pyattacker-0.3.0}/tests/test_declarative.py +0 -0
- {pyattacker-0.2.0 → pyattacker-0.3.0}/tests/test_errors.py +0 -0
- {pyattacker-0.2.0 → pyattacker-0.3.0}/tests/test_lease_safety.py +0 -0
- {pyattacker-0.2.0 → pyattacker-0.3.0}/tests/test_m2.py +0 -0
- {pyattacker-0.2.0 → pyattacker-0.3.0}/tests/test_monitor.py +0 -0
- {pyattacker-0.2.0 → pyattacker-0.3.0}/tests/test_optional_yaml.py +0 -0
- {pyattacker-0.2.0 → pyattacker-0.3.0}/tests/test_pipeline.py +0 -0
- {pyattacker-0.2.0 → pyattacker-0.3.0}/tests/test_plugins.py +0 -0
- {pyattacker-0.2.0 → pyattacker-0.3.0}/tests/test_resume_identity.py +0 -0
- {pyattacker-0.2.0 → pyattacker-0.3.0}/tests/test_retry_policy.py +0 -0
- {pyattacker-0.2.0 → pyattacker-0.3.0}/tests/test_runner.py +0 -0
- {pyattacker-0.2.0 → pyattacker-0.3.0}/tests/test_scheduler.py +0 -0
- {pyattacker-0.2.0 → pyattacker-0.3.0}/tests/test_server.py +0 -0
- {pyattacker-0.2.0 → pyattacker-0.3.0}/tests/test_shard.py +0 -0
- {pyattacker-0.2.0 → pyattacker-0.3.0}/tests/test_subprocess_lifecycle.py +0 -0
- {pyattacker-0.2.0 → pyattacker-0.3.0}/tests/test_task_overrides.py +0 -0
- {pyattacker-0.2.0 → pyattacker-0.3.0}/tests/test_tasks.py +0 -0
- {pyattacker-0.2.0 → pyattacker-0.3.0}/tests/test_tutorial.py +0 -0
|
@@ -55,6 +55,18 @@ jobs:
|
|
|
55
55
|
- name: Build sdist and wheel
|
|
56
56
|
run: uv build
|
|
57
57
|
|
|
58
|
+
- name: Check the sdist stays lean
|
|
59
|
+
# The cap release.yml also enforces before it uploads. Catching a 1.26 MB tarball in the pull
|
|
60
|
+
# request that added the bytes is cheaper than catching it on the tag: 0.2.0 published one.
|
|
61
|
+
run: |
|
|
62
|
+
size=$(stat -c %s dist/*.tar.gz)
|
|
63
|
+
if [ "$size" -gt 1048576 ]; then
|
|
64
|
+
echo "::error::the sdist is $size bytes, over the 1 MiB cap; see the include list in pyproject.toml"
|
|
65
|
+
tar tzvf dist/*.tar.gz | sort -k3 -n -r | head -5
|
|
66
|
+
exit 1
|
|
67
|
+
fi
|
|
68
|
+
echo "sdist is $size bytes, under the 1 MiB cap"
|
|
69
|
+
|
|
58
70
|
no-extra:
|
|
59
71
|
name: base install, no yaml extra
|
|
60
72
|
runs-on: ubuntu-latest
|
|
@@ -1,17 +1,22 @@
|
|
|
1
1
|
name: Release
|
|
2
2
|
|
|
3
3
|
# Publishing is driven by a tag, so the artifact is always traceable to an exact commit.
|
|
4
|
-
# Pushing v0.2.0
|
|
5
|
-
#
|
|
6
|
-
# a
|
|
7
|
-
# published
|
|
4
|
+
# Pushing v0.2.0 verifies and builds the files, publishes them to TestPyPI, installs them back from
|
|
5
|
+
# TestPyPI by name, and only then uploads them to PyPI and attaches them to a GitHub Release. The
|
|
6
|
+
# scratch index is a stage rather than an optional rehearsal because it is the one place the
|
|
7
|
+
# published metadata is exercised the way a user exercises it — `pip install pyattacker==X.Y.Z` —
|
|
8
|
+
# and being scratch it costs nothing, while PyPI refuses a re-upload and a version, once published,
|
|
9
|
+
# can be deleted but never reused.
|
|
10
|
+
#
|
|
11
|
+
# `workflow_dispatch` stops before both uploads unless *Publish to TestPyPI* is checked, which is
|
|
12
|
+
# how you rehearse a release before spending a version number.
|
|
8
13
|
on:
|
|
9
14
|
push:
|
|
10
15
|
tags: ["v*"]
|
|
11
16
|
workflow_dispatch:
|
|
12
17
|
inputs:
|
|
13
18
|
publish_to_testpypi:
|
|
14
|
-
description: "
|
|
19
|
+
description: "Also upload to TestPyPI and install from it (otherwise the run only verifies and builds)"
|
|
15
20
|
type: boolean
|
|
16
21
|
default: false
|
|
17
22
|
|
|
@@ -115,6 +120,21 @@ jobs:
|
|
|
115
120
|
fi
|
|
116
121
|
echo "sdist contains no local tooling state"
|
|
117
122
|
|
|
123
|
+
# The size is a check of its own because 0.2.0 shipped 883 KB of logo PNG in a 1.26 MB sdist:
|
|
124
|
+
# `assets/` was in the include list, PNG does not compress, and nothing that unpacks an sdist
|
|
125
|
+
# needs branding. Source, tests and docs come to about 440 KB, so a 1 MiB cap leaves room for
|
|
126
|
+
# the project to grow while still failing on the next binary that does not belong in a source
|
|
127
|
+
# tree — which is cheaper to find here than in a release.
|
|
128
|
+
- name: Check the sdist stays lean
|
|
129
|
+
run: |
|
|
130
|
+
size=$(stat -c %s dist/*.tar.gz)
|
|
131
|
+
if [ "$size" -gt 1048576 ]; then
|
|
132
|
+
echo "::error::the sdist is $size bytes, over the 1 MiB cap; see the include list in pyproject.toml"
|
|
133
|
+
tar tzvf dist/*.tar.gz | sort -k3 -n -r | head -5
|
|
134
|
+
exit 1
|
|
135
|
+
fi
|
|
136
|
+
echo "sdist is $size bytes, under the 1 MiB cap"
|
|
137
|
+
|
|
118
138
|
- name: Check the sdist rebuilds and passes its own tests
|
|
119
139
|
run: |
|
|
120
140
|
mkdir -p /tmp/sdist && tar xzf dist/*.tar.gz -C /tmp/sdist --strip-components=1
|
|
@@ -184,10 +204,13 @@ jobs:
|
|
|
184
204
|
if-no-files-found: error
|
|
185
205
|
|
|
186
206
|
# Publishing lives in its own job so the OIDC token is scoped to it alone: no test or build step
|
|
187
|
-
# ever holds a credential that can upload a release.
|
|
207
|
+
# ever holds a credential that can upload a release. That is also why the index is verified by the
|
|
208
|
+
# job below rather than by a step in this one.
|
|
188
209
|
testpypi:
|
|
189
210
|
name: publish to TestPyPI
|
|
190
|
-
|
|
211
|
+
# On a tag this is a stage of the release. On a dispatch it runs only when the input asks for it,
|
|
212
|
+
# so that a dry run stays a dry run.
|
|
213
|
+
if: startsWith(github.ref, 'refs/tags/v') || inputs.publish_to_testpypi
|
|
191
214
|
needs: build
|
|
192
215
|
runs-on: ubuntu-latest
|
|
193
216
|
environment: testpypi
|
|
@@ -206,7 +229,10 @@ jobs:
|
|
|
206
229
|
# caching here can only ever warn that the cache will never get invalidated.
|
|
207
230
|
enable-cache: false
|
|
208
231
|
|
|
209
|
-
# --check-url lets a re-run skip files already uploaded instead of failing the job
|
|
232
|
+
# --check-url lets a re-run skip files already uploaded instead of failing the job, which also
|
|
233
|
+
# makes a re-run of a rehearsed release idempotent: identical bytes are skipped and still
|
|
234
|
+
# verified below. Different bytes for a version TestPyPI already holds fail here on purpose —
|
|
235
|
+
# publishing those to PyPI would ship files no index has ever resolved.
|
|
210
236
|
- name: Publish
|
|
211
237
|
run: |
|
|
212
238
|
uv publish \
|
|
@@ -215,10 +241,63 @@ jobs:
|
|
|
215
241
|
--check-url https://test.pypi.org/simple/ \
|
|
216
242
|
dist/*
|
|
217
243
|
|
|
244
|
+
# The check the rehearsal used to leave to a human, and the reason TestPyPI is a stage at all:
|
|
245
|
+
# the files are installed the way a user installs them, by name from an index, with no credential
|
|
246
|
+
# in reach. `--index-url` is what makes the install come from TestPyPI rather than PyPI, and the
|
|
247
|
+
# version being released cannot be on PyPI yet — the upload that would put it there needs this job
|
|
248
|
+
# to pass — so the `--version` assertion proves the TestPyPI files are the ones that got installed.
|
|
249
|
+
# One index is enough because the base package declares no dependencies — `tests/test_packaging.py`
|
|
250
|
+
# enforces that — while TestPyPI's own PyYAML is frozen at 3.11, so a release that grew a runtime
|
|
251
|
+
# dependency fails here rather than publishing metadata no index could resolve. The wider
|
|
252
|
+
# `--extra-index-url` form, for the `yaml` extra and for that case, is in `docs/releasing.md`.
|
|
253
|
+
testpypi_check:
|
|
254
|
+
name: install from TestPyPI
|
|
255
|
+
needs: [build, testpypi]
|
|
256
|
+
runs-on: ubuntu-latest
|
|
257
|
+
steps:
|
|
258
|
+
- name: Install uv
|
|
259
|
+
uses: astral-sh/setup-uv@bec219d24cd3e171d82865faccec33120bb574f4 # v10.1.0
|
|
260
|
+
with:
|
|
261
|
+
# No checkout in this job, so no lockfile to key a cache on — see the note in `testpypi`.
|
|
262
|
+
enable-cache: false
|
|
263
|
+
|
|
264
|
+
- name: Install the published files by name
|
|
265
|
+
env:
|
|
266
|
+
VERSION: ${{ needs.build.outputs.version }}
|
|
267
|
+
run: |
|
|
268
|
+
uv venv /tmp/index-check --python 3.11
|
|
269
|
+
# The simple index is a cached listing per CDN point of presence, so a file that was just
|
|
270
|
+
# uploaded can take a minute to show up in the one this job resolves through — measured at
|
|
271
|
+
# ~30-70s while this was written, on an upload uv itself had already resolved. A release
|
|
272
|
+
# waits that out rather than failing: ten attempts, 15 seconds apart.
|
|
273
|
+
for attempt in 1 2 3 4 5 6 7 8 9 10; do
|
|
274
|
+
if uv pip install --python /tmp/index-check/bin/python \
|
|
275
|
+
--index-url https://test.pypi.org/simple/ \
|
|
276
|
+
"pyattacker==$VERSION"; then
|
|
277
|
+
break
|
|
278
|
+
fi
|
|
279
|
+
if [ "$attempt" = 10 ]; then
|
|
280
|
+
echo "::error::pyattacker $VERSION is not installable from TestPyPI"
|
|
281
|
+
exit 1
|
|
282
|
+
fi
|
|
283
|
+
echo "attempt $attempt could not resolve pyattacker==$VERSION; retrying in 15s"
|
|
284
|
+
sleep 15
|
|
285
|
+
done
|
|
286
|
+
|
|
287
|
+
- name: The index copy is the version being released, and it runs
|
|
288
|
+
env:
|
|
289
|
+
VERSION: ${{ needs.build.outputs.version }}
|
|
290
|
+
run: |
|
|
291
|
+
/tmp/index-check/bin/pyattacker --version | grep -qx "pyattacker $VERSION"
|
|
292
|
+
/tmp/index-check/bin/pyattacker demo --pipelines 10 --store /tmp/index-check/demo.db
|
|
293
|
+
echo "TestPyPI serves pyattacker $VERSION and the installed CLI runs"
|
|
294
|
+
|
|
218
295
|
pypi:
|
|
219
296
|
name: publish to PyPI
|
|
220
297
|
if: startsWith(github.ref, 'refs/tags/v')
|
|
221
|
-
|
|
298
|
+
# Gated on the index check, not just on the build: the files are published for real only after
|
|
299
|
+
# TestPyPI has served them to a clean environment and the CLI there reported this version.
|
|
300
|
+
needs: [build, testpypi_check]
|
|
222
301
|
runs-on: ubuntu-latest
|
|
223
302
|
environment: pypi
|
|
224
303
|
permissions:
|
|
@@ -4,7 +4,216 @@ All notable changes to this project are documented here. The format follows
|
|
|
4
4
|
[Keep a Changelog](https://keepachangelog.com/en/1.1.0/), and this project adheres to
|
|
5
5
|
[Semantic Versioning](https://semver.org/spec/v2.0.0.html).
|
|
6
6
|
|
|
7
|
-
## [
|
|
7
|
+
## [0.3.0] — 2026-09-19
|
|
8
|
+
|
|
9
|
+
### Added
|
|
10
|
+
|
|
11
|
+
* **Simplified Chinese documentation, kept in sync by CI.** `README.zh-CN.md` and `docs/zh-CN/` mirror the
|
|
12
|
+
README, every document in `docs/` (tutorial, reference, design, CLI, benchmark, backward traversal,
|
|
13
|
+
releasing) and the two example READMEs, and each pair carries a language switcher. The translation is
|
|
14
|
+
structural: fenced code blocks are byte-identical to the English ones — so the tests that execute the
|
|
15
|
+
English tutorial, README and CLI examples cover the Chinese documents too — and
|
|
16
|
+
`tests/test_docs_i18n.py` fails when a code block, a heading level, a language switcher or a relative
|
|
17
|
+
link drifts out of sync. Prose quality stays a review question; the machine checks the parts that rot.
|
|
18
|
+
|
|
19
|
+
* **Advanced backward traversal (opt-in, experimental):** `Handoff.rewind(target, value)` uses
|
|
20
|
+
author-selected state; `Handoff.retry_all()` replays immutable bound seed bytes. Separate declarations
|
|
21
|
+
require finite control budgets. Visit-aware atomic stores retain task/attempt/artifact occurrences,
|
|
22
|
+
effective lineage and exact pending inputs across rewind and resume. Revisited RNG includes `ctx.visit`;
|
|
23
|
+
visit-0 IDs and forward-only digests stay compatible. Optional `HistoryArtifact` payloads provide
|
|
24
|
+
detached snapshots, restoration, explicit pruning and registered-subclass codec round trips. Exports
|
|
25
|
+
and `/pipelines` expose visit-aware lineage. Restarting is explicit: `RunConfig.fresh_restart` /
|
|
26
|
+
`--fresh-restart` discards a checkpoint or traversal from the bound seed, resets the control budget and
|
|
27
|
+
keeps the append-only history, and a store that has committed its first revisit records a `base` → `visits-v1`
|
|
28
|
+
feature level (`store.feature_level()`, `StoreFeatureUnsupported`) so lineage-unaware writers are refused
|
|
29
|
+
rather than silently mutating the wrong occurrence. See
|
|
30
|
+
[reference → advanced: backward traversal](docs/reference.md#advanced-backward-traversal-rewind-retry-all-visits).
|
|
31
|
+
|
|
32
|
+
* **Advanced feature: handoffs — a task can skip ahead, on the record (opt-in, experimental).** A task may
|
|
33
|
+
return `Handoff.to(target, value)` to continue at a declared later station, or `Handoff.end(value)` to
|
|
34
|
+
finish the pipeline immediately. Until now the alternatives were to run the remaining stations anyway, to
|
|
35
|
+
fold the branch into one task with `fanout` (losing per-step records), or to raise — which records the
|
|
36
|
+
pipeline as *failed*, which is a lie. A handoff says what actually happened: this row skipped stations 3–5 and
|
|
37
|
+
continued at station 6, or finished here. It is also a **durable checkpoint**: the source task lands
|
|
38
|
+
`handed_off`, its attempt keeps `outcome="handed_off"`, the entry artifact is referenced or stored, a
|
|
39
|
+
`handoffs` ledger row records from/to and why, and the cursor moves — all in **one atomic commit**, so a
|
|
40
|
+
killed process resumes *at the target* with the entry state and never re-runs the source task.
|
|
41
|
+
`Observable changes:` a new append-only `handoffs` table (`store.handoffs(...)`, nested in the `pipelines`
|
|
42
|
+
export row, counted by `stats()["handoffs_total"]`, shown by `report`/`watch`/`/pipelines`); the attempt
|
|
43
|
+
outcome `handed_off`; handoff payload artifacts at `seq >= n_tasks`, so a payload can never overwrite a
|
|
44
|
+
task slot; the `pipeline.handoff` event; `Handoff` and `HandoffRecord` in the public API; and
|
|
45
|
+
`control={"edges": {...}}` on `pipeline(...)` plus the declarative `pipeline.control`. Edges are declared
|
|
46
|
+
rather than derived, so an undeclared or backward target is a fatal error instead of a silent jump, and
|
|
47
|
+
`fanout` rejects a directive returned by one of its branches. **A pipeline without a `control` block is
|
|
48
|
+
untouched**: no new rows, no counter changes, and a byte-identical `spec_digest` (pinned by a literal test),
|
|
49
|
+
so no stored checkpoint, pipeline id or shard assignment is invalidated. Handoffs need a store that can
|
|
50
|
+
commit them atomically; both built-in backends can, and a store that cannot is refused up front with a
|
|
51
|
+
`ConfigError` rather than silently writing a non-durable jump. Marked *experimental until 1.0*;
|
|
52
|
+
the forward declaration excludes backward targets; joins and cross-pipeline jumps are not included. See
|
|
53
|
+
[`docs/design.md` §4.8](docs/design.md#48-advanced-handoffs--declared-forward-jumps-opt-in-experimental),
|
|
54
|
+
the [API reference](docs/reference.md#advanced-handoffs-opt-in) and
|
|
55
|
+
[tutorial step 15](docs/tutorial.md#step-15--advanced-skipping-stations-handoffs).
|
|
56
|
+
|
|
57
|
+
### Fixed
|
|
58
|
+
|
|
59
|
+
* **Documentation drift that no test could see** (issue #60), and a lightweight check for the class of it.
|
|
60
|
+
The design document still opened with "Version: 0.1.0 (M0–M4 complete, M5 … in progress)" while
|
|
61
|
+
`pyproject.toml`, the changelog and the README had been on 0.2.0 since September, and M5 — forward
|
|
62
|
+
handoffs and backward traversal — had shipped; §8 still claimed "v1 only supports asyncio tasks" while the
|
|
63
|
+
first task of the tutorial is a plain `def` and the reference documents sync-task behaviour; the reference
|
|
64
|
+
still called the tutorial "fourteen runnable steps" after three more were added; and the monitoring docs
|
|
65
|
+
promised that `serve` "serves your artifact payloads" although no route returns one — what is actually
|
|
66
|
+
exposed is event `data` and stored `error_message` fields, which is a warning worth keeping, so the
|
|
67
|
+
sentence now names those instead of the wrong thing. All of it is mirrored in the Chinese documents.
|
|
68
|
+
`tests/test_docs_facts.py` now asserts the three claims that have a source of truth — the version header
|
|
69
|
+
against `pyproject.toml`, every "N runnable steps" claim against the number of `# tutorial/` blocks, and
|
|
70
|
+
every route the monitoring endpoint table advertises against the routes `StatsServer` answers — so the next
|
|
71
|
+
release cannot leave the header behind.
|
|
72
|
+
|
|
73
|
+
* **`merge_reports` no longer inflates its counters when it folds a duplicate** (issue #59). Rows were
|
|
74
|
+
always de-duplicated by `pipeline_id`, but `attempts_total` and `handoffs_total` were summed over the
|
|
75
|
+
sources — before de-duplication — so `pyattacker report a.db a.db`, or one pipeline living in two shards
|
|
76
|
+
after a shard-count change, doubled them while `pipelines.total` stayed put, contradicting the documented
|
|
77
|
+
promise that merging is idempotent. Both are now recomputed from the surviving rows (attempts from the
|
|
78
|
+
row's own `attempts_total`, handoffs from the nested ledger it carries), which is also what the module
|
|
79
|
+
docstring, `docs/design.md` §9 and `docs/reference.md` already claimed. The event log genuinely cannot be
|
|
80
|
+
re-derived — an event is not part of a pipeline row — so instead of pretending, it is renamed
|
|
81
|
+
`source_events_total` and documented as a raw per-source total; `summary()` prints it as
|
|
82
|
+
`source_events=`. Regression tests cover the same store passed twice (as two paths and as two objects),
|
|
83
|
+
one pipeline present in two shards, and the run-filtered handoff count; `tutorial.md` step 12 now shows
|
|
84
|
+
the identical `attempts=` on both sides of a duplicated merge (whose printed transcript had also drifted
|
|
85
|
+
from what the program actually prints). Counting from rows strengthens what a custom store must export,
|
|
86
|
+
so the two fields it reads are now explicit: `attempts_total` is required and a row without it raises a
|
|
87
|
+
`ConfigError` naming the row and source instead of quietly counting `0`, while a row that does not nest
|
|
88
|
+
the optional `handoffs` ledger is read through the store's own `handoffs()` capability when it has one
|
|
89
|
+
(a store with neither has no jumps, which is a true `0`).
|
|
90
|
+
|
|
91
|
+
* Backward-traversal follow-up: an explicit fresh start now resets a backward pipeline's control budget
|
|
92
|
+
while preserving durable visit counters and audit history (the counters were never reset — reusing visit
|
|
93
|
+
IDs would collide with historical occurrences), so a pipeline that failed *because* it spent its budget
|
|
94
|
+
can be restarted instead of replaying the same fatal error forever; the forward capability check is
|
|
95
|
+
applied to backward-enabled pipelines too, and `supports_visits` requires the v1
|
|
96
|
+
`commit_handoff`/`reset_pipeline` ledger capability the transition actually uses;
|
|
97
|
+
`MemoryStore.stats()["attempts_total"]` counts attempt rows for the run like `SqliteStore`, instead of
|
|
98
|
+
double-counting a resumed visit's consumed attempts.
|
|
99
|
+
|
|
100
|
+
* Backward-traversal review follow-up. `retry_succeeded=True` no longer restarts an unfinished backward
|
|
101
|
+
pipeline: it is an eligibility switch ("also admit succeeded pipelines"), and an unfinished pipeline still
|
|
102
|
+
owns a recoverable traversal, so replaying it from the seed would repeat external side effects the durable
|
|
103
|
+
checkpoint was about to continue. The operator escape hatch is the new `RunConfig.fresh_restart` /
|
|
104
|
+
`--fresh-restart` / `run.fresh_restart`: it discards the checkpoint or traversal, restarts from the bound
|
|
105
|
+
seed, resets the control budget and invalidates the previous ledger: append-only history (attempts,
|
|
106
|
+
events, handoffs, retained payloads) survives, and a backward pipeline additionally keeps its
|
|
107
|
+
visit-qualified occurrences and counters. It is also the documented recovery for a pipeline whose
|
|
108
|
+
traversal is gone, and it settles every task row it leaves in flight as `interrupted` instead of leaving
|
|
109
|
+
it `running` forever. Reopening a
|
|
110
|
+
`running` backward pipeline is now explicit: without `resume=True` the row is skipped
|
|
111
|
+
(`pipeline.skipped`, `reason="owned_by_another_run"`) instead of silently taking over a traversal another
|
|
112
|
+
run may still own; `resume=True` reclaims it and continues the exact durable visit. A fresh run no longer
|
|
113
|
+
reports `pipeline.checkpoint_missing` for its own seed (a `journal="summary"` store drops that payload by
|
|
114
|
+
design, which used to make every first execution look like a failed checkpoint). `control.retry_all` now
|
|
115
|
+
rejects aliases that resolve to the same source, matching `control.rewind`. A store that has committed a
|
|
116
|
+
revisit records it durably (`store_meta` feature level `visits-v1`, `store.feature_level()`), refuses to
|
|
117
|
+
open an unknown (newer) level with `StoreFeatureUnsupported`, and arms a writer guard so a lineage-unaware
|
|
118
|
+
writer fails loudly instead of mutating the wrong occurrence.
|
|
119
|
+
|
|
120
|
+
* Review follow-up: freeze resolved control topology; atomically reset current task/chain artifact state
|
|
121
|
+
on control-enabled seed replays while retaining history; expose handoff identity/watermark and active
|
|
122
|
+
versus historical API counts; treat unencodable handoff payloads as fatal; correct return-only
|
|
123
|
+
annotation documentation. Handoff-capable custom stores must also provide `reset_pipeline(record)`.
|
|
124
|
+
|
|
125
|
+
* Handoff follow-up: isolate historical ledger rows on whole-pipeline restarts with a durable
|
|
126
|
+
`handoff_floor` watermark (including old-store migration and read-only compatibility); roll back failed
|
|
127
|
+
SQLite handoff commits; select a single final artifact after reruns; reject duplicate source aliases;
|
|
128
|
+
preserve numeric source disambiguation in JSON/TOML/YAML and effective config; restrict the annotation
|
|
129
|
+
escape to return unions containing `Handoff`.
|
|
130
|
+
|
|
131
|
+
* **A worker that dies outside its own handlers no longer hangs the run.** Run completion was tracked purely
|
|
132
|
+
by pipeline counters (`pipelines_done` against `pipelines_admitted`), so a `BaseException` that is not
|
|
133
|
+
`CancelledError` — a custom subclass raised by a store or backend hook, for example — escaped the worker's
|
|
134
|
+
handler chain, ended the task, and left the run waiting forever on a condition the dead worker could no
|
|
135
|
+
longer satisfy: no error, no exit code, no terminal row and no event (only asyncio's "Task exception was
|
|
136
|
+
never retrieved" on stderr). Worker lifetime is now supervised separately from pipeline accounting, by a
|
|
137
|
+
done-callback on every worker task that also sees a death *after* the handler chain (queue or worker
|
|
138
|
+
housekeeping). The run is hard-stopped (`stop_reason == "worker_crashed"`) without manufacturing
|
|
139
|
+
`pipelines_done`; the pipeline the dead worker was holding is recorded `failed` — or `interrupted` when the
|
|
140
|
+
run was already winding down — with the escaping exception on the row, or created if the crash happened
|
|
141
|
+
before its row existed, while a pipeline that was already terminal keeps the state it earned; admission and
|
|
142
|
+
the shutdown sentinels hand items over through an abort-aware put, so a producer parked on the full queue of
|
|
143
|
+
a dead worker is released instead of hanging one step earlier; `runner.worker_crashed` records the pipeline,
|
|
144
|
+
the exception and its traceback; and `run_async` raises the new `WorkerCrashed` (original exception as
|
|
145
|
+
`__cause__`) after closing the run record as `interrupted`, with the CLI exiting 2 on a named error. If the
|
|
146
|
+
store cannot record the crash either, the existing `StoreUnavailable` fatal path is taken instead of
|
|
147
|
+
retrying a broken store. `KeyboardInterrupt`/`SystemExit` still tear the loop down, so a narrow guard inside
|
|
148
|
+
the worker records the row and the event before they continue on their way. Cancellation and the ordinary
|
|
149
|
+
internal-`Exception` recovery path are unchanged.
|
|
150
|
+
* **A pipeline whose cursor already reached the end can be resumed again.** The success path advanced the
|
|
151
|
+
checkpoint cursor and wrote the terminal state as two separate store calls, so a store failure in the
|
|
152
|
+
final write (or a kill between the two) could leave a `failed`/`interrupted` row with
|
|
153
|
+
`n_tasks_done == n_tasks_total`. Every later run then raised `IndexError: tuple index out of range` inside
|
|
154
|
+
`_open_pipeline`, surfaced only as `runner.internal_error`, while overwriting the row's original failure
|
|
155
|
+
with the framework's own crash. Such a row is now settled as completed: the last artifact is verified the
|
|
156
|
+
same way an ordinary resumed checkpoint is (present, payload kept, decodable), marked final, and the
|
|
157
|
+
pipeline is recorded `succeeded` without re-running a task, with `pipeline.terminal_repaired` carrying the
|
|
158
|
+
previous state, error and run. A cursor *past* the end is not a state the Runner can create: it is recorded
|
|
159
|
+
as `CorruptCheckpoint` and reported as `pipeline.corrupt_cursor` instead of being promoted to success, and
|
|
160
|
+
the stored value is left as the evidence; an artifact that is missing, payload-less or undecodable falls
|
|
161
|
+
back to the documented restart-from-zero rule (`pipeline.checkpoint_missing` /
|
|
162
|
+
`pipeline.checkpoint_unusable`). The repair never destroys the failure it is repairing before the outcome is
|
|
163
|
+
durable — `mark_final` runs first, the terminal transition (state, cursor, owning run, cleared failure
|
|
164
|
+
fields) is a single write where the store offers the optional `settle_pipeline` capability, and a repair
|
|
165
|
+
that fails in **either** step — marking the artifact final or settling the row — leaves the row untouched
|
|
166
|
+
(original failure and owning run included) and records `pipeline.terminal_repair_failed` with the phase,
|
|
167
|
+
so the next attempt still reports the original cause instead of the repair's own error. `mark_final` is
|
|
168
|
+
now contractually idempotent and is skipped when the artifact is already final. Because such a row keeps
|
|
169
|
+
its original owner, the failed repair is counted in the new `RunReport.repair_failures` (surfaced by
|
|
170
|
+
`summary()`, `to_dict()` and the CLI exit code) instead of silently leaving the run looking successful;
|
|
171
|
+
and on the two-write fallback for stores without `settle_pipeline`, a failed metadata cleanup after a
|
|
172
|
+
durable `succeeded` is reported as `pipeline.terminal_cleanup_failed` rather than as a failed repair. The final task now also marks its artifact final **before**
|
|
173
|
+
committing the terminal state and advances the cursor in that same call, so a crash in between re-runs the
|
|
174
|
+
final task (the documented at-least-once boundary) rather than leaving a `succeeded` pipeline whose final
|
|
175
|
+
artifact was never marked — a state no later run could repair, because a succeeded pipeline is skipped
|
|
176
|
+
forever.
|
|
177
|
+
|
|
178
|
+
### Changed
|
|
179
|
+
|
|
180
|
+
* **`MergedReport.events_total` is renamed `source_events_total`** (issue #59). The old name sat in the same
|
|
181
|
+
report as the de-duplicated counters without saying that it was the one number which was not de-duplicated;
|
|
182
|
+
the new name states the scope, `stats()` carries it under the same key, and `summary()` labels it
|
|
183
|
+
`source_events=`. Nothing else on `MergedReport` changed name. `MergedReport.events_total` survives as a
|
|
184
|
+
deprecated read-only alias (removed at 1.0) so an attribute read keeps working, but it is deliberately not a
|
|
185
|
+
second key in `stats()`: the point is that the JSON a report emits names one scope per number. Note that
|
|
186
|
+
`store.stats()["events_total"]` is untouched and is not the asymmetry it looks like — for a single store it
|
|
187
|
+
is the same measurement as the merged report's `source_events_total`, which the new test pins.
|
|
188
|
+
|
|
189
|
+
* **The backward-traversal guide is merged into the tutorial and the reference.** `docs/backward.md` was a
|
|
190
|
+
seventh document that a reader had to find before they could use the feature; its usage now lives where
|
|
191
|
+
the rest of the API does. [Tutorial step 16](docs/tutorial.md#step-16--advanced-regenerating-with-rewind-and-retry-all)
|
|
192
|
+
keeps the rewind/retry-all walkthrough and a new step 17 builds an optional `HistoryArtifact` payload,
|
|
193
|
+
while [reference → advanced: backward traversal](docs/reference.md#advanced-backward-traversal-rewind-retry-all-visits)
|
|
194
|
+
gains the `Handoff.rewind`/`Handoff.retry_all` signatures, the backward `control` keys, the visit and
|
|
195
|
+
occurrence model, the budget and recovery rules, and the inspection surface; `HistoryArtifact` is
|
|
196
|
+
documented next to the codecs it belongs to, and the store capability and compatibility/backup rules moved
|
|
197
|
+
into the Stores section. Both steps and the reference examples are marked **advanced**, stay opt-in and
|
|
198
|
+
experimental until 1.0, and are executed by the test suite (`# reference/<name>.py` joins the
|
|
199
|
+
executable-document markers). The Chinese tree mirrors all of it. No API changed.
|
|
200
|
+
|
|
201
|
+
* **The sdist no longer ships the branding images.** `assets/` was in the sdist include list, and PNG
|
|
202
|
+
barely compresses, so the two logo files were 883 KB of the 1.26 MB `pyattacker-0.2.0.tar.gz` that PyPI
|
|
203
|
+
serves — two thirds of the download for an archive whose purpose is to be rebuilt and tested. Nothing
|
|
204
|
+
that unpacks an sdist needs them, and the README loads its banner over an absolute
|
|
205
|
+
`raw.githubusercontent.com` URL, so the PyPI description is unaffected; the tarball is back to 439 KB.
|
|
206
|
+
Both `ci.yml` and `release.yml` now fail the build if the sdist grows past 1 MiB, and
|
|
207
|
+
`tests/test_packaging.py` fails if a README image ever points into the checkout again.
|
|
208
|
+
* **A release goes through TestPyPI before it reaches PyPI.** The scratch-index upload was a rehearsal a
|
|
209
|
+
maintainer had to opt into, and the `v0.2.0` tag skipped it — so the files PyPI received were the first copy
|
|
210
|
+
of that version any index had ever served. Every tag now publishes to TestPyPI, installs those files back by
|
|
211
|
+
name into a clean environment, asserts that the installed CLI reports the version being released, and only
|
|
212
|
+
then uploads to PyPI. The check is a job of its own with no OIDC token, so the job that holds a credential
|
|
213
|
+
that can publish does nothing else, and `pypi` now depends on it. A dispatched run still rehearses when
|
|
214
|
+
*Publish to TestPyPI* is checked and stops before both uploads otherwise. `docs/releasing.md` records the
|
|
215
|
+
consequence this makes load-bearing: a filename an index has seen can never be uploaded again, so a commit
|
|
216
|
+
after a rehearsal means a new version number, not a re-run.
|
|
8
217
|
|
|
9
218
|
## [0.2.0] — 2026-09-17
|
|
10
219
|
|
|
@@ -353,7 +562,7 @@ hard limit; asyncio tasks only (wrap blocking code with `asyncio.to_thread`); on
|
|
|
353
562
|
balance is statistical; a merged report is a union, not a sum; a pipeline is a linear chain; the HTTP endpoint
|
|
354
563
|
is unauthenticated.
|
|
355
564
|
|
|
356
|
-
[
|
|
565
|
+
[0.3.0]: https://github.com/Hazer-BJTU/pyattacker/releases/tag/v0.3.0
|
|
357
566
|
[0.2.0]: https://github.com/Hazer-BJTU/pyattacker/releases/tag/v0.2.0
|
|
358
567
|
[0.1.1]: https://github.com/Hazer-BJTU/pyattacker/releases/tag/v0.1.1
|
|
359
568
|
[0.1.0]: https://github.com/Hazer-BJTU/pyattacker/releases/tag/v0.1.0
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
Metadata-Version: 2.5
|
|
2
2
|
Name: pyattacker
|
|
3
|
-
Version: 0.
|
|
3
|
+
Version: 0.3.0
|
|
4
4
|
Summary: Artifact-centric, resumable async task orchestration: pipeline / task / resource pool.
|
|
5
5
|
Project-URL: Homepage, https://github.com/Hazer-BJTU/pyattacker
|
|
6
6
|
Project-URL: Repository, https://github.com/Hazer-BJTU/pyattacker
|
|
@@ -32,6 +32,8 @@ Description-Content-Type: text/markdown
|
|
|
32
32
|
[](https://pypi.org/project/pyattacker/)
|
|
33
33
|
[](https://github.com/Hazer-BJTU/pyattacker/blob/main/LICENSE)
|
|
34
34
|
|
|
35
|
+
**English** | [简体中文](README.zh-CN.md)
|
|
36
|
+
|
|
35
37
|
> Run tens of thousands of independent tasks to completion — resumably, observably, and without
|
|
36
38
|
> reimplementing endpoint pools, retries and "which rows already ran" for the fifth time.
|
|
37
39
|
|
|
@@ -91,7 +93,7 @@ Five concepts, and that is the whole vocabulary:
|
|
|
91
93
|
| Concept | Meaning | In one line |
|
|
92
94
|
|---|---|---|
|
|
93
95
|
| **artifact** | the persisted state of a task | content-addressed, **persisted as soon as it is produced** → checkpoint granularity = task |
|
|
94
|
-
| **task** | the smallest unit of scheduling | a unary `(artifact) -> artifact` function, sync or async |
|
|
96
|
+
| **task** | the smallest unit of scheduling | a unary `(artifact) -> artifact` function, sync or async — or `-> artifact \| Handoff`, to skip ahead ([advanced](#advanced-handoffs-opt-in)) |
|
|
95
97
|
| **pipeline** | the unit of completion | `fetch \| ask \| judge \| metrics` chained linearly, semantically independent of each other |
|
|
96
98
|
| **resource** | a leasable external capability | one endpoint / one key; once pooled, it can be published and subscribed to concurrency-safely |
|
|
97
99
|
| **algorithm** | the policy for acquiring resources | `wait`, `backoff`, `least_busy`, `failover`, `sticky`, `quota_aware`, `immediate` — orthogonal to "retry on failure" |
|
|
@@ -200,6 +202,8 @@ runner.run(template.map(rows), resume=True) # or pyattacker resume -c config.y
|
|
|
200
202
|
* Failed pipelines → continue from **the first task that produced no artifact**: **if task C died, only task C
|
|
201
203
|
reruns when task B's checkpoint is durable**;
|
|
202
204
|
* The seed artifact is persisted too → recovery **does not depend on the original dataset file**;
|
|
205
|
+
* A pipeline that **handed off** ([advanced](#advanced-handoffs-opt-in)) resumes at the station it jumped to,
|
|
206
|
+
with the entry state the ledger recorded — the task that handed off is not re-run;
|
|
203
207
|
* Changed a task's source code (`spec_digest` includes source digests) → treated as a new pipeline, so old
|
|
204
208
|
results are not incorrectly reused. Factory parameters, fanout children, retry/algorithm policies and
|
|
205
209
|
explicit task `config`/`version` are included too. Explicit keys reject changed definitions or inputs.
|
|
@@ -209,6 +213,106 @@ rerun and old explicit keys conflict. Finish old runs with the old package, then
|
|
|
209
213
|
[resume identity and idempotency reference](https://github.com/Hazer-BJTU/pyattacker/blob/main/docs/reference.md#resume-identity)
|
|
210
214
|
for migration guidance, dynamic functions and external configuration.
|
|
211
215
|
|
|
216
|
+
## Advanced: Handoffs (Opt-In)
|
|
217
|
+
|
|
218
|
+
**Advanced tier: opt-in, changes the execution model, not needed for ordinary pipelines, experimental until
|
|
219
|
+
1.0.** A task can decide that the rest of the chain no longer needs to run, and *say so* instead of inventing a
|
|
220
|
+
failure or hiding the branch inside one step. It **returns** a directive — `Handoff.to(target, value)` to
|
|
221
|
+
continue at a declared later station, `Handoff.end(value)` to finish the pipeline right there:
|
|
222
|
+
|
|
223
|
+
```python
|
|
224
|
+
# example/readme_handoff.py
|
|
225
|
+
"""A gate that skips the stations it does not need, and records why."""
|
|
226
|
+
|
|
227
|
+
from pyattacker import Handoff, Runner, pipeline, task
|
|
228
|
+
|
|
229
|
+
|
|
230
|
+
@task("prepare")
|
|
231
|
+
def prepare(seed: dict) -> dict:
|
|
232
|
+
return {"q": seed["q"], "confidence": seed["confidence"]}
|
|
233
|
+
|
|
234
|
+
|
|
235
|
+
@task("ask")
|
|
236
|
+
async def ask(row: dict, ctx) -> dict:
|
|
237
|
+
await ctx.clock.sleep(0.001) # <- your HTTP call
|
|
238
|
+
return {**row, "answer": f"answer-for:{row['q']}"}
|
|
239
|
+
|
|
240
|
+
|
|
241
|
+
@task("judge")
|
|
242
|
+
def judge(row: dict) -> Handoff | dict:
|
|
243
|
+
if row["confidence"] >= 0.9:
|
|
244
|
+
return Handoff.end({**row, "verdict": "confident"}, reason="already good enough")
|
|
245
|
+
if row["confidence"] >= 0.5:
|
|
246
|
+
return Handoff.to("report", {**row, "verdict": "ok"}, reason="metrics not needed")
|
|
247
|
+
return {**row, "verdict": "needs-metrics"}
|
|
248
|
+
|
|
249
|
+
|
|
250
|
+
@task("metrics")
|
|
251
|
+
def metrics(row: dict) -> dict:
|
|
252
|
+
return {**row, "score": round(row["confidence"] * 10, 1)}
|
|
253
|
+
|
|
254
|
+
|
|
255
|
+
@task("report")
|
|
256
|
+
def report(row: dict) -> dict:
|
|
257
|
+
return {**row, "reported": True}
|
|
258
|
+
|
|
259
|
+
|
|
260
|
+
# Edges are declared, never derived: judge may continue at report, or end the pipeline.
|
|
261
|
+
template = pipeline("qa", prepare | ask | judge | metrics | report,
|
|
262
|
+
control={"edges": {"judge": ["report", "end"]}})
|
|
263
|
+
|
|
264
|
+
with Runner(store=":memory:", concurrency=4) as runner:
|
|
265
|
+
report_obj = runner.run(template.map([
|
|
266
|
+
{"q": "2+2", "confidence": 0.95}, # judge ends the pipeline here
|
|
267
|
+
{"q": "17*23", "confidence": 0.6}, # judge skips metrics, continues at report
|
|
268
|
+
{"q": "prove sqrt(2) is irrational", "confidence": 0.2}, # no handoff: the whole chain runs
|
|
269
|
+
]))
|
|
270
|
+
print(report_obj.summary())
|
|
271
|
+
store = runner.store
|
|
272
|
+
for record in store.pipelines():
|
|
273
|
+
ran = [t.name for t in store.tasks(record.pipeline_id)]
|
|
274
|
+
hops = store.handoffs(pipeline_id=record.pipeline_id)
|
|
275
|
+
print(f"{record.state} ran={ran} skipped={len(template.tasks) - len(ran)} handoffs={len(hops)}")
|
|
276
|
+
```
|
|
277
|
+
|
|
278
|
+
* **Declared forward edges.** A `Handoff` along an undeclared edge, or from a pipeline with no
|
|
279
|
+
`control` block, is a fatal configuration error — never retried, never a silent jump. Destinations must be
|
|
280
|
+
strictly later than their source; `end` from the last task is refused because it would do nothing.
|
|
281
|
+
* **A handoff is a disposition, not a failure.** It is a return value, so the retry policy never sees it and
|
|
282
|
+
a task-side `except Exception:` cannot swallow it; leases are released exactly as on success, and a
|
|
283
|
+
cancelled or timed-out attempt never reaches the return.
|
|
284
|
+
* **It is a durable checkpoint.** The jump is committed atomically (source task, attempt, entry artifact,
|
|
285
|
+
ledger row and cursor together), so a killed process resumes **at the target** with the recorded entry
|
|
286
|
+
state and does not re-run the source task. On a control-enabled pipeline `n_tasks_done` is a *position*,
|
|
287
|
+
not a progress count: skipped stations have no task rows. `report`/`watch` count commits in the selected
|
|
288
|
+
scope (one run when filtered, all history otherwise); `/pipelines.handoffs` counts active execution records and `handoffs_historical` counts all ledger
|
|
289
|
+
records. A resume can use an earlier run's active handoff while recording no new handoffs. Pipeline
|
|
290
|
+
exports include ledger identity and the watermark to distinguish the scopes.
|
|
291
|
+
* **Opt-in and inert.** Without the `control` block nothing changes — not one row, not one counter, and not a
|
|
292
|
+
byte of `spec_digest`.
|
|
293
|
+
* **Backward traversal is separately declared.** `Handoff.rewind(target, value)` sends author-selected
|
|
294
|
+
state to an earlier task; `Handoff.retry_all()` restarts from the original bound seed. Declare
|
|
295
|
+
`control.rewind` / `control.retry_all` and a finite `control.max_handoffs`. Optional `HistoryArtifact`
|
|
296
|
+
payloads provide explicit snapshots and restoration; ordinary dictionaries remain author-controlled.
|
|
297
|
+
Visits and exact artifact occurrences retain history and make recovery safe. The APIs, budgets and
|
|
298
|
+
recovery boundaries are in
|
|
299
|
+
[reference → advanced: backward traversal](https://github.com/Hazer-BJTU/pyattacker/blob/main/docs/reference.md#advanced-backward-traversal-rewind-retry-all-visits),
|
|
300
|
+
with working programs in
|
|
301
|
+
[tutorial steps 16–17](https://github.com/Hazer-BJTU/pyattacker/blob/main/docs/tutorial.md#step-16--advanced-regenerating-with-rewind-and-retry-all).
|
|
302
|
+
|
|
303
|
+
The API is one class ([`Handoff`](https://github.com/Hazer-BJTU/pyattacker/blob/main/docs/reference.md#advanced-handoffs-opt-in)),
|
|
304
|
+
one declaration — `control={"edges": {...}}` for forward jumps, `control.rewind` / `control.retry_all` /
|
|
305
|
+
`control.max_handoffs` for backward traversal — and one optional store capability; a custom store that cannot
|
|
306
|
+
commit a handoff atomically is refused up front instead of writing a jump that would not survive a crash.
|
|
307
|
+
Walkthrough: [tutorial step 15](https://github.com/Hazer-BJTU/pyattacker/blob/main/docs/tutorial.md#step-15--advanced-skipping-stations-handoffs).
|
|
308
|
+
Model and rules: [design §4.8](https://github.com/Hazer-BJTU/pyattacker/blob/main/docs/design.md#48-advanced-handoffs--declared-forward-jumps-opt-in-experimental).
|
|
309
|
+
|
|
310
|
+
When a pipeline restarts from the seed, previous task rows and chain artifacts are reset atomically with the cursor/watermark;
|
|
311
|
+
previous handoffs stay in its history but no longer act as
|
|
312
|
+
checkpoints. A later resume follows only the current execution's handoffs, and completion selects one
|
|
313
|
+
final artifact. Custom stores supporting handoffs must persist the durable `handoff_floor` watermark
|
|
314
|
+
alongside the cursor and implement atomic `reset_pipeline(record)`; see the [recovery contract](docs/reference.md#tables-and-readers).
|
|
315
|
+
|
|
212
316
|
## Running It Across Processes (Sharding)
|
|
213
317
|
|
|
214
318
|
SQLite takes one writer and the kernel is a single event loop, so scale means **processes with their own
|
|
@@ -230,8 +334,8 @@ uv run pyattacker export runs/qa.shard*of4.db runs/all.jsonl
|
|
|
230
334
|
uv run pyattacker export runs/qa.shard*of4.db runs/tasks.csv --rows tasks --format csv
|
|
231
335
|
```
|
|
232
336
|
|
|
233
|
-
`--rows` picks the shape: `pipelines` (nested, default
|
|
234
|
-
`events`, `artifacts`. `--format` picks `jsonl`, `json` or `csv`.
|
|
337
|
+
`--rows` picks the shape: `pipelines` (nested, default — its tasks, artifacts and handoffs included), `tasks`,
|
|
338
|
+
`attempts` (including each retry `decision`), `events`, `artifacts`. `--format` picks `jsonl`, `json` or `csv`.
|
|
235
339
|
|
|
236
340
|
## Declarative (Simple Tasks)
|
|
237
341
|
|
|
@@ -281,9 +385,10 @@ Every flag of every subcommand: [`docs/cli.md`](https://github.com/Hazer-BJTU/py
|
|
|
281
385
|
|
|
282
386
|
## Records and Monitoring
|
|
283
387
|
|
|
284
|
-
The complete story of one pipeline =
|
|
388
|
+
The complete story of one pipeline = six tables queried by `pipeline_id`: state and checkpoint, the final
|
|
285
389
|
state of each task, the full attempt history (including **every retry decision**
|
|
286
|
-
`{retry, reason, delay_s, error_class}`), intermediate and final artifacts,
|
|
390
|
+
`{retry, reason, delay_s, error_class}`), intermediate and final artifacts, the control-flow ledger
|
|
391
|
+
(`handoffs`, empty for an ordinary pipeline), and the structured event stream.
|
|
287
392
|
`pyattacker report/watch` consumes these facts directly.
|
|
288
393
|
|
|
289
394
|
## Extending It
|
|
@@ -326,8 +431,9 @@ uv run pyattacker serve runs/qa.db # http://127.0.0.1:8787
|
|
|
326
431
|
```
|
|
327
432
|
|
|
328
433
|
It opens a fresh read-only connection per request, so it runs happily beside a live run. It has **no
|
|
329
|
-
authentication** and binds to loopback:
|
|
330
|
-
|
|
434
|
+
authentication** and binds to loopback: there is no artifact route, but it does expose what your run recorded —
|
|
435
|
+
event `data` and stored error messages — so treat it as a debug view over your own data and do not put it on a
|
|
436
|
+
public interface without your own proxy in front.
|
|
331
437
|
|
|
332
438
|
**Branching inside a step** — `fanout(a, b)` runs several tasks on the same input concurrently and
|
|
333
439
|
returns `{task_name: value}`. Retry granularity becomes the group, which is the honest price of not
|
|
@@ -368,13 +474,14 @@ what they do not say.
|
|
|
368
474
|
|
|
369
475
|
| Document | What is in it |
|
|
370
476
|
|---|---|
|
|
371
|
-
| [`docs/tutorial.md`](https://github.com/Hazer-BJTU/pyattacker/blob/main/docs/tutorial.md) |
|
|
372
|
-
| [`docs/reference.md`](https://github.com/Hazer-BJTU/pyattacker/blob/main/docs/reference.md) | every public class and function: signatures, parameters, examples |
|
|
477
|
+
| [`docs/tutorial.md`](https://github.com/Hazer-BJTU/pyattacker/blob/main/docs/tutorial.md) | seventeen runnable steps, from "one task" to sharded evaluation and the advanced handoff/backward tiers — rewinds, retry-all and payload history included; each one executed by the test suite |
|
|
478
|
+
| [`docs/reference.md`](https://github.com/Hazer-BJTU/pyattacker/blob/main/docs/reference.md) | every public class and function: signatures, parameters, examples — including the advanced backward-traversal tier and `HistoryArtifact` |
|
|
373
479
|
| [`docs/cli.md`](https://github.com/Hazer-BJTU/pyattacker/blob/main/docs/cli.md) | every subcommand, every flag, exit codes, config reference |
|
|
374
480
|
| [`docs/design.md`](https://github.com/Hazer-BJTU/pyattacker/blob/main/docs/design.md) | conceptual model, the six invariants, the lease contract, data model, tradeoffs |
|
|
375
481
|
| [`docs/benchmark.md`](https://github.com/Hazer-BJTU/pyattacker/blob/main/docs/benchmark.md) | the algorithm benchmark: what the scenarios assume, what the metrics mean, how to read the table |
|
|
376
482
|
| [`CHANGELOG.md`](https://github.com/Hazer-BJTU/pyattacker/blob/main/CHANGELOG.md) | what changed, release by release |
|
|
377
483
|
| [`docs/releasing.md`](https://github.com/Hazer-BJTU/pyattacker/blob/main/docs/releasing.md) | for maintainers: how a release is cut and published |
|
|
484
|
+
| [`docs/zh-CN/`](https://github.com/Hazer-BJTU/pyattacker/tree/main/docs/zh-CN) | 简体中文: the same documents, with code blocks that `tests/test_docs_i18n.py` keeps byte-identical to these |
|
|
378
485
|
|
|
379
486
|
## Examples
|
|
380
487
|
|
|
@@ -400,7 +507,10 @@ These are design decisions, not missing features:
|
|
|
400
507
|
* **Network requests** — you write the openai/anthropic protocols yourself. The kernel never opens a socket.
|
|
401
508
|
* **Semantic reduction** — accuracy, pass@k, F1 and any cross-pipeline aggregation. Export the artifacts and
|
|
402
509
|
compute it outside, or write a sink pipeline out of the primitives.
|
|
403
|
-
* **DAG orchestration** — a pipeline is a linear chain; branch inside a task with `fanout`.
|
|
510
|
+
* **DAG orchestration** — a pipeline is a linear chain; branch inside a task with `fanout`. The one
|
|
511
|
+
qualification is the opt-in [handoff](#advanced-handoffs-opt-in): it changes the traversal of
|
|
512
|
+
the chain along declared edges, never its topology (no joins, no second entry point, no cross-pipeline
|
|
513
|
+
jumps).
|
|
404
514
|
* **A serving gateway** — the only HTTP surface is the read-only debug endpoint above.
|
|
405
515
|
* **Distributed scheduling** — scale out with `--shard`; multi-process is the ceiling.
|
|
406
516
|
|
|
@@ -409,6 +519,17 @@ every known tradeoff with its reason.
|
|
|
409
519
|
|
|
410
520
|
## Status
|
|
411
521
|
|
|
522
|
+
**0.3.0 — advanced control flow (opt-in), and the documentation in Chinese.** Tasks can now
|
|
523
|
+
[hand off](#advanced-handoffs-opt-in): return a `Handoff` to skip declared stations or finish the pipeline
|
|
524
|
+
early, recorded in a durable ledger that recovery resumes from. It is opt-in and inert — a pipeline without a
|
|
525
|
+
`control` block writes no new rows and keeps a byte-identical `spec_digest` — and marked experimental until
|
|
526
|
+
1.0. Backward traversal now adds declared rewind and retry-all with visits and optional payload history. The
|
|
527
|
+
whole documentation set also ships in Simplified Chinese ([`README.zh-CN.md`](README.zh-CN.md),
|
|
528
|
+
[`docs/zh-CN/`](https://github.com/Hazer-BJTU/pyattacker/tree/main/docs/zh-CN)), kept in sync by CI. One
|
|
529
|
+
rename to know about when upgrading: `MergedReport.events_total` is now `source_events_total`, because that
|
|
530
|
+
is the one counter a merged report does **not** de-duplicate (the old name still works as a deprecated
|
|
531
|
+
alias).
|
|
532
|
+
|
|
412
533
|
**0.2.0 — a benchmark, stricter identity, and three correctness fixes.** New: `pyattacker bench`, a
|
|
413
534
|
simulated provider world that compares the acquire algorithms on a vector of metrics instead of a weighted
|
|
414
535
|
score ([`docs/benchmark.md`](https://github.com/Hazer-BJTU/pyattacker/blob/main/docs/benchmark.md));
|