cow-backtester 0.9.0__tar.gz
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- cow_backtester-0.9.0/CHANGELOG.md +136 -0
- cow_backtester-0.9.0/CONTRIBUTING.md +25 -0
- cow_backtester-0.9.0/LICENSE +21 -0
- cow_backtester-0.9.0/MANIFEST.in +5 -0
- cow_backtester-0.9.0/PKG-INFO +411 -0
- cow_backtester-0.9.0/README.md +385 -0
- cow_backtester-0.9.0/cow_backtester/__init__.py +6 -0
- cow_backtester-0.9.0/cow_backtester/__main__.py +11 -0
- cow_backtester-0.9.0/cow_backtester/_version.py +1 -0
- cow_backtester-0.9.0/cow_backtester/backtest.py +1700 -0
- cow_backtester-0.9.0/cow_backtester/cache.py +65 -0
- cow_backtester-0.9.0/cow_backtester/competition.py +147 -0
- cow_backtester-0.9.0/cow_backtester/economics.py +172 -0
- cow_backtester-0.9.0/cow_backtester/report.py +204 -0
- cow_backtester-0.9.0/cow_backtester/scorer.py +518 -0
- cow_backtester-0.9.0/cow_backtester.egg-info/PKG-INFO +411 -0
- cow_backtester-0.9.0/cow_backtester.egg-info/SOURCES.txt +27 -0
- cow_backtester-0.9.0/cow_backtester.egg-info/dependency_links.txt +1 -0
- cow_backtester-0.9.0/cow_backtester.egg-info/entry_points.txt +2 -0
- cow_backtester-0.9.0/cow_backtester.egg-info/requires.txt +5 -0
- cow_backtester-0.9.0/cow_backtester.egg-info/top_level.txt +1 -0
- cow_backtester-0.9.0/fixtures/body_8339027_trimmed.json +75 -0
- cow_backtester-0.9.0/fixtures/settle_8339027_calldata.hex +1 -0
- cow_backtester-0.9.0/fixtures/settlement_8339027.json +18 -0
- cow_backtester-0.9.0/mock_solver.py +88 -0
- cow_backtester-0.9.0/pyproject.toml +52 -0
- cow_backtester-0.9.0/requirements.txt +1 -0
- cow_backtester-0.9.0/setup.cfg +4 -0
- cow_backtester-0.9.0/tests/test_offline.py +802 -0
|
@@ -0,0 +1,136 @@
|
|
|
1
|
+
# Changelog
|
|
2
|
+
|
|
3
|
+
## 0.9.0 — 2026-08-18
|
|
4
|
+
|
|
5
|
+
**`--reward-ev`: CIP-85 v2 consistency economics.** Capture ratios and ranks
|
|
6
|
+
measure competitiveness; solver income on most CoW chains is the consistency
|
|
7
|
+
pool. This release computes the actual v2 metric
|
|
8
|
+
(`Σ executed orders: your_best_fair_surplus / Σ all_solvers_surplus`) from
|
|
9
|
+
the same competition records `--compete` already fetches:
|
|
10
|
+
|
|
11
|
+
- the challenger's counterfactual metric, per-order surpluses inserted into
|
|
12
|
+
the historical denominators (conservatively — field terms keep their
|
|
13
|
+
historical values, so the reported share is a floor);
|
|
14
|
+
- every FIELD solver's real historical metric — a consistency leaderboard
|
|
15
|
+
for the window (also useful for spotting pool-dilution patterns);
|
|
16
|
+
- with `--self-address`, your actual historical metric side by side with
|
|
17
|
+
the replayed one;
|
|
18
|
+
- the win floor, reported explicitly: v2 pays ZERO on a chain where the
|
|
19
|
+
solver won nothing in the period, so the estimate never hides it;
|
|
20
|
+
- `--consistency-budget <COW>` converts the share into a COW/week estimate.
|
|
21
|
+
|
|
22
|
+
Honesty labels throughout: field surpluses come from record amounts (net of
|
|
23
|
+
protocol fees) while a replayed challenger's are gross (a few bps flattering,
|
|
24
|
+
labeled `basis`); challenger fairness is assumed (the field uses CoW's own
|
|
25
|
+
`filteredOut` flags); success_rate is a flag, not a silent multiplier.
|
|
26
|
+
|
|
27
|
+
|
|
28
|
+
## 0.8.0 — 2026-08-17
|
|
29
|
+
|
|
30
|
+
**`--compete`: rank your solver against the historical field.** The v2
|
|
31
|
+
competition endpoint serves per-auction records by auction id (probed:
|
|
32
|
+
retention ≥ 2 months) carrying every submitted solution — solver, score,
|
|
33
|
+
ranking, CoW's own fairness-filtering outcome (`filteredOut`), per-solution
|
|
34
|
+
clearing prices, reference scores, and the auction's start/deadline blocks.
|
|
35
|
+
|
|
36
|
+
- `--compete` fetches each scored auction's record and inserts the
|
|
37
|
+
challenger's surplus into the fairness-surviving score list: per-auction
|
|
38
|
+
`field_rank` (rank, field size, gap-to-winner bps, winner solver), and a
|
|
39
|
+
per-solver summary — rank-1 %, top-3 %, median rank/gap, and a rivals
|
|
40
|
+
table (who beat you, how often, by how much). Honesty label everywhere:
|
|
41
|
+
`rank_basis: surplus_vs_score_proxy` — historical scores include protocol
|
|
42
|
+
fees, challenger surplus does not, so the rank is a floor.
|
|
43
|
+
- `--archive-dir DIR` persists every fetched record as
|
|
44
|
+
`DIR/<chain>/<auction_id>.json.gz` (idempotent) — a local competition
|
|
45
|
+
dataset that outlives the API's retention window and feeds future
|
|
46
|
+
fairness-simulation / reward-EV work.
|
|
47
|
+
- `--self-address 0x…` adds shadow-vs-actual: when your historical
|
|
48
|
+
solverAddress appears in a record, the row carries its actual score,
|
|
49
|
+
ranking, and fairness outcome next to the replayed one.
|
|
50
|
+
- Rows gain `auction_start_block` / `auction_deadline_block` — the auction
|
|
51
|
+
CUT context (what bidders saw), complementing `settlement_block`.
|
|
52
|
+
|
|
53
|
+
## 0.7.2 — 2026-08-17
|
|
54
|
+
|
|
55
|
+
Scoring-fidelity batch from an independent line-level audit. Every item
|
|
56
|
+
below changes reported numbers or their labels — see README caveats.
|
|
57
|
+
|
|
58
|
+
- **Coverage-adjusted capture is the new headline.** Solver errors now keep
|
|
59
|
+
the historical winner's surplus in the denominator (previously an errored
|
|
60
|
+
auction vanished from the ratio — a solver could look better by failing
|
|
61
|
+
hard auctions). The old ratio is kept as `capture_conditional_pct`
|
|
62
|
+
("when it answered, how competitive was it"), and a new
|
|
63
|
+
`lost_to_errors_wei` reports winner surplus forfeited to errors/timeouts.
|
|
64
|
+
- **Exact best-combination optimizer.** The greedy highest-first selection
|
|
65
|
+
of internally compatible solutions undercounted challengers (A=10 on two
|
|
66
|
+
pairs beats B+C=6+6 split — greedy picked 10, optimum is 12). Now solved
|
|
67
|
+
exactly (branch-and-bound; greedy fallback above 20 candidates, labeled
|
|
68
|
+
via `combiner`).
|
|
69
|
+
- **U256 range enforcement.** `to_u256` accepted arbitrarily large Python
|
|
70
|
+
ints; values above 2^256-1 are now rejected as invalid.
|
|
71
|
+
- **Scoring-basis labels.** Every row carries `baseline_quality`
|
|
72
|
+
(`exact_uniform` / `wrapper_lower_bound` / `mixed`); a same-basis
|
|
73
|
+
`capture_exact_basis_pct` (direct settlements only) is reported alongside
|
|
74
|
+
the mixed-basis proxy, and the HTML cards are relabeled accordingly
|
|
75
|
+
("Winner fee take" → "Known direct-path fee wedge").
|
|
76
|
+
- **Per-transaction wrapper attribution.** UID-overlap proof is now enforced
|
|
77
|
+
per settlement transaction, not per auction group (one attributed tx can
|
|
78
|
+
no longer vouch for an unrelated one; drops count as `tx_uid_mismatch`).
|
|
79
|
+
- **A/B on byte-identical bodies.** One shared `deadline` per auction is
|
|
80
|
+
computed before the solver loop (was regenerated per request); rows carry
|
|
81
|
+
`replay_deadline`.
|
|
82
|
+
- **Readiness semantics.** `{"solutions": []}` counts as a healthy answer
|
|
83
|
+
(schema-legitimate abstention) with bid coverage reported separately;
|
|
84
|
+
"valid" wording clarified to auctions with ≥1 valid solution; READY now
|
|
85
|
+
requires `--min-evidence` attempted auctions (default 10) — below that the
|
|
86
|
+
best verdict is REVIEW. Readiness is now rendered in the HTML report.
|
|
87
|
+
- **`settlement_block`** replaces the row field `block` (kept one release as
|
|
88
|
+
a deprecated alias): it is the settlement's block, NOT the auction cut
|
|
89
|
+
block — README fork guidance corrected to match.
|
|
90
|
+
- Docs: `--max-auctions` keeps the **newest** auctions (behavior since
|
|
91
|
+
0.7.1; help/README said oldest).
|
|
92
|
+
|
|
93
|
+
## 0.7.1 — 2026-08-16
|
|
94
|
+
|
|
95
|
+
- Wrapper-routed settlements (solver router contracts, ~43% of mainnet /
|
|
96
|
+
~40% of Base settlements) are attributed via the v2 by-tx-hash endpoint,
|
|
97
|
+
proven by UID overlap against the S3 auction body, and scored on a
|
|
98
|
+
clearly-labeled delivered basis (`entry: "wrapper"`). Previously excluded
|
|
99
|
+
silently.
|
|
100
|
+
- Submitter identification prefers the competition `solverAddress` over the
|
|
101
|
+
relay EOA. Coverage disclosure in the scorecard, `--readiness`, and a
|
|
102
|
+
`_meta` line in `--json-out`. `--max-auctions` keeps the newest auctions.
|
|
103
|
+
|
|
104
|
+
## 0.7.0 — 2026-08-12
|
|
105
|
+
|
|
106
|
+
- `--readiness`: a one-screen pre-production readiness check for a solver
|
|
107
|
+
endpoint. Replays recent auctions against the endpoint and reports answer
|
|
108
|
+
rate, latency (p50/p95/max vs the solve budget), solution validity, and
|
|
109
|
+
surplus captured against the on-chain winners, with a READY / REVIEW /
|
|
110
|
+
NOT READY verdict and a copy-paste command to reproduce the exact run.
|
|
111
|
+
A focused alternative to the full field scorecard for anyone evaluating a
|
|
112
|
+
solver before staging or shadow. Emitted in the `--json-out`/`--html-out`
|
|
113
|
+
summary under `readiness`.
|
|
114
|
+
|
|
115
|
+
## 0.6.0 — 2026-08-07
|
|
116
|
+
|
|
117
|
+
First public release.
|
|
118
|
+
|
|
119
|
+
- Replays archived CoW auctions (the public S3 instance bucket) against any
|
|
120
|
+
solver's `/solve` endpoint, offline.
|
|
121
|
+
- Reconstructs each auction's winning set from on-chain settlement calldata
|
|
122
|
+
and Trade events; groups multiple winning transactions per auction (CIP-67);
|
|
123
|
+
optional cross-check against the v2 competition API (`--verify-api`).
|
|
124
|
+
- Scores both sides on the same basis: before-fee surplus over signed limits
|
|
125
|
+
at uniform clearing prices, converted at the auction's reference prices.
|
|
126
|
+
- Validates solver responses the way the settlement layer would: per-order
|
|
127
|
+
fill accounting, fill-or-kill exactness, fee-adjusted limit feasibility,
|
|
128
|
+
strict U256 parsing; infeasible solutions are excluded and tallied.
|
|
129
|
+
- A/B mode: repeatable `--solver-url`, rotating call order, head-to-head
|
|
130
|
+
panel, per-pair breakdown, solve-latency stats.
|
|
131
|
+
- Ten chains, content cache, concurrent fetching, streamed JSONL rows,
|
|
132
|
+
single-file HTML report, watch mode.
|
|
133
|
+
- 37 offline tests against pinned fixtures; CI runs them plus lint and a
|
|
134
|
+
packaging check on Python 3.10 and 3.12.
|
|
135
|
+
|
|
136
|
+
Pre-release development history is summarized in `docs/DESIGN_NOTES.md`.
|
|
@@ -0,0 +1,25 @@
|
|
|
1
|
+
# Contributing
|
|
2
|
+
|
|
3
|
+
Corrections are the most valuable contribution — especially anywhere this
|
|
4
|
+
tool's accounting diverges from what the protocol actually does. Every scoring
|
|
5
|
+
rule cites its source (GPv2 contracts or `cowprotocol/services`); if you can
|
|
6
|
+
show a rule is wrong, please open an issue with the primary-source reference.
|
|
7
|
+
|
|
8
|
+
## Development
|
|
9
|
+
|
|
10
|
+
```bash
|
|
11
|
+
pip install -e ".[dev]"
|
|
12
|
+
python3 -m pytest # offline suite — no network, must always pass
|
|
13
|
+
python3 -m ruff check . # lint — must be clean
|
|
14
|
+
python3 -m cow_backtester.scorer --selftest && python3 -m cow_backtester --selftest # network tests
|
|
15
|
+
```
|
|
16
|
+
|
|
17
|
+
Ground rules:
|
|
18
|
+
- The offline suite stays offline. New scoring behavior needs a fixture-based
|
|
19
|
+
test, and exploit-class regressions (duplicate trades, fee padding, negative
|
|
20
|
+
numerics, malformed shapes) must keep passing.
|
|
21
|
+
- No silent drops: anything excluded from scoring gets a tallied reason that
|
|
22
|
+
reaches the scorecard.
|
|
23
|
+
- No new runtime dependencies without strong cause (`eth_abi` is the only one).
|
|
24
|
+
- The pinned reference values (auction 8339027: event 2385773, surplus 1471,
|
|
25
|
+
fee 805) are load-bearing; never adjust them to make a test pass.
|
|
@@ -0,0 +1,21 @@
|
|
|
1
|
+
MIT License
|
|
2
|
+
|
|
3
|
+
Copyright (c) 2026 kaisersolver
|
|
4
|
+
|
|
5
|
+
Permission is hereby granted, free of charge, to any person obtaining a copy
|
|
6
|
+
of this software and associated documentation files (the "Software"), to deal
|
|
7
|
+
in the Software without restriction, including without limitation the rights
|
|
8
|
+
to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
|
|
9
|
+
copies of the Software, and to permit persons to whom the Software is
|
|
10
|
+
furnished to do so, subject to the following conditions:
|
|
11
|
+
|
|
12
|
+
The above copyright notice and this permission notice shall be included in all
|
|
13
|
+
copies or substantial portions of the Software.
|
|
14
|
+
|
|
15
|
+
THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
|
|
16
|
+
IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
|
|
17
|
+
FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
|
|
18
|
+
AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
|
|
19
|
+
LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
|
|
20
|
+
OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
|
|
21
|
+
SOFTWARE.
|
|
@@ -0,0 +1,411 @@
|
|
|
1
|
+
Metadata-Version: 2.4
|
|
2
|
+
Name: cow-backtester
|
|
3
|
+
Version: 0.9.0
|
|
4
|
+
Summary: Offline backtester, A/B harness, and counterfactual scorecard for CoW Protocol solvers
|
|
5
|
+
Author: kaisersolver
|
|
6
|
+
License-Expression: MIT
|
|
7
|
+
Project-URL: Repository, https://github.com/kaisersolver/cow-backtester
|
|
8
|
+
Project-URL: Issues, https://github.com/kaisersolver/cow-backtester/issues
|
|
9
|
+
Project-URL: Changelog, https://github.com/kaisersolver/cow-backtester/blob/main/CHANGELOG.md
|
|
10
|
+
Keywords: cow-protocol,solver,backtesting,mev,defi
|
|
11
|
+
Classifier: Development Status :: 4 - Beta
|
|
12
|
+
Classifier: Environment :: Console
|
|
13
|
+
Classifier: Intended Audience :: Developers
|
|
14
|
+
Classifier: Programming Language :: Python :: 3.10
|
|
15
|
+
Classifier: Programming Language :: Python :: 3.11
|
|
16
|
+
Classifier: Programming Language :: Python :: 3.12
|
|
17
|
+
Classifier: Programming Language :: Python :: 3.13
|
|
18
|
+
Requires-Python: >=3.10
|
|
19
|
+
Description-Content-Type: text/markdown
|
|
20
|
+
License-File: LICENSE
|
|
21
|
+
Requires-Dist: eth_abi<6,>=5
|
|
22
|
+
Provides-Extra: dev
|
|
23
|
+
Requires-Dist: pytest>=7; extra == "dev"
|
|
24
|
+
Requires-Dist: ruff>=0.4; extra == "dev"
|
|
25
|
+
Dynamic: license-file
|
|
26
|
+
|
|
27
|
+
# cow-backtester
|
|
28
|
+
|
|
29
|
+
Offline backtester, A/B harness, and counterfactual scorecard for CoW Protocol
|
|
30
|
+
solvers.
|
|
31
|
+
|
|
32
|
+
Point it at your own solver's `/solve` endpoint and it replays recent CoW
|
|
33
|
+
auctions against it, then scores each of your solutions against the set of
|
|
34
|
+
solutions that actually won on-chain. Point it at two endpoints and it runs
|
|
35
|
+
a head-to-head A/B on identical auctions, so a routing or config change can
|
|
36
|
+
be judged before it goes to production.
|
|
37
|
+
|
|
38
|
+
Nothing touches production: no shadow mode, no staging deployment, no keys.
|
|
39
|
+
|
|
40
|
+
## Why this exists
|
|
41
|
+
|
|
42
|
+
Today a solver can only be evaluated *live*: shadow mode consumes the
|
|
43
|
+
production auction stream, and the local playground runs against a chain fork.
|
|
44
|
+
Neither lets you take a fixed set of recent auctions, run your solver against
|
|
45
|
+
them offline, and ask "would I have out-surplused the winners, and by how
|
|
46
|
+
much?" the way you'd backtest a trading strategy. This tool does.
|
|
47
|
+
|
|
48
|
+
## Quick start
|
|
49
|
+
|
|
50
|
+
Python 3.10+.
|
|
51
|
+
|
|
52
|
+
```bash
|
|
53
|
+
pip install . # installs the `cow-backtester` command (one dependency: eth_abi)
|
|
54
|
+
|
|
55
|
+
# 1. Baseline — what the field actually captured (no solver needed)
|
|
56
|
+
cow-backtester --chain base --blocks 2000 --rpc-url <your-rpc>
|
|
57
|
+
|
|
58
|
+
# 2. Counterfactual — replay through your solver
|
|
59
|
+
cow-backtester --chain base --blocks 2000 --rpc-url <your-rpc> \
|
|
60
|
+
--solver-url http://localhost:8080 --json-out results.jsonl
|
|
61
|
+
|
|
62
|
+
# 3. A/B — two solvers, same auctions, head-to-head
|
|
63
|
+
cow-backtester --chain base --blocks 5000 --rpc-url <your-rpc> \
|
|
64
|
+
--solver-url http://localhost:8080 --solver-name baseline \
|
|
65
|
+
--solver-url http://localhost:8081 --solver-name candidate \
|
|
66
|
+
--html-out ab.html
|
|
67
|
+
```
|
|
68
|
+
|
|
69
|
+
From a source checkout without installing, `python3 -m cow_backtester ...`
|
|
70
|
+
works identically. Your endpoint only needs the standard CoW solver-engine
|
|
71
|
+
API (`POST /solve`).
|
|
72
|
+
|
|
73
|
+
## Readiness check
|
|
74
|
+
|
|
75
|
+
If you are bringing up a new solver and want a fast "is this endpoint healthy
|
|
76
|
+
enough to face production auctions?" read, `--readiness` prints a one-screen
|
|
77
|
+
report instead of the full field scorecard:
|
|
78
|
+
|
|
79
|
+
```bash
|
|
80
|
+
cow-backtester --chain base --blocks 2000 --rpc-url <your-rpc> \
|
|
81
|
+
--solver-url http://localhost:8080 --solver-name mine --readiness
|
|
82
|
+
```
|
|
83
|
+
|
|
84
|
+
It replays recent auctions against your endpoint and reports the four things
|
|
85
|
+
that gate a pre-prod solver — does it answer, is it fast enough, are its
|
|
86
|
+
solutions valid, and are they competitive with the on-chain winners — as
|
|
87
|
+
pass/warn checks with a `READY` / `REVIEW` / `NOT READY` verdict:
|
|
88
|
+
|
|
89
|
+
```
|
|
90
|
+
====================================================================
|
|
91
|
+
READINESS — mine [REVIEW]
|
|
92
|
+
base · prod · blocks 49506127..49510000
|
|
93
|
+
====================================================================
|
|
94
|
+
[PASS] reached auctions 50 auctions attempted
|
|
95
|
+
[WARN] no transport errors 2/50 errored
|
|
96
|
+
[PASS] answers reliably 96% returned a solution
|
|
97
|
+
[PASS] inside the deadline 0 past deadline
|
|
98
|
+
[PASS] latency headroom p95 420 ms of a 15000 ms budget
|
|
99
|
+
[PASS] solutions are valid 98% passed limit/fee checks
|
|
100
|
+
[PASS] competitive vs winners 61% of winner surplus captured
|
|
101
|
+
```
|
|
102
|
+
|
|
103
|
+
It prints the exact `--from-block/--to-block` command to reproduce the run,
|
|
104
|
+
and the same data lands in `--json-out`/`--html-out` under `readiness`. Works
|
|
105
|
+
on any of the supported chains, so you can readiness-check an endpoint for a
|
|
106
|
+
chain you are not yet onboarded on. It is a signal, not a settlement
|
|
107
|
+
guarantee — pair it with a self-hosted shadow run before going to production.
|
|
108
|
+
|
|
109
|
+
## Consistency economics (`--reward-ev`)
|
|
110
|
+
|
|
111
|
+
On most CoW chains the money is not in winning — it is in the CIP-85
|
|
112
|
+
consistency pool, which pays
|
|
113
|
+
`success_rate × Σ executed orders ( your best fair bid's surplus / everyone's )`.
|
|
114
|
+
`--reward-ev` computes exactly that metric from the competition records:
|
|
115
|
+
your solver's counterfactual metric and pool share, every field solver's
|
|
116
|
+
real historical metric (a consistency leaderboard for the window), and,
|
|
117
|
+
with `--consistency-budget <COW>`, a COW/week estimate. The win floor is
|
|
118
|
+
reported explicitly — v2 pays zero on a chain where a solver won nothing in
|
|
119
|
+
the period — and every number carries its basis label (field surpluses are
|
|
120
|
+
net of protocol fees, a replayed challenger's are gross; challenger
|
|
121
|
+
fairness is assumed while the field uses CoW's own `filteredOut` flags).
|
|
122
|
+
|
|
123
|
+
## Rank against the historical field (`--compete`)
|
|
124
|
+
|
|
125
|
+
Capture ratio tells you how much surplus you generate; `--compete` tells you
|
|
126
|
+
**where you would have ranked**. For every scored auction it fetches the
|
|
127
|
+
historical competition record (every submitted solution with its score,
|
|
128
|
+
CoW's own fairness-filtering outcome, and the winner) and inserts your
|
|
129
|
+
solver's result into the fairness-surviving score list:
|
|
130
|
+
|
|
131
|
+
```bash
|
|
132
|
+
cow-backtester --chain base --blocks 2000 --rpc-url <your-rpc> \
|
|
133
|
+
--solver-url http://localhost:8080 --solver-name mine --compete \
|
|
134
|
+
--archive-dir ./competition-data
|
|
135
|
+
```
|
|
136
|
+
|
|
137
|
+
The scorecard gains a field-rank line (rank-1 %, top-3 %, median rank,
|
|
138
|
+
median gap to the winner in bps) and a rivals table — which solvers beat
|
|
139
|
+
you, how often, and by how much. Per-auction `field_rank` lands in the JSON
|
|
140
|
+
rows. Two honesty notes: historical scores include protocol fees while your
|
|
141
|
+
replayed surplus does not, so the reported rank is a **floor** (labeled
|
|
142
|
+
`surplus_vs_score_proxy`); and records exist only for auctions that had a
|
|
143
|
+
winner. `--archive-dir` keeps every fetched record on disk — the API serves
|
|
144
|
+
roughly two months of history, so an archive you build today is a dataset
|
|
145
|
+
you keep. `--self-address` marks your historical solverAddress so rows show
|
|
146
|
+
shadow-vs-actual side by side.
|
|
147
|
+
|
|
148
|
+
### Try it without a solver
|
|
149
|
+
|
|
150
|
+
The bundled mock solver lets you see the full counterfactual/A/B output in
|
|
151
|
+
about a minute, before wiring up your own engine:
|
|
152
|
+
|
|
153
|
+
```bash
|
|
154
|
+
MODE=limit python3 mock_solver.py 8901 & # fills exactly at the limit
|
|
155
|
+
MODE=better python3 mock_solver.py 8902 & # fills at 1.5x the limit
|
|
156
|
+
|
|
157
|
+
cow-backtester --chain base --blocks 1500 --rpc-url <your-rpc> \
|
|
158
|
+
--solver-url http://127.0.0.1:8901 --solver-name at-limit \
|
|
159
|
+
--solver-url http://127.0.0.1:8902 --solver-name better --html-out demo.html
|
|
160
|
+
kill %1 %2
|
|
161
|
+
```
|
|
162
|
+
|
|
163
|
+
You get the per-solver panels, the head-to-head table, and the pair
|
|
164
|
+
breakdown. The `implausible_surplus` guard usually fires on the "better"
|
|
165
|
+
mock: claiming 1.5x the limit on real order flow is exactly what the
|
|
166
|
+
validator exists to flag.
|
|
167
|
+
|
|
168
|
+
## How it works (all public data)
|
|
169
|
+
|
|
170
|
+
1. Input: the exact `/solve` bodies CoW sent solvers, from the public S3
|
|
171
|
+
instance bucket (`solver-instances.s3.amazonaws.com/<env>/<chain>/auction/<id>.json`).
|
|
172
|
+
The archived file *is* a valid `/solve` request body.
|
|
173
|
+
2. Winner reconstruction: from the on-chain settlements. (The v1
|
|
174
|
+
`/solver_competition` endpoints were removed in the competition-data
|
|
175
|
+
migration; v2 exists but is retention-bounded. Reconstructing from chain
|
|
176
|
+
data is trustless and independent of API retention — and `--verify-api`
|
|
177
|
+
cross-checks it against v2 where available.) Settlements are grouped by
|
|
178
|
+
auction: since combinatorial auctions (CIP-67), one auction can settle
|
|
179
|
+
through several winning transactions, so the baseline is the combined
|
|
180
|
+
winning set.
|
|
181
|
+
3. Replay and score: each auction body is POSTed to each solver; responses
|
|
182
|
+
are validated (below) and scored with the same function, the same
|
|
183
|
+
limits, and the same price convention as the winners. Your valid solutions
|
|
184
|
+
are combined the CIP-67 way (best-first, disjoint directed token pairs).
|
|
185
|
+
|
|
186
|
+
## The comparison basis (read this before quoting numbers)
|
|
187
|
+
|
|
188
|
+
Both sides are scored on **before-fee surplus over the signed order limits, at
|
|
189
|
+
uniform clearing prices**, converted to the chain's native token at the
|
|
190
|
+
auction's own reference prices.
|
|
191
|
+
|
|
192
|
+
* Signed limits: CoW's accounting scores against the order's signed amounts
|
|
193
|
+
(`fullSellAmount`/`fullBuyAmount`), not the remaining/fee-adjusted amounts.
|
|
194
|
+
* Uniform prices: on-chain settlements carry per-trade post-fee "custom"
|
|
195
|
+
prices; the first occurrence of each token in the price vector is the uniform
|
|
196
|
+
(pre-fee) price, which is what autopilot itself resolves. Scoring the winner
|
|
197
|
+
post-fee but a challenger pre-fee would bias every comparison.
|
|
198
|
+
* Surplus token: buy token for sell orders, sell token for buy orders,
|
|
199
|
+
converted at that token's `referencePrice` (matches official accounting).
|
|
200
|
+
* What the number is: user surplus + protocol fees + network fee. CoW's
|
|
201
|
+
official ranking score is user surplus + protocol fees only, so absolute
|
|
202
|
+
levels here overstate the official score by the network fee (on the pinned
|
|
203
|
+
reference trade, a $2.4 fill where the network fee dominates: 1471 vs an
|
|
204
|
+
official 666 atoms; the gap shrinks as trades grow). When two solutions carry
|
|
205
|
+
very different fees this basis can order them differently than CoW would;
|
|
206
|
+
gas-heavy solutions are flattered. It is the only basis we found that can
|
|
207
|
+
be computed symmetrically offline for both sides.
|
|
208
|
+
|
|
209
|
+
## Response validation (your solver can't accidentally cheat)
|
|
210
|
+
|
|
211
|
+
A solution is scored only if **every** fulfillment trade is feasible:
|
|
212
|
+
|
|
213
|
+
* known order uid (case-insensitive) and strict U256 numerics (decimal or
|
|
214
|
+
0x-hex; negatives and malformed values rejected);
|
|
215
|
+
* prices present for both tokens;
|
|
216
|
+
* at most one fulfillment per order (GPv2 accumulates fills and the driver
|
|
217
|
+
rejects duplicate trades, so N copies of a trade can't score N times the
|
|
218
|
+
surplus);
|
|
219
|
+
* no over-fill (`executed + fee` vs `sellAmount` for sell orders, `executed`
|
|
220
|
+
vs `buyAmount` for buy orders; exact for fill-or-kill);
|
|
221
|
+
* on-chain feasibility at the fee-adjusted terms: the net delivery must still
|
|
222
|
+
cover the gross-scaled limit, which is what GPv2 enforces, so a padded `fee`
|
|
223
|
+
cannot manufacture surplus.
|
|
224
|
+
|
|
225
|
+
Violations mark the whole **solution INVALID** (tallied with a reason), never a
|
|
226
|
+
silent zero. Solutions scoring >10× the winning set (or >0.001 native when the
|
|
227
|
+
winning set scored 0) are flagged `implausible_surplus`. Malformed responses
|
|
228
|
+
are tallied, never crash a run. **Prices are claimed, not simulated** — the
|
|
229
|
+
tool checks feasibility; it does not execute routes.
|
|
230
|
+
|
|
231
|
+
## What a run prints
|
|
232
|
+
|
|
233
|
+
* COVERAGE: settlements found, auctions formed/scored, every skip with its
|
|
234
|
+
reason, unscanned block spans, auction ages, cache stats, `--verify-api` results.
|
|
235
|
+
* FIELD SURPLUS: winner surplus by USD size bucket, plus the winners' fee take.
|
|
236
|
+
* WINNING SUBMITTERS: who is actually winning (label them with `--solver-map`).
|
|
237
|
+
* COUNTERFACTUAL (per solver): returned / valid / positive / beat the
|
|
238
|
+
winning set, surplus sums, capture ratio, solve latency p50+p95, invalid
|
|
239
|
+
reasons, error classes.
|
|
240
|
+
* HEAD TO HEAD (exactly two solvers): side-by-side metrics, per-auction
|
|
241
|
+
win counts, and the surplus delta.
|
|
242
|
+
* TOP PAIRS: winner surplus by token pair with each solver's surplus
|
|
243
|
+
beside it: *where* you win and lose, not just by how much.
|
|
244
|
+
|
|
245
|
+
A real single-solver run (Arbitrum, eight auctions, trimmed):
|
|
246
|
+
|
|
247
|
+
```
|
|
248
|
+
====================================================================
|
|
249
|
+
COUNTERFACTUAL — kaisersolver (8 replayed)
|
|
250
|
+
====================================================================
|
|
251
|
+
returned a solution : 7/8
|
|
252
|
+
valid solutions : 7/8
|
|
253
|
+
positive surplus : 7/8
|
|
254
|
+
beat the winning set: 0/8
|
|
255
|
+
our surplus (sum) : 0.001643 ETH
|
|
256
|
+
winners (sum) : 0.001748 ETH (replayed auctions only)
|
|
257
|
+
capture ratio : 94.0% of the winning set
|
|
258
|
+
solve latency : p50 1174 ms / p95 2471 ms
|
|
259
|
+
|
|
260
|
+
====================================================================
|
|
261
|
+
TOP PAIRS BY WINNER SURPLUS (top 2)
|
|
262
|
+
====================================================================
|
|
263
|
+
pair trades winner kaisersolver
|
|
264
|
+
WETH->MOR 7 0.001722 0.001643
|
|
265
|
+
USDC->USD₮0 1 0.000026 0.000000
|
|
266
|
+
```
|
|
267
|
+
|
|
268
|
+
Eight auctions is a small sample, but the shape is the point: this solver is
|
|
269
|
+
consistently a few percent behind one competitor on one pair, and the table
|
|
270
|
+
says which pair. That took minutes to learn here; it took weeks from logs.
|
|
271
|
+
|
|
272
|
+
`--json-out` streams one row per auction (block, timestamp, age, winner txs +
|
|
273
|
+
submitters + surplus/fees, per-solver validation detail and latency,
|
|
274
|
+
`expired_orders_pct`, flags). `--html-out` writes a single self-contained HTML
|
|
275
|
+
report: no external assets, light/dark aware, fine to attach to a PR or post.
|
|
276
|
+
|
|
277
|
+
## Flags
|
|
278
|
+
|
|
279
|
+
| flag | meaning |
|
|
280
|
+
|---|---|
|
|
281
|
+
| `--chain` | `mainnet`, `arbitrum-one`, `base`, `xdai`, `polygon`, `bnb`, `avalanche`, `linea`, `ink`, `plasma` |
|
|
282
|
+
| `--env` | `prod` (default) or `staging` |
|
|
283
|
+
| `--blocks N` | scan the most recent N blocks |
|
|
284
|
+
| `--from-block/--to-block` | absolute, reproducible window |
|
|
285
|
+
| `--max-auctions N` | cap auctions scored, newest first (default 25; 0 = all) |
|
|
286
|
+
| `--rpc-url` | your RPC (recommended — public defaults rot and rate-limit) |
|
|
287
|
+
| `--solver-url` / `--solver-name` | repeatable; two of them = A/B |
|
|
288
|
+
| `--solve-timeout S` | deadline advertised to the solver (HTTP waits S+5) |
|
|
289
|
+
| `--workers N` | concurrent RPC/S3 fetches (default 8) |
|
|
290
|
+
| `--cache-dir` / `--no-cache` | content cache (default `.cowbt-cache`) |
|
|
291
|
+
| `--json-out` / `--html-out` | machine-readable rows / HTML report |
|
|
292
|
+
| `--solver-map FILE` | JSON `{address: name}` to label winning submitters |
|
|
293
|
+
| `--verify-api` | cross-check winner txs against the v2 competition API |
|
|
294
|
+
| `--compete` | rank the challenger against the historical fairness-surviving field (rank / gap / rivals) |
|
|
295
|
+
| `--archive-dir DIR` | persist competition records as `DIR/<chain>/<id>.json.gz` (local dataset) |
|
|
296
|
+
| `--self-address 0x…` | your historical solverAddress → shadow-vs-actual comparison |
|
|
297
|
+
| `--reward-ev` | CIP-85 v2 consistency economics: counterfactual metric, field leaderboard, COW estimate (implies `--compete`) |
|
|
298
|
+
| `--consistency-budget N` | the chain's weekly consistency pool in COW, to convert share → COW/week |
|
|
299
|
+
| `--min-evidence N` | attempted-auction floor before `--readiness` may say READY (default 10) |
|
|
300
|
+
| `--clamp-validto` | extend expired `validTo` so engines that filter them still solve |
|
|
301
|
+
| `--max-age-hours H` | warn when replayed auctions are older than this |
|
|
302
|
+
| `--watch N` | continuous mode: rescan every N seconds from the last block |
|
|
303
|
+
| `--quiet` | suppress progress (stderr); the scorecard still prints to stdout |
|
|
304
|
+
| `--version` | print version |
|
|
305
|
+
|
|
306
|
+
Progress goes to **stderr**, the scorecard to **stdout** (`2>/dev/null` gives
|
|
307
|
+
clean results; `--json-out`/`--html-out` are unaffected). Exit codes: **0**
|
|
308
|
+
success, **1** runtime error, **2** usage error, **130** interrupted mid-run
|
|
309
|
+
(a partial scorecard is printed when auctions had already been processed;
|
|
310
|
+
stopping `--watch` during its idle sleep is a clean stop and exits 0).
|
|
311
|
+
|
|
312
|
+
## Caching
|
|
313
|
+
|
|
314
|
+
Only **immutable** facts are cached: archived auction bodies, settlements below
|
|
315
|
+
the reorg margin, and block timestamps. Nothing derived from live liquidity is
|
|
316
|
+
ever cached, so a warm re-run is faster (3.8x on a measured six-auction,
|
|
317
|
+
two-solver run) without changing a single number. Delete `.cowbt-cache` any time, or pass `--no-cache`.
|
|
318
|
+
|
|
319
|
+
## Limitations
|
|
320
|
+
|
|
321
|
+
* **Replay uses live liquidity.** Your solver quotes against *current* chain
|
|
322
|
+
state, not the historical block. The counterfactual is **indicative** — it
|
|
323
|
+
answers "how does my solver handle this real order flow", not "the exact
|
|
324
|
+
outcome at that block". Ages are printed and a warning fires beyond
|
|
325
|
+
`--max-age-hours` (default 6h). Prefer recent, short windows. Each row
|
|
326
|
+
carries `settlement_block` — the block the winning settlement landed in,
|
|
327
|
+
which is *later* than the auction cut block the bidders actually saw. A
|
|
328
|
+
fork pinned there is an approximation, not a faithful auction replay (it
|
|
329
|
+
can even include the settlement itself); true historical-fork replay needs
|
|
330
|
+
the auction cut block, which this tool does not yet reconstruct.
|
|
331
|
+
* **The S3 bucket retains roughly one month** of auctions (measured Aug 2026).
|
|
332
|
+
This is a recent-window backtester, not an archive.
|
|
333
|
+
* **Entrypoint coverage.** Direct `settle()` calls are attributed by the
|
|
334
|
+
auction id appended to their calldata. Wrapper-routed settlements (solver
|
|
335
|
+
router contracts — measured Aug 2026 at **~43% of mainnet, ~40% of Base,
|
|
336
|
+
~2% of Arbitrum** settlements, and *larger* than direct ones at the median,
|
|
337
|
+
so they are not a random slice) are attributed via the v2
|
|
338
|
+
`solver_competition/by_tx_hash` endpoint and **proven by uid overlap with
|
|
339
|
+
the S3 auction body** before scoring; they score on a DELIVERED basis
|
|
340
|
+
(Trade events vs signed limits), which understates the direct-path
|
|
341
|
+
before-fee basis by the settlement's fee wedge (typically a few bps) —
|
|
342
|
+
rows carry `entry: "wrapper"` so the bases are distinguishable. Settlements
|
|
343
|
+
the endpoint cannot resolve are counted under `wrapper_unattributed` and
|
|
344
|
+
disclosed in the coverage block, `--readiness`, and the `_meta` JSON line.
|
|
345
|
+
* **"Beat the winning set" is necessary, not sufficient.** Real winner
|
|
346
|
+
selection also applies fairness filters, and bids score net of gas; the tool
|
|
347
|
+
also takes your solutions at face value while the driver merges and simulates.
|
|
348
|
+
* **Orders keep their historical `validTo`** (~1%/day expire; more near the
|
|
349
|
+
retention edge). Engines that filter expired orders look age-degraded — see
|
|
350
|
+
`expired_orders_pct` per row, or pass `--clamp-validto`.
|
|
351
|
+
* **JIT orders.** The winner baseline credits surplus-capturing JIT trades
|
|
352
|
+
(owners in `surplusCapturingJitOrderOwners`), per official accounting; the
|
|
353
|
+
challenger side never credits JIT (a response's JIT order has no verifiable
|
|
354
|
+
owner). On auctions where winner surplus comes from CoW-AMM-style JIT, the
|
|
355
|
+
comparison is conservative *against* the challenger, never in its favor.
|
|
356
|
+
* **Bring your own RPC.** Public defaults rot and rate-limit `eth_getLogs`; the
|
|
357
|
+
tool splits and retries failing ranges and always **reports** any span it
|
|
358
|
+
could not scan rather than under-counting silently. `eth_chainId` is checked
|
|
359
|
+
against `--chain` on startup.
|
|
360
|
+
|
|
361
|
+
## Provenance of the reconstruction
|
|
362
|
+
|
|
363
|
+
* Auction id: CoW's driver appends it to `settle()` calldata; autopilot reads
|
|
364
|
+
back exactly the **last 8 bytes** (`META_DATA_LEN = 8` in
|
|
365
|
+
`cowprotocol/services`). The tool applies the same rule, only when the
|
|
366
|
+
calldata prefix re-encodes canonically, cross-checked by requiring **at least
|
|
367
|
+
one** settled order uid to appear in the fetched body (`body_uid_mismatch`
|
|
368
|
+
guard — this also catches staging/prod contamination, which shares the
|
|
369
|
+
settlement contract).
|
|
370
|
+
* Trade events are filtered by **emitting address + topic** and must align 1:1
|
|
371
|
+
with calldata trades, else the settlement is skipped (`trade_event_mismatch`).
|
|
372
|
+
* Fill-or-kill: GPv2 ignores the calldata `executedAmount` and uses the signed
|
|
373
|
+
amount — the tool applies the same substitution.
|
|
374
|
+
* Response-schema authority: `crates/solvers-dto` / the solver-engine OpenAPI
|
|
375
|
+
in `cowprotocol/services`.
|
|
376
|
+
* `--verify-api` compares the reconstructed winning-tx set per auction against
|
|
377
|
+
the v2 competition endpoint (match / mismatch / unavailable counters).
|
|
378
|
+
|
|
379
|
+
## Tests
|
|
380
|
+
|
|
381
|
+
```bash
|
|
382
|
+
pip install -e ".[dev]" # pytest + ruff (one-time)
|
|
383
|
+
python3 -m pytest # OFFLINE suite (fixtures; no network)
|
|
384
|
+
python3 -m cow_backtester.scorer --unittest # offline core checks
|
|
385
|
+
python3 -m cow_backtester.scorer --selftest # network: winner path on the reference settlement
|
|
386
|
+
python3 -m cow_backtester --selftest # network: counterfactual == winner baseline +
|
|
387
|
+
# every exploit class as regression cases
|
|
388
|
+
```
|
|
389
|
+
|
|
390
|
+
The pinned reference (Arbitrum auction 8339027): decode integrity is exact
|
|
391
|
+
against the GPv2 `Trade` event (2385773 atoms), before-fee surplus 1471 atoms,
|
|
392
|
+
fee take 805. `mock_solver.py` (modes: empty / limit / better / hex / garbage /
|
|
393
|
+
invalid) exercises the wire path including validation and implausibility
|
|
394
|
+
guards. CI runs the offline suite on Python 3.10 and 3.12.
|
|
395
|
+
|
|
396
|
+
## Files
|
|
397
|
+
|
|
398
|
+
| file | role |
|
|
399
|
+
|---|---|
|
|
400
|
+
| `cow_backtester/scorer.py` | settlement decode: auction id + winner surplus |
|
|
401
|
+
| `cow_backtester/backtest.py` | enumerate, group, replay, validate, scorecard |
|
|
402
|
+
| `cow_backtester/cache.py` | immutable-fact content cache |
|
|
403
|
+
| `cow_backtester/report.py` | single-file HTML report |
|
|
404
|
+
| `mock_solver.py` | test/demo solver (six modes) |
|
|
405
|
+
| `fixtures/` | pinned reference data for the offline tests |
|
|
406
|
+
| `tests/` | offline pytest suite, no network |
|
|
407
|
+
| `docs/DESIGN_NOTES.md` | design decisions and validation history |
|
|
408
|
+
|
|
409
|
+
MIT licensed. Contributions and corrections welcome — especially from CoW core
|
|
410
|
+
devs on anything where this tool's accounting diverges from the protocol's.
|
|
411
|
+
</content>
|