cow-backtester 0.9.0__tar.gz

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (29) hide show
  1. cow_backtester-0.9.0/CHANGELOG.md +136 -0
  2. cow_backtester-0.9.0/CONTRIBUTING.md +25 -0
  3. cow_backtester-0.9.0/LICENSE +21 -0
  4. cow_backtester-0.9.0/MANIFEST.in +5 -0
  5. cow_backtester-0.9.0/PKG-INFO +411 -0
  6. cow_backtester-0.9.0/README.md +385 -0
  7. cow_backtester-0.9.0/cow_backtester/__init__.py +6 -0
  8. cow_backtester-0.9.0/cow_backtester/__main__.py +11 -0
  9. cow_backtester-0.9.0/cow_backtester/_version.py +1 -0
  10. cow_backtester-0.9.0/cow_backtester/backtest.py +1700 -0
  11. cow_backtester-0.9.0/cow_backtester/cache.py +65 -0
  12. cow_backtester-0.9.0/cow_backtester/competition.py +147 -0
  13. cow_backtester-0.9.0/cow_backtester/economics.py +172 -0
  14. cow_backtester-0.9.0/cow_backtester/report.py +204 -0
  15. cow_backtester-0.9.0/cow_backtester/scorer.py +518 -0
  16. cow_backtester-0.9.0/cow_backtester.egg-info/PKG-INFO +411 -0
  17. cow_backtester-0.9.0/cow_backtester.egg-info/SOURCES.txt +27 -0
  18. cow_backtester-0.9.0/cow_backtester.egg-info/dependency_links.txt +1 -0
  19. cow_backtester-0.9.0/cow_backtester.egg-info/entry_points.txt +2 -0
  20. cow_backtester-0.9.0/cow_backtester.egg-info/requires.txt +5 -0
  21. cow_backtester-0.9.0/cow_backtester.egg-info/top_level.txt +1 -0
  22. cow_backtester-0.9.0/fixtures/body_8339027_trimmed.json +75 -0
  23. cow_backtester-0.9.0/fixtures/settle_8339027_calldata.hex +1 -0
  24. cow_backtester-0.9.0/fixtures/settlement_8339027.json +18 -0
  25. cow_backtester-0.9.0/mock_solver.py +88 -0
  26. cow_backtester-0.9.0/pyproject.toml +52 -0
  27. cow_backtester-0.9.0/requirements.txt +1 -0
  28. cow_backtester-0.9.0/setup.cfg +4 -0
  29. cow_backtester-0.9.0/tests/test_offline.py +802 -0
@@ -0,0 +1,136 @@
1
+ # Changelog
2
+
3
+ ## 0.9.0 — 2026-08-18
4
+
5
+ **`--reward-ev`: CIP-85 v2 consistency economics.** Capture ratios and ranks
6
+ measure competitiveness; solver income on most CoW chains is the consistency
7
+ pool. This release computes the actual v2 metric
8
+ (`Σ executed orders: your_best_fair_surplus / Σ all_solvers_surplus`) from
9
+ the same competition records `--compete` already fetches:
10
+
11
+ - the challenger's counterfactual metric, per-order surpluses inserted into
12
+ the historical denominators (conservatively — field terms keep their
13
+ historical values, so the reported share is a floor);
14
+ - every FIELD solver's real historical metric — a consistency leaderboard
15
+ for the window (also useful for spotting pool-dilution patterns);
16
+ - with `--self-address`, your actual historical metric side by side with
17
+ the replayed one;
18
+ - the win floor, reported explicitly: v2 pays ZERO on a chain where the
19
+ solver won nothing in the period, so the estimate never hides it;
20
+ - `--consistency-budget <COW>` converts the share into a COW/week estimate.
21
+
22
+ Honesty labels throughout: field surpluses come from record amounts (net of
23
+ protocol fees) while a replayed challenger's are gross (a few bps flattering,
24
+ labeled `basis`); challenger fairness is assumed (the field uses CoW's own
25
+ `filteredOut` flags); success_rate is a flag, not a silent multiplier.
26
+
27
+
28
+ ## 0.8.0 — 2026-08-17
29
+
30
+ **`--compete`: rank your solver against the historical field.** The v2
31
+ competition endpoint serves per-auction records by auction id (probed:
32
+ retention ≥ 2 months) carrying every submitted solution — solver, score,
33
+ ranking, CoW's own fairness-filtering outcome (`filteredOut`), per-solution
34
+ clearing prices, reference scores, and the auction's start/deadline blocks.
35
+
36
+ - `--compete` fetches each scored auction's record and inserts the
37
+ challenger's surplus into the fairness-surviving score list: per-auction
38
+ `field_rank` (rank, field size, gap-to-winner bps, winner solver), and a
39
+ per-solver summary — rank-1 %, top-3 %, median rank/gap, and a rivals
40
+ table (who beat you, how often, by how much). Honesty label everywhere:
41
+ `rank_basis: surplus_vs_score_proxy` — historical scores include protocol
42
+ fees, challenger surplus does not, so the rank is a floor.
43
+ - `--archive-dir DIR` persists every fetched record as
44
+ `DIR/<chain>/<auction_id>.json.gz` (idempotent) — a local competition
45
+ dataset that outlives the API's retention window and feeds future
46
+ fairness-simulation / reward-EV work.
47
+ - `--self-address 0x…` adds shadow-vs-actual: when your historical
48
+ solverAddress appears in a record, the row carries its actual score,
49
+ ranking, and fairness outcome next to the replayed one.
50
+ - Rows gain `auction_start_block` / `auction_deadline_block` — the auction
51
+ CUT context (what bidders saw), complementing `settlement_block`.
52
+
53
+ ## 0.7.2 — 2026-08-17
54
+
55
+ Scoring-fidelity batch from an independent line-level audit. Every item
56
+ below changes reported numbers or their labels — see README caveats.
57
+
58
+ - **Coverage-adjusted capture is the new headline.** Solver errors now keep
59
+ the historical winner's surplus in the denominator (previously an errored
60
+ auction vanished from the ratio — a solver could look better by failing
61
+ hard auctions). The old ratio is kept as `capture_conditional_pct`
62
+ ("when it answered, how competitive was it"), and a new
63
+ `lost_to_errors_wei` reports winner surplus forfeited to errors/timeouts.
64
+ - **Exact best-combination optimizer.** The greedy highest-first selection
65
+ of internally compatible solutions undercounted challengers (A=10 on two
66
+ pairs beats B+C=6+6 split — greedy picked 10, optimum is 12). Now solved
67
+ exactly (branch-and-bound; greedy fallback above 20 candidates, labeled
68
+ via `combiner`).
69
+ - **U256 range enforcement.** `to_u256` accepted arbitrarily large Python
70
+ ints; values above 2^256-1 are now rejected as invalid.
71
+ - **Scoring-basis labels.** Every row carries `baseline_quality`
72
+ (`exact_uniform` / `wrapper_lower_bound` / `mixed`); a same-basis
73
+ `capture_exact_basis_pct` (direct settlements only) is reported alongside
74
+ the mixed-basis proxy, and the HTML cards are relabeled accordingly
75
+ ("Winner fee take" → "Known direct-path fee wedge").
76
+ - **Per-transaction wrapper attribution.** UID-overlap proof is now enforced
77
+ per settlement transaction, not per auction group (one attributed tx can
78
+ no longer vouch for an unrelated one; drops count as `tx_uid_mismatch`).
79
+ - **A/B on byte-identical bodies.** One shared `deadline` per auction is
80
+ computed before the solver loop (was regenerated per request); rows carry
81
+ `replay_deadline`.
82
+ - **Readiness semantics.** `{"solutions": []}` counts as a healthy answer
83
+ (schema-legitimate abstention) with bid coverage reported separately;
84
+ "valid" wording clarified to auctions with ≥1 valid solution; READY now
85
+ requires `--min-evidence` attempted auctions (default 10) — below that the
86
+ best verdict is REVIEW. Readiness is now rendered in the HTML report.
87
+ - **`settlement_block`** replaces the row field `block` (kept one release as
88
+ a deprecated alias): it is the settlement's block, NOT the auction cut
89
+ block — README fork guidance corrected to match.
90
+ - Docs: `--max-auctions` keeps the **newest** auctions (behavior since
91
+ 0.7.1; help/README said oldest).
92
+
93
+ ## 0.7.1 — 2026-08-16
94
+
95
+ - Wrapper-routed settlements (solver router contracts, ~43% of mainnet /
96
+ ~40% of Base settlements) are attributed via the v2 by-tx-hash endpoint,
97
+ proven by UID overlap against the S3 auction body, and scored on a
98
+ clearly-labeled delivered basis (`entry: "wrapper"`). Previously excluded
99
+ silently.
100
+ - Submitter identification prefers the competition `solverAddress` over the
101
+ relay EOA. Coverage disclosure in the scorecard, `--readiness`, and a
102
+ `_meta` line in `--json-out`. `--max-auctions` keeps the newest auctions.
103
+
104
+ ## 0.7.0 — 2026-08-12
105
+
106
+ - `--readiness`: a one-screen pre-production readiness check for a solver
107
+ endpoint. Replays recent auctions against the endpoint and reports answer
108
+ rate, latency (p50/p95/max vs the solve budget), solution validity, and
109
+ surplus captured against the on-chain winners, with a READY / REVIEW /
110
+ NOT READY verdict and a copy-paste command to reproduce the exact run.
111
+ A focused alternative to the full field scorecard for anyone evaluating a
112
+ solver before staging or shadow. Emitted in the `--json-out`/`--html-out`
113
+ summary under `readiness`.
114
+
115
+ ## 0.6.0 — 2026-08-07
116
+
117
+ First public release.
118
+
119
+ - Replays archived CoW auctions (the public S3 instance bucket) against any
120
+ solver's `/solve` endpoint, offline.
121
+ - Reconstructs each auction's winning set from on-chain settlement calldata
122
+ and Trade events; groups multiple winning transactions per auction (CIP-67);
123
+ optional cross-check against the v2 competition API (`--verify-api`).
124
+ - Scores both sides on the same basis: before-fee surplus over signed limits
125
+ at uniform clearing prices, converted at the auction's reference prices.
126
+ - Validates solver responses the way the settlement layer would: per-order
127
+ fill accounting, fill-or-kill exactness, fee-adjusted limit feasibility,
128
+ strict U256 parsing; infeasible solutions are excluded and tallied.
129
+ - A/B mode: repeatable `--solver-url`, rotating call order, head-to-head
130
+ panel, per-pair breakdown, solve-latency stats.
131
+ - Ten chains, content cache, concurrent fetching, streamed JSONL rows,
132
+ single-file HTML report, watch mode.
133
+ - 37 offline tests against pinned fixtures; CI runs them plus lint and a
134
+ packaging check on Python 3.10 and 3.12.
135
+
136
+ Pre-release development history is summarized in `docs/DESIGN_NOTES.md`.
@@ -0,0 +1,25 @@
1
+ # Contributing
2
+
3
+ Corrections are the most valuable contribution — especially anywhere this
4
+ tool's accounting diverges from what the protocol actually does. Every scoring
5
+ rule cites its source (GPv2 contracts or `cowprotocol/services`); if you can
6
+ show a rule is wrong, please open an issue with the primary-source reference.
7
+
8
+ ## Development
9
+
10
+ ```bash
11
+ pip install -e ".[dev]"
12
+ python3 -m pytest # offline suite — no network, must always pass
13
+ python3 -m ruff check . # lint — must be clean
14
+ python3 -m cow_backtester.scorer --selftest && python3 -m cow_backtester --selftest # network tests
15
+ ```
16
+
17
+ Ground rules:
18
+ - The offline suite stays offline. New scoring behavior needs a fixture-based
19
+ test, and exploit-class regressions (duplicate trades, fee padding, negative
20
+ numerics, malformed shapes) must keep passing.
21
+ - No silent drops: anything excluded from scoring gets a tallied reason that
22
+ reaches the scorecard.
23
+ - No new runtime dependencies without strong cause (`eth_abi` is the only one).
24
+ - The pinned reference values (auction 8339027: event 2385773, surplus 1471,
25
+ fee 805) are load-bearing; never adjust them to make a test pass.
@@ -0,0 +1,21 @@
1
+ MIT License
2
+
3
+ Copyright (c) 2026 kaisersolver
4
+
5
+ Permission is hereby granted, free of charge, to any person obtaining a copy
6
+ of this software and associated documentation files (the "Software"), to deal
7
+ in the Software without restriction, including without limitation the rights
8
+ to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
9
+ copies of the Software, and to permit persons to whom the Software is
10
+ furnished to do so, subject to the following conditions:
11
+
12
+ The above copyright notice and this permission notice shall be included in all
13
+ copies or substantial portions of the Software.
14
+
15
+ THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
16
+ IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
17
+ FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
18
+ AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
19
+ LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
20
+ OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
21
+ SOFTWARE.
@@ -0,0 +1,5 @@
1
+ graft fixtures
2
+ graft tests
3
+ include mock_solver.py
4
+ include CHANGELOG.md CONTRIBUTING.md LICENSE requirements.txt
5
+ global-exclude __pycache__ *.pyc
@@ -0,0 +1,411 @@
1
+ Metadata-Version: 2.4
2
+ Name: cow-backtester
3
+ Version: 0.9.0
4
+ Summary: Offline backtester, A/B harness, and counterfactual scorecard for CoW Protocol solvers
5
+ Author: kaisersolver
6
+ License-Expression: MIT
7
+ Project-URL: Repository, https://github.com/kaisersolver/cow-backtester
8
+ Project-URL: Issues, https://github.com/kaisersolver/cow-backtester/issues
9
+ Project-URL: Changelog, https://github.com/kaisersolver/cow-backtester/blob/main/CHANGELOG.md
10
+ Keywords: cow-protocol,solver,backtesting,mev,defi
11
+ Classifier: Development Status :: 4 - Beta
12
+ Classifier: Environment :: Console
13
+ Classifier: Intended Audience :: Developers
14
+ Classifier: Programming Language :: Python :: 3.10
15
+ Classifier: Programming Language :: Python :: 3.11
16
+ Classifier: Programming Language :: Python :: 3.12
17
+ Classifier: Programming Language :: Python :: 3.13
18
+ Requires-Python: >=3.10
19
+ Description-Content-Type: text/markdown
20
+ License-File: LICENSE
21
+ Requires-Dist: eth_abi<6,>=5
22
+ Provides-Extra: dev
23
+ Requires-Dist: pytest>=7; extra == "dev"
24
+ Requires-Dist: ruff>=0.4; extra == "dev"
25
+ Dynamic: license-file
26
+
27
+ # cow-backtester
28
+
29
+ Offline backtester, A/B harness, and counterfactual scorecard for CoW Protocol
30
+ solvers.
31
+
32
+ Point it at your own solver's `/solve` endpoint and it replays recent CoW
33
+ auctions against it, then scores each of your solutions against the set of
34
+ solutions that actually won on-chain. Point it at two endpoints and it runs
35
+ a head-to-head A/B on identical auctions, so a routing or config change can
36
+ be judged before it goes to production.
37
+
38
+ Nothing touches production: no shadow mode, no staging deployment, no keys.
39
+
40
+ ## Why this exists
41
+
42
+ Today a solver can only be evaluated *live*: shadow mode consumes the
43
+ production auction stream, and the local playground runs against a chain fork.
44
+ Neither lets you take a fixed set of recent auctions, run your solver against
45
+ them offline, and ask "would I have out-surplused the winners, and by how
46
+ much?" the way you'd backtest a trading strategy. This tool does.
47
+
48
+ ## Quick start
49
+
50
+ Python 3.10+.
51
+
52
+ ```bash
53
+ pip install . # installs the `cow-backtester` command (one dependency: eth_abi)
54
+
55
+ # 1. Baseline — what the field actually captured (no solver needed)
56
+ cow-backtester --chain base --blocks 2000 --rpc-url <your-rpc>
57
+
58
+ # 2. Counterfactual — replay through your solver
59
+ cow-backtester --chain base --blocks 2000 --rpc-url <your-rpc> \
60
+ --solver-url http://localhost:8080 --json-out results.jsonl
61
+
62
+ # 3. A/B — two solvers, same auctions, head-to-head
63
+ cow-backtester --chain base --blocks 5000 --rpc-url <your-rpc> \
64
+ --solver-url http://localhost:8080 --solver-name baseline \
65
+ --solver-url http://localhost:8081 --solver-name candidate \
66
+ --html-out ab.html
67
+ ```
68
+
69
+ From a source checkout without installing, `python3 -m cow_backtester ...`
70
+ works identically. Your endpoint only needs the standard CoW solver-engine
71
+ API (`POST /solve`).
72
+
73
+ ## Readiness check
74
+
75
+ If you are bringing up a new solver and want a fast "is this endpoint healthy
76
+ enough to face production auctions?" read, `--readiness` prints a one-screen
77
+ report instead of the full field scorecard:
78
+
79
+ ```bash
80
+ cow-backtester --chain base --blocks 2000 --rpc-url <your-rpc> \
81
+ --solver-url http://localhost:8080 --solver-name mine --readiness
82
+ ```
83
+
84
+ It replays recent auctions against your endpoint and reports the four things
85
+ that gate a pre-prod solver — does it answer, is it fast enough, are its
86
+ solutions valid, and are they competitive with the on-chain winners — as
87
+ pass/warn checks with a `READY` / `REVIEW` / `NOT READY` verdict:
88
+
89
+ ```
90
+ ====================================================================
91
+ READINESS — mine [REVIEW]
92
+ base · prod · blocks 49506127..49510000
93
+ ====================================================================
94
+ [PASS] reached auctions 50 auctions attempted
95
+ [WARN] no transport errors 2/50 errored
96
+ [PASS] answers reliably 96% returned a solution
97
+ [PASS] inside the deadline 0 past deadline
98
+ [PASS] latency headroom p95 420 ms of a 15000 ms budget
99
+ [PASS] solutions are valid 98% passed limit/fee checks
100
+ [PASS] competitive vs winners 61% of winner surplus captured
101
+ ```
102
+
103
+ It prints the exact `--from-block/--to-block` command to reproduce the run,
104
+ and the same data lands in `--json-out`/`--html-out` under `readiness`. Works
105
+ on any of the supported chains, so you can readiness-check an endpoint for a
106
+ chain you are not yet onboarded on. It is a signal, not a settlement
107
+ guarantee — pair it with a self-hosted shadow run before going to production.
108
+
109
+ ## Consistency economics (`--reward-ev`)
110
+
111
+ On most CoW chains the money is not in winning — it is in the CIP-85
112
+ consistency pool, which pays
113
+ `success_rate × Σ executed orders ( your best fair bid's surplus / everyone's )`.
114
+ `--reward-ev` computes exactly that metric from the competition records:
115
+ your solver's counterfactual metric and pool share, every field solver's
116
+ real historical metric (a consistency leaderboard for the window), and,
117
+ with `--consistency-budget <COW>`, a COW/week estimate. The win floor is
118
+ reported explicitly — v2 pays zero on a chain where a solver won nothing in
119
+ the period — and every number carries its basis label (field surpluses are
120
+ net of protocol fees, a replayed challenger's are gross; challenger
121
+ fairness is assumed while the field uses CoW's own `filteredOut` flags).
122
+
123
+ ## Rank against the historical field (`--compete`)
124
+
125
+ Capture ratio tells you how much surplus you generate; `--compete` tells you
126
+ **where you would have ranked**. For every scored auction it fetches the
127
+ historical competition record (every submitted solution with its score,
128
+ CoW's own fairness-filtering outcome, and the winner) and inserts your
129
+ solver's result into the fairness-surviving score list:
130
+
131
+ ```bash
132
+ cow-backtester --chain base --blocks 2000 --rpc-url <your-rpc> \
133
+ --solver-url http://localhost:8080 --solver-name mine --compete \
134
+ --archive-dir ./competition-data
135
+ ```
136
+
137
+ The scorecard gains a field-rank line (rank-1 %, top-3 %, median rank,
138
+ median gap to the winner in bps) and a rivals table — which solvers beat
139
+ you, how often, and by how much. Per-auction `field_rank` lands in the JSON
140
+ rows. Two honesty notes: historical scores include protocol fees while your
141
+ replayed surplus does not, so the reported rank is a **floor** (labeled
142
+ `surplus_vs_score_proxy`); and records exist only for auctions that had a
143
+ winner. `--archive-dir` keeps every fetched record on disk — the API serves
144
+ roughly two months of history, so an archive you build today is a dataset
145
+ you keep. `--self-address` marks your historical solverAddress so rows show
146
+ shadow-vs-actual side by side.
147
+
148
+ ### Try it without a solver
149
+
150
+ The bundled mock solver lets you see the full counterfactual/A/B output in
151
+ about a minute, before wiring up your own engine:
152
+
153
+ ```bash
154
+ MODE=limit python3 mock_solver.py 8901 & # fills exactly at the limit
155
+ MODE=better python3 mock_solver.py 8902 & # fills at 1.5x the limit
156
+
157
+ cow-backtester --chain base --blocks 1500 --rpc-url <your-rpc> \
158
+ --solver-url http://127.0.0.1:8901 --solver-name at-limit \
159
+ --solver-url http://127.0.0.1:8902 --solver-name better --html-out demo.html
160
+ kill %1 %2
161
+ ```
162
+
163
+ You get the per-solver panels, the head-to-head table, and the pair
164
+ breakdown. The `implausible_surplus` guard usually fires on the "better"
165
+ mock: claiming 1.5x the limit on real order flow is exactly what the
166
+ validator exists to flag.
167
+
168
+ ## How it works (all public data)
169
+
170
+ 1. Input: the exact `/solve` bodies CoW sent solvers, from the public S3
171
+ instance bucket (`solver-instances.s3.amazonaws.com/<env>/<chain>/auction/<id>.json`).
172
+ The archived file *is* a valid `/solve` request body.
173
+ 2. Winner reconstruction: from the on-chain settlements. (The v1
174
+ `/solver_competition` endpoints were removed in the competition-data
175
+ migration; v2 exists but is retention-bounded. Reconstructing from chain
176
+ data is trustless and independent of API retention — and `--verify-api`
177
+ cross-checks it against v2 where available.) Settlements are grouped by
178
+ auction: since combinatorial auctions (CIP-67), one auction can settle
179
+ through several winning transactions, so the baseline is the combined
180
+ winning set.
181
+ 3. Replay and score: each auction body is POSTed to each solver; responses
182
+ are validated (below) and scored with the same function, the same
183
+ limits, and the same price convention as the winners. Your valid solutions
184
+ are combined the CIP-67 way (best-first, disjoint directed token pairs).
185
+
186
+ ## The comparison basis (read this before quoting numbers)
187
+
188
+ Both sides are scored on **before-fee surplus over the signed order limits, at
189
+ uniform clearing prices**, converted to the chain's native token at the
190
+ auction's own reference prices.
191
+
192
+ * Signed limits: CoW's accounting scores against the order's signed amounts
193
+ (`fullSellAmount`/`fullBuyAmount`), not the remaining/fee-adjusted amounts.
194
+ * Uniform prices: on-chain settlements carry per-trade post-fee "custom"
195
+ prices; the first occurrence of each token in the price vector is the uniform
196
+ (pre-fee) price, which is what autopilot itself resolves. Scoring the winner
197
+ post-fee but a challenger pre-fee would bias every comparison.
198
+ * Surplus token: buy token for sell orders, sell token for buy orders,
199
+ converted at that token's `referencePrice` (matches official accounting).
200
+ * What the number is: user surplus + protocol fees + network fee. CoW's
201
+ official ranking score is user surplus + protocol fees only, so absolute
202
+ levels here overstate the official score by the network fee (on the pinned
203
+ reference trade, a $2.4 fill where the network fee dominates: 1471 vs an
204
+ official 666 atoms; the gap shrinks as trades grow). When two solutions carry
205
+ very different fees this basis can order them differently than CoW would;
206
+ gas-heavy solutions are flattered. It is the only basis we found that can
207
+ be computed symmetrically offline for both sides.
208
+
209
+ ## Response validation (your solver can't accidentally cheat)
210
+
211
+ A solution is scored only if **every** fulfillment trade is feasible:
212
+
213
+ * known order uid (case-insensitive) and strict U256 numerics (decimal or
214
+ 0x-hex; negatives and malformed values rejected);
215
+ * prices present for both tokens;
216
+ * at most one fulfillment per order (GPv2 accumulates fills and the driver
217
+ rejects duplicate trades, so N copies of a trade can't score N times the
218
+ surplus);
219
+ * no over-fill (`executed + fee` vs `sellAmount` for sell orders, `executed`
220
+ vs `buyAmount` for buy orders; exact for fill-or-kill);
221
+ * on-chain feasibility at the fee-adjusted terms: the net delivery must still
222
+ cover the gross-scaled limit, which is what GPv2 enforces, so a padded `fee`
223
+ cannot manufacture surplus.
224
+
225
+ Violations mark the whole **solution INVALID** (tallied with a reason), never a
226
+ silent zero. Solutions scoring >10× the winning set (or >0.001 native when the
227
+ winning set scored 0) are flagged `implausible_surplus`. Malformed responses
228
+ are tallied, never crash a run. **Prices are claimed, not simulated** — the
229
+ tool checks feasibility; it does not execute routes.
230
+
231
+ ## What a run prints
232
+
233
+ * COVERAGE: settlements found, auctions formed/scored, every skip with its
234
+ reason, unscanned block spans, auction ages, cache stats, `--verify-api` results.
235
+ * FIELD SURPLUS: winner surplus by USD size bucket, plus the winners' fee take.
236
+ * WINNING SUBMITTERS: who is actually winning (label them with `--solver-map`).
237
+ * COUNTERFACTUAL (per solver): returned / valid / positive / beat the
238
+ winning set, surplus sums, capture ratio, solve latency p50+p95, invalid
239
+ reasons, error classes.
240
+ * HEAD TO HEAD (exactly two solvers): side-by-side metrics, per-auction
241
+ win counts, and the surplus delta.
242
+ * TOP PAIRS: winner surplus by token pair with each solver's surplus
243
+ beside it: *where* you win and lose, not just by how much.
244
+
245
+ A real single-solver run (Arbitrum, eight auctions, trimmed):
246
+
247
+ ```
248
+ ====================================================================
249
+ COUNTERFACTUAL — kaisersolver (8 replayed)
250
+ ====================================================================
251
+ returned a solution : 7/8
252
+ valid solutions : 7/8
253
+ positive surplus : 7/8
254
+ beat the winning set: 0/8
255
+ our surplus (sum) : 0.001643 ETH
256
+ winners (sum) : 0.001748 ETH (replayed auctions only)
257
+ capture ratio : 94.0% of the winning set
258
+ solve latency : p50 1174 ms / p95 2471 ms
259
+
260
+ ====================================================================
261
+ TOP PAIRS BY WINNER SURPLUS (top 2)
262
+ ====================================================================
263
+ pair trades winner kaisersolver
264
+ WETH->MOR 7 0.001722 0.001643
265
+ USDC->USD₮0 1 0.000026 0.000000
266
+ ```
267
+
268
+ Eight auctions is a small sample, but the shape is the point: this solver is
269
+ consistently a few percent behind one competitor on one pair, and the table
270
+ says which pair. That took minutes to learn here; it took weeks from logs.
271
+
272
+ `--json-out` streams one row per auction (block, timestamp, age, winner txs +
273
+ submitters + surplus/fees, per-solver validation detail and latency,
274
+ `expired_orders_pct`, flags). `--html-out` writes a single self-contained HTML
275
+ report: no external assets, light/dark aware, fine to attach to a PR or post.
276
+
277
+ ## Flags
278
+
279
+ | flag | meaning |
280
+ |---|---|
281
+ | `--chain` | `mainnet`, `arbitrum-one`, `base`, `xdai`, `polygon`, `bnb`, `avalanche`, `linea`, `ink`, `plasma` |
282
+ | `--env` | `prod` (default) or `staging` |
283
+ | `--blocks N` | scan the most recent N blocks |
284
+ | `--from-block/--to-block` | absolute, reproducible window |
285
+ | `--max-auctions N` | cap auctions scored, newest first (default 25; 0 = all) |
286
+ | `--rpc-url` | your RPC (recommended — public defaults rot and rate-limit) |
287
+ | `--solver-url` / `--solver-name` | repeatable; two of them = A/B |
288
+ | `--solve-timeout S` | deadline advertised to the solver (HTTP waits S+5) |
289
+ | `--workers N` | concurrent RPC/S3 fetches (default 8) |
290
+ | `--cache-dir` / `--no-cache` | content cache (default `.cowbt-cache`) |
291
+ | `--json-out` / `--html-out` | machine-readable rows / HTML report |
292
+ | `--solver-map FILE` | JSON `{address: name}` to label winning submitters |
293
+ | `--verify-api` | cross-check winner txs against the v2 competition API |
294
+ | `--compete` | rank the challenger against the historical fairness-surviving field (rank / gap / rivals) |
295
+ | `--archive-dir DIR` | persist competition records as `DIR/<chain>/<id>.json.gz` (local dataset) |
296
+ | `--self-address 0x…` | your historical solverAddress → shadow-vs-actual comparison |
297
+ | `--reward-ev` | CIP-85 v2 consistency economics: counterfactual metric, field leaderboard, COW estimate (implies `--compete`) |
298
+ | `--consistency-budget N` | the chain's weekly consistency pool in COW, to convert share → COW/week |
299
+ | `--min-evidence N` | attempted-auction floor before `--readiness` may say READY (default 10) |
300
+ | `--clamp-validto` | extend expired `validTo` so engines that filter them still solve |
301
+ | `--max-age-hours H` | warn when replayed auctions are older than this |
302
+ | `--watch N` | continuous mode: rescan every N seconds from the last block |
303
+ | `--quiet` | suppress progress (stderr); the scorecard still prints to stdout |
304
+ | `--version` | print version |
305
+
306
+ Progress goes to **stderr**, the scorecard to **stdout** (`2>/dev/null` gives
307
+ clean results; `--json-out`/`--html-out` are unaffected). Exit codes: **0**
308
+ success, **1** runtime error, **2** usage error, **130** interrupted mid-run
309
+ (a partial scorecard is printed when auctions had already been processed;
310
+ stopping `--watch` during its idle sleep is a clean stop and exits 0).
311
+
312
+ ## Caching
313
+
314
+ Only **immutable** facts are cached: archived auction bodies, settlements below
315
+ the reorg margin, and block timestamps. Nothing derived from live liquidity is
316
+ ever cached, so a warm re-run is faster (3.8x on a measured six-auction,
317
+ two-solver run) without changing a single number. Delete `.cowbt-cache` any time, or pass `--no-cache`.
318
+
319
+ ## Limitations
320
+
321
+ * **Replay uses live liquidity.** Your solver quotes against *current* chain
322
+ state, not the historical block. The counterfactual is **indicative** — it
323
+ answers "how does my solver handle this real order flow", not "the exact
324
+ outcome at that block". Ages are printed and a warning fires beyond
325
+ `--max-age-hours` (default 6h). Prefer recent, short windows. Each row
326
+ carries `settlement_block` — the block the winning settlement landed in,
327
+ which is *later* than the auction cut block the bidders actually saw. A
328
+ fork pinned there is an approximation, not a faithful auction replay (it
329
+ can even include the settlement itself); true historical-fork replay needs
330
+ the auction cut block, which this tool does not yet reconstruct.
331
+ * **The S3 bucket retains roughly one month** of auctions (measured Aug 2026).
332
+ This is a recent-window backtester, not an archive.
333
+ * **Entrypoint coverage.** Direct `settle()` calls are attributed by the
334
+ auction id appended to their calldata. Wrapper-routed settlements (solver
335
+ router contracts — measured Aug 2026 at **~43% of mainnet, ~40% of Base,
336
+ ~2% of Arbitrum** settlements, and *larger* than direct ones at the median,
337
+ so they are not a random slice) are attributed via the v2
338
+ `solver_competition/by_tx_hash` endpoint and **proven by uid overlap with
339
+ the S3 auction body** before scoring; they score on a DELIVERED basis
340
+ (Trade events vs signed limits), which understates the direct-path
341
+ before-fee basis by the settlement's fee wedge (typically a few bps) —
342
+ rows carry `entry: "wrapper"` so the bases are distinguishable. Settlements
343
+ the endpoint cannot resolve are counted under `wrapper_unattributed` and
344
+ disclosed in the coverage block, `--readiness`, and the `_meta` JSON line.
345
+ * **"Beat the winning set" is necessary, not sufficient.** Real winner
346
+ selection also applies fairness filters, and bids score net of gas; the tool
347
+ also takes your solutions at face value while the driver merges and simulates.
348
+ * **Orders keep their historical `validTo`** (~1%/day expire; more near the
349
+ retention edge). Engines that filter expired orders look age-degraded — see
350
+ `expired_orders_pct` per row, or pass `--clamp-validto`.
351
+ * **JIT orders.** The winner baseline credits surplus-capturing JIT trades
352
+ (owners in `surplusCapturingJitOrderOwners`), per official accounting; the
353
+ challenger side never credits JIT (a response's JIT order has no verifiable
354
+ owner). On auctions where winner surplus comes from CoW-AMM-style JIT, the
355
+ comparison is conservative *against* the challenger, never in its favor.
356
+ * **Bring your own RPC.** Public defaults rot and rate-limit `eth_getLogs`; the
357
+ tool splits and retries failing ranges and always **reports** any span it
358
+ could not scan rather than under-counting silently. `eth_chainId` is checked
359
+ against `--chain` on startup.
360
+
361
+ ## Provenance of the reconstruction
362
+
363
+ * Auction id: CoW's driver appends it to `settle()` calldata; autopilot reads
364
+ back exactly the **last 8 bytes** (`META_DATA_LEN = 8` in
365
+ `cowprotocol/services`). The tool applies the same rule, only when the
366
+ calldata prefix re-encodes canonically, cross-checked by requiring **at least
367
+ one** settled order uid to appear in the fetched body (`body_uid_mismatch`
368
+ guard — this also catches staging/prod contamination, which shares the
369
+ settlement contract).
370
+ * Trade events are filtered by **emitting address + topic** and must align 1:1
371
+ with calldata trades, else the settlement is skipped (`trade_event_mismatch`).
372
+ * Fill-or-kill: GPv2 ignores the calldata `executedAmount` and uses the signed
373
+ amount — the tool applies the same substitution.
374
+ * Response-schema authority: `crates/solvers-dto` / the solver-engine OpenAPI
375
+ in `cowprotocol/services`.
376
+ * `--verify-api` compares the reconstructed winning-tx set per auction against
377
+ the v2 competition endpoint (match / mismatch / unavailable counters).
378
+
379
+ ## Tests
380
+
381
+ ```bash
382
+ pip install -e ".[dev]" # pytest + ruff (one-time)
383
+ python3 -m pytest # OFFLINE suite (fixtures; no network)
384
+ python3 -m cow_backtester.scorer --unittest # offline core checks
385
+ python3 -m cow_backtester.scorer --selftest # network: winner path on the reference settlement
386
+ python3 -m cow_backtester --selftest # network: counterfactual == winner baseline +
387
+ # every exploit class as regression cases
388
+ ```
389
+
390
+ The pinned reference (Arbitrum auction 8339027): decode integrity is exact
391
+ against the GPv2 `Trade` event (2385773 atoms), before-fee surplus 1471 atoms,
392
+ fee take 805. `mock_solver.py` (modes: empty / limit / better / hex / garbage /
393
+ invalid) exercises the wire path including validation and implausibility
394
+ guards. CI runs the offline suite on Python 3.10 and 3.12.
395
+
396
+ ## Files
397
+
398
+ | file | role |
399
+ |---|---|
400
+ | `cow_backtester/scorer.py` | settlement decode: auction id + winner surplus |
401
+ | `cow_backtester/backtest.py` | enumerate, group, replay, validate, scorecard |
402
+ | `cow_backtester/cache.py` | immutable-fact content cache |
403
+ | `cow_backtester/report.py` | single-file HTML report |
404
+ | `mock_solver.py` | test/demo solver (six modes) |
405
+ | `fixtures/` | pinned reference data for the offline tests |
406
+ | `tests/` | offline pytest suite, no network |
407
+ | `docs/DESIGN_NOTES.md` | design decisions and validation history |
408
+
409
+ MIT licensed. Contributions and corrections welcome — especially from CoW core
410
+ devs on anything where this tool's accounting diverges from the protocol's.
411
+ </content>