pitbacktest 0.2.0__tar.gz
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- pitbacktest-0.2.0/LICENSE +21 -0
- pitbacktest-0.2.0/PKG-INFO +395 -0
- pitbacktest-0.2.0/README.md +355 -0
- pitbacktest-0.2.0/pitbacktest/__init__.py +27 -0
- pitbacktest-0.2.0/pitbacktest/adapters/__init__.py +0 -0
- pitbacktest-0.2.0/pitbacktest/adapters/krx.py +249 -0
- pitbacktest-0.2.0/pitbacktest/adapters/long_format.py +87 -0
- pitbacktest-0.2.0/pitbacktest/adapters/tiingo.py +236 -0
- pitbacktest-0.2.0/pitbacktest/adapters/yfinance.py +243 -0
- pitbacktest-0.2.0/pitbacktest/analytics.py +234 -0
- pitbacktest-0.2.0/pitbacktest/core/__init__.py +0 -0
- pitbacktest-0.2.0/pitbacktest/core/controls.py +126 -0
- pitbacktest-0.2.0/pitbacktest/core/costs.py +144 -0
- pitbacktest-0.2.0/pitbacktest/core/estimators.py +216 -0
- pitbacktest-0.2.0/pitbacktest/core/gates.py +175 -0
- pitbacktest-0.2.0/pitbacktest/core/panel.py +306 -0
- pitbacktest-0.2.0/pitbacktest/crypto/__init__.py +12 -0
- pitbacktest-0.2.0/pitbacktest/crypto/binance_archive.py +287 -0
- pitbacktest-0.2.0/pitbacktest/crypto/costs.py +80 -0
- pitbacktest-0.2.0/pitbacktest/crypto/intraday.py +326 -0
- pitbacktest-0.2.0/pitbacktest/crypto/panel.py +120 -0
- pitbacktest-0.2.0/pitbacktest/equity/__init__.py +10 -0
- pitbacktest-0.2.0/pitbacktest/equity/master.py +60 -0
- pitbacktest-0.2.0/pitbacktest/equity/scenarios.py +104 -0
- pitbacktest-0.2.0/pitbacktest/event.py +206 -0
- pitbacktest-0.2.0/pitbacktest/execution.py +310 -0
- pitbacktest-0.2.0/pitbacktest/ledger.py +128 -0
- pitbacktest-0.2.0/pitbacktest/portfolio.py +373 -0
- pitbacktest-0.2.0/pitbacktest/screen.py +150 -0
- pitbacktest-0.2.0/pitbacktest/shorting.py +46 -0
- pitbacktest-0.2.0/pitbacktest/validation.py +118 -0
- pitbacktest-0.2.0/pitbacktest/weights.py +271 -0
- pitbacktest-0.2.0/pitbacktest.egg-info/PKG-INFO +395 -0
- pitbacktest-0.2.0/pitbacktest.egg-info/SOURCES.txt +60 -0
- pitbacktest-0.2.0/pitbacktest.egg-info/dependency_links.txt +1 -0
- pitbacktest-0.2.0/pitbacktest.egg-info/requires.txt +18 -0
- pitbacktest-0.2.0/pitbacktest.egg-info/top_level.txt +1 -0
- pitbacktest-0.2.0/pyproject.toml +41 -0
- pitbacktest-0.2.0/setup.cfg +4 -0
- pitbacktest-0.2.0/tests/test_analytics.py +285 -0
- pitbacktest-0.2.0/tests/test_annualization.py +39 -0
- pitbacktest-0.2.0/tests/test_argument_checks.py +122 -0
- pitbacktest-0.2.0/tests/test_binance_archive.py +219 -0
- pitbacktest-0.2.0/tests/test_causality.py +113 -0
- pitbacktest-0.2.0/tests/test_costs_events.py +332 -0
- pitbacktest-0.2.0/tests/test_crypto.py +184 -0
- pitbacktest-0.2.0/tests/test_delisting_and_guards.py +118 -0
- pitbacktest-0.2.0/tests/test_docs_scripts.py +93 -0
- pitbacktest-0.2.0/tests/test_equity.py +117 -0
- pitbacktest-0.2.0/tests/test_execution.py +524 -0
- pitbacktest-0.2.0/tests/test_intraday.py +380 -0
- pitbacktest-0.2.0/tests/test_krx.py +187 -0
- pitbacktest-0.2.0/tests/test_ledger.py +100 -0
- pitbacktest-0.2.0/tests/test_packaging.py +51 -0
- pitbacktest-0.2.0/tests/test_reconcile.py +230 -0
- pitbacktest-0.2.0/tests/test_side_costs_shorting.py +288 -0
- pitbacktest-0.2.0/tests/test_symbol_rules.py +120 -0
- pitbacktest-0.2.0/tests/test_synthetic.py +156 -0
- pitbacktest-0.2.0/tests/test_tiingo.py +256 -0
- pitbacktest-0.2.0/tests/test_validation.py +65 -0
- pitbacktest-0.2.0/tests/test_weights.py +225 -0
- pitbacktest-0.2.0/tests/test_yfinance_adapter.py +147 -0
|
@@ -0,0 +1,21 @@
|
|
|
1
|
+
MIT License
|
|
2
|
+
|
|
3
|
+
Copyright (c) 2026 Janghyuk Choi
|
|
4
|
+
|
|
5
|
+
Permission is hereby granted, free of charge, to any person obtaining a copy
|
|
6
|
+
of this software and associated documentation files (the "Software"), to deal
|
|
7
|
+
in the Software without restriction, including without limitation the rights
|
|
8
|
+
to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
|
|
9
|
+
copies of the Software, and to permit persons to whom the Software is
|
|
10
|
+
furnished to do so, subject to the following conditions:
|
|
11
|
+
|
|
12
|
+
The above copyright notice and this permission notice shall be included in all
|
|
13
|
+
copies or substantial portions of the Software.
|
|
14
|
+
|
|
15
|
+
THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
|
|
16
|
+
IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
|
|
17
|
+
FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
|
|
18
|
+
AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
|
|
19
|
+
LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
|
|
20
|
+
OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
|
|
21
|
+
SOFTWARE.
|
|
@@ -0,0 +1,395 @@
|
|
|
1
|
+
Metadata-Version: 2.4
|
|
2
|
+
Name: pitbacktest
|
|
3
|
+
Version: 0.2.0
|
|
4
|
+
Summary: Backtest harness for factor screening, event signals and portfolio alphas, built to make common backtest errors hard
|
|
5
|
+
Author: Janghyuk Choi
|
|
6
|
+
License-Expression: MIT
|
|
7
|
+
Project-URL: Repository, https://github.com/JanghyukChoi/quantbacktest
|
|
8
|
+
Project-URL: Issues, https://github.com/JanghyukChoi/quantbacktest/issues
|
|
9
|
+
Keywords: backtest,quant,survivorship-bias,overfitting,point-in-time,deflated-sharpe,factor
|
|
10
|
+
Classifier: Development Status :: 3 - Alpha
|
|
11
|
+
Classifier: Intended Audience :: Financial and Insurance Industry
|
|
12
|
+
Classifier: Intended Audience :: Science/Research
|
|
13
|
+
Classifier: Programming Language :: Python :: 3
|
|
14
|
+
Classifier: Programming Language :: Python :: 3.10
|
|
15
|
+
Classifier: Programming Language :: Python :: 3.11
|
|
16
|
+
Classifier: Programming Language :: Python :: 3.12
|
|
17
|
+
Classifier: Programming Language :: Python :: 3.13
|
|
18
|
+
Classifier: Programming Language :: Python :: 3.14
|
|
19
|
+
Classifier: Topic :: Office/Business :: Financial :: Investment
|
|
20
|
+
Classifier: Topic :: Scientific/Engineering
|
|
21
|
+
Requires-Python: >=3.10
|
|
22
|
+
Description-Content-Type: text/markdown
|
|
23
|
+
License-File: LICENSE
|
|
24
|
+
Requires-Dist: pandas>=2.0
|
|
25
|
+
Requires-Dist: numpy>=1.24
|
|
26
|
+
Provides-Extra: yfinance
|
|
27
|
+
Requires-Dist: yfinance; extra == "yfinance"
|
|
28
|
+
Provides-Extra: test
|
|
29
|
+
Requires-Dist: pytest>=7.0; extra == "test"
|
|
30
|
+
Requires-Dist: statsmodels; extra == "test"
|
|
31
|
+
Requires-Dist: scipy; extra == "test"
|
|
32
|
+
Provides-Extra: dev
|
|
33
|
+
Requires-Dist: pytest>=7.0; extra == "dev"
|
|
34
|
+
Requires-Dist: statsmodels; extra == "dev"
|
|
35
|
+
Requires-Dist: scipy; extra == "dev"
|
|
36
|
+
Requires-Dist: ruff; extra == "dev"
|
|
37
|
+
Requires-Dist: build; extra == "dev"
|
|
38
|
+
Requires-Dist: twine; extra == "dev"
|
|
39
|
+
Dynamic: license-file
|
|
40
|
+
|
|
41
|
+
# quantbacktest
|
|
42
|
+
|
|
43
|
+
[](https://github.com/JanghyukChoi/quantbacktest/actions/workflows/ci.yml)
|
|
44
|
+
|
|
45
|
+
The Python package is named `pitbacktest` (`import pitbacktest`); the repository is `quantbacktest`.
|
|
46
|
+
|
|
47
|
+
**A backtest toolkit that measures its own biases.** Most backtest libraries compute a return and leave you to wonder how much
|
|
48
|
+
of it is survivorship, look-ahead, a flattering cost assumption, or luck from trying many variants. Here each of those is either a
|
|
49
|
+
measured number or a refused input, for US stocks, Korean stocks and crypto perpetuals. Core dependencies are `pandas` and `numpy`.
|
|
50
|
+
It is not a strategy and ships no data.
|
|
51
|
+
|
|
52
|
+
What is different:
|
|
53
|
+
|
|
54
|
+
- **Survivorship is measured, not assumed.** A free yfinance download returned a correct history for 0 of 41 well-known delisted or
|
|
55
|
+
acquired US stocks; a free Tiingo key returned 23 of 29 takeover and rename cases but **none of 12 bankruptcies and rescue sales**
|
|
56
|
+
(`docs/survivorship.md`). On Korean stocks, where the official KRX data is survivorship-free, restricting to today's survivors moved
|
|
57
|
+
the net Sharpe of four factors by -0.23 to +0.20 depending on the factor, averaging about zero, and the understatement of low
|
|
58
|
+
volatility survived neutralisation against the other styles (`studies/korea_survivorship`).
|
|
59
|
+
- **Preregistered studies, with their corrections kept.** Rules are committed before results; when a cost model turned out wrong, the
|
|
60
|
+
original results stayed, an amendment explained it, and both are reported.
|
|
61
|
+
- **Overfitting is counted, not recalled.** Deflated Sharpe, PBO and a permutation test, plus a hash-chained trial ledger that records
|
|
62
|
+
every run so the number of variants you tried cannot be understated from memory.
|
|
63
|
+
- **One timing convention**, point-in-time universes, and firm-characteristic controls by default: the test statistic that decides `screen`'s first
|
|
64
|
+
gate is computed with controls, and the factor is re-tested after neutralising it, with the uncontrolled figures labelled beside them.
|
|
65
|
+
- **Costs and capacity:** spread, borrow and square-root market impact for weights you supply, a capacity curve, and an intraday layer
|
|
66
|
+
(1-minute Binance bars) that shows how fast an edge decays with latency and at what cost it stops paying.
|
|
67
|
+
- **Checked against independent implementations**, not just against itself: the portfolio engine equals a plain-loop reimplementation to
|
|
68
|
+
1e-17, alpha/beta equal statsmodels' Newey-West regression and IC equals scipy's Spearman, the intraday parser reproduces Binance's own
|
|
69
|
+
daily files, and over 40 deliberately planted bugs are each caught by the tests.
|
|
70
|
+
|
|
71
|
+
What it cannot do, stated up front: it has no order book (so no queue, partial fills or impact below a bar), it does not ship an
|
|
72
|
+
optimiser or point-in-time fundamentals, and only Linux is tested. See **Limits**.
|
|
73
|
+
|
|
74
|
+
## Markets
|
|
75
|
+
|
|
76
|
+
The engine works on any date x security panel. What differs between markets is the **data path** and whether it is
|
|
77
|
+
free of survivorship bias.
|
|
78
|
+
|
|
79
|
+
| Market | Data path | Survivorship-free? | Status |
|
|
80
|
+
|---|---|---|---|
|
|
81
|
+
| Crypto perpetuals (Binance) | `pitbacktest.crypto`, public archive | **Yes**: 900 contracts ever listed, 376 of them delisted or halted. Funding charged, delistings explicit | Tested; preregistered study in `studies/crypto_cross_section` |
|
|
82
|
+
| Korean stocks | `pitbacktest.adapters.krx`, official KRX OpenAPI (free key) | **Yes**: the API returns every stock listed on each day, so later delistings are inside the history. Adjusted returns come from the change versus the reference price, no price-adjustment table needed | Return formula checked on live data (all 953 KOSPI names on 2024-01-02, Samsung's 50:1 split day); panel build tested offline. Measured on 13 years of data: the survivors-only shortcut moves a factor's Sharpe by up to 0.23 in either direction, averaging about zero (`studies/korea_survivorship`) |
|
|
83
|
+
| US stocks | `adapters.yfinance`, or your own point-in-time data through `adapters.long_format` | **Not with yfinance**: it returned a correct history for 0 of 41 well-known delisted or acquired stocks (`docs/survivorship.md`). Yes if you bring CRSP, Sharadar or Norgate data | Detector, coverage report and delisting scenarios. With a free Tiingo key a random 466-ticker sample gives a **lower bound**: the survivors-only difference is indistinguishable from zero (intervals about 0.29 Sharpe wide), which is not evidence of no bias, because free data holds few bankruptcies (`studies/us_survivorship`) |
|
|
84
|
+
|
|
85
|
+
## Three entry points
|
|
86
|
+
|
|
87
|
+
| Function | Question it answers |
|
|
88
|
+
|---|---|
|
|
89
|
+
| `screen(panel, factors)` | Does a factor survive costs, sub-periods, monotonicity and neutralisation? (research stage) |
|
|
90
|
+
| `backtest_portfolio(panel, factor)` | What does the factor do as a long-short portfolio? CAGR, Sharpe, max drawdown, turnover |
|
|
91
|
+
| `backtest_event(panel, signal)` | Does a boolean event signal work as an individual alert? Win rate vs base rate, mean vs median |
|
|
92
|
+
|
|
93
|
+
## What it enforces
|
|
94
|
+
|
|
95
|
+
| Rule | Why |
|
|
96
|
+
|---|---|
|
|
97
|
+
| One timing convention, checked by `assert_timing()` | An entry that is one day late (or early) silently changes results, most of all for 1-day mean reversion |
|
|
98
|
+
| `panel.assert_causal(make_signal)` rebuilds a signal from the panel cut at random dates and requires the last row to match the full-data row | The exact test for look-ahead in how a signal is computed: a negative shift, a centred window, a mean or z-score over the whole history all fail it, a past-only signal always passes. The older heuristic `assert_no_lookahead` has blind spots (it passes a signal that is the future return itself); do not rely on it |
|
|
99
|
+
| Firm-characteristic controls are the default (size, book-to-market, momentum, ROA, asset growth) | Without them a "new alpha" is often a known factor in disguise. It warns when characteristics are missing |
|
|
100
|
+
| `screen` decides on a t-statistic computed **with controls** and re-tests after neutralising; uncontrolled figures (`excess_bp`, `coef_bp_raw`) sit beside them, labelled | A large raw number should not be the thing that passes a factor |
|
|
101
|
+
| Win rate is reported with its base rate: `lift = win rate - base rate` | With a longer holding period both rise together; only the lift says anything about the signal |
|
|
102
|
+
| Mean and median are both reported (`skew_warning`) | If the signs differ, a few big winners hide many small losses. Fine for a portfolio, bad for an alert |
|
|
103
|
+
| Multiple-testing thresholds come from a shuffled null, not Bonferroni | Bonferroni ignores the correlation between tests |
|
|
104
|
+
| Parameter grids return the whole distribution | The user can see whether the best point is a plateau or a spike |
|
|
105
|
+
| `backtest_portfolio(..., ledger=Ledger(dir), family=..., name=...)` records every run; `ledger.deflated_sharpe(family)` takes the trial count from the record | People under-report how many variants they tried. A rerun of the same configuration counts once; any change is a new trial. The file is hash-chained, so editing or deleting a line in the middle is detected. It sees only runs that go through it |
|
|
106
|
+
|
|
107
|
+
## Reading a result
|
|
108
|
+
|
|
109
|
+
```python
|
|
110
|
+
r = q.backtest_portfolio(panel, factor, long_q=0.2, short_q=0.2, hold=5, spread_bp=10)
|
|
111
|
+
r.alpha_beta() # regress net returns on the benchmark: alpha, beta, R2, Newey-West t
|
|
112
|
+
r.sharpe_ci() # stationary block bootstrap interval for the Sharpe ratio
|
|
113
|
+
q.analytics.ic_report(panel, factor, horizons=(1, 5, 20), delist_return=-0.3) # rank IC, ICIR, NW t, hit rate
|
|
114
|
+
q.analytics.sharpe_diff_ci(result_a.net_returns, result_b.net_returns) # paired: is B really different from A?
|
|
115
|
+
```
|
|
116
|
+
|
|
117
|
+
| Tool | Default that avoids a common mistake |
|
|
118
|
+
|---|---|
|
|
119
|
+
| `alpha_beta` | Newey-West errors with the plug-in lag; refuses constant or collinear factors; dates aligned on the intersection. Its t-statistics still over-reject in finite samples (9% instead of 5% in the A4 test), so read |t| below about 2.5 as no evidence |
|
|
120
|
+
| `information_coefficient`, `ic_report` | A name that stops trading earns 0 after its last bar, or `delist_return` if given, instead of a NaN that silently removes the losers from the test |
|
|
121
|
+
| `sharpe_ci` | Block bootstrap keeps autocorrelation (an AR(1) of 0.3 widens the standard error 1.36x, matching theory). Intervals cover about 93% at a nominal 95% on 500 days |
|
|
122
|
+
| `sharpe_diff_ci` | Both series are resampled on the same dates, so a small real difference is detectable (standard error 0.04 against 0.57 unpaired in the B3 test) |
|
|
123
|
+
|
|
124
|
+
## Your own weights, costs and capacity
|
|
125
|
+
|
|
126
|
+
`backtest_portfolio` builds quantile portfolios from a factor. Institutions separate the steps: signal, portfolio
|
|
127
|
+
construction (an optimiser with risk and turnover limits), simulation. `backtest_weights` is the third step: give it the
|
|
128
|
+
weights your optimiser produced. The library does not ship an optimiser on purpose.
|
|
129
|
+
|
|
130
|
+
```python
|
|
131
|
+
w = q.backtest_weights(panel, weights, spread_bp=10, borrow_bp=300,
|
|
132
|
+
impact=q.ImpactModel(aum=50e6, y=1.0)) # square-root impact: Y * sigma * sqrt(|trade| * AUM / ADV)
|
|
133
|
+
q.capacity_curve(panel, weights, aums=[1e6, 5e6, 25e6, 100e6], y_values=(0.5, 1.0, 2.0), spread_bp=10)
|
|
134
|
+
```
|
|
135
|
+
|
|
136
|
+
| Choice | Why |
|
|
137
|
+
|---|---|
|
|
138
|
+
| Opening or increasing a position in a name that is not eligible that day raises | An optimiser that buys outside the point-in-time universe is using information it should not have. Holding a name that left is allowed |
|
|
139
|
+
| The impact coefficient is a parameter and the capacity function takes a list of them | Y is of order 1 in the literature but unknown for a given market; one capacity number would be false precision |
|
|
140
|
+
| Unknown volatility or volume is charged the cap (100 bp per unit traded by default), never zero | A name you cannot size is not free to trade |
|
|
141
|
+
| It reports participation (p99, max, share of trades above 10% of ADV) next to the cost | The square-root law is least reliable at high participation; look at both |
|
|
142
|
+
| Costs are on the net trade per name | `backtest_portfolio` charges its two legs as separate sleeves; with overlapping tranches the two can differ by the netting saving |
|
|
143
|
+
|
|
144
|
+
## Taxes on one side, and short-selling limits
|
|
145
|
+
|
|
146
|
+
A flat round-trip spread cannot say that a sales tax is paid when a position is **sold**, including when a short is opened, or that
|
|
147
|
+
the rate changed on a date. Both engines take `buy_bp` and `sell_bp` for that, on top of the spread:
|
|
148
|
+
|
|
149
|
+
```python
|
|
150
|
+
tax = pd.Series([15.0, 20.0], index=[panel.dates[0], pd.Timestamp("2021-01-04")]) # a rate that changes on a date (example numbers)
|
|
151
|
+
r = q.backtest_portfolio(panel, factor, long_q=0.2, short_q=0.2, hold=5, spread_bp=10, sell_bp=tax)
|
|
152
|
+
|
|
153
|
+
ban = q.shortable_from_bans(panel.dates, panel.tickers, [("2020-03-16", "2020-09-15")]) # your checked ban periods
|
|
154
|
+
r = q.backtest_portfolio(replace(panel, shortable=ban), factor, sell_bp=tax) # the short leg is picked among shortable names
|
|
155
|
+
tax_panel = krx.sell_tax_panel(krx_panel, {"KOSPI": [...], "KOSDAQ": [...]}) # rates by market and effective date
|
|
156
|
+
```
|
|
157
|
+
|
|
158
|
+
| Choice | Why |
|
|
159
|
+
|---|---|
|
|
160
|
+
| A missing rate (NaN, or a date before the first entry) raises, unlike a NaN in a spread panel | A rate you do not know is not zero; a zero would look like a market without the tax |
|
|
161
|
+
| Opening a short pays `sell_bp`, covering it pays `buy_bp` | That is how a sales tax works |
|
|
162
|
+
| `Panel.shortable` is a bool frame; what is missing from it counts as **not** shortable | An unknown is not permission |
|
|
163
|
+
| `backtest_portfolio` ranks the short leg among shortable names; with none, that day is long-only (`short_leg_empty_days`) | It cannot short what it cannot borrow |
|
|
164
|
+
| `backtest_weights` raises when a short is opened or increased where it cannot be sold short (`check_shortable`) | A weight file that shorts during a ban is not a result |
|
|
165
|
+
| **No rate table and no ban calendar ship with the library** | Rates, effective dates and exemptions are law and change; a wrong table is worse than none. Check them against the tax authority and the regulator's notices |
|
|
166
|
+
|
|
167
|
+
## What can actually be traded
|
|
168
|
+
|
|
169
|
+
Trading every position change at the close assumes the security was open, that its price was not locked at a daily limit, and that you
|
|
170
|
+
can hold any fraction of a share. `Panel.can_buy` and `Panel.can_sell` state those assumptions as data, and both engines read them:
|
|
171
|
+
|
|
172
|
+
```python
|
|
173
|
+
from pitbacktest import execution as ex
|
|
174
|
+
can_buy, can_sell = ex.tradability(panel, limit=0.30) # halted = no price or no volume; a move of 29%+ locks the name (your market's limit)
|
|
175
|
+
panel = replace(panel, can_buy=can_buy, can_sell=can_sell)
|
|
176
|
+
r = q.backtest_portfolio(panel, factor) # a blocked trade does not happen; r.metrics["mean_stuck_weight"], ["longest_freeze_days"]
|
|
177
|
+
w = q.backtest_weights(panel, weights, capital=1e8, price=real_prices, lot=1.0, min_trade_value=0.0) # whole lots at the real price level
|
|
178
|
+
r_open = q.backtest_portfolio(ex.at_prices(panel, "open"), factor) # enter and mark at the open
|
|
179
|
+
```
|
|
180
|
+
|
|
181
|
+
| Choice | Why |
|
|
182
|
+
|---|---|
|
|
183
|
+
| Increasing a position needs `can_buy` and decreasing it needs `can_sell`, read on the execution day | A name locked at its upper limit has no sellers; one locked down has no buyers; a halted name has neither |
|
|
184
|
+
| A blocked trade leaves the position as it was | You cannot get out of a name you cannot sell. The frozen weight is reported (`mean_stuck_weight`, `longest_freeze_days`) |
|
|
185
|
+
| After a flagged delisting, and after `max_gap_days` (60) with no price at all, a position can be closed either way | The engines settle it (`delist_return`); refusing the exit froze a large part of one real short leg for ever |
|
|
186
|
+
| A name with a price but no volume stays frozen | A Korean trading suspension is exactly that. Its return is 0 here, which is **optimistic** if it ends in a delisting |
|
|
187
|
+
| Whole-lot sizes need the **real** price level (`meta["raw_close"]` in the KRX and Tiingo panels) | A back-adjusted series has an arbitrary level and gives wrong share counts |
|
|
188
|
+
| **No limit, lot size or minimum order ships with the library** | They differ by market and change by rule. You pass the ones you have checked |
|
|
189
|
+
|
|
190
|
+
## Intraday bars (Binance, 1 minute and up)
|
|
191
|
+
|
|
192
|
+
```python
|
|
193
|
+
from pitbacktest.crypto import intraday as ib, ArchiveStore
|
|
194
|
+
store = ArchiveStore()
|
|
195
|
+
ib.fetch_minutes(store, ["BTCUSDT", "ETHUSDT", ...], "2024-01-01", "2024-06-30") # monthly zips, resumable, no key
|
|
196
|
+
panel = ib.build_intraday_panel(store, symbols, "2024-01-01", "2024-06-30", bar="5min", top_n=30)
|
|
197
|
+
ib.latency_sweep(panel, factor, lags=(1, 2, 5, 15), one_way_bp=4) # the same signal entered k bars late
|
|
198
|
+
ib.breakeven_cost(panel, factor) # one-way bp at which the net mean is zero
|
|
199
|
+
```
|
|
200
|
+
|
|
201
|
+
What it answers: how fast an edge decays with latency, and at what cost it stops paying. What it cannot: whether the signal can
|
|
202
|
+
be traded. The archive has bars, not an order book, so there is no queue, no partial fill and no impact below the bar. See
|
|
203
|
+
`docs/intraday_design.md`.
|
|
204
|
+
|
|
205
|
+
| Choice | Why |
|
|
206
|
+
|---|---|
|
|
207
|
+
| Rows are labelled by bar **end**, and `entry_lag` below 1 is refused | A bar's close is only known when it ends; trading on it is look-ahead |
|
|
208
|
+
| Eligibility for every bar of day D uses days up to D - 1 | Day D's own volume must not decide whether D's bars are in the universe |
|
|
209
|
+
| Only the terminal run of zero-volume bars is removed; a no-trade stretch inside a contract's life keeps its price | Deleting it leaves a hole whose crossing return the engine would drop |
|
|
210
|
+
| Funding keeps its settlement instant (daily sums lose it) | A position held through the settlement pays it, and a flat one does not |
|
|
211
|
+
| `one_way_bp` instead of the engine's round-trip scalar | One unit trap fewer |
|
|
212
|
+
| A memory estimate, measured and not assumed, and a refusal with a bar length that fits | A year of 1-minute bars for 40 contracts needs about 3.4 GB |
|
|
213
|
+
| Annualised metrics from under half a year warn instead of returning NaN silently | Intraday studies are often a few months |
|
|
214
|
+
|
|
215
|
+
## Install
|
|
216
|
+
|
|
217
|
+
```bash
|
|
218
|
+
pip install pitbacktest # from PyPI; import it as `pitbacktest`. pandas and numpy are the only requirements
|
|
219
|
+
```
|
|
220
|
+
|
|
221
|
+
Status: **alpha**. No human has reviewed it yet; read `docs/verification_status.md` first. From a clone:
|
|
222
|
+
|
|
223
|
+
```bash
|
|
224
|
+
pip install -e . # pandas and numpy are the only requirements
|
|
225
|
+
pip install -e ".[yfinance]" # the public-data adapter
|
|
226
|
+
pip install -e ".[test]" # pytest, plus statsmodels and scipy for the cross-check tests
|
|
227
|
+
```
|
|
228
|
+
|
|
229
|
+
Python 3.10 to 3.14, pandas 2.0 to 3.0. CI runs every test file on each of them, including the oldest pandas and numpy the
|
|
230
|
+
package declares. Only Linux is tested.
|
|
231
|
+
|
|
232
|
+
## Example (public data)
|
|
233
|
+
|
|
234
|
+
```python
|
|
235
|
+
import pitbacktest as q
|
|
236
|
+
from pitbacktest.adapters.yfinance import load_panel
|
|
237
|
+
|
|
238
|
+
panel = load_panel(["AAPL", "MSFT", "NVDA", "AMZN", "GOOGL", "META", "TSLA", "JPM", "XOM", "JNJ"],
|
|
239
|
+
"2018-01-01", "2024-12-31")
|
|
240
|
+
|
|
241
|
+
print(panel.audit()) # data integrity audit
|
|
242
|
+
print(q.assert_timing(panel)) # timing self-check -> {'pass': True, ...}
|
|
243
|
+
|
|
244
|
+
# 5-day reversal as a long-short portfolio
|
|
245
|
+
r = q.backtest_portfolio(panel, -panel.close.pct_change(5), long_q=0.2, short_q=0.2, hold=5, spread_bp=5)
|
|
246
|
+
print(r.metrics["CAGR"], r.metrics["Sharpe"], r.metrics["MDD"])
|
|
247
|
+
```
|
|
248
|
+
|
|
249
|
+
This is a demonstration of the API on ten of today's large caps, not a result. The yfinance adapter is **not
|
|
250
|
+
point-in-time**: only tickers that exist today are returned, so delisted names are missing and performance is
|
|
251
|
+
biased upward. The adapter says so when it loads.
|
|
252
|
+
|
|
253
|
+
`examples/quickstart.py` runs all three entry points on synthetic data. It uses random factors on purpose; they
|
|
254
|
+
should fail every gate, and they do.
|
|
255
|
+
|
|
256
|
+
## Crypto
|
|
257
|
+
|
|
258
|
+
Two layers. `market="CRYPTO"` in the yfinance adapter (or `periods_per_year=365` on a `Panel`) fixes the annualisation:
|
|
259
|
+
with 252 days the CAGR, Sharpe and volatility of a 24/7 market come out wrong (on a 4.6-year sample, 252 reported 6.6
|
|
260
|
+
years of data). On top of that, `pitbacktest.crypto` builds a panel for Binance USDT-margined perpetuals from the public
|
|
261
|
+
archive and removes or measures the biases that usually flatter a crypto backtest:
|
|
262
|
+
|
|
263
|
+
| bias | what is done |
|
|
264
|
+
|---|---|
|
|
265
|
+
| Survivorship | the universe is every contract that ever traded (900, of which 376 are delisted or halted), not today's survivors. `survivors_only=True` reproduces the shortcut so its effect can be measured |
|
|
266
|
+
| Universe look-ahead | eligibility on day t uses only data up to t: minimum age, trailing median turnover, a valid and non-zero-volume bar |
|
|
267
|
+
| Stale prices | zero-volume bars (frozen prices after a halt) are dropped |
|
|
268
|
+
| Funding | the daily funding rate is charged by `backtest_portfolio` (long pays a positive rate, short receives) |
|
|
269
|
+
| Delisting | the last real bar is marked; `delist_return` sets the explicit return on the next day, and the events held are counted |
|
|
270
|
+
| Costs | taker fee plus a fixed half spread plus a thin-contract penalty, and a participation report instead of an invented impact model |
|
|
271
|
+
|
|
272
|
+
```python
|
|
273
|
+
import pitbacktest as q
|
|
274
|
+
from pitbacktest.crypto import fetch_all, build_panel, liquidity_cost_bp, participation_report
|
|
275
|
+
|
|
276
|
+
fetch_all() # once: ~900 contracts into ~/.cache/quantbt
|
|
277
|
+
panel = build_panel(start="2020-08-01", min_adv_usd=2e7, min_age_days=90)
|
|
278
|
+
cost = liquidity_cost_bp(panel) # one-way bp, date x ticker
|
|
279
|
+
res = q.backtest_portfolio(panel, factor, long_q=0.2, short_q=0.2, hold=5, spread_bp=cost, delist_return=-0.3)
|
|
280
|
+
print(res.metrics["funding_annual_bp"], res.metrics["delist_events_held"])
|
|
281
|
+
print(participation_report(panel, res.holdings, aum_usd=10e6))
|
|
282
|
+
```
|
|
283
|
+
|
|
284
|
+
`pitbacktest.validation` has the deflated Sharpe, PBO (CSCV) and a permutation test for the selection you ran.
|
|
285
|
+
`studies/crypto_cross_section/` is a worked, preregistered example: four factors, 12 trials, five gates. It is **not
|
|
286
|
+
validated**, and the README there explains what the biases changed and what went wrong along the way.
|
|
287
|
+
|
|
288
|
+
Still not modelled for crypto: market impact beyond the thin-contract penalty (use `backtest_weights` with `ImpactModel` for a square-root estimate), the settlement price of a
|
|
289
|
+
delisted contract after its last bar, borrow limits and margin, and anything outside Binance.
|
|
290
|
+
|
|
291
|
+
## Equities: Korea and the US
|
|
292
|
+
|
|
293
|
+
**Korea.** `pitbacktest.adapters.krx` builds a point-in-time panel from the cached daily files of the official KRX OpenAPI
|
|
294
|
+
(`fetch_days`, resumable, about two calls per trading day). Eligibility on day t uses only data up to t; securities that stop
|
|
295
|
+
trading are kept for the days they traded and flagged in `delist_after`. A code that vanishes for 120+ days and returns is
|
|
296
|
+
treated as a different security. Prices are adjusted by compounding the reference-price returns (price return only,
|
|
297
|
+
no dividends). You need a KRX OpenAPI key (`KRX_OPENAPI_KEY`).
|
|
298
|
+
|
|
299
|
+
```python
|
|
300
|
+
from pitbacktest.adapters.krx import fetch_days, build_krx_panel, load_key
|
|
301
|
+
fetch_days("2013-01-01", "2026-10-06", "~/.cache/quantbt/krx", key=load_key(".env"))
|
|
302
|
+
panel = build_krx_panel("~/.cache/quantbt/krx", start="2013-01-01")
|
|
303
|
+
```
|
|
304
|
+
|
|
305
|
+
**US.** Free data cannot remove the bias, so the package detects it, measures it, and shows what it could be worth:
|
|
306
|
+
|
|
307
|
+
| tool | what it does |
|
|
308
|
+
|---|---|
|
|
309
|
+
| `Panel.audit()` | `survivorship_suspected` is true when a panel of 30+ names over 3+ years has almost nothing that stops trading. The yfinance adapter warns |
|
|
310
|
+
| `equity.universe_coverage` | how many US stocks were listed each year and how many of the ones that stopped trading are in your panel (uses a free Tiingo ticker list; lower bound, it is thin before 2013) |
|
|
311
|
+
| `equity.survivorship_scenarios` | puts delistings back at random (uniform, or tilted to volatile or illiquid names) with an explicit delisting return, and reports the spread of your result. A what-if, not a correction |
|
|
312
|
+
| `equity.survivors_only` | the usual shortcut made explicit, so a result can be compared with and without it |
|
|
313
|
+
| `adapters.long_format.panel_from_long` | a strict door for point-in-time data from CRSP, Sharadar, Norgate and the like: permanent ids, duplicate checks, delisting returns compounded into the last close, ticker-reuse detection |
|
|
314
|
+
|
|
315
|
+
A ticker is not an identifier: in the probe, five tickers returned the history of a *different* company that later reused the
|
|
316
|
+
symbol, with no error. See `docs/survivorship.md`.
|
|
317
|
+
|
|
318
|
+
## What is verified, and what is not
|
|
319
|
+
|
|
320
|
+
[`docs/verification_status.md`](https://github.com/JanghyukChoi/quantbacktest/blob/main/docs/verification_status.md) lists the evidence behind each part and, more important, what is not established (no human review yet, Korean dividends, the impact coefficient, the weaker
|
|
321
|
+
parts such as `screen` and `backtest_event`). Read it before relying on a result.
|
|
322
|
+
|
|
323
|
+
## Tests: known answers, not real data
|
|
324
|
+
|
|
325
|
+
Every test uses data where the right answer is known, so none of them needs a network or real prices.
|
|
326
|
+
|
|
327
|
+
A second kind of check runs the engines on real data in three markets: `python docs/market_validation.py all` (needs the local caches; about 15 minutes, one market at a time) writes
|
|
328
|
+
`docs/market_validation.md`. It checks the method (a shuffled factor is rejected about as often as claimed, a perfect-foresight factor explodes and the day before entry does not pay, costs lower
|
|
329
|
+
the Sharpe, runs repeat to the last bit, the realism options equal the plain result when they have nothing to do) and reports no strategy's performance.
|
|
330
|
+
|
|
331
|
+
```bash
|
|
332
|
+
python tests/test_synthetic.py # core engine: look-ahead, neutralisation, base rates (10 tests)
|
|
333
|
+
python tests/test_annualization.py # 252 versus 365 days
|
|
334
|
+
python tests/test_validation.py # deflated Sharpe, PBO, permutation test: noise must fail, a real edge must pass
|
|
335
|
+
python tests/test_crypto.py # survivorship, point-in-time eligibility, stale bars, delisting, funding sign, costs
|
|
336
|
+
python tests/test_krx.py # Korea: delisted names kept, split-day return, listing day, code reuse, resumable fetch
|
|
337
|
+
python tests/test_equity.py # equity tools: delisting scenarios (known answer), coverage, survivors_only, long-format checks
|
|
338
|
+
python tests/test_reconcile.py # the engine against an independent loop implementation (agrees to 1e-17), and DSR/permutation false-positive rates on noise
|
|
339
|
+
python tests/test_ledger.py # trial ledger: distinct configurations, DSR count from the record, tamper detection
|
|
340
|
+
python tests/test_delisting_and_guards.py # a window that runs into a delisting keeps its loss; screen and backtest_event count delisted trades; intraday and cost guards
|
|
341
|
+
python tests/test_side_costs_shorting.py # buy and sell costs, dated rates, short-selling limits, KRX sell-tax panel
|
|
342
|
+
python tests/test_execution.py # halts, price limits, whole lots, minimum trade, trading at the open, frozen positions
|
|
343
|
+
python tests/test_argument_checks.py # bad arguments raise a clear error instead of quietly running something else
|
|
344
|
+
python tests/test_costs_events.py # spread estimators, crypto costs, event signals and their neutralisation, FM with missing returns, panel checks
|
|
345
|
+
python tests/test_yfinance_adapter.py # the free-data adapter against a fake yfinance: when it warns about survivorship, market-cap paths, errors
|
|
346
|
+
python tests/test_binance_archive.py # the downloader on a fake network: retries, 404, listings, daily bars, funding sums, cache
|
|
347
|
+
python tests/test_intraday.py # intraday layer on a fake archive: parsing, aggregation, point in time, latency, funding, break-even, memory, download
|
|
348
|
+
python tests/test_packaging.py # version, Python floor, CI matrix and classifiers agree
|
|
349
|
+
python tests/smoke_installed.py # run from outside the repo against an installed wheel (what the CI package job does)
|
|
350
|
+
python tests/test_weights.py # weights: equals the engine, plain-loop reference with all costs, known answers, guard rails, capacity curve
|
|
351
|
+
python tests/test_tiingo.py # Tiingo adapter offline: delisting kept, windows cut, resumable, quota stop, point in time, loud failures
|
|
352
|
+
python tests/test_analytics.py # alpha/beta, IC, bootstrap: against statsmodels and scipy when installed, known answers, error rates
|
|
353
|
+
```
|
|
354
|
+
|
|
355
|
+
| Test | Expectation |
|
|
356
|
+
|---|---|
|
|
357
|
+
| T1 perfect-foresight factor | large positive result (timing is aligned) |
|
|
358
|
+
| T2 random factor | near zero (no false positives) |
|
|
359
|
+
| T3 look-ahead factor | caught by the look-ahead check |
|
|
360
|
+
| T4 random factor vs look-ahead check | passes (no false alarm) |
|
|
361
|
+
| T5 shuffled-null threshold | produced and sensible |
|
|
362
|
+
| T6 neutralisation | a control used as the factor disappears after neutralising |
|
|
363
|
+
| T7 base rate | random firing gives a lift near zero |
|
|
364
|
+
| T8 to T10 | small universes, small Fama-MacBeth regressions, perfectly collinear controls |
|
|
365
|
+
| T11 | annualisation scales by exactly sqrt(365/252) |
|
|
366
|
+
| V1 to V3 | deflated Sharpe, PBO and permutation p-value: noise looks like noise, a real edge is caught |
|
|
367
|
+
| K1 to K5 | Korean adapter: a stock gone by the end is in the panel and flagged, a 50:1 split leaves the return unchanged, the listing-day move is dropped, a returning code becomes a new security, fetching resumes and stops cleanly on a quota error |
|
|
368
|
+
| E1 to E6 | injecting a 5% yearly delisting rate at -30% lowers an equal-weight long book by 1.5% a year (the known answer), coverage counts, scenarios, strict long-format input |
|
|
369
|
+
| R1 to R7 | portfolio returns agree with a separate plain-loop implementation (long-short, long-only, delisting, funding), CAGR, Sharpe, drawdown and Sortino match textbook definitions, cost units are pinned, DSR and the permutation test do not reject noise more than they claim |
|
|
370
|
+
| L1 to L4 | the ledger counts a repeated run once and any change as a new trial, its DSR equals the direct computation, editing or deleting a line breaks the hash chain, recording changes no number |
|
|
371
|
+
| S1 to S9 | Buy and sell costs equal an independent loop for weights that change sign (to 1e-18); `backtest_portfolio` and `backtest_weights` agree with them, also when the legs empty and refill (the case where swapping the two rates matters); a dated rate is read per date; an unknown rate raises; the short leg is chosen among shortable names; shorts cannot be opened where they cannot be sold; a long turned into a smaller short outside the universe is caught; `shortable_from_bans` and `krx.sell_tax_panel` build the right frames. Eleven planted bugs are all caught |
|
|
372
|
+
| X1 to X9 | Positions and net returns equal an independent loop when trades are blocked (buys and sells separately, at lag 0, 1 and 2); nothing blocked is bit for bit the old result; a frozen position is measured as counted by hand; a long or short in a delisted name is closed instead of frozen; whole lots at the real price equal a loop; a minimum trade value skips and counts trades; trading at the open equals the engine on a panel built from the open. Nine planted bugs are caught |
|
|
373
|
+
| DG1 to DG3 | `Panel.forward` keeps the loss of a trade that runs into a delisting (hand-computed, with and without `delist_return`), `backtest_event` and `screen` count those trades, and the intraday, cost, hazard and short-panel guards hold. Ten planted removals are all caught |
|
|
374
|
+
| AC1 to AC4 | bad arguments stop with a clear error (mistyped weighting or benchmark, quantiles outside (0, 1], short_q=0, hold below 1, a negative cost that used to turn a Sharpe of -1.01 into +0.96, a boolean factor, horizons, an impact model that makes no sense), and the `screen` default of 20 shuffled-null repetitions, since two repetitions give a threshold biased far below pure noise. Twenty planted removals are all caught |
|
|
375
|
+
| CE1 to CE11 | Roll recovers a 2 cent spread, Corwin-Schultz a 40 bp one, the spread model its coefficients, the crypto cost model is exact, an event signal that only echoes a control keeps -8% of its effect after controls while a real 100 bp effect keeps 102%, Fama-MacBeth with missing returns equals a per-day least squares, the paired difference, panel validation and the point-in-time mask follow their documentation |
|
|
376
|
+
| Y1 to Y5 | the yfinance adapter warns that it is not point in time every time, warns about survivors-only only with 30+ tickers none of which end early, reports partial market-cap coverage with counts, and refuses nothing-eligible and one-ticker panels |
|
|
377
|
+
| B1 to B7 | the Binance downloader retries a 503 and not a 403, treats 404 as missing, follows paginated listings, sums 8-hour funding settlements per day in both timestamp units, drops today's unfinished bar, caches, and survives one failing symbol |
|
|
378
|
+
| I1 to I10, I11 | 1-minute parsing in both timestamp units, aggregation equal to an independent loop, eligibility one day behind (first eligible bar is the listing day + 5 whole days), a knows-one-bar-ahead signal earns at lag 1 and nothing at lag 2, funding paid exactly once per settlement, break-even cost makes the net mean zero, memory guard, resumable download with recorded 404s, halts and no-trade stretches. Planted bugs (shifted labels, same-day eligibility, wrong funding bar, dropped partial bars and others) are caught |
|
|
379
|
+
| P1, P2 | the version is the same in pyproject, `__version__` and the changelog; the declared Python floor, the classifiers and the CI matrix agree, and the matrix tests the oldest declared pandas and numpy |
|
|
380
|
+
| R8 | `neutralize` equals a least-squares solution; a factor inside the controls' span gives NaN even with rounding noise (found because the old behaviour made one test fail on Python 3.10 only) |
|
|
381
|
+
| W1 to W6 | `backtest_weights` equals the engine on the engine's own holdings (exactly, except for a netting saving it documents), equals a plain-loop implementation with spread, borrow and impact, impact matches a hand calculation and scales as sqrt(AUM) and linearly in Y, borrow follows its formula, the untradable cap is exact, guard rails, capacity curve shape. Ten planted bugs are all caught |
|
|
382
|
+
| U1 to U8 | Tiingo adapter against a fake API: a delisted security is kept and flagged, a split does not move the return, a sub-dollar stock is never eligible, a reused ticker becomes two securities, a quota stops cleanly and resumes, eligibility ignores the future, the sample draw is seeded, malformed answers leave no file. Nine planted bugs are all caught |
|
|
383
|
+
| A1 to A8, B1 to B3 | alpha and beta equal statsmodels' Newey-West regression to 1e-12 and IC equals scipy's Spearman (when installed), known answers and invariances, rejection rates on noise, forward returns equal a loop implementation with and without delisting returns, bootstrap coverage and standard errors against theory, a paired difference of identical series is exactly zero. Planting eight bugs in the module (wrong taper, shifted window, ignored delisting, unpaired resampling and others) is caught by these tests every time |
|
|
384
|
+
| C1 to C8 | delisted contracts included, eligibility unchanged by future data, frozen bars dropped, new listings wait `min_age_days`, funding sign and size, explicit delisting return, bounded costs |
|
|
385
|
+
|
|
386
|
+
## Limits
|
|
387
|
+
|
|
388
|
+
- No bundled data. The yfinance adapter is not point-in-time; the Binance archive and KRX paths are. US stocks need data you bring.
|
|
389
|
+
- Fundamentals (`chars`) must be supplied by the user for the controls to be complete; without them it warns.
|
|
390
|
+
- Everything is in English except the Korean security-name patterns and the quota message the KRX adapter has to match.
|
|
391
|
+
- For research and education. Nothing here is investment advice.
|
|
392
|
+
|
|
393
|
+
## License
|
|
394
|
+
|
|
395
|
+
MIT
|