topstep-backtest 0.3.0__tar.gz → 0.4.0__tar.gz
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/.gitignore +2 -0
- {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/AGENTS.md +122 -26
- {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/CHANGELOG.md +149 -0
- {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/PKG-INFO +45 -8
- {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/README.md +44 -7
- {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/docs/DESIGN.md +40 -1
- {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/docs/INDICATORS.md +61 -0
- {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/docs/ROADMAP.md +6 -4
- {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/docs/TUTORIAL_EMA_CROSSOVER.md +45 -6
- {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/examples/run_montecarlo.py +30 -7
- {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/examples/run_real_data.py +81 -4
- topstep_backtest-0.4.0/examples/run_spaced.py +125 -0
- {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/examples/run_windows.py +71 -8
- topstep_backtest-0.4.0/examples/session_scoped.py +151 -0
- {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/src/topstep_backtest/__init__.py +6 -0
- topstep_backtest-0.4.0/src/topstep_backtest/core/sessions.py +143 -0
- topstep_backtest-0.4.0/src/topstep_backtest/data/loaders.py +117 -0
- {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/src/topstep_backtest/data/synthetic.py +57 -21
- {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/src/topstep_backtest/harness.py +39 -7
- {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/src/topstep_backtest/metrics/__init__.py +31 -1
- topstep_backtest-0.4.0/src/topstep_backtest/metrics/confidence.py +536 -0
- {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/src/topstep_backtest/metrics/economics.py +6 -3
- {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/src/topstep_backtest/metrics/montecarlo.py +65 -34
- topstep_backtest-0.4.0/src/topstep_backtest/metrics/windows.py +629 -0
- {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/src/topstep_backtest/strategy/symbol.py +214 -22
- {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/src/topstep_backtest/tearsheet/__init__.py +282 -17
- topstep_backtest-0.4.0/src/topstep_backtest/tearsheet/_assets/sweep.css +154 -0
- topstep_backtest-0.4.0/src/topstep_backtest/tearsheet/_assets/sweep.js +42 -0
- {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/src/topstep_backtest/tearsheet/_assets/tearsheet.css +42 -22
- {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/src/topstep_backtest/tearsheet/_assets/tearsheet.js +94 -27
- topstep_backtest-0.4.0/src/topstep_backtest/tearsheet/sweep.py +522 -0
- topstep_backtest-0.4.0/tests/unit/test_confidence.py +342 -0
- topstep_backtest-0.4.0/tests/unit/test_loaders.py +176 -0
- {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/tests/unit/test_montecarlo.py +17 -5
- topstep_backtest-0.4.0/tests/unit/test_sessions.py +138 -0
- topstep_backtest-0.4.0/tests/unit/test_sweep_tearsheet.py +131 -0
- {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/tests/unit/test_symbol_strategy.py +216 -2
- {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/tests/unit/test_synthetic.py +75 -0
- {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/tests/unit/test_tearsheet.py +77 -0
- {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/tests/unit/test_windows.py +103 -1
- topstep_backtest-0.3.0/src/topstep_backtest/metrics/windows.py +0 -314
- {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/LICENSE +0 -0
- {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/data/sample_mnq_1m.csv +0 -0
- {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/docs/topstep-rules.md +0 -0
- {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/examples/ema_cross.py +0 -0
- {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/examples/hand_wired.py +0 -0
- {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/examples/run_combine.py +0 -0
- {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/examples/run_replay.py +0 -0
- {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/examples/run_tearsheet.py +0 -0
- {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/examples/sma_cross.py +0 -0
- {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/examples/talib_macd.py +0 -0
- {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/pyproject.toml +0 -0
- {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/src/topstep_backtest/_render.py +0 -0
- {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/src/topstep_backtest/clock/__init__.py +0 -0
- {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/src/topstep_backtest/clock/live_clock.py +0 -0
- {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/src/topstep_backtest/clock/test_clock.py +0 -0
- {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/src/topstep_backtest/core/__init__.py +0 -0
- {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/src/topstep_backtest/core/ids.py +0 -0
- {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/src/topstep_backtest/core/instruments.py +0 -0
- {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/src/topstep_backtest/core/money.py +0 -0
- {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/src/topstep_backtest/core/time.py +0 -0
- {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/src/topstep_backtest/data/__init__.py +0 -0
- {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/src/topstep_backtest/data/clean.py +0 -0
- {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/src/topstep_backtest/data/continuous.py +0 -0
- {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/src/topstep_backtest/data/feed.py +0 -0
- {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/src/topstep_backtest/data/validator.py +0 -0
- {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/src/topstep_backtest/data/wrangler.py +0 -0
- {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/src/topstep_backtest/engine/__init__.py +0 -0
- {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/src/topstep_backtest/engine/backtest.py +0 -0
- {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/src/topstep_backtest/execution/__init__.py +0 -0
- {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/src/topstep_backtest/execution/rejections.py +0 -0
- {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/src/topstep_backtest/execution/sim_broker.py +0 -0
- {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/src/topstep_backtest/fills/__init__.py +0 -0
- {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/src/topstep_backtest/fills/bar_fill.py +0 -0
- {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/src/topstep_backtest/fills/fees.py +0 -0
- {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/src/topstep_backtest/fills/path.py +0 -0
- {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/src/topstep_backtest/indicators/__init__.py +0 -0
- {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/src/topstep_backtest/indicators/base.py +0 -0
- {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/src/topstep_backtest/indicators/library.py +0 -0
- {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/src/topstep_backtest/indicators/talib_adapter.py +0 -0
- {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/src/topstep_backtest/metrics/overfitting.py +0 -0
- {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/src/topstep_backtest/metrics/stats.py +0 -0
- {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/src/topstep_backtest/metrics/walkforward.py +0 -0
- {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/src/topstep_backtest/protocols.py +0 -0
- {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/src/topstep_backtest/py.typed +0 -0
- {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/src/topstep_backtest/replay.py +0 -0
- {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/src/topstep_backtest/rules/__init__.py +0 -0
- {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/src/topstep_backtest/rules/kernel.py +0 -0
- {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/src/topstep_backtest/rules/params.py +0 -0
- {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/src/topstep_backtest/strategy/__init__.py +0 -0
- {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/src/topstep_backtest/strategy/base.py +0 -0
- {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/src/topstep_backtest/strategy/tracker.py +0 -0
- {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/src/topstep_backtest/tearsheet/_assets/lightweight-charts.LICENSE +0 -0
- {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/src/topstep_backtest/tearsheet/_assets/lightweight-charts.standalone.production.js +0 -0
- {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/tests/__init__.py +0 -0
- {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/tests/conftest.py +0 -0
- {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/tests/golden/__init__.py +0 -0
- {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/tests/golden/artifacts/verdict_failed_mll_s50k.json +0 -0
- {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/tests/golden/artifacts/verdict_passed_s50k.json +0 -0
- {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/tests/golden/test_combine_kernel.py +0 -0
- {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/tests/golden/test_facade_equivalence.py +0 -0
- {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/tests/golden/test_replay_goldens.py +0 -0
- {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/tests/golden/test_sugar_equivalence.py +0 -0
- {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/tests/golden/test_verdict_goldens.py +0 -0
- {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/tests/parity/__init__.py +0 -0
- {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/tests/parity/test_broker_conformance.py +0 -0
- {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/tests/property/__init__.py +0 -0
- {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/tests/property/test_indicator_props.py +0 -0
- {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/tests/property/test_kernel_props.py +0 -0
- {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/tests/property/test_money_props.py +0 -0
- {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/tests/property/test_replay_props.py +0 -0
- {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/tests/unit/__init__.py +0 -0
- {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/tests/unit/test_bar_fill.py +0 -0
- {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/tests/unit/test_clean.py +0 -0
- {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/tests/unit/test_clock.py +0 -0
- {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/tests/unit/test_continuous.py +0 -0
- {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/tests/unit/test_data_feed.py +0 -0
- {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/tests/unit/test_economics.py +0 -0
- {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/tests/unit/test_engine.py +0 -0
- {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/tests/unit/test_fees.py +0 -0
- {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/tests/unit/test_harness.py +0 -0
- {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/tests/unit/test_indicators.py +0 -0
- {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/tests/unit/test_instruments.py +0 -0
- {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/tests/unit/test_overfitting.py +0 -0
- {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/tests/unit/test_path.py +0 -0
- {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/tests/unit/test_replay.py +0 -0
- {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/tests/unit/test_replay_tearsheet.py +0 -0
- {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/tests/unit/test_sim_broker.py +0 -0
- {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/tests/unit/test_stats.py +0 -0
- {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/tests/unit/test_talib_adapter_hardening.py +0 -0
- {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/tests/unit/test_time.py +0 -0
- {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/tests/unit/test_tracker.py +0 -0
- {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/tests/unit/test_validator.py +0 -0
- {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/tests/unit/test_walkforward.py +0 -0
- {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/tests/unit/test_wrangler.py +0 -0
|
@@ -127,7 +127,7 @@ Generated from the live objects — every signature below is real.
|
|
|
127
127
|
|
|
128
128
|
| Name | Signature | What it does |
|
|
129
129
|
|---|---|---|
|
|
130
|
-
| `SymbolStrategy` | `(contract_id: 'str', require_ready: 'bool' = True, warmup: 'int \| None' = None)` | Base for strategies trading exactly one contract. |
|
|
130
|
+
| `SymbolStrategy` | `(contract_id: 'str', require_ready: 'bool' = True, warmup: 'int \| None' = None, trade_sessions: 'Sequence[Session] \| None' = None)` | Base for strategies trading exactly one contract. |
|
|
131
131
|
|
|
132
132
|
**Position / order views**
|
|
133
133
|
|
|
@@ -144,11 +144,17 @@ Generated from the live objects — every signature below is real.
|
|
|
144
144
|
| `bars_from_dataframe` | `(df: 'Any', *, contract_id: 'str', spec: 'InstrumentSpec', unit: 'AggregateBarUnit', unit_number: 'int', stamp: "Literal['open', 'close']") -> 'tuple[Bar, ...]'` | Build ``Bar`` objects from a pandas DataFrame of OHLCV candles. |
|
|
145
145
|
| `bars_from_records` | `(rows: 'Iterable[tuple[object, ...]]', *, contract_id: 'str', spec: 'InstrumentSpec', unit: 'AggregateBarUnit', unit_number: 'int', stamp: "Literal['open', 'close']") -> 'tuple[Bar, ...]'` | Build tick-grid-validated ``Bar`` objects from ``(ts, o, h, l, c, v)`` rows. |
|
|
146
146
|
|
|
147
|
+
**Data in (Parquet export)**
|
|
148
|
+
|
|
149
|
+
| Name | Signature | What it does |
|
|
150
|
+
|---|---|---|
|
|
151
|
+
| `load_bars` | `(path: 'str \| Path') -> 'tuple[tuple[Bar, ...], InstrumentSpec, dict[str, Any]]'` | Read the export -> (bars, spec, metadata), ready for ``Backtest(...)``. |
|
|
152
|
+
|
|
147
153
|
**Data in (synthetic)**
|
|
148
154
|
|
|
149
155
|
| Name | Signature | What it does |
|
|
150
156
|
|---|---|---|
|
|
151
|
-
| `synthetic_bars` | `(*, contract_id: 'str', spec: 'InstrumentSpec', start_day: 'date', days: 'int', seed: 'int', start_price: 'Decimal', bars_per_day: 'int' =
|
|
157
|
+
| `synthetic_bars` | `(*, contract_id: 'str', spec: 'InstrumentSpec', start_day: 'date', days: 'int', seed: 'int', start_price: 'Decimal', bars_per_day: 'int \| None' = None, unit: 'AggregateBarUnit' = <AggregateBarUnit.MINUTE: 2>, unit_number: 'int' = 1, drift_ticks_per_day: 'int' = 0, vol_ticks: 'int' = 8, hours: 'Hours' = 'rth') -> 'tuple[Bar, ...]'` | Generate ``days`` trading sessions of consistent, on-grid OHLCV bars. |
|
|
152
158
|
|
|
153
159
|
**Instruments**
|
|
154
160
|
|
|
@@ -158,6 +164,15 @@ Generated from the live objects — every signature below is real.
|
|
|
158
164
|
| `symbol_of_contract_id` | `(contract_id: 'str') -> 'str'` | Extract the product symbol from a gateway contract id. |
|
|
159
165
|
| `InstrumentSpec` | `(*args, **kwargs)` | Frozen per-product economics and session metadata. |
|
|
160
166
|
|
|
167
|
+
**Sessions**
|
|
168
|
+
|
|
169
|
+
| Name | Signature | What it does |
|
|
170
|
+
|---|---|---|
|
|
171
|
+
| `Session` | `(*args, **kwargs)` | A named intraday window, defined in its own local timezone. |
|
|
172
|
+
| `ASIA` | `Session(name='ASIA', tz='Asia/Tokyo', start=09:00:00, end=15:00:00)` | A named intraday window, defined in its own local timezone. |
|
|
173
|
+
| `LONDON` | `Session(name='LONDON', tz='Europe/London', start=08:00:00, end=16:30:00)` | A named intraday window, defined in its own local timezone. |
|
|
174
|
+
| `NEW_YORK` | `Session(name='NEW_YORK', tz='America/New_York', start=09:30:00, end=16:00:00)` | A named intraday window, defined in its own local timezone. |
|
|
175
|
+
|
|
161
176
|
**Results**
|
|
162
177
|
|
|
163
178
|
| Name | Signature | What it does |
|
|
@@ -170,7 +185,8 @@ Generated from the live objects — every signature below is real.
|
|
|
170
185
|
|
|
171
186
|
| Name | Signature | What it does |
|
|
172
187
|
|---|---|---|
|
|
173
|
-
| `render_html` | `(report: 'Report', *, replay: 'ReplaySpec' = 'auto') -> 'str'` | Render ``report`` as one self-contained interactive HTML document. |
|
|
188
|
+
| `render_html` | `(report: 'Report', *, replay: 'ReplaySpec' = 'auto', confidence: 'MonteCarloConfidence \| None' = None, crosscheck: 'CrossCheck \| None' = None) -> 'str'` | Render ``report`` as one self-contained interactive HTML document. |
|
|
189
|
+
| `render_sweep_html` | `(sweep: 'WindowSweep \| SpacedSweep') -> 'str'` | Render a sweep as one self-contained HTML document. |
|
|
174
190
|
|
|
175
191
|
**Replay recording**
|
|
176
192
|
|
|
@@ -199,12 +215,29 @@ Generated from the live objects — every signature below is real.
|
|
|
199
215
|
| `MonteCarloResult` | `(*args, **kwargs)` | Outcome distribution over ``paths`` synthetic Combine attempts. |
|
|
200
216
|
| `FailureMode` | `FailureMode.MLL_BREACH \| FailureMode.CONSISTENCY_BLOCKED \| FailureMode.TARGET_NOT_REACHED` | Why a simulated path did not pass. Ordered by when it is decided. |
|
|
201
217
|
|
|
218
|
+
**Monte-Carlo confidence**
|
|
219
|
+
|
|
220
|
+
| Name | Signature | What it does |
|
|
221
|
+
|---|---|---|
|
|
222
|
+
| `mc_confidence` | `(result: 'BacktestResult', *, params: 'CombineParams', paths: 'int' = 2000, horizon_days: 'int \| None' = None, block_length: 'int' = 5, seed: 'int' = 0, outer: 'int' = 200, inner_paths: 'int' = 200, lengths: 'Sequence[int]' = (1, 5, 10, 20)) -> 'MonteCarloConfidence'` | One call: the estimate plus its CI, sensitivity row, and year strata. |
|
|
223
|
+
| `MonteCarloConfidence` | `(*args, **kwargs)` | The point estimate and every qualifier this module can attach to it. |
|
|
224
|
+
| `pass_probability_ci` | `(result: 'BacktestResult', *, params: 'CombineParams', paths: 'int' = 2000, horizon_days: 'int \| None' = None, block_length: 'int' = 5, seed: 'int' = 0, outer: 'int' = 200, inner_paths: 'int' = 200) -> 'PassProbabilityCI'` | Double bootstrap: a confidence band for the pass probability. |
|
|
225
|
+
| `PassProbabilityCI` | `(*args, **kwargs)` | A pass probability with the error bar its sample size actually earns. |
|
|
226
|
+
| `block_length_sensitivity` | `(result: 'BacktestResult', *, params: 'CombineParams', lengths: 'Sequence[int]' = (1, 5, 10, 20), paths: 'int' = 1000, horizon_days: 'int \| None' = None, seed: 'int' = 0) -> 'BlockLengthSensitivity'` | Re-run the Monte Carlo across block lengths and report the swing. |
|
|
227
|
+
| `BlockLengthSensitivity` | `(*args, **kwargs)` | The same estimate at several block lengths, plus how far it moved. |
|
|
228
|
+
| `monte_carlo_by_year` | `(result: 'BacktestResult', *, params: 'CombineParams', paths: 'int' = 1000, horizon_days: 'int \| None' = None, block_length: 'int' = 5, seed: 'int' = 0) -> 'YearStratification'` | One Monte Carlo per calendar year of the source run. |
|
|
229
|
+
| `YearStratification` | `(*args, **kwargs)` | Per-year estimates, ascending by year. |
|
|
230
|
+
| `crosscheck` | `(mc: 'MonteCarloResult', sweep: 'WindowSweep') -> 'CrossCheck'` | Compare a Monte Carlo against a window sweep, outcome by outcome. |
|
|
231
|
+
| `CrossCheck` | `(*args, **kwargs)` | The bootstrap and the window sweep, forced to answer side by side. |
|
|
232
|
+
|
|
202
233
|
**Sequential Combines**
|
|
203
234
|
|
|
204
235
|
| Name | Signature | What it does |
|
|
205
236
|
|---|---|---|
|
|
206
237
|
| `sequential_combines` | `(bars: 'Sequence[Bar]', factory: 'Callable[[], Strategy]', *, window_days: 'int', account: 'AccountSize' = <AccountSize.S50K: '50K'>, dll_enabled: 'bool' = False, warm_start: 'bool' = True, validate: 'bool' = True) -> 'WindowSweep'` | Run one fresh Combine per non-overlapping ``window_days``-day window. |
|
|
238
|
+
| `spaced_combines` | `(bars: 'Sequence[Bar]', factory: 'Callable[[], Strategy]', *, window_days: 'int', periods: 'int', account: 'AccountSize' = <AccountSize.S50K: '50K'>, dll_enabled: 'bool' = False, warm_start: 'bool' = True, validate: 'bool' = True) -> 'SpacedSweep'` | Run ``periods`` fresh Combines with start days spread evenly over the tape. |
|
|
207
239
|
| `WindowSweep` | `(*args, **kwargs)` | Every window's attempt, plus the rates over them. |
|
|
240
|
+
| `SpacedSweep` | `(*args, **kwargs)` | A requested number of periods, start days spread evenly, overlap allowed. |
|
|
208
241
|
| `WindowResult` | `(*args, **kwargs)` | One window's Combine attempt, resolved on its own merits. |
|
|
209
242
|
|
|
210
243
|
**Overfitting guards**
|
|
@@ -249,7 +282,7 @@ Generated from the live objects — every signature below is real.
|
|
|
249
282
|
|
|
250
283
|
| Name | Signature | What it does |
|
|
251
284
|
|---|---|---|
|
|
252
|
-
| `SymbolStrategy.use` | `(self, indicator: 'T') -> 'T'` | Register an indicator: auto-updated on every matching bar and |
|
|
285
|
+
| `SymbolStrategy.use` | `(self, indicator: 'T', *, session: 'Session \| None' = None) -> 'T'` | Register an indicator: auto-updated on every matching bar and |
|
|
253
286
|
| `SymbolStrategy.buy` | `(self, size: 'int', *, stop_loss_ticks: 'int \| None' = None, take_profit_ticks: 'int \| None' = None, limit_price: 'Decimal \| None' = None, stop_price: 'Decimal \| None' = None, custom_tag: 'str \| None' = None) -> 'int \| None'` | Buy this contract (market unless a price kwarg implies otherwise); |
|
|
254
287
|
| `SymbolStrategy.sell` | `(self, size: 'int', *, stop_loss_ticks: 'int \| None' = None, take_profit_ticks: 'int \| None' = None, limit_price: 'Decimal \| None' = None, stop_price: 'Decimal \| None' = None, custom_tag: 'str \| None' = None) -> 'int \| None'` | Sell this contract (market unless a price kwarg implies otherwise); |
|
|
255
288
|
| `SymbolStrategy.close` | `(self) -> 'None'` | Flatten this contract's position; a rejection goes to ``on_reject``. |
|
|
@@ -418,6 +451,9 @@ Rules here; tables, per-function warmups and the refused-function list in `docs/
|
|
|
418
451
|
tradeable price. Never route one into grid math without an explicit `round_to_tick`. They are
|
|
419
452
|
float64-precise, not Decimal-exact: deterministic across reruns, but a `Cross` on a Bollinger
|
|
420
453
|
edge can flip on `STDDEV`'s cancellation noise. A band touch is not exact.
|
|
454
|
+
- **`use(ind, session=…)` scopes the DATA; `trade_sessions=` scopes the DECISION.** Two
|
|
455
|
+
independent switches — see §5.11. Scoping an indicator does not restrict trading, and
|
|
456
|
+
restricting trading does not starve an indicator.
|
|
421
457
|
- One indicator instance per thread; a parameter sweep gets one per worker.
|
|
422
458
|
|
|
423
459
|
### 5.3 Data
|
|
@@ -560,12 +596,23 @@ observed trading days and replays each synthetic sequence through a fresh `Combi
|
|
|
560
596
|
- **`block_length=1` is a footgun.** It degenerates to an i.i.d. resample, destroys the
|
|
561
597
|
losing streaks that actually blow accounts, and will report a pass probability that is far
|
|
562
598
|
too kind. Default is 5.
|
|
563
|
-
-
|
|
564
|
-
|
|
565
|
-
`
|
|
566
|
-
|
|
567
|
-
|
|
599
|
+
- **`horizon_days` defaults to `BILLING_MONTH_DAYS` (21), never the observed day count.** A
|
|
600
|
+
Combine has no time limit, only a monthly fee, so "one attempt" defaults to one fee cycle —
|
|
601
|
+
the same unit `EvalEconomics` bills in. The fixed default is also the safety property: a
|
|
602
|
+
blown run stops recording days at the breach, so its count is a survival time, not a
|
|
603
|
+
Combine length, and is never used as a horizon. `source_truncated` still flags that such a
|
|
604
|
+
sample is survivorship-biased by construction — the days after the blow-up do not exist,
|
|
605
|
+
every figure is conditioned on having survived, and nothing can repair that; do not quote
|
|
606
|
+
such a result without the caveat.
|
|
568
607
|
- **`provisional` below 30 source days.** Resampling cannot create information.
|
|
608
|
+
- **The point estimate ships with its own cross-examination** (`metrics/confidence.py`).
|
|
609
|
+
`pass_probability_ci` double-bootstraps the source days themselves — the error bar the
|
|
610
|
+
day count earns, which the path count never was. `block_length_sensitivity` shows whether
|
|
611
|
+
the streak assumption is load-bearing. `monte_carlo_by_year` refuses to average a hostile
|
|
612
|
+
year against a kind one. `crosscheck` compares against `sequential_combines` under a
|
|
613
|
+
binomial 2×SE null — when they disagree, the disagreement is the finding. Quote a pass
|
|
614
|
+
probability with its CI, not alone; `mc_confidence` bundles the lot and
|
|
615
|
+
`report.to_html(path, confidence=..., crosscheck=...)` renders the cards.
|
|
569
616
|
- It cannot invent a regime the tape never contained, and it inherits every uncalibrated
|
|
570
617
|
constant (§5.4). A probability to three decimals from unverified inputs is precise, not
|
|
571
618
|
accurate.
|
|
@@ -656,25 +703,70 @@ account.
|
|
|
656
703
|
- Still your problem on a multi-year tape: **exchange holidays** (~20/year, no calendar
|
|
657
704
|
ships — filter upstream) and the uncalibrated constants (§5.4).
|
|
658
705
|
|
|
706
|
+
### 5.11 Sessions scope indicator DATA and trading DECISIONS, separately
|
|
707
|
+
|
|
708
|
+
`ASIA` / `LONDON` / `NEW_YORK` live in `core/sessions.py`; the worked example is
|
|
709
|
+
`examples/session_scoped.py`.
|
|
710
|
+
|
|
711
|
+
- **The two switches are independent, and conflating them is the bug.**
|
|
712
|
+
`use(Atr(14), session=NEW_YORK)` restricts which bars that indicator is computed from.
|
|
713
|
+
`SymbolStrategy(..., trade_sessions=(NEW_YORK,))` restricts when `on_bar` may fire. Indicators
|
|
714
|
+
advance regardless of `trade_sessions` — an indicator fed only the tradable window develops
|
|
715
|
+
gaps and computes a different value from the same tape. Both default to today's behaviour, and
|
|
716
|
+
an unscoped strategy is byte-identical to one written before sessions existed.
|
|
717
|
+
- **Which indicators to scope is a modelling decision, and it is not uniform.** Dispersion
|
|
718
|
+
measures (`Atr`, `StdDev`, `Rsi`, `Stoch`, `BBands`) describe how much price moves PER BAR, and
|
|
719
|
+
that is session-dependent: on a 24h feed an `Atr(14)` read at 09:30 ET is computed almost
|
|
720
|
+
entirely from thin pre-market bars, so it understates NY volatility exactly when stop distance
|
|
721
|
+
is being sized. Level measures (`Sma`, `Ema`) answer where price IS, and the overnight move is
|
|
722
|
+
real — an NY-only `Ema` is anchored to yesterday's 16:00 close. **Levels are continuous across
|
|
723
|
+
sessions; dispersion is not.**
|
|
724
|
+
- **A scoped indicator warms in ITS OWN cadence.** It needs `history_bars` bars *of its session*,
|
|
725
|
+
so on 5-minute bars an NY-scoped `Sma(30)` spans ~25 trading days against ~7 unscoped. Read
|
|
726
|
+
`strategy.warm` (per-indicator update counts) or `history_bars_by_session` — never
|
|
727
|
+
`bars_seen >= history_bars`, which cannot express two cadences and OVERSTATES warmth.
|
|
728
|
+
`sequential_combines` still SIZES its preload slice from the unscoped `history_bars`, so a
|
|
729
|
+
scoped strategy will honestly report `fully_warm=False` rather than silently lying.
|
|
730
|
+
- **A `Cross` inherits its inputs' scope** and refuses a conflicting `session=`. Mixed-scope
|
|
731
|
+
inputs are refused outright: they advance on different bars, so comparing them compares values
|
|
732
|
+
sampled at unrelated instants.
|
|
733
|
+
- **Sessions are defined in their own timezone, not as fixed ET offsets.** DST comes from the
|
|
734
|
+
IANA database, so nothing rots — London and New York switch on different dates, and the London
|
|
735
|
+
window really is 04:00 ET rather than 03:00 for ~3 weeks each spring and ~1 each autumn. A
|
|
736
|
+
fixed ET block is still one line: `Session("LONDON_ET", ET, time(3), time(11))`.
|
|
737
|
+
- **Session membership is a DIFFERENT axis from `trading_day_of()`.** Asia sits after the 18:00
|
|
738
|
+
ET rollover, so its bars belong to the NEXT trading day. Membership is tested on `ts_event`
|
|
739
|
+
(the bar's OPEN) over a half-open `[start, end)` window, so the 09:29→09:30 bar — every print
|
|
740
|
+
of it pre-market — is not New York.
|
|
741
|
+
- **`trade_sessions` only narrows what THIS strategy does.** It never widens what the venue
|
|
742
|
+
permits: the 16:10 ET flatten and the 16:10–18:00 no-trade window apply either way.
|
|
743
|
+
- **You need 24h data.** The shipped `data/sample_mnq_1m.csv` is RTH-only (390 bars/day,
|
|
744
|
+
09:30–15:59 ET) and contains no Asia or London bars, so nothing session-scoped is observable on
|
|
745
|
+
it. For synthetic bars pass `synthetic_bars(..., hours="globex")`; the default `"rth"` mode is
|
|
746
|
+
the New York session and nothing else.
|
|
747
|
+
- **Not built: per-session performance attribution.** Nothing in `SummaryStats` splits P&L by
|
|
748
|
+
session, so scoping is currently a modelling choice you make, not one the report scores.
|
|
749
|
+
|
|
659
750
|
## 6. The end-to-end workflow
|
|
660
751
|
|
|
661
752
|
What using this framework actually looks like, in the order you do it:
|
|
662
753
|
|
|
663
754
|
```python
|
|
664
755
|
from topstep_backtest import AccountSize, Backtest
|
|
665
|
-
from topstep_backtest.metrics import
|
|
666
|
-
from topstep_backtest.rules.kernel import Verdict
|
|
756
|
+
from topstep_backtest.metrics import crosscheck, mc_confidence, sequential_combines
|
|
667
757
|
from topstep_backtest.rules.params import combine_params
|
|
668
758
|
|
|
669
|
-
# 1. Write the strategy (§2), then run ONE backtest
|
|
670
|
-
|
|
759
|
+
# 1. Write the strategy (§2), then run ONE backtest — recorded, so the replay
|
|
760
|
+
# tab can show intent against execution.
|
|
761
|
+
report = Backtest(bars, MyStrategy(CONTRACT), account=AccountSize.S50K, record=True).run()
|
|
671
762
|
print(report) # verdict, day trail, four statistics blocks
|
|
672
763
|
|
|
673
764
|
# 2. Sanity-check the wiring BEFORE reading any number.
|
|
674
765
|
assert not report.result.rejections # zero trades + rejections = broken wiring (§5.5)
|
|
675
766
|
assert report.bars_gated != len(bars) # warmup longer than your data
|
|
676
767
|
|
|
677
|
-
# 3. Read the run, minding the basis (§5.6)
|
|
768
|
+
# 3. Read the run, minding the basis (§5.6) — and step the replay tab to each
|
|
769
|
+
# entry: the bracket where you meant it, the notes matching the fills.
|
|
678
770
|
s = report.stats
|
|
679
771
|
if s.provisional:
|
|
680
772
|
... # under 200 closes: estimates, not findings
|
|
@@ -684,16 +776,19 @@ s.drawdown.eod_trailing # Topstep's actual MLL mechanic
|
|
|
684
776
|
s.drawdown.min_floor_headroom # closest the account came to death
|
|
685
777
|
s.daily.p05 # the bad day to size against
|
|
686
778
|
|
|
687
|
-
# 4. Stop trusting one sample
|
|
688
|
-
|
|
689
|
-
|
|
690
|
-
|
|
691
|
-
)
|
|
779
|
+
# 4. Stop trusting one sample — real attempts first, resampled ones second.
|
|
780
|
+
sweep = sequential_combines(bars, lambda: MyStrategy(CONTRACT), window_days=21)
|
|
781
|
+
c = mc_confidence(report.result, params=combine_params(AccountSize.S50K), seed=7)
|
|
782
|
+
check = crosscheck(c.mc, sweep) # same billing-month horizon on both sides (§5.7)
|
|
692
783
|
|
|
693
|
-
# 5. Act on the AUTOPSY,
|
|
784
|
+
# 5. Act on the AUTOPSY, and quote the CI, never the bare point (§5.7).
|
|
694
785
|
# mll_breach -> resize
|
|
695
786
|
# consistency_blocked -> throttle the outsized day; the edge is fine
|
|
696
787
|
# target_not_reached -> the edge is too slow; nothing risk-side helps
|
|
788
|
+
c.ci.p05, c.ci.p95 # the error bar the source-day count earns
|
|
789
|
+
c.sensitivity.spread # wide = the streak assumption is doing the work
|
|
790
|
+
check.divergent # bootstrap vs real windows: disagreement is the finding
|
|
791
|
+
report.to_html("run.html", confidence=c, crosscheck=check) # the cards, archived
|
|
697
792
|
|
|
698
793
|
# 6. Was it an edge, or did you search until something looked good? The trial
|
|
699
794
|
# count must be RECORDED — a remembered one is always too low, because the
|
|
@@ -706,7 +801,7 @@ dsr.expected_max_sharpe > dsr.sharpe # the search alone explains the result
|
|
|
706
801
|
|
|
707
802
|
# 7. Is the attempt worth its price? Every figure is yours; none is baked in.
|
|
708
803
|
ev = evaluate_ev(
|
|
709
|
-
mc,
|
|
804
|
+
c.mc,
|
|
710
805
|
economics=EvalEconomics(
|
|
711
806
|
monthly_fee=Decimal("149"), pass_value=Decimal("2000"), reset_fee=Decimal("99")
|
|
712
807
|
),
|
|
@@ -724,9 +819,10 @@ that. **EV's `pass_value` is an assumption you supply and this package cannot ch
|
|
|
724
819
|
dominates the answer, so quote `breakeven_pass_value` ("a pass must be worth at least $X")
|
|
725
820
|
rather than `ev` unless you can defend the input.
|
|
726
821
|
|
|
727
|
-
Runnable end to end: `examples/run_combine.py` (steps 1–3) and
|
|
728
|
-
`examples/run_montecarlo.py` (steps 4–5, at two position sizes so the autopsy
|
|
729
|
-
discriminates). `docs/TUTORIAL_EMA_CROSSOVER.md` walks all of it line by line
|
|
822
|
+
Runnable end to end: `examples/run_combine.py` (steps 1–3), `examples/run_windows.py` and
|
|
823
|
+
`examples/run_montecarlo.py` (steps 4–5, the latter at two position sizes so the autopsy
|
|
824
|
+
visibly discriminates). `docs/TUTORIAL_EMA_CROSSOVER.md` walks all of it line by line, and
|
|
825
|
+
`website/workflow.md` is the narrative version with the gate each stage must pass.
|
|
730
826
|
|
|
731
827
|
## 7. What `Backtest()` refuses, and what to use instead
|
|
732
828
|
|
|
@@ -844,8 +940,8 @@ moving one toward optimism is a probe to report next to the baseline, not a new
|
|
|
844
940
|
`docs/ROADMAP.md` is the authority. These names appear in older notes and other frameworks and
|
|
845
941
|
**do not exist here**: `LiveBroker`, `RecordingLiveBroker`, `DataEngine`, `TrailingMaxLossLimit`,
|
|
846
942
|
`prob_fill_on_limit` (the real field is `BarFillConfig.fill_limit_on_touch`). Also absent: a
|
|
847
|
-
Parquet/Arrow data catalog, multi-timeframe resampling, a MessageBus, per-year / per-regime
|
|
848
|
-
breakdowns, higher fill tiers, an XFA rule set, and any holiday calendar.
|
|
943
|
+
Parquet/Arrow data catalog, multi-timeframe resampling, a MessageBus, per-year / per-regime /
|
|
944
|
+
per-session breakdowns (§5.11), higher fill tiers, an XFA rule set, and any holiday calendar.
|
|
849
945
|
|
|
850
946
|
**Do** write code against these — they used to be on the list above and now ship: Monte-Carlo
|
|
851
947
|
(§5.7), `optimize()` and the overfitting guards (§5.9), the sequential-Combine sweep (§5.8),
|
|
@@ -4,6 +4,154 @@ All notable changes to this project are documented here. This project adheres to
|
|
|
4
4
|
[Semantic Versioning](https://semver.org/spec/v2.0.0.html). While the version is
|
|
5
5
|
below 1.0, minor releases may contain breaking changes.
|
|
6
6
|
|
|
7
|
+
## [0.4.0] — 2026-08-25
|
|
8
|
+
|
|
9
|
+
A minor: new features throughout. One behavioural default changed — the Monte-Carlo
|
|
10
|
+
horizon now defaults to one billing month rather than the observed day count.
|
|
11
|
+
|
|
12
|
+
### Added
|
|
13
|
+
|
|
14
|
+
- **Regional sessions** (`core/sessions.py`: `Session`, `ASIA`, `LONDON`, `NEW_YORK`) and the two
|
|
15
|
+
independent switches that use them. `use(indicator, session=...)` scopes an indicator's INPUT
|
|
16
|
+
DATA; `SymbolStrategy(trade_sessions=...)` scopes DECISIONS. Keeping them separate is the
|
|
17
|
+
whole feature — a strategy can hold a continuous 24h `Ema` beside an NY-only `Atr` and still
|
|
18
|
+
trade only New York. Indicators advance regardless of `trade_sessions`, because starving one
|
|
19
|
+
outside the tradable window would leave it with gaps and a different value than the same
|
|
20
|
+
indicator on the same tape. Both default to today's behaviour and an unscoped run is
|
|
21
|
+
byte-identical.
|
|
22
|
+
|
|
23
|
+
Why you would scope at all: dispersion measures (`Atr`, `StdDev`, `Rsi`) describe how much
|
|
24
|
+
price moves *per bar*, and that is session-dependent — on 24h bars an `Atr(14)` read at 09:30
|
|
25
|
+
ET is computed almost entirely from thin pre-market bars, understating NY volatility exactly
|
|
26
|
+
when stop distance is being sized. Level measures (`Sma`, `Ema`) answer where price *is*, and
|
|
27
|
+
the overnight move is real: an NY-only `Ema` is anchored to yesterday's 16:00 close and blind
|
|
28
|
+
to a London repricing. Levels are continuous across sessions; dispersion is not.
|
|
29
|
+
|
|
30
|
+
**Sessions are defined in their own local timezone**, not as fixed ET offsets, so DST comes
|
|
31
|
+
from the IANA database and there is no hand-maintained table to rot — the same reasoning that
|
|
32
|
+
dropped the exchange calendar. London and New York switch on different dates, so the London
|
|
33
|
+
window really is 04:00 ET rather than 03:00 for about three weeks a year, and the tests assert
|
|
34
|
+
that against concrete 2026 dates. A fixed-ET block remains one line away
|
|
35
|
+
(`Session("LONDON_ET", ET, time(3), time(11))`), midnight-wrapping included.
|
|
36
|
+
|
|
37
|
+
Session membership is a **different axis** from `trading_day_of()`: Asia sits after the 18:00
|
|
38
|
+
ET rollover, so its bars belong to the *next* trading day. Membership is tested on `ts_event`
|
|
39
|
+
(the bar's open) over a half-open window, so the 09:29->09:30 bar — every print of it
|
|
40
|
+
pre-market — is not New York.
|
|
41
|
+
|
|
42
|
+
- `SymbolStrategy.warm`, `.registrations`, `.history_bars_by_session`, `.bars_out_of_session`,
|
|
43
|
+
`.bars_seen`, `.in_trade_session()` — the diagnostics scoping makes necessary. A scoped
|
|
44
|
+
indicator warms in its OWN cadence, so it needs `history_bars` bars *of its session*: on
|
|
45
|
+
5-minute bars an NY-scoped `Sma(30)` spans ~25 trading days against ~7 for the same indicator
|
|
46
|
+
on a 24h feed. `warm` counts each indicator's own updates, which a single bar count cannot
|
|
47
|
+
express, and `sequential_combines` now asks the strategy rather than comparing counts — the
|
|
48
|
+
old `preload >= needed` silently OVERSTATED warmth for a scoped indicator.
|
|
49
|
+
|
|
50
|
+
- `synthetic_bars(hours="globex")` — the full 23-hour electronic session (18:00 ET previous
|
|
51
|
+
calendar day to 17:00 ET), reusing the canonical open from `core/time.py`. The default
|
|
52
|
+
`"rth"` mode is unchanged. Without this nothing session-scoped is testable: the shipped
|
|
53
|
+
`data/sample_mnq_1m.csv` is RTH-only (exactly 390 bars/day, 09:30-15:59 ET) and contains no
|
|
54
|
+
Asia or London bars at all.
|
|
55
|
+
|
|
56
|
+
- A `Cross` now **inherits its inputs' data scope** and refuses a conflicting one, alongside the
|
|
57
|
+
existing refusal of unregistered inputs. Crossing a 24h `Ema` against an NY-scoped `Ema`
|
|
58
|
+
compares values sampled on unrelated cadences; leaving the `Cross` continuous over scoped
|
|
59
|
+
inputs would have it re-read unchanged values on most bars. Both were silent before.
|
|
60
|
+
|
|
61
|
+
- **`metrics.spaced_combines` — pick how many Combine attempts to simulate.**
|
|
62
|
+
Where `sequential_combines` lets the tape dictate the attempt count (disjoint windows),
|
|
63
|
+
the new sweep takes `periods` and `window_days` and spreads that many start days evenly
|
|
64
|
+
across the tape, overlapping as much as the arithmetic requires — 100 days, 10 periods
|
|
65
|
+
of 40 days starts an attempt roughly every week. Each period is still a completely fresh
|
|
66
|
+
Combine (same shared implementation: day-aligned slicing, prewarmed indicators, the
|
|
67
|
+
kernel's own verdict, `classify_failure` attribution), so the sweep measures how much
|
|
68
|
+
passing depends on *when* the attempt starts. Because overlapping periods are not
|
|
69
|
+
independent samples, the result carries `effective_independent_windows`,
|
|
70
|
+
`stride_days` and `overlap_fraction` alongside its rates, and asking for more periods
|
|
71
|
+
than the tape has distinct start days is refused rather than replaying identical
|
|
72
|
+
windows as new observations. `examples/run_spaced.py` renders a sweep as a text
|
|
73
|
+
report — the rates with their effective-sample caveat in the header, then one line
|
|
74
|
+
per period in start order so start-date sensitivity is visible at a glance.
|
|
75
|
+
|
|
76
|
+
- **An HTML tearsheet for sweeps.** `sweep.to_html(path)` / `sweep.show()` on both
|
|
77
|
+
`WindowSweep` and `SpacedSweep` (or `tearsheet.render_sweep_html`) render one
|
|
78
|
+
self-contained page: a calendar timeline with each attempt drawn over its actual
|
|
79
|
+
dates and stacked into lanes where they overlap — the reason overlapping attempts
|
|
80
|
+
are not independent samples made visible rather than footnoted — the outcome
|
|
81
|
+
autopsy as a stacked bar with each failure mode's prescription, every attempt's
|
|
82
|
+
cumulative P&L overlaid from $0, a closest-approach-to-the-MLL-floor strip (a pass
|
|
83
|
+
with $40 of headroom looks identical to a robust one in the pass rate; not here),
|
|
84
|
+
and a sortable per-attempt table with P&L sparklines, all hover-linked. To feed the
|
|
85
|
+
charts, `WindowResult` now carries `daily_pnl` (and a `cumulative_pnl` property),
|
|
86
|
+
the same per-day series walk-forward's trials keep. No charting library and no
|
|
87
|
+
embedded JSON: the charts are SVG generated in Python, so the render stays a pure
|
|
88
|
+
byte-identical function of the frozen sweep data, and the evidence caveat is
|
|
89
|
+
stamped in the page header where a screenshot cannot shed it.
|
|
90
|
+
|
|
91
|
+
- **Confidence instruments for the Monte-Carlo pass probability**
|
|
92
|
+
(`metrics/confidence.py`). The point estimate's *simulation* error was never the real
|
|
93
|
+
uncertainty; these quantify what is. `pass_probability_ci` double-bootstraps the observed
|
|
94
|
+
day set itself and reports a 5th–95th percentile band — the error bar the source-day
|
|
95
|
+
count earns (a 65% from 500 days and a 65% from 40 now read differently).
|
|
96
|
+
`block_length_sensitivity` re-runs the estimate across block lengths and reports the
|
|
97
|
+
spread: wide means the streak-clustering assumption is doing the work.
|
|
98
|
+
`monte_carlo_by_year` stratifies by calendar year so a hostile year is never averaged
|
|
99
|
+
against a kind one — the spread across strata is the error bar non-stationarity imposes.
|
|
100
|
+
`crosscheck` forces the bootstrap and `sequential_combines` (opposite biases, shared
|
|
101
|
+
`classify_failure` on purpose) to answer side by side, flagging any outcome where the
|
|
102
|
+
real windows sit outside 2×SE of what the bootstrap's own probability would produce;
|
|
103
|
+
when they disagree, the disagreement is the finding. `mc_confidence` bundles the first
|
|
104
|
+
three, and every estimate runs through the same `monte_carlo_from_blocks` core as the
|
|
105
|
+
number it qualifies. All deterministic per seed, like everything else in `metrics/`.
|
|
106
|
+
- **Monte-Carlo cards on the tearsheet.** `report.to_html(path, confidence=...,
|
|
107
|
+
crosscheck=...)` (and `show` / `to_timestamped_html`) render the estimate with its CI,
|
|
108
|
+
the block-sensitivity row, the per-year strata, and the bootstrap-vs-windows card.
|
|
109
|
+
Caller-computed on purpose: writing a file never triggers thousands of simulations as a
|
|
110
|
+
side effect, and the render stays a pure, byte-identical function of its inputs.
|
|
111
|
+
- **The workflow, documented** (`website/workflow.md`, on the site nav as "The workflow").
|
|
112
|
+
The stages in the order the questions become answerable — envelope, data, strategy, one
|
|
113
|
+
recorded run, real attempts, the distribution with its confidence, the search guards,
|
|
114
|
+
the price — each ending with the gate that must pass before the next stage's number
|
|
115
|
+
means anything. `AGENTS.md` §6 is the code-first version and now walks the same loop
|
|
116
|
+
(recorded run, `sequential_combines`, `mc_confidence` + `crosscheck`, guards, EV).
|
|
117
|
+
|
|
118
|
+
- **A native loader for databento-data-playground Parquet exports.**
|
|
119
|
+
`topstep_backtest.data.loaders.load_bars(path)` reads an export produced by that project's
|
|
120
|
+
`convert.py` and returns `(bars, spec, meta)` ready for `Backtest(...)`. The product, bar
|
|
121
|
+
span and timestamp stamping are read from the file's embedded Parquet metadata rather than
|
|
122
|
+
passed in — each is a silent corruption when guessed — and the file's copy of the
|
|
123
|
+
instrument economics is cross-checked against the engine's `InstrumentSpec`, failing loudly
|
|
124
|
+
on any mismatch instead of picking one. Files without the metadata blob are refused.
|
|
125
|
+
Requires the existing `[data]` extra (pandas + pyarrow); nothing new to install.
|
|
126
|
+
`examples/run_real_data.py` auto-detects such exports and loads them with no declaration
|
|
127
|
+
flags (and refuses the flags on one — re-declaring what the file states is the
|
|
128
|
+
contradiction they exist to prevent), and the module joins the site's API reference and
|
|
129
|
+
quickstart.
|
|
130
|
+
|
|
131
|
+
### Changed
|
|
132
|
+
|
|
133
|
+
- **The Monte-Carlo horizon defaults to one billing month** (`BILLING_MONTH_DAYS = 21`),
|
|
134
|
+
not the observed day count. A Combine has no time limit, only a monthly fee, so "one
|
|
135
|
+
attempt" defaults to one fee cycle — the same unit `EvalEconomics` bills in (its
|
|
136
|
+
`trading_days_per_month` default now *is* this constant). The fixed default also
|
|
137
|
+
retires the guard that made `horizon_days` mandatory for a blown source run: the old
|
|
138
|
+
hazard was defaulting to a survival time, and the new default cannot. A failed run
|
|
139
|
+
still reports `source_truncated`, and the survivorship-bias caveat stands unchanged.
|
|
140
|
+
- `sample_day_path`, `nearest_rank`, and `monte_carlo_from_blocks` are module-level
|
|
141
|
+
names in `metrics/montecarlo.py` (previously underscore-private) so the confidence
|
|
142
|
+
instruments run through the identical core; they remain outside `__all__`.
|
|
143
|
+
- **The replay cockpit is easier to read.** The running-stats panel no longer stacks all
|
|
144
|
+
~50 rows: snapshots are regrouped into by-type cards — *status*, *trades (gross)*,
|
|
145
|
+
*round trips (net)*, *risk & quant*, *daily P&L* — shown one at a time behind filter
|
|
146
|
+
chips, with a `•` on any chip whose hidden card the newest snapshot moved. The values are
|
|
147
|
+
the same rows the results tab renders, relocated rather than rebuilt (a stat missing from
|
|
148
|
+
the regrouping map fails loudly instead of silently vanishing from the replay). The event
|
|
149
|
+
log moved from the cramped right column to a full-width strip under the charts — one
|
|
150
|
+
event per line, heading and kind filters on one row — and the panel text was tightened
|
|
151
|
+
throughout. Both tabs' price charts also gained a marker legend (entry/exit arrows, and
|
|
152
|
+
in the replay the cross fires and the working stop/target/avg-entry price lines), so the
|
|
153
|
+
glyphs on the tape no longer require guessing.
|
|
154
|
+
|
|
7
155
|
## [0.3.0] — 2026-08-20
|
|
8
156
|
|
|
9
157
|
A minor: new features, no intended breaking changes.
|
|
@@ -422,5 +570,6 @@ rule and fee constants are **not yet calibrated against a live account**, so a
|
|
|
422
570
|
are not exercised end to end.
|
|
423
571
|
- Tier-0 bar fills only. No quote, depth or MBO tiers.
|
|
424
572
|
|
|
573
|
+
[0.4.0]: https://pypi.org/project/topstep-backtest/0.4.0/
|
|
425
574
|
[0.3.0]: https://pypi.org/project/topstep-backtest/0.3.0/
|
|
426
575
|
[0.1.0]: https://pypi.org/project/topstep-backtest/0.1.0/
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
Metadata-Version: 2.5
|
|
2
2
|
Name: topstep-backtest
|
|
3
|
-
Version: 0.
|
|
3
|
+
Version: 0.4.0
|
|
4
4
|
Summary: Event-driven backtesting framework for Topstep Trading Combine strategies, with backtest/live parity against topstep-sdk.
|
|
5
5
|
Author-email: Tarric Sookdeo <tarricsookdeo@outlook.com>
|
|
6
6
|
License-Expression: MIT
|
|
@@ -84,6 +84,22 @@ indicator formula, so there is no second implementation to drift from the refere
|
|
|
84
84
|
`Cross` is the deliberate exception — TA-Lib has no crossover primitive, so it is a
|
|
85
85
|
framework helper that compares two TA-Lib outputs rather than computing anything.
|
|
86
86
|
|
|
87
|
+
**Sessions are two independent switches.** On a 24h tape you can scope each indicator's input
|
|
88
|
+
data and, separately, restrict when the strategy may trade:
|
|
89
|
+
|
|
90
|
+
```python
|
|
91
|
+
super().__init__(contract_id, trade_sessions=(NEW_YORK,)) # trade New York only
|
|
92
|
+
self.trend = self.use(Ema(50)) # ...but see the whole tape
|
|
93
|
+
self.atr = self.use(Atr(14), session=NEW_YORK) # ...while this sees only NY
|
|
94
|
+
```
|
|
95
|
+
|
|
96
|
+
The `Ema` keeps consuming Asia and London — an indicator fed only the tradable window would
|
|
97
|
+
develop gaps — while every decision happens in New York. Which indicators to scope is a
|
|
98
|
+
modelling choice with a rule behind it: *levels are continuous across sessions, dispersion is
|
|
99
|
+
not*. `ASIA`/`LONDON`/`NEW_YORK` are defined in their own timezones, so daylight saving comes
|
|
100
|
+
from the IANA database rather than a table that rots. See
|
|
101
|
+
[`examples/session_scoped.py`](examples/session_scoped.py).
|
|
102
|
+
|
|
87
103
|
## Install
|
|
88
104
|
|
|
89
105
|
```bash
|
|
@@ -184,8 +200,9 @@ cannot disagree. For a sheet from *every* run without naming a file each time,
|
|
|
184
200
|
both the report and the path.
|
|
185
201
|
|
|
186
202
|
**Replay a run bar by bar.** Pass `record=True` and the tearsheet grows a second tab: a
|
|
187
|
-
replay cockpit that fills the window — its own charts
|
|
188
|
-
running stats
|
|
203
|
+
replay cockpit that fills the window — its own charts beside the settled state and the
|
|
204
|
+
running stats (by-type cards behind filter chips), the event log across the bottom, so
|
|
205
|
+
nothing has to be scrolled between. Step
|
|
189
206
|
through the run (buttons, slider, arrow keys, autoplay) with everything after the cursor
|
|
190
207
|
veiled, and watch the position, working orders (drawn as price lines), balance, floor
|
|
191
208
|
headroom, indicator values — named by the attributes your strategy stores them under — and
|
|
@@ -231,6 +248,15 @@ imply three different fixes:
|
|
|
231
248
|
| `consistency_blocked` | money made in too few days — throttle the outsized day; the edge is fine |
|
|
232
249
|
| `target_not_reached` | the edge is too slow for the window — nothing risk-side helps |
|
|
233
250
|
|
|
251
|
+
The horizon defaults to one billing month (21 sessions) — a Combine has no time limit, only
|
|
252
|
+
a monthly fee, so "one attempt" means one fee cycle, the same unit `evaluate_ev` bills in.
|
|
253
|
+
And never quote the pass probability alone: `metrics.confidence.mc_confidence` attaches the
|
|
254
|
+
error bar the *source-day count* earns (a double-bootstrap CI — the path count was never the
|
|
255
|
+
real uncertainty), a block-length sensitivity row, and per-year strata, while
|
|
256
|
+
`metrics.confidence.crosscheck` compares the bootstrap against `sequential_combines`' real
|
|
257
|
+
windows — when the two disagree, the disagreement is the finding. The whole bundle renders
|
|
258
|
+
on the tearsheet via `report.to_html(path, confidence=..., crosscheck=...)`.
|
|
259
|
+
|
|
234
260
|
```bash
|
|
235
261
|
uv run python examples/run_montecarlo.py # the same edge at two sizes, failing two ways
|
|
236
262
|
```
|
|
@@ -253,6 +279,10 @@ reruns — but read these before trusting a number:
|
|
|
253
279
|
`metrics/economics.py`), but none of them escapes your sample: the Monte-Carlo resamples a
|
|
254
280
|
strategy's own observed days, so it cannot invent a market regime your tape never contained
|
|
255
281
|
and will understate tail risk on a short or single-regime sample.
|
|
282
|
+
- **No per-year, per-regime or per-session performance breakdown.** Sessions scope which bars
|
|
283
|
+
an indicator is computed from and when a strategy may trade, but nothing in `SummaryStats`
|
|
284
|
+
splits P&L by session, so a scoping choice is one you make on reasoning rather than one the
|
|
285
|
+
report scores for you.
|
|
256
286
|
- **No exchange holiday calendar ships with this package.** Bars on market holidays and
|
|
257
287
|
past early-close halts are not detected, flagged or filtered anywhere — filter them
|
|
258
288
|
upstream. (A built-in calendar was removed in 0.1.0: it disagreed with CME on several
|
|
@@ -286,6 +316,7 @@ uv run pytest && uv run ruff check . && uv run ruff format --check . && uv run p
|
|
|
286
316
|
uv run python scripts/gen_api_surface.py --check # AGENTS.md §3 must not be stale
|
|
287
317
|
uv run python examples/run_combine.py # end-to-end combine verdict
|
|
288
318
|
uv run python examples/run_montecarlo.py # outcome distribution + autopsy
|
|
319
|
+
uv run python examples/session_scoped.py # session-scoped indicators + trade gate
|
|
289
320
|
```
|
|
290
321
|
|
|
291
322
|
Those checks are the gate: CI runs all of them on Python 3.12/3.13/3.14, then installs the
|
|
@@ -314,23 +345,29 @@ environment. There is no public issue tracker.
|
|
|
314
345
|
|
|
315
346
|
- **`docs/TUTORIAL_EMA_CROSSOVER.md` — start here**: one strategy end to end, raw candles
|
|
316
347
|
to Combine verdict.
|
|
348
|
+
- `website/workflow.md` — **the workflow**: the stages in the order the questions become
|
|
349
|
+
answerable — data, mechanics, real attempts, the distribution with its confidence, the
|
|
350
|
+
search guards, the price — and the gate each must pass before the next number means
|
|
351
|
+
anything.
|
|
317
352
|
- `website/results.md` — **reading the report**: what every figure means, what basis it is
|
|
318
353
|
on, which ones mislead alone, and what this framework deliberately does *not* report.
|
|
319
354
|
- `AGENTS.md` — dense reference for *using* the framework: strategy dialect, module map, the
|
|
320
|
-
invariants you must not break, and **§6, the end-to-end workflow** (run → check
|
|
321
|
-
→ read the run minding each metric's basis → Monte-Carlo
|
|
355
|
+
invariants you must not break, and **§6, the end-to-end workflow** (run recorded → check
|
|
356
|
+
the wiring → read the run minding each metric's basis → real windows + Monte-Carlo with
|
|
357
|
+
its CI → act on the autopsy → deflate the search → price the attempt).
|
|
322
358
|
- `docs/INDICATORS.md` — the TA-Lib indicator surface, wrapper by wrapper.
|
|
323
359
|
- `docs/topstep-rules.md` — the rulebook being enforced, with sources, confidence levels,
|
|
324
360
|
and a verify-before-trusting checklist.
|
|
325
361
|
- `docs/DESIGN.md` — architecture contract, for *modifying* the framework.
|
|
326
362
|
- `docs/ROADMAP.md` — what is built, partial, and not started. Next: **live adapter +
|
|
327
|
-
calibration** → per-
|
|
328
|
-
(XFA) modeling is deliberately parked.
|
|
363
|
+
calibration** → per-regime (non-calendar) breakdowns → L1/L2/MBO fill tiers;
|
|
364
|
+
funded-account (XFA) modeling is deliberately parked.
|
|
329
365
|
- `examples/` — runnable: `run_real_data.py` (your CSV/Parquet → verdict), `run_combine.py`
|
|
330
366
|
(synthetic end to end), `run_tearsheet.py` (the same run as one HTML file),
|
|
331
367
|
`run_replay.py` (a recorded run with the bar-by-bar scrubber), `run_montecarlo.py`
|
|
332
368
|
(outcome distribution + autopsy), `run_windows.py` (a long tape replayed as consecutive
|
|
333
|
-
independent Combine attempts), `
|
|
369
|
+
independent Combine attempts), `run_spaced.py` (a chosen number of attempts, start days
|
|
370
|
+
spread evenly, overlap allowed), `ema_cross.py`, `sma_cross.py`, `talib_macd.py`,
|
|
334
371
|
`hand_wired.py` (what the facade assembles).
|
|
335
372
|
|
|
336
373
|
## Stack
|