topstep-backtest 0.3.0__tar.gz → 0.4.0__tar.gz

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (135) hide show
  1. {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/.gitignore +2 -0
  2. {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/AGENTS.md +122 -26
  3. {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/CHANGELOG.md +149 -0
  4. {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/PKG-INFO +45 -8
  5. {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/README.md +44 -7
  6. {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/docs/DESIGN.md +40 -1
  7. {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/docs/INDICATORS.md +61 -0
  8. {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/docs/ROADMAP.md +6 -4
  9. {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/docs/TUTORIAL_EMA_CROSSOVER.md +45 -6
  10. {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/examples/run_montecarlo.py +30 -7
  11. {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/examples/run_real_data.py +81 -4
  12. topstep_backtest-0.4.0/examples/run_spaced.py +125 -0
  13. {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/examples/run_windows.py +71 -8
  14. topstep_backtest-0.4.0/examples/session_scoped.py +151 -0
  15. {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/src/topstep_backtest/__init__.py +6 -0
  16. topstep_backtest-0.4.0/src/topstep_backtest/core/sessions.py +143 -0
  17. topstep_backtest-0.4.0/src/topstep_backtest/data/loaders.py +117 -0
  18. {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/src/topstep_backtest/data/synthetic.py +57 -21
  19. {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/src/topstep_backtest/harness.py +39 -7
  20. {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/src/topstep_backtest/metrics/__init__.py +31 -1
  21. topstep_backtest-0.4.0/src/topstep_backtest/metrics/confidence.py +536 -0
  22. {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/src/topstep_backtest/metrics/economics.py +6 -3
  23. {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/src/topstep_backtest/metrics/montecarlo.py +65 -34
  24. topstep_backtest-0.4.0/src/topstep_backtest/metrics/windows.py +629 -0
  25. {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/src/topstep_backtest/strategy/symbol.py +214 -22
  26. {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/src/topstep_backtest/tearsheet/__init__.py +282 -17
  27. topstep_backtest-0.4.0/src/topstep_backtest/tearsheet/_assets/sweep.css +154 -0
  28. topstep_backtest-0.4.0/src/topstep_backtest/tearsheet/_assets/sweep.js +42 -0
  29. {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/src/topstep_backtest/tearsheet/_assets/tearsheet.css +42 -22
  30. {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/src/topstep_backtest/tearsheet/_assets/tearsheet.js +94 -27
  31. topstep_backtest-0.4.0/src/topstep_backtest/tearsheet/sweep.py +522 -0
  32. topstep_backtest-0.4.0/tests/unit/test_confidence.py +342 -0
  33. topstep_backtest-0.4.0/tests/unit/test_loaders.py +176 -0
  34. {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/tests/unit/test_montecarlo.py +17 -5
  35. topstep_backtest-0.4.0/tests/unit/test_sessions.py +138 -0
  36. topstep_backtest-0.4.0/tests/unit/test_sweep_tearsheet.py +131 -0
  37. {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/tests/unit/test_symbol_strategy.py +216 -2
  38. {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/tests/unit/test_synthetic.py +75 -0
  39. {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/tests/unit/test_tearsheet.py +77 -0
  40. {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/tests/unit/test_windows.py +103 -1
  41. topstep_backtest-0.3.0/src/topstep_backtest/metrics/windows.py +0 -314
  42. {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/LICENSE +0 -0
  43. {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/data/sample_mnq_1m.csv +0 -0
  44. {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/docs/topstep-rules.md +0 -0
  45. {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/examples/ema_cross.py +0 -0
  46. {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/examples/hand_wired.py +0 -0
  47. {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/examples/run_combine.py +0 -0
  48. {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/examples/run_replay.py +0 -0
  49. {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/examples/run_tearsheet.py +0 -0
  50. {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/examples/sma_cross.py +0 -0
  51. {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/examples/talib_macd.py +0 -0
  52. {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/pyproject.toml +0 -0
  53. {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/src/topstep_backtest/_render.py +0 -0
  54. {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/src/topstep_backtest/clock/__init__.py +0 -0
  55. {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/src/topstep_backtest/clock/live_clock.py +0 -0
  56. {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/src/topstep_backtest/clock/test_clock.py +0 -0
  57. {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/src/topstep_backtest/core/__init__.py +0 -0
  58. {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/src/topstep_backtest/core/ids.py +0 -0
  59. {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/src/topstep_backtest/core/instruments.py +0 -0
  60. {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/src/topstep_backtest/core/money.py +0 -0
  61. {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/src/topstep_backtest/core/time.py +0 -0
  62. {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/src/topstep_backtest/data/__init__.py +0 -0
  63. {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/src/topstep_backtest/data/clean.py +0 -0
  64. {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/src/topstep_backtest/data/continuous.py +0 -0
  65. {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/src/topstep_backtest/data/feed.py +0 -0
  66. {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/src/topstep_backtest/data/validator.py +0 -0
  67. {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/src/topstep_backtest/data/wrangler.py +0 -0
  68. {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/src/topstep_backtest/engine/__init__.py +0 -0
  69. {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/src/topstep_backtest/engine/backtest.py +0 -0
  70. {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/src/topstep_backtest/execution/__init__.py +0 -0
  71. {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/src/topstep_backtest/execution/rejections.py +0 -0
  72. {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/src/topstep_backtest/execution/sim_broker.py +0 -0
  73. {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/src/topstep_backtest/fills/__init__.py +0 -0
  74. {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/src/topstep_backtest/fills/bar_fill.py +0 -0
  75. {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/src/topstep_backtest/fills/fees.py +0 -0
  76. {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/src/topstep_backtest/fills/path.py +0 -0
  77. {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/src/topstep_backtest/indicators/__init__.py +0 -0
  78. {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/src/topstep_backtest/indicators/base.py +0 -0
  79. {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/src/topstep_backtest/indicators/library.py +0 -0
  80. {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/src/topstep_backtest/indicators/talib_adapter.py +0 -0
  81. {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/src/topstep_backtest/metrics/overfitting.py +0 -0
  82. {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/src/topstep_backtest/metrics/stats.py +0 -0
  83. {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/src/topstep_backtest/metrics/walkforward.py +0 -0
  84. {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/src/topstep_backtest/protocols.py +0 -0
  85. {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/src/topstep_backtest/py.typed +0 -0
  86. {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/src/topstep_backtest/replay.py +0 -0
  87. {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/src/topstep_backtest/rules/__init__.py +0 -0
  88. {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/src/topstep_backtest/rules/kernel.py +0 -0
  89. {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/src/topstep_backtest/rules/params.py +0 -0
  90. {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/src/topstep_backtest/strategy/__init__.py +0 -0
  91. {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/src/topstep_backtest/strategy/base.py +0 -0
  92. {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/src/topstep_backtest/strategy/tracker.py +0 -0
  93. {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/src/topstep_backtest/tearsheet/_assets/lightweight-charts.LICENSE +0 -0
  94. {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/src/topstep_backtest/tearsheet/_assets/lightweight-charts.standalone.production.js +0 -0
  95. {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/tests/__init__.py +0 -0
  96. {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/tests/conftest.py +0 -0
  97. {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/tests/golden/__init__.py +0 -0
  98. {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/tests/golden/artifacts/verdict_failed_mll_s50k.json +0 -0
  99. {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/tests/golden/artifacts/verdict_passed_s50k.json +0 -0
  100. {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/tests/golden/test_combine_kernel.py +0 -0
  101. {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/tests/golden/test_facade_equivalence.py +0 -0
  102. {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/tests/golden/test_replay_goldens.py +0 -0
  103. {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/tests/golden/test_sugar_equivalence.py +0 -0
  104. {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/tests/golden/test_verdict_goldens.py +0 -0
  105. {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/tests/parity/__init__.py +0 -0
  106. {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/tests/parity/test_broker_conformance.py +0 -0
  107. {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/tests/property/__init__.py +0 -0
  108. {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/tests/property/test_indicator_props.py +0 -0
  109. {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/tests/property/test_kernel_props.py +0 -0
  110. {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/tests/property/test_money_props.py +0 -0
  111. {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/tests/property/test_replay_props.py +0 -0
  112. {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/tests/unit/__init__.py +0 -0
  113. {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/tests/unit/test_bar_fill.py +0 -0
  114. {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/tests/unit/test_clean.py +0 -0
  115. {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/tests/unit/test_clock.py +0 -0
  116. {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/tests/unit/test_continuous.py +0 -0
  117. {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/tests/unit/test_data_feed.py +0 -0
  118. {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/tests/unit/test_economics.py +0 -0
  119. {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/tests/unit/test_engine.py +0 -0
  120. {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/tests/unit/test_fees.py +0 -0
  121. {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/tests/unit/test_harness.py +0 -0
  122. {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/tests/unit/test_indicators.py +0 -0
  123. {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/tests/unit/test_instruments.py +0 -0
  124. {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/tests/unit/test_overfitting.py +0 -0
  125. {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/tests/unit/test_path.py +0 -0
  126. {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/tests/unit/test_replay.py +0 -0
  127. {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/tests/unit/test_replay_tearsheet.py +0 -0
  128. {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/tests/unit/test_sim_broker.py +0 -0
  129. {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/tests/unit/test_stats.py +0 -0
  130. {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/tests/unit/test_talib_adapter_hardening.py +0 -0
  131. {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/tests/unit/test_time.py +0 -0
  132. {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/tests/unit/test_tracker.py +0 -0
  133. {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/tests/unit/test_validator.py +0 -0
  134. {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/tests/unit/test_walkforward.py +0 -0
  135. {topstep_backtest-0.3.0 → topstep_backtest-0.4.0}/tests/unit/test_wrangler.py +0 -0
@@ -45,3 +45,5 @@ site/
45
45
  # Tearsheet output from examples/run_tearsheet.py / run_replay.py (written to the CWD)
46
46
  tearsheet.html
47
47
  replay-tearsheet.html
48
+ mc-confidence-tearsheet.html
49
+ spaced-tearsheet.html
@@ -127,7 +127,7 @@ Generated from the live objects — every signature below is real.
127
127
 
128
128
  | Name | Signature | What it does |
129
129
  |---|---|---|
130
- | `SymbolStrategy` | `(contract_id: 'str', require_ready: 'bool' = True, warmup: 'int \| None' = None)` | Base for strategies trading exactly one contract. |
130
+ | `SymbolStrategy` | `(contract_id: 'str', require_ready: 'bool' = True, warmup: 'int \| None' = None, trade_sessions: 'Sequence[Session] \| None' = None)` | Base for strategies trading exactly one contract. |
131
131
 
132
132
  **Position / order views**
133
133
 
@@ -144,11 +144,17 @@ Generated from the live objects — every signature below is real.
144
144
  | `bars_from_dataframe` | `(df: 'Any', *, contract_id: 'str', spec: 'InstrumentSpec', unit: 'AggregateBarUnit', unit_number: 'int', stamp: "Literal['open', 'close']") -> 'tuple[Bar, ...]'` | Build ``Bar`` objects from a pandas DataFrame of OHLCV candles. |
145
145
  | `bars_from_records` | `(rows: 'Iterable[tuple[object, ...]]', *, contract_id: 'str', spec: 'InstrumentSpec', unit: 'AggregateBarUnit', unit_number: 'int', stamp: "Literal['open', 'close']") -> 'tuple[Bar, ...]'` | Build tick-grid-validated ``Bar`` objects from ``(ts, o, h, l, c, v)`` rows. |
146
146
 
147
+ **Data in (Parquet export)**
148
+
149
+ | Name | Signature | What it does |
150
+ |---|---|---|
151
+ | `load_bars` | `(path: 'str \| Path') -> 'tuple[tuple[Bar, ...], InstrumentSpec, dict[str, Any]]'` | Read the export -> (bars, spec, metadata), ready for ``Backtest(...)``. |
152
+
147
153
  **Data in (synthetic)**
148
154
 
149
155
  | Name | Signature | What it does |
150
156
  |---|---|---|
151
- | `synthetic_bars` | `(*, contract_id: 'str', spec: 'InstrumentSpec', start_day: 'date', days: 'int', seed: 'int', start_price: 'Decimal', bars_per_day: 'int' = 390, unit: 'AggregateBarUnit' = <AggregateBarUnit.MINUTE: 2>, unit_number: 'int' = 1, drift_ticks_per_day: 'int' = 0, vol_ticks: 'int' = 8) -> 'tuple[Bar, ...]'` | Generate ``days`` trading sessions of consistent, on-grid OHLCV bars. |
157
+ | `synthetic_bars` | `(*, contract_id: 'str', spec: 'InstrumentSpec', start_day: 'date', days: 'int', seed: 'int', start_price: 'Decimal', bars_per_day: 'int \| None' = None, unit: 'AggregateBarUnit' = <AggregateBarUnit.MINUTE: 2>, unit_number: 'int' = 1, drift_ticks_per_day: 'int' = 0, vol_ticks: 'int' = 8, hours: 'Hours' = 'rth') -> 'tuple[Bar, ...]'` | Generate ``days`` trading sessions of consistent, on-grid OHLCV bars. |
152
158
 
153
159
  **Instruments**
154
160
 
@@ -158,6 +164,15 @@ Generated from the live objects — every signature below is real.
158
164
  | `symbol_of_contract_id` | `(contract_id: 'str') -> 'str'` | Extract the product symbol from a gateway contract id. |
159
165
  | `InstrumentSpec` | `(*args, **kwargs)` | Frozen per-product economics and session metadata. |
160
166
 
167
+ **Sessions**
168
+
169
+ | Name | Signature | What it does |
170
+ |---|---|---|
171
+ | `Session` | `(*args, **kwargs)` | A named intraday window, defined in its own local timezone. |
172
+ | `ASIA` | `Session(name='ASIA', tz='Asia/Tokyo', start=09:00:00, end=15:00:00)` | A named intraday window, defined in its own local timezone. |
173
+ | `LONDON` | `Session(name='LONDON', tz='Europe/London', start=08:00:00, end=16:30:00)` | A named intraday window, defined in its own local timezone. |
174
+ | `NEW_YORK` | `Session(name='NEW_YORK', tz='America/New_York', start=09:30:00, end=16:00:00)` | A named intraday window, defined in its own local timezone. |
175
+
161
176
  **Results**
162
177
 
163
178
  | Name | Signature | What it does |
@@ -170,7 +185,8 @@ Generated from the live objects — every signature below is real.
170
185
 
171
186
  | Name | Signature | What it does |
172
187
  |---|---|---|
173
- | `render_html` | `(report: 'Report', *, replay: 'ReplaySpec' = 'auto') -> 'str'` | Render ``report`` as one self-contained interactive HTML document. |
188
+ | `render_html` | `(report: 'Report', *, replay: 'ReplaySpec' = 'auto', confidence: 'MonteCarloConfidence \| None' = None, crosscheck: 'CrossCheck \| None' = None) -> 'str'` | Render ``report`` as one self-contained interactive HTML document. |
189
+ | `render_sweep_html` | `(sweep: 'WindowSweep \| SpacedSweep') -> 'str'` | Render a sweep as one self-contained HTML document. |
174
190
 
175
191
  **Replay recording**
176
192
 
@@ -199,12 +215,29 @@ Generated from the live objects — every signature below is real.
199
215
  | `MonteCarloResult` | `(*args, **kwargs)` | Outcome distribution over ``paths`` synthetic Combine attempts. |
200
216
  | `FailureMode` | `FailureMode.MLL_BREACH \| FailureMode.CONSISTENCY_BLOCKED \| FailureMode.TARGET_NOT_REACHED` | Why a simulated path did not pass. Ordered by when it is decided. |
201
217
 
218
+ **Monte-Carlo confidence**
219
+
220
+ | Name | Signature | What it does |
221
+ |---|---|---|
222
+ | `mc_confidence` | `(result: 'BacktestResult', *, params: 'CombineParams', paths: 'int' = 2000, horizon_days: 'int \| None' = None, block_length: 'int' = 5, seed: 'int' = 0, outer: 'int' = 200, inner_paths: 'int' = 200, lengths: 'Sequence[int]' = (1, 5, 10, 20)) -> 'MonteCarloConfidence'` | One call: the estimate plus its CI, sensitivity row, and year strata. |
223
+ | `MonteCarloConfidence` | `(*args, **kwargs)` | The point estimate and every qualifier this module can attach to it. |
224
+ | `pass_probability_ci` | `(result: 'BacktestResult', *, params: 'CombineParams', paths: 'int' = 2000, horizon_days: 'int \| None' = None, block_length: 'int' = 5, seed: 'int' = 0, outer: 'int' = 200, inner_paths: 'int' = 200) -> 'PassProbabilityCI'` | Double bootstrap: a confidence band for the pass probability. |
225
+ | `PassProbabilityCI` | `(*args, **kwargs)` | A pass probability with the error bar its sample size actually earns. |
226
+ | `block_length_sensitivity` | `(result: 'BacktestResult', *, params: 'CombineParams', lengths: 'Sequence[int]' = (1, 5, 10, 20), paths: 'int' = 1000, horizon_days: 'int \| None' = None, seed: 'int' = 0) -> 'BlockLengthSensitivity'` | Re-run the Monte Carlo across block lengths and report the swing. |
227
+ | `BlockLengthSensitivity` | `(*args, **kwargs)` | The same estimate at several block lengths, plus how far it moved. |
228
+ | `monte_carlo_by_year` | `(result: 'BacktestResult', *, params: 'CombineParams', paths: 'int' = 1000, horizon_days: 'int \| None' = None, block_length: 'int' = 5, seed: 'int' = 0) -> 'YearStratification'` | One Monte Carlo per calendar year of the source run. |
229
+ | `YearStratification` | `(*args, **kwargs)` | Per-year estimates, ascending by year. |
230
+ | `crosscheck` | `(mc: 'MonteCarloResult', sweep: 'WindowSweep') -> 'CrossCheck'` | Compare a Monte Carlo against a window sweep, outcome by outcome. |
231
+ | `CrossCheck` | `(*args, **kwargs)` | The bootstrap and the window sweep, forced to answer side by side. |
232
+
202
233
  **Sequential Combines**
203
234
 
204
235
  | Name | Signature | What it does |
205
236
  |---|---|---|
206
237
  | `sequential_combines` | `(bars: 'Sequence[Bar]', factory: 'Callable[[], Strategy]', *, window_days: 'int', account: 'AccountSize' = <AccountSize.S50K: '50K'>, dll_enabled: 'bool' = False, warm_start: 'bool' = True, validate: 'bool' = True) -> 'WindowSweep'` | Run one fresh Combine per non-overlapping ``window_days``-day window. |
238
+ | `spaced_combines` | `(bars: 'Sequence[Bar]', factory: 'Callable[[], Strategy]', *, window_days: 'int', periods: 'int', account: 'AccountSize' = <AccountSize.S50K: '50K'>, dll_enabled: 'bool' = False, warm_start: 'bool' = True, validate: 'bool' = True) -> 'SpacedSweep'` | Run ``periods`` fresh Combines with start days spread evenly over the tape. |
207
239
  | `WindowSweep` | `(*args, **kwargs)` | Every window's attempt, plus the rates over them. |
240
+ | `SpacedSweep` | `(*args, **kwargs)` | A requested number of periods, start days spread evenly, overlap allowed. |
208
241
  | `WindowResult` | `(*args, **kwargs)` | One window's Combine attempt, resolved on its own merits. |
209
242
 
210
243
  **Overfitting guards**
@@ -249,7 +282,7 @@ Generated from the live objects — every signature below is real.
249
282
 
250
283
  | Name | Signature | What it does |
251
284
  |---|---|---|
252
- | `SymbolStrategy.use` | `(self, indicator: 'T') -> 'T'` | Register an indicator: auto-updated on every matching bar and |
285
+ | `SymbolStrategy.use` | `(self, indicator: 'T', *, session: 'Session \| None' = None) -> 'T'` | Register an indicator: auto-updated on every matching bar and |
253
286
  | `SymbolStrategy.buy` | `(self, size: 'int', *, stop_loss_ticks: 'int \| None' = None, take_profit_ticks: 'int \| None' = None, limit_price: 'Decimal \| None' = None, stop_price: 'Decimal \| None' = None, custom_tag: 'str \| None' = None) -> 'int \| None'` | Buy this contract (market unless a price kwarg implies otherwise); |
254
287
  | `SymbolStrategy.sell` | `(self, size: 'int', *, stop_loss_ticks: 'int \| None' = None, take_profit_ticks: 'int \| None' = None, limit_price: 'Decimal \| None' = None, stop_price: 'Decimal \| None' = None, custom_tag: 'str \| None' = None) -> 'int \| None'` | Sell this contract (market unless a price kwarg implies otherwise); |
255
288
  | `SymbolStrategy.close` | `(self) -> 'None'` | Flatten this contract's position; a rejection goes to ``on_reject``. |
@@ -418,6 +451,9 @@ Rules here; tables, per-function warmups and the refused-function list in `docs/
418
451
  tradeable price. Never route one into grid math without an explicit `round_to_tick`. They are
419
452
  float64-precise, not Decimal-exact: deterministic across reruns, but a `Cross` on a Bollinger
420
453
  edge can flip on `STDDEV`'s cancellation noise. A band touch is not exact.
454
+ - **`use(ind, session=…)` scopes the DATA; `trade_sessions=` scopes the DECISION.** Two
455
+ independent switches — see §5.11. Scoping an indicator does not restrict trading, and
456
+ restricting trading does not starve an indicator.
421
457
  - One indicator instance per thread; a parameter sweep gets one per worker.
422
458
 
423
459
  ### 5.3 Data
@@ -560,12 +596,23 @@ observed trading days and replays each synthetic sequence through a fresh `Combi
560
596
  - **`block_length=1` is a footgun.** It degenerates to an i.i.d. resample, destroys the
561
597
  losing streaks that actually blow accounts, and will report a pass probability that is far
562
598
  too kind. Default is 5.
563
- - **It refuses to default `horizon_days` when the source run FAILED** — a blown run stops
564
- recording days at the breach, so that count is a survival time, not a Combine length.
565
- `source_truncated` then flags that the sample is survivorship-biased by construction: the
566
- days after the blow-up do not exist, so every figure is conditioned on having survived.
567
- Nothing can repair that; do not quote such a result without the caveat.
599
+ - **`horizon_days` defaults to `BILLING_MONTH_DAYS` (21), never the observed day count.** A
600
+ Combine has no time limit, only a monthly fee, so "one attempt" defaults to one fee cycle —
601
+ the same unit `EvalEconomics` bills in. The fixed default is also the safety property: a
602
+ blown run stops recording days at the breach, so its count is a survival time, not a
603
+ Combine length, and is never used as a horizon. `source_truncated` still flags that such a
604
+ sample is survivorship-biased by construction — the days after the blow-up do not exist,
605
+ every figure is conditioned on having survived, and nothing can repair that; do not quote
606
+ such a result without the caveat.
568
607
  - **`provisional` below 30 source days.** Resampling cannot create information.
608
+ - **The point estimate ships with its own cross-examination** (`metrics/confidence.py`).
609
+ `pass_probability_ci` double-bootstraps the source days themselves — the error bar the
610
+ day count earns, which the path count never was. `block_length_sensitivity` shows whether
611
+ the streak assumption is load-bearing. `monte_carlo_by_year` refuses to average a hostile
612
+ year against a kind one. `crosscheck` compares against `sequential_combines` under a
613
+ binomial 2×SE null — when they disagree, the disagreement is the finding. Quote a pass
614
+ probability with its CI, not alone; `mc_confidence` bundles the lot and
615
+ `report.to_html(path, confidence=..., crosscheck=...)` renders the cards.
569
616
  - It cannot invent a regime the tape never contained, and it inherits every uncalibrated
570
617
  constant (§5.4). A probability to three decimals from unverified inputs is precise, not
571
618
  accurate.
@@ -656,25 +703,70 @@ account.
656
703
  - Still your problem on a multi-year tape: **exchange holidays** (~20/year, no calendar
657
704
  ships — filter upstream) and the uncalibrated constants (§5.4).
658
705
 
706
+ ### 5.11 Sessions scope indicator DATA and trading DECISIONS, separately
707
+
708
+ `ASIA` / `LONDON` / `NEW_YORK` live in `core/sessions.py`; the worked example is
709
+ `examples/session_scoped.py`.
710
+
711
+ - **The two switches are independent, and conflating them is the bug.**
712
+ `use(Atr(14), session=NEW_YORK)` restricts which bars that indicator is computed from.
713
+ `SymbolStrategy(..., trade_sessions=(NEW_YORK,))` restricts when `on_bar` may fire. Indicators
714
+ advance regardless of `trade_sessions` — an indicator fed only the tradable window develops
715
+ gaps and computes a different value from the same tape. Both default to today's behaviour, and
716
+ an unscoped strategy is byte-identical to one written before sessions existed.
717
+ - **Which indicators to scope is a modelling decision, and it is not uniform.** Dispersion
718
+ measures (`Atr`, `StdDev`, `Rsi`, `Stoch`, `BBands`) describe how much price moves PER BAR, and
719
+ that is session-dependent: on a 24h feed an `Atr(14)` read at 09:30 ET is computed almost
720
+ entirely from thin pre-market bars, so it understates NY volatility exactly when stop distance
721
+ is being sized. Level measures (`Sma`, `Ema`) answer where price IS, and the overnight move is
722
+ real — an NY-only `Ema` is anchored to yesterday's 16:00 close. **Levels are continuous across
723
+ sessions; dispersion is not.**
724
+ - **A scoped indicator warms in ITS OWN cadence.** It needs `history_bars` bars *of its session*,
725
+ so on 5-minute bars an NY-scoped `Sma(30)` spans ~25 trading days against ~7 unscoped. Read
726
+ `strategy.warm` (per-indicator update counts) or `history_bars_by_session` — never
727
+ `bars_seen >= history_bars`, which cannot express two cadences and OVERSTATES warmth.
728
+ `sequential_combines` still SIZES its preload slice from the unscoped `history_bars`, so a
729
+ scoped strategy will honestly report `fully_warm=False` rather than silently lying.
730
+ - **A `Cross` inherits its inputs' scope** and refuses a conflicting `session=`. Mixed-scope
731
+ inputs are refused outright: they advance on different bars, so comparing them compares values
732
+ sampled at unrelated instants.
733
+ - **Sessions are defined in their own timezone, not as fixed ET offsets.** DST comes from the
734
+ IANA database, so nothing rots — London and New York switch on different dates, and the London
735
+ window really is 04:00 ET rather than 03:00 for ~3 weeks each spring and ~1 each autumn. A
736
+ fixed ET block is still one line: `Session("LONDON_ET", ET, time(3), time(11))`.
737
+ - **Session membership is a DIFFERENT axis from `trading_day_of()`.** Asia sits after the 18:00
738
+ ET rollover, so its bars belong to the NEXT trading day. Membership is tested on `ts_event`
739
+ (the bar's OPEN) over a half-open `[start, end)` window, so the 09:29→09:30 bar — every print
740
+ of it pre-market — is not New York.
741
+ - **`trade_sessions` only narrows what THIS strategy does.** It never widens what the venue
742
+ permits: the 16:10 ET flatten and the 16:10–18:00 no-trade window apply either way.
743
+ - **You need 24h data.** The shipped `data/sample_mnq_1m.csv` is RTH-only (390 bars/day,
744
+ 09:30–15:59 ET) and contains no Asia or London bars, so nothing session-scoped is observable on
745
+ it. For synthetic bars pass `synthetic_bars(..., hours="globex")`; the default `"rth"` mode is
746
+ the New York session and nothing else.
747
+ - **Not built: per-session performance attribution.** Nothing in `SummaryStats` splits P&L by
748
+ session, so scoping is currently a modelling choice you make, not one the report scores.
749
+
659
750
  ## 6. The end-to-end workflow
660
751
 
661
752
  What using this framework actually looks like, in the order you do it:
662
753
 
663
754
  ```python
664
755
  from topstep_backtest import AccountSize, Backtest
665
- from topstep_backtest.metrics import monte_carlo
666
- from topstep_backtest.rules.kernel import Verdict
756
+ from topstep_backtest.metrics import crosscheck, mc_confidence, sequential_combines
667
757
  from topstep_backtest.rules.params import combine_params
668
758
 
669
- # 1. Write the strategy (§2), then run ONE backtest.
670
- report = Backtest(bars, MyStrategy(CONTRACT), account=AccountSize.S50K).run()
759
+ # 1. Write the strategy (§2), then run ONE backtest — recorded, so the replay
760
+ # tab can show intent against execution.
761
+ report = Backtest(bars, MyStrategy(CONTRACT), account=AccountSize.S50K, record=True).run()
671
762
  print(report) # verdict, day trail, four statistics blocks
672
763
 
673
764
  # 2. Sanity-check the wiring BEFORE reading any number.
674
765
  assert not report.result.rejections # zero trades + rejections = broken wiring (§5.5)
675
766
  assert report.bars_gated != len(bars) # warmup longer than your data
676
767
 
677
- # 3. Read the run, minding the basis (§5.6).
768
+ # 3. Read the run, minding the basis (§5.6) — and step the replay tab to each
769
+ # entry: the bracket where you meant it, the notes matching the fills.
678
770
  s = report.stats
679
771
  if s.provisional:
680
772
  ... # under 200 closes: estimates, not findings
@@ -684,16 +776,19 @@ s.drawdown.eod_trailing # Topstep's actual MLL mechanic
684
776
  s.drawdown.min_floor_headroom # closest the account came to death
685
777
  s.daily.p05 # the bad day to size against
686
778
 
687
- # 4. Stop trusting one sample. A FAILED run needs an explicit horizon (§5.7).
688
- horizon = None if report.result.verdict is not Verdict.FAILED else 40
689
- mc = monte_carlo(
690
- report.result, params=combine_params(AccountSize.S50K), paths=3000, horizon_days=horizon, seed=7
691
- )
779
+ # 4. Stop trusting one sample — real attempts first, resampled ones second.
780
+ sweep = sequential_combines(bars, lambda: MyStrategy(CONTRACT), window_days=21)
781
+ c = mc_confidence(report.result, params=combine_params(AccountSize.S50K), seed=7)
782
+ check = crosscheck(c.mc, sweep) # same billing-month horizon on both sides (§5.7)
692
783
 
693
- # 5. Act on the AUTOPSY, not the headline.
784
+ # 5. Act on the AUTOPSY, and quote the CI, never the bare point (§5.7).
694
785
  # mll_breach -> resize
695
786
  # consistency_blocked -> throttle the outsized day; the edge is fine
696
787
  # target_not_reached -> the edge is too slow; nothing risk-side helps
788
+ c.ci.p05, c.ci.p95 # the error bar the source-day count earns
789
+ c.sensitivity.spread # wide = the streak assumption is doing the work
790
+ check.divergent # bootstrap vs real windows: disagreement is the finding
791
+ report.to_html("run.html", confidence=c, crosscheck=check) # the cards, archived
697
792
 
698
793
  # 6. Was it an edge, or did you search until something looked good? The trial
699
794
  # count must be RECORDED — a remembered one is always too low, because the
@@ -706,7 +801,7 @@ dsr.expected_max_sharpe > dsr.sharpe # the search alone explains the result
706
801
 
707
802
  # 7. Is the attempt worth its price? Every figure is yours; none is baked in.
708
803
  ev = evaluate_ev(
709
- mc,
804
+ c.mc,
710
805
  economics=EvalEconomics(
711
806
  monthly_fee=Decimal("149"), pass_value=Decimal("2000"), reset_fee=Decimal("99")
712
807
  ),
@@ -724,9 +819,10 @@ that. **EV's `pass_value` is an assumption you supply and this package cannot ch
724
819
  dominates the answer, so quote `breakeven_pass_value` ("a pass must be worth at least $X")
725
820
  rather than `ev` unless you can defend the input.
726
821
 
727
- Runnable end to end: `examples/run_combine.py` (steps 1–3) and
728
- `examples/run_montecarlo.py` (steps 4–5, at two position sizes so the autopsy visibly
729
- discriminates). `docs/TUTORIAL_EMA_CROSSOVER.md` walks all of it line by line.
822
+ Runnable end to end: `examples/run_combine.py` (steps 1–3), `examples/run_windows.py` and
823
+ `examples/run_montecarlo.py` (steps 4–5, the latter at two position sizes so the autopsy
824
+ visibly discriminates). `docs/TUTORIAL_EMA_CROSSOVER.md` walks all of it line by line, and
825
+ `website/workflow.md` is the narrative version with the gate each stage must pass.
730
826
 
731
827
  ## 7. What `Backtest()` refuses, and what to use instead
732
828
 
@@ -844,8 +940,8 @@ moving one toward optimism is a probe to report next to the baseline, not a new
844
940
  `docs/ROADMAP.md` is the authority. These names appear in older notes and other frameworks and
845
941
  **do not exist here**: `LiveBroker`, `RecordingLiveBroker`, `DataEngine`, `TrailingMaxLossLimit`,
846
942
  `prob_fill_on_limit` (the real field is `BarFillConfig.fill_limit_on_touch`). Also absent: a
847
- Parquet/Arrow data catalog, multi-timeframe resampling, a MessageBus, per-year / per-regime
848
- breakdowns, higher fill tiers, an XFA rule set, and any holiday calendar.
943
+ Parquet/Arrow data catalog, multi-timeframe resampling, a MessageBus, per-year / per-regime /
944
+ per-session breakdowns (§5.11), higher fill tiers, an XFA rule set, and any holiday calendar.
849
945
 
850
946
  **Do** write code against these — they used to be on the list above and now ship: Monte-Carlo
851
947
  (§5.7), `optimize()` and the overfitting guards (§5.9), the sequential-Combine sweep (§5.8),
@@ -4,6 +4,154 @@ All notable changes to this project are documented here. This project adheres to
4
4
  [Semantic Versioning](https://semver.org/spec/v2.0.0.html). While the version is
5
5
  below 1.0, minor releases may contain breaking changes.
6
6
 
7
+ ## [0.4.0] — 2026-08-25
8
+
9
+ A minor: new features throughout. One behavioural default changed — the Monte-Carlo
10
+ horizon now defaults to one billing month rather than the observed day count.
11
+
12
+ ### Added
13
+
14
+ - **Regional sessions** (`core/sessions.py`: `Session`, `ASIA`, `LONDON`, `NEW_YORK`) and the two
15
+ independent switches that use them. `use(indicator, session=...)` scopes an indicator's INPUT
16
+ DATA; `SymbolStrategy(trade_sessions=...)` scopes DECISIONS. Keeping them separate is the
17
+ whole feature — a strategy can hold a continuous 24h `Ema` beside an NY-only `Atr` and still
18
+ trade only New York. Indicators advance regardless of `trade_sessions`, because starving one
19
+ outside the tradable window would leave it with gaps and a different value than the same
20
+ indicator on the same tape. Both default to today's behaviour and an unscoped run is
21
+ byte-identical.
22
+
23
+ Why you would scope at all: dispersion measures (`Atr`, `StdDev`, `Rsi`) describe how much
24
+ price moves *per bar*, and that is session-dependent — on 24h bars an `Atr(14)` read at 09:30
25
+ ET is computed almost entirely from thin pre-market bars, understating NY volatility exactly
26
+ when stop distance is being sized. Level measures (`Sma`, `Ema`) answer where price *is*, and
27
+ the overnight move is real: an NY-only `Ema` is anchored to yesterday's 16:00 close and blind
28
+ to a London repricing. Levels are continuous across sessions; dispersion is not.
29
+
30
+ **Sessions are defined in their own local timezone**, not as fixed ET offsets, so DST comes
31
+ from the IANA database and there is no hand-maintained table to rot — the same reasoning that
32
+ dropped the exchange calendar. London and New York switch on different dates, so the London
33
+ window really is 04:00 ET rather than 03:00 for about three weeks a year, and the tests assert
34
+ that against concrete 2026 dates. A fixed-ET block remains one line away
35
+ (`Session("LONDON_ET", ET, time(3), time(11))`), midnight-wrapping included.
36
+
37
+ Session membership is a **different axis** from `trading_day_of()`: Asia sits after the 18:00
38
+ ET rollover, so its bars belong to the *next* trading day. Membership is tested on `ts_event`
39
+ (the bar's open) over a half-open window, so the 09:29->09:30 bar — every print of it
40
+ pre-market — is not New York.
41
+
42
+ - `SymbolStrategy.warm`, `.registrations`, `.history_bars_by_session`, `.bars_out_of_session`,
43
+ `.bars_seen`, `.in_trade_session()` — the diagnostics scoping makes necessary. A scoped
44
+ indicator warms in its OWN cadence, so it needs `history_bars` bars *of its session*: on
45
+ 5-minute bars an NY-scoped `Sma(30)` spans ~25 trading days against ~7 for the same indicator
46
+ on a 24h feed. `warm` counts each indicator's own updates, which a single bar count cannot
47
+ express, and `sequential_combines` now asks the strategy rather than comparing counts — the
48
+ old `preload >= needed` silently OVERSTATED warmth for a scoped indicator.
49
+
50
+ - `synthetic_bars(hours="globex")` — the full 23-hour electronic session (18:00 ET previous
51
+ calendar day to 17:00 ET), reusing the canonical open from `core/time.py`. The default
52
+ `"rth"` mode is unchanged. Without this nothing session-scoped is testable: the shipped
53
+ `data/sample_mnq_1m.csv` is RTH-only (exactly 390 bars/day, 09:30-15:59 ET) and contains no
54
+ Asia or London bars at all.
55
+
56
+ - A `Cross` now **inherits its inputs' data scope** and refuses a conflicting one, alongside the
57
+ existing refusal of unregistered inputs. Crossing a 24h `Ema` against an NY-scoped `Ema`
58
+ compares values sampled on unrelated cadences; leaving the `Cross` continuous over scoped
59
+ inputs would have it re-read unchanged values on most bars. Both were silent before.
60
+
61
+ - **`metrics.spaced_combines` — pick how many Combine attempts to simulate.**
62
+ Where `sequential_combines` lets the tape dictate the attempt count (disjoint windows),
63
+ the new sweep takes `periods` and `window_days` and spreads that many start days evenly
64
+ across the tape, overlapping as much as the arithmetic requires — 100 days, 10 periods
65
+ of 40 days starts an attempt roughly every week. Each period is still a completely fresh
66
+ Combine (same shared implementation: day-aligned slicing, prewarmed indicators, the
67
+ kernel's own verdict, `classify_failure` attribution), so the sweep measures how much
68
+ passing depends on *when* the attempt starts. Because overlapping periods are not
69
+ independent samples, the result carries `effective_independent_windows`,
70
+ `stride_days` and `overlap_fraction` alongside its rates, and asking for more periods
71
+ than the tape has distinct start days is refused rather than replaying identical
72
+ windows as new observations. `examples/run_spaced.py` renders a sweep as a text
73
+ report — the rates with their effective-sample caveat in the header, then one line
74
+ per period in start order so start-date sensitivity is visible at a glance.
75
+
76
+ - **An HTML tearsheet for sweeps.** `sweep.to_html(path)` / `sweep.show()` on both
77
+ `WindowSweep` and `SpacedSweep` (or `tearsheet.render_sweep_html`) render one
78
+ self-contained page: a calendar timeline with each attempt drawn over its actual
79
+ dates and stacked into lanes where they overlap — the reason overlapping attempts
80
+ are not independent samples made visible rather than footnoted — the outcome
81
+ autopsy as a stacked bar with each failure mode's prescription, every attempt's
82
+ cumulative P&L overlaid from $0, a closest-approach-to-the-MLL-floor strip (a pass
83
+ with $40 of headroom looks identical to a robust one in the pass rate; not here),
84
+ and a sortable per-attempt table with P&L sparklines, all hover-linked. To feed the
85
+ charts, `WindowResult` now carries `daily_pnl` (and a `cumulative_pnl` property),
86
+ the same per-day series walk-forward's trials keep. No charting library and no
87
+ embedded JSON: the charts are SVG generated in Python, so the render stays a pure
88
+ byte-identical function of the frozen sweep data, and the evidence caveat is
89
+ stamped in the page header where a screenshot cannot shed it.
90
+
91
+ - **Confidence instruments for the Monte-Carlo pass probability**
92
+ (`metrics/confidence.py`). The point estimate's *simulation* error was never the real
93
+ uncertainty; these quantify what is. `pass_probability_ci` double-bootstraps the observed
94
+ day set itself and reports a 5th–95th percentile band — the error bar the source-day
95
+ count earns (a 65% from 500 days and a 65% from 40 now read differently).
96
+ `block_length_sensitivity` re-runs the estimate across block lengths and reports the
97
+ spread: wide means the streak-clustering assumption is doing the work.
98
+ `monte_carlo_by_year` stratifies by calendar year so a hostile year is never averaged
99
+ against a kind one — the spread across strata is the error bar non-stationarity imposes.
100
+ `crosscheck` forces the bootstrap and `sequential_combines` (opposite biases, shared
101
+ `classify_failure` on purpose) to answer side by side, flagging any outcome where the
102
+ real windows sit outside 2×SE of what the bootstrap's own probability would produce;
103
+ when they disagree, the disagreement is the finding. `mc_confidence` bundles the first
104
+ three, and every estimate runs through the same `monte_carlo_from_blocks` core as the
105
+ number it qualifies. All deterministic per seed, like everything else in `metrics/`.
106
+ - **Monte-Carlo cards on the tearsheet.** `report.to_html(path, confidence=...,
107
+ crosscheck=...)` (and `show` / `to_timestamped_html`) render the estimate with its CI,
108
+ the block-sensitivity row, the per-year strata, and the bootstrap-vs-windows card.
109
+ Caller-computed on purpose: writing a file never triggers thousands of simulations as a
110
+ side effect, and the render stays a pure, byte-identical function of its inputs.
111
+ - **The workflow, documented** (`website/workflow.md`, on the site nav as "The workflow").
112
+ The stages in the order the questions become answerable — envelope, data, strategy, one
113
+ recorded run, real attempts, the distribution with its confidence, the search guards,
114
+ the price — each ending with the gate that must pass before the next stage's number
115
+ means anything. `AGENTS.md` §6 is the code-first version and now walks the same loop
116
+ (recorded run, `sequential_combines`, `mc_confidence` + `crosscheck`, guards, EV).
117
+
118
+ - **A native loader for databento-data-playground Parquet exports.**
119
+ `topstep_backtest.data.loaders.load_bars(path)` reads an export produced by that project's
120
+ `convert.py` and returns `(bars, spec, meta)` ready for `Backtest(...)`. The product, bar
121
+ span and timestamp stamping are read from the file's embedded Parquet metadata rather than
122
+ passed in — each is a silent corruption when guessed — and the file's copy of the
123
+ instrument economics is cross-checked against the engine's `InstrumentSpec`, failing loudly
124
+ on any mismatch instead of picking one. Files without the metadata blob are refused.
125
+ Requires the existing `[data]` extra (pandas + pyarrow); nothing new to install.
126
+ `examples/run_real_data.py` auto-detects such exports and loads them with no declaration
127
+ flags (and refuses the flags on one — re-declaring what the file states is the
128
+ contradiction they exist to prevent), and the module joins the site's API reference and
129
+ quickstart.
130
+
131
+ ### Changed
132
+
133
+ - **The Monte-Carlo horizon defaults to one billing month** (`BILLING_MONTH_DAYS = 21`),
134
+ not the observed day count. A Combine has no time limit, only a monthly fee, so "one
135
+ attempt" defaults to one fee cycle — the same unit `EvalEconomics` bills in (its
136
+ `trading_days_per_month` default now *is* this constant). The fixed default also
137
+ retires the guard that made `horizon_days` mandatory for a blown source run: the old
138
+ hazard was defaulting to a survival time, and the new default cannot. A failed run
139
+ still reports `source_truncated`, and the survivorship-bias caveat stands unchanged.
140
+ - `sample_day_path`, `nearest_rank`, and `monte_carlo_from_blocks` are module-level
141
+ names in `metrics/montecarlo.py` (previously underscore-private) so the confidence
142
+ instruments run through the identical core; they remain outside `__all__`.
143
+ - **The replay cockpit is easier to read.** The running-stats panel no longer stacks all
144
+ ~50 rows: snapshots are regrouped into by-type cards — *status*, *trades (gross)*,
145
+ *round trips (net)*, *risk & quant*, *daily P&L* — shown one at a time behind filter
146
+ chips, with a `•` on any chip whose hidden card the newest snapshot moved. The values are
147
+ the same rows the results tab renders, relocated rather than rebuilt (a stat missing from
148
+ the regrouping map fails loudly instead of silently vanishing from the replay). The event
149
+ log moved from the cramped right column to a full-width strip under the charts — one
150
+ event per line, heading and kind filters on one row — and the panel text was tightened
151
+ throughout. Both tabs' price charts also gained a marker legend (entry/exit arrows, and
152
+ in the replay the cross fires and the working stop/target/avg-entry price lines), so the
153
+ glyphs on the tape no longer require guessing.
154
+
7
155
  ## [0.3.0] — 2026-08-20
8
156
 
9
157
  A minor: new features, no intended breaking changes.
@@ -422,5 +570,6 @@ rule and fee constants are **not yet calibrated against a live account**, so a
422
570
  are not exercised end to end.
423
571
  - Tier-0 bar fills only. No quote, depth or MBO tiers.
424
572
 
573
+ [0.4.0]: https://pypi.org/project/topstep-backtest/0.4.0/
425
574
  [0.3.0]: https://pypi.org/project/topstep-backtest/0.3.0/
426
575
  [0.1.0]: https://pypi.org/project/topstep-backtest/0.1.0/
@@ -1,6 +1,6 @@
1
1
  Metadata-Version: 2.5
2
2
  Name: topstep-backtest
3
- Version: 0.3.0
3
+ Version: 0.4.0
4
4
  Summary: Event-driven backtesting framework for Topstep Trading Combine strategies, with backtest/live parity against topstep-sdk.
5
5
  Author-email: Tarric Sookdeo <tarricsookdeo@outlook.com>
6
6
  License-Expression: MIT
@@ -84,6 +84,22 @@ indicator formula, so there is no second implementation to drift from the refere
84
84
  `Cross` is the deliberate exception — TA-Lib has no crossover primitive, so it is a
85
85
  framework helper that compares two TA-Lib outputs rather than computing anything.
86
86
 
87
+ **Sessions are two independent switches.** On a 24h tape you can scope each indicator's input
88
+ data and, separately, restrict when the strategy may trade:
89
+
90
+ ```python
91
+ super().__init__(contract_id, trade_sessions=(NEW_YORK,)) # trade New York only
92
+ self.trend = self.use(Ema(50)) # ...but see the whole tape
93
+ self.atr = self.use(Atr(14), session=NEW_YORK) # ...while this sees only NY
94
+ ```
95
+
96
+ The `Ema` keeps consuming Asia and London — an indicator fed only the tradable window would
97
+ develop gaps — while every decision happens in New York. Which indicators to scope is a
98
+ modelling choice with a rule behind it: *levels are continuous across sessions, dispersion is
99
+ not*. `ASIA`/`LONDON`/`NEW_YORK` are defined in their own timezones, so daylight saving comes
100
+ from the IANA database rather than a table that rots. See
101
+ [`examples/session_scoped.py`](examples/session_scoped.py).
102
+
87
103
  ## Install
88
104
 
89
105
  ```bash
@@ -184,8 +200,9 @@ cannot disagree. For a sheet from *every* run without naming a file each time,
184
200
  both the report and the path.
185
201
 
186
202
  **Replay a run bar by bar.** Pass `record=True` and the tearsheet grows a second tab: a
187
- replay cockpit that fills the window — its own charts on one side, the settled state, the
188
- running stats and the event log on the other, so nothing has to be scrolled between. Step
203
+ replay cockpit that fills the window — its own charts beside the settled state and the
204
+ running stats (by-type cards behind filter chips), the event log across the bottom, so
205
+ nothing has to be scrolled between. Step
189
206
  through the run (buttons, slider, arrow keys, autoplay) with everything after the cursor
190
207
  veiled, and watch the position, working orders (drawn as price lines), balance, floor
191
208
  headroom, indicator values — named by the attributes your strategy stores them under — and
@@ -231,6 +248,15 @@ imply three different fixes:
231
248
  | `consistency_blocked` | money made in too few days — throttle the outsized day; the edge is fine |
232
249
  | `target_not_reached` | the edge is too slow for the window — nothing risk-side helps |
233
250
 
251
+ The horizon defaults to one billing month (21 sessions) — a Combine has no time limit, only
252
+ a monthly fee, so "one attempt" means one fee cycle, the same unit `evaluate_ev` bills in.
253
+ And never quote the pass probability alone: `metrics.confidence.mc_confidence` attaches the
254
+ error bar the *source-day count* earns (a double-bootstrap CI — the path count was never the
255
+ real uncertainty), a block-length sensitivity row, and per-year strata, while
256
+ `metrics.confidence.crosscheck` compares the bootstrap against `sequential_combines`' real
257
+ windows — when the two disagree, the disagreement is the finding. The whole bundle renders
258
+ on the tearsheet via `report.to_html(path, confidence=..., crosscheck=...)`.
259
+
234
260
  ```bash
235
261
  uv run python examples/run_montecarlo.py # the same edge at two sizes, failing two ways
236
262
  ```
@@ -253,6 +279,10 @@ reruns — but read these before trusting a number:
253
279
  `metrics/economics.py`), but none of them escapes your sample: the Monte-Carlo resamples a
254
280
  strategy's own observed days, so it cannot invent a market regime your tape never contained
255
281
  and will understate tail risk on a short or single-regime sample.
282
+ - **No per-year, per-regime or per-session performance breakdown.** Sessions scope which bars
283
+ an indicator is computed from and when a strategy may trade, but nothing in `SummaryStats`
284
+ splits P&L by session, so a scoping choice is one you make on reasoning rather than one the
285
+ report scores for you.
256
286
  - **No exchange holiday calendar ships with this package.** Bars on market holidays and
257
287
  past early-close halts are not detected, flagged or filtered anywhere — filter them
258
288
  upstream. (A built-in calendar was removed in 0.1.0: it disagreed with CME on several
@@ -286,6 +316,7 @@ uv run pytest && uv run ruff check . && uv run ruff format --check . && uv run p
286
316
  uv run python scripts/gen_api_surface.py --check # AGENTS.md §3 must not be stale
287
317
  uv run python examples/run_combine.py # end-to-end combine verdict
288
318
  uv run python examples/run_montecarlo.py # outcome distribution + autopsy
319
+ uv run python examples/session_scoped.py # session-scoped indicators + trade gate
289
320
  ```
290
321
 
291
322
  Those checks are the gate: CI runs all of them on Python 3.12/3.13/3.14, then installs the
@@ -314,23 +345,29 @@ environment. There is no public issue tracker.
314
345
 
315
346
  - **`docs/TUTORIAL_EMA_CROSSOVER.md` — start here**: one strategy end to end, raw candles
316
347
  to Combine verdict.
348
+ - `website/workflow.md` — **the workflow**: the stages in the order the questions become
349
+ answerable — data, mechanics, real attempts, the distribution with its confidence, the
350
+ search guards, the price — and the gate each must pass before the next number means
351
+ anything.
317
352
  - `website/results.md` — **reading the report**: what every figure means, what basis it is
318
353
  on, which ones mislead alone, and what this framework deliberately does *not* report.
319
354
  - `AGENTS.md` — dense reference for *using* the framework: strategy dialect, module map, the
320
- invariants you must not break, and **§6, the end-to-end workflow** (run → check the wiring
321
- → read the run minding each metric's basis → Monte-Carlo → act on the autopsy).
355
+ invariants you must not break, and **§6, the end-to-end workflow** (run recorded → check
356
+ the wiring → read the run minding each metric's basis → real windows + Monte-Carlo with
357
+ its CI → act on the autopsy → deflate the search → price the attempt).
322
358
  - `docs/INDICATORS.md` — the TA-Lib indicator surface, wrapper by wrapper.
323
359
  - `docs/topstep-rules.md` — the rulebook being enforced, with sources, confidence levels,
324
360
  and a verify-before-trusting checklist.
325
361
  - `docs/DESIGN.md` — architecture contract, for *modifying* the framework.
326
362
  - `docs/ROADMAP.md` — what is built, partial, and not started. Next: **live adapter +
327
- calibration** → per-year / per-regime breakdowns → L1/L2/MBO fill tiers; funded-account
328
- (XFA) modeling is deliberately parked.
363
+ calibration** → per-regime (non-calendar) breakdowns → L1/L2/MBO fill tiers;
364
+ funded-account (XFA) modeling is deliberately parked.
329
365
  - `examples/` — runnable: `run_real_data.py` (your CSV/Parquet → verdict), `run_combine.py`
330
366
  (synthetic end to end), `run_tearsheet.py` (the same run as one HTML file),
331
367
  `run_replay.py` (a recorded run with the bar-by-bar scrubber), `run_montecarlo.py`
332
368
  (outcome distribution + autopsy), `run_windows.py` (a long tape replayed as consecutive
333
- independent Combine attempts), `ema_cross.py`, `sma_cross.py`, `talib_macd.py`,
369
+ independent Combine attempts), `run_spaced.py` (a chosen number of attempts, start days
370
+ spread evenly, overlap allowed), `ema_cross.py`, `sma_cross.py`, `talib_macd.py`,
334
371
  `hand_wired.py` (what the facade assembles).
335
372
 
336
373
  ## Stack