create-tradejs 3.1.29-beta.260 → 3.1.29-beta.262

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/dist/index.js CHANGED
@@ -450,8 +450,12 @@ On the first launch, TradeJS asks you to create the local root password.
450
450
  The project includes the complete checksum-managed TradeJS skill set under
451
451
  \`.codex/skills\`. Use \`$strategy-improvement-research\` to orchestrate a full
452
452
  bounded lineage, \`$strategy-backtest-research\` for one preregistered core
453
- experiment, and \`$ai-train-local-research\` for the frozen gate stage. For
454
- example:
453
+ experiment, and \`$ai-train-local-research\` for the frozen gate stage.
454
+ AI-gate research uses outer 60/40 by timestamp groups: discover on all first
455
+ 60%, diagnose three temporal stability blocks inside development, and keep the
456
+ last 40% sealed until candidate rules are frozen. Internal blocks are not
457
+ independent validation. Already opened history remains retrospective evidence.
458
+ For example:
455
459
 
456
460
  \`$strategy-candidate-report MarketFlushReversal\`
457
461
 
@@ -385,23 +385,21 @@ Interpretation:
385
385
  windows. Use `both` (default) or `direction_aware` for policies that retain
386
386
  both approved sides.
387
387
  - `yarn ai-pocket-search` is the default pocket discovery tool for future AI-gate rules. It reconstructs current strategy AI payloads, excludes outcome/current gate-output fields by default, shows progress bars, deduplicates equivalent row-selection pockets, and writes a Markdown report under `data/ai/output`.
388
- - `ai-pocket-search` uses a time-ordered row holdout by default
389
- (`--validationSplit 0.25`) and ranks candidates on that holdout. Treat it as
390
- tuning evidence, not an untouched test. For direction-specific discovery use
391
- `--direction LONG` or `--direction SHORT`. For release research, reserve a
392
- timestamp-grouped tail with `--testSplit ... --sealTest`; the report retains
393
- only its timestamp/count bounds and excludes its rows from current-gate and
394
- pocket economics. Open it once later with a frozen `ai-gate-ablation.mjs`
395
- spec. Plain `--testSplit` still prints test metrics and is therefore an opened
396
- historical test, not sealed release evidence. Use `--validationSplit 0` only
397
- for legacy full-sample exploration.
398
- - `ai-pocket-search` uses `--coverageMode auto` by default. It keeps the
399
- full-history search for non-provider features and runs separate CMC and
400
- Coinalyze cohorts over rows where that context is usable. Each cohort gets
401
- its own timestamp-grouped train/tuning/test split, and every reported cohort
402
- pocket must contain a predicate from that provider family. Use
403
- `--coverageMode full` only when intentionally reproducing the legacy single
404
- full-period search.
388
+ - `ai-pocket-search` defaults to outer 60/40 (`--validationSplit 0
389
+ --testSplit 0.4 --sealTest`). Search and ranking use the whole first 60%.
390
+ Three consecutive timestamp-grouped blocks inside that development partition
391
+ diagnose temporal stability; they are not independent holdouts and do not
392
+ refit pocket discovery. Stable candidates rank before sparse and unstable
393
+ ones within the bounded shortlist. Sparse support remains research-only.
394
+ Freeze expressions before opening the last 40% once with native ablation.
395
+ Never refine thresholds on that tail; previously inspected history remains
396
+ retrospective evidence. Explicit nonzero validationSplit is a legacy option,
397
+ not the standard. For direction-specific discovery use LONG or SHORT.
398
+ - Provider coverage cohorts inherit the global development/test boundaries.
399
+ Filter CMC/Coinalyze coverage within each partition; never resplit a later
400
+ provider window so that global test rows enter discovery. Every cohort pocket
401
+ must include a predicate from that provider family. `--coverageMode full`
402
+ reproduces a single full-history search.
405
403
  - `ai-pocket-search` uses `--cadenceMode auto` by default. For a sparse train
406
404
  partition it scales discovery-only `minSupport` / `minEvents` down from the
407
405
  legacy 20 / 10 defaults using the number of independent timestamp events.
@@ -82,8 +82,8 @@ node .codex/skills/ai-train-local-research/scripts/ai-gate-ablation.mjs \
82
82
  --file data/ai/export/ai-dataset-liquiditytails-merged-1784296244106-part1.jsonl \
83
83
  --variant 'near-ma-and-zone::filter::additionalIndicators.baseContext.regime.trend.priceDistanceToMaSlowAtr <= 1.2 && additionalIndicators.baseContext.structure.liquidityZones.activeCount >= 1' \
84
84
  --featurePattern 'priceDistanceToMaSlowAtr|liquidityZones.activeCount' \
85
- --validationSplit 0.2 \
86
- --testSplit 0.2 \
85
+ --validationSplit 0 \
86
+ --testSplit 0.4 \
87
87
  --output data/ai/output/liquiditytails-near-ma-and-zone.md
88
88
  ```
89
89
 
@@ -206,18 +206,19 @@ Use `--featurePattern '<regex>'` to print matching causal paths, availability,
206
206
  ranges, and categories. Do not use `--includeGateContext` for discovery; it is
207
207
  only for auditing current gate output fields.
208
208
 
209
- For direction-specific discovery, use `yarn ai-pocket-search --direction LONG`
210
- or `--direction SHORT`. For release evidence reserve an untouched chronological
211
- tail with `--testSplit ... --sealTest`; the search ranks pockets using only the
212
- preceding train and validation rows and reports only the sealed test bounds.
213
- Plain `--testSplit` exposes test metrics and cannot be called untouched after
214
- the report is read. Open the sealed tail once with the frozen fixed-rule
215
- ablation.
209
+ The default is timestamp-grouped outer 60/40: `--validationSplit 0
210
+ --testSplit 0.4`. Discovery seals the last 40% by default. Search uses the whole
211
+ first 60%; three consecutive development blocks diagnose stability. Fixed-rule
212
+ ablation reports the same three blocks and opens the outer test only after the
213
+ candidate expressions are frozen. These internal blocks overlap development;
214
+ they are not independent holdouts. Do not adjust rules on the outer test.
215
+ Explicit nonzero validationSplit remains a labelled legacy option.
216
216
 
217
- When several core candidates must be compared, pass the same exact UTC
218
- `--tuningSince` and `--testSince` boundaries to every candidate ablation.
219
- Exact boundaries take precedence over ratio splits and keep sparse candidates
220
- on one calendar partition contract.
217
+ For direction-specific discovery use LONG or SHORT. For multiple core exports,
218
+ freeze one UTC `--testSince` boundary once and pass it to every ablation.
219
+ With no tuningSince, all earlier rows form development and tuning is empty.
220
+ Exact boundaries override ratios; use the parent development timestamp groups
221
+ to choose the shared 60/40 calendar boundary, rather than moving it per candidate.
221
222
 
222
223
  Use `--windowStart <UTC> --windowEnd <UTC>` to compare candidates over the same
223
224
  calendar window. The start is inclusive, and the end is exclusive. Full-period
@@ -233,8 +234,8 @@ available strategy contains shared LONG or SHORT approval/block pockets:
233
234
  ```bash
234
235
  node .codex/skills/ai-train-local-research/scripts/ai-gate-ablation.mjs \
235
236
  --crossStrategy \
236
- --validationSplit 0.2 \
237
- --testSplit 0.2 \
237
+ --validationSplit 0 \
238
+ --testSplit 0.4 \
238
239
  --portfolioCapacity 5 \
239
240
  --output data/ai/output/cross-strategy-shared-pockets.md
240
241
  ```
@@ -323,8 +324,8 @@ node .codex/skills/ai-train-local-research/scripts/ai-gate-ablation.mjs \
323
324
  --strategy LiquidityTails \
324
325
  --movingAverageStudy \
325
326
  --maPeriods 5,10,15,20,25,30,35,40,45,50,55,60,65,70,75,80,85,90,95,100 \
326
- --validationSplit 0.2 \
327
- --testSplit 0.2 \
327
+ --validationSplit 0 \
328
+ --testSplit 0.4 \
328
329
  --json \
329
330
  --output data/ai/output/liquiditytails-ma-grid.json
330
331
  ```
@@ -346,7 +347,7 @@ or parity is incomplete. Candidate ranking uses train and tuning only; the
346
347
  timestamp-grouped test tail is reported after selection and remains exposed
347
348
  historical evidence after the first run.
348
349
 
349
- `--crossStrategy` requires positive `--validationSplit` and `--testSplit`.
350
+ `--crossStrategy` requires positive `--testSplit` and non-negative `--validationSplit` (standard: 0 and 0.4).
350
351
  Opening the historical test tail makes it exposed evidence. Re-running the tool
351
352
  on the same cutoff does not make it untouched again. Every candidate remains
352
353
  research-only until the exact frozen rule survives timestamps strictly after
@@ -360,7 +361,7 @@ Every report contains:
360
361
 
361
362
  - baseline and candidate tables for full history, `180d`, `90d`, `30d`, `7d`;
362
363
  - q3+/q4+/q5+ summaries, configurable with `--qualityThresholds`;
363
- - timestamp-grouped, time-ordered train/tuning/untouched-test splits;
364
+ - timestamp-grouped outer development/test split (60/40 by default), with three temporal stability blocks inside development;
364
365
  - direction and monthly stability;
365
366
  - matched, removed, and added slices;
366
367
  - PnL, winrate, PF, Sharpe, Sortino, Calmar, max drawdown, DD ratios, strict
@@ -396,3 +397,8 @@ Run the tool tests after every change:
396
397
  ```bash
397
398
  node --test .codex/skills/ai-train-local-research/scripts/ai-gate-ablation.test.mjs
398
399
  ```
400
+
401
+ With the standard 60/40 cross-strategy split, the field `tuning` is the last
402
+ third of development, overlapping the discovery partition. It is a diagnostic
403
+ stability block, not an independent validation set. All three development
404
+ blocks are reported separately; the outer test remains distinct.
@@ -101,16 +101,22 @@ of implying that every approval filled.
101
101
 
102
102
  ### 7. Validation
103
103
 
104
- Do not rename row-based candidate-selection output to an untouched test.
104
+ Default: outer 60/40 by whole timestamp groups. Search uses the entire first
105
+ 60%; report three stability blocks within it. Internal blocks overlap
106
+ development and are not independent validation. Label outer test as historical
107
+ when it has already been inspected. Legacy explicit partitions must name their
108
+ actual boundaries. Do not rename candidate-selection output to an untouched test.
105
109
 
106
110
  ```md
107
111
  ### Validation
108
112
 
109
113
  | Partition | Rows | Events | Approved N | WR | PF | PnL | MaxDD | Max batch |
110
114
  | -------------- | ---: | -----: | ---------: | --: | --: | --: | ----: | --------: |
111
- | train | | | | | | | | |
112
- | tuning | | | | | | | | |
113
- | untouched test | | | | | | | | |
115
+ | development 60% | | | | | | | | |
116
+ | development block 1 | | | | | | | | |
117
+ | development block 2 | | | | | | | | |
118
+ | development block 3 | | | | | | | | |
119
+ | outer test 40% | | | | | | | | |
114
120
  ```
115
121
 
116
122
  ### 8. Acceptance checks
@@ -32,8 +32,8 @@ Options:
32
32
  --minQuality <n> Main baseline threshold (default: 4)
33
33
  --qualityThresholds <list> qN+ summaries (default: 3,4,5)
34
34
  --terminalWindows <list> Terminal windows in days (default: 180,90,30,7)
35
- --validationSplit <ratio> Trailing timestamp-grouped tuning share (default: 0.25)
36
- --testSplit <ratio> Later timestamp-grouped test share (default: 0)
35
+ --validationSplit <ratio> Optional legacy tuning share (default: 0; stability stays in development)
36
+ --testSplit <ratio> Outer timestamp-grouped test share (default: 0.4)
37
37
  --tuningSince <timestamp> Exact UTC boundary where tuning starts
38
38
  --testSince <timestamp> Exact UTC boundary where test starts
39
39
  --windowStart <timestamp> Common comparison start, inclusive
@@ -91,8 +91,8 @@ export const parseCliArgs = (argv) => {
91
91
  minQuality: 4,
92
92
  qualityThresholds: DEFAULT_QUALITY_THRESHOLDS,
93
93
  terminalWindows: DEFAULT_WINDOWS,
94
- validationSplit: 0.25,
95
- testSplit: 0,
94
+ validationSplit: 0,
95
+ testSplit: 0.4,
96
96
  tuningSince: null,
97
97
  testSince: null,
98
98
  capacities: DEFAULT_CAPACITIES,
@@ -1193,6 +1193,13 @@ export const splitRowsByTimestamp = (rows, validationSplit, testSplit = 0) => {
1193
1193
  };
1194
1194
 
1195
1195
  export const splitRowsByTimestampBounds = (rows, tuningSince, testSince) => {
1196
+ if (tuningSince == null && Number.isFinite(testSince)) {
1197
+ return {
1198
+ train: rows.filter((row) => row.timestamp < testSince),
1199
+ tuning: [],
1200
+ test: rows.filter((row) => row.timestamp >= testSince),
1201
+ };
1202
+ }
1196
1203
  if (!Number.isFinite(tuningSince) || !Number.isFinite(testSince)) {
1197
1204
  throw new Error(
1198
1205
  'Exact calendar partitions require both tuningSince and testSince',
@@ -1210,6 +1217,21 @@ export const splitRowsByTimestampBounds = (rows, tuningSince, testSince) => {
1210
1217
  };
1211
1218
  };
1212
1219
 
1220
+ // Three diagnostics wholly inside the development partition, never test selection.
1221
+ export const developmentBlocks = (rows) => {
1222
+ const timestamps = [...new Set(rows.map((row) => row.timestamp))].sort(
1223
+ (a, b) => a - b,
1224
+ );
1225
+ return [0, 1, 2].map((index) => {
1226
+ const from = timestamps[Math.floor((timestamps.length * index) / 3)];
1227
+ const until =
1228
+ index === 2
1229
+ ? Infinity
1230
+ : timestamps[Math.floor((timestamps.length * (index + 1)) / 3)];
1231
+ return rows.filter((row) => row.timestamp >= from && row.timestamp < until);
1232
+ });
1233
+ };
1234
+
1213
1235
  const summarizeSplit = (rows, selector, summaryOptions) =>
1214
1236
  summarizeRows(selectRows(rows, selector), getPeriodDays(rows), {
1215
1237
  ...summaryOptions,
@@ -3034,6 +3056,9 @@ export const evaluateCrossPocket = ({
3034
3056
  expectedSign,
3035
3057
  );
3036
3058
  const test = summarizeCrossSlice(split.test, pocket.predicates, expectedSign);
3059
+ const developmentStability = developmentBlocks(split.train).map((block) =>
3060
+ summarizeCrossSlice(block, pocket.predicates, expectedSign),
3061
+ );
3037
3062
  const negativeControl = evaluateNegativeControl(
3038
3063
  split.test,
3039
3064
  pocket.predicates,
@@ -3109,6 +3134,7 @@ export const evaluateCrossPocket = ({
3109
3134
  return {
3110
3135
  condition: pocket.condition,
3111
3136
  predicates: pocket.predicates,
3137
+ developmentStability,
3112
3138
  train,
3113
3139
  tuning,
3114
3140
  test,
@@ -3594,9 +3620,9 @@ export const buildCrossStrategyReport = async ({
3594
3620
  if (groups.length < 2) {
3595
3621
  throw new Error('Cross-strategy research requires at least two exports');
3596
3622
  }
3597
- if (testSplit <= 0 || validationSplit <= 0) {
3623
+ if (testSplit <= 0 || validationSplit < 0) {
3598
3624
  throw new Error(
3599
- '--crossStrategy requires positive --validationSplit and --testSplit',
3625
+ '--crossStrategy requires positive --testSplit and non-negative --validationSplit',
3600
3626
  );
3601
3627
  }
3602
3628
  let searchAiPockets = searchAiPocketsOverride;
@@ -3731,6 +3757,10 @@ export const buildCrossStrategyReport = async ({
3731
3757
  left.timestamp - right.timestamp || left.sequence - right.sequence,
3732
3758
  );
3733
3759
  const split = splitRowsByTimestamp(rows, validationSplit, testSplit);
3760
+ if (validationSplit === 0) {
3761
+ // This overlaps development deliberately: diagnostic stability, not holdout tuning.
3762
+ split.tuning = developmentBlocks(split.train).at(-1);
3763
+ }
3734
3764
  const minSharedStrategies = Math.min(
3735
3765
  groups.length,
3736
3766
  Math.max(5, Math.ceil(minFeatureStrategies * 0.6)),
@@ -3968,6 +3998,10 @@ export const buildCrossStrategyReport = async ({
3968
3998
  },
3969
3999
  validationSplit,
3970
4000
  testSplit,
4001
+ tuningRole:
4002
+ validationSplit === 0
4003
+ ? 'last development block, overlapping diagnostic; not independent holdout'
4004
+ : 'legacy separate tuning',
3971
4005
  minFeatureStrategies,
3972
4006
  acceptance: {
3973
4007
  minSharedStrategies,
@@ -4015,8 +4049,8 @@ export const buildAblationReport = ({
4015
4049
  minQuality,
4016
4050
  qualityThresholds,
4017
4051
  terminalWindows,
4018
- validationSplit,
4019
- testSplit = 0,
4052
+ validationSplit = 0,
4053
+ testSplit = 0.4,
4020
4054
  tuningSince = null,
4021
4055
  testSince = null,
4022
4056
  windowStart = null,
@@ -4054,12 +4088,12 @@ export const buildAblationReport = ({
4054
4088
  }
4055
4089
  const minTimestamp = windowStart ?? rows[0].timestamp;
4056
4090
  const maxTimestamp = windowEnd ?? rows.at(-1).timestamp;
4057
- if ((tuningSince == null) !== (testSince == null)) {
4091
+ if (tuningSince != null && testSince == null) {
4058
4092
  throw new Error(
4059
4093
  'Exact calendar partitions require both tuningSince and testSince',
4060
4094
  );
4061
4095
  }
4062
- const exactCalendarPartitions = tuningSince != null;
4096
+ const exactCalendarPartitions = testSince != null;
4063
4097
  const split = exactCalendarPartitions
4064
4098
  ? splitRowsByTimestampBounds(rows, tuningSince, testSince)
4065
4099
  : splitRowsByTimestamp(rows, validationSplit, testSplit);
@@ -4088,6 +4122,12 @@ export const buildAblationReport = ({
4088
4122
  }
4089
4123
  : {}),
4090
4124
  };
4125
+ const stabilityBlocks = developmentBlocks([...split.train, ...split.tuning]);
4126
+ const stabilitySummary = (selector) =>
4127
+ stabilityBlocks.map((block) => ({
4128
+ ...partitionEvidence(block),
4129
+ metrics: summarizeSplit(block, selector, summaryOptions),
4130
+ }));
4091
4131
  const baselineSelector = (row) => baselineSelectedAt(row, minQuality);
4092
4132
  const approvedSignalTrace = (selector) =>
4093
4133
  selectRows(rows, selector).map((row) => ({
@@ -4121,6 +4161,7 @@ export const buildAblationReport = ({
4121
4161
  maxTimestamp,
4122
4162
  summaryOptions,
4123
4163
  }),
4164
+ developmentStability: stabilitySummary(baselineSelector),
4124
4165
  train: summarizeSplit(split.train, baselineSelector, summaryOptions),
4125
4166
  tuning: summarizeSplit(split.tuning, baselineSelector, summaryOptions),
4126
4167
  test: summarizeSplit(split.test, baselineSelector, summaryOptions),
@@ -4184,6 +4225,7 @@ export const buildAblationReport = ({
4184
4225
  maxTimestamp,
4185
4226
  summaryOptions,
4186
4227
  }),
4228
+ developmentStability: stabilitySummary(candidateSelector),
4187
4229
  train: summarizeSplit(split.train, candidateSelector, summaryOptions),
4188
4230
  tuning: summarizeSplit(split.tuning, candidateSelector, summaryOptions),
4189
4231
  test: summarizeSplit(split.test, candidateSelector, summaryOptions),
@@ -4235,12 +4277,14 @@ export const buildAblationReport = ({
4235
4277
  validationSplit,
4236
4278
  testSplit,
4237
4279
  partitionMode: exactCalendarPartitions ? 'exact-calendar' : 'ratio',
4280
+ developmentStabilityRole:
4281
+ 'diagnostic inside development; never outer test selection',
4282
+ developmentStabilityFolds: 3,
4238
4283
  comparisonWindow: explicitWindow
4239
4284
  ? { start: windowStart, end: windowEnd, interval: '[start, end)' }
4240
4285
  : null,
4241
- tuningSince: exactCalendarPartitions
4242
- ? new Date(tuningSince).toISOString()
4243
- : null,
4286
+ tuningSince:
4287
+ tuningSince != null ? new Date(tuningSince).toISOString() : null,
4244
4288
  testSince: exactCalendarPartitions
4245
4289
  ? new Date(testSince).toISOString()
4246
4290
  : null,
@@ -4566,6 +4610,14 @@ export const formatMarkdownReport = (report) => {
4566
4610
  report.run.trainEvents,
4567
4611
  report.baseline.train,
4568
4612
  ),
4613
+ ...(report.baseline.developmentStability ?? []).map((block, index) =>
4614
+ validationSummaryRow(
4615
+ `development block ${index + 1} (diagnostic)`,
4616
+ block.rows,
4617
+ block.events,
4618
+ block.metrics,
4619
+ ),
4620
+ ),
4569
4621
  validationSummaryRow(
4570
4622
  'tuning',
4571
4623
  report.run.tuningRows,
@@ -4573,7 +4625,7 @@ export const formatMarkdownReport = (report) => {
4573
4625
  report.baseline.tuning,
4574
4626
  ),
4575
4627
  validationSummaryRow(
4576
- 'untouched test',
4628
+ 'outer test (opened)',
4577
4629
  report.run.testRows,
4578
4630
  report.run.testEvents,
4579
4631
  report.baseline.test,
@@ -4685,13 +4737,20 @@ export const formatMarkdownReport = (report) => {
4685
4737
  report.baseline.train,
4686
4738
  variant.train,
4687
4739
  ),
4740
+ ...(variant.developmentStability ?? []).map((block, index) =>
4741
+ validationComparisonRow(
4742
+ `development block ${index + 1} (diagnostic)`,
4743
+ report.baseline.developmentStability[index].metrics,
4744
+ block.metrics,
4745
+ ),
4746
+ ),
4688
4747
  validationComparisonRow(
4689
4748
  'tuning',
4690
4749
  report.baseline.tuning,
4691
4750
  variant.tuning,
4692
4751
  ),
4693
4752
  validationComparisonRow(
4694
- 'untouched test',
4753
+ 'outer test (opened)',
4695
4754
  report.baseline.test,
4696
4755
  variant.test,
4697
4756
  ),
@@ -11,6 +11,7 @@ import {
11
11
  applyBenchmarkEventSnapshots,
12
12
  balanceCrossStrategyRows,
13
13
  buildAblationReport,
14
+ developmentBlocks,
14
15
  buildCrossStrategyReport,
15
16
  buildEquitySeries,
16
17
  buildMovingAverageVariants,
@@ -1453,3 +1454,34 @@ test('uses inclusive UTC calendar days for terminal active-day ratio', () => {
1453
1454
  assert.equal(report.baseline.periods['7d'].activeDays, 2);
1454
1455
  assert.equal(report.baseline.periods['7d'].activeDayRatio, 0.25);
1455
1456
  });
1457
+
1458
+ test('outer 60/40 is the CLI default and stability blocks preserve whole events', () => {
1459
+ const options = parseCliArgs([]);
1460
+ assert.equal(options.validationSplit, 0);
1461
+ assert.equal(options.testSplit, 0.4);
1462
+ const rows = Array.from({ length: 30 }, (_, i) => ({
1463
+ timestamp: Math.floor(i / 2),
1464
+ }));
1465
+ const blocks = developmentBlocks(rows);
1466
+ assert.deepEqual(
1467
+ blocks.map((block) => block.length),
1468
+ [10, 10, 10],
1469
+ );
1470
+ assert.equal(new Set(blocks.flat().map((row) => row.timestamp)).size, 15);
1471
+ for (let i = 0; i < 2; i++)
1472
+ assert.ok(blocks[i].at(-1).timestamp < blocks[i + 1][0].timestamp);
1473
+ });
1474
+
1475
+ test('shared outer boundary permits development/test without separate tuning', () => {
1476
+ const rows = [1, 2, 2, 3, 4].map((timestamp) => ({ timestamp }));
1477
+ const split = splitRowsByTimestampBounds(rows, null, 3);
1478
+ assert.deepEqual(
1479
+ split.train.map((row) => row.timestamp),
1480
+ [1, 2, 2],
1481
+ );
1482
+ assert.deepEqual(split.tuning, []);
1483
+ assert.deepEqual(
1484
+ split.test.map((row) => row.timestamp),
1485
+ [3, 4],
1486
+ );
1487
+ });
@@ -245,3 +245,9 @@ After creating or editing notes, run:
245
245
  ```bash
246
246
  node .codex/skills/strategy-backtest-research/scripts/research-notes-check.mjs
247
247
  ```
248
+
249
+ For AI-gate studies record the standard outer 60/40 timestamp-group split,
250
+ three temporal stability blocks inside the full first 60%, and the frozen
251
+ outer test boundary. The internal blocks are development diagnostics, not
252
+ independent validation. Note every previous exposure of the outer history;
253
+ changing the split cannot make previously inspected data untouched again.
@@ -93,7 +93,7 @@ composition ledger, and the mandatory visual artifacts.
93
93
  `production core + current AI-gate` composition is the mandatory baseline;
94
94
  the production core with its newly rebuilt gate is a separate research
95
95
  candidate and must not replace the current-gate baseline.
96
- Use common calendar train/tuning/test boundaries and open every sealed gate
96
+ Use common calendar development/test boundaries (outer 60/40; three stability blocks inside development) and open every sealed gate
97
97
  tail together only after all per-core gate variants are frozen. Run every
98
98
  gate board through `$ai-train-local-research`; do not retune or relabel the
99
99
  raw-core result inside gate tooling.
@@ -23,15 +23,14 @@ or a gate tuned on another core's signal population.
23
23
  is the baseline; the rebuilt gate is a research candidate. A no-op,
24
24
  zero-approval, or rejected gate still receives a checksum-bound report and
25
25
  explicit disposition; it is not silently omitted.
26
- 4. Freeze common calendar train, tuning, and test boundaries once in the parent
27
- lineage. Keep all rows sharing a timestamp in one partition. Candidate
28
- exports may have different row counts, but final comparison windows and test
29
- timestamps must not move to make one candidate look better. Discovery may
30
- inspect only train and tuning. Open all candidate test tails together after
31
- every gate spec is immutable.
32
- Use `ai-gate-ablation.mjs --tuningSince <UTC> --testSince <UTC>` with the
33
- same values for every candidate; ratio-only splits are not comparable when
34
- core event cadence differs.
26
+ 4. Freeze common calendar development/test boundaries once in the parent
27
+ lineage: outer 60/40, preserving whole timestamp groups. Search the whole
28
+ development partition and diagnose three consecutive blocks inside it.
29
+ These blocks overlap development, not independent validation. Candidate
30
+ exports may have different cadence, but the outer boundary must not move.
31
+ Freeze every gate spec before opening all outer test tails together.
32
+ Use `ai-gate-ablation.mjs --testSince <UTC>` with the same boundary for every
33
+ candidate; ratio-only splits are not comparable across different cadence.
35
34
  5. Count every inspected core and gate behavior in the final trial ledger. The
36
35
  selected strategy is `core identity + candidate-specific gate identity +
37
36
  context identity + direction policy + quality threshold`. Reusing economics
@@ -79,7 +79,7 @@ Retain:
79
79
  attempt over full/3y/4y/max-covered/365d/180d/90d/30d/7d windows;
80
80
  - control/candidate matched, removed, added, changed-outcome, occupancy, regime,
81
81
  month, symbol, event, concentration, and capacity evidence;
82
- - gate train/tuning/test boundaries, feature provenance, threshold rounding,
82
+ - gate development/test boundaries (outer 60/40), three internal stability blocks, feature provenance, threshold rounding,
83
83
  sealed-test flag/open timestamp, ablation, support, and one-round selection
84
84
  decision; record any accidental early test exposure permanently;
85
85
  - the complete `$ai-train-local-research` report sections and structured
@@ -427,7 +427,7 @@ with a future unexposed tail.
427
427
  ## 7. Use one gate tuning round
428
428
 
429
429
  Freeze the isolated finalist's raw-core export and the current deterministic
430
- gate as control. Use one time-grouped, time-ordered train/tuning/test design.
430
+ gate as control. Use one timestamp-grouped outer 60/40 design, with three temporal stability blocks inside the first 60%.
431
431
  Audit existing gate rules, run pocket discovery/ablation without outcome or
432
432
  execution leakage, and preregister rounded thresholds before opening the test.
433
433
  `ai-pocket-search` must reserve the test with `--sealTest`; its discovery report
@@ -710,8 +710,8 @@ yarn ai-train --strategy <Strategy> --file <merged-export-part1.jsonl> \
710
710
  --terminalWindows=1460,1095,365,180,90,30,7
711
711
 
712
712
  yarn ai-pocket-search --strategy <Strategy> \
713
- --file <merged-export-part1.jsonl> -n 0 --validationSplit 0.2 \
714
- --testSplit 0.2 --sealTest --maxDepth 2 --minSupport 25
713
+ --file <merged-export-part1.jsonl> -n 0 --validationSplit 0 \
714
+ --testSplit 0.4 --sealTest --maxDepth 2 --minSupport 25
715
715
 
716
716
  yarn strategy:release profile --input <trades.jsonl> --variant <finalist-id> \
717
717
  --startTime <start-ms> --endTime <end-ms> --days 7,30,90 \
@@ -1,7 +1,7 @@
1
1
  {
2
2
  "schema": "tradejs-skill-bundle/v1",
3
3
  "source": "TradeJS-Dev/TradeJS:.codex/skills",
4
- "bundleSha256": "f31f221397fa300534800c147d1ea57e549627083723acc7c97b75364d31da11",
4
+ "bundleSha256": "650502f532888b4d8e63ae8c52745bdf8d114feb2398831858d7e9f035b1e4bd",
5
5
  "skills": [
6
6
  "ai-train-local-research",
7
7
  "backtest-config-redis",
@@ -19,18 +19,18 @@
19
19
  "strategy-release"
20
20
  ],
21
21
  "files": {
22
- ".codex/skills/ai-train-local-research/references/gate-ablation.md": "afc8334de0282fd0fac3886f467a1fca8e67c624a48cb5de66351635a039421e",
22
+ ".codex/skills/ai-train-local-research/references/gate-ablation.md": "2fff61953de59fc0030004a29a3a9dcee2e5783dd47e1ab6dcfac73c0eaae756",
23
23
  ".codex/skills/ai-train-local-research/references/pocket-feature-exclusion.md": "c7af65fcfcd404d5324f0222678126d5e166fb77c4c27afea96369ef427f4bb0",
24
- ".codex/skills/ai-train-local-research/references/reporting.md": "a5a5cc6438a8cc228560e302e944f1ff320f4c963988af83e1ba1443728f7149",
25
- ".codex/skills/ai-train-local-research/scripts/ai-gate-ablation.mjs": "f68a1a7c898f4c2f46860709a52cabd3abd4cc8cd14443de2b5f827d261be935",
26
- ".codex/skills/ai-train-local-research/scripts/ai-gate-ablation.test.mjs": "f13d2284a4c04a5220588652be81d104d2acde9650481cf569fa0efe56118f08",
27
- ".codex/skills/ai-train-local-research/SKILL.md": "479e44aad9cb6ab6b5c8931c807c63cf2a1938a59d01f553c20f44645716a23a",
24
+ ".codex/skills/ai-train-local-research/references/reporting.md": "ae0faaab3b3dd0f4bd27142fa2c9aa889c5e7572a10e9fe1ee79ff57ebbd8653",
25
+ ".codex/skills/ai-train-local-research/scripts/ai-gate-ablation.mjs": "09336b11919c0d93f5b3916f88a546b2b1a9c7a2e26c30dca90c610a9114d8b0",
26
+ ".codex/skills/ai-train-local-research/scripts/ai-gate-ablation.test.mjs": "c21f36e88888ad5f444e615b5ddf3419e834d945ff9517318288516d7dd489bf",
27
+ ".codex/skills/ai-train-local-research/SKILL.md": "b75a1e737a13f71c96a8c45ff11891dc4d1c415c784d4146dab9cdcbe376f49a",
28
28
  ".codex/skills/backtest-config-redis/scripts/get_backtest_config.sh": "833e950d6348c5b7a4bd3f00d25af60f6394190bff744bd5ab64db9e43d826d5",
29
29
  ".codex/skills/backtest-config-redis/SKILL.md": "85e7b1fa425aa23dfa7950de05c1d3ce046d044dc3872f3665705e9e2f2a1977",
30
30
  ".codex/skills/runtime-parity-mismatch-analysis/SKILL.md": "289925c0069d1933ec3b09d5d6a020632170716c61aafb7031a0998acf3736f6",
31
31
  ".codex/skills/save-strategy-config-from-backtest/agents/openai.yaml": "ba163f9449562241ed954ca2e0f12e22562390740cdcf0d09d7f99f0ef791741",
32
32
  ".codex/skills/save-strategy-config-from-backtest/SKILL.md": "a73cddc29fb8fa741fe6226c051e19b0f9a19747babf96d336e7f165fb1fd298",
33
- ".codex/skills/strategy-backtest-research/references/research-notes.md": "e65eb2124115a11a5be637964efed06961c7a90eff33d95eaf238b521d1f505a",
33
+ ".codex/skills/strategy-backtest-research/references/research-notes.md": "e7b85dce73c612ba6821857e511b63754a574cc14d8ba5224466002fe27509aa",
34
34
  ".codex/skills/strategy-backtest-research/scripts/backtest-run-metrics.mjs": "dbcc6bb7cab4e0108ed11df503b8a0a337d1a01b77c5b56554ce1264a00f5d57",
35
35
  ".codex/skills/strategy-backtest-research/scripts/backtest-run-metrics.test.mjs": "a00d0ae5fb0037cc554282dbe22109a9e34cb64b93d26a6feab69005a2ba637f",
36
36
  ".codex/skills/strategy-backtest-research/scripts/fast-ai-export-metrics.mjs": "1d1747d34269ad15a610dcd687334c3d47ed9460d891285ad1dd851c7658d1af",
@@ -42,24 +42,24 @@
42
42
  ".codex/skills/strategy-forward-start/SKILL.md": "a86311aeb4c5b46e43011806ffa0f665c4ec0fddce8d7f8c1232184e4a14bfa7",
43
43
  ".codex/skills/strategy-forward-status/SKILL.md": "b41f117680259fd7b4d5cbe3c0a1be75a42f08c8cf78f9ad7ccb0d90bfc66269",
44
44
  ".codex/skills/strategy-improvement-plan/SKILL.md": "f5f0f390941dd3332890a2f0ffae1248abde4925db2cf49a41c0c0c1b47e7864",
45
- ".codex/skills/strategy-improvement-research/references/final-composition-board.md": "bc88878e088744ff933f06b54f73af96231d1f22dc162a44c9eb2e81d6cb51aa",
45
+ ".codex/skills/strategy-improvement-research/references/final-composition-board.md": "d7c4c9575af37fdff60548754da434511d2ef1ab9f0fca8dd211edf45e73e8d3",
46
46
  ".codex/skills/strategy-improvement-research/scripts/build-final-composition-spec.mjs": "b51873691fafbcf2dd70d8c9a6da842fec5f8358346c632cfe79211db08b18ba",
47
47
  ".codex/skills/strategy-improvement-research/scripts/build-final-composition-spec.test.mjs": "85d06b8fe09ee7611a3fb989173c361ba32f0e9364422925ff221bfe127b2bfa",
48
48
  ".codex/skills/strategy-improvement-research/scripts/final-composition-board.mjs": "c7834c9d567d3f02c39fcca2b0116a0097bc0384891a5185edbcafa23eeea1db",
49
49
  ".codex/skills/strategy-improvement-research/scripts/final-composition-board.test.mjs": "14572bbb8083f83b7e05b6dad2a9c93a88c212425b09cc31ee15ab7a684ea8c5",
50
50
  ".codex/skills/strategy-improvement-research/scripts/freeze-gate-variants.mjs": "17b3c868880ddd5b326d309e6e42eaeef1ea22d625a5db947997ac73b221ad6d",
51
51
  ".codex/skills/strategy-improvement-research/scripts/freeze-gate-variants.test.mjs": "963e1c64ab3deef11fd48a96e10f3827af38036dfe2eb28daf7849f19dd414d8",
52
- ".codex/skills/strategy-improvement-research/SKILL.md": "6964eb154e40b0b775aa2735b2ebbe998984479f46333ed03eab88d2ee150d48",
52
+ ".codex/skills/strategy-improvement-research/SKILL.md": "75a0f2420305b269fec0f71f63904b5bf04be00612b015855f0a0f2beface5b7",
53
53
  ".codex/skills/strategy-period-revalidate/SKILL.md": "f367ba1c632506a7d9bf92662a40b61df55f7a56ad04f3600a475b3c365bbbcb",
54
54
  ".codex/skills/strategy-release/agents/openai.yaml": "46ddb09119b08ae3be3b9509121b57492342012e774789b615767b24270dba17",
55
55
  ".codex/skills/strategy-release/references/diagnose-live.md": "dc2effa6eba34e3cfcc1eaf4d2fd1c96d6ccdda971a0edec76c9eed1ec293a3d",
56
56
  ".codex/skills/strategy-release/references/direction-policy.md": "e7635e709dfc7301b94f3dbc1421c6a779c7c39ce4f6066b2eece38e4e600bec",
57
57
  ".codex/skills/strategy-release/references/directional-parameter-split.md": "c303fbd15183b4b8c5662526d5bc762edc0a21a42c45e141d0dc5efdd3b52d8c",
58
58
  ".codex/skills/strategy-release/references/evidence-limitations.md": "523d6541a230b51c2b0b81a6c11c99e3aef69203f22b77e94fda1e3b15321e20",
59
- ".codex/skills/strategy-release/references/evidence-retention.md": "8fc5ccbc1003e8dc524f12dd56592eb72839c251245c6f68fd495fd4871e4993",
59
+ ".codex/skills/strategy-release/references/evidence-retention.md": "ed39682e5d1cf33b5df0c1fb4896a889e0f53c4dc27ef37efe3d35d2f08a205c",
60
60
  ".codex/skills/strategy-release/references/historical-hypothesis-audit.md": "3c4d579a0c097e974592b855891446060eb70d8bddf8558d765052a92509b516",
61
61
  ".codex/skills/strategy-release/references/professional-research-loop.md": "e440f05dd6b5c803ba20da9f019f745f89477724d46551c53dbb8dab9cffe40f",
62
- ".codex/skills/strategy-release/references/release-workflow.md": "aeda374c8b0a5d14629b2a1476a9c6e67e556e7ff7709b566d532334f39863eb",
62
+ ".codex/skills/strategy-release/references/release-workflow.md": "a103416334b5605addd856f105462b53602d29ace2816f96162b7e2ecd1a0bb6",
63
63
  ".codex/skills/strategy-release/references/research-objective.md": "650a29260ff7a2ee5dcac89500f59b4c81c2fbd650d6d247a1c799e45814798a",
64
64
  ".codex/skills/strategy-release/references/verdict-contract.md": "7cd2d0ae8c8264da36319a78ff80d76893d5233cbceba4d076973f32b2bb2fe9",
65
65
  ".codex/skills/strategy-release/scripts/direction-policy-checkpoint.mjs": "739240707323ec80b16191b447e76dd2295e946564c3a67f8472949d6129b3a1",
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "create-tradejs",
3
- "version": "3.1.29-beta.260",
3
+ "version": "3.1.29-beta.262",
4
4
  "description": "Create a ready-to-run TradeJS project with local infrastructure and the Web UI.",
5
5
  "keywords": [
6
6
  "tradejs",