merge-cli 3.7__tar.gz → 3.9.0__tar.gz

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (50) hide show
  1. merge_cli-3.9.0/MANIFEST.in +6 -0
  2. merge_cli-3.9.0/PKG-INFO +232 -0
  3. merge_cli-3.9.0/README.md +199 -0
  4. merge_cli-3.9.0/TUTORIAL.md +613 -0
  5. merge_cli-3.9.0/merge_cli/__init__.py +7 -0
  6. merge_cli-3.9.0/merge_cli/api.py +204 -0
  7. {merge_cli-3.7 → merge_cli-3.9.0}/merge_cli/cli.py +492 -151
  8. {merge_cli-3.7 → merge_cli-3.9.0}/merge_cli/config.py +12 -1
  9. merge_cli-3.9.0/merge_cli/data/models/MERGE_coding.pkl +0 -0
  10. merge_cli-3.9.0/merge_cli/data/models/MERGE_noncoding.pkl +0 -0
  11. merge_cli-3.9.0/merge_cli/data/models/MERGE_nonsnv.pkl +0 -0
  12. merge_cli-3.9.0/merge_cli/data/models/MERGE_splicing.pkl +0 -0
  13. merge_cli-3.9.0/merge_cli/data/models/MODEL_CARD.md +40 -0
  14. merge_cli-3.9.0/merge_cli/data/models/ensemble_predict.py +596 -0
  15. merge_cli-3.9.0/merge_cli/local_engine.py +3233 -0
  16. merge_cli-3.9.0/merge_cli/mcp_install.py +414 -0
  17. merge_cli-3.9.0/merge_cli/mcp_server.py +313 -0
  18. {merge_cli-3.7 → merge_cli-3.9.0}/merge_cli/output.py +93 -10
  19. merge_cli-3.9.0/merge_cli.egg-info/PKG-INFO +232 -0
  20. merge_cli-3.9.0/merge_cli.egg-info/SOURCES.txt +33 -0
  21. {merge_cli-3.7 → merge_cli-3.9.0}/merge_cli.egg-info/entry_points.txt +1 -0
  22. {merge_cli-3.7 → merge_cli-3.9.0}/merge_cli.egg-info/requires.txt +11 -2
  23. merge_cli-3.9.0/pyproject.toml +86 -0
  24. merge_cli-3.9.0/tests/test_ep_module_isolation.py +108 -0
  25. merge_cli-3.9.0/tests/test_evo2_polling.py +148 -0
  26. merge_cli-3.9.0/tests/test_mcp_install.py +349 -0
  27. merge_cli-3.9.0/tests/test_semicolon_features.py +96 -0
  28. merge_cli-3.9.0/tests/test_sklearn_compat.py +109 -0
  29. merge_cli-3.9.0/tests/test_smoke_v38.py +262 -0
  30. merge_cli-3.7/MANIFEST.in +0 -1
  31. merge_cli-3.7/PKG-INFO +0 -88
  32. merge_cli-3.7/README.md +0 -61
  33. merge_cli-3.7/merge_cli/__init__.py +0 -1
  34. merge_cli-3.7/merge_cli/api.py +0 -153
  35. merge_cli-3.7/merge_cli/cli_env_patch.py +0 -132
  36. merge_cli-3.7/merge_cli/data/models/BestModel_coding.pkl +0 -0
  37. merge_cli-3.7/merge_cli/data/models/BestModel_noncoding.pkl +0 -0
  38. merge_cli-3.7/merge_cli/data/models/BestModel_splice.pkl +0 -0
  39. merge_cli-3.7/merge_cli/data/models/ensemble_predict.py +0 -752
  40. merge_cli-3.7/merge_cli/ensemble_predict.py +0 -752
  41. merge_cli-3.7/merge_cli/local_engine.py +0 -1810
  42. merge_cli-3.7/merge_cli.egg-info/PKG-INFO +0 -88
  43. merge_cli-3.7/merge_cli.egg-info/SOURCES.txt +0 -24
  44. merge_cli-3.7/pyproject.toml +0 -46
  45. {merge_cli-3.7 → merge_cli-3.9.0}/merge_cli/data/__init__.py +0 -0
  46. {merge_cli-3.7 → merge_cli-3.9.0}/merge_cli/data/models/.gitkeep +0 -0
  47. {merge_cli-3.7 → merge_cli-3.9.0}/merge_cli/data/models/__init__.py +0 -0
  48. {merge_cli-3.7 → merge_cli-3.9.0}/merge_cli.egg-info/dependency_links.txt +0 -0
  49. {merge_cli-3.7 → merge_cli-3.9.0}/merge_cli.egg-info/top_level.txt +0 -0
  50. {merge_cli-3.7 → merge_cli-3.9.0}/setup.cfg +0 -0
@@ -0,0 +1,6 @@
1
+ recursive-include merge_cli/data/models *.pkl *.py *.md .gitkeep
2
+ # README is included automatically as the long_description; TUTORIAL is not, and
3
+ # README links to it — a source distribution without it has a dead link.
4
+ include TUTORIAL.md
5
+ include MODEL_CARD.md
6
+ recursive-include tests *.py
@@ -0,0 +1,232 @@
1
+ Metadata-Version: 2.4
2
+ Name: merge-cli
3
+ Version: 3.9.0
4
+ Summary: MERGE variant pathogenicity prediction CLI: 4 variant types, 14 model features, local/remote modes, MCP server
5
+ Author: MERGE Team
6
+ License-Expression: LicenseRef-Proprietary
7
+ Project-URL: Homepage, https://merge.fanglab.cn
8
+ Keywords: bioinformatics,variant,pathogenicity,genomics,cli
9
+ Classifier: Programming Language :: Python :: 3
10
+ Classifier: Programming Language :: Python :: 3.11
11
+ Classifier: Programming Language :: Python :: 3.12
12
+ Classifier: Programming Language :: Python :: 3.13
13
+ Classifier: Environment :: Console
14
+ Classifier: Topic :: Scientific/Engineering :: Bio-Informatics
15
+ Requires-Python: >=3.11
16
+ Description-Content-Type: text/markdown
17
+ Requires-Dist: click>=8.1
18
+ Requires-Dist: rich>=13.0
19
+ Requires-Dist: requests>=2.31
20
+ Requires-Dist: numpy>=1.26
21
+ Requires-Dist: pandas>=2.1
22
+ Requires-Dist: scikit-learn>=1.8
23
+ Requires-Dist: joblib>=1.3
24
+ Requires-Dist: matplotlib>=3.8
25
+ Requires-Dist: shap>=0.44
26
+ Requires-Dist: catboost>=1.2
27
+ Provides-Extra: mcp
28
+ Requires-Dist: mcp<2,>=1.2; extra == "mcp"
29
+ Provides-Extra: local
30
+ Requires-Dist: pysam>=0.22; extra == "local"
31
+ Provides-Extra: all
32
+ Requires-Dist: merge-cli[local,mcp]; extra == "all"
33
+
34
+ # merge-cli 3.8
35
+
36
+ MERGE variant pathogenicity prediction CLI — the professional-user counterpart to
37
+ [merge.fanglab.cn](https://merge.fanglab.cn), with the same models and the same
38
+ ensemble, runnable entirely on your own machine.
39
+
40
+ **New here? Read [TUTORIAL.md](TUTORIAL.md)** — install, both modes, per-model
41
+ setup, scripting and troubleshooting.
42
+
43
+ ## Highlights
44
+
45
+ - **Four variant types**: coding, splice, noncoding and **non-SNV** (indel / MNV / delins).
46
+ - **13 models**: AlphaGenome, HyenaDNA, NT, AlphaMissense, ESM-1b, GPN-MSA,
47
+ Evo2-7B, Evo2-7B-base, Enformer, GENERATOR, GENERATOR-v2, NT-v2, Carbon-3B.
48
+ (Evo2-1B-base is also supported locally where the GPU allows it.)
49
+ - **Remote and local modes.** Local mode deploys each model into its own conda
50
+ environment and serves it over HTTP — no Docker, no gateway required.
51
+ - Local precomputed VCF cache lookup for hg38/hg19 coding and splicing SNVs.
52
+ - Bundled MERGE ensemble models (the same four deployment bundles the website runs).
53
+
54
+ ## Quick Start
55
+
56
+ ```bash
57
+ pip install merge-cli
58
+ merge --help
59
+ merge predict --chrom chr1 --pos 69428 --ref T --alt G --genome hg38
60
+ ```
61
+
62
+ Optional extras: `merge-cli[mcp]` to ask for predictions in chat (see below),
63
+ `merge-cli[local]` for local mode (adds pysam, which has no Windows wheels —
64
+ which is why it is not in the base install), or `merge-cli[all]` for both.
65
+
66
+ ## Ask in chat (MCP)
67
+
68
+ ```bash
69
+ pip install "merge-cli[mcp]"
70
+ merge mcp install # detects Claude Desktop / Claude Code, writes the config
71
+ merge mcp status
72
+ ```
73
+
74
+ Then restart your client and ask *"Is chr1:1040819 G>GC pathogenic?"*. The MCP
75
+ server runs in remote mode only — no GPU needed. Details in
76
+ [TUTORIAL.md §11](TUTORIAL.md).
77
+
78
+ ## Variant types and ensemble models
79
+
80
+ The task is chosen automatically: a variant whose REF or ALT is not a single base
81
+ is scored as non-SNV; otherwise ANNOVAR decides coding / splice / noncoding.
82
+ Override with `--ensemble-type coding|splice|noncoding|nonsnv`.
83
+
84
+ | Task | Bundle | Algorithm | Features |
85
+ |---|---|---|---|
86
+ | coding | `MERGE_coding.pkl` | LR_L1 | 6 |
87
+ | splice | `MERGE_splicing.pkl` | CatBoost | 10 |
88
+ | noncoding | `MERGE_noncoding.pkl` | LR_L2 | 10 |
89
+ | non-SNV | `MERGE_nonsnv.pkl` | CatBoost | 10 |
90
+
91
+ Each bundle embeds its own preprocessing (median imputation → z-scoring →
92
+ classifier), so raw model scores are fed straight in and any model you did not
93
+ run is median-imputed. See `merge_cli/data/models/MODEL_CARD.md` for the exact
94
+ feature panels.
95
+
96
+ ## Models
97
+
98
+ All models run by default — each one feeds at least one ensemble panel, and a
99
+ model you skip has its feature median-imputed instead. Skip any with `--no-<model>`:
100
+
101
+ ```bash
102
+ merge predict --chrom chr1 --pos 69428 --ref T --alt G --no-carbon --no-nt-v2
103
+ ```
104
+
105
+ Available switches: `--no-alphagenome --no-hyenadna --no-nt --no-alphamissense
106
+ --no-esm1b --no-gpn-msa --no-evo2 --no-enformer --no-generator
107
+ --no-generator-v2 --no-nt-v2 --no-carbon`.
108
+
109
+ In local mode every model needs its service running; the CLI reports which
110
+ features were computed and which were imputed on every prediction.
111
+
112
+ ### Evo2 variants
113
+
114
+ `--evo2-model` picks the weights for the primary Evo2 score (remote and local):
115
+
116
+ ```bash
117
+ merge predict --chrom chr1 --pos 69428 --ref T --alt G --evo2-model evo2_7b_base
118
+ ```
119
+
120
+ All three Evo2 variants are scored by default in local mode —
121
+ `evo2_7b_score`, `evo2_7b_base_score` and `evo2_1b_base_score` are separate
122
+ ensemble features. They share one GPU-resident service and are scored one at a
123
+ time, costing a few seconds each. Narrow the set to trade coverage for speed:
124
+
125
+ ```bash
126
+ merge local predict --chrom chr1 --pos 69428 --ref T --alt G \
127
+ --evo2-models evo2_7b # primary variant only, faster
128
+ ```
129
+
130
+ `evo2_1b_base` needs a GPU with FP8 (compute capability ≥ 8.9: RTX 4000/5000 Ada,
131
+ L40S, H100). Elsewhere — including in remote mode — its feature is left missing and
132
+ median-imputed by the ensemble, which is what the website does too.
133
+
134
+ ## Local mode
135
+
136
+ ```bash
137
+ merge doctor # check GPU and services
138
+ merge local setup # configure paths
139
+ merge local env setup --model evo2 # deploy one model
140
+ merge local env setup --model carbon # ... or another
141
+ merge local env setup --model all # everything (large!)
142
+ merge local env start --model evo2
143
+ merge local env status
144
+ merge local predict --chrom chr17 --pos 43092919 --ref A --alt G
145
+ ```
146
+
147
+ Setting a reference FASTA is strongly recommended — the local services use it to
148
+ cut the exact sequence windows the models were scored with during training:
149
+
150
+ ```bash
151
+ merge local setup # answer the "Reference genome FASTA path" prompt
152
+ # or
153
+ merge local predict ... --genome-ref /path/to/hg38.fa
154
+ ```
155
+
156
+ Without it the services fall back to fetching sequence context from the MERGE
157
+ server, which requires network access.
158
+
159
+ ### Service ports
160
+
161
+ | Model | Env | Port |
162
+ |---|---|---|
163
+ | AlphaGenome | `alphagenome` | 5000 |
164
+ | HyenaDNA | `hyenadna` | 5001 |
165
+ | NT | `nt` | 5002 |
166
+ | Evo2 (all variants) | `evo2` | 5003 |
167
+ | Enformer | `enformer` | 5004 |
168
+ | Carbon | `carbon` | 5005 |
169
+ | GENERATOR | `generator` | 5006 |
170
+ | GENERATOR-v2 | `generator_v2` | 5007 |
171
+ | NT-v2 | `nt_v2` | 5008 |
172
+
173
+ Each service exposes `GET /health` and `POST /predict` (`{chrom, pos, ref, alt,
174
+ genome_version}`); the DNA-LM services also expose `POST /batch_predict`.
175
+
176
+ Each model gets its own conda env, service and port by default, and there is no
177
+ gateway to stand up — so you can deploy only the models you actually need.
178
+
179
+ Environments may also be **shared**: any environment with the right dependencies
180
+ (torch + `transformers<5` + pyfaidx) can serve several of these models, which
181
+ saves tens of GB. Start the service with that environment's interpreter, e.g.
182
+ `~/miniconda3/envs/NT/bin/python ~/.merge-local-servers/generator_server.py`.
183
+ `merge local env status` shows which environment is really behind each port.
184
+
185
+ ## Precomputed VCF Cache
186
+
187
+ ```bash
188
+ merge precomputed configure --data-dir /path/to/precomputed
189
+ merge precomputed status
190
+ merge precomputed download --genome all --variant-type all
191
+ ```
192
+
193
+ Expected files: `coding_merged.vcf.gz`, `splicing_merged.vcf.gz`, their hg19
194
+ counterparts, and `.tbi` indices. The cache only covers SNVs and the older
195
+ feature set, so it is bypassed for non-SNVs and whenever an opt-in DNA-LM is
196
+ requested.
197
+
198
+ ## Scoring conventions
199
+
200
+ The local services reproduce the exact quantities the ensemble was trained on.
201
+ Changing any of them silently invalidates the MERGE score:
202
+
203
+ | Model | Score | Window | VRAM |
204
+ |---|---|---|---|
205
+ | Evo2-7B / -7B-base | mean_PLL(REF) − mean_PLL(ALT) | 8192 bp | ~20 GB |
206
+ | Evo2-1B-base | mean_PLL(REF) − mean_PLL(ALT) | 8192 bp | ~8 GB, needs FP8 |
207
+ | Carbon-3B | logP_sum(REF) − logP_sum(ALT) | 24576 bp | **~20 GB** |
208
+ | GENERATOR / -v2 | NLL(ALT) − NLL(REF) | 600 bp, max_length 128 | ~6 GB |
209
+ | NT-v2 | ‖emb(ALT) − emb(REF)‖₂ | 600 bp, max_length 128 | ~3 GB |
210
+ | NT | ‖emb(ALT) − emb(REF)‖₂ | 8192 bp | ~3 GB |
211
+ | HyenaDNA | log p(ALT) − log p(REF), next token | 1000 bp upstream | ~2 GB |
212
+
213
+ These windows are not tunable, on purpose. The MERGE bundles were fitted on
214
+ features computed exactly this way, so a shorter window does not merely cost
215
+ accuracy — it puts the feature on a different scale than the model expects, and
216
+ the resulting MERGE score is wrong without anything looking wrong. If a model
217
+ does not fit on your GPU, leave it off: the ensemble median-imputes what is
218
+ missing, which is the honest fallback.
219
+
220
+ Carbon in particular needs roughly 20 GB of VRAM at the 24576 bp training window.
221
+
222
+ The `transformers` version is pinned to 4.x in the generated environments:
223
+ NT-v2's remote code does not load under transformers 5.x. `alphagenome` is
224
+ pinned to 0.5.1 to match the MERGE server, so both modes return the same
225
+ AlphaGenome features.
226
+
227
+ AlphaGenome is the one model whose features are **not** reproducible over time:
228
+ it runs on Google's servers and is updated there, so today's values differ from
229
+ those in the training tables. Every other model reproduces bit-for-bit.
230
+
231
+ If local services are unavailable, the errors are reported under
232
+ `prediction.errors` and the CLI continues with the models it could reach.
@@ -0,0 +1,199 @@
1
+ # merge-cli 3.8
2
+
3
+ MERGE variant pathogenicity prediction CLI — the professional-user counterpart to
4
+ [merge.fanglab.cn](https://merge.fanglab.cn), with the same models and the same
5
+ ensemble, runnable entirely on your own machine.
6
+
7
+ **New here? Read [TUTORIAL.md](TUTORIAL.md)** — install, both modes, per-model
8
+ setup, scripting and troubleshooting.
9
+
10
+ ## Highlights
11
+
12
+ - **Four variant types**: coding, splice, noncoding and **non-SNV** (indel / MNV / delins).
13
+ - **13 models**: AlphaGenome, HyenaDNA, NT, AlphaMissense, ESM-1b, GPN-MSA,
14
+ Evo2-7B, Evo2-7B-base, Enformer, GENERATOR, GENERATOR-v2, NT-v2, Carbon-3B.
15
+ (Evo2-1B-base is also supported locally where the GPU allows it.)
16
+ - **Remote and local modes.** Local mode deploys each model into its own conda
17
+ environment and serves it over HTTP — no Docker, no gateway required.
18
+ - Local precomputed VCF cache lookup for hg38/hg19 coding and splicing SNVs.
19
+ - Bundled MERGE ensemble models (the same four deployment bundles the website runs).
20
+
21
+ ## Quick Start
22
+
23
+ ```bash
24
+ pip install merge-cli
25
+ merge --help
26
+ merge predict --chrom chr1 --pos 69428 --ref T --alt G --genome hg38
27
+ ```
28
+
29
+ Optional extras: `merge-cli[mcp]` to ask for predictions in chat (see below),
30
+ `merge-cli[local]` for local mode (adds pysam, which has no Windows wheels —
31
+ which is why it is not in the base install), or `merge-cli[all]` for both.
32
+
33
+ ## Ask in chat (MCP)
34
+
35
+ ```bash
36
+ pip install "merge-cli[mcp]"
37
+ merge mcp install # detects Claude Desktop / Claude Code, writes the config
38
+ merge mcp status
39
+ ```
40
+
41
+ Then restart your client and ask *"Is chr1:1040819 G>GC pathogenic?"*. The MCP
42
+ server runs in remote mode only — no GPU needed. Details in
43
+ [TUTORIAL.md §11](TUTORIAL.md).
44
+
45
+ ## Variant types and ensemble models
46
+
47
+ The task is chosen automatically: a variant whose REF or ALT is not a single base
48
+ is scored as non-SNV; otherwise ANNOVAR decides coding / splice / noncoding.
49
+ Override with `--ensemble-type coding|splice|noncoding|nonsnv`.
50
+
51
+ | Task | Bundle | Algorithm | Features |
52
+ |---|---|---|---|
53
+ | coding | `MERGE_coding.pkl` | LR_L1 | 6 |
54
+ | splice | `MERGE_splicing.pkl` | CatBoost | 10 |
55
+ | noncoding | `MERGE_noncoding.pkl` | LR_L2 | 10 |
56
+ | non-SNV | `MERGE_nonsnv.pkl` | CatBoost | 10 |
57
+
58
+ Each bundle embeds its own preprocessing (median imputation → z-scoring →
59
+ classifier), so raw model scores are fed straight in and any model you did not
60
+ run is median-imputed. See `merge_cli/data/models/MODEL_CARD.md` for the exact
61
+ feature panels.
62
+
63
+ ## Models
64
+
65
+ All models run by default — each one feeds at least one ensemble panel, and a
66
+ model you skip has its feature median-imputed instead. Skip any with `--no-<model>`:
67
+
68
+ ```bash
69
+ merge predict --chrom chr1 --pos 69428 --ref T --alt G --no-carbon --no-nt-v2
70
+ ```
71
+
72
+ Available switches: `--no-alphagenome --no-hyenadna --no-nt --no-alphamissense
73
+ --no-esm1b --no-gpn-msa --no-evo2 --no-enformer --no-generator
74
+ --no-generator-v2 --no-nt-v2 --no-carbon`.
75
+
76
+ In local mode every model needs its service running; the CLI reports which
77
+ features were computed and which were imputed on every prediction.
78
+
79
+ ### Evo2 variants
80
+
81
+ `--evo2-model` picks the weights for the primary Evo2 score (remote and local):
82
+
83
+ ```bash
84
+ merge predict --chrom chr1 --pos 69428 --ref T --alt G --evo2-model evo2_7b_base
85
+ ```
86
+
87
+ All three Evo2 variants are scored by default in local mode —
88
+ `evo2_7b_score`, `evo2_7b_base_score` and `evo2_1b_base_score` are separate
89
+ ensemble features. They share one GPU-resident service and are scored one at a
90
+ time, costing a few seconds each. Narrow the set to trade coverage for speed:
91
+
92
+ ```bash
93
+ merge local predict --chrom chr1 --pos 69428 --ref T --alt G \
94
+ --evo2-models evo2_7b # primary variant only, faster
95
+ ```
96
+
97
+ `evo2_1b_base` needs a GPU with FP8 (compute capability ≥ 8.9: RTX 4000/5000 Ada,
98
+ L40S, H100). Elsewhere — including in remote mode — its feature is left missing and
99
+ median-imputed by the ensemble, which is what the website does too.
100
+
101
+ ## Local mode
102
+
103
+ ```bash
104
+ merge doctor # check GPU and services
105
+ merge local setup # configure paths
106
+ merge local env setup --model evo2 # deploy one model
107
+ merge local env setup --model carbon # ... or another
108
+ merge local env setup --model all # everything (large!)
109
+ merge local env start --model evo2
110
+ merge local env status
111
+ merge local predict --chrom chr17 --pos 43092919 --ref A --alt G
112
+ ```
113
+
114
+ Setting a reference FASTA is strongly recommended — the local services use it to
115
+ cut the exact sequence windows the models were scored with during training:
116
+
117
+ ```bash
118
+ merge local setup # answer the "Reference genome FASTA path" prompt
119
+ # or
120
+ merge local predict ... --genome-ref /path/to/hg38.fa
121
+ ```
122
+
123
+ Without it the services fall back to fetching sequence context from the MERGE
124
+ server, which requires network access.
125
+
126
+ ### Service ports
127
+
128
+ | Model | Env | Port |
129
+ |---|---|---|
130
+ | AlphaGenome | `alphagenome` | 5000 |
131
+ | HyenaDNA | `hyenadna` | 5001 |
132
+ | NT | `nt` | 5002 |
133
+ | Evo2 (all variants) | `evo2` | 5003 |
134
+ | Enformer | `enformer` | 5004 |
135
+ | Carbon | `carbon` | 5005 |
136
+ | GENERATOR | `generator` | 5006 |
137
+ | GENERATOR-v2 | `generator_v2` | 5007 |
138
+ | NT-v2 | `nt_v2` | 5008 |
139
+
140
+ Each service exposes `GET /health` and `POST /predict` (`{chrom, pos, ref, alt,
141
+ genome_version}`); the DNA-LM services also expose `POST /batch_predict`.
142
+
143
+ Each model gets its own conda env, service and port by default, and there is no
144
+ gateway to stand up — so you can deploy only the models you actually need.
145
+
146
+ Environments may also be **shared**: any environment with the right dependencies
147
+ (torch + `transformers<5` + pyfaidx) can serve several of these models, which
148
+ saves tens of GB. Start the service with that environment's interpreter, e.g.
149
+ `~/miniconda3/envs/NT/bin/python ~/.merge-local-servers/generator_server.py`.
150
+ `merge local env status` shows which environment is really behind each port.
151
+
152
+ ## Precomputed VCF Cache
153
+
154
+ ```bash
155
+ merge precomputed configure --data-dir /path/to/precomputed
156
+ merge precomputed status
157
+ merge precomputed download --genome all --variant-type all
158
+ ```
159
+
160
+ Expected files: `coding_merged.vcf.gz`, `splicing_merged.vcf.gz`, their hg19
161
+ counterparts, and `.tbi` indices. The cache only covers SNVs and the older
162
+ feature set, so it is bypassed for non-SNVs and whenever an opt-in DNA-LM is
163
+ requested.
164
+
165
+ ## Scoring conventions
166
+
167
+ The local services reproduce the exact quantities the ensemble was trained on.
168
+ Changing any of them silently invalidates the MERGE score:
169
+
170
+ | Model | Score | Window | VRAM |
171
+ |---|---|---|---|
172
+ | Evo2-7B / -7B-base | mean_PLL(REF) − mean_PLL(ALT) | 8192 bp | ~20 GB |
173
+ | Evo2-1B-base | mean_PLL(REF) − mean_PLL(ALT) | 8192 bp | ~8 GB, needs FP8 |
174
+ | Carbon-3B | logP_sum(REF) − logP_sum(ALT) | 24576 bp | **~20 GB** |
175
+ | GENERATOR / -v2 | NLL(ALT) − NLL(REF) | 600 bp, max_length 128 | ~6 GB |
176
+ | NT-v2 | ‖emb(ALT) − emb(REF)‖₂ | 600 bp, max_length 128 | ~3 GB |
177
+ | NT | ‖emb(ALT) − emb(REF)‖₂ | 8192 bp | ~3 GB |
178
+ | HyenaDNA | log p(ALT) − log p(REF), next token | 1000 bp upstream | ~2 GB |
179
+
180
+ These windows are not tunable, on purpose. The MERGE bundles were fitted on
181
+ features computed exactly this way, so a shorter window does not merely cost
182
+ accuracy — it puts the feature on a different scale than the model expects, and
183
+ the resulting MERGE score is wrong without anything looking wrong. If a model
184
+ does not fit on your GPU, leave it off: the ensemble median-imputes what is
185
+ missing, which is the honest fallback.
186
+
187
+ Carbon in particular needs roughly 20 GB of VRAM at the 24576 bp training window.
188
+
189
+ The `transformers` version is pinned to 4.x in the generated environments:
190
+ NT-v2's remote code does not load under transformers 5.x. `alphagenome` is
191
+ pinned to 0.5.1 to match the MERGE server, so both modes return the same
192
+ AlphaGenome features.
193
+
194
+ AlphaGenome is the one model whose features are **not** reproducible over time:
195
+ it runs on Google's servers and is updated there, so today's values differ from
196
+ those in the training tables. Every other model reproduces bit-for-bit.
197
+
198
+ If local services are unavailable, the errors are reported under
199
+ `prediction.errors` and the CLI continues with the models it could reach.