merge-cli 3.7__tar.gz → 3.9.0__tar.gz
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- merge_cli-3.9.0/MANIFEST.in +6 -0
- merge_cli-3.9.0/PKG-INFO +232 -0
- merge_cli-3.9.0/README.md +199 -0
- merge_cli-3.9.0/TUTORIAL.md +613 -0
- merge_cli-3.9.0/merge_cli/__init__.py +7 -0
- merge_cli-3.9.0/merge_cli/api.py +204 -0
- {merge_cli-3.7 → merge_cli-3.9.0}/merge_cli/cli.py +492 -151
- {merge_cli-3.7 → merge_cli-3.9.0}/merge_cli/config.py +12 -1
- merge_cli-3.9.0/merge_cli/data/models/MERGE_coding.pkl +0 -0
- merge_cli-3.9.0/merge_cli/data/models/MERGE_noncoding.pkl +0 -0
- merge_cli-3.9.0/merge_cli/data/models/MERGE_nonsnv.pkl +0 -0
- merge_cli-3.9.0/merge_cli/data/models/MERGE_splicing.pkl +0 -0
- merge_cli-3.9.0/merge_cli/data/models/MODEL_CARD.md +40 -0
- merge_cli-3.9.0/merge_cli/data/models/ensemble_predict.py +596 -0
- merge_cli-3.9.0/merge_cli/local_engine.py +3233 -0
- merge_cli-3.9.0/merge_cli/mcp_install.py +414 -0
- merge_cli-3.9.0/merge_cli/mcp_server.py +313 -0
- {merge_cli-3.7 → merge_cli-3.9.0}/merge_cli/output.py +93 -10
- merge_cli-3.9.0/merge_cli.egg-info/PKG-INFO +232 -0
- merge_cli-3.9.0/merge_cli.egg-info/SOURCES.txt +33 -0
- {merge_cli-3.7 → merge_cli-3.9.0}/merge_cli.egg-info/entry_points.txt +1 -0
- {merge_cli-3.7 → merge_cli-3.9.0}/merge_cli.egg-info/requires.txt +11 -2
- merge_cli-3.9.0/pyproject.toml +86 -0
- merge_cli-3.9.0/tests/test_ep_module_isolation.py +108 -0
- merge_cli-3.9.0/tests/test_evo2_polling.py +148 -0
- merge_cli-3.9.0/tests/test_mcp_install.py +349 -0
- merge_cli-3.9.0/tests/test_semicolon_features.py +96 -0
- merge_cli-3.9.0/tests/test_sklearn_compat.py +109 -0
- merge_cli-3.9.0/tests/test_smoke_v38.py +262 -0
- merge_cli-3.7/MANIFEST.in +0 -1
- merge_cli-3.7/PKG-INFO +0 -88
- merge_cli-3.7/README.md +0 -61
- merge_cli-3.7/merge_cli/__init__.py +0 -1
- merge_cli-3.7/merge_cli/api.py +0 -153
- merge_cli-3.7/merge_cli/cli_env_patch.py +0 -132
- merge_cli-3.7/merge_cli/data/models/BestModel_coding.pkl +0 -0
- merge_cli-3.7/merge_cli/data/models/BestModel_noncoding.pkl +0 -0
- merge_cli-3.7/merge_cli/data/models/BestModel_splice.pkl +0 -0
- merge_cli-3.7/merge_cli/data/models/ensemble_predict.py +0 -752
- merge_cli-3.7/merge_cli/ensemble_predict.py +0 -752
- merge_cli-3.7/merge_cli/local_engine.py +0 -1810
- merge_cli-3.7/merge_cli.egg-info/PKG-INFO +0 -88
- merge_cli-3.7/merge_cli.egg-info/SOURCES.txt +0 -24
- merge_cli-3.7/pyproject.toml +0 -46
- {merge_cli-3.7 → merge_cli-3.9.0}/merge_cli/data/__init__.py +0 -0
- {merge_cli-3.7 → merge_cli-3.9.0}/merge_cli/data/models/.gitkeep +0 -0
- {merge_cli-3.7 → merge_cli-3.9.0}/merge_cli/data/models/__init__.py +0 -0
- {merge_cli-3.7 → merge_cli-3.9.0}/merge_cli.egg-info/dependency_links.txt +0 -0
- {merge_cli-3.7 → merge_cli-3.9.0}/merge_cli.egg-info/top_level.txt +0 -0
- {merge_cli-3.7 → merge_cli-3.9.0}/setup.cfg +0 -0
|
@@ -0,0 +1,6 @@
|
|
|
1
|
+
recursive-include merge_cli/data/models *.pkl *.py *.md .gitkeep
|
|
2
|
+
# README is included automatically as the long_description; TUTORIAL is not, and
|
|
3
|
+
# README links to it — a source distribution without it has a dead link.
|
|
4
|
+
include TUTORIAL.md
|
|
5
|
+
include MODEL_CARD.md
|
|
6
|
+
recursive-include tests *.py
|
merge_cli-3.9.0/PKG-INFO
ADDED
|
@@ -0,0 +1,232 @@
|
|
|
1
|
+
Metadata-Version: 2.4
|
|
2
|
+
Name: merge-cli
|
|
3
|
+
Version: 3.9.0
|
|
4
|
+
Summary: MERGE variant pathogenicity prediction CLI: 4 variant types, 14 model features, local/remote modes, MCP server
|
|
5
|
+
Author: MERGE Team
|
|
6
|
+
License-Expression: LicenseRef-Proprietary
|
|
7
|
+
Project-URL: Homepage, https://merge.fanglab.cn
|
|
8
|
+
Keywords: bioinformatics,variant,pathogenicity,genomics,cli
|
|
9
|
+
Classifier: Programming Language :: Python :: 3
|
|
10
|
+
Classifier: Programming Language :: Python :: 3.11
|
|
11
|
+
Classifier: Programming Language :: Python :: 3.12
|
|
12
|
+
Classifier: Programming Language :: Python :: 3.13
|
|
13
|
+
Classifier: Environment :: Console
|
|
14
|
+
Classifier: Topic :: Scientific/Engineering :: Bio-Informatics
|
|
15
|
+
Requires-Python: >=3.11
|
|
16
|
+
Description-Content-Type: text/markdown
|
|
17
|
+
Requires-Dist: click>=8.1
|
|
18
|
+
Requires-Dist: rich>=13.0
|
|
19
|
+
Requires-Dist: requests>=2.31
|
|
20
|
+
Requires-Dist: numpy>=1.26
|
|
21
|
+
Requires-Dist: pandas>=2.1
|
|
22
|
+
Requires-Dist: scikit-learn>=1.8
|
|
23
|
+
Requires-Dist: joblib>=1.3
|
|
24
|
+
Requires-Dist: matplotlib>=3.8
|
|
25
|
+
Requires-Dist: shap>=0.44
|
|
26
|
+
Requires-Dist: catboost>=1.2
|
|
27
|
+
Provides-Extra: mcp
|
|
28
|
+
Requires-Dist: mcp<2,>=1.2; extra == "mcp"
|
|
29
|
+
Provides-Extra: local
|
|
30
|
+
Requires-Dist: pysam>=0.22; extra == "local"
|
|
31
|
+
Provides-Extra: all
|
|
32
|
+
Requires-Dist: merge-cli[local,mcp]; extra == "all"
|
|
33
|
+
|
|
34
|
+
# merge-cli 3.8
|
|
35
|
+
|
|
36
|
+
MERGE variant pathogenicity prediction CLI — the professional-user counterpart to
|
|
37
|
+
[merge.fanglab.cn](https://merge.fanglab.cn), with the same models and the same
|
|
38
|
+
ensemble, runnable entirely on your own machine.
|
|
39
|
+
|
|
40
|
+
**New here? Read [TUTORIAL.md](TUTORIAL.md)** — install, both modes, per-model
|
|
41
|
+
setup, scripting and troubleshooting.
|
|
42
|
+
|
|
43
|
+
## Highlights
|
|
44
|
+
|
|
45
|
+
- **Four variant types**: coding, splice, noncoding and **non-SNV** (indel / MNV / delins).
|
|
46
|
+
- **13 models**: AlphaGenome, HyenaDNA, NT, AlphaMissense, ESM-1b, GPN-MSA,
|
|
47
|
+
Evo2-7B, Evo2-7B-base, Enformer, GENERATOR, GENERATOR-v2, NT-v2, Carbon-3B.
|
|
48
|
+
(Evo2-1B-base is also supported locally where the GPU allows it.)
|
|
49
|
+
- **Remote and local modes.** Local mode deploys each model into its own conda
|
|
50
|
+
environment and serves it over HTTP — no Docker, no gateway required.
|
|
51
|
+
- Local precomputed VCF cache lookup for hg38/hg19 coding and splicing SNVs.
|
|
52
|
+
- Bundled MERGE ensemble models (the same four deployment bundles the website runs).
|
|
53
|
+
|
|
54
|
+
## Quick Start
|
|
55
|
+
|
|
56
|
+
```bash
|
|
57
|
+
pip install merge-cli
|
|
58
|
+
merge --help
|
|
59
|
+
merge predict --chrom chr1 --pos 69428 --ref T --alt G --genome hg38
|
|
60
|
+
```
|
|
61
|
+
|
|
62
|
+
Optional extras: `merge-cli[mcp]` to ask for predictions in chat (see below),
|
|
63
|
+
`merge-cli[local]` for local mode (adds pysam, which has no Windows wheels —
|
|
64
|
+
which is why it is not in the base install), or `merge-cli[all]` for both.
|
|
65
|
+
|
|
66
|
+
## Ask in chat (MCP)
|
|
67
|
+
|
|
68
|
+
```bash
|
|
69
|
+
pip install "merge-cli[mcp]"
|
|
70
|
+
merge mcp install # detects Claude Desktop / Claude Code, writes the config
|
|
71
|
+
merge mcp status
|
|
72
|
+
```
|
|
73
|
+
|
|
74
|
+
Then restart your client and ask *"Is chr1:1040819 G>GC pathogenic?"*. The MCP
|
|
75
|
+
server runs in remote mode only — no GPU needed. Details in
|
|
76
|
+
[TUTORIAL.md §11](TUTORIAL.md).
|
|
77
|
+
|
|
78
|
+
## Variant types and ensemble models
|
|
79
|
+
|
|
80
|
+
The task is chosen automatically: a variant whose REF or ALT is not a single base
|
|
81
|
+
is scored as non-SNV; otherwise ANNOVAR decides coding / splice / noncoding.
|
|
82
|
+
Override with `--ensemble-type coding|splice|noncoding|nonsnv`.
|
|
83
|
+
|
|
84
|
+
| Task | Bundle | Algorithm | Features |
|
|
85
|
+
|---|---|---|---|
|
|
86
|
+
| coding | `MERGE_coding.pkl` | LR_L1 | 6 |
|
|
87
|
+
| splice | `MERGE_splicing.pkl` | CatBoost | 10 |
|
|
88
|
+
| noncoding | `MERGE_noncoding.pkl` | LR_L2 | 10 |
|
|
89
|
+
| non-SNV | `MERGE_nonsnv.pkl` | CatBoost | 10 |
|
|
90
|
+
|
|
91
|
+
Each bundle embeds its own preprocessing (median imputation → z-scoring →
|
|
92
|
+
classifier), so raw model scores are fed straight in and any model you did not
|
|
93
|
+
run is median-imputed. See `merge_cli/data/models/MODEL_CARD.md` for the exact
|
|
94
|
+
feature panels.
|
|
95
|
+
|
|
96
|
+
## Models
|
|
97
|
+
|
|
98
|
+
All models run by default — each one feeds at least one ensemble panel, and a
|
|
99
|
+
model you skip has its feature median-imputed instead. Skip any with `--no-<model>`:
|
|
100
|
+
|
|
101
|
+
```bash
|
|
102
|
+
merge predict --chrom chr1 --pos 69428 --ref T --alt G --no-carbon --no-nt-v2
|
|
103
|
+
```
|
|
104
|
+
|
|
105
|
+
Available switches: `--no-alphagenome --no-hyenadna --no-nt --no-alphamissense
|
|
106
|
+
--no-esm1b --no-gpn-msa --no-evo2 --no-enformer --no-generator
|
|
107
|
+
--no-generator-v2 --no-nt-v2 --no-carbon`.
|
|
108
|
+
|
|
109
|
+
In local mode every model needs its service running; the CLI reports which
|
|
110
|
+
features were computed and which were imputed on every prediction.
|
|
111
|
+
|
|
112
|
+
### Evo2 variants
|
|
113
|
+
|
|
114
|
+
`--evo2-model` picks the weights for the primary Evo2 score (remote and local):
|
|
115
|
+
|
|
116
|
+
```bash
|
|
117
|
+
merge predict --chrom chr1 --pos 69428 --ref T --alt G --evo2-model evo2_7b_base
|
|
118
|
+
```
|
|
119
|
+
|
|
120
|
+
All three Evo2 variants are scored by default in local mode —
|
|
121
|
+
`evo2_7b_score`, `evo2_7b_base_score` and `evo2_1b_base_score` are separate
|
|
122
|
+
ensemble features. They share one GPU-resident service and are scored one at a
|
|
123
|
+
time, costing a few seconds each. Narrow the set to trade coverage for speed:
|
|
124
|
+
|
|
125
|
+
```bash
|
|
126
|
+
merge local predict --chrom chr1 --pos 69428 --ref T --alt G \
|
|
127
|
+
--evo2-models evo2_7b # primary variant only, faster
|
|
128
|
+
```
|
|
129
|
+
|
|
130
|
+
`evo2_1b_base` needs a GPU with FP8 (compute capability ≥ 8.9: RTX 4000/5000 Ada,
|
|
131
|
+
L40S, H100). Elsewhere — including in remote mode — its feature is left missing and
|
|
132
|
+
median-imputed by the ensemble, which is what the website does too.
|
|
133
|
+
|
|
134
|
+
## Local mode
|
|
135
|
+
|
|
136
|
+
```bash
|
|
137
|
+
merge doctor # check GPU and services
|
|
138
|
+
merge local setup # configure paths
|
|
139
|
+
merge local env setup --model evo2 # deploy one model
|
|
140
|
+
merge local env setup --model carbon # ... or another
|
|
141
|
+
merge local env setup --model all # everything (large!)
|
|
142
|
+
merge local env start --model evo2
|
|
143
|
+
merge local env status
|
|
144
|
+
merge local predict --chrom chr17 --pos 43092919 --ref A --alt G
|
|
145
|
+
```
|
|
146
|
+
|
|
147
|
+
Setting a reference FASTA is strongly recommended — the local services use it to
|
|
148
|
+
cut the exact sequence windows the models were scored with during training:
|
|
149
|
+
|
|
150
|
+
```bash
|
|
151
|
+
merge local setup # answer the "Reference genome FASTA path" prompt
|
|
152
|
+
# or
|
|
153
|
+
merge local predict ... --genome-ref /path/to/hg38.fa
|
|
154
|
+
```
|
|
155
|
+
|
|
156
|
+
Without it the services fall back to fetching sequence context from the MERGE
|
|
157
|
+
server, which requires network access.
|
|
158
|
+
|
|
159
|
+
### Service ports
|
|
160
|
+
|
|
161
|
+
| Model | Env | Port |
|
|
162
|
+
|---|---|---|
|
|
163
|
+
| AlphaGenome | `alphagenome` | 5000 |
|
|
164
|
+
| HyenaDNA | `hyenadna` | 5001 |
|
|
165
|
+
| NT | `nt` | 5002 |
|
|
166
|
+
| Evo2 (all variants) | `evo2` | 5003 |
|
|
167
|
+
| Enformer | `enformer` | 5004 |
|
|
168
|
+
| Carbon | `carbon` | 5005 |
|
|
169
|
+
| GENERATOR | `generator` | 5006 |
|
|
170
|
+
| GENERATOR-v2 | `generator_v2` | 5007 |
|
|
171
|
+
| NT-v2 | `nt_v2` | 5008 |
|
|
172
|
+
|
|
173
|
+
Each service exposes `GET /health` and `POST /predict` (`{chrom, pos, ref, alt,
|
|
174
|
+
genome_version}`); the DNA-LM services also expose `POST /batch_predict`.
|
|
175
|
+
|
|
176
|
+
Each model gets its own conda env, service and port by default, and there is no
|
|
177
|
+
gateway to stand up — so you can deploy only the models you actually need.
|
|
178
|
+
|
|
179
|
+
Environments may also be **shared**: any environment with the right dependencies
|
|
180
|
+
(torch + `transformers<5` + pyfaidx) can serve several of these models, which
|
|
181
|
+
saves tens of GB. Start the service with that environment's interpreter, e.g.
|
|
182
|
+
`~/miniconda3/envs/NT/bin/python ~/.merge-local-servers/generator_server.py`.
|
|
183
|
+
`merge local env status` shows which environment is really behind each port.
|
|
184
|
+
|
|
185
|
+
## Precomputed VCF Cache
|
|
186
|
+
|
|
187
|
+
```bash
|
|
188
|
+
merge precomputed configure --data-dir /path/to/precomputed
|
|
189
|
+
merge precomputed status
|
|
190
|
+
merge precomputed download --genome all --variant-type all
|
|
191
|
+
```
|
|
192
|
+
|
|
193
|
+
Expected files: `coding_merged.vcf.gz`, `splicing_merged.vcf.gz`, their hg19
|
|
194
|
+
counterparts, and `.tbi` indices. The cache only covers SNVs and the older
|
|
195
|
+
feature set, so it is bypassed for non-SNVs and whenever an opt-in DNA-LM is
|
|
196
|
+
requested.
|
|
197
|
+
|
|
198
|
+
## Scoring conventions
|
|
199
|
+
|
|
200
|
+
The local services reproduce the exact quantities the ensemble was trained on.
|
|
201
|
+
Changing any of them silently invalidates the MERGE score:
|
|
202
|
+
|
|
203
|
+
| Model | Score | Window | VRAM |
|
|
204
|
+
|---|---|---|---|
|
|
205
|
+
| Evo2-7B / -7B-base | mean_PLL(REF) − mean_PLL(ALT) | 8192 bp | ~20 GB |
|
|
206
|
+
| Evo2-1B-base | mean_PLL(REF) − mean_PLL(ALT) | 8192 bp | ~8 GB, needs FP8 |
|
|
207
|
+
| Carbon-3B | logP_sum(REF) − logP_sum(ALT) | 24576 bp | **~20 GB** |
|
|
208
|
+
| GENERATOR / -v2 | NLL(ALT) − NLL(REF) | 600 bp, max_length 128 | ~6 GB |
|
|
209
|
+
| NT-v2 | ‖emb(ALT) − emb(REF)‖₂ | 600 bp, max_length 128 | ~3 GB |
|
|
210
|
+
| NT | ‖emb(ALT) − emb(REF)‖₂ | 8192 bp | ~3 GB |
|
|
211
|
+
| HyenaDNA | log p(ALT) − log p(REF), next token | 1000 bp upstream | ~2 GB |
|
|
212
|
+
|
|
213
|
+
These windows are not tunable, on purpose. The MERGE bundles were fitted on
|
|
214
|
+
features computed exactly this way, so a shorter window does not merely cost
|
|
215
|
+
accuracy — it puts the feature on a different scale than the model expects, and
|
|
216
|
+
the resulting MERGE score is wrong without anything looking wrong. If a model
|
|
217
|
+
does not fit on your GPU, leave it off: the ensemble median-imputes what is
|
|
218
|
+
missing, which is the honest fallback.
|
|
219
|
+
|
|
220
|
+
Carbon in particular needs roughly 20 GB of VRAM at the 24576 bp training window.
|
|
221
|
+
|
|
222
|
+
The `transformers` version is pinned to 4.x in the generated environments:
|
|
223
|
+
NT-v2's remote code does not load under transformers 5.x. `alphagenome` is
|
|
224
|
+
pinned to 0.5.1 to match the MERGE server, so both modes return the same
|
|
225
|
+
AlphaGenome features.
|
|
226
|
+
|
|
227
|
+
AlphaGenome is the one model whose features are **not** reproducible over time:
|
|
228
|
+
it runs on Google's servers and is updated there, so today's values differ from
|
|
229
|
+
those in the training tables. Every other model reproduces bit-for-bit.
|
|
230
|
+
|
|
231
|
+
If local services are unavailable, the errors are reported under
|
|
232
|
+
`prediction.errors` and the CLI continues with the models it could reach.
|
|
@@ -0,0 +1,199 @@
|
|
|
1
|
+
# merge-cli 3.8
|
|
2
|
+
|
|
3
|
+
MERGE variant pathogenicity prediction CLI — the professional-user counterpart to
|
|
4
|
+
[merge.fanglab.cn](https://merge.fanglab.cn), with the same models and the same
|
|
5
|
+
ensemble, runnable entirely on your own machine.
|
|
6
|
+
|
|
7
|
+
**New here? Read [TUTORIAL.md](TUTORIAL.md)** — install, both modes, per-model
|
|
8
|
+
setup, scripting and troubleshooting.
|
|
9
|
+
|
|
10
|
+
## Highlights
|
|
11
|
+
|
|
12
|
+
- **Four variant types**: coding, splice, noncoding and **non-SNV** (indel / MNV / delins).
|
|
13
|
+
- **13 models**: AlphaGenome, HyenaDNA, NT, AlphaMissense, ESM-1b, GPN-MSA,
|
|
14
|
+
Evo2-7B, Evo2-7B-base, Enformer, GENERATOR, GENERATOR-v2, NT-v2, Carbon-3B.
|
|
15
|
+
(Evo2-1B-base is also supported locally where the GPU allows it.)
|
|
16
|
+
- **Remote and local modes.** Local mode deploys each model into its own conda
|
|
17
|
+
environment and serves it over HTTP — no Docker, no gateway required.
|
|
18
|
+
- Local precomputed VCF cache lookup for hg38/hg19 coding and splicing SNVs.
|
|
19
|
+
- Bundled MERGE ensemble models (the same four deployment bundles the website runs).
|
|
20
|
+
|
|
21
|
+
## Quick Start
|
|
22
|
+
|
|
23
|
+
```bash
|
|
24
|
+
pip install merge-cli
|
|
25
|
+
merge --help
|
|
26
|
+
merge predict --chrom chr1 --pos 69428 --ref T --alt G --genome hg38
|
|
27
|
+
```
|
|
28
|
+
|
|
29
|
+
Optional extras: `merge-cli[mcp]` to ask for predictions in chat (see below),
|
|
30
|
+
`merge-cli[local]` for local mode (adds pysam, which has no Windows wheels —
|
|
31
|
+
which is why it is not in the base install), or `merge-cli[all]` for both.
|
|
32
|
+
|
|
33
|
+
## Ask in chat (MCP)
|
|
34
|
+
|
|
35
|
+
```bash
|
|
36
|
+
pip install "merge-cli[mcp]"
|
|
37
|
+
merge mcp install # detects Claude Desktop / Claude Code, writes the config
|
|
38
|
+
merge mcp status
|
|
39
|
+
```
|
|
40
|
+
|
|
41
|
+
Then restart your client and ask *"Is chr1:1040819 G>GC pathogenic?"*. The MCP
|
|
42
|
+
server runs in remote mode only — no GPU needed. Details in
|
|
43
|
+
[TUTORIAL.md §11](TUTORIAL.md).
|
|
44
|
+
|
|
45
|
+
## Variant types and ensemble models
|
|
46
|
+
|
|
47
|
+
The task is chosen automatically: a variant whose REF or ALT is not a single base
|
|
48
|
+
is scored as non-SNV; otherwise ANNOVAR decides coding / splice / noncoding.
|
|
49
|
+
Override with `--ensemble-type coding|splice|noncoding|nonsnv`.
|
|
50
|
+
|
|
51
|
+
| Task | Bundle | Algorithm | Features |
|
|
52
|
+
|---|---|---|---|
|
|
53
|
+
| coding | `MERGE_coding.pkl` | LR_L1 | 6 |
|
|
54
|
+
| splice | `MERGE_splicing.pkl` | CatBoost | 10 |
|
|
55
|
+
| noncoding | `MERGE_noncoding.pkl` | LR_L2 | 10 |
|
|
56
|
+
| non-SNV | `MERGE_nonsnv.pkl` | CatBoost | 10 |
|
|
57
|
+
|
|
58
|
+
Each bundle embeds its own preprocessing (median imputation → z-scoring →
|
|
59
|
+
classifier), so raw model scores are fed straight in and any model you did not
|
|
60
|
+
run is median-imputed. See `merge_cli/data/models/MODEL_CARD.md` for the exact
|
|
61
|
+
feature panels.
|
|
62
|
+
|
|
63
|
+
## Models
|
|
64
|
+
|
|
65
|
+
All models run by default — each one feeds at least one ensemble panel, and a
|
|
66
|
+
model you skip has its feature median-imputed instead. Skip any with `--no-<model>`:
|
|
67
|
+
|
|
68
|
+
```bash
|
|
69
|
+
merge predict --chrom chr1 --pos 69428 --ref T --alt G --no-carbon --no-nt-v2
|
|
70
|
+
```
|
|
71
|
+
|
|
72
|
+
Available switches: `--no-alphagenome --no-hyenadna --no-nt --no-alphamissense
|
|
73
|
+
--no-esm1b --no-gpn-msa --no-evo2 --no-enformer --no-generator
|
|
74
|
+
--no-generator-v2 --no-nt-v2 --no-carbon`.
|
|
75
|
+
|
|
76
|
+
In local mode every model needs its service running; the CLI reports which
|
|
77
|
+
features were computed and which were imputed on every prediction.
|
|
78
|
+
|
|
79
|
+
### Evo2 variants
|
|
80
|
+
|
|
81
|
+
`--evo2-model` picks the weights for the primary Evo2 score (remote and local):
|
|
82
|
+
|
|
83
|
+
```bash
|
|
84
|
+
merge predict --chrom chr1 --pos 69428 --ref T --alt G --evo2-model evo2_7b_base
|
|
85
|
+
```
|
|
86
|
+
|
|
87
|
+
All three Evo2 variants are scored by default in local mode —
|
|
88
|
+
`evo2_7b_score`, `evo2_7b_base_score` and `evo2_1b_base_score` are separate
|
|
89
|
+
ensemble features. They share one GPU-resident service and are scored one at a
|
|
90
|
+
time, costing a few seconds each. Narrow the set to trade coverage for speed:
|
|
91
|
+
|
|
92
|
+
```bash
|
|
93
|
+
merge local predict --chrom chr1 --pos 69428 --ref T --alt G \
|
|
94
|
+
--evo2-models evo2_7b # primary variant only, faster
|
|
95
|
+
```
|
|
96
|
+
|
|
97
|
+
`evo2_1b_base` needs a GPU with FP8 (compute capability ≥ 8.9: RTX 4000/5000 Ada,
|
|
98
|
+
L40S, H100). Elsewhere — including in remote mode — its feature is left missing and
|
|
99
|
+
median-imputed by the ensemble, which is what the website does too.
|
|
100
|
+
|
|
101
|
+
## Local mode
|
|
102
|
+
|
|
103
|
+
```bash
|
|
104
|
+
merge doctor # check GPU and services
|
|
105
|
+
merge local setup # configure paths
|
|
106
|
+
merge local env setup --model evo2 # deploy one model
|
|
107
|
+
merge local env setup --model carbon # ... or another
|
|
108
|
+
merge local env setup --model all # everything (large!)
|
|
109
|
+
merge local env start --model evo2
|
|
110
|
+
merge local env status
|
|
111
|
+
merge local predict --chrom chr17 --pos 43092919 --ref A --alt G
|
|
112
|
+
```
|
|
113
|
+
|
|
114
|
+
Setting a reference FASTA is strongly recommended — the local services use it to
|
|
115
|
+
cut the exact sequence windows the models were scored with during training:
|
|
116
|
+
|
|
117
|
+
```bash
|
|
118
|
+
merge local setup # answer the "Reference genome FASTA path" prompt
|
|
119
|
+
# or
|
|
120
|
+
merge local predict ... --genome-ref /path/to/hg38.fa
|
|
121
|
+
```
|
|
122
|
+
|
|
123
|
+
Without it the services fall back to fetching sequence context from the MERGE
|
|
124
|
+
server, which requires network access.
|
|
125
|
+
|
|
126
|
+
### Service ports
|
|
127
|
+
|
|
128
|
+
| Model | Env | Port |
|
|
129
|
+
|---|---|---|
|
|
130
|
+
| AlphaGenome | `alphagenome` | 5000 |
|
|
131
|
+
| HyenaDNA | `hyenadna` | 5001 |
|
|
132
|
+
| NT | `nt` | 5002 |
|
|
133
|
+
| Evo2 (all variants) | `evo2` | 5003 |
|
|
134
|
+
| Enformer | `enformer` | 5004 |
|
|
135
|
+
| Carbon | `carbon` | 5005 |
|
|
136
|
+
| GENERATOR | `generator` | 5006 |
|
|
137
|
+
| GENERATOR-v2 | `generator_v2` | 5007 |
|
|
138
|
+
| NT-v2 | `nt_v2` | 5008 |
|
|
139
|
+
|
|
140
|
+
Each service exposes `GET /health` and `POST /predict` (`{chrom, pos, ref, alt,
|
|
141
|
+
genome_version}`); the DNA-LM services also expose `POST /batch_predict`.
|
|
142
|
+
|
|
143
|
+
Each model gets its own conda env, service and port by default, and there is no
|
|
144
|
+
gateway to stand up — so you can deploy only the models you actually need.
|
|
145
|
+
|
|
146
|
+
Environments may also be **shared**: any environment with the right dependencies
|
|
147
|
+
(torch + `transformers<5` + pyfaidx) can serve several of these models, which
|
|
148
|
+
saves tens of GB. Start the service with that environment's interpreter, e.g.
|
|
149
|
+
`~/miniconda3/envs/NT/bin/python ~/.merge-local-servers/generator_server.py`.
|
|
150
|
+
`merge local env status` shows which environment is really behind each port.
|
|
151
|
+
|
|
152
|
+
## Precomputed VCF Cache
|
|
153
|
+
|
|
154
|
+
```bash
|
|
155
|
+
merge precomputed configure --data-dir /path/to/precomputed
|
|
156
|
+
merge precomputed status
|
|
157
|
+
merge precomputed download --genome all --variant-type all
|
|
158
|
+
```
|
|
159
|
+
|
|
160
|
+
Expected files: `coding_merged.vcf.gz`, `splicing_merged.vcf.gz`, their hg19
|
|
161
|
+
counterparts, and `.tbi` indices. The cache only covers SNVs and the older
|
|
162
|
+
feature set, so it is bypassed for non-SNVs and whenever an opt-in DNA-LM is
|
|
163
|
+
requested.
|
|
164
|
+
|
|
165
|
+
## Scoring conventions
|
|
166
|
+
|
|
167
|
+
The local services reproduce the exact quantities the ensemble was trained on.
|
|
168
|
+
Changing any of them silently invalidates the MERGE score:
|
|
169
|
+
|
|
170
|
+
| Model | Score | Window | VRAM |
|
|
171
|
+
|---|---|---|---|
|
|
172
|
+
| Evo2-7B / -7B-base | mean_PLL(REF) − mean_PLL(ALT) | 8192 bp | ~20 GB |
|
|
173
|
+
| Evo2-1B-base | mean_PLL(REF) − mean_PLL(ALT) | 8192 bp | ~8 GB, needs FP8 |
|
|
174
|
+
| Carbon-3B | logP_sum(REF) − logP_sum(ALT) | 24576 bp | **~20 GB** |
|
|
175
|
+
| GENERATOR / -v2 | NLL(ALT) − NLL(REF) | 600 bp, max_length 128 | ~6 GB |
|
|
176
|
+
| NT-v2 | ‖emb(ALT) − emb(REF)‖₂ | 600 bp, max_length 128 | ~3 GB |
|
|
177
|
+
| NT | ‖emb(ALT) − emb(REF)‖₂ | 8192 bp | ~3 GB |
|
|
178
|
+
| HyenaDNA | log p(ALT) − log p(REF), next token | 1000 bp upstream | ~2 GB |
|
|
179
|
+
|
|
180
|
+
These windows are not tunable, on purpose. The MERGE bundles were fitted on
|
|
181
|
+
features computed exactly this way, so a shorter window does not merely cost
|
|
182
|
+
accuracy — it puts the feature on a different scale than the model expects, and
|
|
183
|
+
the resulting MERGE score is wrong without anything looking wrong. If a model
|
|
184
|
+
does not fit on your GPU, leave it off: the ensemble median-imputes what is
|
|
185
|
+
missing, which is the honest fallback.
|
|
186
|
+
|
|
187
|
+
Carbon in particular needs roughly 20 GB of VRAM at the 24576 bp training window.
|
|
188
|
+
|
|
189
|
+
The `transformers` version is pinned to 4.x in the generated environments:
|
|
190
|
+
NT-v2's remote code does not load under transformers 5.x. `alphagenome` is
|
|
191
|
+
pinned to 0.5.1 to match the MERGE server, so both modes return the same
|
|
192
|
+
AlphaGenome features.
|
|
193
|
+
|
|
194
|
+
AlphaGenome is the one model whose features are **not** reproducible over time:
|
|
195
|
+
it runs on Google's servers and is updated there, so today's values differ from
|
|
196
|
+
those in the training tables. Every other model reproduces bit-for-bit.
|
|
197
|
+
|
|
198
|
+
If local services are unavailable, the errors are reported under
|
|
199
|
+
`prediction.errors` and the CLI continues with the models it could reach.
|