rfmix-reader 0.2.0__tar.gz → 0.3.0__tar.gz
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- rfmix_reader-0.3.0/PKG-INFO +498 -0
- rfmix_reader-0.3.0/README.md +453 -0
- {rfmix_reader-0.2.0 → rfmix_reader-0.3.0}/pyproject.toml +42 -18
- rfmix_reader-0.3.0/rfmix_reader/__init__.py +103 -0
- {rfmix_reader-0.2.0 → rfmix_reader-0.3.0}/rfmix_reader/_testing.py +5 -5
- rfmix_reader-0.3.0/rfmix_reader/cli/__init__.py +4 -0
- rfmix_reader-0.2.0/rfmix_reader/_cli.py → rfmix_reader-0.3.0/rfmix_reader/cli/create_binaries.py +13 -3
- rfmix_reader-0.3.0/rfmix_reader/cli/merge_phased_zarrs.py +44 -0
- rfmix_reader-0.3.0/rfmix_reader/cli/prepare_reference.py +71 -0
- rfmix_reader-0.3.0/rfmix_reader/io/__init__.py +14 -0
- rfmix_reader-0.2.0/rfmix_reader/_loci_bed.py → rfmix_reader-0.3.0/rfmix_reader/io/loci_bed.py +7 -11
- rfmix_reader-0.3.0/rfmix_reader/io/prepare_reference.py +107 -0
- rfmix_reader-0.2.0/rfmix_reader/_write_data.py → rfmix_reader-0.3.0/rfmix_reader/io/write_data.py +83 -46
- rfmix_reader-0.3.0/rfmix_reader/processing/__init__.py +20 -0
- rfmix_reader-0.3.0/rfmix_reader/processing/_hmm_lai.py +333 -0
- rfmix_reader-0.3.0/rfmix_reader/processing/imputation.py +335 -0
- rfmix_reader-0.3.0/rfmix_reader/processing/phase.py +1249 -0
- rfmix_reader-0.3.0/rfmix_reader/readers/__init__.py +8 -0
- rfmix_reader-0.2.0/rfmix_reader/_fb_read.py → rfmix_reader-0.3.0/rfmix_reader/readers/fb_read.py +2 -3
- rfmix_reader-0.3.0/rfmix_reader/readers/read_flare.py +458 -0
- rfmix_reader-0.2.0/rfmix_reader/_read_rfmix.py → rfmix_reader-0.3.0/rfmix_reader/readers/read_rfmix.py +156 -142
- rfmix_reader-0.3.0/rfmix_reader/readers/read_simu.py +423 -0
- rfmix_reader-0.2.0/rfmix_reader/_utils.py → rfmix_reader-0.3.0/rfmix_reader/utils.py +197 -74
- rfmix_reader-0.3.0/rfmix_reader/viz/__init__.py +17 -0
- rfmix_reader-0.2.0/rfmix_reader/_tagore.py → rfmix_reader-0.3.0/rfmix_reader/viz/tagore.py +1 -5
- rfmix_reader-0.2.0/rfmix_reader/_visualization.py → rfmix_reader-0.3.0/rfmix_reader/viz/visualization.py +19 -26
- rfmix_reader-0.2.0/PKG-INFO +0 -278
- rfmix_reader-0.2.0/README.md +0 -247
- rfmix_reader-0.2.0/rfmix_reader/.local.py +0 -648
- rfmix_reader-0.2.0/rfmix_reader/__init__.py +0 -53
- rfmix_reader-0.2.0/rfmix_reader/_imputation.py +0 -207
- {rfmix_reader-0.2.0 → rfmix_reader-0.3.0}/LICENSE +0 -0
- {rfmix_reader-0.2.0 → rfmix_reader-0.3.0}/rfmix_reader/base.svg.p +0 -0
- /rfmix_reader-0.2.0/rfmix_reader/_chunk.py → /rfmix_reader-0.3.0/rfmix_reader/io/chunk.py +0 -0
- /rfmix_reader-0.2.0/rfmix_reader/_errorhandling.py → /rfmix_reader-0.3.0/rfmix_reader/io/errors.py +0 -0
- /rfmix_reader-0.2.0/rfmix_reader/_constants.py → /rfmix_reader-0.3.0/rfmix_reader/processing/constants.py +0 -0
|
@@ -0,0 +1,498 @@
|
|
|
1
|
+
Metadata-Version: 2.4
|
|
2
|
+
Name: rfmix-reader
|
|
3
|
+
Version: 0.3.0
|
|
4
|
+
Summary: RFMix-reader is a Python package designed to efficiently read and process output files generated by RFMix, a popular tool for estimating local ancestry in admixed populations. The package employs a lazy loading approach, which minimizes memory consumption by reading only the loci that are accessed by the user, rather than loading the entire dataset into memory at once.
|
|
5
|
+
License: GPL-3.0-or-later
|
|
6
|
+
License-File: LICENSE
|
|
7
|
+
Keywords: file parser,rfmix,gpu acceleration,local ancestry
|
|
8
|
+
Author: Kynon J.M. Benjamin
|
|
9
|
+
Author-email: kj.benjamin90@gmail.com
|
|
10
|
+
Maintainer: Kynon J.M. Benjamin
|
|
11
|
+
Maintainer-email: kj.benjamin90@gmail.com
|
|
12
|
+
Requires-Python: >=3.11,<3.15
|
|
13
|
+
Classifier: License :: OSI Approved :: GNU General Public License v3 or later (GPLv3+)
|
|
14
|
+
Classifier: Operating System :: OS Independent
|
|
15
|
+
Classifier: Programming Language :: Python :: 3
|
|
16
|
+
Classifier: Programming Language :: Python :: 3.11
|
|
17
|
+
Classifier: Programming Language :: Python :: 3.12
|
|
18
|
+
Classifier: Programming Language :: Python :: 3.13
|
|
19
|
+
Classifier: Programming Language :: Python :: 3.14
|
|
20
|
+
Provides-Extra: docs
|
|
21
|
+
Provides-Extra: gpu
|
|
22
|
+
Provides-Extra: io
|
|
23
|
+
Provides-Extra: tests
|
|
24
|
+
Provides-Extra: viz
|
|
25
|
+
Requires-Dist: cairosvg (>=2.7,<3.0) ; extra == "viz"
|
|
26
|
+
Requires-Dist: cyvcf2 (>=0.31)
|
|
27
|
+
Requires-Dist: dask (>=2025.1,<2026.0)
|
|
28
|
+
Requires-Dist: fonttools (>=4.61.0,<5.0.0)
|
|
29
|
+
Requires-Dist: matplotlib (>=3.10,<4.0) ; extra == "viz"
|
|
30
|
+
Requires-Dist: numpy (>=1.23,<3)
|
|
31
|
+
Requires-Dist: pandas (>=2.0)
|
|
32
|
+
Requires-Dist: psutil (>=7,<8)
|
|
33
|
+
Requires-Dist: seaborn (>=0.13,<0.14) ; extra == "viz"
|
|
34
|
+
Requires-Dist: sphinx (>=7,<9) ; extra == "docs"
|
|
35
|
+
Requires-Dist: sphinx-autodoc-typehints (>=2,<3) ; extra == "docs"
|
|
36
|
+
Requires-Dist: sphinx-copybutton (>=0.5,<0.6) ; extra == "docs"
|
|
37
|
+
Requires-Dist: sphinx-rtd-theme (>=2,<3) ; extra == "docs"
|
|
38
|
+
Requires-Dist: torch (>=2.8) ; extra == "gpu"
|
|
39
|
+
Requires-Dist: tqdm (>=4.66)
|
|
40
|
+
Project-URL: Bug Tracker, https://github.com/heart-gen/rfmix_reader/issues
|
|
41
|
+
Project-URL: homepage, https://rfmix-reader.readthedocs.io/en/latest/
|
|
42
|
+
Project-URL: repository, https://github.com/heart-gen/rfmix_reader.git
|
|
43
|
+
Description-Content-Type: text/markdown
|
|
44
|
+
|
|
45
|
+
# RFMix-reader
|
|
46
|
+
`RFMix-reader` is a Python package for efficiently reading and processing output
|
|
47
|
+
files generated by [`RFMix`](https://github.com/slowkoni/rfmix), a widely used tool
|
|
48
|
+
for estimating local ancestry in admixed populations.
|
|
49
|
+
It employs a **lazy loading approach** to minimize memory usage, and leverages **GPU acceleration**
|
|
50
|
+
for major speedups when available.
|
|
51
|
+
|
|
52
|
+
---
|
|
53
|
+
|
|
54
|
+
## Installation
|
|
55
|
+
|
|
56
|
+
`rfmix-reader` requires **Python 3.10+**. Install from PyPI:
|
|
57
|
+
|
|
58
|
+
```bash
|
|
59
|
+
pip install rfmix-reader
|
|
60
|
+
````
|
|
61
|
+
|
|
62
|
+
### Installation Options
|
|
63
|
+
|
|
64
|
+
* **Basic install** (CPU only):
|
|
65
|
+
|
|
66
|
+
```bash
|
|
67
|
+
pip install rfmix-reader
|
|
68
|
+
```
|
|
69
|
+
|
|
70
|
+
* **With GPU acceleration** (`cupy`, `cudf`, `dask-cudf`):
|
|
71
|
+
|
|
72
|
+
```bash
|
|
73
|
+
pip install rfmix-reader[gpu]
|
|
74
|
+
```
|
|
75
|
+
|
|
76
|
+
* **With documentation tools** (`sphinx`, `sphinx-rtd-theme`):
|
|
77
|
+
|
|
78
|
+
```bash
|
|
79
|
+
pip install rfmix-reader[docs]
|
|
80
|
+
```
|
|
81
|
+
|
|
82
|
+
* **With testing tools** (`pytest`):
|
|
83
|
+
|
|
84
|
+
```bash
|
|
85
|
+
pip install rfmix-reader[tests]
|
|
86
|
+
```
|
|
87
|
+
|
|
88
|
+
### GPU Notes
|
|
89
|
+
|
|
90
|
+
* `torch` is installed automatically.
|
|
91
|
+
* For CUDA builds, install a matching GPU-enabled wheel for your system following the [PyTorch guide](https://pytorch.org/get-started/locally/).
|
|
92
|
+
* RAPIDS (`cudf`, `cupy`) wheels are version- and CUDA-specific. See the [RAPIDS install guide](https://docs.rapids.ai/install).
|
|
93
|
+
* CPU-only installations will still run efficiently, just without GPU acceleration.
|
|
94
|
+
|
|
95
|
+
---
|
|
96
|
+
|
|
97
|
+
## Quickstart
|
|
98
|
+
|
|
99
|
+
```python
|
|
100
|
+
from rfmix_reader import read_rfmix
|
|
101
|
+
|
|
102
|
+
# Load RFMix outputs (two-population admixture example)
|
|
103
|
+
file_path = "examples/two_populations/out/"
|
|
104
|
+
loci_df, g_anc, local_array = read_rfmix(file_path)
|
|
105
|
+
|
|
106
|
+
print(loci_df.head())
|
|
107
|
+
print(g_anc.head())
|
|
108
|
+
print(local_array.shape)
|
|
109
|
+
"""See the phasing section below for how to phase per-chromosome outputs and
|
|
110
|
+
write them to Zarr with `phase_rfmix_chromosome_to_zarr`."""
|
|
111
|
+
```
|
|
112
|
+
|
|
113
|
+
---
|
|
114
|
+
|
|
115
|
+
## Key Features
|
|
116
|
+
|
|
117
|
+
* **Lazy Loading**: Reads data on-the-fly, reducing memory footprint.
|
|
118
|
+
* **Efficient Access**: Query specific loci or regions of interest.
|
|
119
|
+
* **Seamless Integration**: Works smoothly with `pandas`, `dask`, and other analysis tools.
|
|
120
|
+
* **Loci Imputation**: Impute local ancestry loci to dense genotype variant sites.
|
|
121
|
+
* **GPU Acceleration**: Automatic CUDA acceleration via PyTorch/CuPy when available.
|
|
122
|
+
|
|
123
|
+
---
|
|
124
|
+
|
|
125
|
+
## Simulation Data
|
|
126
|
+
|
|
127
|
+
Test datasets for two- and three-population admixture are available on Synapse:
|
|
128
|
+
[Synapse Project syn61691659](https://www.synapse.org/Synapse:syn61691659).
|
|
129
|
+
|
|
130
|
+
---
|
|
131
|
+
|
|
132
|
+
## Usage
|
|
133
|
+
|
|
134
|
+
### Binary Conversion
|
|
135
|
+
|
|
136
|
+
RFMix does not generate binary files directly.
|
|
137
|
+
Use `create_binaries` to generate them (also available as a CLI):
|
|
138
|
+
|
|
139
|
+
```bash
|
|
140
|
+
create-binaries two_pops/out/
|
|
141
|
+
```
|
|
142
|
+
|
|
143
|
+
```python
|
|
144
|
+
from rfmix_reader import create_binaries
|
|
145
|
+
|
|
146
|
+
create_binaries("two_pops/out/", binary_dir="./binary_files")
|
|
147
|
+
```
|
|
148
|
+
|
|
149
|
+
### Preparing Reference Data for Phasing
|
|
150
|
+
|
|
151
|
+
Use `prepare-reference` to convert bgzipped, indexed reference VCF/BCF files
|
|
152
|
+
into per-chromosome VCF-Zarr stores that the phasing pipeline consumes.
|
|
153
|
+
The command writes one `<chrom>.zarr` directory per input file.
|
|
154
|
+
|
|
155
|
+
```
|
|
156
|
+
prepare-reference -h
|
|
157
|
+
```
|
|
158
|
+
|
|
159
|
+
```
|
|
160
|
+
usage: prepare-reference [-h] [--chunk-length CHUNK_LENGTH]
|
|
161
|
+
[--samples-chunk-size SAMPLES_CHUNK_SIZE]
|
|
162
|
+
[--worker-processes WORKER_PROCESSES]
|
|
163
|
+
[--verbose | --no-verbose] [--version]
|
|
164
|
+
output_dir vcf_paths [vcf_paths ...]
|
|
165
|
+
|
|
166
|
+
Convert one or more bgzipped reference VCF/BCF files into Zarr stores.
|
|
167
|
+
|
|
168
|
+
positional arguments:
|
|
169
|
+
output_dir Directory where the Zarr outputs will be written.
|
|
170
|
+
vcf_paths Paths to reference VCF/BCF files (bgzipped and
|
|
171
|
+
indexed).
|
|
172
|
+
|
|
173
|
+
options:
|
|
174
|
+
-h, --help show this help message and exit
|
|
175
|
+
--chunk-length CHUNK_LENGTH
|
|
176
|
+
Genomic chunk size for the output Zarr stores
|
|
177
|
+
(default: 100000).
|
|
178
|
+
--samples-chunk-size SAMPLES_CHUNK_SIZE
|
|
179
|
+
Chunk size for samples in the output Zarr stores
|
|
180
|
+
(default: library chosen).
|
|
181
|
+
--worker-processes WORKER_PROCESSES
|
|
182
|
+
Number of worker processes to use for conversion
|
|
183
|
+
(default: 0, use library default).
|
|
184
|
+
--verbose, --no-verbose
|
|
185
|
+
Print progress messages (default: enabled).
|
|
186
|
+
--version Show the version of the program and exit.
|
|
187
|
+
```
|
|
188
|
+
|
|
189
|
+
Example data preparation:
|
|
190
|
+
|
|
191
|
+
```bash
|
|
192
|
+
# Sample annotations: two columns (no header): sample_id<TAB>group
|
|
193
|
+
cat > sample_annotations.tsv <<'EOF'
|
|
194
|
+
NA19700 AFR
|
|
195
|
+
NA19701 AFR
|
|
196
|
+
NA20847 EUR
|
|
197
|
+
EOF
|
|
198
|
+
|
|
199
|
+
# Convert chromosome VCFs into a reference store directory
|
|
200
|
+
prepare-reference refs/ 1kg_chr20.vcf.gz 1kg_chr21.vcf.gz \
|
|
201
|
+
--chunk-length 50000 --samples-chunk-size 512
|
|
202
|
+
```
|
|
203
|
+
|
|
204
|
+
### Main Function
|
|
205
|
+
|
|
206
|
+
Once binaries are available, process RFMix results:
|
|
207
|
+
|
|
208
|
+
```python
|
|
209
|
+
from rfmix_reader import read_rfmix
|
|
210
|
+
|
|
211
|
+
loci, g_anc, admix = read_rfmix("two_pops/out/")
|
|
212
|
+
```
|
|
213
|
+
|
|
214
|
+
### Three Population Example
|
|
215
|
+
|
|
216
|
+
Binaries can also be generated on-the-fly within `read_rfmix` with
|
|
217
|
+
`generate_binary` set to `True`.
|
|
218
|
+
|
|
219
|
+
```python
|
|
220
|
+
loci, g_anc, admix = read_rfmix("examples/three_populations/out/",
|
|
221
|
+
binary_dir="./binary_files",
|
|
222
|
+
generate_binary=True)
|
|
223
|
+
```
|
|
224
|
+
|
|
225
|
+
### Phasing data
|
|
226
|
+
|
|
227
|
+
For optimal memory and computational speed, phasing is done per
|
|
228
|
+
chromosome.
|
|
229
|
+
|
|
230
|
+
```python
|
|
231
|
+
from rfmix_reader import phase_rfmix_chromosome_to_zarr
|
|
232
|
+
|
|
233
|
+
# Use the reference store + annotations during phasing
|
|
234
|
+
admix = phase_rfmix_chromosome_to_zarr(
|
|
235
|
+
file_prefix="two_pops/out/",
|
|
236
|
+
ref_zarr_root="refs",
|
|
237
|
+
sample_annot_path="sample_annotations.tsv",
|
|
238
|
+
output_path="./phased_chr21.zarr",
|
|
239
|
+
chrom="21",
|
|
240
|
+
)
|
|
241
|
+
```
|
|
242
|
+
|
|
243
|
+
The chunking is suboptimal for phasing, so remember to
|
|
244
|
+
rechunk before using for optimal processing.
|
|
245
|
+
|
|
246
|
+
```python
|
|
247
|
+
local_array = admix["local_ancestry"].chunk({"variant": 20000, "sample": 50})
|
|
248
|
+
# Compute into memory, if needed
|
|
249
|
+
local_array = local_array.compute()
|
|
250
|
+
```
|
|
251
|
+
|
|
252
|
+
This also saves the data to Zarr for later merging or data processing.
|
|
253
|
+
|
|
254
|
+
```bash
|
|
255
|
+
merge-phased-zarrs ./phased_all.zarr ./phased_chr21.zarr ./phased_chr22.zarr
|
|
256
|
+
```
|
|
257
|
+
### Loci Imputation
|
|
258
|
+
|
|
259
|
+
The imputation workflow now lives in ``rfmix_reader.processing.imputation`` and
|
|
260
|
+
is exported as ``interpolate_array``. It interpolates the local ancestry matrix
|
|
261
|
+
onto a denser variant grid and writes the result to ``<zarr_outdir>/local-ancestry.zarr``
|
|
262
|
+
as a Zarr array shaped ``(variants, samples, ancestries)``.
|
|
263
|
+
|
|
264
|
+
**Inputs**
|
|
265
|
+
|
|
266
|
+
* ``variant_loci_df``: a pandas DataFrame defining the variant grid. Provide at
|
|
267
|
+
least ``chrom``/``pos`` and an ``i`` column that points to the source RFMix
|
|
268
|
+
row index; rows with ``i`` set to ``NaN`` are treated as missing loci to
|
|
269
|
+
interpolate. Sort the frame by genomic coordinate, and include ``pos`` if you
|
|
270
|
+
plan to interpolate in base-pair space.
|
|
271
|
+
* ``admix``: the local ancestry Dask array returned by ``read_rfmix`` (shape
|
|
272
|
+
``(loci, samples, ancestries)``).
|
|
273
|
+
* ``zarr_outdir``: an output directory where the new ``local-ancestry.zarr``
|
|
274
|
+
store will be created.
|
|
275
|
+
|
|
276
|
+
**Key options**
|
|
277
|
+
|
|
278
|
+
* ``interpolation``: ``"linear"`` (default), ``"nearest"``, or ``"stepwise"``.
|
|
279
|
+
* ``use_bp_positions``: set to ``True`` to interpolate along ``variant_loci_df['pos']``
|
|
280
|
+
rather than treating loci as equally spaced indices.
|
|
281
|
+
* ``chunk_size``/``batch_size``: tune how many rows are materialized at a time
|
|
282
|
+
when filling and interpolating the Zarr array.
|
|
283
|
+
|
|
284
|
+
**Workflow example**
|
|
285
|
+
|
|
286
|
+
```python
|
|
287
|
+
import pandas as pd
|
|
288
|
+
from pathlib import Path
|
|
289
|
+
from rfmix_reader import interpolate_array, read_rfmix
|
|
290
|
+
|
|
291
|
+
# Load RFMix loci and local ancestry
|
|
292
|
+
loci_df, _, admix = read_rfmix("two_pops/out/", binary_dir="./binary_files")
|
|
293
|
+
|
|
294
|
+
# Build the variant grid by merging genotype sites with the RFMix loci index
|
|
295
|
+
variants = pd.read_parquet("genotypes/variants.parquet") # must include chrom/pos
|
|
296
|
+
variants = variants.drop_duplicates(subset=["chrom", "pos"]).sort_values("pos")
|
|
297
|
+
variant_loci_df = (
|
|
298
|
+
variants.merge(loci_df.to_pandas(), on=["chrom", "pos"], how="outer", indicator=True)
|
|
299
|
+
.loc[:, ["chrom", "pos", "i", "_merge"]]
|
|
300
|
+
)
|
|
301
|
+
|
|
302
|
+
z = interpolate_array(
|
|
303
|
+
variant_loci_df,
|
|
304
|
+
admix,
|
|
305
|
+
zarr_outdir=Path("./imputed_local_ancestry"),
|
|
306
|
+
interpolation="linear",
|
|
307
|
+
use_bp_positions=True,
|
|
308
|
+
chunk_size=50_000,
|
|
309
|
+
)
|
|
310
|
+
print(z)
|
|
311
|
+
```
|
|
312
|
+
|
|
313
|
+
The interpolator uses GPU acceleration transparently when ``cupy`` and a CUDA
|
|
314
|
+
PyTorch build are available; otherwise it falls back to NumPy. All methods other
|
|
315
|
+
than ``"hmm"`` operate on diploid-summed trajectories and preserve the original
|
|
316
|
+
ancestry dimension.
|
|
317
|
+
|
|
318
|
+
### Phasing workflow (`rfmix_reader.processing.phase`)
|
|
319
|
+
|
|
320
|
+
`rfmix_reader.processing.phase` implements gnomix-style tail-flip corrections
|
|
321
|
+
for local ancestry haplotypes. Phasing now lives outside `read_rfmix` so you can
|
|
322
|
+
process each chromosome independently and write outputs directly to Zarr.
|
|
323
|
+
|
|
324
|
+
#### Required reference inputs
|
|
325
|
+
|
|
326
|
+
* **VCF-Zarr reference** (`ref_zarr_root`): either a single `*.zarr` store or a
|
|
327
|
+
directory containing per-chromosome stores (e.g., `1.zarr`, `chr1.zarr`).
|
|
328
|
+
* **Sample annotations** (`sample_annot_path`): two-column file mapping
|
|
329
|
+
`sample_id` to `group` (ancestry label). One representative sample per group
|
|
330
|
+
is pulled to build reference haplotypes.
|
|
331
|
+
|
|
332
|
+
#### Per-chromosome pipeline
|
|
333
|
+
|
|
334
|
+
1. (Optional) generate RFMix binary caches with `create_binaries`.
|
|
335
|
+
2. Call `phase_rfmix_chromosome_to_zarr` for each chromosome you want to
|
|
336
|
+
process.
|
|
337
|
+
3. Optionally concatenate those per-chromosome Zarr stores with
|
|
338
|
+
`merge_phased_zarrs`.
|
|
339
|
+
|
|
340
|
+
```python
|
|
341
|
+
from rfmix_reader.processing.phase import (
|
|
342
|
+
PhasingConfig,
|
|
343
|
+
merge_phased_zarrs,
|
|
344
|
+
phase_rfmix_chromosome_to_zarr,
|
|
345
|
+
)
|
|
346
|
+
|
|
347
|
+
# Phase chromosome 21 and write to Zarr
|
|
348
|
+
dataset = phase_rfmix_chromosome_to_zarr(
|
|
349
|
+
file_prefix="examples/two_populations/out/",
|
|
350
|
+
ref_zarr_root="/refs/1kg_chr_zarr/",
|
|
351
|
+
sample_annot_path="/refs/1kg_annotations.tsv",
|
|
352
|
+
output_path="/tmp/phased_chr21.zarr",
|
|
353
|
+
chrom="21",
|
|
354
|
+
)
|
|
355
|
+
|
|
356
|
+
# Merge multiple per-chromosome Zarr stores
|
|
357
|
+
merged = merge_phased_zarrs(
|
|
358
|
+
["/tmp/phased_chr21.zarr", "/tmp/phased_chr22.zarr"],
|
|
359
|
+
output_path="/tmp/phased_all.zarr",
|
|
360
|
+
)
|
|
361
|
+
```
|
|
362
|
+
|
|
363
|
+
If you need fine-grained control, you can still start from unphased outputs and
|
|
364
|
+
call `phase_admix_dask_with_index` directly:
|
|
365
|
+
|
|
366
|
+
```python
|
|
367
|
+
from rfmix_reader import read_rfmix
|
|
368
|
+
from rfmix_reader.processing.phase import PhasingConfig, phase_admix_dask_with_index
|
|
369
|
+
|
|
370
|
+
loci_df, g_anc, admix, X_raw = read_rfmix(
|
|
371
|
+
"examples/two_populations/out/",
|
|
372
|
+
return_original=True,
|
|
373
|
+
chrom="21",
|
|
374
|
+
)
|
|
375
|
+
|
|
376
|
+
config = PhasingConfig(window_size=100, min_block_len=10, max_mismatch_frac=0.3)
|
|
377
|
+
phased = phase_admix_dask_with_index(
|
|
378
|
+
admix=admix,
|
|
379
|
+
X_raw=X_raw,
|
|
380
|
+
positions=loci_df.physical_position.to_numpy(),
|
|
381
|
+
chrom=str(loci_df.chromosome.iloc[0]),
|
|
382
|
+
ref_zarr_root="/refs/1kg_chr_zarr/",
|
|
383
|
+
sample_annot_path="/refs/1kg_annotations.tsv",
|
|
384
|
+
config=config,
|
|
385
|
+
)
|
|
386
|
+
```
|
|
387
|
+
|
|
388
|
+
### Reading Haptools simulations
|
|
389
|
+
|
|
390
|
+
Use `read_simu` to load BGZF-compressed VCF files created by
|
|
391
|
+
`haptools simgenotype --pop_field`:
|
|
392
|
+
|
|
393
|
+
```python
|
|
394
|
+
from rfmix_reader import read_simu
|
|
395
|
+
|
|
396
|
+
loci_df, g_anc, admix = read_simu("/path/to/simulations/")
|
|
397
|
+
```
|
|
398
|
+
|
|
399
|
+
Haptools does **not** include the chromosome length in the `##contig`
|
|
400
|
+
header lines, but `read_simu` requires that metadata to index each VCF.
|
|
401
|
+
Copy the `contigs.txt` file Haptools generates from the FASTA you used
|
|
402
|
+
for simulation and reheader every file with the appropriate contig entry
|
|
403
|
+
before calling `read_simu`. The following snippet shows one approach
|
|
404
|
+
using `bcftools` and `tabix`:
|
|
405
|
+
|
|
406
|
+
```bash
|
|
407
|
+
CONTIGS="../../three_populations/_m/contigs.txt"
|
|
408
|
+
VCFDIR="gt-files"
|
|
409
|
+
CHR="chr${SLURM_ARRAY_TASK_ID}"
|
|
410
|
+
OUT="${VCFDIR}/${CHR}.vcf.gz"
|
|
411
|
+
IN="${VCFDIR}/back/${CHR}.vcf.gz"
|
|
412
|
+
|
|
413
|
+
CONTIG_LINE=$(grep -w "ID=${CHR}" "$CONTIGS")
|
|
414
|
+
if [[ -z "$CONTIG_LINE" ]]; then
|
|
415
|
+
echo "ERROR: No contig line found for ${CHR} in $CONTIGS"
|
|
416
|
+
exit 1
|
|
417
|
+
fi
|
|
418
|
+
|
|
419
|
+
bcftools view -h "$IN" \
|
|
420
|
+
| sed "s/^##contig=<ID=${CHR}>.*/${CONTIG_LINE}/" > header.${CHR}.tmp
|
|
421
|
+
bcftools reheader -h header.${CHR}.tmp -o "$OUT" "$IN"
|
|
422
|
+
tabix -p vcf "$OUT"
|
|
423
|
+
```
|
|
424
|
+
|
|
425
|
+
### Visualization
|
|
426
|
+
|
|
427
|
+
`read_rfmix`, `read_flare`, and `read_simu` all return the same
|
|
428
|
+
`(loci_df, g_anc, admix)` tuple, so the plotting utilities in
|
|
429
|
+
`rfmix_reader._visualization` work identically for RFMix, FLARE, and
|
|
430
|
+
Haptools-simulated inputs. The snippet below shows the typical workflow
|
|
431
|
+
for each reader:
|
|
432
|
+
|
|
433
|
+
```python
|
|
434
|
+
from rfmix_reader import (
|
|
435
|
+
plot_ancestry_by_chromosome,
|
|
436
|
+
plot_global_ancestry,
|
|
437
|
+
read_flare,
|
|
438
|
+
read_rfmix,
|
|
439
|
+
read_simu,
|
|
440
|
+
)
|
|
441
|
+
|
|
442
|
+
# RFMix run directory
|
|
443
|
+
loci_df, g_anc, admix = read_rfmix("two_pops/out/")
|
|
444
|
+
plot_global_ancestry(g_anc, save_path="rfmix_global.png")
|
|
445
|
+
plot_ancestry_by_chromosome(loci_df, admix, save_path="rfmix_local.png")
|
|
446
|
+
|
|
447
|
+
# FLARE output directory (contains *.anc.vcf.gz + global.anc.gz)
|
|
448
|
+
loci_df, g_anc, admix = read_flare("flare_runs/chr1/")
|
|
449
|
+
plot_global_ancestry(g_anc, save_path="flare_global.png")
|
|
450
|
+
plot_ancestry_by_chromosome(loci_df, admix, save_path="flare_local.png")
|
|
451
|
+
|
|
452
|
+
# Haptools simulations (after reheadering contigs)
|
|
453
|
+
loci_df, g_anc, admix = read_simu("/path/to/simulations/")
|
|
454
|
+
plot_global_ancestry(g_anc, save_path="simu_global.png")
|
|
455
|
+
plot_ancestry_by_chromosome(loci_df, admix, save_path="simu_local.png")
|
|
456
|
+
```
|
|
457
|
+
|
|
458
|
+
`plot_global_ancestry` builds per-individual stacked bars of global
|
|
459
|
+
ancestry while `plot_ancestry_by_chromosome` summarizes local ancestry
|
|
460
|
+
along each chromosome, giving you quick visual QC for every supported
|
|
461
|
+
input format.
|
|
462
|
+
|
|
463
|
+
---
|
|
464
|
+
|
|
465
|
+
## Development Install
|
|
466
|
+
|
|
467
|
+
For contributors:
|
|
468
|
+
|
|
469
|
+
```bash
|
|
470
|
+
git clone https://github.com/heart-gen/rfmix_reader.git
|
|
471
|
+
cd rfmix_reader
|
|
472
|
+
pip install -e ".[gpu,docs,tests]"
|
|
473
|
+
```
|
|
474
|
+
|
|
475
|
+
---
|
|
476
|
+
|
|
477
|
+
## Citation
|
|
478
|
+
|
|
479
|
+
If you use this software, please cite:
|
|
480
|
+
|
|
481
|
+
[](https://zenodo.org/doi/10.5281/zenodo.12629787)
|
|
482
|
+
|
|
483
|
+
Benjamin, K. J. M. (2024). **RFMix-reader (Version 0.2.0)** \[Computer software].
|
|
484
|
+
[https://github.com/heart-gen/rfmix\_reader](https://github.com/heart-gen/rfmix_reader)
|
|
485
|
+
|
|
486
|
+
Kynon JM Benjamin. *"RFMix-reader: Accelerated reading and processing for local ancestry studies."*
|
|
487
|
+
**bioRxiv** (2024).
|
|
488
|
+
DOI: [10.1101/2024.07.13.603370](https://www.biorxiv.org/content/10.1101/2024.07.13.603370v2).
|
|
489
|
+
|
|
490
|
+
---
|
|
491
|
+
|
|
492
|
+
## Funding
|
|
493
|
+
|
|
494
|
+
This work was supported by the National Institutes of Health,
|
|
495
|
+
National Institute on Minority Health and Health Disparities (NIMHD)
|
|
496
|
+
K99MD016964 / R00MD016964.
|
|
497
|
+
|
|
498
|
+
|