afquery 0.3.1__tar.gz → 0.3.2__tar.gz

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (163) hide show
  1. afquery-0.3.2/PKG-INFO +256 -0
  2. afquery-0.3.2/README.md +226 -0
  3. {afquery-0.3.1 → afquery-0.3.2}/src/afquery/_version.py +2 -2
  4. afquery-0.3.1/PKG-INFO +0 -145
  5. afquery-0.3.1/README.md +0 -115
  6. {afquery-0.3.1 → afquery-0.3.2}/.dockerignore +0 -0
  7. {afquery-0.3.1 → afquery-0.3.2}/.github/workflows/ci.yml +0 -0
  8. {afquery-0.3.1 → afquery-0.3.2}/.github/workflows/docs.yml +0 -0
  9. {afquery-0.3.1 → afquery-0.3.2}/.github/workflows/release.yml +0 -0
  10. {afquery-0.3.1 → afquery-0.3.2}/.gitignore +0 -0
  11. {afquery-0.3.1 → afquery-0.3.2}/CONTRIBUTING.md +0 -0
  12. {afquery-0.3.1 → afquery-0.3.2}/Dockerfile +0 -0
  13. {afquery-0.3.1 → afquery-0.3.2}/LICENSE +0 -0
  14. {afquery-0.3.1 → afquery-0.3.2}/benchmarks/.gitignore +0 -0
  15. {afquery-0.3.1 → afquery-0.3.2}/benchmarks/INSTALL.md +0 -0
  16. {afquery-0.3.1 → afquery-0.3.2}/benchmarks/README.md +0 -0
  17. {afquery-0.3.1 → afquery-0.3.2}/benchmarks/Snakefile +0 -0
  18. {afquery-0.3.1 → afquery-0.3.2}/benchmarks/capture_kit/.gitignore +0 -0
  19. {afquery-0.3.1 → afquery-0.3.2}/benchmarks/capture_kit/01a_assign_samples.py +0 -0
  20. {afquery-0.3.1 → afquery-0.3.2}/benchmarks/capture_kit/01b_write_manifests.py +0 -0
  21. {afquery-0.3.1 → afquery-0.3.2}/benchmarks/capture_kit/02_build_databases.py +0 -0
  22. {afquery-0.3.1 → afquery-0.3.2}/benchmarks/capture_kit/03_compute_metrics.py +0 -0
  23. {afquery-0.3.1 → afquery-0.3.2}/benchmarks/capture_kit/04_classify_acmg.py +0 -0
  24. {afquery-0.3.1 → afquery-0.3.2}/benchmarks/capture_kit/05_plot_figures.py +0 -0
  25. {afquery-0.3.1 → afquery-0.3.2}/benchmarks/capture_kit/README.md +0 -0
  26. {afquery-0.3.1 → afquery-0.3.2}/benchmarks/capture_kit/Snakefile +0 -0
  27. {afquery-0.3.1 → afquery-0.3.2}/benchmarks/capture_kit/__init__.py +0 -0
  28. {afquery-0.3.1 → afquery-0.3.2}/benchmarks/capture_kit/beds/SureSelect_v5.bed +0 -0
  29. {afquery-0.3.1 → afquery-0.3.2}/benchmarks/capture_kit/beds/SureSelect_v6.bed +0 -0
  30. {afquery-0.3.1 → afquery-0.3.2}/benchmarks/capture_kit/beds/SureSelect_v7.bed +0 -0
  31. {afquery-0.3.1 → afquery-0.3.2}/benchmarks/capture_kit/config.py +0 -0
  32. {afquery-0.3.1 → afquery-0.3.2}/benchmarks/config.yaml +0 -0
  33. {afquery-0.3.1 → afquery-0.3.2}/benchmarks/envs/benchmark.yaml +0 -0
  34. {afquery-0.3.1 → afquery-0.3.2}/benchmarks/performance/.gitignore +0 -0
  35. {afquery-0.3.1 → afquery-0.3.2}/benchmarks/performance/01_prepare_data.py +0 -0
  36. {afquery-0.3.1 → afquery-0.3.2}/benchmarks/performance/02_query_scaling.py +0 -0
  37. {afquery-0.3.1 → afquery-0.3.2}/benchmarks/performance/03_build.py +0 -0
  38. {afquery-0.3.1 → afquery-0.3.2}/benchmarks/performance/04_annotate.py +0 -0
  39. {afquery-0.3.1 → afquery-0.3.2}/benchmarks/performance/05_vs_bcftools.py +0 -0
  40. {afquery-0.3.1 → afquery-0.3.2}/benchmarks/performance/06_plot.py +0 -0
  41. {afquery-0.3.1 → afquery-0.3.2}/benchmarks/performance/README.md +0 -0
  42. {afquery-0.3.1 → afquery-0.3.2}/benchmarks/performance/Snakefile +0 -0
  43. {afquery-0.3.1 → afquery-0.3.2}/benchmarks/performance/collect_annotate.py +0 -0
  44. {afquery-0.3.1 → afquery-0.3.2}/benchmarks/performance/collect_bcftools.py +0 -0
  45. {afquery-0.3.1 → afquery-0.3.2}/benchmarks/performance/collect_build_perf.py +0 -0
  46. {afquery-0.3.1 → afquery-0.3.2}/benchmarks/performance/collect_prepare.py +0 -0
  47. {afquery-0.3.1 → afquery-0.3.2}/benchmarks/performance/collect_query_scaling.py +0 -0
  48. {afquery-0.3.1 → afquery-0.3.2}/benchmarks/performance/config.py +0 -0
  49. {afquery-0.3.1 → afquery-0.3.2}/benchmarks/performance/config_smoke.py +0 -0
  50. {afquery-0.3.1 → afquery-0.3.2}/benchmarks/shared/__init__.py +0 -0
  51. {afquery-0.3.1 → afquery-0.3.2}/benchmarks/shared/config.py +0 -0
  52. {afquery-0.3.1 → afquery-0.3.2}/benchmarks/shared/rules/download_1kg.smk +0 -0
  53. {afquery-0.3.1 → afquery-0.3.2}/benchmarks/shared/utils.py +0 -0
  54. {afquery-0.3.1 → afquery-0.3.2}/docs/advanced/benchmarking.md +0 -0
  55. {afquery-0.3.1 → afquery-0.3.2}/docs/advanced/coverage-evidence.md +0 -0
  56. {afquery-0.3.1 → afquery-0.3.2}/docs/advanced/debugging-results.md +0 -0
  57. {afquery-0.3.1 → afquery-0.3.2}/docs/advanced/filter-pass-tracking.md +0 -0
  58. {afquery-0.3.1 → afquery-0.3.2}/docs/advanced/multi-cohort-strategies.md +0 -0
  59. {afquery-0.3.1 → afquery-0.3.2}/docs/advanced/performance.md +0 -0
  60. {afquery-0.3.1 → afquery-0.3.2}/docs/advanced/pipeline-integration.md +0 -0
  61. {afquery-0.3.1 → afquery-0.3.2}/docs/advanced/ploidy-and-sex-chroms.md +0 -0
  62. {afquery-0.3.1 → afquery-0.3.2}/docs/assets/img/gap2_mixed_technologies.png +0 -0
  63. {afquery-0.3.1 → afquery-0.3.2}/docs/faq.md +0 -0
  64. {afquery-0.3.1 → afquery-0.3.2}/docs/getting-started/concepts.md +0 -0
  65. {afquery-0.3.1 → afquery-0.3.2}/docs/getting-started/installation.md +0 -0
  66. {afquery-0.3.1 → afquery-0.3.2}/docs/getting-started/motivation.md +0 -0
  67. {afquery-0.3.1 → afquery-0.3.2}/docs/getting-started/preprocessing.md +0 -0
  68. {afquery-0.3.1 → afquery-0.3.2}/docs/getting-started/quickstart.md +0 -0
  69. {afquery-0.3.1 → afquery-0.3.2}/docs/getting-started/tutorial.md +0 -0
  70. {afquery-0.3.1 → afquery-0.3.2}/docs/getting-started/understanding-output.md +0 -0
  71. {afquery-0.3.1 → afquery-0.3.2}/docs/guides/annotate-vcf.md +0 -0
  72. {afquery-0.3.1 → afquery-0.3.2}/docs/guides/create-database.md +0 -0
  73. {afquery-0.3.1 → afquery-0.3.2}/docs/guides/dump-export.md +0 -0
  74. {afquery-0.3.1 → afquery-0.3.2}/docs/guides/manifest-format.md +0 -0
  75. {afquery-0.3.1 → afquery-0.3.2}/docs/guides/query.md +0 -0
  76. {afquery-0.3.1 → afquery-0.3.2}/docs/guides/sample-filtering.md +0 -0
  77. {afquery-0.3.1 → afquery-0.3.2}/docs/guides/update-database.md +0 -0
  78. {afquery-0.3.1 → afquery-0.3.2}/docs/guides/variant-info.md +0 -0
  79. {afquery-0.3.1 → afquery-0.3.2}/docs/index.md +0 -0
  80. {afquery-0.3.1 → afquery-0.3.2}/docs/reference/cli.md +0 -0
  81. {afquery-0.3.1 → afquery-0.3.2}/docs/reference/data-model.md +0 -0
  82. {afquery-0.3.1 → afquery-0.3.2}/docs/reference/glossary.md +0 -0
  83. {afquery-0.3.1 → afquery-0.3.2}/docs/reference/python-api.md +0 -0
  84. {afquery-0.3.1 → afquery-0.3.2}/docs/scripts/gen_gap2_figure.py +0 -0
  85. {afquery-0.3.1 → afquery-0.3.2}/docs/stylesheets/extra.css +0 -0
  86. {afquery-0.3.1 → afquery-0.3.2}/docs/troubleshooting.md +0 -0
  87. {afquery-0.3.1 → afquery-0.3.2}/docs/use-cases/acmg-use-cases.md +0 -0
  88. {afquery-0.3.1 → afquery-0.3.2}/docs/use-cases/clinical-prioritization.md +0 -0
  89. {afquery-0.3.1 → afquery-0.3.2}/docs/use-cases/cohort-stratification.md +0 -0
  90. {afquery-0.3.1 → afquery-0.3.2}/docs/use-cases/population-specific-af.md +0 -0
  91. {afquery-0.3.1 → afquery-0.3.2}/docs/use-cases/pseudo-controls.md +0 -0
  92. {afquery-0.3.1 → afquery-0.3.2}/docs/use-cases/sex-specific-af.md +0 -0
  93. {afquery-0.3.1 → afquery-0.3.2}/docs/use-cases/technology-integration.md +0 -0
  94. {afquery-0.3.1 → afquery-0.3.2}/examples/demo/README.md +0 -0
  95. {afquery-0.3.1 → afquery-0.3.2}/examples/demo/create_demo_data.py +0 -0
  96. {afquery-0.3.1 → afquery-0.3.2}/mkdocs.yml +0 -0
  97. {afquery-0.3.1 → afquery-0.3.2}/pyproject.toml +0 -0
  98. {afquery-0.3.1 → afquery-0.3.2}/recipes/afquery/meta.yaml +0 -0
  99. {afquery-0.3.1 → afquery-0.3.2}/resources/normalize_vcf.sh +0 -0
  100. {afquery-0.3.1 → afquery-0.3.2}/src/afquery/__init__.py +0 -0
  101. {afquery-0.3.1 → afquery-0.3.2}/src/afquery/annotate.py +0 -0
  102. {afquery-0.3.1 → afquery-0.3.2}/src/afquery/benchmark.py +0 -0
  103. {afquery-0.3.1 → afquery-0.3.2}/src/afquery/bitmaps.py +0 -0
  104. {afquery-0.3.1 → afquery-0.3.2}/src/afquery/capture.py +0 -0
  105. {afquery-0.3.1 → afquery-0.3.2}/src/afquery/cli.py +0 -0
  106. {afquery-0.3.1 → afquery-0.3.2}/src/afquery/cli.pyr +0 -0
  107. {afquery-0.3.1 → afquery-0.3.2}/src/afquery/constants.py +0 -0
  108. {afquery-0.3.1 → afquery-0.3.2}/src/afquery/database.py +0 -0
  109. {afquery-0.3.1 → afquery-0.3.2}/src/afquery/dump.py +0 -0
  110. {afquery-0.3.1 → afquery-0.3.2}/src/afquery/models.py +0 -0
  111. {afquery-0.3.1 → afquery-0.3.2}/src/afquery/ploidy.py +0 -0
  112. {afquery-0.3.1 → afquery-0.3.2}/src/afquery/preprocess/__init__.py +0 -0
  113. {afquery-0.3.1 → afquery-0.3.2}/src/afquery/preprocess/build.py +0 -0
  114. {afquery-0.3.1 → afquery-0.3.2}/src/afquery/preprocess/compact.py +0 -0
  115. {afquery-0.3.1 → afquery-0.3.2}/src/afquery/preprocess/ingest.py +0 -0
  116. {afquery-0.3.1 → afquery-0.3.2}/src/afquery/preprocess/manifest.py +0 -0
  117. {afquery-0.3.1 → afquery-0.3.2}/src/afquery/preprocess/regions.py +0 -0
  118. {afquery-0.3.1 → afquery-0.3.2}/src/afquery/preprocess/synth.py +0 -0
  119. {afquery-0.3.1 → afquery-0.3.2}/src/afquery/preprocess/update.py +0 -0
  120. {afquery-0.3.1 → afquery-0.3.2}/src/afquery/query.py +0 -0
  121. {afquery-0.3.1 → afquery-0.3.2}/src/afquery/variant_info.py +0 -0
  122. {afquery-0.3.1 → afquery-0.3.2}/tests/conftest.py +0 -0
  123. {afquery-0.3.1 → afquery-0.3.2}/tests/data/annotate_input.vcf +0 -0
  124. {afquery-0.3.1 → afquery-0.3.2}/tests/data/annotate_multi_bucket.vcf +0 -0
  125. {afquery-0.3.1 → afquery-0.3.2}/tests/data/annotate_multi_chrom.vcf +0 -0
  126. {afquery-0.3.1 → afquery-0.3.2}/tests/data/beds/wes_kit_a.bed +0 -0
  127. {afquery-0.3.1 → afquery-0.3.2}/tests/data/beds/wes_kit_b.bed +0 -0
  128. {afquery-0.3.1 → afquery-0.3.2}/tests/data/expected_results.json +0 -0
  129. {afquery-0.3.1 → afquery-0.3.2}/tests/data/manifest.tsv +0 -0
  130. {afquery-0.3.1 → afquery-0.3.2}/tests/data/vcfs/S00.vcf +0 -0
  131. {afquery-0.3.1 → afquery-0.3.2}/tests/data/vcfs/S01.vcf +0 -0
  132. {afquery-0.3.1 → afquery-0.3.2}/tests/data/vcfs/S02.vcf +0 -0
  133. {afquery-0.3.1 → afquery-0.3.2}/tests/data/vcfs/S03.vcf +0 -0
  134. {afquery-0.3.1 → afquery-0.3.2}/tests/data/vcfs/S04.vcf +0 -0
  135. {afquery-0.3.1 → afquery-0.3.2}/tests/data/vcfs/S05.vcf +0 -0
  136. {afquery-0.3.1 → afquery-0.3.2}/tests/data/vcfs/S06.vcf +0 -0
  137. {afquery-0.3.1 → afquery-0.3.2}/tests/data/vcfs/S07.vcf +0 -0
  138. {afquery-0.3.1 → afquery-0.3.2}/tests/data/vcfs/S08.vcf +0 -0
  139. {afquery-0.3.1 → afquery-0.3.2}/tests/data/vcfs/S09.vcf +0 -0
  140. {afquery-0.3.1 → afquery-0.3.2}/tests/test_annotate.py +0 -0
  141. {afquery-0.3.1 → afquery-0.3.2}/tests/test_batch.py +0 -0
  142. {afquery-0.3.1 → afquery-0.3.2}/tests/test_benchmark.py +0 -0
  143. {afquery-0.3.1 → afquery-0.3.2}/tests/test_bitmaps.py +0 -0
  144. {afquery-0.3.1 → afquery-0.3.2}/tests/test_capture.py +0 -0
  145. {afquery-0.3.1 → afquery-0.3.2}/tests/test_cli.py +0 -0
  146. {afquery-0.3.1 → afquery-0.3.2}/tests/test_cli_docs_consistency.py +0 -0
  147. {afquery-0.3.1 → afquery-0.3.2}/tests/test_compact.py +0 -0
  148. {afquery-0.3.1 → afquery-0.3.2}/tests/test_constants.py +0 -0
  149. {afquery-0.3.1 → afquery-0.3.2}/tests/test_dump.py +0 -0
  150. {afquery-0.3.1 → afquery-0.3.2}/tests/test_haploid_stats.py +0 -0
  151. {afquery-0.3.1 → afquery-0.3.2}/tests/test_info.py +0 -0
  152. {afquery-0.3.1 → afquery-0.3.2}/tests/test_no_coverage.py +0 -0
  153. {afquery-0.3.1 → afquery-0.3.2}/tests/test_pass_filter.py +0 -0
  154. {afquery-0.3.1 → afquery-0.3.2}/tests/test_ploidy.py +0 -0
  155. {afquery-0.3.1 → afquery-0.3.2}/tests/test_preprocess.py +0 -0
  156. {afquery-0.3.1 → afquery-0.3.2}/tests/test_query.py +0 -0
  157. {afquery-0.3.1 → afquery-0.3.2}/tests/test_sample_filter.py +0 -0
  158. {afquery-0.3.1 → afquery-0.3.2}/tests/test_synth.py +0 -0
  159. {afquery-0.3.1 → afquery-0.3.2}/tests/test_synthetic_stats.py +0 -0
  160. {afquery-0.3.1 → afquery-0.3.2}/tests/test_update.py +0 -0
  161. {afquery-0.3.1 → afquery-0.3.2}/tests/test_update_metadata.py +0 -0
  162. {afquery-0.3.1 → afquery-0.3.2}/tests/test_variant_info.py +0 -0
  163. {afquery-0.3.1 → afquery-0.3.2}/tests/test_warnings.py +0 -0
afquery-0.3.2/PKG-INFO ADDED
@@ -0,0 +1,256 @@
1
+ Metadata-Version: 2.4
2
+ Name: afquery
3
+ Version: 0.3.2
4
+ Summary: Genomic allele frequency query engine with bitmap-encoded genotypes
5
+ License: MIT
6
+ License-File: LICENSE
7
+ Classifier: Programming Language :: Python :: 3
8
+ Classifier: Programming Language :: Python :: 3.10
9
+ Classifier: Programming Language :: Python :: 3.11
10
+ Classifier: Programming Language :: Python :: 3.12
11
+ Requires-Python: >=3.10
12
+ Requires-Dist: click>=8.1
13
+ Requires-Dist: cyvcf2>=0.30
14
+ Requires-Dist: duckdb>=0.10
15
+ Requires-Dist: pyarrow>=14.0
16
+ Requires-Dist: pyranges<0.2,>=0.1.2
17
+ Requires-Dist: pyroaring>=0.4.8
18
+ Requires-Dist: tqdm>=4.60
19
+ Provides-Extra: dev
20
+ Requires-Dist: pytest-cov>=5.0; extra == 'dev'
21
+ Requires-Dist: pytest>=8.0; extra == 'dev'
22
+ Provides-Extra: docs
23
+ Requires-Dist: matplotlib>=3.8; extra == 'docs'
24
+ Requires-Dist: mike>=2.0; extra == 'docs'
25
+ Requires-Dist: mkdocs-material>=9.5; extra == 'docs'
26
+ Requires-Dist: mkdocs-print-site-plugin>=2.6.1; extra == 'docs'
27
+ Requires-Dist: mkdocstrings[python]>=0.25; extra == 'docs'
28
+ Requires-Dist: pymdown-extensions>=10.7; extra == 'docs'
29
+ Description-Content-Type: text/markdown
30
+
31
+ # AFQuery
32
+
33
+ [![CI](https://github.com/babelomics/afquery/actions/workflows/ci.yml/badge.svg)](https://github.com/babelomics/afquery/actions/workflows/ci.yml)
34
+ [![Coverage](https://codecov.io/gh/babelomics/afquery/graph/badge.svg)](https://codecov.io/gh/babelomics/afquery)
35
+ [![Docs](https://img.shields.io/website?url=https%3A%2F%2Fbabelomics.github.io%2Fafquery%2F&label=docs)](https://babelomics.github.io/afquery/)
36
+ <br>
37
+ [![PyPI](https://img.shields.io/pypi/v/afquery.svg?color=blue)](https://pypi.org/project/afquery/)
38
+ [![Bioconda](https://img.shields.io/conda/vn/bioconda/afquery.svg)](https://anaconda.org/bioconda/afquery)
39
+ [![Docker](https://img.shields.io/badge/ghcr.io-afquery-blue?logo=docker)](https://github.com/babelomics/afquery/pkgs/container/afquery)
40
+ [![Python](https://img.shields.io/pypi/pyversions/afquery.svg)](https://pypi.org/project/afquery/)
41
+ [![License: MIT](https://img.shields.io/github/license/babelomics/afquery)](https://github.com/babelomics/afquery/blob/master/LICENSE)
42
+
43
+ **Fast, capture-aware allele frequency queries on local genomic cohorts — without re-scanning VCFs.**
44
+
45
+ AFQuery is a bitmap-indexed engine that recomputes AC/AN/AF for arbitrary subcohorts (by phenotype, sex, sequencing technology, or any combination) in tens of milliseconds, independently of cohort size. It accounts for capture-kit heterogeneity, ploidy on sex chromosomes, and FILTER/coverage evidence, and runs locally as a file-based system (Parquet + SQLite) with no server or cluster required.
46
+
47
+ [Full documentation →](https://babelomics.github.io/afquery/)
48
+
49
+ ---
50
+
51
+ ## Headline results
52
+
53
+ - **~14 ms point queries**, constant from 1,000 to 50,000 samples (O(1) scaling).
54
+ - **33× faster than bcftools** for full-chromosome bulk export at 2,504 samples; **R² > 0.99999** AF concordance over 1.1M common variants.
55
+ - **~25,000 variants/s** VCF annotation on 4 cores — a typical exome annotates in **~1 second**.
56
+ - **Up to 45× reduction** in toward-pathogenic ACMG classification errors vs. naive AN on mixed-capture-kit cohorts.
57
+ - **9.5×–13.4× storage compression** vs. input single-sample VCFs.
58
+
59
+ ## Why AFQuery
60
+
61
+ Local allele frequencies are a core input to ACMG/AMP variant classification (criteria BA1, BS1, PM2), and global resources like gnomAD systematically underrepresent local ancestries and disease-enriched institutional cohorts. Computing AF on your own cohort sounds simple — until the cohort mixes WGS, WES, and panels, or several versions of the same capture kit. Successive Agilent SureSelect versions (v5/v6/v7), for example, only share ~57% of their targets on chromosome 22. Naive AN counting then *inflates* AN at positions outside some kits' targets, *deflates* AF, and systematically shifts variants toward "pathogenic" by ACMG criteria.
62
+
63
+ AFQuery solves this with per-position, per-technology capture-aware AN, ploidy-aware sex chromosome handling, and an explicit `N_NO_COVERAGE` channel that separates trusted hom-ref from "we cannot tell". Queries on subcohorts are answered by intersecting precomputed Roaring Bitmaps, so latency is constant in cohort size.
64
+
65
+ ## When to use AFQuery
66
+
67
+ - You need allele frequencies for phenotype-defined or arbitrarily filtered subcohorts.
68
+ - Your cohort mixes sequencing technologies (WGS, WES, panels, multiple capture-kit versions).
69
+ - You want fast, repeated, interactive queries instead of one-off VCF re-scans.
70
+ - You need a local, reproducible workflow — no cloud, no Spark cluster.
71
+
72
+ ## Features
73
+
74
+ - **Constant-time subcohort queries** — bitmap intersections at query time; no per-query VCF re-scan.
75
+ - **Capture-aware AN** — per-position eligibility from each technology's BED, eliminating systematic AF bias when mixing WGS / WES / panels and kit versions.
76
+ - **Ploidy-aware sex chromosomes** — correct AN on chrX PAR / non-PAR, chrY, chrM, by sample sex.
77
+ - **Coverage evidence model** — `N_NO_COVERAGE` separates trusted hom-ref from samples lacking sufficient evidence; query-time gates (`--min-pass`, `--min-observed`, `--min-quality-evidence`) keep AF conservative.
78
+ - **ACMG-compatible AC/AN/AF** — per-standard definitions, exposed as both query output and VCF INFO fields.
79
+ - **Flexible metadata filtering** — arbitrary phenotype labels (ICD-10, HPO, OMIM, custom tags), inclusion or exclusion (`^` prefix), combined with sex and technology.
80
+ - **Parallel VCF annotation** — multi-threaded; adds `AFQUERY_AC/AN/AF/N_HET/N_HOM_ALT/N_HOM_REF/N_FAIL/N_NO_COVERAGE` INFO fields.
81
+ - **Bulk CSV export** — per-variant frequencies with optional disaggregation by sex, technology, or phenotype.
82
+ - **Incremental updates** — add or remove samples, edit phenotype/sex metadata, compact storage, without full rebuilds.
83
+ - **Audit changelog** — every database operation is logged with timestamps and operator notes.
84
+ - **Database validation** — `afquery check` with scripted exit codes.
85
+ - **Serverless** — Parquet + SQLite on disk; no daemon, no Java, no Spark.
86
+
87
+ ## Installation
88
+
89
+ ```bash
90
+ # PyPI
91
+ pip install afquery
92
+
93
+ # Bioconda
94
+ conda install -c bioconda afquery
95
+
96
+ # Docker (linux/amd64, linux/arm64)
97
+ docker pull ghcr.io/babelomics/afquery:latest
98
+
99
+ # From source
100
+ git clone https://github.com/babelomics/afquery.git
101
+ cd afquery
102
+ pip install -e .
103
+ ```
104
+
105
+ Requires Python ≥ 3.10. Core dependencies: `pyroaring`, `pyarrow`, `duckdb`, `pyranges`, `cyvcf2`, `click`, `tqdm`.
106
+
107
+ ## Quickstart
108
+
109
+ ### 1. Prepare a manifest
110
+
111
+ One row per sample (TSV, header required):
112
+
113
+ ```tsv
114
+ sample_name vcf_path sex tech_name phenotype_codes
115
+ SAMP_001 /data/vcfs/SAMP_001.vcf.gz female wgs E11.9,I10
116
+ SAMP_002 /data/vcfs/SAMP_002.vcf.gz male wes_v6 E11.9
117
+ SAMP_003 /data/vcfs/SAMP_003.vcf.gz female panel_card I42.0
118
+ ```
119
+
120
+ For every non-WGS technology, place a 3-column BED in `--bed-dir` named `<tech_name>.bed`.
121
+
122
+ ### 2. Build the database
123
+
124
+ ```bash
125
+ afquery create-db \
126
+ --manifest manifest.tsv \
127
+ --output-dir ./db/ \
128
+ --genome-build GRCh38 \
129
+ --bed-dir ./beds/
130
+ ```
131
+
132
+ ### 3. Query
133
+
134
+ ```bash
135
+ # Single locus
136
+ afquery query --db ./db/ --locus chr1:925952
137
+
138
+ # Locus filtered to a phenotype-and-sex subcohort
139
+ afquery query --db ./db/ --locus chr1:925952 --phenotype E11.9 --sex female
140
+
141
+ # Genomic region
142
+ afquery query --db ./db/ --region chr1:900000-1000000
143
+
144
+ # Batch from file (chrom pos [ref [alt]] per line)
145
+ afquery query --db ./db/ --from-file variants.tsv
146
+ ```
147
+
148
+ ### 4. Inspect carriers of a variant
149
+
150
+ ```bash
151
+ afquery variant-info --db ./db/ --locus chr17:43093454
152
+ ```
153
+
154
+ ### 5. Annotate a VCF
155
+
156
+ ```bash
157
+ afquery annotate \
158
+ --db ./db/ \
159
+ --input patient.vcf \
160
+ --output patient.annotated.vcf \
161
+ --threads 4
162
+ ```
163
+
164
+ ### 6. Export
165
+
166
+ ```bash
167
+ # Region export, disaggregated by sex
168
+ afquery dump --db ./db/ --chrom chr17 --start 43044292 --end 43170327 \
169
+ --output brca1.csv --by-sex
170
+ ```
171
+
172
+ ### 7. Update
173
+
174
+ ```bash
175
+ afquery update-db --db ./db/ --add-samples new_batch.tsv
176
+ afquery update-db --db ./db/ --update-sample SAMP_007 --set-phenotype I42.0
177
+ ```
178
+
179
+ ## Output fields
180
+
181
+ Every query and annotated VCF reports:
182
+
183
+ | Field | Meaning |
184
+ |---|---|
185
+ | `AC` | Alt-allele count over eligible samples (FILTER=PASS) |
186
+ | `AN` | Total alleles considered, ploidy- and capture-aware |
187
+ | `AF` | `AC / AN` |
188
+ | `N_HET` | Heterozygous PASS carriers |
189
+ | `N_HOM_ALT` | Homozygous-alt PASS carriers |
190
+ | `N_HOM_REF` | Samples trusted as homozygous reference |
191
+ | `N_FAIL` | Carriers with FILTER ≠ PASS (excluded from AC/AN) |
192
+ | `N_NO_COVERAGE` | Non-carriers on a partially-covered tech without sufficient evidence to call hom-ref |
193
+
194
+ VCF INFO field names use the `AFQUERY_` prefix (e.g. `AFQUERY_AF`, `AFQUERY_N_NO_COVERAGE`).
195
+
196
+ ## CLI commands
197
+
198
+ | Command | Purpose |
199
+ |---|---|
200
+ | `create-db` | Build a database from a manifest of single-sample VCFs |
201
+ | `query` | Point / region / batch AF queries |
202
+ | `variant-info` | List samples carrying a variant, with metadata |
203
+ | `annotate` | Annotate a VCF with cohort `AFQUERY_*` INFO fields |
204
+ | `dump` | Bulk CSV export, optionally disaggregated |
205
+ | `update-db` | Add / remove samples, edit metadata, compact |
206
+ | `info` | Show database metadata, sample list, changelog |
207
+ | `check` | Validate database integrity (scripted exit code) |
208
+ | `version show` / `version set` | Inspect or set the database version label |
209
+ | `benchmark` | Run synthetic or on-database performance benchmarks |
210
+
211
+ See the [CLI reference](https://babelomics.github.io/afquery/reference/cli/) for all options.
212
+
213
+ ## How it works
214
+
215
+ AFQuery indexes each variant as three Roaring Bitmaps — heterozygous PASS carriers, homozygous-alt PASS carriers, and FILTER≠PASS carriers — stored in Apache Parquet, partitioned by chromosome and 1-Mbp positional buckets. Sample metadata (sex, technology, phenotype) is precomputed as bitmaps in SQLite. A query resolves its sample filter into a single candidate bitmap by intersection/difference in microseconds, then intersects it against each variant's genotype bitmaps to compute AC. AN is computed per position from the same candidate bitmap, restricted to samples whose technology actually covers the position (via BED-derived capture indices) and adjusted for ploidy on sex chromosomes. See the [data model reference](https://babelomics.github.io/afquery/reference/data-model/) for details.
216
+
217
+ ## How AFQuery compares
218
+
219
+ | | AFQuery | bcftools | GATK GenomicsDB | Hail |
220
+ |---|---|---|---|---|
221
+ | Capture-aware AN | **Yes** | No | No | No |
222
+ | Metadata filtering | **Arbitrary labels** | No | No | Custom code |
223
+ | Ploidy-aware sex chromosomes | **Yes** | Manual | No | Manual |
224
+ | Dynamic subcohort queries | **Yes** | No | Limited | Requires code |
225
+ | FILTER / coverage tracking | **Per variant** | Manual | No | Manual |
226
+ | Incremental updates | **Yes** | No | **Yes** | No |
227
+ | Infrastructure required | **None** | **None** | Java/server | Spark cluster |
228
+
229
+ ### Benchmarks vs. bcftools (1000 Genomes Phase 3, n = 2,504, chr22)
230
+
231
+ | Workload | AFQuery | bcftools | Speedup |
232
+ |---|---|---|---|
233
+ | Full-chromosome AC/AN/AF export | ~7.0 s | ~3.8 min | **~33×** |
234
+ | AF concordance over 1,106,181 common variants | — | — | **R² > 0.99999** |
235
+
236
+ Point-query latency on AFQuery is ~14 ms and constant from 1K to 50K samples (median over 50 replicates, warm cache).
237
+
238
+ ## Documentation
239
+
240
+ - [Full documentation](https://babelomics.github.io/afquery/)
241
+ - [Getting started](https://babelomics.github.io/afquery/getting-started/)
242
+ - [Why local allele frequencies?](https://babelomics.github.io/afquery/getting-started/motivation/)
243
+ - [Manifest format](https://babelomics.github.io/afquery/guides/manifest-format/)
244
+ - [CLI reference](https://babelomics.github.io/afquery/reference/cli/)
245
+ - [FAQ](https://babelomics.github.io/afquery/faq/)
246
+
247
+ ## Citation
248
+
249
+ If you use AFQuery in your work, please cite:
250
+
251
+ > AFQuery: fast, capture-aware allele frequency queries on local genomic cohorts.
252
+ > *(manuscript in preparation)*
253
+
254
+ ## License
255
+
256
+ [MIT](LICENSE)
@@ -0,0 +1,226 @@
1
+ # AFQuery
2
+
3
+ [![CI](https://github.com/babelomics/afquery/actions/workflows/ci.yml/badge.svg)](https://github.com/babelomics/afquery/actions/workflows/ci.yml)
4
+ [![Coverage](https://codecov.io/gh/babelomics/afquery/graph/badge.svg)](https://codecov.io/gh/babelomics/afquery)
5
+ [![Docs](https://img.shields.io/website?url=https%3A%2F%2Fbabelomics.github.io%2Fafquery%2F&label=docs)](https://babelomics.github.io/afquery/)
6
+ <br>
7
+ [![PyPI](https://img.shields.io/pypi/v/afquery.svg?color=blue)](https://pypi.org/project/afquery/)
8
+ [![Bioconda](https://img.shields.io/conda/vn/bioconda/afquery.svg)](https://anaconda.org/bioconda/afquery)
9
+ [![Docker](https://img.shields.io/badge/ghcr.io-afquery-blue?logo=docker)](https://github.com/babelomics/afquery/pkgs/container/afquery)
10
+ [![Python](https://img.shields.io/pypi/pyversions/afquery.svg)](https://pypi.org/project/afquery/)
11
+ [![License: MIT](https://img.shields.io/github/license/babelomics/afquery)](https://github.com/babelomics/afquery/blob/master/LICENSE)
12
+
13
+ **Fast, capture-aware allele frequency queries on local genomic cohorts — without re-scanning VCFs.**
14
+
15
+ AFQuery is a bitmap-indexed engine that recomputes AC/AN/AF for arbitrary subcohorts (by phenotype, sex, sequencing technology, or any combination) in tens of milliseconds, independently of cohort size. It accounts for capture-kit heterogeneity, ploidy on sex chromosomes, and FILTER/coverage evidence, and runs locally as a file-based system (Parquet + SQLite) with no server or cluster required.
16
+
17
+ [Full documentation →](https://babelomics.github.io/afquery/)
18
+
19
+ ---
20
+
21
+ ## Headline results
22
+
23
+ - **~14 ms point queries**, constant from 1,000 to 50,000 samples (O(1) scaling).
24
+ - **33× faster than bcftools** for full-chromosome bulk export at 2,504 samples; **R² > 0.99999** AF concordance over 1.1M common variants.
25
+ - **~25,000 variants/s** VCF annotation on 4 cores — a typical exome annotates in **~1 second**.
26
+ - **Up to 45× reduction** in toward-pathogenic ACMG classification errors vs. naive AN on mixed-capture-kit cohorts.
27
+ - **9.5×–13.4× storage compression** vs. input single-sample VCFs.
28
+
29
+ ## Why AFQuery
30
+
31
+ Local allele frequencies are a core input to ACMG/AMP variant classification (criteria BA1, BS1, PM2), and global resources like gnomAD systematically underrepresent local ancestries and disease-enriched institutional cohorts. Computing AF on your own cohort sounds simple — until the cohort mixes WGS, WES, and panels, or several versions of the same capture kit. Successive Agilent SureSelect versions (v5/v6/v7), for example, only share ~57% of their targets on chromosome 22. Naive AN counting then *inflates* AN at positions outside some kits' targets, *deflates* AF, and systematically shifts variants toward "pathogenic" by ACMG criteria.
32
+
33
+ AFQuery solves this with per-position, per-technology capture-aware AN, ploidy-aware sex chromosome handling, and an explicit `N_NO_COVERAGE` channel that separates trusted hom-ref from "we cannot tell". Queries on subcohorts are answered by intersecting precomputed Roaring Bitmaps, so latency is constant in cohort size.
34
+
35
+ ## When to use AFQuery
36
+
37
+ - You need allele frequencies for phenotype-defined or arbitrarily filtered subcohorts.
38
+ - Your cohort mixes sequencing technologies (WGS, WES, panels, multiple capture-kit versions).
39
+ - You want fast, repeated, interactive queries instead of one-off VCF re-scans.
40
+ - You need a local, reproducible workflow — no cloud, no Spark cluster.
41
+
42
+ ## Features
43
+
44
+ - **Constant-time subcohort queries** — bitmap intersections at query time; no per-query VCF re-scan.
45
+ - **Capture-aware AN** — per-position eligibility from each technology's BED, eliminating systematic AF bias when mixing WGS / WES / panels and kit versions.
46
+ - **Ploidy-aware sex chromosomes** — correct AN on chrX PAR / non-PAR, chrY, chrM, by sample sex.
47
+ - **Coverage evidence model** — `N_NO_COVERAGE` separates trusted hom-ref from samples lacking sufficient evidence; query-time gates (`--min-pass`, `--min-observed`, `--min-quality-evidence`) keep AF conservative.
48
+ - **ACMG-compatible AC/AN/AF** — per-standard definitions, exposed as both query output and VCF INFO fields.
49
+ - **Flexible metadata filtering** — arbitrary phenotype labels (ICD-10, HPO, OMIM, custom tags), inclusion or exclusion (`^` prefix), combined with sex and technology.
50
+ - **Parallel VCF annotation** — multi-threaded; adds `AFQUERY_AC/AN/AF/N_HET/N_HOM_ALT/N_HOM_REF/N_FAIL/N_NO_COVERAGE` INFO fields.
51
+ - **Bulk CSV export** — per-variant frequencies with optional disaggregation by sex, technology, or phenotype.
52
+ - **Incremental updates** — add or remove samples, edit phenotype/sex metadata, compact storage, without full rebuilds.
53
+ - **Audit changelog** — every database operation is logged with timestamps and operator notes.
54
+ - **Database validation** — `afquery check` with scripted exit codes.
55
+ - **Serverless** — Parquet + SQLite on disk; no daemon, no Java, no Spark.
56
+
57
+ ## Installation
58
+
59
+ ```bash
60
+ # PyPI
61
+ pip install afquery
62
+
63
+ # Bioconda
64
+ conda install -c bioconda afquery
65
+
66
+ # Docker (linux/amd64, linux/arm64)
67
+ docker pull ghcr.io/babelomics/afquery:latest
68
+
69
+ # From source
70
+ git clone https://github.com/babelomics/afquery.git
71
+ cd afquery
72
+ pip install -e .
73
+ ```
74
+
75
+ Requires Python ≥ 3.10. Core dependencies: `pyroaring`, `pyarrow`, `duckdb`, `pyranges`, `cyvcf2`, `click`, `tqdm`.
76
+
77
+ ## Quickstart
78
+
79
+ ### 1. Prepare a manifest
80
+
81
+ One row per sample (TSV, header required):
82
+
83
+ ```tsv
84
+ sample_name vcf_path sex tech_name phenotype_codes
85
+ SAMP_001 /data/vcfs/SAMP_001.vcf.gz female wgs E11.9,I10
86
+ SAMP_002 /data/vcfs/SAMP_002.vcf.gz male wes_v6 E11.9
87
+ SAMP_003 /data/vcfs/SAMP_003.vcf.gz female panel_card I42.0
88
+ ```
89
+
90
+ For every non-WGS technology, place a 3-column BED in `--bed-dir` named `<tech_name>.bed`.
91
+
92
+ ### 2. Build the database
93
+
94
+ ```bash
95
+ afquery create-db \
96
+ --manifest manifest.tsv \
97
+ --output-dir ./db/ \
98
+ --genome-build GRCh38 \
99
+ --bed-dir ./beds/
100
+ ```
101
+
102
+ ### 3. Query
103
+
104
+ ```bash
105
+ # Single locus
106
+ afquery query --db ./db/ --locus chr1:925952
107
+
108
+ # Locus filtered to a phenotype-and-sex subcohort
109
+ afquery query --db ./db/ --locus chr1:925952 --phenotype E11.9 --sex female
110
+
111
+ # Genomic region
112
+ afquery query --db ./db/ --region chr1:900000-1000000
113
+
114
+ # Batch from file (chrom pos [ref [alt]] per line)
115
+ afquery query --db ./db/ --from-file variants.tsv
116
+ ```
117
+
118
+ ### 4. Inspect carriers of a variant
119
+
120
+ ```bash
121
+ afquery variant-info --db ./db/ --locus chr17:43093454
122
+ ```
123
+
124
+ ### 5. Annotate a VCF
125
+
126
+ ```bash
127
+ afquery annotate \
128
+ --db ./db/ \
129
+ --input patient.vcf \
130
+ --output patient.annotated.vcf \
131
+ --threads 4
132
+ ```
133
+
134
+ ### 6. Export
135
+
136
+ ```bash
137
+ # Region export, disaggregated by sex
138
+ afquery dump --db ./db/ --chrom chr17 --start 43044292 --end 43170327 \
139
+ --output brca1.csv --by-sex
140
+ ```
141
+
142
+ ### 7. Update
143
+
144
+ ```bash
145
+ afquery update-db --db ./db/ --add-samples new_batch.tsv
146
+ afquery update-db --db ./db/ --update-sample SAMP_007 --set-phenotype I42.0
147
+ ```
148
+
149
+ ## Output fields
150
+
151
+ Every query and annotated VCF reports:
152
+
153
+ | Field | Meaning |
154
+ |---|---|
155
+ | `AC` | Alt-allele count over eligible samples (FILTER=PASS) |
156
+ | `AN` | Total alleles considered, ploidy- and capture-aware |
157
+ | `AF` | `AC / AN` |
158
+ | `N_HET` | Heterozygous PASS carriers |
159
+ | `N_HOM_ALT` | Homozygous-alt PASS carriers |
160
+ | `N_HOM_REF` | Samples trusted as homozygous reference |
161
+ | `N_FAIL` | Carriers with FILTER ≠ PASS (excluded from AC/AN) |
162
+ | `N_NO_COVERAGE` | Non-carriers on a partially-covered tech without sufficient evidence to call hom-ref |
163
+
164
+ VCF INFO field names use the `AFQUERY_` prefix (e.g. `AFQUERY_AF`, `AFQUERY_N_NO_COVERAGE`).
165
+
166
+ ## CLI commands
167
+
168
+ | Command | Purpose |
169
+ |---|---|
170
+ | `create-db` | Build a database from a manifest of single-sample VCFs |
171
+ | `query` | Point / region / batch AF queries |
172
+ | `variant-info` | List samples carrying a variant, with metadata |
173
+ | `annotate` | Annotate a VCF with cohort `AFQUERY_*` INFO fields |
174
+ | `dump` | Bulk CSV export, optionally disaggregated |
175
+ | `update-db` | Add / remove samples, edit metadata, compact |
176
+ | `info` | Show database metadata, sample list, changelog |
177
+ | `check` | Validate database integrity (scripted exit code) |
178
+ | `version show` / `version set` | Inspect or set the database version label |
179
+ | `benchmark` | Run synthetic or on-database performance benchmarks |
180
+
181
+ See the [CLI reference](https://babelomics.github.io/afquery/reference/cli/) for all options.
182
+
183
+ ## How it works
184
+
185
+ AFQuery indexes each variant as three Roaring Bitmaps — heterozygous PASS carriers, homozygous-alt PASS carriers, and FILTER≠PASS carriers — stored in Apache Parquet, partitioned by chromosome and 1-Mbp positional buckets. Sample metadata (sex, technology, phenotype) is precomputed as bitmaps in SQLite. A query resolves its sample filter into a single candidate bitmap by intersection/difference in microseconds, then intersects it against each variant's genotype bitmaps to compute AC. AN is computed per position from the same candidate bitmap, restricted to samples whose technology actually covers the position (via BED-derived capture indices) and adjusted for ploidy on sex chromosomes. See the [data model reference](https://babelomics.github.io/afquery/reference/data-model/) for details.
186
+
187
+ ## How AFQuery compares
188
+
189
+ | | AFQuery | bcftools | GATK GenomicsDB | Hail |
190
+ |---|---|---|---|---|
191
+ | Capture-aware AN | **Yes** | No | No | No |
192
+ | Metadata filtering | **Arbitrary labels** | No | No | Custom code |
193
+ | Ploidy-aware sex chromosomes | **Yes** | Manual | No | Manual |
194
+ | Dynamic subcohort queries | **Yes** | No | Limited | Requires code |
195
+ | FILTER / coverage tracking | **Per variant** | Manual | No | Manual |
196
+ | Incremental updates | **Yes** | No | **Yes** | No |
197
+ | Infrastructure required | **None** | **None** | Java/server | Spark cluster |
198
+
199
+ ### Benchmarks vs. bcftools (1000 Genomes Phase 3, n = 2,504, chr22)
200
+
201
+ | Workload | AFQuery | bcftools | Speedup |
202
+ |---|---|---|---|
203
+ | Full-chromosome AC/AN/AF export | ~7.0 s | ~3.8 min | **~33×** |
204
+ | AF concordance over 1,106,181 common variants | — | — | **R² > 0.99999** |
205
+
206
+ Point-query latency on AFQuery is ~14 ms and constant from 1K to 50K samples (median over 50 replicates, warm cache).
207
+
208
+ ## Documentation
209
+
210
+ - [Full documentation](https://babelomics.github.io/afquery/)
211
+ - [Getting started](https://babelomics.github.io/afquery/getting-started/)
212
+ - [Why local allele frequencies?](https://babelomics.github.io/afquery/getting-started/motivation/)
213
+ - [Manifest format](https://babelomics.github.io/afquery/guides/manifest-format/)
214
+ - [CLI reference](https://babelomics.github.io/afquery/reference/cli/)
215
+ - [FAQ](https://babelomics.github.io/afquery/faq/)
216
+
217
+ ## Citation
218
+
219
+ If you use AFQuery in your work, please cite:
220
+
221
+ > AFQuery: fast, capture-aware allele frequency queries on local genomic cohorts.
222
+ > *(manuscript in preparation)*
223
+
224
+ ## License
225
+
226
+ [MIT](LICENSE)
@@ -18,7 +18,7 @@ version_tuple: tuple[int | str, ...]
18
18
  commit_id: str | None
19
19
  __commit_id__: str | None
20
20
 
21
- __version__ = version = '0.3.1'
22
- __version_tuple__ = version_tuple = (0, 3, 1)
21
+ __version__ = version = '0.3.2'
22
+ __version_tuple__ = version_tuple = (0, 3, 2)
23
23
 
24
24
  __commit_id__ = commit_id = None
afquery-0.3.1/PKG-INFO DELETED
@@ -1,145 +0,0 @@
1
- Metadata-Version: 2.4
2
- Name: afquery
3
- Version: 0.3.1
4
- Summary: Genomic allele frequency query engine with bitmap-encoded genotypes
5
- License: MIT
6
- License-File: LICENSE
7
- Classifier: Programming Language :: Python :: 3
8
- Classifier: Programming Language :: Python :: 3.10
9
- Classifier: Programming Language :: Python :: 3.11
10
- Classifier: Programming Language :: Python :: 3.12
11
- Requires-Python: >=3.10
12
- Requires-Dist: click>=8.1
13
- Requires-Dist: cyvcf2>=0.30
14
- Requires-Dist: duckdb>=0.10
15
- Requires-Dist: pyarrow>=14.0
16
- Requires-Dist: pyranges<0.2,>=0.1.2
17
- Requires-Dist: pyroaring>=0.4.8
18
- Requires-Dist: tqdm>=4.60
19
- Provides-Extra: dev
20
- Requires-Dist: pytest-cov>=5.0; extra == 'dev'
21
- Requires-Dist: pytest>=8.0; extra == 'dev'
22
- Provides-Extra: docs
23
- Requires-Dist: matplotlib>=3.8; extra == 'docs'
24
- Requires-Dist: mike>=2.0; extra == 'docs'
25
- Requires-Dist: mkdocs-material>=9.5; extra == 'docs'
26
- Requires-Dist: mkdocs-print-site-plugin>=2.6.1; extra == 'docs'
27
- Requires-Dist: mkdocstrings[python]>=0.25; extra == 'docs'
28
- Requires-Dist: pymdown-extensions>=10.7; extra == 'docs'
29
- Description-Content-Type: text/markdown
30
-
31
- # AFQuery
32
-
33
- [![CI](https://github.com/babelomics/afquery/actions/workflows/ci.yml/badge.svg)](https://github.com/babelomics/afquery/actions/workflows/ci.yml)
34
- [![Coverage](https://codecov.io/gh/babelomics/afquery/graph/badge.svg)](https://codecov.io/gh/babelomics/afquery)
35
- [![Docs](https://img.shields.io/website?url=https%3A%2F%2Fbabelomics.github.io%2Fafquery%2F&label=docs)](https://babelomics.github.io/afquery/)
36
- <br>
37
- [![PyPI](https://img.shields.io/pypi/v/afquery.svg?color=blue)](https://pypi.org/project/afquery/)
38
- [![Bioconda](https://img.shields.io/conda/vn/bioconda/afquery.svg)](https://anaconda.org/bioconda/afquery)
39
- [![Docker](https://img.shields.io/badge/ghcr.io-afquery-blue?logo=docker)](https://github.com/babelomics/afquery/pkgs/container/afquery)
40
- [![Python](https://img.shields.io/pypi/pyversions/afquery.svg)](https://pypi.org/project/afquery/)
41
- [![License: MIT](https://img.shields.io/github/license/babelomics/afquery)](https://github.com/babelomics/afquery/blob/master/LICENSE)
42
-
43
- AFQuery enables fast allele frequency queries on user-defined subsets of local genomic cohorts, without rescanning VCFs.
44
-
45
- AFQuery is a bitmap-indexed engine that efficiently recomputes AC/AN/AF for dynamically defined subcohorts (e.g., by phenotype, sex, or sequencing technology), a common requirement in ACMG/AMP variant classification. It stores per-variant genotype data as Roaring Bitmaps in Parquet files and resolves sample filters into bitmaps that can be intersected in microseconds, enabling sub-100 ms queries on large cohorts. The system accounts for ploidy in sex chromosomes, adjusts AN based on sequencing technology, supports incremental updates, and runs locally using a file-based setup (Parquet + SQLite) without requiring server or cloud infrastructure.
46
-
47
- [Full Documentation→](https://babelomics.github.io/afquery/)
48
-
49
- ## When to use AFQuery
50
-
51
- - You need allele frequencies for phenotype or user-defined subcohorts
52
- - You work with mixed sequencing technologies or capture kits versions (WGS, WES, targeted panels)
53
- - You require fast, repeated queries without rescanning VCFs
54
- - You want a local, reproducible workflow without cloud or cluster dependencies
55
-
56
- ## Features
57
-
58
- - **Dynamic subcohort queries (<100 ms)** — bitmap intersections at query time; no VCF re-scan required
59
- - **Technology-aware** — avoids bias when mixing WGS, WES, and panels using different BED capture indexes
60
- - **Ploidy-aware** — correct handling of sex chromosomes (PAR/non-PAR, chrX, chrY)
61
- - **ACMG-compatible allele counting** — AC/AN/AF computed per standard definitions
62
- - **Flexible metadata filtering** — arbitrary labels (ICD-10, HPO, custom fields) with inclusion/exclusion rules
63
- - **Incremental updates** — add or remove samples and update metadata without rebuilding the database
64
- - **VCF annotation** — annotate variants using subcohort-specific frequencies
65
- - **FILTER/call quality tracking** — failed calls (FILTER!=PASS) tracked per variant and reported as N_FAIL
66
- - **Batch and region queries** — query a single locus, a genomic region, or a list of variants from a file
67
- - **Bulk CSV export** — export all variant frequencies with optional disaggregation by sex, technology, or phenotype
68
- - **Audit changelog** — all database operations logged with timestamps and operator notes
69
- - **Database validation** — integrity checks with scripted exit codes
70
- - **Portable and serverless** — file-based system, no infrastructure required
71
-
72
- ## Performance
73
-
74
- - Query latency: <100 ms (tested up to 50,000 samples)
75
- - Storage: ~2 bytes/sample/variant
76
- - Scales to millions of variants per chromosome
77
-
78
- ## Comparison with Alternative Tools
79
-
80
- | | AFQuery | bcftools | GATK GenomicsDB | Hail |
81
- |---|---|---|---|---|
82
- | Technology-aware AN | **Yes** | No | No | No |
83
- | Metadata filtering | **Arbitrary labels** | No | No | Custom code |
84
- | Ploidy-aware sex chromosomes | **Yes** | Manual | No | Manual |
85
- | Dynamic subcohort queries | **Yes** | No | Limited | Requires code |
86
- | FILTER/call quality tracking | **Per variant** | Manual | No | Manual |
87
- | Incremental updates | **Yes** | No | **Yes** | No |
88
- | Infrastructure required | **None** | **None** | Java/server | Spark cluster |
89
- | Query latency (50K samples) | **<100 ms** | ~5 min | <1 min | 1–2 min |
90
-
91
- ## Algorithm Overview
92
-
93
- AFQuery pre-indexes per-variant genotype data as [Roaring Bitmaps](https://roaringbitmap.org/) stored in Parquet files. Each variant row holds three bitmaps: heterozygous carriers, homozygous alt carriers, and samples with FILTER!=PASS. Sample metadata (sex, phenotype, technology) is pre-serialized as bitmaps in SQLite.
94
-
95
- At query time, the requested sample filter is resolved to a single candidate bitmap via bitmap intersections and differences — taking microseconds regardless of cohort size. For each variant, the candidate bitmap is intersected with the genotype bitmaps to compute AC/AN/AF. AN accounts for WES capture regions (via BED-indexed interval trees) and for ploidy on sex chromosomes (males are haploid on non-PAR chrX and chrY).
96
-
97
- ## Input Requirements
98
-
99
- - VCF files: normalized and consistent with the selected genome build (GRCh37 or GRCh38)
100
- - Sample metadata: must include sex, sequencing technology, and any fields used for filtering (e.g., phenotype)
101
- - BED files (optional): define capture regions for each sequencing technology
102
-
103
- ## Quick Start
104
-
105
- Example workflow from raw VCFs to query, export, and annotation:
106
-
107
- ```bash
108
- pip install afquery
109
- # Docker: see Installation docs for docker pull / run usage
110
-
111
- # Build the database
112
- afquery create-db --manifest samples.tsv --output-dir ./db/ --genome-build GRCh38
113
-
114
- # Inspect the database
115
- afquery info --db ./db/
116
-
117
- # Query a single position, filtered to a phenotype
118
- afquery query --db ./db/ --locus chr1:925952 --phenotype E11.9 --sex female
119
-
120
- # Query a genomic region
121
- afquery query --db ./db/ --region chr1:900000-1000000
122
-
123
- # Export BRCA1 variant frequencies to CSV
124
- afquery dump --db ./db/ --output all_variants.csv --chrom chr17 --start 43044292 --end 43170327
125
-
126
- # Annotate a VCF with cohort frequencies
127
- afquery annotate --db ./db/ --input patient.vcf --output annotated.vcf --threads 12
128
-
129
- # Add new samples to an existing database
130
- afquery update-db --db ./db/ --add-samples new_samples.tsv
131
- ```
132
-
133
- ## Documentation
134
-
135
- - [Full Documentation](https://babelomics.github.io/afquery/)
136
- - [Getting Started](https://babelomics.github.io/afquery/getting-started/)
137
- - [Why local allele frequencies?](https://babelomics.github.io/afquery/getting-started/motivation/)
138
- - [CLI Reference](https://babelomics.github.io/afquery/reference/cli/)
139
-
140
- ## Citation
141
-
142
- If you use AFQuery, please cite:
143
-
144
- > AFQuery: fast, metadata-aware allele frequency queries on local genomic cohorts.
145
- > *(manuscript in preparation)*