blast-align-tree 1.0.0__tar.gz

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -0,0 +1,21 @@
1
+ MIT License
2
+
3
+ Copyright (c) 2026 Adam Steinbrenner
4
+
5
+ Permission is hereby granted, free of charge, to any person obtaining a copy
6
+ of this software and associated documentation files (the "Software"), to deal
7
+ in the Software without restriction, including without limitation the rights
8
+ to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
9
+ copies of the Software, and to permit persons to whom the Software is
10
+ furnished to do so, subject to the following conditions:
11
+
12
+ The above copyright notice and this permission notice shall be included in all
13
+ copies or substantial portions of the Software.
14
+
15
+ THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
16
+ IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
17
+ FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
18
+ AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
19
+ LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
20
+ OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
21
+ SOFTWARE.
@@ -0,0 +1,433 @@
1
+ Metadata-Version: 2.4
2
+ Name: blast-align-tree
3
+ Version: 1.0.0
4
+ Summary: BLAST, align, and build phylogenetic trees from genomic databases
5
+ Author-email: Adam Steinbrenner <astein10@uw.edu>
6
+ License-Expression: MIT
7
+ Project-URL: Homepage, https://github.com/steinbrennerlab/blast-align-tree
8
+ Project-URL: Issues, https://github.com/steinbrennerlab/blast-align-tree/issues
9
+ Classifier: Development Status :: 4 - Beta
10
+ Classifier: Intended Audience :: Science/Research
11
+ Classifier: Topic :: Scientific/Engineering :: Bio-Informatics
12
+ Classifier: Programming Language :: Python :: 3
13
+ Classifier: Programming Language :: Python :: 3 :: Only
14
+ Classifier: Operating System :: OS Independent
15
+ Requires-Python: >=3.9
16
+ Description-Content-Type: text/markdown
17
+ License-File: LICENSE
18
+ Requires-Dist: biopython
19
+ Dynamic: license-file
20
+
21
+ # blast-align-tree
22
+
23
+ A pipeline to identify BLAST hits and perform phylogenetic analysis across
24
+ multiple queries and local genome databases.
25
+
26
+ [![DOI](https://zenodo.org/badge/374224275.svg)](https://zenodo.org/doi/10.5281/zenodo.10888646)
27
+
28
+ ## Introduction
29
+
30
+ A common task in bioinformatics is to find similar genes across a set of
31
+ genomes and compare them using phylogenetic methods. As an alternative to
32
+ using online tools for such analyses, researchers may wish to download
33
+ genomes of interest for local BLAST and downstream analyses. Homolog
34
+ curation, tree construction, header parsing, and visualization alongside
35
+ other datasets (e.g. gene expression) can give quick insights into a gene
36
+ family of interest.
37
+
38
+ <!-- TODO: regenerate images/flowchart.png if the pipeline layout has changed -->
39
+ ![](images/flowchart.png)
40
+
41
+ ## Installation
42
+
43
+ ### Python package
44
+
45
+ ```
46
+ pip install blast-align-tree
47
+ ```
48
+
49
+ This installs three console commands:
50
+
51
+ | Command | Purpose |
52
+ |---|---|
53
+ | `blast-align-tree` | Run the pipeline (BLAST → align → tree → visualize) |
54
+ | `bat-genome-selector` | Tkinter GUI for building `blast-align-tree` commands |
55
+ | `blast-align-tree-fetch` | Download bundled genome FASTAs into `./genomes/` |
56
+
57
+ From a fresh clone of this repo you can do an editable install instead:
58
+
59
+ ```
60
+ pip install -e .
61
+ ```
62
+
63
+ ### External tools (must be on PATH)
64
+
65
+ `pip` cannot install these — use conda/mamba or your system package
66
+ manager.
67
+
68
+ - **BLAST+** (`makeblastdb`, `blastdbcmd`, `tblastn`, `blastp`, `psiblast`)
69
+ - **MAFFT** (≥ 7) — default aligner
70
+ - **Clustal Omega** (`clustalo`) — alternative aligner
71
+ - **trimAl** — alignment cleanup
72
+ - **FastTree** and/or **RAxML-NG** — tree inference
73
+ - **R** with `ggtree`, `ape`, `phytools`, `ggplot2`, `optparse`, `treeio`,
74
+ `tidytree`, `broom`
75
+ - **HMMER** (`hmmscan`, `hmmpress`) — optional, only needed for `--hmm`
76
+
77
+ Conda YAML files covering all of the above live under `environments/`.
78
+ Pick the one for your platform and create an env:
79
+
80
+ **conda (Linux):**
81
+
82
+ ```
83
+ conda env create -f environments/bat-environment-linux.yml
84
+ conda activate bat
85
+ pip install blast-align-tree
86
+ ```
87
+
88
+ **mamba / micromamba (Linux — faster solver):**
89
+
90
+ ```
91
+ mamba env create -f environments/bat-environment-linux.yml
92
+ mamba activate bat
93
+ pip install blast-align-tree
94
+ ```
95
+
96
+ **macOS (Apple Silicon or Intel):**
97
+
98
+ ```
99
+ mamba env create -f environments/bat-environment-ARMorIntel-mac.yml
100
+ mamba activate bat
101
+ pip install blast-align-tree
102
+ ```
103
+
104
+ **Windows:**
105
+
106
+ ```
107
+ mamba env create -f environments/bat-environment-windows.yml
108
+ mamba activate bat
109
+ pip install blast-align-tree
110
+ ```
111
+
112
+ Verify everything is wired up:
113
+
114
+ ```
115
+ blast-align-tree --check-env
116
+ ```
117
+
118
+ ### HMM profiles need `hmmpress`
119
+
120
+ The repo ships `.hmm` input files under `hmm_files/` (e.g.
121
+ `hmm_files/kinase.hmm`) but **not** the `.h3*` binary indices — those are
122
+ build artifacts and are deleted from the repo. Before using `--hmm`, build
123
+ the index locally:
124
+
125
+ ```
126
+ hmmpress hmm_files/kinase.hmm
127
+ ```
128
+
129
+ This produces `kinase.hmm.h3f`, `.h3i`, `.h3m`, `.h3p` alongside the input.
130
+
131
+ ### Fetching genome databases
132
+
133
+ Genome FASTAs are too large to bundle in the pip package. The downloader
134
+ reads a manifest and drops the selected genomes into `./genomes/`:
135
+
136
+ ```
137
+ blast-align-tree-fetch # default set 🌱 🌿 🫘
138
+ blast-align-tree-fetch --all # everything listed in the manifest 🌍
139
+ blast-align-tree-fetch --list # show available genomes + sizes
140
+ ```
141
+
142
+ The default set 🌱🌿🫘🍅 is:
143
+
144
+ - 🌱 **TAIR10 CDS** — *Arabidopsis thaliana* coding sequences
145
+ - 🌿 **TAIR10 proteins** — *Arabidopsis thaliana* proteome
146
+ - 🫘 **Pvul218 CDS** — *Phaseolus vulgaris* (common bean) coding sequences
147
+ - 🍅 **Niben261 proteins** — *Nicotiana benthamiana* proteome (v2.6.1)
148
+
149
+ `--all` 🌍 additionally pulls the rest of the hosted lineup:
150
+
151
+ - Other plants: 🫛 cowpea CDS
152
+ - Animals / fungus (fetched into `./genomes/animals/`):
153
+ 🧑 human CDS, 🐭 mouse CDS, 🐀 rat CDS, 🐒 chimp CDS,
154
+ 🐟 zebrafish CDS, 🪰 fruit fly CDS, 🪱 *C. elegans* CDS,
155
+ 🍞 yeast ORFs (*S. cerevisiae* S288C)
156
+
157
+ Run `blast-align-tree-fetch --list` for the full lineup and sizes.
158
+
159
+ All pipeline runs should happen in the directory that contains
160
+ `./genomes/`.
161
+
162
+ ## GUI: `bat-genome-selector`
163
+
164
+ The easiest way to build a valid `blast-align-tree` invocation is the
165
+ Tkinter GUI. Launch it from a directory that contains a `genomes/` folder:
166
+
167
+ ```
168
+ bat-genome-selector
169
+ ```
170
+
171
+ <!-- TODO: add screenshot images/gui.png once the GUI is regenerated -->
172
+
173
+ ### Key features
174
+
175
+ - **Auto-discovery.** Scans `./genomes/` (recursively) for `.fa`, `.faa`,
176
+ `.fas`, `.fasta`, `.fna` files and ignores BLAST index sidecars.
177
+ - **Header auto-detection.** Peeks at the first FASTA record in each
178
+ database and suggests a plausible `-hdr` token (e.g. `gene:`, `locus=`,
179
+ `polypeptide=`), with a live "Parsed name" preview so you can see
180
+ exactly what will appear on the tree.
181
+ - **Per-row controls.** One row per genome: include checkbox, query
182
+ column, `-hdr`, `-hdr_sfx`, `-n` (hits to keep), nucleotide/protein
183
+ type, and a "build DB" shortcut that runs `makeblastdb` when the BLAST
184
+ indices are missing.
185
+ - **Bulk actions.** *Select All Hits*, *Deselect All Hits*, *Clear
186
+ Fields*, *Refresh*, plus a **Default -n** spinbox and **Set All -n**
187
+ button.
188
+ - **Options panel.** Aligner (Clustal Omega or MAFFT + mode), tree builder
189
+ (FastTree or RAxML), BLAST type (tblastn/blastp), thread count.
190
+ - **Advanced panel** (collapsible): outgroups (`-add`, `-add_db`), AA
191
+ slice (`-aa`), motif patterns (regex or PROSITE, overlap toggle), and
192
+ HMM profiles (`--hmm`).
193
+ - **Generate Command / Copy to Clipboard.** Produces a ready-to-paste
194
+ `blast-align-tree …` command.
195
+ - **Recent Runs tab.** Lists past `ENTRY/runs/TIMESTAMP/` directories in
196
+ the current working directory so you can quickly jump back to prior
197
+ results.
198
+
199
+ ## Example output
200
+
201
+ The default genome set includes Arabidopsis TAIR10 CDS. The example below
202
+ runs the pipeline for a SERK query and redraws the resulting tree with a
203
+ new subnode/outgroup:
204
+
205
+ ```
206
+ blast-align-tree -q AT4G33430.1 -qdbs TAIR10cds.fa \
207
+ -n 15 15 15 \
208
+ -dbs TAIR10cds.fa Pvul218cds.fa Vung469cds.fa \
209
+ -hdr gene: polypeptide= locus=
210
+ ```
211
+
212
+ The pipeline creates a folder `AT4G33430.1/` in your working directory.
213
+ The timestamped run root keeps the tree PDFs. Newick tree files, gene
214
+ lists, alignment FASTAs, mappings, features, BLAST hit FASTAs, and
215
+ per-genome summaries go under
216
+ `AT4G33430.1/runs/<TIMESTAMP>/genes_alignments_trees/`.
217
+
218
+ After the run finishes, the pipeline prints a re-draw hint. For example,
219
+ to reroot on an outgroup (`-a AT5G10290`) and zoom in on a subnode
220
+ (`-n 45`):
221
+
222
+ ```
223
+ Rscript "<bundled-visualize_tree.r>" -e AT4G33430.1 -b SERK_tree \
224
+ --subdir "runs/<TIMESTAMP>" -a AT5G10290 -n 45
225
+ ```
226
+
227
+ The bundled path is printed for you at the end of each pipeline run,
228
+ already wrapped in double quotes — keep the quotes when copy-pasting,
229
+ especially on Windows, where unquoted paths can cause `Rscript` to
230
+ segfault if they contain spaces or backslashes that the shell misparses.
231
+
232
+ <!-- TODO: regenerate images/tree.png from the SERK run above -->
233
+ ![](images/tree.png)
234
+
235
+ ## Tutorial
236
+
237
+ ### Run blast-align-tree for ACC Oxidase
238
+
239
+ Find 15 homologs of Arabidopsis ACC Oxidase 1 from three plant genomes
240
+ using `tblastn` against complete CDS databases. `-q` specifies the query
241
+ locus, `-qdbs` the database it lives in, `-dbs` the databases to search,
242
+ and `-hdr` the regex tokens used to parse gene names out of each
243
+ database's FASTA headers.
244
+
245
+ ```
246
+ blast-align-tree -q AT2G19590.1 -qdbs TAIR10cds.fa \
247
+ -n 15 15 15 \
248
+ -dbs TAIR10cds.fa Pvul218cds.fa Vung469cds.fa \
249
+ -hdr gene: polypeptide= locus=
250
+ ```
251
+
252
+ This creates `AT2G19590.1/` with tree PDFs at the timestamped run root and
253
+ Newick tree files, alignment files, BLAST hit FASTAs, and per-genome
254
+ summaries under `genes_alignments_trees/`.
255
+
256
+ A powerful feature of `ggtree` is the ability to plot associated data.
257
+ Each run produces two complementary tree PDFs:
258
+
259
+ 1. **Text version** — gene symbols and dataset values printed as labels
260
+ next to each tip.
261
+ 2. **Heatmap version** — the same tree with associated data rendered as
262
+ a coloured heatmap alongside the tips.
263
+
264
+ By default, both include expression data from the [Klepikova *Arabidopsis*
265
+ expression atlas](https://pubmed.ncbi.nlm.nih.gov/26923014/) (headers are
266
+ matched to the AtGenExpress / eFP browser tissue naming). The screenshot
267
+ below shows the heatmap version:
268
+
269
+ <!-- TODO: regenerate images/ACO-tree-1.png (heatmap version) from the ACO run above -->
270
+ ![](images/ACO-tree-1.png)
271
+
272
+ A separate PDF with `.MSA.pdf` appended shows a cartoon alignment — useful
273
+ for spotting large differences in domain architecture. Open the
274
+ underlying FASTA files in `genes_alignments_trees/` to inspect the
275
+ alignment in detail.
276
+
277
+ <!-- TODO: regenerate images/ACO-tree-2.png (.MSA.pdf cartoon alignment) -->
278
+ ![](images/ACO-tree-2.png)
279
+
280
+ ### Redraw the ACC Oxidase tree
281
+
282
+ You can re-run `visualize_tree.r` at any time to produce new PDFs. The
283
+ pipeline prints a ready-to-edit `Rscript …` command at the end of each
284
+ run; copy it and tweak options such as:
285
+
286
+ - `-b <NAME>` — filename stem for the new PDFs
287
+ - `-a <ID>` — reroot on this outgroup
288
+ - `-n <NODE>` — draw a subtree at this node (use `--help` for the full
289
+ option list)
290
+ - `-k 1` — show bootstraps
291
+ - `-l 0` — hide node number labels
292
+ - `-m 2` — enlarge gene-symbol text
293
+
294
+ For example, reroot the default ACO tree on JRG21 (AT2G38240) and zoom
295
+ into the ACO clade at node 58:
296
+
297
+ ```
298
+ Rscript "<bundled-visualize_tree.r>" -e AT2G19590.1 -b ACO_v3 \
299
+ --subdir "runs/<TIMESTAMP>" -a AT2G38240 -n 58 -k 1 -l 0 -m 2
300
+ ```
301
+
302
+ <!-- TODO: regenerate images/ACO-tree-3.png -->
303
+ ![](images/ACO-tree-3.png)
304
+
305
+ The `-n` option is especially helpful for extracting a subset of the tree
306
+ as a FASTA. Sequences are listed in the FASTA in the same order as the
307
+ tree, and `trimAl` is used to strip blank-only alignment columns — useful
308
+ for a quick view of conserved residues (e.g. the ACO active site) in a
309
+ viewer like AliView.
310
+
311
+ ### BLASTP instead of TBLASTN
312
+
313
+ Use `--blast_type blastp` against protein databases. The example below
314
+ pulls 10 NIMIN-1 homologs from the Arabidopsis and *Nicotiana
315
+ benthamiana* proteomes:
316
+
317
+ ```
318
+ blast-align-tree --blast_type blastp \
319
+ -q AT1G02450.1 -qdbs TAIR10protein.fa \
320
+ -n 10 10 \
321
+ -dbs TAIR10protein.fa Niben261_genome.annotation.proteins.fasta \
322
+ -hdr gene: id
323
+ ```
324
+
325
+ ### Multiple queries
326
+
327
+ You can pass several query sequences with `-q`; the pipeline extracts
328
+ each from the database listed at the matching position in `-qdbs`,
329
+ de-duplicates, then searches each `-dbs` entry. The example below uses
330
+ three queries drawn from two databases and searches two other databases:
331
+
332
+ ```
333
+ blast-align-tree -q AT5G45250.1 Phvul.007G077500.1 AT5G17890.1 \
334
+ -qdbs TAIR10cds.fa Pvul218cds.fa TAIR10cds.fa \
335
+ -n 3 4 \
336
+ -dbs TAIR10cds.fa Vung469cds.fa \
337
+ -hdr gene: locus=
338
+ ```
339
+
340
+ ### Adding a new genome
341
+
342
+ You can drop additional genomes into `./genomes/` and use them alongside
343
+ the bundled ones. For each new genome you need a local BLAST database.
344
+
345
+ For CDS files:
346
+
347
+ ```
348
+ makeblastdb -in GenomeCDS.fa -parse_seqids -dbtype nucl
349
+ ```
350
+
351
+ For protein files:
352
+
353
+ ```
354
+ makeblastdb -in GenomeProteins.fa -parse_seqids -dbtype prot
355
+ ```
356
+
357
+ #### Example: add *Nicotiana tabacum* and rebuild an earlier tree
358
+
359
+ Download an annotated tobacco proteome into `./genomes/`:
360
+
361
+ - [*N. tabacum* v4.5 from Sol Genomics](https://solgenomics.net/ftp/ftp/genomes/Nicotiana_tabacum/edwards_et_al_2017/annotation/)
362
+ → `Nitab-v4.5_proteins_Edwards2017.fasta`
363
+
364
+ Build the BLAST database:
365
+
366
+ ```
367
+ cd genomes
368
+ makeblastdb -in Nitab-v4.5_proteins_Edwards2017.fasta -parse_seqids -dbtype prot
369
+ cd ..
370
+ ```
371
+
372
+ Inspecting the first record of `Nitab-v4.5_proteins_Edwards2017.fasta`
373
+ shows that headers use a gene id followed by a description — `-hdr id`
374
+ keeps just the first token.
375
+
376
+ Now build a SOBIR1 homolog tree across Arabidopsis, *N. benthamiana*,
377
+ and the freshly added tobacco proteome:
378
+
379
+ ```
380
+ blast-align-tree --blast_type blastp \
381
+ -q AT2G31880.1 -qdbs TAIR10protein.fa \
382
+ -n 10 10 10 \
383
+ -dbs TAIR10protein.fa \
384
+ Niben261_genome.annotation.proteins.fasta \
385
+ Nitab-v4.5_proteins_Edwards2017.fasta \
386
+ -hdr gene: id id
387
+ ```
388
+
389
+ The run produces a tree PDF with tobacco SOBIR1 homologs slotted in
390
+ alongside the *N. benthamiana* and Arabidopsis sequences.
391
+
392
+ <!-- TODO: regenerate images/SOBIR1_with_ntab.png from the run above -->
393
+ ![](images/SOBIR1_with_ntab.png)
394
+
395
+ ### Rebuild the SOBIR1 tree with MAFFT and RAxML
396
+
397
+ By default the pipeline aligns with Clustal Omega and infers the tree
398
+ with FastTree, but both are swappable. The command below rebuilds the
399
+ same SOBIR1 tree (Arabidopsis + *N. benthamiana* + tobacco) using
400
+ **MAFFT** in `linsi` mode and **RAxML-NG** for the tree inference:
401
+
402
+ ```
403
+ blast-align-tree --blast_type blastp \
404
+ --aligner mafft --mafft_mode linsi \
405
+ --tree_builder RAxML \
406
+ -q AT2G31880.1 -qdbs TAIR10protein.fa \
407
+ -n 10 10 10 \
408
+ -dbs TAIR10protein.fa \
409
+ Niben261_genome.annotation.proteins.fasta \
410
+ Nitab-v4.5_proteins_Edwards2017.fasta \
411
+ -hdr gene: id id
412
+ ```
413
+
414
+ RAxML-NG is noticeably slower than FastTree but provides maximum-
415
+ likelihood branch support via bootstrapping. Comparing the two trees is
416
+ a quick sanity check that any clades you care about are stable across
417
+ inference methods.
418
+
419
+ Use repeated rounds of querying to refine your trees, search different
420
+ genome versions, and compare aligners / tree builders before drawing
421
+ strong conclusions.
422
+
423
+ ## Future features
424
+
425
+ We are currently working on:
426
+
427
+ 1. Iterative BLAST using a first set of hits as secondary queries
428
+ 2. Better organization of query and sub-query folders
429
+ 3. Displaying a subsequence of the MSA in the alignment PDF
430
+ 4. Richer motif / HMM visualization overlaid on the MSA
431
+
432
+ If you'd like to contribute, reach out to Ben and Adam:
433
+ `bdshep@uw.edu` and `astein10@uw.edu`.