blast-align-tree 1.0.0__tar.gz
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- blast_align_tree-1.0.0/LICENSE +21 -0
- blast_align_tree-1.0.0/PKG-INFO +433 -0
- blast_align_tree-1.0.0/README.md +413 -0
- blast_align_tree-1.0.0/blast_align_tree/__init__.py +2 -0
- blast_align_tree-1.0.0/blast_align_tree/cli.py +1463 -0
- blast_align_tree-1.0.0/blast_align_tree/data/gene_symbols.txt +9500 -0
- blast_align_tree-1.0.0/blast_align_tree/data/genomes_manifest.json +123 -0
- blast_align_tree-1.0.0/blast_align_tree/data/scripts/extract_seq.py +71 -0
- blast_align_tree-1.0.0/blast_align_tree/data/scripts/remove_header.py +15 -0
- blast_align_tree-1.0.0/blast_align_tree/data/visualize_tree.r +1058 -0
- blast_align_tree-1.0.0/blast_align_tree/fetch.py +148 -0
- blast_align_tree-1.0.0/blast_align_tree/genome_selector.py +1074 -0
- blast_align_tree-1.0.0/blast_align_tree.egg-info/PKG-INFO +433 -0
- blast_align_tree-1.0.0/blast_align_tree.egg-info/SOURCES.txt +18 -0
- blast_align_tree-1.0.0/blast_align_tree.egg-info/dependency_links.txt +1 -0
- blast_align_tree-1.0.0/blast_align_tree.egg-info/entry_points.txt +4 -0
- blast_align_tree-1.0.0/blast_align_tree.egg-info/requires.txt +1 -0
- blast_align_tree-1.0.0/blast_align_tree.egg-info/top_level.txt +1 -0
- blast_align_tree-1.0.0/pyproject.toml +42 -0
- blast_align_tree-1.0.0/setup.cfg +4 -0
|
@@ -0,0 +1,21 @@
|
|
|
1
|
+
MIT License
|
|
2
|
+
|
|
3
|
+
Copyright (c) 2026 Adam Steinbrenner
|
|
4
|
+
|
|
5
|
+
Permission is hereby granted, free of charge, to any person obtaining a copy
|
|
6
|
+
of this software and associated documentation files (the "Software"), to deal
|
|
7
|
+
in the Software without restriction, including without limitation the rights
|
|
8
|
+
to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
|
|
9
|
+
copies of the Software, and to permit persons to whom the Software is
|
|
10
|
+
furnished to do so, subject to the following conditions:
|
|
11
|
+
|
|
12
|
+
The above copyright notice and this permission notice shall be included in all
|
|
13
|
+
copies or substantial portions of the Software.
|
|
14
|
+
|
|
15
|
+
THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
|
|
16
|
+
IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
|
|
17
|
+
FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
|
|
18
|
+
AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
|
|
19
|
+
LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
|
|
20
|
+
OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
|
|
21
|
+
SOFTWARE.
|
|
@@ -0,0 +1,433 @@
|
|
|
1
|
+
Metadata-Version: 2.4
|
|
2
|
+
Name: blast-align-tree
|
|
3
|
+
Version: 1.0.0
|
|
4
|
+
Summary: BLAST, align, and build phylogenetic trees from genomic databases
|
|
5
|
+
Author-email: Adam Steinbrenner <astein10@uw.edu>
|
|
6
|
+
License-Expression: MIT
|
|
7
|
+
Project-URL: Homepage, https://github.com/steinbrennerlab/blast-align-tree
|
|
8
|
+
Project-URL: Issues, https://github.com/steinbrennerlab/blast-align-tree/issues
|
|
9
|
+
Classifier: Development Status :: 4 - Beta
|
|
10
|
+
Classifier: Intended Audience :: Science/Research
|
|
11
|
+
Classifier: Topic :: Scientific/Engineering :: Bio-Informatics
|
|
12
|
+
Classifier: Programming Language :: Python :: 3
|
|
13
|
+
Classifier: Programming Language :: Python :: 3 :: Only
|
|
14
|
+
Classifier: Operating System :: OS Independent
|
|
15
|
+
Requires-Python: >=3.9
|
|
16
|
+
Description-Content-Type: text/markdown
|
|
17
|
+
License-File: LICENSE
|
|
18
|
+
Requires-Dist: biopython
|
|
19
|
+
Dynamic: license-file
|
|
20
|
+
|
|
21
|
+
# blast-align-tree
|
|
22
|
+
|
|
23
|
+
A pipeline to identify BLAST hits and perform phylogenetic analysis across
|
|
24
|
+
multiple queries and local genome databases.
|
|
25
|
+
|
|
26
|
+
[](https://zenodo.org/doi/10.5281/zenodo.10888646)
|
|
27
|
+
|
|
28
|
+
## Introduction
|
|
29
|
+
|
|
30
|
+
A common task in bioinformatics is to find similar genes across a set of
|
|
31
|
+
genomes and compare them using phylogenetic methods. As an alternative to
|
|
32
|
+
using online tools for such analyses, researchers may wish to download
|
|
33
|
+
genomes of interest for local BLAST and downstream analyses. Homolog
|
|
34
|
+
curation, tree construction, header parsing, and visualization alongside
|
|
35
|
+
other datasets (e.g. gene expression) can give quick insights into a gene
|
|
36
|
+
family of interest.
|
|
37
|
+
|
|
38
|
+
<!-- TODO: regenerate images/flowchart.png if the pipeline layout has changed -->
|
|
39
|
+

|
|
40
|
+
|
|
41
|
+
## Installation
|
|
42
|
+
|
|
43
|
+
### Python package
|
|
44
|
+
|
|
45
|
+
```
|
|
46
|
+
pip install blast-align-tree
|
|
47
|
+
```
|
|
48
|
+
|
|
49
|
+
This installs three console commands:
|
|
50
|
+
|
|
51
|
+
| Command | Purpose |
|
|
52
|
+
|---|---|
|
|
53
|
+
| `blast-align-tree` | Run the pipeline (BLAST → align → tree → visualize) |
|
|
54
|
+
| `bat-genome-selector` | Tkinter GUI for building `blast-align-tree` commands |
|
|
55
|
+
| `blast-align-tree-fetch` | Download bundled genome FASTAs into `./genomes/` |
|
|
56
|
+
|
|
57
|
+
From a fresh clone of this repo you can do an editable install instead:
|
|
58
|
+
|
|
59
|
+
```
|
|
60
|
+
pip install -e .
|
|
61
|
+
```
|
|
62
|
+
|
|
63
|
+
### External tools (must be on PATH)
|
|
64
|
+
|
|
65
|
+
`pip` cannot install these — use conda/mamba or your system package
|
|
66
|
+
manager.
|
|
67
|
+
|
|
68
|
+
- **BLAST+** (`makeblastdb`, `blastdbcmd`, `tblastn`, `blastp`, `psiblast`)
|
|
69
|
+
- **MAFFT** (≥ 7) — default aligner
|
|
70
|
+
- **Clustal Omega** (`clustalo`) — alternative aligner
|
|
71
|
+
- **trimAl** — alignment cleanup
|
|
72
|
+
- **FastTree** and/or **RAxML-NG** — tree inference
|
|
73
|
+
- **R** with `ggtree`, `ape`, `phytools`, `ggplot2`, `optparse`, `treeio`,
|
|
74
|
+
`tidytree`, `broom`
|
|
75
|
+
- **HMMER** (`hmmscan`, `hmmpress`) — optional, only needed for `--hmm`
|
|
76
|
+
|
|
77
|
+
Conda YAML files covering all of the above live under `environments/`.
|
|
78
|
+
Pick the one for your platform and create an env:
|
|
79
|
+
|
|
80
|
+
**conda (Linux):**
|
|
81
|
+
|
|
82
|
+
```
|
|
83
|
+
conda env create -f environments/bat-environment-linux.yml
|
|
84
|
+
conda activate bat
|
|
85
|
+
pip install blast-align-tree
|
|
86
|
+
```
|
|
87
|
+
|
|
88
|
+
**mamba / micromamba (Linux — faster solver):**
|
|
89
|
+
|
|
90
|
+
```
|
|
91
|
+
mamba env create -f environments/bat-environment-linux.yml
|
|
92
|
+
mamba activate bat
|
|
93
|
+
pip install blast-align-tree
|
|
94
|
+
```
|
|
95
|
+
|
|
96
|
+
**macOS (Apple Silicon or Intel):**
|
|
97
|
+
|
|
98
|
+
```
|
|
99
|
+
mamba env create -f environments/bat-environment-ARMorIntel-mac.yml
|
|
100
|
+
mamba activate bat
|
|
101
|
+
pip install blast-align-tree
|
|
102
|
+
```
|
|
103
|
+
|
|
104
|
+
**Windows:**
|
|
105
|
+
|
|
106
|
+
```
|
|
107
|
+
mamba env create -f environments/bat-environment-windows.yml
|
|
108
|
+
mamba activate bat
|
|
109
|
+
pip install blast-align-tree
|
|
110
|
+
```
|
|
111
|
+
|
|
112
|
+
Verify everything is wired up:
|
|
113
|
+
|
|
114
|
+
```
|
|
115
|
+
blast-align-tree --check-env
|
|
116
|
+
```
|
|
117
|
+
|
|
118
|
+
### HMM profiles need `hmmpress`
|
|
119
|
+
|
|
120
|
+
The repo ships `.hmm` input files under `hmm_files/` (e.g.
|
|
121
|
+
`hmm_files/kinase.hmm`) but **not** the `.h3*` binary indices — those are
|
|
122
|
+
build artifacts and are deleted from the repo. Before using `--hmm`, build
|
|
123
|
+
the index locally:
|
|
124
|
+
|
|
125
|
+
```
|
|
126
|
+
hmmpress hmm_files/kinase.hmm
|
|
127
|
+
```
|
|
128
|
+
|
|
129
|
+
This produces `kinase.hmm.h3f`, `.h3i`, `.h3m`, `.h3p` alongside the input.
|
|
130
|
+
|
|
131
|
+
### Fetching genome databases
|
|
132
|
+
|
|
133
|
+
Genome FASTAs are too large to bundle in the pip package. The downloader
|
|
134
|
+
reads a manifest and drops the selected genomes into `./genomes/`:
|
|
135
|
+
|
|
136
|
+
```
|
|
137
|
+
blast-align-tree-fetch # default set 🌱 🌿 🫘
|
|
138
|
+
blast-align-tree-fetch --all # everything listed in the manifest 🌍
|
|
139
|
+
blast-align-tree-fetch --list # show available genomes + sizes
|
|
140
|
+
```
|
|
141
|
+
|
|
142
|
+
The default set 🌱🌿🫘🍅 is:
|
|
143
|
+
|
|
144
|
+
- 🌱 **TAIR10 CDS** — *Arabidopsis thaliana* coding sequences
|
|
145
|
+
- 🌿 **TAIR10 proteins** — *Arabidopsis thaliana* proteome
|
|
146
|
+
- 🫘 **Pvul218 CDS** — *Phaseolus vulgaris* (common bean) coding sequences
|
|
147
|
+
- 🍅 **Niben261 proteins** — *Nicotiana benthamiana* proteome (v2.6.1)
|
|
148
|
+
|
|
149
|
+
`--all` 🌍 additionally pulls the rest of the hosted lineup:
|
|
150
|
+
|
|
151
|
+
- Other plants: 🫛 cowpea CDS
|
|
152
|
+
- Animals / fungus (fetched into `./genomes/animals/`):
|
|
153
|
+
🧑 human CDS, 🐭 mouse CDS, 🐀 rat CDS, 🐒 chimp CDS,
|
|
154
|
+
🐟 zebrafish CDS, 🪰 fruit fly CDS, 🪱 *C. elegans* CDS,
|
|
155
|
+
🍞 yeast ORFs (*S. cerevisiae* S288C)
|
|
156
|
+
|
|
157
|
+
Run `blast-align-tree-fetch --list` for the full lineup and sizes.
|
|
158
|
+
|
|
159
|
+
All pipeline runs should happen in the directory that contains
|
|
160
|
+
`./genomes/`.
|
|
161
|
+
|
|
162
|
+
## GUI: `bat-genome-selector`
|
|
163
|
+
|
|
164
|
+
The easiest way to build a valid `blast-align-tree` invocation is the
|
|
165
|
+
Tkinter GUI. Launch it from a directory that contains a `genomes/` folder:
|
|
166
|
+
|
|
167
|
+
```
|
|
168
|
+
bat-genome-selector
|
|
169
|
+
```
|
|
170
|
+
|
|
171
|
+
<!-- TODO: add screenshot images/gui.png once the GUI is regenerated -->
|
|
172
|
+
|
|
173
|
+
### Key features
|
|
174
|
+
|
|
175
|
+
- **Auto-discovery.** Scans `./genomes/` (recursively) for `.fa`, `.faa`,
|
|
176
|
+
`.fas`, `.fasta`, `.fna` files and ignores BLAST index sidecars.
|
|
177
|
+
- **Header auto-detection.** Peeks at the first FASTA record in each
|
|
178
|
+
database and suggests a plausible `-hdr` token (e.g. `gene:`, `locus=`,
|
|
179
|
+
`polypeptide=`), with a live "Parsed name" preview so you can see
|
|
180
|
+
exactly what will appear on the tree.
|
|
181
|
+
- **Per-row controls.** One row per genome: include checkbox, query
|
|
182
|
+
column, `-hdr`, `-hdr_sfx`, `-n` (hits to keep), nucleotide/protein
|
|
183
|
+
type, and a "build DB" shortcut that runs `makeblastdb` when the BLAST
|
|
184
|
+
indices are missing.
|
|
185
|
+
- **Bulk actions.** *Select All Hits*, *Deselect All Hits*, *Clear
|
|
186
|
+
Fields*, *Refresh*, plus a **Default -n** spinbox and **Set All -n**
|
|
187
|
+
button.
|
|
188
|
+
- **Options panel.** Aligner (Clustal Omega or MAFFT + mode), tree builder
|
|
189
|
+
(FastTree or RAxML), BLAST type (tblastn/blastp), thread count.
|
|
190
|
+
- **Advanced panel** (collapsible): outgroups (`-add`, `-add_db`), AA
|
|
191
|
+
slice (`-aa`), motif patterns (regex or PROSITE, overlap toggle), and
|
|
192
|
+
HMM profiles (`--hmm`).
|
|
193
|
+
- **Generate Command / Copy to Clipboard.** Produces a ready-to-paste
|
|
194
|
+
`blast-align-tree …` command.
|
|
195
|
+
- **Recent Runs tab.** Lists past `ENTRY/runs/TIMESTAMP/` directories in
|
|
196
|
+
the current working directory so you can quickly jump back to prior
|
|
197
|
+
results.
|
|
198
|
+
|
|
199
|
+
## Example output
|
|
200
|
+
|
|
201
|
+
The default genome set includes Arabidopsis TAIR10 CDS. The example below
|
|
202
|
+
runs the pipeline for a SERK query and redraws the resulting tree with a
|
|
203
|
+
new subnode/outgroup:
|
|
204
|
+
|
|
205
|
+
```
|
|
206
|
+
blast-align-tree -q AT4G33430.1 -qdbs TAIR10cds.fa \
|
|
207
|
+
-n 15 15 15 \
|
|
208
|
+
-dbs TAIR10cds.fa Pvul218cds.fa Vung469cds.fa \
|
|
209
|
+
-hdr gene: polypeptide= locus=
|
|
210
|
+
```
|
|
211
|
+
|
|
212
|
+
The pipeline creates a folder `AT4G33430.1/` in your working directory.
|
|
213
|
+
The timestamped run root keeps the tree PDFs. Newick tree files, gene
|
|
214
|
+
lists, alignment FASTAs, mappings, features, BLAST hit FASTAs, and
|
|
215
|
+
per-genome summaries go under
|
|
216
|
+
`AT4G33430.1/runs/<TIMESTAMP>/genes_alignments_trees/`.
|
|
217
|
+
|
|
218
|
+
After the run finishes, the pipeline prints a re-draw hint. For example,
|
|
219
|
+
to reroot on an outgroup (`-a AT5G10290`) and zoom in on a subnode
|
|
220
|
+
(`-n 45`):
|
|
221
|
+
|
|
222
|
+
```
|
|
223
|
+
Rscript "<bundled-visualize_tree.r>" -e AT4G33430.1 -b SERK_tree \
|
|
224
|
+
--subdir "runs/<TIMESTAMP>" -a AT5G10290 -n 45
|
|
225
|
+
```
|
|
226
|
+
|
|
227
|
+
The bundled path is printed for you at the end of each pipeline run,
|
|
228
|
+
already wrapped in double quotes — keep the quotes when copy-pasting,
|
|
229
|
+
especially on Windows, where unquoted paths can cause `Rscript` to
|
|
230
|
+
segfault if they contain spaces or backslashes that the shell misparses.
|
|
231
|
+
|
|
232
|
+
<!-- TODO: regenerate images/tree.png from the SERK run above -->
|
|
233
|
+

|
|
234
|
+
|
|
235
|
+
## Tutorial
|
|
236
|
+
|
|
237
|
+
### Run blast-align-tree for ACC Oxidase
|
|
238
|
+
|
|
239
|
+
Find 15 homologs of Arabidopsis ACC Oxidase 1 from three plant genomes
|
|
240
|
+
using `tblastn` against complete CDS databases. `-q` specifies the query
|
|
241
|
+
locus, `-qdbs` the database it lives in, `-dbs` the databases to search,
|
|
242
|
+
and `-hdr` the regex tokens used to parse gene names out of each
|
|
243
|
+
database's FASTA headers.
|
|
244
|
+
|
|
245
|
+
```
|
|
246
|
+
blast-align-tree -q AT2G19590.1 -qdbs TAIR10cds.fa \
|
|
247
|
+
-n 15 15 15 \
|
|
248
|
+
-dbs TAIR10cds.fa Pvul218cds.fa Vung469cds.fa \
|
|
249
|
+
-hdr gene: polypeptide= locus=
|
|
250
|
+
```
|
|
251
|
+
|
|
252
|
+
This creates `AT2G19590.1/` with tree PDFs at the timestamped run root and
|
|
253
|
+
Newick tree files, alignment files, BLAST hit FASTAs, and per-genome
|
|
254
|
+
summaries under `genes_alignments_trees/`.
|
|
255
|
+
|
|
256
|
+
A powerful feature of `ggtree` is the ability to plot associated data.
|
|
257
|
+
Each run produces two complementary tree PDFs:
|
|
258
|
+
|
|
259
|
+
1. **Text version** — gene symbols and dataset values printed as labels
|
|
260
|
+
next to each tip.
|
|
261
|
+
2. **Heatmap version** — the same tree with associated data rendered as
|
|
262
|
+
a coloured heatmap alongside the tips.
|
|
263
|
+
|
|
264
|
+
By default, both include expression data from the [Klepikova *Arabidopsis*
|
|
265
|
+
expression atlas](https://pubmed.ncbi.nlm.nih.gov/26923014/) (headers are
|
|
266
|
+
matched to the AtGenExpress / eFP browser tissue naming). The screenshot
|
|
267
|
+
below shows the heatmap version:
|
|
268
|
+
|
|
269
|
+
<!-- TODO: regenerate images/ACO-tree-1.png (heatmap version) from the ACO run above -->
|
|
270
|
+

|
|
271
|
+
|
|
272
|
+
A separate PDF with `.MSA.pdf` appended shows a cartoon alignment — useful
|
|
273
|
+
for spotting large differences in domain architecture. Open the
|
|
274
|
+
underlying FASTA files in `genes_alignments_trees/` to inspect the
|
|
275
|
+
alignment in detail.
|
|
276
|
+
|
|
277
|
+
<!-- TODO: regenerate images/ACO-tree-2.png (.MSA.pdf cartoon alignment) -->
|
|
278
|
+

|
|
279
|
+
|
|
280
|
+
### Redraw the ACC Oxidase tree
|
|
281
|
+
|
|
282
|
+
You can re-run `visualize_tree.r` at any time to produce new PDFs. The
|
|
283
|
+
pipeline prints a ready-to-edit `Rscript …` command at the end of each
|
|
284
|
+
run; copy it and tweak options such as:
|
|
285
|
+
|
|
286
|
+
- `-b <NAME>` — filename stem for the new PDFs
|
|
287
|
+
- `-a <ID>` — reroot on this outgroup
|
|
288
|
+
- `-n <NODE>` — draw a subtree at this node (use `--help` for the full
|
|
289
|
+
option list)
|
|
290
|
+
- `-k 1` — show bootstraps
|
|
291
|
+
- `-l 0` — hide node number labels
|
|
292
|
+
- `-m 2` — enlarge gene-symbol text
|
|
293
|
+
|
|
294
|
+
For example, reroot the default ACO tree on JRG21 (AT2G38240) and zoom
|
|
295
|
+
into the ACO clade at node 58:
|
|
296
|
+
|
|
297
|
+
```
|
|
298
|
+
Rscript "<bundled-visualize_tree.r>" -e AT2G19590.1 -b ACO_v3 \
|
|
299
|
+
--subdir "runs/<TIMESTAMP>" -a AT2G38240 -n 58 -k 1 -l 0 -m 2
|
|
300
|
+
```
|
|
301
|
+
|
|
302
|
+
<!-- TODO: regenerate images/ACO-tree-3.png -->
|
|
303
|
+

|
|
304
|
+
|
|
305
|
+
The `-n` option is especially helpful for extracting a subset of the tree
|
|
306
|
+
as a FASTA. Sequences are listed in the FASTA in the same order as the
|
|
307
|
+
tree, and `trimAl` is used to strip blank-only alignment columns — useful
|
|
308
|
+
for a quick view of conserved residues (e.g. the ACO active site) in a
|
|
309
|
+
viewer like AliView.
|
|
310
|
+
|
|
311
|
+
### BLASTP instead of TBLASTN
|
|
312
|
+
|
|
313
|
+
Use `--blast_type blastp` against protein databases. The example below
|
|
314
|
+
pulls 10 NIMIN-1 homologs from the Arabidopsis and *Nicotiana
|
|
315
|
+
benthamiana* proteomes:
|
|
316
|
+
|
|
317
|
+
```
|
|
318
|
+
blast-align-tree --blast_type blastp \
|
|
319
|
+
-q AT1G02450.1 -qdbs TAIR10protein.fa \
|
|
320
|
+
-n 10 10 \
|
|
321
|
+
-dbs TAIR10protein.fa Niben261_genome.annotation.proteins.fasta \
|
|
322
|
+
-hdr gene: id
|
|
323
|
+
```
|
|
324
|
+
|
|
325
|
+
### Multiple queries
|
|
326
|
+
|
|
327
|
+
You can pass several query sequences with `-q`; the pipeline extracts
|
|
328
|
+
each from the database listed at the matching position in `-qdbs`,
|
|
329
|
+
de-duplicates, then searches each `-dbs` entry. The example below uses
|
|
330
|
+
three queries drawn from two databases and searches two other databases:
|
|
331
|
+
|
|
332
|
+
```
|
|
333
|
+
blast-align-tree -q AT5G45250.1 Phvul.007G077500.1 AT5G17890.1 \
|
|
334
|
+
-qdbs TAIR10cds.fa Pvul218cds.fa TAIR10cds.fa \
|
|
335
|
+
-n 3 4 \
|
|
336
|
+
-dbs TAIR10cds.fa Vung469cds.fa \
|
|
337
|
+
-hdr gene: locus=
|
|
338
|
+
```
|
|
339
|
+
|
|
340
|
+
### Adding a new genome
|
|
341
|
+
|
|
342
|
+
You can drop additional genomes into `./genomes/` and use them alongside
|
|
343
|
+
the bundled ones. For each new genome you need a local BLAST database.
|
|
344
|
+
|
|
345
|
+
For CDS files:
|
|
346
|
+
|
|
347
|
+
```
|
|
348
|
+
makeblastdb -in GenomeCDS.fa -parse_seqids -dbtype nucl
|
|
349
|
+
```
|
|
350
|
+
|
|
351
|
+
For protein files:
|
|
352
|
+
|
|
353
|
+
```
|
|
354
|
+
makeblastdb -in GenomeProteins.fa -parse_seqids -dbtype prot
|
|
355
|
+
```
|
|
356
|
+
|
|
357
|
+
#### Example: add *Nicotiana tabacum* and rebuild an earlier tree
|
|
358
|
+
|
|
359
|
+
Download an annotated tobacco proteome into `./genomes/`:
|
|
360
|
+
|
|
361
|
+
- [*N. tabacum* v4.5 from Sol Genomics](https://solgenomics.net/ftp/ftp/genomes/Nicotiana_tabacum/edwards_et_al_2017/annotation/)
|
|
362
|
+
→ `Nitab-v4.5_proteins_Edwards2017.fasta`
|
|
363
|
+
|
|
364
|
+
Build the BLAST database:
|
|
365
|
+
|
|
366
|
+
```
|
|
367
|
+
cd genomes
|
|
368
|
+
makeblastdb -in Nitab-v4.5_proteins_Edwards2017.fasta -parse_seqids -dbtype prot
|
|
369
|
+
cd ..
|
|
370
|
+
```
|
|
371
|
+
|
|
372
|
+
Inspecting the first record of `Nitab-v4.5_proteins_Edwards2017.fasta`
|
|
373
|
+
shows that headers use a gene id followed by a description — `-hdr id`
|
|
374
|
+
keeps just the first token.
|
|
375
|
+
|
|
376
|
+
Now build a SOBIR1 homolog tree across Arabidopsis, *N. benthamiana*,
|
|
377
|
+
and the freshly added tobacco proteome:
|
|
378
|
+
|
|
379
|
+
```
|
|
380
|
+
blast-align-tree --blast_type blastp \
|
|
381
|
+
-q AT2G31880.1 -qdbs TAIR10protein.fa \
|
|
382
|
+
-n 10 10 10 \
|
|
383
|
+
-dbs TAIR10protein.fa \
|
|
384
|
+
Niben261_genome.annotation.proteins.fasta \
|
|
385
|
+
Nitab-v4.5_proteins_Edwards2017.fasta \
|
|
386
|
+
-hdr gene: id id
|
|
387
|
+
```
|
|
388
|
+
|
|
389
|
+
The run produces a tree PDF with tobacco SOBIR1 homologs slotted in
|
|
390
|
+
alongside the *N. benthamiana* and Arabidopsis sequences.
|
|
391
|
+
|
|
392
|
+
<!-- TODO: regenerate images/SOBIR1_with_ntab.png from the run above -->
|
|
393
|
+

|
|
394
|
+
|
|
395
|
+
### Rebuild the SOBIR1 tree with MAFFT and RAxML
|
|
396
|
+
|
|
397
|
+
By default the pipeline aligns with Clustal Omega and infers the tree
|
|
398
|
+
with FastTree, but both are swappable. The command below rebuilds the
|
|
399
|
+
same SOBIR1 tree (Arabidopsis + *N. benthamiana* + tobacco) using
|
|
400
|
+
**MAFFT** in `linsi` mode and **RAxML-NG** for the tree inference:
|
|
401
|
+
|
|
402
|
+
```
|
|
403
|
+
blast-align-tree --blast_type blastp \
|
|
404
|
+
--aligner mafft --mafft_mode linsi \
|
|
405
|
+
--tree_builder RAxML \
|
|
406
|
+
-q AT2G31880.1 -qdbs TAIR10protein.fa \
|
|
407
|
+
-n 10 10 10 \
|
|
408
|
+
-dbs TAIR10protein.fa \
|
|
409
|
+
Niben261_genome.annotation.proteins.fasta \
|
|
410
|
+
Nitab-v4.5_proteins_Edwards2017.fasta \
|
|
411
|
+
-hdr gene: id id
|
|
412
|
+
```
|
|
413
|
+
|
|
414
|
+
RAxML-NG is noticeably slower than FastTree but provides maximum-
|
|
415
|
+
likelihood branch support via bootstrapping. Comparing the two trees is
|
|
416
|
+
a quick sanity check that any clades you care about are stable across
|
|
417
|
+
inference methods.
|
|
418
|
+
|
|
419
|
+
Use repeated rounds of querying to refine your trees, search different
|
|
420
|
+
genome versions, and compare aligners / tree builders before drawing
|
|
421
|
+
strong conclusions.
|
|
422
|
+
|
|
423
|
+
## Future features
|
|
424
|
+
|
|
425
|
+
We are currently working on:
|
|
426
|
+
|
|
427
|
+
1. Iterative BLAST using a first set of hits as secondary queries
|
|
428
|
+
2. Better organization of query and sub-query folders
|
|
429
|
+
3. Displaying a subsequence of the MSA in the alignment PDF
|
|
430
|
+
4. Richer motif / HMM visualization overlaid on the MSA
|
|
431
|
+
|
|
432
|
+
If you'd like to contribute, reach out to Ben and Adam:
|
|
433
|
+
`bdshep@uw.edu` and `astein10@uw.edu`.
|