minimap2 0.2.30.3 → 1.2.31.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (72) hide show
  1. checksums.yaml +4 -4
  2. data/README.md +48 -63
  3. data/ext/minimap2/Makefile +1 -1
  4. data/ext/minimap2/format.c +12 -2
  5. data/ext/minimap2/hit.c +19 -8
  6. data/ext/minimap2/ksw2_ll_sse.c +2 -2
  7. data/ext/minimap2/minimap.h +2 -1
  8. data/ext/minimap2/pe.c +7 -2
  9. data/ext/ruby_minimap2/extconf.rb +47 -0
  10. data/ext/ruby_minimap2/native_main.c +15 -0
  11. data/ext/ruby_minimap2/ruby_minimap2.c +1134 -0
  12. data/lib/minimap2/aligner.rb +66 -374
  13. data/lib/minimap2/alignment.rb +22 -11
  14. data/lib/minimap2/version.rb +2 -1
  15. data/lib/minimap2.rb +39 -110
  16. metadata +8 -89
  17. data/ext/Rakefile +0 -60
  18. data/ext/cmappy/cmappy.c +0 -134
  19. data/ext/cmappy/cmappy.h +0 -46
  20. data/ext/minimap2/FAQ.md +0 -46
  21. data/ext/minimap2/MANIFEST.in +0 -10
  22. data/ext/minimap2/Makefile.simde +0 -97
  23. data/ext/minimap2/NEWS.md +0 -989
  24. data/ext/minimap2/README.md +0 -428
  25. data/ext/minimap2/code_of_conduct.md +0 -30
  26. data/ext/minimap2/cookbook.md +0 -243
  27. data/ext/minimap2/minimap2.1 +0 -844
  28. data/ext/minimap2/misc/README.md +0 -180
  29. data/ext/minimap2/misc/pafcluster.js +0 -241
  30. data/ext/minimap2/misc/paftools.js +0 -3734
  31. data/ext/minimap2/pyproject.toml +0 -2
  32. data/ext/minimap2/python/README.rst +0 -198
  33. data/ext/minimap2/python/cmappy.h +0 -152
  34. data/ext/minimap2/python/cmappy.pxd +0 -156
  35. data/ext/minimap2/python/mappy.pyx +0 -289
  36. data/ext/minimap2/python/minimap2.py +0 -41
  37. data/ext/minimap2/setup.py +0 -55
  38. data/ext/minimap2/test/MT-human.fa +0 -278
  39. data/ext/minimap2/test/MT-orang.fa +0 -276
  40. data/ext/minimap2/test/q-inv.fa +0 -4
  41. data/ext/minimap2/test/q2.fa +0 -2
  42. data/ext/minimap2/test/t-inv.fa +0 -127
  43. data/ext/minimap2/test/t2.fa +0 -2
  44. data/ext/minimap2/test/x3s-aln.txt +0 -5
  45. data/ext/minimap2/test/x3s-qry.fa +0 -5
  46. data/ext/minimap2/test/x3s-ref.fa +0 -10
  47. data/ext/minimap2/tex/Makefile +0 -21
  48. data/ext/minimap2/tex/bioinfo.cls +0 -930
  49. data/ext/minimap2/tex/blasr-mc.eval +0 -17
  50. data/ext/minimap2/tex/bowtie2-s3.sam.eval +0 -28
  51. data/ext/minimap2/tex/bwa-s3.sam.eval +0 -52
  52. data/ext/minimap2/tex/bwa.eval +0 -55
  53. data/ext/minimap2/tex/eval2roc.pl +0 -33
  54. data/ext/minimap2/tex/graphmap.eval +0 -4
  55. data/ext/minimap2/tex/hs38-simu.sh +0 -10
  56. data/ext/minimap2/tex/minialign.eval +0 -49
  57. data/ext/minimap2/tex/minimap2.bib +0 -460
  58. data/ext/minimap2/tex/minimap2.tex +0 -724
  59. data/ext/minimap2/tex/mm2-s3.sam.eval +0 -62
  60. data/ext/minimap2/tex/mm2-update.tex +0 -240
  61. data/ext/minimap2/tex/mm2.approx.eval +0 -12
  62. data/ext/minimap2/tex/mm2.eval +0 -13
  63. data/ext/minimap2/tex/natbib.bst +0 -1288
  64. data/ext/minimap2/tex/natbib.sty +0 -803
  65. data/ext/minimap2/tex/ngmlr.eval +0 -38
  66. data/ext/minimap2/tex/roc.gp +0 -60
  67. data/ext/minimap2/tex/snap-s3.sam.eval +0 -62
  68. data/ext/minimap2.patch +0 -19
  69. data/lib/minimap2/ffi/constants.rb +0 -267
  70. data/lib/minimap2/ffi/functions.rb +0 -239
  71. data/lib/minimap2/ffi/mappy.rb +0 -104
  72. data/lib/minimap2/ffi.rb +0 -27
@@ -1,428 +0,0 @@
1
- [![GitHub Downloads](https://img.shields.io/github/downloads/lh3/minimap2/total.svg?style=social&logo=github&label=Download)](https://github.com/lh3/minimap2/releases)
2
- [![BioConda Install](https://img.shields.io/conda/dn/bioconda/minimap2.svg?style=flag&label=BioConda%20install)](https://anaconda.org/bioconda/minimap2)
3
- [![PyPI](https://img.shields.io/pypi/v/mappy.svg?style=flat)](https://pypi.python.org/pypi/mappy)
4
- [![Build Status](https://github.com/lh3/minimap2/actions/workflows/ci.yaml/badge.svg)](https://github.com/lh3/minimap2/actions)
5
- ## <a name="started"></a>Getting Started
6
- ```sh
7
- git clone https://github.com/lh3/minimap2
8
- cd minimap2 && make
9
- # long sequences against a reference genome
10
- ./minimap2 -a test/MT-human.fa test/MT-orang.fa > test.sam
11
- # create an index first and then map
12
- ./minimap2 -x map-ont -d MT-human-ont.mmi test/MT-human.fa
13
- ./minimap2 -a MT-human-ont.mmi test/MT-orang.fa > test.sam
14
- # use presets (no test data)
15
- ./minimap2 -ax map-pb ref.fa pacbio.fq.gz > aln.sam # PacBio CLR genomic reads
16
- ./minimap2 -ax map-ont ref.fa ont.fq.gz > aln.sam # Oxford Nanopore genomic reads
17
- ./minimap2 -ax map-hifi ref.fa pacbio-ccs.fq.gz > aln.sam # PacBio HiFi/CCS genomic reads (v2.19+)
18
- ./minimap2 -ax lr:hq ref.fa ont-Q20.fq.gz > aln.sam # Nanopore Q20 genomic reads (v2.27+)
19
- ./minimap2 -ax sr ref.fa read1.fa read2.fa > aln.sam # short genomic paired-end reads
20
- ./minimap2 -ax splice ref.fa rna-reads.fa > aln.sam # spliced long reads (strand unknown)
21
- ./minimap2 -ax splice -uf -k14 ref.fa reads.fa > aln.sam # noisy Nanopore direct RNA-seq
22
- ./minimap2 -ax splice:hq -uf ref.fa query.fa > aln.sam # PacBio Kinnex/Iso-seq (RNA-seq)
23
- ./minimap2 -ax splice --junc-bed=anno.bed12 ref.fa query.fa > aln.sam # use annotated junctions
24
- ./minimap2 -ax splice:sr ref.fa r1.fq r2.fq > aln.sam # short-read RNA-seq (v2.29+)
25
- ./minimap2 -ax splice:sr -j anno.bed12 ref.fa r1.fq r2.fq > aln.sam
26
- ./minimap2 -cx asm5 asm1.fa asm2.fa > aln.paf # intra-species asm-to-asm alignment
27
- ./minimap2 -x ava-pb reads.fa reads.fa > overlaps.paf # PacBio read overlap
28
- ./minimap2 -x ava-ont reads.fa reads.fa > overlaps.paf # Nanopore read overlap
29
- # man page for detailed command line options
30
- man ./minimap2.1
31
- ```
32
-
33
- ## Table of Contents
34
-
35
- - [Getting Started](#started)
36
- - [Users' Guide](#uguide)
37
- - [Installation](#install)
38
- - [General usage](#general)
39
- - [Use cases](#cases)
40
- - [Map long noisy genomic reads](#map-long-genomic)
41
- - [Map long mRNA/cDNA reads](#map-long-splice)
42
- - [Find overlaps between long reads](#long-overlap)
43
- - [Map short genomic reads](#short-genomic)
44
- - [Map short RNA-seq reads](#short-rna-seq)
45
- - [Full genome/assembly alignment](#full-genome)
46
- - [Advanced features](#advanced)
47
- - [Working with >65535 CIGAR operations](#long-cigar)
48
- - [The cs optional tag](#cs)
49
- - [Working with the PAF format](#paftools)
50
- - [Algorithm overview](#algo)
51
- - [Getting help](#help)
52
- - [Citing minimap2](#cite)
53
- - [Developers' Guide](#dguide)
54
- - [Limitations](#limit)
55
-
56
- ## <a name="uguide"></a>Users' Guide
57
-
58
- Minimap2 is a versatile sequence alignment program that aligns DNA or mRNA
59
- sequences against a large reference database. Typical use cases include: (1)
60
- mapping PacBio or Oxford Nanopore genomic reads to the human genome; (2)
61
- finding overlaps between long reads with error rate up to ~15%; (3)
62
- splice-aware alignment of PacBio Iso-Seq or Nanopore cDNA or Direct RNA reads
63
- against a reference genome; (4) aligning Illumina single- or paired-end reads;
64
- (5) assembly-to-assembly alignment; (6) full-genome alignment between two
65
- closely related species with divergence below ~15%.
66
-
67
- For ~10kb noisy reads sequences, minimap2 is tens of times faster than
68
- mainstream long-read mappers such as BLASR, BWA-MEM, NGMLR and GMAP. It is more
69
- accurate on simulated long reads and produces biologically meaningful alignment
70
- ready for downstream analyses. For >100bp Illumina short reads, minimap2 is
71
- three times as fast as BWA-MEM and Bowtie2, and as accurate on simulated data.
72
- Detailed evaluations are available from the [minimap2 paper][doi] or the
73
- [preprint][preprint].
74
-
75
- ### <a name="install"></a>Installation
76
-
77
- Minimap2 is optimized for x86-64 CPUs. You can acquire precompiled binaries from
78
- the [release page][release] with:
79
- ```sh
80
- curl -L https://github.com/lh3/minimap2/releases/download/v2.30/minimap2-2.30_x64-linux.tar.bz2 | tar -jxvf -
81
- ./minimap2-2.30_x64-linux/minimap2
82
- ```
83
- If you want to compile from the source, you need to have a C compiler, GNU make
84
- and zlib development files installed. Then type `make` in the source code
85
- directory to compile. If you see compilation errors, try `make sse2only=1`
86
- to disable SSE4 code, which will make minimap2 slightly slower.
87
-
88
- Minimap2 also works with ARM CPUs supporting the NEON instruction sets. To
89
- compile for 32 bit ARM architectures (such as ARMv7), use `make arm_neon=1`. To
90
- compile for for 64 bit ARM architectures (such as ARMv8), use `make arm_neon=1
91
- aarch64=1`.
92
-
93
- Minimap2 can use [SIMD Everywhere (SIMDe)][simde] library for porting
94
- implementation to the different SIMD instruction sets. To compile using SIMDe,
95
- use `make -f Makefile.simde`. To compile for ARM CPUs, use `Makefile.simde`
96
- with the ARM related command lines given above.
97
-
98
- ### <a name="general"></a>General usage
99
-
100
- Without any options, minimap2 takes a reference database and a query sequence
101
- file as input and produce approximate mapping, without base-level alignment
102
- (i.e. coordinates are only approximate and no CIGAR in output), in the [PAF format][paf]:
103
- ```sh
104
- minimap2 ref.fa query.fq > approx-mapping.paf
105
- ```
106
- You can ask minimap2 to generate CIGAR at the `cg` tag of PAF with:
107
- ```sh
108
- minimap2 -c ref.fa query.fq > alignment.paf
109
- ```
110
- or to output alignments in the [SAM format][sam]:
111
- ```sh
112
- minimap2 -a ref.fa query.fq > alignment.sam
113
- ```
114
- Minimap2 seamlessly works with gzip'd FASTA and FASTQ formats as input. You
115
- don't need to convert between FASTA and FASTQ or decompress gzip'd files first.
116
-
117
- For the human reference genome, minimap2 takes a few minutes to generate a
118
- minimizer index for the reference before mapping. To reduce indexing time, you
119
- can optionally save the index with option **-d** and replace the reference
120
- sequence file with the index file on the minimap2 command line:
121
- ```sh
122
- minimap2 -d ref.mmi ref.fa # indexing
123
- minimap2 -a ref.mmi reads.fq > alignment.sam # alignment
124
- ```
125
- ***Importantly***, it should be noted that once you build the index, indexing
126
- parameters such as **-k**, **-w**, **-H** and **-I** can't be changed during
127
- mapping. If you are running minimap2 for different data types, you will
128
- probably need to keep multiple indexes generated with different parameters.
129
- This makes minimap2 different from BWA which always uses the same index
130
- regardless of query data types.
131
-
132
- ### <a name="cases"></a>Use cases
133
-
134
- Minimap2 uses the same base algorithm for all applications. However, due to the
135
- different data types it supports (e.g. short vs long reads; DNA vs mRNA reads),
136
- minimap2 needs to be tuned for optimal performance and accuracy. It is usually
137
- recommended to choose a preset with option **-x**, which sets multiple
138
- parameters at the same time. The default setting is the same as `map-ont`.
139
-
140
- #### <a name="map-long-genomic"></a>Map long noisy genomic reads
141
-
142
- ```sh
143
- minimap2 -ax map-pb ref.fa pacbio-reads.fq > aln.sam # for PacBio CLR reads
144
- minimap2 -ax map-ont ref.fa ont-reads.fq > aln.sam # for Oxford Nanopore reads
145
- minimap2 -ax map-iclr ref.fa iclr-reads.fq > aln.sam # for Illumina Complete Long Reads
146
- ```
147
- The difference between `map-pb` and `map-ont` is that `map-pb` uses
148
- homopolymer-compressed (HPC) minimizers as seeds, while `map-ont` uses ordinary
149
- minimizers as seeds. Empirical evaluation suggests HPC minimizers improve
150
- performance and sensitivity when aligning PacBio CLR reads, but hurt when aligning
151
- Nanopore reads. `map-iclr` uses an adjusted alignment scoring matrix that
152
- accounts for the low overall error rate in the reads, with transversion errors
153
- being less frequent than transitions.
154
-
155
- #### <a name="map-long-splice"></a>Map long mRNA/cDNA reads
156
-
157
- ```sh
158
- minimap2 -ax splice:hq -uf ref.fa iso-seq.fq > aln.sam # PacBio Iso-seq/traditional cDNA
159
- minimap2 -ax splice ref.fa nanopore-cdna.fa > aln.sam # Nanopore 2D cDNA-seq
160
- minimap2 -ax splice -uf -k14 ref.fa direct-rna.fq > aln.sam # Nanopore Direct RNA-seq
161
- minimap2 -ax splice --splice-flank=no SIRV.fa SIRV-seq.fa # mapping against SIRV control
162
- ```
163
- There are different long-read RNA-seq technologies, including tranditional
164
- full-length cDNA, EST, PacBio Iso-seq, Nanopore 2D cDNA-seq and Direct RNA-seq.
165
- They produce data of varying quality and properties. By default, `-x splice`
166
- assumes the read orientation relative to the transcript strand is unknown. It
167
- tries two rounds of alignment to infer the orientation and write the strand to
168
- the `ts` SAM/PAF tag if possible. For Iso-seq, Direct RNA-seq and tranditional
169
- full-length cDNAs, it would be desired to apply `-u f` to force minimap2 to
170
- consider the forward transcript strand only. This speeds up alignment with
171
- slight improvement to accuracy. For noisy Nanopore Direct RNA-seq reads, it is
172
- recommended to use a smaller k-mer size for increased sensitivity to the first
173
- or the last exons.
174
-
175
- Minimap2 rates an alignment by the score of the max-scoring sub-segment,
176
- *excluding* introns, and marks the best alignment as primary in SAM. When a
177
- spliced gene also has unspliced pseudogenes, minimap2 slightly prefers
178
- the spliced alignment. By default, minimap2 outputs up to five secondary
179
- alignments (i.e. likely pseudogenes in the context of RNA-seq mapping). This
180
- can be tuned with option **-N**.
181
-
182
- For long RNA-seq reads, minimap2 may produce chimeric alignments potentially
183
- caused by gene fusions/structural variations or by an intron longer than the
184
- max intron length **-G** (200k by default). For now, it is not recommended to
185
- apply an excessively large **-G** as this slows down minimap2 and sometimes
186
- leads to false alignments.
187
-
188
- It is worth noting that by default `-x splice` prefers GT[A/G]..[C/T]AG
189
- over GT[C/T]..[A/G]AG, and then over other splicing signals. Considering
190
- one additional base improves the junction accuracy for noisy reads, but
191
- reduces the accuracy when aligning against the widely used SIRV control data.
192
- This is because SIRV does not honor the evolutionarily conservative splicing
193
- signal. If you are studying SIRV, you may apply `--splice-flank=no` to let
194
- minimap2 only model GT..AG, ignoring the additional base.
195
-
196
- Since v2.17, minimap2 can optionally take annotated genes as input and
197
- prioritize on annotated splice junctions. To use this feature, you can
198
- ```sh
199
- paftools.js gff2bed anno.gff > anno.bed
200
- minimap2 -ax splice --junc-bed anno.bed ref.fa query.fa > aln.sam
201
- ```
202
- Here, `anno.gff` is the gene annotation in the GTF or GFF3 format (`gff2bed`
203
- automatically tests the format). The output of `gff2bed` is in the 12-column
204
- BED format, or the BED12 format. With the `--junc-bed` option, minimap2 adds a
205
- bonus score (tuned by `--junc-bonus`) if an aligned junction matches a junction
206
- in the annotation. Option `--junc-bed` also takes 5-column BED, including the
207
- strand field. In this case, each line indicates an oriented junction.
208
-
209
- **Note:** `--junc-bed` is intended for long noisy RNA-seq reads only.
210
- Applying the option to short RNA-seq reads would increase run time with little
211
- improvement to junction accuracy.
212
-
213
- #### <a name="long-overlap"></a>Find overlaps between long reads
214
-
215
- ```sh
216
- minimap2 -x ava-pb reads.fq reads.fq > ovlp.paf # PacBio CLR read overlap
217
- minimap2 -x ava-ont reads.fq reads.fq > ovlp.paf # Oxford Nanopore read overlap
218
- ```
219
- Similarly, `ava-pb` uses HPC minimizers while `ava-ont` uses ordinary
220
- minimizers. It is usually not recommended to perform base-level alignment in
221
- the overlapping mode because it is slow and may produce false positive
222
- overlaps. However, if performance is not a concern, you may try to add `-a` or
223
- `-c` anyway.
224
-
225
- #### <a name="short-genomic"></a>Map short genomic reads
226
-
227
- ```sh
228
- minimap2 -ax sr ref.fa reads-se.fq > aln.sam # single-end alignment
229
- minimap2 -ax sr ref.fa read1.fq read2.fq > aln.sam # paired-end alignment
230
- minimap2 -ax sr ref.fa reads-interleaved.fq > aln.sam # paired-end alignment
231
- ```
232
- When two read files are specified, minimap2 reads from each file in turn and
233
- merge them into an interleaved stream internally. Two reads are considered to
234
- be paired if they are adjacent in the input stream and have the same name (with
235
- the `/[0-9]` suffix trimmed if present). Single- and paired-end reads can be
236
- mixed.
237
-
238
- #### <a name="short-rna-seq"></a>Map short RNA-seq reads
239
-
240
- ```sh
241
- minimap2 -ax splice:sr ref.fa reads-se.fq.gz > aln.sam # single-end
242
- minimap2 -ax splice:sr ref.fa r1.fq.gz r2.fq.gz > aln.sam # paired-end
243
- minimap2 -ax splice:sr -j anno.bed ref.fa r1.fq r2.fq > aln.sam # use annotation
244
- # 2-pass alignment
245
- minimap2 -x splice:sr -j anno.bed --write-junc ref.fa r1.fq r2.fq > junc.bed
246
- minimap2 -ax splice:sr -j anno.bed --pass1=junc.bed ref.fa r1.fq r2.fq > aln.sam
247
- ```
248
- The new preset `splice:sr` was added in v2.29. It functions similarly to `sr`
249
- except that it performs spliced alignment.
250
-
251
- #### <a name="full-genome"></a>Full genome/assembly alignment
252
-
253
- ```sh
254
- minimap2 -ax asm5 ref.fa asm.fa > aln.sam # assembly to assembly/ref alignment
255
- ```
256
- For cross-species full-genome alignment, the scoring system needs to be tuned
257
- according to the sequence divergence.
258
-
259
- ### <a name="advanced"></a>Advanced features
260
-
261
- #### <a name="long-cigar"></a>Working with >65535 CIGAR operations
262
-
263
- Due to a design flaw, BAM does not work with CIGAR strings with >65535
264
- operations (SAM and CRAM work). However, for ultra-long nanopore reads minimap2
265
- may align ~1% of read bases with long CIGARs beyond the capability of BAM. If
266
- you convert such SAM/CRAM to BAM, Picard and recent samtools will throw an
267
- error and abort. Older samtools and other tools may create corrupted BAM.
268
-
269
- To avoid this issue, you can add option `-L` at the minimap2 command line.
270
- This option moves a long CIGAR to the `CG` tag and leaves a fully clipped CIGAR
271
- at the SAM CIGAR column. Current tools that don't read CIGAR (e.g. merging and
272
- sorting) still work with such BAM records; tools that read CIGAR will
273
- effectively ignore these records. It has been decided that future tools
274
- will seamlessly recognize long-cigar records generated by option `-L`.
275
-
276
- **TL;DR**: if you work with ultra-long reads and use tools that only process
277
- BAM files, please add option `-L`.
278
-
279
- #### <a name="cs"></a>The cs optional tag
280
-
281
- The `cs` SAM/PAF tag encodes bases at mismatches and INDELs. It matches regular
282
- expression `/(:[0-9]+|\*[a-z][a-z]|[=\+\-][A-Za-z]+)+/`. Like CIGAR, `cs`
283
- consists of series of operations. Each leading character specifies the
284
- operation; the following sequence is the one involved in the operation.
285
-
286
- The `cs` tag is enabled by command line option `--cs`. The following alignment,
287
- for example:
288
- ```txt
289
- CGATCGATAAATAGAGTAG---GAATAGCA
290
- |||||| |||||||||| |||| |||
291
- CGATCG---AATAGAGTAGGTCGAATtGCA
292
- ```
293
- is represented as `:6-ata:10+gtc:4*at:3`, where `:[0-9]+` represents an
294
- identical block, `-ata` represents a deletion, `+gtc` an insertion and `*at`
295
- indicates reference base `a` is substituted with a query base `t`. It is
296
- similar to the `MD` SAM tag but is standalone and easier to parse.
297
-
298
- If `--cs=long` is used, the `cs` string also contains identical sequences in
299
- the alignment. The above example will become
300
- `=CGATCG-ata=AATAGAGTAG+gtc=GAAT*at=GCA`. The long form of `cs` encodes both
301
- reference and query sequences in one string. The `cs` tag also encodes intron
302
- positions and splicing signals (see the [minimap2 manpage][manpage-cs] for
303
- details).
304
-
305
- #### <a name="paftools"></a>Working with the PAF format
306
-
307
- Minimap2 also comes with a (java)script [paftools.js](misc/paftools.js) that
308
- processes alignments in the PAF format. It calls variants from
309
- assembly-to-reference alignment, lifts over BED files based on alignment,
310
- converts between formats and provides utilities for various evaluations. For
311
- details, please see [misc/README.md](misc/README.md).
312
-
313
- ### <a name="algo"></a>Algorithm overview
314
-
315
- In the following, minimap2 command line options have a dash ahead and are
316
- highlighted in bold. The description may help to tune minimap2 parameters.
317
-
318
- 1. Read **-I** [=*4G*] reference bases, extract (**-k**,**-w**)-minimizers and
319
- index them in a hash table.
320
-
321
- 2. Read **-K** [=*200M*] query bases. For each query sequence, do step 3
322
- through 7:
323
-
324
- 3. For each (**-k**,**-w**)-minimizer on the query, check against the reference
325
- index. If a reference minimizer is not among the top **-f** [=*2e-4*] most
326
- frequent, collect its the occurrences in the reference, which are called
327
- *seeds*.
328
-
329
- 4. Sort seeds by position in the reference. Chain them with dynamic
330
- programming. Each chain represents a potential mapping. For read
331
- overlapping, report all chains and then go to step 8. For reference mapping,
332
- do step 5 through 7:
333
-
334
- 5. Let *P* be the set of primary mappings, which is an empty set initially. For
335
- each chain from the best to the worst according to their chaining scores: if
336
- on the query, the chain overlaps with a chain in *P* by **--mask-level**
337
- [=*0.5*] or higher fraction of the shorter chain, mark the chain as
338
- *secondary* to the chain in *P*; otherwise, add the chain to *P*.
339
-
340
- 6. Retain all primary mappings. Also retain up to **-N** [=*5*] top secondary
341
- mappings if their chaining scores are higher than **-p** [=*0.8*] of their
342
- corresponding primary mappings.
343
-
344
- 7. If alignment is requested, filter out an internal seed if it potentially
345
- leads to both a long insertion and a long deletion. Extend from the
346
- left-most seed. Perform global alignments between internal seeds. Split the
347
- chain if the accumulative score along the global alignment drops by **-z**
348
- [=*400*], disregarding long gaps. Extend from the right-most seed. Output
349
- chains and their alignments.
350
-
351
- 8. If there are more query sequences in the input, go to step 2 until no more
352
- queries are left.
353
-
354
- 9. If there are more reference sequences, reopen the query file from the start
355
- and go to step 1; otherwise stop.
356
-
357
- ### <a name="help"></a>Getting help
358
-
359
- Manpage [minimap2.1][manpage] provides detailed description of minimap2
360
- command line options and optional tags. The [FAQ](FAQ.md) page answers several
361
- frequently asked questions. If you encounter bugs or have further questions or
362
- requests, you can raise an issue at the [issue page][issue]. There is not a
363
- specific mailing list for the time being.
364
-
365
- ### <a name="cite"></a>Citing minimap2
366
-
367
- If you use minimap2 in your work, please cite:
368
-
369
- > Li, H. (2018). Minimap2: pairwise alignment for nucleotide sequences.
370
- > *Bioinformatics*, **34**:3094-3100. [doi:10.1093/bioinformatics/bty191][doi]
371
-
372
- and/or:
373
-
374
- > Li, H. (2021). New strategies to improve minimap2 alignment accuracy.
375
- > *Bioinformatics*, **37**:4572-4574. [doi:10.1093/bioinformatics/btab705][doi2]
376
-
377
- ## <a name="dguide"></a>Developers' Guide
378
-
379
- Minimap2 is not only a command line tool, but also a programming library.
380
- It provides C APIs to build/load index and to align sequences against the
381
- index. File [example.c](example.c) demonstrates typical uses of C APIs. Header
382
- file [minimap.h](minimap.h) gives more detailed API documentation. Minimap2
383
- aims to keep APIs in this header stable. File [mmpriv.h](mmpriv.h) contains
384
- additional private APIs which may be subjected to changes frequently.
385
-
386
- This repository also provides Python bindings to a subset of C APIs. File
387
- [python/README.rst](python/README.rst) gives the full documentation;
388
- [python/minimap2.py](python/minimap2.py) shows an example. This Python
389
- extension, mappy, is also [available from PyPI][mappypypi] via `pip install
390
- mappy` or [from BioConda][mappyconda] via `conda install -c bioconda mappy`.
391
-
392
- ## <a name="limit"></a>Limitations
393
-
394
- * Minimap2 may produce suboptimal alignments through long low-complexity
395
- regions where seed positions may be suboptimal. This should not be a big
396
- concern because even the optimal alignment may be wrong in such regions.
397
-
398
- * Minimap2 requires SSE2 instructions on x86 CPUs or NEON on ARM CPUs. It is
399
- possible to add non-SIMD support, but it would make minimap2 slower by
400
- several times.
401
-
402
- * Minimap2 does not work with a single query or database sequence ~2
403
- billion bases or longer (2,147,483,647 to be exact). The total length of all
404
- sequences can well exceed this threshold.
405
-
406
- * Minimap2 often misses small exons.
407
-
408
-
409
-
410
- [paf]: https://github.com/lh3/miniasm/blob/master/PAF.md
411
- [sam]: https://samtools.github.io/hts-specs/SAMv1.pdf
412
- [minimap]: https://github.com/lh3/minimap
413
- [smartdenovo]: https://github.com/ruanjue/smartdenovo
414
- [longislnd]: https://www.ncbi.nlm.nih.gov/pubmed/27667791
415
- [gaba]: https://github.com/ocxtal/libgaba
416
- [ksw2]: https://github.com/lh3/ksw2
417
- [preprint]: https://arxiv.org/abs/1708.01492
418
- [release]: https://github.com/lh3/minimap2/releases
419
- [mappypypi]: https://pypi.python.org/pypi/mappy
420
- [mappyconda]: https://anaconda.org/bioconda/mappy
421
- [issue]: https://github.com/lh3/minimap2/issues
422
- [k8]: https://github.com/attractivechaos/k8
423
- [manpage]: https://lh3.github.io/minimap2/minimap2.html
424
- [manpage-cs]: https://lh3.github.io/minimap2/minimap2.html#10
425
- [doi]: https://doi.org/10.1093/bioinformatics/bty191
426
- [doi2]: https://doi.org/10.1093/bioinformatics/btab705
427
- [simde]: https://github.com/nemequ/simde
428
- [unimap]: https://github.com/lh3/unimap
@@ -1,30 +0,0 @@
1
- ## Contributor Code of Conduct
2
-
3
- As contributors and maintainers of this project, we pledge to respect all
4
- people who contribute through reporting issues, posting feature requests,
5
- updating documentation, submitting pull requests or patches, and other
6
- activities.
7
-
8
- We are committed to making participation in this project a harassment-free
9
- experience for everyone, regardless of level of experience, gender, gender
10
- identity and expression, sexual orientation, disability, personal appearance,
11
- body size, race, age, or religion.
12
-
13
- Examples of unacceptable behavior by participants include the use of sexual
14
- language or imagery, derogatory comments or personal attacks, trolling, public
15
- or private harassment, insults, or other unprofessional conduct.
16
-
17
- Project maintainers have the right and responsibility to remove, edit, or
18
- reject comments, commits, code, wiki edits, issues, and other contributions
19
- that are not aligned to this Code of Conduct. Project maintainers or
20
- contributors who do not follow the Code of Conduct may be removed from the
21
- project team.
22
-
23
- Instances of abusive, harassing, or otherwise unacceptable behavior may be
24
- reported by opening an issue or contacting the maintainer via email.
25
-
26
- This Code of Conduct is adapted from the [Contributor Covenant][cc], [version
27
- 1.0.0][v1].
28
-
29
- [cc]: http://contributor-covenant.org/
30
- [v1]: http://contributor-covenant.org/version/1/0/0/