minimap2 0.2.30.3 → 1.2.31.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (72) hide show
  1. checksums.yaml +4 -4
  2. data/README.md +48 -63
  3. data/ext/minimap2/Makefile +1 -1
  4. data/ext/minimap2/format.c +12 -2
  5. data/ext/minimap2/hit.c +19 -8
  6. data/ext/minimap2/ksw2_ll_sse.c +2 -2
  7. data/ext/minimap2/minimap.h +2 -1
  8. data/ext/minimap2/pe.c +7 -2
  9. data/ext/ruby_minimap2/extconf.rb +47 -0
  10. data/ext/ruby_minimap2/native_main.c +15 -0
  11. data/ext/ruby_minimap2/ruby_minimap2.c +1134 -0
  12. data/lib/minimap2/aligner.rb +66 -374
  13. data/lib/minimap2/alignment.rb +22 -11
  14. data/lib/minimap2/version.rb +2 -1
  15. data/lib/minimap2.rb +39 -110
  16. metadata +8 -89
  17. data/ext/Rakefile +0 -60
  18. data/ext/cmappy/cmappy.c +0 -134
  19. data/ext/cmappy/cmappy.h +0 -46
  20. data/ext/minimap2/FAQ.md +0 -46
  21. data/ext/minimap2/MANIFEST.in +0 -10
  22. data/ext/minimap2/Makefile.simde +0 -97
  23. data/ext/minimap2/NEWS.md +0 -989
  24. data/ext/minimap2/README.md +0 -428
  25. data/ext/minimap2/code_of_conduct.md +0 -30
  26. data/ext/minimap2/cookbook.md +0 -243
  27. data/ext/minimap2/minimap2.1 +0 -844
  28. data/ext/minimap2/misc/README.md +0 -180
  29. data/ext/minimap2/misc/pafcluster.js +0 -241
  30. data/ext/minimap2/misc/paftools.js +0 -3734
  31. data/ext/minimap2/pyproject.toml +0 -2
  32. data/ext/minimap2/python/README.rst +0 -198
  33. data/ext/minimap2/python/cmappy.h +0 -152
  34. data/ext/minimap2/python/cmappy.pxd +0 -156
  35. data/ext/minimap2/python/mappy.pyx +0 -289
  36. data/ext/minimap2/python/minimap2.py +0 -41
  37. data/ext/minimap2/setup.py +0 -55
  38. data/ext/minimap2/test/MT-human.fa +0 -278
  39. data/ext/minimap2/test/MT-orang.fa +0 -276
  40. data/ext/minimap2/test/q-inv.fa +0 -4
  41. data/ext/minimap2/test/q2.fa +0 -2
  42. data/ext/minimap2/test/t-inv.fa +0 -127
  43. data/ext/minimap2/test/t2.fa +0 -2
  44. data/ext/minimap2/test/x3s-aln.txt +0 -5
  45. data/ext/minimap2/test/x3s-qry.fa +0 -5
  46. data/ext/minimap2/test/x3s-ref.fa +0 -10
  47. data/ext/minimap2/tex/Makefile +0 -21
  48. data/ext/minimap2/tex/bioinfo.cls +0 -930
  49. data/ext/minimap2/tex/blasr-mc.eval +0 -17
  50. data/ext/minimap2/tex/bowtie2-s3.sam.eval +0 -28
  51. data/ext/minimap2/tex/bwa-s3.sam.eval +0 -52
  52. data/ext/minimap2/tex/bwa.eval +0 -55
  53. data/ext/minimap2/tex/eval2roc.pl +0 -33
  54. data/ext/minimap2/tex/graphmap.eval +0 -4
  55. data/ext/minimap2/tex/hs38-simu.sh +0 -10
  56. data/ext/minimap2/tex/minialign.eval +0 -49
  57. data/ext/minimap2/tex/minimap2.bib +0 -460
  58. data/ext/minimap2/tex/minimap2.tex +0 -724
  59. data/ext/minimap2/tex/mm2-s3.sam.eval +0 -62
  60. data/ext/minimap2/tex/mm2-update.tex +0 -240
  61. data/ext/minimap2/tex/mm2.approx.eval +0 -12
  62. data/ext/minimap2/tex/mm2.eval +0 -13
  63. data/ext/minimap2/tex/natbib.bst +0 -1288
  64. data/ext/minimap2/tex/natbib.sty +0 -803
  65. data/ext/minimap2/tex/ngmlr.eval +0 -38
  66. data/ext/minimap2/tex/roc.gp +0 -60
  67. data/ext/minimap2/tex/snap-s3.sam.eval +0 -62
  68. data/ext/minimap2.patch +0 -19
  69. data/lib/minimap2/ffi/constants.rb +0 -267
  70. data/lib/minimap2/ffi/functions.rb +0 -239
  71. data/lib/minimap2/ffi/mappy.rb +0 -104
  72. data/lib/minimap2/ffi.rb +0 -27
@@ -1,844 +0,0 @@
1
- .TH minimap2 1 "15 June 2025" "minimap2-2.30 (r1287)" "Bioinformatics tools"
2
- .SH NAME
3
- .PP
4
- minimap2 - mapping and alignment between collections of DNA sequences
5
- .SH SYNOPSIS
6
- * Indexing the target sequences (optional):
7
- .RS 4
8
- minimap2
9
- .RB [ -x
10
- .IR preset ]
11
- .B -d
12
- .I target.mmi
13
- .I target.fa
14
- .br
15
- minimap2
16
- .RB [ -H ]
17
- .RB [ -k
18
- .IR kmer ]
19
- .RB [ -w
20
- .IR miniWinSize ]
21
- .RB [ -I
22
- .IR batchSize ]
23
- .B -d
24
- .I target.mmi
25
- .I target.fa
26
- .RE
27
-
28
- * Long-read alignment with CIGAR:
29
- .RS 4
30
- minimap2
31
- .B -a
32
- .RB [ -x
33
- .IR preset ]
34
- .I target.mmi
35
- .I query.fa
36
- >
37
- .I output.sam
38
- .br
39
- minimap2
40
- .B -c
41
- .RB [ -H ]
42
- .RB [ -k
43
- .IR kmer ]
44
- .RB [ -w
45
- .IR miniWinSize ]
46
- .RB [ ... ]
47
- .I target.fa
48
- .I query.fa
49
- >
50
- .I output.paf
51
- .RE
52
-
53
- * Long-read overlap without CIGAR:
54
- .RS 4
55
- minimap2
56
- .B -x
57
- ava-ont
58
- .RB [ -t
59
- .IR nThreads ]
60
- .I target.fa
61
- .I query.fa
62
- >
63
- .I output.paf
64
- .RE
65
- .SH DESCRIPTION
66
- .PP
67
- Minimap2 is a fast sequence mapping and alignment program that can find
68
- overlaps between long noisy reads, or map long reads or their assemblies to a
69
- reference genome optionally with detailed alignment (i.e. CIGAR). At present,
70
- it works efficiently with query sequences from a few kilobases to ~100
71
- megabases in length at a error rate ~15%. Minimap2 outputs in the PAF or the
72
- SAM format.
73
- .SH OPTIONS
74
- .SS Indexing options
75
- .TP 10
76
- .BI -k \ INT
77
- Minimizer k-mer length [15]
78
- .TP
79
- .BI -w \ INT
80
- Minimizer window size [10]. A minimizer is the smallest k-mer
81
- in a window of w consecutive k-mers.
82
- .TP
83
- .B -H
84
- Use homopolymer-compressed (HPC) minimizers. An HPC sequence is constructed by
85
- contracting homopolymer runs to a single base. An HPC minimizer is a minimizer
86
- on the HPC sequence.
87
- .TP
88
- .BI -I \ NUM
89
- Load at most
90
- .I NUM
91
- target bases into RAM for indexing [8G]. If there are more than
92
- .I NUM
93
- bases in
94
- .IR target.fa ,
95
- minimap2 needs to read
96
- .I query.fa
97
- multiple times to map it against each batch of target sequences. This would create a multi-part index.
98
- .I NUM
99
- may be ending with k/K/m/M/g/G. NB: mapping quality is incorrect given a
100
- multi-part index. See also option
101
- .BR --split-prefix .
102
- .TP
103
- .B --idx-no-seq
104
- Don't store target sequences in the index. It saves disk space and memory but
105
- the index generated with this option will not work with
106
- .B -a
107
- or
108
- .BR -c .
109
- When base-level alignment is not requested, this option is automatically applied.
110
- .TP
111
- .BI -d \ FILE
112
- Save the minimizer index of
113
- .I target.fa
114
- to
115
- .I FILE
116
- [no dump]. Minimap2 indexing is fast. It can index the human genome in a couple
117
- of minutes. If even shorter startup time is desired, use this option to save
118
- the index. Indexing options are fixed in the index file. When an index file is
119
- provided as the target sequences, options
120
- .BR -H ,
121
- .BR -k ,
122
- .BR -w ,
123
- .B -I
124
- will be effectively overridden by the options stored in the index file.
125
- .TP
126
- .BI --alt \ FILE
127
- List of ALT contigs [null]
128
- .TP
129
- .BI --alt-drop \ FLOAT
130
- Drop ALT hits by
131
- .I FLOAT
132
- fraction when ranking and computing mapping quality [0.15]
133
- .SS Mapping options
134
- .TP 10
135
- .BI -f \ FLOAT | INT1 [, INT2 ]
136
- If fraction, ignore top
137
- .I FLOAT
138
- fraction of most frequent minimizers [0.0002]. If integer,
139
- ignore minimizers occuring more than
140
- .I INT1
141
- times.
142
- .I INT2
143
- is only effective in the
144
- .B --sr
145
- or
146
- .B -xsr
147
- mode, which sets the threshold for a second round of seeding.
148
- .TP
149
- .BI -U \ INT1 [, INT2 ]
150
- Lower and upper bounds of k-mer occurrences [10,1000000]. The final k-mer occurrence threshold is
151
- .RI max{ INT1 ,\ min{ INT2 ,
152
- .BR -f }}.
153
- This option prevents excessively small or large
154
- .B -f
155
- estimated from the input reference. Available since r1034 and deprecating
156
- .B --min-occ-floor
157
- in earlier versions of minimap2.
158
- .TP
159
- .BI --q-occ-frac \ FLOAT
160
- Discard a query minimizer if its occurrence is higher than
161
- .I FLOAT
162
- fraction of query minimizers and than the reference occurrence threshold
163
- [0.01]. Set 0 to disable. Available since r1105.
164
- .TP
165
- .BI -e \ INT
166
- Sample a high-frequency minimizer every
167
- .I INT
168
- basepairs [500].
169
- .TP
170
- .BI -g \ NUM
171
- Stop chain enlongation if there are no minimizers within
172
- .IR NUM -bp
173
- [10k].
174
- .TP
175
- .BI -r \ NUM1 [, NUM2 ]
176
- Bandwidth for chaining and base alignment [500,20k].
177
- .I NUM1
178
- is used for initial chaining and alignment extension;
179
- .I NUM2
180
- for RMQ-based re-chaining and closing gaps in alignments.
181
- .TP
182
- .BI -n \ INT
183
- Discard chains consisting of
184
- .RI < INT
185
- number of minimizers [3]
186
- .TP
187
- .BI -m \ INT
188
- Discard chains with chaining score
189
- .RI < INT
190
- [40]. Chaining score equals the approximate number of matching bases minus a
191
- concave gap penalty. It is computed with dynamic programming.
192
- .TP
193
- .B -D
194
- If query sequence name/length are identical to the target name/length, ignore
195
- diagonal anchors. This option also reduces DP-based extension along the
196
- diagonal.
197
- .TP
198
- .B -P
199
- Retain all chains and don't attempt to set primary chains. Options
200
- .B -p
201
- and
202
- .B -N
203
- have no effect when this option is in use.
204
- .TP
205
- .BR --dual = yes | no
206
- If
207
- .BR no ,
208
- skip query-target pairs wherein the query name is lexicographically greater
209
- than the target name [yes]
210
- .TP
211
- .B -X
212
- Equivalent to
213
- .RB ' -DP
214
- .BR --dual = no
215
- .BR --no-long-join '.
216
- Primarily used for all-vs-all read overlapping.
217
- .TP
218
- .BI -p \ FLOAT
219
- Minimal secondary-to-primary score ratio to output secondary mappings [0.8].
220
- Between two chains overlaping over half of the shorter chain (controlled by
221
- .BR -M ),
222
- the chain with a lower score is secondary to the chain with a higher score.
223
- If the ratio of the scores is below
224
- .IR FLOAT ,
225
- the secondary chain will not be outputted or extended with DP alignment later.
226
- This option has no effect when
227
- .B -X
228
- is applied.
229
- .TP
230
- .BI -N \ INT
231
- Output at most
232
- .I INT
233
- secondary alignments [5]. This option has no effect when
234
- .B -X
235
- is applied.
236
- .TP
237
- .BI -G \ NUM
238
- Maximum gap on the reference (effective with
239
- .BR -xsplice / --splice ).
240
- This option also changes the chaining and alignment band width to
241
- .IR NUM .
242
- Increasing this option slows down spliced alignment. [200k]
243
- .TP
244
- .BI -F \ NUM
245
- Maximum fragment length (aka insert size; effective with
246
- .BR -xsr / --frag = yes )
247
- [800]
248
- .TP
249
- .BI -M \ FLOAT
250
- Mark as secondary a chain that overlaps with a better chain by
251
- .I FLOAT
252
- or more of the shorter chain [0.5]
253
- .TP
254
- .BR --rmq = no | yes
255
- Use the minigraph chaining algorithm [no]. The minigraph algorithm is better
256
- for aligning contigs through long INDELs.
257
- .TP
258
- .BI --rmq-inner \ NUM
259
- Apply full dynamic programming for anchors within distance
260
- .I NUM
261
- [1000].
262
- .TP
263
- .B --hard-mask-level
264
- Honor option
265
- .B -M
266
- and disable a heurstic to save unmapped subsequences and disables
267
- .BR --mask-len .
268
- .TP
269
- .BI --mask-len \ NUM
270
- Keep an alignment if dropping it leaves an unaligned region on query longer than
271
- .IR INT
272
- [inf]. Effective without
273
- .BR --hard-mask-level .
274
- .TP
275
- .BI --max-chain-skip \ INT
276
- A heuristics that stops chaining early [25]. Minimap2 uses dynamic programming
277
- for chaining. The time complexity is quadratic in the number of seeds. This
278
- option makes minimap2 exits the inner loop if it repeatedly sees seeds already
279
- on chains. Set
280
- .I INT
281
- to a large number to switch off this heurstics.
282
- .TP
283
- .BI --max-chain-iter \ INT
284
- Check up to
285
- .I INT
286
- partial chains during chaining [5000]. This is a heuristic to avoid quadratic
287
- time complexity in the worst case.
288
- .TP
289
- .BI --chain-gap-scale \ FLOAT
290
- Scale of gap cost during chaining [1.0]
291
- .TP
292
- .B --no-long-join
293
- Disable the long gap patching heuristic. When this option is applied, the
294
- maximum alignment gap is mostly controlled by
295
- .BR -r .
296
- .TP
297
- .B --splice
298
- Enable the splice alignment mode.
299
- .TP
300
- .BR --sr [= no | dna | rna ]
301
- Enable short-read alignment heuristics [no]. If this option is used with no argument,
302
- .RB ` dna '
303
- is set. In the DNA short-read mode, minimap2 applies a second round of chaining
304
- with a higher minimizer occurrence threshold if no good chain is found. In
305
- addition, minimap2 attempts to patch gaps between seeds with ungapped
306
- alignment.
307
- .TP
308
- .BI --split-prefix \ STR
309
- Prefix to create temporary files. Typically used for a multi-part index.
310
- .TP
311
- .BR --frag = no | yes
312
- Whether to enable the fragment mode [no]
313
- .TP
314
- .B --for-only
315
- Only map to the forward strand of the reference sequences. For paired-end
316
- reads in the forward-reverse orientation, the first read is mapped to forward
317
- strand of the reference and the second read to the reverse stand.
318
- .TP
319
- .B --rev-only
320
- Only map to the reverse complement strand of the reference sequences.
321
- .TP
322
- .BR --heap-sort = no | yes
323
- If yes, sort anchors with heap merge, instead of radix sort. Heap merge is
324
- faster for short reads, but slower for long reads. [no]
325
- .TP
326
- .B --no-hash-name
327
- Produce the same alignment for identical sequences regardless of their sequence names.
328
- .SS Alignment options
329
- .TP 10
330
- .BI -A \ INT
331
- Matching score [2]
332
- .TP
333
- .BI -B \ INT
334
- Mismatching penalty [4]
335
- .TP
336
- .BI -b \ INT
337
- Mismatching penalty for transitions [same as
338
- .BR -B ].
339
- .TP
340
- .BI -O \ INT1[,INT2]
341
- Gap open penalty [4,24]. If
342
- .I INT2
343
- is not specified, it is set to
344
- .IR INT1 .
345
- .TP
346
- .BI -E \ INT1[,INT2]
347
- Gap extension penalty [2,1]. A gap of length
348
- .I k
349
- costs
350
- .RI min{ O1 + k * E1 , O2 + k * E2 }.
351
- In the splice mode, the second gap penalties are not used.
352
- .TP
353
- .BI -J \ INT
354
- Splice model [1]. 0 for the original minimap2 splice model that always penalizes non-GT-AG splicing;
355
- 1 for the miniprot model that considers non-GT-AG. Option
356
- .B -C
357
- has no effect with the default
358
- .BR -J1 .
359
- .TP
360
- .BR -j \ FILE
361
- Junctions used to extend alignment towards ends of reads [].
362
- .I FILE
363
- can be gene annotations in the BED12 format (aka 12-column BED), or intron
364
- positions in 5-column BED with the strand column required. BED12 file can be
365
- converted from GTF/GFF3 with `paftools.js gff2bed anno.gtf'. This option is
366
- intended for short RNA-seq reads, while
367
- .B --junc-bed
368
- for long noisy RNA-seq reads.
369
- .TP
370
- .BI -C \ INT
371
- Cost for a non-canonical GT-AG splicing (effective with
372
- .B --splice
373
- .BR -J0 )
374
- [0].
375
- .TP
376
- .BI -z \ INT1[,INT2]
377
- Truncate an alignment if the running alignment score drops too quickly along
378
- the diagonal of the DP matrix (diagonal X-drop, or Z-drop) [400,200]. If the
379
- drop of score is above
380
- .IR INT2 ,
381
- minimap2 will reverse complement the query in the related region and align
382
- again to test small inversions. Minimap2 truncates alignment if there is an
383
- inversion or the drop of score is greater than
384
- .IR INT1 .
385
- Decrease
386
- .I INT2
387
- to find small inversions at the cost of performance and false positives.
388
- Increase
389
- .I INT1
390
- to improves the contiguity of alignment at the cost of poor alignment in the
391
- middle.
392
- .TP
393
- .BI -s \ INT
394
- Minimal peak DP alignment score to output [40]. The peak score is computed from
395
- the final CIGAR. It is the score of the max scoring segment in the alignment
396
- and may be different from the total alignment score.
397
- .TP
398
- .BI -u \ CHAR
399
- How to find canonical splicing sites GT-AG -
400
- .BR f :
401
- transcript strand;
402
- .BR b :
403
- both strands;
404
- .BR n :
405
- no attempt to match GT-AG [n]
406
- .TP
407
- .BI --end-bonus \ INT
408
- Score bonus when alignment extends to the end of the query sequence [0].
409
- .TP
410
- .BI --score-N \ INT
411
- Penalty of a mismatch involving ambiguous bases [1].
412
- .TP
413
- .BR --pairing = strong | weak | no
414
- How to pair paired-end reads [strong].
415
- .RB ` no '
416
- for aligning the two ends in a pair independently with no `properly paired' set.
417
- .RB ` weak '
418
- for aligning the two ends independently and then pairing the hits.
419
- .RB ` strong '
420
- for jointly aligning and pairing the two ends.
421
- .TP
422
- .BR --splice-flank = yes | no
423
- Assume the next base to a
424
- .B GT
425
- donor site tends to be A/G (91% in human and 92% in mouse) and the preceding
426
- base to a
427
- .B AG
428
- acceptor tends to be C/T [no].
429
- This trend is evolutionarily conservative, all the way to S. cerevisiae
430
- (PMID:18688272). Specifying this option generally leads to higher junction
431
- accuracy by several percents, so it is applied by default with
432
- .BR --splice .
433
- However, the SIRV control does not honor this trend
434
- (only ~60%). This option reduces accuracy. If you are benchmarking minimap2
435
- on SIRV data, please add
436
- .B --splice-flank=no
437
- to the command line.
438
- .TP
439
- .BR --spsc \ FILE
440
- Splice scores []. Each line consists of five fields: 1) contig, 2) offset, 3) `+' or `-', 4) `D' or `A', and 5) score,
441
- where offset is the number of bases before a splice junction, `D' indicates the
442
- line corresponds to a donor site and `A' for an acceptor site.
443
- A positive score suggests the junction is preferred and a negative score
444
- suggests the junction is not preferred.
445
- .TP
446
- .BR --spsc0 \ INT
447
- Penalty for positions not in
448
- .I FILE
449
- specified by
450
- .B --spsc
451
- [5]. Effective with
452
- .B --spsc
453
- but not
454
- .BR --junc-bed .
455
- .TP
456
- .BR --spsc-scale \ FLOAT
457
- Scale splice scores in
458
- .B --spsc
459
- by
460
- .IR FLOAT
461
- rounded to the nearest integer [0.7].
462
- .TP
463
- .BR --junc-bed \ FILE
464
- Junctions to prefer during base alignment [].
465
- Same format as
466
- .BR -j .
467
- It is
468
- .I NOT
469
- recommended to apply this option to short RNA-seq reads. This would increase
470
- run time with little improvement to junction accuracy.
471
- .TP
472
- .BR --junc-bonus \ INT
473
- Score bonus for a splice donor or acceptor found in annotation [9]. Effective with
474
- .B --junc-bed
475
- but not
476
- .BR --spsc .
477
- .TP
478
- .BR --jump-min-match \ INT
479
- Minimum matching length to create a jump [3]. Equivalent to
480
- .B STAR
481
- .BR --alignSJDBoverhangMin .
482
- .TP
483
- .BI --end-seed-pen \ INT
484
- Drop a terminal anchor if
485
- .IR s <log( g )+ INT ,
486
- where
487
- .I s
488
- is the local alignment score around the anchor and
489
- .I g
490
- the length of the terminal gap in the chain. This option is only effective
491
- with
492
- .BR --splice .
493
- It helps to avoid tiny terminal exons. [6]
494
- .TP
495
- .B --no-end-flt
496
- Don't filter seeds towards the ends of chains before performing base-level
497
- alignment.
498
- .TP
499
- .BI --cap-sw-mem \ NUM
500
- Skip alignment if the DP matrix size is above
501
- .IR NUM .
502
- Set 0 to disable [100m].
503
- .TP
504
- .BI --cap-kalloc \ NUM
505
- Free thread-local kalloc memory reservoir if after the alignment the size of the reservoir above
506
- .IR NUM .
507
- Set 0 to disable [500m].
508
- .SS Input/output options
509
- .TP 10
510
- .B -a
511
- Generate CIGAR and output alignments in the SAM format. Minimap2 outputs in PAF
512
- by default.
513
- .TP
514
- .BI -o \ FILE
515
- Output alignments to
516
- .I FILE
517
- [stdout].
518
- .TP
519
- .B -Q
520
- Ignore base quality in the input file.
521
- .TP
522
- .B -L
523
- Write CIGAR with >65535 operators at the CG tag. Older tools are unable to
524
- convert alignments with >65535 CIGAR ops to BAM. This option makes minimap2 SAM
525
- compatible with older tools. Newer tools recognizes this tag and reconstruct
526
- the real CIGAR in memory.
527
- .TP
528
- .BI -R \ STR
529
- SAM read group line in a format like
530
- .B @RG\\\\tID:foo\\\\tSM:bar
531
- [].
532
- .TP
533
- .B -y
534
- Copy input FASTA/Q comments to output.
535
- .TP
536
- .B -c
537
- Generate CIGAR. In PAF, the CIGAR is written to the `cg' custom tag.
538
- .TP
539
- .BR --cs [= short | long ]
540
- Output the
541
- .B cs
542
- tag.
543
- If no argument is given,
544
- .RB ` short '
545
- is set. [none]
546
- .TP
547
- .B --MD
548
- Output the MD tag (see the SAM spec).
549
- .TP
550
- .B --eqx
551
- Output =/X CIGAR operators for sequence match/mismatch.
552
- .TP
553
- .B -Y
554
- In SAM output, use soft clipping for supplementary alignments.
555
- .TP
556
- .B --secondary-seq
557
- In SAM output, show query sequences for secondary alignments.
558
- .TP
559
- .B --write-junc
560
- Output splice junctions in 6-column BED: contig name, start, end,
561
- read name, score and strand. Score is the sum of donor and acceptor scores,
562
- where GT gets 3, GC gets 2 and AT gets 1 at donor sites,
563
- while AG gets 3 and AC gets 1 at acceptor sites.
564
- Alignments with mapping quality below 10 are ignored.
565
- .TP
566
- .BI --pass1 \ FILE
567
- Junctions BED file outputted by
568
- .B --write-junc
569
- []. Rows with scores lower than 5 are ignored. When both
570
- .B -j
571
- and
572
- .B --pass1
573
- are present, junctions in
574
- .B -j
575
- are preferred over in
576
- .BR --pass1
577
- when there is ambiguity.
578
- .TP
579
- .BI --seed \ INT
580
- Integer seed for randomizing equally best hits. Minimap2 hashes
581
- .I INT
582
- and read name when choosing between equally best hits. [11]
583
- .TP
584
- .BI -t \ INT
585
- Number of threads [3]. Minimap2 uses at most three threads when indexing target
586
- sequences, and uses up to
587
- .IR INT +1
588
- threads when mapping (the extra thread is for I/O, which is frequently idle and
589
- takes little CPU time).
590
- .TP
591
- .B -2
592
- Use two I/O threads during mapping. By default, minimap2 uses one I/O thread.
593
- When I/O is slow (e.g. piping to gzip, or reading from a slow pipe), the I/O
594
- thread may become the bottleneck. Apply this option to use one thread for input
595
- and another thread for output, at the cost of increased peak RAM.
596
- .TP
597
- .BI -K \ NUM
598
- Number of bases loaded into memory to process in a mini-batch [500M].
599
- Similar to option
600
- .BR -I ,
601
- K/M/G/k/m/g suffix is accepted. A large
602
- .I NUM
603
- helps load balancing in the multi-threading mode, at the cost of increased
604
- memory.
605
- .TP
606
- .BR --secondary = yes | no
607
- Whether to output secondary alignments [yes]
608
- .TP
609
- .BI --max-qlen \ NUM
610
- Filter out query sequences longer than
611
- .IR NUM .
612
- .TP
613
- .B --paf-no-hit
614
- In PAF, output unmapped queries; the strand and the reference name fields are
615
- set to `*'. Warning: some paftools.js commands may not work with such output
616
- for the moment.
617
- .TP
618
- .B --sam-hit-only
619
- In SAM, don't output unmapped reads.
620
- .TP
621
- .B --version
622
- Print version number to stdout
623
- .SS Preset options
624
- .TP 10
625
- .BI -x \ STR
626
- Preset []. This option applies multiple options at the same time. It should be
627
- applied before other options because options applied later will overwrite the
628
- values set by
629
- .BR -x .
630
- Available
631
- .I STR
632
- are:
633
- .RS
634
- .TP 10
635
- .B map-ont
636
- Align noisy long reads of ~10% error rate to a reference genome. This is the
637
- default mode.
638
- .TP
639
- .B lr:hq
640
- Align accurate long reads (error rate <1%) to a reference genome
641
- .RB ( -k19
642
- .B -w19 -U50,500
643
- .BR -g10k ).
644
- This was recommended by ONT developers for recent Nanopore reads
645
- produced with chemistry v14 that can reach ~99% in accuracy.
646
- It was shown to work better for accurate Nanopore reads
647
- than
648
- .BR map-hifi .
649
- .TP
650
- .B map-hifi
651
- Align PacBio high-fidelity (HiFi) reads to a reference genome
652
- .RB ( -xlr:hq
653
- .B -A1 -B4 -O6,26 -E2,1
654
- .BR -s200 ).
655
- It differs from
656
- .B lr:hq
657
- only in scoring. It has not been tested whether
658
- .B lr:hq
659
- would work better for PacBio HiFi reads.
660
- .TP
661
- .B map-pb
662
- Align older PacBio continuous long (CLR) reads to a reference genome
663
- .RB ( -Hk19 ).
664
- Note that this data type is effectively deprecated by HiFi.
665
- Unless you work on very old data, you probably want to use
666
- .B map-hifi
667
- or
668
- .BR lr:hq .
669
- .TP
670
- .B map-iclr
671
- Align Illumina Complete Long Reads (ICLR) to a reference genome
672
- .RB ( -k19
673
- .B -B6 -b4
674
- .BR -O10,50 ).
675
- This was recommended by Illumina developers.
676
- .TP
677
- .B asm5
678
- Long assembly to reference mapping
679
- .RB ( -k19
680
- .B -w19 -U50,500 --rmq -r1k,100k -g10k -A1 -B19 -O39,81 -E3,1 -s200 -z200
681
- .BR -N50 ).
682
- Typically, the alignment will not extend to regions with 5% or higher sequence
683
- divergence. Use this preset if the average divergence is not much higher than 0.1%.
684
- .TP
685
- .B asm10
686
- Long assembly to reference mapping
687
- .RB ( -k19
688
- .B -w19 -U50,500 --rmq -r1k,100k -g10k -A1 -B9 -O16,41 -E2,1 -s200 -z200
689
- .BR -N50 ).
690
- Use this if the average divergence is around 1%.
691
- .TP
692
- .B asm20
693
- Long assembly to reference mapping
694
- .RB ( -k19
695
- .B -w10 -U50,500 --rmq -r1k,100k -g10k -A1 -B4 -O6,26 -E2,1 -s200 -z200
696
- .BR -N50 ).
697
- Use this if the average divergence is around several percent.
698
- .TP
699
- .B splice
700
- Long-read spliced alignment
701
- .RB ( -k15
702
- .B -w5 --splice -g2k -G200k -A1 -B2 -O2,32 -E1,0 -C9 -z200 -ub --junc-bonus=9 --cap-sw-mem=0
703
- .BR --splice-flank=yes ).
704
- In the splice mode, 1) long deletions are taken as introns and represented as
705
- the
706
- .RB ` N '
707
- CIGAR operator; 2) long insertions are disabled; 3) deletion and insertion gap
708
- costs are different during chaining; 4) the computation of the
709
- .RB ` ms '
710
- tag ignores introns to demote hits to pseudogenes.
711
- .TP
712
- .B splice:hq
713
- Spliced alignment for accurate long RNA-seq reads such as PacBio iso-seq
714
- .RB ( -xsplice
715
- .B -C5 -O6,24
716
- .BR -B4 ).
717
- .TP
718
- .B splice:sr
719
- Spliced alignment for short RNA-seq reads
720
- .RB ( -xsplice:hq
721
- .B --frag=yes -m25 -s40 -2K100m --heap-sort=yes --pairing=weak --sr=rna --min-dp-len=20
722
- .BR --secondary=no ).
723
- .TP
724
- .B sr
725
- Short-read alignment without splicing
726
- .RB ( -k21
727
- .B -w11 --sr --frag=yes -A2 -B8 -O12,32 -E2,1 -r100 -p.5 -N20 -f1000,5000 -n2 -m25
728
- .B -s40 -g100 -2K50m --heap-sort=yes
729
- .BR --secondary=no ).
730
- .TP
731
- .B ava-pb
732
- PacBio CLR all-vs-all overlap mapping
733
- .RB ( -Hk19
734
- .B -Xw5 -e0
735
- .BR -m100 ).
736
- .TP
737
- .B ava-ont
738
- Oxford Nanopore all-vs-all overlap mapping
739
- .RB ( -k15
740
- .B -Xw5 -e0 -m100
741
- .BR -r2k ).
742
- .RE
743
- .SS Miscellaneous options
744
- .TP 10
745
- .B --no-kalloc
746
- Use the libc default allocator instead of the kalloc thread-local allocator.
747
- This debugging option is mostly used with Valgrind to detect invalid memory
748
- accesses. Minimap2 runs slower with this option, especially in the
749
- multi-threading mode.
750
- .TP
751
- .B --print-qname
752
- Print query names to stderr, mostly to see which query is crashing minimap2.
753
- .TP
754
- .B --print-seeds
755
- Print seed positions to stderr, for debugging only.
756
- .SH OUTPUT FORMAT
757
- .PP
758
- Minimap2 outputs mapping positions in the Pairwise mApping Format (PAF) by
759
- default. PAF is a TAB-delimited text format with each line consisting of at
760
- least 12 fields as are described in the following table:
761
- .TS
762
- center box;
763
- cb | cb | cb
764
- r | c | l .
765
- Col Type Description
766
- _
767
- 1 string Query sequence name
768
- 2 int Query sequence length
769
- 3 int Query start coordinate (0-based)
770
- 4 int Query end coordinate (0-based)
771
- 5 char `+' if query/target on the same strand; `-' if opposite
772
- 6 string Target sequence name
773
- 7 int Target sequence length
774
- 8 int Target start coordinate on the original strand
775
- 9 int Target end coordinate on the original strand
776
- 10 int Number of matching bases in the mapping
777
- 11 int Number bases, including gaps, in the mapping
778
- 12 int Mapping quality (0-255 with 255 for missing)
779
- .TE
780
-
781
- .PP
782
- When alignment is available, column 11 gives the total number of sequence
783
- matches, mismatches and gaps in the alignment; column 10 divided by column 11
784
- gives the BLAST-like alignment identity. When alignment is unavailable,
785
- these two columns are approximate. PAF may optionally have additional fields in
786
- the SAM-like typed key-value format. Minimap2 may output the following tags:
787
- .TS
788
- center box;
789
- cb | cb | cb
790
- r | c | l .
791
- Tag Type Description
792
- _
793
- tp A Type of aln: P/primary, S/secondary and I,i/inversion
794
- cm i Number of minimizers on the chain
795
- s1 i Chaining score
796
- s2 i Chaining score of the best secondary chain
797
- NM i Total number of mismatches and gaps in the alignment
798
- MD Z To generate the ref sequence in the alignment
799
- AS i DP alignment score
800
- SA Z List of other supplementary alignments (with approximate CIGAR strings)
801
- ms i DP score of the max scoring segment in the alignment
802
- nn i Number of ambiguous bases in the alignment
803
- ts A Transcript strand (splice mode only)
804
- cg Z CIGAR string (only in PAF)
805
- cs Z Difference string
806
- dv f Approximate per-base sequence divergence
807
- de f Gap-compressed per-base sequence divergence
808
- rl i Length of query regions harboring repetitive seeds
809
- zd i Alignment broken due to Z-drop; bit 1: left broken; bit 2: right broken
810
- .TE
811
-
812
- .PP
813
- The
814
- .B cs
815
- tag encodes difference sequences in the short form or the entire query
816
- .I AND
817
- reference sequences in the long form. It consists of a series of operations:
818
- .TS
819
- center box;
820
- cb | cb |cb
821
- r | l | l .
822
- Op Regex Description
823
- _
824
- = [ACGTN]+ Identical sequence (long form)
825
- : [0-9]+ Identical sequence length
826
- * [acgtn][acgtn] Substitution: ref to query
827
- + [acgtn]+ Insertion to the reference
828
- - [acgtn]+ Deletion from the reference
829
- ~ [acgtn]{2}[0-9]+[acgtn]{2} Intron length and splice signal
830
- .TE
831
-
832
- .SH LIMITATIONS
833
- .TP 2
834
- *
835
- Minimap2 may produce suboptimal alignments through long low-complexity regions
836
- where seed positions may be suboptimal. This should not be a big concern
837
- because even the optimal alignment may be wrong in such regions.
838
- .TP
839
- *
840
- Minimap2 requires SSE2 or NEON instructions to compile. It is possible to add
841
- non-SSE2/NEON support, but it would make minimap2 slower by several times.
842
- .SH SEE ALSO
843
- .PP
844
- miniasm(1), minimap(1), bwa(1).