@gmod/bam 8.5.0 → 8.5.1
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +49 -226
- package/package.json +1 -1
package/README.md
CHANGED
|
@@ -21,18 +21,51 @@ const records = await bam.getRecordsForRange('ctgA', 0, 50000)
|
|
|
21
21
|
```
|
|
22
22
|
|
|
23
23
|
Coordinates are 0-based half-open (not the same as `samtools view` inputs).
|
|
24
|
-
`bamPath` reads a local file, so it is node-only; in the browser pass a
|
|
25
|
-
filehandle
|
|
24
|
+
`bamPath` reads a local file, so it is node-only; in the browser pass a URL or a
|
|
25
|
+
generic-filehandle2 filehandle instead:
|
|
26
26
|
|
|
27
27
|
```typescript
|
|
28
|
-
import { BamFile } from '@gmod/bam'
|
|
29
|
-
|
|
30
28
|
const bam = new BamFile({
|
|
31
29
|
bamUrl: 'https://example.com/yourfile.bam',
|
|
32
30
|
baiUrl: 'https://example.com/yourfile.bam.bai',
|
|
33
31
|
})
|
|
34
32
|
```
|
|
35
33
|
|
|
34
|
+
Records come back unfiltered, and are shared between overlapping queries — treat
|
|
35
|
+
them as read-only. Filter them yourself with the flag helpers and `getTag`,
|
|
36
|
+
which decodes one tag instead of all of them:
|
|
37
|
+
|
|
38
|
+
```typescript
|
|
39
|
+
const records = (await bam.getRecordsForRange('chr1', 0, 100000)).filter(
|
|
40
|
+
r => r.isProperlyPaired() && !r.isSecondary() && r.getTag('RG') === 'rg1',
|
|
41
|
+
)
|
|
42
|
+
```
|
|
43
|
+
|
|
44
|
+
## Decompressing on a worker pool
|
|
45
|
+
|
|
46
|
+
BGZF decompression is 70-90% of a cold query, and BGZF blocks are independently
|
|
47
|
+
inflatable. Hand `BamFile` a
|
|
48
|
+
[`@gmod/bgzf-filehandle`](https://github.com/GMOD/bgzf-filehandle) worker pool
|
|
49
|
+
and it inflates chunks there instead of on the calling thread — measured
|
|
50
|
+
2.7-4.1x on the pool's own fixtures.
|
|
51
|
+
|
|
52
|
+
```typescript
|
|
53
|
+
import { getSharedWorkerPool } from '@gmod/bgzf-filehandle'
|
|
54
|
+
|
|
55
|
+
const bam = new BamFile({
|
|
56
|
+
bamUrl: 'https://example.com/yourfile.bam',
|
|
57
|
+
// the pending promise is fine — it is awaited at the point of use
|
|
58
|
+
bgzfWorkerPool: getSharedWorkerPool(),
|
|
59
|
+
})
|
|
60
|
+
```
|
|
61
|
+
|
|
62
|
+
No cross-origin isolation is needed. `getSharedWorkerPool()` gives back
|
|
63
|
+
`undefined` under node, or anywhere Workers cannot be created, which keeps the
|
|
64
|
+
in-process path — so this is safe to pass unconditionally. bam-js never creates
|
|
65
|
+
a pool on its own: the thread budget belongs to the consumer. For worker counts,
|
|
66
|
+
lifecycle and the pool's own benchmarks, see
|
|
67
|
+
[bgzf-filehandle's worker pool docs](https://github.com/GMOD/bgzf-filehandle/blob/main/docs/worker-pool.md).
|
|
68
|
+
|
|
36
69
|
## Usage with htsget
|
|
37
70
|
|
|
38
71
|
```typescript
|
|
@@ -46,10 +79,8 @@ const records = await bam.getRecordsForRange('1', 2000000, 2000001)
|
|
|
46
79
|
```
|
|
47
80
|
|
|
48
81
|
htsget fetches the server's range as-is, so `viewAsPairs`, `pairAcrossChr` and
|
|
49
|
-
`maxInsertSize` are ignored.
|
|
50
|
-
|
|
51
|
-
For a server that requires authentication, pass a `fetch` that adds the bearer
|
|
52
|
-
token the spec calls for:
|
|
82
|
+
`maxInsertSize` are ignored. For a server that requires authentication, pass a
|
|
83
|
+
`fetch` that adds the bearer token:
|
|
53
84
|
|
|
54
85
|
```typescript
|
|
55
86
|
const bam = new HtsgetFile({
|
|
@@ -63,227 +94,19 @@ const bam = new HtsgetFile({
|
|
|
63
94
|
})
|
|
64
95
|
```
|
|
65
96
|
|
|
66
|
-
Your `fetch` is called for the ticket request
|
|
67
|
-
ticket points at,
|
|
68
|
-
|
|
69
|
-
in each url's own `headers` field, which is applied either way.
|
|
70
|
-
|
|
71
|
-
## Documentation
|
|
72
|
-
|
|
73
|
-
### BamFile constructor
|
|
74
|
-
|
|
75
|
-
- `bamPath`/`bamUrl`/`bamFilehandle` - local path, remote URL, or a
|
|
76
|
-
generic-filehandle2 object
|
|
77
|
-
- `baiPath`/`baiUrl`/`baiFilehandle` - BAI index. Defaults to the `.bai` sibling
|
|
78
|
-
of `bamPath`/`bamUrl`
|
|
79
|
-
- `csiPath`/`csiUrl`/`csiFilehandle` - CSI index, required for chromosomes
|
|
80
|
-
longer than 2^29
|
|
81
|
-
- `renameRefSeqs` - `(refName: string) => string` applied to header ref names
|
|
82
|
-
- `recordClass` - custom class extending `BamRecord` (see below)
|
|
83
|
-
- `maxCacheBytes` - ceiling for the parsed-chunk cache, in decompressed bytes.
|
|
84
|
-
default: 1GB, `Infinity` for none. See [Caching](#caching)
|
|
85
|
-
- `cacheIdleTimeoutMs` - drop a cached chunk once nothing has read it for this
|
|
86
|
-
long. default: 3 minutes; `0` disables it. See [Caching](#caching)
|
|
87
|
-
|
|
88
|
-
The `path`/`url` forms are convenience wrappers for generic-filehandle2's
|
|
89
|
-
`LocalFile` and `RemoteFile`.
|
|
90
|
-
|
|
91
|
-
### Caching
|
|
92
|
-
|
|
93
|
-
Parsed chunks — the unit the BAM index hands out — are cached, so overlapping
|
|
94
|
-
and adjacent queries reuse decompressed records instead of re-fetching them. Two
|
|
95
|
-
options bound that cache, and they answer different questions.
|
|
96
|
-
|
|
97
|
-
**`maxCacheBytes` is a ceiling under load, not a limit on what you can ask
|
|
98
|
-
for.** Nothing is ever refused for being too large: a chunk bigger than the
|
|
99
|
-
whole budget is still cached, reads in flight are never evicted, and eviction
|
|
100
|
-
only drops a value that has already been returned once. The worst a budget can
|
|
101
|
-
cost you is a re-read. It can make a query slower; it can never make one fail or
|
|
102
|
-
come back short.
|
|
103
|
-
|
|
104
|
-
**It binds less often than its size suggests.** On the deepest data we measure —
|
|
105
|
-
1000x coverage long reads, 240 windows over six laps — the cache settles at
|
|
106
|
-
573MB across 60 entries and eviction never runs at the 1GB default. Treat it as
|
|
107
|
-
a backstop against a session that pans forever, not as an operating constraint.
|
|
108
|
-
|
|
109
|
-
**Don't pick a number between one query and several.** Below one query's working
|
|
110
|
-
set the cache inverts: each chunk is evicted before the next pan can reuse it,
|
|
111
|
-
so the hit rate is zero, the full re-decompress is paid every time, and the
|
|
112
|
-
unevictable entries retain the memory anyway. At a 200MB budget on that same
|
|
113
|
-
file the cache holds exactly one entry, because one chunk there decompresses to
|
|
114
|
-
181MB. Either size it above the working set or pass `Infinity` and bound memory
|
|
115
|
-
some other way.
|
|
116
|
-
|
|
117
|
-
**`cacheIdleTimeoutMs` is the only thing that gives memory back.**
|
|
118
|
-
`maxCacheBytes` is enforced when a read settles, so an idle cache sits at
|
|
119
|
-
whatever it reached and never lowers — and for a page that holds a `BamFile` for
|
|
120
|
-
the life of a track, that resting level is the number that actually matters. The
|
|
121
|
-
idle sweep is what makes a generous ceiling affordable, by turning it into a
|
|
122
|
-
peak under panning rather than a level a parked tab holds indefinitely. The
|
|
123
|
-
clock runs from the last _read_ of a chunk, or from its parse landing if nothing
|
|
124
|
-
has read it since, so panning back and forth over one region never expires it
|
|
125
|
-
and a slow chunk still gets the full timeout to be reused in. Measured on a pan
|
|
126
|
-
that held 331MB: 0MB once idle.
|
|
127
|
-
|
|
128
|
-
**Neither bounds peak memory**, and on deep data the gap is not small:
|
|
129
|
-
|
|
130
|
-
- Up to six chunk reads run at once and an in-flight read is never evicted. At
|
|
131
|
-
181MB a chunk, that is ~476MB in flight before retention holds anything.
|
|
132
|
-
- A query holds every chunk it parsed until it returns, whether or not the cache
|
|
133
|
-
still does.
|
|
134
|
-
- An entry is weighed once, when its read settles, and records grow after that.
|
|
135
|
-
`end`, `CIGAR` and `tags` each memoize onto the record the first time they are
|
|
136
|
-
read, which is what a renderer does to every visible read — measured at +38%
|
|
137
|
-
over the weighed size.
|
|
138
|
-
|
|
139
|
-
So these are bounds on retained decompressed bytes, not on the heap. Size
|
|
140
|
-
against what you want to keep, and bound total memory at a level that can see
|
|
141
|
-
the whole process.
|
|
142
|
-
|
|
143
|
-
### HtsgetFile constructor
|
|
144
|
-
|
|
145
|
-
- `baseUrl` - htsget reads endpoint, e.g. `https://htsget.example.com/reads`
|
|
146
|
-
- `trackId` - id of the resource under `baseUrl`
|
|
147
|
-
- `fetch` - `fetch` replacement for adding auth headers (see above)
|
|
148
|
-
- `recordClass` - custom class extending `BamRecord` (see below)
|
|
149
|
-
|
|
150
|
-
### async getRecordsForRange(refName, start, end, opts?)
|
|
151
|
-
|
|
152
|
-
- `refName` - chromosome to fetch from
|
|
153
|
-
- `start`/`end` - 0-based half-open coordinates
|
|
154
|
-
- `opts.signal` - `AbortSignal` to stop processing
|
|
155
|
-
- `opts.viewAsPairs` - re-dispatch requests to find mate pairs. default: false
|
|
156
|
-
- `opts.pairAcrossChr` - let `viewAsPairs` pair across chromosomes. default:
|
|
157
|
-
false
|
|
158
|
-
- `opts.maxInsertSize` - distance limit for `viewAsPairs` within a chromosome.
|
|
159
|
-
default: 200kb
|
|
160
|
-
- `opts.onProgress` - `(bytesDownloaded, totalBytes?) => void`, called per BGZF
|
|
161
|
-
chunk for a determinate progress bar
|
|
162
|
-
|
|
163
|
-
Returned records are cached and shared between overlapping queries, so treat
|
|
164
|
-
them as read-only — attaching your own fields to a record mutates it for every
|
|
165
|
-
other query holding it.
|
|
166
|
-
|
|
167
|
-
Records come back unfiltered. Filter them yourself with the flag helpers and
|
|
168
|
-
`getTag`, which decodes one tag instead of all of them:
|
|
169
|
-
|
|
170
|
-
```typescript
|
|
171
|
-
const records = (await bam.getRecordsForRange('chr1', 0, 100000)).filter(
|
|
172
|
-
r => r.isProperlyPaired() && !r.isSecondary() && r.getTag('RG') === 'rg1',
|
|
173
|
-
)
|
|
174
|
-
```
|
|
175
|
-
|
|
176
|
-
### async getHeader(opts?)
|
|
177
|
-
|
|
178
|
-
Returns the parsed SAM header. Called automatically by the query methods and
|
|
179
|
-
cached, so you only need it when you want the header itself.
|
|
180
|
-
`getHeaderText(opts?)` returns the raw header string.
|
|
181
|
-
|
|
182
|
-
### async indexCov(refName, start?, end?)
|
|
183
|
-
|
|
184
|
-
Returns `{start, end, score}` features estimating read density over 16kb
|
|
185
|
-
windows, derived from the BAI linear index. CSI has no linear index, so a
|
|
186
|
-
CSI-indexed file returns `[]`.
|
|
97
|
+
Your `fetch` is called for the ticket request _and_ for the data-block urls the
|
|
98
|
+
ticket points at, which may live on a third-party host — so only attach
|
|
99
|
+
credentials to hosts you trust.
|
|
187
100
|
|
|
188
|
-
|
|
189
|
-
|
|
190
|
-
Number of records on `refName` from the index's pseudo-bin (bin 37450 in BAI,
|
|
191
|
-
`n_mapped` in the SAM spec), or 0 if `refName` is absent.
|
|
192
|
-
|
|
193
|
-
### async hasRefSeq(refName)
|
|
194
|
-
|
|
195
|
-
Whether `refName` is present in the file.
|
|
196
|
-
|
|
197
|
-
### async estimatedBytesForRegions(regions, opts?)
|
|
198
|
-
|
|
199
|
-
Compressed bytes the given `{refName, start, end}[]` would fetch — useful for
|
|
200
|
-
warning before a large query.
|
|
201
|
-
|
|
202
|
-
### clearFeatureCache()
|
|
203
|
-
|
|
204
|
-
Drops the parsed-chunk cache immediately, rather than waiting for
|
|
205
|
-
`maxCacheBytes` or `cacheIdleTimeoutMs` to reclaim it. See [Caching](#caching).
|
|
206
|
-
|
|
207
|
-
### BamRecord
|
|
208
|
-
|
|
209
|
-
```typescript
|
|
210
|
-
// Core alignment fields
|
|
211
|
-
record.fileOffset // "file offset" based id -- not a true file offset
|
|
212
|
-
record.ref_id // numerical sequence id from SAM header
|
|
213
|
-
record.start // 0-based start coordinate
|
|
214
|
-
record.end // 0-based end coordinate
|
|
215
|
-
record.name // QNAME
|
|
216
|
-
record.seq // sequence string
|
|
217
|
-
record.qual // Uint8Array of quality scores (null if SEQ is empty)
|
|
218
|
-
record.CIGAR // CIGAR string e.g. "50M2I48M"
|
|
219
|
-
record.flags // SAM flags integer
|
|
220
|
-
record.mq // mapping quality (undefined if 255)
|
|
221
|
-
record.strand // 1 or -1
|
|
222
|
-
record.template_length // TLEN
|
|
101
|
+
## Docs
|
|
223
102
|
|
|
224
|
-
|
|
225
|
-
record
|
|
226
|
-
|
|
227
|
-
|
|
228
|
-
|
|
229
|
-
|
|
230
|
-
record.getTag('MD') // one tag, without decoding the rest
|
|
231
|
-
record.getTagRaw('MD') // string tag as Uint8Array, skipping string conversion
|
|
232
|
-
|
|
233
|
-
// Typed-array views, for rendering without allocating strings
|
|
234
|
-
record.NUMERIC_MD // MD tag as Uint8Array
|
|
235
|
-
record.NUMERIC_CIGAR // Uint32Array of packed CIGAR operations
|
|
236
|
-
record.NUMERIC_SEQ // Uint8Array of 4-bit encoded sequence
|
|
237
|
-
|
|
238
|
-
// Flag methods
|
|
239
|
-
record.isPaired()
|
|
240
|
-
record.isProperlyPaired()
|
|
241
|
-
record.isSegmentUnmapped()
|
|
242
|
-
record.isMateUnmapped()
|
|
243
|
-
record.isReverseComplemented()
|
|
244
|
-
record.isMateReverseComplemented()
|
|
245
|
-
record.isRead1()
|
|
246
|
-
record.isRead2()
|
|
247
|
-
record.isSecondary()
|
|
248
|
-
record.isFailedQc()
|
|
249
|
-
record.isDuplicate()
|
|
250
|
-
record.isSupplementary()
|
|
251
|
-
|
|
252
|
-
// Utility
|
|
253
|
-
record.seqAt(idx) // single base at position
|
|
254
|
-
record.toJSON()
|
|
255
|
-
```
|
|
256
|
-
|
|
257
|
-
### Custom BamRecord class
|
|
258
|
-
|
|
259
|
-
```typescript
|
|
260
|
-
import { BamFile, BamRecord } from '@gmod/bam'
|
|
261
|
-
|
|
262
|
-
class CustomBamRecord extends BamRecord {
|
|
263
|
-
get customProperty() {
|
|
264
|
-
return `custom-${this.name}`
|
|
265
|
-
}
|
|
266
|
-
}
|
|
267
|
-
|
|
268
|
-
const bam = new BamFile<CustomBamRecord>({
|
|
269
|
-
bamPath: 'test.bam',
|
|
270
|
-
recordClass: CustomBamRecord,
|
|
271
|
-
})
|
|
272
|
-
|
|
273
|
-
// records are typed as CustomBamRecord[]
|
|
274
|
-
const records = await bam.getRecordsForRange('ctgA', 0, 50000)
|
|
275
|
-
console.log(records[0].customProperty)
|
|
276
|
-
```
|
|
103
|
+
- [docs/api.md](docs/api.md) — every constructor option, method and `BamRecord`
|
|
104
|
+
field, plus custom record classes
|
|
105
|
+
- [docs/caching.md](docs/caching.md) — sizing the parsed-chunk cache
|
|
106
|
+
- [agent-docs/adr/](agent-docs/adr/) — the measurements behind the performance
|
|
107
|
+
and caching decisions
|
|
108
|
+
- [CONTRIBUTING.md](CONTRIBUTING.md) — development and release steps
|
|
277
109
|
|
|
278
110
|
## License
|
|
279
111
|
|
|
280
112
|
MIT © [Colin Diesh](https://github.com/cmdcolin)
|
|
281
|
-
|
|
282
|
-
## Publishing
|
|
283
|
-
|
|
284
|
-
[Trusted publishing](https://docs.npmjs.com/about-trusted-publishing) via GitHub
|
|
285
|
-
Actions.
|
|
286
|
-
|
|
287
|
-
```bash
|
|
288
|
-
pnpm version patch # or minor/major
|
|
289
|
-
```
|