@gmod/bam 8.5.0 → 8.5.1

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (2) hide show
  1. package/README.md +49 -226
  2. package/package.json +1 -1
package/README.md CHANGED
@@ -21,18 +21,51 @@ const records = await bam.getRecordsForRange('ctgA', 0, 50000)
21
21
  ```
22
22
 
23
23
  Coordinates are 0-based half-open (not the same as `samtools view` inputs).
24
- `bamPath` reads a local file, so it is node-only; in the browser pass a
25
- filehandle or URL instead:
24
+ `bamPath` reads a local file, so it is node-only; in the browser pass a URL or a
25
+ generic-filehandle2 filehandle instead:
26
26
 
27
27
  ```typescript
28
- import { BamFile } from '@gmod/bam'
29
-
30
28
  const bam = new BamFile({
31
29
  bamUrl: 'https://example.com/yourfile.bam',
32
30
  baiUrl: 'https://example.com/yourfile.bam.bai',
33
31
  })
34
32
  ```
35
33
 
34
+ Records come back unfiltered, and are shared between overlapping queries — treat
35
+ them as read-only. Filter them yourself with the flag helpers and `getTag`,
36
+ which decodes one tag instead of all of them:
37
+
38
+ ```typescript
39
+ const records = (await bam.getRecordsForRange('chr1', 0, 100000)).filter(
40
+ r => r.isProperlyPaired() && !r.isSecondary() && r.getTag('RG') === 'rg1',
41
+ )
42
+ ```
43
+
44
+ ## Decompressing on a worker pool
45
+
46
+ BGZF decompression is 70-90% of a cold query, and BGZF blocks are independently
47
+ inflatable. Hand `BamFile` a
48
+ [`@gmod/bgzf-filehandle`](https://github.com/GMOD/bgzf-filehandle) worker pool
49
+ and it inflates chunks there instead of on the calling thread — measured
50
+ 2.7-4.1x on the pool's own fixtures.
51
+
52
+ ```typescript
53
+ import { getSharedWorkerPool } from '@gmod/bgzf-filehandle'
54
+
55
+ const bam = new BamFile({
56
+ bamUrl: 'https://example.com/yourfile.bam',
57
+ // the pending promise is fine — it is awaited at the point of use
58
+ bgzfWorkerPool: getSharedWorkerPool(),
59
+ })
60
+ ```
61
+
62
+ No cross-origin isolation is needed. `getSharedWorkerPool()` gives back
63
+ `undefined` under node, or anywhere Workers cannot be created, which keeps the
64
+ in-process path — so this is safe to pass unconditionally. bam-js never creates
65
+ a pool on its own: the thread budget belongs to the consumer. For worker counts,
66
+ lifecycle and the pool's own benchmarks, see
67
+ [bgzf-filehandle's worker pool docs](https://github.com/GMOD/bgzf-filehandle/blob/main/docs/worker-pool.md).
68
+
36
69
  ## Usage with htsget
37
70
 
38
71
  ```typescript
@@ -46,10 +79,8 @@ const records = await bam.getRecordsForRange('1', 2000000, 2000001)
46
79
  ```
47
80
 
48
81
  htsget fetches the server's range as-is, so `viewAsPairs`, `pairAcrossChr` and
49
- `maxInsertSize` are ignored.
50
-
51
- For a server that requires authentication, pass a `fetch` that adds the bearer
52
- token the spec calls for:
82
+ `maxInsertSize` are ignored. For a server that requires authentication, pass a
83
+ `fetch` that adds the bearer token:
53
84
 
54
85
  ```typescript
55
86
  const bam = new HtsgetFile({
@@ -63,227 +94,19 @@ const bam = new HtsgetFile({
63
94
  })
64
95
  ```
65
96
 
66
- Your `fetch` is called for the ticket request and for the data-block urls the
67
- ticket points at, so only attach credentials to hosts you trust — data blocks
68
- may live on a third-party host, and the spec has servers put whatever those need
69
- in each url's own `headers` field, which is applied either way.
70
-
71
- ## Documentation
72
-
73
- ### BamFile constructor
74
-
75
- - `bamPath`/`bamUrl`/`bamFilehandle` - local path, remote URL, or a
76
- generic-filehandle2 object
77
- - `baiPath`/`baiUrl`/`baiFilehandle` - BAI index. Defaults to the `.bai` sibling
78
- of `bamPath`/`bamUrl`
79
- - `csiPath`/`csiUrl`/`csiFilehandle` - CSI index, required for chromosomes
80
- longer than 2^29
81
- - `renameRefSeqs` - `(refName: string) => string` applied to header ref names
82
- - `recordClass` - custom class extending `BamRecord` (see below)
83
- - `maxCacheBytes` - ceiling for the parsed-chunk cache, in decompressed bytes.
84
- default: 1GB, `Infinity` for none. See [Caching](#caching)
85
- - `cacheIdleTimeoutMs` - drop a cached chunk once nothing has read it for this
86
- long. default: 3 minutes; `0` disables it. See [Caching](#caching)
87
-
88
- The `path`/`url` forms are convenience wrappers for generic-filehandle2's
89
- `LocalFile` and `RemoteFile`.
90
-
91
- ### Caching
92
-
93
- Parsed chunks — the unit the BAM index hands out — are cached, so overlapping
94
- and adjacent queries reuse decompressed records instead of re-fetching them. Two
95
- options bound that cache, and they answer different questions.
96
-
97
- **`maxCacheBytes` is a ceiling under load, not a limit on what you can ask
98
- for.** Nothing is ever refused for being too large: a chunk bigger than the
99
- whole budget is still cached, reads in flight are never evicted, and eviction
100
- only drops a value that has already been returned once. The worst a budget can
101
- cost you is a re-read. It can make a query slower; it can never make one fail or
102
- come back short.
103
-
104
- **It binds less often than its size suggests.** On the deepest data we measure —
105
- 1000x coverage long reads, 240 windows over six laps — the cache settles at
106
- 573MB across 60 entries and eviction never runs at the 1GB default. Treat it as
107
- a backstop against a session that pans forever, not as an operating constraint.
108
-
109
- **Don't pick a number between one query and several.** Below one query's working
110
- set the cache inverts: each chunk is evicted before the next pan can reuse it,
111
- so the hit rate is zero, the full re-decompress is paid every time, and the
112
- unevictable entries retain the memory anyway. At a 200MB budget on that same
113
- file the cache holds exactly one entry, because one chunk there decompresses to
114
- 181MB. Either size it above the working set or pass `Infinity` and bound memory
115
- some other way.
116
-
117
- **`cacheIdleTimeoutMs` is the only thing that gives memory back.**
118
- `maxCacheBytes` is enforced when a read settles, so an idle cache sits at
119
- whatever it reached and never lowers — and for a page that holds a `BamFile` for
120
- the life of a track, that resting level is the number that actually matters. The
121
- idle sweep is what makes a generous ceiling affordable, by turning it into a
122
- peak under panning rather than a level a parked tab holds indefinitely. The
123
- clock runs from the last _read_ of a chunk, or from its parse landing if nothing
124
- has read it since, so panning back and forth over one region never expires it
125
- and a slow chunk still gets the full timeout to be reused in. Measured on a pan
126
- that held 331MB: 0MB once idle.
127
-
128
- **Neither bounds peak memory**, and on deep data the gap is not small:
129
-
130
- - Up to six chunk reads run at once and an in-flight read is never evicted. At
131
- 181MB a chunk, that is ~476MB in flight before retention holds anything.
132
- - A query holds every chunk it parsed until it returns, whether or not the cache
133
- still does.
134
- - An entry is weighed once, when its read settles, and records grow after that.
135
- `end`, `CIGAR` and `tags` each memoize onto the record the first time they are
136
- read, which is what a renderer does to every visible read — measured at +38%
137
- over the weighed size.
138
-
139
- So these are bounds on retained decompressed bytes, not on the heap. Size
140
- against what you want to keep, and bound total memory at a level that can see
141
- the whole process.
142
-
143
- ### HtsgetFile constructor
144
-
145
- - `baseUrl` - htsget reads endpoint, e.g. `https://htsget.example.com/reads`
146
- - `trackId` - id of the resource under `baseUrl`
147
- - `fetch` - `fetch` replacement for adding auth headers (see above)
148
- - `recordClass` - custom class extending `BamRecord` (see below)
149
-
150
- ### async getRecordsForRange(refName, start, end, opts?)
151
-
152
- - `refName` - chromosome to fetch from
153
- - `start`/`end` - 0-based half-open coordinates
154
- - `opts.signal` - `AbortSignal` to stop processing
155
- - `opts.viewAsPairs` - re-dispatch requests to find mate pairs. default: false
156
- - `opts.pairAcrossChr` - let `viewAsPairs` pair across chromosomes. default:
157
- false
158
- - `opts.maxInsertSize` - distance limit for `viewAsPairs` within a chromosome.
159
- default: 200kb
160
- - `opts.onProgress` - `(bytesDownloaded, totalBytes?) => void`, called per BGZF
161
- chunk for a determinate progress bar
162
-
163
- Returned records are cached and shared between overlapping queries, so treat
164
- them as read-only — attaching your own fields to a record mutates it for every
165
- other query holding it.
166
-
167
- Records come back unfiltered. Filter them yourself with the flag helpers and
168
- `getTag`, which decodes one tag instead of all of them:
169
-
170
- ```typescript
171
- const records = (await bam.getRecordsForRange('chr1', 0, 100000)).filter(
172
- r => r.isProperlyPaired() && !r.isSecondary() && r.getTag('RG') === 'rg1',
173
- )
174
- ```
175
-
176
- ### async getHeader(opts?)
177
-
178
- Returns the parsed SAM header. Called automatically by the query methods and
179
- cached, so you only need it when you want the header itself.
180
- `getHeaderText(opts?)` returns the raw header string.
181
-
182
- ### async indexCov(refName, start?, end?)
183
-
184
- Returns `{start, end, score}` features estimating read density over 16kb
185
- windows, derived from the BAI linear index. CSI has no linear index, so a
186
- CSI-indexed file returns `[]`.
97
+ Your `fetch` is called for the ticket request _and_ for the data-block urls the
98
+ ticket points at, which may live on a third-party host — so only attach
99
+ credentials to hosts you trust.
187
100
 
188
- ### async lineCount(refName)
189
-
190
- Number of records on `refName` from the index's pseudo-bin (bin 37450 in BAI,
191
- `n_mapped` in the SAM spec), or 0 if `refName` is absent.
192
-
193
- ### async hasRefSeq(refName)
194
-
195
- Whether `refName` is present in the file.
196
-
197
- ### async estimatedBytesForRegions(regions, opts?)
198
-
199
- Compressed bytes the given `{refName, start, end}[]` would fetch — useful for
200
- warning before a large query.
201
-
202
- ### clearFeatureCache()
203
-
204
- Drops the parsed-chunk cache immediately, rather than waiting for
205
- `maxCacheBytes` or `cacheIdleTimeoutMs` to reclaim it. See [Caching](#caching).
206
-
207
- ### BamRecord
208
-
209
- ```typescript
210
- // Core alignment fields
211
- record.fileOffset // "file offset" based id -- not a true file offset
212
- record.ref_id // numerical sequence id from SAM header
213
- record.start // 0-based start coordinate
214
- record.end // 0-based end coordinate
215
- record.name // QNAME
216
- record.seq // sequence string
217
- record.qual // Uint8Array of quality scores (null if SEQ is empty)
218
- record.CIGAR // CIGAR string e.g. "50M2I48M"
219
- record.flags // SAM flags integer
220
- record.mq // mapping quality (undefined if 255)
221
- record.strand // 1 or -1
222
- record.template_length // TLEN
101
+ ## Docs
223
102
 
224
- // Mate info
225
- record.next_refid
226
- record.next_pos
227
-
228
- // Auxiliary data
229
- record.tags // all aux tags e.g. {MD: "100", NM: 0}
230
- record.getTag('MD') // one tag, without decoding the rest
231
- record.getTagRaw('MD') // string tag as Uint8Array, skipping string conversion
232
-
233
- // Typed-array views, for rendering without allocating strings
234
- record.NUMERIC_MD // MD tag as Uint8Array
235
- record.NUMERIC_CIGAR // Uint32Array of packed CIGAR operations
236
- record.NUMERIC_SEQ // Uint8Array of 4-bit encoded sequence
237
-
238
- // Flag methods
239
- record.isPaired()
240
- record.isProperlyPaired()
241
- record.isSegmentUnmapped()
242
- record.isMateUnmapped()
243
- record.isReverseComplemented()
244
- record.isMateReverseComplemented()
245
- record.isRead1()
246
- record.isRead2()
247
- record.isSecondary()
248
- record.isFailedQc()
249
- record.isDuplicate()
250
- record.isSupplementary()
251
-
252
- // Utility
253
- record.seqAt(idx) // single base at position
254
- record.toJSON()
255
- ```
256
-
257
- ### Custom BamRecord class
258
-
259
- ```typescript
260
- import { BamFile, BamRecord } from '@gmod/bam'
261
-
262
- class CustomBamRecord extends BamRecord {
263
- get customProperty() {
264
- return `custom-${this.name}`
265
- }
266
- }
267
-
268
- const bam = new BamFile<CustomBamRecord>({
269
- bamPath: 'test.bam',
270
- recordClass: CustomBamRecord,
271
- })
272
-
273
- // records are typed as CustomBamRecord[]
274
- const records = await bam.getRecordsForRange('ctgA', 0, 50000)
275
- console.log(records[0].customProperty)
276
- ```
103
+ - [docs/api.md](docs/api.md) — every constructor option, method and `BamRecord`
104
+ field, plus custom record classes
105
+ - [docs/caching.md](docs/caching.md) — sizing the parsed-chunk cache
106
+ - [agent-docs/adr/](agent-docs/adr/) — the measurements behind the performance
107
+ and caching decisions
108
+ - [CONTRIBUTING.md](CONTRIBUTING.md) — development and release steps
277
109
 
278
110
  ## License
279
111
 
280
112
  MIT © [Colin Diesh](https://github.com/cmdcolin)
281
-
282
- ## Publishing
283
-
284
- [Trusted publishing](https://docs.npmjs.com/about-trusted-publishing) via GitHub
285
- Actions.
286
-
287
- ```bash
288
- pnpm version patch # or minor/major
289
- ```
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "@gmod/bam",
3
- "version": "8.5.0",
3
+ "version": "8.5.1",
4
4
  "packageManager": "pnpm@11.15.1",
5
5
  "description": "Parser for BAM and BAM index (bai) files",
6
6
  "license": "MIT",