@gmod/tabix 3.8.1 → 3.8.2

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -40,6 +40,23 @@ const custom = new TabixIndexedFile({
40
40
  })
41
41
  ```
42
42
 
43
+ Over HTTP, swapping in
44
+ [`@gmod/range-cache-filehandle`](https://github.com/GMOD/range-cache-filehandle)
45
+ is usually worth it: a query reads the index and then a scattered set of BGZF
46
+ blocks, and a byte-range cache coalesces those into one request per contiguous
47
+ run and serves an overlapping query from memory.
48
+
49
+ ```typescript
50
+ import { RemoteFileWithRangeCache } from '@gmod/range-cache-filehandle'
51
+
52
+ const cached = new TabixIndexedFile({
53
+ filehandle: new RemoteFileWithRangeCache('https://example.com/file.vcf.gz'),
54
+ tbiFilehandle: new RemoteFileWithRangeCache(
55
+ 'https://example.com/file.vcf.gz.tbi',
56
+ ),
57
+ })
58
+ ```
59
+
43
60
  ### getLines
44
61
 
45
62
  Fetches lines overlapping a region. `start`/`end` are 0-based half-open
@@ -73,19 +90,19 @@ await file.getLines('chr1', 200, 300, {
73
90
  })
74
91
  ```
75
92
 
76
- `onProgress` ticks once per compressed block, including instant ticks for cache
77
- hits, and `totalBytes` is known up front from the index — enough for a
78
- determinate progress bar.
93
+ `onProgress` ticks once per chunk — the run of BGZF blocks the index resolves a
94
+ query to — including instant ticks for chunks already cached, and the index
95
+ supplies `totalBytes` up front, which is enough for a determinate progress bar.
79
96
 
80
97
  Notes:
81
98
 
82
- - Meta/comment lines are skipped
99
+ - The scan skips meta/comment lines
83
100
  - Line strings have no trailing whitespace
84
101
  - Pass `undefined` for `end` to read to the end of the contig
85
102
  - A `refName` that is not in the index yields no lines and no error, so a
86
103
  `chr1`/`1` naming mismatch looks like an empty region. Check against
87
- [`getReferenceSequenceNames`](#getreferencesequencenamesopts-promisestring) if
88
- a query comes back unexpectedly empty
104
+ [`getReferenceSequenceNames`](docs/api.md#getreferencesequencenamesopts-promisestring)
105
+ if a query comes back unexpectedly empty
89
106
  - `start > end` throws a `TypeError`; `start === end` returns without reading
90
107
 
91
108
  ### Without NPM (CDN)
@@ -98,93 +115,77 @@ See [example/index.html](example/index.html) for a working demo. It fetches the
98
115
  VCF over HTTP, so serve the directory (e.g. `npx serve example`) rather than
99
116
  opening the file directly.
100
117
 
101
- ## API
102
-
103
- ### `new TabixIndexedFile(args)`
104
-
105
- | Arg | Type | Description |
106
- | ------------------------- | -------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
107
- | `path` | `string?` | Local file path |
108
- | `url` | `string?` | Remote URL |
109
- | `filehandle` | `GenericFilehandle?` | Custom filehandle (from [generic-filehandle2](https://github.com/GMOD/generic-filehandle2)) |
110
- | `tbiPath` | `string?` | TBI index path (defaults to `path + '.tbi'`) |
111
- | `tbiUrl` | `string?` | TBI index URL |
112
- | `tbiFilehandle` | `GenericFilehandle?` | TBI index filehandle |
113
- | `csiPath` | `string?` | CSI index path |
114
- | `csiUrl` | `string?` | CSI index URL |
115
- | `csiFilehandle` | `GenericFilehandle?` | CSI index filehandle |
116
- | `chunkCacheSize` | `number?` | Chunk LRU cache budget, in _decompressed_ bytes (default 1 GiB). A retention bound, not a bound on peak memory. Size it to hold several queries: below one query's working set the hit rate drops to zero while the memory is retained anyway |
117
- | `chunkCacheIdleTimeoutMs` | `number?` | Drop a cached chunk once nothing has read it for this long (default 3 minutes, `0` disables). The only thing that lowers the cache while nothing is happening, and what makes the budget above a peak rather than a resting level |
118
-
119
- ### `getLines(refName, start, end, opts)`
120
-
121
- Calls the line callback for each line overlapping `[start, end)`. `start`
122
- defaults to `0` and `end` to the end of the contig when `undefined`. `opts` is
123
- either the callback itself or an object:
124
-
125
- | Option | Type | Description |
126
- | -------------- | -------------------------------------------------------- | --------------------------------------- |
127
- | `lineCallback` | `(line, fileOffset, start, end) => void` | Required |
128
- | `signal` | `AbortSignal?` | Aborts the in-flight reads |
129
- | `onProgress` | `(bytesDownloaded: number, totalBytes?: number) => void` | Called as compressed blocks are fetched |
130
-
131
- ### `getHeader(opts?): Promise<string>`
132
-
133
- Returns all comment/meta lines before the first data line as a string, matching
134
- what `tabix -H` prints. A header row that is not commented is therefore not
135
- included, even when the index counted it as a line to skip — see
136
- `getSkippedLines`.
118
+ ## How a query flows
137
119
 
138
- ### `getHeaderBuffer(opts?): Promise<Uint8Array>`
120
+ `getLines` turns a region into BGZF chunks through the index and decompresses
121
+ each one in wasm — index reads included, since `.tbi` and `.csi` are bgzipped
122
+ too. The rest is ordinary JS: it matches lines as bytes and decodes only the
123
+ ones you asked for. [docs/dataflow.md](docs/dataflow.md) has the diagram and
124
+ walks it through.
139
125
 
140
- Returns the header as raw bytes.
126
+ The file then holds on to those decompressed chunks, so overlapping and adjacent
127
+ queries reuse them instead of inflating again — up to 1GB per file, dropped
128
+ after three idle minutes. A consumer holding one file per track should bound
129
+ them together with a shared `chunkCacheBudget` rather than shrinking each file's
130
+ own ceiling: [docs/caching.md](docs/caching.md).
141
131
 
142
- ### `getSkippedLines(opts?): Promise<string[]>`
132
+ ## Decompressing on a worker pool
143
133
 
144
- Returns the leading lines the index says to skip (`tabix -S N`), or `[]` when it
145
- records none. This is where a file whose header row is not commented keeps it,
146
- which PLINK `.ld`, bedGraph and BED deflines routinely are.
134
+ BGZF blocks inflate independently, so that decompression can spread across
135
+ threads.
147
136
 
148
- Separate from `getHeader` because htslib treats the two differently: a line is
149
- not data when the index's skip count covers it **or** it starts with the meta
150
- character, but `tabix -H` prints only the latter.
151
-
152
- ### `getHeaderLines(opts?): Promise<string[]>`
153
-
154
- Returns the file's header lines however the file keeps them: the commented block
155
- when there is one, and otherwise the rows the index counted. Empty lines are
156
- dropped.
157
-
158
- This is usually the one you want. Reading `getHeader` alone cannot tell a file
159
- that has no header from one whose header is not commented — both come back as
160
- the empty string — so callers fall back to an assumed column layout and quietly
161
- mis-name columns. Both forms are parsed from a single read of the file's leading
162
- blocks, which is also shared with `getHeader` and `getSkippedLines`.
163
-
164
- ### `getReferenceSequenceNames(opts?): Promise<string[]>`
165
-
166
- Returns reference sequence names in index order.
167
-
168
- ### `lineCount(refName, opts?): Promise<number>`
169
-
170
- Returns the number of data lines on the given reference, or `-1` if the
171
- reference is not in the index.
172
-
173
- ### `bytesForRegions(regions, opts?): Promise<number>`
174
-
175
- Estimates the compressed byte size of index chunks covering the given regions.
176
- Useful for deciding whether a request is too large before calling `getLines`.
137
+ ```typescript
138
+ import { getSharedWorkerPool } from '@gmod/bgzf-filehandle'
177
139
 
178
- ## Contributing
140
+ const file = new TabixIndexedFile({
141
+ url: 'https://example.com/yourfile.vcf.gz',
142
+ // the pending promise is fine — it is awaited at the point of use
143
+ bgzfWorkerPool: getSharedWorkerPool(),
144
+ })
145
+ ```
179
146
 
180
- See [CONTRIBUTING.md](CONTRIBUTING.md) for development and release steps.
147
+ Safe to pass unconditionally: `getSharedWorkerPool()` returns `undefined` under
148
+ node, or anywhere the host forbids Workers, which keeps the in-process path. No
149
+ cross-origin isolation needed. tabix-js never creates a pool on its own — the
150
+ thread budget belongs to the consumer.
151
+
152
+ **Worth about 1.4x here, against the 1.95x a BAM reader reports.** Measured in
153
+ jbrowse-components on `test/data/1kg.chr1.subset.vcf.gz` — 213MB of 1000
154
+ Genomes, headless Chrome, real HTTP, four workers, arms interleaved, both
155
+ returning the same record count: **1.34-1.46x** across five window sizes and a
156
+ twelve-step pan.
157
+
158
+ The decompression itself moves **1.83x**. What holds the end-to-end figure below
159
+ that is a **28% floor of per-line byte scanning and string decoding**, which no
160
+ worker count reaches — and that floor is at its worst on multi-sample VCF, whose
161
+ records carry a genotype field per sample and run to ~60KB a line. A format with
162
+ narrower lines sits closer to BAM. If you want more than ~1.5x on a multi-sample
163
+ VCF, the scan is what is left to attack, not the decompression.
164
+
165
+ Worker counts, lifecycle and benchmarks:
166
+ [bgzf-filehandle's worker pool docs](https://github.com/GMOD/bgzf-filehandle/blob/main/docs/worker-pool.md);
167
+ the end-to-end numbers above, and how to confirm a pool is really engaging in
168
+ production rather than quietly falling back, are in jbrowse-components'
169
+ [BGZF_WORKER_POOL.md](https://github.com/GMOD/jbrowse-components/blob/main/agent-docs/reference/BGZF_WORKER_POOL.md).
170
+
171
+ ## Docs
172
+
173
+ - [docs/api.md](docs/api.md) — every constructor arg and method
174
+ - [docs/dataflow.md](docs/dataflow.md) — a query end to end, diagrammed
175
+ - [docs/optimizations.md](docs/optimizations.md) — why each step of that path
176
+ looks the way it does, and what measured it
177
+ - [docs/caching.md](docs/caching.md) — sizing the decompressed-chunk cache, and
178
+ bounding many files together
179
+ - [agent-docs/adr/](agent-docs/adr/) — the measurements behind those decisions
180
+ - [agent-docs/TODO.md](agent-docs/TODO.md) — what is worth doing next, and what
181
+ has to be measured before it
182
+ - [CONTRIBUTING.md](CONTRIBUTING.md) — development and release steps
181
183
 
182
184
  ## Academic Use
183
185
 
184
- This package was written with funding from the [NHGRI](http://genome.gov) as
185
- part of the [JBrowse](http://jbrowse.org) project. If you use it in an academic
186
- project that you publish, please cite the most recent JBrowse paper, which will
187
- be linked from [jbrowse.org](http://jbrowse.org).
186
+ Written with [NHGRI](http://genome.gov) funding as part of
187
+ [JBrowse](http://jbrowse.org). If you use this in a publication, please cite the
188
+ most recent JBrowse paper at [jbrowse.org](http://jbrowse.org).
188
189
 
189
190
  ## License
190
191