@gmod/tabix 3.8.0 → 3.8.2
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +84 -83
- package/dist/tabix-bundle.js +1 -1
- package/dist/tabixIndexedFile.d.ts +24 -14
- package/dist/tabixIndexedFile.js.map +1 -1
- package/esm/tabixIndexedFile.d.ts +24 -14
- package/esm/tabixIndexedFile.js.map +1 -1
- package/package.json +3 -3
- package/src/tabixIndexedFile.ts +24 -14
package/README.md
CHANGED
|
@@ -40,6 +40,23 @@ const custom = new TabixIndexedFile({
|
|
|
40
40
|
})
|
|
41
41
|
```
|
|
42
42
|
|
|
43
|
+
Over HTTP, swapping in
|
|
44
|
+
[`@gmod/range-cache-filehandle`](https://github.com/GMOD/range-cache-filehandle)
|
|
45
|
+
is usually worth it: a query reads the index and then a scattered set of BGZF
|
|
46
|
+
blocks, and a byte-range cache coalesces those into one request per contiguous
|
|
47
|
+
run and serves an overlapping query from memory.
|
|
48
|
+
|
|
49
|
+
```typescript
|
|
50
|
+
import { RemoteFileWithRangeCache } from '@gmod/range-cache-filehandle'
|
|
51
|
+
|
|
52
|
+
const cached = new TabixIndexedFile({
|
|
53
|
+
filehandle: new RemoteFileWithRangeCache('https://example.com/file.vcf.gz'),
|
|
54
|
+
tbiFilehandle: new RemoteFileWithRangeCache(
|
|
55
|
+
'https://example.com/file.vcf.gz.tbi',
|
|
56
|
+
),
|
|
57
|
+
})
|
|
58
|
+
```
|
|
59
|
+
|
|
43
60
|
### getLines
|
|
44
61
|
|
|
45
62
|
Fetches lines overlapping a region. `start`/`end` are 0-based half-open
|
|
@@ -73,19 +90,19 @@ await file.getLines('chr1', 200, 300, {
|
|
|
73
90
|
})
|
|
74
91
|
```
|
|
75
92
|
|
|
76
|
-
`onProgress` ticks once per
|
|
77
|
-
|
|
78
|
-
determinate progress bar.
|
|
93
|
+
`onProgress` ticks once per chunk — the run of BGZF blocks the index resolves a
|
|
94
|
+
query to — including instant ticks for chunks already cached, and the index
|
|
95
|
+
supplies `totalBytes` up front, which is enough for a determinate progress bar.
|
|
79
96
|
|
|
80
97
|
Notes:
|
|
81
98
|
|
|
82
|
-
-
|
|
99
|
+
- The scan skips meta/comment lines
|
|
83
100
|
- Line strings have no trailing whitespace
|
|
84
101
|
- Pass `undefined` for `end` to read to the end of the contig
|
|
85
102
|
- A `refName` that is not in the index yields no lines and no error, so a
|
|
86
103
|
`chr1`/`1` naming mismatch looks like an empty region. Check against
|
|
87
|
-
[`getReferenceSequenceNames`](#getreferencesequencenamesopts-promisestring)
|
|
88
|
-
a query comes back unexpectedly empty
|
|
104
|
+
[`getReferenceSequenceNames`](docs/api.md#getreferencesequencenamesopts-promisestring)
|
|
105
|
+
if a query comes back unexpectedly empty
|
|
89
106
|
- `start > end` throws a `TypeError`; `start === end` returns without reading
|
|
90
107
|
|
|
91
108
|
### Without NPM (CDN)
|
|
@@ -98,93 +115,77 @@ See [example/index.html](example/index.html) for a working demo. It fetches the
|
|
|
98
115
|
VCF over HTTP, so serve the directory (e.g. `npx serve example`) rather than
|
|
99
116
|
opening the file directly.
|
|
100
117
|
|
|
101
|
-
##
|
|
102
|
-
|
|
103
|
-
### `new TabixIndexedFile(args)`
|
|
104
|
-
|
|
105
|
-
| Arg | Type | Description |
|
|
106
|
-
| ------------------------- | -------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
|
|
107
|
-
| `path` | `string?` | Local file path |
|
|
108
|
-
| `url` | `string?` | Remote URL |
|
|
109
|
-
| `filehandle` | `GenericFilehandle?` | Custom filehandle (from [generic-filehandle2](https://github.com/GMOD/generic-filehandle2)) |
|
|
110
|
-
| `tbiPath` | `string?` | TBI index path (defaults to `path + '.tbi'`) |
|
|
111
|
-
| `tbiUrl` | `string?` | TBI index URL |
|
|
112
|
-
| `tbiFilehandle` | `GenericFilehandle?` | TBI index filehandle |
|
|
113
|
-
| `csiPath` | `string?` | CSI index path |
|
|
114
|
-
| `csiUrl` | `string?` | CSI index URL |
|
|
115
|
-
| `csiFilehandle` | `GenericFilehandle?` | CSI index filehandle |
|
|
116
|
-
| `chunkCacheSize` | `number?` | Chunk LRU cache budget, in _decompressed_ bytes (default 1 GiB). A retention bound, not a bound on peak memory. Size it to hold several queries: below one query's working set the hit rate drops to zero while the memory is retained anyway |
|
|
117
|
-
| `chunkCacheIdleTimeoutMs` | `number?` | Drop a cached chunk once nothing has read it for this long (default 3 minutes, `0` disables). The only thing that lowers the cache while nothing is happening, and what makes the budget above a peak rather than a resting level |
|
|
118
|
-
|
|
119
|
-
### `getLines(refName, start, end, opts)`
|
|
120
|
-
|
|
121
|
-
Calls the line callback for each line overlapping `[start, end)`. `start`
|
|
122
|
-
defaults to `0` and `end` to the end of the contig when `undefined`. `opts` is
|
|
123
|
-
either the callback itself or an object:
|
|
124
|
-
|
|
125
|
-
| Option | Type | Description |
|
|
126
|
-
| -------------- | -------------------------------------------------------- | --------------------------------------- |
|
|
127
|
-
| `lineCallback` | `(line, fileOffset, start, end) => void` | Required |
|
|
128
|
-
| `signal` | `AbortSignal?` | Aborts the in-flight reads |
|
|
129
|
-
| `onProgress` | `(bytesDownloaded: number, totalBytes?: number) => void` | Called as compressed blocks are fetched |
|
|
130
|
-
|
|
131
|
-
### `getHeader(opts?): Promise<string>`
|
|
132
|
-
|
|
133
|
-
Returns all comment/meta lines before the first data line as a string, matching
|
|
134
|
-
what `tabix -H` prints. A header row that is not commented is therefore not
|
|
135
|
-
included, even when the index counted it as a line to skip — see
|
|
136
|
-
`getSkippedLines`.
|
|
118
|
+
## How a query flows
|
|
137
119
|
|
|
138
|
-
|
|
120
|
+
`getLines` turns a region into BGZF chunks through the index and decompresses
|
|
121
|
+
each one in wasm — index reads included, since `.tbi` and `.csi` are bgzipped
|
|
122
|
+
too. The rest is ordinary JS: it matches lines as bytes and decodes only the
|
|
123
|
+
ones you asked for. [docs/dataflow.md](docs/dataflow.md) has the diagram and
|
|
124
|
+
walks it through.
|
|
139
125
|
|
|
140
|
-
|
|
126
|
+
The file then holds on to those decompressed chunks, so overlapping and adjacent
|
|
127
|
+
queries reuse them instead of inflating again — up to 1GB per file, dropped
|
|
128
|
+
after three idle minutes. A consumer holding one file per track should bound
|
|
129
|
+
them together with a shared `chunkCacheBudget` rather than shrinking each file's
|
|
130
|
+
own ceiling: [docs/caching.md](docs/caching.md).
|
|
141
131
|
|
|
142
|
-
|
|
132
|
+
## Decompressing on a worker pool
|
|
143
133
|
|
|
144
|
-
|
|
145
|
-
|
|
146
|
-
which PLINK `.ld`, bedGraph and BED deflines routinely are.
|
|
134
|
+
BGZF blocks inflate independently, so that decompression can spread across
|
|
135
|
+
threads.
|
|
147
136
|
|
|
148
|
-
|
|
149
|
-
|
|
150
|
-
character, but `tabix -H` prints only the latter.
|
|
151
|
-
|
|
152
|
-
### `getHeaderLines(opts?): Promise<string[]>`
|
|
153
|
-
|
|
154
|
-
Returns the file's header lines however the file keeps them: the commented block
|
|
155
|
-
when there is one, and otherwise the rows the index counted. Empty lines are
|
|
156
|
-
dropped.
|
|
157
|
-
|
|
158
|
-
This is usually the one you want. Reading `getHeader` alone cannot tell a file
|
|
159
|
-
that has no header from one whose header is not commented — both come back as
|
|
160
|
-
the empty string — so callers fall back to an assumed column layout and quietly
|
|
161
|
-
mis-name columns. Both forms are parsed from a single read of the file's leading
|
|
162
|
-
blocks, which is also shared with `getHeader` and `getSkippedLines`.
|
|
163
|
-
|
|
164
|
-
### `getReferenceSequenceNames(opts?): Promise<string[]>`
|
|
165
|
-
|
|
166
|
-
Returns reference sequence names in index order.
|
|
167
|
-
|
|
168
|
-
### `lineCount(refName, opts?): Promise<number>`
|
|
169
|
-
|
|
170
|
-
Returns the number of data lines on the given reference, or `-1` if the
|
|
171
|
-
reference is not in the index.
|
|
172
|
-
|
|
173
|
-
### `bytesForRegions(regions, opts?): Promise<number>`
|
|
174
|
-
|
|
175
|
-
Estimates the compressed byte size of index chunks covering the given regions.
|
|
176
|
-
Useful for deciding whether a request is too large before calling `getLines`.
|
|
137
|
+
```typescript
|
|
138
|
+
import { getSharedWorkerPool } from '@gmod/bgzf-filehandle'
|
|
177
139
|
|
|
178
|
-
|
|
140
|
+
const file = new TabixIndexedFile({
|
|
141
|
+
url: 'https://example.com/yourfile.vcf.gz',
|
|
142
|
+
// the pending promise is fine — it is awaited at the point of use
|
|
143
|
+
bgzfWorkerPool: getSharedWorkerPool(),
|
|
144
|
+
})
|
|
145
|
+
```
|
|
179
146
|
|
|
180
|
-
|
|
147
|
+
Safe to pass unconditionally: `getSharedWorkerPool()` returns `undefined` under
|
|
148
|
+
node, or anywhere the host forbids Workers, which keeps the in-process path. No
|
|
149
|
+
cross-origin isolation needed. tabix-js never creates a pool on its own — the
|
|
150
|
+
thread budget belongs to the consumer.
|
|
151
|
+
|
|
152
|
+
**Worth about 1.4x here, against the 1.95x a BAM reader reports.** Measured in
|
|
153
|
+
jbrowse-components on `test/data/1kg.chr1.subset.vcf.gz` — 213MB of 1000
|
|
154
|
+
Genomes, headless Chrome, real HTTP, four workers, arms interleaved, both
|
|
155
|
+
returning the same record count: **1.34-1.46x** across five window sizes and a
|
|
156
|
+
twelve-step pan.
|
|
157
|
+
|
|
158
|
+
The decompression itself moves **1.83x**. What holds the end-to-end figure below
|
|
159
|
+
that is a **28% floor of per-line byte scanning and string decoding**, which no
|
|
160
|
+
worker count reaches — and that floor is at its worst on multi-sample VCF, whose
|
|
161
|
+
records carry a genotype field per sample and run to ~60KB a line. A format with
|
|
162
|
+
narrower lines sits closer to BAM. If you want more than ~1.5x on a multi-sample
|
|
163
|
+
VCF, the scan is what is left to attack, not the decompression.
|
|
164
|
+
|
|
165
|
+
Worker counts, lifecycle and benchmarks:
|
|
166
|
+
[bgzf-filehandle's worker pool docs](https://github.com/GMOD/bgzf-filehandle/blob/main/docs/worker-pool.md);
|
|
167
|
+
the end-to-end numbers above, and how to confirm a pool is really engaging in
|
|
168
|
+
production rather than quietly falling back, are in jbrowse-components'
|
|
169
|
+
[BGZF_WORKER_POOL.md](https://github.com/GMOD/jbrowse-components/blob/main/agent-docs/reference/BGZF_WORKER_POOL.md).
|
|
170
|
+
|
|
171
|
+
## Docs
|
|
172
|
+
|
|
173
|
+
- [docs/api.md](docs/api.md) — every constructor arg and method
|
|
174
|
+
- [docs/dataflow.md](docs/dataflow.md) — a query end to end, diagrammed
|
|
175
|
+
- [docs/optimizations.md](docs/optimizations.md) — why each step of that path
|
|
176
|
+
looks the way it does, and what measured it
|
|
177
|
+
- [docs/caching.md](docs/caching.md) — sizing the decompressed-chunk cache, and
|
|
178
|
+
bounding many files together
|
|
179
|
+
- [agent-docs/adr/](agent-docs/adr/) — the measurements behind those decisions
|
|
180
|
+
- [agent-docs/TODO.md](agent-docs/TODO.md) — what is worth doing next, and what
|
|
181
|
+
has to be measured before it
|
|
182
|
+
- [CONTRIBUTING.md](CONTRIBUTING.md) — development and release steps
|
|
181
183
|
|
|
182
184
|
## Academic Use
|
|
183
185
|
|
|
184
|
-
|
|
185
|
-
|
|
186
|
-
|
|
187
|
-
be linked from [jbrowse.org](http://jbrowse.org).
|
|
186
|
+
Written with [NHGRI](http://genome.gov) funding as part of
|
|
187
|
+
[JBrowse](http://jbrowse.org). If you use this in a publication, please cite the
|
|
188
|
+
most recent JBrowse paper at [jbrowse.org](http://jbrowse.org).
|
|
188
189
|
|
|
189
190
|
## License
|
|
190
191
|
|