qvdjs 0.10.1 → 0.11.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +286 -56
- package/dist/index.cjs +690 -210
- package/dist/index.cjs.map +1 -1
- package/dist/index.js +690 -210
- package/dist/index.js.map +1 -1
- package/img/logo/qvdjs_logo-512.png +0 -0
- package/package.json +5 -2
package/README.md
CHANGED
|
@@ -1,3 +1,10 @@
|
|
|
1
|
+
<!-- Served from the npm tarball via jsDelivr rather than by relative path: this
|
|
2
|
+
repository is private, and npmjs.com resolves relative README images against
|
|
3
|
+
raw.githubusercontent.com, where they 404. See CLAUDE.md, "Logo assets". -->
|
|
4
|
+
<p align="center">
|
|
5
|
+
<img src="https://cdn.jsdelivr.net/npm/qvdjs/img/logo/qvdjs_logo-512.png" alt="qvdjs logo" width="200" />
|
|
6
|
+
</p>
|
|
7
|
+
|
|
1
8
|
# qvdjs
|
|
2
9
|
|
|
3
10
|
> Utility library for reading/writing Qlik Sense and QlikView (QVD) files in JavaScript/Node.js
|
|
@@ -49,6 +56,7 @@ It parses the binary QVD format into a JavaScript object structure and back agai
|
|
|
49
56
|
- [⚠️ Important Disclaimer](#️-important-disclaimer)
|
|
50
57
|
- [Install](#install)
|
|
51
58
|
- [Usage](#usage)
|
|
59
|
+
- [Four ways to open a file](#four-ways-to-open-a-file)
|
|
52
60
|
- [Lazy Loading](#lazy-loading)
|
|
53
61
|
- [Important: Symbol Table and High-Cardinality Fields](#important-symbol-table-and-high-cardinality-fields)
|
|
54
62
|
- [Performance Optimizations](#performance-optimizations)
|
|
@@ -68,6 +76,12 @@ It parses the binary QVD format into a JavaScript object structure and back agai
|
|
|
68
76
|
- [API Documentation](#api-documentation)
|
|
69
77
|
- [QvdDataFrame](#qvddataframe)
|
|
70
78
|
- [`static fromQvd(path: string, options?: object): Promise<QvdDataFrame>`](#static-fromqvdpath-string-options-object-promiseqvddataframe)
|
|
79
|
+
- [Reading part of a file](#reading-part-of-a-file)
|
|
80
|
+
- [`maxRows` and `limit` are one option](#maxrows-and-limit-are-one-option)
|
|
81
|
+
- [`fields` skips the columns you did not ask for](#fields-skips-the-columns-you-did-not-ask-for)
|
|
82
|
+
- [`static iterate(path: string, options?: object): AsyncGenerator<QvdDataFrame>`](#static-iteratepath-string-options-object-asyncgeneratorqvddataframe)
|
|
83
|
+
- [Read progress and cancellation](#read-progress-and-cancellation)
|
|
84
|
+
- [`static readMetadata(path: string, options?: object): Promise<object>`](#static-readmetadatapath-string-options-object-promiseobject)
|
|
71
85
|
- [`static fromDict(dict: object): Promise<QvdDataFrame>`](#static-fromdictdict-object-promiseqvddataframe)
|
|
72
86
|
- [`head(n: number): QvdDataFrame`](#headn-number-qvddataframe)
|
|
73
87
|
- [`tail(n: number): QvdDataFrame`](#tailn-number-qvddataframe)
|
|
@@ -80,6 +94,8 @@ It parses the binary QVD format into a JavaScript object structure and back agai
|
|
|
80
94
|
- [`getAllFieldMetadata(): object[]`](#getallfieldmetadata-object)
|
|
81
95
|
- [`setFileMetadata(metadata: object): void`](#setfilemetadatametadata-object-void)
|
|
82
96
|
- [`setFieldMetadata(fieldName: string, metadata: object): void`](#setfieldmetadatafieldname-string-metadata-object-void)
|
|
97
|
+
- [QvdColumnTable](#qvdcolumntable)
|
|
98
|
+
- [The low-level exports](#the-low-level-exports)
|
|
83
99
|
- [Documentation](#documentation)
|
|
84
100
|
- [For Users](#for-users)
|
|
85
101
|
- [For Contributors](#for-contributors)
|
|
@@ -135,17 +151,20 @@ console.log(df.head(5));
|
|
|
135
151
|
The above example loads the _qvdjs_ library and parses an example QVD file. A QVD file is typically loaded using the static
|
|
136
152
|
`QvdDataFrame.fromQvd` function of the `QvdDataFrame` class itself. After loading the file's content, numerous methods and properties are available to work with the parsed data.
|
|
137
153
|
|
|
138
|
-
###
|
|
154
|
+
### Four ways to open a file
|
|
139
155
|
|
|
140
156
|
`fromQvd` is the general one, and often not the one you want. A QVD is an XML header, then a
|
|
141
|
-
symbol table of every distinct value, then a bit-packed index table of one code per cell — so
|
|
142
|
-
far into the file a read has to go
|
|
143
|
-
|
|
144
|
-
|
|
145
|
-
|
|
146
|
-
|
|
|
147
|
-
|
|
|
148
|
-
|
|
|
157
|
+
symbol table of every distinct value, then a bit-packed index table of one code per cell — so two
|
|
158
|
+
things separate these: how far into the file a read has to go, and how much of it the read then
|
|
159
|
+
holds. `iterate` differs from `fromQvd` only on the second, which is why the column below answers
|
|
160
|
+
both.
|
|
161
|
+
|
|
162
|
+
| You want | Call | How far it reads, and what it holds |
|
|
163
|
+
| ---------------------------------------- | ----------------------------------------- | ----------------------------------------------------------------------------------------------- |
|
|
164
|
+
| Rows, to index, iterate or write back | `QvdDataFrame.fromQvd(path)` | Everything, and materialises every row |
|
|
165
|
+
| Rows from a file too large to hold | `QvdDataFrame.iterate(path, {chunkSize})` | Everything, but holds two chunks of rows — a 96 MB heap against 512 MB, at `chunkSize: 125_000` |
|
|
166
|
+
| A few columns of a large file | `QvdColumnTable.fromQvd(path)` | Everything, but stops before building rows — 141 MiB against 385 MiB |
|
|
167
|
+
| Only the schema: names, row count, types | `QvdDataFrame.readMetadata(path)` | The header alone. Constant cost, whatever the file's size |
|
|
149
168
|
|
|
150
169
|
```javascript
|
|
151
170
|
import {QvdDataFrame, QvdColumnTable} from 'qvdjs';
|
|
@@ -161,16 +180,28 @@ for (const value of table.column('amount')) {
|
|
|
161
180
|
if (typeof value === 'number') total += value;
|
|
162
181
|
}
|
|
163
182
|
|
|
183
|
+
// Every row of a file that will not fit, a chunk at a time.
|
|
184
|
+
for await (const chunk of QvdDataFrame.iterate('sales.qvd', {chunkSize: 50_000})) {
|
|
185
|
+
process(chunk.data);
|
|
186
|
+
}
|
|
187
|
+
|
|
164
188
|
// Rows, when you want rows.
|
|
165
189
|
const df = await QvdDataFrame.fromQvd('sales.qvd');
|
|
166
190
|
console.log(df.head(5));
|
|
167
191
|
```
|
|
168
192
|
|
|
169
|
-
|
|
193
|
+
The three that read data take the same options — `offset`, `limit` (or `maxRows`), `fields`,
|
|
194
|
+
`onProgress` and `signal` — because they are three answers about the same file rather than three
|
|
195
|
+
features. See [Reading part of a file](#reading-part-of-a-file). `readMetadata` shares
|
|
196
|
+
`allowedDir`, `onProgress` and `signal`, and ignores the window and field options — a read that
|
|
197
|
+
stops at the XML header has no rows to window and no field's symbols to skip.
|
|
198
|
+
|
|
199
|
+
Reaching for `fromQvd` when you wanted one of the other three is the common mistake, and
|
|
170
200
|
`fromQvd(path, {maxRows: 0})` is not a substitute for `readMetadata`: it loads no rows but still
|
|
171
|
-
reads and parses the whole symbol table, which grows with the data
|
|
172
|
-
[QvdColumnTable](#qvdcolumntable) and the
|
|
173
|
-
details and the measurements behind the table
|
|
201
|
+
reads and parses the whole symbol table, which grows with the data — and neither does `iterate`,
|
|
202
|
+
for the same reason. See [QvdColumnTable](#qvdcolumntable) and the
|
|
203
|
+
[API Documentation](#api-documentation) for the details and the measurements behind the table
|
|
204
|
+
above.
|
|
174
205
|
|
|
175
206
|
### Lazy Loading
|
|
176
207
|
|
|
@@ -194,9 +225,11 @@ console.log(df.shape); // [1000, numberOfColumns]
|
|
|
194
225
|
The QVD format stores data in two parts:
|
|
195
226
|
|
|
196
227
|
1. **Symbol table**: Contains every unique value for every field. It has to be _scanned_ in full, because a
|
|
197
|
-
symbol's length is only known once the previous one has been read
|
|
198
|
-
only the symbols the requested rows actually reference and steps
|
|
199
|
-
is what costs memory.
|
|
228
|
+
symbol's length is only known once the previous one has been read. Above `symbolFilteringThreshold` — 50 MB
|
|
229
|
+
by default — a bounded read also parses only the symbols the requested rows actually reference and steps
|
|
230
|
+
over the rest. Scanning is cheap; parsing is what costs memory. **Below that threshold every symbol is
|
|
231
|
+
parsed**, however few rows you asked for, because the extra analysis pass costs more than it saves on a
|
|
232
|
+
small table.
|
|
200
233
|
2. **Index table**: Contains row-by-row indices into the symbol table (only the requested rows are read)
|
|
201
234
|
|
|
202
235
|
⚠️ **Performance Impact of High-Cardinality Fields:**
|
|
@@ -207,22 +240,29 @@ If your QVD file contains fields with many unique values (high cardinality), suc
|
|
|
207
240
|
- Timestamps with millisecond precision
|
|
208
241
|
- Unique text fields
|
|
209
242
|
|
|
210
|
-
The symbol table then becomes very large, and has to be scanned end to end even when using `maxRows`.
|
|
211
|
-
is skipped for symbols the requested rows do not use, so the cost
|
|
243
|
+
The symbol table then becomes very large, and has to be scanned end to end even when using `maxRows`. Once it
|
|
244
|
+
passes `symbolFilteringThreshold`, parsing is skipped for symbols the requested rows do not use, so the cost
|
|
245
|
+
is I/O rather than memory:
|
|
212
246
|
|
|
213
247
|
- **Small symbol table** (fields with reusable values): Fast loading regardless of file size
|
|
214
248
|
- **Large symbol table** (fields with unique values per row): Slower, because the scan is proportional to the
|
|
215
|
-
symbol table's size — but memory
|
|
249
|
+
symbol table's size — but once filtering engages, memory scales with the rows you asked for rather than with
|
|
250
|
+
the file's cardinality. Under the threshold it still scales with the file, which is the case the default is
|
|
251
|
+
chosen for: a table that small costs little to parse whole.
|
|
216
252
|
|
|
217
253
|
You can see what happened on any load through `loadStats` (see [QvdDataFrame](#qvddataframe)):
|
|
218
254
|
|
|
219
255
|
```javascript
|
|
220
|
-
|
|
256
|
+
// This fixture's symbol table is 0.44 MB, well under the 50 MB default, so filtering does not
|
|
257
|
+
// engage on its own — the threshold is lowered here to show the path. See below.
|
|
258
|
+
const df = await QvdDataFrame.fromQvd('large.qvd', {maxRows: 1000, symbolFilteringThreshold: 0});
|
|
221
259
|
console.log(df.loadStats);
|
|
222
|
-
// { symbolTableBytes: 451704, totalRows: 1705805, rowsLoaded: 1000,
|
|
260
|
+
// { symbolTableBytes: 451704, totalRows: 1705805, rowsLoaded: 1000, offset: 0,
|
|
223
261
|
// symbolFiltering: true, symbolsKept: 3898 }
|
|
224
262
|
//
|
|
225
263
|
// 3,898 symbols parsed for 1,000 rows, out of a file holding 1.7 million.
|
|
264
|
+
// Without symbolFilteringThreshold: 0 the same call reports symbolFiltering: false,
|
|
265
|
+
// symbolsKept: null, and parses every symbol in the file.
|
|
226
266
|
```
|
|
227
267
|
|
|
228
268
|
To check your QVD's symbol table size, look at the `NoOfSymbols` in field metadata - values close to the total row count indicate high cardinality.
|
|
@@ -654,13 +694,13 @@ try {
|
|
|
654
694
|
|
|
655
695
|
Honest boundaries rather than an issue list — these are the ones that change what you can do.
|
|
656
696
|
|
|
657
|
-
| Limitation | What it means in practice
|
|
658
|
-
| ------------------------------------------------ |
|
|
659
|
-
| **Full loads stop at 2 GiB** | Use `maxRows` above that; the lazy path reads in chunks and has no such ceiling. Failure is a raw Node `RangeError`, not a `QvdError`.
|
|
660
|
-
| **A full load materialises every row in memory** | Rows are stored as one JavaScript array per row with a boxed value per cell, measured at roughly 11 bytes per cell. That, not file size, sets the ceiling in the table above. `
|
|
661
|
-
| **Dual values are not round-tripped as pairs** | Qlik stores dates and timestamps as a number plus a display string. A data frame exposes one of the two, so a read-modify-write does not reproduce the original dual — a `BirthDate` field comes back as `24205`, not as a formatted date.
|
|
662
|
-
| **Writes are not atomic** | `toQvd()` writes in place. A failure part-way through, or two writers targeting one path, leaves the previous file damaged rather than intact. Write to a temporary path and rename if that matters.
|
|
663
|
-
| **Writer input is not validated** | Unsupported value types can produce a raw `TypeError`, or a file that is written but does not read back as intended. Feed `fromDict()` numbers, strings and `null`.
|
|
697
|
+
| Limitation | What it means in practice | Tracked as |
|
|
698
|
+
| ------------------------------------------------ | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | --------------------------------------------------------- |
|
|
699
|
+
| **Full loads stop at 2 GiB** | Use `maxRows` above that; the lazy path reads in chunks and has no such ceiling. Failure is a raw Node `RangeError`, not a `QvdError`. | [#122](https://github.com/ptarmiganlabs/qvdjs/issues/122) |
|
|
700
|
+
| **A full load materialises every row in memory** | Rows are stored as one JavaScript array per row with a boxed value per cell, measured at roughly 11 bytes per cell. That, not file size, sets the ceiling in the table above. `{offset, limit}` reads an arbitrary window, `iterate()` walks a file in bounded chunks, and `QvdColumnTable` avoids row materialisation entirely — but **the symbol table is still parsed in full on every one of those paths**, so none of them is constant-memory in the size of a high-cardinality file. | [#140](https://github.com/ptarmiganlabs/qvdjs/issues/140) |
|
|
701
|
+
| **Dual values are not round-tripped as pairs** | Qlik stores dates and timestamps as a number plus a display string. A data frame exposes one of the two, so a read-modify-write does not reproduce the original dual — a `BirthDate` field comes back as `24205`, not as a formatted date. | [#138](https://github.com/ptarmiganlabs/qvdjs/issues/138) |
|
|
702
|
+
| **Writes are not atomic** | `toQvd()` writes in place. A failure part-way through, or two writers targeting one path, leaves the previous file damaged rather than intact. Write to a temporary path and rename if that matters. | [#129](https://github.com/ptarmiganlabs/qvdjs/issues/129) |
|
|
703
|
+
| **Writer input is not validated** | Unsupported value types can produce a raw `TypeError`, or a file that is written but does not read back as intended. Feed `fromDict()` numbers, strings and `null`. | [#130](https://github.com/ptarmiganlabs/qvdjs/issues/130) |
|
|
664
704
|
|
|
665
705
|
Numeric-looking strings are currently coerced to numbers on read, which is being reconsidered as a breaking
|
|
666
706
|
change in [#120](https://github.com/ptarmiganlabs/qvdjs/issues/120). Empty and whitespace-only strings are _not_
|
|
@@ -779,14 +819,14 @@ Empty QVDs are fully supported for both reading and writing, maintaining compati
|
|
|
779
819
|
The `QvdDataFrame` class represents the data frame stored inside of a finally parsed QVD file. It provides a high-level
|
|
780
820
|
abstraction access to the QVD file content. This includes meta information as well as access to the actual data records.
|
|
781
821
|
|
|
782
|
-
| Property | Type | Description
|
|
783
|
-
| -------------- | ---------- |
|
|
784
|
-
| `shape` | `number[]` | The shape of the data table. The first element is the number of rows, the second element is the number of columns.
|
|
785
|
-
| `data` | `any[][]` | The actual data records of the QVD file. The first dimension represents the single rows.
|
|
786
|
-
| `columns` | `string[]` | The names of the fields that are contained in the QVD file.
|
|
787
|
-
| `metadata` | `object` | The complete metadata object from the QVD file header, or null if not loaded from a QVD file.
|
|
788
|
-
| `fileMetadata` | `object` | File-level metadata from the QVD header (qvBuildNo, tableName, createUtcTime, etc.).
|
|
789
|
-
| `loadStats` | `object` | What the read did: `symbolTableBytes`, `totalRows`, `rowsLoaded`, `symbolFiltering`, `symbolsKept`. `
|
|
822
|
+
| Property | Type | Description |
|
|
823
|
+
| -------------- | ---------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
|
|
824
|
+
| `shape` | `number[]` | The shape of the data table. The first element is the number of rows, the second element is the number of columns. |
|
|
825
|
+
| `data` | `any[][]` | The actual data records of the QVD file. The first dimension represents the single rows. |
|
|
826
|
+
| `columns` | `string[]` | The names of the fields that are contained in the QVD file. |
|
|
827
|
+
| `metadata` | `object` | The complete metadata object from the QVD file header, or null if not loaded from a QVD file. |
|
|
828
|
+
| `fileMetadata` | `object` | File-level metadata from the QVD header (qvBuildNo, tableName, createUtcTime, etc.). |
|
|
829
|
+
| `loadStats` | `object` | What the read did: `symbolTableBytes`, `totalRows`, `rowsLoaded`, `offset`, `symbolFiltering`, `symbolsKept`. Carried by `fromQvd()`, by every chunk `iterate()` yields, and by `QvdColumnTable`; `null` on frames from `fromDict()`, `head()`, `tail()`, `rows()` and `select()`, which describe no particular read. |
|
|
790
830
|
|
|
791
831
|
#### `static fromQvd(path: string, options?: object): Promise<QvdDataFrame>`
|
|
792
832
|
|
|
@@ -797,9 +837,14 @@ to a `QvdDataFrame` instance.
|
|
|
797
837
|
|
|
798
838
|
- `path` (string): The path to the QVD file.
|
|
799
839
|
- `options` (object, optional): Loading options
|
|
800
|
-
- `maxRows` (number, optional): Maximum number of rows to load. If not specified, all rows are loaded. This is useful for loading only a subset of data from large QVD files to improve performance and reduce memory usage.
|
|
840
|
+
- `maxRows` (number, optional): Maximum number of rows to load. If not specified, all rows are loaded. This is useful for loading only a subset of data from large QVD files to improve performance and reduce memory usage. **`maxRows` and `limit` are the same option under two names** — see [Reading part of a file](#reading-part-of-a-file). Passing both throws.
|
|
841
|
+
- `limit` (number, optional): Rows to read, counting from `offset`. The same number as `maxRows`, spelled so that it reads correctly beside an offset.
|
|
842
|
+
- `offset` (number, optional): File row to start at. Defaults to 0. An offset past the end of the file returns **no rows** rather than throwing, so a paging loop terminates on its own.
|
|
843
|
+
- `fields` (string[], optional): Field names to read, in the order they should appear in the result. Unselected fields have their symbols **skipped entirely** rather than parsed and discarded. An unknown or repeated name throws.
|
|
844
|
+
- `onProgress` (function, optional): Called with `{stage, current, total, percent}` as the read proceeds — the same shape `toQvd`'s callback receives. See [Read progress and cancellation](#read-progress-and-cancellation).
|
|
845
|
+
- `signal` (AbortSignal, optional): Cancels the read. The rejection is `signal.reason`.
|
|
801
846
|
- `allowedDir` (string, optional): Base directory for file access validation. Defaults to current working directory (CWD). The file path must resolve to a location within this directory to prevent path traversal attacks. Set to a specific directory in production environments with user-provided paths.
|
|
802
|
-
- `memorySafetyFactor` (number, optional): Fraction (0.0-1.0) of the memory budget a load may use. Default is 0.8 (80%). The budget is the smaller of the V8 heap limit and any container memory limit; see [QVD File Size Limitations](#qvd-file-size-limitations).
|
|
847
|
+
- `memorySafetyFactor` (number, optional): Fraction (0.0-1.0) of the memory budget a load may use. Default is 0.8 (80%). The budget is the smaller of the V8 heap limit and any container memory limit; see [QVD File Size Limitations](#qvd-file-size-limitations). Lower it (e.g. to 0.5) to leave the garbage collector more room; raise it toward 1.0 to use more of the budget. The figures either side of 0.8 are decreases and increases respectively — an earlier default of 0.3 is why this line used to call 0.5 an increase. **`0` disables the memory check entirely.**
|
|
803
848
|
- `symbolFilteringThreshold` (number, optional): Symbol table size, in bytes, above which a lazy load switches to the two-pass filtering path. Defaults to 50 MB, the point where the extra analysis pass pays for itself. Lower it to use filtering on smaller files, raise it to keep the simpler single-pass read for longer.
|
|
804
849
|
|
|
805
850
|
**Example:**
|
|
@@ -828,8 +873,152 @@ const dfUnchecked = await QvdDataFrame.fromQvd('large-file.qvd', {memorySafetyFa
|
|
|
828
873
|
const preview = await QvdDataFrame.fromQvd('large-file.qvd', {maxRows: 500});
|
|
829
874
|
console.log(preview.loadStats.symbolFiltering); // true when the two-pass path ran
|
|
830
875
|
console.log(preview.loadStats.symbolsKept); // how many symbols it kept
|
|
876
|
+
|
|
877
|
+
// An arbitrary row window, and only the columns you need
|
|
878
|
+
const page = await QvdDataFrame.fromQvd('trips.qvd', {offset: 1_000_000, limit: 500});
|
|
879
|
+
const narrow = await QvdDataFrame.fromQvd('trips.qvd', {fields: ['fare', 'trip_miles']});
|
|
880
|
+
```
|
|
881
|
+
|
|
882
|
+
#### Reading part of a file
|
|
883
|
+
|
|
884
|
+
Four options control _which_ part of a file is read. They compose: any combination of a window and
|
|
885
|
+
a projection is a valid read.
|
|
886
|
+
|
|
887
|
+
| Option | Meaning |
|
|
888
|
+
| --------- | ------------------------------------------------------------ |
|
|
889
|
+
| `offset` | File row to start at. Default 0. |
|
|
890
|
+
| `limit` | Rows to read from `offset`. Default: to the end of the file. |
|
|
891
|
+
| `maxRows` | The older name for `limit`. Identical; passing both throws. |
|
|
892
|
+
| `fields` | Field names to read, in the order they should appear. |
|
|
893
|
+
|
|
894
|
+
##### `maxRows` and `limit` are one option
|
|
895
|
+
|
|
896
|
+
`maxRows` came first and every released version documents it. `limit` is the spelling that reads
|
|
897
|
+
correctly beside `offset`, because "the maximum number of rows" says nothing about where they
|
|
898
|
+
start. **Neither is deprecated, and they are the same number**: `5`, `{maxRows: 5}` and
|
|
899
|
+
`{offset: 0, limit: 5}` are one read, resolved in one place so they cannot drift.
|
|
900
|
+
|
|
901
|
+
Passing both throws, even when they agree. Any rule for picking a winner is a rule you would have
|
|
902
|
+
to remember, and getting it wrong returns a plausible number of rows rather than an error.
|
|
903
|
+
|
|
904
|
+
```javascript
|
|
905
|
+
await QvdDataFrame.fromQvd('trips.qvd', {maxRows: 100}); // first 100 rows
|
|
906
|
+
await QvdDataFrame.fromQvd('trips.qvd', {limit: 100}); // the same read
|
|
907
|
+
await QvdDataFrame.fromQvd('trips.qvd', {offset: 500, limit: 100}); // rows 500-599
|
|
908
|
+
await QvdDataFrame.fromQvd('trips.qvd', {offset: 500}); // row 500 to the end
|
|
909
|
+
await QvdDataFrame.fromQvd('trips.qvd', {maxRows: 5, limit: 5}); // throws
|
|
910
|
+
```
|
|
911
|
+
|
|
912
|
+
An offset past the end of the file returns **no rows** rather than throwing — the way
|
|
913
|
+
`Array.prototype.slice` answers it — so a paging loop stops without a pre-check.
|
|
914
|
+
|
|
915
|
+
A window deep in a large file costs what the window costs. On the bundled 1.7 M × 20 taxi fixture
|
|
916
|
+
the last thousand rows take **31 ms and read 0.46 MiB**, against 663 ms and 37.9 MiB for the only
|
|
917
|
+
thing that was possible before — reading it all and slicing.
|
|
918
|
+
|
|
919
|
+
##### `fields` skips the columns you did not ask for
|
|
920
|
+
|
|
921
|
+
Unselected fields have their symbol areas **skipped, not parsed and discarded**. That is where the
|
|
922
|
+
saving is: the index table is row-major, so every record carries every field's bits and no
|
|
923
|
+
projection can narrow the bytes read from disk. What it avoids is parsing symbols and decoding
|
|
924
|
+
columns.
|
|
925
|
+
|
|
926
|
+
- Taxi fixture, 2 of 20 columns: **162 ms against 674 ms**, 4.2×.
|
|
927
|
+
- A file whose symbol table is 94 % of it, 2 of 8 columns: **96 ms against 290 ms**, 3.0×.
|
|
928
|
+
|
|
929
|
+
```javascript
|
|
930
|
+
const df = await QvdDataFrame.fromQvd('trips.qvd', {fields: ['fare', 'taxi_id']});
|
|
931
|
+
console.log(df.columns); // ['fare', 'taxi_id'] — your order, not the file's
|
|
932
|
+
```
|
|
933
|
+
|
|
934
|
+
- **The result is in your order**, not the file's, matching `select('b', 'a')`.
|
|
935
|
+
- **An unknown field name throws**, and so does a repeated one or an empty list. Silently returning
|
|
936
|
+
fewer columns than you asked for is the bug this library is designed against.
|
|
937
|
+
- **`metadata` still describes the file**, not the projection: `getAllFieldMetadata()` reports every
|
|
938
|
+
field whether or not it was read, exactly as `NoOfRecords` stays the file's row count under
|
|
939
|
+
`maxRows`. `columns` is what says what was read.
|
|
940
|
+
- **A projection cannot make a corrupt file readable.** Every field's metadata is validated whether
|
|
941
|
+
or not it was selected.
|
|
942
|
+
|
|
943
|
+
#### `static iterate(path: string, options?: object): AsyncGenerator<QvdDataFrame>`
|
|
944
|
+
|
|
945
|
+
Reads a QVD in chunks. Takes every option `fromQvd` takes, plus `chunkSize` (default 100 000).
|
|
946
|
+
|
|
947
|
+
```javascript
|
|
948
|
+
for await (const chunk of QvdDataFrame.iterate('big.qvd', {chunkSize: 50_000})) {
|
|
949
|
+
process(chunk.data); // at most 50,000 rows
|
|
950
|
+
console.log(chunk.loadStats.offset); // where in the file this chunk starts
|
|
951
|
+
}
|
|
952
|
+
```
|
|
953
|
+
|
|
954
|
+
The file is opened, read and parsed **once**; only the index decode and the row building happen per
|
|
955
|
+
chunk. That is why this is a method rather than a loop of `fromQvd({offset, limit})` calls at your
|
|
956
|
+
call site: the symbol table has to be parsed in full whatever the chunk size — a stored index in the
|
|
957
|
+
last chunk can address the first symbol — so re-parsing it per chunk costs more than a plain load
|
|
958
|
+
rather than less.
|
|
959
|
+
|
|
960
|
+
What it bounds is **row materialisation**, which is what dominates a large read's heap. Measured on
|
|
961
|
+
the taxi fixture by binary-searching the smallest `--max-old-space-size` at which each read
|
|
962
|
+
completes, in a fresh process per rung:
|
|
963
|
+
|
|
964
|
+
| | Smallest heap it runs in | Smallest heap it is allowed in |
|
|
965
|
+
| ------------------------------- | -----------------------: | -----------------------------: |
|
|
966
|
+
| `fromQvd()`, all 1.7 M rows | **384 MB** | **512 MB** |
|
|
967
|
+
| `iterate({chunkSize: 125_000})` | **80 MB** | **96 MB** |
|
|
968
|
+
| `iterate({chunkSize: 250_000})` | 128 MB | 192 MB |
|
|
969
|
+
| `iterate({chunkSize: 500_000})` | 224 MB | 320 MB |
|
|
970
|
+
|
|
971
|
+
Every row is read in all four. The two columns answer different questions and both are worth
|
|
972
|
+
having: the first is measured with `memorySafetyFactor: 0`, so it is what the read physically
|
|
973
|
+
needs, and the second is the one you actually hit, because on default settings the memory guard
|
|
974
|
+
refuses below it. Quoting only the first would name a heap no default caller can reproduce.
|
|
975
|
+
|
|
976
|
+
Time is not the reason to use it — walking the whole file in chunks of 100 000 takes 593 ms against
|
|
977
|
+
674 ms for a single frame.
|
|
978
|
+
|
|
979
|
+
⚠️ **This is not constant-memory reading of an arbitrarily large file.** The symbol table is parsed
|
|
980
|
+
in full whatever the chunk size. On a high-cardinality file that table is the bulk of the cost, and
|
|
981
|
+
`readMetadata` is the only read that avoids it.
|
|
982
|
+
|
|
983
|
+
A window covering no rows yields nothing at all, rather than one empty frame, so a loop over an
|
|
984
|
+
exhausted offset simply does not run its body.
|
|
985
|
+
|
|
986
|
+
#### Read progress and cancellation
|
|
987
|
+
|
|
988
|
+
`onProgress` and `signal` work on `fromQvd`, `iterate` and `QvdColumnTable.fromQvd`. The progress
|
|
989
|
+
object is the same `{stage, current, total, percent}` shape that `toQvd` emits, so one progress bar
|
|
990
|
+
serves both directions.
|
|
991
|
+
|
|
992
|
+
```javascript
|
|
993
|
+
const controller = new AbortController();
|
|
994
|
+
setTimeout(() => controller.abort(), 30_000);
|
|
995
|
+
|
|
996
|
+
const df = await QvdDataFrame.fromQvd('huge.qvd', {
|
|
997
|
+
signal: controller.signal,
|
|
998
|
+
onProgress: ({stage, percent}) => process.stdout.write(`\r${stage}: ${percent}% `),
|
|
999
|
+
});
|
|
831
1000
|
```
|
|
832
1001
|
|
|
1002
|
+
Read stages, in order:
|
|
1003
|
+
|
|
1004
|
+
| Stage | Units | Notes |
|
|
1005
|
+
| ----------------- | -------------- | -------------------------------------------------- |
|
|
1006
|
+
| `read` | 0 → 1 | Pulling the bytes off disk. |
|
|
1007
|
+
| `header` | 0 → 1 | Parsing the XML header. |
|
|
1008
|
+
| `symbol-analysis` | fields | Only when the two-pass symbol-filtering path runs. |
|
|
1009
|
+
| `symbol-table` | fields | Parsing symbols. Counts selected fields only. |
|
|
1010
|
+
| `index-table` | fields | Decoding the bit-packed columns. |
|
|
1011
|
+
| `rows` | rows, every 1% | Building row arrays. Absent on a columnar read. |
|
|
1012
|
+
|
|
1013
|
+
Under `iterate`, `index-table` and `rows` repeat per chunk, and `rows` reports the position in the
|
|
1014
|
+
**whole window** rather than in the chunk — so it works as a progress bar across the iteration.
|
|
1015
|
+
|
|
1016
|
+
Cancellation rejects with `signal.reason`, exactly as `signal.throwIfAborted()` produces it: a
|
|
1017
|
+
`DOMException` named `AbortError` unless you called `abort(reason)` with your own error. That is
|
|
1018
|
+
deliberately **not** a `QvdError` — it is what `AbortSignal` means everywhere else in Node, and a
|
|
1019
|
+
library-specific abort error would be the one your code has to special-case. An already-aborted
|
|
1020
|
+
signal stops the read before it opens the file.
|
|
1021
|
+
|
|
833
1022
|
#### `static readMetadata(path: string, options?: object): Promise<object>`
|
|
834
1023
|
|
|
835
1024
|
Reads a QVD file's schema and header metadata without reading its data. The cost is the same
|
|
@@ -848,6 +1037,11 @@ grows.
|
|
|
848
1037
|
- `options` (object, optional):
|
|
849
1038
|
- `allowedDir` (string, optional): Base directory for file access validation, applied exactly as
|
|
850
1039
|
it is for `fromQvd`.
|
|
1040
|
+
- `onProgress` (function, optional): As on `fromQvd`. Only the `read` and `header` stages occur
|
|
1041
|
+
here — there are no symbols to parse and no rows to build.
|
|
1042
|
+
- `signal` (AbortSignal, optional): Cancels the read, rejecting with `signal.reason`.
|
|
1043
|
+
- `offset`, `limit`, `maxRows` and `fields` are accepted and ignored, so one options object can
|
|
1044
|
+
be passed to this and to a data read without stripping it.
|
|
851
1045
|
|
|
852
1046
|
**Returns** a plain object — deliberately not a `QvdDataFrame`, since one with `data: []` would be
|
|
853
1047
|
indistinguishable from an empty file at the call site:
|
|
@@ -925,8 +1119,24 @@ for (const value of fare) {
|
|
|
925
1119
|
console.log(fare.at(0), fare.toArray().length);
|
|
926
1120
|
```
|
|
927
1121
|
|
|
928
|
-
**`QvdColumnTable`** — `static fromQvd(path, options)`
|
|
929
|
-
`
|
|
1122
|
+
**`QvdColumnTable`** — `static fromQvd(path, options)`, `column(name)`, `columns`, `rowCount`,
|
|
1123
|
+
`shape`, `metadata`, `loadStats`.
|
|
1124
|
+
|
|
1125
|
+
`fromQvd` takes **exactly the same options** as `QvdDataFrame.fromQvd` — `offset`, `limit`,
|
|
1126
|
+
`maxRows`, `fields`, `onProgress`, `signal`, `allowedDir`, `memorySafetyFactor`,
|
|
1127
|
+
`symbolFilteringThreshold` — with the same meanings. They are two answers about the same file, not
|
|
1128
|
+
two features, so `{offset, limit}` is how you page through a file columnwise:
|
|
1129
|
+
|
|
1130
|
+
```javascript
|
|
1131
|
+
const table = await QvdColumnTable.fromQvd('trips.qvd', {
|
|
1132
|
+
fields: ['fare', 'tips'],
|
|
1133
|
+
offset: 1_000_000,
|
|
1134
|
+
limit: 100_000,
|
|
1135
|
+
});
|
|
1136
|
+
```
|
|
1137
|
+
|
|
1138
|
+
There is no columnar `iterate()`, deliberately: chunking exists to bound row materialisation, and a
|
|
1139
|
+
columnar read materialises no rows. `{offset, limit}` covers paging.
|
|
930
1140
|
|
|
931
1141
|
**`QvdColumn`** — what `column(name)` returns, frozen:
|
|
932
1142
|
|
|
@@ -1102,6 +1312,25 @@ Immutable properties (cannot be modified):
|
|
|
1102
1312
|
- `bias`: Bias value
|
|
1103
1313
|
- `noOfSymbols`: Number of symbols
|
|
1104
1314
|
|
|
1315
|
+
### The low-level exports
|
|
1316
|
+
|
|
1317
|
+
`src/index.js` also exports `QvdFileReader`, `QvdFileWriter` and `QvdSymbol`. **`QvdDataFrame` and
|
|
1318
|
+
`QvdColumnTable` are the surface you should use**; these three are the building blocks they are
|
|
1319
|
+
built from, and they are documented here because an exported class that appears nowhere in the
|
|
1320
|
+
documentation is the worst of both worlds — nobody can tell whether it is supported, and removing
|
|
1321
|
+
it later breaks whoever guessed that it was.
|
|
1322
|
+
|
|
1323
|
+
They stay exported. Removing them would be a breaking change for a saving of nothing, and
|
|
1324
|
+
`QvdFileReader` in particular is where chunked iteration actually lives.
|
|
1325
|
+
|
|
1326
|
+
| Class | What it is |
|
|
1327
|
+
| --------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
|
|
1328
|
+
| `QvdFileReader` | The reader `fromQvd`, `iterate` and `readMetadata` are thin wrappers over. `new QvdFileReader(path, options)` takes the reader-shaped options — `allowedDir`, `memorySafetyFactor`, `symbolFilteringThreshold`, `fields`, `onProgress`, `signal`, and `materialisesRows` — and exposes `load(window)`, `loadColumnar(window)`, `iterateRows(window, chunkSize)` and `loadMetadata()`. **Pass `materialisesRows: false` whenever you call `loadColumnar` directly**, as `QvdColumnTable.fromQvd` does: it defaults to true, so without it the memory guard charges a columnar read for rows it never builds and refuses reads the high-level call completes. `window` is a number, `null`, or `{offset, limit}`. |
|
|
1329
|
+
| `QvdFileWriter` | What `toQvd` uses. `new QvdFileWriter(path, dataFrame, {allowedDir, onProgress})` then `save()`. |
|
|
1330
|
+
| `QvdSymbol` | One value as the format stores it — an integer, a double, a string, or a dual pairing a number with its display string. `toPrimaryValue()` is what the reader calls to turn one into a JavaScript value. Useful for inspecting the format; the read paths resolve symbols for you. |
|
|
1331
|
+
|
|
1332
|
+
Everything else under `src/` is internal, whatever its visibility.
|
|
1333
|
+
|
|
1105
1334
|
## Documentation
|
|
1106
1335
|
|
|
1107
1336
|
Comprehensive documentation is available to help you understand, use, and contribute to qvdjs:
|
|
@@ -1109,7 +1338,7 @@ Comprehensive documentation is available to help you understand, use, and contri
|
|
|
1109
1338
|
### For Users
|
|
1110
1339
|
|
|
1111
1340
|
- **[README.md](./README.md)** (this file) - Quick start guide and API reference
|
|
1112
|
-
- **[QVD_FORMAT.md](./QVD_FORMAT.md)** - Complete QVD file format specification
|
|
1341
|
+
- **[QVD_FORMAT.md](./docs/QVD_FORMAT.md)** - Complete QVD file format specification
|
|
1113
1342
|
- Binary structure details
|
|
1114
1343
|
- Symbol type encodings
|
|
1115
1344
|
- Bit packing algorithms
|
|
@@ -1117,14 +1346,14 @@ Comprehensive documentation is available to help you understand, use, and contri
|
|
|
1117
1346
|
|
|
1118
1347
|
### For Contributors
|
|
1119
1348
|
|
|
1120
|
-
- **[CONTRIBUTING.md](./CONTRIBUTING.md)** - How to contribute to the project
|
|
1349
|
+
- **[CONTRIBUTING.md](./docs/CONTRIBUTING.md)** - How to contribute to the project
|
|
1121
1350
|
- Development setup
|
|
1122
1351
|
- Code style guidelines
|
|
1123
1352
|
- Commit message conventions
|
|
1124
1353
|
- Pull request process
|
|
1125
1354
|
- Bug reporting and feature requests
|
|
1126
1355
|
|
|
1127
|
-
- **[DEVELOPMENT.md](./DEVELOPMENT.md)** - Technical development guide
|
|
1356
|
+
- **[DEVELOPMENT.md](./docs/DEVELOPMENT.md)** - Technical development guide
|
|
1128
1357
|
- Architecture overview
|
|
1129
1358
|
- Implementation details
|
|
1130
1359
|
- Design patterns
|
|
@@ -1132,7 +1361,7 @@ Comprehensive documentation is available to help you understand, use, and contri
|
|
|
1132
1361
|
- Error handling strategies
|
|
1133
1362
|
- Debugging tips
|
|
1134
1363
|
|
|
1135
|
-
- **[ARCHITECTURE.md](./ARCHITECTURE.md)** - High-level architecture
|
|
1364
|
+
- **[ARCHITECTURE.md](./docs/ARCHITECTURE.md)** - High-level architecture
|
|
1136
1365
|
- Component diagrams
|
|
1137
1366
|
- Class relationships
|
|
1138
1367
|
- Data flow visualization
|
|
@@ -1149,16 +1378,16 @@ Comprehensive documentation is available to help you understand, use, and contri
|
|
|
1149
1378
|
|
|
1150
1379
|
### Quick Links by Task
|
|
1151
1380
|
|
|
1152
|
-
| I want to... | See...
|
|
1153
|
-
| ----------------------- |
|
|
1154
|
-
| Use the library | [README.md](./README.md) - Usage section
|
|
1155
|
-
| Understand QVD format | [QVD_FORMAT.md](./QVD_FORMAT.md)
|
|
1156
|
-
| Report a bug | [CONTRIBUTING.md](./CONTRIBUTING.md#reporting-bugs)
|
|
1157
|
-
| Suggest a feature | [CONTRIBUTING.md](./CONTRIBUTING.md#suggesting-features)
|
|
1158
|
-
| Contribute code | [CONTRIBUTING.md](./CONTRIBUTING.md) + [DEVELOPMENT.md](./DEVELOPMENT.md) |
|
|
1159
|
-
| Understand architecture | [ARCHITECTURE.md](./ARCHITECTURE.md)
|
|
1160
|
-
| Write tests | [docs/TESTING.md](./docs/TESTING.md)
|
|
1161
|
-
| Debug an issue | [DEVELOPMENT.md](./DEVELOPMENT.md#debugging-tips)
|
|
1381
|
+
| I want to... | See... |
|
|
1382
|
+
| ----------------------- | ----------------------------------------------------------------------------------- |
|
|
1383
|
+
| Use the library | [README.md](./README.md) - Usage section |
|
|
1384
|
+
| Understand QVD format | [QVD_FORMAT.md](./docs/QVD_FORMAT.md) |
|
|
1385
|
+
| Report a bug | [CONTRIBUTING.md](./docs/CONTRIBUTING.md#reporting-bugs) |
|
|
1386
|
+
| Suggest a feature | [CONTRIBUTING.md](./docs/CONTRIBUTING.md#suggesting-features) |
|
|
1387
|
+
| Contribute code | [CONTRIBUTING.md](./docs/CONTRIBUTING.md) + [DEVELOPMENT.md](./docs/DEVELOPMENT.md) |
|
|
1388
|
+
| Understand architecture | [ARCHITECTURE.md](./docs/ARCHITECTURE.md) |
|
|
1389
|
+
| Write tests | [docs/TESTING.md](./docs/TESTING.md) |
|
|
1390
|
+
| Debug an issue | [DEVELOPMENT.md](./docs/DEVELOPMENT.md#debugging-tips) |
|
|
1162
1391
|
|
|
1163
1392
|
## Testing
|
|
1164
1393
|
|
|
@@ -1174,8 +1403,9 @@ See the **[Testing Documentation](./docs/README.md)** in the `docs/` directory.
|
|
|
1174
1403
|
|
|
1175
1404
|
Quick links:
|
|
1176
1405
|
|
|
1177
|
-
- **[
|
|
1178
|
-
- **[
|
|
1406
|
+
- **[TESTING.md](./docs/TESTING.md)** - Test architecture, coverage and the multi-platform setup
|
|
1407
|
+
- **[TESTING_WITH_LARGE_FILES.md](./docs/TESTING_WITH_LARGE_FILES.md)** - Working with the bundled
|
|
1408
|
+
multi-megabyte fixtures
|
|
1179
1409
|
|
|
1180
1410
|
### Running Tests
|
|
1181
1411
|
|