qvdjs 0.9.4 → 0.10.1

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -30,9 +30,10 @@ It parses the binary QVD format into a JavaScript object structure and back agai
30
30
  callbacks for long writes. Files it produces open in Qlik Sense and QlikView.
31
31
  - **Refusing rather than crashing.** A load too large for the process throws a catchable `QvdValidationError`
32
32
  naming the limit it hit and a row count that would fit, instead of a `FATAL ERROR: Reached heap limit` that no
33
- `try`/`catch` can intercept. The budget comes from the V8 heap ceiling and any container memory limit — the
34
- two things that actually kill a process — so it neither refuses work the machine can do nor waves through work
35
- it cannot. See [QVD File Size Limitations](#qvd-file-size-limitations).
33
+ `try`/`catch` can intercept. The estimate accounts for the rows and columns being materialised, not just the
34
+ symbol table, and the budget comes from the V8 heap ceiling and any container memory limit — the two things
35
+ that actually kill a process. The suggested row count is checked against the same estimate before being
36
+ offered, so following it works. See [QVD File Size Limitations](#qvd-file-size-limitations).
36
37
  - **Corrupt files are detected, not silently misread.** Truncated index tables, missing header delimiters and
37
38
  out-of-range offsets raise typed errors rather than returning short or fabricated data.
38
39
  - **Path traversal protection by default.** File access is confined to the working directory unless you widen it,
@@ -134,6 +135,43 @@ console.log(df.head(5));
134
135
  The above example loads the _qvdjs_ library and parses an example QVD file. A QVD file is typically loaded using the static
135
136
  `QvdDataFrame.fromQvd` function of the `QvdDataFrame` class itself. After loading the file's content, numerous methods and properties are available to work with the parsed data.
136
137
 
138
+ ### Three ways to open a file
139
+
140
+ `fromQvd` is the general one, and often not the one you want. A QVD is an XML header, then a
141
+ symbol table of every distinct value, then a bit-packed index table of one code per cell — so how
142
+ far into the file a read has to go is what separates these, and all three take the same options.
143
+
144
+ | You want | Call | How far it reads |
145
+ | --------------------------------------- | --------------------------------- | ---------------------------------------------------------------------- |
146
+ | Rows, to index, iterate or write back | `QvdDataFrame.fromQvd(path)` | Everything, and materialises every row |
147
+ | A few columns of a large file | `QvdColumnTable.fromQvd(path)` | Everything, but stops before building rows — 141 MiB against 385 MiB |
148
+ | Only the schema: names, row count, types | `QvdDataFrame.readMetadata(path)` | The header alone. Constant cost, whatever the file's size |
149
+
150
+ ```javascript
151
+ import {QvdDataFrame, QvdColumnTable} from 'qvdjs';
152
+
153
+ // What is in this file? Costs the same whether it is 20 KB or 20 GB.
154
+ const {columns, rowCount} = await QvdDataFrame.readMetadata('sales.qvd');
155
+ console.log(`${rowCount} rows x ${columns.length} columns`);
156
+
157
+ // Sum one column without ever building a row.
158
+ const table = await QvdColumnTable.fromQvd('sales.qvd');
159
+ let total = 0;
160
+ for (const value of table.column('amount')) {
161
+ if (typeof value === 'number') total += value;
162
+ }
163
+
164
+ // Rows, when you want rows.
165
+ const df = await QvdDataFrame.fromQvd('sales.qvd');
166
+ console.log(df.head(5));
167
+ ```
168
+
169
+ Reaching for `fromQvd` when you wanted one of the other two is the common mistake, and
170
+ `fromQvd(path, {maxRows: 0})` is not a substitute for `readMetadata`: it loads no rows but still
171
+ reads and parses the whole symbol table, which grows with the data. See
172
+ [QvdColumnTable](#qvdcolumntable) and the [API Documentation](#api-documentation) for the
173
+ details and the measurements behind the table above.
174
+
137
175
  ### Lazy Loading
138
176
 
139
177
  For large QVD files, you can load only a specific number of rows to improve performance and reduce memory usage. The library implements **lazy loading** - it reads only the necessary portions of the file from disk, not the entire file.
@@ -287,11 +325,12 @@ The maximum QVD file size you can handle with qvdjs depends on several factors a
287
325
 
288
326
  - **Node.js Memory Limits**: By default, Node.js limits heap memory to approximately 4GB (varies by architecture and Node.js version). The library **automatically detects your configured heap size** and scales its safety limits accordingly. You can increase heap size using the `--max-old-space-size` flag (e.g., `node --max-old-space-size=16384 script.js` for 16GB), and qvdjs will automatically allow larger files.
289
327
 
290
- For larger heap configurations, you can also adjust the `memorySafetyFactor` option (default 0.3 = 30%) to make more efficient use of available memory:
328
+ You can also adjust the `memorySafetyFactor` option (default 0.8 = 80% of the budget) to trade
329
+ headroom for reach:
291
330
 
292
331
  ```javascript
293
332
  const df = await QvdDataFrame.fromQvd('large-file.qvd', {
294
- memorySafetyFactor: 0.5, // Use 50% of heap instead of default 30%
333
+ memorySafetyFactor: 0.9, // Use 90% of the budget instead of the default 80%
295
334
  });
296
335
  ```
297
336
 
@@ -326,8 +365,32 @@ The maximum QVD file size you can handle with qvdjs depends on several factors a
326
365
  - **Operation Type**: Reading typically uses less memory than writing, especially when using lazy loading (`maxRows` option). Writing requires building complete symbol and index tables in memory.
327
366
 
328
367
  - **Practical Guidance**:
329
- - For typical business data with moderate cardinality, files up to 1-2GB usually work well with default Node.js settings
330
- - High-cardinality data (unique values in most rows) may limit you to smaller files (hundreds of MB)
368
+ - **Count cells, not megabytes.** What exhausts the heap on a full load is one array per row plus
369
+ one slot per cell, so a file's size on disk predicts very little. On a default 4 GB heap the
370
+ guard currently accepts roughly:
371
+
372
+ | Columns | Rows | Cells |
373
+ | ------- | ------ | ----- |
374
+ | 5 | 30.5 M | 153 M |
375
+ | 10 | 22.5 M | 225 M |
376
+ | 20 | 14.7 M | 295 M |
377
+ | 40 | 8.7 M | 349 M |
378
+
379
+ These are about 4× what they were before 2026-09-11, for two reasons: the bitwise decode
380
+ (#134) cut what a read actually needs by roughly 2.4×, and the guard's model was recalibrated
381
+ against that — it had been over-estimating a real read by 4.8×. The guard still errs high,
382
+ admitting a load only at 1.3–2× the heap it measurably needs, because under-estimating means
383
+ an uncatchable abort while over-estimating means a catchable refusal. Doubling the heap with
384
+ `--max-old-space-size` roughly doubles these numbers.
385
+
386
+ - **A columnar read is not bounded by any of this.** `QvdColumnTable.fromQvd()` stores codes in
387
+ typed arrays, which live outside the V8 heap, so the heap ceiling barely applies: the 38 MB
388
+ taxi fixture reads columnar in a 15 MB heap, where the same file as rows needs 367 MB. If you
389
+ are hitting the table above, reading the columns you need is usually the answer rather than a
390
+ bigger heap.
391
+
392
+ - High-cardinality data (unique values in most rows) also loads the symbol table, so the ceiling
393
+ drops further — subtract about 90 MB of heap per million distinct values
331
394
  - **With increased heap** (e.g., 16GB+), you can handle proportionally larger files by adjusting `memorySafetyFactor`
332
395
  - Use lazy loading (`maxRows` option) when possible to reduce memory footprint when reading
333
396
  - Monitor memory usage with tools like `process.memoryUsage()` for your specific use cases
@@ -432,7 +495,21 @@ These optimizations can reduce write times by 80-90% for large datasets (100K+ r
432
495
 
433
496
  ### Working with Metadata
434
497
 
435
- The library provides full access to QVD file and field metadata:
498
+ To inspect a file's schema without reading its data, use `readMetadata`. It stops at the XML
499
+ header, so it costs the same for a 40 MB file as for a 40 GB one:
500
+
501
+ ```javascript
502
+ import {QvdDataFrame} from 'qvdjs';
503
+
504
+ const {columns, rowCount, fields} = await QvdDataFrame.readMetadata('path/to/file.qvd');
505
+
506
+ console.log(`${rowCount} rows x ${columns.length} columns`);
507
+ ```
508
+
509
+ Note that `fromQvd(path, {maxRows: 0})` is not equivalent — it loads no rows, but still reads and
510
+ parses the whole symbol table.
511
+
512
+ For metadata alongside the data, every accessor below is available on a loaded frame:
436
513
 
437
514
  ```javascript
438
515
  import {QvdDataFrame} from 'qvdjs';
@@ -577,13 +654,13 @@ try {
577
654
 
578
655
  Honest boundaries rather than an issue list — these are the ones that change what you can do.
579
656
 
580
- | Limitation | What it means in practice | Tracked as |
581
- | ---------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | --------------------------------------------------------- |
582
- | **Full loads stop at 2 GiB** | Use `maxRows` above that; the lazy path reads in chunks and has no such ceiling. Failure is a raw Node `RangeError`, not a `QvdError`. | [#122](https://github.com/ptarmiganlabs/qvdjs/issues/122) |
583
- | **The memory estimate is based on symbol table size, not on rows × columns** | A file with a modest symbol table but very many rows can pass every check and then exhaust the heap while materialising rows. Loading tens of millions of rows, size the heap for the result rather than for the file. | [#121](https://github.com/ptarmiganlabs/qvdjs/issues/121) |
584
- | **Dual values are not round-tripped as pairs** | Qlik stores dates and timestamps as a number plus a display string. A data frame exposes one of the two, so a read-modify-write does not reproduce the original dual — a `BirthDate` field comes back as `24205`, not as a formatted date. | [#138](https://github.com/ptarmiganlabs/qvdjs/issues/138) |
585
- | **Writes are not atomic** | `toQvd()` writes in place. A failure part-way through, or two writers targeting one path, leaves the previous file damaged rather than intact. Write to a temporary path and rename if that matters. | [#129](https://github.com/ptarmiganlabs/qvdjs/issues/129) |
586
- | **Writer input is not validated** | Unsupported value types can produce a raw `TypeError`, or a file that is written but does not read back as intended. Feed `fromDict()` numbers, strings and `null`. | [#130](https://github.com/ptarmiganlabs/qvdjs/issues/130) |
657
+ | Limitation | What it means in practice | Tracked as |
658
+ | ------------------------------------------------ | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | --------------------------------------------------------- |
659
+ | **Full loads stop at 2 GiB** | Use `maxRows` above that; the lazy path reads in chunks and has no such ceiling. Failure is a raw Node `RangeError`, not a `QvdError`. | [#122](https://github.com/ptarmiganlabs/qvdjs/issues/122) |
660
+ | **A full load materialises every row in memory** | Rows are stored as one JavaScript array per row with a boxed value per cell, measured at roughly 11 bytes per cell. That, not file size, sets the ceiling in the table above. `maxRows` reads a prefix, and `QvdColumnTable` avoids row materialisation entirely — but there is still no streaming or chunked read of an arbitrary window. | [#140](https://github.com/ptarmiganlabs/qvdjs/issues/140) |
661
+ | **Dual values are not round-tripped as pairs** | Qlik stores dates and timestamps as a number plus a display string. A data frame exposes one of the two, so a read-modify-write does not reproduce the original dual — a `BirthDate` field comes back as `24205`, not as a formatted date. | [#138](https://github.com/ptarmiganlabs/qvdjs/issues/138) |
662
+ | **Writes are not atomic** | `toQvd()` writes in place. A failure part-way through, or two writers targeting one path, leaves the previous file damaged rather than intact. Write to a temporary path and rename if that matters. | [#129](https://github.com/ptarmiganlabs/qvdjs/issues/129) |
663
+ | **Writer input is not validated** | Unsupported value types can produce a raw `TypeError`, or a file that is written but does not read back as intended. Feed `fromDict()` numbers, strings and `null`. | [#130](https://github.com/ptarmiganlabs/qvdjs/issues/130) |
587
664
 
588
665
  Numeric-looking strings are currently coerced to numbers on read, which is being reconsidered as a breaking
589
666
  change in [#120](https://github.com/ptarmiganlabs/qvdjs/issues/120). Empty and whitespace-only strings are _not_
@@ -722,7 +799,7 @@ to a `QvdDataFrame` instance.
722
799
  - `options` (object, optional): Loading options
723
800
  - `maxRows` (number, optional): Maximum number of rows to load. If not specified, all rows are loaded. This is useful for loading only a subset of data from large QVD files to improve performance and reduce memory usage.
724
801
  - `allowedDir` (string, optional): Base directory for file access validation. Defaults to current working directory (CWD). The file path must resolve to a location within this directory to prevent path traversal attacks. Set to a specific directory in production environments with user-provided paths.
725
- - `memorySafetyFactor` (number, optional): Fraction (0.0-1.0) of the memory budget a load may use. Default is 0.3 (30%). The budget is the smaller of the V8 heap limit and any container memory limit; see [QVD File Size Limitations](#qvd-file-size-limitations). Increase this (e.g., to 0.5 or 0.7) when running with a larger heap via `--max-old-space-size`. **`0` disables the memory check entirely.**
802
+ - `memorySafetyFactor` (number, optional): Fraction (0.0-1.0) of the memory budget a load may use. Default is 0.8 (80%). The budget is the smaller of the V8 heap limit and any container memory limit; see [QVD File Size Limitations](#qvd-file-size-limitations). Increase this (e.g., to 0.5 or 0.7) when running with a larger heap via `--max-old-space-size`. **`0` disables the memory check entirely.**
726
803
  - `symbolFilteringThreshold` (number, optional): Symbol table size, in bytes, above which a lazy load switches to the two-pass filtering path. Defaults to 50 MB, the point where the extra analysis pass pays for itself. Lower it to use filtering on smaller files, raise it to keep the simpler single-pass read for longer.
727
804
 
728
805
  **Example:**
@@ -741,7 +818,7 @@ const dfSecure = await QvdDataFrame.fromQvd('reports/sales.qvd', {
741
818
 
742
819
  // Load with increased memory usage for large heap configurations
743
820
  const dfLarge = await QvdDataFrame.fromQvd('large-file.qvd', {
744
- memorySafetyFactor: 0.5, // Use 50% of the budget instead of the default 30%
821
+ memorySafetyFactor: 0.9, // Use 90% of the budget instead of the default 80%
745
822
  });
746
823
 
747
824
  // Manage memory yourself: skip the check entirely
@@ -753,6 +830,124 @@ console.log(preview.loadStats.symbolFiltering); // true when the two-pass path r
753
830
  console.log(preview.loadStats.symbolsKept); // how many symbols it kept
754
831
  ```
755
832
 
833
+ #### `static readMetadata(path: string, options?: object): Promise<object>`
834
+
835
+ Reads a QVD file's schema and header metadata without reading its data. The cost is the same
836
+ whatever the file's size, because it stops at the XML header and never touches the symbol or
837
+ index tables.
838
+
839
+ This is not the same as `fromQvd(path, {maxRows: 0})`. That loads no rows but still reads and
840
+ parses the entire symbol table — 0.4 MB on a 38 MB file, but 15 MB on a high-cardinality one, and
841
+ it grows with the data. Measured on a 200,000-row file where every value is distinct,
842
+ `readMetadata` is **352× faster**, and unlike `{maxRows: 0}` it does not get slower as the file
843
+ grows.
844
+
845
+ **Parameters:**
846
+
847
+ - `path` (string): The path to the QVD file.
848
+ - `options` (object, optional):
849
+ - `allowedDir` (string, optional): Base directory for file access validation, applied exactly as
850
+ it is for `fromQvd`.
851
+
852
+ **Returns** a plain object — deliberately not a `QvdDataFrame`, since one with `data: []` would be
853
+ indistinguishable from an empty file at the call site:
854
+
855
+ | Property | Type | Description |
856
+ | -------------- | ---------- | ----------------------------------------------------------------------------------- |
857
+ | `columns` | `string[]` | Field names, in file order. |
858
+ | `rowCount` | `number` | Rows the **file** declares. Nothing was loaded; this is not a count of rows read. |
859
+ | `columnCount` | `number` | Number of fields. |
860
+ | `fields` | `object[]` | Per-field metadata, same shape and order as `getFieldMetadata()` on a loaded frame. |
861
+ | `fileMetadata` | `object` | Same shape as the `fileMetadata` accessor on a loaded frame. |
862
+ | `metadata` | `object` | The raw `QvdTableHeader`, as `metadata` gives it. |
863
+
864
+ **Example:**
865
+
866
+ ```javascript
867
+ // What is in this file, without reading any of it
868
+ const {columns, rowCount, fields} = await QvdDataFrame.readMetadata('sales.qvd');
869
+
870
+ console.log(`${rowCount} rows x ${columns.length} columns`);
871
+
872
+ for (const field of fields) {
873
+ console.log(`${field.fieldName}: ${field.noOfSymbols} distinct values`);
874
+ }
875
+
876
+ // Decide whether it is worth loading at all
877
+ if (rowCount < 1_000_000) {
878
+ const df = await QvdDataFrame.fromQvd('sales.qvd');
879
+ }
880
+ ```
881
+
882
+ ### QvdColumnTable
883
+
884
+ `QvdColumnTable` reads a QVD **as columns instead of rows**. It uses the same decoder and the
885
+ same symbol resolution as `QvdDataFrame.fromQvd` — it simply stops before building rows, and
886
+ keeps what the decoder already produced: one `Int32Array` of codes per field, and one resolved
887
+ value per _distinct_ symbol.
888
+
889
+ Measured on the bundled 1.7 M × 20 taxi fixture, each in its own process so only one
890
+ representation is alive:
891
+
892
+ | | Live memory | Sum one column |
893
+ | ------------------------ | ----------- | -------------- |
894
+ | `QvdDataFrame.fromQvd` | 385 MiB | 26.8 ms |
895
+ | `QvdColumnTable.fromQvd` | **141 MiB** | **3.6 ms** |
896
+
897
+ Use it when you want to scan or aggregate a few columns of a large file. Use `QvdDataFrame` when
898
+ you want rows. It does not convert to a data frame, deliberately — holding both representations
899
+ is the one configuration in which this costs more than it saves, and re-reading a file as rows
900
+ costs no more than reading it as rows always did.
901
+
902
+ ```javascript
903
+ import {QvdColumnTable} from 'qvdjs';
904
+
905
+ const table = await QvdColumnTable.fromQvd('trips.qvd');
906
+ const fare = table.column('fare');
907
+
908
+ // The fastest scan: both sides are contiguous typed arrays and the dictionary fits in cache.
909
+ const codes = fare.codes;
910
+ const values = fare.numericSymbols();
911
+ let total = 0;
912
+
913
+ for (let row = 0; row < codes.length; row++) {
914
+ const code = codes[row];
915
+ if (code >= 0) {
916
+ const value = values[code];
917
+ if (!Number.isNaN(value)) total += value;
918
+ }
919
+ }
920
+
921
+ // Or, more simply
922
+ for (const value of fare) {
923
+ /* number | string | null */
924
+ }
925
+ console.log(fare.at(0), fare.toArray().length);
926
+ ```
927
+
928
+ **`QvdColumnTable`** — `static fromQvd(path, options)` (same options as `QvdDataFrame.fromQvd`),
929
+ `column(name)`, `columns`, `rowCount`, `shape`, `metadata`, `loadStats`.
930
+
931
+ **`QvdColumn`** — what `column(name)` returns, frozen:
932
+
933
+ | Member | Type | Description |
934
+ | -------------------------- | -------------------- | ---------------------------------------------------------------------------------------------------------------- |
935
+ | `name` | `string` | The field name. |
936
+ | `length` | `number` | Rows in the column. |
937
+ | `codes` | `Int32Array` | One stored index per row. **Negative means NULL.** The table's own array — treat as read-only. |
938
+ | `symbols` | `ReadonlyArray<any>` | The distinct values, indexed by the codes. One entry per distinct value, not per row. |
939
+ | `at(row)` | `any` | The value of one row, or `null`. |
940
+ | `[Symbol.iterator]` | | Iterates values without materialising the column. |
941
+ | `toArray()` | `Array<any>` | Lossless; cell-for-cell what `data[row][column]` holds. Caller-owned. |
942
+ | `numericSymbols()` | `Float64Array` | The **dictionary** as numbers, NaN for non-numeric. One entry per distinct value — 17 KB for a 1.7 M-row column. |
943
+ | `toFloat64Array(options?)` | `Float64Array` | One number **per row**. Lossy, so it throws on a non-numeric value unless you pass `{onNonNumeric: 'nan'}`. |
944
+
945
+ `toFloat64Array` refuses by default rather than writing NaN because on real QVDs the lossy case
946
+ is common, not exceptional: in the bundled taxi fixture **no column is strictly numeric** —
947
+ `dropoff_census_tract` is 43 % empty strings, and four columns are 100 % strings. Silently
948
+ turning two fifths of a column into NaN would erase Qlik's distinction between a blank and a
949
+ number. `numericSymbols()` is the cheap conversion; `toFloat64Array()` is the expensive one.
950
+
756
951
  #### `static fromDict(dict: object): Promise<QvdDataFrame>`
757
952
 
758
953
  The static method `QvdDataFrame.fromDict` constructs a data frame from a dictionary. The dictionary must contain the columns and