qvdjs 0.6.2 → 0.8.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -1,6 +1,6 @@
1
1
  # qvdjs
2
2
 
3
- > Utility library for reading/writing Qlik View Data (QVD) files in JavaScript.
3
+ > Utility library for reading/writing Qlik Sense and QlikView (QVD) files in JavaScript/Node.js
4
4
 
5
5
  ## ⚠️ Important Disclaimer
6
6
 
@@ -28,6 +28,10 @@ structure and vice versa. The library is written to be used in a Node.js environ
28
28
  - [Install](#install)
29
29
  - [Usage](#usage)
30
30
  - [Lazy Loading](#lazy-loading)
31
+ - [Important: Symbol Table and High-Cardinality Fields](#important-symbol-table-and-high-cardinality-fields)
32
+ - [Performance Optimizations](#performance-optimizations)
33
+ - [QVD File Size Limitations](#qvd-file-size-limitations)
34
+ - [Why Safety Limits Exist](#why-safety-limits-exist)
31
35
  - [Progress Tracking for Large File Writes](#progress-tracking-for-large-file-writes)
32
36
  - [Working with Metadata](#working-with-metadata)
33
37
  - [Security Considerations](#security-considerations)
@@ -123,12 +127,162 @@ console.log(df.shape); // [1000, numberOfColumns]
123
127
  - The library reads only the header, symbol table, and the first N rows from the index table
124
128
  - This provides significant memory savings and faster loading times for large files
125
129
 
126
- This is particularly useful for:
130
+ #### Important: Symbol Table and High-Cardinality Fields
131
+
132
+ The QVD format stores data in two parts:
133
+
134
+ 1. **Symbol table**: Contains ALL unique values for ALL fields (must be fully loaded)
135
+ 2. **Index table**: Contains row-by-row indices into the symbol table (can be partially loaded with `maxRows`)
136
+
137
+ ⚠️ **Performance Impact of High-Cardinality Fields:**
138
+
139
+ If your QVD file contains fields with many unique values (high cardinality), such as:
140
+
141
+ - Unique IDs (OrderID, TransactionID, UUID)
142
+ - Timestamps with millisecond precision
143
+ - Unique text fields
144
+
145
+ The symbol table can become very large and **must be read completely** even when using `maxRows`. This means:
146
+
147
+ - **Small symbol table** (fields with reusable values): Fast loading regardless of file size
148
+ - Example: A 500MB file with only 100KB symbol table loads in ~100ms for 5000 rows ✅
149
+ - **Large symbol table** (fields with unique values per row): Slow loading even with `maxRows`
150
+ - Example: A 500MB file with 266MB symbol table takes ~6 seconds for any row count ❌
151
+
152
+ To check your QVD's symbol table size, look at the `NoOfSymbols` in field metadata - values close to the total row count indicate high cardinality.
153
+
154
+ **When lazy loading works best:**
127
155
 
128
156
  - Previewing data from very large QVD files without loading the entire file into memory
129
- - Reducing memory consumption when working with large files
130
- - Faster loading times when you only need a subset of the data
157
+ - Files where most fields have reusable values (low cardinality)
131
158
  - Data exploration and schema inspection of large datasets
159
+ - Faster loading times when you only need a subset of the data
160
+
161
+ ### Performance Optimizations
162
+
163
+ The library includes intelligent symbol table parsing that dramatically improves performance when loading partial data with `maxRows`:
164
+
165
+ **Smart Symbol Loading:**
166
+
167
+ When you specify `maxRows`, the library:
168
+
169
+ 1. Analyzes which symbols are actually needed for the requested rows
170
+ 2. Parses **only** those symbols from the symbol table
171
+ 3. Skips parsing unused symbols entirely (not just filtering after parsing)
172
+
173
+ **Performance Impact:**
174
+
175
+ Compared to parsing all symbols regardless of `maxRows`:
176
+
177
+ - **~10x faster load times** - Loading 5,000 rows from a 20M row file: 8.3 seconds → 0.9 seconds
178
+ - **~10x less memory usage** - Same operation: 1,616 MB → 180 MB
179
+ - **Native or better efficiency** - Memory overhead reduced from 6.0x to 0.7x of raw data size
180
+
181
+ **Real-World Example:**
182
+
183
+ ```javascript
184
+ // File: orders_20m.qvd (20M rows, 266MB symbol table)
185
+ const df = await QvdDataFrame.fromQvd('orders_20m.qvd', {maxRows: 5000});
186
+ // Load time: ~0.9 seconds
187
+ // Memory usage: ~180 MB
188
+ // Only 17,228 out of millions of symbols parsed
189
+ ```
190
+
191
+ **Key Benefits:**
192
+
193
+ - Much faster previews of large QVD files
194
+ - Lower memory footprint for data exploration
195
+ - Efficient handling of files with large symbol tables
196
+ - Automatic optimization - no configuration needed
197
+
198
+ This optimization is particularly effective for files with many unique values (high cardinality) where the symbol table is large but you only need to preview a small portion of the data.
199
+
200
+ ### QVD File Size Limitations
201
+
202
+ **Simple Explanation:**
203
+
204
+ Node.js has memory limits that affect how large QVD files you can work with. By default, Node.js can use up to about 4GB of memory. This means if you try to load a very large QVD file, you might run out of memory and get an error. Think of it like trying to open a very large document on a computer with limited RAM - if the document is too big, it won't open.
205
+
206
+ **What happens when files are too large:**
207
+
208
+ - **When opening large files**: If a QVD file exceeds available memory, Node.js will throw an out-of-memory error (typically "JavaScript heap out of memory" or "FATAL ERROR: Reached heap limit"). The process will crash before completing the file load.
209
+ - **When saving large files**: Writing very large QVD files can similarly exhaust memory during symbol table and index table construction, causing the same out-of-memory errors before the file is written to disk.
210
+ - **Performance degradation**: Even before running out of memory completely, you may notice significant slowdowns, high memory usage, and system swapping as files approach memory limits.
211
+
212
+ The good news is that Node.js memory limits can be increased (though there will always be some limit), and the actual file size you can handle depends on your data.
213
+
214
+ The most common question at this point is usually:
215
+
216
+ > "How large of a QVD file can I work with using qvdjs?"
217
+
218
+ There is unfortunately no simple answer to this question, as it very much depends on the characteristics of what data is inside the QVD file. See below for more details.
219
+
220
+ **Technical Details:**
221
+
222
+ The maximum QVD file size you can handle with qvdjs depends on several factors and there is no single fixed limit:
223
+
224
+ - **Node.js Memory Limits**: By default, Node.js limits heap memory to approximately 4GB (varies by architecture and Node.js version). The library **automatically detects your configured heap size** and scales its safety limits accordingly. You can increase heap size using the `--max-old-space-size` flag (e.g., `node --max-old-space-size=16384 script.js` for 16GB), and qvdjs will automatically allow larger files.
225
+
226
+ For larger heap configurations, you can also adjust the `memorySafetyFactor` option (default 0.3 = 30%) to make more efficient use of available memory:
227
+
228
+ ```javascript
229
+ const df = await QvdDataFrame.fromQvd('large-file.qvd', {
230
+ memorySafetyFactor: 0.5, // Use 50% of heap instead of default 30%
231
+ });
232
+ ```
233
+
234
+ To phrase it differently: There is no magic here - viewing a 40 GB QVD on a laptop with 24 GB RAM will not work. That laptop may in fact struggle with QVD files larger than 4-6 GB depending on data characteristics - or happily work with 10+ GB files if the data is very friendly.
235
+
236
+ - **Data Characteristics**: The actual memory consumption depends _heavily_ on what's inside your QVD:
237
+ - **Field Cardinality**: Files with high-cardinality fields (many unique values per field, like unique IDs or timestamps) require more memory for symbol tables. This is usually the biggest factor, at least if there are many rows too.
238
+ - **Number of Rows**: More rows mean larger index tables in memory
239
+ - **Number of Fields**: More columns increase overall memory requirements
240
+ - **Data Types**: String data generally uses more memory than numeric data
241
+
242
+ - **Operation Type**: Reading typically uses less memory than writing, especially when using lazy loading (`maxRows` option). Writing requires building complete symbol and index tables in memory.
243
+
244
+ - **Practical Guidance**:
245
+ - For typical business data with moderate cardinality, files up to 1-2GB usually work well with default Node.js settings
246
+ - High-cardinality data (unique values in most rows) may limit you to smaller files (hundreds of MB)
247
+ - **With increased heap** (e.g., 16GB+), you can handle proportionally larger files by adjusting `memorySafetyFactor`
248
+ - Use lazy loading (`maxRows` option) when possible to reduce memory footprint when reading
249
+ - Monitor memory usage with tools like `process.memoryUsage()` for your specific use cases
250
+ - Consider processing large datasets in chunks or using streaming approaches if you hit memory limits. Clever things can be done by doing multiple passes over the file instead of loading everything at once.
251
+
252
+ If you consistently work with very large QVD files, consider increasing Node.js memory limits (with matching `memorySafetyFactor` adjustment) or splitting your data into multiple smaller QVD files.
253
+
254
+ #### Why Safety Limits Exist
255
+
256
+ The library implements **dynamic safety limits** to prevent catastrophic crashes. Without these limits:
257
+
258
+ - **Your application will crash hard** - Node.js terminates with `FATAL ERROR: Reached heap limit` when attempting to load files that are too large
259
+ - **No error handling is possible** - JavaScript try-catch blocks cannot intercept out-of-memory (OOM) crashes at the V8 engine level
260
+ - **The entire process dies** - Not just the QVD operation, but your entire application terminates ungracefully
261
+
262
+ The safety limits **prevent these crashes** by checking available memory _before_ attempting to load files, throwing graceful `QvdValidationError` exceptions that you can catch and handle. While the multi-tier safety system may seem complex, it ensures your application stays running and provides helpful error messages with recommendations (like using `maxRows` parameter or increasing heap size) instead of cryptic fatal errors.
263
+
264
+ **Example without safety limits:**
265
+
266
+ ```javascript
267
+ // Process crashes with no chance to recover
268
+ const df = await QvdDataFrame.fromQvd('huge-file.qvd'); // 💥 FATAL ERROR
269
+ // Your application is now terminated
270
+ ```
271
+
272
+ **Example with safety limits:**
273
+
274
+ ```javascript
275
+ try {
276
+ const df = await QvdDataFrame.fromQvd('huge-file.qvd');
277
+ } catch (error) {
278
+ if (error.name === 'QvdValidationError') {
279
+ console.log('File too large, trying with maxRows:', error.context.recommendedMaxRows);
280
+ // ✅ Your application continues running
281
+ }
282
+ }
283
+ ```
284
+
285
+ For more technical details about the memory safety system, including specific thresholds and the four-tier protection model, see [docs/DYNAMIC_SAFETY_LIMITS.md](docs/DYNAMIC_SAFETY_LIMITS.md). For practical examples, see [docs/examples/heap-scaling-example.md](docs/examples/heap-scaling-example.md).
132
286
 
133
287
  ### Progress Tracking for Large File Writes
134
288
 
@@ -173,7 +327,7 @@ The library uses optimized algorithms for large dataset processing:
173
327
  - **Map-based lookups**: O(1) symbol index lookups instead of O(n) findIndex operations
174
328
  - **Reduced algorithmic complexity**: From O(n×m×s) to O(n×m) where n=rows, m=columns, s=symbols
175
329
 
176
- These optimizations can reduce write times by 80-90% for large datasets (100K+ rows).
330
+ These optimizations can reduce write times by 80-90% for large datasets (100K+ rows), compared to earlier versions of the library.
177
331
 
178
332
  ### Working with Metadata
179
333
 
@@ -319,7 +473,11 @@ and the data types of the fields.
319
473
 
320
474
  The symbol table contains the distinct/unique values of the fields and is located directly after the XML header. The order
321
475
  of columns in the symbol table corresponds to the order of the fields in the XML header. The length and offset of the
322
- symbol sections of each column are also stored in the XML header. Each symbol section consist of the unique symbols of the
476
+ symbol sections of each column are also stored in the XML header.
477
+
478
+ **Important**: The offset values in the XML header are **relative to the start of the symbol table section**, not absolute file positions. For example, if a field has `Offset=1143`, this means its symbol data starts 1143 bytes after the symbol table section begins (which itself starts immediately after the XML header ends).
479
+
480
+ Each symbol section consists of the unique symbols of the
323
481
  respective column. The type of a single symbol is determined by a type byte prefixed to the respective symbol value. The
324
482
  following type of symbols are supported:
325
483
 
@@ -428,6 +586,7 @@ to a `QvdDataFrame` instance.
428
586
  - `options` (object, optional): Loading options
429
587
  - `maxRows` (number, optional): Maximum number of rows to load. If not specified, all rows are loaded. This is useful for loading only a subset of data from large QVD files to improve performance and reduce memory usage.
430
588
  - `allowedDir` (string, optional): Base directory for file access validation. Defaults to current working directory (CWD). The file path must resolve to a location within this directory to prevent path traversal attacks. Set to a specific directory in production environments with user-provided paths.
589
+ - `memorySafetyFactor` (number, optional): Memory safety factor (0.0-1.0). Default is 0.3 (30%). Determines what percentage of available memory (or V8 heap limit, whichever is smaller) can be used. Increase this (e.g., to 0.5 or 0.7) when running with larger heap sizes via `--max-old-space-size` to allow processing of proportionally larger files.
431
590
 
432
591
  **Example:**
433
592
 
@@ -442,6 +601,11 @@ const dfLazy = await QvdDataFrame.fromQvd('path/to/file.qvd', {maxRows: 1000});
442
601
  const dfSecure = await QvdDataFrame.fromQvd('reports/sales.qvd', {
443
602
  allowedDir: '/var/data/qvd-files',
444
603
  });
604
+
605
+ // Load with increased memory usage for large heap configurations
606
+ const dfLarge = await QvdDataFrame.fromQvd('large-file.qvd', {
607
+ memorySafetyFactor: 0.5, // Use 50% of heap instead of default 30%
608
+ });
445
609
  ```
446
610
 
447
611
  #### `static fromDict(dict: object): Promise<QvdDataFrame>`