qvdjs 0.6.2 → 0.7.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +121 -6
- package/dist/index.cjs +617 -276
- package/dist/index.cjs.map +1 -1
- package/dist/index.js +616 -276
- package/dist/index.js.map +1 -1
- package/package.json +6 -4
package/README.md
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# qvdjs
|
|
2
2
|
|
|
3
|
-
> Utility library for reading/writing Qlik
|
|
3
|
+
> Utility library for reading/writing Qlik Sense and QlikView (QVD) files in JavaScript/Node.js
|
|
4
4
|
|
|
5
5
|
## ⚠️ Important Disclaimer
|
|
6
6
|
|
|
@@ -28,6 +28,9 @@ structure and vice versa. The library is written to be used in a Node.js environ
|
|
|
28
28
|
- [Install](#install)
|
|
29
29
|
- [Usage](#usage)
|
|
30
30
|
- [Lazy Loading](#lazy-loading)
|
|
31
|
+
- [Important: Symbol Table and High-Cardinality Fields](#important-symbol-table-and-high-cardinality-fields)
|
|
32
|
+
- [Performance Optimizations](#performance-optimizations)
|
|
33
|
+
- [QVD File Size Limitations](#qvd-file-size-limitations)
|
|
31
34
|
- [Progress Tracking for Large File Writes](#progress-tracking-for-large-file-writes)
|
|
32
35
|
- [Working with Metadata](#working-with-metadata)
|
|
33
36
|
- [Security Considerations](#security-considerations)
|
|
@@ -123,12 +126,120 @@ console.log(df.shape); // [1000, numberOfColumns]
|
|
|
123
126
|
- The library reads only the header, symbol table, and the first N rows from the index table
|
|
124
127
|
- This provides significant memory savings and faster loading times for large files
|
|
125
128
|
|
|
126
|
-
|
|
129
|
+
#### Important: Symbol Table and High-Cardinality Fields
|
|
130
|
+
|
|
131
|
+
The QVD format stores data in two parts:
|
|
132
|
+
|
|
133
|
+
1. **Symbol table**: Contains ALL unique values for ALL fields (must be fully loaded)
|
|
134
|
+
2. **Index table**: Contains row-by-row indices into the symbol table (can be partially loaded with `maxRows`)
|
|
135
|
+
|
|
136
|
+
⚠️ **Performance Impact of High-Cardinality Fields:**
|
|
137
|
+
|
|
138
|
+
If your QVD file contains fields with many unique values (high cardinality), such as:
|
|
139
|
+
|
|
140
|
+
- Unique IDs (OrderID, TransactionID, UUID)
|
|
141
|
+
- Timestamps with millisecond precision
|
|
142
|
+
- Unique text fields
|
|
143
|
+
|
|
144
|
+
The symbol table can become very large and **must be read completely** even when using `maxRows`. This means:
|
|
145
|
+
|
|
146
|
+
- **Small symbol table** (fields with reusable values): Fast loading regardless of file size
|
|
147
|
+
- Example: A 500MB file with only 100KB symbol table loads in ~100ms for 5000 rows ✅
|
|
148
|
+
- **Large symbol table** (fields with unique values per row): Slow loading even with `maxRows`
|
|
149
|
+
- Example: A 500MB file with 266MB symbol table takes ~6 seconds for any row count ❌
|
|
150
|
+
|
|
151
|
+
To check your QVD's symbol table size, look at the `NoOfSymbols` in field metadata - values close to the total row count indicate high cardinality.
|
|
152
|
+
|
|
153
|
+
**When lazy loading works best:**
|
|
127
154
|
|
|
128
155
|
- Previewing data from very large QVD files without loading the entire file into memory
|
|
129
|
-
-
|
|
130
|
-
- Faster loading times when you only need a subset of the data
|
|
156
|
+
- Files where most fields have reusable values (low cardinality)
|
|
131
157
|
- Data exploration and schema inspection of large datasets
|
|
158
|
+
- Faster loading times when you only need a subset of the data
|
|
159
|
+
|
|
160
|
+
### Performance Optimizations
|
|
161
|
+
|
|
162
|
+
The library includes intelligent symbol table parsing that dramatically improves performance when loading partial data with `maxRows`:
|
|
163
|
+
|
|
164
|
+
**Smart Symbol Loading:**
|
|
165
|
+
|
|
166
|
+
When you specify `maxRows`, the library:
|
|
167
|
+
|
|
168
|
+
1. Analyzes which symbols are actually needed for the requested rows
|
|
169
|
+
2. Parses **only** those symbols from the symbol table
|
|
170
|
+
3. Skips parsing unused symbols entirely (not just filtering after parsing)
|
|
171
|
+
|
|
172
|
+
**Performance Impact:**
|
|
173
|
+
|
|
174
|
+
Compared to parsing all symbols regardless of `maxRows`:
|
|
175
|
+
|
|
176
|
+
- **~10x faster load times** - Loading 5,000 rows from a 20M row file: 8.3 seconds → 0.9 seconds
|
|
177
|
+
- **~10x less memory usage** - Same operation: 1,616 MB → 180 MB
|
|
178
|
+
- **Native or better efficiency** - Memory overhead reduced from 6.0x to 0.7x of raw data size
|
|
179
|
+
|
|
180
|
+
**Real-World Example:**
|
|
181
|
+
|
|
182
|
+
```javascript
|
|
183
|
+
// File: orders_20m.qvd (20M rows, 266MB symbol table)
|
|
184
|
+
const df = await QvdDataFrame.fromQvd('orders_20m.qvd', {maxRows: 5000});
|
|
185
|
+
// Load time: ~0.9 seconds
|
|
186
|
+
// Memory usage: ~180 MB
|
|
187
|
+
// Only 17,228 out of millions of symbols parsed
|
|
188
|
+
```
|
|
189
|
+
|
|
190
|
+
**Key Benefits:**
|
|
191
|
+
|
|
192
|
+
- Much faster previews of large QVD files
|
|
193
|
+
- Lower memory footprint for data exploration
|
|
194
|
+
- Efficient handling of files with large symbol tables
|
|
195
|
+
- Automatic optimization - no configuration needed
|
|
196
|
+
|
|
197
|
+
This optimization is particularly effective for files with many unique values (high cardinality) where the symbol table is large but you only need to preview a small portion of the data.
|
|
198
|
+
|
|
199
|
+
### QVD File Size Limitations
|
|
200
|
+
|
|
201
|
+
**Simple Explanation:**
|
|
202
|
+
|
|
203
|
+
Node.js has memory limits that affect how large QVD files you can work with. By default, Node.js can use up to about 4GB of memory. This means if you try to load a very large QVD file, you might run out of memory and get an error. Think of it like trying to open a very large document on a computer with limited RAM - if the document is too big, it won't open.
|
|
204
|
+
|
|
205
|
+
**What happens when files are too large:**
|
|
206
|
+
|
|
207
|
+
- **When opening large files**: If a QVD file exceeds available memory, Node.js will throw an out-of-memory error (typically "JavaScript heap out of memory" or "FATAL ERROR: Reached heap limit"). The process will crash before completing the file load.
|
|
208
|
+
- **When saving large files**: Writing very large QVD files can similarly exhaust memory during symbol table and index table construction, causing the same out-of-memory errors before the file is written to disk.
|
|
209
|
+
- **Performance degradation**: Even before running out of memory completely, you may notice significant slowdowns, high memory usage, and system swapping as files approach memory limits.
|
|
210
|
+
|
|
211
|
+
The good news is that Node.js memory limits can be increased (though there will always be some limit), and the actual file size you can handle depends on your data.
|
|
212
|
+
|
|
213
|
+
The most common question at this point is usually:
|
|
214
|
+
|
|
215
|
+
> "How large of a QVD file can I work with using qvdjs?"
|
|
216
|
+
|
|
217
|
+
There is unfortunately no simple answer to this question, as it very much depends on the characteristics of what data is inside the QVD file. See below for more details.
|
|
218
|
+
|
|
219
|
+
**Technical Details:**
|
|
220
|
+
|
|
221
|
+
The maximum QVD file size you can handle with qvdjs depends on several factors and there is no single fixed limit:
|
|
222
|
+
|
|
223
|
+
- **Node.js Memory Limits**: By default, Node.js limits heap memory to approximately 4GB (varies by architecture and Node.js version). You can increase this using the `--max-old-space-size` flag (e.g., `node --max-old-space-size=8192 script.js` for 8GB), but physical RAM and system architecture will ultimately constrain you.
|
|
224
|
+
To phrase it differently: There is no magic here - viewing a 40 GB QVD on a laptop with 24 GB RAM will not work.
|
|
225
|
+
That laptop may in fact struggle with QVD files larger than 4-6 GB depending on data characteristics - or happily work with 10+ GB files if the data is very friendly.
|
|
226
|
+
|
|
227
|
+
- **Data Characteristics**: The actual memory consumption depends _heavily_ on what's inside your QVD:
|
|
228
|
+
- **Field Cardinality**: Files with high-cardinality fields (many unique values per field, like unique IDs or timestamps) require more memory for symbol tables. This is usually the biggest factor, at least if there are many rows too.
|
|
229
|
+
- **Number of Rows**: More rows mean larger index tables in memory
|
|
230
|
+
- **Number of Fields**: More columns increase overall memory requirements
|
|
231
|
+
- **Data Types**: String data generally uses more memory than numeric data
|
|
232
|
+
|
|
233
|
+
- **Operation Type**: Reading typically uses less memory than writing, especially when using lazy loading (`maxRows` option). Writing requires building complete symbol and index tables in memory.
|
|
234
|
+
|
|
235
|
+
- **Practical Guidance**:
|
|
236
|
+
- For typical business data with moderate cardinality, files up to 1-2GB usually work well with default Node.js settings
|
|
237
|
+
- High-cardinality data (unique values in most rows) may limit you to smaller files (hundreds of MB)
|
|
238
|
+
- Use lazy loading (`maxRows` option) when possible to reduce memory footprint when reading
|
|
239
|
+
- Monitor memory usage with tools like `process.memoryUsage()` for your specific use cases
|
|
240
|
+
- Consider processing large datasets in chunks or using streaming approaches if you hit memory limits. Clever things can be done by doing multiple passes over the file instead of loading everything at once.
|
|
241
|
+
|
|
242
|
+
If you consistently work with very large QVD files, consider increasing Node.js memory limits or splitting your data into multiple smaller QVD files.
|
|
132
243
|
|
|
133
244
|
### Progress Tracking for Large File Writes
|
|
134
245
|
|
|
@@ -173,7 +284,7 @@ The library uses optimized algorithms for large dataset processing:
|
|
|
173
284
|
- **Map-based lookups**: O(1) symbol index lookups instead of O(n) findIndex operations
|
|
174
285
|
- **Reduced algorithmic complexity**: From O(n×m×s) to O(n×m) where n=rows, m=columns, s=symbols
|
|
175
286
|
|
|
176
|
-
These optimizations can reduce write times by 80-90% for large datasets (100K+ rows).
|
|
287
|
+
These optimizations can reduce write times by 80-90% for large datasets (100K+ rows), compared to earlier versions of the library.
|
|
177
288
|
|
|
178
289
|
### Working with Metadata
|
|
179
290
|
|
|
@@ -319,7 +430,11 @@ and the data types of the fields.
|
|
|
319
430
|
|
|
320
431
|
The symbol table contains the distinct/unique values of the fields and is located directly after the XML header. The order
|
|
321
432
|
of columns in the symbol table corresponds to the order of the fields in the XML header. The length and offset of the
|
|
322
|
-
symbol sections of each column are also stored in the XML header.
|
|
433
|
+
symbol sections of each column are also stored in the XML header.
|
|
434
|
+
|
|
435
|
+
**Important**: The offset values in the XML header are **relative to the start of the symbol table section**, not absolute file positions. For example, if a field has `Offset=1143`, this means its symbol data starts 1143 bytes after the symbol table section begins (which itself starts immediately after the XML header ends).
|
|
436
|
+
|
|
437
|
+
Each symbol section consists of the unique symbols of the
|
|
323
438
|
respective column. The type of a single symbol is determined by a type byte prefixed to the respective symbol value. The
|
|
324
439
|
following type of symbols are supported:
|
|
325
440
|
|