@reportwright/engine 0.11.0 → 0.12.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/CHANGELOG.md CHANGED
@@ -2,6 +2,33 @@
2
2
 
3
3
  All `@reportwright/*` packages (formerly `@pagewrightjs/*`) share one version. The versioning policy is in [CONTRIBUTING.md](CONTRIBUTING.md#versioning-semver).
4
4
 
5
+ ## 0.12.0 — 2026-10-09
6
+
7
+ Streaming PDF: very large reports are written page by page, with flat memory.
8
+
9
+ ### Added
10
+ - **`exportPdfStream(definition, options, sink)`** in `@reportwright/engine`: a report whose body is a flowing table
11
+ (plain or grouped: repeating group headers, subtotals, `RowNumber()`, widow control, keep-together, page breaks) is laid
12
+ out a window of rows at a time and each finished page goes straight to the sink. Peak memory stays flat: about 260 MB
13
+ for the whole render process from 20,000 to 100,000 rows (projected about 250–275 MB at 1,000,000).
14
+ "Page N of M" and footer totals work (a counting pass first). PDF/A-2b, PDF/UA-1 and PDF/A-2a + PDF/UA-1 stream too
15
+ (veraPDF passes). Anything else (images, charts, matrices, rich text, subreports, signing, encryption) falls back to
16
+ `exportPdf`, with the same output.
17
+ - **The server streams large PDFs** (`/api/reports/<id>/pdf`, `/api/v1/render`) with the same row caps, deadline, stall
18
+ timeout and abort handling as the CSV/Excel streaming; the viewer's large-report button gets the streamed file.
19
+ - **`parallel: N`** (Node, opt-in): page ranges painted on worker threads, byte-identical to the serial file. Measure on
20
+ your hardware before relying on it; each worker costs about 150 MB.
21
+ - `examples/stream-1m.mjs` in the engine package: streams a generated million-row ledger to a file and prints peak memory.
22
+ - `planPaged` exported for custom pipelines.
23
+
24
+ ### Changed
25
+ - The PDF painter is shared between `exportPdf` and the streaming writer; `exportPdf` output is byte-for-byte unchanged.
26
+ - The engine package ships its readable bundle without a source map (1.3 MB unpacked).
27
+
28
+ ### Known limits
29
+ - In the npm engine, grouped reports sort rows in memory (the server spills them to disk), so a 1M-row grouped report is
30
+ not flat outside the server yet.
31
+
5
32
  ## 0.11.0 — 2026-10-08
6
33
 
7
34
  Everything since 0.10.0. Items marked **Behaviour change** alter output that a pixel-locked layout or a locale-sensitive
package/README.md CHANGED
@@ -46,6 +46,44 @@ const bundledFonts = defaultFontStore({ fontsUrl: '/pw/fonts/' }); // a missing
46
46
  `timeoutMs` as well; the PDF exporter checks every PNG before decoding it and leaves a damaged one out (`onWarning`).
47
47
  - **Currency.** `C2` with no `currency` on the report uses the locale's own (`en-IN` ₹, `de-DE` €, `en-US` $).
48
48
 
49
+ ## Streaming PDF for large reports
50
+
51
+ `exportPdfStream(definition, options, sink)` writes the PDF as it is laid out: a flowing table is laid out a window of
52
+ rows at a time and each page leaves memory once it is written, so a million-row report's PDF needs about the memory of a
53
+ few pages (measured: peak RSS flat from 20,000 to 100,000 rows, about 270 MB for the whole Node process). The file reads as
54
+ `exportPdf`'s: the same pages, text and positions.
55
+
56
+ ```js
57
+ import fs from 'node:fs';
58
+ import { exportPdfStream } from '@reportwright/engine';
59
+
60
+ const out = fs.createWriteStream('ledger.pdf');
61
+ const r = await exportPdfStream(definition, { parameters: { year: 2026 }, title: 'Ledger' }, out);
62
+ out.end();
63
+ console.log(r.streamed ? `${r.pages} pages, streamed` : `written whole: ${r.why}`);
64
+ ```
65
+
66
+ The sink is a Node `Writable` or anything with `write(bytes)` that may return a promise (back-pressure); it is not
67
+ closed for you. Options are `render()`'s and `exportPdf`'s (`fonts`, `parameters`, `title`, `pdfa`, `tagged`, …), plus
68
+ `window` (rows laid out at once, 1,000), `maxRows` (rows read, 5,000,000) and `parallel` (Node: that many worker threads
69
+ lay out and paint the pages; the file is byte for byte the same; each thread loads the engine and fonts, so it costs
70
+ memory: about 150 MB a thread). Rows come from the report's source a batch at a time: CSV and REST (JSON) sources are read
71
+ as streams, and `sqlRows(source, params)`, an async iterable of row batches, stands in for a database cursor.
72
+
73
+ **What streams:** a body that is one table (with items above or below it), its rows' expressions row by row (fields,
74
+ parameters, functions of the row, `RowNumber()`), groups with headers that repeat, subtotals (`Sum`, `Count`, `Avg`,
75
+ `Min`, `Max` of a row expression, `CountRows()`), keep-together and page breaks, filters and sorts, page headers and
76
+ footers without aggregates ("Page N of M" included: the pages are counted first, so the source is read twice), PDF/A-2b,
77
+ and accessible PDFs (PDF/UA-1, PDF/A-2a; validated with veraPDF).
78
+
79
+ **What is written whole** (`exportPdf`, the result says why): matrices, charts, lists, subreports, images, rich text,
80
+ lookups and aggregates outside footers, more than one data set, group sections (`newSection`), signed or encrypted PDFs.
81
+ Grouped and sorted tables hold their rows in the sort unless `sorter` spills them to disk (the ReportWright server does).
82
+
83
+ `examples/stream-1m.mjs` streams a generated million-row ledger to a file and prints its peak memory:
84
+ `node node_modules/@reportwright/engine/examples/stream-1m.mjs 1000000 ledger.pdf` (`--grouped`, `--tagged`,
85
+ `--parallel=4`, `--plain`).
86
+
49
87
  ## Many reports at once (Node): a worker pool
50
88
 
51
89
  ```js