@reportwright/engine 0.11.0 → 0.12.1
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +81 -0
- package/README.md +82 -2
- package/dist/index.js +4476 -1786
- package/dist/pdfstreamworker.js +38 -0
- package/dist/types/packages/engine/entry.d.ts +64 -0
- package/dist/types/src/engine/data/guard.d.ts +4 -1
- package/dist/types/src/engine/data/stream.d.ts +48 -0
- package/dist/types/src/engine/index.d.ts +16 -1
- package/dist/types/src/engine/items/chart-kit.d.ts +6 -0
- package/dist/types/src/engine/paged.d.ts +19 -0
- package/dist/types/src/engine/stream.d.ts +56 -0
- package/dist/types/src/engine/text/measure.d.ts +2 -0
- package/dist/types/src/exporters/docx.d.ts +5 -1
- package/dist/types/src/exporters/html.d.ts +9 -5
- package/dist/types/src/exporters/pdfa.d.ts +11 -0
- package/dist/types/src/exporters/pdfpaint.d.ts +99 -0
- package/dist/types/src/exporters/pdfstream.d.ts +199 -0
- package/dist/types/src/exporters/pdfstreamtags.d.ts +45 -0
- package/dist/types/src/exporters/spill.d.ts +20 -0
- package/dist/types/src/exporters/subset.d.ts +1 -0
- package/examples/stream-1m.mjs +100 -0
- package/package.json +3 -2
- package/pool/worker.js +1 -1
- package/dist/index.js.map +0 -6
package/CHANGELOG.md
CHANGED
|
@@ -2,6 +2,87 @@
|
|
|
2
2
|
|
|
3
3
|
All `@reportwright/*` packages (formerly `@pagewrightjs/*`) share one version. The versioning policy is in [CONTRIBUTING.md](CONTRIBUTING.md#versioning-semver).
|
|
4
4
|
|
|
5
|
+
## 0.12.1 — 2026-10-08
|
|
6
|
+
|
|
7
|
+
### Fixed
|
|
8
|
+
- **Word export: long reports.** The document XML goes to the zip in pieces instead of one string: a 500k-row report
|
|
9
|
+
hit V8's string limit (RangeError: Invalid string length). A report with no tables exports without `exportData`.
|
|
10
|
+
- **`exportHtml(model, { fonts, ...options })`** takes the options object like the other exporters; the 0.7 form
|
|
11
|
+
`exportHtml(model, fontStore, options)` still works with a deprecation warning. It threw before ("reading 'get'").
|
|
12
|
+
- **Charts: scatter and bubble axes** no longer pad past zero the data does not stay above (0-100 data gives 0 to 120,
|
|
13
|
+
not -50 to 150). Warnings: a row one series leaves blank is drawn by the others, so it is not reported as not drawn;
|
|
14
|
+
a gantt chart has no series by design and gets no "has no series" warning.
|
|
15
|
+
- **Variable CJK fonts** are cut at Regular (wght 400), not their default Thin instance, in the PDF and HTML subsets.
|
|
16
|
+
- **Data fetch:** gzip, deflate and br bodies are decoded (the byte cap counts decoded bytes); `file:` and `data:` URLs
|
|
17
|
+
are refused with the scheme named, not "its host is made by a parameter".
|
|
18
|
+
- **Format characters** (RLM, LRE/LRI, ZWJ/ZWNJ, soft hyphen) need no glyph: no "No font has these characters" warning,
|
|
19
|
+
and no notdef box in the PDF.
|
|
20
|
+
- **Warnings:** an unnamed item is named by its type (no "undefined:"), and a sub-report warning that validate and the
|
|
21
|
+
render both say is said once.
|
|
22
|
+
- **`parameterOptions()`** lists a parameter's fixed options first, then the fetched ones, each value once (as the viewer does).
|
|
23
|
+
- **HTML Tables: date sort keys** use the wall clock, so the same report gives the same keys in every host time zone.
|
|
24
|
+
|
|
25
|
+
- **Streaming PDF (`exportPdfStream`), from an independent 1M-row verification:**
|
|
26
|
+
- Rows given as an array are read where they are (never copied whole or sorted, tested); JSON given as text is read
|
|
27
|
+
row by row instead of parsed whole (a 300,000-row text source: 910 → 500 MB peak). The README says why a large array
|
|
28
|
+
still shows a large RSS (it is the caller's, and V8 lets garbage grow beside it) and to pass a cursor or text.
|
|
29
|
+
- Grouped and sorted tables: in Node the engine package sorts through a disk spill by default (`spill: { dir, maxMB }`,
|
|
30
|
+
`spill: false`), so a grouped export's memory is flat (300,000 rows: 300 MB, was 670 MB and growing); `presorted:
|
|
31
|
+
true` reads rows that come in the report's order without sorting, checking each against the one before (a row out of
|
|
32
|
+
order stops the export with an error).
|
|
33
|
+
- A data set without declared `fields` streams: its fields are detected from the first 50 rows, as the whole render
|
|
34
|
+
detects them. A report written whole now says why in its first warning, not only in `why`.
|
|
35
|
+
- The window that reaches the last rows is the last one: a table whose footer keeps with its rows could stop with
|
|
36
|
+
"The data changed while the PDF was written".
|
|
37
|
+
- Tagged PDFs are about 27% smaller (a page's borders are one artifact and its cells' text one text object; plain cells
|
|
38
|
+
carry no attributes; the xref stream is predicted). Past a million structure elements the result warns that tools
|
|
39
|
+
which load every object (qpdf, pikepdf) may run out of memory; `tagged.maxElements` caps it.
|
|
40
|
+
- `parallel: N` is documented with its memory (150–300 MB a thread) and when it helps; the example runs serial.
|
|
41
|
+
- The comments of `exportPdfStream` said tagged and PDF/A files do not stream; they do.
|
|
42
|
+
- Note: the non-streaming `exportPdf` still uses pdf-lib, which stays a dependency.
|
|
43
|
+
- **BIRT import: SUM and other aggregates total their column.** A bound column with `aggregateFunction` (SUM, AVE,
|
|
44
|
+
MIN, MAX, FIRST, LAST…) had no listed argument, so it fell back to a row count: `Sum(1, …)` showed 3 for a total of
|
|
45
|
+
250.75. The column's own expression is now the argument. An aggregate with no expression is counted as approximated.
|
|
46
|
+
The Jasper and SSRS importers were checked: they already take the expression.
|
|
47
|
+
- **Docs: no npm provenance badge.** The repository is private, so npm provenance cannot be generated. The badges and
|
|
48
|
+
the claims are removed from the package READMEs and `docs/RELEASING.md`; the publish workflow stays, and its notes say
|
|
49
|
+
provenance needs a public repository.
|
|
50
|
+
- **Docs: the viewer's `data` example** says which shape matches the data set's path: `{ orders: { rows: [...] } }` for
|
|
51
|
+
`"path": "$.rows"`, `{ orders: [...] }` for a data set with no path. A mismatch gives empty tables, silently.
|
|
52
|
+
- **Docs: a missing font** stops the viewer's run with a message naming the font and a Retry button; the README said the
|
|
53
|
+
report draws in a fallback font, which it does not.
|
|
54
|
+
- **Docs: the engine README** says `pw` is in `@reportwright/cli`, and its quick start has a `statement.pw.json` to save.
|
|
55
|
+
- **CLI help** marks `pw audit` and `pw migrate-storage` as server only; the npm CLI does not include them.
|
|
56
|
+
- **Engine text:** the CJK font warning names `npx reportwright-fonts-cjk` first (the app's `npm run fonts:cjk` second),
|
|
57
|
+
and the deprecated positional export forms say they are removed in 1.0.
|
|
58
|
+
|
|
59
|
+
## 0.12.0 — 2026-10-08
|
|
60
|
+
|
|
61
|
+
Streaming PDF: very large reports are written page by page, with flat memory.
|
|
62
|
+
|
|
63
|
+
### Added
|
|
64
|
+
- **`exportPdfStream(definition, options, sink)`** in `@reportwright/engine`: a report whose body is a flowing table
|
|
65
|
+
(plain or grouped: repeating group headers, subtotals, `RowNumber()`, widow control, keep-together, page breaks) is laid
|
|
66
|
+
out a window of rows at a time and each finished page goes straight to the sink. Peak memory stays flat: about 260 MB
|
|
67
|
+
for the whole render process from 20,000 to 100,000 rows (projected about 250–275 MB at 1,000,000).
|
|
68
|
+
"Page N of M" and footer totals work (a counting pass first). PDF/A-2b, PDF/UA-1 and PDF/A-2a + PDF/UA-1 stream too
|
|
69
|
+
(veraPDF passes). Anything else (images, charts, matrices, rich text, subreports, signing, encryption) falls back to
|
|
70
|
+
`exportPdf`, with the same output.
|
|
71
|
+
- **The server streams large PDFs** (`/api/reports/<id>/pdf`, `/api/v1/render`) with the same row caps, deadline, stall
|
|
72
|
+
timeout and abort handling as the CSV/Excel streaming; the viewer's large-report button gets the streamed file.
|
|
73
|
+
- **`parallel: N`** (Node, opt-in): page ranges painted on worker threads, byte-identical to the serial file. Measure on
|
|
74
|
+
your hardware before relying on it; each worker costs about 150 MB.
|
|
75
|
+
- `examples/stream-1m.mjs` in the engine package: streams a generated million-row ledger to a file and prints peak memory.
|
|
76
|
+
- `planPaged` exported for custom pipelines.
|
|
77
|
+
|
|
78
|
+
### Changed
|
|
79
|
+
- The PDF painter is shared between `exportPdf` and the streaming writer; `exportPdf` output is byte-for-byte unchanged.
|
|
80
|
+
- The engine package ships its readable bundle without a source map (1.3 MB unpacked).
|
|
81
|
+
|
|
82
|
+
### Known limits
|
|
83
|
+
- In the npm engine, grouped reports sort rows in memory (the server spills them to disk), so a 1M-row grouped report is
|
|
84
|
+
not flat outside the server yet.
|
|
85
|
+
|
|
5
86
|
## 0.11.0 — 2026-10-08
|
|
6
87
|
|
|
7
88
|
Everything since 0.10.0. Items marked **Behaviour change** alter output that a pixel-locked layout or a locale-sensitive
|
package/README.md
CHANGED
|
@@ -1,13 +1,34 @@
|
|
|
1
1
|
# @reportwright/engine
|
|
2
2
|
|
|
3
|
-
[](https://www.npmjs.com/package/@reportwright/engine) [](https://github.com/MrArun005/reportwright/blob/main/LICENSE) [](https://socket.dev/npm/package/@reportwright/engine)
|
|
3
|
+
[](https://www.npmjs.com/package/@reportwright/engine) [](https://github.com/MrArun005/reportwright/blob/main/LICENSE) [](https://socket.dev/npm/package/@reportwright/engine)
|
|
4
4
|
|
|
5
|
-
The ReportWright report engine for Node and browsers: a JSON report definition and data in, pages out, then PDF (also PDF/A-2b), Word, Excel, PowerPoint (exportPptx), HTML or CSV.
|
|
5
|
+
The ReportWright report engine for Node and browsers: a JSON report definition and data in, pages out, then PDF (also PDF/A-2b), Word, Excel, PowerPoint (exportPptx), HTML or CSV. The command line is `@reportwright/cli` (`pw`).
|
|
6
6
|
|
|
7
7
|
```bash
|
|
8
8
|
npm i @reportwright/engine @reportwright/fonts
|
|
9
9
|
```
|
|
10
10
|
|
|
11
|
+
Save this as `statement.pw.json` (a minimal report; the repository's `examples/` has more):
|
|
12
|
+
|
|
13
|
+
```json
|
|
14
|
+
{
|
|
15
|
+
"$schema": "pagewright/report@1",
|
|
16
|
+
"name": "Statement",
|
|
17
|
+
"page": { "size": "A4", "margins": [36, 36, 36, 36] },
|
|
18
|
+
"parameters": [{ "name": "accountId", "type": "string" }],
|
|
19
|
+
"dataSources": [{ "name": "ledger", "type": "json", "data": { "rows": [{ "item": "Design", "amount": 1200 }, { "item": "Hosting", "amount": 99.5 }] } }],
|
|
20
|
+
"dataSets": [{ "name": "Lines", "source": "ledger", "path": "$.rows" }],
|
|
21
|
+
"sections": { "body": { "items": [
|
|
22
|
+
{ "type": "textbox", "name": "Title", "x": 0, "y": 0, "w": 400, "h": 24, "value": "Statement", "style": { "fontSize": 18, "fontWeight": "bold" } },
|
|
23
|
+
{ "type": "table", "name": "Table", "x": 0, "y": 40, "w": 400, "h": 60, "dataSet": "Lines",
|
|
24
|
+
"columns": [{ "width": 250 }, { "width": 150 }],
|
|
25
|
+
"header": [{ "height": 18, "cells": [{ "value": "Item" }, { "value": "Amount", "style": { "textAlign": "right" } }] }],
|
|
26
|
+
"detail": [{ "height": 18, "cells": [{ "value": "=Fields.item" }, { "value": "=Fields.amount", "style": { "format": "C2", "textAlign": "right" } }] }],
|
|
27
|
+
"footer": [{ "height": 18, "cells": [{ "value": "Total" }, { "value": "=Sum(Fields.amount)", "style": { "format": "C2", "textAlign": "right", "fontWeight": "bold" } }] }] }
|
|
28
|
+
] } }
|
|
29
|
+
}
|
|
30
|
+
```
|
|
31
|
+
|
|
11
32
|
```js
|
|
12
33
|
import { render, defaultFontStore, exportPdf } from '@reportwright/engine';
|
|
13
34
|
import fs from 'node:fs';
|
|
@@ -46,6 +67,65 @@ const bundledFonts = defaultFontStore({ fontsUrl: '/pw/fonts/' }); // a missing
|
|
|
46
67
|
`timeoutMs` as well; the PDF exporter checks every PNG before decoding it and leaves a damaged one out (`onWarning`).
|
|
47
68
|
- **Currency.** `C2` with no `currency` on the report uses the locale's own (`en-IN` ₹, `de-DE` €, `en-US` $).
|
|
48
69
|
|
|
70
|
+
## Streaming PDF for large reports
|
|
71
|
+
|
|
72
|
+
`exportPdfStream(definition, options, sink)` writes the PDF as it is laid out: a flowing table is laid out a window of
|
|
73
|
+
rows at a time and each page leaves memory once it is written, so a million-row report's PDF needs about the memory of a
|
|
74
|
+
few pages (measured: peak RSS flat from 20,000 to 300,000 rows, about 250 MB for the whole Node process; most of it is
|
|
75
|
+
garbage V8 has not collected yet, not rows: it does not shrink with a smaller `window`, and
|
|
76
|
+
`node --max-old-space-size=128` holds a 300,000-row export to about 180 MB). The file reads as
|
|
77
|
+
`exportPdf`'s: the same pages, text and positions.
|
|
78
|
+
|
|
79
|
+
```js
|
|
80
|
+
import fs from 'node:fs';
|
|
81
|
+
import { exportPdfStream } from '@reportwright/engine';
|
|
82
|
+
|
|
83
|
+
const out = fs.createWriteStream('ledger.pdf');
|
|
84
|
+
const r = await exportPdfStream(definition, { parameters: { year: 2026 }, title: 'Ledger' }, out);
|
|
85
|
+
out.end();
|
|
86
|
+
console.log(r.streamed ? `${r.pages} pages, streamed` : `written whole: ${r.why}`);
|
|
87
|
+
```
|
|
88
|
+
|
|
89
|
+
The sink is a Node `Writable` or anything with `write(bytes)` that may return a promise (back-pressure); it is not
|
|
90
|
+
closed for you. Options are `render()`'s and `exportPdf`'s (`fonts`, `parameters`, `title`, `pdfa`, `tagged`, …), plus
|
|
91
|
+
`window` (rows laid out at once, 1,000), `maxRows` (rows read, 5,000,000), `presorted`, `spill` and `parallel`. Rows come
|
|
92
|
+
from the report's source a batch at a time: CSV and REST (JSON) sources are read as streams, JSON given as text is read
|
|
93
|
+
row by row, an array given in the definition (or in `sources`) is read where it is, never copied or sorted, and
|
|
94
|
+
`sqlRows(source, params)`, an async iterable of row batches, stands in for a database cursor. An array you pass is yours
|
|
95
|
+
and stays in memory for the export: the export's own memory stays flat beside it (measured: the array plus about
|
|
96
|
+
150 MB), but V8 lets garbage grow with a large live heap, so pass a cursor (`sqlRows`) or JSON text when the rows are
|
|
97
|
+
many.
|
|
98
|
+
|
|
99
|
+
- **Sorts and groups.** A table sort or groups put the rows in order first. In Node the sort spills to the system temp
|
|
100
|
+
folder (`spill: { dir, maxMB }`; 2 GB at most), so memory stays flat; `spill: false`, and in a
|
|
101
|
+
browser, it is in memory (all the rows: measured about 1.3 KB a row, so keep browser exports under about 100,000 sorted or
|
|
102
|
+
grouped rows). When the source already gives the rows in the report's order (`order by region, type` for groups by
|
|
103
|
+
region then type), pass `presorted: true`: nothing is sorted or held, each row is checked against the one before as it
|
|
104
|
+
passes, and a row out of order stops the export with an error.
|
|
105
|
+
- **parallel: N** (Node, off by default) lays out and paints with N worker threads; the file is byte for byte the same.
|
|
106
|
+
Each thread loads the engine and fonts and lays out its own windows: about 150–300 MB a thread (measured: 260 MB
|
|
107
|
+
serial, 590 MB with 2, 890 MB with 4, for 100,000 rows). It helps only with idle cores and a report whose layout, not
|
|
108
|
+
its source, is the slow part: 15% faster with 2 threads on a 2-core machine. Leave it off unless you have measured.
|
|
109
|
+
- **Tagged files.** A tagged (PDF/UA) table has a structure element a cell: a million-row ledger has some 8 million,
|
|
110
|
+
about 200 MB of file. Readers open it; tools that load every object (qpdf, pikepdf) may need gigabytes. Past a
|
|
111
|
+
million elements the result warns; `tagged: { lang, maxElements }` caps it (the export stops with an error).
|
|
112
|
+
|
|
113
|
+
**What streams:** a body that is one table (with items above or below it), its rows' expressions row by row (fields,
|
|
114
|
+
parameters, functions of the row, `RowNumber()`), groups with headers that repeat, subtotals (`Sum`, `Count`, `Avg`,
|
|
115
|
+
`Min`, `Max` of a row expression, `CountRows()`), keep-together and page breaks, filters and sorts, page headers and
|
|
116
|
+
footers without aggregates ("Page N of M" included: the pages are counted first, so the source is read twice), PDF/A-2b,
|
|
117
|
+
and accessible PDFs (PDF/UA-1, PDF/A-2a; validated with veraPDF). A data set without declared `fields` streams too: its fields are
|
|
118
|
+
detected from its first 50 rows, as the whole render detects them (declare them when a column first appears later).
|
|
119
|
+
|
|
120
|
+
**What is written whole** (`exportPdf`; the result says why, in `why` and in its first warning, since its memory grows
|
|
121
|
+
with its pages): matrices, charts, lists, subreports, images, rich text,
|
|
122
|
+
lookups and aggregates outside footers, more than one data set, group sections (`newSection`), signed or encrypted PDFs.
|
|
123
|
+
Grouped and sorted tables sort through a disk spill in Node (above), in memory in a browser.
|
|
124
|
+
|
|
125
|
+
`examples/stream-1m.mjs` streams a generated million-row ledger to a file and prints its peak memory:
|
|
126
|
+
`node node_modules/@reportwright/engine/examples/stream-1m.mjs 1000000 ledger.pdf` (`--grouped`, `--presorted`,
|
|
127
|
+
`--tagged`, `--plain`, `--array`, `--array=text`; serial unless `--parallel=N`).
|
|
128
|
+
|
|
49
129
|
## Many reports at once (Node): a worker pool
|
|
50
130
|
|
|
51
131
|
```js
|