@reportwright/engine 0.12.0 → 0.13.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (35) hide show
  1. package/CHANGELOG.md +133 -1
  2. package/README.md +133 -12
  3. package/THIRD-PARTY-NOTICES.md +52 -0
  4. package/dist/index.js +6066 -4454
  5. package/dist/pdfstreamworker.js +2 -2
  6. package/dist/types/packages/engine/entry.d.ts +87 -5
  7. package/dist/types/src/engine/data/guard.d.ts +4 -1
  8. package/dist/types/src/engine/data/index.d.ts +14 -1
  9. package/dist/types/src/engine/expr/evaluate.d.ts +1 -0
  10. package/dist/types/src/engine/index.d.ts +11 -3
  11. package/dist/types/src/engine/items/chart-kit.d.ts +6 -0
  12. package/dist/types/src/engine/lazylibs.d.ts +46 -0
  13. package/dist/types/src/engine/notices.d.ts +1 -0
  14. package/dist/types/src/engine/options.d.ts +10 -0
  15. package/dist/types/src/engine/schema/report.schema.d.ts +3 -0
  16. package/dist/types/src/engine/stream.d.ts +32 -35
  17. package/dist/types/src/engine/text/fonts.d.ts +30 -5
  18. package/dist/types/src/engine/text/measure.d.ts +6 -0
  19. package/dist/types/src/engine/text/standard.d.ts +42 -0
  20. package/dist/types/src/exporters/docx.d.ts +19 -1
  21. package/dist/types/src/exporters/html.d.ts +9 -5
  22. package/dist/types/src/exporters/officestream.d.ts +27 -0
  23. package/dist/types/src/exporters/pdf.d.ts +5 -3
  24. package/dist/types/src/exporters/pdfpaint.d.ts +3 -11
  25. package/dist/types/src/exporters/pdfstandard.d.ts +23 -0
  26. package/dist/types/src/exporters/pdfstream.d.ts +16 -8
  27. package/dist/types/src/exporters/pdfstreamtags.d.ts +8 -1
  28. package/dist/types/src/exporters/spill.d.ts +20 -0
  29. package/dist/types/src/exporters/subset.d.ts +15 -0
  30. package/dist/types/src/exporters/xlsxstream.d.ts +22 -0
  31. package/examples/stream-1m.mjs +21 -8
  32. package/package.json +5 -3
  33. package/pool/worker.js +1 -1
  34. package/report.d.ts +1 -0
  35. package/schema.json +3 -0
package/CHANGELOG.md CHANGED
@@ -2,7 +2,139 @@
2
2
 
3
3
  All `@reportwright/*` packages (formerly `@pagewrightjs/*`) share one version. The versioning policy is in [CONTRIBUTING.md](CONTRIBUTING.md#versioning-semver).
4
4
 
5
- ## 0.12.0 — 2026-10-09
5
+ ## 0.13.0 — 2026-10-09
6
+
7
+ A minor release: new options (`largeReportRows`, `signal`, `warnOnLeave`, `onDirtyChange`), the 14 standard PDF fonts, streamed Excel and Word exports, and warnings where mistakes used to be silent. The tester's 0.12.1 tickets are fixed except where noted; fonts are still TTF (WOFF2 is not done).
8
+
9
+ ### Fixed
10
+ - **Cold Helvetica invoice (A1.1):** a report that names only the standard PDF fonts loads no font file and no fontkit (Inter loads in a second layout pass only when some text needs it, such as a character WinAnsi lacks), and `exportPdf` writes such a model (standard-font text, boxes, lines, links) with the engine's own PDF writer instead of pdf-lib. A new process making a 1-page Helvetica invoice PDF: about 115 ms instead of 224 ms on an M2, with the same text, fonts and page content. Embedded fonts, images, form fields, PDF/A, tagged, encrypted and signed files still go through pdf-lib.
11
+ - **Cold start (A1.1):** pdf-lib and fontkit load on first use (in Node through `require`, elsewhere `import()`), not
12
+ when the engine is imported: `import('@reportwright/engine')` takes about 42 ms instead of 175 ms on an M2, a render
13
+ loads only fontkit, and the streaming PDF writer never loads pdf-lib. A new process making a 1-page invoice PDF:
14
+ about 318 ms instead of 364 ms (most of what is left is compiling harfbuzz-subset.wasm, about 150 ms).
15
+ - **Large classic renders (A1.4, A1.5):** `render()` past 100,000 rows warns once (`Large report (N rows):
16
+ exportPdfStream writes it in about 320 MB; render + exportPdf holds it all in memory`), in the model's warnings and
17
+ once per process on the console. The threshold is the render option `largeReportRows` (100000; `0` or `false`: off). `exportPdf` takes a rendered model, so it cannot switch to streaming itself; the
18
+ engine README shows `exportPdfStream` with a sink that keeps the bytes.
19
+ - **Excel streams; Word is written page by page (A1.9):** `exportXlsxStream` keeps memory flat; `exportDocxStream` lowers it, but a Word export's memory still grows slowly with the rows, so it is not yet a flat stream. `exportDocxStream` and `exportXlsxStream` (engine package and `src/exporters/officestream.js`) write a report's Word or Excel file to a sink as its pages or rows are made, with `exportPdfStream`'s sink, `timeoutMs` and `signal`; a report that cannot stream is written whole (`why`). `exportDocx` writes the document XML into the zip page by page instead of holding it. Peak RSS of a Word export, 20k / 50k rows: 329 / 499 MB before (render + exportDocx), 259 / 286 MB streamed with a flat live heap (17.5 / 17.3 MB after GC); Excel 251 / 432 MB before, 125 / 126 MB streamed (`scripts/mem-office.mjs`).
20
+ - **Viewer first load (A1.6):** the embedded viewer no longer loads the engine (layout, expressions, charts, maps, barcodes) on the main thread up front: the worker renders, and the engine loads on the main thread on the first export, parameter list or main-thread render. First load 827 → 652 KB gzipped (about 262 KB once fontkit loads lazily); the render worker 908 → 579 KB, with bwip-js in `engine.worker.bwip.js` beside it, imported for a report with a barcode (a worker copied alone renders such a report on the main thread). `tests/engine/b10-bundle-size.test.js` holds both to a budget.
21
+ - **`pw render -f pdf` streams (A1.4):** into a file, a flowing report is written by `exportPdfStream` (the message says
22
+ `streamed`; the text is the whole-model PDF's), through a temporary file renamed on success, so a failed run leaves
23
+ no damaged PDF. Encrypted files, reports that cannot stream and `-o -` render the whole model as before. A streamed
24
+ grouped report no longer warns that the group totals' internal fields (`__g0_0`) are missing.
25
+ - **First PDF of a cold process (Node):** harfbuzz-subset.wasm is compiled synchronously, which V8 does lazily; the async compile also optimized the whole module in the background, and the first export waited for it. A cold 1-page invoice in Inter: export 215 ms → 40 ms, whole process 498 ms → 305 ms (median of 10, M2). Browsers keep the async compile. A PDF in the standard fonts never loads the subsetter.
26
+ - **CJK fonts survive upgrades (N10):** `@reportwright/fonts-cjk` downloads into a cache folder outside the package (`REPORTWRIGHT_FONTS_DIR`, else `$XDG_CACHE_HOME/reportwright/fonts-cjk`, else `%LOCALAPPDATA%\reportwright\fonts-cjk` on Windows or `~/.cache/reportwright/fonts-cjk`), so upgrading the package no longer deletes them; `defaultFontStore()` reads that folder, then the package's old `fonts/` folder, so existing installs keep working. New exports: `cacheDir()`, `fontPath(name)`.
27
+ - **The 14 standard PDF fonts (A1.2):** `fontFamily` `Helvetica`, `Times-Roman` and `Courier` (four faces each, PostScript names such as `Helvetica-BoldOblique`), `Symbol` and `ZapfDingbats` are named in the PDF and not embedded (Type1, WinAnsiEncoding, the Adobe AFM widths from `@pdf-lib/standard-fonts`); layout measures with the same widths. Aliases: Arial → Helvetica, Times New Roman → Times-Roman, Courier New → Courier. An invoice in Helvetica is about 1.6 KB. Characters WinAnsi lacks are drawn in Inter with a warning; PDF/A and PDF/UA embed Inter (JetBrains Mono for Courier) at the standard font's glyph positions, with a warning; `exportPdfStream` does the same. SVG/HTML name the font with a CSS fallback list, Word and PowerPoint name Arial, Times New Roman or Courier New. `@pdf-lib/standard-fonts` (already installed with pdf-lib) is now a declared dependency, kept external in the packages.
28
+ - **Unknown fonts (N39):** a `fontFamily` the engine does not have is no longer swapped silently: the render warns once per name (`Font "NoSuchFont" is not available; using Inter.`), and monospace names (Consolas, Menlo, Monaco…) use the shipped JetBrains Mono instead of Inter, so aligned columns stay aligned. Family names match in any case, a CSS list (`"Brand, Inter"`) uses its first available font, and the generic names `sans-serif` and `monospace` map to Inter and JetBrains Mono without a warning.
29
+ - **Streaming tickets, from the 0.12.1 retest (N23-N29):** `exportPdfStream` with `encrypt` or `prepare` and `sqlRows` no
30
+ longer fails with "Save the report first" (the rows are gathered for the whole-model render); a misspelled option
31
+ key (`encrpyt`, `windw`) warns with "Did you mean" in `render()`'s model warnings, `exportPdfStream`'s result and
32
+ `exportPdf`'s `onWarning`, and a misspelling of `encrypt` or `prepare` throws so no file is written without its
33
+ password or signature; `sqlRows` yielding single rows says "sqlRows must yield arrays of rows (got an object with
34
+ keys …)" and sync iterables are accepted; `spill.dir` that does not exist says so; `pw render --base` on a private
35
+ address names `--allow-internal`; README: the temp-file + rename pattern for a failed stream and a `pg` +
36
+ `pg-cursor` `sqlRows` example; the `exportPdf` option docs (pdfa, prepare, save) say what the code does.
37
+ - **Data and form-field retest tickets (B9):**
38
+ - A `number` (or `date`) field with text in it (N26) warns once per field with the first bad row and how many values
39
+ are empty, in a render and in a streamed export. Numeric strings such as "12.34" stay silent.
40
+ - A data set whose path matches an object that holds the rows (`data: { orders: { rows: [...] } }`) warns and names
41
+ the array key (N8). An object of fields, such as `$.meta`, does not.
42
+ - `validate()` warns on an aggregate with no scope outside a table, list or data-set item (it uses the first data
43
+ set), and on `nullable: false` or `allowBlank: false` with no default on a parameter that is not required.
44
+ - A form field has a `label` (falls back to `tooltip`); a tagged PDF uses it as the field's accessible name (`/TU`)
45
+ and warns for each field without one (N30).
46
+ - **Tagged PDF table of contents (N32):** a TOC is tagged `TOC > TOCI > Reference > Link`, with the entry text, its page number and the link annotation inside one Link (it was plain `P` elements with a Link around the number only), and its title is an `H1`. veraPDF PDF/UA-1: compliant. The streaming PDF's tagging is unchanged.
47
+ - **Cascading parameters in the viewer:** when a parent parameter changes, a dependent value its list no longer offers is replaced (by the default if still offered, else the first choice) in the state the run uses, and the report re-runs with it. Before, the panel dropped the value but the run kept the old one ("Delhi, South region" with no rows).
48
+ - **Viewer page thumbnails (A11Y-3):** `.vthumb` no longer has an `aria-label` that differs from its visible number (axe `label-content-name-mismatch`); its name "Page 1" comes from hidden text beside the visible 1.
49
+ - **Viewer worker fallback (N1):** when the render worker cannot load (Vite 7 without `worker: { format: 'es' }`, or esbuild not following `new Worker(new URL(...))`), the viewer logs once `[reportwright] worker did not load (<reason>); rendering on the main thread`, and `handle.usesWorker` says which one you got. The viewer README now starts with the bundler setup. No Vite plugin: the setting is one line.
50
+ - **Worker file (N18):** `engine.worker.js` is one self-contained file (no `chunk-*.js` imports), so copying it alone next to a page works.
51
+ - **Viewer fonts (N31):** the viewer fetches a font face only when the layout first needs it, not all seven up front. A report in Inter Regular and Bold now downloads 670 KB (two files) instead of 4.9 MB. Serve `fontsUrl` with `Cache-Control: public, max-age=31536000, immutable` (in the viewer README). WOFF2 for the browser is not done yet.
52
+ - **Streaming PDF (`exportPdfStream`) failures, from the 0.12.1 retest:**
53
+ - A sink that stops reading no longer hangs the export: `timeoutMs` covers the waits for drain (504, `code:
54
+ 'ETIMEDOUT'`), a sink's `error` or `close` ends the wait, and the new `signal` option (an `AbortSignal`) stops the
55
+ export at its next wait (rows, sink, sort, page) with an `AbortError`. `sqlRows` gets `{ signal }` as its third
56
+ argument. On timeout or abort the rows' cursors are closed and spill files removed.
57
+ - **Polar chart without `xValue`** (the angle) says so in validate() and the render ("a polar chart needs xValue"), instead of
58
+ "N of N rows have no x or y value". A valid polar chart gets no warnings.
59
+ - **Variable font subsets say their weight:** a variable font cut at Regular is named Regular (family, subfamily, full and PostScript names) with usWeightClass 400, and its variation tables are dropped, in the PDF and HTML subsets.
60
+ - **Deprecation messages** all say "will be removed in 1.0" (exportHtml's said 0.9). A test greps the source for a deprecation string that names 0.9.
61
+ - **Rich text warnings** say "Rich text:" once (an unnamed item was named "Rich text" and the parser repeated it). The legend warning says what is accepted: a chart's legend is "bottom" or "none", and it is drawn at the bottom.
62
+
63
+ - **Designer (b9):**
64
+ - Embedded designer: `mountDesigner` asks before the page is left with unsaved changes (`warnOnLeave`, default
65
+ true; false leaves it to the host), and the handle has `isDirty()` and an `onDirtyChange(dirty)` option. The guard is
66
+ removed on save and on `unmount()` (N34). A blank report is named "Untitled report" (it was "New report").
67
+ - Keyboard: the outline rows are a listbox with one Tab stop (Up/Down/Home/End move, Enter or Space selects, Shift
68
+ adds; then Delete and the arrows act on the selection); the left rail's tabs take the focus to their section;
69
+ canvas items are Tab stops named "Text box T1, x 10, y 10" that Enter selects (N35).
70
+ - Paste uses the paste event's clipboard (no clipboard-read permission), else the copy made in this tab, and says
71
+ when there is nothing to paste (N36).
72
+ - Width and height are at least 1 pt (a line may be 0 one way) (N33). Ctrl+D offsets the copy by 12 pt, as a paste
73
+ does (N37).
74
+ - Opening a report gives an item without an `id` a stable one (a hash of its band, place and name), so the same
75
+ file always reads back the same; the other normalisation (names, band heights, a table's size) is unchanged (N38).
76
+ - Embedded designer (shadow root): the table cell menu acts again, and Esc and Ctrl+S work after a click on a band
77
+ label, a matrix or a drill-through in Preview.
78
+ - A table that gets taller (Add group) grows its band; a drop into a 0 pt band grows it. A number field dropped on a
79
+ group header cell gives `=Sum(...)`, as on a footer. A matrix value shown as a percentage gets the format P1 while
80
+ its format is still the default N0. The matrix's drill-down toggle starts closed, as the table's.
81
+ - Accessibility: the fx button is named "fx: Open the expression editor"; table row tags have 4.85:1 contrast.
82
+
83
+ ## 0.12.1 — 2026-10-08
84
+
85
+ ### Fixed
86
+ - **Word export: long reports.** The document XML goes to the zip in pieces instead of one string: a 500k-row report
87
+ hit V8's string limit (RangeError: Invalid string length). A report with no tables exports without `exportData`.
88
+ - **`exportHtml(model, { fonts, ...options })`** takes the options object like the other exporters; the 0.7 form
89
+ `exportHtml(model, fontStore, options)` still works with a deprecation warning. It threw before ("reading 'get'").
90
+ - **Charts: scatter and bubble axes** no longer pad past zero the data does not stay above (0-100 data gives 0 to 120,
91
+ not -50 to 150). Warnings: a row one series leaves blank is drawn by the others, so it is not reported as not drawn;
92
+ a gantt chart has no series by design and gets no "has no series" warning.
93
+ - **Variable CJK fonts** are cut at Regular (wght 400), not their default Thin instance, in the PDF and HTML subsets.
94
+ - **Data fetch:** gzip, deflate and br bodies are decoded (the byte cap counts decoded bytes); `file:` and `data:` URLs
95
+ are refused with the scheme named, not "its host is made by a parameter".
96
+ - **Format characters** (RLM, LRE/LRI, ZWJ/ZWNJ, soft hyphen) need no glyph: no "No font has these characters" warning,
97
+ and no notdef box in the PDF.
98
+ - **Warnings:** an unnamed item is named by its type (no "undefined:"), and a sub-report warning that validate and the
99
+ render both say is said once.
100
+ - **`parameterOptions()`** lists a parameter's fixed options first, then the fetched ones, each value once (as the viewer does).
101
+ - **HTML Tables: date sort keys** use the wall clock, so the same report gives the same keys in every host time zone.
102
+
103
+ - **Streaming PDF (`exportPdfStream`), from an independent 1M-row verification:**
104
+ - Rows given as an array are read where they are (never copied whole or sorted, tested); JSON given as text is read
105
+ row by row instead of parsed whole (a 300,000-row text source: 910 → 500 MB peak). The README says why a large array
106
+ still shows a large RSS (it is the caller's, and V8 lets garbage grow beside it) and to pass a cursor or text.
107
+ - Grouped and sorted tables: in Node the engine package sorts through a disk spill by default (`spill: { dir, maxMB }`,
108
+ `spill: false`), so a grouped export's memory is flat (300,000 rows: 300 MB, was 670 MB and growing); `presorted:
109
+ true` reads rows that come in the report's order without sorting, checking each against the one before (a row out of
110
+ order stops the export with an error).
111
+ - A data set without declared `fields` streams: its fields are detected from the first 50 rows, as the whole render
112
+ detects them. A report written whole now says why in its first warning, not only in `why`.
113
+ - The window that reaches the last rows is the last one: a table whose footer keeps with its rows could stop with
114
+ "The data changed while the PDF was written".
115
+ - Tagged PDFs are about 27% smaller (a page's borders are one artifact and its cells' text one text object; plain cells
116
+ carry no attributes; the xref stream is predicted). Past a million structure elements the result warns that tools
117
+ which load every object (qpdf, pikepdf) may run out of memory; `tagged.maxElements` caps it.
118
+ - `parallel: N` is documented with its memory (150–300 MB a thread) and when it helps; the example runs serial.
119
+ - The comments of `exportPdfStream` said tagged and PDF/A files do not stream; they do.
120
+ - Note: the non-streaming `exportPdf` still uses pdf-lib, which stays a dependency.
121
+ - **BIRT import: SUM and other aggregates total their column.** A bound column with `aggregateFunction` (SUM, AVE,
122
+ MIN, MAX, FIRST, LAST…) had no listed argument, so it fell back to a row count: `Sum(1, …)` showed 3 for a total of
123
+ 250.75. The column's own expression is now the argument. An aggregate with no expression is counted as approximated.
124
+ The Jasper and SSRS importers were checked: they already take the expression.
125
+ - **Docs: no npm provenance badge.** The repository is private, so npm provenance cannot be generated. The badges and
126
+ the claims are removed from the package READMEs and `docs/RELEASING.md`; the publish workflow stays, and its notes say
127
+ provenance needs a public repository.
128
+ - **Docs: the viewer's `data` example** says which shape matches the data set's path: `{ orders: { rows: [...] } }` for
129
+ `"path": "$.rows"`, `{ orders: [...] }` for a data set with no path. A mismatch gives empty tables, silently.
130
+ - **Docs: a missing font** stops the viewer's run with a message naming the font and a Retry button; the README said the
131
+ report draws in a fallback font, which it does not.
132
+ - **Docs: the engine README** says `pw` is in `@reportwright/cli`, and its quick start has a `statement.pw.json` to save.
133
+ - **CLI help** marks `pw audit` and `pw migrate-storage` as server only; the npm CLI does not include them.
134
+ - **Engine text:** the CJK font warning names `npx reportwright-fonts-cjk` first (the app's `npm run fonts:cjk` second),
135
+ and the deprecated positional export forms say they are removed in 1.0.
136
+
137
+ ## 0.12.0 — 2026-10-08
6
138
 
7
139
  Streaming PDF: very large reports are written page by page, with flat memory.
8
140
 
package/README.md CHANGED
@@ -1,13 +1,34 @@
1
1
  # @reportwright/engine
2
2
 
3
- [![npm](https://img.shields.io/npm/v/%40reportwright%2Fengine)](https://www.npmjs.com/package/@reportwright/engine) [![license](https://img.shields.io/badge/license-MIT-blue)](https://github.com/MrArun005/reportwright/blob/main/LICENSE) [![Socket](https://socket.dev/api/badge/npm/package/@reportwright/engine)](https://socket.dev/npm/package/@reportwright/engine) [![provenance](https://img.shields.io/badge/npm-provenance-green)](https://docs.npmjs.com/generating-provenance-statements)
3
+ [![npm](https://img.shields.io/npm/v/%40reportwright%2Fengine)](https://www.npmjs.com/package/@reportwright/engine) [![license](https://img.shields.io/badge/license-MIT-blue)](https://github.com/MrArun005/reportwright/blob/main/LICENSE) [![Socket](https://socket.dev/api/badge/npm/package/@reportwright/engine)](https://socket.dev/npm/package/@reportwright/engine)
4
4
 
5
- The ReportWright report engine for Node and browsers: a JSON report definition and data in, pages out, then PDF (also PDF/A-2b), Word, Excel, PowerPoint (exportPptx), HTML or CSV. Includes the `pw` command line.
5
+ The ReportWright report engine for Node and browsers: a JSON report definition and data in, pages out, then PDF (also PDF/A-2b), Word, Excel, PowerPoint (exportPptx), HTML or CSV. The command line is `@reportwright/cli` (`pw`).
6
6
 
7
7
  ```bash
8
8
  npm i @reportwright/engine @reportwright/fonts
9
9
  ```
10
10
 
11
+ Save this as `statement.pw.json` (a minimal report; the repository's `examples/` has more):
12
+
13
+ ```json
14
+ {
15
+ "$schema": "pagewright/report@1",
16
+ "name": "Statement",
17
+ "page": { "size": "A4", "margins": [36, 36, 36, 36] },
18
+ "parameters": [{ "name": "accountId", "type": "string" }],
19
+ "dataSources": [{ "name": "ledger", "type": "json", "data": { "rows": [{ "item": "Design", "amount": 1200 }, { "item": "Hosting", "amount": 99.5 }] } }],
20
+ "dataSets": [{ "name": "Lines", "source": "ledger", "path": "$.rows" }],
21
+ "sections": { "body": { "items": [
22
+ { "type": "textbox", "name": "Title", "x": 0, "y": 0, "w": 400, "h": 24, "value": "Statement", "style": { "fontSize": 18, "fontWeight": "bold" } },
23
+ { "type": "table", "name": "Table", "x": 0, "y": 40, "w": 400, "h": 60, "dataSet": "Lines",
24
+ "columns": [{ "width": 250 }, { "width": 150 }],
25
+ "header": [{ "height": 18, "cells": [{ "value": "Item" }, { "value": "Amount", "style": { "textAlign": "right" } }] }],
26
+ "detail": [{ "height": 18, "cells": [{ "value": "=Fields.item" }, { "value": "=Fields.amount", "style": { "format": "C2", "textAlign": "right" } }] }],
27
+ "footer": [{ "height": 18, "cells": [{ "value": "Total" }, { "value": "=Sum(Fields.amount)", "style": { "format": "C2", "textAlign": "right", "fontWeight": "bold" } }] }] }
28
+ ] } }
29
+ }
30
+ ```
31
+
11
32
  ```js
12
33
  import { render, defaultFontStore, exportPdf } from '@reportwright/engine';
13
34
  import fs from 'node:fs';
@@ -50,7 +71,9 @@ const bundledFonts = defaultFontStore({ fontsUrl: '/pw/fonts/' }); // a missing
50
71
 
51
72
  `exportPdfStream(definition, options, sink)` writes the PDF as it is laid out: a flowing table is laid out a window of
52
73
  rows at a time and each page leaves memory once it is written, so a million-row report's PDF needs about the memory of a
53
- few pages (measured: peak RSS flat from 20,000 to 100,000 rows, about 270 MB for the whole Node process). The file reads as
74
+ few pages (measured: peak RSS flat from 20,000 to 300,000 rows, about 250 MB for the whole Node process; most of it is
75
+ garbage V8 has not collected yet, not rows: it does not shrink with a smaller `window`, and
76
+ `node --max-old-space-size=128` holds a 300,000-row export to about 180 MB). The file reads as
54
77
  `exportPdf`'s: the same pages, text and positions.
55
78
 
56
79
  ```js
@@ -63,26 +86,124 @@ out.end();
63
86
  console.log(r.streamed ? `${r.pages} pages, streamed` : `written whole: ${r.why}`);
64
87
  ```
65
88
 
89
+ **Bytes in memory instead of a file.** `exportPdf(model)` takes a page model that `render()` has already built, so it
90
+ cannot stream: render + exportPdf holds every page (measured on a flowing 4-column table: 435 MB peak RSS at 100,000
91
+ rows, against 273 MB streamed with the rows themselves in the definition). For a large report that you want as bytes, give
92
+ `exportPdfStream` a sink that keeps the chunks; it renders the report whole itself when it cannot stream (`r.why`).
93
+ `render()` warns once per render when a data set has more rows than its `largeReportRows` option (default 100,000; `0`
94
+ or `false` turns it off): `Large report (N rows): …` in `model.warnings`, and on the console once per process.
95
+
96
+ ```js
97
+ const parts = [];
98
+ const r = await exportPdfStream(definition, { title: 'Ledger' }, { write: (b) => { parts.push(b); } });
99
+ const bytes = Buffer.concat(parts); // the same pages as render + exportPdf
100
+ ```
101
+
102
+ **A failed export leaves a damaged file.** The PDF is written as it goes, so an error, a timeout or an abort halfway
103
+ leaves a partial file at the path. Write to a temporary name and rename it when the export has succeeded:
104
+
105
+ ```js
106
+ import fs from 'node:fs';
107
+ import { finished } from 'node:stream/promises';
108
+
109
+ const tmp = `${file}.part`;
110
+ const out = fs.createWriteStream(tmp);
111
+ try {
112
+ await exportPdfStream(def, opts, out);
113
+ await finished(out.end());
114
+ await fs.promises.rename(tmp, file);
115
+ } catch (e) {
116
+ out.destroy();
117
+ await fs.promises.rm(tmp, { force: true });
118
+ throw e;
119
+ }
120
+ ```
121
+
122
+ **Rows from a database (`sqlRows`).** `sqlRows(source, params, { signal })` returns an async (or sync) iterable of
123
+ arrays of rows, a batch at a time; a single row is an error. This one reads a PostgreSQL cursor with `pg` and
124
+ `pg-cursor`, releases the client when the export ends or stops, and checks `signal` between batches:
125
+
126
+ ```js
127
+ import pg from 'pg';
128
+ import Cursor from 'pg-cursor';
129
+
130
+ const pool = new pg.Pool();
131
+ const sqlRows = async function* (source, params, { signal }) {
132
+ const client = await pool.connect();
133
+ const cursor = client.query(new Cursor(source.query, Object.values(params)));
134
+ try {
135
+ for (;;) {
136
+ signal?.throwIfAborted(); // stop between batches when the export is cancelled or times out
137
+ const rows = await cursor.read(1000); // one batch of up to 1,000 rows
138
+ if (!rows.length) break;
139
+ yield rows;
140
+ }
141
+ } finally {
142
+ await cursor.close().catch(() => {});
143
+ client.release();
144
+ }
145
+ };
146
+ await exportPdfStream(def, { sqlRows, timeoutMs: 120_000 }, out);
147
+ ```
148
+
149
+ **Word and Excel, streamed.** `exportDocxStream(definition, options, sink)` and `exportXlsxStream(definition, options,
150
+ sink)` take the same sink, `timeoutMs`, `signal`, `sqlRows`, `maxRows` and `spill`: Word lays out a window of rows at a
151
+ time and writes each page's XML into the zip as it comes, Excel writes the sheet row by row, so neither holds the rows,
152
+ the pages or the XML (measured: live heap 17 MB after 20,000 and after 50,000 rows; a 50,000-row Word export runs in
153
+ `node --max-old-space-size=48`). A report that cannot stream is rendered whole and written in one piece (`streamed:
154
+ false`, and `why`). `exportDocx` and `exportXlsx` still take a rendered model.
155
+
156
+ Options that are not options are reported: a misspelled key (`{ windw: 500 }`) is ignored with a warning in the
157
+ result's `warnings` (`render()`: in the model's; `exportPdf`: through `onWarning`) naming the nearest option, and a
158
+ misspelling of `encrypt` or `prepare` throws, so a file is never written without the password or signature you asked
159
+ for. `encrypt`, `prepare`, `save` and PDF/A without a profile need the whole model, `sqlRows` included: the rows are
160
+ gathered, and the result says `streamed: false`.
161
+
66
162
  The sink is a Node `Writable` or anything with `write(bytes)` that may return a promise (back-pressure); it is not
67
163
  closed for you. Options are `render()`'s and `exportPdf`'s (`fonts`, `parameters`, `title`, `pdfa`, `tagged`, …), plus
68
- `window` (rows laid out at once, 1,000), `maxRows` (rows read, 5,000,000) and `parallel` (Node: that many worker threads
69
- lay out and paint the pages; the file is byte for byte the same; each thread loads the engine and fonts, so it costs
70
- memory: about 150 MB a thread). Rows come from the report's source a batch at a time: CSV and REST (JSON) sources are read
71
- as streams, and `sqlRows(source, params)`, an async iterable of row batches, stands in for a database cursor.
164
+ `window` (rows laid out at once, 1,000), `maxRows` (rows read, 5,000,000), `presorted`, `spill` and `parallel`. Rows come
165
+ from the report's source a batch at a time: CSV and REST (JSON) sources are read as streams, JSON given as text is read
166
+ row by row, an array given in the definition (or in `sources`) is read where it is, never copied or sorted, and
167
+ `sqlRows(source, params)`, an async iterable of row batches, stands in for a database cursor. An array you pass is yours
168
+ and stays in memory for the export: the export's own memory stays flat beside it (measured: the array plus about
169
+ 150 MB), but V8 lets garbage grow with a large live heap, so pass a cursor (`sqlRows`) or JSON text when the rows are
170
+ many.
171
+
172
+ - **Sorts and groups.** A table sort or groups put the rows in order first. In Node the sort spills to the system temp
173
+ folder (`spill: { dir, maxMB }`; 2 GB at most), so memory stays flat; `spill: false`, and in a
174
+ browser, it is in memory (all the rows: measured about 1.3 KB a row, so keep browser exports under about 100,000 sorted or
175
+ grouped rows). When the source already gives the rows in the report's order (`order by region, type` for groups by
176
+ region then type), pass `presorted: true`: nothing is sorted or held, each row is checked against the one before as it
177
+ passes, and a row out of order stops the export with an error.
178
+ - **parallel: N** (Node, off by default) lays out and paints with N worker threads; the file is byte for byte the same.
179
+ Each thread loads the engine and fonts and lays out its own windows: about 150–300 MB a thread (measured: 260 MB
180
+ serial, 590 MB with 2, 890 MB with 4, for 100,000 rows). It helps only with idle cores and a report whose layout, not
181
+ its source, is the slow part: 15% faster with 2 threads on a 2-core machine. Leave it off unless you have measured.
182
+ - **Time limit and cancelling.** `timeoutMs` bounds the whole export, including waits for a sink that stops reading
183
+ (an HTTP client gone quiet, a `Writable` that never drains): past it the export rejects with status 504 and
184
+ `code: 'ETIMEDOUT'`. `signal` (an `AbortSignal`) stops it at its next wait (a batch of rows, the sink, a sort, a
185
+ page) and rejects with the signal's reason (an `AbortError`). `sqlRows(source, params, { signal })` gets the signal
186
+ too, so a database cursor can cancel its query. Either way the rows' cursors are closed and a sort's temporary
187
+ files removed.
188
+ - **Tagged files.** A tagged (PDF/UA) table has a structure element a cell: a million-row ledger has some 8 million,
189
+ about 200 MB of file. Readers open it; tools that load every object (qpdf, pikepdf) may need gigabytes. Past a
190
+ million elements the result warns; `tagged: { lang, maxElements }` caps it (the export stops with an error).
72
191
 
73
192
  **What streams:** a body that is one table (with items above or below it), its rows' expressions row by row (fields,
74
193
  parameters, functions of the row, `RowNumber()`), groups with headers that repeat, subtotals (`Sum`, `Count`, `Avg`,
75
194
  `Min`, `Max` of a row expression, `CountRows()`), keep-together and page breaks, filters and sorts, page headers and
76
195
  footers without aggregates ("Page N of M" included: the pages are counted first, so the source is read twice), PDF/A-2b,
77
- and accessible PDFs (PDF/UA-1, PDF/A-2a; validated with veraPDF).
196
+ and accessible PDFs (PDF/UA-1, PDF/A-2a; validated with veraPDF). A data set without declared `fields` streams too: its fields are
197
+ detected from its first 50 rows, as the whole render detects them (declare them when a column first appears later).
78
198
 
79
- **What is written whole** (`exportPdf`, the result says why): matrices, charts, lists, subreports, images, rich text,
199
+ **What is written whole** (`exportPdf`; the result says why, in `why` and in its first warning, since its memory grows
200
+ with its pages): matrices, charts, lists, subreports, images, rich text,
80
201
  lookups and aggregates outside footers, more than one data set, group sections (`newSection`), signed or encrypted PDFs.
81
- Grouped and sorted tables hold their rows in the sort unless `sorter` spills them to disk (the ReportWright server does).
202
+ Grouped and sorted tables sort through a disk spill in Node (above), in memory in a browser.
82
203
 
83
204
  `examples/stream-1m.mjs` streams a generated million-row ledger to a file and prints its peak memory:
84
- `node node_modules/@reportwright/engine/examples/stream-1m.mjs 1000000 ledger.pdf` (`--grouped`, `--tagged`,
85
- `--parallel=4`, `--plain`).
205
+ `node node_modules/@reportwright/engine/examples/stream-1m.mjs 1000000 ledger.pdf` (`--grouped`, `--presorted`,
206
+ `--tagged`, `--plain`, `--array`, `--array=text`; serial unless `--parallel=N`).
86
207
 
87
208
  ## Many reports at once (Node): a worker pool
88
209
 
@@ -0,0 +1,52 @@
1
+ # Third-party notices
2
+
3
+ ReportWright is MIT licensed (see LICENSE). It ships or bundles the following third-party work, under their own licenses.
4
+
5
+ ## Bundled code
6
+
7
+ | Component | Version | License | Used for |
8
+ |---|---|---|---|
9
+ | [HarfBuzz](https://github.com/harfbuzz/harfbuzz), compiled to WebAssembly by [harfbuzzjs](https://github.com/harfbuzz/harfbuzzjs) | harfbuzzjs 1.6.2 | HarfBuzz: "Old MIT" (below); harfbuzzjs: MIT, Copyright (c) 2019-2026 The harfbuzzjs project authors | text shaping (`harfbuzz.wasm`) and font subsetting (`harfbuzz-subset.wasm`) |
10
+ | [pdf-lib](https://github.com/Hopding/pdf-lib) | 1.17.1 | MIT, Copyright (c) 2019 Andrew Dillon | the classic PDF writer (encryption, `prepare`, `save`) |
11
+ | [@pdf-lib/fontkit](https://github.com/Hopding/fontkit) | 1.1.1 | MIT | reading font files |
12
+ | [@pdf-lib/standard-fonts](https://github.com/Hopding/standard-fonts) | 1.0.0 | MIT; the metrics are Adobe's AFM files for the 14 standard PDF fonts | Helvetica, Times and Courier widths |
13
+ | [bwip-js](https://github.com/metafloor/bwip-js) | 4.11.4 | MIT | barcodes |
14
+ | [fflate](https://github.com/101arrowz/fflate) | 0.8.3 | MIT | zip and deflate (Excel, Word, PowerPoint, PDF streams) |
15
+ | [bidi-js](https://github.com/lojjic/bidi-js) | 1.0.3 | MIT | right-to-left text |
16
+
17
+ The full license text of each npm dependency is in its package under `node_modules`.
18
+
19
+ ### HarfBuzz ("Old MIT" license)
20
+
21
+ HarfBuzz is licensed under the so-called "Old MIT" license. Its copyright holders are listed in
22
+ [COPYING](https://github.com/harfbuzz/harfbuzz/blob/main/COPYING) in the HarfBuzz repository.
23
+
24
+ > Permission is hereby granted, without written agreement and without license or royalty fees, to use, copy, modify,
25
+ > and distribute this software and its documentation for any purpose, provided that the above copyright notice and
26
+ > the following two paragraphs appear in all copies of this software.
27
+ >
28
+ > IN NO EVENT SHALL THE COPYRIGHT HOLDER BE LIABLE TO ANY PARTY FOR DIRECT, INDIRECT, SPECIAL, INCIDENTAL, OR
29
+ > CONSEQUENTIAL DAMAGES ARISING OUT OF THE USE OF THIS SOFTWARE AND ITS DOCUMENTATION, EVEN IF THE COPYRIGHT HOLDER HAS
30
+ > BEEN ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
31
+ >
32
+ > THE COPYRIGHT HOLDER SPECIFICALLY DISCLAIMS ANY WARRANTIES, INCLUDING, BUT NOT LIMITED TO, THE IMPLIED WARRANTIES OF
33
+ > MERCHANTABILITY AND FITNESS FOR A PARTICULAR PURPOSE. THE SOFTWARE PROVIDED HEREUNDER IS ON AN "AS IS" BASIS, AND THE
34
+ > COPYRIGHT HOLDER HAS NO OBLIGATION TO PROVIDE MAINTENANCE, SUPPORT, UPDATES, ENHANCEMENTS, OR MODIFICATIONS.
35
+
36
+ ## Fonts (in @reportwright/fonts and @reportwright/fonts-cjk)
37
+
38
+ All fonts are under the [SIL Open Font License 1.1](https://openfontlicense.org); each family's license file ships next
39
+ to it in `public/fonts` (OFL-Inter.txt, OFL-JetBrainsMono.txt, OFL-NotoSans*.txt).
40
+
41
+ | Family | Source |
42
+ |---|---|
43
+ | Inter | [rsms/inter](https://github.com/rsms/inter) |
44
+ | JetBrains Mono | [JetBrains/JetBrainsMono](https://github.com/JetBrains/JetBrainsMono) |
45
+ | Noto Sans Arabic, Bengali, Devanagari, Gujarati, Hebrew, Kannada, Khmer, Lao, Myanmar, Tamil, Telugu, Thai | [notofonts](https://github.com/notofonts) |
46
+ | Noto Sans CJK and Noto Emoji (fonts-cjk, downloaded on demand) | [notofonts](https://github.com/notofonts) |
47
+
48
+ ## Colour profile
49
+
50
+ `sRGB-v2-magic.icc` (the PDF/A output intent) is from
51
+ [Compact-ICC-Profiles](https://github.com/saucecontrol/Compact-ICC-Profiles) by Clinton Ingram, released under
52
+ CC0 1.0 (public domain).