@reportwright/pdf 0.1.0-beta.1 → 0.1.0-beta.10

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -1,32 +1,369 @@
1
1
  # @reportwright/pdf
2
2
 
3
- A streaming PDF writer for Node and browsers, with no runtime dependencies. Pages are written to your sink as
4
- they end, so memory stays flat however many pages you write. It is the writer inside ReportWright, extracted.
3
+ **Create, comply, sign, light edit: a streaming PDF library for Node and browsers with zero runtime dependencies.**
4
+
5
+ [![npm](https://img.shields.io/npm/v/@reportwright/pdf?color=1f4e79)](https://www.npmjs.com/package/@reportwright/pdf)
6
+ [![license: MIT](https://img.shields.io/badge/license-MIT-blue)](./LICENSE)
7
+ [![dependencies: 0](https://img.shields.io/badge/runtime%20dependencies-0-brightgreen)](./package.json)
8
+ [![node >= 20.16](https://img.shields.io/badge/node-%3E%3D20.16-339933)](./package.json)
9
+ [![PDF/A + PDF/UA: veraPDF](https://img.shields.io/badge/PDF%2FA%201b%E2%80%933a%20%2B%20UA--1-veraPDF%20clean-6f42c1)](./conformance)
10
+
11
+ **Live demo:** [mrarun005.github.io/reportwright-demo](https://mrarun005.github.io/reportwright-demo) ·
12
+ [PDF benchmarks](https://mrarun005.github.io/reportwright-demo/bench.html) ·
13
+ [playground](https://mrarun005.github.io/reportwright-demo/play.html)
14
+
15
+ Pages are written to your sink as they end, so memory stays flat however many pages you write. It is the writer inside
16
+ ReportWright, extracted, plus a hardened reader for merging, splitting, filling and signing existing files.
17
+
18
+ What it is for: **create, comply, sign, light edit**. It does not render PDFs: for rendering use
19
+ [pdf.js](https://mozilla.github.io/pdf.js/) or [PDFium](https://pdfium.googlesource.com/pdfium/). For heavy
20
+ rewriting, repair or raw speed, MuPDF and qpdf are the established tools; this library does not claim to be faster
21
+ than them.
22
+
23
+ - **Write**: text in the 14 standard fonts or embedded TrueType/OpenType (subset, kerned, any Unicode), complex
24
+ scripts through a HarfBuzz shaper, JPEG/PNG images, Canvas-style vector paths, gradients, patterns, layers.
25
+ - **Navigate**: links, named destinations, bookmarks, page labels, annotations.
26
+ - **Forms**: every AcroForm field type with appearance streams; fill and flatten existing forms.
27
+ - **Protect**: AES-256 / AES-128 encryption with permissions; PAdES B-B / B-T signatures while streaming.
28
+ - **Comply**: PDF/A-1b, 2b/2u/2a, 3b/3u/3a and PDF/UA-1 (veraPDF clean), Factur-X / ZUGFeRD / XRechnung.
29
+ - **Read and edit**: load (also damaged or encrypted files), merge, split, reorder, extract text, fill, stamp,
30
+ `saveIncremental` that keeps existing signatures valid.
31
+ - **Safe by default**: no JavaScript or launch actions can be written; active content is stripped on read; every
32
+ input is budgeted.
33
+
34
+ ## Contents
35
+
36
+ [Install](#install) · [Quick start](#quick-start-30-seconds) · [Guide](#guide) · [Examples](#examples) ·
37
+ [Comparison](#comparison) · [Options reference](#options-reference) · [Security model](#security-model) ·
38
+ [Known limits](#known-limits) · [Changes](#changes)
39
+
40
+ ## Install
41
+
42
+ ```sh
43
+ npm i @reportwright/pdf
44
+ ```
45
+
46
+ ESM only, Node 20.16 or later, and modern browsers (the same `dist/index.js`; crypto comes from WebCrypto). Types are
47
+ included (`index.d.ts`). Optional, for complex scripts only: `npm i harfbuzzjs`.
5
48
 
6
- ## Quickstart
49
+ ## Quick start (30 seconds)
7
50
 
8
51
  ```js
52
+ // hello.mjs → node hello.mjs
9
53
  import fs from 'node:fs';
10
54
  import { createPdf } from '@reportwright/pdf';
11
55
 
12
- const pdf = createPdf(fs.createWriteStream('hello.pdf'), { title: 'Hello', creationDate: new Date('2026-01-01') });
13
- const bold = pdf.standardFont('Helvetica-Bold');
14
- const page = pdf.addPage({ width: 595.28, height: 841.89 }); // A4, in points
15
- page.text('Hello, world', { x: 72, y: 72, font: bold, size: 24, color: '#1f4e79' });
56
+ const pdf = createPdf(fs.createWriteStream('hello.pdf'), { title: 'Hello' });
57
+ const page = pdf.addPage(); // A4, points, origin top-left
58
+ page.text('Hello, world', { x: 72, y: 72, font: pdf.standardFont('Helvetica-Bold'), size: 24, color: '#1f4e79' });
16
59
  page.line(72, 104, 523, 104, { color: 0.6, width: 0.5 });
60
+ await page.end(); // the page is written and freed
61
+ console.log(await pdf.end()); // { pages: 1, bytes: …, warnings: [] }
62
+ ```
63
+
64
+ In memory instead of a stream: `const bytes = await toBytes(async (pdf) => { … }, options)`.
65
+
66
+ ## Guide
67
+
68
+ Every snippet below is cut from a file in [`examples/`](./examples), each of which was run and produces a PDF. The
69
+ examples import `../dist/index.js`; in your code import `@reportwright/pdf`. They also ship in the package
70
+ (`node_modules/@reportwright/pdf/examples/`).
71
+
72
+ ### Write: text, fonts, images, tables
73
+
74
+ ```js
75
+ const pdf = createPdf(fs.createWriteStream('report.pdf'), { title: 'Quarterly report', author: 'Example Ltd' });
76
+ const regular = pdf.standardFont('Helvetica'), bold = pdf.standardFont('Helvetica-Bold');
77
+ const inter = await pdf.embedFont(fs.readFileSync('Inter-Regular.ttf')); // subset, any Unicode
78
+ const logo = await pdf.embedImage(fs.readFileSync('logo.png')); // alpha becomes a soft mask
79
+
80
+ const page = pdf.addPage();
81
+ page.image(logo, { x: 40, y: 40, width: 32 });
82
+ page.text('Quarterly report', { x: 84, y: 44, font: bold, size: 22, color: '#1f4e79' });
83
+ const box = page.textBox(longText, { x: 40, y: 90, width: 515, font: inter, size: 11, align: 'justify' });
84
+ // box = { height, lines, overflow }: continue `overflow` in the next box or page
85
+ ```
86
+
87
+ **Tables are paths and text**: there is no table primitive, so you control every rule and alignment.
88
+
89
+ ```js
90
+ const rows = [['Region', 'Q1', 'Q2'], ['North', '1,204', '1,390'], ['South', '980', '1,122']];
91
+ const cols = [40, 300, 430], right = 555, rowH = 22, top = 100 + box.height;
92
+ page.fill(page.path().rect(40, top, right - 40, rowH), { fill: '#1f4e79' }); // header band
93
+ rows.forEach((r, i) => {
94
+ const y = top + i * rowH;
95
+ r.forEach((cell, c) => {
96
+ const font = i === 0 ? bold : regular;
97
+ const x = c === 0 ? cols[c] + 6 : (cols[c + 1] ?? right) - 6 - font.widthOfText(cell, 11); // right-align numbers
98
+ page.text(cell, { x, y: y + 5, font, size: 11, color: i === 0 ? '#ffffff' : 0 });
99
+ });
100
+ if (i > 0) page.stroke(page.path().moveTo(40, y + rowH).lineTo(right, y + rowH), { stroke: 0.75, width: 0.5 });
101
+ });
102
+ await page.end();
103
+ await pdf.end();
104
+ ```
105
+
106
+ Full file: [`examples/write.mjs`](./examples/write.mjs) (it also generates its PNG, so it needs no input files).
107
+
108
+ ### Indic and other complex scripts (HarfBuzz)
109
+
110
+ Pass a shaper to `embedFont` and conjuncts, reordered matras, Arabic joining and marks are drawn correctly; copied
111
+ and extracted text stays the source text, in logical order. The adapter for harfbuzzjs ships as
112
+ [`examples/harfbuzz-shaper.mjs`](./examples/harfbuzz-shaper.mjs).
113
+
114
+ ```js
115
+ import { harfbuzzShaper } from '@reportwright/pdf/examples/harfbuzz-shaper.mjs'; // needs: npm i harfbuzzjs
116
+
117
+ const bytes = fs.readFileSync('NotoSansDevanagari-Regular.ttf');
118
+ const hindi = await pdf.embedFont(bytes, { shaper: harfbuzzShaper(bytes) });
119
+ page.text('किताब हिन्दी क्षत्रिय', { x: 40, y: 40, font: hindi, size: 24 });
120
+ page.textBox('इस किताब में हिन्दी और English दोनों हैं।', { x: 40, y: 90, width: 400, font: hindi, fallback: [latin] });
121
+ ```
122
+
123
+ `node examples/indic.mjs NotoSansDevanagari-Regular.ttf`. Right-to-left: `{ direction: 'rtl' }` (there is no bidi
124
+ algorithm: split mixed-direction lines into runs yourself).
125
+
126
+ ### Links and bookmarks
127
+
128
+ ```js
129
+ page.link(40, 300, 160, 14, { url: 'https://example.com/report' }); // http, https, mailto by default
130
+ pdf.destination('appendix', 2, { top: 40 }); // a named destination
131
+ page.link(40, 320, 160, 14, { dest: 'appendix' }, { alt: 'Go to the appendix' });
132
+ page.link(40, 340, 160, 14, { page: 2, fit: 'FitH', top: 0 });
133
+ pdf.outline([
134
+ { title: 'Quarterly report', page: 1 },
135
+ { title: 'Appendix', page: 2, level: 1 }, // level nests under the previous entry
136
+ ]);
137
+ pdf.pageLabels([{ start: 0, style: 'r' }, { start: 2, style: 'D' }]); // i, ii, 1, 2, …
138
+ ```
139
+
140
+ ### Forms: create, fill, flatten
141
+
142
+ ```js
143
+ import { toBytes, loadPdf } from '@reportwright/pdf';
144
+
145
+ const form = await toBytes(async (pdf) => {
146
+ const font = pdf.standardFont('Helvetica'), page = pdf.addPage();
147
+ pdf.form.textField('name', { page, x: 40, y: 80, width: 240, height: 22, font, tooltip: 'Full name', required: true });
148
+ pdf.form.comboBox('type', { page, x: 40, y: 115, width: 160, height: 22, font, options: ['Annual', 'Sick', 'Unpaid'], value: 'Annual', tooltip: 'Leave type' });
149
+ pdf.form.checkbox('approved', { page, x: 40, y: 150, width: 14, height: 14, tooltip: 'Approved by manager' });
150
+ await page.end();
151
+ });
152
+
153
+ const doc = await loadPdf(form);
154
+ doc.fields; // [{ name: 'name', type: 'text', … }, …]
155
+ doc.fill({ name: 'Aanya Sharma', type: 'Sick', approved: true }); // appearances redrawn by the writer
156
+ const filled = (await doc.save()).bytes;
157
+
158
+ const flattened = (await (await loadPdf(form)).fill({ name: 'Aanya Sharma' }).flatten().save()).bytes; // no AcroForm left
159
+ ```
160
+
161
+ [`examples/fill-form.mjs`](./examples/fill-form.mjs) writes `form.pdf`, `filled.pdf` and `flattened.pdf`. To
162
+ flatten while writing: `createPdf(sink, { form: { flatten: true } })`.
163
+
164
+ ### Encryption
165
+
166
+ ```js
167
+ const bytes = await toBytes(build, {
168
+ encrypt: {
169
+ userPassword: 'open me', ownerPassword: 'full access',
170
+ algorithm: 'aes-256', // default; 'aes-128' for very old readers
171
+ permissions: { print: true, copy: false, modify: false }, // each defaults to true
172
+ },
173
+ });
174
+ const doc = await loadPdf(bytes, { password: 'open me' }); // doc.encryption.algorithm === 'aes-256'
175
+ ```
176
+
177
+ [`examples/encrypt.mjs`](./examples/encrypt.mjs). RC4 is read, never written.
178
+
179
+ ### Signing (PAdES)
180
+
181
+ ```js
182
+ import { createPdf, loadPdf, nodeSigner } from '@reportwright/pdf';
183
+
184
+ const signer = nodeSigner({ key: fs.readFileSync('key.pem', 'utf8'), certs: fs.readFileSync('cert.pem', 'utf8') });
185
+ const pdf = createPdf(sink, { title: 'Offer letter', sign: true });
186
+ const page = pdf.addPage();
187
+ const hr = pdf.form.signature('hr', { page, x: 40, y: 700, width: 200, height: 40, tooltip: 'HR signature' });
188
+ pdf.form.signature('candidate', { page, x: 300, y: 700, width: 200, height: 40, tooltip: 'Candidate signature' });
189
+ await pdf.sign(hr, { signer, reason: 'Issued', location: 'Bengaluru', certify: 2 });
17
190
  await page.end();
18
- await pdf.end(); // ends the stream too
191
+ await pdf.end(); // the signer runs here
192
+
193
+ // Countersign later: an incremental update, so the first signature stays valid.
194
+ const { bytes } = await (await loadPdf(signedBytes)).saveIncremental({ sign: { field: 'candidate', signer } });
195
+ ```
196
+
197
+ [`examples/sign.mjs`](./examples/sign.mjs) (make a test key with the `openssl` line at its top). Its output checks with
198
+ `pdfsig`: both signatures "Signature is Valid", the second "Total document signed". `signer` can be any async function
199
+ returning CMS bytes (an HSM, a cloud KMS: see `cmsSigner`); add `timestamp` for PAdES B-T.
200
+
201
+ ### PDF/A, PDF/UA and Factur-X
202
+
203
+ ```js
204
+ const pdf = createPdf(sink, {
205
+ title: 'Leave policy',
206
+ tagged: { lang: 'en-GB' }, // PDF/UA-1: structure tree, title, embedded fonts
207
+ pdfa: { part: 2, conformance: 'a' }, // or true (2b), { part: 1, conformance: 'b' }, { part: 3, … }
208
+ });
209
+ const font = await pdf.embedFont(fs.readFileSync('Inter-Regular.ttf'));
210
+ const page = pdf.addPage();
211
+ page.artifact({ type: 'Pagination', subtype: 'Header' }, () => page.text('HR policies', { x: 40, y: 20, font, size: 8 }));
212
+ page.tag('H1', () => page.text('Leave policy', { x: 40, y: 50, font, size: 22 }));
213
+ page.figure({ alt: 'Bar chart: annual leave 18 days' }, () => page.fill(page.path().rect(40, 180, 180, 16), { fill: '#1f4e79' }));
214
+ ```
215
+
216
+ A Factur-X / ZUGFeRD invoice is PDF/A-3 with the CII XML attached:
217
+
218
+ ```js
219
+ const pdf = createPdf(sink, { title: 'Invoice INV-2026-001', pdfa: { part: 3, conformance: 'a' }, tagged: { lang: 'en' } });
220
+ // … draw the invoice from the same data as the XML …
221
+ await pdf.facturX(xmlBytes, { profile: 'EN 16931' }); // MINIMUM, BASIC WL, BASIC, EN 16931, EXTENDED, XRECHNUNG
222
+ ```
223
+
224
+ [`examples/accessible.mjs`](./examples/accessible.mjs) (PDF/A-2a + PDF/UA-1) and
225
+ [`examples/invoice-facturx.mjs`](./examples/invoice-facturx.mjs) (self-contained PDF/A-3a invoice; `factur-x.mjs`
226
+ does the same from your own XML) both pass veraPDF with 0 failed rules (`2a` + `ua1`, `3a` + `ua1`).
227
+
228
+ ### Read: merge, split, extractText, saveIncremental
229
+
230
+ ```js
231
+ import { loadPdf, mergePdfs } from '@reportwright/pdf';
232
+
233
+ const a = await loadPdf(fs.readFileSync('contract.pdf')), b = await loadPdf(fs.readFileSync('annexure.pdf'));
234
+ const merged = await mergePdfs([a, b]); // { bytes, pages, stripped, warnings }
235
+
236
+ const all = await loadPdf(merged.bytes);
237
+ all.extractText(); // every page, joined by \f
238
+ all.page(0).extractText(); // one page, in lines
239
+
240
+ for (let i = 0; i < all.pageCount; i++) { // split: one file per page
241
+ fs.writeFileSync(`page-${i + 1}.pdf`, (await all.fork().reorder([i]).save()).bytes);
242
+ }
243
+
244
+ const doc = await loadPdf(merged.bytes);
245
+ doc.setInfo({ title: 'Contract (reviewed)' });
246
+ const { bytes } = await doc.saveIncremental(); // the original bytes + an appended update
19
247
  ```
20
248
 
21
- In memory: `const bytes = await toBytes((pdf) => { ... }, options)`.
249
+ [`examples/merge-split.mjs`](./examples/merge-split.mjs). `save()` rewrites the file (and strips active content);
250
+ `saveIncremental()` appends (and refuses files with active content): see [Security model](#security-model).
22
251
 
23
- ## Coordinates
252
+ ### Bulk: many files in worker threads
253
+
254
+ ```js
255
+ import { bulk } from '@reportwright/pdf';
256
+
257
+ const files = fs.readdirSync('in').map((f) => fs.readFileSync(`in/${f}`));
258
+ const texts = await bulk(files, 'extractText', { workers: 4 }); // [{ ok: true, value } | { ok: false, error }]
259
+ const locked = await bulk(files, 'encrypt', { encrypt: { userPassword: 'u', ownerPassword: 'o' } });
260
+ const pages = await bulk(files, 'split'); // value: one Uint8Array per page
261
+ const joined = await bulk([[a, b], [c, d]], 'merge'); // each input: the files to merge
262
+ ```
263
+
264
+ Tasks are names (`extractText`, `merge`, `split`, `encrypt`; `BULK_TASKS`), never functions, so only bytes and plain
265
+ options cross to a worker. Results come back in input order; a file that fails (malformed, wrong password, over the
266
+ budget) gives `{ ok: false, error: { name, code, message } }` and the batch goes on. Each worker's heap is capped
267
+ (`memoryMb`, default 512): a file that exhausts it kills only that worker, its item fails with `E_WORKER`, and a new
268
+ worker takes the next. `workers` defaults to `os.availableParallelism() - 1`. Node only: in browsers (no
269
+ `node:worker_threads`), and with `workers: 1`, the same tasks run one at a time in the calling thread, without the
270
+ heap cap. [`examples/bulk.mjs`](./examples/bulk.mjs).
271
+
272
+ Workers stay warm (unref'd, so they never keep the process alive) for 2 s after a call, so the next call skips their
273
+ start-up; small files are sent several at a time (up to 8, or 256 KB, a worker). Each item has `itemTimeoutMs` (default 120000) in a worker: past it the worker is ended and the item fails with `E_TIMEOUT`; a failed item is never retried, so a batch of stuck or worker-killing files ends with each failing alone (`workers` is 1–64, `memoryMb` 16–16384; with `workers: 1` there is no item timeout or heap cap). Pass `signal` to stop a call (or
274
+ `AbortSignal.timeout(ms)` for a time limit): its workers are ended and `bulk` rejects with the signal's reason. When
275
+ to use it: CPU-heavy items (encrypt, the text of long files) gain from the first call; cheap items on small files
276
+ (extractText or merge of 1–2 page files) gain only once the workers are warm: the first call pays ~25 ms a worker to
277
+ start them, so a one-off batch of under ~100 ms of work is quicker with `workers: 1`.
278
+
279
+ Measured on an Apple M2 (4 performance + 4 efficiency cores, 8 GB, busy laptop), Node 22, files made with
280
+ `toBytes` (small: 200 files of 2 pages, 0.4 MB in all; large: 20 files of 200 pages of text, 2.3 MB in all). Items/s
281
+ from the best of 3 calls in one process (warm), and the first call's ms (cold); results equal to `workers: 1`'s
282
+ (encrypt: the same text once opened). Before: one worker started per call, one item a round trip.
283
+
284
+ | files | task | workers | items/s before | items/s now | first call ms before | now | peak RSS MB |
285
+ |---|---|---|---|---|---|---|---|
286
+ | small | extractText | 1 | 4762 | 4762 | 80 | 78 | 135 |
287
+ | small | extractText | 2 | 1923 | 5882 | 106 | 99 | 246 |
288
+ | small | extractText | 4 | 1754 | 6452 | 123 | 100 | 351 |
289
+ | small | merge | 1 | 2174 | 2128 | 137 | 136 | 148 |
290
+ | small | merge | 4 | 1299 | 3175 | 171 | 142 | 429 |
291
+ | small | encrypt | 1 | 284 | 288 | 733 | 738 | 171 |
292
+ | small | encrypt | 4 | 563 | 813 | 370 | 373 | 465 |
293
+ | large | extractText | 1 | 11 | 11 | 1906 | 1891 | 218 |
294
+ | large | extractText | 4 | 29 | 37 | 684 | 653 | 422 |
295
+ | large | merge | 1 | 143 | 135 | 224 | 220 | 176 |
296
+ | large | merge | 4 | 77 | 196 | 276 | 249 | 429 |
297
+ | large | encrypt | 1 | 149 | 155 | 213 | 212 | 181 |
298
+ | large | encrypt | 4 | 93 | 217 | 218 | 210 | 416 |
299
+
300
+ ## Examples
301
+
302
+ | file | what it shows | run |
303
+ |---|---|---|
304
+ | [`write.mjs`](./examples/write.mjs) | text, a PNG, a table from paths, a link, bookmarks | `node examples/write.mjs out.pdf` |
305
+ | [`indic.mjs`](./examples/indic.mjs) | Devanagari shaped by HarfBuzz, Latin fallback | `node examples/indic.mjs NotoSansDevanagari-Regular.ttf out.pdf` |
306
+ | [`harfbuzz-shaper.mjs`](./examples/harfbuzz-shaper.mjs) | the harfbuzzjs shaper adapter (a module) | imported by `indic.mjs` |
307
+ | [`fill-form.mjs`](./examples/fill-form.mjs) | create, fill and flatten a form | `node examples/fill-form.mjs outDir` |
308
+ | [`encrypt.mjs`](./examples/encrypt.mjs) | AES-256 with permissions, read back | `node examples/encrypt.mjs out.pdf` |
309
+ | [`sign.mjs`](./examples/sign.mjs) | PAdES signature while streaming + countersignature | `node examples/sign.mjs key.pem cert.pem out.pdf` |
310
+ | [`accessible.mjs`](./examples/accessible.mjs) | PDF/UA-1 + PDF/A-2a | `node examples/accessible.mjs font.ttf out.pdf` |
311
+ | [`invoice-facturx.mjs`](./examples/invoice-facturx.mjs) | self-contained Factur-X PDF/A-3a invoice | `node examples/invoice-facturx.mjs font.ttf out.pdf` |
312
+ | [`factur-x.mjs`](./examples/factur-x.mjs) | Factur-X from your own CII XML | `node examples/factur-x.mjs invoice.xml font.ttf out.pdf` |
313
+ | [`merge-split.mjs`](./examples/merge-split.mjs) | merge, split, extractText, saveIncremental | `node examples/merge-split.mjs outDir` |
314
+ | [`bulk.mjs`](./examples/bulk.mjs) | extractText and encrypt over many files in worker threads | `node examples/bulk.mjs outDir` |
315
+
316
+ Run from `node_modules/@reportwright/pdf` (or the package folder of the repository).
317
+
318
+ ## Comparison
319
+
320
+ Positioning: **create, comply, sign, light edit**. Only measured facts: `bench/compare.mjs` (Node 20, Apple Silicon laptop, median of 10;
321
+ reading cases median of 3) and the beta.0 measurements. Where something was not measured the cell says so. MuPDF and qpdf are faster at rewriting and repair: this table does not compare against them.
322
+
323
+ | | @reportwright/pdf | pdf-lib 1.17.1 | PDFKit | jsPDF |
324
+ |---|---|---|---|---|
325
+ | Runtime dependencies | 0 | 4 (pako, tslib, @pdf-lib/standard-fonts, @pdf-lib/upng) | not measured | not measured |
326
+ | Package size | 161 KB gzipped (dist/index.js, unminified, beta.2) | not measured | not measured | not measured |
327
+ | 10,000 pages | 3.2 s, 123 MB (beta.0); heap 5.0 MB at 20,000 pages (streaming) | not measured | not measured | not measured |
328
+ | Cold start, new process | 91 ms (beta.0); 43 ms for a 1-page invoice now | 96 ms (1-page invoice) | not measured | not measured |
329
+ | 1,000-page text: bytes / warm ms | 461,320 / 134 | 1,421,783 / 918 | not measured | not measured |
330
+ | 1,000 vector shapes: bytes / warm ms | 8,353 / 5.3 | 22,273 / 20.5 | not measured | not measured |
331
+ | 200-field form, filled: bytes / warm ms | 63,095 / 15.4 | 165,910 / 97.0 | not measured | not measured |
332
+ | Merge 10 ten-page PDFs: warm ms | 11.6 | 14.5 | not measured | not measured |
333
+ | Split 100 pages into 100 files: warm ms | 33.7 | 30.3 (faster) | not measured | not measured |
334
+ | Real-world PDFs read | 196 of 206 (beta.0 corpus) | not measured | not measured | not measured |
335
+ | AES encryption | AES-256, AES-128 | cannot encrypt | not measured | not measured |
336
+ | Digital signatures | PAdES B-B / B-T, incremental countersign | cannot sign | not measured | not measured |
337
+ | PDF/A, PDF/UA | 1b, 2b/2u/2a, 3b/3u/3a, UA-1: veraPDF 0 failures | not measured | not measured | not measured |
338
+ | Factur-X / ZUGFeRD | yes | not measured | not measured | not measured |
339
+
340
+ The split files are larger than pdf-lib's (each carries ~1 KB of XMP and Info). PDFKit and jsPDF were not installed
341
+ where the benchmark ran, so they have no numbers here; run `bench/compare.mjs` where they are. Full tables and method: `NOTES.md` ("Measurements") and the [benchmarks page](https://mrarun005.github.io/reportwright-demo/bench.html).
342
+
343
+ ## Options reference
344
+
345
+ The sections below are the full API: every option of `createPdf`, the drawing methods, forms, encryption, signatures,
346
+ PDF/A levels and the reader. The most used:
347
+
348
+ | function | key options |
349
+ |---|---|
350
+ | `createPdf(sink, o)` / `toBytes(build, o)` | `title`, `author`, `creationDate`, `compress`, `metadata`, `pdfa`, `tagged`, `encrypt`, `sign`, `form.flatten`, `signal`, `timeoutMs`, `id` ([table](#createpdfsink-options)) |
351
+ | `pdf.embedFont(bytes, o)` | `subset`, `shaper`, `textMapping` ([Fonts](#fonts), [Shaping](#shaping)) |
352
+ | `page.text` / `page.textBox` | `x`, `y`, `font`, `size`, `color`, `width`, `align`, `maxLines`, `fallback`, `direction`, `kerning` ([Text options](#pages)) |
353
+ | `pdf.embedImage` / `page.image` | `colorSpace`, `maxDecodedBytes`, `compression` / `width`, `height`, `fit`, `opacity`, `alt` ([Images](#images)) |
354
+ | `page.fill` / `stroke` | `fill`, `stroke`, `width`, `dash`, `cap`, `join`, `opacity`, `blendMode`, `rule` ([Vector graphics](#vector-graphics)) |
355
+ | `encrypt` | `userPassword`, `ownerPassword`, `algorithm`, `permissions`, `encryptMetadata` ([Encryption](#encryption-1)) |
356
+ | `pdf.sign(field, o)` | `signer`, `reason`, `location`, `name`, `certify`, `timestamp`, `subFilter`, `reserve` ([Signatures](#signatures)) |
357
+ | `loadPdf(bytes, o)` | `password`, `budget`, `timeoutMs`, `signal` ([Reading](#reading-and-modifying)) |
358
+ | `doc.save(o)` / `mergePdfs(docs, o)` | `compress`, `keepActiveContent`, `encrypt`, `pdfa`, `keepStructure`, `creationDate` |
359
+ | `doc.saveIncremental(o)` | `sign`, `keepActiveContent`, `allowFillAfterSigning`, `allowChangesAfterSigning` |
360
+
361
+ ### Coordinates
24
362
 
25
363
  PDF points (1/72 inch), **origin at the top-left corner of the page, y growing downwards** (as in CSS and canvas;
26
364
  PDF itself counts from the bottom-left — the library converts). For `page.text`, `y` is the **top of the line box**:
27
365
  the baseline sits at `y + font.ascent * size`.
28
366
 
29
- ## API
30
367
 
31
368
  ### `createPdf(sink, options?)`
32
369
 
@@ -54,7 +391,7 @@ writer to be ready. `pdf.end()` ends a Node stream and closes a `WritableStream`
54
391
  `Helvetica-BoldOblique`, `Times-Roman`, `Times-Bold`, `Times-Italic`, `Times-BoldItalic`, `Courier`,
55
392
  `Courier-Bold`, `Courier-Oblique`, `Courier-BoldOblique`, `Symbol`, `ZapfDingbats`). Not embedded; text is WinAnsi
56
393
  (Latin-1 plus €, curly quotes, dashes…). Not allowed with `pdfa` or `tagged`.
57
- - `await pdf.embedFont(bytes, { subset, shaper })` — a TrueType (`glyf`) or OpenType/CFF (`.otf`) font. Any Unicode
394
+ - `await pdf.embedFont(bytes, { subset, shaper, textMapping })` — a TrueType (`glyf`) or OpenType/CFF (`.otf`) font. Any Unicode
58
395
  text (Identity-H with a ToUnicode map, so text can be copied and searched). TrueType embeds as `FontFile2`
59
396
  (CIDFontType2); OpenType/CFF as `FontFile3 /Subtype /OpenType` (CIDFontType0; `pdffonts` says "CID Type 0C (OT)").
60
397
  `subset`: `true` (default) keeps only the glyphs used, both kinds built in (CFF: unused glyphs' charstrings become
@@ -134,6 +471,10 @@ dots) has its text on one glyph and WJ (U+2060) on the others, which are drawn a
134
471
  `CIDToGIDMap`; a CFF font cannot, so there it gets an ActualText span). The shaper's output is checked (glyph IDs in
135
472
  the font, finite numbers, clusters inside the text, at most 4 glyphs a character).
136
473
 
474
+ `textMapping` (left-to-right clusters in an ActualText span): `'carrier'` (default) also maps the cluster's text on its
475
+ carrier glyph, which suits Firefox, pdf.js and search (MuPDF reads such a cluster twice); `'actualText'` maps the
476
+ carrier to WJ so the text is only in the ActualText, which suits MuPDF/PyMuPDF and AI pipelines (pdf.js loses those clusters).
477
+
137
478
  An adapter for harfbuzzjs v1 (`npm i harfbuzzjs`; not a dependency of this package). The same file ships in the package
138
479
  as `examples/harfbuzz-shaper.mjs`:
139
480
 
@@ -260,10 +601,12 @@ const logo = await pdf.embedImage(fs.readFileSync('logo.png')); // alpha
260
601
  for (const y of [400, 500, 600]) page.image(logo, { x: 40, y, height: 24, opacity: 0.6 }); // written once
261
602
  ```
262
603
 
263
- - `await pdf.embedImage(bytes, { colorSpace, maxDecodedBytes })` — a JPEG or PNG, written to the file at once as an image XObject;
604
+ - `await pdf.embedImage(bytes, { colorSpace, maxDecodedBytes, compression })` — a JPEG or PNG, written to the file at once as an image XObject;
264
605
  returns a handle with `width` and `height` (pixels, after the EXIF orientation). `colorSpace`: an ICC colour space
265
606
  (`pdf.iccColorSpace`) with the image's number of components, instead of the device space. `maxDecodedBytes`: the
266
- decoded size allowed for a PNG that needs decoding (default 64 MB, 67,108,864 bytes).
607
+ decoded size allowed for a PNG that needs decoding (default 64 MB, 67,108,864 bytes). `compression`: the zlib level
608
+ for a PNG that is re-compressed (alpha or interlacing): `'default'` is 6, or 3 above 2 megapixels (about half the
609
+ time for 2–3% more bytes on photos); `'fast'` is 1; or a level 1–9. JPEGs and passed-through PNGs are not affected.
267
610
  - **JPEG**: baseline and progressive, gray, RGB and CMYK, embedded unchanged (DCTDecode, no re-encoding). Adobe's
268
611
  inverted CMYK (an APP14 "Adobe" segment) is drawn with `/Decode [1 0 1 0 1 0 1 0]`. The EXIF orientation (1–8)
269
612
  is honoured. 12-bit, lossless, hierarchical and arithmetic-coded JPEGs are refused.
@@ -281,6 +624,10 @@ for (const y of [400, 500, 600]) page.image(logo, { x: 40, y, height: 24, opacit
281
624
  `height`: the other keeps the aspect ratio; neither: one point per pixel. `fit`: `'fill'` (default: stretch to the
282
625
  box), `'contain'` (all of the image, centred), `'cover'` (fills the box, centred, cut to it). `opacity` 0–1. Tagged
283
626
  PDFs: `alt` makes the image a `Figure` with that alternate text; without `alt` it is an artifact (decoration).
627
+ - Images are deduplicated automatically: `embedImage` called again with the same bytes (and the same options) in the
628
+ same document returns the same handle and writes nothing new (keyed by a SHA-256 of the bytes; Node's `node:crypto`,
629
+ else WebCrypto, else a fast hash confirmed byte for byte). A different option (`colorSpace`, `maxDecodedBytes`, `compression`)
630
+ embeds it again.
284
631
  - An image drawn any number of times, on any pages, is one XObject in the file. Images are always XObjects, never
285
632
  inline images: an inline image would be repeated in every content stream and cannot have a soft mask.
286
633
  - Image handles, like fonts and gradients, only work in the PDF that made them.
@@ -686,9 +1033,23 @@ const signed = await doc.saveIncremental({ sign: { field: { name: 'approval', pa
686
1033
  - Errors from the input are `PdfReadError` with a `code`: `E_MALFORMED`, `E_BUDGET`, `E_DEPTH`, `E_CYCLE`,
687
1034
  `E_TIMEOUT`, `E_ABORTED`, `E_PASSWORD`, `E_UNSUPPORTED`, `E_ACTIVE`.
688
1035
 
689
- ## Security
1036
+ ## Security model
1037
+
1038
+ In short:
690
1039
 
691
- What the library refuses, and why:
1040
+ - **No active content can be written.** There is no API for JavaScript, Launch, SubmitForm, ImportData, GoToR or any
1041
+ other action, no `/AA`, no form scripts; URLs are limited to http, https and mailto (widen per link with
1042
+ `allowSchemes`, never to `javascript:`, `data:`, `file:` or `vbscript:`). Every caller string is written escaped.
1043
+ - **Active content is stripped on read.** `save()` and `mergePdfs()` drop JavaScript, launch and submit actions, `/AA`,
1044
+ XFA, RichMedia and every annotation or action outside an allowlist, and report it in `stripped`
1045
+ (`keepActiveContent: true` opts out). `saveIncremental()` cannot strip (the old bytes stay), so it scans the whole
1046
+ file and refuses (`E_ACTIVE`) instead.
1047
+ - **Every input is budgeted.** Decoded bytes, parse steps, objects, nesting, reference chains, text items and time
1048
+ have caps (`budget`, `timeoutMs`, `signal`); exceeding one is a `PdfReadError` (`E_BUDGET`, `E_TIMEOUT`…), never
1049
+ a hang or an out-of-memory crash. Images, attachments and Factur-X XML have their own size caps.
1050
+ - **Platform crypto only**, no RC4 written, signed documents' DocMDP/FieldMDP permissions enforced on update.
1051
+
1052
+ The details, and what the library refuses and why:
692
1053
 
693
1054
  | Refused | Why |
694
1055
  |---|---|
@@ -742,6 +1103,12 @@ any active content, or if anything cannot be read (fail closed), unless `keepAct
742
1103
  strangers, `save()` is the safe path** (it rewrites the file without the active content; this invalidates existing
743
1104
  signatures).
744
1105
 
1106
+ **This is why `saveIncremental` is slower by design.** Appending a few objects could be done without reading the
1107
+ original, but then active content hidden in it (a shadow attack: an earlier revision, an unreferenced object, a
1108
+ duplicate, an object stream the latest cross-reference does not point at) would ride along under your new signature.
1109
+ So it reads and scans the whole file first, every time: a large file costs a full parse before a byte is appended.
1110
+ Safer, slower, on purpose.
1111
+
745
1112
  **Signed documents (signature policy).** `saveIncremental` reads every signature's permissions and takes the
746
1113
  strictest: DocMDP from the catalog `/Perms` and from each signature's `/Reference` in every revision (no `/P` means
747
1114
  P=2), FieldMDP from each `/Reference` and each signed field's `/Lock`, and ReadOnly from any ancestor field. Anything it
@@ -759,13 +1126,28 @@ after signing", and refilling widgets that existed before signing is how "shadow
759
1126
  Security 2021) make a signed document look different while its signature still verifies. As a second guard the filled
760
1127
  widgets may cover at most 40% of a page in all, and a widget with an opaque background may not lie over the page's text.
761
1128
 
762
- ## Not supported yet
1129
+ ## Known limits
763
1130
 
764
1131
  Image masks (stencils) and colour-managed images (embedded ICC profiles of PNG/JPEG files are ignored); the Unicode
765
- bidi algorithm, hyphenation, kerning of the standard fonts, WOFF fonts; popup annotations; RC4 encryption (refused; read only), public-key (certificate) encryption (neither written nor read), verifying signatures' cryptography (they are listed with their coverage); PDF/A-1a, PDF/X, attachment annotations in PDF/A-3 (use `pdf.attach`), validating a Factur-X invoice's content; layout-perfect text extraction; page edits in an incremental update (use `save()`); reading encrypted files in browsers (WebCrypto has no synchronous AES; Node only). See NOTES.md for the plan. Extending the library: ARCHITECTURE.md.
1132
+ bidi algorithm, hyphenation, kerning of the standard fonts, WOFF fonts; popup annotations; RC4 encryption (refused; read only), public-key (certificate) encryption (neither written nor read), verifying signatures' cryptography (they are listed with their coverage); PDF/A-1a, PDF/X, attachment annotations in PDF/A-3 (use `pdf.attach`), validating a Factur-X invoice's content; layout-perfect text extraction; page edits in an incremental update (use `save()`); reading encrypted files in browsers (WebCrypto has no synchronous AES; Node only); with `textMapping: 'actualText'`, readers that ignore ActualText or close it early (pdf.js; a tester saw an older MuPDF drop a line's last Bengali cluster) lose the spanned clusters: the writer closes every span after its cluster's last glyph, and MuPDF 1.28.2 (PyMuPDF) reads Bengali, Devanagari and the other test scripts whole. See NOTES.md for the plan. Extending the library: ARCHITECTURE.md.
766
1133
 
767
1134
  ## Changes
768
1135
 
1136
+ 0.1.0-beta.10: a WritableStream's writer lock is released when the export ends, fails or is stopped (`timeoutMs`,
1137
+ `signal`), so the caller can reuse or cancel the stream; the abort reason is still the original error.
1138
+
1139
+ 0.1.0-beta.9: stamping every page (`page.draw` then `save()`) is about 35% faster on a 1,000-page file: a page's open
1140
+ graphics states are counted with a byte scan instead of a full parse, and FlateDecode inflates whole, valid zlib
1141
+ streams with Node's zlib (damaged, truncated, raw or over-budget streams still go through fflate as before); the output
1142
+ is byte-identical. `bulk` starts an item's `itemTimeoutMs` clock when its worker is ready, so a respawned worker's
1143
+ start-up no longer fails good files; a stopped export (`timeoutMs` or `signal`) aborts its WritableStream even when
1144
+ the stop lands between two calls.
1145
+
1146
+ 0.1.0-beta.2: a left-to-right shaped cluster in an ActualText span has its text only in the ActualText, so MuPDF
1147
+ no longer reads it twice (pdf.js, which ignores ActualText, now misses those clusters; poppler, PDFium and
1148
+ `extractText` read them); `extractText` drops WJ (U+2060) and reads a code a ToUnicode skips as nothing (U+FFFD
1149
+ only when the font has no mapping at all); `@reportwright/pdf/examples/*` can be imported.
1150
+
769
1151
  0.1.0-beta.1: shaped text (Indic, Arabic) extracts once and in order in pdf.js, poppler and `extractText`;
770
1152
  `extractText` honours ActualText, reads unmapped glyphs as U+FFFD, falls back to an embedded TrueType cmap, spaces
771
1153
  separate text objects (spreadsheet cells), orders a line along its baseline (rotated text too), and stops a
@@ -4,7 +4,7 @@
4
4
 
5
5
  | Component | Version | License | Used for |
6
6
  |---|---|---|---|
7
- | [fflate](https://github.com/101arrowz/fflate) by Arjun Barrett | 0.8.3 | MIT | Flate (zlib) compression and decompression, bundled into `dist/index.js` |
7
+ | [fflate](https://github.com/101arrowz/fflate) by Arjun Barrett | 0.8.3 | MIT | Flate (zlib) compression and decompression, bundled into `dist/fflate.js` |
8
8
  | Adobe Font Metrics (AFM) for the 14 standard PDF fonts, via [@pdf-lib/standard-fonts](https://github.com/Hopding/standard-fonts) | 1.0.0 | MIT (the package); the AFM files are © Adobe, distributed with permission to use, copy and distribute | widths and metrics of Helvetica, Times, Courier, Symbol and ZapfDingbats (`src/fonts/standard-metrics.js`, generated) |
9
9
  | `sRGB-v2-magic.icc` from [Compact-ICC-Profiles](https://github.com/saucecontrol/Compact-ICC-Profiles) by Clinton Ingram | — | CC0 1.0 (public domain) | the sRGB output intent of PDF/A files |
10
10