@reportwright/pdf 0.1.0-beta.0 → 0.1.0-beta.2

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -1,32 +1,320 @@
1
1
  # @reportwright/pdf
2
2
 
3
- A streaming PDF writer for Node and browsers, with no runtime dependencies. Pages are written to your sink as
4
- they end, so memory stays flat however many pages you write. It is the writer inside ReportWright, extracted.
3
+ **Create, comply, sign, light edit: a streaming PDF library for Node and browsers with zero runtime dependencies.**
4
+
5
+ [![npm](https://img.shields.io/npm/v/@reportwright/pdf?color=1f4e79)](https://www.npmjs.com/package/@reportwright/pdf)
6
+ [![license: MIT](https://img.shields.io/badge/license-MIT-blue)](./LICENSE)
7
+ [![dependencies: 0](https://img.shields.io/badge/runtime%20dependencies-0-brightgreen)](./package.json)
8
+ [![node >= 20.16](https://img.shields.io/badge/node-%3E%3D20.16-339933)](./package.json)
9
+ [![PDF/A + PDF/UA: veraPDF](https://img.shields.io/badge/PDF%2FA%201b%E2%80%933a%20%2B%20UA--1-veraPDF%20clean-6f42c1)](./conformance)
10
+
11
+ **Live demo:** [mrarun005.github.io/reportwright-demo](https://mrarun005.github.io/reportwright-demo) ·
12
+ [PDF benchmarks](https://mrarun005.github.io/reportwright-demo/bench.html) ·
13
+ [playground](https://mrarun005.github.io/reportwright-demo/play.html)
14
+
15
+ Pages are written to your sink as they end, so memory stays flat however many pages you write. It is the writer inside
16
+ ReportWright, extracted, plus a hardened reader for merging, splitting, filling and signing existing files.
17
+
18
+ What it is for: **create, comply, sign, light edit**. It does not render PDFs: for rendering use
19
+ [pdf.js](https://mozilla.github.io/pdf.js/) or [PDFium](https://pdfium.googlesource.com/pdfium/). For heavy
20
+ rewriting, repair or raw speed, MuPDF and qpdf are the established tools; this library does not claim to be faster
21
+ than them.
22
+
23
+ - **Write**: text in the 14 standard fonts or embedded TrueType/OpenType (subset, kerned, any Unicode), complex
24
+ scripts through a HarfBuzz shaper, JPEG/PNG images, Canvas-style vector paths, gradients, patterns, layers.
25
+ - **Navigate**: links, named destinations, bookmarks, page labels, annotations.
26
+ - **Forms**: every AcroForm field type with appearance streams; fill and flatten existing forms.
27
+ - **Protect**: AES-256 / AES-128 encryption with permissions; PAdES B-B / B-T signatures while streaming.
28
+ - **Comply**: PDF/A-1b, 2b/2u/2a, 3b/3u/3a and PDF/UA-1 (veraPDF clean), Factur-X / ZUGFeRD / XRechnung.
29
+ - **Read and edit**: load (also damaged or encrypted files), merge, split, reorder, extract text, fill, stamp,
30
+ `saveIncremental` that keeps existing signatures valid.
31
+ - **Safe by default**: no JavaScript or launch actions can be written; active content is stripped on read; every
32
+ input is budgeted.
33
+
34
+ ## Contents
35
+
36
+ [Install](#install) · [Quick start](#quick-start-30-seconds) · [Guide](#guide) · [Examples](#examples) ·
37
+ [Comparison](#comparison) · [Options reference](#options-reference) · [Security model](#security-model) ·
38
+ [Known limits](#known-limits) · [Changes](#changes)
39
+
40
+ ## Install
41
+
42
+ ```sh
43
+ npm i @reportwright/pdf
44
+ ```
45
+
46
+ ESM only, Node 20.16 or later, and modern browsers (the same `dist/index.js`; crypto comes from WebCrypto). Types are
47
+ included (`index.d.ts`). Optional, for complex scripts only: `npm i harfbuzzjs`.
5
48
 
6
- ## Quickstart
49
+ ## Quick start (30 seconds)
7
50
 
8
51
  ```js
52
+ // hello.mjs → node hello.mjs
9
53
  import fs from 'node:fs';
10
54
  import { createPdf } from '@reportwright/pdf';
11
55
 
12
- const pdf = createPdf(fs.createWriteStream('hello.pdf'), { title: 'Hello', creationDate: new Date('2026-01-01') });
13
- const bold = pdf.standardFont('Helvetica-Bold');
14
- const page = pdf.addPage({ width: 595.28, height: 841.89 }); // A4, in points
15
- page.text('Hello, world', { x: 72, y: 72, font: bold, size: 24, color: '#1f4e79' });
56
+ const pdf = createPdf(fs.createWriteStream('hello.pdf'), { title: 'Hello' });
57
+ const page = pdf.addPage(); // A4, points, origin top-left
58
+ page.text('Hello, world', { x: 72, y: 72, font: pdf.standardFont('Helvetica-Bold'), size: 24, color: '#1f4e79' });
16
59
  page.line(72, 104, 523, 104, { color: 0.6, width: 0.5 });
60
+ await page.end(); // the page is written and freed
61
+ console.log(await pdf.end()); // { pages: 1, bytes: …, warnings: [] }
62
+ ```
63
+
64
+ In memory instead of a stream: `const bytes = await toBytes(async (pdf) => { … }, options)`.
65
+
66
+ ## Guide
67
+
68
+ Every snippet below is cut from a file in [`examples/`](./examples), each of which was run and produces a PDF. The
69
+ examples import `../dist/index.js`; in your code import `@reportwright/pdf`. They also ship in the package
70
+ (`node_modules/@reportwright/pdf/examples/`).
71
+
72
+ ### Write: text, fonts, images, tables
73
+
74
+ ```js
75
+ const pdf = createPdf(fs.createWriteStream('report.pdf'), { title: 'Quarterly report', author: 'Example Ltd' });
76
+ const regular = pdf.standardFont('Helvetica'), bold = pdf.standardFont('Helvetica-Bold');
77
+ const inter = await pdf.embedFont(fs.readFileSync('Inter-Regular.ttf')); // subset, any Unicode
78
+ const logo = await pdf.embedImage(fs.readFileSync('logo.png')); // alpha becomes a soft mask
79
+
80
+ const page = pdf.addPage();
81
+ page.image(logo, { x: 40, y: 40, width: 32 });
82
+ page.text('Quarterly report', { x: 84, y: 44, font: bold, size: 22, color: '#1f4e79' });
83
+ const box = page.textBox(longText, { x: 40, y: 90, width: 515, font: inter, size: 11, align: 'justify' });
84
+ // box = { height, lines, overflow }: continue `overflow` in the next box or page
85
+ ```
86
+
87
+ **Tables are paths and text**: there is no table primitive, so you control every rule and alignment.
88
+
89
+ ```js
90
+ const rows = [['Region', 'Q1', 'Q2'], ['North', '1,204', '1,390'], ['South', '980', '1,122']];
91
+ const cols = [40, 300, 430], right = 555, rowH = 22, top = 100 + box.height;
92
+ page.fill(page.path().rect(40, top, right - 40, rowH), { fill: '#1f4e79' }); // header band
93
+ rows.forEach((r, i) => {
94
+ const y = top + i * rowH;
95
+ r.forEach((cell, c) => {
96
+ const font = i === 0 ? bold : regular;
97
+ const x = c === 0 ? cols[c] + 6 : (cols[c + 1] ?? right) - 6 - font.widthOfText(cell, 11); // right-align numbers
98
+ page.text(cell, { x, y: y + 5, font, size: 11, color: i === 0 ? '#ffffff' : 0 });
99
+ });
100
+ if (i > 0) page.stroke(page.path().moveTo(40, y + rowH).lineTo(right, y + rowH), { stroke: 0.75, width: 0.5 });
101
+ });
17
102
  await page.end();
18
- await pdf.end(); // ends the stream too
103
+ await pdf.end();
19
104
  ```
20
105
 
21
- In memory: `const bytes = await toBytes((pdf) => { ... }, options)`.
106
+ Full file: [`examples/write.mjs`](./examples/write.mjs) (it also generates its PNG, so it needs no input files).
107
+
108
+ ### Indic and other complex scripts (HarfBuzz)
109
+
110
+ Pass a shaper to `embedFont` and conjuncts, reordered matras, Arabic joining and marks are drawn correctly; copied
111
+ and extracted text stays the source text, in logical order. The adapter for harfbuzzjs ships as
112
+ [`examples/harfbuzz-shaper.mjs`](./examples/harfbuzz-shaper.mjs).
113
+
114
+ ```js
115
+ import { harfbuzzShaper } from '@reportwright/pdf/examples/harfbuzz-shaper.mjs'; // needs: npm i harfbuzzjs
116
+
117
+ const bytes = fs.readFileSync('NotoSansDevanagari-Regular.ttf');
118
+ const hindi = await pdf.embedFont(bytes, { shaper: harfbuzzShaper(bytes) });
119
+ page.text('किताब हिन्दी क्षत्रिय', { x: 40, y: 40, font: hindi, size: 24 });
120
+ page.textBox('इस किताब में हिन्दी और English दोनों हैं।', { x: 40, y: 90, width: 400, font: hindi, fallback: [latin] });
121
+ ```
122
+
123
+ `node examples/indic.mjs NotoSansDevanagari-Regular.ttf`. Right-to-left: `{ direction: 'rtl' }` (there is no bidi
124
+ algorithm: split mixed-direction lines into runs yourself).
125
+
126
+ ### Links and bookmarks
127
+
128
+ ```js
129
+ page.link(40, 300, 160, 14, { url: 'https://example.com/report' }); // http, https, mailto by default
130
+ pdf.destination('appendix', 2, { top: 40 }); // a named destination
131
+ page.link(40, 320, 160, 14, { dest: 'appendix' }, { alt: 'Go to the appendix' });
132
+ page.link(40, 340, 160, 14, { page: 2, fit: 'FitH', top: 0 });
133
+ pdf.outline([
134
+ { title: 'Quarterly report', page: 1 },
135
+ { title: 'Appendix', page: 2, level: 1 }, // level nests under the previous entry
136
+ ]);
137
+ pdf.pageLabels([{ start: 0, style: 'r' }, { start: 2, style: 'D' }]); // i, ii, 1, 2, …
138
+ ```
139
+
140
+ ### Forms: create, fill, flatten
141
+
142
+ ```js
143
+ import { toBytes, loadPdf } from '@reportwright/pdf';
144
+
145
+ const form = await toBytes(async (pdf) => {
146
+ const font = pdf.standardFont('Helvetica'), page = pdf.addPage();
147
+ pdf.form.textField('name', { page, x: 40, y: 80, width: 240, height: 22, font, tooltip: 'Full name', required: true });
148
+ pdf.form.comboBox('type', { page, x: 40, y: 115, width: 160, height: 22, font, options: ['Annual', 'Sick', 'Unpaid'], value: 'Annual', tooltip: 'Leave type' });
149
+ pdf.form.checkbox('approved', { page, x: 40, y: 150, width: 14, height: 14, tooltip: 'Approved by manager' });
150
+ await page.end();
151
+ });
152
+
153
+ const doc = await loadPdf(form);
154
+ doc.fields; // [{ name: 'name', type: 'text', … }, …]
155
+ doc.fill({ name: 'Aanya Sharma', type: 'Sick', approved: true }); // appearances redrawn by the writer
156
+ const filled = (await doc.save()).bytes;
157
+
158
+ const flattened = (await (await loadPdf(form)).fill({ name: 'Aanya Sharma' }).flatten().save()).bytes; // no AcroForm left
159
+ ```
160
+
161
+ [`examples/fill-form.mjs`](./examples/fill-form.mjs) writes `form.pdf`, `filled.pdf` and `flattened.pdf`. To
162
+ flatten while writing: `createPdf(sink, { form: { flatten: true } })`.
163
+
164
+ ### Encryption
165
+
166
+ ```js
167
+ const bytes = await toBytes(build, {
168
+ encrypt: {
169
+ userPassword: 'open me', ownerPassword: 'full access',
170
+ algorithm: 'aes-256', // default; 'aes-128' for very old readers
171
+ permissions: { print: true, copy: false, modify: false }, // each defaults to true
172
+ },
173
+ });
174
+ const doc = await loadPdf(bytes, { password: 'open me' }); // doc.encryption.algorithm === 'aes-256'
175
+ ```
22
176
 
23
- ## Coordinates
177
+ [`examples/encrypt.mjs`](./examples/encrypt.mjs). RC4 is read, never written.
178
+
179
+ ### Signing (PAdES)
180
+
181
+ ```js
182
+ import { createPdf, loadPdf, nodeSigner } from '@reportwright/pdf';
183
+
184
+ const signer = nodeSigner({ key: fs.readFileSync('key.pem', 'utf8'), certs: fs.readFileSync('cert.pem', 'utf8') });
185
+ const pdf = createPdf(sink, { title: 'Offer letter', sign: true });
186
+ const page = pdf.addPage();
187
+ const hr = pdf.form.signature('hr', { page, x: 40, y: 700, width: 200, height: 40, tooltip: 'HR signature' });
188
+ pdf.form.signature('candidate', { page, x: 300, y: 700, width: 200, height: 40, tooltip: 'Candidate signature' });
189
+ await pdf.sign(hr, { signer, reason: 'Issued', location: 'Bengaluru', certify: 2 });
190
+ await page.end();
191
+ await pdf.end(); // the signer runs here
192
+
193
+ // Countersign later: an incremental update, so the first signature stays valid.
194
+ const { bytes } = await (await loadPdf(signedBytes)).saveIncremental({ sign: { field: 'candidate', signer } });
195
+ ```
196
+
197
+ [`examples/sign.mjs`](./examples/sign.mjs) (make a test key with the `openssl` line at its top). Its output checks with
198
+ `pdfsig`: both signatures "Signature is Valid", the second "Total document signed". `signer` can be any async function
199
+ returning CMS bytes (an HSM, a cloud KMS: see `cmsSigner`); add `timestamp` for PAdES B-T.
200
+
201
+ ### PDF/A, PDF/UA and Factur-X
202
+
203
+ ```js
204
+ const pdf = createPdf(sink, {
205
+ title: 'Leave policy',
206
+ tagged: { lang: 'en-GB' }, // PDF/UA-1: structure tree, title, embedded fonts
207
+ pdfa: { part: 2, conformance: 'a' }, // or true (2b), { part: 1, conformance: 'b' }, { part: 3, … }
208
+ });
209
+ const font = await pdf.embedFont(fs.readFileSync('Inter-Regular.ttf'));
210
+ const page = pdf.addPage();
211
+ page.artifact({ type: 'Pagination', subtype: 'Header' }, () => page.text('HR policies', { x: 40, y: 20, font, size: 8 }));
212
+ page.tag('H1', () => page.text('Leave policy', { x: 40, y: 50, font, size: 22 }));
213
+ page.figure({ alt: 'Bar chart: annual leave 18 days' }, () => page.fill(page.path().rect(40, 180, 180, 16), { fill: '#1f4e79' }));
214
+ ```
215
+
216
+ A Factur-X / ZUGFeRD invoice is PDF/A-3 with the CII XML attached:
217
+
218
+ ```js
219
+ const pdf = createPdf(sink, { title: 'Invoice INV-2026-001', pdfa: { part: 3, conformance: 'a' }, tagged: { lang: 'en' } });
220
+ // … draw the invoice from the same data as the XML …
221
+ await pdf.facturX(xmlBytes, { profile: 'EN 16931' }); // MINIMUM, BASIC WL, BASIC, EN 16931, EXTENDED, XRECHNUNG
222
+ ```
223
+
224
+ [`examples/accessible.mjs`](./examples/accessible.mjs) (PDF/A-2a + PDF/UA-1) and
225
+ [`examples/invoice-facturx.mjs`](./examples/invoice-facturx.mjs) (self-contained PDF/A-3a invoice; `factur-x.mjs`
226
+ does the same from your own XML) both pass veraPDF with 0 failed rules (`2a` + `ua1`, `3a` + `ua1`).
227
+
228
+ ### Read: merge, split, extractText, saveIncremental
229
+
230
+ ```js
231
+ import { loadPdf, mergePdfs } from '@reportwright/pdf';
232
+
233
+ const a = await loadPdf(fs.readFileSync('contract.pdf')), b = await loadPdf(fs.readFileSync('annexure.pdf'));
234
+ const merged = await mergePdfs([a, b]); // { bytes, pages, stripped, warnings }
235
+
236
+ const all = await loadPdf(merged.bytes);
237
+ all.extractText(); // every page, joined by \f
238
+ all.page(0).extractText(); // one page, in lines
239
+
240
+ for (let i = 0; i < all.pageCount; i++) { // split: one file per page
241
+ fs.writeFileSync(`page-${i + 1}.pdf`, (await all.fork().reorder([i]).save()).bytes);
242
+ }
243
+
244
+ const doc = await loadPdf(merged.bytes);
245
+ doc.setInfo({ title: 'Contract (reviewed)' });
246
+ const { bytes } = await doc.saveIncremental(); // the original bytes + an appended update
247
+ ```
248
+
249
+ [`examples/merge-split.mjs`](./examples/merge-split.mjs). `save()` rewrites the file (and strips active content);
250
+ `saveIncremental()` appends (and refuses files with active content): see [Security model](#security-model).
251
+
252
+ ## Examples
253
+
254
+ | file | what it shows | run |
255
+ |---|---|---|
256
+ | [`write.mjs`](./examples/write.mjs) | text, a PNG, a table from paths, a link, bookmarks | `node examples/write.mjs out.pdf` |
257
+ | [`indic.mjs`](./examples/indic.mjs) | Devanagari shaped by HarfBuzz, Latin fallback | `node examples/indic.mjs NotoSansDevanagari-Regular.ttf out.pdf` |
258
+ | [`harfbuzz-shaper.mjs`](./examples/harfbuzz-shaper.mjs) | the harfbuzzjs shaper adapter (a module) | imported by `indic.mjs` |
259
+ | [`fill-form.mjs`](./examples/fill-form.mjs) | create, fill and flatten a form | `node examples/fill-form.mjs outDir` |
260
+ | [`encrypt.mjs`](./examples/encrypt.mjs) | AES-256 with permissions, read back | `node examples/encrypt.mjs out.pdf` |
261
+ | [`sign.mjs`](./examples/sign.mjs) | PAdES signature while streaming + countersignature | `node examples/sign.mjs key.pem cert.pem out.pdf` |
262
+ | [`accessible.mjs`](./examples/accessible.mjs) | PDF/UA-1 + PDF/A-2a | `node examples/accessible.mjs font.ttf out.pdf` |
263
+ | [`invoice-facturx.mjs`](./examples/invoice-facturx.mjs) | self-contained Factur-X PDF/A-3a invoice | `node examples/invoice-facturx.mjs font.ttf out.pdf` |
264
+ | [`factur-x.mjs`](./examples/factur-x.mjs) | Factur-X from your own CII XML | `node examples/factur-x.mjs invoice.xml font.ttf out.pdf` |
265
+ | [`merge-split.mjs`](./examples/merge-split.mjs) | merge, split, extractText, saveIncremental | `node examples/merge-split.mjs outDir` |
266
+
267
+ Run from `node_modules/@reportwright/pdf` (or the package folder of the repository).
268
+
269
+ ## Comparison
270
+
271
+ Positioning: **create, comply, sign, light edit**. Only measured facts: `bench/compare.mjs` (Node 20, Apple Silicon laptop, median of 10;
272
+ reading cases median of 3) and the beta.0 measurements. Where something was not measured the cell says so. MuPDF and qpdf are faster at rewriting and repair: this table does not compare against them.
273
+
274
+ | | @reportwright/pdf | pdf-lib 1.17.1 | PDFKit | jsPDF |
275
+ |---|---|---|---|---|
276
+ | Runtime dependencies | 0 | 4 (pako, tslib, @pdf-lib/standard-fonts, @pdf-lib/upng) | not measured | not measured |
277
+ | Package size | 161 KB gzipped (dist/index.js, unminified, beta.2) | not measured | not measured | not measured |
278
+ | 10,000 pages | 3.2 s, 123 MB (beta.0); heap 5.0 MB at 20,000 pages (streaming) | not measured | not measured | not measured |
279
+ | Cold start, new process | 91 ms (beta.0); 43 ms for a 1-page invoice now | 96 ms (1-page invoice) | not measured | not measured |
280
+ | 1,000-page text: bytes / warm ms | 461,320 / 134 | 1,421,783 / 918 | not measured | not measured |
281
+ | 1,000 vector shapes: bytes / warm ms | 8,353 / 5.3 | 22,273 / 20.5 | not measured | not measured |
282
+ | 200-field form, filled: bytes / warm ms | 63,095 / 15.4 | 165,910 / 97.0 | not measured | not measured |
283
+ | Merge 10 ten-page PDFs: warm ms | 11.6 | 14.5 | not measured | not measured |
284
+ | Split 100 pages into 100 files: warm ms | 33.7 | 30.3 (faster) | not measured | not measured |
285
+ | Real-world PDFs read | 196 of 206 (beta.0 corpus) | not measured | not measured | not measured |
286
+ | AES encryption | AES-256, AES-128 | cannot encrypt | not measured | not measured |
287
+ | Digital signatures | PAdES B-B / B-T, incremental countersign | cannot sign | not measured | not measured |
288
+ | PDF/A, PDF/UA | 1b, 2b/2u/2a, 3b/3u/3a, UA-1: veraPDF 0 failures | not measured | not measured | not measured |
289
+ | Factur-X / ZUGFeRD | yes | not measured | not measured | not measured |
290
+
291
+ The split files are larger than pdf-lib's (each carries ~1 KB of XMP and Info). PDFKit and jsPDF were not installed
292
+ where the benchmark ran, so they have no numbers here; run `bench/compare.mjs` where they are. Full tables and method: `NOTES.md` ("Measurements") and the [benchmarks page](https://mrarun005.github.io/reportwright-demo/bench.html).
293
+
294
+ ## Options reference
295
+
296
+ The sections below are the full API: every option of `createPdf`, the drawing methods, forms, encryption, signatures,
297
+ PDF/A levels and the reader. The most used:
298
+
299
+ | function | key options |
300
+ |---|---|
301
+ | `createPdf(sink, o)` / `toBytes(build, o)` | `title`, `author`, `creationDate`, `compress`, `metadata`, `pdfa`, `tagged`, `encrypt`, `sign`, `form.flatten`, `signal`, `timeoutMs`, `id` ([table](#createpdfsink-options)) |
302
+ | `pdf.embedFont(bytes, o)` | `subset`, `shaper`, `textMapping` ([Fonts](#fonts), [Shaping](#shaping)) |
303
+ | `page.text` / `page.textBox` | `x`, `y`, `font`, `size`, `color`, `width`, `align`, `maxLines`, `fallback`, `direction`, `kerning` ([Text options](#pages)) |
304
+ | `pdf.embedImage` / `page.image` | `colorSpace`, `maxDecodedBytes` / `width`, `height`, `fit`, `opacity`, `alt` ([Images](#images)) |
305
+ | `page.fill` / `stroke` | `fill`, `stroke`, `width`, `dash`, `cap`, `join`, `opacity`, `blendMode`, `rule` ([Vector graphics](#vector-graphics)) |
306
+ | `encrypt` | `userPassword`, `ownerPassword`, `algorithm`, `permissions`, `encryptMetadata` ([Encryption](#encryption-1)) |
307
+ | `pdf.sign(field, o)` | `signer`, `reason`, `location`, `name`, `certify`, `timestamp`, `subFilter`, `reserve` ([Signatures](#signatures)) |
308
+ | `loadPdf(bytes, o)` | `password`, `budget`, `timeoutMs`, `signal` ([Reading](#reading-and-modifying)) |
309
+ | `doc.save(o)` / `mergePdfs(docs, o)` | `compress`, `keepActiveContent`, `encrypt`, `pdfa`, `keepStructure`, `creationDate` |
310
+ | `doc.saveIncremental(o)` | `sign`, `keepActiveContent`, `allowFillAfterSigning`, `allowChangesAfterSigning` |
311
+
312
+ ### Coordinates
24
313
 
25
314
  PDF points (1/72 inch), **origin at the top-left corner of the page, y growing downwards** (as in CSS and canvas;
26
315
  PDF itself counts from the bottom-left — the library converts). For `page.text`, `y` is the **top of the line box**:
27
316
  the baseline sits at `y + font.ascent * size`.
28
317
 
29
- ## API
30
318
 
31
319
  ### `createPdf(sink, options?)`
32
320
 
@@ -38,6 +326,7 @@ writer to be ready. `pdf.end()` ends a Node stream and closes a `WritableStream`
38
326
  | option | meaning |
39
327
  |---|---|
40
328
  | `title`, `author`, `subject`, `keywords` | document info, mirrored in the XMP metadata (every value escaped) |
329
+ | `metadata` | `true` (default): an XMP metadata stream mirrors the info. `false`: the Info dictionary only, no XMP (a one-page invoice goes from 2.5 KB to 1.6 KB). Refused with `pdfa` or `tagged`: PDF/A and PDF/UA require the XMP |
41
330
  | `creationDate` | a `Date` (default: now). With the same input and date the output is byte-for-byte identical |
42
331
  | `compress` | `true` (default): content, fonts, object streams and the xref stream are Flate-compressed |
43
332
  | `pdfa` | `true`: PDF/A-2b (sRGB output intent bundled, XMP identification). Fonts must be embedded. `{ part, conformance, outputIntent }`: a level (see PDF/A levels) and/or your RGB or CMYK output profile instead of sRGB (a CMYK intent allows CMYK colour and refuses RGB) |
@@ -53,7 +342,7 @@ writer to be ready. `pdf.end()` ends a Node stream and closes a `WritableStream`
53
342
  `Helvetica-BoldOblique`, `Times-Roman`, `Times-Bold`, `Times-Italic`, `Times-BoldItalic`, `Courier`,
54
343
  `Courier-Bold`, `Courier-Oblique`, `Courier-BoldOblique`, `Symbol`, `ZapfDingbats`). Not embedded; text is WinAnsi
55
344
  (Latin-1 plus €, curly quotes, dashes…). Not allowed with `pdfa` or `tagged`.
56
- - `await pdf.embedFont(bytes, { subset, shaper })` — a TrueType (`glyf`) or OpenType/CFF (`.otf`) font. Any Unicode
345
+ - `await pdf.embedFont(bytes, { subset, shaper, textMapping })` — a TrueType (`glyf`) or OpenType/CFF (`.otf`) font. Any Unicode
57
346
  text (Identity-H with a ToUnicode map, so text can be copied and searched). TrueType embeds as `FontFile2`
58
347
  (CIDFontType2); OpenType/CFF as `FontFile3 /Subtype /OpenType` (CIDFontType0; `pdffonts` says "CID Type 0C (OT)").
59
348
  `subset`: `true` (default) keeps only the glyphs used, both kinds built in (CFF: unused glyphs' charstrings become
@@ -61,6 +350,9 @@ writer to be ready. `pdf.end()` ends a Node stream and closes a `WritableStream`
61
350
  CID-keyed CFF gets an identity charset so every reader maps its glyphs alike); `false` embeds the whole font; or a
62
351
  function `(font, codePoints, { glyphs }) => Promise<Uint8Array>` (e.g. a HarfBuzz subsetter, which also
63
352
  desubroutinises CFF) that must keep glyph IDs. `shaper`: see Shaping.
353
+ A variable font (an `fvar` table) is drawn at its default instance, which may be its thinnest weight (Noto Sans JP's
354
+ is Thin), with a warning naming the axes' values: for another weight, embed a static instance of it (fonttools
355
+ `varLib.instancer`, or `hb-subset --instance`).
64
356
  - `font.widthOfText(text, size, { kerning, characterSpacing, wordSpacing, horizontalScaling, direction })` — the width
65
357
  text is drawn at (kerned, shaped, spaced); `font.widthOf(text, size)` — plain advances; `font.lineHeight(size)` —
66
358
  ascent − descent + line gap (1.2 × size for standard fonts); `font.ascent`, `font.descent` (fractions of the size).
@@ -123,14 +415,35 @@ Text options (`page.text` and `page.textBox`):
123
415
  returns the glyphs in visual order, in font units: `[{ g, cl, ax, dx, dy }]` (glyph ID, cluster = UTF-16 index of
124
416
  its first character, x advance, offsets: what HarfBuzz's `hb_shape` gives). So Arabic, Indic and other complex
125
417
  scripts draw correctly. The glyphs are drawn by ID with their advances and offsets (TJ adjustments, `Ts` for vertical
126
- offsets); a glyph standing for its cluster's text maps it in ToUnicode (a ligature maps to all its letters); a cluster
127
- of several glyphs (reordered matras, conjuncts) or a glyph already standing for other text is wrapped in an
128
- `/ActualText` span, so copied text is the source text. The shaper's output is checked (glyph IDs in the font, finite
129
- numbers, clusters inside the text, at most 4 glyphs a character). `examples/harfbuzz-shaper.mjs` is an adapter for
130
- harfbuzzjs (not a dependency):
418
+ offsets); copied and extracted text is the source text, once, in logical order (checked with Acrobat-style
419
+ ActualText readers, poppler and pdf.js): a cluster of several glyphs (a conjunct, a reordered matra, a letter and its
420
+ dots) has its text on one glyph and WJ (U+2060) on the others, which are drawn after the run or inside an
421
+ `/ActualText` span, and a glyph that stands for different text in different places gets a code for each (a
422
+ `CIDToGIDMap`; a CFF font cannot, so there it gets an ActualText span). The shaper's output is checked (glyph IDs in
423
+ the font, finite numbers, clusters inside the text, at most 4 glyphs a character).
424
+
425
+ `textMapping` (left-to-right clusters in an ActualText span): `'carrier'` (default) also maps the cluster's text on its
426
+ carrier glyph, which suits Firefox, pdf.js and search (MuPDF reads such a cluster twice); `'actualText'` maps the
427
+ carrier to WJ so the text is only in the ActualText, which suits MuPDF/PyMuPDF and AI pipelines (pdf.js loses those clusters).
428
+
429
+ An adapter for harfbuzzjs v1 (`npm i harfbuzzjs`; not a dependency of this package). The same file ships in the package
430
+ as `examples/harfbuzz-shaper.mjs`:
131
431
 
132
432
  ```js
133
- import { harfbuzzShaper } from './harfbuzz-shaper.mjs'; // copy it from examples/
433
+ import * as hb from 'harfbuzzjs';
434
+ function harfbuzzShaper(bytes) {
435
+ const font = new hb.Font(new hb.Face(new hb.Blob(bytes))); // positions in font units
436
+ return (text, _font, { direction, kerning }) => {
437
+ const buf = new hb.Buffer();
438
+ buf.addText(text); // clusters are UTF-16 indices
439
+ buf.guessSegmentProperties(); // script and language from the text
440
+ buf.setDirection(direction === 'rtl' ? hb.Direction.RTL : hb.Direction.LTR);
441
+ hb.shape(font, buf, kerning ? [] : [new hb.Feature('kern', 0)]);
442
+ const pos = buf.getGlyphPositions();
443
+ return buf.getGlyphInfos().map((g, i) => ({ g: g.codepoint, cl: g.cluster, ax: pos[i].xAdvance, dx: pos[i].xOffset, dy: pos[i].yOffset }));
444
+ };
445
+ }
446
+
134
447
  const bytes = fs.readFileSync('NotoSansDevanagari-Regular.ttf');
135
448
  const hindi = await pdf.embedFont(bytes, { shaper: harfbuzzShaper(bytes) });
136
449
  page.text('किताब हिन्दी', { x: 40, y: 40, font: hindi, size: 18 });
@@ -260,6 +573,10 @@ for (const y of [400, 500, 600]) page.image(logo, { x: 40, y, height: 24, opacit
260
573
  `height`: the other keeps the aspect ratio; neither: one point per pixel. `fit`: `'fill'` (default: stretch to the
261
574
  box), `'contain'` (all of the image, centred), `'cover'` (fills the box, centred, cut to it). `opacity` 0–1. Tagged
262
575
  PDFs: `alt` makes the image a `Figure` with that alternate text; without `alt` it is an artifact (decoration).
576
+ - Images are deduplicated automatically: `embedImage` called again with the same bytes (and the same options) in the
577
+ same document returns the same handle and writes nothing new (keyed by a SHA-256 of the bytes; Node's `node:crypto`,
578
+ else WebCrypto, else a fast hash confirmed byte for byte). A different option (`colorSpace`, `maxDecodedBytes`)
579
+ embeds it again.
263
580
  - An image drawn any number of times, on any pages, is one XObject in the file. Images are always XObjects, never
264
581
  inline images: an inline image would be repeated in every content stream and cannot have a soft mask.
265
582
  - Image handles, like fonts and gradients, only work in the PDF that made them.
@@ -657,13 +974,31 @@ const signed = await doc.saveIncremental({ sign: { field: { name: 'approval', pa
657
974
  sign: an empty signature field by name, or a new one (`{ name, page, x, y, width, height }`; no size: invisible), with
658
975
  group F's signers, the `/ByteRange` over the whole file including the original bytes. An encrypted file's new objects
659
976
  use its own AES key; RC4 files cannot be updated (writing RC4 is refused).
660
- - `mergePdfs(docs, options)`: every page of each, into a new file.
977
+ - `mergePdfs(docs, options)`: every page of each, into a new file. **A merged PDF/UA (or PDF/A-a) file is no longer
978
+ tagged**: the structure tree is kept only when every page comes from one document (and the form is not flattened),
979
+ so a merge drops it with `/MarkInfo` and `/Lang`, the XMP's PDF/UA identification (`pdfuaid`) goes too, PDF/A-a is
980
+ claimed as PDF/A-b, and `warnings` says "the structure tree (tags) was dropped … the result is not tagged (no
981
+ longer PDF/UA …)". Tag the merged content again with the writer if the result must be PDF/UA.
661
982
  - Errors from the input are `PdfReadError` with a `code`: `E_MALFORMED`, `E_BUDGET`, `E_DEPTH`, `E_CYCLE`,
662
983
  `E_TIMEOUT`, `E_ABORTED`, `E_PASSWORD`, `E_UNSUPPORTED`, `E_ACTIVE`.
663
984
 
664
- ## Security
985
+ ## Security model
986
+
987
+ In short:
665
988
 
666
- What the library refuses, and why:
989
+ - **No active content can be written.** There is no API for JavaScript, Launch, SubmitForm, ImportData, GoToR or any
990
+ other action, no `/AA`, no form scripts; URLs are limited to http, https and mailto (widen per link with
991
+ `allowSchemes`, never to `javascript:`, `data:`, `file:` or `vbscript:`). Every caller string is written escaped.
992
+ - **Active content is stripped on read.** `save()` and `mergePdfs()` drop JavaScript, launch and submit actions, `/AA`,
993
+ XFA, RichMedia and every annotation or action outside an allowlist, and report it in `stripped`
994
+ (`keepActiveContent: true` opts out). `saveIncremental()` cannot strip (the old bytes stay), so it scans the whole
995
+ file and refuses (`E_ACTIVE`) instead.
996
+ - **Every input is budgeted.** Decoded bytes, parse steps, objects, nesting, reference chains, text items and time
997
+ have caps (`budget`, `timeoutMs`, `signal`); exceeding one is a `PdfReadError` (`E_BUDGET`, `E_TIMEOUT`…), never
998
+ a hang or an out-of-memory crash. Images, attachments and Factur-X XML have their own size caps.
999
+ - **Platform crypto only**, no RC4 written, signed documents' DocMDP/FieldMDP permissions enforced on update.
1000
+
1001
+ The details, and what the library refuses and why:
667
1002
 
668
1003
  | Refused | Why |
669
1004
  |---|---|
@@ -689,7 +1024,7 @@ JavaScript, launch and submit actions cannot be written at all (see Forms and Li
689
1024
  |---|---|
690
1025
  | Decoded stream bytes, the whole document / one stream (every filter and predictor counts as it produces, so a zip bomb stops at the cap: a 302-byte stream that inflates to 10 GB stops at 128 MB) | `maxDecoded` 512 MB / `maxStream` 128 MB |
691
1026
  | Parsing steps (tokens, xref rows, tree nodes, operators); time | `maxSteps` 200 M; `timeoutMs`, `signal` |
692
- | Cross-reference entries and scanned objects held; one page's text pieces | `maxObjects` 5 M; `maxTextItems` 2 M |
1027
+ | Cross-reference entries and scanned objects held; one page's text pieces | `maxObjects` 5 M; `maxTextItems` 1 M (and 16 M characters a page, ActualText included) |
693
1028
  | Nesting of arrays and dictionaries (an explicit stack, never recursion) | 256 |
694
1029
  | A chain of indirect references | 32 |
695
1030
 
@@ -717,6 +1052,12 @@ any active content, or if anything cannot be read (fail closed), unless `keepAct
717
1052
  strangers, `save()` is the safe path** (it rewrites the file without the active content; this invalidates existing
718
1053
  signatures).
719
1054
 
1055
+ **This is why `saveIncremental` is slower by design.** Appending a few objects could be done without reading the
1056
+ original, but then active content hidden in it (a shadow attack: an earlier revision, an unreferenced object, a
1057
+ duplicate, an object stream the latest cross-reference does not point at) would ride along under your new signature.
1058
+ So it reads and scans the whole file first, every time: a large file costs a full parse before a byte is appended.
1059
+ Safer, slower, on purpose.
1060
+
720
1061
  **Signed documents (signature policy).** `saveIncremental` reads every signature's permissions and takes the
721
1062
  strictest: DocMDP from the catalog `/Perms` and from each signature's `/Reference` in every revision (no `/P` means
722
1063
  P=2), FieldMDP from each `/Reference` and each signed field's `/Lock`, and ReadOnly from any ancestor field. Anything it
@@ -734,7 +1075,20 @@ after signing", and refilling widgets that existed before signing is how "shadow
734
1075
  Security 2021) make a signed document look different while its signature still verifies. As a second guard the filled
735
1076
  widgets may cover at most 40% of a page in all, and a widget with an opaque background may not lie over the page's text.
736
1077
 
737
- ## Not supported yet
1078
+ ## Known limits
738
1079
 
739
1080
  Image masks (stencils) and colour-managed images (embedded ICC profiles of PNG/JPEG files are ignored); the Unicode
740
1081
  bidi algorithm, hyphenation, kerning of the standard fonts, WOFF fonts; popup annotations; RC4 encryption (refused; read only), public-key (certificate) encryption (neither written nor read), verifying signatures' cryptography (they are listed with their coverage); PDF/A-1a, PDF/X, attachment annotations in PDF/A-3 (use `pdf.attach`), validating a Factur-X invoice's content; layout-perfect text extraction; page edits in an incremental update (use `save()`); reading encrypted files in browsers (WebCrypto has no synchronous AES; Node only). See NOTES.md for the plan. Extending the library: ARCHITECTURE.md.
1082
+
1083
+ ## Changes
1084
+
1085
+ 0.1.0-beta.2: a left-to-right shaped cluster in an ActualText span has its text only in the ActualText, so MuPDF
1086
+ no longer reads it twice (pdf.js, which ignores ActualText, now misses those clusters; poppler, PDFium and
1087
+ `extractText` read them); `extractText` drops WJ (U+2060) and reads a code a ToUnicode skips as nothing (U+FFFD
1088
+ only when the font has no mapping at all); `@reportwright/pdf/examples/*` can be imported.
1089
+
1090
+ 0.1.0-beta.1: shaped text (Indic, Arabic) extracts once and in order in pdf.js, poppler and `extractText`;
1091
+ `extractText` honours ActualText, reads unmapped glyphs as U+FFFD, falls back to an embedded TrueType cmap, spaces
1092
+ separate text objects (spreadsheet cells), orders a line along its baseline (rotated text too), and stops a
1093
+ 2 M-operator page at 170 MB (`maxTextItems` 1 M); standard-font text writes 40% faster; variable fonts warn; the
1094
+ package ships `examples/`. The list with details: NOTES.md ("Changes") in the repository.