pdf-codec 1.6.0 → 1.8.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +75 -12
- package/dist/cff-bounds.cjs +432 -0
- package/dist/cff-bounds.d.cts +9 -0
- package/dist/cff-bounds.d.ts +9 -0
- package/dist/cff-bounds.js +431 -0
- package/dist/cff-probe.cjs +8 -106
- package/dist/cff-probe.js +9 -107
- package/dist/cff.cjs +158 -0
- package/dist/cff.d.cts +17 -0
- package/dist/cff.d.ts +17 -0
- package/dist/cff.js +150 -0
- package/dist/filters.cjs +45 -2
- package/dist/filters.d.cts +4 -3
- package/dist/filters.d.ts +4 -3
- package/dist/filters.js +45 -2
- package/dist/formula.d.cts +1 -1
- package/dist/formula.d.ts +1 -1
- package/dist/glyf.cjs +62 -6
- package/dist/glyf.d.cts +2 -0
- package/dist/glyf.d.ts +2 -0
- package/dist/glyf.js +62 -6
- package/dist/glyph-bounds-BV_Z40KS.d.cts +10 -0
- package/dist/glyph-bounds-BV_Z40KS.d.ts +10 -0
- package/dist/glyph-bounds.cjs +12 -0
- package/dist/glyph-bounds.d.cts +2 -0
- package/dist/glyph-bounds.d.ts +2 -0
- package/dist/glyph-bounds.js +11 -0
- package/dist/image/jbig2-arith.cjs +432 -0
- package/dist/image/jbig2-arith.d.cts +23 -0
- package/dist/image/jbig2-arith.d.ts +23 -0
- package/dist/image/jbig2-arith.js +426 -0
- package/dist/image/jbig2-bitmap.cjs +80 -0
- package/dist/image/jbig2-bitmap.d.cts +20 -0
- package/dist/image/jbig2-bitmap.d.ts +20 -0
- package/dist/image/jbig2-bitmap.js +72 -0
- package/dist/image/jbig2-errors.cjs +17 -0
- package/dist/image/jbig2-errors.d.cts +2 -0
- package/dist/image/jbig2-errors.d.ts +2 -0
- package/dist/image/jbig2-errors.js +15 -0
- package/dist/image/jbig2-generic.cjs +242 -0
- package/dist/image/jbig2-generic.d.cts +27 -0
- package/dist/image/jbig2-generic.d.ts +27 -0
- package/dist/image/jbig2-generic.js +237 -0
- package/dist/image/jbig2-text.cjs +175 -0
- package/dist/image/jbig2-text.d.cts +58 -0
- package/dist/image/jbig2-text.d.ts +58 -0
- package/dist/image/jbig2-text.js +170 -0
- package/dist/image/jbig2.cjs +333 -0
- package/dist/image/jbig2.d.cts +15 -0
- package/dist/image/jbig2.d.ts +15 -0
- package/dist/image/jbig2.js +332 -0
- package/dist/image/jp2-boxes.cjs +148 -0
- package/dist/image/jp2-boxes.d.cts +26 -0
- package/dist/image/jp2-boxes.d.ts +26 -0
- package/dist/image/jp2-boxes.js +146 -0
- package/dist/image/jpeg2000-codestream.cjs +309 -0
- package/dist/image/jpeg2000-codestream.d.cts +78 -0
- package/dist/image/jpeg2000-codestream.d.ts +78 -0
- package/dist/image/jpeg2000-codestream.js +308 -0
- package/dist/image/jpeg2000-dwt.cjs +148 -0
- package/dist/image/jpeg2000-dwt.d.cts +34 -0
- package/dist/image/jpeg2000-dwt.d.ts +34 -0
- package/dist/image/jpeg2000-dwt.js +145 -0
- package/dist/image/jpeg2000-errors.cjs +17 -0
- package/dist/image/jpeg2000-errors.d.cts +9 -0
- package/dist/image/jpeg2000-errors.d.ts +9 -0
- package/dist/image/jpeg2000-errors.js +15 -0
- package/dist/image/jpeg2000-t1.cjs +254 -0
- package/dist/image/jpeg2000-t1.d.cts +18 -0
- package/dist/image/jpeg2000-t1.d.ts +18 -0
- package/dist/image/jpeg2000-t1.js +253 -0
- package/dist/image/jpeg2000-t2.cjs +316 -0
- package/dist/image/jpeg2000-t2.d.cts +79 -0
- package/dist/image/jpeg2000-t2.d.ts +79 -0
- package/dist/image/jpeg2000-t2.js +311 -0
- package/dist/image/jpeg2000-tagtree.cjs +105 -0
- package/dist/image/jpeg2000-tagtree.d.cts +29 -0
- package/dist/image/jpeg2000-tagtree.d.ts +29 -0
- package/dist/image/jpeg2000-tagtree.js +103 -0
- package/dist/image/jpeg2000.cjs +322 -0
- package/dist/image/jpeg2000.d.cts +54 -0
- package/dist/image/jpeg2000.d.ts +54 -0
- package/dist/image/jpeg2000.js +320 -0
- package/dist/images-read.cjs +78 -1
- package/dist/images-read.js +78 -1
- package/dist/index.cjs +17 -0
- package/dist/index.d.cts +11 -2
- package/dist/index.d.ts +11 -2
- package/dist/index.js +7 -1
- package/dist/jbig2-errors-Bj6MPxrx.d.cts +9 -0
- package/dist/jbig2-errors-Bj6MPxrx.d.ts +9 -0
- package/dist/math-content-write.cjs +37 -1
- package/dist/math-content-write.d.cts +1 -1
- package/dist/math-content-write.d.ts +1 -1
- package/dist/math-content-write.js +38 -2
- package/dist/math-font-write.cjs +6 -1
- package/dist/math-font-write.d.cts +1 -1
- package/dist/math-font-write.d.ts +1 -1
- package/dist/math-font-write.js +6 -1
- package/dist/math-font.cjs +91 -18
- package/dist/math-font.d.cts +8 -1
- package/dist/math-font.d.ts +8 -1
- package/dist/math-font.js +91 -18
- package/dist/math-stretch-Bwk_bxh1.d.ts +24 -0
- package/dist/math-stretch-DXWjs8gL.d.cts +24 -0
- package/dist/math-stretch.cjs +96 -0
- package/dist/math-stretch.d.cts +2 -0
- package/dist/math-stretch.d.ts +2 -0
- package/dist/math-stretch.js +94 -0
- package/dist/math-table-DhPrNoPa.d.ts +70 -0
- package/dist/math-table-gFNB8DmQ.d.cts +70 -0
- package/dist/math-table.cjs +69 -1
- package/dist/math-table.d.cts +2 -45
- package/dist/math-table.d.ts +2 -45
- package/dist/math-table.js +69 -1
- package/dist/{math-types-Ba8qAf1p.d.ts → math-types-BZkWO5Ow.d.cts} +30 -2
- package/dist/{math-types-Ba8qAf1p.d.cts → math-types-Bq7tyV4g.d.ts} +30 -2
- package/dist/math-types.d.cts +3 -2
- package/dist/math-types.d.ts +3 -2
- package/package.json +1 -1
package/README.md
CHANGED
|
@@ -98,7 +98,27 @@ const { box } = layoutFormula(mathml, { metrics: metricsAt(12), sizePt: 12, colo
|
|
|
98
98
|
const pdfBytes = writePdf(doc, { formulas: [{ pageIndex: 0, xPt: 50, yPt: 700, box }] });
|
|
99
99
|
```
|
|
100
100
|
|
|
101
|
-
Because `MathBox` and its own constituent types (`MathGlyphRun`/`MathRule`/`MathStroke`/`MathColor`) are plain, structurally-typed data — not a class, not branded — any producer whose output happens to match the shape works here with no cast, no wrapper, and no transformation, whether or not it imports this package at all.
|
|
101
|
+
Because `MathBox` and its own constituent types (`MathGlyphRun`/`MathRule`/`MathStroke`/`MathAssembledGlyphs`/`MathColor`) are plain, structurally-typed data — not a class, not branded — any producer whose output happens to match the shape works here with no cast, no wrapper, and no transformation, whether or not it imports this package at all.
|
|
102
|
+
|
|
103
|
+
Sizing a stretchy glyph — a parenthesis or brace tall enough to wrap a big fraction, a radical sign sized to its own radicand, an over-brace as wide as the content under it — is the OpenType `MATH` table's `MathVariants` job, and `loadMathFont()` exposes it directly. `stretchGlyph` takes a target extent in points and returns the glyph(s) to draw with each one's own offset along the stretch axis, already in points at the requested font size:
|
|
104
|
+
|
|
105
|
+
```ts
|
|
106
|
+
import { loadMathFont } from 'pdf-codec';
|
|
107
|
+
|
|
108
|
+
const { stretchGlyph } = loadMathFont();
|
|
109
|
+
const paren = stretchGlyph(0x28, 'vertical', 40, 12); // a '(' stretched to 40pt, set at 12pt
|
|
110
|
+
// paren.kind -> 'assembly' (no single pre-built variant reaches 40pt)
|
|
111
|
+
// paren.size -> the extent actually achieved, >= 40 whenever the font can reach it
|
|
112
|
+
// paren.placements -> [{ glyphId, offset, advance }, ...], bottom to top, seams already overlapped
|
|
113
|
+
```
|
|
114
|
+
|
|
115
|
+
Three outcomes are possible and `kind` records which happened: `'base'` when the unstretched glyph was already big enough, `'variant'` when one of the font's own pre-drawn larger glyphs was selected (always preferred — a hand-drawn glyph beats a glued-together one), and `'assembly'` when the construction was genuinely built by repeating the font's own extender piece between its end pieces, overlapping each seam by as much as both sides' declared connector lengths allow so the joined outlines meet cleanly. `offset` is measured along the stretch axis from the construction's own start — the bottom for a vertical construction, the left for a horizontal one. For a caller working in design units rather than points, `MathFont.stretchyConstruction(codePoint, axis)` returns the raw parsed `MathVariants` data and `assembleStretchyGlyph` performs the same computation over it.
|
|
116
|
+
|
|
117
|
+
A stretched construction is genuinely drawable, through `MathBox`'s own `MathAssembledGlyphs` item kind: a list of `{ glyphId, xPt, yPt }` placements addressed by **glyph ID**, not by Unicode text. That distinction is the whole point rather than a convenience — most of the glyphs a `MathVariants` construction names have no Unicode code point at all, so they could never travel through `MathGlyphRun.text`, which the writer resolves through the font's `cmap`. Every pre-built larger variant is unencoded, and so are the radical's and the over-brace's assembly pieces; the bracket family is the one exception, since Unicode gives its pieces dedicated code points (the U+239B–U+23AD block — which is also an independent confirmation that assembly parts really are listed bottom-first, given `LEFT PARENTHESIS LOWER HOOK` comes first and `UPPER HOOK` last). Drawing by glyph ID works regardless, because the composite font this package embeds is Identity-H with CID == GID (see [Gotchas](#gotchas-and-quirks)), so a bare glyph ID is directly showable with no `cmap` involvement.
|
|
118
|
+
|
|
119
|
+
The one thing a glyph ID cannot carry is meaning: an unencoded glyph gets no ToUnicode entry, so `MathAssembledGlyphs` also carries the operator's own original `text` (`"("`, `"["`), which `math-content-write.ts` emits as an `/ActualText` marked-content span around the whole construction. A tall assembled bracket therefore still extracts, searches, and copies as `(`.
|
|
120
|
+
|
|
121
|
+
`MathFontMetrics.stretch` is the layout-facing form of all this, and the one a layout engine actually calls: it resolves a construction at a target size and additionally **measures** it — `inkAscentPt`/`inkDescentPt` are the whole construction's real ink extent about its drawing origin, taken from actual glyph outlines. Without that a caller cannot place the result, since a construction's ink neither starts at its drawing origin (a large parenthesis variant straddles the baseline) nor is bounded by its advance-derived `size`.
|
|
102
122
|
|
|
103
123
|
Building a layout engine on top of this codec (this is what `documents.js`'s own `src/layout/` does for docx/pptx/odt/odp/ods/odg): `TextMeasurer`/`createStandardFontMeasurer` and `wrapRunsToWidth` answer "how wide does this text render, and where does this line break" against the standard-14 metrics; `resolveStandardFont`/`STANDARD_METRICS` map an arbitrary requested family/weight/style onto one of the 14 standard PDF faces and that face's own AFM-derived metrics; `rotatePointAboutCenter` handles shape-rotation placement math. None of these do any PDF I/O themselves — they're the same primitives `writePdf`/`readPdf` use internally, exported so a caller assembling its own `LayoutDocument` (from any source format) can measure and wrap text identically to how this package will actually render it.
|
|
104
124
|
|
|
@@ -121,7 +141,21 @@ Every text run whose family resolves to a real face is subsetted to the glyphs t
|
|
|
121
141
|
|
|
122
142
|
`createFontMeasurer`'s second argument carries `verticalMetrics`, a `VerticalMetricPolicy` of `'hhea'` (the default) / `'os2Typo'` / `'os2Win'`, deciding which of the three competing ascent/descent/line-gap sets an sfnt declares should drive line height for an embedded face. It is an explicit, overridable option rather than a baked-in rule because no specification settles the question — see `src/measure.ts` for what each policy reads and why `'hhea'` is the default.
|
|
123
143
|
|
|
124
|
-
|
|
144
|
+
Reading a JPEG 2000 image directly, either as pixels or as metadata alone:
|
|
145
|
+
|
|
146
|
+
```ts
|
|
147
|
+
import { decodeJpeg2000, readJpeg2000Metadata } from 'pdf-codec';
|
|
148
|
+
|
|
149
|
+
// Works on any conforming codestream -- including one whose pixels this decoder refuses, which is what `decodable`/`undecodableReason` are for.
|
|
150
|
+
const metadata = readJpeg2000Metadata(jp2OrCodestreamBytes);
|
|
151
|
+
console.log(metadata.width, metadata.height, metadata.transform, metadata.layers, metadata.decodable, metadata.undecodableReason);
|
|
152
|
+
|
|
153
|
+
const image = decodeJpeg2000(jp2OrCodestreamBytes); // -> { width, height, bitDepth, components: Int32Array[] }, one plane per component
|
|
154
|
+
```
|
|
155
|
+
|
|
156
|
+
Both accept a whole JP2 file or a bare codestream; `decodeJpeg2000` takes an optional `onWarning` for the recoverable cases (a truncated codestream decodes to whatever packets did arrive) and throws `Jpeg2000UnsupportedError` naming the feature for anything outside [JPEG 2000 scope](#jpeg-2000-scope). Inside a PDF none of this needs calling: `readPdf` decodes a `/JPXDecode` image XObject through the same path automatically.
|
|
157
|
+
|
|
158
|
+
The full `src/bytes/`/`src/image/` surface is exported too — `crc32`, `deflate`/`inflate`/`inflateTolerant`, `ByteReader`/`ByteWriter`/`concatBytes`, `readJpegInfo`, `decodePng`/`encodePng`, `unfilterScanlines`/`filterScanlines`, `decodeCcittFax`, `decodeJbig2Embedded`, `decodeJpeg2000`/`readJpeg2000Metadata`/`parseJp2Container` — generic byte- and image-container primitives with zero PDF-specific knowledge of their own, useful independently of anything PDF-related.
|
|
125
159
|
|
|
126
160
|
Every module under `src/` is also deep-importable directly by its own subpath, for a caller that wants one internal module (not part of the curated barrel above) without pulling in the rest:
|
|
127
161
|
|
|
@@ -137,17 +171,17 @@ This works via package.json's `"./*"` wildcard export, resolving any `pdf-codec/
|
|
|
137
171
|
The package is layered from generic primitives outward to the codec itself:
|
|
138
172
|
|
|
139
173
|
- **`src/math-types.ts`** and **`src/formula.ts`** — a local, structurally-compatible mirror of `documents.js`'s own `src/mathml/layout-types.ts` + `src/mathml/metrics.ts` (`MathColor`/`MathGlyphRun`/`MathRule`/`MathStroke`/`MathLayoutItem`/`MathBox`/`MathGlyphMetrics`/`MathFontMetrics`, including the latter's own `glyph()` method signature) and `src/model/formula.ts`'s `PositionedFormula`. Deliberately not imported from `documents.js` — that would be a circular dependency once `documents.js` depends on this package for `readPdf`/`writePdf` — the same "mirror the shape, don't import the package" trick this whole family already uses elsewhere (`odf.js`'s own `MathMlNode` mirrors `ooxml.js`'s `XmlNode` rather than importing it). Because every one of these types is plain data (only `MathFontMetrics` carries a method), a real `MathBox` value `documents.js`'s own MathML layout engine produces passes into `writePdf({ formulas })` with zero cast, zero wrapper, and zero transformation.
|
|
140
|
-
- **`src/bytes/`** and **`src/image/`** — generic byte and image-container primitives with zero PDF-specific knowledge: a chunked byte writer, a backtracking byte reader, CRC32, and a hand-written PNG decoder/encoder (palette/gray/RGB/alpha, multi-`IDAT` files, all five scanline filters) plus JPEG marker scanning for dimensions only — JPEG's compressed bytes pass through completely unchanged in both directions — and a hand-written CCITT Group 3/Group 4 fax decoder (ITU-T T.4/T.6), which knows nothing of PDF: `src/filters.ts` reads the `/CCITTFaxDecode` parameter dictionary and hands it plain options, and the identical bitstreams are what TIFF's own Group3Options/Group4Options describe. `src/bytes/flate.ts` is the only file that imports `fflate`.
|
|
174
|
+
- **`src/bytes/`** and **`src/image/`** — generic byte and image-container primitives with zero PDF-specific knowledge: a chunked byte writer, a backtracking byte reader, CRC32, and a hand-written PNG decoder/encoder (palette/gray/RGB/alpha, multi-`IDAT` files, all five scanline filters) plus JPEG marker scanning for dimensions only — JPEG's compressed bytes pass through completely unchanged in both directions — and a hand-written CCITT Group 3/Group 4 fax decoder (ITU-T T.4/T.6), which knows nothing of PDF: `src/filters.ts` reads the `/CCITTFaxDecode` parameter dictionary and hands it plain options, and the identical bitstreams are what TIFF's own Group3Options/Group4Options describe. `src/image/jbig2*.ts` is a hand-written JBIG2 decoder (ITU-T T.88) built to the same rule — `jbig2-arith.ts` (the MQ arithmetic decoder of Annex E plus the integer and symbol-ID procedures of Annex A), `jbig2-bitmap.ts` (the bi-level bitmap and the five region composition operators), `jbig2-generic.ts` (generic and refinement region decoding), `jbig2-text.ts` (symbol dictionaries and text regions), and `jbig2.ts` (segment framing and page composition) — with `src/filters.ts` again owning every piece of PDF knowledge involved, namely resolving `/JBIG2Globals` and inverting JBIG2's own black-is-1 polarity to what a 1-bit `/DeviceGray` image expects. `src/image/jp2-boxes.ts` and `src/image/jpeg2000*.ts` are a hand-written JPEG 2000 decoder (ISO/IEC 15444-1 / ITU-T T.800) built the same way — `jp2-boxes.ts` (the JP2 file format's box structure of Annex I), `jpeg2000-codestream.ts` (the marker segments of Annex A), `jpeg2000-tagtree.ts` (the stuffed-bit packet-header reader and the tag tree of B.10), `jpeg2000-t2.ts` (the tile/resolution/precinct/code-block geometry of Annex B and packet decoding), `jpeg2000-t1.ts` (the EBCOT bit-plane coding passes of Annex D, driving the same MQ decoder `jbig2-arith.ts` already owns, since T.800 Annex C and T.88 Annex E specify one identical coder), `jpeg2000-dwt.ts` (the inverse wavelet of Annex F), and `jpeg2000.ts` (the whole pipeline plus the component transform and DC level shift of Annex G) — with `src/images-read.ts` rather than `src/filters.ts` owning the PDF knowledge this time, because a JPEG 2000 codestream's component count and sample depth come from the codestream rather than from the image dictionary and only the image layer has anywhere to put them. `src/bytes/flate.ts` is the only file that imports `fflate`.
|
|
141
175
|
- **`src/util/`** — two small, independently-duplicated copies of logic that lives elsewhere in the family for a reason narrow enough not to warrant a shared dependency: `base64.ts` (isomorphic base64 ⇄ `Uint8Array`, a verbatim copy of `odf.js`'s own `src/util/base64.ts`, replacing a dependency this codec used to have on `ooxml.js` purely for this one helper pair) and `abort.ts` (`throwIfAborted`, a four-line signal-check helper called at every page loop boundary in `write.ts`/`read.ts` — there is no `await` point in this package's synchronous reader/writer pipeline for cancellation to hook into implicitly, so every long-running loop checks explicitly instead; a duplicate of `documents.js`'s own `src/ports/abort.ts`, which stays there since other, non-PDF consumers still depend on it in that repository).
|
|
142
176
|
- **`src/crypto/`** — MD5, SHA-256/384/512, RC4, and AES-CBC, hand-written with zero local imports, exactly like `src/bytes/`. Not a preference: ISO 32000-1's own key-derivation algorithms name MD5 and RC4 directly, neither of which any platform crypto API this package could portably reach still offers, and `crypto.subtle` is asynchronous where this codec's read path is synchronous end to end. Reaching for `node:crypto` would put a Node builtin inside a `src/` tree that deliberately has none and is built with `platform: 'neutral'`, breaking the browser bundle its downstream consumer depends on. Each module cites the specification it implements (RFC 1321, FIPS 180-4, FIPS 197) and is tested against that specification's own published conformance vectors. MD5 and RC4 are cryptographically broken and are here solely to *read* files that already exist and whose format mandates them.
|
|
143
177
|
- **The codec itself, importing only `math-types`/`formula`/`bytes`/`image`/`crypto` (no OOXML or ODF knowledge at all):**
|
|
144
178
|
- **Write**: `objects.ts` (the `PdfObject` discriminated union), `afm-widths.ts`/`encoding.ts`/`winansi.ts`/`fonts.ts` (standard-14 metrics, WinAnsi encoding, family resolution), `font-registry.ts` (the source-document → caller-supplied → vendored-substitute → standard-14 resolution port sitting in front of `resolveStandardFont`, plus `resolveFaceWithRegistry`, the one step both the measurer and the writer resolve through so they can never disagree about which face a `LayoutFont` means), `measure.ts`/`text-layout.ts` (greedy line-wrapping, measured against either a standard-14 face's AFM widths plus a per-family correction or a resolved face's own real `hmtx` advances — never both, see [Fidelity](#fidelity)), `matrix.ts`, `content-write.ts` (`LayoutItem[]` → content-stream operators, with text branching on whether its font resolved to a standard-14 face shown as a WinAnsi byte string or an embedded one shown as Identity-H 2-byte CIDs), `write.ts` (the full object graph, classic cross-reference table, trailer, and — when `WritePdfOptions.formulas` is non-empty — one embedded math composite font group, plus one embedded text font group per face `WritePdfOptions.fonts` resolved, each allocated once in a fixed sorted order and shared across pages).
|
|
145
|
-
- **sfnt font tables**: `sfnt.ts` (a bounds-checked sfnt table-directory reader, big-endian primitive readers, and the `hasBytes` range check every table parser below pre-checks with), `cmap-table.ts` (Unicode → glyph ID, formats 4/12/6), `hmtx-table.ts` (per-glyph advance widths), `font-tables.ts` (`head`/`maxp`/`OS/2`/`post`/`name` — design grid, glyph count, vertical metrics and style bits, italic angle and underline geometry, PostScript and family names), `glyf.ts` (the `loca` offset index, per-glyph headers,
|
|
179
|
+
- **sfnt font tables**: `sfnt.ts` (a bounds-checked sfnt table-directory reader, big-endian primitive readers, and the `hasBytes` range check every table parser below pre-checks with), `cmap-table.ts` (Unicode → glyph ID, formats 4/12/6), `hmtx-table.ts` (per-glyph advance widths), `font-tables.ts` (`head`/`maxp`/`OS/2`/`post`/`name` — design grid, glyph count, vertical metrics and style bits, italic angle and underline geometry, PostScript and family names), `glyf.ts` (the `loca` offset index, per-glyph headers, a composite glyph's own component records — a composite refers to its base letter and combining marks by glyph ID, and those references nest, so subsetting one safely means taking the transitive closure over this walk — and `glyphInkBounds`, a glyph's own tight ink box, read straight out of a simple glyph's header where the format already states it and unioned from a composite's own transformed, placed components where it does not), and `math-table.ts` (the OpenType `MATH` table's constants, glyph-info, and variants subtables). Every one of these parsers degrades to `undefined` on a missing or truncated table rather than throwing: the vendored fonts are trusted, but a font extracted from an arbitrary source document is not.
|
|
146
180
|
- **sfnt subsetting**: `sfnt-subset.ts` (a TrueType-outline glyph subsetter: Unicode code points → glyph IDs through `cmap-table.ts`, the transitive closure over `glyf.ts`'s composite-component walk, then a rebuilt sfnt container carrying only those glyphs' outlines). **Glyph IDs are preserved, never renumbered** — an unused ID below the highest used one survives as an empty `loca` entry rather than being squeezed out — which keeps every composite's own component references correct inside bytes copied verbatim, keeps a caller's already-resolved glyph IDs valid against the subset, and makes CID == GID trivially true for the embedded program (a `/CIDFontType2` then needs only `/CIDToGIDMap /Identity`). The output rebuilds `head`/`hhea`/`maxp`/`loca`/`glyf`/`hmtx` (with `indexToLocFormat` forced long, one always-legal code path), copies the hinting programs (`cvt `/`fpgm`/`prep`) verbatim, stubs `post` as a version 3.0 "no glyph names" header, and omits `cmap`/`name`/`OS/2`/`GSUB`/`GPOS`/`kern` — none of which an embedded `CIDFontType2` program is read through (ISO 32000-1 9.9). The honest cost of preserving IDs: `loca` and `hmtx` stay proportional to the highest used glyph ID rather than to the number of glyphs kept, so for a document touching one glyph near the end of a large font's glyph order, most of the (already much smaller) output is those two index tables rather than outline data. Applies to `glyf`-flavoured fonts only; a CFF-flavoured one returns `undefined`, the same scope boundary [Fidelity](#fidelity) states for the embedded math font.
|
|
147
|
-
- **Embedded math font**: `math-font.ts` (parses and caches the vendored STIX Two Math font once per process, exposing a size-specific `MathFontMetrics` implementation), `math-font-write.ts` (builds the `/Type0`/`/CIDFontType0`/`/FontDescriptor`/`/FontFile3`/ToUnicode object group), `math-content-write.ts` (a `PositionedFormula[]` → PDF content-stream bytes, Identity-H 2-byte CIDs for text-showing, `re`/`m`/`l` operators for rules and the radical hook). See [Fidelity](#fidelity) for the CFF-full-embed (not glyph-subsetted) simplification this makes.
|
|
181
|
+
- **Embedded math font**: `math-font.ts` (parses and caches the vendored STIX Two Math font once per process, exposing a size-specific `MathFontMetrics` implementation and a points-in/points-out stretchy-glyph entry point), `math-stretch.ts` (the OpenType MATH two-stage stretching model over that parsed `MathVariants` data: pick the smallest pre-built variant that reaches the target, else assemble from repeated parts with every seam overlapped inside both sides' own declared connector lengths -- unit-agnostic, so it works in design units or points alike), `math-font-write.ts` (builds the `/Type0`/`/CIDFontType0`/`/FontDescriptor`/`/FontFile3`/ToUnicode object group), `math-content-write.ts` (a `PositionedFormula[]` → PDF content-stream bytes, Identity-H 2-byte CIDs for text-showing, `re`/`m`/`l` operators for rules and the radical hook, and one glyph-ID-addressed text object per placement of a stretched construction, wrapped in an `/ActualText` marked-content span). See [Fidelity](#fidelity) for the CFF-full-embed (not glyph-subsetted) simplification this makes.
|
|
148
182
|
- **Embedded text faces**: `embedded-font.ts` (parses and caches one TrueType-outline face's `cmap`/`hmtx`/`hhea`/`head`/`OS/2`/`post`/`name` metrics, and `encodeForShowEmbedded`, the single code path both measurement and text-showing go through — the same reason `winansi.ts`'s `encodeForShow` exists, since encoding and measuring separately lets the two disagree about which characters resolved to which glyph and silently desyncs a computed wrap point from the drawn line). Every geometry field it exposes is converted into PDF's 1000-units-per-em glyph space (ISO 32000-1 9.8.1), which a font's own design grid frequently is not: STIX Two Math is drawn on a 1000-unit em so that factor is an identity for the math font, but Carlito is drawn on a 2048-unit em, where getting the conversion wrong is silent rather than loud — the font simply renders with roughly twice its intended metrics and nothing anywhere reports an error. The serif flag a `/FontDescriptor` needs is read off the face's own PANOSE classification rather than guessed from its family name. `embedded-font-write.ts` builds the `/Type0`/`/CIDFontType2`/`/FontDescriptor`/`/FontFile2`/ToUnicode object group, with `/CIDToGIDMap /Identity` written explicitly (stating outright the CID == GID invariant `sfnt-subset.ts`'s GID-preserving design guarantees) and `/Length1` set to the **uncompressed** subset length — the single most commonly mis-set key in TrueType embedding, since the obvious-looking value, the stream's own `/Length`, is silently accepted by lenient readers and rejected by strict ones. Its subset tag (`ABCDEF+Carlito-Regular`, six uppercase letters per 9.6.4) is a CRC32 over the face's PostScript name and its exact ascending glyph-ID list, so identical input yields byte-identical output and two subsets of one face carrying different glyphs can never be mistaken for one another.
|
|
149
|
-
- **ToUnicode CMaps**: `tounicode.ts`, shared by both embedded-font writers above rather than duplicated in each — a character code → Unicode code point mapping written as a bfchar CMap (9.10.3), with supplementary-plane code points encoded as genuine UTF-16BE surrogate pairs and entries emitted in blocks of at most 100, the limit the CMap syntax sets and one a subsetted text face routinely exceeds.
|
|
150
|
-
- **CFF
|
|
183
|
+
- **ToUnicode CMaps**: `tounicode.ts`, shared by both embedded-font writers above rather than duplicated in each — a character code → Unicode code point mapping written as a bfchar CMap (9.10.3), with supplementary-plane code points encoded as genuine UTF-16BE surrogate pairs and entries emitted in blocks of at most 100, the limit the CMap syntax sets and one a subsetted text face routinely exceeds. A glyph with no code point to map back to (a stretchy construction's own unencoded pieces) is dropped from the CMap rather than mapped to a stand-in that would extract as the wrong character; the `/ActualText` span around such a construction carries its real text instead.
|
|
184
|
+
- **CFF reading**: `cff.ts` holds the two container structures every CFF program is built out of — the `INDEX` and the `DICT` — shared by the two readers built on it rather than hand-rolled twice. `cff-bounds.ts` is a Type 2 charstring interpreter that computes each glyph's own tight ink bounding box: a CFF glyph, unlike a TrueType one, stores no bounding box anywhere, so the only way to know what area it covers is to run its outline program. It is a path *walker*, not a rasteriser — it tracks the current point through every path-construction operator and solves each cubic's real extrema from the roots of its own derivative, so a bound it reports is genuinely tight rather than a control-point hull. Hint operators are decoded only far enough to know how many bytes a following `hintmask` consumes. Verified against the vendored STIX Two Math font's whole 5,543-glyph repertoire: every glyph matches fontTools' own `BoundsPen` to within 0.01 design units, and the union of all 5,543 computed boxes lands exactly on the font's own `head` table `FontBBox`. Out of scope, each reported as `undefined` rather than guessed at: a CID-keyed program (its local subroutines live per-FD behind `FDArray`/`FDSelect`), `endchar` in its four-argument seac-like form (which needs the charset and Standard Encoding to resolve into two other glyphs), and the arithmetic/storage/conditional escaped operators. `cff-probe.ts` reads a bare CFF program's header, Name INDEX, and just enough of its Top DICT to detect the `ROS` operator (the escaped `12 30` whose presence *is* the definition of a CID-keyed font, CFF 1.0 Appendix H). It exists to make the embedding path **refuse** such a font rather than mis-embed it: a CID-keyed CFF carries its own charset mapping CIDs onto glyph indices, so CID == GID does not hold, and showing text through Identity-H against one anyway produces no error anywhere — the file is structurally valid, every reader accepts it, and the page simply renders the wrong glyphs. Not wired into any write path yet; it is the guard a later source-embedded-font phase needs. It correctly reports the vendored STIX Two Math font's own `CFF ` table as *not* CID-keyed, which is what makes `math-font-write.ts`'s existing embedding sound.
|
|
151
185
|
- **Read**: `lexer.ts`/`parse.ts` (byte tokenizer and tokens → `PdfObject`), `filters.ts`/`predictors.ts` (Flate/LZW/ASCII85/ASCIIHex/RunLength/CCITTFax, TIFF/PNG predictors), `xref.ts`/`document.ts` (classic and cross-reference-stream resolution, object streams, `/Prev` chains, linear-scan recovery, the page tree with attribute inheritance), `encrypt.ts` (the standard security handler: `/Encrypt` parsing, empty-user-password key derivation and `/U` verification, per-object keys, and the string/stream decryption `document.ts` applies transparently as each indirect object is fetched — see [Gotchas](#gotchas-and-quirks)), `content-read.ts`/`interpret.ts` (the content-stream tokenizer and graphics/text state machine, including form-XObject recursion and general vector-path tracking — see [Gotchas](#gotchas-and-quirks)), `cmap.ts`/`font-style.ts`/`font-read.ts` (`/ToUnicode` CMaps, font-dictionary resolution), `images-read.ts` (Image XObjects → PNG/JPEG bytes), `read.ts` (`readPdf`, assembling all of the above into a `LayoutDocument`).
|
|
152
186
|
- `codec.ts` — `pdfCodec`, a `z.codec()` pair over `readPdf`/`writePdf`, plus a standalone, ~20-line local copy of just the `%PDF-` header check `PdfBytesSchema` needs (`documents.js`'s own equivalent schema lives in a file that also carries unrelated docx/pptx/odt schemas that have no place here).
|
|
153
187
|
- **`src/test-support/pdf.ts`** — hand-built PDF fixtures for the parser's own tests (a classic-xref file, an xref-stream-with-object-streams file, a broken-`startxref` file needing linear-scan recovery, an incremental update, a trailer naming a security handler no password could open, and more), built by literal byte/string concatenation and deliberately importing NOTHING from this package's own writer — a fixture built by calling `writePdf` would let a writer bug hide from the corresponding reader test and vice versa. Not part of the public surface; test-only.
|
|
@@ -171,18 +205,47 @@ Dependency direction is strictly downward and checkable: `math-types`/`formula`/
|
|
|
171
205
|
- **Reading arbitrary real-world PDFs is the single largest risk surface in this package**, and the parser is honest about its design target: cleanly-generated output from mainstream producers (Word, PowerPoint, Chrome, LibreOffice, Acrobat), recovering from the malformations those producers and their downstream tooling actually create, and failing loudly and specifically on anything else — not matching a mature library's robustness against adversarial input.
|
|
172
206
|
- **An encrypted PDF is readable when, and only when, it opens without a password.** That is the overwhelmingly common real-world case — a permissions-only file, exported with "no printing" or "no copying" set, whose owner password may well be set but whose *user* password is empty. `readPdf` derives the file key from the empty user password, verifies it against the `/Encrypt` dictionary's own `/U` entry, and decrypts every string and stream transparently; nothing downstream of the object store knows the file was encrypted at all. Supported: `/Filter /Standard` at `/V` 1, 2, 4, and 5 — RC4-40, RC4-128, AES-128 (`/CFM /AESV2`) and AES-256 (`/CFM /AESV3`) — including `/EncryptMetadata false` and `/Identity` crypt filters. A file that genuinely needs a user password throws `PdfPasswordRequiredError`, its own distinct error, because "supply the password" and "this codec cannot read this at all" are different things to tell a user and only one of them can be acted on. Anything else — a non-standard (public-key) security handler, the unpublished `/V 3` algorithm, an unrecognised `/CFM` — still throws `PdfEncryptedError`.
|
|
173
207
|
- **Nothing in this codec accepts, prompts for, or guesses a password.** There is no password parameter on `readPdf`, and no owner-password path: authenticating as owner is a permissions escalation, not a way to read a file you were already allowed to read. Decryption is here to open files that are already open, not to get into ones that are not.
|
|
174
|
-
- **`CCITTFaxDecode`
|
|
208
|
+
- **`CCITTFaxDecode`, `JBIG2Decode` and `JPXDecode` images all decode for real.** `src/image/ccitt.ts` is a hand-written ITU-T T.4/T.6 fax decoder — Group 4 (`/K < 0`, the modern default and the overwhelming majority of scanned-PDF usage), Group 3 one-dimensional (`/K = 0`), and Group 3 mixed (`/K > 0`) — producing a real packed 1-bit-per-pixel bitmap that the rest of `images-read.ts` then treats exactly like any other 1-bit `/DeviceGray` raster, `/Decode` inversion and all. `/Columns`, `/Rows` (falling back to the image's own `/Height`), `/BlackIs1`, and `/EncodedByteAlign` are honoured; the uncompressed-mode extension code is not, and a stream that stops making sense degrades to the rows already recovered plus a `pdf/ccitt-fax-degraded` diagnostic rather than throwing. `src/image/jbig2*.ts` is a hand-written ITU-T T.88 decoder covering what real scanned PDFs actually contain — see [JBIG2 scope](#jbig2-scope) below for exactly what is and is not implemented. `src/image/jpeg2000*.ts` is a hand-written ISO/IEC 15444-1 decoder covering both wavelets and the tiling, layering and precinct structure real files use — see [JPEG 2000 scope](#jpeg-2000-scope) below for exactly what is and is not implemented, and for the parts that are refused by name rather than approximated. JPEG images (`DCTDecode`) pass through completely losslessly in both directions; PNG-sourced images go through a real, narrowly-scoped hand-written codec.
|
|
175
209
|
- **`interpret.ts` tracks general vector paths, not just axis-aligned `re` rectangles.** `m`/`l`/`c`/`v`/`y`/`h` (and `re` itself, per its own ISO 32000-1 definition as a 4-point rectangle subpath) accumulate real subpaths — CTM-transformed line/cubic segments, open or closed — and any paint operator (`f`/`F`/`f*`/`S`/`s`/`B`/`B*`/`b`/`b*`) emits an item built from them. Verified both by dedicated tests and by a genuine `writePath` → `writePdf` → `readPdf` round trip recovering the original `LayoutPath` value exactly. This is the shared infrastructure a caller reconstructing structure from a `LayoutDocument` (a spreadsheet grid, a vector drawing) builds on.
|
|
176
210
|
- **A recovered path that matches one of three characteristic shape patterns comes back as that shape's own kind, not as a generic `LayoutPath`.** PDF has exactly one shape operator, `re`, and no ellipse or line operator at all, so a writer has no way to record what a path *was* — `interpret.ts` recovers it from the geometry instead. A single closed four-corner subpath whose every edge runs along one axis is a `LayoutRect` (so `re` under a non-rotated *or* 90°-rotated CTM, a hand-built `m`/`l`/`l`/`l`/`h` rectangle, and any combination of fill and stroke all reach it — not just the fill-only single-`re` case an earlier fast path covered); a closed subpath of exactly four cubic segments meeting at its bounding box's four cardinal points, with all eight control points at the standard kappa offset (`BEZIER_KAPPA`, 4/3·(√2−1)) those extremes imply, is a `LayoutEllipse` — precisely what `writeEllipse` emits; an open single-straight-segment stroke-only subpath is a `LayoutLine`. Tolerance is `max(1e-3pt, 1e-4 × the shape's own extent)`: the absolute floor is twenty times the 5e-5pt quantisation `formatNumber`'s 4-decimal-place rounding imposes, and the relative term is what lets a large ellipse from a producer that rounded its kappa constant more coarsely (`0.5523`) still match. **These are deliberate, bounded heuristics, not certainties** — a hand-authored freeform path that happens to consist of four kappa-ratio cubics between its bounding box's cardinal points is indistinguishable from a "real" ellipse in the PDF bytes, because at that point it geometrically *is* one, whatever the author called it. What a false positive can never do is misreport geometry: every detected shape reproduces its source path's own points exactly, so it changes an item's kind, never where or how big it is. Anything the patterns don't cover — a non-90° rotation, a curve that isn't the four-quadrant construction, a polygon that isn't a rectangle, multiple subpaths — stays a `LayoutPath`, and a rotated ellipse deliberately does too, since `LayoutEllipse` carries no rotation to report one with.
|
|
177
211
|
- **`writePdf`/`readPdf` round-trip a page's own `notes` field via a hidden annotation, not any real PDF feature.** PDF has no native concept of hidden presenter notes, so a page's `LayoutPage.notes` (when present) is written as a `/Subtype /Text` annotation (the same construct Acrobat's own sticky-note tool uses) with the `Hidden` annotation flag set so it never renders or prints, and `readPdf` reads it back via an internal author marker that distinguishes this package's own notes annotation from a genuine third-party sticky note. This is a round-trip mechanism specific to this package's own writer/reader pair — a PDF produced by anything else will never carry it, and a PDF consumer other than this package's own `readPdf` will never see it as anything but an invisible, empty sticky note. `documents.js` uses this to carry pptx/odp speaker notes through `pptxToPdf`/`pdfToPptx` and `odpToPdf`/`pdfToOdp`.
|
|
178
|
-
- **STIX Two Math (the embedded formula font) is a CFF-flavoured OpenType font (an `OTTO` sfnt wrapping a `CFF ` table), not TrueType/glyf** — confirmed by inspecting the vendored font's own sfnt table directory while `math-font.ts` was built. Genuine Type2-charstring glyph subsetting (re-encoding charstrings, rebuilding the CFF `INDEX` structures with a renumbered, minimal glyph set) is a substantially larger undertaking than TrueType glyf/loca subsetting, and is out of scope: the **entire** `CFF ` table is embedded verbatim, unmodified, as a single `/FontFile3` `/Subtype /CIDFontType0C` stream — a real, correct, working embedded font, just not glyph-subsetted. Everything else genuinely IS built from a targeted parse of only what's used: `cmap` resolves exactly the Unicode code points a document's formulas actually reference to glyph IDs, and the emitted `/W` widths array
|
|
179
|
-
- **The OpenType `MATH` table's `MathVariants` subtable
|
|
180
|
-
-
|
|
212
|
+
- **STIX Two Math (the embedded formula font) is a CFF-flavoured OpenType font (an `OTTO` sfnt wrapping a `CFF ` table), not TrueType/glyf** — confirmed by inspecting the vendored font's own sfnt table directory while `math-font.ts` was built. Genuine Type2-charstring glyph subsetting (re-encoding charstrings, rebuilding the CFF `INDEX` structures with a renumbered, minimal glyph set) is a substantially larger undertaking than TrueType glyf/loca subsetting, and is out of scope: the **entire** `CFF ` table is embedded verbatim, unmodified, as a single `/FontFile3` `/Subtype /CIDFontType0C` stream — a real, correct, working embedded font, just not glyph-subsetted. Everything else genuinely IS built from a targeted parse of only what's used: `cmap` resolves exactly the Unicode code points a document's formulas actually reference to glyph IDs, and the emitted `/W` widths array covers only the glyph IDs actually drawn (those, plus any unencoded pieces a stretchy construction contributed), not the font's full ~5,500-glyph repertoire; the ToUnicode CMap covers the subset of those that have a code point at all. A CID-keyed composite font built this way needs no `/CIDToGIDMap` at all (that key exists only for `/CIDFontType2`): per ISO 32000-1 9.7.4.2, a `/CIDFontType0` whose `/FontFile3` is a "bare" (non-CID-keyed) CFF program is read with CID treated as directly indexing the CFF's own `CharStrings` INDEX by glyph order — i.e. CID == GID, exactly the numbering `cmap`-derived glyph IDs already use, so Identity-H text-showing needs no further remapping anywhere in the write path.
|
|
213
|
+
- **The OpenType `MATH` table's `MathVariants` subtable is parsed, its stretchy-glyph assembly implemented (`math-table.ts`/`math-stretch.ts`), and the result genuinely drawable (`math-types.ts`'s `MathAssembledGlyphs`, `math-content-write.ts`).** Variant selection and part assembly are real, tested computation — `loadMathFont().stretchGlyph(...)` gives back the glyph IDs and offsets for a parenthesis, brace, radical sign, or over-brace at any target size, with every seam overlapped inside the parts' own declared connector lengths — and `MathFontMetrics.stretch` wraps that with the real ink measurement a layout engine needs to place it. Drawing goes by **glyph ID**, because most of the glyphs a construction names have no Unicode code point at all: every pre-built larger variant is unencoded, as are the radical's and the over-brace's assembly pieces, with the bracket family the one exception (Unicode's own U+239B–U+23AD piece block covers it). That works because CID == GID here, so a glyph ID is shown directly with no `cmap` involvement. Two real consequences follow, both handled rather than left silent: `collectUsedGlyphs` now maps a glyph ID to `number | undefined`, so an unencoded glyph still gets its `/W` width but contributes no ToUnicode entry, and the construction is wrapped in an `/ActualText` marked-content span carrying the operator's own text so it still extracts as `(` rather than as nothing. `MathConstants` (every fraction/radical/script-positioning constant `MathFontMetrics` exposes) and `MathGlyphInfo` (italics correction, top-accent attachment) are genuinely parsed in full too.
|
|
214
|
+
- **What this package draws for a stretchy glyph is decided entirely by its caller.** This package resolves and draws whatever construction it is asked for, on either axis; which operators a document actually stretches, and to what, is a layout-engine decision — `documents.js`'s own `src/mathml/layout.ts` currently stretches vertical fences in an `mrow` and nothing else. See that package's own README for which constructions genuinely stretch today and which still render at a fixed size.
|
|
215
|
+
- **Real per-glyph ink bounds are now measured from the outline, but `ascentPerEm`/`descentPerEm` still exist alongside them and a caller has to choose.** `MathGlyphMetrics.inkAscentPt`/`inkDescentPt` (and `MathFont.glyphInkBounds`, the same thing in design units) carry each glyph's own tight ink extent, computed by walking its Type 2 charstring — a full stop measures 0.12 em tall against the 1.0 em the font's nominal `hhea` metrics claim for every glyph alike. `ascentPerEm`/`descentPerEm` remain what they always were: one uniform figure for the whole face, still the right measure for anything sized against the font rather than against particular characters, and still the fallback for a glyph with no outline to measure (a space) or one `cff-bounds.ts` declines to walk, where both ink fields come back `undefined` together. The consumer that motivated this — `documents.js`'s own MathML `layoutToken` — now uses them: it takes the max ink ascent and max ink descent across a run's glyphs, falling back to the nominal metrics per glyph that carries no bounds.
|
|
216
|
+
- **An ink box is genuinely tight, which for a math font means it is often *larger* than the nominal metrics, not smaller.** The "ink is a fraction of the nominal extent" intuition holds for text-like glyphs (a full stop, a parenthesis, an `x`) and fails for the extension pieces, display-size operators, and pre-built large variants a math font is full of: over a tenth of STIX Two Math's repertoire draws above its own nominal ascent, and another tenth below its nominal descent, reaching 2.6 em up and 1.6 em down at the extremes. Sizing those from `ascentPerEm` under-reports them exactly as badly as it over-reports a full stop, which is the whole reason the per-glyph measurement exists.
|
|
217
|
+
- **A glyph's ink descent is negative where its lowest ink sits above the baseline.** `inkDescentPt` follows `descentPerEm`'s own sign convention (ink below the baseline is a positive descent), so a superscript-height glyph honestly reports a negative descent rather than a clamped zero. A consumer that needs a box which never crosses the baseline clamps at its own layer, where it can see what the box is for.
|
|
181
218
|
- **An embedded `CIDFontType2` program needs no `cmap` table of its own, and `sfnt-subset.ts`'s output doesn't carry one — this is a property of the spec, not an oversight this package works around.** `cmap` maps a character code to a glyph ID for a *simple* font; a `Type0` composite font never asks the embedded font program to do that lookup at all. Character code → CID goes through the `Type0` font's own `/Encoding` (Identity-H here, so CID == character code by construction for the 2-byte codes this package writes), and CID → GID goes through `/CIDToGIDMap` (`/Identity` here, matching `sfnt-subset.ts`'s own GID-preserving design). Both steps happen inside the PDF's own object graph, entirely before the embedded font program is ever consulted — ISO 32000-1 9.7.4.2. This is exactly why the subset output can safely omit `cmap` alongside `name`/`OS/2`/`GSUB`/`GPOS`/`kern` (see Architecture above): none of the five is on the code-path a `CIDFontType2` reader actually walks.
|
|
182
219
|
- **No GPOS/GSUB/kern support: an embedded face's own kerning pairs, ligature substitutions, and other OpenType layout features are never read, subsetted, or applied.** `sfnt-subset.ts` strips `GSUB`/`GPOS`/`kern` from its output entirely, and nothing upstream of that — `measure.ts`'s line-wrapping, `content-write.ts`'s glyph placement — ever consults them either, subset or not. Every embedded run is placed glyph-by-glyph at that glyph's own bare `hmtx` advance width, with no per-pair kerning adjustment. See the acceptance-bar paragraph under [Fidelity](#fidelity) for what this bounds and doesn't bound.
|
|
183
220
|
- **`font-substitutes.ts` maps both `Calibri` and `Calibri Light` onto the same, ordinary-weight Carlito face.** Carlito ships only one weight per style axis (regular/bold/italic/bolditalic) — there is no distinct Light design to embed — so `Calibri Light` substitutes to standard Carlito rather than a genuinely lighter face. An honest, documented approximation (see that file's own top-of-file comment), not a faithful weight match: a caller relying on Calibri Light's visibly thinner strokes will not see them, only its width metrics.
|
|
184
221
|
- **`cff-probe.ts`'s CID-keyed CFF guard exists for a source-embedded-font phase this package hasn't built yet — it is not wired into any write path today, and today's embedding never needs it.** Every face this package currently embeds — the vendored Carlito/Caladea substitutes, and any caller-supplied face via `sourceFonts`/`fonts` — is `glyf`-flavoured TrueType, and `sfnt-subset.ts` already refuses (returns `undefined` for) anything that isn't before `cff-probe.ts` would ever run against it. The guard is what a future phase embedding a real, subsetted CFF program (rather than the whole-table CFF embed `math-font-write.ts` already does for STIX Two Math) will need before it can trust CID == GID against an arbitrary caller-supplied font: a CID-keyed CFF carries its own CID → glyph-index charset, so that identity does not hold for one, and nothing about the file signals the mismatch to a reader — it just renders the wrong glyphs.
|
|
185
222
|
|
|
223
|
+
## JBIG2 scope
|
|
224
|
+
|
|
225
|
+
`src/image/jbig2*.ts` is a hand-written ITU-T T.88 decoder, built to the same rule as everything else here: no external library, every layer written against the specification. What it covers is what real scanned PDFs actually contain, and the boundary is stated precisely rather than left to be discovered.
|
|
226
|
+
|
|
227
|
+
**Implemented.** The MQ arithmetic decoder (Annex E) and the arithmetic integer and symbol-ID procedures (Annex A). Generic region decoding (6.2) for all four templates, with adaptive (AT) pixels at any offset, typical prediction (TPGDON), and the MMR variant — which is a plain ITU-T T.6 bitstream, so it routes through `src/image/ccitt.ts` rather than duplicating a Group 4 decoder. Generic refinement region decoding (6.3) for both templates. Symbol dictionaries (6.5) and text regions (6.4) in their arithmetic form, covering height classes, the export-flag runs, every reference corner, transposed regions, multi-row strips, a non-zero `SBDSOFFSET`, and refined symbol instances. Segment framing (clause 7) including the long referred-to-segment form, page composition with all five combination operators, and the `/JBIG2Globals` stream a PDF uses to share one symbol dictionary across images.
|
|
228
|
+
|
|
229
|
+
**Not implemented, and each says so by name rather than guessing.** The Huffman-coded forms of symbol dictionaries and text regions (`SDHUFF`/`SBHUFF` set) and the custom Huffman table segments that go with them — a wholly separate coding path that no mainstream encoder targeting PDF emits. Halftone regions and pattern dictionaries. Intermediate regions, which are retained in an auxiliary buffer rather than composed onto the page. Segments of unknown length (7.2.7). The `EXTTEMPLATE` twelve-adaptive-pixel template of Amendment 2. A symbol dictionary that imports another's arithmetic coding contexts. Typical prediction in a *refinement* region (`TPGRON`) — see below for why that one is a refusal rather than an omission. The aggregate (`REFAGGNINST > 1`) form of refinement/aggregate symbol coding. Any of these raises `Jbig2UnsupportedError` naming itself; `src/filters.ts` turns that into a `pdf/jbig2-undecodable` diagnostic and leaves the image's bytes undecoded, so the image is skipped and the rest of the page still reads — the same degradation an unimplemented filter already got.
|
|
230
|
+
|
|
231
|
+
**How it is verified.** `src/test-support/jbig2.ts` holds real embedded streams, regenerable by `scripts/generate-jbig2-fixtures.mjs`, from three producers none of which is this package: jbig2enc (the encoder behind essentially every JBIG2-in-PDF in the wild) for the generic-region and symbol/text-region fixtures, libtiff for the MMR payload, and a hand-written T.88 Annex E arithmetic *encoder* in the generator script for the templates and coding options jbig2enc will not emit. Every stream — hand-encoded ones included — is decoded by jbig2dec (Ghostscript's independent implementation) before being written out, and the bitmap recorded as each fixture's expected output is jbig2dec's, not this package's. The symbol-mode fixtures additionally exist in six variants with only the text region's `REFCORNER`/`TRANSPOSED` bits rewritten: those two fields change nothing about the arithmetic bitstream, so the patched stream is still genuinely jbig2enc's encoding, but every symbol instance lands somewhere different — which is what turns jbig2dec's output into a real differential test of the placement rules for the corners jbig2enc itself never emits.
|
|
232
|
+
|
|
233
|
+
**What that does and does not establish, stated precisely because it is easy to overclaim.** A differential test against another decoder pins the *set* of template positions and their offsets: a decoder reading a different set of neighbours cannot track the encoder's adaptive state at all. It does **not** pin the *order* those positions are concatenated into a context index, and nothing can — a context index is only a label for a neighbourhood pattern, so any consistent permutation cancels out between an encoder and a decoder that each use their own consistently. The fixed typical-prediction pseudo-contexts are subject to the same caveat, because they share one adaptive state array with the real pattern contexts: a wrong constant still round-trips whenever it happens not to collide with a pattern the test image actually produces. The fixtures that genuinely pin a pseudo-context are therefore only the ones jbig2enc produced itself — `jbig2 -d`, which sets TPGDON — and only for GBTEMPLATE 0, the one template jbig2enc emits.
|
|
234
|
+
|
|
235
|
+
**Why TPGRON is refused rather than shipped unverified.** Refinement itself is fixture-verified — both templates, arbitrary reference offsets, and refined symbol instances inside a text region, all agreeing with jbig2dec. Typical prediction inside a refinement region is the one part that is not, and cannot be here: jbig2enc's refinement support is disabled upstream ("Refinement broke in recent releases since it's rarely used"), so the only stream available to test against is one this package encoded itself, and by the argument above that cannot pin the pseudo-context constant even in principle. Brute-forcing all 1024 ten-bit candidates for GRTEMPLATE 1 against jbig2dec made this concrete: a different, unrelated band of constants passes depending on which test image is used, which is the signature of a test measuring collision luck rather than correctness. So a refinement region that sets TPGRON raises `Jbig2UnsupportedError` naming the flag. Everything else about refinement works, and T.88 6.4.11 fixes TPGRON at 0 for the symbol-instance refinement that is where refinement actually appears in practice.
|
|
236
|
+
|
|
237
|
+
## JPEG 2000 scope
|
|
238
|
+
|
|
239
|
+
`src/image/jp2-boxes.ts` and `src/image/jpeg2000*.ts` are a hand-written ISO/IEC 15444-1 (ITU-T T.800) decoder, built to the same rule as everything else here: no external library, every layer written against the specification. The MQ arithmetic decoder is not written twice — T.800 Annex C and T.88 Annex E specify one identical coder, so `src/image/jbig2-arith.ts`'s `MqDecoder` is reused verbatim, with only JPEG 2000's own three non-zero initial context states (Table D.7) applied on top.
|
|
240
|
+
|
|
241
|
+
**Implemented.** The JP2 file format (Annex I): the box structure, the image header, enumerated and ICC colour specifications, channel definitions, and the contiguous codestream box — as well as the bare codestream a PDF `/JPXDecode` stream may carry instead (ISO 32000-1 7.4.9 permits either). The codestream syntax (Annex A): SIZ, COD, COC, QCD, QCC, POC, RGN, COM, SOT and SOD, with tile-part header overrides resolving against the main header in the precedence A.6 defines. Tier-2 packet decoding (Annex B): the stuffed-bit packet-header reader, tag trees, code-block inclusion across quality layers, zero-bit-plane signalling, the coding-pass prefix code, `Lblock` growth and segment lengths, precinct partitions at any size, SOP and EPH markers, and the tile/resolution/subband/precinct/code-block geometry those index into — at any image and tile origin on the reference grid, not only at zero. Tier-1 EBCOT (Annex D): the three coding passes over every bit-plane, the zero-coding context tables for all four subband orientations, sign coding with its XOR bit, magnitude refinement, cleanup with run-length mode, and the vertically-causal-context, reset-contexts and segmentation-symbol code-block styles. Both wavelets (Annex F): the reversible 5-3 integer lifting and the irreversible 9-7 floating-point lifting, with whole-sample symmetric extension. Dequantization (Annex E) for no-quantization, scalar-derived and scalar-expounded styles. Both component transforms and the DC level shift (Annex G). LRCP and RLCP progression in general; RPCL, PCRL and CPRL when every resolution level holds a single precinct, which is where those three collapse to a plain loop.
|
|
242
|
+
|
|
243
|
+
**Not implemented, and each says so by name rather than guessing.** Sub-sampled components (`XRsiz`/`YRsiz` other than 1), which would need resampling this package does not do. Regions of interest (RGN), whose coefficient upshift is not undone. Progression-order changes (POC). Packed packet headers, in either the main header (PPM) or a tile-part header (PPT). The selective arithmetic coding bypass ("lazy") and terminate-on-every-pass code-block styles, both of which split a code-block into segments this decoder does not read. A JP2 palette (`pclr`/`cmap`) box. A codestream mixing component bit depths or signedness. Each of these raises `Jpeg2000UnsupportedError` naming itself; `src/images-read.ts` turns that into an `image/jpx-undecodable` diagnostic and skips the image, so the rest of the page still reads. `readJpeg2000Metadata` reports the same reason ahead of time as `undecodableReason`, alongside the geometry, component, tile, wavelet, layer and quantization parameters it reads from **any** conforming codestream — including one it cannot decode the pixels of.
|
|
244
|
+
|
|
245
|
+
**How it is verified.** `src/test-support/jpeg2000.ts` holds real codestreams, regenerable by `scripts/generate-jpeg2000-fixtures.mjs`, all produced by OpenJPEG's own `opj_compress` from deterministic PGM/PPM sources. Nothing in this repository influences a byte of them. For every reversible fixture the generator first proves the configuration round-trips byte-identically through `opj_decompress`, and then records **the source image** as the expected output — not any decoder's. That makes the oracle the original integers the encoder was handed, which no shared mistake between an encoder and a decoder can fake, and the test asserts exact equality against it. The fixture set spans odd dimensions, a 1x1 image, an image smaller than one code-block, a non-zero image origin (so resolution levels start at odd coordinates), 12-bit samples, both colour-transform settings, multiple tiles, multiple quality layers, small code-blocks, subdivided precincts, four progression orders, SOP/EPH framing, three code-block styles, and a JP2 container.
|
|
246
|
+
|
|
247
|
+
**What the irreversible fixtures do and do not establish, stated precisely because it is easy to overclaim.** The 9-7 wavelet is lossy by construction, so there is no exact answer to reproduce and no oracle of the kind above: their expected samples are `opj_decompress`'s own output, and the test asserts every sample within one of it with under 1% differing, rather than equality. In practice the observed disagreement is a handful of samples per image, all by exactly one, which is the signature of floating-point rounding at a round-to-nearest boundary (OpenJPEG carries a slightly truncated normalisation constant where this package uses the specification's own) rather than of a decoding difference. That is real evidence — a wrong context label, a wrong subband gain or a misplaced `K` lands orders of magnitude outside a bound like this — but it pins this decoder against OpenJPEG's arithmetic, where the reversible fixtures pin it against the specification absolutely. The `inverseDwt97Level` unit test adds one specification-side check the fixtures cannot: the 9-7 analysis filter maps a constant signal onto a constant low-pass band and an identically zero high-pass band, so the synthesis has to send that straight back, which fails for any wrong lifting constant, step order or `K` placement.
|
|
248
|
+
|
|
186
249
|
## Fidelity
|
|
187
250
|
|
|
188
251
|
**Ordinary text in PDF output uses the standard 14 fonts only, unless a caller supplies `WritePdfOptions.fonts`.** Without a registry, Helvetica/Times-Roman are metric-compatible substitutes for Arial/Times New Roman, but a modern default like Calibri, Cambria, or Aptos is not, so a caller's own line wrapping and pagination (built against this package's `TextMeasurer`) will drift slightly from what the original authoring application would itself produce. `measure.ts` narrows that gap with a small per-family width-correction table — Calibri measures 8% narrower than Helvetica, Verdana 9% wider, and so on — and `content-write.ts` draws the glyphs at the matching `Tz` horizontal scale so the measurement and the drawing agree, but it remains a stretched standard-14 face rather than the real one.
|
|
@@ -191,7 +254,7 @@ Dependency direction is strictly downward and checkable: `math-types`/`formula`/
|
|
|
191
254
|
|
|
192
255
|
**The acceptance bar for embedded-font fidelity is deliberately "no page-count drift on a real corpus", not "line-identical".** Every glyph in an embedded run is placed at its own bare `hmtx` advance width and nothing else — there is no pair-kerning adjustment anywhere in this package's measurement or content-stream writing path (see [Gotchas](#gotchas-and-quirks)). Reaching genuine line-identical parity with the original authoring application would additionally need a GPOS pair-kerning reader on top of the outline subsetter this package already has, a materially larger, separate undertaking this phase deliberately did not attempt. What the current bar does guarantee: a real document measured and drawn through the same resolved face's own advances will not silently reflow onto a different number of pages the way a width-corrected standard-14 substitute occasionally can.
|
|
193
256
|
|
|
194
|
-
**The one exception is math-formula rendering (`WritePdfOptions.formulas`): this genuinely embeds a real, hand-parsed font.** Real box-model glyph runs are shown through the embedded STIX Two Math font with genuine per-glyph metrics (advance width, italic correction, top-accent attachment) and font-wide layout constants (axis height, fraction/radical rule thickness and gaps, script shift amounts) parsed directly from that font's own `MATH` table — not approximated or hand-tuned. See [Gotchas](#gotchas-and-quirks) for the exact boundary of what this package's own font parsing does and doesn't cover (the CFF-full-embed simplification,
|
|
257
|
+
**The one exception is math-formula rendering (`WritePdfOptions.formulas`): this genuinely embeds a real, hand-parsed font.** Real box-model glyph runs are shown through the embedded STIX Two Math font with genuine per-glyph metrics (advance width, italic correction, top-accent attachment) and font-wide layout constants (axis height, fraction/radical rule thickness and gaps, script shift amounts) parsed directly from that font's own `MATH` table — not approximated or hand-tuned. Stretchy constructions are real too: a `MathVariants` variant or part assembly is resolved, measured against actual glyph outlines, and drawn by glyph ID. See [Gotchas](#gotchas-and-quirks) for the exact boundary of what this package's own font parsing does and doesn't cover (the CFF-full-embed simplification, and what a stretched construction costs in ToUnicode terms).
|
|
195
258
|
|
|
196
259
|
**`readPdf(writePdf(doc))` is not guaranteed to reproduce `doc` exactly, and `writePdf(readPdf(bytes))` is not guaranteed to reproduce `bytes` exactly — this package makes no round-trip-losslessness claim in either direction.** A PDF page is fundamentally a stream of positioned drawing operators, not a structured document: a rectangle, an ellipse, and a line are recovered as their own kinds only because each is *always* written as one characteristic operator pattern this package recognises (see the shape-detection gotcha above) — a shape drawn any other way, or rotated off-axis, still collapses to a generic `LayoutPath`, and text is recovered as positioned glyph runs with no guarantee the original run boundaries (which characters were grouped into one `Tj` versus several) survive identically. This is a deliberate, permanent contrast with format-preserving codecs like `ooxml.js`'s own `packageCodec`. `pdfCodec` shares `z.codec()`'s *mechanism* (schema-validated both ways) but not that *guarantee* — wrapping this round trip in `z.codec()` validates the shape of what comes out, not its fidelity to what went in.
|
|
197
260
|
|