pdf-parser.js 4.8.16 → 4.8.17
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +1 -1
- package/package.json +1 -1
package/README.md
CHANGED
|
@@ -251,7 +251,7 @@ The package is layered from generic primitives outward to the codec itself:
|
|
|
251
251
|
- **`src/layout.ts`** — this package's own native document model: the `LayoutDocument` item family — Zod schemas and inferred types for text/image/rect/line/ellipse/path/link items, pages (with the hidden speaker-notes channel), the image-asset registry, and `LAYOUT_FORMAT_VERSION`. Ported verbatim from `document-schema.js`, which carried it until its 4.0.0 dropped it — a codec's native model lives in the codec, like `ooxml.js`'s `Package`/`XmlElement` and `markdown-codec`'s AST; only PDF's native model was ever a public shared-schema export, an accident of this package predating the content pivot. The item layer remains the honest boundary between what the format says (positions) and what we think it means (structure): when reconstruction misjudges a wrapped paragraph, the items stay inspectable as the PDF's actual testimony. The shared leaf shapes the family composes from (`Color`, `ContentStrokeStyleSchema`, `LayoutFont`, `LayoutMetadata`) stay in `document-schema.js` and are imported, keeping one definition of each across content and layout.
|
|
252
252
|
- **The math port types** — `MathColor`/`MathGlyphRun`/`MathRule`/`MathStroke`/`MathLayoutItem`/`MathBox`/`MathGlyphMetrics`/`MathFontMetrics` and `PositionedFormula`, sourced from `document-schema.js`'s math layout port (one shared definition across the family, not a local mirror). Deliberately not imported from `documents.js` — that would be a circular dependency once `documents.js` depends on this package. Because every one of these types is plain data (only `MathFontMetrics` carries a method), a real `MathBox` value `documents.js` produces passes into `writePdf({ formulas })` with zero cast, zero wrapper, and zero transformation.
|
|
253
253
|
- **`src/bytes/`** and **`src/image/`** — generic byte and image-container primitives with zero PDF-specific knowledge: a chunked byte writer, backtracking byte reader, CRC32, a hand-written PNG decoder/encoder, JPEG marker scanning for dimensions only (compressed bytes pass through unchanged), a hand-written CCITT Group 3/Group 4 fax decoder (ITU-T T.4/T.6), a hand-written JBIG2 decoder (ITU-T T.88 — `jbig2-arith.ts` MQ decoder, `jbig2-bitmap.ts`, `jbig2-generic.ts`, `jbig2-text.ts`, `jbig2.ts`), and a hand-written JPEG 2000 decoder (ISO/IEC 15444-1 — `jp2-boxes.ts`, `jpeg2000-codestream.ts`, `jpeg2000-tagtree.ts`, `jpeg2000-t2.ts`, `jpeg2000-t1.ts`, `jpeg2000-dwt.ts`, `jpeg2000.ts`). `src/filters.ts` owns all PDF knowledge for CCITT/JBIG2 (resolving parameters, `/JBIG2Globals`, inverting polarity); `src/images-read.ts` owns the JPEG 2000 PDF integration. `src/bytes/flate.ts` is the only file that imports `fflate`.
|
|
254
|
-
- **`src/util/`** —
|
|
254
|
+
- **`src/util/`** — `abort.ts` (`throwIfAborted`, called at every page loop boundary — there is no `await` point in this synchronous pipeline for cancellation to hook into implicitly). The base64 helper that used to sit beside it, as a verbatim copy of `odf.js`'s own, is `byte-codec`'s now ([ExaDev/documents.js#1282](https://github.com/ExaDev/documents.js/issues/1282)): one implementation for the whole family, reached through the dependency this package already had.
|
|
255
255
|
- **`src/crypto/`** — MD5, SHA-256/384/512, RC4, and AES-CBC, hand-written with zero local imports, plus a thin `random.ts` wrapper around `globalThis.crypto.getRandomValues` for the writer's own salts/file keys/IVs. Not a preference: ISO 32000-1's key-derivation algorithms name MD5 and RC4 directly, neither offered by any portable platform crypto API, and `crypto.subtle` is asynchronous where this codec's read and write paths are both synchronous end to end — `getRandomValues` itself, unlike `crypto.subtle`, is synchronous and portable, so it is used directly rather than hand-written. Reaching for `node:crypto` would break the browser bundle. Each hash/cipher module cites its specification (RFC 1321, FIPS 180-4, FIPS 197) and is tested against published conformance vectors.
|
|
256
256
|
- **The codec itself, importing only `layout`/`bytes`/`image`/`crypto`/`util` plus `document-schema.js`'s port types (no OOXML or ODF knowledge):**
|
|
257
257
|
- **Write**: `objects.ts` (the `PdfObject` discriminated union), `afm-widths.ts`/`encoding.ts`/`winansi.ts`/`fonts.ts` (standard-14 metrics, WinAnsi encoding, family resolution), `font-registry.ts` (resolution port plus `resolveFaceWithRegistry`, the one step both measurer and writer resolve through so they can never disagree about which face a `LayoutFont` means), `font-face.ts` (`readFontFace`, reading a standalone font file's family/bold/italic triple off its `name`/`OS/2`/`head` tables), `measure.ts`/`text-layout.ts` (greedy line-wrapping against either standard-14 AFM widths plus per-family correction or a resolved face's own real `hmtx` advances — never both), `content-write.ts` (`LayoutItem[]` → content-stream operators, with text branching on standard-14 vs embedded face encoding, pair-kerning split into `TJ` arrays, stroke `style` becoming real dash/line-cap state, and `/P <</MCID n /OC R>> BDC`…`EMC` spans around items carrying a structure owner or an optional-content layer — the one marked-content spelling both read-side channels resolve through), `write.ts` (the full object graph, cross-reference table, trailer, and embedded font groups, plus the document-level round-trip surfaces: the `/Names /EmbeddedFiles` attachments tree and `/Outlines` bookmark tree, `/OCProperties` optional-content groups with each layer stated explicitly into the default configuration's `/ON` or `/OFF` list, the `/AcroForm` field tree with the merged-field/widget spelling (a single widget merges into its field dict; a multi-widget field's widgets are separate `/Subtype /Widget` kids, each also referenced from its page's `/Annots` — the same annotation object in both places, so a viewer reading only page-level annotations still renders every widget — and fully-qualified names decompose back into the `/T` chain), the `/StructTreeRoot` element tree with its `/ParentTree` number tree associating each marked item's MCID to its owning element, and the package-level residue rows restored inline onto the Catalog or trailer — the XMP packet as an uncompressed `/Metadata` stream, and any row whose serialisation references source-file objects skipped rather than emitted as a dangling reference).
|
package/package.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "pdf-parser.js",
|
|
3
|
-
"version": "4.8.
|
|
3
|
+
"version": "4.8.17",
|
|
4
4
|
"description": "Hand-written, dependency-minimal PDF codec: parses arbitrary real-world PDFs and generates new ones, built on its own codec-owned LayoutDocument item model and Zod 4 codecs.",
|
|
5
5
|
"type": "module",
|
|
6
6
|
"repository": {
|