officeparser 7.0.0 → 7.0.2

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -1,6 +1,10 @@
1
- # officeParser 📄🚀 - The Most Versatile Office Parser & Generator
1
+ # officeParser — Universal Office Document Parser & Generator
2
2
 
3
- A robust, strictly-typed Node.js and Browser library for parsing and generating office files. It not only extracts content from [`docx`](https://en.wikipedia.org/wiki/Office_Open_XML), [`pptx`](https://en.wikipedia.org/wiki/Office_Open_XML), [`xlsx`](https://en.wikipedia.org/wiki/Office_Open_XML), [`odt`](https://en.wikipedia.org/wiki/OpenDocument), [`odp`](https://en.wikipedia.org/wiki/OpenDocument), [`ods`](https://en.wikipedia.org/wiki/OpenDocument), [`pdf`](https://en.wikipedia.org/wiki/PDF), [`rtf`](https://en.wikipedia.org/wiki/Rich_Text_Format), [`csv`](https://en.wikipedia.org/wiki/Comma-separated_values), [`md`](https://en.wikipedia.org/wiki/Markdown), and [`html`](https://en.wikipedia.org/wiki/HTML) into a rich Abstract Syntax Tree (AST), but also provides a powerful generation engine to convert that AST into formats like **Markdown**, **HTML**, **CSV**, **RTF**, **Text**, **PDF**, and **JSON**, including native **RAG-focused chunking** support.
3
+ A robust, strictly-typed **Node.js and Browser** library for parsing office files into a rich **Abstract Syntax Tree (AST)** and generating high-fidelity output in multiple formats.
4
+
5
+ **Parses:** [`docx`](https://en.wikipedia.org/wiki/Office_Open_XML) · [`pptx`](https://en.wikipedia.org/wiki/Office_Open_XML) · [`xlsx`](https://en.wikipedia.org/wiki/Office_Open_XML) · [`odt`](https://en.wikipedia.org/wiki/OpenDocument) · [`odp`](https://en.wikipedia.org/wiki/OpenDocument) · [`ods`](https://en.wikipedia.org/wiki/OpenDocument) · [`pdf`](https://en.wikipedia.org/wiki/PDF) · [`rtf`](https://en.wikipedia.org/wiki/Rich_Text_Format) · [`csv`](https://en.wikipedia.org/wiki/Comma-separated_values) · [`md`](https://en.wikipedia.org/wiki/Markdown) · [`html`](https://en.wikipedia.org/wiki/HTML)
6
+
7
+ **Generates:** `Markdown` · `HTML` · `CSV` · `RTF` · `PDF` · `Plain Text` · `RAG Chunks`
4
8
 
5
9
  [![npm version](https://badge.fury.io/js/officeparser.svg)](https://badge.fury.io/js/officeparser)
6
10
  [![Total Downloads](https://img.shields.io/npm/dt/officeparser.svg)](https://www.npmjs.com/package/officeparser)
@@ -10,21 +14,53 @@ A robust, strictly-typed Node.js and Browser library for parsing and generating
10
14
  ---
11
15
 
12
16
  ### 🌟 [Live Interactive AST Visualizer & Documentation](https://harshankur.github.io/officeParser/) 🌟
13
- *Test any office file in your browser and see the extracted AST, text, and preview in real-time which is rebuilt from the AST!*
17
+ *Upload any office file in your browser — inspect the AST, tweak config, and preview generated output in real-time.*
14
18
 
15
- **What you can do there:**
16
- - **AST Visualizer**: Upload any office file and inspect the hierarchical AST structure, metadata, and raw content.
17
- - **Config Configurator**: Tweak parsing options (like `ignoreNotes`, `ocr`, `newlineDelimiter`) and see the results instantly.
18
- - **Debugging**: Use the visualizer to debug parsing issues by inspecting exactly how nodes are interpreted.
19
- - **Format Specs**: Read detailed specifications for the AST structure and configuration options.
19
+ - **AST Visualizer**: Inspect the hierarchical node tree, metadata, and raw content
20
+ - **Config Configurator**: Tweak options (`ignoreNotes`, `ocr`, `newlineDelimiter`) and see results instantly
21
+ - **Debugging**: Identify exactly how nodes are interpreted
22
+ - **Format Specs**: Read detailed specs for the AST structure and all config options
20
23
 
21
24
  ---
22
25
 
26
+ ### 📝 [Changelog](CHANGELOG.md)
23
27
 
24
28
  ---
25
29
 
26
- ### 📝 [Changelog](CHANGELOG.md)
27
- *Detailed release notes and the full history of updates are available in the project changelog.*
30
+ ## Table of Contents
31
+ - [Install](#install-via-npm)
32
+ - [Command Line Usage](#command-line-usage)
33
+ - [Quick Decision Guide](#quick-decision-guide)
34
+ - [Library Usage: Parsing](#library-usage-parsing)
35
+ - [Async/Await](#asyncawait)
36
+ - [Callback (Backward Compat)](#callback-backward-compat)
37
+ - [File Buffers & ArrayBuffers](#file-buffers--arraybuffers)
38
+ - [`ast.to()` — Generate from AST](#astto--generate-from-ast)
39
+ - [`ast.toText()` — Quick Text Extraction](#asttotext--quick-text-extraction)
40
+ - [OfficeGenerator](#officegenerator)
41
+ - [OfficeConverter — One-Step API](#officeconverter--one-step-api)
42
+ - [Native RAG Chunking](#native-rag-chunking)
43
+ - [The AST Structure](#the-ast-structure)
44
+ - [Deep Dive: Document Components](#deep-dive-document-components)
45
+ - [Performance Highlights](#performance-highlights)
46
+ - [Advanced AST Usage](#advanced-ast-usage)
47
+ - [Configuration Reference](#configuration-reference)
48
+ - [OfficeParserConfig](#officeparserconfig)
49
+ - [GeneratorConfig (Common)](#generatorconfig-common)
50
+ - [onNode Callback](#onnode-callback--advanced-node-manipulation)
51
+ - [styleMap — Semantic Style Mapping](#stylemap--semantic-style-mapping)
52
+ - [HtmlGeneratorConfig](#htmlgeneratorconfig)
53
+ - [MdGeneratorConfig](#mdgeneratorconfig)
54
+ - [PdfGeneratorConfig](#pdfgeneratorconfig)
55
+ - [CsvGeneratorConfig](#csvgeneratorconfig)
56
+ - [TextGeneratorConfig](#textgeneratorconfig)
57
+ - [OfficeConverterConfig](#officeconverterconfig)
58
+ - [ChunkingConfig](#chunkingconfig)
59
+ - [OCR Scheduler & Resource Management](#ocr-scheduler--resource-management)
60
+ - [Browser Usage](#browser-usage)
61
+ - [Troubleshooting & Common Issues](#troubleshooting--common-issues)
62
+ - [Known Limitations](#known-limitations)
63
+ - [Contributing](#contributing)
28
64
 
29
65
  ---
30
66
 
@@ -34,690 +70,807 @@ A robust, strictly-typed Node.js and Browser library for parsing and generating
34
70
  npm i officeparser
35
71
  ```
36
72
 
37
- ## Command Line usage
38
- You can use `officeparser` directly from the terminal to extract content as JSON AST, plain text, or generate new formats like Markdown and HTML.
73
+ ---
74
+
75
+ ## Command Line Usage
39
76
 
40
77
  ```bash
41
- # Get full AST as JSON (default)
42
- npx officeparser /path/to/officeFile.docx
78
+ # Full AST as JSON (default)
79
+ npx officeparser /path/to/file.docx
43
80
 
44
- # Get plain text only
45
- npx officeparser /path/to/officeFile.docx --toText=true
81
+ # Plain text output
82
+ npx officeparser /path/to/file.docx --format=text
46
83
 
47
- # Generate Markdown file
84
+ # Convert DOCX to Markdown and save
48
85
  npx officeparser report.docx --format=md --output=report.md
49
86
 
50
- # Generate HTML file with specific output
87
+ # Convert PPTX to HTML
51
88
  npx officeparser presentation.pptx --format=html --output=preview.html
52
89
 
53
- # Convert spreadsheet to CSV
90
+ # Convert XLSX to CSV
54
91
  npx officeparser data.xlsx --format=csv
92
+
93
+ # Generate RAG chunks
94
+ npx officeparser document.pdf --format=chunks
55
95
  ```
56
96
 
57
- ### Config Options:
58
- - `--format=[json|text|md|html|csv|rtf|pdf|chunks]` The output format. Default is `json`.
59
- - `--output=[path]` Optional file path to write the output to.
60
- - `--toText=[true|false]` Legacy flag to output only plain text. Use `--format=text` instead.
61
- - `--ignoreNotes=[true|false]` Flag to ignore notes from files like PowerPoint. Default is false.
62
- - `--newlineDelimiter=[delimiter]` The delimiter to use for new lines. Default is `\n`.
63
- - `--putNotesAtLast=[true|false]` Flag to collect notes at the end of files like PowerPoint. Default is false.
64
- - `--outputErrorToConsole=[true|false]` **(Deprecated)** Flag to output errors to the console. Use `onWarning` callback in library usage.
65
- - `--extractAttachments=[true|false]` Flag to extract images/charts as Base64. Default is false.
66
- - `--ocr=[true|false]` Flag to enable OCR for extracted images. Default is false.
67
- - `--includeRawContent=[true|false]` Flag to include raw XML/RTF content in nodes. Default is false.
68
- - `--includeBreakNodes=[true|false]` Flag to include break nodes. Currently only available for DOCX documents.
69
- - `--verbose=[true|false]` Show full error stack traces.
97
+ ### CLI Options
70
98
 
99
+ | Flag | Values | Default | Description |
100
+ |------|--------|---------|-------------|
101
+ | `--format` | `json\|text\|md\|html\|csv\|rtf\|pdf\|chunks` | `json` | Output format |
102
+ | `--output` | path | — | Write output to a file |
103
+ | `--toText` | `true\|false` | `false` | **Deprecated.** Use `--format=text` |
104
+ | `--ignoreNotes` | `true\|false` | `false` | Ignore speaker notes (PPTX/ODP) |
105
+ | `--putNotesAtLast` | `true\|false` | `false` | Collect notes at end of output |
106
+ | `--newlineDelimiter` | string | `\n` | Delimiter between lines |
107
+ | `--extractAttachments` | `true\|false` | `false` | Extract images/charts as Base64 |
108
+ | `--ocr` | `true\|false` | `false` | Enable OCR for images |
109
+ | `--includeRawContent` | `true\|false` | `false` | Include raw XML/RTF in nodes |
110
+ | `--includeBreakNodes` | `true\|false` | `false` | Include break nodes (DOCX only) |
111
+ | `--outputErrorToConsole` | `true\|false` | `false` | **Deprecated.** Use `onWarning` callback |
112
+ | `--verbose` | `true\|false` | `false` | Show full error stack traces |
71
113
 
72
- ## Library Usage
73
- In **v7.0.0**, the library has evolved into a dual-purpose **Parser** and **Generator**. You can first parse any office file into a structured AST and then use the `OfficeGenerator` to transform that AST into various formats or chunks.
114
+ ---
115
+
116
+ ## Quick Decision Guide
117
+
118
+ | Goal | API to use |
119
+ |------|-----------|
120
+ | Extract text / AST from a file | `OfficeParser.parseOffice(file)` |
121
+ | Convert directly to another format | `OfficeConverter.convert(file, 'md')` |
122
+ | Parse first, then generate | `parseOffice()` → `OfficeGenerator.generate(ast, 'html')` |
123
+ | Convert on the AST itself (shorthand) | `ast.to('md')` |
124
+ | RAG pipeline chunking | `OfficeConverter.convert(file, 'chunks', {...})` |
125
+
126
+ ---
127
+
128
+ ## Library Usage: Parsing
129
+
130
+ ### Async/Await
74
131
 
75
- ### Getting Started (Async/Await)
76
132
  ```js
77
133
  const officeParser = require('officeparser');
78
134
 
79
- async function parseMyFile() {
80
- try {
81
- // parseOffice returns an OfficeParserAST object
82
- const ast = await officeParser.parseOffice("/path/to/officeFile.docx");
83
-
84
- // Use the built-in helper to get plain text (similar to old behavior)
85
- const text = ast.toText();
86
- console.log(text);
87
-
88
- // Access structured content
89
- console.log(ast.content); // Array of hierarchical nodes (paragraphs, tables, etc.)
90
- console.log(ast.metadata); // Document properties (author, title, etc.)
91
- } catch (err) {
92
- console.error(err);
93
- }
94
- }
135
+ const ast = await officeParser.parseOffice('/path/to/file.docx');
136
+
137
+ console.log(ast.type); // 'docx'
138
+ console.log(ast.metadata); // { author, title, created, ... }
139
+ console.log(ast.content); // Array of hierarchical nodes
140
+ console.log(ast.attachments);// Images/charts (if extractAttachments: true)
141
+ console.log(ast.warnings); // Non-fatal issues from parsing phase
142
+ ```
143
+
144
+ **TypeScript (named import):**
145
+ ```ts
146
+ import { OfficeParser } from 'officeparser';
147
+
148
+ const ast = await OfficeParser.parseOffice('report.docx', {
149
+ extractAttachments: true,
150
+ ocr: true,
151
+ });
152
+ ```
153
+
154
+ ### Callback (Backward Compat)
155
+
156
+ ```js
157
+ officeParser.parseOffice('/path/to/file.docx', function(ast, err) {
158
+ if (err) { console.error(err); return; }
159
+ console.log(ast.toText());
160
+ });
95
161
  ```
96
162
 
97
- ### Helper Function for Text Extraction (Modern simple way)
98
- If you only need the text and want to maintain a simple one-liner, you can use this pattern:
163
+ ### File Buffers & ArrayBuffers
164
+
165
+ Pass a `Buffer`, `ArrayBuffer`, or `Uint8Array` instead of a file path:
166
+
99
167
  ```js
100
- // Simple helper to get text directly
101
- const getText = async (file, config) => (await officeParser.parseOffice(file, config)).toText();
168
+ const fs = require('fs');
169
+ const buffer = fs.readFileSync('/path/to/file.pdf');
170
+ const ast = await officeParser.parseOffice(buffer);
171
+ ```
172
+
173
+ > [!IMPORTANT]
174
+ > **Text-based formats from buffers need a `fileType` hint.**
175
+ > Formats like `md`, `html`, and `csv` have no magic bytes, so the parser cannot
176
+ > auto-detect them from a buffer. You **must** provide `fileType` in that case:
177
+ > ```js
178
+ > const ast = await officeParser.parseOffice(markdownBuffer, { fileType: 'md' });
179
+ > ```
180
+
181
+ ### `ast.to()` — Generate from AST
182
+
183
+ The preferred way to convert a parsed AST to another format. Returns a `ConversionResult`.
184
+
185
+ ```ts
186
+ // ConversionResult shape:
187
+ // { value: string | Uint8Array | OfficeChunk[], messages: OfficeIssue[] }
188
+
189
+ const { value: markdown, messages } = await ast.to('md');
190
+ const { value: html } = await ast.to('html', { includeFormatting: false });
191
+ const { value: chunks } = await ast.to('chunks', { strategy: 'fixed-size', chunkSize: 800 });
192
+ const { value: pdfBytes } = await ast.to('pdf'); // Uint8Array
193
+ ```
194
+
195
+ ### `ast.toText()` — Quick Text Extraction
196
+
197
+ > [!NOTE]
198
+ > `toText()` is **synchronous** and deprecated in favour of the async `ast.to('text')`.
199
+ > It remains available for backward compatibility.
102
200
 
103
- // usage
104
- const text = await getText("/path/to/officeFile.docx");
105
- console.log(text);
201
+ ```js
202
+ const text = ast.toText(); // synchronous, returns plain string
106
203
  ```
107
204
 
108
- ## Using the OfficeGenerator
109
- The `OfficeGenerator` is a powerful tool to convert your AST into human-readable formats or structured data.
205
+ ---
206
+
207
+ ## OfficeGenerator
208
+
209
+ Use `OfficeGenerator.generate(ast, format, config?)` when you need to produce output from an already-parsed AST:
110
210
 
111
- ```typescript
211
+ ```ts
112
212
  import { OfficeParser, OfficeGenerator } from 'officeparser';
113
213
 
114
214
  const ast = await OfficeParser.parseOffice('report.docx');
115
215
 
116
- // 1. Convert to Markdown
117
- const md = await OfficeGenerator.generate(ast, 'md');
118
- console.log(md.value);
216
+ // Convert to Markdown
217
+ const { value: md } = await OfficeGenerator.generate(ast, 'md');
119
218
 
120
- // 2. Convert to HTML with structured style mapping (Recommended)
121
- const html = await OfficeGenerator.generate(ast, 'html', {
219
+ // Convert to HTML with style mapping
220
+ const { value: html } = await OfficeGenerator.generate(ast, 'html', {
122
221
  includeFormatting: true,
123
222
  styleMap: [
124
- {
125
- selector: { nodeType: 'paragraph', attributes: { style: 'Heading 1' } },
126
- output: { tag: 'h1', classes: ['main-title'] }
223
+ {
224
+ selector: { nodeType: 'paragraph', attributes: { style: 'Heading 1' } },
225
+ output: { tag: 'h1', classes: ['main-title'] }
127
226
  }
128
227
  ]
129
228
  });
130
- console.log(html.value);
131
229
 
132
- // 3. Convert to CSV (for spreadsheets)
133
- const csv = await OfficeGenerator.generate(ast, 'csv');
134
- console.log(csv.value);
230
+ // Convert to CSV (spreadsheets)
231
+ const { value: csv } = await OfficeGenerator.generate(ast, 'csv');
135
232
  ```
136
233
 
137
- ## The New "One-Step" API: `OfficeConverter`
138
- In **v7.0.0**, we introduced the `OfficeConverter.convert` method. This is the new high-level API designed for one-step transformations where you don't need to manually interact with the AST. It automatically handles parser and generator configuration synchronization.
234
+ **Supported destinations:** `'text'` · `'md'` · `'html'` · `'csv'` · `'rtf'` · `'pdf'` · `'chunks'`
235
+
236
+ > [!NOTE]
237
+ > **PDF generation** requires the optional `puppeteer` peer dependency:
238
+ > ```bash
239
+ > npm install puppeteer
240
+ > ```
139
241
 
140
- ```typescript
242
+ ---
243
+
244
+ ## OfficeConverter — One-Step API
245
+
246
+ `OfficeConverter.convert()` combines parsing and generation in a single call. It automatically syncs parser options from generator config (e.g., enables `extractAttachments` when images are requested).
247
+
248
+ ```ts
141
249
  import { OfficeConverter } from 'officeparser';
142
250
 
143
- // One-step conversion from DOCX to Markdown
144
- const result = await OfficeConverter.convert('report.docx', 'md');
145
- console.log(result.value); // The generated Markdown string
146
- console.log(result.messages); // Array of warnings/info (e.g., "Skipped unsupported drawing")
251
+ // Minimal usage
252
+ const { value: markdown } = await OfficeConverter.convert('report.docx', 'md');
147
253
 
148
- // Complex conversion with nested configuration
149
- const htmlResult = await OfficeConverter.convert('data.xlsx', 'html', {
150
- parseConfig: {
151
- ignoreNotes: true
254
+ // With config
255
+ const { value: html, messages } = await OfficeConverter.convert('data.xlsx', 'html', {
256
+ parseConfig: {
257
+ ignoreNotes: true,
258
+ newlineDelimiter: '\n\n',
152
259
  },
153
260
  generatorConfig: {
154
261
  includeFormatting: true,
155
262
  styleMap: [
156
- {
157
- selector: { attributes: { style: { value: 'Header', operator: '~=' } } },
158
- output: { tag: 'h2', classes: ['data-header'] }
263
+ {
264
+ selector: { attributes: { style: { value: 'Header', operator: '~=' } } },
265
+ output: { tag: 'h2', classes: ['data-header'] }
159
266
  }
160
267
  ]
161
268
  },
162
- onWarning: (msg) => console.warn("Conversion Warning:", msg)
269
+ onWarning: (issue) => console.warn(`[${issue.code}] ${issue.message}`)
163
270
  });
164
271
  ```
165
272
 
273
+ > [!IMPORTANT]
274
+ > The `OfficeConverterConfig` shape uses **nested** `parseConfig` and `generatorConfig` sub-objects.
275
+ > Do **not** put parser or generator options at the top level — only `onWarning` lives there.
276
+
277
+ ---
278
+
166
279
  ## Native RAG Chunking
167
- `officeParser` provides native support for document chunking, specifically designed for Retrieval-Augmented Generation (RAG) workflows. It offers three distinct strategies to split your documents while maintaining context and metadata.
168
-
169
- ### 1. Fixed-Size Strategy (Recursive)
170
- Splits text into chunks based on character count with a specified overlap. It uses smart boundary detection to avoid cutting in the middle of sentences or paragraphs.
171
-
172
- ### 2. Document Structure Strategy
173
- Splits the document at natural structural boundaries like pages (PDF/Word), slides (PPTX), or high-level headings. This preserves the logical flow of the document.
174
-
175
- ### 3. Semantic Strategy
176
- Uses cosine similarity between sentence embeddings to identify coherent topic boundaries. This ensures that each chunk contains semantically related content (requires an embedding function).
177
-
178
- ### The `OfficeChunk` Interface
179
- Every chunk produced contains not just text, but rich metadata to help your RAG pipeline:
180
- ```typescript
181
- {
182
- text: string; // The chunk content
183
- metadata: {
184
- sourceType: string; // e.g., "docx", "pdf"
185
- pageNumber?: number; // Current page
186
- slideNumber?: number; // Current slide
187
- closestHeading?: string; // The heading this chunk belongs to
188
- chunkIndex: number; // Sequential index
189
- }
190
- }
280
+
281
+ `officeParser` provides native document chunking for Retrieval-Augmented Generation (RAG) pipelines with three strategies:
282
+
283
+ ### Strategy 1: Document Structure (Default)
284
+ Splits at natural AST boundaries (paragraphs, headings, pages, slides, sheets). Preserves logical flow.
285
+
286
+ ```ts
287
+ const { value: chunks } = await OfficeConverter.convert('report.docx', 'chunks', {
288
+ generatorConfig: {
289
+ chunksConfig: {
290
+ strategy: 'document-structure',
291
+ splitBy: 'heading', // 'paragraph' | 'heading' | 'page' | 'slide' | 'sheet'
292
+ maxChunkSize: 1500,
293
+ tableSplitStrategy: 'row', // repeats header row in every chunk — ideal for RAG
294
+ }
295
+ }
296
+ });
191
297
  ```
192
298
 
193
- #### Example: Generating Chunks
194
- ```typescript
195
- const chunks = await OfficeGenerator.generate(ast, 'chunks', {
196
- strategy: 'fixed-size',
197
- maxChunkSize: 1000,
198
- chunkOverlap: 200
299
+ ### Strategy 2: Fixed-Size (Recursive)
300
+ Splits by character count with overlap. Equivalent to LangChain's `RecursiveCharacterTextSplitter`.
301
+
302
+ ```ts
303
+ const { value: chunks } = await OfficeConverter.convert('report.docx', 'chunks', {
304
+ generatorConfig: {
305
+ chunksConfig: {
306
+ strategy: 'fixed-size',
307
+ chunkSize: 1000,
308
+ chunkOverlap: 200,
309
+ }
310
+ }
199
311
  });
200
- console.log(`Generated ${chunks.value.length} chunks`);
312
+ console.log(`Generated ${chunks.length} chunks`);
201
313
  ```
202
314
 
203
- ### Using Callbacks (Backward Compatibility Support)
204
- Callbacks are still supported for those preferred, but the data returned is now the AST object.
205
- ```js
206
- const officeParser = require('officeparser');
315
+ ### Strategy 3: Semantic
316
+ Uses cosine similarity between sentence embeddings to find topic boundaries. Requires you to provide an `embeddingFunction`.
317
+
318
+ ```ts
319
+ import OpenAI from 'openai';
320
+ const openai = new OpenAI();
207
321
 
208
- officeParser.parseOffice("/path/to/officeFile.docx", function(ast, err) {
209
- if (err) {
210
- console.error(err);
211
- return;
322
+ const { value: chunks } = await OfficeConverter.convert('report.docx', 'chunks', {
323
+ generatorConfig: {
324
+ chunksConfig: {
325
+ strategy: 'semantic',
326
+ embeddingFunction: async (text) => {
327
+ const res = await openai.embeddings.create({
328
+ input: text, model: 'text-embedding-3-small'
329
+ });
330
+ return res.data[0].embedding;
331
+ },
332
+ similarityThreshold: 0.8,
333
+ maxChunkSize: 2000,
334
+ }
212
335
  }
213
- // Get text from AST
214
- console.log(ast.toText());
215
336
  });
216
337
  ```
217
338
 
218
- ### Using File Buffers or ArrayBuffers
219
- You can pass a file path string, a Node.js `Buffer`, or an `ArrayBuffer`.
220
- ```js
221
- const fs = require('fs');
222
- const officeParser = require('officeparser');
223
- const buffer = fs.readFileSync("/path/to/officeFile.pdf");
339
+ ### The `OfficeChunk` Object
340
+
341
+ Every chunk contains text and rich metadata for citations and filtered retrieval:
224
342
 
225
- officeParser.parseOffice(buffer)
226
- .then(ast => console.log(ast.toText()))
227
- .catch(console.error);
343
+ ```ts
344
+ interface OfficeChunk {
345
+ text: string;
346
+ /** Rich metadata for filtered retrieval */
347
+ metadata: {
348
+ sourceType: string; // e.g., 'docx', 'pdf'
349
+ pageNumber?: number; // (PDF only)
350
+ slideNumber?: number; // (PPTX only)
351
+ sheetName?: string; // (XLSX only)
352
+ closestHeading?: string; // Nearest heading above this chunk
353
+ isTableChunk?: boolean; // True if part of a split table
354
+ };
355
+ startIndex?: number; // Character offset (if addStartIndex: true)
356
+ endIndex?: number; // End character offset (if addStartIndex: true)
357
+ }
228
358
  ```
229
359
 
360
+ ---
361
+
230
362
  ## The AST Structure
231
- The `OfficeParserAST` provides a format-agnostic representation of your document, allowing you to traverse and manipulate content as a tree.
232
363
 
233
- ### Visualizing the AST
234
- The `OfficeParserAST` provides a format-agnostic representation of your document. Below is a simplified visualization of how the tree is structured:
364
+ `OfficeParserAST` is a format-agnostic document representation:
235
365
 
236
366
  ```text
237
367
  OfficeParserAST
238
- ├── type: "docx" | "pdf" | "xlsx" | "csv" | "md" | ... (11 formats supported)
239
- ├── metadata: { author, title, created, modified, ..., customProperties }
368
+ ├── type: 'docx' | 'pdf' | 'xlsx' | 'csv' | 'md' | ... (11 formats)
369
+ ├── metadata: { author, title, created, modified, customProperties, styleMap, ... }
240
370
  ├── content: [ OfficeContentNode ]
241
- │ ├── type: "paragraph" | "heading" | "table" | "list" | ...
242
- │ ├── text: "Concatenated text of this node and all children"
243
- │ ├── children: [ OfficeContentNode ] (recursive)
244
- │ ├── formatting: { bold, italic, color, size, font, ... }
245
- │ ├── metadata: { level, listId, paragraphIndentation, row, col, ... }
246
- │ └── rawContent: "<xml>...</xml>" (if enabled)
247
- ├── attachments: [ OfficeAttachment ]
248
- │ ├── type: "image" | "chart"
249
- │ ├── name: "image1.png"
250
- │ ├── data: "base64..."
251
- │ ├── ocrText: "Text extracted via OCR"
252
- │ └── chartData: { title, dataSets, labels, ... }
253
- └── toText(): Function -> returns full plain text
254
- ```
255
-
256
- #### Representative JSON Snippet
257
- ```json
258
- {
259
- "type": "docx",
260
- "metadata": { "author": "John Doe", "title": "Annual Report", "customProperties": { "Department": "Finance" } },
261
- "content": [
262
- {
263
- "type": "heading",
264
- "text": "Introduction",
265
- "metadata": { "level": 1 },
266
- "children": [
267
- { "type": "text", "text": "Introduction", "formatting": { "bold": true } }
268
- ]
269
- },
270
- {
271
- "type": "paragraph",
272
- "text": "This is a report with an image.",
273
- "children": [
274
- { "type": "text", "text": "This is a report with an " },
275
- { "type": "image", "metadata": { "attachmentName": "img1.png" } }
276
- ]
277
- }
278
- ],
279
- "attachments": [
280
- { "name": "img1.png", "type": "image", "data": "iVBOR...", "ocrText": "Extracted Text" }
281
- ]
371
+ │ ├── type: 'paragraph' | 'heading' | 'table' | 'list' | 'image' | 'chart' | ...
372
+ │ ├── text: string (concatenated text of node + all descendants)
373
+ │ ├── children: [ OfficeContentNode ] (recursive)
374
+ │ ├── formatting: { bold, italic, underline, color, size, font, alignment, ... }
375
+ │ └── metadata: { level, listId, row, col, rowSpan, colSpan, style, ... }
376
+ ├── attachments: [ OfficeAttachment ] (populated when extractAttachments: true)
377
+ │ ├── type: 'image' | 'chart'
378
+ │ ├── name: string
379
+ │ ├── mimeType: string
380
+ │ ├── data: string (Base64)
381
+ │ ├── ocrText?: string (if ocr: true)
382
+ │ └── chartData?: { title, dataSets, labels }
383
+ ├── warnings: OfficeIssue[] (non-fatal issues from the parsing phase)
384
+ ├── to(format, config?) (format: 'html'|'md'|'text'|'csv'|'rtf'|'pdf'|'chunks', returns { value, messages })
385
+ └── toText() (Deprecated: use .to('text') instead)
386
+ ```
387
+
388
+ ### `OfficeIssue` — Warning / Error Object
389
+
390
+ All warnings and errors (from both parsing and generation) use this shape:
391
+
392
+ ```ts
393
+ interface OfficeIssue {
394
+ type: 'warning' | 'info' | 'error';
395
+ code: OfficeWarningType | OfficeErrorType; // typed enum, e.g. 'OCR_FAILED'
396
+ message: string;
397
+ node?: OfficeContentNode; // the node that triggered the issue, if any
398
+ details?: any; // original error or extra context
282
399
  }
283
400
  ```
284
401
 
402
+ ---
403
+
285
404
  ## Deep Dive: Document Components
286
405
 
287
- ### 1. Working with Lists
288
- Lists are represented as sequential `list` nodes. To reconstruct or track a list, use the `metadata` fields:
406
+ ### 1. Lists
289
407
 
290
408
  ```text
291
409
  List Node
292
- ├── type: "list"
293
- ├── metadata: {
294
- listId: "1",
295
- listType: "ordered",
296
- indentation: 0,
297
- paragraphIndentation: { left: 720, hanging: 360 },
298
- itemIndex: 0
299
- }
300
- └── children: [ Text Content... ]
410
+ ├── type: 'list'
411
+ ├── metadata: {
412
+ │ listId: '1', // items with the same listId belong to one logical list
413
+ │ listType: 'ordered' | 'unordered',
414
+ │ indentation: 0, // nesting level (0-based)
415
+ │ itemIndex: 0, // sequential position within the list level
416
+ │ paragraphIndentation: { left, hanging, right, firstLine }
417
+ │ }
418
+ └── children: [ Text content ]
301
419
  ```
302
420
 
303
- - **`listId`**: A unique identifier for the list definition. Multiple items with the same `listId` belong to the same logical list.
304
- - **`indentation`**: The structural nesting level (0-based).
305
- - **`paragraphIndentation`**: The physical indentation formatting in twentieths of a point (twips) (e.g., `left`, `right`, `firstLine`, `hanging`).
306
- - **`itemIndex`**: The sequential position within that list level.
307
- - **`listType`**: Either `ordered` (numbered) or `unordered` (bulleted).
308
-
309
421
  > [!TIP]
310
- > Even if a list is interrupted by a regular paragraph, the `itemIndex` will continue to increment for the same `listId`, allowing you to maintain correct numbering.
422
+ > Even if a list is interrupted by a regular paragraph, `itemIndex` keeps incrementing for the same `listId`, so numbering stays correct.
423
+
424
+ ### 2. Tables
311
425
 
312
- ### 2. Navigating Tables
313
- Tables follow a strict hierarchy: `table` -> `row` -> `cell`.
426
+ Tables follow a strict `table → row → cell` hierarchy:
314
427
 
315
428
  ```text
316
- Table Node
317
- ├── type: "table"
318
- └── children: [ Row Node ]
319
- ├── type: "row"
320
- └── children: [ Cell Node ]
321
- ├── type: "cell"
322
- ├── metadata: { row, col, rowSpan, colSpan }
323
- └── children: [ Paragraph/List/etc. ]
429
+ Table Node (type: 'table')
430
+ └── children: Row Nodes (type: 'row')
431
+ └── children: Cell Nodes (type: 'cell')
432
+ ├── metadata: { row, col, rowSpan?, colSpan? }
433
+ └── children: [ Paragraph | List | Table | ... ]
324
434
  ```
325
435
 
326
- - **`row` / `col`**: Zero-based indices for grid positioning.
327
- - **`rowSpan` / `colSpan`** (Optional): Integer values indicating merged cells (primarily in ODF formats). If absent, the cell is not merged.
328
- - **Recursive Content**: Cells contain their own `children` array, which can include paragraphs, lists, or even other nested tables.
436
+ - `row` / `col`: zero-based grid position
437
+ - `rowSpan` / `colSpan`: merged cells (primarily ODF formats)
438
+ - Cells can contain nested tables
329
439
 
330
- ### 3. Charts & Data
331
- When a chart is discovered, it's added as a `chart` node in the content and a corresponding `OfficeAttachment`.
440
+ ### 3. Images & OCR
332
441
 
333
442
  ```text
334
- Chart Node
335
- ├── type: "chart"
336
- ├── metadata: { attachmentName: "chart1.xml" }
337
- └── Attachment (Linked)
338
- └── chartData: { title, dataSets: [...], labels: [...] }
443
+ Image Node (type: 'image')
444
+ ├── metadata: { attachmentName: 'img1.png', altText: '...' }
445
+ └── → Attachment: { data: 'base64...', ocrText: '...' }
339
446
  ```
340
447
 
341
- - **`attachmentName`**: Links the content node to the `attachments` array.
342
- - **`chartData`**: A structured object containing titles, axis labels, and category/series data.
448
+ - Set `extractAttachments: true` to populate `attachment.data`
449
+ - Set `ocr: true` (requires `extractAttachments: true`) to populate `ocrText`
343
450
 
344
- ### 4. Images, OCR & Alt Text
345
- Images are linked via `attachmentName` and can contain valuable metadata:
451
+ ### 4. Charts
346
452
 
347
453
  ```text
348
- Image Node
349
- ├── type: "image"
350
- ├── metadata: { attachmentName: "img1.png", altText: "..." }
351
- └── Attachment (Linked)
352
- ├── data: "base64..."
353
- └── ocrText: "Extracted via OCR"
454
+ Chart Node (type: 'chart')
455
+ ├── metadata: { attachmentName: 'chart1.xml' }
456
+ └── → Attachment: { chartData: { title, dataSets, labels } }
354
457
  ```
355
458
 
356
- - **OCR Text**: If `ocr: true` is set in config, `ocrText` will contain the text found within the image.
357
- - **Alt Text**: Extracted from the document's internal image descriptions.
358
- - **Formatting**: `OfficeContentNode` images may also have parent alignment metadata.
359
-
360
459
  ### 5. Text Formatting
361
- Each `OfficeContentNode` can have a `formatting` object that defines how the text should be styled.
362
460
 
363
- ```text
364
- Text Node
365
- └── formatting: {
366
- bold: boolean,
367
- italic: boolean,
368
- underline: boolean,
369
- strikethrough: boolean,
370
- color: "#hex",
371
- backgroundColor: "#hex",
372
- size: "12pt",
373
- font: "Arial",
374
- subscript: boolean,
375
- superscript: boolean,
376
- alignment: "left" | "center" | "right" | "justify"
461
+ ```ts
462
+ formatting: {
463
+ bold?: boolean
464
+ italic?: boolean
465
+ underline?: boolean
466
+ strikethrough?: boolean
467
+ color?: string // '#RRGGBB'
468
+ backgroundColor?: string
469
+ size?: string // e.g. '12pt'
470
+ font?: string
471
+ subscript?: boolean
472
+ superscript?: boolean
473
+ alignment?: 'left' | 'center' | 'right' | 'justify'
377
474
  }
378
475
  ```
379
476
 
380
- Formatting can be found at two levels:
381
- 1. **Node Level**: Applied directly to a text run or paragraph.
382
- 2. **Document Level**: Found in `ast.metadata.formatting` (defaults) or `ast.metadata.styleMap` (named styles).
477
+ ### 6. Break Nodes (DOCX only)
383
478
 
384
- ### 6. Breaks
385
- Breaks are currently only supported when parsing DOCX-documents. Breaks are added as a node of type `break` and carry metadata of the type `BreakMetadata`. When `includeRawContent` is enabled, they also include the `rawContent` string from the original XML.
479
+ When `includeBreakNodes: true`, break elements appear as nodes:
386
480
 
387
481
  ```text
388
- Break Node
389
- ├── type: "break"
482
+ Break Node (type: 'break')
390
483
  └── metadata: {
391
- breakType: "textWrapping" | "page" | "column" | "lastRenderedPage" | "carriageReturn",
392
- clear?: "all" | "left" | "none" | "right"
484
+ breakType: 'textWrapping' | 'page' | 'column' | 'lastRenderedPage' | 'carriageReturn',
485
+ clear?: 'all' | 'left' | 'none' | 'right'
393
486
  }
394
487
  ```
395
488
 
396
- - `breakType`: Type of break. `textWrapping` (default) is a standard line break, `page` is a page break, `column` is a break to the next column, `lastRenderedPage` is a soft break inserted by Word, and `carriageReturn` is an explicit carriage return (`w:cr`).
397
- - `clear`: Relevant for `textWrapping`. Indicates if text should wrap around floating objects.
398
-
399
489
  > [!NOTE]
400
- > Even though break nodes don't have a `text` property, the `ast.toText()` method will automatically convert them to newlines (`\n`) or the configured delimiter in the final string output.
401
-
402
- ### 7. Advanced Metadata
403
- The `ast.metadata` object provides document-wide context:
404
- - **`styleMap`**: A dictionary of style names to their `TextFormatting` definitions found in the document.
405
- - **`formatting`**: Document-wide default settings (e.g., default font or font size).
406
- - **`customProperties`**: A dictionary of user-defined metadata embedded in the document (OOXML `custom.xml`, ODF `meta:user-defined`, or PDF Info dictionary).
407
-
408
- ### 8. Custom Properties
409
- You can access custom user-defined metadata that might be embedded in the document:
410
-
411
- ```javascript
412
- const ast = await officeParser.parseOffice("contract.docx");
413
- console.log("Custom Metadata:", ast.metadata.customProperties);
414
- // Output: { "ProjectID": "ABC-123", "InternalReview": true }
415
- ```
416
-
417
- ## Performance & Fidelity Highlights (v7.0.0)
418
- The v7.0.0 release brings significant internal optimizations and fidelity improvements:
419
- - **OpenOffice Speedups**: Up to **23x faster** parsing for ODP presentations thanks to optimized XML caching.
420
- - **Excel Memory Efficiency**: Resolved $O(n)$ memory overhead issues for large spreadsheets (#91) by switching to iterative stream-based parsing.
421
- - **RTF Performance**: Rewritten core loop to resolve $O(n^2)$ bottlenecks during string accumulation.
422
- - **Advanced Table Fidelity**: Native support for **vertical cell merging** (`vMerge`) and **horizontal spanning** (`gridSpan`) in DOCX, ensuring complex tables look exactly as they do in Word.
423
- - **Parser Extensions**: You can now parse `CSV`, `Markdown`, and `HTML` files *into* the unified Office AST, allowing you to use the `OfficeGenerator` on them just like any other format.
424
-
425
- ### Advanced AST Usage
426
- Beyond using `ast.toText()`, you can interact with the structural data directly:
427
-
428
- #### 1. Extract all images and their OCR text
429
- ```javascript
430
- const ast = await officeParser.parseOffice("report.docx", { ocr: true });
431
- const images = ast.attachments.filter(a => a.mimeType.startsWith('image/'));
432
- images.forEach(img => {
433
- console.log(`Image: ${img.name} (OCR: ${img.ocrText || 'N/A'})`);
434
- });
490
+ > Break nodes have no `text` property, but `ast.toText()` and `ast.to('text')` automatically convert them to the configured newline delimiter.
491
+
492
+ ### 7. Document Metadata
493
+
494
+ ```ts
495
+ ast.metadata = {
496
+ author?: string
497
+ title?: string
498
+ created?: Date
499
+ modified?: Date
500
+ description?: string
501
+ customProperties?: Record<string, any> // user-defined metadata from the document
502
+ styleMap?: Record<string, TextFormatting> // named styles → formatting definitions
503
+ formatting?: TextFormatting // document-wide defaults
504
+ }
435
505
  ```
436
506
 
437
- #### 2. Find specific headings
438
- ```javascript
439
- const headings = ast.content.filter(node => node.type === 'heading' && node.metadata?.level === 1);
440
- console.log("Main Chapters:", headings.map(h => h.text));
507
+ **Accessing custom properties:**
508
+ ```js
509
+ const ast = await officeParser.parseOffice('contract.docx');
510
+ console.log(ast.metadata.customProperties);
511
+ // { "ProjectID": "ABC-123", "InternalReview": true }
441
512
  ```
442
513
 
443
- #### 3. Custom output (e.g., Simple Markdown conversion)
444
- ```javascript
445
- const toMarkdown = (nodes) => {
446
- return nodes.map(node => {
447
- if (node.type === 'heading') return `${'#'.repeat(node.metadata?.level || 1)} ${node.text}`;
448
- if (node.type === 'list') return `- ${node.text}`;
449
- if (node.type === 'table') return "[Table Data]"; // expand children for actual table
450
- return node.text;
451
- }).join('\n\n');
452
- };
453
- console.log(toMarkdown(ast.content));
514
+ ---
515
+
516
+ ## Performance Highlights
517
+
518
+ Key internal optimizations shipped in recent versions:
519
+
520
+ - **OpenOffice (ODP)**: Up to **23× faster** parsing via optimized XML pre-parsing and style caching
521
+ - **Excel Memory**: Resolved O(n) memory overhead on large sparse spreadsheets using iterative stream-based parsing
522
+ - **RTF Parser**: Rewrote string accumulation loop to eliminate O(n²) bottleneck in large files
523
+ - **Table Fidelity (DOCX)**: Native support for vertical cell merging (`vMerge`) and horizontal spanning (`gridSpan`)
524
+
525
+ ---
526
+
527
+ ## Advanced AST Usage
528
+
529
+ ### Extract all headings
530
+ ```js
531
+ const headings = ast.content.filter(n => n.type === 'heading' && n.metadata?.level === 1);
532
+ console.log(headings.map(h => h.text));
533
+ ```
534
+
535
+ ### Extract images with OCR text
536
+ ```js
537
+ const ast = await officeParser.parseOffice('report.docx', { extractAttachments: true, ocr: true });
538
+ ast.attachments.filter(a => a.mimeType?.startsWith('image/')).forEach(img => {
539
+ console.log(`${img.name}: ${img.ocrText ?? 'no OCR'}`);
540
+ });
454
541
  ```
455
542
 
456
- #### 4. Extracting Tables to CSV
457
- Iterate through table nodes and their children (rows -> cells) to build a CSV string.
458
- ```javascript
459
- const tables = ast.content.filter(node => node.type === 'table');
460
- tables.forEach((table, index) => {
543
+ ### Extract tables to CSV manually
544
+ ```js
545
+ ast.content.filter(n => n.type === 'table').forEach((table, i) => {
461
546
  const csv = table.children
462
- .filter(row => row.type === 'row')
463
- .map(row =>
464
- row.children
465
- .filter(cell => cell.type === 'cell')
466
- .map(cell => `"${cell.text.replace(/"/g, '""')}"`) // Escape quotes
467
- .join(',')
468
- )
547
+ .filter(r => r.type === 'row')
548
+ .map(r => r.children.filter(c => c.type === 'cell')
549
+ .map(c => `"${c.text.replace(/"/g, '""')}"`)
550
+ .join(','))
469
551
  .join('\n');
470
- console.log(`Table ${index + 1} CSV:\n${csv}`);
552
+ console.log(`Table ${i + 1}:\n${csv}`);
471
553
  });
472
554
  ```
473
555
 
474
- #### 5. Filtering by Formatting (e.g., Bold Text)
475
- Find all text nodes that have specific formatting applied.
476
- ```javascript
477
- function findBoldText(nodes) {
478
- let results = [];
479
- nodes.forEach(node => {
480
- if (node.type === 'text' && node.formatting?.bold) {
481
- results.push(node.text);
482
- }
483
- if (node.children) {
484
- results = results.concat(findBoldText(node.children));
485
- }
486
- });
487
- return results;
556
+ ### Find all bold text runs
557
+ ```js
558
+ function findBold(nodes) {
559
+ return nodes.flatMap(n => [
560
+ ...(n.type === 'text' && n.formatting?.bold ? [n.text] : []),
561
+ ...(n.children ? findBold(n.children) : [])
562
+ ]);
488
563
  }
489
-
490
- const boldStrings = findBoldText(ast.content);
491
- console.log("Bold Text Found:", boldStrings);
564
+ console.log(findBold(ast.content));
492
565
  ```
493
566
 
494
- #### 6. Processing Footnotes/Endnotes
495
- If you kept notes inline (default behavior), you can extract them into a separate list for processing.
496
- ```javascript
567
+ ### Extract footnotes / endnotes
568
+ ```js
497
569
  function extractNotes(nodes) {
498
- let notes = [];
499
- nodes.forEach(node => {
500
- if (node.type === 'note') {
501
- notes.push({ id: node.metadata.noteId, text: node.text, type: node.metadata.noteType });
502
- }
503
- if (node.children) {
504
- notes = notes.concat(extractNotes(node.children));
505
- }
506
- });
507
- return notes;
570
+ return nodes.flatMap(n => [
571
+ ...(n.type === 'note' ? [{ id: n.metadata.noteId, text: n.text, type: n.metadata.noteType }] : []),
572
+ ...(n.children ? extractNotes(n.children) : [])
573
+ ]);
574
+ }
575
+ console.log(extractNotes(ast.content));
576
+ ```
577
+
578
+ ### Search for a term (TypeScript)
579
+ ```ts
580
+ import { OfficeParser } from 'officeparser';
581
+
582
+ async function contains(filePath: string, term: string): Promise<boolean> {
583
+ const ast = await OfficeParser.parseOffice(filePath);
584
+ return (await ast.to('text')).value.includes(term);
508
585
  }
586
+ ```
587
+
588
+ ---
589
+
590
+ ## Configuration Reference
591
+
592
+ ### OfficeParserConfig
593
+
594
+ Pass as the second argument to `parseOffice(file, config)`.
595
+
596
+ | Option | Type | Default | Description |
597
+ |--------|------|---------|-------------|
598
+ | `newlineDelimiter` | `string` | `'\n'` | Delimiter inserted between lines in text output |
599
+ | `ignoreNotes` | `boolean` | `false` | Ignore speaker notes (PPTX/ODP) |
600
+ | `putNotesAtLast` | `boolean` | `false` | Collect all notes at the end instead of inline |
601
+ | `extractAttachments` | `boolean` | `false` | Populate `ast.attachments` with Base64 images/charts |
602
+ | `ocr` | `boolean` | `false` | Run Tesseract OCR on images (requires `extractAttachments: true`) |
603
+ | `ocrConfig` | `OcrConfig` | `{}` | OCR worker pool settings — see [OCR section](#ocr-scheduler--resource-management) |
604
+ | `includeRawContent` | `boolean` | `false` | Attach raw XML/RTF source to each node |
605
+ | `serializeRawContent` | `boolean` | `true` | Re-serialize XML to clean strings (only if `includeRawContent: true`) |
606
+ | `preserveXmlWhitespace` | `boolean` | `false` | Preserve original XML whitespace during serialization |
607
+ | `includeBreakNodes` | `boolean` | `false` | Include `w:br` / `w:cr` as typed break nodes (DOCX only) |
608
+ | `ignoreInternalLinks` | `boolean` | `false` | Strip bookmarks and internal cross-references from AST |
609
+ | `fileType` | `SupportedFileType \| null` | `null` | **Required for text-based buffers** (`'md'`, `'html'`, `'csv'`) with no magic bytes |
610
+ | `csvDelimiter` | `string` | `','` | Input delimiter when parsing CSV files |
611
+ | `pdfWorkerSrc` | `string` | CDN (jsDelivr) | Path/URL to `pdf.worker.min.mjs` (required in browser) |
612
+ | `onWarning` | `(issue: OfficeIssue) => void` | — | Callback for non-fatal parsing issues |
613
+ | `outputErrorToConsole` | `boolean` | `false` | **Deprecated.** Use `onWarning` instead |
614
+
615
+ ---
616
+
617
+ ### GeneratorConfig (Common)
618
+
619
+ Options shared by all generator formats. Pass to `OfficeGenerator.generate(ast, format, config)` or `ast.to(format, config)`.
620
+
621
+ | Option | Type | Default | Description |
622
+ |--------|------|---------|-------------|
623
+ | `includeFormatting` | `boolean` | `true` | Include bold/italic/colors/sizes in output |
624
+ | `generateIds` | `boolean` | `true` | Add slug-based `id` attributes to headings |
625
+ | `renderMetadata` | `boolean` | `false` | Render title/author as visible header block |
626
+ | `includeImages` | `boolean` | `true` | Include image nodes in output |
627
+ | `includeCharts` | `boolean` | `true` | Include interactive charts (HTML only) |
628
+ | `ignoreInternalLinks` | `boolean` | `false` | Strip bookmarks and internal anchors from output |
629
+ | `ignoreDefaultStyleMap` | `boolean` | `false` | Disable built-in style mappings (e.g., "Heading 1" → h1) |
630
+ | `styleMap` | `string[] \| StructuredStyleMapping[]` | `[]` | Custom semantic style mappings |
631
+ | `onNode` | `(node) => string \| false \| void` | — | Per-node callback for filtering, overriding, or mutating |
632
+ | `onWarning` | `(issue: OfficeIssue) => void` | — | Callback for non-fatal generation issues |
633
+
634
+ ---
635
+
636
+ ### `onNode` Callback — Advanced Node Manipulation
637
+
638
+ Called for **every node** in the AST during generation. Can be `async`.
639
+
640
+ | Return value | Effect |
641
+ |---|---|
642
+ | `false` | Skip this node and all its children |
643
+ | `string` | Use this string as the output for this node, skip default logic |
644
+ | `void` | Proceed with default rendering (mutations to `node` are applied) |
509
645
 
510
- const allNotes = extractNotes(ast.content);
511
- console.log("Document Notes:", allNotes);
512
- ```
513
-
514
- ## Configuration Object: OfficeParserConfig
515
- Pass an optional config object as the second argument to `parseOffice`.
516
-
517
- | Flag | DataType | Default | Explanation |
518
- |------|----------|---------|-------------|
519
- | `outputErrorToConsole` | boolean | `false` | **Deprecated**: Use `onWarning` instead. Show logs to console in case of an error. |
520
- | `newlineDelimiter` | string | `\n` | Delimiter for new lines in text output. |
521
- | `ignoreNotes` | boolean | `false` | Ignore notes in files like PowerPoint/ODP. |
522
- | `putNotesAtLast` | boolean | `false` | Put notes text at the end of the document. |
523
- | `extractAttachments` | boolean | `false` | Extract images and charts as Base64. |
524
- | `includeRawContent` | boolean | `false` | Include raw XML/RTF markup in the nodes. |
525
- | `serializeRawContent` | boolean | `true` | Re-serializes raw XML to clean strings. |
526
- | `preserveXmlWhitespace` | boolean | `false` | Preserves original XML whitespace. |
527
- | `ocr` | boolean | `false` | Enable OCR for images (requires `extractAttachments: true`). |
528
- | `pdfWorkerSrc` | string | `(see below)` | Path to PDF.js worker. |
529
- | `ocrConfig` | object | `{}` | OCR Scheduler configuration. |
530
- | `includeBreakNodes` | boolean | `false` | Include `w:br`, `w:cr` nodes (DOCX only).|
531
- | `ignoreInternalLinks` | boolean | `false` | Remove all bookmarks and internal jumps. |
532
- | `csvDelimiter` | string | `,` | Custom delimiter for parsing CSV files. |
533
- | `fileType` | string | `null` | Manual format override (authoritative). |
534
-
535
- ## Generator Configuration: GeneratorConfig
536
- Configuration options for `OfficeGenerator.generate`.
537
-
538
- | Flag | DataType | Default | Explanation |
539
- |------|----------|---------|-------------|
540
- | `includeFormatting` | boolean | `false` | Whether to include semantic styles (bold, italic) in output. |
541
- | `styleMap` | string[] \| array | `[]` | Array of style mappings (DSL strings or structured objects). |
542
- | `ignoreDefaultStyleMap`| boolean | `false` | Ignore the library's default style mappings. |
543
- | `includeMetadata` | boolean | `false` | Include document metadata in the output (e.g., as frontmatter). |
544
- | `onNode` | function | `undefined` | Callback to intercept/modify any node during generation. |
545
-
546
- ### 🛠️ Advanced Node Manipulation (Pro Users)
547
- The `onNode` callback is a powerful tool that gives you complete control over the generation process. It is called for **every single node** in the AST before it is rendered.
548
-
549
- #### Callback Capabilities:
550
- 1. **Filter/Remove Nodes**: Return `false` to skip a node and all its children.
551
- 2. **Override Rendering**: Return a `string` to use that exact text as the output, bypassing default logic and recursion.
552
- 3. **Mutate Nodes**: Modify the `node` object directly (e.g., changing `node.text`) and return `void` to let the generator proceed with your changes.
553
- 4. **Async Support**: The callback can be `async`, allowing you to fetch external data or perform complex logic during generation.
554
-
555
- #### Pro Example:
556
- ```typescript
557
- const result = await ast.to('md', {
646
+ ```ts
647
+ const { value: md } = await ast.to('md', {
558
648
  onNode: async (node) => {
559
- // 1. Skip all images
649
+ // Skip all images
560
650
  if (node.type === 'image') return false;
561
651
 
562
- // 2. Redact sensitive info by mutating the node
652
+ // Redact secrets (mutate then proceed)
563
653
  if (node.text?.includes('SECRET_KEY')) {
564
654
  node.text = node.text.replace(/SECRET_KEY: \w+/, 'SECRET_KEY: [REDACTED]');
565
655
  }
566
656
 
567
- // 3. Custom rendering for specific styles
657
+ // Custom rendering for a specific style
568
658
  if (node.metadata?.style === 'Callout') {
569
659
  return `> [!INFO]\n> ${node.text}`;
570
660
  }
571
-
572
- // 4. Proceed with default rendering (implicitly returns void)
573
661
  }
574
662
  });
575
663
  ```
576
664
 
577
- ### Advanced Style Mapping (Semantic Translation)
578
- The `styleMap` configuration is the primary way to define the "semantic meaning" of document styles. We recommend using **Structured Style Mappings** for full type safety and power.
665
+ ---
666
+
667
+ ### `styleMap` — Semantic Style Mapping
579
668
 
580
- #### 1. Structured Style Mappings (Recommended)
581
- Use structured objects to match nodes based on type and attributes, and specify detailed output properties like classes and custom attributes.
669
+ Maps document style names to semantic output elements. Two formats supported:
582
670
 
583
- ```typescript
671
+ #### Structured Objects (Recommended)
672
+
673
+ ```ts
584
674
  styleMap: [
585
- {
586
- selector: {
587
- nodeType: 'paragraph',
588
- attributes: { style: 'Heading 1' }
589
- },
590
- output: {
591
- tag: 'h1',
592
- classes: ['main-title'],
593
- attributes: { id: 'top' }
594
- }
675
+ {
676
+ selector: { nodeType: 'paragraph', attributes: { style: 'Heading 1' } },
677
+ output: { tag: 'h1', classes: ['main-title'], attributes: { id: 'top' } }
595
678
  },
596
679
  {
597
- // Use operators like '~=' for partial matches
680
+ // '~=' operator matches if the word 'Quote' appears anywhere in the style name
598
681
  selector: { attributes: { style: { value: 'Quote', operator: '~=' } } },
599
- output: { tag: 'blockquote' }
682
+ output: { tag: 'blockquote', fresh: true }
600
683
  }
601
684
  ]
602
685
  ```
603
686
 
604
- #### 2. Legacy String DSL
605
- The library also maintains support for a simple string-based DSL, highly compatible with `mammoth.js`.
687
+ `fresh: true` prevents the generator from merging adjacent nodes of the same tag into one block.
606
688
 
607
- - **Literal Matching**: `"p[style-name='Heading 1'] => h1"`
608
- - **Regex-like Matching**: `"p[style~='Title'] => h2"`
609
- - **Attribute Filters**: `"p[style-name='Quote'][lang='en'] => blockquote"`
689
+ #### Legacy String DSL
610
690
 
611
- ## Chunking Configuration: ChunkingConfig
612
- Specific options when using `format: 'chunks'`.
691
+ Compatible with `mammoth.js` style maps:
613
692
 
614
- | Flag | DataType | Default | Explanation |
615
- |------|----------|---------|-------------|
616
- | `strategy` | string | `'fixed-size'`| The chunking strategy (`fixed-size`, `document-structure`, `semantic`). |
617
- | `maxChunkSize` | number | `1000` | Maximum characters per chunk. |
618
- | `chunkOverlap` | number | `200` | Overlap between consecutive chunks. |
619
- | `similarityThreshold`| number | `0.5` | Threshold for semantic splitting (0.0 to 1.0). |
620
- | `embedBatchSize` | number | `50` | Batch size for embedding requests. |
693
+ ```js
694
+ styleMap: [
695
+ "p[style-name='Heading 1'] => h1",
696
+ "p[style~='Title'] => h2",
697
+ "p[style-name='Quote'][lang='en'] => blockquote"
698
+ ]
699
+ ```
621
700
 
622
- ### OCR Scheduler & Resource Management
623
- If your application uses OCR, `officeParser` utilizes an intelligent **Smart Worker Pool** to maintain a background worker pool and optimize repeated parse requests.
701
+ ---
624
702
 
625
- - **Dynamic Affinity**: Workers in the pool persist with their last used language affinity.
626
- - **LRU Re-allocation**: If a new language is requested and the pool is full, the manager identifies the **Least Recently Used (LRU)** idle worker and re-initializes it for the new language. This avoids the overhead of destroying and recreating workers.
627
- - **Auto-Termination**: Workers are automatically cleaned up after 10 seconds of inactivity (configurable via `ocrConfig.autoTerminateTimeout`).
703
+ ### HtmlGeneratorConfig
628
704
 
629
- #### `OfficeParser.terminateOcr()`
630
- If you have used OCR (`{ ocr: true }`) in a short-lived script (like CLI tools or one-off automation), we recommend explicitly calling `terminateOcr()` after your processing is finished. This bypasses the 10-second idle timer and allows the process to return to the terminal prompt immediately.
705
+ Pass as `htmlConfig` inside `GeneratorConfig`.
631
706
 
632
- > [!NOTE]
633
- > If OCR was not used, this function is a no-op and does not need to be called.
707
+ | Option | Type | Default | Description |
708
+ |--------|------|---------|-------------|
709
+ | `standalone` | `boolean` | `true` | Wrap output in a full `<html>` document with CSS |
710
+ | `chartJsSrc` | `string` | jsDelivr CDN | URL for the Chart.js library |
634
711
 
635
- ```js
636
- const officeParser = require('officeparser');
712
+ ### MdGeneratorConfig
637
713
 
638
- async function runCleaner() {
639
- await officeParser.parseOffice("file.pdf", { ocr: true });
640
- // ... process results ...
714
+ Pass as `mdConfig` inside `GeneratorConfig`.
641
715
 
642
- // Manually kill OCR workers for an immediate exit
643
- await officeParser.terminateOcr();
644
- }
645
- ```
716
+ | Option | Type | Default | Description |
717
+ |--------|------|---------|-------------|
718
+ | `fallbackToHtml` | `boolean` | `true` | Use HTML tags for features Markdown cannot represent (underlines, merged table cells, etc.) |
646
719
 
647
- > [!TIP]
648
- > This is handled automatically in the built-in CLI (`npx officeparser ...`). You only need to call this manually if you are using the library in your own custom script and want a snappy exit.
720
+ ### PdfGeneratorConfig
649
721
 
650
- ```js
651
- const config = {
652
- newlineDelimiter: "\n\n",
653
- extractAttachments: true,
654
- ocr: true,
655
- ocrLanguage: 'eng+fra+esp' // Supports English, French, and Spanish simultaneously
656
- };
722
+ Pass as `pdfConfig` inside `GeneratorConfig`. Requires the optional `puppeteer` peer dependency.
657
723
 
658
- const ast = await officeParser.parseOffice("report.docx", config);
659
- console.log(`Extracted ${ast.attachments.length} images`);
660
- ```
724
+ | Option | Type | Default | Description |
725
+ |--------|------|---------|-------------|
726
+ | `format` | `string` | `'A4'` | Paper format (`'A4'`, `'Letter'`, `'Legal'`, etc.) |
727
+ | `landscape` | `boolean` | `false` | Landscape page orientation |
728
+ | `printBackground` | `boolean` | `true` | Print background graphics |
729
+ | `margin` | `object` | `{0,0,0,0}` | Page margins (`top`, `right`, `bottom`, `left`) |
730
+ | `displayHeaderFooter` | `boolean` | `false` | Show print header/footer |
731
+ | `headerTemplate` | `string` | `''` | HTML template for the print header |
732
+ | `footerTemplate` | `string` | `''` | HTML template for the print footer |
733
+ | `scale` | `number` | `1` | Rendering scale factor |
734
+ | `launchOptions` | `object` | headless defaults | Puppeteer launch options (e.g., `executablePath`) |
661
735
 
662
- ## Examples
736
+ ### CsvGeneratorConfig
663
737
 
664
- **Search for a term in a document (TypeScript)**
665
- ```ts
666
- import { OfficeParser } from 'officeparser';
738
+ Pass as `csvConfig` inside `GeneratorConfig`.
667
739
 
668
- async function hasSearchTerm(filePath: string, term: string): Promise<boolean> {
669
- const ast = await OfficeParser.parseOffice(filePath);
670
- return ast.toText().includes(term);
671
- }
672
- ```
740
+ | Option | Type | Default | Description |
741
+ |--------|------|---------|-------------|
742
+ | `sheets` | `string` | `''` | Sheet range to export: `'1'`, `'1-3'`, `'1,3'` (1-based). Empty = all sheets |
743
+ | `mergeSheets` | `boolean` | `true` | Merge all sheets into one CSV. If `false`, returns a ZIP archive |
744
+ | `columnDelimiter` | `string` | `','` | Output column delimiter |
745
+
746
+ ### TextGeneratorConfig
747
+
748
+ Pass as `textConfig` inside `GeneratorConfig`.
749
+
750
+ | Option | Type | Default | Description |
751
+ |--------|------|---------|-------------|
752
+ | `newlineDelimiter` | `string` | `'\n'` | String inserted between structural blocks |
753
+ | `preserveLayout` | `boolean` | `false` | Render tables with aligned columns using whitespace |
754
+
755
+ ---
756
+
757
+ ### OfficeConverterConfig
758
+
759
+ Configuration for `OfficeConverter.convert(file, format, config)`.
760
+
761
+ | Option | Type | Description |
762
+ |--------|------|-------------|
763
+ | `parseConfig` | `OfficeParserConfig` | Settings for the parsing phase |
764
+ | `generatorConfig` | `GeneratorConfig` | Settings for the generation phase |
765
+ | `onWarning` | `(issue: OfficeIssue) => void` | Global warning callback (overrides phase-specific ones) |
766
+
767
+ ---
768
+
769
+ ### ChunkingConfig
770
+
771
+ `ChunkingConfig` is a **discriminated union** — the available options depend on the `strategy` field.
772
+
773
+ #### Common Options (all strategies)
774
+
775
+ | Option | Type | Default | Description |
776
+ |--------|------|---------|-------------|
777
+ | `strategy` | `string` | `'document-structure'` | Chunking strategy |
778
+ | `stripWhitespace` | `boolean` | `true` | Trim leading/trailing whitespace from each chunk |
779
+ | `includeMetadata` | `boolean` | `true` | Include page/slide/heading metadata in each chunk |
780
+ | `addStartIndex` | `boolean` | `false` | Add `startIndex` character offset to chunk metadata |
781
+ | `lengthFunction` | `(text) => number` | `text.length` | Custom size measurer (e.g., token counter) |
782
+ | `sentenceBoundaryRegex` | `string \| RegExp` | `/[.!?。!?]/` | Custom regex for sentence boundary detection |
783
+ | `abbreviations` | `string[]` | common list | Abbreviations to skip when splitting on `.` |
784
+
785
+ #### `strategy: 'fixed-size'`
786
+
787
+ | Option | Type | Default | Description |
788
+ |--------|------|---------|-------------|
789
+ | `chunkSize` | `number` | `1000` | Maximum characters per chunk |
790
+ | `chunkOverlap` | `number` | `200` | Character overlap between consecutive chunks |
791
+ | `separators` | `string[]` | `['\n\n','\n',' ','']` | Ordered list of separators to try |
792
+
793
+ #### `strategy: 'document-structure'`
794
+
795
+ | Option | Type | Default | Description |
796
+ |--------|------|---------|-------------|
797
+ | `splitBy` | `string` | `'paragraph'` | `'paragraph'` · `'heading'` · `'page'` · `'slide'` · `'sheet'` |
798
+ | `maxChunkSize` | `number` | `1000` | Max characters per chunk (oversized units are split recursively) |
799
+ | `tableSplitStrategy` | `string` | `'row'` | `'row'` (repeats header in each chunk) or `'flatten'` |
800
+
801
+ #### `strategy: 'semantic'`
802
+
803
+ | Option | Type | Default | Description |
804
+ |--------|------|---------|-------------|
805
+ | `embeddingFunction` | `(text) => Promise<number[]>` | **required** | Async embedding function |
806
+ | `similarityThreshold` | `number` | `0.8` | Cosine similarity threshold; lower = fewer boundaries |
807
+ | `maxChunkSize` | `number` | `2000` | Max characters even if similarity stays high |
808
+ | `bufferSize` | `number` | `1` | Surrounding sentences used when computing similarity |
809
+ | `embeddingBatchSize` | `number` | `50` | Sentences per embedding API batch |
810
+
811
+ ---
812
+
813
+ ## OCR Scheduler & Resource Management
814
+
815
+ When `ocr: true` is set, `officeParser` maintains an intelligent **Smart Worker Pool** backed by Tesseract.js:
816
+
817
+ - **Dynamic Affinity**: Workers persist with their last-used language, avoiding re-initialization overhead.
818
+ - **LRU Re-allocation**: When a new language is requested and the pool is full, the Least Recently Used idle worker is re-initialized.
819
+ - **Auto-Termination**: Workers shut down after 10 seconds of inactivity (configurable via `ocrConfig.autoTerminateTimeout`).
820
+
821
+ ### OCR Config (`ocrConfig`)
822
+
823
+ | Option | Type | Default | Description |
824
+ |--------|------|---------|-------------|
825
+ | `language` | `string` | `'eng'` | Tesseract language code(s), e.g. `'eng+fra'` |
826
+ | `workerPath` | `string` | `''` | Custom path to Tesseract worker script |
827
+ | `corePath` | `string` | `''` | Custom path to Tesseract core script |
828
+ | `langPath` | `string` | `''` | Custom path for language data files |
829
+ | `autoTerminateTimeout` | `number` | `10000` | Inactivity timeout in ms before auto-teardown (0 = disabled) |
830
+
831
+ See all language codes at [tesseract-ocr.github.io](https://tesseract-ocr.github.io/tessdoc/Data-Files).
832
+
833
+ ### `OfficeParser.terminateOcr()`
834
+
835
+ In **short-lived scripts** (CLI tools, one-off automation), call `terminateOcr()` after processing to bypass the idle timer and exit immediately:
673
836
 
674
- **Extracting Images and their OCR text**
675
837
  ```js
676
838
  const officeParser = require('officeparser');
677
839
 
678
- const config = { extractAttachments: true, ocr: true };
679
- officeParser.parseOffice("presentation.pptx", config).then(ast => {
680
- ast.attachments.forEach(attachment => {
681
- if (attachment.type === 'image') {
682
- console.log(`Image: ${attachment.name}`);
683
- console.log(`OCR Text: ${attachment.ocrText}`);
684
- fs.writeFileSync(attachment.name, Buffer.from(attachment.data, 'base64'));
685
- }
686
- });
687
- });
840
+ const ast = await officeParser.parseOffice('file.pdf', { ocr: true });
841
+ // ... process results ...
842
+ await officeParser.terminateOcr(); // immediate exit
688
843
  ```
689
844
 
845
+ > [!TIP]
846
+ > The built-in CLI (`npx officeparser ...`) handles this automatically.
847
+ > Only call it manually in your own scripts.
848
+
849
+ ---
850
+
690
851
  ## Browser Usage
691
- The library provides two types of browser bundles in the `dist/` directory:
692
- 1. **`officeparser.browser.iife.js`**: Standard IIFE bundle for direct `<script>` tag usage. Exposes the global `officeParser` namespace.
693
- 2. **`officeparser.browser.mjs`**: Modern ESM bundle for use with `import` statements or modern bundlers.
694
852
 
695
- ### Usage (ESM)
696
- If you are using a modern bundler like **Vite**, **Webpack**, or **Next.js**:
853
+ Two bundles are available in the `dist/` directory:
697
854
 
698
- ```javascript
855
+ | Bundle | Usage |
856
+ |--------|-------|
857
+ | `officeparser.browser.mjs` | ESM — use with `import` statements or modern bundlers (Vite, Webpack, Next.js) |
858
+ | `officeparser.browser.iife.js` | IIFE — use with a `<script>` tag; exposes the global `officeParser` object |
859
+
860
+ ### ESM (Vite / Webpack / Next.js)
861
+
862
+ ```js
699
863
  import { OfficeParser } from 'officeparser';
700
864
 
701
865
  const handleFile = async (event) => {
702
866
  const file = event.target.files[0];
703
867
  const buffer = await file.arrayBuffer();
704
-
705
- try {
706
- // Pass the Buffer or Uint8Array directly
707
- const ast = await OfficeParser.parseOffice(new Uint8Array(buffer));
708
- console.log(ast.toText());
709
- } catch (err) {
710
- console.error(err);
711
- }
868
+ const ast = await OfficeParser.parseOffice(new Uint8Array(buffer));
869
+ console.log(ast.toText());
712
870
  };
713
871
  ```
714
872
 
715
- > [!NOTE]
716
- > **Why `fs` fails in the browser**: Browsers do not have a built-in file system. If you try to pass a file path string in the browser, `officeParser` will throw a descriptive "Fail-Fast" error instead of crashing mysteriously:
717
- > `[officeparser] Node.js 'fs' module is not available in the browser. Please pass a Buffer or Uint8Array instead.`
718
-
719
- ### Usage (Script Tag)
720
- Include the IIFE bundle available in the release assets or your `dist/` folder. This exposes the global `officeParser` object.
873
+ ### Script Tag
721
874
 
722
875
  ```html
723
876
  <script src="dist/officeparser.browser.iife.js"></script>
@@ -725,60 +878,78 @@ Include the IIFE bundle available in the release assets or your `dist/` folder.
725
878
  async function handleFile(event) {
726
879
  const file = event.target.files[0];
727
880
  const buffer = await file.arrayBuffer();
728
-
729
- try {
730
- // Reconstruct as Uint8Array for the parser
731
- const ast = await officeParser.parseOffice(new Uint8Array(buffer));
732
- console.log(ast.toText());
733
- } catch (error) {
734
- console.error("Parsing failed:", error);
735
- }
881
+ const ast = await officeParser.parseOffice(new Uint8Array(buffer));
882
+ console.log(ast.toText());
736
883
  }
737
884
  </script>
738
885
  ```
739
886
 
740
- ### PDF Worker Configuration in Browser
741
- When using `officeparser` in a browser environment to parse PDF files, you may provide the `pdfWorkerSrc` configuration option. If not provided, it defaults to a CDN link for `pdfjs-dist@5.6.205`.
887
+ > [!NOTE]
888
+ > **File paths don't work in the browser.** Always pass a `Buffer`, `ArrayBuffer`, or `Uint8Array`.
889
+ > Passing a path string will throw a descriptive `FEATURE_NOT_SUPPORTED_IN_BROWSER` error.
742
890
 
743
- ```javascript
744
- const file = ...; // File object or ArrayBuffer
891
+ ### PDF Worker Configuration
745
892
 
746
- // It will use the default CDN worker if pdfWorkerSrc is omitted
747
- const ast = await officeParser.parseOffice(file);
893
+ When parsing PDFs in the browser, a Web Worker is required. If `pdfWorkerSrc` is omitted, a jsDelivr CDN link is used automatically:
748
894
 
749
- // Or override it with your own path or a different version:
750
- const ast2 = await officeParser.parseOffice(file, {
751
- pdfWorkerSrc: "https://cdn.jsdelivr.net/npm/pdfjs-dist@5.6.205/build/pdf.worker.min.mjs"
895
+ ```js
896
+ // Uses default CDN worker:
897
+ const ast = await officeParser.parseOffice(pdfArrayBuffer);
898
+
899
+ // Or specify your own:
900
+ const ast = await officeParser.parseOffice(pdfArrayBuffer, {
901
+ pdfWorkerSrc: 'https://cdn.jsdelivr.net/npm/pdfjs-dist@5.6.205/build/pdf.worker.min.mjs'
752
902
  });
753
903
  ```
754
904
 
755
- > **Note:** The version of `pdfjs-dist` in the worker source should match the version used by `officeparser` (currently `5.6.205`).
905
+ > [!NOTE]
906
+ > The `pdfjs-dist` worker version must match the version bundled with `officeparser` (currently **5.6.205**).
907
+
908
+ ---
756
909
 
757
910
  ## Troubleshooting & Common Issues
758
911
 
759
- - **Node.js process stays alive after finishing**: If using OCR, the worker pool stays warm for 10s by default. Use `await terminateOcr()` at the end of your script for a snappy exit.
760
- - **"Worker not found" in Browser**: Ensure `pdfWorkerSrc` is correctly pointed to the `pdf.worker.min.mjs` file matching version `5.6.205`.
761
- - **OCR accuracy is low**: Verify your `ocrConfig.language` matches the document content. Note that OCR quality depends on image resolution.
762
- - **Out of memory on large files**: For massive spreadsheets, consider using `ast.toText()` early and allowing the full AST object to be garbage-collected.
912
+ | Symptom | Fix |
913
+ |---------|-----|
914
+ | Node.js process stays alive after finishing | Call `await officeParser.terminateOcr()` at end of script when OCR was used |
915
+ | `"Worker not found"` in browser for PDF | Verify `pdfWorkerSrc` points to `pdf.worker.min.mjs` matching version `5.6.205` |
916
+ | Low OCR accuracy | Verify `ocrConfig.language` matches the document language; quality depends on image resolution |
917
+ | Out of memory on large Excel files | Call `ast.toText()` early and discard the AST object to allow garbage collection |
918
+ | `md`/`html`/`csv` buffer not detected | Add `fileType: 'md'` (or `'html'`, `'csv'`) to config — these formats have no magic bytes |
919
+ | `IMPROPER_BUFFERS` error | Usually means no file extension and no `fileType` hint was provided for a buffer input |
920
+ | PDF generation fails | Install the optional peer dependency: `npm install puppeteer` |
763
921
 
764
- For a comprehensive guide, visit our [Debugging & Troubleshooting Documentation](https://harshankur.github.io/officeParser/#spec/debugging).
922
+ For a full debugging guide, visit the [Live Documentation](https://harshankur.github.io/officeParser/#spec/debugging).
765
923
 
924
+ ---
766
925
 
767
926
  ## Known Limitations
768
- 1. **ODT/ODS Charts**: Extraction may occasionally show inaccurate data when referencing external cell ranges or complex layout-based data.
769
- 2. **PDF Images**: PDF images are extracted as BMP files in the browser for compatibility. This conversion happens automatically.
770
- 3. **RTF Footnotes**: The `putNotesAtLast` configuration is currently not supported for RTF files; footnotes and endnotes are always collected and appended to the end of the content.
771
927
 
772
- ----------
928
+ 1. **ODT/ODS Charts**: May show inaccurate data when the chart references external cell ranges or uses complex layout-based data.
929
+ 2. **PDF Images (Browser)**: Extracted as BMP files for cross-platform compatibility. Conversion is automatic.
930
+ 3. **RTF Notes**: `putNotesAtLast` has no effect for RTF files; footnotes and endnotes are always appended at the end.
931
+
932
+ ---
773
933
 
774
934
  **npm**: [https://npmjs.com/package/officeparser](https://npmjs.com/package/officeparser)
775
935
 
776
936
  **github**: [https://github.com/harshankur/officeParser](https://github.com/harshankur/officeParser)
777
937
 
938
+ ## Support the Project
939
+
940
+ If `officeParser` has helped you save time, consider supporting its continued development. Your sponsorship helps maintain the project, add new features, and keep it robust for everyone.
941
+
942
+ <a href="https://github.com/sponsors/harshankur">
943
+ <img src="https://img.shields.io/badge/Sponsor-GitHub-ea4aaa?style=for-the-badge&logo=github-sponsors" height="36">
944
+ </a>
945
+ <a href="https://www.buymeacoffee.com/harshankur">
946
+ <img src="https://cdn.buymeacoffee.com/buttons/v2/default-yellow.png" height="36" alt="Buy Me A Coffee">
947
+ </a>
948
+
778
949
  ## Contributing
779
950
 
780
- Contributions are welcome! Please see [CONTRIBUTING.md](CONTRIBUTING.md) for details on how to get started.
951
+ Contributions are welcome! Please see [CONTRIBUTING.md](CONTRIBUTING.md) for details.
781
952
 
782
953
  ## License
783
954
 
784
- This project is licensed under the MIT License - see the [LICENSE](LICENSE) file for details.
955
+ This project is licensed under the MIT License — see the [LICENSE](LICENSE) file for details.