officeparser 7.0.1 → 7.0.3

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -1,6 +1,10 @@
1
- # officeParser 📄🚀 - The Most Versatile Office Parser & Generator
1
+ # officeParser — Universal Office Document Parser & Generator
2
2
 
3
- A robust, strictly-typed Node.js and Browser library for parsing and generating office files. It not only extracts content from [`docx`](https://en.wikipedia.org/wiki/Office_Open_XML), [`pptx`](https://en.wikipedia.org/wiki/Office_Open_XML), [`xlsx`](https://en.wikipedia.org/wiki/Office_Open_XML), [`odt`](https://en.wikipedia.org/wiki/OpenDocument), [`odp`](https://en.wikipedia.org/wiki/OpenDocument), [`ods`](https://en.wikipedia.org/wiki/OpenDocument), [`pdf`](https://en.wikipedia.org/wiki/PDF), [`rtf`](https://en.wikipedia.org/wiki/Rich_Text_Format), [`csv`](https://en.wikipedia.org/wiki/Comma-separated_values), [`md`](https://en.wikipedia.org/wiki/Markdown), and [`html`](https://en.wikipedia.org/wiki/HTML) into a rich Abstract Syntax Tree (AST), but also provides a powerful generation engine to convert that AST into formats like **Markdown**, **HTML**, **CSV**, **RTF**, **Text**, **PDF**, and **JSON**, including native **RAG-focused chunking** support.
3
+ A robust, strictly-typed **Node.js and Browser** library for parsing office files into a rich **Abstract Syntax Tree (AST)** and generating high-fidelity output in multiple formats.
4
+
5
+ **Parses:** [`docx`](https://en.wikipedia.org/wiki/Office_Open_XML) · [`pptx`](https://en.wikipedia.org/wiki/Office_Open_XML) · [`xlsx`](https://en.wikipedia.org/wiki/Office_Open_XML) · [`odt`](https://en.wikipedia.org/wiki/OpenDocument) · [`odp`](https://en.wikipedia.org/wiki/OpenDocument) · [`ods`](https://en.wikipedia.org/wiki/OpenDocument) · [`pdf`](https://en.wikipedia.org/wiki/PDF) · [`rtf`](https://en.wikipedia.org/wiki/Rich_Text_Format) · [`csv`](https://en.wikipedia.org/wiki/Comma-separated_values) · [`md`](https://en.wikipedia.org/wiki/Markdown) · [`html`](https://en.wikipedia.org/wiki/HTML)
6
+
7
+ **Generates:** `Markdown` · `HTML` · `CSV` · `RTF` · `PDF` · `Plain Text` · `RAG Chunks`
4
8
 
5
9
  [![npm version](https://badge.fury.io/js/officeparser.svg)](https://badge.fury.io/js/officeparser)
6
10
  [![Total Downloads](https://img.shields.io/npm/dt/officeparser.svg)](https://www.npmjs.com/package/officeparser)
@@ -10,45 +14,53 @@ A robust, strictly-typed Node.js and Browser library for parsing and generating
10
14
  ---
11
15
 
12
16
  ### 🌟 [Live Interactive AST Visualizer & Documentation](https://harshankur.github.io/officeParser/) 🌟
13
- *Test any office file in your browser and see the extracted AST, text, and preview in real-time which is rebuilt from the AST!*
14
-
15
- **What you can do there:**
16
- - **AST Visualizer**: Upload any office file and inspect the hierarchical AST structure, metadata, and raw content.
17
- - **Config Configurator**: Tweak parsing options (like `ignoreNotes`, `ocr`, `newlineDelimiter`) and see the results instantly.
18
- - **Debugging**: Use the visualizer to debug parsing issues by inspecting exactly how nodes are interpreted.
19
- - **Format Specs**: Read detailed specifications for the AST structure and configuration options.
20
-
21
- ---
17
+ *Upload any office file in your browser — inspect the AST, tweak config, and preview generated output in real-time.*
22
18
 
19
+ - **AST Visualizer**: Inspect the hierarchical node tree, metadata, and raw content
20
+ - **Config Configurator**: Tweak options (`ignoreNotes`, `ocr`, `newlineDelimiter`) and see results instantly
21
+ - **Debugging**: Identify exactly how nodes are interpreted
22
+ - **Format Specs**: Read detailed specs for the AST structure and all config options
23
23
 
24
24
  ---
25
25
 
26
26
  ### 📝 [Changelog](CHANGELOG.md)
27
- *Detailed release notes and the full history of updates are available in the project changelog.*
28
27
 
29
28
  ---
30
29
 
31
30
  ## Table of Contents
32
- - [Install via npm](#install-via-npm)
31
+ - [Install](#install-via-npm)
33
32
  - [Command Line Usage](#command-line-usage)
34
- - [Library Usage](#library-usage)
35
- - [Using the OfficeGenerator](#using-the-officegenerator)
36
- - [The New "One-Step" API: OfficeConverter](#the-new-one-step-api-officeconverter)
33
+ - [Quick Decision Guide](#quick-decision-guide)
34
+ - [Library Usage: Parsing](#library-usage-parsing)
35
+ - [Async/Await](#asyncawait)
36
+ - [Callback (Backward Compat)](#callback-backward-compat)
37
+ - [File Buffers & ArrayBuffers](#file-buffers--arraybuffers)
38
+ - [`ast.to()` — Generate from AST](#astto--generate-from-ast)
39
+ - [`ast.toText()` — Quick Text Extraction](#asttotext--quick-text-extraction)
40
+ - [OfficeGenerator](#officegenerator)
41
+ - [OfficeConverter — One-Step API](#officeconverter--one-step-api)
37
42
  - [Native RAG Chunking](#native-rag-chunking)
38
43
  - [The AST Structure](#the-ast-structure)
39
44
  - [Deep Dive: Document Components](#deep-dive-document-components)
40
- - [Performance & Fidelity Highlights (v7.0.0)](#performance--fidelity-highlights-v700)
45
+ - [Performance Highlights](#performance-highlights)
41
46
  - [Advanced AST Usage](#advanced-ast-usage)
42
- - [Configuration Object: OfficeParserConfig](#configuration-object-officeparserconfig)
43
- - [Generator Configuration: GeneratorConfig](#generator-configuration-generatorconfig)
44
- - [Format-Specific Generator Configuration](#format-specific-generator-configuration)
45
- - [One-Step Conversion: OfficeConverterConfig](#one-step-conversion-officeconverterconfig)
46
- - [Chunking Configuration: ChunkingConfig](#chunking-configuration-chunkingconfig)
47
+ - [Configuration Reference](#configuration-reference)
48
+ - [OfficeParserConfig](#officeparserconfig)
49
+ - [GeneratorConfig (Common)](#generatorconfig-common)
50
+ - [onNode Callback](#onnode-callback--advanced-node-manipulation)
51
+ - [styleMap — Semantic Style Mapping](#stylemap--semantic-style-mapping)
52
+ - [HtmlGeneratorConfig](#htmlgeneratorconfig)
53
+ - [MdGeneratorConfig](#mdgeneratorconfig)
54
+ - [PdfGeneratorConfig](#pdfgeneratorconfig)
55
+ - [CsvGeneratorConfig](#csvgeneratorconfig)
56
+ - [TextGeneratorConfig](#textgeneratorconfig)
57
+ - [OfficeConverterConfig](#officeconverterconfig)
58
+ - [ChunkingConfig](#chunkingconfig)
47
59
  - [OCR Scheduler & Resource Management](#ocr-scheduler--resource-management)
48
- - [Examples](#examples)
49
60
  - [Browser Usage](#browser-usage)
50
61
  - [Troubleshooting & Common Issues](#troubleshooting--common-issues)
51
62
  - [Known Limitations](#known-limitations)
63
+ - [Contributing](#contributing)
52
64
 
53
65
  ---
54
66
 
@@ -58,748 +70,807 @@ A robust, strictly-typed Node.js and Browser library for parsing and generating
58
70
  npm i officeparser
59
71
  ```
60
72
 
61
- ## Command Line usage
62
- You can use `officeparser` directly from the terminal to extract content as JSON AST, plain text, or generate new formats like Markdown and HTML.
73
+ ---
74
+
75
+ ## Command Line Usage
63
76
 
64
77
  ```bash
65
- # Get full AST as JSON (default)
66
- npx officeparser /path/to/officeFile.docx
78
+ # Full AST as JSON (default)
79
+ npx officeparser /path/to/file.docx
67
80
 
68
- # Get plain text only
69
- npx officeparser /path/to/officeFile.docx --toText=true
81
+ # Plain text output
82
+ npx officeparser /path/to/file.docx --format=text
70
83
 
71
- # Generate Markdown file
84
+ # Convert DOCX to Markdown and save
72
85
  npx officeparser report.docx --format=md --output=report.md
73
86
 
74
- # Generate HTML file with specific output
87
+ # Convert PPTX to HTML
75
88
  npx officeparser presentation.pptx --format=html --output=preview.html
76
89
 
77
- # Convert spreadsheet to CSV
90
+ # Convert XLSX to CSV
78
91
  npx officeparser data.xlsx --format=csv
92
+
93
+ # Generate RAG chunks
94
+ npx officeparser document.pdf --format=chunks
79
95
  ```
80
96
 
81
- ### Config Options:
82
- - `--format=[json|text|md|html|csv|rtf|pdf|chunks]` The output format. Default is `json`.
83
- - `--output=[path]` Optional file path to write the output to.
84
- - `--toText=[true|false]` Legacy flag to output only plain text. Use `--format=text` instead.
85
- - `--ignoreNotes=[true|false]` Flag to ignore notes from files like PowerPoint. Default is false.
86
- - `--newlineDelimiter=[delimiter]` The delimiter to use for new lines. Default is `\n`.
87
- - `--putNotesAtLast=[true|false]` Flag to collect notes at the end of files like PowerPoint. Default is false.
88
- - `--outputErrorToConsole=[true|false]` **(Deprecated)** Flag to output errors to the console. Use `onWarning` callback in library usage.
89
- - `--extractAttachments=[true|false]` Flag to extract images/charts as Base64. Default is false.
90
- - `--ocr=[true|false]` Flag to enable OCR for extracted images. Default is false.
91
- - `--includeRawContent=[true|false]` Flag to include raw XML/RTF content in nodes. Default is false.
92
- - `--includeBreakNodes=[true|false]` Flag to include break nodes. Currently only available for DOCX documents.
93
- - `--verbose=[true|false]` Show full error stack traces.
97
+ ### CLI Options
94
98
 
99
+ | Flag | Values | Default | Description |
100
+ |------|--------|---------|-------------|
101
+ | `--format` | `json\|text\|md\|html\|csv\|rtf\|pdf\|chunks` | `json` | Output format |
102
+ | `--output` | path | — | Write output to a file |
103
+ | `--toText` | `true\|false` | `false` | **Deprecated.** Use `--format=text` |
104
+ | `--ignoreNotes` | `true\|false` | `false` | Ignore speaker notes (PPTX/ODP) |
105
+ | `--putNotesAtLast` | `true\|false` | `false` | Collect notes at end of output |
106
+ | `--newlineDelimiter` | string | `\n` | Delimiter between lines |
107
+ | `--extractAttachments` | `true\|false` | `false` | Extract images/charts as Base64 |
108
+ | `--ocr` | `true\|false` | `false` | Enable OCR for images |
109
+ | `--includeRawContent` | `true\|false` | `false` | Include raw XML/RTF in nodes |
110
+ | `--includeBreakNodes` | `true\|false` | `false` | Include break nodes (DOCX only) |
111
+ | `--outputErrorToConsole` | `true\|false` | `false` | **Deprecated.** Use `onWarning` callback |
112
+ | `--verbose` | `true\|false` | `false` | Show full error stack traces |
95
113
 
96
- ## Library Usage
97
- In **v7.0.0**, the library has evolved into a dual-purpose **Parser** and **Generator**. You can first parse any office file into a structured AST and then use the `OfficeGenerator` to transform that AST into various formats or chunks.
114
+ ---
115
+
116
+ ## Quick Decision Guide
117
+
118
+ | Goal | API to use |
119
+ |------|-----------|
120
+ | Extract text / AST from a file | `OfficeParser.parseOffice(file)` |
121
+ | Convert directly to another format | `OfficeConverter.convert(file, 'md')` |
122
+ | Parse first, then generate | `parseOffice()` → `OfficeGenerator.generate(ast, 'html')` |
123
+ | Convert on the AST itself (shorthand) | `ast.to('md')` |
124
+ | RAG pipeline chunking | `OfficeConverter.convert(file, 'chunks', {...})` |
125
+
126
+ ---
127
+
128
+ ## Library Usage: Parsing
129
+
130
+ ### Async/Await
98
131
 
99
- ### Getting Started (Async/Await)
100
132
  ```js
101
133
  const officeParser = require('officeparser');
102
134
 
103
- async function parseMyFile() {
104
- try {
105
- // parseOffice returns an OfficeParserAST object
106
- const ast = await officeParser.parseOffice("/path/to/officeFile.docx");
107
-
108
- // Use the built-in helper to get plain text (similar to old behavior)
109
- const text = ast.toText();
110
- console.log(text);
111
-
112
- // Access structured content
113
- console.log(ast.content); // Array of hierarchical nodes (paragraphs, tables, etc.)
114
- console.log(ast.metadata); // Document properties (author, title, etc.)
115
- } catch (err) {
116
- console.error(err);
117
- }
118
- }
135
+ const ast = await officeParser.parseOffice('/path/to/file.docx');
136
+
137
+ console.log(ast.type); // 'docx'
138
+ console.log(ast.metadata); // { author, title, created, ... }
139
+ console.log(ast.content); // Array of hierarchical nodes
140
+ console.log(ast.attachments);// Images/charts (if extractAttachments: true)
141
+ console.log(ast.warnings); // Non-fatal issues from parsing phase
119
142
  ```
120
143
 
121
- ### Helper Function for Text Extraction (Modern simple way)
122
- If you only need the text and want to maintain a simple one-liner, you can use this pattern:
144
+ **TypeScript (named import):**
145
+ ```ts
146
+ import { OfficeParser } from 'officeparser';
147
+
148
+ const ast = await OfficeParser.parseOffice('report.docx', {
149
+ extractAttachments: true,
150
+ ocr: true,
151
+ });
152
+ ```
153
+
154
+ ### Callback (Backward Compat)
155
+
123
156
  ```js
124
- // Simple helper to get text directly
125
- const getText = async (file, config) => (await officeParser.parseOffice(file, config)).toText();
157
+ officeParser.parseOffice('/path/to/file.docx', function(ast, err) {
158
+ if (err) { console.error(err); return; }
159
+ console.log(ast.toText());
160
+ });
161
+ ```
162
+
163
+ ### File Buffers & ArrayBuffers
126
164
 
127
- // usage
128
- const text = await getText("/path/to/officeFile.docx");
129
- console.log(text);
165
+ Pass a `Buffer`, `ArrayBuffer`, or `Uint8Array` instead of a file path:
166
+
167
+ ```js
168
+ const fs = require('fs');
169
+ const buffer = fs.readFileSync('/path/to/file.pdf');
170
+ const ast = await officeParser.parseOffice(buffer);
171
+ ```
172
+
173
+ > [!IMPORTANT]
174
+ > **Text-based formats from buffers need a `fileType` hint.**
175
+ > Formats like `md`, `html`, and `csv` have no magic bytes, so the parser cannot
176
+ > auto-detect them from a buffer. You **must** provide `fileType` in that case:
177
+ > ```js
178
+ > const ast = await officeParser.parseOffice(markdownBuffer, { fileType: 'md' });
179
+ > ```
180
+
181
+ ### `ast.to()` — Generate from AST
182
+
183
+ The preferred way to convert a parsed AST to another format. Returns a `ConversionResult`.
184
+
185
+ ```ts
186
+ // ConversionResult shape:
187
+ // { value: string | Uint8Array | OfficeChunk[], messages: OfficeIssue[] }
188
+
189
+ const { value: markdown, messages } = await ast.to('md');
190
+ const { value: html } = await ast.to('html', { includeFormatting: false });
191
+ const { value: chunks } = await ast.to('chunks', { strategy: 'fixed-size', chunkSize: 800 });
192
+ const { value: pdfBytes } = await ast.to('pdf'); // Uint8Array
193
+ ```
194
+
195
+ ### `ast.toText()` — Quick Text Extraction
196
+
197
+ > [!NOTE]
198
+ > `toText()` is **synchronous** and deprecated in favour of the async `ast.to('text')`.
199
+ > It remains available for backward compatibility.
200
+
201
+ ```js
202
+ const text = ast.toText(); // synchronous, returns plain string
130
203
  ```
131
204
 
132
- ## Using the OfficeGenerator
133
- The `OfficeGenerator` is a powerful tool to convert your AST into human-readable formats or structured data.
205
+ ---
206
+
207
+ ## OfficeGenerator
208
+
209
+ Use `OfficeGenerator.generate(ast, format, config?)` when you need to produce output from an already-parsed AST:
134
210
 
135
- ```typescript
211
+ ```ts
136
212
  import { OfficeParser, OfficeGenerator } from 'officeparser';
137
213
 
138
214
  const ast = await OfficeParser.parseOffice('report.docx');
139
215
 
140
- // 1. Convert to Markdown
141
- const md = await OfficeGenerator.generate(ast, 'md');
142
- console.log(md.value);
216
+ // Convert to Markdown
217
+ const { value: md } = await OfficeGenerator.generate(ast, 'md');
143
218
 
144
- // 2. Convert to HTML with structured style mapping (Recommended)
145
- const html = await OfficeGenerator.generate(ast, 'html', {
219
+ // Convert to HTML with style mapping
220
+ const { value: html } = await OfficeGenerator.generate(ast, 'html', {
146
221
  includeFormatting: true,
147
222
  styleMap: [
148
- {
149
- selector: { nodeType: 'paragraph', attributes: { style: 'Heading 1' } },
150
- output: { tag: 'h1', classes: ['main-title'] }
223
+ {
224
+ selector: { nodeType: 'paragraph', attributes: { style: 'Heading 1' } },
225
+ output: { tag: 'h1', classes: ['main-title'] }
151
226
  }
152
227
  ]
153
228
  });
154
- console.log(html.value);
155
229
 
156
- // 3. Convert to CSV (for spreadsheets)
157
- const csv = await OfficeGenerator.generate(ast, 'csv');
158
- console.log(csv.value);
230
+ // Convert to CSV (spreadsheets)
231
+ const { value: csv } = await OfficeGenerator.generate(ast, 'csv');
159
232
  ```
160
233
 
161
- ## The New "One-Step" API: `OfficeConverter`
162
- In **v7.0.0**, we introduced the `OfficeConverter.convert` method. This is the new high-level API designed for one-step transformations where you don't need to manually interact with the AST. It automatically handles parser and generator configuration synchronization.
234
+ **Supported destinations:** `'text'` · `'md'` · `'html'` · `'csv'` · `'rtf'` · `'pdf'` · `'chunks'`
235
+
236
+ > [!NOTE]
237
+ > **PDF generation** requires the optional `puppeteer` peer dependency:
238
+ > ```bash
239
+ > npm install puppeteer
240
+ > ```
163
241
 
164
- ```typescript
242
+ ---
243
+
244
+ ## OfficeConverter — One-Step API
245
+
246
+ `OfficeConverter.convert()` combines parsing and generation in a single call. It automatically syncs parser options from generator config (e.g., enables `extractAttachments` when images are requested).
247
+
248
+ ```ts
165
249
  import { OfficeConverter } from 'officeparser';
166
250
 
167
- // One-step conversion from DOCX to Markdown
168
- const result = await OfficeConverter.convert('report.docx', 'md');
169
- console.log(result.value); // The generated Markdown string
170
- console.log(result.messages); // Array of warnings/info (e.g., "Skipped unsupported drawing")
251
+ // Minimal usage
252
+ const { value: markdown } = await OfficeConverter.convert('report.docx', 'md');
171
253
 
172
- // Complex conversion with nested configuration
173
- const htmlResult = await OfficeConverter.convert('data.xlsx', 'html', {
174
- parseConfig: {
175
- ignoreNotes: true
254
+ // With config
255
+ const { value: html, messages } = await OfficeConverter.convert('data.xlsx', 'html', {
256
+ parseConfig: {
257
+ ignoreNotes: true,
258
+ newlineDelimiter: '\n\n',
176
259
  },
177
260
  generatorConfig: {
178
261
  includeFormatting: true,
179
262
  styleMap: [
180
- {
181
- selector: { attributes: { style: { value: 'Header', operator: '~=' } } },
182
- output: { tag: 'h2', classes: ['data-header'] }
263
+ {
264
+ selector: { attributes: { style: { value: 'Header', operator: '~=' } } },
265
+ output: { tag: 'h2', classes: ['data-header'] }
183
266
  }
184
267
  ]
185
268
  },
186
- onWarning: (msg) => console.warn("Conversion Warning:", msg)
269
+ onWarning: (issue) => console.warn(`[${issue.code}] ${issue.message}`)
187
270
  });
188
271
  ```
189
272
 
273
+ > [!IMPORTANT]
274
+ > The `OfficeConverterConfig` shape uses **nested** `parseConfig` and `generatorConfig` sub-objects.
275
+ > Do **not** put parser or generator options at the top level — only `onWarning` lives there.
276
+
277
+ ---
278
+
190
279
  ## Native RAG Chunking
191
- `officeParser` provides native support for document chunking, specifically designed for Retrieval-Augmented Generation (RAG) workflows. It offers three distinct strategies to split your documents while maintaining context and metadata.
192
-
193
- ### 1. Fixed-Size Strategy (Recursive)
194
- Splits text into chunks based on character count with a specified overlap. It uses smart boundary detection to avoid cutting in the middle of sentences or paragraphs.
195
-
196
- ### 2. Document Structure Strategy
197
- Splits the document at natural structural boundaries like pages (PDF/Word), slides (PPTX), or high-level headings. This preserves the logical flow of the document.
198
-
199
- ### 3. Semantic Strategy
200
- Uses cosine similarity between sentence embeddings to identify coherent topic boundaries. This ensures that each chunk contains semantically related content (requires an embedding function).
201
-
202
- ### The `OfficeChunk` Interface
203
- Every chunk produced contains not just text, but rich metadata to help your RAG pipeline:
204
- ```typescript
205
- {
206
- text: string; // The chunk content
207
- metadata: {
208
- sourceType: string; // e.g., "docx", "pdf"
209
- pageNumber?: number; // Current page
210
- slideNumber?: number; // Current slide
211
- closestHeading?: string; // The heading this chunk belongs to
212
- chunkIndex: number; // Sequential index
213
- }
214
- }
280
+
281
+ `officeParser` provides native document chunking for Retrieval-Augmented Generation (RAG) pipelines with three strategies:
282
+
283
+ ### Strategy 1: Document Structure (Default)
284
+ Splits at natural AST boundaries (paragraphs, headings, pages, slides, sheets). Preserves logical flow.
285
+
286
+ ```ts
287
+ const { value: chunks } = await OfficeConverter.convert('report.docx', 'chunks', {
288
+ generatorConfig: {
289
+ chunksConfig: {
290
+ strategy: 'document-structure',
291
+ splitBy: 'heading', // 'paragraph' | 'heading' | 'page' | 'slide' | 'sheet'
292
+ maxChunkSize: 1500,
293
+ tableSplitStrategy: 'row', // repeats header row in every chunk — ideal for RAG
294
+ }
295
+ }
296
+ });
215
297
  ```
216
298
 
217
- #### Example: Generating Chunks
218
- ```typescript
219
- const chunks = await OfficeGenerator.generate(ast, 'chunks', {
220
- strategy: 'fixed-size',
221
- maxChunkSize: 1000,
222
- chunkOverlap: 200
299
+ ### Strategy 2: Fixed-Size (Recursive)
300
+ Splits by character count with overlap. Equivalent to LangChain's `RecursiveCharacterTextSplitter`.
301
+
302
+ ```ts
303
+ const { value: chunks } = await OfficeConverter.convert('report.docx', 'chunks', {
304
+ generatorConfig: {
305
+ chunksConfig: {
306
+ strategy: 'fixed-size',
307
+ chunkSize: 1000,
308
+ chunkOverlap: 200,
309
+ }
310
+ }
223
311
  });
224
- console.log(`Generated ${chunks.value.length} chunks`);
312
+ console.log(`Generated ${chunks.length} chunks`);
225
313
  ```
226
314
 
227
- ### Using Callbacks (Backward Compatibility Support)
228
- Callbacks are still supported for those preferred, but the data returned is now the AST object.
229
- ```js
230
- const officeParser = require('officeparser');
315
+ ### Strategy 3: Semantic
316
+ Uses cosine similarity between sentence embeddings to find topic boundaries. Requires you to provide an `embeddingFunction`.
231
317
 
232
- officeParser.parseOffice("/path/to/officeFile.docx", function(ast, err) {
233
- if (err) {
234
- console.error(err);
235
- return;
318
+ ```ts
319
+ import OpenAI from 'openai';
320
+ const openai = new OpenAI();
321
+
322
+ const { value: chunks } = await OfficeConverter.convert('report.docx', 'chunks', {
323
+ generatorConfig: {
324
+ chunksConfig: {
325
+ strategy: 'semantic',
326
+ embeddingFunction: async (text) => {
327
+ const res = await openai.embeddings.create({
328
+ input: text, model: 'text-embedding-3-small'
329
+ });
330
+ return res.data[0].embedding;
331
+ },
332
+ similarityThreshold: 0.8,
333
+ maxChunkSize: 2000,
334
+ }
236
335
  }
237
- // Get text from AST
238
- console.log(ast.toText());
239
336
  });
240
337
  ```
241
338
 
242
- ### Using File Buffers or ArrayBuffers
243
- You can pass a file path string, a Node.js `Buffer`, or an `ArrayBuffer`.
244
- ```js
245
- const fs = require('fs');
246
- const officeParser = require('officeparser');
247
- const buffer = fs.readFileSync("/path/to/officeFile.pdf");
339
+ ### The `OfficeChunk` Object
248
340
 
249
- officeParser.parseOffice(buffer)
250
- .then(ast => console.log(ast.toText()))
251
- .catch(console.error);
341
+ Every chunk contains text and rich metadata for citations and filtered retrieval:
342
+
343
+ ```ts
344
+ interface OfficeChunk {
345
+ text: string;
346
+ /** Rich metadata for filtered retrieval */
347
+ metadata: {
348
+ sourceType: string; // e.g., 'docx', 'pdf'
349
+ pageNumber?: number; // (PDF only)
350
+ slideNumber?: number; // (PPTX only)
351
+ sheetName?: string; // (XLSX only)
352
+ closestHeading?: string; // Nearest heading above this chunk
353
+ isTableChunk?: boolean; // True if part of a split table
354
+ };
355
+ startIndex?: number; // Character offset (if addStartIndex: true)
356
+ endIndex?: number; // End character offset (if addStartIndex: true)
357
+ }
252
358
  ```
253
359
 
360
+ ---
361
+
254
362
  ## The AST Structure
255
- The `OfficeParserAST` provides a format-agnostic representation of your document, allowing you to traverse and manipulate content as a tree.
256
363
 
257
- ### Visualizing the AST
258
- The `OfficeParserAST` provides a format-agnostic representation of your document. Below is a simplified visualization of how the tree is structured:
364
+ `OfficeParserAST` is a format-agnostic document representation:
259
365
 
260
366
  ```text
261
367
  OfficeParserAST
262
- ├── type: "docx" | "pdf" | "xlsx" | "csv" | "md" | ... (11 formats supported)
263
- ├── metadata: { author, title, created, modified, ..., customProperties }
368
+ ├── type: 'docx' | 'pdf' | 'xlsx' | 'csv' | 'md' | ... (11 formats)
369
+ ├── metadata: { author, title, created, modified, customProperties, styleMap, ... }
264
370
  ├── content: [ OfficeContentNode ]
265
- │ ├── type: "paragraph" | "heading" | "table" | "list" | ...
266
- │ ├── text: "Concatenated text of this node and all children"
267
- │ ├── children: [ OfficeContentNode ] (recursive)
268
- │ ├── formatting: { bold, italic, color, size, font, ... }
269
- │ ├── metadata: { level, listId, paragraphIndentation, row, col, ... }
270
- │ └── rawContent: "<xml>...</xml>" (if enabled)
271
- ├── attachments: [ OfficeAttachment ]
272
- │ ├── type: "image" | "chart"
273
- │ ├── name: "image1.png"
274
- │ ├── data: "base64..."
275
- │ ├── ocrText: "Text extracted via OCR"
276
- │ └── chartData: { title, dataSets, labels, ... }
277
- └── toText(): Function -> returns full plain text
278
- ```
279
-
280
- #### Representative JSON Snippet
281
- ```json
282
- {
283
- "type": "docx",
284
- "metadata": { "author": "John Doe", "title": "Annual Report", "customProperties": { "Department": "Finance" } },
285
- "content": [
286
- {
287
- "type": "heading",
288
- "text": "Introduction",
289
- "metadata": { "level": 1 },
290
- "children": [
291
- { "type": "text", "text": "Introduction", "formatting": { "bold": true } }
292
- ]
293
- },
294
- {
295
- "type": "paragraph",
296
- "text": "This is a report with an image.",
297
- "children": [
298
- { "type": "text", "text": "This is a report with an " },
299
- { "type": "image", "metadata": { "attachmentName": "img1.png" } }
300
- ]
301
- }
302
- ],
303
- "attachments": [
304
- { "name": "img1.png", "type": "image", "data": "iVBOR...", "ocrText": "Extracted Text" }
305
- ]
371
+ │ ├── type: 'paragraph' | 'heading' | 'table' | 'list' | 'image' | 'chart' | ...
372
+ │ ├── text: string (concatenated text of node + all descendants)
373
+ │ ├── children: [ OfficeContentNode ] (recursive)
374
+ │ ├── formatting: { bold, italic, underline, color, size, font, alignment, ... }
375
+ │ └── metadata: { level, listId, row, col, rowSpan, colSpan, style, ... }
376
+ ├── attachments: [ OfficeAttachment ] (populated when extractAttachments: true)
377
+ │ ├── type: 'image' | 'chart'
378
+ │ ├── name: string
379
+ │ ├── mimeType: string
380
+ │ ├── data: string (Base64)
381
+ │ ├── ocrText?: string (if ocr: true)
382
+ │ └── chartData?: { title, dataSets, labels }
383
+ ├── warnings: OfficeIssue[] (non-fatal issues from the parsing phase)
384
+ ├── to(format, config?) (format: 'html'|'md'|'text'|'csv'|'rtf'|'pdf'|'chunks', returns { value, messages })
385
+ └── toText() (Deprecated: use .to('text') instead)
386
+ ```
387
+
388
+ ### `OfficeIssue` — Warning / Error Object
389
+
390
+ All warnings and errors (from both parsing and generation) use this shape:
391
+
392
+ ```ts
393
+ interface OfficeIssue {
394
+ type: 'warning' | 'info' | 'error';
395
+ code: OfficeWarningType | OfficeErrorType; // typed enum, e.g. 'OCR_FAILED'
396
+ message: string;
397
+ node?: OfficeContentNode; // the node that triggered the issue, if any
398
+ details?: any; // original error or extra context
306
399
  }
307
400
  ```
308
401
 
402
+ ---
403
+
309
404
  ## Deep Dive: Document Components
310
405
 
311
- ### 1. Working with Lists
312
- Lists are represented as sequential `list` nodes. To reconstruct or track a list, use the `metadata` fields:
406
+ ### 1. Lists
313
407
 
314
408
  ```text
315
409
  List Node
316
- ├── type: "list"
317
- ├── metadata: {
318
- listId: "1",
319
- listType: "ordered",
320
- indentation: 0,
321
- paragraphIndentation: { left: 720, hanging: 360 },
322
- itemIndex: 0
323
- }
324
- └── children: [ Text Content... ]
410
+ ├── type: 'list'
411
+ ├── metadata: {
412
+ │ listId: '1', // items with the same listId belong to one logical list
413
+ │ listType: 'ordered' | 'unordered',
414
+ │ indentation: 0, // nesting level (0-based)
415
+ │ itemIndex: 0, // sequential position within the list level
416
+ │ paragraphIndentation: { left, hanging, right, firstLine }
417
+ │ }
418
+ └── children: [ Text content ]
325
419
  ```
326
420
 
327
- - **`listId`**: A unique identifier for the list definition. Multiple items with the same `listId` belong to the same logical list.
328
- - **`indentation`**: The structural nesting level (0-based).
329
- - **`paragraphIndentation`**: The physical indentation formatting in twentieths of a point (twips) (e.g., `left`, `right`, `firstLine`, `hanging`).
330
- - **`itemIndex`**: The sequential position within that list level.
331
- - **`listType`**: Either `ordered` (numbered) or `unordered` (bulleted).
332
-
333
421
  > [!TIP]
334
- > Even if a list is interrupted by a regular paragraph, the `itemIndex` will continue to increment for the same `listId`, allowing you to maintain correct numbering.
422
+ > Even if a list is interrupted by a regular paragraph, `itemIndex` keeps incrementing for the same `listId`, so numbering stays correct.
335
423
 
336
- ### 2. Navigating Tables
337
- Tables follow a strict hierarchy: `table` -> `row` -> `cell`.
424
+ ### 2. Tables
425
+
426
+ Tables follow a strict `table → row → cell` hierarchy:
338
427
 
339
428
  ```text
340
- Table Node
341
- ├── type: "table"
342
- └── children: [ Row Node ]
343
- ├── type: "row"
344
- └── children: [ Cell Node ]
345
- ├── type: "cell"
346
- ├── metadata: { row, col, rowSpan, colSpan }
347
- └── children: [ Paragraph/List/etc. ]
429
+ Table Node (type: 'table')
430
+ └── children: Row Nodes (type: 'row')
431
+ └── children: Cell Nodes (type: 'cell')
432
+ ├── metadata: { row, col, rowSpan?, colSpan? }
433
+ └── children: [ Paragraph | List | Table | ... ]
348
434
  ```
349
435
 
350
- - **`row` / `col`**: Zero-based indices for grid positioning.
351
- - **`rowSpan` / `colSpan`** (Optional): Integer values indicating merged cells (primarily in ODF formats). If absent, the cell is not merged.
352
- - **Recursive Content**: Cells contain their own `children` array, which can include paragraphs, lists, or even other nested tables.
436
+ - `row` / `col`: zero-based grid position
437
+ - `rowSpan` / `colSpan`: merged cells (primarily ODF formats)
438
+ - Cells can contain nested tables
353
439
 
354
- ### 3. Charts & Data
355
- When a chart is discovered, it's added as a `chart` node in the content and a corresponding `OfficeAttachment`.
440
+ ### 3. Images & OCR
356
441
 
357
442
  ```text
358
- Chart Node
359
- ├── type: "chart"
360
- ├── metadata: { attachmentName: "chart1.xml" }
361
- └── Attachment (Linked)
362
- └── chartData: { title, dataSets: [...], labels: [...] }
443
+ Image Node (type: 'image')
444
+ ├── metadata: { attachmentName: 'img1.png', altText: '...' }
445
+ └── → Attachment: { data: 'base64...', ocrText: '...' }
363
446
  ```
364
447
 
365
- - **`attachmentName`**: Links the content node to the `attachments` array.
366
- - **`chartData`**: A structured object containing titles, axis labels, and category/series data.
448
+ - Set `extractAttachments: true` to populate `attachment.data`
449
+ - Set `ocr: true` (requires `extractAttachments: true`) to populate `ocrText`
367
450
 
368
- ### 4. Images, OCR & Alt Text
369
- Images are linked via `attachmentName` and can contain valuable metadata:
451
+ ### 4. Charts
370
452
 
371
453
  ```text
372
- Image Node
373
- ├── type: "image"
374
- ├── metadata: { attachmentName: "img1.png", altText: "..." }
375
- └── Attachment (Linked)
376
- ├── data: "base64..."
377
- └── ocrText: "Extracted via OCR"
454
+ Chart Node (type: 'chart')
455
+ ├── metadata: { attachmentName: 'chart1.xml' }
456
+ └── → Attachment: { chartData: { title, dataSets, labels } }
378
457
  ```
379
458
 
380
- - **OCR Text**: If `ocr: true` is set in config, `ocrText` will contain the text found within the image.
381
- - **Alt Text**: Extracted from the document's internal image descriptions.
382
- - **Formatting**: `OfficeContentNode` images may also have parent alignment metadata.
383
-
384
459
  ### 5. Text Formatting
385
- Each `OfficeContentNode` can have a `formatting` object that defines how the text should be styled.
386
460
 
387
- ```text
388
- Text Node
389
- └── formatting: {
390
- bold: boolean,
391
- italic: boolean,
392
- underline: boolean,
393
- strikethrough: boolean,
394
- color: "#hex",
395
- backgroundColor: "#hex",
396
- size: "12pt",
397
- font: "Arial",
398
- subscript: boolean,
399
- superscript: boolean,
400
- alignment: "left" | "center" | "right" | "justify"
461
+ ```ts
462
+ formatting: {
463
+ bold?: boolean
464
+ italic?: boolean
465
+ underline?: boolean
466
+ strikethrough?: boolean
467
+ color?: string // '#RRGGBB'
468
+ backgroundColor?: string
469
+ size?: string // e.g. '12pt'
470
+ font?: string
471
+ subscript?: boolean
472
+ superscript?: boolean
473
+ alignment?: 'left' | 'center' | 'right' | 'justify'
401
474
  }
402
475
  ```
403
476
 
404
- Formatting can be found at two levels:
405
- 1. **Node Level**: Applied directly to a text run or paragraph.
406
- 2. **Document Level**: Found in `ast.metadata.formatting` (defaults) or `ast.metadata.styleMap` (named styles).
477
+ ### 6. Break Nodes (DOCX only)
407
478
 
408
- ### 6. Breaks
409
- Breaks are currently only supported when parsing DOCX-documents. Breaks are added as a node of type `break` and carry metadata of the type `BreakMetadata`. When `includeRawContent` is enabled, they also include the `rawContent` string from the original XML.
479
+ When `includeBreakNodes: true`, break elements appear as nodes:
410
480
 
411
481
  ```text
412
- Break Node
413
- ├── type: "break"
482
+ Break Node (type: 'break')
414
483
  └── metadata: {
415
- breakType: "textWrapping" | "page" | "column" | "lastRenderedPage" | "carriageReturn",
416
- clear?: "all" | "left" | "none" | "right"
484
+ breakType: 'textWrapping' | 'page' | 'column' | 'lastRenderedPage' | 'carriageReturn',
485
+ clear?: 'all' | 'left' | 'none' | 'right'
417
486
  }
418
487
  ```
419
488
 
420
- - `breakType`: Type of break. `textWrapping` (default) is a standard line break, `page` is a page break, `column` is a break to the next column, `lastRenderedPage` is a soft break inserted by Word, and `carriageReturn` is an explicit carriage return (`w:cr`).
421
- - `clear`: Relevant for `textWrapping`. Indicates if text should wrap around floating objects.
422
-
423
489
  > [!NOTE]
424
- > Even though break nodes don't have a `text` property, the `ast.toText()` method will automatically convert them to newlines (`\n`) or the configured delimiter in the final string output.
425
-
426
- ### 7. Advanced Metadata
427
- The `ast.metadata` object provides document-wide context:
428
- - **`styleMap`**: A dictionary of style names to their `TextFormatting` definitions found in the document.
429
- - **`formatting`**: Document-wide default settings (e.g., default font or font size).
430
- - **`customProperties`**: A dictionary of user-defined metadata embedded in the document (OOXML `custom.xml`, ODF `meta:user-defined`, or PDF Info dictionary).
431
-
432
- ### 8. Custom Properties
433
- You can access custom user-defined metadata that might be embedded in the document:
434
-
435
- ```javascript
436
- const ast = await officeParser.parseOffice("contract.docx");
437
- console.log("Custom Metadata:", ast.metadata.customProperties);
438
- // Output: { "ProjectID": "ABC-123", "InternalReview": true }
439
- ```
440
-
441
- ## Performance & Fidelity Highlights (v7.0.0)
442
- The v7.0.0 release brings significant internal optimizations and fidelity improvements:
443
- - **OpenOffice Speedups**: Up to **23x faster** parsing for ODP presentations thanks to optimized XML caching.
444
- - **Excel Memory Efficiency**: Resolved $O(n)$ memory overhead issues for large spreadsheets (#91) by switching to iterative stream-based parsing.
445
- - **RTF Performance**: Rewritten core loop to resolve $O(n^2)$ bottlenecks during string accumulation.
446
- - **Advanced Table Fidelity**: Native support for **vertical cell merging** (`vMerge`) and **horizontal spanning** (`gridSpan`) in DOCX, ensuring complex tables look exactly as they do in Word.
447
- - **Parser Extensions**: You can now parse `CSV`, `Markdown`, and `HTML` files *into* the unified Office AST, allowing you to use the `OfficeGenerator` on them just like any other format.
448
-
449
- ### Advanced AST Usage
450
- Beyond using `ast.toText()`, you can interact with the structural data directly:
451
-
452
- #### 1. Extract all images and their OCR text
453
- ```javascript
454
- const ast = await officeParser.parseOffice("report.docx", { ocr: true });
455
- const images = ast.attachments.filter(a => a.mimeType.startsWith('image/'));
456
- images.forEach(img => {
457
- console.log(`Image: ${img.name} (OCR: ${img.ocrText || 'N/A'})`);
458
- });
490
+ > Break nodes have no `text` property, but `ast.toText()` and `ast.to('text')` automatically convert them to the configured newline delimiter.
491
+
492
+ ### 7. Document Metadata
493
+
494
+ ```ts
495
+ ast.metadata = {
496
+ author?: string
497
+ title?: string
498
+ created?: Date
499
+ modified?: Date
500
+ description?: string
501
+ customProperties?: Record<string, any> // user-defined metadata from the document
502
+ styleMap?: Record<string, TextFormatting> // named styles → formatting definitions
503
+ formatting?: TextFormatting // document-wide defaults
504
+ }
459
505
  ```
460
506
 
461
- #### 2. Find specific headings
462
- ```javascript
463
- const headings = ast.content.filter(node => node.type === 'heading' && node.metadata?.level === 1);
464
- console.log("Main Chapters:", headings.map(h => h.text));
507
+ **Accessing custom properties:**
508
+ ```js
509
+ const ast = await officeParser.parseOffice('contract.docx');
510
+ console.log(ast.metadata.customProperties);
511
+ // { "ProjectID": "ABC-123", "InternalReview": true }
465
512
  ```
466
513
 
467
- #### 3. Custom output (e.g., Simple Markdown conversion)
468
- ```javascript
469
- const toMarkdown = (nodes) => {
470
- return nodes.map(node => {
471
- if (node.type === 'heading') return `${'#'.repeat(node.metadata?.level || 1)} ${node.text}`;
472
- if (node.type === 'list') return `- ${node.text}`;
473
- if (node.type === 'table') return "[Table Data]"; // expand children for actual table
474
- return node.text;
475
- }).join('\n\n');
476
- };
477
- console.log(toMarkdown(ast.content));
514
+ ---
515
+
516
+ ## Performance Highlights
517
+
518
+ Key internal optimizations shipped in recent versions:
519
+
520
+ - **OpenOffice (ODP)**: Up to **23× faster** parsing via optimized XML pre-parsing and style caching
521
+ - **Excel Memory**: Resolved O(n) memory overhead on large sparse spreadsheets using iterative stream-based parsing
522
+ - **RTF Parser**: Rewrote string accumulation loop to eliminate O(n²) bottleneck in large files
523
+ - **Table Fidelity (DOCX)**: Native support for vertical cell merging (`vMerge`) and horizontal spanning (`gridSpan`)
524
+
525
+ ---
526
+
527
+ ## Advanced AST Usage
528
+
529
+ ### Extract all headings
530
+ ```js
531
+ const headings = ast.content.filter(n => n.type === 'heading' && n.metadata?.level === 1);
532
+ console.log(headings.map(h => h.text));
478
533
  ```
479
534
 
480
- #### 4. Extracting Tables to CSV
481
- Iterate through table nodes and their children (rows -> cells) to build a CSV string.
482
- ```javascript
483
- const tables = ast.content.filter(node => node.type === 'table');
484
- tables.forEach((table, index) => {
535
+ ### Extract images with OCR text
536
+ ```js
537
+ const ast = await officeParser.parseOffice('report.docx', { extractAttachments: true, ocr: true });
538
+ ast.attachments.filter(a => a.mimeType?.startsWith('image/')).forEach(img => {
539
+ console.log(`${img.name}: ${img.ocrText ?? 'no OCR'}`);
540
+ });
541
+ ```
542
+
543
+ ### Extract tables to CSV manually
544
+ ```js
545
+ ast.content.filter(n => n.type === 'table').forEach((table, i) => {
485
546
  const csv = table.children
486
- .filter(row => row.type === 'row')
487
- .map(row =>
488
- row.children
489
- .filter(cell => cell.type === 'cell')
490
- .map(cell => `"${cell.text.replace(/"/g, '""')}"`) // Escape quotes
491
- .join(',')
492
- )
547
+ .filter(r => r.type === 'row')
548
+ .map(r => r.children.filter(c => c.type === 'cell')
549
+ .map(c => `"${c.text.replace(/"/g, '""')}"`)
550
+ .join(','))
493
551
  .join('\n');
494
- console.log(`Table ${index + 1} CSV:\n${csv}`);
552
+ console.log(`Table ${i + 1}:\n${csv}`);
495
553
  });
496
554
  ```
497
555
 
498
- #### 5. Filtering by Formatting (e.g., Bold Text)
499
- Find all text nodes that have specific formatting applied.
500
- ```javascript
501
- function findBoldText(nodes) {
502
- let results = [];
503
- nodes.forEach(node => {
504
- if (node.type === 'text' && node.formatting?.bold) {
505
- results.push(node.text);
506
- }
507
- if (node.children) {
508
- results = results.concat(findBoldText(node.children));
509
- }
510
- });
511
- return results;
556
+ ### Find all bold text runs
557
+ ```js
558
+ function findBold(nodes) {
559
+ return nodes.flatMap(n => [
560
+ ...(n.type === 'text' && n.formatting?.bold ? [n.text] : []),
561
+ ...(n.children ? findBold(n.children) : [])
562
+ ]);
512
563
  }
513
-
514
- const boldStrings = findBoldText(ast.content);
515
- console.log("Bold Text Found:", boldStrings);
564
+ console.log(findBold(ast.content));
516
565
  ```
517
566
 
518
- #### 6. Processing Footnotes/Endnotes
519
- If you kept notes inline (default behavior), you can extract them into a separate list for processing.
520
- ```javascript
567
+ ### Extract footnotes / endnotes
568
+ ```js
521
569
  function extractNotes(nodes) {
522
- let notes = [];
523
- nodes.forEach(node => {
524
- if (node.type === 'note') {
525
- notes.push({ id: node.metadata.noteId, text: node.text, type: node.metadata.noteType });
526
- }
527
- if (node.children) {
528
- notes = notes.concat(extractNotes(node.children));
529
- }
530
- });
531
- return notes;
570
+ return nodes.flatMap(n => [
571
+ ...(n.type === 'note' ? [{ id: n.metadata.noteId, text: n.text, type: n.metadata.noteType }] : []),
572
+ ...(n.children ? extractNotes(n.children) : [])
573
+ ]);
532
574
  }
575
+ console.log(extractNotes(ast.content));
576
+ ```
577
+
578
+ ### Search for a term (TypeScript)
579
+ ```ts
580
+ import { OfficeParser } from 'officeparser';
581
+
582
+ async function contains(filePath: string, term: string): Promise<boolean> {
583
+ const ast = await OfficeParser.parseOffice(filePath);
584
+ return (await ast.to('text')).value.includes(term);
585
+ }
586
+ ```
587
+
588
+ ---
589
+
590
+ ## Configuration Reference
591
+
592
+ ### OfficeParserConfig
593
+
594
+ Pass as the second argument to `parseOffice(file, config)`.
595
+
596
+ | Option | Type | Default | Description |
597
+ |--------|------|---------|-------------|
598
+ | `newlineDelimiter` | `string` | `'\n'` | Delimiter inserted between lines in text output |
599
+ | `ignoreNotes` | `boolean` | `false` | Ignore speaker notes (PPTX/ODP) |
600
+ | `putNotesAtLast` | `boolean` | `false` | Collect all notes at the end instead of inline |
601
+ | `extractAttachments` | `boolean` | `false` | Populate `ast.attachments` with Base64 images/charts |
602
+ | `ocr` | `boolean` | `false` | Run Tesseract OCR on images (requires `extractAttachments: true`) |
603
+ | `ocrConfig` | `OcrConfig` | `{}` | OCR worker pool settings — see [OCR section](#ocr-scheduler--resource-management) |
604
+ | `includeRawContent` | `boolean` | `false` | Attach raw XML/RTF source to each node |
605
+ | `serializeRawContent` | `boolean` | `true` | Re-serialize XML to clean strings (only if `includeRawContent: true`) |
606
+ | `preserveXmlWhitespace` | `boolean` | `false` | Preserve original XML whitespace during serialization |
607
+ | `includeBreakNodes` | `boolean` | `false` | Include `w:br` / `w:cr` as typed break nodes (DOCX only) |
608
+ | `ignoreInternalLinks` | `boolean` | `false` | Strip bookmarks and internal cross-references from AST |
609
+ | `fileType` | `SupportedFileType \| null` | `null` | **Required for text-based binary data** (`'md'`, `'html'`, `'csv'`) as these lack magic bytes. |
610
+ | `csvDelimiter` | `string` | `','` | Input delimiter when parsing CSV files |
611
+ | `pdfWorkerSrc` | `string` | CDN (jsDelivr) | Path/URL to `pdf.worker.min.mjs` (required in browser) |
612
+ | `onWarning` | `(issue: OfficeIssue) => void` | — | Callback for non-fatal parsing issues |
613
+ | `outputErrorToConsole` | `boolean` | `false` | **Deprecated.** Use `onWarning` instead |
614
+
615
+ ---
616
+
617
+ ### GeneratorConfig (Common)
618
+
619
+ Options shared by all generator formats. Pass to `OfficeGenerator.generate(ast, format, config)` or `ast.to(format, config)`.
620
+
621
+ | Option | Type | Default | Description |
622
+ |--------|------|---------|-------------|
623
+ | `includeFormatting` | `boolean` | `true` | Include bold/italic/colors/sizes in output |
624
+ | `generateIds` | `boolean` | `true` | Add slug-based `id` attributes to headings |
625
+ | `renderMetadata` | `boolean` | `false` | Render title/author as visible header block |
626
+ | `includeImages` | `boolean` | `true` | Include image nodes in output |
627
+ | `includeCharts` | `boolean` | `true` | Include interactive charts (HTML only) |
628
+ | `ignoreInternalLinks` | `boolean` | `false` | Strip bookmarks and internal anchors from output |
629
+ | `ignoreDefaultStyleMap` | `boolean` | `false` | Disable built-in style mappings (e.g., "Heading 1" → h1) |
630
+ | `styleMap` | `string[] \| StructuredStyleMapping[]` | `[]` | Custom semantic style mappings |
631
+ | `onNode` | `(node) => string \| false \| void` | — | Per-node callback for filtering, overriding, or mutating |
632
+ | `onWarning` | `(issue: OfficeIssue) => void` | — | Callback for non-fatal generation issues |
633
+
634
+ ---
635
+
636
+ ### `onNode` Callback — Advanced Node Manipulation
637
+
638
+ Called for **every node** in the AST during generation. Can be `async`.
533
639
 
534
- const allNotes = extractNotes(ast.content);
535
- console.log("Document Notes:", allNotes);
536
- ```
537
-
538
- ## Configuration Object: OfficeParserConfig
539
- Pass an optional config object as the second argument to `parseOffice`.
540
-
541
- | Flag | DataType | Default | Explanation |
542
- |------|----------|---------|-------------|
543
- | `outputErrorToConsole` | boolean | `false` | **Deprecated**: Use `onWarning` instead. Show logs to console in case of an error. |
544
- | `newlineDelimiter` | string | `\n` | Delimiter for new lines in text output. |
545
- | `ignoreNotes` | boolean | `false` | Ignore notes in files like PowerPoint/ODP. |
546
- | `putNotesAtLast` | boolean | `false` | Put notes text at the end of the document. |
547
- | `extractAttachments` | boolean | `false` | Extract images and charts as Base64. |
548
- | `includeRawContent` | boolean | `false` | Include raw XML/RTF markup in the nodes. |
549
- | `serializeRawContent` | boolean | `true` | Re-serializes raw XML to clean strings. |
550
- | `preserveXmlWhitespace` | boolean | `false` | Preserves original XML whitespace. |
551
- | `ocr` | boolean | `false` | Enable OCR for images (requires `extractAttachments: true`). |
552
- | `pdfWorkerSrc` | string | `(see below)` | Path to PDF.js worker. |
553
- | `ocrConfig` | object | `{}` | OCR Scheduler configuration. |
554
- | `includeBreakNodes` | boolean | `false` | Include `w:br`, `w:cr` nodes (DOCX only).|
555
- | `ignoreInternalLinks` | boolean | `false` | Remove all bookmarks and internal jumps. |
556
- | `csvDelimiter` | string | `,` | Custom delimiter for parsing CSV files. |
557
- | `fileType` | string | `null` | Manual format override (authoritative). |
558
-
559
- ## Generator Configuration: GeneratorConfig
560
- Configuration options for `OfficeGenerator.generate`.
561
-
562
- | Flag | DataType | Default | Explanation |
563
- |------|----------|---------|-------------|
564
- | `includeFormatting` | boolean | `true` | Whether to include semantic styles (bold, italic, colors, sizes) in output. |
565
- | `generateIds` | boolean | `true` | Automatically generates unique slug-based IDs for heading nodes. |
566
- | `renderMetadata` | boolean | `false` | Renders document metadata (Title, Author) as a visible header block. |
567
- | `includeImages` | boolean | `true` | Whether to include image nodes in the generated output. |
568
- | `includeCharts` | boolean | `true` | Whether to include interactive charts (HTML only). |
569
- | `ignoreInternalLinks`| boolean | `false` | Suppresses all internal bookmarks and anchor references. |
570
- | `ignoreDefaultStyleMap`| boolean | `false` | Ignore the library's default style mappings. |
571
- | `styleMap` | string[] \| array | `[]` | Array of style mappings (DSL strings or structured objects). |
572
- | `onNode` | function | `undefined` | Callback to filter, override, or mutate any node during generation. |
573
- | `onWarning` | function | `undefined` | Callback for generation-phase warnings. |
574
- | `htmlConfig` | object | `{}` | Format-specific settings for HTML generation. |
575
- | `mdConfig` | object | `{}` | Format-specific settings for Markdown generation. |
576
- | `pdfConfig` | object | `{}` | Format-specific settings for PDF generation. |
577
- | `csvConfig` | object | `{}` | Format-specific settings for CSV generation. |
578
- | `textConfig` | object | `{}` | Format-specific settings for Plain Text generation. |
579
- | `rtfConfig` | object | `{}` | Format-specific settings for RTF generation. |
580
- | `chunksConfig` | object | `(doc-struct)` | Settings for RAG chunking strategies. |
581
-
582
- ### 🛠️ Advanced Node Manipulation (Pro Users)
583
- The `onNode` callback is a powerful tool that gives you complete control over the generation process. It is called for **every single node** in the AST before it is rendered.
584
-
585
- #### Callback Capabilities:
586
- 1. **Filter/Remove Nodes**: Return `false` to skip a node and all its children.
587
- 2. **Override Rendering**: Return a `string` to use that exact text as the output, bypassing default logic and recursion.
588
- 3. **Mutate Nodes**: Modify the `node` object directly (e.g., changing `node.text`) and return `void` to let the generator proceed with your changes.
589
- 4. **Async Support**: The callback can be `async`, allowing you to fetch external data or perform complex logic during generation.
590
-
591
- #### Pro Example:
592
- ```typescript
593
- const result = await ast.to('md', {
640
+ | Return value | Effect |
641
+ |---|---|
642
+ | `false` | Skip this node and all its children |
643
+ | `string` | Use this string as the output for this node, skip default logic |
644
+ | `void` | Proceed with default rendering (mutations to `node` are applied) |
645
+
646
+ ```ts
647
+ const { value: md } = await ast.to('md', {
594
648
  onNode: async (node) => {
595
- // 1. Skip all images
649
+ // Skip all images
596
650
  if (node.type === 'image') return false;
597
651
 
598
- // 2. Redact sensitive info by mutating the node
652
+ // Redact secrets (mutate then proceed)
599
653
  if (node.text?.includes('SECRET_KEY')) {
600
654
  node.text = node.text.replace(/SECRET_KEY: \w+/, 'SECRET_KEY: [REDACTED]');
601
655
  }
602
656
 
603
- // 3. Custom rendering for specific styles
657
+ // Custom rendering for a specific style
604
658
  if (node.metadata?.style === 'Callout') {
605
659
  return `> [!INFO]\n> ${node.text}`;
606
660
  }
607
-
608
- // 4. Proceed with default rendering (implicitly returns void)
609
661
  }
610
662
  });
611
663
  ```
612
664
 
613
- ### Advanced Style Mapping (Semantic Translation)
614
- The `styleMap` configuration is the primary way to define the "semantic meaning" of document styles. We recommend using **Structured Style Mappings** for full type safety and power.
665
+ ---
666
+
667
+ ### `styleMap` — Semantic Style Mapping
668
+
669
+ Maps document style names to semantic output elements. Two formats supported:
615
670
 
616
- #### 1. Structured Style Mappings (Recommended)
617
- Use structured objects to match nodes based on type and attributes, and specify detailed output properties like classes and custom attributes.
671
+ #### Structured Objects (Recommended)
618
672
 
619
- ```typescript
673
+ ```ts
620
674
  styleMap: [
621
- {
622
- selector: {
623
- nodeType: 'paragraph',
624
- attributes: { style: 'Heading 1' }
625
- },
626
- output: {
627
- tag: 'h1',
628
- classes: ['main-title'],
629
- attributes: { id: 'top' }
630
- }
675
+ {
676
+ selector: { nodeType: 'paragraph', attributes: { style: 'Heading 1' } },
677
+ output: { tag: 'h1', classes: ['main-title'], attributes: { id: 'top' } }
631
678
  },
632
679
  {
633
- // Use operators like '~=' for partial matches
680
+ // '~=' operator matches if the word 'Quote' appears anywhere in the style name
634
681
  selector: { attributes: { style: { value: 'Quote', operator: '~=' } } },
635
- output: { tag: 'blockquote' }
682
+ output: { tag: 'blockquote', fresh: true }
636
683
  }
637
684
  ]
638
685
  ```
639
686
 
640
- #### 2. Legacy String DSL
641
- The library also maintains support for a simple string-based DSL, highly compatible with `mammoth.js`.
642
-
643
- - **Literal Matching**: `"p[style-name='Heading 1'] => h1"`
644
- - **Regex-like Matching**: `"p[style~='Title'] => h2"`
645
- - **Attribute Filters**: `"p[style-name='Quote'][lang='en'] => blockquote"`
646
-
647
- ## Format-Specific Generator Configuration
648
- Each destination format has its own specialized sub-configuration object nested within the main `GeneratorConfig`.
649
-
650
- ### 1. HtmlGeneratorConfig (`htmlConfig`)
651
- | Flag | DataType | Default | Explanation |
652
- |------|----------|---------|-------------|
653
- | `standalone` | boolean | `true` | Wraps output in a full `<html>` document with CSS and metadata. |
654
- | `chartJsSrc` | string | `(CDN)` | URL for the Chart.js library used for interactive charts. |
655
-
656
- ### 2. MdGeneratorConfig (`mdConfig`)
657
- | Flag | DataType | Default | Explanation |
658
- |------|----------|---------|-------------|
659
- | `fallbackToHtml` | boolean | `true` | Uses HTML tags for features not supported by Markdown (underlines, complex tables). |
660
-
661
- ### 3. PdfGeneratorConfig (`pdfConfig`)
662
- | Flag | DataType | Default | Explanation |
663
- |------|----------|---------|-------------|
664
- | `format` | string | `'A4'` | Paper format (e.g., 'Letter', 'A4', 'Legal'). |
665
- | `landscape` | boolean | `false` | Page orientation. |
666
- | `margin` | object | `{0,0,0,0}` | Top, right, bottom, left margins. |
667
- | `displayHeaderFooter`| boolean | `false` | Whether to display print headers and footers. |
668
- | `headerTemplate` | string | `''` | HTML template for the print header. |
669
- | `footerTemplate` | string | `''` | HTML template for the print footer. |
670
-
671
- ### 4. CsvGeneratorConfig (`csvConfig`)
672
- | Flag | DataType | Default | Explanation |
673
- |------|----------|---------|-------------|
674
- | `sheets` | string | `''` | Range of sheets to export (e.g., "1", "1-3", "1,3"). |
675
- | `mergeSheets` | boolean | `true` | Merges all sheets into one CSV string. If false, returns a ZIP. |
676
- | `columnDelimiter` | string | `','` | Custom delimiter for the CSV output. |
677
-
678
- ### 5. TextGeneratorConfig (`textConfig`)
679
- | Flag | DataType | Default | Explanation |
680
- |------|----------|---------|-------------|
681
- | `newlineDelimiter` | string | `\n` | String inserted between structural blocks. |
682
- | `preserveLayout` | boolean | `false` | Attempts to maintain table structures using whitespace. |
683
-
684
- ## One-Step Conversion: OfficeConverterConfig
685
- Configuration for the `OfficeConverter.convert()` API.
686
-
687
- | Flag | DataType | Default | Explanation |
688
- |------|----------|---------|-------------|
689
- | `parseConfig` | object | `{}` | Settings for the `OfficeParser` phase. |
690
- | `generatorConfig` | object | `{}` | Settings for the `OfficeGenerator` phase. |
691
- | `onWarning` | function | `undefined` | Global callback for issues in either phase. Overrides phase-specific callbacks. |
692
-
693
- ## Chunking Configuration: ChunkingConfig
694
- Specific options when using `format: 'chunks'`.
695
-
696
- | Flag | DataType | Default | Explanation |
697
- |------|----------|---------|-------------|
698
- | `strategy` | string | `'fixed-size'`| The chunking strategy (`fixed-size`, `document-structure`, `semantic`). |
699
- | `maxChunkSize` | number | `1000` | Maximum characters per chunk. |
700
- | `chunkOverlap` | number | `200` | Overlap between consecutive chunks. |
701
- | `similarityThreshold`| number | `0.5` | Threshold for semantic splitting (0.0 to 1.0). |
702
- | `embedBatchSize` | number | `50` | Batch size for embedding requests. |
703
-
704
- ### OCR Scheduler & Resource Management
705
- If your application uses OCR, `officeParser` utilizes an intelligent **Smart Worker Pool** to maintain a background worker pool and optimize repeated parse requests.
706
-
707
- - **Dynamic Affinity**: Workers in the pool persist with their last used language affinity.
708
- - **LRU Re-allocation**: If a new language is requested and the pool is full, the manager identifies the **Least Recently Used (LRU)** idle worker and re-initializes it for the new language. This avoids the overhead of destroying and recreating workers.
709
- - **Auto-Termination**: Workers are automatically cleaned up after 10 seconds of inactivity (configurable via `ocrConfig.autoTerminateTimeout`).
710
-
711
- #### `OfficeParser.terminateOcr()`
712
- If you have used OCR (`{ ocr: true }`) in a short-lived script (like CLI tools or one-off automation), we recommend explicitly calling `terminateOcr()` after your processing is finished. This bypasses the 10-second idle timer and allows the process to return to the terminal prompt immediately.
687
+ `fresh: true` prevents the generator from merging adjacent nodes of the same tag into one block.
713
688
 
714
- > [!NOTE]
715
- > If OCR was not used, this function is a no-op and does not need to be called.
689
+ #### Legacy String DSL
690
+
691
+ Compatible with `mammoth.js` style maps:
716
692
 
717
693
  ```js
718
- const officeParser = require('officeparser');
694
+ styleMap: [
695
+ "p[style-name='Heading 1'] => h1",
696
+ "p[style~='Title'] => h2",
697
+ "p[style-name='Quote'][lang='en'] => blockquote"
698
+ ]
699
+ ```
700
+
701
+ ---
719
702
 
720
- async function runCleaner() {
721
- await officeParser.parseOffice("file.pdf", { ocr: true });
722
- // ... process results ...
703
+ ### HtmlGeneratorConfig
723
704
 
724
- // Manually kill OCR workers for an immediate exit
725
- await officeParser.terminateOcr();
726
- }
727
- ```
705
+ Pass as `htmlConfig` inside `GeneratorConfig`.
728
706
 
729
- > [!TIP]
730
- > This is handled automatically in the built-in CLI (`npx officeparser ...`). You only need to call this manually if you are using the library in your own custom script and want a snappy exit.
707
+ | Option | Type | Default | Description |
708
+ |--------|------|---------|-------------|
709
+ | `standalone` | `boolean` | `true` | Wrap output in a full `<html>` document with CSS |
710
+ | `chartJsSrc` | `string` | jsDelivr CDN | URL for the Chart.js library |
731
711
 
732
- ```js
733
- const config = {
734
- newlineDelimiter: "\n\n",
735
- extractAttachments: true,
736
- ocr: true,
737
- ocrLanguage: 'eng+fra+esp' // Supports English, French, and Spanish simultaneously
738
- };
712
+ ### MdGeneratorConfig
739
713
 
740
- const ast = await officeParser.parseOffice("report.docx", config);
741
- console.log(`Extracted ${ast.attachments.length} images`);
742
- ```
714
+ Pass as `mdConfig` inside `GeneratorConfig`.
743
715
 
744
- ## Examples
716
+ | Option | Type | Default | Description |
717
+ |--------|------|---------|-------------|
718
+ | `fallbackToHtml` | `boolean` | `true` | Use HTML tags for features Markdown cannot represent (underlines, merged table cells, etc.) |
745
719
 
746
- **Search for a term in a document (TypeScript)**
747
- ```ts
748
- import { OfficeParser } from 'officeparser';
720
+ ### PdfGeneratorConfig
749
721
 
750
- async function hasSearchTerm(filePath: string, term: string): Promise<boolean> {
751
- const ast = await OfficeParser.parseOffice(filePath);
752
- return ast.toText().includes(term);
753
- }
754
- ```
722
+ Pass as `pdfConfig` inside `GeneratorConfig`. Requires the optional `puppeteer` peer dependency.
723
+
724
+ | Option | Type | Default | Description |
725
+ |--------|------|---------|-------------|
726
+ | `format` | `string` | `'A4'` | Paper format (`'A4'`, `'Letter'`, `'Legal'`, etc.) |
727
+ | `landscape` | `boolean` | `false` | Landscape page orientation |
728
+ | `printBackground` | `boolean` | `true` | Print background graphics |
729
+ | `margin` | `object` | `{0,0,0,0}` | Page margins (`top`, `right`, `bottom`, `left`) |
730
+ | `displayHeaderFooter` | `boolean` | `false` | Show print header/footer |
731
+ | `headerTemplate` | `string` | `''` | HTML template for the print header |
732
+ | `footerTemplate` | `string` | `''` | HTML template for the print footer |
733
+ | `scale` | `number` | `1` | Rendering scale factor |
734
+ | `launchOptions` | `object` | headless defaults | Puppeteer launch options (e.g., `executablePath`) |
735
+
736
+ ### CsvGeneratorConfig
737
+
738
+ Pass as `csvConfig` inside `GeneratorConfig`.
739
+
740
+ | Option | Type | Default | Description |
741
+ |--------|------|---------|-------------|
742
+ | `sheets` | `string` | `''` | Sheet range to export: `'1'`, `'1-3'`, `'1,3'` (1-based). Empty = all sheets |
743
+ | `mergeSheets` | `boolean` | `true` | Merge all sheets into one CSV. If `false`, returns a ZIP archive |
744
+ | `columnDelimiter` | `string` | `','` | Output column delimiter |
745
+
746
+ ### TextGeneratorConfig
747
+
748
+ Pass as `textConfig` inside `GeneratorConfig`.
749
+
750
+ | Option | Type | Default | Description |
751
+ |--------|------|---------|-------------|
752
+ | `newlineDelimiter` | `string` | `'\n'` | String inserted between structural blocks |
753
+ | `preserveLayout` | `boolean` | `false` | Render tables with aligned columns using whitespace |
754
+
755
+ ---
756
+
757
+ ### OfficeConverterConfig
758
+
759
+ Configuration for `OfficeConverter.convert(file, format, config)`.
760
+
761
+ | Option | Type | Description |
762
+ |--------|------|-------------|
763
+ | `parseConfig` | `OfficeParserConfig` | Settings for the parsing phase |
764
+ | `generatorConfig` | `GeneratorConfig` | Settings for the generation phase |
765
+ | `onWarning` | `(issue: OfficeIssue) => void` | Global warning callback (overrides phase-specific ones) |
766
+
767
+ ---
768
+
769
+ ### ChunkingConfig
770
+
771
+ `ChunkingConfig` is a **discriminated union** — the available options depend on the `strategy` field.
772
+
773
+ #### Common Options (all strategies)
774
+
775
+ | Option | Type | Default | Description |
776
+ |--------|------|---------|-------------|
777
+ | `strategy` | `string` | `'document-structure'` | Chunking strategy |
778
+ | `stripWhitespace` | `boolean` | `true` | Trim leading/trailing whitespace from each chunk |
779
+ | `includeMetadata` | `boolean` | `true` | Include page/slide/heading metadata in each chunk |
780
+ | `addStartIndex` | `boolean` | `false` | Add `startIndex` character offset to chunk metadata |
781
+ | `lengthFunction` | `(text) => number` | `text.length` | Custom size measurer (e.g., token counter) |
782
+ | `sentenceBoundaryRegex` | `string \| RegExp` | `/[.!?。!?]/` | Custom regex for sentence boundary detection |
783
+ | `abbreviations` | `string[]` | common list | Abbreviations to skip when splitting on `.` |
784
+
785
+ #### `strategy: 'fixed-size'`
786
+
787
+ | Option | Type | Default | Description |
788
+ |--------|------|---------|-------------|
789
+ | `chunkSize` | `number` | `1000` | Maximum characters per chunk |
790
+ | `chunkOverlap` | `number` | `200` | Character overlap between consecutive chunks |
791
+ | `separators` | `string[]` | `['\n\n','\n',' ','']` | Ordered list of separators to try |
792
+
793
+ #### `strategy: 'document-structure'`
794
+
795
+ | Option | Type | Default | Description |
796
+ |--------|------|---------|-------------|
797
+ | `splitBy` | `string` | `'paragraph'` | `'paragraph'` · `'heading'` · `'page'` · `'slide'` · `'sheet'` |
798
+ | `maxChunkSize` | `number` | `1000` | Max characters per chunk (oversized units are split recursively) |
799
+ | `tableSplitStrategy` | `string` | `'row'` | `'row'` (repeats header in each chunk) or `'flatten'` |
800
+
801
+ #### `strategy: 'semantic'`
802
+
803
+ | Option | Type | Default | Description |
804
+ |--------|------|---------|-------------|
805
+ | `embeddingFunction` | `(text) => Promise<number[]>` | **required** | Async embedding function |
806
+ | `similarityThreshold` | `number` | `0.8` | Cosine similarity threshold; lower = fewer boundaries |
807
+ | `maxChunkSize` | `number` | `2000` | Max characters even if similarity stays high |
808
+ | `bufferSize` | `number` | `1` | Surrounding sentences used when computing similarity |
809
+ | `embeddingBatchSize` | `number` | `50` | Sentences per embedding API batch |
810
+
811
+ ---
812
+
813
+ ## OCR Scheduler & Resource Management
814
+
815
+ When `ocr: true` is set, `officeParser` maintains an intelligent **Smart Worker Pool** backed by Tesseract.js:
816
+
817
+ - **Dynamic Affinity**: Workers persist with their last-used language, avoiding re-initialization overhead.
818
+ - **LRU Re-allocation**: When a new language is requested and the pool is full, the Least Recently Used idle worker is re-initialized.
819
+ - **Auto-Termination**: Workers shut down after 10 seconds of inactivity (configurable via `ocrConfig.autoTerminateTimeout`).
820
+
821
+ ### OCR Config (`ocrConfig`)
822
+
823
+ | Option | Type | Default | Description |
824
+ |--------|------|---------|-------------|
825
+ | `language` | `string` | `'eng'` | Tesseract language code(s), e.g. `'eng+fra'` |
826
+ | `workerPath` | `string` | `''` | Custom path to Tesseract worker script |
827
+ | `corePath` | `string` | `''` | Custom path to Tesseract core script |
828
+ | `langPath` | `string` | `''` | Custom path for language data files |
829
+ | `autoTerminateTimeout` | `number` | `10000` | Inactivity timeout in ms before auto-teardown (0 = disabled) |
830
+
831
+ See all language codes at [tesseract-ocr.github.io](https://tesseract-ocr.github.io/tessdoc/Data-Files).
832
+
833
+ ### `OfficeParser.terminateOcr()`
834
+
835
+ In **short-lived scripts** (CLI tools, one-off automation), call `terminateOcr()` after processing to bypass the idle timer and exit immediately:
755
836
 
756
- **Extracting Images and their OCR text**
757
837
  ```js
758
838
  const officeParser = require('officeparser');
759
839
 
760
- const config = { extractAttachments: true, ocr: true };
761
- officeParser.parseOffice("presentation.pptx", config).then(ast => {
762
- ast.attachments.forEach(attachment => {
763
- if (attachment.type === 'image') {
764
- console.log(`Image: ${attachment.name}`);
765
- console.log(`OCR Text: ${attachment.ocrText}`);
766
- fs.writeFileSync(attachment.name, Buffer.from(attachment.data, 'base64'));
767
- }
768
- });
769
- });
840
+ const ast = await officeParser.parseOffice('file.pdf', { ocr: true });
841
+ // ... process results ...
842
+ await officeParser.terminateOcr(); // immediate exit
770
843
  ```
771
844
 
845
+ > [!TIP]
846
+ > The built-in CLI (`npx officeparser ...`) handles this automatically.
847
+ > Only call it manually in your own scripts.
848
+
849
+ ---
850
+
772
851
  ## Browser Usage
773
- The library provides two types of browser bundles in the `dist/` directory:
774
- 1. **`officeparser.browser.iife.js`**: Standard IIFE bundle for direct `<script>` tag usage. Exposes the global `officeParser` namespace.
775
- 2. **`officeparser.browser.mjs`**: Modern ESM bundle for use with `import` statements or modern bundlers.
776
852
 
777
- ### Usage (ESM)
778
- If you are using a modern bundler like **Vite**, **Webpack**, or **Next.js**:
853
+ Two bundles are available in the `dist/` directory:
854
+
855
+ | Bundle | Usage |
856
+ |--------|-------|
857
+ | `officeparser.browser.mjs` | ESM — use with `import` statements or modern bundlers (Vite, Webpack, Next.js) |
858
+ | `officeparser.browser.iife.js` | IIFE — use with a `<script>` tag; exposes the global `officeParser` object |
779
859
 
780
- ```javascript
860
+ ### ESM (Vite / Webpack / Next.js)
861
+
862
+ ```js
781
863
  import { OfficeParser } from 'officeparser';
782
864
 
783
865
  const handleFile = async (event) => {
784
866
  const file = event.target.files[0];
785
867
  const buffer = await file.arrayBuffer();
786
-
787
- try {
788
- // Pass the Buffer or Uint8Array directly
789
- const ast = await OfficeParser.parseOffice(new Uint8Array(buffer));
790
- console.log(ast.toText());
791
- } catch (err) {
792
- console.error(err);
793
- }
868
+ const ast = await OfficeParser.parseOffice(new Uint8Array(buffer));
869
+ console.log(ast.toText());
794
870
  };
795
871
  ```
796
872
 
797
- > [!NOTE]
798
- > **Why `fs` fails in the browser**: Browsers do not have a built-in file system. If you try to pass a file path string in the browser, `officeParser` will throw a descriptive "Fail-Fast" error instead of crashing mysteriously:
799
- > `[officeparser] Node.js 'fs' module is not available in the browser. Please pass a Buffer or Uint8Array instead.`
800
-
801
- ### Usage (Script Tag)
802
- Include the IIFE bundle available in the release assets or your `dist/` folder. This exposes the global `officeParser` object.
873
+ ### Script Tag
803
874
 
804
875
  ```html
805
876
  <script src="dist/officeparser.browser.iife.js"></script>
@@ -807,60 +878,78 @@ Include the IIFE bundle available in the release assets or your `dist/` folder.
807
878
  async function handleFile(event) {
808
879
  const file = event.target.files[0];
809
880
  const buffer = await file.arrayBuffer();
810
-
811
- try {
812
- // Reconstruct as Uint8Array for the parser
813
- const ast = await officeParser.parseOffice(new Uint8Array(buffer));
814
- console.log(ast.toText());
815
- } catch (error) {
816
- console.error("Parsing failed:", error);
817
- }
881
+ const ast = await officeParser.parseOffice(new Uint8Array(buffer));
882
+ console.log(ast.toText());
818
883
  }
819
884
  </script>
820
885
  ```
821
886
 
822
- ### PDF Worker Configuration in Browser
823
- When using `officeparser` in a browser environment to parse PDF files, you may provide the `pdfWorkerSrc` configuration option. If not provided, it defaults to a CDN link for `pdfjs-dist@5.6.205`.
887
+ > [!NOTE]
888
+ > **File paths don't work in the browser.** Always pass a `Buffer`, `ArrayBuffer`, or `Uint8Array`.
889
+ > Passing a path string will throw a descriptive `FEATURE_NOT_SUPPORTED_IN_BROWSER` error.
890
+
891
+ ### PDF Worker Configuration
824
892
 
825
- ```javascript
826
- const file = ...; // File object or ArrayBuffer
893
+ When parsing PDFs in the browser, a Web Worker is required. If `pdfWorkerSrc` is omitted, a jsDelivr CDN link is used automatically:
827
894
 
828
- // It will use the default CDN worker if pdfWorkerSrc is omitted
829
- const ast = await officeParser.parseOffice(file);
895
+ ```js
896
+ // Uses default CDN worker:
897
+ const ast = await officeParser.parseOffice(pdfArrayBuffer);
830
898
 
831
- // Or override it with your own path or a different version:
832
- const ast2 = await officeParser.parseOffice(file, {
833
- pdfWorkerSrc: "https://cdn.jsdelivr.net/npm/pdfjs-dist@5.6.205/build/pdf.worker.min.mjs"
899
+ // Or specify your own:
900
+ const ast = await officeParser.parseOffice(pdfArrayBuffer, {
901
+ pdfWorkerSrc: 'https://cdn.jsdelivr.net/npm/pdfjs-dist@5.6.205/build/pdf.worker.min.mjs'
834
902
  });
835
903
  ```
836
904
 
837
- > **Note:** The version of `pdfjs-dist` in the worker source should match the version used by `officeparser` (currently `5.6.205`).
905
+ > [!NOTE]
906
+ > The `pdfjs-dist` worker version must match the version bundled with `officeparser` (currently **5.6.205**).
907
+
908
+ ---
838
909
 
839
910
  ## Troubleshooting & Common Issues
840
911
 
841
- - **Node.js process stays alive after finishing**: If using OCR, the worker pool stays warm for 10s by default. Use `await terminateOcr()` at the end of your script for a snappy exit.
842
- - **"Worker not found" in Browser**: Ensure `pdfWorkerSrc` is correctly pointed to the `pdf.worker.min.mjs` file matching version `5.6.205`.
843
- - **OCR accuracy is low**: Verify your `ocrConfig.language` matches the document content. Note that OCR quality depends on image resolution.
844
- - **Out of memory on large files**: For massive spreadsheets, consider using `ast.toText()` early and allowing the full AST object to be garbage-collected.
912
+ | Symptom | Fix |
913
+ |---------|-----|
914
+ | Node.js process stays alive after finishing | Call `await officeParser.terminateOcr()` at end of script when OCR was used |
915
+ | `"Worker not found"` in browser for PDF | Verify `pdfWorkerSrc` points to `pdf.worker.min.mjs` matching version `5.6.205` |
916
+ | Low OCR accuracy | Verify `ocrConfig.language` matches the document language; quality depends on image resolution |
917
+ | Out of memory on large Excel files | Call `ast.toText()` early and discard the AST object to allow garbage collection |
918
+ | `md`/`html`/`csv` buffer not detected | Add `fileType: 'md'` (or `'html'`, `'csv'`) to config — these formats have no magic bytes |
919
+ | `IMPROPER_BUFFERS` error | Usually means no file extension and no `fileType` hint was provided for a buffer input |
920
+ | PDF generation fails | Install the optional peer dependency: `npm install puppeteer` |
845
921
 
846
- For a comprehensive guide, visit our [Debugging & Troubleshooting Documentation](https://harshankur.github.io/officeParser/#spec/debugging).
922
+ For a full debugging guide, visit the [Live Documentation](https://harshankur.github.io/officeParser/#spec/debugging).
847
923
 
924
+ ---
848
925
 
849
926
  ## Known Limitations
850
- 1. **ODT/ODS Charts**: Extraction may occasionally show inaccurate data when referencing external cell ranges or complex layout-based data.
851
- 2. **PDF Images**: PDF images are extracted as BMP files in the browser for compatibility. This conversion happens automatically.
852
- 3. **RTF Footnotes**: The `putNotesAtLast` configuration is currently not supported for RTF files; footnotes and endnotes are always collected and appended to the end of the content.
853
927
 
854
- ----------
928
+ 1. **ODT/ODS Charts**: May show inaccurate data when the chart references external cell ranges or uses complex layout-based data.
929
+ 2. **PDF Images (Browser)**: Extracted as BMP files for cross-platform compatibility. Conversion is automatic.
930
+ 3. **RTF Notes**: `putNotesAtLast` has no effect for RTF files; footnotes and endnotes are always appended at the end.
931
+
932
+ ---
855
933
 
856
934
  **npm**: [https://npmjs.com/package/officeparser](https://npmjs.com/package/officeparser)
857
935
 
858
936
  **github**: [https://github.com/harshankur/officeParser](https://github.com/harshankur/officeParser)
859
937
 
938
+ ## Support the Project
939
+
940
+ If `officeParser` has helped you save time, consider supporting its continued development. Your sponsorship helps maintain the project, add new features, and keep it robust for everyone.
941
+
942
+ <a href="https://github.com/sponsors/harshankur">
943
+ <img src="https://img.shields.io/badge/Sponsor-GitHub-ea4aaa?style=for-the-badge&logo=github-sponsors" height="36">
944
+ </a>
945
+ <a href="https://www.buymeacoffee.com/harshankur">
946
+ <img src="https://cdn.buymeacoffee.com/buttons/v2/default-yellow.png" height="36" alt="Buy Me A Coffee">
947
+ </a>
948
+
860
949
  ## Contributing
861
950
 
862
- Contributions are welcome! Please see [CONTRIBUTING.md](CONTRIBUTING.md) for details on how to get started.
951
+ Contributions are welcome! Please see [CONTRIBUTING.md](CONTRIBUTING.md) for details.
863
952
 
864
953
  ## License
865
954
 
866
- This project is licensed under the MIT License - see the [LICENSE](LICENSE) file for details.
955
+ This project is licensed under the MIT License — see the [LICENSE](LICENSE) file for details.