officeparser 7.0.0 → 7.0.2
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +710 -539
- package/dist/OfficeConverter.d.ts +4 -3
- package/dist/OfficeConverter.js +4 -3
- package/dist/OfficeParser.d.ts +1 -1
- package/dist/OfficeParser.js +27 -10
- package/dist/officeparser.browser.d.ts +6 -3
- package/dist/officeparser.browser.iife.js +28 -28
- package/dist/officeparser.browser.mjs +28 -28
- package/dist/sbom.cdx.json +98 -98
- package/dist/types.d.ts +2 -0
- package/dist/types.js +2 -0
- package/dist/utils/envUtils.d.ts +8 -3
- package/dist/utils/envUtils.js +113 -84
- package/dist/utils/errorUtils.js +2 -1
- package/dist/utils/moduleLoader.js +4 -2
- package/package.json +2 -2
package/README.md
CHANGED
|
@@ -1,6 +1,10 @@
|
|
|
1
|
-
# officeParser
|
|
1
|
+
# officeParser — Universal Office Document Parser & Generator
|
|
2
2
|
|
|
3
|
-
A robust, strictly-typed Node.js and Browser library for parsing
|
|
3
|
+
A robust, strictly-typed **Node.js and Browser** library for parsing office files into a rich **Abstract Syntax Tree (AST)** and generating high-fidelity output in multiple formats.
|
|
4
|
+
|
|
5
|
+
**Parses:** [`docx`](https://en.wikipedia.org/wiki/Office_Open_XML) · [`pptx`](https://en.wikipedia.org/wiki/Office_Open_XML) · [`xlsx`](https://en.wikipedia.org/wiki/Office_Open_XML) · [`odt`](https://en.wikipedia.org/wiki/OpenDocument) · [`odp`](https://en.wikipedia.org/wiki/OpenDocument) · [`ods`](https://en.wikipedia.org/wiki/OpenDocument) · [`pdf`](https://en.wikipedia.org/wiki/PDF) · [`rtf`](https://en.wikipedia.org/wiki/Rich_Text_Format) · [`csv`](https://en.wikipedia.org/wiki/Comma-separated_values) · [`md`](https://en.wikipedia.org/wiki/Markdown) · [`html`](https://en.wikipedia.org/wiki/HTML)
|
|
6
|
+
|
|
7
|
+
**Generates:** `Markdown` · `HTML` · `CSV` · `RTF` · `PDF` · `Plain Text` · `RAG Chunks`
|
|
4
8
|
|
|
5
9
|
[](https://badge.fury.io/js/officeparser)
|
|
6
10
|
[](https://www.npmjs.com/package/officeparser)
|
|
@@ -10,21 +14,53 @@ A robust, strictly-typed Node.js and Browser library for parsing and generating
|
|
|
10
14
|
---
|
|
11
15
|
|
|
12
16
|
### 🌟 [Live Interactive AST Visualizer & Documentation](https://harshankur.github.io/officeParser/) 🌟
|
|
13
|
-
*
|
|
17
|
+
*Upload any office file in your browser — inspect the AST, tweak config, and preview generated output in real-time.*
|
|
14
18
|
|
|
15
|
-
**
|
|
16
|
-
- **
|
|
17
|
-
- **
|
|
18
|
-
- **
|
|
19
|
-
- **Format Specs**: Read detailed specifications for the AST structure and configuration options.
|
|
19
|
+
- **AST Visualizer**: Inspect the hierarchical node tree, metadata, and raw content
|
|
20
|
+
- **Config Configurator**: Tweak options (`ignoreNotes`, `ocr`, `newlineDelimiter`) and see results instantly
|
|
21
|
+
- **Debugging**: Identify exactly how nodes are interpreted
|
|
22
|
+
- **Format Specs**: Read detailed specs for the AST structure and all config options
|
|
20
23
|
|
|
21
24
|
---
|
|
22
25
|
|
|
26
|
+
### 📝 [Changelog](CHANGELOG.md)
|
|
23
27
|
|
|
24
28
|
---
|
|
25
29
|
|
|
26
|
-
|
|
27
|
-
|
|
30
|
+
## Table of Contents
|
|
31
|
+
- [Install](#install-via-npm)
|
|
32
|
+
- [Command Line Usage](#command-line-usage)
|
|
33
|
+
- [Quick Decision Guide](#quick-decision-guide)
|
|
34
|
+
- [Library Usage: Parsing](#library-usage-parsing)
|
|
35
|
+
- [Async/Await](#asyncawait)
|
|
36
|
+
- [Callback (Backward Compat)](#callback-backward-compat)
|
|
37
|
+
- [File Buffers & ArrayBuffers](#file-buffers--arraybuffers)
|
|
38
|
+
- [`ast.to()` — Generate from AST](#astto--generate-from-ast)
|
|
39
|
+
- [`ast.toText()` — Quick Text Extraction](#asttotext--quick-text-extraction)
|
|
40
|
+
- [OfficeGenerator](#officegenerator)
|
|
41
|
+
- [OfficeConverter — One-Step API](#officeconverter--one-step-api)
|
|
42
|
+
- [Native RAG Chunking](#native-rag-chunking)
|
|
43
|
+
- [The AST Structure](#the-ast-structure)
|
|
44
|
+
- [Deep Dive: Document Components](#deep-dive-document-components)
|
|
45
|
+
- [Performance Highlights](#performance-highlights)
|
|
46
|
+
- [Advanced AST Usage](#advanced-ast-usage)
|
|
47
|
+
- [Configuration Reference](#configuration-reference)
|
|
48
|
+
- [OfficeParserConfig](#officeparserconfig)
|
|
49
|
+
- [GeneratorConfig (Common)](#generatorconfig-common)
|
|
50
|
+
- [onNode Callback](#onnode-callback--advanced-node-manipulation)
|
|
51
|
+
- [styleMap — Semantic Style Mapping](#stylemap--semantic-style-mapping)
|
|
52
|
+
- [HtmlGeneratorConfig](#htmlgeneratorconfig)
|
|
53
|
+
- [MdGeneratorConfig](#mdgeneratorconfig)
|
|
54
|
+
- [PdfGeneratorConfig](#pdfgeneratorconfig)
|
|
55
|
+
- [CsvGeneratorConfig](#csvgeneratorconfig)
|
|
56
|
+
- [TextGeneratorConfig](#textgeneratorconfig)
|
|
57
|
+
- [OfficeConverterConfig](#officeconverterconfig)
|
|
58
|
+
- [ChunkingConfig](#chunkingconfig)
|
|
59
|
+
- [OCR Scheduler & Resource Management](#ocr-scheduler--resource-management)
|
|
60
|
+
- [Browser Usage](#browser-usage)
|
|
61
|
+
- [Troubleshooting & Common Issues](#troubleshooting--common-issues)
|
|
62
|
+
- [Known Limitations](#known-limitations)
|
|
63
|
+
- [Contributing](#contributing)
|
|
28
64
|
|
|
29
65
|
---
|
|
30
66
|
|
|
@@ -34,690 +70,807 @@ A robust, strictly-typed Node.js and Browser library for parsing and generating
|
|
|
34
70
|
npm i officeparser
|
|
35
71
|
```
|
|
36
72
|
|
|
37
|
-
|
|
38
|
-
|
|
73
|
+
---
|
|
74
|
+
|
|
75
|
+
## Command Line Usage
|
|
39
76
|
|
|
40
77
|
```bash
|
|
41
|
-
#
|
|
42
|
-
npx officeparser /path/to/
|
|
78
|
+
# Full AST as JSON (default)
|
|
79
|
+
npx officeparser /path/to/file.docx
|
|
43
80
|
|
|
44
|
-
#
|
|
45
|
-
npx officeparser /path/to/
|
|
81
|
+
# Plain text output
|
|
82
|
+
npx officeparser /path/to/file.docx --format=text
|
|
46
83
|
|
|
47
|
-
#
|
|
84
|
+
# Convert DOCX to Markdown and save
|
|
48
85
|
npx officeparser report.docx --format=md --output=report.md
|
|
49
86
|
|
|
50
|
-
#
|
|
87
|
+
# Convert PPTX to HTML
|
|
51
88
|
npx officeparser presentation.pptx --format=html --output=preview.html
|
|
52
89
|
|
|
53
|
-
# Convert
|
|
90
|
+
# Convert XLSX to CSV
|
|
54
91
|
npx officeparser data.xlsx --format=csv
|
|
92
|
+
|
|
93
|
+
# Generate RAG chunks
|
|
94
|
+
npx officeparser document.pdf --format=chunks
|
|
55
95
|
```
|
|
56
96
|
|
|
57
|
-
###
|
|
58
|
-
- `--format=[json|text|md|html|csv|rtf|pdf|chunks]` The output format. Default is `json`.
|
|
59
|
-
- `--output=[path]` Optional file path to write the output to.
|
|
60
|
-
- `--toText=[true|false]` Legacy flag to output only plain text. Use `--format=text` instead.
|
|
61
|
-
- `--ignoreNotes=[true|false]` Flag to ignore notes from files like PowerPoint. Default is false.
|
|
62
|
-
- `--newlineDelimiter=[delimiter]` The delimiter to use for new lines. Default is `\n`.
|
|
63
|
-
- `--putNotesAtLast=[true|false]` Flag to collect notes at the end of files like PowerPoint. Default is false.
|
|
64
|
-
- `--outputErrorToConsole=[true|false]` **(Deprecated)** Flag to output errors to the console. Use `onWarning` callback in library usage.
|
|
65
|
-
- `--extractAttachments=[true|false]` Flag to extract images/charts as Base64. Default is false.
|
|
66
|
-
- `--ocr=[true|false]` Flag to enable OCR for extracted images. Default is false.
|
|
67
|
-
- `--includeRawContent=[true|false]` Flag to include raw XML/RTF content in nodes. Default is false.
|
|
68
|
-
- `--includeBreakNodes=[true|false]` Flag to include break nodes. Currently only available for DOCX documents.
|
|
69
|
-
- `--verbose=[true|false]` Show full error stack traces.
|
|
97
|
+
### CLI Options
|
|
70
98
|
|
|
99
|
+
| Flag | Values | Default | Description |
|
|
100
|
+
|------|--------|---------|-------------|
|
|
101
|
+
| `--format` | `json\|text\|md\|html\|csv\|rtf\|pdf\|chunks` | `json` | Output format |
|
|
102
|
+
| `--output` | path | — | Write output to a file |
|
|
103
|
+
| `--toText` | `true\|false` | `false` | **Deprecated.** Use `--format=text` |
|
|
104
|
+
| `--ignoreNotes` | `true\|false` | `false` | Ignore speaker notes (PPTX/ODP) |
|
|
105
|
+
| `--putNotesAtLast` | `true\|false` | `false` | Collect notes at end of output |
|
|
106
|
+
| `--newlineDelimiter` | string | `\n` | Delimiter between lines |
|
|
107
|
+
| `--extractAttachments` | `true\|false` | `false` | Extract images/charts as Base64 |
|
|
108
|
+
| `--ocr` | `true\|false` | `false` | Enable OCR for images |
|
|
109
|
+
| `--includeRawContent` | `true\|false` | `false` | Include raw XML/RTF in nodes |
|
|
110
|
+
| `--includeBreakNodes` | `true\|false` | `false` | Include break nodes (DOCX only) |
|
|
111
|
+
| `--outputErrorToConsole` | `true\|false` | `false` | **Deprecated.** Use `onWarning` callback |
|
|
112
|
+
| `--verbose` | `true\|false` | `false` | Show full error stack traces |
|
|
71
113
|
|
|
72
|
-
|
|
73
|
-
|
|
114
|
+
---
|
|
115
|
+
|
|
116
|
+
## Quick Decision Guide
|
|
117
|
+
|
|
118
|
+
| Goal | API to use |
|
|
119
|
+
|------|-----------|
|
|
120
|
+
| Extract text / AST from a file | `OfficeParser.parseOffice(file)` |
|
|
121
|
+
| Convert directly to another format | `OfficeConverter.convert(file, 'md')` |
|
|
122
|
+
| Parse first, then generate | `parseOffice()` → `OfficeGenerator.generate(ast, 'html')` |
|
|
123
|
+
| Convert on the AST itself (shorthand) | `ast.to('md')` |
|
|
124
|
+
| RAG pipeline chunking | `OfficeConverter.convert(file, 'chunks', {...})` |
|
|
125
|
+
|
|
126
|
+
---
|
|
127
|
+
|
|
128
|
+
## Library Usage: Parsing
|
|
129
|
+
|
|
130
|
+
### Async/Await
|
|
74
131
|
|
|
75
|
-
### Getting Started (Async/Await)
|
|
76
132
|
```js
|
|
77
133
|
const officeParser = require('officeparser');
|
|
78
134
|
|
|
79
|
-
|
|
80
|
-
|
|
81
|
-
|
|
82
|
-
|
|
83
|
-
|
|
84
|
-
|
|
85
|
-
|
|
86
|
-
|
|
87
|
-
|
|
88
|
-
|
|
89
|
-
|
|
90
|
-
|
|
91
|
-
|
|
92
|
-
|
|
93
|
-
|
|
94
|
-
|
|
135
|
+
const ast = await officeParser.parseOffice('/path/to/file.docx');
|
|
136
|
+
|
|
137
|
+
console.log(ast.type); // 'docx'
|
|
138
|
+
console.log(ast.metadata); // { author, title, created, ... }
|
|
139
|
+
console.log(ast.content); // Array of hierarchical nodes
|
|
140
|
+
console.log(ast.attachments);// Images/charts (if extractAttachments: true)
|
|
141
|
+
console.log(ast.warnings); // Non-fatal issues from parsing phase
|
|
142
|
+
```
|
|
143
|
+
|
|
144
|
+
**TypeScript (named import):**
|
|
145
|
+
```ts
|
|
146
|
+
import { OfficeParser } from 'officeparser';
|
|
147
|
+
|
|
148
|
+
const ast = await OfficeParser.parseOffice('report.docx', {
|
|
149
|
+
extractAttachments: true,
|
|
150
|
+
ocr: true,
|
|
151
|
+
});
|
|
152
|
+
```
|
|
153
|
+
|
|
154
|
+
### Callback (Backward Compat)
|
|
155
|
+
|
|
156
|
+
```js
|
|
157
|
+
officeParser.parseOffice('/path/to/file.docx', function(ast, err) {
|
|
158
|
+
if (err) { console.error(err); return; }
|
|
159
|
+
console.log(ast.toText());
|
|
160
|
+
});
|
|
95
161
|
```
|
|
96
162
|
|
|
97
|
-
###
|
|
98
|
-
|
|
163
|
+
### File Buffers & ArrayBuffers
|
|
164
|
+
|
|
165
|
+
Pass a `Buffer`, `ArrayBuffer`, or `Uint8Array` instead of a file path:
|
|
166
|
+
|
|
99
167
|
```js
|
|
100
|
-
|
|
101
|
-
const
|
|
168
|
+
const fs = require('fs');
|
|
169
|
+
const buffer = fs.readFileSync('/path/to/file.pdf');
|
|
170
|
+
const ast = await officeParser.parseOffice(buffer);
|
|
171
|
+
```
|
|
172
|
+
|
|
173
|
+
> [!IMPORTANT]
|
|
174
|
+
> **Text-based formats from buffers need a `fileType` hint.**
|
|
175
|
+
> Formats like `md`, `html`, and `csv` have no magic bytes, so the parser cannot
|
|
176
|
+
> auto-detect them from a buffer. You **must** provide `fileType` in that case:
|
|
177
|
+
> ```js
|
|
178
|
+
> const ast = await officeParser.parseOffice(markdownBuffer, { fileType: 'md' });
|
|
179
|
+
> ```
|
|
180
|
+
|
|
181
|
+
### `ast.to()` — Generate from AST
|
|
182
|
+
|
|
183
|
+
The preferred way to convert a parsed AST to another format. Returns a `ConversionResult`.
|
|
184
|
+
|
|
185
|
+
```ts
|
|
186
|
+
// ConversionResult shape:
|
|
187
|
+
// { value: string | Uint8Array | OfficeChunk[], messages: OfficeIssue[] }
|
|
188
|
+
|
|
189
|
+
const { value: markdown, messages } = await ast.to('md');
|
|
190
|
+
const { value: html } = await ast.to('html', { includeFormatting: false });
|
|
191
|
+
const { value: chunks } = await ast.to('chunks', { strategy: 'fixed-size', chunkSize: 800 });
|
|
192
|
+
const { value: pdfBytes } = await ast.to('pdf'); // Uint8Array
|
|
193
|
+
```
|
|
194
|
+
|
|
195
|
+
### `ast.toText()` — Quick Text Extraction
|
|
196
|
+
|
|
197
|
+
> [!NOTE]
|
|
198
|
+
> `toText()` is **synchronous** and deprecated in favour of the async `ast.to('text')`.
|
|
199
|
+
> It remains available for backward compatibility.
|
|
102
200
|
|
|
103
|
-
|
|
104
|
-
const text =
|
|
105
|
-
console.log(text);
|
|
201
|
+
```js
|
|
202
|
+
const text = ast.toText(); // synchronous, returns plain string
|
|
106
203
|
```
|
|
107
204
|
|
|
108
|
-
|
|
109
|
-
|
|
205
|
+
---
|
|
206
|
+
|
|
207
|
+
## OfficeGenerator
|
|
208
|
+
|
|
209
|
+
Use `OfficeGenerator.generate(ast, format, config?)` when you need to produce output from an already-parsed AST:
|
|
110
210
|
|
|
111
|
-
```
|
|
211
|
+
```ts
|
|
112
212
|
import { OfficeParser, OfficeGenerator } from 'officeparser';
|
|
113
213
|
|
|
114
214
|
const ast = await OfficeParser.parseOffice('report.docx');
|
|
115
215
|
|
|
116
|
-
//
|
|
117
|
-
const md = await OfficeGenerator.generate(ast, 'md');
|
|
118
|
-
console.log(md.value);
|
|
216
|
+
// Convert to Markdown
|
|
217
|
+
const { value: md } = await OfficeGenerator.generate(ast, 'md');
|
|
119
218
|
|
|
120
|
-
//
|
|
121
|
-
const html = await OfficeGenerator.generate(ast, 'html', {
|
|
219
|
+
// Convert to HTML with style mapping
|
|
220
|
+
const { value: html } = await OfficeGenerator.generate(ast, 'html', {
|
|
122
221
|
includeFormatting: true,
|
|
123
222
|
styleMap: [
|
|
124
|
-
{
|
|
125
|
-
selector: { nodeType: 'paragraph', attributes: { style: 'Heading 1' } },
|
|
126
|
-
output: { tag: 'h1', classes: ['main-title'] }
|
|
223
|
+
{
|
|
224
|
+
selector: { nodeType: 'paragraph', attributes: { style: 'Heading 1' } },
|
|
225
|
+
output: { tag: 'h1', classes: ['main-title'] }
|
|
127
226
|
}
|
|
128
227
|
]
|
|
129
228
|
});
|
|
130
|
-
console.log(html.value);
|
|
131
229
|
|
|
132
|
-
//
|
|
133
|
-
const csv = await OfficeGenerator.generate(ast, 'csv');
|
|
134
|
-
console.log(csv.value);
|
|
230
|
+
// Convert to CSV (spreadsheets)
|
|
231
|
+
const { value: csv } = await OfficeGenerator.generate(ast, 'csv');
|
|
135
232
|
```
|
|
136
233
|
|
|
137
|
-
|
|
138
|
-
|
|
234
|
+
**Supported destinations:** `'text'` · `'md'` · `'html'` · `'csv'` · `'rtf'` · `'pdf'` · `'chunks'`
|
|
235
|
+
|
|
236
|
+
> [!NOTE]
|
|
237
|
+
> **PDF generation** requires the optional `puppeteer` peer dependency:
|
|
238
|
+
> ```bash
|
|
239
|
+
> npm install puppeteer
|
|
240
|
+
> ```
|
|
139
241
|
|
|
140
|
-
|
|
242
|
+
---
|
|
243
|
+
|
|
244
|
+
## OfficeConverter — One-Step API
|
|
245
|
+
|
|
246
|
+
`OfficeConverter.convert()` combines parsing and generation in a single call. It automatically syncs parser options from generator config (e.g., enables `extractAttachments` when images are requested).
|
|
247
|
+
|
|
248
|
+
```ts
|
|
141
249
|
import { OfficeConverter } from 'officeparser';
|
|
142
250
|
|
|
143
|
-
//
|
|
144
|
-
const
|
|
145
|
-
console.log(result.value); // The generated Markdown string
|
|
146
|
-
console.log(result.messages); // Array of warnings/info (e.g., "Skipped unsupported drawing")
|
|
251
|
+
// Minimal usage
|
|
252
|
+
const { value: markdown } = await OfficeConverter.convert('report.docx', 'md');
|
|
147
253
|
|
|
148
|
-
//
|
|
149
|
-
const
|
|
150
|
-
parseConfig: {
|
|
151
|
-
ignoreNotes: true
|
|
254
|
+
// With config
|
|
255
|
+
const { value: html, messages } = await OfficeConverter.convert('data.xlsx', 'html', {
|
|
256
|
+
parseConfig: {
|
|
257
|
+
ignoreNotes: true,
|
|
258
|
+
newlineDelimiter: '\n\n',
|
|
152
259
|
},
|
|
153
260
|
generatorConfig: {
|
|
154
261
|
includeFormatting: true,
|
|
155
262
|
styleMap: [
|
|
156
|
-
{
|
|
157
|
-
selector: { attributes: { style: { value: 'Header', operator: '~=' } } },
|
|
158
|
-
output: { tag: 'h2', classes: ['data-header'] }
|
|
263
|
+
{
|
|
264
|
+
selector: { attributes: { style: { value: 'Header', operator: '~=' } } },
|
|
265
|
+
output: { tag: 'h2', classes: ['data-header'] }
|
|
159
266
|
}
|
|
160
267
|
]
|
|
161
268
|
},
|
|
162
|
-
onWarning: (
|
|
269
|
+
onWarning: (issue) => console.warn(`[${issue.code}] ${issue.message}`)
|
|
163
270
|
});
|
|
164
271
|
```
|
|
165
272
|
|
|
273
|
+
> [!IMPORTANT]
|
|
274
|
+
> The `OfficeConverterConfig` shape uses **nested** `parseConfig` and `generatorConfig` sub-objects.
|
|
275
|
+
> Do **not** put parser or generator options at the top level — only `onWarning` lives there.
|
|
276
|
+
|
|
277
|
+
---
|
|
278
|
+
|
|
166
279
|
## Native RAG Chunking
|
|
167
|
-
|
|
168
|
-
|
|
169
|
-
|
|
170
|
-
|
|
171
|
-
|
|
172
|
-
|
|
173
|
-
|
|
174
|
-
|
|
175
|
-
|
|
176
|
-
|
|
177
|
-
|
|
178
|
-
|
|
179
|
-
|
|
180
|
-
|
|
181
|
-
|
|
182
|
-
|
|
183
|
-
|
|
184
|
-
sourceType: string; // e.g., "docx", "pdf"
|
|
185
|
-
pageNumber?: number; // Current page
|
|
186
|
-
slideNumber?: number; // Current slide
|
|
187
|
-
closestHeading?: string; // The heading this chunk belongs to
|
|
188
|
-
chunkIndex: number; // Sequential index
|
|
189
|
-
}
|
|
190
|
-
}
|
|
280
|
+
|
|
281
|
+
`officeParser` provides native document chunking for Retrieval-Augmented Generation (RAG) pipelines with three strategies:
|
|
282
|
+
|
|
283
|
+
### Strategy 1: Document Structure (Default)
|
|
284
|
+
Splits at natural AST boundaries (paragraphs, headings, pages, slides, sheets). Preserves logical flow.
|
|
285
|
+
|
|
286
|
+
```ts
|
|
287
|
+
const { value: chunks } = await OfficeConverter.convert('report.docx', 'chunks', {
|
|
288
|
+
generatorConfig: {
|
|
289
|
+
chunksConfig: {
|
|
290
|
+
strategy: 'document-structure',
|
|
291
|
+
splitBy: 'heading', // 'paragraph' | 'heading' | 'page' | 'slide' | 'sheet'
|
|
292
|
+
maxChunkSize: 1500,
|
|
293
|
+
tableSplitStrategy: 'row', // repeats header row in every chunk — ideal for RAG
|
|
294
|
+
}
|
|
295
|
+
}
|
|
296
|
+
});
|
|
191
297
|
```
|
|
192
298
|
|
|
193
|
-
|
|
194
|
-
|
|
195
|
-
|
|
196
|
-
|
|
197
|
-
|
|
198
|
-
|
|
299
|
+
### Strategy 2: Fixed-Size (Recursive)
|
|
300
|
+
Splits by character count with overlap. Equivalent to LangChain's `RecursiveCharacterTextSplitter`.
|
|
301
|
+
|
|
302
|
+
```ts
|
|
303
|
+
const { value: chunks } = await OfficeConverter.convert('report.docx', 'chunks', {
|
|
304
|
+
generatorConfig: {
|
|
305
|
+
chunksConfig: {
|
|
306
|
+
strategy: 'fixed-size',
|
|
307
|
+
chunkSize: 1000,
|
|
308
|
+
chunkOverlap: 200,
|
|
309
|
+
}
|
|
310
|
+
}
|
|
199
311
|
});
|
|
200
|
-
console.log(`Generated ${chunks.
|
|
312
|
+
console.log(`Generated ${chunks.length} chunks`);
|
|
201
313
|
```
|
|
202
314
|
|
|
203
|
-
###
|
|
204
|
-
|
|
205
|
-
|
|
206
|
-
|
|
315
|
+
### Strategy 3: Semantic
|
|
316
|
+
Uses cosine similarity between sentence embeddings to find topic boundaries. Requires you to provide an `embeddingFunction`.
|
|
317
|
+
|
|
318
|
+
```ts
|
|
319
|
+
import OpenAI from 'openai';
|
|
320
|
+
const openai = new OpenAI();
|
|
207
321
|
|
|
208
|
-
|
|
209
|
-
|
|
210
|
-
|
|
211
|
-
|
|
322
|
+
const { value: chunks } = await OfficeConverter.convert('report.docx', 'chunks', {
|
|
323
|
+
generatorConfig: {
|
|
324
|
+
chunksConfig: {
|
|
325
|
+
strategy: 'semantic',
|
|
326
|
+
embeddingFunction: async (text) => {
|
|
327
|
+
const res = await openai.embeddings.create({
|
|
328
|
+
input: text, model: 'text-embedding-3-small'
|
|
329
|
+
});
|
|
330
|
+
return res.data[0].embedding;
|
|
331
|
+
},
|
|
332
|
+
similarityThreshold: 0.8,
|
|
333
|
+
maxChunkSize: 2000,
|
|
334
|
+
}
|
|
212
335
|
}
|
|
213
|
-
// Get text from AST
|
|
214
|
-
console.log(ast.toText());
|
|
215
336
|
});
|
|
216
337
|
```
|
|
217
338
|
|
|
218
|
-
###
|
|
219
|
-
|
|
220
|
-
|
|
221
|
-
const fs = require('fs');
|
|
222
|
-
const officeParser = require('officeparser');
|
|
223
|
-
const buffer = fs.readFileSync("/path/to/officeFile.pdf");
|
|
339
|
+
### The `OfficeChunk` Object
|
|
340
|
+
|
|
341
|
+
Every chunk contains text and rich metadata for citations and filtered retrieval:
|
|
224
342
|
|
|
225
|
-
|
|
226
|
-
|
|
227
|
-
|
|
343
|
+
```ts
|
|
344
|
+
interface OfficeChunk {
|
|
345
|
+
text: string;
|
|
346
|
+
/** Rich metadata for filtered retrieval */
|
|
347
|
+
metadata: {
|
|
348
|
+
sourceType: string; // e.g., 'docx', 'pdf'
|
|
349
|
+
pageNumber?: number; // (PDF only)
|
|
350
|
+
slideNumber?: number; // (PPTX only)
|
|
351
|
+
sheetName?: string; // (XLSX only)
|
|
352
|
+
closestHeading?: string; // Nearest heading above this chunk
|
|
353
|
+
isTableChunk?: boolean; // True if part of a split table
|
|
354
|
+
};
|
|
355
|
+
startIndex?: number; // Character offset (if addStartIndex: true)
|
|
356
|
+
endIndex?: number; // End character offset (if addStartIndex: true)
|
|
357
|
+
}
|
|
228
358
|
```
|
|
229
359
|
|
|
360
|
+
---
|
|
361
|
+
|
|
230
362
|
## The AST Structure
|
|
231
|
-
The `OfficeParserAST` provides a format-agnostic representation of your document, allowing you to traverse and manipulate content as a tree.
|
|
232
363
|
|
|
233
|
-
|
|
234
|
-
The `OfficeParserAST` provides a format-agnostic representation of your document. Below is a simplified visualization of how the tree is structured:
|
|
364
|
+
`OfficeParserAST` is a format-agnostic document representation:
|
|
235
365
|
|
|
236
366
|
```text
|
|
237
367
|
OfficeParserAST
|
|
238
|
-
├── type:
|
|
239
|
-
├── metadata: { author, title, created, modified,
|
|
368
|
+
├── type: 'docx' | 'pdf' | 'xlsx' | 'csv' | 'md' | ... (11 formats)
|
|
369
|
+
├── metadata: { author, title, created, modified, customProperties, styleMap, ... }
|
|
240
370
|
├── content: [ OfficeContentNode ]
|
|
241
|
-
│ ├── type:
|
|
242
|
-
│ ├── text:
|
|
243
|
-
│ ├── children: [ OfficeContentNode ]
|
|
244
|
-
│ ├── formatting: { bold, italic, color, size, font, ... }
|
|
245
|
-
│
|
|
246
|
-
|
|
247
|
-
├──
|
|
248
|
-
│ ├──
|
|
249
|
-
│ ├──
|
|
250
|
-
│ ├── data:
|
|
251
|
-
│ ├── ocrText
|
|
252
|
-
│ └── chartData
|
|
253
|
-
|
|
254
|
-
|
|
255
|
-
|
|
256
|
-
|
|
257
|
-
|
|
258
|
-
|
|
259
|
-
|
|
260
|
-
|
|
261
|
-
|
|
262
|
-
|
|
263
|
-
|
|
264
|
-
|
|
265
|
-
|
|
266
|
-
|
|
267
|
-
|
|
268
|
-
|
|
269
|
-
},
|
|
270
|
-
{
|
|
271
|
-
"type": "paragraph",
|
|
272
|
-
"text": "This is a report with an image.",
|
|
273
|
-
"children": [
|
|
274
|
-
{ "type": "text", "text": "This is a report with an " },
|
|
275
|
-
{ "type": "image", "metadata": { "attachmentName": "img1.png" } }
|
|
276
|
-
]
|
|
277
|
-
}
|
|
278
|
-
],
|
|
279
|
-
"attachments": [
|
|
280
|
-
{ "name": "img1.png", "type": "image", "data": "iVBOR...", "ocrText": "Extracted Text" }
|
|
281
|
-
]
|
|
371
|
+
│ ├── type: 'paragraph' | 'heading' | 'table' | 'list' | 'image' | 'chart' | ...
|
|
372
|
+
│ ├── text: string (concatenated text of node + all descendants)
|
|
373
|
+
│ ├── children: [ OfficeContentNode ] (recursive)
|
|
374
|
+
│ ├── formatting: { bold, italic, underline, color, size, font, alignment, ... }
|
|
375
|
+
│ └── metadata: { level, listId, row, col, rowSpan, colSpan, style, ... }
|
|
376
|
+
├── attachments: [ OfficeAttachment ] (populated when extractAttachments: true)
|
|
377
|
+
│ ├── type: 'image' | 'chart'
|
|
378
|
+
│ ├── name: string
|
|
379
|
+
│ ├── mimeType: string
|
|
380
|
+
│ ├── data: string (Base64)
|
|
381
|
+
│ ├── ocrText?: string (if ocr: true)
|
|
382
|
+
│ └── chartData?: { title, dataSets, labels }
|
|
383
|
+
├── warnings: OfficeIssue[] (non-fatal issues from the parsing phase)
|
|
384
|
+
├── to(format, config?) (format: 'html'|'md'|'text'|'csv'|'rtf'|'pdf'|'chunks', returns { value, messages })
|
|
385
|
+
└── toText() (Deprecated: use .to('text') instead)
|
|
386
|
+
```
|
|
387
|
+
|
|
388
|
+
### `OfficeIssue` — Warning / Error Object
|
|
389
|
+
|
|
390
|
+
All warnings and errors (from both parsing and generation) use this shape:
|
|
391
|
+
|
|
392
|
+
```ts
|
|
393
|
+
interface OfficeIssue {
|
|
394
|
+
type: 'warning' | 'info' | 'error';
|
|
395
|
+
code: OfficeWarningType | OfficeErrorType; // typed enum, e.g. 'OCR_FAILED'
|
|
396
|
+
message: string;
|
|
397
|
+
node?: OfficeContentNode; // the node that triggered the issue, if any
|
|
398
|
+
details?: any; // original error or extra context
|
|
282
399
|
}
|
|
283
400
|
```
|
|
284
401
|
|
|
402
|
+
---
|
|
403
|
+
|
|
285
404
|
## Deep Dive: Document Components
|
|
286
405
|
|
|
287
|
-
### 1.
|
|
288
|
-
Lists are represented as sequential `list` nodes. To reconstruct or track a list, use the `metadata` fields:
|
|
406
|
+
### 1. Lists
|
|
289
407
|
|
|
290
408
|
```text
|
|
291
409
|
List Node
|
|
292
|
-
├── type:
|
|
293
|
-
├── metadata: {
|
|
294
|
-
|
|
295
|
-
|
|
296
|
-
|
|
297
|
-
|
|
298
|
-
|
|
299
|
-
}
|
|
300
|
-
└── children: [ Text
|
|
410
|
+
├── type: 'list'
|
|
411
|
+
├── metadata: {
|
|
412
|
+
│ listId: '1', // items with the same listId belong to one logical list
|
|
413
|
+
│ listType: 'ordered' | 'unordered',
|
|
414
|
+
│ indentation: 0, // nesting level (0-based)
|
|
415
|
+
│ itemIndex: 0, // sequential position within the list level
|
|
416
|
+
│ paragraphIndentation: { left, hanging, right, firstLine }
|
|
417
|
+
│ }
|
|
418
|
+
└── children: [ Text content ]
|
|
301
419
|
```
|
|
302
420
|
|
|
303
|
-
- **`listId`**: A unique identifier for the list definition. Multiple items with the same `listId` belong to the same logical list.
|
|
304
|
-
- **`indentation`**: The structural nesting level (0-based).
|
|
305
|
-
- **`paragraphIndentation`**: The physical indentation formatting in twentieths of a point (twips) (e.g., `left`, `right`, `firstLine`, `hanging`).
|
|
306
|
-
- **`itemIndex`**: The sequential position within that list level.
|
|
307
|
-
- **`listType`**: Either `ordered` (numbered) or `unordered` (bulleted).
|
|
308
|
-
|
|
309
421
|
> [!TIP]
|
|
310
|
-
> Even if a list is interrupted by a regular paragraph,
|
|
422
|
+
> Even if a list is interrupted by a regular paragraph, `itemIndex` keeps incrementing for the same `listId`, so numbering stays correct.
|
|
423
|
+
|
|
424
|
+
### 2. Tables
|
|
311
425
|
|
|
312
|
-
|
|
313
|
-
Tables follow a strict hierarchy: `table` -> `row` -> `cell`.
|
|
426
|
+
Tables follow a strict `table → row → cell` hierarchy:
|
|
314
427
|
|
|
315
428
|
```text
|
|
316
|
-
Table Node
|
|
317
|
-
|
|
318
|
-
└── children:
|
|
319
|
-
|
|
320
|
-
|
|
321
|
-
├── type: "cell"
|
|
322
|
-
├── metadata: { row, col, rowSpan, colSpan }
|
|
323
|
-
└── children: [ Paragraph/List/etc. ]
|
|
429
|
+
Table Node (type: 'table')
|
|
430
|
+
└── children: Row Nodes (type: 'row')
|
|
431
|
+
└── children: Cell Nodes (type: 'cell')
|
|
432
|
+
├── metadata: { row, col, rowSpan?, colSpan? }
|
|
433
|
+
└── children: [ Paragraph | List | Table | ... ]
|
|
324
434
|
```
|
|
325
435
|
|
|
326
|
-
-
|
|
327
|
-
-
|
|
328
|
-
-
|
|
436
|
+
- `row` / `col`: zero-based grid position
|
|
437
|
+
- `rowSpan` / `colSpan`: merged cells (primarily ODF formats)
|
|
438
|
+
- Cells can contain nested tables
|
|
329
439
|
|
|
330
|
-
### 3.
|
|
331
|
-
When a chart is discovered, it's added as a `chart` node in the content and a corresponding `OfficeAttachment`.
|
|
440
|
+
### 3. Images & OCR
|
|
332
441
|
|
|
333
442
|
```text
|
|
334
|
-
|
|
335
|
-
├──
|
|
336
|
-
|
|
337
|
-
└── Attachment (Linked)
|
|
338
|
-
└── chartData: { title, dataSets: [...], labels: [...] }
|
|
443
|
+
Image Node (type: 'image')
|
|
444
|
+
├── metadata: { attachmentName: 'img1.png', altText: '...' }
|
|
445
|
+
└── → Attachment: { data: 'base64...', ocrText: '...' }
|
|
339
446
|
```
|
|
340
447
|
|
|
341
|
-
-
|
|
342
|
-
-
|
|
448
|
+
- Set `extractAttachments: true` to populate `attachment.data`
|
|
449
|
+
- Set `ocr: true` (requires `extractAttachments: true`) to populate `ocrText`
|
|
343
450
|
|
|
344
|
-
### 4.
|
|
345
|
-
Images are linked via `attachmentName` and can contain valuable metadata:
|
|
451
|
+
### 4. Charts
|
|
346
452
|
|
|
347
453
|
```text
|
|
348
|
-
|
|
349
|
-
├──
|
|
350
|
-
|
|
351
|
-
└── Attachment (Linked)
|
|
352
|
-
├── data: "base64..."
|
|
353
|
-
└── ocrText: "Extracted via OCR"
|
|
454
|
+
Chart Node (type: 'chart')
|
|
455
|
+
├── metadata: { attachmentName: 'chart1.xml' }
|
|
456
|
+
└── → Attachment: { chartData: { title, dataSets, labels } }
|
|
354
457
|
```
|
|
355
458
|
|
|
356
|
-
- **OCR Text**: If `ocr: true` is set in config, `ocrText` will contain the text found within the image.
|
|
357
|
-
- **Alt Text**: Extracted from the document's internal image descriptions.
|
|
358
|
-
- **Formatting**: `OfficeContentNode` images may also have parent alignment metadata.
|
|
359
|
-
|
|
360
459
|
### 5. Text Formatting
|
|
361
|
-
Each `OfficeContentNode` can have a `formatting` object that defines how the text should be styled.
|
|
362
460
|
|
|
363
|
-
```
|
|
364
|
-
|
|
365
|
-
|
|
366
|
-
|
|
367
|
-
|
|
368
|
-
|
|
369
|
-
|
|
370
|
-
|
|
371
|
-
|
|
372
|
-
|
|
373
|
-
|
|
374
|
-
|
|
375
|
-
|
|
376
|
-
alignment: "left" | "center" | "right" | "justify"
|
|
461
|
+
```ts
|
|
462
|
+
formatting: {
|
|
463
|
+
bold?: boolean
|
|
464
|
+
italic?: boolean
|
|
465
|
+
underline?: boolean
|
|
466
|
+
strikethrough?: boolean
|
|
467
|
+
color?: string // '#RRGGBB'
|
|
468
|
+
backgroundColor?: string
|
|
469
|
+
size?: string // e.g. '12pt'
|
|
470
|
+
font?: string
|
|
471
|
+
subscript?: boolean
|
|
472
|
+
superscript?: boolean
|
|
473
|
+
alignment?: 'left' | 'center' | 'right' | 'justify'
|
|
377
474
|
}
|
|
378
475
|
```
|
|
379
476
|
|
|
380
|
-
|
|
381
|
-
1. **Node Level**: Applied directly to a text run or paragraph.
|
|
382
|
-
2. **Document Level**: Found in `ast.metadata.formatting` (defaults) or `ast.metadata.styleMap` (named styles).
|
|
477
|
+
### 6. Break Nodes (DOCX only)
|
|
383
478
|
|
|
384
|
-
|
|
385
|
-
Breaks are currently only supported when parsing DOCX-documents. Breaks are added as a node of type `break` and carry metadata of the type `BreakMetadata`. When `includeRawContent` is enabled, they also include the `rawContent` string from the original XML.
|
|
479
|
+
When `includeBreakNodes: true`, break elements appear as nodes:
|
|
386
480
|
|
|
387
481
|
```text
|
|
388
|
-
Break Node
|
|
389
|
-
├── type: "break"
|
|
482
|
+
Break Node (type: 'break')
|
|
390
483
|
└── metadata: {
|
|
391
|
-
breakType:
|
|
392
|
-
clear?:
|
|
484
|
+
breakType: 'textWrapping' | 'page' | 'column' | 'lastRenderedPage' | 'carriageReturn',
|
|
485
|
+
clear?: 'all' | 'left' | 'none' | 'right'
|
|
393
486
|
}
|
|
394
487
|
```
|
|
395
488
|
|
|
396
|
-
- `breakType`: Type of break. `textWrapping` (default) is a standard line break, `page` is a page break, `column` is a break to the next column, `lastRenderedPage` is a soft break inserted by Word, and `carriageReturn` is an explicit carriage return (`w:cr`).
|
|
397
|
-
- `clear`: Relevant for `textWrapping`. Indicates if text should wrap around floating objects.
|
|
398
|
-
|
|
399
489
|
> [!NOTE]
|
|
400
|
-
>
|
|
401
|
-
|
|
402
|
-
### 7.
|
|
403
|
-
|
|
404
|
-
|
|
405
|
-
|
|
406
|
-
|
|
407
|
-
|
|
408
|
-
|
|
409
|
-
|
|
410
|
-
|
|
411
|
-
|
|
412
|
-
|
|
413
|
-
|
|
414
|
-
|
|
415
|
-
```
|
|
416
|
-
|
|
417
|
-
## Performance & Fidelity Highlights (v7.0.0)
|
|
418
|
-
The v7.0.0 release brings significant internal optimizations and fidelity improvements:
|
|
419
|
-
- **OpenOffice Speedups**: Up to **23x faster** parsing for ODP presentations thanks to optimized XML caching.
|
|
420
|
-
- **Excel Memory Efficiency**: Resolved $O(n)$ memory overhead issues for large spreadsheets (#91) by switching to iterative stream-based parsing.
|
|
421
|
-
- **RTF Performance**: Rewritten core loop to resolve $O(n^2)$ bottlenecks during string accumulation.
|
|
422
|
-
- **Advanced Table Fidelity**: Native support for **vertical cell merging** (`vMerge`) and **horizontal spanning** (`gridSpan`) in DOCX, ensuring complex tables look exactly as they do in Word.
|
|
423
|
-
- **Parser Extensions**: You can now parse `CSV`, `Markdown`, and `HTML` files *into* the unified Office AST, allowing you to use the `OfficeGenerator` on them just like any other format.
|
|
424
|
-
|
|
425
|
-
### Advanced AST Usage
|
|
426
|
-
Beyond using `ast.toText()`, you can interact with the structural data directly:
|
|
427
|
-
|
|
428
|
-
#### 1. Extract all images and their OCR text
|
|
429
|
-
```javascript
|
|
430
|
-
const ast = await officeParser.parseOffice("report.docx", { ocr: true });
|
|
431
|
-
const images = ast.attachments.filter(a => a.mimeType.startsWith('image/'));
|
|
432
|
-
images.forEach(img => {
|
|
433
|
-
console.log(`Image: ${img.name} (OCR: ${img.ocrText || 'N/A'})`);
|
|
434
|
-
});
|
|
490
|
+
> Break nodes have no `text` property, but `ast.toText()` and `ast.to('text')` automatically convert them to the configured newline delimiter.
|
|
491
|
+
|
|
492
|
+
### 7. Document Metadata
|
|
493
|
+
|
|
494
|
+
```ts
|
|
495
|
+
ast.metadata = {
|
|
496
|
+
author?: string
|
|
497
|
+
title?: string
|
|
498
|
+
created?: Date
|
|
499
|
+
modified?: Date
|
|
500
|
+
description?: string
|
|
501
|
+
customProperties?: Record<string, any> // user-defined metadata from the document
|
|
502
|
+
styleMap?: Record<string, TextFormatting> // named styles → formatting definitions
|
|
503
|
+
formatting?: TextFormatting // document-wide defaults
|
|
504
|
+
}
|
|
435
505
|
```
|
|
436
506
|
|
|
437
|
-
|
|
438
|
-
```
|
|
439
|
-
const
|
|
440
|
-
console.log(
|
|
507
|
+
**Accessing custom properties:**
|
|
508
|
+
```js
|
|
509
|
+
const ast = await officeParser.parseOffice('contract.docx');
|
|
510
|
+
console.log(ast.metadata.customProperties);
|
|
511
|
+
// { "ProjectID": "ABC-123", "InternalReview": true }
|
|
441
512
|
```
|
|
442
513
|
|
|
443
|
-
|
|
444
|
-
|
|
445
|
-
|
|
446
|
-
|
|
447
|
-
|
|
448
|
-
|
|
449
|
-
|
|
450
|
-
|
|
451
|
-
|
|
452
|
-
|
|
453
|
-
|
|
514
|
+
---
|
|
515
|
+
|
|
516
|
+
## Performance Highlights
|
|
517
|
+
|
|
518
|
+
Key internal optimizations shipped in recent versions:
|
|
519
|
+
|
|
520
|
+
- **OpenOffice (ODP)**: Up to **23× faster** parsing via optimized XML pre-parsing and style caching
|
|
521
|
+
- **Excel Memory**: Resolved O(n) memory overhead on large sparse spreadsheets using iterative stream-based parsing
|
|
522
|
+
- **RTF Parser**: Rewrote string accumulation loop to eliminate O(n²) bottleneck in large files
|
|
523
|
+
- **Table Fidelity (DOCX)**: Native support for vertical cell merging (`vMerge`) and horizontal spanning (`gridSpan`)
|
|
524
|
+
|
|
525
|
+
---
|
|
526
|
+
|
|
527
|
+
## Advanced AST Usage
|
|
528
|
+
|
|
529
|
+
### Extract all headings
|
|
530
|
+
```js
|
|
531
|
+
const headings = ast.content.filter(n => n.type === 'heading' && n.metadata?.level === 1);
|
|
532
|
+
console.log(headings.map(h => h.text));
|
|
533
|
+
```
|
|
534
|
+
|
|
535
|
+
### Extract images with OCR text
|
|
536
|
+
```js
|
|
537
|
+
const ast = await officeParser.parseOffice('report.docx', { extractAttachments: true, ocr: true });
|
|
538
|
+
ast.attachments.filter(a => a.mimeType?.startsWith('image/')).forEach(img => {
|
|
539
|
+
console.log(`${img.name}: ${img.ocrText ?? 'no OCR'}`);
|
|
540
|
+
});
|
|
454
541
|
```
|
|
455
542
|
|
|
456
|
-
|
|
457
|
-
|
|
458
|
-
|
|
459
|
-
const tables = ast.content.filter(node => node.type === 'table');
|
|
460
|
-
tables.forEach((table, index) => {
|
|
543
|
+
### Extract tables to CSV manually
|
|
544
|
+
```js
|
|
545
|
+
ast.content.filter(n => n.type === 'table').forEach((table, i) => {
|
|
461
546
|
const csv = table.children
|
|
462
|
-
.filter(
|
|
463
|
-
.map(
|
|
464
|
-
|
|
465
|
-
|
|
466
|
-
.map(cell => `"${cell.text.replace(/"/g, '""')}"`) // Escape quotes
|
|
467
|
-
.join(',')
|
|
468
|
-
)
|
|
547
|
+
.filter(r => r.type === 'row')
|
|
548
|
+
.map(r => r.children.filter(c => c.type === 'cell')
|
|
549
|
+
.map(c => `"${c.text.replace(/"/g, '""')}"`)
|
|
550
|
+
.join(','))
|
|
469
551
|
.join('\n');
|
|
470
|
-
console.log(`Table ${
|
|
552
|
+
console.log(`Table ${i + 1}:\n${csv}`);
|
|
471
553
|
});
|
|
472
554
|
```
|
|
473
555
|
|
|
474
|
-
|
|
475
|
-
|
|
476
|
-
|
|
477
|
-
|
|
478
|
-
|
|
479
|
-
|
|
480
|
-
|
|
481
|
-
results.push(node.text);
|
|
482
|
-
}
|
|
483
|
-
if (node.children) {
|
|
484
|
-
results = results.concat(findBoldText(node.children));
|
|
485
|
-
}
|
|
486
|
-
});
|
|
487
|
-
return results;
|
|
556
|
+
### Find all bold text runs
|
|
557
|
+
```js
|
|
558
|
+
function findBold(nodes) {
|
|
559
|
+
return nodes.flatMap(n => [
|
|
560
|
+
...(n.type === 'text' && n.formatting?.bold ? [n.text] : []),
|
|
561
|
+
...(n.children ? findBold(n.children) : [])
|
|
562
|
+
]);
|
|
488
563
|
}
|
|
489
|
-
|
|
490
|
-
const boldStrings = findBoldText(ast.content);
|
|
491
|
-
console.log("Bold Text Found:", boldStrings);
|
|
564
|
+
console.log(findBold(ast.content));
|
|
492
565
|
```
|
|
493
566
|
|
|
494
|
-
|
|
495
|
-
|
|
496
|
-
```javascript
|
|
567
|
+
### Extract footnotes / endnotes
|
|
568
|
+
```js
|
|
497
569
|
function extractNotes(nodes) {
|
|
498
|
-
|
|
499
|
-
|
|
500
|
-
|
|
501
|
-
|
|
502
|
-
|
|
503
|
-
|
|
504
|
-
|
|
505
|
-
|
|
506
|
-
|
|
507
|
-
|
|
570
|
+
return nodes.flatMap(n => [
|
|
571
|
+
...(n.type === 'note' ? [{ id: n.metadata.noteId, text: n.text, type: n.metadata.noteType }] : []),
|
|
572
|
+
...(n.children ? extractNotes(n.children) : [])
|
|
573
|
+
]);
|
|
574
|
+
}
|
|
575
|
+
console.log(extractNotes(ast.content));
|
|
576
|
+
```
|
|
577
|
+
|
|
578
|
+
### Search for a term (TypeScript)
|
|
579
|
+
```ts
|
|
580
|
+
import { OfficeParser } from 'officeparser';
|
|
581
|
+
|
|
582
|
+
async function contains(filePath: string, term: string): Promise<boolean> {
|
|
583
|
+
const ast = await OfficeParser.parseOffice(filePath);
|
|
584
|
+
return (await ast.to('text')).value.includes(term);
|
|
508
585
|
}
|
|
586
|
+
```
|
|
587
|
+
|
|
588
|
+
---
|
|
589
|
+
|
|
590
|
+
## Configuration Reference
|
|
591
|
+
|
|
592
|
+
### OfficeParserConfig
|
|
593
|
+
|
|
594
|
+
Pass as the second argument to `parseOffice(file, config)`.
|
|
595
|
+
|
|
596
|
+
| Option | Type | Default | Description |
|
|
597
|
+
|--------|------|---------|-------------|
|
|
598
|
+
| `newlineDelimiter` | `string` | `'\n'` | Delimiter inserted between lines in text output |
|
|
599
|
+
| `ignoreNotes` | `boolean` | `false` | Ignore speaker notes (PPTX/ODP) |
|
|
600
|
+
| `putNotesAtLast` | `boolean` | `false` | Collect all notes at the end instead of inline |
|
|
601
|
+
| `extractAttachments` | `boolean` | `false` | Populate `ast.attachments` with Base64 images/charts |
|
|
602
|
+
| `ocr` | `boolean` | `false` | Run Tesseract OCR on images (requires `extractAttachments: true`) |
|
|
603
|
+
| `ocrConfig` | `OcrConfig` | `{}` | OCR worker pool settings — see [OCR section](#ocr-scheduler--resource-management) |
|
|
604
|
+
| `includeRawContent` | `boolean` | `false` | Attach raw XML/RTF source to each node |
|
|
605
|
+
| `serializeRawContent` | `boolean` | `true` | Re-serialize XML to clean strings (only if `includeRawContent: true`) |
|
|
606
|
+
| `preserveXmlWhitespace` | `boolean` | `false` | Preserve original XML whitespace during serialization |
|
|
607
|
+
| `includeBreakNodes` | `boolean` | `false` | Include `w:br` / `w:cr` as typed break nodes (DOCX only) |
|
|
608
|
+
| `ignoreInternalLinks` | `boolean` | `false` | Strip bookmarks and internal cross-references from AST |
|
|
609
|
+
| `fileType` | `SupportedFileType \| null` | `null` | **Required for text-based buffers** (`'md'`, `'html'`, `'csv'`) with no magic bytes |
|
|
610
|
+
| `csvDelimiter` | `string` | `','` | Input delimiter when parsing CSV files |
|
|
611
|
+
| `pdfWorkerSrc` | `string` | CDN (jsDelivr) | Path/URL to `pdf.worker.min.mjs` (required in browser) |
|
|
612
|
+
| `onWarning` | `(issue: OfficeIssue) => void` | — | Callback for non-fatal parsing issues |
|
|
613
|
+
| `outputErrorToConsole` | `boolean` | `false` | **Deprecated.** Use `onWarning` instead |
|
|
614
|
+
|
|
615
|
+
---
|
|
616
|
+
|
|
617
|
+
### GeneratorConfig (Common)
|
|
618
|
+
|
|
619
|
+
Options shared by all generator formats. Pass to `OfficeGenerator.generate(ast, format, config)` or `ast.to(format, config)`.
|
|
620
|
+
|
|
621
|
+
| Option | Type | Default | Description |
|
|
622
|
+
|--------|------|---------|-------------|
|
|
623
|
+
| `includeFormatting` | `boolean` | `true` | Include bold/italic/colors/sizes in output |
|
|
624
|
+
| `generateIds` | `boolean` | `true` | Add slug-based `id` attributes to headings |
|
|
625
|
+
| `renderMetadata` | `boolean` | `false` | Render title/author as visible header block |
|
|
626
|
+
| `includeImages` | `boolean` | `true` | Include image nodes in output |
|
|
627
|
+
| `includeCharts` | `boolean` | `true` | Include interactive charts (HTML only) |
|
|
628
|
+
| `ignoreInternalLinks` | `boolean` | `false` | Strip bookmarks and internal anchors from output |
|
|
629
|
+
| `ignoreDefaultStyleMap` | `boolean` | `false` | Disable built-in style mappings (e.g., "Heading 1" → h1) |
|
|
630
|
+
| `styleMap` | `string[] \| StructuredStyleMapping[]` | `[]` | Custom semantic style mappings |
|
|
631
|
+
| `onNode` | `(node) => string \| false \| void` | — | Per-node callback for filtering, overriding, or mutating |
|
|
632
|
+
| `onWarning` | `(issue: OfficeIssue) => void` | — | Callback for non-fatal generation issues |
|
|
633
|
+
|
|
634
|
+
---
|
|
635
|
+
|
|
636
|
+
### `onNode` Callback — Advanced Node Manipulation
|
|
637
|
+
|
|
638
|
+
Called for **every node** in the AST during generation. Can be `async`.
|
|
639
|
+
|
|
640
|
+
| Return value | Effect |
|
|
641
|
+
|---|---|
|
|
642
|
+
| `false` | Skip this node and all its children |
|
|
643
|
+
| `string` | Use this string as the output for this node, skip default logic |
|
|
644
|
+
| `void` | Proceed with default rendering (mutations to `node` are applied) |
|
|
509
645
|
|
|
510
|
-
|
|
511
|
-
|
|
512
|
-
```
|
|
513
|
-
|
|
514
|
-
## Configuration Object: OfficeParserConfig
|
|
515
|
-
Pass an optional config object as the second argument to `parseOffice`.
|
|
516
|
-
|
|
517
|
-
| Flag | DataType | Default | Explanation |
|
|
518
|
-
|------|----------|---------|-------------|
|
|
519
|
-
| `outputErrorToConsole` | boolean | `false` | **Deprecated**: Use `onWarning` instead. Show logs to console in case of an error. |
|
|
520
|
-
| `newlineDelimiter` | string | `\n` | Delimiter for new lines in text output. |
|
|
521
|
-
| `ignoreNotes` | boolean | `false` | Ignore notes in files like PowerPoint/ODP. |
|
|
522
|
-
| `putNotesAtLast` | boolean | `false` | Put notes text at the end of the document. |
|
|
523
|
-
| `extractAttachments` | boolean | `false` | Extract images and charts as Base64. |
|
|
524
|
-
| `includeRawContent` | boolean | `false` | Include raw XML/RTF markup in the nodes. |
|
|
525
|
-
| `serializeRawContent` | boolean | `true` | Re-serializes raw XML to clean strings. |
|
|
526
|
-
| `preserveXmlWhitespace` | boolean | `false` | Preserves original XML whitespace. |
|
|
527
|
-
| `ocr` | boolean | `false` | Enable OCR for images (requires `extractAttachments: true`). |
|
|
528
|
-
| `pdfWorkerSrc` | string | `(see below)` | Path to PDF.js worker. |
|
|
529
|
-
| `ocrConfig` | object | `{}` | OCR Scheduler configuration. |
|
|
530
|
-
| `includeBreakNodes` | boolean | `false` | Include `w:br`, `w:cr` nodes (DOCX only).|
|
|
531
|
-
| `ignoreInternalLinks` | boolean | `false` | Remove all bookmarks and internal jumps. |
|
|
532
|
-
| `csvDelimiter` | string | `,` | Custom delimiter for parsing CSV files. |
|
|
533
|
-
| `fileType` | string | `null` | Manual format override (authoritative). |
|
|
534
|
-
|
|
535
|
-
## Generator Configuration: GeneratorConfig
|
|
536
|
-
Configuration options for `OfficeGenerator.generate`.
|
|
537
|
-
|
|
538
|
-
| Flag | DataType | Default | Explanation |
|
|
539
|
-
|------|----------|---------|-------------|
|
|
540
|
-
| `includeFormatting` | boolean | `false` | Whether to include semantic styles (bold, italic) in output. |
|
|
541
|
-
| `styleMap` | string[] \| array | `[]` | Array of style mappings (DSL strings or structured objects). |
|
|
542
|
-
| `ignoreDefaultStyleMap`| boolean | `false` | Ignore the library's default style mappings. |
|
|
543
|
-
| `includeMetadata` | boolean | `false` | Include document metadata in the output (e.g., as frontmatter). |
|
|
544
|
-
| `onNode` | function | `undefined` | Callback to intercept/modify any node during generation. |
|
|
545
|
-
|
|
546
|
-
### 🛠️ Advanced Node Manipulation (Pro Users)
|
|
547
|
-
The `onNode` callback is a powerful tool that gives you complete control over the generation process. It is called for **every single node** in the AST before it is rendered.
|
|
548
|
-
|
|
549
|
-
#### Callback Capabilities:
|
|
550
|
-
1. **Filter/Remove Nodes**: Return `false` to skip a node and all its children.
|
|
551
|
-
2. **Override Rendering**: Return a `string` to use that exact text as the output, bypassing default logic and recursion.
|
|
552
|
-
3. **Mutate Nodes**: Modify the `node` object directly (e.g., changing `node.text`) and return `void` to let the generator proceed with your changes.
|
|
553
|
-
4. **Async Support**: The callback can be `async`, allowing you to fetch external data or perform complex logic during generation.
|
|
554
|
-
|
|
555
|
-
#### Pro Example:
|
|
556
|
-
```typescript
|
|
557
|
-
const result = await ast.to('md', {
|
|
646
|
+
```ts
|
|
647
|
+
const { value: md } = await ast.to('md', {
|
|
558
648
|
onNode: async (node) => {
|
|
559
|
-
//
|
|
649
|
+
// Skip all images
|
|
560
650
|
if (node.type === 'image') return false;
|
|
561
651
|
|
|
562
|
-
//
|
|
652
|
+
// Redact secrets (mutate then proceed)
|
|
563
653
|
if (node.text?.includes('SECRET_KEY')) {
|
|
564
654
|
node.text = node.text.replace(/SECRET_KEY: \w+/, 'SECRET_KEY: [REDACTED]');
|
|
565
655
|
}
|
|
566
656
|
|
|
567
|
-
//
|
|
657
|
+
// Custom rendering for a specific style
|
|
568
658
|
if (node.metadata?.style === 'Callout') {
|
|
569
659
|
return `> [!INFO]\n> ${node.text}`;
|
|
570
660
|
}
|
|
571
|
-
|
|
572
|
-
// 4. Proceed with default rendering (implicitly returns void)
|
|
573
661
|
}
|
|
574
662
|
});
|
|
575
663
|
```
|
|
576
664
|
|
|
577
|
-
|
|
578
|
-
|
|
665
|
+
---
|
|
666
|
+
|
|
667
|
+
### `styleMap` — Semantic Style Mapping
|
|
579
668
|
|
|
580
|
-
|
|
581
|
-
Use structured objects to match nodes based on type and attributes, and specify detailed output properties like classes and custom attributes.
|
|
669
|
+
Maps document style names to semantic output elements. Two formats supported:
|
|
582
670
|
|
|
583
|
-
|
|
671
|
+
#### Structured Objects (Recommended)
|
|
672
|
+
|
|
673
|
+
```ts
|
|
584
674
|
styleMap: [
|
|
585
|
-
{
|
|
586
|
-
selector: {
|
|
587
|
-
|
|
588
|
-
attributes: { style: 'Heading 1' }
|
|
589
|
-
},
|
|
590
|
-
output: {
|
|
591
|
-
tag: 'h1',
|
|
592
|
-
classes: ['main-title'],
|
|
593
|
-
attributes: { id: 'top' }
|
|
594
|
-
}
|
|
675
|
+
{
|
|
676
|
+
selector: { nodeType: 'paragraph', attributes: { style: 'Heading 1' } },
|
|
677
|
+
output: { tag: 'h1', classes: ['main-title'], attributes: { id: 'top' } }
|
|
595
678
|
},
|
|
596
679
|
{
|
|
597
|
-
//
|
|
680
|
+
// '~=' operator matches if the word 'Quote' appears anywhere in the style name
|
|
598
681
|
selector: { attributes: { style: { value: 'Quote', operator: '~=' } } },
|
|
599
|
-
output: { tag: 'blockquote' }
|
|
682
|
+
output: { tag: 'blockquote', fresh: true }
|
|
600
683
|
}
|
|
601
684
|
]
|
|
602
685
|
```
|
|
603
686
|
|
|
604
|
-
|
|
605
|
-
The library also maintains support for a simple string-based DSL, highly compatible with `mammoth.js`.
|
|
687
|
+
`fresh: true` prevents the generator from merging adjacent nodes of the same tag into one block.
|
|
606
688
|
|
|
607
|
-
|
|
608
|
-
- **Regex-like Matching**: `"p[style~='Title'] => h2"`
|
|
609
|
-
- **Attribute Filters**: `"p[style-name='Quote'][lang='en'] => blockquote"`
|
|
689
|
+
#### Legacy String DSL
|
|
610
690
|
|
|
611
|
-
|
|
612
|
-
Specific options when using `format: 'chunks'`.
|
|
691
|
+
Compatible with `mammoth.js` style maps:
|
|
613
692
|
|
|
614
|
-
|
|
615
|
-
|
|
616
|
-
|
|
617
|
-
|
|
618
|
-
|
|
619
|
-
|
|
620
|
-
|
|
693
|
+
```js
|
|
694
|
+
styleMap: [
|
|
695
|
+
"p[style-name='Heading 1'] => h1",
|
|
696
|
+
"p[style~='Title'] => h2",
|
|
697
|
+
"p[style-name='Quote'][lang='en'] => blockquote"
|
|
698
|
+
]
|
|
699
|
+
```
|
|
621
700
|
|
|
622
|
-
|
|
623
|
-
If your application uses OCR, `officeParser` utilizes an intelligent **Smart Worker Pool** to maintain a background worker pool and optimize repeated parse requests.
|
|
701
|
+
---
|
|
624
702
|
|
|
625
|
-
|
|
626
|
-
- **LRU Re-allocation**: If a new language is requested and the pool is full, the manager identifies the **Least Recently Used (LRU)** idle worker and re-initializes it for the new language. This avoids the overhead of destroying and recreating workers.
|
|
627
|
-
- **Auto-Termination**: Workers are automatically cleaned up after 10 seconds of inactivity (configurable via `ocrConfig.autoTerminateTimeout`).
|
|
703
|
+
### HtmlGeneratorConfig
|
|
628
704
|
|
|
629
|
-
|
|
630
|
-
If you have used OCR (`{ ocr: true }`) in a short-lived script (like CLI tools or one-off automation), we recommend explicitly calling `terminateOcr()` after your processing is finished. This bypasses the 10-second idle timer and allows the process to return to the terminal prompt immediately.
|
|
705
|
+
Pass as `htmlConfig` inside `GeneratorConfig`.
|
|
631
706
|
|
|
632
|
-
|
|
633
|
-
|
|
707
|
+
| Option | Type | Default | Description |
|
|
708
|
+
|--------|------|---------|-------------|
|
|
709
|
+
| `standalone` | `boolean` | `true` | Wrap output in a full `<html>` document with CSS |
|
|
710
|
+
| `chartJsSrc` | `string` | jsDelivr CDN | URL for the Chart.js library |
|
|
634
711
|
|
|
635
|
-
|
|
636
|
-
const officeParser = require('officeparser');
|
|
712
|
+
### MdGeneratorConfig
|
|
637
713
|
|
|
638
|
-
|
|
639
|
-
await officeParser.parseOffice("file.pdf", { ocr: true });
|
|
640
|
-
// ... process results ...
|
|
714
|
+
Pass as `mdConfig` inside `GeneratorConfig`.
|
|
641
715
|
|
|
642
|
-
|
|
643
|
-
|
|
644
|
-
|
|
645
|
-
```
|
|
716
|
+
| Option | Type | Default | Description |
|
|
717
|
+
|--------|------|---------|-------------|
|
|
718
|
+
| `fallbackToHtml` | `boolean` | `true` | Use HTML tags for features Markdown cannot represent (underlines, merged table cells, etc.) |
|
|
646
719
|
|
|
647
|
-
|
|
648
|
-
> This is handled automatically in the built-in CLI (`npx officeparser ...`). You only need to call this manually if you are using the library in your own custom script and want a snappy exit.
|
|
720
|
+
### PdfGeneratorConfig
|
|
649
721
|
|
|
650
|
-
|
|
651
|
-
const config = {
|
|
652
|
-
newlineDelimiter: "\n\n",
|
|
653
|
-
extractAttachments: true,
|
|
654
|
-
ocr: true,
|
|
655
|
-
ocrLanguage: 'eng+fra+esp' // Supports English, French, and Spanish simultaneously
|
|
656
|
-
};
|
|
722
|
+
Pass as `pdfConfig` inside `GeneratorConfig`. Requires the optional `puppeteer` peer dependency.
|
|
657
723
|
|
|
658
|
-
|
|
659
|
-
|
|
660
|
-
|
|
724
|
+
| Option | Type | Default | Description |
|
|
725
|
+
|--------|------|---------|-------------|
|
|
726
|
+
| `format` | `string` | `'A4'` | Paper format (`'A4'`, `'Letter'`, `'Legal'`, etc.) |
|
|
727
|
+
| `landscape` | `boolean` | `false` | Landscape page orientation |
|
|
728
|
+
| `printBackground` | `boolean` | `true` | Print background graphics |
|
|
729
|
+
| `margin` | `object` | `{0,0,0,0}` | Page margins (`top`, `right`, `bottom`, `left`) |
|
|
730
|
+
| `displayHeaderFooter` | `boolean` | `false` | Show print header/footer |
|
|
731
|
+
| `headerTemplate` | `string` | `''` | HTML template for the print header |
|
|
732
|
+
| `footerTemplate` | `string` | `''` | HTML template for the print footer |
|
|
733
|
+
| `scale` | `number` | `1` | Rendering scale factor |
|
|
734
|
+
| `launchOptions` | `object` | headless defaults | Puppeteer launch options (e.g., `executablePath`) |
|
|
661
735
|
|
|
662
|
-
|
|
736
|
+
### CsvGeneratorConfig
|
|
663
737
|
|
|
664
|
-
|
|
665
|
-
```ts
|
|
666
|
-
import { OfficeParser } from 'officeparser';
|
|
738
|
+
Pass as `csvConfig` inside `GeneratorConfig`.
|
|
667
739
|
|
|
668
|
-
|
|
669
|
-
|
|
670
|
-
|
|
671
|
-
|
|
672
|
-
|
|
740
|
+
| Option | Type | Default | Description |
|
|
741
|
+
|--------|------|---------|-------------|
|
|
742
|
+
| `sheets` | `string` | `''` | Sheet range to export: `'1'`, `'1-3'`, `'1,3'` (1-based). Empty = all sheets |
|
|
743
|
+
| `mergeSheets` | `boolean` | `true` | Merge all sheets into one CSV. If `false`, returns a ZIP archive |
|
|
744
|
+
| `columnDelimiter` | `string` | `','` | Output column delimiter |
|
|
745
|
+
|
|
746
|
+
### TextGeneratorConfig
|
|
747
|
+
|
|
748
|
+
Pass as `textConfig` inside `GeneratorConfig`.
|
|
749
|
+
|
|
750
|
+
| Option | Type | Default | Description |
|
|
751
|
+
|--------|------|---------|-------------|
|
|
752
|
+
| `newlineDelimiter` | `string` | `'\n'` | String inserted between structural blocks |
|
|
753
|
+
| `preserveLayout` | `boolean` | `false` | Render tables with aligned columns using whitespace |
|
|
754
|
+
|
|
755
|
+
---
|
|
756
|
+
|
|
757
|
+
### OfficeConverterConfig
|
|
758
|
+
|
|
759
|
+
Configuration for `OfficeConverter.convert(file, format, config)`.
|
|
760
|
+
|
|
761
|
+
| Option | Type | Description |
|
|
762
|
+
|--------|------|-------------|
|
|
763
|
+
| `parseConfig` | `OfficeParserConfig` | Settings for the parsing phase |
|
|
764
|
+
| `generatorConfig` | `GeneratorConfig` | Settings for the generation phase |
|
|
765
|
+
| `onWarning` | `(issue: OfficeIssue) => void` | Global warning callback (overrides phase-specific ones) |
|
|
766
|
+
|
|
767
|
+
---
|
|
768
|
+
|
|
769
|
+
### ChunkingConfig
|
|
770
|
+
|
|
771
|
+
`ChunkingConfig` is a **discriminated union** — the available options depend on the `strategy` field.
|
|
772
|
+
|
|
773
|
+
#### Common Options (all strategies)
|
|
774
|
+
|
|
775
|
+
| Option | Type | Default | Description |
|
|
776
|
+
|--------|------|---------|-------------|
|
|
777
|
+
| `strategy` | `string` | `'document-structure'` | Chunking strategy |
|
|
778
|
+
| `stripWhitespace` | `boolean` | `true` | Trim leading/trailing whitespace from each chunk |
|
|
779
|
+
| `includeMetadata` | `boolean` | `true` | Include page/slide/heading metadata in each chunk |
|
|
780
|
+
| `addStartIndex` | `boolean` | `false` | Add `startIndex` character offset to chunk metadata |
|
|
781
|
+
| `lengthFunction` | `(text) => number` | `text.length` | Custom size measurer (e.g., token counter) |
|
|
782
|
+
| `sentenceBoundaryRegex` | `string \| RegExp` | `/[.!?。!?]/` | Custom regex for sentence boundary detection |
|
|
783
|
+
| `abbreviations` | `string[]` | common list | Abbreviations to skip when splitting on `.` |
|
|
784
|
+
|
|
785
|
+
#### `strategy: 'fixed-size'`
|
|
786
|
+
|
|
787
|
+
| Option | Type | Default | Description |
|
|
788
|
+
|--------|------|---------|-------------|
|
|
789
|
+
| `chunkSize` | `number` | `1000` | Maximum characters per chunk |
|
|
790
|
+
| `chunkOverlap` | `number` | `200` | Character overlap between consecutive chunks |
|
|
791
|
+
| `separators` | `string[]` | `['\n\n','\n',' ','']` | Ordered list of separators to try |
|
|
792
|
+
|
|
793
|
+
#### `strategy: 'document-structure'`
|
|
794
|
+
|
|
795
|
+
| Option | Type | Default | Description |
|
|
796
|
+
|--------|------|---------|-------------|
|
|
797
|
+
| `splitBy` | `string` | `'paragraph'` | `'paragraph'` · `'heading'` · `'page'` · `'slide'` · `'sheet'` |
|
|
798
|
+
| `maxChunkSize` | `number` | `1000` | Max characters per chunk (oversized units are split recursively) |
|
|
799
|
+
| `tableSplitStrategy` | `string` | `'row'` | `'row'` (repeats header in each chunk) or `'flatten'` |
|
|
800
|
+
|
|
801
|
+
#### `strategy: 'semantic'`
|
|
802
|
+
|
|
803
|
+
| Option | Type | Default | Description |
|
|
804
|
+
|--------|------|---------|-------------|
|
|
805
|
+
| `embeddingFunction` | `(text) => Promise<number[]>` | **required** | Async embedding function |
|
|
806
|
+
| `similarityThreshold` | `number` | `0.8` | Cosine similarity threshold; lower = fewer boundaries |
|
|
807
|
+
| `maxChunkSize` | `number` | `2000` | Max characters even if similarity stays high |
|
|
808
|
+
| `bufferSize` | `number` | `1` | Surrounding sentences used when computing similarity |
|
|
809
|
+
| `embeddingBatchSize` | `number` | `50` | Sentences per embedding API batch |
|
|
810
|
+
|
|
811
|
+
---
|
|
812
|
+
|
|
813
|
+
## OCR Scheduler & Resource Management
|
|
814
|
+
|
|
815
|
+
When `ocr: true` is set, `officeParser` maintains an intelligent **Smart Worker Pool** backed by Tesseract.js:
|
|
816
|
+
|
|
817
|
+
- **Dynamic Affinity**: Workers persist with their last-used language, avoiding re-initialization overhead.
|
|
818
|
+
- **LRU Re-allocation**: When a new language is requested and the pool is full, the Least Recently Used idle worker is re-initialized.
|
|
819
|
+
- **Auto-Termination**: Workers shut down after 10 seconds of inactivity (configurable via `ocrConfig.autoTerminateTimeout`).
|
|
820
|
+
|
|
821
|
+
### OCR Config (`ocrConfig`)
|
|
822
|
+
|
|
823
|
+
| Option | Type | Default | Description |
|
|
824
|
+
|--------|------|---------|-------------|
|
|
825
|
+
| `language` | `string` | `'eng'` | Tesseract language code(s), e.g. `'eng+fra'` |
|
|
826
|
+
| `workerPath` | `string` | `''` | Custom path to Tesseract worker script |
|
|
827
|
+
| `corePath` | `string` | `''` | Custom path to Tesseract core script |
|
|
828
|
+
| `langPath` | `string` | `''` | Custom path for language data files |
|
|
829
|
+
| `autoTerminateTimeout` | `number` | `10000` | Inactivity timeout in ms before auto-teardown (0 = disabled) |
|
|
830
|
+
|
|
831
|
+
See all language codes at [tesseract-ocr.github.io](https://tesseract-ocr.github.io/tessdoc/Data-Files).
|
|
832
|
+
|
|
833
|
+
### `OfficeParser.terminateOcr()`
|
|
834
|
+
|
|
835
|
+
In **short-lived scripts** (CLI tools, one-off automation), call `terminateOcr()` after processing to bypass the idle timer and exit immediately:
|
|
673
836
|
|
|
674
|
-
**Extracting Images and their OCR text**
|
|
675
837
|
```js
|
|
676
838
|
const officeParser = require('officeparser');
|
|
677
839
|
|
|
678
|
-
const
|
|
679
|
-
|
|
680
|
-
|
|
681
|
-
if (attachment.type === 'image') {
|
|
682
|
-
console.log(`Image: ${attachment.name}`);
|
|
683
|
-
console.log(`OCR Text: ${attachment.ocrText}`);
|
|
684
|
-
fs.writeFileSync(attachment.name, Buffer.from(attachment.data, 'base64'));
|
|
685
|
-
}
|
|
686
|
-
});
|
|
687
|
-
});
|
|
840
|
+
const ast = await officeParser.parseOffice('file.pdf', { ocr: true });
|
|
841
|
+
// ... process results ...
|
|
842
|
+
await officeParser.terminateOcr(); // immediate exit
|
|
688
843
|
```
|
|
689
844
|
|
|
845
|
+
> [!TIP]
|
|
846
|
+
> The built-in CLI (`npx officeparser ...`) handles this automatically.
|
|
847
|
+
> Only call it manually in your own scripts.
|
|
848
|
+
|
|
849
|
+
---
|
|
850
|
+
|
|
690
851
|
## Browser Usage
|
|
691
|
-
The library provides two types of browser bundles in the `dist/` directory:
|
|
692
|
-
1. **`officeparser.browser.iife.js`**: Standard IIFE bundle for direct `<script>` tag usage. Exposes the global `officeParser` namespace.
|
|
693
|
-
2. **`officeparser.browser.mjs`**: Modern ESM bundle for use with `import` statements or modern bundlers.
|
|
694
852
|
|
|
695
|
-
|
|
696
|
-
If you are using a modern bundler like **Vite**, **Webpack**, or **Next.js**:
|
|
853
|
+
Two bundles are available in the `dist/` directory:
|
|
697
854
|
|
|
698
|
-
|
|
855
|
+
| Bundle | Usage |
|
|
856
|
+
|--------|-------|
|
|
857
|
+
| `officeparser.browser.mjs` | ESM — use with `import` statements or modern bundlers (Vite, Webpack, Next.js) |
|
|
858
|
+
| `officeparser.browser.iife.js` | IIFE — use with a `<script>` tag; exposes the global `officeParser` object |
|
|
859
|
+
|
|
860
|
+
### ESM (Vite / Webpack / Next.js)
|
|
861
|
+
|
|
862
|
+
```js
|
|
699
863
|
import { OfficeParser } from 'officeparser';
|
|
700
864
|
|
|
701
865
|
const handleFile = async (event) => {
|
|
702
866
|
const file = event.target.files[0];
|
|
703
867
|
const buffer = await file.arrayBuffer();
|
|
704
|
-
|
|
705
|
-
|
|
706
|
-
// Pass the Buffer or Uint8Array directly
|
|
707
|
-
const ast = await OfficeParser.parseOffice(new Uint8Array(buffer));
|
|
708
|
-
console.log(ast.toText());
|
|
709
|
-
} catch (err) {
|
|
710
|
-
console.error(err);
|
|
711
|
-
}
|
|
868
|
+
const ast = await OfficeParser.parseOffice(new Uint8Array(buffer));
|
|
869
|
+
console.log(ast.toText());
|
|
712
870
|
};
|
|
713
871
|
```
|
|
714
872
|
|
|
715
|
-
|
|
716
|
-
> **Why `fs` fails in the browser**: Browsers do not have a built-in file system. If you try to pass a file path string in the browser, `officeParser` will throw a descriptive "Fail-Fast" error instead of crashing mysteriously:
|
|
717
|
-
> `[officeparser] Node.js 'fs' module is not available in the browser. Please pass a Buffer or Uint8Array instead.`
|
|
718
|
-
|
|
719
|
-
### Usage (Script Tag)
|
|
720
|
-
Include the IIFE bundle available in the release assets or your `dist/` folder. This exposes the global `officeParser` object.
|
|
873
|
+
### Script Tag
|
|
721
874
|
|
|
722
875
|
```html
|
|
723
876
|
<script src="dist/officeparser.browser.iife.js"></script>
|
|
@@ -725,60 +878,78 @@ Include the IIFE bundle available in the release assets or your `dist/` folder.
|
|
|
725
878
|
async function handleFile(event) {
|
|
726
879
|
const file = event.target.files[0];
|
|
727
880
|
const buffer = await file.arrayBuffer();
|
|
728
|
-
|
|
729
|
-
|
|
730
|
-
// Reconstruct as Uint8Array for the parser
|
|
731
|
-
const ast = await officeParser.parseOffice(new Uint8Array(buffer));
|
|
732
|
-
console.log(ast.toText());
|
|
733
|
-
} catch (error) {
|
|
734
|
-
console.error("Parsing failed:", error);
|
|
735
|
-
}
|
|
881
|
+
const ast = await officeParser.parseOffice(new Uint8Array(buffer));
|
|
882
|
+
console.log(ast.toText());
|
|
736
883
|
}
|
|
737
884
|
</script>
|
|
738
885
|
```
|
|
739
886
|
|
|
740
|
-
|
|
741
|
-
|
|
887
|
+
> [!NOTE]
|
|
888
|
+
> **File paths don't work in the browser.** Always pass a `Buffer`, `ArrayBuffer`, or `Uint8Array`.
|
|
889
|
+
> Passing a path string will throw a descriptive `FEATURE_NOT_SUPPORTED_IN_BROWSER` error.
|
|
742
890
|
|
|
743
|
-
|
|
744
|
-
const file = ...; // File object or ArrayBuffer
|
|
891
|
+
### PDF Worker Configuration
|
|
745
892
|
|
|
746
|
-
|
|
747
|
-
const ast = await officeParser.parseOffice(file);
|
|
893
|
+
When parsing PDFs in the browser, a Web Worker is required. If `pdfWorkerSrc` is omitted, a jsDelivr CDN link is used automatically:
|
|
748
894
|
|
|
749
|
-
|
|
750
|
-
|
|
751
|
-
|
|
895
|
+
```js
|
|
896
|
+
// Uses default CDN worker:
|
|
897
|
+
const ast = await officeParser.parseOffice(pdfArrayBuffer);
|
|
898
|
+
|
|
899
|
+
// Or specify your own:
|
|
900
|
+
const ast = await officeParser.parseOffice(pdfArrayBuffer, {
|
|
901
|
+
pdfWorkerSrc: 'https://cdn.jsdelivr.net/npm/pdfjs-dist@5.6.205/build/pdf.worker.min.mjs'
|
|
752
902
|
});
|
|
753
903
|
```
|
|
754
904
|
|
|
755
|
-
>
|
|
905
|
+
> [!NOTE]
|
|
906
|
+
> The `pdfjs-dist` worker version must match the version bundled with `officeparser` (currently **5.6.205**).
|
|
907
|
+
|
|
908
|
+
---
|
|
756
909
|
|
|
757
910
|
## Troubleshooting & Common Issues
|
|
758
911
|
|
|
759
|
-
|
|
760
|
-
|
|
761
|
-
|
|
762
|
-
|
|
912
|
+
| Symptom | Fix |
|
|
913
|
+
|---------|-----|
|
|
914
|
+
| Node.js process stays alive after finishing | Call `await officeParser.terminateOcr()` at end of script when OCR was used |
|
|
915
|
+
| `"Worker not found"` in browser for PDF | Verify `pdfWorkerSrc` points to `pdf.worker.min.mjs` matching version `5.6.205` |
|
|
916
|
+
| Low OCR accuracy | Verify `ocrConfig.language` matches the document language; quality depends on image resolution |
|
|
917
|
+
| Out of memory on large Excel files | Call `ast.toText()` early and discard the AST object to allow garbage collection |
|
|
918
|
+
| `md`/`html`/`csv` buffer not detected | Add `fileType: 'md'` (or `'html'`, `'csv'`) to config — these formats have no magic bytes |
|
|
919
|
+
| `IMPROPER_BUFFERS` error | Usually means no file extension and no `fileType` hint was provided for a buffer input |
|
|
920
|
+
| PDF generation fails | Install the optional peer dependency: `npm install puppeteer` |
|
|
763
921
|
|
|
764
|
-
For a
|
|
922
|
+
For a full debugging guide, visit the [Live Documentation](https://harshankur.github.io/officeParser/#spec/debugging).
|
|
765
923
|
|
|
924
|
+
---
|
|
766
925
|
|
|
767
926
|
## Known Limitations
|
|
768
|
-
1. **ODT/ODS Charts**: Extraction may occasionally show inaccurate data when referencing external cell ranges or complex layout-based data.
|
|
769
|
-
2. **PDF Images**: PDF images are extracted as BMP files in the browser for compatibility. This conversion happens automatically.
|
|
770
|
-
3. **RTF Footnotes**: The `putNotesAtLast` configuration is currently not supported for RTF files; footnotes and endnotes are always collected and appended to the end of the content.
|
|
771
927
|
|
|
772
|
-
|
|
928
|
+
1. **ODT/ODS Charts**: May show inaccurate data when the chart references external cell ranges or uses complex layout-based data.
|
|
929
|
+
2. **PDF Images (Browser)**: Extracted as BMP files for cross-platform compatibility. Conversion is automatic.
|
|
930
|
+
3. **RTF Notes**: `putNotesAtLast` has no effect for RTF files; footnotes and endnotes are always appended at the end.
|
|
931
|
+
|
|
932
|
+
---
|
|
773
933
|
|
|
774
934
|
**npm**: [https://npmjs.com/package/officeparser](https://npmjs.com/package/officeparser)
|
|
775
935
|
|
|
776
936
|
**github**: [https://github.com/harshankur/officeParser](https://github.com/harshankur/officeParser)
|
|
777
937
|
|
|
938
|
+
## Support the Project
|
|
939
|
+
|
|
940
|
+
If `officeParser` has helped you save time, consider supporting its continued development. Your sponsorship helps maintain the project, add new features, and keep it robust for everyone.
|
|
941
|
+
|
|
942
|
+
<a href="https://github.com/sponsors/harshankur">
|
|
943
|
+
<img src="https://img.shields.io/badge/Sponsor-GitHub-ea4aaa?style=for-the-badge&logo=github-sponsors" height="36">
|
|
944
|
+
</a>
|
|
945
|
+
<a href="https://www.buymeacoffee.com/harshankur">
|
|
946
|
+
<img src="https://cdn.buymeacoffee.com/buttons/v2/default-yellow.png" height="36" alt="Buy Me A Coffee">
|
|
947
|
+
</a>
|
|
948
|
+
|
|
778
949
|
## Contributing
|
|
779
950
|
|
|
780
|
-
Contributions are welcome! Please see [CONTRIBUTING.md](CONTRIBUTING.md) for details
|
|
951
|
+
Contributions are welcome! Please see [CONTRIBUTING.md](CONTRIBUTING.md) for details.
|
|
781
952
|
|
|
782
953
|
## License
|
|
783
954
|
|
|
784
|
-
This project is licensed under the MIT License
|
|
955
|
+
This project is licensed under the MIT License — see the [LICENSE](LICENSE) file for details.
|