officeparser 7.0.1 → 7.0.2
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +704 -615
- package/dist/OfficeConverter.d.ts +4 -3
- package/dist/OfficeConverter.js +4 -3
- package/dist/OfficeParser.d.ts +1 -1
- package/dist/OfficeParser.js +1 -1
- package/dist/officeparser.browser.d.ts +4 -3
- package/dist/sbom.cdx.json +98 -98
- package/package.json +1 -1
package/README.md
CHANGED
|
@@ -1,6 +1,10 @@
|
|
|
1
|
-
# officeParser
|
|
1
|
+
# officeParser — Universal Office Document Parser & Generator
|
|
2
2
|
|
|
3
|
-
A robust, strictly-typed Node.js and Browser library for parsing
|
|
3
|
+
A robust, strictly-typed **Node.js and Browser** library for parsing office files into a rich **Abstract Syntax Tree (AST)** and generating high-fidelity output in multiple formats.
|
|
4
|
+
|
|
5
|
+
**Parses:** [`docx`](https://en.wikipedia.org/wiki/Office_Open_XML) · [`pptx`](https://en.wikipedia.org/wiki/Office_Open_XML) · [`xlsx`](https://en.wikipedia.org/wiki/Office_Open_XML) · [`odt`](https://en.wikipedia.org/wiki/OpenDocument) · [`odp`](https://en.wikipedia.org/wiki/OpenDocument) · [`ods`](https://en.wikipedia.org/wiki/OpenDocument) · [`pdf`](https://en.wikipedia.org/wiki/PDF) · [`rtf`](https://en.wikipedia.org/wiki/Rich_Text_Format) · [`csv`](https://en.wikipedia.org/wiki/Comma-separated_values) · [`md`](https://en.wikipedia.org/wiki/Markdown) · [`html`](https://en.wikipedia.org/wiki/HTML)
|
|
6
|
+
|
|
7
|
+
**Generates:** `Markdown` · `HTML` · `CSV` · `RTF` · `PDF` · `Plain Text` · `RAG Chunks`
|
|
4
8
|
|
|
5
9
|
[](https://badge.fury.io/js/officeparser)
|
|
6
10
|
[](https://www.npmjs.com/package/officeparser)
|
|
@@ -10,45 +14,53 @@ A robust, strictly-typed Node.js and Browser library for parsing and generating
|
|
|
10
14
|
---
|
|
11
15
|
|
|
12
16
|
### 🌟 [Live Interactive AST Visualizer & Documentation](https://harshankur.github.io/officeParser/) 🌟
|
|
13
|
-
*
|
|
14
|
-
|
|
15
|
-
**What you can do there:**
|
|
16
|
-
- **AST Visualizer**: Upload any office file and inspect the hierarchical AST structure, metadata, and raw content.
|
|
17
|
-
- **Config Configurator**: Tweak parsing options (like `ignoreNotes`, `ocr`, `newlineDelimiter`) and see the results instantly.
|
|
18
|
-
- **Debugging**: Use the visualizer to debug parsing issues by inspecting exactly how nodes are interpreted.
|
|
19
|
-
- **Format Specs**: Read detailed specifications for the AST structure and configuration options.
|
|
20
|
-
|
|
21
|
-
---
|
|
17
|
+
*Upload any office file in your browser — inspect the AST, tweak config, and preview generated output in real-time.*
|
|
22
18
|
|
|
19
|
+
- **AST Visualizer**: Inspect the hierarchical node tree, metadata, and raw content
|
|
20
|
+
- **Config Configurator**: Tweak options (`ignoreNotes`, `ocr`, `newlineDelimiter`) and see results instantly
|
|
21
|
+
- **Debugging**: Identify exactly how nodes are interpreted
|
|
22
|
+
- **Format Specs**: Read detailed specs for the AST structure and all config options
|
|
23
23
|
|
|
24
24
|
---
|
|
25
25
|
|
|
26
26
|
### 📝 [Changelog](CHANGELOG.md)
|
|
27
|
-
*Detailed release notes and the full history of updates are available in the project changelog.*
|
|
28
27
|
|
|
29
28
|
---
|
|
30
29
|
|
|
31
30
|
## Table of Contents
|
|
32
|
-
- [Install
|
|
31
|
+
- [Install](#install-via-npm)
|
|
33
32
|
- [Command Line Usage](#command-line-usage)
|
|
34
|
-
- [
|
|
35
|
-
- [
|
|
36
|
-
- [
|
|
33
|
+
- [Quick Decision Guide](#quick-decision-guide)
|
|
34
|
+
- [Library Usage: Parsing](#library-usage-parsing)
|
|
35
|
+
- [Async/Await](#asyncawait)
|
|
36
|
+
- [Callback (Backward Compat)](#callback-backward-compat)
|
|
37
|
+
- [File Buffers & ArrayBuffers](#file-buffers--arraybuffers)
|
|
38
|
+
- [`ast.to()` — Generate from AST](#astto--generate-from-ast)
|
|
39
|
+
- [`ast.toText()` — Quick Text Extraction](#asttotext--quick-text-extraction)
|
|
40
|
+
- [OfficeGenerator](#officegenerator)
|
|
41
|
+
- [OfficeConverter — One-Step API](#officeconverter--one-step-api)
|
|
37
42
|
- [Native RAG Chunking](#native-rag-chunking)
|
|
38
43
|
- [The AST Structure](#the-ast-structure)
|
|
39
44
|
- [Deep Dive: Document Components](#deep-dive-document-components)
|
|
40
|
-
- [Performance
|
|
45
|
+
- [Performance Highlights](#performance-highlights)
|
|
41
46
|
- [Advanced AST Usage](#advanced-ast-usage)
|
|
42
|
-
- [Configuration
|
|
43
|
-
- [
|
|
44
|
-
- [
|
|
45
|
-
- [
|
|
46
|
-
- [
|
|
47
|
+
- [Configuration Reference](#configuration-reference)
|
|
48
|
+
- [OfficeParserConfig](#officeparserconfig)
|
|
49
|
+
- [GeneratorConfig (Common)](#generatorconfig-common)
|
|
50
|
+
- [onNode Callback](#onnode-callback--advanced-node-manipulation)
|
|
51
|
+
- [styleMap — Semantic Style Mapping](#stylemap--semantic-style-mapping)
|
|
52
|
+
- [HtmlGeneratorConfig](#htmlgeneratorconfig)
|
|
53
|
+
- [MdGeneratorConfig](#mdgeneratorconfig)
|
|
54
|
+
- [PdfGeneratorConfig](#pdfgeneratorconfig)
|
|
55
|
+
- [CsvGeneratorConfig](#csvgeneratorconfig)
|
|
56
|
+
- [TextGeneratorConfig](#textgeneratorconfig)
|
|
57
|
+
- [OfficeConverterConfig](#officeconverterconfig)
|
|
58
|
+
- [ChunkingConfig](#chunkingconfig)
|
|
47
59
|
- [OCR Scheduler & Resource Management](#ocr-scheduler--resource-management)
|
|
48
|
-
- [Examples](#examples)
|
|
49
60
|
- [Browser Usage](#browser-usage)
|
|
50
61
|
- [Troubleshooting & Common Issues](#troubleshooting--common-issues)
|
|
51
62
|
- [Known Limitations](#known-limitations)
|
|
63
|
+
- [Contributing](#contributing)
|
|
52
64
|
|
|
53
65
|
---
|
|
54
66
|
|
|
@@ -58,748 +70,807 @@ A robust, strictly-typed Node.js and Browser library for parsing and generating
|
|
|
58
70
|
npm i officeparser
|
|
59
71
|
```
|
|
60
72
|
|
|
61
|
-
|
|
62
|
-
|
|
73
|
+
---
|
|
74
|
+
|
|
75
|
+
## Command Line Usage
|
|
63
76
|
|
|
64
77
|
```bash
|
|
65
|
-
#
|
|
66
|
-
npx officeparser /path/to/
|
|
78
|
+
# Full AST as JSON (default)
|
|
79
|
+
npx officeparser /path/to/file.docx
|
|
67
80
|
|
|
68
|
-
#
|
|
69
|
-
npx officeparser /path/to/
|
|
81
|
+
# Plain text output
|
|
82
|
+
npx officeparser /path/to/file.docx --format=text
|
|
70
83
|
|
|
71
|
-
#
|
|
84
|
+
# Convert DOCX to Markdown and save
|
|
72
85
|
npx officeparser report.docx --format=md --output=report.md
|
|
73
86
|
|
|
74
|
-
#
|
|
87
|
+
# Convert PPTX to HTML
|
|
75
88
|
npx officeparser presentation.pptx --format=html --output=preview.html
|
|
76
89
|
|
|
77
|
-
# Convert
|
|
90
|
+
# Convert XLSX to CSV
|
|
78
91
|
npx officeparser data.xlsx --format=csv
|
|
92
|
+
|
|
93
|
+
# Generate RAG chunks
|
|
94
|
+
npx officeparser document.pdf --format=chunks
|
|
79
95
|
```
|
|
80
96
|
|
|
81
|
-
###
|
|
82
|
-
- `--format=[json|text|md|html|csv|rtf|pdf|chunks]` The output format. Default is `json`.
|
|
83
|
-
- `--output=[path]` Optional file path to write the output to.
|
|
84
|
-
- `--toText=[true|false]` Legacy flag to output only plain text. Use `--format=text` instead.
|
|
85
|
-
- `--ignoreNotes=[true|false]` Flag to ignore notes from files like PowerPoint. Default is false.
|
|
86
|
-
- `--newlineDelimiter=[delimiter]` The delimiter to use for new lines. Default is `\n`.
|
|
87
|
-
- `--putNotesAtLast=[true|false]` Flag to collect notes at the end of files like PowerPoint. Default is false.
|
|
88
|
-
- `--outputErrorToConsole=[true|false]` **(Deprecated)** Flag to output errors to the console. Use `onWarning` callback in library usage.
|
|
89
|
-
- `--extractAttachments=[true|false]` Flag to extract images/charts as Base64. Default is false.
|
|
90
|
-
- `--ocr=[true|false]` Flag to enable OCR for extracted images. Default is false.
|
|
91
|
-
- `--includeRawContent=[true|false]` Flag to include raw XML/RTF content in nodes. Default is false.
|
|
92
|
-
- `--includeBreakNodes=[true|false]` Flag to include break nodes. Currently only available for DOCX documents.
|
|
93
|
-
- `--verbose=[true|false]` Show full error stack traces.
|
|
97
|
+
### CLI Options
|
|
94
98
|
|
|
99
|
+
| Flag | Values | Default | Description |
|
|
100
|
+
|------|--------|---------|-------------|
|
|
101
|
+
| `--format` | `json\|text\|md\|html\|csv\|rtf\|pdf\|chunks` | `json` | Output format |
|
|
102
|
+
| `--output` | path | — | Write output to a file |
|
|
103
|
+
| `--toText` | `true\|false` | `false` | **Deprecated.** Use `--format=text` |
|
|
104
|
+
| `--ignoreNotes` | `true\|false` | `false` | Ignore speaker notes (PPTX/ODP) |
|
|
105
|
+
| `--putNotesAtLast` | `true\|false` | `false` | Collect notes at end of output |
|
|
106
|
+
| `--newlineDelimiter` | string | `\n` | Delimiter between lines |
|
|
107
|
+
| `--extractAttachments` | `true\|false` | `false` | Extract images/charts as Base64 |
|
|
108
|
+
| `--ocr` | `true\|false` | `false` | Enable OCR for images |
|
|
109
|
+
| `--includeRawContent` | `true\|false` | `false` | Include raw XML/RTF in nodes |
|
|
110
|
+
| `--includeBreakNodes` | `true\|false` | `false` | Include break nodes (DOCX only) |
|
|
111
|
+
| `--outputErrorToConsole` | `true\|false` | `false` | **Deprecated.** Use `onWarning` callback |
|
|
112
|
+
| `--verbose` | `true\|false` | `false` | Show full error stack traces |
|
|
95
113
|
|
|
96
|
-
|
|
97
|
-
|
|
114
|
+
---
|
|
115
|
+
|
|
116
|
+
## Quick Decision Guide
|
|
117
|
+
|
|
118
|
+
| Goal | API to use |
|
|
119
|
+
|------|-----------|
|
|
120
|
+
| Extract text / AST from a file | `OfficeParser.parseOffice(file)` |
|
|
121
|
+
| Convert directly to another format | `OfficeConverter.convert(file, 'md')` |
|
|
122
|
+
| Parse first, then generate | `parseOffice()` → `OfficeGenerator.generate(ast, 'html')` |
|
|
123
|
+
| Convert on the AST itself (shorthand) | `ast.to('md')` |
|
|
124
|
+
| RAG pipeline chunking | `OfficeConverter.convert(file, 'chunks', {...})` |
|
|
125
|
+
|
|
126
|
+
---
|
|
127
|
+
|
|
128
|
+
## Library Usage: Parsing
|
|
129
|
+
|
|
130
|
+
### Async/Await
|
|
98
131
|
|
|
99
|
-
### Getting Started (Async/Await)
|
|
100
132
|
```js
|
|
101
133
|
const officeParser = require('officeparser');
|
|
102
134
|
|
|
103
|
-
|
|
104
|
-
|
|
105
|
-
|
|
106
|
-
|
|
107
|
-
|
|
108
|
-
|
|
109
|
-
|
|
110
|
-
console.log(text);
|
|
111
|
-
|
|
112
|
-
// Access structured content
|
|
113
|
-
console.log(ast.content); // Array of hierarchical nodes (paragraphs, tables, etc.)
|
|
114
|
-
console.log(ast.metadata); // Document properties (author, title, etc.)
|
|
115
|
-
} catch (err) {
|
|
116
|
-
console.error(err);
|
|
117
|
-
}
|
|
118
|
-
}
|
|
135
|
+
const ast = await officeParser.parseOffice('/path/to/file.docx');
|
|
136
|
+
|
|
137
|
+
console.log(ast.type); // 'docx'
|
|
138
|
+
console.log(ast.metadata); // { author, title, created, ... }
|
|
139
|
+
console.log(ast.content); // Array of hierarchical nodes
|
|
140
|
+
console.log(ast.attachments);// Images/charts (if extractAttachments: true)
|
|
141
|
+
console.log(ast.warnings); // Non-fatal issues from parsing phase
|
|
119
142
|
```
|
|
120
143
|
|
|
121
|
-
|
|
122
|
-
|
|
144
|
+
**TypeScript (named import):**
|
|
145
|
+
```ts
|
|
146
|
+
import { OfficeParser } from 'officeparser';
|
|
147
|
+
|
|
148
|
+
const ast = await OfficeParser.parseOffice('report.docx', {
|
|
149
|
+
extractAttachments: true,
|
|
150
|
+
ocr: true,
|
|
151
|
+
});
|
|
152
|
+
```
|
|
153
|
+
|
|
154
|
+
### Callback (Backward Compat)
|
|
155
|
+
|
|
123
156
|
```js
|
|
124
|
-
|
|
125
|
-
|
|
157
|
+
officeParser.parseOffice('/path/to/file.docx', function(ast, err) {
|
|
158
|
+
if (err) { console.error(err); return; }
|
|
159
|
+
console.log(ast.toText());
|
|
160
|
+
});
|
|
161
|
+
```
|
|
162
|
+
|
|
163
|
+
### File Buffers & ArrayBuffers
|
|
126
164
|
|
|
127
|
-
|
|
128
|
-
|
|
129
|
-
|
|
165
|
+
Pass a `Buffer`, `ArrayBuffer`, or `Uint8Array` instead of a file path:
|
|
166
|
+
|
|
167
|
+
```js
|
|
168
|
+
const fs = require('fs');
|
|
169
|
+
const buffer = fs.readFileSync('/path/to/file.pdf');
|
|
170
|
+
const ast = await officeParser.parseOffice(buffer);
|
|
171
|
+
```
|
|
172
|
+
|
|
173
|
+
> [!IMPORTANT]
|
|
174
|
+
> **Text-based formats from buffers need a `fileType` hint.**
|
|
175
|
+
> Formats like `md`, `html`, and `csv` have no magic bytes, so the parser cannot
|
|
176
|
+
> auto-detect them from a buffer. You **must** provide `fileType` in that case:
|
|
177
|
+
> ```js
|
|
178
|
+
> const ast = await officeParser.parseOffice(markdownBuffer, { fileType: 'md' });
|
|
179
|
+
> ```
|
|
180
|
+
|
|
181
|
+
### `ast.to()` — Generate from AST
|
|
182
|
+
|
|
183
|
+
The preferred way to convert a parsed AST to another format. Returns a `ConversionResult`.
|
|
184
|
+
|
|
185
|
+
```ts
|
|
186
|
+
// ConversionResult shape:
|
|
187
|
+
// { value: string | Uint8Array | OfficeChunk[], messages: OfficeIssue[] }
|
|
188
|
+
|
|
189
|
+
const { value: markdown, messages } = await ast.to('md');
|
|
190
|
+
const { value: html } = await ast.to('html', { includeFormatting: false });
|
|
191
|
+
const { value: chunks } = await ast.to('chunks', { strategy: 'fixed-size', chunkSize: 800 });
|
|
192
|
+
const { value: pdfBytes } = await ast.to('pdf'); // Uint8Array
|
|
193
|
+
```
|
|
194
|
+
|
|
195
|
+
### `ast.toText()` — Quick Text Extraction
|
|
196
|
+
|
|
197
|
+
> [!NOTE]
|
|
198
|
+
> `toText()` is **synchronous** and deprecated in favour of the async `ast.to('text')`.
|
|
199
|
+
> It remains available for backward compatibility.
|
|
200
|
+
|
|
201
|
+
```js
|
|
202
|
+
const text = ast.toText(); // synchronous, returns plain string
|
|
130
203
|
```
|
|
131
204
|
|
|
132
|
-
|
|
133
|
-
|
|
205
|
+
---
|
|
206
|
+
|
|
207
|
+
## OfficeGenerator
|
|
208
|
+
|
|
209
|
+
Use `OfficeGenerator.generate(ast, format, config?)` when you need to produce output from an already-parsed AST:
|
|
134
210
|
|
|
135
|
-
```
|
|
211
|
+
```ts
|
|
136
212
|
import { OfficeParser, OfficeGenerator } from 'officeparser';
|
|
137
213
|
|
|
138
214
|
const ast = await OfficeParser.parseOffice('report.docx');
|
|
139
215
|
|
|
140
|
-
//
|
|
141
|
-
const md = await OfficeGenerator.generate(ast, 'md');
|
|
142
|
-
console.log(md.value);
|
|
216
|
+
// Convert to Markdown
|
|
217
|
+
const { value: md } = await OfficeGenerator.generate(ast, 'md');
|
|
143
218
|
|
|
144
|
-
//
|
|
145
|
-
const html = await OfficeGenerator.generate(ast, 'html', {
|
|
219
|
+
// Convert to HTML with style mapping
|
|
220
|
+
const { value: html } = await OfficeGenerator.generate(ast, 'html', {
|
|
146
221
|
includeFormatting: true,
|
|
147
222
|
styleMap: [
|
|
148
|
-
{
|
|
149
|
-
selector: { nodeType: 'paragraph', attributes: { style: 'Heading 1' } },
|
|
150
|
-
output: { tag: 'h1', classes: ['main-title'] }
|
|
223
|
+
{
|
|
224
|
+
selector: { nodeType: 'paragraph', attributes: { style: 'Heading 1' } },
|
|
225
|
+
output: { tag: 'h1', classes: ['main-title'] }
|
|
151
226
|
}
|
|
152
227
|
]
|
|
153
228
|
});
|
|
154
|
-
console.log(html.value);
|
|
155
229
|
|
|
156
|
-
//
|
|
157
|
-
const csv = await OfficeGenerator.generate(ast, 'csv');
|
|
158
|
-
console.log(csv.value);
|
|
230
|
+
// Convert to CSV (spreadsheets)
|
|
231
|
+
const { value: csv } = await OfficeGenerator.generate(ast, 'csv');
|
|
159
232
|
```
|
|
160
233
|
|
|
161
|
-
|
|
162
|
-
|
|
234
|
+
**Supported destinations:** `'text'` · `'md'` · `'html'` · `'csv'` · `'rtf'` · `'pdf'` · `'chunks'`
|
|
235
|
+
|
|
236
|
+
> [!NOTE]
|
|
237
|
+
> **PDF generation** requires the optional `puppeteer` peer dependency:
|
|
238
|
+
> ```bash
|
|
239
|
+
> npm install puppeteer
|
|
240
|
+
> ```
|
|
163
241
|
|
|
164
|
-
|
|
242
|
+
---
|
|
243
|
+
|
|
244
|
+
## OfficeConverter — One-Step API
|
|
245
|
+
|
|
246
|
+
`OfficeConverter.convert()` combines parsing and generation in a single call. It automatically syncs parser options from generator config (e.g., enables `extractAttachments` when images are requested).
|
|
247
|
+
|
|
248
|
+
```ts
|
|
165
249
|
import { OfficeConverter } from 'officeparser';
|
|
166
250
|
|
|
167
|
-
//
|
|
168
|
-
const
|
|
169
|
-
console.log(result.value); // The generated Markdown string
|
|
170
|
-
console.log(result.messages); // Array of warnings/info (e.g., "Skipped unsupported drawing")
|
|
251
|
+
// Minimal usage
|
|
252
|
+
const { value: markdown } = await OfficeConverter.convert('report.docx', 'md');
|
|
171
253
|
|
|
172
|
-
//
|
|
173
|
-
const
|
|
174
|
-
parseConfig: {
|
|
175
|
-
ignoreNotes: true
|
|
254
|
+
// With config
|
|
255
|
+
const { value: html, messages } = await OfficeConverter.convert('data.xlsx', 'html', {
|
|
256
|
+
parseConfig: {
|
|
257
|
+
ignoreNotes: true,
|
|
258
|
+
newlineDelimiter: '\n\n',
|
|
176
259
|
},
|
|
177
260
|
generatorConfig: {
|
|
178
261
|
includeFormatting: true,
|
|
179
262
|
styleMap: [
|
|
180
|
-
{
|
|
181
|
-
selector: { attributes: { style: { value: 'Header', operator: '~=' } } },
|
|
182
|
-
output: { tag: 'h2', classes: ['data-header'] }
|
|
263
|
+
{
|
|
264
|
+
selector: { attributes: { style: { value: 'Header', operator: '~=' } } },
|
|
265
|
+
output: { tag: 'h2', classes: ['data-header'] }
|
|
183
266
|
}
|
|
184
267
|
]
|
|
185
268
|
},
|
|
186
|
-
onWarning: (
|
|
269
|
+
onWarning: (issue) => console.warn(`[${issue.code}] ${issue.message}`)
|
|
187
270
|
});
|
|
188
271
|
```
|
|
189
272
|
|
|
273
|
+
> [!IMPORTANT]
|
|
274
|
+
> The `OfficeConverterConfig` shape uses **nested** `parseConfig` and `generatorConfig` sub-objects.
|
|
275
|
+
> Do **not** put parser or generator options at the top level — only `onWarning` lives there.
|
|
276
|
+
|
|
277
|
+
---
|
|
278
|
+
|
|
190
279
|
## Native RAG Chunking
|
|
191
|
-
|
|
192
|
-
|
|
193
|
-
|
|
194
|
-
|
|
195
|
-
|
|
196
|
-
|
|
197
|
-
|
|
198
|
-
|
|
199
|
-
|
|
200
|
-
|
|
201
|
-
|
|
202
|
-
|
|
203
|
-
|
|
204
|
-
|
|
205
|
-
|
|
206
|
-
|
|
207
|
-
|
|
208
|
-
sourceType: string; // e.g., "docx", "pdf"
|
|
209
|
-
pageNumber?: number; // Current page
|
|
210
|
-
slideNumber?: number; // Current slide
|
|
211
|
-
closestHeading?: string; // The heading this chunk belongs to
|
|
212
|
-
chunkIndex: number; // Sequential index
|
|
213
|
-
}
|
|
214
|
-
}
|
|
280
|
+
|
|
281
|
+
`officeParser` provides native document chunking for Retrieval-Augmented Generation (RAG) pipelines with three strategies:
|
|
282
|
+
|
|
283
|
+
### Strategy 1: Document Structure (Default)
|
|
284
|
+
Splits at natural AST boundaries (paragraphs, headings, pages, slides, sheets). Preserves logical flow.
|
|
285
|
+
|
|
286
|
+
```ts
|
|
287
|
+
const { value: chunks } = await OfficeConverter.convert('report.docx', 'chunks', {
|
|
288
|
+
generatorConfig: {
|
|
289
|
+
chunksConfig: {
|
|
290
|
+
strategy: 'document-structure',
|
|
291
|
+
splitBy: 'heading', // 'paragraph' | 'heading' | 'page' | 'slide' | 'sheet'
|
|
292
|
+
maxChunkSize: 1500,
|
|
293
|
+
tableSplitStrategy: 'row', // repeats header row in every chunk — ideal for RAG
|
|
294
|
+
}
|
|
295
|
+
}
|
|
296
|
+
});
|
|
215
297
|
```
|
|
216
298
|
|
|
217
|
-
|
|
218
|
-
|
|
219
|
-
|
|
220
|
-
|
|
221
|
-
|
|
222
|
-
|
|
299
|
+
### Strategy 2: Fixed-Size (Recursive)
|
|
300
|
+
Splits by character count with overlap. Equivalent to LangChain's `RecursiveCharacterTextSplitter`.
|
|
301
|
+
|
|
302
|
+
```ts
|
|
303
|
+
const { value: chunks } = await OfficeConverter.convert('report.docx', 'chunks', {
|
|
304
|
+
generatorConfig: {
|
|
305
|
+
chunksConfig: {
|
|
306
|
+
strategy: 'fixed-size',
|
|
307
|
+
chunkSize: 1000,
|
|
308
|
+
chunkOverlap: 200,
|
|
309
|
+
}
|
|
310
|
+
}
|
|
223
311
|
});
|
|
224
|
-
console.log(`Generated ${chunks.
|
|
312
|
+
console.log(`Generated ${chunks.length} chunks`);
|
|
225
313
|
```
|
|
226
314
|
|
|
227
|
-
###
|
|
228
|
-
|
|
229
|
-
```js
|
|
230
|
-
const officeParser = require('officeparser');
|
|
315
|
+
### Strategy 3: Semantic
|
|
316
|
+
Uses cosine similarity between sentence embeddings to find topic boundaries. Requires you to provide an `embeddingFunction`.
|
|
231
317
|
|
|
232
|
-
|
|
233
|
-
|
|
234
|
-
|
|
235
|
-
|
|
318
|
+
```ts
|
|
319
|
+
import OpenAI from 'openai';
|
|
320
|
+
const openai = new OpenAI();
|
|
321
|
+
|
|
322
|
+
const { value: chunks } = await OfficeConverter.convert('report.docx', 'chunks', {
|
|
323
|
+
generatorConfig: {
|
|
324
|
+
chunksConfig: {
|
|
325
|
+
strategy: 'semantic',
|
|
326
|
+
embeddingFunction: async (text) => {
|
|
327
|
+
const res = await openai.embeddings.create({
|
|
328
|
+
input: text, model: 'text-embedding-3-small'
|
|
329
|
+
});
|
|
330
|
+
return res.data[0].embedding;
|
|
331
|
+
},
|
|
332
|
+
similarityThreshold: 0.8,
|
|
333
|
+
maxChunkSize: 2000,
|
|
334
|
+
}
|
|
236
335
|
}
|
|
237
|
-
// Get text from AST
|
|
238
|
-
console.log(ast.toText());
|
|
239
336
|
});
|
|
240
337
|
```
|
|
241
338
|
|
|
242
|
-
###
|
|
243
|
-
You can pass a file path string, a Node.js `Buffer`, or an `ArrayBuffer`.
|
|
244
|
-
```js
|
|
245
|
-
const fs = require('fs');
|
|
246
|
-
const officeParser = require('officeparser');
|
|
247
|
-
const buffer = fs.readFileSync("/path/to/officeFile.pdf");
|
|
339
|
+
### The `OfficeChunk` Object
|
|
248
340
|
|
|
249
|
-
|
|
250
|
-
|
|
251
|
-
|
|
341
|
+
Every chunk contains text and rich metadata for citations and filtered retrieval:
|
|
342
|
+
|
|
343
|
+
```ts
|
|
344
|
+
interface OfficeChunk {
|
|
345
|
+
text: string;
|
|
346
|
+
/** Rich metadata for filtered retrieval */
|
|
347
|
+
metadata: {
|
|
348
|
+
sourceType: string; // e.g., 'docx', 'pdf'
|
|
349
|
+
pageNumber?: number; // (PDF only)
|
|
350
|
+
slideNumber?: number; // (PPTX only)
|
|
351
|
+
sheetName?: string; // (XLSX only)
|
|
352
|
+
closestHeading?: string; // Nearest heading above this chunk
|
|
353
|
+
isTableChunk?: boolean; // True if part of a split table
|
|
354
|
+
};
|
|
355
|
+
startIndex?: number; // Character offset (if addStartIndex: true)
|
|
356
|
+
endIndex?: number; // End character offset (if addStartIndex: true)
|
|
357
|
+
}
|
|
252
358
|
```
|
|
253
359
|
|
|
360
|
+
---
|
|
361
|
+
|
|
254
362
|
## The AST Structure
|
|
255
|
-
The `OfficeParserAST` provides a format-agnostic representation of your document, allowing you to traverse and manipulate content as a tree.
|
|
256
363
|
|
|
257
|
-
|
|
258
|
-
The `OfficeParserAST` provides a format-agnostic representation of your document. Below is a simplified visualization of how the tree is structured:
|
|
364
|
+
`OfficeParserAST` is a format-agnostic document representation:
|
|
259
365
|
|
|
260
366
|
```text
|
|
261
367
|
OfficeParserAST
|
|
262
|
-
├── type:
|
|
263
|
-
├── metadata: { author, title, created, modified,
|
|
368
|
+
├── type: 'docx' | 'pdf' | 'xlsx' | 'csv' | 'md' | ... (11 formats)
|
|
369
|
+
├── metadata: { author, title, created, modified, customProperties, styleMap, ... }
|
|
264
370
|
├── content: [ OfficeContentNode ]
|
|
265
|
-
│ ├── type:
|
|
266
|
-
│ ├── text:
|
|
267
|
-
│ ├── children: [ OfficeContentNode ]
|
|
268
|
-
│ ├── formatting: { bold, italic, color, size, font, ... }
|
|
269
|
-
│
|
|
270
|
-
|
|
271
|
-
├──
|
|
272
|
-
│ ├──
|
|
273
|
-
│ ├──
|
|
274
|
-
│ ├── data:
|
|
275
|
-
│ ├── ocrText
|
|
276
|
-
│ └── chartData
|
|
277
|
-
|
|
278
|
-
|
|
279
|
-
|
|
280
|
-
|
|
281
|
-
|
|
282
|
-
|
|
283
|
-
|
|
284
|
-
|
|
285
|
-
|
|
286
|
-
|
|
287
|
-
|
|
288
|
-
|
|
289
|
-
|
|
290
|
-
|
|
291
|
-
|
|
292
|
-
|
|
293
|
-
},
|
|
294
|
-
{
|
|
295
|
-
"type": "paragraph",
|
|
296
|
-
"text": "This is a report with an image.",
|
|
297
|
-
"children": [
|
|
298
|
-
{ "type": "text", "text": "This is a report with an " },
|
|
299
|
-
{ "type": "image", "metadata": { "attachmentName": "img1.png" } }
|
|
300
|
-
]
|
|
301
|
-
}
|
|
302
|
-
],
|
|
303
|
-
"attachments": [
|
|
304
|
-
{ "name": "img1.png", "type": "image", "data": "iVBOR...", "ocrText": "Extracted Text" }
|
|
305
|
-
]
|
|
371
|
+
│ ├── type: 'paragraph' | 'heading' | 'table' | 'list' | 'image' | 'chart' | ...
|
|
372
|
+
│ ├── text: string (concatenated text of node + all descendants)
|
|
373
|
+
│ ├── children: [ OfficeContentNode ] (recursive)
|
|
374
|
+
│ ├── formatting: { bold, italic, underline, color, size, font, alignment, ... }
|
|
375
|
+
│ └── metadata: { level, listId, row, col, rowSpan, colSpan, style, ... }
|
|
376
|
+
├── attachments: [ OfficeAttachment ] (populated when extractAttachments: true)
|
|
377
|
+
│ ├── type: 'image' | 'chart'
|
|
378
|
+
│ ├── name: string
|
|
379
|
+
│ ├── mimeType: string
|
|
380
|
+
│ ├── data: string (Base64)
|
|
381
|
+
│ ├── ocrText?: string (if ocr: true)
|
|
382
|
+
│ └── chartData?: { title, dataSets, labels }
|
|
383
|
+
├── warnings: OfficeIssue[] (non-fatal issues from the parsing phase)
|
|
384
|
+
├── to(format, config?) (format: 'html'|'md'|'text'|'csv'|'rtf'|'pdf'|'chunks', returns { value, messages })
|
|
385
|
+
└── toText() (Deprecated: use .to('text') instead)
|
|
386
|
+
```
|
|
387
|
+
|
|
388
|
+
### `OfficeIssue` — Warning / Error Object
|
|
389
|
+
|
|
390
|
+
All warnings and errors (from both parsing and generation) use this shape:
|
|
391
|
+
|
|
392
|
+
```ts
|
|
393
|
+
interface OfficeIssue {
|
|
394
|
+
type: 'warning' | 'info' | 'error';
|
|
395
|
+
code: OfficeWarningType | OfficeErrorType; // typed enum, e.g. 'OCR_FAILED'
|
|
396
|
+
message: string;
|
|
397
|
+
node?: OfficeContentNode; // the node that triggered the issue, if any
|
|
398
|
+
details?: any; // original error or extra context
|
|
306
399
|
}
|
|
307
400
|
```
|
|
308
401
|
|
|
402
|
+
---
|
|
403
|
+
|
|
309
404
|
## Deep Dive: Document Components
|
|
310
405
|
|
|
311
|
-
### 1.
|
|
312
|
-
Lists are represented as sequential `list` nodes. To reconstruct or track a list, use the `metadata` fields:
|
|
406
|
+
### 1. Lists
|
|
313
407
|
|
|
314
408
|
```text
|
|
315
409
|
List Node
|
|
316
|
-
├── type:
|
|
317
|
-
├── metadata: {
|
|
318
|
-
|
|
319
|
-
|
|
320
|
-
|
|
321
|
-
|
|
322
|
-
|
|
323
|
-
}
|
|
324
|
-
└── children: [ Text
|
|
410
|
+
├── type: 'list'
|
|
411
|
+
├── metadata: {
|
|
412
|
+
│ listId: '1', // items with the same listId belong to one logical list
|
|
413
|
+
│ listType: 'ordered' | 'unordered',
|
|
414
|
+
│ indentation: 0, // nesting level (0-based)
|
|
415
|
+
│ itemIndex: 0, // sequential position within the list level
|
|
416
|
+
│ paragraphIndentation: { left, hanging, right, firstLine }
|
|
417
|
+
│ }
|
|
418
|
+
└── children: [ Text content ]
|
|
325
419
|
```
|
|
326
420
|
|
|
327
|
-
- **`listId`**: A unique identifier for the list definition. Multiple items with the same `listId` belong to the same logical list.
|
|
328
|
-
- **`indentation`**: The structural nesting level (0-based).
|
|
329
|
-
- **`paragraphIndentation`**: The physical indentation formatting in twentieths of a point (twips) (e.g., `left`, `right`, `firstLine`, `hanging`).
|
|
330
|
-
- **`itemIndex`**: The sequential position within that list level.
|
|
331
|
-
- **`listType`**: Either `ordered` (numbered) or `unordered` (bulleted).
|
|
332
|
-
|
|
333
421
|
> [!TIP]
|
|
334
|
-
> Even if a list is interrupted by a regular paragraph,
|
|
422
|
+
> Even if a list is interrupted by a regular paragraph, `itemIndex` keeps incrementing for the same `listId`, so numbering stays correct.
|
|
335
423
|
|
|
336
|
-
### 2.
|
|
337
|
-
|
|
424
|
+
### 2. Tables
|
|
425
|
+
|
|
426
|
+
Tables follow a strict `table → row → cell` hierarchy:
|
|
338
427
|
|
|
339
428
|
```text
|
|
340
|
-
Table Node
|
|
341
|
-
|
|
342
|
-
└── children:
|
|
343
|
-
|
|
344
|
-
|
|
345
|
-
├── type: "cell"
|
|
346
|
-
├── metadata: { row, col, rowSpan, colSpan }
|
|
347
|
-
└── children: [ Paragraph/List/etc. ]
|
|
429
|
+
Table Node (type: 'table')
|
|
430
|
+
└── children: Row Nodes (type: 'row')
|
|
431
|
+
└── children: Cell Nodes (type: 'cell')
|
|
432
|
+
├── metadata: { row, col, rowSpan?, colSpan? }
|
|
433
|
+
└── children: [ Paragraph | List | Table | ... ]
|
|
348
434
|
```
|
|
349
435
|
|
|
350
|
-
-
|
|
351
|
-
-
|
|
352
|
-
-
|
|
436
|
+
- `row` / `col`: zero-based grid position
|
|
437
|
+
- `rowSpan` / `colSpan`: merged cells (primarily ODF formats)
|
|
438
|
+
- Cells can contain nested tables
|
|
353
439
|
|
|
354
|
-
### 3.
|
|
355
|
-
When a chart is discovered, it's added as a `chart` node in the content and a corresponding `OfficeAttachment`.
|
|
440
|
+
### 3. Images & OCR
|
|
356
441
|
|
|
357
442
|
```text
|
|
358
|
-
|
|
359
|
-
├──
|
|
360
|
-
|
|
361
|
-
└── Attachment (Linked)
|
|
362
|
-
└── chartData: { title, dataSets: [...], labels: [...] }
|
|
443
|
+
Image Node (type: 'image')
|
|
444
|
+
├── metadata: { attachmentName: 'img1.png', altText: '...' }
|
|
445
|
+
└── → Attachment: { data: 'base64...', ocrText: '...' }
|
|
363
446
|
```
|
|
364
447
|
|
|
365
|
-
-
|
|
366
|
-
-
|
|
448
|
+
- Set `extractAttachments: true` to populate `attachment.data`
|
|
449
|
+
- Set `ocr: true` (requires `extractAttachments: true`) to populate `ocrText`
|
|
367
450
|
|
|
368
|
-
### 4.
|
|
369
|
-
Images are linked via `attachmentName` and can contain valuable metadata:
|
|
451
|
+
### 4. Charts
|
|
370
452
|
|
|
371
453
|
```text
|
|
372
|
-
|
|
373
|
-
├──
|
|
374
|
-
|
|
375
|
-
└── Attachment (Linked)
|
|
376
|
-
├── data: "base64..."
|
|
377
|
-
└── ocrText: "Extracted via OCR"
|
|
454
|
+
Chart Node (type: 'chart')
|
|
455
|
+
├── metadata: { attachmentName: 'chart1.xml' }
|
|
456
|
+
└── → Attachment: { chartData: { title, dataSets, labels } }
|
|
378
457
|
```
|
|
379
458
|
|
|
380
|
-
- **OCR Text**: If `ocr: true` is set in config, `ocrText` will contain the text found within the image.
|
|
381
|
-
- **Alt Text**: Extracted from the document's internal image descriptions.
|
|
382
|
-
- **Formatting**: `OfficeContentNode` images may also have parent alignment metadata.
|
|
383
|
-
|
|
384
459
|
### 5. Text Formatting
|
|
385
|
-
Each `OfficeContentNode` can have a `formatting` object that defines how the text should be styled.
|
|
386
460
|
|
|
387
|
-
```
|
|
388
|
-
|
|
389
|
-
|
|
390
|
-
|
|
391
|
-
|
|
392
|
-
|
|
393
|
-
|
|
394
|
-
|
|
395
|
-
|
|
396
|
-
|
|
397
|
-
|
|
398
|
-
|
|
399
|
-
|
|
400
|
-
alignment: "left" | "center" | "right" | "justify"
|
|
461
|
+
```ts
|
|
462
|
+
formatting: {
|
|
463
|
+
bold?: boolean
|
|
464
|
+
italic?: boolean
|
|
465
|
+
underline?: boolean
|
|
466
|
+
strikethrough?: boolean
|
|
467
|
+
color?: string // '#RRGGBB'
|
|
468
|
+
backgroundColor?: string
|
|
469
|
+
size?: string // e.g. '12pt'
|
|
470
|
+
font?: string
|
|
471
|
+
subscript?: boolean
|
|
472
|
+
superscript?: boolean
|
|
473
|
+
alignment?: 'left' | 'center' | 'right' | 'justify'
|
|
401
474
|
}
|
|
402
475
|
```
|
|
403
476
|
|
|
404
|
-
|
|
405
|
-
1. **Node Level**: Applied directly to a text run or paragraph.
|
|
406
|
-
2. **Document Level**: Found in `ast.metadata.formatting` (defaults) or `ast.metadata.styleMap` (named styles).
|
|
477
|
+
### 6. Break Nodes (DOCX only)
|
|
407
478
|
|
|
408
|
-
|
|
409
|
-
Breaks are currently only supported when parsing DOCX-documents. Breaks are added as a node of type `break` and carry metadata of the type `BreakMetadata`. When `includeRawContent` is enabled, they also include the `rawContent` string from the original XML.
|
|
479
|
+
When `includeBreakNodes: true`, break elements appear as nodes:
|
|
410
480
|
|
|
411
481
|
```text
|
|
412
|
-
Break Node
|
|
413
|
-
├── type: "break"
|
|
482
|
+
Break Node (type: 'break')
|
|
414
483
|
└── metadata: {
|
|
415
|
-
breakType:
|
|
416
|
-
clear?:
|
|
484
|
+
breakType: 'textWrapping' | 'page' | 'column' | 'lastRenderedPage' | 'carriageReturn',
|
|
485
|
+
clear?: 'all' | 'left' | 'none' | 'right'
|
|
417
486
|
}
|
|
418
487
|
```
|
|
419
488
|
|
|
420
|
-
- `breakType`: Type of break. `textWrapping` (default) is a standard line break, `page` is a page break, `column` is a break to the next column, `lastRenderedPage` is a soft break inserted by Word, and `carriageReturn` is an explicit carriage return (`w:cr`).
|
|
421
|
-
- `clear`: Relevant for `textWrapping`. Indicates if text should wrap around floating objects.
|
|
422
|
-
|
|
423
489
|
> [!NOTE]
|
|
424
|
-
>
|
|
425
|
-
|
|
426
|
-
### 7.
|
|
427
|
-
|
|
428
|
-
|
|
429
|
-
|
|
430
|
-
|
|
431
|
-
|
|
432
|
-
|
|
433
|
-
|
|
434
|
-
|
|
435
|
-
|
|
436
|
-
|
|
437
|
-
|
|
438
|
-
|
|
439
|
-
```
|
|
440
|
-
|
|
441
|
-
## Performance & Fidelity Highlights (v7.0.0)
|
|
442
|
-
The v7.0.0 release brings significant internal optimizations and fidelity improvements:
|
|
443
|
-
- **OpenOffice Speedups**: Up to **23x faster** parsing for ODP presentations thanks to optimized XML caching.
|
|
444
|
-
- **Excel Memory Efficiency**: Resolved $O(n)$ memory overhead issues for large spreadsheets (#91) by switching to iterative stream-based parsing.
|
|
445
|
-
- **RTF Performance**: Rewritten core loop to resolve $O(n^2)$ bottlenecks during string accumulation.
|
|
446
|
-
- **Advanced Table Fidelity**: Native support for **vertical cell merging** (`vMerge`) and **horizontal spanning** (`gridSpan`) in DOCX, ensuring complex tables look exactly as they do in Word.
|
|
447
|
-
- **Parser Extensions**: You can now parse `CSV`, `Markdown`, and `HTML` files *into* the unified Office AST, allowing you to use the `OfficeGenerator` on them just like any other format.
|
|
448
|
-
|
|
449
|
-
### Advanced AST Usage
|
|
450
|
-
Beyond using `ast.toText()`, you can interact with the structural data directly:
|
|
451
|
-
|
|
452
|
-
#### 1. Extract all images and their OCR text
|
|
453
|
-
```javascript
|
|
454
|
-
const ast = await officeParser.parseOffice("report.docx", { ocr: true });
|
|
455
|
-
const images = ast.attachments.filter(a => a.mimeType.startsWith('image/'));
|
|
456
|
-
images.forEach(img => {
|
|
457
|
-
console.log(`Image: ${img.name} (OCR: ${img.ocrText || 'N/A'})`);
|
|
458
|
-
});
|
|
490
|
+
> Break nodes have no `text` property, but `ast.toText()` and `ast.to('text')` automatically convert them to the configured newline delimiter.
|
|
491
|
+
|
|
492
|
+
### 7. Document Metadata
|
|
493
|
+
|
|
494
|
+
```ts
|
|
495
|
+
ast.metadata = {
|
|
496
|
+
author?: string
|
|
497
|
+
title?: string
|
|
498
|
+
created?: Date
|
|
499
|
+
modified?: Date
|
|
500
|
+
description?: string
|
|
501
|
+
customProperties?: Record<string, any> // user-defined metadata from the document
|
|
502
|
+
styleMap?: Record<string, TextFormatting> // named styles → formatting definitions
|
|
503
|
+
formatting?: TextFormatting // document-wide defaults
|
|
504
|
+
}
|
|
459
505
|
```
|
|
460
506
|
|
|
461
|
-
|
|
462
|
-
```
|
|
463
|
-
const
|
|
464
|
-
console.log(
|
|
507
|
+
**Accessing custom properties:**
|
|
508
|
+
```js
|
|
509
|
+
const ast = await officeParser.parseOffice('contract.docx');
|
|
510
|
+
console.log(ast.metadata.customProperties);
|
|
511
|
+
// { "ProjectID": "ABC-123", "InternalReview": true }
|
|
465
512
|
```
|
|
466
513
|
|
|
467
|
-
|
|
468
|
-
|
|
469
|
-
|
|
470
|
-
|
|
471
|
-
|
|
472
|
-
|
|
473
|
-
|
|
474
|
-
|
|
475
|
-
|
|
476
|
-
|
|
477
|
-
|
|
514
|
+
---
|
|
515
|
+
|
|
516
|
+
## Performance Highlights
|
|
517
|
+
|
|
518
|
+
Key internal optimizations shipped in recent versions:
|
|
519
|
+
|
|
520
|
+
- **OpenOffice (ODP)**: Up to **23× faster** parsing via optimized XML pre-parsing and style caching
|
|
521
|
+
- **Excel Memory**: Resolved O(n) memory overhead on large sparse spreadsheets using iterative stream-based parsing
|
|
522
|
+
- **RTF Parser**: Rewrote string accumulation loop to eliminate O(n²) bottleneck in large files
|
|
523
|
+
- **Table Fidelity (DOCX)**: Native support for vertical cell merging (`vMerge`) and horizontal spanning (`gridSpan`)
|
|
524
|
+
|
|
525
|
+
---
|
|
526
|
+
|
|
527
|
+
## Advanced AST Usage
|
|
528
|
+
|
|
529
|
+
### Extract all headings
|
|
530
|
+
```js
|
|
531
|
+
const headings = ast.content.filter(n => n.type === 'heading' && n.metadata?.level === 1);
|
|
532
|
+
console.log(headings.map(h => h.text));
|
|
478
533
|
```
|
|
479
534
|
|
|
480
|
-
|
|
481
|
-
|
|
482
|
-
|
|
483
|
-
|
|
484
|
-
|
|
535
|
+
### Extract images with OCR text
|
|
536
|
+
```js
|
|
537
|
+
const ast = await officeParser.parseOffice('report.docx', { extractAttachments: true, ocr: true });
|
|
538
|
+
ast.attachments.filter(a => a.mimeType?.startsWith('image/')).forEach(img => {
|
|
539
|
+
console.log(`${img.name}: ${img.ocrText ?? 'no OCR'}`);
|
|
540
|
+
});
|
|
541
|
+
```
|
|
542
|
+
|
|
543
|
+
### Extract tables to CSV manually
|
|
544
|
+
```js
|
|
545
|
+
ast.content.filter(n => n.type === 'table').forEach((table, i) => {
|
|
485
546
|
const csv = table.children
|
|
486
|
-
.filter(
|
|
487
|
-
.map(
|
|
488
|
-
|
|
489
|
-
|
|
490
|
-
.map(cell => `"${cell.text.replace(/"/g, '""')}"`) // Escape quotes
|
|
491
|
-
.join(',')
|
|
492
|
-
)
|
|
547
|
+
.filter(r => r.type === 'row')
|
|
548
|
+
.map(r => r.children.filter(c => c.type === 'cell')
|
|
549
|
+
.map(c => `"${c.text.replace(/"/g, '""')}"`)
|
|
550
|
+
.join(','))
|
|
493
551
|
.join('\n');
|
|
494
|
-
console.log(`Table ${
|
|
552
|
+
console.log(`Table ${i + 1}:\n${csv}`);
|
|
495
553
|
});
|
|
496
554
|
```
|
|
497
555
|
|
|
498
|
-
|
|
499
|
-
|
|
500
|
-
|
|
501
|
-
|
|
502
|
-
|
|
503
|
-
|
|
504
|
-
|
|
505
|
-
results.push(node.text);
|
|
506
|
-
}
|
|
507
|
-
if (node.children) {
|
|
508
|
-
results = results.concat(findBoldText(node.children));
|
|
509
|
-
}
|
|
510
|
-
});
|
|
511
|
-
return results;
|
|
556
|
+
### Find all bold text runs
|
|
557
|
+
```js
|
|
558
|
+
function findBold(nodes) {
|
|
559
|
+
return nodes.flatMap(n => [
|
|
560
|
+
...(n.type === 'text' && n.formatting?.bold ? [n.text] : []),
|
|
561
|
+
...(n.children ? findBold(n.children) : [])
|
|
562
|
+
]);
|
|
512
563
|
}
|
|
513
|
-
|
|
514
|
-
const boldStrings = findBoldText(ast.content);
|
|
515
|
-
console.log("Bold Text Found:", boldStrings);
|
|
564
|
+
console.log(findBold(ast.content));
|
|
516
565
|
```
|
|
517
566
|
|
|
518
|
-
|
|
519
|
-
|
|
520
|
-
```javascript
|
|
567
|
+
### Extract footnotes / endnotes
|
|
568
|
+
```js
|
|
521
569
|
function extractNotes(nodes) {
|
|
522
|
-
|
|
523
|
-
|
|
524
|
-
|
|
525
|
-
|
|
526
|
-
}
|
|
527
|
-
if (node.children) {
|
|
528
|
-
notes = notes.concat(extractNotes(node.children));
|
|
529
|
-
}
|
|
530
|
-
});
|
|
531
|
-
return notes;
|
|
570
|
+
return nodes.flatMap(n => [
|
|
571
|
+
...(n.type === 'note' ? [{ id: n.metadata.noteId, text: n.text, type: n.metadata.noteType }] : []),
|
|
572
|
+
...(n.children ? extractNotes(n.children) : [])
|
|
573
|
+
]);
|
|
532
574
|
}
|
|
575
|
+
console.log(extractNotes(ast.content));
|
|
576
|
+
```
|
|
577
|
+
|
|
578
|
+
### Search for a term (TypeScript)
|
|
579
|
+
```ts
|
|
580
|
+
import { OfficeParser } from 'officeparser';
|
|
581
|
+
|
|
582
|
+
async function contains(filePath: string, term: string): Promise<boolean> {
|
|
583
|
+
const ast = await OfficeParser.parseOffice(filePath);
|
|
584
|
+
return (await ast.to('text')).value.includes(term);
|
|
585
|
+
}
|
|
586
|
+
```
|
|
587
|
+
|
|
588
|
+
---
|
|
589
|
+
|
|
590
|
+
## Configuration Reference
|
|
591
|
+
|
|
592
|
+
### OfficeParserConfig
|
|
593
|
+
|
|
594
|
+
Pass as the second argument to `parseOffice(file, config)`.
|
|
595
|
+
|
|
596
|
+
| Option | Type | Default | Description |
|
|
597
|
+
|--------|------|---------|-------------|
|
|
598
|
+
| `newlineDelimiter` | `string` | `'\n'` | Delimiter inserted between lines in text output |
|
|
599
|
+
| `ignoreNotes` | `boolean` | `false` | Ignore speaker notes (PPTX/ODP) |
|
|
600
|
+
| `putNotesAtLast` | `boolean` | `false` | Collect all notes at the end instead of inline |
|
|
601
|
+
| `extractAttachments` | `boolean` | `false` | Populate `ast.attachments` with Base64 images/charts |
|
|
602
|
+
| `ocr` | `boolean` | `false` | Run Tesseract OCR on images (requires `extractAttachments: true`) |
|
|
603
|
+
| `ocrConfig` | `OcrConfig` | `{}` | OCR worker pool settings — see [OCR section](#ocr-scheduler--resource-management) |
|
|
604
|
+
| `includeRawContent` | `boolean` | `false` | Attach raw XML/RTF source to each node |
|
|
605
|
+
| `serializeRawContent` | `boolean` | `true` | Re-serialize XML to clean strings (only if `includeRawContent: true`) |
|
|
606
|
+
| `preserveXmlWhitespace` | `boolean` | `false` | Preserve original XML whitespace during serialization |
|
|
607
|
+
| `includeBreakNodes` | `boolean` | `false` | Include `w:br` / `w:cr` as typed break nodes (DOCX only) |
|
|
608
|
+
| `ignoreInternalLinks` | `boolean` | `false` | Strip bookmarks and internal cross-references from AST |
|
|
609
|
+
| `fileType` | `SupportedFileType \| null` | `null` | **Required for text-based buffers** (`'md'`, `'html'`, `'csv'`) with no magic bytes |
|
|
610
|
+
| `csvDelimiter` | `string` | `','` | Input delimiter when parsing CSV files |
|
|
611
|
+
| `pdfWorkerSrc` | `string` | CDN (jsDelivr) | Path/URL to `pdf.worker.min.mjs` (required in browser) |
|
|
612
|
+
| `onWarning` | `(issue: OfficeIssue) => void` | — | Callback for non-fatal parsing issues |
|
|
613
|
+
| `outputErrorToConsole` | `boolean` | `false` | **Deprecated.** Use `onWarning` instead |
|
|
614
|
+
|
|
615
|
+
---
|
|
616
|
+
|
|
617
|
+
### GeneratorConfig (Common)
|
|
618
|
+
|
|
619
|
+
Options shared by all generator formats. Pass to `OfficeGenerator.generate(ast, format, config)` or `ast.to(format, config)`.
|
|
620
|
+
|
|
621
|
+
| Option | Type | Default | Description |
|
|
622
|
+
|--------|------|---------|-------------|
|
|
623
|
+
| `includeFormatting` | `boolean` | `true` | Include bold/italic/colors/sizes in output |
|
|
624
|
+
| `generateIds` | `boolean` | `true` | Add slug-based `id` attributes to headings |
|
|
625
|
+
| `renderMetadata` | `boolean` | `false` | Render title/author as visible header block |
|
|
626
|
+
| `includeImages` | `boolean` | `true` | Include image nodes in output |
|
|
627
|
+
| `includeCharts` | `boolean` | `true` | Include interactive charts (HTML only) |
|
|
628
|
+
| `ignoreInternalLinks` | `boolean` | `false` | Strip bookmarks and internal anchors from output |
|
|
629
|
+
| `ignoreDefaultStyleMap` | `boolean` | `false` | Disable built-in style mappings (e.g., "Heading 1" → h1) |
|
|
630
|
+
| `styleMap` | `string[] \| StructuredStyleMapping[]` | `[]` | Custom semantic style mappings |
|
|
631
|
+
| `onNode` | `(node) => string \| false \| void` | — | Per-node callback for filtering, overriding, or mutating |
|
|
632
|
+
| `onWarning` | `(issue: OfficeIssue) => void` | — | Callback for non-fatal generation issues |
|
|
633
|
+
|
|
634
|
+
---
|
|
635
|
+
|
|
636
|
+
### `onNode` Callback — Advanced Node Manipulation
|
|
637
|
+
|
|
638
|
+
Called for **every node** in the AST during generation. Can be `async`.
|
|
533
639
|
|
|
534
|
-
|
|
535
|
-
|
|
536
|
-
|
|
537
|
-
|
|
538
|
-
|
|
539
|
-
|
|
540
|
-
|
|
541
|
-
|
|
542
|
-
|------|----------|---------|-------------|
|
|
543
|
-
| `outputErrorToConsole` | boolean | `false` | **Deprecated**: Use `onWarning` instead. Show logs to console in case of an error. |
|
|
544
|
-
| `newlineDelimiter` | string | `\n` | Delimiter for new lines in text output. |
|
|
545
|
-
| `ignoreNotes` | boolean | `false` | Ignore notes in files like PowerPoint/ODP. |
|
|
546
|
-
| `putNotesAtLast` | boolean | `false` | Put notes text at the end of the document. |
|
|
547
|
-
| `extractAttachments` | boolean | `false` | Extract images and charts as Base64. |
|
|
548
|
-
| `includeRawContent` | boolean | `false` | Include raw XML/RTF markup in the nodes. |
|
|
549
|
-
| `serializeRawContent` | boolean | `true` | Re-serializes raw XML to clean strings. |
|
|
550
|
-
| `preserveXmlWhitespace` | boolean | `false` | Preserves original XML whitespace. |
|
|
551
|
-
| `ocr` | boolean | `false` | Enable OCR for images (requires `extractAttachments: true`). |
|
|
552
|
-
| `pdfWorkerSrc` | string | `(see below)` | Path to PDF.js worker. |
|
|
553
|
-
| `ocrConfig` | object | `{}` | OCR Scheduler configuration. |
|
|
554
|
-
| `includeBreakNodes` | boolean | `false` | Include `w:br`, `w:cr` nodes (DOCX only).|
|
|
555
|
-
| `ignoreInternalLinks` | boolean | `false` | Remove all bookmarks and internal jumps. |
|
|
556
|
-
| `csvDelimiter` | string | `,` | Custom delimiter for parsing CSV files. |
|
|
557
|
-
| `fileType` | string | `null` | Manual format override (authoritative). |
|
|
558
|
-
|
|
559
|
-
## Generator Configuration: GeneratorConfig
|
|
560
|
-
Configuration options for `OfficeGenerator.generate`.
|
|
561
|
-
|
|
562
|
-
| Flag | DataType | Default | Explanation |
|
|
563
|
-
|------|----------|---------|-------------|
|
|
564
|
-
| `includeFormatting` | boolean | `true` | Whether to include semantic styles (bold, italic, colors, sizes) in output. |
|
|
565
|
-
| `generateIds` | boolean | `true` | Automatically generates unique slug-based IDs for heading nodes. |
|
|
566
|
-
| `renderMetadata` | boolean | `false` | Renders document metadata (Title, Author) as a visible header block. |
|
|
567
|
-
| `includeImages` | boolean | `true` | Whether to include image nodes in the generated output. |
|
|
568
|
-
| `includeCharts` | boolean | `true` | Whether to include interactive charts (HTML only). |
|
|
569
|
-
| `ignoreInternalLinks`| boolean | `false` | Suppresses all internal bookmarks and anchor references. |
|
|
570
|
-
| `ignoreDefaultStyleMap`| boolean | `false` | Ignore the library's default style mappings. |
|
|
571
|
-
| `styleMap` | string[] \| array | `[]` | Array of style mappings (DSL strings or structured objects). |
|
|
572
|
-
| `onNode` | function | `undefined` | Callback to filter, override, or mutate any node during generation. |
|
|
573
|
-
| `onWarning` | function | `undefined` | Callback for generation-phase warnings. |
|
|
574
|
-
| `htmlConfig` | object | `{}` | Format-specific settings for HTML generation. |
|
|
575
|
-
| `mdConfig` | object | `{}` | Format-specific settings for Markdown generation. |
|
|
576
|
-
| `pdfConfig` | object | `{}` | Format-specific settings for PDF generation. |
|
|
577
|
-
| `csvConfig` | object | `{}` | Format-specific settings for CSV generation. |
|
|
578
|
-
| `textConfig` | object | `{}` | Format-specific settings for Plain Text generation. |
|
|
579
|
-
| `rtfConfig` | object | `{}` | Format-specific settings for RTF generation. |
|
|
580
|
-
| `chunksConfig` | object | `(doc-struct)` | Settings for RAG chunking strategies. |
|
|
581
|
-
|
|
582
|
-
### 🛠️ Advanced Node Manipulation (Pro Users)
|
|
583
|
-
The `onNode` callback is a powerful tool that gives you complete control over the generation process. It is called for **every single node** in the AST before it is rendered.
|
|
584
|
-
|
|
585
|
-
#### Callback Capabilities:
|
|
586
|
-
1. **Filter/Remove Nodes**: Return `false` to skip a node and all its children.
|
|
587
|
-
2. **Override Rendering**: Return a `string` to use that exact text as the output, bypassing default logic and recursion.
|
|
588
|
-
3. **Mutate Nodes**: Modify the `node` object directly (e.g., changing `node.text`) and return `void` to let the generator proceed with your changes.
|
|
589
|
-
4. **Async Support**: The callback can be `async`, allowing you to fetch external data or perform complex logic during generation.
|
|
590
|
-
|
|
591
|
-
#### Pro Example:
|
|
592
|
-
```typescript
|
|
593
|
-
const result = await ast.to('md', {
|
|
640
|
+
| Return value | Effect |
|
|
641
|
+
|---|---|
|
|
642
|
+
| `false` | Skip this node and all its children |
|
|
643
|
+
| `string` | Use this string as the output for this node, skip default logic |
|
|
644
|
+
| `void` | Proceed with default rendering (mutations to `node` are applied) |
|
|
645
|
+
|
|
646
|
+
```ts
|
|
647
|
+
const { value: md } = await ast.to('md', {
|
|
594
648
|
onNode: async (node) => {
|
|
595
|
-
//
|
|
649
|
+
// Skip all images
|
|
596
650
|
if (node.type === 'image') return false;
|
|
597
651
|
|
|
598
|
-
//
|
|
652
|
+
// Redact secrets (mutate then proceed)
|
|
599
653
|
if (node.text?.includes('SECRET_KEY')) {
|
|
600
654
|
node.text = node.text.replace(/SECRET_KEY: \w+/, 'SECRET_KEY: [REDACTED]');
|
|
601
655
|
}
|
|
602
656
|
|
|
603
|
-
//
|
|
657
|
+
// Custom rendering for a specific style
|
|
604
658
|
if (node.metadata?.style === 'Callout') {
|
|
605
659
|
return `> [!INFO]\n> ${node.text}`;
|
|
606
660
|
}
|
|
607
|
-
|
|
608
|
-
// 4. Proceed with default rendering (implicitly returns void)
|
|
609
661
|
}
|
|
610
662
|
});
|
|
611
663
|
```
|
|
612
664
|
|
|
613
|
-
|
|
614
|
-
|
|
665
|
+
---
|
|
666
|
+
|
|
667
|
+
### `styleMap` — Semantic Style Mapping
|
|
668
|
+
|
|
669
|
+
Maps document style names to semantic output elements. Two formats supported:
|
|
615
670
|
|
|
616
|
-
####
|
|
617
|
-
Use structured objects to match nodes based on type and attributes, and specify detailed output properties like classes and custom attributes.
|
|
671
|
+
#### Structured Objects (Recommended)
|
|
618
672
|
|
|
619
|
-
```
|
|
673
|
+
```ts
|
|
620
674
|
styleMap: [
|
|
621
|
-
{
|
|
622
|
-
selector: {
|
|
623
|
-
|
|
624
|
-
attributes: { style: 'Heading 1' }
|
|
625
|
-
},
|
|
626
|
-
output: {
|
|
627
|
-
tag: 'h1',
|
|
628
|
-
classes: ['main-title'],
|
|
629
|
-
attributes: { id: 'top' }
|
|
630
|
-
}
|
|
675
|
+
{
|
|
676
|
+
selector: { nodeType: 'paragraph', attributes: { style: 'Heading 1' } },
|
|
677
|
+
output: { tag: 'h1', classes: ['main-title'], attributes: { id: 'top' } }
|
|
631
678
|
},
|
|
632
679
|
{
|
|
633
|
-
//
|
|
680
|
+
// '~=' operator matches if the word 'Quote' appears anywhere in the style name
|
|
634
681
|
selector: { attributes: { style: { value: 'Quote', operator: '~=' } } },
|
|
635
|
-
output: { tag: 'blockquote' }
|
|
682
|
+
output: { tag: 'blockquote', fresh: true }
|
|
636
683
|
}
|
|
637
684
|
]
|
|
638
685
|
```
|
|
639
686
|
|
|
640
|
-
|
|
641
|
-
The library also maintains support for a simple string-based DSL, highly compatible with `mammoth.js`.
|
|
642
|
-
|
|
643
|
-
- **Literal Matching**: `"p[style-name='Heading 1'] => h1"`
|
|
644
|
-
- **Regex-like Matching**: `"p[style~='Title'] => h2"`
|
|
645
|
-
- **Attribute Filters**: `"p[style-name='Quote'][lang='en'] => blockquote"`
|
|
646
|
-
|
|
647
|
-
## Format-Specific Generator Configuration
|
|
648
|
-
Each destination format has its own specialized sub-configuration object nested within the main `GeneratorConfig`.
|
|
649
|
-
|
|
650
|
-
### 1. HtmlGeneratorConfig (`htmlConfig`)
|
|
651
|
-
| Flag | DataType | Default | Explanation |
|
|
652
|
-
|------|----------|---------|-------------|
|
|
653
|
-
| `standalone` | boolean | `true` | Wraps output in a full `<html>` document with CSS and metadata. |
|
|
654
|
-
| `chartJsSrc` | string | `(CDN)` | URL for the Chart.js library used for interactive charts. |
|
|
655
|
-
|
|
656
|
-
### 2. MdGeneratorConfig (`mdConfig`)
|
|
657
|
-
| Flag | DataType | Default | Explanation |
|
|
658
|
-
|------|----------|---------|-------------|
|
|
659
|
-
| `fallbackToHtml` | boolean | `true` | Uses HTML tags for features not supported by Markdown (underlines, complex tables). |
|
|
660
|
-
|
|
661
|
-
### 3. PdfGeneratorConfig (`pdfConfig`)
|
|
662
|
-
| Flag | DataType | Default | Explanation |
|
|
663
|
-
|------|----------|---------|-------------|
|
|
664
|
-
| `format` | string | `'A4'` | Paper format (e.g., 'Letter', 'A4', 'Legal'). |
|
|
665
|
-
| `landscape` | boolean | `false` | Page orientation. |
|
|
666
|
-
| `margin` | object | `{0,0,0,0}` | Top, right, bottom, left margins. |
|
|
667
|
-
| `displayHeaderFooter`| boolean | `false` | Whether to display print headers and footers. |
|
|
668
|
-
| `headerTemplate` | string | `''` | HTML template for the print header. |
|
|
669
|
-
| `footerTemplate` | string | `''` | HTML template for the print footer. |
|
|
670
|
-
|
|
671
|
-
### 4. CsvGeneratorConfig (`csvConfig`)
|
|
672
|
-
| Flag | DataType | Default | Explanation |
|
|
673
|
-
|------|----------|---------|-------------|
|
|
674
|
-
| `sheets` | string | `''` | Range of sheets to export (e.g., "1", "1-3", "1,3"). |
|
|
675
|
-
| `mergeSheets` | boolean | `true` | Merges all sheets into one CSV string. If false, returns a ZIP. |
|
|
676
|
-
| `columnDelimiter` | string | `','` | Custom delimiter for the CSV output. |
|
|
677
|
-
|
|
678
|
-
### 5. TextGeneratorConfig (`textConfig`)
|
|
679
|
-
| Flag | DataType | Default | Explanation |
|
|
680
|
-
|------|----------|---------|-------------|
|
|
681
|
-
| `newlineDelimiter` | string | `\n` | String inserted between structural blocks. |
|
|
682
|
-
| `preserveLayout` | boolean | `false` | Attempts to maintain table structures using whitespace. |
|
|
683
|
-
|
|
684
|
-
## One-Step Conversion: OfficeConverterConfig
|
|
685
|
-
Configuration for the `OfficeConverter.convert()` API.
|
|
686
|
-
|
|
687
|
-
| Flag | DataType | Default | Explanation |
|
|
688
|
-
|------|----------|---------|-------------|
|
|
689
|
-
| `parseConfig` | object | `{}` | Settings for the `OfficeParser` phase. |
|
|
690
|
-
| `generatorConfig` | object | `{}` | Settings for the `OfficeGenerator` phase. |
|
|
691
|
-
| `onWarning` | function | `undefined` | Global callback for issues in either phase. Overrides phase-specific callbacks. |
|
|
692
|
-
|
|
693
|
-
## Chunking Configuration: ChunkingConfig
|
|
694
|
-
Specific options when using `format: 'chunks'`.
|
|
695
|
-
|
|
696
|
-
| Flag | DataType | Default | Explanation |
|
|
697
|
-
|------|----------|---------|-------------|
|
|
698
|
-
| `strategy` | string | `'fixed-size'`| The chunking strategy (`fixed-size`, `document-structure`, `semantic`). |
|
|
699
|
-
| `maxChunkSize` | number | `1000` | Maximum characters per chunk. |
|
|
700
|
-
| `chunkOverlap` | number | `200` | Overlap between consecutive chunks. |
|
|
701
|
-
| `similarityThreshold`| number | `0.5` | Threshold for semantic splitting (0.0 to 1.0). |
|
|
702
|
-
| `embedBatchSize` | number | `50` | Batch size for embedding requests. |
|
|
703
|
-
|
|
704
|
-
### OCR Scheduler & Resource Management
|
|
705
|
-
If your application uses OCR, `officeParser` utilizes an intelligent **Smart Worker Pool** to maintain a background worker pool and optimize repeated parse requests.
|
|
706
|
-
|
|
707
|
-
- **Dynamic Affinity**: Workers in the pool persist with their last used language affinity.
|
|
708
|
-
- **LRU Re-allocation**: If a new language is requested and the pool is full, the manager identifies the **Least Recently Used (LRU)** idle worker and re-initializes it for the new language. This avoids the overhead of destroying and recreating workers.
|
|
709
|
-
- **Auto-Termination**: Workers are automatically cleaned up after 10 seconds of inactivity (configurable via `ocrConfig.autoTerminateTimeout`).
|
|
710
|
-
|
|
711
|
-
#### `OfficeParser.terminateOcr()`
|
|
712
|
-
If you have used OCR (`{ ocr: true }`) in a short-lived script (like CLI tools or one-off automation), we recommend explicitly calling `terminateOcr()` after your processing is finished. This bypasses the 10-second idle timer and allows the process to return to the terminal prompt immediately.
|
|
687
|
+
`fresh: true` prevents the generator from merging adjacent nodes of the same tag into one block.
|
|
713
688
|
|
|
714
|
-
|
|
715
|
-
|
|
689
|
+
#### Legacy String DSL
|
|
690
|
+
|
|
691
|
+
Compatible with `mammoth.js` style maps:
|
|
716
692
|
|
|
717
693
|
```js
|
|
718
|
-
|
|
694
|
+
styleMap: [
|
|
695
|
+
"p[style-name='Heading 1'] => h1",
|
|
696
|
+
"p[style~='Title'] => h2",
|
|
697
|
+
"p[style-name='Quote'][lang='en'] => blockquote"
|
|
698
|
+
]
|
|
699
|
+
```
|
|
700
|
+
|
|
701
|
+
---
|
|
719
702
|
|
|
720
|
-
|
|
721
|
-
await officeParser.parseOffice("file.pdf", { ocr: true });
|
|
722
|
-
// ... process results ...
|
|
703
|
+
### HtmlGeneratorConfig
|
|
723
704
|
|
|
724
|
-
|
|
725
|
-
await officeParser.terminateOcr();
|
|
726
|
-
}
|
|
727
|
-
```
|
|
705
|
+
Pass as `htmlConfig` inside `GeneratorConfig`.
|
|
728
706
|
|
|
729
|
-
|
|
730
|
-
|
|
707
|
+
| Option | Type | Default | Description |
|
|
708
|
+
|--------|------|---------|-------------|
|
|
709
|
+
| `standalone` | `boolean` | `true` | Wrap output in a full `<html>` document with CSS |
|
|
710
|
+
| `chartJsSrc` | `string` | jsDelivr CDN | URL for the Chart.js library |
|
|
731
711
|
|
|
732
|
-
|
|
733
|
-
const config = {
|
|
734
|
-
newlineDelimiter: "\n\n",
|
|
735
|
-
extractAttachments: true,
|
|
736
|
-
ocr: true,
|
|
737
|
-
ocrLanguage: 'eng+fra+esp' // Supports English, French, and Spanish simultaneously
|
|
738
|
-
};
|
|
712
|
+
### MdGeneratorConfig
|
|
739
713
|
|
|
740
|
-
|
|
741
|
-
console.log(`Extracted ${ast.attachments.length} images`);
|
|
742
|
-
```
|
|
714
|
+
Pass as `mdConfig` inside `GeneratorConfig`.
|
|
743
715
|
|
|
744
|
-
|
|
716
|
+
| Option | Type | Default | Description |
|
|
717
|
+
|--------|------|---------|-------------|
|
|
718
|
+
| `fallbackToHtml` | `boolean` | `true` | Use HTML tags for features Markdown cannot represent (underlines, merged table cells, etc.) |
|
|
745
719
|
|
|
746
|
-
|
|
747
|
-
```ts
|
|
748
|
-
import { OfficeParser } from 'officeparser';
|
|
720
|
+
### PdfGeneratorConfig
|
|
749
721
|
|
|
750
|
-
|
|
751
|
-
|
|
752
|
-
|
|
753
|
-
|
|
754
|
-
|
|
722
|
+
Pass as `pdfConfig` inside `GeneratorConfig`. Requires the optional `puppeteer` peer dependency.
|
|
723
|
+
|
|
724
|
+
| Option | Type | Default | Description |
|
|
725
|
+
|--------|------|---------|-------------|
|
|
726
|
+
| `format` | `string` | `'A4'` | Paper format (`'A4'`, `'Letter'`, `'Legal'`, etc.) |
|
|
727
|
+
| `landscape` | `boolean` | `false` | Landscape page orientation |
|
|
728
|
+
| `printBackground` | `boolean` | `true` | Print background graphics |
|
|
729
|
+
| `margin` | `object` | `{0,0,0,0}` | Page margins (`top`, `right`, `bottom`, `left`) |
|
|
730
|
+
| `displayHeaderFooter` | `boolean` | `false` | Show print header/footer |
|
|
731
|
+
| `headerTemplate` | `string` | `''` | HTML template for the print header |
|
|
732
|
+
| `footerTemplate` | `string` | `''` | HTML template for the print footer |
|
|
733
|
+
| `scale` | `number` | `1` | Rendering scale factor |
|
|
734
|
+
| `launchOptions` | `object` | headless defaults | Puppeteer launch options (e.g., `executablePath`) |
|
|
735
|
+
|
|
736
|
+
### CsvGeneratorConfig
|
|
737
|
+
|
|
738
|
+
Pass as `csvConfig` inside `GeneratorConfig`.
|
|
739
|
+
|
|
740
|
+
| Option | Type | Default | Description |
|
|
741
|
+
|--------|------|---------|-------------|
|
|
742
|
+
| `sheets` | `string` | `''` | Sheet range to export: `'1'`, `'1-3'`, `'1,3'` (1-based). Empty = all sheets |
|
|
743
|
+
| `mergeSheets` | `boolean` | `true` | Merge all sheets into one CSV. If `false`, returns a ZIP archive |
|
|
744
|
+
| `columnDelimiter` | `string` | `','` | Output column delimiter |
|
|
745
|
+
|
|
746
|
+
### TextGeneratorConfig
|
|
747
|
+
|
|
748
|
+
Pass as `textConfig` inside `GeneratorConfig`.
|
|
749
|
+
|
|
750
|
+
| Option | Type | Default | Description |
|
|
751
|
+
|--------|------|---------|-------------|
|
|
752
|
+
| `newlineDelimiter` | `string` | `'\n'` | String inserted between structural blocks |
|
|
753
|
+
| `preserveLayout` | `boolean` | `false` | Render tables with aligned columns using whitespace |
|
|
754
|
+
|
|
755
|
+
---
|
|
756
|
+
|
|
757
|
+
### OfficeConverterConfig
|
|
758
|
+
|
|
759
|
+
Configuration for `OfficeConverter.convert(file, format, config)`.
|
|
760
|
+
|
|
761
|
+
| Option | Type | Description |
|
|
762
|
+
|--------|------|-------------|
|
|
763
|
+
| `parseConfig` | `OfficeParserConfig` | Settings for the parsing phase |
|
|
764
|
+
| `generatorConfig` | `GeneratorConfig` | Settings for the generation phase |
|
|
765
|
+
| `onWarning` | `(issue: OfficeIssue) => void` | Global warning callback (overrides phase-specific ones) |
|
|
766
|
+
|
|
767
|
+
---
|
|
768
|
+
|
|
769
|
+
### ChunkingConfig
|
|
770
|
+
|
|
771
|
+
`ChunkingConfig` is a **discriminated union** — the available options depend on the `strategy` field.
|
|
772
|
+
|
|
773
|
+
#### Common Options (all strategies)
|
|
774
|
+
|
|
775
|
+
| Option | Type | Default | Description |
|
|
776
|
+
|--------|------|---------|-------------|
|
|
777
|
+
| `strategy` | `string` | `'document-structure'` | Chunking strategy |
|
|
778
|
+
| `stripWhitespace` | `boolean` | `true` | Trim leading/trailing whitespace from each chunk |
|
|
779
|
+
| `includeMetadata` | `boolean` | `true` | Include page/slide/heading metadata in each chunk |
|
|
780
|
+
| `addStartIndex` | `boolean` | `false` | Add `startIndex` character offset to chunk metadata |
|
|
781
|
+
| `lengthFunction` | `(text) => number` | `text.length` | Custom size measurer (e.g., token counter) |
|
|
782
|
+
| `sentenceBoundaryRegex` | `string \| RegExp` | `/[.!?。!?]/` | Custom regex for sentence boundary detection |
|
|
783
|
+
| `abbreviations` | `string[]` | common list | Abbreviations to skip when splitting on `.` |
|
|
784
|
+
|
|
785
|
+
#### `strategy: 'fixed-size'`
|
|
786
|
+
|
|
787
|
+
| Option | Type | Default | Description |
|
|
788
|
+
|--------|------|---------|-------------|
|
|
789
|
+
| `chunkSize` | `number` | `1000` | Maximum characters per chunk |
|
|
790
|
+
| `chunkOverlap` | `number` | `200` | Character overlap between consecutive chunks |
|
|
791
|
+
| `separators` | `string[]` | `['\n\n','\n',' ','']` | Ordered list of separators to try |
|
|
792
|
+
|
|
793
|
+
#### `strategy: 'document-structure'`
|
|
794
|
+
|
|
795
|
+
| Option | Type | Default | Description |
|
|
796
|
+
|--------|------|---------|-------------|
|
|
797
|
+
| `splitBy` | `string` | `'paragraph'` | `'paragraph'` · `'heading'` · `'page'` · `'slide'` · `'sheet'` |
|
|
798
|
+
| `maxChunkSize` | `number` | `1000` | Max characters per chunk (oversized units are split recursively) |
|
|
799
|
+
| `tableSplitStrategy` | `string` | `'row'` | `'row'` (repeats header in each chunk) or `'flatten'` |
|
|
800
|
+
|
|
801
|
+
#### `strategy: 'semantic'`
|
|
802
|
+
|
|
803
|
+
| Option | Type | Default | Description |
|
|
804
|
+
|--------|------|---------|-------------|
|
|
805
|
+
| `embeddingFunction` | `(text) => Promise<number[]>` | **required** | Async embedding function |
|
|
806
|
+
| `similarityThreshold` | `number` | `0.8` | Cosine similarity threshold; lower = fewer boundaries |
|
|
807
|
+
| `maxChunkSize` | `number` | `2000` | Max characters even if similarity stays high |
|
|
808
|
+
| `bufferSize` | `number` | `1` | Surrounding sentences used when computing similarity |
|
|
809
|
+
| `embeddingBatchSize` | `number` | `50` | Sentences per embedding API batch |
|
|
810
|
+
|
|
811
|
+
---
|
|
812
|
+
|
|
813
|
+
## OCR Scheduler & Resource Management
|
|
814
|
+
|
|
815
|
+
When `ocr: true` is set, `officeParser` maintains an intelligent **Smart Worker Pool** backed by Tesseract.js:
|
|
816
|
+
|
|
817
|
+
- **Dynamic Affinity**: Workers persist with their last-used language, avoiding re-initialization overhead.
|
|
818
|
+
- **LRU Re-allocation**: When a new language is requested and the pool is full, the Least Recently Used idle worker is re-initialized.
|
|
819
|
+
- **Auto-Termination**: Workers shut down after 10 seconds of inactivity (configurable via `ocrConfig.autoTerminateTimeout`).
|
|
820
|
+
|
|
821
|
+
### OCR Config (`ocrConfig`)
|
|
822
|
+
|
|
823
|
+
| Option | Type | Default | Description |
|
|
824
|
+
|--------|------|---------|-------------|
|
|
825
|
+
| `language` | `string` | `'eng'` | Tesseract language code(s), e.g. `'eng+fra'` |
|
|
826
|
+
| `workerPath` | `string` | `''` | Custom path to Tesseract worker script |
|
|
827
|
+
| `corePath` | `string` | `''` | Custom path to Tesseract core script |
|
|
828
|
+
| `langPath` | `string` | `''` | Custom path for language data files |
|
|
829
|
+
| `autoTerminateTimeout` | `number` | `10000` | Inactivity timeout in ms before auto-teardown (0 = disabled) |
|
|
830
|
+
|
|
831
|
+
See all language codes at [tesseract-ocr.github.io](https://tesseract-ocr.github.io/tessdoc/Data-Files).
|
|
832
|
+
|
|
833
|
+
### `OfficeParser.terminateOcr()`
|
|
834
|
+
|
|
835
|
+
In **short-lived scripts** (CLI tools, one-off automation), call `terminateOcr()` after processing to bypass the idle timer and exit immediately:
|
|
755
836
|
|
|
756
|
-
**Extracting Images and their OCR text**
|
|
757
837
|
```js
|
|
758
838
|
const officeParser = require('officeparser');
|
|
759
839
|
|
|
760
|
-
const
|
|
761
|
-
|
|
762
|
-
|
|
763
|
-
if (attachment.type === 'image') {
|
|
764
|
-
console.log(`Image: ${attachment.name}`);
|
|
765
|
-
console.log(`OCR Text: ${attachment.ocrText}`);
|
|
766
|
-
fs.writeFileSync(attachment.name, Buffer.from(attachment.data, 'base64'));
|
|
767
|
-
}
|
|
768
|
-
});
|
|
769
|
-
});
|
|
840
|
+
const ast = await officeParser.parseOffice('file.pdf', { ocr: true });
|
|
841
|
+
// ... process results ...
|
|
842
|
+
await officeParser.terminateOcr(); // immediate exit
|
|
770
843
|
```
|
|
771
844
|
|
|
845
|
+
> [!TIP]
|
|
846
|
+
> The built-in CLI (`npx officeparser ...`) handles this automatically.
|
|
847
|
+
> Only call it manually in your own scripts.
|
|
848
|
+
|
|
849
|
+
---
|
|
850
|
+
|
|
772
851
|
## Browser Usage
|
|
773
|
-
The library provides two types of browser bundles in the `dist/` directory:
|
|
774
|
-
1. **`officeparser.browser.iife.js`**: Standard IIFE bundle for direct `<script>` tag usage. Exposes the global `officeParser` namespace.
|
|
775
|
-
2. **`officeparser.browser.mjs`**: Modern ESM bundle for use with `import` statements or modern bundlers.
|
|
776
852
|
|
|
777
|
-
|
|
778
|
-
|
|
853
|
+
Two bundles are available in the `dist/` directory:
|
|
854
|
+
|
|
855
|
+
| Bundle | Usage |
|
|
856
|
+
|--------|-------|
|
|
857
|
+
| `officeparser.browser.mjs` | ESM — use with `import` statements or modern bundlers (Vite, Webpack, Next.js) |
|
|
858
|
+
| `officeparser.browser.iife.js` | IIFE — use with a `<script>` tag; exposes the global `officeParser` object |
|
|
779
859
|
|
|
780
|
-
|
|
860
|
+
### ESM (Vite / Webpack / Next.js)
|
|
861
|
+
|
|
862
|
+
```js
|
|
781
863
|
import { OfficeParser } from 'officeparser';
|
|
782
864
|
|
|
783
865
|
const handleFile = async (event) => {
|
|
784
866
|
const file = event.target.files[0];
|
|
785
867
|
const buffer = await file.arrayBuffer();
|
|
786
|
-
|
|
787
|
-
|
|
788
|
-
// Pass the Buffer or Uint8Array directly
|
|
789
|
-
const ast = await OfficeParser.parseOffice(new Uint8Array(buffer));
|
|
790
|
-
console.log(ast.toText());
|
|
791
|
-
} catch (err) {
|
|
792
|
-
console.error(err);
|
|
793
|
-
}
|
|
868
|
+
const ast = await OfficeParser.parseOffice(new Uint8Array(buffer));
|
|
869
|
+
console.log(ast.toText());
|
|
794
870
|
};
|
|
795
871
|
```
|
|
796
872
|
|
|
797
|
-
|
|
798
|
-
> **Why `fs` fails in the browser**: Browsers do not have a built-in file system. If you try to pass a file path string in the browser, `officeParser` will throw a descriptive "Fail-Fast" error instead of crashing mysteriously:
|
|
799
|
-
> `[officeparser] Node.js 'fs' module is not available in the browser. Please pass a Buffer or Uint8Array instead.`
|
|
800
|
-
|
|
801
|
-
### Usage (Script Tag)
|
|
802
|
-
Include the IIFE bundle available in the release assets or your `dist/` folder. This exposes the global `officeParser` object.
|
|
873
|
+
### Script Tag
|
|
803
874
|
|
|
804
875
|
```html
|
|
805
876
|
<script src="dist/officeparser.browser.iife.js"></script>
|
|
@@ -807,60 +878,78 @@ Include the IIFE bundle available in the release assets or your `dist/` folder.
|
|
|
807
878
|
async function handleFile(event) {
|
|
808
879
|
const file = event.target.files[0];
|
|
809
880
|
const buffer = await file.arrayBuffer();
|
|
810
|
-
|
|
811
|
-
|
|
812
|
-
// Reconstruct as Uint8Array for the parser
|
|
813
|
-
const ast = await officeParser.parseOffice(new Uint8Array(buffer));
|
|
814
|
-
console.log(ast.toText());
|
|
815
|
-
} catch (error) {
|
|
816
|
-
console.error("Parsing failed:", error);
|
|
817
|
-
}
|
|
881
|
+
const ast = await officeParser.parseOffice(new Uint8Array(buffer));
|
|
882
|
+
console.log(ast.toText());
|
|
818
883
|
}
|
|
819
884
|
</script>
|
|
820
885
|
```
|
|
821
886
|
|
|
822
|
-
|
|
823
|
-
|
|
887
|
+
> [!NOTE]
|
|
888
|
+
> **File paths don't work in the browser.** Always pass a `Buffer`, `ArrayBuffer`, or `Uint8Array`.
|
|
889
|
+
> Passing a path string will throw a descriptive `FEATURE_NOT_SUPPORTED_IN_BROWSER` error.
|
|
890
|
+
|
|
891
|
+
### PDF Worker Configuration
|
|
824
892
|
|
|
825
|
-
|
|
826
|
-
const file = ...; // File object or ArrayBuffer
|
|
893
|
+
When parsing PDFs in the browser, a Web Worker is required. If `pdfWorkerSrc` is omitted, a jsDelivr CDN link is used automatically:
|
|
827
894
|
|
|
828
|
-
|
|
829
|
-
|
|
895
|
+
```js
|
|
896
|
+
// Uses default CDN worker:
|
|
897
|
+
const ast = await officeParser.parseOffice(pdfArrayBuffer);
|
|
830
898
|
|
|
831
|
-
// Or
|
|
832
|
-
const
|
|
833
|
-
pdfWorkerSrc:
|
|
899
|
+
// Or specify your own:
|
|
900
|
+
const ast = await officeParser.parseOffice(pdfArrayBuffer, {
|
|
901
|
+
pdfWorkerSrc: 'https://cdn.jsdelivr.net/npm/pdfjs-dist@5.6.205/build/pdf.worker.min.mjs'
|
|
834
902
|
});
|
|
835
903
|
```
|
|
836
904
|
|
|
837
|
-
>
|
|
905
|
+
> [!NOTE]
|
|
906
|
+
> The `pdfjs-dist` worker version must match the version bundled with `officeparser` (currently **5.6.205**).
|
|
907
|
+
|
|
908
|
+
---
|
|
838
909
|
|
|
839
910
|
## Troubleshooting & Common Issues
|
|
840
911
|
|
|
841
|
-
|
|
842
|
-
|
|
843
|
-
|
|
844
|
-
|
|
912
|
+
| Symptom | Fix |
|
|
913
|
+
|---------|-----|
|
|
914
|
+
| Node.js process stays alive after finishing | Call `await officeParser.terminateOcr()` at end of script when OCR was used |
|
|
915
|
+
| `"Worker not found"` in browser for PDF | Verify `pdfWorkerSrc` points to `pdf.worker.min.mjs` matching version `5.6.205` |
|
|
916
|
+
| Low OCR accuracy | Verify `ocrConfig.language` matches the document language; quality depends on image resolution |
|
|
917
|
+
| Out of memory on large Excel files | Call `ast.toText()` early and discard the AST object to allow garbage collection |
|
|
918
|
+
| `md`/`html`/`csv` buffer not detected | Add `fileType: 'md'` (or `'html'`, `'csv'`) to config — these formats have no magic bytes |
|
|
919
|
+
| `IMPROPER_BUFFERS` error | Usually means no file extension and no `fileType` hint was provided for a buffer input |
|
|
920
|
+
| PDF generation fails | Install the optional peer dependency: `npm install puppeteer` |
|
|
845
921
|
|
|
846
|
-
For a
|
|
922
|
+
For a full debugging guide, visit the [Live Documentation](https://harshankur.github.io/officeParser/#spec/debugging).
|
|
847
923
|
|
|
924
|
+
---
|
|
848
925
|
|
|
849
926
|
## Known Limitations
|
|
850
|
-
1. **ODT/ODS Charts**: Extraction may occasionally show inaccurate data when referencing external cell ranges or complex layout-based data.
|
|
851
|
-
2. **PDF Images**: PDF images are extracted as BMP files in the browser for compatibility. This conversion happens automatically.
|
|
852
|
-
3. **RTF Footnotes**: The `putNotesAtLast` configuration is currently not supported for RTF files; footnotes and endnotes are always collected and appended to the end of the content.
|
|
853
927
|
|
|
854
|
-
|
|
928
|
+
1. **ODT/ODS Charts**: May show inaccurate data when the chart references external cell ranges or uses complex layout-based data.
|
|
929
|
+
2. **PDF Images (Browser)**: Extracted as BMP files for cross-platform compatibility. Conversion is automatic.
|
|
930
|
+
3. **RTF Notes**: `putNotesAtLast` has no effect for RTF files; footnotes and endnotes are always appended at the end.
|
|
931
|
+
|
|
932
|
+
---
|
|
855
933
|
|
|
856
934
|
**npm**: [https://npmjs.com/package/officeparser](https://npmjs.com/package/officeparser)
|
|
857
935
|
|
|
858
936
|
**github**: [https://github.com/harshankur/officeParser](https://github.com/harshankur/officeParser)
|
|
859
937
|
|
|
938
|
+
## Support the Project
|
|
939
|
+
|
|
940
|
+
If `officeParser` has helped you save time, consider supporting its continued development. Your sponsorship helps maintain the project, add new features, and keep it robust for everyone.
|
|
941
|
+
|
|
942
|
+
<a href="https://github.com/sponsors/harshankur">
|
|
943
|
+
<img src="https://img.shields.io/badge/Sponsor-GitHub-ea4aaa?style=for-the-badge&logo=github-sponsors" height="36">
|
|
944
|
+
</a>
|
|
945
|
+
<a href="https://www.buymeacoffee.com/harshankur">
|
|
946
|
+
<img src="https://cdn.buymeacoffee.com/buttons/v2/default-yellow.png" height="36" alt="Buy Me A Coffee">
|
|
947
|
+
</a>
|
|
948
|
+
|
|
860
949
|
## Contributing
|
|
861
950
|
|
|
862
|
-
Contributions are welcome! Please see [CONTRIBUTING.md](CONTRIBUTING.md) for details
|
|
951
|
+
Contributions are welcome! Please see [CONTRIBUTING.md](CONTRIBUTING.md) for details.
|
|
863
952
|
|
|
864
953
|
## License
|
|
865
954
|
|
|
866
|
-
This project is licensed under the MIT License
|
|
955
|
+
This project is licensed under the MIT License — see the [LICENSE](LICENSE) file for details.
|