officeparser 5.2.2 → 6.0.1

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -1,19 +1,26 @@
1
- # officeParser
2
- A Node.js library to parse text out of any office file.
1
+ # officeParser 📄🚀
3
2
 
4
- ### Supported File Types
3
+ A robust, strictly-typed Node.js and Browser library for parsing office files ([`docx`](https://en.wikipedia.org/wiki/Office_Open_XML), [`pptx`](https://en.wikipedia.org/wiki/Office_Open_XML), [`xlsx`](https://en.wikipedia.org/wiki/Office_Open_XML), [`odt`](https://en.wikipedia.org/wiki/OpenDocument), [`odp`](https://en.wikipedia.org/wiki/OpenDocument), [`ods`](https://en.wikipedia.org/wiki/OpenDocument), [`pdf`](https://en.wikipedia.org/wiki/PDF), [`rtf`](https://en.wikipedia.org/wiki/Rich_Text_Format)). It produces a clean, hierarchical Abstract Syntax Tree (AST) with rich metadata, text formatting, and full attachment support.
5
4
 
6
- - [`docx`](https://en.wikipedia.org/wiki/Office_Open_XML)
7
- - [`pptx`](https://en.wikipedia.org/wiki/Office_Open_XML)
8
- - [`xlsx`](https://en.wikipedia.org/wiki/Office_Open_XML)
9
- - [`odt`](https://en.wikipedia.org/wiki/OpenDocument)
10
- - [`odp`](https://en.wikipedia.org/wiki/OpenDocument)
11
- - [`ods`](https://en.wikipedia.org/wiki/OpenDocument)
12
- - [`pdf`](https://en.wikipedia.org/wiki/PDF)
5
+ ---
6
+
7
+ ### 🌟 [Live Interactive AST Visualizer](https://harshankur.github.io/officeParser/) 🌟
8
+ *Test any office file in your browser and see the extracted AST, text, and preview in real-time which is rebuilt from the AST!*
9
+
10
+ ---
13
11
 
14
12
 
15
13
  #### Update
16
- * 2024/11/12 - Added ArrayBuffer as a type of file input. Generating bundle files now which exposes namespace officeParser to be able to access parseOffice and parseOfficeAsync directly on the browser. Extracting text out of pdf files does not work currently in browser bundles.
14
+ * 2025/12/29 - **v6.0.0 Release**: Major overhaul of the library. Transitioned from simple text extraction to a rich **Abstract Syntax Tree (AST)** output.
15
+ - Simplified API: Use `parseOffice` for all parsing needs (returns a Promise).
16
+ - Structured Output: Access hierarchical document structure (paragraphs, headings, tables, lists, etc.).
17
+ - Rich Metadata: Extracted document properties (author, title, creation date).
18
+ - Enhanced Formatting: Support for bold, italic, colors, fonts, alignment, etc.
19
+ - Attachment Handling: Extract images, charts, and embedded files as Base64.
20
+ - OCR Integration: Optional OCR for images using Tesseract.js.
21
+ - RTF Support: Added full support for Rich Text Format files.
22
+ - Improved Type Definitions: Full TypeScript support with detailed interfaces.
23
+ * 2024/11/12 - Added ArrayBuffer as a type of file input. Generating bundle files now which exposes namespace officeParser to be able to access parseOffice directly on the browser.
17
24
  * 2024/10/21 - Replaced extracting zip files from decompress to yauzl. This means that we now extract files in memory and we no longer need to write them to disk. Removed config flags related to extracted files. Added flags for CLI execution.
18
25
  * 2024/10/15 - Fixed erroring out while deleting temp files when multiple worker threads make parallel executions resulting in same file name for multiple files. Fixed erroring out when multiple executions are made without waiting for the previous execution to finish which resulted in deleting the file from other execution. Upgraded dependencies.
19
26
  * 2024/10/13 - Fixed parsing text from xlsx files which contain no shared strings file and files which have inlineStr based strings.
@@ -35,206 +42,447 @@ A Node.js library to parse text out of any office file.
35
42
  * 2019/04/19 - Support added for *.xlsx files.
36
43
  * 2019/04/18 - Support added for *.pptx files.
37
44
 
38
-
39
-
40
45
  ## Install via npm
41
46
 
42
- ```
47
+ ```bash
43
48
  npm i officeparser
44
49
  ```
45
50
 
46
51
  ## Command Line usage
47
- If you want to call the installed officeParser.js file, use below command
48
- ```
49
- node <path/to/officeParser.js> [--configOption=value] [FILE_PATH]
50
- node officeparser [--configOption=value] [FILE_PATH]
51
- ```
52
+ You can use `officeparser` directly from the terminal to get either the full AST (as JSON) or plain text.
52
53
 
53
- Otherwise, you can simply use npx without installing the node module to instantly extract parsed data.
54
- ```
55
- npx officeparser [--configOption=value] [FILE_PATH]
54
+ ```bash
55
+ # Get full AST as JSON (default)
56
+ npx officeparser /path/to/officeFile.docx
57
+
58
+ # Get plain text only
59
+ npx officeparser /path/to/officeFile.docx --toText=true
60
+
61
+ # Use configuration options
62
+ npx officeparser /path/to/officeFile.docx --ignoreNotes=true --newlineDelimiter=" "
56
63
  ```
57
64
 
58
65
  ### Config Options:
66
+ - `--toText=[true|false]` Flag to output only plain text instead of JSON AST.
59
67
  - `--ignoreNotes=[true|false]` Flag to ignore notes from files like PowerPoint. Default is false.
60
68
  - `--newlineDelimiter=[delimiter]` The delimiter to use for new lines. Default is `\n`.
61
69
  - `--putNotesAtLast=[true|false]` Flag to collect notes at the end of files like PowerPoint. Default is false.
62
70
  - `--outputErrorToConsole=[true|false]` Flag to output errors to the console. Default is false.
71
+ - `--extractAttachments=[true|false]` Flag to extract images/charts as Base64. Default is false.
72
+ - `--ocr=[true|false]` Flag to enable OCR for extracted images. Default is false.
73
+ - `--includeRawContent=[true|false]` Flag to include raw XML/RTF content in nodes. Default is false.
74
+
63
75
 
64
76
  ## Library Usage
77
+ In **v6.0.0**, the library has moved to a structured AST output. While this is a change for those expecting a string directly, it provides significantly more power and flexibility.
78
+
79
+ ### Getting Started (Async/Await)
65
80
  ```js
66
81
  const officeParser = require('officeparser');
67
82
 
68
- // callback
69
- officeParser.parseOffice("/path/to/officeFile", function(data, err) {
70
- // "data" string in the callback here is the text parsed from the office file passed in the first argument above
71
- if (err) {
72
- console.log(err);
73
- return;
83
+ async function parseMyFile() {
84
+ try {
85
+ // parseOffice returns an OfficeParserAST object
86
+ const ast = await officeParser.parseOffice("/path/to/officeFile.docx");
87
+
88
+ // Use the built-in helper to get plain text (similar to old behavior)
89
+ const text = ast.toText();
90
+ console.log(text);
91
+
92
+ // Access structured content
93
+ console.log(ast.content); // Array of hierarchical nodes (paragraphs, tables, etc.)
94
+ console.log(ast.metadata); // Document properties (author, title, etc.)
95
+ } catch (err) {
96
+ console.error(err);
74
97
  }
75
- console.log(data);
76
- })
77
-
78
- // promise
79
- officeParser.parseOfficeAsync("/path/to/officeFile");
80
- // "data" string in the promise here is the text parsed from the office file passed in the argument above
81
- .then(data => console.log(data))
82
- .catch(err => console.error(err))
83
-
84
- // async/await
85
- try {
86
- // "data" string returned from promise here is the text parsed from the office file passed in the argument
87
- const data = await officeParser.parseOfficeAsync("/path/to/officeFile");
88
- console.log(data);
89
- } catch (err) {
90
- // resolve error
91
- console.log(err);
92
98
  }
99
+ ```
93
100
 
94
- // USING FILE BUFFERS
95
- // instead of file path, you can also pass file buffers of one of the supported files
96
- // on parseOffice or parseOfficeAsync functions.
97
-
98
- // get file buffers
99
- const fileBuffers = fs.readFileSync("/path/to/officeFile");
100
- // get parsed text from officeParser
101
- // NOTE: Only works with parseOffice. Old functions are not supported.
102
- officeParser.parseOfficeAsync(fileBuffers);
103
- .then(data => console.log(data))
104
- .catch(err => console.error(err))
105
- ```
106
-
107
- ### Configuration Object: OfficeParserConfig
108
- *Optionally add a config object as 3rd variable to parseOffice for the following configurations*
109
- | Flag | DataType | Default | Explanation |
110
- |----------------------|----------|------------------|-------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|
111
- | outputErrorToConsole | boolean | false | Flag to show all the logs to console in case of an error. Default is false. |
112
- | newlineDelimiter | string | \n | The delimiter used for every new line in places that allow multiline text like word. Default is \n. |
113
- | ignoreNotes | boolean | false | Flag to ignore notes from parsing in files like powerpoint. Default is false. It includes notes in the parsed text by default. |
114
- | putNotesAtLast | boolean | false | Flag, if set to true, will collectively put all the parsed text from notes at last in files like powerpoint. Default is false. It puts each notes right after its main slide content. If ignoreNotes is set to true, this flag is also ignored. |
115
- <br>
101
+ ### Helper Function for Text Extraction (Modern simple way)
102
+ If you only need the text and want to maintain a simple one-liner, you can use this pattern:
103
+ ```js
104
+ // Simple helper to get text directly
105
+ const getText = async (file, config) => (await officeParser.parseOffice(file, config)).toText();
106
+
107
+ // usage
108
+ const text = await getText("/path/to/officeFile.docx");
109
+ console.log(text);
110
+ ```
116
111
 
112
+ ### Using Callbacks (Backward Compatibility Support)
113
+ We still support callbacks, but the data returned is now the AST object.
117
114
  ```js
118
- const config = {
119
- newlineDelimiter: " ", // Separate new lines with a space instead of the default \n.
120
- ignoreNotes: true // Ignore notes while parsing presentation files like pptx or odp.
121
- }
115
+ const officeParser = require('officeparser');
122
116
 
123
- // callback
124
- officeParser.parseOffice("/path/to/officeFile", function(data, err){
117
+ officeParser.parseOffice("/path/to/officeFile.docx", function(ast, err) {
125
118
  if (err) {
126
- console.log(err);
119
+ console.error(err);
127
120
  return;
128
121
  }
129
- console.log(data);
130
- }, config)
131
-
132
- // promise
133
- officeParser.parseOfficeAsync("/path/to/officeFile", config);
134
- .then(data => console.log(data))
135
- .catch(err => console.error(err))
122
+ // Get text from AST
123
+ console.log(ast.toText());
124
+ });
136
125
  ```
137
126
 
138
- **Example - JavaScript**
127
+ ### Using File Buffers or ArrayBuffers
128
+ You can pass a file path string, a Node.js `Buffer`, or an `ArrayBuffer`.
139
129
  ```js
130
+ const fs = require('fs');
140
131
  const officeParser = require('officeparser');
132
+ const buffer = fs.readFileSync("/path/to/officeFile.pdf");
141
133
 
142
- const config = {
143
- newlineDelimiter: " ", // Separate new lines with a space instead of the default \n.
144
- ignoreNotes: true // Ignore notes while parsing presentation files like pptx or odp.
134
+ officeParser.parseOffice(buffer)
135
+ .then(ast => console.log(ast.toText()))
136
+ .catch(console.error);
137
+ ```
138
+
139
+ ## The AST Structure
140
+ The `OfficeParserAST` provides a format-agnostic representation of your document, allowing you to traverse and manipulate content as a tree.
141
+
142
+ ### Visualizing the AST
143
+ The `OfficeParserAST` provides a format-agnostic representation of your document. Below is a simplified visualization of how the tree is structured:
144
+
145
+ ```text
146
+ OfficeParserAST
147
+ ├── type: "docx" | "pptx" | "xlsx" | ...
148
+ ├── metadata: { author, title, created, modified, ... }
149
+ ├── content: [ OfficeContentNode ]
150
+ │ ├── type: "paragraph" | "heading" | "table" | "list" | ...
151
+ │ ├── text: "Concatenated text of this node and all children"
152
+ │ ├── children: [ OfficeContentNode ] (recursive)
153
+ │ ├── formatting: { bold, italic, color, size, font, ... }
154
+ │ ├── metadata: { level, listId, row, col, ... }
155
+ │ └── rawContent: "<xml>...</xml>" (if enabled)
156
+ ├── attachments: [ OfficeAttachment ]
157
+ │ ├── type: "image" | "chart"
158
+ │ ├── name: "image1.png"
159
+ │ ├── data: "base64..."
160
+ │ ├── ocrText: "Text extracted via OCR"
161
+ │ └── chartData: { title, dataSets, labels, ... }
162
+ └── toText(): Function -> returns full plain text
163
+ ```
164
+
165
+ #### Representative JSON Snippet
166
+ ```json
167
+ {
168
+ "type": "docx",
169
+ "metadata": { "author": "John Doe", "title": "Annual Report" },
170
+ "content": [
171
+ {
172
+ "type": "heading",
173
+ "text": "Introduction",
174
+ "metadata": { "level": 1 },
175
+ "children": [
176
+ { "type": "text", "text": "Introduction", "formatting": { "bold": true } }
177
+ ]
178
+ },
179
+ {
180
+ "type": "paragraph",
181
+ "text": "This is a report with an image.",
182
+ "children": [
183
+ { "type": "text", "text": "This is a report with an " },
184
+ { "type": "image", "metadata": { "attachmentName": "img1.png" } }
185
+ ]
186
+ }
187
+ ],
188
+ "attachments": [
189
+ { "name": "img1.png", "type": "image", "data": "iVBOR...", "ocrText": "Extracted Text" }
190
+ ]
145
191
  }
192
+ ```
193
+
194
+ ## Deep Dive: Document Components
146
195
 
147
- // relative path is also fine => eg: files/myWorkSheet.ods
148
- officeParser.parseOfficeAsync("/Users/harsh/Desktop/files/mySlides.pptx", config);
149
- .then(data => {
150
- const newText = data + " look, I can parse a powerpoint file";
151
- callSomeOtherFunction(newText);
152
- })
153
- .catch(err => console.error(err));
154
-
155
- // Search for a term in the parsed text.
156
- function searchForTermInOfficeFile(searchterm, filepath) {
157
- return officeParser.parseOfficeAsync(filepath)
158
- .then(data => data.indexOf(searchterm) != -1)
196
+ ### 1. Working with Lists
197
+ Lists are represented as sequential `list` nodes. To reconstruct or track a list, use the `metadata` fields:
198
+
199
+ ```text
200
+ List Node
201
+ ├── type: "list"
202
+ ├── metadata: {
203
+ listId: "1",
204
+ listType: "ordered",
205
+ indentation: 0,
206
+ itemIndex: 0
159
207
  }
208
+ └── children: [ Text Content... ]
160
209
  ```
161
210
 
211
+ - **`listId`**: A unique identifier for the list definition. Multiple items with the same `listId` belong to the same logical list.
212
+ - **`indentation`**: The nesting level (0-based).
213
+ - **`itemIndex`**: The sequential position within that list level.
214
+ - **`listType`**: Either `ordered` (numbered) or `unordered` (bulleted).
215
+
216
+ > [!TIP]
217
+ > Even if a list is interrupted by a regular paragraph, the `itemIndex` will continue to increment for the same `listId`, allowing you to maintain correct numbering.
218
+
219
+ ### 2. Navigating Tables
220
+ Tables follow a strict hierarchy: `table` -> `row` -> `cell`.
221
+
222
+ ```text
223
+ Table Node
224
+ ├── type: "table"
225
+ └── children: [ Row Node ]
226
+ ├── type: "row"
227
+ └── children: [ Cell Node ]
228
+ ├── type: "cell"
229
+ ├── metadata: { row, col, rowSpan, colSpan }
230
+ └── children: [ Paragraph/List/etc. ]
231
+ ```
162
232
 
163
- **Example - TypeScript**
164
- ```ts
165
- import { OfficeParserConfig, parseOfficeAsync } from 'officeparser';
233
+ - **`row` / `col`**: Zero-based indices for grid positioning.
234
+ - **`rowSpan` / `colSpan`** (Optional): Integer values indicating merged cells (primarily in ODF formats). If absent, the cell is not merged.
235
+ - **Recursive Content**: Cells contain their own `children` array, which can include paragraphs, lists, or even other nested tables.
236
+
237
+ ### 3. Charts & Data
238
+ When a chart is discovered, it's added as a `chart` node in the content and a corresponding `OfficeAttachment`.
239
+
240
+ ```text
241
+ Chart Node
242
+ ├── type: "chart"
243
+ ├── metadata: { attachmentName: "chart1.xml" }
244
+ └── Attachment (Linked)
245
+ └── chartData: { title, dataSets: [...], labels: [...] }
246
+ ```
247
+
248
+ - **`attachmentName`**: Links the content node to the `attachments` array.
249
+ - **`chartData`**: A structured object containing titles, axis labels, and category/series data.
250
+
251
+ ### 4. Images, OCR & Alt Text
252
+ Images are linked via `attachmentName` and can contain valuable metadata:
253
+
254
+ ```text
255
+ Image Node
256
+ ├── type: "image"
257
+ ├── metadata: { attachmentName: "img1.png", altText: "..." }
258
+ └── Attachment (Linked)
259
+ ├── data: "base64..."
260
+ └── ocrText: "Extracted via OCR"
261
+ ```
166
262
 
167
- const config: OfficeParserConfig = {
168
- newlineDelimiter: " ", // Separate new lines with a space instead of the default \n.
169
- ignoreNotes: true // Ignore notes while parsing presentation files like pptx or odp.
263
+ - **OCR Text**: If `ocr: true` is set in config, `ocrText` will contain the text found within the image.
264
+ - **Alt Text**: Extracted from the document's internal image descriptions.
265
+ - **Formatting**: `OfficeContentNode` images may also have parent alignment metadata.
266
+
267
+ ### 5. Text Formatting
268
+ Each `OfficeContentNode` can have a `formatting` object that defines how the text should be styled.
269
+
270
+ ```text
271
+ Text Node
272
+ └── formatting: {
273
+ bold: boolean,
274
+ italic: boolean,
275
+ underline: boolean,
276
+ strikethrough: boolean,
277
+ color: "#hex",
278
+ backgroundColor: "#hex",
279
+ size: "12pt",
280
+ font: "Arial",
281
+ subscript: boolean,
282
+ superscript: boolean,
283
+ alignment: "left" | "center" | "right" | "justify"
170
284
  }
285
+ ```
286
+
287
+ Formatting can be found at two levels:
288
+ 1. **Node Level**: Applied directly to a text run or paragraph.
289
+ 2. **Document Level**: Found in `ast.metadata.formatting` (defaults) or `ast.metadata.styleMap` (named styles).
290
+
291
+ ### 6. Advanced Metadata
292
+ The `ast.metadata` object provides document-wide context:
293
+ - **`styleMap`**: A dictionary of style names to their `TextFormatting` definitions found in the document.
294
+ - **`formatting`**: Document-wide default settings (e.g., default font or font size).
295
+
296
+ ### Advanced AST Usage
297
+ Beyond using `ast.toText()`, you can interact with the structural data directly:
298
+
299
+ #### 1. Extract all images and their OCR text
300
+ ```javascript
301
+ const ast = await officeParser.parseOffice("report.docx", { ocr: true });
302
+ const images = ast.attachments.filter(a => a.mimeType.startsWith('image/'));
303
+ images.forEach(img => {
304
+ console.log(`Image: ${img.name} (OCR: ${img.ocrText || 'N/A'})`);
305
+ });
306
+ ```
307
+
308
+ #### 2. Find specific headings
309
+ ```javascript
310
+ const headings = ast.content.filter(node => node.type === 'heading' && node.metadata?.level === 1);
311
+ console.log("Main Chapters:", headings.map(h => h.text));
312
+ ```
313
+
314
+ #### 3. Custom output (e.g., Simple Markdown conversion)
315
+ ```javascript
316
+ const toMarkdown = (nodes) => {
317
+ return nodes.map(node => {
318
+ if (node.type === 'heading') return `${'#'.repeat(node.metadata?.level || 1)} ${node.text}`;
319
+ if (node.type === 'list') return `- ${node.text}`;
320
+ if (node.type === 'table') return "[Table Data]"; // expand children for actual table
321
+ return node.text;
322
+ }).join('\n\n');
323
+ };
324
+ console.log(toMarkdown(ast.content));
325
+ ```
326
+
327
+ #### 4. Extracting Tables to CSV
328
+ Iterate through table nodes and their children (rows -> cells) to build a CSV string.
329
+ ```javascript
330
+ const tables = ast.content.filter(node => node.type === 'table');
331
+ tables.forEach((table, index) => {
332
+ const csv = table.children
333
+ .filter(row => row.type === 'row')
334
+ .map(row =>
335
+ row.children
336
+ .filter(cell => cell.type === 'cell')
337
+ .map(cell => `"${cell.text.replace(/"/g, '""')}"`) // Escape quotes
338
+ .join(',')
339
+ )
340
+ .join('\n');
341
+ console.log(`Table ${index + 1} CSV:\n${csv}`);
342
+ });
343
+ ```
344
+
345
+ #### 5. Filtering by Formatting (e.g., Bold Text)
346
+ Find all text nodes that have specific formatting applied.
347
+ ```javascript
348
+ function findBoldText(nodes) {
349
+ let results = [];
350
+ nodes.forEach(node => {
351
+ if (node.type === 'text' && node.formatting?.bold) {
352
+ results.push(node.text);
353
+ }
354
+ if (node.children) {
355
+ results = results.concat(findBoldText(node.children));
356
+ }
357
+ });
358
+ return results;
359
+ }
360
+
361
+ const boldStrings = findBoldText(ast.content);
362
+ console.log("Bold Text Found:", boldStrings);
363
+ ```
364
+
365
+ #### 6. Processing Footnotes/Endnotes
366
+ If you kept notes inline (default behavior), you can extract them into a separate list for processing.
367
+ ```javascript
368
+ function extractNotes(nodes) {
369
+ let notes = [];
370
+ nodes.forEach(node => {
371
+ if (node.type === 'note') {
372
+ notes.push({ id: node.metadata.noteId, text: node.text, type: node.metadata.noteType });
373
+ }
374
+ if (node.children) {
375
+ notes = notes.concat(extractNotes(node.children));
376
+ }
377
+ });
378
+ return notes;
379
+ }
380
+
381
+ const allNotes = extractNotes(ast.content);
382
+ console.log("Document Notes:", allNotes);
383
+ ```
384
+
385
+ ## Configuration Object: OfficeParserConfig
386
+ Pass an optional config object as the second argument to `parseOffice`.
387
+
388
+ | Flag | DataType | Default | Explanation |
389
+ |------|----------|---------|-------------|
390
+ | `outputErrorToConsole` | boolean | `false` | Show logs to console in case of an error. |
391
+ | `newlineDelimiter` | string | `\n` | Delimiter for new lines in text output. |
392
+ | `ignoreNotes` | boolean | `false` | Ignore notes in files like PowerPoint/ODP. |
393
+ | `putNotesAtLast` | boolean | `false` | Put notes text at the end of the document. (Note: Does not work for RTF. It is treated as true always.) |
394
+ | `extractAttachments` | boolean | `false` | Extract images and charts as Base64. |
395
+ | `ocr` | boolean | `false` | Enable OCR for images (requires `extractAttachments: true`). |
396
+ | `ocrLanguage` | string | `eng` | Language for OCR (e.g., 'eng', 'fra'). Supports multiple languages with '+'. See [Language Codes](https://tesseract-ocr.github.io/tessdoc/Data-Files#data-files-for-version-400-november-29-2016). |
397
+ | `includeRawContent` | boolean | `false` | Include raw XML/RTF markup in the nodes. |
398
+ | `pdfWorkerSrc` | string | `(see below)` | Path to PDF.js worker. Defaults to a CDN link if not provided. |
399
+
400
+ ```js
401
+ const config = {
402
+ newlineDelimiter: "\n\n",
403
+ extractAttachments: true,
404
+ ocr: true,
405
+ ocrLanguage: 'eng+fra+esp' // Supports English, French, and Spanish simultaneously
406
+ };
407
+
408
+ const ast = await officeParser.parseOffice("report.docx", config);
409
+ console.log(`Extracted ${ast.attachments.length} images`);
410
+ ```
411
+
412
+ ## Examples
413
+
414
+ **Search for a term in a document (TypeScript)**
415
+ ```ts
416
+ import { OfficeParser } from 'officeparser';
171
417
 
172
- // relative path is also fine => eg: files/myWorkSheet.ods
173
- parseOfficeAsync("/Users/harsh/Desktop/files/mySlides.pptx", config);
174
- .then(data => {
175
- const newText = data + " look, I can parse a powerpoint file";
176
- callSomeOtherFunction(newText);
177
- })
178
- .catch(err => console.error(err));
179
-
180
- // Search for a term in the parsed text.
181
- function searchForTermInOfficeFile(searchterm: string, filepath: string): Promise<boolean> {
182
- return parseOfficeAsync(filepath)
183
- .then(data => data.indexOf(searchterm) != -1)
418
+ async function hasSearchTerm(filePath: string, term: string): Promise<boolean> {
419
+ const ast = await OfficeParser.parseOffice(filePath);
420
+ return ast.toText().includes(term);
184
421
  }
185
422
  ```
186
- \
187
- **Please take note: I have breached convention in placing err as second argument in my callback but please understand that I had to do it to not break other people's existing modules.**
423
+
424
+ **Extracting Images and their OCR text**
425
+ ```js
426
+ const officeParser = require('officeparser');
427
+
428
+ const config = { extractAttachments: true, ocr: true };
429
+ officeParser.parseOffice("presentation.pptx", config).then(ast => {
430
+ ast.attachments.forEach(attachment => {
431
+ if (attachment.type === 'image') {
432
+ console.log(`Image: ${attachment.name}`);
433
+ console.log(`OCR Text: ${attachment.ocrText}`);
434
+ fs.writeFileSync(attachment.name, Buffer.from(attachment.data, 'base64'));
435
+ }
436
+ });
437
+ });
438
+ ```
188
439
 
189
440
  ## Browser Usage
190
- Download the bundle file available as part of the release asset.
191
- Include this bundle file in your browser html file and access `parseOffice` and `parseOfficeAsync` under the **`officeParser`** namespace.
441
+ The browser bundle exposes the `officeParser` namespace. Include the bundle file available in the release assets.
192
442
 
193
- **Example**
194
443
  ```html
195
- <head>
196
- ...
197
- <!-- Include bundle file in the script tag. -->
198
- <script src="officeParserBundle@5.1.0.js"></script>
199
- </head>
200
- <body>
201
- ...
202
- <input type="file" id="fileInput" />
203
- ...
204
- <script>
205
- document.getElementById('fileInput').addEventListener('change', async function(event) {
206
- const outputDiv = document.getElementById('output');
207
- const file = event.target.files[0];
208
- try {
209
- // Your configuration options for officeParser
210
- const config = {
211
- outputErrorToConsole: false,
212
- newlineDelimiter: '\n',
213
- ignoreNotes: false,
214
- putNotesAtLast: false
215
- };
216
-
217
- const arrayBuffer = await file.arrayBuffer();
218
- const result = await officeParser.parseOfficeAsync(arrayBuffer, config);
219
- // result contains the extracted text.
220
- }
221
- catch (error) {
222
- // Handle error
223
- }
224
- });
225
- </script>
226
- </body>
227
- ```
228
-
229
-
230
- ## Known Bugs
231
- 1. Inconsistency and incorrectness in the positioning of footnotes and endnotes in .docx files where the footnotes and endnotes would end up at the end of the parsed text whereas it would be positioned exactly after the referenced word in .odt files.
232
- 2. The charts and objects information of .odt files are not accurate and may end up showing a few NaN in some cases.
233
- 3. Extracting texts in browser bundles does not work for pdf files.
444
+ <script src="dist/officeparser.browser.js"></script>
445
+ <script>
446
+ async function handleFile(file) {
447
+ // file can be a File object from an input element or an ArrayBuffer
448
+ // The browser bundle exposes the global variable `officeParser`
449
+ // which contains the `OfficeParser` class.
450
+
451
+ try {
452
+ const ast = await officeParser.parseOffice(file, { ocr: true });
453
+ console.log(ast.toText());
454
+ console.log("Metadata:", ast.metadata);
455
+ } catch (error) {
456
+ console.error(error);
457
+ }
458
+ }
459
+ </script>
460
+ ```
461
+
462
+ ### PDF Worker Configuration in Browser
463
+ When using `officeparser` in a browser environment to parse PDF files, you may provide the `pdfWorkerSrc` configuration option. If not provided, it defaults to a CDN link for `pdfjs-dist@5.4.530`.
464
+
465
+ ```javascript
466
+ const file = ...; // File object or ArrayBuffer
467
+
468
+ // It will use the default CDN worker if pdfWorkerSrc is omitted
469
+ const ast = await officeParser.parseOffice(file);
470
+
471
+ // Or override it with your own path or a different version:
472
+ const ast2 = await officeParser.parseOffice(file, {
473
+ pdfWorkerSrc: "https://unpkg.com/pdfjs-dist@5.4.530/build/pdf.worker.min.mjs"
474
+ });
475
+ ```
476
+
477
+ > **Note:** The version of `pdfjs-dist` in the worker source should match the version used by `officeparser` (currently `5.4.530`).
478
+
479
+ ## Known Limitations
480
+ 1. **ODT/ODS Charts**: Extraction may occasionally show inaccurate data when referencing external cell ranges or complex layout-based data.
481
+ 2. **PDF Images**: PDF images are extracted as BMP files in the browser for compatibility. This conversion happens automatically.
482
+ 3. **RTF Footnotes**: The `putNotesAtLast` configuration is currently not supported for RTF files; footnotes and endnotes are always collected and appended to the end of the content.
483
+
234
484
  ----------
235
485
 
236
- **npm**
237
- https://npmjs.com/package/officeparser
486
+ **npm**: [https://npmjs.com/package/officeparser](https://npmjs.com/package/officeparser)
238
487
 
239
- **github**
240
- https://github.com/harshankur/officeParser
488
+ **github**: [https://github.com/harshankur/officeParser](https://github.com/harshankur/officeParser)