officeparser 6.1.0 → 6.1.1

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -23,44 +23,12 @@ A robust, strictly-typed Node.js and Browser library for parsing office files ([
23
23
  ---
24
24
 
25
25
 
26
- #### Update
27
- * 2026-04-14 - **v6.1.0 Release**: Major Infrastructure & Resource Stability. (Incremental since v6.0.0)
28
- - **OCR Scheduler**: Intelligent worker pool that optimizes Tesseract lifecycle across parallel requests. **Note**: By default, Node.js processes stay active for 10s after OCR to keep workers warm (configurable via `ocrConfig.autoTerminateTimeout`); use `terminateOcr()` for immediate CLI/script exit.
29
- - **Core Engine**: Replaced legacy zip extraction with `fflate` for significant performance gains and robust browser/edge compatibility.
30
- - **Module System**: Full native ESM support with `Node16` resolution and verified browser bundles (Vite/Angular compatible).
31
- - **Format Refinements**: Hierarchical PDF coordinate alignment and ODT/RTF list parsing stability.
32
- - **Custom Properties**: Added support for extracting custom document metadata across OOXML, ODF, and PDF formats.
33
- - **Sponsorship**: Integrated `funding.json` manifest and GitHub Sponsors support.
34
- * 2025/12/29 - **v6.0.0 Release**: Major overhaul of the library. Transitioned from simple text extraction to a rich **Abstract Syntax Tree (AST)** output.
35
- - Simplified API: Use `parseOffice` for all parsing needs (returns a Promise).
36
- - Structured Output: Access hierarchical document structure (paragraphs, headings, tables, lists, etc.).
37
- - Rich Metadata: Extracted document properties (author, title, creation date).
38
- - Enhanced Formatting: Support for bold, italic, colors, fonts, alignment, etc.
39
- - Attachment Handling: Extract images, charts, and embedded files as Base64.
40
- - OCR Integration: Optional OCR for images using Tesseract.js.
41
- - RTF Support: Added full support for Rich Text Format files.
42
- - Improved Type Definitions: Full TypeScript support with detailed interfaces.
43
- * 2024/11/12 - Added ArrayBuffer as a type of file input. Generating bundle files now which exposes namespace officeParser to be able to access parseOffice directly on the browser.
44
- * 2024/10/21 - Replaced extracting zip files from decompress to yauzl. This means that we now extract files in memory and we no longer need to write them to disk. Removed config flags related to extracted files. Added flags for CLI execution.
45
- * 2024/10/15 - Fixed erroring out while deleting temp files when multiple worker threads make parallel executions resulting in same file name for multiple files. Fixed erroring out when multiple executions are made without waiting for the previous execution to finish which resulted in deleting the file from other execution. Upgraded dependencies.
46
- * 2024/10/13 - Fixed parsing text from xlsx files which contain no shared strings file and files which have inlineStr based strings.
47
- * 2024/05/06 - Replaced pdf parsing support from pdf-parse library to natively building it using pdf.js library from Mozilla by analyzing its output. Added pdfjs-dist build as a local library.
48
- * 2023/11/25 - Fixed error catching when an error occurs within the parsing of a file, especially after decompressing it. Also fixed the problem with parallel parsing of files as we were using only timestamp in file names.
49
- * 2023/10/24 - Revamped content parsing code. Fixed order of content in files, especially in word files where table information would always land up at the end of the text. Added config object as argument for parseOffice which can be used to set new line delimiter and multiple other configurations. Added support for parsing pdf files using the popular npm library pdf-parse. Removed support for individual file parsing functions.
50
- * 2023/04/26 - Added support for file buffers as argument for filepath for parseOffice and parseOfficeAsync
51
- * 2023/04/07 - Added typings to methods to help with Typescript projects.
52
- * 2022/12/28 - Added command line method to use officeParser with or without installing it and instantly get parsed content on the console.
53
- * 2022/12/10 - Fixed memory leak issues, bugs related to parsing open document files and improved error handling.
54
- * 2021/11/21 - Added promise way to existing callback functions.
55
- * 2020/06/01 - Added error handling and console.log enable/disable methods. Default is set at enabled. Everything backward compatible.
56
- * 2019/06/17 - Added method to change location for decompressing office files in places with restricted write access.
57
- * 2019/04/30 - Removed case sensitive file extension bug. File names with capital lettered extensions now supported.
58
- * 2019/04/23 - Added support for open office files *.odt, *.odp, *.ods through parseOffice function. Created a new method parseOpenOffice for those who prefer targetted functions.
59
- * 2019/04/23 - Added feature to delete the generated dist folder after function callback.
60
- * 2019/04/22 - Added parseOffice method to avoid confusion between type of file and their extension.
61
- * 2019/04/22 - Added file extension validations. Removed errors for excel files with no drawing elements.
62
- * 2019/04/19 - Support added for *.xlsx files.
63
- * 2019/04/18 - Support added for *.pptx files.
26
+ ---
27
+
28
+ ### 📝 [Changelog](CHANGELOG.md)
29
+ *Detailed release notes and the full history of updates are available in the project changelog.*
30
+
31
+ ---
64
32
 
65
33
  ## Install via npm
66
34
 
@@ -91,6 +59,7 @@ npx officeparser /path/to/officeFile.docx --ignoreNotes=true --newlineDelimiter=
91
59
  - `--extractAttachments=[true|false]` Flag to extract images/charts as Base64. Default is false.
92
60
  - `--ocr=[true|false]` Flag to enable OCR for extracted images. Default is false.
93
61
  - `--includeRawContent=[true|false]` Flag to include raw XML/RTF content in nodes. Default is false.
62
+ - `--includeBreakNodes=[true|false]` Flag to include break nodes. Currently only available for DOCX documents
94
63
  - `--verbose=[true|false]` Show full error stack traces.
95
64
 
96
65
 
@@ -131,7 +100,7 @@ console.log(text);
131
100
  ```
132
101
 
133
102
  ### Using Callbacks (Backward Compatibility Support)
134
- We still support callbacks, but the data returned is now the AST object.
103
+ Callbacks are still supported for those preferred, but the data returned is now the AST object.
135
104
  ```js
136
105
  const officeParser = require('officeparser');
137
106
 
@@ -172,7 +141,7 @@ OfficeParserAST
172
141
  │ ├── text: "Concatenated text of this node and all children"
173
142
  │ ├── children: [ OfficeContentNode ] (recursive)
174
143
  │ ├── formatting: { bold, italic, color, size, font, ... }
175
- │ ├── metadata: { level, listId, row, col, ... }
144
+ │ ├── metadata: { level, listId, paragraphIndentation, row, col, ... }
176
145
  │ └── rawContent: "<xml>...</xml>" (if enabled)
177
146
  ├── attachments: [ OfficeAttachment ]
178
147
  │ ├── type: "image" | "chart"
@@ -224,13 +193,15 @@ List Node
224
193
  listId: "1",
225
194
  listType: "ordered",
226
195
  indentation: 0,
196
+ paragraphIndentation: { left: 720, hanging: 360 },
227
197
  itemIndex: 0
228
198
  }
229
199
  └── children: [ Text Content... ]
230
200
  ```
231
201
 
232
202
  - **`listId`**: A unique identifier for the list definition. Multiple items with the same `listId` belong to the same logical list.
233
- - **`indentation`**: The nesting level (0-based).
203
+ - **`indentation`**: The structural nesting level (0-based).
204
+ - **`paragraphIndentation`**: The physical indentation formatting in twentieths of a point (twips) (e.g., `left`, `right`, `firstLine`, `hanging`).
234
205
  - **`itemIndex`**: The sequential position within that list level.
235
206
  - **`listType`**: Either `ordered` (numbered) or `unordered` (bulleted).
236
207
 
@@ -309,13 +280,31 @@ Formatting can be found at two levels:
309
280
  1. **Node Level**: Applied directly to a text run or paragraph.
310
281
  2. **Document Level**: Found in `ast.metadata.formatting` (defaults) or `ast.metadata.styleMap` (named styles).
311
282
 
312
- ### 6. Advanced Metadata
283
+ ### 6. Breaks
284
+ Breaks are currently only supported when parsing DOCX-documents. Breaks are added as a node of type `break` and carry metadata of the type `BreakMetadata`. When `includeRawContent` is enabled, they also include the `rawContent` string from the original XML.
285
+
286
+ ```text
287
+ Break Node
288
+ ├── type: "break"
289
+ └── metadata: {
290
+ breakType: "textWrapping" | "page" | "column" | "lastRenderedPage" | "carriageReturn",
291
+ clear?: "all" | "left" | "none" | "right"
292
+ }
293
+ ```
294
+
295
+ - `breakType`: Type of break. `textWrapping` (default) is a standard line break, `page` is a page break, `column` is a break to the next column, `lastRenderedPage` is a soft break inserted by Word, and `carriageReturn` is an explicit carriage return (`w:cr`).
296
+ - `clear`: Relevant for `textWrapping`. Indicates if text should wrap around floating objects.
297
+
298
+ > [!NOTE]
299
+ > Even though break nodes don't have a `text` property, the `ast.toText()` method will automatically convert them to newlines (`\n`) or the configured delimiter in the final string output.
300
+
301
+ ### 7. Advanced Metadata
313
302
  The `ast.metadata` object provides document-wide context:
314
303
  - **`styleMap`**: A dictionary of style names to their `TextFormatting` definitions found in the document.
315
304
  - **`formatting`**: Document-wide default settings (e.g., default font or font size).
316
305
  - **`customProperties`**: A dictionary of user-defined metadata embedded in the document (OOXML `custom.xml`, ODF `meta:user-defined`, or PDF Info dictionary).
317
306
 
318
- ### 7. Custom Properties
307
+ ### 8. Custom Properties
319
308
  You can access custom user-defined metadata that might be embedded in the document:
320
309
 
321
310
  ```javascript
@@ -435,12 +424,13 @@ Pass an optional config object as the second argument to `parseOffice`.
435
424
  | `ocrConfig.workerPath` | string | `undefined` | Path to Tesseract worker script (for offline use). |
436
425
  | `ocrConfig.corePath` | string | `undefined` | Path to Tesseract core script (for offline use). |
437
426
  | `ocrConfig.langPath` | string | `undefined` | Path for Tesseract language files (for offline use). |
427
+ | `includeBreakNodes` | boolean | `false` | Specifically targets Word documents (DOCX). When set to true, officeParser will also parse `w:br`, `w:cr` and `w:lastRenderedPageBreak` nodes.|
438
428
 
439
429
  ### OCR Scheduler & Resource Management
440
430
  If your application uses OCR, `officeParser` utilizes an intelligent **Smart Worker Pool** to maintain a background worker pool and optimize repeated parse requests.
441
431
 
442
432
  - **Dynamic Affinity**: Workers in the pool persist with their last used language affinity.
443
- - **Smart Re-initialization**: If a new language is requested and the pool is full, the manager identifies the **Least Recently Used (LRU)** idle worker and re-initializes it for the new language using the Tesseract.js v5 API. This avoids the overhead of destroying and recreating workers.
433
+ - **LRU Re-allocation**: If a new language is requested and the pool is full, the manager identifies the **Least Recently Used (LRU)** idle worker and re-initializes it for the new language. This avoids the overhead of destroying and recreating workers.
444
434
  - **Auto-Termination**: Workers are automatically cleaned up after 10 seconds of inactivity (configurable via `ocrConfig.autoTerminateTimeout`).
445
435
 
446
436
  #### `OfficeParser.terminateOcr()`
@@ -462,7 +452,7 @@ async function runCleaner() {
462
452
  ```
463
453
 
464
454
  > [!TIP]
465
- > This is handled automatically in our own CLI (`npx officeparser ...`). You only need to call this manually if you are using the library in your own custom script and want a snappy exit.
455
+ > This is handled automatically in the built-in CLI (`npx officeparser ...`). You only need to call this manually if you are using the library in your own custom script and want a snappy exit.
466
456
 
467
457
  ```js
468
458
  const config = {
@@ -509,36 +499,51 @@ The library provides two types of browser bundles in the `dist/` directory:
509
499
  1. **`officeparser.browser.iife.js`**: Standard IIFE bundle for direct `<script>` tag usage. Exposes the global `officeParser` namespace.
510
500
  2. **`officeparser.browser.mjs`**: Modern ESM bundle for use with `import` statements or modern bundlers.
511
501
 
502
+ ### Usage (ESM)
503
+ If you are using a modern bundler like **Vite**, **Webpack**, or **Next.js**:
504
+
505
+ ```javascript
506
+ import { OfficeParser } from 'officeparser';
507
+
508
+ const handleFile = async (event) => {
509
+ const file = event.target.files[0];
510
+ const buffer = await file.arrayBuffer();
511
+
512
+ try {
513
+ // Pass the Buffer or Uint8Array directly
514
+ const ast = await OfficeParser.parseOffice(new Uint8Array(buffer));
515
+ console.log(ast.toText());
516
+ } catch (err) {
517
+ console.error(err);
518
+ }
519
+ };
520
+ ```
521
+
522
+ > [!NOTE]
523
+ > **Why `fs` fails in the browser**: Browsers do not have a built-in file system. If you try to pass a file path string in the browser, `officeParser` will throw a descriptive "Fail-Fast" error instead of crashing mysteriously:
524
+ > `[officeparser] Node.js 'fs' module is not available in the browser. Please pass a Buffer or Uint8Array instead.`
525
+
512
526
  ### Usage (Script Tag)
513
- Include the IIFE bundle file available in the release assets.
527
+ Include the IIFE bundle available in the release assets or your `dist/` folder. This exposes the global `officeParser` object.
514
528
 
515
529
  ```html
516
530
  <script src="dist/officeparser.browser.iife.js"></script>
517
531
  <script>
518
- async function handleFile(file) {
519
- // file can be a File object from an input element or an ArrayBuffer
532
+ async function handleFile(event) {
533
+ const file = event.target.files[0];
534
+ const buffer = await file.arrayBuffer();
535
+
520
536
  try {
521
- const ast = await officeParser.parseOffice(file, { ocr: true });
537
+ // Reconstruct as Uint8Array for the parser
538
+ const ast = await officeParser.parseOffice(new Uint8Array(buffer));
522
539
  console.log(ast.toText());
523
540
  } catch (error) {
524
- console.error(error);
541
+ console.error("Parsing failed:", error);
525
542
  }
526
543
  }
527
544
  </script>
528
545
  ```
529
546
 
530
- ### Usage (ESM)
531
- If you are using a modern browser that supports modules or a dev server like Vite:
532
-
533
- ```html
534
- <script type="module">
535
- import { OfficeParser } from './dist/officeparser.browser.mjs';
536
-
537
- const ast = await OfficeParser.parseOffice(fileBuffer);
538
- console.log(ast.metadata);
539
- </script>
540
- ```
541
-
542
547
  ### PDF Worker Configuration in Browser
543
548
  When using `officeparser` in a browser environment to parse PDF files, you may provide the `pdfWorkerSrc` configuration option. If not provided, it defaults to a CDN link for `pdfjs-dist@5.6.205`.
544
549
 
@@ -579,7 +584,7 @@ For a comprehensive guide, visit our [Debugging & Troubleshooting Documentation]
579
584
 
580
585
  ## Contributing
581
586
 
582
- We welcome contributions! Please see [CONTRIBUTING.md](CONTRIBUTING.md) for details on how to get started.
587
+ Contributions are welcome! Please see [CONTRIBUTING.md](CONTRIBUTING.md) for details on how to get started.
583
588
 
584
589
  ## License
585
590
 
@@ -120,6 +120,7 @@ class OfficeParser {
120
120
  preserveXmlWhitespace: false,
121
121
  pdfWorkerSrc: '',
122
122
  ocrConfig: {},
123
+ includeBreakNodes: false,
123
124
  ...actualConfig
124
125
  };
125
126
  let buffer = Buffer.alloc(0);
package/dist/cli.js CHANGED
@@ -106,6 +106,7 @@ else {
106
106
  console.log(' --includeRawContent=true Include raw content in AST');
107
107
  console.log(' --serializeRawContent=true Serialize raw XML content (default: true)');
108
108
  console.log(' --preserveXmlWhitespace=true Preserve whitespace in serialized XML (default: false)');
109
+ console.log(' --includeBreakNodes=false Include break nodes (DOCX only, default: false)');
109
110
  console.log(' --verbose=true Show full error stack traces');
110
111
  console.log('');
111
112
  console.log('Examples:');
package/dist/index.d.ts CHANGED
@@ -44,8 +44,8 @@
44
44
  * @module officeparser
45
45
  */
46
46
  import { OfficeParser } from './OfficeParser.js';
47
- import { OfficeParserConfig, OfficeParserAST, OfficeContentNode, OfficeAttachment, OfficeMetadata, TextFormatting, SupportedFileType, OfficeContentNodeType, OfficeMimeType, SlideMetadata, SheetMetadata, HeadingMetadata, ListMetadata, CellMetadata, ImageMetadata, PageMetadata, ContentMetadata } from './types';
47
+ import { OfficeParserConfig, OfficeParserAST, OfficeContentNode, OfficeAttachment, OfficeMetadata, TextFormatting, SupportedFileType, OfficeContentNodeType, OfficeMimeType, SlideMetadata, SheetMetadata, HeadingMetadata, ListMetadata, CellMetadata, ImageMetadata, PageMetadata, ContentMetadata, BreakMetadata } from './types';
48
48
  declare const parseOffice: typeof OfficeParser.parseOffice;
49
49
  declare const terminateOcr: typeof OfficeParser.terminateOcr;
50
- export { OfficeParser, parseOffice, terminateOcr, OfficeParserConfig, OfficeParserAST, OfficeContentNode, OfficeAttachment, OfficeMetadata, TextFormatting, SupportedFileType, OfficeContentNodeType, OfficeMimeType, SlideMetadata, SheetMetadata, HeadingMetadata, ListMetadata, CellMetadata, ImageMetadata, PageMetadata, ContentMetadata };
50
+ export { OfficeParser, parseOffice, terminateOcr, OfficeParserConfig, OfficeParserAST, OfficeContentNode, OfficeAttachment, OfficeMetadata, TextFormatting, SupportedFileType, OfficeContentNodeType, OfficeMimeType, SlideMetadata, SheetMetadata, HeadingMetadata, ListMetadata, CellMetadata, ImageMetadata, PageMetadata, ContentMetadata, BreakMetadata, };
51
51
  export default OfficeParser;
@@ -118,6 +118,13 @@ export interface OfficeParserConfig {
118
118
  * You can override this with your own local path or a different CDN link.
119
119
  */
120
120
  pdfWorkerSrc?: string;
121
+ /**
122
+ * Flag to include break nodes in the AST.
123
+ * This is currently only supported for Word documents. (w:br nodes)
124
+ *
125
+ * Default is false
126
+ */
127
+ includeBreakNodes?: boolean;
121
128
  }
122
129
  /**
123
130
  * Supported file types for parsing.
@@ -126,7 +133,7 @@ export type SupportedFileType = "docx" | "pptx" | "xlsx" | "odt" | "odp" | "ods"
126
133
  /**
127
134
  * Types of content nodes in the AST.
128
135
  */
129
- export type OfficeContentNodeType = "paragraph" | "heading" | "table" | "list" | "text" | "image" | "chart" | "drawing" | "slide" | "note" | "sheet" | "row" | "cell" | "page";
136
+ export type OfficeContentNodeType = "paragraph" | "heading" | "table" | "list" | "text" | "image" | "chart" | "drawing" | "slide" | "note" | "sheet" | "row" | "cell" | "page" | "break";
130
137
  /**
131
138
  * Supported MIME types for attachments.
132
139
  */
@@ -229,6 +236,20 @@ export interface SheetMetadata {
229
236
  /** The style of the sheet. */
230
237
  style?: string;
231
238
  }
239
+ /**
240
+ * Detailed indentation information for paragraphs and headings.
241
+ * Values are typically in twentieths of a point (twips) in OOXML.
242
+ */
243
+ export interface IndentationMetadata {
244
+ /** Left indentation. */
245
+ left?: number;
246
+ /** Right indentation. */
247
+ right?: number;
248
+ /** First line indentation. */
249
+ firstLine?: number;
250
+ /** Hanging indentation. */
251
+ hanging?: number;
252
+ }
232
253
  /**
233
254
  * Metadata for a heading.
234
255
  */
@@ -239,6 +260,8 @@ export interface HeadingMetadata {
239
260
  alignment?: "left" | "center" | "right" | "justify";
240
261
  /** The style of the heading. */
241
262
  style?: string;
263
+ /** Detailed indentation information. */
264
+ paragraphIndentation?: IndentationMetadata;
242
265
  }
243
266
  /**
244
267
  * Metadata for a paragraph.
@@ -248,6 +271,8 @@ export interface ParagraphMetadata {
248
271
  alignment?: "left" | "center" | "right" | "justify";
249
272
  /** The style of the paragraph. */
250
273
  style?: string;
274
+ /** Detailed indentation information. */
275
+ paragraphIndentation?: IndentationMetadata;
251
276
  }
252
277
  /**
253
278
  * Metadata for a list item.
@@ -263,6 +288,8 @@ export interface ListMetadata {
263
288
  * @example 0 for top-level items, 1 for first nested level
264
289
  */
265
290
  indentation: number;
291
+ /** Detailed indentation information. */
292
+ paragraphIndentation?: IndentationMetadata;
266
293
  /**
267
294
  * Text alignment of the list item.
268
295
  * @example 'left', 'center', 'right', 'justify'
@@ -389,10 +416,35 @@ export interface NoteMetadata {
389
416
  */
390
417
  noteId?: string;
391
418
  }
419
+ /**
420
+ * Metadata for break nodes.
421
+ * Used in DOCX files to track line and page breaks.
422
+ */
423
+ export interface BreakMetadata {
424
+ /**
425
+ * Type of break. The break type determines the next location where
426
+ * text shall be placed.
427
+ * - 'column': The next text will be placed in the next column.
428
+ * - 'page': The next text will be placed on the next page.
429
+ * - 'lastRenderedPage': The editing application has inserted a soft break on the last save.
430
+ * - 'textWrapping' (default, assumed when not specified): The next text will be placed on the next line.
431
+ * - 'carriageReturn': An explicit carriage return (w:cr) equivalent to a hard line break.
432
+ */
433
+ breakType: "column" | "page" | "lastRenderedPage" | "textWrapping" | "carriageReturn";
434
+ /**
435
+ * Specifies the location which shall be used as the next available line when breakType
436
+ * has a value of 'textWrapping'. Should be ignored for other break types.
437
+ * - 'all': text wrapping break shall advance the text to the next line which spans the full width of the line
438
+ * - 'left': text wrapping break shall restart in next text region unblocked on the left
439
+ * - 'none': text wrapping break shall advance the text to the next line regardless of any floating objects
440
+ * - 'right': text wrapping break shall restart in next text region unblocked on the right
441
+ */
442
+ clear?: "all" | "left" | "none" | "right";
443
+ }
392
444
  /**
393
445
  * Union type for content metadata.
394
446
  */
395
- export type ContentMetadata = SlideMetadata | SheetMetadata | HeadingMetadata | ListMetadata | CellMetadata | ImageMetadata | ChartMetadata | PageMetadata | ParagraphMetadata | TextMetadata | NoteMetadata | undefined;
447
+ export type ContentMetadata = SlideMetadata | SheetMetadata | HeadingMetadata | ListMetadata | CellMetadata | ImageMetadata | ChartMetadata | PageMetadata | ParagraphMetadata | TextMetadata | NoteMetadata | BreakMetadata | undefined;
396
448
  /**
397
449
  * Represents a node in the document content tree.
398
450
  * This is the core building block of the parsed document structure.