officeparser 6.1.0 → 6.1.1
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +68 -63
- package/dist/OfficeParser.js +1 -0
- package/dist/cli.js +1 -0
- package/dist/index.d.ts +2 -2
- package/dist/officeparser.browser.d.ts +54 -2
- package/dist/officeparser.browser.iife.js +49 -46
- package/dist/officeparser.browser.mjs +49 -46
- package/dist/parsers/ExcelParser.js +5 -5
- package/dist/parsers/OpenOfficeParser.js +97 -49
- package/dist/parsers/PowerPointParser.js +112 -100
- package/dist/parsers/RtfParser.d.ts +20 -0
- package/dist/parsers/RtfParser.js +107 -42
- package/dist/parsers/WordParser.d.ts +1 -0
- package/dist/parsers/WordParser.js +107 -24
- package/dist/sbom.cdx.json +102 -102
- package/dist/types.d.ts +54 -2
- package/package.json +2 -2
package/README.md
CHANGED
|
@@ -23,44 +23,12 @@ A robust, strictly-typed Node.js and Browser library for parsing office files ([
|
|
|
23
23
|
---
|
|
24
24
|
|
|
25
25
|
|
|
26
|
-
|
|
27
|
-
|
|
28
|
-
|
|
29
|
-
|
|
30
|
-
|
|
31
|
-
|
|
32
|
-
- **Custom Properties**: Added support for extracting custom document metadata across OOXML, ODF, and PDF formats.
|
|
33
|
-
- **Sponsorship**: Integrated `funding.json` manifest and GitHub Sponsors support.
|
|
34
|
-
* 2025/12/29 - **v6.0.0 Release**: Major overhaul of the library. Transitioned from simple text extraction to a rich **Abstract Syntax Tree (AST)** output.
|
|
35
|
-
- Simplified API: Use `parseOffice` for all parsing needs (returns a Promise).
|
|
36
|
-
- Structured Output: Access hierarchical document structure (paragraphs, headings, tables, lists, etc.).
|
|
37
|
-
- Rich Metadata: Extracted document properties (author, title, creation date).
|
|
38
|
-
- Enhanced Formatting: Support for bold, italic, colors, fonts, alignment, etc.
|
|
39
|
-
- Attachment Handling: Extract images, charts, and embedded files as Base64.
|
|
40
|
-
- OCR Integration: Optional OCR for images using Tesseract.js.
|
|
41
|
-
- RTF Support: Added full support for Rich Text Format files.
|
|
42
|
-
- Improved Type Definitions: Full TypeScript support with detailed interfaces.
|
|
43
|
-
* 2024/11/12 - Added ArrayBuffer as a type of file input. Generating bundle files now which exposes namespace officeParser to be able to access parseOffice directly on the browser.
|
|
44
|
-
* 2024/10/21 - Replaced extracting zip files from decompress to yauzl. This means that we now extract files in memory and we no longer need to write them to disk. Removed config flags related to extracted files. Added flags for CLI execution.
|
|
45
|
-
* 2024/10/15 - Fixed erroring out while deleting temp files when multiple worker threads make parallel executions resulting in same file name for multiple files. Fixed erroring out when multiple executions are made without waiting for the previous execution to finish which resulted in deleting the file from other execution. Upgraded dependencies.
|
|
46
|
-
* 2024/10/13 - Fixed parsing text from xlsx files which contain no shared strings file and files which have inlineStr based strings.
|
|
47
|
-
* 2024/05/06 - Replaced pdf parsing support from pdf-parse library to natively building it using pdf.js library from Mozilla by analyzing its output. Added pdfjs-dist build as a local library.
|
|
48
|
-
* 2023/11/25 - Fixed error catching when an error occurs within the parsing of a file, especially after decompressing it. Also fixed the problem with parallel parsing of files as we were using only timestamp in file names.
|
|
49
|
-
* 2023/10/24 - Revamped content parsing code. Fixed order of content in files, especially in word files where table information would always land up at the end of the text. Added config object as argument for parseOffice which can be used to set new line delimiter and multiple other configurations. Added support for parsing pdf files using the popular npm library pdf-parse. Removed support for individual file parsing functions.
|
|
50
|
-
* 2023/04/26 - Added support for file buffers as argument for filepath for parseOffice and parseOfficeAsync
|
|
51
|
-
* 2023/04/07 - Added typings to methods to help with Typescript projects.
|
|
52
|
-
* 2022/12/28 - Added command line method to use officeParser with or without installing it and instantly get parsed content on the console.
|
|
53
|
-
* 2022/12/10 - Fixed memory leak issues, bugs related to parsing open document files and improved error handling.
|
|
54
|
-
* 2021/11/21 - Added promise way to existing callback functions.
|
|
55
|
-
* 2020/06/01 - Added error handling and console.log enable/disable methods. Default is set at enabled. Everything backward compatible.
|
|
56
|
-
* 2019/06/17 - Added method to change location for decompressing office files in places with restricted write access.
|
|
57
|
-
* 2019/04/30 - Removed case sensitive file extension bug. File names with capital lettered extensions now supported.
|
|
58
|
-
* 2019/04/23 - Added support for open office files *.odt, *.odp, *.ods through parseOffice function. Created a new method parseOpenOffice for those who prefer targetted functions.
|
|
59
|
-
* 2019/04/23 - Added feature to delete the generated dist folder after function callback.
|
|
60
|
-
* 2019/04/22 - Added parseOffice method to avoid confusion between type of file and their extension.
|
|
61
|
-
* 2019/04/22 - Added file extension validations. Removed errors for excel files with no drawing elements.
|
|
62
|
-
* 2019/04/19 - Support added for *.xlsx files.
|
|
63
|
-
* 2019/04/18 - Support added for *.pptx files.
|
|
26
|
+
---
|
|
27
|
+
|
|
28
|
+
### 📝 [Changelog](CHANGELOG.md)
|
|
29
|
+
*Detailed release notes and the full history of updates are available in the project changelog.*
|
|
30
|
+
|
|
31
|
+
---
|
|
64
32
|
|
|
65
33
|
## Install via npm
|
|
66
34
|
|
|
@@ -91,6 +59,7 @@ npx officeparser /path/to/officeFile.docx --ignoreNotes=true --newlineDelimiter=
|
|
|
91
59
|
- `--extractAttachments=[true|false]` Flag to extract images/charts as Base64. Default is false.
|
|
92
60
|
- `--ocr=[true|false]` Flag to enable OCR for extracted images. Default is false.
|
|
93
61
|
- `--includeRawContent=[true|false]` Flag to include raw XML/RTF content in nodes. Default is false.
|
|
62
|
+
- `--includeBreakNodes=[true|false]` Flag to include break nodes. Currently only available for DOCX documents
|
|
94
63
|
- `--verbose=[true|false]` Show full error stack traces.
|
|
95
64
|
|
|
96
65
|
|
|
@@ -131,7 +100,7 @@ console.log(text);
|
|
|
131
100
|
```
|
|
132
101
|
|
|
133
102
|
### Using Callbacks (Backward Compatibility Support)
|
|
134
|
-
|
|
103
|
+
Callbacks are still supported for those preferred, but the data returned is now the AST object.
|
|
135
104
|
```js
|
|
136
105
|
const officeParser = require('officeparser');
|
|
137
106
|
|
|
@@ -172,7 +141,7 @@ OfficeParserAST
|
|
|
172
141
|
│ ├── text: "Concatenated text of this node and all children"
|
|
173
142
|
│ ├── children: [ OfficeContentNode ] (recursive)
|
|
174
143
|
│ ├── formatting: { bold, italic, color, size, font, ... }
|
|
175
|
-
│ ├── metadata: { level, listId, row, col, ... }
|
|
144
|
+
│ ├── metadata: { level, listId, paragraphIndentation, row, col, ... }
|
|
176
145
|
│ └── rawContent: "<xml>...</xml>" (if enabled)
|
|
177
146
|
├── attachments: [ OfficeAttachment ]
|
|
178
147
|
│ ├── type: "image" | "chart"
|
|
@@ -224,13 +193,15 @@ List Node
|
|
|
224
193
|
listId: "1",
|
|
225
194
|
listType: "ordered",
|
|
226
195
|
indentation: 0,
|
|
196
|
+
paragraphIndentation: { left: 720, hanging: 360 },
|
|
227
197
|
itemIndex: 0
|
|
228
198
|
}
|
|
229
199
|
└── children: [ Text Content... ]
|
|
230
200
|
```
|
|
231
201
|
|
|
232
202
|
- **`listId`**: A unique identifier for the list definition. Multiple items with the same `listId` belong to the same logical list.
|
|
233
|
-
- **`indentation`**: The nesting level (0-based).
|
|
203
|
+
- **`indentation`**: The structural nesting level (0-based).
|
|
204
|
+
- **`paragraphIndentation`**: The physical indentation formatting in twentieths of a point (twips) (e.g., `left`, `right`, `firstLine`, `hanging`).
|
|
234
205
|
- **`itemIndex`**: The sequential position within that list level.
|
|
235
206
|
- **`listType`**: Either `ordered` (numbered) or `unordered` (bulleted).
|
|
236
207
|
|
|
@@ -309,13 +280,31 @@ Formatting can be found at two levels:
|
|
|
309
280
|
1. **Node Level**: Applied directly to a text run or paragraph.
|
|
310
281
|
2. **Document Level**: Found in `ast.metadata.formatting` (defaults) or `ast.metadata.styleMap` (named styles).
|
|
311
282
|
|
|
312
|
-
### 6.
|
|
283
|
+
### 6. Breaks
|
|
284
|
+
Breaks are currently only supported when parsing DOCX-documents. Breaks are added as a node of type `break` and carry metadata of the type `BreakMetadata`. When `includeRawContent` is enabled, they also include the `rawContent` string from the original XML.
|
|
285
|
+
|
|
286
|
+
```text
|
|
287
|
+
Break Node
|
|
288
|
+
├── type: "break"
|
|
289
|
+
└── metadata: {
|
|
290
|
+
breakType: "textWrapping" | "page" | "column" | "lastRenderedPage" | "carriageReturn",
|
|
291
|
+
clear?: "all" | "left" | "none" | "right"
|
|
292
|
+
}
|
|
293
|
+
```
|
|
294
|
+
|
|
295
|
+
- `breakType`: Type of break. `textWrapping` (default) is a standard line break, `page` is a page break, `column` is a break to the next column, `lastRenderedPage` is a soft break inserted by Word, and `carriageReturn` is an explicit carriage return (`w:cr`).
|
|
296
|
+
- `clear`: Relevant for `textWrapping`. Indicates if text should wrap around floating objects.
|
|
297
|
+
|
|
298
|
+
> [!NOTE]
|
|
299
|
+
> Even though break nodes don't have a `text` property, the `ast.toText()` method will automatically convert them to newlines (`\n`) or the configured delimiter in the final string output.
|
|
300
|
+
|
|
301
|
+
### 7. Advanced Metadata
|
|
313
302
|
The `ast.metadata` object provides document-wide context:
|
|
314
303
|
- **`styleMap`**: A dictionary of style names to their `TextFormatting` definitions found in the document.
|
|
315
304
|
- **`formatting`**: Document-wide default settings (e.g., default font or font size).
|
|
316
305
|
- **`customProperties`**: A dictionary of user-defined metadata embedded in the document (OOXML `custom.xml`, ODF `meta:user-defined`, or PDF Info dictionary).
|
|
317
306
|
|
|
318
|
-
###
|
|
307
|
+
### 8. Custom Properties
|
|
319
308
|
You can access custom user-defined metadata that might be embedded in the document:
|
|
320
309
|
|
|
321
310
|
```javascript
|
|
@@ -435,12 +424,13 @@ Pass an optional config object as the second argument to `parseOffice`.
|
|
|
435
424
|
| `ocrConfig.workerPath` | string | `undefined` | Path to Tesseract worker script (for offline use). |
|
|
436
425
|
| `ocrConfig.corePath` | string | `undefined` | Path to Tesseract core script (for offline use). |
|
|
437
426
|
| `ocrConfig.langPath` | string | `undefined` | Path for Tesseract language files (for offline use). |
|
|
427
|
+
| `includeBreakNodes` | boolean | `false` | Specifically targets Word documents (DOCX). When set to true, officeParser will also parse `w:br`, `w:cr` and `w:lastRenderedPageBreak` nodes.|
|
|
438
428
|
|
|
439
429
|
### OCR Scheduler & Resource Management
|
|
440
430
|
If your application uses OCR, `officeParser` utilizes an intelligent **Smart Worker Pool** to maintain a background worker pool and optimize repeated parse requests.
|
|
441
431
|
|
|
442
432
|
- **Dynamic Affinity**: Workers in the pool persist with their last used language affinity.
|
|
443
|
-
- **
|
|
433
|
+
- **LRU Re-allocation**: If a new language is requested and the pool is full, the manager identifies the **Least Recently Used (LRU)** idle worker and re-initializes it for the new language. This avoids the overhead of destroying and recreating workers.
|
|
444
434
|
- **Auto-Termination**: Workers are automatically cleaned up after 10 seconds of inactivity (configurable via `ocrConfig.autoTerminateTimeout`).
|
|
445
435
|
|
|
446
436
|
#### `OfficeParser.terminateOcr()`
|
|
@@ -462,7 +452,7 @@ async function runCleaner() {
|
|
|
462
452
|
```
|
|
463
453
|
|
|
464
454
|
> [!TIP]
|
|
465
|
-
> This is handled automatically in
|
|
455
|
+
> This is handled automatically in the built-in CLI (`npx officeparser ...`). You only need to call this manually if you are using the library in your own custom script and want a snappy exit.
|
|
466
456
|
|
|
467
457
|
```js
|
|
468
458
|
const config = {
|
|
@@ -509,36 +499,51 @@ The library provides two types of browser bundles in the `dist/` directory:
|
|
|
509
499
|
1. **`officeparser.browser.iife.js`**: Standard IIFE bundle for direct `<script>` tag usage. Exposes the global `officeParser` namespace.
|
|
510
500
|
2. **`officeparser.browser.mjs`**: Modern ESM bundle for use with `import` statements or modern bundlers.
|
|
511
501
|
|
|
502
|
+
### Usage (ESM)
|
|
503
|
+
If you are using a modern bundler like **Vite**, **Webpack**, or **Next.js**:
|
|
504
|
+
|
|
505
|
+
```javascript
|
|
506
|
+
import { OfficeParser } from 'officeparser';
|
|
507
|
+
|
|
508
|
+
const handleFile = async (event) => {
|
|
509
|
+
const file = event.target.files[0];
|
|
510
|
+
const buffer = await file.arrayBuffer();
|
|
511
|
+
|
|
512
|
+
try {
|
|
513
|
+
// Pass the Buffer or Uint8Array directly
|
|
514
|
+
const ast = await OfficeParser.parseOffice(new Uint8Array(buffer));
|
|
515
|
+
console.log(ast.toText());
|
|
516
|
+
} catch (err) {
|
|
517
|
+
console.error(err);
|
|
518
|
+
}
|
|
519
|
+
};
|
|
520
|
+
```
|
|
521
|
+
|
|
522
|
+
> [!NOTE]
|
|
523
|
+
> **Why `fs` fails in the browser**: Browsers do not have a built-in file system. If you try to pass a file path string in the browser, `officeParser` will throw a descriptive "Fail-Fast" error instead of crashing mysteriously:
|
|
524
|
+
> `[officeparser] Node.js 'fs' module is not available in the browser. Please pass a Buffer or Uint8Array instead.`
|
|
525
|
+
|
|
512
526
|
### Usage (Script Tag)
|
|
513
|
-
Include the IIFE bundle
|
|
527
|
+
Include the IIFE bundle available in the release assets or your `dist/` folder. This exposes the global `officeParser` object.
|
|
514
528
|
|
|
515
529
|
```html
|
|
516
530
|
<script src="dist/officeparser.browser.iife.js"></script>
|
|
517
531
|
<script>
|
|
518
|
-
async function handleFile(
|
|
519
|
-
|
|
532
|
+
async function handleFile(event) {
|
|
533
|
+
const file = event.target.files[0];
|
|
534
|
+
const buffer = await file.arrayBuffer();
|
|
535
|
+
|
|
520
536
|
try {
|
|
521
|
-
|
|
537
|
+
// Reconstruct as Uint8Array for the parser
|
|
538
|
+
const ast = await officeParser.parseOffice(new Uint8Array(buffer));
|
|
522
539
|
console.log(ast.toText());
|
|
523
540
|
} catch (error) {
|
|
524
|
-
console.error(error);
|
|
541
|
+
console.error("Parsing failed:", error);
|
|
525
542
|
}
|
|
526
543
|
}
|
|
527
544
|
</script>
|
|
528
545
|
```
|
|
529
546
|
|
|
530
|
-
### Usage (ESM)
|
|
531
|
-
If you are using a modern browser that supports modules or a dev server like Vite:
|
|
532
|
-
|
|
533
|
-
```html
|
|
534
|
-
<script type="module">
|
|
535
|
-
import { OfficeParser } from './dist/officeparser.browser.mjs';
|
|
536
|
-
|
|
537
|
-
const ast = await OfficeParser.parseOffice(fileBuffer);
|
|
538
|
-
console.log(ast.metadata);
|
|
539
|
-
</script>
|
|
540
|
-
```
|
|
541
|
-
|
|
542
547
|
### PDF Worker Configuration in Browser
|
|
543
548
|
When using `officeparser` in a browser environment to parse PDF files, you may provide the `pdfWorkerSrc` configuration option. If not provided, it defaults to a CDN link for `pdfjs-dist@5.6.205`.
|
|
544
549
|
|
|
@@ -579,7 +584,7 @@ For a comprehensive guide, visit our [Debugging & Troubleshooting Documentation]
|
|
|
579
584
|
|
|
580
585
|
## Contributing
|
|
581
586
|
|
|
582
|
-
|
|
587
|
+
Contributions are welcome! Please see [CONTRIBUTING.md](CONTRIBUTING.md) for details on how to get started.
|
|
583
588
|
|
|
584
589
|
## License
|
|
585
590
|
|
package/dist/OfficeParser.js
CHANGED
package/dist/cli.js
CHANGED
|
@@ -106,6 +106,7 @@ else {
|
|
|
106
106
|
console.log(' --includeRawContent=true Include raw content in AST');
|
|
107
107
|
console.log(' --serializeRawContent=true Serialize raw XML content (default: true)');
|
|
108
108
|
console.log(' --preserveXmlWhitespace=true Preserve whitespace in serialized XML (default: false)');
|
|
109
|
+
console.log(' --includeBreakNodes=false Include break nodes (DOCX only, default: false)');
|
|
109
110
|
console.log(' --verbose=true Show full error stack traces');
|
|
110
111
|
console.log('');
|
|
111
112
|
console.log('Examples:');
|
package/dist/index.d.ts
CHANGED
|
@@ -44,8 +44,8 @@
|
|
|
44
44
|
* @module officeparser
|
|
45
45
|
*/
|
|
46
46
|
import { OfficeParser } from './OfficeParser.js';
|
|
47
|
-
import { OfficeParserConfig, OfficeParserAST, OfficeContentNode, OfficeAttachment, OfficeMetadata, TextFormatting, SupportedFileType, OfficeContentNodeType, OfficeMimeType, SlideMetadata, SheetMetadata, HeadingMetadata, ListMetadata, CellMetadata, ImageMetadata, PageMetadata, ContentMetadata } from './types';
|
|
47
|
+
import { OfficeParserConfig, OfficeParserAST, OfficeContentNode, OfficeAttachment, OfficeMetadata, TextFormatting, SupportedFileType, OfficeContentNodeType, OfficeMimeType, SlideMetadata, SheetMetadata, HeadingMetadata, ListMetadata, CellMetadata, ImageMetadata, PageMetadata, ContentMetadata, BreakMetadata } from './types';
|
|
48
48
|
declare const parseOffice: typeof OfficeParser.parseOffice;
|
|
49
49
|
declare const terminateOcr: typeof OfficeParser.terminateOcr;
|
|
50
|
-
export { OfficeParser, parseOffice, terminateOcr, OfficeParserConfig, OfficeParserAST, OfficeContentNode, OfficeAttachment, OfficeMetadata, TextFormatting, SupportedFileType, OfficeContentNodeType, OfficeMimeType, SlideMetadata, SheetMetadata, HeadingMetadata, ListMetadata, CellMetadata, ImageMetadata, PageMetadata, ContentMetadata };
|
|
50
|
+
export { OfficeParser, parseOffice, terminateOcr, OfficeParserConfig, OfficeParserAST, OfficeContentNode, OfficeAttachment, OfficeMetadata, TextFormatting, SupportedFileType, OfficeContentNodeType, OfficeMimeType, SlideMetadata, SheetMetadata, HeadingMetadata, ListMetadata, CellMetadata, ImageMetadata, PageMetadata, ContentMetadata, BreakMetadata, };
|
|
51
51
|
export default OfficeParser;
|
|
@@ -118,6 +118,13 @@ export interface OfficeParserConfig {
|
|
|
118
118
|
* You can override this with your own local path or a different CDN link.
|
|
119
119
|
*/
|
|
120
120
|
pdfWorkerSrc?: string;
|
|
121
|
+
/**
|
|
122
|
+
* Flag to include break nodes in the AST.
|
|
123
|
+
* This is currently only supported for Word documents. (w:br nodes)
|
|
124
|
+
*
|
|
125
|
+
* Default is false
|
|
126
|
+
*/
|
|
127
|
+
includeBreakNodes?: boolean;
|
|
121
128
|
}
|
|
122
129
|
/**
|
|
123
130
|
* Supported file types for parsing.
|
|
@@ -126,7 +133,7 @@ export type SupportedFileType = "docx" | "pptx" | "xlsx" | "odt" | "odp" | "ods"
|
|
|
126
133
|
/**
|
|
127
134
|
* Types of content nodes in the AST.
|
|
128
135
|
*/
|
|
129
|
-
export type OfficeContentNodeType = "paragraph" | "heading" | "table" | "list" | "text" | "image" | "chart" | "drawing" | "slide" | "note" | "sheet" | "row" | "cell" | "page";
|
|
136
|
+
export type OfficeContentNodeType = "paragraph" | "heading" | "table" | "list" | "text" | "image" | "chart" | "drawing" | "slide" | "note" | "sheet" | "row" | "cell" | "page" | "break";
|
|
130
137
|
/**
|
|
131
138
|
* Supported MIME types for attachments.
|
|
132
139
|
*/
|
|
@@ -229,6 +236,20 @@ export interface SheetMetadata {
|
|
|
229
236
|
/** The style of the sheet. */
|
|
230
237
|
style?: string;
|
|
231
238
|
}
|
|
239
|
+
/**
|
|
240
|
+
* Detailed indentation information for paragraphs and headings.
|
|
241
|
+
* Values are typically in twentieths of a point (twips) in OOXML.
|
|
242
|
+
*/
|
|
243
|
+
export interface IndentationMetadata {
|
|
244
|
+
/** Left indentation. */
|
|
245
|
+
left?: number;
|
|
246
|
+
/** Right indentation. */
|
|
247
|
+
right?: number;
|
|
248
|
+
/** First line indentation. */
|
|
249
|
+
firstLine?: number;
|
|
250
|
+
/** Hanging indentation. */
|
|
251
|
+
hanging?: number;
|
|
252
|
+
}
|
|
232
253
|
/**
|
|
233
254
|
* Metadata for a heading.
|
|
234
255
|
*/
|
|
@@ -239,6 +260,8 @@ export interface HeadingMetadata {
|
|
|
239
260
|
alignment?: "left" | "center" | "right" | "justify";
|
|
240
261
|
/** The style of the heading. */
|
|
241
262
|
style?: string;
|
|
263
|
+
/** Detailed indentation information. */
|
|
264
|
+
paragraphIndentation?: IndentationMetadata;
|
|
242
265
|
}
|
|
243
266
|
/**
|
|
244
267
|
* Metadata for a paragraph.
|
|
@@ -248,6 +271,8 @@ export interface ParagraphMetadata {
|
|
|
248
271
|
alignment?: "left" | "center" | "right" | "justify";
|
|
249
272
|
/** The style of the paragraph. */
|
|
250
273
|
style?: string;
|
|
274
|
+
/** Detailed indentation information. */
|
|
275
|
+
paragraphIndentation?: IndentationMetadata;
|
|
251
276
|
}
|
|
252
277
|
/**
|
|
253
278
|
* Metadata for a list item.
|
|
@@ -263,6 +288,8 @@ export interface ListMetadata {
|
|
|
263
288
|
* @example 0 for top-level items, 1 for first nested level
|
|
264
289
|
*/
|
|
265
290
|
indentation: number;
|
|
291
|
+
/** Detailed indentation information. */
|
|
292
|
+
paragraphIndentation?: IndentationMetadata;
|
|
266
293
|
/**
|
|
267
294
|
* Text alignment of the list item.
|
|
268
295
|
* @example 'left', 'center', 'right', 'justify'
|
|
@@ -389,10 +416,35 @@ export interface NoteMetadata {
|
|
|
389
416
|
*/
|
|
390
417
|
noteId?: string;
|
|
391
418
|
}
|
|
419
|
+
/**
|
|
420
|
+
* Metadata for break nodes.
|
|
421
|
+
* Used in DOCX files to track line and page breaks.
|
|
422
|
+
*/
|
|
423
|
+
export interface BreakMetadata {
|
|
424
|
+
/**
|
|
425
|
+
* Type of break. The break type determines the next location where
|
|
426
|
+
* text shall be placed.
|
|
427
|
+
* - 'column': The next text will be placed in the next column.
|
|
428
|
+
* - 'page': The next text will be placed on the next page.
|
|
429
|
+
* - 'lastRenderedPage': The editing application has inserted a soft break on the last save.
|
|
430
|
+
* - 'textWrapping' (default, assumed when not specified): The next text will be placed on the next line.
|
|
431
|
+
* - 'carriageReturn': An explicit carriage return (w:cr) equivalent to a hard line break.
|
|
432
|
+
*/
|
|
433
|
+
breakType: "column" | "page" | "lastRenderedPage" | "textWrapping" | "carriageReturn";
|
|
434
|
+
/**
|
|
435
|
+
* Specifies the location which shall be used as the next available line when breakType
|
|
436
|
+
* has a value of 'textWrapping'. Should be ignored for other break types.
|
|
437
|
+
* - 'all': text wrapping break shall advance the text to the next line which spans the full width of the line
|
|
438
|
+
* - 'left': text wrapping break shall restart in next text region unblocked on the left
|
|
439
|
+
* - 'none': text wrapping break shall advance the text to the next line regardless of any floating objects
|
|
440
|
+
* - 'right': text wrapping break shall restart in next text region unblocked on the right
|
|
441
|
+
*/
|
|
442
|
+
clear?: "all" | "left" | "none" | "right";
|
|
443
|
+
}
|
|
392
444
|
/**
|
|
393
445
|
* Union type for content metadata.
|
|
394
446
|
*/
|
|
395
|
-
export type ContentMetadata = SlideMetadata | SheetMetadata | HeadingMetadata | ListMetadata | CellMetadata | ImageMetadata | ChartMetadata | PageMetadata | ParagraphMetadata | TextMetadata | NoteMetadata | undefined;
|
|
447
|
+
export type ContentMetadata = SlideMetadata | SheetMetadata | HeadingMetadata | ListMetadata | CellMetadata | ImageMetadata | ChartMetadata | PageMetadata | ParagraphMetadata | TextMetadata | NoteMetadata | BreakMetadata | undefined;
|
|
396
448
|
/**
|
|
397
449
|
* Represents a node in the document content tree.
|
|
398
450
|
* This is the core building block of the parsed document structure.
|