officeparser 7.2.0 → 7.2.2

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -1,4 +1,4 @@
1
- # officeParser — Universal Office Document Parser & Generator
1
+ # officeParser: Universal Office Document Parser & Generator
2
2
 
3
3
  A robust, strictly-typed **Node.js and Browser** library for parsing office files into a rich **Abstract Syntax Tree (AST)** and generating high-fidelity output in multiple formats.
4
4
 
@@ -14,7 +14,7 @@ A robust, strictly-typed **Node.js and Browser** library for parsing office file
14
14
  ---
15
15
 
16
16
  ### 🌟 [Live Interactive AST Visualizer & Documentation](https://harshankur.github.io/officeParser/) 🌟
17
- *Upload any office file in your browser — inspect the AST, tweak config, and preview generated output in real-time.*
17
+ *Upload any office file in your browser: inspect the AST, tweak config, and preview generated output in real-time.*
18
18
 
19
19
  - **AST Visualizer**: Inspect the hierarchical node tree, metadata, and raw content
20
20
  - **Config Configurator**: Tweak options (`ignoreNotes`, `ocr`, `newlineDelimiter`) and see results instantly
@@ -35,10 +35,10 @@ A robust, strictly-typed **Node.js and Browser** library for parsing office file
35
35
  - [Async/Await](#asyncawait)
36
36
  - [Callback (Backward Compat)](#callback-backward-compat)
37
37
  - [File Buffers & ArrayBuffers](#file-buffers--arraybuffers)
38
- - [`ast.to()` — Generate from AST](#astto--generate-from-ast)
39
- - [`ast.toText()` — Quick Text Extraction](#asttotext--quick-text-extraction)
38
+ - [`ast.to()`: Generate from AST](#astto-generate-from-ast)
39
+ - [`ast.toText()`: Quick Text Extraction](#asttotext-quick-text-extraction)
40
40
  - [OfficeGenerator](#officegenerator)
41
- - [OfficeConverter — One-Step API](#officeconverter--one-step-api)
41
+ - [OfficeConverter: One-Step API](#officeconverter-one-step-api)
42
42
  - [Native RAG Chunking](#native-rag-chunking)
43
43
  - [The AST Structure](#the-ast-structure)
44
44
  - [Deep Dive: Document Components](#deep-dive-document-components)
@@ -47,8 +47,8 @@ A robust, strictly-typed **Node.js and Browser** library for parsing office file
47
47
  - [Configuration Reference](#configuration-reference)
48
48
  - [OfficeParserConfig](#officeparserconfig)
49
49
  - [GeneratorConfig (Common)](#generatorconfig-common)
50
- - [onNode Callback](#onnode-callback--advanced-node-manipulation)
51
- - [styleMap — Semantic Style Mapping](#stylemap--semantic-style-mapping)
50
+ - [onNode Callback](#onnode-callback-advanced-node-manipulation)
51
+ - [styleMap: Semantic Style Mapping](#stylemap-semantic-style-mapping)
52
52
  - [HtmlGeneratorConfig](#htmlgeneratorconfig)
53
53
  - [MdGeneratorConfig](#mdgeneratorconfig)
54
54
  - [PdfGeneratorConfig](#pdfgeneratorconfig)
@@ -79,37 +79,58 @@ npm i officeparser
79
79
  npx officeparser /path/to/file.docx
80
80
 
81
81
  # Plain text output
82
- npx officeparser /path/to/file.docx --format=text
82
+ npx officeparser /path/to/file.docx --to=text
83
83
 
84
84
  # Convert DOCX to Markdown and save
85
- npx officeparser report.docx --format=md --output=report.md
85
+ npx officeparser report.docx --to=md --output=report.md
86
86
 
87
- # Convert PPTX to HTML
88
- npx officeparser presentation.pptx --format=html --output=preview.html
87
+ # Convert PPTX to HTML (using a bare flag for ocr)
88
+ npx officeparser presentation.pptx --to=html --output=preview.html --ocr
89
89
 
90
- # Convert XLSX to CSV
91
- npx officeparser data.xlsx --format=csv
90
+ # Convert XLSX to CSV with a custom delimiter
91
+ npx officeparser data.xlsx --to=csv --csvDelimiter=";"
92
92
 
93
93
  # Generate RAG chunks
94
- npx officeparser document.pdf --format=chunks
94
+ npx officeparser document.pdf --to=chunks
95
+
96
+ # Overriding file extension mapping
97
+ npx officeparser my_document --fileType=docx --to=json
95
98
  ```
96
99
 
100
+ ### CLI Syntax
101
+ - **Values:** Flags can be passed as `--flag=value` or `--flag value`.
102
+ - **Booleans:** Bare flags imply `true` (e.g. `--ocr` is equivalent to `--ocr=true`). Negation flags start with `no-` (e.g. `--no-ocr` is equivalent to `--ocr=false`).
103
+ - **Nested Objects:** You can pass nested properties directly using JSON dot-notation (e.g. `--ocrConfig.language=fra` or `--htmlConfig.containerWidth=900px`).
104
+
97
105
  ### CLI Options
98
106
 
99
107
  | Flag | Values | Default | Description |
100
108
  |------|--------|---------|-------------|
101
- | `--format` | `json\|text\|md\|html\|csv\|rtf\|pdf\|chunks` | `json` | Output format |
109
+ | `--to` | `json\|text\|md\|html\|csv\|rtf\|pdf\|chunks` | `json` | Output format |
102
110
  | `--output` | path | — | Write output to a file |
103
- | ~~`--toText`~~ | `true\|false` | `false` | **Deprecated.** Use `--format=text` |
104
- | `--ignoreNotes` | `true\|false` | `false` | Ignore speaker notes (PPTX/ODP) |
105
- | `--putNotesAtLast` | `true\|false` | `false` | Collect notes at end of output |
106
- | `--newlineDelimiter` | string | `\n` | Delimiter between lines |
107
- | `--extractAttachments` | `true\|false` | `false` | Extract images/charts as Base64 |
108
- | `--ocr` | `true\|false` | `false` | Enable OCR for images |
109
- | `--includeRawContent` | `true\|false` | `false` | Include raw XML/RTF in nodes |
110
- | `--includeBreakNodes` | `true\|false` | `false` | Include break nodes (DOCX only) |
111
- | ~~`--outputErrorToConsole`~~ | `true\|false` | `false` | **Deprecated.** Use `onWarning` callback |
112
- | `--verbose` | `true\|false` | `false` | Show full error stack traces |
111
+ | `--fileType` | `docx\|xlsx\|pptx\|odt\|odp\|ods\|pdf\|rtf\|csv\|md\|html` | — | Explicitly override input file type detection |
112
+ | `--ocr` | boolean | `false` | Enable OCR for images |
113
+ | `--extractAttachments` | boolean | `false` | Extract images/charts as Base64 |
114
+ | `--ignoreNotes` | boolean | `false` | Ignore footnotes/endnotes/speaker notes |
115
+ | `--ignoreComments` | boolean | `false` | Ignore inline comments |
116
+ | `--ignoreHeadersAndFooters` | boolean | `false` | Ignore headers and footers |
117
+ | `--ignoreSlideMasters` | boolean | `false` | Ignore slide masters |
118
+ | `--ignoreInternalLinks` | boolean | `false` | Ignore internal links |
119
+ | `--newlineDelimiter` | string | `\n` | Delimiter between lines/blocks in plaintext outputs |
120
+ | `--csvDelimiter` | string | `,` | Custom delimiter for CSV files |
121
+ | `--includeRawContent` | boolean | `false` | Include raw XML/RTF in nodes |
122
+ | `--serializeRawContent` | boolean | `true` | Include stringified XML in metadata |
123
+ | `--preserveXmlWhitespace` | boolean | `false` | Keep raw formatting space |
124
+ | `--includeBreakNodes` | boolean | `false` | Include break nodes (DOCX only) |
125
+ | `--verbose` | boolean | `false` | Show full error stack traces and warning logs |
126
+ | `--includeFormatting` | boolean | `true` | Include formatting style map matching |
127
+ | `--renderMetadata` | boolean | `false` | Render metadata as visible content in the generated output |
128
+ | `--htmlConfig.containerWidth` | string \| number | `auto` | HTML output container width (e.g. `900px`, `100%`) |
129
+ | ~~`--format`~~ | `json\|text\|md\|html\|csv\|rtf\|pdf\|chunks` | `json` | **Deprecated.** Use `--to` |
130
+ | ~~`--toText`~~ | `true\|false` | `false` | **Deprecated.** Use `--to=text` |
131
+ | ~~`--ocrLanguage`~~ | string | `eng` | **Deprecated.** Use `--ocrConfig.language` |
132
+ | ~~`--putNotesAtLast`~~ | `true\|false` | `false` | **Deprecated and ignored.** Notes are attached structurally to their nodes. |
133
+ | ~~`--outputErrorToConsole`~~ | `true\|false` | `false` | **Deprecated.** Use `--verbose` |
113
134
 
114
135
  ---
115
136
 
@@ -232,7 +253,7 @@ const ast = await officeParser.parseOffice('scanned_document.pdf', {
232
253
  > **Non-Fatal Timeout Recovery**
233
254
  > If `workerLoad` or `recognition` timeouts are exceeded, the parser will log a warning in `ast.warnings` and **continue parsing the rest of the document**. The overall promise resolves successfully with the text extracted from the document layers (rather than failing the entire parse).
234
255
 
235
- ### `ast.to()` — Generate from AST
256
+ ### `ast.to()`: Generate from AST
236
257
 
237
258
  The preferred way to convert a parsed AST to another format. Returns a `ConversionResult`.
238
259
 
@@ -246,7 +267,7 @@ const { value: chunks } = await ast.to('chunks', { strategy: 'fixed-
246
267
  const { value: pdfBytes } = await ast.to('pdf'); // Uint8Array
247
268
  ```
248
269
 
249
- ### `ast.toText()` — Quick Text Extraction
270
+ ### `ast.toText()`: Quick Text Extraction
250
271
 
251
272
  > [!NOTE]
252
273
  > `toText()` is **synchronous** and deprecated in favour of the async `ast.to('text')`.
@@ -295,7 +316,7 @@ const { value: csv } = await OfficeGenerator.generate(ast, 'csv');
295
316
 
296
317
  ---
297
318
 
298
- ## OfficeConverter — One-Step API
319
+ ## OfficeConverter: One-Step API
299
320
 
300
321
  `OfficeConverter.convert()` combines parsing and generation in a single call. It automatically syncs parser options from generator config (e.g., enables `extractAttachments` when images are requested).
301
322
 
@@ -326,7 +347,7 @@ const { value: html, messages } = await OfficeConverter.convert('data.xlsx', 'ht
326
347
 
327
348
  > [!IMPORTANT]
328
349
  > The `OfficeConverterConfig` shape uses **nested** `parseConfig` and `generatorConfig` sub-objects.
329
- > Do **not** put parser or generator options at the top level — only `onWarning` lives there.
350
+ > Do **not** put parser or generator options at the top level; only `onWarning` lives there.
330
351
 
331
352
  ---
332
353
 
@@ -344,7 +365,7 @@ const { value: chunks } = await OfficeConverter.convert('report.docx', 'chunks',
344
365
  strategy: 'document-structure',
345
366
  splitBy: 'heading', // 'paragraph' | 'heading' | 'page' | 'slide' | 'sheet'
346
367
  maxChunkSize: 1500,
347
- tableSplitStrategy: 'row', // repeats header row in every chunk — ideal for RAG
368
+ tableSplitStrategy: 'row', // repeats header row in every chunk, ideal for RAG
348
369
  }
349
370
  }
350
371
  });
@@ -445,7 +466,7 @@ OfficeParserAST
445
466
  └── ~~toText()~~ (Deprecated: use .to('text') instead)
446
467
  ```
447
468
 
448
- ### `OfficeIssue` — Warning / Error Object
469
+ ### `OfficeIssue`: Warning / Error Object
449
470
 
450
471
  All warnings and errors (from both parsing and generation) use this shape:
451
472
 
@@ -639,11 +660,11 @@ printNotes(ast.content);
639
660
  ```
640
661
 
641
662
  > [!IMPORTANT]
642
- > `putNotesAtLast` is **deprecated**. Notes are always attached via `node.notes` — this flag has no effect and will be removed in a future major version.
663
+ > `putNotesAtLast` is **deprecated**. Notes are always attached via `node.notes`; this flag has no effect and will be removed in a future major version.
643
664
 
644
665
  ### Access headers, footers & slide masters
645
666
  ```ts
646
- // These are NOT in ast.content — use ast.auxiliary
667
+ // These are NOT in ast.content; use ast.auxiliary
647
668
  console.log(ast.auxiliary?.headers?.map(h => h.text)); // DOCX headers
648
669
  console.log(ast.auxiliary?.footers?.map(f => f.text)); // DOCX footers
649
670
  console.log(ast.auxiliary?.slideMasters?.length); // PPTX slide masters
@@ -716,13 +737,13 @@ Pass as the second argument to `parseOffice(file, config)`.
716
737
  |--------|------|---------|-------------|
717
738
  | `newlineDelimiter` | `string` | `'\n'` | Delimiter inserted between lines in text output |
718
739
  | `ignoreNotes` | `boolean` | `false` | Ignore footnotes/endnotes (DOCX, RTF) and speaker notes (PPTX/ODP) |
719
- | `ignoreComments` | `boolean` | `false` | **New**: Ignore inline comments/annotations (DOCX, XLSX, PPTX) — by default attached via `node.comments[]` |
740
+ | `ignoreComments` | `boolean` | `false` | **New**: Ignore inline comments/annotations (DOCX, XLSX, PPTX), attached by default via `node.comments[]` |
720
741
  | `ignoreHeadersAndFooters` | `boolean` | `false` | **New**: Skip DOCX headers & footers (populated in `ast.auxiliary.headers/footers` by default) |
721
742
  | `ignoreSlideMasters` | `boolean` | `false` | **New**: Skip PPTX slide masters (populated in `ast.auxiliary.slideMasters` by default) |
722
743
  | ~~`putNotesAtLast`~~ | `boolean` | `false` | **Deprecated**: Notes are now attached via `node.notes[]`. This flag has no effect |
723
744
  | `extractAttachments` | `boolean` | `false` | Populate `ast.attachments` with Base64 images/charts |
724
745
  | `ocr` | `boolean` | `false` | Run Tesseract OCR on images (requires `extractAttachments: true`) |
725
- | `ocrConfig` | `OcrConfig` | `{}` | OCR worker pool settings — see [OCR section](#ocr-scheduler--resource-management) |
746
+ | `ocrConfig` | `OcrConfig` | `{}` | OCR worker pool settings (see [OCR section](#ocr-scheduler--resource-management)) |
726
747
  | `includeRawContent` | `boolean` | `false` | Attach raw XML/RTF source to each node |
727
748
  | `serializeRawContent` | `boolean` | `true` | Re-serialize XML to clean strings (only if `includeRawContent: true`) |
728
749
  | `preserveXmlWhitespace` | `boolean` | `false` | Preserve original XML whitespace during serialization |
@@ -730,6 +751,7 @@ Pass as the second argument to `parseOffice(file, config)`.
730
751
  | `ignoreInternalLinks` | `boolean` | `false` | Strip bookmarks and internal cross-references from AST |
731
752
  | `fileType` | `SupportedFileType \| null` | `null` | **Required for text-based binary data** (`'md'`, `'html'`, `'csv'`) as these lack magic bytes. |
732
753
  | `csvDelimiter` | `string` | `','` | Input delimiter when parsing CSV files |
754
+ | `decompressionLimits` | `DecompressionLimits` | `{ maxUncompressedBytes: 512MB, maxZipEntries: 10000 }` | **New**: Limits applied during ZIP extraction to protect against excessive memory and resource usage |
733
755
  | `pdfWorkerSrc` | `string` | CDN (jsDelivr) | Path/URL to `pdf.worker.min.mjs` (required in browser) |
734
756
  | `onWarning` | `(issue: OfficeIssue) => void` | — | Callback for non-fatal parsing issues |
735
757
  | `abortSignal` | `AbortSignal \| null` | `null` | Optional signal to cancel parsing (rejects with AbortError) |
@@ -757,7 +779,7 @@ Options shared by all generator formats. Pass to `OfficeGenerator.generate(ast,
757
779
 
758
780
  ---
759
781
 
760
- ### `onNode` Callback — Advanced Node Manipulation
782
+ ### `onNode` Callback: Advanced Node Manipulation
761
783
 
762
784
  Called for **every node** in the AST during generation. Can be `async`.
763
785
 
@@ -788,7 +810,7 @@ const { value: md } = await ast.to('md', {
788
810
 
789
811
  ---
790
812
 
791
- ### `styleMap` — Semantic Style Mapping
813
+ ### `styleMap`: Semantic Style Mapping
792
814
 
793
815
  Maps document style names to semantic output elements. Two formats supported:
794
816
 
@@ -833,7 +855,7 @@ Pass as `htmlConfig` inside `GeneratorConfig`.
833
855
  | `standalone` | `boolean` | `true` | Wrap output in a full `<html>` document with CSS |
834
856
  | `chartJsSrc` | `string` | jsDelivr CDN | URL for the Chart.js library |
835
857
  | `containerWidth` | `string \| number` | `'auto'` | Max width of the content container. Positive number (px), CSS length string (`'900px'`, `'100%'`, `'60vw'`), or `'auto'`. Invalid values fall back to `'auto'` with an `INVALID_CONTAINER_WIDTH` warning |
836
- | `customCss` | `string` | `''` | Raw CSS injected into the `<style>` block — use to override built-in styles |
858
+ | `customCss` | `string` | `''` | Raw CSS injected into the `<style>` block; use this to override built-in styles |
837
859
  | `injections.headStart` | `string` | `''` | Raw HTML injected after `<head>` |
838
860
  | `injections.headEnd` | `string` | `''` | Raw HTML injected before `</head>` |
839
861
  | `injections.bodyStart` | `string` | `''` | Raw HTML injected after `<body>` |
@@ -883,7 +905,7 @@ Pass as `textConfig` inside `GeneratorConfig`.
883
905
  | Option | Type | Default | Description |
884
906
  |--------|------|---------|-------------|
885
907
  | `newlineDelimiter` | `string` | `'\n'` | String inserted between structural blocks |
886
- | `preserveLayout` | `boolean` | `false` | Render tables with aligned columns using whitespace |
908
+ | `preserveLayout` | `boolean` | `true` | Render tables with aligned columns using whitespace |
887
909
 
888
910
  ---
889
911
 
@@ -901,7 +923,7 @@ Configuration for `OfficeConverter.convert(file, format, config)`.
901
923
 
902
924
  ### ChunkingConfig
903
925
 
904
- `ChunkingConfig` is a **discriminated union** — the available options depend on the `strategy` field.
926
+ `ChunkingConfig` is a **discriminated union**: the available options depend on the `strategy` field.
905
927
 
906
928
  #### Common Options (all strategies)
907
929
 
@@ -989,8 +1011,8 @@ Two bundles are available in the `dist/` directory:
989
1011
 
990
1012
  | Bundle | Usage |
991
1013
  |--------|-------|
992
- | `officeparser.browser.mjs` | ESM — use with `import` statements or modern bundlers (Vite, Webpack, Next.js) |
993
- | `officeparser.browser.iife.js` | IIFE — use with a `<script>` tag; exposes the global `officeParser` object |
1014
+ | `officeparser.browser.mjs` | ESM, use with `import` statements or modern bundlers (Vite, Webpack, Next.js) |
1015
+ | `officeparser.browser.iife.js` | IIFE, use with a `<script>` tag; exposes the global `officeParser` object |
994
1016
 
995
1017
  ### ESM (Vite / Webpack / Next.js)
996
1018
 
@@ -1050,7 +1072,7 @@ const ast = await officeParser.parseOffice(pdfArrayBuffer, {
1050
1072
  | `"Worker not found"` in browser for PDF | Verify `pdfWorkerSrc` points to `pdf.worker.min.mjs` matching version `5.6.205` |
1051
1073
  | Low OCR accuracy | Verify `ocrConfig.language` matches the document language; quality depends on image resolution |
1052
1074
  | Out of memory on large Excel files | Call `ast.toText()` early and discard the AST object to allow garbage collection |
1053
- | `md`/`html`/`csv` buffer not detected | Add `fileType: 'md'` (or `'html'`, `'csv'`) to config — these formats have no magic bytes |
1075
+ | `md`/`html`/`csv` buffer not detected | Add `fileType: 'md'` (or `'html'`, `'csv'`) to config (these formats have no magic bytes) |
1054
1076
  | `IMPROPER_BUFFERS` error | Usually means no file extension and no `fileType` hint was provided for a buffer input |
1055
1077
  | PDF generation fails | Install the optional peer dependency: `npm install puppeteer` |
1056
1078
 
@@ -1086,4 +1108,4 @@ Contributions are welcome! Please see [CONTRIBUTING.md](CONTRIBUTING.md) for det
1086
1108
 
1087
1109
  ## License
1088
1110
 
1089
- This project is licensed under the MIT License — see the [LICENSE](LICENSE) file for details.
1111
+ This project is licensed under the MIT License; see the [LICENSE](LICENSE) file for details.
@@ -1,8 +1,12 @@
1
- import { ConversionResult, GeneratorConfig, OfficeParserAST, SupportedDestination, SupportedFileType } from './types.js';
1
+ import { ConversionResult, GeneratorConfig, OfficeParserAST, SupportedDestination, SupportedFileType, UniversalGeneratorFormat } from './types.js';
2
2
  /**
3
3
  * Main generator class providing document conversion functionality.
4
4
  */
5
5
  export declare class OfficeGenerator {
6
+ /**
7
+ * Normalizes format aliases (e.g., 'txt' to 'text', 'markdown' to 'md') to standard internal formats.
8
+ */
9
+ static normalizeDestination(dest: string): UniversalGeneratorFormat;
6
10
  /**
7
11
  * Generates a file of the specified type from an AST.
8
12
  * This is the single source of truth for generation logic.
@@ -14,6 +14,17 @@ const errorUtils_js_1 = require("./utils/errorUtils.js");
14
14
  * Main generator class providing document conversion functionality.
15
15
  */
16
16
  class OfficeGenerator {
17
+ /**
18
+ * Normalizes format aliases (e.g., 'txt' to 'text', 'markdown' to 'md') to standard internal formats.
19
+ */
20
+ static normalizeDestination(dest) {
21
+ const d = dest?.toLowerCase();
22
+ if (d === 'txt')
23
+ return 'text';
24
+ if (d === 'markdown')
25
+ return 'md';
26
+ return d;
27
+ }
17
28
  /**
18
29
  * Generates a file of the specified type from an AST.
19
30
  * This is the single source of truth for generation logic.
@@ -26,7 +37,8 @@ class OfficeGenerator {
26
37
  */
27
38
  static async generate(ast, destination, config) {
28
39
  let generator;
29
- switch (destination.toLowerCase()) {
40
+ const normalizedDestination = OfficeGenerator.normalizeDestination(destination);
41
+ switch (normalizedDestination) {
30
42
  case 'text':
31
43
  generator = new TextGenerator_js_1.TextGenerator(ast, config);
32
44
  break;
@@ -49,7 +61,7 @@ class OfficeGenerator {
49
61
  generator = new ChunkingGenerator_js_1.ChunkingGenerator(ast, config);
50
62
  break;
51
63
  default:
52
- throw (0, errorUtils_js_1.getOfficeError)(types_js_1.OfficeErrorType.EXTENSION_UNSUPPORTED, undefined, destination);
64
+ throw (0, errorUtils_js_1.getOfficeError)(types_js_1.OfficeErrorType.FORMAT_UNSUPPORTED, undefined, destination);
53
65
  }
54
66
  return generator.generate();
55
67
  }
@@ -81,7 +81,7 @@ export declare class OfficeParser {
81
81
  * });
82
82
  *
83
83
  * // Parse a Buffer with OCR enabled
84
- * const buffer = await fetch('document.pdf').then(r => r.arrayBuffer());
84
+ * const buffer = await retrieveData('document.pdf').then(r => r.arrayBuffer());
85
85
  * const ast = await OfficeParser.parseOffice(buffer, {
86
86
  * ocr: true,
87
87
  * ocrLanguage: 'eng+fra'
@@ -98,7 +98,7 @@ class OfficeParser {
98
98
  * });
99
99
  *
100
100
  * // Parse a Buffer with OCR enabled
101
- * const buffer = await fetch('document.pdf').then(r => r.arrayBuffer());
101
+ * const buffer = await retrieveData('document.pdf').then(r => r.arrayBuffer());
102
102
  * const ast = await OfficeParser.parseOffice(buffer, {
103
103
  * ocr: true,
104
104
  * ocrLanguage: 'eng+fra'
package/dist/cli.d.ts CHANGED
@@ -4,23 +4,25 @@
4
4
  *
5
5
  * Allows running officeparser from the command line:
6
6
  * npx officeparser file.docx
7
- * officeparser file.docx --toText=true
8
- * officeparser file.docx --ocr=true --extractAttachments=true
7
+ * officeparser file.docx --to=text
8
+ * officeparser file.docx --ocr --extractAttachments
9
9
  *
10
- * Options (--key=value):
11
- * --format=json|text|md|html|csv|rtf|pdf|chunks Convert AST to specified format
10
+ * Options (--key=value, --key value, or bare flags):
11
+ * --to=json|text|md|html|csv|rtf|pdf|chunks Convert AST to specified format (default: json)
12
12
  * --output=path Save result to a file
13
- * --toText=true Legacy flag for plain text output
14
- * --ocr=true Enable OCR for images
15
- * --ocrLanguage=eng OCR language (default: eng)
16
- * --extractAttachments=true Extract embedded attachments
17
- * --ignoreNotes=true Ignore footnotes/endnotes
18
- * --ignoreComments=true Ignore inline comments
19
- * --ignoreHeadersAndFooters=true Ignore headers and footers
20
- * --ignoreSlideMasters=true Ignore slide masters
21
- * --ignoreInternalLinks=true Ignore internal links
22
- * --putNotesAtLast=true Move notes to end of document
23
- * --includeRawContent=true Include raw content in AST
24
- * --outputErrorToConsole=true Log errors to console
13
+ * --fileType=docx|xlsx|... Override file type detection
14
+ * --ocr Enable OCR for images (default: false)
15
+ * --ocrConfig.language=eng OCR language (default: eng)
16
+ * --extractAttachments Extract embedded attachments (default: false)
17
+ * --ignoreNotes Ignore footnotes/endnotes/speaker notes (default: false)
18
+ * --ignoreComments Ignore inline comments (default: false)
19
+ * --ignoreHeadersAndFooters Ignore headers and footers (default: false)
20
+ * --ignoreSlideMasters Ignore slide masters (default: false)
21
+ * --ignoreInternalLinks Ignore internal links (default: false)
22
+ * --includeRawContent Include raw content in AST (default: false)
23
+ * --serializeRawContent Include stringified XML in metadata (default: true)
24
+ * --preserveXmlWhitespace Keep raw formatting space (default: false)
25
+ * --includeBreakNodes Include break nodes (DOCX only, default: false)
26
+ * --verbose Show full error stack traces and warning logs
25
27
  */
26
28
  export {};