officeparser 4.1.2 → 5.0.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -13,6 +13,9 @@ A Node.js library to parse text out of any office file.
13
13
 
14
14
 
15
15
  #### Update
16
+ * 2024/10/21 - Replaced extracting zip files from decompress to yauzl. This means that we now extract files in memory and we no longer need to write them to disk. Removed config flags related to extracted files. Added flags for CLI execution.
17
+ * 2024/10/15 - Fixed erroring out while deleting temp files when multiple worker threads make parallel executions resulting in same file name for multiple files. Fixed erroring out when multiple executions are made without waiting for the previous execution to finish which resulted in deleting the file from other execution. Upgraded dependencies.
18
+ * 2024/10/13 - Fixed parsing text from xlsx files which contain no shared strings file and files which have inlineStr based strings.
16
19
  * 2024/05/06 - Replaced pdf parsing support from pdf-parse library to natively building it using pdf.js library from Mozilla by analyzing its output. Added pdfjs-dist build as a local library.
17
20
  * 2023/11/25 - Fixed error catching when an error occurs within the parsing of a file, especially after decompressing it. Also fixed the problem with parallel parsing of files as we were using only timestamp in file names.
18
21
  * 2023/10/24 - Revamped content parsing code. Fixed order of content in files, especially in word files where table information would always land up at the end of the text. Added config object as argument for parseOffice which can be used to set new line delimiter and multiple other configurations. Added support for parsing pdf files using the popular npm library pdf-parse. Removed support for individual file parsing functions.
@@ -35,7 +38,6 @@ A Node.js library to parse text out of any office file.
35
38
 
36
39
  ## Install via npm
37
40
 
38
-
39
41
  ```
40
42
  npm i officeparser
41
43
  ```
@@ -43,14 +45,20 @@ npm i officeparser
43
45
  ## Command Line usage
44
46
  If you want to call the installed officeParser.js file, use below command
45
47
  ```
46
- node </path/to/officeParser.js> <fileName>
48
+ node <path/to/officeParser.js> [--configOption=value] [FILE_PATH]
49
+ node officeparser [--configOption=value] [FILE_PATH]
47
50
  ```
48
51
 
49
- Otherwise, you can simply use npx to instantly extract parsed data.
52
+ Otherwise, you can simply use npx without installing the node module to instantly extract parsed data.
50
53
  ```
51
- npx officeparser <fileName>
54
+ npx officeparser [--configOption=value] [FILE_PATH]
52
55
  ```
53
56
 
57
+ ### Config Options:
58
+ - `--ignoreNotes=[true|false]` Flag to ignore notes from files like PowerPoint. Default is false.
59
+ - `--newlineDelimiter=[delimiter]` The delimiter to use for new lines. Default is `\n`.
60
+ - `--putNotesAtLast=[true|false]` Flag to collect notes at the end of files like PowerPoint. Default is false.
61
+ - `--outputErrorToConsole=[true|false]` Flag to output errors to the console. Default is false.
54
62
 
55
63
  ## Library Usage
56
64
  ```js
@@ -99,8 +107,6 @@ officeParser.parseOfficeAsync(fileBuffers);
99
107
  *Optionally add a config object as 3rd variable to parseOffice for the following configurations*
100
108
  | Flag | DataType | Default | Explanation |
101
109
  |----------------------|----------|------------------|-------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|
102
- | tempFilesLocation | string | officeParserTemp | The directory where officeparser stores the temp files . The final decompressed data will be put inside officeParserTemp folder within your directory. **Please ensure that this directory actually exists.** Default is officeParserTemp. |
103
- | preserveTempFiles | boolean | false | Flag to not delete the internal content files and the possible duplicate temp files that it uses after unzipping office files. Default is false. It always deletes all of those files. |
104
110
  | outputErrorToConsole | boolean | false | Flag to show all the logs to console in case of an error. Default is false. |
105
111
  | newlineDelimiter | string | \n | The delimiter used for every new line in places that allow multiline text like word. Default is \n. |
106
112
  | ignoreNotes | boolean | false | Flag to ignore notes from parsing in files like powerpoint. Default is false. It includes notes in the parsed text by default. |
@@ -155,7 +161,7 @@ function searchForTermInOfficeFile(searchterm, filepath) {
155
161
 
156
162
  **Example - TypeScript**
157
163
  ```ts
158
- const officeParser = require('officeparser');
164
+ import { OfficeParserConfig, parseOfficeAsync } from 'officeparser';
159
165
 
160
166
  const config: OfficeParserConfig = {
161
167
  newlineDelimiter: " ", // Separate new lines with a space instead of the default \n.
@@ -163,7 +169,7 @@ const config: OfficeParserConfig = {
163
169
  }
164
170
 
165
171
  // relative path is also fine => eg: files/myWorkSheet.ods
166
- officeParser.parseOfficeAsync("/Users/harsh/Desktop/files/mySlides.pptx", config);
172
+ parseOfficeAsync("/Users/harsh/Desktop/files/mySlides.pptx", config);
167
173
  .then(data => {
168
174
  const newText = data + " look, I can parse a powerpoint file";
169
175
  callSomeOtherFunction(newText);
@@ -172,7 +178,7 @@ officeParser.parseOfficeAsync("/Users/harsh/Desktop/files/mySlides.pptx", config
172
178
 
173
179
  // Search for a term in the parsed text.
174
180
  function searchForTermInOfficeFile(searchterm: string, filepath: string): Promise<boolean> {
175
- return officeParser.parseOfficeAsync(filepath)
181
+ return parseOfficeAsync(filepath)
176
182
  .then(data => data.indexOf(searchterm) != -1)
177
183
  }
178
184
  ```