officeparser 3.2.2 → 4.0.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -9,9 +9,12 @@ A Node.js library to parse text out of any office file.
9
9
  - [`odt`](https://en.wikipedia.org/wiki/OpenDocument)
10
10
  - [`odp`](https://en.wikipedia.org/wiki/OpenDocument)
11
11
  - [`ods`](https://en.wikipedia.org/wiki/OpenDocument)
12
+ - [`pdf`](https://en.wikipedia.org/wiki/PDF)
12
13
 
13
14
 
14
15
  #### Update
16
+ * 2023/10/24 - Revamped content parsing code. Fixed order of content in files, especially in word files where table information would always land up at the end of the text. Added config object as argument for parseOffice which can be used to set new line delimiter and multiple other configurations. Added support for parsing pdf files using the popular npm library pdf-parse. Removed support for individual file parsing functions.
17
+ * 2023/04/26 - Added support for file buffers as argument for filepath for parseOffice and parseOfficeAsync
15
18
  * 2023/04/07 - Added typings to methods to help with Typescript projects.
16
19
  * 2022/12/28 - Added command line method to use officeParser with or without installing it and instantly get parsed content on the console.
17
20
  * 2022/12/10 - Fixed memory leak issues, bugs related to parsing open document files and improved error handling.
@@ -46,28 +49,26 @@ Otherwise, you can simply use npx to instantly extract parsed data.
46
49
  npx officeparser <fileName>
47
50
  ```
48
51
 
49
- ----------
50
52
 
51
- **Library Usage**
53
+ ## Library Usage
52
54
  ```js
53
55
  const officeParser = require('officeparser');
54
56
 
55
57
  // callback
56
- officeParser.parseOffice("/path/to/officeFile", function(data, err){
58
+ officeParser.parseOffice("/path/to/officeFile", function(data, err) {
57
59
  // "data" string in the callback here is the text parsed from the office file passed in the first argument above
58
- if (err) return console.log(err);
59
- console.log(data)
60
+ if (err) {
61
+ console.log(err);
62
+ return;
63
+ }
64
+ console.log(data);
60
65
  })
61
66
 
62
67
  // promise
63
68
  officeParser.parseOfficeAsync("/path/to/officeFile");
64
69
  // "data" string in the promise here is the text parsed from the office file passed in the argument above
65
- .then((data) => {
66
- console.log(data)
67
- })
68
- .catch(err) => {
69
- console.log(err)
70
- }
70
+ .then(data => console.log(data))
71
+ .catch(err => console.error(err))
71
72
 
72
73
  // async/await
73
74
  try {
@@ -78,260 +79,135 @@ try {
78
79
  // resolve error
79
80
  console.log(err);
80
81
  }
81
- ```
82
-
83
- **Please take note: I have breached convention in placing err as second argument in my callback but please understand that I had to do it to not break other people's existing modules.**
84
82
 
85
- *Optionally change decompression location for office Files at personalised locations for environments with restricted write access*
86
-
87
- ```js
88
- const officeParser = require('officeparser');
89
-
90
- // Default decompress location for office Files is "officeDist" in the directory where Node is started.
91
- // Put this file before parseOffice method to take effect.
92
- officeParser.setDecompressionLocation("/tmp"); // New decompression location would be "/tmp/officeDist"
93
-
94
- // P.S.: Setting location on a Windows environment with '\' hierarchy requires to be entered twice '\\'
95
- officeParser.setDecompressionLocation("C:\\tmp"); // New decompression location would be "C:\tmp\officeDist"
96
-
97
-
98
- officeParser.parseOffice("/path/to/officeFile", function(data, err){
99
- // "data" string in the callback here is the text parsed from the office file passed in the first argument above
100
- if (err) return console.log(err);
101
- console.log(data)
102
- })
83
+ // USING FILE BUFFERS
84
+ // instead of file path, you can also pass file buffers of one of the supported files
85
+ // on parseOffice or parseOfficeAsync functions.
86
+
87
+ // get file buffers
88
+ const fileBuffers = fs.readFileSync("/path/to/officeFile");
89
+ // get parsed text from officeParser
90
+ // NOTE: Only works with parseOffice. Old functions are not supported.
91
+ officeParser.parseOfficeAsync(fileBuffers);
92
+ .then(data => console.log(data))
93
+ .catch(err => console.error(err))
103
94
  ```
104
95
 
105
- *Optionally add false as 3rd variable to parseOffice to not delete the generated officeDist folder*
96
+ ### Configuration Object: OfficeParserConfig
97
+ *Optionally add a config object as 3rd variable to parseOffice for the following configurations*
98
+ | flag | datatype | explanation |
99
+ |----------------------|----------|-------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|
100
+ | preserveTempFiles | boolean | Flag to not delete the internal content files and the possible duplicate temp files that it uses after unzipping office files. Default is false. It always deletes all of those files. |
101
+ | outputErrorToConsole | boolean | Flag to show all the logs to console in case of an error. |
102
+ | newlineDelimiter | string | The delimiter used for every new line in places that allow multiline text like word. Default is \n. |
103
+ | ignoreNotes | boolean | Flag to ignore notes from parsing in files like powerpoint. Default is false. It includes notes in the parsed text by default. |
104
+ | putNotesAtLast | boolean | Flag, if set to true, will collectively put all the parsed text from notes at last in files like powerpoint. Default is false. It puts each notes right after its main slide content. If ignoreNotes is set to true, this flag is also ignored. |
106
105
 
107
106
  ```js
107
+ const config = {
108
+ newlineDelimiter: " ", // Separate new lines with a space instead of the default \n.
109
+ ignoreNotes: true // Ignore notes while parsing presentation files like pptx or odp.
110
+ }
111
+
108
112
  // callback
109
113
  officeParser.parseOffice("/path/to/officeFile", function(data, err){
110
- if (err) return console.log(err);
111
- console.log(data)
112
- }, false)
114
+ if (err) {
115
+ console.log(err);
116
+ return;
117
+ }
118
+ console.log(data);
119
+ }, config)
113
120
 
114
121
  // promise
115
- officeParser.parseOfficeAsync("/path/to/officeFile", false);
116
- .then((data) => {
117
- console.log(data)
118
- })
119
- .catch(err) => {
120
- console.log(err)
121
- }
122
-
123
- // async/await
124
- try {
125
- const data = await officeParser.parseOfficeAsync("/path/to/officeFile", false);
126
- console.log(data);
127
- } catch (err) {
128
- // resolve error
129
- console.log(err);
130
- }
122
+ officeParser.parseOfficeAsync("/path/to/officeFile", config);
123
+ .then((data) => console.log(data))
124
+ .catch((err) => console.error(err))
131
125
  ```
132
126
 
133
- **Example**
127
+ **Example - JavaScript**
134
128
  ```js
135
129
  const officeParser = require('officeparser');
136
130
 
137
- // callback
138
- officeParser.parseOffice("C:\\files\\myText.docx", function(data, err){
139
- if (err) return console.log(err);
140
- var newText = data + "look, I can parse a word file"
141
- callSomeOtherFunction(newText);
142
- })
143
-
144
-
145
- // promise
146
- officeParser.parseOfficeAsync("/Users/harsh/Desktop/files/mySlides.pptx");
147
- .then((data) => {
148
- var newText = data + "look, I can parse a powerpoint file"
149
- callSomeOtherFunction(newText);
150
- })
151
- .catch(err) => {
152
- console.log(err)
131
+ const config = {
132
+ newlineDelimiter: " ", // Separate new lines with a space instead of the default \n.
133
+ ignoreNotes: true // Ignore notes while parsing presentation files like pptx or odp.
153
134
  }
154
135
 
155
- // Using relative path for file is also fine
156
- officeParser.parseOffice("files/myWorkSheet.ods", function(data, err){
157
- if (err) return console.log(err);
158
- var newText = data + "look, I can parse an excel file"
159
- callSomeOtherFunction(newText);
160
- })
161
-
162
- // async/await
163
- try {
164
- const data = await officeParser.parseOfficeAsync("/Users/harsh/Desktop/files/mySlides.pptx");
165
- let newText = data + "look, I can parse a powerpoint file";
166
- await callSomeOtherFunction(newText);
167
- } catch (err) {
168
- // resolve error
169
- console.log(err);
136
+ // relative path is also fine => eg: files/myWorkSheet.ods
137
+ officeParser.parseOfficeAsync("/Users/harsh/Desktop/files/mySlides.pptx", config);
138
+ .then(data => {
139
+ const newText = data + " look, I can parse a powerpoint file";
140
+ callSomeOtherFunction(newText);
141
+ })
142
+ .catch((err) => console.error(err));
143
+
144
+ // Search for a term in the parsed text.
145
+ function searchForTermInOfficeFile(searchterm, filepath) {
146
+ return officeParser.parseOfficeAsync(filepath)
147
+ .then(data => data.indexOf(searchterm) != -1)
170
148
  }
171
149
  ```
172
150
 
173
151
 
174
- ----------
175
-
176
- ### Old but functional way of extracting text from word, powerpoint and excel files
177
- *These were the initial methods of parsing text till parseOffice method came into existence. These still exist and form the skeleton to this module as parseOffice redirects the below functions anyway. These functions will forever remain available to guarantee long-term usage of this module. I will ensure backward-compatibility with all previous versions.*
178
-
179
- **Usage**
180
- ```js
152
+ **Example - TypeScript**
153
+ ```ts
181
154
  const officeParser = require('officeparser');
182
155
 
183
- // callback
184
- officeParser.parseWord("/path/to/word.docx", function(data, err){
185
- // "data" string in the callback here is the text parsed from the word file passed in the first argument above
186
- if (err) return console.log(err);
187
- console.log(data)
188
- })
189
-
190
- officeParser.parsePowerPoint("/path/to/powerpoint.pptx", function(data, err){
191
- // "data" string in the callback here is the text parsed from the powerpoint file passed in the first argument above
192
- if (err) return console.log(err);
193
- console.log(data)
194
- })
195
-
196
- officeParser.parseExcel("/path/to/excel.xlsx", function(data, err){
197
- // "data" string in the callback here is the text parsed from the excel file passed in the first argument above
198
- if (err) return console.log(err);
199
- console.log(data)
200
- })
201
-
202
- officeParser.parseOpenOffice("/path/to/writer.odt", function(data, err){
203
- // "data" string in the callback here is the text parsed from the writer file passed in the first argument above
204
- if (err) return console.log(err);
205
- console.log(data)
206
- })
207
-
208
- // promise
209
- officeParser.parseWordAsync("/path/to/word.docx");
210
- .then((data) => {
211
- // data is the parsed text
212
- })
213
- officeParser.parsePowerPointAsync("/path/to/powerpoint.pptx");
214
- .then((data) => {
215
- // data is the parsed text
216
- })
217
- officeParser.parseExcelAsync("/path/to/excel.xlsx");
218
- .then((data) => {
219
- // data is the parsed text
220
- })
221
- officeParser.parseOpenOfficeAsync("/path/to/writer.odt");
222
- .then((data) => {
223
- // data is the parsed text
224
- })
225
-
226
- // async/await
227
- try {
228
- // "data" string returned from promise here is the text parsed from the office file passed in the first argument
229
- const data1 = await officeParser.parseWordAsync("/path/to/word.docx");
230
-
231
- const data2 = await officeParser.parsePowerPointAsync("/path/to/powerpoint.pptx");
232
-
233
- const data3 = await officeParser.parseExcelAsync("/path/to/excel.xlsx");
156
+ const config: OfficeParserConfig = {
157
+ newlineDelimiter: " ", // Separate new lines with a space instead of the default \n.
158
+ ignoreNotes: true // Ignore notes while parsing presentation files like pptx or odp.
159
+ }
234
160
 
235
- const data3 = await officeParser.parseOpenOfficeAsync("/path/to/writer.odt");
236
- } catch (err) {
237
- // resolve error
238
- console.log(err);
161
+ // relative path is also fine => eg: files/myWorkSheet.ods
162
+ officeParser.parseOfficeAsync("/Users/harsh/Desktop/files/mySlides.pptx", config);
163
+ .then(data => {
164
+ const newText = data + " look, I can parse a powerpoint file";
165
+ callSomeOtherFunction(newText);
166
+ })
167
+ .catch((err) => console.error(err));
168
+
169
+ // Search for a term in the parsed text.
170
+ function searchForTermInOfficeFile(searchterm, filepath): Promise<boolean> {
171
+ return officeParser.parseOfficeAsync(filepath)
172
+ .then(data => data.indexOf(searchterm) != -1)
239
173
  }
240
174
  ```
241
175
 
242
- **Example**
243
- ```js
244
- const officeParser = require('officeparser');
245
-
246
- // callback
247
- officeParser.parseWord("C:\\files\\myText.docx", function(data, err){
248
- if (err) return console.log(err);
249
- var newText = data + "look, I can parse a word file"
250
- callSomeOtherFunction(newText);
251
- })
252
176
 
253
- officeParser.parsePowerPoint("/Users/harsh/Desktop/files/mySlides.pptx", function(data, err){
254
- if (err) return console.log(err);
255
- var newText = data + "look, I can parse a powerpoint file"
256
- callSomeOtherFunction(newText);
257
- })
177
+ \
178
+ \
179
+ **Please take note: I have breached convention in placing err as second argument in my callback but please understand that I had to do it to not break other people's existing modules.**
258
180
 
259
- // Using relative path for file is also fine
260
- officeParser.parseExcel("files/myWorkSheet.xlsx", function(data, err){
261
- if (err) return console.log(err);
262
- var newText = data + "look, I can parse an excel file"
263
- callSomeOtherFunction(newText);
264
- })
181
+ *Optionally change decompression location for office Files at personalised locations for environments with restricted write access*
265
182
 
266
- officeParser.parseOpenOffice("files/myDocument.odt", function(data, err){
267
- if (err) return console.log(err);
268
- var newText = data + "look, I can parse an OpenOffice file"
269
- callSomeOtherFunction(newText);
270
- })
183
+ ```js
184
+ const officeParser = require('officeparser');
271
185
 
272
- // promise
273
- officeParser.parseWordAsync("C:\\files\\myText.docx");
274
- .then((data) => {
275
- let newText1 = data1 + "look, I can parse a word file";
276
- callSomeOtherFunction(newText1);
277
- })
278
- .catch(err) => {
279
- console.log(err)
280
- }
186
+ // Default decompress location for office Files is "officeDist" in the directory where Node is started.
187
+ // Put this file before parseOffice method to take effect.
188
+ officeParser.setDecompressionLocation("/tmp"); // New decompression location would be "/tmp/officeDist"
281
189
 
282
- officeParser.parsePowerPointAsync("/Users/harsh/Desktop/files/mySlides.pptx");
283
- .then((data) => {
284
- let newText2 = data2 + "look, I can parse a powerpoint file";
285
- callSomeOtherFunction(newText2);
286
- })
287
- .catch(err) => {
288
- console.log(err)
289
- }
190
+ // P.S.: Setting location on a Windows environment with '\' hierarchy requires to be entered twice '\\'
191
+ officeParser.setDecompressionLocation("C:\\tmp"); // New decompression location would be "C:\tmp\officeDist"
290
192
 
291
- officeParser.parseExcelAsync("files/myWorkSheet.xlsx");
292
- .then((data) => {
293
- let newText3 = data3 + "look, I can parse an excel file";
294
- callSomeOtherFunction(newText3);
295
- })
296
- .catch(err) => {
297
- console.log(err)
298
- }
299
193
 
300
- officeParser.parseOpenOfficeAsync("files/myDocument.odt");
301
- .then((data) => {
302
- let newText4 = data4 + "look, I can parse an OpenOffice file";
303
- callSomeOtherFunction(newText4);
194
+ officeParser.parseOffice("/path/to/officeFile", function(data, err){
195
+ // "data" string in the callback here is the text parsed from the office file passed in the first argument above
196
+ if (err) {
197
+ console.log(err);
198
+ return;
199
+ }
200
+ console.log(data);
304
201
  })
305
- .catch(err) => {
306
- console.log(err)
307
- }
308
-
309
-
310
- // async/await
311
- try {
312
- const data1 = await officeParser.parseWordAsync("C:\\files\\myText.docx");
313
- let newText1 = data1 + "look, I can parse a word file";
314
- await callSomeOtherFunction(newText1);
315
-
316
- const data2 = await officeParser.parsePowerPointAsync("/Users/harsh/Desktop/files/mySlides.pptx");
317
- let newText2 = data2 + "look, I can parse a powerpoint file";
318
- await callSomeOtherFunction(newText2);
319
-
320
- // Using relative path for file is also fine
321
- const data3 = await officeParser.parseExcelAsync("files/myWorkSheet.xlsx");
322
- let newText3 = data3 + "look, I can parse an excel file";
323
- await callSomeOtherFunction(newText3);
324
-
325
- const data4 = await officeParser.parseOpenOfficeAsync("files/myDocument.odt");
326
- let newText4 = data4 + "look, I can parse an OpenOffice file";
327
- await callSomeOtherFunction(newText4);
328
- } catch (err) {
329
- // resolve error
330
- console.log(err);
331
- }
332
202
  ```
333
203
 
204
+ ## Known Bugs
205
+ 1. Inconsistency and incorrectness in the positioning of footnotes and endnotes in .docx files where the footnotes and endnotes would end up at the end of the parsed text whereas it would be positioned exactly after the referenced word in .odt files.
206
+ 2. The charts and objects information of .odt files are not accurate and may end up showing a few NaN in some cases.
334
207
  ----------
335
208
 
209
+ **npm**
210
+ https://npmjs.com/package/officeparser
211
+
336
212
  **github**
337
213
  https://github.com/harshankur/officeParser