officeparser 4.0.4 → 4.0.6

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (3) hide show
  1. package/README.md +15 -13
  2. package/officeParser.js +68 -20
  3. package/package.json +1 -1
package/README.md CHANGED
@@ -13,6 +13,7 @@ A Node.js library to parse text out of any office file.
13
13
 
14
14
 
15
15
  #### Update
16
+ * 2023/11/25 - Fixed error catching when an error occurs within the parsing of a file, especially after decompressing it. Also fixed the problem with parallel parsing of files as we were using only timestamp in file names.
16
17
  * 2023/10/24 - Revamped content parsing code. Fixed order of content in files, especially in word files where table information would always land up at the end of the text. Added config object as argument for parseOffice which can be used to set new line delimiter and multiple other configurations. Added support for parsing pdf files using the popular npm library pdf-parse. Removed support for individual file parsing functions.
17
18
  * 2023/04/26 - Added support for file buffers as argument for filepath for parseOffice and parseOfficeAsync
18
19
  * 2023/04/07 - Added typings to methods to help with Typescript projects.
@@ -95,14 +96,15 @@ officeParser.parseOfficeAsync(fileBuffers);
95
96
 
96
97
  ### Configuration Object: OfficeParserConfig
97
98
  *Optionally add a config object as 3rd variable to parseOffice for the following configurations*
98
- | flag | datatype | explanation |
99
- |----------------------|----------|-------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|
100
- | tempFilesLocation | string | The directory where officeparser stores the temp files . The final decompressed data will be put inside officeParserTemp folder within your directory. **Please ensure that this directory actually exists.** Default is officeParsertemp. |
101
- | preserveTempFiles | boolean | Flag to not delete the internal content files and the possible duplicate temp files that it uses after unzipping office files. Default is false. It always deletes all of those files. |
102
- | outputErrorToConsole | boolean | Flag to show all the logs to console in case of an error. |
103
- | newlineDelimiter | string | The delimiter used for every new line in places that allow multiline text like word. Default is \n. |
104
- | ignoreNotes | boolean | Flag to ignore notes from parsing in files like powerpoint. Default is false. It includes notes in the parsed text by default. |
105
- | putNotesAtLast | boolean | Flag, if set to true, will collectively put all the parsed text from notes at last in files like powerpoint. Default is false. It puts each notes right after its main slide content. If ignoreNotes is set to true, this flag is also ignored. |
99
+ | Flag | DataType | Default | Explanation |
100
+ |----------------------|----------|------------------|-------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|
101
+ | tempFilesLocation | string | officeParserTemp | The directory where officeparser stores the temp files . The final decompressed data will be put inside officeParserTemp folder within your directory. **Please ensure that this directory actually exists.** Default is officeParserTemp. |
102
+ | preserveTempFiles | boolean | false | Flag to not delete the internal content files and the possible duplicate temp files that it uses after unzipping office files. Default is false. It always deletes all of those files. |
103
+ | outputErrorToConsole | boolean | false | Flag to show all the logs to console in case of an error. Default is false. |
104
+ | newlineDelimiter | string | \n | The delimiter used for every new line in places that allow multiline text like word. Default is \n. |
105
+ | ignoreNotes | boolean | false | Flag to ignore notes from parsing in files like powerpoint. Default is false. It includes notes in the parsed text by default. |
106
+ | putNotesAtLast | boolean | false | Flag, if set to true, will collectively put all the parsed text from notes at last in files like powerpoint. Default is false. It puts each notes right after its main slide content. If ignoreNotes is set to true, this flag is also ignored. |
107
+ <br>
106
108
 
107
109
  ```js
108
110
  const config = {
@@ -121,8 +123,8 @@ officeParser.parseOffice("/path/to/officeFile", function(data, err){
121
123
 
122
124
  // promise
123
125
  officeParser.parseOfficeAsync("/path/to/officeFile", config);
124
- .then((data) => console.log(data))
125
- .catch((err) => console.error(err))
126
+ .then(data => console.log(data))
127
+ .catch(err => console.error(err))
126
128
  ```
127
129
 
128
130
  **Example - JavaScript**
@@ -140,7 +142,7 @@ officeParser.parseOfficeAsync("/Users/harsh/Desktop/files/mySlides.pptx", config
140
142
  const newText = data + " look, I can parse a powerpoint file";
141
143
  callSomeOtherFunction(newText);
142
144
  })
143
- .catch((err) => console.error(err));
145
+ .catch(err => console.error(err));
144
146
 
145
147
  // Search for a term in the parsed text.
146
148
  function searchForTermInOfficeFile(searchterm, filepath) {
@@ -165,10 +167,10 @@ officeParser.parseOfficeAsync("/Users/harsh/Desktop/files/mySlides.pptx", config
165
167
  const newText = data + " look, I can parse a powerpoint file";
166
168
  callSomeOtherFunction(newText);
167
169
  })
168
- .catch((err) => console.error(err));
170
+ .catch(err => console.error(err));
169
171
 
170
172
  // Search for a term in the parsed text.
171
- function searchForTermInOfficeFile(searchterm, filepath): Promise<boolean> {
173
+ function searchForTermInOfficeFile(searchterm: string, filepath: string): Promise<boolean> {
172
174
  return officeParser.parseOfficeAsync(filepath)
173
175
  .then(data => data.indexOf(searchterm) != -1)
174
176
  }
package/officeParser.js CHANGED
@@ -69,10 +69,11 @@ function parseWord(filepath, callback, config) {
69
69
  { filter: x => [mainContentFile, footnotesFile, endnotesFile].includes(x.path) }
70
70
  )
71
71
  .then(files => {
72
- if (files.length == 0)
72
+ // Verify if atleast the document xml file exists in the extracted files list.
73
+ if (!files.map(file => file.path).includes(mainContentFile))
73
74
  throw ERRORMSG.fileCorrupted(filepath);
74
75
 
75
- return [...files.filter(file => file.path == mainContentFile),
76
+ return [...files.filter(file => file.path == mainContentFile),
76
77
  ...files.filter(file => file.path == footnotesFile),
77
78
  ...files.filter(file => file.path == endnotesFile)
78
79
  ]
@@ -99,7 +100,10 @@ function parseWord(filepath, callback, config) {
99
100
  // Find text nodes with w:t tags
100
101
  const xmlTextNodeList = paragraphNode.getElementsByTagName("w:t");
101
102
  // Join the texts within this paragraph node without any spaces or delimiters.
102
- return Array.from(xmlTextNodeList).map(textNode => textNode.childNodes[0].nodeValue).join("");
103
+ return Array.from(xmlTextNodeList)
104
+ .filter(textNode => textNode.childNodes[0] && textNode.childNodes[0].nodeValue)
105
+ .map(textNode => textNode.childNodes[0].nodeValue)
106
+ .join("");
103
107
  })
104
108
  // Join each paragraph text with a new line delimiter.
105
109
  .join(config.newlineDelimiter ?? "\n")
@@ -132,8 +136,8 @@ function parsePowerPoint(filepath, callback, config) {
132
136
  { filter: x => x.path.match(config.ignoreNotes ? slidesRegex : allFilesRegex) }
133
137
  )
134
138
  .then(files => {
135
- // Check if files is corrupted
136
- if (files.length == 0)
139
+ // Verify if atleast the slides xml files exist in the extracted files list.
140
+ if (files.length == 0 || !files.map(file => file.path).some(filename => filename.match(slidesRegex)))
137
141
  throw ERRORMSG.fileCorrupted(filepath);
138
142
 
139
143
  // Check if any sorting is required.
@@ -166,7 +170,10 @@ function parsePowerPoint(filepath, callback, config) {
166
170
  .map(paragraphNode => {
167
171
  /** Find text nodes with a:t tags */
168
172
  const xmlTextNodeList = paragraphNode.getElementsByTagName("a:t");
169
- return Array.from(xmlTextNodeList).map(textNode => textNode.childNodes[0].nodeValue).join("");
173
+ return Array.from(xmlTextNodeList)
174
+ .filter(textNode => textNode.childNodes[0] && textNode.childNodes[0].nodeValue)
175
+ .map(textNode => textNode.childNodes[0].nodeValue)
176
+ .join("");
170
177
  })
171
178
  .join(config.newlineDelimiter ?? "\n")
172
179
  );
@@ -200,7 +207,8 @@ function parseExcel(filepath, callback, config) {
200
207
  { filter: x => ([sheetsRegex, drawingsRegex, chartsRegex].findIndex(fileRegex => x.path.match(fileRegex)) > -1) || (x.path == stringsFilePath )}
201
208
  )
202
209
  .then(files => {
203
- if (files.length == 0)
210
+ // Verify if atleast the slides xml files exist in the extracted files list.
211
+ if (files.length == 0 || !files.map(file => file.path).some(filename => filename.match(sheetsRegex)))
204
212
  throw ERRORMSG.fileCorrupted(filepath);
205
213
 
206
214
  return {
@@ -225,7 +233,9 @@ function parseExcel(filepath, callback, config) {
225
233
  /** Find text nodes with t tags in sharedStrings xml file */
226
234
  const sharedStringsXmlTNodesList = parseString(xmlContentFilesObject.sharedStringsFile).getElementsByTagName("t");
227
235
  /** Create shared string array. This will be used as a map to get strings from within sheet files. */
228
- const sharedStrings = Array.from(sharedStringsXmlTNodesList).map(tNode => tNode.childNodes[0].nodeValue);
236
+ const sharedStrings = Array.from(sharedStringsXmlTNodesList)
237
+ .filter(tNode => tNode.childNodes[0] && tNode.childNodes[0].nodeValue)
238
+ .map(tNode => tNode.childNodes[0].nodeValue);
229
239
 
230
240
  // Parse Sheet files
231
241
  xmlContentFilesObject.sheetFiles.forEach(sheetXmlContent => {
@@ -234,8 +244,10 @@ function parseExcel(filepath, callback, config) {
234
244
  // Traverse through the nodes list and fill responseText with either the number value in its v node or find a mapped string from sharedStrings.
235
245
  responseText.push(
236
246
  Array.from(sheetsXmlCNodesList)
237
- // Filter c nodes than do not have any v nodes
238
- .filter(cNode => cNode.getElementsByTagName("v").length != 0)
247
+ // Filter c nodes than do not have any valid v nodes
248
+ .filter(cNode => cNode.getElementsByTagName("v")[0]
249
+ && cNode.getElementsByTagName("v")[0].childNodes[0]
250
+ && cNode.getElementsByTagName("v")[0].childNodes[0].nodeValue)
239
251
  .map(cNode => {
240
252
  /** Flag whether this node's value represents a string index */
241
253
  const isString = cNode.getAttribute("t") == "s";
@@ -266,7 +278,10 @@ function parseExcel(filepath, callback, config) {
266
278
  .map(paragraphNode => {
267
279
  /** Find text nodes with a:t tags */
268
280
  const xmlTextNodeList = paragraphNode.getElementsByTagName("a:t");
269
- return Array.from(xmlTextNodeList).map(textNode => textNode.childNodes[0].nodeValue).join("");
281
+ return Array.from(xmlTextNodeList)
282
+ .filter(textNode => textNode.childNodes[0] && textNode.childNodes[0].nodeValue)
283
+ .map(textNode => textNode.childNodes[0].nodeValue)
284
+ .join("");
270
285
  })
271
286
  .join(config.newlineDelimiter ?? "\n")
272
287
  );
@@ -279,6 +294,7 @@ function parseExcel(filepath, callback, config) {
279
294
  /** Store all the text content to respond */
280
295
  responseText.push(
281
296
  Array.from(chartsXmlCVNodesList)
297
+ .filter(cVNode => cVNode.childNodes[0] && cVNode.childNodes[0].nodeValue)
282
298
  .map(cVNode => cVNode.childNodes[0].nodeValue)
283
299
  .join(config.newlineDelimiter ?? "\n")
284
300
  );
@@ -311,7 +327,8 @@ function parseOpenOffice(filepath, callback, config) {
311
327
  { filter: x => x.path == mainContentFilePath || x.path.match(objectContentFilesRegex) }
312
328
  )
313
329
  .then(files => {
314
- if (files.length == 0)
330
+ // Verify if atleast the content xml file exists in the extracted files list.
331
+ if (!files.map(file => file.path).includes(mainContentFilePath))
315
332
  throw ERRORMSG.fileCorrupted(filepath);
316
333
 
317
334
  return {
@@ -475,7 +492,7 @@ function parseOffice(file, callback, config = {}) {
475
492
  .then(data =>
476
493
  {
477
494
  // temp file name
478
- const newfilepath = `${internalConfig.tempFilesLocation}/tempfiles/${new Date().getTime().toString()}.${data.ext.toLowerCase()}`;
495
+ const newfilepath = getNewFileName(internalConfig.tempFilesLocation, data.ext.toLowerCase());
479
496
  // write new file
480
497
  fs.writeFileSync(newfilepath, file);
481
498
  // resolve promise
@@ -492,7 +509,7 @@ function parseOffice(file, callback, config = {}) {
492
509
  throw ERRORMSG.fileDoesNotExist(file);
493
510
 
494
511
  // temp file name
495
- const newfilepath = `${internalConfig.tempFilesLocation}/tempfiles/${new Date().getTime().toString()}.${file.split(".").pop().toLowerCase()}`;
512
+ const newfilepath = getNewFileName(internalConfig.tempFilesLocation, file.split(".").pop().toLowerCase());
496
513
  // Copy the file into a temp location with the temp name
497
514
  fs.copyFileSync(file, newfilepath)
498
515
  // resolve promise
@@ -538,16 +555,13 @@ function parseOffice(file, callback, config = {}) {
538
555
 
539
556
  // Check if there is an error. Throw if there is an error.
540
557
  if (err)
541
- throw err;
558
+ return handleError(err, callback, internalConfig.outputErrorToConsole);
542
559
 
543
560
  // Call the original callback
544
561
  callback(data, undefined);
545
562
  }
546
563
  })
547
- .catch(error => {
548
- consoleError(error, internalConfig.outputErrorToConsole);
549
- callback(undefined, ERRORHEADER + error);
550
- });
564
+ .catch(error => handleError(error, callback, internalConfig.outputErrorToConsole));
551
565
  }
552
566
 
553
567
  /**
@@ -556,7 +570,7 @@ function parseOffice(file, callback, config = {}) {
556
570
  * @param {OfficeParserConfig} [config={}] [OPTIONAL]: Config Object for officeParser
557
571
  * @returns {Promise<string>}
558
572
  */
559
- function parseOfficeAsync (file, config = {}) {
573
+ function parseOfficeAsync(file, config = {}) {
560
574
  return new Promise((res, rej) => {
561
575
  parseOffice(file, function (data, err) {
562
576
  if (err)
@@ -566,6 +580,40 @@ function parseOfficeAsync (file, config = {}) {
566
580
  });
567
581
  }
568
582
 
583
+ /** Global file name iterator. */
584
+ let globalFileNameIterator = 0;
585
+ /**
586
+ * File Name generator that takes the extension as an input and returns a file name that comprises a timestamp and an incrementing number
587
+ * to allow the files to be sorted in chronological order
588
+ * @param {string} tempFilesLocation Directory whether this new file needs to be stored
589
+ * @param {string} ext File extension for this new generated file name
590
+ * @returns {string}
591
+ */
592
+ function getNewFileName(tempFilesLocation, ext) {
593
+ // Get the iterator part of the file name
594
+ let iteratorPart = (globalFileNameIterator++).toString().padStart(5, '0');
595
+ // We want the iterator part of the file name to be of 5 digits.
596
+ // Therefore, when the iterator crosses into 6 digits, we reset it to 0.
597
+ if (globalFileNameIterator > 99999)
598
+ globalFileNameIterator = 0;
599
+
600
+ // Return the file name
601
+ return `${tempFilesLocation}/tempfiles/${new Date().getTime().toString() + iteratorPart}.${ext}`;
602
+ }
603
+
604
+ /**
605
+ * Handle error by logging it to console if permitted by the config.
606
+ * And after that, trigger the callback function with the error value.
607
+ * @param {string} error Error text
608
+ * @param {function} callback Callback function provided by the caller
609
+ * @param {boolean} outputErrorToConsole Flag to log error to console.
610
+ * @returns {void}
611
+ */
612
+ function handleError(error, callback, outputErrorToConsole) {
613
+ consoleError(error, outputErrorToConsole);
614
+ callback(undefined, ERRORHEADER + error);
615
+ }
616
+
569
617
 
570
618
  // Export functions
571
619
  module.exports.parseOffice = parseOffice;
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "officeparser",
3
- "version": "4.0.4",
3
+ "version": "4.0.6",
4
4
  "description": "A Node.js library to parse text out of any office file. Currently supports docx, pptx, xlsx, odt, odp, ods, pdf files.",
5
5
  "main": "officeParser.js",
6
6
  "files": [