officeparser 5.0.0 → 5.1.1

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -13,6 +13,7 @@ A Node.js library to parse text out of any office file.
13
13
 
14
14
 
15
15
  #### Update
16
+ * 2024/11/12 - Added ArrayBuffer as a type of file input. Generating bundle files now which exposes namespace officeParser to be able to access parseOffice and parseOfficeAsync directly on the browser. Extracting text out of pdf files does not work currently in browser bundles.
16
17
  * 2024/10/21 - Replaced extracting zip files from decompress to yauzl. This means that we now extract files in memory and we no longer need to write them to disk. Removed config flags related to extracted files. Added flags for CLI execution.
17
18
  * 2024/10/15 - Fixed erroring out while deleting temp files when multiple worker threads make parallel executions resulting in same file name for multiple files. Fixed erroring out when multiple executions are made without waiting for the previous execution to finish which resulted in deleting the file from other execution. Upgraded dependencies.
18
19
  * 2024/10/13 - Fixed parsing text from xlsx files which contain no shared strings file and files which have inlineStr based strings.
@@ -185,10 +186,51 @@ function searchForTermInOfficeFile(searchterm: string, filepath: string): Promis
185
186
  \
186
187
  **Please take note: I have breached convention in placing err as second argument in my callback but please understand that I had to do it to not break other people's existing modules.**
187
188
 
189
+ ## Browser Usage
190
+ Download the bundle file available as part of the release asset.
191
+ Include this bundle file in your browser html file and access `parseOffice` and `parseOfficeAsync` under the **`officeParser`** namespace.
192
+
193
+ **Example**
194
+ ```html
195
+ <head>
196
+ ...
197
+ <!-- Include bundle file in the script tag. -->
198
+ <script src="officeParserBundle@5.1.0.js"></script>
199
+ </head>
200
+ <body>
201
+ ...
202
+ <input type="file" id="fileInput" />
203
+ ...
204
+ <script>
205
+ document.getElementById('fileInput').addEventListener('change', async function(event) {
206
+ const outputDiv = document.getElementById('output');
207
+ const file = event.target.files[0];
208
+ try {
209
+ // Your configuration options for officeParser
210
+ const config = {
211
+ outputErrorToConsole: false,
212
+ newlineDelimiter: '\n',
213
+ ignoreNotes: false,
214
+ putNotesAtLast: false
215
+ };
216
+
217
+ const arrayBuffer = await file.arrayBuffer();
218
+ const result = await officeParser.parseOfficeAsync(arrayBuffer, config);
219
+ // result contains the extracted text.
220
+ }
221
+ catch (error) {
222
+ // Handle error
223
+ }
224
+ });
225
+ </script>
226
+ </body>
227
+ ```
228
+
188
229
 
189
230
  ## Known Bugs
190
231
  1. Inconsistency and incorrectness in the positioning of footnotes and endnotes in .docx files where the footnotes and endnotes would end up at the end of the parsed text whereas it would be positioned exactly after the referenced word in .odt files.
191
232
  2. The charts and objects information of .odt files are not accurate and may end up showing a few NaN in some cases.
233
+ 3. Extracting texts in browser bundles does not work for pdf files.
192
234
  ----------
193
235
 
194
236
  **npm**
package/officeParser.js CHANGED
@@ -479,12 +479,12 @@ function parsePdf(file, callback, config) {
479
479
  }
480
480
 
481
481
  /** Main async function with callback to execute parseOffice for supported files
482
- * @param {string | Buffer} file File path or file buffers
483
- * @param {function} callback Callback function that returns value or error
484
- * @param {OfficeParserConfig} [config={}] [OPTIONAL]: Config Object for officeParser
482
+ * @param {string | Buffer | ArrayBuffer} srcFile File path or file buffers or Javascript ArrayBuffer
483
+ * @param {function} callback Callback function that returns value or error
484
+ * @param {OfficeParserConfig} [config={}] [OPTIONAL]: Config Object for officeParser
485
485
  * @returns {void}
486
486
  */
487
- function parseOffice(file, callback, config = {}) {
487
+ function parseOffice(srcFile, callback, config = {}) {
488
488
  // Make a clone of the config with default values such that none of the config flags are undefined.
489
489
  /** @type {OfficeParserConfig} */
490
490
  const internalConfig = {
@@ -494,6 +494,12 @@ function parseOffice(file, callback, config = {}) {
494
494
  outputErrorToConsole: false,
495
495
  ...config
496
496
  };
497
+
498
+ // Our internal code can process regular node Buffers or file path.
499
+ // So, if the src file was presented as ArrayBuffers, we create Buffers from them.
500
+ let file = srcFile instanceof ArrayBuffer ? Buffer.from(srcFile)
501
+ : srcFile;
502
+
497
503
  /**
498
504
  * Prepare file for processing
499
505
  * @type {Promise<{ file:string | Buffer, ext: string}>}
@@ -559,13 +565,13 @@ function parseOffice(file, callback, config = {}) {
559
565
  }
560
566
 
561
567
  /** Main async function that can be used with await to execute parseOffice. Or it can be used with promises.
562
- * @param {string | Buffer} file File path or file buffers
563
- * @param {OfficeParserConfig} [config={}] [OPTIONAL]: Config Object for officeParser
568
+ * @param {string | Buffer | ArrayBuffer} srcFile File path or file buffers or Javascript ArrayBuffer
569
+ * @param {OfficeParserConfig} [config={}] [OPTIONAL]: Config Object for officeParser
564
570
  * @returns {Promise<string>}
565
571
  */
566
- function parseOfficeAsync(file, config = {}) {
572
+ function parseOfficeAsync(srcFile, config = {}) {
567
573
  return new Promise((res, rej) => {
568
- parseOffice(file, function (data, err) {
574
+ parseOffice(srcFile, function (data, err) {
569
575
  if (err)
570
576
  return rej(err);
571
577
  return res(data);
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "officeparser",
3
- "version": "5.0.0",
3
+ "version": "5.1.1",
4
4
  "description": "A Node.js library to parse text out of any office file. Currently supports docx, pptx, xlsx, odt, odp, ods, pdf files.",
5
5
  "main": "officeParser.js",
6
6
  "files": [