officeparser 2.2.2 → 3.0.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -1,172 +1,324 @@
1
- # officeParser
2
- A Node.js library to parse text out of any office file.
3
-
4
- ### Supported File Types
5
-
6
- - [`docx`](https://en.wikipedia.org/wiki/Office_Open_XML)
7
- - [`pptx`](https://en.wikipedia.org/wiki/Office_Open_XML)
8
- - [`xlsx`](https://en.wikipedia.org/wiki/Office_Open_XML)
9
- - [`odt`](https://en.wikipedia.org/wiki/OpenDocument)
10
- - [`odp`](https://en.wikipedia.org/wiki/OpenDocument)
11
- - [`ods`](https://en.wikipedia.org/wiki/OpenDocument)
12
-
13
-
14
- #### Update
15
- * 2020/06/01 - Added error handling and console.log enable/disable methods. Default is set at enabled. Everything backward compatible.
16
- * 2019/06/17 - Added method to change location for decompressing office files in places with restricted write access.
17
- * 2019/04/30 - Removed case sensitive file extension bug. File names with capital lettered extensions now supported.
18
- * 2019/04/23 - Added support for open office files *.odt, *.odp, *.ods through parseOffice function. Created a new method parseOpenOffice for those who prefer targetted functions.
19
- * 2019/04/23 - Added feature to delete the generated dist folder after function callback
20
- * 2019/04/22 - Added parseOffice method to avoid confusion between type of file and their extension
21
- * 2019/04/22 - Added file extension validations. Removed errors for excel files with no drawing elements.
22
- * 2019/04/19 - Support added for *.xlsx files.
23
- * 2019/04/18 - Support added for *.pptx files.
24
-
25
-
26
-
27
- ## Install via npm
28
-
29
-
30
- ```
31
- npm i officeparser
32
- ```
33
-
34
- ----------
35
-
36
- **Usage**
37
- ```
38
- const officeParser = require('officeparser');
39
-
40
- officeParser.parseOffice("/path/to/officeFile", function(data, err){
41
- // "data" string in the callback here is the text parsed from the office file passed in the first argument above
42
- if (err) return console.log(err);
43
- console.log(data)
44
- })
45
-
46
- ```
47
-
48
- **Please take note: I have breached convention in placing err as second argument in my callback but please understand that I had to do it to not break other people's existing modules.**
49
-
50
- *Optionally change decompression location for office Files at personalised locations for environments with restricted write access*
51
-
52
- ```
53
- const officeParser = require('officeparser');
54
-
55
- // Default decompress location for office Files is "officeDist" in the directory where Node is started.
56
- // Put this file before parseOffice method to take effect.
57
- officeParser.setDecompressionLocation("/tmp"); // New decompression location would be "/tmp/officeDist"
58
-
59
- // P.S.: Setting location on a Windows environment with '\' heirarchy requires to be entered twice '\\'
60
- officeParser.setDecompressionLocation("C:\\tmp"); // New decompression location would be "C:\tmp\officeDist"
61
-
62
-
63
- officeParser.parseOffice("/path/to/officeFile", function(data, err){
64
- // "data" string in the callback here is the text parsed from the office file passed in the first argument above
65
- if (err) return console.log(err);
66
- console.log(data)
67
- })
68
- ```
69
-
70
- *Optionally add false as 3rd variable to parseOffice to not delete the generated officeDist folder*
71
-
72
- ```
73
- officeParser.parseOffice("/path/to/officeFile", function(data, err){
74
- // "data" string in the callback here is the text parsed from the office file passed in the first argument above
75
- if (err) return console.log(err);
76
- console.log(data)
77
- }, false)
78
- ```
79
-
80
- **Example**
81
- ```
82
- const officeParser = require('officeparser');
83
-
84
- officeParser.parseOffice("C:\\files\\myText.docx", function(data, err){
85
- if (err) return console.log(err);
86
- var newText = data + "look, I can parse a word file"
87
- callSomeOtherFunction(newText);
88
- })
89
-
90
- officeParser.parseOffice("/Users/harsh/Desktop/files/mySlides.pptx", function(data, err){
91
- if (err) return console.log(err);
92
- var newText = data + "look, I can parse a powerpoint file"
93
- callSomeOtherFunction(newText);
94
- })
95
-
96
- // Using relative path for file is also fine
97
- officeParser.parseOffice("files/myWorkSheet.ods", function(data, err){
98
- if (err) return console.log(err);
99
- var newText = data + "look, I can parse an excel file"
100
- callSomeOtherFunction(newText);
101
- })
102
- ```
103
-
104
-
105
- ----------
106
-
107
- ### Old but functional way of extracting text from word, powerpoint and excel files
108
- *These were the initial methods of parsing text till parseOffice method came into existence. These still exist and form the skeleton to this module as parseOffice redirects the below functions anyway. These functions will forever remain available to guarantee long-term usage of this module. I will ensure backward-compatibility with all previous versions.*
109
-
110
- **Usage**
111
- ```
112
- const officeParser = require('officeparser');
113
-
114
- officeParser.parseWord("/path/to/word.docx", function(data, err){
115
- // "data" string in the callback here is the text parsed from the word file passed in the first argument above
116
- if (err) return console.log(err);
117
- console.log(data)
118
- })
119
-
120
- officeParser.parsePowerPoint("/path/to/powerpoint.pptx", function(data, err){
121
- // "data" string in the callback here is the text parsed from the powerpoint file passed in the first argument above
122
- if (err) return console.log(err);
123
- console.log(data)
124
- })
125
-
126
- officeParser.parseExcel("/path/to/excel.xlsx", function(data, err){
127
- // "data" string in the callback here is the text parsed from the excel file passed in the first argument above
128
- if (err) return console.log(err);
129
- console.log(data)
130
- })
131
-
132
- officeParser.parseOpenOffice("/path/to/writer.odt", function(data, err){
133
- // "data" string in the callback here is the text parsed from the writer file passed in the first argument above
134
- if (err) return console.log(err);
135
- console.log(data)
136
- })
137
- ```
138
-
139
- **Example**
140
- ```
141
- const officeParser = require('officeparser');
142
-
143
- officeParser.parseWord("C:\\files\\myText.docx", function(data, err){
144
- if (err) return console.log(err);
145
- var newText = data + "look, I can parse a word file"
146
- callSomeOtherFunction(newText);
147
- })
148
-
149
- officeParser.parsePowerPoint("/Users/harsh/Desktop/files/mySlides.pptx", function(data, err){
150
- if (err) return console.log(err);
151
- var newText = data + "look, I can parse a powerpoint file"
152
- callSomeOtherFunction(newText);
153
- })
154
-
155
- // Using relative path for file is also fine
156
- officeParser.parseExcel("files/myWorkSheet.xlsx", function(data, err){
157
- if (err) return console.log(err);
158
- var newText = data + "look, I can parse an excel file"
159
- callSomeOtherFunction(newText);
160
- })
161
-
162
- officeParser.parseOpenOffice("files/myDocument.odt", function(data, err){
163
- if (err) return console.log(err);
164
- var newText = data + "look, I can parse an OpenOffice file"
165
- callSomeOtherFunction(newText);
166
- })
167
- ```
168
-
169
- ----------
170
-
171
- **github**
172
- https://github.com/harshankur/officeParser
1
+ # officeParser
2
+ A Node.js library to parse text out of any office file.
3
+
4
+ ### Supported File Types
5
+
6
+ - [`docx`](https://en.wikipedia.org/wiki/Office_Open_XML)
7
+ - [`pptx`](https://en.wikipedia.org/wiki/Office_Open_XML)
8
+ - [`xlsx`](https://en.wikipedia.org/wiki/Office_Open_XML)
9
+ - [`odt`](https://en.wikipedia.org/wiki/OpenDocument)
10
+ - [`odp`](https://en.wikipedia.org/wiki/OpenDocument)
11
+ - [`ods`](https://en.wikipedia.org/wiki/OpenDocument)
12
+
13
+
14
+ #### Update
15
+ * 2022/12/10 - Fixed memory leak issues, bugs related to parsing open document files and improved error handling
16
+ * 2021/11/21 - Added promise way to existing callback functions
17
+ * 2020/06/01 - Added error handling and console.log enable/disable methods. Default is set at enabled. Everything backward compatible.
18
+ * 2019/06/17 - Added method to change location for decompressing office files in places with restricted write access.
19
+ * 2019/04/30 - Removed case sensitive file extension bug. File names with capital lettered extensions now supported.
20
+ * 2019/04/23 - Added support for open office files *.odt, *.odp, *.ods through parseOffice function. Created a new method parseOpenOffice for those who prefer targetted functions.
21
+ * 2019/04/23 - Added feature to delete the generated dist folder after function callback
22
+ * 2019/04/22 - Added parseOffice method to avoid confusion between type of file and their extension
23
+ * 2019/04/22 - Added file extension validations. Removed errors for excel files with no drawing elements.
24
+ * 2019/04/19 - Support added for *.xlsx files.
25
+ * 2019/04/18 - Support added for *.pptx files.
26
+
27
+
28
+
29
+ ## Install via npm
30
+
31
+
32
+ ```
33
+ npm i officeparser
34
+ ```
35
+
36
+ ----------
37
+
38
+ **Usage**
39
+ ```js
40
+ const officeParser = require('officeparser');
41
+
42
+ // callback
43
+ officeParser.parseOffice("/path/to/officeFile", function(data, err){
44
+ // "data" string in the callback here is the text parsed from the office file passed in the first argument above
45
+ if (err) return console.log(err);
46
+ console.log(data)
47
+ })
48
+
49
+ // promise
50
+ officeParser.parseOfficeAsync("/path/to/officeFile");
51
+ // "data" string in the promise here is the text parsed from the office file passed in the argument above
52
+ .then((data) => {
53
+ console.log(data)
54
+ })
55
+ .catch(err) => {
56
+ console.log(err)
57
+ }
58
+
59
+ // async/await
60
+ try {
61
+ // "data" string returned from promise here is the text parsed from the office file passed in the argument
62
+ const data = await officeParser.parseOfficeAsync("/path/to/officeFile");
63
+ console.log(data);
64
+ } catch (err) {
65
+ // resolve error
66
+ console.log(err);
67
+ }
68
+ ```
69
+
70
+ **Please take note: I have breached convention in placing err as second argument in my callback but please understand that I had to do it to not break other people's existing modules.**
71
+
72
+ *Optionally change decompression location for office Files at personalised locations for environments with restricted write access*
73
+
74
+ ```js
75
+ const officeParser = require('officeparser');
76
+
77
+ // Default decompress location for office Files is "officeDist" in the directory where Node is started.
78
+ // Put this file before parseOffice method to take effect.
79
+ officeParser.setDecompressionLocation("/tmp"); // New decompression location would be "/tmp/officeDist"
80
+
81
+ // P.S.: Setting location on a Windows environment with '\' hierarchy requires to be entered twice '\\'
82
+ officeParser.setDecompressionLocation("C:\\tmp"); // New decompression location would be "C:\tmp\officeDist"
83
+
84
+
85
+ officeParser.parseOffice("/path/to/officeFile", function(data, err){
86
+ // "data" string in the callback here is the text parsed from the office file passed in the first argument above
87
+ if (err) return console.log(err);
88
+ console.log(data)
89
+ })
90
+ ```
91
+
92
+ *Optionally add false as 3rd variable to parseOffice to not delete the generated officeDist folder*
93
+
94
+ ```js
95
+ // callback
96
+ officeParser.parseOffice("/path/to/officeFile", function(data, err){
97
+ if (err) return console.log(err);
98
+ console.log(data)
99
+ }, false)
100
+
101
+ // promise
102
+ officeParser.parseOfficeAsync("/path/to/officeFile", false);
103
+ .then((data) => {
104
+ console.log(data)
105
+ })
106
+ .catch(err) => {
107
+ console.log(err)
108
+ }
109
+
110
+ // async/await
111
+ try {
112
+ const data = await officeParser.parseOfficeAsync("/path/to/officeFile", false);
113
+ console.log(data);
114
+ } catch (err) {
115
+ // resolve error
116
+ console.log(err);
117
+ }
118
+ ```
119
+
120
+ **Example**
121
+ ```js
122
+ const officeParser = require('officeparser');
123
+
124
+ // callback
125
+ officeParser.parseOffice("C:\\files\\myText.docx", function(data, err){
126
+ if (err) return console.log(err);
127
+ var newText = data + "look, I can parse a word file"
128
+ callSomeOtherFunction(newText);
129
+ })
130
+
131
+
132
+ // promise
133
+ officeParser.parseOfficeAsync("/Users/harsh/Desktop/files/mySlides.pptx");
134
+ .then((data) => {
135
+ var newText = data + "look, I can parse a powerpoint file"
136
+ callSomeOtherFunction(newText);
137
+ })
138
+ .catch(err) => {
139
+ console.log(err)
140
+ }
141
+
142
+ // Using relative path for file is also fine
143
+ officeParser.parseOffice("files/myWorkSheet.ods", function(data, err){
144
+ if (err) return console.log(err);
145
+ var newText = data + "look, I can parse an excel file"
146
+ callSomeOtherFunction(newText);
147
+ })
148
+
149
+ // async/await
150
+ try {
151
+ const data = await officeParser.parseOfficeAsync("/Users/harsh/Desktop/files/mySlides.pptx");
152
+ let newText = data + "look, I can parse a powerpoint file";
153
+ await callSomeOtherFunction(newText);
154
+ } catch (err) {
155
+ // resolve error
156
+ console.log(err);
157
+ }
158
+ ```
159
+
160
+
161
+ ----------
162
+
163
+ ### Old but functional way of extracting text from word, powerpoint and excel files
164
+ *These were the initial methods of parsing text till parseOffice method came into existence. These still exist and form the skeleton to this module as parseOffice redirects the below functions anyway. These functions will forever remain available to guarantee long-term usage of this module. I will ensure backward-compatibility with all previous versions.*
165
+
166
+ **Usage**
167
+ ```js
168
+ const officeParser = require('officeparser');
169
+
170
+ // callback
171
+ officeParser.parseWord("/path/to/word.docx", function(data, err){
172
+ // "data" string in the callback here is the text parsed from the word file passed in the first argument above
173
+ if (err) return console.log(err);
174
+ console.log(data)
175
+ })
176
+
177
+ officeParser.parsePowerPoint("/path/to/powerpoint.pptx", function(data, err){
178
+ // "data" string in the callback here is the text parsed from the powerpoint file passed in the first argument above
179
+ if (err) return console.log(err);
180
+ console.log(data)
181
+ })
182
+
183
+ officeParser.parseExcel("/path/to/excel.xlsx", function(data, err){
184
+ // "data" string in the callback here is the text parsed from the excel file passed in the first argument above
185
+ if (err) return console.log(err);
186
+ console.log(data)
187
+ })
188
+
189
+ officeParser.parseOpenOffice("/path/to/writer.odt", function(data, err){
190
+ // "data" string in the callback here is the text parsed from the writer file passed in the first argument above
191
+ if (err) return console.log(err);
192
+ console.log(data)
193
+ })
194
+
195
+ // promise
196
+ officeParser.parseWordAsync("/path/to/word.docx");
197
+ .then((data) => {
198
+ // data is the parsed text
199
+ })
200
+ officeParser.parsePowerPointAsync("/path/to/powerpoint.pptx");
201
+ .then((data) => {
202
+ // data is the parsed text
203
+ })
204
+ officeParser.parseExcelAsync("/path/to/excel.xlsx");
205
+ .then((data) => {
206
+ // data is the parsed text
207
+ })
208
+ officeParser.parseOpenOfficeAsync("/path/to/writer.odt");
209
+ .then((data) => {
210
+ // data is the parsed text
211
+ })
212
+
213
+ // async/await
214
+ try {
215
+ // "data" string returned from promise here is the text parsed from the office file passed in the first argument
216
+ const data1 = await officeParser.parseWordAsync("/path/to/word.docx");
217
+
218
+ const data2 = await officeParser.parsePowerPointAsync("/path/to/powerpoint.pptx");
219
+
220
+ const data3 = await officeParser.parseExcelAsync("/path/to/excel.xlsx");
221
+
222
+ const data3 = await officeParser.parseOpenOfficeAsync("/path/to/writer.odt");
223
+ } catch (err) {
224
+ // resolve error
225
+ console.log(err);
226
+ }
227
+ ```
228
+
229
+ **Example**
230
+ ```js
231
+ const officeParser = require('officeparser');
232
+
233
+ // callback
234
+ officeParser.parseWord("C:\\files\\myText.docx", function(data, err){
235
+ if (err) return console.log(err);
236
+ var newText = data + "look, I can parse a word file"
237
+ callSomeOtherFunction(newText);
238
+ })
239
+
240
+ officeParser.parsePowerPoint("/Users/harsh/Desktop/files/mySlides.pptx", function(data, err){
241
+ if (err) return console.log(err);
242
+ var newText = data + "look, I can parse a powerpoint file"
243
+ callSomeOtherFunction(newText);
244
+ })
245
+
246
+ // Using relative path for file is also fine
247
+ officeParser.parseExcel("files/myWorkSheet.xlsx", function(data, err){
248
+ if (err) return console.log(err);
249
+ var newText = data + "look, I can parse an excel file"
250
+ callSomeOtherFunction(newText);
251
+ })
252
+
253
+ officeParser.parseOpenOffice("files/myDocument.odt", function(data, err){
254
+ if (err) return console.log(err);
255
+ var newText = data + "look, I can parse an OpenOffice file"
256
+ callSomeOtherFunction(newText);
257
+ })
258
+
259
+ // promise
260
+ officeParser.parseWordAsync("C:\\files\\myText.docx");
261
+ .then((data) => {
262
+ let newText1 = data1 + "look, I can parse a word file";
263
+ callSomeOtherFunction(newText1);
264
+ })
265
+ .catch(err) => {
266
+ console.log(err)
267
+ }
268
+
269
+ officeParser.parsePowerPointAsync("/Users/harsh/Desktop/files/mySlides.pptx");
270
+ .then((data) => {
271
+ let newText2 = data2 + "look, I can parse a powerpoint file";
272
+ callSomeOtherFunction(newText2);
273
+ })
274
+ .catch(err) => {
275
+ console.log(err)
276
+ }
277
+
278
+ officeParser.parseExcelAsync("files/myWorkSheet.xlsx");
279
+ .then((data) => {
280
+ let newText3 = data3 + "look, I can parse an excel file";
281
+ callSomeOtherFunction(newText3);
282
+ })
283
+ .catch(err) => {
284
+ console.log(err)
285
+ }
286
+
287
+ officeParser.parseOpenOfficeAsync("files/myDocument.odt");
288
+ .then((data) => {
289
+ let newText4 = data4 + "look, I can parse an OpenOffice file";
290
+ callSomeOtherFunction(newText4);
291
+ })
292
+ .catch(err) => {
293
+ console.log(err)
294
+ }
295
+
296
+
297
+ // async/await
298
+ try {
299
+ const data1 = await officeParser.parseWordAsync("C:\\files\\myText.docx");
300
+ let newText1 = data1 + "look, I can parse a word file";
301
+ await callSomeOtherFunction(newText1);
302
+
303
+ const data2 = await officeParser.parsePowerPointAsync("/Users/harsh/Desktop/files/mySlides.pptx");
304
+ let newText2 = data2 + "look, I can parse a powerpoint file";
305
+ await callSomeOtherFunction(newText2);
306
+
307
+ // Using relative path for file is also fine
308
+ const data3 = await officeParser.parseExcelAsync("files/myWorkSheet.xlsx");
309
+ let newText3 = data3 + "look, I can parse an excel file";
310
+ await callSomeOtherFunction(newText3);
311
+
312
+ const data4 = await officeParser.parseOpenOfficeAsync("files/myDocument.odt");
313
+ let newText4 = data4 + "look, I can parse an OpenOffice file";
314
+ await callSomeOtherFunction(newText4);
315
+ } catch (err) {
316
+ // resolve error
317
+ console.log(err);
318
+ }
319
+ ```
320
+
321
+ ----------
322
+
323
+ **github**
324
+ https://github.com/harshankur/officeParser