officeparser 4.2.0 → 5.1.1
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +55 -9
- package/officeParser.js +592 -556
- package/package.json +5 -4
- package/typings/officeParser.d.ts +2 -11
package/README.md
CHANGED
|
@@ -13,6 +13,8 @@ A Node.js library to parse text out of any office file.
|
|
|
13
13
|
|
|
14
14
|
|
|
15
15
|
#### Update
|
|
16
|
+
* 2024/11/12 - Added ArrayBuffer as a type of file input. Generating bundle files now which exposes namespace officeParser to be able to access parseOffice and parseOfficeAsync directly on the browser. Extracting text out of pdf files does not work currently in browser bundles.
|
|
17
|
+
* 2024/10/21 - Replaced extracting zip files from decompress to yauzl. This means that we now extract files in memory and we no longer need to write them to disk. Removed config flags related to extracted files. Added flags for CLI execution.
|
|
16
18
|
* 2024/10/15 - Fixed erroring out while deleting temp files when multiple worker threads make parallel executions resulting in same file name for multiple files. Fixed erroring out when multiple executions are made without waiting for the previous execution to finish which resulted in deleting the file from other execution. Upgraded dependencies.
|
|
17
19
|
* 2024/10/13 - Fixed parsing text from xlsx files which contain no shared strings file and files which have inlineStr based strings.
|
|
18
20
|
* 2024/05/06 - Replaced pdf parsing support from pdf-parse library to natively building it using pdf.js library from Mozilla by analyzing its output. Added pdfjs-dist build as a local library.
|
|
@@ -37,7 +39,6 @@ A Node.js library to parse text out of any office file.
|
|
|
37
39
|
|
|
38
40
|
## Install via npm
|
|
39
41
|
|
|
40
|
-
|
|
41
42
|
```
|
|
42
43
|
npm i officeparser
|
|
43
44
|
```
|
|
@@ -45,14 +46,20 @@ npm i officeparser
|
|
|
45
46
|
## Command Line usage
|
|
46
47
|
If you want to call the installed officeParser.js file, use below command
|
|
47
48
|
```
|
|
48
|
-
node
|
|
49
|
+
node <path/to/officeParser.js> [--configOption=value] [FILE_PATH]
|
|
50
|
+
node officeparser [--configOption=value] [FILE_PATH]
|
|
49
51
|
```
|
|
50
52
|
|
|
51
|
-
Otherwise, you can simply use npx to instantly extract parsed data.
|
|
53
|
+
Otherwise, you can simply use npx without installing the node module to instantly extract parsed data.
|
|
52
54
|
```
|
|
53
|
-
npx officeparser
|
|
55
|
+
npx officeparser [--configOption=value] [FILE_PATH]
|
|
54
56
|
```
|
|
55
57
|
|
|
58
|
+
### Config Options:
|
|
59
|
+
- `--ignoreNotes=[true|false]` Flag to ignore notes from files like PowerPoint. Default is false.
|
|
60
|
+
- `--newlineDelimiter=[delimiter]` The delimiter to use for new lines. Default is `\n`.
|
|
61
|
+
- `--putNotesAtLast=[true|false]` Flag to collect notes at the end of files like PowerPoint. Default is false.
|
|
62
|
+
- `--outputErrorToConsole=[true|false]` Flag to output errors to the console. Default is false.
|
|
56
63
|
|
|
57
64
|
## Library Usage
|
|
58
65
|
```js
|
|
@@ -101,8 +108,6 @@ officeParser.parseOfficeAsync(fileBuffers);
|
|
|
101
108
|
*Optionally add a config object as 3rd variable to parseOffice for the following configurations*
|
|
102
109
|
| Flag | DataType | Default | Explanation |
|
|
103
110
|
|----------------------|----------|------------------|-------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|
|
|
104
|
-
| tempFilesLocation | string | officeParserTemp | The directory where officeparser stores the temp files . The final decompressed data will be put inside officeParserTemp folder within your directory. **Please ensure that this directory actually exists.** Default is officeParserTemp. |
|
|
105
|
-
| preserveTempFiles | boolean | false | Flag to not delete the internal content files and the possible duplicate temp files that it uses after unzipping office files. Default is false. It always deletes all of those files. |
|
|
106
111
|
| outputErrorToConsole | boolean | false | Flag to show all the logs to console in case of an error. Default is false. |
|
|
107
112
|
| newlineDelimiter | string | \n | The delimiter used for every new line in places that allow multiline text like word. Default is \n. |
|
|
108
113
|
| ignoreNotes | boolean | false | Flag to ignore notes from parsing in files like powerpoint. Default is false. It includes notes in the parsed text by default. |
|
|
@@ -157,7 +162,7 @@ function searchForTermInOfficeFile(searchterm, filepath) {
|
|
|
157
162
|
|
|
158
163
|
**Example - TypeScript**
|
|
159
164
|
```ts
|
|
160
|
-
|
|
165
|
+
import { OfficeParserConfig, parseOfficeAsync } from 'officeparser';
|
|
161
166
|
|
|
162
167
|
const config: OfficeParserConfig = {
|
|
163
168
|
newlineDelimiter: " ", // Separate new lines with a space instead of the default \n.
|
|
@@ -165,7 +170,7 @@ const config: OfficeParserConfig = {
|
|
|
165
170
|
}
|
|
166
171
|
|
|
167
172
|
// relative path is also fine => eg: files/myWorkSheet.ods
|
|
168
|
-
|
|
173
|
+
parseOfficeAsync("/Users/harsh/Desktop/files/mySlides.pptx", config);
|
|
169
174
|
.then(data => {
|
|
170
175
|
const newText = data + " look, I can parse a powerpoint file";
|
|
171
176
|
callSomeOtherFunction(newText);
|
|
@@ -174,17 +179,58 @@ officeParser.parseOfficeAsync("/Users/harsh/Desktop/files/mySlides.pptx", config
|
|
|
174
179
|
|
|
175
180
|
// Search for a term in the parsed text.
|
|
176
181
|
function searchForTermInOfficeFile(searchterm: string, filepath: string): Promise<boolean> {
|
|
177
|
-
return
|
|
182
|
+
return parseOfficeAsync(filepath)
|
|
178
183
|
.then(data => data.indexOf(searchterm) != -1)
|
|
179
184
|
}
|
|
180
185
|
```
|
|
181
186
|
\
|
|
182
187
|
**Please take note: I have breached convention in placing err as second argument in my callback but please understand that I had to do it to not break other people's existing modules.**
|
|
183
188
|
|
|
189
|
+
## Browser Usage
|
|
190
|
+
Download the bundle file available as part of the release asset.
|
|
191
|
+
Include this bundle file in your browser html file and access `parseOffice` and `parseOfficeAsync` under the **`officeParser`** namespace.
|
|
192
|
+
|
|
193
|
+
**Example**
|
|
194
|
+
```html
|
|
195
|
+
<head>
|
|
196
|
+
...
|
|
197
|
+
<!-- Include bundle file in the script tag. -->
|
|
198
|
+
<script src="officeParserBundle@5.1.0.js"></script>
|
|
199
|
+
</head>
|
|
200
|
+
<body>
|
|
201
|
+
...
|
|
202
|
+
<input type="file" id="fileInput" />
|
|
203
|
+
...
|
|
204
|
+
<script>
|
|
205
|
+
document.getElementById('fileInput').addEventListener('change', async function(event) {
|
|
206
|
+
const outputDiv = document.getElementById('output');
|
|
207
|
+
const file = event.target.files[0];
|
|
208
|
+
try {
|
|
209
|
+
// Your configuration options for officeParser
|
|
210
|
+
const config = {
|
|
211
|
+
outputErrorToConsole: false,
|
|
212
|
+
newlineDelimiter: '\n',
|
|
213
|
+
ignoreNotes: false,
|
|
214
|
+
putNotesAtLast: false
|
|
215
|
+
};
|
|
216
|
+
|
|
217
|
+
const arrayBuffer = await file.arrayBuffer();
|
|
218
|
+
const result = await officeParser.parseOfficeAsync(arrayBuffer, config);
|
|
219
|
+
// result contains the extracted text.
|
|
220
|
+
}
|
|
221
|
+
catch (error) {
|
|
222
|
+
// Handle error
|
|
223
|
+
}
|
|
224
|
+
});
|
|
225
|
+
</script>
|
|
226
|
+
</body>
|
|
227
|
+
```
|
|
228
|
+
|
|
184
229
|
|
|
185
230
|
## Known Bugs
|
|
186
231
|
1. Inconsistency and incorrectness in the positioning of footnotes and endnotes in .docx files where the footnotes and endnotes would end up at the end of the parsed text whereas it would be positioned exactly after the referenced word in .odt files.
|
|
187
232
|
2. The charts and objects information of .odt files are not accurate and may end up showing a few NaN in some cases.
|
|
233
|
+
3. Extracting texts in browser bundles does not work for pdf files.
|
|
188
234
|
----------
|
|
189
235
|
|
|
190
236
|
**npm**
|