extract-hwp 0.1.0__tar.gz
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- extract_hwp-0.1.0/.gitignore +13 -0
- extract_hwp-0.1.0/.python-version +1 -0
- extract_hwp-0.1.0/CLAUDE.md +74 -0
- extract_hwp-0.1.0/LICENSE +29 -0
- extract_hwp-0.1.0/PKG-INFO +197 -0
- extract_hwp-0.1.0/README.md +169 -0
- extract_hwp-0.1.0/examples/README.md +174 -0
- extract_hwp-0.1.0/examples/basic_usage.py +144 -0
- extract_hwp-0.1.0/examples/batch_processing.py +183 -0
- extract_hwp-0.1.0/examples/cli_tool.py +212 -0
- extract_hwp-0.1.0/examples/protected_document.hwp +0 -0
- extract_hwp-0.1.0/examples/sample_document.hwp +0 -0
- extract_hwp-0.1.0/examples/sample_document.hwpx +0 -0
- extract_hwp-0.1.0/pyproject.toml +51 -0
- extract_hwp-0.1.0/src/__init__.py +1 -0
- extract_hwp-0.1.0/src/extract_hwp/__init__.py +28 -0
- extract_hwp-0.1.0/src/extract_hwp/core.py +60 -0
- extract_hwp-0.1.0/src/extract_hwp/hwp5.py +184 -0
- extract_hwp-0.1.0/src/extract_hwp/hwpx.py +63 -0
- extract_hwp-0.1.0/src/extract_hwp/password.py +89 -0
|
@@ -0,0 +1 @@
|
|
|
1
|
+
3.13
|
|
@@ -0,0 +1,74 @@
|
|
|
1
|
+
# CLAUDE.md
|
|
2
|
+
|
|
3
|
+
This file provides guidance to Claude Code (claude.ai/code) when working with code in this repository.
|
|
4
|
+
|
|
5
|
+
## Project Overview
|
|
6
|
+
|
|
7
|
+
This is a Python library for extracting text from Korean HWP (Hangul Word Processor) files, supporting both HWP 5.0 (OLE format) and HWPX (ZIP-based XML format) files. The project includes password protection detection and structured text extraction capabilities.
|
|
8
|
+
|
|
9
|
+
## Development Commands
|
|
10
|
+
|
|
11
|
+
### Package Management
|
|
12
|
+
- **Install dependencies**: `uv sync` (uses uv for dependency management)
|
|
13
|
+
- **Run main script**: `python main.py` (basic hello world entry point)
|
|
14
|
+
- **Run extraction**: `python -m extract_hwp` or import functions from `extract_hwp.py`
|
|
15
|
+
|
|
16
|
+
### Testing and Quality
|
|
17
|
+
No specific test commands are configured yet - tests would need to be added to pyproject.toml.
|
|
18
|
+
|
|
19
|
+
## Code Architecture
|
|
20
|
+
|
|
21
|
+
### Core Components
|
|
22
|
+
|
|
23
|
+
**extract_hwp.py**: Main extraction module with three key functions:
|
|
24
|
+
- `extract_text_from_hwp()`: Unified interface that routes to appropriate extractor based on file extension
|
|
25
|
+
- `extract_text_from_hwpx()`: HWPX file processor (ZIP-based XML format)
|
|
26
|
+
- `extract_text_from_hwp5()`: HWP 5.0 processor (OLE compound document format)
|
|
27
|
+
|
|
28
|
+
**Password Protection Detection**:
|
|
29
|
+
- `is_hwpx_password_protected()`: Checks HWPX files via META-INF/manifest.xml encryption data
|
|
30
|
+
- `is_hwp5_password_protected()`: Checks HWP 5.0 files via FileHeader stream bit flags
|
|
31
|
+
- `is_hwp_file_password_protected()`: Unified interface for both formats
|
|
32
|
+
|
|
33
|
+
### File Format Handling
|
|
34
|
+
|
|
35
|
+
**HWPX Files**:
|
|
36
|
+
- ZIP archives containing XML sections in `Contents/section*.xml`
|
|
37
|
+
- Text extracted from `<p>` (paragraph) elements containing `<t>` (text) nodes
|
|
38
|
+
- Preserves paragraph structure with newlines
|
|
39
|
+
|
|
40
|
+
**HWP 5.0 Files**:
|
|
41
|
+
- OLE compound documents with compressed BodyText streams
|
|
42
|
+
- Requires decompression (zlib) and binary parsing of structured records
|
|
43
|
+
- Text stored in PARA_TEXT records (tag_id: 67) as Unicode sequences
|
|
44
|
+
- Supports multiple sections (`BodyText/Section0`, `BodyText/Section1`, etc.)
|
|
45
|
+
|
|
46
|
+
### Dependencies
|
|
47
|
+
|
|
48
|
+
**External Libraries**:
|
|
49
|
+
- `olefile`: OLE compound document parsing for HWP 5.0 files
|
|
50
|
+
- `superclaude>=3.0.0.2`: Framework dependency (external utility framework)
|
|
51
|
+
|
|
52
|
+
**Framework Integration**:
|
|
53
|
+
The code imports from `core.util_text` and `core.util_extraction` modules, indicating this is part of a larger text extraction framework. These modules provide:
|
|
54
|
+
- Text cleaning/indexing utilities
|
|
55
|
+
- Error handling decorators (@handle_extraction_errors)
|
|
56
|
+
- File validation and logging utilities
|
|
57
|
+
|
|
58
|
+
### Error Handling Strategy
|
|
59
|
+
|
|
60
|
+
The codebase uses a defensive approach:
|
|
61
|
+
- Password-protected files are detected and skipped (not extracted)
|
|
62
|
+
- File validation occurs before processing
|
|
63
|
+
- Graceful degradation for parsing errors
|
|
64
|
+
- Comprehensive logging for troubleshooting
|
|
65
|
+
- Returns empty strings rather than throwing exceptions for most extraction failures
|
|
66
|
+
|
|
67
|
+
### Character Encoding Considerations
|
|
68
|
+
|
|
69
|
+
HWP 5.0 extraction includes specific Unicode range validation:
|
|
70
|
+
- Basic Latin (0x0020-0x007E)
|
|
71
|
+
- Korean syllables (0xAC00-0xD7AF)
|
|
72
|
+
- Korean Jamo (0x3130-0x318F)
|
|
73
|
+
- Full-width characters (0xFF00-0xFFEF)
|
|
74
|
+
- General punctuation (0x2000-0x206F)
|
|
@@ -0,0 +1,29 @@
|
|
|
1
|
+
BSD 3-Clause License
|
|
2
|
+
|
|
3
|
+
Copyright (c) 2025, thlee
|
|
4
|
+
All rights reserved.
|
|
5
|
+
|
|
6
|
+
Redistribution and use in source and binary forms, with or without
|
|
7
|
+
modification, are permitted provided that the following conditions are met:
|
|
8
|
+
|
|
9
|
+
1. Redistributions of source code must retain the above copyright notice, this
|
|
10
|
+
list of conditions and the following disclaimer.
|
|
11
|
+
|
|
12
|
+
2. Redistributions in binary form must reproduce the above copyright notice,
|
|
13
|
+
this list of conditions and the following disclaimer in the documentation
|
|
14
|
+
and/or other materials provided with the distribution.
|
|
15
|
+
|
|
16
|
+
3. Neither the name of the copyright holder nor the names of its
|
|
17
|
+
contributors may be used to endorse or promote products derived from
|
|
18
|
+
this software without specific prior written permission.
|
|
19
|
+
|
|
20
|
+
THIS SOFTWARE IS PROVIDED BY THE COPYRIGHT HOLDERS AND CONTRIBUTORS "AS IS"
|
|
21
|
+
AND ANY EXPRESS OR IMPLIED WARRANTIES, INCLUDING, BUT NOT LIMITED TO, THE
|
|
22
|
+
IMPLIED WARRANTIES OF MERCHANTABILITY AND FITNESS FOR A PARTICULAR PURPOSE ARE
|
|
23
|
+
DISCLAIMED. IN NO EVENT SHALL THE COPYRIGHT HOLDER OR CONTRIBUTORS BE LIABLE
|
|
24
|
+
FOR ANY DIRECT, INDIRECT, INCIDENTAL, SPECIAL, EXEMPLARY, OR CONSEQUENTIAL
|
|
25
|
+
DAMAGES (INCLUDING, BUT NOT LIMITED TO, PROCUREMENT OF SUBSTITUTE GOODS OR
|
|
26
|
+
SERVICES; LOSS OF USE, DATA, OR PROFITS; OR BUSINESS INTERRUPTION) HOWEVER
|
|
27
|
+
CAUSED AND ON ANY THEORY OF LIABILITY, WHETHER IN CONTRACT, STRICT LIABILITY,
|
|
28
|
+
OR TORT (INCLUDING NEGLIGENCE OR OTHERWISE) ARISING IN ANY WAY OUT OF THE USE
|
|
29
|
+
OF THIS SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
|
|
@@ -0,0 +1,197 @@
|
|
|
1
|
+
Metadata-Version: 2.4
|
|
2
|
+
Name: extract-hwp
|
|
3
|
+
Version: 0.1.0
|
|
4
|
+
Summary: Python library for extracting text from Korean HWP files (HWP 5.0 and HWPX formats)
|
|
5
|
+
Project-URL: Homepage, https://github.com/thlee/extract-hwp
|
|
6
|
+
Project-URL: Bug Reports, https://github.com/thlee/extract-hwp/issues
|
|
7
|
+
Project-URL: Source, https://github.com/thlee/extract-hwp
|
|
8
|
+
Author-email: extract-hwp <extract-hwp@example.com>
|
|
9
|
+
License: BSD-3-Clause
|
|
10
|
+
License-File: LICENSE
|
|
11
|
+
Keywords: document,hwp,hwpx,korean,text-extraction
|
|
12
|
+
Classifier: Development Status :: 4 - Beta
|
|
13
|
+
Classifier: Intended Audience :: Developers
|
|
14
|
+
Classifier: License :: OSI Approved :: BSD License
|
|
15
|
+
Classifier: Operating System :: OS Independent
|
|
16
|
+
Classifier: Programming Language :: Python :: 3
|
|
17
|
+
Classifier: Programming Language :: Python :: 3.8
|
|
18
|
+
Classifier: Programming Language :: Python :: 3.9
|
|
19
|
+
Classifier: Programming Language :: Python :: 3.10
|
|
20
|
+
Classifier: Programming Language :: Python :: 3.11
|
|
21
|
+
Classifier: Programming Language :: Python :: 3.12
|
|
22
|
+
Classifier: Programming Language :: Python :: 3.13
|
|
23
|
+
Classifier: Topic :: Software Development :: Libraries :: Python Modules
|
|
24
|
+
Classifier: Topic :: Text Processing
|
|
25
|
+
Requires-Python: >=3.8
|
|
26
|
+
Requires-Dist: olefile>=0.46
|
|
27
|
+
Description-Content-Type: text/markdown
|
|
28
|
+
|
|
29
|
+
# extract-hwp
|
|
30
|
+
|
|
31
|
+
한글과컴퓨터의 HWP 파일(HWP 5.0 및 HWPX 형식)에서 텍스트를 추출하는 Python 라이브러리입니다.
|
|
32
|
+
|
|
33
|
+
## 특징
|
|
34
|
+
|
|
35
|
+
- **다중 포맷 지원**: HWP 5.0 (OLE) 및 HWPX (ZIP/XML) 파일 모두 지원
|
|
36
|
+
- **암호화 파일 감지**: 처리하기 전에 암호로 보호된 파일을 감지
|
|
37
|
+
- **구조화된 추출**: 텍스트 추출 시 단락 구조 보존
|
|
38
|
+
- **견고한 오류 처리**: 손상되거나 잘못된 파일에 대한 방어적 처리
|
|
39
|
+
- **유니코드 지원**: 한글 및 다국어 텍스트 완전 지원
|
|
40
|
+
|
|
41
|
+
## 설치
|
|
42
|
+
|
|
43
|
+
```bash
|
|
44
|
+
pip install extract-hwp
|
|
45
|
+
```
|
|
46
|
+
|
|
47
|
+
## 사용법
|
|
48
|
+
|
|
49
|
+
### 기본 사용법
|
|
50
|
+
|
|
51
|
+
```python
|
|
52
|
+
from extract_hwp import extract_text_from_hwp
|
|
53
|
+
|
|
54
|
+
# HWP 또는 HWPX 파일에서 텍스트 추출
|
|
55
|
+
text, error = extract_text_from_hwp("document.hwp")
|
|
56
|
+
if error is None:
|
|
57
|
+
print(text)
|
|
58
|
+
else:
|
|
59
|
+
print(f"오류: {error}")
|
|
60
|
+
```
|
|
61
|
+
|
|
62
|
+
### 포맷별 추출
|
|
63
|
+
|
|
64
|
+
```python
|
|
65
|
+
from extract_hwp import extract_text_from_hwpx, extract_text_from_hwp5
|
|
66
|
+
|
|
67
|
+
# HWPX 파일 전용
|
|
68
|
+
hwpx_text = extract_text_from_hwpx("document.hwpx")
|
|
69
|
+
|
|
70
|
+
# HWP 5.0 파일 전용
|
|
71
|
+
hwp5_text = extract_text_from_hwp5("document.hwp")
|
|
72
|
+
```
|
|
73
|
+
|
|
74
|
+
### 암호화 파일 감지
|
|
75
|
+
|
|
76
|
+
```python
|
|
77
|
+
from extract_hwp import is_hwp_file_password_protected
|
|
78
|
+
|
|
79
|
+
if is_hwp_file_password_protected("document.hwp"):
|
|
80
|
+
print("파일이 암호로 보호되어 있습니다.")
|
|
81
|
+
else:
|
|
82
|
+
text, error = extract_text_from_hwp("document.hwp")
|
|
83
|
+
```
|
|
84
|
+
|
|
85
|
+
## API 참조
|
|
86
|
+
|
|
87
|
+
### 핵심 함수
|
|
88
|
+
|
|
89
|
+
#### `extract_text_from_hwp(filepath)`
|
|
90
|
+
|
|
91
|
+
HWP/HWPX 파일에서 텍스트를 추출합니다 (통합 인터페이스).
|
|
92
|
+
|
|
93
|
+
**매개변수:**
|
|
94
|
+
- `filepath` (str): HWP 또는 HWPX 파일 경로
|
|
95
|
+
|
|
96
|
+
**반환값:**
|
|
97
|
+
- `tuple`: (추출된_텍스트, 오류_메시지). 성공시 오류_메시지는 None
|
|
98
|
+
|
|
99
|
+
**예외:**
|
|
100
|
+
- `FileNotFoundError`: 파일을 찾을 수 없음
|
|
101
|
+
- `PermissionError`: 파일 접근 권한이 없음
|
|
102
|
+
- `ValueError`: 지원하지 않는 파일 형식
|
|
103
|
+
|
|
104
|
+
### 포맷별 함수
|
|
105
|
+
|
|
106
|
+
#### `extract_text_from_hwpx(hwpx_file_path)`
|
|
107
|
+
|
|
108
|
+
HWPX 파일에서 텍스트를 추출합니다.
|
|
109
|
+
|
|
110
|
+
**매개변수:**
|
|
111
|
+
- `hwpx_file_path` (str): HWPX 파일 경로
|
|
112
|
+
|
|
113
|
+
**반환값:**
|
|
114
|
+
- `str`: 추출된 텍스트 (오류 시 빈 문자열)
|
|
115
|
+
|
|
116
|
+
#### `extract_text_from_hwp5(filepath)`
|
|
117
|
+
|
|
118
|
+
HWP 5.0 (OLE) 파일에서 텍스트를 추출합니다.
|
|
119
|
+
|
|
120
|
+
**매개변수:**
|
|
121
|
+
- `filepath` (str): HWP 파일 경로
|
|
122
|
+
|
|
123
|
+
**반환값:**
|
|
124
|
+
- `str`: 추출된 텍스트 (오류 시 빈 문자열)
|
|
125
|
+
|
|
126
|
+
### 암호화 감지 함수
|
|
127
|
+
|
|
128
|
+
#### `is_hwp_file_password_protected(filepath)`
|
|
129
|
+
|
|
130
|
+
HWP/HWPX 파일이 암호로 보호되어 있는지 확인합니다.
|
|
131
|
+
|
|
132
|
+
**매개변수:**
|
|
133
|
+
- `filepath` (str): 확인할 파일 경로
|
|
134
|
+
|
|
135
|
+
**반환값:**
|
|
136
|
+
- `bool`: 암호로 보호된 경우 True, 그렇지 않으면 False
|
|
137
|
+
|
|
138
|
+
## 지원 포맷
|
|
139
|
+
|
|
140
|
+
### HWP 5.0 (OLE 포맷)
|
|
141
|
+
- 확장자: `.hwp`
|
|
142
|
+
- 구조: OLE 복합 문서 형식
|
|
143
|
+
- 압축: zlib 압축 지원
|
|
144
|
+
- 특징: 바이너리 구조 분석을 통한 텍스트 추출
|
|
145
|
+
|
|
146
|
+
### HWPX (ZIP/XML 포맷)
|
|
147
|
+
- 확장자: `.hwpx`
|
|
148
|
+
- 구조: XML 문서가 포함된 ZIP 아카이브
|
|
149
|
+
- 특징: 구조화된 텍스트 추출을 위한 XML 파싱
|
|
150
|
+
|
|
151
|
+
## 의존성
|
|
152
|
+
|
|
153
|
+
- `olefile>=0.46`: HWP 5.0 OLE 파일 처리
|
|
154
|
+
|
|
155
|
+
## 개발
|
|
156
|
+
|
|
157
|
+
### 개발 환경 설정
|
|
158
|
+
|
|
159
|
+
```bash
|
|
160
|
+
# 저장소 복제
|
|
161
|
+
git clone https://github.com/thlee/extract-hwp.git
|
|
162
|
+
cd extract-hwp
|
|
163
|
+
|
|
164
|
+
# 의존성 설치
|
|
165
|
+
uv sync
|
|
166
|
+
|
|
167
|
+
# 개발 의존성 포함 설치
|
|
168
|
+
uv sync --extra dev
|
|
169
|
+
```
|
|
170
|
+
|
|
171
|
+
### 테스트
|
|
172
|
+
|
|
173
|
+
```bash
|
|
174
|
+
# 테스트 실행
|
|
175
|
+
pytest
|
|
176
|
+
|
|
177
|
+
# 커버리지 포함
|
|
178
|
+
pytest --cov=src/extract_hwp
|
|
179
|
+
```
|
|
180
|
+
|
|
181
|
+
### 코드 품질
|
|
182
|
+
|
|
183
|
+
```bash
|
|
184
|
+
# 코드 포맷팅
|
|
185
|
+
black src/ tests/
|
|
186
|
+
|
|
187
|
+
# 타입 검사
|
|
188
|
+
mypy src/
|
|
189
|
+
```
|
|
190
|
+
|
|
191
|
+
## 라이선스
|
|
192
|
+
|
|
193
|
+
BSD 3-Clause License - 자세한 내용은 [LICENSE](LICENSE) 파일을 참조하세요.
|
|
194
|
+
|
|
195
|
+
## 변경사항
|
|
196
|
+
|
|
197
|
+
버전 히스토리는 [CHANGELOG.md](CHANGELOG.md)에서 확인할 수 있습니다.
|
|
@@ -0,0 +1,169 @@
|
|
|
1
|
+
# extract-hwp
|
|
2
|
+
|
|
3
|
+
한글과컴퓨터의 HWP 파일(HWP 5.0 및 HWPX 형식)에서 텍스트를 추출하는 Python 라이브러리입니다.
|
|
4
|
+
|
|
5
|
+
## 특징
|
|
6
|
+
|
|
7
|
+
- **다중 포맷 지원**: HWP 5.0 (OLE) 및 HWPX (ZIP/XML) 파일 모두 지원
|
|
8
|
+
- **암호화 파일 감지**: 처리하기 전에 암호로 보호된 파일을 감지
|
|
9
|
+
- **구조화된 추출**: 텍스트 추출 시 단락 구조 보존
|
|
10
|
+
- **견고한 오류 처리**: 손상되거나 잘못된 파일에 대한 방어적 처리
|
|
11
|
+
- **유니코드 지원**: 한글 및 다국어 텍스트 완전 지원
|
|
12
|
+
|
|
13
|
+
## 설치
|
|
14
|
+
|
|
15
|
+
```bash
|
|
16
|
+
pip install extract-hwp
|
|
17
|
+
```
|
|
18
|
+
|
|
19
|
+
## 사용법
|
|
20
|
+
|
|
21
|
+
### 기본 사용법
|
|
22
|
+
|
|
23
|
+
```python
|
|
24
|
+
from extract_hwp import extract_text_from_hwp
|
|
25
|
+
|
|
26
|
+
# HWP 또는 HWPX 파일에서 텍스트 추출
|
|
27
|
+
text, error = extract_text_from_hwp("document.hwp")
|
|
28
|
+
if error is None:
|
|
29
|
+
print(text)
|
|
30
|
+
else:
|
|
31
|
+
print(f"오류: {error}")
|
|
32
|
+
```
|
|
33
|
+
|
|
34
|
+
### 포맷별 추출
|
|
35
|
+
|
|
36
|
+
```python
|
|
37
|
+
from extract_hwp import extract_text_from_hwpx, extract_text_from_hwp5
|
|
38
|
+
|
|
39
|
+
# HWPX 파일 전용
|
|
40
|
+
hwpx_text = extract_text_from_hwpx("document.hwpx")
|
|
41
|
+
|
|
42
|
+
# HWP 5.0 파일 전용
|
|
43
|
+
hwp5_text = extract_text_from_hwp5("document.hwp")
|
|
44
|
+
```
|
|
45
|
+
|
|
46
|
+
### 암호화 파일 감지
|
|
47
|
+
|
|
48
|
+
```python
|
|
49
|
+
from extract_hwp import is_hwp_file_password_protected
|
|
50
|
+
|
|
51
|
+
if is_hwp_file_password_protected("document.hwp"):
|
|
52
|
+
print("파일이 암호로 보호되어 있습니다.")
|
|
53
|
+
else:
|
|
54
|
+
text, error = extract_text_from_hwp("document.hwp")
|
|
55
|
+
```
|
|
56
|
+
|
|
57
|
+
## API 참조
|
|
58
|
+
|
|
59
|
+
### 핵심 함수
|
|
60
|
+
|
|
61
|
+
#### `extract_text_from_hwp(filepath)`
|
|
62
|
+
|
|
63
|
+
HWP/HWPX 파일에서 텍스트를 추출합니다 (통합 인터페이스).
|
|
64
|
+
|
|
65
|
+
**매개변수:**
|
|
66
|
+
- `filepath` (str): HWP 또는 HWPX 파일 경로
|
|
67
|
+
|
|
68
|
+
**반환값:**
|
|
69
|
+
- `tuple`: (추출된_텍스트, 오류_메시지). 성공시 오류_메시지는 None
|
|
70
|
+
|
|
71
|
+
**예외:**
|
|
72
|
+
- `FileNotFoundError`: 파일을 찾을 수 없음
|
|
73
|
+
- `PermissionError`: 파일 접근 권한이 없음
|
|
74
|
+
- `ValueError`: 지원하지 않는 파일 형식
|
|
75
|
+
|
|
76
|
+
### 포맷별 함수
|
|
77
|
+
|
|
78
|
+
#### `extract_text_from_hwpx(hwpx_file_path)`
|
|
79
|
+
|
|
80
|
+
HWPX 파일에서 텍스트를 추출합니다.
|
|
81
|
+
|
|
82
|
+
**매개변수:**
|
|
83
|
+
- `hwpx_file_path` (str): HWPX 파일 경로
|
|
84
|
+
|
|
85
|
+
**반환값:**
|
|
86
|
+
- `str`: 추출된 텍스트 (오류 시 빈 문자열)
|
|
87
|
+
|
|
88
|
+
#### `extract_text_from_hwp5(filepath)`
|
|
89
|
+
|
|
90
|
+
HWP 5.0 (OLE) 파일에서 텍스트를 추출합니다.
|
|
91
|
+
|
|
92
|
+
**매개변수:**
|
|
93
|
+
- `filepath` (str): HWP 파일 경로
|
|
94
|
+
|
|
95
|
+
**반환값:**
|
|
96
|
+
- `str`: 추출된 텍스트 (오류 시 빈 문자열)
|
|
97
|
+
|
|
98
|
+
### 암호화 감지 함수
|
|
99
|
+
|
|
100
|
+
#### `is_hwp_file_password_protected(filepath)`
|
|
101
|
+
|
|
102
|
+
HWP/HWPX 파일이 암호로 보호되어 있는지 확인합니다.
|
|
103
|
+
|
|
104
|
+
**매개변수:**
|
|
105
|
+
- `filepath` (str): 확인할 파일 경로
|
|
106
|
+
|
|
107
|
+
**반환값:**
|
|
108
|
+
- `bool`: 암호로 보호된 경우 True, 그렇지 않으면 False
|
|
109
|
+
|
|
110
|
+
## 지원 포맷
|
|
111
|
+
|
|
112
|
+
### HWP 5.0 (OLE 포맷)
|
|
113
|
+
- 확장자: `.hwp`
|
|
114
|
+
- 구조: OLE 복합 문서 형식
|
|
115
|
+
- 압축: zlib 압축 지원
|
|
116
|
+
- 특징: 바이너리 구조 분석을 통한 텍스트 추출
|
|
117
|
+
|
|
118
|
+
### HWPX (ZIP/XML 포맷)
|
|
119
|
+
- 확장자: `.hwpx`
|
|
120
|
+
- 구조: XML 문서가 포함된 ZIP 아카이브
|
|
121
|
+
- 특징: 구조화된 텍스트 추출을 위한 XML 파싱
|
|
122
|
+
|
|
123
|
+
## 의존성
|
|
124
|
+
|
|
125
|
+
- `olefile>=0.46`: HWP 5.0 OLE 파일 처리
|
|
126
|
+
|
|
127
|
+
## 개발
|
|
128
|
+
|
|
129
|
+
### 개발 환경 설정
|
|
130
|
+
|
|
131
|
+
```bash
|
|
132
|
+
# 저장소 복제
|
|
133
|
+
git clone https://github.com/thlee/extract-hwp.git
|
|
134
|
+
cd extract-hwp
|
|
135
|
+
|
|
136
|
+
# 의존성 설치
|
|
137
|
+
uv sync
|
|
138
|
+
|
|
139
|
+
# 개발 의존성 포함 설치
|
|
140
|
+
uv sync --extra dev
|
|
141
|
+
```
|
|
142
|
+
|
|
143
|
+
### 테스트
|
|
144
|
+
|
|
145
|
+
```bash
|
|
146
|
+
# 테스트 실행
|
|
147
|
+
pytest
|
|
148
|
+
|
|
149
|
+
# 커버리지 포함
|
|
150
|
+
pytest --cov=src/extract_hwp
|
|
151
|
+
```
|
|
152
|
+
|
|
153
|
+
### 코드 품질
|
|
154
|
+
|
|
155
|
+
```bash
|
|
156
|
+
# 코드 포맷팅
|
|
157
|
+
black src/ tests/
|
|
158
|
+
|
|
159
|
+
# 타입 검사
|
|
160
|
+
mypy src/
|
|
161
|
+
```
|
|
162
|
+
|
|
163
|
+
## 라이선스
|
|
164
|
+
|
|
165
|
+
BSD 3-Clause License - 자세한 내용은 [LICENSE](LICENSE) 파일을 참조하세요.
|
|
166
|
+
|
|
167
|
+
## 변경사항
|
|
168
|
+
|
|
169
|
+
버전 히스토리는 [CHANGELOG.md](CHANGELOG.md)에서 확인할 수 있습니다.
|
|
@@ -0,0 +1,174 @@
|
|
|
1
|
+
# extract-hwp 사용 예제
|
|
2
|
+
|
|
3
|
+
이 디렉토리에는 extract-hwp 라이브러리를 사용하는 다양한 예제들이 포함되어 있습니다.
|
|
4
|
+
|
|
5
|
+
## 예제 파일들
|
|
6
|
+
|
|
7
|
+
### 1. `basic_usage.py` - 기본 사용법
|
|
8
|
+
라이브러리의 기본적인 사용 방법을 보여주는 예제입니다.
|
|
9
|
+
|
|
10
|
+
```bash
|
|
11
|
+
python basic_usage.py
|
|
12
|
+
```
|
|
13
|
+
|
|
14
|
+
**주요 기능:**
|
|
15
|
+
- 단일 파일 텍스트 추출
|
|
16
|
+
- 암호화 파일 감지
|
|
17
|
+
- 오류 처리 방법
|
|
18
|
+
- 포맷별 추출 함수 사용법
|
|
19
|
+
|
|
20
|
+
### 2. `batch_processing.py` - 일괄 처리
|
|
21
|
+
디렉토리 내의 모든 HWP 파일을 일괄 처리하는 예제입니다.
|
|
22
|
+
|
|
23
|
+
```bash
|
|
24
|
+
python batch_processing.py <디렉토리> [출력형식]
|
|
25
|
+
```
|
|
26
|
+
|
|
27
|
+
**사용 예제:**
|
|
28
|
+
```bash
|
|
29
|
+
# 현재 디렉토리의 모든 HWP 파일을 텍스트 파일로 저장
|
|
30
|
+
python batch_processing.py ./documents txt
|
|
31
|
+
|
|
32
|
+
# JSON 형식으로 결과 저장
|
|
33
|
+
python batch_processing.py ./documents json
|
|
34
|
+
```
|
|
35
|
+
|
|
36
|
+
**주요 기능:**
|
|
37
|
+
- 재귀적 파일 검색 (하위 디렉토리 포함)
|
|
38
|
+
- 일괄 처리 진행 상황 표시
|
|
39
|
+
- 처리 결과 통계
|
|
40
|
+
- 다양한 출력 형식 지원
|
|
41
|
+
|
|
42
|
+
### 3. `cli_tool.py` - 명령줄 도구
|
|
43
|
+
HWP 파일 처리를 위한 완전한 명령줄 인터페이스입니다.
|
|
44
|
+
|
|
45
|
+
```bash
|
|
46
|
+
python cli_tool.py <파일경로> [옵션]
|
|
47
|
+
```
|
|
48
|
+
|
|
49
|
+
**사용 예제:**
|
|
50
|
+
```bash
|
|
51
|
+
# 기본 사용 (표준 출력)
|
|
52
|
+
python cli_tool.py document.hwp
|
|
53
|
+
|
|
54
|
+
# 파일로 저장
|
|
55
|
+
python cli_tool.py document.hwp --output extracted.txt
|
|
56
|
+
|
|
57
|
+
# 암호화 여부만 확인
|
|
58
|
+
python cli_tool.py document.hwp --check-password
|
|
59
|
+
|
|
60
|
+
# 조용한 모드로 실행
|
|
61
|
+
python cli_tool.py document.hwpx --quiet --output-dir ./outputs
|
|
62
|
+
```
|
|
63
|
+
|
|
64
|
+
**주요 옵션:**
|
|
65
|
+
- `-o, --output`: 출력 파일 지정
|
|
66
|
+
- `--output-dir`: 출력 디렉토리 지정
|
|
67
|
+
- `-c, --check-password`: 암호화 여부만 확인
|
|
68
|
+
- `-q, --quiet`: 진행 메시지 숨기기
|
|
69
|
+
- `--encoding`: 출력 파일 인코딩 지정
|
|
70
|
+
|
|
71
|
+
## 실행 방법
|
|
72
|
+
|
|
73
|
+
### 개발 환경에서 실행
|
|
74
|
+
```bash
|
|
75
|
+
# extract-hwp 프로젝트 루트 디렉토리에서
|
|
76
|
+
cd examples
|
|
77
|
+
python basic_usage.py
|
|
78
|
+
python batch_processing.py ./test_files
|
|
79
|
+
python cli_tool.py test.hwp
|
|
80
|
+
```
|
|
81
|
+
|
|
82
|
+
### 설치된 패키지로 실행
|
|
83
|
+
패키지가 설치된 상태에서는 import 경로 수정이 필요합니다:
|
|
84
|
+
|
|
85
|
+
```python
|
|
86
|
+
# 예제 파일 상단의 다음 부분을 제거하거나 주석 처리
|
|
87
|
+
# sys.path.insert(0, str(Path(__file__).parent.parent / "src"))
|
|
88
|
+
```
|
|
89
|
+
|
|
90
|
+
## 테스트 파일 준비
|
|
91
|
+
|
|
92
|
+
예제를 실제로 테스트해보려면 HWP/HWPX 파일이 필요합니다:
|
|
93
|
+
|
|
94
|
+
1. **샘플 파일 생성**: 한글(HWP)에서 간단한 문서를 작성하고 `.hwp` 및 `.hwpx` 형식으로 저장
|
|
95
|
+
2. **테스트 디렉토리**: `./test_files/` 디렉토리를 만들고 샘플 파일들을 넣기
|
|
96
|
+
3. **암호화 파일**: 암호가 설정된 HWP 파일도 테스트용으로 준비
|
|
97
|
+
|
|
98
|
+
## 고급 사용법
|
|
99
|
+
|
|
100
|
+
### 커스텀 처리 로직
|
|
101
|
+
```python
|
|
102
|
+
from extract_hwp import extract_text_from_hwp, is_hwp_file_password_protected
|
|
103
|
+
|
|
104
|
+
def custom_processor(file_path):
|
|
105
|
+
"""커스텀 HWP 파일 처리기"""
|
|
106
|
+
|
|
107
|
+
# 1. 암호화 확인
|
|
108
|
+
if is_hwp_file_password_protected(file_path):
|
|
109
|
+
return None, "암호화된 파일"
|
|
110
|
+
|
|
111
|
+
# 2. 텍스트 추출
|
|
112
|
+
text, error = extract_text_from_hwp(file_path)
|
|
113
|
+
if error:
|
|
114
|
+
return None, error
|
|
115
|
+
|
|
116
|
+
# 3. 후처리 (예: 텍스트 정리)
|
|
117
|
+
processed_text = text.strip()
|
|
118
|
+
processed_text = '\n'.join(line.strip() for line in processed_text.split('\n') if line.strip())
|
|
119
|
+
|
|
120
|
+
return processed_text, None
|
|
121
|
+
```
|
|
122
|
+
|
|
123
|
+
### 에러 처리 패턴
|
|
124
|
+
```python
|
|
125
|
+
import logging
|
|
126
|
+
|
|
127
|
+
def safe_extract(file_path):
|
|
128
|
+
"""안전한 텍스트 추출 (로깅 포함)"""
|
|
129
|
+
try:
|
|
130
|
+
text, error = extract_text_from_hwp(file_path)
|
|
131
|
+
if error:
|
|
132
|
+
logging.warning(f"추출 실패: {file_path} - {error}")
|
|
133
|
+
return ""
|
|
134
|
+
|
|
135
|
+
logging.info(f"추출 성공: {file_path} ({len(text)} 문자)")
|
|
136
|
+
return text
|
|
137
|
+
|
|
138
|
+
except Exception as e:
|
|
139
|
+
logging.error(f"예외 발생: {file_path} - {e}")
|
|
140
|
+
return ""
|
|
141
|
+
```
|
|
142
|
+
|
|
143
|
+
## 성능 최적화 팁
|
|
144
|
+
|
|
145
|
+
1. **대용량 파일**: 매우 큰 HWP 파일의 경우 메모리 사용량에 주의
|
|
146
|
+
2. **일괄 처리**: 많은 파일을 처리할 때는 멀티프로세싱 사용 고려
|
|
147
|
+
3. **암호화 확인**: 처리 전에 암호화 여부를 먼저 확인하여 불필요한 처리 방지
|
|
148
|
+
|
|
149
|
+
## 문제 해결
|
|
150
|
+
|
|
151
|
+
### 일반적인 문제들
|
|
152
|
+
|
|
153
|
+
1. **ImportError**: 패키지가 설치되지 않은 경우
|
|
154
|
+
```bash
|
|
155
|
+
pip install extract-hwp
|
|
156
|
+
```
|
|
157
|
+
|
|
158
|
+
2. **파일을 찾을 수 없음**: 파일 경로 확인
|
|
159
|
+
```python
|
|
160
|
+
import os
|
|
161
|
+
print(os.path.exists("your_file.hwp"))
|
|
162
|
+
```
|
|
163
|
+
|
|
164
|
+
3. **인코딩 오류**: 파일명에 특수문자가 있는 경우
|
|
165
|
+
```python
|
|
166
|
+
file_path = file_path.encode('utf-8').decode('utf-8')
|
|
167
|
+
```
|
|
168
|
+
|
|
169
|
+
4. **권한 오류**: 파일 읽기 권한 확인
|
|
170
|
+
```bash
|
|
171
|
+
ls -la your_file.hwp
|
|
172
|
+
```
|
|
173
|
+
|
|
174
|
+
더 자세한 정보는 프로젝트의 메인 README.md를 참조하세요.
|