extract-hwp 0.1.0__tar.gz

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -0,0 +1,13 @@
1
+ # Python-generated files
2
+ __pycache__/
3
+ *.py[oc]
4
+ build/
5
+ dist/
6
+ wheels/
7
+ *.egg-info
8
+
9
+ # Virtual environments
10
+ .venv
11
+
12
+ # UV
13
+ uv.lock
@@ -0,0 +1 @@
1
+ 3.13
@@ -0,0 +1,74 @@
1
+ # CLAUDE.md
2
+
3
+ This file provides guidance to Claude Code (claude.ai/code) when working with code in this repository.
4
+
5
+ ## Project Overview
6
+
7
+ This is a Python library for extracting text from Korean HWP (Hangul Word Processor) files, supporting both HWP 5.0 (OLE format) and HWPX (ZIP-based XML format) files. The project includes password protection detection and structured text extraction capabilities.
8
+
9
+ ## Development Commands
10
+
11
+ ### Package Management
12
+ - **Install dependencies**: `uv sync` (uses uv for dependency management)
13
+ - **Run main script**: `python main.py` (basic hello world entry point)
14
+ - **Run extraction**: `python -m extract_hwp` or import functions from `extract_hwp.py`
15
+
16
+ ### Testing and Quality
17
+ No specific test commands are configured yet - tests would need to be added to pyproject.toml.
18
+
19
+ ## Code Architecture
20
+
21
+ ### Core Components
22
+
23
+ **extract_hwp.py**: Main extraction module with three key functions:
24
+ - `extract_text_from_hwp()`: Unified interface that routes to appropriate extractor based on file extension
25
+ - `extract_text_from_hwpx()`: HWPX file processor (ZIP-based XML format)
26
+ - `extract_text_from_hwp5()`: HWP 5.0 processor (OLE compound document format)
27
+
28
+ **Password Protection Detection**:
29
+ - `is_hwpx_password_protected()`: Checks HWPX files via META-INF/manifest.xml encryption data
30
+ - `is_hwp5_password_protected()`: Checks HWP 5.0 files via FileHeader stream bit flags
31
+ - `is_hwp_file_password_protected()`: Unified interface for both formats
32
+
33
+ ### File Format Handling
34
+
35
+ **HWPX Files**:
36
+ - ZIP archives containing XML sections in `Contents/section*.xml`
37
+ - Text extracted from `<p>` (paragraph) elements containing `<t>` (text) nodes
38
+ - Preserves paragraph structure with newlines
39
+
40
+ **HWP 5.0 Files**:
41
+ - OLE compound documents with compressed BodyText streams
42
+ - Requires decompression (zlib) and binary parsing of structured records
43
+ - Text stored in PARA_TEXT records (tag_id: 67) as Unicode sequences
44
+ - Supports multiple sections (`BodyText/Section0`, `BodyText/Section1`, etc.)
45
+
46
+ ### Dependencies
47
+
48
+ **External Libraries**:
49
+ - `olefile`: OLE compound document parsing for HWP 5.0 files
50
+ - `superclaude>=3.0.0.2`: Framework dependency (external utility framework)
51
+
52
+ **Framework Integration**:
53
+ The code imports from `core.util_text` and `core.util_extraction` modules, indicating this is part of a larger text extraction framework. These modules provide:
54
+ - Text cleaning/indexing utilities
55
+ - Error handling decorators (@handle_extraction_errors)
56
+ - File validation and logging utilities
57
+
58
+ ### Error Handling Strategy
59
+
60
+ The codebase uses a defensive approach:
61
+ - Password-protected files are detected and skipped (not extracted)
62
+ - File validation occurs before processing
63
+ - Graceful degradation for parsing errors
64
+ - Comprehensive logging for troubleshooting
65
+ - Returns empty strings rather than throwing exceptions for most extraction failures
66
+
67
+ ### Character Encoding Considerations
68
+
69
+ HWP 5.0 extraction includes specific Unicode range validation:
70
+ - Basic Latin (0x0020-0x007E)
71
+ - Korean syllables (0xAC00-0xD7AF)
72
+ - Korean Jamo (0x3130-0x318F)
73
+ - Full-width characters (0xFF00-0xFFEF)
74
+ - General punctuation (0x2000-0x206F)
@@ -0,0 +1,29 @@
1
+ BSD 3-Clause License
2
+
3
+ Copyright (c) 2025, thlee
4
+ All rights reserved.
5
+
6
+ Redistribution and use in source and binary forms, with or without
7
+ modification, are permitted provided that the following conditions are met:
8
+
9
+ 1. Redistributions of source code must retain the above copyright notice, this
10
+ list of conditions and the following disclaimer.
11
+
12
+ 2. Redistributions in binary form must reproduce the above copyright notice,
13
+ this list of conditions and the following disclaimer in the documentation
14
+ and/or other materials provided with the distribution.
15
+
16
+ 3. Neither the name of the copyright holder nor the names of its
17
+ contributors may be used to endorse or promote products derived from
18
+ this software without specific prior written permission.
19
+
20
+ THIS SOFTWARE IS PROVIDED BY THE COPYRIGHT HOLDERS AND CONTRIBUTORS "AS IS"
21
+ AND ANY EXPRESS OR IMPLIED WARRANTIES, INCLUDING, BUT NOT LIMITED TO, THE
22
+ IMPLIED WARRANTIES OF MERCHANTABILITY AND FITNESS FOR A PARTICULAR PURPOSE ARE
23
+ DISCLAIMED. IN NO EVENT SHALL THE COPYRIGHT HOLDER OR CONTRIBUTORS BE LIABLE
24
+ FOR ANY DIRECT, INDIRECT, INCIDENTAL, SPECIAL, EXEMPLARY, OR CONSEQUENTIAL
25
+ DAMAGES (INCLUDING, BUT NOT LIMITED TO, PROCUREMENT OF SUBSTITUTE GOODS OR
26
+ SERVICES; LOSS OF USE, DATA, OR PROFITS; OR BUSINESS INTERRUPTION) HOWEVER
27
+ CAUSED AND ON ANY THEORY OF LIABILITY, WHETHER IN CONTRACT, STRICT LIABILITY,
28
+ OR TORT (INCLUDING NEGLIGENCE OR OTHERWISE) ARISING IN ANY WAY OUT OF THE USE
29
+ OF THIS SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
@@ -0,0 +1,197 @@
1
+ Metadata-Version: 2.4
2
+ Name: extract-hwp
3
+ Version: 0.1.0
4
+ Summary: Python library for extracting text from Korean HWP files (HWP 5.0 and HWPX formats)
5
+ Project-URL: Homepage, https://github.com/thlee/extract-hwp
6
+ Project-URL: Bug Reports, https://github.com/thlee/extract-hwp/issues
7
+ Project-URL: Source, https://github.com/thlee/extract-hwp
8
+ Author-email: extract-hwp <extract-hwp@example.com>
9
+ License: BSD-3-Clause
10
+ License-File: LICENSE
11
+ Keywords: document,hwp,hwpx,korean,text-extraction
12
+ Classifier: Development Status :: 4 - Beta
13
+ Classifier: Intended Audience :: Developers
14
+ Classifier: License :: OSI Approved :: BSD License
15
+ Classifier: Operating System :: OS Independent
16
+ Classifier: Programming Language :: Python :: 3
17
+ Classifier: Programming Language :: Python :: 3.8
18
+ Classifier: Programming Language :: Python :: 3.9
19
+ Classifier: Programming Language :: Python :: 3.10
20
+ Classifier: Programming Language :: Python :: 3.11
21
+ Classifier: Programming Language :: Python :: 3.12
22
+ Classifier: Programming Language :: Python :: 3.13
23
+ Classifier: Topic :: Software Development :: Libraries :: Python Modules
24
+ Classifier: Topic :: Text Processing
25
+ Requires-Python: >=3.8
26
+ Requires-Dist: olefile>=0.46
27
+ Description-Content-Type: text/markdown
28
+
29
+ # extract-hwp
30
+
31
+ 한글과컴퓨터의 HWP 파일(HWP 5.0 및 HWPX 형식)에서 텍스트를 추출하는 Python 라이브러리입니다.
32
+
33
+ ## 특징
34
+
35
+ - **다중 포맷 지원**: HWP 5.0 (OLE) 및 HWPX (ZIP/XML) 파일 모두 지원
36
+ - **암호화 파일 감지**: 처리하기 전에 암호로 보호된 파일을 감지
37
+ - **구조화된 추출**: 텍스트 추출 시 단락 구조 보존
38
+ - **견고한 오류 처리**: 손상되거나 잘못된 파일에 대한 방어적 처리
39
+ - **유니코드 지원**: 한글 및 다국어 텍스트 완전 지원
40
+
41
+ ## 설치
42
+
43
+ ```bash
44
+ pip install extract-hwp
45
+ ```
46
+
47
+ ## 사용법
48
+
49
+ ### 기본 사용법
50
+
51
+ ```python
52
+ from extract_hwp import extract_text_from_hwp
53
+
54
+ # HWP 또는 HWPX 파일에서 텍스트 추출
55
+ text, error = extract_text_from_hwp("document.hwp")
56
+ if error is None:
57
+ print(text)
58
+ else:
59
+ print(f"오류: {error}")
60
+ ```
61
+
62
+ ### 포맷별 추출
63
+
64
+ ```python
65
+ from extract_hwp import extract_text_from_hwpx, extract_text_from_hwp5
66
+
67
+ # HWPX 파일 전용
68
+ hwpx_text = extract_text_from_hwpx("document.hwpx")
69
+
70
+ # HWP 5.0 파일 전용
71
+ hwp5_text = extract_text_from_hwp5("document.hwp")
72
+ ```
73
+
74
+ ### 암호화 파일 감지
75
+
76
+ ```python
77
+ from extract_hwp import is_hwp_file_password_protected
78
+
79
+ if is_hwp_file_password_protected("document.hwp"):
80
+ print("파일이 암호로 보호되어 있습니다.")
81
+ else:
82
+ text, error = extract_text_from_hwp("document.hwp")
83
+ ```
84
+
85
+ ## API 참조
86
+
87
+ ### 핵심 함수
88
+
89
+ #### `extract_text_from_hwp(filepath)`
90
+
91
+ HWP/HWPX 파일에서 텍스트를 추출합니다 (통합 인터페이스).
92
+
93
+ **매개변수:**
94
+ - `filepath` (str): HWP 또는 HWPX 파일 경로
95
+
96
+ **반환값:**
97
+ - `tuple`: (추출된_텍스트, 오류_메시지). 성공시 오류_메시지는 None
98
+
99
+ **예외:**
100
+ - `FileNotFoundError`: 파일을 찾을 수 없음
101
+ - `PermissionError`: 파일 접근 권한이 없음
102
+ - `ValueError`: 지원하지 않는 파일 형식
103
+
104
+ ### 포맷별 함수
105
+
106
+ #### `extract_text_from_hwpx(hwpx_file_path)`
107
+
108
+ HWPX 파일에서 텍스트를 추출합니다.
109
+
110
+ **매개변수:**
111
+ - `hwpx_file_path` (str): HWPX 파일 경로
112
+
113
+ **반환값:**
114
+ - `str`: 추출된 텍스트 (오류 시 빈 문자열)
115
+
116
+ #### `extract_text_from_hwp5(filepath)`
117
+
118
+ HWP 5.0 (OLE) 파일에서 텍스트를 추출합니다.
119
+
120
+ **매개변수:**
121
+ - `filepath` (str): HWP 파일 경로
122
+
123
+ **반환값:**
124
+ - `str`: 추출된 텍스트 (오류 시 빈 문자열)
125
+
126
+ ### 암호화 감지 함수
127
+
128
+ #### `is_hwp_file_password_protected(filepath)`
129
+
130
+ HWP/HWPX 파일이 암호로 보호되어 있는지 확인합니다.
131
+
132
+ **매개변수:**
133
+ - `filepath` (str): 확인할 파일 경로
134
+
135
+ **반환값:**
136
+ - `bool`: 암호로 보호된 경우 True, 그렇지 않으면 False
137
+
138
+ ## 지원 포맷
139
+
140
+ ### HWP 5.0 (OLE 포맷)
141
+ - 확장자: `.hwp`
142
+ - 구조: OLE 복합 문서 형식
143
+ - 압축: zlib 압축 지원
144
+ - 특징: 바이너리 구조 분석을 통한 텍스트 추출
145
+
146
+ ### HWPX (ZIP/XML 포맷)
147
+ - 확장자: `.hwpx`
148
+ - 구조: XML 문서가 포함된 ZIP 아카이브
149
+ - 특징: 구조화된 텍스트 추출을 위한 XML 파싱
150
+
151
+ ## 의존성
152
+
153
+ - `olefile>=0.46`: HWP 5.0 OLE 파일 처리
154
+
155
+ ## 개발
156
+
157
+ ### 개발 환경 설정
158
+
159
+ ```bash
160
+ # 저장소 복제
161
+ git clone https://github.com/thlee/extract-hwp.git
162
+ cd extract-hwp
163
+
164
+ # 의존성 설치
165
+ uv sync
166
+
167
+ # 개발 의존성 포함 설치
168
+ uv sync --extra dev
169
+ ```
170
+
171
+ ### 테스트
172
+
173
+ ```bash
174
+ # 테스트 실행
175
+ pytest
176
+
177
+ # 커버리지 포함
178
+ pytest --cov=src/extract_hwp
179
+ ```
180
+
181
+ ### 코드 품질
182
+
183
+ ```bash
184
+ # 코드 포맷팅
185
+ black src/ tests/
186
+
187
+ # 타입 검사
188
+ mypy src/
189
+ ```
190
+
191
+ ## 라이선스
192
+
193
+ BSD 3-Clause License - 자세한 내용은 [LICENSE](LICENSE) 파일을 참조하세요.
194
+
195
+ ## 변경사항
196
+
197
+ 버전 히스토리는 [CHANGELOG.md](CHANGELOG.md)에서 확인할 수 있습니다.
@@ -0,0 +1,169 @@
1
+ # extract-hwp
2
+
3
+ 한글과컴퓨터의 HWP 파일(HWP 5.0 및 HWPX 형식)에서 텍스트를 추출하는 Python 라이브러리입니다.
4
+
5
+ ## 특징
6
+
7
+ - **다중 포맷 지원**: HWP 5.0 (OLE) 및 HWPX (ZIP/XML) 파일 모두 지원
8
+ - **암호화 파일 감지**: 처리하기 전에 암호로 보호된 파일을 감지
9
+ - **구조화된 추출**: 텍스트 추출 시 단락 구조 보존
10
+ - **견고한 오류 처리**: 손상되거나 잘못된 파일에 대한 방어적 처리
11
+ - **유니코드 지원**: 한글 및 다국어 텍스트 완전 지원
12
+
13
+ ## 설치
14
+
15
+ ```bash
16
+ pip install extract-hwp
17
+ ```
18
+
19
+ ## 사용법
20
+
21
+ ### 기본 사용법
22
+
23
+ ```python
24
+ from extract_hwp import extract_text_from_hwp
25
+
26
+ # HWP 또는 HWPX 파일에서 텍스트 추출
27
+ text, error = extract_text_from_hwp("document.hwp")
28
+ if error is None:
29
+ print(text)
30
+ else:
31
+ print(f"오류: {error}")
32
+ ```
33
+
34
+ ### 포맷별 추출
35
+
36
+ ```python
37
+ from extract_hwp import extract_text_from_hwpx, extract_text_from_hwp5
38
+
39
+ # HWPX 파일 전용
40
+ hwpx_text = extract_text_from_hwpx("document.hwpx")
41
+
42
+ # HWP 5.0 파일 전용
43
+ hwp5_text = extract_text_from_hwp5("document.hwp")
44
+ ```
45
+
46
+ ### 암호화 파일 감지
47
+
48
+ ```python
49
+ from extract_hwp import is_hwp_file_password_protected
50
+
51
+ if is_hwp_file_password_protected("document.hwp"):
52
+ print("파일이 암호로 보호되어 있습니다.")
53
+ else:
54
+ text, error = extract_text_from_hwp("document.hwp")
55
+ ```
56
+
57
+ ## API 참조
58
+
59
+ ### 핵심 함수
60
+
61
+ #### `extract_text_from_hwp(filepath)`
62
+
63
+ HWP/HWPX 파일에서 텍스트를 추출합니다 (통합 인터페이스).
64
+
65
+ **매개변수:**
66
+ - `filepath` (str): HWP 또는 HWPX 파일 경로
67
+
68
+ **반환값:**
69
+ - `tuple`: (추출된_텍스트, 오류_메시지). 성공시 오류_메시지는 None
70
+
71
+ **예외:**
72
+ - `FileNotFoundError`: 파일을 찾을 수 없음
73
+ - `PermissionError`: 파일 접근 권한이 없음
74
+ - `ValueError`: 지원하지 않는 파일 형식
75
+
76
+ ### 포맷별 함수
77
+
78
+ #### `extract_text_from_hwpx(hwpx_file_path)`
79
+
80
+ HWPX 파일에서 텍스트를 추출합니다.
81
+
82
+ **매개변수:**
83
+ - `hwpx_file_path` (str): HWPX 파일 경로
84
+
85
+ **반환값:**
86
+ - `str`: 추출된 텍스트 (오류 시 빈 문자열)
87
+
88
+ #### `extract_text_from_hwp5(filepath)`
89
+
90
+ HWP 5.0 (OLE) 파일에서 텍스트를 추출합니다.
91
+
92
+ **매개변수:**
93
+ - `filepath` (str): HWP 파일 경로
94
+
95
+ **반환값:**
96
+ - `str`: 추출된 텍스트 (오류 시 빈 문자열)
97
+
98
+ ### 암호화 감지 함수
99
+
100
+ #### `is_hwp_file_password_protected(filepath)`
101
+
102
+ HWP/HWPX 파일이 암호로 보호되어 있는지 확인합니다.
103
+
104
+ **매개변수:**
105
+ - `filepath` (str): 확인할 파일 경로
106
+
107
+ **반환값:**
108
+ - `bool`: 암호로 보호된 경우 True, 그렇지 않으면 False
109
+
110
+ ## 지원 포맷
111
+
112
+ ### HWP 5.0 (OLE 포맷)
113
+ - 확장자: `.hwp`
114
+ - 구조: OLE 복합 문서 형식
115
+ - 압축: zlib 압축 지원
116
+ - 특징: 바이너리 구조 분석을 통한 텍스트 추출
117
+
118
+ ### HWPX (ZIP/XML 포맷)
119
+ - 확장자: `.hwpx`
120
+ - 구조: XML 문서가 포함된 ZIP 아카이브
121
+ - 특징: 구조화된 텍스트 추출을 위한 XML 파싱
122
+
123
+ ## 의존성
124
+
125
+ - `olefile>=0.46`: HWP 5.0 OLE 파일 처리
126
+
127
+ ## 개발
128
+
129
+ ### 개발 환경 설정
130
+
131
+ ```bash
132
+ # 저장소 복제
133
+ git clone https://github.com/thlee/extract-hwp.git
134
+ cd extract-hwp
135
+
136
+ # 의존성 설치
137
+ uv sync
138
+
139
+ # 개발 의존성 포함 설치
140
+ uv sync --extra dev
141
+ ```
142
+
143
+ ### 테스트
144
+
145
+ ```bash
146
+ # 테스트 실행
147
+ pytest
148
+
149
+ # 커버리지 포함
150
+ pytest --cov=src/extract_hwp
151
+ ```
152
+
153
+ ### 코드 품질
154
+
155
+ ```bash
156
+ # 코드 포맷팅
157
+ black src/ tests/
158
+
159
+ # 타입 검사
160
+ mypy src/
161
+ ```
162
+
163
+ ## 라이선스
164
+
165
+ BSD 3-Clause License - 자세한 내용은 [LICENSE](LICENSE) 파일을 참조하세요.
166
+
167
+ ## 변경사항
168
+
169
+ 버전 히스토리는 [CHANGELOG.md](CHANGELOG.md)에서 확인할 수 있습니다.
@@ -0,0 +1,174 @@
1
+ # extract-hwp 사용 예제
2
+
3
+ 이 디렉토리에는 extract-hwp 라이브러리를 사용하는 다양한 예제들이 포함되어 있습니다.
4
+
5
+ ## 예제 파일들
6
+
7
+ ### 1. `basic_usage.py` - 기본 사용법
8
+ 라이브러리의 기본적인 사용 방법을 보여주는 예제입니다.
9
+
10
+ ```bash
11
+ python basic_usage.py
12
+ ```
13
+
14
+ **주요 기능:**
15
+ - 단일 파일 텍스트 추출
16
+ - 암호화 파일 감지
17
+ - 오류 처리 방법
18
+ - 포맷별 추출 함수 사용법
19
+
20
+ ### 2. `batch_processing.py` - 일괄 처리
21
+ 디렉토리 내의 모든 HWP 파일을 일괄 처리하는 예제입니다.
22
+
23
+ ```bash
24
+ python batch_processing.py <디렉토리> [출력형식]
25
+ ```
26
+
27
+ **사용 예제:**
28
+ ```bash
29
+ # 현재 디렉토리의 모든 HWP 파일을 텍스트 파일로 저장
30
+ python batch_processing.py ./documents txt
31
+
32
+ # JSON 형식으로 결과 저장
33
+ python batch_processing.py ./documents json
34
+ ```
35
+
36
+ **주요 기능:**
37
+ - 재귀적 파일 검색 (하위 디렉토리 포함)
38
+ - 일괄 처리 진행 상황 표시
39
+ - 처리 결과 통계
40
+ - 다양한 출력 형식 지원
41
+
42
+ ### 3. `cli_tool.py` - 명령줄 도구
43
+ HWP 파일 처리를 위한 완전한 명령줄 인터페이스입니다.
44
+
45
+ ```bash
46
+ python cli_tool.py <파일경로> [옵션]
47
+ ```
48
+
49
+ **사용 예제:**
50
+ ```bash
51
+ # 기본 사용 (표준 출력)
52
+ python cli_tool.py document.hwp
53
+
54
+ # 파일로 저장
55
+ python cli_tool.py document.hwp --output extracted.txt
56
+
57
+ # 암호화 여부만 확인
58
+ python cli_tool.py document.hwp --check-password
59
+
60
+ # 조용한 모드로 실행
61
+ python cli_tool.py document.hwpx --quiet --output-dir ./outputs
62
+ ```
63
+
64
+ **주요 옵션:**
65
+ - `-o, --output`: 출력 파일 지정
66
+ - `--output-dir`: 출력 디렉토리 지정
67
+ - `-c, --check-password`: 암호화 여부만 확인
68
+ - `-q, --quiet`: 진행 메시지 숨기기
69
+ - `--encoding`: 출력 파일 인코딩 지정
70
+
71
+ ## 실행 방법
72
+
73
+ ### 개발 환경에서 실행
74
+ ```bash
75
+ # extract-hwp 프로젝트 루트 디렉토리에서
76
+ cd examples
77
+ python basic_usage.py
78
+ python batch_processing.py ./test_files
79
+ python cli_tool.py test.hwp
80
+ ```
81
+
82
+ ### 설치된 패키지로 실행
83
+ 패키지가 설치된 상태에서는 import 경로 수정이 필요합니다:
84
+
85
+ ```python
86
+ # 예제 파일 상단의 다음 부분을 제거하거나 주석 처리
87
+ # sys.path.insert(0, str(Path(__file__).parent.parent / "src"))
88
+ ```
89
+
90
+ ## 테스트 파일 준비
91
+
92
+ 예제를 실제로 테스트해보려면 HWP/HWPX 파일이 필요합니다:
93
+
94
+ 1. **샘플 파일 생성**: 한글(HWP)에서 간단한 문서를 작성하고 `.hwp` 및 `.hwpx` 형식으로 저장
95
+ 2. **테스트 디렉토리**: `./test_files/` 디렉토리를 만들고 샘플 파일들을 넣기
96
+ 3. **암호화 파일**: 암호가 설정된 HWP 파일도 테스트용으로 준비
97
+
98
+ ## 고급 사용법
99
+
100
+ ### 커스텀 처리 로직
101
+ ```python
102
+ from extract_hwp import extract_text_from_hwp, is_hwp_file_password_protected
103
+
104
+ def custom_processor(file_path):
105
+ """커스텀 HWP 파일 처리기"""
106
+
107
+ # 1. 암호화 확인
108
+ if is_hwp_file_password_protected(file_path):
109
+ return None, "암호화된 파일"
110
+
111
+ # 2. 텍스트 추출
112
+ text, error = extract_text_from_hwp(file_path)
113
+ if error:
114
+ return None, error
115
+
116
+ # 3. 후처리 (예: 텍스트 정리)
117
+ processed_text = text.strip()
118
+ processed_text = '\n'.join(line.strip() for line in processed_text.split('\n') if line.strip())
119
+
120
+ return processed_text, None
121
+ ```
122
+
123
+ ### 에러 처리 패턴
124
+ ```python
125
+ import logging
126
+
127
+ def safe_extract(file_path):
128
+ """안전한 텍스트 추출 (로깅 포함)"""
129
+ try:
130
+ text, error = extract_text_from_hwp(file_path)
131
+ if error:
132
+ logging.warning(f"추출 실패: {file_path} - {error}")
133
+ return ""
134
+
135
+ logging.info(f"추출 성공: {file_path} ({len(text)} 문자)")
136
+ return text
137
+
138
+ except Exception as e:
139
+ logging.error(f"예외 발생: {file_path} - {e}")
140
+ return ""
141
+ ```
142
+
143
+ ## 성능 최적화 팁
144
+
145
+ 1. **대용량 파일**: 매우 큰 HWP 파일의 경우 메모리 사용량에 주의
146
+ 2. **일괄 처리**: 많은 파일을 처리할 때는 멀티프로세싱 사용 고려
147
+ 3. **암호화 확인**: 처리 전에 암호화 여부를 먼저 확인하여 불필요한 처리 방지
148
+
149
+ ## 문제 해결
150
+
151
+ ### 일반적인 문제들
152
+
153
+ 1. **ImportError**: 패키지가 설치되지 않은 경우
154
+ ```bash
155
+ pip install extract-hwp
156
+ ```
157
+
158
+ 2. **파일을 찾을 수 없음**: 파일 경로 확인
159
+ ```python
160
+ import os
161
+ print(os.path.exists("your_file.hwp"))
162
+ ```
163
+
164
+ 3. **인코딩 오류**: 파일명에 특수문자가 있는 경우
165
+ ```python
166
+ file_path = file_path.encode('utf-8').decode('utf-8')
167
+ ```
168
+
169
+ 4. **권한 오류**: 파일 읽기 권한 확인
170
+ ```bash
171
+ ls -la your_file.hwp
172
+ ```
173
+
174
+ 더 자세한 정보는 프로젝트의 메인 README.md를 참조하세요.