pyxtxt 0.3.4.2__tar.gz → 0.3.6__tar.gz
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- pyxtxt-0.3.6/PKG-INFO +448 -0
- pyxtxt-0.3.6/README.md +363 -0
- {pyxtxt-0.3.4.2 → pyxtxt-0.3.6}/pyproject.toml +14 -5
- pyxtxt-0.3.6/src/pyxtxt/__init__.py +59 -0
- pyxtxt-0.3.6/src/pyxtxt/core.py +131 -0
- pyxtxt-0.3.6/src/pyxtxt/estrattori/__init__.py +25 -0
- pyxtxt-0.3.6/src/pyxtxt/estrattori/exif.py +183 -0
- {pyxtxt-0.3.4.2 → pyxtxt-0.3.6}/src/pyxtxt/estrattori/md.py +2 -5
- {pyxtxt-0.3.4.2 → pyxtxt-0.3.6}/src/pyxtxt/estrattori/ocr_ollama.py +4 -0
- {pyxtxt-0.3.4.2 → pyxtxt-0.3.6}/src/pyxtxt/estrattori/pdf.py +5 -2
- {pyxtxt-0.3.4.2 → pyxtxt-0.3.6}/src/pyxtxt/estrattori/rtf.py +3 -5
- {pyxtxt-0.3.4.2 → pyxtxt-0.3.6}/src/pyxtxt/estrattori/tex.py +3 -5
- {pyxtxt-0.3.4.2 → pyxtxt-0.3.6}/src/pyxtxt/estrattori/xls.py +7 -2
- {pyxtxt-0.3.4.2 → pyxtxt-0.3.6}/src/pyxtxt/estrattori/xlsx.py +4 -2
- pyxtxt-0.3.6/src/pyxtxt.egg-info/PKG-INFO +448 -0
- {pyxtxt-0.3.4.2 → pyxtxt-0.3.6}/src/pyxtxt.egg-info/SOURCES.txt +4 -3
- pyxtxt-0.3.6/tests/test_extraction.py +122 -0
- pyxtxt-0.3.6/tests/test_import.py +49 -0
- pyxtxt-0.3.4.2/MANIFEST.in +0 -3
- pyxtxt-0.3.4.2/PKG-INFO +0 -585
- pyxtxt-0.3.4.2/README.md +0 -483
- pyxtxt-0.3.4.2/src/pyxtxt/__init__.py +0 -19
- pyxtxt-0.3.4.2/src/pyxtxt/core.py +0 -125
- pyxtxt-0.3.4.2/src/pyxtxt/estrattori/__init__.py +0 -18
- pyxtxt-0.3.4.2/src/pyxtxt/pyxtxt.py +0 -280
- pyxtxt-0.3.4.2/src/pyxtxt.egg-info/PKG-INFO +0 -585
- {pyxtxt-0.3.4.2 → pyxtxt-0.3.6}/LICENSE +0 -0
- {pyxtxt-0.3.4.2 → pyxtxt-0.3.6}/setup.cfg +0 -0
- {pyxtxt-0.3.4.2 → pyxtxt-0.3.6}/src/pyxtxt/estrattori/audio.py +0 -0
- {pyxtxt-0.3.4.2 → pyxtxt-0.3.6}/src/pyxtxt/estrattori/doc.py +0 -0
- {pyxtxt-0.3.4.2 → pyxtxt-0.3.6}/src/pyxtxt/estrattori/docx.py +0 -0
- {pyxtxt-0.3.4.2 → pyxtxt-0.3.6}/src/pyxtxt/estrattori/eml.py +0 -0
- {pyxtxt-0.3.4.2 → pyxtxt-0.3.6}/src/pyxtxt/estrattori/epub.py +0 -0
- {pyxtxt-0.3.4.2 → pyxtxt-0.3.6}/src/pyxtxt/estrattori/html.py +0 -0
- {pyxtxt-0.3.4.2 → pyxtxt-0.3.6}/src/pyxtxt/estrattori/msg.py +0 -0
- {pyxtxt-0.3.4.2 → pyxtxt-0.3.6}/src/pyxtxt/estrattori/ocr.py +0 -0
- {pyxtxt-0.3.4.2 → pyxtxt-0.3.6}/src/pyxtxt/estrattori/odt.py +0 -0
- {pyxtxt-0.3.4.2 → pyxtxt-0.3.6}/src/pyxtxt/estrattori/pptx.py +0 -0
- {pyxtxt-0.3.4.2 → pyxtxt-0.3.6}/src/pyxtxt/estrattori/svg.py +0 -0
- {pyxtxt-0.3.4.2 → pyxtxt-0.3.6}/src/pyxtxt/estrattori/txt.py +0 -0
- {pyxtxt-0.3.4.2 → pyxtxt-0.3.6}/src/pyxtxt/estrattori/xml.py +0 -0
- {pyxtxt-0.3.4.2 → pyxtxt-0.3.6}/src/pyxtxt/examples.py +0 -0
- {pyxtxt-0.3.4.2 → pyxtxt-0.3.6}/src/pyxtxt.egg-info/dependency_links.txt +0 -0
- {pyxtxt-0.3.4.2 → pyxtxt-0.3.6}/src/pyxtxt.egg-info/requires.txt +0 -0
- {pyxtxt-0.3.4.2 → pyxtxt-0.3.6}/src/pyxtxt.egg-info/top_level.txt +0 -0
pyxtxt-0.3.6/PKG-INFO
ADDED
|
@@ -0,0 +1,448 @@
|
|
|
1
|
+
Metadata-Version: 2.4
|
|
2
|
+
Name: pyxtxt
|
|
3
|
+
Version: 0.3.6
|
|
4
|
+
Summary: A Python library for extracting text from different types of files (PDF, DOCX, PPTX, XLSX, ODT, etc.).
|
|
5
|
+
Author-email: Giuseppe Levi <giuseppe.levi@gmail.com>
|
|
6
|
+
License-Expression: MIT
|
|
7
|
+
Project-URL: Homepage, https://github.com/GiuseppeLeviBo/pyxtxt
|
|
8
|
+
Project-URL: Repository, https://github.com/GiuseppeLeviBo/pyxtxt
|
|
9
|
+
Project-URL: Issues, https://github.com/GiuseppeLeviBo/pyxtxt/issues
|
|
10
|
+
Classifier: Development Status :: 4 - Beta
|
|
11
|
+
Classifier: Intended Audience :: Developers
|
|
12
|
+
Classifier: Operating System :: OS Independent
|
|
13
|
+
Classifier: Programming Language :: Python :: 3
|
|
14
|
+
Classifier: Programming Language :: Python :: 3.7
|
|
15
|
+
Classifier: Programming Language :: Python :: 3.8
|
|
16
|
+
Classifier: Programming Language :: Python :: 3.9
|
|
17
|
+
Classifier: Programming Language :: Python :: 3.10
|
|
18
|
+
Classifier: Programming Language :: Python :: 3.11
|
|
19
|
+
Classifier: Programming Language :: Python :: 3.12
|
|
20
|
+
Classifier: Programming Language :: Python :: 3.13
|
|
21
|
+
Classifier: Topic :: Text Processing
|
|
22
|
+
Classifier: Topic :: Utilities
|
|
23
|
+
Requires-Python: >=3.7
|
|
24
|
+
Description-Content-Type: text/markdown
|
|
25
|
+
License-File: LICENSE
|
|
26
|
+
Requires-Dist: python-magic; sys_platform != "win32"
|
|
27
|
+
Requires-Dist: python-magic-bin; sys_platform == "win32"
|
|
28
|
+
Provides-Extra: pdf
|
|
29
|
+
Requires-Dist: PyMuPDF; extra == "pdf"
|
|
30
|
+
Provides-Extra: docx
|
|
31
|
+
Requires-Dist: python-docx; extra == "docx"
|
|
32
|
+
Provides-Extra: presentation
|
|
33
|
+
Requires-Dist: python-pptx; extra == "presentation"
|
|
34
|
+
Provides-Extra: spreadsheet
|
|
35
|
+
Requires-Dist: openpyxl; extra == "spreadsheet"
|
|
36
|
+
Requires-Dist: xlrd; extra == "spreadsheet"
|
|
37
|
+
Provides-Extra: odf
|
|
38
|
+
Requires-Dist: odfpy; extra == "odf"
|
|
39
|
+
Provides-Extra: html
|
|
40
|
+
Requires-Dist: beautifulsoup4; extra == "html"
|
|
41
|
+
Requires-Dist: lxml; extra == "html"
|
|
42
|
+
Provides-Extra: doc
|
|
43
|
+
Provides-Extra: markdown
|
|
44
|
+
Requires-Dist: markdown; extra == "markdown"
|
|
45
|
+
Requires-Dist: beautifulsoup4; extra == "markdown"
|
|
46
|
+
Provides-Extra: epub
|
|
47
|
+
Requires-Dist: ebooklib; extra == "epub"
|
|
48
|
+
Requires-Dist: beautifulsoup4; extra == "epub"
|
|
49
|
+
Provides-Extra: rtf
|
|
50
|
+
Requires-Dist: striprtf; extra == "rtf"
|
|
51
|
+
Provides-Extra: email
|
|
52
|
+
Requires-Dist: beautifulsoup4; extra == "email"
|
|
53
|
+
Provides-Extra: outlook
|
|
54
|
+
Requires-Dist: extract-msg; extra == "outlook"
|
|
55
|
+
Requires-Dist: beautifulsoup4; extra == "outlook"
|
|
56
|
+
Provides-Extra: latex
|
|
57
|
+
Requires-Dist: pylatexenc; extra == "latex"
|
|
58
|
+
Provides-Extra: audio
|
|
59
|
+
Requires-Dist: openai-whisper; extra == "audio"
|
|
60
|
+
Provides-Extra: ocr
|
|
61
|
+
Requires-Dist: easyocr; extra == "ocr"
|
|
62
|
+
Requires-Dist: pillow; extra == "ocr"
|
|
63
|
+
Provides-Extra: ocr-ollama
|
|
64
|
+
Requires-Dist: ollama; extra == "ocr-ollama"
|
|
65
|
+
Requires-Dist: pillow; extra == "ocr-ollama"
|
|
66
|
+
Provides-Extra: all
|
|
67
|
+
Requires-Dist: PyMuPDF; extra == "all"
|
|
68
|
+
Requires-Dist: python-docx; extra == "all"
|
|
69
|
+
Requires-Dist: python-pptx; extra == "all"
|
|
70
|
+
Requires-Dist: openpyxl; extra == "all"
|
|
71
|
+
Requires-Dist: xlrd; extra == "all"
|
|
72
|
+
Requires-Dist: odfpy; extra == "all"
|
|
73
|
+
Requires-Dist: beautifulsoup4; extra == "all"
|
|
74
|
+
Requires-Dist: lxml; extra == "all"
|
|
75
|
+
Requires-Dist: markdown; extra == "all"
|
|
76
|
+
Requires-Dist: ebooklib; extra == "all"
|
|
77
|
+
Requires-Dist: striprtf; extra == "all"
|
|
78
|
+
Requires-Dist: extract-msg; extra == "all"
|
|
79
|
+
Requires-Dist: pylatexenc; extra == "all"
|
|
80
|
+
Requires-Dist: openai-whisper; extra == "all"
|
|
81
|
+
Requires-Dist: easyocr; extra == "all"
|
|
82
|
+
Requires-Dist: pillow; extra == "all"
|
|
83
|
+
Requires-Dist: ollama; extra == "all"
|
|
84
|
+
Dynamic: license-file
|
|
85
|
+
|
|
86
|
+
# PyxTxt
|
|
87
|
+
|
|
88
|
+
[](https://pypi.org/project/pyxtxt/)
|
|
89
|
+
[](https://pypi.org/project/pyxtxt/)
|
|
90
|
+
[](https://github.com/GiuseppeLeviBo/pyxtxt/actions/workflows/ci.yml)
|
|
91
|
+
[](https://opensource.org/licenses/MIT)
|
|
92
|
+
|
|
93
|
+
**PyxTxt** is a small Python library that extracts plain text from many file formats through a single function, `xtxt()`.
|
|
94
|
+
It detects the file type automatically (via `libmagic`) and dispatches to the right extractor. Extractors are optional:
|
|
95
|
+
install only the ones you need.
|
|
96
|
+
|
|
97
|
+
```python
|
|
98
|
+
from pyxtxt import xtxt
|
|
99
|
+
|
|
100
|
+
text = xtxt("report.pdf")
|
|
101
|
+
```
|
|
102
|
+
|
|
103
|
+
---
|
|
104
|
+
|
|
105
|
+
## ✨ Features
|
|
106
|
+
|
|
107
|
+
- **One function for everything**: `xtxt()` accepts a file path, an `io.BytesIO` buffer, raw `bytes` or a `requests.Response`
|
|
108
|
+
- **Automatic type detection** with `python-magic`, refined by the file extension when libmagic is not specific enough (e.g. Markdown)
|
|
109
|
+
- **Modular dependencies**: each format is an optional extra, the core only needs `python-magic`
|
|
110
|
+
- **Office, web and document formats**: PDF, DOCX, PPTX, XLSX, XLS, ODT, HTML, XML, SVG, Markdown, EPUB, RTF, EML, MSG, LaTeX, DOC, TXT
|
|
111
|
+
- **Audio and video transcription** with OpenAI Whisper
|
|
112
|
+
- **OCR from images** with EasyOCR or with a local multimodal LLM through Ollama
|
|
113
|
+
- **EXIF metadata** extraction from photos
|
|
114
|
+
|
|
115
|
+
---
|
|
116
|
+
|
|
117
|
+
## 📄 Supported formats
|
|
118
|
+
|
|
119
|
+
| Format | Install extra | Notes |
|
|
120
|
+
|---|---|---|
|
|
121
|
+
| PDF | `pdf` | PyMuPDF |
|
|
122
|
+
| DOCX | `docx` | Paragraph text (tables are not extracted yet) |
|
|
123
|
+
| PPTX | `presentation` | Text of all slide shapes |
|
|
124
|
+
| XLSX, XLS | `spreadsheet` | Every row of every visible sheet, cells joined with ` \| ` |
|
|
125
|
+
| ODT | `odf` | |
|
|
126
|
+
| HTML | `html` | |
|
|
127
|
+
| XML, SVG | `html` | Both use `lxml`, installed by the `html` extra |
|
|
128
|
+
| Markdown | `markdown` | Detected by the `.md` / `.markdown` extension |
|
|
129
|
+
| EPUB | `epub` | |
|
|
130
|
+
| RTF | `rtf` | |
|
|
131
|
+
| EML | `email` | Plain-text and HTML parts |
|
|
132
|
+
| MSG (Outlook) | `outlook` | |
|
|
133
|
+
| LaTeX | `latex` | |
|
|
134
|
+
| DOC (legacy Word) | — | Needs the `antiword` system tool |
|
|
135
|
+
| TXT and other `text/*` | — | Always available, decoded as UTF-8 |
|
|
136
|
+
| Audio and video | `audio` | Whisper, needs `ffmpeg`; heavy download |
|
|
137
|
+
| Images (OCR) | `ocr` or `ocr-ollama` | See [OCR from images](#-ocr-from-images) |
|
|
138
|
+
|
|
139
|
+
To list what is available in your installation:
|
|
140
|
+
|
|
141
|
+
```python
|
|
142
|
+
from pyxtxt import extxt_available_formats
|
|
143
|
+
|
|
144
|
+
print(extxt_available_formats()) # MIME types
|
|
145
|
+
print(extxt_available_formats(pretty=True)) # Short names
|
|
146
|
+
```
|
|
147
|
+
|
|
148
|
+
---
|
|
149
|
+
|
|
150
|
+
## 📦 Installation
|
|
151
|
+
|
|
152
|
+
Install every extractor (this includes the heavy audio and OCR dependencies):
|
|
153
|
+
|
|
154
|
+
```bash
|
|
155
|
+
pip install "pyxtxt[all]"
|
|
156
|
+
```
|
|
157
|
+
|
|
158
|
+
or only the formats you need:
|
|
159
|
+
|
|
160
|
+
```bash
|
|
161
|
+
pip install "pyxtxt[pdf,docx,presentation,spreadsheet,html,markdown,epub,email]"
|
|
162
|
+
```
|
|
163
|
+
|
|
164
|
+
Heavy optional extras:
|
|
165
|
+
|
|
166
|
+
```bash
|
|
167
|
+
pip install "pyxtxt[audio]" # Whisper transcription (~2 GB with models, pulls in PyTorch)
|
|
168
|
+
pip install "pyxtxt[ocr]" # EasyOCR (~1 GB with models, pulls in PyTorch)
|
|
169
|
+
pip install "pyxtxt[ocr-ollama]" # OCR through a local Ollama server
|
|
170
|
+
```
|
|
171
|
+
|
|
172
|
+
### System dependencies
|
|
173
|
+
|
|
174
|
+
**libmagic** (required by `python-magic`):
|
|
175
|
+
|
|
176
|
+
```bash
|
|
177
|
+
sudo apt install libmagic1 # Ubuntu / Debian
|
|
178
|
+
brew install libmagic # macOS
|
|
179
|
+
```
|
|
180
|
+
|
|
181
|
+
On Windows the `python-magic-bin` package, which bundles libmagic, is installed automatically.
|
|
182
|
+
|
|
183
|
+
**antiword** (only for legacy `.doc` files):
|
|
184
|
+
|
|
185
|
+
```bash
|
|
186
|
+
sudo apt install antiword # Ubuntu / Debian
|
|
187
|
+
brew install antiword # macOS
|
|
188
|
+
```
|
|
189
|
+
|
|
190
|
+
**ffmpeg** (only for audio/video transcription):
|
|
191
|
+
|
|
192
|
+
```bash
|
|
193
|
+
sudo apt install ffmpeg # Ubuntu / Debian
|
|
194
|
+
brew install ffmpeg # macOS
|
|
195
|
+
# Windows: https://ffmpeg.org/download.html
|
|
196
|
+
```
|
|
197
|
+
|
|
198
|
+
---
|
|
199
|
+
|
|
200
|
+
## 📚 Usage
|
|
201
|
+
|
|
202
|
+
### Basic usage
|
|
203
|
+
|
|
204
|
+
```python
|
|
205
|
+
import io
|
|
206
|
+
from pyxtxt import xtxt
|
|
207
|
+
|
|
208
|
+
# From a file path
|
|
209
|
+
text = xtxt("document.pdf")
|
|
210
|
+
|
|
211
|
+
# From an in-memory buffer
|
|
212
|
+
with open("document.docx", "rb") as f:
|
|
213
|
+
buffer = io.BytesIO(f.read())
|
|
214
|
+
text = xtxt(buffer)
|
|
215
|
+
|
|
216
|
+
# Give the buffer a name to help type detection (useful for Markdown, LaTeX, RTF)
|
|
217
|
+
buffer = io.BytesIO(markdown_bytes)
|
|
218
|
+
buffer.name = "notes.md"
|
|
219
|
+
text = xtxt(buffer)
|
|
220
|
+
```
|
|
221
|
+
|
|
222
|
+
`xtxt()` returns the extracted text as a `str`, or `None` when the file cannot be read or its type is not supported.
|
|
223
|
+
|
|
224
|
+
### Web content
|
|
225
|
+
|
|
226
|
+
`xtxt_from_url()` and `requests.Response` support need the `requests` package (`pip install requests`).
|
|
227
|
+
|
|
228
|
+
```python
|
|
229
|
+
import requests
|
|
230
|
+
from pyxtxt import xtxt, xtxt_from_url
|
|
231
|
+
|
|
232
|
+
response = requests.get("https://example.com/document.pdf")
|
|
233
|
+
text = xtxt(response.content) # from bytes
|
|
234
|
+
text = xtxt(response) # from the Response object
|
|
235
|
+
|
|
236
|
+
text = xtxt_from_url("https://example.com/document.pdf", timeout=10)
|
|
237
|
+
```
|
|
238
|
+
|
|
239
|
+
Extra keyword arguments of `xtxt_from_url()` are passed to `requests.get()`.
|
|
240
|
+
|
|
241
|
+
Typical uses:
|
|
242
|
+
|
|
243
|
+
```python
|
|
244
|
+
# File uploads (Flask / Django)
|
|
245
|
+
text = xtxt(request.files["document"].read())
|
|
246
|
+
|
|
247
|
+
# Email attachments
|
|
248
|
+
text = xtxt(attachment.get_payload(decode=True))
|
|
249
|
+
```
|
|
250
|
+
|
|
251
|
+
### Audio and video transcription
|
|
252
|
+
|
|
253
|
+
```python
|
|
254
|
+
from pyxtxt import xtxt
|
|
255
|
+
|
|
256
|
+
text = xtxt("meeting_recording.mp3")
|
|
257
|
+
text = xtxt("interview.wav")
|
|
258
|
+
text = xtxt("presentation.mp4") # the audio track is extracted automatically
|
|
259
|
+
```
|
|
260
|
+
|
|
261
|
+
The Whisper `base` model is downloaded on first use and cached for the following calls.
|
|
262
|
+
|
|
263
|
+
### 🖼 OCR from images
|
|
264
|
+
|
|
265
|
+
Two OCR back-ends are available. If both are installed, **Ollama takes precedence** for `xtxt()` on images.
|
|
266
|
+
|
|
267
|
+
**EasyOCR** (`pip install "pyxtxt[ocr]"`) runs locally on CPU, recognising Italian and English:
|
|
268
|
+
|
|
269
|
+
```python
|
|
270
|
+
text = xtxt("scanned_document.png")
|
|
271
|
+
```
|
|
272
|
+
|
|
273
|
+
**Ollama** (`pip install "pyxtxt[ocr-ollama]"`) uses a multimodal LLM served by a local
|
|
274
|
+
[Ollama](https://ollama.com) instance. Start the server and pull a model first (`ollama pull gemma3:4b`).
|
|
275
|
+
|
|
276
|
+
```python
|
|
277
|
+
from pyxtxt import (
|
|
278
|
+
xtxt, xtxt_image_describe,
|
|
279
|
+
set_ollama_model, set_ollama_config, get_ollama_config, reset_ollama_config,
|
|
280
|
+
)
|
|
281
|
+
|
|
282
|
+
set_ollama_model("gemma3:12b") # default: gemma3:4b; also llava:7b, llava:13b, gemma3:27b
|
|
283
|
+
|
|
284
|
+
set_ollama_config(
|
|
285
|
+
language="italian", # language hint
|
|
286
|
+
caption_length="long", # short, medium, long
|
|
287
|
+
style="detailed", # descriptive, technical, simple, detailed
|
|
288
|
+
context="document", # general, document, handwriting, technical, cookbook, ...
|
|
289
|
+
temperature=0.2,
|
|
290
|
+
max_tokens=2000,
|
|
291
|
+
auto_fallback=False, # by default other models are tried when the result looks poor
|
|
292
|
+
)
|
|
293
|
+
|
|
294
|
+
text = xtxt("complex_document.png") # text only
|
|
295
|
+
analysis = xtxt_image_describe("scientific_diagram.png")
|
|
296
|
+
# TEXT: ...
|
|
297
|
+
# DESCRIPTION: ...
|
|
298
|
+
|
|
299
|
+
print(get_ollama_config())
|
|
300
|
+
reset_ollama_config()
|
|
301
|
+
```
|
|
302
|
+
|
|
303
|
+
#### Confidence score
|
|
304
|
+
|
|
305
|
+
```python
|
|
306
|
+
from pyxtxt import xtxt_image_with_confidence, set_ollama_config
|
|
307
|
+
|
|
308
|
+
set_ollama_config(confidence_threshold=0.8)
|
|
309
|
+
text, confidence = xtxt_image_with_confidence("document.png", mode="ocr")
|
|
310
|
+
```
|
|
311
|
+
|
|
312
|
+
The score is a **heuristic** computed on the model's answer: it rewards structured text (numbers, punctuation) and
|
|
313
|
+
penalises vague language and typical hallucination keywords (e.g. "ancient", "papyrus", "painting", "dragon").
|
|
314
|
+
It is not a calibrated probability. In OCR mode, results below `confidence_threshold` (default 0.7) are discarded
|
|
315
|
+
and an empty string is returned.
|
|
316
|
+
|
|
317
|
+
### EXIF metadata
|
|
318
|
+
|
|
319
|
+
Requires Pillow (installed by the `ocr` or `ocr-ollama` extras, or `pip install pillow`).
|
|
320
|
+
|
|
321
|
+
```python
|
|
322
|
+
from pyxtxt import xtxt_exif
|
|
323
|
+
|
|
324
|
+
print(xtxt_exif("vacation_photo.jpg"))
|
|
325
|
+
# Camera make/model, shooting settings, date/time, GPS coordinates, image size...
|
|
326
|
+
```
|
|
327
|
+
|
|
328
|
+
### More examples
|
|
329
|
+
|
|
330
|
+
An examples script is installed with the package:
|
|
331
|
+
|
|
332
|
+
```bash
|
|
333
|
+
python -m pyxtxt.examples
|
|
334
|
+
```
|
|
335
|
+
|
|
336
|
+
or, to read its source:
|
|
337
|
+
|
|
338
|
+
```python
|
|
339
|
+
from importlib.resources import files
|
|
340
|
+
print((files("pyxtxt") / "examples.py").read_text())
|
|
341
|
+
```
|
|
342
|
+
|
|
343
|
+
---
|
|
344
|
+
|
|
345
|
+
## ⚠️ Known limitations
|
|
346
|
+
|
|
347
|
+
- **Supported inputs**: file paths, `io.BytesIO`, `bytes` and `requests.Response`. A file object returned by
|
|
348
|
+
`open()` must be read first (`xtxt(f.read())`).
|
|
349
|
+
- **Type detection without a file name**: libmagic cannot tell apart some formats from raw bytes (legacy Office files
|
|
350
|
+
share the same signature; Markdown looks like plain text). Pass a file path, or set `buffer.name`, when possible.
|
|
351
|
+
- **Legacy PowerPoint (`.ppt`)** is not supported.
|
|
352
|
+
- **DOCX**: text inside tables, headers and footers is not extracted yet. **SVG**: text inside `<tspan>` elements is not extracted yet.
|
|
353
|
+
- Errors are reported with messages printed to standard output and the functions return `None` or an empty string;
|
|
354
|
+
they do not raise exceptions.
|
|
355
|
+
|
|
356
|
+
### 🤖 AI-powered features
|
|
357
|
+
|
|
358
|
+
OCR through Ollama and Whisper transcription rely on machine-learning models and can produce **hallucinations**
|
|
359
|
+
(text or content that is not there), misinterpretations and language errors. Results vary between models and versions.
|
|
360
|
+
|
|
361
|
+
**Do not use them for critical applications** — medical diagnosis or medical image interpretation, legal or financial
|
|
362
|
+
documents where errors matter, safety systems — without human verification. Validate the results against the source,
|
|
363
|
+
use the confidence score as a hint only, and keep a traditional OCR (EasyOCR) as a cross-check when accuracy matters.
|
|
364
|
+
|
|
365
|
+
---
|
|
366
|
+
|
|
367
|
+
## 🛠 Development
|
|
368
|
+
|
|
369
|
+
```bash
|
|
370
|
+
git clone https://github.com/GiuseppeLeviBo/pyxtxt
|
|
371
|
+
cd pyxtxt
|
|
372
|
+
python -m venv .venv && source .venv/bin/activate
|
|
373
|
+
pip install -e ".[pdf,docx,presentation,spreadsheet,odf,html,markdown,epub,rtf,email,latex]" pytest
|
|
374
|
+
pytest
|
|
375
|
+
```
|
|
376
|
+
|
|
377
|
+
Tests for formats whose libraries are not installed are skipped. CI runs the test suite on Python 3.10–3.13.
|
|
378
|
+
|
|
379
|
+
### Releasing
|
|
380
|
+
|
|
381
|
+
Releases are published to PyPI by GitHub Actions (`.github/workflows/publish.yml`) through
|
|
382
|
+
[Trusted Publishing](https://docs.pypi.org/trusted-publishers/), so no API token is needed:
|
|
383
|
+
|
|
384
|
+
1. Update `version` in `pyproject.toml` and the changelog below; merge to `main`.
|
|
385
|
+
2. Tag and push: `git tag v0.3.6 && git push origin v0.3.6`.
|
|
386
|
+
|
|
387
|
+
The workflow checks that the tag matches the version, builds sdist and wheel, and uploads them.
|
|
388
|
+
|
|
389
|
+
One-time setup: on PyPI, open the project's *Settings → Publishing* and add a
|
|
390
|
+
GitHub publisher with owner `GiuseppeLeviBo`, repository `pyxtxt`, workflow `publish.yml` and environment `pypi`.
|
|
391
|
+
|
|
392
|
+
---
|
|
393
|
+
|
|
394
|
+
## 🔒 License
|
|
395
|
+
|
|
396
|
+
Distributed under the MIT License. See [LICENSE](LICENSE).
|
|
397
|
+
|
|
398
|
+
## 🤝 Contributing
|
|
399
|
+
|
|
400
|
+
Pull requests, issues and feedback are welcome.
|
|
401
|
+
|
|
402
|
+
- **Bug reports**: include a sample file (or how to create one) and the full error message
|
|
403
|
+
- **Feature requests**: describe your use case and the expected behaviour
|
|
404
|
+
- **Code**: follow the existing patterns and add a test in `tests/`
|
|
405
|
+
|
|
406
|
+
---
|
|
407
|
+
|
|
408
|
+
## 📊 Changelog
|
|
409
|
+
|
|
410
|
+
### v0.3.6
|
|
411
|
+
- **FIXED**: `import pyxtxt` crashed with `AttributeError` unless both `ollama` and Pillow were installed (regression in 0.3.4.2 and 0.3.5)
|
|
412
|
+
- **FIXED**: Markdown, RTF and LaTeX files were returned as raw source instead of being converted to text
|
|
413
|
+
- **FIXED**: a failed extraction (e.g. a corrupted PDF) returned the string `"None"` instead of `None`
|
|
414
|
+
- **FIXED**: XLSX and XLS files were silently truncated to 200 and 100 rows per sheet; all rows are now extracted
|
|
415
|
+
(`max_rows_per_sheet` is still available when calling the extractors directly)
|
|
416
|
+
- **FIXED**: when both EasyOCR and Ollama were installed, the OCR back-end used for images was random; Ollama now always takes precedence
|
|
417
|
+
- **FIXED**: a single extractor failing to load no longer prevents the whole package from importing
|
|
418
|
+
- **FIXED**: `BytesIO` buffers keep their `name`, which is now used to refine type detection
|
|
419
|
+
- PyMuPDF is imported as `pymupdf`, removing the deprecation warning printed at every import
|
|
420
|
+
- Added the missing `LICENSE` file, SPDX license metadata and project URLs
|
|
421
|
+
- Removed the obsolete `pyxtxt/pyxtxt.py` module
|
|
422
|
+
- Added a test suite and GitHub Actions workflows for CI and PyPI publishing
|
|
423
|
+
|
|
424
|
+
### v0.3.0 – v0.3.5
|
|
425
|
+
- **NEW**: OCR through Ollama multimodal models (`set_ollama_model`, `xtxt_image_describe`) — 0.3.0
|
|
426
|
+
- **NEW**: `set_ollama_config()`, `get_ollama_config()`, `reset_ollama_config()` — 0.3.2
|
|
427
|
+
- **NEW**: confidence score and hallucination detection (`xtxt_image_with_confidence`) — 0.3.4
|
|
428
|
+
- **NEW**: image enhancement before OCR, automatic model fallback, context presets — 0.3.4.2
|
|
429
|
+
- **NEW**: EXIF metadata extraction (`xtxt_exif`) — 0.3.5
|
|
430
|
+
|
|
431
|
+
### v0.2.4
|
|
432
|
+
- **NEW**: video transcription support (MP4, MOV, AVI, WebM, MKV) via Whisper
|
|
433
|
+
|
|
434
|
+
### v0.2.3
|
|
435
|
+
- **NEW**: audio transcription (MP3, WAV, M4A, FLAC, ...) with Whisper
|
|
436
|
+
- **NEW**: OCR from images (JPEG, PNG, TIFF, BMP, WebP) with EasyOCR
|
|
437
|
+
- **NEW**: `audio`, `ocr` and `all` installation extras
|
|
438
|
+
|
|
439
|
+
### v0.2.0 – v0.2.2
|
|
440
|
+
- **NEW**: automatic extractor registration
|
|
441
|
+
- **NEW**: Markdown, EPUB, RTF, EML, MSG and LaTeX extractors
|
|
442
|
+
|
|
443
|
+
### v0.1.24
|
|
444
|
+
- **NEW**: support for `bytes` and `requests.Response` inputs, `xtxt_from_url()` helper
|
|
445
|
+
|
|
446
|
+
### v0.1.0 – v0.1.23
|
|
447
|
+
- Initial releases: modular extractors for PDF, DOCX, PPTX, XLSX, ODT, HTML, XML, TXT and legacy Office files,
|
|
448
|
+
MIME detection with python-magic, `BytesIO` support
|