pyxtxt 0.3.4.2__tar.gz → 0.3.6__tar.gz

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (45) hide show
  1. pyxtxt-0.3.6/PKG-INFO +448 -0
  2. pyxtxt-0.3.6/README.md +363 -0
  3. {pyxtxt-0.3.4.2 → pyxtxt-0.3.6}/pyproject.toml +14 -5
  4. pyxtxt-0.3.6/src/pyxtxt/__init__.py +59 -0
  5. pyxtxt-0.3.6/src/pyxtxt/core.py +131 -0
  6. pyxtxt-0.3.6/src/pyxtxt/estrattori/__init__.py +25 -0
  7. pyxtxt-0.3.6/src/pyxtxt/estrattori/exif.py +183 -0
  8. {pyxtxt-0.3.4.2 → pyxtxt-0.3.6}/src/pyxtxt/estrattori/md.py +2 -5
  9. {pyxtxt-0.3.4.2 → pyxtxt-0.3.6}/src/pyxtxt/estrattori/ocr_ollama.py +4 -0
  10. {pyxtxt-0.3.4.2 → pyxtxt-0.3.6}/src/pyxtxt/estrattori/pdf.py +5 -2
  11. {pyxtxt-0.3.4.2 → pyxtxt-0.3.6}/src/pyxtxt/estrattori/rtf.py +3 -5
  12. {pyxtxt-0.3.4.2 → pyxtxt-0.3.6}/src/pyxtxt/estrattori/tex.py +3 -5
  13. {pyxtxt-0.3.4.2 → pyxtxt-0.3.6}/src/pyxtxt/estrattori/xls.py +7 -2
  14. {pyxtxt-0.3.4.2 → pyxtxt-0.3.6}/src/pyxtxt/estrattori/xlsx.py +4 -2
  15. pyxtxt-0.3.6/src/pyxtxt.egg-info/PKG-INFO +448 -0
  16. {pyxtxt-0.3.4.2 → pyxtxt-0.3.6}/src/pyxtxt.egg-info/SOURCES.txt +4 -3
  17. pyxtxt-0.3.6/tests/test_extraction.py +122 -0
  18. pyxtxt-0.3.6/tests/test_import.py +49 -0
  19. pyxtxt-0.3.4.2/MANIFEST.in +0 -3
  20. pyxtxt-0.3.4.2/PKG-INFO +0 -585
  21. pyxtxt-0.3.4.2/README.md +0 -483
  22. pyxtxt-0.3.4.2/src/pyxtxt/__init__.py +0 -19
  23. pyxtxt-0.3.4.2/src/pyxtxt/core.py +0 -125
  24. pyxtxt-0.3.4.2/src/pyxtxt/estrattori/__init__.py +0 -18
  25. pyxtxt-0.3.4.2/src/pyxtxt/pyxtxt.py +0 -280
  26. pyxtxt-0.3.4.2/src/pyxtxt.egg-info/PKG-INFO +0 -585
  27. {pyxtxt-0.3.4.2 → pyxtxt-0.3.6}/LICENSE +0 -0
  28. {pyxtxt-0.3.4.2 → pyxtxt-0.3.6}/setup.cfg +0 -0
  29. {pyxtxt-0.3.4.2 → pyxtxt-0.3.6}/src/pyxtxt/estrattori/audio.py +0 -0
  30. {pyxtxt-0.3.4.2 → pyxtxt-0.3.6}/src/pyxtxt/estrattori/doc.py +0 -0
  31. {pyxtxt-0.3.4.2 → pyxtxt-0.3.6}/src/pyxtxt/estrattori/docx.py +0 -0
  32. {pyxtxt-0.3.4.2 → pyxtxt-0.3.6}/src/pyxtxt/estrattori/eml.py +0 -0
  33. {pyxtxt-0.3.4.2 → pyxtxt-0.3.6}/src/pyxtxt/estrattori/epub.py +0 -0
  34. {pyxtxt-0.3.4.2 → pyxtxt-0.3.6}/src/pyxtxt/estrattori/html.py +0 -0
  35. {pyxtxt-0.3.4.2 → pyxtxt-0.3.6}/src/pyxtxt/estrattori/msg.py +0 -0
  36. {pyxtxt-0.3.4.2 → pyxtxt-0.3.6}/src/pyxtxt/estrattori/ocr.py +0 -0
  37. {pyxtxt-0.3.4.2 → pyxtxt-0.3.6}/src/pyxtxt/estrattori/odt.py +0 -0
  38. {pyxtxt-0.3.4.2 → pyxtxt-0.3.6}/src/pyxtxt/estrattori/pptx.py +0 -0
  39. {pyxtxt-0.3.4.2 → pyxtxt-0.3.6}/src/pyxtxt/estrattori/svg.py +0 -0
  40. {pyxtxt-0.3.4.2 → pyxtxt-0.3.6}/src/pyxtxt/estrattori/txt.py +0 -0
  41. {pyxtxt-0.3.4.2 → pyxtxt-0.3.6}/src/pyxtxt/estrattori/xml.py +0 -0
  42. {pyxtxt-0.3.4.2 → pyxtxt-0.3.6}/src/pyxtxt/examples.py +0 -0
  43. {pyxtxt-0.3.4.2 → pyxtxt-0.3.6}/src/pyxtxt.egg-info/dependency_links.txt +0 -0
  44. {pyxtxt-0.3.4.2 → pyxtxt-0.3.6}/src/pyxtxt.egg-info/requires.txt +0 -0
  45. {pyxtxt-0.3.4.2 → pyxtxt-0.3.6}/src/pyxtxt.egg-info/top_level.txt +0 -0
pyxtxt-0.3.6/PKG-INFO ADDED
@@ -0,0 +1,448 @@
1
+ Metadata-Version: 2.4
2
+ Name: pyxtxt
3
+ Version: 0.3.6
4
+ Summary: A Python library for extracting text from different types of files (PDF, DOCX, PPTX, XLSX, ODT, etc.).
5
+ Author-email: Giuseppe Levi <giuseppe.levi@gmail.com>
6
+ License-Expression: MIT
7
+ Project-URL: Homepage, https://github.com/GiuseppeLeviBo/pyxtxt
8
+ Project-URL: Repository, https://github.com/GiuseppeLeviBo/pyxtxt
9
+ Project-URL: Issues, https://github.com/GiuseppeLeviBo/pyxtxt/issues
10
+ Classifier: Development Status :: 4 - Beta
11
+ Classifier: Intended Audience :: Developers
12
+ Classifier: Operating System :: OS Independent
13
+ Classifier: Programming Language :: Python :: 3
14
+ Classifier: Programming Language :: Python :: 3.7
15
+ Classifier: Programming Language :: Python :: 3.8
16
+ Classifier: Programming Language :: Python :: 3.9
17
+ Classifier: Programming Language :: Python :: 3.10
18
+ Classifier: Programming Language :: Python :: 3.11
19
+ Classifier: Programming Language :: Python :: 3.12
20
+ Classifier: Programming Language :: Python :: 3.13
21
+ Classifier: Topic :: Text Processing
22
+ Classifier: Topic :: Utilities
23
+ Requires-Python: >=3.7
24
+ Description-Content-Type: text/markdown
25
+ License-File: LICENSE
26
+ Requires-Dist: python-magic; sys_platform != "win32"
27
+ Requires-Dist: python-magic-bin; sys_platform == "win32"
28
+ Provides-Extra: pdf
29
+ Requires-Dist: PyMuPDF; extra == "pdf"
30
+ Provides-Extra: docx
31
+ Requires-Dist: python-docx; extra == "docx"
32
+ Provides-Extra: presentation
33
+ Requires-Dist: python-pptx; extra == "presentation"
34
+ Provides-Extra: spreadsheet
35
+ Requires-Dist: openpyxl; extra == "spreadsheet"
36
+ Requires-Dist: xlrd; extra == "spreadsheet"
37
+ Provides-Extra: odf
38
+ Requires-Dist: odfpy; extra == "odf"
39
+ Provides-Extra: html
40
+ Requires-Dist: beautifulsoup4; extra == "html"
41
+ Requires-Dist: lxml; extra == "html"
42
+ Provides-Extra: doc
43
+ Provides-Extra: markdown
44
+ Requires-Dist: markdown; extra == "markdown"
45
+ Requires-Dist: beautifulsoup4; extra == "markdown"
46
+ Provides-Extra: epub
47
+ Requires-Dist: ebooklib; extra == "epub"
48
+ Requires-Dist: beautifulsoup4; extra == "epub"
49
+ Provides-Extra: rtf
50
+ Requires-Dist: striprtf; extra == "rtf"
51
+ Provides-Extra: email
52
+ Requires-Dist: beautifulsoup4; extra == "email"
53
+ Provides-Extra: outlook
54
+ Requires-Dist: extract-msg; extra == "outlook"
55
+ Requires-Dist: beautifulsoup4; extra == "outlook"
56
+ Provides-Extra: latex
57
+ Requires-Dist: pylatexenc; extra == "latex"
58
+ Provides-Extra: audio
59
+ Requires-Dist: openai-whisper; extra == "audio"
60
+ Provides-Extra: ocr
61
+ Requires-Dist: easyocr; extra == "ocr"
62
+ Requires-Dist: pillow; extra == "ocr"
63
+ Provides-Extra: ocr-ollama
64
+ Requires-Dist: ollama; extra == "ocr-ollama"
65
+ Requires-Dist: pillow; extra == "ocr-ollama"
66
+ Provides-Extra: all
67
+ Requires-Dist: PyMuPDF; extra == "all"
68
+ Requires-Dist: python-docx; extra == "all"
69
+ Requires-Dist: python-pptx; extra == "all"
70
+ Requires-Dist: openpyxl; extra == "all"
71
+ Requires-Dist: xlrd; extra == "all"
72
+ Requires-Dist: odfpy; extra == "all"
73
+ Requires-Dist: beautifulsoup4; extra == "all"
74
+ Requires-Dist: lxml; extra == "all"
75
+ Requires-Dist: markdown; extra == "all"
76
+ Requires-Dist: ebooklib; extra == "all"
77
+ Requires-Dist: striprtf; extra == "all"
78
+ Requires-Dist: extract-msg; extra == "all"
79
+ Requires-Dist: pylatexenc; extra == "all"
80
+ Requires-Dist: openai-whisper; extra == "all"
81
+ Requires-Dist: easyocr; extra == "all"
82
+ Requires-Dist: pillow; extra == "all"
83
+ Requires-Dist: ollama; extra == "all"
84
+ Dynamic: license-file
85
+
86
+ # PyxTxt
87
+
88
+ [![PyPI version](https://img.shields.io/pypi/v/pyxtxt.svg)](https://pypi.org/project/pyxtxt/)
89
+ [![Python versions](https://img.shields.io/pypi/pyversions/pyxtxt.svg)](https://pypi.org/project/pyxtxt/)
90
+ [![CI](https://github.com/GiuseppeLeviBo/pyxtxt/actions/workflows/ci.yml/badge.svg)](https://github.com/GiuseppeLeviBo/pyxtxt/actions/workflows/ci.yml)
91
+ [![License: MIT](https://img.shields.io/badge/License-MIT-yellow.svg)](https://opensource.org/licenses/MIT)
92
+
93
+ **PyxTxt** is a small Python library that extracts plain text from many file formats through a single function, `xtxt()`.
94
+ It detects the file type automatically (via `libmagic`) and dispatches to the right extractor. Extractors are optional:
95
+ install only the ones you need.
96
+
97
+ ```python
98
+ from pyxtxt import xtxt
99
+
100
+ text = xtxt("report.pdf")
101
+ ```
102
+
103
+ ---
104
+
105
+ ## ✨ Features
106
+
107
+ - **One function for everything**: `xtxt()` accepts a file path, an `io.BytesIO` buffer, raw `bytes` or a `requests.Response`
108
+ - **Automatic type detection** with `python-magic`, refined by the file extension when libmagic is not specific enough (e.g. Markdown)
109
+ - **Modular dependencies**: each format is an optional extra, the core only needs `python-magic`
110
+ - **Office, web and document formats**: PDF, DOCX, PPTX, XLSX, XLS, ODT, HTML, XML, SVG, Markdown, EPUB, RTF, EML, MSG, LaTeX, DOC, TXT
111
+ - **Audio and video transcription** with OpenAI Whisper
112
+ - **OCR from images** with EasyOCR or with a local multimodal LLM through Ollama
113
+ - **EXIF metadata** extraction from photos
114
+
115
+ ---
116
+
117
+ ## 📄 Supported formats
118
+
119
+ | Format | Install extra | Notes |
120
+ |---|---|---|
121
+ | PDF | `pdf` | PyMuPDF |
122
+ | DOCX | `docx` | Paragraph text (tables are not extracted yet) |
123
+ | PPTX | `presentation` | Text of all slide shapes |
124
+ | XLSX, XLS | `spreadsheet` | Every row of every visible sheet, cells joined with ` \| ` |
125
+ | ODT | `odf` | |
126
+ | HTML | `html` | |
127
+ | XML, SVG | `html` | Both use `lxml`, installed by the `html` extra |
128
+ | Markdown | `markdown` | Detected by the `.md` / `.markdown` extension |
129
+ | EPUB | `epub` | |
130
+ | RTF | `rtf` | |
131
+ | EML | `email` | Plain-text and HTML parts |
132
+ | MSG (Outlook) | `outlook` | |
133
+ | LaTeX | `latex` | |
134
+ | DOC (legacy Word) | — | Needs the `antiword` system tool |
135
+ | TXT and other `text/*` | — | Always available, decoded as UTF-8 |
136
+ | Audio and video | `audio` | Whisper, needs `ffmpeg`; heavy download |
137
+ | Images (OCR) | `ocr` or `ocr-ollama` | See [OCR from images](#-ocr-from-images) |
138
+
139
+ To list what is available in your installation:
140
+
141
+ ```python
142
+ from pyxtxt import extxt_available_formats
143
+
144
+ print(extxt_available_formats()) # MIME types
145
+ print(extxt_available_formats(pretty=True)) # Short names
146
+ ```
147
+
148
+ ---
149
+
150
+ ## 📦 Installation
151
+
152
+ Install every extractor (this includes the heavy audio and OCR dependencies):
153
+
154
+ ```bash
155
+ pip install "pyxtxt[all]"
156
+ ```
157
+
158
+ or only the formats you need:
159
+
160
+ ```bash
161
+ pip install "pyxtxt[pdf,docx,presentation,spreadsheet,html,markdown,epub,email]"
162
+ ```
163
+
164
+ Heavy optional extras:
165
+
166
+ ```bash
167
+ pip install "pyxtxt[audio]" # Whisper transcription (~2 GB with models, pulls in PyTorch)
168
+ pip install "pyxtxt[ocr]" # EasyOCR (~1 GB with models, pulls in PyTorch)
169
+ pip install "pyxtxt[ocr-ollama]" # OCR through a local Ollama server
170
+ ```
171
+
172
+ ### System dependencies
173
+
174
+ **libmagic** (required by `python-magic`):
175
+
176
+ ```bash
177
+ sudo apt install libmagic1 # Ubuntu / Debian
178
+ brew install libmagic # macOS
179
+ ```
180
+
181
+ On Windows the `python-magic-bin` package, which bundles libmagic, is installed automatically.
182
+
183
+ **antiword** (only for legacy `.doc` files):
184
+
185
+ ```bash
186
+ sudo apt install antiword # Ubuntu / Debian
187
+ brew install antiword # macOS
188
+ ```
189
+
190
+ **ffmpeg** (only for audio/video transcription):
191
+
192
+ ```bash
193
+ sudo apt install ffmpeg # Ubuntu / Debian
194
+ brew install ffmpeg # macOS
195
+ # Windows: https://ffmpeg.org/download.html
196
+ ```
197
+
198
+ ---
199
+
200
+ ## 📚 Usage
201
+
202
+ ### Basic usage
203
+
204
+ ```python
205
+ import io
206
+ from pyxtxt import xtxt
207
+
208
+ # From a file path
209
+ text = xtxt("document.pdf")
210
+
211
+ # From an in-memory buffer
212
+ with open("document.docx", "rb") as f:
213
+ buffer = io.BytesIO(f.read())
214
+ text = xtxt(buffer)
215
+
216
+ # Give the buffer a name to help type detection (useful for Markdown, LaTeX, RTF)
217
+ buffer = io.BytesIO(markdown_bytes)
218
+ buffer.name = "notes.md"
219
+ text = xtxt(buffer)
220
+ ```
221
+
222
+ `xtxt()` returns the extracted text as a `str`, or `None` when the file cannot be read or its type is not supported.
223
+
224
+ ### Web content
225
+
226
+ `xtxt_from_url()` and `requests.Response` support need the `requests` package (`pip install requests`).
227
+
228
+ ```python
229
+ import requests
230
+ from pyxtxt import xtxt, xtxt_from_url
231
+
232
+ response = requests.get("https://example.com/document.pdf")
233
+ text = xtxt(response.content) # from bytes
234
+ text = xtxt(response) # from the Response object
235
+
236
+ text = xtxt_from_url("https://example.com/document.pdf", timeout=10)
237
+ ```
238
+
239
+ Extra keyword arguments of `xtxt_from_url()` are passed to `requests.get()`.
240
+
241
+ Typical uses:
242
+
243
+ ```python
244
+ # File uploads (Flask / Django)
245
+ text = xtxt(request.files["document"].read())
246
+
247
+ # Email attachments
248
+ text = xtxt(attachment.get_payload(decode=True))
249
+ ```
250
+
251
+ ### Audio and video transcription
252
+
253
+ ```python
254
+ from pyxtxt import xtxt
255
+
256
+ text = xtxt("meeting_recording.mp3")
257
+ text = xtxt("interview.wav")
258
+ text = xtxt("presentation.mp4") # the audio track is extracted automatically
259
+ ```
260
+
261
+ The Whisper `base` model is downloaded on first use and cached for the following calls.
262
+
263
+ ### 🖼 OCR from images
264
+
265
+ Two OCR back-ends are available. If both are installed, **Ollama takes precedence** for `xtxt()` on images.
266
+
267
+ **EasyOCR** (`pip install "pyxtxt[ocr]"`) runs locally on CPU, recognising Italian and English:
268
+
269
+ ```python
270
+ text = xtxt("scanned_document.png")
271
+ ```
272
+
273
+ **Ollama** (`pip install "pyxtxt[ocr-ollama]"`) uses a multimodal LLM served by a local
274
+ [Ollama](https://ollama.com) instance. Start the server and pull a model first (`ollama pull gemma3:4b`).
275
+
276
+ ```python
277
+ from pyxtxt import (
278
+ xtxt, xtxt_image_describe,
279
+ set_ollama_model, set_ollama_config, get_ollama_config, reset_ollama_config,
280
+ )
281
+
282
+ set_ollama_model("gemma3:12b") # default: gemma3:4b; also llava:7b, llava:13b, gemma3:27b
283
+
284
+ set_ollama_config(
285
+ language="italian", # language hint
286
+ caption_length="long", # short, medium, long
287
+ style="detailed", # descriptive, technical, simple, detailed
288
+ context="document", # general, document, handwriting, technical, cookbook, ...
289
+ temperature=0.2,
290
+ max_tokens=2000,
291
+ auto_fallback=False, # by default other models are tried when the result looks poor
292
+ )
293
+
294
+ text = xtxt("complex_document.png") # text only
295
+ analysis = xtxt_image_describe("scientific_diagram.png")
296
+ # TEXT: ...
297
+ # DESCRIPTION: ...
298
+
299
+ print(get_ollama_config())
300
+ reset_ollama_config()
301
+ ```
302
+
303
+ #### Confidence score
304
+
305
+ ```python
306
+ from pyxtxt import xtxt_image_with_confidence, set_ollama_config
307
+
308
+ set_ollama_config(confidence_threshold=0.8)
309
+ text, confidence = xtxt_image_with_confidence("document.png", mode="ocr")
310
+ ```
311
+
312
+ The score is a **heuristic** computed on the model's answer: it rewards structured text (numbers, punctuation) and
313
+ penalises vague language and typical hallucination keywords (e.g. "ancient", "papyrus", "painting", "dragon").
314
+ It is not a calibrated probability. In OCR mode, results below `confidence_threshold` (default 0.7) are discarded
315
+ and an empty string is returned.
316
+
317
+ ### EXIF metadata
318
+
319
+ Requires Pillow (installed by the `ocr` or `ocr-ollama` extras, or `pip install pillow`).
320
+
321
+ ```python
322
+ from pyxtxt import xtxt_exif
323
+
324
+ print(xtxt_exif("vacation_photo.jpg"))
325
+ # Camera make/model, shooting settings, date/time, GPS coordinates, image size...
326
+ ```
327
+
328
+ ### More examples
329
+
330
+ An examples script is installed with the package:
331
+
332
+ ```bash
333
+ python -m pyxtxt.examples
334
+ ```
335
+
336
+ or, to read its source:
337
+
338
+ ```python
339
+ from importlib.resources import files
340
+ print((files("pyxtxt") / "examples.py").read_text())
341
+ ```
342
+
343
+ ---
344
+
345
+ ## ⚠️ Known limitations
346
+
347
+ - **Supported inputs**: file paths, `io.BytesIO`, `bytes` and `requests.Response`. A file object returned by
348
+ `open()` must be read first (`xtxt(f.read())`).
349
+ - **Type detection without a file name**: libmagic cannot tell apart some formats from raw bytes (legacy Office files
350
+ share the same signature; Markdown looks like plain text). Pass a file path, or set `buffer.name`, when possible.
351
+ - **Legacy PowerPoint (`.ppt`)** is not supported.
352
+ - **DOCX**: text inside tables, headers and footers is not extracted yet. **SVG**: text inside `<tspan>` elements is not extracted yet.
353
+ - Errors are reported with messages printed to standard output and the functions return `None` or an empty string;
354
+ they do not raise exceptions.
355
+
356
+ ### 🤖 AI-powered features
357
+
358
+ OCR through Ollama and Whisper transcription rely on machine-learning models and can produce **hallucinations**
359
+ (text or content that is not there), misinterpretations and language errors. Results vary between models and versions.
360
+
361
+ **Do not use them for critical applications** — medical diagnosis or medical image interpretation, legal or financial
362
+ documents where errors matter, safety systems — without human verification. Validate the results against the source,
363
+ use the confidence score as a hint only, and keep a traditional OCR (EasyOCR) as a cross-check when accuracy matters.
364
+
365
+ ---
366
+
367
+ ## 🛠 Development
368
+
369
+ ```bash
370
+ git clone https://github.com/GiuseppeLeviBo/pyxtxt
371
+ cd pyxtxt
372
+ python -m venv .venv && source .venv/bin/activate
373
+ pip install -e ".[pdf,docx,presentation,spreadsheet,odf,html,markdown,epub,rtf,email,latex]" pytest
374
+ pytest
375
+ ```
376
+
377
+ Tests for formats whose libraries are not installed are skipped. CI runs the test suite on Python 3.10–3.13.
378
+
379
+ ### Releasing
380
+
381
+ Releases are published to PyPI by GitHub Actions (`.github/workflows/publish.yml`) through
382
+ [Trusted Publishing](https://docs.pypi.org/trusted-publishers/), so no API token is needed:
383
+
384
+ 1. Update `version` in `pyproject.toml` and the changelog below; merge to `main`.
385
+ 2. Tag and push: `git tag v0.3.6 && git push origin v0.3.6`.
386
+
387
+ The workflow checks that the tag matches the version, builds sdist and wheel, and uploads them.
388
+
389
+ One-time setup: on PyPI, open the project's *Settings → Publishing* and add a
390
+ GitHub publisher with owner `GiuseppeLeviBo`, repository `pyxtxt`, workflow `publish.yml` and environment `pypi`.
391
+
392
+ ---
393
+
394
+ ## 🔒 License
395
+
396
+ Distributed under the MIT License. See [LICENSE](LICENSE).
397
+
398
+ ## 🤝 Contributing
399
+
400
+ Pull requests, issues and feedback are welcome.
401
+
402
+ - **Bug reports**: include a sample file (or how to create one) and the full error message
403
+ - **Feature requests**: describe your use case and the expected behaviour
404
+ - **Code**: follow the existing patterns and add a test in `tests/`
405
+
406
+ ---
407
+
408
+ ## 📊 Changelog
409
+
410
+ ### v0.3.6
411
+ - **FIXED**: `import pyxtxt` crashed with `AttributeError` unless both `ollama` and Pillow were installed (regression in 0.3.4.2 and 0.3.5)
412
+ - **FIXED**: Markdown, RTF and LaTeX files were returned as raw source instead of being converted to text
413
+ - **FIXED**: a failed extraction (e.g. a corrupted PDF) returned the string `"None"` instead of `None`
414
+ - **FIXED**: XLSX and XLS files were silently truncated to 200 and 100 rows per sheet; all rows are now extracted
415
+ (`max_rows_per_sheet` is still available when calling the extractors directly)
416
+ - **FIXED**: when both EasyOCR and Ollama were installed, the OCR back-end used for images was random; Ollama now always takes precedence
417
+ - **FIXED**: a single extractor failing to load no longer prevents the whole package from importing
418
+ - **FIXED**: `BytesIO` buffers keep their `name`, which is now used to refine type detection
419
+ - PyMuPDF is imported as `pymupdf`, removing the deprecation warning printed at every import
420
+ - Added the missing `LICENSE` file, SPDX license metadata and project URLs
421
+ - Removed the obsolete `pyxtxt/pyxtxt.py` module
422
+ - Added a test suite and GitHub Actions workflows for CI and PyPI publishing
423
+
424
+ ### v0.3.0 – v0.3.5
425
+ - **NEW**: OCR through Ollama multimodal models (`set_ollama_model`, `xtxt_image_describe`) — 0.3.0
426
+ - **NEW**: `set_ollama_config()`, `get_ollama_config()`, `reset_ollama_config()` — 0.3.2
427
+ - **NEW**: confidence score and hallucination detection (`xtxt_image_with_confidence`) — 0.3.4
428
+ - **NEW**: image enhancement before OCR, automatic model fallback, context presets — 0.3.4.2
429
+ - **NEW**: EXIF metadata extraction (`xtxt_exif`) — 0.3.5
430
+
431
+ ### v0.2.4
432
+ - **NEW**: video transcription support (MP4, MOV, AVI, WebM, MKV) via Whisper
433
+
434
+ ### v0.2.3
435
+ - **NEW**: audio transcription (MP3, WAV, M4A, FLAC, ...) with Whisper
436
+ - **NEW**: OCR from images (JPEG, PNG, TIFF, BMP, WebP) with EasyOCR
437
+ - **NEW**: `audio`, `ocr` and `all` installation extras
438
+
439
+ ### v0.2.0 – v0.2.2
440
+ - **NEW**: automatic extractor registration
441
+ - **NEW**: Markdown, EPUB, RTF, EML, MSG and LaTeX extractors
442
+
443
+ ### v0.1.24
444
+ - **NEW**: support for `bytes` and `requests.Response` inputs, `xtxt_from_url()` helper
445
+
446
+ ### v0.1.0 – v0.1.23
447
+ - Initial releases: modular extractors for PDF, DOCX, PPTX, XLSX, ODT, HTML, XML, TXT and legacy Office files,
448
+ MIME detection with python-magic, `BytesIO` support