papero-extract 3.0.0__tar.gz

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (51) hide show
  1. papero_extract-3.0.0/.gitignore +20 -0
  2. papero_extract-3.0.0/LICENSE +21 -0
  3. papero_extract-3.0.0/PKG-INFO +322 -0
  4. papero_extract-3.0.0/README.md +279 -0
  5. papero_extract-3.0.0/pyproject.toml +77 -0
  6. papero_extract-3.0.0/src/pdf_text_api/__init__.py +44 -0
  7. papero_extract-3.0.0/src/pdf_text_api/__main__.py +3 -0
  8. papero_extract-3.0.0/src/pdf_text_api/api.py +578 -0
  9. papero_extract-3.0.0/src/pdf_text_api/cleaning.py +114 -0
  10. papero_extract-3.0.0/src/pdf_text_api/cli.py +200 -0
  11. papero_extract-3.0.0/src/pdf_text_api/columns.py +120 -0
  12. papero_extract-3.0.0/src/pdf_text_api/config.py +63 -0
  13. papero_extract-3.0.0/src/pdf_text_api/extractor.py +493 -0
  14. papero_extract-3.0.0/src/pdf_text_api/layout.py +2620 -0
  15. papero_extract-3.0.0/src/pdf_text_api/model.py +207 -0
  16. papero_extract-3.0.0/src/pdf_text_api/render.py +302 -0
  17. papero_extract-3.0.0/src/pdf_text_api/symbols.py +189 -0
  18. papero_extract-3.0.0/src/pdf_text_api/tex_fonts.py +99 -0
  19. papero_extract-3.0.0/src/pdf_text_api/tika_client.py +219 -0
  20. papero_extract-3.0.0/src/pdf_text_api/tika_xhtml.py +237 -0
  21. papero_extract-3.0.0/tests/conftest.py +68 -0
  22. papero_extract-3.0.0/tests/fixtures.py +428 -0
  23. papero_extract-3.0.0/tests/js/expected.py +23 -0
  24. papero_extract-3.0.0/tests/js/package-lock.json +289 -0
  25. papero_extract-3.0.0/tests/js/package.json +7 -0
  26. papero_extract-3.0.0/tests/js/parity.mjs +42 -0
  27. papero_extract-3.0.0/tests/test_api.py +168 -0
  28. papero_extract-3.0.0/tests/test_cleaning.py +55 -0
  29. papero_extract-3.0.0/tests/test_extractor.py +81 -0
  30. papero_extract-3.0.0/tests/test_layout.py +231 -0
  31. papero_extract-3.0.0/web/assets/app.js +487 -0
  32. papero_extract-3.0.0/web/assets/badges/license.svg +1 -0
  33. papero_extract-3.0.0/web/assets/badges/ml-models.svg +1 -0
  34. papero_extract-3.0.0/web/assets/badges/python.svg +1 -0
  35. papero_extract-3.0.0/web/assets/columns.js +97 -0
  36. papero_extract-3.0.0/web/assets/demo.gif +0 -0
  37. papero_extract-3.0.0/web/assets/demo.mp4 +0 -0
  38. papero_extract-3.0.0/web/assets/engine.js +2051 -0
  39. papero_extract-3.0.0/web/assets/exemplo.pdf +0 -0
  40. papero_extract-3.0.0/web/assets/export.js +515 -0
  41. papero_extract-3.0.0/web/assets/favicon.png +0 -0
  42. papero_extract-3.0.0/web/assets/i18n.js +224 -0
  43. papero_extract-3.0.0/web/assets/icon.png +0 -0
  44. papero_extract-3.0.0/web/assets/icon.svg +15 -0
  45. papero_extract-3.0.0/web/assets/logo-dark.png +0 -0
  46. papero_extract-3.0.0/web/assets/logo-light.png +0 -0
  47. papero_extract-3.0.0/web/assets/logo.png +0 -0
  48. papero_extract-3.0.0/web/assets/screenshot.png +0 -0
  49. papero_extract-3.0.0/web/assets/style.css +323 -0
  50. papero_extract-3.0.0/web/assets/texfonts.js +119 -0
  51. papero_extract-3.0.0/web/index.html +226 -0
@@ -0,0 +1,20 @@
1
+ __pycache__/
2
+ *.py[cod]
3
+ *.egg-info/
4
+ build/
5
+ dist/
6
+ .venv/
7
+ venv/
8
+ .env
9
+ .pytest_cache/
10
+ .ruff_cache/
11
+ .coverage
12
+ *.log
13
+ benchmarks/dataset/
14
+ pdf_dificil/
15
+
16
+
17
+ # JS engine parity test
18
+ tests/js/node_modules/
19
+ tests/js/out/
20
+ tests/js/.engine.node.mjs
@@ -0,0 +1,21 @@
1
+ MIT License
2
+
3
+ Copyright (c) 2026 Beatriz Almeida
4
+
5
+ Permission is hereby granted, free of charge, to any person obtaining a copy
6
+ of this software and associated documentation files (the "Software"), to deal
7
+ in the Software without restriction, including without limitation the rights
8
+ to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
9
+ copies of the Software, and to permit persons to whom the Software is
10
+ furnished to do so, subject to the following conditions:
11
+
12
+ The above copyright notice and this permission notice shall be included in all
13
+ copies or substantial portions of the Software.
14
+
15
+ THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
16
+ IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
17
+ FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
18
+ AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
19
+ LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
20
+ OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
21
+ SOFTWARE.
@@ -0,0 +1,322 @@
1
+ Metadata-Version: 2.5
2
+ Name: papero-extract
3
+ Version: 3.0.0
4
+ Summary: papero — PDF to Markdown, JSON and tables: structured PDF extraction (reading order, tables, formulas, images, bounding boxes) on Apache Tika. Python library, CLI, REST API and browser app.
5
+ Project-URL: Homepage, https://github.com/beatrizalmeidaf/papero-pdf-text-extractor
6
+ Project-URL: Issues, https://github.com/beatrizalmeidaf/papero-pdf-text-extractor/issues
7
+ Project-URL: Documentation, https://github.com/beatrizalmeidaf/papero-pdf-text-extractor#readme
8
+ Project-URL: Web app, https://beatrizalmeidaf.github.io/papero-pdf-text-extractor/
9
+ Author: Beatriz Almeida
10
+ License-Expression: MIT
11
+ License-File: LICENSE
12
+ Keywords: apache-tika,document-parsing,docx,fastapi,formula-extraction,latex,layout-analysis,llm,markdown,ocr,pdf,pdf-extraction,pdf-parser,pdf-tables,pdf-to-json,pdf-to-markdown,pdf-to-text,pdfium,pptx,rag,reading-order,table-extraction,tika
13
+ Classifier: Development Status :: 4 - Beta
14
+ Classifier: Framework :: FastAPI
15
+ Classifier: Intended Audience :: Developers
16
+ Classifier: Natural Language :: Portuguese (Brazilian)
17
+ Classifier: Operating System :: OS Independent
18
+ Classifier: Programming Language :: Python :: 3
19
+ Classifier: Programming Language :: Python :: 3.10
20
+ Classifier: Programming Language :: Python :: 3.11
21
+ Classifier: Programming Language :: Python :: 3.12
22
+ Classifier: Programming Language :: Python :: 3.13
23
+ Classifier: Topic :: Scientific/Engineering :: Information Analysis
24
+ Classifier: Topic :: Text Processing
25
+ Classifier: Topic :: Text Processing :: Markup :: Markdown
26
+ Classifier: Topic :: Utilities
27
+ Requires-Python: >=3.10
28
+ Requires-Dist: pillow>=10
29
+ Requires-Dist: pypdfium2>=4.30
30
+ Requires-Dist: requests>=2.31
31
+ Requires-Dist: tika>=2.6.0
32
+ Provides-Extra: api
33
+ Requires-Dist: fastapi>=0.110; extra == 'api'
34
+ Requires-Dist: python-multipart>=0.0.9; extra == 'api'
35
+ Requires-Dist: uvicorn[standard]>=0.29; extra == 'api'
36
+ Provides-Extra: dev
37
+ Requires-Dist: fpdf2>=2.7; extra == 'dev'
38
+ Requires-Dist: httpx>=0.27; extra == 'dev'
39
+ Requires-Dist: pdf-text-api[api]; extra == 'dev'
40
+ Requires-Dist: pytest>=8; extra == 'dev'
41
+ Requires-Dist: ruff>=0.5; extra == 'dev'
42
+ Description-Content-Type: text/markdown
43
+
44
+ <p align="center">
45
+ <picture>
46
+ <source media="(prefers-color-scheme: dark)" srcset="web/assets/logo-dark.png">
47
+ <img src="web/assets/logo-light.png" alt="papero" width="420">
48
+ </picture>
49
+ </p>
50
+
51
+ <h3 align="center">Document structure extraction without the heavyweight stack.</h3>
52
+
53
+ <p align="center">
54
+ PDF → Markdown · JSON · Word · Excel — with reading order, tables, formulas, figures and the position of every block.<br>
55
+ CPU only. No ML models. Runs in your browser, in Python, or as an API.
56
+ </p>
57
+
58
+ <p align="center">
59
+ <a href="https://github.com/beatrizalmeidaf/papero-pdf-text-extractor/actions/workflows/ci.yml"><img src="https://github.com/beatrizalmeidaf/papero-pdf-text-extractor/actions/workflows/ci.yml/badge.svg" alt="CI"></a>
60
+ <img src="web/assets/badges/python.svg" alt="Python 3.10–3.13">
61
+ <a href="LICENSE"><img src="web/assets/badges/license.svg" alt="MIT license"></a>
62
+ <img src="web/assets/badges/ml-models.svg" alt="No ML models">
63
+ </p>
64
+
65
+ <p align="center">
66
+ <a href="https://beatrizalmeidaf.github.io/papero-pdf-text-extractor/"><b>▶ Try it in your browser</b></a>
67
+ &nbsp;·&nbsp;
68
+ <a href="#quick-start"><b>Quick start</b></a>
69
+ &nbsp;·&nbsp;
70
+ <a href="#benchmarks"><b>Benchmarks</b></a>
71
+ </p>
72
+
73
+ <p align="center">
74
+ <a href="https://beatrizalmeidaf.github.io/papero-pdf-text-extractor/">
75
+ <img src="web/assets/demo.gif" alt="papero demo: load a PDF, every block outlined on the page, click a table to inspect it, see tables and LaTeX formulas, export to Word" width="900">
76
+ </a>
77
+ <br>
78
+ <sub>30 seconds in the browser app: load a PDF, inspect any block, check tables and formulas, export to Word. Your PDF never leaves your machine. (<a href="web/assets/demo.mp4">MP4</a>)</sub>
79
+ </p>
80
+
81
+ ---
82
+
83
+ ## Why papero
84
+
85
+ Getting the *text* out of a PDF is easy. Getting its **structure** back — which column comes first, which lines are a table, where the formula is — is what makes the output usable for RAG, search and LLMs. papero does that with plain geometry, so it stays fast on a laptop CPU.
86
+
87
+ <table>
88
+ <tr>
89
+ <td width="33%" valign="top">
90
+ <b>📖 Reading order</b><br>
91
+ Two- and three-column papers read column by column. Headers, footers, page numbers and repeated logos are set aside.
92
+ </td>
93
+ <td width="33%" valign="top">
94
+ <b>▦ Real tables</b><br>
95
+ Ruled, borderless and LaTeX <i>booktabs</i> tables come back as rows and columns — multi-line cells included. Export to CSV or Excel.
96
+ </td>
97
+ <td width="33%" valign="top">
98
+ <b>∑ Formulas</b><br>
99
+ Superscripts, subscripts and math symbols become LaTeX (<code>E = mc^{2}</code>), plus a cropped image of the formula.
100
+ </td>
101
+ </tr>
102
+ <tr>
103
+ <td valign="top">
104
+ <b>📍 Position of everything</b><br>
105
+ Every block has a bounding box — cite the exact spot in a RAG answer, draw over the page, or crop it.
106
+ </td>
107
+ <td valign="top">
108
+ <b>🖼 Figures &amp; charts</b><br>
109
+ Images and vector charts are cropped to PNG, with their caption, axis labels and legend kept together.
110
+ </td>
111
+ <td valign="top">
112
+ <b>📝 Back to Word</b><br>
113
+ Alignment, indents, line spacing, bold runs and fonts are kept, so a <code>.docx</code> export looks like the original page.
114
+ </td>
115
+ </tr>
116
+ </table>
117
+
118
+ Also: accents drawn as separate glyphs in LaTeX PDFs (`Computa¸ca˜o` → `Computação`), invisible white text used by form generators is dropped, scanned pages go through OCR, and DOCX/PPTX/XLSX/EPUB/HTML are read through Apache Tika.
119
+
120
+ ## Quick start
121
+
122
+ ```bash
123
+ pip install pdf-text-api
124
+ ```
125
+
126
+ ```python
127
+ from pdf_text_api import extract
128
+
129
+ doc = extract("paper.pdf")
130
+ print(doc.to_markdown())
131
+ ```
132
+
133
+ Or skip the install: **[open the browser app](https://beatrizalmeidaf.github.io/papero-pdf-text-extractor/)**, drop a PDF, export to the format you need.
134
+
135
+ <details>
136
+ <summary><b>More Python</b> — tables, formulas, positions, images, options</summary>
137
+
138
+ ```python
139
+ from pdf_text_api import extract, extract_text
140
+
141
+ doc = extract("paper.pdf", images=True)
142
+
143
+ doc.tables[0].rows # [["Model", "Accuracy"], ["Base", "0.81"], ...]
144
+ doc.formulas[0].latex # "E = mc^{2}"
145
+ doc.figures[0].image.data # PNG bytes
146
+
147
+ for block in doc.pages[0].blocks: # reading order, with positions
148
+ print(block.type, block.bbox, block.text[:60])
149
+
150
+ doc.to_html() # keeps alignment and indents
151
+ doc.to_dict() # the full JSON
152
+
153
+ extract("slides.pptx").to_markdown() # any format Apache Tika reads
154
+ extract_text("contract.pdf").text # fastest: clean text only
155
+ ```
156
+
157
+ | Option | Default | |
158
+ |---|---|---|
159
+ | `pages` | all | `"1-3,5,10-"` |
160
+ | `images` | `False` | crop figures, tables and formulas to PNG |
161
+ | `tables` / `formulas` | `True` | detection on/off |
162
+ | `ocr` | `"auto"` | `"auto"` (scanned pages only), `"force"`, `"off"` |
163
+ | `ocr_language` | `"por+eng"` | Tesseract languages |
164
+ | `tika` | `True` | `False` runs the layout engine alone (no Java) |
165
+ | `workers` | `1` | processes for long documents |
166
+
167
+ </details>
168
+
169
+ <details>
170
+ <summary><b>CLI</b></summary>
171
+
172
+ ```bash
173
+ pdf-text-api extract paper.pdf -o paper.md --images # Markdown + images/ folder
174
+ pdf-text-api extract paper.pdf -o paper.json # format from the extension
175
+ pdf-text-api extract paper.pdf -f csv -o tables.csv # tables only
176
+ pdf-text-api extract paper.pdf -p 1-5 -f html
177
+ pdf-text-api extract paper.pdf --fast # clean text only
178
+ pdf-text-api serve --port 8000 # API + browser app
179
+ ```
180
+
181
+ </details>
182
+
183
+ <details>
184
+ <summary><b>REST API &amp; Docker</b></summary>
185
+
186
+ ```bash
187
+ docker compose up # API + Apache Tika + Tesseract + browser app on :8000
188
+ ```
189
+
190
+ ```bash
191
+ curl -F "file=@paper.pdf" "localhost:8000/v1/extract?format=markdown"
192
+ curl -F "file=@paper.pdf" "localhost:8000/v1/extract?format=zip&images=true" -o paper.zip
193
+ curl -F "file=@paper.pdf" "localhost:8000/v1/extract?per_page=true" # blocks + positions
194
+ ```
195
+
196
+ One endpoint, `POST /v1/extract`; interactive docs at `/docs`.
197
+
198
+ | Parameter | Default | |
199
+ |---|---|---|
200
+ | `mode` | `structured` | `structured` (layout + Tika) or `fast` (text only) |
201
+ | `format` | `json` | `json`, `markdown`, `text`, `html`, `csv`, `zip` |
202
+ | `pages` | all | `1-3,5,10-` |
203
+ | `per_page` | `false` | include pages, blocks and positions in the JSON |
204
+ | `images` | `false` | crop figures, tables and formulas |
205
+ | `ocr` | `auto` | `auto`, `force`, `off` |
206
+
207
+ Configuration through environment variables — see [`.env.example`](.env.example).
208
+
209
+ </details>
210
+
211
+ ## What comes out
212
+
213
+ Every block knows what it is and where it was:
214
+
215
+ ```json
216
+ {
217
+ "type": "table",
218
+ "bbox": [56.7, 294.8, 481.9, 374.2],
219
+ "rows": [["Model", "Accuracy"], ["Base", "0.81"]],
220
+ "caption": "Table 1: Comparison between models."
221
+ }
222
+ ```
223
+
224
+ | Output | Python · CLI · API | Browser app |
225
+ |---|:-:|:-:|
226
+ | Markdown, plain text, JSON | ✓ | ✓ |
227
+ | HTML (keeps alignment and indents) | ✓ | ✓ |
228
+ | CSV of the tables, ZIP with images | ✓ | ✓ |
229
+ | Word `.docx` that keeps the page's look | — | ✓ |
230
+ | Excel `.xlsx`, one sheet per table | — | ✓ |
231
+
232
+ <details>
233
+ <summary><b>Full JSON schema and block types</b></summary>
234
+
235
+ ```json
236
+ {
237
+ "schema": "pdf-text-api/document@1",
238
+ "engine": "tika+pdfium",
239
+ "page_count": 12,
240
+ "metadata": { "title": "...", "author": "...", "language": "en" },
241
+ "pages": [{
242
+ "number": 1, "width": 595.3, "height": 841.9,
243
+ "blocks": [{
244
+ "id": "p1-b4", "type": "paragraph", "bbox": [74.0, 217.0, 522.0, 275.0],
245
+ "text": "Atestamos que a estudante ...",
246
+ "style": { "pt": 11.0, "font": "Arial", "bold": false },
247
+ "format": { "align": "justify", "first_line": 42.7, "line_spacing": 1.8 },
248
+ "runs": [{ "text": "FULANA DE TAL", "bold": true, "italic": false, "script": null }]
249
+ }]
250
+ }]
251
+ }
252
+ ```
253
+
254
+ Block types: `heading` (with `level`), `paragraph`, `list_item` (with `marker`), `table` (with `rows`), `figure`, `formula` (with `latex`), `caption`, `code`, and — kept apart from the text — `header`, `footer`, `page_number`. Bounding boxes are `[x0, y0, x1, y1]` in points, origin at the top-left of the page.
255
+
256
+ </details>
257
+
258
+ ## Benchmarks
259
+
260
+ <p align="center">
261
+ <img src="benchmarks/latency.svg" alt="Average extraction time per PDF on a log scale: PyMuPDF 97 ms, papero fast 136 ms, papero structured 543 ms, pypdf 1.52 s, pdfplumber 3.56 s, Docling 82.9 s" width="760">
262
+ </p>
263
+
264
+ Dense arXiv papers (multi-column, formulas, tables, figures) on one laptop CPU, no GPU. `papero · fast` returns clean text; `papero · structured` also rebuilds reading order, tables, formulas and figures — **0 failures on 54 papers, 39 ms per page (median)**. Reproduce with [`benchmarks/`](benchmarks/).
265
+
266
+ | | papero | PyMuPDF | pdfplumber | pypdf | Docling | Marker |
267
+ |---|:-:|:-:|:-:|:-:|:-:|:-:|
268
+ | License | MIT | AGPL | MIT | BSD | MIT | GPL |
269
+ | Needs ML models / PyTorch | no | no | no | no | yes | yes |
270
+ | Multi-column reading order | ✓ | partial | — | — | ✓ | ✓ |
271
+ | Structured tables | ✓ | ✓ | ✓ | — | ✓ | ✓ |
272
+ | Formulas | LaTeX from glyphs + image | — | — | — | ✓ | ✓ |
273
+ | Bounding boxes | ✓ | ✓ | ✓ | — | ✓ | ✓ |
274
+ | DOCX / PPTX / XLSX / EPUB | ✓ | partial | — | — | ✓ | partial |
275
+ | Runs entirely in the browser | ✓ | — | — | — | — | — |
276
+
277
+ ML-based tools still win on very irregular layouts and complex math (stacked fractions, matrices) — papero gives you the formula as approximate LaTeX **and** as an image so nothing is lost.
278
+
279
+ ## How it works
280
+
281
+ Two engines run on the same file **at the same time**:
282
+
283
+ - **A layout engine on PDFium** reads every glyph with its position, font and size, plus every rule and image, and rebuilds columns, tables, formulas, lists and figures with a column-aware XY-cut.
284
+ - **Apache Tika** adds metadata, tagged-PDF headings, OCR (Tesseract) and every non-PDF format.
285
+
286
+ The browser app runs the same algorithm ported to JavaScript on pdf.js, and CI checks block by block that both engines agree.
287
+
288
+ <details>
289
+ <summary><b>Limitations</b></summary>
290
+
291
+ - **Math:** LaTeX is rebuilt from glyphs — stacked fractions, matrices and big radicals come out linear (the cropped image is always there).
292
+ - **Borderless tables** with very narrow gaps between columns can read as text.
293
+ - **Scanned PDFs** need OCR, which runs on the server path (Tesseract is in the Docker image).
294
+ - **Word/Excel export** is in the browser app for now.
295
+
296
+ </details>
297
+
298
+ <details>
299
+ <summary><b>Development</b></summary>
300
+
301
+ ```bash
302
+ git clone https://github.com/beatrizalmeidaf/papero-pdf-text-extractor.git && cd pdf-text-extractor
303
+ pip install -e ".[dev]"
304
+ pytest -q # includes real-world regressions
305
+ ruff check src tests && ruff format --check src tests
306
+ npm install --prefix tests/js && python tests/js/expected.py tests/js/out && node tests/js/parity.mjs tests/js/out
307
+ python -m http.server -d web # browser app at http://localhost:8000
308
+ ```
309
+
310
+ `src/pdf_text_api/` is the Python engine, API and CLI · `web/` is the browser app (GitHub Pages) · `tests/js/` checks the two engines agree · `benchmarks/` downloads the dataset and draws the chart.
311
+
312
+ </details>
313
+
314
+ ## Contributing
315
+
316
+ Found a PDF papero gets wrong? **That's the most useful issue you can open** — attach the file (or a page of it) and say what you expected. Reading order, tables, formulas, encoding, OCR and browser/server differences are all fair game.
317
+
318
+ If papero saves you time, **a ⭐ helps other people find it.**
319
+
320
+ <sub>**Keywords:** PDF to Markdown · PDF to JSON · PDF to Word · PDF to Excel · PDF table extraction · PDF parser · document parsing · layout analysis · reading order · multi-column PDF · formula extraction · LaTeX · bounding boxes · OCR · Apache Tika · PDFium · pdf.js · RAG preprocessing · LLM document loader · Docling alternative · PyMuPDF alternative · converter PDF para Markdown, Word e Excel · extrair tabelas de PDF · extrair texto de PDF mantendo a formatação · OCR de PDF escaneado</sub>
321
+
322
+ <sub>MIT © Beatriz Almeida · package and imports keep the name `pdf-text-api` / `pdf_text_api` for compatibility.</sub>
@@ -0,0 +1,279 @@
1
+ <p align="center">
2
+ <picture>
3
+ <source media="(prefers-color-scheme: dark)" srcset="web/assets/logo-dark.png">
4
+ <img src="web/assets/logo-light.png" alt="papero" width="420">
5
+ </picture>
6
+ </p>
7
+
8
+ <h3 align="center">Document structure extraction without the heavyweight stack.</h3>
9
+
10
+ <p align="center">
11
+ PDF → Markdown · JSON · Word · Excel — with reading order, tables, formulas, figures and the position of every block.<br>
12
+ CPU only. No ML models. Runs in your browser, in Python, or as an API.
13
+ </p>
14
+
15
+ <p align="center">
16
+ <a href="https://github.com/beatrizalmeidaf/papero-pdf-text-extractor/actions/workflows/ci.yml"><img src="https://github.com/beatrizalmeidaf/papero-pdf-text-extractor/actions/workflows/ci.yml/badge.svg" alt="CI"></a>
17
+ <img src="web/assets/badges/python.svg" alt="Python 3.10–3.13">
18
+ <a href="LICENSE"><img src="web/assets/badges/license.svg" alt="MIT license"></a>
19
+ <img src="web/assets/badges/ml-models.svg" alt="No ML models">
20
+ </p>
21
+
22
+ <p align="center">
23
+ <a href="https://beatrizalmeidaf.github.io/papero-pdf-text-extractor/"><b>▶ Try it in your browser</b></a>
24
+ &nbsp;·&nbsp;
25
+ <a href="#quick-start"><b>Quick start</b></a>
26
+ &nbsp;·&nbsp;
27
+ <a href="#benchmarks"><b>Benchmarks</b></a>
28
+ </p>
29
+
30
+ <p align="center">
31
+ <a href="https://beatrizalmeidaf.github.io/papero-pdf-text-extractor/">
32
+ <img src="web/assets/demo.gif" alt="papero demo: load a PDF, every block outlined on the page, click a table to inspect it, see tables and LaTeX formulas, export to Word" width="900">
33
+ </a>
34
+ <br>
35
+ <sub>30 seconds in the browser app: load a PDF, inspect any block, check tables and formulas, export to Word. Your PDF never leaves your machine. (<a href="web/assets/demo.mp4">MP4</a>)</sub>
36
+ </p>
37
+
38
+ ---
39
+
40
+ ## Why papero
41
+
42
+ Getting the *text* out of a PDF is easy. Getting its **structure** back — which column comes first, which lines are a table, where the formula is — is what makes the output usable for RAG, search and LLMs. papero does that with plain geometry, so it stays fast on a laptop CPU.
43
+
44
+ <table>
45
+ <tr>
46
+ <td width="33%" valign="top">
47
+ <b>📖 Reading order</b><br>
48
+ Two- and three-column papers read column by column. Headers, footers, page numbers and repeated logos are set aside.
49
+ </td>
50
+ <td width="33%" valign="top">
51
+ <b>▦ Real tables</b><br>
52
+ Ruled, borderless and LaTeX <i>booktabs</i> tables come back as rows and columns — multi-line cells included. Export to CSV or Excel.
53
+ </td>
54
+ <td width="33%" valign="top">
55
+ <b>∑ Formulas</b><br>
56
+ Superscripts, subscripts and math symbols become LaTeX (<code>E = mc^{2}</code>), plus a cropped image of the formula.
57
+ </td>
58
+ </tr>
59
+ <tr>
60
+ <td valign="top">
61
+ <b>📍 Position of everything</b><br>
62
+ Every block has a bounding box — cite the exact spot in a RAG answer, draw over the page, or crop it.
63
+ </td>
64
+ <td valign="top">
65
+ <b>🖼 Figures &amp; charts</b><br>
66
+ Images and vector charts are cropped to PNG, with their caption, axis labels and legend kept together.
67
+ </td>
68
+ <td valign="top">
69
+ <b>📝 Back to Word</b><br>
70
+ Alignment, indents, line spacing, bold runs and fonts are kept, so a <code>.docx</code> export looks like the original page.
71
+ </td>
72
+ </tr>
73
+ </table>
74
+
75
+ Also: accents drawn as separate glyphs in LaTeX PDFs (`Computa¸ca˜o` → `Computação`), invisible white text used by form generators is dropped, scanned pages go through OCR, and DOCX/PPTX/XLSX/EPUB/HTML are read through Apache Tika.
76
+
77
+ ## Quick start
78
+
79
+ ```bash
80
+ pip install pdf-text-api
81
+ ```
82
+
83
+ ```python
84
+ from pdf_text_api import extract
85
+
86
+ doc = extract("paper.pdf")
87
+ print(doc.to_markdown())
88
+ ```
89
+
90
+ Or skip the install: **[open the browser app](https://beatrizalmeidaf.github.io/papero-pdf-text-extractor/)**, drop a PDF, export to the format you need.
91
+
92
+ <details>
93
+ <summary><b>More Python</b> — tables, formulas, positions, images, options</summary>
94
+
95
+ ```python
96
+ from pdf_text_api import extract, extract_text
97
+
98
+ doc = extract("paper.pdf", images=True)
99
+
100
+ doc.tables[0].rows # [["Model", "Accuracy"], ["Base", "0.81"], ...]
101
+ doc.formulas[0].latex # "E = mc^{2}"
102
+ doc.figures[0].image.data # PNG bytes
103
+
104
+ for block in doc.pages[0].blocks: # reading order, with positions
105
+ print(block.type, block.bbox, block.text[:60])
106
+
107
+ doc.to_html() # keeps alignment and indents
108
+ doc.to_dict() # the full JSON
109
+
110
+ extract("slides.pptx").to_markdown() # any format Apache Tika reads
111
+ extract_text("contract.pdf").text # fastest: clean text only
112
+ ```
113
+
114
+ | Option | Default | |
115
+ |---|---|---|
116
+ | `pages` | all | `"1-3,5,10-"` |
117
+ | `images` | `False` | crop figures, tables and formulas to PNG |
118
+ | `tables` / `formulas` | `True` | detection on/off |
119
+ | `ocr` | `"auto"` | `"auto"` (scanned pages only), `"force"`, `"off"` |
120
+ | `ocr_language` | `"por+eng"` | Tesseract languages |
121
+ | `tika` | `True` | `False` runs the layout engine alone (no Java) |
122
+ | `workers` | `1` | processes for long documents |
123
+
124
+ </details>
125
+
126
+ <details>
127
+ <summary><b>CLI</b></summary>
128
+
129
+ ```bash
130
+ pdf-text-api extract paper.pdf -o paper.md --images # Markdown + images/ folder
131
+ pdf-text-api extract paper.pdf -o paper.json # format from the extension
132
+ pdf-text-api extract paper.pdf -f csv -o tables.csv # tables only
133
+ pdf-text-api extract paper.pdf -p 1-5 -f html
134
+ pdf-text-api extract paper.pdf --fast # clean text only
135
+ pdf-text-api serve --port 8000 # API + browser app
136
+ ```
137
+
138
+ </details>
139
+
140
+ <details>
141
+ <summary><b>REST API &amp; Docker</b></summary>
142
+
143
+ ```bash
144
+ docker compose up # API + Apache Tika + Tesseract + browser app on :8000
145
+ ```
146
+
147
+ ```bash
148
+ curl -F "file=@paper.pdf" "localhost:8000/v1/extract?format=markdown"
149
+ curl -F "file=@paper.pdf" "localhost:8000/v1/extract?format=zip&images=true" -o paper.zip
150
+ curl -F "file=@paper.pdf" "localhost:8000/v1/extract?per_page=true" # blocks + positions
151
+ ```
152
+
153
+ One endpoint, `POST /v1/extract`; interactive docs at `/docs`.
154
+
155
+ | Parameter | Default | |
156
+ |---|---|---|
157
+ | `mode` | `structured` | `structured` (layout + Tika) or `fast` (text only) |
158
+ | `format` | `json` | `json`, `markdown`, `text`, `html`, `csv`, `zip` |
159
+ | `pages` | all | `1-3,5,10-` |
160
+ | `per_page` | `false` | include pages, blocks and positions in the JSON |
161
+ | `images` | `false` | crop figures, tables and formulas |
162
+ | `ocr` | `auto` | `auto`, `force`, `off` |
163
+
164
+ Configuration through environment variables — see [`.env.example`](.env.example).
165
+
166
+ </details>
167
+
168
+ ## What comes out
169
+
170
+ Every block knows what it is and where it was:
171
+
172
+ ```json
173
+ {
174
+ "type": "table",
175
+ "bbox": [56.7, 294.8, 481.9, 374.2],
176
+ "rows": [["Model", "Accuracy"], ["Base", "0.81"]],
177
+ "caption": "Table 1: Comparison between models."
178
+ }
179
+ ```
180
+
181
+ | Output | Python · CLI · API | Browser app |
182
+ |---|:-:|:-:|
183
+ | Markdown, plain text, JSON | ✓ | ✓ |
184
+ | HTML (keeps alignment and indents) | ✓ | ✓ |
185
+ | CSV of the tables, ZIP with images | ✓ | ✓ |
186
+ | Word `.docx` that keeps the page's look | — | ✓ |
187
+ | Excel `.xlsx`, one sheet per table | — | ✓ |
188
+
189
+ <details>
190
+ <summary><b>Full JSON schema and block types</b></summary>
191
+
192
+ ```json
193
+ {
194
+ "schema": "pdf-text-api/document@1",
195
+ "engine": "tika+pdfium",
196
+ "page_count": 12,
197
+ "metadata": { "title": "...", "author": "...", "language": "en" },
198
+ "pages": [{
199
+ "number": 1, "width": 595.3, "height": 841.9,
200
+ "blocks": [{
201
+ "id": "p1-b4", "type": "paragraph", "bbox": [74.0, 217.0, 522.0, 275.0],
202
+ "text": "Atestamos que a estudante ...",
203
+ "style": { "pt": 11.0, "font": "Arial", "bold": false },
204
+ "format": { "align": "justify", "first_line": 42.7, "line_spacing": 1.8 },
205
+ "runs": [{ "text": "FULANA DE TAL", "bold": true, "italic": false, "script": null }]
206
+ }]
207
+ }]
208
+ }
209
+ ```
210
+
211
+ Block types: `heading` (with `level`), `paragraph`, `list_item` (with `marker`), `table` (with `rows`), `figure`, `formula` (with `latex`), `caption`, `code`, and — kept apart from the text — `header`, `footer`, `page_number`. Bounding boxes are `[x0, y0, x1, y1]` in points, origin at the top-left of the page.
212
+
213
+ </details>
214
+
215
+ ## Benchmarks
216
+
217
+ <p align="center">
218
+ <img src="benchmarks/latency.svg" alt="Average extraction time per PDF on a log scale: PyMuPDF 97 ms, papero fast 136 ms, papero structured 543 ms, pypdf 1.52 s, pdfplumber 3.56 s, Docling 82.9 s" width="760">
219
+ </p>
220
+
221
+ Dense arXiv papers (multi-column, formulas, tables, figures) on one laptop CPU, no GPU. `papero · fast` returns clean text; `papero · structured` also rebuilds reading order, tables, formulas and figures — **0 failures on 54 papers, 39 ms per page (median)**. Reproduce with [`benchmarks/`](benchmarks/).
222
+
223
+ | | papero | PyMuPDF | pdfplumber | pypdf | Docling | Marker |
224
+ |---|:-:|:-:|:-:|:-:|:-:|:-:|
225
+ | License | MIT | AGPL | MIT | BSD | MIT | GPL |
226
+ | Needs ML models / PyTorch | no | no | no | no | yes | yes |
227
+ | Multi-column reading order | ✓ | partial | — | — | ✓ | ✓ |
228
+ | Structured tables | ✓ | ✓ | ✓ | — | ✓ | ✓ |
229
+ | Formulas | LaTeX from glyphs + image | — | — | — | ✓ | ✓ |
230
+ | Bounding boxes | ✓ | ✓ | ✓ | — | ✓ | ✓ |
231
+ | DOCX / PPTX / XLSX / EPUB | ✓ | partial | — | — | ✓ | partial |
232
+ | Runs entirely in the browser | ✓ | — | — | — | — | — |
233
+
234
+ ML-based tools still win on very irregular layouts and complex math (stacked fractions, matrices) — papero gives you the formula as approximate LaTeX **and** as an image so nothing is lost.
235
+
236
+ ## How it works
237
+
238
+ Two engines run on the same file **at the same time**:
239
+
240
+ - **A layout engine on PDFium** reads every glyph with its position, font and size, plus every rule and image, and rebuilds columns, tables, formulas, lists and figures with a column-aware XY-cut.
241
+ - **Apache Tika** adds metadata, tagged-PDF headings, OCR (Tesseract) and every non-PDF format.
242
+
243
+ The browser app runs the same algorithm ported to JavaScript on pdf.js, and CI checks block by block that both engines agree.
244
+
245
+ <details>
246
+ <summary><b>Limitations</b></summary>
247
+
248
+ - **Math:** LaTeX is rebuilt from glyphs — stacked fractions, matrices and big radicals come out linear (the cropped image is always there).
249
+ - **Borderless tables** with very narrow gaps between columns can read as text.
250
+ - **Scanned PDFs** need OCR, which runs on the server path (Tesseract is in the Docker image).
251
+ - **Word/Excel export** is in the browser app for now.
252
+
253
+ </details>
254
+
255
+ <details>
256
+ <summary><b>Development</b></summary>
257
+
258
+ ```bash
259
+ git clone https://github.com/beatrizalmeidaf/papero-pdf-text-extractor.git && cd pdf-text-extractor
260
+ pip install -e ".[dev]"
261
+ pytest -q # includes real-world regressions
262
+ ruff check src tests && ruff format --check src tests
263
+ npm install --prefix tests/js && python tests/js/expected.py tests/js/out && node tests/js/parity.mjs tests/js/out
264
+ python -m http.server -d web # browser app at http://localhost:8000
265
+ ```
266
+
267
+ `src/pdf_text_api/` is the Python engine, API and CLI · `web/` is the browser app (GitHub Pages) · `tests/js/` checks the two engines agree · `benchmarks/` downloads the dataset and draws the chart.
268
+
269
+ </details>
270
+
271
+ ## Contributing
272
+
273
+ Found a PDF papero gets wrong? **That's the most useful issue you can open** — attach the file (or a page of it) and say what you expected. Reading order, tables, formulas, encoding, OCR and browser/server differences are all fair game.
274
+
275
+ If papero saves you time, **a ⭐ helps other people find it.**
276
+
277
+ <sub>**Keywords:** PDF to Markdown · PDF to JSON · PDF to Word · PDF to Excel · PDF table extraction · PDF parser · document parsing · layout analysis · reading order · multi-column PDF · formula extraction · LaTeX · bounding boxes · OCR · Apache Tika · PDFium · pdf.js · RAG preprocessing · LLM document loader · Docling alternative · PyMuPDF alternative · converter PDF para Markdown, Word e Excel · extrair tabelas de PDF · extrair texto de PDF mantendo a formatação · OCR de PDF escaneado</sub>
278
+
279
+ <sub>MIT © Beatriz Almeida · package and imports keep the name `pdf-text-api` / `pdf_text_api` for compatibility.</sub>