papero-extract 3.0.0__tar.gz
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- papero_extract-3.0.0/.gitignore +20 -0
- papero_extract-3.0.0/LICENSE +21 -0
- papero_extract-3.0.0/PKG-INFO +322 -0
- papero_extract-3.0.0/README.md +279 -0
- papero_extract-3.0.0/pyproject.toml +77 -0
- papero_extract-3.0.0/src/pdf_text_api/__init__.py +44 -0
- papero_extract-3.0.0/src/pdf_text_api/__main__.py +3 -0
- papero_extract-3.0.0/src/pdf_text_api/api.py +578 -0
- papero_extract-3.0.0/src/pdf_text_api/cleaning.py +114 -0
- papero_extract-3.0.0/src/pdf_text_api/cli.py +200 -0
- papero_extract-3.0.0/src/pdf_text_api/columns.py +120 -0
- papero_extract-3.0.0/src/pdf_text_api/config.py +63 -0
- papero_extract-3.0.0/src/pdf_text_api/extractor.py +493 -0
- papero_extract-3.0.0/src/pdf_text_api/layout.py +2620 -0
- papero_extract-3.0.0/src/pdf_text_api/model.py +207 -0
- papero_extract-3.0.0/src/pdf_text_api/render.py +302 -0
- papero_extract-3.0.0/src/pdf_text_api/symbols.py +189 -0
- papero_extract-3.0.0/src/pdf_text_api/tex_fonts.py +99 -0
- papero_extract-3.0.0/src/pdf_text_api/tika_client.py +219 -0
- papero_extract-3.0.0/src/pdf_text_api/tika_xhtml.py +237 -0
- papero_extract-3.0.0/tests/conftest.py +68 -0
- papero_extract-3.0.0/tests/fixtures.py +428 -0
- papero_extract-3.0.0/tests/js/expected.py +23 -0
- papero_extract-3.0.0/tests/js/package-lock.json +289 -0
- papero_extract-3.0.0/tests/js/package.json +7 -0
- papero_extract-3.0.0/tests/js/parity.mjs +42 -0
- papero_extract-3.0.0/tests/test_api.py +168 -0
- papero_extract-3.0.0/tests/test_cleaning.py +55 -0
- papero_extract-3.0.0/tests/test_extractor.py +81 -0
- papero_extract-3.0.0/tests/test_layout.py +231 -0
- papero_extract-3.0.0/web/assets/app.js +487 -0
- papero_extract-3.0.0/web/assets/badges/license.svg +1 -0
- papero_extract-3.0.0/web/assets/badges/ml-models.svg +1 -0
- papero_extract-3.0.0/web/assets/badges/python.svg +1 -0
- papero_extract-3.0.0/web/assets/columns.js +97 -0
- papero_extract-3.0.0/web/assets/demo.gif +0 -0
- papero_extract-3.0.0/web/assets/demo.mp4 +0 -0
- papero_extract-3.0.0/web/assets/engine.js +2051 -0
- papero_extract-3.0.0/web/assets/exemplo.pdf +0 -0
- papero_extract-3.0.0/web/assets/export.js +515 -0
- papero_extract-3.0.0/web/assets/favicon.png +0 -0
- papero_extract-3.0.0/web/assets/i18n.js +224 -0
- papero_extract-3.0.0/web/assets/icon.png +0 -0
- papero_extract-3.0.0/web/assets/icon.svg +15 -0
- papero_extract-3.0.0/web/assets/logo-dark.png +0 -0
- papero_extract-3.0.0/web/assets/logo-light.png +0 -0
- papero_extract-3.0.0/web/assets/logo.png +0 -0
- papero_extract-3.0.0/web/assets/screenshot.png +0 -0
- papero_extract-3.0.0/web/assets/style.css +323 -0
- papero_extract-3.0.0/web/assets/texfonts.js +119 -0
- papero_extract-3.0.0/web/index.html +226 -0
|
@@ -0,0 +1,20 @@
|
|
|
1
|
+
__pycache__/
|
|
2
|
+
*.py[cod]
|
|
3
|
+
*.egg-info/
|
|
4
|
+
build/
|
|
5
|
+
dist/
|
|
6
|
+
.venv/
|
|
7
|
+
venv/
|
|
8
|
+
.env
|
|
9
|
+
.pytest_cache/
|
|
10
|
+
.ruff_cache/
|
|
11
|
+
.coverage
|
|
12
|
+
*.log
|
|
13
|
+
benchmarks/dataset/
|
|
14
|
+
pdf_dificil/
|
|
15
|
+
|
|
16
|
+
|
|
17
|
+
# JS engine parity test
|
|
18
|
+
tests/js/node_modules/
|
|
19
|
+
tests/js/out/
|
|
20
|
+
tests/js/.engine.node.mjs
|
|
@@ -0,0 +1,21 @@
|
|
|
1
|
+
MIT License
|
|
2
|
+
|
|
3
|
+
Copyright (c) 2026 Beatriz Almeida
|
|
4
|
+
|
|
5
|
+
Permission is hereby granted, free of charge, to any person obtaining a copy
|
|
6
|
+
of this software and associated documentation files (the "Software"), to deal
|
|
7
|
+
in the Software without restriction, including without limitation the rights
|
|
8
|
+
to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
|
|
9
|
+
copies of the Software, and to permit persons to whom the Software is
|
|
10
|
+
furnished to do so, subject to the following conditions:
|
|
11
|
+
|
|
12
|
+
The above copyright notice and this permission notice shall be included in all
|
|
13
|
+
copies or substantial portions of the Software.
|
|
14
|
+
|
|
15
|
+
THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
|
|
16
|
+
IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
|
|
17
|
+
FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
|
|
18
|
+
AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
|
|
19
|
+
LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
|
|
20
|
+
OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
|
|
21
|
+
SOFTWARE.
|
|
@@ -0,0 +1,322 @@
|
|
|
1
|
+
Metadata-Version: 2.5
|
|
2
|
+
Name: papero-extract
|
|
3
|
+
Version: 3.0.0
|
|
4
|
+
Summary: papero — PDF to Markdown, JSON and tables: structured PDF extraction (reading order, tables, formulas, images, bounding boxes) on Apache Tika. Python library, CLI, REST API and browser app.
|
|
5
|
+
Project-URL: Homepage, https://github.com/beatrizalmeidaf/papero-pdf-text-extractor
|
|
6
|
+
Project-URL: Issues, https://github.com/beatrizalmeidaf/papero-pdf-text-extractor/issues
|
|
7
|
+
Project-URL: Documentation, https://github.com/beatrizalmeidaf/papero-pdf-text-extractor#readme
|
|
8
|
+
Project-URL: Web app, https://beatrizalmeidaf.github.io/papero-pdf-text-extractor/
|
|
9
|
+
Author: Beatriz Almeida
|
|
10
|
+
License-Expression: MIT
|
|
11
|
+
License-File: LICENSE
|
|
12
|
+
Keywords: apache-tika,document-parsing,docx,fastapi,formula-extraction,latex,layout-analysis,llm,markdown,ocr,pdf,pdf-extraction,pdf-parser,pdf-tables,pdf-to-json,pdf-to-markdown,pdf-to-text,pdfium,pptx,rag,reading-order,table-extraction,tika
|
|
13
|
+
Classifier: Development Status :: 4 - Beta
|
|
14
|
+
Classifier: Framework :: FastAPI
|
|
15
|
+
Classifier: Intended Audience :: Developers
|
|
16
|
+
Classifier: Natural Language :: Portuguese (Brazilian)
|
|
17
|
+
Classifier: Operating System :: OS Independent
|
|
18
|
+
Classifier: Programming Language :: Python :: 3
|
|
19
|
+
Classifier: Programming Language :: Python :: 3.10
|
|
20
|
+
Classifier: Programming Language :: Python :: 3.11
|
|
21
|
+
Classifier: Programming Language :: Python :: 3.12
|
|
22
|
+
Classifier: Programming Language :: Python :: 3.13
|
|
23
|
+
Classifier: Topic :: Scientific/Engineering :: Information Analysis
|
|
24
|
+
Classifier: Topic :: Text Processing
|
|
25
|
+
Classifier: Topic :: Text Processing :: Markup :: Markdown
|
|
26
|
+
Classifier: Topic :: Utilities
|
|
27
|
+
Requires-Python: >=3.10
|
|
28
|
+
Requires-Dist: pillow>=10
|
|
29
|
+
Requires-Dist: pypdfium2>=4.30
|
|
30
|
+
Requires-Dist: requests>=2.31
|
|
31
|
+
Requires-Dist: tika>=2.6.0
|
|
32
|
+
Provides-Extra: api
|
|
33
|
+
Requires-Dist: fastapi>=0.110; extra == 'api'
|
|
34
|
+
Requires-Dist: python-multipart>=0.0.9; extra == 'api'
|
|
35
|
+
Requires-Dist: uvicorn[standard]>=0.29; extra == 'api'
|
|
36
|
+
Provides-Extra: dev
|
|
37
|
+
Requires-Dist: fpdf2>=2.7; extra == 'dev'
|
|
38
|
+
Requires-Dist: httpx>=0.27; extra == 'dev'
|
|
39
|
+
Requires-Dist: pdf-text-api[api]; extra == 'dev'
|
|
40
|
+
Requires-Dist: pytest>=8; extra == 'dev'
|
|
41
|
+
Requires-Dist: ruff>=0.5; extra == 'dev'
|
|
42
|
+
Description-Content-Type: text/markdown
|
|
43
|
+
|
|
44
|
+
<p align="center">
|
|
45
|
+
<picture>
|
|
46
|
+
<source media="(prefers-color-scheme: dark)" srcset="web/assets/logo-dark.png">
|
|
47
|
+
<img src="web/assets/logo-light.png" alt="papero" width="420">
|
|
48
|
+
</picture>
|
|
49
|
+
</p>
|
|
50
|
+
|
|
51
|
+
<h3 align="center">Document structure extraction without the heavyweight stack.</h3>
|
|
52
|
+
|
|
53
|
+
<p align="center">
|
|
54
|
+
PDF → Markdown · JSON · Word · Excel — with reading order, tables, formulas, figures and the position of every block.<br>
|
|
55
|
+
CPU only. No ML models. Runs in your browser, in Python, or as an API.
|
|
56
|
+
</p>
|
|
57
|
+
|
|
58
|
+
<p align="center">
|
|
59
|
+
<a href="https://github.com/beatrizalmeidaf/papero-pdf-text-extractor/actions/workflows/ci.yml"><img src="https://github.com/beatrizalmeidaf/papero-pdf-text-extractor/actions/workflows/ci.yml/badge.svg" alt="CI"></a>
|
|
60
|
+
<img src="web/assets/badges/python.svg" alt="Python 3.10–3.13">
|
|
61
|
+
<a href="LICENSE"><img src="web/assets/badges/license.svg" alt="MIT license"></a>
|
|
62
|
+
<img src="web/assets/badges/ml-models.svg" alt="No ML models">
|
|
63
|
+
</p>
|
|
64
|
+
|
|
65
|
+
<p align="center">
|
|
66
|
+
<a href="https://beatrizalmeidaf.github.io/papero-pdf-text-extractor/"><b>▶ Try it in your browser</b></a>
|
|
67
|
+
·
|
|
68
|
+
<a href="#quick-start"><b>Quick start</b></a>
|
|
69
|
+
·
|
|
70
|
+
<a href="#benchmarks"><b>Benchmarks</b></a>
|
|
71
|
+
</p>
|
|
72
|
+
|
|
73
|
+
<p align="center">
|
|
74
|
+
<a href="https://beatrizalmeidaf.github.io/papero-pdf-text-extractor/">
|
|
75
|
+
<img src="web/assets/demo.gif" alt="papero demo: load a PDF, every block outlined on the page, click a table to inspect it, see tables and LaTeX formulas, export to Word" width="900">
|
|
76
|
+
</a>
|
|
77
|
+
<br>
|
|
78
|
+
<sub>30 seconds in the browser app: load a PDF, inspect any block, check tables and formulas, export to Word. Your PDF never leaves your machine. (<a href="web/assets/demo.mp4">MP4</a>)</sub>
|
|
79
|
+
</p>
|
|
80
|
+
|
|
81
|
+
---
|
|
82
|
+
|
|
83
|
+
## Why papero
|
|
84
|
+
|
|
85
|
+
Getting the *text* out of a PDF is easy. Getting its **structure** back — which column comes first, which lines are a table, where the formula is — is what makes the output usable for RAG, search and LLMs. papero does that with plain geometry, so it stays fast on a laptop CPU.
|
|
86
|
+
|
|
87
|
+
<table>
|
|
88
|
+
<tr>
|
|
89
|
+
<td width="33%" valign="top">
|
|
90
|
+
<b>📖 Reading order</b><br>
|
|
91
|
+
Two- and three-column papers read column by column. Headers, footers, page numbers and repeated logos are set aside.
|
|
92
|
+
</td>
|
|
93
|
+
<td width="33%" valign="top">
|
|
94
|
+
<b>▦ Real tables</b><br>
|
|
95
|
+
Ruled, borderless and LaTeX <i>booktabs</i> tables come back as rows and columns — multi-line cells included. Export to CSV or Excel.
|
|
96
|
+
</td>
|
|
97
|
+
<td width="33%" valign="top">
|
|
98
|
+
<b>∑ Formulas</b><br>
|
|
99
|
+
Superscripts, subscripts and math symbols become LaTeX (<code>E = mc^{2}</code>), plus a cropped image of the formula.
|
|
100
|
+
</td>
|
|
101
|
+
</tr>
|
|
102
|
+
<tr>
|
|
103
|
+
<td valign="top">
|
|
104
|
+
<b>📍 Position of everything</b><br>
|
|
105
|
+
Every block has a bounding box — cite the exact spot in a RAG answer, draw over the page, or crop it.
|
|
106
|
+
</td>
|
|
107
|
+
<td valign="top">
|
|
108
|
+
<b>🖼 Figures & charts</b><br>
|
|
109
|
+
Images and vector charts are cropped to PNG, with their caption, axis labels and legend kept together.
|
|
110
|
+
</td>
|
|
111
|
+
<td valign="top">
|
|
112
|
+
<b>📝 Back to Word</b><br>
|
|
113
|
+
Alignment, indents, line spacing, bold runs and fonts are kept, so a <code>.docx</code> export looks like the original page.
|
|
114
|
+
</td>
|
|
115
|
+
</tr>
|
|
116
|
+
</table>
|
|
117
|
+
|
|
118
|
+
Also: accents drawn as separate glyphs in LaTeX PDFs (`Computa¸ca˜o` → `Computação`), invisible white text used by form generators is dropped, scanned pages go through OCR, and DOCX/PPTX/XLSX/EPUB/HTML are read through Apache Tika.
|
|
119
|
+
|
|
120
|
+
## Quick start
|
|
121
|
+
|
|
122
|
+
```bash
|
|
123
|
+
pip install pdf-text-api
|
|
124
|
+
```
|
|
125
|
+
|
|
126
|
+
```python
|
|
127
|
+
from pdf_text_api import extract
|
|
128
|
+
|
|
129
|
+
doc = extract("paper.pdf")
|
|
130
|
+
print(doc.to_markdown())
|
|
131
|
+
```
|
|
132
|
+
|
|
133
|
+
Or skip the install: **[open the browser app](https://beatrizalmeidaf.github.io/papero-pdf-text-extractor/)**, drop a PDF, export to the format you need.
|
|
134
|
+
|
|
135
|
+
<details>
|
|
136
|
+
<summary><b>More Python</b> — tables, formulas, positions, images, options</summary>
|
|
137
|
+
|
|
138
|
+
```python
|
|
139
|
+
from pdf_text_api import extract, extract_text
|
|
140
|
+
|
|
141
|
+
doc = extract("paper.pdf", images=True)
|
|
142
|
+
|
|
143
|
+
doc.tables[0].rows # [["Model", "Accuracy"], ["Base", "0.81"], ...]
|
|
144
|
+
doc.formulas[0].latex # "E = mc^{2}"
|
|
145
|
+
doc.figures[0].image.data # PNG bytes
|
|
146
|
+
|
|
147
|
+
for block in doc.pages[0].blocks: # reading order, with positions
|
|
148
|
+
print(block.type, block.bbox, block.text[:60])
|
|
149
|
+
|
|
150
|
+
doc.to_html() # keeps alignment and indents
|
|
151
|
+
doc.to_dict() # the full JSON
|
|
152
|
+
|
|
153
|
+
extract("slides.pptx").to_markdown() # any format Apache Tika reads
|
|
154
|
+
extract_text("contract.pdf").text # fastest: clean text only
|
|
155
|
+
```
|
|
156
|
+
|
|
157
|
+
| Option | Default | |
|
|
158
|
+
|---|---|---|
|
|
159
|
+
| `pages` | all | `"1-3,5,10-"` |
|
|
160
|
+
| `images` | `False` | crop figures, tables and formulas to PNG |
|
|
161
|
+
| `tables` / `formulas` | `True` | detection on/off |
|
|
162
|
+
| `ocr` | `"auto"` | `"auto"` (scanned pages only), `"force"`, `"off"` |
|
|
163
|
+
| `ocr_language` | `"por+eng"` | Tesseract languages |
|
|
164
|
+
| `tika` | `True` | `False` runs the layout engine alone (no Java) |
|
|
165
|
+
| `workers` | `1` | processes for long documents |
|
|
166
|
+
|
|
167
|
+
</details>
|
|
168
|
+
|
|
169
|
+
<details>
|
|
170
|
+
<summary><b>CLI</b></summary>
|
|
171
|
+
|
|
172
|
+
```bash
|
|
173
|
+
pdf-text-api extract paper.pdf -o paper.md --images # Markdown + images/ folder
|
|
174
|
+
pdf-text-api extract paper.pdf -o paper.json # format from the extension
|
|
175
|
+
pdf-text-api extract paper.pdf -f csv -o tables.csv # tables only
|
|
176
|
+
pdf-text-api extract paper.pdf -p 1-5 -f html
|
|
177
|
+
pdf-text-api extract paper.pdf --fast # clean text only
|
|
178
|
+
pdf-text-api serve --port 8000 # API + browser app
|
|
179
|
+
```
|
|
180
|
+
|
|
181
|
+
</details>
|
|
182
|
+
|
|
183
|
+
<details>
|
|
184
|
+
<summary><b>REST API & Docker</b></summary>
|
|
185
|
+
|
|
186
|
+
```bash
|
|
187
|
+
docker compose up # API + Apache Tika + Tesseract + browser app on :8000
|
|
188
|
+
```
|
|
189
|
+
|
|
190
|
+
```bash
|
|
191
|
+
curl -F "file=@paper.pdf" "localhost:8000/v1/extract?format=markdown"
|
|
192
|
+
curl -F "file=@paper.pdf" "localhost:8000/v1/extract?format=zip&images=true" -o paper.zip
|
|
193
|
+
curl -F "file=@paper.pdf" "localhost:8000/v1/extract?per_page=true" # blocks + positions
|
|
194
|
+
```
|
|
195
|
+
|
|
196
|
+
One endpoint, `POST /v1/extract`; interactive docs at `/docs`.
|
|
197
|
+
|
|
198
|
+
| Parameter | Default | |
|
|
199
|
+
|---|---|---|
|
|
200
|
+
| `mode` | `structured` | `structured` (layout + Tika) or `fast` (text only) |
|
|
201
|
+
| `format` | `json` | `json`, `markdown`, `text`, `html`, `csv`, `zip` |
|
|
202
|
+
| `pages` | all | `1-3,5,10-` |
|
|
203
|
+
| `per_page` | `false` | include pages, blocks and positions in the JSON |
|
|
204
|
+
| `images` | `false` | crop figures, tables and formulas |
|
|
205
|
+
| `ocr` | `auto` | `auto`, `force`, `off` |
|
|
206
|
+
|
|
207
|
+
Configuration through environment variables — see [`.env.example`](.env.example).
|
|
208
|
+
|
|
209
|
+
</details>
|
|
210
|
+
|
|
211
|
+
## What comes out
|
|
212
|
+
|
|
213
|
+
Every block knows what it is and where it was:
|
|
214
|
+
|
|
215
|
+
```json
|
|
216
|
+
{
|
|
217
|
+
"type": "table",
|
|
218
|
+
"bbox": [56.7, 294.8, 481.9, 374.2],
|
|
219
|
+
"rows": [["Model", "Accuracy"], ["Base", "0.81"]],
|
|
220
|
+
"caption": "Table 1: Comparison between models."
|
|
221
|
+
}
|
|
222
|
+
```
|
|
223
|
+
|
|
224
|
+
| Output | Python · CLI · API | Browser app |
|
|
225
|
+
|---|:-:|:-:|
|
|
226
|
+
| Markdown, plain text, JSON | ✓ | ✓ |
|
|
227
|
+
| HTML (keeps alignment and indents) | ✓ | ✓ |
|
|
228
|
+
| CSV of the tables, ZIP with images | ✓ | ✓ |
|
|
229
|
+
| Word `.docx` that keeps the page's look | — | ✓ |
|
|
230
|
+
| Excel `.xlsx`, one sheet per table | — | ✓ |
|
|
231
|
+
|
|
232
|
+
<details>
|
|
233
|
+
<summary><b>Full JSON schema and block types</b></summary>
|
|
234
|
+
|
|
235
|
+
```json
|
|
236
|
+
{
|
|
237
|
+
"schema": "pdf-text-api/document@1",
|
|
238
|
+
"engine": "tika+pdfium",
|
|
239
|
+
"page_count": 12,
|
|
240
|
+
"metadata": { "title": "...", "author": "...", "language": "en" },
|
|
241
|
+
"pages": [{
|
|
242
|
+
"number": 1, "width": 595.3, "height": 841.9,
|
|
243
|
+
"blocks": [{
|
|
244
|
+
"id": "p1-b4", "type": "paragraph", "bbox": [74.0, 217.0, 522.0, 275.0],
|
|
245
|
+
"text": "Atestamos que a estudante ...",
|
|
246
|
+
"style": { "pt": 11.0, "font": "Arial", "bold": false },
|
|
247
|
+
"format": { "align": "justify", "first_line": 42.7, "line_spacing": 1.8 },
|
|
248
|
+
"runs": [{ "text": "FULANA DE TAL", "bold": true, "italic": false, "script": null }]
|
|
249
|
+
}]
|
|
250
|
+
}]
|
|
251
|
+
}
|
|
252
|
+
```
|
|
253
|
+
|
|
254
|
+
Block types: `heading` (with `level`), `paragraph`, `list_item` (with `marker`), `table` (with `rows`), `figure`, `formula` (with `latex`), `caption`, `code`, and — kept apart from the text — `header`, `footer`, `page_number`. Bounding boxes are `[x0, y0, x1, y1]` in points, origin at the top-left of the page.
|
|
255
|
+
|
|
256
|
+
</details>
|
|
257
|
+
|
|
258
|
+
## Benchmarks
|
|
259
|
+
|
|
260
|
+
<p align="center">
|
|
261
|
+
<img src="benchmarks/latency.svg" alt="Average extraction time per PDF on a log scale: PyMuPDF 97 ms, papero fast 136 ms, papero structured 543 ms, pypdf 1.52 s, pdfplumber 3.56 s, Docling 82.9 s" width="760">
|
|
262
|
+
</p>
|
|
263
|
+
|
|
264
|
+
Dense arXiv papers (multi-column, formulas, tables, figures) on one laptop CPU, no GPU. `papero · fast` returns clean text; `papero · structured` also rebuilds reading order, tables, formulas and figures — **0 failures on 54 papers, 39 ms per page (median)**. Reproduce with [`benchmarks/`](benchmarks/).
|
|
265
|
+
|
|
266
|
+
| | papero | PyMuPDF | pdfplumber | pypdf | Docling | Marker |
|
|
267
|
+
|---|:-:|:-:|:-:|:-:|:-:|:-:|
|
|
268
|
+
| License | MIT | AGPL | MIT | BSD | MIT | GPL |
|
|
269
|
+
| Needs ML models / PyTorch | no | no | no | no | yes | yes |
|
|
270
|
+
| Multi-column reading order | ✓ | partial | — | — | ✓ | ✓ |
|
|
271
|
+
| Structured tables | ✓ | ✓ | ✓ | — | ✓ | ✓ |
|
|
272
|
+
| Formulas | LaTeX from glyphs + image | — | — | — | ✓ | ✓ |
|
|
273
|
+
| Bounding boxes | ✓ | ✓ | ✓ | — | ✓ | ✓ |
|
|
274
|
+
| DOCX / PPTX / XLSX / EPUB | ✓ | partial | — | — | ✓ | partial |
|
|
275
|
+
| Runs entirely in the browser | ✓ | — | — | — | — | — |
|
|
276
|
+
|
|
277
|
+
ML-based tools still win on very irregular layouts and complex math (stacked fractions, matrices) — papero gives you the formula as approximate LaTeX **and** as an image so nothing is lost.
|
|
278
|
+
|
|
279
|
+
## How it works
|
|
280
|
+
|
|
281
|
+
Two engines run on the same file **at the same time**:
|
|
282
|
+
|
|
283
|
+
- **A layout engine on PDFium** reads every glyph with its position, font and size, plus every rule and image, and rebuilds columns, tables, formulas, lists and figures with a column-aware XY-cut.
|
|
284
|
+
- **Apache Tika** adds metadata, tagged-PDF headings, OCR (Tesseract) and every non-PDF format.
|
|
285
|
+
|
|
286
|
+
The browser app runs the same algorithm ported to JavaScript on pdf.js, and CI checks block by block that both engines agree.
|
|
287
|
+
|
|
288
|
+
<details>
|
|
289
|
+
<summary><b>Limitations</b></summary>
|
|
290
|
+
|
|
291
|
+
- **Math:** LaTeX is rebuilt from glyphs — stacked fractions, matrices and big radicals come out linear (the cropped image is always there).
|
|
292
|
+
- **Borderless tables** with very narrow gaps between columns can read as text.
|
|
293
|
+
- **Scanned PDFs** need OCR, which runs on the server path (Tesseract is in the Docker image).
|
|
294
|
+
- **Word/Excel export** is in the browser app for now.
|
|
295
|
+
|
|
296
|
+
</details>
|
|
297
|
+
|
|
298
|
+
<details>
|
|
299
|
+
<summary><b>Development</b></summary>
|
|
300
|
+
|
|
301
|
+
```bash
|
|
302
|
+
git clone https://github.com/beatrizalmeidaf/papero-pdf-text-extractor.git && cd pdf-text-extractor
|
|
303
|
+
pip install -e ".[dev]"
|
|
304
|
+
pytest -q # includes real-world regressions
|
|
305
|
+
ruff check src tests && ruff format --check src tests
|
|
306
|
+
npm install --prefix tests/js && python tests/js/expected.py tests/js/out && node tests/js/parity.mjs tests/js/out
|
|
307
|
+
python -m http.server -d web # browser app at http://localhost:8000
|
|
308
|
+
```
|
|
309
|
+
|
|
310
|
+
`src/pdf_text_api/` is the Python engine, API and CLI · `web/` is the browser app (GitHub Pages) · `tests/js/` checks the two engines agree · `benchmarks/` downloads the dataset and draws the chart.
|
|
311
|
+
|
|
312
|
+
</details>
|
|
313
|
+
|
|
314
|
+
## Contributing
|
|
315
|
+
|
|
316
|
+
Found a PDF papero gets wrong? **That's the most useful issue you can open** — attach the file (or a page of it) and say what you expected. Reading order, tables, formulas, encoding, OCR and browser/server differences are all fair game.
|
|
317
|
+
|
|
318
|
+
If papero saves you time, **a ⭐ helps other people find it.**
|
|
319
|
+
|
|
320
|
+
<sub>**Keywords:** PDF to Markdown · PDF to JSON · PDF to Word · PDF to Excel · PDF table extraction · PDF parser · document parsing · layout analysis · reading order · multi-column PDF · formula extraction · LaTeX · bounding boxes · OCR · Apache Tika · PDFium · pdf.js · RAG preprocessing · LLM document loader · Docling alternative · PyMuPDF alternative · converter PDF para Markdown, Word e Excel · extrair tabelas de PDF · extrair texto de PDF mantendo a formatação · OCR de PDF escaneado</sub>
|
|
321
|
+
|
|
322
|
+
<sub>MIT © Beatriz Almeida · package and imports keep the name `pdf-text-api` / `pdf_text_api` for compatibility.</sub>
|
|
@@ -0,0 +1,279 @@
|
|
|
1
|
+
<p align="center">
|
|
2
|
+
<picture>
|
|
3
|
+
<source media="(prefers-color-scheme: dark)" srcset="web/assets/logo-dark.png">
|
|
4
|
+
<img src="web/assets/logo-light.png" alt="papero" width="420">
|
|
5
|
+
</picture>
|
|
6
|
+
</p>
|
|
7
|
+
|
|
8
|
+
<h3 align="center">Document structure extraction without the heavyweight stack.</h3>
|
|
9
|
+
|
|
10
|
+
<p align="center">
|
|
11
|
+
PDF → Markdown · JSON · Word · Excel — with reading order, tables, formulas, figures and the position of every block.<br>
|
|
12
|
+
CPU only. No ML models. Runs in your browser, in Python, or as an API.
|
|
13
|
+
</p>
|
|
14
|
+
|
|
15
|
+
<p align="center">
|
|
16
|
+
<a href="https://github.com/beatrizalmeidaf/papero-pdf-text-extractor/actions/workflows/ci.yml"><img src="https://github.com/beatrizalmeidaf/papero-pdf-text-extractor/actions/workflows/ci.yml/badge.svg" alt="CI"></a>
|
|
17
|
+
<img src="web/assets/badges/python.svg" alt="Python 3.10–3.13">
|
|
18
|
+
<a href="LICENSE"><img src="web/assets/badges/license.svg" alt="MIT license"></a>
|
|
19
|
+
<img src="web/assets/badges/ml-models.svg" alt="No ML models">
|
|
20
|
+
</p>
|
|
21
|
+
|
|
22
|
+
<p align="center">
|
|
23
|
+
<a href="https://beatrizalmeidaf.github.io/papero-pdf-text-extractor/"><b>▶ Try it in your browser</b></a>
|
|
24
|
+
·
|
|
25
|
+
<a href="#quick-start"><b>Quick start</b></a>
|
|
26
|
+
·
|
|
27
|
+
<a href="#benchmarks"><b>Benchmarks</b></a>
|
|
28
|
+
</p>
|
|
29
|
+
|
|
30
|
+
<p align="center">
|
|
31
|
+
<a href="https://beatrizalmeidaf.github.io/papero-pdf-text-extractor/">
|
|
32
|
+
<img src="web/assets/demo.gif" alt="papero demo: load a PDF, every block outlined on the page, click a table to inspect it, see tables and LaTeX formulas, export to Word" width="900">
|
|
33
|
+
</a>
|
|
34
|
+
<br>
|
|
35
|
+
<sub>30 seconds in the browser app: load a PDF, inspect any block, check tables and formulas, export to Word. Your PDF never leaves your machine. (<a href="web/assets/demo.mp4">MP4</a>)</sub>
|
|
36
|
+
</p>
|
|
37
|
+
|
|
38
|
+
---
|
|
39
|
+
|
|
40
|
+
## Why papero
|
|
41
|
+
|
|
42
|
+
Getting the *text* out of a PDF is easy. Getting its **structure** back — which column comes first, which lines are a table, where the formula is — is what makes the output usable for RAG, search and LLMs. papero does that with plain geometry, so it stays fast on a laptop CPU.
|
|
43
|
+
|
|
44
|
+
<table>
|
|
45
|
+
<tr>
|
|
46
|
+
<td width="33%" valign="top">
|
|
47
|
+
<b>📖 Reading order</b><br>
|
|
48
|
+
Two- and three-column papers read column by column. Headers, footers, page numbers and repeated logos are set aside.
|
|
49
|
+
</td>
|
|
50
|
+
<td width="33%" valign="top">
|
|
51
|
+
<b>▦ Real tables</b><br>
|
|
52
|
+
Ruled, borderless and LaTeX <i>booktabs</i> tables come back as rows and columns — multi-line cells included. Export to CSV or Excel.
|
|
53
|
+
</td>
|
|
54
|
+
<td width="33%" valign="top">
|
|
55
|
+
<b>∑ Formulas</b><br>
|
|
56
|
+
Superscripts, subscripts and math symbols become LaTeX (<code>E = mc^{2}</code>), plus a cropped image of the formula.
|
|
57
|
+
</td>
|
|
58
|
+
</tr>
|
|
59
|
+
<tr>
|
|
60
|
+
<td valign="top">
|
|
61
|
+
<b>📍 Position of everything</b><br>
|
|
62
|
+
Every block has a bounding box — cite the exact spot in a RAG answer, draw over the page, or crop it.
|
|
63
|
+
</td>
|
|
64
|
+
<td valign="top">
|
|
65
|
+
<b>🖼 Figures & charts</b><br>
|
|
66
|
+
Images and vector charts are cropped to PNG, with their caption, axis labels and legend kept together.
|
|
67
|
+
</td>
|
|
68
|
+
<td valign="top">
|
|
69
|
+
<b>📝 Back to Word</b><br>
|
|
70
|
+
Alignment, indents, line spacing, bold runs and fonts are kept, so a <code>.docx</code> export looks like the original page.
|
|
71
|
+
</td>
|
|
72
|
+
</tr>
|
|
73
|
+
</table>
|
|
74
|
+
|
|
75
|
+
Also: accents drawn as separate glyphs in LaTeX PDFs (`Computa¸ca˜o` → `Computação`), invisible white text used by form generators is dropped, scanned pages go through OCR, and DOCX/PPTX/XLSX/EPUB/HTML are read through Apache Tika.
|
|
76
|
+
|
|
77
|
+
## Quick start
|
|
78
|
+
|
|
79
|
+
```bash
|
|
80
|
+
pip install pdf-text-api
|
|
81
|
+
```
|
|
82
|
+
|
|
83
|
+
```python
|
|
84
|
+
from pdf_text_api import extract
|
|
85
|
+
|
|
86
|
+
doc = extract("paper.pdf")
|
|
87
|
+
print(doc.to_markdown())
|
|
88
|
+
```
|
|
89
|
+
|
|
90
|
+
Or skip the install: **[open the browser app](https://beatrizalmeidaf.github.io/papero-pdf-text-extractor/)**, drop a PDF, export to the format you need.
|
|
91
|
+
|
|
92
|
+
<details>
|
|
93
|
+
<summary><b>More Python</b> — tables, formulas, positions, images, options</summary>
|
|
94
|
+
|
|
95
|
+
```python
|
|
96
|
+
from pdf_text_api import extract, extract_text
|
|
97
|
+
|
|
98
|
+
doc = extract("paper.pdf", images=True)
|
|
99
|
+
|
|
100
|
+
doc.tables[0].rows # [["Model", "Accuracy"], ["Base", "0.81"], ...]
|
|
101
|
+
doc.formulas[0].latex # "E = mc^{2}"
|
|
102
|
+
doc.figures[0].image.data # PNG bytes
|
|
103
|
+
|
|
104
|
+
for block in doc.pages[0].blocks: # reading order, with positions
|
|
105
|
+
print(block.type, block.bbox, block.text[:60])
|
|
106
|
+
|
|
107
|
+
doc.to_html() # keeps alignment and indents
|
|
108
|
+
doc.to_dict() # the full JSON
|
|
109
|
+
|
|
110
|
+
extract("slides.pptx").to_markdown() # any format Apache Tika reads
|
|
111
|
+
extract_text("contract.pdf").text # fastest: clean text only
|
|
112
|
+
```
|
|
113
|
+
|
|
114
|
+
| Option | Default | |
|
|
115
|
+
|---|---|---|
|
|
116
|
+
| `pages` | all | `"1-3,5,10-"` |
|
|
117
|
+
| `images` | `False` | crop figures, tables and formulas to PNG |
|
|
118
|
+
| `tables` / `formulas` | `True` | detection on/off |
|
|
119
|
+
| `ocr` | `"auto"` | `"auto"` (scanned pages only), `"force"`, `"off"` |
|
|
120
|
+
| `ocr_language` | `"por+eng"` | Tesseract languages |
|
|
121
|
+
| `tika` | `True` | `False` runs the layout engine alone (no Java) |
|
|
122
|
+
| `workers` | `1` | processes for long documents |
|
|
123
|
+
|
|
124
|
+
</details>
|
|
125
|
+
|
|
126
|
+
<details>
|
|
127
|
+
<summary><b>CLI</b></summary>
|
|
128
|
+
|
|
129
|
+
```bash
|
|
130
|
+
pdf-text-api extract paper.pdf -o paper.md --images # Markdown + images/ folder
|
|
131
|
+
pdf-text-api extract paper.pdf -o paper.json # format from the extension
|
|
132
|
+
pdf-text-api extract paper.pdf -f csv -o tables.csv # tables only
|
|
133
|
+
pdf-text-api extract paper.pdf -p 1-5 -f html
|
|
134
|
+
pdf-text-api extract paper.pdf --fast # clean text only
|
|
135
|
+
pdf-text-api serve --port 8000 # API + browser app
|
|
136
|
+
```
|
|
137
|
+
|
|
138
|
+
</details>
|
|
139
|
+
|
|
140
|
+
<details>
|
|
141
|
+
<summary><b>REST API & Docker</b></summary>
|
|
142
|
+
|
|
143
|
+
```bash
|
|
144
|
+
docker compose up # API + Apache Tika + Tesseract + browser app on :8000
|
|
145
|
+
```
|
|
146
|
+
|
|
147
|
+
```bash
|
|
148
|
+
curl -F "file=@paper.pdf" "localhost:8000/v1/extract?format=markdown"
|
|
149
|
+
curl -F "file=@paper.pdf" "localhost:8000/v1/extract?format=zip&images=true" -o paper.zip
|
|
150
|
+
curl -F "file=@paper.pdf" "localhost:8000/v1/extract?per_page=true" # blocks + positions
|
|
151
|
+
```
|
|
152
|
+
|
|
153
|
+
One endpoint, `POST /v1/extract`; interactive docs at `/docs`.
|
|
154
|
+
|
|
155
|
+
| Parameter | Default | |
|
|
156
|
+
|---|---|---|
|
|
157
|
+
| `mode` | `structured` | `structured` (layout + Tika) or `fast` (text only) |
|
|
158
|
+
| `format` | `json` | `json`, `markdown`, `text`, `html`, `csv`, `zip` |
|
|
159
|
+
| `pages` | all | `1-3,5,10-` |
|
|
160
|
+
| `per_page` | `false` | include pages, blocks and positions in the JSON |
|
|
161
|
+
| `images` | `false` | crop figures, tables and formulas |
|
|
162
|
+
| `ocr` | `auto` | `auto`, `force`, `off` |
|
|
163
|
+
|
|
164
|
+
Configuration through environment variables — see [`.env.example`](.env.example).
|
|
165
|
+
|
|
166
|
+
</details>
|
|
167
|
+
|
|
168
|
+
## What comes out
|
|
169
|
+
|
|
170
|
+
Every block knows what it is and where it was:
|
|
171
|
+
|
|
172
|
+
```json
|
|
173
|
+
{
|
|
174
|
+
"type": "table",
|
|
175
|
+
"bbox": [56.7, 294.8, 481.9, 374.2],
|
|
176
|
+
"rows": [["Model", "Accuracy"], ["Base", "0.81"]],
|
|
177
|
+
"caption": "Table 1: Comparison between models."
|
|
178
|
+
}
|
|
179
|
+
```
|
|
180
|
+
|
|
181
|
+
| Output | Python · CLI · API | Browser app |
|
|
182
|
+
|---|:-:|:-:|
|
|
183
|
+
| Markdown, plain text, JSON | ✓ | ✓ |
|
|
184
|
+
| HTML (keeps alignment and indents) | ✓ | ✓ |
|
|
185
|
+
| CSV of the tables, ZIP with images | ✓ | ✓ |
|
|
186
|
+
| Word `.docx` that keeps the page's look | — | ✓ |
|
|
187
|
+
| Excel `.xlsx`, one sheet per table | — | ✓ |
|
|
188
|
+
|
|
189
|
+
<details>
|
|
190
|
+
<summary><b>Full JSON schema and block types</b></summary>
|
|
191
|
+
|
|
192
|
+
```json
|
|
193
|
+
{
|
|
194
|
+
"schema": "pdf-text-api/document@1",
|
|
195
|
+
"engine": "tika+pdfium",
|
|
196
|
+
"page_count": 12,
|
|
197
|
+
"metadata": { "title": "...", "author": "...", "language": "en" },
|
|
198
|
+
"pages": [{
|
|
199
|
+
"number": 1, "width": 595.3, "height": 841.9,
|
|
200
|
+
"blocks": [{
|
|
201
|
+
"id": "p1-b4", "type": "paragraph", "bbox": [74.0, 217.0, 522.0, 275.0],
|
|
202
|
+
"text": "Atestamos que a estudante ...",
|
|
203
|
+
"style": { "pt": 11.0, "font": "Arial", "bold": false },
|
|
204
|
+
"format": { "align": "justify", "first_line": 42.7, "line_spacing": 1.8 },
|
|
205
|
+
"runs": [{ "text": "FULANA DE TAL", "bold": true, "italic": false, "script": null }]
|
|
206
|
+
}]
|
|
207
|
+
}]
|
|
208
|
+
}
|
|
209
|
+
```
|
|
210
|
+
|
|
211
|
+
Block types: `heading` (with `level`), `paragraph`, `list_item` (with `marker`), `table` (with `rows`), `figure`, `formula` (with `latex`), `caption`, `code`, and — kept apart from the text — `header`, `footer`, `page_number`. Bounding boxes are `[x0, y0, x1, y1]` in points, origin at the top-left of the page.
|
|
212
|
+
|
|
213
|
+
</details>
|
|
214
|
+
|
|
215
|
+
## Benchmarks
|
|
216
|
+
|
|
217
|
+
<p align="center">
|
|
218
|
+
<img src="benchmarks/latency.svg" alt="Average extraction time per PDF on a log scale: PyMuPDF 97 ms, papero fast 136 ms, papero structured 543 ms, pypdf 1.52 s, pdfplumber 3.56 s, Docling 82.9 s" width="760">
|
|
219
|
+
</p>
|
|
220
|
+
|
|
221
|
+
Dense arXiv papers (multi-column, formulas, tables, figures) on one laptop CPU, no GPU. `papero · fast` returns clean text; `papero · structured` also rebuilds reading order, tables, formulas and figures — **0 failures on 54 papers, 39 ms per page (median)**. Reproduce with [`benchmarks/`](benchmarks/).
|
|
222
|
+
|
|
223
|
+
| | papero | PyMuPDF | pdfplumber | pypdf | Docling | Marker |
|
|
224
|
+
|---|:-:|:-:|:-:|:-:|:-:|:-:|
|
|
225
|
+
| License | MIT | AGPL | MIT | BSD | MIT | GPL |
|
|
226
|
+
| Needs ML models / PyTorch | no | no | no | no | yes | yes |
|
|
227
|
+
| Multi-column reading order | ✓ | partial | — | — | ✓ | ✓ |
|
|
228
|
+
| Structured tables | ✓ | ✓ | ✓ | — | ✓ | ✓ |
|
|
229
|
+
| Formulas | LaTeX from glyphs + image | — | — | — | ✓ | ✓ |
|
|
230
|
+
| Bounding boxes | ✓ | ✓ | ✓ | — | ✓ | ✓ |
|
|
231
|
+
| DOCX / PPTX / XLSX / EPUB | ✓ | partial | — | — | ✓ | partial |
|
|
232
|
+
| Runs entirely in the browser | ✓ | — | — | — | — | — |
|
|
233
|
+
|
|
234
|
+
ML-based tools still win on very irregular layouts and complex math (stacked fractions, matrices) — papero gives you the formula as approximate LaTeX **and** as an image so nothing is lost.
|
|
235
|
+
|
|
236
|
+
## How it works
|
|
237
|
+
|
|
238
|
+
Two engines run on the same file **at the same time**:
|
|
239
|
+
|
|
240
|
+
- **A layout engine on PDFium** reads every glyph with its position, font and size, plus every rule and image, and rebuilds columns, tables, formulas, lists and figures with a column-aware XY-cut.
|
|
241
|
+
- **Apache Tika** adds metadata, tagged-PDF headings, OCR (Tesseract) and every non-PDF format.
|
|
242
|
+
|
|
243
|
+
The browser app runs the same algorithm ported to JavaScript on pdf.js, and CI checks block by block that both engines agree.
|
|
244
|
+
|
|
245
|
+
<details>
|
|
246
|
+
<summary><b>Limitations</b></summary>
|
|
247
|
+
|
|
248
|
+
- **Math:** LaTeX is rebuilt from glyphs — stacked fractions, matrices and big radicals come out linear (the cropped image is always there).
|
|
249
|
+
- **Borderless tables** with very narrow gaps between columns can read as text.
|
|
250
|
+
- **Scanned PDFs** need OCR, which runs on the server path (Tesseract is in the Docker image).
|
|
251
|
+
- **Word/Excel export** is in the browser app for now.
|
|
252
|
+
|
|
253
|
+
</details>
|
|
254
|
+
|
|
255
|
+
<details>
|
|
256
|
+
<summary><b>Development</b></summary>
|
|
257
|
+
|
|
258
|
+
```bash
|
|
259
|
+
git clone https://github.com/beatrizalmeidaf/papero-pdf-text-extractor.git && cd pdf-text-extractor
|
|
260
|
+
pip install -e ".[dev]"
|
|
261
|
+
pytest -q # includes real-world regressions
|
|
262
|
+
ruff check src tests && ruff format --check src tests
|
|
263
|
+
npm install --prefix tests/js && python tests/js/expected.py tests/js/out && node tests/js/parity.mjs tests/js/out
|
|
264
|
+
python -m http.server -d web # browser app at http://localhost:8000
|
|
265
|
+
```
|
|
266
|
+
|
|
267
|
+
`src/pdf_text_api/` is the Python engine, API and CLI · `web/` is the browser app (GitHub Pages) · `tests/js/` checks the two engines agree · `benchmarks/` downloads the dataset and draws the chart.
|
|
268
|
+
|
|
269
|
+
</details>
|
|
270
|
+
|
|
271
|
+
## Contributing
|
|
272
|
+
|
|
273
|
+
Found a PDF papero gets wrong? **That's the most useful issue you can open** — attach the file (or a page of it) and say what you expected. Reading order, tables, formulas, encoding, OCR and browser/server differences are all fair game.
|
|
274
|
+
|
|
275
|
+
If papero saves you time, **a ⭐ helps other people find it.**
|
|
276
|
+
|
|
277
|
+
<sub>**Keywords:** PDF to Markdown · PDF to JSON · PDF to Word · PDF to Excel · PDF table extraction · PDF parser · document parsing · layout analysis · reading order · multi-column PDF · formula extraction · LaTeX · bounding boxes · OCR · Apache Tika · PDFium · pdf.js · RAG preprocessing · LLM document loader · Docling alternative · PyMuPDF alternative · converter PDF para Markdown, Word e Excel · extrair tabelas de PDF · extrair texto de PDF mantendo a formatação · OCR de PDF escaneado</sub>
|
|
278
|
+
|
|
279
|
+
<sub>MIT © Beatriz Almeida · package and imports keep the name `pdf-text-api` / `pdf_text_api` for compatibility.</sub>
|