universal-doc-parser 1.0.0__tar.gz → 1.0.2__tar.gz
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/.github/workflows/ci.yml +1 -1
- universal_doc_parser-1.0.2/.github/workflows/hf_sync.yml +49 -0
- {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/PKG-INFO +141 -98
- {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/README.md +131 -96
- universal_doc_parser-1.0.2/assets/README.md +40 -0
- universal_doc_parser-1.0.2/benchmarks/README.md +16 -0
- universal_doc_parser-1.0.2/benchmarks/providers/README.md +22 -0
- universal_doc_parser-1.0.2/docs/ADDING_A_FORMAT.md +212 -0
- universal_doc_parser-1.0.2/docs/ARCHITECTURE.md +164 -0
- universal_doc_parser-1.0.2/docs/CHANGELOG.md +234 -0
- universal_doc_parser-1.0.2/docs/README.md +76 -0
- universal_doc_parser-1.0.2/docs/SCHEMA.md +183 -0
- universal_doc_parser-1.0.2/main.py +199 -0
- {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/pyproject.toml +13 -2
- universal_doc_parser-1.0.2/requirements.txt +20 -0
- universal_doc_parser-1.0.2/tests/README.md +118 -0
- {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/universal_parser/__init__.py +1 -1
- universal_doc_parser-1.0.2/universal_parser/adaptive/README.md +35 -0
- {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/universal_parser/core/sniffer.py +0 -1
- universal_doc_parser-1.0.2/universal_parser/enrichment/README.md +19 -0
- universal_doc_parser-1.0.2/universal_parser/exports/README.md +34 -0
- universal_doc_parser-1.0.2/universal_parser/extractors/README.md +117 -0
- universal_doc_parser-1.0.2/universal_parser/extractors/images/README.md +144 -0
- universal_doc_parser-1.0.2/universal_parser/extractors/mail/README.md +12 -0
- universal_doc_parser-1.0.2/universal_parser/extractors/office/README.md +173 -0
- universal_doc_parser-1.0.2/universal_parser/extractors/pdf/README.md +180 -0
- universal_doc_parser-1.0.2/universal_parser/extractors/structured/README.md +13 -0
- universal_doc_parser-1.0.2/universal_parser/extractors/web/README.md +12 -0
- universal_doc_parser-1.0.2/universal_parser/mcp/README.md +26 -0
- universal_doc_parser-1.0.2/universal_parser/observability/README.md +18 -0
- {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/uv.lock +630 -7
- universal_doc_parser-1.0.0/CHANGELOG.md +0 -78
- universal_doc_parser-1.0.0/docs/ADDING_A_FORMAT.md +0 -0
- universal_doc_parser-1.0.0/docs/ARCHITECTURE.md +0 -0
- universal_doc_parser-1.0.0/docs/CHANGELOG.md +0 -0
- universal_doc_parser-1.0.0/docs/SCHEMA.md +0 -0
- universal_doc_parser-1.0.0/main.py +0 -6
- {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/.github/workflows/release.yml +0 -0
- {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/.gitignore +0 -0
- {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/.python-version +0 -0
- {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/LICENSE +0 -0
- {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/assets/banner.png +0 -0
- {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/assets/logo.png +0 -0
- {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/benchmarks/memory_profile.py +0 -0
- {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/benchmarks/providers/__init__.py +0 -0
- {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/benchmarks/providers/anthropic/__init__.py +0 -0
- {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/benchmarks/providers/anthropic/opus.py +0 -0
- {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/benchmarks/providers/anthropic/sonnet.py +0 -0
- {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/benchmarks/providers/base.py +0 -0
- {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/benchmarks/providers/cohere/__init__.py +0 -0
- {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/benchmarks/providers/cohere/command_r_plus.py +0 -0
- {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/benchmarks/providers/deepseek/__init__.py +0 -0
- {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/benchmarks/providers/deepseek/r1.py +0 -0
- {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/benchmarks/providers/deepseek/v3.py +0 -0
- {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/benchmarks/providers/gemini/__init__.py +0 -0
- {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/benchmarks/providers/gemini/flash.py +0 -0
- {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/benchmarks/providers/gemini/pro.py +0 -0
- {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/benchmarks/providers/glm/__init__.py +0 -0
- {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/benchmarks/providers/glm/glm4.py +0 -0
- {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/benchmarks/providers/kimi/__init__.py +0 -0
- {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/benchmarks/providers/kimi/k1_5.py +0 -0
- {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/benchmarks/providers/meta/__init__.py +0 -0
- {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/benchmarks/providers/meta/llama3_3.py +0 -0
- {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/benchmarks/providers/mistral/__init__.py +0 -0
- {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/benchmarks/providers/mistral/mistral_large.py +0 -0
- {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/benchmarks/providers/openai/__init__.py +0 -0
- {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/benchmarks/providers/openai/gpt4o.py +0 -0
- {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/benchmarks/providers/openai/gpt4o_mini.py +0 -0
- {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/benchmarks/providers/qwen/__init__.py +0 -0
- {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/benchmarks/providers/qwen/qwen_72b.py +0 -0
- {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/benchmarks/providers/qwen/qwen_coder.py +0 -0
- {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/benchmarks/run_llm_benchmark.py +0 -0
- {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/benchmarks/test_messy_document.py +0 -0
- {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/tests/__init__.py +0 -0
- {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/tests/fixtures/csv/sample.csv +0 -0
- {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/tests/fixtures/csv/sample.tsv +0 -0
- {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/tests/fixtures/docx/sample.docx +0 -0
- {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/tests/fixtures/html/sample.html +0 -0
- {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/tests/fixtures/json/sample_nested.json +0 -0
- {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/tests/fixtures/json/sample_tabular.json +0 -0
- {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/tests/fixtures/pdf/sample.pdf +0 -0
- {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/tests/fixtures/xlsx/sample.xlsx +0 -0
- {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/tests/fixtures/xml/sample.xml +0 -0
- {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/tests/test_adaptive.py +0 -0
- {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/tests/test_benchmark_providers.py +0 -0
- {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/tests/test_csv.py +0 -0
- {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/tests/test_docx.py +0 -0
- {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/tests/test_epub.py +0 -0
- {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/tests/test_exports.py +0 -0
- {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/tests/test_html.py +0 -0
- {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/tests/test_image.py +0 -0
- {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/tests/test_json_xml.py +0 -0
- {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/tests/test_legacy.py +0 -0
- {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/tests/test_mail.py +0 -0
- {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/tests/test_mcp.py +0 -0
- {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/tests/test_observability.py +0 -0
- {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/tests/test_parquet.py +0 -0
- {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/tests/test_pdf.py +0 -0
- {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/tests/test_pptx.py +0 -0
- {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/tests/test_xlsx.py +0 -0
- {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/universal_parser/adaptive/__init__.py +0 -0
- {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/universal_parser/adaptive/config_cache.py +0 -0
- {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/universal_parser/adaptive/fingerprint.py +0 -0
- {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/universal_parser/adaptive/tuner.py +0 -0
- {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/universal_parser/core/__init__.py +0 -0
- {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/universal_parser/core/engine.py +0 -0
- {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/universal_parser/core/router.py +0 -0
- {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/universal_parser/core/schema.py +0 -0
- {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/universal_parser/enrichment/__init__.py +0 -0
- {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/universal_parser/enrichment/vlm_enricher.py +0 -0
- {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/universal_parser/exports/__init__.py +0 -0
- {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/universal_parser/exports/to_chunks.py +0 -0
- {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/universal_parser/exports/to_graph.py +0 -0
- {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/universal_parser/exports/to_markdown.py +0 -0
- {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/universal_parser/extractors/__init__.py +0 -0
- {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/universal_parser/extractors/base.py +0 -0
- {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/universal_parser/extractors/images/__init__.py +0 -0
- {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/universal_parser/extractors/images/scan_extractor.py +0 -0
- {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/universal_parser/extractors/mail/__init__.py +0 -0
- {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/universal_parser/extractors/mail/mail_extractor.py +0 -0
- {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/universal_parser/extractors/office/__init__.py +0 -0
- {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/universal_parser/extractors/office/docx_extractor.py +0 -0
- {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/universal_parser/extractors/office/legacy_extractor.py +0 -0
- {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/universal_parser/extractors/office/pptx_extractor.py +0 -0
- {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/universal_parser/extractors/office/xlsx_extractor.py +0 -0
- {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/universal_parser/extractors/pdf/__init__.py +0 -0
- {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/universal_parser/extractors/pdf/native.py +0 -0
- {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/universal_parser/extractors/pdf/tables.py +0 -0
- {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/universal_parser/extractors/pdf/visual_onnx.py +0 -0
- {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/universal_parser/extractors/structured/__init__.py +0 -0
- {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/universal_parser/extractors/structured/csv_extractor.py +0 -0
- {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/universal_parser/extractors/structured/json_xml_extractor.py +0 -0
- {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/universal_parser/extractors/structured/parquet_extractor.py +0 -0
- {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/universal_parser/extractors/web/__init__.py +0 -0
- {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/universal_parser/extractors/web/epub_extractor.py +0 -0
- {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/universal_parser/extractors/web/html_extractor.py +0 -0
- {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/universal_parser/mcp/__init__.py +0 -0
- {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/universal_parser/mcp/server.py +0 -0
- {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/universal_parser/observability/__init__.py +0 -0
- {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/universal_parser/observability/dashboard.py +0 -0
- {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/universal_parser/observability/logger.py +0 -0
- {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/universal_parser/observability/metrics.py +0 -0
|
@@ -0,0 +1,49 @@
|
|
|
1
|
+
name: Sync to Hugging Face Spaces
|
|
2
|
+
|
|
3
|
+
on:
|
|
4
|
+
push:
|
|
5
|
+
branches:
|
|
6
|
+
- main
|
|
7
|
+
|
|
8
|
+
jobs:
|
|
9
|
+
sync-to-hub:
|
|
10
|
+
runs-on: ubuntu-latest
|
|
11
|
+
steps:
|
|
12
|
+
- name: Checkout repository
|
|
13
|
+
uses: actions/checkout@v4
|
|
14
|
+
|
|
15
|
+
- name: Prepare Hugging Face README Header & Deploy Clean Tree
|
|
16
|
+
env:
|
|
17
|
+
HF_TOKEN: ${{ secrets.HF_TOKEN }}
|
|
18
|
+
run: |
|
|
19
|
+
cat << 'EOF' > hf_header.txt
|
|
20
|
+
---
|
|
21
|
+
title: Universal Document Parser
|
|
22
|
+
emoji: 📄
|
|
23
|
+
colorFrom: blue
|
|
24
|
+
colorTo: indigo
|
|
25
|
+
sdk: gradio
|
|
26
|
+
sdk_version: 5.9.0
|
|
27
|
+
app_file: main.py
|
|
28
|
+
pinned: false
|
|
29
|
+
license: mit
|
|
30
|
+
short_description: Zero-GPU CPU document ingestion engine for RAG pipelines.
|
|
31
|
+
---
|
|
32
|
+
|
|
33
|
+
EOF
|
|
34
|
+
cat hf_header.txt README.md > README_HF.md
|
|
35
|
+
mv README_HF.md README.md
|
|
36
|
+
rm hf_header.txt
|
|
37
|
+
|
|
38
|
+
# Configure git
|
|
39
|
+
git config --global user.email "actions@github.com"
|
|
40
|
+
git config --global user.name "GitHub Actions"
|
|
41
|
+
|
|
42
|
+
# Remove binary files and test fixtures for HF deployment
|
|
43
|
+
rm -rf assets/ tests/fixtures/ .git/
|
|
44
|
+
|
|
45
|
+
# Initialize a fresh single-commit git repository to strip any historical binary blobs
|
|
46
|
+
git init -b main
|
|
47
|
+
git add README.md main.py pyproject.toml requirements.txt universal_parser/ benchmarks/ LICENSE
|
|
48
|
+
git commit -m "deploy: live Hugging Face Space application"
|
|
49
|
+
git push -f https://Karan6124:$HF_TOKEN@huggingface.co/spaces/Karan6124/universal-doc-parser main
|
|
@@ -1,7 +1,12 @@
|
|
|
1
1
|
Metadata-Version: 2.5
|
|
2
2
|
Name: universal-doc-parser
|
|
3
|
-
Version: 1.0.
|
|
4
|
-
Summary: A zero-GPU, CPU-only
|
|
3
|
+
Version: 1.0.2
|
|
4
|
+
Summary: A zero-GPU, CPU-only document ingestion engine for RAG pipelines and AI agents.
|
|
5
|
+
Project-URL: Homepage, https://github.com/Edge-Explorer/Parse-Anything-
|
|
6
|
+
Project-URL: Live Demo, https://huggingface.co/spaces/Karan6124/universal-doc-parser
|
|
7
|
+
Project-URL: Repository, https://github.com/Edge-Explorer/Parse-Anything-
|
|
8
|
+
Project-URL: Documentation, https://github.com/Edge-Explorer/Parse-Anything-#readme
|
|
9
|
+
Project-URL: Issues, https://github.com/Edge-Explorer/Parse-Anything-/issues
|
|
5
10
|
License: MIT
|
|
6
11
|
License-File: LICENSE
|
|
7
12
|
Requires-Python: >=3.11
|
|
@@ -26,6 +31,9 @@ Requires-Dist: rapidocr-onnxruntime>=1.3.0
|
|
|
26
31
|
Requires-Dist: selectolax>=0.4.11
|
|
27
32
|
Requires-Dist: xlrd>=2.0.1
|
|
28
33
|
Requires-Dist: xlwt>=1.3.0
|
|
34
|
+
Provides-Extra: demo
|
|
35
|
+
Requires-Dist: gradio>=4.0.0; extra == 'demo'
|
|
36
|
+
Requires-Dist: plotly>=5.0.0; extra == 'demo'
|
|
29
37
|
Provides-Extra: dev
|
|
30
38
|
Requires-Dist: hypothesis; extra == 'dev'
|
|
31
39
|
Requires-Dist: psutil>=5.9.0; extra == 'dev'
|
|
@@ -40,69 +48,100 @@ Requires-Dist: rapidocr-onnxruntime>=1.3.0; extra == 'ocr'
|
|
|
40
48
|
Description-Content-Type: text/markdown
|
|
41
49
|
|
|
42
50
|
<p align="center">
|
|
43
|
-
<img src="assets/banner.png"
|
|
51
|
+
<img src="https://raw.githubusercontent.com/Edge-Explorer/Parse-Anything-/main/assets/banner.png" alt="Universal Document Parser Banner" width="800" />
|
|
44
52
|
</p>
|
|
45
53
|
|
|
46
54
|
# UNIVERSAL PARSER
|
|
47
55
|
|
|
48
56
|
[](https://github.com/Edge-Explorer/Parse-Anything-/actions/workflows/ci.yml)
|
|
49
|
-
[](https://pypi.org/project/universal-doc-parser/)
|
|
58
|
+
[](https://huggingface.co/spaces/Karan6124/universal-doc-parser)
|
|
59
|
+
[](https://pypi.org/project/universal-doc-parser/)
|
|
51
60
|
[](LICENSE)
|
|
52
61
|
[](#memory-and-performance)
|
|
53
62
|
|
|
54
|
-
A
|
|
63
|
+
A CPU-only document ingestion engine for RAG pipelines and AI agents. Parses 15+ file formats into a unified Pydantic schema with hierarchical chunking, adaptive layout fingerprinting, and a FastMCP server interface.
|
|
55
64
|
|
|
56
|
-
|
|
65
|
+
**[Try the Live Interactive Demo on Hugging Face Spaces](https://huggingface.co/spaces/Karan6124/universal-doc-parser)**
|
|
66
|
+
|
|
67
|
+
No GPU. No paid API. No recurring cost.
|
|
57
68
|
|
|
58
69
|
---
|
|
59
70
|
|
|
60
71
|
## Table of Contents
|
|
61
72
|
|
|
62
|
-
- [
|
|
63
|
-
- [
|
|
73
|
+
- [Problem and Scope](#problem-and-scope)
|
|
74
|
+
- [Related Work](#related-work)
|
|
75
|
+
- [The One Original Contribution](#the-one-original-contribution)
|
|
64
76
|
- [Supported Formats](#supported-formats)
|
|
65
77
|
- [Installation](#installation)
|
|
66
78
|
- [Quickstart](#quickstart)
|
|
67
79
|
- [Output Schema](#output-schema)
|
|
68
80
|
- [API Reference](#api-reference)
|
|
69
81
|
- [Architecture](#architecture)
|
|
82
|
+
- [Adaptive Layout Fingerprinting](#adaptive-layout-fingerprinting)
|
|
70
83
|
- [FastMCP Server](#fastmcp-server)
|
|
71
84
|
- [Memory and Performance](#memory-and-performance)
|
|
72
|
-
- [
|
|
85
|
+
- [Benchmarks](#benchmarks)
|
|
86
|
+
- [Limitations](#limitations)
|
|
73
87
|
- [Contributing](#contributing)
|
|
74
88
|
- [License](#license)
|
|
75
89
|
|
|
76
90
|
---
|
|
77
91
|
|
|
78
|
-
##
|
|
92
|
+
## Problem and Scope
|
|
93
|
+
|
|
94
|
+
Feeding real-world documents into a RAG pipeline is harder than it looks.
|
|
95
|
+
|
|
96
|
+
A PDF is a stream of positioned drawing commands. Reading order breaks on multi-column layouts. Tables without visible borders are invisible to naive text extraction. Scanned pages have no machine-readable text. Every file format requires a different library, and those libraries return different data structures — making a consistent, type-safe downstream pipeline difficult to build.
|
|
97
|
+
|
|
98
|
+
This library handles the extraction and normalization layer: MIME sniffing, format routing, multi-column reading order reconstruction, two-pass table detection, OCR fallback, and output to a single validated Pydantic schema — for 15+ file formats, from a single `parse(path)` call, running entirely on CPU.
|
|
99
|
+
|
|
100
|
+
It does **not** do layout model inference, deep learning-based element classification, or PDF reconstruction at the visual rendering layer. For those capabilities, see Docling.
|
|
101
|
+
|
|
102
|
+
---
|
|
103
|
+
|
|
104
|
+
## Related Work
|
|
79
105
|
|
|
80
|
-
|
|
106
|
+
Honest comparison with the tools a practitioner would actually evaluate before using this one.
|
|
81
107
|
|
|
82
|
-
|
|
108
|
+
**[Docling](https://github.com/DS4SD/docling)** (IBM Research / Linux Foundation, MIT-adjacent license)
|
|
109
|
+
Docling ships a trained document layout analysis model and a PDF table structure recognition model. Its classification quality on complex PDFs — dense academic papers, financial reports with borderless tables — is substantially higher than any heuristic-based approach including this one. If layout accuracy on complex PDFs is your primary concern, evaluate Docling first. It is CPU-capable and actively maintained by a funded team.
|
|
83
110
|
|
|
84
|
-
|
|
111
|
+
**[Marker](https://github.com/VikParuchuri/marker)** (Apache-2.0)
|
|
112
|
+
Marker uses a fine-tuned Surya OCR model and a layout segmentation model. It produces high-quality Markdown from PDFs, including scanned documents. Its OCR and rendering quality on scientific and academic PDFs is superior to the RapidOCR fallback used here. If your pipeline is primarily scientific PDFs, evaluate Marker first.
|
|
85
113
|
|
|
86
|
-
|
|
87
|
-
|
|
88
|
-
- **Single-format tools (pdfminer, mammoth, etc.):** Each handles one format. Building a multi-format pipeline requires a new library, a new schema, and new tests for every format.
|
|
114
|
+
**[Unstructured](https://github.com/Unstructured-IO/unstructured)** (Apache-2.0, managed API available)
|
|
115
|
+
Unstructured supports the broadest format coverage in the space (40+ formats) with optional hi-res partition mode and a managed API. If format breadth or managed infrastructure is a priority, evaluate Unstructured first.
|
|
89
116
|
|
|
90
|
-
|
|
117
|
+
**Where this library differs:**
|
|
118
|
+
|
|
119
|
+
| Criterion | Docling | Marker | Unstructured | Universal Parser |
|
|
120
|
+
|---|---|---|---|---|
|
|
121
|
+
| Layout model inference | Yes (trained model) | Yes (Surya) | Optional (hi-res mode) | No (heuristics only) |
|
|
122
|
+
| Table structure recognition | Trained model | Limited | Optional | Heuristic (lattice + stream) |
|
|
123
|
+
| OCR quality | Good | Excellent | Good | Adequate (RapidOCR) |
|
|
124
|
+
| Format coverage | PDF, DOCX, XLSX, PPTX, HTML | PDF, images | 40+ formats | 15+ formats |
|
|
125
|
+
| Python version | 3.9+ | 3.9+ | 3.9+ | 3.11+ |
|
|
126
|
+
| Memory footprint | Moderate (model weights) | Higher (model weights) | Varies | <250 MB RSS (no model weights) |
|
|
127
|
+
| Per-template auto-tuning | No | No | No | Yes (see below) |
|
|
128
|
+
| MCP server interface | No | No | No | Yes |
|
|
129
|
+
|
|
130
|
+
The heuristic approach used here extracts less accurately on complex layouts than Docling or Marker. The trade-off is zero model weights, lower memory, and a per-template configuration learning mechanism described in the next section.
|
|
91
131
|
|
|
92
132
|
---
|
|
93
133
|
|
|
94
|
-
##
|
|
134
|
+
## The One Original Contribution
|
|
135
|
+
|
|
136
|
+
The piece of this library that does not exist in the same form in Docling, Marker, or Unstructured is the **adaptive layout fingerprinting and per-template auto-tuner**.
|
|
137
|
+
|
|
138
|
+
Many real RAG pipelines process the same document template repeatedly — the same invoice format thousands of times, the same SEC 10-K filing structure across years, the same internal report template across departments. In these workloads, the failure mode of heuristic extractors is predictable and reproducible: the same column threshold is wrong on the same template, every time.
|
|
95
139
|
|
|
96
|
-
-
|
|
97
|
-
|
|
98
|
-
|
|
99
|
-
|
|
100
|
-
|
|
101
|
-
- **Streams large files page-by-page** through generator pipelines without loading the full document into memory
|
|
102
|
-
- Exports directly to **Markdown**, **RAG-ready hierarchical chunks with breadcrumb context**, and **knowledge graph triples**
|
|
103
|
-
- Fingerprints document layouts using a **content-agnostic 2D spatial histogram** and caches optimal extraction parameters per template
|
|
104
|
-
- Exposes a **FastMCP server interface** for Claude Desktop, Cursor, and AI agent frameworks
|
|
105
|
-
- Collects ingestion **observability telemetry** and exports a standalone interactive HTML dashboard
|
|
140
|
+
The fingerprinter computes a content-agnostic 10x10 spatial occupancy grid from element bounding boxes, hashes it with SHA-256, and uses it as a stable template identity. The auto-tuner runs coordinate descent over the parameter space against a reference ground-truth document for that template, then persists the optimal configuration to a JSON cache. On subsequent files matching the same fingerprint (exact or fuzzy Cosine-Jaccard similarity above a configurable threshold), the cached configuration is applied automatically.
|
|
141
|
+
|
|
142
|
+
This is a focused, narrow contribution: it does not make the base extraction better than Docling on arbitrary documents. It reduces error variance on recurring templates where a heuristic extractor's default parameters are consistently wrong.
|
|
143
|
+
|
|
144
|
+
The ablation study validating this claim on a real document set is [planned and tracked here](#benchmarks).
|
|
106
145
|
|
|
107
146
|
---
|
|
108
147
|
|
|
@@ -122,7 +161,8 @@ This library solves all of it from a single function call, running entirely on C
|
|
|
122
161
|
| Structured | CSV / TSV | `.csv`, `.tsv` | Python `csv.Sniffer` dialect auto-detection |
|
|
123
162
|
| Structured | Parquet | `.parquet` | `pyarrow` zero-copy columnar record batch streaming |
|
|
124
163
|
| Structured | JSON / XML | `.json`, `.xml` | Recursive tree flattening and normalized schema mapping |
|
|
125
|
-
| Email | EML / MBOX
|
|
164
|
+
| Email | EML / MBOX | `.eml`, `.mbox` | `email` stdlib with recursive embedded attachment routing |
|
|
165
|
+
| Email | MSG (Outlook) | `.msg` | `extract-msg` (GPL-3.0 — see license section) with recursive attachment routing |
|
|
126
166
|
|
|
127
167
|
---
|
|
128
168
|
|
|
@@ -131,39 +171,42 @@ This library solves all of it from a single function call, running entirely on C
|
|
|
131
171
|
Requires Python 3.11 or later.
|
|
132
172
|
|
|
133
173
|
```bash
|
|
134
|
-
pip install universal-parser
|
|
174
|
+
pip install universal-doc-parser
|
|
135
175
|
```
|
|
136
176
|
|
|
137
177
|
With OCR support for scanned PDFs and raster images:
|
|
138
178
|
|
|
139
179
|
```bash
|
|
140
|
-
pip install "universal-parser[ocr]"
|
|
180
|
+
pip install "universal-doc-parser[ocr]"
|
|
141
181
|
```
|
|
142
182
|
|
|
143
183
|
With development tooling:
|
|
144
184
|
|
|
145
185
|
```bash
|
|
146
|
-
pip install "universal-parser[dev]"
|
|
186
|
+
pip install "universal-doc-parser[dev]"
|
|
147
187
|
```
|
|
148
188
|
|
|
149
189
|
Using uv:
|
|
150
190
|
|
|
151
191
|
```bash
|
|
152
|
-
uv add universal-parser
|
|
153
|
-
uv add "universal-parser[ocr]"
|
|
192
|
+
uv add universal-doc-parser
|
|
193
|
+
uv add "universal-doc-parser[ocr]"
|
|
154
194
|
```
|
|
155
195
|
|
|
156
|
-
System dependencies on Linux
|
|
196
|
+
System dependencies on Linux (for magic-byte MIME detection — optional, falls back to extension sniffing if absent):
|
|
157
197
|
|
|
158
198
|
```bash
|
|
159
|
-
# Ubuntu / Debian
|
|
160
|
-
sudo apt-get install -y
|
|
199
|
+
# Ubuntu / Debian Bookworm (Python 3.10 era containers)
|
|
200
|
+
sudo apt-get install -y libmagic1t64
|
|
201
|
+
|
|
202
|
+
# Ubuntu / Debian Bullseye and earlier
|
|
203
|
+
sudo apt-get install -y libmagic1
|
|
161
204
|
|
|
162
205
|
# Fedora / RHEL
|
|
163
|
-
sudo dnf install -y file-libs
|
|
206
|
+
sudo dnf install -y file-libs
|
|
164
207
|
```
|
|
165
208
|
|
|
166
|
-
On macOS and Windows
|
|
209
|
+
On macOS and Windows, MIME detection is handled through the Python package layer automatically.
|
|
167
210
|
|
|
168
211
|
---
|
|
169
212
|
|
|
@@ -216,7 +259,7 @@ for chunk in chunks:
|
|
|
216
259
|
print("---")
|
|
217
260
|
```
|
|
218
261
|
|
|
219
|
-
Each chunk carries its full heading ancestry (H1 > H2 > H3) prepended as context. Retrieval models
|
|
262
|
+
Each chunk carries its full heading ancestry (H1 > H2 > H3) prepended as context. Retrieval models receive structurally anchored chunks rather than arbitrary token windows.
|
|
220
263
|
|
|
221
264
|
### Export a knowledge graph
|
|
222
265
|
|
|
@@ -304,7 +347,7 @@ Every supported format produces the same output structure.
|
|
|
304
347
|
| `metadata.has_scanned_pages` | `bool` | `true` if any page required OCR processing. |
|
|
305
348
|
| `element_id` | `string` | UUID v4 per element. |
|
|
306
349
|
| `type` | `enum` | One of: `heading`, `paragraph`, `table`, `figure`, `list_item`, `code_block`. |
|
|
307
|
-
| `level` | `int or null` | Heading depth 1
|
|
350
|
+
| `level` | `int or null` | Heading depth 1-6. `null` for non-heading elements. |
|
|
308
351
|
| `text` | `string or null` | Plain text content. `null` for pure table elements. |
|
|
309
352
|
| `page` | `int or null` | 1-indexed page number. `null` for formats without page structure. |
|
|
310
353
|
| `bbox` | `BBox or null` | `{x0, y0, x1, y1}` in PDF points (72 pt = 1 inch). `null` for non-spatial formats. |
|
|
@@ -327,7 +370,7 @@ doc: Document = parse(path)
|
|
|
327
370
|
|
|
328
371
|
The main entry point. Accepts `str` or `pathlib.Path`. Performs format sniffing, routing, extraction, and Pydantic validation. Returns a fully validated `Document`.
|
|
329
372
|
|
|
330
|
-
Raises `FileNotFoundError` if the path does not exist. Raises `ValueError` for corrupt, unreadable, or unrecognized files. All internal extractor errors are caught and surfaced as `ValueError` with context.
|
|
373
|
+
Raises `FileNotFoundError` if the path does not exist. Raises `ValueError` for corrupt, unreadable, or unrecognized files. All internal extractor errors are caught and surfaced as `ValueError` with context.
|
|
331
374
|
|
|
332
375
|
---
|
|
333
376
|
|
|
@@ -354,8 +397,8 @@ chunks: list[Chunk] = to_chunks(doc, max_tokens=512, overlap_tokens=50)
|
|
|
354
397
|
Hierarchical token-aware chunker.
|
|
355
398
|
|
|
356
399
|
Parameters:
|
|
357
|
-
- `max_tokens`
|
|
358
|
-
- `overlap_tokens`
|
|
400
|
+
- `max_tokens` - Maximum estimated tokens per chunk. Default: `512`. Estimated at approximately 4 characters per token.
|
|
401
|
+
- `overlap_tokens` - Reserved for future sliding window chunking.
|
|
359
402
|
|
|
360
403
|
Chunk fields:
|
|
361
404
|
|
|
@@ -433,7 +476,7 @@ Adding a new format requires creating one file in `universal_parser/extractors/`
|
|
|
433
476
|
|
|
434
477
|
PDF extraction runs in two cooperative passes over each page.
|
|
435
478
|
|
|
436
|
-
**Pass 1
|
|
479
|
+
**Pass 1 - Table Detection** (`extractors/pdf/tables.py`)
|
|
437
480
|
|
|
438
481
|
Two strategies are attempted in sequence:
|
|
439
482
|
|
|
@@ -442,7 +485,7 @@ Two strategies are attempted in sequence:
|
|
|
442
485
|
|
|
443
486
|
Confidence scores are assigned based on cell uniformity and structural regularity.
|
|
444
487
|
|
|
445
|
-
**Pass 2
|
|
488
|
+
**Pass 2 - Text and Heading Extraction** (`extractors/pdf/native.py`)
|
|
446
489
|
|
|
447
490
|
- Character positions are read from the PDFium character map for every non-table region.
|
|
448
491
|
- Font sizes across the page are collected and percentile thresholds computed. Elements at or above the 95th percentile are classified H1, the 85th percentile H2, and the 75th percentile H3. All remaining text is classified as paragraphs.
|
|
@@ -451,9 +494,9 @@ Confidence scores are assigned based on cell uniformity and structural regularit
|
|
|
451
494
|
|
|
452
495
|
---
|
|
453
496
|
|
|
454
|
-
|
|
497
|
+
## Adaptive Layout Fingerprinting
|
|
455
498
|
|
|
456
|
-
Documents that follow recurring templates
|
|
499
|
+
Documents that follow recurring templates - invoices, financial reports, regulatory filings - can be registered once and reused across thousands of files with cached optimal extraction parameters.
|
|
457
500
|
|
|
458
501
|
**Fingerprinting** (`adaptive/fingerprint.py`)
|
|
459
502
|
|
|
@@ -464,7 +507,7 @@ from universal_parser.adaptive import compute_fingerprint
|
|
|
464
507
|
|
|
465
508
|
doc = parse("invoice_template.pdf")
|
|
466
509
|
fp = compute_fingerprint(doc)
|
|
467
|
-
print(fp.hash_digest)
|
|
510
|
+
print(fp.hash_digest) # SHA-256 hex string
|
|
468
511
|
print(fp.spatial_grid) # 10x10 occupancy matrix
|
|
469
512
|
```
|
|
470
513
|
|
|
@@ -494,6 +537,8 @@ optimal_config = auto_tune(
|
|
|
494
537
|
cache.set(fp.hash_digest, optimal_config)
|
|
495
538
|
```
|
|
496
539
|
|
|
540
|
+
The ablation study measuring extraction accuracy with auto-tuning on vs off across a real set of recurring templates is in progress and will be published here when complete with methodology and sample sizes disclosed.
|
|
541
|
+
|
|
497
542
|
---
|
|
498
543
|
|
|
499
544
|
### Observability System
|
|
@@ -543,12 +588,12 @@ Configure in Claude Desktop (`claude_desktop_config.json`):
|
|
|
543
588
|
```json
|
|
544
589
|
{
|
|
545
590
|
"mcpServers": {
|
|
546
|
-
"universal-parser": {
|
|
591
|
+
"universal-doc-parser": {
|
|
547
592
|
"command": "uv",
|
|
548
593
|
"args": [
|
|
549
594
|
"run",
|
|
550
595
|
"--with",
|
|
551
|
-
"universal-parser",
|
|
596
|
+
"universal-doc-parser",
|
|
552
597
|
"python",
|
|
553
598
|
"-m",
|
|
554
599
|
"universal_parser.mcp.server"
|
|
@@ -566,15 +611,15 @@ After restarting Claude Desktop, Claude can invoke `parse_document` against any
|
|
|
566
611
|
|
|
567
612
|
The streaming generator architecture bounds memory consumption regardless of document length. No page is held in memory after it is yielded.
|
|
568
613
|
|
|
569
|
-
|
|
614
|
+
The numbers below are from the memory benchmark suite (`benchmarks/memory_profile.py`) running on an Intel Core i7, 16 GB RAM, no GPU. These are heuristic estimates from a controlled synthetic document, not profiled runs on varied real-world corpora. Real-world peak RSS will vary depending on document complexity, table density, and OCR engagement.
|
|
570
615
|
|
|
571
616
|
| Document Size | Peak RSS Memory | Processing Time |
|
|
572
617
|
|---|---|---|
|
|
573
|
-
| 100 pages |
|
|
574
|
-
| 500 pages | 165 MB |
|
|
575
|
-
| 1,000 pages |
|
|
618
|
+
| 100 pages | ~150 MB | ~5 seconds |
|
|
619
|
+
| 500 pages | ~165 MB | ~24 seconds |
|
|
620
|
+
| 1,000 pages | ~180 MB | ~49 seconds |
|
|
576
621
|
|
|
577
|
-
The 250 MB RSS hard limit is asserted in the
|
|
622
|
+
The 250 MB RSS hard limit is asserted in the benchmark suite on every CI push:
|
|
578
623
|
|
|
579
624
|
```bash
|
|
580
625
|
uv run python benchmarks/memory_profile.py
|
|
@@ -582,50 +627,48 @@ uv run python benchmarks/memory_profile.py
|
|
|
582
627
|
|
|
583
628
|
---
|
|
584
629
|
|
|
585
|
-
##
|
|
586
|
-
|
|
587
|
-
|
|
588
|
-
|
|
589
|
-
|
|
590
|
-
|
|
591
|
-
|
|
592
|
-
|
|
593
|
-
|
|
594
|
-
|
|
595
|
-
| Engine / Model | Provider | Latency | Cost / 10-Page Doc | Table Accuracy | RAG Faithfulness |
|
|
596
|
-
|---|---|---|---|---|---|
|
|
597
|
-
| **Universal Parser (CPU)** | **Local** | **~450 ms** | **$0.00000** | **98.5%** | **99.0%** |
|
|
598
|
-
| Gemini 2.5 Flash | Google | 1,450 ms | $0.00075 | 96.0% | 97.5% |
|
|
599
|
-
| Gemini 2.5 Pro | Google | 2,850 ms | $0.00350 | 97.5% | 98.5% |
|
|
600
|
-
| GPT-4o | OpenAI | 2,100 ms | $0.01250 | 96.5% | 98.0% |
|
|
601
|
-
| GPT-4o Mini | OpenAI | 1,250 ms | $0.00075 | 93.0% | 95.0% |
|
|
602
|
-
| Claude 3.5 Sonnet | Anthropic | 2,400 ms | $0.01500 | 97.0% | 98.5% |
|
|
603
|
-
| Claude 3 Opus | Anthropic | 3,900 ms | $0.07500 | 98.0% | 99.0% |
|
|
604
|
-
| DeepSeek V3 | DeepSeek | 1,600 ms | $0.00085 | 95.5% | 97.0% |
|
|
605
|
-
| DeepSeek R1 | DeepSeek | 3,200 ms | $0.00280 | 97.0% | 98.0% |
|
|
606
|
-
| Qwen 2.5 72B | Alibaba | 1,750 ms | $0.00180 | 95.0% | 96.5% |
|
|
607
|
-
| Qwen 2.5 Coder | Alibaba | 1,650 ms | $0.00150 | 94.5% | 96.0% |
|
|
608
|
-
| Llama 3.3 70B | Meta | 1,350 ms | $0.00190 | 94.0% | 97.0% |
|
|
609
|
-
| Mistral Large 2411 | Mistral AI | 1,680 ms | $0.01000 | 94.0% | 97.0% |
|
|
610
|
-
| Kimi k1.5 | Moonshot AI | 1,420 ms | $0.00600 | 93.0% | 95.0% |
|
|
611
|
-
| GLM-4 9B | Zhipu AI | 1,150 ms | $0.00050 | 91.0% | 94.0% |
|
|
612
|
-
| Command R+ | Cohere | 1,550 ms | $0.01250 | 93.0% | 96.0% |
|
|
613
|
-
|
|
614
|
-
Cloud LLMs predict every character autoregressively from a visual or token representation of the document. Universal Parser reads the underlying binary vector streams and coordinate data directly. Table borders, cell boundaries, and reading order are computed geometrically from exact floating-point positions — there is no prediction step and therefore no hallucination risk at the extraction layer.
|
|
615
|
-
|
|
616
|
-
The RAG faithfulness score follows from the hierarchical chunker. Every chunk carries its full heading ancestry prepended as context. Retrieval models and downstream LLMs receive structurally anchored chunks rather than arbitrary token windows, eliminating the most common source of retrieval hallucination.
|
|
617
|
-
|
|
618
|
-
Run the benchmark suite:
|
|
630
|
+
## Benchmarks
|
|
631
|
+
|
|
632
|
+
### What exists today
|
|
633
|
+
|
|
634
|
+
The benchmark suite at `benchmarks/run_llm_benchmark.py` compares this library's extraction output against live API calls to Gemini 2.5 Flash, GPT-4o, DeepSeek V3, Qwen 2.5 72B, Llama 3.3 70B, and other models accessible via the Google AI Studio and OpenRouter free tiers.
|
|
635
|
+
|
|
636
|
+
14 of the 15 rows in the original benchmark table were live API results. The two Claude rows (Claude 3.5 Sonnet and Claude 3 Opus) were simulated estimates — Anthropic does not expose Claude on any free API tier, and the original README did not disclose this distinction. Those rows have been removed from published tables until a properly labeled live run can be completed.
|
|
637
|
+
|
|
638
|
+
Run the benchmark suite yourself:
|
|
619
639
|
|
|
620
640
|
```bash
|
|
621
|
-
# Offline
|
|
641
|
+
# Offline mode — simulates responses, zero cost, no API keys required
|
|
622
642
|
uv run python benchmarks/run_llm_benchmark.py
|
|
623
643
|
|
|
624
|
-
# Live mode
|
|
644
|
+
# Live mode — runs real API calls against Gemini and OpenRouter models
|
|
625
645
|
GEMINI_API_KEY=your_key OPENROUTER_API_KEY=your_key \
|
|
626
646
|
uv run python benchmarks/run_llm_benchmark.py --live
|
|
627
647
|
```
|
|
628
648
|
|
|
649
|
+
### What is planned
|
|
650
|
+
|
|
651
|
+
The benchmark work that would make this project defensible — and which does not yet exist — is:
|
|
652
|
+
|
|
653
|
+
1. **Head-to-head against Docling, Marker, and Unstructured** on the same document corpus (target: SEC EDGAR 10-K filings and PubTables-1M) with disclosed sample sizes and a documented scoring methodology.
|
|
654
|
+
|
|
655
|
+
2. **Auto-tuning ablation study:** Extraction accuracy with fingerprint-based auto-tuning ON vs OFF, across a set of 20-30 recurring invoice and filing templates, with sample sizes and error metric definition stated explicitly.
|
|
656
|
+
|
|
657
|
+
These are the two experiments that would either validate or invalidate the claims this project is making. Until they exist, treat the current benchmark numbers as directional indicators, not validated results.
|
|
658
|
+
|
|
659
|
+
---
|
|
660
|
+
|
|
661
|
+
## Limitations
|
|
662
|
+
|
|
663
|
+
These are known failure modes, not edge cases.
|
|
664
|
+
|
|
665
|
+
- **Complex PDF layouts:** On multi-column academic papers and dense financial reports with borderless tables, Docling's trained layout model will outperform the heuristic approach used here. If layout accuracy on complex PDFs is the primary requirement, use Docling.
|
|
666
|
+
- **Scanned document quality:** The RapidOCR fallback performs adequately on clean scans. On degraded, skewed, or low-resolution scans, Marker's Surya-based OCR pipeline will produce substantially better results.
|
|
667
|
+
- **Python version requirement:** This library requires Python 3.11+, which excludes some deployment environments. Docling, Marker, and Unstructured support Python 3.9+.
|
|
668
|
+
- **MSG parsing license:** The `extract-msg` dependency carries a GPL-3.0 license. MSG parsing is therefore subject to GPL-3.0 copyleft terms — see the license section for the full implication.
|
|
669
|
+
- **Memory numbers are heuristic estimates:** The memory table above was produced from a controlled synthetic document. Real-world peak RSS will vary.
|
|
670
|
+
- **Benchmark numbers are not externally validated:** No one outside of the author has run the full benchmark suite on the full dataset yet. Treat published numbers accordingly.
|
|
671
|
+
|
|
629
672
|
---
|
|
630
673
|
|
|
631
674
|
## Contributing
|
|
@@ -640,14 +683,14 @@ Adding a new format:
|
|
|
640
683
|
4. Import the module in `universal_parser/__init__.py`
|
|
641
684
|
5. Add fixture files in `tests/fixtures/<format>/` — at minimum three samples including one deliberately complex or malformed file
|
|
642
685
|
6. Write tests in `tests/test_<format>.py` asserting schema correctness, content accuracy, and graceful error handling
|
|
643
|
-
7. Add a `CHANGELOG.md` entry
|
|
686
|
+
7. Add a `docs/CHANGELOG.md` entry
|
|
644
687
|
|
|
645
688
|
Pull requests without fixture files and corresponding tests will not be reviewed.
|
|
646
689
|
|
|
647
690
|
Code standards:
|
|
648
691
|
- Pass `uv run ruff check .` with zero errors
|
|
649
692
|
- Format with `uv run ruff format .`
|
|
650
|
-
- No AGPL-licensed dependencies. All additions must carry MIT, Apache-2.0, or BSD licenses
|
|
693
|
+
- No new AGPL-licensed dependencies. All additions must carry MIT, Apache-2.0, or BSD licenses
|
|
651
694
|
|
|
652
695
|
Full local quality gate:
|
|
653
696
|
|
|
@@ -667,7 +710,7 @@ uv run python benchmarks/run_llm_benchmark.py
|
|
|
667
710
|
|
|
668
711
|
This project is licensed under the **MIT License**. See [LICENSE](LICENSE) for the full text.
|
|
669
712
|
|
|
670
|
-
All runtime dependencies carry permissive,
|
|
713
|
+
**License note on MSG support:** The `extract-msg` dependency used for Outlook `.msg` parsing is licensed under **GPL-3.0**. If you parse `.msg` files in a closed-source product, the GPL-3.0 copyleft terms apply to that use. All other runtime dependencies carry permissive licenses (MIT, Apache-2.0, BSD, LGPL-3.0). If your use case requires a fully permissive dependency tree, you can exclude `.msg` parsing by not calling `parse()` on `.msg` files and removing `extract-msg` from your installation.
|
|
671
714
|
|
|
672
715
|
| Dependency | License | Purpose |
|
|
673
716
|
|---|---|---|
|
|
@@ -685,8 +728,8 @@ All runtime dependencies carry permissive, commercially compatible licenses. The
|
|
|
685
728
|
| `python-pptx` | MIT | PowerPoint parsing |
|
|
686
729
|
| `xlrd` | BSD-3-Clause | Legacy XLS binary parsing |
|
|
687
730
|
| `olefile` | BSD-2-Clause | OLE compound file parsing |
|
|
688
|
-
| `extract-msg` | GPL-3.0 | Outlook MSG email parsing |
|
|
731
|
+
| `extract-msg` | **GPL-3.0** | Outlook MSG email parsing -- see note above |
|
|
689
732
|
| `beautifulsoup4` | MIT | HTML fallback parser |
|
|
690
733
|
| `mcp` | MIT | FastMCP server interface |
|
|
691
734
|
|
|
692
|
-
The previous dependency on `PyMuPDF` / `fitz` (AGPL-3.0) was removed in v0.1.0 and replaced with `pypdfium2` (Apache-2.0).
|
|
735
|
+
The previous dependency on `PyMuPDF` / `fitz` (AGPL-3.0) was removed in v0.1.0 and replaced with `pypdfium2` (Apache-2.0).
|