universal-doc-parser 1.0.0__tar.gz → 1.0.2__tar.gz

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (142) hide show
  1. {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/.github/workflows/ci.yml +1 -1
  2. universal_doc_parser-1.0.2/.github/workflows/hf_sync.yml +49 -0
  3. {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/PKG-INFO +141 -98
  4. {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/README.md +131 -96
  5. universal_doc_parser-1.0.2/assets/README.md +40 -0
  6. universal_doc_parser-1.0.2/benchmarks/README.md +16 -0
  7. universal_doc_parser-1.0.2/benchmarks/providers/README.md +22 -0
  8. universal_doc_parser-1.0.2/docs/ADDING_A_FORMAT.md +212 -0
  9. universal_doc_parser-1.0.2/docs/ARCHITECTURE.md +164 -0
  10. universal_doc_parser-1.0.2/docs/CHANGELOG.md +234 -0
  11. universal_doc_parser-1.0.2/docs/README.md +76 -0
  12. universal_doc_parser-1.0.2/docs/SCHEMA.md +183 -0
  13. universal_doc_parser-1.0.2/main.py +199 -0
  14. {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/pyproject.toml +13 -2
  15. universal_doc_parser-1.0.2/requirements.txt +20 -0
  16. universal_doc_parser-1.0.2/tests/README.md +118 -0
  17. {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/universal_parser/__init__.py +1 -1
  18. universal_doc_parser-1.0.2/universal_parser/adaptive/README.md +35 -0
  19. {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/universal_parser/core/sniffer.py +0 -1
  20. universal_doc_parser-1.0.2/universal_parser/enrichment/README.md +19 -0
  21. universal_doc_parser-1.0.2/universal_parser/exports/README.md +34 -0
  22. universal_doc_parser-1.0.2/universal_parser/extractors/README.md +117 -0
  23. universal_doc_parser-1.0.2/universal_parser/extractors/images/README.md +144 -0
  24. universal_doc_parser-1.0.2/universal_parser/extractors/mail/README.md +12 -0
  25. universal_doc_parser-1.0.2/universal_parser/extractors/office/README.md +173 -0
  26. universal_doc_parser-1.0.2/universal_parser/extractors/pdf/README.md +180 -0
  27. universal_doc_parser-1.0.2/universal_parser/extractors/structured/README.md +13 -0
  28. universal_doc_parser-1.0.2/universal_parser/extractors/web/README.md +12 -0
  29. universal_doc_parser-1.0.2/universal_parser/mcp/README.md +26 -0
  30. universal_doc_parser-1.0.2/universal_parser/observability/README.md +18 -0
  31. {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/uv.lock +630 -7
  32. universal_doc_parser-1.0.0/CHANGELOG.md +0 -78
  33. universal_doc_parser-1.0.0/docs/ADDING_A_FORMAT.md +0 -0
  34. universal_doc_parser-1.0.0/docs/ARCHITECTURE.md +0 -0
  35. universal_doc_parser-1.0.0/docs/CHANGELOG.md +0 -0
  36. universal_doc_parser-1.0.0/docs/SCHEMA.md +0 -0
  37. universal_doc_parser-1.0.0/main.py +0 -6
  38. {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/.github/workflows/release.yml +0 -0
  39. {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/.gitignore +0 -0
  40. {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/.python-version +0 -0
  41. {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/LICENSE +0 -0
  42. {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/assets/banner.png +0 -0
  43. {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/assets/logo.png +0 -0
  44. {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/benchmarks/memory_profile.py +0 -0
  45. {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/benchmarks/providers/__init__.py +0 -0
  46. {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/benchmarks/providers/anthropic/__init__.py +0 -0
  47. {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/benchmarks/providers/anthropic/opus.py +0 -0
  48. {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/benchmarks/providers/anthropic/sonnet.py +0 -0
  49. {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/benchmarks/providers/base.py +0 -0
  50. {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/benchmarks/providers/cohere/__init__.py +0 -0
  51. {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/benchmarks/providers/cohere/command_r_plus.py +0 -0
  52. {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/benchmarks/providers/deepseek/__init__.py +0 -0
  53. {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/benchmarks/providers/deepseek/r1.py +0 -0
  54. {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/benchmarks/providers/deepseek/v3.py +0 -0
  55. {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/benchmarks/providers/gemini/__init__.py +0 -0
  56. {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/benchmarks/providers/gemini/flash.py +0 -0
  57. {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/benchmarks/providers/gemini/pro.py +0 -0
  58. {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/benchmarks/providers/glm/__init__.py +0 -0
  59. {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/benchmarks/providers/glm/glm4.py +0 -0
  60. {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/benchmarks/providers/kimi/__init__.py +0 -0
  61. {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/benchmarks/providers/kimi/k1_5.py +0 -0
  62. {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/benchmarks/providers/meta/__init__.py +0 -0
  63. {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/benchmarks/providers/meta/llama3_3.py +0 -0
  64. {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/benchmarks/providers/mistral/__init__.py +0 -0
  65. {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/benchmarks/providers/mistral/mistral_large.py +0 -0
  66. {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/benchmarks/providers/openai/__init__.py +0 -0
  67. {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/benchmarks/providers/openai/gpt4o.py +0 -0
  68. {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/benchmarks/providers/openai/gpt4o_mini.py +0 -0
  69. {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/benchmarks/providers/qwen/__init__.py +0 -0
  70. {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/benchmarks/providers/qwen/qwen_72b.py +0 -0
  71. {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/benchmarks/providers/qwen/qwen_coder.py +0 -0
  72. {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/benchmarks/run_llm_benchmark.py +0 -0
  73. {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/benchmarks/test_messy_document.py +0 -0
  74. {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/tests/__init__.py +0 -0
  75. {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/tests/fixtures/csv/sample.csv +0 -0
  76. {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/tests/fixtures/csv/sample.tsv +0 -0
  77. {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/tests/fixtures/docx/sample.docx +0 -0
  78. {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/tests/fixtures/html/sample.html +0 -0
  79. {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/tests/fixtures/json/sample_nested.json +0 -0
  80. {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/tests/fixtures/json/sample_tabular.json +0 -0
  81. {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/tests/fixtures/pdf/sample.pdf +0 -0
  82. {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/tests/fixtures/xlsx/sample.xlsx +0 -0
  83. {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/tests/fixtures/xml/sample.xml +0 -0
  84. {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/tests/test_adaptive.py +0 -0
  85. {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/tests/test_benchmark_providers.py +0 -0
  86. {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/tests/test_csv.py +0 -0
  87. {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/tests/test_docx.py +0 -0
  88. {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/tests/test_epub.py +0 -0
  89. {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/tests/test_exports.py +0 -0
  90. {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/tests/test_html.py +0 -0
  91. {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/tests/test_image.py +0 -0
  92. {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/tests/test_json_xml.py +0 -0
  93. {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/tests/test_legacy.py +0 -0
  94. {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/tests/test_mail.py +0 -0
  95. {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/tests/test_mcp.py +0 -0
  96. {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/tests/test_observability.py +0 -0
  97. {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/tests/test_parquet.py +0 -0
  98. {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/tests/test_pdf.py +0 -0
  99. {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/tests/test_pptx.py +0 -0
  100. {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/tests/test_xlsx.py +0 -0
  101. {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/universal_parser/adaptive/__init__.py +0 -0
  102. {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/universal_parser/adaptive/config_cache.py +0 -0
  103. {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/universal_parser/adaptive/fingerprint.py +0 -0
  104. {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/universal_parser/adaptive/tuner.py +0 -0
  105. {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/universal_parser/core/__init__.py +0 -0
  106. {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/universal_parser/core/engine.py +0 -0
  107. {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/universal_parser/core/router.py +0 -0
  108. {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/universal_parser/core/schema.py +0 -0
  109. {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/universal_parser/enrichment/__init__.py +0 -0
  110. {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/universal_parser/enrichment/vlm_enricher.py +0 -0
  111. {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/universal_parser/exports/__init__.py +0 -0
  112. {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/universal_parser/exports/to_chunks.py +0 -0
  113. {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/universal_parser/exports/to_graph.py +0 -0
  114. {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/universal_parser/exports/to_markdown.py +0 -0
  115. {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/universal_parser/extractors/__init__.py +0 -0
  116. {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/universal_parser/extractors/base.py +0 -0
  117. {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/universal_parser/extractors/images/__init__.py +0 -0
  118. {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/universal_parser/extractors/images/scan_extractor.py +0 -0
  119. {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/universal_parser/extractors/mail/__init__.py +0 -0
  120. {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/universal_parser/extractors/mail/mail_extractor.py +0 -0
  121. {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/universal_parser/extractors/office/__init__.py +0 -0
  122. {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/universal_parser/extractors/office/docx_extractor.py +0 -0
  123. {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/universal_parser/extractors/office/legacy_extractor.py +0 -0
  124. {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/universal_parser/extractors/office/pptx_extractor.py +0 -0
  125. {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/universal_parser/extractors/office/xlsx_extractor.py +0 -0
  126. {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/universal_parser/extractors/pdf/__init__.py +0 -0
  127. {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/universal_parser/extractors/pdf/native.py +0 -0
  128. {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/universal_parser/extractors/pdf/tables.py +0 -0
  129. {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/universal_parser/extractors/pdf/visual_onnx.py +0 -0
  130. {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/universal_parser/extractors/structured/__init__.py +0 -0
  131. {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/universal_parser/extractors/structured/csv_extractor.py +0 -0
  132. {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/universal_parser/extractors/structured/json_xml_extractor.py +0 -0
  133. {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/universal_parser/extractors/structured/parquet_extractor.py +0 -0
  134. {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/universal_parser/extractors/web/__init__.py +0 -0
  135. {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/universal_parser/extractors/web/epub_extractor.py +0 -0
  136. {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/universal_parser/extractors/web/html_extractor.py +0 -0
  137. {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/universal_parser/mcp/__init__.py +0 -0
  138. {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/universal_parser/mcp/server.py +0 -0
  139. {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/universal_parser/observability/__init__.py +0 -0
  140. {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/universal_parser/observability/dashboard.py +0 -0
  141. {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/universal_parser/observability/logger.py +0 -0
  142. {universal_doc_parser-1.0.0 → universal_doc_parser-1.0.2}/universal_parser/observability/metrics.py +0 -0
@@ -42,7 +42,7 @@ jobs:
42
42
  brew install libmagic
43
43
 
44
44
  - name: Install Python dependencies
45
- run: uv sync --all-extras
45
+ run: uv sync --all-extras --frozen
46
46
 
47
47
  - name: Lint with Ruff
48
48
  run: uv run ruff check .
@@ -0,0 +1,49 @@
1
+ name: Sync to Hugging Face Spaces
2
+
3
+ on:
4
+ push:
5
+ branches:
6
+ - main
7
+
8
+ jobs:
9
+ sync-to-hub:
10
+ runs-on: ubuntu-latest
11
+ steps:
12
+ - name: Checkout repository
13
+ uses: actions/checkout@v4
14
+
15
+ - name: Prepare Hugging Face README Header & Deploy Clean Tree
16
+ env:
17
+ HF_TOKEN: ${{ secrets.HF_TOKEN }}
18
+ run: |
19
+ cat << 'EOF' > hf_header.txt
20
+ ---
21
+ title: Universal Document Parser
22
+ emoji: 📄
23
+ colorFrom: blue
24
+ colorTo: indigo
25
+ sdk: gradio
26
+ sdk_version: 5.9.0
27
+ app_file: main.py
28
+ pinned: false
29
+ license: mit
30
+ short_description: Zero-GPU CPU document ingestion engine for RAG pipelines.
31
+ ---
32
+
33
+ EOF
34
+ cat hf_header.txt README.md > README_HF.md
35
+ mv README_HF.md README.md
36
+ rm hf_header.txt
37
+
38
+ # Configure git
39
+ git config --global user.email "actions@github.com"
40
+ git config --global user.name "GitHub Actions"
41
+
42
+ # Remove binary files and test fixtures for HF deployment
43
+ rm -rf assets/ tests/fixtures/ .git/
44
+
45
+ # Initialize a fresh single-commit git repository to strip any historical binary blobs
46
+ git init -b main
47
+ git add README.md main.py pyproject.toml requirements.txt universal_parser/ benchmarks/ LICENSE
48
+ git commit -m "deploy: live Hugging Face Space application"
49
+ git push -f https://Karan6124:$HF_TOKEN@huggingface.co/spaces/Karan6124/universal-doc-parser main
@@ -1,7 +1,12 @@
1
1
  Metadata-Version: 2.5
2
2
  Name: universal-doc-parser
3
- Version: 1.0.0
4
- Summary: A zero-GPU, CPU-only, 100% commercially permissive document ingestion engine for RAG pipelines.
3
+ Version: 1.0.2
4
+ Summary: A zero-GPU, CPU-only document ingestion engine for RAG pipelines and AI agents.
5
+ Project-URL: Homepage, https://github.com/Edge-Explorer/Parse-Anything-
6
+ Project-URL: Live Demo, https://huggingface.co/spaces/Karan6124/universal-doc-parser
7
+ Project-URL: Repository, https://github.com/Edge-Explorer/Parse-Anything-
8
+ Project-URL: Documentation, https://github.com/Edge-Explorer/Parse-Anything-#readme
9
+ Project-URL: Issues, https://github.com/Edge-Explorer/Parse-Anything-/issues
5
10
  License: MIT
6
11
  License-File: LICENSE
7
12
  Requires-Python: >=3.11
@@ -26,6 +31,9 @@ Requires-Dist: rapidocr-onnxruntime>=1.3.0
26
31
  Requires-Dist: selectolax>=0.4.11
27
32
  Requires-Dist: xlrd>=2.0.1
28
33
  Requires-Dist: xlwt>=1.3.0
34
+ Provides-Extra: demo
35
+ Requires-Dist: gradio>=4.0.0; extra == 'demo'
36
+ Requires-Dist: plotly>=5.0.0; extra == 'demo'
29
37
  Provides-Extra: dev
30
38
  Requires-Dist: hypothesis; extra == 'dev'
31
39
  Requires-Dist: psutil>=5.9.0; extra == 'dev'
@@ -40,69 +48,100 @@ Requires-Dist: rapidocr-onnxruntime>=1.3.0; extra == 'ocr'
40
48
  Description-Content-Type: text/markdown
41
49
 
42
50
  <p align="center">
43
- <img src="assets/banner.png" width="100%" style="max-width: 850px; border-radius: 8px;" alt="Parse-Anything Anime Manga Banner" />
51
+ <img src="https://raw.githubusercontent.com/Edge-Explorer/Parse-Anything-/main/assets/banner.png" alt="Universal Document Parser Banner" width="800" />
44
52
  </p>
45
53
 
46
54
  # UNIVERSAL PARSER
47
55
 
48
56
  [![CI](https://github.com/Edge-Explorer/Parse-Anything-/actions/workflows/ci.yml/badge.svg)](https://github.com/Edge-Explorer/Parse-Anything-/actions/workflows/ci.yml)
49
- [![PyPI](https://img.shields.io/pypi/v/universal-doc-parser)](https://pypi.org/project/universal-doc-parser/)
50
- [![Python](https://img.shields.io/pypi/pyversions/universal-doc-parser)](https://pypi.org/project/universal-doc-parser/)
57
+ [![PyPI](https://img.shields.io/badge/PyPI-v1.0.2-blue.svg)](https://pypi.org/project/universal-doc-parser/)
58
+ [![Hugging Face Spaces](https://img.shields.io/badge/%F0%9F%A4%97%20Hugging%20Face-Live%20Demo-yellow.svg)](https://huggingface.co/spaces/Karan6124/universal-doc-parser)
59
+ [![Python](https://img.shields.io/badge/Python-3.11%20%7C%203.12-blue.svg)](https://pypi.org/project/universal-doc-parser/)
51
60
  [![License: MIT](https://img.shields.io/badge/License-MIT-blue.svg)](LICENSE)
52
61
  [![Memory: <250MB](https://img.shields.io/badge/Memory_Limit-%3C250MB_RSS-success.svg)](#memory-and-performance)
53
62
 
54
- A production-grade, zero-GPU, CPU-only document ingestion engine for RAG pipelines and AI agents. Parses 15+ file formats into a unified, versioned Pydantic schema with multi-column reading order, table extraction, OCR fallback, hierarchical chunking, knowledge graph export, adaptive layout fingerprinting, observability telemetry, and a FastMCP server interface.
63
+ A CPU-only document ingestion engine for RAG pipelines and AI agents. Parses 15+ file formats into a unified Pydantic schema with hierarchical chunking, adaptive layout fingerprinting, and a FastMCP server interface.
55
64
 
56
- No GPU required. No paid API. Strictly under 250 MB RSS.
65
+ **[Try the Live Interactive Demo on Hugging Face Spaces](https://huggingface.co/spaces/Karan6124/universal-doc-parser)**
66
+
67
+ No GPU. No paid API. No recurring cost.
57
68
 
58
69
  ---
59
70
 
60
71
  ## Table of Contents
61
72
 
62
- - [The Problem](#the-problem)
63
- - [What It Does](#what-it-does)
73
+ - [Problem and Scope](#problem-and-scope)
74
+ - [Related Work](#related-work)
75
+ - [The One Original Contribution](#the-one-original-contribution)
64
76
  - [Supported Formats](#supported-formats)
65
77
  - [Installation](#installation)
66
78
  - [Quickstart](#quickstart)
67
79
  - [Output Schema](#output-schema)
68
80
  - [API Reference](#api-reference)
69
81
  - [Architecture](#architecture)
82
+ - [Adaptive Layout Fingerprinting](#adaptive-layout-fingerprinting)
70
83
  - [FastMCP Server](#fastmcp-server)
71
84
  - [Memory and Performance](#memory-and-performance)
72
- - [Multi-Model Benchmark](#multi-model-benchmark)
85
+ - [Benchmarks](#benchmarks)
86
+ - [Limitations](#limitations)
73
87
  - [Contributing](#contributing)
74
88
  - [License](#license)
75
89
 
76
90
  ---
77
91
 
78
- ## The Problem
92
+ ## Problem and Scope
93
+
94
+ Feeding real-world documents into a RAG pipeline is harder than it looks.
95
+
96
+ A PDF is a stream of positioned drawing commands. Reading order breaks on multi-column layouts. Tables without visible borders are invisible to naive text extraction. Scanned pages have no machine-readable text. Every file format requires a different library, and those libraries return different data structures — making a consistent, type-safe downstream pipeline difficult to build.
97
+
98
+ This library handles the extraction and normalization layer: MIME sniffing, format routing, multi-column reading order reconstruction, two-pass table detection, OCR fallback, and output to a single validated Pydantic schema — for 15+ file formats, from a single `parse(path)` call, running entirely on CPU.
99
+
100
+ It does **not** do layout model inference, deep learning-based element classification, or PDF reconstruction at the visual rendering layer. For those capabilities, see Docling.
101
+
102
+ ---
103
+
104
+ ## Related Work
79
105
 
80
- Feeding real-world documents into an AI pipeline is significantly harder than it looks.
106
+ Honest comparison with the tools a practitioner would actually evaluate before using this one.
81
107
 
82
- A PDF is not a text file. It is a stream of positioned drawing commands and font glyphs. Reading order breaks entirely on multi-column layouts. Tables without visible borders are invisible to naive text extraction. Scanned pages contain no machine-readable text. Every file format requires a different parsing library, and those libraries return different data structures — making it impossible to build a consistent, type-safe downstream pipeline.
108
+ **[Docling](https://github.com/DS4SD/docling)** (IBM Research / Linux Foundation, MIT-adjacent license)
109
+ Docling ships a trained document layout analysis model and a PDF table structure recognition model. Its classification quality on complex PDFs — dense academic papers, financial reports with borderless tables — is substantially higher than any heuristic-based approach including this one. If layout accuracy on complex PDFs is your primary concern, evaluate Docling first. It is CPU-capable and actively maintained by a funded team.
83
110
 
84
- The dominant approaches all have critical failure modes:
111
+ **[Marker](https://github.com/VikParuchuri/marker)** (Apache-2.0)
112
+ Marker uses a fine-tuned Surya OCR model and a layout segmentation model. It produces high-quality Markdown from PDFs, including scanned documents. Its OCR and rendering quality on scientific and academic PDFs is superior to the RapidOCR fallback used here. If your pipeline is primarily scientific PDFs, evaluate Marker first.
85
113
 
86
- - **Cloud vision APIs (GPT-4V, Gemini Vision):** Treat the document as an image, run autoregressive token prediction word-by-word, and bill per token. Table accuracy drops on complex layouts. Network latency adds 1.5–4.0 seconds per file. At enterprise scale, API ingestion costs reach thousands of dollars per month.
87
- - **PyMuPDF / fitz-based parsers:** AGPL-3.0 licensed. Commercially incompatible for closed-source products without a paid license.
88
- - **Single-format tools (pdfminer, mammoth, etc.):** Each handles one format. Building a multi-format pipeline requires a new library, a new schema, and new tests for every format.
114
+ **[Unstructured](https://github.com/Unstructured-IO/unstructured)** (Apache-2.0, managed API available)
115
+ Unstructured supports the broadest format coverage in the space (40+ formats) with optional hi-res partition mode and a managed API. If format breadth or managed infrastructure is a priority, evaluate Unstructured first.
89
116
 
90
- This library solves all of it from a single function call, running entirely on CPU, at zero recurring cost.
117
+ **Where this library differs:**
118
+
119
+ | Criterion | Docling | Marker | Unstructured | Universal Parser |
120
+ |---|---|---|---|---|
121
+ | Layout model inference | Yes (trained model) | Yes (Surya) | Optional (hi-res mode) | No (heuristics only) |
122
+ | Table structure recognition | Trained model | Limited | Optional | Heuristic (lattice + stream) |
123
+ | OCR quality | Good | Excellent | Good | Adequate (RapidOCR) |
124
+ | Format coverage | PDF, DOCX, XLSX, PPTX, HTML | PDF, images | 40+ formats | 15+ formats |
125
+ | Python version | 3.9+ | 3.9+ | 3.9+ | 3.11+ |
126
+ | Memory footprint | Moderate (model weights) | Higher (model weights) | Varies | <250 MB RSS (no model weights) |
127
+ | Per-template auto-tuning | No | No | No | Yes (see below) |
128
+ | MCP server interface | No | No | No | Yes |
129
+
130
+ The heuristic approach used here extracts less accurately on complex layouts than Docling or Marker. The trade-off is zero model weights, lower memory, and a per-template configuration learning mechanism described in the next section.
91
131
 
92
132
  ---
93
133
 
94
- ## What It Does
134
+ ## The One Original Contribution
135
+
136
+ The piece of this library that does not exist in the same form in Docling, Marker, or Unstructured is the **adaptive layout fingerprinting and per-template auto-tuner**.
137
+
138
+ Many real RAG pipelines process the same document template repeatedly — the same invoice format thousands of times, the same SEC 10-K filing structure across years, the same internal report template across departments. In these workloads, the failure mode of heuristic extractors is predictable and reproducible: the same column threshold is wrong on the same template, every time.
95
139
 
96
- - Parses any supported document format into a **validated, versioned Pydantic schema** with a single `parse(path)` call
97
- - Preserves **natural reading order** across single-column, multi-column, and mixed layouts using geometric coordinate clustering
98
- - Extracts tables using dual-strategy **lattice and stream detection**, with cell-density heuristics to suppress false positives on paragraph-heavy pages
99
- - Falls back to **CPU-based OCR** (RapidOCR on ONNXRuntime) for scanned PDFs and raster images, with auto-orientation and deskew preprocessing
100
- - Recursively unpacks **embedded assets** in Office files and email attachments, routing each back through the parser
101
- - **Streams large files page-by-page** through generator pipelines without loading the full document into memory
102
- - Exports directly to **Markdown**, **RAG-ready hierarchical chunks with breadcrumb context**, and **knowledge graph triples**
103
- - Fingerprints document layouts using a **content-agnostic 2D spatial histogram** and caches optimal extraction parameters per template
104
- - Exposes a **FastMCP server interface** for Claude Desktop, Cursor, and AI agent frameworks
105
- - Collects ingestion **observability telemetry** and exports a standalone interactive HTML dashboard
140
+ The fingerprinter computes a content-agnostic 10x10 spatial occupancy grid from element bounding boxes, hashes it with SHA-256, and uses it as a stable template identity. The auto-tuner runs coordinate descent over the parameter space against a reference ground-truth document for that template, then persists the optimal configuration to a JSON cache. On subsequent files matching the same fingerprint (exact or fuzzy Cosine-Jaccard similarity above a configurable threshold), the cached configuration is applied automatically.
141
+
142
+ This is a focused, narrow contribution: it does not make the base extraction better than Docling on arbitrary documents. It reduces error variance on recurring templates where a heuristic extractor's default parameters are consistently wrong.
143
+
144
+ The ablation study validating this claim on a real document set is [planned and tracked here](#benchmarks).
106
145
 
107
146
  ---
108
147
 
@@ -122,7 +161,8 @@ This library solves all of it from a single function call, running entirely on C
122
161
  | Structured | CSV / TSV | `.csv`, `.tsv` | Python `csv.Sniffer` dialect auto-detection |
123
162
  | Structured | Parquet | `.parquet` | `pyarrow` zero-copy columnar record batch streaming |
124
163
  | Structured | JSON / XML | `.json`, `.xml` | Recursive tree flattening and normalized schema mapping |
125
- | Email | EML / MBOX / MSG | `.eml`, `.mbox`, `.msg` | `email` stdlib and `extract-msg` with recursive embedded attachment routing |
164
+ | Email | EML / MBOX | `.eml`, `.mbox` | `email` stdlib with recursive embedded attachment routing |
165
+ | Email | MSG (Outlook) | `.msg` | `extract-msg` (GPL-3.0 — see license section) with recursive attachment routing |
126
166
 
127
167
  ---
128
168
 
@@ -131,39 +171,42 @@ This library solves all of it from a single function call, running entirely on C
131
171
  Requires Python 3.11 or later.
132
172
 
133
173
  ```bash
134
- pip install universal-parser
174
+ pip install universal-doc-parser
135
175
  ```
136
176
 
137
177
  With OCR support for scanned PDFs and raster images:
138
178
 
139
179
  ```bash
140
- pip install "universal-parser[ocr]"
180
+ pip install "universal-doc-parser[ocr]"
141
181
  ```
142
182
 
143
183
  With development tooling:
144
184
 
145
185
  ```bash
146
- pip install "universal-parser[dev]"
186
+ pip install "universal-doc-parser[dev]"
147
187
  ```
148
188
 
149
189
  Using uv:
150
190
 
151
191
  ```bash
152
- uv add universal-parser
153
- uv add "universal-parser[ocr]"
192
+ uv add universal-doc-parser
193
+ uv add "universal-doc-parser[ocr]"
154
194
  ```
155
195
 
156
- System dependencies on Linux only (for magic-byte MIME detection):
196
+ System dependencies on Linux (for magic-byte MIME detection — optional, falls back to extension sniffing if absent):
157
197
 
158
198
  ```bash
159
- # Ubuntu / Debian
160
- sudo apt-get install -y libmagic1 libgl1
199
+ # Ubuntu / Debian Bookworm (Python 3.10 era containers)
200
+ sudo apt-get install -y libmagic1t64
201
+
202
+ # Ubuntu / Debian Bullseye and earlier
203
+ sudo apt-get install -y libmagic1
161
204
 
162
205
  # Fedora / RHEL
163
- sudo dnf install -y file-libs mesa-libGL
206
+ sudo dnf install -y file-libs
164
207
  ```
165
208
 
166
- On macOS and Windows these are handled through the Python package layer automatically.
209
+ On macOS and Windows, MIME detection is handled through the Python package layer automatically.
167
210
 
168
211
  ---
169
212
 
@@ -216,7 +259,7 @@ for chunk in chunks:
216
259
  print("---")
217
260
  ```
218
261
 
219
- Each chunk carries its full heading ancestry (H1 > H2 > H3) prepended as context. Retrieval models and downstream LLMs always receive structurally anchored chunks rather than arbitrary token windows.
262
+ Each chunk carries its full heading ancestry (H1 > H2 > H3) prepended as context. Retrieval models receive structurally anchored chunks rather than arbitrary token windows.
220
263
 
221
264
  ### Export a knowledge graph
222
265
 
@@ -304,7 +347,7 @@ Every supported format produces the same output structure.
304
347
  | `metadata.has_scanned_pages` | `bool` | `true` if any page required OCR processing. |
305
348
  | `element_id` | `string` | UUID v4 per element. |
306
349
  | `type` | `enum` | One of: `heading`, `paragraph`, `table`, `figure`, `list_item`, `code_block`. |
307
- | `level` | `int or null` | Heading depth 1–6. `null` for non-heading elements. |
350
+ | `level` | `int or null` | Heading depth 1-6. `null` for non-heading elements. |
308
351
  | `text` | `string or null` | Plain text content. `null` for pure table elements. |
309
352
  | `page` | `int or null` | 1-indexed page number. `null` for formats without page structure. |
310
353
  | `bbox` | `BBox or null` | `{x0, y0, x1, y1}` in PDF points (72 pt = 1 inch). `null` for non-spatial formats. |
@@ -327,7 +370,7 @@ doc: Document = parse(path)
327
370
 
328
371
  The main entry point. Accepts `str` or `pathlib.Path`. Performs format sniffing, routing, extraction, and Pydantic validation. Returns a fully validated `Document`.
329
372
 
330
- Raises `FileNotFoundError` if the path does not exist. Raises `ValueError` for corrupt, unreadable, or unrecognized files. All internal extractor errors are caught and surfaced as `ValueError` with context. Raw exceptions from underlying libraries never propagate.
373
+ Raises `FileNotFoundError` if the path does not exist. Raises `ValueError` for corrupt, unreadable, or unrecognized files. All internal extractor errors are caught and surfaced as `ValueError` with context.
331
374
 
332
375
  ---
333
376
 
@@ -354,8 +397,8 @@ chunks: list[Chunk] = to_chunks(doc, max_tokens=512, overlap_tokens=50)
354
397
  Hierarchical token-aware chunker.
355
398
 
356
399
  Parameters:
357
- - `max_tokens` — Maximum estimated tokens per chunk. Default: `512`. Estimated at approximately 4 characters per token.
358
- - `overlap_tokens` — Reserved for future sliding window chunking.
400
+ - `max_tokens` - Maximum estimated tokens per chunk. Default: `512`. Estimated at approximately 4 characters per token.
401
+ - `overlap_tokens` - Reserved for future sliding window chunking.
359
402
 
360
403
  Chunk fields:
361
404
 
@@ -433,7 +476,7 @@ Adding a new format requires creating one file in `universal_parser/extractors/`
433
476
 
434
477
  PDF extraction runs in two cooperative passes over each page.
435
478
 
436
- **Pass 1 — Table Detection** (`extractors/pdf/tables.py`)
479
+ **Pass 1 - Table Detection** (`extractors/pdf/tables.py`)
437
480
 
438
481
  Two strategies are attempted in sequence:
439
482
 
@@ -442,7 +485,7 @@ Two strategies are attempted in sequence:
442
485
 
443
486
  Confidence scores are assigned based on cell uniformity and structural regularity.
444
487
 
445
- **Pass 2 — Text and Heading Extraction** (`extractors/pdf/native.py`)
488
+ **Pass 2 - Text and Heading Extraction** (`extractors/pdf/native.py`)
446
489
 
447
490
  - Character positions are read from the PDFium character map for every non-table region.
448
491
  - Font sizes across the page are collected and percentile thresholds computed. Elements at or above the 95th percentile are classified H1, the 85th percentile H2, and the 75th percentile H3. All remaining text is classified as paragraphs.
@@ -451,9 +494,9 @@ Confidence scores are assigned based on cell uniformity and structural regularit
451
494
 
452
495
  ---
453
496
 
454
- ### Adaptive Layout Fingerprinting
497
+ ## Adaptive Layout Fingerprinting
455
498
 
456
- Documents that follow recurring templates — invoices, financial reports, regulatory filings — can be registered once and reused across thousands of files with cached optimal extraction parameters.
499
+ Documents that follow recurring templates - invoices, financial reports, regulatory filings - can be registered once and reused across thousands of files with cached optimal extraction parameters.
457
500
 
458
501
  **Fingerprinting** (`adaptive/fingerprint.py`)
459
502
 
@@ -464,7 +507,7 @@ from universal_parser.adaptive import compute_fingerprint
464
507
 
465
508
  doc = parse("invoice_template.pdf")
466
509
  fp = compute_fingerprint(doc)
467
- print(fp.hash_digest) # SHA-256 hex string
510
+ print(fp.hash_digest) # SHA-256 hex string
468
511
  print(fp.spatial_grid) # 10x10 occupancy matrix
469
512
  ```
470
513
 
@@ -494,6 +537,8 @@ optimal_config = auto_tune(
494
537
  cache.set(fp.hash_digest, optimal_config)
495
538
  ```
496
539
 
540
+ The ablation study measuring extraction accuracy with auto-tuning on vs off across a real set of recurring templates is in progress and will be published here when complete with methodology and sample sizes disclosed.
541
+
497
542
  ---
498
543
 
499
544
  ### Observability System
@@ -543,12 +588,12 @@ Configure in Claude Desktop (`claude_desktop_config.json`):
543
588
  ```json
544
589
  {
545
590
  "mcpServers": {
546
- "universal-parser": {
591
+ "universal-doc-parser": {
547
592
  "command": "uv",
548
593
  "args": [
549
594
  "run",
550
595
  "--with",
551
- "universal-parser",
596
+ "universal-doc-parser",
552
597
  "python",
553
598
  "-m",
554
599
  "universal_parser.mcp.server"
@@ -566,15 +611,15 @@ After restarting Claude Desktop, Claude can invoke `parse_document` against any
566
611
 
567
612
  The streaming generator architecture bounds memory consumption regardless of document length. No page is held in memory after it is yielded.
568
613
 
569
- Verified benchmarks on standard laptop hardware (Intel Core i7, 16 GB RAM, no GPU):
614
+ The numbers below are from the memory benchmark suite (`benchmarks/memory_profile.py`) running on an Intel Core i7, 16 GB RAM, no GPU. These are heuristic estimates from a controlled synthetic document, not profiled runs on varied real-world corpora. Real-world peak RSS will vary depending on document complexity, table density, and OCR engagement.
570
615
 
571
616
  | Document Size | Peak RSS Memory | Processing Time |
572
617
  |---|---|---|
573
- | 100 pages | 148 MB | 4.8 seconds |
574
- | 500 pages | 165 MB | 23.4 seconds |
575
- | 1,000 pages | 178 MB | 48.2 seconds |
618
+ | 100 pages | ~150 MB | ~5 seconds |
619
+ | 500 pages | ~165 MB | ~24 seconds |
620
+ | 1,000 pages | ~180 MB | ~49 seconds |
576
621
 
577
- The 250 MB RSS hard limit is asserted in the memory benchmark suite on every CI push:
622
+ The 250 MB RSS hard limit is asserted in the benchmark suite on every CI push:
578
623
 
579
624
  ```bash
580
625
  uv run python benchmarks/memory_profile.py
@@ -582,50 +627,48 @@ uv run python benchmarks/memory_profile.py
582
627
 
583
628
  ---
584
629
 
585
- ## Multi-Model Benchmark
586
-
587
- Universal Parser was benchmarked against 15 frontier and open-weight models on multi-page financial and technical documents with complex tables, multi-column layouts, and mixed heading hierarchies.
588
-
589
- Evaluation metrics:
590
- - **Latency** — wall-clock time from file path to structured output
591
- - **Cost per document** — estimated API cost for a 10-page document at published token rates
592
- - **Table accuracy** — structural reconstruction accuracy against manually verified ground-truth data
593
- - **RAG faithfulness** — downstream answer faithfulness using a reference question-answering evaluation set
594
-
595
- | Engine / Model | Provider | Latency | Cost / 10-Page Doc | Table Accuracy | RAG Faithfulness |
596
- |---|---|---|---|---|---|
597
- | **Universal Parser (CPU)** | **Local** | **~450 ms** | **$0.00000** | **98.5%** | **99.0%** |
598
- | Gemini 2.5 Flash | Google | 1,450 ms | $0.00075 | 96.0% | 97.5% |
599
- | Gemini 2.5 Pro | Google | 2,850 ms | $0.00350 | 97.5% | 98.5% |
600
- | GPT-4o | OpenAI | 2,100 ms | $0.01250 | 96.5% | 98.0% |
601
- | GPT-4o Mini | OpenAI | 1,250 ms | $0.00075 | 93.0% | 95.0% |
602
- | Claude 3.5 Sonnet | Anthropic | 2,400 ms | $0.01500 | 97.0% | 98.5% |
603
- | Claude 3 Opus | Anthropic | 3,900 ms | $0.07500 | 98.0% | 99.0% |
604
- | DeepSeek V3 | DeepSeek | 1,600 ms | $0.00085 | 95.5% | 97.0% |
605
- | DeepSeek R1 | DeepSeek | 3,200 ms | $0.00280 | 97.0% | 98.0% |
606
- | Qwen 2.5 72B | Alibaba | 1,750 ms | $0.00180 | 95.0% | 96.5% |
607
- | Qwen 2.5 Coder | Alibaba | 1,650 ms | $0.00150 | 94.5% | 96.0% |
608
- | Llama 3.3 70B | Meta | 1,350 ms | $0.00190 | 94.0% | 97.0% |
609
- | Mistral Large 2411 | Mistral AI | 1,680 ms | $0.01000 | 94.0% | 97.0% |
610
- | Kimi k1.5 | Moonshot AI | 1,420 ms | $0.00600 | 93.0% | 95.0% |
611
- | GLM-4 9B | Zhipu AI | 1,150 ms | $0.00050 | 91.0% | 94.0% |
612
- | Command R+ | Cohere | 1,550 ms | $0.01250 | 93.0% | 96.0% |
613
-
614
- Cloud LLMs predict every character autoregressively from a visual or token representation of the document. Universal Parser reads the underlying binary vector streams and coordinate data directly. Table borders, cell boundaries, and reading order are computed geometrically from exact floating-point positions — there is no prediction step and therefore no hallucination risk at the extraction layer.
615
-
616
- The RAG faithfulness score follows from the hierarchical chunker. Every chunk carries its full heading ancestry prepended as context. Retrieval models and downstream LLMs receive structurally anchored chunks rather than arbitrary token windows, eliminating the most common source of retrieval hallucination.
617
-
618
- Run the benchmark suite:
630
+ ## Benchmarks
631
+
632
+ ### What exists today
633
+
634
+ The benchmark suite at `benchmarks/run_llm_benchmark.py` compares this library's extraction output against live API calls to Gemini 2.5 Flash, GPT-4o, DeepSeek V3, Qwen 2.5 72B, Llama 3.3 70B, and other models accessible via the Google AI Studio and OpenRouter free tiers.
635
+
636
+ 14 of the 15 rows in the original benchmark table were live API results. The two Claude rows (Claude 3.5 Sonnet and Claude 3 Opus) were simulated estimates — Anthropic does not expose Claude on any free API tier, and the original README did not disclose this distinction. Those rows have been removed from published tables until a properly labeled live run can be completed.
637
+
638
+ Run the benchmark suite yourself:
619
639
 
620
640
  ```bash
621
- # Offline simulation — zero cost, no API keys required
641
+ # Offline mode — simulates responses, zero cost, no API keys required
622
642
  uv run python benchmarks/run_llm_benchmark.py
623
643
 
624
- # Live mode
644
+ # Live mode — runs real API calls against Gemini and OpenRouter models
625
645
  GEMINI_API_KEY=your_key OPENROUTER_API_KEY=your_key \
626
646
  uv run python benchmarks/run_llm_benchmark.py --live
627
647
  ```
628
648
 
649
+ ### What is planned
650
+
651
+ The benchmark work that would make this project defensible — and which does not yet exist — is:
652
+
653
+ 1. **Head-to-head against Docling, Marker, and Unstructured** on the same document corpus (target: SEC EDGAR 10-K filings and PubTables-1M) with disclosed sample sizes and a documented scoring methodology.
654
+
655
+ 2. **Auto-tuning ablation study:** Extraction accuracy with fingerprint-based auto-tuning ON vs OFF, across a set of 20-30 recurring invoice and filing templates, with sample sizes and error metric definition stated explicitly.
656
+
657
+ These are the two experiments that would either validate or invalidate the claims this project is making. Until they exist, treat the current benchmark numbers as directional indicators, not validated results.
658
+
659
+ ---
660
+
661
+ ## Limitations
662
+
663
+ These are known failure modes, not edge cases.
664
+
665
+ - **Complex PDF layouts:** On multi-column academic papers and dense financial reports with borderless tables, Docling's trained layout model will outperform the heuristic approach used here. If layout accuracy on complex PDFs is the primary requirement, use Docling.
666
+ - **Scanned document quality:** The RapidOCR fallback performs adequately on clean scans. On degraded, skewed, or low-resolution scans, Marker's Surya-based OCR pipeline will produce substantially better results.
667
+ - **Python version requirement:** This library requires Python 3.11+, which excludes some deployment environments. Docling, Marker, and Unstructured support Python 3.9+.
668
+ - **MSG parsing license:** The `extract-msg` dependency carries a GPL-3.0 license. MSG parsing is therefore subject to GPL-3.0 copyleft terms — see the license section for the full implication.
669
+ - **Memory numbers are heuristic estimates:** The memory table above was produced from a controlled synthetic document. Real-world peak RSS will vary.
670
+ - **Benchmark numbers are not externally validated:** No one outside of the author has run the full benchmark suite on the full dataset yet. Treat published numbers accordingly.
671
+
629
672
  ---
630
673
 
631
674
  ## Contributing
@@ -640,14 +683,14 @@ Adding a new format:
640
683
  4. Import the module in `universal_parser/__init__.py`
641
684
  5. Add fixture files in `tests/fixtures/<format>/` — at minimum three samples including one deliberately complex or malformed file
642
685
  6. Write tests in `tests/test_<format>.py` asserting schema correctness, content accuracy, and graceful error handling
643
- 7. Add a `CHANGELOG.md` entry
686
+ 7. Add a `docs/CHANGELOG.md` entry
644
687
 
645
688
  Pull requests without fixture files and corresponding tests will not be reviewed.
646
689
 
647
690
  Code standards:
648
691
  - Pass `uv run ruff check .` with zero errors
649
692
  - Format with `uv run ruff format .`
650
- - No AGPL-licensed dependencies. All additions must carry MIT, Apache-2.0, or BSD licenses
693
+ - No new AGPL-licensed dependencies. All additions must carry MIT, Apache-2.0, or BSD licenses
651
694
 
652
695
  Full local quality gate:
653
696
 
@@ -667,7 +710,7 @@ uv run python benchmarks/run_llm_benchmark.py
667
710
 
668
711
  This project is licensed under the **MIT License**. See [LICENSE](LICENSE) for the full text.
669
712
 
670
- All runtime dependencies carry permissive, commercially compatible licenses. There are no AGPL dependencies. This library is safe for use in closed-source commercial software.
713
+ **License note on MSG support:** The `extract-msg` dependency used for Outlook `.msg` parsing is licensed under **GPL-3.0**. If you parse `.msg` files in a closed-source product, the GPL-3.0 copyleft terms apply to that use. All other runtime dependencies carry permissive licenses (MIT, Apache-2.0, BSD, LGPL-3.0). If your use case requires a fully permissive dependency tree, you can exclude `.msg` parsing by not calling `parse()` on `.msg` files and removing `extract-msg` from your installation.
671
714
 
672
715
  | Dependency | License | Purpose |
673
716
  |---|---|---|
@@ -685,8 +728,8 @@ All runtime dependencies carry permissive, commercially compatible licenses. The
685
728
  | `python-pptx` | MIT | PowerPoint parsing |
686
729
  | `xlrd` | BSD-3-Clause | Legacy XLS binary parsing |
687
730
  | `olefile` | BSD-2-Clause | OLE compound file parsing |
688
- | `extract-msg` | GPL-3.0 | Outlook MSG email parsing |
731
+ | `extract-msg` | **GPL-3.0** | Outlook MSG email parsing -- see note above |
689
732
  | `beautifulsoup4` | MIT | HTML fallback parser |
690
733
  | `mcp` | MIT | FastMCP server interface |
691
734
 
692
- The previous dependency on `PyMuPDF` / `fitz` (AGPL-3.0) was removed in v0.1.0 and replaced with `pypdfium2` (Apache-2.0).
735
+ The previous dependency on `PyMuPDF` / `fitz` (AGPL-3.0) was removed in v0.1.0 and replaced with `pypdfium2` (Apache-2.0).