doc-redaction 2.2.2__tar.gz → 2.2.4__tar.gz

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (112) hide show
  1. doc_redaction-2.2.4/MANIFEST.in +4 -0
  2. {doc_redaction-2.2.2/doc_redaction.egg-info → doc_redaction-2.2.4}/PKG-INFO +90 -101
  3. {doc_redaction-2.2.2 → doc_redaction-2.2.4}/README.md +86 -97
  4. {doc_redaction-2.2.2 → doc_redaction-2.2.4}/README_PYPI.md +87 -98
  5. {doc_redaction-2.2.2 → doc_redaction-2.2.4}/app.py +207 -17
  6. doc_redaction-2.2.4/doc_redaction/example_data/Bold minimalist professional cover letter.docx +0 -0
  7. doc_redaction-2.2.4/doc_redaction/example_data/Difficult handwritten note.jpg +0 -0
  8. doc_redaction-2.2.4/doc_redaction/example_data/Example-cv-university-graduaty-hr-role-with-photo-2.pdf +0 -0
  9. doc_redaction-2.2.4/doc_redaction/example_data/Lambeth_2030-Our_Future_Our_Lambeth.pdf.csv +295 -0
  10. doc_redaction-2.2.4/doc_redaction/example_data/Partnership-Agreement-Toolkit_0_0.pdf +0 -0
  11. doc_redaction-2.2.4/doc_redaction/example_data/Partnership-Agreement-Toolkit_test_deny_list_para_single_spell.csv +2 -0
  12. doc_redaction-2.2.4/doc_redaction/example_data/combined_case_notes.csv +19 -0
  13. doc_redaction-2.2.4/doc_redaction/example_data/combined_case_notes.xlsx +0 -0
  14. doc_redaction-2.2.4/doc_redaction/example_data/doubled_output_joined.pdf +0 -0
  15. doc_redaction-2.2.4/doc_redaction/example_data/example_complaint_letter.jpg +0 -0
  16. doc_redaction-2.2.4/doc_redaction/example_data/example_of_emails_sent_to_a_professor_before_applying.pdf +0 -0
  17. doc_redaction-2.2.4/doc_redaction/example_data/example_outputs/Partnership-Agreement-Toolkit_0_0.pdf_ocr_output.csv +277 -0
  18. doc_redaction-2.2.4/doc_redaction/example_data/example_outputs/Partnership-Agreement-Toolkit_0_0.pdf_review_file.csv +77 -0
  19. doc_redaction-2.2.4/doc_redaction/example_data/example_outputs/Partnership-Agreement-Toolkit_0_0_ocr_results_with_words_textract.csv +2438 -0
  20. doc_redaction-2.2.4/doc_redaction/example_data/example_outputs/doubled_output_joined.pdf_ocr_output.csv +923 -0
  21. doc_redaction-2.2.4/doc_redaction/example_data/example_outputs/example_of_emails_sent_to_a_professor_before_applying_ocr_output_textract.csv +40 -0
  22. doc_redaction-2.2.4/doc_redaction/example_data/example_outputs/example_of_emails_sent_to_a_professor_before_applying_ocr_results_with_words_textract.csv +432 -0
  23. doc_redaction-2.2.4/doc_redaction/example_data/example_outputs/example_of_emails_sent_to_a_professor_before_applying_review_file.csv +15 -0
  24. doc_redaction-2.2.4/doc_redaction/example_data/graduate-job-example-cover-letter.pdf +0 -0
  25. doc_redaction-2.2.4/doc_redaction/example_data/partnership_toolkit_redact_custom_deny_list.csv +2 -0
  26. doc_redaction-2.2.4/doc_redaction/example_data/partnership_toolkit_redact_some_pages.csv +2 -0
  27. doc_redaction-2.2.4/doc_redaction/example_data/test_allow_list_graduate.csv +1 -0
  28. doc_redaction-2.2.4/doc_redaction/example_data/test_allow_list_partnership.csv +1 -0
  29. {doc_redaction-2.2.2 → doc_redaction-2.2.4/doc_redaction.egg-info}/PKG-INFO +90 -101
  30. {doc_redaction-2.2.2 → doc_redaction-2.2.4}/doc_redaction.egg-info/SOURCES.txt +25 -1
  31. {doc_redaction-2.2.2 → doc_redaction-2.2.4}/doc_redaction.egg-info/requires.txt +2 -2
  32. {doc_redaction-2.2.2 → doc_redaction-2.2.4}/pyproject.toml +10 -5
  33. {doc_redaction-2.2.2 → doc_redaction-2.2.4}/test/test_package_api_smoke.py +9 -9
  34. {doc_redaction-2.2.2 → doc_redaction-2.2.4}/tools/config.py +27 -7
  35. {doc_redaction-2.2.2 → doc_redaction-2.2.4}/tools/custom_image_analyser_engine.py +31 -4
  36. {doc_redaction-2.2.2 → doc_redaction-2.2.4}/tools/file_redaction.py +23 -1
  37. doc_redaction-2.2.4/tools/preview_redaction_boxes.py +323 -0
  38. {doc_redaction-2.2.2 → doc_redaction-2.2.4}/tools/simplified_api.py +99 -0
  39. doc_redaction-2.2.2/MANIFEST.in +0 -3
  40. {doc_redaction-2.2.2 → doc_redaction-2.2.4}/agent_routes.py +0 -0
  41. {doc_redaction-2.2.2 → doc_redaction-2.2.4}/cli_redact.py +0 -0
  42. {doc_redaction-2.2.2 → doc_redaction-2.2.4}/doc_redaction/__init__.py +0 -0
  43. {doc_redaction-2.2.2 → doc_redaction-2.2.4}/doc_redaction/api.py +0 -0
  44. {doc_redaction-2.2.2 → doc_redaction-2.2.4/doc_redaction/assets}/favicon.png +0 -0
  45. {doc_redaction-2.2.2 → doc_redaction-2.2.4}/doc_redaction/cli_api.py +0 -0
  46. {doc_redaction-2.2.2 → doc_redaction-2.2.4}/doc_redaction/cli_redact.py +0 -0
  47. {doc_redaction-2.2.2 → doc_redaction-2.2.4}/doc_redaction/data_anonymise.py +0 -0
  48. {doc_redaction-2.2.2 → doc_redaction-2.2.4}/doc_redaction/file_conversion.py +0 -0
  49. {doc_redaction-2.2.2 → doc_redaction-2.2.4}/doc_redaction/file_redaction.py +0 -0
  50. {doc_redaction-2.2.2 → doc_redaction-2.2.4}/doc_redaction/find_duplicate_pages.py +0 -0
  51. {doc_redaction-2.2.2 → doc_redaction-2.2.4}/doc_redaction/find_duplicate_tabular.py +0 -0
  52. {doc_redaction-2.2.2 → doc_redaction-2.2.4}/doc_redaction/gradio_app.py +0 -0
  53. {doc_redaction-2.2.2 → doc_redaction-2.2.4}/doc_redaction/helper_functions.py +0 -0
  54. {doc_redaction-2.2.2 → doc_redaction-2.2.4}/doc_redaction/install_deps.py +0 -0
  55. {doc_redaction-2.2.2 → doc_redaction-2.2.4}/doc_redaction/lambda_entrypoint.py +0 -0
  56. {doc_redaction-2.2.2 → doc_redaction-2.2.4}/doc_redaction/redaction_review.py +0 -0
  57. {doc_redaction-2.2.2 → doc_redaction-2.2.4}/doc_redaction/summaries.py +0 -0
  58. {doc_redaction-2.2.2 → doc_redaction-2.2.4}/doc_redaction.egg-info/dependency_links.txt +0 -0
  59. {doc_redaction-2.2.2 → doc_redaction-2.2.4}/doc_redaction.egg-info/entry_points.txt +0 -0
  60. {doc_redaction-2.2.2 → doc_redaction-2.2.4}/doc_redaction.egg-info/top_level.txt +0 -0
  61. {doc_redaction-2.2.2 → doc_redaction-2.2.4}/intros/long_intro.txt +0 -0
  62. {doc_redaction-2.2.2 → doc_redaction-2.2.4}/intros/short_intro.txt +0 -0
  63. {doc_redaction-2.2.2 → doc_redaction-2.2.4}/intros/short_intro_responsible.txt +0 -0
  64. {doc_redaction-2.2.2 → doc_redaction-2.2.4}/lambda_entrypoint.py +0 -0
  65. {doc_redaction-2.2.2 → doc_redaction-2.2.4}/load_dynamo_logs.py +0 -0
  66. {doc_redaction-2.2.2 → doc_redaction-2.2.4}/load_s3_logs.py +0 -0
  67. {doc_redaction-2.2.2 → doc_redaction-2.2.4}/mcp_doc_redaction/__init__.py +0 -0
  68. {doc_redaction-2.2.2 → doc_redaction-2.2.4}/mcp_doc_redaction/artifact_bundle.py +0 -0
  69. {doc_redaction-2.2.2 → doc_redaction-2.2.4}/mcp_doc_redaction/gradio_transport.py +0 -0
  70. {doc_redaction-2.2.2 → doc_redaction-2.2.4}/mcp_doc_redaction/schemas.py +0 -0
  71. {doc_redaction-2.2.2 → doc_redaction-2.2.4}/mcp_doc_redaction/server.py +0 -0
  72. {doc_redaction-2.2.2 → doc_redaction-2.2.4}/setup.cfg +0 -0
  73. {doc_redaction-2.2.2 → doc_redaction-2.2.4}/test/test_agent_apply_review_redactions.py +0 -0
  74. {doc_redaction-2.2.2 → doc_redaction-2.2.4}/test/test_annotation_color_parsing.py +0 -0
  75. {doc_redaction-2.2.2 → doc_redaction-2.2.4}/test/test_cli_smoke.py +0 -0
  76. {doc_redaction-2.2.2 → doc_redaction-2.2.4}/test/test_gradio_doc_redact_simple.py +0 -0
  77. {doc_redaction-2.2.2 → doc_redaction-2.2.4}/test/test_gradio_summarise_simple.py +0 -0
  78. {doc_redaction-2.2.2 → doc_redaction-2.2.4}/test/test_gradio_transport_sse.py +0 -0
  79. {doc_redaction-2.2.2 → doc_redaction-2.2.4}/test/test_gradio_upload_staging.py +0 -0
  80. {doc_redaction-2.2.2 → doc_redaction-2.2.4}/test/test_gui_only.py +0 -0
  81. {doc_redaction-2.2.2 → doc_redaction-2.2.4}/test/test_mcp_doc_redaction_bundle.py +0 -0
  82. {doc_redaction-2.2.2 → doc_redaction-2.2.4}/test/test_mcp_doc_redaction_extract_paths.py +0 -0
  83. {doc_redaction-2.2.2 → doc_redaction-2.2.4}/test/test_placeholder_bbox_scaling.py +0 -0
  84. {doc_redaction-2.2.2 → doc_redaction-2.2.4}/test/test_redaction_overlay_export.py +0 -0
  85. {doc_redaction-2.2.2 → doc_redaction-2.2.4}/test/test_redaction_types.py +0 -0
  86. {doc_redaction-2.2.2 → doc_redaction-2.2.4}/test/test_review_ocr_visualisation_export.py +0 -0
  87. {doc_redaction-2.2.2 → doc_redaction-2.2.4}/tools/__init__.py +0 -0
  88. {doc_redaction-2.2.2 → doc_redaction-2.2.4}/tools/apply_hf_zero_gpu_readme_frontmatter.py +0 -0
  89. {doc_redaction-2.2.2 → doc_redaction-2.2.4}/tools/auth.py +0 -0
  90. {doc_redaction-2.2.2 → doc_redaction-2.2.4}/tools/aws_functions.py +0 -0
  91. {doc_redaction-2.2.2 → doc_redaction-2.2.4}/tools/aws_textract.py +0 -0
  92. {doc_redaction-2.2.2 → doc_redaction-2.2.4}/tools/cli_usage_logger.py +0 -0
  93. {doc_redaction-2.2.2 → doc_redaction-2.2.4}/tools/custom_csvlogger.py +0 -0
  94. {doc_redaction-2.2.2 → doc_redaction-2.2.4}/tools/data_anonymise.py +0 -0
  95. {doc_redaction-2.2.2 → doc_redaction-2.2.4}/tools/file_conversion.py +0 -0
  96. {doc_redaction-2.2.2 → doc_redaction-2.2.4}/tools/find_duplicate_pages.py +0 -0
  97. {doc_redaction-2.2.2 → doc_redaction-2.2.4}/tools/find_duplicate_tabular.py +0 -0
  98. {doc_redaction-2.2.2 → doc_redaction-2.2.4}/tools/helper_functions.py +0 -0
  99. {doc_redaction-2.2.2 → doc_redaction-2.2.4}/tools/llm_entity_detection.py +0 -0
  100. {doc_redaction-2.2.2 → doc_redaction-2.2.4}/tools/llm_entity_detection_prompts.py +0 -0
  101. {doc_redaction-2.2.2 → doc_redaction-2.2.4}/tools/llm_funcs.py +0 -0
  102. {doc_redaction-2.2.2 → doc_redaction-2.2.4}/tools/load_spacy_model_custom_recognisers.py +0 -0
  103. {doc_redaction-2.2.2 → doc_redaction-2.2.4}/tools/presidio_analyzer_custom.py +0 -0
  104. {doc_redaction-2.2.2 → doc_redaction-2.2.4}/tools/quickstart.py +0 -0
  105. {doc_redaction-2.2.2 → doc_redaction-2.2.4}/tools/redaction_review.py +0 -0
  106. {doc_redaction-2.2.2 → doc_redaction-2.2.4}/tools/redaction_types.py +0 -0
  107. {doc_redaction-2.2.2 → doc_redaction-2.2.4}/tools/run_vlm.py +0 -0
  108. {doc_redaction-2.2.2 → doc_redaction-2.2.4}/tools/secure_path_utils.py +0 -0
  109. {doc_redaction-2.2.2 → doc_redaction-2.2.4}/tools/secure_regex_utils.py +0 -0
  110. {doc_redaction-2.2.2 → doc_redaction-2.2.4}/tools/summaries.py +0 -0
  111. {doc_redaction-2.2.2 → doc_redaction-2.2.4}/tools/textract_batch_call.py +0 -0
  112. {doc_redaction-2.2.2 → doc_redaction-2.2.4}/tools/word_segmenter.py +0 -0
@@ -0,0 +1,4 @@
1
+ recursive-include doc_redaction/assets *.png
2
+ recursive-include doc_redaction/example_data *
3
+ recursive-include intros *.txt
4
+
@@ -1,6 +1,6 @@
1
1
  Metadata-Version: 2.4
2
2
  Name: doc_redaction
3
- Version: 2.2.2
3
+ Version: 2.2.4
4
4
  Summary: Redact PDF/image-based documents, Word, or CSV/XLSX files using a Gradio-based GUI interface
5
5
  Author-email: Sean Pedrick-Case <spedrickcase@lambeth.gov.uk>
6
6
  Maintainer-email: Sean Pedrick-Case <spedrickcase@lambeth.gov.uk>
@@ -32,7 +32,7 @@ Requires-Dist: pikepdf<=10.3.0
32
32
  Requires-Dist: pandas<=2.3.3
33
33
  Requires-Dist: scikit-learn<=1.8.0
34
34
  Requires-Dist: spacy<=3.8.14
35
- Requires-Dist: gradio<=6.10.0
35
+ Requires-Dist: gradio<=6.10.0,>=6.9.0
36
36
  Requires-Dist: boto3<=1.42.91
37
37
  Requires-Dist: pyarrow<=23.0.1
38
38
  Requires-Dist: openpyxl<=3.1.5
@@ -70,7 +70,7 @@ Requires-Dist: accelerate<=1.13.0; extra == "vlm"
70
70
  Requires-Dist: bitsandbytes<=0.49.2; extra == "vlm"
71
71
  Requires-Dist: sentencepiece<=0.2.1; extra == "vlm"
72
72
  Provides-Extra: mcp
73
- Requires-Dist: gradio[mcp]<=6.10.0; extra == "mcp"
73
+ Requires-Dist: gradio[mcp]<=6.10.0,>=6.9.0; extra == "mcp"
74
74
 
75
75
  # Document redaction (doc_redaction)
76
76
 
@@ -84,81 +84,39 @@ Redact personally identifiable information (PII) from documents (PDF, PNG, JPG),
84
84
 
85
85
  Follow these instructions to get the document redaction application running on your local machine.
86
86
 
87
- ### 1. Prerequisites: System Dependencies
87
+ ### 1. Package installation
88
88
 
89
- This application relies on two external tools for OCR (Tesseract) and PDF processing (Poppler). Please install them on your system before proceeding.
90
-
91
- ---
92
-
93
- #### Automated dependency setup (recommended)
94
-
95
- If you **don’t have admin rights** (or you just want the simplest setup), you can have the project download and configure **Tesseract** and **Poppler** into a local `redaction_deps/` folder inside the doc_redaction folder.
96
-
97
- You need the installer script available first, which means either:
98
-
99
- - **Repository checkout**: `git clone ...` and run the command from the repo root (recommended for the web UI), or
100
- - **PyPI install**: `pip install doc_redaction` and run from a writable folder where you want `redaction_deps/` and `config/app_config.env` to be created/updated.
89
+ #### Install from source repo (recommended for full features)
101
90
 
102
- From the repository root (or your chosen working folder) after creating/activating your venv and installing Python requirements:
91
+ Clone the repository and install in editable mode:
103
92
 
104
93
  ```bash
105
- python -m doc_redaction.install_deps
94
+ git clone https://github.com/seanpedrick-case/doc_redaction.git
95
+ cd doc_redaction
96
+ pip install -e .
106
97
  ```
107
98
 
108
- This writes `TESSERACT_FOLDER` / `POPPLER_FOLDER` into `config/app_config.env` so the app can find the binaries without you editing your system PATH.
109
-
110
- To just check whether your machine can already see the tools:
99
+ ##### Full install from source (Paddle and VLM)
111
100
 
112
101
  ```bash
113
- python -m doc_redaction.install_deps --verify-only
102
+ pip install -e ".[paddle,vlm]"
114
103
  ```
115
104
 
116
-
117
- #### **On Windows**
118
-
119
- If you don’t use the automated setup above, you can install the dependencies manually by downloading installers and adding the programs to your system's PATH.
120
-
121
- 1. **Install Tesseract OCR:**
122
- * Download the installer from the official Tesseract at [UB Mannheim page](https://github.com/UB-Mannheim/tesseract/wiki) (e.g., `tesseract-ocr-w64-setup-v5.X.X...exe`).
123
- * Run the installer.
124
- * **IMPORTANT:** During installation, ensure you select the option to "Add Tesseract to system PATH for all users" or a similar option. This is crucial for the application to find the Tesseract executable.
125
-
126
-
127
- 2. **Install Poppler:**
128
- * Download the latest Poppler binary for Windows. A common source is the [Poppler for Windows](https://github.com/oschwartz10612/poppler-windows) GitHub releases page. Download the `.zip` file (e.g., `poppler-25.07.0-win.zip`).
129
- * Extract the contents of the zip file to a permanent location on your computer, for example, `C:\Program Files\poppler\`.
130
- * You must add the `bin` folder from your Poppler installation to your system's PATH environment variable.
131
- * Search for "Edit the system environment variables" in the Windows Start Menu and open it.
132
- * Click the "Environment Variables..." button.
133
- * In the "System variables" section, find and select the `Path` variable, then click "Edit...".
134
- * Click "New" and add the full path to the `bin` directory inside your Poppler folder (e.g., `C:\Program Files\poppler\poppler-24.02.0\bin`).
135
- * Click OK on all windows to save the changes.
136
-
137
- To verify, open a new Command Prompt and run `tesseract --version` and `pdftoppm -v`. If they both return version information, you have successfully installed the prerequisites.
138
-
139
- ---
140
-
141
- #### **On Linux (Debian/Ubuntu)**
142
-
143
- Open your terminal and run the following command to install Tesseract and Poppler:
144
-
105
+ Note that the versions of both PaddleOCR and Torch installed by default are the CPU-only versions. If you want to install the equivalent GPU versions, you will need to run the following commands:
145
106
  ```bash
146
- sudo apt-get update && sudo apt-get install -y tesseract-ocr poppler-utils
107
+ pip install paddlepaddle-gpu==3.2.1 --index-url https://www.paddlepaddle.org.cn/packages/stable/cu129/
147
108
  ```
148
109
 
149
- #### **On Linux (Fedora/CentOS/RHEL)**
150
-
151
- Open your terminal and use the `dnf` or `yum` package manager:
110
+ **Note:** It is difficult to get paddlepaddle gpu working in an environment alongside torch. You may well need to reinstall the cpu version to ensure compatibility, and run paddlepaddle-gpu in a separate environment without torch installed. If you get errors related to .dll files following paddle gpu install, you may need to install the latest c++ redistributables. For Windows, you can find them [here](https://learn.microsoft.com/en-us/cpp/windows/latest-supported-vc-redist?view=msvc-170)
152
111
 
153
112
  ```bash
154
- sudo dnf install -y tesseract poppler-utils
113
+ pip install torch==2.8.0 --index-url https://download.pytorch.org/whl/cu129
114
+ pip install torchvision --index-url https://download.pytorch.org/whl/cu129
155
115
  ```
156
- ---
157
-
158
116
 
159
- ### 2. Installation: Python packages
117
+ #### Install from PyPI (recommended for users and library use)
160
118
 
161
- Once the system prerequisites are installed, create a virtual environment (recommended) and install **doc_redaction**.
119
+ Create a virtual environment (recommended) and install **doc_redaction**.
162
120
 
163
121
  ```bash
164
122
  python -m venv venv
@@ -168,8 +126,6 @@ python -m venv venv
168
126
  source venv/bin/activate
169
127
  ```
170
128
 
171
- #### Install from PyPI (recommended for users and library use)
172
-
173
129
  The package is published on PyPI as **`doc-redaction`** (import name **`doc_redaction`**):
174
130
 
175
131
  ```bash
@@ -182,7 +138,7 @@ Optional extras (same as in `pyproject.toml`):
182
138
  pip install "doc_redaction[paddle,vlm]"
183
139
  ```
184
140
 
185
- For programmatic use (CLI-first API matching Gradio `api_name` routes), see **[Package API usage (Python)](https://seanpedrick-case.github.io/doc_redaction/src/package_api_usage.html)**. The console script **`cli_redact`** is available after install.
141
+ For programmatic use (CLI-first API matching Gradio `api_name` routes), see **[Python Package usage (Python)](https://seanpedrick-case.github.io/doc_redaction/src/python_package_usage.html)**. The console script **`cli_redact`** is available after install.
186
142
 
187
143
  **Web UI from a PyPI install:** You *can* start the Gradio UI after `pip install doc_redaction` by running:
188
144
 
@@ -192,71 +148,100 @@ python -m app
192
148
 
193
149
  **Important: your working directory matters.** When you run `python -m app`, the app treats your *current folder* as the “app folder”:
194
150
 
195
- - It will load configuration from `config/app_config.env` *relative to the folder you run it from* (and `python -m doc_redaction.install_deps` will also create/update that file there).
151
+ - It will look for configuration at `config/app_config.env` *relative to the folder you run it from* (and `python -m doc_redaction.install_deps` will also write `config/app_config.env` there).
196
152
  - It may create new folders in that location (for example `config/`, `output/`, `input/`, `logs/`, `usage/`, `feedback/`, and temporary/cache folders depending on your settings).
197
- - The full set of bundled UI example files (`example_data/`) is part of the **Git repository checkout** rather than the PyPI wheel. If you run from a “random” directory after a PyPI install, you should expect the Examples section to be missing unless you provide your own `example_data/` folder.
153
+ - The UI example files and bundled assets are packaged with the PyPI install (they live inside the installed `doc_redaction` package). If you run from a “random” directory after a PyPI install, the app can still locate its packaged examples; your working directory mainly affects where `config/`, `input/`, `output/`, logs, and temp folders are created.
198
154
 
199
155
  In practice, the **smoothest UI experience** (examples, bundled assets, docs links, predictable relative paths) is still usually via a **repository checkout** or **Docker**, but PyPI install is sufficient to launch the UI as long as you run it from a suitable working folder and have the system dependencies available (or run `python -m doc_redaction.install_deps` first).
200
156
 
201
- #### Install from source (repository checkout / development)
157
+ #### Docker installation
202
158
 
203
- Clone the repository and install in editable mode:
159
+ The doc_redaction Redaction app can be installed by using the [Dockerfile](https://github.com/seanpedrick-case/doc_redaction/blob/main/Dockerfile) or Docker compose files ([llama.cpp](https://github.com/ggml-org/llama.cpp), [vLLM](https://docs.vllm.ai/en/stable/)) provided in the repo.
204
160
 
205
- ```bash
206
- git clone https://github.com/seanpedrick-case/doc_redaction.git
207
- cd doc_redaction
208
- pip install -e .
209
- ```
161
+ ##### Without Llama.cpp / vLLM inference server
210
162
 
211
- From the same checkout you can use `requirements_lightweight.txt` instead of editable install if you prefer:
163
+ If you want a working Docker installation without GPU support, you can install from the [Dockerfile](https://github.com/seanpedrick-case/doc_redaction/blob/main/Dockerfile) in the repo. A working example of this, with the CPU version of PaddleOCR, can be found on [Hugging Face](https://huggingface.co/spaces/seanpedrickcase/document_redaction). You can adjust the INSTALL_PADDLEOCR, PADDLE_GPU_ENABLED, INSTALL_VLM, and TORCH_GPU_ENABLED config variables to adjust for PaddleOCR and Transformers packages for local VLM support. Note that GPU-enabled PaddleOCR, and GPU-enabled Transformers/Torch often don't work well together, which is one reason why a Llama.cpp/vLLM inference server Docker installation option is provided below.
212
164
 
213
- ```bash
214
- pip install -r requirements_lightweight.txt
215
- ```
165
+ ##### With Llama.cpp / vLLM inference server
216
166
 
217
- ##### Full install from source (Paddle and VLM)
167
+ The project now has Docker and Docker compose files available to pair running the Redaction app with local inference servers powered by [llama.cpp](https://github.com/ggml-org/llama.cpp), or [vLLM](https://docs.vllm.ai/en/stable/). Llama.cpp is more flexible than vLLM for low VRAM systems, as Llama.cpp will offload to cpu/system RAM automatically rather than failing as vLLM tends to do.
218
168
 
219
- ```bash
220
- pip install -e ".[paddle,vlm]"
221
- ```
169
+ For Llama.cpp, you can use the [docker-compose_llama.yml](https://github.com/seanpedrick-case/doc_redaction/blob/main/docker-compose_llama.yml) file, and for vLLM, you can use the [docker-compose_vllm.yml](https://github.com/seanpedrick-case/doc_redaction/blob/main/docker-compose_vllm.yml) file. To run, Docker / Docker Desktop should be installed, and then you can run the commands suggested in the top of the files to run the servers.
222
170
 
223
- Alternatively, use the full `requirements.txt` (includes PaddleOCR and Torch/transformers references for CUDA 12.9):
171
+ You will need ~40-50GB of disk space to run everything depending on the model chosen from the compose file. For the vLLM server, you will need 24 GB VRAM. For the Llama.cpp server, 24 GB VRAM is needed to run at full speed, but the n-gpu-layers and n-cpu-moe parameters in the Docker compose file can be adjusted to fit into your system. I would suggest that 8 GB VRAM is needed as a bare minimum for decent inference speed. See the [Unsloth guide](https://unsloth.ai/docs/models/qwen3.5) for more details on working with GGUF files for Qwen 3.5.
224
172
 
225
- ```bash
226
- pip install -r requirements.txt
227
- ```
173
+ ### 2. Install prerequisites: Tesseract and Poppler
174
+
175
+ This application relies on two external tools for OCR (Tesseract) and PDF processing (Poppler). Please install them on your system before proceeding.
176
+
177
+ ---
178
+
179
+ #### Automated dependency setup (recommended)
180
+
181
+ If you **don’t have admin rights** (or you just want the simplest setup), you can have the project download and configure **Tesseract** and **Poppler** into a local `redaction_deps/` folder inside the doc_redaction folder.
182
+
183
+ You need the installer script available first, which means either:
184
+
185
+ - **Repository checkout**: `git clone ...` and run the command from the repo root (recommended for the web UI), or
186
+ - **PyPI install**: `pip install doc_redaction` and run from a writable folder where you want `redaction_deps/` and `config/app_config.env` to be created/updated.
187
+
188
+ From the repository root (or your chosen working folder) after creating/activating your venv and installing Python requirements:
228
189
 
229
- Note that the versions of both PaddleOCR and Torch installed by default are the CPU-only versions. If you want to install the equivalent GPU versions, you will need to run the following commands:
230
190
  ```bash
231
- pip install paddlepaddle-gpu==3.2.1 --index-url https://www.paddlepaddle.org.cn/packages/stable/cu129/
191
+ python -m doc_redaction.install_deps
232
192
  ```
233
193
 
234
- **Note:** It is difficult to get paddlepaddle gpu working in an environment alongside torch. You may well need to reinstall the cpu version to ensure compatibility, and run paddlepaddle-gpu in a separate environment without torch installed. If you get errors related to .dll files following paddle gpu install, you may need to install the latest c++ redistributables. For Windows, you can find them [here](https://learn.microsoft.com/en-us/cpp/windows/latest-supported-vc-redist?view=msvc-170)
194
+ This writes `TESSERACT_FOLDER` / `POPPLER_FOLDER` into `config/app_config.env` so the app can find the binaries without you editing your system PATH.
195
+
196
+ To just check whether your machine can already see the tools:
235
197
 
236
198
  ```bash
237
- pip install torch==2.8.0 --index-url https://download.pytorch.org/whl/cu129
238
- pip install torchvision --index-url https://download.pytorch.org/whl/cu129
199
+ python -m doc_redaction.install_deps --verify-only
239
200
  ```
240
201
 
241
- #### Docker installation
202
+ #### **On Windows**
242
203
 
243
- The doc_redaction Redaction app can be installed by using the [Dockerfile](https://github.com/seanpedrick-case/doc_redaction/blob/main/Dockerfile) or Docker compose files ([llama.cpp](https://github.com/ggml-org/llama.cpp), [vLLM](https://docs.vllm.ai/en/stable/)) provided in the repo.
204
+ If you don’t use the automated setup above, you can install the dependencies manually by downloading installers and adding the programs to your system's PATH.
244
205
 
245
- ##### Without Llama.cpp / vLLM inference server
206
+ 1. **Install Tesseract OCR:**
207
+ * Download the installer from the official Tesseract at [UB Mannheim page](https://github.com/UB-Mannheim/tesseract/wiki) (e.g., `tesseract-ocr-w64-setup-v5.X.X...exe`).
208
+ * Run the installer.
209
+ * **IMPORTANT:** During installation, ensure you select the option to "Add Tesseract to system PATH for all users" or a similar option. This is crucial for the application to find the Tesseract executable.
246
210
 
247
- If you want a working Docker installation without GPU support, you can install from the [Dockerfile](https://github.com/seanpedrick-case/doc_redaction/blob/main/Dockerfile) in the repo. A working example of this, with the CPU version of PaddleOCR, can be found on [Hugging Face](https://huggingface.co/spaces/seanpedrickcase/document_redaction). You can adjust the INSTALL_PADDLEOCR, PADDLE_GPU_ENABLED, INSTALL_VLM, and TORCH_GPU_ENABLED config variables to adjust for PaddleOCR and Transformers packages for local VLM support. Note that GPU-enabled PaddleOCR, and GPU-enabled Transformers/Torch often don't work well together, which is one reason why a Llama.cpp/vLLM inference server Docker installation option is provided below.
248
211
 
249
- ##### With Llama.cpp / vLLM inference server
212
+ 2. **Install Poppler:**
213
+ * Download the latest Poppler binary for Windows. A common source is the [Poppler for Windows](https://github.com/oschwartz10612/poppler-windows) GitHub releases page. Download the `.zip` file (e.g., `poppler-25.07.0-win.zip`).
214
+ * Extract the contents of the zip file to a permanent location on your computer, for example, `C:\Program Files\poppler\`.
215
+ * You must add the `bin` folder from your Poppler installation to your system's PATH environment variable.
216
+ * Search for "Edit the system environment variables" in the Windows Start Menu and open it.
217
+ * Click the "Environment Variables..." button.
218
+ * In the "System variables" section, find and select the `Path` variable, then click "Edit...".
219
+ * Click "New" and add the full path to the `bin` directory inside your Poppler folder (e.g., `C:\Program Files\poppler\poppler-24.02.0\bin`).
220
+ * Click OK on all windows to save the changes.
250
221
 
251
- The project now has Docker and Docker compose files available to pair running the Redaction app with local inference servers powered by [llama.cpp](https://github.com/ggml-org/llama.cpp), or [vLLM](https://docs.vllm.ai/en/stable/). Llama.cpp is more flexible than vLLM for low VRAM systems, as Llama.cpp will offload to cpu/system RAM automatically rather than failing as vLLM tends to do.
222
+ To verify, open a new Command Prompt and run `tesseract --version` and `pdftoppm -v`. If they both return version information, you have successfully installed the prerequisites.
223
+ ---
252
224
 
253
- For Llama.cpp, you can use the [docker-compose_llama.yml](https://github.com/seanpedrick-case/doc_redaction/blob/main/docker-compose_llama.yml) file, and for vLLM, you can use the [docker-compose_vllm.yml](https://github.com/seanpedrick-case/doc_redaction/blob/main/docker-compose_vllm.yml) file. To run, Docker / Docker Desktop should be installed, and then you can run the commands suggested in the top of the files to run the servers.
225
+ #### **On Linux (Debian/Ubuntu)**
254
226
 
255
- You will need ~40-50GB of disk space to run everything depending on the model chosen from the compose file. For the vLLM server, you will need 24 GB VRAM. For the Llama.cpp server, 24 GB VRAM is needed to run at full speed, but the n-gpu-layers and n-cpu-moe parameters in the Docker compose file can be adjusted to fit into your system. I would suggest that 8 GB VRAM is needed as a bare minimum for decent inference speed. See the [Unsloth guide](https://unsloth.ai/docs/models/qwen3.5) for more details on working with GGUF files for Qwen 3.5.
227
+ Open your terminal and run the following command to install Tesseract and Poppler:
228
+
229
+ ```bash
230
+ sudo apt-get update && sudo apt-get install -y tesseract-ocr poppler-utils
231
+ ```
232
+
233
+ #### **On Linux (Fedora/CentOS/RHEL)**
234
+
235
+ Open your terminal and use the `dnf` or `yum` package manager:
236
+
237
+ ```bash
238
+ sudo dnf install -y tesseract poppler-utils
239
+ ```
240
+ ---
256
241
 
257
242
  ### 3. Run the Application
258
243
 
259
- With all dependencies installed, you can now start the Gradio application.
244
+ With all dependencies installed, you can now start the Gradio application GUI. For a guide on how to use this, please go [here](https://seanpedrick-case.github.io/doc_redaction/src/user_guide.html).
260
245
 
261
246
  ```bash
262
247
  python app.py
@@ -268,6 +253,8 @@ Open this URL in your web browser to use the document redaction tool
268
253
 
269
254
  #### Command line interface
270
255
 
256
+ For example CLI commands, please refer to [this guide](https://seanpedrick-case.github.io/doc_redaction/src/user_guide.html#command-line-interface-cli) or the examples in [cli_redact.py](https://github.com/seanpedrick-case/doc_redaction/blob/main/cli_redact.py#L321)
257
+
271
258
  If you installed from **PyPI**, use the installed console script:
272
259
 
273
260
  ```bash
@@ -280,7 +267,9 @@ From a **repository checkout**, you can also run:
280
267
  python cli_redact.py --help
281
268
  ```
282
269
 
283
- For Python examples that mirror each Gradio `api_name`, see [Package API usage (Python)](https://seanpedrick-case.github.io/doc_redaction/src/package_api_usage.html) (source: [src/package_api_usage.qmd](src/package_api_usage.qmd)).
270
+ #### Python package commands
271
+
272
+ For Python examples in using the Python package, please see [Python Package usage (Python)](https://seanpedrick-case.github.io/doc_redaction/src/python_package_usage.html).
284
273
 
285
274
  ---
286
275
 
@@ -387,9 +376,9 @@ If those endpoints are not present in your deployment, fall back to the long UI-
387
376
 
388
377
  ### Optional: MCP server
389
378
 
390
- If you want external agents to call this app reliably without re-implementing Gradio upload/call/poll/download details, consider an **MCP server** that wraps the main tasks (`redact_document`, `apply_review_redactions`, `redact_tabular`, `summarise_document`) behind a small tool interface. See [src/agent_mcp.md](src/agent_mcp.md).
379
+ If you want external agents to call this app reliably without re-implementing Gradio upload/call/poll/download details, consider an **MCP server** that wraps the main tasks (`redact_document`, `apply_review_redactions`, `redact_tabular`, `summarise_document`) behind a small tool interface. See the [relevant documentation](https://github.com/seanpedrick-case/doc_redaction/blob/main/mcp_doc_redaction/README.md).
391
380
 
392
- **Use as a library:** After installing from [PyPI](https://pypi.org/project/doc-redaction/) (`pip install doc_redaction`), you can call the same workflows as the Gradio `api_name` routes from Python. See the documentation: [Package API usage (Python)](https://seanpedrick-case.github.io/doc_redaction/src/package_api_usage.html) (source: [src/package_api_usage.qmd](src/package_api_usage.qmd)).
381
+ **Use as a library:** After installing from [PyPI](https://pypi.org/project/doc-redaction/) (`pip install doc_redaction`), you can call the same workflows as the Gradio `api_name` routes from Python. See the documentation: [Python Package usage (Python)](https://seanpedrick-case.github.io/doc_redaction/src/python_package_usage.html).
393
382
 
394
383
  To extract text from documents, the 'Local' options are PikePDF for PDFs with selectable text, and OCR with Tesseract. Use AWS Textract to extract more complex elements e.g. handwriting, signatures, or unclear text. PaddleOCR and VLM support is also provided (see the installation instructions below).
395
384
 
@@ -21,81 +21,39 @@ Redact personally identifiable information (PII) from documents (PDF, PNG, JPG),
21
21
 
22
22
  Follow these instructions to get the document redaction application running on your local machine.
23
23
 
24
- ### 1. Prerequisites: System Dependencies
24
+ ### 1. Package installation
25
25
 
26
- This application relies on two external tools for OCR (Tesseract) and PDF processing (Poppler). Please install them on your system before proceeding.
27
-
28
- ---
29
-
30
- #### Automated dependency setup (recommended)
31
-
32
- If you **don’t have admin rights** (or you just want the simplest setup), you can have the project download and configure **Tesseract** and **Poppler** into a local `redaction_deps/` folder inside the doc_redaction folder.
33
-
34
- You need the installer script available first, which means either:
35
-
36
- - **Repository checkout**: `git clone ...` and run the command from the repo root (recommended for the web UI), or
37
- - **PyPI install**: `pip install doc_redaction` and run from a writable folder where you want `redaction_deps/` and `config/app_config.env` to be created/updated.
26
+ #### Install from source repo (recommended for full features)
38
27
 
39
- From the repository root (or your chosen working folder) after creating/activating your venv and installing Python requirements:
28
+ Clone the repository and install in editable mode:
40
29
 
41
30
  ```bash
42
- python -m doc_redaction.install_deps
31
+ git clone https://github.com/seanpedrick-case/doc_redaction.git
32
+ cd doc_redaction
33
+ pip install -e .
43
34
  ```
44
35
 
45
- This writes `TESSERACT_FOLDER` / `POPPLER_FOLDER` into `config/app_config.env` so the app can find the binaries without you editing your system PATH.
46
-
47
- To just check whether your machine can already see the tools:
36
+ ##### Full install from source (Paddle and VLM)
48
37
 
49
38
  ```bash
50
- python -m doc_redaction.install_deps --verify-only
39
+ pip install -e ".[paddle,vlm]"
51
40
  ```
52
41
 
53
-
54
- #### **On Windows**
55
-
56
- If you don’t use the automated setup above, you can install the dependencies manually by downloading installers and adding the programs to your system's PATH.
57
-
58
- 1. **Install Tesseract OCR:**
59
- * Download the installer from the official Tesseract at [UB Mannheim page](https://github.com/UB-Mannheim/tesseract/wiki) (e.g., `tesseract-ocr-w64-setup-v5.X.X...exe`).
60
- * Run the installer.
61
- * **IMPORTANT:** During installation, ensure you select the option to "Add Tesseract to system PATH for all users" or a similar option. This is crucial for the application to find the Tesseract executable.
62
-
63
-
64
- 2. **Install Poppler:**
65
- * Download the latest Poppler binary for Windows. A common source is the [Poppler for Windows](https://github.com/oschwartz10612/poppler-windows) GitHub releases page. Download the `.zip` file (e.g., `poppler-25.07.0-win.zip`).
66
- * Extract the contents of the zip file to a permanent location on your computer, for example, `C:\Program Files\poppler\`.
67
- * You must add the `bin` folder from your Poppler installation to your system's PATH environment variable.
68
- * Search for "Edit the system environment variables" in the Windows Start Menu and open it.
69
- * Click the "Environment Variables..." button.
70
- * In the "System variables" section, find and select the `Path` variable, then click "Edit...".
71
- * Click "New" and add the full path to the `bin` directory inside your Poppler folder (e.g., `C:\Program Files\poppler\poppler-24.02.0\bin`).
72
- * Click OK on all windows to save the changes.
73
-
74
- To verify, open a new Command Prompt and run `tesseract --version` and `pdftoppm -v`. If they both return version information, you have successfully installed the prerequisites.
75
-
76
- ---
77
-
78
- #### **On Linux (Debian/Ubuntu)**
79
-
80
- Open your terminal and run the following command to install Tesseract and Poppler:
81
-
42
+ Note that the versions of both PaddleOCR and Torch installed by default are the CPU-only versions. If you want to install the equivalent GPU versions, you will need to run the following commands:
82
43
  ```bash
83
- sudo apt-get update && sudo apt-get install -y tesseract-ocr poppler-utils
44
+ pip install paddlepaddle-gpu==3.2.1 --index-url https://www.paddlepaddle.org.cn/packages/stable/cu129/
84
45
  ```
85
46
 
86
- #### **On Linux (Fedora/CentOS/RHEL)**
87
-
88
- Open your terminal and use the `dnf` or `yum` package manager:
47
+ **Note:** It is difficult to get paddlepaddle gpu working in an environment alongside torch. You may well need to reinstall the cpu version to ensure compatibility, and run paddlepaddle-gpu in a separate environment without torch installed. If you get errors related to .dll files following paddle gpu install, you may need to install the latest c++ redistributables. For Windows, you can find them [here](https://learn.microsoft.com/en-us/cpp/windows/latest-supported-vc-redist?view=msvc-170)
89
48
 
90
49
  ```bash
91
- sudo dnf install -y tesseract poppler-utils
50
+ pip install torch==2.8.0 --index-url https://download.pytorch.org/whl/cu129
51
+ pip install torchvision --index-url https://download.pytorch.org/whl/cu129
92
52
  ```
93
- ---
94
-
95
53
 
96
- ### 2. Installation: Python packages
54
+ #### Install from PyPI (recommended for users and library use)
97
55
 
98
- Once the system prerequisites are installed, create a virtual environment (recommended) and install **doc_redaction**.
56
+ Create a virtual environment (recommended) and install **doc_redaction**.
99
57
 
100
58
  ```bash
101
59
  python -m venv venv
@@ -105,8 +63,6 @@ python -m venv venv
105
63
  source venv/bin/activate
106
64
  ```
107
65
 
108
- #### Install from PyPI (recommended for users and library use)
109
-
110
66
  The package is published on PyPI as **`doc-redaction`** (import name **`doc_redaction`**):
111
67
 
112
68
  ```bash
@@ -119,7 +75,7 @@ Optional extras (same as in `pyproject.toml`):
119
75
  pip install "doc_redaction[paddle,vlm]"
120
76
  ```
121
77
 
122
- For programmatic use (CLI-first API matching Gradio `api_name` routes), see **[Package API usage (Python)](https://seanpedrick-case.github.io/doc_redaction/src/package_api_usage.html)**. The console script **`cli_redact`** is available after install.
78
+ For programmatic use (CLI-first API matching Gradio `api_name` routes), see **[Python Package usage (Python)](https://seanpedrick-case.github.io/doc_redaction/src/python_package_usage.html)**. The console script **`cli_redact`** is available after install.
123
79
 
124
80
  **Web UI from a PyPI install:** You *can* start the Gradio UI after `pip install doc_redaction` by running:
125
81
 
@@ -131,69 +87,98 @@ python -m app
131
87
 
132
88
  - It will look for configuration at `config/app_config.env` *relative to the folder you run it from* (and `python -m doc_redaction.install_deps` will also write `config/app_config.env` there).
133
89
  - It may create new folders in that location (for example `config/`, `output/`, `input/`, `logs/`, `usage/`, `feedback/`, and temporary/cache folders depending on your settings).
134
- - The full set of bundled UI example files (`example_data/`) is part of the **Git repository checkout** rather than the PyPI wheel. If you run from a “random” directory after a PyPI install, you should expect the Examples section to be missing unless you provide your own `example_data/` folder.
90
+ - The UI example files and bundled assets are packaged with the PyPI install (they live inside the installed `doc_redaction` package). If you run from a “random” directory after a PyPI install, the app can still locate its packaged examples; your working directory mainly affects where `config/`, `input/`, `output/`, logs, and temp folders are created.
135
91
 
136
92
  In practice, the **smoothest UI experience** (examples, bundled assets, docs links, predictable relative paths) is still usually via a **repository checkout** or **Docker**, but PyPI install is sufficient to launch the UI as long as you run it from a suitable working folder and have the system dependencies available (or run `python -m doc_redaction.install_deps` first).
137
93
 
138
- #### Install from source (repository checkout / development)
94
+ #### Docker installation
139
95
 
140
- Clone the repository and install in editable mode:
96
+ The doc_redaction Redaction app can be installed by using the [Dockerfile](https://github.com/seanpedrick-case/doc_redaction/blob/main/Dockerfile) or Docker compose files ([llama.cpp](https://github.com/ggml-org/llama.cpp), [vLLM](https://docs.vllm.ai/en/stable/)) provided in the repo.
141
97
 
142
- ```bash
143
- git clone https://github.com/seanpedrick-case/doc_redaction.git
144
- cd doc_redaction
145
- pip install -e .
146
- ```
98
+ ##### Without Llama.cpp / vLLM inference server
147
99
 
148
- From the same checkout you can use `requirements_lightweight.txt` instead of editable install if you prefer:
100
+ If you want a working Docker installation without GPU support, you can install from the [Dockerfile](https://github.com/seanpedrick-case/doc_redaction/blob/main/Dockerfile) in the repo. A working example of this, with the CPU version of PaddleOCR, can be found on [Hugging Face](https://huggingface.co/spaces/seanpedrickcase/document_redaction). You can adjust the INSTALL_PADDLEOCR, PADDLE_GPU_ENABLED, INSTALL_VLM, and TORCH_GPU_ENABLED config variables to adjust for PaddleOCR and Transformers packages for local VLM support. Note that GPU-enabled PaddleOCR, and GPU-enabled Transformers/Torch often don't work well together, which is one reason why a Llama.cpp/vLLM inference server Docker installation option is provided below.
149
101
 
150
- ```bash
151
- pip install -r requirements_lightweight.txt
152
- ```
102
+ ##### With Llama.cpp / vLLM inference server
153
103
 
154
- ##### Full install from source (Paddle and VLM)
104
+ The project now has Docker and Docker compose files available to pair running the Redaction app with local inference servers powered by [llama.cpp](https://github.com/ggml-org/llama.cpp), or [vLLM](https://docs.vllm.ai/en/stable/). Llama.cpp is more flexible than vLLM for low VRAM systems, as Llama.cpp will offload to cpu/system RAM automatically rather than failing as vLLM tends to do.
155
105
 
156
- ```bash
157
- pip install -e ".[paddle,vlm]"
158
- ```
106
+ For Llama.cpp, you can use the [docker-compose_llama.yml](https://github.com/seanpedrick-case/doc_redaction/blob/main/docker-compose_llama.yml) file, and for vLLM, you can use the [docker-compose_vllm.yml](https://github.com/seanpedrick-case/doc_redaction/blob/main/docker-compose_vllm.yml) file. To run, Docker / Docker Desktop should be installed, and then you can run the commands suggested in the top of the files to run the servers.
159
107
 
160
- Alternatively, use the full `requirements.txt` (includes PaddleOCR and Torch/transformers references for CUDA 12.9):
108
+ You will need ~40-50GB of disk space to run everything depending on the model chosen from the compose file. For the vLLM server, you will need 24 GB VRAM. For the Llama.cpp server, 24 GB VRAM is needed to run at full speed, but the n-gpu-layers and n-cpu-moe parameters in the Docker compose file can be adjusted to fit into your system. I would suggest that 8 GB VRAM is needed as a bare minimum for decent inference speed. See the [Unsloth guide](https://unsloth.ai/docs/models/qwen3.5) for more details on working with GGUF files for Qwen 3.5.
161
109
 
162
- ```bash
163
- pip install -r requirements.txt
164
- ```
110
+ ### 2. Install prerequisites: Tesseract and Poppler
111
+
112
+ This application relies on two external tools for OCR (Tesseract) and PDF processing (Poppler). Please install them on your system before proceeding.
113
+
114
+ ---
115
+
116
+ #### Automated dependency setup (recommended)
117
+
118
+ If you **don’t have admin rights** (or you just want the simplest setup), you can have the project download and configure **Tesseract** and **Poppler** into a local `redaction_deps/` folder inside the doc_redaction folder.
119
+
120
+ You need the installer script available first, which means either:
121
+
122
+ - **Repository checkout**: `git clone ...` and run the command from the repo root (recommended for the web UI), or
123
+ - **PyPI install**: `pip install doc_redaction` and run from a writable folder where you want `redaction_deps/` and `config/app_config.env` to be created/updated.
124
+
125
+ From the repository root (or your chosen working folder) after creating/activating your venv and installing Python requirements:
165
126
 
166
- Note that the versions of both PaddleOCR and Torch installed by default are the CPU-only versions. If you want to install the equivalent GPU versions, you will need to run the following commands:
167
127
  ```bash
168
- pip install paddlepaddle-gpu==3.2.1 --index-url https://www.paddlepaddle.org.cn/packages/stable/cu129/
128
+ python -m doc_redaction.install_deps
169
129
  ```
170
130
 
171
- **Note:** It is difficult to get paddlepaddle gpu working in an environment alongside torch. You may well need to reinstall the cpu version to ensure compatibility, and run paddlepaddle-gpu in a separate environment without torch installed. If you get errors related to .dll files following paddle gpu install, you may need to install the latest c++ redistributables. For Windows, you can find them [here](https://learn.microsoft.com/en-us/cpp/windows/latest-supported-vc-redist?view=msvc-170)
131
+ This writes `TESSERACT_FOLDER` / `POPPLER_FOLDER` into `config/app_config.env` so the app can find the binaries without you editing your system PATH.
132
+
133
+ To just check whether your machine can already see the tools:
172
134
 
173
135
  ```bash
174
- pip install torch==2.8.0 --index-url https://download.pytorch.org/whl/cu129
175
- pip install torchvision --index-url https://download.pytorch.org/whl/cu129
136
+ python -m doc_redaction.install_deps --verify-only
176
137
  ```
177
138
 
178
- #### Docker installation
139
+ #### **On Windows**
179
140
 
180
- The doc_redaction Redaction app can be installed by using the [Dockerfile](https://github.com/seanpedrick-case/doc_redaction/blob/main/Dockerfile) or Docker compose files ([llama.cpp](https://github.com/ggml-org/llama.cpp), [vLLM](https://docs.vllm.ai/en/stable/)) provided in the repo.
141
+ If you don’t use the automated setup above, you can install the dependencies manually by downloading installers and adding the programs to your system's PATH.
181
142
 
182
- ##### Without Llama.cpp / vLLM inference server
143
+ 1. **Install Tesseract OCR:**
144
+ * Download the installer from the official Tesseract at [UB Mannheim page](https://github.com/UB-Mannheim/tesseract/wiki) (e.g., `tesseract-ocr-w64-setup-v5.X.X...exe`).
145
+ * Run the installer.
146
+ * **IMPORTANT:** During installation, ensure you select the option to "Add Tesseract to system PATH for all users" or a similar option. This is crucial for the application to find the Tesseract executable.
183
147
 
184
- If you want a working Docker installation without GPU support, you can install from the [Dockerfile](https://github.com/seanpedrick-case/doc_redaction/blob/main/Dockerfile) in the repo. A working example of this, with the CPU version of PaddleOCR, can be found on [Hugging Face](https://huggingface.co/spaces/seanpedrickcase/document_redaction). You can adjust the INSTALL_PADDLEOCR, PADDLE_GPU_ENABLED, INSTALL_VLM, and TORCH_GPU_ENABLED config variables to adjust for PaddleOCR and Transformers packages for local VLM support. Note that GPU-enabled PaddleOCR, and GPU-enabled Transformers/Torch often don't work well together, which is one reason why a Llama.cpp/vLLM inference server Docker installation option is provided below.
185
148
 
186
- ##### With Llama.cpp / vLLM inference server
149
+ 2. **Install Poppler:**
150
+ * Download the latest Poppler binary for Windows. A common source is the [Poppler for Windows](https://github.com/oschwartz10612/poppler-windows) GitHub releases page. Download the `.zip` file (e.g., `poppler-25.07.0-win.zip`).
151
+ * Extract the contents of the zip file to a permanent location on your computer, for example, `C:\Program Files\poppler\`.
152
+ * You must add the `bin` folder from your Poppler installation to your system's PATH environment variable.
153
+ * Search for "Edit the system environment variables" in the Windows Start Menu and open it.
154
+ * Click the "Environment Variables..." button.
155
+ * In the "System variables" section, find and select the `Path` variable, then click "Edit...".
156
+ * Click "New" and add the full path to the `bin` directory inside your Poppler folder (e.g., `C:\Program Files\poppler\poppler-24.02.0\bin`).
157
+ * Click OK on all windows to save the changes.
187
158
 
188
- The project now has Docker and Docker compose files available to pair running the Redaction app with local inference servers powered by [llama.cpp](https://github.com/ggml-org/llama.cpp), or [vLLM](https://docs.vllm.ai/en/stable/). Llama.cpp is more flexible than vLLM for low VRAM systems, as Llama.cpp will offload to cpu/system RAM automatically rather than failing as vLLM tends to do.
159
+ To verify, open a new Command Prompt and run `tesseract --version` and `pdftoppm -v`. If they both return version information, you have successfully installed the prerequisites.
160
+ ---
189
161
 
190
- For Llama.cpp, you can use the [docker-compose_llama.yml](https://github.com/seanpedrick-case/doc_redaction/blob/main/docker-compose_llama.yml) file, and for vLLM, you can use the [docker-compose_vllm.yml](https://github.com/seanpedrick-case/doc_redaction/blob/main/docker-compose_vllm.yml) file. To run, Docker / Docker Desktop should be installed, and then you can run the commands suggested in the top of the files to run the servers.
162
+ #### **On Linux (Debian/Ubuntu)**
191
163
 
192
- You will need ~40-50GB of disk space to run everything depending on the model chosen from the compose file. For the vLLM server, you will need 24 GB VRAM. For the Llama.cpp server, 24 GB VRAM is needed to run at full speed, but the n-gpu-layers and n-cpu-moe parameters in the Docker compose file can be adjusted to fit into your system. I would suggest that 8 GB VRAM is needed as a bare minimum for decent inference speed. See the [Unsloth guide](https://unsloth.ai/docs/models/qwen3.5) for more details on working with GGUF files for Qwen 3.5.
164
+ Open your terminal and run the following command to install Tesseract and Poppler:
165
+
166
+ ```bash
167
+ sudo apt-get update && sudo apt-get install -y tesseract-ocr poppler-utils
168
+ ```
169
+
170
+ #### **On Linux (Fedora/CentOS/RHEL)**
171
+
172
+ Open your terminal and use the `dnf` or `yum` package manager:
173
+
174
+ ```bash
175
+ sudo dnf install -y tesseract poppler-utils
176
+ ```
177
+ ---
193
178
 
194
179
  ### 3. Run the Application
195
180
 
196
- With all dependencies installed, you can now start the Gradio application.
181
+ With all dependencies installed, you can now start the Gradio application GUI. For a guide on how to use this, please go [here](https://seanpedrick-case.github.io/doc_redaction/src/user_guide.html).
197
182
 
198
183
  ```bash
199
184
  python app.py
@@ -205,6 +190,8 @@ Open this URL in your web browser to use the document redaction tool
205
190
 
206
191
  #### Command line interface
207
192
 
193
+ For example CLI commands, please refer to [this guide](https://seanpedrick-case.github.io/doc_redaction/src/user_guide.html#command-line-interface-cli) or the examples in [cli_redact.py](https://github.com/seanpedrick-case/doc_redaction/blob/main/cli_redact.py#L321)
194
+
208
195
  If you installed from **PyPI**, use the installed console script:
209
196
 
210
197
  ```bash
@@ -217,7 +204,9 @@ From a **repository checkout**, you can also run:
217
204
  python cli_redact.py --help
218
205
  ```
219
206
 
220
- For Python examples that mirror each Gradio `api_name`, see [Package API usage (Python)](https://seanpedrick-case.github.io/doc_redaction/src/package_api_usage.html) (source: [src/package_api_usage.qmd](src/package_api_usage.qmd)).
207
+ #### Python package commands
208
+
209
+ For Python examples in using the Python package, please see [Python Package usage (Python)](https://seanpedrick-case.github.io/doc_redaction/src/python_package_usage.html).
221
210
 
222
211
  ---
223
212
 
@@ -329,9 +318,9 @@ If those endpoints are not present in your deployment, fall back to the long UI-
329
318
 
330
319
  ### Optional: MCP server
331
320
 
332
- If you want external agents to call this app reliably without re-implementing Gradio upload/call/poll/download details, consider an **MCP server** that wraps the main tasks (`redact_document`, `apply_review_redactions`, `redact_tabular`, `summarise_document`) behind a small tool interface. See [src/agent_mcp.md](src/agent_mcp.md).
321
+ If you want external agents to call this app reliably without re-implementing Gradio upload/call/poll/download details, consider an **MCP server** that wraps the main tasks (`redact_document`, `apply_review_redactions`, `redact_tabular`, `summarise_document`) behind a small tool interface. See the [relevant documentation](https://github.com/seanpedrick-case/doc_redaction/blob/main/mcp_doc_redaction/README.md).
333
322
 
334
- **Use as a library:** After installing from [PyPI](https://pypi.org/project/doc-redaction/) (`pip install doc_redaction`), you can call the same workflows as the Gradio `api_name` routes from Python. See the documentation: [Package API usage (Python)](https://seanpedrick-case.github.io/doc_redaction/src/package_api_usage.html) (source: [src/package_api_usage.qmd](src/package_api_usage.qmd)).
323
+ **Use as a library:** After installing from [PyPI](https://pypi.org/project/doc-redaction/) (`pip install doc_redaction`), you can call the same workflows as the Gradio `api_name` routes from Python. See the documentation: [Python Package usage (Python)](https://seanpedrick-case.github.io/doc_redaction/src/python_package_usage.html).
335
324
 
336
325
  To extract text from documents, the 'Local' options are PikePDF for PDFs with selectable text, and OCR with Tesseract. Use AWS Textract to extract more complex elements e.g. handwriting, signatures, or unclear text. PaddleOCR and VLM support is also provided (see the installation instructions below).
337
326