doc-redaction 2.2.2__tar.gz → 2.2.4__tar.gz
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- doc_redaction-2.2.4/MANIFEST.in +4 -0
- {doc_redaction-2.2.2/doc_redaction.egg-info → doc_redaction-2.2.4}/PKG-INFO +90 -101
- {doc_redaction-2.2.2 → doc_redaction-2.2.4}/README.md +86 -97
- {doc_redaction-2.2.2 → doc_redaction-2.2.4}/README_PYPI.md +87 -98
- {doc_redaction-2.2.2 → doc_redaction-2.2.4}/app.py +207 -17
- doc_redaction-2.2.4/doc_redaction/example_data/Bold minimalist professional cover letter.docx +0 -0
- doc_redaction-2.2.4/doc_redaction/example_data/Difficult handwritten note.jpg +0 -0
- doc_redaction-2.2.4/doc_redaction/example_data/Example-cv-university-graduaty-hr-role-with-photo-2.pdf +0 -0
- doc_redaction-2.2.4/doc_redaction/example_data/Lambeth_2030-Our_Future_Our_Lambeth.pdf.csv +295 -0
- doc_redaction-2.2.4/doc_redaction/example_data/Partnership-Agreement-Toolkit_0_0.pdf +0 -0
- doc_redaction-2.2.4/doc_redaction/example_data/Partnership-Agreement-Toolkit_test_deny_list_para_single_spell.csv +2 -0
- doc_redaction-2.2.4/doc_redaction/example_data/combined_case_notes.csv +19 -0
- doc_redaction-2.2.4/doc_redaction/example_data/combined_case_notes.xlsx +0 -0
- doc_redaction-2.2.4/doc_redaction/example_data/doubled_output_joined.pdf +0 -0
- doc_redaction-2.2.4/doc_redaction/example_data/example_complaint_letter.jpg +0 -0
- doc_redaction-2.2.4/doc_redaction/example_data/example_of_emails_sent_to_a_professor_before_applying.pdf +0 -0
- doc_redaction-2.2.4/doc_redaction/example_data/example_outputs/Partnership-Agreement-Toolkit_0_0.pdf_ocr_output.csv +277 -0
- doc_redaction-2.2.4/doc_redaction/example_data/example_outputs/Partnership-Agreement-Toolkit_0_0.pdf_review_file.csv +77 -0
- doc_redaction-2.2.4/doc_redaction/example_data/example_outputs/Partnership-Agreement-Toolkit_0_0_ocr_results_with_words_textract.csv +2438 -0
- doc_redaction-2.2.4/doc_redaction/example_data/example_outputs/doubled_output_joined.pdf_ocr_output.csv +923 -0
- doc_redaction-2.2.4/doc_redaction/example_data/example_outputs/example_of_emails_sent_to_a_professor_before_applying_ocr_output_textract.csv +40 -0
- doc_redaction-2.2.4/doc_redaction/example_data/example_outputs/example_of_emails_sent_to_a_professor_before_applying_ocr_results_with_words_textract.csv +432 -0
- doc_redaction-2.2.4/doc_redaction/example_data/example_outputs/example_of_emails_sent_to_a_professor_before_applying_review_file.csv +15 -0
- doc_redaction-2.2.4/doc_redaction/example_data/graduate-job-example-cover-letter.pdf +0 -0
- doc_redaction-2.2.4/doc_redaction/example_data/partnership_toolkit_redact_custom_deny_list.csv +2 -0
- doc_redaction-2.2.4/doc_redaction/example_data/partnership_toolkit_redact_some_pages.csv +2 -0
- doc_redaction-2.2.4/doc_redaction/example_data/test_allow_list_graduate.csv +1 -0
- doc_redaction-2.2.4/doc_redaction/example_data/test_allow_list_partnership.csv +1 -0
- {doc_redaction-2.2.2 → doc_redaction-2.2.4/doc_redaction.egg-info}/PKG-INFO +90 -101
- {doc_redaction-2.2.2 → doc_redaction-2.2.4}/doc_redaction.egg-info/SOURCES.txt +25 -1
- {doc_redaction-2.2.2 → doc_redaction-2.2.4}/doc_redaction.egg-info/requires.txt +2 -2
- {doc_redaction-2.2.2 → doc_redaction-2.2.4}/pyproject.toml +10 -5
- {doc_redaction-2.2.2 → doc_redaction-2.2.4}/test/test_package_api_smoke.py +9 -9
- {doc_redaction-2.2.2 → doc_redaction-2.2.4}/tools/config.py +27 -7
- {doc_redaction-2.2.2 → doc_redaction-2.2.4}/tools/custom_image_analyser_engine.py +31 -4
- {doc_redaction-2.2.2 → doc_redaction-2.2.4}/tools/file_redaction.py +23 -1
- doc_redaction-2.2.4/tools/preview_redaction_boxes.py +323 -0
- {doc_redaction-2.2.2 → doc_redaction-2.2.4}/tools/simplified_api.py +99 -0
- doc_redaction-2.2.2/MANIFEST.in +0 -3
- {doc_redaction-2.2.2 → doc_redaction-2.2.4}/agent_routes.py +0 -0
- {doc_redaction-2.2.2 → doc_redaction-2.2.4}/cli_redact.py +0 -0
- {doc_redaction-2.2.2 → doc_redaction-2.2.4}/doc_redaction/__init__.py +0 -0
- {doc_redaction-2.2.2 → doc_redaction-2.2.4}/doc_redaction/api.py +0 -0
- {doc_redaction-2.2.2 → doc_redaction-2.2.4/doc_redaction/assets}/favicon.png +0 -0
- {doc_redaction-2.2.2 → doc_redaction-2.2.4}/doc_redaction/cli_api.py +0 -0
- {doc_redaction-2.2.2 → doc_redaction-2.2.4}/doc_redaction/cli_redact.py +0 -0
- {doc_redaction-2.2.2 → doc_redaction-2.2.4}/doc_redaction/data_anonymise.py +0 -0
- {doc_redaction-2.2.2 → doc_redaction-2.2.4}/doc_redaction/file_conversion.py +0 -0
- {doc_redaction-2.2.2 → doc_redaction-2.2.4}/doc_redaction/file_redaction.py +0 -0
- {doc_redaction-2.2.2 → doc_redaction-2.2.4}/doc_redaction/find_duplicate_pages.py +0 -0
- {doc_redaction-2.2.2 → doc_redaction-2.2.4}/doc_redaction/find_duplicate_tabular.py +0 -0
- {doc_redaction-2.2.2 → doc_redaction-2.2.4}/doc_redaction/gradio_app.py +0 -0
- {doc_redaction-2.2.2 → doc_redaction-2.2.4}/doc_redaction/helper_functions.py +0 -0
- {doc_redaction-2.2.2 → doc_redaction-2.2.4}/doc_redaction/install_deps.py +0 -0
- {doc_redaction-2.2.2 → doc_redaction-2.2.4}/doc_redaction/lambda_entrypoint.py +0 -0
- {doc_redaction-2.2.2 → doc_redaction-2.2.4}/doc_redaction/redaction_review.py +0 -0
- {doc_redaction-2.2.2 → doc_redaction-2.2.4}/doc_redaction/summaries.py +0 -0
- {doc_redaction-2.2.2 → doc_redaction-2.2.4}/doc_redaction.egg-info/dependency_links.txt +0 -0
- {doc_redaction-2.2.2 → doc_redaction-2.2.4}/doc_redaction.egg-info/entry_points.txt +0 -0
- {doc_redaction-2.2.2 → doc_redaction-2.2.4}/doc_redaction.egg-info/top_level.txt +0 -0
- {doc_redaction-2.2.2 → doc_redaction-2.2.4}/intros/long_intro.txt +0 -0
- {doc_redaction-2.2.2 → doc_redaction-2.2.4}/intros/short_intro.txt +0 -0
- {doc_redaction-2.2.2 → doc_redaction-2.2.4}/intros/short_intro_responsible.txt +0 -0
- {doc_redaction-2.2.2 → doc_redaction-2.2.4}/lambda_entrypoint.py +0 -0
- {doc_redaction-2.2.2 → doc_redaction-2.2.4}/load_dynamo_logs.py +0 -0
- {doc_redaction-2.2.2 → doc_redaction-2.2.4}/load_s3_logs.py +0 -0
- {doc_redaction-2.2.2 → doc_redaction-2.2.4}/mcp_doc_redaction/__init__.py +0 -0
- {doc_redaction-2.2.2 → doc_redaction-2.2.4}/mcp_doc_redaction/artifact_bundle.py +0 -0
- {doc_redaction-2.2.2 → doc_redaction-2.2.4}/mcp_doc_redaction/gradio_transport.py +0 -0
- {doc_redaction-2.2.2 → doc_redaction-2.2.4}/mcp_doc_redaction/schemas.py +0 -0
- {doc_redaction-2.2.2 → doc_redaction-2.2.4}/mcp_doc_redaction/server.py +0 -0
- {doc_redaction-2.2.2 → doc_redaction-2.2.4}/setup.cfg +0 -0
- {doc_redaction-2.2.2 → doc_redaction-2.2.4}/test/test_agent_apply_review_redactions.py +0 -0
- {doc_redaction-2.2.2 → doc_redaction-2.2.4}/test/test_annotation_color_parsing.py +0 -0
- {doc_redaction-2.2.2 → doc_redaction-2.2.4}/test/test_cli_smoke.py +0 -0
- {doc_redaction-2.2.2 → doc_redaction-2.2.4}/test/test_gradio_doc_redact_simple.py +0 -0
- {doc_redaction-2.2.2 → doc_redaction-2.2.4}/test/test_gradio_summarise_simple.py +0 -0
- {doc_redaction-2.2.2 → doc_redaction-2.2.4}/test/test_gradio_transport_sse.py +0 -0
- {doc_redaction-2.2.2 → doc_redaction-2.2.4}/test/test_gradio_upload_staging.py +0 -0
- {doc_redaction-2.2.2 → doc_redaction-2.2.4}/test/test_gui_only.py +0 -0
- {doc_redaction-2.2.2 → doc_redaction-2.2.4}/test/test_mcp_doc_redaction_bundle.py +0 -0
- {doc_redaction-2.2.2 → doc_redaction-2.2.4}/test/test_mcp_doc_redaction_extract_paths.py +0 -0
- {doc_redaction-2.2.2 → doc_redaction-2.2.4}/test/test_placeholder_bbox_scaling.py +0 -0
- {doc_redaction-2.2.2 → doc_redaction-2.2.4}/test/test_redaction_overlay_export.py +0 -0
- {doc_redaction-2.2.2 → doc_redaction-2.2.4}/test/test_redaction_types.py +0 -0
- {doc_redaction-2.2.2 → doc_redaction-2.2.4}/test/test_review_ocr_visualisation_export.py +0 -0
- {doc_redaction-2.2.2 → doc_redaction-2.2.4}/tools/__init__.py +0 -0
- {doc_redaction-2.2.2 → doc_redaction-2.2.4}/tools/apply_hf_zero_gpu_readme_frontmatter.py +0 -0
- {doc_redaction-2.2.2 → doc_redaction-2.2.4}/tools/auth.py +0 -0
- {doc_redaction-2.2.2 → doc_redaction-2.2.4}/tools/aws_functions.py +0 -0
- {doc_redaction-2.2.2 → doc_redaction-2.2.4}/tools/aws_textract.py +0 -0
- {doc_redaction-2.2.2 → doc_redaction-2.2.4}/tools/cli_usage_logger.py +0 -0
- {doc_redaction-2.2.2 → doc_redaction-2.2.4}/tools/custom_csvlogger.py +0 -0
- {doc_redaction-2.2.2 → doc_redaction-2.2.4}/tools/data_anonymise.py +0 -0
- {doc_redaction-2.2.2 → doc_redaction-2.2.4}/tools/file_conversion.py +0 -0
- {doc_redaction-2.2.2 → doc_redaction-2.2.4}/tools/find_duplicate_pages.py +0 -0
- {doc_redaction-2.2.2 → doc_redaction-2.2.4}/tools/find_duplicate_tabular.py +0 -0
- {doc_redaction-2.2.2 → doc_redaction-2.2.4}/tools/helper_functions.py +0 -0
- {doc_redaction-2.2.2 → doc_redaction-2.2.4}/tools/llm_entity_detection.py +0 -0
- {doc_redaction-2.2.2 → doc_redaction-2.2.4}/tools/llm_entity_detection_prompts.py +0 -0
- {doc_redaction-2.2.2 → doc_redaction-2.2.4}/tools/llm_funcs.py +0 -0
- {doc_redaction-2.2.2 → doc_redaction-2.2.4}/tools/load_spacy_model_custom_recognisers.py +0 -0
- {doc_redaction-2.2.2 → doc_redaction-2.2.4}/tools/presidio_analyzer_custom.py +0 -0
- {doc_redaction-2.2.2 → doc_redaction-2.2.4}/tools/quickstart.py +0 -0
- {doc_redaction-2.2.2 → doc_redaction-2.2.4}/tools/redaction_review.py +0 -0
- {doc_redaction-2.2.2 → doc_redaction-2.2.4}/tools/redaction_types.py +0 -0
- {doc_redaction-2.2.2 → doc_redaction-2.2.4}/tools/run_vlm.py +0 -0
- {doc_redaction-2.2.2 → doc_redaction-2.2.4}/tools/secure_path_utils.py +0 -0
- {doc_redaction-2.2.2 → doc_redaction-2.2.4}/tools/secure_regex_utils.py +0 -0
- {doc_redaction-2.2.2 → doc_redaction-2.2.4}/tools/summaries.py +0 -0
- {doc_redaction-2.2.2 → doc_redaction-2.2.4}/tools/textract_batch_call.py +0 -0
- {doc_redaction-2.2.2 → doc_redaction-2.2.4}/tools/word_segmenter.py +0 -0
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
Metadata-Version: 2.4
|
|
2
2
|
Name: doc_redaction
|
|
3
|
-
Version: 2.2.
|
|
3
|
+
Version: 2.2.4
|
|
4
4
|
Summary: Redact PDF/image-based documents, Word, or CSV/XLSX files using a Gradio-based GUI interface
|
|
5
5
|
Author-email: Sean Pedrick-Case <spedrickcase@lambeth.gov.uk>
|
|
6
6
|
Maintainer-email: Sean Pedrick-Case <spedrickcase@lambeth.gov.uk>
|
|
@@ -32,7 +32,7 @@ Requires-Dist: pikepdf<=10.3.0
|
|
|
32
32
|
Requires-Dist: pandas<=2.3.3
|
|
33
33
|
Requires-Dist: scikit-learn<=1.8.0
|
|
34
34
|
Requires-Dist: spacy<=3.8.14
|
|
35
|
-
Requires-Dist: gradio<=6.10.0
|
|
35
|
+
Requires-Dist: gradio<=6.10.0,>=6.9.0
|
|
36
36
|
Requires-Dist: boto3<=1.42.91
|
|
37
37
|
Requires-Dist: pyarrow<=23.0.1
|
|
38
38
|
Requires-Dist: openpyxl<=3.1.5
|
|
@@ -70,7 +70,7 @@ Requires-Dist: accelerate<=1.13.0; extra == "vlm"
|
|
|
70
70
|
Requires-Dist: bitsandbytes<=0.49.2; extra == "vlm"
|
|
71
71
|
Requires-Dist: sentencepiece<=0.2.1; extra == "vlm"
|
|
72
72
|
Provides-Extra: mcp
|
|
73
|
-
Requires-Dist: gradio[mcp]<=6.10.0; extra == "mcp"
|
|
73
|
+
Requires-Dist: gradio[mcp]<=6.10.0,>=6.9.0; extra == "mcp"
|
|
74
74
|
|
|
75
75
|
# Document redaction (doc_redaction)
|
|
76
76
|
|
|
@@ -84,81 +84,39 @@ Redact personally identifiable information (PII) from documents (PDF, PNG, JPG),
|
|
|
84
84
|
|
|
85
85
|
Follow these instructions to get the document redaction application running on your local machine.
|
|
86
86
|
|
|
87
|
-
### 1.
|
|
87
|
+
### 1. Package installation
|
|
88
88
|
|
|
89
|
-
|
|
90
|
-
|
|
91
|
-
---
|
|
92
|
-
|
|
93
|
-
#### Automated dependency setup (recommended)
|
|
94
|
-
|
|
95
|
-
If you **don’t have admin rights** (or you just want the simplest setup), you can have the project download and configure **Tesseract** and **Poppler** into a local `redaction_deps/` folder inside the doc_redaction folder.
|
|
96
|
-
|
|
97
|
-
You need the installer script available first, which means either:
|
|
98
|
-
|
|
99
|
-
- **Repository checkout**: `git clone ...` and run the command from the repo root (recommended for the web UI), or
|
|
100
|
-
- **PyPI install**: `pip install doc_redaction` and run from a writable folder where you want `redaction_deps/` and `config/app_config.env` to be created/updated.
|
|
89
|
+
#### Install from source repo (recommended for full features)
|
|
101
90
|
|
|
102
|
-
|
|
91
|
+
Clone the repository and install in editable mode:
|
|
103
92
|
|
|
104
93
|
```bash
|
|
105
|
-
|
|
94
|
+
git clone https://github.com/seanpedrick-case/doc_redaction.git
|
|
95
|
+
cd doc_redaction
|
|
96
|
+
pip install -e .
|
|
106
97
|
```
|
|
107
98
|
|
|
108
|
-
|
|
109
|
-
|
|
110
|
-
To just check whether your machine can already see the tools:
|
|
99
|
+
##### Full install from source (Paddle and VLM)
|
|
111
100
|
|
|
112
101
|
```bash
|
|
113
|
-
|
|
102
|
+
pip install -e ".[paddle,vlm]"
|
|
114
103
|
```
|
|
115
104
|
|
|
116
|
-
|
|
117
|
-
#### **On Windows**
|
|
118
|
-
|
|
119
|
-
If you don’t use the automated setup above, you can install the dependencies manually by downloading installers and adding the programs to your system's PATH.
|
|
120
|
-
|
|
121
|
-
1. **Install Tesseract OCR:**
|
|
122
|
-
* Download the installer from the official Tesseract at [UB Mannheim page](https://github.com/UB-Mannheim/tesseract/wiki) (e.g., `tesseract-ocr-w64-setup-v5.X.X...exe`).
|
|
123
|
-
* Run the installer.
|
|
124
|
-
* **IMPORTANT:** During installation, ensure you select the option to "Add Tesseract to system PATH for all users" or a similar option. This is crucial for the application to find the Tesseract executable.
|
|
125
|
-
|
|
126
|
-
|
|
127
|
-
2. **Install Poppler:**
|
|
128
|
-
* Download the latest Poppler binary for Windows. A common source is the [Poppler for Windows](https://github.com/oschwartz10612/poppler-windows) GitHub releases page. Download the `.zip` file (e.g., `poppler-25.07.0-win.zip`).
|
|
129
|
-
* Extract the contents of the zip file to a permanent location on your computer, for example, `C:\Program Files\poppler\`.
|
|
130
|
-
* You must add the `bin` folder from your Poppler installation to your system's PATH environment variable.
|
|
131
|
-
* Search for "Edit the system environment variables" in the Windows Start Menu and open it.
|
|
132
|
-
* Click the "Environment Variables..." button.
|
|
133
|
-
* In the "System variables" section, find and select the `Path` variable, then click "Edit...".
|
|
134
|
-
* Click "New" and add the full path to the `bin` directory inside your Poppler folder (e.g., `C:\Program Files\poppler\poppler-24.02.0\bin`).
|
|
135
|
-
* Click OK on all windows to save the changes.
|
|
136
|
-
|
|
137
|
-
To verify, open a new Command Prompt and run `tesseract --version` and `pdftoppm -v`. If they both return version information, you have successfully installed the prerequisites.
|
|
138
|
-
|
|
139
|
-
---
|
|
140
|
-
|
|
141
|
-
#### **On Linux (Debian/Ubuntu)**
|
|
142
|
-
|
|
143
|
-
Open your terminal and run the following command to install Tesseract and Poppler:
|
|
144
|
-
|
|
105
|
+
Note that the versions of both PaddleOCR and Torch installed by default are the CPU-only versions. If you want to install the equivalent GPU versions, you will need to run the following commands:
|
|
145
106
|
```bash
|
|
146
|
-
|
|
107
|
+
pip install paddlepaddle-gpu==3.2.1 --index-url https://www.paddlepaddle.org.cn/packages/stable/cu129/
|
|
147
108
|
```
|
|
148
109
|
|
|
149
|
-
|
|
150
|
-
|
|
151
|
-
Open your terminal and use the `dnf` or `yum` package manager:
|
|
110
|
+
**Note:** It is difficult to get paddlepaddle gpu working in an environment alongside torch. You may well need to reinstall the cpu version to ensure compatibility, and run paddlepaddle-gpu in a separate environment without torch installed. If you get errors related to .dll files following paddle gpu install, you may need to install the latest c++ redistributables. For Windows, you can find them [here](https://learn.microsoft.com/en-us/cpp/windows/latest-supported-vc-redist?view=msvc-170)
|
|
152
111
|
|
|
153
112
|
```bash
|
|
154
|
-
|
|
113
|
+
pip install torch==2.8.0 --index-url https://download.pytorch.org/whl/cu129
|
|
114
|
+
pip install torchvision --index-url https://download.pytorch.org/whl/cu129
|
|
155
115
|
```
|
|
156
|
-
---
|
|
157
|
-
|
|
158
116
|
|
|
159
|
-
|
|
117
|
+
#### Install from PyPI (recommended for users and library use)
|
|
160
118
|
|
|
161
|
-
|
|
119
|
+
Create a virtual environment (recommended) and install **doc_redaction**.
|
|
162
120
|
|
|
163
121
|
```bash
|
|
164
122
|
python -m venv venv
|
|
@@ -168,8 +126,6 @@ python -m venv venv
|
|
|
168
126
|
source venv/bin/activate
|
|
169
127
|
```
|
|
170
128
|
|
|
171
|
-
#### Install from PyPI (recommended for users and library use)
|
|
172
|
-
|
|
173
129
|
The package is published on PyPI as **`doc-redaction`** (import name **`doc_redaction`**):
|
|
174
130
|
|
|
175
131
|
```bash
|
|
@@ -182,7 +138,7 @@ Optional extras (same as in `pyproject.toml`):
|
|
|
182
138
|
pip install "doc_redaction[paddle,vlm]"
|
|
183
139
|
```
|
|
184
140
|
|
|
185
|
-
For programmatic use (CLI-first API matching Gradio `api_name` routes), see **[Package
|
|
141
|
+
For programmatic use (CLI-first API matching Gradio `api_name` routes), see **[Python Package usage (Python)](https://seanpedrick-case.github.io/doc_redaction/src/python_package_usage.html)**. The console script **`cli_redact`** is available after install.
|
|
186
142
|
|
|
187
143
|
**Web UI from a PyPI install:** You *can* start the Gradio UI after `pip install doc_redaction` by running:
|
|
188
144
|
|
|
@@ -192,71 +148,100 @@ python -m app
|
|
|
192
148
|
|
|
193
149
|
**Important: your working directory matters.** When you run `python -m app`, the app treats your *current folder* as the “app folder”:
|
|
194
150
|
|
|
195
|
-
- It will
|
|
151
|
+
- It will look for configuration at `config/app_config.env` *relative to the folder you run it from* (and `python -m doc_redaction.install_deps` will also write `config/app_config.env` there).
|
|
196
152
|
- It may create new folders in that location (for example `config/`, `output/`, `input/`, `logs/`, `usage/`, `feedback/`, and temporary/cache folders depending on your settings).
|
|
197
|
-
- The
|
|
153
|
+
- The UI example files and bundled assets are packaged with the PyPI install (they live inside the installed `doc_redaction` package). If you run from a “random” directory after a PyPI install, the app can still locate its packaged examples; your working directory mainly affects where `config/`, `input/`, `output/`, logs, and temp folders are created.
|
|
198
154
|
|
|
199
155
|
In practice, the **smoothest UI experience** (examples, bundled assets, docs links, predictable relative paths) is still usually via a **repository checkout** or **Docker**, but PyPI install is sufficient to launch the UI as long as you run it from a suitable working folder and have the system dependencies available (or run `python -m doc_redaction.install_deps` first).
|
|
200
156
|
|
|
201
|
-
####
|
|
157
|
+
#### Docker installation
|
|
202
158
|
|
|
203
|
-
|
|
159
|
+
The doc_redaction Redaction app can be installed by using the [Dockerfile](https://github.com/seanpedrick-case/doc_redaction/blob/main/Dockerfile) or Docker compose files ([llama.cpp](https://github.com/ggml-org/llama.cpp), [vLLM](https://docs.vllm.ai/en/stable/)) provided in the repo.
|
|
204
160
|
|
|
205
|
-
|
|
206
|
-
git clone https://github.com/seanpedrick-case/doc_redaction.git
|
|
207
|
-
cd doc_redaction
|
|
208
|
-
pip install -e .
|
|
209
|
-
```
|
|
161
|
+
##### Without Llama.cpp / vLLM inference server
|
|
210
162
|
|
|
211
|
-
|
|
163
|
+
If you want a working Docker installation without GPU support, you can install from the [Dockerfile](https://github.com/seanpedrick-case/doc_redaction/blob/main/Dockerfile) in the repo. A working example of this, with the CPU version of PaddleOCR, can be found on [Hugging Face](https://huggingface.co/spaces/seanpedrickcase/document_redaction). You can adjust the INSTALL_PADDLEOCR, PADDLE_GPU_ENABLED, INSTALL_VLM, and TORCH_GPU_ENABLED config variables to adjust for PaddleOCR and Transformers packages for local VLM support. Note that GPU-enabled PaddleOCR, and GPU-enabled Transformers/Torch often don't work well together, which is one reason why a Llama.cpp/vLLM inference server Docker installation option is provided below.
|
|
212
164
|
|
|
213
|
-
|
|
214
|
-
pip install -r requirements_lightweight.txt
|
|
215
|
-
```
|
|
165
|
+
##### With Llama.cpp / vLLM inference server
|
|
216
166
|
|
|
217
|
-
|
|
167
|
+
The project now has Docker and Docker compose files available to pair running the Redaction app with local inference servers powered by [llama.cpp](https://github.com/ggml-org/llama.cpp), or [vLLM](https://docs.vllm.ai/en/stable/). Llama.cpp is more flexible than vLLM for low VRAM systems, as Llama.cpp will offload to cpu/system RAM automatically rather than failing as vLLM tends to do.
|
|
218
168
|
|
|
219
|
-
|
|
220
|
-
pip install -e ".[paddle,vlm]"
|
|
221
|
-
```
|
|
169
|
+
For Llama.cpp, you can use the [docker-compose_llama.yml](https://github.com/seanpedrick-case/doc_redaction/blob/main/docker-compose_llama.yml) file, and for vLLM, you can use the [docker-compose_vllm.yml](https://github.com/seanpedrick-case/doc_redaction/blob/main/docker-compose_vllm.yml) file. To run, Docker / Docker Desktop should be installed, and then you can run the commands suggested in the top of the files to run the servers.
|
|
222
170
|
|
|
223
|
-
|
|
171
|
+
You will need ~40-50GB of disk space to run everything depending on the model chosen from the compose file. For the vLLM server, you will need 24 GB VRAM. For the Llama.cpp server, 24 GB VRAM is needed to run at full speed, but the n-gpu-layers and n-cpu-moe parameters in the Docker compose file can be adjusted to fit into your system. I would suggest that 8 GB VRAM is needed as a bare minimum for decent inference speed. See the [Unsloth guide](https://unsloth.ai/docs/models/qwen3.5) for more details on working with GGUF files for Qwen 3.5.
|
|
224
172
|
|
|
225
|
-
|
|
226
|
-
|
|
227
|
-
|
|
173
|
+
### 2. Install prerequisites: Tesseract and Poppler
|
|
174
|
+
|
|
175
|
+
This application relies on two external tools for OCR (Tesseract) and PDF processing (Poppler). Please install them on your system before proceeding.
|
|
176
|
+
|
|
177
|
+
---
|
|
178
|
+
|
|
179
|
+
#### Automated dependency setup (recommended)
|
|
180
|
+
|
|
181
|
+
If you **don’t have admin rights** (or you just want the simplest setup), you can have the project download and configure **Tesseract** and **Poppler** into a local `redaction_deps/` folder inside the doc_redaction folder.
|
|
182
|
+
|
|
183
|
+
You need the installer script available first, which means either:
|
|
184
|
+
|
|
185
|
+
- **Repository checkout**: `git clone ...` and run the command from the repo root (recommended for the web UI), or
|
|
186
|
+
- **PyPI install**: `pip install doc_redaction` and run from a writable folder where you want `redaction_deps/` and `config/app_config.env` to be created/updated.
|
|
187
|
+
|
|
188
|
+
From the repository root (or your chosen working folder) after creating/activating your venv and installing Python requirements:
|
|
228
189
|
|
|
229
|
-
Note that the versions of both PaddleOCR and Torch installed by default are the CPU-only versions. If you want to install the equivalent GPU versions, you will need to run the following commands:
|
|
230
190
|
```bash
|
|
231
|
-
|
|
191
|
+
python -m doc_redaction.install_deps
|
|
232
192
|
```
|
|
233
193
|
|
|
234
|
-
|
|
194
|
+
This writes `TESSERACT_FOLDER` / `POPPLER_FOLDER` into `config/app_config.env` so the app can find the binaries without you editing your system PATH.
|
|
195
|
+
|
|
196
|
+
To just check whether your machine can already see the tools:
|
|
235
197
|
|
|
236
198
|
```bash
|
|
237
|
-
|
|
238
|
-
pip install torchvision --index-url https://download.pytorch.org/whl/cu129
|
|
199
|
+
python -m doc_redaction.install_deps --verify-only
|
|
239
200
|
```
|
|
240
201
|
|
|
241
|
-
####
|
|
202
|
+
#### **On Windows**
|
|
242
203
|
|
|
243
|
-
|
|
204
|
+
If you don’t use the automated setup above, you can install the dependencies manually by downloading installers and adding the programs to your system's PATH.
|
|
244
205
|
|
|
245
|
-
|
|
206
|
+
1. **Install Tesseract OCR:**
|
|
207
|
+
* Download the installer from the official Tesseract at [UB Mannheim page](https://github.com/UB-Mannheim/tesseract/wiki) (e.g., `tesseract-ocr-w64-setup-v5.X.X...exe`).
|
|
208
|
+
* Run the installer.
|
|
209
|
+
* **IMPORTANT:** During installation, ensure you select the option to "Add Tesseract to system PATH for all users" or a similar option. This is crucial for the application to find the Tesseract executable.
|
|
246
210
|
|
|
247
|
-
If you want a working Docker installation without GPU support, you can install from the [Dockerfile](https://github.com/seanpedrick-case/doc_redaction/blob/main/Dockerfile) in the repo. A working example of this, with the CPU version of PaddleOCR, can be found on [Hugging Face](https://huggingface.co/spaces/seanpedrickcase/document_redaction). You can adjust the INSTALL_PADDLEOCR, PADDLE_GPU_ENABLED, INSTALL_VLM, and TORCH_GPU_ENABLED config variables to adjust for PaddleOCR and Transformers packages for local VLM support. Note that GPU-enabled PaddleOCR, and GPU-enabled Transformers/Torch often don't work well together, which is one reason why a Llama.cpp/vLLM inference server Docker installation option is provided below.
|
|
248
211
|
|
|
249
|
-
|
|
212
|
+
2. **Install Poppler:**
|
|
213
|
+
* Download the latest Poppler binary for Windows. A common source is the [Poppler for Windows](https://github.com/oschwartz10612/poppler-windows) GitHub releases page. Download the `.zip` file (e.g., `poppler-25.07.0-win.zip`).
|
|
214
|
+
* Extract the contents of the zip file to a permanent location on your computer, for example, `C:\Program Files\poppler\`.
|
|
215
|
+
* You must add the `bin` folder from your Poppler installation to your system's PATH environment variable.
|
|
216
|
+
* Search for "Edit the system environment variables" in the Windows Start Menu and open it.
|
|
217
|
+
* Click the "Environment Variables..." button.
|
|
218
|
+
* In the "System variables" section, find and select the `Path` variable, then click "Edit...".
|
|
219
|
+
* Click "New" and add the full path to the `bin` directory inside your Poppler folder (e.g., `C:\Program Files\poppler\poppler-24.02.0\bin`).
|
|
220
|
+
* Click OK on all windows to save the changes.
|
|
250
221
|
|
|
251
|
-
|
|
222
|
+
To verify, open a new Command Prompt and run `tesseract --version` and `pdftoppm -v`. If they both return version information, you have successfully installed the prerequisites.
|
|
223
|
+
---
|
|
252
224
|
|
|
253
|
-
|
|
225
|
+
#### **On Linux (Debian/Ubuntu)**
|
|
254
226
|
|
|
255
|
-
|
|
227
|
+
Open your terminal and run the following command to install Tesseract and Poppler:
|
|
228
|
+
|
|
229
|
+
```bash
|
|
230
|
+
sudo apt-get update && sudo apt-get install -y tesseract-ocr poppler-utils
|
|
231
|
+
```
|
|
232
|
+
|
|
233
|
+
#### **On Linux (Fedora/CentOS/RHEL)**
|
|
234
|
+
|
|
235
|
+
Open your terminal and use the `dnf` or `yum` package manager:
|
|
236
|
+
|
|
237
|
+
```bash
|
|
238
|
+
sudo dnf install -y tesseract poppler-utils
|
|
239
|
+
```
|
|
240
|
+
---
|
|
256
241
|
|
|
257
242
|
### 3. Run the Application
|
|
258
243
|
|
|
259
|
-
With all dependencies installed, you can now start the Gradio application.
|
|
244
|
+
With all dependencies installed, you can now start the Gradio application GUI. For a guide on how to use this, please go [here](https://seanpedrick-case.github.io/doc_redaction/src/user_guide.html).
|
|
260
245
|
|
|
261
246
|
```bash
|
|
262
247
|
python app.py
|
|
@@ -268,6 +253,8 @@ Open this URL in your web browser to use the document redaction tool
|
|
|
268
253
|
|
|
269
254
|
#### Command line interface
|
|
270
255
|
|
|
256
|
+
For example CLI commands, please refer to [this guide](https://seanpedrick-case.github.io/doc_redaction/src/user_guide.html#command-line-interface-cli) or the examples in [cli_redact.py](https://github.com/seanpedrick-case/doc_redaction/blob/main/cli_redact.py#L321)
|
|
257
|
+
|
|
271
258
|
If you installed from **PyPI**, use the installed console script:
|
|
272
259
|
|
|
273
260
|
```bash
|
|
@@ -280,7 +267,9 @@ From a **repository checkout**, you can also run:
|
|
|
280
267
|
python cli_redact.py --help
|
|
281
268
|
```
|
|
282
269
|
|
|
283
|
-
|
|
270
|
+
#### Python package commands
|
|
271
|
+
|
|
272
|
+
For Python examples in using the Python package, please see [Python Package usage (Python)](https://seanpedrick-case.github.io/doc_redaction/src/python_package_usage.html).
|
|
284
273
|
|
|
285
274
|
---
|
|
286
275
|
|
|
@@ -387,9 +376,9 @@ If those endpoints are not present in your deployment, fall back to the long UI-
|
|
|
387
376
|
|
|
388
377
|
### Optional: MCP server
|
|
389
378
|
|
|
390
|
-
If you want external agents to call this app reliably without re-implementing Gradio upload/call/poll/download details, consider an **MCP server** that wraps the main tasks (`redact_document`, `apply_review_redactions`, `redact_tabular`, `summarise_document`) behind a small tool interface. See [
|
|
379
|
+
If you want external agents to call this app reliably without re-implementing Gradio upload/call/poll/download details, consider an **MCP server** that wraps the main tasks (`redact_document`, `apply_review_redactions`, `redact_tabular`, `summarise_document`) behind a small tool interface. See the [relevant documentation](https://github.com/seanpedrick-case/doc_redaction/blob/main/mcp_doc_redaction/README.md).
|
|
391
380
|
|
|
392
|
-
**Use as a library:** After installing from [PyPI](https://pypi.org/project/doc-redaction/) (`pip install doc_redaction`), you can call the same workflows as the Gradio `api_name` routes from Python. See the documentation: [Package
|
|
381
|
+
**Use as a library:** After installing from [PyPI](https://pypi.org/project/doc-redaction/) (`pip install doc_redaction`), you can call the same workflows as the Gradio `api_name` routes from Python. See the documentation: [Python Package usage (Python)](https://seanpedrick-case.github.io/doc_redaction/src/python_package_usage.html).
|
|
393
382
|
|
|
394
383
|
To extract text from documents, the 'Local' options are PikePDF for PDFs with selectable text, and OCR with Tesseract. Use AWS Textract to extract more complex elements e.g. handwriting, signatures, or unclear text. PaddleOCR and VLM support is also provided (see the installation instructions below).
|
|
395
384
|
|
|
@@ -21,81 +21,39 @@ Redact personally identifiable information (PII) from documents (PDF, PNG, JPG),
|
|
|
21
21
|
|
|
22
22
|
Follow these instructions to get the document redaction application running on your local machine.
|
|
23
23
|
|
|
24
|
-
### 1.
|
|
24
|
+
### 1. Package installation
|
|
25
25
|
|
|
26
|
-
|
|
27
|
-
|
|
28
|
-
---
|
|
29
|
-
|
|
30
|
-
#### Automated dependency setup (recommended)
|
|
31
|
-
|
|
32
|
-
If you **don’t have admin rights** (or you just want the simplest setup), you can have the project download and configure **Tesseract** and **Poppler** into a local `redaction_deps/` folder inside the doc_redaction folder.
|
|
33
|
-
|
|
34
|
-
You need the installer script available first, which means either:
|
|
35
|
-
|
|
36
|
-
- **Repository checkout**: `git clone ...` and run the command from the repo root (recommended for the web UI), or
|
|
37
|
-
- **PyPI install**: `pip install doc_redaction` and run from a writable folder where you want `redaction_deps/` and `config/app_config.env` to be created/updated.
|
|
26
|
+
#### Install from source repo (recommended for full features)
|
|
38
27
|
|
|
39
|
-
|
|
28
|
+
Clone the repository and install in editable mode:
|
|
40
29
|
|
|
41
30
|
```bash
|
|
42
|
-
|
|
31
|
+
git clone https://github.com/seanpedrick-case/doc_redaction.git
|
|
32
|
+
cd doc_redaction
|
|
33
|
+
pip install -e .
|
|
43
34
|
```
|
|
44
35
|
|
|
45
|
-
|
|
46
|
-
|
|
47
|
-
To just check whether your machine can already see the tools:
|
|
36
|
+
##### Full install from source (Paddle and VLM)
|
|
48
37
|
|
|
49
38
|
```bash
|
|
50
|
-
|
|
39
|
+
pip install -e ".[paddle,vlm]"
|
|
51
40
|
```
|
|
52
41
|
|
|
53
|
-
|
|
54
|
-
#### **On Windows**
|
|
55
|
-
|
|
56
|
-
If you don’t use the automated setup above, you can install the dependencies manually by downloading installers and adding the programs to your system's PATH.
|
|
57
|
-
|
|
58
|
-
1. **Install Tesseract OCR:**
|
|
59
|
-
* Download the installer from the official Tesseract at [UB Mannheim page](https://github.com/UB-Mannheim/tesseract/wiki) (e.g., `tesseract-ocr-w64-setup-v5.X.X...exe`).
|
|
60
|
-
* Run the installer.
|
|
61
|
-
* **IMPORTANT:** During installation, ensure you select the option to "Add Tesseract to system PATH for all users" or a similar option. This is crucial for the application to find the Tesseract executable.
|
|
62
|
-
|
|
63
|
-
|
|
64
|
-
2. **Install Poppler:**
|
|
65
|
-
* Download the latest Poppler binary for Windows. A common source is the [Poppler for Windows](https://github.com/oschwartz10612/poppler-windows) GitHub releases page. Download the `.zip` file (e.g., `poppler-25.07.0-win.zip`).
|
|
66
|
-
* Extract the contents of the zip file to a permanent location on your computer, for example, `C:\Program Files\poppler\`.
|
|
67
|
-
* You must add the `bin` folder from your Poppler installation to your system's PATH environment variable.
|
|
68
|
-
* Search for "Edit the system environment variables" in the Windows Start Menu and open it.
|
|
69
|
-
* Click the "Environment Variables..." button.
|
|
70
|
-
* In the "System variables" section, find and select the `Path` variable, then click "Edit...".
|
|
71
|
-
* Click "New" and add the full path to the `bin` directory inside your Poppler folder (e.g., `C:\Program Files\poppler\poppler-24.02.0\bin`).
|
|
72
|
-
* Click OK on all windows to save the changes.
|
|
73
|
-
|
|
74
|
-
To verify, open a new Command Prompt and run `tesseract --version` and `pdftoppm -v`. If they both return version information, you have successfully installed the prerequisites.
|
|
75
|
-
|
|
76
|
-
---
|
|
77
|
-
|
|
78
|
-
#### **On Linux (Debian/Ubuntu)**
|
|
79
|
-
|
|
80
|
-
Open your terminal and run the following command to install Tesseract and Poppler:
|
|
81
|
-
|
|
42
|
+
Note that the versions of both PaddleOCR and Torch installed by default are the CPU-only versions. If you want to install the equivalent GPU versions, you will need to run the following commands:
|
|
82
43
|
```bash
|
|
83
|
-
|
|
44
|
+
pip install paddlepaddle-gpu==3.2.1 --index-url https://www.paddlepaddle.org.cn/packages/stable/cu129/
|
|
84
45
|
```
|
|
85
46
|
|
|
86
|
-
|
|
87
|
-
|
|
88
|
-
Open your terminal and use the `dnf` or `yum` package manager:
|
|
47
|
+
**Note:** It is difficult to get paddlepaddle gpu working in an environment alongside torch. You may well need to reinstall the cpu version to ensure compatibility, and run paddlepaddle-gpu in a separate environment without torch installed. If you get errors related to .dll files following paddle gpu install, you may need to install the latest c++ redistributables. For Windows, you can find them [here](https://learn.microsoft.com/en-us/cpp/windows/latest-supported-vc-redist?view=msvc-170)
|
|
89
48
|
|
|
90
49
|
```bash
|
|
91
|
-
|
|
50
|
+
pip install torch==2.8.0 --index-url https://download.pytorch.org/whl/cu129
|
|
51
|
+
pip install torchvision --index-url https://download.pytorch.org/whl/cu129
|
|
92
52
|
```
|
|
93
|
-
---
|
|
94
|
-
|
|
95
53
|
|
|
96
|
-
|
|
54
|
+
#### Install from PyPI (recommended for users and library use)
|
|
97
55
|
|
|
98
|
-
|
|
56
|
+
Create a virtual environment (recommended) and install **doc_redaction**.
|
|
99
57
|
|
|
100
58
|
```bash
|
|
101
59
|
python -m venv venv
|
|
@@ -105,8 +63,6 @@ python -m venv venv
|
|
|
105
63
|
source venv/bin/activate
|
|
106
64
|
```
|
|
107
65
|
|
|
108
|
-
#### Install from PyPI (recommended for users and library use)
|
|
109
|
-
|
|
110
66
|
The package is published on PyPI as **`doc-redaction`** (import name **`doc_redaction`**):
|
|
111
67
|
|
|
112
68
|
```bash
|
|
@@ -119,7 +75,7 @@ Optional extras (same as in `pyproject.toml`):
|
|
|
119
75
|
pip install "doc_redaction[paddle,vlm]"
|
|
120
76
|
```
|
|
121
77
|
|
|
122
|
-
For programmatic use (CLI-first API matching Gradio `api_name` routes), see **[Package
|
|
78
|
+
For programmatic use (CLI-first API matching Gradio `api_name` routes), see **[Python Package usage (Python)](https://seanpedrick-case.github.io/doc_redaction/src/python_package_usage.html)**. The console script **`cli_redact`** is available after install.
|
|
123
79
|
|
|
124
80
|
**Web UI from a PyPI install:** You *can* start the Gradio UI after `pip install doc_redaction` by running:
|
|
125
81
|
|
|
@@ -131,69 +87,98 @@ python -m app
|
|
|
131
87
|
|
|
132
88
|
- It will look for configuration at `config/app_config.env` *relative to the folder you run it from* (and `python -m doc_redaction.install_deps` will also write `config/app_config.env` there).
|
|
133
89
|
- It may create new folders in that location (for example `config/`, `output/`, `input/`, `logs/`, `usage/`, `feedback/`, and temporary/cache folders depending on your settings).
|
|
134
|
-
- The
|
|
90
|
+
- The UI example files and bundled assets are packaged with the PyPI install (they live inside the installed `doc_redaction` package). If you run from a “random” directory after a PyPI install, the app can still locate its packaged examples; your working directory mainly affects where `config/`, `input/`, `output/`, logs, and temp folders are created.
|
|
135
91
|
|
|
136
92
|
In practice, the **smoothest UI experience** (examples, bundled assets, docs links, predictable relative paths) is still usually via a **repository checkout** or **Docker**, but PyPI install is sufficient to launch the UI as long as you run it from a suitable working folder and have the system dependencies available (or run `python -m doc_redaction.install_deps` first).
|
|
137
93
|
|
|
138
|
-
####
|
|
94
|
+
#### Docker installation
|
|
139
95
|
|
|
140
|
-
|
|
96
|
+
The doc_redaction Redaction app can be installed by using the [Dockerfile](https://github.com/seanpedrick-case/doc_redaction/blob/main/Dockerfile) or Docker compose files ([llama.cpp](https://github.com/ggml-org/llama.cpp), [vLLM](https://docs.vllm.ai/en/stable/)) provided in the repo.
|
|
141
97
|
|
|
142
|
-
|
|
143
|
-
git clone https://github.com/seanpedrick-case/doc_redaction.git
|
|
144
|
-
cd doc_redaction
|
|
145
|
-
pip install -e .
|
|
146
|
-
```
|
|
98
|
+
##### Without Llama.cpp / vLLM inference server
|
|
147
99
|
|
|
148
|
-
|
|
100
|
+
If you want a working Docker installation without GPU support, you can install from the [Dockerfile](https://github.com/seanpedrick-case/doc_redaction/blob/main/Dockerfile) in the repo. A working example of this, with the CPU version of PaddleOCR, can be found on [Hugging Face](https://huggingface.co/spaces/seanpedrickcase/document_redaction). You can adjust the INSTALL_PADDLEOCR, PADDLE_GPU_ENABLED, INSTALL_VLM, and TORCH_GPU_ENABLED config variables to adjust for PaddleOCR and Transformers packages for local VLM support. Note that GPU-enabled PaddleOCR, and GPU-enabled Transformers/Torch often don't work well together, which is one reason why a Llama.cpp/vLLM inference server Docker installation option is provided below.
|
|
149
101
|
|
|
150
|
-
|
|
151
|
-
pip install -r requirements_lightweight.txt
|
|
152
|
-
```
|
|
102
|
+
##### With Llama.cpp / vLLM inference server
|
|
153
103
|
|
|
154
|
-
|
|
104
|
+
The project now has Docker and Docker compose files available to pair running the Redaction app with local inference servers powered by [llama.cpp](https://github.com/ggml-org/llama.cpp), or [vLLM](https://docs.vllm.ai/en/stable/). Llama.cpp is more flexible than vLLM for low VRAM systems, as Llama.cpp will offload to cpu/system RAM automatically rather than failing as vLLM tends to do.
|
|
155
105
|
|
|
156
|
-
|
|
157
|
-
pip install -e ".[paddle,vlm]"
|
|
158
|
-
```
|
|
106
|
+
For Llama.cpp, you can use the [docker-compose_llama.yml](https://github.com/seanpedrick-case/doc_redaction/blob/main/docker-compose_llama.yml) file, and for vLLM, you can use the [docker-compose_vllm.yml](https://github.com/seanpedrick-case/doc_redaction/blob/main/docker-compose_vllm.yml) file. To run, Docker / Docker Desktop should be installed, and then you can run the commands suggested in the top of the files to run the servers.
|
|
159
107
|
|
|
160
|
-
|
|
108
|
+
You will need ~40-50GB of disk space to run everything depending on the model chosen from the compose file. For the vLLM server, you will need 24 GB VRAM. For the Llama.cpp server, 24 GB VRAM is needed to run at full speed, but the n-gpu-layers and n-cpu-moe parameters in the Docker compose file can be adjusted to fit into your system. I would suggest that 8 GB VRAM is needed as a bare minimum for decent inference speed. See the [Unsloth guide](https://unsloth.ai/docs/models/qwen3.5) for more details on working with GGUF files for Qwen 3.5.
|
|
161
109
|
|
|
162
|
-
|
|
163
|
-
|
|
164
|
-
|
|
110
|
+
### 2. Install prerequisites: Tesseract and Poppler
|
|
111
|
+
|
|
112
|
+
This application relies on two external tools for OCR (Tesseract) and PDF processing (Poppler). Please install them on your system before proceeding.
|
|
113
|
+
|
|
114
|
+
---
|
|
115
|
+
|
|
116
|
+
#### Automated dependency setup (recommended)
|
|
117
|
+
|
|
118
|
+
If you **don’t have admin rights** (or you just want the simplest setup), you can have the project download and configure **Tesseract** and **Poppler** into a local `redaction_deps/` folder inside the doc_redaction folder.
|
|
119
|
+
|
|
120
|
+
You need the installer script available first, which means either:
|
|
121
|
+
|
|
122
|
+
- **Repository checkout**: `git clone ...` and run the command from the repo root (recommended for the web UI), or
|
|
123
|
+
- **PyPI install**: `pip install doc_redaction` and run from a writable folder where you want `redaction_deps/` and `config/app_config.env` to be created/updated.
|
|
124
|
+
|
|
125
|
+
From the repository root (or your chosen working folder) after creating/activating your venv and installing Python requirements:
|
|
165
126
|
|
|
166
|
-
Note that the versions of both PaddleOCR and Torch installed by default are the CPU-only versions. If you want to install the equivalent GPU versions, you will need to run the following commands:
|
|
167
127
|
```bash
|
|
168
|
-
|
|
128
|
+
python -m doc_redaction.install_deps
|
|
169
129
|
```
|
|
170
130
|
|
|
171
|
-
|
|
131
|
+
This writes `TESSERACT_FOLDER` / `POPPLER_FOLDER` into `config/app_config.env` so the app can find the binaries without you editing your system PATH.
|
|
132
|
+
|
|
133
|
+
To just check whether your machine can already see the tools:
|
|
172
134
|
|
|
173
135
|
```bash
|
|
174
|
-
|
|
175
|
-
pip install torchvision --index-url https://download.pytorch.org/whl/cu129
|
|
136
|
+
python -m doc_redaction.install_deps --verify-only
|
|
176
137
|
```
|
|
177
138
|
|
|
178
|
-
####
|
|
139
|
+
#### **On Windows**
|
|
179
140
|
|
|
180
|
-
|
|
141
|
+
If you don’t use the automated setup above, you can install the dependencies manually by downloading installers and adding the programs to your system's PATH.
|
|
181
142
|
|
|
182
|
-
|
|
143
|
+
1. **Install Tesseract OCR:**
|
|
144
|
+
* Download the installer from the official Tesseract at [UB Mannheim page](https://github.com/UB-Mannheim/tesseract/wiki) (e.g., `tesseract-ocr-w64-setup-v5.X.X...exe`).
|
|
145
|
+
* Run the installer.
|
|
146
|
+
* **IMPORTANT:** During installation, ensure you select the option to "Add Tesseract to system PATH for all users" or a similar option. This is crucial for the application to find the Tesseract executable.
|
|
183
147
|
|
|
184
|
-
If you want a working Docker installation without GPU support, you can install from the [Dockerfile](https://github.com/seanpedrick-case/doc_redaction/blob/main/Dockerfile) in the repo. A working example of this, with the CPU version of PaddleOCR, can be found on [Hugging Face](https://huggingface.co/spaces/seanpedrickcase/document_redaction). You can adjust the INSTALL_PADDLEOCR, PADDLE_GPU_ENABLED, INSTALL_VLM, and TORCH_GPU_ENABLED config variables to adjust for PaddleOCR and Transformers packages for local VLM support. Note that GPU-enabled PaddleOCR, and GPU-enabled Transformers/Torch often don't work well together, which is one reason why a Llama.cpp/vLLM inference server Docker installation option is provided below.
|
|
185
148
|
|
|
186
|
-
|
|
149
|
+
2. **Install Poppler:**
|
|
150
|
+
* Download the latest Poppler binary for Windows. A common source is the [Poppler for Windows](https://github.com/oschwartz10612/poppler-windows) GitHub releases page. Download the `.zip` file (e.g., `poppler-25.07.0-win.zip`).
|
|
151
|
+
* Extract the contents of the zip file to a permanent location on your computer, for example, `C:\Program Files\poppler\`.
|
|
152
|
+
* You must add the `bin` folder from your Poppler installation to your system's PATH environment variable.
|
|
153
|
+
* Search for "Edit the system environment variables" in the Windows Start Menu and open it.
|
|
154
|
+
* Click the "Environment Variables..." button.
|
|
155
|
+
* In the "System variables" section, find and select the `Path` variable, then click "Edit...".
|
|
156
|
+
* Click "New" and add the full path to the `bin` directory inside your Poppler folder (e.g., `C:\Program Files\poppler\poppler-24.02.0\bin`).
|
|
157
|
+
* Click OK on all windows to save the changes.
|
|
187
158
|
|
|
188
|
-
|
|
159
|
+
To verify, open a new Command Prompt and run `tesseract --version` and `pdftoppm -v`. If they both return version information, you have successfully installed the prerequisites.
|
|
160
|
+
---
|
|
189
161
|
|
|
190
|
-
|
|
162
|
+
#### **On Linux (Debian/Ubuntu)**
|
|
191
163
|
|
|
192
|
-
|
|
164
|
+
Open your terminal and run the following command to install Tesseract and Poppler:
|
|
165
|
+
|
|
166
|
+
```bash
|
|
167
|
+
sudo apt-get update && sudo apt-get install -y tesseract-ocr poppler-utils
|
|
168
|
+
```
|
|
169
|
+
|
|
170
|
+
#### **On Linux (Fedora/CentOS/RHEL)**
|
|
171
|
+
|
|
172
|
+
Open your terminal and use the `dnf` or `yum` package manager:
|
|
173
|
+
|
|
174
|
+
```bash
|
|
175
|
+
sudo dnf install -y tesseract poppler-utils
|
|
176
|
+
```
|
|
177
|
+
---
|
|
193
178
|
|
|
194
179
|
### 3. Run the Application
|
|
195
180
|
|
|
196
|
-
With all dependencies installed, you can now start the Gradio application.
|
|
181
|
+
With all dependencies installed, you can now start the Gradio application GUI. For a guide on how to use this, please go [here](https://seanpedrick-case.github.io/doc_redaction/src/user_guide.html).
|
|
197
182
|
|
|
198
183
|
```bash
|
|
199
184
|
python app.py
|
|
@@ -205,6 +190,8 @@ Open this URL in your web browser to use the document redaction tool
|
|
|
205
190
|
|
|
206
191
|
#### Command line interface
|
|
207
192
|
|
|
193
|
+
For example CLI commands, please refer to [this guide](https://seanpedrick-case.github.io/doc_redaction/src/user_guide.html#command-line-interface-cli) or the examples in [cli_redact.py](https://github.com/seanpedrick-case/doc_redaction/blob/main/cli_redact.py#L321)
|
|
194
|
+
|
|
208
195
|
If you installed from **PyPI**, use the installed console script:
|
|
209
196
|
|
|
210
197
|
```bash
|
|
@@ -217,7 +204,9 @@ From a **repository checkout**, you can also run:
|
|
|
217
204
|
python cli_redact.py --help
|
|
218
205
|
```
|
|
219
206
|
|
|
220
|
-
|
|
207
|
+
#### Python package commands
|
|
208
|
+
|
|
209
|
+
For Python examples in using the Python package, please see [Python Package usage (Python)](https://seanpedrick-case.github.io/doc_redaction/src/python_package_usage.html).
|
|
221
210
|
|
|
222
211
|
---
|
|
223
212
|
|
|
@@ -329,9 +318,9 @@ If those endpoints are not present in your deployment, fall back to the long UI-
|
|
|
329
318
|
|
|
330
319
|
### Optional: MCP server
|
|
331
320
|
|
|
332
|
-
If you want external agents to call this app reliably without re-implementing Gradio upload/call/poll/download details, consider an **MCP server** that wraps the main tasks (`redact_document`, `apply_review_redactions`, `redact_tabular`, `summarise_document`) behind a small tool interface. See [
|
|
321
|
+
If you want external agents to call this app reliably without re-implementing Gradio upload/call/poll/download details, consider an **MCP server** that wraps the main tasks (`redact_document`, `apply_review_redactions`, `redact_tabular`, `summarise_document`) behind a small tool interface. See the [relevant documentation](https://github.com/seanpedrick-case/doc_redaction/blob/main/mcp_doc_redaction/README.md).
|
|
333
322
|
|
|
334
|
-
**Use as a library:** After installing from [PyPI](https://pypi.org/project/doc-redaction/) (`pip install doc_redaction`), you can call the same workflows as the Gradio `api_name` routes from Python. See the documentation: [Package
|
|
323
|
+
**Use as a library:** After installing from [PyPI](https://pypi.org/project/doc-redaction/) (`pip install doc_redaction`), you can call the same workflows as the Gradio `api_name` routes from Python. See the documentation: [Python Package usage (Python)](https://seanpedrick-case.github.io/doc_redaction/src/python_package_usage.html).
|
|
335
324
|
|
|
336
325
|
To extract text from documents, the 'Local' options are PikePDF for PDFs with selectable text, and OCR with Tesseract. Use AWS Textract to extract more complex elements e.g. handwriting, signatures, or unclear text. PaddleOCR and VLM support is also provided (see the installation instructions below).
|
|
337
326
|
|