pdf-anonymizer-cli 0.3.0__tar.gz → 0.3.2__tar.gz
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- pdf_anonymizer_cli-0.3.2/PKG-INFO +133 -0
- pdf_anonymizer_cli-0.3.2/README.md +120 -0
- {pdf_anonymizer_cli-0.3.0 → pdf_anonymizer_cli-0.3.2}/pyproject.toml +1 -1
- {pdf_anonymizer_cli-0.3.0 → pdf_anonymizer_cli-0.3.2}/src/pdf_anonymizer_cli/cli.py +9 -10
- pdf_anonymizer_cli-0.3.2/src/pdf_anonymizer_cli.egg-info/PKG-INFO +133 -0
- pdf_anonymizer_cli-0.3.0/PKG-INFO +0 -100
- pdf_anonymizer_cli-0.3.0/README.md +0 -87
- pdf_anonymizer_cli-0.3.0/src/pdf_anonymizer_cli.egg-info/PKG-INFO +0 -100
- {pdf_anonymizer_cli-0.3.0 → pdf_anonymizer_cli-0.3.2}/setup.cfg +0 -0
- {pdf_anonymizer_cli-0.3.0 → pdf_anonymizer_cli-0.3.2}/src/pdf_anonymizer_cli/__init__.py +0 -0
- {pdf_anonymizer_cli-0.3.0 → pdf_anonymizer_cli-0.3.2}/src/pdf_anonymizer_cli/main.py +0 -0
- {pdf_anonymizer_cli-0.3.0 → pdf_anonymizer_cli-0.3.2}/src/pdf_anonymizer_cli.egg-info/SOURCES.txt +0 -0
- {pdf_anonymizer_cli-0.3.0 → pdf_anonymizer_cli-0.3.2}/src/pdf_anonymizer_cli.egg-info/dependency_links.txt +0 -0
- {pdf_anonymizer_cli-0.3.0 → pdf_anonymizer_cli-0.3.2}/src/pdf_anonymizer_cli.egg-info/entry_points.txt +0 -0
- {pdf_anonymizer_cli-0.3.0 → pdf_anonymizer_cli-0.3.2}/src/pdf_anonymizer_cli.egg-info/requires.txt +0 -0
- {pdf_anonymizer_cli-0.3.0 → pdf_anonymizer_cli-0.3.2}/src/pdf_anonymizer_cli.egg-info/top_level.txt +0 -0
|
@@ -0,0 +1,133 @@
|
|
|
1
|
+
Metadata-Version: 2.4
|
|
2
|
+
Name: pdf-anonymizer-cli
|
|
3
|
+
Version: 0.3.2
|
|
4
|
+
Summary: CLI for a tool to anonymize PDF, Markdown, and plain text files using LLMs.
|
|
5
|
+
Author-email: Leonid Ganeline <leo.gan.57@gmail.com>
|
|
6
|
+
License: MIT
|
|
7
|
+
Project-URL: repository, https://github.com/leo-gan/anonymizer
|
|
8
|
+
Requires-Python: >=3.10
|
|
9
|
+
Description-Content-Type: text/markdown
|
|
10
|
+
Requires-Dist: typer
|
|
11
|
+
Requires-Dist: rich
|
|
12
|
+
Requires-Dist: pdf-anonymizer-core
|
|
13
|
+
|
|
14
|
+
# 🦉🫥 PDF Anonymizer CLI
|
|
15
|
+
|
|
16
|
+
A command-line interface for anonymizing PDF, Markdown, and plain text files using LLMs.
|
|
17
|
+
|
|
18
|
+
- **High-Quality Anonymization**: Leverages LLMs to identify and replace Personally Identifiable Information (PII) with high accuracy.
|
|
19
|
+
- **Large File Support**: Consistently anonymizes large files (tested up to 1GB).
|
|
20
|
+
- **Multi-Provider & Cost-Effective**: Free to use with local [Ollama](https://ollama.com/) models. It also supports major providers like [OpenAI](https://openai.com/), [Anthropic](https://www.anthropic.com/), [Google](https://ai.google.com/), [Hugging Face](https://huggingface.co/), and [OpenRouter](https://openrouter.ai/).
|
|
21
|
+
- **Reversible**: Supports deanonymization to recover original data when needed.
|
|
22
|
+
- **Multi-Format**: Works with PDF, Markdown, and plain text files.
|
|
23
|
+
|
|
24
|
+
|
|
25
|
+
## Installation
|
|
26
|
+
|
|
27
|
+
Install the CLI with your favorite package manager. To use a specific LLM provider, you must install the corresponding extra.
|
|
28
|
+
|
|
29
|
+
- **Google**: `pip install "pdf-anonymizer-cli[google]"`
|
|
30
|
+
- **Ollama**: `pip install "pdf-anonymizer-cli[ollama]"`
|
|
31
|
+
- **Hugging Face**: `pip install "pdf-anonymizer-cli[huggingface]"`
|
|
32
|
+
- **OpenRouter**: `pip install "pdf-anonymizer-cli[openrouter]"`
|
|
33
|
+
- **OpenAI**: `pip install "pdf-anonymizer-cli[openai]"`
|
|
34
|
+
- **Anthropic**: `pip install "pdf-anonymizer-cli[anthropic]"`
|
|
35
|
+
|
|
36
|
+
You can also install multiple extras at once:
|
|
37
|
+
|
|
38
|
+
```bash
|
|
39
|
+
pip install "pdf-anonymizer-cli[google,openrouter]"
|
|
40
|
+
```
|
|
41
|
+
|
|
42
|
+
This installs the `pdf-anonymizer` executable.
|
|
43
|
+
|
|
44
|
+
|
|
45
|
+
## Environment Variables
|
|
46
|
+
|
|
47
|
+
The CLI will automatically load a `.env` file from the current directory or any parent directory. For consistency, it's recommended to place a single `.env` file at the root of the repository.
|
|
48
|
+
|
|
49
|
+
- `GOOGLE_API_KEY`: Required when using Google models.
|
|
50
|
+
- `HUGGING_FACE_TOKEN`: Required when using Hugging Face models. You can get a token from [here](https://huggingface.co/docs/hub/security-tokens).
|
|
51
|
+
- `OPENROUTER_API_KEY`: Required when using OpenRouter models.
|
|
52
|
+
- `OPENAI_API_KEY`: Required when using OpenAI models.
|
|
53
|
+
- `ANTHROPIC_API_KEY`: Required when using Anthropic models.
|
|
54
|
+
- `OLLAMA_HOST`: Optional, defaults to `http://localhost:11434` when using Ollama models.
|
|
55
|
+
|
|
56
|
+
Example `.env` file:
|
|
57
|
+
```env
|
|
58
|
+
GOOGLE_API_KEY="YOUR_API_KEY_HERE"
|
|
59
|
+
HUGGING_FACE_TOKEN="YOUR_HF_TOKEN_HERE"
|
|
60
|
+
OPENROUTER_API_KEY="YOUR_OPENROUTER_KEY"
|
|
61
|
+
```
|
|
62
|
+
|
|
63
|
+
## Usage
|
|
64
|
+
|
|
65
|
+
### Anonymize
|
|
66
|
+
|
|
67
|
+
The `run` command anonymizes one or more files.
|
|
68
|
+
|
|
69
|
+
```bash
|
|
70
|
+
pdf-anonymizer run FILE_PATH [FILE_PATH ...] \
|
|
71
|
+
[--characters-to-anonymize INTEGER] \
|
|
72
|
+
[--prompt-name {simple|detailed}] \
|
|
73
|
+
[--model-name TEXT] \
|
|
74
|
+
[--anonymized-entities PATH]
|
|
75
|
+
```
|
|
76
|
+
|
|
77
|
+
**Arguments**:
|
|
78
|
+
- `FILE_PATH`: Path to one or several PDF, Markdown, or text files for anonymization.
|
|
79
|
+
|
|
80
|
+
**Options**:
|
|
81
|
+
- `--characters-to-anonymize INTEGER`: Number of characters to process in each chunk (default: `100000`).
|
|
82
|
+
- `--prompt-name [simple|detailed]`: The prompt template to use (default: `detailed`).
|
|
83
|
+
- `--model-name TEXT`: The language model to use.
|
|
84
|
+
- `--anonymized-entities PATH`: Path to a file with a list of entities to anonymize.
|
|
85
|
+
|
|
86
|
+
**Models**:
|
|
87
|
+
You can use any of the predefined models below, or specify a new model using the format `"provider/model-name"`.
|
|
88
|
+
For example: `--model-name "google/gemini-flash-latest"`.
|
|
89
|
+
|
|
90
|
+
- **Google**: `gemini-2.5-pro`, `gemini-2.5-flash` (default), `gemini-2.5-flash-lite`.
|
|
91
|
+
- **Ollama**: `gemma:7b`, `phi4-mini`.
|
|
92
|
+
- **Hugging Face**: `openai/gpt-oss-20b`, `mistralai/Mistral-7B-Instruct-v0.1`, `HuggingFaceH4/zephyr-7b-beta`.
|
|
93
|
+
- **OpenRouter**: `openai/gpt-4o`, `google/gemini-pro`.
|
|
94
|
+
- **OpenAI**: `gpt-4o`, `gpt-5`.
|
|
95
|
+
- **Anthropic**: `claude-4-sonet`, `claude-4.5-sonet`.
|
|
96
|
+
|
|
97
|
+
### Examples
|
|
98
|
+
|
|
99
|
+
**Basic anonymization with the default model (Google)**:
|
|
100
|
+
```bash
|
|
101
|
+
pdf-anonymizer run document.pdf
|
|
102
|
+
```
|
|
103
|
+
|
|
104
|
+
**A new model (Google) and a simple prompt**:
|
|
105
|
+
```bash
|
|
106
|
+
pdf-anonymizer run notes.md --model-name "google/gemini-flash-latest" --prompt-name simple
|
|
107
|
+
```
|
|
108
|
+
|
|
109
|
+
**Using an OpenRouter model**:
|
|
110
|
+
```bash
|
|
111
|
+
pdf-anonymizer run report.pdf --model-name "openai/gpt-4o"
|
|
112
|
+
```
|
|
113
|
+
|
|
114
|
+
### Deanonymize
|
|
115
|
+
|
|
116
|
+
The `deanonymize` command reverts anonymization using a mapping file.
|
|
117
|
+
|
|
118
|
+
```bash
|
|
119
|
+
pdf-anonymizer deanonymize ANONYMIZED_FILE MAPPING_FILE
|
|
120
|
+
```
|
|
121
|
+
|
|
122
|
+
**Arguments**:
|
|
123
|
+
- `ANONYMIZED_FILE`: Path to the anonymized text file.
|
|
124
|
+
- `MAPPING_FILE`: Path to the JSON mapping file.
|
|
125
|
+
|
|
126
|
+
**Example**:
|
|
127
|
+
```bash
|
|
128
|
+
pdf-anonymizer deanonymize \
|
|
129
|
+
data/anonymized/document.anonymized.md \
|
|
130
|
+
data/mappings/document.mapping.json
|
|
131
|
+
```
|
|
132
|
+
|
|
133
|
+
This will create a deanonymized version of the file at `data/deanonymized/document.deanonymized.md`.
|
|
@@ -0,0 +1,120 @@
|
|
|
1
|
+
# 🦉🫥 PDF Anonymizer CLI
|
|
2
|
+
|
|
3
|
+
A command-line interface for anonymizing PDF, Markdown, and plain text files using LLMs.
|
|
4
|
+
|
|
5
|
+
- **High-Quality Anonymization**: Leverages LLMs to identify and replace Personally Identifiable Information (PII) with high accuracy.
|
|
6
|
+
- **Large File Support**: Consistently anonymizes large files (tested up to 1GB).
|
|
7
|
+
- **Multi-Provider & Cost-Effective**: Free to use with local [Ollama](https://ollama.com/) models. It also supports major providers like [OpenAI](https://openai.com/), [Anthropic](https://www.anthropic.com/), [Google](https://ai.google.com/), [Hugging Face](https://huggingface.co/), and [OpenRouter](https://openrouter.ai/).
|
|
8
|
+
- **Reversible**: Supports deanonymization to recover original data when needed.
|
|
9
|
+
- **Multi-Format**: Works with PDF, Markdown, and plain text files.
|
|
10
|
+
|
|
11
|
+
|
|
12
|
+
## Installation
|
|
13
|
+
|
|
14
|
+
Install the CLI with your favorite package manager. To use a specific LLM provider, you must install the corresponding extra.
|
|
15
|
+
|
|
16
|
+
- **Google**: `pip install "pdf-anonymizer-cli[google]"`
|
|
17
|
+
- **Ollama**: `pip install "pdf-anonymizer-cli[ollama]"`
|
|
18
|
+
- **Hugging Face**: `pip install "pdf-anonymizer-cli[huggingface]"`
|
|
19
|
+
- **OpenRouter**: `pip install "pdf-anonymizer-cli[openrouter]"`
|
|
20
|
+
- **OpenAI**: `pip install "pdf-anonymizer-cli[openai]"`
|
|
21
|
+
- **Anthropic**: `pip install "pdf-anonymizer-cli[anthropic]"`
|
|
22
|
+
|
|
23
|
+
You can also install multiple extras at once:
|
|
24
|
+
|
|
25
|
+
```bash
|
|
26
|
+
pip install "pdf-anonymizer-cli[google,openrouter]"
|
|
27
|
+
```
|
|
28
|
+
|
|
29
|
+
This installs the `pdf-anonymizer` executable.
|
|
30
|
+
|
|
31
|
+
|
|
32
|
+
## Environment Variables
|
|
33
|
+
|
|
34
|
+
The CLI will automatically load a `.env` file from the current directory or any parent directory. For consistency, it's recommended to place a single `.env` file at the root of the repository.
|
|
35
|
+
|
|
36
|
+
- `GOOGLE_API_KEY`: Required when using Google models.
|
|
37
|
+
- `HUGGING_FACE_TOKEN`: Required when using Hugging Face models. You can get a token from [here](https://huggingface.co/docs/hub/security-tokens).
|
|
38
|
+
- `OPENROUTER_API_KEY`: Required when using OpenRouter models.
|
|
39
|
+
- `OPENAI_API_KEY`: Required when using OpenAI models.
|
|
40
|
+
- `ANTHROPIC_API_KEY`: Required when using Anthropic models.
|
|
41
|
+
- `OLLAMA_HOST`: Optional, defaults to `http://localhost:11434` when using Ollama models.
|
|
42
|
+
|
|
43
|
+
Example `.env` file:
|
|
44
|
+
```env
|
|
45
|
+
GOOGLE_API_KEY="YOUR_API_KEY_HERE"
|
|
46
|
+
HUGGING_FACE_TOKEN="YOUR_HF_TOKEN_HERE"
|
|
47
|
+
OPENROUTER_API_KEY="YOUR_OPENROUTER_KEY"
|
|
48
|
+
```
|
|
49
|
+
|
|
50
|
+
## Usage
|
|
51
|
+
|
|
52
|
+
### Anonymize
|
|
53
|
+
|
|
54
|
+
The `run` command anonymizes one or more files.
|
|
55
|
+
|
|
56
|
+
```bash
|
|
57
|
+
pdf-anonymizer run FILE_PATH [FILE_PATH ...] \
|
|
58
|
+
[--characters-to-anonymize INTEGER] \
|
|
59
|
+
[--prompt-name {simple|detailed}] \
|
|
60
|
+
[--model-name TEXT] \
|
|
61
|
+
[--anonymized-entities PATH]
|
|
62
|
+
```
|
|
63
|
+
|
|
64
|
+
**Arguments**:
|
|
65
|
+
- `FILE_PATH`: Path to one or several PDF, Markdown, or text files for anonymization.
|
|
66
|
+
|
|
67
|
+
**Options**:
|
|
68
|
+
- `--characters-to-anonymize INTEGER`: Number of characters to process in each chunk (default: `100000`).
|
|
69
|
+
- `--prompt-name [simple|detailed]`: The prompt template to use (default: `detailed`).
|
|
70
|
+
- `--model-name TEXT`: The language model to use.
|
|
71
|
+
- `--anonymized-entities PATH`: Path to a file with a list of entities to anonymize.
|
|
72
|
+
|
|
73
|
+
**Models**:
|
|
74
|
+
You can use any of the predefined models below, or specify a new model using the format `"provider/model-name"`.
|
|
75
|
+
For example: `--model-name "google/gemini-flash-latest"`.
|
|
76
|
+
|
|
77
|
+
- **Google**: `gemini-2.5-pro`, `gemini-2.5-flash` (default), `gemini-2.5-flash-lite`.
|
|
78
|
+
- **Ollama**: `gemma:7b`, `phi4-mini`.
|
|
79
|
+
- **Hugging Face**: `openai/gpt-oss-20b`, `mistralai/Mistral-7B-Instruct-v0.1`, `HuggingFaceH4/zephyr-7b-beta`.
|
|
80
|
+
- **OpenRouter**: `openai/gpt-4o`, `google/gemini-pro`.
|
|
81
|
+
- **OpenAI**: `gpt-4o`, `gpt-5`.
|
|
82
|
+
- **Anthropic**: `claude-4-sonet`, `claude-4.5-sonet`.
|
|
83
|
+
|
|
84
|
+
### Examples
|
|
85
|
+
|
|
86
|
+
**Basic anonymization with the default model (Google)**:
|
|
87
|
+
```bash
|
|
88
|
+
pdf-anonymizer run document.pdf
|
|
89
|
+
```
|
|
90
|
+
|
|
91
|
+
**A new model (Google) and a simple prompt**:
|
|
92
|
+
```bash
|
|
93
|
+
pdf-anonymizer run notes.md --model-name "google/gemini-flash-latest" --prompt-name simple
|
|
94
|
+
```
|
|
95
|
+
|
|
96
|
+
**Using an OpenRouter model**:
|
|
97
|
+
```bash
|
|
98
|
+
pdf-anonymizer run report.pdf --model-name "openai/gpt-4o"
|
|
99
|
+
```
|
|
100
|
+
|
|
101
|
+
### Deanonymize
|
|
102
|
+
|
|
103
|
+
The `deanonymize` command reverts anonymization using a mapping file.
|
|
104
|
+
|
|
105
|
+
```bash
|
|
106
|
+
pdf-anonymizer deanonymize ANONYMIZED_FILE MAPPING_FILE
|
|
107
|
+
```
|
|
108
|
+
|
|
109
|
+
**Arguments**:
|
|
110
|
+
- `ANONYMIZED_FILE`: Path to the anonymized text file.
|
|
111
|
+
- `MAPPING_FILE`: Path to the JSON mapping file.
|
|
112
|
+
|
|
113
|
+
**Example**:
|
|
114
|
+
```bash
|
|
115
|
+
pdf-anonymizer deanonymize \
|
|
116
|
+
data/anonymized/document.anonymized.md \
|
|
117
|
+
data/mappings/document.mapping.json
|
|
118
|
+
```
|
|
119
|
+
|
|
120
|
+
This will create a deanonymized version of the file at `data/deanonymized/document.deanonymized.md`.
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
[project]
|
|
2
2
|
name = "pdf-anonymizer-cli"
|
|
3
|
-
version = "0.3.
|
|
3
|
+
version = "0.3.2"
|
|
4
4
|
description = "CLI for a tool to anonymize PDF, Markdown, and plain text files using LLMs."
|
|
5
5
|
authors = [{ name = "Leonid Ganeline", email = "leo.gan.57@gmail.com" }]
|
|
6
6
|
license = { text = "MIT" }
|
|
@@ -10,10 +10,9 @@ from pdf_anonymizer_core.conf import (
|
|
|
10
10
|
DEFAULT_CHARACTERS_TO_ANONYMIZE,
|
|
11
11
|
DEFAULT_MODEL_NAME,
|
|
12
12
|
DEFAULT_PROMPT_NAME,
|
|
13
|
-
ModelName,
|
|
14
|
-
ModelProvider,
|
|
15
13
|
PromptEnum,
|
|
16
14
|
get_enum_value,
|
|
15
|
+
get_provider_and_model_name,
|
|
17
16
|
)
|
|
18
17
|
from pdf_anonymizer_core.core import anonymize_file
|
|
19
18
|
from pdf_anonymizer_core.prompts import detailed, simple
|
|
@@ -66,12 +65,11 @@ def run(
|
|
|
66
65
|
),
|
|
67
66
|
] = get_enum_value(PromptEnum, DEFAULT_PROMPT_NAME),
|
|
68
67
|
model_name: Annotated[
|
|
69
|
-
|
|
68
|
+
str,
|
|
70
69
|
typer.Option(
|
|
71
|
-
help="The name of the model to use for anonymization.",
|
|
72
|
-
case_sensitive=False,
|
|
70
|
+
help="The name of the model to use for anonymization (supports enum values or 'provider/model').",
|
|
73
71
|
),
|
|
74
|
-
] =
|
|
72
|
+
] = DEFAULT_MODEL_NAME,
|
|
75
73
|
anonymized_entities: Annotated[
|
|
76
74
|
Optional[Path],
|
|
77
75
|
typer.Option(
|
|
@@ -98,8 +96,9 @@ def run(
|
|
|
98
96
|
"""
|
|
99
97
|
load_environment()
|
|
100
98
|
|
|
101
|
-
|
|
102
|
-
|
|
99
|
+
provider_name, _ = get_provider_and_model_name(model_name)
|
|
100
|
+
if provider_name == "google":
|
|
101
|
+
if "gemini" in model_name and not os.getenv("GOOGLE_API_KEY"):
|
|
103
102
|
logging.error(
|
|
104
103
|
"Error: GOOGLE_API_KEY not found. Please set it in the .env file."
|
|
105
104
|
)
|
|
@@ -107,7 +106,7 @@ def run(
|
|
|
107
106
|
|
|
108
107
|
logging.info(f" --file-paths: {file_paths}")
|
|
109
108
|
logging.info(f" --characters-to-anonymize: {characters_to_anonymize}")
|
|
110
|
-
logging.info(f" --model-name: {model_name
|
|
109
|
+
logging.info(f" --model-name: {model_name}")
|
|
111
110
|
|
|
112
111
|
# Select the appropriate prompt template
|
|
113
112
|
prompt_templates: Dict[str, str] = {
|
|
@@ -132,7 +131,7 @@ def run(
|
|
|
132
131
|
str(file_path),
|
|
133
132
|
characters_to_anonymize,
|
|
134
133
|
prompt_template,
|
|
135
|
-
model_name
|
|
134
|
+
model_name,
|
|
136
135
|
entities_to_anonymize,
|
|
137
136
|
)
|
|
138
137
|
|
|
@@ -0,0 +1,133 @@
|
|
|
1
|
+
Metadata-Version: 2.4
|
|
2
|
+
Name: pdf-anonymizer-cli
|
|
3
|
+
Version: 0.3.2
|
|
4
|
+
Summary: CLI for a tool to anonymize PDF, Markdown, and plain text files using LLMs.
|
|
5
|
+
Author-email: Leonid Ganeline <leo.gan.57@gmail.com>
|
|
6
|
+
License: MIT
|
|
7
|
+
Project-URL: repository, https://github.com/leo-gan/anonymizer
|
|
8
|
+
Requires-Python: >=3.10
|
|
9
|
+
Description-Content-Type: text/markdown
|
|
10
|
+
Requires-Dist: typer
|
|
11
|
+
Requires-Dist: rich
|
|
12
|
+
Requires-Dist: pdf-anonymizer-core
|
|
13
|
+
|
|
14
|
+
# 🦉🫥 PDF Anonymizer CLI
|
|
15
|
+
|
|
16
|
+
A command-line interface for anonymizing PDF, Markdown, and plain text files using LLMs.
|
|
17
|
+
|
|
18
|
+
- **High-Quality Anonymization**: Leverages LLMs to identify and replace Personally Identifiable Information (PII) with high accuracy.
|
|
19
|
+
- **Large File Support**: Consistently anonymizes large files (tested up to 1GB).
|
|
20
|
+
- **Multi-Provider & Cost-Effective**: Free to use with local [Ollama](https://ollama.com/) models. It also supports major providers like [OpenAI](https://openai.com/), [Anthropic](https://www.anthropic.com/), [Google](https://ai.google.com/), [Hugging Face](https://huggingface.co/), and [OpenRouter](https://openrouter.ai/).
|
|
21
|
+
- **Reversible**: Supports deanonymization to recover original data when needed.
|
|
22
|
+
- **Multi-Format**: Works with PDF, Markdown, and plain text files.
|
|
23
|
+
|
|
24
|
+
|
|
25
|
+
## Installation
|
|
26
|
+
|
|
27
|
+
Install the CLI with your favorite package manager. To use a specific LLM provider, you must install the corresponding extra.
|
|
28
|
+
|
|
29
|
+
- **Google**: `pip install "pdf-anonymizer-cli[google]"`
|
|
30
|
+
- **Ollama**: `pip install "pdf-anonymizer-cli[ollama]"`
|
|
31
|
+
- **Hugging Face**: `pip install "pdf-anonymizer-cli[huggingface]"`
|
|
32
|
+
- **OpenRouter**: `pip install "pdf-anonymizer-cli[openrouter]"`
|
|
33
|
+
- **OpenAI**: `pip install "pdf-anonymizer-cli[openai]"`
|
|
34
|
+
- **Anthropic**: `pip install "pdf-anonymizer-cli[anthropic]"`
|
|
35
|
+
|
|
36
|
+
You can also install multiple extras at once:
|
|
37
|
+
|
|
38
|
+
```bash
|
|
39
|
+
pip install "pdf-anonymizer-cli[google,openrouter]"
|
|
40
|
+
```
|
|
41
|
+
|
|
42
|
+
This installs the `pdf-anonymizer` executable.
|
|
43
|
+
|
|
44
|
+
|
|
45
|
+
## Environment Variables
|
|
46
|
+
|
|
47
|
+
The CLI will automatically load a `.env` file from the current directory or any parent directory. For consistency, it's recommended to place a single `.env` file at the root of the repository.
|
|
48
|
+
|
|
49
|
+
- `GOOGLE_API_KEY`: Required when using Google models.
|
|
50
|
+
- `HUGGING_FACE_TOKEN`: Required when using Hugging Face models. You can get a token from [here](https://huggingface.co/docs/hub/security-tokens).
|
|
51
|
+
- `OPENROUTER_API_KEY`: Required when using OpenRouter models.
|
|
52
|
+
- `OPENAI_API_KEY`: Required when using OpenAI models.
|
|
53
|
+
- `ANTHROPIC_API_KEY`: Required when using Anthropic models.
|
|
54
|
+
- `OLLAMA_HOST`: Optional, defaults to `http://localhost:11434` when using Ollama models.
|
|
55
|
+
|
|
56
|
+
Example `.env` file:
|
|
57
|
+
```env
|
|
58
|
+
GOOGLE_API_KEY="YOUR_API_KEY_HERE"
|
|
59
|
+
HUGGING_FACE_TOKEN="YOUR_HF_TOKEN_HERE"
|
|
60
|
+
OPENROUTER_API_KEY="YOUR_OPENROUTER_KEY"
|
|
61
|
+
```
|
|
62
|
+
|
|
63
|
+
## Usage
|
|
64
|
+
|
|
65
|
+
### Anonymize
|
|
66
|
+
|
|
67
|
+
The `run` command anonymizes one or more files.
|
|
68
|
+
|
|
69
|
+
```bash
|
|
70
|
+
pdf-anonymizer run FILE_PATH [FILE_PATH ...] \
|
|
71
|
+
[--characters-to-anonymize INTEGER] \
|
|
72
|
+
[--prompt-name {simple|detailed}] \
|
|
73
|
+
[--model-name TEXT] \
|
|
74
|
+
[--anonymized-entities PATH]
|
|
75
|
+
```
|
|
76
|
+
|
|
77
|
+
**Arguments**:
|
|
78
|
+
- `FILE_PATH`: Path to one or several PDF, Markdown, or text files for anonymization.
|
|
79
|
+
|
|
80
|
+
**Options**:
|
|
81
|
+
- `--characters-to-anonymize INTEGER`: Number of characters to process in each chunk (default: `100000`).
|
|
82
|
+
- `--prompt-name [simple|detailed]`: The prompt template to use (default: `detailed`).
|
|
83
|
+
- `--model-name TEXT`: The language model to use.
|
|
84
|
+
- `--anonymized-entities PATH`: Path to a file with a list of entities to anonymize.
|
|
85
|
+
|
|
86
|
+
**Models**:
|
|
87
|
+
You can use any of the predefined models below, or specify a new model using the format `"provider/model-name"`.
|
|
88
|
+
For example: `--model-name "google/gemini-flash-latest"`.
|
|
89
|
+
|
|
90
|
+
- **Google**: `gemini-2.5-pro`, `gemini-2.5-flash` (default), `gemini-2.5-flash-lite`.
|
|
91
|
+
- **Ollama**: `gemma:7b`, `phi4-mini`.
|
|
92
|
+
- **Hugging Face**: `openai/gpt-oss-20b`, `mistralai/Mistral-7B-Instruct-v0.1`, `HuggingFaceH4/zephyr-7b-beta`.
|
|
93
|
+
- **OpenRouter**: `openai/gpt-4o`, `google/gemini-pro`.
|
|
94
|
+
- **OpenAI**: `gpt-4o`, `gpt-5`.
|
|
95
|
+
- **Anthropic**: `claude-4-sonet`, `claude-4.5-sonet`.
|
|
96
|
+
|
|
97
|
+
### Examples
|
|
98
|
+
|
|
99
|
+
**Basic anonymization with the default model (Google)**:
|
|
100
|
+
```bash
|
|
101
|
+
pdf-anonymizer run document.pdf
|
|
102
|
+
```
|
|
103
|
+
|
|
104
|
+
**A new model (Google) and a simple prompt**:
|
|
105
|
+
```bash
|
|
106
|
+
pdf-anonymizer run notes.md --model-name "google/gemini-flash-latest" --prompt-name simple
|
|
107
|
+
```
|
|
108
|
+
|
|
109
|
+
**Using an OpenRouter model**:
|
|
110
|
+
```bash
|
|
111
|
+
pdf-anonymizer run report.pdf --model-name "openai/gpt-4o"
|
|
112
|
+
```
|
|
113
|
+
|
|
114
|
+
### Deanonymize
|
|
115
|
+
|
|
116
|
+
The `deanonymize` command reverts anonymization using a mapping file.
|
|
117
|
+
|
|
118
|
+
```bash
|
|
119
|
+
pdf-anonymizer deanonymize ANONYMIZED_FILE MAPPING_FILE
|
|
120
|
+
```
|
|
121
|
+
|
|
122
|
+
**Arguments**:
|
|
123
|
+
- `ANONYMIZED_FILE`: Path to the anonymized text file.
|
|
124
|
+
- `MAPPING_FILE`: Path to the JSON mapping file.
|
|
125
|
+
|
|
126
|
+
**Example**:
|
|
127
|
+
```bash
|
|
128
|
+
pdf-anonymizer deanonymize \
|
|
129
|
+
data/anonymized/document.anonymized.md \
|
|
130
|
+
data/mappings/document.mapping.json
|
|
131
|
+
```
|
|
132
|
+
|
|
133
|
+
This will create a deanonymized version of the file at `data/deanonymized/document.deanonymized.md`.
|
|
@@ -1,100 +0,0 @@
|
|
|
1
|
-
Metadata-Version: 2.4
|
|
2
|
-
Name: pdf-anonymizer-cli
|
|
3
|
-
Version: 0.3.0
|
|
4
|
-
Summary: CLI for a tool to anonymize PDF, Markdown, and plain text files using LLMs.
|
|
5
|
-
Author-email: Leonid Ganeline <leo.gan.57@gmail.com>
|
|
6
|
-
License: MIT
|
|
7
|
-
Project-URL: repository, https://github.com/leo-gan/anonymizer
|
|
8
|
-
Requires-Python: >=3.10
|
|
9
|
-
Description-Content-Type: text/markdown
|
|
10
|
-
Requires-Dist: typer
|
|
11
|
-
Requires-Dist: rich
|
|
12
|
-
Requires-Dist: pdf-anonymizer-core
|
|
13
|
-
|
|
14
|
-
# PDF Anonymizer CLI
|
|
15
|
-
|
|
16
|
-
A command-line interface for anonymizing PDF, Markdown, and plain text files using LLMs.
|
|
17
|
-
|
|
18
|
-
## Installation
|
|
19
|
-
|
|
20
|
-
This project uses `uv` and is structured as a monorepo. The dependencies for the CLI and its core library are managed at the root of the project.
|
|
21
|
-
|
|
22
|
-
1. **Install `uv`**: Follow the [official installation instructions](https://astral.sh/docs/uv#installation).
|
|
23
|
-
2. **Install dependencies from the repository root**:
|
|
24
|
-
```bash
|
|
25
|
-
# From the repository root
|
|
26
|
-
uv sync
|
|
27
|
-
```
|
|
28
|
-
This installs the `pdf-anonymizer` executable.
|
|
29
|
-
|
|
30
|
-
## Environment Variables
|
|
31
|
-
|
|
32
|
-
The CLI will automatically load a `.env` file from the current directory or any parent directory. For consistency, it's recommended to place a single `.env` file at the root of the repository.
|
|
33
|
-
|
|
34
|
-
- `GOOGLE_API_KEY`: Required when using Google's Gemini models.
|
|
35
|
-
- `OLLAMA_HOST`: Optional, defaults to `http://localhost:11434` when using local Ollama models.
|
|
36
|
-
|
|
37
|
-
Example `.env` file:
|
|
38
|
-
```env
|
|
39
|
-
GOOGLE_API_KEY="YOUR_API_KEY_HERE"
|
|
40
|
-
```
|
|
41
|
-
|
|
42
|
-
## Usage
|
|
43
|
-
|
|
44
|
-
### Anonymize
|
|
45
|
-
|
|
46
|
-
The `run` command anonymizes one or more files.
|
|
47
|
-
|
|
48
|
-
```bash
|
|
49
|
-
pdf-anonymizer run FILE_PATH [FILE_PATH ...] \
|
|
50
|
-
[--characters-to-anonymize INTEGER] \
|
|
51
|
-
[--prompt-name {simple|detailed}] \
|
|
52
|
-
[--model-name TEXT] \
|
|
53
|
-
[--anonymized-entities PATH]
|
|
54
|
-
```
|
|
55
|
-
|
|
56
|
-
**Arguments**:
|
|
57
|
-
- `FILE_PATH`: Path to one or several PDF, Markdown, or text files for anonymization.
|
|
58
|
-
|
|
59
|
-
**Options**:
|
|
60
|
-
- `--characters-to-anonymize INTEGER`: Number of characters to process in each chunk (default: `100000`).
|
|
61
|
-
- `--prompt-name [simple|detailed]`: The prompt template to use (default: `detailed`).
|
|
62
|
-
- `--model-name TEXT`: The language model to use.
|
|
63
|
-
- `--anonymized-entities PATH`: Path to a file with a list of entities to anonymize.
|
|
64
|
-
|
|
65
|
-
**Models**:
|
|
66
|
-
- **Google**: `gemini-2.5-pro`, `gemini-2.5-flash` (default), `gemini-2.5-flash-lite`.
|
|
67
|
-
- **Ollama**: `gemma:7b`, `phi4-mini`.
|
|
68
|
-
|
|
69
|
-
### Examples
|
|
70
|
-
|
|
71
|
-
**Basic anonymization**:
|
|
72
|
-
```bash
|
|
73
|
-
pdf-anonymizer run document.pdf
|
|
74
|
-
```
|
|
75
|
-
|
|
76
|
-
**Custom model and prompt**:
|
|
77
|
-
```bash
|
|
78
|
-
pdf-anonymizer run notes.md --model-name phi4-mini --prompt-name simple
|
|
79
|
-
```
|
|
80
|
-
|
|
81
|
-
### Deanonymize
|
|
82
|
-
|
|
83
|
-
The `deanonymize` command reverts anonymization using a mapping file.
|
|
84
|
-
|
|
85
|
-
```bash
|
|
86
|
-
pdf-anonymizer deanonymize ANONYMIZED_FILE MAPPING_FILE
|
|
87
|
-
```
|
|
88
|
-
|
|
89
|
-
**Arguments**:
|
|
90
|
-
- `ANONYMIZED_FILE`: Path to the anonymized text file.
|
|
91
|
-
- `MAPPING_FILE`: Path to the JSON mapping file.
|
|
92
|
-
|
|
93
|
-
**Example**:
|
|
94
|
-
```bash
|
|
95
|
-
pdf-anonymizer deanonymize \
|
|
96
|
-
data/anonymized/document.anonymized.md \
|
|
97
|
-
data/mappings/document.mapping.json
|
|
98
|
-
```
|
|
99
|
-
|
|
100
|
-
This will create a deanonymized version of the file at `data/deanonymized/document.deanonymized.md`.
|
|
@@ -1,87 +0,0 @@
|
|
|
1
|
-
# PDF Anonymizer CLI
|
|
2
|
-
|
|
3
|
-
A command-line interface for anonymizing PDF, Markdown, and plain text files using LLMs.
|
|
4
|
-
|
|
5
|
-
## Installation
|
|
6
|
-
|
|
7
|
-
This project uses `uv` and is structured as a monorepo. The dependencies for the CLI and its core library are managed at the root of the project.
|
|
8
|
-
|
|
9
|
-
1. **Install `uv`**: Follow the [official installation instructions](https://astral.sh/docs/uv#installation).
|
|
10
|
-
2. **Install dependencies from the repository root**:
|
|
11
|
-
```bash
|
|
12
|
-
# From the repository root
|
|
13
|
-
uv sync
|
|
14
|
-
```
|
|
15
|
-
This installs the `pdf-anonymizer` executable.
|
|
16
|
-
|
|
17
|
-
## Environment Variables
|
|
18
|
-
|
|
19
|
-
The CLI will automatically load a `.env` file from the current directory or any parent directory. For consistency, it's recommended to place a single `.env` file at the root of the repository.
|
|
20
|
-
|
|
21
|
-
- `GOOGLE_API_KEY`: Required when using Google's Gemini models.
|
|
22
|
-
- `OLLAMA_HOST`: Optional, defaults to `http://localhost:11434` when using local Ollama models.
|
|
23
|
-
|
|
24
|
-
Example `.env` file:
|
|
25
|
-
```env
|
|
26
|
-
GOOGLE_API_KEY="YOUR_API_KEY_HERE"
|
|
27
|
-
```
|
|
28
|
-
|
|
29
|
-
## Usage
|
|
30
|
-
|
|
31
|
-
### Anonymize
|
|
32
|
-
|
|
33
|
-
The `run` command anonymizes one or more files.
|
|
34
|
-
|
|
35
|
-
```bash
|
|
36
|
-
pdf-anonymizer run FILE_PATH [FILE_PATH ...] \
|
|
37
|
-
[--characters-to-anonymize INTEGER] \
|
|
38
|
-
[--prompt-name {simple|detailed}] \
|
|
39
|
-
[--model-name TEXT] \
|
|
40
|
-
[--anonymized-entities PATH]
|
|
41
|
-
```
|
|
42
|
-
|
|
43
|
-
**Arguments**:
|
|
44
|
-
- `FILE_PATH`: Path to one or several PDF, Markdown, or text files for anonymization.
|
|
45
|
-
|
|
46
|
-
**Options**:
|
|
47
|
-
- `--characters-to-anonymize INTEGER`: Number of characters to process in each chunk (default: `100000`).
|
|
48
|
-
- `--prompt-name [simple|detailed]`: The prompt template to use (default: `detailed`).
|
|
49
|
-
- `--model-name TEXT`: The language model to use.
|
|
50
|
-
- `--anonymized-entities PATH`: Path to a file with a list of entities to anonymize.
|
|
51
|
-
|
|
52
|
-
**Models**:
|
|
53
|
-
- **Google**: `gemini-2.5-pro`, `gemini-2.5-flash` (default), `gemini-2.5-flash-lite`.
|
|
54
|
-
- **Ollama**: `gemma:7b`, `phi4-mini`.
|
|
55
|
-
|
|
56
|
-
### Examples
|
|
57
|
-
|
|
58
|
-
**Basic anonymization**:
|
|
59
|
-
```bash
|
|
60
|
-
pdf-anonymizer run document.pdf
|
|
61
|
-
```
|
|
62
|
-
|
|
63
|
-
**Custom model and prompt**:
|
|
64
|
-
```bash
|
|
65
|
-
pdf-anonymizer run notes.md --model-name phi4-mini --prompt-name simple
|
|
66
|
-
```
|
|
67
|
-
|
|
68
|
-
### Deanonymize
|
|
69
|
-
|
|
70
|
-
The `deanonymize` command reverts anonymization using a mapping file.
|
|
71
|
-
|
|
72
|
-
```bash
|
|
73
|
-
pdf-anonymizer deanonymize ANONYMIZED_FILE MAPPING_FILE
|
|
74
|
-
```
|
|
75
|
-
|
|
76
|
-
**Arguments**:
|
|
77
|
-
- `ANONYMIZED_FILE`: Path to the anonymized text file.
|
|
78
|
-
- `MAPPING_FILE`: Path to the JSON mapping file.
|
|
79
|
-
|
|
80
|
-
**Example**:
|
|
81
|
-
```bash
|
|
82
|
-
pdf-anonymizer deanonymize \
|
|
83
|
-
data/anonymized/document.anonymized.md \
|
|
84
|
-
data/mappings/document.mapping.json
|
|
85
|
-
```
|
|
86
|
-
|
|
87
|
-
This will create a deanonymized version of the file at `data/deanonymized/document.deanonymized.md`.
|
|
@@ -1,100 +0,0 @@
|
|
|
1
|
-
Metadata-Version: 2.4
|
|
2
|
-
Name: pdf-anonymizer-cli
|
|
3
|
-
Version: 0.3.0
|
|
4
|
-
Summary: CLI for a tool to anonymize PDF, Markdown, and plain text files using LLMs.
|
|
5
|
-
Author-email: Leonid Ganeline <leo.gan.57@gmail.com>
|
|
6
|
-
License: MIT
|
|
7
|
-
Project-URL: repository, https://github.com/leo-gan/anonymizer
|
|
8
|
-
Requires-Python: >=3.10
|
|
9
|
-
Description-Content-Type: text/markdown
|
|
10
|
-
Requires-Dist: typer
|
|
11
|
-
Requires-Dist: rich
|
|
12
|
-
Requires-Dist: pdf-anonymizer-core
|
|
13
|
-
|
|
14
|
-
# PDF Anonymizer CLI
|
|
15
|
-
|
|
16
|
-
A command-line interface for anonymizing PDF, Markdown, and plain text files using LLMs.
|
|
17
|
-
|
|
18
|
-
## Installation
|
|
19
|
-
|
|
20
|
-
This project uses `uv` and is structured as a monorepo. The dependencies for the CLI and its core library are managed at the root of the project.
|
|
21
|
-
|
|
22
|
-
1. **Install `uv`**: Follow the [official installation instructions](https://astral.sh/docs/uv#installation).
|
|
23
|
-
2. **Install dependencies from the repository root**:
|
|
24
|
-
```bash
|
|
25
|
-
# From the repository root
|
|
26
|
-
uv sync
|
|
27
|
-
```
|
|
28
|
-
This installs the `pdf-anonymizer` executable.
|
|
29
|
-
|
|
30
|
-
## Environment Variables
|
|
31
|
-
|
|
32
|
-
The CLI will automatically load a `.env` file from the current directory or any parent directory. For consistency, it's recommended to place a single `.env` file at the root of the repository.
|
|
33
|
-
|
|
34
|
-
- `GOOGLE_API_KEY`: Required when using Google's Gemini models.
|
|
35
|
-
- `OLLAMA_HOST`: Optional, defaults to `http://localhost:11434` when using local Ollama models.
|
|
36
|
-
|
|
37
|
-
Example `.env` file:
|
|
38
|
-
```env
|
|
39
|
-
GOOGLE_API_KEY="YOUR_API_KEY_HERE"
|
|
40
|
-
```
|
|
41
|
-
|
|
42
|
-
## Usage
|
|
43
|
-
|
|
44
|
-
### Anonymize
|
|
45
|
-
|
|
46
|
-
The `run` command anonymizes one or more files.
|
|
47
|
-
|
|
48
|
-
```bash
|
|
49
|
-
pdf-anonymizer run FILE_PATH [FILE_PATH ...] \
|
|
50
|
-
[--characters-to-anonymize INTEGER] \
|
|
51
|
-
[--prompt-name {simple|detailed}] \
|
|
52
|
-
[--model-name TEXT] \
|
|
53
|
-
[--anonymized-entities PATH]
|
|
54
|
-
```
|
|
55
|
-
|
|
56
|
-
**Arguments**:
|
|
57
|
-
- `FILE_PATH`: Path to one or several PDF, Markdown, or text files for anonymization.
|
|
58
|
-
|
|
59
|
-
**Options**:
|
|
60
|
-
- `--characters-to-anonymize INTEGER`: Number of characters to process in each chunk (default: `100000`).
|
|
61
|
-
- `--prompt-name [simple|detailed]`: The prompt template to use (default: `detailed`).
|
|
62
|
-
- `--model-name TEXT`: The language model to use.
|
|
63
|
-
- `--anonymized-entities PATH`: Path to a file with a list of entities to anonymize.
|
|
64
|
-
|
|
65
|
-
**Models**:
|
|
66
|
-
- **Google**: `gemini-2.5-pro`, `gemini-2.5-flash` (default), `gemini-2.5-flash-lite`.
|
|
67
|
-
- **Ollama**: `gemma:7b`, `phi4-mini`.
|
|
68
|
-
|
|
69
|
-
### Examples
|
|
70
|
-
|
|
71
|
-
**Basic anonymization**:
|
|
72
|
-
```bash
|
|
73
|
-
pdf-anonymizer run document.pdf
|
|
74
|
-
```
|
|
75
|
-
|
|
76
|
-
**Custom model and prompt**:
|
|
77
|
-
```bash
|
|
78
|
-
pdf-anonymizer run notes.md --model-name phi4-mini --prompt-name simple
|
|
79
|
-
```
|
|
80
|
-
|
|
81
|
-
### Deanonymize
|
|
82
|
-
|
|
83
|
-
The `deanonymize` command reverts anonymization using a mapping file.
|
|
84
|
-
|
|
85
|
-
```bash
|
|
86
|
-
pdf-anonymizer deanonymize ANONYMIZED_FILE MAPPING_FILE
|
|
87
|
-
```
|
|
88
|
-
|
|
89
|
-
**Arguments**:
|
|
90
|
-
- `ANONYMIZED_FILE`: Path to the anonymized text file.
|
|
91
|
-
- `MAPPING_FILE`: Path to the JSON mapping file.
|
|
92
|
-
|
|
93
|
-
**Example**:
|
|
94
|
-
```bash
|
|
95
|
-
pdf-anonymizer deanonymize \
|
|
96
|
-
data/anonymized/document.anonymized.md \
|
|
97
|
-
data/mappings/document.mapping.json
|
|
98
|
-
```
|
|
99
|
-
|
|
100
|
-
This will create a deanonymized version of the file at `data/deanonymized/document.deanonymized.md`.
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
{pdf_anonymizer_cli-0.3.0 → pdf_anonymizer_cli-0.3.2}/src/pdf_anonymizer_cli.egg-info/SOURCES.txt
RENAMED
|
File without changes
|
|
File without changes
|
|
File without changes
|
{pdf_anonymizer_cli-0.3.0 → pdf_anonymizer_cli-0.3.2}/src/pdf_anonymizer_cli.egg-info/requires.txt
RENAMED
|
File without changes
|
{pdf_anonymizer_cli-0.3.0 → pdf_anonymizer_cli-0.3.2}/src/pdf_anonymizer_cli.egg-info/top_level.txt
RENAMED
|
File without changes
|