pdf-anonymizer-cli 0.3.0__tar.gz → 0.3.2__tar.gz

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -0,0 +1,133 @@
1
+ Metadata-Version: 2.4
2
+ Name: pdf-anonymizer-cli
3
+ Version: 0.3.2
4
+ Summary: CLI for a tool to anonymize PDF, Markdown, and plain text files using LLMs.
5
+ Author-email: Leonid Ganeline <leo.gan.57@gmail.com>
6
+ License: MIT
7
+ Project-URL: repository, https://github.com/leo-gan/anonymizer
8
+ Requires-Python: >=3.10
9
+ Description-Content-Type: text/markdown
10
+ Requires-Dist: typer
11
+ Requires-Dist: rich
12
+ Requires-Dist: pdf-anonymizer-core
13
+
14
+ # 🦉🫥 PDF Anonymizer CLI
15
+
16
+ A command-line interface for anonymizing PDF, Markdown, and plain text files using LLMs.
17
+
18
+ - **High-Quality Anonymization**: Leverages LLMs to identify and replace Personally Identifiable Information (PII) with high accuracy.
19
+ - **Large File Support**: Consistently anonymizes large files (tested up to 1GB).
20
+ - **Multi-Provider & Cost-Effective**: Free to use with local [Ollama](https://ollama.com/) models. It also supports major providers like [OpenAI](https://openai.com/), [Anthropic](https://www.anthropic.com/), [Google](https://ai.google.com/), [Hugging Face](https://huggingface.co/), and [OpenRouter](https://openrouter.ai/).
21
+ - **Reversible**: Supports deanonymization to recover original data when needed.
22
+ - **Multi-Format**: Works with PDF, Markdown, and plain text files.
23
+
24
+
25
+ ## Installation
26
+
27
+ Install the CLI with your favorite package manager. To use a specific LLM provider, you must install the corresponding extra.
28
+
29
+ - **Google**: `pip install "pdf-anonymizer-cli[google]"`
30
+ - **Ollama**: `pip install "pdf-anonymizer-cli[ollama]"`
31
+ - **Hugging Face**: `pip install "pdf-anonymizer-cli[huggingface]"`
32
+ - **OpenRouter**: `pip install "pdf-anonymizer-cli[openrouter]"`
33
+ - **OpenAI**: `pip install "pdf-anonymizer-cli[openai]"`
34
+ - **Anthropic**: `pip install "pdf-anonymizer-cli[anthropic]"`
35
+
36
+ You can also install multiple extras at once:
37
+
38
+ ```bash
39
+ pip install "pdf-anonymizer-cli[google,openrouter]"
40
+ ```
41
+
42
+ This installs the `pdf-anonymizer` executable.
43
+
44
+
45
+ ## Environment Variables
46
+
47
+ The CLI will automatically load a `.env` file from the current directory or any parent directory. For consistency, it's recommended to place a single `.env` file at the root of the repository.
48
+
49
+ - `GOOGLE_API_KEY`: Required when using Google models.
50
+ - `HUGGING_FACE_TOKEN`: Required when using Hugging Face models. You can get a token from [here](https://huggingface.co/docs/hub/security-tokens).
51
+ - `OPENROUTER_API_KEY`: Required when using OpenRouter models.
52
+ - `OPENAI_API_KEY`: Required when using OpenAI models.
53
+ - `ANTHROPIC_API_KEY`: Required when using Anthropic models.
54
+ - `OLLAMA_HOST`: Optional, defaults to `http://localhost:11434` when using Ollama models.
55
+
56
+ Example `.env` file:
57
+ ```env
58
+ GOOGLE_API_KEY="YOUR_API_KEY_HERE"
59
+ HUGGING_FACE_TOKEN="YOUR_HF_TOKEN_HERE"
60
+ OPENROUTER_API_KEY="YOUR_OPENROUTER_KEY"
61
+ ```
62
+
63
+ ## Usage
64
+
65
+ ### Anonymize
66
+
67
+ The `run` command anonymizes one or more files.
68
+
69
+ ```bash
70
+ pdf-anonymizer run FILE_PATH [FILE_PATH ...] \
71
+ [--characters-to-anonymize INTEGER] \
72
+ [--prompt-name {simple|detailed}] \
73
+ [--model-name TEXT] \
74
+ [--anonymized-entities PATH]
75
+ ```
76
+
77
+ **Arguments**:
78
+ - `FILE_PATH`: Path to one or several PDF, Markdown, or text files for anonymization.
79
+
80
+ **Options**:
81
+ - `--characters-to-anonymize INTEGER`: Number of characters to process in each chunk (default: `100000`).
82
+ - `--prompt-name [simple|detailed]`: The prompt template to use (default: `detailed`).
83
+ - `--model-name TEXT`: The language model to use.
84
+ - `--anonymized-entities PATH`: Path to a file with a list of entities to anonymize.
85
+
86
+ **Models**:
87
+ You can use any of the predefined models below, or specify a new model using the format `"provider/model-name"`.
88
+ For example: `--model-name "google/gemini-flash-latest"`.
89
+
90
+ - **Google**: `gemini-2.5-pro`, `gemini-2.5-flash` (default), `gemini-2.5-flash-lite`.
91
+ - **Ollama**: `gemma:7b`, `phi4-mini`.
92
+ - **Hugging Face**: `openai/gpt-oss-20b`, `mistralai/Mistral-7B-Instruct-v0.1`, `HuggingFaceH4/zephyr-7b-beta`.
93
+ - **OpenRouter**: `openai/gpt-4o`, `google/gemini-pro`.
94
+ - **OpenAI**: `gpt-4o`, `gpt-5`.
95
+ - **Anthropic**: `claude-4-sonet`, `claude-4.5-sonet`.
96
+
97
+ ### Examples
98
+
99
+ **Basic anonymization with the default model (Google)**:
100
+ ```bash
101
+ pdf-anonymizer run document.pdf
102
+ ```
103
+
104
+ **A new model (Google) and a simple prompt**:
105
+ ```bash
106
+ pdf-anonymizer run notes.md --model-name "google/gemini-flash-latest" --prompt-name simple
107
+ ```
108
+
109
+ **Using an OpenRouter model**:
110
+ ```bash
111
+ pdf-anonymizer run report.pdf --model-name "openai/gpt-4o"
112
+ ```
113
+
114
+ ### Deanonymize
115
+
116
+ The `deanonymize` command reverts anonymization using a mapping file.
117
+
118
+ ```bash
119
+ pdf-anonymizer deanonymize ANONYMIZED_FILE MAPPING_FILE
120
+ ```
121
+
122
+ **Arguments**:
123
+ - `ANONYMIZED_FILE`: Path to the anonymized text file.
124
+ - `MAPPING_FILE`: Path to the JSON mapping file.
125
+
126
+ **Example**:
127
+ ```bash
128
+ pdf-anonymizer deanonymize \
129
+ data/anonymized/document.anonymized.md \
130
+ data/mappings/document.mapping.json
131
+ ```
132
+
133
+ This will create a deanonymized version of the file at `data/deanonymized/document.deanonymized.md`.
@@ -0,0 +1,120 @@
1
+ # 🦉🫥 PDF Anonymizer CLI
2
+
3
+ A command-line interface for anonymizing PDF, Markdown, and plain text files using LLMs.
4
+
5
+ - **High-Quality Anonymization**: Leverages LLMs to identify and replace Personally Identifiable Information (PII) with high accuracy.
6
+ - **Large File Support**: Consistently anonymizes large files (tested up to 1GB).
7
+ - **Multi-Provider & Cost-Effective**: Free to use with local [Ollama](https://ollama.com/) models. It also supports major providers like [OpenAI](https://openai.com/), [Anthropic](https://www.anthropic.com/), [Google](https://ai.google.com/), [Hugging Face](https://huggingface.co/), and [OpenRouter](https://openrouter.ai/).
8
+ - **Reversible**: Supports deanonymization to recover original data when needed.
9
+ - **Multi-Format**: Works with PDF, Markdown, and plain text files.
10
+
11
+
12
+ ## Installation
13
+
14
+ Install the CLI with your favorite package manager. To use a specific LLM provider, you must install the corresponding extra.
15
+
16
+ - **Google**: `pip install "pdf-anonymizer-cli[google]"`
17
+ - **Ollama**: `pip install "pdf-anonymizer-cli[ollama]"`
18
+ - **Hugging Face**: `pip install "pdf-anonymizer-cli[huggingface]"`
19
+ - **OpenRouter**: `pip install "pdf-anonymizer-cli[openrouter]"`
20
+ - **OpenAI**: `pip install "pdf-anonymizer-cli[openai]"`
21
+ - **Anthropic**: `pip install "pdf-anonymizer-cli[anthropic]"`
22
+
23
+ You can also install multiple extras at once:
24
+
25
+ ```bash
26
+ pip install "pdf-anonymizer-cli[google,openrouter]"
27
+ ```
28
+
29
+ This installs the `pdf-anonymizer` executable.
30
+
31
+
32
+ ## Environment Variables
33
+
34
+ The CLI will automatically load a `.env` file from the current directory or any parent directory. For consistency, it's recommended to place a single `.env` file at the root of the repository.
35
+
36
+ - `GOOGLE_API_KEY`: Required when using Google models.
37
+ - `HUGGING_FACE_TOKEN`: Required when using Hugging Face models. You can get a token from [here](https://huggingface.co/docs/hub/security-tokens).
38
+ - `OPENROUTER_API_KEY`: Required when using OpenRouter models.
39
+ - `OPENAI_API_KEY`: Required when using OpenAI models.
40
+ - `ANTHROPIC_API_KEY`: Required when using Anthropic models.
41
+ - `OLLAMA_HOST`: Optional, defaults to `http://localhost:11434` when using Ollama models.
42
+
43
+ Example `.env` file:
44
+ ```env
45
+ GOOGLE_API_KEY="YOUR_API_KEY_HERE"
46
+ HUGGING_FACE_TOKEN="YOUR_HF_TOKEN_HERE"
47
+ OPENROUTER_API_KEY="YOUR_OPENROUTER_KEY"
48
+ ```
49
+
50
+ ## Usage
51
+
52
+ ### Anonymize
53
+
54
+ The `run` command anonymizes one or more files.
55
+
56
+ ```bash
57
+ pdf-anonymizer run FILE_PATH [FILE_PATH ...] \
58
+ [--characters-to-anonymize INTEGER] \
59
+ [--prompt-name {simple|detailed}] \
60
+ [--model-name TEXT] \
61
+ [--anonymized-entities PATH]
62
+ ```
63
+
64
+ **Arguments**:
65
+ - `FILE_PATH`: Path to one or several PDF, Markdown, or text files for anonymization.
66
+
67
+ **Options**:
68
+ - `--characters-to-anonymize INTEGER`: Number of characters to process in each chunk (default: `100000`).
69
+ - `--prompt-name [simple|detailed]`: The prompt template to use (default: `detailed`).
70
+ - `--model-name TEXT`: The language model to use.
71
+ - `--anonymized-entities PATH`: Path to a file with a list of entities to anonymize.
72
+
73
+ **Models**:
74
+ You can use any of the predefined models below, or specify a new model using the format `"provider/model-name"`.
75
+ For example: `--model-name "google/gemini-flash-latest"`.
76
+
77
+ - **Google**: `gemini-2.5-pro`, `gemini-2.5-flash` (default), `gemini-2.5-flash-lite`.
78
+ - **Ollama**: `gemma:7b`, `phi4-mini`.
79
+ - **Hugging Face**: `openai/gpt-oss-20b`, `mistralai/Mistral-7B-Instruct-v0.1`, `HuggingFaceH4/zephyr-7b-beta`.
80
+ - **OpenRouter**: `openai/gpt-4o`, `google/gemini-pro`.
81
+ - **OpenAI**: `gpt-4o`, `gpt-5`.
82
+ - **Anthropic**: `claude-4-sonet`, `claude-4.5-sonet`.
83
+
84
+ ### Examples
85
+
86
+ **Basic anonymization with the default model (Google)**:
87
+ ```bash
88
+ pdf-anonymizer run document.pdf
89
+ ```
90
+
91
+ **A new model (Google) and a simple prompt**:
92
+ ```bash
93
+ pdf-anonymizer run notes.md --model-name "google/gemini-flash-latest" --prompt-name simple
94
+ ```
95
+
96
+ **Using an OpenRouter model**:
97
+ ```bash
98
+ pdf-anonymizer run report.pdf --model-name "openai/gpt-4o"
99
+ ```
100
+
101
+ ### Deanonymize
102
+
103
+ The `deanonymize` command reverts anonymization using a mapping file.
104
+
105
+ ```bash
106
+ pdf-anonymizer deanonymize ANONYMIZED_FILE MAPPING_FILE
107
+ ```
108
+
109
+ **Arguments**:
110
+ - `ANONYMIZED_FILE`: Path to the anonymized text file.
111
+ - `MAPPING_FILE`: Path to the JSON mapping file.
112
+
113
+ **Example**:
114
+ ```bash
115
+ pdf-anonymizer deanonymize \
116
+ data/anonymized/document.anonymized.md \
117
+ data/mappings/document.mapping.json
118
+ ```
119
+
120
+ This will create a deanonymized version of the file at `data/deanonymized/document.deanonymized.md`.
@@ -1,6 +1,6 @@
1
1
  [project]
2
2
  name = "pdf-anonymizer-cli"
3
- version = "0.3.0"
3
+ version = "0.3.2"
4
4
  description = "CLI for a tool to anonymize PDF, Markdown, and plain text files using LLMs."
5
5
  authors = [{ name = "Leonid Ganeline", email = "leo.gan.57@gmail.com" }]
6
6
  license = { text = "MIT" }
@@ -10,10 +10,9 @@ from pdf_anonymizer_core.conf import (
10
10
  DEFAULT_CHARACTERS_TO_ANONYMIZE,
11
11
  DEFAULT_MODEL_NAME,
12
12
  DEFAULT_PROMPT_NAME,
13
- ModelName,
14
- ModelProvider,
15
13
  PromptEnum,
16
14
  get_enum_value,
15
+ get_provider_and_model_name,
17
16
  )
18
17
  from pdf_anonymizer_core.core import anonymize_file
19
18
  from pdf_anonymizer_core.prompts import detailed, simple
@@ -66,12 +65,11 @@ def run(
66
65
  ),
67
66
  ] = get_enum_value(PromptEnum, DEFAULT_PROMPT_NAME),
68
67
  model_name: Annotated[
69
- ModelName,
68
+ str,
70
69
  typer.Option(
71
- help="The name of the model to use for anonymization.",
72
- case_sensitive=False,
70
+ help="The name of the model to use for anonymization (supports enum values or 'provider/model').",
73
71
  ),
74
- ] = get_enum_value(ModelName, DEFAULT_MODEL_NAME),
72
+ ] = DEFAULT_MODEL_NAME,
75
73
  anonymized_entities: Annotated[
76
74
  Optional[Path],
77
75
  typer.Option(
@@ -98,8 +96,9 @@ def run(
98
96
  """
99
97
  load_environment()
100
98
 
101
- if model_name.provider == ModelProvider.GOOGLE:
102
- if "gemini" in model_name.value and not os.getenv("GOOGLE_API_KEY"):
99
+ provider_name, _ = get_provider_and_model_name(model_name)
100
+ if provider_name == "google":
101
+ if "gemini" in model_name and not os.getenv("GOOGLE_API_KEY"):
103
102
  logging.error(
104
103
  "Error: GOOGLE_API_KEY not found. Please set it in the .env file."
105
104
  )
@@ -107,7 +106,7 @@ def run(
107
106
 
108
107
  logging.info(f" --file-paths: {file_paths}")
109
108
  logging.info(f" --characters-to-anonymize: {characters_to_anonymize}")
110
- logging.info(f" --model-name: {model_name.value}")
109
+ logging.info(f" --model-name: {model_name}")
111
110
 
112
111
  # Select the appropriate prompt template
113
112
  prompt_templates: Dict[str, str] = {
@@ -132,7 +131,7 @@ def run(
132
131
  str(file_path),
133
132
  characters_to_anonymize,
134
133
  prompt_template,
135
- model_name.value,
134
+ model_name,
136
135
  entities_to_anonymize,
137
136
  )
138
137
 
@@ -0,0 +1,133 @@
1
+ Metadata-Version: 2.4
2
+ Name: pdf-anonymizer-cli
3
+ Version: 0.3.2
4
+ Summary: CLI for a tool to anonymize PDF, Markdown, and plain text files using LLMs.
5
+ Author-email: Leonid Ganeline <leo.gan.57@gmail.com>
6
+ License: MIT
7
+ Project-URL: repository, https://github.com/leo-gan/anonymizer
8
+ Requires-Python: >=3.10
9
+ Description-Content-Type: text/markdown
10
+ Requires-Dist: typer
11
+ Requires-Dist: rich
12
+ Requires-Dist: pdf-anonymizer-core
13
+
14
+ # 🦉🫥 PDF Anonymizer CLI
15
+
16
+ A command-line interface for anonymizing PDF, Markdown, and plain text files using LLMs.
17
+
18
+ - **High-Quality Anonymization**: Leverages LLMs to identify and replace Personally Identifiable Information (PII) with high accuracy.
19
+ - **Large File Support**: Consistently anonymizes large files (tested up to 1GB).
20
+ - **Multi-Provider & Cost-Effective**: Free to use with local [Ollama](https://ollama.com/) models. It also supports major providers like [OpenAI](https://openai.com/), [Anthropic](https://www.anthropic.com/), [Google](https://ai.google.com/), [Hugging Face](https://huggingface.co/), and [OpenRouter](https://openrouter.ai/).
21
+ - **Reversible**: Supports deanonymization to recover original data when needed.
22
+ - **Multi-Format**: Works with PDF, Markdown, and plain text files.
23
+
24
+
25
+ ## Installation
26
+
27
+ Install the CLI with your favorite package manager. To use a specific LLM provider, you must install the corresponding extra.
28
+
29
+ - **Google**: `pip install "pdf-anonymizer-cli[google]"`
30
+ - **Ollama**: `pip install "pdf-anonymizer-cli[ollama]"`
31
+ - **Hugging Face**: `pip install "pdf-anonymizer-cli[huggingface]"`
32
+ - **OpenRouter**: `pip install "pdf-anonymizer-cli[openrouter]"`
33
+ - **OpenAI**: `pip install "pdf-anonymizer-cli[openai]"`
34
+ - **Anthropic**: `pip install "pdf-anonymizer-cli[anthropic]"`
35
+
36
+ You can also install multiple extras at once:
37
+
38
+ ```bash
39
+ pip install "pdf-anonymizer-cli[google,openrouter]"
40
+ ```
41
+
42
+ This installs the `pdf-anonymizer` executable.
43
+
44
+
45
+ ## Environment Variables
46
+
47
+ The CLI will automatically load a `.env` file from the current directory or any parent directory. For consistency, it's recommended to place a single `.env` file at the root of the repository.
48
+
49
+ - `GOOGLE_API_KEY`: Required when using Google models.
50
+ - `HUGGING_FACE_TOKEN`: Required when using Hugging Face models. You can get a token from [here](https://huggingface.co/docs/hub/security-tokens).
51
+ - `OPENROUTER_API_KEY`: Required when using OpenRouter models.
52
+ - `OPENAI_API_KEY`: Required when using OpenAI models.
53
+ - `ANTHROPIC_API_KEY`: Required when using Anthropic models.
54
+ - `OLLAMA_HOST`: Optional, defaults to `http://localhost:11434` when using Ollama models.
55
+
56
+ Example `.env` file:
57
+ ```env
58
+ GOOGLE_API_KEY="YOUR_API_KEY_HERE"
59
+ HUGGING_FACE_TOKEN="YOUR_HF_TOKEN_HERE"
60
+ OPENROUTER_API_KEY="YOUR_OPENROUTER_KEY"
61
+ ```
62
+
63
+ ## Usage
64
+
65
+ ### Anonymize
66
+
67
+ The `run` command anonymizes one or more files.
68
+
69
+ ```bash
70
+ pdf-anonymizer run FILE_PATH [FILE_PATH ...] \
71
+ [--characters-to-anonymize INTEGER] \
72
+ [--prompt-name {simple|detailed}] \
73
+ [--model-name TEXT] \
74
+ [--anonymized-entities PATH]
75
+ ```
76
+
77
+ **Arguments**:
78
+ - `FILE_PATH`: Path to one or several PDF, Markdown, or text files for anonymization.
79
+
80
+ **Options**:
81
+ - `--characters-to-anonymize INTEGER`: Number of characters to process in each chunk (default: `100000`).
82
+ - `--prompt-name [simple|detailed]`: The prompt template to use (default: `detailed`).
83
+ - `--model-name TEXT`: The language model to use.
84
+ - `--anonymized-entities PATH`: Path to a file with a list of entities to anonymize.
85
+
86
+ **Models**:
87
+ You can use any of the predefined models below, or specify a new model using the format `"provider/model-name"`.
88
+ For example: `--model-name "google/gemini-flash-latest"`.
89
+
90
+ - **Google**: `gemini-2.5-pro`, `gemini-2.5-flash` (default), `gemini-2.5-flash-lite`.
91
+ - **Ollama**: `gemma:7b`, `phi4-mini`.
92
+ - **Hugging Face**: `openai/gpt-oss-20b`, `mistralai/Mistral-7B-Instruct-v0.1`, `HuggingFaceH4/zephyr-7b-beta`.
93
+ - **OpenRouter**: `openai/gpt-4o`, `google/gemini-pro`.
94
+ - **OpenAI**: `gpt-4o`, `gpt-5`.
95
+ - **Anthropic**: `claude-4-sonet`, `claude-4.5-sonet`.
96
+
97
+ ### Examples
98
+
99
+ **Basic anonymization with the default model (Google)**:
100
+ ```bash
101
+ pdf-anonymizer run document.pdf
102
+ ```
103
+
104
+ **A new model (Google) and a simple prompt**:
105
+ ```bash
106
+ pdf-anonymizer run notes.md --model-name "google/gemini-flash-latest" --prompt-name simple
107
+ ```
108
+
109
+ **Using an OpenRouter model**:
110
+ ```bash
111
+ pdf-anonymizer run report.pdf --model-name "openai/gpt-4o"
112
+ ```
113
+
114
+ ### Deanonymize
115
+
116
+ The `deanonymize` command reverts anonymization using a mapping file.
117
+
118
+ ```bash
119
+ pdf-anonymizer deanonymize ANONYMIZED_FILE MAPPING_FILE
120
+ ```
121
+
122
+ **Arguments**:
123
+ - `ANONYMIZED_FILE`: Path to the anonymized text file.
124
+ - `MAPPING_FILE`: Path to the JSON mapping file.
125
+
126
+ **Example**:
127
+ ```bash
128
+ pdf-anonymizer deanonymize \
129
+ data/anonymized/document.anonymized.md \
130
+ data/mappings/document.mapping.json
131
+ ```
132
+
133
+ This will create a deanonymized version of the file at `data/deanonymized/document.deanonymized.md`.
@@ -1,100 +0,0 @@
1
- Metadata-Version: 2.4
2
- Name: pdf-anonymizer-cli
3
- Version: 0.3.0
4
- Summary: CLI for a tool to anonymize PDF, Markdown, and plain text files using LLMs.
5
- Author-email: Leonid Ganeline <leo.gan.57@gmail.com>
6
- License: MIT
7
- Project-URL: repository, https://github.com/leo-gan/anonymizer
8
- Requires-Python: >=3.10
9
- Description-Content-Type: text/markdown
10
- Requires-Dist: typer
11
- Requires-Dist: rich
12
- Requires-Dist: pdf-anonymizer-core
13
-
14
- # PDF Anonymizer CLI
15
-
16
- A command-line interface for anonymizing PDF, Markdown, and plain text files using LLMs.
17
-
18
- ## Installation
19
-
20
- This project uses `uv` and is structured as a monorepo. The dependencies for the CLI and its core library are managed at the root of the project.
21
-
22
- 1. **Install `uv`**: Follow the [official installation instructions](https://astral.sh/docs/uv#installation).
23
- 2. **Install dependencies from the repository root**:
24
- ```bash
25
- # From the repository root
26
- uv sync
27
- ```
28
- This installs the `pdf-anonymizer` executable.
29
-
30
- ## Environment Variables
31
-
32
- The CLI will automatically load a `.env` file from the current directory or any parent directory. For consistency, it's recommended to place a single `.env` file at the root of the repository.
33
-
34
- - `GOOGLE_API_KEY`: Required when using Google's Gemini models.
35
- - `OLLAMA_HOST`: Optional, defaults to `http://localhost:11434` when using local Ollama models.
36
-
37
- Example `.env` file:
38
- ```env
39
- GOOGLE_API_KEY="YOUR_API_KEY_HERE"
40
- ```
41
-
42
- ## Usage
43
-
44
- ### Anonymize
45
-
46
- The `run` command anonymizes one or more files.
47
-
48
- ```bash
49
- pdf-anonymizer run FILE_PATH [FILE_PATH ...] \
50
- [--characters-to-anonymize INTEGER] \
51
- [--prompt-name {simple|detailed}] \
52
- [--model-name TEXT] \
53
- [--anonymized-entities PATH]
54
- ```
55
-
56
- **Arguments**:
57
- - `FILE_PATH`: Path to one or several PDF, Markdown, or text files for anonymization.
58
-
59
- **Options**:
60
- - `--characters-to-anonymize INTEGER`: Number of characters to process in each chunk (default: `100000`).
61
- - `--prompt-name [simple|detailed]`: The prompt template to use (default: `detailed`).
62
- - `--model-name TEXT`: The language model to use.
63
- - `--anonymized-entities PATH`: Path to a file with a list of entities to anonymize.
64
-
65
- **Models**:
66
- - **Google**: `gemini-2.5-pro`, `gemini-2.5-flash` (default), `gemini-2.5-flash-lite`.
67
- - **Ollama**: `gemma:7b`, `phi4-mini`.
68
-
69
- ### Examples
70
-
71
- **Basic anonymization**:
72
- ```bash
73
- pdf-anonymizer run document.pdf
74
- ```
75
-
76
- **Custom model and prompt**:
77
- ```bash
78
- pdf-anonymizer run notes.md --model-name phi4-mini --prompt-name simple
79
- ```
80
-
81
- ### Deanonymize
82
-
83
- The `deanonymize` command reverts anonymization using a mapping file.
84
-
85
- ```bash
86
- pdf-anonymizer deanonymize ANONYMIZED_FILE MAPPING_FILE
87
- ```
88
-
89
- **Arguments**:
90
- - `ANONYMIZED_FILE`: Path to the anonymized text file.
91
- - `MAPPING_FILE`: Path to the JSON mapping file.
92
-
93
- **Example**:
94
- ```bash
95
- pdf-anonymizer deanonymize \
96
- data/anonymized/document.anonymized.md \
97
- data/mappings/document.mapping.json
98
- ```
99
-
100
- This will create a deanonymized version of the file at `data/deanonymized/document.deanonymized.md`.
@@ -1,87 +0,0 @@
1
- # PDF Anonymizer CLI
2
-
3
- A command-line interface for anonymizing PDF, Markdown, and plain text files using LLMs.
4
-
5
- ## Installation
6
-
7
- This project uses `uv` and is structured as a monorepo. The dependencies for the CLI and its core library are managed at the root of the project.
8
-
9
- 1. **Install `uv`**: Follow the [official installation instructions](https://astral.sh/docs/uv#installation).
10
- 2. **Install dependencies from the repository root**:
11
- ```bash
12
- # From the repository root
13
- uv sync
14
- ```
15
- This installs the `pdf-anonymizer` executable.
16
-
17
- ## Environment Variables
18
-
19
- The CLI will automatically load a `.env` file from the current directory or any parent directory. For consistency, it's recommended to place a single `.env` file at the root of the repository.
20
-
21
- - `GOOGLE_API_KEY`: Required when using Google's Gemini models.
22
- - `OLLAMA_HOST`: Optional, defaults to `http://localhost:11434` when using local Ollama models.
23
-
24
- Example `.env` file:
25
- ```env
26
- GOOGLE_API_KEY="YOUR_API_KEY_HERE"
27
- ```
28
-
29
- ## Usage
30
-
31
- ### Anonymize
32
-
33
- The `run` command anonymizes one or more files.
34
-
35
- ```bash
36
- pdf-anonymizer run FILE_PATH [FILE_PATH ...] \
37
- [--characters-to-anonymize INTEGER] \
38
- [--prompt-name {simple|detailed}] \
39
- [--model-name TEXT] \
40
- [--anonymized-entities PATH]
41
- ```
42
-
43
- **Arguments**:
44
- - `FILE_PATH`: Path to one or several PDF, Markdown, or text files for anonymization.
45
-
46
- **Options**:
47
- - `--characters-to-anonymize INTEGER`: Number of characters to process in each chunk (default: `100000`).
48
- - `--prompt-name [simple|detailed]`: The prompt template to use (default: `detailed`).
49
- - `--model-name TEXT`: The language model to use.
50
- - `--anonymized-entities PATH`: Path to a file with a list of entities to anonymize.
51
-
52
- **Models**:
53
- - **Google**: `gemini-2.5-pro`, `gemini-2.5-flash` (default), `gemini-2.5-flash-lite`.
54
- - **Ollama**: `gemma:7b`, `phi4-mini`.
55
-
56
- ### Examples
57
-
58
- **Basic anonymization**:
59
- ```bash
60
- pdf-anonymizer run document.pdf
61
- ```
62
-
63
- **Custom model and prompt**:
64
- ```bash
65
- pdf-anonymizer run notes.md --model-name phi4-mini --prompt-name simple
66
- ```
67
-
68
- ### Deanonymize
69
-
70
- The `deanonymize` command reverts anonymization using a mapping file.
71
-
72
- ```bash
73
- pdf-anonymizer deanonymize ANONYMIZED_FILE MAPPING_FILE
74
- ```
75
-
76
- **Arguments**:
77
- - `ANONYMIZED_FILE`: Path to the anonymized text file.
78
- - `MAPPING_FILE`: Path to the JSON mapping file.
79
-
80
- **Example**:
81
- ```bash
82
- pdf-anonymizer deanonymize \
83
- data/anonymized/document.anonymized.md \
84
- data/mappings/document.mapping.json
85
- ```
86
-
87
- This will create a deanonymized version of the file at `data/deanonymized/document.deanonymized.md`.
@@ -1,100 +0,0 @@
1
- Metadata-Version: 2.4
2
- Name: pdf-anonymizer-cli
3
- Version: 0.3.0
4
- Summary: CLI for a tool to anonymize PDF, Markdown, and plain text files using LLMs.
5
- Author-email: Leonid Ganeline <leo.gan.57@gmail.com>
6
- License: MIT
7
- Project-URL: repository, https://github.com/leo-gan/anonymizer
8
- Requires-Python: >=3.10
9
- Description-Content-Type: text/markdown
10
- Requires-Dist: typer
11
- Requires-Dist: rich
12
- Requires-Dist: pdf-anonymizer-core
13
-
14
- # PDF Anonymizer CLI
15
-
16
- A command-line interface for anonymizing PDF, Markdown, and plain text files using LLMs.
17
-
18
- ## Installation
19
-
20
- This project uses `uv` and is structured as a monorepo. The dependencies for the CLI and its core library are managed at the root of the project.
21
-
22
- 1. **Install `uv`**: Follow the [official installation instructions](https://astral.sh/docs/uv#installation).
23
- 2. **Install dependencies from the repository root**:
24
- ```bash
25
- # From the repository root
26
- uv sync
27
- ```
28
- This installs the `pdf-anonymizer` executable.
29
-
30
- ## Environment Variables
31
-
32
- The CLI will automatically load a `.env` file from the current directory or any parent directory. For consistency, it's recommended to place a single `.env` file at the root of the repository.
33
-
34
- - `GOOGLE_API_KEY`: Required when using Google's Gemini models.
35
- - `OLLAMA_HOST`: Optional, defaults to `http://localhost:11434` when using local Ollama models.
36
-
37
- Example `.env` file:
38
- ```env
39
- GOOGLE_API_KEY="YOUR_API_KEY_HERE"
40
- ```
41
-
42
- ## Usage
43
-
44
- ### Anonymize
45
-
46
- The `run` command anonymizes one or more files.
47
-
48
- ```bash
49
- pdf-anonymizer run FILE_PATH [FILE_PATH ...] \
50
- [--characters-to-anonymize INTEGER] \
51
- [--prompt-name {simple|detailed}] \
52
- [--model-name TEXT] \
53
- [--anonymized-entities PATH]
54
- ```
55
-
56
- **Arguments**:
57
- - `FILE_PATH`: Path to one or several PDF, Markdown, or text files for anonymization.
58
-
59
- **Options**:
60
- - `--characters-to-anonymize INTEGER`: Number of characters to process in each chunk (default: `100000`).
61
- - `--prompt-name [simple|detailed]`: The prompt template to use (default: `detailed`).
62
- - `--model-name TEXT`: The language model to use.
63
- - `--anonymized-entities PATH`: Path to a file with a list of entities to anonymize.
64
-
65
- **Models**:
66
- - **Google**: `gemini-2.5-pro`, `gemini-2.5-flash` (default), `gemini-2.5-flash-lite`.
67
- - **Ollama**: `gemma:7b`, `phi4-mini`.
68
-
69
- ### Examples
70
-
71
- **Basic anonymization**:
72
- ```bash
73
- pdf-anonymizer run document.pdf
74
- ```
75
-
76
- **Custom model and prompt**:
77
- ```bash
78
- pdf-anonymizer run notes.md --model-name phi4-mini --prompt-name simple
79
- ```
80
-
81
- ### Deanonymize
82
-
83
- The `deanonymize` command reverts anonymization using a mapping file.
84
-
85
- ```bash
86
- pdf-anonymizer deanonymize ANONYMIZED_FILE MAPPING_FILE
87
- ```
88
-
89
- **Arguments**:
90
- - `ANONYMIZED_FILE`: Path to the anonymized text file.
91
- - `MAPPING_FILE`: Path to the JSON mapping file.
92
-
93
- **Example**:
94
- ```bash
95
- pdf-anonymizer deanonymize \
96
- data/anonymized/document.anonymized.md \
97
- data/mappings/document.mapping.json
98
- ```
99
-
100
- This will create a deanonymized version of the file at `data/deanonymized/document.deanonymized.md`.