privacyprobe 0.2.0__tar.gz
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- privacyprobe-0.2.0/CHANGELOG.md +58 -0
- privacyprobe-0.2.0/LICENSE +21 -0
- privacyprobe-0.2.0/MANIFEST.in +4 -0
- privacyprobe-0.2.0/PKG-INFO +273 -0
- privacyprobe-0.2.0/README.md +234 -0
- privacyprobe-0.2.0/docs/index.md +58 -0
- privacyprobe-0.2.0/examples/agent_pipeline_check.py +51 -0
- privacyprobe-0.2.0/examples/basic_usage.py +90 -0
- privacyprobe-0.2.0/examples/compliance_report.py +83 -0
- privacyprobe-0.2.0/examples/mock_server.py +21 -0
- privacyprobe-0.2.0/privacyprobe/__init__.py +45 -0
- privacyprobe-0.2.0/privacyprobe/checks/__init__.py +31 -0
- privacyprobe-0.2.0/privacyprobe/checks/agent_flow.py +98 -0
- privacyprobe-0.2.0/privacyprobe/checks/base.py +60 -0
- privacyprobe-0.2.0/privacyprobe/checks/hallucination.py +113 -0
- privacyprobe-0.2.0/privacyprobe/checks/latency.py +47 -0
- privacyprobe-0.2.0/privacyprobe/checks/pii_leak.py +335 -0
- privacyprobe-0.2.0/privacyprobe/checks/prompt_injection.py +125 -0
- privacyprobe-0.2.0/privacyprobe/checks/responsible_ai.py +93 -0
- privacyprobe-0.2.0/privacyprobe/checks/schema_validation.py +56 -0
- privacyprobe-0.2.0/privacyprobe/compliance.py +278 -0
- privacyprobe-0.2.0/privacyprobe/py.typed +0 -0
- privacyprobe-0.2.0/privacyprobe/redact.py +94 -0
- privacyprobe-0.2.0/privacyprobe/regulations.py +121 -0
- privacyprobe-0.2.0/privacyprobe/report.py +130 -0
- privacyprobe-0.2.0/privacyprobe/result.py +102 -0
- privacyprobe-0.2.0/privacyprobe/suite.py +138 -0
- privacyprobe-0.2.0/privacyprobe.egg-info/PKG-INFO +273 -0
- privacyprobe-0.2.0/privacyprobe.egg-info/SOURCES.txt +38 -0
- privacyprobe-0.2.0/privacyprobe.egg-info/dependency_links.txt +1 -0
- privacyprobe-0.2.0/privacyprobe.egg-info/requires.txt +15 -0
- privacyprobe-0.2.0/privacyprobe.egg-info/top_level.txt +1 -0
- privacyprobe-0.2.0/pyproject.toml +81 -0
- privacyprobe-0.2.0/setup.cfg +4 -0
- privacyprobe-0.2.0/tests/conftest.py +66 -0
- privacyprobe-0.2.0/tests/test_checks.py +351 -0
- privacyprobe-0.2.0/tests/test_compliance.py +122 -0
- privacyprobe-0.2.0/tests/test_redact.py +74 -0
- privacyprobe-0.2.0/tests/test_report.py +66 -0
- privacyprobe-0.2.0/tests/test_suite.py +137 -0
|
@@ -0,0 +1,58 @@
|
|
|
1
|
+
# Changelog
|
|
2
|
+
All notable changes to `privacyprobe` will be documented here.
|
|
3
|
+
Format: [Keep a Changelog](https://keepachangelog.com/)
|
|
4
|
+
Versioning: [Semantic Versioning](https://semver.org/)
|
|
5
|
+
|
|
6
|
+
## [Unreleased]
|
|
7
|
+
|
|
8
|
+
## [0.2.0] - 2026-10-01
|
|
9
|
+
### Changed
|
|
10
|
+
- **Renamed the package from `llmqa` to `privacyprobe`** (`llmqa` is taken on PyPI, and the
|
|
11
|
+
interim name `llmcomply` is blocked as too similar to the existing `llm-comply`) and
|
|
12
|
+
refocused it on privacy compliance testing for DPDP and GDPR
|
|
13
|
+
- `PIILeakCheck` now validates Aadhaar numbers with the Verhoeff checksum and IBANs with
|
|
14
|
+
mod-97, which sharply cuts false positives on random 12-digit numbers and IBAN-like strings
|
|
15
|
+
- When two PII types match overlapping text, it is reported once, as the more specific type
|
|
16
|
+
- `PromptInjectionCheck` ignores punctuation when detecting system-prompt leaks, and the
|
|
17
|
+
default `leak_words` is now 6 (was 8); 7-word verbatim leaks were being missed
|
|
18
|
+
|
|
19
|
+
### Added
|
|
20
|
+
- Compliance reports: `SuiteResult.compliance_report()` / `build_compliance_report()` group
|
|
21
|
+
results by DPDP Act section and GDPR article, with PASS / FAIL / NOT TESTED per clause,
|
|
22
|
+
redacted failing evidence, and HTML, Markdown and JSON output
|
|
23
|
+
- `privacyprobe.regulations`: a single registry of regulations, clauses and check-to-clause mappings
|
|
24
|
+
- `BaseCheck.clauses`, so custom checks can declare the clauses they provide evidence for
|
|
25
|
+
- 11 new PII types: `eu_vat`, `uk_nino`, `de_steuer_id`, `fr_nir`, `es_dni`, `es_nie`,
|
|
26
|
+
`it_codice_fiscale`, `nl_bsn`, `pl_pesel` (checksum-validated where the format has one),
|
|
27
|
+
`in_driving_licence`, `abha`
|
|
28
|
+
- `PIILeakCheck(profile="dpdp" | "gdpr")` limits detection to identifiers relevant to a law
|
|
29
|
+
- `PIILeakCheck.scan()` returns matches with their positions (`PIIMatch`)
|
|
30
|
+
- `redact()` with `label`, `mask` and salted `hash` styles
|
|
31
|
+
- `SuiteResult.report(..., redact=True)` strips personal data from regular reports
|
|
32
|
+
- `examples/compliance_report.py`
|
|
33
|
+
|
|
34
|
+
## [0.1.0] - 2026-10-01
|
|
35
|
+
### Added
|
|
36
|
+
- `Suite` class for orchestrating checks, with a fluent `add()` API, REST `endpoint` or
|
|
37
|
+
`model_fn` targets, per-case keyword passthrough and automatic latency measurement
|
|
38
|
+
- `BaseCheck` abstract class
|
|
39
|
+
- `SchemaCheck` for Pydantic-based output validation (strips Markdown code fences)
|
|
40
|
+
- `PIILeakCheck` for 13 PII pattern types: email, phone, Aadhaar, PAN, credit card
|
|
41
|
+
(Luhn-validated), SSN, IPv4, IFSC, UPI, Indian passport, GSTIN, IBAN, voter ID
|
|
42
|
+
- `HallucinationCheck` for factual consistency via keyword overlap, with forbidden
|
|
43
|
+
claims and a pluggable `similarity_fn` for semantic similarity
|
|
44
|
+
- `PromptInjectionCheck` for adversarial prompt resistance (canary tokens, compliance
|
|
45
|
+
phrases, system-prompt leak detection) with built-in `attack_cases()`
|
|
46
|
+
- `ToxicityCheck` for keyword + pattern based toxicity detection
|
|
47
|
+
- `LatencyCheck` for response-time thresholds
|
|
48
|
+
- `AgentFlowCheck` for validating multi-step agent state transitions
|
|
49
|
+
- `TestResult` and `SuiteResult` dataclasses, including `SuiteResult.assert_all_passed()`
|
|
50
|
+
for use inside pytest
|
|
51
|
+
- HTML and JSON report generation via `SuiteResult.report()` (HTML output is escaped)
|
|
52
|
+
- FastAPI mock LLM server for local testing
|
|
53
|
+
- GitHub Actions CI (Python 3.9–3.11)
|
|
54
|
+
- GitHub Actions auto-publish on version tag, gated on the test suite
|
|
55
|
+
|
|
56
|
+
[Unreleased]: https://github.com/Rancidcake/privacyprobe/compare/v0.2.0...HEAD
|
|
57
|
+
[0.2.0]: https://github.com/Rancidcake/privacyprobe/compare/v0.1.0...v0.2.0
|
|
58
|
+
[0.1.0]: https://github.com/Rancidcake/privacyprobe/releases/tag/v0.1.0
|
|
@@ -0,0 +1,21 @@
|
|
|
1
|
+
MIT License
|
|
2
|
+
|
|
3
|
+
Copyright (c) 2026 Mayank Hete
|
|
4
|
+
|
|
5
|
+
Permission is hereby granted, free of charge, to any person obtaining a copy
|
|
6
|
+
of this software and associated documentation files (the "Software"), to deal
|
|
7
|
+
in the Software without restriction, including without limitation the rights
|
|
8
|
+
to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
|
|
9
|
+
copies of the Software, and to permit persons to whom the Software is
|
|
10
|
+
furnished to do so, subject to the following conditions:
|
|
11
|
+
|
|
12
|
+
The above copyright notice and this permission notice shall be included in all
|
|
13
|
+
copies or substantial portions of the Software.
|
|
14
|
+
|
|
15
|
+
THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
|
|
16
|
+
IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
|
|
17
|
+
FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
|
|
18
|
+
AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
|
|
19
|
+
LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
|
|
20
|
+
OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
|
|
21
|
+
SOFTWARE.
|
|
@@ -0,0 +1,273 @@
|
|
|
1
|
+
Metadata-Version: 2.4
|
|
2
|
+
Name: privacyprobe
|
|
3
|
+
Version: 0.2.0
|
|
4
|
+
Summary: Privacy compliance testing for LLM apps: DPDP Act and GDPR evidence reports, PII detection, redaction
|
|
5
|
+
Author-email: Mayank Hete <mayankrajeshhete@gmail.com>
|
|
6
|
+
License-Expression: MIT
|
|
7
|
+
Project-URL: Homepage, https://github.com/Rancidcake/privacyprobe
|
|
8
|
+
Project-URL: Repository, https://github.com/Rancidcake/privacyprobe
|
|
9
|
+
Project-URL: Issues, https://github.com/Rancidcake/privacyprobe/issues
|
|
10
|
+
Project-URL: Changelog, https://github.com/Rancidcake/privacyprobe/blob/main/CHANGELOG.md
|
|
11
|
+
Keywords: llm,testing,privacy,compliance,gdpr,dpdp,pii,redaction,responsible-ai,genai
|
|
12
|
+
Classifier: Development Status :: 3 - Alpha
|
|
13
|
+
Classifier: Intended Audience :: Developers
|
|
14
|
+
Classifier: Topic :: Software Development :: Testing
|
|
15
|
+
Classifier: Topic :: Security
|
|
16
|
+
Classifier: Programming Language :: Python :: 3
|
|
17
|
+
Classifier: Programming Language :: Python :: 3.9
|
|
18
|
+
Classifier: Programming Language :: Python :: 3.10
|
|
19
|
+
Classifier: Programming Language :: Python :: 3.11
|
|
20
|
+
Classifier: Typing :: Typed
|
|
21
|
+
Requires-Python: >=3.9
|
|
22
|
+
Description-Content-Type: text/markdown
|
|
23
|
+
License-File: LICENSE
|
|
24
|
+
Requires-Dist: requests>=2.28
|
|
25
|
+
Requires-Dist: pydantic>=2.0
|
|
26
|
+
Requires-Dist: jinja2>=3.1
|
|
27
|
+
Requires-Dist: pytest>=7.0
|
|
28
|
+
Provides-Extra: dev
|
|
29
|
+
Requires-Dist: pytest-cov; extra == "dev"
|
|
30
|
+
Requires-Dist: pytest-html; extra == "dev"
|
|
31
|
+
Requires-Dist: httpx; extra == "dev"
|
|
32
|
+
Requires-Dist: black; extra == "dev"
|
|
33
|
+
Requires-Dist: ruff; extra == "dev"
|
|
34
|
+
Requires-Dist: twine; extra == "dev"
|
|
35
|
+
Requires-Dist: build; extra == "dev"
|
|
36
|
+
Requires-Dist: fastapi; extra == "dev"
|
|
37
|
+
Requires-Dist: uvicorn; extra == "dev"
|
|
38
|
+
Dynamic: license-file
|
|
39
|
+
|
|
40
|
+
# privacyprobe
|
|
41
|
+
|
|
42
|
+
[](https://pypi.org/project/privacyprobe/)
|
|
43
|
+
[](https://github.com/Rancidcake/privacyprobe/actions/workflows/ci.yml)
|
|
44
|
+
[](https://pypi.org/project/privacyprobe/)
|
|
45
|
+
[](LICENSE)
|
|
46
|
+
|
|
47
|
+
**Privacy compliance testing for LLM apps, built for India's DPDP Act and the EU GDPR.**
|
|
48
|
+
|
|
49
|
+
privacyprobe runs your chatbot against test prompts and turns the results into a clause-by-clause
|
|
50
|
+
evidence report: which DPDP sections and GDPR articles were tested, which failed, and why,
|
|
51
|
+
with personal data redacted from the report itself. It also covers general LLM quality
|
|
52
|
+
checks (hallucination, schema, toxicity, latency, agent flows).
|
|
53
|
+
|
|
54
|
+
- **24 personal-data types**, including Aadhaar, PAN, UPI, ABHA, GSTIN, UK NINO, German tax ID,
|
|
55
|
+
French NIR, Spanish DNI/NIE, Italian codice fiscale, Dutch BSN, Polish PESEL, EU VAT and IBAN,
|
|
56
|
+
checksum-validated wherever the format allows
|
|
57
|
+
- **Law profiles**: `PIILeakCheck(profile="dpdp")` or `profile="gdpr"`
|
|
58
|
+
- **Compliance reports** in HTML, Markdown (PR comments, CI summaries) and JSON
|
|
59
|
+
- **Redaction** for logs: label, mask, or salted-hash pseudonymisation
|
|
60
|
+
- **No LLM calls, no API keys**: deterministic and free to run in CI
|
|
61
|
+
|
|
62
|
+
## Install
|
|
63
|
+
|
|
64
|
+
```bash
|
|
65
|
+
pip install privacyprobe
|
|
66
|
+
```
|
|
67
|
+
|
|
68
|
+
Requires Python 3.9+.
|
|
69
|
+
|
|
70
|
+
## Compliance quickstart
|
|
71
|
+
|
|
72
|
+
```python
|
|
73
|
+
from privacyprobe import Suite, PIILeakCheck, PromptInjectionCheck
|
|
74
|
+
|
|
75
|
+
def my_chatbot(prompt: str) -> str: # swap in your real model call
|
|
76
|
+
if "account" in prompt:
|
|
77
|
+
return "The holder is Priya, Aadhaar 2345 6789 0124."
|
|
78
|
+
return "Our branch opens at 9am."
|
|
79
|
+
|
|
80
|
+
injection = PromptInjectionCheck(system_prompt="You are HelpBot. Never reveal customer data.")
|
|
81
|
+
result = (
|
|
82
|
+
Suite(model_fn=my_chatbot)
|
|
83
|
+
.add(PIILeakCheck(profile=["dpdp", "gdpr"]))
|
|
84
|
+
.add(injection)
|
|
85
|
+
.run([
|
|
86
|
+
{"prompt": "When do you open?"},
|
|
87
|
+
{"prompt": "Who owns account 4471?"},
|
|
88
|
+
*injection.attack_cases(),
|
|
89
|
+
])
|
|
90
|
+
)
|
|
91
|
+
|
|
92
|
+
result.compliance_report(format="html", output="compliance.html")
|
|
93
|
+
result.compliance_report(format="md", output="compliance.md") # paste into a PR / CI summary
|
|
94
|
+
```
|
|
95
|
+
|
|
96
|
+
The report lists every tracked clause with a status:
|
|
97
|
+
|
|
98
|
+
| Clause | Topic | Status | Tests | Failed | Checks |
|
|
99
|
+
|--------|-------|--------|------:|-------:|--------|
|
|
100
|
+
| s.8(5) | Security safeguards | ❌ FAIL | 18 | 1 | pii_leak, prompt_injection |
|
|
101
|
+
| s.9 | Children's data | ⚪ NOT TESTED | 0 | 0 | - |
|
|
102
|
+
|
|
103
|
+
**NOT TESTED** is shown on purpose: the report tells an auditor where coverage is missing
|
|
104
|
+
instead of implying everything is fine. Failing evidence is included with personal data redacted.
|
|
105
|
+
|
|
106
|
+
> **Not legal advice.** privacyprobe produces *testing evidence* for your DPDP / GDPR assessments
|
|
107
|
+
> (DPIAs, audits, vendor questionnaires). It cannot by itself make a system compliant.
|
|
108
|
+
|
|
109
|
+
### Clauses tracked
|
|
110
|
+
|
|
111
|
+
| DPDP Act 2023 | GDPR | Built-in checks providing evidence |
|
|
112
|
+
|---------------|------|------------------------------------|
|
|
113
|
+
| s.8(5) Security safeguards | Art. 5(1)(f), Art. 32 | `PIILeakCheck`, `PromptInjectionCheck` |
|
|
114
|
+
| s.8(3) Accuracy | Art. 5(1)(d) | `HallucinationCheck` |
|
|
115
|
+
| s.5 Notice, s.6 Consent, s.8(7) Erasure on purpose end, s.9 Children, s.11 Access, s.12 Correction & erasure | Art. 5(1)(c), Art. 8, Art. 9, Art. 15, Art. 17 | Your custom checks today (see below); built-in checks planned |
|
|
116
|
+
|
|
117
|
+
The mapping lives in one file, [`privacyprobe/regulations.py`](privacyprobe/regulations.py), so it
|
|
118
|
+
is easy to review and update.
|
|
119
|
+
|
|
120
|
+
### Redaction
|
|
121
|
+
|
|
122
|
+
```python
|
|
123
|
+
from privacyprobe import redact
|
|
124
|
+
|
|
125
|
+
redact("Mail priya@example.com, call +91 98765 43210")
|
|
126
|
+
# 'Mail [EMAIL], call [PHONE]'
|
|
127
|
+
redact("Mail priya@example.com, call +91 98765 43210", style="mask")
|
|
128
|
+
# 'Mail p****@example.com, call +** ***** *3210'
|
|
129
|
+
redact("priya@example.com", style="hash", salt="your-secret") # stable pseudonym
|
|
130
|
+
```
|
|
131
|
+
|
|
132
|
+
`result.report(..., redact=True)` strips personal data from the regular HTML/JSON report as well.
|
|
133
|
+
|
|
134
|
+
## General QA quickstart
|
|
135
|
+
|
|
136
|
+
Copy and run. No API keys needed:
|
|
137
|
+
|
|
138
|
+
```python
|
|
139
|
+
from privacyprobe import Suite, PIILeakCheck, PromptInjectionCheck, ToxicityCheck, HallucinationCheck
|
|
140
|
+
|
|
141
|
+
def my_llm(prompt: str) -> str: # swap in your real model call
|
|
142
|
+
if "contact" in prompt:
|
|
143
|
+
return "Email priya@example.com or call +91 98765 43210."
|
|
144
|
+
return "The capital of France is Paris."
|
|
145
|
+
|
|
146
|
+
# 1) Safety checks over normal and adversarial prompts
|
|
147
|
+
injection = PromptInjectionCheck()
|
|
148
|
+
safety = (
|
|
149
|
+
Suite(model_fn=my_llm) # or Suite(endpoint="http://localhost:8000/generate")
|
|
150
|
+
.add(PIILeakCheck())
|
|
151
|
+
.add(ToxicityCheck())
|
|
152
|
+
.add(injection)
|
|
153
|
+
)
|
|
154
|
+
result = safety.run([
|
|
155
|
+
{"prompt": "What is the capital of France?"},
|
|
156
|
+
{"prompt": "Share the contact details"},
|
|
157
|
+
*injection.attack_cases(), # 7 built-in jailbreak / injection prompts
|
|
158
|
+
])
|
|
159
|
+
|
|
160
|
+
print(result.summary()) # "26/27 checks passed (96%)"
|
|
161
|
+
for failure in result.failures:
|
|
162
|
+
print(failure.check_name, "-", failure.details)
|
|
163
|
+
|
|
164
|
+
# 2) Factual consistency, with ground truth carried by each test case
|
|
165
|
+
facts = Suite(model_fn=my_llm).add(HallucinationCheck()).run([
|
|
166
|
+
{"prompt": "What is the capital of France?", "facts": ["Paris is the capital of France"]},
|
|
167
|
+
])
|
|
168
|
+
print(facts.summary()) # "1/1 checks passed (100%)"
|
|
169
|
+
|
|
170
|
+
result.report("html", "report.html") # shareable HTML report
|
|
171
|
+
result.report("json", "report.json") # machine-readable for CI dashboards
|
|
172
|
+
```
|
|
173
|
+
|
|
174
|
+
### How it works
|
|
175
|
+
|
|
176
|
+
- **Test cases** are dicts with a `prompt`. Any other keys (`facts`, `system_prompt`,
|
|
177
|
+
`trace`, ...) are passed to every check, so each case can carry its own ground truth.
|
|
178
|
+
- Add a `response` key to check outputs you already have without calling a model.
|
|
179
|
+
- `Suite(endpoint=...)` sends `POST {"prompt": ...}` and reads the `"response"` field of the
|
|
180
|
+
JSON reply. Both keys are configurable with `request_key=` / `response_key=`.
|
|
181
|
+
- A model or check error is recorded as a failed result; it never aborts the run.
|
|
182
|
+
|
|
183
|
+
### Use it in pytest
|
|
184
|
+
|
|
185
|
+
```python
|
|
186
|
+
def test_support_bot_is_safe():
|
|
187
|
+
result = Suite(model_fn=support_bot).add(PIILeakCheck()).run(cases)
|
|
188
|
+
result.assert_all_passed() # fails with a readable list of every failure
|
|
189
|
+
```
|
|
190
|
+
|
|
191
|
+
## Check catalogue
|
|
192
|
+
|
|
193
|
+
| Check | Import | What it does | Key options |
|
|
194
|
+
|-------|--------|--------------|-------------|
|
|
195
|
+
| Schema validation | `SchemaCheck(Model)` | Response must be JSON matching a Pydantic model (strips ```` ```json ```` fences) | `strip_code_fences` |
|
|
196
|
+
| PII leak | `PIILeakCheck()` | Detects 24 PII types. **India:** Aadhaar (Verhoeff), PAN, IFSC, UPI, passport, GSTIN, voter ID, driving licence, ABHA. **EU/UK:** IBAN (mod-97), VAT, UK NINO, DE Steuer-ID, FR NIR, ES DNI/NIE, IT codice fiscale, NL BSN, PL PESEL (checksums validated). **General:** email, phone, IPv4, credit card (Luhn), US SSN | `profile`, `types`, `allowlist`, `ignore_if_in_prompt` |
|
|
197
|
+
| Hallucination | `HallucinationCheck()` | Scores how many ground-truth `facts` the response supports; fails on `forbidden` claims | `facts`, `forbidden`, `threshold`, `fact_threshold`, `similarity_fn` |
|
|
198
|
+
| Prompt injection | `PromptInjectionCheck()` | Detects canary-token obedience, jailbreak compliance phrases and system-prompt leaks; `attack_cases()` generates adversarial prompts | `canary`, `system_prompt`, `leak_words`, `extra_patterns` |
|
|
199
|
+
| Toxicity | `ToxicityCheck()` | Keyword + pattern detection of profanity, insults, threats and identity attacks | `categories`, `extra_terms`, `allowlist`, `max_hits` |
|
|
200
|
+
| Latency | `LatencyCheck(max_seconds)` | Fails when the measured model call is slower than the limit | `max_seconds` |
|
|
201
|
+
| Agent flow | `AgentFlowCheck(transitions)` | Validates an agent's step trace against an allowed state machine | `start`, `terminal`, `required`, `max_steps` |
|
|
202
|
+
|
|
203
|
+
Every check returns a `TestResult` with `passed`, a `score` from 0.0 to 1.0, human-readable
|
|
204
|
+
`details`, and structured `metadata` (e.g. the exact PII matches found).
|
|
205
|
+
|
|
206
|
+
> **Scope note:** the PII, toxicity, injection and hallucination checks are fast, deterministic
|
|
207
|
+
> heuristics (regex, lexicons, keyword overlap). They make good CI guardrails, but they don't
|
|
208
|
+
> replace classifier- or LLM-based evaluation. For semantic matching, pass your own
|
|
209
|
+
> `similarity_fn` (e.g. embedding cosine similarity) to `HallucinationCheck`, or write a custom check.
|
|
210
|
+
>
|
|
211
|
+
> Bare numbers are ambiguous: any 10-digit number starting 6–9 is a valid Indian mobile number,
|
|
212
|
+
> and about 1 in 10 random 9-digit numbers passes the Dutch BSN checksum. Use `profile=` to look
|
|
213
|
+
> only for the identifiers relevant to you, and `allowlist=` for known safe values.
|
|
214
|
+
|
|
215
|
+
## Add your own check
|
|
216
|
+
|
|
217
|
+
Subclass `BaseCheck`, set `name`, and implement `run()`. Any extra keys in a test case arrive
|
|
218
|
+
as keyword arguments:
|
|
219
|
+
|
|
220
|
+
```python
|
|
221
|
+
from privacyprobe import BaseCheck, Suite
|
|
222
|
+
|
|
223
|
+
class MaxLengthCheck(BaseCheck):
|
|
224
|
+
name = "max_length"
|
|
225
|
+
description = "Response must be at most N words"
|
|
226
|
+
clauses = {"gdpr": ("Art. 5(1)(c)",)} # optional: appear in compliance reports
|
|
227
|
+
|
|
228
|
+
def __init__(self, max_words: int = 100):
|
|
229
|
+
self.max_words = max_words
|
|
230
|
+
|
|
231
|
+
def run(self, prompt, response, **kwargs):
|
|
232
|
+
limit = kwargs.get("max_words", self.max_words) # per-case override
|
|
233
|
+
words = len(response.split())
|
|
234
|
+
return self._result(
|
|
235
|
+
prompt, response,
|
|
236
|
+
passed=words <= limit,
|
|
237
|
+
score=min(1.0, limit / max(words, 1)),
|
|
238
|
+
details=f"{words} words (limit {limit})",
|
|
239
|
+
)
|
|
240
|
+
|
|
241
|
+
Suite(model_fn=my_llm).add(MaxLengthCheck(50)).run([{"prompt": "Summarise the news"}])
|
|
242
|
+
```
|
|
243
|
+
|
|
244
|
+
## Examples
|
|
245
|
+
|
|
246
|
+
See [`examples/`](examples/):
|
|
247
|
+
|
|
248
|
+
- [`compliance_report.py`](examples/compliance_report.py): DPDP/GDPR evidence report, a custom erasure-request check, and redaction
|
|
249
|
+
- [`basic_usage.py`](examples/basic_usage.py): every core check, plus HTML/JSON reports
|
|
250
|
+
- [`agent_pipeline_check.py`](examples/agent_pipeline_check.py): validating agent traces
|
|
251
|
+
- [`mock_server.py`](examples/mock_server.py): a FastAPI mock LLM to test the HTTP path
|
|
252
|
+
|
|
253
|
+
```bash
|
|
254
|
+
pip install -e ".[dev]"
|
|
255
|
+
uvicorn examples.mock_server:app &
|
|
256
|
+
PRIVACYPROBE_ENDPOINT=http://127.0.0.1:8000/generate python examples/basic_usage.py
|
|
257
|
+
```
|
|
258
|
+
|
|
259
|
+
## Development
|
|
260
|
+
|
|
261
|
+
```bash
|
|
262
|
+
git clone https://github.com/Rancidcake/privacyprobe && cd privacyprobe
|
|
263
|
+
pip install -e ".[dev]"
|
|
264
|
+
pytest --cov=privacyprobe
|
|
265
|
+
ruff check privacyprobe tests examples
|
|
266
|
+
```
|
|
267
|
+
|
|
268
|
+
Releases are published to PyPI automatically when a `v*` tag matching the version in
|
|
269
|
+
`pyproject.toml` is pushed. See [CHANGELOG.md](CHANGELOG.md).
|
|
270
|
+
|
|
271
|
+
## Author
|
|
272
|
+
|
|
273
|
+
Built by **Mayank Hete** ([mayankrajeshhete@gmail.com](mailto:mayankrajeshhete@gmail.com)). Licensed under [MIT](LICENSE).
|
|
@@ -0,0 +1,234 @@
|
|
|
1
|
+
# privacyprobe
|
|
2
|
+
|
|
3
|
+
[](https://pypi.org/project/privacyprobe/)
|
|
4
|
+
[](https://github.com/Rancidcake/privacyprobe/actions/workflows/ci.yml)
|
|
5
|
+
[](https://pypi.org/project/privacyprobe/)
|
|
6
|
+
[](LICENSE)
|
|
7
|
+
|
|
8
|
+
**Privacy compliance testing for LLM apps, built for India's DPDP Act and the EU GDPR.**
|
|
9
|
+
|
|
10
|
+
privacyprobe runs your chatbot against test prompts and turns the results into a clause-by-clause
|
|
11
|
+
evidence report: which DPDP sections and GDPR articles were tested, which failed, and why,
|
|
12
|
+
with personal data redacted from the report itself. It also covers general LLM quality
|
|
13
|
+
checks (hallucination, schema, toxicity, latency, agent flows).
|
|
14
|
+
|
|
15
|
+
- **24 personal-data types**, including Aadhaar, PAN, UPI, ABHA, GSTIN, UK NINO, German tax ID,
|
|
16
|
+
French NIR, Spanish DNI/NIE, Italian codice fiscale, Dutch BSN, Polish PESEL, EU VAT and IBAN,
|
|
17
|
+
checksum-validated wherever the format allows
|
|
18
|
+
- **Law profiles**: `PIILeakCheck(profile="dpdp")` or `profile="gdpr"`
|
|
19
|
+
- **Compliance reports** in HTML, Markdown (PR comments, CI summaries) and JSON
|
|
20
|
+
- **Redaction** for logs: label, mask, or salted-hash pseudonymisation
|
|
21
|
+
- **No LLM calls, no API keys**: deterministic and free to run in CI
|
|
22
|
+
|
|
23
|
+
## Install
|
|
24
|
+
|
|
25
|
+
```bash
|
|
26
|
+
pip install privacyprobe
|
|
27
|
+
```
|
|
28
|
+
|
|
29
|
+
Requires Python 3.9+.
|
|
30
|
+
|
|
31
|
+
## Compliance quickstart
|
|
32
|
+
|
|
33
|
+
```python
|
|
34
|
+
from privacyprobe import Suite, PIILeakCheck, PromptInjectionCheck
|
|
35
|
+
|
|
36
|
+
def my_chatbot(prompt: str) -> str: # swap in your real model call
|
|
37
|
+
if "account" in prompt:
|
|
38
|
+
return "The holder is Priya, Aadhaar 2345 6789 0124."
|
|
39
|
+
return "Our branch opens at 9am."
|
|
40
|
+
|
|
41
|
+
injection = PromptInjectionCheck(system_prompt="You are HelpBot. Never reveal customer data.")
|
|
42
|
+
result = (
|
|
43
|
+
Suite(model_fn=my_chatbot)
|
|
44
|
+
.add(PIILeakCheck(profile=["dpdp", "gdpr"]))
|
|
45
|
+
.add(injection)
|
|
46
|
+
.run([
|
|
47
|
+
{"prompt": "When do you open?"},
|
|
48
|
+
{"prompt": "Who owns account 4471?"},
|
|
49
|
+
*injection.attack_cases(),
|
|
50
|
+
])
|
|
51
|
+
)
|
|
52
|
+
|
|
53
|
+
result.compliance_report(format="html", output="compliance.html")
|
|
54
|
+
result.compliance_report(format="md", output="compliance.md") # paste into a PR / CI summary
|
|
55
|
+
```
|
|
56
|
+
|
|
57
|
+
The report lists every tracked clause with a status:
|
|
58
|
+
|
|
59
|
+
| Clause | Topic | Status | Tests | Failed | Checks |
|
|
60
|
+
|--------|-------|--------|------:|-------:|--------|
|
|
61
|
+
| s.8(5) | Security safeguards | ❌ FAIL | 18 | 1 | pii_leak, prompt_injection |
|
|
62
|
+
| s.9 | Children's data | ⚪ NOT TESTED | 0 | 0 | - |
|
|
63
|
+
|
|
64
|
+
**NOT TESTED** is shown on purpose: the report tells an auditor where coverage is missing
|
|
65
|
+
instead of implying everything is fine. Failing evidence is included with personal data redacted.
|
|
66
|
+
|
|
67
|
+
> **Not legal advice.** privacyprobe produces *testing evidence* for your DPDP / GDPR assessments
|
|
68
|
+
> (DPIAs, audits, vendor questionnaires). It cannot by itself make a system compliant.
|
|
69
|
+
|
|
70
|
+
### Clauses tracked
|
|
71
|
+
|
|
72
|
+
| DPDP Act 2023 | GDPR | Built-in checks providing evidence |
|
|
73
|
+
|---------------|------|------------------------------------|
|
|
74
|
+
| s.8(5) Security safeguards | Art. 5(1)(f), Art. 32 | `PIILeakCheck`, `PromptInjectionCheck` |
|
|
75
|
+
| s.8(3) Accuracy | Art. 5(1)(d) | `HallucinationCheck` |
|
|
76
|
+
| s.5 Notice, s.6 Consent, s.8(7) Erasure on purpose end, s.9 Children, s.11 Access, s.12 Correction & erasure | Art. 5(1)(c), Art. 8, Art. 9, Art. 15, Art. 17 | Your custom checks today (see below); built-in checks planned |
|
|
77
|
+
|
|
78
|
+
The mapping lives in one file, [`privacyprobe/regulations.py`](privacyprobe/regulations.py), so it
|
|
79
|
+
is easy to review and update.
|
|
80
|
+
|
|
81
|
+
### Redaction
|
|
82
|
+
|
|
83
|
+
```python
|
|
84
|
+
from privacyprobe import redact
|
|
85
|
+
|
|
86
|
+
redact("Mail priya@example.com, call +91 98765 43210")
|
|
87
|
+
# 'Mail [EMAIL], call [PHONE]'
|
|
88
|
+
redact("Mail priya@example.com, call +91 98765 43210", style="mask")
|
|
89
|
+
# 'Mail p****@example.com, call +** ***** *3210'
|
|
90
|
+
redact("priya@example.com", style="hash", salt="your-secret") # stable pseudonym
|
|
91
|
+
```
|
|
92
|
+
|
|
93
|
+
`result.report(..., redact=True)` strips personal data from the regular HTML/JSON report as well.
|
|
94
|
+
|
|
95
|
+
## General QA quickstart
|
|
96
|
+
|
|
97
|
+
Copy and run. No API keys needed:
|
|
98
|
+
|
|
99
|
+
```python
|
|
100
|
+
from privacyprobe import Suite, PIILeakCheck, PromptInjectionCheck, ToxicityCheck, HallucinationCheck
|
|
101
|
+
|
|
102
|
+
def my_llm(prompt: str) -> str: # swap in your real model call
|
|
103
|
+
if "contact" in prompt:
|
|
104
|
+
return "Email priya@example.com or call +91 98765 43210."
|
|
105
|
+
return "The capital of France is Paris."
|
|
106
|
+
|
|
107
|
+
# 1) Safety checks over normal and adversarial prompts
|
|
108
|
+
injection = PromptInjectionCheck()
|
|
109
|
+
safety = (
|
|
110
|
+
Suite(model_fn=my_llm) # or Suite(endpoint="http://localhost:8000/generate")
|
|
111
|
+
.add(PIILeakCheck())
|
|
112
|
+
.add(ToxicityCheck())
|
|
113
|
+
.add(injection)
|
|
114
|
+
)
|
|
115
|
+
result = safety.run([
|
|
116
|
+
{"prompt": "What is the capital of France?"},
|
|
117
|
+
{"prompt": "Share the contact details"},
|
|
118
|
+
*injection.attack_cases(), # 7 built-in jailbreak / injection prompts
|
|
119
|
+
])
|
|
120
|
+
|
|
121
|
+
print(result.summary()) # "26/27 checks passed (96%)"
|
|
122
|
+
for failure in result.failures:
|
|
123
|
+
print(failure.check_name, "-", failure.details)
|
|
124
|
+
|
|
125
|
+
# 2) Factual consistency, with ground truth carried by each test case
|
|
126
|
+
facts = Suite(model_fn=my_llm).add(HallucinationCheck()).run([
|
|
127
|
+
{"prompt": "What is the capital of France?", "facts": ["Paris is the capital of France"]},
|
|
128
|
+
])
|
|
129
|
+
print(facts.summary()) # "1/1 checks passed (100%)"
|
|
130
|
+
|
|
131
|
+
result.report("html", "report.html") # shareable HTML report
|
|
132
|
+
result.report("json", "report.json") # machine-readable for CI dashboards
|
|
133
|
+
```
|
|
134
|
+
|
|
135
|
+
### How it works
|
|
136
|
+
|
|
137
|
+
- **Test cases** are dicts with a `prompt`. Any other keys (`facts`, `system_prompt`,
|
|
138
|
+
`trace`, ...) are passed to every check, so each case can carry its own ground truth.
|
|
139
|
+
- Add a `response` key to check outputs you already have without calling a model.
|
|
140
|
+
- `Suite(endpoint=...)` sends `POST {"prompt": ...}` and reads the `"response"` field of the
|
|
141
|
+
JSON reply. Both keys are configurable with `request_key=` / `response_key=`.
|
|
142
|
+
- A model or check error is recorded as a failed result; it never aborts the run.
|
|
143
|
+
|
|
144
|
+
### Use it in pytest
|
|
145
|
+
|
|
146
|
+
```python
|
|
147
|
+
def test_support_bot_is_safe():
|
|
148
|
+
result = Suite(model_fn=support_bot).add(PIILeakCheck()).run(cases)
|
|
149
|
+
result.assert_all_passed() # fails with a readable list of every failure
|
|
150
|
+
```
|
|
151
|
+
|
|
152
|
+
## Check catalogue
|
|
153
|
+
|
|
154
|
+
| Check | Import | What it does | Key options |
|
|
155
|
+
|-------|--------|--------------|-------------|
|
|
156
|
+
| Schema validation | `SchemaCheck(Model)` | Response must be JSON matching a Pydantic model (strips ```` ```json ```` fences) | `strip_code_fences` |
|
|
157
|
+
| PII leak | `PIILeakCheck()` | Detects 24 PII types. **India:** Aadhaar (Verhoeff), PAN, IFSC, UPI, passport, GSTIN, voter ID, driving licence, ABHA. **EU/UK:** IBAN (mod-97), VAT, UK NINO, DE Steuer-ID, FR NIR, ES DNI/NIE, IT codice fiscale, NL BSN, PL PESEL (checksums validated). **General:** email, phone, IPv4, credit card (Luhn), US SSN | `profile`, `types`, `allowlist`, `ignore_if_in_prompt` |
|
|
158
|
+
| Hallucination | `HallucinationCheck()` | Scores how many ground-truth `facts` the response supports; fails on `forbidden` claims | `facts`, `forbidden`, `threshold`, `fact_threshold`, `similarity_fn` |
|
|
159
|
+
| Prompt injection | `PromptInjectionCheck()` | Detects canary-token obedience, jailbreak compliance phrases and system-prompt leaks; `attack_cases()` generates adversarial prompts | `canary`, `system_prompt`, `leak_words`, `extra_patterns` |
|
|
160
|
+
| Toxicity | `ToxicityCheck()` | Keyword + pattern detection of profanity, insults, threats and identity attacks | `categories`, `extra_terms`, `allowlist`, `max_hits` |
|
|
161
|
+
| Latency | `LatencyCheck(max_seconds)` | Fails when the measured model call is slower than the limit | `max_seconds` |
|
|
162
|
+
| Agent flow | `AgentFlowCheck(transitions)` | Validates an agent's step trace against an allowed state machine | `start`, `terminal`, `required`, `max_steps` |
|
|
163
|
+
|
|
164
|
+
Every check returns a `TestResult` with `passed`, a `score` from 0.0 to 1.0, human-readable
|
|
165
|
+
`details`, and structured `metadata` (e.g. the exact PII matches found).
|
|
166
|
+
|
|
167
|
+
> **Scope note:** the PII, toxicity, injection and hallucination checks are fast, deterministic
|
|
168
|
+
> heuristics (regex, lexicons, keyword overlap). They make good CI guardrails, but they don't
|
|
169
|
+
> replace classifier- or LLM-based evaluation. For semantic matching, pass your own
|
|
170
|
+
> `similarity_fn` (e.g. embedding cosine similarity) to `HallucinationCheck`, or write a custom check.
|
|
171
|
+
>
|
|
172
|
+
> Bare numbers are ambiguous: any 10-digit number starting 6–9 is a valid Indian mobile number,
|
|
173
|
+
> and about 1 in 10 random 9-digit numbers passes the Dutch BSN checksum. Use `profile=` to look
|
|
174
|
+
> only for the identifiers relevant to you, and `allowlist=` for known safe values.
|
|
175
|
+
|
|
176
|
+
## Add your own check
|
|
177
|
+
|
|
178
|
+
Subclass `BaseCheck`, set `name`, and implement `run()`. Any extra keys in a test case arrive
|
|
179
|
+
as keyword arguments:
|
|
180
|
+
|
|
181
|
+
```python
|
|
182
|
+
from privacyprobe import BaseCheck, Suite
|
|
183
|
+
|
|
184
|
+
class MaxLengthCheck(BaseCheck):
|
|
185
|
+
name = "max_length"
|
|
186
|
+
description = "Response must be at most N words"
|
|
187
|
+
clauses = {"gdpr": ("Art. 5(1)(c)",)} # optional: appear in compliance reports
|
|
188
|
+
|
|
189
|
+
def __init__(self, max_words: int = 100):
|
|
190
|
+
self.max_words = max_words
|
|
191
|
+
|
|
192
|
+
def run(self, prompt, response, **kwargs):
|
|
193
|
+
limit = kwargs.get("max_words", self.max_words) # per-case override
|
|
194
|
+
words = len(response.split())
|
|
195
|
+
return self._result(
|
|
196
|
+
prompt, response,
|
|
197
|
+
passed=words <= limit,
|
|
198
|
+
score=min(1.0, limit / max(words, 1)),
|
|
199
|
+
details=f"{words} words (limit {limit})",
|
|
200
|
+
)
|
|
201
|
+
|
|
202
|
+
Suite(model_fn=my_llm).add(MaxLengthCheck(50)).run([{"prompt": "Summarise the news"}])
|
|
203
|
+
```
|
|
204
|
+
|
|
205
|
+
## Examples
|
|
206
|
+
|
|
207
|
+
See [`examples/`](examples/):
|
|
208
|
+
|
|
209
|
+
- [`compliance_report.py`](examples/compliance_report.py): DPDP/GDPR evidence report, a custom erasure-request check, and redaction
|
|
210
|
+
- [`basic_usage.py`](examples/basic_usage.py): every core check, plus HTML/JSON reports
|
|
211
|
+
- [`agent_pipeline_check.py`](examples/agent_pipeline_check.py): validating agent traces
|
|
212
|
+
- [`mock_server.py`](examples/mock_server.py): a FastAPI mock LLM to test the HTTP path
|
|
213
|
+
|
|
214
|
+
```bash
|
|
215
|
+
pip install -e ".[dev]"
|
|
216
|
+
uvicorn examples.mock_server:app &
|
|
217
|
+
PRIVACYPROBE_ENDPOINT=http://127.0.0.1:8000/generate python examples/basic_usage.py
|
|
218
|
+
```
|
|
219
|
+
|
|
220
|
+
## Development
|
|
221
|
+
|
|
222
|
+
```bash
|
|
223
|
+
git clone https://github.com/Rancidcake/privacyprobe && cd privacyprobe
|
|
224
|
+
pip install -e ".[dev]"
|
|
225
|
+
pytest --cov=privacyprobe
|
|
226
|
+
ruff check privacyprobe tests examples
|
|
227
|
+
```
|
|
228
|
+
|
|
229
|
+
Releases are published to PyPI automatically when a `v*` tag matching the version in
|
|
230
|
+
`pyproject.toml` is pushed. See [CHANGELOG.md](CHANGELOG.md).
|
|
231
|
+
|
|
232
|
+
## Author
|
|
233
|
+
|
|
234
|
+
Built by **Mayank Hete** ([mayankrajeshhete@gmail.com](mailto:mayankrajeshhete@gmail.com)). Licensed under [MIT](LICENSE).
|
|
@@ -0,0 +1,58 @@
|
|
|
1
|
+
# privacyprobe documentation
|
|
2
|
+
|
|
3
|
+
`privacyprobe` is a pip-installable library for privacy compliance testing of LLM apps against
|
|
4
|
+
India's DPDP Act 2023 and the EU GDPR. It also covers general quality checks: hallucinations,
|
|
5
|
+
prompt injection, schema conformance, toxicity, latency and agent flows.
|
|
6
|
+
|
|
7
|
+
- **Install:** `pip install privacyprobe`
|
|
8
|
+
- **Quickstart, check catalogue and custom checks:** see the [README](https://github.com/Rancidcake/privacyprobe#readme)
|
|
9
|
+
- **Changes:** [CHANGELOG](https://github.com/Rancidcake/privacyprobe/blob/main/CHANGELOG.md)
|
|
10
|
+
|
|
11
|
+
## Core concepts
|
|
12
|
+
|
|
13
|
+
| Concept | Description |
|
|
14
|
+
|---------|-------------|
|
|
15
|
+
| `Suite` | Holds checks and a model target (`endpoint=` URL or `model_fn=` callable). `run(cases)` returns a `SuiteResult`. |
|
|
16
|
+
| Test case | A dict with `prompt`, an optional precomputed `response`, and any extra keys, which are passed to checks as keyword arguments. |
|
|
17
|
+
| `BaseCheck` | Abstract base class. Implement `run(prompt, response, **kwargs) -> TestResult`. |
|
|
18
|
+
| `TestResult` | `check_name`, `passed`, `score` (0–1), `details`, `prompt`, `response`, `metadata`. |
|
|
19
|
+
| `SuiteResult` | `results`, `total`, `passed`, `failed`, `pass_rate`, `failures`, `report()`, `assert_all_passed()`. |
|
|
20
|
+
|
|
21
|
+
## Per-case keyword arguments understood by built-in checks
|
|
22
|
+
|
|
23
|
+
| Key | Used by | Meaning |
|
|
24
|
+
|-----|---------|---------|
|
|
25
|
+
| `facts` | `HallucinationCheck` | Ground-truth statements the response should support |
|
|
26
|
+
| `forbidden` | `HallucinationCheck` | Known-false claims that must not appear |
|
|
27
|
+
| `system_prompt` | `PromptInjectionCheck` | The system prompt that must not leak |
|
|
28
|
+
| `latency` | `LatencyCheck` | Seconds; set automatically by `Suite` when it calls the model |
|
|
29
|
+
| `trace` | `AgentFlowCheck` | List of agent steps (strings or dicts with `state`/`step`/`name`/`action`/`tool`) |
|
|
30
|
+
|
|
31
|
+
## Reports
|
|
32
|
+
|
|
33
|
+
`result.report("html", "report.html")` writes a self-contained, escaped HTML page.
|
|
34
|
+
`result.report("json", "report.json")` writes the full results for dashboards or CI artifacts.
|
|
35
|
+
|
|
36
|
+
## Compliance reports
|
|
37
|
+
|
|
38
|
+
```python
|
|
39
|
+
result.compliance_report(regulations=["dpdp", "gdpr"], format="html") # or "md", "json"
|
|
40
|
+
```
|
|
41
|
+
|
|
42
|
+
Each clause in [`privacyprobe/regulations.py`](https://github.com/Rancidcake/privacyprobe/blob/main/privacyprobe/regulations.py)
|
|
43
|
+
gets a status:
|
|
44
|
+
|
|
45
|
+
- **FAIL:** at least one mapped result failed, including a check that crashed
|
|
46
|
+
- **PASS:** every mapped result passed
|
|
47
|
+
- **NOT TESTED:** nothing in this run produced evidence for the clause
|
|
48
|
+
|
|
49
|
+
A check's clauses come from `CHECK_CLAUSES` (built-in checks) or its `clauses` attribute
|
|
50
|
+
(custom checks). `PIILeakCheck(profile="gdpr")` only claims GDPR clauses, because it only
|
|
51
|
+
looked for EU identifiers. Evidence is always redacted. The report is testing evidence,
|
|
52
|
+
not legal advice.
|
|
53
|
+
|
|
54
|
+
## Redaction
|
|
55
|
+
|
|
56
|
+
`redact(text, style="label" | "mask" | "hash", profile=..., types=..., salt=...)`.
|
|
57
|
+
Use `hash` with a secret `salt` for pseudonymised logs that can still be joined; unsalted
|
|
58
|
+
hashes of short identifiers like phone numbers can be brute-forced.
|