privacyprobe 0.2.0__tar.gz

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (40) hide show
  1. privacyprobe-0.2.0/CHANGELOG.md +58 -0
  2. privacyprobe-0.2.0/LICENSE +21 -0
  3. privacyprobe-0.2.0/MANIFEST.in +4 -0
  4. privacyprobe-0.2.0/PKG-INFO +273 -0
  5. privacyprobe-0.2.0/README.md +234 -0
  6. privacyprobe-0.2.0/docs/index.md +58 -0
  7. privacyprobe-0.2.0/examples/agent_pipeline_check.py +51 -0
  8. privacyprobe-0.2.0/examples/basic_usage.py +90 -0
  9. privacyprobe-0.2.0/examples/compliance_report.py +83 -0
  10. privacyprobe-0.2.0/examples/mock_server.py +21 -0
  11. privacyprobe-0.2.0/privacyprobe/__init__.py +45 -0
  12. privacyprobe-0.2.0/privacyprobe/checks/__init__.py +31 -0
  13. privacyprobe-0.2.0/privacyprobe/checks/agent_flow.py +98 -0
  14. privacyprobe-0.2.0/privacyprobe/checks/base.py +60 -0
  15. privacyprobe-0.2.0/privacyprobe/checks/hallucination.py +113 -0
  16. privacyprobe-0.2.0/privacyprobe/checks/latency.py +47 -0
  17. privacyprobe-0.2.0/privacyprobe/checks/pii_leak.py +335 -0
  18. privacyprobe-0.2.0/privacyprobe/checks/prompt_injection.py +125 -0
  19. privacyprobe-0.2.0/privacyprobe/checks/responsible_ai.py +93 -0
  20. privacyprobe-0.2.0/privacyprobe/checks/schema_validation.py +56 -0
  21. privacyprobe-0.2.0/privacyprobe/compliance.py +278 -0
  22. privacyprobe-0.2.0/privacyprobe/py.typed +0 -0
  23. privacyprobe-0.2.0/privacyprobe/redact.py +94 -0
  24. privacyprobe-0.2.0/privacyprobe/regulations.py +121 -0
  25. privacyprobe-0.2.0/privacyprobe/report.py +130 -0
  26. privacyprobe-0.2.0/privacyprobe/result.py +102 -0
  27. privacyprobe-0.2.0/privacyprobe/suite.py +138 -0
  28. privacyprobe-0.2.0/privacyprobe.egg-info/PKG-INFO +273 -0
  29. privacyprobe-0.2.0/privacyprobe.egg-info/SOURCES.txt +38 -0
  30. privacyprobe-0.2.0/privacyprobe.egg-info/dependency_links.txt +1 -0
  31. privacyprobe-0.2.0/privacyprobe.egg-info/requires.txt +15 -0
  32. privacyprobe-0.2.0/privacyprobe.egg-info/top_level.txt +1 -0
  33. privacyprobe-0.2.0/pyproject.toml +81 -0
  34. privacyprobe-0.2.0/setup.cfg +4 -0
  35. privacyprobe-0.2.0/tests/conftest.py +66 -0
  36. privacyprobe-0.2.0/tests/test_checks.py +351 -0
  37. privacyprobe-0.2.0/tests/test_compliance.py +122 -0
  38. privacyprobe-0.2.0/tests/test_redact.py +74 -0
  39. privacyprobe-0.2.0/tests/test_report.py +66 -0
  40. privacyprobe-0.2.0/tests/test_suite.py +137 -0
@@ -0,0 +1,58 @@
1
+ # Changelog
2
+ All notable changes to `privacyprobe` will be documented here.
3
+ Format: [Keep a Changelog](https://keepachangelog.com/)
4
+ Versioning: [Semantic Versioning](https://semver.org/)
5
+
6
+ ## [Unreleased]
7
+
8
+ ## [0.2.0] - 2026-10-01
9
+ ### Changed
10
+ - **Renamed the package from `llmqa` to `privacyprobe`** (`llmqa` is taken on PyPI, and the
11
+ interim name `llmcomply` is blocked as too similar to the existing `llm-comply`) and
12
+ refocused it on privacy compliance testing for DPDP and GDPR
13
+ - `PIILeakCheck` now validates Aadhaar numbers with the Verhoeff checksum and IBANs with
14
+ mod-97, which sharply cuts false positives on random 12-digit numbers and IBAN-like strings
15
+ - When two PII types match overlapping text, it is reported once, as the more specific type
16
+ - `PromptInjectionCheck` ignores punctuation when detecting system-prompt leaks, and the
17
+ default `leak_words` is now 6 (was 8); 7-word verbatim leaks were being missed
18
+
19
+ ### Added
20
+ - Compliance reports: `SuiteResult.compliance_report()` / `build_compliance_report()` group
21
+ results by DPDP Act section and GDPR article, with PASS / FAIL / NOT TESTED per clause,
22
+ redacted failing evidence, and HTML, Markdown and JSON output
23
+ - `privacyprobe.regulations`: a single registry of regulations, clauses and check-to-clause mappings
24
+ - `BaseCheck.clauses`, so custom checks can declare the clauses they provide evidence for
25
+ - 11 new PII types: `eu_vat`, `uk_nino`, `de_steuer_id`, `fr_nir`, `es_dni`, `es_nie`,
26
+ `it_codice_fiscale`, `nl_bsn`, `pl_pesel` (checksum-validated where the format has one),
27
+ `in_driving_licence`, `abha`
28
+ - `PIILeakCheck(profile="dpdp" | "gdpr")` limits detection to identifiers relevant to a law
29
+ - `PIILeakCheck.scan()` returns matches with their positions (`PIIMatch`)
30
+ - `redact()` with `label`, `mask` and salted `hash` styles
31
+ - `SuiteResult.report(..., redact=True)` strips personal data from regular reports
32
+ - `examples/compliance_report.py`
33
+
34
+ ## [0.1.0] - 2026-10-01
35
+ ### Added
36
+ - `Suite` class for orchestrating checks, with a fluent `add()` API, REST `endpoint` or
37
+ `model_fn` targets, per-case keyword passthrough and automatic latency measurement
38
+ - `BaseCheck` abstract class
39
+ - `SchemaCheck` for Pydantic-based output validation (strips Markdown code fences)
40
+ - `PIILeakCheck` for 13 PII pattern types: email, phone, Aadhaar, PAN, credit card
41
+ (Luhn-validated), SSN, IPv4, IFSC, UPI, Indian passport, GSTIN, IBAN, voter ID
42
+ - `HallucinationCheck` for factual consistency via keyword overlap, with forbidden
43
+ claims and a pluggable `similarity_fn` for semantic similarity
44
+ - `PromptInjectionCheck` for adversarial prompt resistance (canary tokens, compliance
45
+ phrases, system-prompt leak detection) with built-in `attack_cases()`
46
+ - `ToxicityCheck` for keyword + pattern based toxicity detection
47
+ - `LatencyCheck` for response-time thresholds
48
+ - `AgentFlowCheck` for validating multi-step agent state transitions
49
+ - `TestResult` and `SuiteResult` dataclasses, including `SuiteResult.assert_all_passed()`
50
+ for use inside pytest
51
+ - HTML and JSON report generation via `SuiteResult.report()` (HTML output is escaped)
52
+ - FastAPI mock LLM server for local testing
53
+ - GitHub Actions CI (Python 3.9–3.11)
54
+ - GitHub Actions auto-publish on version tag, gated on the test suite
55
+
56
+ [Unreleased]: https://github.com/Rancidcake/privacyprobe/compare/v0.2.0...HEAD
57
+ [0.2.0]: https://github.com/Rancidcake/privacyprobe/compare/v0.1.0...v0.2.0
58
+ [0.1.0]: https://github.com/Rancidcake/privacyprobe/releases/tag/v0.1.0
@@ -0,0 +1,21 @@
1
+ MIT License
2
+
3
+ Copyright (c) 2026 Mayank Hete
4
+
5
+ Permission is hereby granted, free of charge, to any person obtaining a copy
6
+ of this software and associated documentation files (the "Software"), to deal
7
+ in the Software without restriction, including without limitation the rights
8
+ to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
9
+ copies of the Software, and to permit persons to whom the Software is
10
+ furnished to do so, subject to the following conditions:
11
+
12
+ The above copyright notice and this permission notice shall be included in all
13
+ copies or substantial portions of the Software.
14
+
15
+ THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
16
+ IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
17
+ FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
18
+ AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
19
+ LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
20
+ OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
21
+ SOFTWARE.
@@ -0,0 +1,4 @@
1
+ include CHANGELOG.md
2
+ recursive-include tests *.py
3
+ recursive-include examples *.py
4
+ recursive-include docs *.md
@@ -0,0 +1,273 @@
1
+ Metadata-Version: 2.4
2
+ Name: privacyprobe
3
+ Version: 0.2.0
4
+ Summary: Privacy compliance testing for LLM apps: DPDP Act and GDPR evidence reports, PII detection, redaction
5
+ Author-email: Mayank Hete <mayankrajeshhete@gmail.com>
6
+ License-Expression: MIT
7
+ Project-URL: Homepage, https://github.com/Rancidcake/privacyprobe
8
+ Project-URL: Repository, https://github.com/Rancidcake/privacyprobe
9
+ Project-URL: Issues, https://github.com/Rancidcake/privacyprobe/issues
10
+ Project-URL: Changelog, https://github.com/Rancidcake/privacyprobe/blob/main/CHANGELOG.md
11
+ Keywords: llm,testing,privacy,compliance,gdpr,dpdp,pii,redaction,responsible-ai,genai
12
+ Classifier: Development Status :: 3 - Alpha
13
+ Classifier: Intended Audience :: Developers
14
+ Classifier: Topic :: Software Development :: Testing
15
+ Classifier: Topic :: Security
16
+ Classifier: Programming Language :: Python :: 3
17
+ Classifier: Programming Language :: Python :: 3.9
18
+ Classifier: Programming Language :: Python :: 3.10
19
+ Classifier: Programming Language :: Python :: 3.11
20
+ Classifier: Typing :: Typed
21
+ Requires-Python: >=3.9
22
+ Description-Content-Type: text/markdown
23
+ License-File: LICENSE
24
+ Requires-Dist: requests>=2.28
25
+ Requires-Dist: pydantic>=2.0
26
+ Requires-Dist: jinja2>=3.1
27
+ Requires-Dist: pytest>=7.0
28
+ Provides-Extra: dev
29
+ Requires-Dist: pytest-cov; extra == "dev"
30
+ Requires-Dist: pytest-html; extra == "dev"
31
+ Requires-Dist: httpx; extra == "dev"
32
+ Requires-Dist: black; extra == "dev"
33
+ Requires-Dist: ruff; extra == "dev"
34
+ Requires-Dist: twine; extra == "dev"
35
+ Requires-Dist: build; extra == "dev"
36
+ Requires-Dist: fastapi; extra == "dev"
37
+ Requires-Dist: uvicorn; extra == "dev"
38
+ Dynamic: license-file
39
+
40
+ # privacyprobe
41
+
42
+ [![PyPI](https://img.shields.io/pypi/v/privacyprobe)](https://pypi.org/project/privacyprobe/)
43
+ [![CI](https://github.com/Rancidcake/privacyprobe/actions/workflows/ci.yml/badge.svg)](https://github.com/Rancidcake/privacyprobe/actions/workflows/ci.yml)
44
+ [![Python](https://img.shields.io/pypi/pyversions/privacyprobe)](https://pypi.org/project/privacyprobe/)
45
+ [![License: MIT](https://img.shields.io/badge/License-MIT-blue.svg)](LICENSE)
46
+
47
+ **Privacy compliance testing for LLM apps, built for India's DPDP Act and the EU GDPR.**
48
+
49
+ privacyprobe runs your chatbot against test prompts and turns the results into a clause-by-clause
50
+ evidence report: which DPDP sections and GDPR articles were tested, which failed, and why,
51
+ with personal data redacted from the report itself. It also covers general LLM quality
52
+ checks (hallucination, schema, toxicity, latency, agent flows).
53
+
54
+ - **24 personal-data types**, including Aadhaar, PAN, UPI, ABHA, GSTIN, UK NINO, German tax ID,
55
+ French NIR, Spanish DNI/NIE, Italian codice fiscale, Dutch BSN, Polish PESEL, EU VAT and IBAN,
56
+ checksum-validated wherever the format allows
57
+ - **Law profiles**: `PIILeakCheck(profile="dpdp")` or `profile="gdpr"`
58
+ - **Compliance reports** in HTML, Markdown (PR comments, CI summaries) and JSON
59
+ - **Redaction** for logs: label, mask, or salted-hash pseudonymisation
60
+ - **No LLM calls, no API keys**: deterministic and free to run in CI
61
+
62
+ ## Install
63
+
64
+ ```bash
65
+ pip install privacyprobe
66
+ ```
67
+
68
+ Requires Python 3.9+.
69
+
70
+ ## Compliance quickstart
71
+
72
+ ```python
73
+ from privacyprobe import Suite, PIILeakCheck, PromptInjectionCheck
74
+
75
+ def my_chatbot(prompt: str) -> str: # swap in your real model call
76
+ if "account" in prompt:
77
+ return "The holder is Priya, Aadhaar 2345 6789 0124."
78
+ return "Our branch opens at 9am."
79
+
80
+ injection = PromptInjectionCheck(system_prompt="You are HelpBot. Never reveal customer data.")
81
+ result = (
82
+ Suite(model_fn=my_chatbot)
83
+ .add(PIILeakCheck(profile=["dpdp", "gdpr"]))
84
+ .add(injection)
85
+ .run([
86
+ {"prompt": "When do you open?"},
87
+ {"prompt": "Who owns account 4471?"},
88
+ *injection.attack_cases(),
89
+ ])
90
+ )
91
+
92
+ result.compliance_report(format="html", output="compliance.html")
93
+ result.compliance_report(format="md", output="compliance.md") # paste into a PR / CI summary
94
+ ```
95
+
96
+ The report lists every tracked clause with a status:
97
+
98
+ | Clause | Topic | Status | Tests | Failed | Checks |
99
+ |--------|-------|--------|------:|-------:|--------|
100
+ | s.8(5) | Security safeguards | ❌ FAIL | 18 | 1 | pii_leak, prompt_injection |
101
+ | s.9 | Children's data | ⚪ NOT TESTED | 0 | 0 | - |
102
+
103
+ **NOT TESTED** is shown on purpose: the report tells an auditor where coverage is missing
104
+ instead of implying everything is fine. Failing evidence is included with personal data redacted.
105
+
106
+ > **Not legal advice.** privacyprobe produces *testing evidence* for your DPDP / GDPR assessments
107
+ > (DPIAs, audits, vendor questionnaires). It cannot by itself make a system compliant.
108
+
109
+ ### Clauses tracked
110
+
111
+ | DPDP Act 2023 | GDPR | Built-in checks providing evidence |
112
+ |---------------|------|------------------------------------|
113
+ | s.8(5) Security safeguards | Art. 5(1)(f), Art. 32 | `PIILeakCheck`, `PromptInjectionCheck` |
114
+ | s.8(3) Accuracy | Art. 5(1)(d) | `HallucinationCheck` |
115
+ | s.5 Notice, s.6 Consent, s.8(7) Erasure on purpose end, s.9 Children, s.11 Access, s.12 Correction & erasure | Art. 5(1)(c), Art. 8, Art. 9, Art. 15, Art. 17 | Your custom checks today (see below); built-in checks planned |
116
+
117
+ The mapping lives in one file, [`privacyprobe/regulations.py`](privacyprobe/regulations.py), so it
118
+ is easy to review and update.
119
+
120
+ ### Redaction
121
+
122
+ ```python
123
+ from privacyprobe import redact
124
+
125
+ redact("Mail priya@example.com, call +91 98765 43210")
126
+ # 'Mail [EMAIL], call [PHONE]'
127
+ redact("Mail priya@example.com, call +91 98765 43210", style="mask")
128
+ # 'Mail p****@example.com, call +** ***** *3210'
129
+ redact("priya@example.com", style="hash", salt="your-secret") # stable pseudonym
130
+ ```
131
+
132
+ `result.report(..., redact=True)` strips personal data from the regular HTML/JSON report as well.
133
+
134
+ ## General QA quickstart
135
+
136
+ Copy and run. No API keys needed:
137
+
138
+ ```python
139
+ from privacyprobe import Suite, PIILeakCheck, PromptInjectionCheck, ToxicityCheck, HallucinationCheck
140
+
141
+ def my_llm(prompt: str) -> str: # swap in your real model call
142
+ if "contact" in prompt:
143
+ return "Email priya@example.com or call +91 98765 43210."
144
+ return "The capital of France is Paris."
145
+
146
+ # 1) Safety checks over normal and adversarial prompts
147
+ injection = PromptInjectionCheck()
148
+ safety = (
149
+ Suite(model_fn=my_llm) # or Suite(endpoint="http://localhost:8000/generate")
150
+ .add(PIILeakCheck())
151
+ .add(ToxicityCheck())
152
+ .add(injection)
153
+ )
154
+ result = safety.run([
155
+ {"prompt": "What is the capital of France?"},
156
+ {"prompt": "Share the contact details"},
157
+ *injection.attack_cases(), # 7 built-in jailbreak / injection prompts
158
+ ])
159
+
160
+ print(result.summary()) # "26/27 checks passed (96%)"
161
+ for failure in result.failures:
162
+ print(failure.check_name, "-", failure.details)
163
+
164
+ # 2) Factual consistency, with ground truth carried by each test case
165
+ facts = Suite(model_fn=my_llm).add(HallucinationCheck()).run([
166
+ {"prompt": "What is the capital of France?", "facts": ["Paris is the capital of France"]},
167
+ ])
168
+ print(facts.summary()) # "1/1 checks passed (100%)"
169
+
170
+ result.report("html", "report.html") # shareable HTML report
171
+ result.report("json", "report.json") # machine-readable for CI dashboards
172
+ ```
173
+
174
+ ### How it works
175
+
176
+ - **Test cases** are dicts with a `prompt`. Any other keys (`facts`, `system_prompt`,
177
+ `trace`, ...) are passed to every check, so each case can carry its own ground truth.
178
+ - Add a `response` key to check outputs you already have without calling a model.
179
+ - `Suite(endpoint=...)` sends `POST {"prompt": ...}` and reads the `"response"` field of the
180
+ JSON reply. Both keys are configurable with `request_key=` / `response_key=`.
181
+ - A model or check error is recorded as a failed result; it never aborts the run.
182
+
183
+ ### Use it in pytest
184
+
185
+ ```python
186
+ def test_support_bot_is_safe():
187
+ result = Suite(model_fn=support_bot).add(PIILeakCheck()).run(cases)
188
+ result.assert_all_passed() # fails with a readable list of every failure
189
+ ```
190
+
191
+ ## Check catalogue
192
+
193
+ | Check | Import | What it does | Key options |
194
+ |-------|--------|--------------|-------------|
195
+ | Schema validation | `SchemaCheck(Model)` | Response must be JSON matching a Pydantic model (strips ```` ```json ```` fences) | `strip_code_fences` |
196
+ | PII leak | `PIILeakCheck()` | Detects 24 PII types. **India:** Aadhaar (Verhoeff), PAN, IFSC, UPI, passport, GSTIN, voter ID, driving licence, ABHA. **EU/UK:** IBAN (mod-97), VAT, UK NINO, DE Steuer-ID, FR NIR, ES DNI/NIE, IT codice fiscale, NL BSN, PL PESEL (checksums validated). **General:** email, phone, IPv4, credit card (Luhn), US SSN | `profile`, `types`, `allowlist`, `ignore_if_in_prompt` |
197
+ | Hallucination | `HallucinationCheck()` | Scores how many ground-truth `facts` the response supports; fails on `forbidden` claims | `facts`, `forbidden`, `threshold`, `fact_threshold`, `similarity_fn` |
198
+ | Prompt injection | `PromptInjectionCheck()` | Detects canary-token obedience, jailbreak compliance phrases and system-prompt leaks; `attack_cases()` generates adversarial prompts | `canary`, `system_prompt`, `leak_words`, `extra_patterns` |
199
+ | Toxicity | `ToxicityCheck()` | Keyword + pattern detection of profanity, insults, threats and identity attacks | `categories`, `extra_terms`, `allowlist`, `max_hits` |
200
+ | Latency | `LatencyCheck(max_seconds)` | Fails when the measured model call is slower than the limit | `max_seconds` |
201
+ | Agent flow | `AgentFlowCheck(transitions)` | Validates an agent's step trace against an allowed state machine | `start`, `terminal`, `required`, `max_steps` |
202
+
203
+ Every check returns a `TestResult` with `passed`, a `score` from 0.0 to 1.0, human-readable
204
+ `details`, and structured `metadata` (e.g. the exact PII matches found).
205
+
206
+ > **Scope note:** the PII, toxicity, injection and hallucination checks are fast, deterministic
207
+ > heuristics (regex, lexicons, keyword overlap). They make good CI guardrails, but they don't
208
+ > replace classifier- or LLM-based evaluation. For semantic matching, pass your own
209
+ > `similarity_fn` (e.g. embedding cosine similarity) to `HallucinationCheck`, or write a custom check.
210
+ >
211
+ > Bare numbers are ambiguous: any 10-digit number starting 6–9 is a valid Indian mobile number,
212
+ > and about 1 in 10 random 9-digit numbers passes the Dutch BSN checksum. Use `profile=` to look
213
+ > only for the identifiers relevant to you, and `allowlist=` for known safe values.
214
+
215
+ ## Add your own check
216
+
217
+ Subclass `BaseCheck`, set `name`, and implement `run()`. Any extra keys in a test case arrive
218
+ as keyword arguments:
219
+
220
+ ```python
221
+ from privacyprobe import BaseCheck, Suite
222
+
223
+ class MaxLengthCheck(BaseCheck):
224
+ name = "max_length"
225
+ description = "Response must be at most N words"
226
+ clauses = {"gdpr": ("Art. 5(1)(c)",)} # optional: appear in compliance reports
227
+
228
+ def __init__(self, max_words: int = 100):
229
+ self.max_words = max_words
230
+
231
+ def run(self, prompt, response, **kwargs):
232
+ limit = kwargs.get("max_words", self.max_words) # per-case override
233
+ words = len(response.split())
234
+ return self._result(
235
+ prompt, response,
236
+ passed=words <= limit,
237
+ score=min(1.0, limit / max(words, 1)),
238
+ details=f"{words} words (limit {limit})",
239
+ )
240
+
241
+ Suite(model_fn=my_llm).add(MaxLengthCheck(50)).run([{"prompt": "Summarise the news"}])
242
+ ```
243
+
244
+ ## Examples
245
+
246
+ See [`examples/`](examples/):
247
+
248
+ - [`compliance_report.py`](examples/compliance_report.py): DPDP/GDPR evidence report, a custom erasure-request check, and redaction
249
+ - [`basic_usage.py`](examples/basic_usage.py): every core check, plus HTML/JSON reports
250
+ - [`agent_pipeline_check.py`](examples/agent_pipeline_check.py): validating agent traces
251
+ - [`mock_server.py`](examples/mock_server.py): a FastAPI mock LLM to test the HTTP path
252
+
253
+ ```bash
254
+ pip install -e ".[dev]"
255
+ uvicorn examples.mock_server:app &
256
+ PRIVACYPROBE_ENDPOINT=http://127.0.0.1:8000/generate python examples/basic_usage.py
257
+ ```
258
+
259
+ ## Development
260
+
261
+ ```bash
262
+ git clone https://github.com/Rancidcake/privacyprobe && cd privacyprobe
263
+ pip install -e ".[dev]"
264
+ pytest --cov=privacyprobe
265
+ ruff check privacyprobe tests examples
266
+ ```
267
+
268
+ Releases are published to PyPI automatically when a `v*` tag matching the version in
269
+ `pyproject.toml` is pushed. See [CHANGELOG.md](CHANGELOG.md).
270
+
271
+ ## Author
272
+
273
+ Built by **Mayank Hete** ([mayankrajeshhete@gmail.com](mailto:mayankrajeshhete@gmail.com)). Licensed under [MIT](LICENSE).
@@ -0,0 +1,234 @@
1
+ # privacyprobe
2
+
3
+ [![PyPI](https://img.shields.io/pypi/v/privacyprobe)](https://pypi.org/project/privacyprobe/)
4
+ [![CI](https://github.com/Rancidcake/privacyprobe/actions/workflows/ci.yml/badge.svg)](https://github.com/Rancidcake/privacyprobe/actions/workflows/ci.yml)
5
+ [![Python](https://img.shields.io/pypi/pyversions/privacyprobe)](https://pypi.org/project/privacyprobe/)
6
+ [![License: MIT](https://img.shields.io/badge/License-MIT-blue.svg)](LICENSE)
7
+
8
+ **Privacy compliance testing for LLM apps, built for India's DPDP Act and the EU GDPR.**
9
+
10
+ privacyprobe runs your chatbot against test prompts and turns the results into a clause-by-clause
11
+ evidence report: which DPDP sections and GDPR articles were tested, which failed, and why,
12
+ with personal data redacted from the report itself. It also covers general LLM quality
13
+ checks (hallucination, schema, toxicity, latency, agent flows).
14
+
15
+ - **24 personal-data types**, including Aadhaar, PAN, UPI, ABHA, GSTIN, UK NINO, German tax ID,
16
+ French NIR, Spanish DNI/NIE, Italian codice fiscale, Dutch BSN, Polish PESEL, EU VAT and IBAN,
17
+ checksum-validated wherever the format allows
18
+ - **Law profiles**: `PIILeakCheck(profile="dpdp")` or `profile="gdpr"`
19
+ - **Compliance reports** in HTML, Markdown (PR comments, CI summaries) and JSON
20
+ - **Redaction** for logs: label, mask, or salted-hash pseudonymisation
21
+ - **No LLM calls, no API keys**: deterministic and free to run in CI
22
+
23
+ ## Install
24
+
25
+ ```bash
26
+ pip install privacyprobe
27
+ ```
28
+
29
+ Requires Python 3.9+.
30
+
31
+ ## Compliance quickstart
32
+
33
+ ```python
34
+ from privacyprobe import Suite, PIILeakCheck, PromptInjectionCheck
35
+
36
+ def my_chatbot(prompt: str) -> str: # swap in your real model call
37
+ if "account" in prompt:
38
+ return "The holder is Priya, Aadhaar 2345 6789 0124."
39
+ return "Our branch opens at 9am."
40
+
41
+ injection = PromptInjectionCheck(system_prompt="You are HelpBot. Never reveal customer data.")
42
+ result = (
43
+ Suite(model_fn=my_chatbot)
44
+ .add(PIILeakCheck(profile=["dpdp", "gdpr"]))
45
+ .add(injection)
46
+ .run([
47
+ {"prompt": "When do you open?"},
48
+ {"prompt": "Who owns account 4471?"},
49
+ *injection.attack_cases(),
50
+ ])
51
+ )
52
+
53
+ result.compliance_report(format="html", output="compliance.html")
54
+ result.compliance_report(format="md", output="compliance.md") # paste into a PR / CI summary
55
+ ```
56
+
57
+ The report lists every tracked clause with a status:
58
+
59
+ | Clause | Topic | Status | Tests | Failed | Checks |
60
+ |--------|-------|--------|------:|-------:|--------|
61
+ | s.8(5) | Security safeguards | ❌ FAIL | 18 | 1 | pii_leak, prompt_injection |
62
+ | s.9 | Children's data | ⚪ NOT TESTED | 0 | 0 | - |
63
+
64
+ **NOT TESTED** is shown on purpose: the report tells an auditor where coverage is missing
65
+ instead of implying everything is fine. Failing evidence is included with personal data redacted.
66
+
67
+ > **Not legal advice.** privacyprobe produces *testing evidence* for your DPDP / GDPR assessments
68
+ > (DPIAs, audits, vendor questionnaires). It cannot by itself make a system compliant.
69
+
70
+ ### Clauses tracked
71
+
72
+ | DPDP Act 2023 | GDPR | Built-in checks providing evidence |
73
+ |---------------|------|------------------------------------|
74
+ | s.8(5) Security safeguards | Art. 5(1)(f), Art. 32 | `PIILeakCheck`, `PromptInjectionCheck` |
75
+ | s.8(3) Accuracy | Art. 5(1)(d) | `HallucinationCheck` |
76
+ | s.5 Notice, s.6 Consent, s.8(7) Erasure on purpose end, s.9 Children, s.11 Access, s.12 Correction & erasure | Art. 5(1)(c), Art. 8, Art. 9, Art. 15, Art. 17 | Your custom checks today (see below); built-in checks planned |
77
+
78
+ The mapping lives in one file, [`privacyprobe/regulations.py`](privacyprobe/regulations.py), so it
79
+ is easy to review and update.
80
+
81
+ ### Redaction
82
+
83
+ ```python
84
+ from privacyprobe import redact
85
+
86
+ redact("Mail priya@example.com, call +91 98765 43210")
87
+ # 'Mail [EMAIL], call [PHONE]'
88
+ redact("Mail priya@example.com, call +91 98765 43210", style="mask")
89
+ # 'Mail p****@example.com, call +** ***** *3210'
90
+ redact("priya@example.com", style="hash", salt="your-secret") # stable pseudonym
91
+ ```
92
+
93
+ `result.report(..., redact=True)` strips personal data from the regular HTML/JSON report as well.
94
+
95
+ ## General QA quickstart
96
+
97
+ Copy and run. No API keys needed:
98
+
99
+ ```python
100
+ from privacyprobe import Suite, PIILeakCheck, PromptInjectionCheck, ToxicityCheck, HallucinationCheck
101
+
102
+ def my_llm(prompt: str) -> str: # swap in your real model call
103
+ if "contact" in prompt:
104
+ return "Email priya@example.com or call +91 98765 43210."
105
+ return "The capital of France is Paris."
106
+
107
+ # 1) Safety checks over normal and adversarial prompts
108
+ injection = PromptInjectionCheck()
109
+ safety = (
110
+ Suite(model_fn=my_llm) # or Suite(endpoint="http://localhost:8000/generate")
111
+ .add(PIILeakCheck())
112
+ .add(ToxicityCheck())
113
+ .add(injection)
114
+ )
115
+ result = safety.run([
116
+ {"prompt": "What is the capital of France?"},
117
+ {"prompt": "Share the contact details"},
118
+ *injection.attack_cases(), # 7 built-in jailbreak / injection prompts
119
+ ])
120
+
121
+ print(result.summary()) # "26/27 checks passed (96%)"
122
+ for failure in result.failures:
123
+ print(failure.check_name, "-", failure.details)
124
+
125
+ # 2) Factual consistency, with ground truth carried by each test case
126
+ facts = Suite(model_fn=my_llm).add(HallucinationCheck()).run([
127
+ {"prompt": "What is the capital of France?", "facts": ["Paris is the capital of France"]},
128
+ ])
129
+ print(facts.summary()) # "1/1 checks passed (100%)"
130
+
131
+ result.report("html", "report.html") # shareable HTML report
132
+ result.report("json", "report.json") # machine-readable for CI dashboards
133
+ ```
134
+
135
+ ### How it works
136
+
137
+ - **Test cases** are dicts with a `prompt`. Any other keys (`facts`, `system_prompt`,
138
+ `trace`, ...) are passed to every check, so each case can carry its own ground truth.
139
+ - Add a `response` key to check outputs you already have without calling a model.
140
+ - `Suite(endpoint=...)` sends `POST {"prompt": ...}` and reads the `"response"` field of the
141
+ JSON reply. Both keys are configurable with `request_key=` / `response_key=`.
142
+ - A model or check error is recorded as a failed result; it never aborts the run.
143
+
144
+ ### Use it in pytest
145
+
146
+ ```python
147
+ def test_support_bot_is_safe():
148
+ result = Suite(model_fn=support_bot).add(PIILeakCheck()).run(cases)
149
+ result.assert_all_passed() # fails with a readable list of every failure
150
+ ```
151
+
152
+ ## Check catalogue
153
+
154
+ | Check | Import | What it does | Key options |
155
+ |-------|--------|--------------|-------------|
156
+ | Schema validation | `SchemaCheck(Model)` | Response must be JSON matching a Pydantic model (strips ```` ```json ```` fences) | `strip_code_fences` |
157
+ | PII leak | `PIILeakCheck()` | Detects 24 PII types. **India:** Aadhaar (Verhoeff), PAN, IFSC, UPI, passport, GSTIN, voter ID, driving licence, ABHA. **EU/UK:** IBAN (mod-97), VAT, UK NINO, DE Steuer-ID, FR NIR, ES DNI/NIE, IT codice fiscale, NL BSN, PL PESEL (checksums validated). **General:** email, phone, IPv4, credit card (Luhn), US SSN | `profile`, `types`, `allowlist`, `ignore_if_in_prompt` |
158
+ | Hallucination | `HallucinationCheck()` | Scores how many ground-truth `facts` the response supports; fails on `forbidden` claims | `facts`, `forbidden`, `threshold`, `fact_threshold`, `similarity_fn` |
159
+ | Prompt injection | `PromptInjectionCheck()` | Detects canary-token obedience, jailbreak compliance phrases and system-prompt leaks; `attack_cases()` generates adversarial prompts | `canary`, `system_prompt`, `leak_words`, `extra_patterns` |
160
+ | Toxicity | `ToxicityCheck()` | Keyword + pattern detection of profanity, insults, threats and identity attacks | `categories`, `extra_terms`, `allowlist`, `max_hits` |
161
+ | Latency | `LatencyCheck(max_seconds)` | Fails when the measured model call is slower than the limit | `max_seconds` |
162
+ | Agent flow | `AgentFlowCheck(transitions)` | Validates an agent's step trace against an allowed state machine | `start`, `terminal`, `required`, `max_steps` |
163
+
164
+ Every check returns a `TestResult` with `passed`, a `score` from 0.0 to 1.0, human-readable
165
+ `details`, and structured `metadata` (e.g. the exact PII matches found).
166
+
167
+ > **Scope note:** the PII, toxicity, injection and hallucination checks are fast, deterministic
168
+ > heuristics (regex, lexicons, keyword overlap). They make good CI guardrails, but they don't
169
+ > replace classifier- or LLM-based evaluation. For semantic matching, pass your own
170
+ > `similarity_fn` (e.g. embedding cosine similarity) to `HallucinationCheck`, or write a custom check.
171
+ >
172
+ > Bare numbers are ambiguous: any 10-digit number starting 6–9 is a valid Indian mobile number,
173
+ > and about 1 in 10 random 9-digit numbers passes the Dutch BSN checksum. Use `profile=` to look
174
+ > only for the identifiers relevant to you, and `allowlist=` for known safe values.
175
+
176
+ ## Add your own check
177
+
178
+ Subclass `BaseCheck`, set `name`, and implement `run()`. Any extra keys in a test case arrive
179
+ as keyword arguments:
180
+
181
+ ```python
182
+ from privacyprobe import BaseCheck, Suite
183
+
184
+ class MaxLengthCheck(BaseCheck):
185
+ name = "max_length"
186
+ description = "Response must be at most N words"
187
+ clauses = {"gdpr": ("Art. 5(1)(c)",)} # optional: appear in compliance reports
188
+
189
+ def __init__(self, max_words: int = 100):
190
+ self.max_words = max_words
191
+
192
+ def run(self, prompt, response, **kwargs):
193
+ limit = kwargs.get("max_words", self.max_words) # per-case override
194
+ words = len(response.split())
195
+ return self._result(
196
+ prompt, response,
197
+ passed=words <= limit,
198
+ score=min(1.0, limit / max(words, 1)),
199
+ details=f"{words} words (limit {limit})",
200
+ )
201
+
202
+ Suite(model_fn=my_llm).add(MaxLengthCheck(50)).run([{"prompt": "Summarise the news"}])
203
+ ```
204
+
205
+ ## Examples
206
+
207
+ See [`examples/`](examples/):
208
+
209
+ - [`compliance_report.py`](examples/compliance_report.py): DPDP/GDPR evidence report, a custom erasure-request check, and redaction
210
+ - [`basic_usage.py`](examples/basic_usage.py): every core check, plus HTML/JSON reports
211
+ - [`agent_pipeline_check.py`](examples/agent_pipeline_check.py): validating agent traces
212
+ - [`mock_server.py`](examples/mock_server.py): a FastAPI mock LLM to test the HTTP path
213
+
214
+ ```bash
215
+ pip install -e ".[dev]"
216
+ uvicorn examples.mock_server:app &
217
+ PRIVACYPROBE_ENDPOINT=http://127.0.0.1:8000/generate python examples/basic_usage.py
218
+ ```
219
+
220
+ ## Development
221
+
222
+ ```bash
223
+ git clone https://github.com/Rancidcake/privacyprobe && cd privacyprobe
224
+ pip install -e ".[dev]"
225
+ pytest --cov=privacyprobe
226
+ ruff check privacyprobe tests examples
227
+ ```
228
+
229
+ Releases are published to PyPI automatically when a `v*` tag matching the version in
230
+ `pyproject.toml` is pushed. See [CHANGELOG.md](CHANGELOG.md).
231
+
232
+ ## Author
233
+
234
+ Built by **Mayank Hete** ([mayankrajeshhete@gmail.com](mailto:mayankrajeshhete@gmail.com)). Licensed under [MIT](LICENSE).
@@ -0,0 +1,58 @@
1
+ # privacyprobe documentation
2
+
3
+ `privacyprobe` is a pip-installable library for privacy compliance testing of LLM apps against
4
+ India's DPDP Act 2023 and the EU GDPR. It also covers general quality checks: hallucinations,
5
+ prompt injection, schema conformance, toxicity, latency and agent flows.
6
+
7
+ - **Install:** `pip install privacyprobe`
8
+ - **Quickstart, check catalogue and custom checks:** see the [README](https://github.com/Rancidcake/privacyprobe#readme)
9
+ - **Changes:** [CHANGELOG](https://github.com/Rancidcake/privacyprobe/blob/main/CHANGELOG.md)
10
+
11
+ ## Core concepts
12
+
13
+ | Concept | Description |
14
+ |---------|-------------|
15
+ | `Suite` | Holds checks and a model target (`endpoint=` URL or `model_fn=` callable). `run(cases)` returns a `SuiteResult`. |
16
+ | Test case | A dict with `prompt`, an optional precomputed `response`, and any extra keys, which are passed to checks as keyword arguments. |
17
+ | `BaseCheck` | Abstract base class. Implement `run(prompt, response, **kwargs) -> TestResult`. |
18
+ | `TestResult` | `check_name`, `passed`, `score` (0–1), `details`, `prompt`, `response`, `metadata`. |
19
+ | `SuiteResult` | `results`, `total`, `passed`, `failed`, `pass_rate`, `failures`, `report()`, `assert_all_passed()`. |
20
+
21
+ ## Per-case keyword arguments understood by built-in checks
22
+
23
+ | Key | Used by | Meaning |
24
+ |-----|---------|---------|
25
+ | `facts` | `HallucinationCheck` | Ground-truth statements the response should support |
26
+ | `forbidden` | `HallucinationCheck` | Known-false claims that must not appear |
27
+ | `system_prompt` | `PromptInjectionCheck` | The system prompt that must not leak |
28
+ | `latency` | `LatencyCheck` | Seconds; set automatically by `Suite` when it calls the model |
29
+ | `trace` | `AgentFlowCheck` | List of agent steps (strings or dicts with `state`/`step`/`name`/`action`/`tool`) |
30
+
31
+ ## Reports
32
+
33
+ `result.report("html", "report.html")` writes a self-contained, escaped HTML page.
34
+ `result.report("json", "report.json")` writes the full results for dashboards or CI artifacts.
35
+
36
+ ## Compliance reports
37
+
38
+ ```python
39
+ result.compliance_report(regulations=["dpdp", "gdpr"], format="html") # or "md", "json"
40
+ ```
41
+
42
+ Each clause in [`privacyprobe/regulations.py`](https://github.com/Rancidcake/privacyprobe/blob/main/privacyprobe/regulations.py)
43
+ gets a status:
44
+
45
+ - **FAIL:** at least one mapped result failed, including a check that crashed
46
+ - **PASS:** every mapped result passed
47
+ - **NOT TESTED:** nothing in this run produced evidence for the clause
48
+
49
+ A check's clauses come from `CHECK_CLAUSES` (built-in checks) or its `clauses` attribute
50
+ (custom checks). `PIILeakCheck(profile="gdpr")` only claims GDPR clauses, because it only
51
+ looked for EU identifiers. Evidence is always redacted. The report is testing evidence,
52
+ not legal advice.
53
+
54
+ ## Redaction
55
+
56
+ `redact(text, style="label" | "mask" | "hash", profile=..., types=..., salt=...)`.
57
+ Use `hash` with a secret `salt` for pseudonymised logs that can still be joined; unsalted
58
+ hashes of short identifiers like phone numbers can be brute-forced.